How to find orphan pages on your website — and why a crawler can only say "potential"
Pages nothing links to are hard for visitors and search engines to find. How to find them, why a crawl can only flag "potential" orphans, and what to do with each one.
An orphan page is a page on your site that no other page links to. It may still exist, be listed in your sitemap and even rank for something, but visitors browsing your site can't reach it, and search engines have only the sitemap to tell them it exists. They also get no signal from your own linking that the page matters.
Orphans are rarely deliberate. They're what's left after a redesign drops a menu item, a campaign landing page outlives its campaign, a category is removed, or a blog post only ever got linked from a homepage slot that has since moved on.
Why orphan pages matter
- Discovery. Search engines find most pages by following links. A page reachable only through the sitemap is discovered and recrawled less reliably.
- Importance. Internal links are one of the ways search engines judge which pages you consider important. A page with none looks unimportant.
- Visitors. People can't click their way to it, so useful content goes unread.
- Maintenance. Forgotten pages drift out of date: old prices, dead offers, outdated advice.
How orphan pages are found
To call a page an orphan, you need two lists: every page that exists, and every page that something links to. The orphans are on the first list and missing from the second. The hard part is the first list, because by definition you can't find orphans by following links. The usual sources are:
- your XML sitemap;
- a content export from your CMS;
- pages that get visits in your analytics;
- pages Google knows about in Search Console's Performance and Page indexing reports;
- URLs that appear in your server logs.
The second list comes from a crawl: start at the homepage, follow every internal link, and record which pages were linked to.
Why a crawler says "potential orphan"
A crawler only knows about links on the pages it actually visited. If a crawl stops at 10 or 100 pages, or at a certain click depth, or skips pages blocked by robots.txt, then a sitemap page that looks unlinked might be linked from a page the crawl never reached. That's why an honest tool calls these potential orphans: "no internal link to this page was found in this crawl", not "nothing links to this page".
Other things that can hide a real link from a crawler: links that only appear after JavaScript runs (when the crawler reads raw HTML), links inside forms or search results, and links from subdomains or other hosts, which don't count as internal.
Look for potential orphan pages
Run an audit: a Rudra site scan compares your sitemap with the internal links it found and lists pages nothing in the crawl linked to.
How Rudra finds potential orphans
A Rudra site scan (free account; the free website audit covers a single page) reads your XML sitemap during the crawl. After the crawl, it lists sitemap pages that returned 200 and received no internal link from any other crawled page. The homepage is excluded, only links between pages of the same site count, and URLs are compared after normalising trailing slashes, fragments and http/https or www variants. The check only runs when a sitemap was found and at least two pages were crawled, and the finding states how many pages the crawl covered, so you can judge how complete it is.
The same report lists pages with few internal links — one or none from other crawled pages — as a separate, low-confidence note, and pages four or more clicks from the homepage. Neither is an error; they're prompts to review your linking.
Confirming an orphan
- Search your own site. Use your CMS's search or a database query for the page's URL and slug, and a
site:yourdomain.com "page title"search for mentions. - Check Search Console. Links → Internal links shows how many internal links Google has seen to a URL.
- Check where the crawl stopped. If your crawl hit its page limit, a larger crawl may find the link.
- Look at analytics. A page with visits but no internal referrers is being reached from search, email or ads only.
What to do with each orphan
Decide page by page whether it still deserves to exist:
- Useful and current: link to it from relevant places — a category or hub page, related articles, the navigation or footer if it's important, and breadcrumbs. Use descriptive link text rather than "click here".
- Useful but outdated: update it, then link to it.
- Replaced by a better page: 301-redirect it to the replacement and remove it from the sitemap.
- Obsolete with no replacement: remove it (a 404 or 410 is fine), and remove it from the sitemap.
- Deliberately unlinked (a campaign landing page, a thank-you page): consider
noindexand removing it from the sitemap, so it doesn't compete with your main pages.
While you're there, check the links you add actually work — our guide on how to fix broken links covers that — and keep the sitemap in step, as described in how to diagnose sitemap problems.
Frequently asked questions
Are orphan pages bad for SEO?
They're not penalised, but they're harder for search engines to discover and look less important because nothing on your site links to them. Important pages should always be linked from other relevant pages.
Can a page be an orphan if it's in my sitemap?
Yes. The sitemap tells search engines the page exists, but it isn't a link. An orphan is a page with no internal links pointing to it, whether or not the sitemap lists it.
Why did the scan flag a page I know is linked?
The link is probably on a page the crawl didn't reach because of its page or depth limit, added by JavaScript, or on another host. That's why the finding says "potential" orphan; confirm before acting.
Should I just add every orphan to the navigation?
No. Link each page from the places where it's genuinely relevant. Navigation is for your most important pages; everything else belongs in hubs, categories and related content.