Search Console lists dozens of URLs as discovered or crawled but not indexed. Robots.txt permits them, canonicals are correct, the sitemap is clean and the pages return a healthy status. Everything technical checks out, which is exactly why this is frustrating: the barrier is not technical. Google crawled the page, evaluated it, and decided it was not worth storing.
01Read the status precisely
Discovered but not crawled means the URL is known and has not been fetched, which usually points at crawl budget or weak internal linking. Crawled but not indexed means it was fetched and rejected, which is a quality assessment. These require entirely different responses and are frequently conflated.
Verify the technical layer once, properly, then stop returning to it. Confirm the page renders its content without JavaScript execution, that the canonical points to itself, that it returns a 200, and that it is reachable from your navigation. If all of that holds, the remaining problem is the content.
| Status | Meaning | Address by |
|---|---|---|
| Discovered - not crawled | Known, never fetched | Internal links, sitemap priority |
| Crawled - not indexed | Fetched and rejected | Content depth and uniqueness |
| Duplicate, no canonical | Seen as a copy of another page | Consolidate or differentiate |
| Alternate with canonical | Correctly pointing elsewhere | Usually expected |
| Soft 404 | Looks empty or error-like | Add substantive content |
02Near-duplicate templates are the usual cause
Pages generated from a template with a variable swapped - the same service description repeated across twenty locations, or near-identical product pages differing by one attribute - are recognised as substantially the same document. Google indexes one and discards the rest.
The fix is genuine differentiation rather than synonym substitution. Each page needs material that exists nowhere else on the site: specific examples, distinct data, real detail relevant to that variant. If two pages could be swapped without a reader noticing, only one will be kept.
03Orphaned pages signal unimportance
A page reachable only from the sitemap, with no internal links pointing at it, is implicitly declared unimportant by your own site structure. Sitemaps aid discovery; internal links communicate importance, and the second is the stronger signal.
Link new content from established pages that are already indexed and receiving traffic. Related-article sections, category pages and contextual links within body copy all pass signal, and a page with several internal links from indexed pages is indexed far more reliably than one with none.
| Check | How |
|---|---|
| Renders without JavaScript | View source, look for the body text |
| Internal links point to it | Search your own site for the URL |
| Content is genuinely unique | Compare against your closest page |
| Substantive depth | Does it answer the query completely? |
| Loads quickly on mobile | Field data, not lab scores |
04New sites are evaluated more strictly
A domain with little history and few external links gets less benefit of the doubt. The same page on an established site would be indexed immediately; on a new domain it sits in the queue while the site accumulates evidence of being worth crawling deeply.
This resolves with sustained publishing rather than with technical intervention. Consistent output of genuinely useful pages, internally linked, shifts the assessment over weeks to months. It is slow, and it is the actual mechanism.
05Write for a query someone actually types
Pages that get indexed and ranked answer a question completely enough that the reader stops searching. Pages that restate what a dozen other sites already say, at similar length and depth, add nothing to the index and are treated accordingly.
Target the specific problem phrasing rather than the broad topic. Long-tail queries have far less competition, and a page that comprehensively answers a narrow question will outperform a shallow page targeting a broad one - particularly on a domain with little authority.
Topics
Marcus Hale
Lead Architect · SyncTrix
Writes about the engineering decisions behind production systems - architecture, delivery and the trade-offs that only show up at scale.
Building something like this?
SyncTrix engineers AI, SaaS, platform and cloud systems for enterprises and high-growth teams. Tell us what you're shipping and we'll scope it with you.
Talk to an engineer