Insight

Search Console page indexing statuses: what “discovered” and “crawled – currently not indexed” actually tell you

The page indexing report in Google Search Console lists why pages are not indexed. Some reasons point to real problems, many are expected, and a few are signals about quality rather than technical errors. Reading them correctly prevents wasted effort and panicked site changes.

Published by Somnium Digital

A wireframe of the Insight page: headline, supporting sections and a single call to action. Insight Search Console page indexing stat… Get in touch 01 Crawling, indexing an… 02 Statuses that usually… 03 Statuses that often n…

Crawling, indexing and ranking

Google finds URLs through links, sitemaps and other sources, crawls some of them, processes the content and decides whether to include pages in its index. Only indexed pages can appear in search results, and even then they rank only for queries where Google considers them useful.

Search Console’s page indexing report shows how many known pages are indexed and groups the rest by reason. It is a diagnostic tool, not a score. A healthy website normally has many URLs that are not indexed, such as redirects, duplicates, filtered pages and deliberately excluded pages.

Statuses that usually reflect deliberate choices

Several reasons simply confirm the site is working as intended. They need attention only if important pages appear in them by mistake.

Excluded by noindex tag
The page asks not to be indexed. Correct for thank-you pages, internal search results or drafts; a problem if applied to key pages.
Page with redirect
The URL redirects elsewhere. Normal after migrations and URL changes.
Alternate page with proper canonical tag
The page points to another canonical version, which Google accepted.
Blocked by robots.txt
Crawling is disallowed. Note that blocked URLs can still be indexed without content if linked, so noindex is the tool for keeping pages out of the index.
Not found (404)
The page does not exist. Normal for removed content, unless important pages or linked URLs return 404 unexpectedly.

Statuses that often need fixing

Server error (5xx) means Google could not retrieve the page because the server failed. Recurring server errors reduce crawling and can remove pages from the index, so they should be investigated with hosting logs.

Soft 404 means the page returns a success status but looks like an error or empty page, such as a product page saying the item is unavailable with no content. Either return a proper 404 or 410 status, or give the page real content.

Duplicate, Google chose different canonical than user means Google ignored the canonical tag the site specified, often because the pages are very similar, internal links point elsewhere or the sitemap contradicts the tag. Signals should be aligned so the intended canonical is consistent everywhere.

Duplicate without user-selected canonical means Google found duplicates and no canonical was specified. Adding canonical tags and reducing unnecessary duplicate URLs, such as tracking parameters and printer versions, gives clearer signals.

Blocked due to unauthorised request (401) or access forbidden (403) usually means Google was blocked by authentication or firewall rules, sometimes because security settings treat crawlers as bots.

Discovered – currently not indexed

This status means Google knows the URL but has not crawled it yet. Google’s documentation notes that this commonly happens when crawling the site was expected to overload the server, so crawling was rescheduled.

It is common for new sites, large sites with many low-value URLs, and sites with slow servers. Improving server performance, reducing the number of unnecessary URLs, such as endless filter combinations, and strengthening internal links to important pages help Google prioritise what matters.

Crawled – currently not indexed

This status means Google crawled the page but decided not to index it, at least for now. Google says the page may or may not be indexed in the future, and there is no need to resubmit the URL.

There is no technical error to fix here in most cases. The status is often associated with pages Google considers not valuable enough to index: thin pages, near-duplicates of other pages, templated pages with little unique content, or pages on topics the site does not demonstrate strength in.

Requesting indexing repeatedly rarely changes the outcome. Improving the page’s unique value, consolidating similar pages, and linking to it from relevant, important pages are more effective. For large sites, a pattern of many pages in this status is a useful signal about which templates or content types are not earning their place.

Using the report well

Start with the pages that matter commercially: key service pages, product categories, top products, location pages and important articles. Use the URL Inspection tool to see the current status, the canonical Google selected, and whether the page can be fetched and rendered.

Submit accurate XML sitemaps containing only canonical, indexable URLs. The report can then be filtered by sitemap, which makes it easier to see whether intended pages are indexed rather than scanning all URLs Google has ever found.

Finally, keep indexing and ranking separate in reporting. Getting a page indexed is a prerequisite, not an achievement. Search performance reports showing impressions, clicks and queries tell whether indexed pages actually attract visitors.

Questions

Is it bad to have many non-indexed pages in Search Console?

Not necessarily. Redirects, duplicates, noindex pages and removed pages are normal. Focus on whether important pages are indexed.

What does “Discovered – currently not indexed” mean?

Google knows the URL but has not crawled it yet, often because crawling was rescheduled to avoid overloading the site.

What does “Crawled – currently not indexed” mean?

Google crawled the page but chose not to index it for now; it may be indexed later, and resubmitting is not needed.

Should we keep requesting indexing?

Repeated requests rarely help. Improving page value and internal linking is more effective.

Does robots.txt keep pages out of Google?

It blocks crawling, but blocked URLs can still be indexed without content. Use noindex to keep pages out of the index.

Why did Google choose a different canonical?

Usually because signals conflict, such as similar content, internal links, redirects or sitemaps pointing to other URLs.

Does indexing guarantee rankings?

No. Indexing only makes a page eligible to appear; ranking depends on relevance and quality for each query.

Where this sits in what we do

This article covers one decision inside a wider engagement. The solution page sets out how that engagement runs, what it includes and what it costs to find out.

Important pages missing from Google?

We diagnose indexing issues page by page, clean up duplicates and sitemaps, and improve the pages that matter so they earn their place in search.

Get in touch

Tell us what you are trying to change

Describe the problem rather than the service — the two frequently differ, and working out which is which is the useful part of a first conversation. We reply within one working day, and if it is outside what we do well you will hear that in the reply rather than after a call.

We use what you send to reply to you. Nothing else, and no list.

WhatsApp