Indexability is whether a web page can be added to a search engine’s index, the database of pages it can show in results. A page that is indexable has no technical barriers preventing it from being stored and ranked. A page that is not indexable will not appear in search results, no matter how good its content is. Indexability is closely related to crawlability, but it is a separate question: crawlability is whether a search engine can reach and read a page; indexability is whether it is allowed to, and chooses to, keep it.

That second part matters. A page can be fully indexable in the technical sense and still not be indexed, because Google decided it was not valuable enough to store. Good indexability work covers both: removing the barriers you did not intend, keeping out the pages you do not want in search, and making sure the pages you do want are strong enough to earn their place.

Crawling, indexing, and ranking

Search engines work in stages. Discovery is finding that a URL exists, through links, sitemaps, or other sources. Crawling is fetching the page. Rendering is processing its HTML, CSS, and JavaScript to see the content as a browser would. Indexing is analyzing the page and deciding whether to store it. Ranking is choosing which indexed pages to show for a query. A problem at any earlier stage prevents everything after it, which is why indexability is foundational to all other SEO work.

What makes a page non-indexable

Deliberate signals

Noindex directives. A noindex value in a robots meta tag, or in an X-Robots-Tag HTTP header, tells search engines not to include the page. This is the correct tool for pages you want accessible but not in search, such as thank-you pages, internal search results, or thin archive pages.

Canonical tags pointing elsewhere. A canonical tag that names a different URL tells search engines that another page is the preferred version. Google usually honors it and indexes the canonical instead.

Redirects. A URL that 301 redirects to another is not indexed itself; the destination is.

Accidental barriers

Robots.txt blocks. Disallowing a URL in robots.txt prevents crawling, not indexing. Google can still index a blocked URL if other pages link to it, but without seeing its content, and it cannot see a noindex tag on a page it is not allowed to crawl. Robots.txt is for managing crawl, not for keeping pages out of search.

Leftover noindex tags. The most common and most damaging indexability error is a site launched with the staging environment’s “discourage search engines” setting still on, or a template that applies noindex more broadly than intended.

Server errors and soft 404s. Pages returning 5xx errors, or pages that return a 200 status but look like error pages, are dropped.

Content that only appears with JavaScript. Google can render JavaScript, but rendering may be delayed, and content that depends on user interaction or blocked scripts may never be seen. Many AI crawlers do not render JavaScript at all.

Login walls and blocked resources. Content behind authentication cannot be indexed, and CSS or JavaScript files blocked in robots.txt can prevent Google from rendering the page correctly.

Crawled, currently not indexed

Google Search Console’s page indexing report separates pages that have technical problems from pages Google simply chose not to index. Two statuses in particular cause confusion.

Crawled, currently not indexed means Google fetched the page and decided not to store it. There is no technical error; the page was judged not valuable or distinct enough, or too similar to other pages. The fix is usually to improve the page, consolidate it with a stronger sibling, or accept that it does not need to be in search.

Discovered, currently not indexed means Google knows the URL exists but has not crawled it yet. On smaller or lower-authority sites this often reflects crawl priority: Google does not yet consider the page important enough to fetch. Stronger internal links, inclusion in the sitemap, and requesting indexing in Search Console help.

These reports also retain history. A URL crawled months ago, before you added a redirect or rewrote it, will show its old status until Google recrawls it. Validating the fix in Search Console prompts a recrawl.

How to check indexability

URL Inspection in Search Console shows whether a specific URL is indexed, which canonical Google selected, when it was last crawled, and whether indexing is allowed. It also lets you test the live URL and request indexing.

The page indexing report shows every known URL grouped by status and reason, which is the fastest way to spot patterns, such as a whole template excluded by noindex.

A site crawler such as Screaming Frog lists every page’s status code, robots directives, and canonical, so conflicts are easy to find.

The page source and response headers confirm what search engines actually receive, including any X-Robots-Tag header.

Controlling what gets indexed

Good indexability is not about getting every URL indexed. It is about getting the right ones indexed. Most sites generate URLs that should stay out of search: tag and date archives, filtered and sorted product listings, internal search results, attachment pages, feeds, and staging copies. Leaving them indexable wastes crawl attention and fills the index with thin pages that can dilute a site’s overall quality. The common tools are noindex for pages that should be reachable but not searchable, canonical tags for near-duplicates, redirects for retired URLs, and an XML sitemap that lists only the URLs you want indexed.

A lesson from our own site

In mid-2026 we found that our own XML sitemaps were being served from a stale edge cache. The sitemap itself had been corrected in WordPress, and every check that added a query string to the URL showed the right version, but the bare sitemap URL that Googlebot actually requests kept returning an outdated copy. Purging the host’s edge cache fixed it, and indexing climbed steadily over the following weeks. The lesson we took: always verify what search engines receive at the exact URL they request, not what the CMS says it generated, and re-check after any change that affects sitemaps, redirects, or robots directives. It is now a standing step in every technical audit we run.

Indexability on WordPress

WordPress sites have a few recurring indexability traps. The “Discourage search engines from indexing this site” setting under Reading applies a sitewide noindex and is easy to carry over from staging at launch. SEO plugins can noindex whole content types or taxonomies with a single toggle, which is useful for tag archives and attachment pages but damaging if applied to the wrong post type. Custom post types sometimes launch without being included in the sitemap, and duplicated or migrated pages can end up with stale SEO data that keeps them out of it. After any launch, migration, or plugin change, check the page indexing report and a crawl of the site before assuming everything is visible.

Indexability and site quality

Because Google decides which crawlable pages to keep, indexability is ultimately tied to content quality. Sites with large numbers of thin, overlapping, or low-value pages tend to see more of them excluded, and Google’s systems evaluate quality partly at the site level. Consolidating weak pages into stronger ones, noindexing pages that do not serve searchers, and expanding pages that deserve to rank all improve how much of the rest of the site gets indexed and how well it performs.

Indexability is one of the first areas we review in our technical SEO services and every SEO audit. If important pages are missing from Google, book a discovery call and we will find out why.