What Is Indexing in SEO? How Google Indexes Pages
Indexing is the process by which a search engine fetches, parses, and stores a page in its database so the page can appear in search results. A page that is crawled but not indexed cannot rank.
Explain It Like I'm 5
Indexing is Google writing your page down in its giant library notebook. If Google reads your page but does not write it down, your page is not in the library - and it can never win a book contest it is not entered in.
Understanding Indexing
Indexing sits between crawling and ranking. First Googlebot discovers the URL, then it fetches and renders the page, then it parses the content and adds the URL to the index with all its associated signals. Only indexed pages can appear for searches.
Crawled but not indexed is the most common indexing problem: Google visited the URL, often multiple times, but decided not to store it. Search Console's Pages report lists this status per URL, and causes include thin or duplicate content, low quality signals, crawl-budget prioritization, and newness.
The index is enormous but not infinite. Google chooses what to index based on quality and uniqueness assessments. Sites that publish templated or near-duplicate pages at scale routinely see large crawled-not-indexed counts, which is the system working as designed.
Indexing can be controlled: noindex removes pages deliberately, canonicals consolidate duplicates into one indexed URL, sitemaps hint at what deserves priority, and internal links are the strongest within-site signal of which pages matter most.
Types of Indexing
Crawled, Indexed
Google fetched the page and stored it. It can rank.
Example: A healthy post appearing in Search Console with impressions and clicks.
Crawled, Not Indexed
Google fetched the page but chose not to store it. The page cannot rank.
Example: A thin category page Google visits repeatedly but never indexes.
Discovered, Not Crawled
Google knows the URL exists but has not fetched it yet, often a crawl-budget prioritization.
Example: A deep URL known only through an old sitemap entry.
Why Indexing Matters
Ranking debates are moot for unindexed pages. Every SEO effort presumes the page is in the library; index coverage is the gate everything else passes through.
Best Practices
Submit Sitemaps and Keep Them Clean
An accurate XML sitemap with lastmod dates helps Google prioritize what to crawl and index. Remove redirected, noindexed, and 404 URLs from it.
Build Internal Links to Every Page You Want Indexed
Pages reachable only through deep pagination or orphaned entirely are indexed far less reliably. Strong internal linking is an indexation lever.
Fix Crawled-Not-Indexed by Improving Value
Merge thin duplicates, strengthen content, and remove internal noindex conflicts. Re-crawling after fixes is requestable in Search Console.
Use noindex Deliberately
Tag archives, internal search results, and thin utility pages often should stay out of the index so crawl budget concentrates on pages that matter.
Common Mistakes
Treating crawled-not-indexed as a technical bug
Fix: It is usually Google's quality judgment. Fix the content or consolidate it; there is no tag that forces indexing.
Blocking crawl while expecting indexing
Fix: A page blocked in robots.txt cannot be crawled, so its noindex or canonical cannot be seen. Do not robots-block pages you want deindexed; let Google fetch them.
Index bloat through parameter URLs
Fix: Faceted navigation can multiply indexable URLs into the thousands. Canonicalize, noindex, or robots-disallow variants so the index holds your real pages.
How WPLink Drives Indexation
WPLink strengthens the internal link graph so every important page is well connected from related content, which is the strongest within-site signal Google uses when deciding what belongs in the index.
Frequently Asked Questions
Related Articles
Ready to optimize your internal links?
Get started with WPLink today and see the difference.
Download WPLink