Search readiness
How search engines crawl and index a website
Discovery, crawling, indexing and ranking are four separate stages. Knowing which one has failed is what makes a search problem diagnosable.
In short. Search works in four stages: discovery, crawling, indexing and ranking. Most "we are not on Google" problems are a failure at stage one, two or three, and each has a different cause and a different fix. Diagnosing which stage failed is the difference between fixing a problem and buying a service that cannot address it.
Search works in four separate stages, and knowing which one has failed is what makes a search problem diagnosable rather than a subject for opinion. Google is the search engine most Australian small businesses actually mean when they say “search engines,” so this page uses Google’s own terms — crawling, indexing, Googlebot — throughout, though the same four stages apply to any search engine.
Discovery is a search engine learning that an address exists. Crawling is Googlebot, or another search engine’s crawler, fetching that address. Indexing is the search engine deciding to store and process what it fetched. Ranking is it selecting from what it has indexed when somebody runs a search.
Almost every “we are not appearing on Google” complaint is a failure at one of the first three stages. Almost every response sold to that complaint addresses the fourth.
The four stages of what search engines actually do, and what can go wrong at each
| Stage | What happens | Common failure | Whose problem |
|---|---|---|---|
| Discovery | The search engine learns the URL exists | No links to the page, not in the sitemap | The build |
| Crawling | The search engine’s crawler fetches the page | Blocked by robots.txt, server errors, login required | The build |
| Indexing | The search engine stores and processes the page | Excluded by an instruction, duplicate of another page, judged not worth storing | The build, mostly |
| Ranking | The search engine selects among indexed pages | The page is not the best answer available for that search | Content and competition |
The first three are build responsibilities and are checkable. The fourth is a competition and has no completion state.
Discovery: how search engines find URLs and pages
A search engine like Google finds an address, or URL, in one of three ways: a link to it from a page Google already knows about, an entry in a submitted sitemap file, or a direct submission through Google Search Console.
A page with no internal links and no sitemap entry is invisible to search engines in the most literal sense. That is why an orphan page — one that exists but that nothing links to — is a genuine defect, not just untidiness. And it is why the internal linking structure decided during the sitemap stage is not merely a usability matter for search engines like Google.
Crawling: crawl budgets, crawlers, robots.txt and Googlebot
Crawling is a fetch: Googlebot, or another search engine’s crawler bot, requesting a page from the server. It can fail for reasons that have nothing to do with content.
- The site instructs crawlers not to fetch the address. That is what a robots.txt file, sometimes written “robots txt” without the dot, does — and it is regularly misused by accident, blocking Googlebot and other search engines’ crawlers from pages that should be crawled.
- The server returns an error, or is slow enough that the crawler’s fetch times out.
- The page requires a login, or sits behind a password left over from staging.
- The address redirects in a loop, or redirects to something unrelated.
Crawling is also budgeted. Google will not fetch an unlimited number of URLs from one site in one visit — this is Google’s crawl budget for that site. A site that generates thousands of near-identical URLs spends that crawl budget on pages nobody wants crawled. That is a structural problem produced by filters, session identifiers and parameters. It is the main practical reason to care about how many URLs a site actually generates.
Indexing: crawling, indexing and what gets indexed
Indexing is a decision Google makes, not an automatic consequence of crawling. A page can be crawled by Googlebot and then not indexed.
The causes divide into instructed and judged.
Instructed. The page carries a directive telling search engines not to index it, or points a canonical link at a different page, or is a duplicate of a page already indexed. These are mechanical, they are visible in Google Search Console, and they are fixable.
Judged. Google’s crawler fetched the page, its indexing system saw nothing worth storing, and moved on. Thin pages, near-duplicate pages and pages generated from a template with one word changed are the usual population Google leaves unindexed. There is no setting to flip for this, and no appeal. The fix is that the page has to become worth indexing.
The second category is the reason the location-page caution in the structure guides is a caution rather than a preference. A site can produce forty crawlable, technically indexable pages that Google simply does not index.
Ranking
Ranking is what most people mean by SEO, and it is the stage a build has the least control over. It is Google comparing every other indexed page that could answer the same search query, and choosing an order.
Two honest statements about ranking. First, Google’s ranking factors are not published in full, and anyone presenting a complete list is presenting a model, not a fact. Second, on a new domain there is no history for Google to assess, so early ranking performance says very little about whether the site was built well.
How to tell which stage failed for a specific search query
Google Search Console is the only reliable instrument for diagnosing crawling and indexing problems. It reports what Google’s crawlers actually did, not what a third-party tool inferred. It will tell you whether a specific URL was discovered, when Googlebot last crawled it, whether it was indexed, and if not, which reason Google gives.
That single capability is why account ownership matters, and why it should sit with the business rather than with the supplier. That is covered on Search Console for a new site.
For a site that has never appeared in Google at all, the order of checks is: is the site blocking Google’s crawlers, is there an indexing directive left over from staging, and does the site actually resolve at one address rather than several. The first two are on robots.txt and noindex and taking a website out of staging.
What this means when commissioning
The four stages of crawling and indexing are why the build-side obligations are a finite list. A supplier is responsible for stages one to three — discovery, crawling and indexing — being possible for search engines to complete. Nobody is responsible for stage four, ranking, and nobody can be.
The commercial version of that boundary is on whether SEO is included in a build, and the scope of a build generally is on website design services.
What to do next: checking whether your pages were crawled
Run a Google search for your own business name in quotes. If nothing appears in the search results, the failure is at stage one, two or three, and it is a build problem — discovery, crawling or indexing has failed. If your site appears in Google’s index for its own name but not for anything else, the first three stages are working, and you are looking at stage four, which is a different conversation with a different supplier.
Evidence for this page
This page exists because the demand below was measured, not assumed. The figures are search-market data about the topic — they are not prices.
- Entity this page targets
- how search engines crawl and index a website
- Measured Google volume
- no data
- Keyword difficulty
- no data
- Advertiser cost per click
- no data
- AI assistant volume
- no data
- Advertiser competition
- no data
- Measured on
- 31 July 2026
- Search results inspected for intent
- No
3 other phrasings resolve to this same page
crawling and indexing explained · how does google find my website · website indexing explained
Absent from the measured Australian universe in research/national-volume-au.json. The page is justified by coverage: a buyer cannot judge any of the other decisions in this section without the four stages described here.
Source: research/national-volume-au.json · DataForSEO Labs, location_code 2036 (Australia), language en · pulled 31 July 2026.
Provenance
Written by Australian Website Design. Published 2026-08-03, last updated 2026-08-03.
Sources
- National keyword volume and difficulty, Australia —
research/national-volume-au.json(accessed 2026-07-31)