Australian Website Design Measured figures. Named sources.
Menu Close

Search readiness

How search engines crawl and index a website

Discovery, crawling, indexing and ranking are four separate stages. Knowing which one has failed is what makes a search problem diagnosable.

In short. Search works in four stages: discovery, crawling, indexing and ranking. Most "we are not on Google" problems are a failure at stage one, two or three, and each has a different cause and a different fix. Diagnosing which stage failed is the difference between fixing a problem and buying a service that cannot address it.

Search works in four separate stages, and knowing which one has failed is what makes a search problem diagnosable rather than a subject for opinion. Google is the search engine most Australian small businesses actually mean when they say “search engines,” so this page uses Google’s own terms — crawling, indexing, Googlebot — throughout, though the same four stages apply to any search engine.

Discovery is a search engine learning that an address exists. Crawling is Googlebot, or another search engine’s crawler, fetching that address. Indexing is the search engine deciding to store and process what it fetched. Ranking is it selecting from what it has indexed when somebody runs a search.

Almost every “we are not appearing on Google” complaint is a failure at one of the first three stages. Almost every response sold to that complaint addresses the fourth.

The four stages of what search engines actually do, and what can go wrong at each

StageWhat happensCommon failureWhose problem
DiscoveryThe search engine learns the URL existsNo links to the page, not in the sitemapThe build
CrawlingThe search engine’s crawler fetches the pageBlocked by robots.txt, server errors, login requiredThe build
IndexingThe search engine stores and processes the pageExcluded by an instruction, duplicate of another page, judged not worth storingThe build, mostly
RankingThe search engine selects among indexed pagesThe page is not the best answer available for that searchContent and competition

The first three are build responsibilities and are checkable. The fourth is a competition and has no completion state.

Discovery: how search engines find URLs and pages

A search engine like Google finds an address, or URL, in one of three ways: a link to it from a page Google already knows about, an entry in a submitted sitemap file, or a direct submission through Google Search Console.

A page with no internal links and no sitemap entry is invisible to search engines in the most literal sense. That is why an orphan page — one that exists but that nothing links to — is a genuine defect, not just untidiness. And it is why the internal linking structure decided during the sitemap stage is not merely a usability matter for search engines like Google.

Crawling: crawl budgets, crawlers, robots.txt and Googlebot

Crawling is a fetch: Googlebot, or another search engine’s crawler bot, requesting a page from the server. It can fail for reasons that have nothing to do with content.

  • The site instructs crawlers not to fetch the address. That is what a robots.txt file, sometimes written “robots txt” without the dot, does — and it is regularly misused by accident, blocking Googlebot and other search engines’ crawlers from pages that should be crawled.
  • The server returns an error, or is slow enough that the crawler’s fetch times out.
  • The page requires a login, or sits behind a password left over from staging.
  • The address redirects in a loop, or redirects to something unrelated.

Crawling is also budgeted. Google will not fetch an unlimited number of URLs from one site in one visit — this is Google’s crawl budget for that site. A site that generates thousands of near-identical URLs spends that crawl budget on pages nobody wants crawled. That is a structural problem produced by filters, session identifiers and parameters. It is the main practical reason to care about how many URLs a site actually generates.

Indexing: crawling, indexing and what gets indexed

Indexing is a decision Google makes, not an automatic consequence of crawling. A page can be crawled by Googlebot and then not indexed.

The causes divide into instructed and judged.

Instructed. The page carries a directive telling search engines not to index it, or points a canonical link at a different page, or is a duplicate of a page already indexed. These are mechanical, they are visible in Google Search Console, and they are fixable.

Judged. Google’s crawler fetched the page, its indexing system saw nothing worth storing, and moved on. Thin pages, near-duplicate pages and pages generated from a template with one word changed are the usual population Google leaves unindexed. There is no setting to flip for this, and no appeal. The fix is that the page has to become worth indexing.

The second category is the reason the location-page caution in the structure guides is a caution rather than a preference. A site can produce forty crawlable, technically indexable pages that Google simply does not index.

Ranking

Ranking is what most people mean by SEO, and it is the stage a build has the least control over. It is Google comparing every other indexed page that could answer the same search query, and choosing an order.

Two honest statements about ranking. First, Google’s ranking factors are not published in full, and anyone presenting a complete list is presenting a model, not a fact. Second, on a new domain there is no history for Google to assess, so early ranking performance says very little about whether the site was built well.

How to tell which stage failed for a specific search query

Google Search Console is the only reliable instrument for diagnosing crawling and indexing problems. It reports what Google’s crawlers actually did, not what a third-party tool inferred. It will tell you whether a specific URL was discovered, when Googlebot last crawled it, whether it was indexed, and if not, which reason Google gives.

That single capability is why account ownership matters, and why it should sit with the business rather than with the supplier. That is covered on Search Console for a new site.

For a site that has never appeared in Google at all, the order of checks is: is the site blocking Google’s crawlers, is there an indexing directive left over from staging, and does the site actually resolve at one address rather than several. The first two are on robots.txt and noindex and taking a website out of staging.

What this means when commissioning

The four stages of crawling and indexing are why the build-side obligations are a finite list. A supplier is responsible for stages one to three — discovery, crawling and indexing — being possible for search engines to complete. Nobody is responsible for stage four, ranking, and nobody can be.

The commercial version of that boundary is on whether SEO is included in a build, and the scope of a build generally is on website design services.

What to do next: checking whether your pages were crawled

Run a Google search for your own business name in quotes. If nothing appears in the search results, the failure is at stage one, two or three, and it is a build problem — discovery, crawling or indexing has failed. If your site appears in Google’s index for its own name but not for anything else, the first three stages are working, and you are looking at stage four, which is a different conversation with a different supplier.

Evidence for this page

This page exists because the demand below was measured, not assumed. The figures are search-market data about the topic — they are not prices.

Entity this page targets
how search engines crawl and index a website
Measured Google volume
no data
Keyword difficulty
no data
Advertiser cost per click
no data
AI assistant volume
no data
Advertiser competition
no data
Measured on
31 July 2026
Search results inspected for intent
No
3 other phrasings resolve to this same page

crawling and indexing explained · how does google find my website · website indexing explained

Absent from the measured Australian universe in research/national-volume-au.json. The page is justified by coverage: a buyer cannot judge any of the other decisions in this section without the four stages described here.

Source: research/national-volume-au.json · DataForSEO Labs, location_code 2036 (Australia), language en · pulled 31 July 2026.

Provenance

Written by Australian Website Design. Published 2026-08-03, last updated 2026-08-03.

Sources

  • National keyword volume and difficulty, Australia — research/national-volume-au.json (accessed 2026-07-31)