Analytics
Bots, spam and inflated numbers in Google Analytics
Bot traffic in Google Analytics — why a share of every site's traffic is not human, how GA4 attempts to filter it, and what that means for a raw figure.
No website’s raw traffic figure is a perfectly clean count of real human visitors. Some share is automated. GA4 filters a meaningful portion of it automatically, but treating any traffic number as unquestionably accurate misreads what the figure actually represents.
What GA4 filters automatically: known bots and Google Analytics exclusions
GA4 applies bot filtering against a recognised industry list of known bots and spiders, the IAB/ABC International Spiders and Bots List. It automatically excludes traffic matched against it from standard reports. This catches a meaningful share of identifiable automated crawlers and scrapers without requiring any setup.
What it does not reliably catch: unusual or suspicious bot traffic
Newer or deliberately disguised bots not yet matched against the recognised list, including some automated traffic built specifically to mimic genuine browser behaviour.
Referrer spam is where an automated script sends traffic designed to appear as if it came from a specific website, without ever genuinely visiting the page. It was historically used to get a fake referring domain to appear in someone’s traffic reports. This kind of traffic can arrive through mechanisms that bypass standard bot filtering entirely, depending on how it is generated.
Traffic from automated monitoring and testing tools includes a business’s own uptime monitors or a supplier’s testing scripts. These are not malicious, but they are still not genuine visitors and can inflate numbers if not deliberately excluded.
Why GA4 never sees most bot traffic in the first place
GA4 only records a visit when the tracking code actually executes in a browser. That requires JavaScript to run and the tracking request to fire. The overwhelming majority of automated bots, including most search engine crawlers and simple scrapers, do not execute JavaScript at all. They never trigger a GA4 event in the first place. This is a structurally different situation from a raw web server log. A server log records every single request that reaches the server regardless of whether anything executes afterward, including every bot and crawler hit. A business comparing its GA4 traffic figures against its hosting provider’s raw server log numbers will often see a large gap between the two, for exactly this reason. The gap itself is not evidence of a problem with either measurement. The two are counting genuinely different things.
AI crawlers are a newer category worth knowing about specifically
Beyond traditional search engine crawlers, a newer category of automated traffic comes from crawlers operated by AI companies. They gather content for model training or for generating cited answers in AI assistants. These identify themselves with their own named user agents in server logs, distinct from Googlebot or Bingbot. Like most crawlers, they generally do not execute JavaScript, so they do not appear in GA4 at all. A business specifically interested in whether AI assistants are citing its content needs a different measurement approach entirely from standard analytics. Neither GA4 nor a simple bot-exclusion list is built to answer that question.
Data centre IP ranges are a further identifying signal
Traffic originating from known data centre and cloud-hosting IP address ranges, rather than a residential or mobile internet connection, is a further practical signal of automated traffic — beyond named bots on a recognised list. An ordinary visitor browsing from home or from a phone does not appear from a data centre’s own IP range. This is a supplementary signal worth checking when investigating an unusual traffic pattern, alongside the bounce-rate and engagement-time signatures already covered above.
Why this matters for reading a traffic report’s pageviews and visits honestly
A sudden, unexplained spike in traffic with an unusually high bounce rate, near-zero engagement time, and traffic concentrated from an unfamiliar source is a common signature of bot or spam traffic, rather than a genuine surge in interest. Recognising this pattern prevents mistaking a data-quality issue for a real result. It also prevents mistaking a batch of non-human traffic for a real problem. When investigating why a change “did not work” as the analytics suggested, that non-human traffic can be the actual explanation distorting the comparison.
A short, checkable list of warning signs in your analytics reports
| Signal | What it suggests |
|---|---|
| A sudden traffic spike with near-zero engagement time and 100% bounce | Likely automated, not genuine interest |
| Traffic concentrated from an unfamiliar or irrelevant referring domain | Possible referrer spam |
| Traffic from a single unusual source with no corresponding change in marketing activity | Worth checking the source in GA4’s specific referral report before drawing conclusions |
| Consistent, small, unexplained traffic from a data centre or hosting-provider location | Often automated monitoring or scraping traffic |
What this does not mean for the rest of this cluster’s advice
None of this justifies inflating a genuinely low number by attributing it to “probably bots” without evidence — the same honesty standard that applies to not fabricating a positive result applies here in reverse. The point is to recognise clear, checkable signatures of non-human traffic where they exist, not to explain away disappointing figures without cause.
What to do next: exclude known bots in Google Analytics
Set up internal traffic exclusion (covered in setting up GA4: the first week checklist) as the first, most reliable step, since it removes a known, identifiable source of non-genuine traffic. Beyond that, treat an unusual spike as a prompt to check the source breakdown before treating it as either good or bad news. For the underlying tracking setup this connects to, web development services covers where analytics configuration sits in a build.
Evidence for this page
This page exists because the demand below was measured, not assumed. The figures are search-market data about the topic — they are not prices.
- Entity this page targets
- bot traffic in google analytics
- Measured Google volume
- no data
- Keyword difficulty
- no data
- Advertiser cost per click
- no data
- AI assistant volume
- no data
- Advertiser competition
- no data
- Measured on
- 31 July 2026
- Search results inspected for intent
- No
2 other phrasings resolve to this same page
fake traffic google analytics · spam referral traffic
Not present in the measured keyword set. A genuine null.
Source: research/national-volume-au.json · DataForSEO Labs, location_code 2036 (Australia), language en · pulled 31 July 2026.
Provenance
Written by Australian Website Design. Published 2026-08-03, last updated 2026-08-03.
Sources
- Google Analytics 4 Help — Bot filtering (accessed 2026-08-03)