StatableScanner

Last updated: August 29, 2026

If you found this address in your server logs, StatableScanner requested a page from your site. This page explains what it is, what it asked for, and how to stop it. You do not need to contact us to make it go away — the robots.txt instructions below are enough.

TL;DR

  • It is a research crawler operated by Statable
  • It reads public pages only — no logins, no forms, no paywalled or member content
  • It obeys robots.txt and never tries to work around a block
  • It waits at least 2 seconds between requests to a site
  • It measures which third-party services a page loads, not who your visitors are
  • It never submits data to your site and never attempts to log in

How to identify it

Two builds of the crawler are in use. Either of these user agent strings is us:

Mozilla/5.0 (compatible; StatableScanner/0.1; +https://statable.com/scanner-bot; scanner@statable.com)
StatablePreConsentScan/0.1 (+https://statable.com/scanner; research contact: hello@statable.com)

The second one points at /scanner, which redirects here. Both reach the same team; scanner@statable.com is the address to use.

Anything claiming to be our scanner from a different address is not us. The user agent always carries both this page and a working contact address, and we never disguise it as a browser or as another company's crawler.

What it does

Statable publishes research on what websites load before a visitor has consented to anything — trackers, analytics tags, fonts, embeds and the cookies they set. To measure that, the scanner opens a page the way a browser would and records which third-party servers the page contacts and when.

The subject of the measurement is your page's own behaviour. We do not collect information about your visitors, we do not have access to it, and a scan tells us nothing about who visits your site.

What it requests

  • /robots.txt, before anything else
  • Your home page, and in some runs a small number of additional public pages
  • The sub-resources those pages reference — scripts, stylesheets, images — because that is what the measurement is about

It moves the mouse pointer and scrolls, because a large family of performance plugins defers third-party scripts until one of those events. It never clicks, never presses a key, and never interacts with a consent banner or any other control.

How to block it

Add this to your robots.txt. The scanner reads robots.txt before every site and honours it:

User-agent: StatableScanner
Disallow: /

User-agent: StatablePreConsentScan
Disallow: /

A blocked domain is recorded as blocked and is never loaded. It is not retried under another name, from another address, or with a different user agent, and a block is never treated as a signal to look harder.

If you would rather we skipped your site without editing robots.txt, email us and we will add it to a permanent exclusion list.

What we keep, and for how long

A scan produces a network record of the page load: the addresses contacted, when, the request headers, and the names and attributes of the cookies that were set.

  • Response bodies are not stored. We keep the fact that a request happened, not the content that came back
  • Cookies set during the scan belong to the scanner's own throwaway browser session. No real visitor is involved at any point, so there is no visitor data to hold
  • Network records and screenshots are kept for 90 days, after which only the derived aggregate facts remain
  • Aggregates carry no personal data and are kept for as long as the research is published

Every published figure can be traced back to a specific scan: which build ran it, on what date, from which country, and with which browser locale. If you ask about a result, that is what we check it against.

Rate and load

One page per domain in a normal run, at least two seconds between requests, and a hard time limit per domain. The scanner does not crawl a site breadth-first, does not follow internal links beyond the pages named above, and is not designed to be run repeatedly against one site.

If it ever costs you noticeable resources, that is a bug on our side. Write to us and we will stop and fix it.

If we published something about your site

Our findings are behavioural measurements taken on a specific date — what a browser observed your page load — and they are not legal conclusions about your site. Sites change, and a measurement can also simply be wrong.

If you believe a published result about your site is inaccurate, write to us. We keep the underlying evidence for every scan, including the full network record, and we will check it against your description and correct the record if it is wrong.

Contact

Write to scanner@statable.com for anything about the crawler: exclusion requests, questions about a published result, or a complaint about its behaviour. For everything else, support@statable.com reaches our normal support.

Statable is operated by Key Arg B.V., registered in the Netherlands.