How This Site Sources Its Numbers
The sourcing method behind every figure on this site: which sources qualify, what is excluded, how often numbers are refreshed, and the limits of what any of it proves.
How ageing content is handled
Pages written earlier are preserved as they were written. Each carries a dated context update at the foot: what changed since, added by the writer rather than rewritten into the original.
The old text is the record; the update is the service note.
Rewriting history to look correct is the easiest way to stop being trustworthy. This site would rather show its working.
This site does not measure the internet. It gathers, dates, and contextualises measurements made by others — research institutions, government agencies, peer-reviewed papers, and the security firms that run the infrastructure. Every figure names its source and the month it was published, because every figure will eventually be wrong.
What qualifies as a source
A number ships only if it comes from one of these, and only with attribution and a date attached:
- Primary reports from the organisation that did the measuring — for example Imperva and Thales on bot traffic, or the FBI's Internet Crime Complaint Center on fraud losses.
- Peer-reviewed research and recognised academic benchmarks — for example the RAID detector benchmark presented at ACL, or detector-bias research published in Patterns.
- Government and intergovernmental bodies — agencies, regulators, and organisations such as the International Energy Agency.
- First-hand reporting by established newsrooms, when the underlying study is not public — cited as the reporting, not as the study.
- Direct statements from the organisation concerned, labelled as their claim rather than as independent fact.
What is excluded, always
- Marketing claims from vendors selling the thing being measured — especially AI-detection accuracy figures, which are routinely self-reported.
- SEO content roundups and trade blogs restating a number whose origin they never name.
- Figures with no traceable original source, however often they are repeated. A number repeated a thousand times is still one unverified number.
- Predictions presented as measurements.
How figures are handled
- Dated in place. Each figure carries the source and publication month beside it, not in a footnote you will never open.
- Stated as estimates where they are estimates. Nobody can census the web. Detection-derived percentages are approximations produced by imperfect tools, and this site says so next to the number.
- Counter-signals published alongside. Where evidence cuts against the site's own framing, it appears in the same section, not a quieter one.
- Refreshed quarterly. The record page is re-cut every three months around whatever is newest; superseded figures move to its archive rather than vanishing.
- Failed predictions kept. When a widely repeated forecast does not come true — as with the claim that 90% of online content would be synthetic by 2026 — it stays on the page, marked failed.
WHAT THIS SITE IS NOT
Not original research. Not a census. Not a detector, and not an authority on whether any particular thing was made by a machine. The tools that claim to answer that question are, by independent testing, too unreliable to accuse anyone with — and this site will not do it either.
Who writes this
The site is written under a persona — the writer — by someone who works in generative AI: making images, text, and video with these tools, daily. That is a deliberate choice, and worth stating plainly rather than leaving you to guess.
Why anonymous: for personal privacy, and to keep attention on the sources rather than on a byline. Why it should still be checkable: because nothing here rests on the writer's authority. Every claim traces to a named institution with a date. You are never asked to trust the writer — only to check the arithmetic, which you can.
The perspective matters more than the name: this is written from inside the industry producing the content, not from outside it. That is also the site's conflict of interest, and it is disclosed on .
Corrections
Figures change. When one on this site is wrong, or is superseded, the change is published as a dated note rather than quietly edited away. The running list is on the corrections page.
Verification before publication
Before this site went public, every sourced claim on every page was traced back to its issuing organisation — not to reporting about it, but to the report, the paper, the press release. That process ran six times, and it found real errors:
- A statistic about automated attacks that could not be traced to any source. Removed and replaced with a figure that could.
- A percentage that should have been a count — a secondary write-up had rendered "17 countries out of 194" as "17%", roughly doubling it. Corrected against the issuing body's own release.
- False precision: a figure quoted to one decimal place that the underlying research did not support. Replaced with a range, and the variance stated on the page.
- A quotation attributed to the wrong publication.
- An article that had been published under the wrong page title entirely.
- A widely repeated statistic reproduced from its popular misreading rather than the paper. Coverage rendered a 2024 study as "57% of the web is machine translated"; the paper actually reports that 57.1% of sentences within a large multi-way parallel corpus appear in three or more languages, and that corpus derives from web snapshots collected between 2017 and 2020. The site had repeated the headline version. Corrected to state the measurement, the scope and the date of the underlying data.
None of that appears in the corrections log, because no reader was ever shown it. The log begins at publication. This section exists so the process is on the record rather than implied.
What the process now enforces on every page: trace each figure to the organisation that issued it · cross-check at least two independent accounts · check the shape of a number, not only its source, since a count presented as a percentage looks entirely normal · prefer imprecision to false precision · and mark anything single-sourced, contested or not re-verified as exactly that.