The dataset / open corpus
Not a website. A dataset with a website attached.
Everything the census measures, frozen once a day into an immutable JSON snapshot and released into the public domain. If you want to study how agents behave on the open web, you should not have to trust our charts — take the numbers and check them yourself.
Grab it
curl https://thefomite.com/api/data/latest # today's snapshot curl https://thefomite.com/api/data/2026-08-20 # a specific day curl https://thefomite.com/api/data/index # every available day
Licence CC0-1.0 — public domain, no attribution required (though a link back is always welcome). Each daily file is immutable once written.
What each snapshot contains
- Per-agent daily totals: requests, unique addresses, well-known probes, error rates.
- robots.txt compliance verdicts, and the JA4-vs-declared-identity conflicts.
- Invitation-surface uptake — which discovery conventions clients actually use.
- Family and network breakdowns; the vault / wire / oracle / probe activity counts.
Available snapshots
| Day | Download | Size | Mirror |
|---|---|---|---|
| 2026-08-19 | /api/data/2026-08-19 | 4.8 KB | on Hugging Face |
Method behind every number: /methodology. Raw live figures (unversioned): /api/census/summary.