Methods

How the deposit book is kept, and where it has gaps.

What is recorded

Every GET or HEAD request whose User-Agent string contains one of the crawler names configured on this site is written to a database at the edge, as it happens: the page requested, the time, the name the request gave, the operator that name belongs to, the purpose that operator declares for it, the method, the status the server returned, the doorman mode in force, the verification result described below, the network the request arrived from (the organisation registered for its address block, never the address itself), the User-Agent string, and the article revision served. Requests from people are not examined and not written.

What counts as served

A request is any matching GET or HEAD. Served is narrower: a GET answered 200 with an HTML or plain-text body. HEAD probes, 304 Not Modified responses, errors, and the 402 gate (when enabled) are all recorded but not counted as served. The totals keep both numbers.

Declared purpose

Agents are grouped as training, answering, or search according to the purpose their operators declare for that agent name. That is a declaration by the operator, not something this server can observe, and the grouping is maintained by hand.

Verification

A User-Agent string is a claim. Where an operator publishes the address ranges its agents use, the requesting address is checked against that list (fetched live, cached for six hours). The result is one of: verified — within the ranges published for that agent; verified (operator) — within a list the operator publishes for all of its agents, so the operator is supported but the specific agent name remains a claim; not in ranges — the operator publishes ranges and the address is not among them, so the claim is not supported; no method configured — no verification source is set up here for that name, which is a fact about this site, not about the operator; unchecked — the list could not be fetched at the time. Sources configured at the time of writing:

OperatorAgentsScopeSource
OpenAIGPTBot · OAI-SearchBot · ChatGPT-Userper agentopenai.com/gptbot.json
AnthropicClaudeBot · Claude-User · Claude-SearchBotoperator-wideclaude.com/crawling/bots.json
Common CrawlCCBotper agentindex.commoncrawl.org/ccbot.json
AmazonAmazonbot · Amzn-SearchBot · Amzn-Userper agentdeveloper.amazon.com/amazonbot/ip-addresses/
GoogleGooglebotper agentdevelopers.google.com/static/crawling/ipranges/common-crawlers.json
Microsoftbingbotper agentwww.bing.com/toolbox/bingbot.json
PerplexityPerplexityBot · Perplexity-Userper agentwww.perplexity.ai/perplexitybot.json
AppleApplebotper agentsearch.developer.apple.com/applebot.json

Article revisions

Each article has a revision hash computed from its body at build time, alongside its revision date. The ledger records which revision was served with each request, so that once an article is corrected it is possible to say how many requests received the version before the correction and how many after. The hash identifies the article as built. The response a training crawler receives also carries per-request lines and the terms, so the full response is a different object from the article.

Gaps

The record covers requests that reached this server's worker while its database was writable. It does not include: requests refused at Cloudflare's edge before reaching the worker (bot management, firewall rules, rate limits); requests made after the worker's daily free allowance is exhausted, when pages are served statically without logging; requests made after the database's daily write allowance is exhausted; clients that do not identify themselves, which pass as people; and anything before the first entry. A quiet day in the record may be a quiet day or a gap. Absence from this book is not evidence of absence.

Retention and display

Individual entries are kept for ninety days; daily totals are kept indefinitely. The public page is cached for five minutes. When the database cannot be read the page says so explicitly rather than showing an empty table.

What this does not establish

Whether a page was stored, used in training, used to ground an answer, or read by anyone. Whether any later system received a later revision. Whether a crawler encountered, understood or accepted the terms it was served. The record stops at the edge of this server.