# Threat intelligence

> The public IP data behind AgentGate's country, ASN, cloud, Tor and verified-crawler labels, how it refreshes, and what it does not cover.

## Data sources

AgentGate downloads public, provider-published data and labels requests with
it. There is no CDN-scale telemetry behind these labels: they are exactly as
good as the files below.

| Label / score | Source (URL checked live 2026-09-24) | Licence / terms | Refresh |
| --- | --- | --- | --- |
| `agentgate:ip:asn:<n>`, `agentgate:ip:country:<cc>`; `asn`, `country` statements | iptoasn.com `ip2asn-combined.tsv.gz` | Public domain (PDDL v1.0, stated on iptoasn.com) | daily |
| `agentgate:ip:cloud:aws` | `ip-ranges.amazonaws.com/ip-ranges.json` | published by AWS for allowlisting | daily |
| `agentgate:ip:cloud:gcp` | `www.gstatic.com/ipranges/cloud.json` | published by Google | daily |
| `agentgate:ip:cloud:azure` | Service Tags Public JSON (`AzureCloud`), linked from `microsoft.com/en-us/download/details.aspx?id=56519` | Microsoft download terms | daily |
| `agentgate:ip:cloud:oracle` | `docs.oracle.com/en-us/iaas/tools/public_ip_ranges.json` | published by Oracle | daily |
| `agentgate:ip:cloud:cloudflare` | `www.cloudflare.com/ips-v4`, `ips-v6` | published by Cloudflare | daily |
| `agentgate:ip:cloud:digitalocean` | `digitalocean.com/geo/google.csv` (geofeed) | published by DigitalOcean | daily |
| `agentgate:ip:tor` | `check.torproject.org/torbulkexitlist` | Tor Project | hourly |
| `agentgate:bot:verified:ip` and friends | vendor crawler range files (below) | published by each vendor for verification | daily |
| `ip_reputation`, `agentgate:ip:reputation:high` | this process's own decisions | — | live |

## Refresh and caching

Every feed loads its cached copy from `$STATE_DIR/intel/` (default
`./state/intel`) in the background at startup, so startup never waits; then
it refreshes off the request path when the copy is older than its interval.
A download is size-limited and parsed completely (with a minimum plausible
size, so a truncated file is rejected) before it replaces the old data with
an atomic swap; on any failure the previous data stays and the feed retries
in 30 minutes. The cache is written with write-to-temp and rename.
`AGENTGATE_OFFLINE=1` disables every network fetch and reverse DNS lookup
(cached copies still load); the test suite never touches the network.
Lookups are in-memory: ASN is a binary search over sorted ranges (the full
dataset parses in ~0.3 s, ~15 MB, ~30 ns per lookup), prefixes use one map
probe per prefix length present.

## Verified crawlers by IP

A User-Agent that claims a catalogued crawler is
checked against its operator's published ranges: Google
(`developers.google.com/static/crawling/ipranges/` `common-crawlers.json` →
Googlebot, `special-crawlers.json` → AdsBot/Mediapartners/APIs-Google/
Google-Safety, `user-triggered-fetchers.json` and
`user-triggered-fetchers-google.json` → Google-Read-Aloud and Gemini
Notebook, `user-triggered-agents.json` → Google-Agent), Bing
(`www.bing.com/toolbox/bingbot.json`), OpenAI (`openai.com/gptbot.json`,
`chatgpt-user.json`, `searchbot.json`, `adsbot.json`), Perplexity
(`www.perplexity.com/perplexitybot.json`, `perplexity-user.json`), Apple
(`search.developer.apple.com/applebot.json`) and Anthropic
(`claude.com/crawling/bots.json`, one list for ClaudeBot, Claude-User and
Claude-SearchBot). Each file proves only the bots it is published for: a
ChatGPT-User address does not verify a GPTBot claim. Googlebot
(`googlebot.com`, `google.com`), Applebot (`applebot.apple.com`) and Bingbot
(`search.msn.com`) fall back to forward-confirmed reverse DNS: the PTR name
must be under those domains and resolve back to the address. The lookup runs
in background workers (2 s timeout, results cached 6 h, failures 1 h, bounded
cache), never on the request path: the first request from an unknown address
is labelled `agentgate:bot:ip_pending`. Labels:
`agentgate:bot:verified:ip` (and `agentgate:bot:verified`, replacing
`agentgate:bot:unverified`; plus `agentgate:bot:verified:rdns` for the DNS
path), `agentgate:bot:ip_mismatch` (the claim fails a published check),
`agentgate:bot:ip_unchecked` (no data loaded yet). A bot already proven by a
Web Bot Auth signature is never labelled `ip_mismatch`. Bots whose operators
publish neither ranges nor an rDNS rule (DuckDuckBot, CCBot, ...) get no IP
label. The committed files in `testdata/verifiedbots/` are the real vendor
files of 2026-09-24. Bing's verification page renders client-side and was not
re-read; its JSON URL and `search.msn.com` rule are long-standing but were not
re-checked against current docs.

## IP reputation

Terminating decisions are remembered per address and per /24 (/48 for
IPv6) with a 30-minute half-life: block, drop and a failed visible check
count +1, a challenge +0.5. The score is
`1 - exp(-(ip + 0.25 × network) / 4)`: about 0.27 after one block and 0.8
(the `agentgate:ip:reputation:high` label) after about six recent blocks
from the address. Memory is process-local and bounded (100,000 addresses,
50,000 networks) and kept **per site**: one customer's blocks never raise
an address's score on another customer's site. Every decision made through
the rule engine feeds it. The score is also a bot-model feature
(`ip_reputation`, weight 0 until a model is trained on it).

## Not covered

Commercial VPNs, residential and mobile proxy networks, open
proxies and smaller hosting providers: there is no reliable free source.
Cloud ranges catch VPN exits hosted in the big clouds only; ASN rules can
target specific hosting ASNs by hand. The country is the AS registration's
country from iptoasn, not a geolocation of the address. Reputation is local
to one process and learns only from the site's own traffic.
