Threat intelligence
The public IP data behind AgentGate's country, ASN, cloud, Tor and verified-crawler labels, how it refreshes, and what it does not cover.
Data sources
AgentGate downloads public, provider-published data and labels requests with it. There is no CDN-scale telemetry behind these labels: they are exactly as good as the files below.
| Label / score | Source (URL checked live 2026-09-24) | Licence / terms | Refresh |
|---|---|---|---|
agentgate:ip:asn:<n>, agentgate:ip:country:<cc>; asn, country statements | iptoasn.com ip2asn-combined.tsv.gz | Public domain (PDDL v1.0, stated on iptoasn.com) | daily |
agentgate:ip:cloud:aws | ip-ranges.amazonaws.com/ip-ranges.json | published by AWS for allowlisting | daily |
agentgate:ip:cloud:gcp | www.gstatic.com/ipranges/cloud.json | published by Google | daily |
agentgate:ip:cloud:azure | Service Tags Public JSON (AzureCloud), linked from microsoft.com/en-us/download/details.aspx?id=56519 | Microsoft download terms | daily |
agentgate:ip:cloud:oracle | docs.oracle.com/en-us/iaas/tools/public_ip_ranges.json | published by Oracle | daily |
agentgate:ip:cloud:cloudflare | www.cloudflare.com/ips-v4, ips-v6 | published by Cloudflare | daily |
agentgate:ip:cloud:digitalocean | digitalocean.com/geo/google.csv (geofeed) | published by DigitalOcean | daily |
agentgate:ip:tor | check.torproject.org/torbulkexitlist | Tor Project | hourly |
agentgate:bot:verified:ip and friends | vendor crawler range files (below) | published by each vendor for verification | daily |
ip_reputation, agentgate:ip:reputation:high | this process's own decisions | — | live |
Refresh and caching
Every feed loads its cached copy from $STATE_DIR/intel/ (default ./state/intel) in the background at startup, so startup never waits; then it refreshes off the request path when the copy is older than its interval. A download is size-limited and parsed completely (with a minimum plausible size, so a truncated file is rejected) before it replaces the old data with an atomic swap; on any failure the previous data stays and the feed retries in 30 minutes. The cache is written with write-to-temp and rename. AGENTGATE_OFFLINE=1 disables every network fetch and reverse DNS lookup (cached copies still load); the test suite never touches the network. Lookups are in-memory: ASN is a binary search over sorted ranges (the full dataset parses in ~0.3 s, ~15 MB, ~30 ns per lookup), prefixes use one map probe per prefix length present.
Verified crawlers by IP
A User-Agent that claims a catalogued crawler is checked against its operator's published ranges: Google (developers.google.com/static/crawling/ipranges/ common-crawlers.json → Googlebot, special-crawlers.json → AdsBot/Mediapartners/APIs-Google/ Google-Safety, user-triggered-fetchers.json and user-triggered-fetchers-google.json → Google-Read-Aloud and Gemini Notebook, user-triggered-agents.json → Google-Agent), Bing (www.bing.com/toolbox/bingbot.json), OpenAI (openai.com/gptbot.json, chatgpt-user.json, searchbot.json, adsbot.json), Perplexity (www.perplexity.com/perplexitybot.json, perplexity-user.json), Apple (search.developer.apple.com/applebot.json) and Anthropic (claude.com/crawling/bots.json, one list for ClaudeBot, Claude-User and Claude-SearchBot). Each file proves only the bots it is published for: a ChatGPT-User address does not verify a GPTBot claim. Googlebot (googlebot.com, google.com), Applebot (applebot.apple.com) and Bingbot (search.msn.com) fall back to forward-confirmed reverse DNS: the PTR name must be under those domains and resolve back to the address. The lookup runs in background workers (2 s timeout, results cached 6 h, failures 1 h, bounded cache), never on the request path: the first request from an unknown address is labelled agentgate:bot:ip_pending. Labels: agentgate:bot:verified:ip (and agentgate:bot:verified, replacing agentgate:bot:unverified; plus agentgate:bot:verified:rdns for the DNS path), agentgate:bot:ip_mismatch (the claim fails a published check), agentgate:bot:ip_unchecked (no data loaded yet). A bot already proven by a Web Bot Auth signature is never labelled ip_mismatch. Bots whose operators publish neither ranges nor an rDNS rule (DuckDuckBot, CCBot, ...) get no IP label. The committed files in testdata/verifiedbots/ are the real vendor files of 2026-09-24. Bing's verification page renders client-side and was not re-read; its JSON URL and search.msn.com rule are long-standing but were not re-checked against current docs.
IP reputation
Terminating decisions are remembered per address and per /24 (/48 for IPv6) with a 30-minute half-life: block, drop and a failed visible check count +1, a challenge +0.5. The score is 1 - exp(-(ip + 0.25 × network) / 4): about 0.27 after one block and 0.8 (the agentgate:ip:reputation:high label) after about six recent blocks from the address. Memory is process-local and bounded (100,000 addresses, 50,000 networks) and kept per site: one customer's blocks never raise an address's score on another customer's site. Every decision made through the rule engine feeds it. The score is also a bot-model feature (ip_reputation, weight 0 until a model is trained on it).
Not covered
Commercial VPNs, residential and mobile proxy networks, open proxies and smaller hosting providers: there is no reliable free source. Cloud ranges catch VPN exits hosted in the big clouds only; ASN rules can target specific hosting ASNs by hand. The country is the AS registration's country from iptoasn, not a geolocation of the address. Reputation is local to one process and learns only from the site's own traffic.