Threat intelligence

The public IP data behind AgentGate's country, ASN, cloud, Tor and verified-crawler labels, how it refreshes, and what it does not cover.

Data sources

AgentGate downloads public, provider-published data and labels requests with it. There is no CDN-scale telemetry behind these labels: they are exactly as good as the files below.

Label / scoreSource (URL checked live 2026-09-24)Licence / termsRefresh
agentgate:ip:asn:<n>, agentgate:ip:country:<cc>; asn, country statementsiptoasn.com ip2asn-combined.tsv.gzPublic domain (PDDL v1.0, stated on iptoasn.com)daily
agentgate:ip:cloud:awsip-ranges.amazonaws.com/ip-ranges.jsonpublished by AWS for allowlistingdaily
agentgate:ip:cloud:gcpwww.gstatic.com/ipranges/cloud.jsonpublished by Googledaily
agentgate:ip:cloud:azureService Tags Public JSON (AzureCloud), linked from microsoft.com/en-us/download/details.aspx?id=56519Microsoft download termsdaily
agentgate:ip:cloud:oracledocs.oracle.com/en-us/iaas/tools/public_ip_ranges.jsonpublished by Oracledaily
agentgate:ip:cloud:cloudflarewww.cloudflare.com/ips-v4, ips-v6published by Cloudflaredaily
agentgate:ip:cloud:digitaloceandigitalocean.com/geo/google.csv (geofeed)published by DigitalOceandaily
agentgate:ip:torcheck.torproject.org/torbulkexitlistTor Projecthourly
agentgate:bot:verified:ip and friendsvendor crawler range files (below)published by each vendor for verificationdaily
ip_reputation, agentgate:ip:reputation:highthis process's own decisions—live

Refresh and caching

Every feed loads its cached copy from $STATE_DIR/intel/ (default ./state/intel) in the background at startup, so startup never waits; then it refreshes off the request path when the copy is older than its interval. A download is size-limited and parsed completely (with a minimum plausible size, so a truncated file is rejected) before it replaces the old data with an atomic swap; on any failure the previous data stays and the feed retries in 30 minutes. The cache is written with write-to-temp and rename. AGENTGATE_OFFLINE=1 disables every network fetch and reverse DNS lookup (cached copies still load); the test suite never touches the network. Lookups are in-memory: ASN is a binary search over sorted ranges (the full dataset parses in ~0.3 s, ~15 MB, ~30 ns per lookup), prefixes use one map probe per prefix length present.

Verified crawlers by IP

A User-Agent that claims a catalogued crawler is checked against its operator's published ranges: Google (developers.google.com/static/crawling/ipranges/ common-crawlers.json → Googlebot, special-crawlers.json → AdsBot/Mediapartners/APIs-Google/ Google-Safety, user-triggered-fetchers.json and user-triggered-fetchers-google.json → Google-Read-Aloud and Gemini Notebook, user-triggered-agents.json → Google-Agent), Bing (www.bing.com/toolbox/bingbot.json), OpenAI (openai.com/gptbot.json, chatgpt-user.json, searchbot.json, adsbot.json), Perplexity (www.perplexity.com/perplexitybot.json, perplexity-user.json), Apple (search.developer.apple.com/applebot.json) and Anthropic (claude.com/crawling/bots.json, one list for ClaudeBot, Claude-User and Claude-SearchBot). Each file proves only the bots it is published for: a ChatGPT-User address does not verify a GPTBot claim. Googlebot (googlebot.com, google.com), Applebot (applebot.apple.com) and Bingbot (search.msn.com) fall back to forward-confirmed reverse DNS: the PTR name must be under those domains and resolve back to the address. The lookup runs in background workers (2 s timeout, results cached 6 h, failures 1 h, bounded cache), never on the request path: the first request from an unknown address is labelled agentgate:bot:ip_pending. Labels: agentgate:bot:verified:ip (and agentgate:bot:verified, replacing agentgate:bot:unverified; plus agentgate:bot:verified:rdns for the DNS path), agentgate:bot:ip_mismatch (the claim fails a published check), agentgate:bot:ip_unchecked (no data loaded yet). A bot already proven by a Web Bot Auth signature is never labelled ip_mismatch. Bots whose operators publish neither ranges nor an rDNS rule (DuckDuckBot, CCBot, ...) get no IP label. The committed files in testdata/verifiedbots/ are the real vendor files of 2026-09-24. Bing's verification page renders client-side and was not re-read; its JSON URL and search.msn.com rule are long-standing but were not re-checked against current docs.

IP reputation

Terminating decisions are remembered per address and per /24 (/48 for IPv6) with a 30-minute half-life: block, drop and a failed visible check count +1, a challenge +0.5. The score is 1 - exp(-(ip + 0.25 × network) / 4): about 0.27 after one block and 0.8 (the agentgate:ip:reputation:high label) after about six recent blocks from the address. Memory is process-local and bounded (100,000 addresses, 50,000 networks) and kept per site: one customer's blocks never raise an address's score on another customer's site. Every decision made through the rule engine feeds it. The score is also a bot-model feature (ip_reputation, weight 0 until a model is trained on it).

Not covered

Commercial VPNs, residential and mobile proxy networks, open proxies and smaller hosting providers: there is no reliable free source. Cloud ranges catch VPN exits hosted in the big clouds only; ASN rules can target specific hosting ASNs by hand. The country is the AS registration's country from iptoasn, not a geolocation of the address. Reputation is local to one process and learns only from the site's own traffic.

View as Markdown