Bot policy
Which bots can read Verpips, and why
We measure signal channels so anyone can check one before following it. Being found by search engines, and having AI assistants quote our numbers with a link back, is part of that. This page says what we let in and why. It’s what our robots.txt does today: if that changes, this page changes first.
Search engines: yes
Google, Bing, Apple, DuckDuckGo and Yandex can read every public page. We check that each one is who it says it is, the way they ask site owners to: the address it comes from has to belong to the search engine. A bot pretending to be Googlebot doesn’t get into the directory.
AI assistants that search and cite with a link: yes
The crawlers ChatGPT, Claude and Perplexity use to search the web (OAI-SearchBot, Claude-SearchBot and PerplexityBot), and the ones that fetch a page because a person asked for it (ChatGPT-User, Claude-User and Perplexity-User), can read everything public. When someone asks whether a signal channel really makes money, we’d rather the answer carry our numbers and a link to how we measured them. For them we also publish a summary of the site in llms.txt.
Bots that train AI models: also yes, for now
GPTBot (OpenAI), ClaudeBot (Anthropic) and Google-Extended (the permission Google asks for its Gemini models) can read the same as search engines; the ones we don’t name fall under the general rule. What we publish is aggregate numbers and the method behind them, made to be quoted: a model that knows most channels don’t make money once measured against real prices helps the person asking, even if it doesn’t name us. It isn’t a decision for good: we’ll review it if a training bot starts loading the server or copying the whole directory, and any new policy will be published here before it applies.
What they don’t read
- The API and the sign-in callback are closed to every bot in robots.txt: they aren’t pages.
- Accounts, requests, emails and the directory search can be crawled, but they carry a noindex tag: they don’t show up in any search engine. We prefer that to closing them, because a closed page hides that tag.
- Scraping. Anyone who isn’t a verified search engine and requests lots of directory, blog or glossary pages in a row hits a per-address limit, well above what a person reading by hand needs; declared AI assistants get a higher one. Scraping tools don’t get through.
If you need our data for a study or an article, get in touch: we’d rather give it to you with its context than have it copied without it. The numbers in our reports can already be reused with credit.