Skip to main content

What is a language model's training data?

Training data is the text a language model learned from before release, such as web pages, books and licensed sources. It determines what the model knows about a brand without web search. That knowledge ends at a cutoff date and is only updated with a new model.

Web content reaches training data through crawlers among other routes. OpenAI names GPTBot for this, Anthropic ClaudeBot, and Common Crawl provides an open web archive via CCBot that feeds many training datasets. Google separates use for its AI models from normal search crawling through the control token Google-Extended.

Training data and live retrieval are two different routes into an answer. Excluding a training crawler in robots.txt prevents future use for training, not necessarily retrieval of the page during a web search. Several vendors run separate crawlers for that.

What a model learned about a brand during training cannot be inspected directly. You only see what it answers to questions, and those answers can be outdated or wrong.

What it means for your website

For a website, the decision about training crawlers is a trade-off between control over your content and the chance that future models know the brand. What matters is making it on purpose. Hosting protection rules often block these crawlers without the owner knowing. The free GPTBot check shows robots.txt rules and the real server response.

Related terms

More on deeploupe

Sources

This entry contains no figures and relies on generally documented terms.

Terms help you understand. Whether AI crawlers can reach your site is something you measure.

Check your website for freefree · no signup · no credit card

All terms in the GEO glossary