Blog · How AI models learn about a new brand
How AI Models Learn About a New Brand and How Long It Takes
Priya Bothra · October 4, 2026
AI models do not learn about a new brand the way a person does. A released model's internal knowledge is frozen at its training cutoff, so a brand launched after that date is unknown to the model's memory. New brands become visible to AI through four separate paths: live retrieval from the web at answer time, entity databases such as knowledge graphs, inclusion in the training data of future models, and context a user supplies during a conversation.
Each path runs on a different clock. Retrieval can pick up a new brand within days or weeks of its pages being crawled and indexed, while training data typically takes many months to flow into a released model. This guide is a narrow deep-dive into that timeline for newly launched or recently renamed companies. For the broader question of how to help AI systems describe an established company correctly, see our guide on how to make AI engines understand your company.
Why "learn" is the wrong word for most of what happens
Learning, in the technical sense, happens only during training, when a model adjusts its parameters based on a large corpus of text. Once a model is released, those parameters do not change as it talks to users. A model does not absorb your press release the day it goes live, and it does not remember a correct answer a previous user gave it.
What people usually mean by "the AI learned about us" is one of two different events:
- The AI system found you. A search component retrieved a page about your brand and handed it to the model while it was writing an answer. The model itself is unchanged.
- A newer model was trained on text that mentions you. A future model version included pages about your brand in its training data, so it can describe you without searching.
Keeping these two events apart matters because they respond to different actions and have very different timelines. A brand can be well represented in retrieval-based answers while being completely absent from a model's memory, and the reverse can also be true for a brand that has since rebranded.
Path 1: Live retrieval, the fastest route for a new brand
Live retrieval is the process in which an AI system searches the web or its own index at answer time and gives the results to the model. Outside of a user pasting information into a chat, it is the main path that can make a brand launched last month appear in an AI answer this month.
Every major assistant now retrieves:
- ChatGPT search relies on OpenAI's OAI-SearchBot. OpenAI's crawler documentation states that "sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers." OpenAI's ChatGPT search help page notes that allowing the crawler makes a site eligible but does not guarantee placement.
- Claude uses Claude-SearchBot for search indexing, separate from ClaudeBot for training, and all of Anthropic's bots respect robots.txt, according to Search Engine Land.
- Perplexity runs PerplexityBot to surface and link websites in answers.
- Google AI Overviews and AI Mode draw on Google's index and use a "query fan-out" technique that runs multiple related searches across subtopics, as described in Google's AI features guidance. Google says there are no additional requirements beyond normal search eligibility.
For a new brand, retrieval has one important weakness: the system has to decide to search, and then it has to find you among the results. When a user asks a broad category question such as "What tools help agencies track time?", the retrieval layer pulls pages that already rank or are frequently cited for that topic. A brand with three pages on its own site and no third-party coverage has little chance of appearing, even when it is fully indexed. Retrieval gets a new brand into the candidate pool. It does not by itself make the brand a likely pick. A branded question such as "What is [BrandName]?" will often trigger a search that finds your homepage, while a category question usually will not. For a technical look at the pipeline, see how large language models retrieve information.
Path 2: Entity databases and knowledge graphs
An entity database is a structured store of facts about people, places, organizations and things. Google's Knowledge Graph is the best-known example, and it feeds knowledge panels and other search features that AI answers can draw on.
Google explains on its Knowledge Graph help page that "facts in the Knowledge Graph come from a variety of sources that compile factual information," and that Google also receives "factual information directly from content owners." The process is automated, so a new company is not added because it asks. It is recognized once enough consistent information about it exists across sources Google trusts.
For a new brand, the entity path matters in two ways. First, it helps AI systems tell your company apart from other things with the same or similar names, which is a real problem for startups with generic or reused names. Second, entity recognition tends to be stable once it exists, so it provides a steady baseline that retrieval results can build on. Public reference sites also play a role here. Wikipedia, for instance, is highly cited by ChatGPT: Profound's analysis of 680 million citations found it was ChatGPT's top source at 7.8% of citations. However, Wikipedia's notability guideline for organizations and companies requires significant coverage in independent, reliable sources, which most newly launched companies do not yet have. Creating a promotional article too early usually leads to deletion.
To go deeper on this layer, see our guide on how to build entity authority for your brand.
Path 3: Training data for future models, the slowest route
Training data is the text a model learns from before release, and it is the only path that places a brand inside the model's own memory. For a new brand, this path is slow for two reasons: the cutoff lag and the long-tail problem.
The cutoff lag
Every model has a date after which it has little or no knowledge. Anthropic publishes a "reliable knowledge cutoff" for each Claude model in its models overview. As one example, Claude Haiku 4.5, whose API model identifier is dated October 2025, lists a reliable knowledge cutoff of February 2025. Other providers publish similar cutoffs for their models. The gap between the cutoff and the release date varies by model, but it means a brand launched today will not appear in the memory of any model currently on the market, and it will only appear in future models whose training data was collected after the brand had built a visible footprint.
In practice, a new brand should plan for its first mention in model memory to arrive months after launch at the earliest, and only if enough text about it existed before a given model's cutoff. Nobody outside the AI labs can predict the exact timing, and any vendor promising a specific date is guessing.
The long-tail problem
Even when a brand is in the training data, a model may not reliably learn it. Research by Kandpal and colleagues, Large Language Models Struggle to Learn Long-Tail Knowledge (ICML 2023), found that a model's ability to answer a factual question "relates to how many documents associated with that question were seen during pre-training," with both correlational and causal evidence. The same paper found that retrieval augmentation can reduce this dependence on pre-training data.
For a new brand, that finding translates into a simple rule: a handful of pages is rarely enough. A model is more likely to retain what your company does if many independent documents describe it consistently. A single launch announcement, however well written, is a weak signal on its own.
Where training text comes from
Public web crawls make up a large share of pre-training data for many models. The original GPT-3 paper reported that a filtered version of Common Crawl made up 60% of its weighted training mix, with the rest coming from WebText2, books and Wikipedia. Current labs disclose less detail about their data mixes, so treat that figure as historical context rather than a description of today's models. Common Crawl's crawler, CCBot, obeys robots.txt, and Anthropic's ClaudeBot and OpenAI's GPTBot are separate training crawlers that sites can allow or block independently of the search crawlers.
This creates a real trade-off for new brands. Blocking training crawlers keeps your content out of future models, which some companies prefer for legal or competitive reasons. It also reduces the chance that future models learn what your company does from your own pages. The decision should be deliberate, not a side effect of an old robots.txt template.
Path 4: Context the user supplies
User-supplied context is information a person gives an AI assistant during a conversation, such as a pasted product description, an uploaded file or a URL the assistant is asked to read. It is the most immediate path, but it applies only to that conversation.
AI providers run separate agents for these user-initiated fetches. OpenAI's ChatGPT-User fetches pages when a user asks, and OpenAI notes that robots.txt rules may not apply in the same way to user-initiated requests. Anthropic runs Claude-User, and Perplexity runs Perplexity-User, which Perplexity's documentation says generally ignores robots.txt because a person requested the fetch.
For a new brand, this path matters most late in the buying process. A prospect who already knows your name may paste your pricing page into an assistant and ask for a comparison. If that page is vague, hard to render or blocked, the assistant has little to work with. The same prospect's next conversation starts from scratch, because the model does not keep what it read.
Timeline summary: when each path picks up a new brand
The table below summarizes typical timing and levers. The timings are general guidance based on how each mechanism works, not measured benchmarks, and they vary by platform and by how much coverage a brand earns.
| Path | What it gives the AI | Typical time to take effect | Main levers for a new brand |
|---|---|---|---|
| Live retrieval | Current pages found at answer time | Days to weeks after crawling and indexing | Crawler access, indexable pages, third-party coverage for category queries |
| Entity databases | Structured facts that distinguish your brand | Weeks to months, once consistent sources exist | Consistent name, description and profiles across trusted sources |
| Training data | Built-in memory in future models | Many months, tied to future model releases | Volume of consistent independent mentions, training crawler policy |
| User context | Whatever the user shares in one session | Immediate, single conversation only | Clear, renderable product and pricing pages |
A launch sequence for new brands: framework
The following framework orders the work by which path it feeds, starting with the fastest. It is a recommendation, not a guarantee of outcomes.
1. Open the right doors before launch
Check robots.txt, CDN and firewall rules for Googlebot, Bingbot, OAI-SearchBot, Claude-SearchBot and PerplexityBot. Decide separately whether to allow training crawlers such as GPTBot, ClaudeBot, CCBot and Google-Extended. New sites often inherit bot-blocking defaults from hosting or security tools, which can silently remove them from AI search.
2. Publish one canonical description of the brand
Write a single, plain-language statement of what the company is, what it sells, who it serves and how it differs. Use it on your homepage, about page, docs, social profiles, directories and press materials. Consistency across sources helps retrieval, entity recognition and future training all at once. Our guide on fixing inconsistent brand information covers how to find and correct conflicting descriptions.
3. Make product facts easy to extract
Publish clear pages on pricing, features, integrations and use cases, with direct answers near the top. These pages serve branded retrieval and user-supplied context, which are the paths most likely to work early.
4. Earn independent mentions in the places AI cites
Category recommendations depend heavily on third-party sources. Ahrefs' study of 75,000 brands found branded web mentions had a 0.664 correlation with appearing in Google AI Overviews, higher than backlinks at 0.218, though the authors stress correlation is not causation. Ahrefs' 2026 benchmark found YouTube mentions were the strongest AI visibility signal among the factors it studied. For a new brand, reviews, integration partner pages, podcasts, community discussions and industry coverage all add to the document count that both retrieval and future training depend on. Press can help too: see how press releases influence AI search.
5. Measure recognition by path, not as one number
Test branded prompts ("What is [Brand]?") and category prompts ("Best tools for X") separately, with web search on and off where the interface allows. Run each prompt many times. SparkToro and Gumshoe research found less than a 1 in 100 chance that an AI tool would return the same brand list twice, so a percentage across repeated runs is more meaningful than a single screenshot.
Common mistakes new brands make
Expecting the model to "know" them after launch. A model released before your launch cannot have learned about you. If an assistant describes you accurately, it almost certainly retrieved the information.
Treating retrieval visibility as memory. Appearing in a search-enabled answer today does not mean the next model version will know you. Those are separate events on separate timelines.
Relying on a single launch announcement. Research on long-tail knowledge suggests models learn facts that appear in many documents. One announcement is a start, not a footprint. Chasing shortcuts. Google says no special AI text files or markup are required for its AI features, and Google's John Mueller described llms.txt as "purely speculative for now" in June 2026. Fundamentals come first.
A hypothetical example
Consider a hypothetical invoicing startup that launches in March. By April, ChatGPT with search answers "What is [Startup]?" correctly by citing its homepage. The founders assume ChatGPT "knows" them. Yet when a buyer asks for "invoicing software for freelance designers," the startup never appears, because the retrieved sources for that question are comparison articles and review sites that have not covered it yet.
A path-by-path review would show strong branded retrieval, weak category retrieval, no entity recognition yet and no presence in model memory. The priorities follow directly: earn coverage in the comparison and review sources that category answers cite, keep the brand description identical everywhere, and treat future training inclusion as a long-term result of that coverage rather than a separate project.
How Bob Builds AI helps
Bob Builds AI is an AEO and GEO platform and agency that measures how AI systems describe and recommend brands, which is useful for tracking a new brand's progress along each path. Visibility Monitoring tracks visibility rate, citation rate, competitor positioning, citation sources and recommendation changes over time across ChatGPT, Gemini, Claude, Perplexity, Copilot and Google AI Overviews and AI Mode, and its documentation says it measures the real chat and search interfaces rather than raw model APIs. Brand Memory keeps one source of truth for products, differentiators and messaging, which supports the consistent description new brands need. Prompt Research surfaces the questions buyers ask AI in your category, and Agent Analytics shows how AI crawlers access your site.
FAQ
How long does it take for ChatGPT to know about a new brand?
It depends on which path you mean. ChatGPT search can find and cite a new brand's pages soon after OAI-SearchBot crawls them, often within days or weeks, if the site is not blocked. Knowledge built into the model itself arrives only when OpenAI releases a model trained on data collected after your brand became visible, which typically takes many months. OpenAI does not publish a timeline for either.
Do AI models learn from conversations with users?
A released model's parameters do not change during conversations, so it does not learn a brand from one user's chat. Some assistants offer memory features that recall details for the same user, but that is separate from the model's general knowledge. Providers have their own policies on whether conversation data may be used to train future models, so check each provider's data settings.
Why does an AI assistant know my brand name but describe it wrongly?
Usually the assistant retrieved an outdated or inaccurate source, or it confused your brand with another entity that has a similar name. It can also be answering from older training data that predates a rebrand or pivot. Checking which sources the answer cites, and whether web search was used, helps identify the cause. Consistent descriptions across your site and third-party profiles reduce the problem over time.
Can I submit my brand directly to an AI model?
There is no public submission form that adds a brand to a model's memory. The practical routes are allowing AI search crawlers, publishing clear pages about your company and earning independent coverage. Google says its Knowledge Graph receives some information directly from content owners, but inclusion is automated and depends on consistent information across trusted sources.
Does blocking training crawlers hurt a new brand's AI visibility?
Blocking training crawlers such as GPTBot, ClaudeBot or CCBot does not block AI search crawlers like OAI-SearchBot or Claude-SearchBot, so retrieval-based visibility can continue. It does reduce the chance that future models learn about your brand from your own pages. Many brands will still appear in training data through third-party coverage. The decision is a business and legal trade-off.
Why do AI assistants mention established competitors instead of my new brand?
Established competitors have more independent documents describing them, and research on long-tail knowledge shows models answer more reliably about entities that appear in many pre-training documents. Retrieval also favors pages that already rank and are frequently cited for a category. A new brand closes this gap mainly by earning mentions in the review sites, publications and communities AI answers cite.
Do I need a Wikipedia page for AI models to learn about my brand?
No. Wikipedia is heavily cited by ChatGPT, but most new companies do not meet its notability guideline, which requires significant coverage in independent reliable sources. Creating a promotional page too early often leads to deletion. Focus first on consistent profiles, clear product pages and independent coverage, which are also what a future Wikipedia article would need.
How can I tell whether an AI answer about my brand came from memory or search?
Look for citations or a visible search step, which indicate retrieval. Where an interface lets you turn web search off, compare answers with it on and off. An answer without citations that includes outdated details likely reflects training data. Run each prompt several times, because AI answers vary noticeably between runs.
Conclusion
AI models do not learn about a new brand in one moment. Retrieval can surface a new company within weeks, entity databases recognize it once consistent facts exist across trusted sources, future models absorb it only after it has built a broad footprint, and user-supplied context helps only within a single conversation.
The practical implication for founders and marketing teams is to manage each path on its own timeline. Open access for AI search crawlers now, publish one consistent brand description, make product facts easy to extract, and invest steadily in independent coverage that feeds both today's retrieval and tomorrow's training data. Then measure branded and category prompts separately, across many runs.
A useful next step is to run ten branded and ten category prompts across three assistants and note which path each answer relies on. If you want to track that progress systematically across models, Bob Builds AI can help you set up the monitoring and decide which gaps to close first.