The short answer
It does not help you rank, and it was never proposed as something that would. Google's guidance on AI features and your website tells site owners directly: “You don't need to create new machine readable files, AI text files, or markup to appear in these features.” The same page adds that there is no special schema.org structured data you need to add. If your reason for considering llms.txt is that it might improve your position in Google's AI Overviews or AI Mode, the platform has answered that question and the answer is no.
That is not the same as saying the file is worthless. The proposal describes a genuine and much narrower use case: giving an agent that is already interested in your site a concise, expert-level map of it, available on demand. Whether you have that use case is a real question with a real answer, and it depends on what your site is. A documentation-heavy developer product and a twelve-page consultancy site are not in the same position.
The confusion is worth untangling, because it is a small instance of a pattern that costs marketing teams a great deal of money: a specific technical proposal gets absorbed into the discourse as a general-purpose visibility trick, the original claim gets inflated in the retelling, and by the time it reaches a planning meeting it has become a task on a checklist with no stated purpose. Reading what the proposal says, and what platforms say about it, takes about ten minutes and settles the matter.
References: Google Search Central — AI features and your websitellms.txt — a proposal to standardize LLM-friendly content
What the proposal actually says
The llms.txt proposal was published by Jeremy Howard in September 2024 and has been revised since. Its own summary is modest: “We propose adding a /llms.txt markdown file to websites to provide LLM-friendly content.” The reasoning is about context windows. A language model working on a question has a limited amount of room, and an ordinary marketing site is a poor way to spend it: navigation, scripts, cookie banners and repeated boilerplate arrive alongside the three paragraphs that actually matter.
The file's structure follows from that. It asks for an H1 with the project or site name, which is the only required element, a blockquote with a brief summary, optional prose sections, and H2-delimited lists of links with short descriptions of what each one contains. It is, in effect, a curated table of contents written for a reader who will only look at part of it and who benefits from being told in advance which part to read.
The proposal's own framing of when this gets used is the sentence most summaries omit: agents “are best served by concise, expert-level information gathered in a single, accessible location”, and the information is “used on demand, when an agent needs information”. On demand. Not crawled on a schedule, not submitted anywhere, not evaluated as a quality signal. Something already looking at your site fetches a file that helps it find what it came for.
Read that way, llms.txt is closer to a well-written README than to a piece of SEO infrastructure. That comparison also predicts who benefits: the same kinds of sites where a good README changes the experience.
References: llms.txt — a proposal to standardize LLM-friendly content
Why the platform statement and the proposal do not contradict each other
A lot of the online argument about this file treats two statements as a dispute. They are not in conflict. Google is answering a question about its own ranking and eligibility systems, and saying that adding a file does not create a path into them. The proposal is answering a question about how an agent that has already arrived can orient itself efficiently. Both can be true because they are about different moments.
It is also worth being precise about what Google's position covers. It covers Google. It is not a statement about every assistant, every agent framework or every internal tool your customers might build. An organization writing its own retrieval pipeline over a vendor's documentation is free to look for a file like this and many do, because it is the cheapest available map. That is a private integration decision, not a search-engine behavior, and no search engine's guidance governs it.
The structural argument against it as a ranking signal is straightforward and worth understanding rather than memorizing. A file that a publisher writes about itself cannot distinguish between publishers, because every publisher would write that they are the best source. Self-reported manifests are useful for navigation and useless for adjudication. This is the same reason meta keywords stopped mattering a long time ago, and the reason structured data has to describe content a visitor can actually see.
References: Google Search Central — AI features and your websitellms.txt — a proposal to standardize LLM-friendly content
It is not robots.txt and it is not a sitemap
The proposal distinguishes itself from both neighbors explicitly, and the distinction is load-bearing. Of robots.txt it says that the file “lets automated tools know what access to a site is considered acceptable, such as for search indexing bots”, whereas llms.txt information “is instead used on demand, when an agent needs information”. One governs permission, the other offers orientation. Publishing an llms.txt grants nothing and forbids nothing.
This matters because teams occasionally reach for llms.txt when what they actually need is a crawler decision. If your concern is whether your content may be used for model training, or whether an assistant's search index may include you, those are robots.txt questions answered by naming specific user agents, and they have documented consequences. A summary file has no bearing on either.
Of sitemap.xml the proposal notes that a sitemap is not a substitute, because it generally covers a set of documents that in aggregate will be too large to fit in a context window. That is a fair description of the difference. A sitemap is exhaustive and machine-oriented, built so a crawler can discover every URL. An llms.txt is selective and editorial, built so a reader with limited attention can find the five things worth reading. The selection is the value, which means an llms.txt that lists every page has thrown away the only thing it had.
References: llms.txt — a proposal to standardize LLM-friendly contentGoogle Search Central — Introduction to robots.txt
Where it plausibly earns its place
There is a shape of site where this file does useful work. It has a lot of pages, the pages are genuinely technical, and people regularly point assistants at them to answer questions. Software documentation is the obvious example: an agent asked how to configure something benefits enormously from a curated index that says which page covers configuration, which covers migration and which is a changelog. API references, developer platforms and large knowledge bases fall in the same category.
A second case is internal. If your own team is building retrieval over your own content, the discipline of writing an accurate index of what exists and what each part is for pays off regardless of whether any external system reads the file. Several of the most useful outcomes we have seen from this exercise had nothing to do with AI: the act of writing one paragraph describing each important page surfaced duplicates, outdated pages and sections nobody could explain the purpose of.
A third case is genuine curiosity, honestly labeled. Publishing the file, recording the date, and watching your logs for requests to it is a small, cheap experiment. That is a legitimate thing to do as long as the result is reported as an observation rather than converted into a success story. If nothing ever requests it, that is a finding.
The case where it does not earn its place is the common one: a small commercial site with fifteen pages, no documentation, and a marketing team adding the file because a competitor has one. Fifteen pages fit in a context window. There is nothing to navigate. The file will be created once, drift out of date within a quarter, and contribute nothing except a maintenance obligation nobody assigned.
References: llms.txt — a proposal to standardize LLM-friendly contentOpenAI — Overview of OpenAI Crawlers
The cost nobody budgets
Creating the file is nearly free, which is why the decision gets made carelessly. The cost is not creation, it is truth maintenance. An llms.txt is a set of factual claims about your own business, written in the format most likely to be read in full and quoted confidently. Every claim in it has to stay accurate.
Consider what typically goes in one: what the company does, what it offers, which pages explain each offer, how to get in touch. Now consider what changes in a year at a services company. Offers get renamed. Pages move. A service is retired. Pricing language changes. Contact routing changes. Each of those is a line in a file that no one has looked at since it was created, and each one is now a confident wrong answer sitting at a well-known address, formatted for extraction.
This is materially worse than a stale page buried in a blog archive, for a reason specific to the format. The whole design goal is that a reader takes this file as an authoritative summary and does not go looking further. A stale page competes with other pages. A stale llms.txt is presented as the map.
So the honest version of the decision is not “should we add a file”. It is “who owns this file, what triggers a review of it, and what happens when the person who wrote it leaves”. If the answer is that nobody owns it, do not publish it. An absent file is neutral. A wrong one is not.
References: llms.txt — a proposal to standardize LLM-friendly content
How to decide, in four questions
These questions are ordered so the first no ends the exercise. That is deliberate. Most sites should stop at the first or second.
- Do you have more content than fits comfortably in a context window, organized in a way that genuinely needs a guide? If a capable reader could understand your whole site in ten minutes, an index of it adds nothing.
- Is there a specific audience that points assistants or agents at your content today? Developers reading documentation qualify. A hypothetical future user does not.
- Can you name the person who will review the file when an offer, a page or a contact route changes, and the event that triggers that review? If not, stop here.
- Are you willing to report the outcome honestly, including “nothing requested it”? If the answer is only acceptable when it is positive, this is not an experiment, and you will end up defending a file instead of evaluating it.
References: llms.txt — a proposal to standardize LLM-friendly content
If you publish one, publish it honestly
Assuming you got through those questions with a yes, the format is not the hard part. Name the site, summarize what it is in a sentence a stranger would recognize as true, then list the pages that matter with an honest description of each. Say what a page covers, not how good it is. Descriptions that read as sales copy are both less useful and more obviously self-serving, and the whole premise of the file is that someone is trusting your description instead of reading the page.
Keep it consistent with what a visitor can see. This is the same principle that governs structured data: a summary that describes services you do not offer, scale you do not have or evidence you cannot produce is not a discovery aid, it is a misleading claim in a machine-readable wrapper. The fact that the audience may be a model rather than a person does not change what is true.
Include what you do not do, where it is genuinely useful. A file that says which problems this company does not take on saves everyone time, and it is the kind of statement that reads as credible precisely because nothing else in the format encourages it.
Then put the review on a schedule, alongside whatever cadence you already use for checking crawler directives and metadata. If you are maintaining one machine-readable file thoughtfully, the marginal cost of checking the rest at the same time is small, and that combined review is worth considerably more than any single file in it.
References: llms.txt — a proposal to standardize LLM-friendly contentGoogle Search Central — AI features and your website
What we decided for this site, and why
It would be odd to write this without saying what we do. Cendar Lab serves an llms.txt at its own root. We are also the people telling you that it will not improve your ranking, so the reasoning is worth stating rather than leaving as an apparent contradiction.
We publish it for three reasons, none of them about rankings. The first is that our own site is the place we test the things we recommend, and a recommendation we have never implemented is a weaker recommendation. The second is that the file is generated from the same content the site is generated from, so it cannot drift independently: when a service is renamed, the summary changes with it, because both read the same source. That removes the maintenance risk described above, and it is the condition we would put on any client doing the same thing. The third is that it is a small, honest experiment, and we would rather run it than have an opinion about it.
What we do not do is treat its existence as an achievement. It is not listed as a deliverable, it does not appear in a proposal as a reason to hire us, and if it turns out that nothing ever requests it, we will say so here rather than quietly leaving the file up as decoration.
The generated-from-source approach is the part worth copying. If you are going to publish a machine-readable summary of your business, derive it from the same data that renders your pages. A hand-written file is a second source of truth, and a second source of truth is a future contradiction.
References: llms.txt — a proposal to standardize LLM-friendly content
What would change our mind
A position worth holding should come with the conditions that would overturn it, otherwise it is just a preference. Here are ours.
If a major assistant documented that it reads llms.txt as part of building its search index, and said what it does with the contents, that would move the file from optional orientation to genuine discovery infrastructure, and we would say so. Nothing in the current documentation from the vendors whose crawlers we have reviewed says this. What they document is user agents, crawl behavior and the consequences of blocking them.
If we observed sustained, verified requests for the file from identifiable agents, across a range of sites rather than one, that would be evidence of a real retrieval path even without a vendor statement. Server logs can show this, and the honest version of that observation requires the same discipline as any other measurement: a long enough window, verification that the requester is who it claims to be, and a note of what could not be established.
What would not change our mind is more articles asserting that the file is now essential. The volume of confident writing about a technique is not evidence about the technique. That is worth remembering more broadly: a large share of what circulates about AI search consists of restatements of other restatements, and the original claim, when you trace it back, is frequently either a vendor's documentation saying something narrower or a blog post citing another blog post.
Tracing a claim to its source is slower than repeating it and it is the only part of this work that compounds. When we publish something here, the sources are listed with the date we read them, so that you can check whether the ground has moved since.
References: llms.txt — a proposal to standardize LLM-friendly contentGoogle Search Central — AI features and your websiteOpenAI — Overview of OpenAI Crawlers
What actually moves the needle instead
If the underlying goal was visibility inside AI answers, the work is less novel than a new file and considerably more effective. Make sure the crawlers that build assistant search indexes are allowed to reach you, and that you decided that deliberately rather than inheriting it. Serve your content in HTML that is present in the response rather than assembled afterwards. Answer the actual question near the top of the page, in language that survives being extracted as a single paragraph.
Make your claims checkable. An assistant summarizing your page cannot verify an adjective, but it can carry a specific, sourced statement into an answer without distorting it. Keep the facts about your organization consistent wherever they appear, so that a system assembling a picture of you from several places finds one company rather than three. Give a reader a reason to trust the page that does not depend on trusting you.
None of this is new, which is the point. The arrival of generative answers changed the interface and left the foundations largely intact: reachable, understandable, trustworthy content about questions someone is actually asking. A text file at your site root does not substitute for any of that, and the proposal's own author never suggested it would.
References: Google Search Central — AI features and your websiteOpenAI — Overview of OpenAI Crawlers