SEO & AI Search11 min read
What Is llms.txt? The New File Every Website Should Consider
llms.txt is a root-level markdown file that hands AI systems your site's summary instead of making them excavate it. The format, what belongs in it, platform-by-platform setup, and an honest read on what it can and cannot do.
BurTech Solution
Engineering team

llms.txt is a plain-text file that lives at the root of your website — yoursite.com/llms.txt — and gives AI language models a curated, readable summary of what your site is, what it offers, and which pages matter. Think of it as robots.txt’s younger sibling with the opposite job: robots.txt tells crawlers where they may not go; llms.txt tells AI systems what they should understand.
It is an emerging convention rather than an enforced standard — no law compels ChatGPT, Claude or Perplexity to read it — but it costs twenty minutes, carries zero risk, and puts the first description of your business that a machine encounters under your control instead of leaving it to inference. This guide explains where the idea came from, exactly how to write one, what belongs in it (and what does not), and how to think honestly about what it can and cannot do for you.
The problem it exists to solve
When a language model system wants to understand your website, it faces a document built for human eyeballs: headers, mega-menus, cookie banners, chat widgets, footers with forty links, and — somewhere inside the wrapping — the content. Models cope, but coping costs context: the useful tokens (what you sell, who you serve, what things cost) arrive diluted by boilerplate. Multiply that across a site and the machine’s summary of your business is assembled from fragments it happened to parse well.
llms.txt inverts the deal. Instead of making the machine excavate your message, you hand it the message: a single, clean, markdown-formatted document with the signal and none of the chrome. The model that reads it starts from your framing — your one-sentence description, your service list, your canonical facts — rather than reconstructing you from a cookie-consent-wrapped homepage.
The idea was proposed in late 2024 by Jeremy Howard (of fast.ai and Answer.AI) and spread through 2025–2026 the way robots.txt and sitemaps once did: no central authority, just a sensible convention that tools and sites began adopting because it was cheap and useful. Documentation platforms adopted it fastest — developer-tool sites now routinely ship llms.txt and even llms-full.txt (the entire docs in one file) — and business sites followed.
What the format looks like
The convention is deliberately simple markdown. Four building blocks, in order:
- An H1 with your site or business name —
# BurTech Solution. - A blockquote summary — one to three sentences stating what the site is, written the way you would want an AI to repeat it. This is the highest-value real estate in the file.
- Optional free paragraphs for canonical facts — contact routes, pricing shape, service areas, the details you would put on a business card if it were three sentences long.
- H2-grouped link lists — sections like Services, Guides, Company, each entry as
[Page name](/url/): one-line description. The descriptions matter more than the links: they are the machine-readable index of what each page answers.
That is the whole grammar. No XML, no schema validation, no tooling required — a text editor and honesty about what your site contains. Ours is live at /llms.txt if you want a working example against this structure: business summary, key facts including pricing, then sections for services, workflows, portfolio, guides and company pages, each link carrying a one-line description.
What belongs in yours (and what does not)
Include
- The one-sentence answer to “what is this business.” Write it as the citation you want: “Engineering-first digital studio for ecommerce brands …” — subject, category, audience, differentiator.
- Canonical facts a model should never guess: starting prices, service areas, response promises, contact routes, founding facts. Every fact stated here is a fact a model does not have to infer — and inference is where the embarrassing errors come from.
- Your money pages with honest one-liners. Services, pricing, key guides, case studies. The one-liner should say what question the page answers, not repeat its title.
- Your best explanatory content. The guides you would want an AI to draw on when answering questions in your category — which is also a quiet incentive to have guides worth drawing on.
Leave out
- Marketing adjectives. “World-class,” “cutting-edge” and friends are noise to a machine and lower the file’s signal density. State capabilities, not enthusiasm.
- Every page you have. This is a curated map, not a second sitemap — the sitemap already exists for exhaustiveness. Twenty to forty links with descriptions beats two hundred bare URLs.
- Anything you would not want quoted. The file is public and explicitly written for machines to repeat. Draft-quality claims, expiring promotions and internal jargon do not belong.
- Secrets, obviously — but also anything that contradicts the site. A model that finds two different price lists trusts neither.
Setting it up, platform by platform
Any host with file access: create llms.txt, write the markdown, upload to the web root next to robots.txt. Verify it loads as plain text at /llms.txt. Done.
WordPress: either upload the file via your host’s file manager to the site root, or — if you prefer everything managed in one place — a small snippet can serve the content at the /llms.txt route; several SEO plugins now generate one as well. However it is produced, the test is the same: fetch the URL, see clean markdown, no HTML wrapper.
Shopify: the platform does not expose the root for arbitrary files, so use the proxy/redirect pattern — host the file elsewhere and redirect, or serve it via a template route. Imperfect, but the convention tolerates it; what matters is that the canonical URL responds.
Keep it current: the file is only as trustworthy as its freshest fact. Put “update llms.txt” into the same checklist that publishes a new service page or changes pricing — ours updates whenever the site’s content map changes, as part of the publishing routine rather than as a someday task.
What it can and cannot do — the honest version
What it plausibly does: gives crawling AI systems a clean first impression under your control; concentrates your canonical facts where retrieval can find them; and — for the growing set of tools that explicitly fetch it — measurably shapes how your site is summarised. Server logs on sites we manage show AI user agents fetching llms.txt alongside robots.txt; the audience is real.
What it cannot do: force any provider to read it, override what your actual pages say, or substitute for the fundamentals. A site with thin content, broken schema and blocked crawlers gains nothing from a beautiful llms.txt — the file is a map, and a map of an empty city is still empty. It also is not a rankings lever in classic search; Google has not made it a signal, and nobody credible claims otherwise.
The asymmetry argument: twenty minutes of writing, zero maintenance beyond your normal publishing checklist, no way for it to hurt you — against a plausible improvement in how the fastest-growing discovery surface describes your business. Files with that risk profile do not need a guarantee to justify themselves. This is the same reasoning we laid out in our SEO/AEO/GEO guide: show up early on cheap conventions while competitors wait for proof.
A worked example: writing the file for a fictional business
Templates teach less than watching one get written, so here is the thinking for “Harbourline Physio,” a two-clinic physiotherapy practice.
The H1 and blockquote do the positioning work: # Harbourline Physio followed by > Physiotherapy clinics in Halifax and Dartmouth specialising in post-surgical rehabilitation and sports injury recovery. Direct billing to major insurers; same-week initial assessments. Notice what those two sentences encode: category, locations, two specialisms, and the two facts patients actually ask about (billing and wait time). A model answering “physio near me that does direct billing” now has its answer in one retrieval.
The facts paragraph carries what would otherwise be guessed: Locations: 12 Water St Halifax; 88 Portland St Dartmouth. Hours: Mon–Sat. Initial assessment $110, follow-ups $85. Book online or call. Plain, quotable, unambiguous.
The sections then map the site: a Services section (post-surgical rehab, sports injuries, concussion care — each with a one-liner saying who it is for), a Guides section (the practice’s articles on recovery timelines and exercise progressions — the content a model would cite when answering patient questions), and a Company section (about, team, contact). Twenty-five lines total. Any AI system that reads it can now describe Harbourline accurately, recommend it for the right queries, and quote its prices correctly — which is the entire job.
llms.txt vs robots.txt vs sitemap.xml
| File | Audience | Job | Tone |
|---|---|---|---|
| robots.txt | All crawlers | Permission: where bots may and may not go | Rules |
| sitemap.xml | Search crawlers | Exhaustive inventory of URLs for indexing | List |
| llms.txt | AI/language models | Curated understanding: what matters and what it means | Summary |
They cooperate rather than compete: robots.txt decides access (including whether GPTBot, ClaudeBot and PerplexityBot may crawl at all — a separate decision we covered in the GEO guide), the sitemap ensures completeness, and llms.txt supplies comprehension. A site serious about AI-era discovery ships all three, plus the structured data that does the page-level equivalent of llms.txt’s site-level job.
The deeper shift this file represents
Step back from the mechanics and llms.txt marks a real change in what “being findable” means. For twenty-five years, websites optimised for a reader who clicks: the search engine’s job ended at delivering a visitor, and the page’s job began there. AI systems break that handoff — increasingly, the machine reads on the visitor’s behalf, synthesises an answer, and the visitor may never load your page at all. In that world, the question “what does the machine believe about us?” becomes as commercially important as “where do we rank?”
Most businesses have no answer to the first question because they have never asked it. The machine’s beliefs were assembled from whatever it parsed: an outdated directory listing, a competitor’s comparison page, your homepage minus everything the cookie banner obscured. llms.txt is the first convention that lets a site owner answer the question deliberately — here is what we are, in words we chose, in a format built for you. It will not be the last such convention; structured self-description for machines is a direction, not a file. Getting comfortable with the practice now — maintaining one small, honest, machine-first document — is how a small business builds the muscle that the next decade of discovery will keep exercising.
There is also a competitive-timing argument that is easy to underrate. Conventions like this have adoption curves: early on, presence itself differentiates — the assistant that finds llms.txt on your site and nothing on your competitor’s has one clean source and one guess. By the time adoption is universal, the advantage evaporates and the file is merely table stakes. Sitemaps followed exactly this arc. The window where twenty minutes of markdown constitutes an edge is open now; it will not reopen.
Common mistakes we see in the wild
- The dumped sitemap: three hundred bare URLs, no descriptions. Machines get no meaning; the curation was the point.
- The press release: four paragraphs of brand story, zero facts. Models quote facts; they compress prose.
- The stale file: last year’s prices, a discontinued service. Worse than absence, because it competes with the live site and loses trust for both.
- The HTML response: a “pretty” llms.txt rendered through the theme with header and footer. The entire value is cleanliness; serve text/plain markdown.
- The contradiction: a file saying “from $500” while the pricing page says “from $750.” Pick one truth and enforce it everywhere — entity consistency is the deeper game this file is one move in.
Beyond the basics: three refinements once the file exists
Write descriptions as answers, not labels. Compare “[Pricing](/services/): our pricing page” with “[Pricing](/services/): starting prices for every service; fixed quotes within one business day.” The second version lets a model answer a pricing question without another fetch — and models, like people, prefer sources that save them work. Audit your one-liners with a simple test: could each line, quoted alone, be a useful sentence in an AI’s answer?
Order sections by what you want to be known for. Position in the file is a soft priority signal. If AI-era customers should think of you first for automation and second for design, your sections should run that way — not alphabetically, and not in whatever order the pages were built.
Version it like code. Keep the file in whatever change-tracking you have — even a dated copy in a folder — so you can see what the machines were told and when. When an assistant repeats an outdated fact months from now, knowing what your file said at the time turns a mystery into a diagnosis. Small discipline, disproportionate debugging value.
How to tell if it is working
Direct attribution is genuinely hard — no provider reports “we read your llms.txt.” The practical signals, in rising order of effort: server logs showing AI user agents fetching the file (grep your access log for llms.txt — the visits are usually there within weeks); assistant spot-checks — ask ChatGPT, Perplexity and friends what your business does and watch whether their phrasing drifts toward your blockquote’s phrasing over a quarter; and AI-referral traffic in analytics as a lagging composite of all your AEO work. Treat the file as one instrument in the orchestra and measure the orchestra.
How this fits the rest of your AI-visibility stack
A useful mental model: machines learn about your business at three altitudes, and each has its own instrument. At the site level, llms.txt supplies the summary — who you are, what matters, where to look. At the page level, structured data (Organization, Service, Product, FAQPage) states what each page means in a vocabulary machines share. At the passage level, answer-first writing — headings phrased as questions, resolved immediately in two liftable sentences — gives engines the actual quotable material. Miss the bottom layers and the top one is a lobby with no building; do all three and every altitude reinforces the others: the file points to pages whose schema confirms what their passages deliver.
Practically, that means the twenty-minute file usually belongs at the end of an AEO sprint, not the start: fix crawlability, ship the schema, rewrite the money pages answer-first — then write the llms.txt that proudly maps what now exists. Written in that order, the file is a table of contents for genuine substance. Written first, it risks being the most articulate description of pages that cannot back it up — and machines, increasingly, check.
A final calibration note: treat everything about AI-provider behaviour in this space as dated the day it is written — which providers fetch the file, how they weight it, what new conventions join it. The durable parts of this guide are the ones that were true before AI and will outlast this file format: state your facts once and keep them consistent, describe every page honestly, and make the machine’s job easy wherever a human’s answer depends on it.
The bottom line
llms.txt is the cheapest move on the AI-visibility board: a twenty-minute file that replaces machine guesswork about your business with your own words. It will not rescue weak fundamentals and it does not need to — it needs only to be accurate, current and quotable. Write the summary you want repeated, map the pages that matter, and let every system that cares start from your framing.
Frequently asked questions
Does Google use llms.txt?
Google has not announced support, and llms.txt is not a ranking factor in classic search. The file targets the broader ecosystem of AI systems — chat assistants, answer engines, agent tools — several of which do fetch it. Write it for that audience and let Google keep reading your sitemap and schema.
What is llms-full.txt?
A companion convention: the entire site's meaningful content concatenated into one large markdown file, so a model can ingest everything without crawling. It shines for documentation sites; for a typical business site it is optional — the curated llms.txt plus clean pages covers the need.
Could it leak competitive information?
It contains only what your public site already says, gathered conveniently. If a fact is too sensitive for llms.txt, it was already too sensitive for your website — the file changes packaging, not exposure.
Twenty minutes — really?
For a typical small-business site: five minutes listing your key pages, ten writing one-liners and the blockquote, five uploading and verifying. The hard part is the discipline it quietly enforces — writing one honest sentence about what each page is for. Sites that struggle with that have discovered a content problem, not a file problem.
Should agencies write this file for clients, or should owners?
Ideally together. The owner supplies the one-sentence identity and the facts — nobody else knows which specialisms actually drive the business — while whoever maintains the site owns the mechanics: serving it as plain text, keeping it synced with content changes, versioning it. The failure mode to avoid is the file nobody owns: written once at launch by whoever was configuring the server, never touched again, quietly ageing into misinformation. Like every artefact in this space, it needs a name next to it.
What is llms.txt?
A plain-text markdown file at a website's root that gives AI language models a curated summary of the site: what the business is, canonical facts, and the key pages with one-line descriptions.
How long does llms.txt take to set up?
About twenty minutes for a typical small-business site: list the key pages, write one-line descriptions and a two-sentence business summary, upload to the web root and verify it serves as plain text.
Written by
BurTech Solution
Engineering team
The BurTech Solution engineering team designs, builds and maintains AI automation, ecommerce stores, SaaS and custom software for growing businesses. Everything on this blog comes from work we ship for clients and run ourselves.
Keep reading
More on seo & ai search.

SEO & AI Search10 min read
FAQ Schema in 2026: Still Worth It?
The rich-result dropdowns are history, but FAQ schema quietly became more valuable, not less: structured Q&A is what answer engines parse when choosing whom to quote. The honest 2026 assessment — what it does now, which pages deserve it, and how to…
Read the article →
SEO & AI Search9 min read
Content Refresh: When to Update, Merge or Delete Old Posts
Your archive is quietly split into winners, fixables and dead weight. The four-way decision framework — update, merge, delete or leave alone — with the criteria, the redirect mechanics, and why AI search makes refresh discipline pay twice.
Read the article →
SEO & AI Search10 min read
Local SEO for Service Businesses: The Google Business Profile Playbook
The map pack runs on relevance, distance and prominence — and the Google Business Profile is the lever. The fields that matter, the review system that compounds, local content, and the honest limits.
Read the article →