← Back to Blog

ChatGPT cites your competitor when users ask about your industry. Perplexity surfaces a five-year-old blog post that contradicts your current pricing. It is natural to want a way to point AI tools at your best pages. llms.txt is one proposed way: a curated map of the pages you consider most useful. This guide covers what it is, how to write one, and what it does and does not do, including the fact that Google Search ignores it.

What is llms.txt?

llms.txt is a plain-text Markdown file placed at the root of your domain that lists the URLs and short descriptions of the content you most want large language models to read and cite. Proposed by Jeremy Howard in September 2024, it acts as a curated content map for AI systems - a content-discovery hint, not an access control mechanism.

Think of llms.txt as the AI-era equivalent of an XML sitemap, but for LLMs instead of search engines. Where a sitemap dumps every URL on your site for crawlers to evaluate, llms.txt is a hand-picked, human-readable list of your canonical references - the pages an AI should fetch when it wants to understand what your business does.

llms.txt: Key Facts

What it is: A Markdown file at /llms.txt that lists URLs and descriptions for AI systems

Proposed by: Jeremy Howard (Answer.AI), September 2024

Format: Markdown - H1 site name, blockquote summary, H2 section groupings, bullet links with descriptions

Location: https://yoursite.com/llms.txt (root of domain, like robots.txt)

Not in the proposal: some sites also publish /llms-full.txt with the full text of the listed pages; the current proposal instead suggests a Markdown version of each page at the same URL with .md appended

Who publishes it: thousands of sites according to the proposal, including OpenAI, Anthropic and Gemini for their own developer docs

Effect on Google: none; Google says Google Search ignores llms.txt files

Effort: 30 minutes for a small site, half a day for a large documentation site

How Does llms.txt Actually Work?

A tool that uses llms.txt can read it to find your most useful pages quickly, usually when an agent or assistant needs information about your site while answering a question. The proposal says the file is mainly useful at that point, for inference, rather than for training. How each tool uses the file, and whether it reads it at all, is up to that tool, and few providers document it.

The protocol is intentionally simple. There is no API key, no verification, no signing. You publish a static file at a known location and AI clients that support the spec can choose to honour it. Adoption is driven by the AI tool, not by you - your job is to publish the signal and wait for tools to read it.

llms.txt vs robots.txt vs Sitemap: What Each One Does

llms.txtrobots.txtsitemap.xml
AudienceLarge language modelsSearch engine crawlersSearch engine crawlers
Purpose"Read these pages first""Do not fetch these URLs""These URLs exist, here is the freshness"
FormatMarkdownPlain text directivesXML
EffectDiscovery + curationAccess controlDiscovery
Enforced?Voluntary by AI clientVoluntary by good-citizen botsVoluntary by search engines
Coexist?YesYesYes

All three serve different masters and should be published together. llms.txt does not replace robots.txt or sitemap.xml - it complements them. A site with all three is signalling clearly to traditional search engines AND to AI systems.

What Should Your llms.txt Contain?

The Jeremy Howard spec defines a loose Markdown structure with five elements. Every element is optional except the H1.

Required: H1 Site Name

The single H1 at the top is your domain or product name. AI clients use this as the canonical reference for who you are. Example: # Daylytix.

Recommended: Blockquote Summary

A 2-3 sentence elevator pitch. This is what the AI will quote when it introduces you. Make it a clear definition: classification, mechanism, primary use case. Example: > Daylytix is an AI-powered SEO audit platform built for agencies and in-house SEO teams. It runs 150+ technical, content, and AI-readiness signals per site.

Recommended: H2 Section Groupings

Group your linked URLs by category. Common sections: Documentation, Blog, Product, Pricing, Legal. Each section is just an H2 heading.

Required (within sections): Bullet Links

Inside each H2, list the URLs as Markdown bullets with descriptions. Format: - [Page title](https://yoursite.com/page): Short description of what this page covers.. The description should be human-readable but information-dense - 1-2 sentences max.

Optional: "Optional" Sub-Section

The spec defines an ## Optional H2 where you list secondary references. AI clients may skip these to save tokens. Use this for cross-references and "see also" content rather than core authoritative pages.

How to Build Your llms.txt (Step by Step)

This is a 30-minute exercise for most sites. The hardest part is choosing which URLs to include - keep the list short and signal-dense rather than comprehensive.

List Your Top 15-30 Canonical URLs

Open Google Analytics or your CMS and identify the pages that get the most organic traffic, that explain what your product does, or that answer the most common support questions. Keep the list focused: a few dozen well-described pages are easier to use than hundreds.

Write a 1-2 Sentence Description for Each

For each URL, write a description that an AI could quote verbatim when answering a related question. Lead with the page's purpose, not its title. Example: "Pricing page: Daylytix offers three tiers (Starter €9.99/mo, Pro €22.99/mo, Agency €59.99/mo) with a 7-day free trial, no credit card required."

Group URLs Into Logical H2 Sections

Common groupings: Product (homepage, features, pricing), Docs (getting started, API reference, integrations), Blog (the 5-10 most cited posts), Company (about, contact). Use whatever taxonomy matches your site.

Save as Plain Text at the Domain Root

Create the file /llms.txt and serve it with content-type text/plain or text/markdown. Most CDNs do this automatically based on the .txt extension. Verify by visiting https://yoursite.com/llms.txt in a browser - you should see plain text, not a 404.

(Optional) Generate llms-full.txt

Some sites also publish an llms-full.txt that concatenates the text of every page listed in llms.txt into one Markdown file. It is not part of the llmstxt.org proposal. The current version of the proposal suggests something different: a clean Markdown version of each important page at the same URL with .md appended, linked from the page with rel="alternate" type="text/markdown".

Daylytix audits your llms.txt presence + AI-readiness in every report. Find out what AI crawlers see when they hit your site.
Try it free →

Example: A Minimal Working llms.txt

# Acme Tools

> Acme Tools is a SaaS platform for project managers. It combines task tracking,
> time logging, and client reporting into one workspace.

## Product
- [Homepage](https://acmetools.com/): Overview of features and use cases.
- [Pricing](https://acmetools.com/pricing): Three tiers from $19 to $99 per user / month.
- [Features](https://acmetools.com/features): Detailed list of every feature with screenshots.

## Documentation
- [Quickstart](https://docs.acmetools.com/quickstart): How to create your first project in 5 minutes.
- [API reference](https://docs.acmetools.com/api): REST API documentation with code samples.

## Blog (most-cited posts)
- [Project management for remote teams](https://acmetools.com/blog/remote-pm): 2026 guide.
- [Time tracking ethics](https://acmetools.com/blog/time-tracking): When to track and when to stop.

## Optional
- [About us](https://acmetools.com/about)
- [Press kit](https://acmetools.com/press)

Does llms.txt Actually Work in 2026?

It depends on what you mean by work. Google says Google Search ignores llms.txt files, and that publishing one for other services neither helps nor harms you in Google, so it will not change rankings, AI Overviews or AI Mode. According to the proposal’s author on llmstxt.org, thousands of sites publish an llms.txt file, documentation platforms generate one automatically, and OpenAI, Anthropic and Gemini publish llms.txt files for their own developer docs. Its main real-world use is AI agents, such as coding assistants, reading documentation while they work.

The honest answer: publishing llms.txt is a small, optional courtesy to AI tools, not a visibility lever. It takes little time and does no harm, and it is most useful for sites with documentation or reference content that agents look up. We could not find documentation from OpenAI or Perplexity saying their assistants read other sites’ llms.txt files, so do not expect measurable traffic from it.

Common Mistakes With llms.txt

Mistake 1: Treating It Like a Sitemap

Why it happens: SEOs default to "more URLs = more visibility". Why it backfires: An llms.txt with 500 URLs is just noise - the LLM cannot tell which pages are authoritative. What to do instead: 15-30 hand-picked URLs with thoughtful descriptions.

Mistake 2: Forgetting the Descriptions

Why it happens: Easy to auto-generate a bare list of URLs. Why it backfires: The descriptions are what the LLM quotes. Without them, you lose control of how your site is summarised. What to do instead: Write each description as if it were the answer the AI will give.

Mistake 3: Serving It as text/html

Why it happens: Some CMSes wrap text files in HTML templates by default. Why it backfires: tools that read the file expect plain text or Markdown, not an HTML page. What to do instead: Configure your web server to serve .txt with the correct content type, or use a CDN edge rule to override.

Limitation: No Enforcement

llms.txt is voluntary on the AI client side. There is no way to verify that ChatGPT actually fetched your file or used your descriptions. The current ecosystem is built on goodwill and convention, much like robots.txt in 1994. Expect this to formalise as AI search matures.

TL;DR: llms.txt Summary

What it is: A Markdown file at /llms.txt listing your canonical URLs and descriptions for AI systems.

How it is used: AI agents and assistants that choose to read it use it as a map of your key pages, mainly while answering questions.

vs robots.txt: Different audience, different purpose; both coexist.

Who publishes it: thousands of sites, including OpenAI, Anthropic and Gemini for their developer docs. Google Search ignores it.

Effort to implement: 30 minutes to half a day depending on site size.

Effect on Google rankings and AI Overviews: none.

Bottom line: publish one if it is easy to keep current, especially if you have documentation; spend most of your effort on crawlable, helpful pages and AI crawler access in robots.txt.

Frequently Asked Questions

What is llms.txt?

llms.txt is a plain-text Markdown file placed at the root of your domain that lists the URLs and short descriptions of the content you most want large language models to read and cite. It is a content-discovery hint, not an access control mechanism.

Is llms.txt the same as robots.txt?

No. robots.txt tells crawlers which URLs they are allowed to fetch. llms.txt tells AI systems which URLs you would most like them to read. The two coexist and serve different purposes.

Do ChatGPT and Perplexity actually read llms.txt?

We could not find documentation from OpenAI or Perplexity saying their assistants read other sites’ llms.txt files. OpenAI, Anthropic and Gemini do publish llms.txt files for their own developer docs, and Google says Google Search ignores llms.txt.

Where do I put the llms.txt file?

At the root of your domain, served at https://yoursite.com/llms.txt. It must be a plain-text file in Markdown format, returning a 200 OK with content-type text/plain or text/markdown.

Should I publish llms-full.txt as well?

Only if it is easy for you. llms-full.txt is not part of the llmstxt.org proposal; the current version suggests Markdown versions of individual pages instead. Neither affects Google Search.

Related Topics

Getting Started

Daylytix publishes its own file at daylytix.com/llms.txt, regenerated whenever we add or change pages. We treat it as a courtesy to AI tools, not a ranking lever.

The work that does move AI visibility is the same as for search: pages Google can crawl and index, AI search crawlers allowed in robots.txt, and content worth quoting. Run an audit with Daylytix free for 7 days and it checks your llms.txt, your robots.txt rules for AI crawlers and your AI-readiness signals in one pass.