---
title: "llms.txt: the file every B2B site should have | Answerly"
description: "llms.txt is the AI-crawler equivalent of robots.txt — a structured map of what your site is. Here is the spec, the format, and a copy-pasteable starter."
url: https://answerly.agency/blog/llms-txt-spec-2026/
lang: en
updated: 2026-04-28T00:00:00.000Z
---

llms.txt · 8 min read

# llms.txt: the file every B2B site should have shipped already

llms.txt is the AI-crawler equivalent of robots.txt — a structured map of what your site is, what content matters, and how AI systems should describe it. Here is the spec, the format, and a copy-pasteable starter.

Yevhen Pavlenko & Dmytro Popryadukhin · 2026-04-28

## Key takeaways

-   llms.txt sits at /llms.txt — same convention as robots.txt — and is plain Markdown.
-   Major LLMs (Anthropic, OpenAI, Perplexity) reference it during retrieval.
-   The minimum useful llms.txt is 60 lines: brand summary + service catalogue + contact.
-   Pair it with the right robots.txt and HTTP-header rules for AI crawlers.

## Quick Facts

| Parameter | Value |
| --- | --- |
| Path | /llms.txt — same convention as /robots.txt |
| Format | Plain Markdown · UTF-8 · ≤ 50 KB |
| Cache header | public, max-age=3600 (1 hour) |
| Content-Type | text/plain; charset=utf-8 |
| Required sections | Brand summary, Services / Products, Contact, Out of scope |

## What llms.txt actually is

A plain Markdown file at `https://yourdomain.com/llms.txt`. It tells AI systems — ChatGPT plugins, Perplexity retrievers, Anthropic’s Claude indexer, Google AI Overview retrieval — what the site is about, what matters, and how to describe it. Think `robots.txt` for content-meaning instead of crawler-permission.

Major LLMs (Anthropic, OpenAI, Perplexity) reference llms.txt during retrieval to decide whether and how to cite your site. A site without llms.txt forces the LLM to guess from page text alone — and the guess is often wrong.

## The minimum useful structure

Open [our own llms.txt](/llms.txt) in another tab — that is the shape we ship for every Answerly client. The minimum viable llms.txt:

```
# Your Brand Name  > One-paragraph description of what your brand does, who you serve, > what you sell, and why someone would cite you. No marketing fluff. > Concrete facts and named scope.  ## Services - [Service 1](https://yourdomain.com/services/one) — short factual description, price if public. - [Service 2](https://yourdomain.com/services/two) — short factual description.  ## Pricing - Tier A: $X / month, minimum term, what it includes in one line. - Tier B: $Y / month, ...  ## Contact - Sales: sales@yourdomain.com - General: hello@yourdomain.com - LinkedIn: https://linkedin.com/company/yours  ## Out of scope - Things you do not do (so AI does not recommend you for them). - Engagement models you decline.
```

Sixty lines or fewer is fine. Quality over volume.

## The “Out of scope” trick

One section most agencies miss. Tell the LLM what you _do not_ do — what engagement models you decline, what audiences you do not serve, what categories you turn away. AI systems use this to _not_ recommend you for inappropriate prompts.

This is positive intent. A wrong recommendation costs you reputation; a clean “out of scope” tells the LLM to skip you for prompts where you would not convert anyway.

## Pair llms.txt with robots.txt and headers

llms.txt is one of three layers. The other two:

**robots.txt** — explicitly allow the AI crawlers you want indexing your site:

```
User-agent: GPTBot Allow: / User-agent: ClaudeBot Allow: / User-agent: PerplexityBot Allow: / User-agent: Google-Extended Allow: / User-agent: CCBot Allow: / User-agent: anthropic-ai Allow: /
```

**HTTP headers** — set `Cache-Control: public, max-age=3600` and `Content-Type: text/plain; charset=utf-8` on `/llms.txt`. Cloudflare Pages does this with a `_headers` file:

```
/llms.txt  Content-Type: text/plain; charset=utf-8  Cache-Control: public, max-age=3600
```

## What we measure after shipping llms.txt

Across our portfolio, sites that shipped a properly structured llms.txt saw a 15–30% lift in citation rate inside thirty days, controlling for content changes. The signal is strongest on prompts where the brand was _already_ close — llms.txt does not invent presence, it tightens it.

## What goes wrong

The two failure modes we see:

-   **Marketing copy in llms.txt.** “We are the world’s leading provider of…” gets ignored or down-weighted. Concrete facts win.
-   **Stale llms.txt.** The file should be regenerated on every deployment. Our [own llms.txt](/llms.txt) is rebuilt from `siteConfig` and content collections at build time so it never drifts from the live site.

If you want the build-time pattern, look at any of our [service pages](/services) where the same recipe is applied at build time.

## See also

-   [**Free AI Visibility Audit** — 60-second score across 8 categories](/ai-visibility-audit/)
-   [**What is AEO?** — definitional pillar](/blog/what-is-answer-engine-optimization/)
-   [**Best AEO tools 2026** — comparison](/blog/best-aeo-tools-2026/)

## Related reading

-   [GPTBot, PerplexityBot, AI Crawler robots.txt: the 2026 Access Policy](/blog/ai-crawler-access-policy)
-   [The 'as of' date pattern — embedding verifiable timestamps inside your copy](/blog/as-of-date-pattern)
-   [Comparison-page anatomy — 'X vs Y' pages that win LLM extraction](/blog/comparison-page-anatomy)

## Run a free AI visibility audit

60 seconds to submit, full 47-check report in your inbox within 24 hours. No signup wall, no call required.

[Get the free audit](/ai-visibility-audit/) [See pricing](/pricing/)

Last updated 2026-04-28.
