Markdown for Agents: Serving text/markdown Through Content Negotiation
Serve the same URL as Markdown when an agent asks for it. What we check, how to set it up, and the caching header people forget.
What we check
We request the page you scanned a second time, with the header Accept: text/markdown. The check passes when the response is a 2xx and its Content-Type starts with text/markdown.
A page that ignores the header and returns HTML does not pass, even with a 200 status. Neither does a response that contains Markdown but declares text/plain or text/html. We cache the result per path, because sites often support this on their docs and not on their marketing pages.
Why agents want it
An agent pays for every token it reads. A typical HTML page is mostly navigation, scripts, class names and footer links, and the article itself is a small share of the bytes. The same content as Markdown is a fraction of the size, and the agent is less likely to quote something from your cookie banner.
Content negotiation is the way HTTP has always handled this: one URL, several formats, and the client says which one it wants. Browsers keep sending Accept: text/html and keep getting your normal page.
How to set it up
Check your platform before writing code. Some CDNs and documentation platforms can convert pages to Markdown with a setting, and several docs platforms serve Markdown already.
If you build it yourself:
- Detect
text/markdownin the request'sAcceptheader. - Return the page's main content as Markdown: the title, the body, headings, lists, tables and links. Leave out navigation, footers and widgets.
- Set
Content-Type: text/markdown; charset=utf-8. - Set
Vary: Accepton both the HTML and the Markdown response.
The last step is the one people forget. Without Vary: Accept, a CDN can cache the Markdown version and serve it to the next human visitor, or cache the HTML and serve it to every agent.
If your content is stored as HTML, convert it at publish time and store the Markdown next to it. Converting on every request works too, but costs CPU on each agent visit.
Test it
curl -s -D - -o /dev/null -H "Accept: text/markdown" https://yoursite.com/some-page
# look for: content-type: text/markdown
# vary: Accept
Then request the same URL without the header and confirm you still get HTML.
How it relates to llms.txt
They work together. llms.txt is an index that tells an agent which pages matter, and its links should point at Markdown versions. Content negotiation lets you serve that Markdown at the page's own URL instead of a separate .md path. See the llms.txt article for how the two connect.
๐ก Quick win
Start with the pages agents are most likely to fetch: docs, pricing, and your top articles. You pass the check for any page you scan that supports it, so you do not need full coverage on day one.
