← Back to SEO tools
RESEARCH / 124

We tested 184 popular tool sites.
Average score: 2.3 / 9.

AI agents are learning to read the web — the web isn't ready. This study measures llms.txt adoption, Markdown negotiation, structured data, and crawler rules across 184 reachable online tool websites (October 2026).

How ready is the web for AI agents? Measured across 184 popular tool websites in October 2026: 31.5% publish an llms.txt, but only 1.6% actually serve Markdown when an agent asks (Accept: text/markdown) — and none of the llms.txt adopters do. The average readiness score is 2.3 out of 9; the best sites score 6.

AI agent readiness findings

The headline numbers

SignalSites% of 184
robots.txt allows AI crawlers (GPTBot, ClaudeBot, PerplexityBot, GrokBot, Bingbot)13372.3%
XML sitemap available8144.0%
llms.txt published5831.5%
JSON-LD structured data4423.9%
llms-full.txt published3921.2%
Markdown content negotiation31.6%
rel="alternate" Markdown link relations10.5%

Key findings

Leaderboard (top 10 by score)

SiteScoreMissing
base64encode.org6 / 9Markdown, rel=alternate
imageonline.co6 / 9Markdown, rel=alternate
omnicalculator.com6 / 9Markdown, rel=alternate
quicktype.io6 / 9JSON-LD, Markdown
tableconvert.com6 / 9Markdown, rel=alternate
trydocsy.com6 / 9Markdown, rel=alternate
urlencoder.org6 / 9Markdown, rel=alternate
11zon.com5 / 9Markdown, JSON-LD, rel=alternate
animista.net5 / 9Markdown, JSON-LD, rel=alternate
file.io5 / 9Markdown, JSON-LD, rel=alternate

Methodology

A curated sample of 212 popular browser-based tool websites (image editing, PDF, file conversion, developer tools, calculators, SEO utilities, design tools) was collected in October 2026. 184 sites responded and are included; 28 were excluded as unreachable (bot protection, timeouts, or availability issues). Each site was fetched once with a standard browser user agent over HTTPS (HTTP fallback). Signals: robots.txt parsed for explicit Disallow: / rules targeting GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, anthropic-ai, PerplexityBot, GrokBot, xAI-SearchBot, or Bingbot; llms.txt / llms-full.txt presence (HTTP 200, non-empty); Markdown negotiation tested with Accept: text/markdown, text/html;q=0.9 and a text/markdown response required; JSON-LD detected via application/ld+json in the HTML; sitemap via /sitemap.xml; link relations via rel="alternate" type="text/markdown". Scoring: robots 1, llms.txt 2, llms-full.txt 1, Markdown negotiation 2, JSON-LD 1, sitemap 1, alternate 1 (max 9). Limitations: a single snapshot; CDN and geo-dependent behavior varies; unreachable sites are excluded rather than scored.

Tool Pantry itself scores 9/9 on this checklist — it is the reference implementation behind the open-source kit described below. Reproduce the study with the scan script in the dataset repository.

Advertisement
QUICK ANSWERS
What is Markdown content negotiation?

It is HTTP content negotiation (RFC 9110): an AI agent requests a page with Accept: text/markdown and the server replies with a clean Markdown version instead of full HTML. Same URL, two representations. It saves tokens and improves extraction accuracy for AI assistants — currently only 1.6% of the tested tool sites do it.

Why do llms.txt adopters not serve Markdown?

llms.txt (31.5% adoption in this study) tells AI systems where your content is; it doesn't change what the server returns when your pages are fetched. The two work best together — an index pointing at Markdown URLs — but most sites so far stop at the index.

Does being AI-ready help a tool site get cited?

It is necessary but not sufficient: AI answers cite sources that are crawlable, parseable, structured, and mentioned elsewhere on the web. This checklist covers the technical half; backlinks, lists, and reviews cover the other half.

How do I check my own site?

Use the free checker at toolpantry.app/agent-ready-checker — it runs the same seven checks as this study on any URL and explains exactly how to fix each gap. The fix kit is open source (MIT).

Can I reproduce this study?

Yes. The full dataset (CSV) is linked above, and the scan script plus the domain list are published in the companion GitHub repository. Re-run it any time to track adoption changes.