Leaderboard
Website readability for AI agents
One fetch per page with JavaScript switched off, which is what an AI crawler gets. If the words weren’t in the HTML, they didn’t count.
What we found. The median site scores 74. 19 of 40 reach the strong band, and only 10 bother to publish an llms.txt. The companies building AI came out ahead of the firms funding them, by 7.5 points at the median.
- 1
Bessemer Venture Partnersbvp.com88B
- 2
Lightspeedlsvp.com87B
- 3
Menlo Venturesmenlovc.com85B - 4
Index Venturesindexventures.com82B
- 5
Kleiner Perkinskleinerperkins.com78B - 5
Sequoia Capitalsequoiacap.com78B - 7
Battery Venturesbattery.com75B - 8
Coatuecoatue.com74C - 9
Y Combinatorycombinator.com71C
- 10
Founders Fundfoundersfund.com70C
- 11
Andreessen Horowitza16z.com69C - 12
Greylockgreylock.com68C
- 12
ICONIQiconiq.com68C - 14
General Catalystgeneralcatalyst.com67C - 15
Khosla Ventureskhoslaventures.com65C - 16
SV Angelsvangel.com64C - 17
IVPivp.com60C - 18
Accelaccel.com53D - 19
Spark Capitalsparkcapital.com48D - 20
Benchmarkbenchmark.com37F
- 1
Together AItogether.ai88B - 2
ElevenLabselevenlabs.io84B
- 3
Braintrustbraintrust.dev82B - 3
Gleanglean.com82B - 3
Sierrasierra.ai82B - 6
Cursorcursor.com81B
- 6
Lovablelovable.dev81B
- 8
Anthropicanthropic.com79B - 8
Fireworks AIfireworks.ai79B
- 10
Abridgeabridge.com77B - 10
Applied Computeappliedcompute.com77B
- 10
Harveyharvey.ai77B - 13
Decagondecagon.ai74C
- 14
Mercormercor.com73C - 15
Cognitioncognition.com70C
- 16
Thinking Machines Labthinkingmachines.ai69C - 17
Sunosuno.com68C
- 18
Reflection AIreflection.ai64C
- 19
Figurefigure.ai61C - 20
Physical Intelligencepi.website55D
- OpenAIopenai.com— returns HTTP 403 to automated requests
- Perplexityperplexity.ai— returns HTTP 403 to automated requests
These sites turn automated requests away at the door. They may well wave through the big answer-engine crawlers by name, but anything else gets the refusal we got, so we have no way to see what an agent would see. Better a blank than a number we can’t stand behind.
How we did this
We request each page once, with JavaScript off, and read robots.txt, sitemap.xml and llms.txt alongside it. For every domain we take the homepage plus up to six pages a crawler could actually find from there, and blend them into a single score across five pillars.
None of this is a model’s opinion. The rubric is arithmetic, so the same site on the same day always comes out the same, and each point traces back to a check you can see in the report. Blocking training crawlers costs nothing here. That’s a rights decision, not a readability problem.
What this measures is whether a machine can read you, not whether it trusts you. Reputation drives most real citations, and a single fetch can’t see any of it. A high score is the floor, not a promise that anyone quotes you.
Both lists are judgement calls, and they reflect where things stood on July 27, 2026. Sites change. Open any row and re-run the report yourself if you want to check our work.
Scored July 27, 2026. Run any row again and you’ll get the same number.