The 2026 Agent Readiness Benchmark: what 30 measured homepages actually score
Every number on this page is recomputed at build time from the same cohort file the rubric publishes — not asserted, not rounded from memory.
Published · updated · by Julian Joseph, founder
Get an Agent Trust Gap BriefWhat is the 2026 Agent Readiness Benchmark?
It is a measurement of 30 well-known B2B SaaS and AI-governance vendor homepages against the published ARB/1.0 rubric — 36 deterministic checks across 7 pillars, defined at methodology.json (spec date 2026-08-24).
What is the real score distribution across the 30 sites measured?
Recomputed from cohort.json (generated 2026-08-24, n=30): scores cluster in two bands — 6 sites scored under 20 (blocked or unreachable), and 9 scored 70–79.9. None reached 80.
| Score band | Sites | Share |
|---|---|---|
| 0–20 | 6 | 20.0% |
| 20–30 | 0 | 0.0% |
| 30–40 | 0 | 0.0% |
| 40–50 | 2 | 6.7% |
| 50–60 | 5 | 16.7% |
| 60–70 | 8 | 26.7% |
| 70–80 | 9 | 30.0% |
| 80–90 | 0 | 0.0% |
| 90–100 | 0 | 0.0% |
Source: /var/lib/apexclaw-audit/cohort.json, generated_utc 2026-08-24T16:55:50Z. Recomputed at build time, 2026-08-24.
What is the median score, and how spread out is it from the mean?
The median is 61.8 and the mean is 54.1, a 7.6-point gap. The mean sits lower because 6 of 30 sites were blocked or unreachable and scored near zero, pulling the average down without moving the median.
Which pillars fail most often across the cohort?
Defining "failed" as a site scoring under 50% on that pillar, Agent Interaction Surface fails for 53.3% of the cohort — more than any other pillar, ahead of Structured Data and Evidence & Citability, tied at 46.7%.
| Pillar | Mean % | Median % | Sites failing (<50%) |
|---|---|---|---|
| Agent Interaction Surface | 38.3 | 33.3 | 53.3% (16/30) |
| Structured Data | 35.7 | 61.1 | 46.7% (14/30) |
| Evidence & Citability | 45.3 | 50.0 | 46.7% (14/30) |
| Governance & Trust | 49.0 | 50.0 | 36.7% (11/30) |
| Answer Extractability | 57.1 | 70.3 | 30.0% (9/30) |
| Machine Access | 74.8 | 88.9 | 20.0% (6/30) |
| Entity Clarity | 77.1 | 87.5 | 20.0% (6/30) |
Source: /var/lib/apexclaw-audit/cohort.json, 30 sites, recomputed 2026-08-24. "Failing" is this page's own threshold (pillar % < 50), stated here rather than implied.
Why do Machine Access and Entity Clarity score so much higher than Agent Interaction Surface?
Machine Access and Entity Clarity (means 74.8% and 77.1%) reward decades-old SEO practice: HTTPS, titles, sitemaps. Agent Interaction Surface (mean 38.3%) rewards booking links, discoverable APIs and structured contact data — practices with no legacy incentive yet.
How many of the 30 sites could not be scored properly at all?
5 sites returned HTTP 403 to the audit user-agent and 1 failed to connect — 6 of 30 (20%) were closed to the same request pattern an AI agent would send, before any content check ran.
What this benchmark deliberately does not claim
- Not a full-site crawl. Each site was measured at one URL plus the site-wide files at its root (robots.txt, llms.txt, sitemap.xml, security.txt, ai.txt) — the scope the rubric itself publishes. A different page on the same domain can and does score differently.
- Not a ranking of these companies. The cohort is 30 public B2B homepages chosen to exercise the rubric across a range of real sites, not a competitive leaderboard. A low score reflects agent-readiness signals, not product quality.
- Not permanent. The 5 sites returning HTTP 403 here reflect a bot-detection posture a site can change at any time; a re-run tomorrow would not necessarily reproduce the same 5 of 30 blocked.
- Not a claim about any single check in isolation. The per-pillar failure rates on this page are aggregates; which individual check drives each pillar down is a separate, narrower question — see the single most-failed check.
- Not a statement about how any one agent will actually behave. The rubric measures what a page structurally offers — JSON-LD, dates, a booking link — not whether a specific deployed agent will notice or use it. It is a readiness measure, not a behavior prediction.
- Not a one-time snapshot presented as a forecast. Every figure above carries the exact generation timestamp it was measured at; a claim about tomorrow's web is out of scope for a page that only reports what was true today.
Method note
Every figure above is recomputed directly from /var/lib/apexclaw-audit/cohort.json (30 sites, generated 2026-08-24T16:55:50Z, engine ARB/1.0, source cohort_run.py) and cross-checked against the summary the rubric itself publishes at methodology.json (median 61.8, min 2.0, max 79.5 — matching this page's independent recomputation of 61.75/2.0/79.5). No site named here is a client of ApexClaw or ThruLiquid; the cohort is public homepages measured for methodology demonstration, per the rubric's own published cohort block.
Standards referenced by the Governance & Trust pillar: OWASP Agentic Top 10, NIST AI RMF, ISO/IEC 42001, EU AI Act.
By Julian Joseph, Founder, ApexClaw. Every figure on this page is recomputed at build time from a named source file and date — see the method notes above. Reviewed against the claims policy: sourced, first-party, or labelled.