Research: agent readiness, measured and sourced
Five studies, each built from one of three named data files and recomputed at build time. No number here is asserted from memory.
Published · updated · by Julian Joseph, founder
Get an Agent Trust Gap BriefWhat is this research section, and where does its data come from?
Five studies built from exactly three named files: cohort.json (30 measured sites), the published methodology.json rubric, and aga.py's 27 governance controls — nothing beyond those three.
What are the five studies published here?
A distribution study, a single-check failure study, a mechanics explainer, a standards crosswalk, and a deterministic-vs-judgment comparison — each answering one question the raw data files do not answer on their own.
| Study | What it answers | Headline figure |
|---|---|---|
| The 2026 Agent Readiness Benchmark | Real score distribution, median and per-pillar failure rates across 30 measured B2B homepages. | Median 61.75, mean 54.1, 30 sites, 0 above 80 |
| The One Check Almost Every Site Fails | Which single check fails hardest across the cohort, recomputed check by check, and why it matters. | sd.article fails 100% of the cohort (30/30) |
| How an AI Agent Actually Reads a Website | What the published rubric says an agent actually fetches and scores, with no browser and no guessing. | 6 fetches, 36 checks, 9 free-tier |
| 27 Governance Controls, Mapped to Standards | Every deterministic agent-governance control, its weight, and its OWASP/NIST/ISO/EU AI Act reference. | 27 controls, 7 families, 54 standard refs |
| Deterministic Audit vs. LLM Judgment | Two adversarial fixture trees, re-scored at build time, showing what file-and-line evidence catches. | tricky_clean PASSes, tricky_dirty FAILs, both re-run today |
How is a number on these pages verified?
Every figure is recomputed directly from its named source file at build time and stated inline with that file and its date — never rounded from memory, never carried over from an earlier draft without re-checking.
Will this section be updated as new data comes in?
Yes. A cohort re-measurement or a rubric revision gets a new build date and a fresh recomputation, following the same expiry discipline the trust center applies to every other claim on this site.
Why only three source files, instead of a wider literature review?
Because a narrower, fully-verifiable claim beats a broader, half-sourced one. Every figure here traces to a named file this site can point to directly, rather than to a study this site cannot re-run or re-check on its own.
The five studies
The 2026 Agent Readiness Benchmark
Real score distribution, median and per-pillar failure rates across 30 measured B2B homepages.
Read →The One Check Almost Every Site Fails
Which single check fails hardest across the cohort, recomputed check by check, and why it matters.
Read →How an AI Agent Actually Reads a Website
What the published rubric says an agent actually fetches and scores, with no browser and no guessing.
Read →27 Governance Controls, Mapped to Standards
Every deterministic agent-governance control, its weight, and its OWASP/NIST/ISO/EU AI Act reference.
Read →Deterministic Audit vs. LLM Judgment
Two adversarial fixture trees, re-scored at build time, showing what file-and-line evidence catches.
Read →Five figures, one from each study
- The measured cohort's median score is 61.75 against a possible 100, with none of the 30 sites reaching 80 — the distribution study.
- One single check, Article/TechArticle markup, fails on 100% of the same 30-site cohort — the single-check study.
- The published rubric weights Machine Access and Structured Data at 18% each, more than double Agent Interaction Surface's 12% — the mechanics study.
- NIST references account for 35.2% of the 54 standard citations across all 27 governance controls, more than any other single standard — the standards crosswalk.
- Both controls compared in the fixture study matched their intended verdict on 100% of the four fixture/control pairs tested — the deterministic-vs-judgment study.
Why publish this as five separate pages instead of one long one?
Because each question deserves its own citable URL. An answer engine, or a person, linking to "the distribution study" should not have to link to the same page as "the standards crosswalk" and hope the right section gets quoted. Five distinct pages means five distinct, individually correct citations — each with its own dated method note, its own table, and its own answer, rather than one page trying to be authoritative about five different questions at once.
Sources for this section
/var/lib/apexclaw-audit/cohort.json (30-site measurement, 2026-08-24), methodology.json (the ARB/1.0 rubric as it stood on 2026-08-24, when these five studies were built), and /root/apexclaw-audit/aga.py (27 governance controls). Standards referenced across the five studies: OWASP Agentic Top 10, NIST AI RMF, ISO/IEC 42001, EU AI Act.
The engine itself moved to a new major version after these five studies were published, and as of 2026-08-26 (BAKE8) ARB/1.1 is the live interactive audit tool's default engine — the methodology.json link above now resolves to ARB/1.1, not the ARB/1.0 snapshot these five studies were built from. See the ARB/1.1 scoring update (three-state verdicts, two new pillars, this same cohort re-measured) and the runnable conformance vectors behind it.
By Julian Joseph, Founder, ApexClaw. Every figure on this page is recomputed at build time from a named source file and date — see the method notes above. Reviewed against the claims policy: sourced, first-party, or labelled.