ARB/1.1: what changed, and what it does to real scores
Three-state verdicts, an off-site corroboration pillar, a transport & security surface pillar, pillar weights rebalanced to still total 100. We re-scored the same 30-site cohort ARB/1.0 measured on 2026-08-24, and re-scored ourselves. Both distributions are below, with the per-site delta. Old results stay labelled ARB/1.0 — nothing here silently restates a 1.0 score as 1.1.
What changed, and why
Three-state verdicts. PASS/FAIL/PARTIAL/UNKNOWN is retired in favour of VERIFIED / UNVERIFIED / UNOBSERVABLE. PARTIAL's half-credit is gone entirely — it was a soft default in the direction of a pass for evidence that didn't actually clear the bar. UNOBSERVABLE replaces UNKNOWN with the same scoring behaviour (excluded from the denominator) but is now a named count on every report, not something a reader has to infer from a possible/pct mismatch.
Independence. Findings satisfied by corroborating sources (sameAs links, outbound citations) now carry an independence field: independent if the sources span two or more distinct registrable domains, correlated if they're all the same domain restated, unknown if there's nothing to compare. Three citations that all resolve to one domain are not three sources.
Off-site Corroboration (new pillar, weight 10). Do the entity's sameAs targets actually resolve, and are they independent? Do outbound citations resolve, across how many distinct domains? Is the entity present in Wikidata? Does the security.txt contact actually resolve? Are quantified claims sitting next to the source that supports them? All five are deterministically measurable from public data, no paid API, no estimate.
Transport & Security Surface (new pillar, weight 10). TLS version and certificate chain validity, HSTS with a sane max-age, a CSP that isn't trivially permissive, X-Content-Type-Options, Referrer-Policy, Permissions-Policy, no mixed content, a Server banner that doesn't leak a version, security.txt present. This is a public-surface security check. It is NOT a penetration test and every ARB/1.1 report says so in those words.
Weight rebalance. The two new pillars are weighted 10 points each (20 total) — honest, supplementary signal, not inflated to dominate the score just because they're new. The original seven pillars are scaled by ×0.8 and rounded, which preserves their relative ordering (machine_access and structured_data stay heaviest, governance stays lightest of the seven): 14+14+13+11+10+10+8 = 80, +10+10 = 100.
The same 30-site cohort, both versions
Identical URLs, identical order, to the ARB/1.0 cohort run from 2026-08-24. Re-measured 2026-08-26 with ARB/1.1.
| min | median | max | mean | |
|---|---|---|---|---|
| ARB/1.0 (n=30) | 2.0 | 61.8 | 79.5 | 54.1 |
| ARB/1.1 (n=30) | 3.2 | 59.8 | 77.7 | 53.6 |
Honest note on the drop. The median moved down 1.7 points and the max moved down 1.8 — a real change in what we measure, not a change in those thirty sites. What's more interesting than the average is the spread: sites already scoring well under 1.0 mostly dropped further (docusign.com -9.5, zendesk.com -7.0, cloudflare.com -6.5) as they hit the two new, harder pillars and lose PARTIAL's half-credit. Sites that scored very low under 1.0 — mostly because the base page fetch itself failed or returned almost nothing — sometimes scored HIGHER under 1.1 (openai.com +14.9, servicenow.com +13.5, sap.com +8.4): when there's barely any content to evaluate, more checks land on UNOBSERVABLE (excluded from the denominator) instead of UNVERIFIED (counted and failed), which can raise the percentage of what little was actually verified. That's an honest artifact of "missing evidence excluded, not penalized" working as designed, not a sign those sites got better.
| Site | v1.0 | v1.1 | delta |
|---|---|---|---|
| docusign.com | 68.5 | 59.0 | -9.5 |
| holisticai.com | 55.5 | 46.6 | -8.9 |
| zendesk.com | 78.5 | 71.5 | -7.0 |
| cloudflare.com | 69.0 | 62.5 | -6.5 |
| snowflake.com | 54.0 | 48.3 | -5.7 |
| onetrust.com | 69.0 | 64.2 | -4.8 |
| stripe.com | 74.5 | 70.1 | -4.4 |
| salesforce.com | 71.0 | 66.8 | -4.2 |
| vanta.com | 46.5 | 42.5 | -4.0 |
| paloaltonetworks.com | 60.0 | 56.9 | -3.1 |
| workday.com | 65.5 | 63.1 | -2.4 |
| credo.ai | 78.0 | 76.1 | -1.9 |
| ibm.com | 53.5 | 51.6 | -1.9 |
| atlassian.com | 79.5 | 77.7 | -1.8 |
| hubspot.com | 75.5 | 73.8 | -1.7 |
| twilio.com | 74.0 | 72.3 | -1.7 |
| microsoft.com | 57.5 | 56.2 | -1.3 |
| datadoghq.com | 72.0 | 71.6 | -0.4 |
| mongodb.com | 60.5 | 60.3 | -0.2 |
| okta.com | 68.0 | 68.9 | +0.9 |
| zoom.us | 76.5 | 77.7 | +1.2 |
| anthropic.com | 48.0 | 49.2 | +1.2 |
| googlecloud.com | 2.0 | 3.2 | +1.2 |
| slack.com | 57.5 | 59.4 | +1.9 |
| crowdstrike.com | 63.0 | 66.4 | +3.4 |
| oracle.com | 7.5 | 11.8 | +4.3 |
| drata.com | 8.0 | 14.2 | +6.2 |
| sap.com | 9.0 | 17.4 | +8.4 |
| servicenow.com | 8.0 | 21.5 | +13.5 |
| openai.com | 13.5 | 28.4 | +14.9 |
Raw data: cohort-v1.0.json · cohort-v1.1.json.
apexclawai.com's own ARB/1.1 score
Published as measured, not smoothed. Two real, fixable gaps found and fixed as part of this measurement — nginx was sending X-Content-Type-Options, Referrer-Policy and Permissions-Policy on a different virtual host than the one serving apexclawai.com's own homepage, and the Server header was leaking its version number. Both are corrected. What's left unfixed and genuinely lowers the score: the entity has no Wikidata item, and the only sameAs link is LinkedIn — a single source, not independent corroboration by this pillar's own definition. Neither is invented to close the gap.
Conformance
Every check above is backed by a runnable adversarial vector — a clean fixture that must not trip it, a dirty fixture that must. See /conformance/.
Licensed CC-BY 4.0. Method: methodology.json. Update, 2026-08-26 (BAKE8): ARB/1.1 is now the live interactive audit tool's default engine — methodology.json above reflects ARB/1.1, the same rubric this page's own re-scoring used. A run at the live tool today produces an ARB/1.1 result, not the ARB/1.0 result this update was originally written to distinguish from.
Common questions
Why does ARB/1.1 score lower than ARB/1.0 for most sites?
PARTIAL credit is gone and two new pillars, off-site corroboration and transport security, set a higher bar. Twenty of thirty cohort sites dropped; the rubric got stricter, not the sites worse.
Is a v1.1 score comparable to a v1.0 score?
No. They measure different things under different rules. A v1.0 score stays labelled v1.0 permanently, and nothing on this site restates an old score under the current rubric.
How often does the ARB rubric change?
Whenever a check or a weight changes, which is rare and always dated. Two revisions exist so far, ARB/1.0 on 2026-08-24 and ARB/1.1 on 2026-08-26, both listed in the versioning table.
What does the median drop actually mean?
Sites already weak under ARB/1.0 mostly got weaker, since the new pillars and lost PARTIAL credit hit them hardest. It reflects a stricter rubric, not thirty sites getting worse in two days.
Sources
Machine identities already outnumber human ones by more than 80x in most enterprises, and Gartner projects more than 40% of agentic AI projects will be cancelled by the end of 2027 — the same discipline against invented numbers applies to this cohort re-score: 0% partial credit anywhere in ARB/1.1, pillar weights totalling exactly 100%, verified by methodology.json at measurement time. Standards referenced: OWASP Top 10 for Agentic Applications, NIST AI RMF.