ApexClaw
Agent Governance, Measured

ApexClaw controls what your AI agents can do, and proves what they did.

The wedge: deterministic, rule-based measurement of agent systems — no model sits in the judgment seat. When we scan a customer's own source tree, the scan runs on their machine and never leaves it.

49deterministic checks, ARB/1.1
39real organisations scored
30site benchmark corpus
How a decision actually happens

One request, followed end to end.

Not marketing shapes — the real control path. A deterministic gate decides before anything executes; a checksummed receipt is written after.

Agent requests an action Policy evaluated deterministic rules — no model in the loop same input, same verdict, every time ALLOW in scope, passes every gate ESCALATE requires human approval BLOCK refused before execution Execution Execution receipt written carries a SHA-256 integrity checksum policy version, approval, payload hash, outcome Verifiable offline by anyone holding the receipt file The honest boundary The checksum detects an edited record. It does not prove who produced it. No signature exists on this record today. Signing is planned, not shipped. refusal receipt written no execution if approved Three possible outcomes Agent requests an action Policy evaluated deterministic rules, no model same input, same verdict, every time ALLOW in scope, passes every gate ESCALATE requires human approval BLOCK refused before execution Execution Execution receipt written SHA-256 integrity checksum policy, approval, payload hash, outcome Verifiable offline by anyone holding the receipt file The honest boundary The checksum detects an edited record. It does not prove who produced it. No signature exists on this record today. Signing is planned, not shipped. refusal receipt written no execution If ALLOWED, or ESCALATE approved, continue below

Static picture when motion is reduced: the same diagram, same labels, no animation.

The platform

Five surfaces, one trust spine.

Each is a control that fires before an action, or an evidence object produced after it — the canonical set, from /platform/.

Agent Passport

Every agent carries an identity: owner, mission, permissions, autonomy level, and lifecycle. No anonymous actors.

Agent identity →

Assurance Gateway

Deny-by-default policy gates and payload-bound, expiring approvals decide whether an action executes.

Policy enforcement →

Flight Recorder

Checksummed execution receipts you can replay to reconstruct any action, end to end.

Execution receipts →

MCP Governance

Identity, authorization and tool-exposure control for Model Context Protocol servers and tools.

MCP governance →

Agent Wallet Governance

Hard, fail-closed ceilings on actions, rate, and spend. A cap that refuses, not a throttle that slows.

Wallet governance →
  • Agent Passport — identity, owner, mission, permissions, autonomy level, lifecycle.
  • Assurance Gateway — deny-by-default gate; payload-bound, expiring approvals.
  • Flight Recorder — checksummed execution receipts, replayable end to end.
  • MCP Governance — identity and tool-exposure control for MCP servers.
  • Agent Wallet Governance — hard, fail-closed ceilings on actions, rate and spend.
Self-serve

Run the same engine against your own page.

The Agent Readiness Benchmark (ARB/1.1) fetches one public page and runs 49 deterministic checks across 9 pillars — machine access, structured data, extractability, evidence, entity clarity, agent surface, governance, off-site corroboration, and security surface. Free teaser scan; full findings, remediation and the standards crosswalk on unlock. Full rubric →

Our own homepage, ARB/1.1 — labelled as ours
92.5 / 100

Verified live at audit.apexclawai.com — three-state verdicts, never a silent pass.

45Verified
3Unverified
1Unobservable

ApexClaw also runs the governance it sells: Omega, the autonomous system we operate ourselves under our own controls, has produced 0% real external sends across its entire history — verified at the network layer, not inferred from logs.

30-site public-homepage cohort — real scores, not a mockup

median 59.8 0score100 n = 30 · min 3.2 · max 77.7

Only 26.7% of this cohort scores 70 or above — see the full cohort and methodology.

What's in the full report?

  • Every check, three-state: VERIFIED, UNVERIFIED, or UNOBSERVABLE — never a silent pass.
  • The evidence string each check actually observed, not just a score.
  • A remediation note per finding, and the standards crosswalk mapped to your results.
  • Where you sit against the 30-site cohort — percentile, not just a raw number.

Common questions

Does the audit tool see our source code?

The public-page benchmark only fetches the one page you give it. The governance audit's local scanner runs on your own machine against your own source tree and never leaves it.

Are execution receipts signed?

No. A receipt carries a SHA-256 integrity checksum — it detects an edited record but does not prove who produced it. Signing is planned, not shipped.

What do VERIFIED, UNVERIFIED and UNOBSERVABLE mean?

Evidence was observed and satisfies the check (VERIFIED); evidence was observed and does not (UNVERIFIED); or the evidence needed to decide is absent, which is never scored as a pass (UNOBSERVABLE).

Proof, not assertion

The Agent Readiness Registry.

The same engine, pointed outward: 39 real organisations' public pages scored. The naming rule was fixed before any organisation was scored — only scores of 70 or above are named, so the page can't read as cherry-picking. 3 of the 39 currently clear it, under 8%: by cohort, that's 20% of AI Security and Governance Tools and 8.3% of AI Agent Vendors.

  • Credo AI
  • Arize AI
  • Mistral AI

See the full registry, every cohort, every score →

Standards and method

Mapped to the frameworks auditors actually ask about.

FrameworkOfficial referenceMapped on this site
OWASP ASI Top 10genai.owasp.orgView mapping →
NIST AI RMFnist.govView mapping →
ISO/IEC 42001iso.orgView mapping →
EU AI Acteur-lex.europa.euView mapping →

The rubric is published and versioned — methodology — and the vectors behind it are runnable by anyone with a Python interpreter, against a fixed clean/dirty fixture pair with a fixed expected outcome. Nothing here asks to be taken on faith.

See the full standards crosswalk →  ·  Run the conformance vectors yourself →

Cryptographic agility

A governance question, not a product claim.

Evidence about agent actions gets retained for years, so whatever eventually signs it has to be replaceable before that signature scheme is broken — algorithm agility is the requirement, decided long before the migration is urgent.

ApexClaw does not sign receipts today. Current integrity comes from a SHA-256 checksum only — it detects an edited record, it does not authenticate who produced it. Signing, when it ships, will be designed for algorithm agility from day one rather than retrofitted under deadline.

Read the full position →

Services

Assessment engagements, scoped and delivered directly.

For teams that want the audit run against their own estate, not just a public page — inventory, permissions, MCP exposure, evidence, revocation testing. ApexClaw is the company entity behind this site: ApexClaw on LinkedIn.

Talk to us →