The agent governance evidence report, 2026
Original first-party research in two parts: what AI answer engines actually tell buyers about this category on a dated capture, and what ApexClaw's own governance produced on a live autonomous system - refusals, receipts, a gauntlet PASS, and an honest not-yet-working list.
Get an Agent Trust Gap BriefSummary
This is original, first-party research, run on ApexClaw itself, in two parts. Part one measures what AI answer engines told buyers about the agent-governance category on the capture date. Part two documents what running our own governance against a live autonomous system produced - refusals, receipts, and the parts not yet working. Both are dated and labelled; neither is a claim about anyone else's estate.
FIRST-PARTY. Every figure below is our own measurement, captured 2026-08-13, expiring for re-measurement 2026-11-13. Method stated inline. A past measurement is not proof of present state.
Part 1 - What the engines actually answered
Method: on the capture date we posed two prompt classes to the public web-retrieval layer that AI answer engines draw from - (a) buyer-diligence prompts ("is ApexClaw legitimate", "ApexClaw AI agent governance") and (b) the category shortlist prompt ("agent governance vendors compared 2026"). We recorded which entities the retrieval layer surfaced.
Finding A - the branded query resolves, but the name-space is contested
ApexClaw's own pages were retrievable and served as answer substrate on branded and category prompts. The same retrieval also mixed in unrelated same-string entities (an Oracle database feature, an unrelated "Apex" product), confirming that entity disambiguation - stating "ApexClaw, the AI Agent Trust Layer" identically everywhere - is a live requirement, not a nicety.
Finding B - absent from the shortlist that decides deals
On the category shortlist prompt, the retrieval layer surfaced an established set of governance names and did not surface ApexClaw. This is the single most consequential measured fact in this report: the buyer's shortlist is being assembled by models, and on the capture date it was assembled without us. Every item in the "what changes the answer" section below targets that gap directly.
| Prompt class | What the retrieval layer surfaced (capture date) | ApexClaw present? |
|---|---|---|
| Branded diligence | ApexClaw's own pages, plus same-string unrelated entities | Yes (with disambiguation risk) |
| Category shortlist | An established set of governance names | No |
What changes the answer (and why it is defensible)
Answer presence is downstream of retrievable, corroborated, machine-followable evidence. The moves that shift it are structural, not promotional: an entity stated identically across independent surfaces; comparison pages that answer the shortlist prompt directly; an open, verifiable evidence standard; and original research - this report - that other sources can cite. Each is being executed and dated, so the re-measurement in November is a real before-and-after, not a story.
Part 2 - Running our own governance, failures included
ApexClaw operates an autonomous internal system (Omega) under ApexClaw's own gate architecture. This is what that produced, stated in the same restricted vocabulary the system is held to.
| Observation | Status | What it demonstrates |
|---|---|---|
| Unauthorized external sends refused, receipt emitted for each | PROVEN (first-party) | The gate refuses actions lacking satisfied human approval, and records the refusal |
| Zero real external sends across full history | PROVEN (network-layer verified) | The control holds at the effect, not just in application logs |
| Acceptance gauntlet - core governance laws | PASS 2026-08-13 | Role authority, tool authority, approval-object requirement, canonical-owner effect path, and refusal-on-invalid-auth all verified by the suite |
| Full approved-action success cycle end to end | NOT YET RUN | Published as not-done rather than implied - the success path with human approval has not completed |
| Brain service | MASKED (deliberate) | Stopped on purpose, not crashed - stated plainly |
| Signature scheme | Ed25519 (not post-quantum) | Named as a known limitation, not hidden |
Publishing the NOT-YET-RUN and MASKED rows is the point. The governance claim under test is narrow and survives every limitation above: actions are refused unless authorized, and there is a receipt either way. None of the failures is a claim we made and missed - which is exactly the discipline the evidence layer exists to enforce. Full detail: the governed run.
Why this report exists, and what it is not
It is not a polished case study - those prove nothing. It is the two things a trust vendor can actually stand behind: a dated measurement of the market's AI answers, and its own governance running with the failures visible. The method in Part 1 is offered as a service (the AI Search Visibility Audit); the discipline in Part 2 is the product (the agent trust layer). This page is the reference implementation of both, run on ourselves.
First-party research. Captured 2026-08-13, re-measurement due 2026-11-13. Not legal advice, not independently audited, not a claim about any other organisation. Method stated inline; see the claims policy.