We point the same validation at ourselves.
RocSite runs an internal agent fleet that builds and maintains our own tooling. It operates under the posture we sell to everyone else: independent verification, evidence on failure, and a human on every irreversible decision. This page describes how that works, because a company that asks hospitals to audit its models should be willing to show how it audits itself.
Where the line is.
Regulated device software follows formal design control. ICH Triage and anything else on the SaMD pathway is developed under our documented design-control process with human review at every stage. The fleet accelerates the internal tooling, operations, and research infrastructure around that work. It does not write, review, or release regulated device software. If you are evaluating us as a device manufacturer, that boundary is the first thing to check, and we would rather state it before you ask.
Independent verification, not self-grading.
A deterministic gate is the only path to “done.” Every change an agent makes has to pass a test gate the agent does not control and cannot edit. There is no route to completion that runs through the agent’s own opinion of its work.
Agents never certify themselves. This is the same thesis as AI Governor, turned inward. The useful question is never whether a system can explain itself convincingly. It is whether an independent check can disprove it. A confident agent that fails the gate has produced nothing.
Evidence chains on failure.
A failed attempt writes a record, not just a retry. When work fails the gate, the agent records what it expected, what was actually measured, and which specific check failed. That record is structured, not prose.
The next attempt has to read the last one. Retries are informed by the prior evidence rather than blind repetition, they are capped, and at a fixed ceiling the work stops and escalates to a person. An agent is not permitted to keep trying indefinitely, because a loop that never terminates is indistinguishable from progress until someone checks.
Graduated authority.
Decisions are classed by how reversible they are. Reversible actions that a gate has verified run autonomously. Consequential changes run as canaries with automatic rollback. Irreversible or security-touching decisions are human-gated without exception, and no track record promotes an agent into that tier.
Trust is measured, not assumed. Authority is promoted and demoted on observed track record rather than granted once at setup. An agent that starts failing loses latitude.
Every decision goes to an append-only ledger. Records are hash-chained, so a later edit breaks the chain and is detectable. The agents being recorded cannot modify the ledger that records them. If the ledger cannot be written, the action does not run: no log, no action.
Local first, escalation by evidence.
Work runs on our own models, on our own hardware, first. That is the same offline posture behind our clinical products, applied to our own development. It is also a discipline: default capacity is something we own and can reason about.
Frontier models are an escalation tier, not the default. Escalation happens when logged evidence justifies it, and the escalation itself is part of the record.
A stop button, and a supervisor that is itself audited.
There is a kill switch. A single control halts fleet execution, and using it is logged like any other decision.
The supervisor restarts stalled work and reports what it could not fix. When a component goes quiet, that is surfaced as a human-visible item rather than absorbed silently. Silent failure is the dangerous kind: a fleet that appears healthy because nothing is complaining is the specific failure mode this is built to catch.
What a decision record contains.
Every entry in the ledger carries the same fields, so the record can be replayed by someone who was not there:
The intent. What was proposed, and by which agent.
The classification. Which authority tier it fell into, and therefore whether
a human had to approve it.
The undo. The rollback path, recorded before execution, not after.
The verification. Which independent check ran afterward, and what it
returned.
The chain. The hash of the previous entry, which is what makes the history
tamper-evident rather than merely long.
What we are not claiming.
This page describes architecture, not benchmarks. We are not claiming that local models resolve production issues unaided; that measurement is pending and we will publish it when we have it. We are not claiming an absence of humans. We publish no build-quality metrics we have not measured, on this page or anywhere else.
Autonomy here means governed autonomy. Autonomous execution with human oversight on irreversible decisions. If that distinction sounds like hedging, it is the same distinction we apply to clinical AI, and it is the reason we published a null result rather than a success story.
For technical and diligence reviewers.
If you want to walk through the gate design, the authority tiers, or the ledger format in more detail than is appropriate to publish, ask. Schedule a conversation. No form. No queue. We read everything.

Contact