Evidence ladder
What exists today, in order of claim strength.
Each rung answers a different question. Implementation evidence shows
the mechanism exists. Retrospective evidence calibrates it. Prospective
evidence shows the public runtime seam works. Six comparative studies
show the progression from timeout sensitivity and escalation regressions
to a passed, narrowly calibrated hybrid support envelope.
01 / PRODUCT
Durable operating layer
/do, repository state, campaigns, fleets, recovery, evidence, and handoffs work around the coding agent a developer already uses.
Implemented
02 / HISTORY
Signed 120-cell matrix
Ten frozen scenarios across three repositories and four economic policies preserve 33 verified, 51 failed, and 36 unknown outcomes.
Verified
03 / STACK
Sentient ROMA binding
A thin adapter controlled a pinned recursive solver module by module. Its 24-cell diagnostic passed the evidence gate and failed the efficiency hypothesis.
Verified
04 / RUNTIME
Prospective public task
One preregistered Claude Code operation on a fresh public clone matched model and topology, changed the required artifact, and passed a deterministic verifier outside the model.
Verified
05 / ECONOMICS V1
Adaptive local comparison
Across 12 tasks × 2 policies × 3 timing repetitions, adaptive recorded 27/36 verified cells versus 24/36. The frozen aggregate used 9.9% less GPU energy, but excluding one same-route timeout pair reverses the comparison to 3.5% more.
Verified negative gate
06 / ECONOMICS V2
Capability-profile falsification
A separately frozen 72-cell follow-up matched 24/36 baseline cell completion, but 12 escalations caused 15.7% more GPU energy and 16.4% more modeled GPU cost.
Verified regression
07 / REPOSITORY
Representative fixture shakedown
Six artifact-producing fixture tasks, two policies, and two timing repetitions produced 24 signed cells. Both policies verified 6/12; zero false passes and path violations survived replay, while the 7.1% energy reduction missed the frozen 20% gate.
Integrity passed; economics failed
08 / HYBRID CALIBRATION
Valid baseline, narrow miss
Claude and Citadel each verified 12/12 fresh tasks. Citadel avoided four Claude calls and reduced comparison cost 28.4%, missing the frozen 30% gate.
Quality passed; economics failed
09 / HYBRID V2
Calibrated support envelope
On twelve new tasks, both policies verified 12/12. Citadel used local 3B eight times, recovered once, reduced Claude calls from twelve to five, and reduced comparison cost 38.7%.
Every frozen gate passed
10 / PUBLIC HOLDOUT
Outside-authored baseline falsification
Twenty-four distinct repositories produced eight calibration and sixteen untouched evaluation tasks. Direct Claude verified 2/16 and the sealed controller 3/16 at 1.26% lower comparison cost. The baseline was too weak for a general result.
Evidence complete; baseline invalid