Skip to content
Evaluator evidence index

Inspect the claim.
Then try to break it.

Citadel binds an operation's declared route to what actually ran, its measured and modeled cost lenses, and whether a deterministic verifier outside the routed model accepted the outcome. Failed policies stay failed. Unknown cost stays unknown.

Six separately frozen studies

The failures changed the policy.

Local studies exposed timeout, escalation, and baseline-validity defects. A calibrated synthetic support envelope passed. The outside-authored follow-up then showed that the retrieval/edit substrate and strong baseline were not ready.

V1 · adaptive local

More verified cells; savings not robust.

Across 12 tasks × 2 policies × 3 timing repetitions, adaptive recorded 27/36 verified cells versus 24/36. Its frozen aggregate missed both 30% gates, and excluding one matched 60-second baseline timeout reverses all three economic comparisons.

27/36adaptive verified
+3.5%GPU energy in sensitivity
+5.4%modeled GPU cost in sensitivity
Method + signed result
V2 · capability profile

Same cell completion, economic regression.

Twelve new exact instances, mostly from task templates already seen in v1, routed work to 1.5B, 3B, or 7B. Model-external verification matched the baseline cell completion at higher measured GPU energy and modeled GPU cost.

24/36both policies verified
+15.7%GPU energy
12strong escalations
Signed result Required disclosure
V3 · repository operations

Artifact integrity passed; economics failed.

Six fixture repositories × 2 policies × 2 timing repetitions. Both policies verified 6/12 cells. Citadel used 7.1% less measured GPU energy, below the frozen 20% gate, and 13.2% more tokens.

6/12both policies verified
0false passes or path violations
Failedfrozen evidence result
Method + signed result
V4 · hybrid calibration

Quality valid; economic gate missed narrowly.

Both policies verified 12/12 fresh tasks. The risk-only policy avoided four Claude calls but made four unsupported local attempts, reducing comparison cost 28.4% against a frozen 30% gate.

12/12both policies verified
28.4%comparison-cost reduction
Failedfrozen economic result
Calibration report
V5 · calibrated hybrid v2

Same verified outcomes; bounded economic gate passed.

Twelve new tasks, a valid Claude Sonnet 5 baseline, and a preregistered support envelope. Citadel used eight local attempts, recovered once, reduced Claude calls from twelve to five, and reduced comparison cost 38.7%.

12/12both policies verified
38.7%comparison-cost reduction
Passedevery frozen gate
Passed method + signed result
V6 · outside-authored public holdout

Evidence sequence passed; optimization claim did not.

Twenty-four distinct repositories supplied eight calibration and sixteen untouched evaluation tasks. Routes were published before calls. Qwen verified 1/16, direct Claude 2/16, and the Qwen-first controller 3/16 at 1.26% lower comparison cost. The baseline was too weak for a general result.

24distinct repositories
32/32official verdicts
Invalidstrong-baseline claim
Final report Validation
Evidence ladder

One index. No scavenger hunt.

Each artifact answers a different question. The state at right is the gate's actual outcome, not a maturity badge.

01 / HISTORY

120-cell operation matrix

All 120 signed cells preserve 33 verified, 51 failed, and 36 unknown outcomes; 84 cells reached a model and 36 remained setup-unknown.

Integrity passed
02 / ROMA

Sentient stack binding

A pinned recursive stack consumed the operation contract; control evidence passed and the efficiency hypothesis failed.

Policy failed
03 / RUNTIME

Prospective Claude operation

Requested and observed model/topology matched, the public clone changed as required, and a deterministic repository verifier outside the model passed.

Integration passed
04 / LOCAL V1

Adaptive local calibration

12 tasks × 2 policies × 3 timing repetitions; 27/36 versus 24/36 verified cells; frozen economic gates failed and timeout sensitivity reversed the economic direction.

Gate failed
05 / LOCAL V2

Capability-profile follow-up

12 exact task instances × 2 policies × 3 timing repetitions; matched baseline cell completion; verifier escalation made measured GPU economics worse.

Policy regressed
06 / REPOSITORY

Representative fixture shakedown

Six artifact-producing tasks × 2 policies × 2 timing repetitions; both policies verified 6/12 cells; integrity gates passed and economic gates failed.

Gate failed
07 / HYBRID CAL

Claude plus local calibration

Both policies verified 12/12 fresh tasks; four Claude calls were avoided; 28.4% comparison-cost reduction missed the frozen 30% gate.

Gate missed
08 / HYBRID V2

Calibrated support envelope

Both policies verified 12/12 new tasks; Citadel reduced Claude calls from twelve to five and comparison cost 38.7% with every gate passed.

All gates passed
09 / ONBOARD

Fresh-clone governed path

Five unattended engineering stages completed from a clean local clone in 28.17 seconds; the nested doctor command exited zero but reported semantic health as unknown.

5 stages completed
10 / PUBLIC HOLDOUT

Outside-authored route diagnostic

24 distinct repositories, 16 sealed evaluation routes, 32 official verdicts; direct Claude 2/16 and controller 3/16. Baseline validity failed for a general claim.

Baseline invalid
Claim boundary

What the work permits us to say.

Citadel is not presented as a universally best-in-class model router. It now demonstrates a narrower, useful result: an operation policy can be controlled, observed, graded outside the routed model, rejected when it fails, and accepted when it preserves verified outcomes while clearing a frozen economic gate.

Demonstrated

  • One contract has external-stack adoption evidence in ROMA and contract-layer runtime coverage in Claude Code and Ollama.
  • Requested and observed runtime identity can be reconciled.
  • Deterministic verification outside the routed model rejects failed and adversarial answers.
  • Signed chains preserve every passed, failed, and unknown cell.
  • A plausible policy can be shown to cost more before deployment.
  • On twelve author-selected tasks inside a declared support envelope, Citadel preserved 12/12 verified outcomes while reducing Claude calls from twelve to five and comparison cost by 38.7%.
  • On 24 distinct outside-authored repositories, Citadel sealed routes before model calls and published all 32 official evaluation verdicts.
  • The outside-authored result rejected its own optimization headline because direct Claude verified only 2/16, despite the controller's 3/16 and 1.26% lower comparison cost.
  • Repository changes can be checked for exact artifacts and allowed-path boundaries without trusting the model's own report.

Still open

  • Thirty percent lower actual end-to-end cash cost across outside-authored production tasks with a valid strong baseline.
  • Generalization across repositories, agent stacks, model families, and hardware.
  • Actual cash accounting where subscription allocation or whole-system energy is absent.
  • Learned operation-value prediction that prices likely escalation and recovery.
  • Equivalent prospective actual-run evidence for every supported adapter.
Offline verification

Recompute the proof without asking a model.

This command runs seventeen offline checks across the retained failed calibrations, the passed hybrid result, the public holdout, routes, exact-answer verdicts, model identity, cost derivations, source bindings, artifact digests, receipt chains, Ed25519 signatures, public claims, the application package, and the site story.

Evaluator guide
npm run grant:verify
What the Sentient grant buys

Learn operation value, then test it in public.

Funding expands the controller from frozen local pilots to multiple open agent stacks, task categories, model families, hardware profiles, tool routes, and complete cost lenses. The method and negative results remain public either way. Frontier must first verify at least 80% overall and 70% in every frozen task stratum, or the comparison is invalid.

≥80%absolute verified completion
≥95%of a valid frontier baseline
≥30%lower measured end-to-end cost