Skip to content
Evaluator evidence index

Inspect the claim.
Then try to break it.

Citadel binds an operation's declared route to what actually ran, its measured and modeled cost lenses, and whether a deterministic verifier outside the routed model accepted the outcome. Failed policies stay failed. Unknown cost stays unknown.

Three separately frozen studies

The evidence changed the plan.

V1's aggregate was timeout-sensitive. V2 showed that verification escalation can cost more. The representative shakedown proved repository-artifact replay while still missing its frozen economic gates.

V1 · adaptive local

More verified cells; savings not robust.

Across 12 tasks × 2 policies × 3 timing repetitions, adaptive recorded 27/36 verified cells versus 24/36. Its frozen aggregate missed both 30% gates, and excluding one matched 60-second baseline timeout reverses all three economic comparisons.

27/36adaptive verified
+3.5%GPU energy in sensitivity
+5.4%modeled GPU cost in sensitivity
Method + signed result
V2 · capability profile

Same cell completion, economic regression.

Twelve new exact instances, mostly from task templates already seen in v1, routed work to 1.5B, 3B, or 7B. Model-external verification matched the baseline cell completion at higher measured GPU energy and modeled GPU cost.

24/36both policies verified
+15.7%GPU energy
12strong escalations
Signed result Required disclosure
V3 · repository operations

Artifact integrity passed; economics failed.

Six fixture repositories × 2 policies × 2 timing repetitions. Both policies verified 6/12 cells. Citadel used 7.1% less measured GPU energy, below the frozen 20% gate, and 13.2% more tokens.

6/12both policies verified
0false passes or path violations
Failedfrozen evidence result
Method + signed result
Evidence ladder

One index. No scavenger hunt.

Each artifact answers a different question. The state at right is the gate's actual outcome, not a maturity badge.

Claim boundary

What the work permits us to say.

Citadel is not presented as a best-in-class model router. Its demonstrated advantage is that operation policies can be controlled, observed, graded outside the routed model, and rejected.

Demonstrated

  • One contract has external-stack adoption evidence in ROMA and contract-layer runtime coverage in Claude Code and Ollama.
  • Requested and observed runtime identity can be reconciled.
  • Deterministic verification outside the routed model rejects failed and adversarial answers.
  • Signed chains preserve every passed, failed, and unknown cell.
  • A plausible policy can be shown to cost more before deployment.
  • Repository changes can be checked for exact artifacts and allowed-path boundaries without trusting the model's own report.

Still open

  • Thirty percent lower end-to-end cost at the funded quality target.
  • Generalization across repositories, agent stacks, model families, and hardware.
  • Actual cash accounting where subscription allocation or whole-system energy is absent.
  • Learned operation-value prediction that prices likely escalation and recovery.
  • Equivalent prospective actual-run evidence for every supported adapter.
Offline verification

Recompute the proof without asking a model.

These commands recompute routes, exact-answer verdicts, model identity, cost derivations, source bindings, artifact digests, receipt chains, and Ed25519 signatures from committed evidence.

Evaluator guide
npm run readiness:verify
npm run readiness:v2:verify
npm run representative:v2:verify
npm run operation-proof:verify
npm run application:evidence:check
npm run onboarding:fresh-clone:verify
What the Sentient grant buys

Learn operation value, then test it in public.

Funding expands the controller from frozen local pilots to multiple open agent stacks, task categories, model families, hardware profiles, tool routes, and complete cost lenses. The method and negative results remain public either way. Frontier must first verify at least 80% overall and 70% in every frozen task stratum, or the comparison is invalid.

≥80%absolute verified completion
≥95%of a valid frontier baseline
≥30%lower measured end-to-end cost