Citadel Optimizer chooses and revises the path through models,
agents, topology, tools, retries, and stopping. Receipts and
deterministic verifiers outside the routed model prevent cheap failure from masquerading as savings.
10frozen scenarios
3public repositories
4economic policies
120signed operation-control cells
Prospective local results
A controller must price verification and recovery.
All three studies were frozen before execution under their published identities. V1 fails a timeout sensitivity. V2 shows why model-size routing alone is insufficient. V3 extends the verifier to repository artifacts and misses its economic gates.
V1 · more verified cells, savings not robust
12 tasks × 2 policies × 3 timing repetitions.
Adaptive: 27/36 verified cells; always-7B: 24/36.
Frozen aggregate: 9.9% less GPU energy and 10.3% less modeled GPU cost.
Excluding one same-route 60-second timeout pair: 3.5% more energy and 5.4% more modeled GPU cost.
The benchmark holds task and verification constant while changing
the economic policy. Explore the same Nano ID task under each policy.
The adaptive view uses a deterministic fixture probe here so the UI
remains a demo, not a hidden model call. The actual Nano ID cells
failed setup before any model ran and are not presented as results.
nanoid-size-consistencyfrontier tiersingle agent
Probe, then reserve frontier capacity
Repository evidence shows the change spans secure, browser, non-secure,
tests, types, and docs. The cheap path is not plausible enough, so the
controller chooses frontier before a failed attempt.
03 / ROUTE
Frontier profilePrediction source is visibly a policy assumption.
04 / CONTROL
Verify or stopNo completion means no successful economic result.
Claim discipline
Two ledgers: what exists, what is proven.
Citadel keeps implementation maturity separate from performance
evidence. That boundary is the difference between a credible grant
application and a polished repo making claims it cannot defend.
Implemented and locally verified
Strict cost provenance with unknown never coerced to zero.
Prompt-only and adaptive policy contracts.
Bounded read-only repository reconnaissance.
Continue, escalate, split, and stop decisions.
Holdout rejection during capability learning.
Signed-run tamper rejection and adversarial gates.
4/4 exact model, receipt, and cost-source calibration gates passed; 0/4 original task verifiers passed.
No-model forensics found the original verifier failed before task tests; the matrix now uses a task-focused verifier proven against the bug and a reference repair.
Future failed attempts retain bounded, path- and secret-redacted verifier and patch receipts.
120/120 signed matrix cells are present; 84 reached a model and their actual-run attestations verify.
The engineering gate passed with zero adversarial false passes.
The committed proof bundle passed its clean GitHub-hosted verification job.
Not demonstrated
Any matrix-level cost reduction.
Adaptive routing better than prompt-only routing.
A setup-complete benchmark across all three repositories.
A prospective multi-stack comparison that meets the cost and completion target.
Usefulness beyond the three frozen repositories.
Fixture math exercises a 23.2507% held-out median-cost reduction with no
simulated completion loss. This validates report behavior only. It is
not a result, benchmark score, or grant claim.
Open research gates
The remaining work is value prediction, not another label.
The matrix is complete, but its precommitted performance conditions
are not. A fixture, partial cost ledger, or relabeled failure cannot
turn this result into a savings claim.
ECONOMIC_TARGET_OPENNo frozen local study has reached its preregistered economic target.
ESCALATION_VALUE_REQUIREDThe controller must predict whether a cheaper attempt plus likely verifier escalation costs more than selecting the strong path first.
MULTI_STACK_GENERALIZATION_OPENLocal Qwen exact-answer tasks do not establish performance across repositories, agent stacks, model families, or hardware.
ACTUAL_CASH_UNKNOWNWhole-system energy, setup cost, and subscription allocation remain unknown and are not converted to zero.
The grant thesis
Make open agents cheap by default without grading their own homework.
This directly targets Sentient Foundation's Token and Economic
Optimization for Agents request. Citadel is not claiming to be best
in class: its first controller failed the performance gate. The grant
thesis is an open controller and proof discipline that preserves
negative results, repairs the benchmark under a new identity, and can
establish whether whole-operation optimization works.