Citadel Optimizer chooses and revises the path through models,
agents, topology, tools, retries, and stopping. Receipts and
deterministic verifiers outside the routed model prevent cheap failure from masquerading as savings.
24outside-authored repositories
16untouched evaluation tasks
32/32official route verdicts
17offline verification checks
Prospective economic results
A controller must price verification, recovery, and its support boundary.
Every study was frozen under its published identity. V1 fails a timeout sensitivity. V2 prices escalation. V3 extends verification to repository artifacts. The first hybrid misses narrowly; calibrated hybrid v2 passes synthetically; the outside-authored follow-up invalidates its baseline.
V1 · more verified cells, savings not robust
12 tasks × 2 policies × 3 timing repetitions.
Adaptive: 27/36 verified cells; always-7B: 24/36.
Frozen aggregate: 9.9% less GPU energy and 10.3% less modeled GPU cost.
Excluding one same-route 60-second timeout pair: 3.5% more energy and 5.4% more modeled GPU cost.
The benchmark holds task and verification constant while changing
the economic policy. Explore the same Nano ID task under each policy.
The adaptive view uses a deterministic fixture probe here so the UI
remains a demo, not a hidden model call. The actual Nano ID cells
failed setup before any model ran and are not presented as results.
nanoid-size-consistencyfrontier tiersingle agent
Probe, then reserve frontier capacity
Repository evidence shows the change spans secure, browser, non-secure,
tests, types, and docs. The cheap path is not plausible enough, so the
controller chooses frontier before a failed attempt.
03 / ROUTE
Frontier profilePrediction source is visibly a policy assumption.
04 / CONTROL
Verify or stopNo completion means no successful economic result.
Claim discipline
Two ledgers: what exists, what is proven.
Citadel keeps implementation maturity separate from performance
evidence. That boundary is the difference between a credible grant
application and a polished repo making claims it cannot defend.
Implemented and locally verified
Strict cost provenance with unknown never coerced to zero.
Prompt-only and adaptive policy contracts.
Bounded read-only repository reconnaissance.
Continue, escalate, split, and stop decisions.
Holdout rejection during capability learning.
Signed-run tamper rejection and adversarial gates.
4/4 exact model, receipt, and cost-source calibration gates passed; 0/4 original task verifiers passed.
No-model forensics found the original verifier failed before task tests; the matrix now uses a task-focused verifier proven against the bug and a reference repair.
Future failed attempts retain bounded, path- and secret-redacted verifier and patch receipts.
120/120 signed matrix cells are present; 84 reached a model and their actual-run attestations verify.
The engineering gate passed with zero adversarial false passes.
The committed proof bundle passed its clean GitHub-hosted verification job.
Not demonstrated
A valid strong baseline on outside-authored work.
Useful production-level completion or cost reduction.
A learned policy that discriminates among task routes.
A prospective multi-stack comparison meeting the cost and completion target.
Actual cash savings with complete setup, tool, energy, and human cost.
Fixture math exercises a 23.2507% held-out median-cost reduction with no
simulated completion loss. This validates report behavior only. It is
not a result, benchmark score, or grant claim.
Open research gates
The remaining work is external validity and learned value prediction.
One bounded synthetic policy cleared its precommitted conditions. The
outside-authored follow-up then exposed a weak baseline and nearly
uniform route. That is a research result, not production savings.
VALID_EXTERNAL_BASELINE_OPENThe outside-authored pilot is complete, but direct Claude passed only 2/16. Retrieval, edit representation, and the strong route must clear the frozen validity floor before economic comparison.
LEARNED_VALUE_REQUIREDThe frozen support envelope must become a calibrated policy that predicts whether a cheaper attempt plus likely verifier recovery beats the strong path.
MULTI_STACK_GENERALIZATION_OPENOne Claude-plus-Qwen result does not establish performance across repositories, agent stacks, model families, or hardware.
ACTUAL_CASH_UNKNOWNWhole-system energy, setup cost, and subscription allocation remain unknown and are not converted to zero.
The grant thesis
Make open agents cheap by default without grading their own homework.
This directly targets Sentient Foundation's Token and Economic
Optimization for Agents request. Citadel is not claiming to be best
in class: several controllers failed, a bounded synthetic policy
passed, and the outside-authored follow-up invalidated its baseline.
The grant thesis is an open controller and proof discipline that
identifies exactly why an operation claim fails, then funds the
retrieval, baseline, policy, and cost work needed to retest it.