Skip to content
Outcome-aware economic routing

Route by evidence. Keep unknowns honest.

Citadel Optimizer chooses and revises the path through models, agents, topology, tools, retries, and stopping. Receipts and deterministic verifiers outside the routed model prevent cheap failure from masquerading as savings.

10frozen scenarios
3public repositories
4economic policies
120signed operation-control cells
Prospective local results

A controller must price verification and recovery.

All three studies were frozen before execution under their published identities. V1 fails a timeout sensitivity. V2 shows why model-size routing alone is insufficient. V3 extends the verifier to repository artifacts and misses its economic gates.

V1 · more verified cells, savings not robust

  • 12 tasks × 2 policies × 3 timing repetitions.
  • Adaptive: 27/36 verified cells; always-7B: 24/36.
  • Frozen aggregate: 9.9% less GPU energy and 10.3% less modeled GPU cost.
  • Excluding one same-route 60-second timeout pair: 3.5% more energy and 5.4% more modeled GPU cost.
  • Failed the frozen 30% economic gates.
Inspect signed v1 result

V2 · matched cell completion, cost regressed

  • 12 task instances × 2 policies × 3 timing repetitions.
  • Capability profile: 24/36 verified cells; always-7B: 24/36.
  • 12 failed small-model answers escalated to 7B.
  • 15.7% more measured GPU energy.
  • 16.4% more modeled GPU cost; zero false passes.
Inspect signed v2 result

V3 · artifacts verified, savings gate missed

  • 6 fixture repositories × 2 policies × 2 timing repetitions.
  • Risk profile: 6/12 verified cells; always-7B: 6/12.
  • Zero false passes, path violations, or integrity failures.
  • 7.1% less measured GPU energy, below the frozen 20% gate.
  • 13.2% more tokens; no production-generalization claim.
Inspect signed v3 result
Interactive contract

The route is the product.

The benchmark holds task and verification constant while changing the economic policy. Explore the same Nano ID task under each policy. The adaptive view uses a deterministic fixture probe here so the UI remains a demo, not a hidden model call. The actual Nano ID cells failed setup before any model ran and are not presented as results.

nanoid-size-consistency frontier tier single agent

Probe, then reserve frontier capacity

Repository evidence shows the change spans secure, browser, non-secure, tests, types, and docs. The cheap path is not plausible enough, so the controller chooses frontier before a failed attempt.

01 / INPUT Frozen task Same commit, verifier, tools, and attempt limit.
02 / SIGNAL Bounded probe Six task-relevant files; strong tests; cross-cutting scope.
03 / ROUTE Frontier profile Prediction source is visibly a policy assumption.
04 / CONTROL Verify or stop No completion means no successful economic result.
Claim discipline

Two ledgers: what exists, what is proven.

Citadel keeps implementation maturity separate from performance evidence. That boundary is the difference between a credible grant application and a polished repo making claims it cannot defend.

Implemented and locally verified

  • Strict cost provenance with unknown never coerced to zero.
  • Prompt-only and adaptive policy contracts.
  • Bounded read-only repository reconnaissance.
  • Continue, escalate, split, and stop decisions.
  • Holdout rejection during capability learning.
  • Signed-run tamper rejection and adversarial gates.
  • 4/4 exact model, receipt, and cost-source calibration gates passed; 0/4 original task verifiers passed.
  • No-model forensics found the original verifier failed before task tests; the matrix now uses a task-focused verifier proven against the bug and a reference repair.
  • Future failed attempts retain bounded, path- and secret-redacted verifier and patch receipts.
  • 120/120 signed matrix cells are present; 84 reached a model and their actual-run attestations verify.
  • The engineering gate passed with zero adversarial false passes.
  • The committed proof bundle passed its clean GitHub-hosted verification job.

Not demonstrated

  • Any matrix-level cost reduction.
  • Adaptive routing better than prompt-only routing.
  • A setup-complete benchmark across all three repositories.
  • A prospective multi-stack comparison that meets the cost and completion target.
  • Usefulness beyond the three frozen repositories.
Fixture math exercises a 23.2507% held-out median-cost reduction with no simulated completion loss. This validates report behavior only. It is not a result, benchmark score, or grant claim.
Open research gates

The remaining work is value prediction, not another label.

The matrix is complete, but its precommitted performance conditions are not. A fixture, partial cost ledger, or relabeled failure cannot turn this result into a savings claim.

ECONOMIC_TARGET_OPENNo frozen local study has reached its preregistered economic target.
ESCALATION_VALUE_REQUIREDThe controller must predict whether a cheaper attempt plus likely verifier escalation costs more than selecting the strong path first.
MULTI_STACK_GENERALIZATION_OPENLocal Qwen exact-answer tasks do not establish performance across repositories, agent stacks, model families, or hardware.
ACTUAL_CASH_UNKNOWNWhole-system energy, setup cost, and subscription allocation remain unknown and are not converted to zero.
The grant thesis

Make open agents cheap by default without grading their own homework.

This directly targets Sentient Foundation's Token and Economic Optimization for Agents request. Citadel is not claiming to be best in class: its first controller failed the performance gate. The grant thesis is an open controller and proof discipline that preserves negative results, repairs the benchmark under a new identity, and can establish whether whole-operation optimization works.

Funded target

≥80%
absolute verified completion

≥30%
lower end-to-end cost

≥95%
of a valid frontier baseline

Frontier must first clear 80% overall and 70% per frozen stratum. Targets, not current results.