Citadel is building the evidence layer between an optimization policy
and its claim. It controls what an agent operation may run, records
what actually ran, and asks a deterministic verifier outside the routed model whether the work
counts. A cheaper failure is not a saving. An unknown is not zero.
Optimize the whole operation, then prove the outcome.
Prompt routing is only one decision. Real agent economics also depend
on topology, decomposition, retries, tools, timeouts, local versus
hosted execution, verification, and recovery. Citadel turns those
choices into a bounded contract whose result can be reproduced and
rejected.
What is different
The optimizer does not grade its own homework.
A selected executor cannot become the winner because it sounded
confident, reported a low number, or produced a patch. Requested
and observed runtime facts are reconciled, required artifacts are
checked, and a model-external repository verifier decides completion.
Funded target
A result strong enough to survive a no.
Across a preregistered multi-stack benchmark, the funded controller targets:
≥80%absolute verified completion
≥95%of a valid frontier baseline
≥30%lower measured end-to-end cost
Frontier must first verify at least 80% overall and 70% in every frozen task stratum, or the comparison is baseline-invalid.
Evidence ladder
What exists today, in order of claim strength.
Each rung answers a different question. Implementation evidence shows
the mechanism exists. Retrospective evidence calibrates it. Prospective
evidence shows the public runtime seam works. Three comparative studies
now expose a timeout-sensitive apparent gain, a policy regression, and
repository-artifact integrity without laundering any into a success claim.
01 / PRODUCT
Durable operating layer
/do, repository state, campaigns, fleets, recovery, evidence, and handoffs work around the coding agent a developer already uses.
Implemented
02 / HISTORY
Signed 120-cell matrix
Ten frozen scenarios across three repositories and four economic policies preserve 33 verified, 51 failed, and 36 unknown outcomes.
Verified
03 / STACK
Sentient ROMA binding
A thin adapter controlled a pinned recursive solver module by module. Its 24-cell diagnostic passed the evidence gate and failed the efficiency hypothesis.
Verified
04 / RUNTIME
Prospective public task
One preregistered Claude Code operation on a fresh public clone matched model and topology, changed the required artifact, and passed a deterministic verifier outside the model.
Verified
05 / ECONOMICS V1
Adaptive local comparison
Across 12 tasks × 2 policies × 3 timing repetitions, adaptive recorded 27/36 verified cells versus 24/36. The frozen aggregate used 9.9% less GPU energy, but excluding one same-route timeout pair reverses the comparison to 3.5% more.
Verified negative gate
06 / ECONOMICS V2
Capability-profile falsification
A separately frozen 72-cell follow-up matched 24/36 baseline cell completion, but 12 escalations caused 15.7% more GPU energy and 16.4% more modeled GPU cost.
Verified regression
07 / REPOSITORY
Representative fixture shakedown
Six artifact-producing fixture tasks, two policies, and two timing repetitions produced 24 signed cells. Both policies verified 6/12; zero false passes and path violations survived replay, while the 7.1% energy reduction missed the frozen 20% gate.
Integrity passed; economics failed
Funded work
Four milestones, each with a failure condition.
The grant does not fund a promise to make Citadel look intelligent.
It funds a public research program whose method, artifacts, cost
lenses, and negative results remain inspectable whether the final
performance target passes or fails.
Controller gate
Learn operation value
Choose models, topology, decomposition, retries, tools, and stopping from calibrated outcome evidence instead of prompt difficulty alone.
Exit: held-out decisions are reproducible and every rejected or unknown outcome remains in the ledger.
Economic gate
Measure the complete cost
Separate reported, derived, market-equivalent, marginal, and end-to-end cost. Include local hardware and energy without converting missing evidence to zero.
Exit: every published comparison carries a complete cost basis or is explicitly blocked.
Portability gate
Generalize the adapter
Exercise the same operation contract across multiple open and proprietary agent stacks without replacing their native planning or execution logic.
Exit: requested versus observed identity, artifacts, and outcome receipts reconcile across each stack.
Scale gate
Run the prospective study
Extend the three local studies to multiple stacks, repositories, model families, hardware profiles, and tool routes. Freeze every task, baseline, policy, verifier, stopping rule, and economic target before execution.
Exit: the result reproduces offline and in clean hosted verification, even if the performance hypothesis fails.
Claim boundary
Credibility comes from saying exactly where the proof ends.
Citadel has enough evidence to show that its controller can be held
accountable to economic claims. It does not yet have enough evidence
to claim the funded multi-stack result. The site and repository use
the same boundary.
Demonstrated
Stack-neutral operation contracts can bind real agent runtimes.
Requested and observed model identity can be reconciled without inference.
Model-external deterministic verification can reject false, partial, and adversarial outputs.
Signed evidence can preserve passed, failed, and unknown outcomes.
Clean hosted verification reproduces the committed proof bundle.
Three prospective local studies reject a timeout-sensitive apparent gain, a cost-increasing capability policy, and a representative pilot that missed its economic gates.
Not demonstrated
Best-in-class agent performance or broad benchmark leadership.
A general quality advantage across repositories or model families.
Lower latency than direct or frontier-only execution.
Thirty percent lower end-to-end cost at the quality target.
Production reliability across many external users and environments.
Why Sentient
A native fit for token and economic optimization.
Sentient's product request asks for a layer that can optimize agents
across models and workflows. Citadel contributes the operation-level
control and proof discipline needed to tell a genuine economic win
from a cheaper failed attempt. The first ROMA adapter makes that fit
concrete rather than hypothetical.