Skip to content
Open agent optimization research

Make agent optimization falsifiable.

Citadel is building the evidence layer between an optimization policy and its claim. It controls what an agent operation may run, records what actually ran, and asks a deterministic verifier outside the routed model whether the work counts. A cheaper failure is not a saving. An unknown is not zero.

168signed prospective local cells
3separately frozen studies
14verifier escalations observed
0adversarial false passes
The research thesis

Optimize the whole operation, then prove the outcome.

Prompt routing is only one decision. Real agent economics also depend on topology, decomposition, retries, tools, timeouts, local versus hosted execution, verification, and recovery. Citadel turns those choices into a bounded contract whose result can be reproduced and rejected.

What is different

The optimizer does not grade its own homework.

A selected executor cannot become the winner because it sounded confident, reported a low number, or produced a patch. Requested and observed runtime facts are reconciled, required artifacts are checked, and a model-external repository verifier decides completion.

Funded target

A result strong enough to survive a no.

Across a preregistered multi-stack benchmark, the funded controller targets:

≥80%absolute verified completion
≥95%of a valid frontier baseline
≥30%lower measured end-to-end cost

Frontier must first verify at least 80% overall and 70% in every frozen task stratum, or the comparison is baseline-invalid.

Evidence ladder

What exists today, in order of claim strength.

Each rung answers a different question. Implementation evidence shows the mechanism exists. Retrospective evidence calibrates it. Prospective evidence shows the public runtime seam works. Three comparative studies now expose a timeout-sensitive apparent gain, a policy regression, and repository-artifact integrity without laundering any into a success claim.

01 / PRODUCT

Durable operating layer

/do, repository state, campaigns, fleets, recovery, evidence, and handoffs work around the coding agent a developer already uses.

Implemented
02 / HISTORY

Signed 120-cell matrix

Ten frozen scenarios across three repositories and four economic policies preserve 33 verified, 51 failed, and 36 unknown outcomes.

Verified
03 / STACK

Sentient ROMA binding

A thin adapter controlled a pinned recursive solver module by module. Its 24-cell diagnostic passed the evidence gate and failed the efficiency hypothesis.

Verified
04 / RUNTIME

Prospective public task

One preregistered Claude Code operation on a fresh public clone matched model and topology, changed the required artifact, and passed a deterministic verifier outside the model.

Verified
05 / ECONOMICS V1

Adaptive local comparison

Across 12 tasks × 2 policies × 3 timing repetitions, adaptive recorded 27/36 verified cells versus 24/36. The frozen aggregate used 9.9% less GPU energy, but excluding one same-route timeout pair reverses the comparison to 3.5% more.

Verified negative gate
06 / ECONOMICS V2

Capability-profile falsification

A separately frozen 72-cell follow-up matched 24/36 baseline cell completion, but 12 escalations caused 15.7% more GPU energy and 16.4% more modeled GPU cost.

Verified regression
07 / REPOSITORY

Representative fixture shakedown

Six artifact-producing fixture tasks, two policies, and two timing repetitions produced 24 signed cells. Both policies verified 6/12; zero false passes and path violations survived replay, while the 7.1% energy reduction missed the frozen 20% gate.

Integrity passed; economics failed
Funded work

Four milestones, each with a failure condition.

The grant does not fund a promise to make Citadel look intelligent. It funds a public research program whose method, artifacts, cost lenses, and negative results remain inspectable whether the final performance target passes or fails.

Controller gate

Learn operation value

Choose models, topology, decomposition, retries, tools, and stopping from calibrated outcome evidence instead of prompt difficulty alone.

Exit: held-out decisions are reproducible and every rejected or unknown outcome remains in the ledger.
Economic gate

Measure the complete cost

Separate reported, derived, market-equivalent, marginal, and end-to-end cost. Include local hardware and energy without converting missing evidence to zero.

Exit: every published comparison carries a complete cost basis or is explicitly blocked.
Portability gate

Generalize the adapter

Exercise the same operation contract across multiple open and proprietary agent stacks without replacing their native planning or execution logic.

Exit: requested versus observed identity, artifacts, and outcome receipts reconcile across each stack.
Scale gate

Run the prospective study

Extend the three local studies to multiple stacks, repositories, model families, hardware profiles, and tool routes. Freeze every task, baseline, policy, verifier, stopping rule, and economic target before execution.

Exit: the result reproduces offline and in clean hosted verification, even if the performance hypothesis fails.
Claim boundary

Credibility comes from saying exactly where the proof ends.

Citadel has enough evidence to show that its controller can be held accountable to economic claims. It does not yet have enough evidence to claim the funded multi-stack result. The site and repository use the same boundary.

Demonstrated

  • Stack-neutral operation contracts can bind real agent runtimes.
  • Requested and observed model identity can be reconciled without inference.
  • Model-external deterministic verification can reject false, partial, and adversarial outputs.
  • Signed evidence can preserve passed, failed, and unknown outcomes.
  • Clean hosted verification reproduces the committed proof bundle.
  • Three prospective local studies reject a timeout-sensitive apparent gain, a cost-increasing capability policy, and a representative pilot that missed its economic gates.

Not demonstrated

  • Best-in-class agent performance or broad benchmark leadership.
  • A general quality advantage across repositories or model families.
  • Lower latency than direct or frontier-only execution.
  • Thirty percent lower end-to-end cost at the quality target.
  • Production reliability across many external users and environments.
Why Sentient

A native fit for token and economic optimization.

Sentient's product request asks for a layer that can optimize agents across models and workflows. Citadel contributes the operation-level control and proof discipline needed to tell a genuine economic win from a cheaper failed attempt. The first ROMA adapter makes that fit concrete rather than hypothetical.