Skip to content
Open agent optimization research

Make agent optimization falsifiable.

Citadel is building the evidence layer between an optimization policy and its claim. It controls what an agent operation may run, records what actually ran, and asks a deterministic verifier outside the routed model whether the work counts. A cheaper failure is not a saving. An unknown is not zero.

240signed prospective comparison cells
24outside-authored repositories
32/32official holdout verdicts
38.7%bounded synthetic reduction
The research thesis

Optimize the whole operation, then prove the outcome.

Prompt routing is only one decision. Real agent economics also depend on topology, decomposition, retries, tools, timeouts, local versus hosted execution, verification, and recovery. Citadel turns those choices into a bounded contract whose result can be reproduced and rejected.

What is different

The optimizer does not grade its own homework.

A selected executor cannot become the winner because it sounded confident, reported a low number, or produced a patch. Requested and observed runtime facts are reconciled, required artifacts are checked, and a model-external repository verifier decides completion.

Funded target

A result strong enough to survive a no.

Across a preregistered multi-stack benchmark, the funded controller targets:

≥80%absolute verified completion
≥95%of a valid frontier baseline
≥30%lower measured end-to-end cost

Frontier must first verify at least 80% overall and 70% in every frozen task stratum, or the comparison is baseline-invalid.

Evidence ladder

What exists today, in order of claim strength.

Each rung answers a different question. Implementation evidence shows the mechanism exists. Retrospective evidence calibrates it. Prospective evidence shows the public runtime seam works. Six comparative studies show the progression from timeout sensitivity and escalation regressions to a passed, narrowly calibrated hybrid support envelope.

01 / PRODUCT

Durable operating layer

/do, repository state, campaigns, fleets, recovery, evidence, and handoffs work around the coding agent a developer already uses.

Implemented
02 / HISTORY

Signed 120-cell matrix

Ten frozen scenarios across three repositories and four economic policies preserve 33 verified, 51 failed, and 36 unknown outcomes.

Verified
03 / STACK

Sentient ROMA binding

A thin adapter controlled a pinned recursive solver module by module. Its 24-cell diagnostic passed the evidence gate and failed the efficiency hypothesis.

Verified
04 / RUNTIME

Prospective public task

One preregistered Claude Code operation on a fresh public clone matched model and topology, changed the required artifact, and passed a deterministic verifier outside the model.

Verified
05 / ECONOMICS V1

Adaptive local comparison

Across 12 tasks × 2 policies × 3 timing repetitions, adaptive recorded 27/36 verified cells versus 24/36. The frozen aggregate used 9.9% less GPU energy, but excluding one same-route timeout pair reverses the comparison to 3.5% more.

Verified negative gate
06 / ECONOMICS V2

Capability-profile falsification

A separately frozen 72-cell follow-up matched 24/36 baseline cell completion, but 12 escalations caused 15.7% more GPU energy and 16.4% more modeled GPU cost.

Verified regression
07 / REPOSITORY

Representative fixture shakedown

Six artifact-producing fixture tasks, two policies, and two timing repetitions produced 24 signed cells. Both policies verified 6/12; zero false passes and path violations survived replay, while the 7.1% energy reduction missed the frozen 20% gate.

Integrity passed; economics failed
08 / HYBRID CALIBRATION

Valid baseline, narrow miss

Claude and Citadel each verified 12/12 fresh tasks. Citadel avoided four Claude calls and reduced comparison cost 28.4%, missing the frozen 30% gate.

Quality passed; economics failed
09 / HYBRID V2

Calibrated support envelope

On twelve new tasks, both policies verified 12/12. Citadel used local 3B eight times, recovered once, reduced Claude calls from twelve to five, and reduced comparison cost 38.7%.

Every frozen gate passed
10 / PUBLIC HOLDOUT

Outside-authored baseline falsification

Twenty-four distinct repositories produced eight calibration and sixteen untouched evaluation tasks. Direct Claude verified 2/16 and the sealed controller 3/16 at 1.26% lower comparison cost. The baseline was too weak for a general result.

Evidence complete; baseline invalid
Funded work

Four milestones, each with a failure condition.

The grant does not fund a promise to make Citadel look intelligent. It funds a public research program whose method, artifacts, cost lenses, and negative results remain inspectable whether the final performance target passes or fails.

Controller gate

Learn operation value

Choose models, topology, decomposition, retries, tools, and stopping from calibrated outcome evidence instead of prompt difficulty alone.

Exit: held-out decisions are reproducible and every rejected or unknown outcome remains in the ledger.
Economic gate

Measure the complete cost

Separate reported, derived, market-equivalent, marginal, and end-to-end cost. Include local hardware and energy without converting missing evidence to zero.

Exit: every published comparison carries a complete cost basis or is explicitly blocked.
Portability gate

Generalize the adapter

Exercise the same operation contract across multiple open and proprietary agent stacks without replacing their native planning or execution logic.

Exit: requested versus observed identity, artifacts, and outcome receipts reconcile across each stack.
Scale gate

Run the prospective study

Extend the three local studies to multiple stacks, repositories, model families, hardware profiles, and tool routes. Freeze every task, baseline, policy, verifier, stopping rule, and economic target before execution.

Exit: the result reproduces offline and in clean hosted verification, even if the performance hypothesis fails.
Claim boundary

Credibility comes from saying exactly where the proof ends.

Citadel has a positive synthetic result and a later outside-authored diagnostic that invalidated its baseline. The controller can be held accountable when a policy or comparison fails. It does not yet have the complete-cost, strong-baseline, multi-stack evidence needed for a general claim.

Demonstrated

  • Stack-neutral operation contracts can bind real agent runtimes.
  • Requested and observed model identity can be reconciled without inference.
  • Model-external deterministic verification can reject false, partial, and adversarial outputs.
  • Signed evidence can preserve passed, failed, and unknown outcomes.
  • Clean hosted verification reproduces the committed proof bundle.
  • Four retained calibration studies reject timeout-sensitive, escalation-heavy, baseline-invalid, and narrowly missed policies.
  • A separately frozen hybrid v2 preserved 12/12 completions and reduced comparison cost 38.7% inside a preregistered support envelope.
  • A later pilot sealed routes across 24 outside-authored repositories and published all 32 official evaluation verdicts.
  • The outside-authored result rejected a general claim because direct Claude verified only 2/16.

Not demonstrated

  • Best-in-class agent performance or broad benchmark leadership.
  • A general quality advantage across repositories or model families.
  • Lower latency than direct or frontier-only execution.
  • Thirty percent lower actual end-to-end cash across outside-authored production tasks with a valid strong baseline.
  • Production reliability across many external users and environments.
Why Sentient

A native fit for token and economic optimization.

Sentient's product request asks for a layer that can optimize agents across models and workflows. Citadel contributes the operation-level control and proof discipline needed to tell a genuine economic win from a cheaper failed attempt. The first ROMA adapter makes that fit concrete rather than hypothetical.