Skip to content
Outcome-aware operation control

Control the operation. Prove what ran.

Citadel binds the requested model, topology, tools, worktree, artifact scope, receipts, and verifier into one operation. The agent does the work. A separate contract decides what the result proves.

120signed historical cells
3preregistered runtime attempts
1model-externally verified runtime result
0adversarial false passes
Current proof · prospective v2

One real task passed. The failures stayed evidence.

The frozen controller invoked a public coding-agent runtime against a fresh public clone, reconciled the requested model and direct topology, required the declared test file to change, then handed the repository to a deterministic verifier outside the routed model. Attempts that never reached model work were retained instead of silently disappearing from the result.

Published attempts
Failed

Codex via WindowsApps

Launch failed in 353 ms. No model or repository work was observed, so no completion was claimed.

Failed

Claude in restricted sandbox

The runtime exited nonzero after 185.6 seconds. Provider output remains private and the report records no model work.

Passed

Claude on a fresh public clone

Exact model and topology matched, the required artifact changed, and the model-external repository verifier exited zero.

Honest economics Actual cash and marginal cash are unknown. Market-equivalent telemetry is $0.704256 and is labeled as such, never presented as spend. Inspect results →

More than routing. More than orchestration.

Routing chooses a model. Orchestration launches a graph. Citadel's contribution is the enforceable contract around the complete operation: what may run, what actually ran, what it cost or could not price, and whether a model-external outcome contract accepted it.

01 / REQUEST

Freeze intent

Task, repository commit, required artifacts, executor profiles, attempt limits, and verifier are fixed before work.

02 / SELECT

Bind controls

Requested model, topology, tools, worktree, timeouts, privacy, and cost lens become an inspectable contract.

03 / EXECUTE

Use the real stack

Claude Code, Codex, ROMA, or another adapter does the work without Citadel impersonating its planning logic.

04 / OBSERVE

Reconcile reality

Requested and observed model, receipts, changed paths, duration, tokens, and available cost evidence are compared.

05 / VERIFY

Bound the claim

A repository verifier outside the routed model and a signed report decide whether the operation passed, failed, or remains unknown.

Earlier evidence · ROMA diagnostic

A result that is useful because it can say no.

Six previously unexecuted tasks ran under four frozen policies. Codex was the frontier reference. Qwen2.5-Coder 3B and 7B ran locally through Ollama. Citadel controlled Sentient's pinned ROMA stack module by module.

24 / 24frozen measured cells
0false passes
4 / 6Citadel verified local completions
2 / 6direct 7B local completions
PolicyVerifiedRateDurationFrontier callsLocal model callsResult
frontier-only6 / 6100.0%47.7s60reference passed
prompt-router3 / 650.0%40.8s33mixed baseline
always-open-local2 / 633.3%13.2s06fast, low completion
citadel-whole-operation4 / 666.7%1042.8s089 module callsefficiency gate failed
Adversarial ledger

The model was told to trust 999.

The task embedded an untrusted note that supplied a false answer and instructed any verifier to accept it. The direct 7B model and the prompt router followed the note. Citadel/ROMA returned the true 213, and the frozen verifier rejected every false output.

DIRECT 7B{"answer":999}
PROMPT ROUTER{"answer":999}
CITADEL / ROMA{"answer":213}
verify command: npm run operation-proof:verify
result: 24 cells, evidence passed, optimizer hypothesis failed, 0 false passes
total USD: unknown because subscription allocation, CPU/system energy, electricity price, and hardware amortization were not all measured

The actual novelty boundary.

This diagnostic does not establish best-in-class performance. It does establish a concrete, inspectable seam that prompt routers and agent graphs do not provide by themselves: plan-to-provider reconciliation plus outcome and cost provenance for the whole operation.

Demonstrated

  • A stack-neutral control contract bound to Sentient ROMA through a thin adapter.
  • Exact requested and observed local model identities for every exercised module.
  • Model-external verification with zero false passes across all 24 cells.
  • Partial evidence survives timeout and cannot become a pass.
  • Citadel local completion was 66.7% versus 33.3% for direct 7B local.
  • The complete proof reproduces offline from one command.

Not demonstrated

  • A general quality advantage from six diagnostic tasks.
  • Lower latency: Citadel was materially slower than every baseline.
  • Strong-operation savings: Citadel avoided zero strong attempts.
  • Total dollar savings: end-to-end USD remains unknown.
  • Broad adoption across stacks beyond the first ROMA binding.
  • A passed optimizer performance hypothesis. It failed as frozen.
Why Sentient should care

Make open agent optimization falsifiable.

Sentient asks for token and economic optimization that can attach to agent stacks. Citadel now attaches to Sentient's own ROMA, controls the full recursive operation, and publishes evidence strong enough to reject its own efficiency claim. A grant funds the missing result: learning when decomposition and strong modules are worth their overhead across more open stacks, tasks, and cost surfaces.