Start with the product.
End with the falsifiable claim.
The full Citadel case in one short pass: progressive first use, operation evidence, six separately frozen studies, and the bounded result Sentient funding would generalize.
Watch the evidence reproduce.
This 42-second supplement is rendered from commands actually executed in the release checkout. The committed JSON retains every output line, exit code, output digest, and source revision.
Citadel starts with /do, then adds durable operation control only when the work needs it. The research case asks whether the complete operation—not merely one model call—can be made cheaper without hiding failure.
Read the complete transcriptFour parts · about two minutes
Citadel starts with one command: /do. You describe the engineering outcome, and Citadel chooses the smallest operating lane that can carry it. When work lasts longer, Citadel adds repository state, recovery, coordination, bounded execution, and a concrete next action around Claude Code or Codex.
The research question begins where ordinary orchestration ends. A smaller model is not cheaper if its answer fails, the verifier forces a retry, or the operation hides part of its cost. Citadel binds the declared route to observed execution, a verdict outside the routed model, cost lenses, and a signed receipt.
Earlier prospective studies tested that contract. Three local calibrations failed or missed their gates and remain public. A first Claude-plus-local policy preserved 12/12 outcomes but reduced comparison cost only 28.4%, missing its frozen 30% gate. A separately frozen hybrid v2 used twelve new synthetic tasks, preserved 12/12 outcomes for both policies, reduced Claude calls from twelve to five, and reduced comparison cost 38.7% with every frozen gate passed.
The later public holdout is the current result. It used 24 distinct outside-authored repositories, sealed sixteen evaluation routes before model calls, and published all 32 official verdicts. Direct Claude verified 2/16 and the controller 3/16 at 1.26% lower comparison cost. Because the baseline passed only 12.5%, Citadel does not call that useful quality preservation or savings. Sentient funding buys the retrieval, edit protocol, valid baseline, broader stacks, models, tools, hardware, complete cost, and learned operation-value policy still missing. The public gate remains at least 80% absolute verified completion, at least 95% of a valid frontier baseline, and at least 30% lower measured end-to-end cost; frontier must first clear 80% overall and 70% in every frozen task stratum.