Six separately frozen studies
The failures changed the policy.
Local studies exposed timeout, escalation, and baseline-validity defects. A calibrated synthetic support envelope passed. The outside-authored follow-up then showed that the retrieval/edit substrate and strong baseline were not ready.
V1 · adaptive local
More verified cells; savings not robust.
Across 12 tasks × 2 policies × 3 timing repetitions, adaptive recorded 27/36 verified cells versus 24/36. Its frozen aggregate missed both 30% gates, and excluding one matched 60-second baseline timeout reverses all three economic comparisons.
27/36adaptive verified
+3.5%GPU energy in sensitivity
+5.4%modeled GPU cost in sensitivity
Method + signed result
V2 · capability profile
Same cell completion, economic regression.
Twelve new exact instances, mostly from task templates already seen in v1, routed work to 1.5B, 3B, or 7B. Model-external verification matched the baseline cell completion at higher measured GPU energy and modeled GPU cost.
24/36both policies verified
+15.7%GPU energy
12strong escalations
Signed result
Required disclosure
V3 · repository operations
Artifact integrity passed; economics failed.
Six fixture repositories × 2 policies × 2 timing repetitions. Both policies verified 6/12 cells. Citadel used 7.1% less measured GPU energy, below the frozen 20% gate, and 13.2% more tokens.
6/12both policies verified
0false passes or path violations
Failedfrozen evidence result
Method + signed result
V4 · hybrid calibration
Quality valid; economic gate missed narrowly.
Both policies verified 12/12 fresh tasks. The risk-only policy avoided four Claude calls but made four unsupported local attempts, reducing comparison cost 28.4% against a frozen 30% gate.
12/12both policies verified
28.4%comparison-cost reduction
Failedfrozen economic result
Calibration report
V5 · calibrated hybrid v2
Same verified outcomes; bounded economic gate passed.
Twelve new tasks, a valid Claude Sonnet 5 baseline, and a preregistered support envelope. Citadel used eight local attempts, recovered once, reduced Claude calls from twelve to five, and reduced comparison cost 38.7%.
12/12both policies verified
38.7%comparison-cost reduction
Passedevery frozen gate
Passed method + signed result
V6 · outside-authored public holdout
Evidence sequence passed; optimization claim did not.
Twenty-four distinct repositories supplied eight calibration and sixteen untouched evaluation tasks. Routes were published before calls. Qwen verified 1/16, direct Claude 2/16, and the Qwen-first controller 3/16 at 1.26% lower comparison cost. The baseline was too weak for a general result.
24distinct repositories
32/32official verdicts
Invalidstrong-baseline claim
Final report
Validation