Delta / recorded research
What our first
training comparison shows.
A matched 4B procurement experiment, comparing training on real records with a synthetic control. Every score comes from task-level programmatic grading.
Procurement pilot / recorded checkpoint
Real work. A measurable signal.
+75pp
Development-task pass-rate gap
Same 4B base model, 128 training cases per arm and matched supervised-token budgets. 20 development tasks scored by the programmatic verifier.
Development checkpoint after one matched continuation. The synthetic arm is still below the training-fit gate; the 300-task model test remains sealed.
300
Private test tasks
643
Packages checked
12,359
Invalid mutations rejected