Ship only what beats the baseline.
BrainstormLLM predicts which phases of a software-development pipeline a task actually needs, from specification through documentation, so AI agents can skip the rest. It was trained on 2,203 real Claude Code session turns and had to clear kill gates set before any training began.
Third approach, first pass.
A one-shot model reached a mean F1 of 0.52 and a plan cache reached 0.60. Both failed the gate. The version that passed conditions each phase's prediction on the outcome of the phases before it: the same kind of model, with better features.
Kill gate: mean F1 of at least 0.75
Set in advance. All passed.
Missing any one of these would have ended the project. The thresholds were written down before the first model was trained.
| Gate | Threshold | Result | Status |
|---|---|---|---|
| Enough data | 200+ examples, 3+ task types | 2,203 / 9 | Pass |
| Phase-skip accuracy | Mean F1 of at least 0.75 | 0.796 | Pass |
| Cost reduction | At least 25% against running every phase | 68% | Pass |
| Inference time | Under 10 ms with ONNX | 0.31 ms | Pass |
Sequential predictors
Gradient-boosted models, one per pipeline phase, each conditioned on the phases before it.
Model-tier classifier
Trained separately on 400,000+ RouterBench data points; 93% cross-validated accuracy.
Inline inference
Both models export to ONNX, so BrainstormRouter can call them on every request.
Part of one loop.
BrainstormLLM learns from the trajectories the control plane records, and its predictions decide what the CLI runs next.