AI agents that pick the right model and stay inside the lines.
Brainstorm is my AI research lab: a control plane that connects AI coding assistants to real systems, a router that governs and routes every request an agent makes, and a learned model that predicts which work a task actually needs. Built in my own time, not a commercial product.
Three projects. One loop.
Each one answers a question I kept running into while building with AI agents. How do they act safely on real systems? How does every request get governed? And how does the system learn which work is worth doing?
How the agents act.
Brainstorm
One governed channel between AI coding assistants and real systems. An MCP server with 58+ tools, approvals, cost tracking and an audit trail.
How every request is governed.
BrainstormRouter
Identity, budgets and evidence for AI agents, with Thompson sampling choosing among 55 models from nine providers on every request.
How the system learns.
BrainstormLLM
A model trained on 2,203 real Claude Code session turns that predicts which pipeline phases a task needs, so agents skip the rest.
A system that improves from its own work.
- 01
The CLI builds
Coding assistants work through the control plane, and every session leaves a trajectory: what was asked, which steps ran and what they cost.
- 02
The router routes
Every model call passes through the router, which learns which models win on which kinds of work and keeps agents inside their budgets.
- 03
The model predicts
BrainstormLLM learns from the trajectories which phases a task actually needs, and the CLI runs that plan next time.
Five endpoints make any system operable by agents.
How do you give an assistant such as Claude Code useful access to real infrastructure while keeping every action understandable and controllable? Instead of a custom integration per system, each service implements one small contract. The assistant discovers what it can do at run time, and every tool declares its risk.
- GET/healthStatus, version and product name
- GET/api/v1/…/toolsTool discovery, with a risk level on every tool
- POST/api/v1/…/executeRun a tool, or simulate it and return the change set
- POST/api/v1/platform/eventsSigned events back to the control plane
- POST/api/v1/platform/tenantsTenant lifecycle, so every call is scoped
Agents write most of the code. I set the bar.
The common thread across these projects is a disciplined way of working with AI agents, so their output is judged against something other than their own description of it.
Spec first
Substantial work starts as a written specification with acceptance criteria.
Gates, not vibes
Results have to beat a baseline set in advance. BrainstormLLM shipped only after passing its kill gates.
Review loops
Independent reviewer agents score the work against a stated bar, round after round.
Verify on the real path
Done means proven where it runs: a live response, a database row, an App Store build.
Built by Justin Jilg.
I lead alliances in cloud cybersecurity by day. This is what I build with AI agents in my own time, alongside iPhone apps and sites made with and for my family.