Why We're Building Bandito
Most GenAI features aren't successful. Maybe make it to prod, rarely stay there. Old school machine learning projects had fit this same pattern.
Why? Probabilistics systems will eventually be wrong. Treating them like deterministic software isn't the right pattern.
The solution is to manage the probabilities through analysis, scoring and constant improvement. This is hard! Bandito makes it easier, but not set & forget.
The Problem We're Solving
You developed an AI workflow or (buzzword warning) agent. Maybe you do some internal testing & setup trace logging. Everything seems (vibe) solid.
But cost, latency, edge cases, model degredation and more can torpedo you product/feature.
Bandito wraps up AI engineering basics into an easy to use CLI.
- Run basic diagnostics
Path to Products Persisting in Production
You've thoughtful designed your AI workflow / agent. No "god pattern", solid harness, defined purpose, etc.
You setup trace logging, ran tests offline and validated your best guess at common patterns.
If you're building with LLMs in production, you've already set up tracing. Langfuse, LangSmith, Braintrust — pick your favorite. You can see every span, every token count, every latency number.
And then what?
You stare at dashboards. You spot a bad trace, squint at the spans, form a theory. You open your prompt file, tweak something, redeploy, and hope. Maybe you check back in a week.
This is the gap. Trace loggers answer "what happened." Nobody answers "what to do about it" — let alone does it for you.
The loop every good team runs
The best AI teams we've talked to run the same workflow:
- Observe — pull recent traces, look for patterns
- Grade — manually score a sample for quality
- Judge — scale that grading with LLM-as-judge
- Analyze — find cost/quality/latency tradeoffs
- Improve — test changes offline before deploying
It's the right workflow. It's also tedious, manual, and the first thing that gets skipped when a deadline hits.
Bandito runs the loop for you
Bandito is the intelligence layer that sits on top of your traces. It connects to wherever your traces already live — local files, S3, Postgres, Langfuse — and runs the observe-grade-judge-analyze-improve loop continuously.
- You grade 15 traces. Bandito scales that to 500 with LLM-as-judge.
- Judge finds a pattern. 12% of failures are context window errors.
- Analysis shows a tradeoff. 60% of your traces could use a cheaper model at minimal quality loss.
- Bandito recommends the fix. Route simple queries to mini. Add assertions for max_tokens. Create an eval set from error traces.
Each cycle compounds. The rubric gets better. The judge gets more accurate. The recommendations get more specific. You ship with more confidence.
What's next
We're in early access. The SDK, CLI, TUI, and judge are all working. Analysis surfaces are live. The autonomous AI engineer loop is next.
If you're building LLM-powered products and tired of staring at dashboards, join the waitlist — we'd love to build this with you.