AI reliability & evaluation
Evals, verification gates, tracing and cost tracking, built into the system instead of bolted on.
03 · AI reliability & evaluation
Know when your AI is wrong
It's easy to see what an AI system did. It's much harder to say whether it was right. Without evaluation, every prompt change is a guess and regressions are found by users.
Getting it right
- Evaluation sets built from real failures, run on every change
- Narrow, checkable questions instead of one vague quality score
- Verification gates that stop bad outputs before anyone sees them
- Tracing and cost per request, so failures and spend are visible
What I build
- Evaluation harnesses and regression suites
- Model-graded checks with explicit thresholds
- Verification and refusal gates
- Observability: traces, cost ledgers, budget guards
Proof
Typical stack
Langfuse · OpenTelemetry · Prometheus / Grafana · pytest · TypeSafe · custom harnesses
Questions
Rules or a model as the judge?
Both. Code checks what can be checked deterministically, a model handles judgment, and each has its own threshold.
How big does an evaluation set need to be?
Start small and real: a few dozen labelled cases drawn from actual failures already catch regressions. Grow it every time a new failure appears.