Skip to content
Ayush Gupta

AI reliability & evaluation

Evals, verification gates, tracing and cost tracking, built into the system instead of bolted on.

03 · AI reliability & evaluation

Know when your AI is wrong

It's easy to see what an AI system did. It's much harder to say whether it was right. Without evaluation, every prompt change is a guess and regressions are found by users.

Talk about this

Getting it right

  • Evaluation sets built from real failures, run on every change
  • Narrow, checkable questions instead of one vague quality score
  • Verification gates that stop bad outputs before anyone sees them
  • Tracing and cost per request, so failures and spend are visible

What I build

  • Evaluation harnesses and regression suites
  • Model-graded checks with explicit thresholds
  • Verification and refusal gates
  • Observability: traces, cost ledgers, budget guards

Typical stack

Langfuse · OpenTelemetry · Prometheus / Grafana · pytest · TypeSafe · custom harnesses

Questions

Rules or a model as the judge?

Both. Code checks what can be checked deterministically, a model handles judgment, and each has its own threshold.

How big does an evaluation set need to be?

Start small and real: a few dozen labelled cases drawn from actual failures already catch regressions. Grow it every time a new failure appears.

Working on something like this? Tell me about it.

Get in touch