AI Reliability Consultant · ex-NVIDIA
Reliability engineering for AI systems in production
I help small teams find out why their AI products fail, build the evals and tracing to catch it, and set up the monitoring so it stays fixed.
Your AI product worked in the demo. Then real users showed up. Now a customer has sent you a bad answer nobody can reproduce. Or someone changed a prompt and something unrelated broke, and a user found it before you did. Or maybe the bill tripled last month and nobody can say which feature caused it.
I start by finding out how your product actually fails, then ranking those failures by what each one costs you. Evals come after that, because evals written before you know how a product fails can pass while the product keeps breaking.
Once we know what's worth catching, we build it. Evals built from your real traffic and wired into CI, so a prompt or model change gets checked before it ships. Tracing, so a failure can be reproduced instead of guessed at. Monitoring, so the next one reaches you before it reaches a customer. Correctness, cost and speed measured together, because a change that improves answers and doubles your spend is still a regression.
Not sure where to start? The audit is the usual first step.
Engagements
Reliability audit
|Find out where your AI system stands, and what's most likely to break next.
For teams who have shipped something and aren't sure why it misbehaves, or who want to know before it does.
- Root-cause analysis on a sample of recent production failures
- Failure taxonomy: the categories of things that go wrong, and how often
- Review of your evals, monitoring and release process, with the gaps called out
- Cost, latency and timeout findings
- A prioritized fix list, ranked by what each failure costs you
- A small starter eval set, built from the real failures found here
You get a written report, the eval set, and a call to walk through it.
Reliability build
|Put evals, tracing and monitoring into the system you already have.
For teams who know what they need, or who've done the audit and want it fixed.
- Tracing across every step: prompts, retrieval, tool calls, model responses
- Eval suite built from your real traffic and failure modes, wired into CI so every change gets checked before it ships
- Correctness, cost and latency measured together, with baselines so you can tell whether a new model is actually better for you
- Production monitoring and alerting: what to watch, and what's worth waking someone for
- Runbook for the failures most likely to recur
You get working infrastructure your team owns, plus documentation and a handover session.
Ongoing reliability
|Someone who investigates when it breaks and keeps the evals honest.
For teams shipping regularly who don't want reliability to decay between releases.
- Incident investigation, with written root-cause reports
- After each incident: why it wasn't caught earlier, and what gets added so it can't recur
- Eval maintenance as prompts, models and data change
- Review of the big changes before they ship: model swaps, prompt rewrites, retrieval changes
You get a reliability engineer without a full-time hire.
Why me

I spent over seven years at NVIDIA building diagnostic software for automotive and datacenter hardware, and leading that work for several automotive platforms. That work is deciding what to measure, separating real faults from noise, and tracing a fault through signals like power, thermal and resource utilization to find what else it affected. Then defining what happens when one is found, and making sure it gets caught automatically the next time.
AI systems break in familiar shapes: intermittently, under real load, without much visibility inside. So I work the same way here. Trace the failure, isolate the layer it came from, whether that's retrieval, the model, the prompt, a tool call or the code around it, then turn it into a regression test and leave the observability in place to catch the next one.
Get in touch
Tell me what's breaking. I'll tell you honestly whether I can help.
Book a callor email me at gaurav@goyalgaurav.com