Do I need to know how to code to run evals?
No. Writing jobs, checklists and scenarios takes no code, the AI grader is a prompt you paste into any chat assistant, and the by-hand table works in any tool.
Which AI tool do I need?
Any. The method works the same in ChatGPT, Claude, Gemini or Copilot. Tools that automate parts of it are optional, and they change fast, so learn the method first.
What's the difference between an eval and a test?
A test checks that software does exactly what it was built to do. An eval checks whether an AI's work is good enough, including judgment calls like tone, where there's no single right answer.
How many test scenarios do I need?
Start with five to ten for one job. Then add one every time the agent gets something wrong in real use. The set grows as you learn where it fails.
Where do test scenarios come from?
Real inputs first: past emails, requests or notes the agent has handled. Make up cases only for gaps you can name, like a request it's never seen, and mark them. Either way, the job's owner writes or approves what good looks like for each.
Why run the same scenario more than once?
Because AI doesn't answer the same way twice. A scenario that passes once and fails twice isn't passing.
Can AI grade AI's work?
Yes, if you give it a yes/no checklist. Check it against a person on a handful of examples first, and reword any question where they disagree.
Why yes/no instead of a score from 1 to 10?
Because AI graders hedge. Asked for a score, they drift toward 7 or 8. Yes or no forces a real answer.
What's a trace?
A log of everything an agent did on a real run: every lookup, every tool it used, every message. It's how you find out why it did something wrong.
How do evals help me choose a model?
Run each model you're considering on your test scenarios, then compare how many it passes with what it costs. The right trade-off depends on the job.
When should I fine-tune a smaller model?
Most teams should start by improving the agent's instructions and tools, which is cheaper and faster. For narrow, repetitive jobs, Vivek says some teams are fine-tuning small open-source models on their own data, and seeing results close to top models at around a tenth of the cost.