Key takeaways
- A test set of 50 to 300 real, labelled examples catches most regressions from prompt or model changes.
- Agent evals score task completion and tool-call accuracy, not only the final text.
- RAG systems need separate scores for retrieval quality and for answer faithfulness to the sources.
- LLM-as-a-judge scoring should be checked against human labels before you trust it.
What AI evals are and which LLM evaluation metrics matter
AI evals are repeatable tests for AI systems. You collect realistic inputs, define what a good output looks like, and score the system automatically every time something changes: the prompt, the model, the retrieval index or a tool. Without evals, teams find out about regressions from users. With them, a model upgrade becomes a measured decision instead of a guess, and you can compare providers on your own data rather than public leaderboards.
The right LLM evaluation metrics depend on the task. For extraction, measure field-level accuracy against labelled documents. For classification or routing, measure precision and recall per category. For drafting, use a rubric covering correctness, tone and required content, scored by a judge model and checked by people. Operational metrics matter too: latency, cost per task, refusal rate and how often the system escalates to a person. A short dashboard with five or six of these is more useful than dozens of generic scores.
How to evaluate an AI agent
An AI agent evaluation framework has to judge a sequence of steps, not a single answer. The agent reads an input, chooses tools, calls them with arguments and decides when it is done. Useful AI agent evaluation metrics include task success rate, tool-call accuracy, whether arguments were valid, number of steps compared with the expected path, and whether the agent correctly stopped and asked for approval before an irreversible action. The last one is a safety metric and should be close to 100 percent.
In practice you build scenario tests: a realistic starting state, such as an inbox with a supplier invoice, and an expected end state, such as a draft bill in Xero awaiting approval. Run each scenario several times, because agents are not fully deterministic. AI agent observability and evaluation work together here: traces from tools like Langfuse or LangSmith show each step in production, and failed traces become new test cases, so the test set grows from real failures.
Evaluating RAG, LLM-as-a-judge and how often to run evals
LLM evaluation for RAG splits into two questions. Did retrieval find the right passages, measured by context precision and recall against labelled questions? And did the answer stay faithful to those passages, measured by groundedness or faithfulness scores? Libraries such as Ragas compute these. A low faithfulness score with good retrieval points to the prompt or model. Poor retrieval points to chunking, metadata or the search index, which is a different fix entirely.
LLM-as-a-judge uses a strong model to grade outputs against a rubric. It scales well, but judges have biases, such as favouring longer answers, so calibrate them against a few hundred human labels and keep the rubric specific. Run the offline test set on every change before release, run sampled online scoring daily or weekly depending on volume, and do a human review of 20 to 50 production cases each week. Re-run everything when the provider releases a new model version.
How it works
- 1
Define success per task
We agree with the business owner what a correct outcome looks like for each agent task, including when it must escalate or ask for approval.
- 2
Build the test set
We label 50 to 300 real examples, covering common cases, edge cases and past failures, and store them with expected outputs.
- 3
Automate scoring
We wire deterministic checks, RAG metrics and calibrated LLM-as-a-judge rubrics into a pipeline that runs on every change.
- 4
Monitor production
We trace live runs, sample outputs for scoring and review, and turn every failure into a new test case.
- 5
Release with approval gates
New prompts or models ship only when scores hold, and actions such as sends, payments and postings continue to wait for a person to approve.
Before and after
Typical ranges from comparable deployments. Your baseline is measured before anything is built.
Tools it works with
- Langfuse
- LangSmith
- Braintrust
- Arize Phoenix
- Ragas
- Promptfoo
- OpenAI Evals
- Claude
- Datadog