Skip to content
Velum

Managed AI operations

How to evaluate LLMs and AI agents in production

Short answer

LLM evaluation in production means scoring a live AI system against a fixed test set before every change and sampling real traffic after release, using metrics tied to the task, such as correct extraction, correct tool calls, groundedness and escalation rate. For agents, you measure whether the whole task was completed correctly, not only whether each reply sounded good. Automated checks and LLM-as-a-judge scoring handle volume, and people review a small sample every week.

Key takeaways

  • A test set of 50 to 300 real, labelled examples catches most regressions from prompt or model changes.
  • Agent evals score task completion and tool-call accuracy, not only the final text.
  • RAG systems need separate scores for retrieval quality and for answer faithfulness to the sources.
  • LLM-as-a-judge scoring should be checked against human labels before you trust it.

What AI evals are and which LLM evaluation metrics matter

AI evals are repeatable tests for AI systems. You collect realistic inputs, define what a good output looks like, and score the system automatically every time something changes: the prompt, the model, the retrieval index or a tool. Without evals, teams find out about regressions from users. With them, a model upgrade becomes a measured decision instead of a guess, and you can compare providers on your own data rather than public leaderboards.

The right LLM evaluation metrics depend on the task. For extraction, measure field-level accuracy against labelled documents. For classification or routing, measure precision and recall per category. For drafting, use a rubric covering correctness, tone and required content, scored by a judge model and checked by people. Operational metrics matter too: latency, cost per task, refusal rate and how often the system escalates to a person. A short dashboard with five or six of these is more useful than dozens of generic scores.

How to evaluate an AI agent

An AI agent evaluation framework has to judge a sequence of steps, not a single answer. The agent reads an input, chooses tools, calls them with arguments and decides when it is done. Useful AI agent evaluation metrics include task success rate, tool-call accuracy, whether arguments were valid, number of steps compared with the expected path, and whether the agent correctly stopped and asked for approval before an irreversible action. The last one is a safety metric and should be close to 100 percent.

In practice you build scenario tests: a realistic starting state, such as an inbox with a supplier invoice, and an expected end state, such as a draft bill in Xero awaiting approval. Run each scenario several times, because agents are not fully deterministic. AI agent observability and evaluation work together here: traces from tools like Langfuse or LangSmith show each step in production, and failed traces become new test cases, so the test set grows from real failures.

Evaluating RAG, LLM-as-a-judge and how often to run evals

LLM evaluation for RAG splits into two questions. Did retrieval find the right passages, measured by context precision and recall against labelled questions? And did the answer stay faithful to those passages, measured by groundedness or faithfulness scores? Libraries such as Ragas compute these. A low faithfulness score with good retrieval points to the prompt or model. Poor retrieval points to chunking, metadata or the search index, which is a different fix entirely.

LLM-as-a-judge uses a strong model to grade outputs against a rubric. It scales well, but judges have biases, such as favouring longer answers, so calibrate them against a few hundred human labels and keep the rubric specific. Run the offline test set on every change before release, run sampled online scoring daily or weekly depending on volume, and do a human review of 20 to 50 production cases each week. Re-run everything when the provider releases a new model version.

How it works

  1. 1

    Define success per task

    We agree with the business owner what a correct outcome looks like for each agent task, including when it must escalate or ask for approval.

  2. 2

    Build the test set

    We label 50 to 300 real examples, covering common cases, edge cases and past failures, and store them with expected outputs.

  3. 3

    Automate scoring

    We wire deterministic checks, RAG metrics and calibrated LLM-as-a-judge rubrics into a pipeline that runs on every change.

  4. 4

    Monitor production

    We trace live runs, sample outputs for scoring and review, and turn every failure into a new test case.

  5. 5

    Release with approval gates

    New prompts or models ship only when scores hold, and actions such as sends, payments and postings continue to wait for a person to approve.

Before and after

TaskBy handWith agents
How regressions are foundUser complaints, days or weeks laterTest run before release, within minutes
Model upgrade decisionGut feel from a few manual triesSide-by-side scores on 50 to 300 real cases
Human review effortAd hoc, often none20 to 50 sampled cases per week
Visibility into agent stepsFinal output onlyFull trace of each tool call and decision

Typical ranges from comparable deployments. Your baseline is measured before anything is built.

Tools it works with

  • Langfuse
  • LangSmith
  • Braintrust
  • Arize Phoenix
  • Ragas
  • Promptfoo
  • OpenAI Evals
  • Claude
  • Datadog

Questions people ask

01

What are AI evals?

AI evals are repeatable tests that score an AI system's outputs against expected results or a rubric. They run automatically when prompts, models or tools change, so you can see whether quality went up or down. Good evals use real examples from your own work.

02

How do you evaluate an AI agent?

Run the agent through realistic scenarios and check whether it reached the correct end state, called the right tools with valid arguments and stopped for approval where required. Run each scenario several times because agent behaviour varies. In production, trace every run and turn failures into new test cases.

03

What metrics are used for LLM evaluation?

Common metrics include accuracy against labelled answers, precision and recall for classification, rubric scores for generated text, faithfulness for RAG, and task success rate for agents. Operational metrics such as latency, cost per task and escalation rate matter just as much. Pick the few that match the business outcome.

04

How do you evaluate a RAG system?

Score retrieval and generation separately. Measure whether the right passages were retrieved using context precision and recall, then measure whether the answer is faithful to those passages and relevant to the question. Tools like Ragas automate these metrics on a labelled question set.

05

How often should AI evals run?

Run the offline test set on every change to prompts, models, tools or retrieval, before release. Score a sample of production traffic daily or weekly, depending on volume, and review a small sample by hand each week. Re-run everything when a provider updates the model you use.

06

What is LLM-as-a-judge?

LLM-as-a-judge is using a language model to grade another model's outputs against a written rubric. It lets you score thousands of outputs cheaply, but judges can be biased toward length or style. Check the judge against human labels before relying on it and keep rubrics narrow and specific.

Start with one workflow.

Thirty minutes. One real process. A practical next step.