LLM Evaluation: How to Test AI Agent Responses
AI agent responses can sound confident, polished, and completely reasonable while still being wrong. They may use the wrong source, call the wrong tool, pass invalid arguments, ignore a business rule, or reach a correct answer through an unsafe action.
That is why production AI systems need more than traditional unit tests. They need an evaluation process that measures not only the final text, but also the evidence, decisions, tool calls, safety, latency, and cost behind it.
This guide explains how to design that process from an engineering perspective.
What Is LLM Evaluation?
LLM evaluation is the systematic process of measuring whether a language-model application produces acceptable results for a defined task. An evaluation runs representative inputs through the system, captures its outputs and execution traces, scores them against explicit criteria, and reports whether a change improved or degraded the system.
A practical LLM evaluation process must examine both the final response and the actions the AI agent performed to produce it.
The object under test is usually larger than the model. A useful evaluation may include the prompt, model and parameters, retrieved context, tools, orchestration logic, output schema, safety controls, and post-processing code.
This distinction matters. A weak answer does not automatically mean the model is weak. The retriever may have returned irrelevant documents, the prompt may have omitted a policy, a tool may have failed, or the application may have truncated useful context.
If you are new to the underlying architecture, start with What Is an AI Agent? and What Is Agentic AI?.
Why AI Agent Evaluation Is Harder Than Normal Software Testing
Traditional code is mostly deterministic. Given the same input and environment, a function normally returns the same output. Language models are probabilistic: two answers can use different wording and still be equally correct, while a fluent answer can be factually wrong.
Agents add another layer of variability because they make decisions and interact with external systems. The same user request may produce several valid plans, tool sequences, or response formats.
| Traditional application test | AI agent evaluation |
|---|---|
| Expected value is often exact | Several outputs may be acceptable |
| Behavior is usually deterministic | Behavior may vary across runs |
| Tests focus on code paths | Tests cover prompts, context, model behavior, tools, and code |
| A pass or fail is often sufficient | Quality frequently needs a graded score |
| Unit tests can isolate most logic | End-to-end behavior may depend on live retrieval and services |
An agent can also produce the right final sentence for the wrong reason. For example, it might invent an order status instead of querying the order system. A response-only test may pass; an execution-trace test should fail.
What Should You Evaluate?
A production evaluation should separate quality into dimensions. A single “good answer” score hides the reason a system fails.
1. Task success
Did the agent complete the user’s requested outcome? For a support agent, success may mean answering the policy question and creating a valid ticket. For a scheduling agent, it may mean finding an allowed time and obtaining confirmation before booking.
2. Correctness
Are the factual claims and calculations correct? Correctness can sometimes be checked against a reference answer, database record, calculation, or executable assertion.
3. Relevance and completeness
Does the response address the actual question without unrelated content? Does it include every required element?
4. Groundedness
Are claims supported by the context or authoritative sources provided to the model? Groundedness is critical for retrieval-augmented generation.
5. Tool-use quality
Did the agent choose the correct tool, supply valid arguments, interpret the result correctly, and avoid unnecessary calls? Read Tool Calling in AI Agents Explained for the underlying request-execution pattern.
6. Instruction and policy compliance
Did the agent follow system instructions, authorization rules, business constraints, and required approval steps?
7. Safety and security
Did it resist prompt injection, avoid exposing sensitive data, refuse prohibited actions, and stay within the user’s permissions?
8. Output validity
If another service consumes the result, is it valid JSON and compliant with the required schema? See Structured Output in AI Agents: Why JSON Matters.
9. Operational quality
How long did the request take? How many tokens and tool calls did it consume? What did it cost? A response can be correct but operationally unacceptable.
A Practical LLM Evaluation Architecture
Treat evaluation as a pipeline, not an occasional manual review.
Versioned test cases
↓
Evaluation runner
↓
AI system under test
├── prompt and model
├── retrieval pipeline
├── agent orchestrator
└── tools and policies
↓
Captured result and trace
↓
Scorers
├── deterministic assertions
├── semantic or model-based graders
├── safety checks
└── human review samples
↓
Metrics, failure analysis, and release gate
Each evaluation run should record enough configuration to reproduce the result:
- test case and dataset version;
- prompt or instruction version;
- model, model version, and inference parameters;
- retrieval and reranking configuration;
- tool definitions and relevant service versions;
- final response, citations, tool calls, and errors;
- latency, token usage, and estimated cost;
- individual scores and grader versions.
Without this metadata, a changed score tells you that something moved but not why.
Start With a Representative Evaluation Dataset
The quality of an evaluation depends heavily on its test cases. Twenty convenient prompts written by the development team rarely represent production behavior.
Build a versioned dataset containing:
- Normal cases: common requests the system must handle well.
- Boundary cases: missing fields, long input, conflicting instructions, ambiguous dates, and unusual values.
- Negative cases: requests the agent should decline or escalate.
- Adversarial cases: prompt injection, data-exfiltration attempts, unsafe tool arguments, and permission bypasses.
- Historical failures: sanitized production incidents converted into permanent regression tests.
- Slice labels: language, customer type, task type, risk level, tool, document family, or other segments important to the product.
A test case should define more than a prompt. It can include expected facts, allowed tools, forbidden actions, required source IDs, output schema, maximum latency, and a scoring rubric.
{
"case_id": "refund-approval-014",
"input": "Refund order 8472 to the original payment method.",
"expected": {
"required_tools": ["get_order", "get_refund_policy"],
"forbidden_tools": ["issue_refund"],
"required_behavior": "request_approval",
"must_not_claim": ["refund_completed"]
},
"tags": ["refund", "approval", "high-risk"]
}
In this example, the agent should gather facts but must not execute the refund before approval. Testing only the final wording would miss the most important behavior.
Five Evaluation Methods That Work Together
1. Deterministic checks
Use normal code whenever the expected behavior can be expressed exactly. These checks are fast, inexpensive, and easy to run in CI.
- JSON parses successfully.
- Output validates against a schema.
- Required fields and citations are present.
- Numbers match an authoritative calculation.
- A prohibited phrase or sensitive value is absent.
- Tool names and arguments satisfy business rules.
- Latency and cost stay below defined limits.
Do not use an LLM grader for a rule that a few lines of code can verify more reliably.
2. Reference-based scoring
Compare the result with a known answer or set of required facts. Exact string matching is usually too brittle for prose. Prefer checking normalized values, extracted entities, required concepts, or task-specific facts.
Semantic similarity can help detect broadly equivalent wording, but it should not be treated as proof of factual correctness. Two sentences can be semantically similar while differing on a crucial date, amount, or negation.
3. LLM-as-a-judge
A separate model can grade qualities that are difficult to encode as exact rules, such as relevance, clarity, completeness, or groundedness. Give the judge the user request, response, relevant context, and a precise rubric.
A useful rubric defines each score:
- 4 — Excellent: correct, complete, supported, and directly useful.
- 3 — Acceptable: correct overall with a minor omission that does not change the outcome.
- 2 — Weak: partially correct but missing an important requirement or containing an unsupported claim.
- 1 — Failed: incorrect, unsafe, irrelevant, or unable to complete the task.
Ask for structured scores and short evidence, not an open-ended opinion. Judge outputs should also be calibrated against human-reviewed examples. Models can show position bias, verbosity bias, self-preference, and inconsistency.
4. Pairwise comparison
When choosing between two prompts, models, or retrieval strategies, ask a grader or reviewer which response better satisfies the same rubric. Pairwise evaluation is often easier and more stable than assigning absolute scores.
Randomize which candidate appears first and include a tie option. Otherwise, presentation order can influence the result.
5. Human evaluation
Human review remains necessary for high-risk decisions, subjective quality, newly discovered behaviors, and calibration of automated graders. Reviewers should use the same documented rubric and periodically score overlapping cases so you can measure agreement.
Human evaluation does not need to cover every request. A strong design uses deterministic checks broadly, model-based graders at scale, and targeted human review for calibration, disagreement, and risk.
How to Evaluate Tool-Calling Agents
Tool-calling evaluation must inspect the trajectory, not only the final response. A tool-enabled agent typically has four distinct responsibilities:
- decide whether a tool is needed;
- select the correct tool;
- construct safe, valid arguments;
- use the returned result correctly.
Useful metrics include:
- Tool selection accuracy: percentage of cases using the expected tool when required.
- Argument validity: percentage of calls matching the declared schema and domain constraints.
- Unauthorized action rate: percentage of cases that attempt an action outside policy or user permission.
- Unnecessary call rate: calls that add cost or risk without helping the task.
- Tool-result faithfulness: whether the response accurately represents the returned data.
- End-to-end task success: whether the overall goal was completed correctly.
Mock tools are useful for repeatable offline tests. They let you inject timeouts, malformed data, empty results, authorization failures, and conflicting records without changing production systems. Keep a smaller set of integration tests for real service contracts.
If your agent uses standardized tool connections, What Is MCP? explains how Model Context Protocol fits into the architecture.
How to Evaluate RAG-Based Agent Responses
RAG evaluation should separate retrieval quality from generation quality. If the correct evidence never reaches the model, prompt tuning alone cannot solve the problem.
Retrieval metrics
- Recall at k: did the top k results include the document or passage needed to answer?
- Precision at k: how many retrieved results were actually relevant?
- Ranking quality: did the most useful evidence appear near the top?
- Context relevance: is the supplied context useful for the specific question?
Generation metrics
- Answer correctness: is the response factually correct?
- Faithfulness or groundedness: are its claims supported by retrieved context?
- Answer relevance: does it address the user’s request?
- Citation correctness: do citations support the claims attached to them?
- Abstention quality: does the system say it lacks enough evidence when retrieval cannot support an answer?
For the larger retrieval design, read How RAG Helps AI Agents Use Your Own Data, Vector Databases for AI Agents Explained, and Embeddings Explained for AI Agents and RAG.
A Practical Evaluation Scorecard
A scorecard should reflect business risk rather than averaging every metric equally.
| Dimension | Example measurement | Example release rule |
|---|---|---|
| Task success | Pass rate across representative cases | No more than 1 percentage-point regression |
| Correctness | Deterministic checks plus rubric score | At least 95% on critical fact checks |
| Tool behavior | Selection, arguments, result use | 100% on high-risk authorization cases |
| Groundedness | Supported-claim rate | No critical unsupported claims |
| Safety | Adversarial and policy test pass rate | Zero critical violations |
| Format | Schema-validation success | At least 99.5% |
| Latency | p50 and p95 end-to-end duration | p95 below product objective |
| Cost | Average and p95 cost per task | Within the approved budget envelope |
These numbers are examples, not universal standards. A medical workflow, internal documentation assistant, and marketing-copy generator should not share the same thresholds.
Also report scores by meaningful slices. An overall 92% success rate can hide a 60% success rate for long documents, one language, a particular tool, or high-risk requests.
Put LLM Evaluation Into CI/CD
Run evaluation whenever you change a prompt, model, tool definition, retrieval setting, policy, or orchestration component. A practical pipeline uses several layers:
- Pull request: run a small, fast set of deterministic and high-risk regression cases.
- Pre-release: run the complete offline suite, model-based graders, and cost comparison.
- Staging: run end-to-end tests against realistic integrations.
- Canary: expose a small traffic percentage and compare production signals.
- Full release: promote only when quality, safety, latency, and cost gates pass.
Compare a candidate configuration with a fixed baseline. Absolute scores matter, but the release question is often: Did this change improve the intended cases without creating unacceptable regressions elsewhere?
Because model outputs vary, run important cases multiple times when variance could affect a release decision. Report the distribution and failure rate instead of relying on one lucky response.
Offline Evaluation Is Not Enough
A test dataset cannot capture every real user, document, tool failure, or attack. Production evaluation closes that gap.
Capture privacy-safe traces and monitor:
- task completion and abandonment;
- user corrections, retries, and escalations;
- tool errors and authorization failures;
- schema-validation failures;
- retrieval quality proxies and missing citations;
- policy and safety events;
- latency, tokens, and cost;
- quality scores on sampled interactions.
Do not log sensitive prompts, retrieved documents, or tool results indiscriminately. Redact or tokenize personal data, restrict access, define retention periods, and keep evaluation data aligned with your security and compliance requirements.
Turn confirmed production failures into sanitized regression cases. Over time, this creates an evaluation suite based on how the system actually breaks—not only how the team imagined it might break.
Common LLM Evaluation Mistakes
Using one metric
A single aggregate score cannot explain whether failure came from retrieval, reasoning, tool use, formatting, or safety.
Testing only the final answer
For agents, inspect tool selection, arguments, intermediate results, approvals, and termination behavior.
Using exact-match tests for natural language
Exact matching rejects valid paraphrases. Reserve it for IDs, values, enums, schemas, and other deterministic requirements.
Trusting an LLM judge without calibration
A judge is another probabilistic component. Validate it against human-reviewed examples and monitor disagreement.
Changing several components at once
If you change the model, prompt, chunking, and tool schema together, a score change is difficult to diagnose. Use controlled comparisons where possible.
Ignoring variance
One run is a sample, not a guarantee. Repeat critical tests and track pass rates.
Optimizing the benchmark instead of the product
A narrow dataset can encourage brittle prompt changes. Protect a holdout set, refresh cases, and validate with production outcomes.
Ignoring cost and latency
A configuration that gains one quality point while doubling response time and cost may be the wrong production decision.
A Pragmatic Implementation Plan
You do not need a large evaluation platform on day one.
- Choose one important agent workflow.
- Define task success, critical failures, and operational limits.
- Create 30–50 representative cases, including failures and adversarial inputs.
- Add deterministic assertions for schemas, facts, tool rules, and policies.
- Add one rubric-based grader for qualities that code cannot check.
- Have domain reviewers calibrate a sample and resolve ambiguous criteria.
- Record versions, traces, tokens, latency, and cost for every run.
- Establish a baseline before changing the system.
- Add the fast suite to pull requests and the complete suite before release.
- Convert verified production failures into regression tests.
The goal is not to prove that the agent is universally intelligent. The goal is to gather enough evidence that this version performs defined tasks within acceptable quality, safety, latency, and cost boundaries.
Final Thoughts
LLM evaluation becomes manageable when you stop treating it as a search for one perfect metric.
Use normal assertions for deterministic requirements. Use reference checks for known facts. Use carefully calibrated model graders for nuanced language quality. Use human reviewers for risk, ambiguity, and calibration. For agents, evaluate the execution path as well as the final response. For RAG, evaluate retrieval separately from generation.
Most importantly, make evaluation part of engineering delivery. Prompts, models, tools, retrieval pipelines, and policies are versioned system components. They deserve regression tests and release gates just like application code.
To place evaluation in the broader AI-agent landscape, continue with AI Agent vs Chatbot vs Copilot.
Frequently Asked Questions
What is LLM evaluation?
LLM evaluation is the process of testing a language-model application against defined tasks and quality criteria. It can measure correctness, relevance, groundedness, safety, tool use, formatting, latency, and cost.
How do you evaluate AI agent responses?
Evaluate both the final response and the agent’s execution trace. Check whether it selected the correct tools, passed valid arguments, followed authorization rules, interpreted results accurately, and completed the user’s task.
Can LLM responses be unit tested?
Yes, but exact text matching is rarely appropriate. Unit-test deterministic elements such as schemas, required fields, calculations, tool arguments, policies, and post-processing. Use rubrics, semantic checks, or human review for open-ended language quality.
What is LLM-as-a-judge?
LLM-as-a-judge uses a language model to grade another model’s output against a rubric. It is useful for scalable evaluation of relevance, completeness, or groundedness, but it should be calibrated against human judgments and combined with deterministic checks.
How do you evaluate a RAG system?
Evaluate retrieval and generation separately. Retrieval metrics include recall at k, precision at k, and ranking quality. Generation metrics include correctness, groundedness, relevance, citation accuracy, and appropriate abstention.
How large should an LLM evaluation dataset be?
There is no universal size. A focused workflow can begin with 30–50 carefully chosen cases. Expand the dataset with boundary cases, adversarial tests, production failures, and important user segments. Coverage and representativeness matter more than raw count.
Should LLM evaluation run in CI/CD?
Yes. Run a small, fast regression suite for pull requests and a broader suite before release. Use quality, safety, latency, and cost thresholds as release gates, then monitor sampled interactions in production.
Is human evaluation still necessary?
Yes. Human reviewers are especially important for high-risk use cases, subjective quality, new failure modes, and calibration of automated graders. Automated evaluation reduces the volume of manual review; it does not eliminate the need for judgment. Effective LLM evaluation combines deterministic tests, model-based scoring, human review, and production monitoring.
