SHORT ANSWER
Test an enterprise LLM application in two layers. Before every release, score it offline against a golden dataset of 100 to 300 real questions. After launch, monitor quality, cost and latency on live traffic. Track groundedness, answer relevance, retrieval precision and recall, safety, latency and cost per request, run the checks automatically in CI, and red-team the system before users see it.
You test an enterprise LLM application with two kinds of evaluation. Offline evaluation scores a fixed set of real questions against expected answers before every release. Online evaluation watches quality, safety, cost and latency on live traffic. Add red teaming before launch and regression tests in CI, and you can change prompts, models or data without guessing what broke.
Why do LLM applications need a different kind of testing?
Traditional software is deterministic. The same input gives the same output, so a unit test passes or fails. LLM applications are probabilistic. The same question can come back with different wording, answers depend on which documents the retriever found, and a model update from your provider can change behaviour while your code stays the same.
So testing moves from exact matches to scored quality against a reference, measured over many examples. Thresholds then decide whether a release goes live.
Offline vs. online evaluation
| Offline evaluation | Online evaluation |
|---|---|---|
When | Before release, on every change | After launch, continuously |
Data | Golden dataset of curated questions | Real user traffic and feedback |
Goal | Catch regressions, compare options | Detect drift, failures and new topics |
Scoring | Automated metrics, LLM-as-judge, spot checks | Sampled scoring, user ratings, business KPIs |
Output | Pass or fail gate for deployment | Dashboards, alerts, new test cases |
You need both. Offline evaluation tells you whether a change made things better or worse. Online evaluation tells you what users actually ask, which always differs from what the project team expected.
How do you build a golden dataset?
A golden dataset is a versioned set of questions with reference answers and, for retrieval systems, the source documents the system should find. It’s the most valuable evaluation asset you’ll create.
Collect real questions. Use search logs, support tickets and interviews with future users, not questions the project team made up.
Cover the spread. Include common questions, rare but important ones, multi-step questions and ambiguous ones.
Add negative cases. Include questions the system must refuse or escalate, such as out-of-scope topics or requests for data the user may not see.
Write reference answers with experts. Subject matter experts define what a correct answer contains and which sources support it.
Label by category and risk. Tag each item by topic, difficulty and business impact so you can see where quality drops.
Version it. Keep the dataset in source control and add every production failure you fix as a new test case.
For most enterprise assistants, 100 to 300 items is a practical start. Teams typically build it over one to three weeks with a few hours of expert time per week.
Which metrics should you track?
Metric | What it measures | How it’s usually scored |
|---|---|---|
Groundedness (faithfulness) | Whether every claim in the answer is supported by the retrieved sources | LLM-as-judge per claim, human spot checks |
Answer relevance | Whether the answer addresses the question asked | LLM-as-judge, user ratings |
Answer correctness | Agreement with the reference answer | Comparison with reference, expert review |
Retrieval precision | Share of retrieved chunks that are relevant | Labelled relevant documents per question |
Retrieval recall | Share of relevant documents that were found | Labelled relevant documents per question |
Toxicity and safety | Harmful, biased or policy-breaking output | Safety classifiers, red-team prompts |
Latency | Time to first token and to full answer | Measured per request (p50, p95) |
Cost | Model and infrastructure cost per request or conversation | Token counts times price, plus hosting |
Set thresholds per metric and per risk category. A customer-facing answer about contracts may need near-perfect groundedness, while an internal brainstorming assistant can live with more variation. To see why retrieval drives so many of these scores, read what retrieval-augmented generation is.
What we learned evaluating a GraphRAG chatbot
We built a GraphRAG chatbot on Neo4j for a multi-team enterprise whose managers spent hours each week digging through notes and chat logs before team check-ins. Three evaluation lessons came out of it.
Test the data layer, not only the answers. Our pilot used an automated LLM graph builder. It invented organisational roles that didn’t exist and created duplicate nodes for the same person. We switched to a verified layer of fixed entities (people, roles, projects, departments) and ended with zero entity duplication across the graph.
Put test cases where the system is fragile. Simple questions, such as who leads a project, translated reliably into graph queries. Ambiguous, compound questions sometimes produced wrong queries. Schema constraints and few-shot examples helped, but that’s where most of your golden dataset should focus.
Make the reasoning visible. Leadership needed to see how the system reached each answer, so we built a transparency mode that shows the generated query and the raw results.
Using an LLM as a judge, and its limits
Scoring hundreds of answers by hand after every change isn’t realistic, so most teams use a strong model as a judge. It gets the question, the retrieved context, the answer and a scoring rubric, and returns a score with a short justification. That works for clear criteria such as groundedness and relevance. It also has known weaknesses:
Position and length bias. Judges tend to prefer the first option shown and longer answers.
Inconsistency. Scores can vary between runs, so use low temperature, fixed rubrics and repeated runs for close calls.
Calibration drift. When the judge model changes, its scores can shift. Pin its version and re-check it.
Calibrate the judge by having experts grade 50 to 100 answers and comparing their scores with the judge’s. Keep human review for high-risk categories and for any release where the automated scores are borderline.
How do you red-team an LLM application before launch?
Red teaming means trying to make the system misbehave on purpose. The OWASP Top 10 for LLM Applications gives you a useful list of risk areas. Test at least these:
Prompt injection, both typed directly and hidden in documents, emails or web pages the system reads. See what prompt injection is.
Data leakage, such as revealing documents a user isn’t allowed to see, system prompts or personal data.
Jailbreaks that get around content rules through role play, encoding or long conversations.
Harmful or biased output on sensitive topics relevant to your users.
Tool misuse for agents: actions taken without confirmation, wrong parameters or loops.
Every attack that works becomes a permanent test case in the golden dataset.
How do you run LLM regression tests in CI?
Trigger on every change to prompts, model versions, retrieval settings, chunking, tools or data pipelines.
Run the golden dataset automatically and compute all metrics, broken down by category.
Compare with the last release and block deployment if a metric falls below its threshold or drops by more than an agreed margin.
Keep a fast suite and a full suite. Run 20 to 50 critical questions on each commit and the full set before release, to control cost and time.
What should you monitor in production?
Log questions, retrieved sources, answers, latency, token usage and model version for every request, with access controls on the logs.
Score a daily sample with the same metrics as offline evaluation to detect drift.
Collect user feedback and review negative cases weekly.
Alert on spikes in cost, latency, refusals or safety flags.
Tie quality to a business KPI, such as tickets resolved or time saved per request.
How RUBICON helps with LLM evaluation
We treat evaluation as part of the build, not a step just before launch. Our QA and AI engineers build the golden dataset with your experts, add automated scoring to your CI pipeline and set up production monitoring. RUBICON is ISO 27001:2022 and ISO 9001:2015 certified and a Microsoft Solutions Partner for Cloud & AI Platforms. See our AI agents and QA and software testing services.
If you’re preparing a launch, our engineers can review your evaluation plan with you.
Frequently asked questions
How many test questions do we need to evaluate an LLM application?
Most enterprise teams start with 100 to 300 real questions with reference answers, covering the main topics, difficult edge cases and questions the system should refuse. That's enough to spot regressions between releases. Grow the set with failures found in production, and keep a smaller set of 20 to 50 critical questions that must always pass.
What is the difference between groundedness and answer relevance?
Groundedness, also called faithfulness, checks whether every claim in the answer is supported by the retrieved sources, so it catches hallucinations. Answer relevance checks whether the answer actually addresses the question. An answer can be fully grounded but beside the point, or on point but invented, so you need both metrics.
Can we trust an LLM to grade another LLM?
Partly. LLM-as-judge scales well and agrees reasonably with human reviewers on clear criteria such as groundedness, but it has known biases. It can favour longer answers, its own model family and the first option shown. Calibrate the judge against a sample of human-graded answers and keep people in the loop for high-risk decisions.
How often should we re-run LLM evaluations?
Run the offline suite on every change to prompts, models, retrieval settings or data pipelines, ideally automatically in CI. Re-run it when your model provider releases a new version, and review production samples weekly. Model updates and new documents can change behaviour even when your code hasn't changed.
Related case study

Enterprise GraphRAG Chatbot with Neo4j | Case Study
How RUBICON's Two Layer Fixed Entity Architecture eliminated data bottlenecks for a multi team enterprise, delivering a conversational AI system that gives leadership instant project clarity, without hallucinations.
More resources
