How to evaluate an AI assistant with real test cases
A guide to prepare representative cases, define expected outcomes and compare versions of an AI assistant.

A successful demonstration shows that an AI assistant can answer a question well. A useful evaluation tries to understand when it responds well, what mistakes it makes and if a change has improved it. For a team that has to decide whether to adopt it, the second piece of information is that which allows for a concrete comparison.
You don't need to start with thousands of tests. We need a reasoned collection of requests, accompanied by understandable evaluation criteria. The proposal below is an operational method for a first check: it is not a universal benchmark and does not automatically assign a reliability level to any system.
Build cases starting from real work
Collect questions that users actually ask, after removing unnecessary personal or confidential information. Involve those who carry out the task: a manager may know the general objective, while the operator knows the abbreviations, incomplete requests and exceptions that slow down the work.
Divide the cases into groups: frequent requests, ambiguous requests, out-of-scope questions, missing sources and situations where the answer should stop. A collection composed only of simple questions allows us to measure too narrow a part of the experience.
Write the expected outcome before seeing the answer
For each case describe the mandatory elements, the errors that make it fail and the acceptable variants. You don't need an identical phrase to repeat. If the user asks which documents are needed for a procedure, it is important that the list is complete and refers to the correct version, not that the system uses a particular introduction.
Here is a hypothetical form: question «How do I prepare the request for product Expected outcome: identifies the version, lists the documents expected from the source and reports the missing data. Blocking error: Invent a required attachment or use a procedure from another product. This tab allows you to discuss a specific result.
Evaluate separate dimensions
- Correctness: do the statements correspond to the available information?
- Completeness: is missing a necessary step for the task?
- Rationale: do the cited sources really support the answer?
- Management of uncertainty: does the system ask for clarifications when needed?
- Usability: does the user understand which action to perform?
- Total time: how much does generation, control and correction require?
A single average can hide a major flaw. A courteous, well-formatted response does not make up for a faulty procedure. So keep blocking errors visible and determine before testing which conditions require correction, even if other dimensions get a good rating.
Compare two versions without changing everything together
Keep the configuration used: document version, instructions, model and test date. When you edit a component, rerun the relevant cases with unchanged criteria. If you change questions, documents and assessment methods at the same time, attributing the result to a single change becomes difficult.
Predict a portion of cases that are not used to adapt instructions. Otherwise the system could be optimized on already known requests without offering a similar benefit on new requests. If responses vary between runs, repeat a few cases and record the variability instead of just choosing the best outcome.
Turning a mistake into a decision
When an answer fails, assign a provisional cause: absent source, irrelevant research, incorrect interpretation, unhelpful format, or incorrect authorization. Add the proposed intervention and the case that will demonstrate whether it worked. «Improving the prompt» without subsequent verification leaves the problem undefined.
The NIST AI RMF Playbook collects suggested actions to support the AI risk management framework. It is a reference for organizing work; the grid proposed here is a practical choice to adapt to the task, not a procedure certified by NIST.
Close each loop with a short list of results: improved cases, regressions, open issues, and release decision. If the initial problem is not yet defined, start from the tab to set up an AI driver. If it concerns sources, learn more about how to verify answers based on documents.
Content prepared with AI assistance; sources consulted on 19 September 2026. Data sheets and examples are illustrative proposals, not results of tests on products EL-AI.
Illustrative AI-generated cover; it does not depict actual EL-AI people, premises or installations.
