You evaluate a new AI model in one afternoon with your own test set of about ten real tasks, blind scoring and a look at cost, speed, data terms and typical errors. Benchmarks and rankings say little about how a model performs on your tasks. A small, repeatable test does.
What does the afternoon look like?
The plan fits into four to five hours. You need no programming skills, but someone who knows your tasks.
| Time | Step | Result |
|---|---|---|
| 30 minutes | Set goal and success criteria | One page: task, quality standard, must-have criteria |
| 60 minutes | Collect ten test tasks and note expected results | Test set with scoring grid |
| 60 minutes | Run each task with the candidate and the current model | Answers, several per task |
| 45 minutes | Score blind | Points per task and criterion |
| 30 minutes | Check cost, speed and data terms | Comparison table |
| 30 minutes | Decision and documentation | Switch, keep or pilot |
How do you build a good test set?
Choose tasks that genuinely occur in your daily work: a customer reply, a summary, extracting details from a text. Take six or seven typical cases, two difficult edge cases and one or two traps in which the model should admit it does not know something. For each task, write down what a good answer looks like.
Anonymise all examples. Customer names, addresses and confidential figures do not belong in a test until contract and data protection are settled. How to choose models in general is described in Claude, ChatGPT or Gemini for business use.
How do you score fairly?
Give all candidates the same instruction and the same documents. Run each task two or three times, because answers vary. Hide which model produced which answer and award points from one to five for correctness, completeness, tone and format. The providers' documentation describes the approach in more detail, for example Anthropic on developing tests and OpenAI on evals.
What else belongs in the comparison besides quality?
Cost, speed and data terms often matter more than small quality differences. Count the tokens processed and scale up with the provider's current prices to your expected monthly volume. Measure response time. Clarify whether inputs are stored or used for training, where processing takes place and whether a data processing agreement exists.
Pay particular attention to failure modes: invented facts, ignored instructions, unstable format, unjustified refusals and information lost in long texts. On hallucinations, read AI hallucinations: what they are and how to protect your business. Background on how models work is in What is an LLM?.
- Goal and must-have criteria recorded on one page
- Ten anonymised test tasks with expected results collected
- Each task run several times with candidate and current model
- Answers scored blind, errors noted
- Cost, speed and data terms compared
- Decision documented with reasons
- Date set for the next review
How do you decide in the end?
Decide beforehand what counts as a reason to switch: for example, that the new model must be no worse on the must-have criteria and noticeably better or cheaper in at least one area. A tie favours what you have, because switching costs time for adaptation and training. If you switch, begin with a small pilot, such as one task and two people, and compare with the starting values after four weeks. Document decision, test set and date, and the next evaluation can be repeated in a few hours.
Conclusion: your own test instead of a ranking
A repeatable test set is the best insurance against model hype. Once you have built it, you can assess every new version in a few hours. If you would like help with the test set or the rollout, see our AI automation service or describe your tasks.




