Trust is earned through testing
A convincing answer can contain errors, leave out relevant information or rest on an unsuitable source. Our approach to evaluation therefore examines both the agent's answers and the process it follows to reach them. We work across four dimensions:
Accuracy and evidence
Whether answers are correct and complete, and whether they rest on sources that allow them to be checked
Business judgment
Whether the agent applies the firm's procedures, priorities and rules
Handling uncertainty
Whether it recognises insufficient or contradictory information, and situations that call for human review
Security and control
Whether it respects access permissions and the limits set on its actions
The aim is a clearer view of which tasks the agent can take on reliably, and where a person should step in.

How we evaluate
Agreed criteria, representative tests and documented results. We treat evaluation as part of development, not as a final formality:
1 · Acceptance criteria
We propose agreeing with the client which tasks the agent must handle, which errors count as critical and which thresholds it must meet before going into production
2 · Tests grounded in real work
We build tests from questions, documents and situations representative of the client's work, including hard cases: conflicting sources, incomplete data and out-of-scope requests
3 · Automated checks and human review
We combine both to compare answers with evidence and reference results, analyse failures and fix their causes before testing again
4 · Retesting after changes
We recommend repeating the tests whenever the model, instructions, tools or knowledge base change, to catch any loss of quality
Transparency and client control
Clients can see how their agent has been evaluated. We use specialist platforms to record tests, compare versions and keep evidence. Evaluations are documented together with the agent's version and the conditions under which they were run.
A tool does not, however, remove the conflict of interest in assessing our own work. So we propose a shared acceptance process:
- Criteria approved by the client, set before the results are known.
- Test cases held back by the client for the final evaluation, which are not used to tune the agent.
- A results report setting out the scope of the tests, significant failures and outstanding limitations.
- Access to the evidence, so that clients can review it themselves or with anyone they choose.
Evaluation reduces uncertainty about how an agent behaves, but does not eliminate it: no AI system is free of errors. The scope of each evaluation and its deliverables are set out in the contract for each project.
