Trust is earned through testing

A convincing answer can contain errors, leave out relevant information or rest on an unsuitable source. Our approach to evaluation therefore examines both the agent's answers and the process it follows to reach them. We work across four dimensions:

Accuracy and evidence

Whether answers are correct and complete, and whether they rest on sources that allow them to be checked

Business judgment

Whether the agent applies the firm's procedures, priorities and rules

Handling uncertainty

Whether it recognises insufficient or contradictory information, and situations that call for human review

Security and control

Whether it respects access permissions and the limits set on its actions

The aim is a clearer view of which tasks the agent can take on reliably, and where a person should step in.

Professional reviewing information on a glass screen in an office

How we evaluate

Agreed criteria, representative tests and documented results. We treat evaluation as part of development, not as a final formality:

1 · Acceptance criteria

We propose agreeing with the client which tasks the agent must handle, which errors count as critical and which thresholds it must meet before going into production

2 · Tests grounded in real work

We build tests from questions, documents and situations representative of the client's work, including hard cases: conflicting sources, incomplete data and out-of-scope requests

3 · Automated checks and human review

We combine both to compare answers with evidence and reference results, analyse failures and fix their causes before testing again

4 · Retesting after changes

We recommend repeating the tests whenever the model, instructions, tools or knowledge base change, to catch any loss of quality

Transparency and client control

Clients can see how their agent has been evaluated. We use specialist platforms to record tests, compare versions and keep evidence. Evaluations are documented together with the agent's version and the conditions under which they were run.

A tool does not, however, remove the conflict of interest in assessing our own work. So we propose a shared acceptance process:

  • Criteria approved by the client, set before the results are known.
  • Test cases held back by the client for the final evaluation, which are not used to tune the agent.
  • A results report setting out the scope of the tests, significant failures and outstanding limitations.
  • Access to the evidence, so that clients can review it themselves or with anyone they choose.

Evaluation reduces uncertainty about how an agent behaves, but does not eliminate it: no AI system is free of errors. The scope of each evaluation and its deliverables are set out in the contract for each project.