AI systems

AI agents: what they are genuinely good at, and where they fail

An honest split of the work agents close, the work they can only draft, and the work they should never touch — with the test we use to decide which is which.

· 3 min read

There is a version of this conversation that is entirely marketing, and a version that is useful. This is an attempt at the second.

Start with the transcript

Before designing anything, get two hundred real enquiries and the replies your team actually sent. Read them. Sort them into three piles.

Pile one: the agent can close this. The question is answerable from data the system holds, the answer format is predictable, and being wrong is embarrassing rather than expensive. Opening hours. Availability on a date. What is included. Where to meet.

Pile two: the agent can draft this. The reasoning is routine but the output needs judgement, or the cost of an error is real. A quote assembled from rules. A response to a complaint. A recommendation between two options. The agent does the work, a person approves it, and approval takes fifteen seconds instead of fifteen minutes.

Pile three: the agent must not touch this. Anything where being wrong is expensive, legally significant, or emotionally serious. Cancellation terms in a dispute. A medical or safety question. Anyone who is upset.

The ratio between those three piles is the entire business case. We have seen it come out at 60/30/10 and at 15/25/60, and the second business should not build an agent for that process.

Where agents genuinely earn their cost

Time of day. An enquiry arriving at 2am answered at 2am, rather than 9am, converts differently. This is often the whole return on its own.

Volume with variation. Two hundred enquiries a day that are similar but not identical is precisely the shape traditional automation cannot handle and an agent can.

Structured extraction from unstructured text. Reading a rambling email and reliably pulling out dates, party size, budget and intent is something agents do well and rules-based systems do badly.

First-draft work. Almost anything where a person currently starts from a blank page.

Where they fail

Silent confidence. The failure mode is not refusing to answer. It is answering fluently and being wrong, in a way that reads exactly like being right. Every mitigation for this is architectural rather than a matter of better prompting: retrieval over recall, citations attached to claims, evaluation sets that measure the rate, and constrained output where the shape of the answer is fixed.

Long chains. Reliability compounds downward. A step that is 95% reliable is fine. Eight of them in sequence is 66%, which is not a product. Keep chains short, checkpoint state, and make failure visible rather than letting an agent proceed on a bad intermediate result.

Anything requiring genuine accountability. Not because the technology is insufficient, but because a customer in a dispute wants a person, and giving them a machine is a decision about your business rather than about your software.

Silent drift. The model underneath changes. Behaviour moves. Without an evaluation set you find out from a customer. This is the single most common thing missing from agents built in the last two years.

The test we apply

Before automating a step, ask: if this is wrong, who finds out, how quickly, and what does it cost?

If the answer is the customer, immediately, and it costs a booking — that step gets human review. If it is nobody, eventually, and it costs nothing — automate it and stop thinking about it. Most steps are in between, and that is what the draft-and-approve pattern is for.

It is a duller framework than most of what is written about agents. It is also the one that survives contact with a live business.

Start a conversation

Tell us what the system has to do.

Describe the problem in a paragraph and we will tell you honestly whether we are the right people for it — and roughly what it costs before you spend anything.