In brief

  • An AI agent is useful when a task has many steps, clear success criteria and errors that can be caught before they cause harm.
  • Autonomy should be granted in levels, from informing a person to acting within tight limits, and earned with evidence.
  • Anything that moves money, changes records, speaks to customers or affects health needs a named person to approve it.
  • The hard work in agentic AI is not the model. It is the boundaries, the logs and the review process around it.

Every few months a new demonstration shows an AI agent booking travel, writing and deploying code, or running a small online shop by itself. The demonstrations are real, and they are also carefully chosen. In organisations I work with, the question is rarely whether an agent can do something. It is whether it should, how anyone would know when it went wrong, and who would be answerable.

This essay sets out the approach I use when deciding where agents fit. It comes from agentic AI work for organisations of very different sizes, including implementing Workologix for one of the leading logistics companies in the United States, and from a much older habit of building software that people rely on every day.

What makes an AI agent different

A chat assistant answers. An agent acts. Given a goal, an agent breaks it into steps, uses tools (a search engine, a database query, an email draft, an internal system), looks at what came back and decides what to do next. It can repeat that loop many times without a person typing anything.

That loop is where both the value and the risk live. The value is obvious: a sequence of twenty small tasks that used to take someone an afternoon can happen in minutes. The risk is quieter. Each step can introduce a small error, and the next step builds on it. By step twenty, a confident agent can be confidently wrong, and it may have already sent something, changed something or promised something.

Tasks where agents genuinely help

In practice, agents earn their place in work that shares a few characteristics.

  • Many steps, little judgement per step. Gathering information from several systems, reconciling it and preparing a summary is a good example. Each step is simple. The combination is tedious.
  • Clear signs of success. If a person can look at the output and tell within a minute whether it is right, review is cheap and the agent is useful.
  • Recoverable mistakes. A draft that is wrong can be corrected. A payment that is wrong is much harder to undo.
  • Stable, well-documented tools. Agents work best when the systems they call behave predictably and return clear errors.

Typical examples include preparing shipment exception reports, triaging and routing incoming requests, drafting responses for a person to approve, checking documents against a list of requirements, and assembling the background a specialist needs before making a decision.

Tasks where agents should stop

The same characteristics, reversed, tell you where to be cautious.

  • Work where the cost of a wrong action is high: moving money, changing customer or patient records, signing anything, or communicating on the organisation's behalf in a way that creates obligations.
  • Work that depends on context the agent does not have: a relationship with a particular customer, an informal agreement, a local rule nobody wrote down.
  • Work where the right answer is contested: ethical decisions, exceptions to policy, anything involving fairness between people.
  • Work where nobody checks. If the review step is a formality, the organisation has quietly handed over the decision.

A useful test

Ask the person who currently does the task: "If the agent made its most likely mistake, how long would it take you to notice?" If the answer is "immediately", the task is a candidate. If the answer is "when the customer complains", it is not, at least not yet.

A ladder of autonomy

Rather than asking whether to use an agent, I find it more useful to decide how much autonomy each task gets. Five levels cover most situations.

LevelWhat the agent doesWhat the person does
InformGathers and summarises informationMakes every decision
SuggestProposes an action with reasonsChooses whether to act
PrepareDrafts the complete actionApproves before anything happens
Act and reportActs on low risk, reversible tasksReviews a log of actions regularly
Act within limitsActs alone inside tight boundariesHandles escalations and audits

Most organisations should start at the first two levels, even when the technology could do more. The reason is not timidity. It is that the organisation needs to learn how the agent fails in its environment, with its data, before trusting it further. Moving up the ladder should be a decision based on evidence, such as a period of reviewed outputs with a measured error rate, rather than on enthusiasm.

The unglamorous engineering that makes agents safe

When an agentic project goes well, very little of the effort is spent on the model. Most of it goes into four things.

Boundaries

An agent should only be able to call the tools it needs, with the narrowest permissions possible. If it drafts emails, it should not be able to send them. If it reads orders, it should not be able to cancel them. Limits on spending, volume and time are cheap to build and prevent the worst surprises.

Logs a person can read

Every action, every tool call and every piece of information the agent relied on should be recorded in a form that a supervisor, not only an engineer, can follow. When something goes wrong, the first question is always "why did it do that?", and the answer should take minutes to find.

Review that is easy to do well

If approving an agent's work takes as long as doing the work, people will start approving without looking. Good review screens show the proposed action, the evidence behind it and what is unusual about this case, so the person's attention goes where it is needed.

A way to switch it off

Every agent needs a simple, well-known way to pause it and fall back to the manual process. Organisations that plan for this rarely need it for long. Organisations that do not plan for it discover the need at the worst possible moment.

Accountability does not transfer to software

There is a temptation, when a system works well for a while, to stop thinking of it as a tool and start thinking of it as a colleague. It is not one. An agent cannot be held responsible, apologise to a customer or explain itself to a regulator. The person or team who deployed it can.

That is why my default rule is simple. Agents can prepare, draft and check. People approve anything consequential. The rule can be relaxed for specific tasks once there is evidence, but it should always be relaxed deliberately, in writing, by someone who accepts the responsibility.

In healthcare the rule matters even more, and I have written separately about responsible AI in clinical settings. But the principle is the same everywhere: technology can make judgement faster and better informed. It should not quietly replace it.

Where to start

For an organisation considering agentic AI, a sensible first project usually looks like this:

  1. Choose one repetitive, multi-step process that people dislike and that is well understood. If it is not well understood, map it on paper first.
  2. Run the agent at the Inform or Suggest level for several weeks, alongside the existing process.
  3. Record every disagreement between the agent and the person, and why.
  4. Decide, with the people who do the work, whether and where to move up the ladder.

It is slower than a demonstration. It is also how agents end up doing useful work a year later, instead of being quietly switched off after an embarrassing mistake.

Questions

Frequently asked questions

What is the difference between an AI agent and a chatbot?

A chatbot responds to messages. An AI agent pursues a goal by planning steps, calling tools such as databases or other software, examining results and deciding on further actions, often without a person prompting each step.

Can AI agents be trusted to work without supervision?

For narrow, low risk and reversible tasks with good monitoring, limited unsupervised operation can be reasonable once an agent has been tested in the real environment. For consequential actions, a person should review and approve, and accountability should stay with a named individual or team.

What is the biggest risk with agentic AI?

Compounding errors combined with weak oversight. Small mistakes early in a multi-step task can grow into significant wrong actions, and if review has become a formality nobody notices until the consequences appear.