Your AI Agents Need a Line Manager
An agent matches a £20,000 supplier invoice, approves it and the payment goes out. Whose authority was that, and what could you show an auditor? How to give every agent a manager, an identity and an off switch before it touches real money.

An invoice for £20,000 arrives from a supplier. An agent matches it to the purchase order and the goods receipt, approves it, and the payment goes out that afternoon. Nobody in your finance team touched it. A week later your finance director, your auditor or your board asks three things: was the payment correct, whose authority did it rest on, and what can you show to back that up?
I think most businesses running agents could answer the first on a good day and struggle with the other two. Agents get switched on because the use case is attractive, and the questions about authority and evidence wait until something goes wrong. This article sets out how I'd settle them first. It builds on AI Agents: What They Actually Are and What They Can Do for Your Business, which covers what agents are and where they work. This one covers how to run them.
Start with a manager
The frame I find most useful is to treat an agent as an employee, because executives already know how to manage one. Nobody joins your finance team without a manager who answers for their work, a clear brief and a period on probation. An agent that can approve invoices needs the same, and most don't get it. The owner is a named person, not a team and not another agent, and they put their name to what the agent is allowed to do.
The comparison holds until something goes wrong. A person has common sense, a feel for consequences and the ability to learn from a mistake. An agent has none of these. A new clerk who misreads the tolerance rule on goods receipts approves a handful of wrong invoices before a colleague spots it. An agent applies the same misreading to every invoice that arrives, at any hour, and if nobody is watching the totals the error surfaces at month-end reconciliation.
Accountability also can't move sideways. If one agent hands work to two others, the person answerable is still the human owner, and they need to be able to see what each of them did.
That gives me a test before anything is built. If you can't write a precise brief, covering what the agent may decide, what it may never decide and where a person must step in, don't build the agent.
Four things to settle before go-live
Remit. Write down what the agent is for, which data it may use and which decisions it may take. For the invoice agent that might be: approve matched invoices up to a stated value, never create a supplier, never change bank details. Add the context a new hire would pick up in their first month, such as which suppliers have negotiated terms, what counts as a normal price variance and who to ask when unsure. I'd expect most agent failures to come from missing context like this, with far fewer coming from faulty process logic.
Identity and reach. The shortcut I expect to see most is an agent running on someone's login, because it's quick and it works. It also does the most long-term damage. The agent inherits everything that person can reach, it can read a great deal of data very quickly, and every action is logged under their name. Builders shouldn't lend their own credentials either. Give each agent its own identity from the day it's registered, one that stays the same through new credentials, redeployments and model updates, and tie it to its owner. Grant access for the task and for as long as the task takes, and not as standing access. When one agent passes work to another, the record should show one acting for the other, because passing a person's token down the chain hides who authorised what. Check any outside agent before yours deal with it.
Limits and checks. Decide how far a mistake can travel before something stops it. Keep test and live environments apart, restrict the tools the agent can call, and set spend and volume limits with alerts when they're approached. In the invoice case that means a ceiling per invoice and per day, a person approving anything above it, and a rule that halts the agent if it approves an unusual run. Scale all of this to the risk. An agent that reads and summarises needs little, and one that moves money needs the lot. Review them continuously and not only at go-live, because agents, models and tools all change.
Record and exit. For every action, keep what was done and when, the instruction and inputs, which version of the agent and model took it, what it was permitted to do at that moment, and which process step and owner it sat under. If you can't reconstruct an action you can't fix it or safely scale it, and it won't stand up to an audit. Then the exit. Name the person who can stop the agent, make sure stopping it also cuts its access, and test that before go-live and again on a schedule, timing how long it takes. When an agent is retired, close its access, keep its records and take it off the list. A dormant agent with broad access is a risk nobody is looking at.
Back to the £20,000 payment. Was it correct? Yes if the matching rule and its tolerances are written down and checkable. Whose authority? The owner's signed brief, the agent's own identity and an approval ceiling that the payment sat under. What can you show? The record of the invoice, the order, the receipt, the version of the agent and the rule it applied.
Is the process ready?
Agents work best on processes designed for them. A process built around human habits, such as an informal approval, a note in someone's head or "ask Sam", carries that ambiguity straight into the agent. Before handing one over I'd ask three things.
Can you see the work? There should be a digital trail, and signals that tell you whether it's going well or badly. Can you define right? That means a checkable description of a good result, with tolerances, sign-offs and prohibitions written down. Can you survive being wrong? The damage from a failure should be known and recoverable, and the tasks should be able to run inside the controls you've set.
For the invoice, "matched" needs a rule an auditor could check, and a paid invoice that turns out to be wrong needs a route to get the money back. Where the answer to any question is no, redesign the process first or keep a person in that step.
Inside and outside your walls
Everything above assumes the agent works inside your business. When a supplier's agent deals with yours, you also need to know who it is, who it speaks for and whether they consented to what it asked. The same applies if agents from several companies share a process. That needs agreed identity between organisations and goes beyond what most businesses need today. The rule that carries over is that you can always say who authorised an action and prove it.
Who owns it at executive level
I don't think there's one right answer. Some businesses will assign agents by business area and others by the customer outcome they serve, and either works if it's applied consistently. Separate the person who owns design and build from the one accountable in production, because they're often different people and the gap between them is where agents get orphaned.
Security owns identity and access, HR owns the workforce side, where mixed teams change performance management, role design and skills, and operations owns the outcomes. AI accountability adds to these roles and doesn't need a new one. Even a copilot rollout needs security, HR and technology input, and treated as a pure technology project it tends to fail. Boards need a shared vocabulary for this, built around capability and accountability, with operational sign-off sitting below them.
If your agents run on several vendors' platforms, each shows you only its own. Keep one list, decide who owns it, and go looking for the orphans: agents still running with nobody answerable for them.
What I'd do on Monday
List every agent running, including pilots and those inside vendor platforms, recording what each does, what it can reach and whose credentials it uses. Name one accountable owner for each, and switch off any that don't have one. Then pick a live agent, cut its access and time how long it takes to stop and what you can see afterwards. Put the next agent you plan to give more autonomy through the three process questions first.
Further reading in the playbook
- AI Agents: What They Actually Are and What They Can Do for Your Business covers what agents are and where they work
- Section 8: Risk and Governance covers governance before you deploy autonomous systems
- Section 14: Operating Model covers roles, ownership and accountability