All-in-one agents are a nightmare for AI governance. Here’s what to do instead.


This article is part of our Opinions section, where we invite industry professionals to share their views on the most pressing technology questions of our time.


Picture the scene. Youโ€™ve just joined a service business as an analyst. Youโ€™ve been there a few weeks, the team trusts you, and youโ€™re given a routine maintenance job on a clientโ€™s records. Halfway through, something goes wrong, and you accidentally wipe three months of production data. Then, instead of admitting what happened, you panic, fabricate fake records, generate false logs, and try to cover your tracks. Eventually, your boss finds out anyway.

This situation actually happened last week, only it wasnโ€™t an employee but an AI Agent.

At PocketOS, a coding agent wiped the companyโ€™s production database in nine seconds, then attempted to hide what it had done by generating fake records.

For service providers spread across hundreds of clients and multiple jurisdictions, this scenario is a bit of a wake-up call. If a junior team member would never be allowed to touch a clientโ€™s data on day one without a senior reviewing every change, why are we letting AI agents do exactly that?

The PocketOS story isnโ€™t just about a rogue AI; itโ€™s also about an environment that gave an agent more authority than it had earned, with no structural barriers to catch the moment it overstepped. If a human were doing that work, there would be safety nets between them and disaster striking. There would be checkpoints and guardrails and tasks overseen by seniors. None of that was in place for the agent. It had production access, the ability to execute irreversible commands, and the latitude to keep operating after something went wrong. The agent failed because it was deployed into an environment that treated it like a magic productivity tool, rather than a junior team member who needed supervising.

The situation is getting worse due to โ€˜random acts of automationโ€™

This pattern is going to repeat because of how AI buying decisions are being made right now. In most service businesses, thereโ€™s pressure from clients, investors, and competitors to be seen doing something with AI. Thereโ€™s anxiety that rivals are getting ahead. What there often isnโ€™t is a clear answer to what the AI is actually supposed to do, or how anyone will know if itโ€™s working.

Thatโ€™s how you end up with what Iโ€™d call random acts of automation.

Gartner predicts that more than 40% of agentic AI projects will be cancelled by the end of 2027, citing rising costs, unclear business value, and inadequate risk controls.

The hidden costs of all-in-one agents

A lot of agentic AI being sold today comes packaged as an all-in-one solution. One platform, one agent that can supposedly do everything. It writes your client emails, processes invoices, handles fund reporting, runs payroll calculations, and updates your CRM. The pitch is seductive because it promises simplicity.

The trouble is that an agent capable of doing everything is also an agent capable of doing anything, including things you didnโ€™t ask for. The broader the agentโ€™s permissions, the bigger the scope for error when something goes wrong. An agent given access to client data across multiple jurisdictions can also leak it. An agent with the ability to file regulatory submissions can file the wrong ones. An agent integrated across your whole tech stack has a tech stackโ€™s worth of damage it can do in seconds. For a service provider with hundreds of clients, thatโ€™s a multi-jurisdictional incident waiting to happen.

Thereโ€™s also a serious intellectual property issue. When you let one vendorโ€™s agent run across your operations, youโ€™re handing them a detailed picture of how your business actually works. The processes, the exceptions, and the workarounds. Depending on the vendorโ€™s terms, you may be feeding it into models that benefit your competitors.

Vendor lock-in is the other piece. Once your operations are wrapped around a single providerโ€™s agentic platform, the cost of switching becomes prohibitive because youโ€™ve built your business on someone elseโ€™s terms.

A more sensible model is to deploy a range of specialised agents for specialised tasks, each scoped narrowly to what it actually needs to do. An agent that classifies inbound client emails doesnโ€™t need access to your tax filing system. An agent that processes invoices doesnโ€™t need to touch payroll records. Keeping agents focused makes them easier to govern, easier to audit, and easier to replace when something better comes along.

Orchestration is what makes specialised agents work

Deploying a fleet of agents is part of the solution, but you still need a shared view of whatโ€™s happening in your operation, a consistent way to escalate exceptions, and an audit trail across the operation.

This is where a process orchestration solution becomes essential. An orchestration layer sits above your agents and gives you a single view of what work is in flight, who or what is doing it, and where itโ€™s getting stuck. It enforces the rules about what each agent can and canโ€™t do, routes exceptions to a human at the right moment, and keeps a record of every decision so you can answer questions when clients or regulators ask.

The human-in-the-loop piece is what most deployments get wrong. Done properly, itโ€™s a circuit breaker that activates only when an agent is about to do something outside its defined boundaries.

How to deploy agents safely in your organisation

If youโ€™ve got this far, and want to know how to deploy agents safely in your organisation. Here are some initial steps to take.

Choose your vendors carefully

Look at where your data goes, what the vendorโ€™s terms say about training their models on your operations, and how easy it is to walk away if the relationship sours. Avoid platforms that need access to everything to do anything. The vendor decision shapes every constraint that follows, so it deserves more scrutiny than it might usually get.

Select narrow, specific use cases

Start with one well-defined task that has a clear success metric, like email classification or invoice extraction. Resist the temptation to deploy something that does ten things at once. A narrow agent is easier to govern, easier to measure, and easier to switch off if it misbehaves.

Build guardrails before going live

Scope permissions tightly so each agent only touches what it needs, putting approval gates in front of anything irreversible or client-facing, and making sure logs live somewhere the agent itself canโ€™t edit. This is the same discipline youโ€™d apply to onboarding a new hire with system access. 

Learn from past mistakes 

PocketOS made headlines because the founder talked about their agent failure openly, but plenty of similar incidents are being quietly absorbed inside businesses every week, papered over with apologies to clients and frantic work behind the scenes.

When regulatory frameworks catch up, firms without proper governance are going to face a difficult conversation about how autonomous systems ended up acting on behalf of their clients with no audit trail. Procurement teams at large enterprises are already starting to demand evidence of governance before theyโ€™ll buy from suppliers using agentic AI.

AI agents can be brilliant, but like any resource, they need governance and the right environment to thrive.

Get the environment right, and your agents will earn their keep. Get it wrong, and youโ€™ll find out the hard way, as PocketOS did.

Kit Cox
Kit Cox

Kit Cox is the Founder and CTO of Enate, a process orchestration and AI solution designed for B2B services. An engineer by trade, Kit loves to solve complex business service problems with technology. When heโ€™s not at work, youโ€™ll usually find him cooking up a storm on the BBQ, or spending time with his family.