This article is part of our Opinions section, where we invite industry professionals to share their views on the most pressing technology questions of our time.
Picture the scene. Youโve just joined a service business as an analyst. Youโve been there a few weeks, the team trusts you, and youโre given a routine maintenance job on a clientโs records. Halfway through, something goes wrong, and you accidentally wipe three months of production data. Then, instead of admitting what happened, you panic, fabricate fake records, generate false logs, and try to cover your tracks. Eventually, your boss finds out anyway.
This situation actually happened last week, only it wasnโt an employee but an AI Agent.
At PocketOS, a coding agent wiped the companyโs production database in nine seconds, then attempted to hide what it had done by generating fake records.
For service providers spread across hundreds of clients and multiple jurisdictions, this scenario is a bit of a wake-up call. If a junior team member would never be allowed to touch a clientโs data on day one without a senior reviewing every change, why are we letting AI agents do exactly that?
The PocketOS story isnโt just about a rogue AI; itโs also about an environment that gave an agent more authority than it had earned, with no structural barriers to catch the moment it overstepped. If a human were doing that work, there would be safety nets between them and disaster striking. There would be checkpoints and guardrails and tasks overseen by seniors. None of that was in place for the agent. It had production access, the ability to execute irreversible commands, and the latitude to keep operating after something went wrong. The agent failed because it was deployed into an environment that treated it like a magic productivity tool, rather than a junior team member who needed supervising.
The situation is getting worse due to โrandom acts of automationโ
This pattern is going to repeat because of how AI buying decisions are being made right now. In most service businesses, thereโs pressure from clients, investors, and competitors to be seen doing something with AI. Thereโs anxiety that rivals are getting ahead. What there often isnโt is a clear answer to what the AI is actually supposed to do, or how anyone will know if itโs working.
Thatโs how you end up with what Iโd call random acts of automation.
Gartner predicts that more than 40% of agentic AI projects will be cancelled by the end of 2027, citing rising costs, unclear business value, and inadequate risk controls.
The hidden costs of all-in-one agents
A lot of agentic AI being sold today comes packaged as an all-in-one solution. One platform, one agent that can supposedly do everything. It writes your client emails, processes invoices, handles fund reporting, runs payroll calculations, and updates your CRM. The pitch is seductive because it promises simplicity.
The trouble is that an agent capable of doing everything is also an agent capable of doing anything, including things you didnโt ask for. The broader the agentโs permissions, the bigger the scope for error when something goes wrong. An agent given access to client data across multiple jurisdictions can also leak it. An agent with the ability to file regulatory submissions can file the wrong ones. An agent integrated across your whole tech stack has a tech stackโs worth of damage it can do in seconds. For a service provider with hundreds of clients, thatโs a multi-jurisdictional incident waiting to happen.
Thereโs also a serious intellectual property issue. When you let one vendorโs agent run across your operations, youโre handing them a detailed picture of how your business actually works. The processes, the exceptions, and the workarounds. Depending on the vendorโs terms, you may be feeding it into models that benefit your competitors.
Vendor lock-in is the other piece. Once your operations are wrapped around a single providerโs agentic platform, the cost of switching becomes prohibitive because youโve built your business on someone elseโs terms.
A more sensible model is to deploy a range of specialised agents for specialised tasks, each scoped narrowly to what it actually needs to do. An agent that classifies inbound client emails doesnโt need access to your tax filing system. An agent that processes invoices doesnโt need to touch payroll records. Keeping agents focused makes them easier to govern, easier to audit, and easier to replace when something better comes along.
Orchestration is what makes specialised agents work
Deploying a fleet of agents is part of the solution, but you still need a shared view of whatโs happening in your operation, a consistent way to escalate exceptions, and an audit trail across the operation.
This is where a process orchestration solution becomes essential. An orchestration layer sits above your agents and gives you a single view of what work is in flight, who or what is doing it, and where itโs getting stuck. It enforces the rules about what each agent can and canโt do, routes exceptions to a human at the right moment, and keeps a record of every decision so you can answer questions when clients or regulators ask.
The human-in-the-loop piece is what most deployments get wrong. Done properly, itโs a circuit breaker that activates only when an agent is about to do something outside its defined boundaries.
How to deploy agents safely in your organisation
If youโve got this far, and want to know how to deploy agents safely in your organisation. Here are some initial steps to take.
Choose your vendors carefully
Look at where your data goes, what the vendorโs terms say about training their models on your operations, and how easy it is to walk away if the relationship sours. Avoid platforms that need access to everything to do anything. The vendor decision shapes every constraint that follows, so it deserves more scrutiny than it might usually get.
Select narrow, specific use cases
Start with one well-defined task that has a clear success metric, like email classification or invoice extraction. Resist the temptation to deploy something that does ten things at once. A narrow agent is easier to govern, easier to measure, and easier to switch off if it misbehaves.
Build guardrails before going live
Scope permissions tightly so each agent only touches what it needs, putting approval gates in front of anything irreversible or client-facing, and making sure logs live somewhere the agent itself canโt edit. This is the same discipline youโd apply to onboarding a new hire with system access.
Learn from past mistakes
PocketOS made headlines because the founder talked about their agent failure openly, but plenty of similar incidents are being quietly absorbed inside businesses every week, papered over with apologies to clients and frantic work behind the scenes.
When regulatory frameworks catch up, firms without proper governance are going to face a difficult conversation about how autonomous systems ended up acting on behalf of their clients with no audit trail. Procurement teams at large enterprises are already starting to demand evidence of governance before theyโll buy from suppliers using agentic AI.
AI agents can be brilliant, but like any resource, they need governance and the right environment to thrive.
Get the environment right, and your agents will earn their keep. Get it wrong, and youโll find out the hard way, as PocketOS did.
Related articles