Jon Dedman, Director at Cloudhouse: “If you are down, your customers and your regulator hold you responsible, regardless of whose automation caused it.”

Two years after the CrowdStrike outage brought millions of Windows machines to a standstill, the lessons for IT leaders extend well beyond the need for more cautious software updates. The incident exposed a more fundamental problem: organisations can only manage the risks they can see, yet an increasing proportion of change across modern technology estates happens automatically, outside traditional change-management processes.

That problem is becoming more acute as businesses rely on SaaS platforms, cloud environments, third-party services and automated processes that can alter systems without waiting for an internal change board – or for the IT team to return from a summer holiday. Even a relatively small configuration or policy change can have consequences that remain hidden until they combine with other changes or expose an underlying dependency.

For Jon Dedman, Director at Cloudhouse, this is creating a new challenge for IT leaders. He argues that the industry has focused too heavily on controlling when changes happen, while failing to understand what is actually changing. In his view, the answer is not to try to approve every change in advance, but to build the visibility needed to detect and verify changes as they happen.

Here, Dedman explains why he believes the industry learned the wrong lesson from CrowdStrike, how “invisible change” can quietly undermine an organisation’s understanding of its technology estate, and why AI-driven automation will force businesses to rethink traditional approaches to change management. He also shares practical steps IT teams can take to reduce risk when key staff are away and automated systems continue to evolve around the clock.


Related reading: “Change for the sake of change and rushed modernisation will only increase risk


It’s now two years since the CrowdStrike outage highlighted the risks of a single automated update bringing down millions of systems. Looking back, what do you think the industry has genuinely learned from that incident, and where do you still see organisations repeating the same mistakes? 

I think the industry learned the wrong lesson. The response to CrowdStrike has been about update discipline: deployment rings, phased rollouts, more fault-tolerant procedures. That is real progress and it was overdue. But CrowdStrike was not a faulty binary. It was faulty content, and content updates almost never pass through those rings, because nobody classifies them as risky. Around 8.5 million Windows machines went down because of a change that most change processes would have waved straight through.

That gap is where organisations are still repeating the mistake. The commercial pressure does not help. The risk of an outage from applying a patch gets weighed against the risk of a breach from not applying it. The security risk usually wins, so critical updates bypass the careful process by design.

The deeper issue is classification. A configuration file, a policy tweak, a definition update, these are treated as low risk because they are small, not because anyone has assessed what they touch. Most outages are caused by configuration change, not by new code. The real lesson from CrowdStrike should have been that you need to understand a change well enough to classify it properly. Two years on, most organisations still classify by change type rather than by potential impact.

Many IT leaders have traditionally avoided making major changes on Fridays or before holiday periods because key staff may be unavailable if something goes wrong. As AI agents and automated update processes become more common, does that conventional wisdom still hold, or has the nature of change management fundamentally shifted? 

It still holds, but for a shrinking share of the problem. Read-only Friday and reduced change over holiday periods do exactly what they were designed to do, which is make sure someone competent is around when a human makes a change. The difficulty is that humans are no longer making most of the changes.

Vendors update SaaS platforms on their own schedule, automated processes run without intervention, and none of that pauses because your team has gone home. The freeze can become a comfort blanket: the change calendar looks quiet while the estate carries on changing underneath it.

Change management has had to evolve, because traditional processes cannot keep pace with the volume and rate of change in a modern estate. What is replacing it is change enablement, which in plain terms means loosening the gate at the front and strengthening the checks behind it. Rather than every change waiting on a review board, changes proceed and automation reconciles what actually happened against what was expected.

That is the real shift. You are no longer trying to approve your way to safety. You are trying to see clearly enough to catch a problem when it happens, rather than when it becomes an outage.

You’ve spoken about the rise of “invisible change” across modern IT environments. What do you mean by that, and why is it becoming increasingly difficult for organisations to understand exactly what’s changing across their technology estates? 

Invisible change is any change that happens without being recorded, so afterwards nobody can point to it. It is rarely malicious and usually not even careless. It is the engineer who thinks, while I’m here, I’ll just update this as well, and does not put it in the ticket.

It is getting harder to avoid because the estate has grown well beyond the part you control. SaaS solutions, third-party managed services and cloud environments all change on somebody else’s schedule. Cloud can evolve rapidly, and that is one of its genuine benefits, but it carries the risk of becoming a Wild West of change, impossible to manage despite everyone’s best intentions.

The damage is rarely immediate. A small undocumented change can sit there for months causing nothing at all. The real cost compounds. Once the record of your environment is wrong, every subsequent change is planned against a picture that no longer matches reality, so you are making risk decisions on assumptions that stopped being true some time ago. That is what makes it insidious. It is not the first invisible change that takes you down, it is the fiftieth.

AI is enabling systems to make decisions and implement changes at a speed that humans simply can’t match. How can organisations embrace greater automation without losing the visibility and governance needed to maintain operational resilience? 

The honest starting point is arithmetic. An agent can make a hundred changes in the time it takes a person to review one. Any governance model built on a human approving things in advance stops being possible, not because it is a bad model but because there are not enough hours in the day.

So, the control has to move from approval before the change to verification after it. You let the change happen, and you prove quickly and automatically that the environment is still in the state you expect it to be in.

That only works if you know what state you expect. It means understanding the applications you run, how they interconnect, and what depends on what. Building that into a CMDB is difficult and it never really finishes, because the estate keeps moving underneath you. Discovery tooling helps, but process knowledge matters just as much.

The mistake I would warn against is waiting. Monitoring should not wait until the CMDB is in a good state. Run it in parallel. It feeds useful data back into the CMDB work, and in the meantime it gives you audit coverage across your critical applications. Imperfect visibility today is worth considerably more than a perfect map in eighteen months.

Summer holidays often leave IT teams operating with reduced staffing and key knowledge holders away from the office. What practical steps should technology leaders take to reduce operational risk during these periods, particularly when automated systems continue to make changes around the clock?

Four things, and all of them are easier before people leave than after.

  • Capture a known-good baseline of your critical systems while the people who understand them are still in the building. You cannot tell what has drifted if nobody ever wrote down what right looked like.
  • Monitor against that baseline rather than waiting for users to report a problem. Drift detection gives early warning, which is precisely what a reduced team needs: the chance to fix something small before it escalates into something that needs the person who is away.
  • Agree the escalation path in advance and write it down. Who gets called, for which system, at what threshold. Cover arrangements fail on ambiguity far more often than on absence.
  • Make sure whoever is on cover can see what changed without needing the expert. If something does break, a clear record of recent changes lets them either identify the likely cause or eliminate options, which shrinks the investigation dramatically. That is the difference between a nervous weekend and a twenty-minute fix. 

Looking ahead, as AI agents become responsible for managing increasingly complex technology environments, what new capabilities will organisations need to ensure they can trust automated change without sacrificing control or accountability?

Some control will move away from the organisation, and people should be honest about that. As infrastructure accelerates and applications move onto SaaS platforms or into third-party management, you no longer control how a change is made or when it lands. What does not move is accountability.

If you are down, your customers and your regulator hold you responsible, regardless of whose automation caused it.

Closing that gap needs three capabilities. First, machine-readable evidence of what changed and when, across the systems you own and the ones you do not, because after an incident you need a record rather than a recollection. Second, continuous drift detection, because if change is constant then checking at review points tells you very little. Third, the ability to reconstruct a timeline: what the environment looked like before, what altered, and in what order.

None of that is exotic. It is the same discipline the industry has always relied on, applied at a speed humans are no longer in the loop for.

About The Author

Avatar photo
Ricardo Oliveira

Ricardo Oliveira is a Senior Director at TechFinitive, where he frequently collaborates with TechFinitive's editorial team to write and produce content. He's based in Sydney, Australia.

Read more from this author.

We take journalism seriously. To learn more on why you should trust us, head to our editorial guidelines page or meet our team.