top of page

AI Governance Statement 33: Create business continuity plans

Writer: Bruce Mullan
Bruce Mullan
Aug 20
6 min read

AI Governance Statement 33: Create business continuity plansDigital transformation in its purest form is digitising a paper-based process. Before IT systems, the biggest operational risk to business continuity was running out of paper or a power outage. Technology is now so embedded in how we work today, that any IT outage has the potential to massively disrupt our work and the goods and services we supply. Of course, these things always depend on context. A hospital operating theatre unit has a different continuity risk profile to an accounts payable function.


Traditional business continuity planning has always focuses on a system's availability or "up-time". The AI conversation now shifts slightly beyond just having a system up and running, to "what happens if the AI outputs cannot be trusted or are unsafe?".

An AI failure in an operable AI system, could occur because of a model update, corrupted data, prompt manipulation, hallucinations, a compromised integration, changes to an external AI service, or a failure in an automated AI agent.


Business Continuity Plans now need to be extended to cover more potential outage scenarios due to more possible failure points when using AI systems.


Five reasons AI systems go "down"

The data it needs is not available

The AI system fails because it might be referencing a key data set, waiting for data from another system or an integration point has failed.


Your AI software vendor goes down

Every company is dependent on their vendor software's up-time.  What happens if the AI vendor, cloud provider, integration layer or their underlying AI model experiences an outage?

How many people remember CrowdStrike's blue screen death 2024? CrowdStrike is a 3rd party software used by Microsoft. In July 2024, a flawed update by CrowdStrike caused a massive global IT outage crashing an estimated 8.5 million Microsoft Windows devices worldwide. It was called the largest outage in the history of IT. As an aside, the outage cost tens of billions of dollars, but CrowdStrike liability was capped to "fees paid", effectively providing only a money back

guarantee. See source.


AI output failure

What happens if the AI system is operational but starts producing unreliable, incorrect or unsafe outputs? Ideally, you'll be monitoring for this. In the latter case, a kill switch might need to be activated. Typically, upgrades, model drift, data changes and degradation might be at fault. In any event, until it's fixed, the outputs aren't usable.  (Learn about AI shutdowns and kill switches in my article)


Security and cyber incidents

The AI account, model, prompt, data source or integration is compromised by a cyber-security incident. Ransom-ware attacks are notorious for locking companies out of their own systems.


Presidential directives

In June 2026, the Trump administration used an emergency export-control directive to force Anthropic to shut down its AI models globally. Instantly, companies using Anthropic AI software had no access to it outside of the US. The system was down for three weeks. Source 

A whole new level of geo-political risk emerges for AI systems when foreign governments intervene. Companies using Anthropic software as part of their operations without a Business Continuity Plan would have been left floundering.


Ways to keep working if there's an AI system outage

Manual fallback processes

Recently, the PocketOS AI car rental system went down overnight when a rogue AI deleted all customer records. So how did the company process in-coming rental returns and book new rentals while the system was being restored (which took 3 days)? (See source)


Pull out your pens and notebooks. Tasks that were automated will now need to be processed the "old way". This is not as simple as it sounds because people might have forgotten the old way or the staff have never done this before. So procedures are essential. If the Finance system is down, admin staff will need to pay invoices how? If the company website is down, staff might need to revert to taking telephone orders. The point here is to ensure any manual procedures are documented (up-to-date?), trained and tested. A manual fallback process that doesn't work is a double disaster.


Alternative AI systems

The old Plan B. Establish a secondary AI provider, model or platform that can be used if the primary AI service fails. Again, difficult to implement if the primary AI system has been trained on your data but might be useful for high-risk scenarios where business continuity is critical for human health and safety.


Recovery and testing

The fewer hours you spend without your main system the better. This is where IT earns its stripes. Recovering failed systems from backups while minimising any disruption. This may not just be your internal IT team who is on the hook, but your software vendor too. How quickly can the AI service be restored is usually a question asked at the procurement stage in the non-functional requirements and is tied to contractual up-time and availability performance levels.

This may sound obvious but have the recovery arrangements actually been tested end-to-end? It's a double disaster if they haven't. In the PocketOS AI car rental system catastrophe, the rogue AI deleted all the back ups too ... so the current system state had to be rebuilt manually from bank statements and email records.


The P part of BCP

The AI Governance statement 33 requires you to "develop plans to ensure critical systems remain operational during disruptions." 


This includes (1) identifying and managing potential risks to AI operations.

AI introduced a new range of vulnerabilities and risks, so it is more important to be prepared for an outage than ever before (see CrowdStrike incident above).


The first part of risk planning is to identify critical AI systems and use cases. This will be part of your AI systems inventory (learn more about this here). You must identify AI systems that support critical business processes and understand what happens if they stop working.

The second part is to understand how business decisions are made to suspend or shut down an AI system. Identify the person, the conditions and the process with the authority to suspend AI-generated decisions and revert to human decision-making protocols.


Lastly, define what bad looks like. How long is an acceptable downtime before things get ugly? An outage of 24 hrs might be ok but it could depend on which 24hrs. Similar, in high risk contexts, an outage of 30 minutes could be fatal (just ask Optus or Telstra).


In summary, for every high-risk or business-critical AI use case, document:


  1. AI use case

  2. Business process 

  3. Criticality

  4. Dependencies

  5. Failure scenarios

  6. Business impact

  7. Manual fallback

  8. Alternative solution

  9. Human decision-maker

  10. Recovery objective

  11. Testing frequency

  12. And lastly, define clear triggers for switching the AI off.


If AI-generated outputs fall below an agreed accuracy threshold, critical data becomes unavailable, the AI provider experiences a security incident, or the system behaves unexpectedly, the organisation can immediately suspend the AI process and activate the documented manual procedure.

The standard requires (2) defining disaster recovery, backup and restore, monitoring plans. I covered this above.


It also requires (3) testing business continuity plans for relevance, which I covered earlier.

Lastly, BCP for AI systems is not a set and forget scenario. Constant vigilance is required. The standard requires (4) regularly reviewing and updating objectives, success criteria, failure indicators, plans, processes and procedures to ensure they remain appropriate to the use case and its operating environment.


AI governance builds in contingency planning

Mature AI governance means you will know what models you are using, where their dependencies exist, what data is being processed and which operations will be affected if access were interrupted. 


What people often overlook with good AI governance is that it makes sure the business can keep operating when the AI fails. That's the value of adopting the Australian AI Governance standard because it builds in contingency planning.


Stay safe,


Bruce


AI. Use responsibly.



ABOUT ME

I partner with mid-size companies to confidently adopt AI, prevent high-profile failures and avoid the expensive mistake.


I write all my own content, you can tell by the odd typo and occasional missing word. I use AI for my research.


To learn about my upcoming public AI Governance workshops visit: Public workshops


To learn more about AI Governance, check out my Hitchhikers Guide to AI Governance Podcast.



Bruce Mullan hosts Hitchhikers Guide to AI Governance podcast
Bruce Mullan hosts Hitchhikers Guide to AI Governance


 
 
 

Comments


CONTACT

If you have a question or request  please contact us today!

© 2026 BY TRIPLE P GLOBAL PTY LTD T/AS Ai Governance Partners -

ABN 96 119 485 791

Thanks for contacting us. we'll be in touch.

bottom of page