top of page

AI Governance Statement 12: Define success criteria

Writer: Bruce Mullan
Bruce Mullan
Jul 23
6 min read

If you don't know where you're landing the plane, you'll probably end up somewhere else.

The groundhog day conversation I've had over a dozen times in the last 15 years after a newly launched system failed within 12 months. During the post-mortem coffee conversation:


  1. IT: We desperately needed a new system

  2. ME: What were you seeing to indicate that?

  3. IT: Staff were constantly complaining and the vendor support was hopeless

  4. [ME thinking: Poor data, broken processes, 100+ work-arounds]

  5. IT: We had to modernise. Get the latest systems. We had no choice.

  6. ME: Oh, I see. So you reverted back to the old system when it failed?

  7. IT: Yes.


A new system is always good story to tell. It tells the Board things are moving forward, something positive, modern systems are coming. IT people get experience in contemporary technologies. The latest tech, cloud computing, a contemporary user interface, a vendor with snazzy support tools. What could possibly go wrong?

But once you start implementing, you find where the bodies are buried. Someone believed a new system would fix:


  •  Years of unfettered document and data accumulation,

  •  Poor business processes,

  •  Integration with other ageing systems and

  •  Operating knowledge inside people's heads


Then, the worst part happens. You hit these four big challenges with data, people, processes and integrations and so on, but you've gone too far. You've passed the point of no return. You now have to make it "work".

Been there, done that. Somehow the project goes live, months later. But is anyone better off? If the measure of success was modernisation, of course they are. But was it worth all the disruption ... because the poor data, broken processes, 100+ work-arounds are actually still there. And so are the jaded users.

For a brilliantly written article on IT modernisation projects in the AI era, I highly recommend reading Renee Murphy's piece titled: My Mid-Market Modernization Manifesto for AI Transformation.

What is often overlooked in all the excitement and unfettered optimism that comes with an imperative to change is defining what success looks like.


IMHO I believe most ambitious IT projects would never actually progress to going live if someone defined success criteria upfront and keenly monitored it throughout the pre-go-live phases. Someone might have loudly whispered: "This baby just ain't gonna fly." 


So this is where we land with AI adoption.


One of the biggest mistakes companies make with AI is not defining what success looks like. They make the same modernisation-myth assumptions. 

"Let's just give everyone Microsoft Copilot access and see what happens."

So there are three ways that "What does success look like?" goes wrong:


  1. Not defining it, 

  2. Defining the wrong metrics or 

  3. Not defining enough metrics. 


Many AI projects are declared a success because they save time, reduce costs or generate impressive-looking outputs. But if the AI is inaccurate, biased, unsafe or ignored by staff, has it really succeeded? 


This is exactly what Statement 12 of the Australian Government's DTA AI Technical Standard addresses. Rather than focusing on whether AI is simply "working", Statement 12 requires companies to define what success actually looks like before the system goes live. 

The Standard requires companies to identify appropriate metrics, understand the trade-offs between them, and continually review whether those metrics remain appropriate throughout the AI system's lifecycle.


AI success means much more than measuring ChatGPT accuracy

With a common understanding that AI makes mistakes, many might assume the most important AI metric is accuracy. While accuracy is certainly important, it is rarely enough by itself. Imagine an AI system that drafts responses to customer enquiries with 98% grammatical accuracy. Sounds impressive. However consider these outcomes:


  •  Customers find the tone cold and impersonal.

  • Staff spend longer correcting the responses than writing them themselves.

  • The AI occasionally invents information.

  • People gradually stop trusting the system.


Technically the AI may be "accurate," yet the project has still failed. The goal is to evaluate AI from multiple perspectives, recognising that different measures often compete with each other.


What are appropriate success measures?

Luckily, the DTA Technical Standard provides a broad range of relevant metrics that may be appropriate depending on the AI use case. These include:


  • Business value and benefits realised

  • Productivity improvements

  • Model performance (such as precision or recall)

  • Data quality and diversity

  • Bias and fairness

  • Safety and harmful outputs

  • Reliability and uptime

  • Human-in-the-loop intervention rates

  • User adoption

  • User satisfaction

  • Drift in AI performance over time


The important point is that no single metric tells the whole story, and even a few are not enough. For example, an AI recruitment tool might achieve excellent prediction accuracy but unintentionally screen out certain demographic groups. Similarly, an AI chatbot may respond extremely quickly but provide misleading advice. Measuring the wrong things can create a false sense of confidence.


In 2019, Victoria Police procured an AI facial recognition system for use in law enforcement operations across metropolitan Melbourne. The system had been evaluated at 95% accuracy on the vendor’s test datasets. These datasets were not representative of Melbourne’s diverse population.


By 2021, operational accuracy on Melbourne’s population had dropped to 76%, with systematic failures on darker skin tones. The 19-percentage-point gap between vendor-reported accuracy and operational accuracy represented a fundamental procurement failure: the specification tested a snapshot of performance on non-representative data rather than requiring continuous performance across demographic groups.


AI trade-offs

Perhaps my most valuable advice is that (AI) design is full of compromises. Improving one metric often impacts another. For example:


  • Increasing sensitivity may generate more false positives.

  • Faster responses may reduce answer quality.

  • More complex models may improve prediction accuracy but become difficult to explain and cost more.

  • More automation may reduce human oversight.


Best practice is to identify these trade-offs early and consciously decide which outcomes matter most for each use case. Any decisions should be documented so stakeholders understand why certain design choices were made.


A moveable feast

An AI system that performs well today may drift or degrade over time as:


  • customer behaviour changes

  • legislation changes

  • organisational policies change

  • new products are introduced

  • data quality declines and

  • user expectations evolve


This variation is called model drift or data drift, and is one of the biggest reasons AI systems gradually become unreliable. Regularly reassessing the original success metrics is essential to ensure they reflect current day real-world performance.


Five Practical Tips for Measuring Success


1. Define success early, in the Design stage

Before anyone starts purchasing AI or connecting to an LLM, ask:


  • What business problem are we solving here?

  • How will we know we've solved it?

  • What would failure look like?


Avoid vague objectives such as "improve efficiency." Define each measure with a number in it:


  • Reduce average response time by 40%

  • Maintain customer satisfaction above 90%


2. Measure lots of things

Include measures across multiple categories:


  • Business outcomes

  • User experience

  • Safety

  • Fairness

  • Compliance

  • Operational reliability

  • Human oversight


This provides a much more balanced picture of what's going on.


3. Understand the trade-offs

For each AI system, discuss questions and document choices such as:


  • Is explainability more important than accuracy?

  • Is speed more important than completeness?

  • When must humans intervene?


4. Create an AI scorecard

Develop a simple (automated) dashboard for each AI use case that reports metrics such as:


  • Adoption rates

  • User satisfaction levels

  • Accuracy and error rates

  • Human intervention rates

  • Safety incidents

  • Privacy incidents

  • Productivity gains


Review these results regularly through your leadership meetings.


5. Monitor success metrics continuously (and keep monitoring!)

Success should never be measured only at go live. Schedule periodic reviews to examine:


  • declining accuracy, bias, explainability, auditability and so forth

  • changing user behaviour

  • new risks that may be emerging

  • unexpected outcomes or unintended consequences

  • emerging regulatory requirements


Nurture your AI as a living system that requires ongoing oversight rather than a fixed, set-and-forget system.


Bringing it all together into AI governance

By defining meaningful success criteria upfront, measuring multiple dimensions of performance, documenting trade-offs and continuously monitoring outcomes, companies may discover that it's better to abandon a project rather than live through a failed implementation. Ouch. That's a success in my book. Failed projects are an expensive waste of money, talent and time. Often, they are embarrassing. "You did what?"


If a Board Member ever asks "how's the AI going?". 


You can either say: "It's going well thanks".


Or, you can confidently answer something like: "Based on 12 success metrics, ten are tracking OK within acceptable parameters and two are being closely monitored, accuracy is trending downwards and one specific area of human oversight needs improving." 

What the Board wants to hear is actually: "The AI delivering the outcomes we intended, without creating unacceptable risks."

Moving from assumptions to measurable evidence is what separates mature AI governance from the next AI mistake.


Stay safe,


Bruce


AI. Use responsibly.



ABOUT ME

I partner with mid-size companies to confidently adopt AI, prevent high-profile failures and avoid the expensive mistake.


I write all my own content, you can tell by the odd typo and occasional missing word. I use AI for my research.


To learn about my upcoming public AI Governance workshops visit: Public workshops


To learn more about AI Governance, check out my Hitchhikers Guide to AI Governance Podcast.



Bruce Mullan hosts Hitchhikers Guide to AI Governance podcast
Bruce Mullan hosts Hitchhikers Guide to AI Governance


 
 
 

Comments


CONTACT

If you have a question or request  please contact us today!

© 2026 BY TRIPLE P GLOBAL PTY LTD T/AS Ai Governance Partners -

ABN 96 119 485 791

Thanks for contacting us. we'll be in touch.

bottom of page