AI Governance Statement 12: Define success criteria

If you don't know where you're landing the plane, you'll probably end up somewhere else.
The groundhog day conversation I've had over a dozen times in the last 15 years after a newly launched system failed within 12 months. During the post-mortem coffee conversation:
IT: We desperately needed a new system
ME: What were you seeing to indicate that?
IT: Staff were constantly complaining and the vendor support was hopeless
[ME thinking: Poor data, broken processes, 100+ work-arounds]
IT: We had to modernise. Get the latest systems. We had no choice.
ME: Oh, I see. So you reverted back to the old system when it failed?
IT: Yes.
A new system is always good story to tell. It tells the Board things are moving forward, something positive, modern systems are coming. IT people get experience in contemporary technologies. The latest tech, cloud computing, a contemporary user interface, a vendor with snazzy support tools. What could possibly go wrong?
But once you start implementing, you find where the bodies are buried. Someone believed a new system would fix:
Years of unfettered document and data accumulation,
Poor business processes,
Integration with other ageing systems and
Operating knowledge inside people's heads
Then, the worst part happens. You hit these four big challenges with data, people, processes and integrations and so on, but you've gone too far. You've passed the point of no return. You now have to make it "work".
Been there, done that. Somehow the project goes live, months later. But is anyone better off? If the measure of success was modernisation, of course they are. But was it worth all the disruption ... because the poor data, broken processes, 100+ work-arounds are actually still there. And so are the jaded users.
For a brilliantly written article on IT modernisation projects in the AI era, I highly recommend reading Renee Murphy's piece titled: My Mid-Market Modernization Manifesto for AI Transformation.
What is often overlooked in all the excitement and unfettered optimism that comes with an imperative to change is defining what success looks like.
IMHO I believe most ambitious IT projects would never actually progress to going live if someone defined success criteria upfront and keenly monitored it throughout the pre-go-live phases. Someone might have loudly whispered: "This baby just ain't gonna fly."
So this is where we land with AI adoption.
One of the biggest mistakes companies make with AI is not defining what success looks like. They make the same modernisation-myth assumptions.
"Let's just give everyone Microsoft Copilot access and see what happens."
So there are three ways that "What does success look like?" goes wrong:
Not defining it,
Defining the wrong metrics or
Not defining enough metrics.
Many AI projects are declared a success because they save time, reduce costs or generate impressive-looking outputs. But if the AI is inaccurate, biased, unsafe or ignored by staff, has it really succeeded?
This is exactly what Statement 12 of the Australian Government's DTA AI Technical Standard addresses. Rather than focusing on whether AI is simply "working", Statement 12 requires companies to define what success actually looks like before the system goes live.
The Standard requires companies to identify appropriate metrics, understand the trade-offs between them, and continually review whether those metrics remain appropriate throughout the AI system's lifecycle.
AI success means much more than measuring ChatGPT accuracy
With a common understanding that AI makes mistakes, many might assume the most important AI metric is accuracy. While accuracy is certainly important, it is rarely enough by itself. Imagine an AI system that drafts responses to customer enquiries with 98% grammatical accuracy. Sounds impressive. However consider these outcomes:
Customers find the tone cold and impersonal.
Staff spend longer correcting the responses than writing them themselves.
The AI occasionally invents information.
People gradually stop trusting the system.
Technically the AI may be "accurate," yet the project has still failed. The goal is to evaluate AI from multiple perspectives, recognising that different measures often compete with each other.
What are appropriate success measures?
Luckily, the DTA Technical Standard provides a broad range of relevant metrics that may be appropriate depending on the AI use case. These include:
Business value and benefits realised
Productivity improvements
Model performance (such as precision or recall)
Data quality and diversity
Bias and fairness
Safety and harmful outputs
Reliability and uptime
Human-in-the-loop intervention rates
User adoption
User satisfaction
Drift in AI performance over time
The important point is that no single metric tells the whole story, and even a few are not enough. For example, an AI recruitment tool might achieve excellent prediction accuracy but unintentionally screen out certain demographic groups. Similarly, an AI chatbot may respond extremely quickly but provide misleading advice. Measuring the wrong things can create a false sense of confidence.
In 2019, Victoria Police procured an AI facial recognition system for use in law enforcement operations across metropolitan Melbourne. The system had been evaluated at 95% accuracy on the vendor’s test datasets. These datasets were not representative of Melbourne’s diverse population.
By 2021, operational accuracy on Melbourne’s population had dropped to 76%, with systematic failures on darker skin tones. The 19-percentage-point gap between vendor-reported accuracy and operational accuracy represented a fundamental procurement failure: the specification tested a snapshot of performance on non-representative data rather than requiring continuous performance across demographic groups.
AI trade-offs
Perhaps my most valuable advice is that (AI) design is full of compromises. Improving one metric often impacts another. For example:
Increasing sensitivity may generate more false positives.
Faster responses may reduce answer quality.
More complex models may improve prediction accuracy but become difficult to explain and cost more.
More automation may reduce human oversight.
Best practice is to identify these trade-offs early and consciously decide which outcomes matter most for each use case. Any decisions should be documented so stakeholders understand why certain design choices were made.
A moveable feast
An AI system that performs well today may drift or degrade over time as:
customer behaviour changes
legislation changes
organisational policies change
new products are introduced
data quality declines and
user expectations evolve
This variation is called model drift or data drift, and is one of the biggest reasons AI systems gradually become unreliable. Regularly reassessing the original success metrics is essential to ensure they reflect current day real-world performance.
Five Practical Tips for Measuring Success
1. Define success early, in the Design stage
Before anyone starts purchasing AI or connecting to an LLM, ask:
What business problem are we solving here?
How will we know we've solved it?
What would failure look like?
Avoid vague objectives such as "improve efficiency." Define each measure with a number in it:
Reduce average response time by 40%
Maintain customer satisfaction above 90%
2. Measure lots of things
Include measures across multiple categories:
Business outcomes
User experience
Safety
Fairness
Compliance
Operational reliability
Human oversight
This provides a much more balanced picture of what's going on.
3. Understand the trade-offs
For each AI system, discuss questions and document choices such as:
Is explainability more important than accuracy?
Is speed more important than completeness?
When must humans intervene?
4. Create an AI scorecard
Develop a simple (automated) dashboard for each AI use case that reports metrics such as:
Adoption rates
User satisfaction levels
Accuracy and error rates
Human intervention rates
Safety incidents
Privacy incidents
Productivity gains
Review these results regularly through your leadership meetings.
5. Monitor success metrics continuously (and keep monitoring!)
Success should never be measured only at go live. Schedule periodic reviews to examine:
declining accuracy, bias, explainability, auditability and so forth
changing user behaviour
new risks that may be emerging
unexpected outcomes or unintended consequences
emerging regulatory requirements
Nurture your AI as a living system that requires ongoing oversight rather than a fixed, set-and-forget system.
Bringing it all together into AI governance
By defining meaningful success criteria upfront, measuring multiple dimensions of performance, documenting trade-offs and continuously monitoring outcomes, companies may discover that it's better to abandon a project rather than live through a failed implementation. Ouch. That's a success in my book. Failed projects are an expensive waste of money, talent and time. Often, they are embarrassing. "You did what?"
If a Board Member ever asks "how's the AI going?".
You can either say: "It's going well thanks".
Or, you can confidently answer something like: "Based on 12 success metrics, ten are tracking OK within acceptable parameters and two are being closely monitored, accuracy is trending downwards and one specific area of human oversight needs improving."
What the Board wants to hear is actually: "The AI delivering the outcomes we intended, without creating unacceptable risks."
Moving from assumptions to measurable evidence is what separates mature AI governance from the next AI mistake.
Stay safe,
Bruce
AI. Use responsibly.
ABOUT ME
I partner with mid-size companies to confidently adopt AI, prevent high-profile failures and avoid the expensive mistake.
I write all my own content, you can tell by the odd typo and occasional missing word. I use AI for my research.
To learn about my upcoming public AI Governance workshops visit: Public workshops
To learn more about AI Governance, check out my Hitchhikers Guide to AI Governance Podcast.
To listen visit: Hitchhikers Guide to AI Governance Podcast





Comments