This deep dive expands on the Loop Asia conversation with Karin Verspoor, Executive Dean of Computer Science at RMIT University.
Ask most executives how they manage the risk of a generative AI tool in production and you'll get the same answer: a human's in the loop. A clinician reads the AI scribe's summary and ticks a box. It sounds like governance. Karin Verspoor's warning is that it usually isn't — and the reason should worry any leader who's told a board "human review" closes the risk.
The failure isn't in the ninth review. It's in the tenth.
Verspoor's example: a clinician checks an AI summary, finds it accurate, nine times running. On the tenth, the temptation is to tick the box without reading closely — it's been right every time, why would this be different? That's the default behavior of anyone asked to catch a rare error inside a repetitive task, and why "a human reviews it" is weaker than it sounds: designed to catch the machine's mistakes, in practice it degrades into trusting its track record.
Quality built in, not inspected in
There's an old quality-management principle: you can't reliably inspect quality into a process after the fact. Inspection catches fewer errors the more consistently a process has performed — backwards from what a risk-conscious executive wants, since scrutiny should rise with consequences, not fall as trust builds. Human attention doesn't work that way; it relaxes as confidence grows.
What Verspoor sees developing instead is tooling that flags when something looks anomalous, rather than asking a human to catch every error cold — shifting the job from "read everything closely" to "pay attention when the system flags it." That's the difference between a control that degrades under repetition and one that doesn't.
The other half of her answer is the AI itself: not every part of a workflow needs to be generative. Information extraction, normalization, narrow classification models — older, more predictable, prediction rather than generation — mean the human reviewer isn't the only thing standing between a plausible output and a wrong one.
What a board actually needs to hear
For a business leader carrying an AI mandate — CEO, COO, head of professional services — "a human reviews it" is the natural first answer. The better answer describes the tenth review, not the first: what triggers extra scrutiny, which parts of the pipeline are deliberately non-generative, and what evidence shows the review step still works once the novelty wears off.
That's a harder story than a checkbox, and most organizations haven't written it — AI rollouts get designed around making the model work, not around what happens once everyone stops paying attention. It's not a framework a business-side leader should design from scratch on top of the rest of the mandate. Getting it right before the tenth mistake proves the gap, rather than after, is the difference between an advisory conversation and an incident report.
Telling your board "a human reviews it" and hoping that closes the risk?
The review step that actually holds up under repetition looks different from the one that gets rubber-stamped by month three — worth getting right before the tenth mistake proves the gap.