Written for the CIO / CDIO: board-facing, personally accountable for clinical safety, cyber and responsible AI, technically literate but not hands-on. DTAC, DCB 0129/0160, DSPT and DPIA are used without definition. The piece deliberately contains zero integration acronyms, reflecting the language this audience actually uses; the editorial stance is that rigorous scrutiny is leadership, not obstruction, which positions the reader (and WeHub) on the same side of the table. The three questions are designed to be lifted directly into an evaluation pack, giving the article a practical afterlife beyond the read. Pairs naturally with a Medi-voice LinkedIn take rather than a brand-page broadcast.
Somewhere in your inbox right now is an AI pitch. Probably several. Ambient scribes, triage models, discharge summarisation, flow prediction: nearly every supplier deck that crossed a digital leader's desk this year has AI on it somewhere, and most of them lead with it.
The demos are often genuinely impressive. That's the problem. A polished demo compresses the distance between "this works in a controlled walkthrough" and "this is safe on a Tuesday night ward round", and the person who carries the gap between those two statements isn't the vendor. It's you.
NHS AI governance has plenty of formal machinery: DTAC, clinical safety cases under DCB 0129 and DCB 0160, information governance sign-off, MHRA classification where a product crosses into medical device territory. All necessary. None of it answers the question that actually decides whether an AI product belongs in your estate, because conformance tells you a vendor completed the paperwork. It doesn't tell you how they think.
Three questions do. They're short, they're fair, and the way a vendor answers them tells you more than the rest of the procurement pack combined.
Why NHS AI governance starts in the vendor meeting
Traditional due diligence was built for deterministic software. A rules engine either fires or it doesn't; you test it, you document it, you move on. AI is probabilistic. It will be right most of the time, and "most of the time" is not a technical specification. It's a clinical governance position that someone in your organisation has to own.
That changes what the vendor conversation is for. You're no longer verifying features. You're auditing how a supplier reasons about failure, control and evidence, because those habits, far more than the model architecture, determine what your relationship with them looks like in year two.
The questions below aren't hostile. Good vendors enjoy them. The ones who reroute to the demo are telling you something too.
Question one: what happens when it's wrong, and who's accountable?
Every AI product will be wrong. The only unknowns are how often, in which direction, and whether anyone notices in time.
So ask it plainly, and listen for structure in the answer. A serious vendor can describe their known failure modes without flinching: where the model degrades, which cohorts it performed weakest on in evaluation, what happens at the edges of its training data. They can tell you how errors surface (does the product flag its own uncertainty, or fail silently?), how incidents get reported, and how their hazard log under DCB 0129 maps to the DCB 0160 case your clinical safety officer will have to own.
Then push on the part most pitches skip: change. Models get retrained and updated. Ask how behaviour changes are versioned, communicated and re-assured, because an update that shifts outputs without a change process is a new deployment wearing the old one's safety case.
The answer to walk away from is "it's only decision support, a clinician always reviews it". Sometimes that's true. But human-in-the-loop is an accountability design, not a disclaimer, and if the workflow makes reviewing every output impractical, the loop is decorative. If a vendor's answer to "who's accountable" quietly resolves to "your clinicians are", you've learned what you needed to.
Question two: can we run it on our terms?
This is the control question, and it splits into data and deployment.
Data first: what leaves your boundary, where does inference happen, what's retained, and is your data used to train anyone else's model? You'll need those answers for the DSPT return and the DPIA anyway. Better to hear them in plain English now than excavate them from a contract schedule later.
Then deployment. Cloud-only is a design decision on the vendor's part, and some NHS estates can't accept it, whether for connectivity, risk posture or policy. Ask whether the product can run hybrid or fully inside your own environment, and what capability you'd trade if it did. Deployment control is a hill we've deliberately chosen at WeHub: if a Trust's security posture says a workload runs inside its own walls, the platform should follow the posture, not argue with it.
Next, the supply chain. An AI product is rarely one company. It's a stack of model providers, hosting layers and third-party services, and every one of them is now part of your attack surface. Ask who the second, third and fourth parties are. A vendor who can't name their own dependencies can't help you defend them.
And ask about the exit. Contracts end. What comes back to you, in what format, and what breaks on the way out?
Question three: where's the evidence it saves time without adding risk?
"Hype to value" has attached itself to every NHS AI adoption conversation for a reason: the gap between claimed benefit and realised benefit is where most of the disappointment lives.
So ask for evidence, and be specific about what counts. An evaluation in a setting that resembles yours. A baseline measured before deployment, not reconstructed after it. Outcomes defined in advance rather than harvested from whatever happened to move. Independence matters too: a vendor-run pilot with a happy quote from one enthusiastic site is a testimonial, not evidence.
Ask what didn't work. Every honest evaluation has negative findings, and a vendor with none is either very new or very selective. Ask about time in full, not in part: minutes saved on documentation mean little if they reappear as review burden, correction work or a longer sign-off queue elsewhere in the pathway.
One more filter worth adding, borrowed from how thoughtful digital leaders increasingly frame it: does the product make care fairer, or just cheaper? A model that performs beautifully on the average patient and quietly worse on the cohorts your Trust actually serves isn't a productivity gain. It's a risk transfer.
What good answers sound like
Across all three questions, the pattern you're listening for is the same.
Good vendors answer in plain English before they answer in acronyms. They name accountability instead of diffusing it. They treat your governance requirements as design inputs rather than obstacles, and they volunteer limitations before you find them. When they don't know, they say so and come back with the answer, instead of improvising one in the room.
Weak vendors do the opposite. They answer failure questions with accuracy percentages, control questions with compliance badges, and evidence questions with another demo. None of that is disqualifying on its own. As a pattern, it is.
The decision you're actually making
You aren't really evaluating a product in these meetings. You're choosing what your organisation will be accountable for after the contract is signed, when the deployment is live and the vendor's attention has moved to the next Trust.
The manageable next step: put these three questions, in writing, into your standard AI evaluation pack this week, and make plain-English answers a condition of progressing. It costs nothing, it filters fast, and it tells every supplier who walks through your door what kind of buyer they're dealing with: one who wants the value and intends to govern for it. That, more than any framework, is where NHS AI governance actually starts.



