6 Critical AI Decision Accountability Practices for Regulated Markets
Your commercial team is already using AI decisions your board will eventually ask about. In most organizations, nobody has written down who is accountable when one of those decisions turns out to be wrong.
This is not a hypothetical exposure. It is a documentation gap that becomes visible at exactly the moment it is most expensive: when a payer disputes a claim, when a regulator asks how a commercial position was arrived at, or when a board asks why a launch was funded on the basis it was.
A recommendation is not accountability.
Most executives treat an AI output as a tool result, in the same category as a spreadsheet model or a purchased market report. In a regulated market, it is not the same category at all.
The moment a recommendation informs a commercial decision that a payer, a regulator, or a plaintiff can later examine, it becomes part of a decision record. And in most companies, that record has a blank signature line, because nobody treated the moment as a decision point requiring a name.
The gap is not a policy gap. Most healthcare organizations now have an AI policy. It sits in the handbook, it describes acceptable use, it prohibits entering patient data into consumer tools, and it has rarely been the reason a commercial gate was held.
The gap is that nobody has defined what grade of evidence a decision requires before it is permitted to proceed, and nobody has tied that grade to how difficult the decision is to reverse.
Reversibility sets the evidence bar, not importance.
The most useful reframe available here is that not all AI decisions require the same evidence, and the variable that should determine the requirement is not how important the decision feels. It is how hard it is to undo.
The distinction has recognizable lineage.It is the same logic as the one-way and two-way door framing that has circulated in operating literature for years, and it is closely related to the premortem technique developed by Gary Klein, which uses prospective hindsight to surface the reasons a plan will fail before it has been committed to. Naming the lineage rather than implying novelty is itself a small evidence-standards decision.
Reversible decisions
A strong AI decision that can be unwound within a quarter at contained cost. A pilot in one territory. A message test. An agency engaged on a three-month term. A fractional hire. If the assumption underneath turns out to be wrong, you learn quickly, the cost is bounded, and the learning is usually worth more than the loss.
Committed decisions
A decision requiring eighteen months and organizational credibility to reverse. A country entry. A restructured commercial organization. A pricing architecture. A senior hire with equity attached. A launch sequence. Reversing any of these requires unwinding everything that has been built on top of it in the interim, and the cost compounds with each quarter of delay.Most organizations apply the same evidence standard to both, which means they systematically over-evidence the reversible decisions and under-evidence the committed ones. The second error is the expensive one, and it is rarely visible until the reversal is required, at which point it is framed as a market change rather than a failure.
The Evidence Grade scale
Three tiers on one page. The point is not sophistication. The point is that the grade is named before the AI decision is discussed, so the conversation becomes one about evidence rather than one about confidence, and those two things are routinely confused in senior rooms.
Grade A: decision-grade
Primary, sourced, and adversarially tested. Someone has actively gone looking for the evidence that would contradict the conclusion and has documented what they found, including the case where they found nothing.
Grade A authorizes irreversible and high-commitment decisions. It is expensive to produce, and that expense is appropriate, because the choices it authorizes are expensive to reverse. A company that finds Grade A evidence too costly to produce has usually not compared it to the cost of an unwind.
Grade B: directional
Credible, incomplete, typically resting on a single source or a single method, and not adversarially tested. Grade B authorizes reversible AI decisions only.
It is the correct and sufficient grade for most commercial experimentation, and it becomes dangerous only when it is used to authorize commitment. The most common failure pattern in commercial organizations is a Grade B insight producing a committed decision because the insight was compelling and the room was aligned.
Grade C: indicative
Anecdote, vendor claim, unvalidated model output, or a single customer conversation. Grade C authorizes nothing. Its function is to signal that an inquiry is worth opening.
Treating Grade C evidence as though it were Grade B is the most common failure in commercial decision-making generally, and AI outputs are frequently Grade C delivered with Grade A confidence. That mismatch between the epistemic status of the content and the tone of its delivery is the specific new risk that generative systems introduce.
Why “the model recommended it” fails an audit
The problem with an AI recommendation is not accuracy. Models are frequently right, and in several commercial domains they are more consistently right than the humans they assist.
The problem is that a recommendation carries no provenance a third party can examine. When a choice is challenged twelve or twenty-four months later, the questions are consistent, and they are procedural rather than technical.
• What did you know at the time
• What did you consider and reject
• Who decided
• What would have changed the decision?A model output answers none of these. It cannot state what it discounted. It cannot state what it did not have access to. It cannot be deposed. A named human who recorded five answers before the decision was made can answer all four in ten minutes.
This is not an argument against using AI in commercial decisions. Used well, AI does more of the work below the decision faster, so the human holding the call arrives with better information and more time. Used badly, the AI effectively holds the call, because it presents with confidence and the human is overworked and under-resourced. The difference between those two states is entirely procedural.
The five questions and the one-page Decision Record
Before any commercial gate opens, five answers are written down. Not discussed. Written.
• What exactly is the call, stated in one sentence a person outside the room would understand
• Who owns it, by name and not by function
• What grade of evidence does a decision this hard to reverse require
• What grade did it actually receive
• What would have to become true, or stop being true, for us to revisit this?The fifth question is the one almost always skipped and the one that does the most work. A decision with no named revisit condition cannot be reversed on evidence. It can only be abandoned on politics, usually two quarters late, usually by someone who was not in the original room and therefore has no institutional cost in abandoning it.
The record fits on one page. It adds roughly twenty minutes to a gate review. That is the entire cost, and the objection that it introduces bureaucracy does not survive contact with the actual artifact, which is shorter than the agenda of the meeting it sits inside.
Where AI is safe inside the gate, and where it is not
The useful boundary is not by function but by reversibility, applied one level down.
AI is well suited to the work below the decision. Assembling the evidence base. Summarizing what is known. Generating the disconfirming case, which it does unusually well and which humans find socially difficult. Drafting the first version of the record itself. Stress-testing an assumption by arguing against it.
AI is poorly suited to holding the call where the cost of error is asymmetric, and the synthesis requires knowing what the room is not saying, what the regulator will tolerate, and what the organization can actually execute. Those are not information problems, and more information does not solve them.
A practical rule: AI may produce any input to the record. It may not be the answer to question two.
Where to start, and what not to do first
Do not begin by writing a policy. Policies are the artifact organizations produce when they want to feel governed without changing an AI decision.
Begin by taking your last three commercial decisions and grading them retroactively against the scale. Ask what grade each one required given its reversibility, and what grade it actually received.
The exercise takes under an hour, and it is uncomfortable, which is the signal that it is working. In most organizations, at least one committed decision turns out to have been made on Grade B evidence. Identifying which one, before it needs unwinding, is worth considerably more than any policy document you could write this quarter.
Frequently Asked Questions
Who is accountable when an AI recommendation is wrong?
In a regulated market, accountability rests with the named human who authorized the decision, not with the system that produced the recommendation. Most organizations have not named that person in advance, which means accountability is assigned retroactively and usually attaches to whoever executed rather than whoever decided.
What is an evidence grade?
An evidence grade classifies how strong the evidence behind a decision is and what class of decision that evidence is permitted to authorize. A three-tier scale distinguishes decision-grade evidence, which is primary and adversarially tested, from directional evidence, which is credible but incomplete, and indicative evidence, which authorizes nothing.
Is an AI policy enough to govern AI decisions?
No. An AI policy describes acceptable use. It does not typically define what evidence a decision requires or who is accountable for a specific call. Most organizations with a policy still cannot name the human responsible for their most recent AI-assisted commercial decision.
How does reversibility affect the evidence required for a decision?
Reversibility rather than importance should set the evidence bar. These are AI decisions that can be unwound quickly at contained cost and can proceed on directional evidence. Decisions requiring eighteen months and organizational credibility to reverse require decision-grade evidence and, in practice, rarely receive it.
What is a decision record?
A one-page document capturing five things before an AI decision is made: the call, the named owner, the evidence grade required, the evidence grade received, and the condition that would trigger a revisit. It allows a result to be traced back to the decision and the evidence that produced it.
Why does “the model recommended it” fail an audit?
Because a model output carries no examinable provenance. It cannot state what it discounted, what it lacked access to, or what would have changed its conclusion. A named human who recorded five answers before deciding can answer all of those questions.
Where is AI safe to use inside a commercial decision process?
AI is well suited to the work below the decision: assembling evidence, summarizing what is known, generating the disconfirming case, and drafting the record. It is poorly suited to holding a call where the cost of error is asymmetric, and the synthesis depends on organizational and regulatory judgment.
What is a revisit trigger?
A condition named in advance that, if met, requires the AI decision to be reopened. Without one, a decision cannot be reversed on evidence and can only be abandoned on politics, usually late and usually by someone who was not in the original room.
EXTERNAL CITATIONS
• Gary Klein, Performing a Project Premortem, Harvard Business Review, September 2007