|
Got this forwarded? Subscribe here →
AI in Finance FROM PRACTICE, NOT THEORY Issue №07 - 26 August 2026
What Regulators Actually Think About AI (It's Not What You Expect)
. THREE PRACTITIONER INSIGHTS .
01
I sat across from our C-Management. They didn't want us to stop. They wanted us to show our work.
I'll share this without specifics, but the message was unmistakable. In our last C-level management on AI, the tone wasn't adversarial. It was curious. The management's concern wasn't "are you using AI?" — it was "can you demonstrate that you're using it responsibly?" They wanted to see clear accountability chains, documented decision processes, monitoring evidence, and a credible explanation of what happens when the model is wrong. Banks that approach regulators defensively, hiding deployments, minimizing capabilities are creating the adversarial relationship they fear. Banks that approach proactively, here's what we're doing, here's how we govern it, here's what we're uncertain about — find an ally, not an auditor.
02
Your model risk framework was built for a world that no longer exists.
Most banks' MRM frameworks assume models are deterministic (same input → same output), retrained infrequently, stable between retraining cycles, and numerical in their outputs. GenAI breaks every one of these. Non-deterministic outputs. Frequent model updates by the provider. Behavioral drift without any change on your end. And outputs that are text, not numbers — which means your monitoring needs to catch semantic drift, not just statistical drift. If your MRM framework doesn't have sections on prompt versioning, output quality monitoring, hallucination detection, and the distinction between models that decide and models that assist, it's incomplete. Not wrong — just incomplete for the current reality.
03
The supervisory question that caught us flat-footed wasn't about accuracy. It was about disagreement.
I went into a supervisory dialogue expecting to defend our model performance. I had the accuracy metrics polished, the validation results, the benchmark comparisons. The question that actually mattered never touched any of it. What they asked, roughly: "Show me a case where your human reviewer disagreed with the model, and walk me through what happened next." I had a good answer in principle and a thin answer in evidence. We could produce the model's accuracy to two decimal places. We could not, on the spot, produce a clean record of human override events, when a reviewer overrode the model, why, and whether that disagreement fed back into anything. Our override rate on one assisted-decision tool was around 4%, and when they asked me to explain the pattern in those overrides, I realised I didn't actually know it, we'd been logging that the override happened, not why. That's the gap. Supervisors are not auditing whether your model is good. They are auditing whether your humans are genuinely in control of it, and the evidence for that lives in the disagreements, not the agreements. A 98% agreement rate, which I used to quote with pride, reads very differently to a supervisor, it can mean strong alignment, or it can mean nobody is really looking. The only thing that distinguishes the two is whether you can explain the 2%. We now log every override with a structured reason code. I wish we'd started two years ago. If a supervisor asked you tomorrow to characterise your override patterns, could you? I couldn't, and I run this for a living.
|
- Two Use Cases -
→ WIN We briefed our C-Level Management before we launched. It was the best decision of the year.
Before deploying a GenAI-powered assistant into a client-adjacent workflow, we did something my risk team initially resisted: we proactively briefed our C-Level Management. We presented use case scope, human-in-the-loop design, governance framework, monitoring plan, and a live demo. No surprises. No spin. They gave informal feedback that led to two design improvements we hadn't considered. When we launched, there were zero supervisory concerns, because there were zero surprises. Proactive engagement took courage and two weeks of preparation. It bought us credibility that has made every subsequent conversation smoother. If you're hiding your AI deployments from your supervisor, ask yourself what you're protecting. Because it's probably not the bank.
→ LESSON "The human is in the loop" — but the human was rubber-stamping
A credit risk team deployed a model that recommended approval or rejection of SME loan applications. The "human-in-the-loop" was a credit analyst who reviewed each recommendation and could override. Sounds solid. Except: the analyst saw a green or red recommendation with a confidence score. No explanation of which factors drove the decision. No indication of which inputs would change the outcome. In practice, the override rate was 2% — the analyst agreed with the model 98% of the time. When the supervisor examined this, they didn't call it "human oversight." They called it "automation bias with an extra click." Human-in-the-loop is only meaningful if the human has enough information to genuinely disagree. A green light with no explanation isn't oversight. It's a rubber stamp with a salary.
|
One myth I'd retire
"If the human makes the final decision, the model doesn't need to be explainable."
This is the most dangerous myth in AI governance right now. Regulators are increasingly clear: human oversight means meaningful human oversight. That requires three things: (1) the human can see what the model considered, (2) the human can understand why the model reached its conclusion, and (3) the human can realistically override based on their own judgment. If any of these are missing, you don't have human oversight. You have a human witness. Those are very different things, and regulators know the difference even if your internal governance documentation doesn't yet distinguish them.
◉ THE REGULATORY SIGNAL
[Written July 23th.] If you want to know what supervisors will probe before they probe it, read what they write for each other. The Basel Committee on Banking Supervision published a paper last autumn on how supervisors can address explainability in banks' use of AI — and the reason it matters more than another vendor whitepaper is that it is supervisors telling supervisors what "good" looks like, which is a preview of the questions coming to your model risk team. The direction of travel is consistent with what I described above: the scrutiny is shifting from model performance to the demonstrability of human control and the traceability of decisions. The EBA, for its part, has signalled it sees no immediate need for new AI-specific guidelines — instead it is working through 2026–2027 to build a common supervisory approach across national authorities, which in plain terms means the questions your supervisor asks are converging with the questions every other European supervisor asks. That convergence is good news: it means you can prepare one evidence pack, not twenty-seven. What to do this quarter: assemble an explainability evidence file for each high-risk and client-adjacent system, and build it to answer the supervisor's real question, not the textbook one. Not "how accurate is the model" but "show me a decision, show me what the human saw, show me a case where the human disagreed, and show me where that disagreement went." If you can produce that file in an afternoon, you are ahead of most of the European market. If assembling it would take three weeks, that is your September project.
|
🎁 FREE THIS ISSUE: Supervisory Briefing Template
Proactive supervisory engagement was the best risk decision we made this year. I turned our approach into a reusable template: a 2-page document with a structured briefing agenda (use case scope, governance framework, monitoring plan, human oversight design), the 10 questions our supervisor asked that we should have anticipated, and a preparation checklist for the week before the meeting. After this issue, you'll notice the questions are less about your model and more about your humans — the template reflects that.
.
|
Next issue is about EU AI Act practical implementation, "What I'm Doing About It." With the Digital Omnibus shifting Annex III deadlines to December 2027, this issue becomes even more important: the extra time is only useful if you use it, and most banks won't. I'll share the exact classification work we're doing right now, and the one mistake I see every bank making with Annex III mapping. September 09th.
If this was useful, forward it to one finance leader who'd want it. That's how this newsletter grows.
|
Unsubscribe · Preferences
|