Fireside chat · Altimetrik Executive Exchange · London

What it really takes to run AI in a regulated business?

Fireside Chat
/
Naman Joshi,
Solution Engineering Leader, EMEA, OpenAI
Shoby Abdi,
Senior Growth Partner, Altimetrik
In financial services the debate about whether AI works is mostly settled. The harder question is what it takes to move past pilots and keep AI running in production inside a regulated institution.
This was recorded at the Altimetrik Executive Exchange in London, hosted with OpenAI for a room of senior enterprise leaders under Chatham House rules. Shoby Abdi leads growth for Altimetrik’s AI practice. Naman Joshi leads solution engineering for OpenAI across EMEA. Over about half an hour they get into the adoption paradox, the argument that regulators are no longer the thing holding banks back, and the crawl-walk-run rollout that took one European bank from a handful of scattered experiments to an enterprise-wide programme. The last stretch is audience questions, one on turning productivity into growth and one on getting consistent output from models that are not deterministic.
01 — The starting point

This is about production, not pilots

  • Naman — Good evening, everyone. I lead a team here in Europe and I’ve been with OpenAI for about six months, so it’s been quite a ride. Before OpenAI I spent around four years at Databricks on the data and AI platform side, including work in financial services with an Australian investment bank. I’m genuinely excited to be here.
  • Shoby —And I’m a growth leader within our AI practice at Altimetrik. (After a brief hunt for the right microphone, “one mic to rule them all.”) This, by the way, is how we work with our partners.
02 — The adoption paradox

Everyone has AI. Almost no one has value, wall to wall.

  • Shoby — Quick show of hands. Whose organisation has an AI capability? Who’s actively using AI? Now, who feels it has really gone wall to wall, across the whole organisation, with concrete, measurable value? Far fewer hands. We call this the adoption paradox.
  • Shoby — When I talk to organisations about AI, the conversation often isn’t a strategic discussion about process transformation. It becomes something closer to a therapy session about how AI won’t break the organisation. And the number-one driver of that paradox is fear.
Too often the AI conversation isn’t a strategy session. It’s a therapy session about how AI won’t break the organisation.
Shoby Abdi · Altimetrik
03 — The regulators

You can’t blame the regulators anymore

  • Shoby — Think back to 2024 and 2025. Who was the number-one fear? Usually the regulator. What will they think, what will they do, what will they mandate, especially in Europe, with the ECB and the EU AI Act. But the regulators came out quickly and clearly and said, in effect: “We’re not the problem. Here are defined ways to work with AI inside a regulated institution.”
  • Shoby — Everyone expected the dam to burst on AI across banks. For the most part it hasn’t. Adoption is still a little stifled. So you can’t blame the regulators anymore, and that’s the core theme tonight. They’ve done a solid job defining what safeguarded AI looks like.
  • Shoby — There’s real fragmentation: the EU AI Act, guidance out of Westminster, even state-level rules in the US like New York and Illinois. But a lot of it is advice rather than hard regulation. The ECB and the AI Act have tried to consolidate that into specific, usable guidance.
You really cannot blame the regulators anymore. They’ve done a solid job of defining what safeguarded AI looks like.
Shoby Abdi · Altimetrik
04 — Pilot purgatory

If you focus on the pilot, all you’ll ever build is a pilot

  • Shoby — Who here has built an AI pilot? And who built it hoping it would become production-ready? That’s the trap: focus on the pilot and all you’ll ever do is pilot. The good news is that making something production-ready in 2026 isn’t as hard as it used to be. The technology is there; partners like OpenAI are an accelerant. Strategy isn’t a deck. It’s built into the culture of organisations that actually put things into production.
  • Naman — What’s striking over the last six months is the pace. At OpenAI we talk about a “capability overhang”: the gap between how fast the models are advancing and how fast enterprises adopt them keeps widening. A lot of the go-to-market job is closing that gap, and it’s rarely the technology that’s the blocker. It’s internal culture and executive vision.
  • Naman — A simple example: a German automotive manufacturer that genuinely wants to transform. Just to get started, they spent about four months on NDAs, legals, compliance and architecture, and by the end the pilot we’d scoped was already out of date. The space had moved. The real question is whether you have the conviction, at board level, to treat this as transformation, not just another IT use case.
The gap between what the models can do and what enterprises have adopted keeps getting bigger. That’s the capability overhang.
Naman Joshi · OpenAI
05 — Change management

Find your champion

  • Shoby — So who is that champion in your organisation? If it isn’t you, do you know who it is? One takeaway from tonight: find that person and work out how to support them.
  • Shoby — I loved talking to banks in 2025. You spent real money on AI. But a lot of that year was build-something, shelve-it, build-something, shelve-it; some of it never got past capability because teams assumed the technology wasn’t ready. In 2026 it is. Take Codex: many people see it as a software-development tool. It’s far more than that. It’s a process-transformation and workflow-automation capability. The technology is more advanced than most people assume, so it’s worth revisiting. The remaining gap is organisational: leadership alignment and champions.
Who is the champion in your organisation? If it isn’t you, do you at least know who it is?
Shoby Abdi · Altimetrik
06 — The framework

Crawl, walk, run, then accelerate

  • Naman — There’s a crawl-walk-run pattern that works well. Take a major Spanish bank we work with. They started by crawling: enabling small teams across finance, HR, legal and the retail bank to just get going. Your people already want these tools. One internal stat we saw suggested a large share of 18-to-35-year-olds use them several times a day in their personal lives. They come to work and ask, “Why don’t I have this here?”
  • Naman — Small fires started. The retail team said, “We get about 4,000 requests a year to customise home loans, can we build a custom GPT?” A process that took about three weeks dropped to a few hours, and a junior analyst could demonstrate that value and share it. That bubbled up to the chairman and the C-suite, who recognised a transformational platform.
  • Naman — They launched a programme, internally nicknamed the “Robots Programme,” with a deliberately simple, six-use-case shape: two customer-facing (service/IVR and the mobile app), two internal-workflow (risk-management approvals and banker productivity) and two employee-productivity. Critically, it wasn’t an IT delivery. They built a roughly 200-person team drawn from the lines of business, with ownership on the Chief Risk Officer, Chief Customer Officer and Chief Strategy Officer. Anyone in the organisation could tell you what the Robots Programme was, and that simplicity is what made it land.
The signature pattern

The path from pilot to production

Stage 01

Crawl

Enable small teams across functions and let people use the tools they already know.
  • Finance, HR, legal, retail: just get going
  • Custom GPTs and quick wins
  • Weeks of work become hours; an analyst can demo it

Stage 02

Walk

Turn small fires into a simple, nameable programme owned by the business.
  • A tight set of use cases everyone can name
  • Ownership on the business lines, not IT
  • Sponsored at C-level; guardrails defined as you go

Stage 03

Run

Accelerate to enterprise-scale, production-grade workflows.
  • Agents chained into real processes
  • Across customer, internal and employee journeys
  • Kept safe, repeatable and auditable
Naman’s crawl-walk-run pattern, generalised from a European bank’s rollout. The stages depend on each other, so you don’t skip from crawl to run.
07 — Sandbox to scale

A sandbox isn’t a playground. It’s the path to scale

  • Shoby — Is anyone doing sandboxing with their AI work today? A sandbox lets business users see AI in action even when the data or processes aren’t fully ready, so it removes the fear of the unknown. But the point of a sandbox isn’t experimentation for its own sake; it’s scale. You sandbox so you can govern properly, and governance doesn’t have to live in a 55-page document; it can be expressed as semantic-gateway access and guardian agents. Once you’ve governed and policed the sandbox, you can scale out of it. Even the ECB frames it this way: it’s not about control, it’s “show us, in a safe space, how you’ll use these tools.”
  • Shoby — On that safe space, look at what’s happening with Codex. Internally at OpenAI we moved from ChatGPT to Codex remarkably fast (and yes, the name is rough, we’re rebranding it). It’s not a coding tool anymore; it’s workflow automation. Our finance, HR and go-to-market teams use it. Even as a fairly AI-forward company we still get “wow” moments: something that took a few hours now takes a few minutes, and it’s automated. Agents talk to other agents. I have one that reviews my calendar each day, prepares me for meetings, runs background checks and drafts follow-ups.
  • Shoby — That raises the obvious question: how do I regulate this safely? The only way is to play with it and iterate in environments built for that. With the Spanish bank, Codex lowered the technical bar to build, but approvals through security and compliance, and promoting a change from dev to staging to pre-prod to prod, still took weeks. Meanwhile challenger banks ship in weeks, not months. So we had to get architecture, risk and compliance in the room, show them the tool, and take them on the journey.

Agents talk to other agents now. One of mine reads my calendar every morning, gets me ready for each meeting, and drafts the follow-ups.

Naman Joshi · OpenAI
08 — From the room · Q&A

When does saved time turn into growth?

  • Audience — Most AI conversations right now are about productivity, doing things more efficiently. Has the conversation with your clients started to shift from productivity to growth? You might save hours, but is that saved time translating into growth?
  • Naman — Great question. Take customer service, a big cost for many banks and retailers. One client fields around 20 million calls a year, mostly for a handful of routine things: change a PIN, a declined card. The first use case was pure productivity: deflect and triage those calls. Good value, easy to tick off. But because the foundations and guardrails were already in place, we could turn a negative inbound call into a positive one, proactively surfacing, say, “Did you know you’re pre-approved for a credit card?”, the way a telecom might say “You’re eligible for a new phone.” The same agent that solves the problem can create revenue.
  • Shoby — Almost everything in AI comes back to three metrics: efficiency, revenue, cost. People start with efficiency because it’s the safest. But you can’t start and stop there. If developers are more productive, where does that time go, longer holidays or the next piece of value? Cross-sell and upsell is one of the biggest growth plays. Another is investment-banking-style research in wealth management: decades, sometimes a century, of client and market knowledge trapped in a few experts. The use case is to free that knowledge and arm financial advisers with exactly what they need, on demand, without waiting six weeks for a quant. Break the long task into smaller ones and you can see clearly where the value, the revenue, actually is.
AI almost always comes back to three things: efficiency, revenue, cost. Most people start with efficiency because it’s the safest.
Shoby Abdi · Altimetrik
09 — From the room · Q&A

Same prompt, two answers: how do you build consistency?

  • Audience — I work in banking, and a recurring observation is consistency of output. Put the same prompt into two different chats and you can get two different answers. That’s a problem in a role where consistency is everything, and at the end-user level the prompt is often the only lever we have.
  • Naman — Great question. The simple, technical answer is that every AI project should start with evals and guardrails. Before anything else, define what “good” looks like. Models are non-deterministic; you don’t know exactly what they’ll output. So you build guardrails and evaluations that coach the model on how to navigate the task.
  • Shoby — Exactly. The point of evaluations is that you, the organisation, define what a quality response looks like, independent of which model you use. There are established techniques: a golden-set ratio, where you define quality as a percentage; or LLM-as-a-judge, where a model scores which answers are correct. It usually comes down to a confidence score. Organisations often ask for 99% confidence, which sounds high, until you ask whether one wrong answer in a hundred is acceptable for your use case. Most major platforms, OpenAI included, have evaluations built in, running before anyone sees a result. With one large US bank, the lesson was stark: for every line of AI code we wrote, we wrote roughly ten lines for quality and consistency. If you take one thing back, ask your organisation, “What’s our eval strategy?” That alone will kick-start the right conversation.
The discipline

Most of the work is the code around the model

1
line of AI code
:
10
lines of quality and
consistency code
A ratio Shoby cited from a build with a large US bank, alongside the one question he tells every team to ask: “What’s our eval strategy?”
Shoby Abdi · Altimetrik

Key takeaways

The debate is over; production is the problem.
Leaders have stopped asking whether AI works. The hard part is escaping pilots and running AI in production, safely, inside a regulated business.
You can’t blame the regulators anymore.
The ECB and EU AI Act, amid fragmented advice elsewhere, have defined what safeguarded AI looks like. The blocker now is internal.
Pilot purgatory is self-inflicted.
Focus on a pilot and a pilot is all you’ll get. Production-readiness in 2026 is mostly a culture-and-conviction problem, not a technology one.
Mind the capability overhang.
Models are advancing faster than enterprises adopt them, and the gap is widening. Closing it takes executive vision and a named champion.
Crawl, walk, run, then accelerate.
Start with small fires across the business, shape them into a simple, nameable programme owned by the business lines, then scale to production.
Sandbox to scale, not to play.
A governed sandbox, with semantic-gateway access and guardian agents, is how you satisfy regulators and grow, not just experiment.
Move from efficiency to growth.
Efficiency is the safe start; the prize is revenue: cross-sell, upsell, and freeing trapped expert knowledge for advisers on demand.
Consistency is engineered.
Non-deterministic models demand evals and guardrails. Define quality yourself, and expect to write far more code for consistency than for the feature itself.

The speakers

Naman Joshi
Solution Engineering Leader, EMEA · OpenAI
Naman Joshi leads solution engineering for OpenAI across EMEA, based in London, helping enterprises move AI from pilots into production. He joined from Databricks, and earlier spent seven years at Splunk and more than seven years in enterprise platforms at Macquarie Group. He studied computer and biomedical engineering at UNSW and completed the AI Strategy and Leadership programme at Oxford’s Saïd Business School in 2025.
Shoby Abdi
Senior Growth Partner, Digital Business & AI · Altimetrik
Shoby Abdi is a Senior Growth Partner in Altimetrik’s Digital Business and AI practice, based in Chicago, working with enterprise clients on adopting and scaling AI in regulated settings. He spent most of the previous decade at Salesforce, latterly as a principal architecture evangelist behind the Salesforce Well-Architected framework, after starting out as a Force.com developer. He holds a master’s in software engineering from DePaul University and an MBA from Lake Forest Graduate School of Management.

Contact Us

We'd love to hear from you.
Contact Us