AI Maturity and Transformation: A CTO’s Guide to Joining the 6% That Profit

Rising coin stacks with a 6% label, showing the share of firms that profit from AI maturity and transformation

Nearly nine in ten organisations use AI regularly. Only 6% attribute 5% or more of their EBIT to it, according to McKinsey’s State of AI survey. What separates the two groups is how far each has taken AI maturity and transformation. Most firms are stuck at the same stage, and the way out is an engineering decision. 

TL;DR

  • Where firms stall: In MIT CISR’s Enterprise AI Maturity Model, only Stages 3 and 4 are linked to above-average financial performance. 62% of firms are still in Stages 1 and 2.
  • Why they stall: Pilots prove that a model works. Moving up a stage needs a shared platform, governed data and internal ownership.
  • What to do first: Score six maturity dimensions against evidence, then fix the weakest one before funding another use case.
  • What the 6% do differently: Nearly 75% of them have fundamentally redesigned workflows around AI.
  • What the board should track: Time from sign-off to production, cost per production use case, platform reuse rate and EBIT attributed to AI.

What an AI Maturity Model Measures, and Why Adoption Numbers Mislead Boards

An AI maturity model measures how much AI reaches production, how much of it gets reused, and whether it changes financial results. Adoption numbers measure something else. Licence counts, active users and pilot totals rise quickly, and none of them shows value.

McKinsey’s 2026 State of AI survey shows the gap clearly. 44% of organisations say AI is scaling across their enterprise. Scale also depends on size: 40% of enterprises with revenue above $1bn are scaling AI agents, against 22% of smaller organisations. A board that tracks adoption sees growth. A board that tracks maturity sees the stall.

The Four Stages of Enterprise AI Maturity, and the Two That Pay Off

In MIT CISR’s Enterprise AI Maturity Model, only the last two stages are linked to above-average financial performance. The model comes from a 2022 survey of 721 companies, followed by executive interviews in 2024. The percentages have likely shifted since then. The stage logic still holds.

  • Stage 1, Experiment and Prepare (28% of firms). The focus is AI literacy, policies and building comfort with automation.
  • Stage 2, Build Pilots and Capabilities (34%). Firms run pilots, define metrics, simplify processes and start consolidating data.
  • Stage 3, Industrialise AI Across the Enterprise (31%). Firms build scalable architecture, reuse components and begin deploying agents.
  • Stage 4, AI Future-Ready (7%). AI shapes every decision, and firms sell proprietary AI capabilities as new services.

The researchers name culture as the hardest part of leaving Stage 1. It means moving from command-and-control to coach-and-communicate. Leaving Stage 2 needs investment in enterprise architecture and secure, prepared data. That makes it an engineering problem, and it sits on the CTO’s desk.

Evidence test for the four stages of AI maturity and transformation, from a written AI policy to revenue from AI services

Why AI Pilots Fail to Reach Production: 4 Causes and Fixes

Most programmes stall at the jump from Stage 2 to Stage 3. A pilot has to prove that one model works. Stage 3 has to make every later use case cheaper than the one before it. Four patterns keep firms stuck at that line, and each one can be fixed before the next pilot starts.

Pilot Metrics Stop Short of the P&L

Pilots usually report accuracy, speed or user satisfaction. Finance cannot book any of them. The pilot works, but the funding request fails because nobody can put the result in pounds.

The fix: Before build starts, a business owner signs a financial target, such as cost per claim, hours per month or revenue per account.

Data Built for Dashboards Cannot Feed Models

Reporting estates tolerate monthly batches, loose definitions and manual fixes. A model making live decisions cannot. Pilots get around this with a one-off extract, and that shortcut fails the day real users arrive.

The fix: Treat the pilot’s data feed as a production data product from day one, with an owner, a refresh schedule and documented lineage.

Every Use Case Rebuilds the Same Plumbing

When each pilot ships its own retrieval layer, prompt store and monitoring, the second use case costs as much as the first. Reuse only happens when it is designed in. This shared layer is what separates AI experimentation from AI engineering.

The fix: Fund the shared layer as its own line item, separate from any single use case, so it outlives the pilot that paid for it.

Knowledge Leaves When the Vendor Does

External teams often build the pilot. When the contract ends, the people who understand the prompts, the evaluation logic and the data quirks leave with it. The next team starts from zero.

The fix: Write handover into the contract. Internal engineers pair on the build, and documentation is a deliverable with its own acceptance criteria.

Five warning signs that AI maturity and transformation has stalled, from ageing pilots to an unknown cost per AI request

AI Maturity Assessment: A Six-Dimension Scorecard for CTOs

Score each dimension from 1 to 4, and count only evidence that exists today. A 2 means the capability exists in one project. A 3 means it is shared, but not yet standard. Plans, budgets and slide decks score zero. Your lowest score sets the pace for the whole programme, so it matters more than the average.

DimensionScores a 1Scores a 4Evidence to Ask For
Strategy and Use-Case PortfolioA list of ideas with no ownersA ranked portfolio with an owner and financial target per itemThe signed-off target for your top use case
Data ReadinessEach pilot pulls its own extractGoverned data products that models use directlyLineage for the data behind one live model
AI Platform and ArchitectureEvery use case has its own stackA shared gateway, retrieval layer and component libraryHow many components the last use case reused
MLOps and LLMOpsModels change with no testingAutomated evaluation, monitoring and rollbackThe last time a model change was rolled back
Governance and RiskRisk reviews each project from scratchControls built into the platform and mapped to regulationHow long approval took for the most recent use case
Operating Model and SkillsThe vendor runs production AIInternal teams own and run itWho is on call when a model fails tonight

How to read the result: Any dimension at 1 or 2 keeps the programme at Stage 2, whatever the other scores say.

Internal teams tend to score intent rather than evidence. An external review through AI consulting services gives the score a neutral baseline before budgets are set.

Redesign Before You Automate: How AI High Performers Get Returns

High performers change the work before they automate it. McKinsey’s 2026 data shows how far ahead that habit puts them.

Nearly 75% of AI high performers have fundamentally redesigned workflows because of AI, up from 55% a year earlier. They are also twice as likely to report senior leadership commitment, and twice as likely to spend more than 15% of their ICT budget on AI.

Adding a model to an unchanged process limits the return to one step. Redesign removes steps, handovers and approvals. Insurance claims triage shows the gap:

Automate the step: A model reads each new claim and drafts a summary for the handler. Each case takes a few minutes less, and the queue stays the same length.

Redesign the flow: The model scores each claim for risk. Low-risk claims settle without a handler, and handlers work only the complex cases. The queue for simple claims disappears.

Three questions to ask before automating any workflow:

  1. Which steps exist only because a person used to do the work by hand?
  2. Which approvals could become rules that the system checks automatically?
  3. If AI handled the routine cases, what would your experts spend their time on?

When the answers lead to removing steps, fund the redesign first and the model second.

AI Transformation Roadmap: 4 Phases From Discovery to Production

Deployflow’s AI engineering and automation delivery runs in four phases. Each phase below carries one exit test that is hard to fake. A phase that fails its test should not hand over to the next one.

AI maturity and transformation roadmap in four phases, from discovery to optimisation, with the exit test and failure sign for each phase

The loop is what raises maturity. One use case in production is a delivery win. When each new use case reaches production faster than the last, the firm is operating the way Stage 3 firms do.

Who owns outcomes, models, data, evaluation and risk controls at Stage 2 vs Stage 3 of AI maturity and transformation

UK AI Regulation and the EU AI Act: Governance That Speeds Up Delivery

Governance built into the platform speeds up every launch after the first. Each new use case inherits controls the risk team has already approved. When controls live inside individual pilots, each launch restarts the review.

The Two Rulebooks UK Firms Answer To

  • In the UK: There is no single, cross-sector AI law. Sector regulators such as the FCA and the ICO apply existing rules on consumer outcomes, accountability and data protection to AI systems.
  • In the EU: Firms that serve EU customers fall under the EU AI Act. The AI Omnibus entered into force on 27 July 2026. High-risk obligations for Annex III systems now apply from 2 December 2027, and those for AI embedded in regulated products apply from 2 August 2028.

Build These Controls Once, at Platform Level

  1. An AI system inventory, so every model in use has an owner and a stated purpose.
  2. Risk classification at intake, so each new use case is tagged high-risk or not before build starts.
  3. Logging and evaluation by default, which also supports the record-keeping and accuracy duties for high-risk systems.
  4. Defined human oversight points, set once as policy and reused by every team.

The time before 2 December 2027 is the window to build these controls once. The AI mistakes UK-regulated industries make follow the Stage 2 pattern: controls are added, reviewed and rebuilt one project at a time.

How to Measure AI ROI: 4 Metrics Boards Trust

A board asks four questions about AI spend. Each question has one metric that answers it.

“Is delivery getting faster?”
Time from sign-off to production. Measure it from the day the business case is signed off to the first live user. Report the median across use cases, so one slow project does not skew the trend.

“Is the platform paying for itself?”
Cost per production use case. Include build, run and support costs for the first 12 months. Compare use cases in the order they launched.

“Are teams building on shared foundations?”
Platform reuse rate. Count the components each new use case takes from the shared layer, divided by the total it needs.

“Is AI changing the P&L?”
EBIT attributed to AI. Agree the attribution method with finance before launch. A number the finance team did not help define will not hold up in the boardroom.

Report all four every quarter on one page. Keep accuracy scores and licence counts for engineering reviews. 

Where to Start With AI Maturity and Transformation

Find your lowest-scoring dimension and fix it before you fund another use case. Three recent engagements show what that looks like: 

  • Platform and governance: For a national-scale energy AI client, Deployflow built a platform where every environment inherits security and governance policies automatically, with zero manual provisioning steps. New AI use cases now land on it without re-engineering the core infrastructure.
  • Data readiness: For a multi-billion-dollar UAE public-sector organisation, Deployflow designed AI pipelines to replace manual classification across disconnected surveys and spreadsheets, feeding one unified data layer. 
  • Ownership: For Vodafone, Deployflow oversaw delivery of a platform that moves €1.5bn a month across nine markets. Platform failures fell by 85%, and Deployflow trained the internal team to take over at handover.

“Deployflow has an amazing talent pool who demonstrated an extraordinary level of professionalism and transparency right from our initial interaction.”

Khev P. via G2

Talk to Deployflow about scoring your six dimensions and building a roadmap from your weakest one. 

Frequently Asked Questions: AI Maturity and Transformation

What Are the Five Levels of Gartner’s AI Maturity Model?

Gartner’s AI maturity model has five levels: Awareness, Active, Operational, Systemic and Transformational. 

At Awareness, the organisation discusses AI but has not started using it. Active means experiments and early pilots. Operational means at least one AI project in production, with executive sponsorship and a dedicated budget. Systemic means AI is used across the business and shapes its digital products. Transformational means AI is part of the business model itself. Gartner’s levels map closely onto MIT CISR’s four stages. Either model gives the leadership team and the board a shared language for progress.

What Is the Difference Between AI Readiness and AI Maturity?

AI readiness measures whether an organisation has the conditions to start using AI. AI maturity measures how far it has got once it has started. Readiness looks at inputs such as data quality, infrastructure, skills, budget and leadership support. Maturity looks at outcomes such as use cases in production, reuse across teams and financial impact. A firm can score high on readiness and low on maturity, for example with a modern data platform and no production AI. Readiness assessments suit firms at Stage 1. Once pilots exist, a maturity assessment is more useful because it tests what has actually been delivered.

Who Should Lead AI Transformation: the CTO, the CIO or a Chief AI Officer?

One executive should own the overall outcome, and business functions should own their individual use cases. The right title depends on the firm. In technology-led firms, the CTO often leads, because the platform work sits in engineering. When AI primarily affects internal operations, the CIO may be the better fit. A Chief AI Officer makes sense when AI spans many functions and needs its own budget and a voice at board level. The mandate matters more than the title. The leader needs authority over the shared platform, the data and the governance model, beyond a coordinating role.

How Is AI Transformation Different From Digital Transformation?

Digital transformation moves processes and customer journeys onto software and cloud platforms. AI transformation changes how decisions and work happen on top of those platforms. 

The first is largely about systems of record. The second is about systems that predict, generate and act. AI transformation depends on digital foundations, so gaps in data, cloud and integration from earlier programmes resurface quickly. It also adds new demands, including model evaluation, monitoring for drift and error, and governance over automated decisions. For most enterprises, AI transformation is the next phase of digital work, built on the same stack.

Can Small and Mid-Sized Businesses Use an AI Maturity Model?

Yes. The stages apply at any size, although the path through them is shorter and lighter. 

A mid-sized firm rarely needs its own model gateway or a large platform team. It can reach Stage 3 behaviours by standardising on one cloud AI service, one evaluation method and one owner per use case. The six dimensions still apply, and a single workshop is usually enough to score them. 

Smaller firms also have advantages, such as fewer legacy systems and shorter approval chains. The main risk is buying many separate AI tools with no shared data or governance. That recreates the Stage 2 stall at a smaller scale.