
You don’t need 20 hires to run AI in production. Three to five people can do it, as long as they have the right platform, the right tests and a partner that hands over everything it builds.
Most AI hiring plans start with an org chart and hope a capability appears. This guide works the other way round. It covers which roles to own, which to rent and when to hire the fourth person. It also gives you a month-six test that shows whether you are building a capability or just paying for one.
Executive Summary
- Smaller team, faster output: three to five in-house owners can ship and run enterprise AI.
- Own the judgement, rent the skills: keep architecture, platform and sign-off in-house. Bring in specialists for fixed periods.
- The assets outlast the people: evals, platform, decision rights and runbooks should stay put when engineers leave.
- Choose a partner that hands over: build-operate-transfer pays the partner to become less necessary.
- Track the right metric: time to the second production use case tells you more than headcount.
Why the 20-Person AI Team Is the Wrong Starting Point
The 20-person AI team is a borrowed blueprint. It comes from hyperscalers and AI-native firms, whose core business is training and serving foundation models. That work needs research scientists, training infrastructure and large data labelling operations.
Most enterprises are doing a different job:
| Model Builders | Most Enterprises | |
| Core task | Train and serve foundation models | Connect existing models to internal data and business rules |
| Hardest problem | Model quality at scale | Integration, testing and safe operation |
| Hires that matter | Researchers and training engineers | Engineers who ship and run production systems |
You are applying models. That is engineering work, and it needs far fewer people than the blueprint suggests.
Copying the structure anyway costs you in four ways:
- You pay before you earn. Salaries start on day one. Production value arrives months later, if it arrives at all.
- Coordination eats capacity. Twenty people need managers, ceremonies and handoffs before anything ships.
- Skills miss the work. Research hires sit idle when the job turns out to be integration and operations.
- The people may not be there to hire. In the UK government’s AI Labour Market Survey 2025, 35% of organisations reported struggling to fill AI roles. 28% said technical skills shortages had affected their ability to achieve business goals.
So stop asking how many people you need. Ask what has to exist before your organisation can ship and run AI again and again.
What Is an AI Engineering Capability?
An AI engineering capability is the repeatable ability to take a business problem to a governed, monitored production system. Each time you do it, it should get faster.
Watch out for things that look like a capability but aren’t. A team name, a vendor contract, a platform licence or a Centre of Excellence slide can all exist while your organisation still cannot ship on its own. That gap between trying AI and running it is covered in more depth in a guide about AI experimentation vs AI engineering.
A real capability rests on four assets. Each one has a job to do and a clear cost when it is missing.
The Four Assets That Stay When People Leave
- Evaluation Harness: Test sets, scoring rules and regression checks that prove a change is safe to release.
Without it: every model or prompt update needs manual testing and a judgement call.
- Platform Layer: The shared plumbing: model gateway, retrieval, observability, access controls and cost limits.
Without it: every new use case rebuilds the same foundations from scratch.
- Decision Rights: Named owners for model approval, data access and go-live.
Without it: projects stall in review, or ship without one.
- Operating Knowledge: Runbooks, known failure modes and a log of which model and prompt versions ran when, and why.
Without it: incidents take longer to fix, and nobody can explain past decisions.
Test this against your own team. If your strongest engineer resigned tomorrow, these four assets would decide whether the capability stays or walks out with them.
AI Engineering Team Structure: The Lean Model That Works
Keep the team small and let the four assets carry the load. Your people own the assets, and the assets handle the repeat work.
Three AI Roles You Must Keep In-House
Each role answers a question that no one outside your organisation should answer for you.
AI Engineering Lead: “How should this be built?”
Picks the model and the retrieval pattern, and decides what to build and what to buy. That judgement can’t be outsourced.
Platform or MLOps Engineer: “Will it run safely and affordably?”
Runs the shared layer, keeps it secure, and tracks cost and latency. Without this role, every use case turns into a one-off.
Product Owner With Domain Authority: “Is the output good enough?”
Defines what “good” looks like and signs off on the evaluation criteria. This is usually an existing business leader, not a new hire.
Add one or two AI engineers as live use cases grow. That gives you a core team of three to five.
AI Skills You Can Rent Without Losing Control
Some skills are needed in short bursts. Bring them in when the need arises:
- Retrieval tuning: when a large knowledge base returns weak or irrelevant answers
- Fine-tuning or distillation: when a general model is too slow, too costly or too imprecise for the task
- Red-teaming: before a high-risk or customer-facing launch
- Data engineering: when a source system has to be cleaned or connected before any model can use it
Rent the skill, own the output. Every engagement should end with code, tests and documentation in your repositories, not the supplier’s.

AI Engineer, ML Engineer or Data Scientist: Who to Hire First
| Role | What They Do | Hire First If |
| AI engineer | Builds applications on existing models | Your use cases run on commercial or open models |
| ML engineer | Trains and serves custom models | You need a model no provider offers |
| Data scientist | Analyses data and runs experiments | You need insight, not a production system |
Most organisations should hire an AI engineer or platform engineer first. A data scientist adds more value once live systems are producing data.
Why AI Outsourcing Hands Over the Code but Keeps the Knowledge
Most AI engagements end with a tidy handover of code. The code is the easy part. What your team really needs is the reasoning behind it, and that rarely changes hands.
What transfers easily:
- Source code and infrastructure-as-code
- Architecture diagrams
- Access credentials
What usually stays with the supplier:
- Why one model was chosen over another
- Which prompts were tried, and why they failed
- The edge cases that broke the system in testing
- The judgement to know when a new model is worth the switch
The second list is what keeps an AI system working. Models change every few months, and every change raises the same questions. Only the people who hold the reasoning can answer them quickly. If they sit with a supplier, each model update becomes a new engagement, and you end up paying twice for knowledge you already funded.
How Build-Operate-Transfer Keeps AI Knowledge In-House
Build-operate-transfer is designed around that second list. A partner builds the platform and the first use cases, then runs them jointly with your team. Ownership passes to your team at agreed milestones.
The partner is paid to make itself less necessary. Your engineers take part in the decisions, not just the demos. The reasoning moves across as the work happens, rather than arriving in a handover pack at the end. AI engineering partner vs consultant explains how this differs from a typical consultancy engagement.
How an AI Evaluation Harness Replaces Manual Testing
In a lean AI team, the evaluation harness takes on the checking that a bigger team would do by hand. Picture the same event in two organisations: a provider releases a new model version.
Organisation A has no harness. Engineers test outputs by hand and argue over whether the results are better. Knowledge of edge cases lives in a few people’s heads. The upgrade takes weeks, and when the work grows, the fix is more hires.
Organisation B has a harness. The new model runs against hundreds of known cases in minutes. Regressions appear before customers see them, and a three-person team signs off on the switch with evidence.
What separates them is a system. A useful harness answers four questions before every release:
- Does it still get the right answers? Golden datasets of real inputs with agreed correct outputs.
- Does it fail safely? Adversarial cases such as ambiguous queries, prompt injection attempts and out-of-scope requests.
- Can you afford to run it? Thresholds for cost per request and response time, not just accuracy.
- Who approved it? The product owner signs off the criteria. The engineering lead signs off the release.
Run the harness like a product. Version it, review it and add a new test case after every incident.
A 12-Month AI Hiring Plan for a Lean Team
This plan assumes a build-operate-transfer model. Stretch or shorten the timings to fit your roadmap, but keep the order. Each phase ends with a gate, and you don’t move on until you pass it.
Months 0 to 3: The Partner Leads, You Learn
- Hire: an AI engineering lead who pairs with the partner from day one
- Deliver: the platform layer and the first use case, both built by the partner
- Gate: the product owner has signed off the first evaluation set
Months 4 to 6: Joint Delivery
- Hire: a platform engineer
- Deliver: the second use case, with your lead making the architecture calls
- Gate: your team owns and runs the evaluation harness
Months 7 to 12: You Lead, the Partner Supports
- Hire: one or two AI engineers, if the hiring signals below apply
- Deliver: new use cases led by your team, with the partner on call for specialist problems only
- Gate: a formal handover review confirms that the runbooks, repositories and infrastructure are fully yours
Where you are at month 12: three to five internal people, one platform and several use cases in production. The partner is now an option, not a dependency.

Contract Clauses That Keep AI Knowledge In-House
The contract decides whether you build a capability or rent one. Ask for these four clauses, and have your legal team draft the wording.
- Ownership From Day One
All code, infrastructure-as-code and model configuration live in accounts you own for the whole engagement.
Protects against: a handover that depends on the partner’s goodwill.
- Knowledge as a Deliverable
Evaluation sets and runbooks are named deliverables, accepted against agreed criteria at each milestone.
Protects against: documentation squeezed into the final week, or never written at all.
- Paired Delivery
A named internal engineer works alongside partner staff on every critical component.
Protects against: systems that only the partner can explain.
- Exit-Readiness Reviews
At agreed points, your team shows it can run the system without partner support.
Protects against: discovering the gap on the day the contract ends.
A note for UK buyers: if the model relies on individual contractors, check IR35 status early. Long, embedded engagements under your direction can look like employment for tax purposes. An outcome-based contract with a delivery partner is usually structured differently.

How to Measure AI Engineering Maturity Without Counting Heads
Headcount doesn’t show whether your AI capability works. These five measures do:
| Measure | Healthy Trend |
| Time to second production use case | Falling |
| Model changes shipped without the partner | Rising |
| Evaluation coverage per use case | Rising |
| Cost per request | Falling |
| Incident recovery time | Falling |
Report them to the board every quarter.
Four Mistakes That Sink Lean AI Teams
According to RAND research, more than 80% of AI projects fail by some estimates, twice the rate of IT projects without AI. Two of the root causes RAND identifies, solving the wrong problem and lacking the right data, appear in the list below.
Starting with the showiest use case. Impressive demos rarely have a clear owner or a baseline you can measure against. Start with a dull process where results can be counted.
Waiting for perfect data. Data is never perfect. Pick a use case whose data is good enough today, and improve the data as you go.
Leaving costs uncapped. Model pricing scales with usage, so one popular feature can blow the budget in weeks. Set spending limits for each use case before launch.
Adding governance after the fact. Regulated sectors need decision rights and audit trails from the first use case. AI consulting specialist can build this in before development starts.
How Deployflow Builds AI Capability You Keep
The organisations getting value from AI can ship a second, third, and fourth use case without starting from scratch each time.
Deployflow’s AI engineering and automation services are built around that outcome, and the client results below show it.
A Platform That Takes the Next Use Case Without Rework
For a national-scale energy AI programme, Deployflow worked inside the client’s platform engineering team. Together they built a shared capability. New AI use cases can now be added to the platform without re-engineering the core infrastructure. More than 1PB of subsurface data is processed in real time across H100 GPU clusters. Every environment inherits its security and governance policies with zero manual steps.
A Team Built at Speed, With Knowledge Transferred
Zilch needed a tech team straight away and within budget. Deployflow assembled a dedicated team that delivered complex API integrations in one month.
“They assembled a dedicated workforce, enabling us to transform our vision into reality. Their seamless team-building and thorough knowledge transfer have been instrumental in bringing our product to life.”
Sean Hederman, CIO, Zilch
Delivery That Ends With the Client in Control
Deployflow oversaw delivery of Vodafone‘s SEPA-compliant payment platform, which went live across nine markets and now moves over €1.5bn a month. Platform failures fell by 85%, and uptime reached 99.5%. The internal team was trained throughout, so the handover at the end of the project went smoothly.
If you are weighing up a first AI hire or a first delivery partner, a free consultation with Deployflow is a practical place to start. It covers which roles to keep in-house, which skills to bring in, and a realistic 12-month route to running AI in production yourselves.
Building an AI Engineering Capability: Frequently Asked Questions
What skills should an AI engineer have?
An AI engineer needs strong software engineering first, then AI-specific skills on top. Core skills include Python, API design, cloud services and version control. The AI layer covers prompt design, retrieval-augmented generation, vector databases and working with model APIs.
Evaluation skills matter just as much: building test sets and measuring output quality. Security awareness is now essential, especially around prompt injection and data leakage. Deep maths and model training experience are useful but rarely required for enterprise application work. When hiring, look for evidence of shipped production systems rather than notebooks or demos. Shipping shows the person can handle the messy integration work.
Can existing software engineers become AI engineers?
Yes, and it is often the fastest route. Experienced software engineers already understand testing, deployment, security and production support. Those are the hardest parts of AI engineering to teach.
The gap is usually model behaviour, retrieval design and evaluation methods. Most strong engineers can close it within months through paired work on a real use case. Formal courses help, but pairing with an experienced AI engineer transfers judgement far faster. Upskilling also reduces hiring risk. You keep people who already know your systems, data and business rules. Pick engineers who are curious and comfortable with systems that do not behave the same way every time.
Who should own AI in an organisation: the CTO or a Chief Data Officer?
Production AI usually belongs with the CTO, though no single model fits every organisation.
AI systems run on the same platforms, pipelines and security controls as other software. Placing them under technology leadership keeps architecture, operations and incident response in one line.
A Chief Data Officer often owns data quality, governance and access policies. That role remains essential, because AI is only as good as its data.
What matters most is clarity. Name one executive accountable for AI in production, and define how the data function supports them. Split ownership without clear decision rights is a common cause of stalled AI programmes.
What is the difference between AI engineering and MLOps?
AI engineering builds AI-powered applications. MLOps keeps models running reliably in production. An AI engineer designs how a model fits into a product. That covers retrieval, prompts, integrations and user-facing behaviour. MLOps covers deployment pipelines, monitoring, versioning, retraining and cost control.
In a small team, one person often covers parts of both. The platform engineer role usually carries most MLOps work. The distinction matters as you scale. Without MLOps discipline, AI applications degrade quietly as data, models and usage change. Treat MLOps as the operating backbone that lets AI engineering work repeat safely rather than as a separate project.
How do you retain AI engineers in a competitive market?
You keep AI engineers by giving them production work, time to learn and clear ownership, not just higher pay.
Salary gets them through the door, but it doesn’t always keep them. Engineers stay where their work goes live, so limit proofs of concept that never launch. Set aside regular time for learning, because tools and models change every few months. Make it clear which systems each engineer owns. Don’t let one person become the only expert on anything, because that leads to burnout as well as risk. Documentation and pairing spread the load. Finally, show a visible career path into technical leadership or architecture.

You don’t need 20 hires to run AI in production. Three to five people can...
read full article

Claude went down across every surface on 29 September 2026. From 14:00 to 14:59 UTC,...
read full article

In 2015, Amazon launched its first Prime Day, a global shopping extravaganza created to celebrate...
read full article

