24 AI systems in live use in UK organisations and their peers. What each does, what its owner reports in their own words, and what it takes.
Each entry is in production, not piloting, on the evidence of a public source read on the date given. The outcome is quoted from the owner or its supplier; where no outcome has been published, the entry says so. “What it takes” is our analysis of what a comparable organisation would need.
11 publish a measured outcome
5 publish usage only
8 publish no figure
Professional services
Law, audit and advisory firms.
Garfield.Law
A regulated law firm that runs small debt claims on AI
The problem
For small debt claims, legal costs can be a barrier to recovering an unpaid invoice at all.
In production
Authorised by the Solicitors Regulation Authority in May 2025, the firm produces pre-action letters, claims, witness statements and trial bundles. In May 2026 it took a £7,000 unpaid-invoice claim to trial at Wandsworth County Court and won.
Reported outcome
“The claimant paid around £400 in Garfield AI fees to recover the £7,000 owed.” No figures across all cases are published.
What it takes Vardonne’s analysis
Any firm copying this needs the claims process mapped as rule-bound steps and a clear point where a qualified person reviews what goes out. What made this a law firm rather than a document generator is regulation: SRA authorisation and the accountability that comes with it.
Gemini built into due diligence and case management
The problem
Lawyers spend long hours reading and synthesising large document sets across due diligence, disputes and multi-jurisdiction research.
In production
More than 5,000 people across the firm use tools built on Google’s Gemini models, and over 2,100 regularly use NotebookLM Enterprise. Gemini sits inside the firm’s due diligence, case management and multi-jurisdictional insights platforms, and its own general AI platform.
Outcome
No outcome figure published. The firm says Gemini “is no longer an experiment at Freshfields – it is infrastructure”.
What it takes Vardonne’s analysis
Governed document stores, access controls that respect client confidentiality, and customer-managed encryption keys. The adoption came with champions and a training academy: the change cost is at least the size of the technology cost.
General-purpose assistants bought firm-wide are often used unevenly, so the licences never pay back.
In production
Microsoft Copilot is used in every role for checking emails, summarising, analysing and researching. The firm states the tool “is not legally trained” and keeps it away from work that needs legal expertise.
Usage reported
Staff reached one million prompts four months early, releasing an extra £1m into the bonus pot. That is a usage figure, not an outcome.
What it takes Vardonne’s analysis
Incentives, training and a clear policy on what the tool may be used for; little integration beyond the Microsoft 365 tenant. The risk is rewarding activity rather than value, so set outcome measures beside the prompt count.
KPMG International (KPMG Clara) Outside the UK · Global network; the source does not separate the UK
AI agents for substantive audit procedures
The problem
Audit teams spend much of their time on repetitive, high-volume testing and disclosure checklists.
In production
Agents deployed into the global KPMG Clara platform for expense vouching and the search for unrecorded liabilities, with a report analyser for disclosure checklists. The release “aims to empower more than 95,000 auditors globally”; it does not say how many engagements use the agents.
Outcome
No outcome figure published.
What it takes Vardonne’s analysis
Standardised client data feeds and a methodology that says what the agent does and what the auditor must evidence and sign. Inspection-ready documentation, not the technology, sets the pace.
Call summaries and complaint responses written by AI
The problem
Contact-centre and complaints staff spend a large share of their day writing up calls and drafting responses.
In production
Automated call summaries and simpler complaint responses in Retail; summarisation for relationship managers in Private Banking. About 60,000 colleagues have access to AI tools.
Reported outcome
“In Retail, AI tools have already saved more than 70,000 hours”; in Private Banking, summarisation “freed up 30% more time for customer conversations”.
What it takes Vardonne’s analysis
Reliable transcription tied to the CRM, and quality sampling good enough for a summary to stand as the record of the call. Complaint responses in a regulated bank also need human approval and an audit trail that fits Consumer Duty.
A knowledge assistant for colleagues answering customers
The problem
Colleagues lose time searching scattered policy and procedure while a customer waits.
In production
Athena, an internal search and knowledge assistant, is used by 20,000 colleagues. It sits beside coding assistants for about 5,000 engineers and an HR assistant.
Reported outcome
“It has reduced search times by 66% on average”; across the group, generative AI “delivered around £50 million of value” in 2025.
What it takes Vardonne’s analysis
A maintained, version-controlled policy library with owners and role-based permissions. Governance has to reach the answers colleagues pass on: cited sources and regular accuracy testing.
Medical reports summarised for protection underwriting
The problem
Underwriters read long, complex medical reports before deciding on life and critical illness applications.
In production
A generative tool summarising medical reports went live for life cover in November 2025 and was extended to critical illness in March 2026. The tool supports underwriters, who make the decisions.
Reported outcome
“This has reduced the time underwriters spend reviewing each case by around half.”
What it takes Vardonne’s analysis
Digital intake of GP reports, a summary format built around the underwriting rules, and accuracy testing against human underwriters before use. Special-category health data means firm rules on who decides, audit trails and data protection controls.
Push-payment fraud is a large and costly problem for UK banks and their customers, and a buyer on an online marketplace cannot easily tell a genuine listing or seller from a fake one.
In production
Scam Intelligence launched in October 2025 for personal and business customers who opt in. Customers upload a listing, a message or a text and the tool flags warning signs in seconds. Customer data is not used for model training.
Outcome
No outcome figure published.
What it takes Vardonne’s analysis
A cloud set-up that keeps uploads out of training, and a design that combines the model with the bank’s own risk signals. Governance must settle how an advisory verdict is worded and what happens when the tool misses a scam.
NHS trusts, care providers and the services around them.
NHS England stroke services (Brainomix 360 Stroke)
Stroke scans read with AI decision support
The problem
Stroke patients need fast scan interpretation and transfer decisions so that eligible patients reach thrombectomy in time.
In production
Decision-support tools were rolled out to every stroke centre in England in summer 2024; the Brainomix tool runs across a network of more than 70 hospitals.
Reported outcome
Thrombectomy rates at participating sites doubled “from 2.3% to 4.6%”, and primary stroke centres saw “a 64-minute reduction in door-in-door-out time” (NHS England, citing The Lancet Digital Health).
What it takes Vardonne’s analysis
Integration with scanners and imaging systems, results that reach clinicians across hub-and-spoke networks, and the governance of a regulated medical device: clinical safety cases and monitoring against national audit data.
Falls and unnoticed deterioration among older people cared for at home lead to avoidable admissions.
In production
Carers, family members and care staff record patient updates in an app; the model predicts falls risk and flags deterioration. NHS England says it is used in more than two million home care visits a month, across over two thirds of integrated care systems.
Reported outcome
NHS England reports up to 2,000 falls and admissions prevented a day and hospitalisations reduced “by up to 70%”. These are headline figures; the method is not published.
What it takes Vardonne’s analysis
Consistent recording at every visit, and a defined response pathway for alerts into GPs and community services. Independent checking of the predictive claims, and clear accountability when an alert is raised and not acted on.
Mind Matters Surrey NHS Talking Therapies (Limbic Access)
A referral assistant for NHS talking therapies
The problem
Talking therapies services receive high referral volumes, and clinicians spend time assessing referrals that belong elsewhere.
In production
Listed as live on the government’s AI Knowledge Hub: a chatbot that takes self-referrals, identifies ineligible ones and redirects them to the right service.
Reported outcome
Ineligible referrals were identified and redirected, “creating additional clinician capacity. For Mind Matters Surrey, this applied to 20% of referrals.”
What it takes Vardonne’s analysis
Write-back into the patient record, the service’s eligibility and risk rules, and robust escalation when a patient discloses risk. A clinical safety case, equality monitoring of who completes, and a non-AI route that stays open.
Source: Limbic Access, AI Knowledge Hub (ai.gov.uk), updated 23 May 2025. Confidence high.
The Rotherham NHS Foundation Trust
An AI agent on the staff IT service desk
The problem
Rising demand for IT support with no room to add headcount; most staff phoned for routine queries.
In production
Since January 2026 an autonomous agent resolves routine queries in real time, raises tickets and routes staff to the right channel. An out-of-hours phase is under way.
Reported outcome
Help desk call volumes down 28%; “41% of IT queries are now handled through self-service and AI agents”.
What it takes Vardonne’s analysis
A kept-up IT knowledge base and integration with the service-management tool. Lighter governance than clinical AI, but clear escalation to engineers, above all for clinical-system outages.
Incoming letters scanned for claimants who need help now
The problem
Incoming post was sorted by hand, including the letters from people who need attention at once.
In production
In production on about 25,000 scanned letters a day. One classifier flags possible vulnerability, and an analytics team reviews the daily flags before they reach benefit teams; a second labels every letter by theme so it can be routed.
Outcome
No outcome figure published.
What it takes Vardonne’s analysis
Reliable scanning ahead of the model, a defined set of vulnerability themes, and a route into teams that can act the same day. The weight falls on human review and on testing for missed cases, because a missed flag is the costliest error.
Answers from government guidance, in the GOV.UK app
The problem
People struggle to find and understand the guidance that applies to them across a very large estate.
In production
GOV.UK Chat launched in the GOV.UK app in May 2026. Answers are grounded in official guidance and link to their sources; personal information is filtered out. Safety checks before release involved the AI Security Institute, and the service is monitored and evaluated for accuracy and safety.
Usage reported
Usage, not outcome: in the weeks after soft launch “more than 7,800 people” used it and it “handled more than 15,000 questions”.
What it takes Vardonne’s analysis
An authoritative content base kept current, continuous evaluation of answers as guidance changes, published limits on what the service will not do, and independent safety testing before it scales.
A small team of examiners oversees tens of thousands of garages and needs to visit where non-compliance is likely.
In production
A model rates about 23,000 garages red, amber or green to set the order of visits; around 150 officers use it daily. Enforcement rests on evidence found at the visit, never on the rating.
Outcome
No outcome figure published.
What it takes Vardonne’s analysis
A complete transactional dataset the agency already owns, which is why it works where most risk scoring stalls. The rating only prioritises; periodic checks that it is not skewed against particular kinds of garage.
Source: DVSA: MOT Risk Rating, GOV.UK, Algorithmic Transparency Recording Standard, 10 February 2025. Confidence high.
Newcastle City Council
Live prompts for council contact-centre agents
The problem
Agents must find the right answer across many services during a live call, and supervisors can check only a sample.
In production
Used daily across council contact services: real-time knowledge prompts and suggested responses, with transcription and analytics for quality and coaching. “Staff make all decisions and can ignore suggestions.”
Outcome
No outcome figure published; the record lists expected benefits, not measured ones.
What it takes Vardonne’s analysis
Suggested answers are only as good as the knowledge articles, so an owned content base comes first. Recording and transcribing calls needs a DPIA, clear notice to callers and retention rules.
Contractors, utilities networks, transport assets.
Yorkshire Water
Sewer monitors that warn before a blockage spills
The problem
Blockages are found after they flood or pollute, and crews are sent reactively.
In production
Part of a £53m programme begun in June 2025: 8,500 sewer level monitors and 8,000 customer sewer alarms installed so far, of 92,000 devices planned over five years, read by machine learning alongside rainfall and river levels. The control room is alerted to abnormal behaviour and an operative investigates.
Reported outcome
“It has prevented over 35 pollution incidents, reduced reactive site visits by 15%.”
What it takes Vardonne’s analysis
The model is the smaller part. Dense sensors, telemetry and joined weather feeds are the larger one, and value depends on crews with the capacity to act on alerts and on tracking false alarms so operators keep trusting them.
Site and office time goes on emails, meeting notes and documents rather than delivery.
In production
Microsoft 365 Copilot launched across the business in August 2025 after a 2024 pilot, backed by £7.2m. 6,848 colleagues had used it by the annual report.
Usage reported
“1,535,839 actions taken – assisting 115,178 hours of tasks such as drafting emails, summarising meetings and generating documents.” A usage measure.
What it takes Vardonne’s analysis
Clean permissions in SharePoint and Teams, because the assistant surfaces whatever a user can already reach. Pair the licence roll-out with task-level measurement and training.
The contact centre fields many repetitive enquiries that existing information could answer.
In production
In production: the assistant recognises common questions and answers with links, otherwise hands over to a live agent or opens a case out of hours.
Usage reported
Usage: an average of 2,151 chat sessions a month, of which 371 a month transfer to the contact team.
What it takes Vardonne’s analysis
A curated set of questions and answers kept in step with operations, and a clean hand-over into case management. Monitor low-confidence matches and the share handed to staff.
Source: Network Rail: Digital Assistant, GOV.UK, Algorithmic Transparency Recording Standard, 17 December 2024. Confidence high.
Heathrow Airport
Cameras and AI on aircraft stands to speed turnarounds
The problem
Slow turnarounds between flights cascade into late departures.
In production
A camera network being installed across the stands, with AI analysing the footage to speed turnarounds. Live, not yet at full coverage.
Outcome
No outcome figure published for this system.
What it takes Vardonne’s analysis
Cameras at every stand, integration with operational data and handlers, and agreed definitions of each turnaround milestone. Governance of workforce privacy on the apron and of how the data is used with handlers.
National Energy System Operator, with Open Climate Fix
Solar forecasting in the national control room
The problem
Solar output changes minute by minute; forecast errors mean last-minute backup and imbalance costs.
In production
The Quartz Solar forecast, combining machine learning with satellite imagery and weather data, is used in the Electricity National Control Room up to 36 hours ahead.
Reported outcome
Supplier’s figures: “2.8 times more accurate than previous solar generation forecasting tools”, avoiding “at least £30 million in imbalance costs each year”.
What it takes Vardonne’s analysis
Continuous satellite and weather feeds, a forecast that fits control-room tools and routines, and the old method kept as a fallback. Set your own baseline: the savings are the supplier’s estimate.
Picking tens of thousands of different items, varying in shape, weight and fragility, is hard to automate.
In production
On-grid robotic pick arms work in Ocado’s automated fulfilment, using machine learning to learn how to grasp new products. Sites and arm numbers are not given.
Usage reported
“In 2024, we picked over 30 million items using OGRP.”
What it takes Vardonne’s analysis
A highly standardised automated warehouse and large volumes of pick data to train and retrain on, which is why it does not transfer to a conventional warehouse. Expect a long ramp, with people picking what the arms cannot yet handle.
DHL Supply Chain, with HappyRobot Outside the UK · DHL Supply Chain, across several regions; countries not named
Voice and email agents for logistics coordination
The problem
Warehouses and transport generate a flood of routine calls and emails: bookings, driver follow-ups, status updates.
In production
Agents handle appointment scheduling, driver follow-up and status calls; current deployments “target hundreds of thousands of emails and millions of voice minutes annually”. DHL Supply Chain had spent over 18 months identifying and validating generative and agentic AI use cases.
Outcome
No outcome figure published.
What it takes Vardonne’s analysis
Live access to transport and warehouse systems so answers are accurate, escalation to people for exceptions, telling callers they are speaking to an AI, and monitoring the commitments the agent makes on the company’s behalf.
Nissan North America Outside the UK · United States and Mexico; published 2024
Vision inspection of paint on finished bodies
The problem
Inspectors miss some paint flaws, and finding their root causes is hard.
In production
A self-learning surface inspection system at three plants, learning defect types from technicians’ feedback and supporting inspectors rather than replacing them. It had evaluated over 500,000 vehicles at one plant.
Reported outcome
Defect detection up “by nearly 7%” at Smyrna; the engineer quoted says the system “identifies over 98%” of flaws against 85–95% for the human eye.
What it takes Vardonne’s analysis
Controlled lighting and cameras on the line, labelled defect images, and a feedback loop where technicians confirm or reject. Decide who signs off a body when the system and the inspector disagree.
Reviewed 21 September 2026. Entries are added as evidence appears; one that is retired is noted, not removed. To suggest a case, write to hello@vardonne.com.