A live facial recognition trial at some of London’s busiest railway stations scanned more than half a million people over six months. The number of arrests directly generated by the system was zero.
British Transport Police, or BTP, deployed the technology 18 times between February and July 2026. According to The Guardian, the trial scanned more than 500,000 faces, cost £320,786 and generated only one alert. That alert was an incorrect identification.
The result creates an unusually concrete test of facial recognition’s value. The debate is no longer simply whether the technology can recognize faces accurately. It is whether continuously screening enormous crowds produces enough useful outcomes to justify the infrastructure, police time and biometric processing involved.
The system worked, but found almost nobody
BTP launched the pilot on February 11, 2026, initially at railway hubs including London Bridge. Its stated purpose was to assess how live facial recognition performs in a railway environment and help identify and apprehend individuals wanted for serious criminal offences.
The technology compares faces captured by cameras against a predefined watchlist. BTP uses NEC’s NeoFace M40 algorithm. When the software detects a possible match, it generates an alert; an officer then visually compares the camera image with the person before deciding whether to intervene.
That human-in-the-loop step is important because a facial-recognition score is not itself proof of identity.
Yet the striking feature of the six-month results is not a flood of false alarms. It is the near absence of alerts altogether.
The trial used almost 100 hours of police time, while the one automated alert produced during the initial period was wrong. Other arrests occurred around deployments, but they were not the result of facial-recognition matches and therefore were not counted as direct outcomes of the system.
Accuracy and usefulness are different measurements
Low numbers of false positives can sound like evidence that a system is working well.
Technically, they may be.
Live facial recognition is commonly evaluated using at least two measures. The true positive identification rate measures how often someone actually on a watchlist is correctly recognized. The College of Policing shared that the false positive identification rate measures how often someone who is not on the watchlist is incorrectly flagged.
Independent National Physical Laboratory testing cited by the UK government found that the policing algorithm tested had an 89% chance of identifying someone who was on the watchlist at the assessed settings. For a watchlist containing 10,000 images, the worst measured false-alert probability was about one in 6,000.
But those figures answer a different question from whether deployment is productive.
A system can be highly accurate when a wanted individual passes the camera and still generate almost no useful matches if watchlisted people rarely enter the recognition zone.
That is what makes the London trial important. It separates algorithmic accuracy from operational yield.
Scanning 500,000 people with very few false alerts demonstrates one form of technical performance. Scanning the same population without directly producing an arrest raises a separate question about where and when the technology provides enough benefit to justify deployment.
Thresholds change the balance between misses and false alerts
Facial recognition also involves an engineering trade-off that is easy to overlook.
Algorithms calculate similarity between a captured face and images on a watchlist. Operators set a threshold above which that similarity becomes an alert.
Raising the threshold can reduce false matches but may also make genuine matches harder to detect. Lowering it can capture more possible matches while increasing the chance that innocent people are flagged.
The UK government’s guidance explicitly notes that changing the threshold changes the accuracy of the system.
NIST’s broader facial-recognition evaluations show why this matters. Its research distinguishes false positives—incorrectly associating two different people—from false negatives, where the system fails to associate images of the same person. NIST also finds that performance can vary by algorithm, image quality and demographic characteristics.
The technical question is therefore not simply whether facial recognition is “accurate.” It is accurate under particular configurations, watchlists, cameras, populations and operating conditions.
Public acceptance depends on who operates the technology
Technical performance is only one part of adoption.
A 2026 peer-reviewed study in Computers in Human Behavior surveyed 507 participants about police uses of facial recognition. Researchers found that trust in law-enforcement institutions was the strongest predictor of acceptance, while greater general knowledge about AI was associated with lower trust and support for police facial recognition.
That finding challenges the assumption that skepticism is primarily caused by people not understanding the technology.
A larger UK Home Office survey of 3,920 respondents similarly found that 64% supported police use of facial recognition overall, but support changed sharply according to institutional trust. Among respondents who trusted police completely, 81% supported its use, compared with only 30% among those who did not trust police at all.
The public therefore appears to evaluate facial recognition as more than an algorithm.
People are also judging the institution controlling the watchlist, deciding where cameras are placed and determining what happens after an alert.
BTP extended the trial despite the initial results
The initial six-month period did not end the experiment.
In August, BTP expanded the pilot into selected London Underground stations and extended it until November 2026. The force says the extended period will help assess effectiveness, public-safety impacts and public reaction.
After the extension, three watchlist matches were reportedly generated, but the people involved were found to be complying with their court orders rather than violating them.
That does not prove facial recognition has no policing value. A larger or differently targeted deployment could produce very different results, and preventing or detecting a single serious offence can carry value that simple arrest counts fail to capture.
But the trial exposes the metric technology teams and policymakers now need to confront.
The important number is not only how accurately a machine recognizes a face.
It is how much public-safety value results from recognizing faces at scale—and whether that value is proportionate to scanning hundreds of thousands of people who were never suspected of anything.