Skip to content

How to run a face recognition pilot

A useful face recognition pilot measures the system on the buyer's own cameras, population and thresholds, against criteria agreed before it starts. The essential discipline is writing the acceptance criteria first: a pilot judged after the fact by whoever ran it will always look like a success, because the criteria move.

Why run a pilot rather than compare accuracy figures?

Because accuracy figures are not portable. The same algorithm produces different error rates on visa photographs, mugshots and border-crossing images, and different rates again at a gate where the subject stops than in a corridor where they walk past at an angle. A number quoted without its dataset, its threshold and its capture geometry is not a claim a buyer can act on.

A pilot substitutes the buyer's own conditions for the vendor's. That makes the result specific and usable, and it also makes it uncomfortable β€” which is the reason to insist on it. A vendor who resists a pilot on your cameras is telling you something.

What should be agreed before the pilot starts?

Everything below should be written down and signed by both sides before any equipment is installed. This is the step that decides whether the pilot produces a decision or an argument.

  1. The task. One-to-one verification, one-to-many identification, or watchlist alerting β€” they are different problems with different failure modes, and a pilot that does not name one will end up measuring whichever went best.
  2. The population and its size. How many people are enrolled, and how many distinct individuals will pass the capture point during the pilot.
  3. The capture points, with their geometry and lighting recorded, and agreement that they will not be changed mid-pilot.
  4. The threshold, or the process for setting it. If the vendor tunes the threshold during the pilot, the results before and after are not comparable.
  5. The acceptance criteria, as numbers with a direction: what false-match rate and what missed-match rate are acceptable, at what throughput.
  6. The fallback. What happens when recognition fails, who handles it, and whether the time it takes is part of what is being measured.
  7. Who owns the data collected, how long it is kept, and when it is destroyed.

The threshold item is the one most often left open and most often decisive. Raising a threshold cuts false matches and raises missed matches; lowering it does the reverse. A vendor allowed to choose it after seeing the data can present almost any pilot as a success.

How many subjects and how long?

The honest answer is that it depends on the error rate you are trying to detect, and that most pilots are too small to distinguish a good system from a mediocre one. The arithmetic is worth doing rather than guessing at.

  • To observe a failure mode at all, you need enough attempts that it would be expected to occur several times. If missed matches are expected in roughly 1 in 100 attempts, 50 attempts will frequently show none β€” and prove nothing.
  • False matches need far more data than missed matches, because they are rarer. Detecting a false-match problem usually needs a gallery and a volume of non-matching attempts that a two-week pilot does not naturally produce.
  • Run across the full daily and weekly cycle. A pilot that runs only during office hours misses the lighting conditions that break entrances.
  • Include the population that fails. Children, wheelchair users, people in lawful head coverings, people wearing masks or heavy glasses β€” a pilot drawn from volunteers who are all staff aged 25 to 50 will overstate performance.

Where the volume for a statistically meaningful false-match measurement is genuinely unavailable, say so in the report rather than reporting a number the sample cannot support. A pilot that concludes "we could not measure this" is more useful than one that reports a rate with no power behind it.

What should be measured?

MeasureWhy it mattersCommon mistake
Missed matches (false non-match / false negative identification)Every one is a person the system failed to recognise, who now needs staff time.Counting only attempts that produced a result, and discarding the ones where no face was detected at all.
False matches (false match / false positive identification)Every one is a person wrongly identified as someone else β€” the error with the highest consequence.Measuring on too small a sample to detect a rate this low, then reporting zero as if it meant something.
Failure to acquireThe camera never got a usable face. Not an algorithm failure, but the visitor experiences it identically.Recording it as an algorithm error, which sends the fix to the wrong place.
ThroughputPeople per minute through the capture point, including the fallback queue.Measuring only the successful path, which is not the throughput the site experiences.
Time to resolve a failureHow long a staffed fallback takes. This sizes the staffing and therefore the business case.Not measuring it at all, which is the most common omission in the whole exercise.
Variation across the populationAggregate rates conceal groups the system serves worse, which is usually the number that matters most.Reporting one headline figure for everybody.
The measurements that decide whether a deployment works.

How should the results be reported?

A pilot report that a buyer can act on has a shape. It states the conditions, then the numbers, then what the numbers do not cover β€” in that order, because a number read without its conditions is what produces the accuracy claims this industry is full of.

  1. The conditions: cameras, placement, lighting, threshold, gallery size, dates and times.
  2. The counts, raw: attempts, matches, missed matches, false matches, failures to acquire. Rates are derived from these; publishing only rates hides the sample size.
  3. The breakdown by subgroup, wherever the sample allows it.
  4. What could not be measured, and why.
  5. What would have to change to meet the acceptance criteria, if they were not met.

A pilot result belongs to the site it was run at. It does not transfer to another building, another camera set or another population, and a vendor quoting your pilot result to their next customer is misrepresenting it β€” including if that vendor is Ayonix.

What does a pilot not tell you?

  • How the system behaves at full scale. A gallery of 500 and a gallery of 500,000 are different problems, and identification error rates rise with gallery size.
  • How it degrades. Two weeks does not reveal what happens as enrolment images age, as the population changes, or as a camera drifts out of focus.
  • How it behaves under attack. Presentation-attack resistance needs deliberate adversarial testing, not incidental observation.
  • Whether operators will use it as designed. Automation bias appears under time pressure, which a pilot with observers present tends not to produce.

These are reasons to plan a review after deployment rather than reasons to skip the pilot. The pilot answers whether the system can work here; the review answers whether it still does.

Frequently asked questions

How long should a face recognition pilot run?
Long enough to cover the full daily and weekly cycle of the site, including the lighting conditions at the worst times of day, and long enough to produce enough attempts that the failure modes you care about would be expected to occur several times. Two weeks is common; whether it is sufficient depends on volume, not on the calendar.
How many people do I need in a face recognition pilot?
More than most pilots use. If missed matches are expected in roughly 1 in 100 attempts, 50 attempts will frequently show none and prove nothing. False matches are rarer still and usually need more data than a short pilot naturally produces β€” in which case the report should say the measurement was not possible rather than report an unsupported rate.
Who should set the matching threshold during a pilot?
It should be agreed before the pilot starts, or the process for setting it should be. A vendor allowed to tune the threshold after seeing the data can present almost any pilot as a success, because raising it cuts false matches while lowering it cuts missed matches and there is no single correct setting.
What acceptance criteria should a face recognition pilot have?
Numbers with a direction, written down before the pilot: an acceptable false-match rate, an acceptable missed-match rate, a required throughput, and a maximum time to resolve a failure through the staffed fallback. Criteria agreed afterwards are not criteria.
Can I compare pilot results between vendors?
Only if the conditions were identical β€” same cameras, same placement, same gallery, same threshold policy, same population, same days. Otherwise you are comparing sites rather than systems. Running vendors sequentially on the same installation is closer to a fair comparison than running them at different sites.
Does a successful pilot mean the system will work at full scale?
No. Identification error rates rise with gallery size, so a result against 500 enrolled people does not predict the result against 500,000. A pilot answers whether the system can work at this site under these conditions; scale, ageing enrolment images and operator behaviour under real time pressure need a review after deployment.

Jan Mocary β€” Chief Technology Officer, Ayonix

Leads engineering for Ayonix face recognition and the ATLAS agent platform, including their on-premise and air-gapped deployment modes.