PONOPT FIELD NOTES · Закупка AI

30 Questions to Ask an AI Video Analytics Vendor Before a Pilot

Grouped into six decision areas, these 30 vendor questions turn a strong demo into testable claims: outcomes, accuracy, privacy, security, integration and cost.

Buy an AI video analytics pilot the way you would buy a production system: start with the business outcome, demand evidence of accuracy in your own environment, verify where the video is processed, and put every promise into a written go/no-go plan. The questions below are grouped into six decision areas that convert a persuasive demonstration into claims you can actually test.

Key takeaways

  • Define the single metric the pilot must move and write explicit pass/fail thresholds before the first camera goes live.
  • Accuracy figures from controlled demos mean little; ask for false-positive and false-negative evidence from sites like yours and validate on your own footage.
  • Find out where processing happens — edge or on-premise designs that keep raw video on site cut privacy, latency and bandwidth exposure.
  • Separate compliance from marketing: demand SOC 2/ISO 27001 evidence, DPIA support, clear data ownership and a contractual opt-out from model training.
  • Run the pilot on your hardest site, not your showcase site, and cost the whole lifecycle including compute, storage, annotation labour, support and the exit path.

Anchor the pilot to a business outcome

A pilot that only tests whether a model 'sees things' tends to look good on screen and prove nothing about your operation. Before any vendor demo, agree on the exact metric the pilot should move — near-miss frequency, retail shrinkage, perimeter-incident response time, PPE-compliance rate, queue or dwell time. Make that one outcome the anchor for every later decision, so the pilot becomes a business validation rather than a technology experiment.

Once the outcome is fixed, write down the thresholds that count as success before launch: the minimum acceptable detection performance, the maximum tolerable alert volume per shift, how the baseline will be measured, and who signs off go/no-go. A vendor that cannot connect detection capability to a measurable operational result is proposing a demonstration, not a validation, and scope drift after launch is almost guaranteed without agreed thresholds in writing.

  • Which specific metric is this pilot designed to move, and how will it be measured at baseline and at the end?
  • Can you list exactly which detection events and alert types are in scope — and which are explicitly out of scope?
  • What written pass/fail thresholds define success, and who decides whether the pilot moves to production?
  • What happens if results are mixed, for example privacy and integration pass but night-shift accuracy needs more time?

Test accuracy the way the system will actually run

Accuracy quoted from a controlled demo rarely survives contact with a real site. Ask for measured false-positive rates (alerts for events that did not happen) and false-negative rates (missed events) from deployments with lighting, occlusion, camera height, traffic density and shift patterns comparable to yours. Aggregate accuracy numbers tell you little; figures broken down by failure class tell you what actually breaks.

Agree on how you will validate: use your own held-out clips the vendor has not tuned to, and define the ground-truth or reference data used for the measurement. Ask how performance behaves under your worst conditions and how configurable confidence thresholds change the balance between catching real events and drowning operators in noise. Frameworks such as IEC 62676-6 exist precisely because the industry needed standardised performance testing and grading of real-time video content analysis.

  • What measured false-positive and false-negative rates do you have from sites similar to ours (lighting, occlusion, density, shift patterns)?
  • Can you break down accuracy by failure class instead of giving one overall percentage?
  • How does performance change in our worst conditions — night shift, backlight, rain, crowds, dust or vibration?
  • What confidence thresholds can we configure, and what precision-versus-recall trade-off applies at each setting?
  • Can we validate on our own held-out clips, and what reference or ground-truth data will be used for scoring?

Find out where the video actually goes

Where processing happens determines most of your privacy, latency and bandwidth exposure. Ask whether inference and anonymization run on an edge device or on-premise, or whether raw footage is continuously streamed to a vendor cloud for analysis. In privacy-first designs only event metadata or short anonymized clips leave the site when a configured event occurs, which keeps unnecessary footage inside your control.

Then clarify what is stored, where it resides, how long it is kept, and who can access it. Confirm retention and deletion controls in detail, and address the single most common hidden risk in video procurement: whether your footage or derived data may be used to improve the vendor's product. If you do not want that, the restriction belongs in the pilot contract, not in a spoken assurance.

  • Where does inference and anonymization happen — on the edge device or on-premise — and does raw footage ever leave the site?
  • What data leaves the site (only event metadata or anonymized clips?), what anonymization is applied, and is re-identification possible?
  • Where is event data stored, what data-residency options exist, and how are retention and deletion configured per site?
  • Who can view footage and event clips, and how are access controls and audit trails structured?
  • Will our footage or derived data be used to train your models, and can we opt out in the contract?

Demand evidence, not marketing, on security and compliance

Security claims are only as useful as the documents behind them. Ask for SOC 2 Type II or ISO 27001 evidence, a recent penetration-test summary, and documentation covering multi-factor authentication, role-based access and audit logs before you sign. Watch the nuance between SOC 2 Type I (a point-in-time design of controls) and Type II (operating effectiveness over a period) — if your policy requires Type II, that difference is material.

Governance and ownership must be settled up front: who owns footage, clips, metadata and model outputs, which subprocessors can touch the data, and for which jurisdictions the compliance architecture is built. If you operate under the GDPR, the European Data Protection Board's guidelines on video devices are a natural reference point; ask whether the vendor will support your data-protection impact assessment rather than leaving you to build it alone. Where face recognition or other biometric processing is not genuinely needed, exclude it to shrink the legal surface.

  • Can you supply SOC 2 Type II or ISO 27001 evidence, a recent penetration-test summary, and MFA, RBAC and audit-log documentation before the pilot?
  • Have you supported a DPIA for GDPR-regulated video deployments, and which other rules do you map to — HIPAA, CCPA or local surveillance law — for our geography?
  • Who owns the footage, event clips, metadata and model outputs, and what happens to them if we terminate the agreement?
  • Which subprocessors have access to our data, and under what contractual and security conditions?
  • Does the solution involve biometric data such as facial recognition, do we actually need it, and can it be excluded from scope?
  • Which retention and data-governance defaults apply, and can they be changed per site or use case?

Check what it connects to and how it scales

A vendor that forces a camera rip-and-replace project is quietly adding cost and lock-in to your decision. Ask for the certified camera list and which video management platforms are natively integrated, typically through ONVIF or RTSP protocols and common VMS products. Then examine the API: what events it exposes, what export formats exist, and whether it can feed your ticketing, access control, warehouse or SIEM systems without manual correlation.

Scalability is as much a licensing and architecture question as a technical one. Confirm maximum cameras per site, total cameras, sites per platform and concurrent users, plus behavior during an internet or power outage and the end-to-end latency that includes capture. Finally, pin down how pricing scales — per camera, per site or per use case — because the economics of adding the tenth site differ sharply from the first.

  • Which cameras and VMS platforms are certified for your solution, and can you integrate without replacing existing hardware?
  • Which VMS, access-control, ticketing, WMS/MES or SIEM systems do you connect to natively, and what does your API expose?
  • What happens during an internet or power outage, and what is end-to-end latency including capture at our camera count?
  • What are your limits for cameras per site, total cameras, sites per platform and concurrent users?
  • Is pricing per camera, per site or per use case, and how does it change when we expand?

Plan the workflow, support and economics before day one

A detection that nobody acts on has no value, so define the operating model before launch: who receives each alert, who reviews event clips, how tasks are assigned and how closure is tracked. Ask how the system is recalibrated when site layouts, workflows, lighting or camera views change, because a model frozen at go-live will decay as your operation evolves. Treat pilot support as a preview of the production support model — confirm a named contact and documented response times by severity.

Finally, cost the entire lifecycle rather than the sticker price. Ask for an itemized three-year view covering licence, compute, storage, annotation labour and support, and confirm what triggers a price change. Understand the exit terms before you sign: what data and models you keep, in what format, and what you can export if you terminate after year two. A clear exit path protects you from vendor lock-in when the business case changes.

  • Who receives each alert, who reviews event clips, how are tasks assigned and how is closure tracked?
  • What support do we get during the pilot — a named contact and response times by severity tier — and how does it change in production?
  • How is the model recalibrated when layouts, lighting, workflows or camera views change, and who pays for that work?
  • What is the full itemized three-year cost (licence, compute, storage, annotation labour, support), and what triggers a price change?
  • What is your exit plan — which data and models can we keep, in what format, if we terminate after year two?

Red flags that should slow down any pilot

Some signals predict trouble better than any feature list. Be wary of pricing opacity, closed APIs that force proprietary hardware, sparse documentation, accuracy promises without supporting evidence, and a reference list of pilots rather than live multi-site deployments. A vendor that cannot point to deployments maintained for more than a year in environments resembling yours is asking you to take unproven risk.

Equally telling is how the vendor reacts to your request for evidence and to a site that is not flattering. Willingness to validate on your worst footage and to put thresholds in the agreement is a strong sign of maturity; defensiveness or delay is not. Run the pilot on your hardest site, reserve budget for a second, follow-on model built after the honeymoon of vendor attention, and assign a named owner for data quality.

  • Can you name live deployments in our sector and geography that have been running for more than 12 months, and give us reference contacts?
  • Which of your claims can be documented in writing before the pilot, and which can only be shown after signing?
  • What is the alert volume expected in the first 30 days, and how does alert fatigue get managed during ramp-up?
  • Do you provide a data-flow diagram and a documented ownership statement as part of the pilot package?

Weighted go/no-go scorecard for the AI video analytics pilot

Score each candidate vendor from 1 (weak) to 5 (strong) on every line, multiply by the weight that reflects your priorities, and sum the totals. Use a pass line you set before scoring — for example 80% of the maximum — and treat any vendor who cannot answer a question as scoring a 1 rather than a guess. This card works across shortlists because the same questions and weights are applied to every vendor.

  1. Business alignment: how precisely does the vendor connect detection to the one metric your pilot must move?
  2. Real-world accuracy: documented false-positive/false-negative evidence from sites like yours, broken down by failure class.
  3. Validation honesty: willingness to be scored on your own held-out clips and your worst conditions, not a tuned demo.
  4. Privacy architecture: edge or on-premise processing so raw video stays on site, with clear anonymization and re-identification risk.
  5. Data governance: ownership of footage and outputs, retention and deletion controls, and a contractual model-training opt-out.
  6. Security posture: SOC 2 Type II or ISO 27001 evidence, penetration-test summary, MFA, RBAC and audit logs in scope.
  7. Integration depth: certified cameras and native VMS/API connections to your ticketing, access-control, WMS or SIEM systems.
  8. Scalability and licensing clarity: per-camera/site/use-case economics and documented limits for cameras, sites and users.
  9. Support and SLA: named pilot contact, response times by severity, and a defined transition to production support.
  10. Lifecycle cost and exit: itemized three-year cost and a written plan for data and model export if you terminate.

Questions people ask

How long should an AI video analytics pilot run to give trustworthy results?

Long enough to capture normal variation across shifts, lighting, weather and operating patterns — usually several weeks rather than a few days. Agree on the data-collection period before launch and define how many events or how much baseline footage are needed before scoring starts. Do not approve production based on ideal daytime conditions alone; require evidence covering the night shift and other worst-case operating periods.

What is the most reliable predictor that a video analytics vendor will scale to production?

The vendor's live deployment record in environments similar to yours, maintained for more than 12 months across multiple sites. Pilots and controlled demonstrations are weak evidence; multi-year production deployments show that the vendor has handled alert volume, recalibration, support and integration at operational rigour. Ask for reference contacts at clients in your sector who have used the platform in production for over a year.

Do I need facial recognition to get value from AI video analytics?

Usually not. Most safety, operational and security analytics — detecting people, vehicles, objects, loitering, perimeter intrusion or PPE issues — work without identifying individuals. Excluding face recognition removes a special-category processing burden under GDPR and shrinks the DPIA scope. Note that public testing such as NIST's FRVT work shows accuracy can vary across demographic groups, so biometric features should only be bought when a specific need justifies them.

Why should I ask about false positives and alert fatigue before signing a pilot?

A system that generates constant false alarms trains operators to ignore alerts, which is operationally close to having no alerting at all. False positives raise staffing cost and erode trust, while false negatives leave real events undetected. Ask for both rates from live deployments similar to yours, set a maximum tolerable alert volume per shift, and define how alert fatigue is managed during the first 30 days.

Should the pilot run on my best site or my worst one?

Run it on the hardest site, not the showcase one. A well-lit, cooperative flagship site flatters any vendor and tells you little about production conditions. The pilot is where you discover how the system handles your real constraints — poor lighting, occlusion, crowds, network limits and staff behaviour — and that knowledge is what you need before committing to a multi-site rollout.

What is the difference between a benchmark accuracy figure and operational accuracy?

Benchmark accuracy is measured on controlled datasets that rarely match your cameras, lighting and scene density. Operational accuracy describes how the system performs on your own footage and conditions, including the failure classes that matter to your use case. Standardised performance testing of real-time video analysis exists partly to make such comparisons meaningful, but the decisive test is validation on your own held-out clips in your environment.

Sources and further reading

Sources were checked when this page was generated. Confirm changing dates, rules and prices with the original publisher.

  1. 12 Questions IT Should Ask Before a Pilot (Computer Vision Vendor Due Diligence)Protex AI
  2. How to Evaluate AI Video Surveillance Vendors: A 10-Point Checklist for Enterprise Security ArchitectsLumana
  3. 面向企业采购的计算机视觉供应商(Computer Vision Vendors for Enterprise Procurement)Ultralytics
  4. How to Choose an AI Video Analytics Vendor: The Evaluation Checklist That Actually WorksStaqu
  5. EVS-EN IEC 62676-6:2026 — Video surveillance systems — Part 6: Performance testing and grading of real-time intelligent video content analysisEVS (Estonian Centre for Standardisation)
  6. NIST Study Evaluates Effects of Race, Age, Sex on Face Recognition SoftwareNational Institute of Standards and Technology (NIST)