The short answer
Frame the RFP around an operational problem and a measurable outcome, not a feature list. Specify each detection scenario — what, where and when the system must find — ask every vendor identical questions about real-world accuracy, false positives, deployment and data governance, publish a weighted scoring matrix, and tie acceptance to a pilot on your own cameras with per-scenario recall and precision targets fixed before it starts. That turns a marketing exercise into a comparison you can defend.
Key takeaways
- Start from the operational outcome — what share of significant events must reach operators and how many false alarms a day the team will tolerate — because video analytics output is probabilistic, not deterministic.
- Specify each scenario precisely with the object or event, camera zone, schedule and alarm routing, so vendors cannot reply with generic capability statements.
- Ask every vendor the same questions about the conditions behind measured accuracy, live false-positive rates, deployment architecture, integration, subprocessors and references running for over a year.
- Treat a vendor demo as evidence of almost nothing; the real test is a shadow-mode pilot on your own footage with written criteria and per-scenario metrics agreed before it starts.
- Set recall and precision targets per scenario rather than one blended average, because an average hides misses on rare critical events.
- Publish scoring weights in the RFP so vendors address what actually decides the award, and weight the pilot and post-pilot price lock alongside feature claims.
- Verify data governance, retention, model-retraining cadence and line-item pricing to protect total cost and your regulatory posture.
Start with the outcome, not a feature checklist
A standard procurement template built around a feature checklist suits deterministic systems, but AI video analytics is probabilistic. The same platform that posts impressive accuracy in a curated demo can behave far worse under your site's lighting, camera angles and scene density. So the first thing to fix in the RFP is not the feature list but the operational problem and its measurable result.
Before release, decide what "working well" means in numbers: how many significant events out of a hundred must reach the operator and within what latency, how many false alarms per day the team can absorb, and which scenarios are critical. These figures become the acceptance baseline. Add a clause that the definitions will not change after results arrive — otherwise the evaluation dissolves into a dispute over terms.
Functional and technical requirements to specify
Describe each scenario independently: which object or event the system must detect, in which camera zone, during which hours, and with which reaction — alert to an operator, recording, or export to an adjacent system. A line like "a person enters the warehouse restricted zone between 10 p.m. and 6 a.m." leaves no room for generic promises.
Also ask vendors to state their limitations: how detection degrades in fog, rain, backlight, crowds or darkness. An honest description of the operating envelope is more valuable than a page of feature names.
- Event categories and interest zones per camera, including schedules and exceptions.
- Integration: support for your VMS, ONVIF/RTSP, an API and metadata formats for exporting events.
- Architecture: edge, on-premise or cloud; which data leaves the site and what latency is acceptable.
- Camera requirements: resolution and pixels-on-target in priority zones, plus low-light performance.
- Capacity: channel count, throughput, failover and redundancy.
- Storage and retention: where video and metadata live, default retention periods and who can access them.
Questions that separate a working platform from a demo
Ask every shortlisted vendor the same questions and demand figures from live installations rather than laboratory datasets. Probe the conditions under which accuracy was measured, the false-positive rate in a real crowded scene, whether the platform supports on-premise or edge deployment without sending footage off-site, and how data governance works.
The operational risk that matters most is rarely a single missed event; it is an avalanche of false alarms that trains staff to ignore the system. So always pair accuracy questions with questions about false-positive rates, duplicate events and how the platform reviews and suppresses irrelevant activity.
- Show detection results from deployments similar to yours in industry, camera type and lighting.
- What is the live false-positive rate and how are duplicate events handled?
- Which VMS platforms have native integration and what does the API expose downstream?
- Where and for how long are video and metadata stored, who are the subprocessors, is there a local loop?
- References from clients in your sector running the platform for over a year, with contacts you can call.
- Pricing model — per camera, site or use case — and how it scales as you expand.
Acceptance criteria and how to test them
The strongest evidence is a shadow-mode pilot on your own video streams running beside your current operation, not a vendor demo. Agree in writing before it starts: duration, the number and variety of cameras, the set of test scenarios, named reviewers on both sides, and what counts as success. Include hard cameras as well as easy ones — night, crowds, poor lighting and known nuisance-alert sources.
Measure metrics per scenario rather than as one blended number. The core measures are recall (the share of real events correctly found) and precision (the share of alarms that were actually warranted), together with specificity and overall accuracy. If an impressive average is driven by easy scenarios while a rare critical event is missed, acceptance fails no matter how good the headline looks.
An evaluation should also review what the system suppressed. Sample the filtered activity across cameras, times and conditions to catch potentially important missed events, and agree on the meaning of terms such as trigger, alert, false positive and duplicate before data collection. In some jurisdictions regional testing methodology and terminology can anchor these metrics; where such standards apply, they help make results reproducible and comparable between vendors.
Weighted scoring matrix
Publish a weighted scoring matrix in the RFP so vendors address what actually decides the award. A typical starting set for security is technical fit and accuracy on your footage at 25–35%, total cost of ownership at 15–20%, integration and deployment at 15%, data governance and privacy at 10–15%, SLA and support at 10%, references at 10%, and roadmap at 5%. Adjust the weights to your situation — regulated sectors usually lift governance and explainability.
Give explicit weight to things absent from a demo: written pilot acceptance criteria and a price lock for full deployment after a passing pilot. That prevents the pattern where a vendor wins on evaluation and then renegotiates commercial terms once the data arrives.
Commercial terms, SLA and reference checks
Require line-item pricing across licenses, integration, model retraining, first-year support and renewals. A single blended number hides the levers of total cost. In the SLA, define response times by severity, the cadence and process for retraining when a scene changes or a new event type appears, and what happens when the model drifts from baseline performance.
Treat references as a deployment-record test: ask for live installations in your sector, not marketing case studies, and contact customers with more than a year of production experience. Walk-away signals include guaranteed ROI before anyone inspects your site, refusal to commit acceptance criteria in writing, pricing that excludes integration and data work, and unclear language about whether the vendor may reuse your labeled footage to train models for other customers.
Put it into practice
RFP pre-flight checklist: requirements, questions and acceptance criteria
Use this checklist to review a draft RFP before release. It covers the document sections, the questions every candidate should answer identically, and the items you should fix in the contract before a pilot begins.
- Business problem and measurable success metric defined: events reaching operators, acceptable latency, tolerable false alarms per day.
- Each scenario described with object or event, camera zone, schedule, alarm type and routing.
- Integration requirements stated: VMS, ONVIF/RTSP, API, metadata and export formats.
- Deployment architectures listed (edge, on-premise, cloud) and what may leave the site.
- Capacity, channel count, failover and redundancy requirements written down.
- Storage, retention, access and biometrics handling policy specified.
- Identical accuracy, false-positive, subprocessor and reference questions drafted for all vendors.
- Weighted scoring matrix with shares per criterion published in the RFP.
- Pilot defined in writing: duration, cameras, test scenarios, per-scenario metrics and named reviewers.
- Acceptance criteria, retraining process, SLA and post-pilot price lock captured in the contract.
- Live references in your sector with over a year of production use verified by phone.
- Walk-away red flags recorded: premature ROI guarantees, refusal of written criteria, hidden integration cost, unclear data reuse.
Questions people ask
Which metrics should acceptance criteria use, and why not blend recall and precision into one number?
The core metrics are recall, precision, specificity and overall accuracy. Recall measures the share of real events the system found; precision measures the share of its alerts that were actually warranted. They trade off against each other: a vendor can raise recall by lowering the detection threshold, only to flood operators with false alarms. Set targets per scenario and fix an acceptable daily false-alarm volume. A blended average is misleading because it hides a failure on a rare critical scenario behind strong results on common easy ones. In some regions, testing methodology standards anchor how these figures are computed and reported.
How do I verify a vendor's accuracy claims before buying?
Ask for results from live installations similar to yours in industry, camera type and lighting, and ask under what conditions the numbers were measured; laboratory benchmarks rarely reflect a real scene. The most reliable check is a shadow-mode pilot on your own footage with scenarios, reviewers and per-scenario metrics agreed in writing before it starts. Also probe the live false-positive rate in a crowded scene, because that determines the operational load your team will actually carry.
What should a pilot include to be a fair test?
Agree in writing before it starts on duration (typically two to four weeks), the number and variety of cameras — including hard scenes such as night, crowds, poor lighting and known nuisance-alert sources — the set of test scenarios, named reviewers on both sides, and target metric values per scenario. Review a sample of suppressed activity to find missed events, and add deliberate scenario tests where lawful and safe. State that definitions will not change after results arrive.
What data and privacy requirements belong in an RFP?
Specify where video and metadata are stored, default retention periods, which subprocessors have access, and how the system handles biometrics if used. For regulated sectors, require an on-premise or edge option so footage need not leave the site. Clarify ownership and reuse: whether the vendor may use your labeled footage to train models for other customers. This article is general guidance, not legal advice; check the rules that apply in your jurisdiction with a qualified professional.
How do I compare different vendors objectively?
Publish a weighted scoring matrix before release and score every candidate against the same question set. A typical starting mix for security is technical fit and accuracy on your footage at 25–35%, total cost of ownership at 15–20%, integration at 15%, data governance at 10–15%, SLA at 10%, references at 10% and roadmap at 5%. Compare responses line by line rather than by demo, and use a short targeted addendum to clarify thin sections instead of restarting the whole cycle.
What should I do if a vendor refuses to commit to written acceptance criteria?
Treat the refusal as a serious signal and a reason to deprioritize that candidate. Written pass/fail criteria agreed before a pilot protect both sides and stop the vendor from redefining success after the data arrives. If the criteria are reasonable and the vendor still will not sign them, the commercial risk of proceeding is high. Publish the acceptance framework in the RFP itself, and reference it in the contract before any pilot begins.
Sources and further reading
Sources were checked when this page was generated. Confirm changing dates, rules and prices with the original publisher.
- ГОСТ Р 72563-2026: Ситуационная видеоаналитика. Методология определения значений функциональных характеристикMeganorm.ru (электронный фонд документов)
- ГОСТ Р 59385-2021: ИИ. Ситуационная видеоаналитика. Термины и определенияИнформпроект Групп
- How to Choose an AI Video Analytics Vendor: The Evaluation Checklist That Actually WorksStaqu
- AI Video Analytics for Physical Security: A Buying GuideAmbient.ai
- Don't Trust the AI Demo: Prove Video Monitoring ROI in 14 DaysArcadian.ai
- AI Vision RFP Template & Vendor Scoring Matrix 2026iFactory