PONOPT FIELD NOTES · Бизнес-стратегия и asset management

A Stage-Gate for Scaling Pilots: Evidence Required Before Each New Site

A stage-gate playbook for scaling a pilot to new sites: the evidence, thresholds, and no-go signals each location must clear before it opens.

A stage-gate turns scaling a pilot to new sites into a sequence of evidence-checked decisions rather than one leap of faith. Before each new location opens, a cross-functional gate must confirm that the effect is reproducible beyond the original site, unit economics still hold at real scale, data, integrations, and processes are ready, and risks and operating rules are institutionalized. Pass a gate only on evidence gathered against thresholds set in advance, never on momentum or retrospective rationalization.

Key takeaways

  • A successful pilot proves an outcome in one controlled context; it does not prove systemic reproducibility. Separate transfer evidence from local evidence.
  • Stage-gate logic (after Robert G. Cooper) applies between sites: each new location or tranche of sites is its own gate with pre-agreed criteria.
  • Before each new site, assemble evidence packs: reproducibility, economics at scale, data and integration readiness, and governance/institutionalization.
  • Raise the evidence bar as commitments become less reversible: a decision that cannot be undone should be harder to reach than one that can.
  • No-go signals — manual data fixes, excluded exception paths, hand-held users, fragile integrations, no support owner — mean the location is still a pilot.
  • Scaling is enabled by institutionalization: SLAs, reporting formats, acceptance rules, named support owners, and a workable rollback plan.
  • Codify a repeatable playbook and hold a resource buffer of roughly 20–30% before declaring scale.

Why one successful pilot is not a licence to scale

The gap between an invention and an innovation is easiest to see when a pilot ends. A pilot proves that a solution can deliver an outcome inside one controlled environment; it does not prove the same outcome will appear at the second, fifth, or twentieth site. By the standards used in innovation management and asset-management practice, an innovation exists only when it is embedded in processes and produces a repeatable effect. A single-location result remains local evidence; until it has been shown to travel, it is not transfer evidence.

Scaling fails most often not because growth is impossible but because evidence is misread. Teams treat a successful pilot as permission to commit broadly — a pattern sometimes called evidence overreach. A pilot can look strong in one line, one shift, or one product family and still collapse in production when master data is incomplete, exception paths were excluded, users were hand-held by the project team, or integrations were manually patched. The remedy is to treat each new site as its own test of whether the original result generalizes.

  • Transfer evidence: outcomes that replicate across contexts, teams, and operating conditions.
  • Local evidence: results that depend on a specific team, site, or one-off configuration.
  • Ask before every gate: what about the pilot would not survive contact with a new site?

How stage-gate thinking applies between sites

The stage-gate model, introduced by Robert G. Cooper for new-product development, divides a project into stages separated by gates. At each gate, cross-functional reviewers apply go/kill/hold criteria against a scorecard before the next tranche of resources is released. Its value is discipline: decisions are made on evidence, not momentum, and weak bets are killed early while stopping them is still cheap. The same logic transfers naturally to site rollout — instead of one all-or-nothing launch decision, you install a sequence of gates, one per location or per tranche of sites.

In highly regulated environments — transport, utilities, manufacturing, retail networks — gate criteria must go beyond commercial promise. Infrastructure practitioners who adapt stage-gate add explicit readiness criteria for data, analytics, processes, and safety alongside the business case. A site can be technically ready yet organizationally and institutionally unprepared: responsibilities unclear, contracts absent, acceptance rules unwritten. Each site gate should therefore review four planes: evidence of effect, economics, the operational and data foundation, and the governance or institutional base.

Keep gates light but real. For low-risk digital changes a rigid multi-stage process adds friction; for capital-heavy or regulated rollouts the gate is the cheapest place to stop a mistake. A sensible default is fast, hypothesis-driven iteration inside the preparation stage, with strict gates reserved for the moment a new site commits real capacity and budget.

The evidence packs each new site must clear

Define, before the pilot ends, what a new site must prove to pass. Government guidance on converting pilots into production repeatedly stresses that success criteria must be quantifiable, tied to a dashboard, and set before results arrive, so go/no-go reviews are objective rather than retrospective rationalizations. In practice, five evidence packs cover most cases.

Reproducibility and performance. The model should show stable outcomes across shifts, users, and normal variability — repeatable cycle time, error rate, latency, uptime — not a single good week. If the result was achieved only through project-team heroics or vendor engineering on site, it is not yet reproducible.

Data and integration readiness. Completeness of master data, roles, document versions, and records matters more than teams expect; many pilots fail at scale because data cleanup was deferred. Interfaces to surrounding systems must be tested under realistic load and failure, because manual workarounds that were tolerable in a pilot become bottlenecks in production.

Economics at scale and governance. The business case must survive production overhead: support, validation, licensing, integration maintenance, and rollout labor. In parallel, move the operating agreement onto durable rails — SLAs, reporting formats, acceptance rules, named support owners, and a workable rollback — so the innovation shifts from project mode into regular operations. That institutionalization package is what turns a working pilot into a working norm.

Designing the gate: thresholds, scorecards, and no-go signals

A gate review needs three inputs: a completed outputs package, a weighted scorecard agreed in advance, and a cross-functional panel that includes operations, finance, and the site's own leadership. Scores should be checked against thresholds fixed before results are known, not against a softened target invented afterward. Reviewers need the authority and the habit to issue a genuine hold or no-go, because a gate that rubber-stamps progress stops predicting anything.

Watch for no-go signals that indicate a site is not ready: the rollout depends on manual data fixes or project-team intervention that will not exist in steady state; exception paths such as rework, deviations, or nonconformance were excluded; core integrations remain fragile or unreconciled; validation evidence is incomplete; no clear owner exists for support and configuration; or KPIs improved only because measurement conditions were unrepresentative.

A useful practical rule: if the implementation team stepped away for thirty days, could site operations, quality, and IT run the process, support users, manage changes, and defend the records? If not, the location is still in pilot, however green the dashboard looks. Raise the evidence bar as commitments become less reversible; a decision that can no longer be undone should be harder to reach than one that can. This is general management guidance, not legal or financial advice, and exact thresholds are site- and jurisdiction-specific.

Turning a project into an operating norm

Scaling across sites becomes affordable only when the second site stops being a fresh project and becomes execution of a codified playbook. Capture learnings, resource needs, and delivery processes as templates; record pricing logic, onboarding, escalation, and change-management procedures; and monitor post-rollout with simple red-amber-green dashboards that surface issues early. Lessons from site N should be fed back into the package for site N+1 through loops between engineering, operations, and the new locations.

Replication also needs honest capacity planning. Organizations that scale before production and support can absorb demand damage reputation and customer trust. Practitioners commonly recommend holding a buffer on the order of 20–30% of resources and budget, because a broad rollout cannot always be slowed or stopped mid-flight. And because current metrics and rules optimized for the existing operating model can systematically suppress a new solution, design acceptance criteria, contract forms, and reporting with an eye to procurement and compliance realities rather than bolting them on at the last gate.

New-Site Evidence Gate Scorecard

Complete this scorecard before opening each new site and review it at the gate. Every line is a question to answer with facts and artifacts, not intentions. If three or more lines come back negative or unverified, the default decision is to defer the opening until the evidence is produced.

  1. Reproducibility: stable performance across shifts and teams for at least 4–6 weeks with no project-team heroics.
  2. Transfer vs. local: the pilot factors unlikely to transfer are named, with a detection plan for the new site.
  3. Data readiness: master data, roles, document versions, and records are complete; no manual corrections in steady state.
  4. Integrations tested under realistic load and failure; no manual patches in the main run path.
  5. Economics: unit economics hold after support, licensing, validation, and rollout overhead; ≥20–30% buffer held.
  6. Governance: SLAs, reporting formats, acceptance rules, and the risk register are updated for the new site.
  7. Ownership: named owners for support, master-data stewardship, incidents, changes, and rollback.
  8. Compliance: licenses, security and access controls, and audit trails pass review for the new site and jurisdiction.
  9. Change management: training done, role-based access works, and the escalation path is agreed with site leadership.
  10. Gate discipline: scorecard thresholds were set before results; a genuine go / hold / no-go is recorded.
  11. Playbook updated: lessons from prior sites are folded into this site's package.
  12. Recovery: a workable rollback and business-continuity plan has been tested.

Questions people ask

How many sites should a pilot run on before we trust the model at scale?

There is no fixed number, but a useful benchmark is two to three sites in genuinely different conditions (different teams, volumes, geography) to separate transferable effect from local luck. The first site typically tests the hypothesis, the second tests reproducibility without project-team heroics, and the third tests stability under routine support. Open the next site only after the previous one has passed the gate on all evidence packs, not merely because the calendar says so.

What is the difference between transferable and site-specific evidence, and why does it matter?

Transferable evidence shows an outcome repeats across contexts — other teams, locations, load conditions. Site-specific evidence shows success only in the original setting and may depend on a particular team, a one-off configuration, or manual hand-holding. Confusing the two is a leading cause of premature scaling, because a local win is mistaken for systemic proof. Before each new site, ask what would not transfer from the pilot and how you would detect that early.

What red flags mean a pilot is not ready for the next site?

Key signals include: rollout depends on manual data fixes or project-team intervention that will not exist in steady state; exception paths (rework, deviations, nonconformance) were excluded from testing; core integrations remain fragile or unreconciled; validation and change-control evidence is incomplete; there is no named owner for support and configuration; KPIs improved only because measurement conditions were unrepresentative; or there is no workable rollback. A strong practical test: if the team stepped away for thirty days and the site could not run the process itself, it is still a pilot.

Who should sit on each site's gate review, and what should they check?

Assemble a cross-functional panel: operations and quality, finance, IT or the digital function, and leadership from the site itself. Review the evidence packs rather than a general sense of success — reproducibility, economics under real load, data and integration readiness, and the governance base (SLAs, reporting, acceptance rules, owners, rollback). Scores should be checked against thresholds fixed in advance, and the go/hold/no-go outcome should be documented.

How do we keep every new site from being reinvented as a fresh project?

Codify a repeatability playbook: capture lessons, resource needs, delivery processes, onboarding, escalation, and change management as templates. Each post-pilot review should feed the package for the next site through loops between engineering, operations, and the new locations. In parallel, move the operating agreement onto durable rails — SLAs, reporting formats, acceptance rules — so that the second site is execution of a document rather than a new project built from scratch.

What if unit economics look fine at the pilot but degrade at the next site?

That is a classic sign that the pilot economics did not include production overhead: support, licensing, validation, integration maintenance, rollout labor, and differences in site conditions. Return to the gate and separate what actually degraded — cost to acquire or deliver, margin, throughput, or reproducibility. If the degradation is systemic rather than a one-off, do not expand further: fix the economic model, build in a buffer, and re-validate on one more site before rolling out broadly.

When is it right to skip a pilot and deploy directly?

When the evidence already exists: the model is proven in comparable conditions, the integration path is understood, governance and compliance approvals are in hand, infrastructure is ready, and change-management resources are committed. If all those conditions are met and you still pilot, you are avoiding a decision rather than managing risk. Perpetual piloting consumes resources, delays value, and erodes confidence; the discipline is to name the minimum evidence for deployment and act on it once it exists.

Sources and further reading

Sources were checked when this page was generated. Confirm changing dates, rules and prices with the original publisher.

  1. Track 3: Pilot to Production ChecklistUK Government (GOV.UK)
  2. Механизм масштабирования инноваций в метрополитене: адаптация STAGE-GATEКиберЛенинка (статья Ф. М. Александровского)
  3. Stage-gate process: How gates and stages drive NPDNetguru
  4. AI Pilot vs Full Deployment: The Scaling DecisionAI Advisory Practice
  5. Developing a Scaling Strategy (MicroCanvas Framework)MicroCanvas
  6. What are the key go/no-go criteria to move from pilot to production?Connect 981
  7. Как стартапу понять, что он готов к масштабированиюRusbase (RB.RU)