The short answer
A shared KPI dictionary is a governed catalog where every metric used across your sites has one approved definition, formula, data source, unit, granularity, owner, and refresh rule. Build it before benchmarking: centralize a small non-negotiable core, let each site keep clearly labeled local metrics, and bind every report to the dictionary so that a term like “downtime” means the same thing in every location.
Key takeaways
- Comparability fails at the definition layer, not the formula layer: two sites can use the same-named metric and still be non-comparable when denominators, exclusions, time windows, or source systems differ.
- A dictionary record is more than a name and one-line description: it needs a business definition, formula, inclusions and exclusions, grain, unit, source, refresh rule, owner, certification status, and effective date.
- Standardize a small non-negotiable core of about 10–15 KPIs network-wide and let sites keep clearly labeled local operational metrics outside cross-site rankings.
- Govern definitions through change control and versioning: formula changes need the governance forum, and versioned history separates a metric trend change from a definition change.
- Roll out by pilot site, and gate each KPI on a traceable data source and identical recalculation before it earns a place in the network scorecard.
- Compare like with like: benchmark sites against same-type peers, their own history, and the plan, never against a network-wide average that hides structure.
- Declare the data state (provisional vs settled) in every report so incomplete figures are never ranked against reconciled periods.
Same name, different meaning
In any multi-site network the problem usually starts with interpretation, not arithmetic. One site includes delivery fees in revenue, another does not; one counts scheduled maintenance as downtime, another excludes it; one logs micro-stops, another rounds them away. When every location keeps its own convenient meaning for a word, the group dashboard fills with numbers that cannot be compared, and leadership eventually stops trusting all of them.
A metric that shares a name can be answering different questions. Comparing sites directly makes definition differences look like performance differences: the site with the “bad” number may simply count more strictly. The larger the network, the more this mistake costs, because rankings, bonuses, and replication decisions are built on those figures.
For manufacturing there is a useful reference point: ISO 22400, an industry-neutral standard that presents KPIs through their formula, time behavior, unit, and user group, so indicators can be defined, composed, and exchanged consistently across sites. Retail and logistics have no universal ready-made dictionary, so every chain builds its own governed catalog using the same underlying principles.
Anatomy of a dictionary record
A serviceable dictionary record is not a paragraph of explanation but a structured set of fields that both people and systems can read. Beyond a plain-language business definition, the record needs the formula or calculation rule, explicit inclusions and exclusions (what is counted and what is discarded), the denominator, granularity (site, day, shift, product), the unit and dimension, and data sources traced to system, table, and column.
Responsibility is recorded separately: the owner accountable for the business result, the hands-on steward people contact when the number is wrong, the refresh cadence, and a certification status that moves from draft to verified to deprecated. The rule “one name, one formula” is respected consciously: if the unit or the audience changes, you create a separate explicitly named term instead of editing a shared one, because collapsing different formulas into one entry pollutes the data layer and breaks lineage.
Add synonyms and a version with an effective date. Synonyms let people find a metric under the different working names regions use, while versioning ensures history is never silently rewritten when a formula changes.
- Plain-language business definition plus the full formula and its elements.
- Inclusions, exclusions, and an explicit denominator.
- Granularity: site, date, shift, product, channel.
- Unit and dimension; a change of unit becomes a separate term.
- Sources to system, table, and column, with who pulls the data.
- Refresh cadence and data state (provisional or settled).
- Owner, steward, and certification status: draft, verified, deprecated.
- Version, effective date, and synonyms.
Central core, local margin
Standardization does not mean every site must compute dozens of metrics identically. A pragmatic model is “core plus margin”: a small group of metrics that will be compared network-wide is defined once, while a site’s operational metrics stay local and never enter cross-site rankings without clear labeling.
Keep the mandatory set to roughly 10–15 KPIs tied to strategy: on-time delivery or OTIF, operational accuracy, profitability, equipment reliability. For each, the dictionary holds a single approved definition. Local metrics are allowed but must be labeled, linked to the canonical enterprise term when meaning overlaps, and excluded from ranking unless additional conditions are met.
A workable split is about eighty percent of metrics in the shared template and twenty percent open to regional extensions. Extensions carry their own definition, data source, and expiry so a temporary metric does not quietly become permanent. No one changes a core definition unilaterally.
Governance and versioning
A dictionary stays alive only when it has an owner and a change process. The business owner is accountable for the result, the steward keeps the data healthy and is the first contact for discrepancies, and a cross-functional committee chartered by operations leadership approves structural changes. This prevents any single function from owning definitions unilaterally.
Calibrate change control to the type of edit: a typo needs no approval, a wording change requires the term owner, and a change to formula or scope goes to the governance forum. Review cadence follows importance: core KPIs are re-reviewed quarterly, standard terms yearly, and deprecated terms are archived rather than deleted.
Versioning is non-negotiable. Every change records its date, reason, affected scope, and whether history must be recalculated. Without it, you cannot tell whether a trend moved because sites improved or because the formula changed.
Roll out by pilot, then hold a cadence
A dictionary and its standard should not switch on across the whole network at once. Prove the definitions and tooling on one pilot site, capture the lessons, then replicate site by site using a repeatable playbook. The first site carries the cost of discovery; every later site inherits a finished playbook and reaches a trustworthy baseline faster than the one before it.
Assign roles by level: the site owns its metrics and improvement actions, the region rolls them up, and headquarters owns definitions, cadence, and targets. A short weekly review at site level and a monthly comparison at network level are usually enough, provided the data feeding them is automatic and trustworthy.
Build network-level reporting from day one rather than bolting it on later. When headquarters can compare sites on the same definition from the first location, the standard holds, best practice from a strong site transfers to a weaker one, and the network improves as a network instead of as disconnected projects.
Audit before you benchmark
A correct formula does not mean the figures are comparable. Two sites can compute “average ticket” the same way while one includes delivery fees and the other does not, or one counts orders by payment and the other by fulfillment. Before comparing, check four things: the same accounting object, the same time window, the same data state, and the same exception-handling rules.
If any answer is no, mark the metric as not directly comparable and show the reason instead of a rank. Also separate data states in reports: operational figures can be shown immediately, but conclusions about profit and repeat purchase wait until reconciliation and refunds are closed, otherwise you rank unfinished data.
A network-wide average is misleading: five sites growing forty percent, eight growing ten, and seven shrinking eight still average to a calm twelve percent. Compare using the median for a typical site, quartile intervals, and the share of anomalous sites, and choose a base from same-type peers, the site’s own history, and the plan. Only when those three baselines point the same way should you draw a confident conclusion.
Put it into practice
Pre-launch audit: dictionary records and comparability gate
Run this checklist for every metric before it enters the network scorecard or a ranking. If any item fails, either fix the record or keep the metric out of direct cross-site comparison until it passes.
- The metric has an accountable owner, a hands-on steward, and a certification status (draft, verified, deprecated).
- The record contains a plain-language business definition and the full formula with its elements.
- Inclusions, exclusions, and the denominator are explicit: what is counted and what is discarded.
- Granularity goes down to the level that locates problems: site, date, shift, product, channel.
- Unit and dimension are recorded; a change of unit creates a separate named term.
- Sources are traced to system, table, and column, with who is responsible for extracting data.
- Refresh cadence and data state (provisional vs settled) are declared for the metric.
- A version and effective date are set, and a definition change never silently rewrites history.
- A small core of 10–15 KPIs is designated non-negotiable and identical across all sites.
- Local operational metrics are clearly labeled and excluded from network-wide rankings.
- The comparability gate passes: same accounting object, time window, data state, and exception handling across sites.
- A pilot site ran one full reconciled period before rollout, and all definition changes are logged and approved.
Questions people ask
Who should own and maintain the shared KPI dictionary?
No single function should own definitions unilaterally. Put a hands-on steward, the person users contact when a number is wrong, in the day-to-day role, and place program accountability with a cross-functional committee chartered by operations leadership. The business owner is accountable for the result, the steward keeps records current, and the committee approves formula or scope changes. Without this separation the dictionary decays quickly because day-to-day maintenance has no clear owner. In small networks one analyst can combine steward and some owner duties, but approval of formulas must stay with the governance body.
How many KPIs should be standardized across all sites?
A practical range is about 10–15 core KPIs tied directly to strategy: on-time delivery, operational accuracy, profitability, and equipment reliability. This does not mean sites cannot keep their own operational metrics; they can, but those must be labeled local and excluded from network rankings. The more metrics are made mandatory, the lower the discipline and the more likely sites will compute them differently in practice. Start with a small core, stabilize definitions on a pilot, and only then expand the set.
How can I detect that two sites use different definitions for the same-named KPI?
Check the underlying data, not the formula. Take one identical operation and period at two sites and trace the path from source system to the final figure. Compare four parameters: the accounting object (for example, whether delivery is included), the time window, the data state (provisional or settled), and exception handling (returns, scheduled maintenance, test orders). Also ask each site to describe the metric in its own words; wording differences almost always reveal definition differences that the final number hides.
What should I do when a site legitimately needs a different calculation?
Separate the mandatory core from the local margin. If a site wants a different calculation for a metric that is not in the network rankings, allow a local metric provided it is documented as a separate term with a definition, source, and expiry. If the metric is part of the core, do not change the definition to fit the site. Instead, create an explicitly named variant, such as “revenue including delivery” and “revenue excluding delivery,” and link both to the canonical enterprise term. Direct comparison across sites is only valid on the same version of a definition.
How do I keep history comparable when a definition changes?
Treat any change to a formula or scope as a new version with an effective date, a stated reason, and an affected scope, then archive rather than delete the old version. If the new formula makes past data incomparable, decide whether to recalculate history under one rule and log that decision in the version journal. This lets you distinguish a trend change caused by site performance from one caused by a formula change. Without versioning you cannot tell what changed, and any before-and-after comparison loses meaning.
Do local, site-specific KPIs break network benchmarking?
No, if they are managed correctly. Local operational metrics are useful precisely because a fully standardized metric may not reflect local operational reality. The rule is to keep the standardized enterprise core separate from locally useful metrics and govern how they relate: label local KPIs, link them to canonical enterprise terms when meaning overlaps, and exclude them from cross-site rankings. The dashboard can show both global and local KPIs at once, but only standardized metrics should feed network-level comparisons and rankings.
Why is a network-wide average a dangerous basis for comparing sites?
An average hides structure. Five sites growing forty percent, eight growing ten, and seven shrinking eight still average to a calm twelve percent even though the network is highly uneven. Averages suit estimates of overall scale but not ranking or conclusions about an individual site. Use the median for a typical site, quartile intervals, and the share of anomalous sites, and choose a base from same-type peers, the site's own history, and the plan. When those baselines disagree, explain site structure before rewarding or penalizing anyone.
Sources and further reading
Sources were checked when this page was generated. Confirm changing dates, rules and prices with the original publisher.
- ISO/DIS 22400-2 — Automation systems and integration: KPIs for manufacturing operations management, Part 2ISO (International Organization for Standardization)
- Structure your glossary, taxonomy, and metrics catalogAtlan
- Step 3: Create key metrics and glossary termsCoalesce
- Multi-Site OEE Rollout PlaybookTeepTrak
- How to Standardize Operations Across Multiple Warehouses: Processes, KPIs & ToolsCleverence