The short answer
Start with two layers: a stable classification (type, subcategory, canonical object or address, symptom) and a registration flag saying whether a request is first, repeat, duplicate, or new facts on an old case. Match duplicates by object plus category plus time window, never by exact wording. Merge only true duplicates and keep distinct households apart, then aggregate clean volumes by category, location, and time to expose patterns a noisy inbox hides.
Key takeaways
- Classify every request by type, canonical location or object, and symptom so comparable records stay comparable even as agencies, wards, and systems change over time.
- Distinguish a true duplicate (same author, same object, open case) from a repeat request and from many people reporting one shared problem; each needs different handling.
- Match on object-plus-category-plus-time, never on exact text, because wording varies widely across residents and channels.
- Keep one-problem-many-reporters visible as an issue cluster; merging it into one ticket hides scale and urgency.
- Audit your taxonomy regularly: rename, merge, or retire categories and reconcile them after ward, division, or system changes.
- Watch data quality and equity: complaint-driven systems can be biased, and messy fields bias whatever you infer from them.
Why weak classification hides systemic problems
Every municipality that runs a public feedback channel eventually drowns in near-identical messages: broken lights on the same street, the same pothole reported every rainy week, the same noise complaint from a new resident. Without a shared way to name and count these requests, the inbox inflates with duplicates while the underlying problem stays invisible. Large 311 systems illustrate the scale. A study of New York City 311 records found 210 distinct complaint types, yet the top 20 accounted for roughly 70 percent of all requests, and eight different noise categories alone made up about 22 percent. In Toronto, open 311 data contained more than 850 unique request types, many describing the same issue under different names.
The practical effect is that staff spend effort filing and re-filing while managers cannot tell whether many messages mean one angry person, dozens of households, or one failing asset. That is why classification and duplicate handling are not bookkeeping chores. They are the precondition for seeing which complaints repeat, where they cluster, and which single failures generate cascades of requests.
Build two layers: taxonomy and request status
Treat each record as two independent layers. The first is a stable taxonomy describing what the request is about: a category and subcategory, a canonical object or address, a symptom or requested action, the channel, and the territory or district. The second layer is a status describing its relationship to other records: first request, repeat by the same author, duplicate of an open request, or new facts attached to an old case. Do not mix the two. Many systems fail because they store duplicate as if it were a subject, which destroys every later count.
A good taxonomy stays stable over years even as structures change. Toronto analysts had to harmonize request types after division renames and after the city moved from a 47-ward to a 25-ward model, mapping hundreds of legacy types that described the same work. When you build the taxonomy, choose a single canonical location reference and a single vocabulary, and add an explicit unclassified bucket so you can always see how much still escapes the scheme.
Removing duplicates without losing signal
Remove only true duplicates, meaning the same reporter, the same object or address, the same category, and a case still open within a defined window. Match on object plus category plus time, never on exact wording: residents describe the same hole as a pothole, a hazard, or the thing on Maple. Near-duplicate requests from different households about the same shared problem are not duplicates; they form a cluster, and collapsing them hides how many people are affected.
Agree on merge rules before you merge. Give every merge a human override and log how many records you combined, so the duplicate rate itself becomes a metric of interface quality and responsiveness.
- Merge exact duplicates into the original record and keep one clear audit trail.
- Cluster near-duplicates on the same object or block into an issue counter, not a merged ticket.
- Reopen a closed case only when a new report adds facts or signals that the fix failed.
- Route ambiguous pairs to a reviewer instead of auto-merging them.
- Track the merge rate and watch it drop as you improve self-service and status visibility.
Turn clean records into systemic signals
Once records are clean, the signal appears at the aggregate level. Count requests by category, territory, and time window, then compare with the previous period. A single request to fix one lamp is a task; fifty requests across three blocks about lamps that never stay lit is a systemic failure of maintenance scheduling. The same logic applies to snow removal, graffiti, or housing-code cases that recur on the same assets.
Forecasting work in Toronto shows why this matters. After harmonizing hundreds of request types, analysts could model ward-level demand, and the model explained over 90 percent of variance for the dominant forestry workstream, which represented the large majority of environment-related requests. Extreme weather produced genuine outlier spikes, for example more than fifteen hundred environment requests in a single day after a windstorm, which is a reminder to treat occasional canary events as a separate analytical category from steady chronic complaints.
Bias and data-quality traps to watch
Read the data with two cautions. First, reporting itself is uneven. A city ombudsman in Portland found that complaint-driven property-maintenance enforcement fell disproportionately on homeowners of color in gentrifying neighborhoods and that a substantial share of complaints led to no violation at all. If some groups report more or less than others, volumes are a proxy, not the truth.
Second, the underlying fields are messy. Curation studies of the New York City dataset found duplicate fields, blank values, and inconsistent street names, and removing redundant fields alone trimmed the file by a large share. Cities that publish open data often omit data dictionaries, and service-type names differ across jurisdictions, as seen when one city calls a case abandoned vehicle while another writes vehicle abandoned. Plan cleaning, validation, and reconciliation as permanent work rather than a one-time cleanup, and cross-check request volumes against objective conditions or proactive inspections.
Put it into practice
Quarterly request taxonomy and duplicate audit
Run this checklist every quarter on a sample or the full backlog to keep classification clean, duplicates low, and systemic problems visible. Each item should end in a signed decision by someone on your team.
- Confirm every record carries a type, a canonical object or address, a symptom, and a territory code; export the records that do not.
- Measure the unclassified bucket; if it grows, expand or simplify categories rather than ignoring the residue.
- Run duplicate detection on object plus category plus time window; report the merge rate and show the reviewer queue.
- Separate true duplicates from multi-reporter clusters, and keep cluster counts visible as an issue-severity signal.
- Reconcile categories after any division, ward, or system change by mapping legacy names onto the current taxonomy.
- Have two staff independently classify a sample of 30 to 50 requests to measure inter-rater consistency.
- Compare the top categories month over month and flag any category whose volume moves beyond a set threshold.
- Review the long tail: rename, merge, or retire rare categories so the taxonomy stays usable instead of sprawling.
- Check data-quality fields such as blank addresses and missing closure dates, and fix them at the source before they poison trends.
- Publish a short summary of merged counts and the leading systemic issues for managers and residents.
Questions people ask
What is the difference between a duplicate and a repeat request?
A duplicate is the same person filing the same open request again, usually because they did not see a reply or expected faster service; you merge it into the existing open record. A repeat request is the same author returning about the same object after the case closed, often because the fix did not hold or no answer came; treat it as new information about a failed or unresolved service rather than hide it. The distinction matters because repeats signal quality problems, while duplicates mostly signal communication and interface problems.
When should I merge two requests, and when should I keep them separate?
Merge only exact duplicates: same author, same object or address, same category, and a case still open within your defined window, for example 30 days. Keep separate any two requests that involve different authors even if they concern the same pothole or the same tree, because you need to count how many people or households are affected. If you must group such records, do it in a separate cluster or issue layer so the merge never destroys the underlying counts.
How many categories should a request taxonomy have?
Keep the user-facing taxonomy small enough that a person reliably picks a category and large enough that systemic issues stay distinct, typically a few dozen top-level types with subcategories beneath. Evidence from large systems shows a long tail dominates little volume: a New York study found the top 20 complaint types covered about 70 percent of all requests. Optimize for the categories that actually occur, and retire rarely used types rather than endlessly adding new ones.
How do I avoid bias when reporting itself is uneven?
Treat complaint volume as a proxy, not a census, because willingness to report varies by neighborhood, income, tenure, and language. A Portland ombudsman found complaint-driven enforcement fell disproportionately on homeowners of color and gentrifying areas, and many complaints led to no violation. Cross-check request volumes against objective condition data or proactive inspections, weight for population, and review whether your escalation and enforcement rules concentrate penalties on people who report the least.
Can automation reliably classify and deduplicate resident requests?
Yes for the repetitive bulk, but not alone for judgment calls. Natural-language and fuzzy-matching tools classify many messages and catch near-duplicates quickly, yet accuracy is imperfect; one regional deployment reached roughly 70 percent accuracy in detecting the issue in citizen messages and relied on staff verification before routing. Always keep a human review step for ambiguous records, object boundaries, and systemic flags, and treat automation as a triage layer that improves with clean, well-labeled training data.
Sources and further reading
Sources were checked when this page was generated. Confirm changing dates, rules and prices with the original publisher.
- Федеральный закон от 02.05.2006 № 59-ФЗ «О порядке рассмотрения обращений граждан Российской Федерации» (с изменениями и дополнениями)Система ГАРАНТ
- Сборник методических рекомендаций по работе с обращениями граждан (рабочая группа при Администрации Президента РФ, протокол № 15 от 20.09.2018)Система ГАРАНТ
- Интеллектуальная обработка документов в Правительстве Тюменской областиDirectum
- Guest Post: Forecasting 311 Service Requests to Support a More Proactive TorontoCity of Toronto Open Data Portal
- Principles for Open Data Curation: A Case Study with the New York City 311 Service Request DataarXiv (Tussey, NYC DoITT; Yan, University of Connecticut)
- How Louisville's History of Published Service Calls is Transforming 311Granicus
- Portland Ombudsman says complaint-based rules enforcement most affects diverse and gentrifying neighborhoodsOPB (Oregon Public Broadcasting)