Your backlog is growing. Do not start with a clean-up sprint
A rising ticket count tells you that unfinished work is accumulating. It does not tell you why.
More work may be arriving. Genuine completion may have fallen. Tickets may be blocked by dependencies or returning after an apparent resolution. The increase may even come from a reporting change rather than the operation itself. Several movements can happen at once.
That is why a blanket instruction to close more tickets is risky. It can lower the visible count while increasing reopens and repeat contacts. Hiring will not repair broken routing, and a bot may hide contacts without resolving the underlying problem.
Treat backlog growth as a symptom: validate the signal, locate the movement and segment that changed, test the likely cause, then monitor whether the intervention improves flow without damaging customer outcomes.
What to check first: validate the signal
1. Freeze the denominator
Write down what counts as backlog in this investigation:
- included statuses;
- queues, brands, channels and teams;
- treatment of pending and on-hold work;
- treatment of transferred or externally owned work; and
- reporting period, time zone, and use of calendar or business hours.
Keep that definition fixed during the diagnosis. If you need a second view, such as agent-actionable tickets, report it beside the total rather than silently removing waiting work.
2. Check for measurement changes
Look for a migration, new channel, form change, bot rule, status mapping, merge rule, workflow change or report-filter edit around the point where the trend moved. A time-zone or business-hours change may also shift tickets between periods.
Do not assume familiar metric labels have standard meanings. Zendesk, for example, documents separate unsolved-ticket and reopened-ticket measures; its duration documentation distinguishes requester wait, agent wait and on-hold time. Check the definitions in your own system.
3. Reconcile the movement
For one denominator and reporting period, compare:
closing backlog = opening backlog + work entering - work genuinely completed
This is a diagnostic reconciliation, not a universal platform formula. State how deleted, merged, transferred and reopened tickets are handled, and whether repeat contact arriving as a new ticket is visible as returned demand.
Illustrative calculation: A team starts the week with 420 tickets in its stated backlog. During the week, 310 enter that denominator and 270 genuinely leave it. The expected closing backlog is
420 + 310 - 270 = 460, an increase of 40. If the dashboard shows 445, investigate the missing 15 before diagnosing the operation. Merges, transfers, deletions, reopens or mismatched filters may explain the gap.
Treat “genuinely completed” carefully. A solved count may overstate completion when customers reopen tickets, contact the team again about the same issue or move into an unmeasured escalation route.
4. Find the break point
Compare the growth period with a credible local baseline using the same definitions. Find the first date, queue or segment where the pattern diverged. Where the data is reliable, segment by contact reason, severity, channel, customer group, language, product area, owner group, agent tenure and time of day. A blended total often conceals the change that matters.
What the data may be telling you
Use four movements to organise the diagnosis. They are operational hypotheses, not mutually exclusive categories.
1. More work is entering
Compare arrivals with their normal pattern, then find where the increase sits. Check for concentration around a release, outage, billing event, migration, campaign, policy change, new channel or contact reason.
“Volume is up” is not yet a diagnosis. Establish whether the demand is avoidable, incident-driven or a lasting change in channel use. Stable total arrivals can still hide an increase in difficult or high-severity work.
2. Less work is completing
Compare genuine completion over the same period and for a comparable case mix. Inspect coverage, absence, onboarding, meetings, incident duties, escalation load, tool problems and workflow changes. Unchanged headcount does not mean unchanged capacity when cases become harder.
Do not equate lower throughput with weak agent performance. A team supporting new starters, handling more complex work or compensating for a broken tool may be working appropriately while completing fewer tickets. Separate effective capacity from nominal staffing.
3. More work is blocked
Inspect waiting reason, time in state, current owner, next action and dependency owner. Break out work waiting on Product, Engineering, Billing, Security, a supplier or the customer.
Pending and on-hold tickets may sit outside the agent-action queue, but they remain operational work. If they are hidden, unowned or lack a review date, backlog can grow without an obvious decline in frontline activity. The Kanban Guide treats work-item age, throughput and work in progress as distinct flow measures; one cannot stand in for another.
4. Work is returning after apparent completion
Check reopened tickets, repeat contacts, duplicates, escalations and a sample of quality-assurance findings. A solved-ticket count can rise while customers return because the underlying issue remains unresolved.
Closure pressure and poorly governed automation can therefore create false improvement. Containment, deflection and auto-closure are not equivalent to successful resolution. Check what happened to the customer, not only whether a ticket disappeared from one queue.
Likely causes: match the pattern before acting
Look for evidence that could weaken your preferred explanation as well as evidence that supports it.
| Observed pattern | Causes to test | Evidence that strengthens the hypothesis | Evidence that weakens it |
|---|---|---|---|
| Arrivals rise; completion is stable | Release, outage, campaign, policy, new channel, defect or avoidable demand | Increase clusters around a date, reason, channel or product area | Growth is broad, or arrivals remain within the comparable baseline |
| Arrivals are stable; completion falls | Coverage loss, harder case mix, skill gap, tool or workflow problem, escalation load | Fall begins with a shift in mix, coverage or process | Completion is stable after controlling for case mix |
| Unassigned or reassigned work grows | Routing or ownership failure | Time unassigned and reassignments rise in affected queues | Work is promptly accepted and owned |
| Pending or on-hold work grows | Dependency bottleneck or weak review cadence | Time in state rises; owners or review dates are missing | Waiting work is current, owned and demonstrably progressing |
| Solved count rises; reopens or repeat contacts rise | Premature closure or incomplete resolution | Ticket samples show that the original cause remained unresolved | Reopens are stable and quality checks support genuine resolution |
| Backlog grows in particular periods | Coverage mismatch or variable demand | Arrival patterns and staffed coverage diverge by hour or day | Similar growth appears across periods with different coverage |
Causes can coexist. Prioritise by customer consequence rather than ticket count alone: a small group of severe, breached or commercially sensitive cases may require action before a larger pool of low-consequence work.
Illustrative scenario: A blended dashboard shows an 18% backlog increase, which initially resembles a staffing problem. Segmentation shows that most growth sits in one product area and began immediately after a release. Overall completion is stable, while arrivals for that product area have risen. This does not prove that the release caused the contacts, but it makes product-driven demand a stronger hypothesis than a broad capacity failure. Contact-reason analysis and ticket samples could confirm or disprove it. All figures are illustrative.
Actions to take
Stabilise the next 48 hours
Protect customers while the diagnosis continues:
- surface critical, breached, unassigned and high-consequence work;
- give significant tickets an accountable owner, next action and review time;
- correct obvious routing failures so new work does not remain unowned;
- separate incident-, defect- or campaign-driven demand when it needs a dedicated response; and
- preserve the denominator so a definition change cannot masquerade as recovery.
Service-level agreements can express measurable response and resolution expectations, as Atlassian's SLA guidance explains. Use commitments to expose customer risk, but do not let a target encourage premature closure.
Match the intervention to the cause
| Diagnosed constraint | Appropriate response | Leading signal | Balancing measure |
|---|---|---|---|
| Demand spike | Communicate known issues, correct misleading product or policy messaging, improve the relevant knowledge path, address the root cause and add temporary triage where justified | Incoming work in the implicated segment slows | Repeat contact, escalations and customer comments do not worsen |
| Completion or capacity constraint | Restore lost coverage, protect focus time, repair tools or process, use suitable temporary capacity, and investigate staffing if evidence shows a sustained gap | Genuine completion recovers for a comparable case mix | Quality findings and reopens remain acceptable locally |
| Routing or ownership failure | Repair routing, require acceptance at hand-off, name dependency owners and set review dates | Unassigned time and reassignments fall | Misroutes and missed commitments do not shift elsewhere |
| Changed case mix or skill constraint | Route by skill, pair or coach agents, improve escalation packets, and address the source of complex work | Completion improves in the affected case type | Escalations and QA findings do not deteriorate |
| Hidden rework | Sample reopens and repeat contacts, remove premature-closure incentives and improve diagnostic guidance | Returned work falls | Resolution time is not improved by leaving customers without a viable outcome |
| Stale work | Apply documented follow-up and closure rules, with a clear route to reopen | Owned stale work reduces | Repeat contact and complaints do not rise |
For each intervention, record an owner, expected leading signal, review date and stop or change condition. That makes the action a testable operational decision.
Avoid false fixes
- Do not impose a universal “healthy” backlog target without accounting for ticket mix, severity, channels, support hours, service commitments and customer consequence.
- Do not ask every agent to close more tickets before locating the constraint.
- Do not change statuses or exclusions to improve the dashboard.
- Do not bulk-close old work merely to reduce the count.
- Do not treat bot containment, deflection or auto-closure as resolution without checking the customer outcome.
- Do not move directly from backlog growth to a hiring request. Demand, case mix, coverage, shrinkage, skill distribution, productivity, quality and service expectations all affect the staffing case.
What to monitor afterwards
Choose a review cadence that matches the operation; there is no universal recovery timeline. Track two layers:
- Flow: net backlog change, incoming work, genuinely completed work, unassigned work, waiting-reason mix, age of active work and returned work.
- Customer and quality balancing measures: quality-assurance findings, repeat contact, escalations, missed commitments, customer comments, and customer effort or satisfaction where response volume and response bias are understood.
Retain the segment that implicated the cause. A blended total can improve while one channel, product area, customer group or period deteriorates. Compare like with like when demand mix, severity, agent tenure or coverage changes.
Keep definitions visible beside the chart. The Zendesk Support metrics documentation and its native duration metrics documentation illustrate why status and time measures require precise interpretation. Your platform and operating model may treat them differently.
Do not finish by asking only, “How do we clear it?” Ask: Which movement changed, where did it change, and what evidence would disprove our explanation? The answer should lead to a smaller, safer intervention than a clean-up target applied to the whole queue.

Stephen Wood
Stephen Wood is a customer experience and support operations leader with 20 years of experience leading global CX teams, including roles with Oracle and NICE. At Signals, he focuses on helping organisations improve support performance through clearer operating models, better data, practical automation and responsible AI.
- Customer experience
- Support operations
- Responsible AI
Keep exploring
Continue with practical guidance related to this topic.
Customer Support Backlog: How to Measure, Diagnose and Reduce It
Learn how to define, measure, diagnose and reduce a customer support backlog without hiding risk or lowering resolution quality.
Customer Support Metrics: The Complete Guide for Support Leaders
Learn which customer support metrics matter, how to calculate them, what they hide, and how support leaders can turn them into action.