Recurring incidents
Every month end, the reporting job fails and someone restarts it. It has been restarted forty times and the ticket is closed each time.
- Why it happens
- Support is measured on closing tickets, not on removing causes, and nobody has time or ownership for the root fix.
- What it costs
- Repeated disruption and a team stuck in reactive mode.
- How we approach it
- We classify incidents, run root cause analysis on the top recurring ones, fix causes, add monitoring and track recurrence as a service metric.
- What to measure
- Incidents per month by cause, repeat incident rate and mean time to restore.