MTTR is mean time to R, and the R is the problem. It is used for mean time to repair, mean time to recovery, mean time to respond and mean time to resolve, and those four measure different spans of the same incident.
No standards body arbitrates between them, so two teams can both report an MTTR of four hours while measuring things that barely overlap. Before the number means anything, somebody has to say when the clock starts and when it stops.
The second thing worth knowing is that it is a mean, and incident durations are not the kind of data a mean describes well.
- The R is repair, recovery, respond or resolve, depending on the speaker
- Those four measure different spans of the same incident
- No standard decides which one MTTR means
- Always ask when the clock starts and when it stops
- Report the median beside it, because one long outage owns the mean
On this page
The four RsThe four Rs, and what each one measures
What is MTTR? It is the average time a team takes to deal with a failure, and IT operations, security response teams and maintenance organizations all use it. Each of the four versions below is one of the legitimate incident metrics. The trouble is only that they share a name.
| Metric | Clock starts | Clock stops | What it tells you |
|---|---|---|---|
| Mean time to respond | The alert fires | Somebody begins working | Whether alerts reach a human at all |
| Mean time to repair | The fault is detected | The thing is fixed | How fast the fixing is |
| Mean time to recovery | The failure begins | The system is normal again | What the outage cost in time |
| Mean time to resolve | The failure begins | The cause is addressed too | Whether it will happen again |
Read the two columns in the middle. Respond ends where repair begins. Recovery starts earlier than repair, because it counts the time before anyone noticed. Resolve ends later than recovery, because it includes the work that stops a recurrence.
So one incident produces four different durations, all correctly labeled MTTR. A team measuring response reports a number in minutes. A team measuring resolution reports one in days. Neither is wrong, and comparing the two metrics is meaningless.
Mean time to recovery is usually the one a business cares about, because it is the only one measuring what the business experienced: the downtime, or how long systems were not working. The others measure how the team performed inside that window.
The formulaHow to calculate MTTR
Every version of the MTTR formula has the same shape: the total time spent in the chosen span, divided by the number of incidents in the period.
MTTR = total downtime or repair time, divided by the number of incidents
Take a system with four incidents in a month, lasting 30 minutes, 45 minutes, 90 minutes and 3 hours 15 minutes. Total downtime is 360 minutes, so MTTR is 360 divided by 4, or 90 minutes. The median of the same data is 67.5 minutes, the first hint of the problem described further down.
What goes into the total depends on the R.
- Mean time to repair: total repair time, including testing the fix, divided by the number of repairs.
- Mean time to recovery: total downtime, from failure to normal service, divided by the number of incidents.
- Mean time to respond: total time from alert to the start of work, divided by the number of alerts.
- Mean time to resolve: total time from failure to the cause being addressed, divided by the number of incidents.
MTTR vs MTBF, and what the two say about availability
MTTR vs MTBF is not a choice, because the two metrics combine. Availability is MTBF divided by the sum of MTBF and MTTR. A system that runs 500 hours between failures and takes 2 hours to recover is available 500 hours out of 502, or 99.6 percent of the time.
Halve the MTTR to 1 hour and availability rises to 99.8 percent. Double the MTBF to 1,000 hours instead and the result is the same 99.8 percent. On paper the two are equal. Users see half as many interruptions in the second case, which is why both metrics belong in the same report.
The clockThe question that makes the number usable
There is one question, and asking it takes longer than most people expect because the answer is often unclear internally.
When does the clock start? At the failure, at the alert, or at the ticket? These are very different.
A disk that fails at two in the morning, alerts at two, and gets a ticket at eight when somebody arrives has already produced a six hour gap, and which metric absorbs it depends on where the incident management process says the clock starts.
When does it stop? At service restored, at ticket closed, or at root cause addressed? Tickets are frequently closed after the service is back and before anybody has worked out why, which makes repair look excellent and hides whether it will recur.
What is excluded? Time waiting for a customer, time outside business hours, time waiting on a vendor. Every exclusion is defensible and every exclusion makes the average smaller. Metrics with unstated exclusions are not comparable to anything.
Write those three answers down next to the number. A stated definition with a worse-looking figure is more useful than an impressive one nobody can interpret.
The averageWhy the mean is the wrong average
This is the part that is rarely said and matters more than the choice of R.
Incident durations are heavily skewed. Most incidents are short, a few are very long, and the long ones dominate any arithmetic average. Fifty incidents of twenty minutes and one of forty hours produce a mean of about sixty-seven minutes, which describes none of the fifty-one incidents that actually happened.
That has two practical consequences.
The average moves for the wrong reason. A single bad month can double it while nothing about ordinary incidents changed. Somebody then explains a trend that is one system failure.
Improving the mean can mean the wrong work. The fastest way to move it is to shorten the longest incident, which may be correct, or may be one unrepeatable event that eats the improvement budget.
The fix is not to abandon MTTR but to report it alongside the median, which describes a typical incident, and a high percentile such as the 95th, which describes the bad ones. Three numbers instead of one, and the shape of the distribution becomes visible.
Reducing itHow to reduce MTTR
Reducing MTTR starts with knowing which span of the incident is slow. The four clocks help here: they split total downtime into detection, response, repair and follow up, and each span has a different fix.
- Detect faster. Monitoring and alerting that catches a failure when it begins, not when a user calls. Whether alerts reach a person at night is the NOC question.
- Respond faster. A clear on-call rota, escalation rules and alerts routed to the team that owns the system.
- Diagnose faster. Runbooks for known failures, and a record of what depends on what, which is the job of a CMDB.
- Repair faster. Automation for routine fixes, spare parts on hand, and redundancy or failover so service recovers before the repair is finished.
- Stop the repeat. Root cause analysis after the incident, so the same failure does not return as new downtime next month.
MTTR does not reduce downtime by itself. It shows where the time went, and that data tells operations teams which of the five to work on first.
The familyThe rest of the family
MTTR is one of four incident management metrics that share a shape, and the other three are less ambiguous than it is. Two of them are frequently swapped for each other.
| Metric | What it measures | Use it for |
|---|---|---|
| MTTA, mean time to acknowledge | Alert fired to a human picking it up | Whether alerting reaches anybody, especially at night |
| MTTR, mean time to R | Some span of the incident, depending on the R | How the team performed once the system was broken |
| MTBF, mean time between failures | How long a repairable system runs between faults | Reliability of something you fix and keep |
| MTTF, mean time to failure | How long a non-repairable component lasts | Reliability of something you replace rather than fix |
The bottom two are the pair that gets confused, and the distinction is not academic. MTBF is for systems you repair; MTTF is for parts you replace. A server is repaired, so it has a mean time between failures.
A disk inside it is swapped rather than fixed, so what a manufacturer publishes for it is a mean time to failure. Quoting MTBF for a component you throw away, or MTTF for a system you keep, produces a number that cannot be interpreted.
MTTA is the one worth adding first if you only track one more. It isolates the span nobody owns: the gap between an alert firing and a person starting work. That gap is invisible inside any MTTR whose clock starts at the ticket, and it is where the hours go on an overnight incident.
Gaming itHow the number gets gamed
Not always deliberately. Every one of these is something a reasonable person does for a reasonable reason.
Closing early and reopening. The service is back, the ticket closes, the underlying fault returns next week as a new ticket. Two short incidents look better than one long one.
Splitting an incident. One system outage recorded as four component tickets produces four small durations instead of one large one.
Generous exclusions. Waiting on the customer, waiting on a supplier, out of hours. Each is arguable and together they can remove most of the elapsed time.
Alerting later. Fewer alerts, or alerts that fire only once a fault is severe, shorten every metric that starts at the alert. The service got worse and the metrics improved.
The defense is not suspicion, it is definition. If the clock rules are written down and the exclusions are listed, most of this becomes visible rather than impossible. The second defense is a record nobody types.
Timestamps read out of a log store rather than entered into a ticket cannot be tidied afterwards, which is what turns an arguable exclusion into a checkable one.
PitfallsWhere people go wrong
Comparing MTTR between organizations. Without matching definitions of when the time starts and stops, the comparison is arithmetic on unlike quantities. Benchmark metrics published without their clock rules are decoration.
Reporting the mean alone. It is the single least representative summary of a skewed distribution. Add the median and a high percentile.
Treating MTTR as a target for a team. Metrics that become targets get managed, and this one has at least four honest ways to be managed without improving the system at all.
Confusing it with RTO. One is measured and one is promised, and a report that blurs them will eventually be read by somebody holding the contract.
Ignoring the detection gap. If the clock starts at the alert, everything before the alert is invisible, and monitoring gaps improve the metric while making the outage longer.
Chasing it instead of MTBF. Recovering faster is worth less than failing less often, and a system that fails half as often has done more for customers than one that recovers twice as fast.
ComparisonMTTR, MTBF and RTO, side by side
| Criterion | MTTR | MTBF | RTO |
|---|---|---|---|
| What it is | A measurement of the past | A measurement of the past | A target agreed in advance |
| Measures | How long recovery takes | How long between failures | How long you promised |
| Rises when | Fixing gets slower | Things break less often | Never, it is set |
| Owned by | Operations | Operations, or the vendor | The business, in a contract |
| Failing it means | Slower than before | More frequent failures | A broken commitment |
The last column is the one to keep separate. RTO is a promise about the future written into an agreement. MTTR is a description of what happened. Reporting MTTR against an RTO is fair; treating them as the same kind of number is not, and it is a common way for a service report to look better than the service.
FAQFrequently asked questions
What does MTTR stand for?
Mean time to R, where the R is repair, recovery, respond or resolve. Which one is meant depends entirely on the speaker, and no standard settles it.
What is the difference between mean time to repair and mean time to recovery?
Repair usually starts when the fault is detected and ends when it is fixed. Recovery starts when the failure begins and ends when service is normal, so it includes the time before anybody noticed.
What is mean time to resolve?
The longest of the four. It runs from the failure to the point where the underlying cause has been addressed, not just the point where service returned.
How is MTTR calculated?
Total time across the incidents in a period, divided by the number of incidents. The formula is the easy part; deciding what counts as the start and end of each incident is the whole difficulty.
What is a good MTTR?
Not a question that has an answer without the definitions. A four hour figure could be excellent or terrible depending on what it includes, and it cannot be compared across organizations that define the clock differently.
What is the difference between MTTR and MTBF?
MTTR measures how long recovery takes. MTBF, mean time between failures, measures how often a repairable system fails. Reducing failures usually delivers more than recovering faster.
Is MTTR the same as RTO?
No. RTO is a recovery time objective, a target agreed in advance and usually written into a contract. MTTR is a measurement of what actually happened.
Why should I report the median as well?
Because incident durations are skewed. One very long outage among many short ones drags the average somewhere that describes no real incident, while the median describes a typical one.
Can MTTR be gamed?
Easily, and usually without intent. Closing tickets early, splitting one incident into several, and generous exclusions all shorten it. Written clock rules are the defense.
Should MTTR be a team target?
It is risky. There are several honest ways to improve the number without improving the service, and a metric that becomes a target tends to find them.
What starts the clock in most tools?
Commonly the alert or the ticket, because that is what the incident management tool can see. That silently excludes the detection gap, which is often the largest part of a bad incident.
Which of the four should I use?
Mean time to recovery, if the audience is the business, because it measures what the business experienced. Use the others internally to see which part of the window is slow.
What is the MTTR formula?
The MTTR formula is total repair time divided by the number of repairs in the period. Ten hours spent fixing four incidents gives an MTTR of 2.5 hours. State which meaning of the R you are using, because repair, recovery, respond and resolve start and stop the clock at different points.
Keep readingRelated concepts
Read next · Tools Network Management Software, and the Three Questions It Has to Answer Where the alert that starts most of these clocks comes from, and why a detection gap quietly improves the metric. Open this next11 min- Managed IT · 10 min What Backup and Disaster Recovery Actually Buys You Where RTO comes from, which is the promise this measurement gets reported against and is not the same kind of number.
- Managed IT · 9 min What an MSP PSA Actually Is, and Why Replacing One Is Harder Where the tickets and time entries this is calculated from actually live, and why the clock rules are a configuration decision.
- Operations · 10 min The ELK Stack, and the Licensing Question Underneath It Where the numbers behind that measurement are kept, and under which license.
- Security operations · 8 min SOC vs NOC, the Two Operations Centers Compared One of the numbers that defines whether the NOC is doing its job.