Systems · Concept · 9 min read

MTTR, and Why Two Teams Quoting the Same Number Disagree

Four legitimate metrics share this acronym and no standard decides between them, so the figure matters less than the definition attached to it.

Written by Marko Ristic, Editor Updated Sep 17, 2026
4Things the R stands for, all of them legitimate and all different
0Standards bodies that decide which one the acronym means
3Numbers to report instead of one: mean, median and 95th percentile
MTTAThe span nobody owns, from the alert firing to a person starting
Short answer

MTTR is mean time to R, and the R is the problem. It is used for mean time to repair, mean time to recovery, mean time to respond and mean time to resolve, and those four measure different spans of the same incident.

No standards body arbitrates between them, so two teams can both report an MTTR of four hours while measuring things that barely overlap. Before the number means anything, somebody has to say when the clock starts and when it stops.

The second thing worth knowing is that it is a mean, and incident durations are not the kind of data a mean describes well.

  • The R is repair, recovery, respond or resolve, depending on the speaker
  • Those four measure different spans of the same incident
  • No standard decides which one MTTR means
  • Always ask when the clock starts and when it stops
  • Report the median beside it, because one long outage owns the mean
On this page

The four RsThe four Rs, and what each one measures

What is MTTR? It is the average time a team takes to deal with a failure, and IT operations, security response teams and maintenance organizations all use it. Each of the four versions below is one of the legitimate incident metrics. The trouble is only that they share a name.

MetricClock startsClock stopsWhat it tells you
Mean time to respondThe alert firesSomebody begins workingWhether alerts reach a human at all
Mean time to repairThe fault is detectedThe thing is fixedHow fast the fixing is
Mean time to recoveryThe failure beginsThe system is normal againWhat the outage cost in time
Mean time to resolveThe failure beginsThe cause is addressed tooWhether it will happen again

Read the two columns in the middle. Respond ends where repair begins. Recovery starts earlier than repair, because it counts the time before anyone noticed. Resolve ends later than recovery, because it includes the work that stops a recurrence.

So one incident produces four different durations, all correctly labeled MTTR. A team measuring response reports a number in minutes. A team measuring resolution reports one in days. Neither is wrong, and comparing the two metrics is meaningless.

Mean time to recovery is usually the one a business cares about, because it is the only one measuring what the business experienced: the downtime, or how long systems were not working. The others measure how the team performed inside that window.

The formulaHow to calculate MTTR

Every version of the MTTR formula has the same shape: the total time spent in the chosen span, divided by the number of incidents in the period.

MTTR = total downtime or repair time, divided by the number of incidents

Take a system with four incidents in a month, lasting 30 minutes, 45 minutes, 90 minutes and 3 hours 15 minutes. Total downtime is 360 minutes, so MTTR is 360 divided by 4, or 90 minutes. The median of the same data is 67.5 minutes, the first hint of the problem described further down.

What goes into the total depends on the R.

  • Mean time to repair: total repair time, including testing the fix, divided by the number of repairs.
  • Mean time to recovery: total downtime, from failure to normal service, divided by the number of incidents.
  • Mean time to respond: total time from alert to the start of work, divided by the number of alerts.
  • Mean time to resolve: total time from failure to the cause being addressed, divided by the number of incidents.

MTTR vs MTBF, and what the two say about availability

MTTR vs MTBF is not a choice, because the two metrics combine. Availability is MTBF divided by the sum of MTBF and MTTR. A system that runs 500 hours between failures and takes 2 hours to recover is available 500 hours out of 502, or 99.6 percent of the time.

Halve the MTTR to 1 hour and availability rises to 99.8 percent. Double the MTBF to 1,000 hours instead and the result is the same 99.8 percent. On paper the two are equal. Users see half as many interruptions in the second case, which is why both metrics belong in the same report.

The clockThe question that makes the number usable

There is one question, and asking it takes longer than most people expect because the answer is often unclear internally.

When does the clock start? At the failure, at the alert, or at the ticket? These are very different.

A disk that fails at two in the morning, alerts at two, and gets a ticket at eight when somebody arrives has already produced a six hour gap, and which metric absorbs it depends on where the incident management process says the clock starts.

When does it stop? At service restored, at ticket closed, or at root cause addressed? Tickets are frequently closed after the service is back and before anybody has worked out why, which makes repair look excellent and hides whether it will recur.

What is excluded? Time waiting for a customer, time outside business hours, time waiting on a vendor. Every exclusion is defensible and every exclusion makes the average smaller. Metrics with unstated exclusions are not comparable to anything.

Write those three answers down next to the number. A stated definition with a worse-looking figure is more useful than an impressive one nobody can interpret.

The averageWhy the mean is the wrong average

This is the part that is rarely said and matters more than the choice of R.

Incident durations are heavily skewed. Most incidents are short, a few are very long, and the long ones dominate any arithmetic average. Fifty incidents of twenty minutes and one of forty hours produce a mean of about sixty-seven minutes, which describes none of the fifty-one incidents that actually happened.

That has two practical consequences.

The average moves for the wrong reason. A single bad month can double it while nothing about ordinary incidents changed. Somebody then explains a trend that is one system failure.

Improving the mean can mean the wrong work. The fastest way to move it is to shorten the longest incident, which may be correct, or may be one unrepeatable event that eats the improvement budget.

The fix is not to abandon MTTR but to report it alongside the median, which describes a typical incident, and a high percentile such as the 95th, which describes the bad ones. Three numbers instead of one, and the shape of the distribution becomes visible.

Reducing itHow to reduce MTTR

Reducing MTTR starts with knowing which span of the incident is slow. The four clocks help here: they split total downtime into detection, response, repair and follow up, and each span has a different fix.

  • Detect faster. Monitoring and alerting that catches a failure when it begins, not when a user calls. Whether alerts reach a person at night is the NOC question.
  • Respond faster. A clear on-call rota, escalation rules and alerts routed to the team that owns the system.
  • Diagnose faster. Runbooks for known failures, and a record of what depends on what, which is the job of a CMDB.
  • Repair faster. Automation for routine fixes, spare parts on hand, and redundancy or failover so service recovers before the repair is finished.
  • Stop the repeat. Root cause analysis after the incident, so the same failure does not return as new downtime next month.

MTTR does not reduce downtime by itself. It shows where the time went, and that data tells operations teams which of the five to work on first.

The familyThe rest of the family

MTTR is one of four incident management metrics that share a shape, and the other three are less ambiguous than it is. Two of them are frequently swapped for each other.

MetricWhat it measuresUse it for
MTTA, mean time to acknowledgeAlert fired to a human picking it upWhether alerting reaches anybody, especially at night
MTTR, mean time to RSome span of the incident, depending on the RHow the team performed once the system was broken
MTBF, mean time between failuresHow long a repairable system runs between faultsReliability of something you fix and keep
MTTF, mean time to failureHow long a non-repairable component lastsReliability of something you replace rather than fix

The bottom two are the pair that gets confused, and the distinction is not academic. MTBF is for systems you repair; MTTF is for parts you replace. A server is repaired, so it has a mean time between failures.

A disk inside it is swapped rather than fixed, so what a manufacturer publishes for it is a mean time to failure. Quoting MTBF for a component you throw away, or MTTF for a system you keep, produces a number that cannot be interpreted.

MTTA is the one worth adding first if you only track one more. It isolates the span nobody owns: the gap between an alert firing and a person starting work. That gap is invisible inside any MTTR whose clock starts at the ticket, and it is where the hours go on an overnight incident.

Gaming itHow the number gets gamed

Not always deliberately. Every one of these is something a reasonable person does for a reasonable reason.

Closing early and reopening. The service is back, the ticket closes, the underlying fault returns next week as a new ticket. Two short incidents look better than one long one.

Splitting an incident. One system outage recorded as four component tickets produces four small durations instead of one large one.

Generous exclusions. Waiting on the customer, waiting on a supplier, out of hours. Each is arguable and together they can remove most of the elapsed time.

Alerting later. Fewer alerts, or alerts that fire only once a fault is severe, shorten every metric that starts at the alert. The service got worse and the metrics improved.

The defense is not suspicion, it is definition. If the clock rules are written down and the exclusions are listed, most of this becomes visible rather than impossible. The second defense is a record nobody types.

Timestamps read out of a log store rather than entered into a ticket cannot be tidied afterwards, which is what turns an arguable exclusion into a checkable one.

PitfallsWhere people go wrong

Comparing MTTR between organizations. Without matching definitions of when the time starts and stops, the comparison is arithmetic on unlike quantities. Benchmark metrics published without their clock rules are decoration.

Reporting the mean alone. It is the single least representative summary of a skewed distribution. Add the median and a high percentile.

Treating MTTR as a target for a team. Metrics that become targets get managed, and this one has at least four honest ways to be managed without improving the system at all.

Confusing it with RTO. One is measured and one is promised, and a report that blurs them will eventually be read by somebody holding the contract.

Ignoring the detection gap. If the clock starts at the alert, everything before the alert is invisible, and monitoring gaps improve the metric while making the outage longer.

Chasing it instead of MTBF. Recovering faster is worth less than failing less often, and a system that fails half as often has done more for customers than one that recovers twice as fast.

ONE INCIDENT, FOUR NUMBERS, ALL CALLED MTTRThe bars below are the same event measured four legitimate ways.failurealertwork startsrestoredcause fixedmean time to respondmean time to repairmean time to recoverymean time to resolvethis row is the one the business experiencedNobody arbitrates which of the four MTTR means.So the number is unusable until somebody says where the clock starts and stops.And report the median beside it, because one long outage owns the average.
Four bars, one incident, four legitimate answers. A definition list cannot show that the spans differ; a timeline can, and the difference is the whole point.

ComparisonMTTR, MTBF and RTO, side by side

CriterionMTTRMTBFRTO
What it isA measurement of the pastA measurement of the pastA target agreed in advance
MeasuresHow long recovery takesHow long between failuresHow long you promised
Rises whenFixing gets slowerThings break less oftenNever, it is set
Owned byOperationsOperations, or the vendorThe business, in a contract
Failing it meansSlower than beforeMore frequent failuresA broken commitment

The last column is the one to keep separate. RTO is a promise about the future written into an agreement. MTTR is a description of what happened. Reporting MTTR against an RTO is fair; treating them as the same kind of number is not, and it is a common way for a service report to look better than the service.

FAQFrequently asked questions

What does MTTR stand for?

Mean time to R, where the R is repair, recovery, respond or resolve. Which one is meant depends entirely on the speaker, and no standard settles it.

What is the difference between mean time to repair and mean time to recovery?

Repair usually starts when the fault is detected and ends when it is fixed. Recovery starts when the failure begins and ends when service is normal, so it includes the time before anybody noticed.

What is mean time to resolve?

The longest of the four. It runs from the failure to the point where the underlying cause has been addressed, not just the point where service returned.

How is MTTR calculated?

Total time across the incidents in a period, divided by the number of incidents. The formula is the easy part; deciding what counts as the start and end of each incident is the whole difficulty.

What is a good MTTR?

Not a question that has an answer without the definitions. A four hour figure could be excellent or terrible depending on what it includes, and it cannot be compared across organizations that define the clock differently.

What is the difference between MTTR and MTBF?

MTTR measures how long recovery takes. MTBF, mean time between failures, measures how often a repairable system fails. Reducing failures usually delivers more than recovering faster.

Is MTTR the same as RTO?

No. RTO is a recovery time objective, a target agreed in advance and usually written into a contract. MTTR is a measurement of what actually happened.

Why should I report the median as well?

Because incident durations are skewed. One very long outage among many short ones drags the average somewhere that describes no real incident, while the median describes a typical one.

Can MTTR be gamed?

Easily, and usually without intent. Closing tickets early, splitting one incident into several, and generous exclusions all shorten it. Written clock rules are the defense.

Should MTTR be a team target?

It is risky. There are several honest ways to improve the number without improving the service, and a metric that becomes a target tends to find them.

What starts the clock in most tools?

Commonly the alert or the ticket, because that is what the incident management tool can see. That silently excludes the detection gap, which is often the largest part of a bad incident.

Which of the four should I use?

Mean time to recovery, if the audience is the business, because it measures what the business experienced. Use the others internally to see which part of the window is slow.

What is the MTTR formula?

The MTTR formula is total repair time divided by the number of repairs in the period. Ten hours spent fixing four incidents gives an MTTR of 2.5 hours. State which meaning of the R you are using, because repair, recovery, respond and resolve start and stop the clock at different points.

Read next · Tools Network Management Software, and the Three Questions It Has to Answer Where the alert that starts most of these clocks comes from, and why a detection gap quietly improves the metric. Open this next11 min
Also worth reading
One packet a weekA short, illustrated explainer every Tuesday. No vendor pitches, unsubscribe in one click.