The numbers you see
Durability and forecasts
Whether you hold up late in long work — and why we don't promise a finish time.
9 min read
A month ago you ran for two hours and barely faded over the last twenty minutes. Today, the same two-hour run — and by the end your heart rate had crept up and your pace had slipped down. The difference between those two runs is durability: the ability to still be roughly the same athlete at the end of a long effort as you were at the start.
The fourth quality of endurance
Endurance is usually described by three quantities: the ceiling, VO2max, your threshold, and economy. In recent years a fourth has increasingly been added — durability. Two athletes with the same threshold on fresh legs can look completely different in the third hour, and over long distances that decides a great deal.
In 2025 this stopped being theory: the decoupling of heart rate from work on controlled sessions was checked directly against the drop in power after two and a half hours — the link is strong, and the authors call it a practical way of tracking durability. We build on that work and compute the metric ourselves, from your recordings.
What it looks like in numbers
On a long, steady session the ratio of work to heart rate is calculated: speed per beat in running, corrected for terrain, power per beat on the bike. The warm-up and cool-down — the first and last ten percent of the time — are discarded, and the remaining part is assessed in two ways.
The main answer is the minute from which efficiency slipped. Every minute, the last five minutes of work are compared with the first ten minutes of the evaluated part. As soon as efficiency drops by five percent and does not come back, that minute is the answer: “today it held to minute 78, a month ago to minute 62”. If efficiency is still at its starting level at the end, it held for the whole session.
The second number is the drift between anchors. The first ten and the last ten minutes of the evaluated part are compared: by how many percent the same price in heartbeats bought less movement. That is decoupling of heart rate and pace. Such calculations used to split a session in half; 2025 studies showed that anchor windows at the start and the end separate people more precisely, and that the moment decoupling begins says more about a person than its size. That is why the minute became the main answer, not the percentage.
We call this efficiency drift rather than durability, and we do so deliberately. In research, durability is measured after a genuinely large amount of work — of the order of 1,500–2,500 kilojoules, or a good two hours and more. Our thresholds are lower: they're chosen to fit the training an amateur actually does. So the app won't tell you “your durability has improved” — it shows a trend across comparable sessions, and that's a different claim. Cadence is not part of the calculation: once speed is accounted for, it adds nothing of its own.
Which sessions qualify at all
Far from all of them. Only long, steady aerobic work makes it into the comparison:
- runs from 60 minutes or rides from 90 minutes;
- steady effort with no intervals: minute-by-minute speed or power varying by no more than 10%, at easy or base intensity;
- heart rate and pace recorded together and usable for at least 80% of the time;
- the first and last 10% of the time are discarded — otherwise the warm-up and cool-down would draw the drift by themselves;
- pauses longer than 30 seconds are excluded;
- heart rate across the whole compared series measured by one and the same device.
Pool swimming doesn't get this calculation at all: there's no continuous comparable speed signal there, and we treat heart rate in the water as unusable. We're not going to invent a substitute just so it exists “in all three sports”.
Where to see it
On the activity page — the “Durability” card: up to which minute efficiency held and the drift between anchors; for a session that does not qualify it says why the calculation does not apply. On the fitness page — the list of comparable sessions over twelve weeks with that minute for each.
Your trend only, and only once enough has built up
One session says nothing: cardiac drift happens to everyone, and on its own it is neither good nor bad. Meaning appears when you are compared with yourself.
So only similar sessions are compared: the same sport, similar duration (within ±20%), similar intensity and the same set of significant conditions. A trend appears once there are at least six such sessions over four weeks or more. Until then the honest answer is calibrating.
There will be no “normal” drift value here and no comparison with other people — only the direction of your own curve.
What distorts the picture
Conditions affect drift more than fitness does, and that is the main reason for the strict filters. Heat and dehydration raise heart rate by themselves. Terrain in the second half of a route changes the price of the same speed — in running we correct for gradient, but the correction is not almighty. A skipped breakfast or missed fuelling on a long session produces a fade that has nothing to do with your form. Starting too fast guarantees a “deterioration” in the second half simply by construction.
Hence the rules: the trainer is not compared with the road, a hot day with a cool one, a fuelled session with a fasted one. If a strong condition is unknown, the platform lowers the quality of the record rather than pretending it wasn't there.
Goal readiness: three honest answers
Separate from drift there is the goal preparation status, and it has exactly three answers: on track, at risk and don't know yet. It looks at the training prerequisites: whether the key long sessions and volumes of the current phase are being done, what is happening to your threshold and accumulated fitness. This status is not shown in the app yet: the calculation is specified and approved, but not built.
“On track” means one thing: the confirmed training prerequisites are being met. It is not a guarantee of a race result and not a conclusion about your health — wellbeing questions live separately and do not affect this status.
When there is not enough data, the status says so and explains exactly which data is missing.
A guide from your thresholds instead of a forecast
On the race page every leg carries a range, and below them stands a total — also a range. This is the guide from your thresholds: what your confirmed thresholds turn into on this distance under the listed assumptions. The swim is derived from your critical swim speed, the bike from your threshold power, your riding position calibrated from your own outdoor rides with a power meter, and the course profile if one is uploaded, the run from your critical running speed with an allowance for running after the bike. Transitions are added on top — a standard assumption, not a consequence of your thresholds.
This is arithmetic, not prediction. We take the numbers already known about you and convert them into minutes — at even pacing, without wind, heat, fuelling trouble or a queue at the rack. Next to every leg you see which threshold, from which date, went into it. If a discipline has no fresh confirmed threshold, that leg stays without a number, there is no total at all, and the app says what is missing. We will not substitute an average for people of your age and weight.
The ranges are added honestly: the lower bound of the total is the sum of all lower bounds, the upper bound the sum of all upper bounds. There is a way to add them so the total comes out narrower and prettier, but it assumes the leg errors are independent — and on a real race day heat, a skipped breakfast and a fast start hit all three legs at once. So our total is wide, and that is the content of the answer, not a shortcoming. The actual time may fall outside either bound: the middle of the range is not “the most likely result”, and its width is not precision.
What you will not see here: a single number instead of a range, a percentage “chance of making it”, or a comparison with what the guide showed a month ago. The first two promise precision we do not have. The third turns arithmetic into a statement about how you are progressing towards a result — and that is a forecast.
Why this is not yet a forecast and what has to happen
A forecast is a different statement: “you will finish in such-and-such time, and we know how often we are right”. We do not have the second half. For it to appear we need real finishes — not our calculations, but your results, against which the calculation can be checked after the fact. The conditions are fixed in advance so the model cannot be declared ready just because it looks plausible:
- the calculation is checked forwards, on real finishes, not fitted to past data;
- at least 100 independent finishes are collected, and at least 20 per distance;
- the stated interval must cover at least 75% of actual results on each distance;
- the mean error and systematic bias are published — including by sex, age and level;
- no such number weakens any safety limit.
Until then the word “forecast” will not appear on screen, and the guide remains what it is: your current thresholds translated into minutes. The race results you enter after a race are stored next to the guide that stood on race day — that is how the corpus is built. Your results do not yet change the guide itself: from two or three races one cannot tell your own trait from the course and the weather of that day. How we decide what counts as proven is in the article “What it's all based on”.
Terms in this article
What to read next
- Fitness, fatigue and formThree lines on one chart: what has built up, what hasn't cleared yet, and how fresh you are today.
- How we make methodology decisionsWhy some rules are fixed, some adjustable, and some we don't ship at all.
- How the plan is builtSeason phases, load growth, recovery weeks and the taper into your key race.