← Greg's blog / Training science
Your sensors are excellent. Your scores are worthless.
Smartwatches are often accused of being inaccurate. That's wrong — the opposite, in fact: the sensors have become excellent. The problem sits elsewhere, and it's more troubling. Between the signal your wrist picks up and the number you base tomorrow's session on, there are four layers of interpretation — and scientific validity evaporates a little more at each one.
The paradox: never so measured, never so unsure
Strava reports more than 180 million users across 185+ countries, 14 billion kudos exchanged in 2025 alone, a new route created every 19 seconds. Endurance training has never produced this much data. And yet, in the same survey, 46 % of respondents said they would happily use an AI as a coach.
Translation: people have the numbers, and they are looking for someone to tell them what to do with them. This is not a data-volume problem. It's a conclusion problem.
To see why, we have to stop treating "watch accuracy" as one thing. A watch doesn't produce one number: it stacks four of them, one on top of the other. And they are nowhere near equally solid.
Layer 1 — heart rate: excellent, no argument
A validation study published in Sensors compared wrist optical measurement against a reference electrocardiogram during sleep. Result for heart rate: an intraclass correlation coefficient of 1.00 (95 % CI: 0.99–1.00), with bias below 0.4 %.
That is a laboratory instrument on your wrist. When your watch says 52 beats per minute, it is 52. This needs saying, because what follows is less flattering and it would be dishonest to imply the whole thing is junk.
Layer 2 — heart rate variability: good, but already at the edge
Same study, same night, same sensor — except we're no longer measuring a rhythm, we're measuring the variation in that rhythm. Bias rises to 2–3.5 %, limits of agreement to 6–6.5 %.
That sounds small. It isn't. The authors themselves flag it: this margin of error approaches or exceeds the smallest worthwhile change for an athlete. Put plainly, the daily swing your app reads as a fatigue signal may well be the noise of the measurement itself.
A second caveat matters just as much: this study covers six participants and fifteen nights. It is a serious validation, not large-scale proof. When a manufacturer writes "scientifically validated", this is often the kind of work they mean.
Layer 3 — sleep stages: halfway there, at best
Here we leave measurement for interpretation. Whether you are in light, deep or REM sleep cannot be read off a pulse: it is inferred, by an algorithm.
A multicentre study published in JMIR mHealth and uHealth ran the test properly: 11 consumer sleep trackers, 3,890 hours of recorded sleep, compared against 543 hours of polysomnography — the laboratory gold standard. Total: 349,114 epochs analysed one by one.
The best device reached a macro F1 score of 0.69. The worst: 0.26. Not one of the eleven cleared 0.70.
For scale: on a four-state classification, an F1 of 0.26 is not far from what a slightly biased coin toss would achieve. And that is the number millions of people use, on waking, to decide whether they have "recovered enough" for their threshold session.
Layer 3b — VO2max: 13 to 16 % error
Same story for estimated maximal oxygen uptake. An Apple Watch Series 7 validation published in JMIR Biomedical Engineering measured a mean absolute percentage error of 15.79 % (RMSE 8.85 ml/kg/min) against a laboratory exercise test. A second study, in PLOS ONE, found 13.31 % mean error, with a clear tendency to underestimate.
Brands are not all equal — a Garmin Forerunner 920XT was measured at 7.3 % error on a comparable protocol. But the order of magnitude holds, and the reason is structural: these algorithms are anchored to population averages. They underestimate the highly trained and overestimate the sedentary. Useful for tracking a trend across a year. Useless for calibrating Tuesday evening's session.
Layer 4 — decision scores: where it actually breaks
Top layer: the numbers that claim to tell you what to do. Here it is no longer a question of precision. It is a question of existence.
The acute:chronic workload ratio (ACWR) is probably the most widespread metric in connected sport: the idea that the ratio between your last 7 days of load and your last 28 predicts injury risk, with a "sweet spot" you shouldn't leave. In 2020, Impellizzeri and colleagues published a methodical demolition in the International Journal of Sports Physiology and Performance: mathematical coupling between numerator and denominator, statistical artefacts, spurious correlations. Their conclusion is unambiguous: there is no evidence supporting the use of ACWR in a load-management system, nor for training recommendations aimed at reducing injury risk.
The same authors went as far as asking the British Journal of Sports Medicine to retract the founding figure of that "sweet spot". It is still shipping in dozens of apps.
HRV-guided training meets a related fate, more nuanced but just as instructive. A meta-analysis in Applied Sciences compared HRV-guided training against a conventional predefined plan. Verdict: no significant advantage on maximal aerobic capacity or endurance performance. The only solid effect (SMD = 0.50) is on vagal indices themselves. In plain terms: training by your HRV mostly improves… your HRV.
Why these scores exist anyway
There is no conspiracy here. A composite score is an excellent product: it fits inside a coloured ring, it can be compared with friends, it gives you a reason to open the app every morning. A confidence interval does none of that.
The problem isn't that a manufacturer simplifies. It's that the simplification is presented as a conclusion when it is only a hypothesis — and that it arrives stripped of the context that would give it meaning: your goal, your week, your injury history, what you have planned on Saturday, the fact that you slept on your mother-in-law's sofa.
Meanwhile the real problem hasn't moved. The reference meta-analysis on running injuries reports 17.8 injuries per 1,000 hours of running in novice runners, against 7.7 in recreational runners. Ten years and billions of data points later, that number has not budged.
What's needed instead
Not more measurement. A layer that reasons.
- Go back down to the signal. A good training decision rests on heart rate, duration, pace, actual load and how you felt — not on a score that blends all of it and then flattens it into a percentage.
- A rule beats a number. "You've had three short nights in a row and a threshold session tomorrow" is actionable. "Recovery: 42 %" is not.
- Say the uncertainty out loud. An honest system says when it doesn't know. The studies above don't say "this data is useless": they say "here is the margin of error". Hiding that margin — that's the lie.
- Context before metric. Your race in eleven weeks, your left knee, your three sessions a week: none of that lives in a sensor, and all of it determines the right call.
That is exactly the bet we're making with Greg: not adding one more score to the dashboard, but building the layer that reads your raw data, knows your goal and your history, and tells you what it means for tomorrow's session — including when it doesn't know.
Sources
Every study cited is publicly accessible. If you only open one, make it the second: the table of 11 devices speaks for itself.
- Bellenger CR, Miller DJ, Halson SL, Roach GD, Sargent C. Wrist-Based Photoplethysmography Assessment of Heart Rate and Heart Rate Variability: Validation of WHOOP. Sensors. 2021;21(10):3571. doi.org/10.3390/s21103571
- Lee T, Cho Y, Cha KS, et al. Accuracy of 11 Wearable, Nearable, and Airable Consumer Sleep Trackers: Prospective Multicenter Validation Study. JMIR mHealth and uHealth. 2023;11:e50983. doi.org/10.2196/50983
- Assessing the Accuracy of Smartwatch-Based Estimation of Maximum Oxygen Uptake Using the Apple Watch Series 7: Validation Study. JMIR Biomedical Engineering. 2024;1:e59459. biomedeng.jmir.org
- Investigating the accuracy of Apple Watch VO2 max measurements: A validation study. PLOS ONE. 2025. doi.org/10.1371/journal.pone.0323741
- Impellizzeri FM, Tenan MS, Kempton T, Novak A, Coutts AJ. Acute:Chronic Workload Ratio: Conceptual Issues and Fundamental Pitfalls. International Journal of Sports Physiology and Performance. 2020;15(6):907-913. doi.org/10.1123/ijspp.2019-0864
- Effectiveness of Training Prescription Guided by Heart Rate Variability Versus Predefined Training for Physiological and Aerobic Performance Improvements: A Systematic Review and Meta-Analysis. Applied Sciences. 2020;10(23):8532. doi.org/10.3390/app10238532
- Videbaek S, Bueno AM, Nielsen RO, Rasmussen S. Incidence of Running-Related Injuries Per 1000 h of running in Different Types of Runners: A Systematic Review and Meta-Analysis. Sports Medicine. 2015;45(7):1017-1026. doi.org/10.1007/s40279-015-0333-8
- Strava. 12th Annual Year in Sport Trend Report, 2025. press.strava.com
Frequently asked questions
Is my watch accurate, yes or no?
For heart rate, yes: a validation against ECG gives an intraclass correlation coefficient of 1.00. For the composite scores it derives from that (sleep, recovery, VO2max), it is far weaker — down to 0.26 F1 on sleep staging. The measurement is good; the interpretation much less so.
Is my watch's sleep score useful for anything?
For following a trend across several weeks, yes. For deciding whether to run your threshold session this morning, no: across 11 devices tested against polysomnography, the best reached 0.69 macro F1 and the worst 0.26. Total sleep duration is markedly more reliable than the stage breakdown.
Is the VO2max my watch shows correct?
It is off by 13 to 16 % on average on Apple Watch according to two independent studies, with a tendency to underestimate trained people. Some models do better (7.3 % measured on a Garmin Forerunner 920XT). Useful for spotting progress over a year, not for calibrating a session.
Should I steer my training with heart rate variability (HRV)?
A meta-analysis in Applied Sciences found no significant advantage of HRV-guided training on aerobic capacity or performance compared with a conventional predefined plan. HRV remains an interesting signal, but it does not replace planning.
Does ACWR (acute:chronic workload ratio) predict injuries?
No. The reference review by Impellizzeri and colleagues (2020) concludes there is no evidence supporting its use to reduce injury risk: mathematical coupling, statistical artefacts and spurious correlations. The authors even requested retraction of the 'sweet spot' figure.
“I'm not going to tell you to stop looking at your numbers — that would be silly, they're good. I'm telling you not to ask them what they don't know. Your watch knows your heart beat at 52. It doesn't know your race is in eleven weeks, that your left knee grumbled on Tuesday, and that you train three times a week, not six. That part is my job. And when I don't know, I'll say so.”
Understanding AI sports coaching →
Read next
Ski conditioning: our matrix, and the 53 studies behind it
23 ski sub-dimensions, 18 physical qualities, 100 variables, 24 guardrails. What sits underneath the Winter Season program — and what the research does not say.
Read the article →Race cancelled: the Plan B that saves your season
Your race is off. What you do in the next seven days decides the rest of your season.
Read the article →UpGreg at the Global Tech Summit in Lausanne
UpGreg will be at the Global Tech Summit in Lausanne on September 1. Come meet our AI sports coach.
Read the article →