Part 2 of 3 in Paceline's Wearable Accuracy Series. Part 1: Heart rate and HRV accuracy · Part 3: VO2max, calories, steps and SpO2
Consumer wearables are good at telling whether you're asleep and roughly how long you slept. They're much less reliable at telling deep sleep from light sleep, and their readiness and recovery scores are proprietary formulas that no independent researcher can fully check. Put an Oura Ring, an Apple Watch and a Whoop strap on the same sleeper and you can get three different sleep-stage breakdowns and three different "readiness" verdicts in the morning.
Key takeaways
- Sleep vs wake detection is strong. Devices correctly flag 90%+ of sleep, but they often miss time you spend awake in bed (specificity roughly 18–54%) (Chinoy et al., 2021; Schyvens et al., 2025).
- Sleep stages are the weak spot. In one lab night with six devices, stage-by-stage agreement with polysomnography was only 50% to 65% (Miller et al., 2022).
- Brands disagree about deep sleep. In a 2024 study, Apple Watch Series 8 underestimated deep sleep by 43 minutes and Fitbit Sense 2 by 15 minutes, while Oura Gen 3 showed no significant difference (Robbins et al., 2024).
- Readiness and recovery scores are black boxes. A 2025 review of 14 such scores from 10 manufacturers found none disclosed their exact formula (Doherty et al., 2025).
- The practical rule: use sleep data for trends on one device, and never compare your readiness score to someone else's on a different brand.
How do wearables measure sleep at all?
Clinical sleep studies use polysomnography (PSG): electrodes on the scalp, face and chest that record brain waves, eye movements and muscle tone. Sleep stages are defined by those brain signals.
Even that gold standard involves human judgment, because trained technologists score the recording epoch by epoch, and they don't always agree. In the American Academy of Sleep Medicine's inter-scorer reliability program (more than 2,500 scorers and over 3.2 million scoring decisions), agreement with the majority score averaged 82.6%, and only 67.4% for deep (N3) sleep (Rosenberg & Van Hout, 2013). That puts device results in context: the 50–65% stage agreement you'll see below falls well short of human experts, but no scoring method, human or machine, reaches 100%.
A ring or watch can't see your brain. It infers sleep from movement, heart rate, HRV and sometimes temperature, then runs a machine-learning model trained on people who slept in a lab. The Sleep Research Society's 2024 state-of-the-science review concluded that modern consumer devices beat traditional actigraphy at sleep versus wake. It also flagged clear limits: misclassifying wakefulness during the night, trouble with naps and sleep outside the main nighttime bout, artifacts, and unclear performance in people with health conditions (de Zambotti et al., 2024).
How accurate is total sleep time on Oura, Apple Watch, Whoop, Garmin and Fitbit?
Pretty good on an average night, and less good on a bad one.
- Meta-analysis (24 studies, 798 people): wrist-worn devices differed from PSG by about 17 minutes of total sleep time on average, with significant differences in sleep efficiency, sleep latency and wake after sleep onset (about 13 minutes) (Lee et al., 2025, J Clin Sleep Med).
- Umbrella review: wearables "showed a tendency to overestimate total sleep time", with MAPE typically above 10% (Doherty et al., 2024).
- Disrupted nights are harder. In a sleep-lab study of seven consumer devices, sensitivity for sleep was ≥0.93 but specificity for wake was only 0.18–0.54, and devices "tended to perform worse on nights with poorer/disrupted sleep" (Chinoy et al., 2021). In that study the Garmin Fenix 5S and Vivosmart 3 did worse than research actigraphy.
- Older Fitbit models without sleep staging overestimated total sleep by roughly 7 to 67 minutes. Newer staging models that use heart rate did better (Haghayegh et al., 2019).
The takeaway: total sleep time is useful for trends, but the error grows when you most want an accurate answer, after a restless night.
Are wearable sleep stages (deep, REM, light) accurate?
This is where devices disagree the most.
Miller et al. (2022) had 53 adults sleep one night in a lab wearing six devices at once. Agreement with PSG for sleep stages was 53% for Apple Watch Series 6, 50% for Garmin Forerunner 245, 51% for Polar Vantage V, 61% for Oura Gen 2, 60% for Whoop 3.0 and 65% for the Somfit headband. The authors concluded all six "require improvement for the assessment of specific sleep stages" (Miller et al., 2022). (Their lab receives research support from Whoop, which the paper discloses.)
Robbins et al. (2024), at Brigham and Women's Hospital, tested 35 adults for a single night. Stage sensitivity was 76.0–79.5% for Oura Gen 3, 61.7–78.0% for Fitbit Sense 2 and 50.5–86.1% for Apple Watch Series 8. Apple overestimated light sleep by 45 minutes and underestimated deep sleep by 43. Fitbit overestimated light by 18 and underestimated deep by 15. Oura wasn't significantly different from PSG for any stage (Robbins et al., 2024). The lead author sits on Oura's medical advisory board, as the paper discloses.
Schyvens et al. (2025) tested six wrist devices on 62 adults and found fair-to-moderate agreement (Cohen's kappa 0.21 to 0.53), with Apple Watch Series 8 highest at 0.53, then Fitbit Sense (0.42) and Fitbit Charge 5 (0.41). Most devices differed significantly from PSG on total sleep time, sleep efficiency, wake after sleep onset and light sleep (Schyvens et al., 2025). Note that 52 of the 62 participants were men.
Lee et al. (2023) ran the largest multi-device comparison: 11 trackers, including Google Pixel Watch, Samsung Galaxy Watch 5, Fitbit Sense 2, Apple Watch 8 and Oura Ring 3, across 75 participants and 349,114 scored epochs. Stage-classification performance (macro F1) ranged from 0.26 to 0.69, showing "substantial performance variation" between devices (Lee et al., 2023).
Notice that Apple Watch ranked near the bottom for deep sleep in one study and top for overall agreement in another. Different populations, device generations, algorithm versions and scoring methods produce different rankings. Oura's own-algorithm validation (Gen 3 with its 2.0 staging algorithm, 96 participants) reported stage accuracy from 75.5% (light) to 90.6% (REM). That study was funded by Oura (Svensson et al., 2024). Independent replication of each new algorithm version is still catching up.
Sleep tracking accuracy: what the studies found
| Study (year) | Devices | Participants | Key finding |
|---|---|---|---|
| Miller et al. (2022) | Apple S6, Garmin FR245, Polar Vantage V, Oura Gen 2, Whoop 3.0, Somfit | 53 adults, 1 lab night | Sleep/wake agreement 86–89%; stage agreement 50–65% |
| Robbins et al. (2024) | Oura Gen 3, Fitbit Sense 2, Apple Watch S8 | 35 adults, 1 night | Sleep sensitivity ≥95%; Apple −43 min deep sleep, Fitbit −15 min |
| Schyvens et al. (2025) | Fitbit Charge 5, Fitbit Sense, Withings ScanWatch, Garmin Vivosmart 4, Whoop 4.0, Apple S8 | 62 adults | Kappa 0.21–0.53; specificity 29–52% |
| Lee et al. (2023) | 11 trackers incl. Pixel Watch, Galaxy Watch 5, Apple Watch 8, Oura Ring 3 | 75 adults | Stage macro F1 0.26–0.69 |
| Chinoy et al. (2021) | Fitbit Alta HR, Garmin Fenix 5S, Vivosmart 3 + others | 34 adults, 3 nights | Specificity 0.18–0.54; worse on disrupted nights |
| Lee et al. (2025), meta-analysis | 12+ brands, 24 studies | 798 people | Total sleep time off by ~17 min on average |
Readiness, recovery and Body Battery: why the scores can't be compared
Every major brand now condenses your night into one morning number: Oura Readiness, Whoop Recovery, Garmin Body Battery and Training Readiness, Fitbit Daily Readiness, Polar Nightly Recharge, Samsung Energy Score, Ultrahuman Dynamic Recovery and others.
A 2025 review in Translational Exercise Biomedicine catalogued 14 of these scores across 10 manufacturers. The most common inputs were HRV (86%), resting heart rate (79%), physical activity (71%) and sleep duration (71%). The authors found "significant discrepancies" in data collection timeframes, metric weighting and scoring methods. None of the manufacturers disclosed their exact formulas, and few offered peer-reviewed evidence for the scores' accuracy (Doherty et al., 2025).
The inputs themselves differ too. As Part 1 showed, Polar calculates overnight HRV from the first 4 hours of sleep, Whoop weights toward the last deep-sleep phase, and Oura and Garmin average across the night (Dial et al., 2025). A readiness score built on a different HRV window, a different sleep-staging model and a different weighting scheme isn't the same measurement. An Oura 82 and a Whoop 82 are just two numbers that happen to match.
None of this makes the scores useless. Comparing your own score to your own baseline on the same device can still flag a rough night or a coming cold. The scores just aren't a shared scale, and any cross-device benchmark would need to normalize the underlying inputs rather than the branded score.
What about Samsung, Garmin, Coros and Ultrahuman?
The sleep evidence is uneven across brands. Garmin's older Fenix 5S, Vivosmart 3/4 and Forerunner 245 have been tested; current models largely haven't. Samsung appears through the Galaxy Watch 5 in one multi-device study. We found no independent peer-reviewed PSG validation of Coros or Ultrahuman sleep staging. The Sleep Research Society review notes that consumer devices update their algorithms often, so validation can go out of date quickly (de Zambotti et al., 2024).
Habits for more reliable sleep data
- Watch weekly trends, not nightly stage splits. Total sleep time and consistency are the most dependable signals. Deep and REM minutes are rough estimates.
- Wear the same device the same way every night. Same finger or wrist, snug fit, charged before bed. Switching hands or devices shifts results.
- Don't compare your readiness score to a friend's. Different brands, different formulas.
- Keep your sleep window consistent. Devices struggle with naps and irregular sleep, so regular bed and wake times help both your sleep and your data.
- Log how you feel. A quick morning check-in gives you something to compare the score against. If you feel rested but the app says you slept badly, give your own sense more weight. If you're exhausted day after day while the numbers look fine, that's worth raising with a clinician.
- After a restless night, trust the device less. That's when wake detection is weakest.
- Talk to a clinician about suspected sleep disorders. Consumer wearables aren't a substitute for a sleep study.
Where gear can help: sleep trackers and a better sleep environment
If you want a different way to track sleep, or want to make your nights more consistent, these third-party partner products are in the Paceline Marketplace:
- Garmin Index Sleep Monitor: an upper-arm band for people who'd rather not sleep in a watch.
- Withings ScanWatch 2: a watch that tracks sleep duration, sleep stages, interruptions and overnight heart rate. The original Withings ScanWatch was one of the six wrist devices tested against PSG in Schyvens et al. (2025).
- Withings Sleep Mat: an under-mattress tracker. Withings' mat was among the 11 trackers compared in Lee et al. (2023).
- Nodpod Weighted Sleep Mask and Sijo TempTune Cooling Mattress Pad: for darker, cooler, more consistent nights.
No tracker makes stage data exact, and no product here fixes the cross-brand comparison problem. Consistent conditions do make your own trend easier to read. You can browse more in the Marketplace's Sleep & Recovery collection.
Track sleep consistency, whatever you wear
Sleep data is most useful as a steady trend, not a nightly grade. The Paceline app connects the wearable you already have and focuses on the habits you repeat week to week, whether you sleep in a ring, a watch or a strap.
Previous: Part 1: Heart rate and HRV accuracy · Next: Part 3: VO2max, calories, steps, SpO2 and why normalization matters
FAQ
Which wearable is most accurate for sleep tracking?
No brand wins consistently. Oura Gen 3 matched PSG stages closely in one study (Robbins et al., 2024), and Apple Watch Series 8 had the best agreement in another (Schyvens et al., 2025). Results depend on the model, algorithm version and who is being tested.
Is my deep sleep number accurate?
Treat it as a rough estimate. Devices over- or underestimated deep sleep by 15 to 43 minutes in a recent study, and stage-by-stage agreement is often 50–65%. Even trained human scorers agree on deep sleep only about two-thirds of the time (Rosenberg & Van Hout, 2013).
Can I compare my Oura readiness score to a Whoop recovery score?
No. The scores use different inputs, time windows and undisclosed weightings.
Which sleep metrics from a wearable are most reliable?
Total sleep time and sleep consistency (regular bed and wake times) are the most dependable sleep outputs, especially as weekly averages. Overnight resting heart rate also holds up well: in a 536-night study, Oura, Polar and Whoop nocturnal resting heart rate had a mean absolute error of about 1.7–3% against an ECG reference (Dial et al., 2025). Part 1 covers heart rate accuracy in more detail. Deep and REM minutes are the least reliable.
Should I stop using my sleep tracker if it isn't perfectly accurate?
Not necessarily. A sleep tracker is a trend tool, not a diagnostic one. If it helps you notice that late workouts, alcohol or an irregular schedule change your sleep over several weeks, it's doing its job, even if the nightly stage percentages are rough.
When should I talk to a doctor about my sleep data?
If you regularly wake up unrefreshed despite what looks like enough sleep, feel tired during the day, or suspect a condition such as sleep apnea, talk to a healthcare provider. A wearable can start that conversation, but only a clinical evaluation can diagnose a sleep disorder.
Sources
- Miller DJ, Sargent C, Roach GD. A validation of six wearable devices for estimating sleep, heart rate and heart rate variability in healthy adults. Sensors. 2022;22(16):6317. PubMed 36016077
- Robbins R, et al. Accuracy of three commercial wearable devices for sleep tracking in healthy adults. Sensors. 2024;24(20):6532. PubMed 39460013
- Schyvens AM, et al. A performance validation of six commercial wrist-worn wearable sleep-tracking devices for sleep stage scoring compared to polysomnography. Sleep Adv. 2025;6(2):zpaf021. PubMed 40303381
- Lee T, et al. Accuracy of 11 wearable, nearable, and airable consumer sleep trackers: prospective multicenter validation study. JMIR Mhealth Uhealth. 2023;11:e50983. PubMed 37917155
- Chinoy ED, et al. Performance of seven consumer sleep-tracking devices compared with polysomnography. Sleep. 2021;44(5):zsaa291. PubMed 33378539
- Lee YJ, et al. Performance of consumer wrist-worn sleep tracking devices compared to polysomnography: a meta-analysis. J Clin Sleep Med. 2025;21(3):573-582. PubMed 39484805
- Haghayegh S, et al. Accuracy of wristband Fitbit models in assessing sleep: systematic review and meta-analysis. J Med Internet Res. 2019;21(11):e16273. PubMed 31778122
- Svensson T, et al. Validity and reliability of the Oura Ring Generation 3 with Oura sleep staging algorithm 2.0 when compared to multi-night ambulatory polysomnography. Sleep Med. 2024;115:251-263. PubMed 38382312
- de Zambotti M, et al. State of the science and recommendations for using wearable technology in sleep and circadian research. Sleep. 2024;47(4):zsad325. PubMed 38149978
- Doherty C, Baldwin M, Lambe R, Burke D, Altini M. Readiness, recovery, and strain: an evaluation of composite […] in consumer wearables (review of 14 scores from 10 manufacturers). Translational Exercise Biomedicine. 2025;2:128-144. doi:10.1515/teb-2025-0001
- Doherty C, et al. Keeping pace with wearables: a living umbrella review. Sports Med. 2024;54(11):2907-2926. PubMed 39080098
- Dial MB, et al. Validation of nocturnal resting heart rate and heart rate variability in consumer wearables. Physiol Rep. 2025;13(16):e70527. PubMed 40834291
- Rosenberg RS, Van Hout S. The American Academy of Sleep Medicine inter-scorer reliability program: sleep stage scoring. J Clin Sleep Med. 2013;9(1):81-87. PubMed 23319910
This article is for educational purposes only and isn't medical advice. Consumer sleep trackers can't diagnose sleep disorders such as sleep apnea or insomnia. If you're concerned about your sleep, talk to a qualified clinician.