Body Battery, Recovery, Readiness: 14 “readiness scores” from 10 brands — and not a single published formula
In May Google released the Fitbit Air, a screenless band whose headline number is “readiness”. Garmin, WHOOP, Oura, Polar and Coros all have similar indices. A 2025 review found 14 such scores from 10 manufacturers: none disclosed its formula, and only a handful showed validation against a reference standard. Some of the raw data underneath them, however, is measured quite well.

On 7 May 2026 Google released the Fitbit Air — a $99.99 band with no screen. It weighs 12 grams with the strap and displays nothing itself; the main thing the user sees in the app each morning is the Readiness Score, a measure of “readiness” for training. The same idea lives on in Garmin (Body Battery and Training Readiness), WHOOP (Recovery), Oura (Readiness) and Polar (Nightly Recharge). A single number from 0 to 100 decides whether you run intervals today.
The question worth asking before you rebuild your plan: what is this number, and has anyone checked that it means what it promises?
What was examined
In 2025 Cathal Doherty and colleagues — among them Marco Altini, creator of the HRV4Training app — published the first systematic review of such indices in Translational Exercise Biomedicine. The authors call them composite health scores: a single number built from several sensor signals.
The method is simple and revealing. The researchers collected everything the manufacturers publish themselves: technical white papers, user guides, app screens and scientific papers where they exist. Then they looked at what each score is made of.
They found 14 scores from 10 manufacturers:
- Fitbit — Daily Readiness;
- Garmin — Body Battery and Training Readiness;
- Oura — Readiness and Resilience;
- WHOOP — Strain, Recovery and Stress Monitor;
- Polar — Nightly Recharge;
- Samsung — Energy Score;
- Suunto — Body Resources;
- Ultrahuman — Dynamic Recovery;
- Coros — Daily Stress;
- Withings — Health Improvement Score.
What is inside
The ingredients are similar across the board:
- heart rate variability (HRV) — in 12 of the 14 scores (86%);
- resting heart rate — in 11 of 14 (79%);
- daytime activity — in 10 of 14 (71%);
- sleep duration — in 10 of 14 (71%).
Beyond that, things diverge. Different data windows: some scores are based on the previous night, others average several days, and Garmin's Body Battery even “drains” over the course of the day. Different component weights. Different scoring rules.
And most importantly: no manufacturer disclosed the exact formula. Only a handful provided empirical validation or peer-reviewed data showing that the score is accurate and means something for health or performance. The authors' conclusion: the idea is promising, but the scientific validity, transparency and clinical usefulness of such indices remain unclear.
That does not mean the numbers are random. It means that nobody on the outside can say how far to trust them.
What about the raw data underneath?
Here there is evidence. In the same year, 2025, Michael Dial and colleagues (Physiological Reports) tested nocturnal resting heart rate and HRV from five devices against ECG: 13 healthy adults (6 women) slept wearing a reference ECG and several gadgets at the same time, for a total of 536 nights.
HRV — agreement with ECG (Lin's concordance correlation coefficient, CCC) and mean absolute error:
- Oura Gen 4 — CCC 0.99, error 5.96%;
- Oura Gen 3 — 0.97, 7.15%;
- WHOOP 4.0 — 0.94, 8.17%;
- Garmin Fenix 6 — 0.87, 10.52%;
- Polar Grit X Pro — 0.82, 16.32%.
Resting heart rate is measured better: Oura's error was 1.67–1.94%, Polar's 2.71% and WHOOP's 3.00%. Garmin was excluded from the resting heart rate analysis because of methodological discrepancies.
One point from this study's limitations section is worth paraphrasing: every device has its own algorithm, and almost nothing is known about which metrics, and with what weights, make up the “readiness” or “recovery” the user ultimately sees. On top of that, the algorithms are updated from time to time.
Why a single number misleads
Three layers of uncertainty stack up.
- Sensor error. If nocturnal HRV is off by 10–16% on average, a morning reading of “8% below baseline” may be device noise rather than fatigue.
- Sleep is measured worse than it seems. Trackers are decent at detecting sleep but poor at detecting wakefulness: a night with awakenings looks longer and better than it was. And sleep duration feeds into 71% of the scores.
- Unknown weights and windows. Two scores for the same person can legitimately disagree simply because one looks at last night and the other at the past week.
That is why comparing “readiness” across brands is pointless, and after switching watches or firmware it is better not to treat your old history as comparable.
How to apply this
- Look at the components, not the total. Morning HRV and resting heart rate relative to your personal baseline are what has at least partly been validated. How to train by HRV without fooling yourself is covered in our article on HRV-guided training.
- One day is noise, several days are a signal. HRV below baseline and resting heart rate above it for several mornings in a row mean more than a single red number.
- Weigh the score against how you feel. Red “readiness”, but the warm-up feels easy and your heart rate at a familiar pace is normal — you can do the session, perhaps trimming the volume. A green number with heavy legs and high perceived exertion is no reason to push.
- One device, the same conditions. The same device, the same wrist or finger, measured overnight. Do not compare numbers across brands.
- Choosing a gadget for HRV? In Dial's validation, the Oura ring and WHOOP agreed with ECG better than the Garmin and Polar watches. For the interval sessions themselves, a chest strap still measures heart rate more accurately.
- Set your zones from tests, not from “readiness”. The score does not know your threshold — the zones generator works that out from your test results.
The bottom line
- A 2025 review found 14 composite scores from 10 manufacturers — Fitbit, Garmin, Oura, WHOOP, Polar, Samsung, Suunto, Ultrahuman, Coros, Withings.
- Inside there is almost always HRV (86%), resting heart rate (79%), activity and sleep (71% each), but the windows, weights and formulas all differ and are undisclosed; only a handful provided validation.
- The raw data are measured unevenly: nocturnal HRV error is 6–7% for Oura, 8% for WHOOP, 10.5% for the Garmin Fenix 6 and 16% for the Polar Grit X Pro.
- Use multi-day trends in HRV and resting heart rate plus your own sensations, and treat the “readiness” number as a reminder to look at them, not as a decision.
Sources: Doherty C., Baldwin M., Lambe R., Burke D., Altini M. “Readiness, recovery, and strain: an evaluation of composite health scores in consumer wearables”. Translational Exercise Biomedicine, 2025;2:128–144. DOI: 10.1515/teb-2025-0001 · Dial M.B., Hollander M.E., Vatne E.A., Emerson A.M., Edwards N.A., Hagen J.A. “Validation of nocturnal resting heart rate and heart rate variability in consumer wearables”. Physiological Reports, 2025;13(16):e70527. DOI: 10.14814/phy2.70527