
TL;DR — Apple’s nine-page heart rate study is better than it had to be. It enrolled 1,460 people at five sites, used a chest-strap ECG as reference, and reported its own losses. It also discloses three things worth sitting with: Apple read its own watch through an internal logging app no competitor was given, Apple chose how Samsung’s untimestamped background data would be scored, and the Daily Living protocol — the part underpinning the “60 times more often” marketing claim — had no pre-specified sample size at all. None of that makes the result wrong. It does change what the result means.
The Story
Apple published one new health white paper for the Series 12 launch, and it is not the one the four new algorithms needed. It is a heart rate accuracy study, nine pages, dated September 2026.
The headline is easy to repeat: across 159 activity-by-measure comparisons against six competitors, Apple Watch was more accurate in 139.
The paper is worth reading properly, though, because it is a genuinely unusual document. Most vendor benchmarks are written to be un-checkable. This one hands you the knife. Apple wrote down its exclusion criteria, its statistical model, its data quality thresholds, the exact reason it needed a special tool to read its own watch, and the two comparisons it lost. A company writing pure marketing does not include the sentence “the comparison device’s accuracy was higher in 2.”
So the interesting question is not whether Apple cheated. It is narrower and more useful: given exactly what Apple describes, what does the 139 actually measure?
The Setup
Participants wore an Apple Watch on one wrist and a competitor on the other, with a Polar H10 chest strap ECG as the reference standard. Which wrist got the Apple Watch was randomized, so the dominant hand carried it about half the time — a small detail that matters more than it sounds, because wrist motion is one of the main things that breaks optical heart rate.
The comparison set: Garmin Forerunner 970, Google Pixel Watch 4, Huawei Watch 5, Samsung Galaxy Watch 8, WHOOP 5, and Oura Ring 5 on an index finger. Those brands account for roughly 90% of the worldwide smartwatch category, according to Omdia shipment data cited in the paper — not Apple’s own estimate.
Activities were chosen to break things on purpose. Outdoor runs, treadmill intervals, outdoor cycling, HIIT, strength training, 15 to 20 minutes each. Apple names the failure modes it was hunting: cadence lock in running, loaded forearm grip in strength work and cycling, rapid wrist motion and fast heart rate transitions in HIIT. There was also a separate Daily Living protocol — rest, desk work, meal preparation, walking, about 20 minutes each — to test background heart rate.
Five sites: two near Cupertino, plus San Diego, Austin, and Selangor, Malaysia. Apple ran Cupertino and Austin itself; a third-party contract research organization ran San Diego and Selangor, and 59% of the analyzed participants came from those CRO sites.
Of 1,460 enrolled, 1,254 contributed at least one paired measurement. Twenty percent had Fitzpatrick skin tones V–VI, which is the demographic optical sensors historically fail and wearable studies historically underweight. Mean age 39, mean BMI 26. Thirty-five percent female — described as “representation across sex,” which it is, though it is a roughly two-to-one male skew.
One exclusion criterion is worth quoting because nobody writes this by accident. Alongside beta blockers, equipment allergies, unstable cardiovascular conditions, and tattoos or large moles at the sensor site, Apple excluded people working in tech media.
Everyone Got a Different Pipe
Here is the first thing that changes how you read the result.
Apple could not just ask each device for its data, because each manufacturer exposes data differently. So Apple picked a method per device. For most competitors it used a Bluetooth Low Energy stream from the manufacturer’s own heart rate app into a logging app on a separate iPhone — because, as the paper puts it, “not all manufacturers write high-fidelity data to HealthKit, or permit such data to be exported from their companion apps.”
For the Apple Watch, Apple used an internal data logging application.
The stated reason is specific and checkable: background heart rate is only written to HealthKit every 30 seconds, to conserve device memory. But the value the watch actually shows you — in a watch face complication, in the Workout app — is a single reading every 5 seconds. To evaluate the number the user sees, Apple needed a tool that captures the number the user sees. HealthKit would have understated its own watch.
That logic holds. It is also true that only one company in this study had the ability to build that tool.
And this is where the framing of the whole paper matters. Apple states its selection rules up front, and there are three: the data stream had to be visible to the user in the product experience, it had to carry timestamps, and where a device offered more than one qualifying stream, Apple took the highest frequency available. The second and third rules are the ones that quietly do the work — a device that surfaces a number but will not tell you when it was taken cannot be scored the normal way, and a device that exports a slower stream than it displays gets scored on the slower one. The first rule is a deliberate and defensible choice — arguably the more honest consumer question is “how accurate is the number on my wrist,” not “how good is the raw diode.” But it has a consequence Apple does not spell out:
Part of what this study measures is how much data each company lets out of its own ecosystem.
If a competitor’s sensor is excellent but its app only exports a coarse stream, this methodology scores the coarse stream. That is a real thing that affects real users. It is not the same claim as “Apple’s sensor hardware is more accurate,” and the paper’s conclusion — “Apple Watch offers users the most accurate heart rate sensing among the wearable devices evaluated” — sits right on the seam between the two.
The Samsung Problem
The Galaxy Watch 8 case is the sharpest illustration, and to Apple’s credit it is disclosed in full.
For daily-living background readings, Samsung’s export gave no per-sample timestamps. It gave a start time, an end time, a minimum, a maximum, and one additional heart rate value whose derivation Samsung does not document.
So Apple had to decide what that mystery value should be compared against. It tested three options: the reference at the interval’s start, the reference at the interval’s end, and the mean of the reference across the interval. The interval mean fit best, so that is what Apple used.
Read that again. Apple selected the scoring rule for a competitor’s data by testing which of three candidate alignments matched the reference most closely — the best-fit option rather than an arbitrary one — and then still reported that Samsung lost.
That is the fair-minded reading, and I think it is the right one. But the underlying situation stands regardless of intent: in one head-to-head comparison, one competitor’s grading method was chosen by the other competitor, because the first one shipped an undocumented number.
The Actual Scoreboard
The 139 is real. It is also not the whole line.
Across 159 activity-by-measure comparisons: Apple better at 99% confidence in 139, indeterminate in 18, competitor better in 2. Ignore statistical significance and just count which way the point estimate leaned, and Apple is ahead in 155 of 159. On the overall workout and overall daily living protocols, Apple beat every device on both RMSE and MAE.
That is a strong result by any standard. Two details are worth pulling out of the figures, though.
The Pixel Watch 4 beat the Apple Watch. Twice — RMSE and MAE, in the rest condition of the Daily Living protocol. Apple reports it and adds that the margin was under 0.5 bpm. Pixel Watch 4 also produced an indeterminate result against Apple in cycling, the only workout activity where Apple did not win outright against everyone.
Now notice the asymmetry in how those are written. Apple’s wins are reported as statistically significant, direction only. Apple’s single loss is reported with a magnitude attached — “by less than a margin of 0.5 bpm.” Both statements are accurate. But the effect sizes for the 139 wins live only inside Figures 1 through 3; there is no table of numbers in the text. You are told Apple won and by how little it lost.
For calibration: the study was sized to detect differences in mean absolute error of roughly 2.5 bpm in running and cycling, up to 3.7 bpm in HIIT. So “statistically significant” here is not synonymous with “large enough to feel.” Whether a 2 bpm edge changes anything about your training is a separate question from whether it is real.
The Least-Designed Part Is the Part Marketing Needs Most
This is the finding I did not expect.
Apple pre-specified its sample size: a minimum of 50 participants per workout activity per comparator device, with 80% power, α = 0.05, and a Bonferroni adjustment for the six device comparisons. That is a properly powered design.
And then one clause: sizing “pertains only to workouts and not Daily Living.”
The Daily Living protocol had no pre-specified sample size. It is still analyzed with the same linear mixed model and the same 99% bootstrap confidence intervals, so the reported results are not casual. But it was not designed to a target in advance.
Daily Living is background heart rate. Background heart rate is precisely what the Series 12 marketing claim is about — readings 60 times more often than Series 11, the finer-grained passive data that Readiness and Health Age are built from. The single most-promoted capability of the new sensor is validated in the half of the study that got no power analysis.
What Apple Chose to Tell You
I want to be even-handed, because the easy version of this article is cynical and the easy version is wrong.
Apple disclosed, without being obliged to:
- that it used a tool on its own device that no competitor had, and exactly why
- that it picked the alignment method for a competitor’s undocumented data
- that it lost two comparisons, and to whom
- that 18 more were indeterminate
- that the Daily Living arm was not power-sized
- that its participant pool was 65% male
Each of those is a stick handed to a critic. Most vendor white papers contain none of them.
The gaps that remain are structural rather than sneaky. The protocol was approved by review boards within Apple — the paper names groups covering research ethics, safety, privacy, and data governance, but no external institutional review board. There is no trial registration. The paper is not peer-reviewed, and the underlying data is not available, so nobody outside Apple can re-run the analysis with different reasonable choices and see whether 139 becomes 120. Those are the normal conditions of vendor research, not misconduct — but they are the reason vendor research and independent research are not interchangeable, no matter how carefully the vendor writes.
What This Means If You Wear One
- Heart rate is the one Series 12 capability with fresh evidence behind it. Everything else in the new health stack shipped without a paper. If you are choosing based on documentation, this is the documented part.
- Read the claim as “the number on the screen,” not “the sensor.” That is Apple’s own stated framing, and it is the useful one for a buyer — but it means a competitor’s poor export policy counts against its score.
- Do not translate “more accurate” into “meaningfully different for you.” The design could detect gaps around 2.5 to 3.7 bpm. If you are training by heart rate zones, ask whether a couple of bpm crosses a boundary you care about. Usually it does not.
- The background-data claim is the softest part. Workout accuracy was designed, powered, and won convincingly. Passive all-day accuracy was measured with less advance design — and it is the input to the scores Apple is promoting hardest.
The Takeaway
Grading your own homework is not automatically dishonest. It becomes dishonest when you hide the rubric. Apple published the rubric: which tool read which device, how a competitor’s ambiguous data was handled, where the power analysis applied and where it did not, and which comparisons it lost.
What you are left with is a well-run study that answers a narrower question than its conclusion sentence implies. Apple Watch reports heart rate to its user more accurately than six competitors report heart rate to theirs, under conditions Apple designed, using an extraction method only Apple could build for itself, and with the strongest evidence in workouts rather than in the passive background data its new features actually run on.
That is still a good result. It is just a different sentence than “Apple has the most accurate heart rate sensor,” and the gap between those two sentences is where all the interesting reading is.
Source: Apple Watch Heart Rate Accuracy Study, September 2026
Photo: Nik / Unsplash
댓글 남기기