TraviaTechPie Review

Review Tech, Science, Finance

  • TL;DR — watchOS 27 installs on eight Apple Watches, and five models lose support in one year: Series 6, Series 7, Series 8, SE 2, and the original Ultra. That is the largest single-year cull in the platform’s history, and it retires a four-year-old flagship. But the cut and the features are not the same story, and conflating them is the mistake almost every compatibility article makes. Apple’s own explanation for dropping those five names Siri AI and a tap gesture — not the health system. The health system, readiness and five-second heart rate and recovery HRV, runs on exactly two watches, both of which ship with watchOS 27 already installed. Eight watches get an update. Two get a product.

    The Story

    Every year there is a version of this article and it is always the same article: here are the watches that can install the update, here are the ones that cannot, sorry about your Series 5.

    That framing quietly assumes something that turns out to be false — that installing the update is what determines whether you get the year. It is not. There are three separate floors in watchOS 27, they sit at different heights, and the one everybody writes about is the least consequential of the three.

    Here is the short version. Apple cut five watches loose. Apple’s stated reason for that cut is Siri AI and a new tap gesture. And then Apple put everything it actually spent the keynote on behind hardware that did not exist until this generation. The five watches were not cut so the health system could ship. They were cut so a gesture could.

    Five Models, One Cut

    Here is the damage, plainly.

    Runs watchOS 27:

    ModelReleasedChip
    Apple Watch Series 92023S9
    Apple Watch Ultra 22023S9
    Apple Watch Series 102024S10
    Apple Watch Series 112025S10
    Apple Watch Ultra 32025S10
    Apple Watch SE 32025S10
    Apple Watch Series 122026S11
    Apple Watch Ultra 42026S11

    The last two ship with watchOS 27 preinstalled, so the number of watches that upgrade to it is six.

    Does not run watchOS 27: Series 6, Series 7, Series 8, SE (2nd generation), Apple Watch Ultra (1st generation).

    Five models in one year. The usual pattern is one or two — a single generation ageing out the back of a rolling window.

    The Ultra is the one that should make you pause. It arrived in 2022 at $799, opening a tier above every Apple Watch that came before it and aimed squarely at people who buy equipment rather than accessories: divers, ultrarunners, the ambiently outdoorsy. Every Ultra since has launched at that same $799, so the price of the tier has not moved — only the length of time it buys you. Four years later it cannot install the current OS. A watch sold on ruggedness and longevity got fewer feature updates than the $399 Series 9 sitting above it on the supported list.

    And note the floor. The oldest supported device is from 2023. The platform’s memory is now three model years deep.

    Three Floors, Not One

    This is the part that compatibility coverage collapses into a single list, and it should not be.

    Floor one — can install watchOS 27. Eight models, as above. Cross this and you get the interface: Liquid Glass, the new watch faces, the dynamic app grid, the Find My redesign, custom passes in Wallet, Cycle Tracking additions.

    Floor two — Siri AI and the new tap gesture. Same eight models. This is the floor Apple actually cited when asked why five watches were dropped.

    Floor three — the health sensing system. Readiness, heart rate measured every five seconds all day, recovery HRV. From Apple’s own announcement: “Apple Watch Series 12 and Apple Watch Ultra 4 introduce readiness,” and “Apple Watch Series 12 and Apple Watch Ultra 4 now measure heart rate every five seconds, all day long.” Two models. Both of them ship with watchOS 27 already on them.

    Audio Intelligence — Sound Recognition, faster Shazam, Live Rewind, Siri Recap — sits on floor three as well, restricted to the Series 12 and Ultra 4. I went through why the microphone story there is more complicated than the marketing suggests separately.

    So if you own a Series 9, 10, 11, Ultra 2, Ultra 3 or SE 3, here is your year: a redesigned interface, a smarter Siri, a new gesture, and none of the sensing. The update is real. The product announcement was for somebody else.

    Apple Cut Five Watches for a Gesture

    Now put the two facts next to each other, because separately they are unremarkable and together they are not.

    Apple’s only public explanation for the cut came through an interview rather than a support document, so treat the sourcing accordingly. Cait Dooley, Apple Watch and Health product marketing manager, told 9to5Mac: “The great new features in watchOS, including the capabilities of Siri AI and the new tap gesture, work best with the processing power that is in Apple Watch Series 9 and later, Ultra 2 and later, and SE 3.”

    Read the list at the end of that sentence. Series 9 and later, Ultra 2 and later, SE 3. That is the watchOS 27 compatibility list, exactly — no device above the line that cannot install the OS, no device that installs the OS but sits below the line.

    The compatibility list was not drawn where old watches stop working acceptably. It was drawn at the floor for Siri AI and a tap gesture, and the OS list was made to match it.

    Which is a strange trade once you notice that the health system — the thing the announcement was actually about — did not need that line drawn at all. Readiness and five-second heart rate are on floor three regardless. They require a sensor array, not a chip generation. Cutting the Series 8 loose bought Apple nothing on floor three. It bought a gesture and a Siri.

    The Double Tap Precedent

    Here is why floor three is unlikely to move, for anyone hoping a later point release loosens it.

    In watchOS 10.1, Apple shipped Double Tap and restricted it to the Series 9 and Ultra 2. The wording of the restriction was unusually specific. Apple did not say “newer models” — it described the gesture as powered by the S9 SiP and its new 4-core Neural Engine, naming the silicon directly.

    Three years have passed. Double Tap has not come to the Series 8.

    This is not a case of missing hardware. The Series 8 has an accelerometer, a gyroscope, an optical heart sensor and a Neural Engine. The gesture is a signal-processing problem across sensors the Series 8 already carries. Enthusiasts have made this argument every year since. Apple has never moved, and Apple has never explained further than that original sentence.

    Now hold that against a case where Apple did not name the silicon. Sleep Score arrived in watchOS 26 with no chip in its description, and it went all the way back to the Series 6 — hardware from 2020.

    Two features, two announcements, two very different outcomes. The variable that predicts which happens is not how demanding the feature looks. It is whether Apple’s announcement names a specific piece of hardware.

    For floor three, Apple named the models in the first sentence of the press release. On the historical record, that is not a temporary state of affairs.

    The Second Gate: Your iPhone

    Clear floor two and you can still get nothing, because several of the watchOS 27 features gate on the phone.

    FeatureWatch requirementiPhone requirement
    Liquid Glass, watch faces, Find My, WalletAny watchOS 27 deviceNone
    New tap gestureAny watchOS 27 deviceNone
    Siri AI, Workout Buddy, Call ContextAny watchOS 27 deviceApple Intelligence iPhone
    Sound Recognition, ShazamSeries 12, Ultra 4iPhone 11 or later, or SE 2nd gen or later, on iOS 27
    Live Rewind, Siri RecapSeries 12, Ultra 4iPhone 16 or later — Apple lists models individually, and the 16e is absent while the 17e is present
    Readiness, 5-second heart rate, recovery HRVSeries 12, Ultra 4None

    Work an example. You own a Series 11 — the generation before the Series 12, fully supported, no complaints — and an iPhone 14, which is a perfectly good phone that takes perfectly good photographs. You get the interface, the tap gesture, and better treadmill distances. You do not get Siri AI, Workout Buddy or Call Context, because those want an Apple Intelligence phone. You do not get any of the sensing, because that wants a watch you do not have.

    What is left is a visual refresh and a gesture. The watch was never the constraint on half of that list, and nobody puts the phone in the headline.

    What a Supported-But-Older Watch Actually Gets

    The fair version of this is not that floor-two watches got nothing. They got a real update:

    • Siri AI, the dynamic app grid, and the new tap gesture
    • Workout Buddy, which no longer needs the iPhone physically nearby during the workout, provided the watch has Wi-Fi or cellular
    • Call Context
    • The Find My redesign, custom passes in Wallet, Cycle Tracking additions
    • Liquid Glass and the new watch faces

    And one item deserves more attention than it will get: treadmill distance accuracy. Apple describes it as new machine learning models that improve distance “from the very first step of a walk or run,” and the announcement attaches no model restriction to it at all.

    That is the counter-argument to reading every gate as artificial. Apple demonstrably can improve sensor-derived accuracy on older watches, in software, when it wants to — and when it does, it conspicuously does not name a chip. Which makes the gates where it does name one harder to explain away, not easier.

    How Long Does an Apple Watch Actually Last

    Worth stating plainly: Apple does not publish a number of years of watchOS updates. There is a vintage-and-obsolete policy governing hardware service, and no equivalent commitment anywhere for software. Whatever you believe about how long your watch keeps gaining features, you inferred it. Apple did not tell you.

    The observed record is roughly four to six years, with the recent trend at the low end. The original Ultra got four. The oldest device on the current list, the Series 9, is three years old.

    Compare that to the only rival that publishes a commitment: the Galaxy Watch 9 ships with a stated five years of updates. Whether Samsung honours it is a separate question, and the two platforms are not otherwise alike — but one company put a number in writing and the other did not, and only one of those positions can be checked later. I put that watch against the Series 12 and the Garmin Venu 4 if you are choosing between platforms rather than within one.

    Two cautions while we are here. Apple has not published a security-update commitment for dropped watchOS versions either, so “it will still get patches for a while” is an inference from past behaviour, not a promise. And you will find plenty of forum posts reporting battery degradation after major watchOS updates; those are anecdotal, and Apple’s release notes carry only boilerplate about updates that “may affect performance and/or battery life.” Neither the complaints nor the disclaimer should decide anything for you.

    So Should You Upgrade

    • You have a Series 6, 7, 8, SE 2 or original Ultra. Your watch keeps doing everything it does now. It stops gaining features, and at some unannounced point stops getting patched. There is no urgency. There is a direction of travel.
    • You have a Series 9 or Ultra 2. You are supported and you are the oldest device on the list. Plan on being cut within a cycle or two — not this one.
    • You have a Series 10, 11, Ultra 3 or SE 3. No action, and no envy required: on floor three you are in exactly the same position as the Series 9.
    • The health sensing system is why you are buying. It requires a Series 12 or Ultra 4 and nothing else — no phone requirement attached. That is the cleanest upgrade case in this whole article.
    • Audio Intelligence is why you are buying. Series 12 or Ultra 4 and, for the parts most people want, a phone on Apple’s enumerated list — which starts at the iPhone 16, skips the 16e, and picks the 17e back up. Check the phone before the watch. I broke down what the Ultra 4 adds over the Ultra 3 if that is the tier you are shopping.
    • You want the longest useful life per dollar. Buy the newest sensor generation you can, not the best-specced body. Floor three tracks the sensor array, and every precedent in this article says those floors do not descend.

    The Takeaway

    The headline is that five Apple Watches lost support. The finding is that the cut and the features are two unrelated stories that get told as one.

    Apple drew the compatibility line exactly where Siri AI and a tap gesture needed it — Apple said so — and then put the year’s actual announcement behind a sensor array that no amount of chip generosity would have reached. Six watches upgraded into a visual refresh. Two shipped with a product.

    And the Double Tap precedent tells you how long that lasts. Three years after Apple named the S9 SiP, Double Tap still does not exist on a Series 8 that has every sensor it needs. When Apple names the hardware in the announcement, the gate is a wall.

    Which reframes the question to ask a new Apple Watch. Not what it does — what sensor generation it is, and therefore how many of these announcements it will be on the right side of.

    Source: Apple Newsroom — health and fitness capabilities · About watchOS 27 Updates

    Photo: Simon Daoudi / Unsplash

  • TL;DR — 24 hours, 30 hours, 12 days. Those are the three headline battery figures, and putting them in one table is close to meaningless: Apple publishes the exact test script, Samsung’s number is quoted with the always-on display running, and Garmin’s manual does not say whether the display was on, off, or anything else. Two other things the spec sheet hides: the Venu 4 is a 2025 product competing with two 2026 ones while costing the most, and the Galaxy Watch 9’s headline health features need not just an Android phone but a Samsung phone. The spec that decides this purchase is which phone is already in your pocket.

    The Story

    I set out to build a clean three-way table and could not. Not because the numbers are hard to find — because the numbers are not commensurable, and nobody says so.

    Every roundup you will read puts 24 hours, 30 hours and 12 days in adjacent cells as though they answer the same question. They do not. One of those figures comes with a published, auditable test protocol. One comes with a single sentence of hand-waving. One is measured in a display mode the others were not. The gap between 24 hours and 12 days is real, but it is nothing like 12 times.

    So this comparison is organised differently. Where a number is solid, I will say where it came from. Where it is not, I will say that too — because for a purchase this size, knowing which figures you can lean on matters more than having a full table.

    The Age Gap Nobody Puts in the Table

    Start here, because it reframes everything else.

    • Apple Watch Series 12 — announced September 9, 2026.
    • Galaxy Watch 9 — announced July 22, 2026, on sale August 7.
    • Garmin Venu 4 — announced September 17, 2025.

    The Venu 4 is a full product generation older than the other two. It is also, at $549.99, the most expensive of the three by a margin — $150 more than the Series 12’s $399 starting price, $170 more than the 40mm Bluetooth Galaxy Watch 9 at $379.99. And it has no cellular option at any price.

    That is not automatically a mark against it. Garmin’s release cadence is slower because Garmin’s hardware changes less, and as you will see below, not changing the hardware turns out to have a concrete regulatory payoff. But if you are comparing on price, you are comparing a year-old device against two current ones, and every roundup I checked left that out.

    The Spec Sheet, With Its Sources Shown

    A note on the middle column before you read it. Apple publishes a full specification page that reads like a document. Samsung’s US site carries some figures — thickness, price, the update commitment — but its dedicated specification URL does not resolve and the rest of the sheet is locked inside JavaScript-rendered product pages. Garmin’s product pages are also JavaScript-rendered and its support site refuses automated requests outright, though its owner’s manual is plain HTML and, as you will see, the single most useful document of the three. So the evidence quality is genuinely uneven, and I have marked it rather than smoothing it over.

    Series 12Galaxy Watch 9Venu 4
    AnnouncedSept 2026July 2026 (on sale Aug)Sept 2025
    Sizes42mm, 46mm (ceramic 43/47mm)40mm, 44mm41mm, 45mm
    Thickness9.7mm8.6mm°
    Weight32.2 g (42mm Al, GPS) / 39.5 g (46mm Al, GPS)31.5 g / 34 g °33 g / 38 g, case only °
    CaseAluminum, titanium, or ceramicArmor aluminum °Fiber-reinforced polymer, metal bezel °
    DisplayLTPO3 OLED, 374×446 / 416×496Super AMOLED, 438×438 / 480×480 °AMOLED, 390×390 / 454×454 °
    Peak brightness2,000 nits (min 1 nit)3,000 nits °2,000 nits °
    ChipS11, dual-core, 4-core Neural EngineSnapdragon Wear Elite, penta-core 3nm °Not disclosed by Garmin
    Storage / RAM64 GB / not disclosed32 GB / 2 GB °8 GB ° / not disclosed
    GNSS“Precision L1 GPS” (GPS, GLONASS, Galileo, QZSS, BeiDou) — single-frequencyL1+L5 dual-band °L1+L5 multi-band with SatIQ °
    ECGYesYes (Samsung phone required)Yes, “not available in all regions”
    Blood pressureHypertension notifications — see belowYes, US, wellness not FDA-cleared, cuff calibration requiredNo
    Body compositionNoYes (BIA) °No
    Water50m ISO 22810 + IP6X, depth to 6m5ATM + IP68 + MIL-STD-810H °5 ATM
    Phone supportiPhone onlyAndroid onlyiPhone or Android
    US price from$399$379.99 (Bluetooth)$549.99
    Update promiseNone published5 yearsNone published

    ° = figure could not be confirmed on the manufacturer’s own site and comes from secondary sources. Unmarked rows were read off the manufacturer directly. Treat the marked Samsung and Garmin rows as softer than anything in the Apple column — not wrong, just unverifiable at the source, which is a different kind of number to buy on.

    What Samsung does publish where anyone can read it is instructive: the price, the thickness, and the five-year update commitment — $459.99 for the 44mm LTE sits on its own buy page. The display resolution, the RAM, the battery footnote do not. That tells you which parts of a spec sheet a company considers load-bearing, and it is not the parts an engineer would pick.

    The Battery Numbers Measure Three Different Things

    Here is the part that actually matters, and it takes a minute.

    Apple: 24 hours normal, 38 hours Low Power, 10 hours outdoor workout. Apple publishes the script. This is the specification page’s wording, verbatim:

    > “All-day battery life testing includes: 300 time checks, 90 notifications, 15 minutes of app use, a 60-minute workout with music playback from Apple Watch via Bluetooth, and 6 hours of sleep tracking over the course of 24 hours.”

    You can argue with whether that resembles your day. You cannot argue about what was measured. And a detail worth catching: the 38-hour Low Power figure is not the same test run longer — it is a separate test with more interactions packed into it.

    Garmin: 12 days smartwatch mode (45mm), 10 days (41mm). Garmin’s owner’s manual publishes seven distinct modes, which is more granularity than either rival offers:

    Mode41mm45mm
    Smartwatch modeUp to 10 daysUp to 12 days
    Battery saver watch modeUp to 18 daysUp to 25 days
    GPS only modeUp to 15 hoursUp to 20 hours
    All satellite systems modeUp to 13 hoursUp to 19 hours
    All satellite systems with music modeUp to 6 hoursUp to 9 hours
    All satellite systems plus multi-band modeUp to 12 hoursUp to 17 hours
    All satellite systems plus multi-band with music modeUp to 6 hoursUp to 9 hours

    The two music rows are not a duplication — they are different modes that happen to carry identical figures, which is its own small tell about how finely these numbers are actually resolved.

    And here is Garmin’s statement of conditions, in full:

    > “The actual battery life depends on the features enabled on your watch, such as wrist-based heart rate, phone notifications, GPS, internal sensors, and connected sensors.”

    That is the whole thing. No interaction count, no display setting, no pulse-ox sampling rate. Garmin has not published a battery methodology anywhere a buyer can read it, and the manual is the most detailed document it offers.

    The display setting is the gap that matters most, and I want to be careful about it. Garmin’s specification page does not state whether the always-on display was enabled for the twelve-day figure. Reviewers who have tested the watch with always-on enabled report getting a fraction of the headline number — the figures circulating land around a third of it — but that is measured by third parties, not published by Garmin, and Garmin has not confirmed the test condition either way. Which is exactly the problem: the most consequential variable in the number is the one nobody will put in writing.

    Samsung: up to 30 hours, quoted with the always-on display enabled °. That is the opposite convention to how Garmin’s figure is usually read — but note the marker. I could not retrieve Samsung’s measurement footnote from Samsung, so even the condition attached to the condition is secondhand.

    So the honest summary is this. Garmin genuinely lasts far longer than the other two — that is not in doubt and it is the main reason to buy one. But “12 days versus 24 hours” is not the comparison. If the third-party always-on figures are right, the realistic gap for someone who wants a watch face they can glance at is closer to a few days against one: still a decisive win, and a very different number from the one on the box. The finding is that you cannot check this. Two of the three vendors do not publish enough to let you normalise their own figures.

    The Deciding Spec Is Your Phone

    None of the above will determine your purchase. This will.

    • Apple Watch Series 12 requires an iPhone. iPhone 11 or later on iOS 27. There is no Android support and there never has been. One restriction, clearly stated.
    • Garmin Venu 4 works with either. “A compatible iPhone or Android smartphone.” It is the only genuinely cross-platform option here.
    • Galaxy Watch 9 requires Android — and its health features require a Samsung. This is a two-tier restriction and it is the one people get wrong.

    Unpack that last one, because it is the single most consequential fact in this comparison. There are two separate requirements here and they are easy to run together. The first is the watch itself: the Galaxy Watch 9 runs Wear OS and pairs with a phone on Android 13 or newer, from any manufacturer. The second is Samsung Health Monitor — the app that delivers ECG, irregular heart rhythm notifications and blood pressure, which are the features Samsung markets hardest. That app states its own, different requirement: “a Galaxy smartphone with Android 12 or later.” That number is Samsung’s, and it is also a trap — the Watch 9 will not pair below Android 13 in the first place, so the effective floor for ECG and blood pressure is Android 13 on a Galaxy phone, not Android 12 on anything.

    So the watch works on a Pixel and the headline health features do not. Be precise about the size of that loss, because it is routinely overstated in both directions. What you actually lose is the Samsung Health Monitor set — ECG, irregular rhythm notifications, blood pressure. Samsung Wallet, notifications and the general Samsung Health tracking continue to work in some form on non-Samsung Android, with feature availability varying by region and by what the phone supports. The clean, verifiable statement is the narrow one: buy this watch for its cardiac features on a non-Samsung phone and those specific features will not run.

    So the cheapest watch here carries the most lock-in, and the most expensive carries the least. You are paying roughly a $150 premium for the Venu 4 and part of what you are buying is the freedom to change phones without rebuying a watch.

    Worth saying plainly: Apple’s restriction is stricter in absolute terms — no Android at all — but it is also completely unambiguous. Nobody buys an Apple Watch expecting it to work with a Pixel. People absolutely do buy a Galaxy Watch expecting it to work fully with one.

    The New Apple Watch Is Behind the Old One

    The regulatory picture is the strangest thing in this comparison, and it runs the opposite way to intuition.

    From Apple’s own support document, verbatim:

    > “Hypertension Notifications might not be available on Apple Watch Series 12 or Apple Watch Ultra 4 in your region, as additional regulatory clearances for the feature for these models are in process.”

    Apple redesigned the heart sensor, and a redesigned heart sensor resets the regulatory clock in every market that cleared the old one. So in a substantial number of countries the brand-new watch does less, health-wise, than the one it replaces. I went through the full shape of this in the Series 12 comparison — and the Ultra line inherited the same problem. Apple names no countries; the tallies circulating come from press counts.

    Now hold that next to Garmin. Garmin’s ECG app has accumulated clearances across a long list of markets — the US, the UK, the EU, Australia and Switzerland among many others — on hardware that has not been redesigned. Its stability is a direct consequence of the slower hardware cadence that makes the Venu 4 look dated on paper. Garmin still says “the ECG app is not available in all regions,” so check your own, but it is not waiting on a clock it restarted itself.

    And Samsung’s blood pressure feature, which reached the US market well after its debut elsewhere, is worth reading carefully: Samsung positions it as a wellness feature, not an FDA-cleared medical one, and it requires calibration against a real cuff. A cuff you already own, using a watch to interpolate between calibrations. That is a legitimate design and it is not the same thing as a cleared blood-pressure device.

    If cleared cardiac features are the reason you are buying, that ordering — Garmin stable, Apple in flux, Samsung’s flagship BP explicitly not cleared — is close to the inverse of the brand hierarchy most people assume.

    Nobody Paywalls Your Heart Rate

    Briefly, because it is a common worry and the answer is reassuring.

    • Garmin Connect+ is $6.99/month or $69.99/year, and Garmin’s own release states: “All existing features and data in Garmin Connect will remain free.” The paid tier is AI insights and dashboards.
    • Apple Fitness+ is $9.99/month or $79.99/year and is entirely optional. No Series 12 health feature sits behind it.
    • Samsung Health has no mandatory subscription.

    So none of the three charges you for a core health metric. But notice that Samsung’s gate is not a fee — it is the hardware lock described above. Free, and available only if you keep buying Samsung phones. That is a cost, just not one with a price tag.

    Who Should Actually Buy Which

    • You own an iPhone and want the best all-round watch. Series 12, and it is not close — the competition is iPhone-incompatible or, in Garmin’s case, a deliberately narrower product. Check whether hypertension notifications are cleared in your country first; if they are not and that feature matters, a discounted Series 11 is a rational purchase.
    • You own a Galaxy phone. Galaxy Watch 9. You get the full feature set, the brightest display of the three, body composition nobody else offers, and the only published update commitment — five years. It is also the cheapest. On a Samsung phone this is a genuinely strong buy.
    • You own a non-Samsung Android phone. Venu 4, or a Pixel Watch, but not a Galaxy Watch 9 — you would pay for sensors whose headline app refuses to run.
    • You train seriously, or you resent charging things. Venu 4. Multi-band GNSS, days of battery instead of hours, and the most stable regulatory footing. Accept that you are buying 2025 hardware at the highest price of the three, and that nobody — Garmin included — will tell you what display setting the twelve-day figure assumes.
    • You expect to switch platforms in the next few years. Venu 4 is the only one of the three that survives the move.

    The Takeaway

    The three most-quoted numbers in this comparison — 24 hours, 30 hours, 12 days — are produced by three incompatible methodologies, and only one vendor publishes enough for you to check. That is the practical lesson: a spec sheet is a document written by a marketing department under legal supervision, and comparing across companies means comparing their disclosure policies as much as their hardware.

    The decision itself is simpler than the table makes it look. Your phone picks your watch. Everything above is about what you get to notice once that constraint has already chosen for you — and about which company was willing to show you its work.

    Source: Apple Watch Series 12 Tech Specs · Garmin Venu 4 Owner’s Manual · Samsung Galaxy Watch 9

    Photo: Dawn Patrol Surf Tracking / Unsplash

  • TL;DR — Apple published an eleven-page privacy paper for Audio Intelligence, and it is the only document that tells you two things worth knowing before you buy. Live Rewind is not really a watch feature: press the crown with your phone out of range and the transfer “will fail with an error,” and the fifteen seconds are discarded. And activating it fires a chime from the watch speaker that plays even on silent, even with headphones connected. The paper is also where you learn that raw audio — not a transcript — is what crosses from your wrist to your phone for Siri Recap. None of it is hidden. It is just eleven pages deep instead of on the page you were reading.

    Audio Intelligence is only one of three separate hardware floors in this release — I mapped all of them in which watches get watchOS 27 and which get its features.

    The Story

    The interesting question about a wrist microphone that can replay the last fifteen seconds is usually is the thing always on? I assumed Apple had dodged it, because the product page and the support document both dance around the word “always.”

    Apple did not dodge it. It answered in a document called the Audio Intelligence Privacy Overview, dated September 2026, eleven pages, linked from nowhere you would naturally land. The answer, verbatim:

    > “The S11 chip uses a lightweight AI model to determine whether speech is occurring nearby. This model doesn’t transcribe or record anything. It simply detects whether a conversation has started. If a conversation is detected, audio flows into a protected buffer inside the Secure Exclave on the S11 chip.”

    So: yes, something is listening continuously, and what it is doing is a binary speech/no-speech decision that never becomes text. That is a clear, specific, falsifiable architectural claim, and it is more than most of this industry offers.

    Which makes the rest of the paper the story. Because while Apple was answering the question everyone asks, it also documented several things nobody has asked, and those are the ones that change whether you should buy the watch.

    What’s Actually in the Box

    Apple’s footnote, which appears identically on the spec page, the product page, and both newsroom releases, defines the bundle:

    > “Audio Intelligence includes Live Rewind, Siri Recap, Sound Recognition, and faster Shazam and requires Apple Watch Series 12 or Apple Watch Ultra 4.”

    Four things, not three, and one of them is not new:

    • Live Rewind — double-press the Digital Crown, get the last 15 seconds as text. You can save it or ask Siri about it.
    • Siri Recap — rolling summaries of the conversations you have during the day, in the new Siri app. They auto-delete after seven days unless you save them.
    • Sound Recognition — detects “important sounds like sirens, alarms, and doorbells,” and notifies you “even when your iPhone is not with you.”
    • Faster Shazam — Shazam has been on Apple Watch for years. What changed is that it “instantly detects the music playing around you” and populates the Smart Stack widget “all without a tap.”

    Note the shape of that last one. Same feature, converted from something you invoke into something that runs on its own. MacRumors ran it under the headline “Three All-New ‘Audio Intelligence’ Features,” which is understandable and wrong — it is four, one of which is an old feature that stopped waiting to be asked.

    The hardware requirements are not uniform, and the split is the first clue about what runs where:

    FeatureWatchiPhone
    Sound RecognitionSeries 12 / Ultra 4iPhone 11 or later, or SE (2nd gen or later), on iOS 27 — but alerts fire “even when your iPhone is not with you”
    Faster ShazamSeries 12 / Ultra 4iPhone 11 or later, or SE (2nd gen or later), on iOS 27
    Live RewindSeries 12 / Ultra 4iPhone 16 or later, excluding the 16e — and it must be in wireless range at the moment you press
    Siri RecapSeries 12 / Ultra 4iPhone 16 or later, excluding the 16e

    Every feature lists an iPhone, but they are not the same kind of requirement. Sound Recognition and Shazam need a phone that is merely paired — an iPhone 11 clears the bar, and Apple explicitly says Sound Recognition alerts you when the phone is elsewhere. Live Rewind and Siri Recap need an Apple Intelligence phone, and the cutoff has one oddity: Apple’s list includes the iPhone 17e but leaves out the 16e.

    Older watches running watchOS 27 get none of it. This is a hardware gate, not a software one.

    Live Rewind Is an iPhone Feature Wearing a Watch

    This is the part that should be in every review and is in none of them.

    Apple’s support page does list iPhone models as a requirement, so the dependency is not concealed. What no page outside the privacy paper tells you is what happens when the phone is not there:

    > “After activating, the immediately preceding 15 seconds of buffered audio is sent from the Secure Exclave of the Apple Watch to iPhone. If iPhone is not within wireless range of Apple Watch, the transfer will fail with an error, and the audio will be immediately discarded.”

    The watch holds the buffer. The watch does not do the transcription. Your phone does. Press the crown on a run, on a walk, in a meeting you left your bag for — anywhere your phone is not — and the feature does not degrade gracefully. It throws an error and deletes the fifteen seconds you were trying to recover.

    Compare that to Sound Recognition, which the same document says works “even when your iPhone is not with you,” and which for the deaf and hard of hearing is the feature that matters most. Apple built one Audio Intelligence feature that is genuinely independent and one that is not, and the independent one is the accessibility feature. That is the right way round, and it is worth saying so.

    But if you are buying a Series 12 for Live Rewind, understand that you are buying a two-device feature. The watch is the microphone and the display. The phone is the computer.

    Raw Audio Does Leave Your Wrist

    Here is Apple’s public privacy claim, the one on the marketing pages:

    > “These features do not create or store audio recordings, and raw audio used for processing is completely inaccessible, even to Apple, and is deleted after processing.”

    Every word is true. And almost everyone will read “inaccessible” and “deleted after processing” as “it stays on the watch and dies there,” which is not what it says. From the paper:

    > “For Siri Recap, the Secure Exclaves on Apple Watch and iPhone establish a secure, encrypted channel during the secure pairing process. Then, raw audio is encrypted and transferred from the Secure Exclave on Apple Watch to the Secure Exclave on the paired iPhone.”

    Raw audio — not a transcript, not a summary, the audio — travels from your wrist to your pocket every time Siri Recap decides a conversation is happening.

    Apple’s support article does mention the hop, and it is worth being precise about how. It says “the Secure Exclave on Apple Watch encrypts the detected audio and transmits it to the Secure Exclave on the paired iPhone.” Detected audio. The paper is the document that says raw, and that describes the encrypted channel, the pairing mechanism that constrains where it can go, and what happens when the transfer fails. Three documents, three degrees of specificity, and the product page — the only one most buyers will ever open — carries none of it.

    To be clear about the engineering: this is well built. The audio is encrypted end to end between two hardware compartments neither operating system can address. Apple added an audio-verified pairing mechanism on top of Bluetooth specifically so that “ambient data from the Secure Exclave on Apple Watch can only be sent to the Secure Exclave of the paired iPhone, and no other device.” The keys are device-bound, time-bound, rotate, and expire — and if the transfer cannot complete, “the encryption keys expire and the encrypted audio is permanently deleted.” Find My invalidates the pairing keys immediately.

    That is a better design than “it never leaves the watch” would have been, because the watch does not have the silicon to transcribe speech well and Apple chose not to pretend otherwise. But it is a different sentence from the one on the product page, and the difference is the whole point of reading the paper.

    The Sentence Written for a Courtroom

    In the introduction, this:

    > “Audio Intelligence features do not create audio recordings that can be accessed by the operating system, apps, the user, or Apple. There is no recording to share, forward, or produce if requested by any party, because no recording exists.”

    Read the second sentence again. “Produce if requested by any party” is not consumer language. That is discovery language — the vocabulary of subpoenas, warrants and civil litigation. Apple is telling courts, in a document aimed at customers, that there is nothing to compel.

    It is a legitimate and in my view admirable thing to design for. Architecture beats policy: a company that cannot produce your audio does not have to be trusted not to. But it is striking to see it stated that plainly, and it tells you which risk Apple’s lawyers were thinking about when this feature was specced.

    The Chime You Cannot Turn Off

    I had assumed Apple shipped no bystander mechanism. The opposite is true, and the mechanism is unusually aggressive:

    > “When activated, an audible chime plays from the speaker on your Apple Watch, even if your Apple Watch is on silent or you have headphones connected. A full-display animation and microphone indicator appear on the Apple Watch display, and a distinct double-press gesture is required to activate Live Rewind.”

    Silent mode does not suppress it. Routing audio to headphones does not suppress it. Apple deliberately removed every ordinary way a user could quiet their own device, in order to make the feature impossible to use discreetly.

    Think about what that means in practice. You cannot Live Rewind a lecture without the room hearing it. You cannot Live Rewind in a meeting without your colleagues knowing. That is a real product constraint, chosen on purpose, and it is the single most consequential fact about how this feature will feel to own — and it appears on page nine of a PDF rather than anywhere near the buy button.

    Siri Recap took the opposite approach and is the weaker case for it. It runs in the background with no chime and no indicator. Apple’s argument is that the output is not a recording and not a verbatim transcript:

    > “By design, Siri Recap does not create a recording, does not produce a verbatim transcript, and does not identify and attribute speakers. The output is a brief, high-level summary, comparable to notes a person might write after a conversation.”

    That is a reasonable defence and it is not nothing — a summary of a conversation is categorically different from a recording of one. It is also still a machine-generated account of what the people around you said, produced without their knowledge, and in two-party-consent jurisdictions the legal analysis of “notes a person might write” has not been done yet.

    Where Your Conversation Actually Goes

    The paper includes a stage-by-stage map for Siri Recap, which is worth compressing:

    1. Speech detection — S11 Always On Processor. Raw audio “processed on-device and discarded in real time.” 2. Buffered audio on the watch — Secure Exclave, inaccessible to watchOS. 3. Buffered audio on the phone — encrypted in transit, decrypted only inside the iPhone’s Secure Exclave. 4. On-device transcription and condensing — speech recognition plus an on-device language model that produces “a condensed version that is less than half the length of the original transcript,” stripping tone, filler and redundancy. Then “the raw audio is immediately and permanently deleted and no longer exists on either Secure Exclave.” A safety model screens for harmful terms. 5. Private Cloud Compute — the condensed text is encrypted and sent up, “subject to daily usage limits.” Apple Foundation Models write the title and key points. “The results are immediately deleted from Private Cloud Compute.” 6. The Siri app — end-to-end encrypted across your devices, seven-day auto-delete unless saved.

    Two things in that chain deserve flagging.

    Apple also sends context alongside the text: Now Playing data, calendar data, “high-level location labels such as home, work, and school,” city and state, and point-of-interest categories like “grocery store” or “park.” Precise location is excluded. That is a defensible list for improving a summary, and it is more metadata than “an encrypted transcript” implies.

    And the retention question I thought was unanswered is answered — the summarization “is then encrypted and sent back to iPhone. The results are immediately deleted from Private Cloud Compute.” It just isn’t answered anywhere a buyer will read.

    Sound Recognition Is a Six-Year-Old Accessibility Feature

    Sound Recognition is the oldest thing in the bundle. It shipped in iOS 14 in 2020, buried under Accessibility, detecting alarms, sirens, doorbells and crying babies for people who cannot hear them. What is new is that the watch does it natively, in the Secure Exclave, without an iPhone.

    Google’s Sound Notifications also arrived in 2020, with the flashlight strobe and an on-screen alert as its main outputs and the option to relay to a compatible Wear OS watch. Either way the detection ran on the phone and the watch was a display. Apple’s runs on the watch. On the underlying capability Apple is late; on this specific implementation it is not behind.

    Two practical notes. Samsung’s Hearing Health measures ambient noise levels for hearing protection — dosimetry, not semantic recognition, so it is not a comparison. And Apple’s long-standing guidance for Sound Recognition on iPhone is blunt: “Don’t rely on your iPhone to recognize sounds in circumstances where you may be harmed or injured, in high-risk or emergency situations, or for navigation.” That warning is in the iPhone user guide. The Apple Watch support article for Audio Intelligence carries no equivalent line, and no accuracy or false-positive rates have been published for the watch implementation. If you are deaf or hard of hearing and evaluating this as a safety system, that is the most important paragraph in this article — and the absence of the warning on the watch page is not a reason to assume it does not apply.

    The Part Where It Costs Money

    Buried in the newsroom footnotes:

    > “Certain Audio Intelligence features that rely on server-side models are subject to daily usage limits, including but not limited to Live Rewind and Siri Recap.”

    And in a separate footnote, this one covering Apple Intelligence and Siri broadly rather than Audio Intelligence specifically:

    > “Expanded access to such features will be available for a fee in the future.”

    The privacy paper independently confirms the limit — the Private Cloud Compute stage is described as “subject to daily usage limits” — without quantifying it either. So the cap is real, it is documented in two places, and its size is stated in neither.

    Credit where it is due: Apple disclosed a paywall in advance rather than introducing one quietly later. But if you are buying a Series 12 partly for Live Rewind, you are buying a feature whose free tier has an unstated ceiling.

    What This Means If You Wear One

    • Live Rewind needs your phone in range, every single time. No phone, no feature, and the audio is discarded rather than queued. If you wanted this for phone-free runs, it does not do that.
    • It also makes a noise you cannot mute. Chime plus full-screen animation, overriding silent mode and headphones. Discreet use is not a supported scenario, by design.
    • “Raw audio is inaccessible” is true; “raw audio stays on the watch” is not. For Siri Recap it moves to your phone, encrypted, Exclave to Exclave. Good architecture, different claim.
    • Siri Recap sends more than a transcript. Condensed text plus calendar, Now Playing, and coarse location context go to Private Cloud Compute.
    • Do not treat Sound Recognition as a safety system. Apple’s own guidance says not to, and no accuracy data exists for the watch version.
    • The other people in the room get a chime for Live Rewind and nothing for Siri Recap. Apple thought hard about bystanders in one feature and made a policy argument in the other.

    The Takeaway

    Three documents describe this product, and they get more honest as they get harder to find. The product page says raw audio is inaccessible and deleted after processing, and leaves you to infer that it never moves. The support article admits the hop and calls what moves “detected audio.” The privacy paper calls it raw audio, names the channel, and is the only one of the three that tells you Live Rewind dies without your phone and announces itself with a chime you cannot silence. It is also the one at a URL you will never stumble onto.

    That is the pattern worth naming, because it is now the third time in one launch — after the Readiness score with no published validation and a 60x sampling claim with no denominator — that the interesting part of the story was a document rather than a feature. The difference this time is that the document exists and is good. Apple did the hard engineering, wrote it down honestly, and then filed it where nobody making a purchase decision will find it.

    If a company is going to build ambient microphones into a wrist computer, this is close to how you would want it done. It would just be better if the eleven pages describing it were the thing you read before buying, rather than the thing you find afterwards.

    Source: Apple — Audio Intelligence Privacy Overview (September 2026) · Apple Support — Audio Intelligence on Apple Watch

    Photo: Luke Chesser / Unsplash

  • TL;DR — Same 49mm titanium case, same 3000-nit display, same $799, same everything you can see. What changed is the battery (42 → 50 hours), the chip (S10 → S11), the sensor array, and one thing nobody puts in a comparison table: the Ultra 4’s spec sheet no longer lists GLONASS. The Ultra 3 supported five satellite constellations. The Ultra 4 supports four. Apple has not said why, and the Series 12 kept GLONASS — so this is an Ultra-specific decision. Apple’s own comparison tool cannot show you this, because it names no constellations for either watch.

    If you are weighing the Ultra line against what Samsung and Garmin offer, I put the three side by side in a three-way comparison against the Galaxy Watch 9 and Garmin Venu 4.

    The Story

    Most generational comparisons are an exercise in finding what got added. This one is more interesting for what got removed, and for the fact that Apple built a comparison tool that structurally cannot surface the removal.

    The Ultra 4 is, physically, an Ultra 3. Identical case dimensions, identical display, identical water rating, identical price. Apple spent this generation on three things: a bigger battery, a new chip, and the same redesigned heart sensing system that went into the Series 12. That is a defensible, focused release.

    And somewhere in that process, a satellite constellation fell off the list.

    The Spec Sheet, Honestly

    Ultra 4 figures come from Apple’s spec page. Ultra 3 figures come from Apple’s support tech-specs document, which is where Apple keeps the specifications for models it no longer sells.

    Ultra 3Ultra 4
    Case49 × 44 × 12 mm, Grade 5 titanium49 × 44 × 12 mm, Grade 5 titanium
    Weight (natural / black)61.6 g / 61.8 g63.0 g / 63.1 g
    Display422×514, 326 ppi, 3000 nits peak422×514, 326 ppi, 3000 nits peak
    ChipS10 chipS11 chip, 64-bit dual-core processor
    Storage64 GB64 GB
    Battery, normal useUp to 42 hoursUp to 50 hours
    Battery, Low Power ModeUp to 72 hoursUp to 84 hours
    Battery, outdoor workoutNot listed by AppleUp to 18 hours
    Battery, Extended WorkoutNot listed by AppleUp to 25 hours
    Battery, Max Extended WorkoutNot listed by AppleUp to 45 hours
    GNSS“L1 and L5 precision dual-frequency GPS (GPS, GLONASS, Galileo, QZSS, and BeiDou)”“Precision dual-frequency GPS (GPS, Galileo, QZSS, and BeiDou)”
    Electrical heart sensorNo generation listed by AppleSecond generation
    Optical heart sensorThird generationApple lists no generation number
    Water / dive100m, ISO 22810; EN13319 to 40m100m, ISO 22810; EN13319 to 40m
    US starting price$799 at launch$799

    Two notes before anyone builds a purchase decision on that table.

    The Ultra 3’s workout-specific battery rows are genuinely absent from Apple’s document, not something I failed to find. Apple simply did not publish outdoor-workout hours for that model, so the table cannot be made symmetric without inventing numbers. The Ultra 3’s $799 is from launch coverage rather than from Apple’s own page — Apple strips pricing from the specification pages of models it has stopped selling.

    And the price line is doing something unusual. $799 then, $799 now. This is not a comparison between price tiers. It is a comparison between two watches that cost the same, which means the only question is whether the delta is worth the money you would spend either way.

    The Constellation That Disappeared

    Here are the two strings, copied from Apple’s own pages.

    Ultra 3:

    > “L1 and L5 precision dual-frequency GPS (GPS, GLONASS, Galileo, QZSS, and BeiDou)”

    Ultra 4:

    > “Precision dual-frequency GPS (GPS, Galileo, QZSS, and BeiDou)”

    Four systems remain, in the same order. One token is gone: GLONASS, the Russian global navigation system. Something smaller went with it — the Ultra 3 named its two frequencies (“L1 and L5”) and the Ultra 4 just says “dual-frequency,” so the band detail dropped out of the sentence at the same time. That is the whole change, and Apple documented it in the only place it appears — the fine print of two spec pages.

    Apple has published no explanation. Not in the newsroom release, not in a support document, not in a footnote. GLONASS is the technical outlier of the four: it separates satellites by frequency rather than by code, so a receiver has to carry inter-frequency bias corrections that the other three do not require. That is a real difference and it is the one people reach for when they speculate. It is speculation all the same — nobody has connected it to Apple’s decision, least of all Apple. Nobody outside Apple knows whether this was about accuracy, silicon area, licensing, supply chain, or something duller.

    What makes it worth writing down is the asymmetry: the Series 12 kept GLONASS. Its spec page reads “Precision L1 GPS (GPS, GLONASS, Galileo, QZSS, and BeiDou)” — one frequency, five constellations, where the Ultra 4 has two frequencies and four. So this was not a company-wide decision to simplify the GNSS stack. It was made specifically for the watch Apple sells to people who navigate for a living, or at least buy it believing they might, and the two product lines now sit on opposite sides of the same tradeoff: the cheaper watch sees more satellites, the expensive one hears each of them on more than one channel.

    How much does it matter? Honestly, for most people, probably not much. Four constellations is still a lot of satellites, and dual-frequency L1+L5 reception does more for real-world accuracy than constellation count does — L5’s higher power and longer codes are what cut through urban multipath. If you run in a city, that is the feature carrying your track, and the Ultra 4 is still dual-frequency — though Apple no longer spells out which two bands.

    But “probably not much” is a different sentence from “Apple explained the tradeoff,” and only one of those two things is available to you.

    Why Apple’s Comparison Tool Can’t Show You This

    This is the part I did not expect.

    Go to Apple’s watch comparison page, put the Ultra 3 and Ultra 4 side by side, and look at the navigation row. The Ultra 3 reads “Precision L1 + L5 dual frequency.” The Ultra 4 reads “Precision dual-frequency GPS.” Neither entry names a single constellation.

    So the tool does show you a difference. It shows you the wrong one. A buyer diffing those two lines learns that Apple stopped printing the band names, which is a copywriting change, and learns nothing at all about the constellation that actually left. The one real, documented, Apple-published subtraction in this generation is the single thing the comparison interface has no field for.

    I do not think this is a trick. Comparison tables get abbreviated, and “L1 + L5 dual frequency” is arguably the more meaningful summary for a buyer. But it is a useful reminder about where the truth lives. The compare page is a marketing surface. The spec page is the document of record, and the two do not contain the same facts.

    If you take one practical habit from this article: when a spec matters to you, read the spec pages for both models and diff them yourself. The comparison tool is a convenience, not an authority.

    The Battery Number Has Three Asterisks

    50 hours is real and it is Apple’s own figure — the newsroom release says “For everyday use, Apple Watch Ultra 4 now offers more than two days of battery life, up to 50 hours.” Up from 42. That is a genuine 19% gain in the headline number, in the same case, which almost certainly explains the 1.4 grams the watch gained.

    The workout figures need more care, because Apple published three of them and they are not three tiers of the same thing.

    • 18 hours — outdoor workout, full fidelity. This is the honest number.
    • 25 hours — Extended Workout, which turns on Low Power Mode. Apple’s footnote for this mode says “running form metrics will be unavailable.”
    • 45 hours — Max Extended Workout, which does that and turns off alerts and splits. Its footnote is vaguer than the one above: “certain metrics will be unavailable,” without saying which.

    So “45 hours of GPS tracking” is technically true and practically misleading. You get 45 hours by asking the watch to stop doing most of the things you bought it to do during a workout. For an ultramarathon or a multi-day hike that is exactly the right trade. For comparing against a Garmin’s stated GPS endurance, it is not a like-for-like number.

    Apple’s own test protocol, from its footnote, is also worth knowing: 630 time checks, 190 notifications, 30 minutes of app use, two 60-minute workouts “with music playback from Apple Watch via Bluetooth,” and 12 hours of sleep tracking, across 50 hours. Apple ran it on preproduction hardware paired with an iPhone, on prerelease software. As always, that is a constant for comparing generations, not a forecast of your week.

    The Upgrade That Takes Something Away

    The Ultra 4 inherits the Series 12’s regulatory problem, for the same reason: Apple redesigned the heart sensor, and a heart sensor is the part a medical regulator has jurisdiction over.

    From Apple’s support document, verbatim:

    > “Hypertension Notifications might not be available on Apple Watch Series 12 or Apple Watch Ultra 4 in your region, as additional regulatory clearances for the feature for these models are in process.”

    The restriction names the new models. An Ultra 3 owner in an affected region who upgrades gives up a working hypertension feature until Apple clears the new hardware locally. I went through the full shape of this in the Series 12 comparison, and it applies identically here.

    One correction to the way this gets reported, including by me: Apple’s document names no countries. It points at a general availability page. The widely repeated count of affected regions comes from press tallies, not from Apple. The restriction is real and Apple-stated; the number attached to it is not Apple’s.

    Who Should Actually Upgrade

    • Ultra 3 owner who does multi-day efforts — this is the real case. Eight more hours of normal use and published extended-workout modes are the entire value of the generation, and they land precisely on the people who bought an Ultra for the right reasons.
    • Ultra 3 owner in a region where hypertension notifications are affected — wait. You would pay $799 to lose a cleared health feature for an unspecified period.
    • Ultra 3 owner who wants Readiness — this is a genuine hardware gate, not segmentation; the score needs the higher sampling rate. Just know the score itself has no published validation.
    • Ultra 3 owner who navigates in places where satellite coverage is thin — you are the one person for whom the GLONASS line is not trivia. It probably still does not change your day, but you should make that call with the spec pages open rather than the comparison tool.
    • Coming from an Ultra 2 or a Series watch — the accumulated gap is real and the decision is easy. At the same $799 the Ultra 3 launched at, the Ultra 4 is the one that will be supported longest.

    The Takeaway

    The Ultra 4 is a good, narrow update: more battery, better sensor, same excellent body, same price. If you want the short version, that is it, and it is a fair deal.

    The longer version is about how you find things out. Apple published the GLONASS removal accurately and completely, in the place where specifications belong, and then built a comparison interface that cannot express it. Both of those are normal corporate behaviour. Together they produce a situation where the company is entirely truthful and the average buyer still cannot see what changed.

    That gap is going to keep widening as watches accumulate features that live in footnotes — regulatory clearances that vary by country, scores gated on sampling rates, satellite systems that appear and disappear. The spec sheet is still honest. It is just no longer the surface most people read.

    Source: Apple Watch Ultra 4 Tech Specs · Apple Watch Ultra 3 Tech Specs

    Photo: Shutter Speed / Unsplash

  • TL;DR — The Series 12’s headline sensor claim is “60x more frequent background heart rate readings than Series 11.” Apple has never published how often the Series 11 took a background reading, so the number has no denominator you can check. Apple also describes the physical change in exactly one recurring phrase — “larger, power-efficient green LEDs” — publishes no component counts, and quietly dropped the generation number it used to print on the optical sensor. Meanwhile its own white paper says the five-second readings you see on the watch face reach HealthKit only every 30 seconds. The hardware almost certainly did improve. Apple just replaced the description of it with a ratio.

    The same S11 chip runs the watch’s new ambient microphone features, and Apple documented those far more thoroughly — the Audio Intelligence privacy paper spells out what the mic is doing.

    The Story

    There is a specific way to read a launch page, and it is to notice which claims are structural and which are relative.

    A structural claim tells you what a thing is: this sensor has these emitters, at this wavelength, in this arrangement. You can verify it, or at least someone with a screwdriver can. A relative claim tells you how a thing compares to something else: 60 times more often, 25 percent longer, twice as bright. It is verifiable only if both ends of the comparison are public.

    The Series 12’s sensor is described almost entirely in relative claims. That is what makes it worth a careful read — not because Apple is lying, but because the shape of the disclosure tells you what Apple wanted the conversation to be about.

    One Sentence of Hardware

    Here is essentially everything Apple has said about what physically changed, from the newsroom release:

    > “With larger, power-efficient green LEDs on the optical heart sensor, Apple Watch Series 12 now measures heart rate every five seconds, all day long.”

    The product page adds a second formulation of the same idea — “Redesigned with more capable optical and electrical sensors and larger, more power-efficient green LEDs, the sensing system unlocks deeper health and fitness insights” — and the tech specs reduce it to a list: “second-generation electrical heart sensor, always-on optical heart sensor, blood oxygen sensor, and temperature sensor.”

    That is the complete public description of the hardware. Larger green LEDs, more power-efficient, second-generation electrical sensor. No count of emitters. No count of detectors. No wavelength figures, no sampling architecture, no arrangement.

    Note also how the credit gets distributed. In the sentence that introduces five-second sampling, the cause is optical: with larger, power-efficient green LEDs, the watch now measures every five seconds. A few paragraphs later the silicon gets it instead — “With the new Health Sensing System and S11 chip, Apple Watch Series 12 delivers higher-frequency heart rate and heart rate variability (HRV) measurements.” The same chip is also credited with more accurate step counts and with the new Audio Intelligence features.

    Both components are named, and nowhere does Apple say which one made the frequency increase possible, or in what proportion. That is not a contradiction. It is the absence of a claim — and it happens to be the one claim that would tell you whether this is a sensor upgrade or a compute upgrade, which is the same question as whether any of it could ever reach a Series 11 in software.

    60 Times More Than What

    The number everyone repeated is this one, from the product page:

    > “60x more frequent background heart rate readings than Series 11”

    Two things about it.

    First, it is not in the newsroom release. The press release contains the 24x figure for HRV; the 60x figure for background heart rate lives on the marketing page only. If you see it attributed to Apple’s announcement, that attribution is wrong.

    Second, and more importantly: Apple has never published how often the Series 11 read your heart rate in the background. Not in its specs, not in its support documents, not in the developer documentation. The Series 12 end of the ratio is public — every five seconds. The Series 11 end never was.

    So the claim is unfalsifiable in the plain sense of the word. You cannot check it, and neither can a reviewer, because one of the two numbers required to check it has never existed publicly. You can work backwards and infer that the Series 11 must have sampled around every five minutes, but you are reconstructing Apple’s denominator from Apple’s ratio, which is not verification. It is arithmetic performed on an assumption.

    This is not unique to Apple, and it is not fraud. It is just a category of claim that reads like a specification and functions like an advertisement.

    Apple Stopped Counting

    Here is a small detail I find more telling than the big number.

    On Apple’s own comparison page, the Series 11’s optical sensor is listed as “Third-generation optical heart sensor.” The Series 12’s is listed as “Always-on optical heart sensor.”

    The generation number is gone. Not incremented — removed. Apple counted to three across previous models and then, on the generation it describes as the biggest sensing change in years, stopped counting and substituted a capability word.

    Meanwhile the electrical sensor did get a number: “second-generation.” So this is not a house-style change. Apple is still willing to number a sensor. It stopped numbering this particular one at the exact moment it redesigned it.

    I do not know why, and I am not going to pretend the reason is sinister — “always-on” is genuinely the more meaningful descriptor for a buyer, and generation numbers are a weak signal anyway. But it is a strange place to lose your counter, and it removes the one remaining anchor a reader had for placing this sensor in a lineage.

    Five Seconds on the Screen, Thirty in the Database

    This is the part almost nobody covered, and the source is Apple itself.

    Buried in the methodology of Apple’s heart rate accuracy white paper is an explanation of why Apple needed an internal logging tool to read its own watch during the study:

    > “This data is a single value (not an average) reported every 5 seconds — in watch face complications for background heart rate, and in the Workout app for workout activities. The internal app, which logs the same data surfaced to user, was required because background data is only published to HealthKit every 30 seconds to manage device memory.

    So there are two resolutions, not one. The watch face and the Workout app show you a value every five seconds. HealthKit — the database every third-party app reads, the thing your data exports from, the layer any other developer builds on — receives background heart rate every 30 seconds.

    That is a six-fold gap between what the watch displays and what the watch stores. Both numbers are real, and Apple’s reason is a plain engineering one: writing a sample every five seconds all day would eat memory. Fine. But it means the 60x improvement is 60x at the glass, not throughout.

    If you use a third-party training or sleep app, the resolution it can see is the 30-second one. If you export your health data, that is what you get. The higher-frequency stream feeds Apple’s own features — Readiness, Recovery HRV, the vitals view — and reaches everyone else downsampled.

    I want to be careful here: 30 seconds is still an improvement if the Series 11 wrote less often than that, and Apple hasn’t told us what the Series 11 wrote either. The point is not that Apple degraded anything. It is that “60 times more often” is a statement about a display layer, and the number that governs your actual data is a different, lower one that appears only in a methodology paragraph of a PDF.

    What People With Screwdrivers Think They See

    At the time of writing no retail unit had been torn down — iFixit had not opened one. What exists are hands-on observations from the announcement event, and they are worth reporting precisely because they are the only structural information available — and because they are frequently misquoted.

    Engadget’s hands-on describes “[…] a ring of 12 lights compared to the four LEDs laid out in a diamond shape on last year’s device.” That is a count of emitters, made by eye, at a press event.

    The 5K Runner describes “[…] eight windows with double photodiodes in each, roughly sixteen photodiode elements in total.” That is an estimate of detectors, hedged with “roughly,” and not sourced.

    These two numbers get cited against each other as though the sources disagree. They do not. One is counting lights and the other is counting sensors, and they are different components. What they have in common is more interesting than the discrepancy people imagine: Apple published neither figure. Not the LED count, not the photodiode count, not the arrangement. Every structural number in circulation about this sensor originates outside Apple.

    The same applies to two other claims I would avoid repeating. That the photodiodes are arranged in a ring is a press description, not an Apple statement. That the Series 11 used infrared LEDs for background readings and the Series 12 switched to green appears in one outlet and nowhere in Apple’s materials — Apple says “green LEDs,” and says nothing about what the Series 11 used.

    What This Means If You Wear One

    • The sensor change is real; the size of it is not public. Apple confirmed larger, more efficient green LEDs and a second-generation electrical sensor. Everything more specific than that is somebody’s estimate.
    • Treat 60x as marketing, not as a spec. Not because it is false, but because the number it is measured against has never been published. There is nothing to check it with.
    • If you rely on third-party health apps, your resolution is 30 seconds, not 5. That is from Apple’s own white paper. The five-second figure describes what the watch shows you.
    • Wait for the teardown if the structure matters to you. Every component count you have read so far was made by looking at a demo unit across a table.

    The Takeaway

    Apple made a defensible engineering decision and then described it in the least checkable way available.

    You can see the logic. Component counts invite comparison shopping against Samsung and Garmin on a spec axis Apple does not want to compete on, and they age badly. A ratio — 60 times, 24 times — is vivid, is probably true, and is inert. Nobody can argue with it, because nobody can compute it.

    What you are left with is a sensor that is almost certainly better and is described almost entirely in numbers you cannot verify, sitting under a score that has no published validation, measured by a study that Apple designed and ran itself. Each of those three links in the chain is individually reasonable. It is only when you lay them end to end that you notice there is no point along the way where an outsider gets to check the work.

    Source: Apple Newsroom — Introducing Apple Watch Series 12 · Apple Watch Heart Rate Accuracy Study, September 2026

    Photo: Nik / Unsplash

  • TL;DR — Put the two spec sheets side by side and almost nothing moved. Same display, same resolution, same 2000 nits, same 64GB, same Bluetooth 5.3, same second-generation UWB, same 24-hour battery rating, and — in 2026 — still Wi-Fi 4. What Apple did change is the heart sensing system. That is the entire release. And because a new heart sensor is a new medical device in the eyes of a regulator, the Series 12 shipped unable to do hypertension notifications in 33 countries and regions where the Series 11 can, pending clearances Apple says are in process. Apple changed exactly one subsystem, and it was the one that needs a government’s permission.

    The same sensor redesign landed on the Ultra line, where it came with a quieter subtraction — see the Ultra 4 comparison.

    The Story

    The upgrade question usually sorts into “is the new one enough better.” This year it sorts differently, because the Series 12 is not a broadly improved watch. It is a Series 11 with a new sensor array bolted to the same body, screen, radio stack and battery rating.

    That is not a criticism by itself. The sensor is the interesting part of a watch, and Apple put real work into it — and then published a 1,460-person accuracy study to defend it, which is more than it did for the four new health algorithms that shipped alongside it.

    But it does change how you should read the upgrade. You are not buying a better watch. You are buying a better sensor, and inheriting whatever comes with it.

    The Spec Sheet, Honestly

    Both columns come from Apple. Series 12 from its tech specs page; Series 11 from Apple’s support tech-specs document, because the Series 11 product page now redirects — Apple has discontinued it.

    Series 11Series 12
    Case sizes42mm, 46mm42mm, 46mm aluminum/titanium; 43mm, 47mm ceramic
    Case materialsAluminum, titaniumAluminum, titanium, ceramic
    Weight, 46mm aluminum GPS37.8 g39.5 g
    Weight, 46mm titanium43.1 g44.9 g
    Display resolution416×496 / 374×446416×496 / 374×446
    Max brightness2000 nits2000 nits
    Front glass (aluminum)Ion-X front glass with 2x scratch resistanceCeramic Shield 2
    ChipS10 SiP, dual-core, 4-core Neural EngineS11 SiP, dual-core, 4-core Neural Engine
    Storage64 GB64 GB
    Electrical heart sensorFirst generationSecond generation
    Optical heart sensorThird generationAlways-on (Apple lists no generation number)
    Battery, normal useUp to 24 hoursUp to 24 hours
    Battery, Low Power ModeUp to 38 hoursUp to 38 hours
    Battery, outdoor workoutUp to 8 hoursUp to 10 hours
    15-minute charge yieldsUp to 8 hours normal useUp to 12 hours normal use
    Wi-FiWi-Fi 4 (802.11n)Wi-Fi 4 (802.11n)
    Bluetooth5.35.3
    UWBSecond generationSecond generation
    Water / dust50m, IP6X50m, IP6X
    US starting priceDiscontinuedFrom $399 (smallest aluminum GPS config)

    A few notes on that table, because precision matters more than tidiness.

    The Series 11’s 8-hour outdoor workout figure is not printed on Apple’s support page. It is derived from Apple’s own newsroom claim that the Series 12’s 10 hours is “25 percent more than the previous model.” So treat it as Apple-derived rather than Apple-stated.

    On price: $399 is the only figure Apple publishes, and it is a starting price — it buys the smallest aluminum GPS configuration. The 46mm size that every weight in that table refers to costs more, as do titanium, ceramic and cellular. Apple does not print a public price list for those; the numbers circulating come from retailer listings. So read $399 as the floor, not as the price of the watch described in the rest of this table, and check the configurator for the one you actually want.

    And yes — Wi-Fi 4. The 802.11n standard was ratified in 2009. It is a defensible choice on a device this small, since the watch mostly talks to a phone over Bluetooth and the radio budget matters more than throughput. But it has now survived long enough to be older than some of the people buying the watch.

    What Series 11 Owners Get for Free

    This is the part that quietly deflates most of the upgrade case.

    watchOS 27 supports Series 9, 10, 11 and 12, plus SE 3 and Ultra 2 through 4. Your Series 11 is not being cut off. It gets the dynamic app grid with Siri-suggested apps, the single-handed Smart Stack gesture, the Siri Modular watch face, perimenopause and menopause support in Cycle Tracking, Workout Buddy without a nearby iPhone, and the consolidated Find My.

    The redesigned Health app, the Longevity tab and Health Age all live on the iPhone side. Apple’s footnote for Health Age says it “requires Apple Watch” — no generation named.

    Sleep apnea notifications, which people often assume are a new-watch feature, run on Series 9 or later, Ultra 2, and SE 3. Your Series 11 already qualifies.

    So a large share of what looks like “the new Apple Watch experience” arrives on the old one as a free download.

    What Only New Silicon Can Do

    The genuinely exclusive list is short, and it is all downstream of the sensor:

    • The Health Sensing System itself — second-generation electrical heart sensor, always-on optical sensor, background heart rate every five seconds, HRV as often as every five minutes.
    • Readiness, the 0–10 daily score, which is computed from that higher-frequency data.
    • Recovery HRV and the new daytime vitals view in the Heart Rate app.
    • Audio Intelligence — Sound Recognition, Live Rewind, Siri Recap, Shazam — which runs on the S11.
    • Ceramic Shield 2, the ceramic case option, and the faster charging curve.

    One honesty note on Readiness. Apple has not published a compatibility document saying it requires a Series 12. It names Readiness on the Series 12 product pages and leaves it off the watchOS 27 feature list, which is strong circumstantial positioning but not an explicit statement. If Readiness is your reason to upgrade, that is worth knowing: you are trusting marketing placement, not a support page.

    The Upgrade That Downgrades You

    Here is the part almost no one leads with, and it is the most important thing in this comparison.

    From Apple’s own watchOS feature availability page, verbatim:

    > “Hypertension Notifications may not be available on Apple Watch Series 12 or Ultra 4 in your region, as additional regulatory clearances for these devices are in process.”

    Read that again with the model numbers in mind. The restriction names the new watches. In 33 countries and regions — a count compiled by MacRumors rather than published by Apple, covering Korea, Japan, Canada, Australia, India, Singapore, Taiwan, the UAE and Brazil — a Series 11 owner who upgrades to a Series 12 loses hypertension notifications until Apple clears the new hardware locally. The US, UK and EU are unaffected.

    And this is exactly where the “Apple only changed one subsystem” observation stops being trivia and starts being the whole story.

    A hypertension notification is a regulated medical device function. Clearance is granted for a specific device with a specific sensor, not for a brand. Apple changed the sensor. That reset the clock in every jurisdiction that has to look at it again. The display did not need re-clearance. The Wi-Fi radio did not. The only component Apple touched is the only component a health regulator cares about.

    So the single-subsystem upgrade and the 33-country regression are not two separate facts. They are the same fact told from two ends.

    It should resolve. Apple says clearances are in process, and it has been through this cycle before with ECG and with sleep apnea notifications. But “it should resolve” is not a date, and if you live in one of those 33 regions and you use hypertension alerts, upgrading on launch day means giving up a working health feature for an unspecified number of months in exchange for a score Apple hasn’t published a paper on.

    The Battery Question Nobody Can Answer Yet

    Apple kept the rating at 24 hours while claiming a 60-fold increase in background heart rate sampling. That is the most interesting engineering claim of the release, and at the time of writing no independent battery measurement of a retail unit had been published.

    I want to be blunt about this because the internet will not be: any “real-world battery life” number for the Series 12 published before units are in reviewers’ hands is not a measurement. Apple’s 24-hour figure comes from a defined protocol — 300 time checks, 90 notifications, 15 minutes of app use, a 60-minute workout with Bluetooth audio, and 6 hours of sleep tracking. Useful as a constant across generations, not as a prediction of your Tuesday.

    If battery is your deciding factor, the honest advice is to wait for tests rather than to buy or skip on a number that does not exist yet.

    Who Should Actually Upgrade

    • Coming from a Series 11, in one of the 33 affected regions, and you use hypertension alerts — wait. You would be trading a cleared medical feature for an uncleared one.
    • Coming from a Series 11, everywhere else, and you want Readiness — this is the real case, and it is a sensor-and-score purchase. Just go in knowing the score has no published validation behind it.
    • Coming from a Series 11 for anything other than health sensing — skip. The screen, radios, storage and battery rating are the same, and watchOS 27 brings the interface changes to the watch you already own.
    • Coming from a Series 9, or an older SE, or nothing — the comparison that matters to you is not 11 versus 12. The accumulated gap since Series 9 is real, and at $399 the Series 12 is the one Apple will keep supporting longest.
    • Shopping clearance Series 11 stock — plausible, and in the affected regions it does more health-wise than the new one until those clearances land. Verify the discount is real at the moment you buy; clearance pricing moves daily.

    The Takeaway

    The interesting thing about the Series 12 is not that Apple changed little. It is what the little was.

    Every generation has a couple of components that are hard to change, and for a watch that has quietly become a medical device, the sensor is now the hardest of all — not because the engineering is difficult, but because touching it means asking dozens of governments for permission again. Apple touched it anyway. The 33-country hypertension gap is the receipt.

    That tells you something about where this product is heading. The screen, the chip, the radios are approaching the boring plateau of a mature device. The sensor is the frontier, and the frontier now runs through a regulatory office. Every future Apple Watch that meaningfully improves health sensing will pay this same tax, and buyers in the slower-clearing markets will keep paying it in the form of a new watch that briefly does less than the old one.

    Source: Apple Watch Series 12 Tech Specs · Apple watchOS Feature Availability

    Photo: Simon Daoudi / Unsplash

  • TL;DR — Apple’s nine-page heart rate study is better than it had to be. It enrolled 1,460 people at five sites, used a chest-strap ECG as reference, and reported its own losses. It also discloses three things worth sitting with: Apple read its own watch through an internal logging app no competitor was given, Apple chose how Samsung’s untimestamped background data would be scored, and the Daily Living protocol — the part underpinning the “60 times more often” marketing claim — had no pre-specified sample size at all. None of that makes the result wrong. It does change what the result means.

    For the hardware behind those numbers, see what Apple did and did not disclose about the sensor.

    The Story

    Apple published one new health white paper for the Series 12 launch, and it is not the one the four new algorithms needed. It is a heart rate accuracy study, nine pages, dated September 2026.

    The headline is easy to repeat: across 159 activity-by-measure comparisons against six competitors, Apple Watch was more accurate in 139.

    The paper is worth reading properly, though, because it is a genuinely unusual document. Most vendor benchmarks are written to be un-checkable. This one hands you the knife. Apple wrote down its exclusion criteria, its statistical model, its data quality thresholds, the exact reason it needed a special tool to read its own watch, and the two comparisons it lost. A company writing pure marketing does not include the sentence “the comparison device’s accuracy was higher in 2.”

    So the interesting question is not whether Apple cheated. It is narrower and more useful: given exactly what Apple describes, what does the 139 actually measure?

    The Setup

    Participants wore an Apple Watch on one wrist and a competitor on the other, with a Polar H10 chest strap ECG as the reference standard. Which wrist got the Apple Watch was randomized, so the dominant hand carried it about half the time — a small detail that matters more than it sounds, because wrist motion is one of the main things that breaks optical heart rate.

    The comparison set: Garmin Forerunner 970, Google Pixel Watch 4, Huawei Watch 5, Samsung Galaxy Watch 8, WHOOP 5, and Oura Ring 5 on an index finger. Those brands account for roughly 90% of the worldwide smartwatch category, according to Omdia shipment data cited in the paper — not Apple’s own estimate.

    Activities were chosen to break things on purpose. Outdoor runs, treadmill intervals, outdoor cycling, HIIT, strength training, 15 to 20 minutes each. Apple names the failure modes it was hunting: cadence lock in running, loaded forearm grip in strength work and cycling, rapid wrist motion and fast heart rate transitions in HIIT. There was also a separate Daily Living protocol — rest, desk work, meal preparation, walking, about 20 minutes each — to test background heart rate.

    Five sites: two near Cupertino, plus San Diego, Austin, and Selangor, Malaysia. Apple ran Cupertino and Austin itself; a third-party contract research organization ran San Diego and Selangor, and 59% of the analyzed participants came from those CRO sites.

    Of 1,460 enrolled, 1,254 contributed at least one paired measurement. Twenty percent had Fitzpatrick skin tones V–VI, which is the demographic optical sensors historically fail and wearable studies historically underweight. Mean age 39, mean BMI 26. Thirty-five percent female — described as “representation across sex,” which it is, though it is a roughly two-to-one male skew.

    One exclusion criterion is worth quoting because nobody writes this by accident. Alongside beta blockers, equipment allergies, unstable cardiovascular conditions, and tattoos or large moles at the sensor site, Apple excluded people working in tech media.

    Everyone Got a Different Pipe

    Here is the first thing that changes how you read the result.

    Apple could not just ask each device for its data, because each manufacturer exposes data differently. So Apple picked a method per device. For most competitors it used a Bluetooth Low Energy stream from the manufacturer’s own heart rate app into a logging app on a separate iPhone — because, as the paper puts it, “not all manufacturers write high-fidelity data to HealthKit, or permit such data to be exported from their companion apps.”

    For the Apple Watch, Apple used an internal data logging application.

    The stated reason is specific and checkable: background heart rate is only written to HealthKit every 30 seconds, to conserve device memory. But the value the watch actually shows you — in a watch face complication, in the Workout app — is a single reading every 5 seconds. To evaluate the number the user sees, Apple needed a tool that captures the number the user sees. HealthKit would have understated its own watch.

    That logic holds. It is also true that only one company in this study had the ability to build that tool.

    And this is where the framing of the whole paper matters. Apple states its selection rules up front, and there are three: the data stream had to be visible to the user in the product experience, it had to carry timestamps, and where a device offered more than one qualifying stream, Apple took the highest frequency available. The second and third rules are the ones that quietly do the work — a device that surfaces a number but will not tell you when it was taken cannot be scored the normal way, and a device that exports a slower stream than it displays gets scored on the slower one. The first rule is a deliberate and defensible choice — arguably the more honest consumer question is “how accurate is the number on my wrist,” not “how good is the raw diode.” But it has a consequence Apple does not spell out:

    Part of what this study measures is how much data each company lets out of its own ecosystem.

    If a competitor’s sensor is excellent but its app only exports a coarse stream, this methodology scores the coarse stream. That is a real thing that affects real users. It is not the same claim as “Apple’s sensor hardware is more accurate,” and the paper’s conclusion — “Apple Watch offers users the most accurate heart rate sensing among the wearable devices evaluated” — sits right on the seam between the two.

    The Samsung Problem

    The Galaxy Watch 8 case is the sharpest illustration, and to Apple’s credit it is disclosed in full.

    For daily-living background readings, Samsung’s export gave no per-sample timestamps. It gave a start time, an end time, a minimum, a maximum, and one additional heart rate value whose derivation Samsung does not document.

    So Apple had to decide what that mystery value should be compared against. It tested three options: the reference at the interval’s start, the reference at the interval’s end, and the mean of the reference across the interval. The interval mean fit best, so that is what Apple used.

    Read that again. Apple selected the scoring rule for a competitor’s data by testing which of three candidate alignments matched the reference most closely — the best-fit option rather than an arbitrary one — and then still reported that Samsung lost.

    That is the fair-minded reading, and I think it is the right one. But the underlying situation stands regardless of intent: in one head-to-head comparison, one competitor’s grading method was chosen by the other competitor, because the first one shipped an undocumented number.

    The Actual Scoreboard

    The 139 is real. It is also not the whole line.

    Across 159 activity-by-measure comparisons: Apple better at 99% confidence in 139, indeterminate in 18, competitor better in 2. Ignore statistical significance and just count which way the point estimate leaned, and Apple is ahead in 155 of 159. On the overall workout and overall daily living protocols, Apple beat every device on both RMSE and MAE.

    That is a strong result by any standard. Two details are worth pulling out of the figures, though.

    The Pixel Watch 4 beat the Apple Watch. Twice — RMSE and MAE, in the rest condition of the Daily Living protocol. Apple reports it and adds that the margin was under 0.5 bpm. Pixel Watch 4 also produced an indeterminate result against Apple in cycling, the only workout activity where Apple did not win outright against everyone.

    Now notice the asymmetry in how those are written. Apple’s wins are reported as statistically significant, direction only. Apple’s single loss is reported with a magnitude attached — “by less than a margin of 0.5 bpm.” Both statements are accurate. But the effect sizes for the 139 wins live only inside Figures 1 through 3; there is no table of numbers in the text. You are told Apple won and by how little it lost.

    For calibration: the study was sized to detect differences in mean absolute error of roughly 2.5 bpm in running and cycling, up to 3.7 bpm in HIIT. So “statistically significant” here is not synonymous with “large enough to feel.” Whether a 2 bpm edge changes anything about your training is a separate question from whether it is real.

    The Least-Designed Part Is the Part Marketing Needs Most

    This is the finding I did not expect.

    Apple pre-specified its sample size: a minimum of 50 participants per workout activity per comparator device, with 80% power, α = 0.05, and a Bonferroni adjustment for the six device comparisons. That is a properly powered design.

    And then one clause: sizing “pertains only to workouts and not Daily Living.”

    The Daily Living protocol had no pre-specified sample size. It is still analyzed with the same linear mixed model and the same 99% bootstrap confidence intervals, so the reported results are not casual. But it was not designed to a target in advance.

    Daily Living is background heart rate. Background heart rate is precisely what the Series 12 marketing claim is about — readings 60 times more often than Series 11, the finer-grained passive data that Readiness and Health Age are built from. The single most-promoted capability of the new sensor is validated in the half of the study that got no power analysis.

    What Apple Chose to Tell You

    I want to be even-handed, because the easy version of this article is cynical and the easy version is wrong.

    Apple disclosed, without being obliged to:

    • that it used a tool on its own device that no competitor had, and exactly why
    • that it picked the alignment method for a competitor’s undocumented data
    • that it lost two comparisons, and to whom
    • that 18 more were indeterminate
    • that the Daily Living arm was not power-sized
    • that its participant pool was 65% male

    Each of those is a stick handed to a critic. Most vendor white papers contain none of them.

    The gaps that remain are structural rather than sneaky. The protocol was approved by review boards within Apple — the paper names groups covering research ethics, safety, privacy, and data governance, but no external institutional review board. There is no trial registration. The paper is not peer-reviewed, and the underlying data is not available, so nobody outside Apple can re-run the analysis with different reasonable choices and see whether 139 becomes 120. Those are the normal conditions of vendor research, not misconduct — but they are the reason vendor research and independent research are not interchangeable, no matter how carefully the vendor writes.

    What This Means If You Wear One

    • Heart rate is the one Series 12 capability with fresh evidence behind it. Everything else in the new health stack shipped without a paper. If you are choosing based on documentation, this is the documented part.
    • Read the claim as “the number on the screen,” not “the sensor.” That is Apple’s own stated framing, and it is the useful one for a buyer — but it means a competitor’s poor export policy counts against its score.
    • Do not translate “more accurate” into “meaningfully different for you.” The design could detect gaps around 2.5 to 3.7 bpm. If you are training by heart rate zones, ask whether a couple of bpm crosses a boundary you care about. Usually it does not.
    • The background-data claim is the softest part. Workout accuracy was designed, powered, and won convincingly. Passive all-day accuracy was measured with less advance design — and it is the input to the scores Apple is promoting hardest.

    The Takeaway

    Grading your own homework is not automatically dishonest. It becomes dishonest when you hide the rubric. Apple published the rubric: which tool read which device, how a competitor’s ambiguous data was handled, where the power analysis applied and where it did not, and which comparisons it lost.

    What you are left with is a well-run study that answers a narrower question than its conclusion sentence implies. Apple Watch reports heart rate to its user more accurately than six competitors report heart rate to theirs, under conditions Apple designed, using an extraction method only Apple could build for itself, and with the strongest evidence in workouts rather than in the passive background data its new features actually run on.

    That is still a good result. It is just a different sentence than “Apple has the most accurate heart rate sensor,” and the gap between those two sentences is where all the interesting reading is.

    Source: Apple Watch Heart Rate Accuracy Study, September 2026

    Photo: Nik / Unsplash

  • TL;DR — Apple Watch Series 12 launched with a redesigned sensor array and four new health algorithms: Readiness, Health Age, Movement Evaluations, and a Longevity tab. Apple published exactly one new validation paper alongside them — a heart rate accuracy study — and that paper does not test heart rate variability or nighttime data. Readiness runs on heart rate variability and sleep. Apple has published validation papers for sleep apnea, hypertension, sleep stages, blood oxygen, and arrhythmia. Four of those five needed a regulator’s signature. Blood Oxygen never did — Apple shipped it as an unregulated wellness feature and published a paper on it anyway. That is the standard Apple set for itself, and the four new features do not meet it. The same launch also demonstrates the cost of the alternative: hypertension notifications, the regulated feature, is unavailable on the new hardware in 33 countries because clearances do not transfer.

    If you are weighing an upgrade rather than the algorithms themselves, the full Series 12 vs Series 11 comparison goes through the spec sheet line by line.

    The Story

    Apple published one new health white paper for the Series 12 launch. It is nine pages, dated September 2026, and it is titled Heart Rate Accuracy Study.

    It is a serious document. Apple enrolled 1,460 people across five sites in Cupertino, San Diego, Austin, and Selangor, Malaysia, and analyzed data from 1,254 of them. Twenty percent had Fitzpatrick skin tones V–VI, which matters enormously for optical heart sensors and is a demographic that wearable validation studies have historically underweighted. The reference standard was a Polar H10 chest strap ECG. Apple ran the watch against six competitors — Garmin Forerunner 970, Pixel Watch 4, Huawei Watch 5, Galaxy Watch 8, WHOOP 5, and Oura Ring 5 — across outdoor running, treadmill intervals, cycling, HIIT, strength training, and ordinary daily living. Across 159 activity-by-metric comparisons, Apple Watch came out ahead in 139.

    That is a real study, and the result is probably real too. Hold that thought, because the interesting part is not what the paper contains.

    The paper does not evaluate heart rate variability. It does not evaluate nighttime or sleep data.

    Now look at what Apple built on top of the same sensor. Readiness is the headline software feature of Series 12 — a 0-to-10 score with four verdicts (Recover, Pace Yourself, Ready, Go For It) that updates through the day. Apple’s description of the inputs is one sentence long: recent activity, training load, vitals, and your sleep score. It does not say what sits inside “vitals,” and it does not say how the four are weighted. Separately, Apple says overnight vitals now includes recovery HRV measurements, and that this generation splits heart rate variability into two readings — Recovery HRV and Overall HRV — sampled up to 24 times more often than on Series 11. Put those two statements side by side, which Apple does not do for you, and the score looks like it leans substantially on HRV and sleep.

    So the flagship feature runs on HRV and sleep, and the only validation paper Apple published measures neither.

    Four Algorithms, No Documentation

    Readiness is not alone. Series 12 shipped with three more inference-heavy health features, and none of them has a technical paper:

    FeatureWhat it claimsValidation paper
    Readiness0–10 recovery score from activity, training load, vitals, sleepNone published
    Health AgeA biological age from VO2 max, resting heart rate, sleep, HRV, plus optional lab valuesNone published
    Movement EvaluationsFlexibility, strength, balance and mechanics scored by iPhone camera vision modelsNone published
    Longevity tabSeven-domain longevity viewNone published

    Apple says Readiness was developed using data from the Apple Heart and Movement Study, in collaboration with Apple’s own exercise physiologists and physicians. That is the entire disclosure. No subject count, no study duration, no agreement statistic against any reference, no error bars.

    And there is a smaller tell in the same launch. Movement Evaluations includes a new VO2 max test run through the iPhone camera — it requires an iPhone, an Apple Watch, and either AirPods Pro 3 or a third-party heart rate monitor working together, not the watch on its own. The document Apple still points to for VO2 max estimation, Using Apple Watch to Estimate Cardio Fitness with VO2 max, carries a creation date of May 2021. A new measurement method arrived; the supporting paper did not move.

    Apple Used to Do This Differently

    This is not a company that dislikes publishing. Apple has a standing library of health validation papers, and it is reasonably thorough:

    Line those up against the four new features and a pattern falls out immediately.

    Four of those five are features a regulator had to approve. Atrial fibrillation detection, sleep apnea notifications and hypertension notifications are cleared medical device functions. Hypertension notifications went through the FDA as a 510(k), cleared under product code SFR as a Class II device. Sleep apnea went the same route. When you file a 510(k) you are assembling an evidence package regardless; publishing a consumer-facing summary of it costs you almost nothing.

    The fifth one is the interesting one. Blood Oxygen was never FDA-cleared. Apple deliberately shipped it as a general wellness feature in 2020, outside the regulatory perimeter, explicitly not for medical use. That was a legitimate path, not a trick — classify a measurement as general wellness and no clearance is required. Then Apple published a paper on it anyway, two years later.

    That matters, because it means the standard the four new features are failing is not a regulator’s. It is Apple’s own. Nobody obliged Apple to document Blood Oxygen. Apple documented it because that is what the company did when it shipped a new physiological measurement.

    None of the four new features is a regulated claim either. A recovery score is not a diagnosis. A “health age” is not a medical finding. A flexibility rating from a phone camera is not a clinical assessment. In the United States these sit inside the FDA’s general wellness policy, which carves out low-risk products that promote a healthy lifestyle without claiming to diagnose or treat. No submission, no evidence package, no obligation to publish anything.

    Blood Oxygen sat in that same carve-out and got a paper. Readiness, Health Age, Movement Evaluations and the Longevity tab did not.

    So the regulatory line is not a loophole Apple invented — it is the real boundary, and every wearable company lives on both sides of it. But the boundary does not explain the change. Apple used to publish on both sides of it. What changed is that the papers stopped at precisely the moment the new health features stopped needing anyone’s signature.

    The Same Launch Shows Why

    Here is the part that turns a pattern into a strategy.

    Hypertension notifications — Apple’s most heavily validated recent health feature, with a published paper and an FDA clearance — does not work on Series 12 or Ultra 4 in 33 countries. Apple’s own watchOS feature availability page states it plainly: the feature “may not be available on Apple Watch Series 12 or Ultra 4 in your region, as additional regulatory clearances for these devices are in process.”

    The reason is the sensor. Apple redesigned the optical and electrical heart hardware, and a clearance granted for one sensor configuration does not automatically carry to another. So the regulated feature has to queue up again, country by country, on the new hardware. Buyers in Canada, Australia, Japan, Singapore, India, Brazil, Korea and two dozen other markets upgrade to the newer watch and lose a health feature the older watch had.

    Readiness, meanwhile, shipped everywhere on day one.

    There is also a business reason the unregulated side is worth building out. At the same event Apple announced a Quest Diagnostics blood panel — more than 50 biomarkers for $119, ordered through the Health app, US only, described as arriving later in 2026 and not live yet. Health Age already takes optional lab values as an input. Neither company has disclosed how the money moves, so this is a shape rather than an accusation — but the shape is a wellness score that asks you for lab work, and a paid lab order sitting one tap away inside the Health app. That is a loop Apple can close on its own surface. A regulated diagnostic sends you to a doctor instead.

    That contrast is the whole argument in miniature. The regulated path buys you credibility and costs you years and geography. The wellness path costs you nothing and ships instantly. Apple built a sensor overhaul whose entire justification is finer-grained physiological data, and then routed that data almost entirely into features that no regulator will ever examine.

    Frequency Is Not Accuracy

    There is one more thing worth separating out, because Apple’s own language is careful about it in a way the coverage has not been.

    Apple’s claim for the new sensor is that Series 12 takes background heart rate readings 60 times more often than Series 11, and HRV readings up to 24 times more often. Read those again. They are claims about frequency. Apple did not claim the readings are more accurate, and the accuracy paper it did publish is about heart rate, not HRV.

    The two multipliers also describe two different signals, and they are worth keeping apart. Heart rate moves from an interval Apple’s general support documentation puts at roughly five minutes, varying with activity, down to about five seconds — that is the 60×. HRV moves from spacing on the order of a couple of hours down to roughly five minutes — that is the 24×.

    Neither of those baselines appears on a Series 11 product page. Apple has never published a per-model background sampling interval; the numbers above come from general support documentation and from working backwards through Apple’s own multiplier. That is a reconstruction, not a disclosure, and it means “60 times more often” cannot be checked against anything Apple has actually stated.

    The distinction matters because of how a recovery score works. Sampling an unvalidated signal more often does not make the signal more true. It makes the derived number more stable and more confident-looking. A Readiness score that updates all day off 24× the HRV samples will feel far more authoritative than one computed from a single morning reading — and the underlying question of whether wrist HRV tracks recovery well enough to justify a 0-to-10 verdict is exactly as open as it was before.

    Nobody Documents the Score. Apple Documents Nothing.

    It is worth being precise about how Apple compares here, because the honest version of the comparison is narrower than the easy version.

    Nobody documents their composite recovery score. WHOOP does not publish how Recovery is calculated. Oura does not publish how Readiness is calculated — and yes, Oura’s flagship metric has carried that exact name for years, which makes Apple’s choice of word a small statement in itself. Garmin does not publish how Body Battery is calculated. All three are black boxes at the top, and so is Apple.

    The difference sits one level down, at the components. WHOOP’s menstrual-cycle work went through peer review in npj Digital Medicine, and its Healthspan feature has a published white paper behind it — not a journal article, but a document you can read. Oura’s sleep staging algorithm, OSSA 2.0, was validated against polysomnography in Sleep Medicine, 96 subjects and over 400,000 epochs. Garmin’s physiological metrics rest on the publicly documented Firstbeat research base. Each of them has published something about the parts, even where the whole stays closed.

    Apple has both layers for its regulated features and neither for its new ones. For Readiness, Health Age, Movement Evaluations and the Longevity tab there is no composite documentation and no component documentation. That is the actual gap — not that Apple is uniquely opaque about the score, but that it is the only one of the four with nothing underneath it either.

    What This Means If You Wear One

    None of this makes Readiness useless. A directional signal built from real physiological data, tracked against your own baseline, can be genuinely helpful — the same way a bathroom scale is useful even though it cannot tell you your body composition.

    The practical guidance is narrower than that:

    • Treat Readiness as a trend, not a measurement. Its value is in how it moves relative to your own history. The absolute number has no external reference to be right or wrong against.
    • Heart rate is the part that was actually tested. If you are buying Series 12 for a specific capability, continuous heart rate is the one with a fresh nine-page study behind it and a documented win over six competitors.
    • Check the regulated features against your country before upgrading. Hypertension notifications is the current example, and it is a real downgrade in 33 markets. Apple lists availability by region, and that page is worth reading before an upgrade rather than after.
    • Health Age deserves the most skepticism, and for a practical reason. Compressing VO2 max, resting heart rate, sleep and HRV into one “biological age” is the largest inferential leap of the four. The concrete problem is not that the number might be wrong — it is that you cannot see which input moved it. If your Health Age climbs two years, nothing tells you whether that was sleep, cardio fitness, or a bad week of resting heart rate, so there is nothing in it you can act on.
    • Movement Evaluations is a camera measurement, so read it like one. Flexibility, strength and balance scored from phone video will shift with lighting, camera angle, distance and what you are wearing. Compare sessions you recorded the same way in the same place, or do not compare them at all.

    The Takeaway

    The story of this launch is not that Apple built a worse sensor. By its own testing, it built a better one, and the testing looks credible.

    The story is where the output goes. Apple spent a hardware generation on physiological measurement and then pointed that measurement at four features sitting deliberately outside the regulatory perimeter — while its most rigorously validated feature was stuck at the border in 33 countries. The papers stopped appearing at exactly the point where nobody was requiring them. Apple’s own Blood Oxygen paper is the proof that it did not always need to be required.

    If you want to know how much a health feature has been checked, the most reliable signal is not the marketing. It is whether Apple published something you can read.

    Source: Apple Heart Rate Accuracy Study, September 2026

    Photo: Simon Daoudi / Unsplash

  • 휴머노이드 회사들은 결국 손을 보여 줍니다.

    시연 영상에서 가장 공들이는 장면이 손입니다. 손가락이 머그컵을 감싸고, 케이블을 끼우고, 기계를 보고 있다는 사실을 잠시 잊게 할 만큼 섬세한 동작을 합니다. 그리고 화면에 띄우는 숫자는 거의 항상 같은 종류입니다. 자유도. 22축, 25축. 관절이 이만큼 많다는 이야기입니다.

    그 숫자의 문제는 분명합니다. 손이 무엇을 움직일 수 있는지는 알려 주지만, 손이 무엇을 느낄 수 있는지는 전혀 알려 주지 않습니다.

    1X가 휴머노이드 NEO에 장착할 새 손을 공개했습니다. 이 손을 볼 가치가 있는 이유는 회사가 앞세운 숫자가 자유도가 아니기 때문입니다.

    자유도가 아니라 기어비입니다

    1X 공식 기술 소개에 따르면 이 손은 25 자유도입니다. 손가락과 손바닥에 완전 구동 22축, 손목에 3축이 더 붙습니다. 텐던(힘줄) 구동이고, 모터는 손 안이 아니라 전완에 들어갑니다. 여기까지는 인상적이지만 특이하지는 않습니다.

    특이한 것은 동력 전달부입니다. 감속비가 약 5:1에서 15:1입니다.

    이 숫자가 와닿지 않아도 괜찮습니다. 이 글 전체가 이 숫자에 관한 이야기이기 때문입니다.

    감속기는 속도를 힘으로 바꿉니다. 작은 모터에 큰 감속비를 걸면 큰 힘이 나옵니다. 산업용 로봇 관절이 고감속 하모닉 드라이브를 쓰는 이유이고, 흔히 쓰이는 감속비는 100:1에서 200:1 사이입니다. 한 방향으로는 훌륭하게 작동합니다. 문제는 반대 방향입니다.

    밖에서 그런 관절을 밀면, 그 힘은 기어열 전체를 거꾸로 밀고 나와야 합니다. 마찰은 방향을 가리지 않고 저항합니다. 고감속비 하모닉 드라이브의 역방향 효율은 대체로 30~50퍼센트 구간에 놓이고, 감속비와 윤활 조건에 따라 40퍼센트 아래로 떨어지기도 합니다. 정방향으로 돌릴 때보다 서너 배 큰 토크를 넣어야 겨우 역구동되는 경우도 있습니다.

    여기서 나오는 결론은 잔인한데 잘 언급되지 않습니다. 감속기가 자기 내부 마찰을 이기는 데만 5 Nm를 쓴다면, 그 관절은 5 Nm 이하의 모든 접촉력에 대해 장님입니다. 부정확한 것이 아닙니다. 아예 못 느낍니다. 신호가 약해지는 것이 아니라 도달하지 않습니다.

    두 번째 손실은 어쩌면 더 심각합니다. 반사 관성 — 바깥세상 입장에서 모터가 얼마나 무겁게 느껴지는가 — 은 기어비의 제곱에 비례합니다. 10:1에서 150:1로 가면 관절에서 느껴지는 관성이 15배가 아니라 225배가 됩니다. 그렇게 감속된 손가락은 더 이상 손가락이 아닙니다. 손가락 모양을 한 작은 프레스입니다.

    이 두 가지가 합쳐지면 결과가 분명해집니다. 고감속 관절은 접촉을 감지하는 데 두 번 실패합니다. 작은 힘은 마찰에 묻혀서 못 느끼고, 느낀다 해도 관성 때문에 제때 물러나지 못합니다. 감지와 반응이 동시에 막히는 구조입니다.

    여기서 한 가지 더 짚어 둘 것은 마찰 손실과 반사 관성이 감속비에 대해 같은 속도로 커지지 않는다는 점입니다. 마찰 손실은 대체로 완만하게 늘어나지만, 출력측에서 본 반사 관성은 감속비의 제곱을 타고 올라갑니다. 그래서 감속비를 두 배 올리면 손실이 두 배 늘어나는 것이 아니라, 어느 지점을 넘어서면 역구동 자체가 사실상 성립하지 않는 구간으로 넘어갑니다. 셀프 로킹이라고 부르는 상태입니다. 웜기어가 그 극단적인 예로, 정방향으로는 잘 돌지만 역방향으로는 아예 돌지 않습니다. 고감속 하모닉 드라이브는 그 정도까지 가지는 않지만 같은 방향의 성질을 갖습니다. 감속비를 높이는 선택은 힘과 감각 사이의 완만한 트레이드오프가 아니라, 어느 순간부터 한쪽을 통째로 버리는 선택에 가깝습니다.

    그래서 업계는 오랫동안 우회로를 썼습니다. 손목이나 손가락에 별도의 힘·토크 센서를 붙이는 방식입니다. 다만 이 방식에는 구조적 한계가 있습니다. 센서는 자기가 붙어 있는 지점의 힘만 읽습니다. 센서 바깥쪽 링크의 질량과 마찰은 여전히 측정되지 않은 채로 남습니다. 그리고 센서가 힘을 읽어도 관절이 그 정보에 맞춰 부드럽게 물러나지 못하면, 아는 것과 하는 것 사이의 간극은 그대로입니다.

    “센서를 달면 되지 않느냐”가 왜 해답이 아닌지는 여기서 분명해집니다. 문제는 감지가 아니라 응답입니다. 손끝에 아무리 좋은 센서를 붙여도, 그 신호를 받아 힘을 줄이라고 명령했을 때 관절이 그만큼 물러나 줘야 의미가 생깁니다. 반사 관성이 225배로 부풀어 있는 관절은 명령을 받고도 즉시 방향을 바꾸지 못합니다. 모터가 감속을 시작해도 기어열 전체를 세워야 하고, 그 사이에 접촉면에서는 이미 힘이 올라가 버립니다. 제어 주기를 아무리 빠르게 돌려도 기계적으로 못 따라가는 구간이 남습니다. 감각은 소프트웨어로 덧붙일 수 있지만, 응답 속도는 감속기의 물리가 정합니다.

    유리컵이 깨진 뒤에야 아는 손

    그런 손에 와인잔을 집으라고 하면 무슨 일이 벌어지는지 생각해 볼 만합니다.

    손은 위치 명령을 실행합니다. “이 각도까지 닫아라.” 자기가 얼마나 세게 쥐고 있는지는 의미 있는 수준으로 알지 못합니다. 그 정보가 지나갈 유일한 통로가 마찰과 관성으로 막혀 있기 때문입니다. 잔이 미끄러졌다는 사실은 사람과 똑같은 방식으로, 소리가 나고 나서야 알게 됩니다.

    1X가 선택한 교환이 바로 이 지점입니다. 그리고 이것이 교환이라는 점을 솔직하게 짚어야 합니다.

    감속비를 한 자릿수와 낮은 두 자릿수까지 떨어뜨리면, 25개 관절 전부가 회사 표현대로 네이티브 힘 제어가 되고 완전히 백드라이버블해집니다. 밖에서 손가락을 밀면 그대로 밀려 준다는 뜻입니다. 힘은 물체 쪽으로 흘러 나가고, 접촉 정보는 정확히 같은 기계적 경로를 타고 되돌아옵니다. 나중에 덧붙인 로드셀도 필요 없고, 기어열과 싸우는 모터 전류에서 힘을 추정할 필요도 없습니다. 동력 전달부 자체가 센서입니다. 이것을 ‘힘의 투명성’이라고 부를 만합니다. 손과 세계가 서로를 느끼는 상태입니다.

    ‘네이티브’라는 말에 방점이 있습니다. 힘 제어를 소프트웨어로 흉내 내는 방법은 예전부터 있었습니다. 모터 전류를 읽어 토크를 추정하고 임피던스 제어(힘과 위치를 함께 다루는 제어)를 거는 방식입니다. 문제는 그 추정이 마찰 모델의 정확도에 통째로 의존한다는 점입니다. 마찰은 온도에 따라, 마모에 따라, 심지어 그날 관절이 얼마나 움직였는지에 따라 변합니다. 모델은 계속 어긋나고, 어긋난 만큼이 그대로 오차가 됩니다.

    감속비가 낮으면 추정할 것이 별로 없습니다. 마찰 항이 작아서 모터에서 읽은 값이 손끝에서 벌어지는 일과 거의 그대로 대응합니다. 소프트웨어가 물리를 보정하는 것이 아니라, 물리가 이미 정직한 상태입니다.

    공개된 수치가 오히려 정직합니다

    공개된 수치들은 과장과 거리가 멉니다. 그 점이 오히려 이것이 공짜 점심이 아니라 진짜 선택이었음을 보여 줍니다.

    항목1X NEO 손
    총 자유도25축 (손가락·손바닥 22 + 손목 3)
    구동 방식텐던 구동, 모터는 전완 배치
    감속비약 5:1 ~ 15:1
    엄지 CMC 최대 토크3.5 Nm
    손가락 MCP 최대 토크2.6 Nm
    말단 굴곡력최대 45 N
    손목 토크17.75 Nm
    위치 정밀도±0.2 mm
    방수·방진IP68, 식품 접촉 안전 소재
    내구 시험손가락 수백만 사이클, 손목 200만 사이클 초과
    생산 능력연 최대 1만 개 (회사 주장)
    무게미공개
    소비 전력미공개

    IP68에 식품 접촉 안전 소재라는 조합은 1X가 이 로봇을 어디에 두려는지 말해 줍니다. 부엌 싱크대입니다.

    손끝 45 N은 대략 5킬로그램 남짓의 집기 힘입니다. 머그컵, 접시, 장바구니에는 충분합니다. 무언가를 으스러뜨리기에는 명백히 부족하고, 그것이 의도입니다. 다만 이것은 동시에 천장이기도 하고, 실제로 걸리는 천장입니다. 투명성에 맞춰 튜닝한 손으로 꽉 잠긴 병뚜껑을 열 수는 없습니다. 1X 내부 어딘가에서 누군가는 그 능력이 유리컵이 미끄러지는 것을 아는 능력보다 덜 중요하다고 판단했습니다. 제 생각에는 맞는 판단이지만, 이것은 입증된 사실이 아니라 가설에 건 선택입니다.

    전단력 센서가 혼자서는 아무것도 못 하는 이유

    또 하나의 층은 촉각 스택입니다. 1X는 손끝 피부가 수직 항력, 접촉 위치, 그리고 전단력을 감지한다고 밝혔습니다. 흥미로운 것은 전단력입니다.

    전단력은 옆으로 미끄러지는 힘이고, 파지가 무너지기 시작할 때 나타나는 물리적 신호입니다. 전단력을 충분히 일찍 잡아내면 로봇공학에서 ‘초기 미끄러짐’이라고 부르는 순간을 포착할 수 있습니다. 물체가 손을 떠난 뒤가 아니라 떠나기 전에 힘을 더 줄 수 있습니다.

    그런데 여기에 의존 관계가 있습니다. 관절이 미세한 힘 조정으로 반응하지 못하면 전단력 감지는 거의 쓸모가 없습니다. 센서와 동력 전달부는 짝으로만 작동합니다. 훌륭한 촉각 피부를 150:1 감속기 위에 붙이면, 자기가 어떻게 실패하고 있는지 정확히 알면서 아무것도 못 하는 손이 됩니다.

    부엌을 겨냥한 설계라는 신호

    이 두 사양을 함께 넣었다는 점은 따로 볼 만합니다. 산업용 로봇 손에는 보통 필요 없는 사양입니다. 공장 셀 안에서는 물에 담글 일이 없고, 사람이 먹을 것에 닿을 일도 없습니다.

    이 두 사양이 함께 가리키는 작업은 설거지입니다. 그릇을 물에 담그고, 세제를 묻히고, 음식물이 묻은 표면을 만지는 작업입니다. 그리고 설거지는 힘 제어가 특히 까다로운 일이기도 합니다. 젖은 그릇은 마찰계수가 급격히 떨어져서 미끄러지고, 그릇마다 견디는 힘이 다릅니다. 유리컵과 스테인리스 냄비를 같은 세기로 쥐면 하나는 깨집니다.

    즉 방수 등급과 낮은 기어비는 서로 다른 사양이 아니라 같은 목표를 향한 두 가지 선택입니다. 젖은 손으로 미끄러운 것을 다루겠다는 이야기입니다. 사양표를 이렇게 읽으면 이 회사가 어떤 작업을 먼저 풀려는지가 드러납니다.

    가정이 공장보다 어려운 이유

    로봇 손이 이미 공장에서 잘 쓰이고 있는데 왜 가정용이 이렇게 어려운지는 짚고 갈 필요가 있습니다. 직관과 반대로 보이기 때문입니다.

    산업용 그리퍼가 성립하는 이유는 문제가 미리 제거되어 있기 때문입니다. 공장 셀에서는 집을 물체가 정해져 있습니다. 같은 부품이 같은 무게로 같은 표면 상태로 옵니다. 지그와 컨베이어가 위치까지 알려 줍니다. 그러면 손은 아무것도 느낄 필요가 없습니다. 이 좌표로 가서 이만큼 닫으라는 명령을 반복하면 됩니다. 파지력도 최악의 경우에 맞춰 넉넉하게 잡아 두면 그만입니다. 어차피 대상이 그 부품 하나뿐이라 으스러질 걱정도 없습니다. 산업용 그리퍼가 손가락 두 개로 충분한 이유가 여기 있습니다. 자유도가 필요 없는 것이 아니라, 환경이 자유도를 대신 해결해 준 것입니다.

    가정에는 그 전제가 하나도 없습니다. 싱크대에 쌓인 그릇은 매번 다른 조합이고, 형상도 무게도 제각각입니다. 유리컵과 플라스틱 용기와 무쇠 팬을 같은 손이 연달아 집어야 하는데, 견디는 힘의 범위가 수십 배 차이 납니다. 위치도 알려 주지 않습니다. 한 번 놓았던 자리에 그 컵이 다시 있으리라는 보장이 없습니다.

    가장 성가신 변수는 마찰계수입니다. 물체를 떨어뜨리지 않으려면 필요한 쥐는 힘은 무게만이 아니라 표면 마찰에 달려 있는데, 이 값이 가정에서는 계속 변합니다. 마른 컵과 젖은 컵이 다르고, 세제 거품이 묻으면 또 다릅니다. 기름기가 묻은 접시는 완전히 다른 물건입니다. 사전에 계산할 수 있는 값이 아니라는 뜻입니다. 그래서 미끄러짐을 실시간으로 잡아내고 그 자리에서 힘을 올리는 것 말고는 방법이 없습니다. 앞서 본 전단력 감지와 백드라이버블 관절이 필요한 이유가 정확히 이 지점입니다.

    이 조합을 같은 맥락에서 다시 보면 의미가 달라집니다. 단순히 튼튼하다는 표시가 아닙니다. 물에 담그고, 세제를 묻히고, 사람이 먹을 것에 닿는 작업을 하겠다는 선언입니다. 그리고 그 작업이 바로 마찰계수가 가장 심하게 흔들리는 작업입니다. 손이 젖어도 죽지 않는 것과 젖은 물체를 놓치지 않는 것은 다른 문제이고, 부엌에 들어가려면 둘 다 필요합니다. 이 회사는 두 가지를 한 손에 넣으려 하고 있습니다.

    하드웨어가 데이터의 상한을 정합니다

    여기서 이야기가 손 하나보다 커집니다. 제가 이 글을 쓰고 싶었던 진짜 이유이기도 합니다.

    전에 로봇의 진짜 병목은 하드웨어가 아니라 데이터라고 썼습니다. 이 분야의 무게중심이 “누가 가장 좋은 몸을 만드는가”에서 “누가 가장 많은 경험을 모으는가”로 옮겨 갔다는 이야기였습니다. 지금도 그 판단은 유효하다고 봅니다.

    다만 이 손은 제가 과소평가했던 것을 드러냅니다. 하드웨어가 데이터에 담길 수 있는 것의 상한을 정합니다.

    휴머노이드 조작 데이터가 실제로 어떻게 만들어지는지 짚어 보면 분명해집니다. 사람이 로봇을 원격조종하고, 로봇은 그 과정을 기록하고, 그 기록이 학습용 시연 데이터가 됩니다. 이 과정을 백드라이버블하지 않은 손으로 돌린다면 무엇이 기록되는지 따져 볼 문제입니다.

    관절 각도가 기록됩니다. 타임스탬프가 기록됩니다. 카메라 프레임이 기록됩니다. 그리고 기구적으로 측정 자체가 불가능했기 때문에 기록되지 않은 것이 있습니다. 손가락이 얼마나 세게 눌렀는지, 물체가 움직이는 동안 그 압력이 어떻게 변했는지입니다.

    그런 데이터를 아무리 크게 쌓아 아무리 큰 모델을 학습시켜도 힘 제어는 배우지 못합니다. 모델이 작아서도 아니고 데이터가 짧아서도 아닙니다. 그 변수가 애초에 파일에 없기 때문입니다. 힘을 느끼지 못하는 하드웨어는 힘을 가르칠 수 없는 데이터를 만듭니다.

    원격조종 쪽에서 보면 더 분명해집니다. 백드라이버블하지 않은 손으로 원격조종을 하는 사람은 손끝에서 아무 저항도 느끼지 못합니다. 화면만 보고 조종합니다. 그러면 그 사람이 만들어 내는 시연 자체가 힘 조절이 빠진 시연이 됩니다. 데이터에 힘 정보가 기록되지 않는 것을 넘어, 애초에 사람이 보여 준 동작에 힘 조절이라는 행위가 들어 있지 않습니다. 손이 힘을 되돌려 줘야 조종하는 사람도 힘을 쓸 줄 아는 시연을 만듭니다.

    이 구분은 한 번 더 밀고 갈 가치가 있습니다. 백드라이버블 하드웨어는 데이터를 모으는 쪽과 그 데이터를 재생하는 쪽에서 서로 다른 일을 하기 때문입니다.

    모으는 쪽부터 보면, 조종기가 조종자에게 힘을 되돌려 주지 않는 구조를 로봇공학에서는 단방향 제어라고 부릅니다. 목표 위치 값만 로봇 쪽으로 흘려보내는 방식입니다. 구현이 쉽고 천천히 움직이는 비접촉 작업에는 잘 맞습니다.

    문제는 접촉이 많은 작업에서 드러납니다. 조종자는 로봇 손끝이 물체에 닿았는지조차 손으로는 알 수 없고, 화면에서 물체가 찌그러지는 것을 보고 나서야 압니다. 사람의 촉각 반사는 수십 밀리초 단위로 작동하는데, 눈으로 보고 판단해서 손을 되돌리는 경로는 그보다 한참 느립니다. 그래서 힘 피드백이 없는 조종은 늘 과하게 쥐거나 부족하게 쥐는 쪽으로 치우칩니다.

    여기서 나오는 결론이 중요합니다. 이런 조종으로 얻은 시연은 힘 정보가 빠진 데이터가 아니라, 애초에 힘 조절이 서툰 사람의 데이터입니다. 나중에 손끝에 센서를 붙여 그 시연을 다시 기록한다 해도, 기록되는 것은 조종자가 잘못 준 힘의 궤적입니다. 라벨은 정확해지지만 가르치려는 행동 자체가 틀려 있습니다. 힘 피드백 유무가 시연 품질과 모방 학습 결과에 모두 영향을 준다는 것을 실험으로 확인한 연구가 같은 작업을 여러 조건으로 나눠 비교하는 이유가 이것입니다. 이 연구는 힘 피드백을 준 조건에서 얻은 시연으로 학습한 정책이 힘 데이터를 직접 입력받지 않고도 더 빠르고 안전하게 작동한다는 점도 함께 보고합니다. 데이터의 질은 기록 장치만이 아니라 조종 경험에도 달려 있습니다.

    재생하는 쪽은 별개의 문제입니다. 학습이 끝난 정책이 “여기서 3 N으로 쥐어라”라는 출력을 내놓아도, 그 명령을 실행할 관절이 3 N을 만들어 내지 못하면 아무 소용이 없습니다. 고감속 관절은 내부 마찰이 그 값보다 크기 때문에 명령과 실제 출력 사이에 알 수 없는 오프셋이 생깁니다. 학습된 힘 정책을 실행할 수 없는 몸에 올려놓은 상태입니다.

    그래서 백드라이버블 하드웨어는 양쪽에서 각각 필요합니다. 모으는 쪽에서는 조종자가 힘을 느끼게 해서 애초에 제대로 된 시연이 나오게 하고, 재생하는 쪽에서는 학습된 힘 명령이 실제 힘으로 나가게 합니다. 두 역할 중 하나라도 빠지면 전체가 무너집니다. 힘을 느끼는 하드웨어로 모은 데이터를 느끼지 못하는 로봇에 넣으면 실행이 안 되고, 반대로 좋은 손에 힘 조절이 빠진 데이터를 넣으면 배울 것이 없습니다.

    이 차이는 나중에 메울 수 없다는 점이 중요합니다. 해상도가 낮은 영상은 나중에 다시 찍으면 됩니다. 라벨이 부실하면 다시 붙이면 됩니다. 그런데 측정되지 않은 접촉력은 복원 대상이 아닙니다. 그 순간에 센서가 없었다면 그 값은 어디에도 존재한 적이 없습니다.

    이 관점은 수직 계열화 이야기의 의미도 바꿉니다. 1X는 모터, 전자부품, 내장 센서, 텐던 소재, 폴리머 스킨, 손 전용 펌웨어까지 사내에서 만든다고 밝혔습니다. 보통 “모터를 직접 만든다”는 말은 원가와 공급망 이야기로 읽히고, 실제로 일부는 그렇습니다.

    그런데 동력 전달부와 스킨과 펌웨어를 통제하는 회사는 자기 데이터셋에 어떤 열이 존재할지를 통제합니다. 스펙 시트에는 절대 나타나지 않는 복리형 우위입니다. 그리고 이것은 LG가 몸은 만들고 뇌는 가져다 쓰기로 한 결정과 같은 구조의 논리인데 방향만 반대입니다. LG는 뇌가 범용재라는 쪽에 걸었습니다. 1X는 몸이 뇌가 배울 수 있는 것의 범위를 정한다는 쪽에 걸었습니다.

    텐던도 전완 배치도 1X만의 것이 아닙니다

    차별점은 헤드라인이 암시하는 것보다 좁습니다. 이 점을 정확히 해 둘 필요가 있습니다.

    텐던 구동과 전완 모터 배치는 1X 독점이 아닙니다. 테슬라 옵티머스 V3 특허도 액추에이터를 전완으로 옮기고 케이블을 손목으로 통과시키는 텐던 구동 구조를 담고 있습니다. 손가락과 손바닥에 22 자유도, 손목에 2 자유도를 두는 구성으로, 손목이 3축인 1X와는 이 지점에서 갈립니다. 사람의 악력 근육이 실제로 전완에 있다는 당연한 이유 때문에, 업계 상당수가 전완 쪽으로 가고 있습니다.

    그러니 텐던은 핵심이 아닙니다. 기어비가 핵심입니다. 좁은 주장이고, 좁은 주장이 대체로 참인 주장입니다.

    여기서 한 가지 덧붙일 것이 있습니다. 텐던 구동이라고 해서 자동으로 백드라이버블해지지는 않습니다. 텐던은 힘을 전달하는 방식일 뿐입니다. 전완에 있는 모터에 고감속 기어박스를 물리면, 케이블로 연결했든 아니든 그 관절은 여전히 역구동이 어렵습니다. 케이블 자체의 마찰과 늘어남이 더해지면 오히려 나빠질 수도 있습니다. 백드라이버블 여부를 정하는 것은 전달 방식이 아니라 감속비입니다. 그래서 “텐던 구동 손”이라는 표현만으로는 두 손을 구분할 수 없습니다.

    텐던이 실제로 청구하는 비용

    텐던 구동에는 구조적으로 따라오는 비용이 있고, 스펙 시트에는 잘 적히지 않습니다.

    가장 근본적인 제약은 텐던이 당기기만 할 뿐 밀지 못한다는 점입니다. 줄은 잡아당길 때만 힘을 전달합니다. 밀면 그냥 휘어집니다. 그래서 관절 하나를 양방향으로 움직이려면 방법이 둘뿐입니다. 굽히는 줄과 펴는 줄을 각각 달아 서로 반대로 당기게 하거나(길항 구조), 한쪽만 줄로 당기고 반대 방향은 복원 스프링에 맡기는 방식입니다.

    둘 다 대가가 있습니다. 길항 구조는 관절 하나당 줄이 두 가닥, 사실상 액추에이터가 두 개 필요합니다. 22축을 전부 이렇게 구성하면 전완이 감당해야 할 부품 수가 급증합니다. 스프링 복원 방식은 부품 수를 절반으로 줄이지만, 펴는 방향의 힘이 스프링이 내주는 만큼으로 고정됩니다. 능동적으로 밀어내지 못한다는 뜻입니다. 그리고 스프링은 항상 당기고 있으므로, 관절을 굽힌 상태로 유지하는 동안 모터는 계속 스프링과 싸워야 합니다. 정지 상태에서도 전력을 씁니다.

    두 번째 비용은 정밀도 쪽에서 나옵니다. 줄은 경로를 따라 도르래와 튜브를 지나가는데, 그 경로마다 마찰이 붙습니다. 모터가 10 mm를 감았을 때 손끝이 정확히 그만큼 움직인다는 보장이 없습니다. 마찰이 걸리면 줄이 국소적으로 늘어나면서 감은 만큼이 그대로 전달되지 않고, 방향을 바꿀 때는 히스테리시스(이력 오차)가 생깁니다. 굽힐 때와 펼 때 같은 모터 각도가 서로 다른 손끝 위치에 대응한다는 뜻입니다.

    세 번째는 시간에 따른 변화입니다. 줄은 장력을 받은 상태로 오래 있으면 늘어납니다. 텐던 구동 로보틱스 문헌에서 공통적으로 지적되듯, 소재별로 이 크리프(장력을 받은 상태에서 서서히 늘어나는 성질) 특성이 크게 다릅니다. 폴리머 계열 줄은 가볍고 유연한 대신 크리프와 히스테리시스가 크고, 강선은 튼튼한 대신 굽힘 구간에서 손실이 쌓입니다. 어느 쪽이든 사용 시간이 쌓이면 기준점이 조금씩 밀립니다.

    이 문제가 실제로 어떤 형태로 나타날지는 지켜볼 지점입니다. 기준점이 밀리면 정기적인 장력 재조정이나 재캘리브레이션이 필요해질 수 있습니다. 관절측에 별도 위치 센서를 두고 모터측 값과 대조해 보정하는 방법이 텐던 구동 손의 마찰을 다룬 DLR 연구 등에 제시되어 있고, 1X가 내장 센서를 직접 만든다고 밝힌 만큼 어떤 형태로든 대응책을 넣었을 가능성은 있습니다. 다만 NEO의 손이 이 문제를 어떤 방식으로 다루는지, 사용자가 주기적으로 관리해야 하는 항목이 있는지는 공개된 자료로는 확인되지 않습니다. 텐던 손을 볼 때 자유도보다 먼저 물어야 할 질문은 “1년 뒤에도 같은 정밀도가 나오는가”입니다.

    사람 손과 비교하면 어디쯤인가

    기준점을 하나 잡아 두면 이 숫자들을 읽기가 쉬워집니다.

    사람 손의 자유도는 세는 방식에 따라 21축에서 27축 사이로 봅니다. 손목을 포함하는지, 엄지를 몇 축으로 모델링하는지에 따라 달라집니다. 25축이면 관절 개수만으로는 사람에 꽤 가까워진 셈입니다. 다만 사람 손에서 실제로 중요한 것은 관절 수가 아니라 제어의 질입니다. 사람은 달걀을 깨지 않고 쥐면서 동시에 병뚜껑을 열 수 있습니다. 같은 손으로 힘의 범위를 넓게 오갑니다.

    로봇 손은 아직 그 범위를 한 번에 갖지 못합니다. 힘을 키우면 감속비가 올라가고 감각을 잃습니다. 감각을 살리면 힘의 천장이 낮아집니다. 1X는 후자를 골랐습니다. 사람 손을 흉내 내는 경쟁이라기보다는, 사람 손이 가진 두 가지 능력 중 어느 쪽을 먼저 확보할지 고르는 문제에 가깝습니다.

    그래서 45 N이라는 숫자를 “사람보다 약하다”로만 읽으면 절반만 읽은 것입니다. 사람의 최대 악력에는 한참 못 미치지만, 사람이 일상에서 실제로 쓰는 힘의 대부분은 최대치 근처가 아닙니다. 컵을 들고 문고리를 돌리고 옷을 개는 데 필요한 힘은 그보다 훨씬 작습니다. 문제는 그 작은 힘을 정확히 조절하는 쪽이었고, 1X는 거기에 자원을 몰았습니다.

    회의적으로 볼 지점들

    찬물을 끼얹을 대목이 적지 않습니다.

    낮은 기어비를 택하면 최대 파지력 하나만 포기하는 것이 아닙니다. 텐던은 늘어나고, 닳고, 끊어집니다. 그리고 뻣뻣한 기어 관절보다 훨씬 정교한 제어를 요구합니다. 손을 안전하게 만드는 그 유연함이 정밀한 명령을 어렵게 만들기도 합니다. 1X는 수백만 사이클 시험을 언급했지만, Embodied Global의 기술 분석이 지적하듯 평균 고장 간격, 접촉 수명, 현장 수리 비용을 검증하는 제3자 데이터는 공개되지 않았습니다. 통제된 하중에서 돌린 실험실 사이클과 실제 부엌에서의 1년은 다른 시험입니다.

    모터를 전완으로 옮긴 것은 공간 문제를 풀면서 질량 문제를 만듭니다. 그쪽으로 옮긴 액추에이터 하나하나가 긴 지렛대 끝에서 휘둘리는 부위의 무게와 관성을 늘립니다. 손가락이 가벼워지고 팔이 무거워집니다. 휴머노이드에 공짜 공간은 없습니다. 문제를 옮기는 것이지 없애는 것이 아닙니다.

    연 1만 개라는 숫자는 회사가 밝힌 내부 생산 능력이지 실증된 생산 실적이 아닙니다. 제조 능력 발표는 후공정이 전부 협조한다는 가정 아래 이론적으로 그만큼 돌 수 있는 라인을 묘사해 온 긴 역사가 있습니다.

    수직 계열화도 양면입니다. 데이터 관점에서는 강점이지만, 제조 관점에서는 모든 공정의 수율을 혼자 책임진다는 뜻입니다. 폴리머 스킨 하나가 안 나오면 손 전체가 멈춥니다. 부품을 사 오는 회사는 공급처를 바꾸면 되지만, 직접 만드는 회사는 스스로 고쳐야 합니다. 손 하나에 25축이 들어가고 그 안에 촉각 센서와 텐던과 스킨이 겹겹이 쌓여 있다는 점을 생각하면, 조립 난도와 불량률은 일반적인 로봇 부품보다 높은 쪽에 가깝습니다.

    정밀도 수치도 맥락을 봐야 합니다. ±0.2 mm는 위치 정밀도이고, 무부하에 가까운 조건에서 재는 값입니다. 유연한 관절은 하중이 걸리면 그만큼 휩니다. 힘의 투명성을 얻기 위해 설계에 넣은 그 유연함이, 무거운 것을 들 때는 위치 오차로 나타납니다. 두 수치는 같은 조건에서 동시에 성립하는 값이 아닐 가능성이 큽니다. 이 부분은 공개된 자료로는 확인되지 않습니다.

    그리고 가장 큰 경고 사항은 손에 관한 것이 아닙니다. NEO는 일시불 2만 달러 또는 월 499달러의 얼리 액세스 단계에 있고, TNW 보도를 비롯한 여러 매체는 익숙하지 않은 작업이 여전히 원격조종 ‘엑스퍼트 모드’를 거친다고 전합니다. 원격의 사람이 개입한다는 뜻입니다.

    접촉을 느끼는 손이 장기 자율성을 주지는 않습니다. 실패 복구도 주지 않습니다. 시연과 제품을 실제로 가르는 능력이 바로 이것입니다. 파지에 실패했음을 알고, 왜 실패했는지 파악하고, 시키지 않아도 다른 방법을 시도하는 능력입니다. 이것은 소프트웨어 문제이고, 더 좋은 하드웨어로 출발선이 조금 앞으로 당겨질 뿐입니다.

    이미 나온 가정용 로봇은 전부 손이 없습니다

    한국 독자 입장에서 이 이야기를 가장 빠르게 체감하는 방법이 있습니다. 이미 상용화된 로봇들을 떠올려 보는 것입니다.

    로봇청소기는 여러 가정에 보급되어 있습니다. 식당용 서빙 로봇도 상용화 단계에 들어섰습니다. 건물 안에서 물건을 옮기는 배송 로봇, 물류창고의 이송 로봇도 실제로 돌아갑니다. 이들의 공통점이 있습니다. 전부 다관절 손이 없습니다. 있어도 쟁반을 얹어 두는 선반이거나, 정해진 규격의 상자를 밀어 넣는 단순한 기구입니다. 물류 자동화 쪽에는 물건을 집어 옮기는 로봇 팔이 들어가 있지만, 그 손은 규격이 알려진 물체를 정해진 방식으로 집는 그리퍼에 가깝습니다.

    우연이 아닙니다. 이동은 상당 부분 풀린 문제이고 조작은 아직 풀리지 않은 문제이기 때문입니다. 이동은 잘 정의된 문제입니다. 지도를 만들고, 자기 위치를 알아내고, 장애물을 피해 경로를 짭니다. 센서는 카메라와 라이다면 대체로 충분하고, 바닥은 평평하며, 실패해도 대개는 멈추면 됩니다. 조작은 그렇지 않습니다. 물체마다 조건이 다르고, 성공과 실패가 접촉면에서 밀리미터와 뉴턴 단위로 갈리며, 실패하면 물건이 깨집니다.

    이것은 모라벡의 역설이라고 불려 온 오래된 관찰과 같은 이야기입니다. 사람에게 어려운 일이 기계에는 쉽고, 사람이 생각 없이 해내는 일이 기계에는 어렵다는 관찰입니다. 어른의 계산 능력은 진작 넘어섰지만, 어질러진 식탁에서 처음 보는 물건을 집어 드는 일은 아직입니다. 그 일을 걸음마를 뗀 아이는 아무 생각 없이 해냅니다.

    그래서 국내에서 논의되는 돌봄 로봇이나 가사 지원 로봇이 대체로 이동·말벗·모니터링 기능부터 나오는 것도 같은 맥락입니다. 손이 필요 없는 기능부터 상용화되고 있습니다. 반대로 설거지, 빨래 개기, 침대 정리처럼 실제로 사람의 시간을 많이 잡아먹는 집안일은 전부 손이 필요한 일이라 아직 남아 있습니다. 고령화 속도가 빠른 사회에서 정말 필요한 것은 말벗보다 손 쪽에 가깝습니다.

    이 관점에서 보면 NEO의 손이 왜 주목받는지가 또렷해집니다. 손이 잘 되는지가 가정용 로봇이 청소기 다음 단계로 갈 수 있는지를 가릅니다. 그리고 그 손을 가르는 숫자가 자유도가 아니라 기어비라는 것이 이 글의 요지입니다.

    살 수 있는 물건인가

    기대치를 먼저 짚어 두는 것이 좋습니다.

    NEO는 공개 시점 기준으로 얼리 액세스 단계입니다. 일시불 2만 달러 또는 월 499달러 구독으로 제시되었고, 예약에는 보증금이 붙습니다. 초기 인도는 미국 시장이 먼저이고, 다른 시장으로의 확대는 그다음 단계로 예고되어 있습니다. 국내 정식 유통 계획이나 가격, 전파·전기 인증 절차에 대해서는 공개된 자료로는 확인되지 않습니다. 지금 한국에서 주문해 받을 수 있는 물건이 아니라고 보는 편이 안전합니다.

    그리고 영상을 볼 때 반드시 구분해야 할 것이 하나 있습니다. 로봇이 하는 동작에는 두 종류가 있습니다. 스스로 판단해서 하는 자율 동작과, 원격의 사람이 조종해서 하는 동작입니다. 1X는 이 원격조종을 숨기지 않고 제품 설계의 일부로 밝혔습니다. 익숙하지 않은 작업은 사람이 붙어서 대신 해 주고, 그 과정이 학습 데이터가 됩니다. 문을 열어 주거나 물건을 가져오거나 불을 끄는 정도의 기본 동작은 처음부터 자율로 제시되지만, 새로운 집안일은 아직 사람의 손을 거칩니다.

    그래서 시연 영상에서 손이 유려하게 움직이는 장면을 봤다면, 먼저 물어야 할 질문은 “저게 자율인가”입니다. 이 구분 없이 보면 실제보다 몇 년쯤 앞선 능력으로 오해하기 쉽습니다. 덧붙여 원격조종에는 사생활 문제가 따라옵니다. 원격 조종자가 로봇의 카메라를 통해 집 안을 본다는 뜻이고, 구매자는 이 조건에 동의해야 합니다. 성능 이전에 각자 판단할 문제입니다.

    그럼에도 이 손이 만드는 것

    그래서 어느 해든 가장 흥미로운 로봇은 대체로 가장 인상적인 로봇이 아닙니다. Physical AI에서 조용히 이기고 있는 기계는 병원 복도의 바퀴 달린 캐비닛이라고 썼고, 그 판단은 지금도 유효합니다. 지루하고, 좁고, 실제로 배치되어 있습니다. NEO의 손이 유용성에서 그 캐비닛을 이기려면 한참 걸립니다.

    다만 이 손은 캐비닛이 못 하는 일을 하고 있고, 그것이 무엇인지 정확히 말할 가치가 있습니다.

    업계 하드웨어 대부분이 물리적으로 포착하지 못하는 종류의 데이터를 만들어 내고 있습니다. 손재주가 정말 스케일링 법칙을 따른다면, 이 시기에 힘 정보가 주석된 시연 데이터를 모은 회사들은 나중에 아무도 소급해서 복원할 수 없는 데이터셋을 쥐게 됩니다. 영상은 언제든 더 찍을 수 있습니다. 하지만 그 자리에 센서가 없어서 기록되지 않은 접촉력을 나중에 돌아가서 측정할 방법은 없습니다.

    자유도 경쟁은 처음부터 다소 허영에 가까운 지표였습니다. 관절을 세는 일은 쉽고 사진도 잘 나옵니다. 손이 만지는 것을 느낄 수 있는지는 슬라이드에 담기 어렵습니다. 그리고 그 숫자에 따라 나머지 관절들이 의미를 갖기도 하고 갖지 못하기도 합니다.

    ※ 정보 제공 목적이며 투자 권유가 아닙니다.


    원문 출처 / Source: https://www.1x.tech/discover/neos-hands

    이미지: Franck V. / Unsplash

  • TL;DR — 1X unveiled a new hand for its NEO humanoid: 25 degrees of freedom (22 in the fingers and palm, 3 at the wrist), tendon-driven, with the motors sitting in the forearm instead of the hand. The spec that actually matters isn’t the DOF count — it’s the gear ratio. Most robot joints run 100:1 to 200:1 reduction. 1X runs roughly 5:1 to 15:1, which makes every joint natively force-controlled and fully backdrivable. That’s not a better hand. It’s a hand that flipped the trade-off: it gave up peak grip strength to buy force transparency. And the second-order effect lands somewhere unexpected — on the training data.

    The Story

    Every humanoid company eventually shows you a hand. It’s the money shot of the demo reel — fingers curling around a mug, threading a cable, doing something delicate enough to make you forget you’re watching a machine. And the number they put on the screen is almost always the same kind of number: degrees of freedom. Twenty-two. Twenty-five. Look how many joints.

    Here’s the problem with that number. It tells you what the hand can move. It tells you nothing about whether the hand can feel.

    1X’s new hand for NEO is worth looking at precisely because the company put a different number in front. According to 1X’s own technical writeup, the hand has 25 degrees of freedom — 22 fully actuated in the fingers and palm, plus 3 at the wrist — driven by tendons, with the motors housed in the forearm rather than crammed into the hand itself. Fine. Impressive, but not unusual. What’s unusual is the transmission: gear ratios of roughly 5:1 to 15:1.

    If that means nothing to you, stay with me, because it’s the whole story.

    A gearbox trades speed for torque. Small motor, big reduction, lots of force — that’s why industrial robot joints commonly run high-reduction harmonic drives, often in the 100:1 to 200:1 range. It works beautifully in one direction. The catastrophe is what it does in the other direction. Push back on a joint like that from the outside and the force has to fight its way backward through the entire gear train, against friction that doesn’t care which way you’re going. Reverse efficiency on a high-ratio harmonic drive typically lands in the 30-to-50 percent band, and drops below 40 percent in plenty of configurations. Some of them need three or four times more torque to backdrive than to drive forward.

    The practical consequence is brutal and rarely stated plainly: if the gearbox burns 5 Nm just overcoming its own internal friction, the joint is blind to every contact force below 5 Nm. Not “imprecise.” Blind. The signal doesn’t attenuate — it never arrives.

    And there’s a second penalty that’s arguably worse. Reflected inertia — how heavy the motor feels to the outside world — scales with the square of the gear ratio. Go from 10:1 to 150:1 and the apparent inertia at the joint doesn’t grow 15 times. It grows 225 times. A finger geared that hard isn’t a finger anymore. It’s a small hydraulic press that happens to be finger-shaped.

    So what does a hand like that actually do when you tell it to pick up a wine glass? It executes a position command. “Close to this angle.” It has no meaningful idea how hard it’s squeezing, because the only channel through which that information could travel is clogged with friction and inertia. It finds out the glass was slipping the same way you’d find out — from the sound.

    This is the trade 1X made, and it’s important to name it honestly as a trade. Dropping to single-digit and low-double-digit ratios means every one of those 25 joints is, in 1X’s phrasing, natively force-controlled and fully backdrivable. Force flows outward to the object; contact information flows back along the exact same mechanical path. No load cell bolted on as an afterthought, no inference from motor current fighting through a gear train. The transmission itself is the sensor. Call it “force transparency” — the hand and the world can feel each other.

    The published numbers are modest in a way that tells you this was a real choice, not a free lunch. Peak torque of 3.5 Nm at the thumb CMC joint, 2.6 Nm at the finger MCP joints, distal flexion forces up to 45 N, and 17.75 Nm at the wrist. Positioning accuracy of ±0.2 mm. IP68 rating with food-safe materials — which is 1X telling you where it thinks this robot lives, and the answer is your kitchen sink. 1X also says finger assemblies were tested through millions of cycles and wrist joints beyond 2 million under high load.

    Forty-five newtons at the fingertip is roughly five kilograms of pinch force. That’s enough for a mug, a plate, a bag of groceries. It is emphatically not enough to crush anything, and that’s the point — but it’s also a ceiling, and a real one. You cannot open a badly stuck jar with a hand tuned for transparency. Somewhere in a spreadsheet at 1X, someone decided that mattered less than knowing when a glass is sliding. My read is they’re right, but it’s a bet, not a fact.

    The other layer here is the tactile stack. 1X says the fingertip skin senses normal force, contact location, and shear — and shear is the interesting one. Shear is the sideways sliding force, and it’s the physical signature of a grip beginning to fail. Detect shear early enough and you catch what roboticists call “incipient slip,” which lets the hand tighten before the object is gone rather than after. But notice the dependency: shear sensing is nearly useless if the joints can’t respond with a fine-grained force adjustment. Sensor and transmission only work as a pair. Bolt great tactile skin onto a 150:1 gearbox and you’ve built a hand that knows precisely how it’s failing and can’t do anything about it.

    One more thing that’s easy to skip. 1X claims it makes the whole stack in-house — motors, electronics, embedded sensors, tendon materials, polymer skins, hand-specific firmware — on a dedicated line it says can produce up to 10,000 hands annually. That vertical integration claim is the part I’d flag hardest, and I’ll come back to it.

    The Takeaway

    Here’s where this gets bigger than one hand, and it’s the reason I wanted to write about it at all.

    I’ve argued before that the real bottleneck in robotics isn’t hardware, it’s data — that the field’s center of gravity moved from “who builds the best body” to “who collects the most experience.” I still think that’s right. But this hand exposes something I underweighted: hardware sets the ceiling on what the data can contain.

    Think about how humanoid manipulation data actually gets made. A human teleoperates the robot, the robot records what happened, and that recording becomes a training demonstration. Now run that through a non-backdrivable hand. What got recorded? Joint angles. Timestamps. Camera frames. What did not get recorded, because the mechanism was physically incapable of measuring it? How hard the fingers were pressing, and how that pressure changed moment to moment as the object shifted.

    You can train a very large model on a very large pile of that data and it will never learn force control. Not because the model is too small or the dataset is too short, but because the variable was never in the file. Hardware that cannot feel force generates data that cannot teach it.

    That reframes the vertical integration story. Everyone reads “we make our own motors” as a cost-and-supply-chain claim, and partly it is. But a company that controls the transmission, the skin, and the firmware controls what columns exist in its dataset. That’s a compounding advantage that doesn’t show up on a spec sheet, and it’s the same structural logic behind LG deciding to build the body and license the brain — except pointed the opposite way. LG bet the brain is the commodity. 1X is betting the body determines what the brain can ever learn.

    Worth noting the differentiator is narrower than the headlines suggest. Tendon drive and forearm-mounted motors are not 1X exclusives — Tesla’s Optimus V3 patents describe a tendon-driven hand with actuators moved into the forearm, routing cables through the wrist — 22 DOF in the fingers and palm plus a 2-DOF wrist, against 1X’s 3-DOF wrist. The forearm is where a lot of the industry is going, for the obvious reason that it’s where your own grip muscles are. So the tendons aren’t the story. The gear ratio is. That’s a narrow claim, and narrow claims are usually the true ones.

    Now the cold water, because there’s plenty.

    Low gear ratios cost you more than peak force. Tendon systems stretch, fray, and wear, and they demand far more sophisticated control than a stiff geared joint — the compliance that makes the hand safe also makes it harder to command precisely. 1X reports millions of test cycles, but as the technical writeup at Embodied Global points out, there’s no published third-party data on mean time between failures, contact lifetime, or field repair cost. Lab cycles under controlled load and a year in a real kitchen are different tests.

    Moving motors to the forearm solves a packing problem and creates a mass problem. Every actuator you relocate there adds weight and inertia to a segment that swings at the end of a long lever. Lighter fingers, heavier arm. There’s no free space in a humanoid — you move a problem, you don’t delete it.

    The 10,000-hands figure is a stated internal capacity, not a demonstrated run rate. Manufacturing capacity announcements have a long and unimpressive history of describing a line that could theoretically run at that volume if everything downstream cooperated.

    And the largest caveat isn’t about the hand at all. NEO is in early access at $20,000 outright or $499 a month, and reporting on the launch — TNW’s coverage among others — notes that unfamiliar tasks still route through a teleoperated “Expert Mode,” meaning a remote human. A hand that feels contact does not give you long-horizon autonomy. It doesn’t give you failure recovery, the thing that actually separates a demo from a product: knowing the grasp failed, figuring out why, and trying a different approach without being told. That’s a software problem, and better hardware only moves the starting line.

    Which is why the most interesting robot in any given year is usually not the most impressive one. I’ve made the case that the machine quietly winning Physical AI is a cabinet on wheels in a hospital hallway — boring, narrow, and actually deployed. That’s still true. NEO’s hand won’t beat the cabinet on usefulness for a long time.

    But it’s doing something the cabinet can’t, and it’s worth being precise about what. It’s generating a category of data that most of the industry’s hardware physically cannot capture. If dexterity really does follow a scaling law, then the companies that spent this era collecting force-annotated demonstrations will have a dataset nobody can retroactively reconstruct. You can always collect more video. You cannot go back and measure a contact force that no sensor was there to record.

    The DOF race was always a bit of a vanity metric. Counting joints is easy, and it photographs well. Whether the hand can feel what it’s touching is harder to put on a slide — and it’s the number that decides whether any of the joints matter.

    This article is for informational purposes only and is not investment or purchasing advice.

    Photo: Franck V. / Unsplash