Wiki / concepts / wiki

Quality-gated data omission

The selection-bias mechanism dc-rainmaker uncovered in Apple's heart-rate accuracy study (apple-watch Ultra 4 review): the sensor writes a reading only when an internal confidence bar is met, and writes nothing otherwise. Apple's advertised "every 5 seconds" is an average across a workout β€” sometimes 1-second cadence, sometimes 20-second silences. Apple confirmed the study then grades only the values the watch chose to list against the Polar H10 baseline, so every low-confidence stretch β€” precisely the moments the sensor was implicitly wrong β€” is excluded from scoring, while competitors that plot a value every second get graded on all of them.

His analogy: "If this was a written test, every single time Apple wasn't sure about an answer, they just wouldn't put down the answer. And instead of docking them points… the test proctor said, 'Oh, if there's no answer there, we will consider that right.' Okay, Apple."

Observed gap lengths in his own files: 11s (trail run), 12s (Peloton interval peak), 14/18/20s (road ride), ~40s reading 5-7 bpm high (track). The sensor is still genuinely accurate when it does report β€” the point is not that Apple's sensor is bad, but that this grading rule would put "any other wearable at the very bottom of the chart" if applied asymmetrically.

Transferable lesson: when a benchmark's subject controls which samples enter the benchmark, the missing data is the finding. Parent pattern: manufacturer-study-skepticism.

Source: Ultra 4 in-depth review

Linked from

Apple Watch (Series 12 / Ultra 4)DC Rainmaker (Ray)Manufacturer-study skepticismRLHF is not RL