How Can Health Professionals Navigate Data From Wearables? Thresholds and the Measurement Problem Clinicians Are Not Being Told About.

KLES Digital Health | Clinical and metrological considerations for safe wearable integration

Consumer wearable are available and widely used, making it harder for health professionals to negotiate

It’s becoming increasingly common that people attend consultations with health professionals and in addition to Dr Google, there are an array of figures from a smart watch which had indicated that they are seriously ill. All of which have to be taken seriously by staff and this can lead to secondary investigations and staff time to rule out issues which are not present and wouldn’t have been known about if it were not for the dreaded devices. Hence some professionals roll their eyes and regard the devices as a nuisance which adds unnecessary burden to already busy services, perhaps dismissing something that might have given them some useful information or education for the patient.

One patient presents with a consumer wearable printout showing an irregular heart rhythm notification. Another attends having reviewed their smartwatch data and believes they may have had an unrecognised seizure. A third is convinced their wearable confirms they have a serious sleep disorder. In each case, the device has produced a binary classification - detected or not detected and the patient has treated that binary output as a clinical fact.

The question for the clinician is not simply whether the device is accurate in the population sense. It is whether, for this reading, at this moment, the device's classification is reliable. That question requires a different kind of reasoning than most wearable guidance currently provides.

The Fundamental Metrology Problem: Threshold Classification Under Uncertainty

Every wearable alarm system applies a threshold to a continuous measurement to produce a categorical output. The device measures something a photoplethysmographic signal, an accelerometer trace, a derived inter-beat interval sequence and compares a derived quantity against a cut-off. The output is binary: the threshold was crossed, or it was not.

This is where the central clinical problem lies, and it is a problem of metrology rather than of technology. All physical measurements carry uncertainty. PPG-derived heart rate and rhythm estimates carry uncertainty arising from motion artefact, optical noise, skin perfusion variation, sensor contact, and ambient interference. The proprietary algorithms applied to the raw signal add further uncertainty through their own modelling assumptions, and those assumptions are not published, not independently validated across clinical populations, and not accessible to the clinician interpreting the output.

The relationship between measurement uncertainty and classification error is non-linear at a threshold. When a true physiological value sits well above or below the decision boundary, measurement uncertainty has a negligible effect on classification: the result is stable across the realistic range of measurement error. But as the true value approaches the threshold, a progressively smaller measurement uncertainty produces a progressively larger probability of misclassification. At the threshold itself, any uncertainty at all makes the classification essentially arbitrary.

Consumer wearables present a single binary output with no confidence interval, no indication of proximity to the decision boundary, and no mechanism for the clinician or patient to know whether the classification was stable or marginal. This is not a limitation that will be resolved by improved sensor hardware alone. It is a structural property of threshold-based classification systems applied to uncertain measurements, and it requires explicit clinical management.

Where This Matters in Practice

Three clinical applications in neurology are acutely vulnerable to this problem.

First, seizure-related heart rate detection. Clinical research establishes that most focal impaired aware and generalised tonic-clonic seizures produce a heart rate rise exceeding 10 bpm or 20% from individual baseline. Many patients with epilepsy are unaware of their own seizures, making wearable-derived physiological data potentially valuable as an objective record. But a measured rise of exactly 10 bpm, derived from PPG under ambulatory conditions, carries substantial uncertainty. A true rise of 9 bpm and a true rise of 11 bpm are indistinguishable within the measurement error of a consumer smartwatch. At the threshold, the detection algorithm is resolving ambiguity by fiat, not by measurement. Clinicians advising patients to monitor for unexplained heart rate spikes as a possible seizure correlate must communicate that events near this boundary are clinically ambiguous, not clinically confirmed.

Second, atrial fibrillation classification. Consumer smartwatches and handheld ECG devices classify cardiac rhythms using proprietary algorithms applied to PPG or single-lead ECG signals. The performance of these algorithms in large healthy populations who are younger, with fewer comorbidities than typical neurology patients; shows reasonable positive predictive values in published studies. But population-level sensitivity and specificity statistics describe average behaviour; they do not characterise the uncertainty of any individual classification. A reading classified as irregular by a marginally performing algorithm on a patient with borderline rhythm irregularity, peripheral vascular disease, and an active workday is a very different proposition from a clearly irregular trace at rest. The binary output conceals this entirely.

Third, sleep staging. Consumer sleep wearables derive sleep stage classifications from accelerometry and PPG, without EEG. The boundaries between wake, light NREM, and deep sleep in these algorithms are inherently uncertain and unvalidated in neurologically diseased populations. Clinicians should not accept consumer sleep staging data for patients with suspected sleep disorders as equivalent to polysomnography, not because the devices are unreliable on average, but because the classification uncertainty at sleep stage boundaries is uncharacterised and the algorithms have not been validated in the populations presenting to sleep neurology services.

The Actigraphy Example: When Objective Data Contradicts Assumption

A useful illustration of what rigorous wearable-derived measurement can reveal and how it can overturn clinical assumptions comes from actigraphy research in cluster headache. Studies using the Empatica E4 wristband, a research-grade multimodal device recording accelerometry, PPG, electrodermal activity, and skin temperature, found that energy expenditure decreased in the pre-ictal period and during cluster headache attacks. This contradicted the widely held clinical assumption of patient pacing driven by pain-related agitation.

This finding is instructive on two counts. First, it demonstrates the genuine value of continuous objective physiological monitoring in conditions where patient-reported behaviour is unreliable or counterintuitive. Second, it illustrates the difference between research-grade data collection with characterised sensors and consumer wearable data collected under uncontrolled conditions. The Empatica E4 is not a consumer product. Its measurement properties are documented. The translation of findings from research-grade devices to consumer wearables cannot be assumed.

What Clinicians Should Ask That They Currently Do Not

Current clinical guidance on wearables focuses on whether devices are FDA-cleared or CE-marked, whether validation studies exist, and whether findings should be confirmed by medical-grade testing. These are necessary questions. They are not sufficient.

The questions that metrological rigour demands are:

•       What is the measurement uncertainty of this device for this signal in this patient population under real-world conditions? Not population-average accuracy, but the error distribution applicable to this individual reading.

•       How far from the classification threshold was this reading? A device that cannot answer this question cannot tell you how reliable its own output is.

•       Has the algorithm's behaviour near the decision boundary been characterised and published? For most consumer devices, the answer is no.

•       Was the measurement taken under conditions that inflate uncertainty such as physical activity, cold extremities, poor sensor contact, dark skin tone, peripheral vascular disease? These are not edge cases in clinical populations.

None of these questions are currently answerable from consumer device outputs as they stand. That is not an argument against using wearable data, it is an argument for using it with explicit acknowledgement of what is and is not known about the reliability of any given classification.

A Practical Clinical Framework

Until device manufacturers provide uncertainty quantification alongside their binary outputs, which metrological standards for medical devices should require - clinicians need a working framework for managing this gap.

•       Treat any wearable classification as a prompt for investigation, not a diagnosis. This applies with particular force to readings that are likely to be near a threshold like modestly elevated heart rate, borderline rhythm irregularity, marginal sleep staging.

•       Weight wearable alerts by signal strength, not merely by direction. An alert triggered by a reading well beyond the threshold is more credible than one triggered near it, even though both produce the same binary output.

•       Require confirmatory medical-grade testing before clinical action. The standard of care for AFib diagnosis, seizure evaluation, and sleep disorder assessment has not changed. Consumer wearables alter the trigger for applying that standard, not the standard itself.

•       Select patients explicitly. Anxiety from wearable alerts is a documented adverse outcome, not a theoretical risk. Patients who are hypervigilant, health-anxious, or have limited capacity to contextualise ambiguous data are poor candidates for unsupported wearable monitoring.

•       Assess data completeness before drawing conclusions. Adherence gaps in wearable datasets are common and material. Absence of detected events during periods of non-wear means nothing clinically.

•       Apply additional scepticism to consumer device data from neurologically and medically complex patients. Validation studies are predominantly conducted in healthy adults. The uncertainty characteristics of consumer wearables in diseased populations are largely unknown.

Wearable technology generates data of genuine clinical relevance. The task for clinicians is not to decide whether to engage with it, that decision has already been made by patients, but to engage with it on terms that reflect its actual measurement properties rather than the binary confidence its outputs imply.

Clinical basis: Benish SM et al. Wearable Technology and Its Role in Neurologic Care. Neurology. 2026;106:e214802.

Seizure HR threshold: Opherk C et al. Epilepsy Res. 2002;52:117-127. | Zijlmans M et al. Epilepsia. 2002;43:847-854.

Cluster headache actigraphy: Vandenbussche N et al. Brain Behav. 2024;14:e3360.

AFib detection: Perez MV et al. N Engl J Med. 2019;381:1909-1917. | Lubitz SA et al. Circulation. 2022;146:1415-1424.

Next
Next

From Ambition to Implementation - are Wearables Really Taking Hold in the NHS 10-Year Plan?