WearablesLong read

Wearable Cycle Tracking Accuracy in Irregular Perimenopause Cycles

Devices designed for regular cycles fail to track perimenopause's unpredictable patterns.

Columnist · · 11 min read
Cover illustration for “Wearable Cycle Tracking Accuracy in Irregular Perimenopause Cycles”
Wearables · September 28, 2026 · 11 min read · 2,438 words

Wearable cycle trackers are built and tested on regular, predictable cycles, and perimenopause is defined by the opposite: cycle lengths that swing unpredictably for years before menopause arrives. That mismatch is not a minor calibration issue. It is a structural problem, and understanding exactly where it comes from is what lets someone reading their Oura app or Clue log figure out how much of what it says to actually believe. Wearable Cycle Tracking Accuracy in Irregular Perimenopause Cycles.

Perimenopause, cycle patterns, and tracking tool reliability

Perimenopause is not a switch that flips. The clinical framework used to describe it, called STRAW+10, splits the transition into early and late stages, and the distinction matters because each stage behaves differently. In early perimenopause, cycles start varying by seven days or more from what used to be normal for that person, which is itself the clinical threshold used to define irregularity. By late perimenopause, gaps between periods can stretch to 60 days or beyond.

None of this variability is random static. Symptoms track the same erratic rhythm. A woman might deal with disruptive night sweats for weeks, feel entirely normal for two months, then skip a period altogether. Over 30 documented symptoms span physical, cognitive, and mood territory, and that range matters a great deal for the rest of this piece, because a wearable device only ever sees a narrow slice of that picture go-go-gaia.com.

Why does this matter for a tracking algorithm specifically? Because irregularity is not an edge case sitting outside the perimenopausal picture. It is the picture. A tool built around the assumption of regularity is, from the outset, applying the wrong template to the wrong population. Perimenopause is not a single event but a transition spanning 4 to 10 years, beginning as early as the mid-30s go-go-gaia.com.

Wearable cycle trackers and biological signal reading

Wrist-worn devices lean on a handful of core sensors: skin temperature, heart rate, interbeat interval, and electrodermal activity. Ring devices, Oura being the most cited example, tend to rely more on axillary or finger temperature paired with heart rate, and a 2026 meta-analysis in npj Digital Medicine (Shi et al.) found that ring form factors were more efficient at detecting the fertility window than other wearable types.

Progesterone triggers a small temperature rise after ovulation. Modern wearables try to do more: machine learning models are trained to classify cycle phase (follicular, ovulation, luteal, menses) from the pattern of temperature, heart rate, and related signals over time. One representative study, published in npj Women's Health (Kilungeja et al., 2025), trained a random forest model on skin temperature, EDA, interbeat interval, and heart rate from a wrist device across 65 cycles in 18 subjects, and got 87% accuracy classifying three phases. That is a respectable number. But it comes with a caveat: the subjects in that study had regular cycles, and the authors were explicit that generalizing the model to perimenopause was not established.

That caveat is not an isolated footnote. It reflects how most of this field's training data looks: predominantly reproductive-age women with predictable cycles, with perimenopausal users underrepresented in the datasets that taught these models what a "normal" cycle looks like.

What wearables are genuinely capable of picking up: temperature shifts, even irregular ones, sleep disruption, and changes in heart rate variability, all of which are present during perimenopause even when interpreting them is harder. What they cannot do, under any circumstance, is measure estradiol, FSH, LH, or progesterone directly. Every hormonal inference a wearable makes is downstream, a guess built from physiological echoes rather than the hormone levels themselves.

The accuracy numbers and the population they were measured on

Zoom out to the field-wide numbers, and the picture looks fairly strong. The 2026 Shi et al. systematic review and Bayesian meta-analysis in npj Digital Medicine, drawing on 27 studies through January 2025, found a pooled accuracy of 0.88, sensitivity of 0.79, and specificity of 0.80 npj Digital Medicine (Shi et al.). Those are the best aggregate figures the field has to offer.

But what population produced them? Treating 0.88 as a promise requires knowing what population produced them npj Digital Medicine (Shi et al.). The Kilungeja et al npj Women's Health. A fifteen-point drop, on paper, sounds almost survivable npj Women's Health (Kilungeja et al.). It is not, once the use case narrows to something closer to daily life.

Layer onto that a separate problem: the reference standards researchers use to check wearable predictions against are not standardized across the field. Some studies check against urinary LH, others against PdG, others against ultrasound or serum hormone panels, and a PMC review flagged that this inconsistency makes cross-study comparison unreliable. There is no single agreed gold standard. So the 0.88 pooled figure is not wrong, exactly npj Digital Medicine (Shi et al.). It just describes a population that mostly excludes the women asking, right now, whether they can trust their ring's fertility window prediction. Treat it as a ceiling on what is achievable under favorable conditions, not an expectation for what a perimenopausal user will actually experience npj Digital Medicine (Shi et al.). The Kilungeja et al.. Daily phase tracking compounds the problem: the same npj Women's Health study found accuracy dropped to 68% and AUC-ROC to 0.77 when classifying four phases on a sliding daily window, the most clinically relevant use case.

Diagram: Accuracy Drops When It Matters Most. Visualizes: Show a stepped or bar-style comparison of three accuracy figures drawn from the same npj Women's Health (Kilungeja et al., 2025) study, illustrating how performance degrades as conditions…

Irregular cycles and the broken assumptions of wearable algorithms

Regular-cycle algorithms are built around two anchor points: the onset of menstruation and the post-ovulatory temperature rise. Cycle length variance, in that design, is treated as a parameter to tune, not something that can break the model's basic structure. Perimenopause violates that assumption at the root, because ovulation itself becomes intermittent. Some cycles are anovulatory: there is no LH surge and no resulting progesterone bump, and without that bump, the post-ovulatory temperature shift the algorithm is watching for may never appear.

Meanwhile the hormones themselves are actively misleading in a way regular cycles never are. Estradiol can spike sharply before falling, producing a temperature or physiological pattern that mimics normal follicular activity even though nothing about the cycle is following a normal script. FSH climbs unevenly and LH secretion turns erratic, throwing occasional surges into a backdrop that is fundamentally noisy rather than clean. Sleep disruption, itself one of the defining perimenopause symptoms, degrades the temperature data quality the algorithm depends on, and night sweats introduce further artifact directly into the signal the device needs to be clean.

Cycle length itself stops behaving. The typical 24-to-35-day window that most algorithms are calibrated around expands to a range running from 24 days to 60 or more go-go-gaia.com. Every phase-length assumption the model learned during training starts to drift out from under it. Because ovulation may not happen at all in a given cycle, a device can report a fertile window that corresponds to nothing real, applying a regular-cycle template onto a cycle that never had the structure the template assumes.

None of this is an edge case sitting at the margins of the perimenopausal experience. It is the experience. A model performing at 87.46% on regular cycles and 72.51% on irregular ones is not describing an occasional miss, it is describing its default behavior against the population that will actually be using it during this life stage npj Women's Health (Kilungeja et al.).

Specific devices and apps responding to the irregular-cycle problem

Some of the field is actively working the problem rather than ignoring it. Oura updated its Cycle Insights feature in November 2025 with irregular-cycle performance as an explicit target. The update runs on a new neural network architecture designed to handle cycle variability, and predictions are now available from as little as one night of sleep data, down from the 60 nights previously required, with 12-month forward period and ovulation predictions now available, previously single-cycle only. Oura's underlying ovulation detection algorithm, published in JMIR, is cited above 96% accuracy, though that figure describes the algorithm's general performance, not its performance specifically in perimenopausal cycles Frontiers in Endocrinology Oura Cycle Insights Update. Oura notes that Cycle Insights does not apply to post-menopausal users and is designed for natural cycles without hormonal birth control, and the company has partnered with Stanford on the STIGMA study to study cycle physiology in underrepresented groups.

Clue has built a dedicated Perimenopause mode alongside an irregular period tracker, and syncs wearable data from Oura, WHOOP, and Fitbit, though the depth of that sync varies: Oura and WHOOP hand over temperature, sleep, and heart rate, while Fitbit contributes sleep and heart rate only. Go Go Gaia takes a different approach, keeping tracking active through irregular and skipped cycles and logging over 20 symptoms, including hot flashes and night sweats, then linking those symptoms to sleep, diet, and activity data, with a free core tier and an optional paid Premium layer. Balance, built by Dr. Louise Newson and ORCHA-certified, focuses on menopause more broadly rather than cycle prediction specifically. Midday offers a Mayo Clinic partnership with telehealth access, Health & Her provides a free symptom toolkit built around CBT exercises, pelvic floor training, and meditation, and Flo has added a Perimenopause Score aimed mainly at its existing user base.

The percentage improvements Oura reports (17%, 38%, and so on) are measured against the company's own prior algorithm version, not against a clinical gold standard for perimenopausal accuracy specifically Oura Cycle Insights Update. That is not a knock against the update. It is a reminder that "better than before" and "reliably accurate for this population" are two different claims, and the endocrine validation gap described earlier applies here just as much as anywhere else in the field. Period prediction was 29% more accurate overall 6 days before a period and 21% more accurate for irregular cycles at the same horizon Oura Cycle Insights Update. Ovulation prediction was 10% more accurate for typical cycle lengths and 17% more accurate for irregular cycles, 6 days before ovulation Oura Cycle Insights Update. Ovulation detection was 45% more accurate overall 4 days after ovulation and 38% more accurate for irregular cycles at that same horizon Oura Cycle Insights Update.

What wearables can still tell you during perimenopause

None of this means a wearable is useless during perimenopause. It means the question being asked of it needs to change.

Sleep tracking is one of the strongest cases. Perimenopause disrupts sleep in ways that are genuinely hard to self-report with any accuracy, and a wearable captures wake events, duration, and quality passively, independent of whether its cycle-phase prediction is right. Heart rate variability and resting heart rate reflect autonomic nervous system shifts tied to hormonal change and stress, and tracking how those metrics move over weeks or months can surface patterns to bring to a clinician. Even raw temperature data retains value on its own terms: when the algorithm's phase interpretation is off, the underlying temperature curve itself, erratic as it may look, can still be a signal that a cycle was anovulatory or that hormone levels are swinging hard.

Symptom logging adds another layer. Apps that let a user actively record hot flashes, night sweats, brain fog, and mood changes alongside the passive sensor data build a longitudinal record that holds clinical value regardless of whether the underlying cycle prediction was accurate that month. And dedicated sensors built around a narrower job can perform very well: the Tempdrop axillary armband, tested by Shpaichler et al. in Sensors (published October 13, 2025) across 194 cycles and 125 women, validated against the Clearblue Connected Ovulation Test System at 96.8% sensitivity and 99.1% specificity. That result shows dedicated temperature sensors, checked against an actual hormonal reference standard, can hit high performance, though the regularity of that particular study population still deserves scrutiny before generalizing it to perimenopause broadly.

The more useful frame, then, is not "is this prediction correct?" It is closer to: does this data, gathered over months rather than days, reveal a pattern to flag to a doctor? Wearables are longitudinal observation tools first, and prediction engines a distant second, at least for this population.

The limits of wearable data and the additional picture blood-based testing provides

The root problem was never sensor quality. It is that wearables infer hormonal state from downstream physiological signals instead of measuring the hormones themselves. Blood testing closes that gap directly, capturing estradiol, FSH, LH, progesterone, and AMH, the exact hormones whose erratic behavior is what breaks wearable cycle prediction in the first place.

FSH alone is not a clean answer, though. A single elevated reading does not confirm perimenopause, since levels swing erratically, and a "normal" FSH in a symptomatic woman over 45 does not rule the transition out either, particularly given that most reference ranges were calibrated on younger, reproductive-age populations. Serial measurement over months tells a far more reliable story than any single draw. AMH behaves differently and, in some ways, more usefully: it declines steadily with age rather than swinging with cycle phase, and a 2025 Frontiers in Endocrinology cohort study of 22,920 women found median AMH drops below 1.2 ng/mL by age 36, with the prevalence of diminished ovarian reserve climbing from 15.9% at age 18 to 96% by age 45 Oura Cycle Insights Update. That is ovarian reserve context no wearable can approximate. A comprehensive panel also checks TSH, since hypothyroidism reproduces nearly the entire perimenopause symptom list, fatigue, weight gain, mood changes, hair changes, and skipping that check risks misattributing a thyroid problem to a hormonal transition it has nothing to do with.

The 2025 European Society of Endocrinology guideline advises against routine biochemical testing to confirm perimenopause in symptomatic women over 45. That guidance is about avoiding redundant confirmation of something already clinically obvious, not an argument against testing when the picture is genuinely unclear or when the result would change how the transition gets managed.

A single panel is still a snapshot, and the same variability that makes one FSH reading hard to interpret makes serial panels, tracked across months, far more valuable than any isolated draw. That is the same logic that makes a wearable's longitudinal sleep and HRV data useful even when its cycle-phase prediction is shaky. Neither tool, on its own, resolves the uncertainty perimenopause creates. Together, one tracking daily physiological patterns and the other tracking the hormonal trajectory that produces them, they cover more of the picture than either could alone. Women who have spent months puzzling over what a wearable's fertility window even means deserve better than a shrug, and the biology behind why these tools struggle is, at minimum, a starting point for asking better questions of both the device and the doctor.

Sources

  1. Oura Cycle Insights Update: Faster, More Accurate Period & Ovulation Tracking
  2. Machine learning-based menstrual phase identification using wearable device data | npj Women's Health
  3. The diagnostic accuracy of wearable digital technology in detecting fertility window and menstrual cycles: a systematic review and Bayesian network meta-analysis | npj Digital Medicine
  4. Best Perimenopause App 2026: Hot Flash & HRT Trackers Tested
  5. Fertility Tracking with Wearables: Empowering the Femtech Frontier
Filed underWearables

More in Wearables