The Oura lawsuit is asking the wrong question
On 20 August a proposed class action was filed against Oura in the Northern District of California. It challenges the “95% Sleep Staging Accuracy Compared to clinical sleep lab” figure that has appeared in the company's marketing. The plaintiff's firm puts the real number nearer a coin flip; Oura is contesting it and has published a defence of its science. I am not going to argue either side's characterisation. I want to talk about the number nobody is litigating.
On this page
What the filing actually says
The complaint — Surber v. Oura, brought by the Clarkson Law Firm on behalf of a California buyer — sets two of Oura's own numbers against each other. The 95% figure sits on a product page. A company blog post from 2022 reports 79% agreement with polysomnography for classifying the four stages of sleep. Oura's response points to peer-reviewed validation work, an algorithm developed against more than a thousand recorded nights, and separate figures for the easier task of telling sleep from wake, where it reports roughly 92%. Those are all real numbers measuring different things, which is most of the argument.
The accuracy fight has a ceiling nobody mentions
A sleep lab measures brain waves, eye movement and muscle activity. A ring measures heart rate, temperature, movement and breathing, and infers the rest. That gap is real and worth arguing about.
But the lab is not a perfect ruler either. Put trained human scorers in front of the same overnight recording and ask them to label every thirty-second window, and they do not fully agree. The American Academy of Sleep Medicine runs an inter-scorer reliability programme, and the average agreement between scorers comes out around 83%. It is worst on the lightest stage, where humans agree barely six times in ten.
So the honest version of the accuracy question is not “is the ring lying”. It is “how close can anything get to a target that moves depending on who is scoring it”. No device can be more consistent than the labels it was trained on. That is a genuine limitation and it deserves saying out loud — and it is, to my mind, the less interesting problem.
The number that cannot be wrong
Sleep staging can at least be checked. There is a lab, a protocol, published papers, and now a lawsuit. You can disagree with the figure because there is a figure to disagree with.
Now look at the sleep score. Or readiness. A single number out of 100 that decides how your morning feels.
Nobody publishes how it is weighted — not Oura, not Whoop, not Apple, not any consumer ring I know of. You do not know what share is duration, what share is timing, what share is resting heart rate, what share is temperature deviation, or how any of it is scaled. So when the score says 62 and you feel fine, there is no way to work out which input dragged it down, and no way to establish that it is wrong.
A claim that cannot be contradicted is not a strong claim. It is one that has been placed out of reach. And the practical effect is worse than inaccuracy: people stop checking. If the number cannot be argued with, the number wins, and you slowly stop trusting how you actually feel. Sleep clinicians have a name for where that ends up — orthosomnia. Unfalsifiable is a worse failure mode than wrong. Wrong gets corrected; unfalsifiable just sits there.
What the ring is actually for
I build a desktop app that reads Oura data, so treat me as biased. This is what I have settled on: the ring is a poor instrument for judging one night and a decent instrument for dating a change. Ask it “did I sleep well on Tuesday” and it will mislead you. Ask it “did something shift in the last three weeks, and roughly when” and it earns its place.
That has three consequences. You are not buying a ring, you are buying a baseline — the first couple of weeks are the price of admission, which is also why switching devices costs more than the sticker price suggests. Numbers from different devices agree on direction and never on level, so comparing your score to a friend's watch is wasted effort while comparing this week to your own last month is not. And a missing day has to stay missing: if the ring was on the charger, that day is null, not zero. Averaging a gap in as a bad night is how a tool starts inventing trends.
What I would rather see than a better accuracy number
Publish the weighting. Not the model, not the intellectual property — just what the composite is made of and roughly how much each part counts. That single change would do more for trust than another percentage point of staging agreement, because it would make the score arguable again.
Until then, the reasonable thing is to look past the score at what it was built from, and to keep your own history somewhere it cannot be taken away from you. Vitra reads your Oura data on your own Mac or PC, shows each measurement against your own baseline rather than a verdict, and states what every number is made of. Everything is computed locally; your health data never leaves your machine. Vitra is independent and not affiliated with Oura.
Frequently asked questions
Pedro Thomaz builds Vitra, a desktop app that reads Oura data against your own baseline instead of a population average. He has worn a ring daily for years and reads these same numbers every morning — which is where most of what is written here comes from. Vitra is not a medical device and nothing on this blog is medical advice.
Local AI on your Mac or PC. One-time purchase, 7-day trial, no Vitra subscription.
Download Vitra →