Methodology
How the HealthScore is calculated
This page is the whole model: what each domain is worth and why, what happens to the questions you skip, what can put a ceiling on a score, where the sources come from, and the times we have had to correct ourselves. Nothing here is summarized from somewhere more detailed — this is the detail.
The shape of it
28 self-reported questions produce a score from 0 to 100 across 14 domains. No blood draws, no wearables, nothing measured by us — every answer is something you can honestly give from memory, which is the ceiling on how precise any of this can be and the reason the page you are reading exists.
The domains are combined as a weighted geometric mean, not a weighted average. An average assumes the domains substitute for one another — that a gain in diet exactly offsets an equal loss in smoking, at every level. That is false about health, and it produced absurd results: on an earlier version, someone flawless on every other domain who smoked a pack a day scored 87 out of 100. A geometric mean means weak areas cannot be fully bought off by strong ones. A single domain is floored at 10 before the multiplication, so one zero cannot annihilate everything else — which would be the opposite mistake.
Where the domains came from
The American Heart Association’s Life’s Essential 8 names eight measures of cardiovascular health: diet, physical activity, nicotine, sleep, weight, blood pressure, cholesterol and blood sugar. All eight are scored here. So are the 6the construct does not cover — alcohol, stress, social connection, hearing and mental engagement, the air where you live, and whether anyone has actually looked.
The questions are ours, the weighting is ours and the ceilings are ours: this is not the AHA’s scoring and does not claim to be. The starting list is theirs, and the paper is public. American Heart Association / Circulation: Life's Essential 8
Why a questionnaire rather than a wearable. A sensor reaches two of these 14domains well: movement and sleep. Nothing on a wrist knows what a clinician last told you about your blood pressure, how many days a week you drink, whether you follow conversation in a noisy room, what the air is like where you sleep, or whether anyone has screened you for anything. Those are questions, not readings — which is also the limit of the whole model, since a self-report is only as good as the memory behind it.
The weights
These are ours, and they are a judgment call. These weights are a considered editorial ranking of evidence strength for premature all-cause mortality. They are NOT a published weighting lifted from any one paper, and no paper endorses them. The construct that inspired the score weights its own domains equally; we deliberately do not, because equal weighting is what let an early version treat how often someone manages stress as the peer of smoking a pack a day. Reasonable people would weight these differently, and the argument for each number is printed beside it so you can disagree with a specific one rather than with the whole thing.
| Domain | Share | Why this weight |
|---|---|---|
| Physical activity | 13% | One of the two largest modifiable behavioral contributors to premature death, and the one with the clearest dose-response evidence across the whole range from sedentary upward. |
| Nicotine | 13% | The other of the two largest modifiable behavioral contributors, and the single answer in the model most likely to trip a hard ceiling on its own. |
| Diet | 11% | A large contributor, but spread across many components, and each single self-reported item is a noisier read than “do you smoke.” Weighted below the two behaviors above it for that reason. |
| Blood pressure | 9% | The highest-weighted clinical marker here: strong, graded and treatable, and the one most people can actually recall being told. |
| Sleep | 8% | A large and independent contributor, and partly upstream of the clinical markers below — weighting it at blood pressure’s level would count the same causal pathway twice. |
| Weight | 8% | Large, and partly upstream of blood pressure, blood sugar and cholesterol — so it is weighted alongside sleep rather than alongside the markers it feeds. |
| Blood sugar | 6% | A large effect, and unusually actionable at the prediabetes stage, which is the argument for weighting it above cholesterol. |
| Social connection | 6% | An independent association comparable in magnitude to several classic behavioral risk factors. Omitting it from a product that talks about lifespan would be a real gap, not a conservative choice. |
| Alcohol | 5% | Real and graded, and kept as its own domain rather than merged with nicotine: a never-smoker who binge-drinks and a daily smoker who never drinks are not the same risk, and averaging them into one number says they are. |
| Environment & safety | 5% | Air quality and everyday safety had no home in the other thirteen domains, and both carry real population-level burden. Weighted below the behavioral domains for the same reason as cognitive: the underlying risk is real, our measurement of it is coarse. |
| Cholesterol | 4% | Real and treatable, but the most commonly unknown of the three markers, the one most mediated by diet and weight — both already weighted above — and the one whose “has anyone checked” half is now scored by preventive care instead. |
| Preventive care | 4% | Deliberately small, for three reasons. The evidence for screening is strongest for deaths from the disease screened for and much thinner for lifespan overall, which is what every other weight here is ranked against. The measurement is two coarse self-reports. And part of the benefit is already counted, because getting screened is how a person comes to answer the three marker questions at all. |
| Stress | 4% | A genuine association, and the weakest measurement in the model — one subjective self-report rather than a behavior count or a recalled clinical fact. A weight must not claim more precision than its input. |
| Hearing & mental engagement | 4% | Hearing and mental engagement together account for a meaningful share of the Lancet Commission’s modifiable dementia risk, and hearing is unusually correctable — which is what tips it from interesting to actionable. Weighted below the behavioral domains because our measurement of it is coarser self-report. |
Questions you skip, and the coverage ceiling
A domain you did not answer is excluded, and the remaining weights are re-spread across what is left. It is never filled in with an invented middle value. An early version scored “I don’t know” as 50 out of 100, which is a data point nobody supplied.
Excluding-and-re-spreading was the right instinct with the wrong consequence: re-spreading is a rewardwhenever the missing domain would have scored below your average. Someone answering “I don’t know” to blood pressure, cholesterol and blood sugar could score 100 — a perfect score certifying three things nobody had measured — while the same person who knows their blood pressure is high scored in the seventies. Knowing less scored better.
So there is a coverage ceiling: your score may not exceed the share of the model that produced a real answer. Answer 81% of it and the most we will say is 81. It binds only when the score would otherwise claim more than your answers support, and it lifts by exactly the action anyone would recommend anyway — go and find out.
Ceilings
A few single answers carry enough absolute risk that no combination of good habits should place someone in a reassuring band. Each sets a hard maximum on the score:
| Answer | Score cannot exceed |
|---|---|
| Daily smoking | 70 |
| Unmanaged diabetes | 70 |
| Untreated high blood pressure | 72 |
| BMI of 40 or above | 74 |
| Diagnosed but untreated sleep apnea | 80 |
These ceilings are editorial. They encode a product judgment about what a score is allowed to imply. They are not derived from a published risk equation, and no paper sets them. They are kept few, high-threshold, and in one table so the judgment is inspectable rather than scattered through the code.
They compound rather than the lowest one simply winning, because taking the lowest would say a third catastrophic answer is free once you already have two. Carrying the two lowest ceilings at once caps a score at 49; the three lowest cap it at 35.
Every ceiling targets an untreated state. Treated high blood pressure, managed diabetes and treated sleep apnea set no ceiling at all: someone doing the right thing about a real diagnosis must never be scored as though they had done nothing.
The bands, and one thing they are not
| Band | Range |
|---|---|
| High range | 80–100 |
| Middle range | 50–79 |
| Low range | 0–49 |
These cut points are the American Heart Association’s published recommendation for categorizing a 0–100 cardiovascular-health score. The weighting inside them is ours.
The mortality study we quote beside your score used different groups: its top group starts at 75, not 80, and its middle group runs from 50 to 74. Those cut points were derived by that study’s own authors, who describe the boundary as somewhat arbitrary. So the band you are labeled with follows the AHA’s scale and the mortality figure follows the study’s, and between 75 and 79 those two things differ. We say so rather than papering over it. Until 2026-08-28 we did not, and that is the first entry in the log below.
What moves the number, and what does not
One number, and the thing that keeps it honest
- HealthScore
- Where you stand. It is a measurement, so it changes when you measure again — retake the quiz, or re-answer a single domain. Checking a habit off never moves it, by design.
- What you do daily
- Coach is where the work happens, and it is deliberately not a second score. There is no count to build and no streak to break. What logging does is sharpen the next measurement — it gets built out of the days you actually recorded rather than the ones you can remember.
Sources
The habit library cites 110 sources, and all 110 have now been reviewed by a person — meaning somebody opened the primary source and checked that each figure quoted is verbatim from it and that the copy claims no more than the source found.
17of those sit on publishers that block automated readers — the CDC, the American Heart Association and the Dietary Guidelines all return an error to anything that is not a browser. For those the review is a person’s reading and cannot be re-checked from the codebase: the automated guards that verify every other figure against its source are blind to them. If one of those numbers is ever wrong, a machine will not be what catches it. We would rather say so than let the word “reviewed” carry a weight it does not have.
The library holds 85 entries — 54 recurring habits, 26 one-off moves and 5 reads that are never tracked. Of those 85, 85 have been reviewed. The number that actually governs what the product will claim is stricter: the entry itself and every source it rests on must both have been checked, and today all 85 of the reviewed entries clear that bar. A card states how strong its evidence is only when it does. Everything else says “strength not yet reviewed” and still shows and links its sources, so you can check them yourself.
No clinician has reviewed this product, and we do not claim one has. The review referred to above is editorial: whether each figure is verbatim from the source cited, and whether the copy claims more than the source found.
Corrections
A methodology nobody has ever had to fix is a methodology nobody has ever checked. These are the substantive corrections, most recent first.
- 2026-08-28 — The mortality statistic beside the score was being selected on our display bands rather than on the cited study’s own groups.
The bands are the American Heart Association’s recommended categories for a 0-100 cardiovascular-health score, and its high band starts at 80. The cohort we quote for mortality drew its own groups, and its high group starts at 75. Anyone scoring between those two numbers was shown the middle group’s figure when the study would have counted them in its top group. The statistic now follows the study; the band a person is labeled with did not move, because correcting a citation should not change what anyone is told about themselves. - 2026-08-28 — The score reveal called hearing loss the largest single modifiable dementia risk factor. It is tied, not first.
The Lancet Commission’s own table gives hearing loss a weighted attributable fraction of 7.03% and high LDL cholesterol 6.88%, and rounds both to 7%. A gap of 0.15 of a percentage point on a modeled figure is not a rank order. Our own source record had said so in writing; three habit cards and the score page had not followed it. - 2026-08-28 — Four population claims on the free score reveal named no source at all.
The rule that every user-facing claim traces to a source was enforced by a test that walked the habit catalog, and this copy lives outside it. Each of the four has now been checked against its primary source and names it, and the check was extended to cover the copy that is written in code rather than in content. - 2026-08-28 — Cards showed an evidence-strength label whether or not that strength had been reviewed by a person.
Every habit and every source carries a review status, and nothing was reading it. Content still awaiting review now says so instead of asserting a strength nobody has checked. - 2026-08-27 — Numbers appeared in habit copy that no source cited by that habit reports.
An audit of the whole catalog found figures attributed to records that do not contain them. Every number in a habit’s copy must now appear verbatim in one of that habit’s own sources, or be declared in writing as a design choice of ours with a reason attached. - 2026-08-25 — A perfect answer set with three unknowns scored 100.
Answering “I don’t know” to blood pressure, cholesterol and blood sugar excluded those domains and re-spread their weight over the rest — which is a reward, not a neutral act, whenever the missing domain would have scored below average. A score may no longer exceed the share of the model that produced a real answer. - 2026-08-25 — Good habits could buy off lethal risk factors.
The composite was a weighted average, which assumes a gain anywhere offsets a loss anywhere. Someone flawless on every other domain who smoked a pack a day scored 87. The composite is now a weighted geometric mean, and a few individually catastrophic answers set a hard ceiling on top of it.
What this is not
This is a wellness score, not a medical assessment. It does not diagnose, does not treat, and gives no individual risk estimate. Where it quotes a study it is describing a pattern across a population, never a forecast for you. If something here prompts a question about your own health, the person to ask is a clinician.