Article

Polygenic risk scores in wellness tests: a statistical model is not a diagnosis

A polygenic risk score combines effects from many genetic variants into a model estimate for a defined outcome and population. Its meaning depends on the training data, ancestry, phenotype, validation, calibration, absolute-risk conversion, and whether the score improves a real decision.

3 min read Published Source checked

Hundreds of small genetic points flow into a calibrated risk curve with several uncertainty bands
Treomark editorial illustration

A polygenic risk score is a model output, not a diagnosis or a single disease-causing mutation. It combines weighted effects from many variants for a defined outcome and reference population. To interpret one, identify the phenotype, score version, training and validation cohorts, ancestry performance, calibration, comparison group, absolute-risk method, added value beyond ordinary risk factors, and the decision the result is supposed to improve.123

A percentile alone answers only where a score falls relative to a chosen reference. It does not say that disease is inevitable, explain current symptoms, or tell a person to start or stop a treatment.

Reconstruct the score before reading the color gauge

Model fieldWhy it mattersQuestion to ask
Outcome definitionBroad labels can hide different diagnoses, ages or severitiesWhat exact event was predicted, over what time?
Variant and weight setDifferent score versions can produce different ranksWhich published model and genome build were used?
Training populationWeights reflect the data in which they were estimatedWho was represented, and how similar are they to the tested person?
ValidationPerformance in the training data can be optimisticWas the locked score tested independently and prospectively?
CalibrationA good ranking can still give wrong absolute probabilitiesDid predicted and observed risk agree in the intended population?
Clinical utilityPrediction is not automatically a better decisionDid using the score improve care or outcomes beyond existing tools?

NHGRI’s PRIMED program exists in part because scores have not performed uniformly across populations and because diverse data and methods are essential.2 Do not treat a general “multi-ancestry” label as the end of that audit.

Relative rank and absolute risk are different numbers

A score may report “top 10%,” “twofold genetic risk,” or a projected percentage. The percentile depends on the reference group. A relative risk needs a baseline rate. An absolute-risk estimate additionally needs age, sex when relevant, time horizon, competing events, and often clinical factors.

For a rare outcome, a high relative increase can still correspond to a low absolute probability. For a common outcome, a modest increase can matter more. Require the report to show the baseline and uncertainty rather than only a red category.

A polygenic score is not monogenic testing

A pathogenic variant in a high-penetrance gene and an aggregate polygenic score are different evidence types. One does not rule out or confirm the other. A negative BRCA selected-variant report, for example, cannot be replaced by a breast-cancer PRS, and a PRS does not explain a strong hereditary pattern.

The DTC BRCA guide follows selected variants and clinical confirmation. A polygenic report should state whether known high-impact variants are assessed separately and where genetic counseling enters.

Added prediction must change a real decision

Ask how the score performs compared with age, family history, examination, blood pressure, cholesterol, smoking, imaging, or another established model. Report discrimination, calibration and reclassification in the intended population—not only an area under a curve from development data.34

Then ask what changes at a threshold. If everyone receives the same general lifestyle advice, the score may add curiosity without clinical utility. If it changes screening or treatment, require a guideline, trial or validated pathway showing that the change improves the balance of benefit and harm.

Model updates can move the result

Variant weights, reference cohorts, imputation, phenotype definitions and risk calculators change. Ask whether the company freezes the result, silently recomputes it, or issues a new version. Preserve the original score and method so a later change can be explained.

Genomic information can remain identifiable or become re-identifiable when combined with other data. Review storage, research use, third-party sharing, deletion, family implications and whether results enter a medical record. Privacy consent is separate from model validity.5

The decisive question

Ask: “For this exact score in people like me, how well calibrated is the absolute risk, and what validated decision improves because I know it?” Without both parts, a sophisticated genetic model may remain a wellness metric rather than a clinical tool.

Sources

  1. National Human Genome Research Institute. Polygenic risk score. Current federal definition of polygenic risk scores and core interpretation concepts. Accessed .
  2. National Human Genome Research Institute. Polygenic Risk Methods in Diverse Populations Consortium. Federal research program addressing performance, methods and diversity limitations in polygenic scores. Accessed .
  3. National Library of Medicine. Incorporating polygenic risk scores and social determinants of health across populations: Considerations and best practices in research. Current review and considerations paper on integrating polygenic scores with social determinants of health across populations. Accessed .
  4. National Library of Medicine. Development and Validation of a Clinical Polygenic Risk Report in U.S.-Based Health Systems for 8 Cardiovascular Conditions. Current study developed scores in 245,394 All of Us participants and externally validated them in 53,306 Mass General Brigham participants; it does not endorse a consumer test. Accessed .
  5. National Human Genome Research Institute. Privacy in genomics. Federal overview of genomic identifiability, data sharing, informed consent, and context-dependent privacy protections. Accessed .
Built from the public records listed above. Spot an error? Report a correction