Health evidence is not one ladder where every study occupies an obvious rung. A systematic review can inherit weak studies. A randomized trial can study the wrong population or only a short period. An official product page can accurately state battery life while being the wrong source for accuracy or health benefit.

Our method separates four questions: What is being claimed? What source could reasonably support it? How directly does that source match the current product and user? What important uncertainty remains?

Four labels, each with a reason.

LEVEL AStrong

Consistent high-quality evidence or authoritative guidance directly supports the scoped claim.

LEVEL BModerate

Useful evidence exists, but design, population, duration, or consistency limits certainty.

LEVEL CEmerging

Promising early studies, indirect evidence, or meaningful unresolved disagreement.

LEVEL DUnknown

No adequate independent evidence was identified for the specific model, metric, or claim.

These are editorial communication labels, not a substitute for GRADE, risk-of-bias tools, or statistical certainty. Every label must be accompanied by the reason it was assigned.

Our source hierarchy depends on the claim.

Health effectSTART WITH

Systematic reviews, meta-analyses, guidelines, and well-designed primary studies. Commercial trend reports cannot establish a health effect.

Measurement accuracySTART WITH

Independent validation against an appropriate reference such as ECG or polysomnography, ideally for the exact device generation and population.

Regulatory statusSTART WITH

The regulator database, clearance summary, official labeling, and current regional availability—not a launch headline.

SpecificationSTART WITH

The manufacturer’s current technical and support pages, time-stamped because price, battery claims, features, and subscriptions can change.

ExperienceSTART WITH

Hands-on testing only when we actually performed it. Desk research is always labeled and never rewritten as personal use.

Product reviews have stricter disclosure rules.

  • No affiliate links at launch. Direct product links are included for verification, not commission.
  • No manufacturer payment or supplied unit. If that ever changes, the disclosure appears at the top of the page.
  • No invented testing. Current reviews are evidence-led desk reviews. We do not score comfort, durability, app usability, or real battery performance without direct use.
  • No numeric star score. A precise rating would falsely combine specifications, evidence quality, personal experience, value, and medical relevance.
  • No class-wide inheritance without a warning. Evidence for Oura Gen3, WHOOP 4.0, or an older algorithm does not automatically validate the newest hardware.
  • No “best” without a defined user. Value depends on purpose, region, fit, subscription tolerance, phone ecosystem, and need for training or health features.
Commercial independence

HealthFit currently receives no commission, payment, free product, or editorial approval from the products covered. Product names and marks belong to their owners. Inclusion does not imply partnership or endorsement.

Corrections are part of the product.

Specifications and evidence dates appear on the page. A material change—new independent validation, regulator action, corrected pricing, recall, or feature retirement—updates the modified date and the affected conclusion. Minor copy edits do not silently become a “new” evidence review.

Readers and manufacturers may send a correction with a primary source to hello@healthfit.ai. A correction request does not guarantee the requested wording; it guarantees review against the source.

Health boundary

Research grading describes evidence, not personal suitability. Product pages cannot diagnose, clear someone for exercise, adjust medication, or replace a clinician who can evaluate symptoms and context.