“Validated” is a common word in digital health and an imprecise one. A sensor can be validated for one purpose and not for another. An algorithm can be accurate in healthy adults and unreliable in a patient population.
Sponsors building a digital measure need a clearer picture: which kinds of evidence exist, what each one shows, and how much of each a particular measure needs for its role in a trial.
Why Does Validation Depend on Context of Use?
Validation depends on context of use because evidence only supports the specific way a measure will be used. Validation should always be evaluated within the intended context of use. A wearable sensor validated for healthy adults may not perform similarly in patients with neurological disorders, respiratory disease, or mobility limitations, a point made in how to select wearable sensors for a clinical trial.
The FDA-NIH Biomarkers, EndpointS, and other Tools (BEST) Resource defines context of use as “a statement that fully and clearly describes the way the medical product development tool is to be used and the regulated product development and review-related purpose of the use.”
In practice, evidence does not transfer automatically. Evidence gathered in one population, one setting or one trial role has to be reassessed when any of those change.
What Does a Context of Use Include?
A context of use for a digital measure usually pins down several things at once:
- The concept of interest the measure is meant to reflect
- The population, including age, condition and functional ability
- The setting, such as in clinic, at home or free-living
- The technology, wear location and data collection period
- The role of the measure in the trial, from exploratory to primary endpoint
Changing any one of these can change the evidence a measure needs.
What Is the V3 Framework?
The V3 framework is a three-part structure for the evidence behind sensor-based digital health technologies (DHTs): verification, analytical validation and clinical validation. It was set out in a 2020 paper in npj Digital Medicine by Goldsack and colleagues.
Verification
Verification evaluates the sensor itself. The V3 paper describes it as a systematic evaluation in which sample-level sensor outputs are assessed, typically by the hardware manufacturer. The question is whether the device captures data correctly.
Analytical Validation
Analytical validation evaluates the algorithm that turns sensor data into a measure. The V3 paper places it where engineering and clinical expertise meet, focused on the data processing that converts sample-level measurements into physiological metrics.
Analytical validity demonstrates that a digital biomarker accurately measures its intended target. Sponsors typically look at measurement accuracy, precision, repeatability, and algorithm performance, as described in how sponsors build FDA-ready digital biomarkers.
Clinical Validation
Clinical validation asks whether the measure means something for patients. Clinical validation establishes that the wearable-derived measurements are clinically meaningful and associated with relevant patient outcomes, in a defined context of use.
The BEST Resource defines clinical validation more generally as “a process to establish that the test, tool, or instrument acceptably identifies, measures, or predicts the concept of interest.” That definition ties biomarker validation, including for digital measures, to the concept of interest.
What Does V3+ Add?
V3+ adds a fourth component, usability validation, to the original three. The addition was proposed in a January 2025 V3+ paper in npj Digital Medicine by Bakker and colleagues.
Usability validation asks whether the people using a technology can use it as intended. That includes participants wearing a device at home, care partners helping them, and investigators and research staff setting it up. The authors present usability validation as a way to make digital measurement tools user-centric.
Usability matters for data quality. The V3+ authors note that poor usability can lead to sub-optimal adherence and, in turn, excessive missing data. A device that is uncomfortable, confusing or hard to charge can produce gaps, however accurate its algorithm is. How those gaps show up in wear time and valid days is covered in wear time, valid days and analyzable data.
What Evidence Does Each Component Typically Involve?
Each component answers a different question. In the V3 paper, each of the original three components is usually generated by a different party.
| Component | Question it answers | Typical evidence | Who is involved |
|---|---|---|---|
| Verification | Does the sensor capture data correctly? | Computational (in silico) and bench (in vitro) testing of sample-level sensor outputs against a reference | The hardware manufacturer |
| Analytical validation | Does the algorithm turn sensor data into an accurate measure? | Testing in people (in vivo), comparing algorithm outputs against an appropriate reference standard | The entity that created the algorithm, either the technology vendor or the clinical trial sponsor |
| Clinical validation | Does the measure reflect a meaningful state in the defined context of use? | Studies in cohorts of patients with and without the phenotype of interest | Sponsors of new medical products or clinical researchers |
| Usability validation (V3+) | Can intended users use the technology as intended? | A use specification, use-related risk analysis, and formative and summative evaluations with each user group | Studied with every user group named in the use specification |
Because different parties generate different parts of the evidence, sponsors can often draw on verification data from a manufacturer and analytical validation data from earlier studies. The question is whether that evidence covers the population and setting of the new trial.
What Does FDA’s 2026 Paper on Digitally Derived Measures Emphasize?
On August 20, 2026, FDA published “Key Considerations for the Development and Use of Digitally Derived Measures for Clinical Investigations.” It was authored across the Center for Biologics Evaluation and Research, the Center for Drug Evaluation and Research, the Center for Devices and Radiological Health and the Oncology Center of Excellence.
It is a paper, not a new guidance. In FDA’s words, it “highlights key considerations drawn from existing U.S. Food and Drug Administration (FDA) guidances.”
What a Digitally Derived Measure Is
The paper defines digitally derived measures (DDMs) as “measures derived from data collected using DHTs.” It notes that “DDMs may be used as clinical outcome assessments (COAs), biomarkers, or as part of multicomponent endpoints derived from multimodal data.” The distinctions between those roles are set out in biomarkers vs clinical outcome assessments vs endpoints.
Verification and Validation
The paper states that “a DHT used to generate a DDM should be verified and validated to be considered fit-for-purpose.” It describes validation as having two major parts: analytical validation and clinical validation.
Usability
It says that “evidence for validation should demonstrate that users understand and can follow the instructions for use, typically assessed through human factors or usability studies.”
Sources of Error and Missing Data
It says it is important to “identify potential sources of error and factors that might negatively impact the validity of the DDM,” and notes that “the impact of missing data and variable-quality data may be considered during analytical validation.”
Software Updates During a Trial
The paper addresses technology changes mid-study. If a DHT or associated technology is updated during a clinical investigation, it says sponsors and other relevant parties should confirm that the DHT remains fit-for-purpose and that the updates do not affect the DDMs.
Evidence Matched to Intended Use
It describes a risk-based approach, “tailoring evidence generation commensurate with the intended use of the DDM,” and says the scope of validation “is determined based on the current state of evidence and the intended use of the DDM, e.g. context of use, role in the clinical investigation.”
The paper gives an example of how role changes the expectation: “a DDM serving as a primary endpoint in a pivotal study is expected to have prospective validation with pre-specified performance thresholds in the intended target population, whereas a DDM used as an exploratory endpoint may rely on retrospective or bridging evidence.”
Patient Involvement
It says it is important to involve patients, caregivers and clinicians “in the development process to determine meaningfulness and clinical relevance of the DDM.” VivoSense has written about how that expectation developed in patient-centricity in digital measure development.
Matching Evidence to the Measure’s Role in the Trial
The same technology can need very different evidence depending on how a trial uses it. A measure supporting an exploratory objective is in a different position from one proposed as a key endpoint. FDA’s paper frames this as evidence commensurate with intended use and role in the investigation.
Validation also sits inside a longer development path, often described in five stages: Identify a Meaningful Clinical Concept, Select Appropriate Technologies, Develop Digital Biomarkers, Establish Validation Evidence, Build Digital Endpoints. The stages are set out in digital biomarkers vs digital endpoints. Evidence planning belongs with the concept and technology decisions, not after them. More on the measurement itself is in what digital biomarkers are.
For sponsors, that means scoping validation work early and in writing. Questions to settle before the protocol is final include:
- What is the concept of interest, and why does it matter to patients?
- What role will the measure play in the trial?
- In which population, and in what setting, will it be used?
- What verification evidence exists for the device, and what analytical validation evidence exists for that population?
- What clinical validation evidence exists, and what still needs to be generated?
- How will usability be assessed for participants, caregivers and sites?
- How will software or firmware changes be handled during the study?
A Hypothetical Example: One Measure, Two Roles
The following example is hypothetical. A sponsor plans a Phase 2 trial in adults whose condition limits walking. The concept of interest is how much participants walk in daily life. The candidate measure is daily walking time, derived from a wrist-worn accelerometer worn at home for two weeks at several points in the study.
If the measure supports an exploratory objective, the sponsor might assemble existing verification data from the manufacturer, published analytical validation of the walking algorithm, and any bridging evidence that the algorithm performs in people who walk slowly or use walking aids. A small usability assessment with the target population could show whether participants can wear and charge the device as instructed.
If the sponsor later proposes the same measure as the primary endpoint of a confirmatory Phase 3 trial, the context of use has changed. The primary endpoint example in FDA’s paper calls for prospective validation with pre-specified performance thresholds in the intended target population. The sponsor would plan new analytical and clinical validation work in that population, rather than rely on the exploratory evidence alone.
The device, algorithm and population are the same in both scenarios. The evidence plan is not.
An Example of Analytical Validation Work
In February 2025, VivoSense published an analytical validation of wrist-worn accelerometer-based step count methods during structured and free-living activities, describing “a validated step-counting method optimized for real-world performance.” That work addresses analytical validity for step counting in the studied conditions. Use in another population or context of use would need its own assessment. The step count analytical validation summary describes the study.
Common Mistakes
Assuming Validation Transfers Across Populations
Evidence from healthy adults does not automatically apply to patients with different movement, sleep or breathing patterns.
Presenting Analytical Evidence as Clinical Evidence
An accurate algorithm shows the measure is computed correctly. It does not show the measure matters to patients.
Treating “Fit for Purpose” as a Synonym for “Validated”
Fit for purpose depends on the purpose. The same measure can be fit for one role and not for another.
Carrying Exploratory Evidence Into a Primary Endpoint
Evidence that supported an exploratory measure may not match what FDA’s paper describes for a primary endpoint in a later-stage confirmatory trial. The evidence plan needs revisiting when the role changes.
Leaving Usability Until the Study Starts
Usability problems found after enrollment can show up as missing data. They are easier to address before the protocol is locked.
Ignoring Software Changes
Operating system and firmware updates can change how a device behaves. Plans should say how updates will be assessed during the trial.
Digital Measurement With VivoSense
VivoSense is a wearable sensor contract research organization (CRO) that works alongside the sponsor’s trial team and CRO on the digital measurement workstream. It helps sponsors choose the right device based on the disease state and patient population, and decide which measures to capture. It ships devices and trains sites, monitors real-time wear compliance, and delivers formatted regulatory-ready data packages for the study team. The service areas are described on the VivoSense solutions page.
VivoSense was founded in 2010.
Frequently Asked Questions
What is the V3 framework?
A framework published in npj Digital Medicine in 2020 that describes three components of evidence for sensor-based digital health technologies: verification, analytical validation and clinical validation.
What is the difference between analytical and clinical validation?
Analytical validation shows a measure is computed accurately and reliably from sensor data. Clinical validation shows the measure is clinically meaningful and associated with relevant patient outcomes in a defined context of use.
What is usability validation?
The fourth component added in V3+. It assesses whether intended users, such as participants, care partners and research staff, can use a digital health technology as intended.
What is context of use?
The BEST Resource defines it as a statement that fully and clearly describes how a medical product development tool is to be used and the regulatory purpose of that use. For a digital measure, it covers the concept of interest, population, setting, technology and role in the trial.
What is biomarker validation?
Biomarker validation is the process of establishing that a biomarker performs acceptably for its intended use. For digital biomarkers, that usually means verification of the sensor, analytical validation of the algorithm, and clinical validation against the concept of interest, all within a stated context of use.
Who performs each part of V3 validation?
The V3 paper describes verification as typically done by the hardware manufacturer, analytical validation by the entity that created the algorithm, either the vendor or the clinical trial sponsor, and clinical validation mainly by sponsors of new medical products or clinical researchers.
Is FDA’s August 2026 paper on digitally derived measures a guidance?
No. FDA describes it as a paper that highlights key considerations drawn from existing FDA guidances.
What does fit for purpose mean for a digital health technology?
FDA’s December 2023 guidance, announced in the Federal Register notice for the DHT guidance, describes it as the level of validation being sufficient to support the technology’s use, including the interpretability of its data in the clinical investigation.
Does validation in one study apply to another?
Not automatically. Evidence is tied to a context of use, including the population, setting, and role of the measure in the trial.
Let’s Talk
Schedule a consultation to explore how to design and use validated digital measures in your clinical trials.
