
January 2026
In the FDA context, human factors validation testing serves as a component of design validation in accordance with the FDA design control requirements (formerly 21 CFR Part 820.30, incorporated into ISO 13485:2016, clause 7.3, by the Quality Management System Regulation, QMSR, since February 2026), with the specific objective of demonstrating that the medical device can be used safely and effectively by the intended users in the intended use environments. The FDA has outlined its expectations for this in the guidance document “Applying Human Factors and Usability Engineering to Medical Devices.” In May 2026, the FDA additionally finalized its guidance on the content of human factors information in marketing submissions. A central foundation is the risk‑based identification and categorization of critical tasks based on the severity of potential harm, all of which must be fully covered in the test. The FDA typically expects usability tests to be conducted as simulated‑use testing with the final user interface, realistic use conditions, representative user populations (generally at least 15 participants per relevant user group), and a test execution that enables independent and natural use without moderator influence. Data collection should combine observational data, knowledge task data where applicable, and post‑use interviews. The analysis is qualitative and requires a root cause analysis for use errors, close calls, and use difficulties, as well as a robust justification of residual risk. This article explains in detail how to proceed with this process.
In the context of FDA approval, the FDA Guidances emphasize the risk‑based nature of human factors and usability engineering and focus primarily on ensuring that the user interface is designed such that use errors with potentially harmful consequences are eliminated or reduced as much as possible. Human factors validation testing is conducted at the end of development to demonstrate that the final user interface supports safe and effective use under expected conditions of use and that the implemented risk control measures are effective.

Within the context of FDA approval, a central structural principle is the definition of critical tasks. These are tasks for which incorrect performance or failure to perform the task could result in serious harm to the patient or user, including compromised medical care. The identification and justification of these critical tasks is relevant not only for the test design but also shapes the evaluation approach and the argumentation presented in the report.
With respect to critical tasks, the following principles apply:
The FDA recommends identifying critical tasks through a combination of analytical and empirical approaches:
An important aspect of the FDA’s reasoning is that the probability of occurrence of use errors is difficult to quantify reliably in actual use; therefore, the severity of potential harm is the primary factor used to prioritize tasks.
To plan the study in a methodologically correct manner, the development of a test protocol is required. The FDA typically requires that human factors validation be conducted as a usability test in the form of simulated‑use testing and specifies clear minimum requirements for the study design. The validation test is not an exploratory format; rather, it is a purposefully constructed demonstration intended to show that the final user interface can be used safely and effectively by the intended users in the intended use environments, particularly with respect to the critical tasks.
At its core, the FDA requires that human factors validation testing ensures that:
the test participants represent the intended users,
all critical tasks are actually performed during the test,
the user interface reflects the final design,
and the test conditions are sufficiently realistic so that the findings can be generalized to actual use.
These requirements must not merely be stated in the protocol but defined operationally. This means that it must be documented in a traceable manner which user characteristics are considered representative (and why), how full coverage of all critical tasks is ensured, which design version is considered “final” (including labeling and the instructions for use (IFU)), and which environmental and contextual factors are relevant for achieving realistic test conditions.
Simulated‑use conditions must be realistic enough for the results to be generalizable to actual use. The FDA emphasizes that relevant environmental factors must be incorporated into the simulation based on risk, insofar as they could affect user interaction with the device. The guidance cites examples such as dim lighting, distractions, and multitasking.
For protocol planning, this means that realism is not achieved through superficial “scenery,” but through the targeted representation of those conditions that are plausible contributing factors to use errors. The protocol should therefore explicitly define:
which environmental parameters are relevant for realism (e.g., lighting, noise, spatial constraints, device setup, typical distractions),
which of these parameters will be controlled, which will be varied, and which will intentionally not be addressed in the test,
and how these decisions are justified with respect to the use‑related risk analysis and the critical tasks.
At the same time, the simulation must remain methodologically stable: realism must not introduce uncontrollable confounding variables. Therefore, the FDA requirement focuses on “sufficient realism,” not on a full replication of a clinical environment. What matters is the logical justification that the essential interaction conditions have been adequately represented.
For validation, participants must use the device independently and naturally. The FDA explicitly states that think‑aloud techniques are not acceptable in validation studies, as verbalization alters behavior and therefore does not represent real‑use conditions. The protocol must therefore be structured so that relevant data are derived primarily from observation. Interactions should not be influenced by prompts, questions, or “assistance” from the study environment. If questions are necessary (e.g., for root‑cause clarification), they should generally be shifted to a post‑use interview to avoid distorting actual task performance.
The FDA considers labeling—including the instructions for use (IFU), device labels, packaging, quick‑start guides, and comparable informational components—to be part of the user interface. Accordingly, the final versions of these materials must be used in the validation test. This is methodologically relevant because many critical tasks depend not only on interaction design but also on the clarity, accessibility, and applicability of the information provided. Human factors validation therefore implicitly assesses whether users, under realistic use conditions, can make correct decisions and perform critical tasks accurately using the provided labeling. At the same time, the FDA makes clear that labeling or training must not serve as a default compensation for design‑related use errors on critical tasks. If use errors on critical tasks occur during validation, relying solely on labeling or training modifications—without additional evidence demonstrating their effectiveness—is not acceptable.
Although the guidance does not prescribe a rigid structure, we recommend that the following aspects be described in the test protocol:
test objectives with explicit reference to safe and effective use and the critical tasks,
description of the final user interface, including the labeling set,
use scenarios and task sets from which the test cases are derived to ensure full coverage of the critical tasks,
success criteria for each task, clearly defining what constitutes successful task performance,
training requirements consistent with real‑use conditions,
test environment and realism parameters,
data collection methods (primarily observational, supplemented by interviews and knowledge tasks),
rules for moderator behavior and handling of deviations.
A systematically developed test protocol demonstrates that the study was not merely “performed,” but methodologically constructed to truly meet the FDA’s requirements for human factors validation as a risk‑based demonstration.
In the context of FDA approval, the FDA recommends a practical minimum of 15 participants for human factors validation tests; if multiple relevant user groups are involved, at least 15 participants per user group must be included. User groups differ particularly when characteristics or task responsibilities vary in ways that are likely to affect interaction with the device (e.g., lay users versus healthcare professionals, pediatric versus adult users). When selecting test participants, it is essential that they are U.S. residents.
Participants should represent the relevant range within the user group, particularly with regard to characteristics that may influence interactions (e.g., sensory, physical, or cognitive limitations, literacy/language skills, experience). For certain medical conditions that may cause functional limitations, corresponding representativeness considerations must be explicitly taken into account. Employees of the manufacturer should not serve as participants in human factors validation tests (exceptions apply only in rare, highly specialized cases).
The training approach applied in the testing must reflect real‑use conditions: if actual users receive minimal or no training, test participants must not receive more extensive training. Training immediately before the test is problematic; the FDA highlights training decay and expects a time interval (depending on the device, up to a day) to realistically represent retention.
For successful FDA approval, the FDA expects that data collection in human factors validation testing is based primarily on observable task performance, supplemented by information that cannot be reliably inferred from behavior (e.g., comprehension, interpretation, decision‑making knowledge). Accordingly, the data collection approach in the protocol should clearly define which data sources will be used (observational data, knowledge task data, interviews).
According to the FDA, successful performance of critical tasks must be assessed primarily through observation. Observational data therefore constitute the key source of evidence for determining whether the user interface supports safe and effective use under the intended conditions.
From a methodological perspective, this means that the test record (observation log) must precisely document:
What constitutes task completion (success criteria for each task, including required subtasks, correct system states, and correct decisions).
Which deviations are to be classified as use errors (e.g., incorrect selection, incorrect setting, omission of a safety‑critical action, improper response to feedback/alarms).
Which observable indicators must be captured as close calls or use difficulties (e.g., search behavior, repeated attempts, inconsistent strategies, hesitation at decision points).
In addition, use difficulties must be captured—not as incidental notes, but as systematic indicators, particularly when they occur repeatedly or cluster around safety‑critical points of interaction. The FDA identifies the following typical observable markers, among others:
Repeated attempts
Hesitation
Confusion
Search behavior
Part of the critical use safety depends not only on physical operation but also on knowledge and comprehension, for example with respect to:
contraindications,
warnings,
environmental conditions/vulnerabilities,
maintenance or safety‑critical preparatory steps.
Such aspects cannot often be reliably inferred from observable behavior alone—particularly when the test setup cannot practically represent every rare contextual condition. For this purpose, the FDA recommends knowledge tasks, i.e., knowledge and comprehension questions designed to assess whether participants can correctly interpret relevant information and make appropriate decisions.
Key methodological requirements include:
In the FDA approach, post‑use interviews are a complementary data source. They are intended to deepen the observational data but must not replace it. Interviews serve in particular to:
systematically explore the root causes of observed use errors or problems,
capture subjectively experienced use difficulties that may not have been clearly observable,
and understand the interaction logic from the user’s perspective (e.g., mental models, expectations, interpretation of information).
The FDA recommends conducting interviews after the completion of the use scenarios to avoid influencing task performance. Methodologically, it is essential that the exploration of each observed use error or problem is conducted systematically—using open, neutrally phrased questions (e.g., “What did you expect at this point?”, “How did you interpret this information?”, “What led you to this decision?”) rather than confirmatory or leading formulations.
In addition, interviews should explicitly provide space for:
subjectively reported difficulties (even when no errors were observed),
perceived ambiguities in terminology, symbols, or feedback,
and the reasons for recovery strategies used in close calls.
In the context of FDA approval, the FDA positions human factors validation testing as primarily a qualitative approach: the goal is not to estimate statistical frequencies, but to identify design‑related root causes of use errors or problems and to derive effective risk controls.
The analysis aggregates observational data, knowledge task data, and interview data and derives from them a root cause analysis and prioritization relative to the potential severity of harm. Key FDA expectations include:
The FDA makes it clear that devices cannot be entirely free of risk; some residual risk will remain. However, persistent serious use errors are generally not acceptable in premarket submissions unless there is a robust justification demonstrating that further reduction is not practicable and that the benefits outweigh the risks.
The information presented in this technical article regarding standards and regulatory guidance has been prepared to the best of the author’s knowledge and expertise. It reflects solely the opinion of the author. No guarantee can be given for the completeness, currency, or accuracy of the information provided. Standards and regulations are subject to regular revisions and updates, which may not always be reflected here in real time. This article does not constitute binding advice and does not replace a review of the applicable standards and regulatory requirements by qualified professionals or official bodies. For the application and interpretation of standards and regulatory guidance, the currently valid original documents and the respective authoritative organizations are always decisive.

As specialists in human factors and usability engineering, we at USE‑Ing. are pleased to support you in the planning, execution, and documentation of human factors validation tests in the United States. Do you have questions? Feel free to contact us.