fda zulassung medizinprodukte

FDA approval for medical devices: Human factors validation explained in simple terms

What manufacturers of medical devices need to know for FDA approval

January 2026

Executive Summary

Kopie von auditive Visual Haptical Feedback 3

In the FDA context, human factors validation testing serves as a component of design validation in accordance with the FDA design control requirements (formerly 21 CFR Part 820.30, incorporated into ISO 13485:2016, clause 7.3, by the Quality Management System Regulation, QMSR, since February 2026), with the specific objective of demonstrating that the medical device can be used safely and effectively by the intended users in the intended use environments. The FDA has outlined its expectations for this in the guidance document “Applying Human Factors and Usability Engineering to Medical Devices.” In May 2026, the FDA additionally finalized its guidance on the content of human factors information in marketing submissions. A central foundation is the risk‑based identification and categorization of critical tasks based on the severity of potential harm, all of which must be fully covered in the test. The FDA typically expects usability tests to be conducted as simulated‑use testing with the final user interface, realistic use conditions, representative user populations (generally at least 15 participants per relevant user group), and a test execution that enables independent and natural use without moderator influence. Data collection should combine observational data, knowledge task data where applicable, and post‑use interviews. The analysis is qualitative and requires a root cause analysis for use errors, close calls, and use difficulties, as well as a robust justification of residual risk. This article explains in detail how to proceed with this process.

1Regulatory context and objectives

In the context of FDA approval, the FDA Guidances emphasize the risk‑based nature of human factors and usability engineering and focus primarily on ensuring that the user interface is designed such that use errors with potentially harmful consequences are eliminated or reduced as much as possible. Human factors validation testing is conducted at the end of development to demonstrate that the final user interface supports safe and effective use under expected conditions of use and that the implemented risk control measures are effective.

FDA Devices

2Critical tasks as the guiding test unit in human factors validation

Within the context of FDA approval, a central structural principle is the definition of critical tasks. These are tasks for which incorrect performance or failure to perform the task could result in serious harm to the patient or user, including compromised medical care. The identification and justification of these critical tasks is relevant not only for the test design but also shapes the evaluation approach and the argumentation presented in the report.

With respect to critical tasks, the following principles apply:

  • all critical tasks must be performed during the validation test.
  • critical tasks are derived in advance from the risk analysis.
  • the list is dynamic and may expand as understanding of user–device interaction increases.

The FDA recommends identifying critical tasks through a combination of analytical and empirical approaches:

  • Analytical approaches include task analysis (including PCA‑based methods), failure mode and effects analysis (FMEA), or fault tree analysis (FTA), ideally conducted by a multidisciplinary team.
  • Empirical approaches include contextual inquiry, interviews, and formative evaluations (including cognitive walk‑throughs and simulated‑use formative testing).

An important aspect of the FDA’s reasoning is that the probability of occurrence of use errors is difficult to quantify reliably in actual use; therefore, the severity of potential harm is the primary factor used to prioritize tasks.

3Development of the human factors validation test protocol

To plan the study in a methodologically correct manner, the development of a test protocol is required. The FDA typically requires that human factors validation be conducted as a usability test in the form of simulated‑use testing and specifies clear minimum requirements for the study design. The validation test is not an exploratory format; rather, it is a purposefully constructed demonstration intended to show that the final user interface can be used safely and effectively by the intended users in the intended use environments, particularly with respect to the critical tasks.

Key FDA requirements for the study design

At its core, the FDA requires that human factors validation testing ensures that:

  • the test participants represent the intended users,

  • all critical tasks are actually performed during the test,

  • the user interface reflects the final design,

  • and the test conditions are sufficiently realistic so that the findings can be generalized to actual use.

These requirements must not merely be stated in the protocol but defined operationally. This means that it must be documented in a traceable manner which user characteristics are considered representative (and why), how full coverage of all critical tasks is ensured, which design version is considered “final” (including labeling and the instructions for use (IFU)), and which environmental and contextual factors are relevant for achieving realistic test conditions.

Principle of realism and test environment

Simulated‑use conditions must be realistic enough for the results to be generalizable to actual use. The FDA emphasizes that relevant environmental factors must be incorporated into the simulation based on risk, insofar as they could affect user interaction with the device. The guidance cites examples such as dim lighting, distractions, and multitasking.

For protocol planning, this means that realism is not achieved through superficial “scenery,” but through the targeted representation of those conditions that are plausible contributing factors to use errors. The protocol should therefore explicitly define:

  • which environmental parameters are relevant for realism (e.g., lighting, noise, spatial constraints, device setup, typical distractions),

  • which of these parameters will be controlled, which will be varied, and which will intentionally not be addressed in the test,

  • and how these decisions are justified with respect to the use‑related risk analysis and the critical tasks.

At the same time, the simulation must remain methodologically stable: realism must not introduce uncontrollable confounding variables. Therefore, the FDA requirement focuses on “sufficient realism,” not on a full replication of a clinical environment. What matters is the logical justification that the essential interaction conditions have been adequately represented.

Methodological distinction: No moderator influence, no think‑aloud

For validation, participants must use the device independently and naturally. The FDA explicitly states that think‑aloud techniques are not acceptable in validation studies, as verbalization alters behavior and therefore does not represent real‑use conditions. The protocol must therefore be structured so that relevant data are derived primarily from observation. Interactions should not be influenced by prompts, questions, or “assistance” from the study environment. If questions are necessary (e.g., for root‑cause clarification), they should generally be shifted to a post‑use interview to avoid distorting actual task performance.

Labeling and IFU as part of the final user interface

The FDA considers labeling—including the instructions for use (IFU), device labels, packaging, quick‑start guides, and comparable informational components—to be part of the user interface. Accordingly, the final versions of these materials must be used in the validation test. This is methodologically relevant because many critical tasks depend not only on interaction design but also on the clarity, accessibility, and applicability of the information provided. Human factors validation therefore implicitly assesses whether users, under realistic use conditions, can make correct decisions and perform critical tasks accurately using the provided labeling. At the same time, the FDA makes clear that labeling or training must not serve as a default compensation for design‑related use errors on critical tasks. If use errors on critical tasks occur during validation, relying solely on labeling or training modifications—without additional evidence demonstrating their effectiveness—is not acceptable.

Components of the test protocol typically expected under the FDA’s logic

Although the guidance does not prescribe a rigid structure, we recommend that the following aspects be described in the test protocol:

  • test objectives with explicit reference to safe and effective use and the critical tasks,

  • description of the final user interface, including the labeling set,

  • use scenarios and task sets from which the test cases are derived to ensure full coverage of the critical tasks,

  • success criteria for each task, clearly defining what constitutes successful task performance,

  • training requirements consistent with real‑use conditions,

  • test environment and realism parameters,

  • data collection methods (primarily observational, supplemented by interviews and knowledge tasks),

  • rules for moderator behavior and handling of deviations.

A systematically developed test protocol demonstrates that the study was not merely “performed,” but methodologically constructed to truly meet the FDA’s requirements for human factors validation as a risk‑based demonstration.

4Sample and training

In the context of FDA approval, the FDA recommends a practical minimum of 15 participants for human factors validation tests; if multiple relevant user groups are involved, at least 15 participants per user group must be included. User groups differ particularly when characteristics or task responsibilities vary in ways that are likely to affect interaction with the device (e.g., lay users versus healthcare professionals, pediatric versus adult users). When selecting test participants, it is essential that they are U.S. residents.

Representativeness and relevant variability in user characteristics

Participants should represent the relevant range within the user group, particularly with regard to characteristics that may influence interactions (e.g., sensory, physical, or cognitive limitations, literacy/language skills, experience). For certain medical conditions that may cause functional limitations, corresponding representativeness considerations must be explicitly taken into account. Employees of the manufacturer should not serve as participants in human factors validation tests (exceptions apply only in rare, highly specialized cases).

Training and real‑use consistency (including training decay)

The training approach applied in the testing must reflect real‑use conditions: if actual users receive minimal or no training, test participants must not receive more extensive training. Training immediately before the test is problematic; the FDA highlights training decay and expects a time interval (depending on the device, up to a day) to realistically represent retention.

5Data collection – observational data, knowledge tasks, post‑use interviews

For successful FDA approval, the FDA expects that data collection in human factors validation testing is based primarily on observable task performance, supplemented by information that cannot be reliably inferred from behavior (e.g., comprehension, interpretation, decision‑making knowledge). Accordingly, the data collection approach in the protocol should clearly define which data sources will be used (observational data, knowledge task data, interviews).

Observational data as primary evidence for task performance

According to the FDA, successful performance of critical tasks must be assessed primarily through observation. Observational data therefore constitute the key source of evidence for determining whether the user interface supports safe and effective use under the intended conditions.

From a methodological perspective, this means that the test record (observation log) must precisely document:

  • What constitutes task completion (success criteria for each task, including required subtasks, correct system states, and correct decisions).

  • Which deviations are to be classified as use errors (e.g., incorrect selection, incorrect setting, omission of a safety‑critical action, improper response to feedback/alarms).

  • Which observable indicators must be captured as close calls or use difficulties (e.g., search behavior, repeated attempts, inconsistent strategies, hesitation at decision points).

The FDA also notes that timing measurements are only meaningful when speed is clinically relevant; otherwise, no timing criteria should be introduced, as this can distort the validation logic (e.g., shifting the focus from “correct” to “fast”). 

Systematically capturing use errors, close calls, and use difficulties

During the execution of the simulated‑use test, it is essential to document in detail any problems that occur during user interaction. The FDA categorizes these issues as use errors, close calls, and use difficulties.

Use errors are defined as “a user action or lack of action that was different from that expected by the manufacturer and that resulted in an outcome that (1) was different from the outcome expected by the user, and (2) was not caused solely by device failure, and (3) did or could result in harm.” The term “harm” explicitly includes compromised medical care.

Close calls are errors that did not lead to an incorrect final outcome due to self‑correction or accidental recovery. Close calls are methodologically important because they often indicate an increased likelihood of error, even when the task was formally completed successfully.

In addition, use difficulties must be captured—not as incidental notes, but as systematic indicators, particularly when they occur repeatedly or cluster around safety‑critical points of interaction. The FDA identifies the following typical observable markers, among others:

  • Repeated attempts

  • Hesitation

  • Confusion

  • Search behavior

Knowledge task data for non‑observable comprehension tasks

Part of the critical use safety depends not only on physical operation but also on knowledge and comprehension, for example with respect to:

  • contraindications,

  • warnings,

  • environmental conditions/vulnerabilities,

  • maintenance or safety‑critical preparatory steps.

Such aspects cannot often be reliably inferred from observable behavior alone—particularly when the test setup cannot practically represent every rare contextual condition. For this purpose, the FDA recommends knowledge tasks, i.e., knowledge and comprehension questions designed to assess whether participants can correctly interpret relevant information and make appropriate decisions.

Key methodological requirements include:

  • Use of open‑ended, neutrally phrased questions (no leading questions, no multiple‑choice prompts as the default).
  • Separation of observation and knowledge inquiry: knowledge tasks should be placed so that they do not pre‑condition participant behavior during the use scenarios.
  • Documentation of the information basis: if knowledge related to labeling or the IFU is being assessed, it must be clear whether and how participants had access to the material and whether they were able to find and interpret the information.

Post‑use interview as a complement, not a substitute

In the FDA approach, post‑use interviews are a complementary data source. They are intended to deepen the observational data but must not replace it. Interviews serve in particular to:

  • systematically explore the root causes of observed use errors or problems,

  • capture subjectively experienced use difficulties that may not have been clearly observable,

  • and understand the interaction logic from the user’s perspective (e.g., mental models, expectations, interpretation of information).

The FDA recommends conducting interviews after the completion of the use scenarios to avoid influencing task performance. Methodologically, it is essential that the exploration of each observed use error or problem is conducted systematically—using open, neutrally phrased questions (e.g., “What did you expect at this point?”, “How did you interpret this information?”, “What led you to this decision?”) rather than confirmatory or leading formulations.

In addition, interviews should explicitly provide space for:

  • subjectively reported difficulties (even when no errors were observed),

  • perceived ambiguities in terminology, symbols, or feedback,

  • and the reasons for recovery strategies used in close calls.

6Analysis – root cause analysis and residual risk

In the context of FDA approval, the FDA positions human factors validation testing as primarily a qualitative approach: the goal is not to estimate statistical frequencies, but to identify design‑related root causes of use errors or problems and to derive effective risk controls.

Data aggregation with the aim of root cause analysis

The analysis aggregates observational data, knowledge task data, and interview data and derives from them a root cause analysis and prioritization relative to the potential severity of harm. Key FDA expectations include:

  • Use errors or problems that could result in serious harm require a clear analysis of which user interface component was involved and how the interaction led to the error.
  • Seemingly obvious causes must still be explored in the interview; the user’s perspective can provide critical insights.
  • If modifications are implemented, retesting may be required to demonstrate their effectiveness and to rule out new risks.

Residual risk – when “acceptable” is regulatory defensible

The FDA makes it clear that devices cannot be entirely free of risk; some residual risk will remain. However, persistent serious use errors are generally not acceptable in premarket submissions unless there is a robust justification demonstrating that further reduction is not practicable and that the benefits outweigh the risks.

Disclaimer

The information presented in this technical article regarding standards and regulatory guidance has been prepared to the best of the author’s knowledge and expertise. It reflects solely the opinion of the author. No guarantee can be given for the completeness, currency, or accuracy of the information provided. Standards and regulations are subject to regular revisions and updates, which may not always be reflected here in real time. This article does not constitute binding advice and does not replace a review of the applicable standards and regulatory requirements by qualified professionals or official bodies. For the application and interpretation of standards and regulatory guidance, the currently valid original documents and the respective authoritative organizations are always decisive.

Contact e1749215510462

As specialists in human factors and usability engineering, we at USE‑Ing. are pleased to support you in the planning, execution, and documentation of human factors validation tests in the United States. Do you have questions? Feel free to contact us.

The USE-Ing. Compass

Stay on course with our newsletter