Dr.-Ing. Benedikt JannySenior Usability Engineer | Managing Partner
Last updated: October 2026
Short definition

Human factors validation testing is the test at the end of product development described in the FDA human factors guidance. It is intended to show that the intended users can use a medical device under the expected conditions of use without serious use errors. The FDA classifies it as part of design validation.

Classification and purpose

Human factors validation testing is a term of the US regulatory authority FDA. In its FDA human factors guidance (“Applying Human Factors and Usability Engineering to Medical Devices”, version of August 2026), it defines it as a test at the end of product development. It evaluates the interaction of users with the user interface in order to detect use errors that lead or could lead to serious harm to the patient or user. At the same time, it serves to assess the effectiveness of risk control measures.

According to the guidance, the test is intended to show that the product can be used by the intended users, for the intended uses and under the expected conditions of use without serious use errors or problems. It is to be designed comprehensively, capture use errors that are attributable to the design of the user interface and be conducted so that the results can be transferred to actual use.

The FDA classifies human factors validation as one part of design validation (see also validation). The guidance itself contains non-binding recommendations. In practice, however, the FDA measures submissions against these expectations.

Delimitation from the summative evaluation according to IEC 62366-1

In purpose and position in the process, human factors validation corresponds to the summative evaluation according to IEC 62366-1. The two terms should not be equated, however. The FDA itself points this out in its guidance: Human factors validation testing is sometimes referred to as “summative testing”, but summative tests can be defined differently, and some definitions leave out essential components of human factors validation.

In practice, the following differences in particular carry weight:

  • Scope: The FDA expects all critical tasks to be performed in the test. IEC 62366-1 works with a justified selection of hazard-related use scenarios.
  • Number of participants: As a rule, the FDA names at least 15 participants per distinguishable user group. IEC 62366-1 names no minimum number.
  • Residence of participants: For the evidence toward the FDA, the participants should live in the United States.
  • Documentation: For a submission in the US, the results are summarized in the HFE/UE report.
FDA: human factors validationIEC 62366-1: summative evaluation
ScopeAll critical tasks are performed in the testJustified selection of hazard-related use scenarios
Number of participantsAs a rule, at least 15 participants per distinguishable user groupNo minimum number named

A summative evaluation according to IEC 62366-1 therefore does not automatically meet the expectations of the FDA. Anyone who wants to cover both markets with one study takes these points into account already in planning.

Requirements for the study design

The guidance names four basic requirements:

  • The participants represent the intended users of the product.
  • All critical tasks are performed in the test.
  • The user interface corresponds to the final design.
  • The test conditions are realistic enough to reflect the actual conditions of use.

As a rule, the test takes place under simulated conditions of use (simulated use). Environmental factors that can influence the interaction with the product belong in the simulation. As examples, the guidance names dimmed light, several simultaneous alarms, distractions and multitasking.

The participants should use the product as independently and naturally as possible, without influence from the test facilitator. According to the guidance, the think-aloud method is not acceptable in human factors validation because thinking aloud does not reflect actual usage behavior.

Tasks that follow one another logically in use are combined into use scenarios and described in the test plan (test protocol). Before the test, it must be defined for each task which performance counts as success. Time measurements make sense only if the speed of use is clinically relevant. The FDA encourages manufacturers to submit the draft test plan for feedback via a Pre-Submission before conducting it.

If a simulation is not sufficient to assess the interaction with the product, data can also be collected under actual conditions of use (actual use) or as part of a clinical study.

Participants and user groups

The most important criterion is that the participants represent the intended users (see representative users). Manufacturers determine the number of participants themselves. As a rule, the guidance names at least 15 participants; for certain product types, the recommended minimum can be higher.

If a product has several distinguishable user groups, at least 15 participants from each group should be represented. Groups are considered distinguishable if their characteristics are likely to influence the interaction with the product or if they perform different tasks on the product. Examples in the guidance are different age groups, medical professionals and lay persons, and different roles such as installation or maintenance.

Within a group, the participants should reflect the range of relevant characteristics. If a product is intended for patients whose condition can cause functional limitations, people with such limitations belong in the sample.

Two further requirements concern selection. Employees of the manufacturer should not be used as participants, apart from rare exceptions. In addition, the participants should live in the United States because clinical practice, units of measurement and language in other countries can influence the results. The FDA reviews exceptions case by case on the basis of a well-founded justification.

Labeling, instructions for use and training

The labeling must correspond to the final version in the test. This applies to labels on the product and accessories, information on the display, packaging, instructions for use, manuals, package inserts and quick guides. If the labeling is available to users in real use, it should also be available in the test. The participants decide for themselves, however, whether to use it.

The training in the test should correspond to the training that real users receive. If most users in practice receive little or no training, the participants should not be trained either. Because what has been learned fades over time (training decay), the test should not follow the training immediately. Depending on the product, a break of one hour can be sufficient; in other cases an interval of one or several days is appropriate.

If use errors occur in the test with critical tasks, the FDA does not readily accept a reference to a changed instructions for use or to additional training. Additional data are required that demonstrate that the change effectively reduces the risk to an acceptable level.

Data collection

The test plan defines which data are collected. The guidance distinguishes three sources:

  • Observational data: Whether a critical task was performed successfully is captured by observation and not derived from the participants' assessment alone. Use errors, close calls and anomalies such as repeated attempts or visible confusion are recorded.
  • Knowledge tasks: Some critical tasks depend on the understanding of information, for example contraindications and warnings. This knowledge is tested with open, neutrally worded questions.
  • Interview data: An interview follows after the use scenarios are completed. It supplements the observation but does not replace it. Each observed use error is discussed there with the participant to clarify how and why it arose from their point of view.

According to the guidance, a close call exists when a user has a difficulty or makes a use error that could lead to harm but prevents the harm through their own intervention.

Analysis and residual risk

The results are analyzed qualitatively. Observational data, knowledge tasks and interview data are brought together in order to determine the cause for all use errors and problems, that is, also for close calls and use difficulties (see root cause analysis). The causes are then assessed in connection with the associated risks in order to estimate the potential for harm and to prioritize further risk control measures.

If the results lead to changes to the user interface, a renewed validation can be necessary. If the changes concern only individual aspects of use, the repeated test can be limited to these.

A completely risk-free product does not exist; a residual risk remains (see risk management). If serious use errors persist, the FDA accepts this in a submission only if the results have been analyzed thoroughly and the submission shows that further reduction is not possible or not practicable and that the benefit of the product outweighs the residual risks. The FDA does not accept the announcement to fix identified design flaws only in a later product version.

Documentation

The results of human factors validation belong to the human factors documentation of a submission and are summarized in the HFE/UE report. Which human factors information a submission should contain is described by the FDA in the guidance “Content of Human Factors Information in Medical Device Marketing Submissions”.

How planning, execution and analysis look in detail is described in our technical article “How do you conduct human factors validation tests in compliance with FDA requirements?”.

In brief

Human factors validation is the test expected by the FDA at the end of development. It is intended to show that the intended users can use the product under realistic conditions without serious use errors.

It is closely related to the summative evaluation according to IEC 62366-1 but not congruent with it: All critical tasks must be tested, as a rule with at least 15 participants per user group living in the United States.

Frequently asked questions (FAQ)

Is human factors validation testing the same as a summative evaluation?

No. Both come at the end of development and pursue a comparable purpose. The FDA itself points out, however, that summative tests can be defined differently and that some definitions omit essential components of human factors validation. Differences exist above all in the scope of the tasks tested, the number of participants and the residence of the participants.

How many participants does the FDA expect?

Manufacturers determine the number themselves. As a rule, the guidance names at least 15 participants and, with several distinguishable user groups, at least 15 per group. For certain product types, the recommended minimum can be higher.

Do the participants have to live in the United States?

In principle yes. The FDA justifies this with differences in clinical practice, units of measurement and language that can influence the results. It reviews exceptions case by case on the basis of a well-founded justification.

Is thinking aloud allowed in human factors validation?

No. The guidance describes the think-aloud method in validation as not acceptable because it does not reflect actual usage behavior. In formative evaluations, by contrast, it can be useful.

What applies if use errors occur in the test with critical tasks?

Each error is analyzed for its cause and assessed in connection with the associated risk. The FDA accepts a changed instructions for use or additional training as a measure only with additional data that demonstrate its effectiveness. After changes to the user interface, a renewed validation can be necessary.

Are you planning a human factors validation for the US market? We support you in planning, conducting and documenting human factors validation tests in the United States.

More about our usability engineering

Sources

Related terms

← Back to the wiki overview