The summative usability evaluation is the final assessment with which a manufacturer demonstrates that the intended users can operate a medical device safely. It is carried out with the final or production-equivalent product, representative users and under realistic conditions. In regulatory usage, it corresponds to the validation of the user interface. For the US market, the FDA describes with human factors validation testing a closely related, separately defined format of evidence.
Purpose and position in the process
The summative evaluation comes at the end of the usability engineering process according to IEC 62366-1. It answers a single, clearly defined question: Can the intended users cope safely with the selected hazard-related use scenarios?
It is explicitly not an optimization instrument. Anyone who still wants to collect ideas for improvement in the summative study has set up the process wrongly. That is what formative evaluation is for. It is not designed as an exploratory optimization study; its primary purpose is the final demonstration. Observed use errors, use difficulties or unexpected use patterns must nevertheless be analyzed and, where applicable, fed back into design, risk analysis and further validation activities.
What is tested is not the entire product but a justified selection of the previously identified hazard-related use scenarios. This selection must emerge from the use-related risk analysis and be documented in a comprehensible way.
Delimitation from formative evaluation
| Formative | Summative | |
|---|---|---|
| Timing | Accompanies development | Concludes it |
| Purpose | Improve | Demonstrate |
| Product status | Prototypes, mock-ups or click dummies are sufficient | Final or production-equivalent product required, including final labeling, packaging and instructions for use |
| Criteria | Open and exploratory | Acceptance criteria defined in advance |
| Intervention | Follow-up questions and assistance are permitted and desirable | Any assistance distorts the result |
A late formative study does not become a summative evaluation merely because it takes place with the final device. What matters are planning, sample, acceptance criteria and discipline in execution.
Prerequisites: product maturity, users, environment
Three conditions determine the usability of the study:
- Production-equivalent product: The tested user interface must correspond to the series version. Later changes to safety-relevant elements can make a repetition necessary.
- Representative users: The participants must correspond to the groups described in the user profile. People with development, sales or project-specific insider knowledge are generally unsuitable and devalue the study. Prior product experience as such, by contrast, is not automatically an exclusion criterion. What matters is whether it corresponds to the real user profile, for example when real users typically bring experience with predecessor products or comparable devices.
- Realistic use environment: Relevant use environment factors such as noise, lighting, time pressure, interruptions or gloves must be reproduced. A sterile laboratory does not represent an emergency department.
In addition, it must be clarified whether and how training is represented. If the user receives an introduction in real operation, it may also take place in the test, but then realistically, including the time interval between training and use (decay period).
Sample size and user groups
IEC 62366-1 deliberately does not name a minimum number but requires a justified, representative sample. For the US market, the FDA often sets an order of magnitude of at least 15 participants per essential, separately considered user group; this guidance recommendation is not a rigid legal minimum. Whether groups have to be examined separately depends not on the job title alone but on whether there are relevant differences in tasks, experience, abilities or possible use difficulties.
Practically consequential is the number of user groups: A product for nurses, physicians and family caregivers needs three groups of 15 participants, that is, 45 people. This is exactly where a narrowly defined use specification pays off: Every additional user group included raises the testing effort considerably.
Important: The summative evaluation is not a statistical study. It is designed qualitatively: A single safety-critical use error can trigger a need for action, regardless of its frequency.
Conduct and data collection
The participants work on the selected scenarios independently. The facilitator observes and records and, as a matter of principle, does not intervene to help during task completion, not even when an error becomes obvious. Only in this way can it be assessed whether the product itself catches the error. Safety-related interventions and abort criteria defined in advance remain unaffected; any necessary assistance is to be documented as a relevant study finding.
Collected per task are: success or failure, observed use errors, close calls, use difficulties and comprehension problems with labeling or instructions for use.
The central instrument is the post-task interview: After the tasks are completed, not during them, targeted follow-up questions are asked about why the participant acted as they did, what they had expected and how they had understood the situation. Without this interview, the cause of an error remains speculation. The think-aloud method, by contrast, is unusual in the summative evaluation because thinking aloud changes behavior.
Analysis, root cause analysis and residual risks
A root cause analysis must be carried out for every observed use error and every use difficulty. Counting errors is not enough. What is required is an understanding of whether the cause lay in perception, cognition or action (perception-cognition-action).
Subsequently, it is assessed whether the error can lead to a hazardous situation and whether the resulting residual risk according to ISO 14971 is acceptable. If it is not, measures are required: primarily through design change, secondarily through protective measures and only lastly through information for safety.
Documentation and regulatory use
The results feed into the usability engineering file and must be consistent with the risk management file. For the US market, they are additionally summarized in the HFE/UE report.
The report should contain at least: objective and research question, tested scenarios with derivation from the risk analysis, participant profiles and recruitment criteria, description of the test environment, acceptance criteria defined in advance, results per task, root cause analysis of all incidents and the final residual risk assessment.
The most frequent audit finding is not a poor result but a missing derivation: Tested scenarios without a documented link to the use-related risk analysis make the entire evidence vulnerable.
The summative usability evaluation is the demonstration, not the optimization. It requires a production-equivalent product, representative users, a realistic environment, acceptance criteria defined in advance and a root cause analysis for every observed error. Its value stands or falls with the comprehensible derivation of the tested scenarios from the use-related risk analysis.
Frequently asked questions (FAQ)
How many participants does a summative usability evaluation need?
IEC 62366-1 does not name a fixed number but requires a justified, representative sample. For the US market, the FDA generally expects at least 15 participants per distinguishable user group. With several user groups, the number of participants multiplies accordingly.
May the summative evaluation be carried out with a prototype?
Only if the prototype is production-equivalent, that is, represents the final user interface including labeling, packaging and instructions for use. Early prototypes and click dummies are reserved for formative evaluation.
Does the summative evaluation have to be repeated if the product changes?
That depends on the scope of the change. If it affects safety-relevant elements of the user interface, a renewed or supplementary evaluation is required. If it affects only non-safety-relevant aspects, a justified assessment can be sufficient. The decision must be documented.
What happens if use errors occur?
Use errors do not automatically lead to failure. Each error is analyzed for its cause, and the depth of analysis depends on the risk. Subsequently, it must be assessed whether the resulting residual risk is acceptable. If it is not, design measures are primarily required. Information for safety and training measures are subordinate within the risk control hierarchy.
Is the summative evaluation the same as a clinical investigation?
No. The summative evaluation tests the safe operability of the user interface. A clinical investigation examines clinical benefit and performance in the patient. Both are separate forms of evidence with different goals, methods and regulatory requirements.
Are you planning a summative usability evaluation or do you need support with study design, recruitment and analysis? We conduct summative studies in a standards-compliant way, from scenario selection to the audit-proof report.
More about our usability engineeringSources
- IEC 62366-1:2015+AMD1:2020, Medical devices, Part 1: Application of usability engineering to medical devices
- ISO 14971:2019, Medical devices, Application of risk management to medical devices