
How to conduct formative usability evaluations in accordance with IEC 62366‑1
What manufacturers need to know
February 2026
Executive Summary
Formative usability evaluations are a methodological core component of iterative user interface design in accordance with IEC 62366‑1, the normative reference for implementing the European Medical Device Regulation (MDR). They serve both the continuous improvement of the user interface (design focus) and the early identification of safety‑relevant use errors (safety focus). Suitable methods range from expert reviews and cognitive walkthroughs to early simulated‑use usability tests. A user interface evaluation plan supports consistency by defining objectives, methods, and criteria for formative and summative activities.
The evaluation should systematically link observations of use errors, close calls, and use difficulties with root cause analyses. What is essential is the feedback of the results into design iterations and—if safety‑relevant—into the use‑related risk analysis. Clear documentation ensures transparency and traceability across development stages. This article explains how to proceed in a targeted and efficient manner.
1The role of formative evaluations in the development process
Formative usability evaluations serve the iterative improvement of the user interface and are therefore an integral part of the user interface design methodology. In contrast to summative evidence, formative activities do not focus on final validation but on a structured cycle of learning and optimization. The aim of formative evaluations is to identify interaction problems as early as possible and to derive user interface design improvements before they become entrenched in later development stages and can only be corrected with considerable effort.
Formative evaluations also serve a risk‑oriented function. They support the early identification of potential use errors that could contribute to hazardous situations and thereby enable targeted further development of the user interface before a summative evaluation takes place. In this way, formative evaluations contribute not only to usability in terms of efficient and satisfactory operation, but also to the robustness of use in safety‑critical situations.

2The dual objective framework: design focus and safety focus
Formative evaluations can be methodologically structured along two objectives that complement each other in practice and are often addressed in parallel. Clearly distinguishing these perspectives supports precise study planning, targeted data collection, and a transparent interpretation of the results.
- Design focus
The objective is to gain insights that support improving the user interface during development. The emphasis is on aspects such as clarity, consistency, orientation, information presentation, workflow support, and the reduction of unnecessary cognitive load. The design focus examines, in particular, whether the user interface adequately supports the users’ expected mental models, whether interaction logics are intuitively comprehensible, and whether feedback, status indications, and system responses are designed in a way that enables users to make action decisions satisfactorily, effectively, and efficiently. Results from this perspective typically lead to design adjustments that improve usability, efficiency, and error tolerance, without necessarily requiring a direct safety relation in the narrower sense.
- Safety focus
The objective is the identification of previously unknown use errors that could lead to hazardous situations. While the design focus often aims at optimizing “normal” use, the safety focus specifically examines those points of interaction where incorrect actions are plausible and may have potentially safety‑critical consequences. In this context, not only actually observed use errors are considered, but also close calls, use difficulties, or strategies that indicate an increased likelihood of error. Methodologically relevant here is, in particular, the question of which contributing factors promote the occurrence of the use error (e.g., unclear feedback, ambiguities, unfavorable sequences, insufficient error prevention or detection), and how these factors can be reduced through UI design measures or accompanying risk control measures.
This differentiation is helpful in practice because it guides the selection, depth, and timing of appropriate methods. An early concept review is typically more strongly anchored in the design focus: it examines the fundamental interaction logic, consistency, terminology, information architecture, and potential cognitive loads, often before high‑fidelity prototypes are available. A formative usability test (e.g., simulated use at a more advanced development stage), by contrast, can be more strongly oriented toward the safety focus: it addresses risk‑relevant use paths, evaluates the robustness of the user interface in critical situations, and makes safety‑critical incorrect actions as well as close calls empirically visible.
In practical implementation, this means that formative evaluations should not be understood as “a type of test,” but rather as a methodically tiered sequence of activities. Depending on the development phase, the emphasis may shift between the design focus and the safety focus. An explicit assignment of objectives also facilitates appropriately weighting the results: while design‑related findings often lead to incremental optimizations, safety‑related findings regularly result in prioritized actions that directly feed into risk justification and subsequent evaluation planning.
3Team and set of methods for formative evaluations
The analysis of formative evaluations is typically conducted in a multidisciplinary manner, as use‑related problems are rarely attributable solely to “usability” in the narrow sense, but often arise from an interplay of clinical application context, technical constraints, information presentation, work processes, and training. Accordingly, the involvement of relevant stakeholders—such as engineering, clinical/medical, quality/regulatory, risk management, and design—is methodologically advantageous in order to classify observations correctly and translate them into actionable design and risk measures. In addition, the involvement of representative users is useful and often necessary, provided that the development stage, accessibility, and study setting allow for it. Representative users enable, in particular, the validation or refutation of assumptions regarding mental models, routines, contextual factors, and typical error strategies, which are often only partially identifiable through internal reviews alone.
From a methodological standpoint, it is also important to recognize that formative evaluations can vary in their level of rigor depending on the development phase. In early stages, rapid, analytical methods are often prioritized, as they can be conducted even when prototypes are still low‑fidelity. As the user interface matures, the emphasis shifts toward empirical methods that make interaction patterns observable under more realistic conditions. As a result, a combined approach is often most effective: analytical methods for broadly identifying weaknesses, and empirical methods for confirming, prioritizing, and analyzing root causes.
Suitable methods include, among others:
- Expert reviews for the rapid identification of obvious interaction weaknesses. They enable a structured examination of the user interface for inconsistencies, ambiguities, missing feedback, potential paths to incorrect use, and generally known interaction principles or heuristics.
- Cognitive walkthroughs for analyzing workflows and mental models, particularly for complex operating sequences or safety‑critical decision points. They specifically address the question of whether users can plausibly infer the intended action steps from the user interface.
- Early‑stage usability tests (simulated use) for observing real interaction behavior using prototype setups. Early‑stage tests allow empirical observation of interaction behavior—including use difficulties, close calls, and potential use errors—under realistic (simulated) conditions. They are especially valuable for capturing actual user strategies that are often underestimated in analytical methods (e.g., workarounds, shortcuts, error‑prone routines, responses to stress or distraction). Depending on the development stage, such tests can be conducted with low‑ to high‑fidelity prototypes, with increasing prototype maturity improving the meaningfulness of results regarding timing, feedback, and interaction dynamics. For risk‑relevant use paths in particular, simulated‑use tests provide an opportunity to empirically verify hypotheses derived from task analysis and hazard‑related use scenarios and to derive prioritized design changes.
4The formative workflow model: planning, execution, iteration
An effective workflow model for formative evaluations typically follows a repeatable cycle of planning, execution, analysis, and iteration. What matters is that formative activities are not viewed as isolated actions, but as a guided sequence whose results are transparently fed into design decisions and—where relevant—into risk‑based justification. The methodological aim is to combine the exploratory nature of formative evaluations with sufficient structure so that insights can be reliably compared, prioritized, and documented.
Step 1: Establish the user interface evaluation plan
A user interface evaluation plan provides the organizational and methodological framework for formative (and typically also summative) activities. It should define the objectives of each evaluation, the methods to be used, the evaluation criteria applied, and how the evaluation is embedded within the overall development and risk logic. Particularly important is the explicit connection to risk‑ and scenario‑based requirements, as formative evaluations often focus on specific interaction paths, UI characteristics, or use situations that are relevant in the risk context.
A consistent plan prevents results from remaining isolated findings and instead establishes a traceable structure: Which question was addressed by which method? Which version of the design was examined? Which user groups and use contexts were included? This clarity is essential for subsequent iterations, as it is the only way to determine whether improvements were truly effective or whether problems merely shifted elsewhere.
Step 2: Conducting the evaluation with appropriate realism
Formative tests also benefit from realistic tasks and plausible contextual conditions, though always within an intentionally exploratory logic. Here, realism is not primarily aimed at regulatory demonstration but at validating assumptions about use, decision‑making processes, and environmental conditions. A suitable setup reflects typical workflows and relevant contextual factors sufficiently well without losing the focus on learning and optimization goals.
During the execution of the evaluation, it is especially important that task descriptions are formulated in a way that triggers real decision‑making rather than simply “checking UI functions.” Depending on the prototype’s maturity, realism can be achieved through various means—for example, by simulating relevant use situations, providing appropriate materials (e.g., draft Instructions for Use (IFUs)), incorporating typical time and attention constraints, or designing realistic transitions between steps. At the same time, the formative logic remains flexible: if it becomes apparent during the execution of the evaluation that critical issues lie in a different interaction point than initially expected, adjusting the focus is methodologically acceptable as long as it is transparently documented.
Step 3: Analyzing the identified use problems
The analysis of formative evaluations should not be limited to a simple list of problems, but should instead be consistently root‑cause oriented. The central question is which observations are actually relevant for design decisions and which factors explain the difficulties that were observed. This includes the structured consideration of:
use errors (actions or omissions with potentially risk‑relevant consequences),
close calls (near misses that often indicate an increased likelihood of error),
use difficulties (challenges, delays, misunderstandings, or workarounds).
A robust analysis examines which UI characteristics, which information presentations, which interaction sequences, or which contextual conditions contribute to the observed problems. Particularly valuable is the identification of recurring patterns: Do several participants encounter difficulties at the same point? Do different users exhibit similar error strategies? Are there typical corrective actions (e.g., backtracking, repeated attempts, aborting) that indicate insufficient system transparency or inadequate error prevention?
The cause‑and‑effect reasoning should be formulated in a way that allows concrete design measures to be derived from it. In this way, formative results become a methodological decision‑making tool rather than merely a collection of observations.
Step 4: Design modifications and feedback into risk‑related activities
The central output of formative evaluations consists of prioritized design changes and a transparent rationale explaining why these changes are suitable for reducing the observed problems. In practice, it is helpful to document design actions not merely as “fixes,” but as part of a structured iteration: Which change addresses which root cause? Which assumption does it modify? Which interaction point does it stabilize?
In parallel, feedback into risk activities is required whenever the results indicate new or confirmed use errors that could be safety‑relevant. This particularly applies in cases where close calls or use difficulties can be interpreted as precursors to potential hazardous situations. In such cases, the insights should be fed back into the use‑related risk analysis to maintain consistency in the risk logic and to strengthen the basis for risk‑based planning of subsequent evaluations.
This creates a closed learning cycle: formative evaluations not only improve usability but also strengthen the safety‑related rationale by making risk‑relevant insights visible early and translating them into targeted design and risk measures.
5Documentation requirements
The guideline requires careful selection of participants, training where appropriate, and data collection with a focus on the critical tasks. It also emphasizes that individuals who frequently participate in usability tests of the same device or other devices from the same manufacturer should be excluded. Notably, compared with other guidelines, the NMPA guideline explicitly expects a justification when no device‑specific training is required for test participants. The test reports should include detailed information on the objectives, simulation conditions, results related to use errors and their associated root causes, as well as any deviations.
Formative evaluations must be documented in a way that makes the results traceable, methodologically reproducible, and usable for subsequent activities. In practice, this is typically achieved through evaluation protocols and evaluation reports that not only capture individual observations but also establish a consistent line of reasoning from objectives to methods to resulting actions. Here, the decisive factor is not the sheer level of detail, but structured transparency: third parties should be able to understand why the evaluation was conducted, how it was conducted, what was observed, and which conclusions were methodologically derived from it.
A complete documentation set typically includes the following elements:
Objectives and research questions: Clear specification of whether the evaluation was primarily design‑oriented, safety‑oriented, or a combination of both, and which UI aspects or use scenarios were the focus.
Methodology and study design: Description of the method(s) applied (e.g., expert review, cognitive walkthrough, usability test (simulated‑use)), including justification for the method selection in relation to the maturity level of the user interface.
Sample and representativeness: Characterization of the users or experts involved, including relevant attributes related to the intended user profiles (e.g., experience, qualification, potential limitations), and—where applicable—a justification of the representativeness and its limitations.
Test object and versioning: Clear indication of which version of the user interface or which prototype stage was evaluated, including relevant supporting materials (e.g., draft IFUs, training assumptions).
Tasks and scenarios: Description of the tasks, use scenarios, and conditions used—ideally in a way that makes the logic of task derivation and the intended triggering of specific interaction paths recognizable.
Data collection and analysis logic: Explanation of which data were collected (e.g., observations, error classification, qualitative notes), how observations were structured, and which criteria guided interpretation and prioritization.
Results presented in a structured form: Consistent documentation of use errors, close calls, and use difficulties, including context, observed action sequence, and—where possible—a root cause analysis or contributing factors.
Derived measures and design decisions: Concrete, transparently justified actions derived from the results, including prioritization and rationale explaining why the measure is suitable for addressing the observed root cause.
What is essential here is the ability to build on prior results: the subsequent development documentation should clearly show how formative findings have informed design decisions and the use‑related risk analysis. This includes, in particular, the connection between the observed problems, their interpretation in terms of root causes, and the design changes that have been implemented or planned.
In addition, the documentation plays an important steering function throughout the project. It enables assessment across iterations of whether measures were effective, whether issues have shifted, or whether new risks have emerged. Clear versioning and a consistent results structure also facilitate later preparation for summative evaluation activities, as they make it evident which risk‑relevant interaction points have already been addressed and which aspects remain critical.
Disclaimer
The information presented in this technical article regarding standards and regulations has been prepared to the best of the author’s knowledge and expertise. It reflects solely the author’s opinion. No guarantee can be given for the completeness, currency, or accuracy of the information provided. Standards and regulations are subject to regular revisions and changes, which may not always be reflected here in a timely manner. This article does not constitute binding advice and does not replace the review of the applicable standards and regulations by qualified professionals or official bodies. For the application and interpretation of standards and regulations, the currently valid original documents and the respective competent organizations are authoritative.

As usability engineering specialists, we at USE‑Ing. are happy to support you in the planning, execution, and documentation of formative usability evaluations. Do you have any questions? Feel free to contact us.
