Blog

Tag: AI-enabled Medical Devices

Usability Engineering for AI-enabled Medical Devices

Usability engineering for AI-enabled medical devices

Usability Engineering for AI-enabled Medical Devices

Practical implications for the usability engineering of AI-enabled medical devices

May 2026

Dr. Benedikt Janny, Katrin Gernert

Executive Summary

Kopie von auditive Visual Haptical Feedback 3
Usability engineering for AI-enabled medical devices confronts manufacturers with new challenges, because artificial intelligence (AI) fundamentally changes medical technology, both in how devices work technically and in the interaction between human and device. Usability engineering for AI-enabled medical devices therefore places new demands on human factors, risk management and regulatory processes. While classic medical devices rely on fixed parameters and respond predictably in identical situations, AI-based systems generate probabilistic, varying and context-dependent outputs. This dynamic considerably affects clinical workflows, role models and the cognitive demands on users.

For the safe integration of AI into medical products, the usability engineering process moves further into focus: user profiles change, tasks shift toward the assessment, plausibility check and critical interpretation of AI outputs in the sense of a human-in-the-loop (HITL) and human oversight paradigm, and new potential risks arise, in particular from misinterpretation or missing plausibility checks. At the same time, variable AI outputs and possible system errors make it harder to draw a clear line between a use error and a product error, model limitation or data-induced error.

The article shows how these changes affect all central deliverables of usability engineering, from the use specification as an increasingly dynamic artifact and the task analysis to hazard considerations, usability tests and post-market surveillance. It emphasizes the need for close collaboration between usability experts, clinical departments, risk management and AI development, as well as the importance of continuous monitoring after market launch in the context of Good Machine Learning Practice (GMLP) and AI lifecycle management. The goal is to integrate AI-enabled medical devices into existing clinical workflows in a safe, understandable and reliable way while ensuring patient and user safety as far as possible.

1AI-enabled Medical Devices & Human-AI Interaction

Artificial intelligence has become indispensable in both private and professional everyday life and is increasingly finding its way into medical technology. Alongside numerous benefits, this development also brings challenges. AI can change entire workflows and intervene directly, or through recommended actions, in medical decision-making and medical action.

With regard to the user interface, the key difference between AI-enabled and classic medical devices lies in the kind of information provided to users. Classic medical devices are based on hard-coded parameters and clearly defined limits. If, for example, air bubbles are detected in an ECMO tubing system, the system triggers an alarm and can automatically stop the pump to prevent harm.

With AI-enabled medical devices, by contrast, decisions are not based on rigid parameters but on learned patterns. This enables more complex applications, such as identifying cancerous structures in CT and MRI images, where organ- and patient-dependent tissue variations make rigid parameterization impossible. Through training with positive and negative findings, AI systems can recognize patterns and transfer them to different organ structures. The approach thus resembles human learning. At the same time, the question arises whether AI-supported medical action can be equated with a team member. Like humans, however, AI can make errors that are not solely attributable to technical defects.

A further difference is that classic medical devices always respond identically to identical input, whereas AI medical devices, owing to their ability to learn, can deliver different results for the same situation. This dynamic depends on whether a system is trained only during development or keeps learning after market launch. Because the quality of the training data largely determines the quality of the AI, faulty inputs can lead to wrong conclusions.

The following figure illustrates this human-AI interaction, which, based on the PCA approach (Perception, Cognition, Action), is represented as a closed loop between user and medical device.

Basic user interface schema for AI-enabled medical devices

Basic schema of human-AI interaction, adapted from Engler et al. Usability Engineering for

Medical Devices using Artificial Intelligence and Machine Learning Technology – A Position Paper of DKE UK 811.4 (2024)

The AI-enabled system processes data using machine learning and generates outputs such as predictions, recommendations or decisions, which are presented to users via the user interface. Users perceive the information (perception), interpret it (cognition) and act accordingly (action), which feeds new inputs back into the system. Because of continuous learning, these inputs may lead to system adaptations, which in turn can affect the interaction.

2Influences on the usability engineering process for AI-enabled medical devices

2.1 Use specification

The use of AI can automate manual tasks of clinical staff. This automation changes the demands on users and thus their user profiles (intended user profiles). The elimination of physically demanding tasks, for example, can influence the demographic composition of a user group (intended user group), which in turn affects the requirements for user interface specifications and training measures. In addition, workflows and tasks change, which may require new skills and experience.

Workplace and social exchange can also change significantly through AI-enabled medical devices. If action proposals or automated medical actions come from AI, professional exchange within the team can decrease. Automated medication delivery can also minimize interaction with patients and move workplaces further away from areas close to the patient.

These examples show how closely the context of use, the user interface and training measures are interlinked. The basic approach to determining the context of use remains unchanged, but a particularly careful analysis of users, their tasks and goals, the necessary resources and the use environment is required, because AI brings profound changes.

2.2 Task analysis (PCA task analysis)

AI medical devices can change workflows considerably. Instead of making diagnoses entirely on their own, radiologists, for example, already receive suggestions from the AI that serve as the basis for the final diagnostic decision. This initial orientation makes decisions easier but carries the risk that results are accepted without verification. A careful plausibility check is therefore essential.

PCA task analysis

Despite these changes, it remains worthwhile to analyze existing workflows before product development begins, in order to identify tasks that can be automated and those that newly arise from AI-related challenges. The following example illustrates the differences between the current-state scenarios and the task models of the new product generation combined with AI:

The role of radiologists thus shifts from detailed image analysis to assessing the plausibility of the AI output. If the AI also performs actions automatically, the task focus shifts further toward monitoring instead of active execution.

Basic workflow diagram using the example of an AI-enabled medical device

2.3 Use errors in AI-enabled medical devices & resulting hazardous situations

The definition of a use error remains unchanged when interacting with an AI. Use errors comprise actions that users perform with the medical device, or the omission of a necessary action, that lead to a result different from what the manufacturer intended or the user expected and that may give rise to hazardous situations.

Because of the AI’s upstream medical action (e.g., recommendations, predictions, triage or decision support) and the necessary plausibility check, the cognitive demands on users shift in particular. Compared with classic systems that respond deterministically in identical situations, AI-enabled medical devices can produce varying outputs in similar situations. This variability can foster uncertainty and increase the need for professional judgment. In the use situation, assessing the output itself therefore becomes part of the interaction and a decisive cognitive task.

A key future question is how to assess causes and responsibilities when users follow an AI recommendation that turns out to be medically wrong. The EU AI Act defines actor roles for this purpose, such as provider and deployer (operator or user acting under their own authority). These role definitions structure obligations along the supply and use chain, but they do not automatically resolve the practical question, case by case, of whether a failure was caused by

a) a system or information error,
b) a use error or
c) a combination of both factors.

Several uncertainties arise from this for the use error analysis:
AI outputs are not always “perfect” but can be probabilistic, incomplete or context-sensitive. A wrong or suboptimal output can be interpreted both as an expected residual risk within defined performance limits and as an indication of a product or information error (e.g., insufficient robustness, unclear intended use, missing warnings, lack of transparency about limits).
At the same time, blindly following the AI without a plausibility check can count as a use error, because it is the omission of a necessary action that is required for safe and intended use.
In practice, combined causes are realistic: a suboptimal AI output (system side) plus an insufficient plausibility check (use side) can together contribute to a hazardous situation. The root cause analysis must therefore separate systematically: What was recognizable to users in the use situation, which information and signals were available, and which actions could realistically be expected?
The plausibility check thus becomes a critical user task and at the same time a risk control measure whose effectiveness depends not only on user competence but also on support from the system (e.g., transparency, explanatory notes, warnings, display of uncertainty, logging and traceability). The AI Act addresses this lifecycle perspective through obligations for high-risk systems, in particular through requirements for traceability, documentation and post-market monitoring. Providers of high-risk AI must operate a post-market monitoring system that collects and analyzes performance data over the lifetime of the system in order to detect continued conformity as well as risks and changes in performance.

In summary, intervening or not intervening in the case of faulty, uncertain or ambiguous AI, as well as the final medical action, plays a central role in identifying use errors (and deriving suitable risk controls). Plausibility checks that are not performed or are insufficient are a plausible and probably frequent cause of error in AI-supported use contexts. At the same time, the distinction remains subject to uncertainty in practice: whether an omission counts as a use error depends largely on whether the manufacturer adequately supports the necessity, scope and feasibility of the plausibility check in the specific use situation (among other things through understandable information, a suitable UI, warnings, and traceable limits and uncertainties of the system). The exact implementation will only become clear in the future. The identification of hazard-related use scenarios can be explored in more depth in our knowledge base.

2.4 Usability evaluations

The challenges described in chapter 2.3 directly affect the planning, execution and interpretation of usability evaluations. Classic medical technology tests generally assume that a system delivers consistent and medically correct information for identical input. AI medical devices change this basic assumption: a test participant can interpret the displayed information correctly and apply it formally “right”, and still arrive at a medically wrong result if the AI output itself is faulty. This is where the plausibility check required of users and the related challenges explained in chapter 2.3 come into play. For the assessment in the usability test, the question therefore arises whether an observed problem should be classified as a use error or as a system error (or product error).

Usability teams therefore need extended clinical expertise in the test setting to avoid medical misjudgments by the moderator and note-taker team. This can be supported by the presence of clinical experts, prepared medical reference paths or closely aligned evaluation criteria.
Usability evaluation
A further challenge arises from potentially variable AI outputs: learning or adapting models can deliver different results for the same test input in different test sessions. For usability evaluation, this means that test protocols and acceptance criteria must be designed flexibly so that different system responses can be classified correctly. At the same time, evaluators must define clearly how to handle diverging outputs in order to ensure both repeatability and comparability between tests.
Usability test session

Overall, AI medical devices (AI-enabled medical devices) mean that usability evaluations must distinguish more strongly between interpretation errors by users and content errors of the system. Both can look similar but have different causes, risks and regulatory consequences. As a result, multidisciplinary teams, medical reference standards and precise documentation of system outputs during the test become far more important than in classic usability studies.

2.5 Post-market surveillance

Especially for AI-enabled medical devices that continue to be adapted or updated after market launch, post-market surveillance (PMS) gains considerable importance. Manufacturers must ensure that users can use the product efficiently and safely even when outputs change due to new training data or model updates.

The EU AI Act obliges manufacturers to maintain extended monitoring, reporting and documentation processes in order to continuously check model performance, data drift and possible interactions with other AI-enabled medical devices. These AI-specific requirements extend classic PMS and thus also affect the long-term assurance of high usability.

Usability Engineering for AI-enabled Medical Devices: Conclusion

AI fundamentally changes medical technology: instead of deterministic, clearly predictable system responses, dynamic, context-dependent outputs emerge that place new demands on users and their decision-making. In particular, the necessary plausibility check of AI results becomes the central user task and at the same time a critical risk control measure. Roles, workflows and potential sources of error shift markedly, and the clear distinction between use error and system error becomes more complex.

For manufacturers this means: usability engineering must be thought of earlier, deeper and more interdisciplinary, from the use specification to post-market surveillance. Only an integrated interplay of UX, clinical expertise, risk management and AI development can deliver safe, understandable and standards-compliant AI-enabled medical devices.

Support from USE-Ing.
If you are developing or further developing AI-based medical technology, we are happy to support you, for example with:

  • Adapting your usability engineering process to AI systems
  • Conducting PCA analyses and use error assessments
  • Developing understandable and safe user interfaces for AI outputs
  • Planning and conducting AI-specific usability tests
  • Integrating regulatory requirements (IEC 62366, EU AI Act)
Feel free to get in touch. We will bring your AI-enabled medical devices into practice safely, user-centered and in compliance with standards.

Disclaimer

The information presented in this technical article regarding standards and regulatory guidance has been prepared to the best of the author’s knowledge and expertise. It reflects solely the opinion of the author. No guarantee can be given for the completeness, currency, or accuracy of the information provided. Standards and regulations are subject to regular revisions and updates, which may not always be reflected here in real time. This article does not constitute binding advice and does not replace a review of the applicable standards and regulatory requirements by qualified professionals or official bodies. For the application and interpretation of standards and regulatory guidance, the currently valid original documents and the respective authoritative organizations are always decisive.

Contact e1749215510462

As usability engineering specialists, we at USE‑Ing. support you in developing AI-based medical technology. Benefit from our expertise, from concept to implementation. Do you have questions? Feel free to talk to us.

The USE-Ing. Compass

Stay on course with our newsletter