Evaluating the interpretability of clinical speech AI models: Lessons from two user studies.
Current AI interpretation designs may mislead clinicians and require alignment with clinical reasoning for effective integration into practice.
Where it sits
this study against the rest of the adamax corpusSummary and findings
The study evaluated the interpretability of SHAP-based AI model interpretations in clinical speech applications, specifically dysarthria detection, through two user studies. Eight factors were examined, including faithfulness, cognitive load, and user trust, revealing risks in current interpretation practices. The study found that the bar-chart design often misled speech-language pathology students regarding feature influence and clinical severity.
Abstract
The deployment of Artificial intelligence (AI) in clinical speech applications has been limited in large part by the lack of interpretability, which is essential for establishing clinician trust and enabling effective decision support. Although methods such as SHapley Additive exPlanations (SHAP) aim to improve transparency in many clinical domains, their applicability to clinical speech-language pathology practice is uncertain. Since these methods rely on data modalities like acoustic signal features and spectrograms, which are unfamiliar to clinicians and misaligned with clinical workflows, the resulting interpretations may introduce additional burden and bias rather than provide clinically meaningful insight. To better understand this challenge, we conducted two consecutive user studies to systematically evaluate a commonly used SHAP-based interpretation design (a bar chart showing the influence of acoustic features on AI decisions) in dysarthria detection. Building on our prior works, eight factors were examined: faithfulness, computational efficiency, cognitive load, human-AI task performance, mental model, user trust, clinical understandability, and decision relevance. The results reveal a previously unrecognized risk in current interpretation practices. The seemingly intuitive bar-chart design frequently misled participating speech-language pathology (SLP) students to interpret feature influence as an indicator of clinical severity. Other findings include difficulty understanding AI mechanisms, discrepancies between human and model reasoning, and the limited ability of interpretations to address clinical questions. Through this work, we highlight the need for interpretation designs that are more closely aligned with clinical reasoning patterns and suggest practical considerations for developing speech-based AI systems that can be meaningfully integrated into clinical practice.
Background
The study addresses the challenge of interpretability in AI models used for clinical speech applications, a critical factor for clinician trust and effective decision support. Current methods like SHAP aim to enhance transparency but may not align well with clinical workflows, potentially introducing bias. This research is important as it explores how AI interpretations can be better integrated into clinical practice, particularly in speech-language pathology.
Methods
The study conducted two consecutive user studies to evaluate a SHAP-based interpretation design in dysarthria detection. The design involved a bar chart showing the influence of acoustic features on AI decisions. Eight factors were assessed: faithfulness, computational efficiency, cognitive load, human-AI task performance, mental model, user trust, clinical understandability, and decision relevance. The participants were speech-language pathology students.
Results
The primary finding was that the bar-chart design misled participants to interpret feature influence as an indicator of clinical severity. Participants also faced difficulties in understanding AI mechanisms and noted discrepancies between human reasoning and model outputs. The interpretations provided limited insights into clinical questions, highlighting a need for designs that align with clinical reasoning.
Interpretation
The study suggests that current SHAP-based interpretations may not effectively support clinical decision-making in speech-language pathology. The findings are consistent with prior concerns about the misalignment of AI interpretations with clinical workflows. The study underscores the need for AI interpretation designs that are more intuitive and clinically relevant, though the small and specific sample limits generalizability.
Key findings
- Eight factors were examined: faithfulness, computational efficiency, cognitive load, human-AI task performance, mental model, user trust, clinical understandability, and decision relevance.
- The bar-chart design frequently misled SLP students to interpret feature influence as an indicator of clinical severity.
- There were difficulties in understanding AI mechanisms and discrepancies between human and model reasoning.
- Interpretations had limited ability to address clinical questions.
- The study highlights the need for interpretation designs aligned with clinical reasoning patterns.
Limitations
- Findings based on speech-language pathology students.
- No specific numeric outcomes reported.
- Limited generalizability to broader clinical settings.