Evaluating AI in postoperative rehabilitation of axillary web syndrome after breast cancer surgery: ChatGPT vs DeepSeek.
DeepSeek outperformed ChatGPT in rehabilitation decision support for axillary web syndrome, but findings should be interpreted with caution due to the limited case number.
Where it sits
this study against the rest of the evuzamitide corpusSummary and findings
This study compared the performance of two AI models, ChatGPT and DeepSeek, in the rehabilitation management of axillary web syndrome after breast cancer surgery. Five virtual cases were evaluated by three rehabilitation specialists using a standardized scoring rubric. DeepSeek showed superior performance in several assessment dimensions compared to ChatGPT.
Abstract
<h4>Objective</h4>To systematically compare two generative artificial intelligence (AI) models, ChatGPT and DeepSeek in the rehabilitation management for axillary web syndrome (AWS) following breast cancer surgery; to assess their potential value in supporting clinical rehabilitation decision-making.<h4>Methods</h4>Five virtual cases, based on real clinical scenarios and representing varying levels of complexity and functional impairment, were constructed. Three rehabilitation medicine specialists independently evaluated the models' responses using a standardised scoring rubric. Assessment dimensions included diagnostic accuracy, comprehensiveness of evaluation, individualisation of treatment planning, safety and risk awareness, evidence-based support, and clarity of logic. Additional analyses examined clinical reasoning depth, error patterns, and differences in response style.<h4>Results</h4>Both models demonstrated high inter-rater reliability. For overall clinical competency scores, DeepSeek performed significantly better than ChatGPT. DeepSeek showed superior performance in diagnostic accuracy (<i>P</i> = 0.006), comprehensiveness of evaluation (<i>P</i> = 0.034), and individualisation of therapeutic planning (<i>P</i> = 0.005). DeepSeek maintained robust performance across cases of varying complexity, with particularly strong advantages in diagnostically challenging cases and those involving concomitant lymphedema. ChatGPT more frequently exhibited generalised treatment recommendations and omissions in differential diagnoses, whereas DeepSeek, despite occasionally provided recommendations with lower levels of evidence, had a lower overall error rate. DeepSeek generated more actionable and context-specific responses, often including concrete exercise prescriptions and technical details. It also outperformed ChatGPT in reasoning chain completeness, depth of evidence justification, and mechanistic explanation rationality.<h4>Conclusion</h4>In this exploratory pilot study based on a limited number of virtual cases, DeepSeek demonstrates relatively better performance than ChatGPT in AWS rehabilitation decision support. However, findings should be interpreted cautiously, and further validation in real-world clinical settings is required.