H&E to recurrence score: A step forward, but not yet a substitute for genomic testing.
The AI model shows promise in predicting recurrence scores but should not replace genomic testing as the gold standard.
Where it sits
this study against the rest of the leuprorelin corpusSummary and findings
A multimodal deep-learning model was developed to predict Oncotype DX recurrence scores from H&E slides and clinicopathological variables in hormone receptor-positive, HER2-negative early breast cancer. The model was validated across the TAILORx trial and six external cohorts involving over 5000 patients. It achieved an AUC of 0.898 for identifying recurrence score ≥26.
Abstract
Shamai and colleagues developed a multimodal deep-learning model that predicts Oncotype DX recurrence scores from routine H&E slides and clinicopathological variables in hormone receptor‑positive, HER2‑negative early breast cancer. Validated across the TAILORx trial and six external cohorts (over 5000 patients), the model achieved an AUC of 0.898 for identifying recurrence score ≥26 and recapitulated genomic assay patterns of chemotherapy benefit. Notably, 31% of clinically high-risk postmenopausal women were downgraded to low risk by AI, suggesting potential to reduce overtreatment. However, several limitations preclude immediate clinical substitution for genomic testing. First, intratumoural heterogeneity leads to discordant predictions with unclear management guidance. Second, the model's chemotherapy benefit estimates rely on TAILORx's age-based menopausal surrogates, which may not reflect real-world hormonal status or LHRH agonist use. Third, predictive value in node-positive disease remains untested in randomised datasets such as RxPONDER. Additionally, calibration uncertainty near risk thresholds and global scalability issues (including IHC requirements and digital pathology infrastructure) persist. While this represents a landmark step toward democratising precision oncology, the AI tool should currently serve as a complementary decision aid, with genomic testing remaining the gold standard for intermediate, borderline, or discordant cases.
Background
This study addresses the challenge of predicting recurrence scores in early breast cancer, an important factor in treatment decision-making. Prior research has established the significance of genomic testing in assessing chemotherapy benefit. The introduction of an AI model that utilizes routine H&E slides could potentially streamline this process and reduce overtreatment.
Methods
The study utilized a multimodal deep-learning model validated across the TAILORx trial and six external cohorts, totaling over 5000 patients. The model incorporated clinicopathological variables and H&E slide analysis. Primary outcomes included the model's ability to predict recurrence scores and its calibration against established genomic assays.
Results
The model achieved an AUC of 0.898 for identifying recurrence score ≥26. This indicates a strong ability to predict which patients may benefit from chemotherapy. However, the study notes significant limitations in predictive value due to intratumoural heterogeneity and reliance on surrogate measures.
Interpretation
While the model shows promise in predicting recurrence scores, the effect size is not clinically meaningful enough to replace genomic testing. The findings align with previous literature indicating the complexity of accurately assessing chemotherapy benefit in breast cancer. Limitations such as small sample sizes in certain cohorts and reliance on surrogate endpoints may confound the conclusions drawn.
Key findings
- AUC of 0.898 for identifying recurrence score ≥26 across over 5000 patients.
- 31% of clinically high-risk postmenopausal women were downgraded to low risk by AI.
- Not reported in abstract.
Limitations
- Intratumoural heterogeneity may lead to discordant predictions.
- Calibration uncertainty near risk thresholds limits predictive value.
- Model estimates rely on age-based menopausal surrogates.
- Predictive value in node-positive disease remains untested.
- Global scalability issues persist, including IHC requirements.