Construction of a Benchmark for Breast Ultrasound AI Interpretation and Performance Evaluation of Multimodal AI Models
BUST-AI Bench
Construction of a Standardized Benchmark Evaluation System for Intelligent Breast Ultrasound Image Interpretation and Systematic Performance Assessment of Multimodal Artificial Intelligence Models Based on ACR BI-RADS v2025 Criteria
3 other identifiers
observational
1,380
1 country
1
Brief Summary
This single-center, retrospective, observational study aims to construct a standardized benchmark evaluation system for intelligent breast ultrasound image interpretation and to systematically assess the diagnostic performance of current mainstream multimodal artificial intelligence (AI) models. De-identified B-mode breast ultrasound images with confirmed pathological diagnoses will be retrospectively collected from the institutional archive (2018-2025) and supplemented with images from published open-access datasets. Expert radiologists with varying experience levels will independently annotate all images according to the American College of Radiology (ACR) Breast Imaging Reporting and Data System (BI-RADS) v2025 criteria, including glandular tissue composition, lesion characterization (mass vs. non-mass lesion), morphological descriptors, and final BI-RADS classification. Baseline deep learning models (CNN-based ResNet-50 and Transformer-based USFM) will be trained to establish performance baselines and to stratify cases by diagnostic difficulty through cross-architecture consensus. Multiple multimodal large language models (MLLMs), including both general-purpose and medical-domain models, will then be evaluated via standardized API calls using BI-RADS-guided chain-of-thought prompts at temperature 0 for reproducibility. Primary endpoints include BI-RADS classification accuracy and diagnostic AUC for benign-malignant differentiation. Model robustness and safety will be assessed through out-of-distribution rejection testing, temperature-stability experiments, and thinking-mode ablation studies. This study adheres to the FLAIR and TRIPOD-LLM reporting guidelines.
Trial Health
Trial Health Score
Automated assessment based on enrollment pace, timeline, and geographic reach
participants targeted
Target at P75+ for all trials
Started Mar 2026
Shorter than P25 for all trials
1 active site
Health score is calculated from publicly available data and should be used for screening purposes only.
Trial Relationships
Click on a node to explore related trials.
Study Timeline
Key milestones and dates
Study Start
First participant enrolled
March 12, 2026
CompletedFirst Submitted
Initial submission to the registry
March 24, 2026
CompletedFirst Posted
Study publicly available on registry
March 30, 2026
CompletedPrimary Completion
Last participant's last visit for primary outcome
December 1, 2026
ExpectedStudy Completion
Last participant's last visit for all outcomes
March 1, 2027
March 30, 2026
March 1, 2026
9 months
March 24, 2026
March 24, 2026
Conditions
Keywords
Outcome Measures
Primary Outcomes (2)
Diagnostic Accuracy for Pathological Diagnosis
Sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and F1 score of AI models for benign-malignant classification, with histopathological diagnosis as the gold standard.
At study completion, approximately 12 months
BI-RADS Classification Accuracy
Overall accuracy of AI models in assigning BI-RADS categories (2, 3, 4A, 4B, 4C, 5) to breast ultrasound images, compared with expert consensus annotation as the reference standard.
At study completion, approximately 12 months
Secondary Outcomes (3)
Agreement with Expert Consensus (Cohen's Kappa)
At study completion, approximately 12 months
Out-of-Distribution Rejection Rate
At study completion, approximately 12 months
Sensitivity, Specificity, PPV, NPV, and F1 Score
At study completion, approximately 12 months
Study Arms (3)
Normal Breast
Breast ultrasound images showing normal glandular tissue across different tissue composition types, with no focal lesions identified. Confirmed by senior radiologist review.
Benign Lesion
Breast ultrasound images containing pathologically confirmed benign lesions (BI-RADS 2-4B), including fibroadenoma, cyst, lipoma, sclerosing adenosis, intraductal papilloma, and selected non-mass lesions (NML).
Malignant Lesion
Breast ultrasound images containing pathologically confirmed malignant lesions (BI-RADS 3-5), including invasive ductal carcinoma, invasive lobular carcinoma, mucinous carcinoma, and selected non-mass lesions (NML).
Interventions
Retrospective evaluation of de-identified breast ultrasound images by multiple AI systems, including baseline deep learning models (ResNet-50, USFM) and multimodal large language models, using standardized BI-RADS-guided chain-of-thought prompts via API. No patient contact or clinical decision-making is involved.
Eligibility Criteria
De-identified breast ultrasound images from adult patients who underwent breast ultrasound examination at Peking Union Medical College Hospital between 2018 and 2025 with subsequent pathological confirmation, supplemented by images from published, ethics-approved, open-access breast ultrasound datasets (e.g., BUSI, BrEaST).
You may qualify if:
- B-mode breast ultrasound grayscale images from the institutional PACS database or from published open-access breast ultrasound datasets with documented original institutional ethics approval
- Image quality adequate for clinical diagnosis with clear visualization of the region of interest
- Pathological diagnosis confirmed (for benign and malignant lesion groups), or normal breast status confirmed by a senior radiologist with \>15 years of breast ultrasound experience (for the normal group)
- Complete de-identification with removal of all personally identifiable information
You may not qualify if:
- Severely degraded image quality precluding meaningful BI-RADS assessment
- Duplicate images from the same patient (only the most representative image retained per lesion)
- Images with residual personally identifiable information after de-identification processing
- Cases with ambiguous, disputed, or unavailable pathological results
- Non-B-mode ultrasound images, including elastography, contrast-enhanced ultrasound, and Doppler imaging
Contact the study team to confirm eligibility.
Sponsors & Collaborators
Study Sites (1)
Peking Union Medical College Hospital
Beijing, 100730, China
Related Publications (12)
Bi WL, Hosny A, Schabath MB, Giger ML, Birkbak NJ, Mehrtash A, Allison T, Arnaout O, Abbosh C, Dunn IF, Mak RH, Tamimi RM, Tempany CM, Swanton C, Hoffmann U, Schwartz LH, Gillies RJ, Huang RY, Aerts HJWL. Artificial intelligence in cancer imaging: Clinical challenges and applications. CA Cancer J Clin. 2019 Mar;69(2):127-157. doi: 10.3322/caac.21552. Epub 2019 Feb 5.
PMID: 30720861BACKGROUNDBhayana R, Krishna S, Bleakney RR. Performance of ChatGPT on a Radiology Board-style Examination: Insights into Current Strengths and Limitations. Radiology. 2023 Jun;307(5):e230582. doi: 10.1148/radiol.230582. Epub 2023 May 16.
PMID: 37191485BACKGROUNDClusmann J, Kolbinger FR, Muti HS, Carrero ZI, Eckardt JN, Laleh NG, Loffler CML, Schwarzkopf SC, Unger M, Veldhuizen GP, Wagner SJ, Kather JN. The future landscape of large language models in medicine. Commun Med (Lond). 2023 Oct 10;3(1):141. doi: 10.1038/s43856-023-00370-1.
PMID: 37816837BACKGROUNDSeyyed-Kalantari L, Zhang H, McDermott MBA, Chen IY, Ghassemi M. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nat Med. 2021 Dec;27(12):2176-2182. doi: 10.1038/s41591-021-01595-0. Epub 2021 Dec 10.
PMID: 34893776BACKGROUNDMoor M, Banerjee O, Abad ZSH, Krumholz HM, Leskovec J, Topol EJ, Rajpurkar P. Foundation models for generalist medical artificial intelligence. Nature. 2023 Apr;616(7956):259-265. doi: 10.1038/s41586-023-05881-4. Epub 2023 Apr 12.
PMID: 37045921BACKGROUNDSung H, Ferlay J, Siegel RL, Laversanne M, Soerjomataram I, Jemal A, Bray F. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA Cancer J Clin. 2021 May;71(3):209-249. doi: 10.3322/caac.21660. Epub 2021 Feb 4.
PMID: 33538338BACKGROUNDBenary M, Wang XD, Schmidt M, Soll D, Hilfenhaus G, Nassir M, Sigler C, Knodler M, Keller U, Beule D, Keilholz U, Leser U, Rieke DT. Leveraging Large Language Models for Decision Support in Personalized Oncology. JAMA Netw Open. 2023 Nov 1;6(11):e2343689. doi: 10.1001/jamanetworkopen.2023.43689.
PMID: 37976064BACKGROUNDMiaojiao S, Xia L, Xian Tao Z, Zhi Liang H, Sheng C, Songsong W. Using a Large Language Model for Breast Imaging Reporting and Data System Classification and Malignancy Prediction to Enhance Breast Ultrasound Diagnosis: Retrospective Study. JMIR Med Inform. 2025 Jun 11;13:e70924. doi: 10.2196/70924.
PMID: 40498674BACKGROUNDJiao J, Zhou J, Li X, Xia M, Huang Y, Huang L, Wang N, Zhang X, Zhou S, Wang Y, Guo Y. USFM: A universal ultrasound foundation model generalized to tasks and organs towards label efficient image analysis. Med Image Anal. 2024 Aug;96:103202. doi: 10.1016/j.media.2024.103202. Epub 2024 May 15.
PMID: 38788326BACKGROUNDXiang H, Wang X, Xu M, Zhang Y, Zeng S, Li C, Liu L, Deng T, Tang G, Yan C, Ou J, Lin Q, He J, Sun P, Li A, Chen H, Heng PA, Lin X. Deep Learning-assisted Diagnosis of Breast Lesions on US Images: A Multivendor, Multicenter Study. Radiol Artif Intell. 2023 Jul 12;5(5):e220185. doi: 10.1148/ryai.220185. eCollection 2023 Sep.
PMID: 37795135BACKGROUNDCollins GS, Moons KGM, Dhiman P, Riley RD, Beam AL, Van Calster B, Ghassemi M, Liu X, Reitsma JB, van Smeden M, Boulesteix AL, Camaradou JC, Celi LA, Denaxas S, Denniston AK, Glocker B, Golub RM, Harvey H, Heinze G, Hoffman MM, Kengne AP, Lam E, Lee N, Loder EW, Maier-Hein L, Mateen BA, McCradden MD, Oakden-Rayner L, Ordish J, Parnell R, Rose S, Singh K, Wynants L, Logullo P. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024 Apr 16;385:e078378. doi: 10.1136/bmj-2023-078378.
PMID: 38626948BACKGROUNDKottlors J, Iuga AI, Bluethgen C, Bressem K, Kather JN, Moy L, Wald C, Wang W, Liu T, Ranschaert E, Dratsch T, Kleesiek J, Gertz RJ, Rajpurkar P, Bedayat A, Fink MA, Zeeck A, Chaudhari A, Alkasab T, Wu H, Nensa F, Wang B, Grosse Hokamp N, Laukamp KR, Persigehl T, Maintz D, Truhn D, Lennartz S. Guidelines for Reporting Studies on Large Language Models in Radiology: An International Delphi Expert Survey. Radiology. 2026 Feb;318(2):e250913. doi: 10.1148/radiol.250913.
PMID: 41631991BACKGROUND
MeSH Terms
Conditions
Condition Hierarchy (Ancestors)
Study Officials
- PRINCIPAL INVESTIGATOR
Qingli Zhu, MD
Peking Union Medical College Hospital
Central Study Contacts
Study Design
- Study Type
- observational
- Observational Model
- COHORT
- Time Perspective
- RETROSPECTIVE
- Sponsor Type
- OTHER
- Responsible Party
- PRINCIPAL INVESTIGATOR
- PI Title
- Professor, Department of Ultrasound, Peking Union Medical College Hospital
Study Record Dates
First Submitted
March 24, 2026
First Posted
March 30, 2026
Study Start
March 12, 2026
Primary Completion (Estimated)
December 1, 2026
Study Completion (Estimated)
March 1, 2027
Last Updated
March 30, 2026
Record last verified: 2026-03
Data Sharing
- IPD Sharing
- Will share
- Shared Documents
- STUDY PROTOCOL, SAP, ANALYTIC CODE
- Time Frame
- Within 6 months of primary publication, available indefinitely
- Access Criteria
- Open access via a recognized data repository (to be determined)
The de-identified benchmark evaluation dataset, including expert-annotated breast ultrasound images with paired BI-RADS reading reports, is planned for public release to promote academic reproducibility and collaborative research.