Large Language Models Versus Anesthesiologists for ASA Physical Status Classification
ASA-LLM
Comparison of Clinical Assessment and Large Language Models in Preoperative Risk Classification: A Retrospective Analysis of ChatGPT, DeepSeek, Gemini, and Claude in ASA Physical Status Classification
1 other identifier
observational
350
0 countries
N/A
Brief Summary
The American Society of Anesthesiologists Physical Status (ASA-PS) classification is a cornerstone of preoperative risk assessment, yet interrater variability among clinicians is well documented. Large language models (LLMs) have recently demonstrated expert-level performance in several clinical classification tasks, including ASA-PS assignment. This retrospective observational study evaluates whether four widely used LLMs - ChatGPT, DeepSeek, Gemini, and Claude - can accurately and consistently assign ASA-PS classes from structured, fully anonymized clinical vignettes derived from real preoperative anesthesia evaluations, using a consensus of senior anesthesiologists as the reference standard. No patient data will be transmitted to third-party platforms. Clinical information will be converted by the investigators into de-identified structured vignettes containing only age range, sex, body mass index range, presence or absence of systemic diseases, functional capacity, and the major/minor nature of the planned surgery, in full compliance with national data protection legislation (KVKK).
Trial Health
Trial Health Score
Automated assessment based on enrollment pace, timeline, and geographic reach
participants targeted
Target at P75+ for all trials
Started Jul 2026
Shorter than P25 for all trials
Health score is calculated from publicly available data and should be used for screening purposes only.
Trial Relationships
Click on a node to explore related trials.
Study Timeline
Key milestones and dates
First Submitted
Initial submission to the registry
July 6, 2026
CompletedFirst Posted
Study publicly available on registry
July 10, 2026
CompletedStudy Start
First participant enrolled
July 21, 2026
CompletedPrimary Completion
Last participant's last visit for primary outcome
August 21, 2026
ExpectedStudy Completion
Last participant's last visit for all outcomes
October 21, 2026
July 10, 2026
July 1, 2026
1 month
July 6, 2026
July 6, 2026
Conditions
Keywords
Outcome Measures
Primary Outcomes (1)
Agreement between LLM-assigned and reference-standard ASA-PS class
Quadratic weighted Cohen's kappa between each large language model's ASA-PS assignment (ChatGPT, DeepSeek, Gemini, Claude) and the reference standard defined by consensus of a blinded panel of at least three senior anesthesiologists. Agreement of at least "good" level (weighted kappa ≥ 0.60) is hypothesized.
Through study completion, an average of 3 months
Secondary Outcomes (2)
Overall classification accuracy of each LLM
Through study completion, an average of 3 months
Subgroup error patterns
Through study completion, an average of 3 months
Study Arms (1)
elective surgery patients
Adult patients (≥18 years) who underwent preoperative anesthesia evaluation before elective surgery. Anonymized structured vignettes derived from their records will be classified by four LLMs (ChatGPT, DeepSeek, Gemini, Claude) and by a blinded senior anesthesiologist panel serving as the reference standard.
Eligibility Criteria
Adult patients who underwent preoperative anesthesia evaluation before elective surgery at a tertiary university hospital in Istanbul, Turkey.
You may qualify if:
- Age 18 years or older
- Planned elective surgery
- Completed preoperative anesthesia evaluation
You may not qualify if:
- Emergency surgical procedures
- ASA VI (brain death)
- Incomplete clinical records
Contact the study team to confirm eligibility.
Sponsors & Collaborators
Related Publications (5)
Chen YH, Ruan SJ, Chen PF. Predicting 30-Day Postoperative Mortality and American Society of Anesthesiologists Physical Status Using Retrieval-Augmented Large Language Models: Development and Validation Study. J Med Internet Res. 2025 Jun 3;27:e75052. doi: 10.2196/75052.
PMID: 40460423BACKGROUNDCheng T, Li Y, Gu J, He Y, He G, Zhou P, Li S, Xu H, Bao Y, Wang X. The performance of ChatGPT in day surgery and pre-anesthesia risk assessment: a case-control study of 150 simulated patient presentations. Perioper Med (Lond). 2024 Nov 21;13(1):111. doi: 10.1186/s13741-024-00469-6.
PMID: 39574189BACKGROUNDYoon SB, Lee J, Lee HC, Jung CW, Lee H. Comparison of NLP machine learning models with human physicians for ASA Physical Status classification. NPJ Digit Med. 2024 Sep 28;7(1):259. doi: 10.1038/s41746-024-01259-6.
PMID: 39341936BACKGROUNDChung P, Fong CT, Walters AM, Aghaeepour N, Yetisgen M, O'Reilly-Shah VN. Large Language Model Capabilities in Perioperative Risk Prediction and Prognostication. JAMA Surg. 2024 Aug 1;159(8):928-937. doi: 10.1001/jamasurg.2024.1621.
PMID: 38837145BACKGROUNDTuran EI, Baydemir AE, Ozcan FG, Sahin AS. Evaluating the accuracy of ChatGPT-4 in predicting ASA scores: A prospective multicentric study ChatGPT-4 in ASA score prediction. J Clin Anesth. 2024 Sep;96:111475. doi: 10.1016/j.jclinane.2024.111475. Epub 2024 Apr 23.
PMID: 38657530BACKGROUND
MeSH Terms
Conditions
Condition Hierarchy (Ancestors)
Central Study Contacts
Study Design
- Study Type
- observational
- Observational Model
- COHORT
- Time Perspective
- RETROSPECTIVE
- Sponsor Type
- OTHER
- Responsible Party
- PRINCIPAL INVESTIGATOR
- PI Title
- asistan prof
Study Record Dates
First Submitted
July 6, 2026
First Posted
July 10, 2026
Study Start
July 21, 2026
Primary Completion (Estimated)
August 21, 2026
Study Completion (Estimated)
October 21, 2026
Last Updated
July 10, 2026
Record last verified: 2026-07
Data Sharing
- IPD Sharing
- Will not share