Diagnostic Accuracy of Two Large Language Models in Turkish Emergency Department Anamnesis Notes
LLM-ED-DX-TR
1 other identifier
observational
600
1 country
1
Brief Summary
This retrospective diagnostic accuracy study evaluates two large language models - GPT-4.1 (gpt-4.1-2025-04-14; OpenAI) and Claude Sonnet 4.6 (claude-sonnet-4-6; Anthropic) - as retrospective coding-quality instruments applied to anonymized Turkish-language emergency department anamnesis notes. The reference standard is the majority consensus of three board-certified emergency medicine specialists who independently coded each note in ICD-10, blinded to one another, to the code entered by the treating physician at case closure, and to the subsequent clinical course. Cases without chapter-level majority agreement are excluded without replacement. Both models are queried once per note with a single locked prompt at temperature 0 in stateless application programming interface calls, with no retrieval augmentation, no external tools and no extended-reasoning mode. The primary outcome is the proportion of cases in which each model's rank-1 diagnosis matches the reference standard at ICD-10 chapter level, reported with a Wilson 95% confidence interval. Registered secondary outcome measures are chapter-level Cohen's kappa between each model's rank-1 diagnosis and the reference standard; top-3 chapter accuracy for each model; and chapter-level concordance between the closure ICD-10 code and the reference standard. Additional prespecified analyses set out in the statistical analysis plan (paired between-model difference, three-character accuracy, note-length association, confidence calibration and model-to-model agreement) are reported in the primary publication. The ICD-10 code entered at case closure is characterised against the same reference standard as a description of current documentation practice; it is not a comparator, and no test of superiority or inferiority against model output is performed. The analysis plan was finalised and frozen before any accuracy computation. Reporting follows STARD-AI 2025.
Trial Health
Trial Health Score
Automated assessment based on enrollment pace, timeline, and geographic reach
participants targeted
Target at P75+ for all trials
Started May 2026
Shorter than P25 for all trials
1 active site
Health score is calculated from publicly available data and should be used for screening purposes only.
Trial Relationships
Click on a node to explore related trials.
Study Timeline
Key milestones and dates
Study Start
First participant enrolled
May 1, 2026
CompletedFirst Submitted
Initial submission to the registry
June 3, 2026
CompletedFirst Posted
Study publicly available on registry
June 8, 2026
CompletedPrimary Completion
Last participant's last visit for primary outcome
August 3, 2026
CompletedStudy Completion
Last participant's last visit for all outcomes
August 7, 2026
CompletedAugust 14, 2026
August 1, 2026
3 months
June 3, 2026
August 12, 2026
Conditions
Keywords
Outcome Measures
Primary Outcomes (2)
Diagnostic Accuracy of GPT-4.1 for ICD-10 Chapter-Level Diagnosis
Proportion of cases in which the GPT-4.1 primary (rank 1) diagnosis matches the 3-specialist majority-vote reference standard at the ICD-10 chapter level (22 categories). Range: 0 to 1.00.
At the single index-test run on 3 August 2026
Diagnostic Accuracy of Claude Sonnet 4.6 for ICD-10 Chapter-Level Diagnosis
Proportion of cases in which the Claude Sonnet 4.6 primary (rank 1) diagnosis matches the 3-specialist majority-vote reference standard at the ICD-10 chapter level (22 categories). Range: 0 to 1.00.
At the single index-test run on 3 August 2026
Secondary Outcomes (5)
Cohen's Kappa Between GPT-4.1 Primary Diagnosis and the Reference Standard
At the single index-test run on 3 August 2026
Cohen's Kappa Between Claude Sonnet 4.6 Primary Diagnosis and the Reference Standard
At the single index-test run on 3 August 2026
Top-3 Diagnostic Accuracy of GPT-4.1
At the single index-test run on 3 August 2026
Top-3 Diagnostic Accuracy of Claude Sonnet 4.6
At the single index-test run on 3 August 2026
Chapter-Level Concordance Between the Closure ICD-10 Code and the Reference Standard
At the original clinical encounter (retrospective data spanning 1 May to 3 August 2026)
Study Arms (1)
Emergency Department Patient Cohort
Consecutive adult patients (aged 18 years and older) evaluated in the ambulatory (green/yellow triage) area of the emergency department, who had a free-text electronic anamnesis note recorded at presentation and an ICD-10 code entered by the treating physician at case closure. No note-completeness or minimum-length requirement was applied. The closure code is characterised against the reference standard as a description of current documentation practice; it is not a comparator, and no test of superiority or inferiority against model output is performed.
Eligibility Criteria
The study population comprises consecutive adult patients (aged 18 years and older) evaluated in the ambulatory (green/yellow triage) area of the emergency department of a tertiary care training and research hospital, and whose encounters were documented in the hospital information system (HBYS). Patients triaged to the high-acuity resuscitation area (Emergency Severity Index level 1) were excluded a priori; no Emergency Severity Index level 1 or level 2 presentation occurred within the sampling window, so resuscitation-area presentations are absent from the study population altogether. The findings do not extend to high-acuity emergency care.
You may qualify if:
- Adult patients (aged 18 years and older) presenting to the emergency department, evaluated in the ambulatory (green/yellow triage) area.
- A free-text electronic anamnesis note entered at presentation in the hospital information system (HBYS). No minimum note length and no "sufficient information for diagnosis" requirement was applied, because such a criterion preferentially retains more readily classifiable cases; note length was treated as a covariate rather than as an eligibility threshold. A note was excluded only if all three of the following were absent: any symptom statement, any duration or onset information, and a non-empty anamnesis field.
- An ICD-10 code entered by the treating emergency physician at case closure. Cases in which this entry was absent or did not form a valid ICD-10 code were retained in the analysis set and counted in the denominator of the closure-code analyses.
You may not qualify if:
- Notes lacking all three of the following: any symptom statement, any duration or onset information, and a non-empty anamnesis field.
- Pediatric cases (age under 18 years).
- Patients critically ill and triaged to high-acuity resuscitation areas (Emergency Severity Index \[ESI\] level 1).
- Clinical notes containing residual identifying information that cannot be fully de-identified, preventing compliance with data privacy regulations.
- Non-independent clinical notes consisting solely of a brief cross-reference to a prior hospital visit without a new history entry.
Contact the study team to confirm eligibility.
Sponsors & Collaborators
Study Sites (1)
Marmara University Pendik Training and Research Hospital
Istanbul, Istanbul, 34899, Turkey (Türkiye)
Related Publications (10)
Sounderajah V, Guni A, Liu X, Collins GS, Karthikesalingam A, Markar SR, Golub RM, Denniston AK, Shetty S, Moher D, Bossuyt PM, Darzi A, Ashrafian H; STARD-AI Steering Committee. The STARD-AI reporting guideline for diagnostic accuracy studies using artificial intelligence. Nat Med. 2025 Oct;31(10):3283-3289. doi: 10.1038/s41591-025-03953-8. Epub 2025 Sep 15.
PMID: 40954311BACKGROUNDNewman-Toker DE, Peterson SM, Badihian S, Hassoon A, Nassery N, Parizadeh D, Wilson LM, Jia Y, Omron R, Tharmarajah S, Guerin L, Bastani PB, Fracica EA, Kotwal S, Robinson KA. Diagnostic Errors in the Emergency Department: A Systematic Review [Internet]. Rockville (MD): Agency for Healthcare Research and Quality (US); 2022 Dec. Report No.: 22(23)-EHC043. Available from http://www.ncbi.nlm.nih.gov/books/NBK588118/
PMID: 36574484BACKGROUNDWei J et al. Chain-of-thought prompting elicits reasoning in LLMs. NeurIPS. 2022;35:24824-24837.
BACKGROUNDNiset A, Melot I, Pireau M, Englebert A, Scius N, Flament J, El Hadwe S, Al Barajraji M, Thonon H, Barrit S. Grounded large language models for diagnostic prediction in real-world emergency department settings. JAMIA Open. 2025 Oct 21;8(5):ooaf119. doi: 10.1093/jamiaopen/ooaf119. eCollection 2025 Oct.
PMID: 41127256BACKGROUNDWilliams CYK, Miao BY, Kornblith AE, Butte AJ. Evaluating the use of large language models to provide clinical recommendations in the Emergency Department. Nat Commun. 2024 Oct 8;15(1):8236. doi: 10.1038/s41467-024-52415-1.
PMID: 39379357BACKGROUNDHoppe JM, Auer MK, Struven A, Massberg S, Stremmel C. ChatGPT With GPT-4 Outperforms Emergency Department Physicians in Diagnostic Accuracy: Retrospective Analysis. J Med Internet Res. 2024 Jul 8;26:e56110. doi: 10.2196/56110.
PMID: 38976865BACKGROUNDKanjee Z, Crowe B, Rodman A. Accuracy of a Generative Artificial Intelligence Model in a Complex Diagnostic Challenge. JAMA. 2023 Jul 3;330(1):78-80. doi: 10.1001/jama.2023.8288.
PMID: 37318797BACKGROUNDTakita H, Kabata D, Walston SL, Tatekawa H, Saito K, Tsujimoto Y, Miki Y, Ueda D. A systematic review and meta-analysis of diagnostic performance comparison between generative AI and physicians. NPJ Digit Med. 2025 Mar 22;8(1):175. doi: 10.1038/s41746-025-01543-z.
PMID: 40121370BACKGROUNDShan G, Chen X, Wang C, Liu L, Gu Y, Jiang H, Shi T. Comparing Diagnostic Accuracy of Clinical Professionals and Large Language Models: Systematic Review and Meta-Analysis. JMIR Med Inform. 2025 Apr 25;13:e64963. doi: 10.2196/64963.
PMID: 40279517BACKGROUNDTaylor RA, Sangal RB, Smith ME, Haimovich AD, Rodman A, Iscoe MS, Pavuluri SK, Rose C, Janke AT, Wright DS, Socrates V, Declan A. Leveraging artificial intelligence to reduce diagnostic errors in emergency medicine: Challenges, opportunities, and future directions. Acad Emerg Med. 2025 Mar;32(3):327-339. doi: 10.1111/acem.15066. Epub 2024 Dec 15.
PMID: 39676165BACKGROUND
MeSH Terms
Conditions
Condition Hierarchy (Ancestors)
Study Officials
- PRINCIPAL INVESTIGATOR
Emir Ünal
Marmara University
Study Design
- Study Type
- observational
- Observational Model
- COHORT
- Time Perspective
- RETROSPECTIVE
- Sponsor Type
- OTHER
- Responsible Party
- PRINCIPAL INVESTIGATOR
- PI Title
- MD, Assistant Professor
Study Record Dates
First Submitted
June 3, 2026
First Posted
June 8, 2026
Study Start
May 1, 2026
Primary Completion
August 3, 2026
Study Completion
August 7, 2026
Last Updated
August 14, 2026
Record last verified: 2026-08