NCT07632859

Brief Summary

This retrospective diagnostic accuracy study evaluates two large language models - GPT-4.1 (gpt-4.1-2025-04-14; OpenAI) and Claude Sonnet 4.6 (claude-sonnet-4-6; Anthropic) - as retrospective coding-quality instruments applied to anonymized Turkish-language emergency department anamnesis notes. The reference standard is the majority consensus of three board-certified emergency medicine specialists who independently coded each note in ICD-10, blinded to one another, to the code entered by the treating physician at case closure, and to the subsequent clinical course. Cases without chapter-level majority agreement are excluded without replacement. Both models are queried once per note with a single locked prompt at temperature 0 in stateless application programming interface calls, with no retrieval augmentation, no external tools and no extended-reasoning mode. The primary outcome is the proportion of cases in which each model's rank-1 diagnosis matches the reference standard at ICD-10 chapter level, reported with a Wilson 95% confidence interval. Registered secondary outcome measures are chapter-level Cohen's kappa between each model's rank-1 diagnosis and the reference standard; top-3 chapter accuracy for each model; and chapter-level concordance between the closure ICD-10 code and the reference standard. Additional prespecified analyses set out in the statistical analysis plan (paired between-model difference, three-character accuracy, note-length association, confidence calibration and model-to-model agreement) are reported in the primary publication. The ICD-10 code entered at case closure is characterised against the same reference standard as a description of current documentation practice; it is not a comparator, and no test of superiority or inferiority against model output is performed. The analysis plan was finalised and frozen before any accuracy computation. Reporting follows STARD-AI 2025.

Trial Health

87
On Track

Trial Health Score

Automated assessment based on enrollment pace, timeline, and geographic reach

Enrollment
600

participants targeted

Target at P75+ for all trials

Timeline
Completed

Started May 2026

Shorter than P25 for all trials

Geographic Reach
1 country

1 active site

Status
completed

Health score is calculated from publicly available data and should be used for screening purposes only.

Trial Relationships

Click on a node to explore related trials.

Study Timeline

Key milestones and dates

Study Start

First participant enrolled

May 1, 2026

Completed
1 month until next milestone

First Submitted

Initial submission to the registry

June 3, 2026

Completed
5 days until next milestone

First Posted

Study publicly available on registry

June 8, 2026

Completed
2 months until next milestone

Primary Completion

Last participant's last visit for primary outcome

August 3, 2026

Completed
4 days until next milestone

Study Completion

Last participant's last visit for all outcomes

August 7, 2026

Completed
Last Updated

August 14, 2026

Status Verified

August 1, 2026

Enrollment Period

3 months

First QC Date

June 3, 2026

Last Update Submit

August 12, 2026

Conditions

Keywords

Large Language Model; GPT-4o; Claude 4.6 Sonnet; ICD-10; Clinical Coding; Turkish; Emergency Department; Diagnostic Accuracy; STARD; STARD-AI

Outcome Measures

Primary Outcomes (2)

  • Diagnostic Accuracy of GPT-4.1 for ICD-10 Chapter-Level Diagnosis

    Proportion of cases in which the GPT-4.1 primary (rank 1) diagnosis matches the 3-specialist majority-vote reference standard at the ICD-10 chapter level (22 categories). Range: 0 to 1.00.

    At the single index-test run on 3 August 2026

  • Diagnostic Accuracy of Claude Sonnet 4.6 for ICD-10 Chapter-Level Diagnosis

    Proportion of cases in which the Claude Sonnet 4.6 primary (rank 1) diagnosis matches the 3-specialist majority-vote reference standard at the ICD-10 chapter level (22 categories). Range: 0 to 1.00.

    At the single index-test run on 3 August 2026

Secondary Outcomes (5)

  • Cohen's Kappa Between GPT-4.1 Primary Diagnosis and the Reference Standard

    At the single index-test run on 3 August 2026

  • Cohen's Kappa Between Claude Sonnet 4.6 Primary Diagnosis and the Reference Standard

    At the single index-test run on 3 August 2026

  • Top-3 Diagnostic Accuracy of GPT-4.1

    At the single index-test run on 3 August 2026

  • Top-3 Diagnostic Accuracy of Claude Sonnet 4.6

    At the single index-test run on 3 August 2026

  • Chapter-Level Concordance Between the Closure ICD-10 Code and the Reference Standard

    At the original clinical encounter (retrospective data spanning 1 May to 3 August 2026)

Study Arms (1)

Emergency Department Patient Cohort

Consecutive adult patients (aged 18 years and older) evaluated in the ambulatory (green/yellow triage) area of the emergency department, who had a free-text electronic anamnesis note recorded at presentation and an ICD-10 code entered by the treating physician at case closure. No note-completeness or minimum-length requirement was applied. The closure code is characterised against the reference standard as a description of current documentation practice; it is not a comparator, and no test of superiority or inferiority against model output is performed.

Eligibility Criteria

Age18 Years+
Sexall
Healthy VolunteersNo
Age GroupsAdult (18-64), Older Adult (65+)
Sampling MethodNon-Probability Sample
Study Population

The study population comprises consecutive adult patients (aged 18 years and older) evaluated in the ambulatory (green/yellow triage) area of the emergency department of a tertiary care training and research hospital, and whose encounters were documented in the hospital information system (HBYS). Patients triaged to the high-acuity resuscitation area (Emergency Severity Index level 1) were excluded a priori; no Emergency Severity Index level 1 or level 2 presentation occurred within the sampling window, so resuscitation-area presentations are absent from the study population altogether. The findings do not extend to high-acuity emergency care.

You may qualify if:

  • Adult patients (aged 18 years and older) presenting to the emergency department, evaluated in the ambulatory (green/yellow triage) area.
  • A free-text electronic anamnesis note entered at presentation in the hospital information system (HBYS). No minimum note length and no "sufficient information for diagnosis" requirement was applied, because such a criterion preferentially retains more readily classifiable cases; note length was treated as a covariate rather than as an eligibility threshold. A note was excluded only if all three of the following were absent: any symptom statement, any duration or onset information, and a non-empty anamnesis field.
  • An ICD-10 code entered by the treating emergency physician at case closure. Cases in which this entry was absent or did not form a valid ICD-10 code were retained in the analysis set and counted in the denominator of the closure-code analyses.

You may not qualify if:

  • Notes lacking all three of the following: any symptom statement, any duration or onset information, and a non-empty anamnesis field.
  • Pediatric cases (age under 18 years).
  • Patients critically ill and triaged to high-acuity resuscitation areas (Emergency Severity Index \[ESI\] level 1).
  • Clinical notes containing residual identifying information that cannot be fully de-identified, preventing compliance with data privacy regulations.
  • Non-independent clinical notes consisting solely of a brief cross-reference to a prior hospital visit without a new history entry.

Contact the study team to confirm eligibility.

Sponsors & Collaborators

Study Sites (1)

Marmara University Pendik Training and Research Hospital

Istanbul, Istanbul, 34899, Turkey (Türkiye)

Location

Related Publications (10)

  • Sounderajah V, Guni A, Liu X, Collins GS, Karthikesalingam A, Markar SR, Golub RM, Denniston AK, Shetty S, Moher D, Bossuyt PM, Darzi A, Ashrafian H; STARD-AI Steering Committee. The STARD-AI reporting guideline for diagnostic accuracy studies using artificial intelligence. Nat Med. 2025 Oct;31(10):3283-3289. doi: 10.1038/s41591-025-03953-8. Epub 2025 Sep 15.

    PMID: 40954311BACKGROUND
  • Newman-Toker DE, Peterson SM, Badihian S, Hassoon A, Nassery N, Parizadeh D, Wilson LM, Jia Y, Omron R, Tharmarajah S, Guerin L, Bastani PB, Fracica EA, Kotwal S, Robinson KA. Diagnostic Errors in the Emergency Department: A Systematic Review [Internet]. Rockville (MD): Agency for Healthcare Research and Quality (US); 2022 Dec. Report No.: 22(23)-EHC043. Available from http://www.ncbi.nlm.nih.gov/books/NBK588118/

    PMID: 36574484BACKGROUND
  • Wei J et al. Chain-of-thought prompting elicits reasoning in LLMs. NeurIPS. 2022;35:24824-24837.

    BACKGROUND
  • Niset A, Melot I, Pireau M, Englebert A, Scius N, Flament J, El Hadwe S, Al Barajraji M, Thonon H, Barrit S. Grounded large language models for diagnostic prediction in real-world emergency department settings. JAMIA Open. 2025 Oct 21;8(5):ooaf119. doi: 10.1093/jamiaopen/ooaf119. eCollection 2025 Oct.

    PMID: 41127256BACKGROUND
  • Williams CYK, Miao BY, Kornblith AE, Butte AJ. Evaluating the use of large language models to provide clinical recommendations in the Emergency Department. Nat Commun. 2024 Oct 8;15(1):8236. doi: 10.1038/s41467-024-52415-1.

    PMID: 39379357BACKGROUND
  • Hoppe JM, Auer MK, Struven A, Massberg S, Stremmel C. ChatGPT With GPT-4 Outperforms Emergency Department Physicians in Diagnostic Accuracy: Retrospective Analysis. J Med Internet Res. 2024 Jul 8;26:e56110. doi: 10.2196/56110.

    PMID: 38976865BACKGROUND
  • Kanjee Z, Crowe B, Rodman A. Accuracy of a Generative Artificial Intelligence Model in a Complex Diagnostic Challenge. JAMA. 2023 Jul 3;330(1):78-80. doi: 10.1001/jama.2023.8288.

    PMID: 37318797BACKGROUND
  • Takita H, Kabata D, Walston SL, Tatekawa H, Saito K, Tsujimoto Y, Miki Y, Ueda D. A systematic review and meta-analysis of diagnostic performance comparison between generative AI and physicians. NPJ Digit Med. 2025 Mar 22;8(1):175. doi: 10.1038/s41746-025-01543-z.

    PMID: 40121370BACKGROUND
  • Shan G, Chen X, Wang C, Liu L, Gu Y, Jiang H, Shi T. Comparing Diagnostic Accuracy of Clinical Professionals and Large Language Models: Systematic Review and Meta-Analysis. JMIR Med Inform. 2025 Apr 25;13:e64963. doi: 10.2196/64963.

    PMID: 40279517BACKGROUND
  • Taylor RA, Sangal RB, Smith ME, Haimovich AD, Rodman A, Iscoe MS, Pavuluri SK, Rose C, Janke AT, Wright DS, Socrates V, Declan A. Leveraging artificial intelligence to reduce diagnostic errors in emergency medicine: Challenges, opportunities, and future directions. Acad Emerg Med. 2025 Mar;32(3):327-339. doi: 10.1111/acem.15066. Epub 2024 Dec 15.

    PMID: 39676165BACKGROUND

MeSH Terms

Conditions

Emergencies

Condition Hierarchy (Ancestors)

Disease AttributesPathologic ProcessesPathological Conditions, Signs and Symptoms

Study Officials

  • Emir Ünal

    Marmara University

    PRINCIPAL INVESTIGATOR

Study Design

Study Type
observational
Observational Model
COHORT
Time Perspective
RETROSPECTIVE
Sponsor Type
OTHER
Responsible Party
PRINCIPAL INVESTIGATOR
PI Title
MD, Assistant Professor

Study Record Dates

First Submitted

June 3, 2026

First Posted

June 8, 2026

Study Start

May 1, 2026

Primary Completion

August 3, 2026

Study Completion

August 7, 2026

Last Updated

August 14, 2026

Record last verified: 2026-08

Locations