Comparative Analysis of Diagnostic Accuracy and Case Difficulty Assessment of Three Large Language Models
1 other identifier
observational
342
0 countries
N/A
Brief Summary
This prospective, blinded diagnostic accuracy study aims to compare the performance of three large language models-ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7-in endodontic diagnosis and case difficulty assessment. The models will be evaluated against expert consensus as the reference standard using standardized clinical data and periapical radiographs. Diagnostic accuracy, sensitivity, specificity, and agreement with expert consensus will be assessed to determine the potential of LLMs as clinical decision-support tools in endodontics.
Trial Health
Trial Health Score
Automated assessment based on enrollment pace, timeline, and geographic reach
participants targeted
Target at P75+ for all trials
Started Sep 2026
Shorter than P25 for all trials
Health score is calculated from publicly available data and should be used for screening purposes only.
Trial Relationships
Click on a node to explore related trials.
Study Timeline
Key milestones and dates
First Submitted
Initial submission to the registry
July 10, 2026
CompletedFirst Posted
Study publicly available on registry
July 29, 2026
CompletedStudy Start
First participant enrolled
September 1, 2026
ExpectedPrimary Completion
Last participant's last visit for primary outcome
December 1, 2026
Study Completion
Last participant's last visit for all outcomes
January 1, 2027
July 29, 2026
July 1, 2026
3 months
July 10, 2026
July 23, 2026
Conditions
Outcome Measures
Primary Outcomes (1)
Diagnostic accuracy of ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7 for pulpal and periapical diagnosis
Diagnostic performance of each large language model compared with the expert consensus reference standard. Accuracy, sensitivity, specificity, and agreement (Cohen's kappa) will be calculated according to the American Association of Endodontists (AAE) diagnostic criteria
At baseline
Secondary Outcomes (1)
Accuracy of ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7 in endodontic case difficulty assessment
At baseline
Interventions
Large language model used to analyze standardized clinical information and periapical radiographs to provide pulpal and periapical diagnosis and endodontic case difficulty assessment according to AAE criteria
Eligibility Criteria
Systemically healthy adult patients (≥16 years) requiring primary endodontic treatment or retreatment at the Endodontic Department, Faculty of Dentistry, Cairo University, who provide informed consent.
You may qualify if:
- Age above 16 years old.
- Systemically healthy patient (ASA I or II).
- Requiring endodontic treatment or retreatment
- Patient's acceptance to participate in the study
You may not qualify if:
- Medically compromised patients.
- Pregnant women.
- Traumatic dental injuries
- Low quality periapical radiograph
Contact the study team to confirm eligibility.
Sponsors & Collaborators
- Cairo Universitylead
Central Study Contacts
Study Design
- Study Type
- observational
- Observational Model
- OTHER
- Time Perspective
- PROSPECTIVE
- Sponsor Type
- OTHER
- Responsible Party
- PRINCIPAL INVESTIGATOR
- PI Title
- PhD Candidate
Study Record Dates
First Submitted
July 10, 2026
First Posted
July 29, 2026
Study Start (Estimated)
September 1, 2026
Primary Completion (Estimated)
December 1, 2026
Study Completion (Estimated)
January 1, 2027
Last Updated
July 29, 2026
Record last verified: 2026-07