Large Language Models Versus Human Examiners for Grading Physiotherapy Clinical Cases
PACE-AI
Agreement Between Large Language Models and Faculty Assessment in the Evaluation of Clinical Reasoning Case Examinations in Undergraduate Physiotherapy Education: A Comparative Reliability Study
1 other identifier
observational
65
1 country
1
Brief Summary
This study evaluates whether large language models (LLMs) can reliably assess written clinical-reasoning case examinations completed by undergraduate physiotherapy students, compared with faculty assessment. In the course "Specific Methods in Physiotherapy" (third year of the Physiotherapy Degree), students solve complex clinical cases that require clinical reasoning, technical knowledge, and therapeutic decision-making. These cases are traditionally graded by faculty, a time-consuming process that may show inter-rater variability. A set of de-identified student case examinations will be assessed using the rubric currently applied in the course, which covers clarity and structure of clinical reasoning, integration of the biopsychosocial model (ICF and APTA frameworks), accuracy in identifying pain mechanisms, coherence between diagnosis, hypotheses, and treatment, originality and depth of analysis, and professional writing. Each examination will be scored independently by three LLMs (for example, Claude, ChatGPT, and Gemini), each receiving an identical standardized prompt that embeds the same rubric, and by faculty serving as the reference standard. To avoid overloading faculty, full double human grading may not be feasible; the human reference will therefore consist of expert faculty grading by one independent rater or, when resources allow, two independent raters. In contrast, paired assessment is fully implemented across the AI models: each examination is scored by several LLMs, and each model is queried in duplicate, allowing the study to estimate agreement between models and the test-retest stability of each model. The primary aim is to quantify agreement between LLM-generated scores and the faculty reference score. Secondary aims include agreement among the LLMs, test-retest reliability of each model, criterion-level agreement, the quality and usefulness of the qualitative feedback generated, the time and cost associated with each approach, and students' perceptions of the usefulness of human versus AI feedback. The findings will clarify the strengths and limitations of LLMs as supportive tools for formative assessment in health-professions education and will inform criteria for their responsible and effective use. No LLM output will affect students' official grades, which remain the sole responsibility of faculty.
Trial Health
Trial Health Score
Automated assessment based on enrollment pace, timeline, and geographic reach
participants targeted
Target at P25-P50 for all trials
Started Aug 2026
1 active site
Health score is calculated from publicly available data and should be used for screening purposes only.
Trial Relationships
Click on a node to explore related trials.
Study Timeline
Key milestones and dates
First Submitted
Initial submission to the registry
June 24, 2026
CompletedFirst Posted
Study publicly available on registry
June 30, 2026
CompletedStudy Start
First participant enrolled
August 1, 2026
CompletedPrimary Completion
Last participant's last visit for primary outcome
August 10, 2026
ExpectedStudy Completion
Last participant's last visit for all outcomes
August 10, 2026
June 30, 2026
June 1, 2026
9 days
June 24, 2026
June 24, 2026
Conditions
Keywords
Outcome Measures
Primary Outcomes (1)
Agreement between LLM global scores and the faculty reference global score
Agreement between the global examination score generated by each large language model (LLM) and the faculty reference global score, computed for the same anonymized examinations. Agreement is quantified with the intraclass correlation coefficient (ICC), two-way random-effects model, absolute-agreement definition, single- and average-measures forms \[ICC(2,1) and ICC(2,k)\], with 95% confidence intervals. Systematic bias is examined with Bland-Altman analysis (mean difference and 95% limits of agreement). ICC is interpreted as poor (\<0.50), moderate (0.50-0.75), good (0.75-0.90), or excellent (\>0.90). Pre-specified target: ICC \>= 0.75.
Single cross-sectional assessment during the data-collection period (approximately 2 months)
Secondary Outcomes (5)
Criterion-level agreement between LLM and faculty scores
Single cross-sectional assessment during the data-collection period (approximately 2 months)
Intra-model test-retest reliability of each large language model
Single cross-sectional assessment during the data-collection period (approximately 2 months)
Quality and coverage of the qualitative feedback
Assessed after completion of all evaluations, during the analysis period (approximately 3 months)
Mean evaluation time per examination: faculty versus LLM
Single cross-sectional assessment during the data-collection period (approximately 2 months)
Cost per evaluation: faculty versus LLM
Single cross-sectional assessment during the data-collection period (approximately 2 months)
Study Arms (1)
Anonymized physiotherapy clinical case examinations
Single cohort consisting of de-identified written clinical-reasoning case examinations produced by undergraduate physiotherapy students in the course "Specific Methods in Physiotherapy." Each examination is assessed independently, using the same predefined rubric, by faculty (reference standard) and by three large language models (LLMs), with each model queried in duplicate to assess test-retest reliability. The examination is the unit of analysis; no participant follow-up is performed.
Interventions
Assessment of each anonymized examination by three large language models (for example, Claude, ChatGPT, and Gemini, in the versions available during data collection). Each model receives an identical standardized prompt embedding the study rubric and returns a score per criterion, a global score, and structured qualitative feedback. Each model is queried in duplicate in independent sessions under fixed generation parameters to estimate intra-model (test-retest) reliability, and outputs are compared across models to estimate inter-model agreement.
Assessment of the same anonymized examinations by faculty with expertise in the course, applying the identical rubric, serving as the reference standard. In the preferred scenario, two faculty members score each examination independently (paired human correction); if faculty workload precludes this, a single expert faculty rating, or the official course grade already assigned, is used as the reference. Faculty and LLM raters are blinded to one another's scores.
Eligibility Criteria
The study population comprises undergraduate students enrolled in the course "Specific Methods in Physiotherapy" (third year of the Physiotherapy Degree) during the study period, approximately 60 to 80 students. As part of the course, each student produces a written clinical-reasoning case examination. The de-identified examinations from consenting students constitute the units of analysis and are assessed independently, using the same predefined rubric, by faculty (reference standard) and by three large language models. No clinical intervention is applied and participation does not affect students' official grades, which are determined exclusively by faculty.
You may qualify if:
- Students officially enrolled in the course "Specific Methods in Physiotherapy" (third year of the Physiotherapy Degree) during the study period.
- Submission of a completed written clinical-reasoning case examination as part of the course.
- Provision of informed consent for the anonymized examination to be used for educational-research purposes.
You may not qualify if:
- Refusal to provide, or withdrawal of, informed consent.
- Blank, incomplete, or non-evaluable examinations (e.g., no developed written response).
- Examinations that cannot be reliably de-identified prior to assessment.
Contact the study team to confirm eligibility.
Sponsors & Collaborators
- Neuron, Spainlead
- Centro Universitario La Sallecollaborator
Study Sites (1)
Centro Superior de Estudios Universitarios La Salle
Madrid, Madrid, 28023, Spain
Central Study Contacts
Study Design
- Study Type
- observational
- Observational Model
- COHORT
- Time Perspective
- CROSS SECTIONAL
- Target Duration
- 1 Day
- Sponsor Type
- OTHER
- Responsible Party
- PRINCIPAL INVESTIGATOR
- PI Title
- Mr.
Study Record Dates
First Submitted
June 24, 2026
First Posted
June 30, 2026
Study Start
August 1, 2026
Primary Completion (Estimated)
August 10, 2026
Study Completion (Estimated)
August 10, 2026
Last Updated
June 30, 2026
Record last verified: 2026-06
Data Sharing
- IPD Sharing
- Will share
De-identified individual participant data (anonymized examination scores per rubric criterion and global score from all human and LLM raters, including duplicate LLM runs) and the corresponding data dictionary will be shared. The standardized LLM prompt, the scoring rubric, and the statistical analysis code will also be made available. Data will be deposited in the Zenodo open-access repository and assigned a permanent DOI. No directly identifying information will be shared; all examinations are anonymized prior to assessment.