NCT07677202

Brief Summary

This study evaluates whether large language models (LLMs) can reliably assess written clinical-reasoning case examinations completed by undergraduate physiotherapy students, compared with faculty assessment. In the course "Specific Methods in Physiotherapy" (third year of the Physiotherapy Degree), students solve complex clinical cases that require clinical reasoning, technical knowledge, and therapeutic decision-making. These cases are traditionally graded by faculty, a time-consuming process that may show inter-rater variability. A set of de-identified student case examinations will be assessed using the rubric currently applied in the course, which covers clarity and structure of clinical reasoning, integration of the biopsychosocial model (ICF and APTA frameworks), accuracy in identifying pain mechanisms, coherence between diagnosis, hypotheses, and treatment, originality and depth of analysis, and professional writing. Each examination will be scored independently by three LLMs (for example, Claude, ChatGPT, and Gemini), each receiving an identical standardized prompt that embeds the same rubric, and by faculty serving as the reference standard. To avoid overloading faculty, full double human grading may not be feasible; the human reference will therefore consist of expert faculty grading by one independent rater or, when resources allow, two independent raters. In contrast, paired assessment is fully implemented across the AI models: each examination is scored by several LLMs, and each model is queried in duplicate, allowing the study to estimate agreement between models and the test-retest stability of each model. The primary aim is to quantify agreement between LLM-generated scores and the faculty reference score. Secondary aims include agreement among the LLMs, test-retest reliability of each model, criterion-level agreement, the quality and usefulness of the qualitative feedback generated, the time and cost associated with each approach, and students' perceptions of the usefulness of human versus AI feedback. The findings will clarify the strengths and limitations of LLMs as supportive tools for formative assessment in health-professions education and will inform criteria for their responsible and effective use. No LLM output will affect students' official grades, which remain the sole responsibility of faculty.

Trial Health

63
Monitor

Trial Health Score

Automated assessment based on enrollment pace, timeline, and geographic reach

Enrollment
65

participants targeted

Target at P25-P50 for all trials

Timeline
0mo left

Started Aug 2026

Geographic Reach
1 country

1 active site

Status
not yet recruiting

Health score is calculated from publicly available data and should be used for screening purposes only.

Trial Relationships

Click on a node to explore related trials.

Study Timeline

Key milestones and dates

Study Progress8%
Aug 2026Aug 2026

First Submitted

Initial submission to the registry

June 24, 2026

Completed
6 days until next milestone

First Posted

Study publicly available on registry

June 30, 2026

Completed
1 month until next milestone

Study Start

First participant enrolled

August 1, 2026

Completed
9 days until next milestone

Primary Completion

Last participant's last visit for primary outcome

August 10, 2026

Expected
Same day until next milestone

Study Completion

Last participant's last visit for all outcomes

August 10, 2026

Last Updated

June 30, 2026

Status Verified

June 1, 2026

Enrollment Period

9 days

First QC Date

June 24, 2026

Last Update Submit

June 24, 2026

Conditions

Keywords

Clinical CompetenceEducational, Medical

Outcome Measures

Primary Outcomes (1)

  • Agreement between LLM global scores and the faculty reference global score

    Agreement between the global examination score generated by each large language model (LLM) and the faculty reference global score, computed for the same anonymized examinations. Agreement is quantified with the intraclass correlation coefficient (ICC), two-way random-effects model, absolute-agreement definition, single- and average-measures forms \[ICC(2,1) and ICC(2,k)\], with 95% confidence intervals. Systematic bias is examined with Bland-Altman analysis (mean difference and 95% limits of agreement). ICC is interpreted as poor (\<0.50), moderate (0.50-0.75), good (0.75-0.90), or excellent (\>0.90). Pre-specified target: ICC \>= 0.75.

    Single cross-sectional assessment during the data-collection period (approximately 2 months)

Secondary Outcomes (5)

  • Criterion-level agreement between LLM and faculty scores

    Single cross-sectional assessment during the data-collection period (approximately 2 months)

  • Intra-model test-retest reliability of each large language model

    Single cross-sectional assessment during the data-collection period (approximately 2 months)

  • Quality and coverage of the qualitative feedback

    Assessed after completion of all evaluations, during the analysis period (approximately 3 months)

  • Mean evaluation time per examination: faculty versus LLM

    Single cross-sectional assessment during the data-collection period (approximately 2 months)

  • Cost per evaluation: faculty versus LLM

    Single cross-sectional assessment during the data-collection period (approximately 2 months)

Study Arms (1)

Anonymized physiotherapy clinical case examinations

Single cohort consisting of de-identified written clinical-reasoning case examinations produced by undergraduate physiotherapy students in the course "Specific Methods in Physiotherapy." Each examination is assessed independently, using the same predefined rubric, by faculty (reference standard) and by three large language models (LLMs), with each model queried in duplicate to assess test-retest reliability. The examination is the unit of analysis; no participant follow-up is performed.

Diagnostic Test: LLM-based assessmentDiagnostic Test: Faculty assessment (reference standard)

Interventions

LLM-based assessmentDIAGNOSTIC_TEST

Assessment of each anonymized examination by three large language models (for example, Claude, ChatGPT, and Gemini, in the versions available during data collection). Each model receives an identical standardized prompt embedding the study rubric and returns a score per criterion, a global score, and structured qualitative feedback. Each model is queried in duplicate in independent sessions under fixed generation parameters to estimate intra-model (test-retest) reliability, and outputs are compared across models to estimate inter-model agreement.

Anonymized physiotherapy clinical case examinations

Assessment of the same anonymized examinations by faculty with expertise in the course, applying the identical rubric, serving as the reference standard. In the preferred scenario, two faculty members score each examination independently (paired human correction); if faculty workload precludes this, a single expert faculty rating, or the official course grade already assigned, is used as the reference. Faculty and LLM raters are blinded to one another's scores.

Anonymized physiotherapy clinical case examinations

Eligibility Criteria

Age18 Years+
Sexall
Healthy VolunteersYes
Age GroupsAdult (18-64), Older Adult (65+)
Sampling MethodNon-Probability Sample
Study Population

The study population comprises undergraduate students enrolled in the course "Specific Methods in Physiotherapy" (third year of the Physiotherapy Degree) during the study period, approximately 60 to 80 students. As part of the course, each student produces a written clinical-reasoning case examination. The de-identified examinations from consenting students constitute the units of analysis and are assessed independently, using the same predefined rubric, by faculty (reference standard) and by three large language models. No clinical intervention is applied and participation does not affect students' official grades, which are determined exclusively by faculty.

You may qualify if:

  • Students officially enrolled in the course "Specific Methods in Physiotherapy" (third year of the Physiotherapy Degree) during the study period.
  • Submission of a completed written clinical-reasoning case examination as part of the course.
  • Provision of informed consent for the anonymized examination to be used for educational-research purposes.

You may not qualify if:

  • Refusal to provide, or withdrawal of, informed consent.
  • Blank, incomplete, or non-evaluable examinations (e.g., no developed written response).
  • Examinations that cannot be reliably de-identified prior to assessment.

Contact the study team to confirm eligibility.

Sponsors & Collaborators

Study Sites (1)

Centro Superior de Estudios Universitarios La Salle

Madrid, Madrid, 28023, Spain

Location

Central Study Contacts

Alfredo Lerín Calvo, MSc

CONTACT

Study Design

Study Type
observational
Observational Model
COHORT
Time Perspective
CROSS SECTIONAL
Target Duration
1 Day
Sponsor Type
OTHER
Responsible Party
PRINCIPAL INVESTIGATOR
PI Title
Mr.

Study Record Dates

First Submitted

June 24, 2026

First Posted

June 30, 2026

Study Start

August 1, 2026

Primary Completion (Estimated)

August 10, 2026

Study Completion (Estimated)

August 10, 2026

Last Updated

June 30, 2026

Record last verified: 2026-06

Data Sharing

IPD Sharing
Will share

De-identified individual participant data (anonymized examination scores per rubric criterion and global score from all human and LLM raters, including duplicate LLM runs) and the corresponding data dictionary will be shared. The standardized LLM prompt, the scoring rubric, and the statistical analysis code will also be made available. Data will be deposited in the Zenodo open-access repository and assigned a permanent DOI. No directly identifying information will be shared; all examinations are anonymized prior to assessment.

Locations