Performance of an OCR-Prompt-LLM Integrated Workflow for Extracting Multi-dimensional Clinical Data in Ischemic Heart Disease
OPAL-CAD
1 other identifier
observational
308
1 country
1
Brief Summary
This research aims to evaluate a comprehensive AI-driven workflow for both clinical data extraction and diagnostic classification in coronary artery disease (CAD). Leveraging OCR and Large Language Models (LLMs), the system is designed to extract ten key clinical parameters (such as LVEF and lab results) and provide diagnostic subtypes (UA, STEMI, NSTEMI, CCS) directly from unstructured inpatient records. A man-machine comparative trial will be conducted using a test set of 308 patients, where the performance of the LLM-based workflow will be benchmarked against the average diagnostic accuracy and processing time of seven clinical physicians. The findings will provide evidence for the feasibility of using LLMs to enhance clinical data structuring and diagnostic efficiency in cardiology.
Trial Health
Trial Health Score
Automated assessment based on enrollment pace, timeline, and geographic reach
participants targeted
Target at P75+ for all trials
Started Feb 2026
1 active site
Health score is calculated from publicly available data and should be used for screening purposes only.
Trial Relationships
Click on a node to explore related trials.
Study Timeline
Key milestones and dates
Study Start
First participant enrolled
February 23, 2026
CompletedPrimary Completion
Last participant's last visit for primary outcome
March 1, 2026
CompletedStudy Completion
Last participant's last visit for all outcomes
March 2, 2026
CompletedFirst Submitted
Initial submission to the registry
March 24, 2026
CompletedFirst Posted
Study publicly available on registry
March 30, 2026
CompletedMarch 30, 2026
February 1, 2026
6 days
March 24, 2026
March 24, 2026
Conditions
Outcome Measures
Primary Outcomes (1)
Overall Diagnostic and Extraction Accuracy Rate
To calculate the overall accuracy rate of the LLM-based workflow across 308 cases (including the pilot set, internal validation cohort, and external validation cohort) for 10 clinical indicators (e.g., LVEF, blood glucose, etc.) and 4 diagnostic subtypes of coronary artery disease. Accuracy is defined as the proportion of cases where the LLM's extraction or diagnostic results are perfectly consistent with the 'Gold Standard' established by human clinical experts.
Through study completion, an average of 3 months.
Study Arms (3)
Test Cohort
This group consists of 50 patient records from the AIM-CHD Study at Fuwai Hospital. These data are specifically utilized for refining OCR processing and optimizing Prompt Engineering for the LLM-based workflow.
Internal Validation Cohort
This cohort includes 188 clinical cases sourced from the SMART-CHD Study at Fuwai Hospital. These records serve as the primary internal benchmark to evaluate the diagnostic and extraction accuracy of the LLM workflow against the established ground truth.
External Validation Cohort
This cohort comprises 70 patient records collected from 8 independent sub-centers (excluding Fuwai Hospital) to assess the generalizability and robustness of the model across diverse clinical environments and different medical record formats.
Interventions
The intervention is an automated clinical data management system integrating Optical Character Recognition (OCR), optimized Prompt Engineering, and Large Language Models (LLMs). The workflow processes unstructured inpatient records to extract 10 key clinical indicators (e.g., LVEF, CAD subtypes, medications) and classifies the patient into specific coronary artery disease categories (UA, STEMI, NSTEMI, CCS)
Standard manual process where experienced clinical physicians collect and interpret patient information from medical records. This serves as the human benchmark for comparing diagnostic accuracy and operational efficiency.
Eligibility Criteria
he study population consists of 308 patients diagnosed with various subtypes of coronary artery disease (CAD). The cohort is derived from two major clinical studies: the AIM-CHD study (for pilot testing and prompt optimization) and the SMART-CHD study (for internal validation), both conducted at Fuwai Hospital. Additionally, an external validation cohort is included, comprising patients from 8 independent clinical sub-centers across China to ensure geographical and institutional diversity. The population covers a spectrum of CAD presentations, including Unstable Angina (UA), STEMI, NSTEMI, and Chronic Coronary Syndrome (CCS), providing a robust dataset for evaluating AI-driven diagnostic and data extraction performance.
You may qualify if:
- Patients aged 18 years and older.
- Clinical records of patients who were previously enrolled in the AIM-CHD (for the pilot/prompt optimization set) or SMART-CHD (for the internal validation cohort) studies.
- Patients diagnosed with, or suspected of having, coronary artery disease (CAD), including subtypes: Unstable Angina (UA), STEMI, NSTEMI, and Chronic Coronary Syndrome (CCS).
You may not qualify if:
- Clinical records with severe data fragmentation or missing more than 50% of the key clinical indicators.
- Handwritten medical records or low-quality scans that are illegible for Optical Character Recognition (OCR) processing.
- Duplicate records or records with conflicting "Gold Standard" labels that cannot be reconciled by the expert committee.
Contact the study team to confirm eligibility.
Sponsors & Collaborators
Study Sites (1)
Fuwai Hospital
Beijing, 100037, China
MeSH Terms
Conditions
Condition Hierarchy (Ancestors)
Study Design
- Study Type
- observational
- Observational Model
- COHORT
- Time Perspective
- RETROSPECTIVE
- Sponsor Type
- OTHER GOV
- Responsible Party
- SPONSOR
Study Record Dates
First Submitted
March 24, 2026
First Posted
March 30, 2026
Study Start
February 23, 2026
Primary Completion
March 1, 2026
Study Completion
March 2, 2026
Last Updated
March 30, 2026
Record last verified: 2026-02
Data Sharing
- IPD Sharing
- Will not share
To protect patient privacy and comply with the data management policies of the participating institutions (Fuwai Hospital and sub-centers), individual participant data will not be made publicly available. However, aggregated study results and statistical analyses will be included in the final publication.