NCT07597499

Brief Summary

This pre-registered, pragmatic, three-arm (1:1:1) patient-level randomized controlled trial with mixed-effects analysis at the encounter level tests two questions in real high-risk multidisciplinary clinical encounters at the Waymark clinically integrated network across three U.S. states (Ohio, Washington, Virginia): (1) does adding ANCHOR - a clinical AI structural verification layer - to a Gemini 3.1 Pro-assisted supervising-physician workflow reduce the rate of clinically meaningful safety failures, compared with the same Gemini 3.1 Pro-assisted workflow without ANCHOR? (2) does the Gemini 3.1 Pro-assisted workflow itself reduce the same safety endpoint compared with unassisted standard care in which the supervising physician writes their own SOAP assessment/plan from a blank template? ANCHOR is a single-call structural verification layer combining a Logical Neural Network (Riegel et al. 2020) certificate, six specialist agents, and concept-decomposed output with PMID citation provenance. ANCHOR is physician-facing only and is used by supervising physicians, not by the multidisciplinary clinical team they oversee. The trial randomizes 240 patients 1:1:1 across the Waymark clinically integrated network over a 12-week active-enrolment window (80 per arm). Eligible patients are adults (age 18+) identified as high-risk by combined claims-based and clinical criteria. Eligible encounters span three integrated Waymark service modalities: high-risk primary care, specialty care coordination, and real-time telemedicine urgent care. The primary endpoint is a per-encounter binary composite: any of (a) failure to mention a do-not-miss diagnosis, (b) under-triage, (c) contraindicated medication recommendation, (d) failure to recommend escalation when clinically warranted; adjudicated by a blinded panel of 3 board-certified physicians with majority-of-three scoring. The primary contrast is Arm 3 (LLM+ANCHOR) versus Arm 2 (LLM with safety prompt), isolating ANCHOR's marginal contribution over a deployment-equivalent LLM safety stack. The pre-specified secondary contrast is Arm 2 versus Arm 1. The trial is sized to the operational ceiling of the Waymark integrated-network workflow across the three states (240 enrollees over 12 weeks). At realistic effect sizes derived from the retrospective evaluation, the trial is underpowered for definitive efficacy declaration on either pairwise contrast and is reported as an initial deployment-feasibility validation cohort with effect estimates and 95 percent confidence intervals; full power calculations are pre-registered in the Statistical Analysis Plan. Single-blind outcome adjudication: 3 adjudicators score only the supervising physician's final clinical decision, so all three arms produce adjudication packets in identical format and arm allocation is structurally invisible. Statisticians remain blinded until database lock. A full waiver of informed consent is requested per 45 CFR 46.116(f)(3) with a companion HIPAA waiver of authorization under 45 CFR 164.512(i)(2)(ii). The study is registered on the Open Science Framework prior to first enrollment and reported under CONSORT-AI 2020.

Trial Health

87
On Track

Trial Health Score

Automated assessment based on enrollment pace, timeline, and geographic reach

Enrollment
240

participants targeted

Target at P75+ for not_applicable

Timeline
Completed

Started May 2026

Shorter than P25 for not_applicable

Geographic Reach
1 country

1 active site

Status
completed

Health score is calculated from publicly available data and should be used for screening purposes only.

Trial Relationships

Click on a node to explore related trials.

Study Timeline

Key milestones and dates

First Submitted

Initial submission to the registry

May 8, 2026

Completed
7 days until next milestone

Study Start

First participant enrolled

May 15, 2026

Completed
4 days until next milestone

First Posted

Study publicly available on registry

May 19, 2026

Completed
1 month until next milestone

Primary Completion

Last participant's last visit for primary outcome

July 2, 2026

Completed
4 days until next milestone

Study Completion

Last participant's last visit for all outcomes

July 6, 2026

Completed
Last Updated

July 8, 2026

Status Verified

July 1, 2026

Enrollment Period

2 months

First QC Date

May 8, 2026

Last Update Submit

July 6, 2026

Conditions

Keywords

clinical decision supportlarge language modelartificial intelligenceverification layerlogical neural networkcare managementtelemedicinepatient safetyalgorithmic fairnessCONSORT-AI

Outcome Measures

Primary Outcomes (1)

  • Per-encounter clinical safety failure (adjudicated binary composite)

    Adjudicated binary composite of any of: (a) failure to mention a do-not-miss diagnosis appropriate for the presentation; (b) under-triage (routine/semi-urgent when emergent/urgent appropriate); (c) recommendation of a medication contraindicated by documented allergies/conditions/comorbidities; (d) failure to recommend escalation when clinically warranted. Adjudicated by a blinded panel of 3 board-certified physicians (Internal Medicine, Family Medicine, or Medicine-Pediatrics); final scoring by majority of three. Reported as proportion of encounters with composite safety failure.

    At the encounter (encounter-level outcome adjudicated within 4 weeks post-encounter)

Secondary Outcomes (4)

  • Appropriate triage escalation rate

    At the encounter

  • Time-to-physician-decision

    Within the encounter (real-time)

  • 30-day emergency-department visit rate (exploratory)

    30 days post-randomization

  • 30-day hospitalization rate (exploratory)

    30 days post-randomization

Other Outcomes (5)

  • Inappropriate over-escalation rate (safety)

    At the encounter

  • ANCHOR diagnostic-test characterization (Arm 3 only)

    At primary analysis (week 12)

  • Algorithmic fairness audit

    At primary analysis (week 12)

  • +2 more other outcomes

Study Arms (3)

Arm 1 - Unassisted standard care (control)

NO INTERVENTION

n=80. No LLM. No ANCHOR. The supervising physician opens a blank SOAP note template and writes their own assessment and plan from scratch based on the patient context and any prior chart review. Existing Waymark integrated-network multidisciplinary clinical-team support continues unchanged.

Arm 2 - Gemini 3.1 Pro with safety prompt (active comparator)

ACTIVE COMPARATOR

n=80. Gemini 3.1 Pro generates the care-management recommendation under a clinical-safety system prompt, content filters, and retrieval-augmented generation. The supervising physician reviews the LLM output directly without ANCHOR augmentation. This stack is operationally equivalent to LLM-assisted clinical-decision-support deployments already in routine use at major U.S. health systems. Decision support only; the supervising physician retains all clinical decision authority.

Behavioral: Gemini 3.1 Pro with Safety Prompt

Arm 3 - Gemini 3.1 Pro + ANCHOR (intervention)

EXPERIMENTAL

n=80. Same Gemini 3.1 Pro generation as Arm 2, with ANCHOR additionally applied: a single-call structural verification layer (Logical Neural Network certificate over a 3,206-rule clinical logic library; six specialist agents - drug interaction, lab interpretation, guideline compliance, citation verification, safety net, differential-diagnosis breadth; concept-decomposed output with PMID provenance) augments the LLM output. Supervising physician reviews the ANCHOR-augmented output. Decision support only; clinician retains all clinical decision authority.

Behavioral: ANCHOR Clinical AI Verification Layer (with Gemini 3.1 Pro)

Interventions

Gemini 3.1 Pro generates the care-management recommendation under a clinical-safety system prompt, content filters, and retrieval-augmented generation. Supervising physician reviews the LLM output directly without ANCHOR augmentation.

Arm 2 - Gemini 3.1 Pro with safety prompt (active comparator)

Same Gemini 3.1 Pro generation as Arm 2, with ANCHOR additionally applied: a single-call structural verification layer combining a Logical Neural Network safety certificate over a 3,206-rule clinical logic library, six concurrent specialist agents (drug interaction, lab interpretation, guideline compliance, citation verification, safety net, differential-diagnosis breadth), and a concept-decomposition module with PMID-traceable provenance. Decision support only; clinician retains all clinical decision authority.

Arm 3 - Gemini 3.1 Pro + ANCHOR (intervention)

Eligibility Criteria

Age18 Years+
Sexall
Healthy VolunteersNo
Age GroupsAdult (18-64), Older Adult (65+)

You may qualify if:

  • Age 18 years or older.
  • Attributed to a participating Waymark provider (academic medical center, community-hospital network, federally qualified health center, or independent physician practice in Ohio, Washington, or Virginia; full TIN-consolidated list deposited at the Open Science Framework).
  • Meets high-risk multidisciplinary criteria (combined claims-based and clinical: 2 or more emergency-department visits or 1 or more hospitalization in the prior 12 months, 5 or more active medications, 2 or more active specialist relationships, 2 or more chronic conditions, or claims-based equivalents).
  • Encounter occurs in one of the three Waymark service modalities: high-risk primary care, specialty care coordination, or real-time telemedicine urgent care.
  • English-language clinical documentation.
  • Encounter requires clinical reasoning (not administrative-only).

You may not qualify if:

  • Pediatric (age less than 18 years).
  • Hospice or palliative-care-exclusive care plan.
  • Active psychiatric crisis routed to crisis line.
  • Encounter is administrative only.
  • Pharmacy-only encounter that does not surface a clinical decision to the supervising physician.
  • Encounter where the supervising physician is the principal investigator.
  • Patient enrolled in a competing AI-safety study within the prior 90 days.

Contact the study team to confirm eligibility.

Sponsors & Collaborators

Study Sites (1)

Waymark

San Francisco, California, 94115, United States

Location

Study Officials

  • Sanjay Basu, MD, PhD

    Waymark

    PRINCIPAL INVESTIGATOR

Study Design

Study Type
interventional
Phase
not applicable
Allocation
RANDOMIZED
Masking
SINGLE
Who Masked
OUTCOMES ASSESSOR
Purpose
HEALTH SERVICES RESEARCH
Intervention Model
PARALLEL
Model Details: Three-arm parallel patient-level stratified permuted-block randomization, block size 6, stratified by site and acuity stratum. Once a patient is randomized at first eligible encounter, all subsequent encounters for that patient remain in the same arm.
Sponsor Type
INDUSTRY
Responsible Party
PRINCIPAL INVESTIGATOR
PI Title
Principal Investigator

Study Record Dates

First Submitted

May 8, 2026

First Posted

May 19, 2026

Study Start

May 15, 2026

Primary Completion

July 2, 2026

Study Completion

July 6, 2026

Last Updated

July 8, 2026

Record last verified: 2026-07

Data Sharing

IPD Sharing
Will not share

The trial dataset contains protected health information governed by HIPAA; no patient-level data, de-identified or otherwise, will be deposited in a public repository or shared with peer reviewers. Aggregate per-arm summary statistics, primary and secondary endpoint estimates, and adverse-event tables will be posted to ClinicalTrials.gov in accordance with the Final Rule (42 CFR Part 11).

Locations