Benchmarking Large Language Models Against Tumour Boards for Oncology Treatment Recommendations
BEACON
Benchmarking AI for Clinical Oncology decisioNmaking (BEACON): A Prospective, Multicentre, Blinded Evaluation of Frontier Large Language Models Against Multidisciplinary Tumour Board Recommendations in Oncology Treatment Planning
1 other identifier
observational
100
1 country
1
Brief Summary
BEACON (Benchmarking AI for Clinical Oncology decisioNmaking) is a prospective, multicentre, comparative, blinded, non-interventional benchmark evaluating the treatment recommendations of five frontier large language models (LLMs) against the recommendations of multidisciplinary tumour boards (RCP) in oncology treatment planning. One hundred standardised synthetic cases (20 per localisation, across breast, lung, urological, digestive and gynaecological cancers) are submitted as identical structured input to two independent tumour boards per localisation and to five frontier LLMs. Each recommendation - human or model - is decomposed into five predefined decision domains (intent, surgery, radiotherapy, systemic therapy, work-up and biomarkers) and scored 0/1/2 for concordance against a two-tier reference: the consensus of the two tumour boards, complemented by an a priori locked guideline matrix (ESMO, NCCN). The primary endpoint is domain-level concordance between LLM and RCP consensus, expressed as a linearly weighted Cohen's kappa. A co-primary safety endpoint captures the proportion of recommendations carrying serious harm potential, because concordance alone can conceal dangerous errors. Because expert boards may disagree with one another on identical cases, model performance is always interpreted against the human consensus. BEACON is designed as reusable, openly licensed, pre-registered infrastructure: all synthetic cases, evaluation rubrics, the locked guideline matrix, scoring algorithms and verbatim prompts are released for full reproducibility.
Trial Health
Trial Health Score
Automated assessment based on enrollment pace, timeline, and geographic reach
participants targeted
Target at P50-P75 for all trials
Started May 2026
Shorter than P25 for all trials
1 active site
Health score is calculated from publicly available data and should be used for screening purposes only.
Trial Relationships
Click on a node to explore related trials.
Study Timeline
Key milestones and dates
Study Start
First participant enrolled
May 1, 2026
CompletedFirst Submitted
Initial submission to the registry
July 28, 2026
CompletedFirst Posted
Study publicly available on registry
July 31, 2026
CompletedPrimary Completion
Last participant's last visit for primary outcome
October 1, 2026
ExpectedStudy Completion
Last participant's last visit for all outcomes
October 1, 2026
July 31, 2026
July 1, 2026
5 months
July 28, 2026
July 28, 2026
Conditions
Keywords
Outcome Measures
Primary Outcomes (1)
Domain-level performance between LLM recommendations and the locked guidelines.
For each recommendation domain and each LLM, proportion of LLM recommendation concordant with locked guidelines
Assessed once at central scoring, after data collection (~October 2026)
Secondary Outcomes (7)
Proportion of recommendations carrying serious harm potential ( LLM and tumour boards)
Up to October 2026
Domain-level recommendation concordance between LLM and tumour-boards
Up to October 2026
Inter-tumour board domain-level recommendation concordance
Up to October 2026
Equipoise rate
Up to October 2026
Completeness
Up to October 2026
- +2 more secondary outcomes
Study Arms (5)
Breast cancers
Lung cancers
Urological cancers
Digestive cancers
Gynaecological cancers
Interventions
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case. Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6, Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
Eligibility Criteria
100 synthetic oncology treatment-planning cases (20 per localisation) across five localisations: breast, lung, urological (prostate, bladder / upper-tract urothelial, kidney), digestive and gynaecological. Each case is a structured JSON input specifying UICC 8th-edition stage, biomarkers, ECOG performance status, comorbidities and a standardised clinical question. No human participants, no patient data and no identifiable individuals. Recommendations are produced by two independent tumour boards per localisation and by five frontier LLMs (queried May 2026).
You may qualify if:
- Synthetic oncology case within one of the five predefined localisations (breast, lung, urological, digestive, gynaecological).
- Complete structured schema: UICC 8th-edition stage, biomarkers, ECOG performance status, comorbidities and a standardised clinical question.
- A clinically answerable treatment-planning question that is mappable to the locked guideline matrix.
You may not qualify if:
- Case outside the five predefined localisations.
- Incomplete, internally inconsistent or ambiguous schema.
- Duplicate or near-duplicate of an existing case in the set.
- Question not resolvable by current guidelines.
Contact the study team to confirm eligibility.
Sponsors & Collaborators
Study Sites (1)
Hopital Européen Georges Pompidou
Paris, France
MeSH Terms
Conditions
Condition Hierarchy (Ancestors)
Central Study Contacts
Study Design
- Study Type
- observational
- Observational Model
- COHORT
- Time Perspective
- PROSPECTIVE
- Sponsor Type
- OTHER
- Responsible Party
- SPONSOR
Study Record Dates
First Submitted
July 28, 2026
First Posted
July 31, 2026
Study Start
May 1, 2026
Primary Completion (Estimated)
October 1, 2026
Study Completion (Estimated)
October 1, 2026
Last Updated
July 31, 2026
Record last verified: 2026-07