Skip to main content
medRxiv
  • Home
  • About
  • Submit
  • ALERTS / RSS
Advanced Search

Evaluation of a Generative Medical Event Foundation Model for Predicting Post-Discharge Trajectories in Emergency Department Abdominal Pain

View ORCID ProfileKent A. McCann, View ORCID ProfileDonald S. Wright, View ORCID ProfileMark S. Iscoe, View ORCID ProfileEdward R. Melnick, View ORCID ProfileLucila Ohno-Machado, View ORCID ProfileDaniella Meeker, View ORCID ProfileArjun K. Venkatesh, View ORCID ProfileRohit B. Sangal, View ORCID ProfileAndrew J. Loza
doi: https://doi.org/10.64898/2026.05.18.26353199
Kent A. McCann
1Department of Emergency Medicine, Yale School of Medicine, New Haven, CT
2Veterans Affairs Connecticut Healthcare System, West Haven, CT
MD
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Kent A. McCann
Donald S. Wright
1Department of Emergency Medicine, Yale School of Medicine, New Haven, CT
2Veterans Affairs Connecticut Healthcare System, West Haven, CT
3Department of Biomedical Informatics and Data Science, Yale School of Medicine, New Haven, CT
MD MHS
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Donald S. Wright
Mark S. Iscoe
1Department of Emergency Medicine, Yale School of Medicine, New Haven, CT
3Department of Biomedical Informatics and Data Science, Yale School of Medicine, New Haven, CT
MD MHS
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Mark S. Iscoe
Edward R. Melnick
1Department of Emergency Medicine, Yale School of Medicine, New Haven, CT
2Veterans Affairs Connecticut Healthcare System, West Haven, CT
MD MHS
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Edward R. Melnick
Lucila Ohno-Machado
3Department of Biomedical Informatics and Data Science, Yale School of Medicine, New Haven, CT
MD PhD
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Lucila Ohno-Machado
Daniella Meeker
3Department of Biomedical Informatics and Data Science, Yale School of Medicine, New Haven, CT
4Yale New Haven Health System, New Haven, CT
PhD
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Daniella Meeker
Arjun K. Venkatesh
1Department of Emergency Medicine, Yale School of Medicine, New Haven, CT
MD MBA MHS
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Arjun K. Venkatesh
Rohit B. Sangal
1Department of Emergency Medicine, Yale School of Medicine, New Haven, CT
MD MBA
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Rohit B. Sangal
Andrew J. Loza
2Veterans Affairs Connecticut Healthcare System, West Haven, CT
3Department of Biomedical Informatics and Data Science, Yale School of Medicine, New Haven, CT
4Yale New Haven Health System, New Haven, CT
MD PhD
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Andrew J. Loza
  • For correspondence: andrew.loza{at}yale.edu
  • Abstract
  • Full Text
  • Info/History
  • Metrics
  • Supplementary material
  • Data/Code
  • Preview PDF
Loading

Abstract

Objective Emergency department (ED) risk models predict whether a patient will revisit within a fixed window, but not what the visit will entail. We evaluated whether a medical event foundation model predicts the timing, sequence, and severity of downstream care after ED discharge for abdominal pain.

Materials and Methods Retrospective cohort study in the de-identified Epic Cosmos network. From 150,030 eligible adults, we analyzed a random sample of 3,000, generating 50 simulated 30-day trajectories per patient with Curiosity, a generative medical event foundation model, and compared these with gradient boosted tree baselines. Endpoints were discrimination for three outcome classes (any revisit, discharge-revisit, admit-revisit) at three horizons, 36-class trajectory accuracy, transition-level calibration, and subgroup performance.

Results Curiosity’s 30-day admit-revisit AUROC was 0.83 (95% CI 0.79-0.87) versus 0.76 (0.71-0.81) for XGBoost (P=.002). Differences were significant at five of nine outcome-horizon pairs but not for discharge-revisit at any horizon. Its most likely trajectory (of 36) matched the observed in 45.9% versus 43.0% (P<.001); median absolute calibration error across 45 transitions was 1.30 percentage points. Discrimination and calibration were poorer for Hispanic or Latino, self-pay, and other-race patients.

Discussion Without task-specific training, one pretrained foundation model matched or exceeded a tuned supervised baseline, with significant advantages on the primary endpoint and trajectory prediction but none for discharge-revisit. Subgroup gaps must be resolved before clinical use.

Conclusion A medical event foundation model can represent post-ED-discharge revisit risk as a calibrated trajectory rather than a single probability; prospective evaluation and remediation of subgroup performance gaps remain necessary.

Competing Interest Statement

The authors have declared no competing interest.

Author Declarations

I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.

Yes

The details of the IRB/oversight body that provided approval or exemption for the research described are given below:

This study was exempted from human subjects review by the Yale institutional review board.

I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.

Yes

I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).

Yes

I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.

Yes

Footnotes

  • ↵* Co-Last Author

  • Baseline Strengthening The XGBoost baseline was retrained on 100,000 patients, increased from 10,000, with learning curves reported for each outcome. Each of four XGBoost models, three per-horizon models and the 36-class trajectory model, was independently hyperparameter-tuned via Optuna at the 100,000-patient plateau. Primary Endpoint For 30-day admit-revisit, XGBoost AUROC increased from 0.70 to 0.76, with a 95 percent confidence interval of 0.71 to 0.81. Curiosity remained 0.83. The comparison remains significant, with a P value of .002 rather than less than .001. The AUROC gap was approximately halved, from 0.13 to 0.07. The statement that Curiosity was superior at all nine outcome-by-horizon pairs was revised to numerically higher at all nine and statistically significantly higher at five: all three any-revisit horizons and admit-revisit at 7 and 30 days. The three DC-revisit horizons and admit-revisit at 72 hours no longer reached significance. Trajectory-Level Metrics XGBoost single-best trajectory match increased from 41.0 percent to 43.0 percent and remained significant. Top-3 match increased from 63.9 percent to 65.5 percent and remained significant. Top-5 increased from 77.0 percent to 79.2 percent and was no longer significant, with a P value of .08. XGBoost median edit distance improved from 1.40 to 1.385. Curiosity produced the closer trajectory in 60.6 percent of patients, decreased from 63.9 percent. Paired difference and Wilcoxon results were updated. Figure 3 was updated. XGBoost cohort-referenced relative risks increased from 4.72 and 2.66 to 6.02 and 2.72. Group overlap decreased from 23.3 percent to 18.7 percent, and cross-group DC-revisit relative risk decreased from 1.55 to 1.33. The tuned baseline better identifies high-risk patients but still differentiates revisit subtypes less effectively than Curiosity. Calibration and Subgroup Analysis Calibration slope, calibration-in-the-large, and integrated calibration index are now reported for Curiosity across all outcome-by-horizon combinations. Subgroup analyses of 30-day revisit risk are now reported for all eligible subgroups meeting the Cosmos cell-size requirement of more than 11 events. Magnitude Claims The admit-revisit AUC-PR advantage was revised from approximately 2.8-fold to 1.7-fold after XGBoost AUC-PR increased from 0.13 to 0.216. Text Additions A Baseline Learning Curve section was added. XGBoost discrimination plateaued by approximately 60,000 patients and improved to 0.76 after tuning, compared with 0.83 for Curiosity. The baseline limitation now notes both training to the learning-curve plateau and per-endpoint hyperparameter tuning.

Data availability

Deidentified electronic health record data underlying these analyses reside in Epic Cosmos and are available to investigators with approved access through Epic Systems. The Curiosity model used in this study is described in Waxler et al and is hosted within the Epic Cosmos environment. Analytic code resides within Epic Cosmos and is available to investigators with approved Cosmos access.

Funder Information Declared

ARIA Foundation
National Center for Advancing Translational Sciences, https://ror.org/04pw6fb54, KL2 TR001862, UL1 TR001863
United States National Library of Medicine, https://ror.org/0060t0j89, 5T15LM007056-40
Hartwell Foundation, https://ror.org/038cgyc59
Copyright 
The copyright holder for this preprint is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. It is made available under a CC-BY-NC-ND 4.0 International license.
Back to top
PreviousNext
Posted August 11, 2026.
Download PDF

Supplementary Material

Data/Code
Email

Thank you for your interest in spreading the word about medRxiv.

NOTE: Your email address is requested solely to identify you as the sender of this article.

Enter multiple addresses on separate lines or separate them with commas.
Evaluation of a Generative Medical Event Foundation Model for Predicting Post-Discharge Trajectories in Emergency Department Abdominal Pain
(Your Name) has forwarded a page to you from medRxiv
(Your Name) thought you would like to see this page from the medRxiv website.
CAPTCHA
This question is for testing whether or not you are a human visitor and to prevent automated spam submissions.
Share
Evaluation of a Generative Medical Event Foundation Model for Predicting Post-Discharge Trajectories in Emergency Department Abdominal Pain
Kent A. McCann, Donald S. Wright, Mark S. Iscoe, Edward R. Melnick, Lucila Ohno-Machado, Daniella Meeker, Arjun K. Venkatesh, Rohit B. Sangal, Andrew J. Loza
medRxiv 2026.05.18.26353199; doi: https://doi.org/10.64898/2026.05.18.26353199
Twitter logo Facebook logo LinkedIn logo Mendeley logo
Citation Tools
Evaluation of a Generative Medical Event Foundation Model for Predicting Post-Discharge Trajectories in Emergency Department Abdominal Pain
Kent A. McCann, Donald S. Wright, Mark S. Iscoe, Edward R. Melnick, Lucila Ohno-Machado, Daniella Meeker, Arjun K. Venkatesh, Rohit B. Sangal, Andrew J. Loza
medRxiv 2026.05.18.26353199; doi: https://doi.org/10.64898/2026.05.18.26353199

Citation Manager Formats

  • BibTeX
  • Bookends
  • EasyBib
  • EndNote (tagged)
  • EndNote 8 (xml)
  • Medlars
  • Mendeley
  • Papers
  • RefWorks Tagged
  • Ref Manager
  • RIS
  • Zotero
  • Tweet Widget
  • Facebook Like
  • Google Plus One

Subject Area

  • Emergency Medicine
Subject Areas
All Articles
  • Addiction Medicine (617)
  • Allergy and Immunology (899)
  • Anesthesia (330)
  • Cardiovascular Medicine (4804)
  • Dentistry and Oral Medicine (478)
  • Dermatology (414)
  • Emergency Medicine (648)
  • Endocrinology (including Diabetes Mellitus and Metabolic Disease) (1625)
  • Epidemiology (15925)
  • Forensic Medicine (32)
  • Gastroenterology (1202)
  • Genetic and Genomic Medicine (7038)
  • Geriatric Medicine (732)
  • Health Economics (1064)
  • Health Informatics (5033)
  • Health Policy (1432)
  • Health Systems and Quality Improvement (1758)
  • Hematology (584)
  • HIV/AIDS (1343)
  • Infectious Diseases (except HIV/AIDS) (16282)
  • Intensive Care and Critical Care Medicine (1170)
  • Medical Education (667)
  • Medical Ethics (153)
  • Nephrology (722)
  • Neurology (7260)
  • Nursing (366)
  • Nutrition (1080)
  • Obstetrics and Gynecology (1240)
  • Occupational and Environmental Health (1004)
  • Oncology (3603)
  • Ophthalmology (1049)
  • Orthopedics (397)
  • Otolaryngology (451)
  • Pain Medicine (470)
  • Palliative Medicine (139)
  • Pathology (706)
  • Pediatrics (1795)
  • Pharmacology and Therapeutics (737)
  • Primary Care Research (765)
  • Psychiatry and Clinical Psychology (5886)
  • Public and Global Health (9742)
  • Radiology and Imaging (2412)
  • Rehabilitation Medicine and Physical Therapy (1448)
  • Respiratory Medicine (1247)
  • Rheumatology (640)
  • Sexual and Reproductive Health (771)
  • Sports Medicine (578)
  • Surgery (775)
  • Toxicology (107)
  • Transplantation (304)
  • Urology (290)