Abstract
Objective Emergency department (ED) risk models predict whether a patient will revisit within a fixed window, but not what the visit will entail. We evaluated whether a medical event foundation model predicts the timing, sequence, and severity of downstream care after ED discharge for abdominal pain.
Materials and Methods Retrospective cohort study in the de-identified Epic Cosmos network. From 150,030 eligible adults, we analyzed a random sample of 3,000, generating 50 simulated 30-day trajectories per patient with Curiosity, a generative medical event foundation model, and compared these with gradient boosted tree baselines. Endpoints were discrimination for three outcome classes (any revisit, discharge-revisit, admit-revisit) at three horizons, 36-class trajectory accuracy, transition-level calibration, and subgroup performance.
Results Curiosity’s 30-day admit-revisit AUROC was 0.83 (95% CI 0.79-0.87) versus 0.76 (0.71-0.81) for XGBoost (P=.002). Differences were significant at five of nine outcome-horizon pairs but not for discharge-revisit at any horizon. Its most likely trajectory (of 36) matched the observed in 45.9% versus 43.0% (P<.001); median absolute calibration error across 45 transitions was 1.30 percentage points. Discrimination and calibration were poorer for Hispanic or Latino, self-pay, and other-race patients.
Discussion Without task-specific training, one pretrained foundation model matched or exceeded a tuned supervised baseline, with significant advantages on the primary endpoint and trajectory prediction but none for discharge-revisit. Subgroup gaps must be resolved before clinical use.
Conclusion A medical event foundation model can represent post-ED-discharge revisit risk as a calibrated trajectory rather than a single probability; prospective evaluation and remediation of subgroup performance gaps remain necessary.
Competing Interest Statement
The authors have declared no competing interest.
Author Declarations
I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.
Yes
The details of the IRB/oversight body that provided approval or exemption for the research described are given below:
This study was exempted from human subjects review by the Yale institutional review board.
I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.
Yes
I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).
Yes
I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.
Yes
Footnotes
↵* Co-Last Author
Baseline Strengthening The XGBoost baseline was retrained on 100,000 patients, increased from 10,000, with learning curves reported for each outcome. Each of four XGBoost models, three per-horizon models and the 36-class trajectory model, was independently hyperparameter-tuned via Optuna at the 100,000-patient plateau. Primary Endpoint For 30-day admit-revisit, XGBoost AUROC increased from 0.70 to 0.76, with a 95 percent confidence interval of 0.71 to 0.81. Curiosity remained 0.83. The comparison remains significant, with a P value of .002 rather than less than .001. The AUROC gap was approximately halved, from 0.13 to 0.07. The statement that Curiosity was superior at all nine outcome-by-horizon pairs was revised to numerically higher at all nine and statistically significantly higher at five: all three any-revisit horizons and admit-revisit at 7 and 30 days. The three DC-revisit horizons and admit-revisit at 72 hours no longer reached significance. Trajectory-Level Metrics XGBoost single-best trajectory match increased from 41.0 percent to 43.0 percent and remained significant. Top-3 match increased from 63.9 percent to 65.5 percent and remained significant. Top-5 increased from 77.0 percent to 79.2 percent and was no longer significant, with a P value of .08. XGBoost median edit distance improved from 1.40 to 1.385. Curiosity produced the closer trajectory in 60.6 percent of patients, decreased from 63.9 percent. Paired difference and Wilcoxon results were updated. Figure 3 was updated. XGBoost cohort-referenced relative risks increased from 4.72 and 2.66 to 6.02 and 2.72. Group overlap decreased from 23.3 percent to 18.7 percent, and cross-group DC-revisit relative risk decreased from 1.55 to 1.33. The tuned baseline better identifies high-risk patients but still differentiates revisit subtypes less effectively than Curiosity. Calibration and Subgroup Analysis Calibration slope, calibration-in-the-large, and integrated calibration index are now reported for Curiosity across all outcome-by-horizon combinations. Subgroup analyses of 30-day revisit risk are now reported for all eligible subgroups meeting the Cosmos cell-size requirement of more than 11 events. Magnitude Claims The admit-revisit AUC-PR advantage was revised from approximately 2.8-fold to 1.7-fold after XGBoost AUC-PR increased from 0.13 to 0.216. Text Additions A Baseline Learning Curve section was added. XGBoost discrimination plateaued by approximately 60,000 patients and improved to 0.76 after tuning, compared with 0.83 for Curiosity. The baseline limitation now notes both training to the learning-curve plateau and per-endpoint hyperparameter tuning.
Data availability
Deidentified electronic health record data underlying these analyses reside in Epic Cosmos and are available to investigators with approved access through Epic Systems. The Curiosity model used in this study is described in Waxler et al and is hosted within the Epic Cosmos environment. Analytic code resides within Epic Cosmos and is available to investigators with approved Cosmos access.





