Abstract
In this study, we assess the capacity of the BERT (Bidirectional Encoder Representations from Transformers) framework to predict a 12-month risk for major diabetic complications—retinopathy, nephropathy, neuropathy, and major adverse cardiovascular events (MACE) using a single-center EHR dataset. We introduce a task-oriented predictive (Top)-BERT architecture, which is a unique end-to-end training and evaluation framework utilizing sequential input structure, embedding layer, and encoder stacks inherent to BERT. This enhanced architecture trains and evaluates the model across multiple learning tasks simultaneously, enhancing the model’s ability to learn from a limited amount of data. Our findings demonstrate that this approach can outperform both traditional pretraining-finetuning BERT models and conventional machine learning methods, offering a promising tool for early identification of patients at risk of diabetes-related complications. We also investigate how different temporal embedding strategies affect the model’s predictive capabilities, with simpler designs yielding better performance. The use of Integrated Gradients (IG) augments the explainability of our predictive models, yielding feature attributions that substantiate the clinical significance of this study. Finally, this study also highlights the essential role of proactive symptom assessment and the management of comorbid conditions in preventing the advancement of complications in patients with diabetes.
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
This study did not receive any funding.
Author Declarations
I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.
Yes
The details of the IRB/oversight body that provided approval or exemption for the research described are given below:
Institutional Review Board of University of Missouri-Columbia gave ethical approval for this work.
I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.
Yes
I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).
Yes
I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.
Yes
Data Availability
All data produced are obtained from the University of Missouri PCORnet Common Data Model.