Skip to main content
medRxiv
  • Home
  • About
  • Submit
  • ALERTS / RSS
Advanced Search

Nuclear magnetic resonance-based metabolomics with machine learning for predicting progression from prediabetes to diabetes

Jiang Li, Yuefeng Yu, Ying Sun, Yanqi Fu, Wenqi Shen, Lingli Cai, Xiao Tan, Yan Cai, Ningjian Wang, Yingli Lu, Bin Wang
doi: https://doi.org/10.1101/2024.05.14.24307378
Jiang Li
aInstitute and Department of Endocrinology and Metabolism, Shanghai Ninth People’s Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Yuefeng Yu
aInstitute and Department of Endocrinology and Metabolism, Shanghai Ninth People’s Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Ying Sun
aInstitute and Department of Endocrinology and Metabolism, Shanghai Ninth People’s Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Yanqi Fu
aInstitute and Department of Endocrinology and Metabolism, Shanghai Ninth People’s Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Wenqi Shen
aInstitute and Department of Endocrinology and Metabolism, Shanghai Ninth People’s Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Lingli Cai
aInstitute and Department of Endocrinology and Metabolism, Shanghai Ninth People’s Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Xiao Tan
bDepartment of Medical Sciences, Uppsala University, Uppsala, Sweden
cDepartment of Big Data in Health Science, School of Public Health, Zhejiang University School of Medicine, Hangzhou, China
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Yan Cai
dDepartment of Endocrinology, the Fifth Affiliated Hospital of Kunming Medical University, Yunnan Honghe Prefecture Central Hospital (Ge Jiu People’s Hospital), Yunnan, China
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Ningjian Wang
aInstitute and Department of Endocrinology and Metabolism, Shanghai Ninth People’s Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Yingli Lu
aInstitute and Department of Endocrinology and Metabolism, Shanghai Ninth People’s Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • For correspondence: binwang1126{at}163.com luyingli2008{at}126.com
Bin Wang
aInstitute and Department of Endocrinology and Metabolism, Shanghai Ninth People’s Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • For correspondence: binwang1126{at}163.com luyingli2008{at}126.com
  • Abstract
  • Full Text
  • Info/History
  • Metrics
  • Supplementary material
  • Data/Code
  • Preview PDF
Loading

Abstract

Background Identification of individuals with prediabetes who are at high risk of developing diabetes allows for precise interventions. We aimed to determine the role of nuclear magnetic resonance (NMR)-based metabolomic signature in predicting the progression from prediabetes to diabetes.

Methods This prospective study included 13,489 participants with prediabetes who had metabolomic data from the UK Biobank. Circulating metabolites were quantified via NMR spectroscopy. Cox proportional hazard (CPH) models were performed to estimate the associations between metabolites and diabetes risk. Supporting vector machine, random forest, and extreme gradient boosting were used to select the optimal metabolite panel for prediction. CPH and random survival forest (RSF) models were utilized to validate the predictive ability of the metabolites.

Results During a median follow-up of 13.6 years, 2,525 participants developed diabetes. After adjusting for covariates, 94 of 168 metabolites were associated with risk of progression to diabetes. A panel of nine metabolites, selected by all three machine learning algorithms, was found to significantly improve diabetes risk prediction beyond conventional risk factors in the CPH model (area under the receiver operating characteristic curve [AUROC], 1-year: 0.823 for risk factors + metabolites vs 0.759 for risk factors, 5-year: 0.830 vs 0.798, 10-year: 0.801 vs 0.776, all P <0.05). Similar results were observed from the RSF model. Categorization of participants according to the predicted value thresholds revealed distinct cumulative risk of diabetes.

Conclusions Our study lends support for use of the metabolite markers to help determine individuals with prediabetes who are at high risk of progressing to diabetes and inform targeted and efficient interventions.

Funding Shanghai Municipal Health Commission (2022XD017). Innovative Research Team of High-level Local Universities in Shanghai (SHSMU-ZDCX20212501). Shanghai Municipal Human Resources and Social Security Bureau (2020074). Clinical Research Plan of Shanghai Hospital Development Center (SHDC2020CR4006). Science and Technology Commission of Shanghai Municipality (22015810500).

Graphical abstract.
  • Download figure
  • Open in new tab
Graphical abstract.

CPH, Cox proportional hazard; NMR, nuclear magnetic resonance; RF, random forest; RSF, Random survival forest; SVM, supporting vector machine; XGBoost, extreme gradient boosting.

Introduction

Prediabetes, an intermediate stage of glucose dysregulation that blood glucose levels are elevated but lower than in diabetes, has become a burgeoning global health emergency1. Prediabetes affected approximately 720 million individuals worldwide in 2021, with a project to 1 billion people by 20452. Approximately 5% to 10% of people with prediabetes progress to having diabetes each year and the lifetime conversion rate to diabetes could be as high as 70%3,4. Therefore, preventing or delaying diabetes development among people with prediabetes will have substantial clinical and public health benefits.

Although lifestyle modification and medical therapy have been proven to be effective in preventing or delaying the diabetes onset among people with prediabetes5-7, the substantial cost of modification programs and medications as well as drug-related side effects limit the widespread delivery of such interventions in this large high-risk population8,9. Notably, the progression from prediabetes to diabetes is highly heterogeneous, and a fraction of individuals with prediabetes may regress to normoglycemia without treatment.10 Therefore, identifying targeted population who are at high risk of developing diabetes is the key step to tailor precise and efficient interventions. Glycemic indicators alone for risk stratification are deficient, with fasting glucose and glycosylated hemoglobin A1c (HbA1c) being convenient but less sensitive, while post-load glucose tolerance being sensitive but unfeasible in practice on a large scale11,12. In addition, several risk assessment models based on conventional clinical variables have been developed, but most of which had comparatively low performance and failed to take follow-up time into account13-15.

Plasma metabolomics using high-throughput techniques could provide a comprehensive profiling of small-molecule metabolites in a specific physiological period, which might yield valuable information for risk prediction. Previous studies have implied that incorporating circulating metabolites into basic models with conventional risk factors could improve prediction of diabetes risk 16-18. However, we are aware of only one study that has assessed the relationship between metabolomic profiling and the progression to diabetes among individuals with prediabetes and investigated the predictive values of metabolites19. Nevertheless, it was limited by a nested case-control study design with a relatively short follow-up (median 5 years) and small sample size (n=∼300). Whether addition of metabolic biomarkers improves the ability in predicting the progression from prediabetes to diabetes in prospective settings remains largely unknown.

To address these knowledge gaps, in the current study we aimed to examine the longitudinal associations of circulating metabolic biomarkers, quantified using high-throughput nuclear magnetic resonance (NMR), with the risk of incident diabetes among individuals with prediabetes from the UK Biobank. Moreover, we evaluated whether metabolic signature adds anything to prediction models for diabetes development and risk stratification.

Methods

Study design and participants

The UK Biobank is a large population-based prospective cohort study enrolling more than 500,000 community-dwelling adults from 22 assessment centers across the UK between 2006 and 201020,21. Participants completed touchscreen questionnaires and physical measurements and provided blood samples at baseline. The study was approved by the Northwest Multicenter Research Ethics Committee (REC reference for UK Biobank 11/NW/0382), and all participants provided informed consent.

For the identification of metabolomic biomarkers associated with the progression from prediabetes to diabetes, the current study focused on participants with prediabetes at baseline with available circulating metabolite data. The diagnosis of prediabetes was defined by an HbA1c level of 5.7% to 6.4% (39 to 47 mmol/mol) in participants without diabetes, according to the American Diabetes Association (ADA) criteria22. After excluding individuals who developed diabetes or died within 1 month from the baseline, 13,489 participants with prediabetes were included in the final analyses.

Metabolite quantification

The metabolomics analysis of approximately 118,000 non-fasting ethylenediaminetetraacetic acid (EDTA) plasma samples at baseline was performed using the high-throughput NMR platform in Nightingale Health’s laboratories of Finland. Details of the metabolic profiling platform and experimentation have been described elsewhere23-25. In brief, the EDTA samples were collected and stored at -80°C. Before preparation, frozen samples were slowly thawed at +4°C overnight and were centrifuged (3,400 g) for 3 minutes. Each sample was analyzed with a spectrometer and the metabolic biomarkers were quantified using Nightingale Health’s proprietary software. The quality control procedures were implemented during the whole process and only samples and biomarkers that underwent the quality control process were stored in the UK Biobank dataset (Method S1). The coefficient of variations was below 5% for most of the biomarkers and there were no batch effects26.

A total of 249 metabolic biomarkers (168 directly measured and 81 ratios of these), spanning lipids, lipoprotein subclass, fatty acids, amino acids, ketone bodies, and glycolysis metabolites were quantified for each sample. In the present study, we analyzed 168 metabolic biomarkers that were directly measured (Supplementary Table S1). The values of all metabolites were transformed using natural logarithmic transformation (ln[x+1]) followed by Z-transformation.

Covariate collection

Information on covariates was collected through a self-completed touchscreen questionnaire or verbal interview at baseline, including age, sex, ethnicity, Townsend deprivation index, household income, education, employment status, smoking status, moderate alcohol, physical activity, healthy diet score, healthy sleep score, family history of diabetes, history of cardiovascular disease (CVD), history of hypertension, history of dyslipidemia, history of chronic lung diseases (CLD), and history of cancer.

Physical measurements included systolic (SBP) and diastolic blood pressure (DBP), height, weight, waist circumference (WC), and hip circumference (HC). Body mass index (BMI) was calculated as weight in kilograms divided by the square of height in meters (kg/m²). Missing covariates were imputed by the median value for continuous variables and a missing indicator for categorical variables. More details about covariates collection can be found in Method S2.

Ascertainment of diabetes

Incident diabetes was ascertained from hospital inpatient records, death registers, and primary care records, according to the International Classification of Diseases, 10th revision (ICD-10) codes. Detailed information about the linkage procedure is available from https://content.digital.nhs.uk/services. The follow-up time was calculated from the baseline to the occurrence of diabetes, death, or the censoring date (30 March 2023), whichever came first.

Statistical analyses

Baseline characteristics were presented as numbers (percentages) for categorical variables and means (standard deviations, SDs) for continuous variables, respectively. Continuous variables were assessed for statistical differences using t-test and categorical variables were evaluated using the χ2 test. Overall schematic workflow of the study is shown in Graphical abstract.

Metabolite selection

We first used Cox proportional hazards (CPH) model to assess the associations between individual metabolites and risk of diabetes progression with adjustment for sociodemographic covariates (age, sex, ethnicity, education, Townsend Deprivation Index, employment status and household income), family history of diabetes, health conditions (history of CVD, hypertension, dyslipidemia, CLD and cancer), physical measurements (BMI, WC, HC, SBP and DBP), lifestyle factors (smoking status, moderate alcohol, healthy diet score, healthy sleep score and physical activity), and HbA1c. The potential confounders were selected based on prior knowledge of the risk factors for diabetes. Metabolites that were significantly associated with incident diabetes (P <0.05/168) were retained.

Secondly, we performed priority-Lasso to deal with multicollinearity in high dimensional data and to retain variables with nonzero coefficients. Priority-Lasso is a Lasso-based intuitive analysis strategy, which uses prior knowledge regarding the outcome by defining the blocks of different types of predictor variables27. In this study, we defined the 24 covariates as block 1, while all metabolites significantly associated with diabetes risk in the CPH model were defined as block 2. The penalization parameter λ was determined as values with maximum partial-likelihood in a 10-fold cross-validation.

Thirdly, three machine learning models including supporting vector machine (SVM), random forest (RF), and extreme gradient boosting (XGBoost) were adopted to further evaluate the importance of the Lasso-selected metabolites, as they can model nonlinear and nonadditive relations more flexibly28. Models were built by 10-fold cross-validation through the “caret” package. Common signals detected across diverse approaches are more likely to represent the strongest and true patterns in the data. We chose the intersection set of the top 20 most important variables selected by the three machine learning models, after balancing the performance of the final diabetes risk prediction model and the clinical applicability associated with measurement costs of metabolites.

Model development

Participants were randomly subclassified into a training set and a test set at a ratio of 8:2 and two common algorithms for survival data including CPH model and random survival forest (RSF)29 were adopted for model development. RSF, as a machine learning method, is designed to be used specifically for survival outcome prediction and has shown promising results in various settings30,31. It builds many decision trees using split points based on the log-rank test to identify different survival statuses and produces the predicted probability for an individual derived from the average prediction across all trees32. The RSF model was fitted using the “randomForestSRC” package and the grid search method was used for hyperparameter tuning (number of trees, number of variables to possibly split at each node, and minimum size of terminal node) (See Method S3 for more details).

Model evaluation

The model performance was assessed in the test set. The time-dependent area under the receiver operating characteristic curve (AUROC) was used to evaluate the model’s discrimination ability. Continuous net reclassification improvement (NRI), and absolute integrated discrimination improvement (IDI) were used to assess whether adding the selected metabolites could improve risk discrimination and reclassification for the risk of progression from prediabetes to diabetes over the basic model that was built on 10 conventional clinical variables (age, sex, Townsend Deprivation Index, family history of diabetes mellitus, BMI, WC, HC, SBP, DBP, and HbA1c)33. The calibration ability of the model was estimated using calibration curve. Furthermore, we used decision curve analysis (DCA) to assess the clinical usefulness of prediction model-based guidance for prediabetes management, which calculates a clinical “net benefit” for one or more prediction models in comparison to default strategies of treating all or no patients34. To facilitate risk stratification, we classified participants into two risk groups according to the predictive value using “surv_cutpoint” function in the “survminer” R package35. We also divided participants into three categories according to the tertiles of probability. In addition, we included 90,688 participants with normal glucose from the UK Biobank and divided them into the training and test sets using an 8:2 ratio to further investigate the additive value of the selected metabolites in diabetes prediction among participants with normoglycemia. All analyses were conducted in R software (version 4.2.2). A two-sided P value < 0.05 was considered statistically significant. To control for the false discovery rate in the association between multiple metabolic biomarkers and incident diabetes, Bonferroni correction for P value (P <0.05/168) was used.

Results

Baseline characteristics

Among the 13,489 participants with baseline prediabetes, the mean age was 59.6 (SD, 7.1) years, and 6,166 (45.7%) were males. During a median follow-up of 13.6 (12.3-14.6) years, 2,525 (18.7%) participants progressed to diabetes. Baseline characteristics of the study population stratified by incident diabetes are summarized in Table 1. Participants who developed diabetes were more likely to be male, non-White, less educated, more deprived, and smokers. They also tended to have a family history of diabetes, comorbidities such as CVD, hypertension, dyslipidemia and CLD, and higher levels of BMI, WC, and HC.

View this table:
  • View inline
  • View popup
Table 1. Baseline characteristics of participants with prediabetes stratified by incident diabetes status.

Identification of metabolic biomarkers for progression to diabetes

After adjusting for covariates and correcting for multiple testing, 94 of 168 metabolic biomarkers were significantly associated with the risk of incident diabetes (Figure 1, Table S2). Concentrations of very low-density lipoprotein (VLDL) particles, particularly larger VLDL particles and composition within larger VLDL, were strongly associated with progression to diabetes. Triglyceride in all lipoprotein subclasses also demonstrated strong positive associations with diabetes risk. In contrast, concentrations of larger high-density lipoprotein (HDL) particles and composition within these particles were inversely associated with incident diabetes. For lipoprotein particle diameter, larger HDL and LDL particle sizes were associated with a lower risk of progression to diabetes, while larger VLDL particle size was associated with a higher risk.

Figure 1.
  • Download figure
  • Open in new tab
Figure 1. Associations of 168 metabolic biomarkers with risk of diabetes among 13,489 participants with prediabetes.

Hazard ratios (HR) were presented per 1 standard deviation (SD) higher of metabolic biomarker on the natural log scale and were adjusted for age, sex, ethnicity, education, Townsend Deprivation Index, employment status, household income, family history of diabetes, history of CVD, history of hypertension, history of dyslipidemia, history of CLD, history of cancer, body mass index, waist circumference, hip circumference, smoking status, moderate alcohol, healthy diet score, healthy sleep score, physical activity, systolic blood pressure, diastolic blood pressure and glycated hemoglobin A1c. *False discovery rate controlled P < 0.05/168.

Apo-A1, apolipoprotein A1; Apo-B, apolipoprotein B; Apo-LP, apolipoprotein; BCAA, branched-chain amino acid; BMI, body mass index; CVD, cardiovascular disease; CLD, chronic lung disease; DHA, docosahexaenoic acid; FA, fatty acids; HDL, high-density lipoproteins; HDL-D, high-density lipoprotein particle diameter; IDL, intermediate-density lipoproteins; L, large; LA, linoleic acid; LDL, low-density lipoproteins; LDL-D, low-density lipoprotein particle diameter; LP, lipoprotein; M, medium; MUFA, monounsaturated fatty acids; PUFA, polyunsaturated fatty acids; S, small; SFA, saturated fatty acids; VLDL, very-low-density lipoproteins; VLDL-D, very-low-density lipoprotein particle diameter; XL, very large; XS, very small; XXL, extremely large.

Monounsaturated fatty acids and saturated fatty acids were positively associated with the risk of diabetes, whereas docosahexaenoic acid and the degree of fatty acid unsaturation were negatively associated with diabetes. Among the amino acids, higher concentrations of alanine, tyrosine, and branched-chain amino acid (BCAA) such as leucine and valine were associated with an increased risk of diabetes, but glutamine and glycine were inversely associated with diabetes. Neither of the ketone bodies showed an association with the risk of diabetes.

Of the 94 metabolites that were significantly associated with diabetes, 17 metabolites were selected by priority-Lasso (Table S3). When further evaluating the importance of these metabolites after adjustment for covariates using three machine learning algorithms, the intersection of the top 20 important predictors identified a total of 9 metabolites, namely cholesteryl esters in large HDL, cholesteryl esters in medium VLDL, triglycerides in very large VLDL, average diameter for LDL particles, triglycerides in IDL, glycine, tyrosine, glucose, and docosahexaenoic acid (Figure 2, Table S4).

Figure 2.
  • Download figure
  • Open in new tab
Figure 2. The top 20 important variables selected by three machine learning models: (A) supporting vector machine (SVM); (B) extreme gradient boosting (XGBoost); (C) random forest (RF).

The models were adjusted for age, sex, ethnicity, education, Townsend Deprivation Index, employment status, household income, family history of diabetes, history of CVD, history of hypertension, history of dyslipidemia, history of CLD, history of cancer, body mass index, waist circumference, hip circumference, smoking status, moderate alcohol, healthy diet score, healthy sleep score, physical activity, systolic blood pressure, diastolic blood pressure and glycated hemoglobin A1c. CVD, cardiovascular disease; CLD, chronic lung disease. HDL, high-density lipoproteins; IDL, intermediate-density lipoproteins; LDL, low-density lipoproteins; VLDL, very-low-density lipoproteins.

Model development and evaluation

Build upon the selected 9 metabolites and 10 clinical variables, there was no obvious difference in the AUROC obtained from CPH model (1-year: 0.823 [95% confidence interval, CI 0.702, 0.945]; 5-year: 0.830 [0.797, 0.864]; 10-year: 0.801 [0.778, 0.825]) and RSF model (1-year: 0.828 [0.723, 0.933]; 5-year: 0.820 [0.785, 0.855]; 10-year: 0.802 [0.778, 0.826]). Hence, we chose CPH model as the final model because of its simplicity and interpretability. The addition of selected metabolites consecutively outperformed the basic model with conventional clinical variables in diabetes risk prediction from 1 to 10 years (Figure S1). Specifically, the AUROC increased from 0.759 (95% CI 0.608, 0.911) to 0.823 (0.702, 0.945), 0.798 (0.762, 0.834) to 0.830 (0.797, 0.864), and 0.776 (0.750, 0.801) to 0.801 (0.778, 0.825) for 1-year, 5-year, and 10-year diabetes risk, respectively (Table 2, Figure S2). Results from continuous NRI and absolute IDI also demonstrated improvement in the risk prediction for progression to diabetes (Table 2), although the model calibration was not significantly improved (Figure S3). The decision curve analysis showed that the inclusion of the metabolites had a higher net benefit across the threshold probabilities of 0-0.35 for predicting 5-year diabetes risk and 0-0.55 for predicting 10-year diabetes risk (Figure S4).

View this table:
  • View inline
  • View popup
  • Download powerpoint
Table 2. Performance of Cox proportional hazards regression models in prediction of the progression of prediabetes to diabetes.

We further categorized the participants from the test set into low-risk and high-risk groups according to the optimal threshold of the predicted value (1.02) reflecting the best risk difference. Compared with the low-risk group, participants in the high-risk group had a significantly higher cumulative risk of incident diabetes (log-rank P < 0.0001) (Figure 3). When participants were alternatively classified into low-risk, medium-risk, and high-risk groups according to the tertile cut-off point of the predicted value, the high-risk group showed the highest risk of developing diabetes, followed by the medium-risk and low-risk groups (log-rank P < 0.0001). Similar results were also observed when considering the competing risk from death (Fine-Gray P < 0.0001) (Figure S5). In addition, the predicted risk of diabetes within 1 year (P = 0.001), 5 years (P < 0.001), or 10 years (P < 0.001) was generally higher among participants who progressed to diabetes than those who did not (Figure 4).

Figure 3.
  • Download figure
  • Open in new tab
Figure 3. Cumulative hazard curves for participants with prediabetes with different risks stratified by the Cox model based on clinical variables and 9 metabolites.

The Cox model divided participants with prediabetes in the test set to two categories (A) and three categories (B) with significant differences in cumulative hazard of diabetes during the follow-up (both P <0.0001).

Figure 4.
  • Download figure
  • Open in new tab
Figure 4. The distribution of the predictive probability of developing diabetes among participants with prediabetes by incident diabetes status within 1-year (A), 5-year (B), and 10-year (C).

Among participants with normoglycemia, we also observed a significant improvement in the prediction of diabetes after the addition of metabolic biomarkers to the basic model. The AUROC increased from 0.821 (95% CI 0.736, 0.907) to 0.868 (0.802, 0.934), 0.790 (0.738, 0.842) to 0.811 (0.762, 0.860), and 0.791 (0.765, 0.816) to 0.806 (0.781, 0.831) for 1-year, 5-year, and 10-year diabetes risk, respectively. (Table S5). The increases in NRI and IDI were similar to or slightly lower than those found among participants with prediabetes.

Discussion

By leveraging data from the large UK Biobank cohort, this prospective study provided a comprehensive analysis of the associations of circulating metabolites with the risk of progression to diabetes and predictive ability in participants with prediabetes. We found that lipoprotein particles, lipoprotein particle size and composition, fatty acids, and amino acids were associated with the risk of incident diabetes. More importantly, our findings suggested that adding the selected metabolites (i.e., cholesteryl esters in large HDL, cholesteryl esters in medium VLDL, triglycerides in very large VLDL, average diameter for LDL particles, triglycerides in IDL, glycine, tyrosine, glucose, and docosahexaenoic acid) could significantly improve the risk prediction of progression from prediabetes to diabetes beyond the conventional clinical variables.

In the present study, the association between diabetes risk and lipid and lipoprotein profile, including VLDL particles and composition with larger VLDL, HDL particles and composition within larger HDL, triglyceride, smaller HDL and LDL particle sizes, and larger VLDL particle sizes, were broadly consistent with previous studies in the general population36-39. BCAAs have been widely reported to be involved in the pathogenesis of diabetes, which might impair insulin signaling and lead to increased insulin secretion and pancreatic β-cell exhaustion40. Furthermore, genetic association studies have shown higher BCAAS resulting from insulin resistance, which may in turn cause diabetes41,42. Our study confirmed the vital role of these metabolites in the progression to diabetes among individuals with prediabetes.

Several risk assessment models for predicting the risk of progression from prediabetes to diabetes have been reported13-15. Yokota et al. developed a logistic regression model to predict the risk for conversion from prediabetes to diabetes based on family history of diabetes, sex, SBP, fasting plasma glucose (FPG), HbA1c, and alanine aminotransferase (ALT)13. The model derived from a retrospective longitudinal study design achieved an AUROC of 0.80 (0.70–0.87) but did not take follow-up time into account. Similarly, Liang et al developed a predictive model using three glycemic indicators (FPG, 2-h postprandial blood glucose [2-hPG], and HbA1c) alone15 and obtained a relatively low AUROC of 0.732 (95% CI 0.688-0.776). In a cohort study of 852,454 individuals with prediabetes, a machine-learning model predicting the progression to diabetes within 1-year was established using data from electronic medical records14. The model built on age, gender, BMI, medication usage, and laboratory results achieved a high AUROC of 0.865 (0.860-0.869). However, the model’s performance over a longer follow-up period was unclear and conventional parameters such as lifestyle, family history of diabetes or comorbidities were not taken into account.

Changes in circulating small-molecule metabolites may occur long before the disease onset. Although rapid development in the technology of metabolomics provides a powerful tool for precise disease prediction, few studies have investigated the role of metabolomics-derived metabolic biomarkers in predicting progression from prediabetes to diabetes. To our best knowledge, only one case-control study among 153 individuals with prediabetes and 160 matched controls reported that adding 13 metabolites to conventional clinical variables including BMI, waist-hip ratio, WC, SBP, DBP, triglyceride, LDL, and triglyceride-glucose index improved the risk prediction of diabetes progression within 5 years, with the AUROC increasing from 0.72 to 0.9819. However, the predictive ability of metabolites in prospective settings with large sample size remains uncertain. In this longitudinal study among 13,489 participants with prediabetes, we comprehensively used multiple machine learning algorithms to identify a panel of 9 circulating metabolites that were associated with diabetes incidence during a median follow-up of 13.6 years. The CPH model integrating conventional clinical variables and the selected metabolic signature achieved a comparatively high AUROC of 0.823, 0.830, and 0.801 for 1-year, 5-year, and 10-year diabetes risk, respectively. Importantly, the addition of the metabolites resulted in a significant improvement in the discrimination ability and risk reclassification of diabetes beyond conventional risk factors. Furthermore, we categorized participants according to the optimal threshold points of the predicted value and found that the high-risk group had a significantly higher cumulative incidence of diabetes than the low-risk group. Most importantly, a model with good discrimination does not necessarily have high clinical value. Hence, DCA was used to compare the clinical utility of the model before and after adding the metabolites, and this showed a higher net benefit for the latter than the basic model, suggesting the addition of the metabolites increased the clinical value of prediction, i.e., the potential benefit of guiding management in individuals with prediabetes34,43. These results provided novel evidence supporting the value of metabolic biomarkers in risk prediction and stratification for the progression from prediabetes to diabetes. Considering the epidemic proportion of prediabetes worldwide, even a modest improvement in diabetes risk prediction among individuals with prediabetes will have substantial clinical and public health implications. Early detection of individuals with prediabetes who are at high risk of developing diabetes would not only advance targeted screening initiatives, health management and interventions but also facilitate a rational allocation of medical resources while avoiding disproportionate healthcare expenditure, which could finally translate into precise and efficient prevention of diabetes. The value of the selected metabolic biomarkers in diabetes prediction was also confirmed in individuals with normal glucose.

Our study presents several strengths. Circulating metabolites were quantified via NMR-based metabolome profiling within the UK Biobank, which offers metabolite qualification with relatively lower costs and better reproducibility26. Additional strengths of our study included large sample size, prospective study design with long-term follow-up, and comprehensive control of covariates. Moreover, we used multiple machine learning algorithms to identify the consistently important metabolic biomarkers based on which we developed the predictive models. The final model exhibited relatively high performance for 1-year, 5-year, and 10-year diabetes risk prediction. However, several limitations of our study should be noted. First, since FPG and 2-hPG were not available in the UK Biobank, we defined prediabetes using HbA1c alone and to what extent our results could be extrapolated to other people with prediabetes determined by multiple glycemic indicators requires further investigation. Second, circulating metabolites were measured at baseline, thus their dynamic change over time could not be captured. However, our models showed stable performance in predicting short-term and long-term progression to diabetes (1 to 10 years), indicating the validity of single measurements of metabolic biomarkers for risk prediction. Third, the Nightingale metabolomics platform primarily focused on lipids and lipoprotein sub-fractions, and thus the predictive value of other metabolites in the progression from prediabetes to diabetes warranted further research using an untargeted metabolomics approach. Additionally, the use of non-fasting blood samples might increase inter-individual variation in metabolic biomarker concentrations, however, fasting duration has been reported to account for only a small proportion of variation in plasma metabolic biomarker concentrations44. Therefore, we believe the impact of non-fasting samples on our findings would be minor. Fourth, although incident diabetes cases were ascertained through different data sources, including hospital inpatient records, death registers, and primary care records, some undiagnosed diabetes might have been missed. This misclassification would underestimate the effect of the observed associations between metabolites and diabetes risk. Fifth, we could not draw any conclusion about the causality between the identified metabolites and the risk for progression to diabetes due to the observational nature, which remained to be validated in further experimental studies. Sixth, in this study, the prediction models were established and tested using the UK Biobank dataset, external validation in an independent cohort is warranted to confirm the predictive values of the metabolic biomarkers. Finally, the participants from the UK Biobank were mostly White, which might limit the generalizability of the findings to other populations.

Conclusions

In this large prospective study among individuals with prediabetes, we detected a panel of circulating metabolites that were associated with an increased risk of progressing to diabetes. Use of these metabolites significantly improved the risk prediction of progression from prediabetes to diabetes. Our findings provide evidence that integrating metabolite markers with conventional risk factors is a promising approach to advance effective screening strategies and precise interventions for individuals with prediabetes who are at high risk of developing diabetes.

Availability of data and materials

The data analyzed during this study are available at https://www.ukbiobank.ac.uk/. This research has been conducted using the UK Biobank Resource under application number 77740.

Competing interests

The authors declare no competing interests.

Authors’ contributions

B.W. and Y.L. conceived and designed the study. J.L. and Y.Y. performed the statistical analysis and drafted the manuscript. Y.S., Y.F., W.S., and L.C. participated in data collection. B.W., Y.L., X.T., N.W. critically revised the manuscript. All authors read and approved the final manuscript. B.W. is the guarantor of this work and, as such, had full access to all the data in the study and takes responsibility for the integrity of the data and the accuracy of the data analysis.

Acknowledgements

We thank all participants and staff in the UK Biobank for their dedication and contribution to this study.

Footnotes

  • Table 1 revised; Supplemental files updated.

References

  1. ↵
    Echouffo-Tcheugui, J. B. & Selvin, E. Prediabetes and What It Means: The Epidemiological Evidence. Annu Rev Public Health 42, 59–77 (2021). doi:10.1146/annurev-publhealth-090419-102644
    OpenUrlCrossRef
  2. ↵
    Sun, H. et al. IDF Diabetes Atlas: Global, regional and country-level diabetes prevalence estimates for 2021 and projections for 2045. Diabetes Res Clin Pract 183, 109119 (2022). doi:10.1016/j.diabres.2021.109119
    OpenUrlCrossRefPubMed
  3. ↵
    Tabák, A. G., Herder, C., Rathmann, W., Brunner, E. J. & Kivimäki, M. Prediabetes: a high-risk state for diabetes development. Lancet 379, 2279–2290 (2012). doi:10.1016/s0140-6736(12)60283-9
    OpenUrlCrossRefPubMedWeb of Science
  4. ↵
    Ligthart, S. et al. Lifetime risk of developing impaired glucose metabolism and eventual progression from prediabetes to type 2 diabetes: a prospective cohort study. Lancet Diabetes Endocrinol 4, 44–51 (2016). doi:10.1016/s2213-8587(15)00362-9
    OpenUrlCrossRef
  5. ↵
    Gong, Q. et al. Morbidity and mortality after lifestyle intervention for people with impaired glucose tolerance: 30-year results of the Da Qing Diabetes Prevention Outcome Study. Lancet Diabetes Endocrinol 7, 452–461 (2019). doi:10.1016/s2213-8587(19)30093-2
    OpenUrlCrossRef
  6. DeFronzo, R. A. et al. Pioglitazone for diabetes prevention in impaired glucose tolerance. N Engl J Med 364, 1104–1115 (2011). doi:10.1056/NEJMoa1010949
    OpenUrlCrossRefPubMedWeb of Science
  7. ↵
    Herman, W. H. Prediabetes Diagnosis and Management. JAMA 329, 1157–1159 (2023). doi:10.1001/jama.2023.4406
    OpenUrlCrossRef
  8. ↵
    Roberts, S. et al. Preventing type 2 diabetes: systematic review of studies of cost-effectiveness of lifestyle programmes and metformin, with and without screening, for pre-diabetes. BMJ Open 7, e017184 (2017). doi:10.1136/bmjopen-2017-017184
    OpenUrlCrossRefPubMed
  9. ↵
    Piller, C. Dubious diagnosis. Science 363, 1026–1031 (2019). doi:10.1126/science.363.6431.1026
    OpenUrlAbstract/FREE Full Text
  10. ↵
    Shang, Y. et al. Natural history of prediabetes in older adults from a population-based longitudinal study. J Intern Med 286, 326–340 (2019). doi:10.1111/joim.12920
    OpenUrlCrossRefPubMed
  11. ↵
    Phillips, L. S., Ratner, R. E., Buse, J. B. & Kahn, S. E. We can change the natural history of type 2 diabetes. Diabetes Care 37, 2668–2676 (2014). doi:10.2337/dc14-0817
    OpenUrlAbstract/FREE Full Text
  12. ↵
    Ferrannini, E. Definition of intervention points in prediabetes. Lancet Diabetes Endocrinol 2, 667–675 (2014). doi:10.1016/s2213-8587(13)70175-x
    OpenUrlCrossRef
  13. ↵
    Yokota, N. et al. Predictive models for conversion of prediabetes to diabetes. Journal of Diabetes and its Complications 31, 1266–1271 (2017). doi:10.1016/j.jdiacomp.2017.01.005
    OpenUrlCrossRefPubMed
  14. ↵
    Cahn, A. et al. Prediction of progression from pre-diabetes to diabetes: Development and validation of a machine learning model. Diabetes Metab Res Rev 36, e3252 (2020). doi:10.1002/dmrr.3252
    OpenUrlCrossRef
  15. ↵
    Liang, K. et al. Nomogram Predicting the Risk of Progression from Prediabetes to Diabetes After a 3-Year Follow-Up in Chinese Adults. Diabetes, Metabolic Syndrome and Obesity: Targets and Therapy Volume 14, 2641–2649 (2021). doi:10.2147/dmso.S307456
    OpenUrlCrossRef
  16. ↵
    Merino, J. et al. Metabolomics insights into early type 2 diabetes pathogenesis and detection in individuals with normal fasting glucose. Diabetologia 61, 1315–1324 (2018). doi:10.1007/s00125-018-4599-x
    OpenUrlCrossRefPubMed
  17. Peddinti, G. et al. Early metabolic markers identify potential targets for the prevention of type 2 diabetes. Diabetologia 60, 1740–1750 (2017). doi:10.1007/s00125-017-4325-0
    OpenUrlCrossRef
  18. ↵
    Rebholz, C. M. et al. Serum metabolomic profile of incident diabetes. Diabetologia 61, 1046–1054 (2018). doi:10.1007/s00125-018-4573-7
    OpenUrlCrossRefPubMed
  19. ↵
    Ren, M. et al. Potential Novel Serum Metabolic Markers Associated With Progression of Prediabetes to Overt Diabetes in a Chinese Population. Front Endocrinol (Lausanne) 12, 745214 (2021). doi:10.3389/fendo.2021.745214
    OpenUrlCrossRef
  20. ↵
    Sudlow, C. et al. UK biobank: an open access resource for identifying the causes of a wide range of complex diseases of middle and old age. PLoS Med 12, e1001779 (2015). doi:10.1371/journal.pmed.1001779
    OpenUrlCrossRefPubMed
  21. ↵
    Allen, N. et al. UK Biobank: Current status and what it means for epidemiology. Health Policy and Technology 1, 123–126 (2012).
    OpenUrl
  22. ↵
    ElSayed, N. A. et al. 2. Classification and Diagnosis of Diabetes: Standards of Care in Diabetes-2023. Diabetes Care 46, S19-s40 (2023). doi:10.2337/dc23-S002
    OpenUrlCrossRefPubMed
  23. ↵
    Würtz, P. et al. Quantitative Serum Nuclear Magnetic Resonance Metabolomics in Large-Scale Epidemiology: A Primer on -Omic Technologies. Am J Epidemiol 186, 1084–1096 (2017). doi:10.1093/aje/kwx016
    OpenUrlCrossRefPubMed
  24. Soininen, P., Kangas, A. J., Würtz, P., Suna, T. & Ala-Korpela, M. Quantitative serum nuclear magnetic resonance metabolomics in cardiovascular epidemiology and genetics. Circ Cardiovasc Genet 8, 192–206 (2015). doi:10.1161/circgenetics.114.000216
    OpenUrlAbstract/FREE Full Text
  25. ↵
    Zhang, X. et al. Plasma metabolomic profiles of dementia: a prospective study of 110,655 participants in the UK Biobank. BMC Med 20, 252 (2022). doi:10.1186/s12916-022-02449-3
    OpenUrlCrossRef
  26. ↵
    Geng, T.-T. et al. Nuclear Magnetic Resonance–Based Metabolomics and Risk of CKD. American Journal of Kidney Diseases (2023). doi:10.1053/j.ajkd.2023.05.014
    OpenUrlCrossRef
  27. ↵
    Klau, S., Jurinovic, V., Hornung, R., Herold, T. & Boulesteix, A. L. Priority-Lasso: a simple hierarchical approach to the prediction of clinical outcome using multi-omics data. BMC Bioinformatics 19, 322 (2018). doi:10.1186/s12859-018-2344-6
    OpenUrlCrossRef
  28. ↵
    Morgenstern, J. D., Rosella, L. C., Costa, A. P., de Souza, R. J. & Anderson, L. N. Perspective: Big Data and Machine Learning Could Help Advance Nutritional Epidemiology. Adv Nutr 12, 621–631 (2021). doi:10.1093/advances/nmaa183
    OpenUrlCrossRef
  29. ↵
    Qiu, X. et al. A Comparison Study of Machine Learning (Random Survival Forest) and Classic Statistic (Cox Proportional Hazards) for Predicting Progression in High-Grade Glioma after Proton and Carbon Ion Radiotherapy. Front Oncol 10, 551420 (2020). doi:10.3389/fonc.2020.551420
    OpenUrlCrossRef
  30. ↵
    Rahman, S. A. et al. The AUGIS Survival Predictor: Prediction of Long-Term and Conditional Survival After Esophagectomy Using Random Survival Forests. Ann Surg 277, 267–274 (2023). doi:10.1097/sla.0000000000004794
    OpenUrlCrossRef
  31. ↵
    Kwak, S. et al. Markers of Myocardial Damage Predict Mortality in Patients With Aortic Stenosis. J Am Coll Cardiol 78, 545–558 (2021). doi:10.1016/j.jacc.2021.05.047
    OpenUrlCrossRefPubMed
  32. ↵
    Ishwaran, H., Kogalur, U. B., Blackstone, E. H. & Lauer, M. S. Random survival forests. The Annals of Applied Statistics 2, 841–860, 820 (2008).
    OpenUrl
  33. ↵
    Wilson, P. W. et al. Prediction of incident diabetes mellitus in middle-aged adults: the Framingham Offspring Study. Arch Intern Med 167, 1068–1074 (2007). doi:10.1001/archinte.167.10.1068
    OpenUrlCrossRefPubMedWeb of Science
  34. ↵
    Vickers, A. J., van Calster, B. & Steyerberg, E. W. A simple, step-by-step guide to interpreting decision curve analysis. Diagn Progn Res 3, 18 (2019). doi:10.1186/s41512-019-0064-7
    OpenUrlCrossRefPubMed
  35. ↵
    Fan, X. et al. Noninvasive radiomics model reveals macrophage infiltration in glioma. Cancer Lett 573, 216380 (2023). doi:10.1016/j.canlet.2023.216380
    OpenUrlCrossRef
  36. ↵
    Bragg, F. et al. Predictive value of circulating NMR metabolic biomarkers for type 2 diabetes risk in the UK Biobank study. BMC Med 20, 159 (2022). doi:10.1186/s12916-022-02354-9
    OpenUrlCrossRef
  37. Mackey, R. H. et al. Lipoprotein particles and incident type 2 diabetes in the multi-ethnic study of atherosclerosis. Diabetes Care 38, 628–636 (2015). doi:10.2337/dc14-0645
    OpenUrlAbstract/FREE Full Text
  38. Bragg, F. et al. The role of NMR-based circulating metabolic biomarkers in development and risk prediction of new onset type 2 diabetes. Sci Rep 12, 15071 (2022). doi:10.1038/s41598-022-19159-8
    OpenUrlCrossRef
  39. ↵
    Bragg, F. et al. Circulating Metabolites and the Development of Type 2 Diabetes in Chinese Adults. Diabetes Care 45, 477–480 (2022). doi:10.2337/dc21-1415
    OpenUrlCrossRef
  40. ↵
    Morze, J. et al. Metabolomics and Type 2 Diabetes Risk: An Updated Systematic Review and Meta-analysis of Prospective Cohort Studies. Diabetes Care 45, 1013–1024 (2022). doi:10.2337/dc21-1705
    OpenUrlCrossRef
  41. ↵
    Lotta, L. A. et al. Genetic Predisposition to an Impaired Metabolism of the Branched-Chain Amino Acids and Risk of Type 2 Diabetes: A Mendelian Randomisation Analysis. PLoS Med 13, e1002179 (2016). doi:10.1371/journal.pmed.1002179
    OpenUrlCrossRefPubMed
  42. ↵
    Mahendran, Y. et al. Genetic evidence of a causal effect of insulin resistance on branched-chain amino acid levels. Diabetologia 60, 873–878 (2017). doi:10.1007/s00125-017-4222-6
    OpenUrlCrossRefPubMed
  43. ↵
    Li, J., Xi, F., Yu, W., Sun, C. & Wang, X. Real-Time Prediction of Sepsis in Critical Trauma Patients: Machine Learning-Based Modeling Study. JMIR Form Res 7, e42452 (2023). doi:10.2196/42452
    OpenUrlCrossRef
  44. ↵
    Li-Gao, R. et al. Assessment of reproducibility and biological variability of fasting and postprandial plasma metabolite concentrations using 1H NMR spectroscopy. PLoS One 14, e0218549 (2019). doi:10.1371/journal.pone.0218549
    OpenUrlCrossRef
Back to top
PreviousNext
Posted August 28, 2024.
Download PDF

Supplementary Material

Data/Code
Email

Thank you for your interest in spreading the word about medRxiv.

NOTE: Your email address is requested solely to identify you as the sender of this article.

Enter multiple addresses on separate lines or separate them with commas.
Nuclear magnetic resonance-based metabolomics with machine learning for predicting progression from prediabetes to diabetes
(Your Name) has forwarded a page to you from medRxiv
(Your Name) thought you would like to see this page from the medRxiv website.
CAPTCHA
This question is for testing whether or not you are a human visitor and to prevent automated spam submissions.
Share
Nuclear magnetic resonance-based metabolomics with machine learning for predicting progression from prediabetes to diabetes
Jiang Li, Yuefeng Yu, Ying Sun, Yanqi Fu, Wenqi Shen, Lingli Cai, Xiao Tan, Yan Cai, Ningjian Wang, Yingli Lu, Bin Wang
medRxiv 2024.05.14.24307378; doi: https://doi.org/10.1101/2024.05.14.24307378
Twitter logo Facebook logo LinkedIn logo Mendeley logo
Citation Tools
Nuclear magnetic resonance-based metabolomics with machine learning for predicting progression from prediabetes to diabetes
Jiang Li, Yuefeng Yu, Ying Sun, Yanqi Fu, Wenqi Shen, Lingli Cai, Xiao Tan, Yan Cai, Ningjian Wang, Yingli Lu, Bin Wang
medRxiv 2024.05.14.24307378; doi: https://doi.org/10.1101/2024.05.14.24307378

Citation Manager Formats

  • BibTeX
  • Bookends
  • EasyBib
  • EndNote (tagged)
  • EndNote 8 (xml)
  • Medlars
  • Mendeley
  • Papers
  • RefWorks Tagged
  • Ref Manager
  • RIS
  • Zotero
  • Tweet Widget
  • Facebook Like
  • Google Plus One

Subject Area

  • Endocrinology (including Diabetes Mellitus and Metabolic Disease)
Subject Areas
All Articles
  • Addiction Medicine (430)
  • Allergy and Immunology (755)
  • Anesthesia (221)
  • Cardiovascular Medicine (3288)
  • Dentistry and Oral Medicine (364)
  • Dermatology (277)
  • Emergency Medicine (479)
  • Endocrinology (including Diabetes Mellitus and Metabolic Disease) (1169)
  • Epidemiology (13359)
  • Forensic Medicine (19)
  • Gastroenterology (898)
  • Genetic and Genomic Medicine (5147)
  • Geriatric Medicine (481)
  • Health Economics (782)
  • Health Informatics (3264)
  • Health Policy (1140)
  • Health Systems and Quality Improvement (1190)
  • Hematology (429)
  • HIV/AIDS (1017)
  • Infectious Diseases (except HIV/AIDS) (14622)
  • Intensive Care and Critical Care Medicine (912)
  • Medical Education (476)
  • Medical Ethics (127)
  • Nephrology (522)
  • Neurology (4919)
  • Nursing (262)
  • Nutrition (727)
  • Obstetrics and Gynecology (882)
  • Occupational and Environmental Health (795)
  • Oncology (2519)
  • Ophthalmology (723)
  • Orthopedics (280)
  • Otolaryngology (347)
  • Pain Medicine (323)
  • Palliative Medicine (90)
  • Pathology (543)
  • Pediatrics (1299)
  • Pharmacology and Therapeutics (550)
  • Primary Care Research (556)
  • Psychiatry and Clinical Psychology (4205)
  • Public and Global Health (7499)
  • Radiology and Imaging (1704)
  • Rehabilitation Medicine and Physical Therapy (1011)
  • Respiratory Medicine (980)
  • Rheumatology (479)
  • Sexual and Reproductive Health (497)
  • Sports Medicine (424)
  • Surgery (547)
  • Toxicology (72)
  • Transplantation (235)
  • Urology (205)