ABSTRACT
Despite the abundance of multi-modal data, suitable statistical models that can improve our understanding of diseases with genetic underpinnings are challenging to develop. Here we present SparseGMM, a novel statistical approach for gene regulatory network discovery. SparseGMM uniquely uses latent variable modeling with sparsity constraints regulators to learn gaussian mixtures from multi-omic data. By combining co-expression patterns with a Bayesian framework, sparseGMM quantitatively measures confidence in regulators and uncertainty in target gene assignment by computing gene entropy. We apply SparseGMM to liver cancer and normal liver tissue data and evaluate the discovered gene modules in an independent scRNA-seq dataset. sparseGMM identifies PROCR as a regulator of angiogenesis, and PDCD1LG2 and HNF4A as regulators of immune response and blood coagulation in cancer, respectively. Additionally, we show that more genes have significantly higher entropy in cancer compared to normal liver; among high entropy genes are key multifunctional components shared by critical pathways, such as p53 and estrogen signaling.
Software availability The software is available at https://hub.docker.com/r/shaimaabakr/sparse_gmm
One-sentence summary A novel statistical approach for gene regulatory network discovery recovers modules and corresponding regulators of diverse normal liver functions, important liver cancer processes, as well as shared biology between liver cancer and normal tissue.
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
This work supported by the National Cancer Institute (NCI) under awards: R01 CA260271, U01 CA217851 and U01 CA199241.
Author Declarations
I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.
Yes
The details of the IRB/oversight body that provided approval or exemption for the research described are given below:
TCGA: https://portal.gdc.cancer.gov/ GTEx https://gtexportal.org/home/
I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.
Yes
I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).
Yes
I have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable.
Yes
Data Availability
All data used in the present study were openly available to the public before the initiation of the study and can be accessed online at: https://portal.gdc.cancer.gov/ https://gtexportal.org/home/