TY - JOUR T1 - Unsupervised Learning for Large Scale Data: The ATHLOS Project JF - medRxiv DO - 10.1101/2021.04.01.21254751 SP - 2021.04.01.21254751 AU - Petros Barmpas AU - Sotiris Tasoulis AU - Aristidis G. Vrahatis AU - Panagiotis Anagnostou AU - Spiros Georgakopoulos AU - Matthew Prina AU - José Luis Ayuso-Mateos AU - Jerome Bickenbach AU - Ivet Bayes AU - Martin Bobak AU - Francisco Félix Caballero AU - Somnath Chatterji AU - Laia Egea-Cortés AU - Esther García-Esquinas AU - Matilde Leonardi AU - Seppo Koskinen AU - Ilona Koupil AU - Andrzej Pająk AU - Martin Prince AU - Warren Sanderson AU - Sergei Scherbov AU - Abdonas Tamosiunas AU - Aleksander Galas AU - Josep MariaHaro AU - Albert Sanchez-Niubo AU - Vassilis P. Plagianakos AU - Demosthenes Panagiotakos Y1 - 2021/01/01 UR - http://medrxiv.org/content/early/2021/04/06/2021.04.01.21254751.abstract N2 - Recent technological advancements in various domains, such as the biomedical and health, offer a plethora of big data for analysis. Part of this data pool is the experimental studies that record various and several features for each instance. It creates datasets having very high dimensionality with mixed data types, with both numerical and categorical variables. On the other hand, unsupervised learning has shown to be able to assist in high-dimensional data, allowing the discovery of unknown patterns through clustering, visualization, dimensionality reduction, and in some cases, their combination. This work highlights unsupervised learning methodologies for large-scale, high-dimensional data, providing the potential of a unified framework that combines the knowledge retrieved from clustering and visualization. The main purpose is to uncover hidden patterns in a high-dimensional mixed dataset, which we achieve through our application in a complex, real-world dataset. The experimental analysis indicates the existence of notable information exposing the usefulness of the utilized methodological framework for similar high-dimensional and mixed, real-world applications.Competing Interest StatementThe authors have declared no competing interest.Funding StatementThis work is supported by the ATHLOS (Aging Trajectories of Health: Longitudinal Opportunities and Synergies) project, funded by the European Union's Horizon 2020 Research and Innovation Program under grant agreement number 635316.Author DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesThe details of the IRB/oversight body that provided approval or exemption for the research described are given below:Does not apply in our workAll necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable.YesData sharing is not applicable to this article ER -