RT Journal Article SR Electronic T1 Interpreting chest X-rays via CNNs that exploit disease dependencies and uncertainty labels JF medRxiv FD Cold Spring Harbor Laboratory Press SP 19013342 DO 10.1101/19013342 A1 Hieu H. Pham A1 Tung T. Le A1 Dat Q. Tran A1 Dat T. Ngo A1 Ha Q. Nguyen YR 2019 UL http://medrxiv.org/content/early/2019/11/29/19013342.abstract AB Chest radiography is one of the most common types of diagnostic radiology exams, which is critical for screening and diagnosis of many different thoracic diseases. Specialized algorithms have been developed to detect several specific pathologies such as lung nodule or lung cancer. However, accurately detecting the presence of multiple diseases from chest X-rays (CXRs) is still a challenging task. This paper presents a supervised multi-label classification framework based on deep convolutional neural networks (CNNs) for predicting the risk of 14 common thoracic diseases. We tackle this problem by training state-of-the-art CNNs that exploit dependencies among abnormality labels. We also propose to use the label smoothing technique for a better handling of uncertain samples, which occupy a significant portion of almost every CXR dataset. Our model is trained on over 200,000 CXRs of the recently released CheXpert dataset and achieves a mean area under the curve (AUC) of 0.940 in predicting 5 selected pathologies from the validation set. This is the highest AUC score yet reported to date. The proposed method is also evaluated on the independent test set of the CheXpert competition, which is composed of 500 CXR studies annotated by a panel of 5 experienced radiologists. The performance is on average better than 2.6 out of 3 other individual radiologists with a mean AUC of 0.930, which ranks first on the CheXpert leaderboard at the time of writing this paper.Competing Interest StatementThe authors have declared no competing interest.Funding StatementThis research was supported by the Vingroup Big Data Institute (VinBDI). 380 The authors gratefully acknowledge Jeremy Irvin from the Machine Learning Group, Stanford University for helping us evaluate the proposed method on the hidden test set of CheXpert.Author DeclarationsAll relevant ethical guidelines have been followed; any necessary IRB and/or ethics committee approvals have been obtained and details of the IRB/oversight body are included in the manuscript.YesAll necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable.YesThe dataset used in this research is available at https://stanfordmlgroup.github.io/competitions/chexpert/ https://stanfordmlgroup.github.io/competitions/chexpert/