Skip to main content
medRxiv
  • Home
  • About
  • Submit
  • ALERTS / RSS
Advanced Search

A Unified Framework for Statistical Inference and Study Design of Comparative F1 and Fβ Scores under Paired Classifier Evaluations

View ORCID ProfileChih-Yuan Hsu, Qi Liu, Yu Shyr
doi: https://doi.org/10.64898/2026.07.15.26358166
Chih-Yuan Hsu
1Department of Biostatistics, Vanderbilt University Medical Center, Nashville, TN 37203, USA
2Center for Quantitative Sciences, Vanderbilt University Medical Center, Nashville, TN 37203, USA
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Chih-Yuan Hsu
Qi Liu
1Department of Biostatistics, Vanderbilt University Medical Center, Nashville, TN 37203, USA
2Center for Quantitative Sciences, Vanderbilt University Medical Center, Nashville, TN 37203, USA
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • For correspondence: qi.liu{at}vumc.org yu.shyr{at}vumc.org
Yu Shyr
1Department of Biostatistics, Vanderbilt University Medical Center, Nashville, TN 37203, USA
2Center for Quantitative Sciences, Vanderbilt University Medical Center, Nashville, TN 37203, USA
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • For correspondence: qi.liu{at}vumc.org yu.shyr{at}vumc.org
  • Abstract
  • Full Text
  • Info/History
  • Metrics
  • Supplementary material
  • Data/Code
  • Preview PDF
Loading

Abstract

Background and objective The F1 score and its generalized Fβ score are widely used to evaluate machine learning and artificial intelligence (AI) models in healthcare, particularly for imbalanced clinical datasets. In practice, competing prediction models are commonly evaluated on the same patient cohort, resulting in correlated classifier decisions. However, existing approaches for statistical inference of F1 -related metrics typically assume independent classifier decisions or lack integrated procedures for comparative evaluation, power analysis, and sample size determination in paired validation studies.

Methods We propose psF1pair, a unified framework for confidence interval estimation, hypothesis testing, and power and sample size calculation for comparative F1 and Fβ scores under paired evaluation designs. Dependence between classifiers is modeled using a four-component multinomial representation of the joint decision process, allowing explicit estimation and incorporation of classifier correlations commonly encountered when AI models are evaluated on the same patient cohort. Exact distributions are used for small sample settings, while asymptotic approximations are employed for computational efficiency in large studies.

Results Simulation studies demonstrated that the proposed confidence intervals achieved nominal coverage probabilities across a wide range of sample sizes and correlation settings. Estimated power closely agreed with empirical power, with discrepancies generally below 3%. Compared with existing methods, psF1pair showed competitive or superior statistical power while maintaining appropriate type I error rates across a broad range of scenarios. Applications to skin cancer classification and breast cancer screening demonstrated that accounting for classifier correlation produced narrower confidence intervals and improved statistical efficiency.

Conclusions psF1pair provides a practical and rigorous framework for evaluation and study planning of medical AI systems using F1 and Fβ metrics. The method supports comparative benchmarking, uncertainty quantification, and sample size determination for future validation studies. An open-source R package is freely available.

Highlights

  • Unified statistical inference and study-design framework for comparative F1 and Fβ scores under paired evaluation designs.

  • Explicit modeling of dependence between classifiers evaluated on the same dataset using a multinomial framework.

  • Open-source R package supporting comparative evaluation and validation studies of predictive and AI models.

Competing Interest Statement

The authors have declared no competing interest.

Author Declarations

I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.

Yes

I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.

Yes

I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).

Yes

I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.

Yes

Footnotes

  • The Abstract, Introduction, and Discussion sections were revised, and a real-world application was added.

Data Availability

All data produced in the present work are contained in the manuscript.

https://github.com/cyhsuTN/psF1pair

Funder Information Declared

National Institutes of Health, P50CA236733, U54CA163072
VA Tennessee Valley Healthcare System, VUMC120842
Lifespan Rhode Island Hospital, U01CA287008
Copyright 
The copyright holder for this preprint is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. All rights reserved. No reuse allowed without permission.
Back to top
PreviousNext
Posted August 22, 2026.
Download PDF

Supplementary Material

Data/Code
Email

Thank you for your interest in spreading the word about medRxiv.

NOTE: Your email address is requested solely to identify you as the sender of this article.

Enter multiple addresses on separate lines or separate them with commas.
A Unified Framework for Statistical Inference and Study Design of Comparative F1 and Fβ Scores under Paired Classifier Evaluations
(Your Name) has forwarded a page to you from medRxiv
(Your Name) thought you would like to see this page from the medRxiv website.
CAPTCHA
This question is for testing whether or not you are a human visitor and to prevent automated spam submissions.
Share
A Unified Framework for Statistical Inference and Study Design of Comparative F1 and Fβ Scores under Paired Classifier Evaluations
Chih-Yuan Hsu, Qi Liu, Yu Shyr
medRxiv 2026.07.15.26358166; doi: https://doi.org/10.64898/2026.07.15.26358166
Twitter logo Facebook logo LinkedIn logo Mendeley logo
Citation Tools
A Unified Framework for Statistical Inference and Study Design of Comparative F1 and Fβ Scores under Paired Classifier Evaluations
Chih-Yuan Hsu, Qi Liu, Yu Shyr
medRxiv 2026.07.15.26358166; doi: https://doi.org/10.64898/2026.07.15.26358166

Citation Manager Formats

  • BibTeX
  • Bookends
  • EasyBib
  • EndNote (tagged)
  • EndNote 8 (xml)
  • Medlars
  • Mendeley
  • Papers
  • RefWorks Tagged
  • Ref Manager
  • RIS
  • Zotero
  • Tweet Widget
  • Facebook Like
  • Google Plus One

Subject Area

  • Dermatology
Subject Areas
All Articles
  • Addiction Medicine (617)
  • Allergy and Immunology (901)
  • Anesthesia (332)
  • Cardiovascular Medicine (4814)
  • Dentistry and Oral Medicine (479)
  • Dermatology (415)
  • Emergency Medicine (650)
  • Endocrinology (including Diabetes Mellitus and Metabolic Disease) (1627)
  • Epidemiology (15949)
  • Forensic Medicine (33)
  • Gastroenterology (1203)
  • Genetic and Genomic Medicine (7048)
  • Geriatric Medicine (733)
  • Health Economics (1065)
  • Health Informatics (5047)
  • Health Policy (1434)
  • Health Systems and Quality Improvement (1767)
  • Hematology (585)
  • HIV/AIDS (1344)
  • Infectious Diseases (except HIV/AIDS) (16291)
  • Intensive Care and Critical Care Medicine (1173)
  • Medical Education (668)
  • Medical Ethics (154)
  • Nephrology (724)
  • Neurology (7274)
  • Nursing (366)
  • Nutrition (1082)
  • Obstetrics and Gynecology (1241)
  • Occupational and Environmental Health (1005)
  • Oncology (3610)
  • Ophthalmology (1052)
  • Orthopedics (397)
  • Otolaryngology (456)
  • Pain Medicine (470)
  • Palliative Medicine (139)
  • Pathology (708)
  • Pediatrics (1798)
  • Pharmacology and Therapeutics (737)
  • Primary Care Research (767)
  • Psychiatry and Clinical Psychology (5897)
  • Public and Global Health (9757)
  • Radiology and Imaging (2423)
  • Rehabilitation Medicine and Physical Therapy (1453)
  • Respiratory Medicine (1249)
  • Rheumatology (641)
  • Sexual and Reproductive Health (771)
  • Sports Medicine (578)
  • Surgery (777)
  • Toxicology (107)
  • Transplantation (305)
  • Urology (291)