Abstract
Background and objective The F1 score and its generalized Fβ score are widely used to evaluate machine learning and artificial intelligence (AI) models in healthcare, particularly for imbalanced clinical datasets. In practice, competing prediction models are commonly evaluated on the same patient cohort, resulting in correlated classifier decisions. However, existing approaches for statistical inference of F1 -related metrics typically assume independent classifier decisions or lack integrated procedures for comparative evaluation, power analysis, and sample size determination in paired validation studies.
Methods We propose psF1pair, a unified framework for confidence interval estimation, hypothesis testing, and power and sample size calculation for comparative F1 and Fβ scores under paired evaluation designs. Dependence between classifiers is modeled using a four-component multinomial representation of the joint decision process, allowing explicit estimation and incorporation of classifier correlations commonly encountered when AI models are evaluated on the same patient cohort. Exact distributions are used for small sample settings, while asymptotic approximations are employed for computational efficiency in large studies.
Results Simulation studies demonstrated that the proposed confidence intervals achieved nominal coverage probabilities across a wide range of sample sizes and correlation settings. Estimated power closely agreed with empirical power, with discrepancies generally below 3%. Compared with existing methods, psF1pair showed competitive or superior statistical power while maintaining appropriate type I error rates across a broad range of scenarios. Applications to skin cancer classification and breast cancer screening demonstrated that accounting for classifier correlation produced narrower confidence intervals and improved statistical efficiency.
Conclusions psF1pair provides a practical and rigorous framework for evaluation and study planning of medical AI systems using F1 and Fβ metrics. The method supports comparative benchmarking, uncertainty quantification, and sample size determination for future validation studies. An open-source R package is freely available.
Highlights
Unified statistical inference and study-design framework for comparative F1 and Fβ scores under paired evaluation designs.
Explicit modeling of dependence between classifiers evaluated on the same dataset using a multinomial framework.
Open-source R package supporting comparative evaluation and validation studies of predictive and AI models.
Competing Interest Statement
The authors have declared no competing interest.
Author Declarations
I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.
Yes
I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.
Yes
I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).
Yes
I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.
Yes
Footnotes
The Abstract, Introduction, and Discussion sections were revised, and a real-world application was added.
Data Availability
All data produced in the present work are contained in the manuscript.





