Abstract
Purpose To test whether disease-aware adversarial attacks can disrupt demographic classification in color fundus photographs (CFPs) while retaining glaucoma classification and whether effects transfer across model architectures.
Methods This retrospective study included 13,959 CFPs from 4,271 patients. Vision Transformer (ViT) models classified glaucoma, race, sex, and ethnicity. Standard and disease-aware FGSM, PGD, C&W, and diffusion attacks changed demographic predictions; disease-aware attacks added a glaucoma classification loss term. Fixed ViT-generated perturbations were tested on ResNet50 and EfficientNetB0. Performance was measured using area under the receiver operating characteristic curve (AUC), accuracy, and image-similarity metrics.
Results Baseline glaucoma AUCs were 0.958–0.963 and demographic AUCs were 0.955–0.992. Disease-aware attacks better preserved glaucoma classification while strongly altering demographic predictions. Disease-aware PGD preserved sex-cohort glaucoma AUC at 0.909 (95% CI, 0.894–0.922) while sex AUC fell to 0.000 (95% CI, 0.000–0.000). Disease-aware diffusion preserved ethnicity-cohort glaucoma AUC at 0.946 (95% CI, 0.936–0.956) while ethnicity AUC was 0.001 (95% CI, 0.000–0.002). Demographic changes transferred poorly to ResNet50 and EfficientNetB0.
Conclusions For PGD, C&W, and diffusion, disease-aware attacks retained more glaucoma classification performance than standard attacks while strongly disrupting targeted demographic prediction. Weak cross-architecture transfer indicates model-specific effects rather than demographic information removal.
Translational Relevance Methods intended to reduce sensitive information in retinal images could support multicenter AI development, but failure of one demographic classifier does not establish de-identification. Cross-model evaluation while preserving disease classification may provide a more reliable framework for clinical data sharing.
Competing Interest Statement
The authors have declared no competing interest.
Author Declarations
I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.
Yes
The details of the IRB/oversight body that provided approval or exemption for the research described are given below:
No human subjects were included in this study. This retrospective study was approved by the Mass General Brigham Institutional Review Board (IRB) and conducted in accordance with the principles outlined in the Declaration of Helsinki. The IRB also deemed informed consent to be waived for this study.
I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.
Yes
I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).
Yes
I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.
Yes
Footnotes
Address for Reprints: Mengyu Wang, PhD, Harvard Ophthalmology AI Lab, Schepens Eye Research Institute of Massachusetts Eye and Ear 20 Staniford Street Boston, MA 02114, United States, Email: mengyu_wang{at}meei.harvard.edu
This version has been revised to clarify the interpretation and generalizability of adversarial demographic prediction disruption in retinal fundus images. We added cross architecture evaluation using ResNet50 and EfficientNetB0 to determine whether perturbations generated for a Vision Transformer generalized to unseen model architectures. We also revised the interpretation of very low binary demographic AUC values, emphasizing that near zero AUC can reflect classifier inversion rather than removal of demographic information. The Introduction and Discussion were updated to distinguish targeted classifier failure from true de identification and to place the findings in the context of prior adversarial privacy and concept removal literature. We added discussion of recent retinal privacy work using adversarial perturbations for patient identity protection and highlighted differences in privacy endpoints and evaluation. The manuscript was also revised for clearer terminology, more concise framing, updated references, and improved consistency throughout.
Data Availability
The dataset used in this study is institutionally restricted due to patient privacy considerations. Code for all experiments, including model training and adversarial attack generation, is available from the corresponding author upon request.





