SARATOV FALL MEETING SFM 

© 2026 All Rights Reserved

The Influence of Feature-Space Similarity Between Original and Synthetic Dermoscopic Images on Classification Performance

Nikita K. Zakharov1, Irina A. Matveeva1; 1Samara National Research University, Samara, Russia

Abstract

Synthetic images are increasingly used to mitigate class imbalance in dermoscopic classification, but visual realism and feature-space proximity do not necessarily imply practical value. This study asks whether synthetic dermoscopic images selected by geometry in feature space provide more useful information than a source-matched repeated presentation of real images. Using HAM10000 with lesion-level separation to prevent leakage, we trained ConvNeXt-S under identical compute budgets and random initializations, and evaluated performance with macro-AUPRC, macro-F1, Matthews correlation coefficient (MCC), calibration error, and hierarchical bootstrap confidence intervals. In exploratory analysis, the effect depended on the geometric layer of synthetic selection: strictly in-domain synthetic samples improved MCC, whereas samples drawn from a random remainder subset reduced it. However, in the confirmatory setting, the positive shift in macro-F1 was accompanied by lower macro-AUPRC, worse melanoma AUPRC, and degraded calibration. Multi-model diagnostics showed higher diversity but lower precision, density, and coverage for synthetic data in most encoder-class combinations, indicating that diversity alone was not sufficient. Causal decomposition further revealed that a substantial part of the ranking loss arose already during preprocessing, especially square cropping, while the latent VAE step and single UNet denoising step contributed additional degradation. Taken together, the results show that feature-space closeness, perceptual plausibility, and improved generative metrics are not sufficient criteria for judging the utility of synthetic medical images. A practical evaluation protocol should compare synthetic data against matched real-image repetition, preserve lesion-level separation, and assess ranking, thresholding, and calibration separately.

Speaker

Nikita K. Zakharov
Samara National Research University
Russia

Discussion

Ask question