Speaker
Description
Synthetic biometric data is attractive when privacy constraints, annotation cost, and limited subject availability restrict the collection of large real datasets. This paper presents a diffusion-first synthetic ear generation and benchmarking framework. The proposed pipeline uses Stable Diffusion 2.1 image-to-image synthesis: real ear crops define the source identity structure, multiple candidates are sampled per planned variant, candidates are scored by an ear-recognition encoder, and accepted images are filtered by identity consistency, visual quality, duplicate risk, and privacy diagnostics before verification-oriented adaptation. The completed diffusion run generated 1,048 candidates from 131 source identities and retained 169 images from 90 identities. Under the fixed template-quality protocol, base/adapted score fusion obtains 17.89\% EER, 6.09\% template EER, and 52.74\% Rank-1 on \earvn; 7.34\%, 0.69\%, and 93.47\% on AMI; 4.02\%, 2.76\%, and 93.83\% on IITDelhi; and 33.53\%, 17.88\%, and 7.00\% on UERC. Against the strongest listed non-ours synthetic baseline per target, the proposed diffusion method reduces mean EER from 25.92\% to 15.69\%, a 39.5\% relative reduction. A few-real ablation shows that, with the accepted diffusion set fixed, adding two real images per identity gives the best mean EER (15.61\%) and eight real images per identity gives the best mean Rank-1 (61.76\%).