AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss

2026-08-11Computer Vision and Pattern Recognition

Computer Vision and Pattern Recognition
AI summary

The authors found that using a fixed feature space to measure differences between real and generated data can lead to problems where the model appears to improve but actually doesn't get better visually. To fix this, they created a method called Adversarial Fréchet Distance (AdvFD), which adds a learnable feature space that actively tries to highlight differences between real and fake data. The generator then learns to minimize these highlighted differences, making the training more effective. They also introduced a technique called real-feature whitening to keep the learning stable. Their experiments show AdvFD helps improve generator training across different setups.

Fréchet distancegenerator post-trainingfeature spaceadversarial learningdiscrepancyreal-feature whiteningmin-max optimizationJiT backbonepMF backbonediffusion models
Authors
Mingju Gao, Jingkai Zhou, Kun Gai, Changqian Yu, Hao Tang
Abstract
Fréchet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-level diffusion and flow-matching losses. However, directly optimizing Fréchet objectives can cause Fréchet hacking. The target metrics keep improving, but visual quality and Fréchet alignment in other feature spaces may stagnate or deteriorate. We attribute this failure to the static pretrained feature spaces used by existing Fréchet losses. These feature spaces provide incomplete and fixed views of the differences between real and generated distributions. To address this limitation, we propose Adversarial Fréchet Distance (AdvFD), which complements the static representation targets in FD-Loss with a calibrated adversarially learned representation. AdvFD augments the original static Fréchet objective with a learnable representation that adversarially maximizes the Fréchet discrepancy between real and generated samples, while the generator minimizes the same discrepancy in the resulting adaptive feature space. To prevent the adversarial representation from trivially increasing the objective through feature amplification, we further introduce real-feature whitening, which normalizes its scale and covariance geometry and stabilizes the min--max optimization. Extensive experiments show that AdvFD consistently improves one-step generator post-training across both JiT and pMF backbones and across different model scales.