Learning the Target Priors Before Image Translation: A Decoupled Training Paradigm for Cross-Modal Image Translation in Remote Sensing

2026-08-28Computer Vision and Pattern Recognition

Computer Vision and Pattern Recognition
AI summary

The authors address how to translate images between different types of remote sensing data, like radar to regular photos, while keeping the original content accurate and making the output look realistic. They find that learning the general style of the target images (target prior) should be done separately from learning how to match the specific source image (conditional dependence). Their method, LTP-BIT, first learns the target style from many unpaired images, then adapts to the source images with fewer parameters. This approach works better and needs less paired training data than previous methods.

cross-modal image translationremote sensingtarget priorconditional dependencegenerative modelpaired dataSAR-to-RGBparameter-efficientimage translation benchmarks
Authors
Keyan Hu, Mingtao Wang, Ziyu Zhou, Tiandong Shi, Haifeng Li, Ji Qi, Chao Tao
Abstract
Cross-modal image translation in remote sensing must preserve source-observed content while matching the target-domain distribution. Existing methods jointly learn the target prior and cross-modal dependence from scarce paired data, overlooking a key asymmetry: only the latter intrinsically requires cross-modal correspondence. We formalize this distinction through conditional-score and denoising-risk analyses and propose Learning the Target Priors Before Image Translation (LTP-BIT), a prior-first paradigm that decouples the two learning tasks. LTP-BIT first learns a target-domain generative prior from large-scale unpaired imagery, then retains the pretrained backbone weights and learns source-conditioned control through P-DART, a parameter-efficient dual-stream architecture. Controlled experiments show that prior matching and scaling primarily improve target-domain realism, whereas instance fidelity relies more strongly on conditional adaptation. LTP-BIT achieves state-of-the-art performance across SAR-to-RGB and NIR-to-RGB benchmarks using only 9.81% task-specific parameters. On QXS-SAROPT, it retains near-full-data instance fidelity with only 25% of the paired samples.