AI summaryⓘ
The authors study multimodal language models that are usually built assuming all types of data (modalities) seen in training are also available at test time. They focus on cases where some extra data types are only present during training but not available later, which is common in real life. To handle this, they propose a new method called Mixture of Probes (MoP), which lets the model separately learn signals specific to each modality and general patterns across them. They also introduce a training technique, MoP Cross-modal Training, to improve learning and avoid problems during training. The authors show MoP improves performance across multiple tasks and data types, proving extra training data can help even if missing at test time.
Multimodal Large Language ModelsPrivileged Modality SettingModality-Specific SignalsModality-General SignalsIntermediate RepresentationsCross-modal TrainingProbe Disentanglement LossAuxiliary ModalitiesInferenceRepresentation Learning
Authors
Dominick Reilly, Qiyu Wu, Hiromi Wakaki, Srijan Das, Yuki Mistufuji
Abstract
Multimodal Large Language Models (MLLMs) are typically designed under the assumption that all modalities available during training will also be accessible at inference. However, many real-world settings violate this assumption, requiring models to operate under a privileged modality setting, where auxiliary modalities are available only during training. While these modalities contain valuable information, existing MLLMs largely fail to leverage them effectively, as they treat modalities as interchangeable inputs rather than sources of complementary supervision. We propose Mixture of Probes (MoP), a novel framework that disentangles modality-specific and modality-general signals within the MLLM, allowing the model to preserve modality-dependent structure while learning transferable representations across modalities. At its core, MoP achieves this through a structured probing mechanism that extracts and organizes information from intermediate representations of a shared modality encoder, rather than relying only on final-layer alignment as done in existing MLLMs. To support this disentanglement, we further introduce MoP Cross-modal Training (MoP-X), a training strategy for MoP centered around a probe disentanglement loss that prevents probe collapse and encourages cross-modal learning. We evaluate MoP across two domains spanning eight tasks and four modalities under a comprehensive evaluation protocol tailored to the privileged modality setting, where each modality is independently treated as the sole input at inference time. MoP consistently outperforms strong MLLM baselines, achieving up to 65% relative improvement, demonstrating that auxiliary modalities, even when unavailable at inference, can provide substantial gains when effectively leveraged during training. Code, model checkpoints, and evaluation protocols will be made available at https://github.com/Sony/MoP.