Research arXiv cs.AI

C$^3$PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models

MLLMbenchmarkcross-modal reasoningC3PO

C³PO (arXiv:2608.05381) targets two key abilities: information composition, which requires fusing dispersed evidence across modalities, and counterfactual conflict, which tests how well models handle deliberate contradictions between modalities. The benchmark spans video, audio, image, and text, providing a broader evaluation than many existing benchmarks. The authors argue that current multimodal LLMs are heavily biased toward a dominant modality, leading to brittle cross-modal reasoning, and C³PO aims to expose these weaknesses. With 3,404 samples, it offers a systematic tool for measuring progress in multimodal reasoning and may drive improvements in model training and architecture to achieve more balanced cross-modal understanding.

Read original →

← Back to home