Research arXiv cs.CL

Seeing Through Conflicts: Improving Instruction Hierarchy Alignment in Vision-Language Models

instruction hierarchyvision-language modelsreinforcement learningmultimodal safety

Instruction hierarchy alignment teaches language models to prioritize higher-level instructions when inputs conflict. While prior work has focused mainly on text-only settings, vision-language models introduce new challenges: instructions may be embedded in images, split across modalities, visually transformed, or encountered during agentic tasks.

The authors frame multimodal IH alignment as a reasoning problem. They train VLMs using reinforcement learning with rule-based rewards, comparing text-only, image-only, and mixed-modality supervision.

They find that text-only IH training partially transfers to multimodal attacks, but fails when models must decode, reconstruct, or reason over instructions across modalities. Image-based training improves robustness beyond text-only supervision, while mixed-modality training performs best overall. The benefits generalize beyond the synthetic typographic training setting to real-image and web-agent safety tasks, while largely preserving general multimodal capability. This shows that lightweight, verifiable supervision can meaningfully improve VLM robustness under adversarial, cross-modal, and interactive instruction conflicts.

Read original →

← Back to home