RIG-RoPE: Relation- and Instance-Gated Rotary Positional Encoding with Duration-Aware Temporal Coordinates
The paper identifies two key limitations of static multidimensional position assignment in interleaved multimodal contexts, such as height/width rotations being applied to token pairs that may not be spatially related. RIG-RoPE introduces relation- and instance-gating mechanisms to dynamically adjust positional encoding based on token relationships and instance boundaries, alongside duration-aware temporal coordinates to better handle time-varying content. This work builds on M-RoPE, which splits positional channels into temporal, height, and width subspaces. The proposed method could improve performance on multimodal tasks that involve interleaved images, videos, and text, though experimental results are not yet detailed in the provided abstract.