Omni-IO Skills: Harnessing Your Agent Omni-Native
General-purpose agents can already plan, reason, and act over long horizons, but their production capabilities remain fragmented across text, images, audio, video, documents, 3D assets, and code. Extending a foundation model to additional modalities ties capability growth to costly model updates, while assembling specialist models and tools leaves unresolved how procedures, dependencies, intermediate assets, and cross-turn revisions should be coordinated.
The arXiv paper (2609.31847v1) presents Omni-IO Skills, a plug-and-play Agent Harness that aims to make existing agents omni-native. It combines hierarchical Skills, a standardized multimodal execution interface, dependency-aware orchestration, and a persistent Asset Registry. Multi-asset workflows are represented as Declare Execution Graphs, which schedule independent operations concurrently and register successful outputs for downstream and cross-turn reuse across replaceable execution backends. Its 27 Skills cover 38 representative tasks spanning seven artifact modalities and four capability families: understanding, generation, reasoning, and retrieval.
On UniM-90, the harness raises the input-support rates of GPT-5.6 Sol and Claude Sonnet 5 from 40.00% and 38.89% to 100%. It also increases the relative Semantic--Quality Coupled Score from 26.99 to 74.94 for GPT-5.6 Sol and from 27.82 to 77.78 for Claude Sonnet 5, while Strict Structure Score reaches 100.00 and 99.78, respectively.
The authors conclude that these results establish harness-level capability composition as a practical route to broad, evolvable Omni systems without changing the host agent's reasoning core.