🤖 AI Summary
This study addresses the fragmentation of multimodal capabilities and cross-modal coordination challenges in general-purpose agents by proposing a plug-and-play Agent Harness framework. Methodologically, it introduces a hierarchical skill architecture, standardized interfaces, and a persistent asset registry to achieve native support across all modalities. Furthermore, declarative execution graphs and dependency-aware orchestration mechanisms are developed to extend multimodal capabilities without updating underlying models, while supporting swappable execution backends. Evaluated on the UniM-90 benchmark, the proposed framework elevates the input modality coverage of mainstream models to 100% and yields significant improvements in semantic quality scores. This work provides an efficient and scalable paradigm for constructing unified multimodal agents.
📝 Abstract
General-purpose agents can plan, reason, and act over long horizons, yet their production capabilities remain fragmented across text, images, audio, video, documents, 3D assets, and code. Extending a foundation model to additional modalities ties capability growth to costly model updates, while assembling specialist models and tools leaves unresolved how procedures, dependencies, intermediate assets, and cross-turn revisions should be coordinated. We present Omni-IO Skills, a plug-and-play Agent Harness that makes existing agents omni-native through hierarchical Skills, a standardized multimodal execution interface, dependency-aware orchestration, and a persistent Asset Registry. Multi-asset workflows are represented as Declare Execution Graphs, which schedule independent operations concurrently and register successful outputs for downstream and cross-turn reuse across replaceable execution backends. Its 27 Skills cover 38 representative tasks spanning seven artifact modalities and four capability families: understanding, generation, reasoning, and retrieval. On UniM-90, the harness raises the input-support rates of GPT-5.6 Sol and Claude Sonnet 5 from 40.00% and 38.89% to 100%, while increasing relative Semantic--Quality Coupled Score from 26.99 to 74.94 and from 27.82 to 77.78, respectively; Strict Structure Score reaches 100.00 and 99.78. These results establish harness-level capability composition as a practical route to broad, evolvable Omni systems without changing the host agent's reasoning core.