π€ AI Summary
This study addresses the challenges of context overflow, tool non-stationarity, and heterogeneous multimodal feedback encountered by AI agents operating within large-scale dynamic tool ecosystems. To this end, we propose a general framework grounded in a self-evolving action space. Methodologically, we construct a unified action layer that optimizes hybrid action spaces through hierarchical progressive retrieval and test-time reliability pruning. Additionally, a heterogeneous observation grounding module is designed to align multimodal information and facilitate cost-balanced task orchestration. Experimental evaluations demonstrate that the proposed approach achieves state-of-the-art performance on both the LiveMCPBench and OSMCP benchmarks. Notably, on OSMCP, our method attains a 77.27% success rate within merely 50 steps, effectively doubling operational efficiency compared to existing baselines.
π Abstract
As large language models (LLMs) advance, AI agents are increasingly deployed in open-world environments to tackle complex sequential tasks (e.g., document processing, cross-application collaboration), relying heavily on actions ranging from GUI operations to semantic APIs. However, three core challenges persist: the "scale dilemma" of massive tool ecosystems exceeding LLM context windows, the "non-stationarity" of tool quality due to updates or outages, and the "heterogeneity" of feedback formats (pixels, text, structured data) creating information silos. To address these, we propose AnyAct, a universal action layer that unifies available capabilities into a self-evolving action space, enabling agents to operate efficiently and reliably in large-scale, dynamic tool ecosystems. AnyAct's core design focuses on two objectives: constructing this action space via hierarchical progressive retrieval (filtering task-relevant actions) and test-time reliability evolution (pruning unreliable actions), and enabling reliability-aware action orchestration through a heterogeneous observation grounding module that unifies multi-modal feedback. Additionally, it defines a hybrid action space (primitive + semantic actions) and optimizes for a balance between task success rate and execution cost. Evaluations on LiveMCPBench and OSMCP (a new benchmark we developed for multi-granularity action collaboration) demonstrate state-of-the-art performance. AnyAct delivers substantial performance gains over baseline methods across various LLM base models on LiveMCPBench and improvements are particularly notable for models with constrained native capabilities. On OSMCP, it achieves 77.27% overall success with only 50 steps, which is half the steps required by most competitors.