Don't Blame the Model, Verify the Data: An Evaluation of SMT-based Dataset Verification
研究通过SMT求解方法验证高风险ML系统的数据集质量,评估了数据属性类型、规范风格和编码策略对验证性能的影响。
研究通过SMT求解方法验证高风险ML系统的数据集质量,评估了数据属性类型、规范风格和编码策略对验证性能的影响。
本文针对B5G/6G工业服务中多阶段工作流与网络QoS协调缺失的问题,提出了一种能力代理方法以管理QoS承诺并保证一致性。
This study addresses the unreliability of long-horizon planning in large language models and the absence of automated verification benchmarks for PDDL generation. We propose an iterative agent-based PDDL generation framework and introduce NL-PDDLGym, a novel benchmark supporting executable environment validation. This approach leverages agent feedback mechanisms to ensure reliable natural-language-to-symbolic planning translation with automated verification. Experimental results demonstrate that our method achieves an 89.6% valid plan generation rate on the test set, significantly outperforming existing PDDL generation approaches and direct LLM planning baselines. These findings confirm that integrating iterative feedback with executable verification effectively enhances both the accuracy and robustness of symbolic planning derived from natural language specifications.
Although block encoding theoretically underpins numerous advanced quantum algorithms—such as Quantum Signal Processing (QSP) and Quantum Singular Value Transformation (QSVT)—its intricate implementation has hindered practical adoption. This work introduces, for the first time, a generalized programming interface that abstracts block encoding and integrates it into the Eclipse Qrisp framework. The interface encapsulates key techniques including qubitization, the Childs–Kothari–Somma construction, and arithmetic composition, enabling high-level expression and automated resource estimation for algorithms like matrix inversion, polynomial filtering, and Hamiltonian simulation. By providing clear mechanisms for constructing and composing block encodings, this interface substantially lowers the barrier to using state-of-the-art quantum algorithms, enhances developer productivity, and improves accessibility, as demonstrated through illustrative code examples.
Interactive AI agents introduce emergent systemic risks at the system level—particularly unpredictable failures arising from multi-agent coordination. Method: We propose a scenario-driven risk identification paradigm, construct representative risk-evolution case studies spanning smart grids and social welfare domains, and introduce Agentology—a novel graphical modeling language for capturing complex agent interactions. We systematically identify and categorize emergent behaviors that trigger systemic risks, establishing the first hierarchical risk taxonomy specifically designed for interactive AI systems. Contribution/Results: Our work yields a cross-domain, transferable risk analysis framework and pioneers the research direction of AI system-level safety. It provides both theoretical foundations and methodological tools for designing and governing high-reliability multi-agent systems, advancing rigorous, scalable approaches to AI safety beyond component-level assurance.
研究通过SMT求解方法验证高风险ML系统的数据集质量,评估了数据属性类型、规范风格和编码策略对验证性能的影响。
本文针对B5G/6G工业服务中多阶段工作流与网络QoS协调缺失的问题,提出了一种能力代理方法以管理QoS承诺并保证一致性。
This study addresses the unreliability of long-horizon planning in large language models and the absence of automated verification benchmarks for PDDL generation. We propose an iterative agent-based PDDL generation framework and introduce NL-PDDLGym, a novel benchmark supporting executable environment validation. This approach leverages agent feedback mechanisms to ensure reliable natural-language-to-symbolic planning translation with automated verification. Experimental results demonstrate that our method achieves an 89.6% valid plan generation rate on the test set, significantly outperforming existing PDDL generation approaches and direct LLM planning baselines. These findings confirm that integrating iterative feedback with executable verification effectively enhances both the accuracy and robustness of symbolic planning derived from natural language specifications.
Although block encoding theoretically underpins numerous advanced quantum algorithms—such as Quantum Signal Processing (QSP) and Quantum Singular Value Transformation (QSVT)—its intricate implementation has hindered practical adoption. This work introduces, for the first time, a generalized programming interface that abstracts block encoding and integrates it into the Eclipse Qrisp framework. The interface encapsulates key techniques including qubitization, the Childs–Kothari–Somma construction, and arithmetic composition, enabling high-level expression and automated resource estimation for algorithms like matrix inversion, polynomial filtering, and Hamiltonian simulation. By providing clear mechanisms for constructing and composing block encodings, this interface substantially lowers the barrier to using state-of-the-art quantum algorithms, enhances developer productivity, and improves accessibility, as demonstrated through illustrative code examples.
Interactive AI agents introduce emergent systemic risks at the system level—particularly unpredictable failures arising from multi-agent coordination. Method: We propose a scenario-driven risk identification paradigm, construct representative risk-evolution case studies spanning smart grids and social welfare domains, and introduce Agentology—a novel graphical modeling language for capturing complex agent interactions. We systematically identify and categorize emergent behaviors that trigger systemic risks, establishing the first hierarchical risk taxonomy specifically designed for interactive AI systems. Contribution/Results: Our work yields a cross-domain, transferable risk analysis framework and pioneers the research direction of AI system-level safety. It provides both theoretical foundations and methodological tools for designing and governing high-reliability multi-agent systems, advancing rigorous, scalable approaches to AI safety beyond component-level assurance.