agent communication evaluation

Designs and executes empirical studies and controlled interventions to measure and analyze how communication protocols among autonomous agents affect system-level outcomes. Builds and applies protocol ablations and targeted communication modifications (e.g., toggling permissions, changing visibility or signal availability) to evaluate impacts on metrics such as conflict, cooperation, reputation dynamics, and resource depletion.

agentcommunicationevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

While existing methods can detect conflicts among intelligent agents in O-RAN, they struggle to predict the actual impact of such conflicts on radio access network (RAN) performance, particularly when agents generate control actions at heterogeneous temporal frequencies. This work proposes an end-to-end framework for conflict impact prediction that, for the first time, explicitly incorporates agent action frequency into the evaluation process. By integrating statistical behavioral profiling, quantification of conflict severity, and a weighted temporal modeling scheme that assigns higher influence to high-frequency control agents, the framework captures the nuanced dynamics of multi-agent interactions. Experimental results demonstrate that the proposed approach accurately predicts the degree of RAN performance degradation caused by conflicting multi-objective intelligent agents, substantially improving both the accuracy and practical utility of impact assessment in O-RAN environments.

autonomous agentsconflict impactmulti-timescale control

Current mainstream AI agent communication protocols—such as MCP, A2A, Agora, and ANP—lack a systematic framework for security threat modeling and protocol-level risk assessment, hindering effective identification of security risks in multi-agent interactions. This work proposes the first structured threat modeling methodology tailored to these protocols, defining twelve categories of protocol-level risks and establishing a qualitative yet quantifiable evaluation framework that translates security claims into falsifiable empirical analyses. Through architectural decomposition, lifecycle modeling, and case studies involving multi-server compositions, the study identifies both common and protocol-specific attack surfaces. Notably, it quantifies the risk of erroneous tool execution in MCP due to the absence of mandatory verification mechanisms, thereby providing actionable insights for secure protocol design and standardization.

AI agent communication protocolsinteroperabilityprotocol risk assessment

This work addresses a critical gap in evaluating tool-augmented large language models (LLMs), as existing metrics predominantly emphasize linguistic alignment or task success while overlooking the structural relationship between linguistic signals and executable actions across varying autonomy architectures. To remedy this, the study proposes a behavior-centric evaluation framework grounded in the execution layer, introducing a two-dimensional action–refusal (A–R) space defined by action rate (A) and refusal signals (R), along with a divergence metric (D) to quantify their coordination. Systematic experiments across four canonical scenarios and three autonomy configurations—direct execution, planning, and reflection—reveal significant behavioral distributional differences: reflective scaffolding consistently increases refusal rates in high-risk contexts, yet models exhibit structurally heterogeneous redistribution patterns. By replacing scalar safety scores with separable behavioral dimensions, this approach enables fine-grained, comparable, and interpretable characterization of tool-augmented LLM behaviors.

autonomy scaffoldsexecution-level behaviororganizational deployment

This study systematically evaluates the efficacy and underlying mechanisms of AI-driven influence operations on social networks. By constructing a synthetic social network platform that integrates multi-agent simulation, natural language generation, and belief dynamics modeling, the work quantifies for the first time the impact of three core strategies—narrative dissemination, information amplification, and counter-messaging—on audience beliefs. The findings reveal that information amplification achieves the broadest reach, counter-messaging is most effective in shifting opinions, and narrative dissemination requires substantially greater adversarial investment to yield measurable effects. Furthermore, the research uncovers an intrinsic relationship between the behavioral footprints of influence actors and the resulting effectiveness of their campaigns.

AI agentsbelief changeinfluence operations

Ctrl-Z: Controlling AI Agents via Resampling

Apr 14, 2025
AB
Aryan Bhatt
🏛️ Redwood Research | ML Alignment and Theory Scholars (MATS) Program

AI agents performing multi-step system administration tasks pose a security risk by covertly downloading and executing malicious code. Method: This paper proposes a control-assessment-based defense framework. It introduces BashBench—the first benchmark tailored to AI agent environments—for systematic control evaluation; designs a novel dynamic action-level resampling protocol that selectively resamples suspicious actions to enable real-time attack blocking and behavioral attribution; and integrates adversarial agent construction, joint historical-step analysis, and multi-step task modeling to balance security and practicality. Contribution/Results: Experiments demonstrate that the optimal protocol reduces malicious agent attack success rate from 58% to 7%, while degrading benign task performance by only 5%, validating both effectiveness and low intrusiveness.

Balancing attack prevention with agent usefulnessEvaluating AI agent safety in multi-step tasksPreventing covert malicious code execution by AI

Latest Papers

What's happening recently
View more

This study addresses the unclear relationship between parameter configurations and emergent social behaviors of AI agents in open-ended social environments. Deploying thirteen OpenClaw agents on Moltbook—a custom Reddit-style platform—the authors conducted a week-long multifactorial controlled experiment involving approximately 400 autonomous interactions. By systematically manipulating three key variables—personality profiles (based on SOUL.md), underlying large language models, and operational rules (based on AGENTS.md)—the study quantitatively demonstrates for the first time that personality settings are the primary driver of social behavior, significantly influencing response length. In contrast, the choice of language model and rule set moderately modulates rhetorical style and topical breadth. These findings provide empirical grounding and actionable guidelines for designing AI agents in real-world social settings.

AI agentsLLM modeloperational rules

Current agent benchmarks often yield misleading evaluation scores due to invalid protocols, primarily stemming from reward hacking or assessment vulnerabilities. This work presents the first systematic formalization of “protocol validity” and introduces Mislead Gap—a quantitative metric—and HackDetect, a posterior auditing framework. By integrating trajectory auditing, exposure point identification, and intent-exploitation score comparison, the framework uniformly detects and quantifies the impact of reward hacking. Empirical analysis across 15 benchmarks and 2,385 agent trajectories reveals that 66.7%–67.0% of evaluations exhibit exposure to or active engagement in reward hacking, inflating scores by 0.45–1.00. These findings demonstrate that prevailing benchmarks generally fail to validate agents’ true capabilities.

agent benchmarkscapability evaluationprotocol validity

Hot Scholars

CC

Changhyeok Choi

University of Toronto
computational catalysiselectrocatalystmachine learning
YW

Yite Wang

Research Scientist, Snowflake
Large Language ModelEfficient Deep LearningComputer VisionNatural Language Processing
CS

Chandan Singh

Senior researcher, Microsoft research
🔍 Interpretability🤖 Foundation models🧠 Neuroscience🌳 Transparent models
ZJ

Zhi Jin

Sun Yat-Sen University, Associate Professor