Institution profile

Cisco Systems

Industry researchnorthamerica · us
Official website
Research library329linked papers
Opportunities0open roles
Selected work

Representative Papers

A Survey on Large Language Model-Based Game Agents

Apr 02, 2024arXiv.org

This paper addresses the challenge of enabling human-like decision-making in game agents operating within complex environments. Methodologically, it proposes the first three-dimensional functional architecture—Memory–Reasoning–I/O—for LLM-driven game agents, systematically reviewing over 50 representative works across six game genres, including adventure, communication, and competitive games. The approach integrates multimodal perception, long-term memory, chain-of-thought reasoning, and game API integration to establish cross-genre unified evaluation dimensions. Key contributions include: (1) the first formalization of a functional architecture for LLM-based game agents; (2) the creation of an open-source, structured, and authoritative repository of relevant literature; and (3) an empirical analysis revealing critical performance bottlenecks and generalization limitations, thereby providing both a theoretical framework and a practical roadmap for AGI-oriented game agent research.

45 citations2 influentialRead paper

AssistedDS: Benchmarking How External Domain Knowledge Assists LLMs in Automated Data Science

May 25, 2025arXiv.org

Large language models (LLMs) struggle to critically leverage external domain knowledge in automated data science. Method: We introduce AssistedDS—the first benchmark for domain-knowledge-assisted evaluation—comprising synthetic and real Kaggle datasets paired with beneficial or adversarial domain documents. Our “interpretable synthesis + real-world scenarios” dual-track framework employs multi-stage prompting to assess end-to-end capabilities: document retrieval, knowledge filtering, code generation, and execution validation. Contribution/Results: We uncover a critical “blind adoption” flaw in LLMs: they fail significantly in time-series modeling, cross-fold consistency, and categorical variable handling. Experiments show state-of-the-art models suffer sharp performance degradation under adversarial documents; beneficial knowledge fails to mitigate harmful information, exposing severe deficiencies in domain knowledge discrimination and robust application.

3 citations1 influentialRead paper

Adaptive Memory Crystallization for Autonomous AI Agent Learning in Dynamic Environments

Apr 02, 2026

This work addresses the challenge of catastrophic forgetting and the difficulty in balancing old and new knowledge in autonomous AI agents during continual learning. The authors propose the Adaptive Memory Crystallization (AMC) architecture, which uniquely models memory stability as a continuous liquid–glass–crystal phase transition governed by an Itô stochastic differential equation, coupled with multi-objective utility signals to regulate experience transfer. Theoretical analysis yields a closed-form Beta stationary distribution, rigorously establishing convergence guarantees, error bounds, and memory capacity limits. Empirical evaluations demonstrate that AMC achieves 34–43% improvement in forward transfer, reduces forgetting by 67–80%, and decreases memory footprint by 62% across benchmarks including Meta-World MT50, a 20-task Atari sequence, and MuJoCo environments.

1 citationsRead paper

Toward Quantitative Modeling of Cybersecurity Risks Due to AI Misuse

Dec 09, 2025

This study addresses the challenge of quantifying how AI misuse exacerbates cybersecurity risks. We develop nine quantitative risk models grounded in the MITRE ATT&CK framework to systematically assess AI’s impact on attack scale, frequency, success rate, and impact severity. Methodologically, we propose a novel dual-source uncertainty estimation framework integrating Delphi expert elicitation with LLM-simulated expert reasoning, and map AI benchmark scores (e.g., Cybench, BountyBench) to interpretable risk factors to decompose and attribute cross-dimensional attack efficacy uplift. Using Monte Carlo aggregation, we generate auditable, confidence-interval–bounded results. Our approach enables dynamic defense prioritization and evidence-based AI governance decisions—marking the first step toward verifiable, debatable, and iterative quantitative AI security risk assessment, moving beyond qualitative descriptions.

1 citationsRead paper

A Generic Framework for Conformal Fairness

May 22, 2025International Conference on Learning Representations

This work addresses the unfair coverage problem of conformal prediction (CP) on data containing sensitive attributes. We formally define “conformal fairness” as a constraint on the disparity in marginal coverage across sensitive groups. We propose the first theoretically grounded conformal fairness framework, which relaxes the standard i.i.d. assumption to accommodate non-i.i.d. structured data—such as graphs. Our method integrates exchangeability assumptions, group-wise calibration, and adaptive confidence adjustment to jointly control both coverage validity and fairness. Experiments on graph and tabular datasets demonstrate that our approach strictly satisfies the theoretical coverage guarantee while reducing inter-group coverage disparity to within a user-specified threshold. It consistently outperforms existing baselines in both fairness and calibration fidelity.

1 citationsRead paper
Recent publications

Latest Papers