Institution profile

Workday Inc.

Industry researchnorthamerica · us
Official website
Research library7linked papers
Opportunities0open roles
Selected work

Representative Papers

MemLeak: Cross-User Semantic Leakage in Multi-Tenant AI Agent Memory

Oct 03, 2026

This study addresses the privacy risks arising from cross-user semantic memory leakage when multi-tenant AI agents share vector stores. We formally define cross-user acceptability failure and quantify privacy vulnerabilities under both non-adversarial and adversarial retrieval settings. Through systematic evaluation using MiniLM dense retrieval, TF-IDF sparse retrieval, and cosine similarity, we propose a low-latency hard-ownership gating mechanism. Experimental results demonstrate that unprotected systems exhibit leakage rates of 70%–100% with response contamination scores reaching 5/5. The proposed hard-gating approach emerges as the sole effective mitigation strategy, restoring contamination scores to baseline levels (1.00/5) while introducing only 1.4 ms of additional latency, thereby achieving an optimal balance between security guarantees and real-time performance requirements.

0 citationsRead paper

Does the Readout Bypass Leak the Input? A Feature-Visibility Audit of Hybrid Quantum-Classical Models

Sep 29, 2026

This study addresses the security vulnerability of raw input leakage through readout-side residual shortcuts in hybrid quantum-classical models. Challenging the prevailing misconception that quantum processing inherently guarantees privacy protection, this work proposes a feature visibility-based privacy auditing framework. Through iterative gradient matching, PSNR metric analysis, membership inference attacks, and multi-architecture comparative experiments, it systematically evaluates the exposure risk of original data coordinates via shortcut connections. The results demonstrate that residual shortcuts enable near-perfect reconstruction of raw inputs, whereas purely quantum heads recover only partially encoded coordinates. To our knowledge, this is the first work to quantitatively reveal the privacy bottleneck inherent in hybrid architectures, providing critical insights for their secure design.

0 citationsRead paper

Report: Progressive Disclosure of Agent Skills

Sep 28, 2026

As the skill libraries of large language model (LLM) agents expand, operational costs escalate sharply, yet the effects of progressive disclosure strategies on retrieval quality and latency remain unclear. This study investigates enterprise-grade Workday agents to empirically quantify, for the first time, the performance trade-offs of lazy-loading-based progressive disclosure mechanisms in real-world production environments. Through systematic experimental evaluation, we demonstrate that this strategy significantly enhances skill retrieval quality while introducing only marginal response latency and effectively reducing operational costs. These findings provide critical empirical evidence for optimizing the architectural design of enterprise LLM agents, offering a practical pathway to balance scalability, efficiency, and system responsiveness in large-scale deployments.

0 citationsRead paper

Robust Explanations for User Trust in Enterprise NLP Systems

Apr 13, 2026

This study addresses the lack of effective evaluation for the robustness of token-level explanations in real-world enterprise NLP systems deployed as black boxes under authentic user noise, which undermines user trust. The work proposes the first unified black-box robustness evaluation framework, integrating leave-one-out masking with multiple realistic perturbations—substitution, deletion, shuffling, and back-translation—and introduces the top-token flip rate as a key metric. Large-scale experiments across six models, including BERT, RoBERTa, Qwen, and Llama, enable the first systematic cross-architecture comparison of explanation robustness between encoders and decoder-based large language models (LLMs). Results reveal that decoder LLMs significantly outperform encoder models, exhibiting 73% lower average flip rates, and demonstrate that scaling model size from 7B to 70B parameters improves stability by 44%, alongside a derived cost–robustness trade-off curve.

0 citationsRead paper

LLMs as Orchestrators: Constraint-Compliant Multi-Agent Optimization for Recommendation Systems

Jan 27, 2026

This work addresses the challenge of enforcing hard business constraints—such as fairness and coverage—in recommender systems, which are often softened into penalty terms in existing approaches, leading to frequent violations in deployment. To overcome this limitation, the authors propose DualAgent-Rec, a novel framework that leverages a large language model (LLM) as a coordinator to orchestrate two collaborative agents: one optimizes recommendation accuracy under strict adherence to hard constraints, while the other enhances diversity through unconstrained Pareto search. An adaptive epsilon-relaxation mechanism is integrated to ensure solution feasibility and computational efficiency. Evaluated on the Amazon Reviews 2023 dataset, the method achieves 100% constraint satisfaction, improves Pareto hypervolume by 4–6%, and maintains an excellent trade-off between accuracy and diversity.

0 citationsRead paper
Recent publications

Latest Papers

MemLeak: Cross-User Semantic Leakage in Multi-Tenant AI Agent Memory

Oct 03, 2026

This study addresses the privacy risks arising from cross-user semantic memory leakage when multi-tenant AI agents share vector stores. We formally define cross-user acceptability failure and quantify privacy vulnerabilities under both non-adversarial and adversarial retrieval settings. Through systematic evaluation using MiniLM dense retrieval, TF-IDF sparse retrieval, and cosine similarity, we propose a low-latency hard-ownership gating mechanism. Experimental results demonstrate that unprotected systems exhibit leakage rates of 70%–100% with response contamination scores reaching 5/5. The proposed hard-gating approach emerges as the sole effective mitigation strategy, restoring contamination scores to baseline levels (1.00/5) while introducing only 1.4 ms of additional latency, thereby achieving an optimal balance between security guarantees and real-time performance requirements.

0 citationsRead paper

Does the Readout Bypass Leak the Input? A Feature-Visibility Audit of Hybrid Quantum-Classical Models

Sep 29, 2026

This study addresses the security vulnerability of raw input leakage through readout-side residual shortcuts in hybrid quantum-classical models. Challenging the prevailing misconception that quantum processing inherently guarantees privacy protection, this work proposes a feature visibility-based privacy auditing framework. Through iterative gradient matching, PSNR metric analysis, membership inference attacks, and multi-architecture comparative experiments, it systematically evaluates the exposure risk of original data coordinates via shortcut connections. The results demonstrate that residual shortcuts enable near-perfect reconstruction of raw inputs, whereas purely quantum heads recover only partially encoded coordinates. To our knowledge, this is the first work to quantitatively reveal the privacy bottleneck inherent in hybrid architectures, providing critical insights for their secure design.

0 citationsRead paper

Report: Progressive Disclosure of Agent Skills

Sep 28, 2026

As the skill libraries of large language model (LLM) agents expand, operational costs escalate sharply, yet the effects of progressive disclosure strategies on retrieval quality and latency remain unclear. This study investigates enterprise-grade Workday agents to empirically quantify, for the first time, the performance trade-offs of lazy-loading-based progressive disclosure mechanisms in real-world production environments. Through systematic experimental evaluation, we demonstrate that this strategy significantly enhances skill retrieval quality while introducing only marginal response latency and effectively reducing operational costs. These findings provide critical empirical evidence for optimizing the architectural design of enterprise LLM agents, offering a practical pathway to balance scalability, efficiency, and system responsiveness in large-scale deployments.

0 citationsRead paper

Robust Explanations for User Trust in Enterprise NLP Systems

Apr 13, 2026

This study addresses the lack of effective evaluation for the robustness of token-level explanations in real-world enterprise NLP systems deployed as black boxes under authentic user noise, which undermines user trust. The work proposes the first unified black-box robustness evaluation framework, integrating leave-one-out masking with multiple realistic perturbations—substitution, deletion, shuffling, and back-translation—and introduces the top-token flip rate as a key metric. Large-scale experiments across six models, including BERT, RoBERTa, Qwen, and Llama, enable the first systematic cross-architecture comparison of explanation robustness between encoders and decoder-based large language models (LLMs). Results reveal that decoder LLMs significantly outperform encoder models, exhibiting 73% lower average flip rates, and demonstrate that scaling model size from 7B to 70B parameters improves stability by 44%, alongside a derived cost–robustness trade-off curve.

0 citationsRead paper

LLMs as Orchestrators: Constraint-Compliant Multi-Agent Optimization for Recommendation Systems

Jan 27, 2026

This work addresses the challenge of enforcing hard business constraints—such as fairness and coverage—in recommender systems, which are often softened into penalty terms in existing approaches, leading to frequent violations in deployment. To overcome this limitation, the authors propose DualAgent-Rec, a novel framework that leverages a large language model (LLM) as a coordinator to orchestrate two collaborative agents: one optimizes recommendation accuracy under strict adherence to hard constraints, while the other enhances diversity through unconstrained Pareto search. An adaptive epsilon-relaxation mechanism is integrated to ensure solution feasibility and computational efficiency. Evaluated on the Amazon Reviews 2023 dataset, the method achieves 100% constraint satisfaction, improves Pareto hypervolume by 4–6%, and maintains an excellent trade-off between accuracy and diversity.

0 citationsRead paper