Institution profile

Distyl AI

Industry researchnorthamerica · us
Official website
Research library7linked papers
Opportunities0open roles
Selected work

Representative Papers

AgentMemGate: Addressing Speculation Contamination in Conversational Assistant Memory

Oct 05, 2026

This study addresses the problem of “speculative contamination” in conversational assistants, where tentative user plans are erroneously stored as established facts within long-term memory. To mitigate this, we propose a write-time gating mechanism that leverages an LLM-based classifier to identify utterance intent—distinguishing speculation, completion, and correction—and employs conditional logic to govern memory promotion and deletion, thereby retaining only confirmed events while isolating speculative information. This work introduces the first speculative filtering technique targeting intermediate-state memory and constructs a multi-session speculative dataset to bridge existing evaluation gaps. Experimental results demonstrate that our approach completely eliminates memory contamination on core benchmarks, improving task accuracy from 65% to 95% and significantly outperforming baseline models such as Mem0.

0 citationsRead paper

HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases

Sep 24, 2026

This study addresses the limitation of existing language model evaluation benchmarks in capturing the complexities of enterprise deployment, which leads to distorted performance assessments. To this end, this work proposes HARDEN, a method that leverages constrained evolutionary search along domain-specific complexity axes to automatically generate answer-preserving, high-difficulty evaluation variants while strictly maintaining task semantics, authenticity, and execution validity. Experimental results demonstrate that HARDEN reduces average model accuracy by 22.7%, with maximum declines reaching 49.9%. These findings indicate that the proposed approach significantly enhances evaluation rigor, providing a more realistic and effective tool for assessing the robustness of language models in practical deployment scenarios.

0 citationsRead paper

Environment Maps: Structured Environmental Representations for Long-Horizon Agents

Mar 24, 2026

Although large language models (LLMs) have advanced rapidly, robust automation of complex software workflows remains an open problem. In long-horizon settings, agents frequently suffer from cascading errors and environmental stochasticity; a single misstep in a dynamic interface can lead to task failure, resulting in hallucinations or trial-and-error. This paper introduces $\textit{Environment Maps}$: a persistent, agent-agnostic representation that mitigates these failures by consolidating heterogeneous evidence, such as screen recordings and execution traces, into a structured graph. The representation consists of four core components: (1) Contexts (abstracted locations), (2) Actions (parameterized affordances), (3) Workflows (observed trajectories), and (4) Tacit Knowledge (domain definitions and reusable procedures). We evaluate this framework on the WebArena benchmark across five domains. Agents equipped with environment maps achieve a 28.2% success rate, nearly doubling the performance of baselines limited to session-bound context (14.2%) and outperforming agents that have access to the raw trajectory data used to generate the environment maps (23.3%). By providing a structured interface between the model and the environment, Environment Maps establish a persistent foundation for long-horizon planning that is human-interpretable, editable, and incrementally refinable.

0 citationsRead paper

PrefPO: Pairwise Preference Prompt Optimization

Mar 13, 2026

This work proposes a lightweight, unsupervised prompt optimization framework based on pairwise preferences, addressing key limitations of existing methods that rely on annotated data, produce verbose and repetitive prompts, and require extensive manual tuning. The approach leverages only an initial prompt and natural language criteria, using a large language model as a discriminator to provide preference-based feedback for iterative refinement. This method significantly enhances prompt conciseness and diversity while reducing the risk of “cheating” through overfitting to evaluation metrics. Evaluated across nine Big-Bench Hard (BBH) tasks, it achieves state-of-the-art or comparable performance on six, and matches TextGrad’s results on IFEval-Hard. Moreover, it reduces prompt length and repetition by 3–5×, with both human and model-based evaluations consistently outperforming baseline methods.

0 citationsRead paper
Recent publications

Latest Papers

AgentMemGate: Addressing Speculation Contamination in Conversational Assistant Memory

Oct 05, 2026

This study addresses the problem of “speculative contamination” in conversational assistants, where tentative user plans are erroneously stored as established facts within long-term memory. To mitigate this, we propose a write-time gating mechanism that leverages an LLM-based classifier to identify utterance intent—distinguishing speculation, completion, and correction—and employs conditional logic to govern memory promotion and deletion, thereby retaining only confirmed events while isolating speculative information. This work introduces the first speculative filtering technique targeting intermediate-state memory and constructs a multi-session speculative dataset to bridge existing evaluation gaps. Experimental results demonstrate that our approach completely eliminates memory contamination on core benchmarks, improving task accuracy from 65% to 95% and significantly outperforming baseline models such as Mem0.

0 citationsRead paper

HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases

Sep 24, 2026

This study addresses the limitation of existing language model evaluation benchmarks in capturing the complexities of enterprise deployment, which leads to distorted performance assessments. To this end, this work proposes HARDEN, a method that leverages constrained evolutionary search along domain-specific complexity axes to automatically generate answer-preserving, high-difficulty evaluation variants while strictly maintaining task semantics, authenticity, and execution validity. Experimental results demonstrate that HARDEN reduces average model accuracy by 22.7%, with maximum declines reaching 49.9%. These findings indicate that the proposed approach significantly enhances evaluation rigor, providing a more realistic and effective tool for assessing the robustness of language models in practical deployment scenarios.

0 citationsRead paper

Environment Maps: Structured Environmental Representations for Long-Horizon Agents

Mar 24, 2026

Although large language models (LLMs) have advanced rapidly, robust automation of complex software workflows remains an open problem. In long-horizon settings, agents frequently suffer from cascading errors and environmental stochasticity; a single misstep in a dynamic interface can lead to task failure, resulting in hallucinations or trial-and-error. This paper introduces $\textit{Environment Maps}$: a persistent, agent-agnostic representation that mitigates these failures by consolidating heterogeneous evidence, such as screen recordings and execution traces, into a structured graph. The representation consists of four core components: (1) Contexts (abstracted locations), (2) Actions (parameterized affordances), (3) Workflows (observed trajectories), and (4) Tacit Knowledge (domain definitions and reusable procedures). We evaluate this framework on the WebArena benchmark across five domains. Agents equipped with environment maps achieve a 28.2% success rate, nearly doubling the performance of baselines limited to session-bound context (14.2%) and outperforming agents that have access to the raw trajectory data used to generate the environment maps (23.3%). By providing a structured interface between the model and the environment, Environment Maps establish a persistent foundation for long-horizon planning that is human-interpretable, editable, and incrementally refinable.

0 citationsRead paper

PrefPO: Pairwise Preference Prompt Optimization

Mar 13, 2026

This work proposes a lightweight, unsupervised prompt optimization framework based on pairwise preferences, addressing key limitations of existing methods that rely on annotated data, produce verbose and repetitive prompts, and require extensive manual tuning. The approach leverages only an initial prompt and natural language criteria, using a large language model as a discriminator to provide preference-based feedback for iterative refinement. This method significantly enhances prompt conciseness and diversity while reducing the risk of “cheating” through overfitting to evaluation metrics. Evaluated across nine Big-Bench Hard (BBH) tasks, it achieves state-of-the-art or comparable performance on six, and matches TextGrad’s results on IFEval-Hard. Moreover, it reduces prompt length and repetition by 3–5×, with both human and model-based evaluations consistently outperforming baseline methods.

0 citationsRead paper