ContextWeave: A Real-World Workflow Benchmark

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing evaluation frameworks struggle to assess the memory capabilities of language agents in long-term, stateful real-world workflows. To address this gap, this work proposes ContextWeave—the first longitudinal benchmark constructed from de-identified office workflows of 14 participants over multiple months, comprising 1,005 executable tasks. Leveraging containerized environments, task trajectory logging, and fine-grained scoring, the framework systematically evaluates the impact of diverse memory mechanisms on agent workflow performance. Experiments demonstrate that an optimized memory configuration improves the Workspace Score from 68.08 to 78.20 and elevates the Preference Score from 41.50 to 70.60. Empirical memory outperforms summarization-based approaches but proves more susceptible to misleading recall. This study is the first to quantitatively measure memory’s contribution to agent efficacy in realistic, extended scenarios, introducing preference alignment and multidimensional diagnostic metrics.
📝 Abstract
Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We introduce ContextWeave, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams. ContextWeave reconstructs privacy-preserved, multi-month workflows of 14 participants into 1,005 executable tasks, including 568 core evaluation tasks, with instructions, containerized environments, trajectories, and task-specific rubrics. It measures workspace quality and alignment with participant-specific preferences, complemented by diagnostics of relevance, continuity, solvability, and robustness to misleading recall. Across six memory components under a fixed model, the strongest configuration raises Workspace Score from 68.08 to 78.20 and Preference Score from 41.50 to 70.60. With a fixed memory component, recall improves both outcomes for all five tested base models, although gains vary substantially. Our analysis shows that actionable, experience-rich memory supports workflow continuation and reduces redundant exploration more effectively than compact summaries, while it can also be more susceptible to misleading recall. These findings motivate memory systems that optimize not only retrieval relevance but also reliable use during execution.
Problem

Research questions and friction points this paper is trying to address.

memory
language agents
workflow benchmark
long-horizon tasks
real-world evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

ContextWeave
memory evaluation
long-horizon workflows
language agents
real-world benchmark
🔎 Similar Papers
No similar papers found.
B
Bo Wang
Fudan University
Y
Yuqian Yao
Fudan University; Shanghai Innovation Institute
E
Enxi Wang
Fudan University
L
Luozhijie Jin
Fudan University; Shanghai Innovation Institute
Yang Liu
Yang Liu
AP, Tongji University; Ph.D., Fudan University & University of Toronto; B.E., Nanjing University
Signal processingComputer visionComputing
Y
Yiran Suo
Fudan University; Shanghai Innovation Institute
Y
Yuxuan Cai
Fudan University
E
Enyu Zhou
Fudan University
Yufei Gao
Yufei Gao
Zhengzhou University
Machine learningMedical Image Analysis
Honglin Guo
Honglin Guo
Fudan University
Large Language Model
Tianyu Huai
Tianyu Huai
East China Normal University
Continual Learning
L
Li Ji
Fudan University; Shanghai Innovation Institute
Z
Zhikai Lei
Fudan University
B
Bufan Li
Fudan University
L
Lizhi Lin
Fudan University
J
Jinxiu Liu
Fudan University
J
Jie Yang
Fudan University
J
Jiazheng Zhou
Fudan University
M
Maosen Zhou
Fudan University; Shanghai Innovation Institute
P
Pengfang Qian
Fudan University; Shanghai Innovation Institute
Shichun Liu
Shichun Liu
Fudan University
NLP
G
Guanshan Liu
ByteDance
H
Hao Zheng
ByteDance
Y
Yunhao Yu
ByteDance
Hang Yan
Hang Yan
Computer Science, Fudan University
natural language processing