Training of Scaffolded Language Models with Language Supervision: A Survey

📅 2024-10-21
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the challenges of structural design and continual optimization for scaffolded language models (LMs) in multi-step tasks. We propose a novel paradigm—*language-supervised training*—that enables non-parametric optimization via natural-language instructions, tool-call trajectories, and human-readable/editable linguistic feedback. This framework dynamically adapts external variables—including prompts, toolchains, and scaffolding code—without modifying model parameters. It is compatible with closed-source LMs, mitigates catastrophic forgetting, and supports human-in-the-loop streaming learning. We introduce the first taxonomy of non-parametric variables tailored to language supervision, unifying prompt engineering, multi-step reasoning orchestration, and feedback modeling. Our approach provides both theoretical foundations and a systematic implementation pathway for deploying hybrid autonomous agents in real-world settings.

Technology Category

Natural Language Processing: (Large) Language ModelsMachine Learning: Large Multimodal Models (LMMs)Search and Optimization: Learning to Search

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
📝 Abstract
This survey organizes the intricate literature on the design and optimization of emerging structures around post-trained LMs. We refer to this overarching structure as scaffolded LMs and focus on LMs that are integrated into multi-step processes with tools. We view scaffolded LMs as semi-parametric models wherein we train non-parametric variables, including the prompt, tools, and scaffold's code. In particular, they interpret instructions, use tools, and receive feedback all in language. Recent works use an LM as an optimizer to interpret language supervision and update non-parametric variables according to intricate objectives. In this survey, we refer to this paradigm as training of scaffolded LMs with language supervision. A key feature of non-parametric training is the ability to learn from language. Parametric training excels in learning from demonstration (supervised learning), exploration (reinforcement learning), or observations (unsupervised learning), using well-defined loss functions. Language-based optimization enables rich, interpretable, and expressive objectives, while mitigating issues like catastrophic forgetting and supporting compatibility with closed-source models. Furthermore, agents are increasingly deployed as co-workers in real-world applications such as Copilot in Office tools or software development. In these mixed-autonomy settings, where control and decision-making are shared between human and AI, users point out errors or suggest corrections. Accordingly, we discuss agents that continuously improve by learning from this real-time, language-based feedback and refer to this setting as streaming learning from language supervision.
Problem

Research questions and friction points this paper is trying to address.

Survey on scaffolded LMs design and optimization techniques
Training LMs with language supervision for tool integration
Enhancing LMs via real-time language feedback in mixed-autonomy settings
Innovation

Methods, ideas, or system contributions that make the work stand out.

Scaffolded LMs integrate tools via semi-parametric training
Language supervision optimizes non-parametric variables adaptively
Streaming learning enables real-time feedback from users
🔎 Similar Papers
No similar papers found.
Tsinghua University | Pennsylvania State University | Xi’an Jiaotong University
M
M. Lin
Tsinghua University
J
Jenny Sheng
Tsinghua University
Andrew Zhao
Andrew Zhao
Tsinghua University
Reinforcement LearningLanguage AgentReasoning
S
Shenzhi Wang
Tsinghua University
Y
Yang Yue
Tsinghua University
Y
Yiran Wu
Pennsylvania State University
H
Huan Liu
Xi’an Jiaotong University
J
Jun Liu
Xi’an Jiaotong University
G
Gao Huang
Tsinghua University
Yong-Jin Liu
Yong-Jin Liu
Professor of College of Mathematics and Computer Science at Fuzhou University
Mathematical ProgrammingStatistical OptimizationNumerical Computation