SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文通过SkillGym框架将人类技能转化为可执行的任务环境,用于训练大型语言模型,提高其解决实际问题的能力。
📝 Abstract
Human-written agent skills encode rich workflows for real-world problem solving, but are typically used as external inference-time instructions rather than internalized as reusable model capabilities. We introduce \texttt{SkillGym}, a framework that transforms these skills into executable, verifiable training environments for large language model agents. Its skill-to-task pipeline instantiates concrete tasks, verifies outcomes with code-based checkers, and assesses empirical skill dependence through contrastive executions. We construct and release 2,756 environments across 12 categories and collect 8,364 successful trajectories from multiple models and harnesses, averaging 49 tool calls and over 60k logged text tokens. These resources support supervised fine-tuning on verified workflows and reinforcement learning with outcome-based rewards. Under Claude Code, supervised fine-tuning improves Qwen3.5-35B-A3B by 199 Elo on GDPval-AA v2, 19.10 percentage points on Terminal-Bench 2.1, and 28.13 and 12.38 points on SkillsBench v1.1 with and without skills, respectively. Our 35B \texttt{SkillGym-Agent} reaches 51.47\% on skill-assisted SkillsBench, exceeding reported scores for Claude Sonnet 4.6, GPT-5.4 Mini, and DeepSeek V4 Pro. Without skills, it also surpasses skill-assisted bases under Codex and Claude Code, suggesting reusable procedural competence.
Problem

Research questions and friction points this paper is trying to address.

human skills
large language models
real-world problem solving
internalized capabilities
Innovation

Methods, ideas, or system contributions that make the work stand out.

skill-to-task pipeline
executable training environments
supervised fine-tuning
reinforcement learning
empirical skill dependence
🔎 Similar Papers
No similar papers found.
Z
Zhilong Ge
School of Computer Science and Technology, East China Normal University, Shanghai
Y
Yuting Shao
School of Computer Science and Technology, East China Normal University, Shanghai
Y
Yutao Yang
School of Computer Science and Technology, East China Normal University, Shanghai
Y
Yuxuan Cai
School of Computer Science and Technology, East China Normal University, Shanghai
Jie Zhou
Jie Zhou
East China Normal University
NLPContinuous LearningSentiment AnalysisLLMsInformation Extraction
Kai Chen
Kai Chen
Shanghai AI Laboratory
LLMVLMComputer Vision
B
Bo Zhang
Shanghai AI Laboratory
Qin Chen
Qin Chen
East China Normal University
Natural Language ProcessingQuestion AnsweringLarge Language Model
Liang He
Liang He
East China Normal University
Artificial IntelligenceNatural Language ProcessingHuman-in-the-Loop