Minimal Witness Reinforcement Learning

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of standard reinforcement learning in minimal sufficient witness set discovery, which typically yields only single or redundant solutions. To overcome this bottleneck, we propose the MWRL framework, which achieves credit assignment by evaluating the marginal contribution of proposals to covered groups. Derived directly from the problem definition, this credit mechanism unifies minimality constraints with alternative solution recovery. Algorithmically, MWRL integrates a value iteration planner with policy gradient methods scalable to large language models. Experimental results demonstrate that MWRL efficiently recovers the vast majority of minimal witness sets, significantly outperforming conventional approaches that produce redundant or singular outputs.
📝 Abstract
``What are the irreducible conditions that are sufficient to produce an outcome?'' is one of the most common questions that recur across computation and science. Its answers, the minimal sufficient witnesses, are what we mean by explanations, mechanisms and reasons. These problems usually ask for multiple minimal witnesses, yet standard RL methods may reveal only one solution or redundant ones. We formalize this problem as minimal-witness identification and introduce Minimal-Witness Reinforcement Learning (MWRL). MWRL takes the union of the sets certified by successful proposals sampled from the policy and credits each proposal for the coverage the group union would lose without that proposal. This credit assignment, derived directly from the problem definition, unifies the demands for minimality and recovery of alternatives from a single black-box verifier bit. Under this principle, we derive a value iteration planner that recovers the entire family of witnesses and a policy gradient method that can scale to large language models. Across different experimental settings, MWRL recovers most minimal witnesses, while other methods return redundant supersets or a single witness. By making witness families learnable from verifier feedback, MWRL expands the scope of reinforcement learning beyond single-solution optimization. Our code is available at https://github.com/TSUITUENYUE/MWRL.
Problem

Research questions and friction points this paper is trying to address.

Minimal Witness
Reinforcement Learning
Credit Assignment
Sufficient Conditions
Black-box Verifier
Innovation

Methods, ideas, or system contributions that make the work stand out.

Minimal Witness Reinforcement Learning
Credit Assignment
Policy Gradient
Black-box Verifier
Large Language Models
🔎 Similar Papers
No similar papers found.
T
T. Y. Tsui
University of Pennsylvania
Zihao Ye
Zihao Ye
NVIDIA, University of Washington
CompilersMachine Learning Systems
P
Pengxiang Cai
Shanghai AI Laboratory
Y
Yanchao Li
Shanghai AI Laboratory
Yuqiang Li
Yuqiang Li
Central South University
Internal Combustion EngineCombustionEmissionsMechansim
Z
Zhehong Ai
Shanghai AI Laboratory