Training Witnesses: Trusting the Training without Trusting the Trainer

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the prohibitive costs of machine learning verification, evaluation disparities arising from retraining, and the resulting trust deficit by proposing Witnesses, a method that shifts verification responsibility to trainers. Witnesses authenticates training processes and data usage through rapid behavioral fingerprinting and probabilistic replay challenges, while maintaining compatibility with DDP/FSDP distributed parallel strategies for large-scale model auditing. Experiments demonstrate that this framework achieves minimal verification overhead on language models ranging from 100 million to 2 billion parameters and exponentially amplifies the probability of rejecting substandard training. Furthermore, this project establishes an automated certification leaderboard to provide shared baselines for the community and promote research reproducibility.
📝 Abstract
Progress in machine learning cannot outpace our ability to verify it. With an explosion in papers today, every scientific claim rests initially on trust in the trainer, leading to uneven evaluation, baselines, and forestalling of reliable progress. Traditionally, the burden of verification falls on the reader, who must reproduce expensive training runs. This strategy is impractical due to an explosion in slop contributions, diversity of methods, and the sheer compute required. We put the burden of proof where it belongs, on the trainer, and in the process also cut the overall cost of verification significantly. We introduce Witnesses, a method for certifying training, data usage and evaluation in a neural network training run. Our key insight is that fast behavioral fingerprints with occasional replay challenges are sufficient for auditing neural network training. Our method is applicable at scale with minimal overhead to the trainer, is cheap for the verifier, rejects bad training runs with amplifiable probability, and allows for exact queries of both data inclusion and exclusion. We test our method on language model training runs from 100M to 2B scales, across DDP and FSDP, and demonstrate this minimal overhead. We also introduce a self-regulating leaderboard of"auto-certified"training runs that enables shared baselines and progress. We invite the community to participate in the leaderboard to improve reproducibility in machine learning.
Problem

Research questions and friction points this paper is trying to address.

verification
reproducibility
machine learning
training audit
trust
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training Verification
Behavioral Fingerprints
Replay Challenges
Reproducibility
Auto-certified Leaderboard