Cross-Stack Validation of Language-Model Training: A Clinical Fine-Tuning Case Study

📅 2026-08-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过使用独立实现的训练栈作为差异预言机来验证语言模型训练过程,以解决神经网络训练中的oracle问题。
📝 Abstract
Neural network training has an oracle problem: a run can converge normally and yield a usable model while the software beneath it computes something other than specified. Almost all such work runs on one stack, so there is rarely anything independent to check against. We study whether independently implemented training stacks can serve as differential oracles for a whole fine-tuning pipeline, rather than the operators and inference paths that prior differential testing targets. We define a trajectory-level protocol -- a shared specification, cross-check points spanning arithmetic, model loading, data rendering and the learning trajectory, and a separation of independence of the stack, the orchestration and the language runtime -- and apply it to a LoRA adaptation of Qwen3-0.6B over 168,574 clinical question-answer pairs under PyTorch and under numbat, an independent framework written in Zig, driven natively and through its C interface from six languages. Across 42 paired evaluations spanning a full epoch the two stacks' held-out cross-entropy differs by 0.134% on average, and four implementations end the epoch within 0.15% of one another. The comparison exposed 17 faults that single-implementation development had missed, two of them notable for software engineering. The fault with the largest effect on the trained model lay outside the numerical kernels: a mismatch in how clinical text was rendered moved held-out loss 0.15, some 500 times more than the arithmetic faults found beside it. And four faults were reachable only from a language whose memory model differs from the first two implementations: a scheduler migrating work across threads, a collector blind to device memory, an ownership discipline needing a primitive the interface lacked. Implementation diversity has several axes, and the runtime is one.
Problem

Research questions and friction points this paper is trying to address.

neural network training
oracle problem
independent stacks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Stack Validation
Differential Oracle
Trajectory-Level Protocol
Implementation Diversity
🔎 Similar Papers
No similar papers found.
T
Thang Tran
CloudKites AI Lab, New South Wales, Australia
L
Lan Dang
Monash Business School, Monash University, Victoria, Australia