Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces
Chain-of-Thought (CoT) is frequently regarded as a faithful record of model reasoning, yet its natural language form resists mechanical verification. This study leverages the iGSM benchmark to decouple "correct answers" from "valid reasoning" by programmatically inspecting generated trajectories step-by-step. The findings reveal that even when final answers are correct, 31.6% of the most challenging instances are accompanied by semantically invalid CoTs. Furthermore, accuracy remains robust under non-minimal or perturbed training data. These results expose critical blind spots in CoT monitoring and challenge prevailing assumptions within AI safety that rely on chain-of-thought reasoning for model interpretability and alignment.