Reason in Style: Discovering and Controlling Style in Language Models

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the deep entanglement of content and style in language models, which hinders the identification and control of reasoning styles. We propose an unsupervised content-style disentanglement algorithm that automatically discovers and balances six mathematical reasoning styles. Combined with importance-weighted fine-tuning of a student model, this approach achieves explicit control over reasoning styles for the first time. Evaluated across six mathematical benchmarks, our method significantly improves Pass@k performance while ensuring strong alignment between requested and generated styles. Furthermore, it reveals how stylistic choices influence accuracy and elucidates problem-style adaptation mechanisms. This work establishes a novel paradigm for enhancing the reasoning capabilities of language models through principled style control.
📝 Abstract
Language models learn content and style jointly, making stylistic variation in their outputs difficult to identify and control. We study whether recurring styles in model responses can be discovered without supervision and explicitly controlled. We design an algorithm that learns to separate representations of content and style from language models' outputs and validate its effectiveness on math questions in a controlled setting. By applying this method to over 100K verified traces from nine distinct teacher models, we discover six recurring yet imbalanced styles. We then fine-tune smaller student models to follow these styles when explicitly conditioned on them, using importance weighting to balance the contribution of the styles represented in the corpus. This approach improves Pass@$k$ over standard fine-tuning on the same data across six math reasoning benchmarks, demonstrating that we can diversify the style of answers effectively. We confirm that this also results in strong correspondence between requested and realized styles. We find that style affects correctness: the probability of solving a problem depends on the style we condition on, and different problems benefit from different styles. In summary, our results show that stylistic variation in model-generated data can be discovered in an unsupervised way, and made explicit, providing a source of both control and improved reasoning performance.
Problem

Research questions and friction points this paper is trying to address.

language models
style control
unsupervised discovery
content-style separation
math reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Style Disentanglement
Unsupervised Style Discovery
Importance Weighting
Conditional Fine-tuning
Math Reasoning
🔎 Similar Papers
No similar papers found.