Institution profile

InfiX.ai

Industry research
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning

Oct 04, 2026

This study addresses the training instability arising from learner-sampler misalignment in low-precision policy optimization. We propose a direction-aware stabilization method that employs a piecewise diagnostic mechanism to reveal the interaction between such mismatches and gradient directions, alongside selective rebalancing and bounded repair algorithms to correct severe deviations. This approach enables native NVFP4 (W4A4) quantized training, achieving stable optimization while preserving forward inference efficiency. Experimental results demonstrate that the proposed method attains full-precision performance on mathematical reasoning benchmarks, yielding throughput improvements of up to 2.3× over the BF16 baseline.

0 citationsRead paper

InfiMed2: A Generalist Medical Multimodal Foundation Model from Contextual Evidence and Stability-Aware Supervision

Sep 28, 2026

This study addresses critical bottlenecks in the continual training of medical multimodal large language models, including data inefficiency, insufficient supervision signals, and poor output consistency, by proposing the InfiMed2 framework. Methodologically, we construct a 55.68-billion-token corpus and design a phase-aware data strategy to optimize pretraining. Furthermore, we introduce an answer-stability-based SFT reconstruction mechanism combined with evidence-focused mixing during learning rate decay, alongside RLVR to enhance fine-tuning quality. Ultimately, we release 4B and 27B general-purpose medical multimodal foundation models. The 4B model achieves an average score of 66.73%, surpassing Qwen3.5-9B, while the 27B model attains 73.72%, establishing new state-of-the-art performance among open-source counterparts.

0 citationsRead paper

Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs

Sep 26, 2026

This study addresses whether routing drift in Mixture-of-Experts (MoE) model merging is equivalent to routing failure and when intervention is necessary. We propose redefining routing failure through the lens of task loss recoverability. Leveraging a cross-model routing analysis toolkit, we employ counterfactual interventions and token-level attribution to demonstrate that most expert reassignments stem from input shifts rather than functional degradation. Based on these findings, we introduce a Selective Router Repair (SRR) strategy. This work challenges the prevailing misconception that drift inherently implies failure, showing that naively restoring source routing does not reliably improve performance. By releasing open-source tools and code, we provide theoretical foundations for precise post-merging diagnosis and repair in MoE architectures.

0 citationsRead paper
Recent publications

Latest Papers

TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning

Oct 04, 2026

This study addresses the training instability arising from learner-sampler misalignment in low-precision policy optimization. We propose a direction-aware stabilization method that employs a piecewise diagnostic mechanism to reveal the interaction between such mismatches and gradient directions, alongside selective rebalancing and bounded repair algorithms to correct severe deviations. This approach enables native NVFP4 (W4A4) quantized training, achieving stable optimization while preserving forward inference efficiency. Experimental results demonstrate that the proposed method attains full-precision performance on mathematical reasoning benchmarks, yielding throughput improvements of up to 2.3× over the BF16 baseline.

0 citationsRead paper

InfiMed2: A Generalist Medical Multimodal Foundation Model from Contextual Evidence and Stability-Aware Supervision

Sep 28, 2026

This study addresses critical bottlenecks in the continual training of medical multimodal large language models, including data inefficiency, insufficient supervision signals, and poor output consistency, by proposing the InfiMed2 framework. Methodologically, we construct a 55.68-billion-token corpus and design a phase-aware data strategy to optimize pretraining. Furthermore, we introduce an answer-stability-based SFT reconstruction mechanism combined with evidence-focused mixing during learning rate decay, alongside RLVR to enhance fine-tuning quality. Ultimately, we release 4B and 27B general-purpose medical multimodal foundation models. The 4B model achieves an average score of 66.73%, surpassing Qwen3.5-9B, while the 27B model attains 73.72%, establishing new state-of-the-art performance among open-source counterparts.

0 citationsRead paper

Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs

Sep 26, 2026

This study addresses whether routing drift in Mixture-of-Experts (MoE) model merging is equivalent to routing failure and when intervention is necessary. We propose redefining routing failure through the lens of task loss recoverability. Leveraging a cross-model routing analysis toolkit, we employ counterfactual interventions and token-level attribution to demonstrate that most expert reassignments stem from input shifts rather than functional degradation. Based on these findings, we introduce a Selective Router Repair (SRR) strategy. This work challenges the prevailing misconception that drift inherently implies failure, showing that naively restoring source routing does not reliably improve performance. By releasing open-source tools and code, we provide theoretical foundations for precise post-merging diagnosis and repair in MoE architectures.

0 citationsRead paper