Institution profile

ML Collective

Research institutionnorthamerica · us
Official website
Research library13linked papers
Opportunities0open roles
Selected work

Representative Papers

A Tilted Bowl Is Not a Slippery Slope: Compressing Looped Models

Sep 30, 2026

This study addresses the performance collapse in recurrent model compression caused by the accumulation of misjudged rounding errors. To this end, it proposes the "tilted bowl" theory, which reveals that fixed rounding errors within convergent loops merely shift the stopping point rather than accumulating progressively. Building upon this theoretical insight, the work integrates quantization-aware training with 8-bit mixed-precision inference to construct a dynamic depth controller based on unlabeled measurements, enabling effective error failure prediction and recovery. Evaluated on Sudoku and Maze tasks, the proposed method surpasses fixed-depth inference baselines by up to 15 points while requiring lower weight traffic, thereby significantly enhancing the inference efficiency of recurrent neural networks.

0 citationsRead paper

Spontaneous Context Restoration: How Language Models Recover from Corrupted Inputs

Sep 28, 2026

This study addresses the phenomenon whereby large language models generate correct outputs despite corrupted inputs, a process whose internal mechanisms remain poorly understood. By integrating residual stream analysis, attention probing, and fine-tuning techniques, this work reveals a two-stage repair process within the model and identifies an unsupervised spontaneous recovery mechanism. Building upon these findings, we propose a novel paradigm that leverages the hidden states of the first token for efficient fault triage. Experimental results demonstrate that this approach achieves a fault prediction performance with an ROC-AUC of 0.78. Furthermore, moderate fine-tuning significantly enhances model robustness while reducing nonlinear errors.

0 citationsRead paper

PRISM: A Predictive Protocol for Permutation Optimization via Landscape Diagnostics

Aug 08, 2026

This work addresses the challenge of efficiently determining whether permutation optimization is worthwhile—and which search strategy to employ—in scenarios where system components are fixed but their ordering significantly impacts performance. The authors propose PRISM, a novel protocol that introduces fitness landscape diagnostics into permutation optimization for the first time. By leveraging low-cost first-order autocorrelation and fitness-distance correlation analyses, PRISM predicts search behavior prior to optimization, thereby guiding the selection of appropriate search strategies, operators, and simplification schemes. Empirical validation across diverse tasks—including neural architecture design, scientific machine learning, and large-model instruction sequencing—demonstrates that PRISM accurately forecasts optimization outcomes and confirms the complementary nature of permutation and content optimization.

0 citationsRead paper

Towards Digital Preservation of Efik: TTS for a Low-Resource African Language

Jul 05, 2026

This study addresses the challenges of speech synthesis and digital preservation for Efik, a low-resource African tonal language. The authors present the first end-to-end text-to-speech (TTS) system for Efik, leveraging a newly collected 3-hour single-speaker speech corpus. They establish reproducible TTS baselines using VITS, MMS-TTS, SpeechT5, and Orpheus-TTS, and evaluate performance through subjective assessments including MOS, Nat-MOS, and A-MOS. Among the models, MMS-TTS achieves the highest quality (MOS: 3.80 ± 0.63) and demonstrates greater stability in synthesizing long utterances, though it still exhibits tonal inaccuracies. This work provides the first systematic evaluation framework for TTS in low-resource tonal languages and underscores the need for larger-scale corpora and improved tonal modeling to advance synthesis quality.

0 citationsRead paper
Recent publications

Latest Papers

A Tilted Bowl Is Not a Slippery Slope: Compressing Looped Models

Sep 30, 2026

This study addresses the performance collapse in recurrent model compression caused by the accumulation of misjudged rounding errors. To this end, it proposes the "tilted bowl" theory, which reveals that fixed rounding errors within convergent loops merely shift the stopping point rather than accumulating progressively. Building upon this theoretical insight, the work integrates quantization-aware training with 8-bit mixed-precision inference to construct a dynamic depth controller based on unlabeled measurements, enabling effective error failure prediction and recovery. Evaluated on Sudoku and Maze tasks, the proposed method surpasses fixed-depth inference baselines by up to 15 points while requiring lower weight traffic, thereby significantly enhancing the inference efficiency of recurrent neural networks.

0 citationsRead paper

Spontaneous Context Restoration: How Language Models Recover from Corrupted Inputs

Sep 28, 2026

This study addresses the phenomenon whereby large language models generate correct outputs despite corrupted inputs, a process whose internal mechanisms remain poorly understood. By integrating residual stream analysis, attention probing, and fine-tuning techniques, this work reveals a two-stage repair process within the model and identifies an unsupervised spontaneous recovery mechanism. Building upon these findings, we propose a novel paradigm that leverages the hidden states of the first token for efficient fault triage. Experimental results demonstrate that this approach achieves a fault prediction performance with an ROC-AUC of 0.78. Furthermore, moderate fine-tuning significantly enhances model robustness while reducing nonlinear errors.

0 citationsRead paper

PRISM: A Predictive Protocol for Permutation Optimization via Landscape Diagnostics

Aug 08, 2026

This work addresses the challenge of efficiently determining whether permutation optimization is worthwhile—and which search strategy to employ—in scenarios where system components are fixed but their ordering significantly impacts performance. The authors propose PRISM, a novel protocol that introduces fitness landscape diagnostics into permutation optimization for the first time. By leveraging low-cost first-order autocorrelation and fitness-distance correlation analyses, PRISM predicts search behavior prior to optimization, thereby guiding the selection of appropriate search strategies, operators, and simplification schemes. Empirical validation across diverse tasks—including neural architecture design, scientific machine learning, and large-model instruction sequencing—demonstrates that PRISM accurately forecasts optimization outcomes and confirms the complementary nature of permutation and content optimization.

0 citationsRead paper

Towards Digital Preservation of Efik: TTS for a Low-Resource African Language

Jul 05, 2026

This study addresses the challenges of speech synthesis and digital preservation for Efik, a low-resource African tonal language. The authors present the first end-to-end text-to-speech (TTS) system for Efik, leveraging a newly collected 3-hour single-speaker speech corpus. They establish reproducible TTS baselines using VITS, MMS-TTS, SpeechT5, and Orpheus-TTS, and evaluate performance through subjective assessments including MOS, Nat-MOS, and A-MOS. Among the models, MMS-TTS achieves the highest quality (MOS: 3.80 ± 0.63) and demonstrates greater stability in synthesizing long utterances, though it still exhibits tonal inaccuracies. This work provides the first systematic evaluation framework for TTS in low-resource tonal languages and underscores the need for larger-scale corpora and improved tonal modeling to advance synthesis quality.

0 citationsRead paper