A two-step sequential approach for hyperparameter selection in finite context models

📅 2026-03-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the computationally expensive joint optimization of the context length \(k\) and smoothing parameter \(\alpha\) in finite-context models (FCMs), which typically relies on exhaustive search. The authors propose a two-stage sequential estimation approach: first, they efficiently estimate the optimal \(k\) using sequence dependence measures—such as Cramér’s ν, Cohen’s κ, and partial mutual information—and then estimate \(\alpha\) via conditional maximum likelihood. By decoupling the joint optimization into two independent steps, the method substantially improves tuning efficiency. Experimental results on synthetic symbolic sequences demonstrate that the proposed approach achieves compression performance (in bits per symbol) comparable to exhaustive search while significantly reducing computational cost, with estimation accuracy improving as sample size increases.

Technology Category

Natural Language Processing: Learning & Optimization for NLPSearch and Optimization: Learning to SearchMachine Learning: Optimization

Application Category

User Modeling, Personalization and Recommendation: Practical large-scale studies of user experienceSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
Finite-context models (FCMs) are widely used for compressing symbolic sequences such as DNA, where predictive performance depends critically on the context length k and smoothing parameter α. In practice, these hyperparameters are typically selected through exhaustive search, which is computationally expensive and scales poorly with model complexity. This paper proposes a statistically grounded two-step sequential approach for efficient hyperparameter selection in FCMs. The key idea is to decompose the joint optimization problem into two independent stages. First, the context length k is estimated using categorical serial dependence measures, including Cramér's ν, Cohen's \k{appa} and partial mutual information (pami). Second, the smoothing parameter α is estimated via maximum likelihood conditional on the selected context length k. Simulation experiments were conducted on synthetic symbolic sequences generated by FCMs across multiple (k, α) configurations, considering a four-letter alphabet and different sample sizes. Results show that the dependence measures are substantially more sensitive to variations in k than in α, supporting the sequential estimation strategy. As expected, the accuracy of the hyperparameter estimation improves with increasing sample size. Furthermore, the proposed method achieves compression performance comparable to exhaustive grid search in terms of average bitrate (bits per symbol), while substantially reducing computational cost. Overall, the results on simulated data show that the proposed sequential approach is a practical and computationally efficient alternative to exhaustive hyperparameter tuning in FCMs.
Problem

Research questions and friction points this paper is trying to address.

finite-context models
hyperparameter selection
context length
smoothing parameter
symbolic sequence compression
Innovation

Methods, ideas, or system contributions that make the work stand out.

finite-context models
hyperparameter selection
sequential estimation
context length
smoothing parameter
🔎 Similar Papers
2024-07-08arXiv.orgCitations: 1
J
José Contente
Institute of Electronics and Informatics Engineering of Aveiro (IEETA), Department of Electronics, Telecommunications and Informatics (DETI), University of Aveiro (UA), Aveiro, Portugal
Ana Martins
Ana Martins
Visiting Scholar, UC Berkeley
Systems BiologyMetabolic NetworksEnzymologyDrug-delivery nanocarriers
Armando J. Pinho
Armando J. Pinho
IEETA/DETI, University of Aveiro
Data CompressionImage CompressionKolmogorov ComplexityAlgorithmic Information Theory
Sónia Gouveia
Sónia Gouveia
University of Aveiro