denoising event sequence decoding

Designs and implements decoders that transduce noisy model outputs (e.g., per-entity logits or corrupted token sequences) into ordered, structured event sequences by learning denoising objectives that recover event attributes, order, and timing. Work includes architecting decoding mechanisms and attention/ordering strategies (for example spatial- or entity-first ordering and multi-stage per-entity attention), conditioning on contextual state features, and measuring decoding fidelity and error modes.

denoisingeventsequencedecoding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.57
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the limited robustness of current event detection systems in real-world noisy scenarios for low-resource languages such as Bengali, despite their strong performance on clean text. The authors construct a Bengali news event benchmark dataset comprising 9,979 annotated sentences spanning clean text, ASR transcripts, and spelling-perturbed variants. They systematically evaluate the noise robustness of encoder-based models (e.g., BanglaBERT, XLM-R) and instruction-tuned decoder large language models (e.g., Llama 3, Gemma 2). The work reveals a fundamental difference in noise resilience between the two architectures: decoders exhibit greater robustness when trigger words are corrupted. By integrating annotation guidelines into instruction tuning and training on mixed clean-noisy data, the performance gap is substantially narrowed. Combining model scaling with multi-source joint training yields state-of-the-art results across diverse noise conditions.

Banglaevent detectionlow-resource languages

This work challenges the conventional left-to-right generation order in autoregressive speech synthesis, whose optimality remains unverified. Leveraging a masked diffusion framework, the proposed approach enables arbitrary decoding orders during both training and inference, allowing for a systematic evaluation of fixed versus adaptive strategies on synthesis quality. By incorporating discrete acoustic representations, a position-wise progressive demasking mechanism, and a Top-K adaptive decoding strategy, the study demonstrates that decoding order significantly influences speech fidelity, with adaptive strategies consistently outperforming fixed ones. Notably, high-fidelity speech can still be generated under an extremely low 1-bit quantization condition, underscoring the method’s efficiency and robustness.

adaptive decodingautoregressive speech synthesisdecoding order

This work addresses two key limitations in event detection using decoder-only large language models: their inability to model bidirectional contextual information and the bias of Micro-F1 toward head classes, which obscures performance on long-tailed event categories. To overcome these issues, the authors propose a novel architecture that integrates a context-aware encoder with LoRA-based fine-tuning, adopting Macro-F1 as the primary evaluation metric. This approach represents the first application of combining context-aware encoding with LoRA for event detection, substantially enhancing the model’s capacity to capture bidirectional semantic dependencies. As a result, it effectively mitigates the underrepresentation of minority event types in evaluation and achieves significant improvements in Macro-F1 scores on long-tail classes.

bidirectional contextdecoder-only LLMsevent detection

This work addresses the limitations of traditional autoregressive approaches in dense video captioning, which suffer from low inference efficiency and poor scalability in long videos with high event density. The authors propose a parallelized autoregressive framework that restructures the causal dependency graph to explicitly model weak local dependencies among events. By introducing a latent global planning mechanism, the model automatically learns event structure and aggregates audio-visual semantics, enabling cross-event parallel generation through factorized event decoding. This approach maintains local semantic coherence while significantly improving inference efficiency and temporal grounding accuracy. The method achieves state-of-the-art results across multiple benchmarks, demonstrating particularly strong performance in multimodal event localization and description tasks.

autoregressive decodingdense video captioningevent density

Improving Factuality in Large Language Models via Decoding-Time Hallucinatory and Truthful Comparators

Aug 22, 2024
DY
Dingkang Yang
🏛️ Fudan University | ByteDance | Tencent

Large language models (LLMs) frequently generate factually inconsistent hallucinations, undermining their reliability. To address this, we propose a decoding-time dual-comparator framework—termed the Hallucination/Truth Comparator—that dynamically constrains the token prediction distribution during autoregressive generation, enabling factuality calibration without fine-tuning or architectural modification. Our approach innovatively integrates an instruction-prototype-guided Mixture-of-Experts (MoE) mechanism with a logit-difference comparison strategy, establishing a task-agnostic factual constraint paradigm. Extensive experiments across diverse downstream tasks demonstrate substantial improvements in response factuality and overall performance, while preserving the original model’s capabilities. The method is fully plug-and-play, requiring no retraining or parameter updates.

FactualityLarge Language ModelsReliability

Latest Papers

What's happening recently
View more

Soft-decision decoding remains challenging in terms of universality, efficiency, and the need for signal-to-noise ratio (SNR) estimation. This work proposes a novel framework that, for the first time, formulates error-correcting code decoding as a continuous-time denoising process via a score-matching-based neural probability flow ordinary differential equation (ODE). The approach trains directly on raw signed channel observations without requiring SNR conditioning and leverages ODE solvers to flexibly trade off decoding latency against accuracy. By incorporating parity-check constraints and employing Euler and DPM solvers, the method achieves the lowest bit error rate in 39 out of 42 code–SNR configurations, yielding an average SNR gain of 0.17 dB (up to 0.46 dB). Switching to the DPM solver further reduces decoding time by 8.86% on average (up to 12.82%) while preserving performance.

bit error ratedecoding latencyerror-correcting codes

This work addresses the absence of a unified theoretical framework for autoregressive decoding strategies in speech processing, which has led to ambiguous definitions, inconsistent taxonomies, and difficulties in fair comparison. The paper introduces, for the first time, a general formal framework that precisely specifies inclusion criteria for autoregressive search and systematically categorizes and describes decoding strategies employed in neural speech generation models. By clarifying conceptual boundaries, the framework enhances comparability and evaluation consistency across strategies, streamlines the design of decoding-centric benchmarking protocols, and enables ablation studies focused specifically on search mechanisms. Consequently, it facilitates standardized analysis of inference-stage behavior in speech generation models.

auto-regressive decodingdecoding frameworksearch strategies

This work addresses the challenge that order-agnostic language models (OALMs) exhibit path-dependent artifacts in likelihood estimation across different revelation orders, obscuring true content difficulty. The authors propose the variance of confidence trajectories as a novel diagnostic metric for decoding path quality and theoretically demonstrate that, under fixed total likelihood, uniform per-step confidence maximizes target recoverability. Building upon the discrete diffusion language model (dLLM) framework and integrating confidence-first (CF) decoding with chain-rule bias analysis, they empirically validate on C4 and four downstream tasks that low confidence variance effectively identifies structured decoding paths and exhibits a strong positive correlation with task accuracy.

confidence variancedecoding pathslikelihood inconsistency

Diffusion language models struggle to achieve truly order-agnostic generation in fast parallel decoding due to their sensitivity to denoising order. This work formalizes the resulting “order collapse” problem for the first time as an issue of compatibility among local conditional distributions. Adopting a non-conservative field perspective, it introduces order-induced pseudo-joint distributions and local denoising circulations to uncover the root cause of path dependence. Building upon probabilistic graphical models and circulation decomposition, the paper proposes a theoretical framework that enables diagnosis of order-freedom solely at inference time. This framework effectively disentangles path dependence from two distinct error sources: conditional dependency errors arising from parallel updates and order-specific estimation errors, thereby providing the first quantifiable analytical tool for evaluating the order-agnostic properties of diffusion language models.

denoising compatibilitydiffusion language modelsnon-conservative field

Hot Scholars

YL

Yonghui Li

the University of Sydney
Wireless communicationsChannel codingInternet of ThingsSignal Processing
BV

Branka Vucetic

The University of Sydney
wireless communicationscoding theory
CY

Chentao Yue

The University of Sydney
Coding TheoryInformation Theory
ZY

Zhifei Yang

Peking University
3D GenerationGenerative Models
BJ

Beihong Jin

Institute of Software, Chinese Academy of Sciences
Pervasive ComputingDistributed Computing