pmi-based token reweighting

Designs and implements algorithms that compute pointwise mutual information between tokens and contextual states to produce reweighting functions applied to token scores or probabilities. This includes methods to suppress globally frequent stopwords and fillers, elevate context‑evoked candidate tokens, and reshape model head probabilities prior to truncation.

pmi-basedtokenreweighting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.19
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

PoRe: Position-Reweighted Visual Token Pruning for Vision Language Models

Aug 25, 2025
KZ
Kai Zhao
🏛️ Shanghai University | University of California, Los Angeles

Visual-language models (VLMs) commonly prune visual tokens based on text–vision attention scores; however, this practice suffers from the “recency bias” inherent in sequence models, leading to excessive retention of bottom-region image tokens and thus imbalanced pruning. This bias stems from the coupling between spatial image positions and token ordering in the sequence. To address this, we propose a position-aware attention reweighting mechanism that requires no architectural modification or additional training: it explicitly encodes spatial coordinates and dynamically recalibrates attention scores to suppress position-induced bias. Our method is fully compatible with mainstream token pruning frameworks and is empirically validated across multiple VLMs. It consistently improves post-pruning performance—e.g., boosting VQAv2 accuracy by +1.8%—while incurring negligible computational overhead. The approach offers a simple, general, and plug-and-play optimization for efficient VLM inference.

Addresses recency bias in visual token pruningImproves pruning by reweighting attention scores spatiallyReduces redundant visual tokens in VLMs

Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?

Feb 17, 2025
ZW
Zichen Wen
🏛️ Shanghai Jiao Tong University | Shanghai AI Laboratory | Sun Yat-sen University

Token pruning in multimodal large language models (MLLMs) aims to reduce inference overhead, yet existing methods suffer from fundamental flaws—including unreliable attention-based scoring, limited linguistic contribution, suboptimal redundancy-repetition trade-offs, and biased evaluation protocols. Method: This work is the first to systematically challenge both pruning design principles and evaluation paradigms, proposing a novel assessment framework that jointly ensures interpretability and fairness. We develop a diagnostic evaluation protocol grounded in theoretical analysis, controlled ablation experiments, and cross-model validation (ViT-L/LLaMA-2/3). Contribution/Results: Key findings reveal that most state-of-the-art pruning methods perform no better than random pruning; visual tokens constitute the primary source of redundancy; and linguistic information exhibits sharply diminishing marginal utility under pruning. These insights establish a rigorous theoretical foundation and practical benchmark for efficient MLLM inference.

Analyze language information impactAssess attention-based scoring adequacyEvaluate token pruning efficiency

This work addresses the severe inference latency in high-resolution multimodal large language models caused by the explosion in visual token count. Existing pruning methods rely on iterative optimization, hindering efficient acceleration. To overcome this limitation, the authors propose SFPruner (Single-Forward Pruner), which, for the first time, integrates structured redundancy modeling into visual token scoring. SFPruner employs a semantics-guided ridge leverage mechanism to suppress covariance-dominant directions and incorporates a ranking-based directional mask to enable asymmetric similarity competition. This enables non-iterative pruning in a single forward pass, effectively balancing semantic diversity and instruction relevance. Evaluated on Qwen2.5-VL, the method reduces the selection time for 512 tokens from 112.4 ms to 2.5 ms, achieving substantial inference speedup while maintaining performance comparable to state-of-the-art approaches.

inference latencymultimodal large language modelsstructured redundancy

Latest Papers

What's happening recently
View more

This work addresses the inefficiency of multimodal large language models caused by redundant visual token prefixes. Existing approaches struggle to balance performance and efficiency due to their assumptions of token independence and fixed compression ratios. To overcome these limitations, this study formulates visual token pruning as a conditional, dependency-aware sequential decision problem. It introduces a pointer-based selection mechanism that iteratively identifies high-information tokens and incorporates a learnable termination action to dynamically determine the optimal compression scale. Discrete selection is made differentiable via variance-preserving noisy interpolation, enabling end-to-end training. Experiments on LLaVA-v1.5-7B and Qwen2.5-VL-7B show that retaining only 11.1% of visual tokens achieves 94.6% of the original model’s accuracy while reducing prefill latency by 1.88×, substantially outperforming fixed-ratio baselines.

compression ratioinference efficiencymultimodal large language models

Existing metric-based approaches for detecting machine-generated text are susceptible to the randomness inherent in generative processes, leading to biased token-level detection scores. This work is the first to reveal that these scores exhibit a multi-hop transition property and proposes a multi-level contextual token relationship modeling framework to address this issue. The framework integrates a lightweight Markov information calibration module to correct local biases and combines it with explicit, context-statistics-based logical rules for global reasoning, enabling joint optimization. The proposed method significantly enhances detection performance across diverse large language models and domains while maintaining low computational overhead.

detection biasgeneration randomnessmachine-generated text detection

This study addresses the unclear generalization mechanisms of next-token prediction under Markov data. To this end, it constructs an information-theoretic framework that decouples algorithmic effects from temporal dependencies. By introducing rate-distortion theory to accommodate continuous hypothesis spaces and integrating the Donsker-Varadhan variational representation, McDiarmid’s inequality, and noisy low-dimensional compression techniques, the authors derive generalization bounds for both cross-entropy and margin-based predictions. The work elucidates how mixing properties influence generalization and reveals that extended contexts widen the generalization gap. Furthermore, experiments on the ETTh2 dataset empirically validate 24 hours as an effective memory horizon.

context lengthgeneralization boundinformation theory

This work addresses a critical limitation in existing token-averaging–based text detection methods, which are susceptible to Simpson’s paradox due to their neglect of heterogeneous likelihood score distributions in the latent space, often obscuring strong local signals and impairing discrimination between human- and large language model–generated text. To mitigate this, the authors propose a Bayesian decision–theoretic local calibration mechanism that employs a lightweight conditional distribution predictor to recalibrate latent-space positions and replaces raw scores with calibrated log-likelihood ratios for aggregation. The approach seamlessly integrates into any token-averaging detection pipeline and achieves substantial performance gains across multiple baselines and datasets—for instance, boosting Fast-DetectGPT’s AUROC on GPT-4–generated text from 0.63 to 0.85, establishing a new state-of-the-art detection efficacy.

DetectionLog-LikelihoodMachine-Generated Text

Hot Scholars

AM

Andrew McCallum

Distinguished Professor of Computer Science, University of Massachusetts Amherst
Machine LearningNatural Language ProcessingArtificial IntelligencePeer Review
RL

Ruixuan Li

Professor of Computer Science, Huazhong University of Science and Technology
Distributed systemssecurity and privacydata management
YT

Yuchuang Tong

Institute of Automation Chinese Academy of Sciences
Embodied IntelligenceHumanoid RobotsRobotic Intelligent ControlRobotic Learning
JG

Jindong Gu

Google Research & DeepMind, University of Oxford
Trustworthy AIAI SafetyMultimodal AI
LC

Liuji Chen

Institute of Automation, Chinese Academy of Sciences
LLM AgentTrustworthy AI