Institution profile

Hanbat National University

Academic institutionasia · kr
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

Beyond Reconstruction Loss in Post-Training Quantization: Balanced Fitting for Large Vision-Language Models

Sep 28, 2026

This study addresses the generalization bias in post-training quantization of large vision-language models (LVLMs) caused by an over-reliance on reconstruction loss. To overcome this limitation, we propose Balanced Fitting, a framework that departs from the conventional error-minimization paradigm by exploiting the regularization benefits that quantization confers upon specific layers and modalities. Through fine-grained evaluation of component-wise quantization effects, a hybrid fitting strategy, and joint weight-activation quantization, our approach dynamically balances accuracy preservation with regularization gains. Extensive experiments demonstrate that the proposed method significantly outperforms existing baselines across diverse LVLM architectures. These findings compellingly establish that low reconstruction loss does not necessarily translate to superior downstream performance, thereby introducing a new paradigm for multimodal model quantization.

0 citationsRead paper

When Text Matters: Design Principles for Visual Token Pruning in Vision-Language Model

Sep 28, 2026

This study addresses the significant performance degradation in existing visual token pruning methods caused by early text guidance, which restricts the identification of answer-relevant regions. To overcome this limitation, we propose a training-free two-stage pruning strategy that innovatively decouples vision-guided pruning from delayed text-guided re-selection. Specifically, the method first performs preliminary pruning using visual encoder attention, and subsequently refines the final token set via text-to-vision attention at intermediate decoder layers. Extensive experiments across three models and eight benchmarks demonstrate that our approach recovers an average of 11.10 and 16.84 percentage points in performance at 80% and 90% pruning ratios, respectively, while effectively reducing inference latency.

0 citationsRead paper

CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG

Aug 07, 2026

Traditional RAG systems struggle with information redundancy and noise when processing long contexts, and coarse-grained block-level KV cache reuse fails to simultaneously achieve low prefill latency and high answer accuracy. This work proposes a fine-grained RAG approach that identifies query-relevant semantic units—termed “information nuggets”—through a two-stage retrieval process, then integrates their sliced KV representations with block-level context to construct a compact, semantically focused context representation. The method introduces an offline fine-grained KV cache reuse mechanism, which, under standard fast prefill latency constraints, improves average F1 by 5.3% on LongBench multi-hop question answering tasks while significantly reducing computational overhead, thereby advancing beyond the current Pareto frontier of efficiency and accuracy in RAG systems.

0 citationsRead paper

SPARK: Self-Play with Asymmetric Reward from Knowledge Graphs

May 06, 2026

This work addresses the challenge of automatically generating trustworthy multi-hop reasoning questions from scientific literature, where relationships among multimodal elements are often implicit and difficult to verify. To this end, it introduces knowledge graphs into a self-play framework for scientific documents, constructing a unified graph to generate multi-hop relational questions and providing verifiable reward signals grounded in structured factual knowledge. By leveraging an information asymmetry mechanism, a single small-scale vision-language model alternately assumes the roles of questioner and answerer during training. The proposed approach significantly outperforms text-only self-play baselines on both public benchmarks and a newly curated cross-document multi-hop question answering dataset, with performance gains becoming more pronounced as the number of reasoning hops increases.

0 citationsRead paper
Recent publications

Latest Papers

Beyond Reconstruction Loss in Post-Training Quantization: Balanced Fitting for Large Vision-Language Models

Sep 28, 2026

This study addresses the generalization bias in post-training quantization of large vision-language models (LVLMs) caused by an over-reliance on reconstruction loss. To overcome this limitation, we propose Balanced Fitting, a framework that departs from the conventional error-minimization paradigm by exploiting the regularization benefits that quantization confers upon specific layers and modalities. Through fine-grained evaluation of component-wise quantization effects, a hybrid fitting strategy, and joint weight-activation quantization, our approach dynamically balances accuracy preservation with regularization gains. Extensive experiments demonstrate that the proposed method significantly outperforms existing baselines across diverse LVLM architectures. These findings compellingly establish that low reconstruction loss does not necessarily translate to superior downstream performance, thereby introducing a new paradigm for multimodal model quantization.

0 citationsRead paper

When Text Matters: Design Principles for Visual Token Pruning in Vision-Language Model

Sep 28, 2026

This study addresses the significant performance degradation in existing visual token pruning methods caused by early text guidance, which restricts the identification of answer-relevant regions. To overcome this limitation, we propose a training-free two-stage pruning strategy that innovatively decouples vision-guided pruning from delayed text-guided re-selection. Specifically, the method first performs preliminary pruning using visual encoder attention, and subsequently refines the final token set via text-to-vision attention at intermediate decoder layers. Extensive experiments across three models and eight benchmarks demonstrate that our approach recovers an average of 11.10 and 16.84 percentage points in performance at 80% and 90% pruning ratios, respectively, while effectively reducing inference latency.

0 citationsRead paper

CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG

Aug 07, 2026

Traditional RAG systems struggle with information redundancy and noise when processing long contexts, and coarse-grained block-level KV cache reuse fails to simultaneously achieve low prefill latency and high answer accuracy. This work proposes a fine-grained RAG approach that identifies query-relevant semantic units—termed “information nuggets”—through a two-stage retrieval process, then integrates their sliced KV representations with block-level context to construct a compact, semantically focused context representation. The method introduces an offline fine-grained KV cache reuse mechanism, which, under standard fast prefill latency constraints, improves average F1 by 5.3% on LongBench multi-hop question answering tasks while significantly reducing computational overhead, thereby advancing beyond the current Pareto frontier of efficiency and accuracy in RAG systems.

0 citationsRead paper

SPARK: Self-Play with Asymmetric Reward from Knowledge Graphs

May 06, 2026

This work addresses the challenge of automatically generating trustworthy multi-hop reasoning questions from scientific literature, where relationships among multimodal elements are often implicit and difficult to verify. To this end, it introduces knowledge graphs into a self-play framework for scientific documents, constructing a unified graph to generate multi-hop relational questions and providing verifiable reward signals grounded in structured factual knowledge. By leveraging an information asymmetry mechanism, a single small-scale vision-language model alternately assumes the roles of questioner and answerer during training. The proposed approach significantly outperforms text-only self-play baselines on both public benchmarks and a newly curated cross-document multi-hop question answering dataset, with performance gains becoming more pronounced as the number of reasoning hops increases.

0 citationsRead paper