Institution profile

Korea Telecom

Industry researchasia · kr
Official website
Research library58linked papers
Opportunities0open roles
Selected work

Representative Papers

Mi:dm 2.0 Korea-centric Bilingual Language Models

Jan 14, 2026

This work addresses the limitations of current large language models in handling Korean, which stem from low-quality training data and a lack of cultural alignment, hindering their ability to capture Korea-specific values, commonsense knowledge, and nuanced emotional expressions. To overcome these challenges, we propose Mi:dm 2.0—the first bilingual large language model systematically integrating Korean sociocultural commonsense and reasoning patterns. Through high-quality data curation, synthetic data generation, a curriculum learning–guided data mixing strategy, and a Korean-optimized tokenizer, Mi:dm 2.0 achieves deep contextual understanding of local nuances. Released under the MIT License in both general and lightweight variants, the model attains state-of-the-art zero-shot performance on Korean benchmarks such as KMMLU, significantly outperforming existing models and advancing the development of the K-intelligence ecosystem.

1 citationsRead paper

Guide-to-Explain for Controllable Summarization

Nov 19, 2024arXiv.org

Large language models (LLMs) exhibit limited precision in controlling numerical attributes—such as summary length and extractiveness—in controllable summarization, hindering practical deployment aligned with user preferences. To address this, we propose a Guided Reflection Framework featuring a novel two-stage self-reflective mechanism: “Guide–Explain.” First, attribute-aware bias detection identifies deviations between the initial summary and target constraints; second, an attributional error explanation is generated to guide conditional regeneration. Our approach integrates self-reflective prompt engineering with multi-attribute joint constraints, significantly improving control fidelity and optimization efficiency. Experiments on multidimensional controllable summarization demonstrate substantial gains: constraint satisfaction rates increase markedly, and average iteration counts decrease by over 40% compared to state-of-the-art pure-LLM iterative baselines.

1 citationsRead paper

Multimodal Safety Evaluation Should Measure Controllability Beyond Classification

Oct 05, 2026

This study addresses the limitation of traditional multimodal safety evaluations, which rely solely on behavioral classification and fail to reveal the controllability of models' internal safety mechanisms. To this end, we propose a "controllability profiling" framework that leverages sparse autoencoders (SAEs) to establish internal representation controllability as an independent evaluation dimension for the first time, quantifying both the detectability of safety signals and their intervention sensitivity in vision-language models. Implicit toxicity stress tests conducted on LlavaGuard and Qwen3.5 demonstrate significant discrepancies across models regarding the alignment between internal readout capabilities and selective control. By exposing these divergences, this work provides a critical theoretical foundation for developing next-generation multimodal safety benchmarks.

0 citationsRead paper

Behavior-Preserving KV Cache Compression

Oct 05, 2026

This study addresses the KV cache bottleneck in long-context inference for large language models and the performance degradation caused by existing methods’ reliance on proxy signals. We propose a training-free KV cache compression framework that introduces, for the first time, the principle of preserving model prediction consistency as its core criterion. Specifically, critical cache entries are selected by evaluating the impact of candidate removals on the output distribution. Furthermore, we design a pre-eviction statistic reuse technique to eliminate the overhead of multiple forward passes. Experimental results demonstrate that, under identical cache budgets, the proposed framework significantly improves downstream task quality, with particularly pronounced advantages in aggressive compression scenarios, while simultaneously achieving end-to-end inference acceleration.

0 citationsRead paper

ODDR: One-Step Deshadow Diffusion via Reward Guidance

Oct 01, 2026

This work addresses the limited generalization of shadow removal methods caused by their reliance on costly real paired data by proposing the ODDR framework. Specifically, it trains a one-step shadow removal diffusion model exclusively on synthetic data. To bridge the domain gap between synthetic and real images, this study introduces ShadowReward, a novel reward model that eliminates the need for manual annotations by ranking synthetically degraded images to simulate human perception, thereby guiding reinforcement learning fine-tuning. By effectively aligning synthetic training with real-world distributions, ODDR achieves shadow removal performance comparable to fully supervised approaches while preserving single-step inference efficiency. Consequently, this method significantly improves image restoration quality in scenarios where paired real data is unavailable.

0 citationsRead paper
Recent publications

Latest Papers

Multimodal Safety Evaluation Should Measure Controllability Beyond Classification

Oct 05, 2026

This study addresses the limitation of traditional multimodal safety evaluations, which rely solely on behavioral classification and fail to reveal the controllability of models' internal safety mechanisms. To this end, we propose a "controllability profiling" framework that leverages sparse autoencoders (SAEs) to establish internal representation controllability as an independent evaluation dimension for the first time, quantifying both the detectability of safety signals and their intervention sensitivity in vision-language models. Implicit toxicity stress tests conducted on LlavaGuard and Qwen3.5 demonstrate significant discrepancies across models regarding the alignment between internal readout capabilities and selective control. By exposing these divergences, this work provides a critical theoretical foundation for developing next-generation multimodal safety benchmarks.

0 citationsRead paper

Behavior-Preserving KV Cache Compression

Oct 05, 2026

This study addresses the KV cache bottleneck in long-context inference for large language models and the performance degradation caused by existing methods’ reliance on proxy signals. We propose a training-free KV cache compression framework that introduces, for the first time, the principle of preserving model prediction consistency as its core criterion. Specifically, critical cache entries are selected by evaluating the impact of candidate removals on the output distribution. Furthermore, we design a pre-eviction statistic reuse technique to eliminate the overhead of multiple forward passes. Experimental results demonstrate that, under identical cache budgets, the proposed framework significantly improves downstream task quality, with particularly pronounced advantages in aggressive compression scenarios, while simultaneously achieving end-to-end inference acceleration.

0 citationsRead paper

ODDR: One-Step Deshadow Diffusion via Reward Guidance

Oct 01, 2026

This work addresses the limited generalization of shadow removal methods caused by their reliance on costly real paired data by proposing the ODDR framework. Specifically, it trains a one-step shadow removal diffusion model exclusively on synthetic data. To bridge the domain gap between synthetic and real images, this study introduces ShadowReward, a novel reward model that eliminates the need for manual annotations by ranking synthetically degraded images to simulate human perception, thereby guiding reinforcement learning fine-tuning. By effectively aligning synthetic training with real-world distributions, ODDR achieves shadow removal performance comparable to fully supervised approaches while preserving single-step inference efficiency. Consequently, this method significantly improves image restoration quality in scenarios where paired real data is unavailable.

0 citationsRead paper

Concepts Complement Dense Semantics: Learning Compact Sparse Spaces for Text-Image Retrieval

Sep 26, 2026

This study addresses the limitation of dense embedding spaces in obscuring fine-grained visual-textual information, which leads to insufficient cross-modal retrieval accuracy and a lack of explicit semantic grounding. To overcome this, we propose GRASP, a framework that leverages corpus-level concept mining to extract visual-textual concepts and construct a compact, conceptually sparse learning space. By integrating vision-language pre-trained models with a lightweight sparse prediction head, GRASP augments dense semantic matching with interpretable sparse evidence. This approach circumvents the reliance on redundant token spaces inherent in traditional methods. Extensive experiments demonstrate that GRASP surpasses state-of-the-art baselines in retrieval accuracy, achieving interpretable cross-modal retrieval that simultaneously delivers high precision, structural compactness, and explicit semantic grounding.

0 citationsRead paper