Institution profile

Artificial Intelligence Research Institute

Academic institutionasia · jp
Official website
Research library18linked papers
Opportunities0open roles
Selected work

Representative Papers

The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning

Aug 14, 2026

This study addresses the limitation of existing large language model unlearning methods that overlook fact popularity, rendering high-frequency knowledge difficult to remove. We propose AdaPop, a novel approach that models popularity as a learnable parameter by integrating token confidence with external proxy evaluations. Through a bi-ascent controller, AdaPop dynamically adjusts penalty intensity to achieve an adaptive balance between forgetting and retention. Experiments across three model families and two benchmarks demonstrate that AdaPop reduces content leakage under paraphrased queries by approximately fivefold and under adversarial reconstruction by 1.6 times. Furthermore, internal representation analysis confirms superior forgetting separation, effectively resolving the challenge of unlearning high-frequency facts while preserving general model utility.

0 citationsRead paper
Recent publications

Latest Papers

The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning

Aug 14, 2026

This study addresses the limitation of existing large language model unlearning methods that overlook fact popularity, rendering high-frequency knowledge difficult to remove. We propose AdaPop, a novel approach that models popularity as a learnable parameter by integrating token confidence with external proxy evaluations. Through a bi-ascent controller, AdaPop dynamically adjusts penalty intensity to achieve an adaptive balance between forgetting and retention. Experiments across three model families and two benchmarks demonstrate that AdaPop reduces content leakage under paraphrased queries by approximately fivefold and under adversarial reconstruction by 1.6 times. Furthermore, internal representation analysis confirms superior forgetting separation, effectively resolving the challenge of unlearning high-frequency facts while preserving general model utility.

0 citationsRead paper