LLM unbranding: Erasing Commercial Identity while Preserving Generic Utility

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the risks of trademark dilution and defamation arising from brand descriptions in large language model (LLM) outputs, alongside the challenge of identifying textual trade dress. We formally define the task of LLM de-branding for the first time and construct a multi-domain evaluation benchmark. To this end, we propose MUTE, an inference-time method that requires no parameter updates and eliminates brand leakage by iteratively refining system instructions through an optimization loop. Experiments reveal the limitations of existing state-of-the-art models and demonstrate that MUTE effectively mitigates brand-related risks while preserving general model capabilities, thereby achieving a favorable balance between safety and utility. The code and datasets have been made publicly available.
📝 Abstract
Establishing unbranding as a critical practice to prevent visual logos from acquiring negative connotations is standard in image generation. Large Language Models (LLMs) now face a parallel and emerging challenge. These models frequently generate brand descriptions within diverse contexts. This frequency introduces significant risks, such as trademark dilution, false attribution, and brand defamation. In response, we formally define the novel task of LLM Unbranding. We specifically address the complex challenge of managing trade dress within textual outputs. This involves neutralizing characteristic language, slogans, and stylistic markers that define brand identity. Crucially, these elements are less evident than explicit visual logos. To benchmark this task, we introduce a comprehensive evaluation dataset incorporating prominent brands from multiple commercial domains. We rigorously evaluate existing state-of-the-art machine unlearning models using this benchmark. This evaluation identifies their limitations in selective textual unbranding. Finally, we propose MUTE, a novel inference-time method that effectively neutralizes textual trade dress while preserving the LLM's general capabilities and utility. By leveraging an iterative refinement loop, MUTE systematically optimizes system instructions to safely eliminate brand leakage without requiring fragile parameter updates. Code and dataset: The evaluation dataset and code for LLM Unbranding are available at https://github.com/KajetanOzog/LLM_unbranding. The implementation of MUTE is available at https://github.com/KajetanOzog/MUTE.
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Unbranding
Textual Trade Dress
Inference-time Method
Iterative Refinement
Machine Unlearning
🔎 Similar Papers
No similar papers found.