Score
Designs, builds, or evaluates methods and tools that perform precise, constrained modifications to textual content by identifying minimal edit spans or tokens to change and applying token-level edits. These methods ensure edits preserve the document’s reasoning or discourse coherence, incorporate external knowledge or control signals to constrain scope and semantics, and verify that changes are minimal and consistent with the original intent.
This work addresses the challenge of jointly controlling text editing across multiple scales—from coarse-grained stylistic attributes to fine-grained lexical choices—and diverse semantic attributes (e.g., toxicity, sentiment). We propose the first diffusion-based two-stage controllable editing framework: Stage I employs SDEdit for global, style-level coarse editing; Stage II introduces a self-conditioning mechanism enabling localized, fine-grained semantic refinement. Crucially, our framework is the first to decouple and jointly regulate editing strength and semantic granularity within diffusion language models. Extensive experiments on multiple benchmarks demonstrate substantial improvements in both edit fidelity and attribute control: toxicity reduction increases by 23.6%, and sentiment reversal accuracy improves by 18.4%. The method establishes a novel paradigm for controllable text generation, advancing the state of the art in diffusion-based language modeling.
为解决AI草稿中局部修改导致无关内容变化的问题,PatchWrite通过编译检查和证据锁定方法确保修改的有效性和一致性。
Large language models (LLMs) exhibit significant limitations in precise, context-aware, fine-grained text editing—particularly in preserving deep structural integrity and logical consistency. To address this, we introduce InstrEditBench, the first structured editing benchmark comprising over 20,000 samples spanning Wikipedia, LaTeX, source code, and domain-specific languages (DSLs). We propose an instruction-following–oriented automated framework for editing task generation and evaluation, and design FineEdit: a lightweight fine-tuning paradigm integrating high-quality structured-data supervised fine-tuning, multi-domain consistency constraints, and instruction-driven editing modeling. Experiments demonstrate that FineEdit achieves approximately 10% absolute improvement over Gemini on direct editing tasks, while guaranteeing zero perturbation to non-target content and enabling high-fidelity, verifiable semantic-level modifications.
Large language models (LLMs) often struggle to precisely execute user-specified editing intentions in instruction-driven text editing tasks, frequently over-editing unmodified regions and thereby compromising faithfulness and locality. To address this, we propose a lightweight, fine-grained editing framework: (1) a hypernetwork-based dynamic adaptation mechanism that generates instruction-customized editing policies; and (2) span-level difference-aware regularization, which imposes precise supervision on modified spans to effectively suppress over-editing. The resulting model contains only 3 billion parameters, balancing efficiency and capability. On modified-span BLEU, it outperforms state-of-the-art methods by 9–30%, while maintaining high edit accuracy and minimal contextual interference. Our approach establishes a new paradigm for controllable, localized editing of code and documentation.
This work addresses the high token consumption, latency, energy usage, and low accuracy often incurred by generative AI in code editing due to unstructured feedback. To overcome these limitations, the authors propose FileMark—a structured, line-anchored feedback mechanism that replaces conventional holistic prompting. Through a VS Code extension, multi-model controlled experiments, and function-level automated patch application, the study systematically demonstrates for the first time that line-anchored feedback substantially reduces token usage (by 22%–58%, and up to 24%–80% on large files) while significantly improving correction accuracy, especially for weaker models (gains of +5–7 percentage points, with nearly threefold improvement on large files). The findings further reveal that shifting the editing burden to the feedback mechanism can amplify these benefits.
This study addresses the disconnect between localization prediction and content generation capabilities in diffusion language models for code editing. Building upon masked diffusion language models, this work systematically evaluates four interaction paradigms—including whole-file rewriting and search-and-replace—using the CanItEdit benchmark alongside a Wiki probing framework. Furthermore, it introduces a "composition gap" metric to quantify the degree of separation between localization and generation. The findings reveal that strong infilling proficiency alone is insufficient to guarantee editing reliability. Instead, dependable code editing necessitates simultaneously achieving complete change coverage and minimizing extraneous regeneration. These insights provide a theoretical foundation for designing code editing interfaces that effectively balance precision with safety.
研究提出一种结合直接操作和自然语言编程的框架,通过编辑语言作为接口,支持多种编程交互方式,并探讨了这种结合如何影响编程过程。
This study addresses the lack of coherence and consistency in document-level text simplification for low-resource languages such as Estonian. We evaluate five multilingual large language models using three strategies: single-turn generation, pipeline processing, and guideline-enhanced prompting. Comprehensive evaluation combines automated metrics with human annotation. The primary contributions include a novel metric for document-level coherence, validation of evidence-based prompting strategies, and the release of open-source reproducible resources. Experimental results demonstrate that Gemini-2.0 and LLaMA-3.3 produce simplified texts exhibiting near-native fluency alongside strong semantic preservation.
本文提出一种闭环控制架构,通过生成、评估、调整、存档和分析五个阶段,解决大语言模型在文本生成中满足数值输出约束不可靠的问题。
本文通过调整不同停顿阈值划分微单元和宏单元,评估文本生成过程中的编辑操作分布,认为基于过程的单元更适合。