Score
Designs, builds, and evaluates algorithms and systems that transform textual content to change stylistic attributes—such as tone, formality, politeness, dialect, or bias—while preserving semantic content and fluency. Work includes methods for controlling and generating target style, adapting or preserving communication tone, applying style-transfer procedures (e.g., neutralization of biased text), and creating evaluation techniques to measure rewritten-text quality.
This work addresses the computational modeling and controllable manipulation of textual style, systematically tackling three core tasks: Text Style Transfer (TST), Author Attribution (AA), and Author Verification (AV). We propose a parameter-efficient fine-tuning framework built upon large language models, integrating contrastive learning with instruction tuning to achieve disentangled style representations and explicit content–style separation. Crucially, we unify TST and AV under a single, interpretable contrastive disentanglement paradigm—enhancing transfer fidelity and attribution reliability. Empirically, our method achieves state-of-the-art performance across multiple standard benchmarks: it preserves content fidelity in TST while attaining SOTA accuracy on both AA and AV tasks. Moreover, the approach demonstrates strong generalization across domains and styles, alongside inherent interpretability through disentangled latent representations.
Existing content preservation evaluation methods for text style/attribute transfer—relying on lexical or semantic similarity metrics or current LLM-based evaluators—fail to model stylistic conditionality, resulting in low correlation with human judgments. Method: The authors propose, for the first time, that content preservation assessment must be *conditioned on style transfer*, and introduce a zero-shot evaluation method based on next-token conditional likelihood. They further construct a human-annotated benchmark specifically designed for meta-evaluation alignment to systematically validate the necessity of conditional modeling across multiple style transfer tasks. Contribution/Results: Experiments demonstrate that the proposed method significantly outperforms baseline approaches, achieving an average 23% improvement in correlation with human judgments. This work establishes conditional modeling as essential for accurate, human-aligned content preservation evaluation in style transfer.
Text style transfer (TST) lacks reliable, automated evaluation metrics—especially in multilingual and cross-task settings. To address this, we conduct the first systematic meta-evaluation of TST, covering sentiment transfer and detoxification tasks across English, Hindi, and Bengali. We propose a hybrid metric framework integrating BERTScore, BLEURT, MAUVE, and LLM-based assessment, augmented with an ensemble strategy. Experimental results demonstrate that general-purpose NLP metrics consistently outperform traditional TST-specific metrics; our hybrid approach improves average Spearman correlation with human judgments by 23%, significantly enhancing consistency, accuracy, and reproducibility. This work establishes the first empirically validated, standardized evaluation paradigm for multilingual TST.
Text-driven style transfer faces challenges including reference-style overfitting, insufficient fine-grained control, and text–style misalignment. This paper proposes a fine-grained controllable approach enabling explicit modulation of stylistic elements—such as color, texture, and brushstrokes—while preserving high semantic fidelity between generated content and textual prompts. Our method, built upon a diffusion architecture, integrates adaptive cross-modal instance normalization (for joint style–text modeling), style-conditioned classifier-free guidance (SCFG), and a teacher-model-assisted layout stabilization mechanism. It synergistically combines AdaIN, classifier-free guidance, and knowledge distillation, requiring no fine-tuning and supporting plug-and-play deployment. Quantitative evaluation shows a 23.6% reduction in FID, alongside significant improvements in text alignment and style fidelity. User studies confirm a 41% increase in perceived stylistic controllability.
Existing text-to-image diffusion models often compromise semantic content when performing style editing, struggling to achieve fine-grained, continuous, and disentangled control. This work proposes a method that learns disentangled editing directions from synthetic data, integrating a guidance composition mechanism with a regularized training loss and optimizing the null-text embedding to enhance DDIM inversion. This approach enables parameterized, continuous adjustment of stylistic attributes while preserving semantic consistency. Evaluated on styles such as contour emphasis, local contrast, watercolor effects, and geometric patterns, the method significantly outperforms current text-driven editing techniques, achieving more precise, coherent, and controllable style transfer.
This study addresses the prevalence of aggressive and emotionally harmful content on social media platforms, where current moderation practices predominantly rely on content removal and often overlook opportunities for communicative repair. To bridge this gap, the authors propose a controllable text style transfer–based writing assistance framework that reformulates toxic comments into neutral and温和 expressions while preserving their original semantic meaning. A novel Emotion Drift Index (EDI) is introduced to quantitatively measure emotional shifts before and after rewriting, enabling effective mitigation of harmful interactions without resorting to simple deletion. Experimental results demonstrate that the proposed approach significantly reduces emotional antagonism and social harm in online discourse while maintaining informational integrity.
This study systematically investigates stylistic differences between human- and large language model (LLM)-generated texts across genres, models, and decoding strategies to inform responsible LLM deployment. Leveraging Biber’s multidimensional framework of register variation, the authors conduct a large-scale comparative analysis of texts produced by 11 LLMs across eight genres and four decoding strategies. The findings reveal that model type and genre exert substantially stronger influences on textual style than prompting or decoding choices. Notably, chat-oriented models exhibit pronounced clustering in stylistic space, and key linguistic features of LLM-generated text demonstrate robustness across generation conditions, with genre effects consistently outweighing those of text origin.
This study investigates the impact of large language models (LLMs) on textual style and authorial voice when rewriting personal narratives. Three prompting conditions—generic optimization, rewrite-only, and voice-preserving—were applied to guide three state-of-the-art LLMs in rewriting 300 personal narratives. Quantitative analysis based on 13 computational stylometric features (e.g., function words, lexical diversity, first-person pronouns, and affective terms) reveals that, regardless of prompting strategy, LLMs consistently induce stylistic convergence, manifesting as homogenization and decontextualization. Notably, even under explicit instructions to preserve voice, source-text traceability significantly diminishes, with narrative stance shifting from embedded to distanced and causal expressions becoming more compressed and abstract.
This study addresses the risk of personal writing style erosion when users rely on large language models (LLMs) in style-sensitive writing contexts. Through a preregistered online experiment, it systematically evaluates the effectiveness of post-editing in recovering individual stylistic identity by integrating users’ subjective perceptions with an embedding-based objective measure of stylistic similarity. Comparing texts produced through unaided human writing, LLM generation, and post-edited LLM output, the findings reveal that while post-editing increases stylistic alignment with users’ native writing, the resulting text remains closer to the LLM’s inherent style and exhibits significantly lower stylistic diversity than human-authored text. These results highlight the current limitations of post-editing strategies in preserving authentic personal style and provide empirical grounding for the development of more personalized AI writing assistants.
This work addresses the challenge of scarce parallel corpora in text style transfer by proposing a novel approach that operates without authentic parallel data. The method leverages back-translation to generate neutral-style texts as shared input representations and integrates parameter-efficient fine-tuning (PEFT) with retrieval-augmented generation (RAG) to enhance terminological consistency and stylistic control. Evaluated across four domains, the proposed framework significantly outperforms zero-shot prompting and few-shot in-context learning (ICL), achieving state-of-the-art performance in both BLEU scores and style accuracy. This advancement circumvents the traditional reliance on manually annotated parallel corpora, offering a scalable and effective solution for style transfer tasks under low-resource conditions.