Score
Designs and builds systems that use large language models to generate sentences conditioned on gloss sequences or other gloss-like controls, optionally anchoring outputs to a source corpus and producing novel gloss–sentence pairs. Implements mechanisms to control lexical and syntactic choices, corpus fidelity, and the diversity of generated outputs.
To address the challenge of precisely satisfying syntactic and semantic constraints during generation with large language models (LLMs), this paper proposes a controllable text generation framework based on Sequential Monte Carlo (SMC). The method formalizes constraints as conditional probability distributions and integrates domain knowledge at inference time—without fine-tuning—via constraint-aware particle resampling and dynamic computational resource allocation. This work is the first to systematically introduce SMC into LLM-based controllable generation, offering theoretically grounded improvements in posterior approximation quality. Evaluated on four tasks—Python code generation, text-to-SQL parsing, goal inference, and molecular synthesis—the approach enables small open-source models to outperform both closed-source fine-tuned models and LLMs eight times larger in parameter count, while incurring negligible additional inference overhead.
This study addresses the challenge of balancing novelty, feasibility, and effectiveness—three critical quality dimensions—in scientific idea generation by large language models (LLMs). We propose a two-stage collaborative optimization framework: first, a base generative model is constructed via supervised fine-tuning on paper–idea pairs; second, a fine-grained multi-dimensional reward model—integrated with a dynamic dimension controller and sentence-level decoding coordination mechanism—guides controllable reinforcement learning. To our knowledge, this is the first scientific idea generation method supporting dynamic, real-time adjustment of dimensional weights. Experimental results demonstrate superior trade-off optimization across all three dimensions, yielding significantly improved outputs that better align with domain experts’ evaluation criteria.
This work addresses the challenge of fine-grained, continuous control over textual attributes—such as length, complexity, sentiment, and tone—in large language models. We propose a novel continuous control signal mechanism based on interpolatable embeddings: each attribute dimension is modeled as a linear interpolation vector between “low” and “high” extremal token embeddings in the word embedding space, enabling conditional generation via lightweight fine-tuning. To our knowledge, this is the first approach to achieve spectrum-based, differentiable, and interpolatable textual attribute control. Experiments on response length control demonstrate that our method significantly improves stability and precision over both in-context learning and discrete-label fine-tuning, reducing control error by 37% while exhibiting strong generalization across unseen attribute values. The code and dataset are publicly released.
To address the need for controllable text generation by large language models (LLMs) in safety-critical applications, this paper tackles the challenge of efficiently ensuring semantic safety of generated outputs. Methodologically, it formalizes semantic constraints as linear structures in the model’s latent space and models the generation process as a trajectory evolution; it then introduces the first gradient-free, closed-form geometric intervention strategy for latent-space control. Theoretically, it establishes the first probabilistic guarantee that generated texts provably reside within a pre-specified safe semantic region. Empirical evaluation on toxicity mitigation demonstrates substantial reduction in harmful content generation while preserving linguistic fluency and lexical diversity—achieving a balanced optimization between controllability and generation quality.
This study investigates whether instruction-tuned large language models (LLMs) adhere to human genre conventions in written style. Methodologically, it employs a linguistics-informed quantitative stylistic analysis—integrating corpus-based statistics with controlled comparative experiments—to systematically characterize LLM output against human-authored texts. Results reveal a previously undocumented, intrinsic “noun-dense” bias in LLMs: their generated text consistently deviates from human norms across grammatical complexity, informational density, and contextual appropriateness—even under informal prompting—and fails to replicate human rhetorical pacing and lexical selection patterns. This finding introduces an interpretable, measurable linguistic dimension for LLM text evaluation, moving beyond opaque, black-box assessment paradigms. It provides both theoretical grounding and empirical evidence for advancing controllable stylistic generation and human–AI collaborative writing systems.
Current LLM instruction-following evaluation faces three key challenges: (1) human evaluation is subjective and costly; (2) LLM-as-a-judge introduces systematic biases; and (3) programmatic benchmarks lack expressive power for fine-grained, compositional lexical constraints. To address these, we propose the first formal rule-based framework for fine-grained lexical instruction evaluation. Our method parses complex instructions into verifiable subject-predicate-object triples, constructs a human-in-the-loop, multi-stage data generation pipeline, and integrates both a programmable verification engine and LLM-as-a-judge comparative analysis. We publicly release a high-quality dataset and evaluation toolkit. This work enables the first objective, interpretable, and reproducible automated assessment of compositional lexical instructions—significantly improving evaluation transparency, granularity, and fidelity.
Current large language models (LLMs) ensure syntactic and constraint validity in structured generation but suffer from severely limited output diversity. To address this, we propose an automaton-guided generation mechanism that leverages historical state-transition trajectories—extracted during structured decoding—to dynamically steer the model toward under-explored structural patterns. By tightly integrating automata theory with LLM decoding, our method enhances structural and semantic diversity without compromising validity or inference efficiency. Experimental evaluation on open-source library test-case generation demonstrates a 27.4% improvement in diversity metrics—including structural coverage and semantic dissimilarity—while maintaining a 98.6% compliance rate with syntax and domain constraints. This confirms the method’s effectiveness and practical applicability for diverse, valid structured generation.
This work addresses the challenge of deploying large language models in production settings, where they often fail to meet low-latency requirements, while smaller models typically suffer from limited reasoning capabilities, hallucinations, and insufficient long-context memory. To overcome these limitations, the authors propose supervised fine-tuning small models such as Mistral on domain-specific natural language–code paired data, thereby internalizing domain knowledge directly into model weights and substantially reducing reliance on runtime context. Experimental results demonstrate that the fine-tuned small models outperform larger counterparts in code generation quality while maintaining lower latency. Load testing and real-world deployment confirm their efficiency and stability. Furthermore, the approach supports additional customer-specific fine-tuning without compromising general-purpose capabilities, offering a practical pathway toward efficient and accurate domain-specific code generation.
This study investigates the evolutionary dynamics of Chinese text generation by large language models under memoryless iterative conditions. We formalize this process as a sentence-level Markovian generation chain, where each step depends solely on a fixed prompt template and the output from the previous step. Through restatement and round-trip translation experiments, we analyze the emergent patterns of textual evolution. As the first work to model iterative generation as a Markov chain, we uncover the mechanisms by which temperature settings and initial inputs govern the trajectory of textual diversity: the iterative process either converges to a small cyclic set of sentences or continues producing novel outputs within a finite number of steps, with diversity exhibiting bidirectional trends—increasing or decreasing—depending on parameter configurations.