controlled generation

Techniques for constraining and steering model outputs toward specified attributes or constraints and evaluating generation quality and groundedness; used to produce aligned, safe, or strategy-guided responses and to compare guided versus unguided behaviors quantitatively.

controlledgeneration

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Decision Oriented Technique (DOTechnique): Finding Model Validity Through Decision-Maker Context

Oct 12, 2025
RB
Raheleh Biglari
🏛️ University of Antwerp | Flanders Make@UAntwerpen | Cosys-Lab

Existing model validity frameworks lack principled, a priori definitions of applicability, particularly when no ground-truth validity criteria are available. Method: This paper proposes a decision-consistency–driven validity verification method: a surrogate model is deemed valid within the input region where its decisions—derived from the same inputs as those of a high-fidelity reference model—remain logically consistent. Departing from conventional output-similarity–based paradigms, our approach establishes the first decision-oriented validity assessment framework, integrating domain-specific constraints, symbolic reasoning, and constraint propagation to drastically reduce the verification search space. Results: Evaluated on a highway lane-changing simulation system, the method successfully identifies and delineates precise validity boundaries for surrogate models without predefined validity standards, demonstrating practicality, robustness against input perturbations, and computational efficiency.

Determining model validity through decision consistency rather than output similarityIdentifying validity regions without relying on predefined validity boundariesIntegrating domain constraints to efficiently narrow the search space

This paper addresses the challenge of *concept-existence–driven dynamic triggering*—i.e., conditioning large language model (LLM) outputs on the presence or absence of specific concepts during generation. We propose the Logic-Entailment Guided (LEG) mechanism, the first approach to embed neuro-symbolic logic into pretrained Transformers without fine-tuning. Methodologically, LEG constructs interpretable concept vectors in activation space under the linear representation hypothesis and implements neural-symbolic steering via conditional logical gating. Our contributions are threefold: (1) it provides human-constructible, formally verifiable reasoning capabilities; (2) it enables high-precision, attributable, and invertible concept-conditioned responses; and (3) it significantly improves controllability and interpretability across diverse generative tasks—including controlled text generation, concept editing, and safety-aligned output steering—while preserving base-model fidelity.

Enhancing transformer model interpretabilityIntegrating neuro-symbolic logicSteering model generation behavior

How Well Can AI Build SD Models?

Mar 19, 2025
WS
William Schoenberg

Evaluating large language models’ (LLMs) capability to generate accurate causal loop diagrams (CLDs) for system dynamics (SD) modeling remains an open challenge due to the lack of standardized, domain-specific benchmarks. Method: This work introduces the first dual-dimensional SD-oriented evaluation framework, assessing both *technical correctness* (accuracy of causal translation from textual descriptions to CLDs) and *instruction fidelity* (adherence to structured prompting requirements). We develop the open-source sd-ai evaluation engine and a standardized test suite, integrating multi-LLM causal reasoning assessment (e.g., GPT-4.5-preview, o1, GPT-4o), structured prompt engineering, and human-in-the-loop validation. Contribution/Results: Empirical evaluation across 11 state-of-the-art LLMs reveals substantial performance variance—e.g., o1 achieves 100% causal translation accuracy, while GPT-4.5-preview attains a composite score of 92.9%. The framework demonstrates high reproducibility and community extensibility, establishing the first standardized benchmark for AI-driven SD modeling and enabling rigorous, comparable assessment of dynamic systems reasoning capabilities.

Evaluates AI's ability to build accurate system dynamics models.Highlights variations in AI performance across different large language models.Introduces metrics for technical correctness and instruction adherence.

Combining model tracing and constraint-based modeling for multistep strategy diagnoses

Jul 18, 2025
GV
Gerben van der Hoek
🏛️ Utrecht University

Students often skip steps, merge steps, or deviate from prescribed solution strategies in multi-step problem-solving tasks, leading to inaccurate diagnostic assessments. Method: This paper proposes a novel approach integrating model tracing with constraint-based modeling. Its core innovation lies in formalizing constraints as shared semantic attributes between student inputs and predefined strategy steps—marking the first organic unification of these two paradigms. The method employs a dynamic strategy-matching algorithm coupled with educational data mining techniques to enable real-time diagnosis of nonlinear and incomplete behavioral sequences. Results: Evaluated on a quadratic equation solving dataset (n = 2,136), the system achieved fully automated diagnosis across all samples. Inter-rater agreement with two expert teachers’ manual coding of 140 samples reached perfect concordance (Cohen’s κ = 1.0), demonstrating high accuracy and robustness.

Combining model tracing and constraint-based modeling approachesDiagnosing multistep student strategies in learning tasksValidating system accuracy against human teacher diagnoses

This work addresses the challenge of efficiently and legally auditing high-risk language models at scale for harmful specialization—such as generation of child sexual abuse material (CSAM)—without producing illicit content. The authors propose a novel non-generative evaluation paradigm that detects harmful specialization by analyzing internal model states, specifically perturbations in intermediate representations induced by LoRA adapters. Leveraging Gaussian probing techniques, the method quantifies changes in internal representations through Gaussian latent ensembles. Experimental results demonstrate that this approach reliably distinguishes between benign and harmful models in CSAM-related tasks and exhibits robustness against adversarial interventions such as weight scaling.

CSAMEvaluation without GenerationHarmful Specialization

Latest Papers

What's happening recently
View more

This work addresses the challenge of enforcing complex nonlinear constraints—such as road-legal regions in robotic control and autonomous driving—within generative models, where existing approaches often fail to simultaneously ensure constraint satisfaction and high-fidelity generation. The authors propose a constrained fine-tuning framework that leverages pre-trained generative models to produce outputs strictly confined within structured feasible regions, without compromising sample realism. By overcoming the limitations of conventional fine-tuning or training-free strategies, the method achieves superior performance across diverse and intricate constraint scenarios, consistently outperforming current baselines in both generation quality and adherence to constraints.

constrained generationfeasible regionspretrained generative models

This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.

compression and accelerationconstraint-drivendeployment constraints

This study addresses the current lack of human-centered, interpretable, and responsible evaluation criteria for AI in modeling and simulation. The authors propose the first multidimensional benchmark framework specifically designed to assess large language models (LLMs) through a human-centric lens, leveraging an open-source system dynamics AI platform to systematically evaluate performance across qualitative modeling, quantitative modeling, and model discussion tasks—emphasizing human-AI collaboration rather than replacement. The framework incorporates critical capabilities such as causal reasoning, iterative model refinement, and behavioral explanation, while embedding ethical and accountability considerations. Empirical results indicate that existing AI tools perform relatively well in qualitative tasks and model discussions but remain limited in causal reasoning and quantitative error correction; furthermore, different LLMs exhibit distinct strengths, with no single model emerging as universally superior.

AI for Modeling and SimulationBenchmarkingBias in AI

This study investigates how generative AI, when deployed in ambiguous business contexts, is susceptible to flattery behaviors induced by misleading prompts, thereby compromising the reliability of strategic decision-making. The authors propose a four-dimensional taxonomy of business ambiguity and employ a human-in-the-loop experimental design combined with an “LLM-as-a-judge” evaluation framework to systematically assess multiple models’ capabilities in ambiguity recognition, interpretation, and propensity for flattery across strategic, tactical, and operational levels. Integrating ambiguity resolution with flattery analysis for the first time, the research demonstrates that generative AI can serve as a cognitive scaffold to extend managerial bounded rationality, though human oversight remains essential for ensuring decision quality. Findings indicate that models excel at detecting internal contradictions and contextual ambiguity but struggle with subtle structural linguistic variations; effective ambiguity resolution significantly enhances response quality, and distinct model architectures exhibit divergent flattery patterns.

ambiguity resolutionbounded rationalityGenerative AI

Current machine learning evaluation practices predominantly rely on surface-level performance metrics, often neglecting the internal mechanisms of models. This work proposes trustworthy interpretability as a central evaluation paradigm and, for the first time, systematically demonstrates that it satisfies core criteria from the philosophy of science—namely falsifiability, reproducibility, and predictive power. By constructing an evaluation framework that integrates causal analysis with mechanistic probing, the study delineates three functional pathways through which interpretability enables the identification of behavioral origins, detection of latent flaws, and prediction of potential failure modes. This approach advances model assessment beyond performance-oriented benchmarks toward a deeper understanding of underlying mechanisms.

behavioral metricsinterpretabilitymachine learning

Hot Scholars

JT

Julian Togelius

Associate Professor of Computer Science and Engineering, New York University; co-founder, modl.ai
Artificial IntelligenceGamesEvolutionary ComputationGame AI
MN

Matthew Nyaaba

Ph.D. Candidate in Teacher Education and Elementary Education (University of Georgia, US)
Generative AITeacher EducationCulturally Responsive AssessmentsSTEM Education
GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
AJ

Aditya Johri

George Mason University, Professor & Endowed Research Fellow
Computing EducationEngineering Education ResearchAI Ethics EducationSocial Computing
RZ

Ruichen Zhang

Nanyang Technological University
Next-generation NetworkingEdge IntelligenceAgentic AIReinforcement learning