expert elicitation

Design, implement, and analyze procedures, instruments, and protocols to extract, structure, and validate expert knowledge, professional judgments, requirements, and preferences; produce calibrated quantitative priors, pairwise and scalar preference datasets, categorical requirements, and curated knowledge artifacts together with their qualitative rationales. This work includes creating targeted elicitation queries and workflows, translating observed behaviors into preference or prior models, sequencing and formatting queries for users, and applying validation and curation methods to ensure reliability of elicited responses.

expertelicitation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.38
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$220K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

elicito: A Python Package for Expert Prior Elicitation

Jun 20, 2025
PB
Paul-Christian Burkner
🏛️ TU Dortmund University

Expert priors are typically elicited from domain experts based on observable quantities rather than model parameters, making direct translation into Bayesian prior distributions challenging. Method: We propose a modular, simulation-based framework for expert prior elicitation that systematically maps expert judgments about observables to computationally tractable parameter priors. Our approach features a configurable architecture that decouples the generative model, expert input format, prior assumptions, and loss function—supporting both structured and predictive elicitation paradigms. It integrates simulation-based inference, parametric and nonparametric prior modeling, and interactive interfaces to enhance transparency, reproducibility, and methodological comparability. Contribution/Results: We release *elicito*, an open-source Python package with a comprehensive API and empirical case studies, demonstrating its efficiency and robustness in bridging the gap between expert knowledge and probabilistic modeling in real-world applications.

Bridges gap between observable quantities and model parametersProvides modular framework for customizable prior elicitationTranslates expert knowledge into usable Bayesian priors

A validity-guided workflow for robust large language model research in psychology

Jul 06, 2025
ZL
Zhicheng Lin
🏛️ Yonsei University | University of Science and Technology of China

Large language models (LLMs) are increasingly deployed in psychological research—as tools, targets of assessment, and cognitive models—yet recent evidence reveals severe measurement unreliability: factor structures of personality traits collapse, moral judgments reverse with minor punctuation changes, and theory-of-mind performance fluctuates dramatically under syntactic rephrasing. These “measurement ghosts” reflect statistical artifacts rather than substantive phenomena, threatening construct validity. Method: We propose the first validity-driven, six-stage workflow integrating psychometric principles and causal inference frameworks, dynamically calibrating validation rigor to research objectives and systematically governing the entire LLM psychology research lifecycle. Our approach includes construct validity verification, computational confound control, modeling of non-independent observations, and transparent experimental design. Contribution/Results: Applied to assessing “LLM selfhood,” our framework successfully disentangles genuine computational phenomena from measurement artifacts, establishing a reproducible empirical paradigm for AI psychology.

Addressing measurement unreliability in LLM-based psychological researchDeveloping a validity-guided workflow for robust AI psychology studiesDistinguishing genuine computational phenomena from measurement artifacts

Beyond Quantification: Navigating Uncertainty in Professional AI Systems

Sep 03, 2025
SD
Sylvie Delacroix
🏛️ Dickson Poon School of Law | King’s College London | Centre for Language AI Research | Tohoku University | Department of Computer Science and Technology | University of Cambridge | Center for Data Science | New York University | Laboratoire de Neurosciences Cognitives | École Normale Supérieure | Helmholtz Center Munich German Research Center for Environmental Health | Helmholtz AI | Department of Informatics | AstraZeneca | Department of Computer Science | University of Bath | ETH Zurich | School of Comput

Conventional probabilistic uncertainty quantification in AI systems (e.g., confidence scores) fails to capture inherently non-quantifiable judgmental uncertainty in high-stakes professional domains—such as domestic violence risk assessment, cultural sensitivity evaluation, or conceptual understanding recognition—where uncertainty is epistemically and contextually grounded rather than statistically estimable. Method: We propose a non-quantitative, meaning-oriented uncertainty representation paradigm that reframes uncertainty articulation as a collaborative meaning-making process within professional communities. Integrating human-computer interaction, practice theory, and participatory design, we develop a co-evolving refinement mechanism wherein domain experts iteratively define and optimize uncertainty expression forms. Contribution/Results: We establish the first theoretical framework for non-quantitative uncertainty tailored to professional practice and empirically validate its efficacy across multiple domains, demonstrating significant improvements in expert–AI collaborative decision quality.

Addressing non-quantifiable uncertainties in professional AI systemsCreating participatory processes for professional uncertainty communicationDeveloping richer uncertainty expressions beyond probabilistic measures

PriorWeaver: Prior Elicitation via Iterative Dataset Construction

Oct 07, 2025
YX
Yuwei Xiao
🏛️ University of California, Los Angeles | Aalto University

In Bayesian analysis, translating domain expertise into computationally tractable prior distributions remains challenging due to expressive limitations and cognitive gaps. This paper introduces an interactive visual prior elicitation method that reframes prior specification as a “hypothetical data construction” process: users iteratively generate synthetic datasets aligned with their beliefs via drag-and-drop operations, constraint imposition, and forward simulation; the system then automatically infers the corresponding prior distribution and provides real-time feedback via prior predictive checks. Integrating visual reasoning, probabilistic modeling, and predictive calibration, the approach enhances the intuitiveness, controllability, and credibility of prior encoding. A user study demonstrates that, compared to conventional parametric prior specification, 92% of participants formulated priors more faithfully reflecting their domain beliefs; moreover, prior clarity and debugging efficiency improved significantly.

Facilitates prior elicitation through interactive dataset constructionHelps analysts create better-aligned priors through iterative refinementTranslates visual assumptions into statistical priors automatically

Contextualized Evaluations: Taking the Guesswork Out of Language Model Evaluations

Nov 11, 2024
CM
Chaitanya Malaviya
🏛️ University of Pennsylvania | Allen Institute for AI | UMass Amherst

To address the ambiguity and subjectivity in evaluating language model responses to fuzzy queries (e.g., subjective or open-ended questions), this paper proposes a contextualized evaluation protocol that embeds structured context—such as synthesized user identities, query intents, and utility criteria—into the assessment process. Methodologically, it integrates context synthesis modeling, multi-dimensional human evaluation design, cross-dimensional quality analysis, and bias-sensitivity quantification. Our study is the first to systematically uncover mainstream models’ implicit preference for WEIRD (Western, Educated, Industrialized, Rich, Democratic) contexts and reveal pronounced asymmetry in their contextual adherence capabilities. Experiments demonstrate that the protocol reverses relative model win rates, mitigates superficial stylistic biases, and yields fine-grained behavioral insights—thereby substantially enhancing evaluation objectivity, interpretability, and diagnostic utility.

Current evaluations make arbitrary judgments on response qualityEvaluating underspecified queries lacks explicit context criteriaModels show bias towards WEIRD contexts in default responses

Latest Papers

What's happening recently
View more

This study addresses the lack of executable and verifiable knowledge representations in existing meta-analyses, which hinders the traceability and reproducibility of critical analytical decisions. To overcome this limitation, the authors propose Executable Analytical Knowledge Representation (EAKR) and introduce MetaSynDec, an agent-based framework that, for the first time, enables explicit modeling, machine-actionable execution, and closed-loop validation of meta-analytic decisions. The system leverages large language models to generate structured knowledge and validates and executes it through deterministic, schema- and contract-based services. Evaluated across 58 synthesis units, EAKR successfully constructed all units, achieved exact evidence-set consistency in 75% of cases, and produced confidence intervals overlapping with published results in 98.2% of cases—substantially outperforming direct LLM-generated approaches.

analytical knowledge representationevidence synthesisexecutable knowledge

Traditional questionnaires struggle to simultaneously capture qualitative depth and quantitative structure, limiting comprehensive understanding of complex social phenomena. This study proposes a dynamic survey platform powered by large language models (LLMs) that, for the first time, enables real-time semantic clustering of open-ended responses. Through an interactive feedback mechanism, users can rate, rank, and reflect on these clusters, generating visual reports that integrate qualitative insights with quantitative analysis. Innovatively embedding LLMs within a closed-loop data collection framework, the approach facilitates dynamic comparisons between individual perspectives and group-level trends. Empirical validation across two field studies involving 93 participants demonstrates that the platform significantly enhances data richness and user engagement compared to conventional survey tools, while effectively fostering collaborative sensemaking.

collaborative interactionLLMsqualitative depth

Existing benchmarks for knowledge work evaluation largely adhere to traditional NLP task paradigms, failing to capture systems’ capabilities in real-world knowledge-intensive settings. This work proposes a three-step framework—explicitly defining work activities, establishing realistic test environments, and focusing evaluation on deliverable outputs—and derives 18 core knowledge work activities from the O*NET database. Innovatively integrating role responsibilities, local tool usage, and downstream usability into benchmark design, the approach establishes a coherent “work activity–test setup–scoring artifact” alignment. Validation through three case studies (GDPval, OfficeQA Pro, and APEX-SWE) exposes critical misalignments in current benchmarks between tasks, environments, and actual work objectives, offering a new paradigm for evaluating knowledge work systems in practical, application-oriented contexts.

benchmark designevaluationknowledge work

Hot Scholars

DY

Diyi Yang

Stanford University
Computational Social ScienceNatural Language ProcessingMachine Learning
YC

Yejin Choi

Stanford University / NVIDIA
Natural Language ProcessingDeep LearningArtificial IntelligenceCommonsense Reasoning
IG

Iryna Gurevych

Full Professor, TU Darmstadt; Adjunct Professor, MBZUAI, UAE; Affiliated Professor, INSAIT, Bulgaria
Natural Language ProcessingLarge Language ModelsArtificial Intelligence
SL

Seth Lazar

Australian National University
Ethicspolitical philosophyethics of riskethics of war
HH

Hoda Heidari

Carnegie Mellon University
Responsible AIAI EthicsAI AccountabilityAlgorithmic Fairness