Score
Design, implement, and analyze procedures, instruments, and protocols to extract, structure, and validate expert knowledge, professional judgments, requirements, and preferences; produce calibrated quantitative priors, pairwise and scalar preference datasets, categorical requirements, and curated knowledge artifacts together with their qualitative rationales. This work includes creating targeted elicitation queries and workflows, translating observed behaviors into preference or prior models, sequencing and formatting queries for users, and applying validation and curation methods to ensure reliability of elicited responses.
Expert priors are typically elicited from domain experts based on observable quantities rather than model parameters, making direct translation into Bayesian prior distributions challenging. Method: We propose a modular, simulation-based framework for expert prior elicitation that systematically maps expert judgments about observables to computationally tractable parameter priors. Our approach features a configurable architecture that decouples the generative model, expert input format, prior assumptions, and loss function—supporting both structured and predictive elicitation paradigms. It integrates simulation-based inference, parametric and nonparametric prior modeling, and interactive interfaces to enhance transparency, reproducibility, and methodological comparability. Contribution/Results: We release *elicito*, an open-source Python package with a comprehensive API and empirical case studies, demonstrating its efficiency and robustness in bridging the gap between expert knowledge and probabilistic modeling in real-world applications.
Large language models (LLMs) are increasingly deployed in psychological research—as tools, targets of assessment, and cognitive models—yet recent evidence reveals severe measurement unreliability: factor structures of personality traits collapse, moral judgments reverse with minor punctuation changes, and theory-of-mind performance fluctuates dramatically under syntactic rephrasing. These “measurement ghosts” reflect statistical artifacts rather than substantive phenomena, threatening construct validity. Method: We propose the first validity-driven, six-stage workflow integrating psychometric principles and causal inference frameworks, dynamically calibrating validation rigor to research objectives and systematically governing the entire LLM psychology research lifecycle. Our approach includes construct validity verification, computational confound control, modeling of non-independent observations, and transparent experimental design. Contribution/Results: Applied to assessing “LLM selfhood,” our framework successfully disentangles genuine computational phenomena from measurement artifacts, establishing a reproducible empirical paradigm for AI psychology.
Conventional probabilistic uncertainty quantification in AI systems (e.g., confidence scores) fails to capture inherently non-quantifiable judgmental uncertainty in high-stakes professional domains—such as domestic violence risk assessment, cultural sensitivity evaluation, or conceptual understanding recognition—where uncertainty is epistemically and contextually grounded rather than statistically estimable. Method: We propose a non-quantitative, meaning-oriented uncertainty representation paradigm that reframes uncertainty articulation as a collaborative meaning-making process within professional communities. Integrating human-computer interaction, practice theory, and participatory design, we develop a co-evolving refinement mechanism wherein domain experts iteratively define and optimize uncertainty expression forms. Contribution/Results: We establish the first theoretical framework for non-quantitative uncertainty tailored to professional practice and empirically validate its efficacy across multiple domains, demonstrating significant improvements in expert–AI collaborative decision quality.
In Bayesian analysis, translating domain expertise into computationally tractable prior distributions remains challenging due to expressive limitations and cognitive gaps. This paper introduces an interactive visual prior elicitation method that reframes prior specification as a “hypothetical data construction” process: users iteratively generate synthetic datasets aligned with their beliefs via drag-and-drop operations, constraint imposition, and forward simulation; the system then automatically infers the corresponding prior distribution and provides real-time feedback via prior predictive checks. Integrating visual reasoning, probabilistic modeling, and predictive calibration, the approach enhances the intuitiveness, controllability, and credibility of prior encoding. A user study demonstrates that, compared to conventional parametric prior specification, 92% of participants formulated priors more faithfully reflecting their domain beliefs; moreover, prior clarity and debugging efficiency improved significantly.
To address the ambiguity and subjectivity in evaluating language model responses to fuzzy queries (e.g., subjective or open-ended questions), this paper proposes a contextualized evaluation protocol that embeds structured context—such as synthesized user identities, query intents, and utility criteria—into the assessment process. Methodologically, it integrates context synthesis modeling, multi-dimensional human evaluation design, cross-dimensional quality analysis, and bias-sensitivity quantification. Our study is the first to systematically uncover mainstream models’ implicit preference for WEIRD (Western, Educated, Industrialized, Rich, Democratic) contexts and reveal pronounced asymmetry in their contextual adherence capabilities. Experiments demonstrate that the protocol reverses relative model win rates, mitigates superficial stylistic biases, and yields fine-grained behavioral insights—thereby substantially enhancing evaluation objectivity, interpretability, and diagnostic utility.
This study addresses the lack of executable and verifiable knowledge representations in existing meta-analyses, which hinders the traceability and reproducibility of critical analytical decisions. To overcome this limitation, the authors propose Executable Analytical Knowledge Representation (EAKR) and introduce MetaSynDec, an agent-based framework that, for the first time, enables explicit modeling, machine-actionable execution, and closed-loop validation of meta-analytic decisions. The system leverages large language models to generate structured knowledge and validates and executes it through deterministic, schema- and contract-based services. Evaluated across 58 synthesis units, EAKR successfully constructed all units, achieved exact evidence-set consistency in 75% of cases, and produced confidence intervals overlapping with published results in 98.2% of cases—substantially outperforming direct LLM-generated approaches.
Traditional questionnaires struggle to simultaneously capture qualitative depth and quantitative structure, limiting comprehensive understanding of complex social phenomena. This study proposes a dynamic survey platform powered by large language models (LLMs) that, for the first time, enables real-time semantic clustering of open-ended responses. Through an interactive feedback mechanism, users can rate, rank, and reflect on these clusters, generating visual reports that integrate qualitative insights with quantitative analysis. Innovatively embedding LLMs within a closed-loop data collection framework, the approach facilitates dynamic comparisons between individual perspectives and group-level trends. Empirical validation across two field studies involving 93 participants demonstrates that the platform significantly enhances data richness and user engagement compared to conventional survey tools, while effectively fostering collaborative sensemaking.
Existing benchmarks for knowledge work evaluation largely adhere to traditional NLP task paradigms, failing to capture systems’ capabilities in real-world knowledge-intensive settings. This work proposes a three-step framework—explicitly defining work activities, establishing realistic test environments, and focusing evaluation on deliverable outputs—and derives 18 core knowledge work activities from the O*NET database. Innovatively integrating role responsibilities, local tool usage, and downstream usability into benchmark design, the approach establishes a coherent “work activity–test setup–scoring artifact” alignment. Validation through three case studies (GDPval, OfficeQA Pro, and APEX-SWE) exposes critical misalignments in current benchmarks between tasks, environments, and actual work objectives, offering a new paradigm for evaluating knowledge work systems in practical, application-oriented contexts.