conceptual framing

Defining and operationalizing high-level constructs into measurable, testable components and evaluation targets (e.g., culture, literacy, hallucination sources) so they can be measured, annotated, and used in system design and evaluation.

conceptualframing

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the frequent neglect of local cultural perspectives in existing automated evaluations of AI-generated images, particularly regarding “cultural appropriateness.” It introduces a novel evaluation framework that deeply integrates diverse community participation from the outset, collaborating with blind and visually impaired individuals in the UK and residents of Kerala and Tamil Nadu in India to systematically translate lived cultural experiences and community concerns into actionable assessment dimensions. Leveraging multimodal large language models as judges (LLM-as-a-judge), the approach operationalizes community consensus into structured scoring rules, enabling automated evaluation of cultural appropriateness. The work not only establishes a conceptual framework grounded in community values and demonstrates its feasibility but also exposes critical limitations in current AI models’ understanding of cultural context.

AI-generated imagescommunity-informed evaluationcultural appropriateness

HCI scale development has long suffered from nonstandardized processes, poor construct-theory alignment, and low item reuse rates. This paper introduces the first interactive support system integrating large language models (LLMs) with a structured, empirically grounded measurement knowledge base, enabling a closed-loop workflow: construct identification → theory-informed custom definition → context-aware item generation. The system retrieves theoretically appropriate constructs from a literature-anchored database and leverages LLMs to generate semantically coherent, domain-specific items, supporting human-AI co-refinement. Its key innovation lies in the deep coupling of LLMs with an evidence-validated construct–item relational database, shifting scale development from experience-driven practice toward evidence-enhanced collaborative measurement. Experiments show a 62% reduction in design time, a 3.1× increase in item reuse, and significantly improved theoretical fidelity; expert evaluations across multiple rounds confirm ≥92% contextual appropriateness. The system has been integrated into a prototype HCI research workflow.

Improving rigor and efficiency in HCI measurement designLeveraging LLMs and prior literature for construct developmentStandardizing measurement item design process for researchers

Ambiguity in item difficulty calibration undermines measurement validity and test reusability in data visualization literacy assessment. Method: This paper proposes DRIVE-T—a methodology grounded in semiotic theory’s three-layer framework (syntactic, semantic, pragmatic) to model latent literacy constructs. It integrates structured task annotation, independent multi-rater scoring with inter-rater consistency validation, and the Many-Facets Rasch Model (MFRM) to empirically calibrate item difficulty and precisely align items with examinee ability. Contribution/Results: DRIVE-T enables automated selection of high-discriminating, representative items and supports scalable item bank development. Pilot validation demonstrates significant improvements in structural validity, expressive power, and cross-context reusability of assessments. The framework provides a generalizable methodological and technical foundation for formative literacy evaluation.

Address underspecification of difficulty levels in data visualization literacy assessmentsMeasure syntax, semantics, and pragmatics mastery in data visualizationPropose DRIVE-T for discriminative and representative item selection

This study addresses the current lack of systematic approaches for evaluating artificial intelligence’s adaptability and comprehension across diverse cultural contexts. Drawing on measurement theory, it introduces—for the first time—the validity framework from psychometrics into the assessment of AI cultural competence, thereby disentangling the construct of “cultural intelligence” from its operationalization. The work proposes a modular and extensible evaluation paradigm that integrates cultural dimension modeling, indicator design, data collection, and assessment protocols. By delineating core competency domains and their corresponding measurable indicators, this research establishes a theoretical and methodological foundation for large-scale, systematic evaluation of AI systems’ cultural adaptability.

AI evaluationcross-cultural competencecultural intelligence

This study challenges the prevailing view of language models as passive recorders of cultural phenomena, arguing instead that they function as active apparatuses that co-constitute cultural reality. Drawing on Karen Barad’s concept of “agential cuts” and adopting a material-discursive perspective, the research integrates natural language processing, qualitative analysis, and apparatus critique—illustrated through case studies such as film and television dialogue—to uncover the entanglements inherent in how models delineate cultural boundaries. Emphasizing ethical and theoretical reflexivity in methodological choices for cultural measurement, the work proposes a new paradigm that is both culturally sensitive and theoretically grounded. It further demonstrates how current model designs often erase cultural markers and diminish historical sensitivity, thereby shaping—and potentially distorting—the production and interpretation of cultural structures.

agential cutculturelanguage models

Latest Papers

What's happening recently
View more

This work addresses the lack of a unified semantic foundation in current software systems, which creates comprehension gaps among development, usage, and governance due to deficiencies in usability, modularity, and accountability. To bridge this divide, the paper proposes grounding software semantics in domain behavioral phenomena—specifically individuals, actions, and facts—as a shared conceptual vocabulary for stakeholders. This approach systematically integrates phenomenon-based modeling into software development by organizing behaviors into conceptual units, leveraging large language models (LLMs) to map semantics to modular, readable code, and establishing agent accountability through behavior-oriented norms. Empirical evaluation demonstrates that the proposed method significantly enhances the quality of usability design, improves the modularity and readability of LLM-generated code, and strengthens the accountability of autonomous agent behaviors.

accountabilitymeaningmodularity

Although large language models (LLMs) can achieve agreement with human annotators in text coding, their judgments may rely on superficial features unrelated to the underlying theoretical construct, thereby lacking construct validity. To address this issue, this work proposes a “fine-grained calibration” approach that decomposes theoretical constructs into clause-level components, validates each component against extractive evidence, and aggregates results according to explicit theoretical rules to assess whether LLMs genuinely measure the target construct. This method shifts the validation of construct validity from output consistency to process interpretability, enabling identification of errors stemming either from missing components or confusion with neighboring constructs. It establishes a transparent and interpretable paradigm for trustworthy measurement using LLMs in the social sciences.

coding reliabilityconstruct validitylarge language models

Current evaluations of large language models often conflate benchmark scores with true capabilities, overlooking issues such as test set contamination and annotation errors, thereby undermining construct validity. This work proposes a Structured Capability Model that, for the first time, jointly models the influence of model scale on intrinsic capabilities and the distortion introduced by measurement error within a unified framework, enabling quantitative assessment of evaluation construct validity. By integrating the strengths of scaling laws and latent factor models, our approach extracts interpretable and generalizable latent capability indicators from massive benchmark data. Evaluated on the OpenLLM Leaderboard, the proposed model demonstrates superior parsimonious fit compared to standard latent factor models and outperforms scaling laws in out-of-distribution benchmark prediction, exhibiting enhanced explanatory and predictive power.

benchmark evaluationconstruct validitylarge language models

This study addresses the limitations of existing AI evaluation methods, which often fail to align with real-world user needs, contextual nuances, and local policies, while manual assessment remains difficult to scale. To bridge this gap, the authors propose an auditable and iterative, context-aware evaluation framework that integrates persona-driven test case generation, domain-specific scoring rubrics, and a hybrid adjudication mechanism combining human reviewers and LLM-based judges. Automated scoring is activated only when sufficient agreement between LLM judgments and human annotations is achieved. A three-week pilot across four organizations involving 108 annotated question-answer pairs demonstrates that the approach effectively balances policy alignment with scalable automation, enabling reliable end-to-end evaluation of AI systems.

AI evaluationcontextual alignmenthuman-aligned scoring

This work addresses the limited fine-grained diagnostic capability of existing large-scale multilingual evaluations, which hinders effective model optimization. The authors propose the first reusable multilingual agent-based diagnostic framework, decomposing post-evaluation analysis into five stages: planning, aggregation, instance inspection, cross-cultural reflection, and report generation. Integrating an expert knowledge base with multilingual understanding and culture-aware modules, the framework enables deep attribution across 33 model families, 11 benchmarks, 26 languages, and 34 cultural contexts. Leveraging an expert-driven diagnostic set comprising 54 queries across 15 languages, the approach translates scores into actionable guidance, yielding diagnostic reports that outperform the strongest baseline by 47% in quality and prevail in 87.9% of pairwise comparisons against human experts. The study further distills four key insights regarding deployment strategies, iterative refinement, and cross-cultural risk mitigation.

benchmark insightscross-cultural assessmentfine-grained diagnosis

Hot Scholars

JK

JaeWon Kim

University of Washington
Human-Computer InteractionSocial Computing
YO

Yulia Otmakhova

Research Fellow grade 2 (eq. Assistant Professor), University of Melbourne
NLPbioNLP
LF

Lea Frermann

Computing and Information Systems, The University of Melbourne
Computational Linguisticscomputational cognitive modelingNLP for narrativesbias and fairness
PH

Pan Hui

Chair Professor, Nokia Chair in Data Science, FREng & IEEE Fellow (HKUST & University of Helsinki)
Ubiquitous ComputingMobile ComputingAugmented RealityData Science
PC

Paul C. Parsons

Associate Professor at Purdue University
human-computer interactionvisualizationapplied cognitiondesign