paired response analysis

Design and implement methods to create and validate matched pairs of responses and to compare paired data for alignment, disagreement, or correspondence, using pair-matching and matched-pair analytical techniques. Build and apply statistical and computational comparisons (agreement metrics, difference scoring, visualization, and hypothesis tests) and analyze contextual variables associated with mismatches to quantify, characterize, and explain misalignments and inform strategies to address them.

pairedresponseanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.44
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Text-Based Approaches to Item Alignment to Content Standards in Large-Scale Reading & Writing Tests

Sep 30, 2025
YF
Yanbin Fu
🏛️ University of Maryland, College Park

In large-scale assessment development, aligning test items with content standards (domain/skill-level taxonomies) often relies on subjective, labor-intensive manual annotation. This paper proposes an automated alignment method based on fine-tuned small language models (SLMs), integrating multilingual E5-large-instruct embeddings with supervised learning and conducting semantic analysis via cosine similarity, KL divergence, and 2D projection. Experiments demonstrate that augmenting item text data substantially improves SLM performance, enabling superior skill-level alignment accuracy compared to conventional embedding approaches. Semantic analysis further reveals non-negligible semantic overlap among certain skills in SAT/PSAT assessments, contributing to misclassification. The proposed framework offers a scalable, interpretable, and lightweight technical pathway for enhancing the efficiency and rigor of validity evidence generation in test construction.

Automating alignment of test items to content standardsEvaluating fine-tuned language models for educational assessmentOvercoming subjectivity in human expert alignment process

AlignSurvey: A Comprehensive Benchmark for Human Preferences Alignment in Social Surveys

Nov 11, 2025
CL
Chenxi Lin
🏛️ Zhejiang University | Zhejiang Gongshang University

Traditional social surveys suffer from rigidity, high costs, and poor cross-cultural equivalence, while existing LLM research predominantly focuses on structured questionnaires, neglecting end-to-end modeling and exacerbating underrepresentation of marginalized groups due to data bias. To address these gaps, we propose the first full-pipeline social survey benchmark spanning four stages: role modeling, semi-structured interviewing, attitude/stance inference, and response generation. We introduce a multi-level evaluation framework that jointly optimizes diversity and fairness. Leveraging open-source LLMs, we perform two-stage fine-tuning—integrating expert annotations with nationally representative survey data—to develop the SurveyLM model family. We release a multilevel dataset comprising over 44,000 dialogues and 400,000 questionnaire records, alongside fully open-sourced code, models, and tooling. Our approach significantly improves LLMs’ fidelity, alignment, and representational fairness in social survey tasks.

Lack comprehensive benchmarks for evaluating alignment fidelity and fairness in surveysLLM-based survey simulations overlook full pipeline and risk under-representing marginalized groupsTraditional surveys face fixed formats and high costs with limited adaptability

Estimating the functional relationship between a continuous exposure and a binary outcome is challenging when covariates are measured with error. This study presents the first systematic evaluation of Simulation-Extrapolation, Regression Calibration, multiple imputation, and Bayesian correction methods, each coupled with flexible modeling techniques—including B-splines, P-splines, and fractional polynomials—within a multi-team, fully blinded, neutral simulation framework. By generating 155 distinct simulation scenarios and repeated samples, the research quantifies the bias and variance of each approach, revealing their relative strengths and limitations. The findings not only inform method selection under measurement error but also demonstrate the feasibility and value of this neutral comparative paradigm for rigorous methodological assessment.

covariate adjustmentexposure-outcome relationshipfunctional form

This study addresses the lack of a unified computational, management, and visualization framework for characterizing diverse pairwise associations—such as linear correlation, nonlinear dependence, and Simpson’s paradox—between numeric and categorical variables. Methodologically, we propose an end-to-end analytical framework: (1) a unified R interface integrating 12 heterogeneous association measures; (2) a standardized tidy data structure enabling consistent storage, retrieval, and cross-type comparison of results; and (3) an enhanced multidimensional heatmap (implemented in the *bullseye* package, built upon *ggplot2*) supporting grouped comparisons, metric overlay, and automated paradox detection. Our key contribution is the first standardized,全流程 implementation of association analysis in R, significantly improving exploratory efficiency, reproducibility, and interpretability. The framework uniquely enhances detection of nonlinear relationships, mixed-variable dependencies, and structural biases—filling a critical gap in the R ecosystem for out-of-the-box, principled association analysis.

Enhancing traditional heatmaps with richer relationship visualizationsManaging and visualizing multiple pairwise correlation scoresProviding uniform interface for diverse association calculations

Existing evaluation metrics for survey simulation are fragmented and lack standardization, often overlooking the critical dimension of response option alignment, which hinders meaningful comparison of model performance. To address this gap, this work proposes RADIUS—the first two-dimensional evaluation framework that jointly incorporates rank alignment and distribution alignment. RADIUS systematically assesses the quality of large language models in survey simulation by integrating rank consistency measures, distributional similarity metrics (e.g., KL divergence), and statistical significance testing. The framework not only exposes the limitations of conventional metrics but also establishes a more reliable, comparable, and decision-relevant benchmark. To foster standardized evaluation practices in the research community, the authors open-source the RADIUS implementation.

distribution alignmentevaluation metricsLLM

Latest Papers

What's happening recently
View more

This work addresses the lack of systematic alignment between large language models and human reviewers in survey evaluation, as well as the absence of a multidimensional, quantifiable assessment framework. To bridge this gap, the authors introduce SurveyReview—the first benchmark specifically designed for survey reviewing—comprising 675 survey papers and 1,630 structured review reports, along with standardized data splits and evaluation protocols. Building upon Qwen3-32B with LoRA fine-tuning and external knowledge augmentation, the proposed strong baseline model, SurveyAlign, translates free-form reviews into scores and justifications across four dimensions: readability, criticality, comprehensiveness, and structure. Experimental results demonstrate that SurveyAlign significantly outperforms GPT-5.2 with prompt-based evaluation, reducing average MSE from 2.28 to 1.38 and MAE from 1.15 to 0.69 on the test set, thereby substantially improving alignment with human reviewers.

benchmarklarge language modelspeer review

This study addresses systematic biases in large language models (LLMs) when simulating survey responses, including skewed marginal distributions, poor variance calibration, and attenuated variable relationships. It proposes a novel decomposition of simulation fidelity into three quantifiable dimensions—structural, marginal, and individual—and systematically evaluates the multidimensional fidelity of three mitigation strategies: prompt engineering, output post-processing, and few-shot fine-tuning, using small-scale pilot data. Empirical results demonstrate that few-shot fine-tuning achieves a favorable balance across fidelity dimensions; however, uneven fidelity across subpopulations may compromise the consistent alignment of diverse viewpoints.

LLM-based survey simulationpluralistic alignmentsmall pilot data

This study addresses the pervasive issue of measurement error in both outcome variables and multiple covariates within routinely collected biomedical data, such as electronic health records, which, if uncorrected, can induce analytical bias and misinform clinical decisions. For the first time within a tutorial framework, it systematically reviews and empirically compares several methods capable of simultaneously correcting measurement error in both outcomes and multiple covariates—including regression calibration, SIMEX, instrumental variable approaches, and modeling strategies leveraging validation subsamples. Through a unified illustrative example and publicly available code, the work not only clarifies the relative performance of these methods in real-world data to guide researchers’ methodological choices but also establishes a reproducible end-to-end analytical pipeline and highlights promising directions for future research.

biomedical researchcovariatesmeasurement error

This work addresses the lack of intuitive, immediate feedback on fitting errors in existing model-fitting approaches. It proposes an interactive fitting framework that integrates visual and auditory feedback: as users manipulate parametric curves, the system synthesizes audio in real time, with greater model-data discrepancies producing louder and more dissonant sounds. This is the first approach to incorporate auditory cues into model exploration, enabling multisensory assessment of fit quality. Combining interactive visualization, real-time audio synthesis, and Gaussian process regression, the method demonstrates effectiveness and generalizability across four diverse case studies—golf putting, dilution experiments, cosmological parameter estimation, and temperature data fitting—significantly enhancing users’ intuitive perception of model misfit.

curve fittingdata visualizationmodel exploration

This study addresses the challenges of complex sample size calculations and difficult interpretation of Kappa coefficients in inter-rater agreement analysis, particularly for researchers lacking programming expertise. To this end, we developed an open-source, interactive web application based on R Shiny that integrates the core functionalities of the kappaSize and irr packages into a unified platform for the first time. The tool supports computation of Cohen’s, Fleiss’, and Light’s Kappa statistics and automatically interprets results according to the Landis & Koch benchmark scale. Offering both a graphical user interface and command-line access, it substantially lowers the technical barrier to conducting agreement analyses. Published on CRAN, this application provides an efficient, user-friendly, and reproducible solution for consistency assessment in categorical data.

agreement analysiscategorical datainter-rater agreement

Hot Scholars

JH

Junfeng He

Research Scientist, Google Research
Machine LearningComputer VisionRankingHCI
SM

Sheng Meng

institute of physics, chinese academy of science
first-principles quantum dynamicsdensity functional theory
JB

Jonathan Berant

Professor, Tel-Aviv University, Visiting Faculty Researcher, Google DeepMInd
Natural Language ProcessingMachine Learning
TG

Tanya Goyal

Cornell University
Natural Language Processing