Institution profile

Haverford College

Academic institutionnorthamerica · us
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

Cross-Lingual Alignment for Decoder-Only Models using MoE Routers

Oct 01, 2026

This study addresses the challenge of cross-lingual representation alignment in decoder-only large language models, which arises from tokenization discrepancies across languages. To tackle this issue, this work proposes a novel contrastive learning paradigm that leverages Mixture-of-Experts (MoE) router outputs as alignment anchors. Departing from conventional auxiliary losses applied to hidden states, the method constructs robust sequence-level alignment objectives through pooling and optimizes them via controlled continual pre-training. Experimental results demonstrate that the proposed approach effectively aligns lower-layer hidden representations, substantially enhancing the multilingual performance of open-source MoE models. Overall, this research offers a promising new direction for achieving cross-lingual alignment in large language models.

0 citationsRead paper

Accounting for Stochasticity in Studies of Large Language Model Refusal

Sep 27, 2026

This study addresses the significant limitations of existing single-observation criteria in evaluating the stochasticity of large language model (LLM) refusal behaviors, demonstrating that a single query cannot yield accurate assessments. To overcome this, we propose a longitudinal auditing framework grounded in Bernoulli process modeling, conducting large-scale repeated prompting experiments on GPT-4.1 to quantitatively analyze the stability and decision boundaries of its refusal behavior. Our findings reveal that approximately 20% of queries reside within the decision boundary region, indicating that single observations severely underestimate the sample size required for reliable evaluation. Consequently, we establish a new standard requiring 15 to 25 repeated queries to robustly quantify model refusal behaviors, thereby providing a more rigorous methodological foundation for LLM safety assessment.

0 citationsRead paper

Benchmarking LLM Competence on Logical Inference over Probability Operators

Jul 29, 2026

This study addresses the difficulty large language models (LLMs) face in distinguishing genuine logical reasoning from superficial pattern matching when processing probabilistic expressions such as “might” or “must,” revealing a fundamental deficiency in their capacity for logical inference under uncertainty. The authors present the first systematic benchmark comprising 14,320 synthetically generated English samples across 15 reasoning templates, carefully controlling variables including question formulation, negation strategies, and surface-level content to isolate models’ logical capabilities with modal expressions. Introducing a “reasoning floor” metric—defined as the worst-case accuracy on Yes/No tasks—the work evaluates whether models exhibit true reasoning rather than response biases. Evaluation of 29 prominent LLMs shows only nine surpass random performance, with most exhibiting significant, logic-irrelevant deviations tied to question phrasing, verb animacy, and named entity attributes.

0 citationsRead paper

ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation

Jul 25, 2026

This study addresses the limitations of existing cross-lingual reasoning benchmarks, which rely on translated English datasets and thus introduce linguistic bias while failing to assess models’ analogical reasoning within authentic cultural contexts. To overcome this, the authors propose the first language-agnostic, culturally grounded evaluation framework for analogical reasoning. They construct native, high-difficulty analogy datasets for Arabic, Amharic, and Japanese through a collaborative process involving native speakers and large language models, entirely bypassing translation. Experiments across 14 open-source models reveal a stark performance gap: despite strong results on English proverb-based tasks, model accuracy drops by 12–52 percentage points on the localized benchmarks, exposing significant deficiencies in cultural reasoning. The complete pipeline, datasets, and evaluation suite are publicly released.

0 citationsRead paper

Biases in Large Language Model-Elicited Text: A Case Study in Natural Language Inference

Mar 06, 2025International Conference on Computational Linguistics

This work investigates whether natural language inference (NLI) data generated by large language models (LLMs)—specifically GPT-4, Llama-2-70b, and Mistral-7b—inherit annotation artifacts and societal biases (e.g., gender, race, age) present in human-annotated NLI datasets. We construct LLM-generated NLI subsets and empirically identify severe hypothesis exclusivity bias and stereotypical social biases—first such evidence in synthetic NLI data. To detect these biases systematically, we propose a dual-path framework: (1) a fine-tuned BERT-based hypothesis exclusivity classifier, and (2) pointwise mutual information (PMI) analysis for bias-associated lexical patterns. Experiments show the framework achieves 86–96% classification accuracy on LLM-generated data—significantly outperforming its performance on human-annotated data—and quantitatively identifies multiple bias-correlated lexical terms. Our findings provide both novel methodology and critical empirical evidence for bias assessment and mitigation in LLM-synthesized training data.

0 citationsRead paper
Recent publications

Latest Papers

Cross-Lingual Alignment for Decoder-Only Models using MoE Routers

Oct 01, 2026

This study addresses the challenge of cross-lingual representation alignment in decoder-only large language models, which arises from tokenization discrepancies across languages. To tackle this issue, this work proposes a novel contrastive learning paradigm that leverages Mixture-of-Experts (MoE) router outputs as alignment anchors. Departing from conventional auxiliary losses applied to hidden states, the method constructs robust sequence-level alignment objectives through pooling and optimizes them via controlled continual pre-training. Experimental results demonstrate that the proposed approach effectively aligns lower-layer hidden representations, substantially enhancing the multilingual performance of open-source MoE models. Overall, this research offers a promising new direction for achieving cross-lingual alignment in large language models.

0 citationsRead paper

Accounting for Stochasticity in Studies of Large Language Model Refusal

Sep 27, 2026

This study addresses the significant limitations of existing single-observation criteria in evaluating the stochasticity of large language model (LLM) refusal behaviors, demonstrating that a single query cannot yield accurate assessments. To overcome this, we propose a longitudinal auditing framework grounded in Bernoulli process modeling, conducting large-scale repeated prompting experiments on GPT-4.1 to quantitatively analyze the stability and decision boundaries of its refusal behavior. Our findings reveal that approximately 20% of queries reside within the decision boundary region, indicating that single observations severely underestimate the sample size required for reliable evaluation. Consequently, we establish a new standard requiring 15 to 25 repeated queries to robustly quantify model refusal behaviors, thereby providing a more rigorous methodological foundation for LLM safety assessment.

0 citationsRead paper

Benchmarking LLM Competence on Logical Inference over Probability Operators

Jul 29, 2026

This study addresses the difficulty large language models (LLMs) face in distinguishing genuine logical reasoning from superficial pattern matching when processing probabilistic expressions such as “might” or “must,” revealing a fundamental deficiency in their capacity for logical inference under uncertainty. The authors present the first systematic benchmark comprising 14,320 synthetically generated English samples across 15 reasoning templates, carefully controlling variables including question formulation, negation strategies, and surface-level content to isolate models’ logical capabilities with modal expressions. Introducing a “reasoning floor” metric—defined as the worst-case accuracy on Yes/No tasks—the work evaluates whether models exhibit true reasoning rather than response biases. Evaluation of 29 prominent LLMs shows only nine surpass random performance, with most exhibiting significant, logic-irrelevant deviations tied to question phrasing, verb animacy, and named entity attributes.

0 citationsRead paper

ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation

Jul 25, 2026

This study addresses the limitations of existing cross-lingual reasoning benchmarks, which rely on translated English datasets and thus introduce linguistic bias while failing to assess models’ analogical reasoning within authentic cultural contexts. To overcome this, the authors propose the first language-agnostic, culturally grounded evaluation framework for analogical reasoning. They construct native, high-difficulty analogy datasets for Arabic, Amharic, and Japanese through a collaborative process involving native speakers and large language models, entirely bypassing translation. Experiments across 14 open-source models reveal a stark performance gap: despite strong results on English proverb-based tasks, model accuracy drops by 12–52 percentage points on the localized benchmarks, exposing significant deficiencies in cultural reasoning. The complete pipeline, datasets, and evaluation suite are publicly released.

0 citationsRead paper

Biases in Large Language Model-Elicited Text: A Case Study in Natural Language Inference

Mar 06, 2025International Conference on Computational Linguistics

This work investigates whether natural language inference (NLI) data generated by large language models (LLMs)—specifically GPT-4, Llama-2-70b, and Mistral-7b—inherit annotation artifacts and societal biases (e.g., gender, race, age) present in human-annotated NLI datasets. We construct LLM-generated NLI subsets and empirically identify severe hypothesis exclusivity bias and stereotypical social biases—first such evidence in synthetic NLI data. To detect these biases systematically, we propose a dual-path framework: (1) a fine-tuned BERT-based hypothesis exclusivity classifier, and (2) pointwise mutual information (PMI) analysis for bias-associated lexical patterns. Experiments show the framework achieves 86–96% classification accuracy on LLM-generated data—significantly outperforming its performance on human-annotated data—and quantitatively identifies multiple bias-correlated lexical terms. Our findings provide both novel methodology and critical empirical evidence for bias assessment and mitigation in LLM-synthesized training data.

0 citationsRead paper