Institution profile

Max Planck Institute for Psycholinguistics

Academic institutioneurope · nl
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

Perceptually Grounded and Semantics-Aware Evaluation for Holistic Co-Speech Gesture Generation

Oct 08, 2026

This study addresses the disconnect between objective metrics and human perception in co-speech gesture generation, along with the difficulty of quantifying semantic appropriateness. To this end, it constructs a comprehensive benchmark integrating standardized comparisons, human-centric validation, and fine-grained semantic evaluation. Methodologically, multimodal large language models are leveraged to enhance data annotation, and a Semantic Gesture Preservation (SGP) metric is proposed alongside a perception-aligned composite evaluation framework. Experimental results demonstrate that SGP correlates significantly with human subjective judgments, while the composite metric effectively improves perceptual alignment across all evaluated dimensions. These findings confirm the necessity of calibrating objective metrics against subjective evaluations for more faithful assessment of generated gestures.

0 citationsRead paper

RICE-Alpha: Reliability-Informed Correction with Event Graphs for LLM-Agent Stock Forecasting

Sep 27, 2026

This study addresses the difficulty of large language model (LLM)-based financial agents in modeling the continuity and transition reliability of historical events by proposing a point-in-time constrained stock scoring framework that decouples multi-view base alphas from reliability-calibrated residual correction terms. Methodologically, the framework introduces a hierarchical memory layer and typed event agents to aggregate cross-firm event successor relationships via local pairing. It further integrates LLMs, graph neural networks, and reliability calibration techniques for event state modeling and residual correction. Experimental results demonstrate that the proposed framework doubles the information coefficient information ratio (ICIR), achieving net Sharpe ratios of 1.656 and 1.725 on U.S. and Hong Kong stock markets, respectively, significantly outperforming baseline models.

0 citationsRead paper

Investigating Multimodal Informativity under Different Partner Visibility Conditions in Video-Mediated Dialogue

Aug 09, 2026

This study investigates how interactional visibility modulates the informational contributions of gesture and speech in referential tasks within video-mediated dialogue. By constructing models based on speech transcripts, gesture skeleton sequences, and their multimodal fusion—trained with alignment between representations and referential images—the research systematically evaluates the referential efficacy of each modality under varying visibility conditions. Findings demonstrate that gestures possess independent referential capacity, and multimodal fusion yields the greatest performance gains when speech is highly ambiguous. Moreover, human behavioral analyses reveal that visibility not only regulates the informativeness of gestural production but also elicits cross-turn verbal coordination, underscoring the critical role of pragmatic factors in multimodal reference.

0 citationsRead paper

Entity tracking emerges in sub-billion parameter language models and exceeds human performance in naturalistic narratives

Jun 04, 2026

This study addresses the absence of human baselines and the reliance on manual task construction in evaluating entity tracking capabilities of language models by introducing a fair comparative benchmark grounded in natural narratives. Through multi-complexity narrative testing and controlled human behavioral experiments, we systematically assess entity tracking across language models of varying scales. Our findings reveal that models with merely 410M parameters achieve human-level performance, with larger models significantly surpassing it as scale increases. Furthermore, degradation in entity tracking is primarily driven by narrative complexity rather than length. This work reshapes our understanding of the emergence thresholds for core linguistic capabilities, demonstrating that reliable language comprehension emerges at model scales substantially smaller than previously anticipated.

0 citationsRead paper
Recent publications

Latest Papers

Perceptually Grounded and Semantics-Aware Evaluation for Holistic Co-Speech Gesture Generation

Oct 08, 2026

This study addresses the disconnect between objective metrics and human perception in co-speech gesture generation, along with the difficulty of quantifying semantic appropriateness. To this end, it constructs a comprehensive benchmark integrating standardized comparisons, human-centric validation, and fine-grained semantic evaluation. Methodologically, multimodal large language models are leveraged to enhance data annotation, and a Semantic Gesture Preservation (SGP) metric is proposed alongside a perception-aligned composite evaluation framework. Experimental results demonstrate that SGP correlates significantly with human subjective judgments, while the composite metric effectively improves perceptual alignment across all evaluated dimensions. These findings confirm the necessity of calibrating objective metrics against subjective evaluations for more faithful assessment of generated gestures.

0 citationsRead paper

RICE-Alpha: Reliability-Informed Correction with Event Graphs for LLM-Agent Stock Forecasting

Sep 27, 2026

This study addresses the difficulty of large language model (LLM)-based financial agents in modeling the continuity and transition reliability of historical events by proposing a point-in-time constrained stock scoring framework that decouples multi-view base alphas from reliability-calibrated residual correction terms. Methodologically, the framework introduces a hierarchical memory layer and typed event agents to aggregate cross-firm event successor relationships via local pairing. It further integrates LLMs, graph neural networks, and reliability calibration techniques for event state modeling and residual correction. Experimental results demonstrate that the proposed framework doubles the information coefficient information ratio (ICIR), achieving net Sharpe ratios of 1.656 and 1.725 on U.S. and Hong Kong stock markets, respectively, significantly outperforming baseline models.

0 citationsRead paper

Investigating Multimodal Informativity under Different Partner Visibility Conditions in Video-Mediated Dialogue

Aug 09, 2026

This study investigates how interactional visibility modulates the informational contributions of gesture and speech in referential tasks within video-mediated dialogue. By constructing models based on speech transcripts, gesture skeleton sequences, and their multimodal fusion—trained with alignment between representations and referential images—the research systematically evaluates the referential efficacy of each modality under varying visibility conditions. Findings demonstrate that gestures possess independent referential capacity, and multimodal fusion yields the greatest performance gains when speech is highly ambiguous. Moreover, human behavioral analyses reveal that visibility not only regulates the informativeness of gestural production but also elicits cross-turn verbal coordination, underscoring the critical role of pragmatic factors in multimodal reference.

0 citationsRead paper

Entity tracking emerges in sub-billion parameter language models and exceeds human performance in naturalistic narratives

Jun 04, 2026

This study addresses the absence of human baselines and the reliance on manual task construction in evaluating entity tracking capabilities of language models by introducing a fair comparative benchmark grounded in natural narratives. Through multi-complexity narrative testing and controlled human behavioral experiments, we systematically assess entity tracking across language models of varying scales. Our findings reveal that models with merely 410M parameters achieve human-level performance, with larger models significantly surpassing it as scale increases. Furthermore, degradation in entity tracking is primarily driven by narrative complexity rather than length. This work reshapes our understanding of the emergence thresholds for core linguistic capabilities, demonstrating that reliable language comprehension emerges at model scales substantially smaller than previously anticipated.

0 citationsRead paper