Integrating knowledge graphs and multilingual scholarly corpora for domain-adaptive LLMs in SSH

๐Ÿ“… 2026-07-07
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses key challenges in applying large language models (LLMs) to the social sciences and humanities (SSH), including disciplinary heterogeneity, limited access to multilingual scholarly literature, and insufficient evaluability of outputs. To overcome these issues, the work proposes a domain-adaptive framework that uniquely integrates knowledge graphs with multilingual academic corpora, deeply coupling domain sensitivity, regulatory compliance, and generative AI for the first time. The approach leverages knowledge graph embeddings, retrieval-augmented generation, multilingual fine-tuning, and ethical compliance mechanisms, all aligned with the LLMs4EU evaluation protocol, to develop a trustworthy, traceable, and responsible SSH-specific model. Experimental results demonstrate strong performance across retrieval, summarization, traceability, and hallucination detection metrics, with qualitative validation by digital humanities experts confirming its scholarly applicability and reliability.
๐Ÿ“ Abstract
The integration of Large Language Models (LLMs) into scientific research workflows, particularly for bibliographic discovery and literature synthesis, raises significant methodological, epistemic and regulatory challenges for the Social Sciences and Humanities (SSH), especially with regard to disciplinary diversity, multilingual access to sources and the evaluation of results. This paper presents an on-going use case developed within the European project LLMs4EU and the ALT-EDIC infrastructure, aimed at adapting foundation models to SSH research practices and supporting tasks such as question answering, comparative document analysis and literature review. The evaluation framework follows the LLMs4EU protocol and encompasses both independent quantitative benchmarking (retrieval, summarisation, traceability and hallucination detection) and a qualitative assessment involving a panel of Digital Humanities experts. By embedding model adaptation within research infrastructures and a structured legal and ethical compliance framework, the use case explores how domain-sensitive and regulation-aware generative AI can support SSH scholarship while preserving reliability and epistemic responsibility.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Social Sciences and Humanities
multilingual scholarly corpora
domain adaptation
epistemic responsibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

knowledge graphs
multilingual scholarly corpora
domain-adaptive LLMs
epistemic responsibility
legal and ethical compliance
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
A
Adam Faci
Huma-Num, CNRS, Paris, France
Alessio Miaschi
Alessio Miaschi
Researcher in Computational Linguistics, ItaliaNLP Lab @ CNR-ILC, Pisa
Natural Language ProcessingComputational LinguisticsLanguage ModelsDeep Learning
A
Anne Combe
Inria, France
P
Pascal Cuxac
Inist, CNRS, 2 rue Jean Zay, 54500 Vandoeuvre-lรจs-Nancy, France
Francesca Frontini
Francesca Frontini
Istituto di Linguistica Computazionale A.Zampolli - CNR - Pisa
computational linguisticscomputational lexicographylinguisticsdigital humanities
N
Nicolas Larrousse
Huma-Num, CNRS, Paris, France
S
Stรฉphane Pouyllau
Huma-Num, CNRS, Paris, France