From Policy Documents to Structured Survey Responses: Evaluating Large Language Models for Policy Monitoring

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of science and technology policy monitoring, which traditionally relies on labor-intensive manual processes that are costly and difficult to scale across countries. To overcome these limitations, this work proposes a human-in-the-loop automated workflow powered by large language models (LLMs). The approach leverages long-context learning to generate structured questionnaire responses directly from policy texts, establishing an end-to-end pipeline encompassing data extraction, secondary LLM-based verification, and multi-source comparative analysis. Experimental results demonstrate that the proposed method achieves 84%–95% agreement rates on structured indicators. By substantially enhancing monitoring efficiency and cross-national scalability while preserving essential human oversight, this framework offers a robust and efficient new paradigm for large-scale policy tracking.
📝 Abstract
Science, technology, and innovation policies are crucial for competitiveness, yet their diversity and scale make them difficult to map and monitor consistently. Existing approaches rely heavily on manual survey efforts, which are costly and challenging to scale across countries. Large language models (LLMs) enable new possibilities for extracting and structuring information from long and unstructured policy documents. This paper presents an application of LLMs as "AI respondents" for generating structured survey responses from policy texts. We develop a data extraction pipeline based on long-context in-context learning to map information from public web sources into predefined survey categories, including policy instruments, target groups, and thematic areas. The pipeline integrates a validation step using a secondary LLM to assess relevance and evidence, alongside comparisons with human-provided responses. Using a multi-country dataset, we evaluate the alignment between LLM-generated and human-generated outputs through overlap measures and cross-validation. Results show that LLMs achieve high agreement for structured indicators (84-95%), while differences remain in free-text fields, where models tend to provide more detailed procedural descriptions. These findings highlight the potential of hybrid human-AI workflows for policy monitoring, improving both efficiency and scalability while maintaining the need for human validation and contextual interpretation.
Problem

Research questions and friction points this paper is trying to address.

Policy Monitoring
Large Language Models
Information Extraction
Science and Technology Policy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
In-context Learning
Data Extraction Pipeline
Policy Monitoring
Human-AI Workflow
💼 Related Jobs
No related jobs found.
C
Carolyn Cole
Reliable Intelligence Team, VTT Technical Research Centre of Finland Ltd., 02150 Espoo, Finland
M
Matthias Deschryvere
Reliable Intelligence Team, VTT Technical Research Centre of Finland Ltd., 02150 Espoo, Finland
Toqeer Ehsan
Toqeer Ehsan
Teknologian tutkimuskeskus VTT Oy
Natural Language ProcessingDeep LearningArtificial Intelligence
A
Arash Hajikhani
Reliable Intelligence Team, VTT Technical Research Centre of Finland Ltd., 02150 Espoo, Finland