Score
Designs and executes systematic coding schemes and annotation protocols to extract, label, and quantify framings, claims, risk categories, and governance provisions from policy and governance documents. Builds annotated corpora, document- and clause-level codebooks, and reproducible analysis workflows that produce coded datasets, prevalence counts, and comparative summaries for qualitative and quantitative investigation.
This study addresses the limitations of traditional framing codebooks when applied to large-scale, cross-cultural, or dynamically evolving news corpora—namely, ambiguous rules, difficulty adjudicating boundary cases, and insufficient theoretical adaptability. To overcome these challenges, the authors propose an interactive workflow that integrates theory-driven guidance with data-driven insights, positioning large language models (LLMs) as analytical collaborators. Through iterative human–LLM dialogue, the approach refines codebooks, externalizes decision logic, and surfaces latent framing dimensions. While preserving researchers’ interpretive authority, this method enhances the creativity and adaptability of the coding process. Empirical application to Latin American news corpora demonstrates its capacity to identify novel framing patterns and effectively recalibrate established theoretical frameworks within new cultural and contextual settings.
This study addresses the challenge of efficiently extracting governance variables from unstructured corporate governance documents, such as charters, by constructing the first standardized benchmark dataset and systematically evaluating the performance of various large language models (LLMs) on document-level binary classification tasks. Through comparative analysis of different prompt designs, task decomposition strategies, and document processing approaches, the research finds that most governance provisions can be extracted with high accuracy—achieving median performance close to theoretical upper bounds. While complex processing pipelines do not consistently enhance state-of-the-art LLMs, they substantially narrow the performance gap between advanced and lightweight models. These results demonstrate the feasibility of using LLMs for structured information extraction from complex legal texts and establish a reproducible benchmark and methodological framework for automated corporate governance research.
While large language models (LLMs) can achieve high accuracy in coding political events, they often fail to faithfully adhere to expert-defined coding rules, leading to unreliable behavior. This study systematically evaluates LLMs’ logical consistency under controlled perturbations—such as variations in label names and coding order—by enhancing structured codebooks through precise terminology, illustrative examples, retrieval-augmented context, and rules for challenging cases, combined with prompt engineering. The work reveals, for the first time, a critical disconnect between predictive performance and behavioral reliability, demonstrating that high accuracy does not necessarily imply compliance with social science coding logic. Although the refined codebooks substantially improve fine-grained classification performance, the models remain sensitive to minor codebook modifications, underscoring the necessity of explicit reliability assessment in computational social science applications.
This study addresses a critical limitation in existing document layout analysis methods, which treat figures and tables as generic objects and thus fail to identify semantically valuable, reusable analytical visual content—referred to as “data snapshots”—in institutional documents. The work introduces the novel task of data snapshot extraction, presents a benchmark dataset comprising humanitarian reports and World Bank policy papers, and proposes an evaluation framework that integrates spatial localization with semantic annotation. Systematic evaluation of multiple open-source layout models reveals consistent shortcomings in handling institutional documents, including confusion between analytical and non-analytical content, fragmentation of composite charts, and lack of contextual awareness. By exposing the generalization bottlenecks of current models in operational documents, this research provides a foundation for future advancements through the public release of its dataset and codebase.
Qualitative content analysis of institutional texts—such as legal rules, social norms, and strategic conventions—suffers from high coder subjectivity and a persistent theory-computation gap. Method: This study introduces a computationally grounded analytical framework based on Institutional Grammar 2.0, featuring the IG Parser tool and the first domain-specific formal grammar (IG Script), enabling full computational implementation of institutional grammar theory. Integrating NLP, rule-driven parsing, and a modular architecture, the framework achieves high-fidelity, automated conversion of natural-language texts into structured representations (JSON/XML/CSV). Contribution/Results: The approach significantly improves inter-coder reliability and analytical efficiency while preserving theoretical fidelity and supporting cross-paradigmatic institutional analysis. Its scalability and robustness have been empirically validated across diverse institutional domains, effectively bridging qualitative institutional theory and computational social science.
This study addresses the lack of empirical evaluation regarding whether existing dataset documentation frameworks effectively foster developer reflectivity. Combining mixed-methods thematic analysis with corpus-assisted discourse analysis, the research systematically examines how prevailing documentation frameworks—and their real-world instantiations—cover core dimensions of reflectivity. The findings reveal, for the first time, that current frameworks consistently overlook critical reflective themes. Building on this insight, the authors develop a reflectivity-oriented coding manual and propose an enhanced datasheet template incorporating targeted prompts to elicit deeper reflection. This work offers actionable strategies and practical tools to strengthen the reflective capacity of dataset documentation practices.
This study addresses the lack of systematic approaches for constructing, storing, and sharing high-quality annotated corpora. It proposes a generalizable and reusable end-to-end methodology encompassing annotation guideline development, corpus annotation, data storage, sharing mechanisms, and value realization, with an emphasis on full lifecycle management and cross-domain applicability. Integrating linguistic annotation theory, data management standards, and collaborative research practices, the approach is articulated through a structured framework and illustrative examples to yield a clear and actionable guide. The resulting methodology provides standardized support for diverse research domains, significantly enhancing the efficiency and quality with which researchers can build and utilize annotated textual data.
This study addresses the limitations of manual thematic analysis in clinical qualitative data—namely, poor scalability and low reproducibility—as well as the restricted generalizability and lack of traceability in existing large language model–based approaches. To overcome these challenges, the authors propose a novel automated thematic analysis framework that uniquely integrates iterative codebook refinement with end-to-end auditability. This integration substantially enhances the generalizability, consistency, and auditability of qualitative analyses. Empirical evaluation demonstrates that the method achieves the highest overall quality scores on four out of five datasets, with statistically significant improvements across all four evaluation metrics. Furthermore, in two pediatric cardiology corpora, the automatically generated themes exhibit strong alignment with expert annotations.
This work addresses the absence of structured, auditable, and jurisdiction-aware behavioral governance mechanisms in large language models during inference. We propose the Dynamic Behavioral Constraints (DBC) benchmark, introducing a taxonomy-based hierarchical governance framework that deploys 150 model-agnostic controls at the system prompt layer across 30 risk domains. The framework integrates six risk clusters and five adversarial attack strategies—including role-playing and authority impersonation—to establish a causally attributable, three-tier comparative evaluation protocol. Experimental results demonstrate that DBC reduces overall risk exposure from 7.19% to 4.55% (a 36.8% relative reduction), substantially outperforming standard safety prompts. The model achieves an MDBC compliance score of 8.7/10 and an EU AI Act alignment score of 8.5/10. All code and evaluation artifacts are open-sourced to enable automated compliance assessment and longitudinal tracking.