LLM-Assisted Review Prioritization for German Statutory Health Insurance Websites: A Multi-Stage Corpus Audit

📅 2026-08-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of auditing the vast and complex content of Germany’s statutory health insurance websites, which exceeds manual review capacity and eludes detection by general-purpose AI tools due to nuanced medical, legal, and editorial issues. The authors propose a reproducible, multi-stage auditing framework that integrates deterministic rule-based filtering, large language model–assisted triage, temporal validity checks, and a dual-model comparison mechanism. Crucially, the approach distinguishes genuine content quality issues from superficial AI-generated signals without relying on AI-detection heuristics. Applied to 56,198 webpages, the framework prioritized content for review, producing 35,998 audit records and directing 21,452 pages into in-depth examination. In a stress test of 300 pages, it identified anomalies in 33.3% of cases, with the dual-model component achieving 75.8% agreement (Cohen’s κ = 0.532) across 182 matched instances.
📝 Abstract
Background: German statutory health insurance (SHI) funds publish web portfolios that exceed continuous specialist review capacity. Their content can shape health and benefit expectations. Generic AI-text detection does not identify medical, benefit, legal, or editorial review needs. Objective: To characterize a multi-stage workflow that prioritizes substantive review needs while separating AI-provenance signals from quality claims. Methods: We analyzed 56,198 pages from 84 SHI websites or sub-sites. The workflow combined deterministic screening, model-assisted triage and in-depth review, minimum evidence checks, temporal-validity safeguards, and paired-model comparison. It is reproducibility-bounded, not a validated detector. Production code is proprietary; reproducibility rests on frozen derived tables and paired-comparison artifacts. The 300-page lower-priority check was a single-model, risk-enriched routing stress test, not a human-reference evaluation. Results: All pages received a review state. The workflow generated 35,998 review records and routed 21,452 to case review. The workload concentrated in transparency, legal framing, medical content, contradictions, and AI-related failure-mode signals. A quoted passage was locatable in captured page text for 31,347 records, confirming literal occurrence rather than factual correctness. The routing stress test surfaced a signal on 100/300 pages (33.3% within the sample). Across 182 matched cases, two models agreed in 75.8% (kappa = 0.532; 95% CI 0.415-0.649). Conclusions: The workflow produces a prioritized workload, not error prevalence or final legal, medical, or insurer-level findings. It neither proves AI authorship nor validates autonomous detection. Paired-model agreement quantifies consistency, not correctness or sufficient triage performance; public claims require human adjudication.
Problem

Research questions and friction points this paper is trying to address.

statutory health insurance
review prioritization
AI-generated content
content quality
corpus audit
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-stage workflow
review prioritization
paired-model comparison
AI-provenance signals
corpus audit
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Martin Möller