🤖 AI Summary
This study addresses the challenge that complex event sequences in U.S. employment discrimination complaints are inadequately captured by traditional lexical or embedding-based representations. To this end, it proposes an Event Knowledge Graph (EKG) construction paradigm that integrates structured generation with multi-granularity merging. Specifically, the method designs a source-grounding pipeline based on a 5W1H heuristic schema, synergizing domain-specific legal models with large language models for structured information extraction to generate document-level EKGs that consolidate participants, temporal dynamics, and causal relations, thereby enhancing evidence organization and reasoning capabilities. Experimental results demonstrate that graph-structured classifiers significantly outperform baseline models, and EKG retrieval effectively improves bounded question-answering performance, validating its practical value in legal evidence reasoning.
📝 Abstract
U.S. employment-discrimination complaints describe complex event sequences that are not explicitly captured by lexical or embedding-based representations alone. We present ARGUS, a source-grounded pipeline that combines a 5W1H-inspired schema, legal-domain models, and LLM-based structured generation to construct document-level Event Knowledge Graphs (EKGs) from CourtListener complaints. ARGUS extracts fact-bearing statements, builds chunk-level event graphs with participant, temporal, and causal structure, and merges them into document-level representations. We evaluate graph quality through human and multi-model assessment and test downstream utility on claim classification and legal QA. The graph-structured classifier outperforms raw and linearized baselines on the held-out set, and EKG-only retrieval improves document-scoped QA, while open-retrieval gains remain limited by low first-stage candidate recall. These results suggest that EKGs are most useful for organizing and reasoning over evidence once relevant material has been retrieved.