Sentence Splitter: Uncovering Latent Factual Structure for Self-Supervised Learning
This work proposes the Sentence Splitter framework, which formalizes the discovery of implicit factual structures in natural language sentences—specifically, descriptive prefixes and their corresponding factual completions—as a self-supervised learning task. Built upon the T5 architecture, the approach models sentence splitting as a discrete segmentation problem, generating self-supervised signals through templated symbolic pairs and recovering factual completions via probabilistic sequence generation. A lightweight bootstrapping mechanism further expands plausible prefix-completion structures. Without requiring human annotations, the method extracts structured prefix-tail pairs directly from raw text, effectively bridging symbolic knowledge with natural language. Experimental results demonstrate that the extracted supervision signals substantially enhance performance on downstream tasks such as knowledge graph completion and commonsense question answering, confirming the framework’s effectiveness and generalization capability.