🤖 AI Summary
This study addresses the challenge of transforming unstructured text into interpretable knowledge for precise retrieval and reasoning by constructing an end-to-end pipeline from raw corpora to semantic structuring. Methodologically, it introduces the Binary Bleed algorithm to optimize the search complexity of non-negative matrix factorization (NMF), develops HNMFk for adaptive hierarchical topic modeling, and designs a T-SRAG dynamic routing mechanism. By synergizing knowledge graphs with vector databases, the framework enables link prediction-based reasoning and contrastive learning alignment. Experimental results demonstrate that the proposed system significantly enhances retrieval precision in domains such as cybersecurity, effectively mitigates generative hallucinations, and supports early trend detection and hypothesis generation.
📝 Abstract
This dissertation presents a scalable architecture for transforming unstructured, domain-specific text into structured knowledge for retrieval and reasoning. It integrates semi-automatic corpus curation, semantic structuring, retrieval, and inference into an interpretable pipeline.
The research introduces Binary Bleed, an adapted binary search method that reduces low-rank search complexity for Non-negative Matrix Factorization (NMF), and Hierarchical NMF with automatic latent feature selection (HNMFk), a depth-adaptive topic modeling method that produces interpretable taxonomies guided by subject matter experts. These representations populate a typed Knowledge Graph and a semantically aligned Vector Store containing extracted latent features, synchronized through an event-driven substrate.
Tensor-Structured Retrieval-Augmented Generation (T-SRAG) dynamically routes queries across retrieval paths. Contrastive alignment maps document and query embeddings to hierarchical topic structures to improve semantic fidelity and reduce hallucinations. Beyond retrieval, tensor-based link prediction identifies and completes missing links in the Knowledge Graph, supporting inference grounded in citation structure.
Applications across cybersecurity, law, materials science, and healthcare demonstrate improvements in retrieval precision, early trend detection, hypothesis generation, and hallucination mitigation. The dissertation provides a deployable, modular foundation for trustworthy, domain-specific AI systems that retrieve and reason over structured knowledge.