Leveraging Weighted Syntactic and Semantic Context Assessment Summary (wSSAS) Towards Text Categorization Using LLMs

📅 2026-04-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of large language models in text classification, where stochastic attention mechanisms and sensitivity to noise often compromise accuracy and reproducibility. To mitigate these issues, the authors propose the wSSAS framework, which leverages signal-to-noise ratio (SNR) to identify high-value semantic features and organizes texts into a hierarchical “topic–narrative–cluster” structure. The framework incorporates a deterministic mechanism that jointly evaluates weighted syntactic and semantic contextual cues and employs a Summary-of-Summaries architecture to aggregate salient information. Empirical evaluations on multi-domain review datasets from Google, Amazon, and Goodreads demonstrate that wSSAS significantly reduces classification entropy, enhances clustering completeness, and improves classification accuracy, thereby validating its effectiveness in bolstering result stability and robustness against noise.

Technology Category

Natural Language Processing: SummarizationMachine Learning: Large Multimodal Models (LMMs)Reasoning under Uncertainty: Stochastic Optimization

Application Category

Web Mining and Content Analysis: Normalization, clustering, classification, and summarization of Web textSearch and Retrieval-Augmented AI: Large language models for searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
The use of Large Language Models (LLMs) for reliable, enterprise-grade analytics such as text categorization is often hindered by the stochastic nature of attention mechanisms and sensitivity to noise that compromise their analytical precision and reproducibility. To address these technical frictions, this paper introduces the Weighted Syntactic and Semantic Context Assessment Summary (wSSAS), a deterministic framework designed to enforce data integrity on large-scale, chaotic datasets. We propose a two-phased validation framework that first organizes raw text into a hierarchical classification structure containing Themes, Stories, and Clusters. It then leverages a Signal-to-Noise Ratio (SNR) to prioritize high-value semantic features, ensuring the model's attention remains focused on the most representative data points. By incorporating this scoring mechanism into a Summary-of-Summaries (SoS) architecture, the framework effectively isolates essential information and mitigates background noise during data aggregation. Experimental results using Gemini 2.0 Flash Lite across diverse datasets - including Google Business reviews, Amazon Product reviews, and Goodreads Book reviews - demonstrate that wSSAS significantly improves clustering integrity and categorization accuracy. Our findings indicate that wSSAS reduces categorization entropy and provides a reproducible pathway for improving LLM based summaries based on a high-precision, deterministic process for large-scale text categorization.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
text categorization
attention mechanisms
noise sensitivity
reproducibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

wSSAS
deterministic framework
Signal-to-Noise Ratio (SNR)
hierarchical classification
Summary-of-Summaries (SoS)