🤖 AI Summary
This study addresses the inefficiency and limited accuracy of aspect-based sentiment analysis (ABSA) methods that rely heavily on text generation, proposing a novel "decide rather than generate" paradigm. Methodologically, using a frozen model named Jev, ABSA is decomposed into type decision and scoring tasks, integrating rubric-based scoring, label probabilities, and yes/no judgment mechanisms. Annotation alignment is achieved solely by fitting linear coefficients on a CPU, requiring neither text generation nor backbone fine-tuning. Experimental results demonstrate that this approach attains an RMSE as low as 1.0645 across ten corpora in regression settings, while its F1 scores for triplet and quadruplet extraction surpass those of large language model baselines such as Llama-3.3-70B. These findings validate the effectiveness of supervised calibration and compositional evidence in advancing efficient and accurate ABSA.
📝 Abstract
Aspect-based sentiment analysis (ABSA) has largely turned to text generation. We show that competitive dimensional ABSA does not need it. Using Jev, a frozen model that answers typed questions with rubric scores, label probabilities, and yes/no judgments, we decompose all three tasks of SemEval-2026 Task III Track A into such decisions and align them with the annotation scheme through 488 coefficients fitted on CPU, with no text generation and no backbone tuning. On valence-arousal regression over ten corpora in six languages, the system reaches 1.0645 RMSE, the lowest aggregate error of any participating system. On triplet and quadruplet extraction, it reaches 52.09 and 44.06 continuous F1, above fine-tuned Llama-3.3-70B and GPT-OSS-120B baselines. Analyses and ablations show where the accuracy comes from: supervised calibration roughly halves the raw regression error, exact valence-arousal would add only 4.5 F1 to extraction, and the learned combination of span-boundary evidence, not any single signal, carries the extraction systems.