Score
Designs and builds LLM-based zero-shot classifiers and prompt/inference pipelines that assign rhetorical-stance labels (e.g., support/oppose, hedging/certainty) to textual units without task-specific supervised training. Analyzes label schemas, prompting strategies and evaluation procedures for sentence- or utterance-level and longer-text inputs (e.g., transcripts), including robustness to spurious lexicon correlations and other failure modes.
Existing research lacks a systematic survey of large language models (LLMs) for stance detection. This paper introduces the first three-dimensional taxonomy—spanning learning paradigms, data modalities, and target relations—specifically designed for LLM-based stance detection. We comprehensively review task formalizations, methodological advances—including supervised, few-shot, and zero-shot learning; multimodal fusion; prompt engineering; and instruction tuning—as well as emerging challenges such as implicit stance modeling and mitigation of cultural bias. We conduct unified benchmarking of mainstream LLMs across 12 standard datasets, empirically assessing their efficacy and limitations in real-world applications like fake news detection and sentiment-aware舆情 analysis. Our findings yield a principled roadmap and practical guidelines for both theoretical advancement and industrial deployment of stance detection systems.
Large language models (LLMs) face challenges in political science text classification under few-shot settings—namely, heavy reliance on manual prompt engineering, static in-context example selection, and opaque, uninterpretable predictions. Method: This paper proposes a three-stage, fine-tuning-free framework: (1) task-driven automatic structured prompt generation; (2) query-aware dynamic KNN retrieval of semantically similar examples; and (3) multi-path output aggregation via weighted consensus, emulating collaborative coding by multiple annotators. It innovatively integrates meta-prompt engineering with consensus-based ensemble mechanisms. Contribution/Results: The open-source toolkit PoliPrompt achieves an average 12.7% accuracy gain over human-crafted prompts with fixed examples across sentiment analysis, stance detection, and campaign ad tone classification. It requires no training, operates out-of-the-box, and substantially reduces manual tuning effort while enhancing interpretability and robustness.
This study investigates the extent to which prompt design influences zero-shot large language model (LLM) ranking performance—and whether its impact surpasses that of ranking algorithm selection and underlying model architecture (e.g., GPT-3.5, FLAN-T5). Through large-scale, controlled ablation experiments, we systematically disentangle the effects of prompt components (e.g., role specification phrasing), ranking paradigms (pairwise vs. listwise), and model architecture. Our key finding—quantified for the first time—is that fine-grained prompt design significantly outweighs algorithmic differences across multiple evaluation scenarios; moreover, algorithmic advantages substantially diminish under prompt perturbations. This work challenges the prevailing consensus that attributes ranking performance primarily to algorithm or model choice, and establishes a new, reproducible, and attribution-aware benchmark for LLM-based ranking research.
Large language models (LLMs) exhibit unstable outputs in software applications when prompts undergo minor rephrasings, hindering reliable deployment. Method: This paper introduces two label-free, quantifiable metrics—sensitivity (cross-prompt prediction variance) and consistency (prediction stability across semantically equivalent prompts)—to formally decouple and evaluate LLM robustness to prompt perturbations. Leveraging text classification tasks, we conduct systematic, multi-round prompt rewriting and statistical analysis of prediction distributions. Contribution/Results: Empirical evaluation reveals that mainstream LLMs consistently exhibit high sensitivity and low consistency, exposing a critical robustness gap. Our framework provides a reproducible, ground-truth-label-free diagnostic paradigm for prompt engineering, enabling joint optimization of accuracy and robustness. This work establishes the first formal, measurement-driven approach to assessing and improving LLM resilience against prompt variations.
This work systematically evaluates the efficacy and limitations of large language models (LLMs) as annotators for subjective tasks. We survey 12 prior studies and empirically compare opinion distribution alignment between GPT-series models and human annotators across four subjective datasets—introducing, for the first time, a novel evaluation paradigm centered on *perspective diversity alignment*. Results reveal substantial distributional biases in LLMs, including underrepresentation of minority viewpoints, prompt sensitivity, English-language bias, and embedded societal prejudices; most existing annotation methods overlook such distributional discrepancies, and only a few strategies effectively capture opinion diversity. Our findings expose critical reliability risks in deploying LLMs for subjective annotation and establish a reproducible statistical framework—comprising quantitative metrics and methodological guidelines—for assessing annotation quality in subjective NLP tasks.
This study addresses the challenge of fairly comparing existing stance detection methods, which has been hindered by inconsistent data splits, base models, and evaluation protocols. For the first time, it systematically evaluates five method categories—three prompt-based reasoning approaches and two multi-agent debate frameworks—across 14 subtasks on four benchmark datasets, leveraging 15 large language models from six families with parameter scales ranging from 7B to over 72B. Results show that prompt-based methods consistently outperform multi-agent approaches while reducing API calls by 7–12×. Model scale exerts a stronger influence on performance than method choice, with gains saturating around 32B parameters. Notably, reasoning-enhancement strategies yield no significant improvement in stance detection accuracy. This work establishes a reproducible benchmark and offers practical guidance for method selection in stance detection research.
This study addresses the problem of efficient and accurate user intent classification in large language models to facilitate downstream domain-specific model routing. For the first time, it systematically compares training-free strategies—such as lightweight methods based on internal representation statistics—with training-based approaches, including linear probes and MLP classifiers, evaluating their performance across varying task difficulty, mixed-intent prompts, and adversarial inputs. The results reveal that both paradigms achieve performance saturation on simple tasks; however, training-based methods excel in fine-grained classification (e.g., distinguishing Java from Python), whereas training-free methods demonstrate superior robustness to mixed and adversarial prompts. These findings highlight fundamental differences between the two approaches in terms of accuracy, robustness, and failure modes.
This work addresses the limitations of existing AI-generated text detection methods, which rely on author labels, struggle with zero-shot generalization, and are vulnerable to adversarial attacks. The authors propose an unsupervised style representation learning framework that decouples and captures non-semantic stylistic features by freezing a semantic encoder and using only a style encoder to inversely reconstruct human-like text from machine-generated paraphrases. This approach enables style modeling without any author labels, supporting both zero-shot and few-shot detection settings. Experimental results demonstrate that the method achieves few-shot performance on par with or superior to current state-of-the-art approaches across multiple benchmarks, while its zero-shot effectiveness closely matches that of fully supervised models. Furthermore, it exhibits strong generalization capabilities in author verification and fine-grained style discrimination tasks.
This work addresses the limitations of existing zero-shot detectors, which often fail to robustly identify machine-generated text due to their neglect of the generative mechanisms inherent in large language models. To overcome this, we propose EchoPrompt, a training-free detection method that introduces implicit prompt recovery into the zero-shot detection framework for the first time. EchoPrompt activates latent dependencies embedded in generated text through a unified prefix and quantifies a text’s reliance on its underlying prompt by measuring the likelihood gain discrepancy between an instruction-tuned model and its base counterpart. Extensive experiments demonstrate that EchoPrompt achieves state-of-the-art performance across multiple challenging scenarios while maintaining strong robustness, significantly outperforming current zero-shot detectors.
This study addresses the underperformance of large language models on natural language inference tasks for low-resource African languages—such as Swahili, Yoruba, and Hausa—particularly in zero-shot prompting scenarios where systematic investigation has been lacking. The authors propose a language-aware prompt structure and conduct a systematic evaluation of five prompting strategies (Baseline, Script-Aware, Language-Specific, Contrastive, and NL-STP) using Llama3.2-3B and Gemma3-4B. Experimental results demonstrate that the proposed approach significantly improves overall accuracy and class balance across multilingual and multi-model settings. It not only outperforms existing zero-shot prompting methods but also surpasses strong baselines such as few-shot learning and chain-of-thought prompting, effectively mitigating the issue of class collapse.