zero-shot stance classification

Designs and builds LLM-based zero-shot classifiers and prompt/inference pipelines that assign rhetorical-stance labels (e.g., support/oppose, hedging/certainty) to textual units without task-specific supervised training. Analyzes label schemas, prompting strategies and evaluation procedures for sentence- or utterance-level and longer-text inputs (e.g., transcripts), including robustness to spurious lexicon correlations and other failure modes.

zero-shotstanceclassification

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Large language models (LLMs) face challenges in political science text classification under few-shot settings—namely, heavy reliance on manual prompt engineering, static in-context example selection, and opaque, uninterpretable predictions. Method: This paper proposes a three-stage, fine-tuning-free framework: (1) task-driven automatic structured prompt generation; (2) query-aware dynamic KNN retrieval of semantically similar examples; and (3) multi-path output aggregation via weighted consensus, emulating collaborative coding by multiple annotators. It innovatively integrates meta-prompt engineering with consensus-based ensemble mechanisms. Contribution/Results: The open-source toolkit PoliPrompt achieves an average 12.7% accuracy gain over human-crafted prompts with fixed examples across sentiment analysis, stance detection, and campaign ad tone classification. It requires no training, operates out-of-the-box, and substantially reduces manual tuning effort while enhancing interpretability and robustness.

Dynamically selecting relevant exemplars for few-shot learningEnhancing classification accuracy without task-specific retrainingOptimizing prompts for political science text classification

An Investigation of Prompt Variations for Zero-shot LLM-based Rankers

Jun 20, 2024
SS
Shuoqi Sun
🏛️ RMIT University | CSIRO | The University of Queensland

This study investigates the extent to which prompt design influences zero-shot large language model (LLM) ranking performance—and whether its impact surpasses that of ranking algorithm selection and underlying model architecture (e.g., GPT-3.5, FLAN-T5). Through large-scale, controlled ablation experiments, we systematically disentangle the effects of prompt components (e.g., role specification phrasing), ranking paradigms (pairwise vs. listwise), and model architecture. Our key finding—quantified for the first time—is that fine-grained prompt design significantly outweighs algorithmic differences across multiple evaluation scenarios; moreover, algorithmic advantages substantially diminish under prompt perturbations. This work challenges the prevailing consensus that attributes ranking performance primarily to algorithm or model choice, and establishes a new, reproducible, and attribution-aware benchmark for LLM-based ranking research.

Game RankingLLM-based ModelsPrompt Influence

What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Jun 18, 2024
FE
Federico Errica
🏛️ NEC Italia | NEC Laboratories Europe

Large language models (LLMs) exhibit unstable outputs in software applications when prompts undergo minor rephrasings, hindering reliable deployment. Method: This paper introduces two label-free, quantifiable metrics—sensitivity (cross-prompt prediction variance) and consistency (prediction stability across semantically equivalent prompts)—to formally decouple and evaluate LLM robustness to prompt perturbations. Leveraging text classification tasks, we conduct systematic, multi-round prompt rewriting and statistical analysis of prediction distributions. Contribution/Results: Empirical evaluation reveals that mainstream LLMs consistently exhibit high sensitivity and low consistency, exposing a critical robustness gap. Our framework provides a reproducible, ground-truth-label-free diagnostic paradigm for prompt engineering, enabling joint optimization of accuracy and robustness. This work establishes the first formal, measurement-driven approach to assessing and improving LLM resilience against prompt variations.

Large Language ModelsPrediction InstabilitySemantic Sensitivity

The Effectiveness of LLMs as Annotators: A Comparative Overview and Empirical Analysis of Direct Representation

May 02, 2024
MP
Maja Pavlovic
🏛️ Queen Mary University of London | University of Utrecht

This work systematically evaluates the efficacy and limitations of large language models (LLMs) as annotators for subjective tasks. We survey 12 prior studies and empirically compare opinion distribution alignment between GPT-series models and human annotators across four subjective datasets—introducing, for the first time, a novel evaluation paradigm centered on *perspective diversity alignment*. Results reveal substantial distributional biases in LLMs, including underrepresentation of minority viewpoints, prompt sensitivity, English-language bias, and embedded societal prejudices; most existing annotation methods overlook such distributional discrepancies, and only a few strategies effectively capture opinion diversity. Our findings expose critical reliability risks in deploying LLMs for subjective annotation and establish a reproducible statistical framework—comprising quantitative metrics and methodological guidelines—for assessing annotation quality in subjective NLP tasks.

Addressing limitations like bias and prompt sensitivityComparing human and GPT-generated opinion distributionsEvaluating LLMs' effectiveness in data annotation tasks

Latest Papers

What's happening recently
View more

This study addresses the challenge of fairly comparing existing stance detection methods, which has been hindered by inconsistent data splits, base models, and evaluation protocols. For the first time, it systematically evaluates five method categories—three prompt-based reasoning approaches and two multi-agent debate frameworks—across 14 subtasks on four benchmark datasets, leveraging 15 large language models from six families with parameter scales ranging from 7B to over 72B. Results show that prompt-based methods consistently outperform multi-agent approaches while reducing API calls by 7–12×. Model scale exerts a stronger influence on performance than method choice, with gains saturating around 32B parameters. Notably, reasoning-enhancement strategies yield no significant improvement in stance detection accuracy. This work establishes a reproducible benchmark and offers practical guidance for method selection in stance detection research.

large language modelsmulti-agent methodsprompting methods

This study addresses the problem of efficient and accurate user intent classification in large language models to facilitate downstream domain-specific model routing. For the first time, it systematically compares training-free strategies—such as lightweight methods based on internal representation statistics—with training-based approaches, including linear probes and MLP classifiers, evaluating their performance across varying task difficulty, mixed-intent prompts, and adversarial inputs. The results reveal that both paradigms achieve performance saturation on simple tasks; however, training-based methods excel in fine-grained classification (e.g., distinguishing Java from Python), whereas training-free methods demonstrate superior robustness to mixed and adversarial prompts. These findings highlight fundamental differences between the two approaches in terms of accuracy, robustness, and failure modes.

intent classificationLarge Language Modelsrobustness

This work addresses the limitations of existing AI-generated text detection methods, which rely on author labels, struggle with zero-shot generalization, and are vulnerable to adversarial attacks. The authors propose an unsupervised style representation learning framework that decouples and captures non-semantic stylistic features by freezing a semantic encoder and using only a style encoder to inversely reconstruct human-like text from machine-generated paraphrases. This approach enables style modeling without any author labels, supporting both zero-shot and few-shot detection settings. Experimental results demonstrate that the method achieves few-shot performance on par with or superior to current state-of-the-art approaches across multiple benchmarks, while its zero-shot effectiveness closely matches that of fully supervised models. Furthermore, it exhibits strong generalization capabilities in author verification and fine-grained style discrimination tasks.

AI-text detectionauthorship labelsstyle representation

This work addresses the limitations of existing zero-shot detectors, which often fail to robustly identify machine-generated text due to their neglect of the generative mechanisms inherent in large language models. To overcome this, we propose EchoPrompt, a training-free detection method that introduces implicit prompt recovery into the zero-shot detection framework for the first time. EchoPrompt activates latent dependencies embedded in generated text through a unified prefix and quantifies a text’s reliance on its underlying prompt by measuring the likelihood gain discrepancy between an instruction-tuned model and its base counterpart. Extensive experiments demonstrate that EchoPrompt achieves state-of-the-art performance across multiple challenging scenarios while maintaining strong robustness, significantly outperforming current zero-shot detectors.

latent prompt dependencyLLM-generated text detectionmachine-generated text

This study addresses the underperformance of large language models on natural language inference tasks for low-resource African languages—such as Swahili, Yoruba, and Hausa—particularly in zero-shot prompting scenarios where systematic investigation has been lacking. The authors propose a language-aware prompt structure and conduct a systematic evaluation of five prompting strategies (Baseline, Script-Aware, Language-Specific, Contrastive, and NL-STP) using Llama3.2-3B and Gemma3-4B. Experimental results demonstrate that the proposed approach significantly improves overall accuracy and class balance across multilingual and multi-model settings. It not only outperforms existing zero-shot prompting methods but also surpasses strong baselines such as few-shot learning and chain-of-thought prompting, effectively mitigating the issue of class collapse.

African languagesclass imbalancelow-resource languages

Hot Scholars

ML

Mengyuan Li

University of Southern California
Hardware SecurityTrusted Execution EnvironmentCloud computing
HX

Haoyan Xu

University of Southern California
Machine Learning
JG

Jie Gui

Southeast University, China
Pattern Recognition and Machine LearningArtificial IntelligenceData MiningDeep Learning
YD

Yushun Dong

Assistant Professor, Department of Computer Science, Florida State University
AI SecurityAI IntegrityGraph Machine LearningLLMs
CH

Chaolei Han

Southeast University
Computer VisionVideo AnalysisAction Detection