Score
Designs and implements methods that prompt and parse large language model outputs to extract interpretable linguistic, lexical, syntactic, semantic, and discourse features and convert them into fixed-dimensional vectors. Covers prompted linguistic feature extraction, LLM output parsing, and LLM-augmented linguistics for building analyzers or feature pipelines used in downstream modeling or evaluation.
This study systematically investigates the current state and evolutionary trajectory of large language models (LLMs) in natural language processing (NLP), addressing three core questions: (1) How are LLMs applied to NLP tasks? (2) Have classical NLP tasks been fundamentally resolved? (3) What are the future research directions? To this end, we propose the first unified classification framework for LLM applications in NLP, dichotomizing methodologies into parameter-freezing and parameter-finetuning paradigms. We conduct a comprehensive empirical assessment of LLMs’ coverage and performance bottlenecks across canonical NLP tasks. Furthermore, we synthesize emerging frontiers—including reasoning augmentation and multimodal integration—as well as persistent challenges such as interpretability and long-range dependency modeling. Grounded in a rigorous literature review and technical evolution analysis, we construct the first structured, holistic research landscape map of LLM-NLP, offering both theoretical insight and practical guidance for the community.
High-dimensional, opaque text representations (e.g., embeddings, bag-of-words) hinder interpretable rule learning. To address this, we propose a novel framework leveraging LLaMA-2 via prompt engineering to extract a compact set of 62 low-dimensional, semantically transparent, and human-understandable features—such as “methodological rigor” and “novelty”—directly from scientific literature. This marks the first use of large language models to generate *rule-ready*, inherently interpretable textual features. Our pipeline integrates statistical hypothesis testing, supervised classification (Logistic Regression and Random Forest), and rule extraction algorithms to derive actionable, domain-generalizable decision rules. Evaluated on CORD-19 (binary classification) and M17+ (5-class classification), our approach achieves performance comparable to 768-dimensional SciBERT while substantially enhancing model transparency, interpretability, and practical deployability.
This study systematically investigates core challenges impeding large language model (LLM) industrial deployment, identifying 12 representative bottlenecks across four critical dimensions: data scarcity, inefficient inference, complex deployment, and inaccurate evaluation. Method: We employ a mixed-methods approach—structured interviews with frontline practitioners, a research-question-driven review of 68 industrial practice papers, and qualitative content analysis. Contribution/Results: We propose the first “industry-perspective-driven” taxonomy for LLM deployment challenges; establish a dynamically updated GitHub knowledge repository of industrial LLM literature; and deliver an actionable, lifecycle-spanning optimization roadmap. The framework has been adopted by multiple enterprises and serves as a key reference benchmark for industrial LLM adoption.
Current LLM application development lacks systematic, practice-informed guidelines, leading to a growing gap between academic research and industrial engineering. Method: Drawing on transcribed texts from 189 real-world developer practice videos (2022–2024), we integrate BERTopic-based automated topic modeling with iterative human refinement to construct the first empirically grounded, production-oriented thematic map of LLM application development. Contribution/Results: The map identifies eight core themes—including design & architecture, model enhancement, infrastructure, and ethical risk—spanning 20 key issues. Design & Architecture emerges as the most densely populated theme, with RAG at its architectural center; prompt engineering, fine-tuning, deployment toolchains, and AI ethics are recurrent high-frequency concerns. Critically, the map exposes significant lags in academic research relative to industrial practice and delivers an actionable, empirically validated priority framework—thereby bridging a critical empirical gap in the LLM engineering knowledge base.
Existing feature engineering approaches suffer from three fundamental limitations: poor interpretability, weak generalizability, and inflexible strategies—hindering practical deployment across diverse scenarios. To address these challenges, this paper proposes the first large language model (LLM)-driven dynamic adaptive feature generation paradigm. Our method integrates task-aware prompting with semantic modeling of the feature space, enabling real-time, interpretable, and controllable feature generation tailored to both data characteristics and task requirements. It ensures cross-modal and cross-task generality while maintaining full transparency in the feature generation process. Extensive experiments on multiple structured and unstructured data tasks demonstrate that features generated by our approach improve feature quality by 23.6% and boost downstream model performance by an average of 11.4%, significantly outperforming conventional automated feature engineering methods.
A systematic, data-driven understanding of the relationship between large language model (LLM) architectural configurations and performance remains lacking. Method: This project introduces the first large-scale, open-source LLM architecture–performance benchmark dataset and proposes a data-driven quantification framework integrating multi-benchmark evaluation, statistical modeling, and mechanistic interpretability techniques to perform attribution analysis on key architectural parameters—including number of layers, attention heads, and feed-forward network dimensions. Contribution/Results: Experiments reveal significant, nonlinear causal effects of specific architectural choices on downstream task performance. The project releases a fully reproducible dataset and analytical toolkit, uncovering empirical patterns in architectural evolution. These resources enable accurate performance prediction and efficient model design, establishing a novel paradigm for LLM interpretability and controllable optimization.
This study investigates whether attention heads and input embeddings in large language models (LLMs) genuinely encode human-interpretable semantic information. Addressing concerns that prevailing interpretability methods—such as those based on attention weights or embedding analyses—may be confounded by data artifacts or methodological biases, the authors employ token-level relational structural probes and map human-interpretable attributes onto the embedding space to systematically evaluate the validity of these approaches across multiple Transformer layers. Their findings reveal that both widely adopted explanation techniques fail to reliably reflect the model’s true semantic capabilities, thereby challenging the foundational assumptions underlying current claims about LLMs’ “understanding.” This work carries significant implications for deploying LLMs in edge and distributed computing environments, where interpretability and reliability are critical.
Conventional text preprocessing techniques—such as stopword removal, lemmatization, and stemming—rely heavily on language-specific linguistic rules and ignore contextual information, limiting their generalizability across multilingual settings. Method: This paper pioneers a systematic investigation of large language models (LLMs) as context-aware, universal preprocessors. Leveraging prompt engineering, we uniformly perform the three preprocessing tasks across six European languages without language-specific annotations or handcrafted rules. Contribution/Results: Experiments show LLMs achieve 97%, 82%, and 74% accuracy on stopword removal, lemmatization, and stemming, respectively. Downstream text classification models fed with LLM-preprocessed inputs attain up to a 6-percentage-point improvement in F1 score. This work demonstrates the feasibility and effectiveness of LLM-driven, end-to-end, context-sensitive, and multilingual-compatible text preprocessing—establishing a novel paradigm that reduces reliance on manual linguistic rules and enhances preprocessing robustness.
This work addresses the limited accessibility of existing interpretability tools for language models, which are often too complex for non-expert users to effectively utilize. To bridge this gap, the authors introduce ELIA, an interactive web application that uniquely integrates multiple mechanistic analysis techniques—including Attribution Analysis, Function Vector Analysis, and Circuit Tracing—and innovatively incorporates a vision–language model to automatically generate natural language explanations for complex visualizations. User studies demonstrate that ELIA significantly lowers the barrier to understanding model behavior, with AI-generated explanations effectively mitigating knowledge disparities. Notably, users’ comprehension outcomes show no significant correlation with their prior experience using large language models, underscoring the system’s broad usability and general effectiveness across diverse user backgrounds.