llm pattern extraction

Design and implement methods and tools that prompt and analyze large language model (LLM) outputs to induce, extract, and represent reusable abstract patterns—linguistic, logical, or structural—from model behavior. Build detection and rule-based extraction pipelines and pattern representations for instances, and analyze how patterns form, are induced by prompts, and generalize across LLM responses.

llmpatternextraction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.18
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$198K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Large Language Models (LLMs) for Source Code Analysis: applications, models and datasets

Mar 21, 2025
HJ
Hamed Jelodar
🏛️ Canadian Institute for Cybersecurity | University of New Brunswick

This study addresses fragmented task definitions, inconsistent evaluation protocols, and dataset biases hindering systematic progress in applying large language models (LLMs) to source code analysis. We conduct a systematic review of 217 publications from 2019–2024 and propose the first three-dimensional knowledge graph—spanning *tasks*, *models*, and *datasets*—specifically for code analysis, covering 12 analytical task categories, 8 mainstream LLM variants, and 19 core benchmarks. Methodologically, we introduce a reproducible, standardized evaluation framework grounded in empirical analysis. Through bibliometric analysis and architectural feature modeling, we identify key evolutionary trends and persistent bottlenecks—including limited generalizability and inadequate contextual modeling. Our contributions provide a structured conceptual foundation and methodological guidance for advancing LLM-driven code intelligence research.

Exploring LLMs' role in source code analysis tasksIdentifying challenges in LLM-based code analysis workflowsInvestigating models and datasets for code analysis

Large Language Models for Constructing and Optimizing Machine Learning Workflows: A Survey

Nov 11, 2024
YG
Yang Gu
🏛️ Shanghai Jiao Tong University | Stanford University

This work addresses the intelligent evolution of AutoML by investigating how large language models (LLMs) can optimize the end-to-end machine learning (ML) pipeline. Method: We propose a four-dimensional capability framework—language understanding, reasoning, interaction, and generation—to systematically characterize LLM-driven ML workflow paradigms; integrate prompt engineering, instruction tuning, chain-of-thought reasoning, tool-augmented LLMs, and multi-stage orchestration; and synthesize over 50 state-of-the-art techniques. Contribution/Results: Empirical evaluation demonstrates that LLMs substantially lower modeling barriers, enhance cross-task generalization, and improve human-AI collaboration efficiency—achieving semantic modeling and human-in-the-loop breakthroughs in data preprocessing, feature engineering, model selection, hyperparameter optimization, and workflow evaluation. However, critical challenges remain regarding reliability, interpretability, and computational overhead.

Automated Machine LearningData Processing and Model SelectionLarge Language Models

On the Design and Analysis of LLM-Based Algorithms

Jul 20, 2024
YC
Yanxi Chen
🏛️ Alibaba Group

Current LLM-based algorithm design relies heavily on empirical trial-and-error, lacking formal theoretical foundations to systematically analyze how critical design choices—such as task decomposition strategies and prompt engineering—affect accuracy and computational efficiency. Method: We propose the first formal analytical framework for LLM-invocation algorithms, modeling LLM subroutines as a computational graph, establishing structured principles for task decomposition, and introducing an error propagation model that enables provable analysis of both accuracy and computational complexity. Contribution/Results: Empirically validated across parallel, hierarchical, and recursive paradigms, our framework explains observed empirical phenomena, guides prompt design and granularity selection, predicts performance bottlenecks, and inspires novel robust algorithm designs. The implementation is publicly available.

Analyzing impact of design choices on algorithm accuracyDeveloping analytical framework for LLM-based algorithm designProviding systematic approach for LLM algorithm optimization

Themes of Building LLM-based Applications for Production: A Practitioner's View

Nov 13, 2024
AM
Alina Mailach
🏛️ ScaDS.AI Dresden/Leipzig | Leipzig University

Current LLM application development lacks systematic, practice-informed guidelines, leading to a growing gap between academic research and industrial engineering. Method: Drawing on transcribed texts from 189 real-world developer practice videos (2022–2024), we integrate BERTopic-based automated topic modeling with iterative human refinement to construct the first empirically grounded, production-oriented thematic map of LLM application development. Contribution/Results: The map identifies eight core themes—including design & architecture, model enhancement, infrastructure, and ethical risk—spanning 20 key issues. Design & Architecture emerges as the most densely populated theme, with RAG at its architectural center; prompt engineering, fine-tuning, deployment toolchains, and AI ethics are recurrent high-frequency concerns. Critically, the map exposes significant lags in academic research relative to industrial practice and delivers an actionable, empirically validated priority framework—thereby bridging a critical empirical gap in the LLM engineering knowledge base.

Analyze practitioner discussions on LLM deployment challengesIdentify key considerations for LLM-based system developmentProvide systematic overview of LLM application priorities

LLMs with Industrial Lens: Deciphering the Challenges and Prospects - A Survey

Feb 22, 2024
AU
Ashok Urlana
🏛️ TCS Research | IIIT Hyderabad

This study systematically investigates core challenges impeding large language model (LLM) industrial deployment, identifying 12 representative bottlenecks across four critical dimensions: data scarcity, inefficient inference, complex deployment, and inaccurate evaluation. Method: We employ a mixed-methods approach—structured interviews with frontline practitioners, a research-question-driven review of 68 industrial practice papers, and qualitative content analysis. Contribution/Results: We propose the first “industry-perspective-driven” taxonomy for LLM deployment challenges; establish a dynamically updated GitHub knowledge repository of industrial LLM literature; and deliver an actionable, lifecycle-spanning optimization roadmap. The framework has been adopted by multiple enterprises and serves as a key reference benchmark for industrial LLM adoption.

Exploring challenges in industrial LLM applicationsIdentifying opportunities for enhancing LLM utilizationSurveying industry practices and research on LLMs

Latest Papers

What's happening recently
View more

Existing approaches struggle to systematically compare the outputs of large language models under varying generation conditions. This work proposes a “visual fingerprint” framework that models model outputs as distributions over multidimensional linguistic choices—encompassing content, expression, and structure—and enables cross-condition comparison of generative behavior through an integrated natural language processing pipeline and distribution visualization techniques. For the first time, this method facilitates intuitive, distribution-level insights into stable behavioral patterns that persist across diverse settings yet remain undetectable via single-sample inspection or conventional aggregate metrics. The efficacy of the framework is demonstrated across four distinct application scenarios.

generation conditionslinguistic choicesLLM generation

This work addresses the challenges of applying large language models (LLMs) in modeling and simulation (M&S), where suboptimal prompt design, improper hyperparameter configuration, or inadequate data handling often lead to performance degradation, information loss, and non-deterministic behavior. For the first time, this study systematically identifies latent pitfalls specific to LLM deployment in M&S and proposes a principled framework centered on rigorous design and empirical evaluation. The framework encompasses key techniques including prompt engineering, retrieval-augmented generation (RAG), low-rank adaptation (LoRA), temperature control, and context management. By offering a structured set of practical guidelines, this research enables practitioners to critically assess the suitability and implementation strategies of LLMs in M&S contexts, thereby substantially enhancing their effectiveness and reliability.

Hyper-parameter TuningLarge Language ModelsModeling and Simulation

A Roadmap for Tamed Interactions with Large Language Models

Oct 28, 2025
VS
Vincenzo Scotti
🏛️ Karlsruhe Institute of Technology (KIT) | Ruhr Institute for Software Technology (paluno) | University of Duisburg-Essen

Large language models (LLMs) suffer from hallucination, unreliability, and uncontrolled behavior, hindering their trustworthy deployment in safety-critical workflows; existing reliability-enhancement tools are fragmented and lack a systematic framework. This paper introduces LSL (LLM Scripting Language), a domain-specific scripting language that embeds formal specifications, verifiable constraints, and explainability mechanisms directly into the LLM interaction process—enabling structured output constraints, programmable behavioral control, and decoupled execution governance. LSL unifies domain-specific language (DSL) design, formal verification, and runtime checking, significantly improving output reliability, consistency, and traceability. Experiments demonstrate that LSL effectively mitigates hallucination across diverse tasks, supports safe and controllable LLM integration, and establishes a novel interaction paradigm for trustworthy AI systems.

Addressing LLM unreliability and hallucination issuesDeveloping DSL to control LLM outputs and interactionsIntegrating verification and validation for trustworthy LLM applications

Toward Automated and Trustworthy Scientific Analysis and Visualization with LLM-Generated Code

Nov 26, 2025
AK
Apu Kumar Chakroborti
🏛️ Georgia State University

Domain scientists often lack sufficient programming expertise to conduct data analysis efficiently. This paper addresses the low reliability and poor trustworthiness of large language models (LLMs) in scientific code generation by introducing the first benchmark suite for Python-based data analysis and visualization grounded in real-world research tasks. We propose three synergistic strategies: data-aware prompt disambiguation, retrieval-augmented prompt optimization, and iterative error repair—integrated with retrieval-augmented generation (RAG) and automated execution validation. Experiments demonstrate substantial improvements in code executability and functional correctness. However, domain-context understanding remains a critical bottleneck. This work contributes both a reusable, realistic evaluation benchmark and a systematic technical framework for developing trustworthy AI-powered scientific tools.

LLMs generate code for scientific data analysis and visualizationStrategies improve code reliability but need further refinementTrustworthiness of LLM-generated code is limited without human intervention

To address the challenge of adapting large language models (LLMs) to proprietary industrial programming languages—such as ABB RAPID—in automation domains, this paper proposes a fine-tuning-free, few-shot prompting method enabling locally deployed LLMs to directly comprehend and modify RAPID programs. By eliminating reliance on large-scale annotated datasets or custom model training, the approach preserves data privacy and enhances deployment flexibility. Experimental evaluation demonstrates its effectiveness on elementary tasks including code repair and logical adaptation, substantially lowering the adoption barrier for LLMs in non-general-purpose industrial language settings. Key contributions include: (i) the first systematic investigation into LLM support for closed industrial languages like RAPID; (ii) a lightweight, secure, and plug-and-play prompting framework; and (iii) a low-overhead, highly controllable paradigm for AI-assisted programming tailored to high-sensitivity industrial environments.

Applying LLMs to industrial automation with specialized proprietary languagesEnabling on-premise solutions to protect sensitive industrial company dataExploring few-shot prompting for domain-specific languages without extensive training

Hot Scholars

PL

Peng Liang

School of Computer Science, Wuhan University
Software EngineeringSoftware ArchitectureEmpirical Software Engineering
SS

Shashi Shekhar

McKnight Distinguished University Professor of Computer Science, University of Minnesota
Spatial Big DataSpatial ComputingSpatial DatabasesSpatial Data Mining
DL

Dahua Lin

The Chinese University of Hong Kong
computer visionmachine learningprobabilistic inferencebayesian nonparametrics
SW

Su Wang

Beijing Institute of Technology
Motor ImageryEEG RecognitionNeural Network
ZW

Zhijun Wang

Institute of Physics, Chinese Academy of Sciences
Condensed Matter Physics