llm report generation

Designs and implements systems and pipelines that use large language models to produce structured, constrained summaries and human‑readable discrepancy reports. Ensures outputs obey factual constraints (e.g., reporting deltas), augments deterministic deltas with explanatory text, and formats content to support expert interpretation and review.

llmreportgeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the lack of a systematic review on the application of large language models (LLMs) in software engineering documentation and modeling tasks. Through a comprehensive literature survey, it establishes a multi-dimensional taxonomy that categorizes existing research by task type, offering an in-depth analysis of key technical approaches—including prompt engineering, natural language understanding, and structured language processing. The work further synthesizes the distribution of tasks, evaluation metrics, human assessment methodologies, and commonly used datasets across major conferences in the field. By systematically mapping the research landscape and identifying prevailing technical trends, this paper provides a thorough reference and strategic guidance for future investigations at the intersection of LLMs and software engineering.

Generative AILarge Language ModelsSoftware Documentation

This work addresses the challenges of applying large language models (LLMs) in modeling and simulation (M&S), where suboptimal prompt design, improper hyperparameter configuration, or inadequate data handling often lead to performance degradation, information loss, and non-deterministic behavior. For the first time, this study systematically identifies latent pitfalls specific to LLM deployment in M&S and proposes a principled framework centered on rigorous design and empirical evaluation. The framework encompasses key techniques including prompt engineering, retrieval-augmented generation (RAG), low-rank adaptation (LoRA), temperature control, and context management. By offering a structured set of practical guidelines, this research enables practitioners to critically assess the suitability and implementation strategies of LLMs in M&S contexts, thereby substantially enhancing their effectiveness and reliability.

Hyper-parameter TuningLarge Language ModelsModeling and Simulation

This study addresses the widespread lack of adherence to reporting standards such as RAT-RS in agent-based modelling (ABM) research, largely due to the time-consuming and undervalued nature of manual reporting. It presents the first systematic evaluation of the feasibility of using large language models (LLMs) to automate the generation of RAT-RS–compliant content. The authors propose a practical framework based on supervised information extraction and introduce heuristic rules to delineate model reliability and the boundaries requiring human intervention. Experimental comparisons across four LLMs demonstrate that these models produce coherent and reliable outputs for descriptive tasks, significantly enhancing reporting quality and consistency. However, they exhibit notable limitations in explanatory and evaluative tasks, underscoring the continued need for human oversight in more interpretive aspects of ABM reporting.

Adoption BarrierAgent-Based ModellingDocumentation Standards

Evaluating the Effectiveness of Large Language Models in Automated News Article Summarization

Feb 24, 2025
LR
Lionel Richy Panlap Houamegni
🏛️ University of Applied Sciences Ruhr West

This study presents the first systematic evaluation of large language models (LLMs) for news summarization in supply chain risk analysis. We address three key challenges: difficulty in identifying compliance and operational risks from heterogeneous news sources, low information density, and high redundancy in generated summaries. To this end, we propose an automated summarization framework integrating few-shot prompting, multi-source aggregation, and structured generation. We further introduce a three-dimensional evaluation metric encompassing readability, redundancy suppression, and risk identification accuracy. Experimental results show that Few-Shot GPT-4o mini significantly outperforms mainstream LLMs—achieving +23.6% higher risk identification accuracy, +18.4% improved readability, and −31.2% reduced redundancy. A user study confirms that summaries generated by our method enhance enterprise risk response efficiency by 40%, validating its practical effectiveness in real-world business scenarios.

Assessing LLMs in automated news summarization.Evaluating LLMs' effectiveness in readability and risk detection.Focusing on supply chain risk analysis.

SciDaSynth: Interactive Structured Knowledge Extraction and Synthesis from Scientific Literature with Large Language Model

Apr 21, 2024
XW
Xingbo Wang
🏛️ Weill Cornell Medicine | Cornell University | Hong Kong University of Science and Technology

Scientific literature is inherently multimodal, heterogeneous, and unstructured, posing significant challenges for existing knowledge extraction systems in achieving cross-document consistency and dynamic adaptation to user intent. To address this, we propose the first LLM-driven interactive knowledge structuring paradigm, integrating prompt engineering, structured output control, conversational state management, and multi-granularity visual exploration. This enables researchers to automatically generate structured tables via natural language queries while collaboratively verifying and iteratively refining outputs. Our approach overcomes key limitations of conventional automated systems: it maintains high accuracy and coverage while reducing manual correction effort by over 40%. Empirical evaluation demonstrates substantial improvements in the efficiency of constructing high-quality scientific knowledge bases, offering a novel paradigm for domain-specific knowledge graph construction and reproducible research.

Building scalable interactive systems for literature-based synthesisExtracting structured knowledge from multimodal scientific literatureProcessing inconsistent information across diverse research papers

Latest Papers

What's happening recently
View more

This study addresses the current lack of interdisciplinary understanding regarding the integration pathways, efficacy boundaries, and systemic risks of large language models (LLMs) across natural sciences, social sciences, and humanities. Through a systematic literature review and illustrative case analyses, it critically evaluates the deployment of LLMs throughout the research lifecycle—including hypothesis generation, literature synthesis, data analysis, and scholarly writing. The work identifies ten previously underappreciated systemic risks, such as diminished researcher autonomy, AI-induced confirmation bias, ambiguous authorship, and inequitable access to technology. It further demonstrates how LLMs, while enhancing efficiency, simultaneously introduce challenges like hallucination, irreproducibility, data bias, and model opacity. To guide responsible adoption, the study proposes an interdisciplinary governance framework and a roadmap for explainable AI research in scholarly contexts.

AI ethicsinterdisciplinary integrationLarge Language Models

This work addresses the inefficiency and unreliability of directly deploying large language models (LLMs) or their distilled variants for enterprise tasks, which are typically deterministic, structured, and heavily reliant on domain-specific knowledge under strict constraints of cost, latency, and reliability. To overcome these limitations, the authors propose a modular AI architecture that confines LLMs to structured information extraction while offloading knowledge storage and reasoning to dedicated knowledge bases and symbolic systems. This design circumvents the bottlenecks of monolithic models in terms of interpretability, reliability, and maintainability. Theoretical analysis and system implementation demonstrate that the proposed architecture offers a more efficient, transparent, and sustainable alternative to end-to-end LLM approaches, providing a scalable AI solution tailored for enterprise applications.

deterministic workflowsenterprise tasksknowledge dependency

This work addresses the limitations of existing software documentation, which often suffers from poor consistency, weak relevance, and unclear expression, while purely automated generation methods struggle with insufficient reliability and controllability. To overcome these challenges, the authors propose a human-AI collaborative approach that integrates empirically derived quality guidelines with large language model (LLM) assistance through a generate–evaluate–iterate workflow. This method preserves developers’ domain control while enhancing explanation quality by embedding experience-driven quality criteria directly into the LLM-assisted writing process, enabling controllable, efficient, and high-quality output. Preliminary experiments demonstrate a 24.4% average improvement in authoring efficiency with tool support, and user studies reveal significantly higher satisfaction with the generated explanations compared to purely manual writing (p = 0.003, effect size = 0.86).

explanation qualityLLM-generated textnatural-language explanations

This work addresses the prevailing limitation in large language model (LLM) development, wherein human values are typically incorporated only post-training, lacking systematic integration across the model’s entire lifecycle. To bridge this gap, the paper introduces the Human-Centric Large Language Model (HCLLM) framework, which for the first time deeply integrates natural language processing, human-computer interaction, and responsible AI methodologies throughout all stages—from system design and data collection to training, evaluation, and deployment. The framework harmonizes ethical, economic, and technical objectives, offering developers actionable, principle-based guidance. Its forward-looking applicability and practical utility are demonstrated through a case study situated in future workplace scenarios, thereby advancing LLM development toward a genuinely human-centered paradigm.

Ethical AIHuman-Centered AIHuman-Computer Interaction

This study addresses the heavy reliance of large language models on prompt design for code summarization tasks and the absence of systematic comparisons and unified evaluation standards across diverse prompting strategies. Through a comprehensive literature review, it integrates and categorizes mainstream approaches—including few-shot prompting, chain-of-thought reasoning, retrieval-augmented generation, and zero-shot learning—and analyzes their effectiveness across different models and scenarios. The work highlights the limitations of current evaluations that overly depend on surface-level overlap metrics, delineates the conditions under which each prompting paradigm performs best, and proposes a unified evaluation framework to guide future research and practical applications in this domain.

code summarizationlarge language modelsprompt engineering

Hot Scholars

HS

Holli Sargeant

PhD in Law Candidate, University of Cambridge
LawArtificial IntelligenceMachine LearningEthics
FS

Felix Steffek

Professor of Law, University of Cambridge
Law
CZ

Chengzhi Zhang

Nanjing University of Science and Technology
Text MiningNatural Language ProcessingScience of Science