client-side llm summarization

Designs and builds client-side systems that run (often frozen) large language models locally to convert user interactions and on-device content into compact semantic summaries or vector representations. These systems encode typical interaction patterns and produce shareable semantic signals (embeddings or concise summaries) to support downstream tasks without transmitting raw user data.

client-sidellmsummarization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.37
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Small Models, Big Results: Achieving Superior Intent Extraction through Decomposition

Sep 15, 2025
DC
Danielle Cohen
🏛️ Google | Bar-Ilan University

Resource-constrained edge devices face challenges in accurately understanding user intent from UI interaction traces, while simultaneously ensuring privacy preservation and real-time responsiveness. Method: This paper proposes a two-stage decomposed architecture: (1) generating structured sequential summaries of interaction behaviors, followed by (2) lightweight intent inference based on these summaries. The approach integrates context-aggregated enhancement and task-adaptive fine-tuning to strengthen semantic modeling capabilities of small models. Contribution/Results: Experimental results demonstrate that, under identical privacy guarantees and low-latency constraints, the proposed method achieves higher intent recognition accuracy than state-of-the-art large multimodal language models. It establishes an efficient, privacy-aware, and real-time interaction understanding paradigm for on-device intelligent agents.

Enhancing intent understanding through decomposed structured summarizationImproving intent extraction accuracy in small on-device modelsOvercoming limitations of privacy-preserving low-latency models

Comparative Analysis of Large Language Models for the Machine-Assisted Resolution of User Intentions

Aug 29, 2025
JF
Justus Flerlage
🏛️ Technische Universität Berlin | logsight.ai GmbH

Cloud-hosted proprietary large language models (LLMs) pose significant challenges in user intent parsing—including privacy leakage, lack of user autonomy, and limited scalability. Method: This paper systematically evaluates the feasibility of open-weight LLMs as core components of a localized, intent-driven operating system. We design a lightweight intent parsing framework leveraging multiple open-source models (e.g., Llama-3, Qwen, Phi-3), integrating natural language understanding, workflow generation, and on-device inference. Empirical evaluation is conducted across multi-application collaborative tasks, with GPT-4 as a performance benchmark. Results: Several open-weight models achieve intent recognition accuracy comparable to closed-source counterparts under local deployment—while ensuring data privacy, enabling offline operation, and preserving user control. To our knowledge, this is the first study to empirically validate open LLMs as enablers of decentralized, language-first interaction paradigms. It establishes an open architectural foundation and empirical evidence for trustworthy, scalable intent-oriented operating systems.

Assessing feasibility of local LLMs for privacy-conscious operating systemsComparing performance of open models against proprietary systemsEvaluating open-source LLMs for local user intent resolution

Evaluating Small Language Models for News Summarization: Implications and Factors Influencing Performance

Feb 02, 2025
BX
Borui Xu
🏛️ Shandong University | National University of Singapore | HKUST

This study systematically evaluates the performance of 19 small language models (SLMs) on news summarization under resource-constrained settings, benchmarking against 70B-parameter large language models (LLMs) across summary quality, coherence, factual consistency, and length efficiency. Method: Using a 2,000-sample news corpus, we employ a multi-dimensional evaluation framework combining human assessment with automated metrics—ROUGE, BERTScore, and fact-checking—and conduct controlled prompt experiments, instruction-tuning ablations, and computational complexity analysis. Contribution/Results: We empirically demonstrate that top-performing SLMs (e.g., Phi3-Mini) match 70B LLMs in summary quality while reducing output length by 35% on average. Simpler prompts outperform complex instructions, and instruction tuning yields no significant gains for news summarization. This work establishes, for the first time, the practical efficacy boundary and efficient deployment pathway for SLMs in news summarization.

Resource-constrained EnvironmentSmall Language ModelsText Summarization

A Review on Edge Large Language Models: Design, Execution, and Applications

Sep 29, 2024
YZ
Yue Zheng
🏛️ Zhejiang University of Technology | Zhejiang University

Deploying large language models (LLMs) on edge devices faces fundamental challenges including constrained computational resources, limited memory capacity, and hardware heterogeneity. To address these, this paper systematically surveys the full lifecycle of edge LLMs and introduces a comprehensive, stack-wide technical taxonomy—spanning lightweight model design (e.g., pruning, quantization, distillation), runtime optimization (e.g., memory-aware inference, device-adaptive scheduling), and on-device deployment (e.g., cloud-edge collaborative frameworks). It proposes a novel cross-platform co-deployment paradigm that unifies pre-deployment model compression with dynamic execution optimization. Based on a rigorous synthesis of over 120 state-of-the-art studies, the work identifies five persistent technical bottlenecks and six key future research directions. The resulting methodology provides both a reusable conceptual framework and practical guidelines for deploying AI at the edge.

Challenges in computational limitationsDeploying LLMs on edge devicesOptimizing edge LLMs lifecycle

Language Representations Can be What Recommenders Need: Findings and Potentials

Jul 07, 2024
LS
Leheng Sheng
🏛️ National University of Singapore | University of Science and Technology of China

This work investigates whether pretrained language models (PLMs) implicitly encode user preferences and collaborative signals in their representation spaces. To this end, we propose TextCF—a purely text-driven collaborative filtering framework that eliminates the need for item ID embeddings. TextCF employs lightweight linear projections to map item title representations extracted from PLMs (e.g., LLaMA, BERT) into a recommendation-aware latent space. Theoretically, we establish a homomorphic structure between linguistic and collaborative representation spaces. Empirically, TextCF significantly outperforms state-of-the-art ID-based CF methods across multiple public benchmarks—demonstrating, for the first time, that title text alone suffices to achieve superior recommendation performance. Moreover, TextCF exhibits zero-shot recommendation capability and inherent potential for user intent awareness. Collectively, it introduces a new paradigm for recommender systems that is initialization-friendly, generalizable, and inherently interpretable.

Designs collaborative filtering models using only language representationsExplores if language models encode user preferences for recommendationsTests mapping language representations to item spaces for better performance

Latest Papers

What's happening recently
View more

Semantic Interactivity: leveraging NLP to enable a shared interaction approach for joint activities

Nov 07, 2025
OV
Olaf V. Adan
🏛️ Eindhoven University of Technology

Existing collaborative systems predominantly focus on individual tasks, limiting their capacity to support authentic joint activities and shared experiences. To address this, we propose “semantic interaction”—a novel paradigm that integrates natural language processing (NLP) techniques to enable group co-discussion within co-located spaces. We instantiate this paradigm through CollEagle, an interactive tabletop system that supports low-effort information creation, collective externalization, dynamic reorganization, and structured management of shared artifacts. Crucially, CollEagle embeds deep semantic understanding directly into the collaborative interface—marking a departure from traditional task-centric design approaches. This integration significantly enhances interaction quality and shared cognition among participants. Preliminary empirical evaluation demonstrates that semantic interaction effectively modulates collaborative behaviors, offering both theoretical grounding and practical design guidelines for collaboration interfaces targeting joint activity.

Creating low-effort interaction mechanisms for joint information structuringDeveloping NLP-driven semantic interactivity for shared collaborative experiencesEnabling shared externalization practices in collocated teamwork systems

This work addresses the challenges of deploying automatic summarization for multi-party dialogues in industrial settings, where dynamic shifts in user requirements, highly subjective evaluation criteria, and scarce annotated data render conventional static dataset approaches inadequate. To bridge this gap, the authors propose a practical, lifecycle-oriented framework for adaptive summarization systems, leveraging an agent-based architecture to decompose tasks, integrating large language model prompt engineering with component-level optimization, and introducing a robust evaluation mechanism tailored to handle requirement volatility and subjectivity. Moving beyond static research paradigms, the project delivers a reusable industrial development guideline and uncovers critical practical insights—such as the impact of data quality and risks of vendor lock-in—thereby significantly enhancing system reliability and adaptability in real-world applications.

applied summarizationdialogue summarizationevolving requirements

This work addresses the high computational overhead and scalability challenges inherent in training and inference of large language models (LLMs) by proposing the first cloud-native distributed systems framework specifically designed for LLMs. The framework integrates key technologies including microservices, auto-scaling, cloud-edge collaboration, serverless inference, and federated learning, while also exploring a forward-looking pathway toward quantum computing integration. By systematically optimizing data management and resource scheduling, the proposed architecture significantly enhances deployment efficiency and elasticity. It provides robust technical foundations for the efficient operation and continuous evolution of LLMs, thereby fostering synergies across industry, academia, and research communities and advancing standardization efforts in the field.

Cloud-nativeComputational DemandsDistributed Systems

To address the prohibitively high computational cost and inability of large language models (LLMs) to meet stringent low-latency and high-throughput requirements in LinkedIn’s semantic job search, this paper proposes a co-optimization framework for small language models (SLMs) with pure-text decoders. Our method jointly optimizes structured pruning and semantic-aware context compression: structured pruning reduces model parameters by 40%, while context-aware input sequence compression achieves an average 10× reduction in context length. We further integrate GPU kernel optimization and a lightweight serving architecture. Evaluated in LinkedIn’s production environment, the system achieves a 10× throughput improvement—reaching one million queries per second—with P99 latency consistently below 50 ms, while preserving retrieval quality (Recall@10 degradation < 0.3%). This work establishes an efficient, scalable deployment paradigm for SLMs in large-scale semantic search.

Applying model compression to reduce size while maintaining accuracyOptimizing small language models for efficient semantic job search deploymentScaling serving infrastructure for high-throughput industry applications

This study investigates whether attention heads and input embeddings in large language models (LLMs) genuinely encode human-interpretable semantic information. Addressing concerns that prevailing interpretability methods—such as those based on attention weights or embedding analyses—may be confounded by data artifacts or methodological biases, the authors employ token-level relational structural probes and map human-interpretable attributes onto the embedding space to systematically evaluate the validity of these approaches across multiple Transformer layers. Their findings reveal that both widely adopted explanation techniques fail to reliably reflect the model’s true semantic capabilities, thereby challenging the foundational assumptions underlying current claims about LLMs’ “understanding.” This work carries significant implications for deploying LLMs in edge and distributed computing environments, where interpretability and reliability are critical.

attention mechanismsembedding analysisexplainability

Hot Scholars

TH

Tiansheng Huang

Georgia Institute of Technology
Parallel and Distributed ComputingDistributed machine learningLLM safety
TW

Taro Watanabe

Nara Institute of Science and Technology
Machine TranslationMachine Learning
YL

Yufeng Li

East China Normal University
Artificial Intelligence
JZ

Jiayuan Zhou

Principal Researcher, Waterloo Research Centre, Huawei Canada
OSS VulnerabilitiesCrowdsourced Software EngineeringMining Software RepositoriesEmpirical