parameter-efficient modelling

Designs and builds models, architectures, modules, and layers that achieve target predictive or detection performance while minimizing parameter count, including budgeting and integration of compact components. Also analyzes parameterization choices and per-parameter performance trade‑offs (parameterization analysis) and implements lightweight architectures using techniques such as depthwise‑separable convolutions, small recurrent blocks, or other parameter‑saving modules.

parameter-efficientmodelling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Smaller, Faster, Cheaper: Architectural Designs for Efficient Machine Learning

Jul 26, 2025
SW
Steven Walton
🏛️ University of Oregon

To address excessive computational overhead when deploying vision models on resource-constrained devices, this paper proposes an efficient Vision Transformer (ViT) architecture design framework. Methodologically: (1) it optimizes the input-output data pathway to enhance representational capacity of lightweight models; (2) it restructures the context window of computationally constrained attention mechanisms to improve local-global modeling efficiency; and (3) it leverages the invertibility and explicit probabilistic modeling properties of normalizing flows to enable high-fidelity, low-overhead knowledge distillation. Experiments demonstrate that the proposed approach achieves comparable or superior accuracy on benchmarks such as ImageNet, while requiring significantly fewer parameters and FLOPs. It also substantially reduces inference latency and memory footprint. The framework establishes a scalable new paradigm for efficient visual understanding at the edge.

Design efficient ML architectures for high performance with fewer resourcesImprove vision transformers and normalizing flows for computational efficiencyOptimize data flow in neural units to enhance small model performance

This work proposes a parameterized convolutional accelerator architecture based on high-level synthesis (HLS) to address the limitations of conventional CNN accelerators, which often prioritize peak performance at the expense of critical embedded constraints such as latency, power consumption, area, and cost. By leveraging a hardware-software co-design approach, the proposed architecture enables efficient multi-objective optimization across these dimensions, overcoming the rigidity of fixed architectures. Experimental results demonstrate that, compared to non-parameterized designs, the proposed solution not only meets stringent embedded deployment requirements but also offers superior scalability and energy efficiency. Furthermore, the framework exhibits broad applicability and can be readily extended to other deep learning acceleration scenarios.

CNN acceleratordesign constraintsembedded deep learning

Traditional AI systems rely on fixed monolithic models, which struggle to dynamically allocate resources, decompose tasks, or update knowledge in response to varying inputs, leading to degraded performance and increased costs. This work proposes the first system-level design methodology for distributed composite AI systems, formulating a design space through workflow topologies and configuration choices and identifying eight core design patterns. The framework jointly optimizes model selection and runtime parameters, enabling task decomposition, multi-model orchestration, and explicit control logic, thereby facilitating a shift from static monolithic architectures toward dynamic, composable, and adaptive ones. Evaluated across three case studies, the approach reduces latency by up to 60% and cost by up to 71%, with only a 2.5–4 percentage point drop in accuracy.

Compound AI SystemsDistributed AIModel-Centric Design

MoRe Fine-Tuning with 10x Fewer Parameters

Aug 30, 2024
WT
Wenxuan Tan
🏛️ University of Wisconsin—Madison

Existing parameter-efficient fine-tuning (PEFT) methods—such as LoRA—rely on heuristic adapter architectures, suffering from poor generalization and limited transferability across models. Method: We propose the first learnable rectangular adapter search framework grounded in Monarch matrices—the first application of Monarch structure to PEFT—supported by theoretical analysis demonstrating superior expressivity over LoRA. Our approach employs differentiable neural architecture search to automatically discover optimal lightweight adapter topologies, eliminating manual specification of rank or module shape, and integrates low-parameter adapter design with efficient fine-tuning strategies. Contribution/Results: On multi-task and multi-model benchmarks, our method significantly outperforms state-of-the-art PEFT approaches, achieving comparable or superior performance using only 5% of LoRA’s parameters. It delivers both strong cross-task/model generalization and exceptional parameter efficiency.

Enhancing performance with fewer parameters in PEFT techniquesOptimizing adapter architectures for parameter-efficient fine-tuningReducing reliance on heuristics in low-rank adapters (LoRA)

This work addresses the inefficiencies in large-scale recommendation systems caused by maintaining separate models for different scenarios and objectives, which hinders development velocity and delays technology adoption. To overcome this, the authors propose the Standardized Model Template (SMT) framework, which leverages composable, standardized machine learning components to enable “design once, deploy everywhere,” uniformly accommodating diverse data distributions and optimization objectives. By decoupling model architecture from scenario-specific configurations, SMT reduces the complexity of technology deployment from O(n·2ᵏ) to O(n+k), breaking away from the conventional “one objective, one model” paradigm. Empirical evaluation on Meta’s ad ranking system demonstrates that SMT improves average cross-entropy by 0.63%, reduces engineering time per model iteration by 92%, and increases the throughput of technology-model pair adoption by 6.3×.

computational advertisinglarge-scale ML ecosystemsML technique propagation

Latest Papers

What's happening recently
View more

This work addresses the challenge of achieving real-time performance in edge vision systems, where conventional multi-stage detection-classification pipelines suffer from fully GPU-serialized execution. The authors propose a five-step optimization methodology enabling zero-GPU-fallback INT8 deployment of classification models on NVIDIA Jetson Deep Learning Accelerators (DLAs), and construct a parallel inference pipeline with GPU-based detection and DLA-based classification. Key innovations include the first-ever DLA deployment workflow that entirely avoids GPU fallback, overcoming DLA operator limitations and quantization compatibility bottlenecks through techniques such as manual dynamic range calibration, quantization-aware training, and ONNX graph surgery. Evaluated on a Jetson Orin NX, the dual-head human attribute classifier operating in parallel with the detector incurs only a 0.8 FPS overhead (12.5 vs. 13.3 FPS) and supports cost-free scaling across dual DLAs.

edge inferencehierarchical classificationmodel deployment

This study addresses the pressing environmental sustainability challenges posed by the high computational costs and energy consumption of large-scale AI models. It presents a systematic review of full-stack technical pathways toward greener foundation models, uniquely integrating co-optimization strategies across algorithmic and hardware layers. On the algorithmic side, it encompasses linear-complexity architectures, sparsification, and parameter-efficient fine-tuning; on the hardware side, it includes energy-efficient chips, memory-centric designs, and cross-platform deployment. The work further extends these advances to sustainability-oriented applications such as remote sensing and national infrastructure. By constructing a comprehensive roadmap spanning model design, training, and deployment, this research provides both theoretical grounding and practical guidance for developing large models that are efficient, scalable, and socially responsible.

computational costenergy consumptionenvironmental sustainability

This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.

compression and accelerationconstraint-drivendeployment constraints

本文探讨了RISC-V架构在机器学习应用中的现状与挑战,通过分析其指令集扩展、核心实现及软件工具链等,提出四个研究方向以解决当前局限。

ecosystem fragmentationenergy efficiencymachine learning

This work addresses the challenge of efficiently deploying convolutional neural networks in resource-constrained embedded environments, where high computational and memory demands hinder practical application. To overcome this, the authors present a lightweight object detection system built upon FPGA implementation of YOLOv3-Tiny, incorporating algorithm-hardware co-optimization techniques including low-bit quantization, batch normalization folding, and lookup-table-based activation mapping. A pipelined architecture and on-chip caching mechanism are further designed to minimize off-chip memory access. Evaluated on the ZYNQ-XC7Z035 platform, the proposed system achieves an inference latency of 0.211 seconds—representing a 75.58% speedup over the baseline—while delivering an energy efficiency of 10.11 GOPS/W and reducing hardware resource utilization by up to 51.94%.

computational complexityconvolutional neural networksembedded systems

Hot Scholars

YW

Yunhe Wang

Noah's Ark Lab, Huawei Technologies
Deep LearningLanguage ModelMachine LearningComputer Vision
JL

Jianguo Li

Director, Ant Group
deep learningcomputer visionmachine learningsystem
DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing
HT

Hao Tan

Adobe Research
Vision and Language3D Multimodal
YT

Yehui Tang

Shanghai Jiao Tong University
Machine LearningQuantum AI & AI4Science