A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi

šŸ“… 2026-07-16
šŸ“ˆ Citations: 0
✨ Influential: 0
šŸ“„ PDF
šŸ¤– AI Summary
This work presents the first end-to-end fully on-GPU inference of the modern multimodal large model MiniCPM-V-4.6 on a Fermi-architecture GPU (Tesla C2075) with only 6 GB of memory. By leveraging hand-written CUDA GEMM kernels, 8-bit weight dequantization, an sm_20-compatible SigLIP2 vision encoder, chunked delta-rule recursive computation, and a zero-overhead fused window attention mechanism, the approach overcomes severe memory and compute constraints. Key contributions include the first demonstration of full-GPU multimodal inference on Fermi hardware, the counterintuitive finding that 4-bit quantization degrades decoding speed, mitigation of platform-specific floating-point indexing nondeterminism, and significant performance gains: image encoding completes in just 0.93 seconds, and 10k-token prefill throughput increases 17-fold to 361 tokens per second, with visual forward-pass errors below 1.4e⁻⁵.
šŸ“ Abstract
A companion study ran a 35B mixture-of-experts model on a 2011 NVIDIA Tesla C2075 (Fermi, sm_20, 6GB) as a GPU-prefill/CPU-decode hybrid, because the 4-bit model did not fit in device memory (arXiv:2606.24031). This report keeps the hardware and asks what a model that fits can do: we deploy MiniCPM-V-4.6, a modern multimodal assistant pairing a SigLIP2 vision encoder and window-attention merger (16x visual token compression) with a compact hybrid gated-delta-net backbone, entirely on the GPU. Three results. (i) An all-GPU engine built on measured foundations: projections that dequantize 8-bit weights once and call the vendor SGEMM still in the last Fermi toolchain (64% of FP32 peak; our best hand-written GEMM hit 37%, wrongly called the ceiling); a chunked delta-rule rewrite of the recurrent layers, 2.8x faster than the sequential scan once attribution exposed one bad kernel; and a measured negative: 4-bit weights make decode slower than 8-bit here, since Fermi issues nibble-unpacking shifts at half rate. (ii) The vision side is a port with a proof obligation: we translate tower, merger, and projector to sm_20 CUDA, validating every stage against a locally generated reference forward (full tower 1.4e-5). One failure, position-embedding bucketization differing on exact rational ties, generalizes to a rule: float tie-breaking in index arithmetic is implementation-defined; call the reference operator, do not reimplement it. (iii) Long context exposes an O(N^2) wall short benchmarks hide: prefill falls from 114 tok/s at 2k tokens to 21 at 10k in a naive attention kernel; per-head vendor-GEMM calls writing into the existing score buffer (zero extra memory) restore a flat profile (408 at 2k, 361 at 10k; 17x), verified by exact needle retrieval from 60% depth. The same rewrite cuts image encoding 6x, to 0.93s. The system answers an image question end-to-end in 1.7s.
Problem

Research questions and friction points this paper is trying to address.

multimodal inference
legacy GPU
Fermi architecture
all-GPU deployment
resource-constrained AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

all-GPU inference
Fermi architecture
multimodal LLM
CUDA optimization
long-context attention
šŸ”Ž Similar Papers
A
A. C. Opus
Department of Physics, University of Puerto Rico, Mayaguez, PR 00680, USA
J
J. Q. Lu
Department of Physics, University of Puerto Rico, Mayaguez, PR 00680, USA