Institution profile

Baseten

Industry researchnorthamerica · us
Official website
Research library6linked papers
Opportunities28open roles
Selected work

Representative Papers

WavePP: High-Throughput Pipeline Parallel LLM Prefill under Prefix Reuse

Sep 28, 2026

This study addresses the high coordination overhead of cross-stage prefix reuse and the limited prefill throughput in pipeline parallelism. To mitigate these issues, this work proposes an asynchronous prefix reuse coordination mechanism coupled with a dynamic chunking strategy. By constructing an efficient runtime that overlaps request admission with execution, the system asynchronously synchronizes cache states across all stages and dynamically plans chunk sizes to maximize pipeline utilization, thereby eliminating admission latency bottlenecks. Implemented on TensorRT-LLM, the proposed approach achieves up to a 2.91× throughput improvement on GLM and MiniMax models, outperforming mainstream baseline systems in most scenarios.

0 citationsRead paper

Can a Language Model Learn Facts Continually in Its Weights?

Jul 12, 2026

This study investigates whether large language models can continuously learn and accumulate new factual knowledge through weight updates without catastrophically forgetting previously acquired information. Focusing on the Qwen3 model, the authors sequentially inject fictional facts via weight editing and systematically evaluate the impact of different training formulations on knowledge retention and usability using held-out question answering, KL divergence analysis, and context-based recovery tests. The findings reveal that training with diverse phrasings boosts knowledge retention to 46%, compared to merely 1% with a single statement. Moreover, ostensibly “forgotten” facts remain implicitly encoded in the weights and can be recovered to 77–80% accuracy with appropriate prompting. These results suggest that reliable knowledge transfer depends more on contextual activation than on static weight embedding, indicating that forgetting primarily stems from query redirection masking rather than actual knowledge erasure.

0 citationsRead paper

Still: Amortized KV Cache Compaction in a Single Forward Pass

Jun 05, 2026

This work addresses the severe memory bottleneck caused by KV caching in long-context inference, necessitating compression methods that simultaneously achieve light weight, high expressivity, and cross-trajectory reusability. The authors propose Still, a lightweight Perceiver module embedded at each layer of a frozen base model, which enables synthetic KV cache compression in a single forward pass without per-context optimization. Trained solely on top of frozen foundation models, Still is the first method to jointly satisfy all three desiderata—lightweight design, strong representational capacity, and reusability—enabling efficient inference and flexible summarization even under extreme compression ratios (8–200×). Evaluated on Qwen and Gemma with 128k context lengths, Still significantly outperforms existing approaches, achieving gains of 8–22 points over the strongest baseline on the RULER benchmark and setting new state-of-the-art results on LongBench summarization tasks.

0 citationsRead paper
Recent publications

Latest Papers

WavePP: High-Throughput Pipeline Parallel LLM Prefill under Prefix Reuse

Sep 28, 2026

This study addresses the high coordination overhead of cross-stage prefix reuse and the limited prefill throughput in pipeline parallelism. To mitigate these issues, this work proposes an asynchronous prefix reuse coordination mechanism coupled with a dynamic chunking strategy. By constructing an efficient runtime that overlaps request admission with execution, the system asynchronously synchronizes cache states across all stages and dynamically plans chunk sizes to maximize pipeline utilization, thereby eliminating admission latency bottlenecks. Implemented on TensorRT-LLM, the proposed approach achieves up to a 2.91× throughput improvement on GLM and MiniMax models, outperforming mainstream baseline systems in most scenarios.

0 citationsRead paper

Can a Language Model Learn Facts Continually in Its Weights?

Jul 12, 2026

This study investigates whether large language models can continuously learn and accumulate new factual knowledge through weight updates without catastrophically forgetting previously acquired information. Focusing on the Qwen3 model, the authors sequentially inject fictional facts via weight editing and systematically evaluate the impact of different training formulations on knowledge retention and usability using held-out question answering, KL divergence analysis, and context-based recovery tests. The findings reveal that training with diverse phrasings boosts knowledge retention to 46%, compared to merely 1% with a single statement. Moreover, ostensibly “forgotten” facts remain implicitly encoded in the weights and can be recovered to 77–80% accuracy with appropriate prompting. These results suggest that reliable knowledge transfer depends more on contextual activation than on static weight embedding, indicating that forgetting primarily stems from query redirection masking rather than actual knowledge erasure.

0 citationsRead paper

Still: Amortized KV Cache Compaction in a Single Forward Pass

Jun 05, 2026

This work addresses the severe memory bottleneck caused by KV caching in long-context inference, necessitating compression methods that simultaneously achieve light weight, high expressivity, and cross-trajectory reusability. The authors propose Still, a lightweight Perceiver module embedded at each layer of a frozen base model, which enables synthetic KV cache compression in a single forward pass without per-context optimization. Trained solely on top of frozen foundation models, Still is the first method to jointly satisfy all three desiderata—lightweight design, strong representational capacity, and reusability—enabling efficient inference and flexible summarization even under extreme compression ratios (8–200×). Evaluated on Qwen and Gemma with 128k context lengths, Still significantly outperforms existing approaches, achieving gains of 8–22 points over the strongest baseline on the RULER benchmark and setting new state-of-the-art results on LongBench summarization tasks.

0 citationsRead paper