Institution profile

Perplexity.AI

Industry researchnorthamerica · us
Official website
Research library5linked papers
Opportunities43open roles
Selected work

Representative Papers

A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and Teaching

Oct 07, 2026

This study addresses the challenges of sparse rewards in reinforcement learning and performance degradation caused by teacher-student capability mismatch during self-distillation. To this end, we propose JOLT, a method that jointly trains a single policy to serve as both a privileged teacher and an unprivileged student, ensuring that the guidance remains aligned with the student’s current capabilities. By deriving the necessary and sufficient conditions under which the teacher update constitutes a positive multiple of the student gradient, we design a teacher optimization objective that integrates outcome rewards with KL regularization. Leveraging on-policy distillation and joint optimization, JOLT significantly improves both training efficiency and final performance on tasks such as mathematical reasoning and programming.

0 citationsRead paper

LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization

Sep 29, 2026

This study addresses the issues of error accumulation and precision degradation in recurrent state quantization for linear attention mechanisms. We propose a training-free, 8-bit near-lossless quantization method that introduces a novel skip-window quantization mechanism to suppress rounding error propagation. This is combined with high-precision outlier compensation tokens to handle extreme activation values, alongside a pre-quantization residual smoothing algorithm to further enhance numerical stability. Experimental results demonstrate that the proposed approach preserves FP32-level accuracy while achieving 2.05–3.7× kernel speedup and 1.47× end-to-end inference acceleration, offering an effective solution for the efficient deployment of linear attention models.

0 citationsRead paper

Diffusion-Pretrained Dense and Contextual Embeddings

Feb 11, 2026

This work addresses the challenge of balancing global semantic modeling and computational efficiency in multilingual long-document retrieval by proposing a novel embedding method based on diffusion-based pretrained language models. The approach integrates document-level global context into paragraph representations through a late-chunking strategy and a context-aware bidirectional attention mechanism. High-quality dense vectors are further refined via multi-stage contrastive learning and mean pooling. The resulting model, pplx-embed-v1, achieves strong performance across multilingual and code retrieval benchmarks, including MTEB and MIRACL. Its contextual variant, pplx-embed-context-v1, sets a new state-of-the-art on ConTEB and demonstrates both efficiency and practicality in large-scale production environments with tens of millions of documents.

0 citationsRead paper

RDMA Point-to-Point Communication for LLM Systems

Oct 31, 2025

Existing LLM systems—including decoupled inference, MoE routing, and asynchronous reinforcement learning fine-tuning—rely on flexible point-to-point communication, yet mainstream RDMA implementations are tightly coupled to vendor-specific NICs, hindering portability and integration into inference engines. This work proposes TransferEngine, a universal RDMA communication framework that introduces a unified abstraction layer and novel primitives—WriteImm and ImmCounter—to enable precise, out-of-order transmission completion notification. TransferEngine achieves transparent, hardware-agnostic management across heterogeneous NICs (e.g., NVIDIA ConnectX-7, AWS EFA). Evaluation demonstrates: (1) efficient and reliable dynamic KvCache migration; (2) trillion-parameter RL weight updates completed in just 1.3 seconds; and (3) MoE inference latency lower than DeepEP with peak throughput reaching 400 Gbps.

0 citationsRead paper
Recent publications

Latest Papers

A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and Teaching

Oct 07, 2026

This study addresses the challenges of sparse rewards in reinforcement learning and performance degradation caused by teacher-student capability mismatch during self-distillation. To this end, we propose JOLT, a method that jointly trains a single policy to serve as both a privileged teacher and an unprivileged student, ensuring that the guidance remains aligned with the student’s current capabilities. By deriving the necessary and sufficient conditions under which the teacher update constitutes a positive multiple of the student gradient, we design a teacher optimization objective that integrates outcome rewards with KL regularization. Leveraging on-policy distillation and joint optimization, JOLT significantly improves both training efficiency and final performance on tasks such as mathematical reasoning and programming.

0 citationsRead paper

LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization

Sep 29, 2026

This study addresses the issues of error accumulation and precision degradation in recurrent state quantization for linear attention mechanisms. We propose a training-free, 8-bit near-lossless quantization method that introduces a novel skip-window quantization mechanism to suppress rounding error propagation. This is combined with high-precision outlier compensation tokens to handle extreme activation values, alongside a pre-quantization residual smoothing algorithm to further enhance numerical stability. Experimental results demonstrate that the proposed approach preserves FP32-level accuracy while achieving 2.05–3.7× kernel speedup and 1.47× end-to-end inference acceleration, offering an effective solution for the efficient deployment of linear attention models.

0 citationsRead paper

Diffusion-Pretrained Dense and Contextual Embeddings

Feb 11, 2026

This work addresses the challenge of balancing global semantic modeling and computational efficiency in multilingual long-document retrieval by proposing a novel embedding method based on diffusion-based pretrained language models. The approach integrates document-level global context into paragraph representations through a late-chunking strategy and a context-aware bidirectional attention mechanism. High-quality dense vectors are further refined via multi-stage contrastive learning and mean pooling. The resulting model, pplx-embed-v1, achieves strong performance across multilingual and code retrieval benchmarks, including MTEB and MIRACL. Its contextual variant, pplx-embed-context-v1, sets a new state-of-the-art on ConTEB and demonstrates both efficiency and practicality in large-scale production environments with tens of millions of documents.

0 citationsRead paper

RDMA Point-to-Point Communication for LLM Systems

Oct 31, 2025

Existing LLM systems—including decoupled inference, MoE routing, and asynchronous reinforcement learning fine-tuning—rely on flexible point-to-point communication, yet mainstream RDMA implementations are tightly coupled to vendor-specific NICs, hindering portability and integration into inference engines. This work proposes TransferEngine, a universal RDMA communication framework that introduces a unified abstraction layer and novel primitives—WriteImm and ImmCounter—to enable precise, out-of-order transmission completion notification. TransferEngine achieves transparent, hardware-agnostic management across heterogeneous NICs (e.g., NVIDIA ConnectX-7, AWS EFA). Evaluation demonstrates: (1) efficient and reliable dynamic KvCache migration; (2) trillion-parameter RL weight updates completed in just 1.3 seconds; and (3) MoE inference latency lower than DeepEP with peak throughput reaching 400 Gbps.

0 citationsRead paper