Institution profile

Black Sesame Technologies

Industry researchasia · cn
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

CurveTQ: Rotation-Free Trellis Quantization of LLM Weights via Curvature-Weighted Search

Oct 06, 2026

This study addresses the limitations of existing large language model quantization methods that rely on rotation operations, which increase decoding overhead and fail to fully exploit Hessian curvature information. We establish, for the first time, the equivalence between rotation and weighted search, proposing a rotation-free lattice quantization framework. Specifically, diagonal Hessian weights are incorporated into the Viterbi search metric to achieve curvature-aware encoding, while error-feedback residuals adaptively adjust the lattice initial state to eliminate the need for rotation. Evaluated on 4–8B and 35B mixture-of-experts models, the proposed method outperforms QTIP and Proteus by 1–3 percentage points at 2-bit precision. Furthermore, by obviating inverse rotation computations, it achieves state-of-the-art decoding speed.

0 citationsRead paper

Provenance Tracking in AI Compilers through the Lens of Coalgebra

Jun 09, 2026

This work addresses the challenge of reliably tracing the provenance of tensors and operators through graph rewrites—particularly non-injective transformations—in AI compilers. The authors propose a lightweight, generative provenance method grounded in observational semantics, which infers origins by analyzing the behavioral effects of graph transformations rather than relying on identifier propagation. For the first time, they introduce coalgebraic modeling and bisimulation to this domain, guaranteeing provenance consistency even after intermediate nodes are eliminated. The approach requires no invasive compiler modifications and naturally supports non-injective rewrites. Evaluated within COVAN, a prototype AI compiler, the method demonstrates stable, low-overhead provenance tracking throughout an end-to-end compilation pipeline.

0 citationsRead paper

NTR: Neural Token Reconstruction for Scene Token Bottleneck in End-to-End Driving

May 29, 2026

This work addresses the insufficient visual constraints and representational redundancy in perception-agnostic end-to-end autonomous driving, where scene tokens are supervised solely by planning objectives. To mitigate this, the authors propose a Neural Token Reconstruction (NTR) framework that introduces, for the first time, a self-distilled masked latent reconstruction objective at the scene token bottleneck. This approach leverages compact tokens as memory to reconstruct patch-level image features, thereby enhancing their representational capacity. Guided by semantic priors from foundation models, the reconstruction process focuses on driving-relevant structures without requiring additional modules during inference. Evaluated on Waymo E2E and NavSim1&2 benchmarks, NTR achieves state-of-the-art performance (RFS: 8.0461; PDMS/EPDMS: 94.1/90.9), significantly reducing token redundancy and improving effective rank.

0 citationsRead paper
Recent publications

Latest Papers

CurveTQ: Rotation-Free Trellis Quantization of LLM Weights via Curvature-Weighted Search

Oct 06, 2026

This study addresses the limitations of existing large language model quantization methods that rely on rotation operations, which increase decoding overhead and fail to fully exploit Hessian curvature information. We establish, for the first time, the equivalence between rotation and weighted search, proposing a rotation-free lattice quantization framework. Specifically, diagonal Hessian weights are incorporated into the Viterbi search metric to achieve curvature-aware encoding, while error-feedback residuals adaptively adjust the lattice initial state to eliminate the need for rotation. Evaluated on 4–8B and 35B mixture-of-experts models, the proposed method outperforms QTIP and Proteus by 1–3 percentage points at 2-bit precision. Furthermore, by obviating inverse rotation computations, it achieves state-of-the-art decoding speed.

0 citationsRead paper

Provenance Tracking in AI Compilers through the Lens of Coalgebra

Jun 09, 2026

This work addresses the challenge of reliably tracing the provenance of tensors and operators through graph rewrites—particularly non-injective transformations—in AI compilers. The authors propose a lightweight, generative provenance method grounded in observational semantics, which infers origins by analyzing the behavioral effects of graph transformations rather than relying on identifier propagation. For the first time, they introduce coalgebraic modeling and bisimulation to this domain, guaranteeing provenance consistency even after intermediate nodes are eliminated. The approach requires no invasive compiler modifications and naturally supports non-injective rewrites. Evaluated within COVAN, a prototype AI compiler, the method demonstrates stable, low-overhead provenance tracking throughout an end-to-end compilation pipeline.

0 citationsRead paper

NTR: Neural Token Reconstruction for Scene Token Bottleneck in End-to-End Driving

May 29, 2026

This work addresses the insufficient visual constraints and representational redundancy in perception-agnostic end-to-end autonomous driving, where scene tokens are supervised solely by planning objectives. To mitigate this, the authors propose a Neural Token Reconstruction (NTR) framework that introduces, for the first time, a self-distilled masked latent reconstruction objective at the scene token bottleneck. This approach leverages compact tokens as memory to reconstruct patch-level image features, thereby enhancing their representational capacity. Guided by semantic priors from foundation models, the reconstruction process focuses on driving-relevant structures without requiring additional modules during inference. Evaluated on Waymo E2E and NavSim1&2 benchmarks, NTR achieves state-of-the-art performance (RFS: 8.0461; PDMS/EPDMS: 94.1/90.9), significantly reducing token redundancy and improving effective rank.

0 citationsRead paper