Institution profile

Teradata Corporation

Industry researchnorthamerica · us
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

KathDB-FAO: Synthesized Query Plans in a Multimodal DBMS

Sep 23, 2026

This study addresses the inherent challenge of balancing execution efficiency and accuracy in natural language database querying. To this end, it proposes a multimodal database query subsystem that translates natural language into execution plans composed of dynamically synthesized functions, thereby achieving query-level optimization. The core innovation lies in extracting atomic actions, establishing input-output contracts, and synthesizing functions on the fly, effectively integrating techniques from natural language processing, program synthesis, and database query optimization. Evaluated on the SemBench benchmark, the proposed approach reduces execution costs by an average of 58.8% while maintaining comparable or superior query quality. These results demonstrate that the method effectively enables both efficient and accurate natural language interfaces to databases.

0 citationsRead paper

Puffin-Backed Vector Indexes: Attaching Approximate Nearest Neighbor Indexes to Apache Iceberg Snapshots for Compute-Disaggregated Query Engines

Jun 02, 2026

This work addresses the operational complexity and architectural coupling introduced by conventional vector similarity search systems in compute-storage disaggregated environments, which typically rely on a separate indexing layer. To resolve this, the authors propose deeply integrating a distributed approximate nearest neighbor (ANN) index into the Apache Iceberg table format. The approach leverages Puffin sidecar files to store sharded Vamana graphs and employs Iceberg’s snapshot mechanism to atomically bind indexes with data, enabling versioning and time travel. A coordinator-executor protocol is introduced, featuring a hierarchical probing strategy where coordinators cache compact centroid indexes while executors store large graphs on SSDs. The design reuses Iceberg’s REST catalog optimistic concurrency control for index commits. This solution is the first to host billion-scale vector graph indexes within Puffin, achieving a favorable trade-off between recall and latency while preserving Iceberg’s native transactional capabilities and significantly reducing system complexity.

0 citationsRead paper
Recent publications

Latest Papers

KathDB-FAO: Synthesized Query Plans in a Multimodal DBMS

Sep 23, 2026

This study addresses the inherent challenge of balancing execution efficiency and accuracy in natural language database querying. To this end, it proposes a multimodal database query subsystem that translates natural language into execution plans composed of dynamically synthesized functions, thereby achieving query-level optimization. The core innovation lies in extracting atomic actions, establishing input-output contracts, and synthesizing functions on the fly, effectively integrating techniques from natural language processing, program synthesis, and database query optimization. Evaluated on the SemBench benchmark, the proposed approach reduces execution costs by an average of 58.8% while maintaining comparable or superior query quality. These results demonstrate that the method effectively enables both efficient and accurate natural language interfaces to databases.

0 citationsRead paper

Puffin-Backed Vector Indexes: Attaching Approximate Nearest Neighbor Indexes to Apache Iceberg Snapshots for Compute-Disaggregated Query Engines

Jun 02, 2026

This work addresses the operational complexity and architectural coupling introduced by conventional vector similarity search systems in compute-storage disaggregated environments, which typically rely on a separate indexing layer. To resolve this, the authors propose deeply integrating a distributed approximate nearest neighbor (ANN) index into the Apache Iceberg table format. The approach leverages Puffin sidecar files to store sharded Vamana graphs and employs Iceberg’s snapshot mechanism to atomically bind indexes with data, enabling versioning and time travel. A coordinator-executor protocol is introduced, featuring a hierarchical probing strategy where coordinators cache compact centroid indexes while executors store large graphs on SSDs. The design reuses Iceberg’s REST catalog optimistic concurrency control for index commits. This solution is the first to host billion-scale vector graph indexes within Puffin, achieving a favorable trade-off between recall and latency while preserving Iceberg’s native transactional capabilities and significantly reducing system complexity.

0 citationsRead paper