Institution profile

DatologyAI

Industry research
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

DatBench: Discriminative, Faithful, and Efficient VLM Evaluations

Jan 05, 2026arXiv.org

Current evaluation methods for vision-language models (VLMs) commonly suffer from modality unfaithfulness, insufficient discriminative power, and computational inefficiency. This work is the first to systematically articulate three core desiderata for VLM evaluation: faithfulness, discriminability, and efficiency. We establish a high-quality evaluation pipeline by reformulating multiple-choice tasks as generative ones, filtering out samples amenable to blind guessing (up to 70% of instances), and correcting mislabeled examples (42% of cases). Based on this framework, we introduce DatBench-Full, encompassing 33 datasets, along with its highly discriminative subset, DatBench. Our benchmarks maintain discriminative capacity comparable to original benchmarks while achieving an average 13× and up to 50× acceleration in evaluation speed.

1 citationsRead paper

Zephon: Elastic Determinism for Online, Stateful Foundation Model Data Loading Pipelines

Oct 02, 2026

This study addresses the non-deterministic data loading in online stateful data pipelines caused by topology changes, failure recovery, and backend heterogeneity. It proposes Zephon, a system enabling resilient and deterministic data loading. The core innovations include a novel topology-agnostic stream partitioning scheme and an ordered decision serialization mechanism, which achieve constant-overhead checkpoint recovery by persisting only bounded in-flight states. These techniques are integrated with stateless parallel processing, backend abstraction, and incremental checkpointing to form a comprehensive solution. Experiments demonstrate that Zephon sustains high throughput under both text and multimodal workloads while providing online determinism guarantees unattainable by existing approaches.

0 citationsRead paper

Luxical: High-Speed Lexical-Dense Text Embeddings

Dec 09, 2025

To address the trade-off between speed and flexibility in web-scale text organization, this paper proposes a lightweight “lexical-dense” embedding paradigm. Leveraging knowledge distillation, it transfers semantic capabilities from large language models into a compact architecture that jointly encodes TF-IDF–based sparse lexical features and dense representations via a small ReLU network. The resulting embeddings retain the versatility of dense vectors—supporting retrieval, clustering, classification, and data cleaning—while achieving inference speeds comparable to FastText. Experiments demonstrate 3×–100× higher throughput than neural baselines on document retrieval and LLM data cleaning tasks, with quality matching state-of-the-art embedding models. This work is the first to systematically bridge the efficiency–expressiveness gap between traditional lexical models and Transformer-based embeddings, establishing a scalable, high-performance paradigm for large-scale text preprocessing.

0 citationsRead paper

BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Aug 14, 2025

To address the performance saturation in large language model (LLM) pretraining caused by data bottlenecks, this work proposes BeyondWeb—a framework that systematically investigates the joint impact of model scale, architecture family, and data rewriting strategies on synthetic data quality. It establishes a high-quality synthetic data generation paradigm integrating model feedback, rewrite filtering, and diversity control. Evaluated on trillion-token pretraining, BeyondWeb significantly enhances semantic richness and training efficiency of synthetic data. Experiments show it achieves +5.1 and +2.6 percentage points average accuracy over Cosmopedia and Nemotron-Synth across 14 benchmarks; accelerates training up to 7.7×; and enables a 3B model trained on 180B tokens to outperform an 8B baseline. This is the first work to realize a multi-factor co-optimized synthetic data generation system, providing a reproducible and scalable pathway to overcome pretraining data constraints.

0 citationsRead paper
Recent publications

Latest Papers

Zephon: Elastic Determinism for Online, Stateful Foundation Model Data Loading Pipelines

Oct 02, 2026

This study addresses the non-deterministic data loading in online stateful data pipelines caused by topology changes, failure recovery, and backend heterogeneity. It proposes Zephon, a system enabling resilient and deterministic data loading. The core innovations include a novel topology-agnostic stream partitioning scheme and an ordered decision serialization mechanism, which achieve constant-overhead checkpoint recovery by persisting only bounded in-flight states. These techniques are integrated with stateless parallel processing, backend abstraction, and incremental checkpointing to form a comprehensive solution. Experiments demonstrate that Zephon sustains high throughput under both text and multimodal workloads while providing online determinism guarantees unattainable by existing approaches.

0 citationsRead paper

DatBench: Discriminative, Faithful, and Efficient VLM Evaluations

Jan 05, 2026arXiv.org

Current evaluation methods for vision-language models (VLMs) commonly suffer from modality unfaithfulness, insufficient discriminative power, and computational inefficiency. This work is the first to systematically articulate three core desiderata for VLM evaluation: faithfulness, discriminability, and efficiency. We establish a high-quality evaluation pipeline by reformulating multiple-choice tasks as generative ones, filtering out samples amenable to blind guessing (up to 70% of instances), and correcting mislabeled examples (42% of cases). Based on this framework, we introduce DatBench-Full, encompassing 33 datasets, along with its highly discriminative subset, DatBench. Our benchmarks maintain discriminative capacity comparable to original benchmarks while achieving an average 13× and up to 50× acceleration in evaluation speed.

1 citationsRead paper

Luxical: High-Speed Lexical-Dense Text Embeddings

Dec 09, 2025

To address the trade-off between speed and flexibility in web-scale text organization, this paper proposes a lightweight “lexical-dense” embedding paradigm. Leveraging knowledge distillation, it transfers semantic capabilities from large language models into a compact architecture that jointly encodes TF-IDF–based sparse lexical features and dense representations via a small ReLU network. The resulting embeddings retain the versatility of dense vectors—supporting retrieval, clustering, classification, and data cleaning—while achieving inference speeds comparable to FastText. Experiments demonstrate 3×–100× higher throughput than neural baselines on document retrieval and LLM data cleaning tasks, with quality matching state-of-the-art embedding models. This work is the first to systematically bridge the efficiency–expressiveness gap between traditional lexical models and Transformer-based embeddings, establishing a scalable, high-performance paradigm for large-scale text preprocessing.

0 citationsRead paper

BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Aug 14, 2025

To address the performance saturation in large language model (LLM) pretraining caused by data bottlenecks, this work proposes BeyondWeb—a framework that systematically investigates the joint impact of model scale, architecture family, and data rewriting strategies on synthetic data quality. It establishes a high-quality synthetic data generation paradigm integrating model feedback, rewrite filtering, and diversity control. Evaluated on trillion-token pretraining, BeyondWeb significantly enhances semantic richness and training efficiency of synthetic data. Experiments show it achieves +5.1 and +2.6 percentage points average accuracy over Cosmopedia and Nemotron-Synth across 14 benchmarks; accelerates training up to 7.7×; and enables a 3B model trained on 180B tokens to outperform an 8B baseline. This is the first work to realize a multi-factor co-optimized synthetic data generation system, providing a reproducible and scalable pathway to overcome pretraining data constraints.

0 citationsRead paper