Institution profile

xAI

Industry researchnorthamerica · us
Official website
Research library16linked papers
Opportunities66open roles
Selected work

Representative Papers

TerraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows

Oct 02, 2026

This study addresses the limitation of existing evaluation metrics in capturing visual inconsistencies that violate physical laws, anatomical structures, or spatial relationships in text-to-image models. To this end, we propose TerraVis, a framework that establishes world consistency as an independent evaluation dimension. TerraVis constructs a hierarchical taxonomy of violations spanning object, interaction, and scene levels, and designs a multi-stage automated detection and scoring pipeline leveraging multimodal large language models. Experimental results demonstrate that TerraVis achieves the highest correlation with human judgments on mainstream benchmarks, effectively revealing significant real-world logical deficiencies that persist even in traditionally high-scoring models.

0 citationsRead paper

Gaming Consensus: Coordinated Manipulation in Crowdsourced Fact-Checking

Jul 02, 2026

This study addresses the vulnerability of crowdsourced fact-checking systems to strategic manipulation, wherein coordinated users exploit voting mechanisms to fabricate cross-ideological consensus. The research demonstrates that “unhelpful” ratings can paradoxically inflate a note’s perceived usefulness, and that fewer than ten adversarial ratings suffice to push 10.7% of low-quality notes above the consensus threshold. To quantify this manipulation cost—a first in the literature—the authors integrate matrix factorization, theoretical modeling, and empirical analysis of historical data. Building on these insights, they propose targeted algorithmic interventions to mitigate such attacks. The proposed defenses have been deployed in X’s Community Notes system, significantly enhancing the robustness of its consensus mechanism against coordinated manipulation.

0 citationsRead paper

DeployBench: Benchmarking LLM Agents for Research Artifact Deployment

Jun 03, 2026

This work addresses the frequent failure of research code deployment due to complex environment configurations, heterogeneous toolchains, system dependencies (e.g., GPU/CUDA), and legacy compatibility issues—challenges inadequately captured by existing benchmarks. To bridge this gap, we introduce DeployBench, the first systematic, multidimensional deployment benchmark encompassing 51 research artifacts across AI/ML, computer systems, and scientific computing. DeployBench evaluates the autonomous deployment capabilities of LLM agents through hidden validation pipelines that reproduce experiments and verify outputs. Built upon the OpenHands framework, our evaluation integrates four state-of-the-art LLMs and executes end-to-end deployment and validation in full system environments. Results reveal that even the best-performing agent achieves a success rate of only 7.8%–51.0%, with 63% of failures attributed to premature self-termination or misaligned validation objectives, exposing critical deficiencies in agents’ judgment of task completion.

0 citationsRead paper

Almost-Orthogonality in Lp Spaces: A Case Study with Grok

May 06, 2026

This study investigates the validity of Carbery’s proposed sharpened form of the multilinear triangle inequality in $L^p$ spaces for $p > 2$. By constructing explicit counterexamples, the work disproves the universal validity of this inequality in the superquadratic regime and establishes that any such estimate can hold only if the exponent $c$ satisfies $c \leq p'$, where $p'$ denotes the conjugate exponent of $p$. Moreover, at the critical exponent $c = p'$, the inequality is proven to hold for all integer $p \geq 2$. In the three-function case, the authors obtain a sharp bound with an optimal exponent that improves upon the result of Carlen–Frank–Lieb. The analysis combines tools from functional analysis, $L^p$ space theory, and carefully designed counterexamples, with key lemmas aided by large language model–assisted derivations.

0 citationsRead paper

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

Apr 13, 2026

This work addresses the limited fine-grained alignment between image patches and textual concepts in current vision-language foundation models. To enhance patch-text alignment, the authors propose a novel pretraining framework featuring several key innovations: patch-level knowledge distillation, an improved iBOT++ masked image modeling objective that incorporates unmasked patches directly into the loss computation, an optimized exponential moving average mechanism, and a multi-granularity synthetic caption sampling strategy. Built upon an efficient dual-encoder architecture, the method achieves state-of-the-art or competitive performance across nine task categories and twenty benchmark datasets, demonstrating its effectiveness in advancing visual representation learning.

0 citationsRead paper
Recent publications

Latest Papers

TerraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows

Oct 02, 2026

This study addresses the limitation of existing evaluation metrics in capturing visual inconsistencies that violate physical laws, anatomical structures, or spatial relationships in text-to-image models. To this end, we propose TerraVis, a framework that establishes world consistency as an independent evaluation dimension. TerraVis constructs a hierarchical taxonomy of violations spanning object, interaction, and scene levels, and designs a multi-stage automated detection and scoring pipeline leveraging multimodal large language models. Experimental results demonstrate that TerraVis achieves the highest correlation with human judgments on mainstream benchmarks, effectively revealing significant real-world logical deficiencies that persist even in traditionally high-scoring models.

0 citationsRead paper

Gaming Consensus: Coordinated Manipulation in Crowdsourced Fact-Checking

Jul 02, 2026

This study addresses the vulnerability of crowdsourced fact-checking systems to strategic manipulation, wherein coordinated users exploit voting mechanisms to fabricate cross-ideological consensus. The research demonstrates that “unhelpful” ratings can paradoxically inflate a note’s perceived usefulness, and that fewer than ten adversarial ratings suffice to push 10.7% of low-quality notes above the consensus threshold. To quantify this manipulation cost—a first in the literature—the authors integrate matrix factorization, theoretical modeling, and empirical analysis of historical data. Building on these insights, they propose targeted algorithmic interventions to mitigate such attacks. The proposed defenses have been deployed in X’s Community Notes system, significantly enhancing the robustness of its consensus mechanism against coordinated manipulation.

0 citationsRead paper

DeployBench: Benchmarking LLM Agents for Research Artifact Deployment

Jun 03, 2026

This work addresses the frequent failure of research code deployment due to complex environment configurations, heterogeneous toolchains, system dependencies (e.g., GPU/CUDA), and legacy compatibility issues—challenges inadequately captured by existing benchmarks. To bridge this gap, we introduce DeployBench, the first systematic, multidimensional deployment benchmark encompassing 51 research artifacts across AI/ML, computer systems, and scientific computing. DeployBench evaluates the autonomous deployment capabilities of LLM agents through hidden validation pipelines that reproduce experiments and verify outputs. Built upon the OpenHands framework, our evaluation integrates four state-of-the-art LLMs and executes end-to-end deployment and validation in full system environments. Results reveal that even the best-performing agent achieves a success rate of only 7.8%–51.0%, with 63% of failures attributed to premature self-termination or misaligned validation objectives, exposing critical deficiencies in agents’ judgment of task completion.

0 citationsRead paper

Almost-Orthogonality in Lp Spaces: A Case Study with Grok

May 06, 2026

This study investigates the validity of Carbery’s proposed sharpened form of the multilinear triangle inequality in $L^p$ spaces for $p > 2$. By constructing explicit counterexamples, the work disproves the universal validity of this inequality in the superquadratic regime and establishes that any such estimate can hold only if the exponent $c$ satisfies $c \leq p'$, where $p'$ denotes the conjugate exponent of $p$. Moreover, at the critical exponent $c = p'$, the inequality is proven to hold for all integer $p \geq 2$. In the three-function case, the authors obtain a sharp bound with an optimal exponent that improves upon the result of Carlen–Frank–Lieb. The analysis combines tools from functional analysis, $L^p$ space theory, and carefully designed counterexamples, with key lemmas aided by large language model–assisted derivations.

0 citationsRead paper

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

Apr 13, 2026

This work addresses the limited fine-grained alignment between image patches and textual concepts in current vision-language foundation models. To enhance patch-text alignment, the authors propose a novel pretraining framework featuring several key innovations: patch-level knowledge distillation, an improved iBOT++ masked image modeling objective that incorporates unmasked patches directly into the loss computation, an optimized exponential moving average mechanism, and a multi-granularity synthetic caption sampling strategy. Built upon an efficient dual-encoder architecture, the method achieves state-of-the-art or competitive performance across nine task categories and twenty benchmark datasets, demonstrating its effectiveness in advancing visual representation learning.

0 citationsRead paper