Institution profile

Tetras AI

Industry researchnorthamerica · us
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

Rethinking Visual Token Compression for Video Large Language Models: A Simple Yet Strong Baseline

Sep 28, 2026

This study addresses the inference inefficiency of large video models caused by massive visual tokens generated from long videos. We propose SimpleCluster, a training-free method for efficient visual token compression. By leveraging position-aware cross-frame clustering and feature mean representation, this approach substantially reduces token count while effectively preserving spatiotemporal feature structures. Our findings demonstrate that a straightforward feature distribution preservation strategy outperforms more complex compression paradigms. Extensive experiments show that SimpleCluster surpasses existing methods across four benchmarks and three mainstream models, exhibiting remarkable robustness even at an extremely low retention rate of 1%.

0 citationsRead paper

UniCAD: A Unified Benchmark and Universal Model for Multi-Modal Multi-Task CAD

Jun 03, 2026

Existing CAD research typically addresses individual tasks in isolation, lacking a unified benchmark for multimodal multitask learning. This work introduces the first comprehensive multimodal benchmark encompassing point cloud reconstruction, text- and image-to-CAD generation, and CAD-based question answering. Furthermore, we propose UniCAD-MLLM, an end-to-end general-purpose multimodal large language model that, for the first time, integrates textual, visual, sketch, and point cloud inputs within a single unified framework to enable collaborative modeling and understanding across diverse tasks. Evaluated on both the newly introduced UniCAD benchmark and the established Fusion360 dataset, UniCAD-MLLM consistently outperforms existing specialized and multitask approaches, achieving state-of-the-art performance across all evaluated tasks.

0 citationsRead paper

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model

Jul 15, 2025

Multimodal large language models (MLLMs) remain limited in fine-grained image understanding—particularly for deformable object keypoint localization. To address this, we propose the first general-purpose keypoint understanding framework, introducing a novel “identify-then-detect” paradigm and structured chain-of-thought reasoning to achieve unified, cross-scene and cross-category keypoint localization. Our method jointly leverages instruction-driven semantic parsing and pixel-level keypoint regression, trained on a large-scale, multi-category dataset comprising over 500K samples. Evaluated on multiple benchmarks, it achieves state-of-the-art performance, significantly improving localization accuracy and generalization under complex occlusions and diverse object appearances. Moreover, it enhances semantic controllability in human–AI collaborative interaction by enabling precise, instruction-guided keypoint interpretation.

0 citationsRead paper

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

May 29, 2025

To address the challenge of unifying multimodal understanding and generation within a single framework, this paper introduces OpenUni—the first fully open-source, minimalist unified architecture. Its core design modularly couples off-the-shelf multimodal large language models (MLLMs) and diffusion models via learnable queries and a lightweight Transformer connector, activating only 1.1B or 3.1B parameters. Crucially, OpenUni avoids end-to-end training, drastically reducing computational overhead while enabling strong cross-modal synergy. Experiments demonstrate state-of-the-art performance on instruction-aligned image generation tasks and top-tier results across multiple benchmarks—including GenEval, DPG-Bench, and WISE. To foster reproducibility and community advancement, the project releases all model weights, training code, and a high-quality dataset comprising 23 million image-text pairs.

0 citationsRead paper
Recent publications

Latest Papers

Rethinking Visual Token Compression for Video Large Language Models: A Simple Yet Strong Baseline

Sep 28, 2026

This study addresses the inference inefficiency of large video models caused by massive visual tokens generated from long videos. We propose SimpleCluster, a training-free method for efficient visual token compression. By leveraging position-aware cross-frame clustering and feature mean representation, this approach substantially reduces token count while effectively preserving spatiotemporal feature structures. Our findings demonstrate that a straightforward feature distribution preservation strategy outperforms more complex compression paradigms. Extensive experiments show that SimpleCluster surpasses existing methods across four benchmarks and three mainstream models, exhibiting remarkable robustness even at an extremely low retention rate of 1%.

0 citationsRead paper

UniCAD: A Unified Benchmark and Universal Model for Multi-Modal Multi-Task CAD

Jun 03, 2026

Existing CAD research typically addresses individual tasks in isolation, lacking a unified benchmark for multimodal multitask learning. This work introduces the first comprehensive multimodal benchmark encompassing point cloud reconstruction, text- and image-to-CAD generation, and CAD-based question answering. Furthermore, we propose UniCAD-MLLM, an end-to-end general-purpose multimodal large language model that, for the first time, integrates textual, visual, sketch, and point cloud inputs within a single unified framework to enable collaborative modeling and understanding across diverse tasks. Evaluated on both the newly introduced UniCAD benchmark and the established Fusion360 dataset, UniCAD-MLLM consistently outperforms existing specialized and multitask approaches, achieving state-of-the-art performance across all evaluated tasks.

0 citationsRead paper

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model

Jul 15, 2025

Multimodal large language models (MLLMs) remain limited in fine-grained image understanding—particularly for deformable object keypoint localization. To address this, we propose the first general-purpose keypoint understanding framework, introducing a novel “identify-then-detect” paradigm and structured chain-of-thought reasoning to achieve unified, cross-scene and cross-category keypoint localization. Our method jointly leverages instruction-driven semantic parsing and pixel-level keypoint regression, trained on a large-scale, multi-category dataset comprising over 500K samples. Evaluated on multiple benchmarks, it achieves state-of-the-art performance, significantly improving localization accuracy and generalization under complex occlusions and diverse object appearances. Moreover, it enhances semantic controllability in human–AI collaborative interaction by enabling precise, instruction-guided keypoint interpretation.

0 citationsRead paper

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

May 29, 2025

To address the challenge of unifying multimodal understanding and generation within a single framework, this paper introduces OpenUni—the first fully open-source, minimalist unified architecture. Its core design modularly couples off-the-shelf multimodal large language models (MLLMs) and diffusion models via learnable queries and a lightweight Transformer connector, activating only 1.1B or 3.1B parameters. Crucially, OpenUni avoids end-to-end training, drastically reducing computational overhead while enabling strong cross-modal synergy. Experiments demonstrate state-of-the-art performance on instruction-aligned image generation tasks and top-tier results across multiple benchmarks—including GenEval, DPG-Bench, and WISE. To foster reproducibility and community advancement, the project releases all model weights, training code, and a high-quality dataset comprising 23 million image-text pairs.

0 citationsRead paper