Score
Designs, builds, or analyzes multi-task perception models and inference heads that jointly produce object-level detections (e.g., boxes or centers) and region- or pixel-level segmentation maps, often together with class labels or enhancement outputs. Work in this skill focuses on architectures and fusion strategies that share representations and inject cross-task cues (semantic or motion) to provide unified perception outputs for downstream decision-making.
Existing dense-prediction multi-task learning (MTL) methods primarily model cross-task relationships in 2D image space, lacking 3D geometric awareness—leading to unstructured features and limited scene understanding. To address this, we propose the lightweight Cross-view Module (CvM), the first approach to explicitly incorporate 3D-aware cross-view correlations into MTL frameworks. CvM constructs a cost volume to enforce geometric consistency across views, fuses features from multi-task encoders, and supports both single- and multi-view inputs within a generic architecture. Evaluated on NYUv2 and PASCAL-Context, our method achieves significant improvements in joint semantic segmentation and depth estimation performance. Results demonstrate that geometric consistency provides effective inductive bias for multi-task representation learning. This work establishes a new paradigm for 3D-aware dense-prediction MTL, bridging geometric reasoning with task synergy in a principled, scalable manner.
Existing video understanding methods often rely on modality-specific architectures or parameters, hindering simultaneous cross-modal generalization and multi-task synergy while neglecting modality distribution shifts and task representation heterogeneity. To address this, we propose the first universal framework for video object tracking and segmentation applicable to arbitrary modalities. Our approach introduces a decoupled Mixture-of-Experts (DeMoE) mechanism that explicitly separates cross-modal shared knowledge from task-specific representations. We further design a unified instance encoder, multimodal feature alignment module, and task-aware tracking decoder to enable joint multimodal–multitask optimization. Evaluated on 18 mainstream benchmarks, our method achieves state-of-the-art performance, significantly improving cross-modal transferability and multitask generalization. This work establishes a new paradigm for universal visual modeling.
Existing referring segmentation datasets focus on single-object, object-level understanding and lack support for fine-grained semantic reasoning over multiple objects and their constituent parts. Method: We introduce the Multi-Object Multi-Granularity Referring Segmentation (MMR) task and present the first large-scale benchmark—comprising 194K implicit instructions—that jointly supports object- and part-level recognition and cross-object relational modeling. We formally define the task paradigm, propose hierarchical annotation protocols and layered prompt engineering, and design a lightweight multi-head decoding head for end-to-end joint segmentation. Results: Experiments on MMR reveal a >32% accuracy drop in part-level recognition for current state-of-the-art models, confirming the benchmark’s diagnostic utility. Our method significantly outperforms baselines, establishing a new foundation for fine-grained, embodied visual-language reasoning.
Existing visual representation learning methods struggle to simultaneously achieve global semantic alignment and local spatial precision. To address this challenge, this work proposes MTV, a multi-task vision pretraining framework that systematically integrates vision-language contrastive learning, self-supervised learning, and dense spatial supervision within a shared backbone architecture. The framework leverages foundation models—including CLIP, MAE, and DINO—together with Depth Anything V2 and OWLv2 to generate high-quality dense pseudo-labels, thereby eliminating the need for manual annotation. Through joint optimization of these complementary tasks, MTV uncovers both synergistic and interfering interactions among them, significantly enhancing fine-grained spatial reasoning while preserving robust global semantic understanding. This approach enables scalable, general-purpose visual encoders that achieve a balanced “best-of-both-worlds” representation.
While SAM and SAM 2 excel at segmenting context-agnostic objects (e.g., persons, vehicles), they exhibit significant limitations on context-dependent (CD) concepts—such as visual saliency, camouflaged objects, industrial defects, and medical lesions—primarily due to insufficient global-local semantic co-modeling. Method: We introduce the first comprehensive CD benchmark spanning 11 concept categories across natural, medical, and industrial domains, incorporating 2D/3D images and videos. We propose a unified evaluation framework supporting human annotation, automated metrics, and self-prompted interaction, augmented with prompt robustness testing and SAM 2’s in-context learning analysis. Our method further incorporates multi-granularity prompt generation, cross-modal self-prompting, and context-aware evaluation metrics. Contribution/Results: Experiments reveal fundamental architectural bottlenecks in SAM-series models, delivering the first quantitative analysis of CD segmentation performance and establishing an empirical foundation for designing SAM 3.
This work proposes a unified vision model that transcends task-specific architectures traditionally required in computer vision by formulating diverse visual tasks—including detection, segmentation, and geometric prediction—as multimodal generation problems. The model is driven solely by natural language instructions (optionally augmented with visual prompts) to produce text, images, or hybrid outputs from a single architecture, eliminating the need for specialized heads or modular designs. It introduces the SenseNova-Vision Corpus, a large-scale dataset of vision-language instruction-response pairs, and leverages off-the-shelf pretrained multimodal models refined through instruction tuning and joint multimodal generation strategies. This end-to-end framework supports compositional, language-defined tasks and achieves performance on par with or superior to specialized systems across structured understanding, dense geometric prediction, segmentation, and multi-view geometry benchmarks.
A fundamental paradigm gap exists between SAM2 (prompt-driven) and SAM3 (concept-driven) segmentation, manifesting in divergent architectural designs, training data distributions, optimization objectives, and evaluation logics. Method: We propose a “prompt → concept” paradigm shift framework featuring a unified vision-language architecture that integrates a vision-language encoder, a geometry/exemplar encoder, a DETR-style decoder, object queries, and a Mixture-of-Experts module for ambiguity resolution—trained end-to-end on open-vocabulary annotated data. Contribution/Results: We establish the first evaluation benchmark specifically designed for concept-driven segmentation, quantifying the technical gap between prompt- and concept-based approaches. Our framework advances image segmentation beyond pixel-level prompt responsiveness toward semantic-aware, generalizable, and reasoning-capable segmentation—paving the way for a new generation of foundation models in visual understanding.
Existing 3D vision approaches typically address tasks such as depth estimation, novel view synthesis, and object manipulation in isolation, lacking a unified representation and cross-task transferability. This work proposes the Three-World Model (3WM), a probabilistic graphical framework that represents multimodal scene elements as graph nodes and supports diverse 3D understanding and interaction tasks through composable conditional inference paths. By leveraging only task-specific prompts, 3WM achieves zero-shot generalization without requiring task-specific training or fine-tuning. To our knowledge, this is the first method to unify multiple 3D tasks within a single model, attaining state-of-the-art performance in both novel view synthesis and 3D object manipulation while demonstrating strong geometric consistency, controllability, and robustness in real-world scenarios.
Existing unified multimodal models are typically evaluated in isolation on either visual understanding or generation tasks, lacking assessment of semantic consistency across tasks. To address this gap, this work proposes XTC-Bench, a novel evaluation framework that introduces the Continuous Cross-Task Agreement (CCTA) metric to disentangle a model’s internal consistency from its individual task performance at the atomic fact level. The framework leverages structured scene graphs to align comprehension queries with generation prompts and establishes the first reproducible, model-agnostic benchmark for cross-task consistency through fine-grained matching of objects, attributes, and relationships. Experiments across nine state-of-the-art models reveal that high task accuracy does not guarantee strong cross-task consistency, which is primarily influenced by the degree of coupling in cross-modal learning objectives.