🤖 AI Summary
This study addresses the substantial structural discrepancies between dense and sparse prediction tasks in multi-task vision models, as well as the high computational overhead associated with keypoint detection. To this end, this work proposes a unified encoder-decoder architecture that efficiently integrates five visual tasks, including segmentation and depth estimation, through a shared encoder-decoder backbone coupled with lightweight task-specific projection heads. Furthermore, an auxiliary knowledge distillation strategy is introduced to enable top-down keypoint detection within a single forward pass. Evaluated on the COCO dataset, the proposed method achieves state-of-the-art performance with 53.1 PQ and 66.5 mIoU. The resulting framework is both lightweight and computationally efficient, while demonstrating significant complementary gains across the integrated tasks.
📝 Abstract
Generalist multitasking vision models aim to unify multiple vision tasks within a single framework, enabling more efficient and versatile learning. However, handling diverse vision tasks -- spanning dense and sparse predictions -- remains challenging due to their inherently varying output structures. In this paper, we propose AHMAD, a simple yet effective framework for generalist multitask learning that integrates different key vision tasks: semantic segmentation, instance segmentation, depth estimation, keypoint detection, and object detection. Our approach incorporates these five tasks into a unified structure: a shared encoder-decoder with several lightweight task-specific projectors. Under the multitask learning paradigm, we observed a complementary performance gain, achieving a state-of-the-art PQ of 53.1 and an mIoU of 66.5 for COCO-val panoptic and semantic segmentation, respectively. Additionally, for top-down keypoint detection, which typically incurs high computational overhead due to multiple forward passes, we introduce a knowledge distillation-based method that enables a single forward pass over the entire image, greatly improving efficiency. Ultimately, our model delivers a lightweight yet effective generalist multitask learning framework, demonstrating strong performance across five vision tasks.