BinoGen: Scaling egocentric binocular data for embodied visual perception and learning
为了解决收集大规模第一人称双目视觉数据困难的问题,BinoGen通过生成环境和观察者变化来创建室内双目视觉体验,用于改善视觉感知任务。
为了解决收集大规模第一人称双目视觉数据困难的问题,BinoGen通过生成环境和观察者变化来创建室内双目视觉体验,用于改善视觉感知任务。
研究通过结合细胞注释和基因调控信息,提出scKITE模型,以更少的预训练样本提高单细胞基础模型性能。
研究通过在多智能体系统中交换角色匹配的代理来测试代理间的互换性,发现虽然任务得分影响小,但沟通成本显著增加。
本文提出Speech2MaskTrack方法,通过语音识别、运动中心时间定位、掩模跟踪等步骤,解决语音引导的视频对象分割问题。
This study addresses the issues of coarse boundaries and instance adhesion in remote sensing semantic segmentation caused by multi-target mixing within visual tokens. To overcome these limitations, we propose FIRM, a novel method that innovatively introduces intra-token sub-unit mask representations and a lightweight continuous rendering mechanism. By transcending single-label constraints through sub-unit prediction, lookup table transformation, and soft structural field marginalization, FIRM achieves fine-grained segmentation. Extensive experiments demonstrate state-of-the-art performance across five benchmarks. Notably, on the LASER dataset, FIRM attains GIoU/CIoU scores of 70.5/80.5 and improves the EarthReason metric by 3.0 points, significantly enhancing segmentation accuracy in complex scenes.
为了解决收集大规模第一人称双目视觉数据困难的问题,BinoGen通过生成环境和观察者变化来创建室内双目视觉体验,用于改善视觉感知任务。
研究通过结合细胞注释和基因调控信息,提出scKITE模型,以更少的预训练样本提高单细胞基础模型性能。
研究通过在多智能体系统中交换角色匹配的代理来测试代理间的互换性,发现虽然任务得分影响小,但沟通成本显著增加。
本文提出Speech2MaskTrack方法,通过语音识别、运动中心时间定位、掩模跟踪等步骤,解决语音引导的视频对象分割问题。
This study addresses the issues of coarse boundaries and instance adhesion in remote sensing semantic segmentation caused by multi-target mixing within visual tokens. To overcome these limitations, we propose FIRM, a novel method that innovatively introduces intra-token sub-unit mask representations and a lightweight continuous rendering mechanism. By transcending single-label constraints through sub-unit prediction, lookup table transformation, and soft structural field marginalization, FIRM achieves fine-grained segmentation. Extensive experiments demonstrate state-of-the-art performance across five benchmarks. Notably, on the LASER dataset, FIRM attains GIoU/CIoU scores of 70.5/80.5 and improves the EarthReason metric by 3.0 points, significantly enhancing segmentation accuracy in complex scenes.