Beyond Emotion Prompts: Fine-Grained Text-to-Image Generation Driven by Valence-Arousal-Dominance
本文提出EMOTRANS方法,通过将心理基础的VAD坐标转换为生成条件,解决图像生成中精细情绪控制的问题。
本文提出EMOTRANS方法,通过将心理基础的VAD坐标转换为生成条件,解决图像生成中精细情绪控制的问题。
本文提出了一种新的无奇点引导向量场方法,能够在保持路径误差动态不变的情况下,使机器人物理速度收敛到预设值,解决了传统方法中速度依赖于路径参数化的问题。
为解决远程传感任务中单体决策框架的不稳定性和错误传播问题,提出HiRS-Agent系统,采用分层多代理架构和强化学习策略优化任务执行。
This study addresses the unknown applicability of traditional cartographic color principles to spatial reasoning in foundation models. By constructing a controlled benchmark and integrating multimodal evaluation, LoRA fine-tuning, and factorial experiments, this work systematically quantifies the impact of color variables on model reasoning for the first time. Results indicate that disordered color sequences and low contrast significantly impair performance, and notably, fine-tuning fails to eliminate this sensitivity. Highlighting the critical roles of sequential color ordering and contrast, this research proposes AI-friendly cartographic design guidelines. These findings provide empirical evidence and methodological guidance for optimizing map understanding capabilities in artificial intelligence systems, bridging the gap between classical cartography and modern vision-language models.
Existing point cloud scene generation methods rely on partial scans as conditioning inputs, leading to a mismatch between training and inference, poor handling of sparsity in distant regions and occluded areas, and limited flexibility in generating scenes without LiDAR observations. To address these limitations, this work proposes a unified generation framework that dispenses with partial scans by predicting density, height, and occupancy masks in bird’s-eye view (BEV) to construct structured point sources. Furthermore, it introduces a teacher–student approximate optimal transport mechanism that learns straighter transport paths for efficient single-step point generation. The approach supports both unconditional and multi-cue conditional generation, achieving state-of-the-art Jensen–Shannon divergence (JSD) and voxel IoU on SemanticKITTI, and the best Coverage score on KITTI-360 under unconditional generation.
本文提出EMOTRANS方法,通过将心理基础的VAD坐标转换为生成条件,解决图像生成中精细情绪控制的问题。
本文提出了一种新的无奇点引导向量场方法,能够在保持路径误差动态不变的情况下,使机器人物理速度收敛到预设值,解决了传统方法中速度依赖于路径参数化的问题。
为解决远程传感任务中单体决策框架的不稳定性和错误传播问题,提出HiRS-Agent系统,采用分层多代理架构和强化学习策略优化任务执行。
This study addresses the unknown applicability of traditional cartographic color principles to spatial reasoning in foundation models. By constructing a controlled benchmark and integrating multimodal evaluation, LoRA fine-tuning, and factorial experiments, this work systematically quantifies the impact of color variables on model reasoning for the first time. Results indicate that disordered color sequences and low contrast significantly impair performance, and notably, fine-tuning fails to eliminate this sensitivity. Highlighting the critical roles of sequential color ordering and contrast, this research proposes AI-friendly cartographic design guidelines. These findings provide empirical evidence and methodological guidance for optimizing map understanding capabilities in artificial intelligence systems, bridging the gap between classical cartography and modern vision-language models.
Existing point cloud scene generation methods rely on partial scans as conditioning inputs, leading to a mismatch between training and inference, poor handling of sparsity in distant regions and occluded areas, and limited flexibility in generating scenes without LiDAR observations. To address these limitations, this work proposes a unified generation framework that dispenses with partial scans by predicting density, height, and occupancy masks in bird’s-eye view (BEV) to construct structured point sources. Furthermore, it introduces a teacher–student approximate optimal transport mechanism that learns straighter transport paths for efficient single-step point generation. The approach supports both unconditional and multi-cue conditional generation, achieving state-of-the-art Jensen–Shannon divergence (JSD) and voxel IoU on SemanticKITTI, and the best Coverage score on KITTI-360 under unconditional generation.