"Dear LLaVA, Please Drive": A Depth-Aware Vision-Language Agent for Closed-Loop Robotic Control
本文提出了一种基于指令学习的参数高效微调方法,使预训练的视觉-语言模型能够仅使用立体深度观测生成无碰撞路径,实现自然语言驱动的机器人控制。
本文提出了一种基于指令学习的参数高效微调方法,使预训练的视觉-语言模型能够仅使用立体深度观测生成无碰撞路径,实现自然语言驱动的机器人控制。
本文提出AI暴露度和韧性双维度评估框架,用以解决AI对软件及基于软件的商业模式影响评估问题。
本文通过引入基于LiDAR条件的扩散模型,利用2D基础模型生成的伪标签进行训练,解决了将丰富的2D先验知识转移到稀疏3D LiDAR数据中的难题。
This work addresses the lack of accessible tools for genealogists and local historians to contribute to and discover architectural historical data by proposing a fully client-side, open collaborative platform that requires no custom backend. The platform uniquely integrates Wikidata’s time-qualified property model with OpenHistoricalMap’s historical building footprints, enabling structured crowdsourcing through an interactive map interface. Leveraging public APIs from both services, it performs authentication and read/write operations entirely in the browser, storing all contributions under the CC0 license in a public database. The system supports visual editing and querying of genealogically relevant information—including building occupants, owners, and address changes—thereby significantly enhancing the accessibility of open historical data and fostering greater community engagement.
This work addresses the challenge of coexisting structural ambiguity—arising from non-injectivity of the forward model—and probabilistic ambiguity—stemming from parameter uncertainty—in stochastic inverse problems involving nonlinear parameter dependencies and observational uncertainties. To tackle this, the authors propose a unified modeling framework that couples both types of ambiguity for the first time by representing parameter uncertainty through a mixture of probability densities and formulating a Bayesian inversion-based posterior inference algorithm. The approach is validated on one- and two-dimensional quadratic forward models, demonstrating its ability to accurately resolve probabilistic ambiguity, visualize residual structural ambiguity, and handle both finite and infinite solution sets. Furthermore, it reveals the interaction mechanisms between the two forms of ambiguity within the posterior distribution.
本文提出了一种基于指令学习的参数高效微调方法,使预训练的视觉-语言模型能够仅使用立体深度观测生成无碰撞路径,实现自然语言驱动的机器人控制。
本文提出AI暴露度和韧性双维度评估框架,用以解决AI对软件及基于软件的商业模式影响评估问题。
本文通过引入基于LiDAR条件的扩散模型,利用2D基础模型生成的伪标签进行训练,解决了将丰富的2D先验知识转移到稀疏3D LiDAR数据中的难题。
This work addresses the lack of accessible tools for genealogists and local historians to contribute to and discover architectural historical data by proposing a fully client-side, open collaborative platform that requires no custom backend. The platform uniquely integrates Wikidata’s time-qualified property model with OpenHistoricalMap’s historical building footprints, enabling structured crowdsourcing through an interactive map interface. Leveraging public APIs from both services, it performs authentication and read/write operations entirely in the browser, storing all contributions under the CC0 license in a public database. The system supports visual editing and querying of genealogically relevant information—including building occupants, owners, and address changes—thereby significantly enhancing the accessibility of open historical data and fostering greater community engagement.
This work addresses the challenge of coexisting structural ambiguity—arising from non-injectivity of the forward model—and probabilistic ambiguity—stemming from parameter uncertainty—in stochastic inverse problems involving nonlinear parameter dependencies and observational uncertainties. To tackle this, the authors propose a unified modeling framework that couples both types of ambiguity for the first time by representing parameter uncertainty through a mixture of probability densities and formulating a Bayesian inversion-based posterior inference algorithm. The approach is validated on one- and two-dimensional quadratic forward models, demonstrating its ability to accurately resolve probabilistic ambiguity, visualize residual structural ambiguity, and handle both finite and infinite solution sets. Furthermore, it reveals the interaction mechanisms between the two forms of ambiguity within the posterior distribution.