Institution profile

Army Research Laboratory

Academic institutionnorthamerica · us
Official website
Research library231linked papers
Opportunities0open roles
Selected work

Representative Papers

Enhancing Vision Language Models with Logic Reasoning for Situational Awareness

Jan 16, 2026IEEE Transactions on Artificial Intelligence

This work addresses the limitations of vision-language models (VLMs) in situational awareness—specifically, their poor recognition of infrequent critical events, insufficient detail capture, and low output reliability—by proposing an enhanced framework that integrates traditional computer vision with explicit logical reasoning. The approach introduces fine-grained event parsing and a logic-guided, intelligent fine-tuning strategy, while also generating interpretable justifications for the first time during inference. This significantly improves both the accuracy of rare-event recognition and the trustworthiness of model outputs. By coupling discriminative capabilities with transparent, traceable reasoning chains, the method not only boosts VLM performance but also provides a verifiable basis for validating or challenging its conclusions.

1 citationsRead paper

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models

May 01, 2025

Current large multimodal models (LMMs) suffer from strong 2D biases and insufficient 3D training data, resulting in severely limited 3D spatial reasoning capabilities. To address this, we propose SpatialLLM—the first systematic framework for enhancing 3D spatial understanding in LMMs. Methodologically, we introduce the first VQA dataset integrating real-world images with explicit 3D orientation relations; design a joint data-architecture-training optimization paradigm comprising 3D-aware probing, dialogue-based data construction, multi-stage fine-tuning, and a plug-and-play spatial relation modeling module. Experiments demonstrate that SpatialLLM outperforms GPT-4o by 8.7% on dedicated 3D spatial reasoning benchmarks and significantly improves geometric understanding and reasoning in complex scenarios such as vehicle collision prediction. Our work establishes a novel paradigm for multimodal 3D cognitive modeling.

1 citationsRead paper

Hierarchical Preference Optimization: Learning to achieve goals via feasible subgoals prediction

Nov 01, 2024arXiv.org

Hierarchical reinforcement learning (HRL) suffers from two key challenges: non-stationarity in high-level policies due to evolving low-level policies, and infeasible sub-goals generated by high-level policies that low-level policies cannot execute. To address these, we propose Hierarchical Preference Optimization (HPO), the first framework to integrate token-level direct preference optimization (DPO) into HRL—without requiring a pretrained reference policy. HPO jointly optimizes high-level goal generation and low-level action selection via a bilevel optimization formulation. We introduce a primitive-regularized DPO loss that mathematically enforces sub-goal feasibility and prevents degenerate solutions. Additionally, maximum entropy regularization is incorporated to enhance exploration robustness. Evaluated on robotic navigation and manipulation tasks, HPO achieves an average 35% performance gain over strong baselines, significantly mitigating both non-stationarity and sub-goal infeasibility. Ablation studies and quantitative analysis comprehensively validate its effectiveness.

1 citationsRead paper
Recent publications

Latest Papers