Scholar
Jitesh Jain
Google Scholar ID: nygnfNwAAAAJ
Georgia Tech
Image Segmentation
Multimodal Reasoning
Computer Vision
Follow
Homepage
↗
Google Scholar
↗
Citations & Impact
All-time
Citations
1,022
H-index
7
i10-index
7
Publications
12
Co-authors
10
list available
Publications
6 items
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
2026
Cited
6
SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
2025
Cited
0
AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory
2025
Cited
0
Person Recognition at Altitude and Range: Fusion of Face, Body Shape and Gait
2025
Cited
0
Slow-Fast Architecture for Video Multi-Modal Large Language Models
2025
Cited
0
OLA-VLM: Elevating Visual Perception in Multimodal LLMs with Auxiliary Embedding Distillation
arXiv.org · 2024
Cited
2
Co-authors
8 total
Humphrey Shi
Georgia Tech | UIUC || ...
Jianwei Yang
Research Scientist, Meta SuperIntelligence Lab
Zilong Huang
ByteDance Inc.
Ning Yu
Netflix Eyeline Studios
Yuqian Zhou
Senior Research Scientist at Adobe Research
Zhengyuan Yang
Principal Researcher, Microsoft
Jianfeng Gao
Microsoft Research, Redmond
Anil K. Jain
Michigan State University