Institution profile

Indian Institute of Technology Gandhinagar

Academic institutionasia · in
Official website
Research library102linked papers
Opportunities0open roles
Selected work

Representative Papers

Hand Shadow Art: A Differentiable Rendering Perspective

May 27, 2025

This work addresses the problem of computational hand shadow art generation. We propose the first differentiable rendering-based method for 3D hand deformation inversion: given a target 2D shadow image and illumination conditions, it jointly optimizes the geometry and pose of both hands along with lighting parameters to minimize the discrepancy between rendered and target shadows. Our approach integrates neural implicit hand representations, physically grounded shadow modeling, and a gradient-guided co-optimization framework—enabling simultaneous bilateral hand solving and smooth pose interpolation across semantically distinct shadows. Experiments demonstrate stable reconstruction of high-fidelity hand shadows, precise matching of intricate shadow structures, and seamless temporal transitions. This work establishes a novel paradigm and practical toolkit for applying differentiable graphics to digital artistic creation.

3 citationsRead paper

Where To Look? : Causal Tracing of Vision Encoders in VLM

Aug 11, 2026

This study investigates whether visual language models (VLMs) genuinely rely on visual information localized to target regions when generating answers. Employing causal tracing, the authors systematically analyze which token positions in the visual encoder exert causal influence on model outputs and evaluate the models’ capacity to understand visual structure under conditions where appearance cues are absent. The findings reveal that tokens with high causal impact are often distributed outside target regions, challenging the assumption that strong multimodal performance implies spatially localized causal representations. Furthermore, the work demonstrates that prevailing VLMs predominantly depend on superficial appearance cues to infer structural relationships, exposing a significant gap between their abilities to perceive, utilize, and reason about visual structure. This research establishes a causal framework for analyzing how visual information is transformed, preserved, and leveraged in VLMs.

0 citationsRead paper
Recent publications

Latest Papers

Where To Look? : Causal Tracing of Vision Encoders in VLM

Aug 11, 2026

This study investigates whether visual language models (VLMs) genuinely rely on visual information localized to target regions when generating answers. Employing causal tracing, the authors systematically analyze which token positions in the visual encoder exert causal influence on model outputs and evaluate the models’ capacity to understand visual structure under conditions where appearance cues are absent. The findings reveal that tokens with high causal impact are often distributed outside target regions, challenging the assumption that strong multimodal performance implies spatially localized causal representations. Furthermore, the work demonstrates that prevailing VLMs predominantly depend on superficial appearance cues to infer structural relationships, exposing a significant gap between their abilities to perceive, utilize, and reason about visual structure. This research establishes a causal framework for analyzing how visual information is transformed, preserved, and leveraged in VLMs.

0 citationsRead paper

Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation

Aug 06, 2026

Existing multilingual text embedding models typically employ a single training objective across diverse tasks, overlooking the fundamental differences in their optimization requirements. This work proposes the Task-Conditional Flow Matching (TCFM) framework, which introduces a task-conditional mechanism to tailor optimization objectives according to task-specific learning dynamics: leveraging flow matching for translation while designing more suitable objectives for retrieval, classification, and other tasks. The approach further integrates teacher-guided representations with a three-stage curriculum learning strategy to enable stable and efficient multitask adaptation. Evaluated on the Indic Massive Text Embedding Benchmark, TCFM achieves a new state of the art, significantly enhancing embedding quality across a wide range of multilingual tasks and demonstrating strong generalization across different model families.

0 citationsRead paper