Institution profile

Chubu University

Academic institutionasia · jp
Official website
Research library11linked papers
Opportunities0open roles
Selected work

Representative Papers

PAM-ToD: Plug-and-Play Appearance Modeling for Cross-Time-of-Day 3D Gaussian Splatting

Oct 08, 2026

This study addresses the challenge of adapting pre-trained 3D Gaussian Splatting models to cross-temporal scenes, where maintaining appearance consistency and real-time rendering with limited anchor images remains difficult. To overcome this, the authors propose a lightweight, plug-and-play module that decouples illumination and brightness variations through color scaling and additive terms. By leveraging a simplified imaging model to eliminate albedo dependency and introducing spatial smoothness constraints to guide few-shot learning, the method achieves cross-temporal appearance correction without fixed parameters. Furthermore, this work establishes the CARLA-ToD benchmark dataset. Experimental results demonstrate that, using only single-temporal multi-view anchor images, the proposed approach significantly outperforms baselines in PSNR and LPIPS across both static and dynamic scenes, enabling high-quality, real-time cross-temporal novel view synthesis.

0 citationsRead paper

EviSplat: Preserving Multi-View Evidence in 3D Gaussian Splatting for Open-Vocabulary Segmentation

Sep 28, 2026

This study addresses the loss of critical cues caused by pre-query fusion of multi-view features in open-vocabulary 3D segmentation. To overcome this limitation, we propose a deferred aggregation strategy based on 3D Gaussian Splatting (3DGS). Rather than suppressing information prematurely, our method preserves independent observations from each view as distinct evidence. Specifically, it learns a support distribution for each Gaussian, enabling dynamic, relevance-weighted aggregation of optimal features during the query phase. By integrating class-agnostic instance segmentation with this multi-view evidence preservation mechanism, the proposed approach achieves state-of-the-art performance across multiple benchmark datasets. These results validate the effectiveness of retaining multi-view evidence until query time to facilitate dynamic feature aggregation.

0 citationsRead paper

Structured Pruning of Large Language Models via Power Transformation and Sign-Preserving Score Aggregation with Adaptive Feature Retention

Jul 08, 2026

This work addresses the challenges of transferring adaptive feature retention (AFR) from unstructured to structured pruning, which include heterogeneous pruning score distributions, loss of sign information, and outlier interference. To bridge this gap, the authors propose a unified structured pruning framework that introduces power transformation to align score distributions, designs a sign-preserving aggregation mechanism to maintain consistent optimization directions, and incorporates a percentile-based outlier removal strategy. This approach effectively narrows the performance gap between structured and unstructured pruning, achieving accuracy on Llama-3-8B, Vicuna-v1.5-13B, and LLaVA-v1.5-13B models that closely matches unstructured pruning while delivering substantial real-world inference speedups.

0 citationsRead paper

SPG: Sparse-Projected Guides with Sparse Autoencoders for Zero-Shot Anomaly Detection

Apr 03, 2026

This work addresses zero-shot anomaly detection and segmentation—identifying anomalies on unseen categories without target-domain adaptation—by proposing SPG, a prompt-free framework. Departing from conventional prompting strategies, SPG introduces sparse projection guidance for the first time: leveraging frozen foundation models (e.g., DINOv3 or OpenCLIP ViT-L/14@336px), it learns sparse guidance coefficients in two stages—first training a sparse autoencoder (SAE) and then optimizing only the guidance coefficients to produce normal and anomaly guidance vectors. These coefficients map sparsely to a small set of SAE dictionary atoms, revealing both category-agnostic and category-specific factors. Experiments demonstrate that SPG achieves state-of-the-art image-level detection performance under cross-dataset zero-shot settings on MVTec AD and VisA, and sets new records in pixel-level segmentation AUROC, particularly with the DINOv3 backbone.

0 citationsRead paper

OD-RASE: Ontology-Driven Risk Assessment and Safety Enhancement for Autonomous Driving

Mar 06, 2026

This study addresses the limitations of existing road infrastructure, which is predominantly designed for human drivers and often fails to ensure safety in rare or complex scenarios encountered by autonomous vehicles, with safety enhancements typically lagging behind accident occurrences. To bridge this gap, the authors propose OD-RASE, a novel framework that integrates domain-specific traffic ontologies with large vision-language models (LVLMs) to accurately identify road structures contributing to accidents and automatically generate interpretable infrastructure improvement recommendations. The approach combines ontology-driven data filtering, LVLM-based reasoning, and diffusion-model-enhanced visualization, and introduces the first annotated dataset for this task. Experimental results demonstrate that OD-RASE achieves high precision in predicting high-risk road configurations and produces actionable retrofitting strategies, significantly enhancing the proactive safety of autonomous driving systems.

0 citationsRead paper
Recent publications

Latest Papers

PAM-ToD: Plug-and-Play Appearance Modeling for Cross-Time-of-Day 3D Gaussian Splatting

Oct 08, 2026

This study addresses the challenge of adapting pre-trained 3D Gaussian Splatting models to cross-temporal scenes, where maintaining appearance consistency and real-time rendering with limited anchor images remains difficult. To overcome this, the authors propose a lightweight, plug-and-play module that decouples illumination and brightness variations through color scaling and additive terms. By leveraging a simplified imaging model to eliminate albedo dependency and introducing spatial smoothness constraints to guide few-shot learning, the method achieves cross-temporal appearance correction without fixed parameters. Furthermore, this work establishes the CARLA-ToD benchmark dataset. Experimental results demonstrate that, using only single-temporal multi-view anchor images, the proposed approach significantly outperforms baselines in PSNR and LPIPS across both static and dynamic scenes, enabling high-quality, real-time cross-temporal novel view synthesis.

0 citationsRead paper

EviSplat: Preserving Multi-View Evidence in 3D Gaussian Splatting for Open-Vocabulary Segmentation

Sep 28, 2026

This study addresses the loss of critical cues caused by pre-query fusion of multi-view features in open-vocabulary 3D segmentation. To overcome this limitation, we propose a deferred aggregation strategy based on 3D Gaussian Splatting (3DGS). Rather than suppressing information prematurely, our method preserves independent observations from each view as distinct evidence. Specifically, it learns a support distribution for each Gaussian, enabling dynamic, relevance-weighted aggregation of optimal features during the query phase. By integrating class-agnostic instance segmentation with this multi-view evidence preservation mechanism, the proposed approach achieves state-of-the-art performance across multiple benchmark datasets. These results validate the effectiveness of retaining multi-view evidence until query time to facilitate dynamic feature aggregation.

0 citationsRead paper

Structured Pruning of Large Language Models via Power Transformation and Sign-Preserving Score Aggregation with Adaptive Feature Retention

Jul 08, 2026

This work addresses the challenges of transferring adaptive feature retention (AFR) from unstructured to structured pruning, which include heterogeneous pruning score distributions, loss of sign information, and outlier interference. To bridge this gap, the authors propose a unified structured pruning framework that introduces power transformation to align score distributions, designs a sign-preserving aggregation mechanism to maintain consistent optimization directions, and incorporates a percentile-based outlier removal strategy. This approach effectively narrows the performance gap between structured and unstructured pruning, achieving accuracy on Llama-3-8B, Vicuna-v1.5-13B, and LLaVA-v1.5-13B models that closely matches unstructured pruning while delivering substantial real-world inference speedups.

0 citationsRead paper

SPG: Sparse-Projected Guides with Sparse Autoencoders for Zero-Shot Anomaly Detection

Apr 03, 2026

This work addresses zero-shot anomaly detection and segmentation—identifying anomalies on unseen categories without target-domain adaptation—by proposing SPG, a prompt-free framework. Departing from conventional prompting strategies, SPG introduces sparse projection guidance for the first time: leveraging frozen foundation models (e.g., DINOv3 or OpenCLIP ViT-L/14@336px), it learns sparse guidance coefficients in two stages—first training a sparse autoencoder (SAE) and then optimizing only the guidance coefficients to produce normal and anomaly guidance vectors. These coefficients map sparsely to a small set of SAE dictionary atoms, revealing both category-agnostic and category-specific factors. Experiments demonstrate that SPG achieves state-of-the-art image-level detection performance under cross-dataset zero-shot settings on MVTec AD and VisA, and sets new records in pixel-level segmentation AUROC, particularly with the DINOv3 backbone.

0 citationsRead paper

OD-RASE: Ontology-Driven Risk Assessment and Safety Enhancement for Autonomous Driving

Mar 06, 2026

This study addresses the limitations of existing road infrastructure, which is predominantly designed for human drivers and often fails to ensure safety in rare or complex scenarios encountered by autonomous vehicles, with safety enhancements typically lagging behind accident occurrences. To bridge this gap, the authors propose OD-RASE, a novel framework that integrates domain-specific traffic ontologies with large vision-language models (LVLMs) to accurately identify road structures contributing to accidents and automatically generate interpretable infrastructure improvement recommendations. The approach combines ontology-driven data filtering, LVLM-based reasoning, and diffusion-model-enhanced visualization, and introduces the first annotated dataset for this task. Experimental results demonstrate that OD-RASE achieves high precision in predicting high-risk road configurations and produces actionable retrofitting strategies, significantly enhancing the proactive safety of autonomous driving systems.

0 citationsRead paper