Institution profile

Pinscreen

Industry researchnorthamerica · us
Official website
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

LOCI: Spatial Linear Memory for Streaming World Models

Sep 30, 2026

This study addresses the challenges of low scene revisit fidelity and memory scaling with sequence length in video world models by proposing a hybrid spatial memory architecture. The method pioneers the integration of viewpoint-conditioned read-write operations into recurrent linear attention, alternating camera geometric projections with full-attention blocks and incorporating key-value cache management to enable efficient long-range memory retrieval. Experimental results demonstrate that this architecture significantly improves content reproduction fidelity on the MIND benchmark while reducing peak VRAM consumption by approximately 30%. Furthermore, it supports streaming generation of long videos under constant memory constraints.

0 citationsRead paper

InverseDraping: Recovering Sewing Patterns from 3D Garment Surfaces via BoxMesh Bridging

Apr 03, 2026

This work addresses the highly ill-posed inverse mapping from 3D garments to 2D sewing patterns, which is complicated by geometric–structural coupling induced by wrinkles. To resolve this challenge, the authors propose a two-stage framework that introduces BoxMesh as a structured intermediate representation to explicitly decouple panel intrinsic geometry, stitching topology, and wrinkle deformation. In the first stage, BoxMesh is reconstructed from the 3D garment; in the second, a geometry-driven and semantics-aware autoregressive model parses the BoxMesh into parametric patterns, incorporating physical constraints and supporting variable-length sequence generation. Evaluated on the GarmentCodeData benchmark, the method achieves state-of-the-art performance and demonstrates strong generalization to real-world scans and single-view images.

0 citationsRead paper

HairOrbit: Multi-view Aware 3D Hair Modeling from Single Portraits

Apr 03, 2026

This work addresses the challenge of reconstructing hair-strand-level 3D hair models from a single portrait image, where maintaining realistic detail in occluded regions remains difficult. The authors reformulate the task as a calibrated multi-view reconstruction problem and introduce, for the first time, 3D priors derived from video generation models. They propose a two-stage hair strand growing algorithm driven by a hybrid implicit field, incorporating a neural orientation extractor trained on sparse real-world annotations. This approach significantly outperforms existing methods in both visible and occluded regions, achieving high-fidelity, geometrically consistent hair reconstructions across diverse hairstyles at the individual strand level.

0 citationsRead paper

Steering Video Diffusion Transformers with Massive Activations

Mar 18, 2026

This work proposes Structured Activation Steering (STAS), a training-free, self-guided method that leverages the structured distribution of massive activations in video diffusion Transformers. By uncovering, for the first time, the hierarchical spatiotemporal patterns of massive activations within temporally chunked latent spaces, STAS selectively modulates activation values at critical positions to significantly enhance both visual quality and temporal coherence of generated videos with minimal computational overhead. Extensive experiments demonstrate consistent improvements across diverse text-to-video diffusion models, highlighting STAS’s efficiency, robustness, and broad applicability without requiring model retraining or architectural modifications.

0 citationsRead paper

Spatio-Temporal Garment Reconstruction Using Diffusion Mapping via Pattern Coordinates

Feb 27, 2026

High-fidelity 3D reconstruction of loose garments from monocular images or videos remains challenging. This work proposes a unified framework for both static and dynamic clothing reconstruction, leveraging an Implicit Sewing Pattern (ISP) in UV space to encode garment shape priors. By integrating a spatiotemporal diffusion mechanism with analytical projection constraints, the method establishes precise correspondences among pixels, UV coordinates, and 3D geometry. At inference time, a guidance strategy is introduced to enforce cross-frame consistency and plausibly complete occluded regions. Experiments demonstrate that the approach generalizes robustly to real-world imagery, significantly outperforming existing methods in reconstructing fine geometric details. Furthermore, the reconstructed garments support downstream applications such as texture editing, garment retargeting, and animation.

0 citationsRead paper
Recent publications

Latest Papers

LOCI: Spatial Linear Memory for Streaming World Models

Sep 30, 2026

This study addresses the challenges of low scene revisit fidelity and memory scaling with sequence length in video world models by proposing a hybrid spatial memory architecture. The method pioneers the integration of viewpoint-conditioned read-write operations into recurrent linear attention, alternating camera geometric projections with full-attention blocks and incorporating key-value cache management to enable efficient long-range memory retrieval. Experimental results demonstrate that this architecture significantly improves content reproduction fidelity on the MIND benchmark while reducing peak VRAM consumption by approximately 30%. Furthermore, it supports streaming generation of long videos under constant memory constraints.

0 citationsRead paper

InverseDraping: Recovering Sewing Patterns from 3D Garment Surfaces via BoxMesh Bridging

Apr 03, 2026

This work addresses the highly ill-posed inverse mapping from 3D garments to 2D sewing patterns, which is complicated by geometric–structural coupling induced by wrinkles. To resolve this challenge, the authors propose a two-stage framework that introduces BoxMesh as a structured intermediate representation to explicitly decouple panel intrinsic geometry, stitching topology, and wrinkle deformation. In the first stage, BoxMesh is reconstructed from the 3D garment; in the second, a geometry-driven and semantics-aware autoregressive model parses the BoxMesh into parametric patterns, incorporating physical constraints and supporting variable-length sequence generation. Evaluated on the GarmentCodeData benchmark, the method achieves state-of-the-art performance and demonstrates strong generalization to real-world scans and single-view images.

0 citationsRead paper

HairOrbit: Multi-view Aware 3D Hair Modeling from Single Portraits

Apr 03, 2026

This work addresses the challenge of reconstructing hair-strand-level 3D hair models from a single portrait image, where maintaining realistic detail in occluded regions remains difficult. The authors reformulate the task as a calibrated multi-view reconstruction problem and introduce, for the first time, 3D priors derived from video generation models. They propose a two-stage hair strand growing algorithm driven by a hybrid implicit field, incorporating a neural orientation extractor trained on sparse real-world annotations. This approach significantly outperforms existing methods in both visible and occluded regions, achieving high-fidelity, geometrically consistent hair reconstructions across diverse hairstyles at the individual strand level.

0 citationsRead paper

Steering Video Diffusion Transformers with Massive Activations

Mar 18, 2026

This work proposes Structured Activation Steering (STAS), a training-free, self-guided method that leverages the structured distribution of massive activations in video diffusion Transformers. By uncovering, for the first time, the hierarchical spatiotemporal patterns of massive activations within temporally chunked latent spaces, STAS selectively modulates activation values at critical positions to significantly enhance both visual quality and temporal coherence of generated videos with minimal computational overhead. Extensive experiments demonstrate consistent improvements across diverse text-to-video diffusion models, highlighting STAS’s efficiency, robustness, and broad applicability without requiring model retraining or architectural modifications.

0 citationsRead paper

Spatio-Temporal Garment Reconstruction Using Diffusion Mapping via Pattern Coordinates

Feb 27, 2026

High-fidelity 3D reconstruction of loose garments from monocular images or videos remains challenging. This work proposes a unified framework for both static and dynamic clothing reconstruction, leveraging an Implicit Sewing Pattern (ISP) in UV space to encode garment shape priors. By integrating a spatiotemporal diffusion mechanism with analytical projection constraints, the method establishes precise correspondences among pixels, UV coordinates, and 3D geometry. At inference time, a guidance strategy is introduced to enforce cross-frame consistency and plausibly complete occluded regions. Experiments demonstrate that the approach generalizes robustly to real-world imagery, significantly outperforming existing methods in reconstructing fine geometric details. Furthermore, the reconstructed garments support downstream applications such as texture editing, garment retargeting, and animation.

0 citationsRead paper