Scholar
Cong Wei
Google Scholar ID: y1d5C5YAAAAJ
University of Waterloo
Reasoning
Diffusion
Efficiency
Follow
Homepage
↗
Google Scholar
↗
Citations & Impact
All-time
Citations
2,044
H-index
10
i10-index
11
Publications
14
Co-authors
10
list available
Contact
Email
congwei1230@gmail.com
Twitter
Open ↗
GitHub
Open ↗
LinkedIn
Open ↗
Publications
18 items
InfoAgent: Traceable Generation and Repair of Evidence-Grounded Infographics
2026
Cited
0
PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation
2026
Cited
0
VGI-BENCH: Probing Visual Intelligence in Video Generation Models
2026
Cited
0
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
2026
Cited
0
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation
2026
Cited
0
Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth
2026
Cited
0
RewardHarness: Self-Evolving Agentic Post-Training
2026
Cited
0
Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction
2026
Cited
0
Load more
Resume
Academic Achievements
- MoCha: Towards Movie-Grade Talking Character Synthesis, NeurIPS 2025 (Spotlight Presentation)
- UniVideo: Unified Understanding, Generation, and Editing for Videos, Arxiv 2025
- OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision, ICLR 2025
- Sparsifiner: Learning Sparse Instance-Dependent Attention for Efficient Vision Transformers, CVPR 2023
- UniIR: Training and Benchmarking Universal Multimodal Information Retrievers, ECCV 2024 (Oral Presentation)
- AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks, TMLR 2024 (TMLR Reproducibility Certification)
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI, CVPR 2024 (Oral Presentation, Best Paper Finalist)
Research Experience
- Kuaishou Technology KlingAI, May 2025 - Present, Research Scientist Intern
- Meta GenAI, US, Oct 2024 - Apr 2025, Research Scientist Intern
- ModiFace, Canada, May 2022 - Nov 2022, Machine Learning Researcher Intern
- Vector Institute, Canada, Sep 2020 - Sep 2021, Undergraduate Researcher
Education
- University of Waterloo, Canada
- PhD in Computer Science, May 2023 - Present, Advisor: Wenhu Chen
- University of Toronto, Canada
- Master of Science in Applied Computing, Sep 2021 - Jun 2023, Advisor: Florian Shkurti
- Honours Bachelor of Science, Sep 2017 - May 2021, Majors: Computer Science, Statistics, Minor: Mathematics, Advisors: David Duvenaud
- Vector Institute, Undergraduate Researcher, Advisors: David Duvenaud and Gennady Pekhimenko, Sep 2020 - Sep 2021
Background
- Research Interests: Video generation and multi-modal models
- Field: Computer Science
- Brief Introduction: Building unified models to scale up data usage. Previously, did research on sparse attention.
Co-authors
8 total
Wenhu Chen
Assistant Professor at University of Waterloo
Ge Zhang
M-A-P, Bytedance, University of Waterloo
Xiang Yue
Carnegie Mellon University
Yang Chen
Research Scientist, NVIDIA
Alan Ritter
Georgia Institute of Technology
Jie Fu
Shanghai AI Lab
Florian Shkurti
Assistant Professor, Computer Science, University of Toronto
Graham Taylor
University of Guelph and Vector Institute for Artificial Intelligence