NewXResearch hub
ExploreLibraryProfile
Account
Sign In
Constantin Venhoff
Scholar

Constantin Venhoff

Google Scholar ID: kzHpG-EAAAAJ
University of Oxford
AIMLInterpretabilityMechanistic InterpretabilityAI Alignment
Google Scholar↗
Citations & Impact
All-time
Citations
28
 
H-index
2
 
i10-index
1
 
Publications
8
 
Co-authors
9
list available
Publications
12 items
Weight Oracles: Reading Neural Network Weights with Language Models
2026
Cited
0
Multimodal Model Diffing for Feature Discovery and Control
2026
Cited
0
Probing the Misaligned Thinking Process of Language Models
2026
Cited
0
Towards Understanding Multimodal Fine-Tuning: Spatial Features
2026
Cited
0
Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive Learning
2026
Cited
0
Too Late to Recall: Explaining the Two-Hop Problem in Multimodal Knowledge Retrieval
2025
Cited
0
Base Models Know How to Reason, Thinking Models Learn When
2025
Cited
0
Towards Mechanistic Defenses Against Typographic Attacks in CLIP
2025
Cited
0
Co-authors
8 total
Philip Torr
Philip Torr
Professor, University of Oxford
Neel Nanda
Neel Nanda
Mechanistic Interpretability Team Lead, Google DeepMind
Ashkan Khakzar
Ashkan Khakzar
University of Oxford
Bernhard Rumpe
Bernhard Rumpe
RWTH Aachen University
Christian Schroeder de Witt
Christian Schroeder de Witt
University of Oxford
Iván Arcuschin
Iván Arcuschin
Independent Researcher
Arthur Conmy
Arthur Conmy
Google DeepMind
Xingyi Yang
Xingyi Yang
Assistant Professor, The Hong Kong Polytechnic University