NewXResearch hub
ExploreLibraryProfile
Account
Sign In
Orion Reblitz-Richardson
Scholar

Orion Reblitz-Richardson

Google Scholar ID: FzCPVDMAAAAJ
Applied Research Scientist, Facebook
machine learninginterpretabilitymodel understanding
Homepage↗Google Scholar↗
Citations & Impact
All-time
Citations
1,923
 
H-index
12
 
i10-index
12
 
Publications
20
 
Co-authors
1
 
Publications
6 items
Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment
2026
Cited
0
Refusal Reads Only a Slice of What the Model Knows: Harm-Keyed Routing and Its Exceptions Across Model Families
2026
Cited
0
Calibrating Interpretability Instruments Before Trusting Their Verdicts
2026
Cited
0
How Language Models Organize and Structure Moral Knowledge
2026
Cited
0
Output Dilution: Redundant but Fragile Representations in MoE Models
2026
Cited
0
When Probing Accuracy Saturates, Fragility Resolves: A Complementary Metric for LLM Pre-Training Analysis
2026
Cited
0