🤖 AI Summary
This paper addresses the reliability of instrumental variable (IV) exogeneity tests under joint weak identification and heteroskedasticity. It systematically compares the Kleibergen–Paap (KP) rank test with the conventional J-test, deriving their asymptotic distributions under heteroskedasticity and weak instruments and complementing theoretical analysis with Monte Carlo simulations and empirical application to a life-cycle consumption model. Results show that the J-test suffers from severe size distortion—over-rejecting the null—whereas the KP test achieves superior size control and robustness. The paper makes two key contributions: first, it recommends replacing the J-test with the KP test—or, alternatively, a heteroskedasticity-robust score test based on limited-information maximum likelihood (LIML)—in settings characterized by weak identification and heteroskedasticity; second, it demonstrates that divergent estimates of the intertemporal elasticity of substitution (EIS) stem primarily from model specification differences rather than instrument invalidity. These findings provide a more reliable inferential framework for assessing IV validity.
📝 Abstract
Exogeneity is key for IV estimators, which can assessed via overidentification (OID) tests. We discuss the Kleibergen-Paap (KP) rank test as a heteroskedasticity-robust OID test and compare to the typical J-test. We derive the heteroskedastic weak-instrument limiting distributions for J and KP as special cases of the robust score test estimated via 2SLS and LIML respectively. Monte Carlo simulations show that KP usually performs better than J, which is prone to severe size distortions. Test size depends on model parameters not consistently estimable with weak instruments, so a conservative approach is recommended. This generalises recommendations to use LIML-based OID tests under homoskedasticity. We then revisit the classic problem of estimating the elasticity of intertemporal substitution (EIS) in lifecycle consumption models. Lagged macroeconomic indicators should provide naturally valid but frequently weak instruments. The literature provides a wide range of estimates for this parameter, and J frequently rejects the null of valid instruments. J often rejects the null whereas KP does not; we suggest that J over-rejects, sometimes severely. We argue that KP-test should be used over the J-test. We also argue that instrument invalidity/misspecification is unlikely the cause of the range of EIS estimates in the literature.