🤖 AI Summary
This study addresses the estimation of history-dependent conditional prediction revision scales in sequential models—the magnitude by which predictions are updated upon observing new data—a quantity inherently unobservable due to its dependence on unknown conditional means. We systematically evaluate block bootstrap, conditional heteroskedasticity models, state-space filters, O(1) streaming smoothers, and pretrained RNN forget gates across varying structural assumptions and computational budgets. Theoretical and empirical analyses reveal that lagged smoothers are inconsistent under rapid dynamics, while structurally aligned state-space filters can surpass conventional convergence rate limits. Although neural forget gates do not explicitly encode this scale, it can be effectively decoded via linear probing. These findings motivate a practical guideline: “identify structure, match method, choose minimal cost.” In volatility-driven settings, conditional variance models outperform block bootstrap by orders of magnitude in both speed and accuracy; in state-driven scenarios, only structurally matched filters reliably track abrupt changes.
📝 Abstract
The \emph{conditional forecast-revision scale}
$\It=\{\Var(\E[X_{t+1}\mid\F_t]\mid\F_{t-1})\}^{1/2}$ measures the
history-specific size of the forecast update induced by observing $X_t$.
Because it is a conditional second moment built from two unknown conditional
means, it is not directly observed. We study which estimator of $\It$ should
be used under different structural assumptions and computational budgets. The
comparison includes a block bootstrap, a conditional-variance model, a fitted
state-space model, two $O(1)$ streaming smoothers, and the forget gate of an
already-trained recurrent network. An error decomposition separates
one-step-prediction error from conditional-second-moment tracking error. We
show that externally tuned lag-only smoothers can be inconsistent when $\It$
changes at the sampling scale, although they attain the usual $T^{-2/3}$
mean-squared-error rate ($T^{-1/3}$ for $\It$) under slow variation; a correctly
specified state-space estimator escapes
this limit by using the current state. In volatility-driven designs, a cheap
conditional-variance model is more accurate and over one hundred times cheaper
\emph{as a point estimator} than the implemented block bootstrap, whose value
lies in the sampling distribution it provides rather than in point tracking. In state-driven designs,
only the structurally matched filter recovers the fast variation. Read directly, a
trained network's forget gate does not track $\It$ --- though a supervised linear
probe on the full gate vector does, so $\It$ is linearly decodable but not
available for free. These results yield a practical rule: identify
the conditional-second-moment structure, match the estimator to it, and then
choose the least costly adequate method.