Representable but Unlearned: Encoding Rank and the Interaction-Prediction Floor

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue whereby input encoding constrains predictors from jointly reproducing contrasts, yielding unlearnable interaction errors. To resolve this, we construct a reachable contrast space by computing encoder equivalence classes, enabling the derivation of empirical error lower bounds for arbitrary decoders without training. Furthermore, we propose a rank-based unsupervised diagnostic method that disentangles encoding capacity from actual model performance, revealing rank deficiency induced by feature masking. By integrating linear algebra, graph neural networks, and contrast projection techniques, this work establishes an error lower bound of 0.009980 on siRNA data—accounting for 14.6% of the fitting error—thereby empirically demonstrating the fundamental constraints imposed by input encoding on learnability.
📝 Abstract
Input encodings can restrict which measured contrasts a predictor can jointly reproduce, even when no single contrast is forced to vanish. We compute the attainable contrast space from an encoder's equivalence classes and a fixed contrast design, without labels, loss, or a fitted model; projecting the recorded contrasts onto that space gives an empirical error floor for any unrestricted decoder on those classes. On a 140-rectangle siRNA interaction panel, a graph neural network's training-only feature mask merges 165 endpoint states into 90 classes and cuts the rank of the 140 interaction contrasts to 72. The resulting floor is 0.009980, which is 14.6% of the fitted model's interaction squared error; the fitted model reaches 0.068335, slightly worse than a control predicting no interaction at all. A minimum of three restored chemistry columns recovers full rank. Refitting without the mask removes the floor entirely, yet interaction MSE improves by only 0.000017 under the reported protocol, and the restored columns remain absent from every training input. On a released RNA-splicing predictor, whose encoding is injective on the measured states, the same computation returns the full design rank of 1,986 and a floor of exactly zero. These results separate what an encoding permits from what a fitted model achieves; they do not identify what limits the remaining error. The rank check needs no fits and bounds what any amount of training under a fixed encoding can recover. The project repository is available at https://github.com/shadi97kh/REPRESENTABLE-BUT-UNLEARNED.
Problem

Research questions and friction points this paper is trying to address.

input encoding
contrast rank
prediction error floor
representability
feature masking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Input Encoding
Contrast Rank
Error Floor
Graph Neural Network
Representability
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.