Institution profile

Kalinga Institute of Industrial Technology

Academic institutionasia · in
Official website
Research library63linked papers
Opportunities0open roles
Selected work

Representative Papers

Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees

Aug 08, 2026

This work addresses the vulnerability of reward models in language model alignment to “reward hacking,” which can yield superficially high scores without genuine improvements in response quality. The study introduces a novel covariance-geometric perspective to analyze ensemble evaluators, revealing an intrinsic link between common-mode errors and evaluator disagreement, and formally demonstrates that internal scoring alone cannot detect such errors. To mitigate this, the authors propose a dual-anchor Bernstein error bound, integrating sub-Gaussian joint modeling, covariance decomposition, and adaptive search under conditional calibration, yielding a search error bound proportional to √log K. The theoretical claims are validated through 120 Gaussian stress tests and real-model audits, which also expose the limitations of disagreement-based metrics under strong search pressure.

0 citationsRead paper
Recent publications

Latest Papers

Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees

Aug 08, 2026

This work addresses the vulnerability of reward models in language model alignment to “reward hacking,” which can yield superficially high scores without genuine improvements in response quality. The study introduces a novel covariance-geometric perspective to analyze ensemble evaluators, revealing an intrinsic link between common-mode errors and evaluator disagreement, and formally demonstrates that internal scoring alone cannot detect such errors. To mitigate this, the authors propose a dual-anchor Bernstein error bound, integrating sub-Gaussian joint modeling, covariance decomposition, and adaptive search under conditional calibration, yielding a search error bound proportional to √log K. The theoretical claims are validated through 120 Gaussian stress tests and real-model audits, which also expose the limitations of disagreement-based metrics under strong search pressure.

0 citationsRead paper