🤖 AI Summary
This study addresses the limitation of relying solely on benchmark accuracy to evaluate industrial AI, which fails to reveal reliability risks arising from structural misalignments among physics, data, models, and human intent. To overcome this, we propose a structural alignment framework spanning four worlds—physical, representational, machine, and cognitive—that defines the solution space through digitization and objective-encoding interfaces, thereby facilitating a paradigm shift from model-centric evaluation to system-level alignment. Our analysis demonstrates that heterogeneous failures across domains such as healthcare, energy, and exploration originate from shared mechanisms of structural mismatch. Furthermore, this work establishes a unified theoretical foundation grounded in existence and robustness, offering principled guidance for the governance of industrial AI reliability.
📝 Abstract
Artificial intelligence is increasingly deployed in critical industrial domains, including healthcare, energy grids, subsurface exploration, where failures can have severe consequences for human safety, system stability, and economic outcomes. Yet AI is still evaluated primarily through benchmark accuracy, a model-centric metric that fails to capture the structural complexity and risks of real-world deployment. We propose a framework that views industrial AI reliability as a problem of structural alignment across four interacting worlds: physical, representational, machine, and human cognitive. These worlds are connected through two interfaces: digitalization, linking physical reality to computational representations, and goal encoding, translating human cognition to the machine objectives. Together, they define the space of admissible solutions. We characterize the solution space through four attributes: existence, non-uniqueness, robustness, and interpretability and show how mismatches arise at interfaces and propagate across worlds to produce reliability failures. Applications to healthcare, energy grids, and subsurface exploration illustrate that although dominant failure modes differ across domains, for example, interpretability in healthcare, robustness in energy grids, and non-uniqueness in subsurface exploration, all originate from a shared structural mechanism. By shifting the focus from model-centric evaluation to system-level alignment, this framework offers a principled foundation for assessing and governing reliability in industrial AI systems.