๐ค AI Summary
This study addresses the challenge of composing privacy guarantees for multi-institutional scientific AI operating in mixed-trust environments by proposing a unified privacy framework. Methodologically, it reformulates privacy as a six-element assurance problem and establishes a declarative registry to evaluate risks across the model lifecycle. The framework orchestrates synergistic protections by integrating differential privacy, federated learning, secure multi-party computation, and trusted execution environments. Key contributions include identifying critical research gaps specific to leadership-class facilities, such as metadata leakage and instrument side channels, and translating privacy assurances into a universal format that is declarable, comparable, and auditable. Furthermore, this work delineates six priority directions for privacy research, ultimately providing a systematic governance paradigm for cross-institutional scientific collaboration.
๐ Abstract
Scientific artificial intelligence (AI), spanning foundation models (FMs) to federated data-analysis pipelines, is becoming shared infrastructure across national laboratories, universities, hospitals, and industrial partners. This collaboration creates privacy risks whose natural unit is often an institution's participation, research strategy, or technical capability rather than a single record. Differential privacy (DP), federated learning (FL), secure computation, trusted execution, and provenance each protect parts of the stack, but their guarantees rarely compose across mixed-trust institutions, access tiers, and autonomous agents. This perspective recasts privacy for scientific AI as an assurance problem defined by six elements: protected asset, observer, channel, permitted disclosure, guarantee, and evidence. We demonstrate the framing through a claim register for a composite cross-institutional scenario and use it to assess the model lifecycle. Two of the resulting gaps are specific to leadership-class facilities: scheduler, allocation, and telemetry metadata expose an institution's resource posture, and instrument-attached control loops leak research strategy through timing and contention on shared accelerators. We identify six research priorities: institution-level guarantees, agent-communication privacy, cross-tier information flow, privacy-compatible reproducibility, leadership-scale accounting, and instrument side channels. The contribution is a common form for stating, comparing, and auditing claims whose guarantees otherwise remain fragmented across the scientific AI stack.