PAL*M: Property Attestation for Large Generative Models

📅 2026-01-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes the first attribute attestation framework tailored for the full lifecycle of large generative models—such as large language models—encompassing both training and inference phases. To address the limitations of existing approaches in supporting integrity verification for generative tasks and large-scale datasets, the framework integrates Intel TDX confidential virtual machines with NVIDIA H100 security-aware GPUs, establishing end-to-end hardware-level security across CPU-GPU workflows. Furthermore, it introduces an incremental multiset hashing mechanism based on memory-mapped datasets to efficiently track data integrity. The proposed solution achieves significant advances over prior methods in scalability, generality, and security, effectively overcoming their constraints in scale and applicability.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Computer Vision: Large Vision ModelsNatural Language Processing: (Large) Language Models

Application Category

Security and Privacy: Data transparency and provenanceSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Large language models for search
📝 Abstract
Machine learning property attestations allow provers (e.g., model providers or owners) to attest properties of their models/datasets to verifiers (e.g., regulators, customers), enabling accountability towards regulations and policies. But, current approaches do not support generative models or large datasets. We present PAL*M, a property attestation framework for large generative models, illustrated using large language models. PAL*M defines properties across training and inference, leverages confidential virtual machines with security-aware GPUs for coverage of CPU-GPU operations, and proposes using incremental multiset hashing over memory-mapped datasets to efficiently track their integrity. We implement PAL*M on Intel TDX and NVIDIA H100, showing it is efficient, scalable, versatile, and secure.
Problem

Research questions and friction points this paper is trying to address.

property attestation
generative models
large language models
model accountability
secure verification
Innovation

Methods, ideas, or system contributions that make the work stand out.

property attestation
generative models
confidential computing
incremental multiset hashing
secure GPU
🔎 Similar Papers
No similar papers found.