Robust Decentralized Fairness Auditing

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the trust vulnerability in decentralized auditing of large language models, where malicious nodes forge statistical vectors to execute "fairness washing." To mitigate this, we propose Auditopus, a framework employing a serverless round-based collaboration mechanism wherein auditors exchange only cumulative statistical vectors to simultaneously preserve privacy and enable fairness estimation. Its core innovation lies in a local consistency detection mechanism that automatically down-weights anomalous nodes, effectively defending against both single-point attacks and large-scale collusion. Experimental results demonstrate that Auditopus reduces auditing error by 78% compared to an undefended baseline and accurately identifies unfair models even when malicious nodes constitute up to 49% of the network.
📝 Abstract
Emerging legislation requires large language models (LLMs) to be audited for compliance with regulatory standards, particularly fairness. Such black-box audits typically assume a single auditor with access to a large, representative set of queries. In practice, it can be difficult for an auditor to obtain such a query set, but multiple auditors can together cover the relevant demographic groups by auditing the LLM collaboratively with their individual query sets. However, relying on multiple auditors raises a fundamental trust problem, as they may act on behalf of the LLM provider to portray a misleading appearance of fairness, i.e., fairwashing. We propose Auditopus, a novel approach for robust decentralized fairness auditing. In Auditopus, auditing proceeds in rounds without a central server. In each round, every auditor issues a fixed number of queries to the LLM, and sends only cumulative statistics vectors of its query results to other auditors instead of sensitive queries in clear. The fairness of the audited LLM is then estimated by aggregating all the vectors. We show theoretically and empirically that even a single adversarial auditor in the network can steer this estimate by fabricating the vectors it sends, making an unfair LLM appear fair. To address this threat, Auditopus has each honest auditor locally down-weight any auditor whose cumulative statistics vectors are statistically inconsistent with previous ones. We implement Auditopus and compare it to robust aggregation baselines on two datasets with two pre-trained LLMs. Against an attacker that optimizes the vectors it sends to make the LLM appear fair, Auditopus reduces audit error by up to 78% on average relative to no defense and at least 62% relative to the robust aggregation baselines. Even when 49% of the auditors are adversarial, Auditopus never lets a very unfair or moderately unfair LLM pass as fair.
Problem

Research questions and friction points this paper is trying to address.

Decentralized Fairness Auditing
Large Language Models
Fairwashing
Adversarial Auditors
Robust Aggregation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decentralized Fairness Auditing
Large Language Models
Robust Aggregation
Fairwashing Defense
Adversarial Robustness
🔎 Similar Papers
No similar papers found.