conduct qualitative audits

Designs and conducts qualitative audits and interpretive assessments of algorithms, models, or sociotechnical systems by developing audit protocols, running semi‑structured interviews, triangulating documents and case materials, and carrying out comparative qualitative reviews. Analyzes and synthesizes stakeholder perspectives and contextual evidence to surface nontechnical bias drivers, emergent harms, and interpretive conclusions about system behavior and impacts.

conductqualitativeaudits

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Audit Cards: Contextualizing AI Evaluations

Apr 18, 2025
LS
Leon Staufer
🏛️ Technical University of Munich | University of Pennsylvania | Stanford University | Massachusetts Institute of Technology

Current AI governance relies heavily on auditing, yet evaluation reports frequently omit critical contextual information—such as auditor identity, conflicts of interest, model access privileges, and methodological limitations—undermining interpretability, comparability, and trustworthiness. Method: We introduce the “Audit Card” framework, the first systematic, structured specification of six essential contextual dimensions required for rigorous AI assessment, integrating sociotechnical perspectives into audit practice. Grounded in a comprehensive literature review, multi-stakeholder interviews, and analysis of global governance frameworks—and validated through qualitative coding and empirical audits—we document the pervasive absence of such context in existing reports. Contribution/Results: The framework significantly enhances transparency, interpretability, and cross-report comparability of AI evaluations. It provides a scalable, verifiable disclosure paradigm to advance standardization, credibility, and real-world efficacy in AI auditing.

Inconsistent reporting of key audit detailsLack of context in AI audit evaluationsNeed for standardized audit reporting framework

This study investigates the mechanisms underlying the emergence of social bias in artificial intelligence systems, practitioners’ understandings of these issues, and potential mitigation strategies. Employing a qualitative multiple-case design grounded in an interpretivist paradigm, the research integrates intersectionality theory and cognitive science, drawing on semi-structured interviews, document analysis, and triangulation to examine AI practitioners’ experiences across design, development, and governance. Findings reveal that algorithmic bias is deeply rooted in historical inequities, exclusionary assumptions, and organizational pressures for efficiency, underscoring the insufficiency of purely technical fixes. The study proposes an innovative approach that embeds ethical considerations early in the development lifecycle, strengthens structural accountability, fosters diverse stakeholder participation and cognitive awareness, and actively reshapes organizational culture to cultivate AI systems that are transparent, accountable, and aligned with community values.

AI biasalgorithmic fairnessethical AI

This study addresses the limitations of existing algorithmic registry frameworks, which often fail to adequately capture the embedded sociotechnical systems and are hampered by divergent stakeholder expectations and understandings of transparency, thereby undermining accountability. Integrating System-Theoretic Process Analysis (STPA), participatory systems mapping, interviews, and surveys, the research engages diverse stakeholders to co-construct a governance landscape for welfare eligibility assessment algorithms. By situating algorithmic registries within a broader sociotechnical governance framework, the project identifies critical risks—such as wrongful denials, system performance degradation, and breakdowns in appeal mechanisms—that registries alone cannot reveal. Furthermore, it uncovers overlooked normative and political dimensions of algorithmic governance, advocating for safety analyses that are more inclusive and politically attuned.

accountabilityalgorithm registerssociotechnical systems

This study identifies a severe WEIRD (Western, Educated, Industrialized, Rich, Democratic) bias in algorithmic auditing research: over 60% of studies focus exclusively on the U.S., English-language contexts, and a narrow set of platforms, while 85% examine only reductive demographic attributes (e.g., race, gender), neglecting structural discrimination and non-Western sociotechnical settings. Through a systematic literature review (SLR) of 176 peer-reviewed papers, we conduct metadata coding, geolinguistic distribution analysis, thematic clustering, and bias mapping—quantitatively confirming systemic imbalances in platform selection, linguistic coverage, geographic representation, and operationalization of group attributes. Our key contribution is the proposal of an “Inclusive Algorithmic Auditing” framework, advocating multilingual, multicentric, and multidimensional approaches to auditing—including structural and intersectional attributes—and providing an actionable roadmap for transnational collaboration. This work advances a paradigm shift toward globally representative, contextually embedded, and socially accountable algorithmic fairness research.

Analyzes geographic and linguistic disparities in algorithm audits.Highlights limited diversity in group-based attributes studied.Identifies focus skew towards Western platforms and English data.

This study addresses the lack of auditability in large language models (LLMs) when applied to qualitative data analysis, a limitation stemming from their opaque processing. To remedy this, the authors propose QualAnalyzer, an open-source Chrome extension that atomizes each data segment and independently logs prompts, inputs, and outputs, thereby enabling full traceability and auditability of the LLM analytical process for the first time. Integrating browser-based functionality, Google Workspace compatibility, and systematic prompt engineering, QualAnalyzer facilitates rigorous comparative analysis between LLM-generated and human judgments. Empirical validation through two case studies—essay scoring and interview thematic coding—demonstrates that this approach substantially enhances transparency and methodological robustness in LLM-assisted qualitative research.

analytic transparencylarge language modelsLLM-assisted analysis

Latest Papers

What's happening recently
View more

This study addresses the challenges of traditional semi-structured interviews in empirical software engineering, which are often resource-intensive and hindered by cross-time-zone coordination and multilingual barriers. The authors propose a self-administered AI interview approach based on a customized MyGPT model, enabling participants to complete unmoderated interviews via voice in their preferred language, with the system automatically generating structured summaries according to a predefined protocol. As the first work to demonstrate the feasibility of AI-conducted, short-duration, low-risk interviews in this domain, the evaluation shows that 92.4% of 66 submissions met formatting requirements; 90.9% of participants reported a positive experience, 95.5% found the questions clear, and 89.4% expressed willingness to participate again, indicating high acceptability and effectiveness. The study also identifies limitations concerning interview depth and privacy concerns.

cross-language coordinationempirical software engineeringinterview logistics

This study addresses the limitations of current AI auditing practices, which predominantly focus on individual models while overlooking integration risks arising from interactions among system components and between systems and their environments. Through a scoping review and reflexive thematic analysis of 58 studies, the work systematically codes existing literature to delineate, for the first time, three distinct domains of AI integration auditing: inter-component, system–environment, and multi-system. It further introduces domain-specific evaluation dimensions—compatibility, completeness, and oversight—that capture unique aspects of integrated AI systems. The findings reveal that current auditing practices remain fragmented and nascent, underscoring the critical role of accessible information and resource support in effective audit design. The paper calls for novel auditing frameworks capable of spanning components, environments, and systems to enable systematic exploration, identification, coordination, and standardization of integration-related risks.

AI auditingaudit gapscomponent interaction

This study addresses the high cost and limited scalability of manual compliance audits under Germany’s IT-Grundschutz framework, which pose a significant burden on small and medium-sized enterprises. To partially automate the certification process—encompassing structural analysis, protection requirements assessment, modeling, and compliance verification—the authors propose a multi-agent system (MAS) integrated with a hybrid retrieval-augmented generation (HybridRAG) approach. The method innovatively incorporates a hypothesis-validation loop to mitigate agent hallucinations and employs a decoupled reasoning pipeline that separates semantic extraction from deterministic inheritance of protection requirements, thereby enhancing compliance rigor. Experimental results demonstrate a substantial reduction in human effort for semantic tasks; however, stages relying on deterministic logic remain constrained by the inherent probabilistic nature of large language models.

Automated AuditsDeterministic ComplianceIT-Grundschutz

Hot Scholars

WJ

Wendy Ju

Cornell Tech
Interaction DesignHuman Robot InteractionDesignAutomotive Interaction
SL

Seth Lazar

Australian National University
Ethicspolitical philosophyethics of riskethics of war
ZZ

Zhaoxiang Zhang

Institute of Automation, Chinese Academy of Sciences
Computer VisionPattern RecognitionBiologically-inspired Learning
PC

Preetha Chatterjee

Assistant Professor at Drexel University
Software EngineeringEmpirical Software EngineeringSoftware AnalyticsAI4SE
KD

Kostadin Damevski

Professor of Computer Science, Virginia Commonwealth University
Software EngineeringMining Software RepositoriesNatural Language Processing