attribution-guided data sharing

Designs and builds systems that compute and aggregate per-example and per-feature contribution scores at scale and use those scores to select minimal subsets of data or features for sharing between parties. Analyzes and implements attribution-guided policies that align client contributions with a global task and prioritize data for aggregation to improve model accuracy and convergence while managing communication and scalability constraints.

attribution-guideddatasharing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of decision threshold invalidation caused by score distribution shifts during model updates in multi-tenant Model-as-a-Service (MaaS) environments, which hinders rapid model iteration. To resolve this, the authors propose MUSE, a novel framework featuring a two-stage score transformation mechanism that maps model outputs to a stable reference distribution. Integrated with dynamic intent-based routing, MUSE enables seamless model updates and efficient sharing across tenants while preserving strict isolation guarantees. The architecture supports highly available, low-latency deployment and has been successfully deployed at scale by Feedzai, processing over 55 billion events annually. This deployment reduces model rollout time from weeks to minutes, significantly decreasing both fraud losses and operational costs.

decision thresholdmodel servingmodel update

Recommendation systems face challenges including deploying ultra-large-scale models, continual learning on online streaming data, adapting to heterogeneous scenarios, and meeting stringent latency and computational constraints. This paper proposes the Foundation-Expert paradigm: a centralized foundation model learns cross-scenario general-purpose representations, which are efficiently transferred to lightweight expert models via target-aware embedding and a decoupled training/inference architecture. It is the first work to deploy this paradigm at production scale—handling trillions of daily requests—enabling lifelong learning, multimodal fusion, and low-overhead knowledge transfer. Based on this, we design HyperCast, a system that rearchitects training, serving, and iterative deployment pipelines. Evaluated in Meta’s production environment—processing数千亿 requests daily—HyperCast significantly improves online metrics over single-stage baselines, while enhancing R&D efficiency and maintaining high computational resource utilization.

Adapting to diverse tasks with strict latency constraints is difficultDeploying hyperscale recommender models efficiently remains unresolvedRecommender systems face unique challenges like shifting data distributions

Aggregated Individual Reporting for Post-Deployment Evaluation

Jun 22, 2025
JD
Jessica Dai
🏛️ University of California, Berkeley | University of California, San Francisco

This study addresses the lack of fine-grained, dynamic evaluation mechanisms for AI systems in the post-deployment phase. We propose the Aggregated Individual Reporting (AIR) framework, the first to systematically incorporate qualitative, user-generated feedback from real-world interactions into AI evaluation. AIR employs structured reporting collection, temporal aggregation, and thematic analysis to enable early detection of performance degradation and emergent safety risks. Grounded in democratic AI principles, it specifies an actionable reporting interface, a scalable data aggregation pipeline, and a responsive decision-making pathway—thereby bridging critical gaps in user-centricity and real-time governance within existing evaluation paradigms. Empirical validation demonstrates that individual reports effectively surface previously unanticipated safety issues in black-box models and facilitate targeted interventions. We further formalize a standardized AIR workflow and outline future research directions toward cross-platform collaborative AI governance.

Aggregate individual reports to identify safety and performance issuesDevelop post-deployment AI evaluation beyond static benchmarksEnable public reporting for democratic AI system assessment

Data Shapley in One Training Run

Jun 16, 2024
JT
Jiachen T. Wang
🏛️ Princeton University | University of California, Berkeley | Virginia Tech

Existing Data Shapley methods require repeated retraining of data subsets, incurring prohibitive computational overhead, and yield generic contribution scores that cannot be tailored to specific target models. Method: We propose In-Run Data Shapley—the first framework enabling efficient, model-specific data contribution attribution for a single training run, including large language models. It embeds Shapley value theory directly into the dynamic parameter update process via gradient tracing and stochastic linear approximation, eliminating the need for auxiliary training. Contribution/Results: The method incurs negligible attribution overhead and supports fine-grained, pretraining-stage quantification of data value. Experiments demonstrate its interpretability and practical utility in copyright provenance and data curation. By bypassing iterative retraining, In-Run Data Shapley overcomes the computational bottleneck hindering data value assessment in large-scale models.

Efficient computation without retraining subsetsInsights into pretraining data contributionsScalable data attribution for target models

This study addresses the issue that standard evaluation mechanisms—such as win rate—can induce model homogenization in AI markets, thereby undermining consumer utility. To counter this, the authors propose a weighted win rate mechanism that incentivizes model specialization by offering differentiated rewards for high-quality responses. Drawing on game-theoretic and mechanism design frameworks, the work combines theoretical analysis with empirical validation using real-world benchmark data. The results demonstrate that the proposed mechanism effectively promotes model diversity while significantly enhancing consumer welfare, offering a principled approach to aligning model development incentives with user interests in competitive AI ecosystems.

AI marketplaceconsumer utilityincentive design

Latest Papers

What's happening recently
View more

This study addresses the challenges of excessive discounting, inefficient manual review, and inconsistent pricing decisions in enterprise software contract negotiations, which often stem from a lack of data-driven benchmarks. To overcome these issues, this work proposes an adaptive nearest-neighbor-based peer comparison scoring system that embeds dynamic peer benchmarking directly into the contract design workflow. By leveraging tree-based ensemble models to learn similarity from historical contracts and defining data-driven proximity through shared leaf nodes, the system generates real-time letter-grade ratings and product-line-level insights for new contracts. The approach enables auditable, real-time pricing decisions and centralized oversight, significantly enhancing discount discipline post-deployment and driving commercially meaningful revenue uplift across scored contract portfolios.

algorithmic governancecontract designdiscount discipline

This work addresses the lack of a reproducible evaluation framework for meta-decision strategies—such as task decomposition and tool invocation—in existing agent systems. We introduce MetaRoute-Bench, the first open benchmark enabling fine-grained analysis of meta-decision routing, comprising 180 synthetic tasks, 8 distinct strategies, and 30 random seeds per configuration. Evaluation employs offline seeded execution and multidimensional metrics—including success rate, cost, and latency—to ensure fair comparison. Experiments demonstrate that task-aware compositional strategies achieve a significantly higher success rate (79.4%) compared to static strategies (76.7%), single-step routing (67.4%), and direct answering (52.9%), with only marginal increases in cost (4.7%) and latency (6.4%). Ablation studies further confirm the critical contributions of compositional operations and verification mechanisms. Code and execution trajectories are publicly released.

agentic workflow routingevaluation frameworkmeta-decision

This study addresses the lack of systematic understanding regarding the implementation and maintenance of the Model Context Protocol (MCP) in real-world open-source projects. To bridge this gap, we introduce a transparent, reproducible multi-stage validation pipeline that integrates GitHub REST/GraphQL APIs with custom Python scripts to systematically annotate structural evidence, classify repository roles, and filter out non-functional examples from 3,238 candidate repositories. This process yields a high-quality dataset of 2,297 verified MCP projects, achieving a validation precision of 83% at 95% confidence. Our analysis reveals Python and TypeScript as the dominant implementation languages and identifies hybrid architecture as the most prevalent design pattern, thereby establishing the first large-scale empirical benchmark for MCP ecosystem research.

GitHublarge-scale datasetMCP implementation

Hot Scholars

VN

Viswa Narayanan Sankaranarayanan

PhD Candidate, Lulea University of Technology, Sweden
Adaptive controlEuler-Lagrangian systemsAerial robotsBarrier functions
FK

Frederick Klauschen

Institute of Pathology, University of Munich (LMU)
PathologyDigital Pathology/AIPrecision MedicineMolecular Diagnostics
OE

Oliver Eberle

TU Berlin
Explainable AIInterpretabilityDeep LearningMachine Learning
JC

Jimmy Chiun

PhD Student at National University of Singapore
Robot learningMulti-robot systemsVisual-Language Navigation
YC

Yuhong Cao

National University of Singapore
Robot learningPath Planing