slice-based evaluation

Design and implement evaluation protocols and diagnostic tooling that partition inputs into interpretable "slices" (capability slices) defined by background conditions, compute and aggregate per-slice metrics for stable measurement, and localize specific model weaknesses to individual slices to guide debugging and remediation.

slice-basedevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Dynamic SLA-aware Network Slice Monitoring

Dec 12, 2025
NS
Niloy Saha
🏛️ University of Waterloo | University of Regina

To address the dual challenges of end-to-end performance monitoring gaps and SLA-aware resource allocation under constrained telemetry budgets in 6G network slicing, this paper models slice monitoring as a closed-loop control problem. It introduces the novel concept of “telemetry primitive contracts” to formally specify minimal data-plane capabilities required for SLA compliance. We further propose an SLA-criticality-driven dynamic resource scheduling mechanism and design a change-triggered In-band Network Telemetry (INT) coordination architecture. Evaluated on programmable switches and large-scale simulations, our approach achieves four times higher monitoring accuracy for critical slices compared to static baselines. The change-triggered INT scheme significantly outperforms existing telemetry primitives while strictly adhering to contract constraints. To the best of our knowledge, this is the first solution enabling SLA-sensitive, end-to-end visible, and resource-adaptive real-time slice monitoring.

Addresses lack of control mechanisms in existing network slice monitoring solutions.Ensures real-time SLA compliance for network slices within limited monitoring budgets.Provides end-to-end visibility and dynamic resource allocation for slice monitoring.

AttributionScanner: A Visual Analytics System for Model Validation with Metadata-Free Slice Finding

Jan 12, 2024
XX
Xiwei Xuan
🏛️ University of California | Bosch | Splunk

Automatically identifying and attributing performance-deficient data slices (i.e., subpopulations) in unlabeled, unstructured image data remains challenging due to the absence of metadata and interpretable diagnostics. Method: This paper proposes a metadata-agnostic data slicing framework that jointly leverages gradient-based class attribution maps and clustering-driven slice discovery; introduces Attribution Mosaic—a novel visual analytics technique for slice-level attribution interpretation; and integrates a human-in-the-loop analysis pipeline with a plug-and-play attribution-consistency regularization mechanism for end-to-end model repair. Results: Evaluated on two benchmark vision datasets, the method achieves an average 5.2% improvement in slice-level accuracy, enables users to complete bias diagnosis and mitigation within 10 minutes, and reduces reliance on manual annotations by over 90%. Its core contributions are the first metadata-free, interpretable slice discovery method, attribution-driven visual analytics, and a deployable, end-to-end repair pipeline.

Detect and mitigate model issues like biases.Identify poor performance subgroups in datasets.Validate vision models without additional metadata.

Current evaluation methods for large language models (LLMs) primarily identify failing samples or categories but struggle to uncover underlying capability deficiencies, thereby limiting targeted model improvement. This work proposes CRAFT, a novel framework that diagnoses model weaknesses at the scoring-criterion level. CRAFT constructs a hierarchical capability tree by extracting capability descriptions and applying hierarchical clustering, then dynamically identifies low-performance nodes across multiple granularities to generate targeted fine-tuning data. Evaluated on financial and legal domains as well as 13 standard benchmarks, CRAFT significantly outperforms prompt-clustering and random data generation baselines. Fine-tuning four open-source LLMs with CRAFT-generated data consistently enhances their performance, demonstrating more precise localization of capability gaps and enabling efficient, targeted model refinement.

capability diagnosisevaluationfine-tuning data

Current vision-language models (VLMs) for autonomous driving suffer from sparse validation coverage within their operational design domain (ODD), leading to unreliable empirical failure rates. To address this, this work proposes SliceScorer—a scoring mechanism—and SliceNav, a validation pipeline that jointly incorporates exposure frequency priors and proximity-based failure propagation priors to proactively identify and recommend high-risk, under-tested scenario slices. The approach leverages large language models (LLMs) to orchestrate an interpretable, deterministic, and end-to-end validation workflow. For the first time, it integrates deterministic risk scoring with LLM-driven validation scheduling. Experiments on WiseAD, DriveMM, and Cosmos-Reason2-2B demonstrate that SliceNav more efficiently uncovers high-risk coverage gaps and yields greater recommendation diversity compared to existing methods, with ablation studies confirming the contribution of each component.

coverage gapdriving VLMsOperational Design Domain

LADDER: Language Driven Slice Discovery and Error Rectification

Jul 31, 2024
SG
Shantanu Ghosh
🏛️ Boston University

Existing error diagnosis methods for machine learning models struggle to identify semantically coherent error patterns, disentangle deep-rooted bias sources, and rely heavily on manual annotations and predefined attributes—limiting their applicability in specialized domains such as medical imaging. This paper proposes a language-driven error diagnosis paradigm: leveraging large language models (LLMs) to autonomously generate verifiable hypotheses from textual model outputs, enabling unsupervised, prior-free error slice discovery; introducing pseudo-attribute synthesis and cross-subgroup collaborative correction to internalize LLMs’ implicit domain knowledge for fine-grained bias attribution; and establishing a multimodal bias evaluation framework. Evaluated across six diverse datasets—including natural and radiological imaging—the method consistently outperforms over 200 baseline approaches in error slice completeness, semantic coherence, and error mitigation efficacy.

Bias UnderstandingDomain-specific KnowledgeModel Error Identification

Latest Papers

What's happening recently
View more

This work addresses the lack of automated, structured failure-recovery mechanisms in current software engineering agents, which struggle to translate heterogeneous runtime evidence into actionable repair guidance. The paper proposes PROBE, a novel framework that introduces a failure-anchored, structured recovery paradigm. PROBE employs a three-layer architecture—telemetry, diagnosis, and guidance gate—to decouple yet coordinate diagnosis and recovery, enabling non-intrusive integration. By integrating runtime telemetry, multi-signal diagnosis, and evidence-driven bounded guidance generation, PROBE constructs an end-to-end recovery pipeline. Evaluated on 257 unresolved cases, PROBE achieves a Top-1 diagnostic accuracy of 65.37% and a recovery success rate of 21.79%, significantly outperforming the strongest baseline. Its practical feasibility has been validated through deployment in Microsoft’s IcM system.

failure diagnosispost-failure recoveryruntime telemetry

Traditional end-to-end evaluation struggles to pinpoint the specific layer responsible for regressions in large language model (LLM) agents. This work proposes a hierarchical regression detection method that decomposes production-grade LLM agents into functionally distinct layers and constructs LLM-free, deterministic test suites. By integrating coverage-honest test adequacy criteria with layered assertion slicing, the approach enables sub-second, isolated regression validation. Evaluated on 238 test cases, the method accurately identifies seven categories of manually injected regressions, with per-layer pass rates dropping by 25%–91%, while aggregate metrics decline only marginally by 1.7–5.9 percentage points. These results demonstrate substantially higher sensitivity and precise localization capability compared to conventional end-to-end evaluation.

component-level evaluationdeterministic testinglayer-isolated evaluation

This work addresses the semantic gap between evaluation metrics and training data in large model pretraining, which hinders precise diagnosis and remediation of capability deficiencies. The authors propose “capability slices” as fundamental units aligning evaluation and data, establishing a bidirectional classification framework that links evaluation tasks with non-instructional training data through explicit mapping rules. This enables a closed-loop pipeline from evaluation failures to targeted data interventions. For the first time, the approach supports auditable and systematic reasoning that translates evaluation signals into data corrections, moving beyond intuition-driven tuning paradigms. Experiments demonstrate its efficacy in both directions: repairing specific training loss components restores BBH performance to 66.44, while targeted data sampling boosts AIME2025/2026 Pass@128 from 6.67/0.00 to 26.67.

capability slicedata-evaluation gapevaluation-to-data inference

Hot Scholars

RP

Rohit Parikh

Distinguished Professor, CS, Math, Philosophy, Brooklyn College and CUNY Graduate Center
logicgame theoryphilosophy of language