Score
Design and implement evaluation protocols and diagnostic tooling that partition inputs into interpretable "slices" (capability slices) defined by background conditions, compute and aggregate per-slice metrics for stable measurement, and localize specific model weaknesses to individual slices to guide debugging and remediation.
To address the dual challenges of end-to-end performance monitoring gaps and SLA-aware resource allocation under constrained telemetry budgets in 6G network slicing, this paper models slice monitoring as a closed-loop control problem. It introduces the novel concept of “telemetry primitive contracts” to formally specify minimal data-plane capabilities required for SLA compliance. We further propose an SLA-criticality-driven dynamic resource scheduling mechanism and design a change-triggered In-band Network Telemetry (INT) coordination architecture. Evaluated on programmable switches and large-scale simulations, our approach achieves four times higher monitoring accuracy for critical slices compared to static baselines. The change-triggered INT scheme significantly outperforms existing telemetry primitives while strictly adhering to contract constraints. To the best of our knowledge, this is the first solution enabling SLA-sensitive, end-to-end visible, and resource-adaptive real-time slice monitoring.
Automatically identifying and attributing performance-deficient data slices (i.e., subpopulations) in unlabeled, unstructured image data remains challenging due to the absence of metadata and interpretable diagnostics. Method: This paper proposes a metadata-agnostic data slicing framework that jointly leverages gradient-based class attribution maps and clustering-driven slice discovery; introduces Attribution Mosaic—a novel visual analytics technique for slice-level attribution interpretation; and integrates a human-in-the-loop analysis pipeline with a plug-and-play attribution-consistency regularization mechanism for end-to-end model repair. Results: Evaluated on two benchmark vision datasets, the method achieves an average 5.2% improvement in slice-level accuracy, enables users to complete bias diagnosis and mitigation within 10 minutes, and reduces reliance on manual annotations by over 90%. Its core contributions are the first metadata-free, interpretable slice discovery method, attribution-driven visual analytics, and a deployable, end-to-end repair pipeline.
Current evaluation methods for large language models (LLMs) primarily identify failing samples or categories but struggle to uncover underlying capability deficiencies, thereby limiting targeted model improvement. This work proposes CRAFT, a novel framework that diagnoses model weaknesses at the scoring-criterion level. CRAFT constructs a hierarchical capability tree by extracting capability descriptions and applying hierarchical clustering, then dynamically identifies low-performance nodes across multiple granularities to generate targeted fine-tuning data. Evaluated on financial and legal domains as well as 13 standard benchmarks, CRAFT significantly outperforms prompt-clustering and random data generation baselines. Fine-tuning four open-source LLMs with CRAFT-generated data consistently enhances their performance, demonstrating more precise localization of capability gaps and enabling efficient, targeted model refinement.
Current vision-language models (VLMs) for autonomous driving suffer from sparse validation coverage within their operational design domain (ODD), leading to unreliable empirical failure rates. To address this, this work proposes SliceScorer—a scoring mechanism—and SliceNav, a validation pipeline that jointly incorporates exposure frequency priors and proximity-based failure propagation priors to proactively identify and recommend high-risk, under-tested scenario slices. The approach leverages large language models (LLMs) to orchestrate an interpretable, deterministic, and end-to-end validation workflow. For the first time, it integrates deterministic risk scoring with LLM-driven validation scheduling. Experiments on WiseAD, DriveMM, and Cosmos-Reason2-2B demonstrate that SliceNav more efficiently uncovers high-risk coverage gaps and yields greater recommendation diversity compared to existing methods, with ablation studies confirming the contribution of each component.
Existing error diagnosis methods for machine learning models struggle to identify semantically coherent error patterns, disentangle deep-rooted bias sources, and rely heavily on manual annotations and predefined attributes—limiting their applicability in specialized domains such as medical imaging. This paper proposes a language-driven error diagnosis paradigm: leveraging large language models (LLMs) to autonomously generate verifiable hypotheses from textual model outputs, enabling unsupervised, prior-free error slice discovery; introducing pseudo-attribute synthesis and cross-subgroup collaborative correction to internalize LLMs’ implicit domain knowledge for fine-grained bias attribution; and establishing a multimodal bias evaluation framework. Evaluated across six diverse datasets—including natural and radiological imaging—the method consistently outperforms over 200 baseline approaches in error slice completeness, semantic coherence, and error mitigation efficacy.
This work addresses the lack of automated, structured failure-recovery mechanisms in current software engineering agents, which struggle to translate heterogeneous runtime evidence into actionable repair guidance. The paper proposes PROBE, a novel framework that introduces a failure-anchored, structured recovery paradigm. PROBE employs a three-layer architecture—telemetry, diagnosis, and guidance gate—to decouple yet coordinate diagnosis and recovery, enabling non-intrusive integration. By integrating runtime telemetry, multi-signal diagnosis, and evidence-driven bounded guidance generation, PROBE constructs an end-to-end recovery pipeline. Evaluated on 257 unresolved cases, PROBE achieves a Top-1 diagnostic accuracy of 65.37% and a recovery success rate of 21.79%, significantly outperforming the strongest baseline. Its practical feasibility has been validated through deployment in Microsoft’s IcM system.
Traditional end-to-end evaluation struggles to pinpoint the specific layer responsible for regressions in large language model (LLM) agents. This work proposes a hierarchical regression detection method that decomposes production-grade LLM agents into functionally distinct layers and constructs LLM-free, deterministic test suites. By integrating coverage-honest test adequacy criteria with layered assertion slicing, the approach enables sub-second, isolated regression validation. Evaluated on 238 test cases, the method accurately identifies seven categories of manually injected regressions, with per-layer pass rates dropping by 25%–91%, while aggregate metrics decline only marginally by 1.7–5.9 percentage points. These results demonstrate substantially higher sensitivity and precise localization capability compared to conventional end-to-end evaluation.
This work addresses the semantic gap between evaluation metrics and training data in large model pretraining, which hinders precise diagnosis and remediation of capability deficiencies. The authors propose “capability slices” as fundamental units aligning evaluation and data, establishing a bidirectional classification framework that links evaluation tasks with non-instructional training data through explicit mapping rules. This enables a closed-loop pipeline from evaluation failures to targeted data interventions. For the first time, the approach supports auditable and systematic reasoning that translates evaluation signals into data corrections, moving beyond intuition-driven tuning paradigms. Experiments demonstrate its efficacy in both directions: repairing specific training loss components restores BBH performance to 66.44, while targeted data sampling boosts AIME2025/2026 Pass@128 from 6.67/0.00 to 26.67.