design error taxonomy

Design and construct hierarchical taxonomies of errors and failures that define categories, class relationships, and priority/order rules to support fine‑grained diagnostic annotation and systematic analysis. Develop, analyze, and maintain the taxonomy structure, category definitions, and assignment criteria so it can be used for error labeling, reporting, and downstream evaluation.

designerrortaxonomy

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.29
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of efficiently constructing a comprehensive and well-structured taxonomy of artificial intelligence skills and tasks from massive hiring data. To this end, the authors propose TaxonomyBuilder, a framework that integrates systematic data filtering, clustering algorithms, and large language model–enhanced hierarchical label generation to automatically derive domain-specific taxonomies from curated, high-quality data subsets. Experimental results demonstrate that taxonomies built from filtered data exhibit significantly broader coverage and superior structural coherence compared to those generated from raw, unfiltered data using existing methods. The study thus establishes a novel paradigm for data-driven, automated taxonomy construction in specialized domains.

AI skillsdata filteringhierarchical taxonomy

This work addresses the challenge of effectively reusing failure feedback from existing agent execution trajectories, which are often lengthy, instance-specific, and lack standardized failure descriptions. To overcome this, the authors propose an unsupervised method that automatically distills raw trajectories into a structured, evidence-backed failure taxonomy. This taxonomy forms an adaptive failure glossary organized along three axes—system-level, role-level, and domain-level—and serves as a unified feedback interface integrated into trajectory selection, runtime monitoring, and system search processes. Requiring no manual annotation, the glossary achieves a 10× compression ratio while exhibiting semantics closely aligned with expert annotations. Empirical results demonstrate significant performance gains across multiple benchmarks: SWE-agent’s resolution rate improves from 60% to 70%, Claude Code reaches 70.7%, and Terminal-Bench 2.0 accuracy increases by 8–15 points.

adaptive taxonomiesagent systemsexecution traces

This study addresses the limitations of existing learner corpora, which typically employ flat annotation schemes that hinder the disentanglement of multiple linguistic dimensions and impede fine-grained analysis of error origins. To overcome this, the authors propose a linguistically grounded multidimensional taxonomy and develop a semi-automatic annotation expansion framework that transforms flat labels into rich, multifaceted structures integrating linguistic features and metadata. As a first application of this approach, they construct a Turkish learner corpus conforming to the proposed taxonomy, accompanied by detailed annotation guidelines and supporting tools to enable cross-dimensional error pattern mining. Experimental results demonstrate that the method achieves 95.86% accuracy at the facet level, substantially enhancing both query capabilities and analytical depth of the corpus.

annotationerror analysisflat labels

This study addresses the challenges of extracting structured risk factors from corporate 10-K filings, particularly inconsistencies with predefined hierarchical taxonomies and the absence of continuous optimization mechanisms. The authors propose an end-to-end, three-stage framework: first, a large language model (LLM) extracts risk factors with source citations; second, semantic embeddings map these factors to a taxonomy; and third, an LLM-as-a-judge mechanism filters erroneous matches. Additionally, an AI agent is introduced to autonomously diagnose and iteratively refine the taxonomy. Experiments on S&P 500 company filings demonstrate a 63% increase in within-industry risk similarity (Cohen’s d = 1.06, AUC = 0.82) and a 104.7% improvement in embedding separation between categories, confirming the method’s effectiveness and generalizability.

10-K filingsautonomous taxonomy maintenancerisk factor extraction

Multi-Label Requirements Classification with Large Taxonomies

Jun 07, 2024
WA
Waleed Abdeen
🏛️ Blekinge Institute of Technology | HOCHTIEF ViCon GmbH

Addressing the challenge of multi-label automatic annotation for large-scale, hierarchical classification systems in software requirements engineering, this study proposes a sentence-level zero-shot classification paradigm to circumvent the high annotation costs associated with supervised training. We introduce the first industrial-scale requirements annotation benchmark comprising 769 taxonomy labels and systematically demonstrate a strong negative correlation between the number of taxonomy leaf nodes and classification recall. We further propose a zero-shot multi-label classification method leveraging SBERT sentence embeddings, achieving significant improvements in recall. Empirical evaluation reveals that hierarchical strategies yield no consistent performance gain across settings. Our work validates the effectiveness and feasibility of zero-shot learning for large-scale requirements classification, offering a scalable, low-human-effort automation solution for requirements tracing. (138 words)

Analyzes classifier types and taxonomy structures impact on classification performanceEvaluates zero-shot learning feasibility for cost-effective multi-label classificationInvestigates multi-label classification for software requirements with large taxonomies

Latest Papers

What's happening recently
View more

This study addresses the high cost and expert dependency of manual taxonomy construction in software engineering (SE) by conducting the first systematic, multi-dimensional empirical evaluation of large language model (LLM)-driven automatic classification in this domain. Leveraging two representative approaches—TnT-LLM and CLIMB—and five state-of-the-art LLMs across seven human-annotated SE paper datasets, the work analyzes performance along key dimensions including classification quality, alignment with expert judgments, reliability, and efficiency. Results reveal that TnT-LLM achieves near-human classification quality but incurs high computational cost and structural complexity, whereas CLIMB offers 15–40× faster inference and 8–49× lower cost at the expense of reduced accuracy in tasks requiring deep technical reasoning. The findings elucidate critical trade-offs among quality, cost, and complexity, providing actionable guidance for method selection in practice.

automated methodsempirical evaluationLarge Language Models

This work addresses fine-grained hierarchical classification in regulatory-intensive domains such as customs tariff codes and export controls, where predictions must strictly adhere to hierarchical structures and rule-based boundaries. Existing approaches struggle to simultaneously ensure hierarchical validity, rule consistency, and boundary-aware reasoning. To bridge this gap, the paper formalizes, for the first time, a regulation-driven hierarchical classification task and introduces a constraint-aware hierarchical search framework. The method parses regulatory documents into a searchable tree and, at each step, retrieves only legally permissible candidate nodes, guiding path decisions through structured fields and evidential text snippets. Evaluated on four expert-validated datasets, the approach significantly outperforms baselines in average accuracy—particularly excelling in distinguishing adjacent categories and handling boundary cases—while producing auditable and traceable decision paths.

constraint-aware reasoningfine-grained classificationhierarchical text classification

This work addresses under-classification errors in ordinal multi-class classification—specifically, the misclassification of high-priority instances as lower-priority classes—by proposing a hierarchical Neyman-Pearson (H-NP) classification framework that rigorously controls such error rates. The method introduces the first flexible H-NP classifier capable of integrating diverse built-in scoring functions (e.g., logistic regression, random forests, support vector machines) or user-defined ones, explicitly designed for ordered multi-class settings. Empirical evaluations demonstrate that the proposed approach satisfies user-specified under-classification error constraints with high probability, offering both practical utility and strong scalability while maintaining strict statistical guarantees on error control.

error controlhierarchical classificationNeyman-Pearson classification

This work addresses the challenge of invalid bug reports, which typically do not require code changes yet incur high manual costs for root cause identification and non-code remediation. The study introduces the first root-cause–oriented, fine-grained taxonomy for such reports, constructs a high-quality benchmark dataset, and systematically evaluates the performance of large language models (LLMs), retrieval-augmented generation (RAG), and agent-driven web search in both subclass identification and generation of non-code repair suggestions. Experimental results demonstrate that RAG achieves a weighted F1 score of 0.66 in subclass identification, while agent-based web search yields the best performance in generating non-code fixes, attaining a success rate of 68.9% as assessed by a Judge LLM.

bug report triageinvalid bug reportsno-code fix generation

Current automated formalization evaluations lack interpretable diagnostics for semantic errors, hindering both system optimization and human understanding. This work proposes FormalRx, a novel framework that introduces the first fine-grained, hierarchical taxonomy of Semantic Correctness Issues (SCI) comprising 28 error categories, and develops an end-to-end diagnostic model, FormalRx-8B, capable of aligning, classifying, localizing, and correcting formalization errors. Trained on 56,287 fine-grained annotated samples, the model achieves strong performance across four diagnostic tasks—0.88 F1 for alignment, 0.71 F1 for classification, 0.75 accuracy for localization, and 0.73 accuracy for correction—significantly outperforming both general-purpose large language models and specialized baselines. The study also releases FormalRx-Test, the first fine-grained diagnostic benchmark, thereby establishing a closed-loop pipeline from opaque evaluation to actionable feedback.

autoformalizationerror diagnosisevaluation framework

Hot Scholars

WZ

Wentao Zhang

Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsc
DS

Dawn Song

Professor of Computer Science, UC Berkeley
Computer Security and Privacy
YH

Yintong Huo

Singapore Management University
AI4SEAIOpsLog analysisMLLM for SE
LZ

Lingling Zhang

Assistant Professor, Xi'an Jiaotong University
Computer visionFew-shot learningZero-shot learning
AC

Aman Chadha

GenAI Leadership @ Apple • Stanford AI • UW-Madison ECE • Ex: Apple, AWS, Alexa, Nvidia
Multimodal AINatural Language ProcessingComputer VisionSpeech Processing