From Prompts to Trees: Effective LLM-Guided Tree Generation for Few-Shot Tabular Classification

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational cost and poor interpretability of large language models (LLMs) in tabular classification, alongside the limited few-shot performance of conventional decision trees. To overcome these challenges, this work proposes a three-stage knowledge distillation paradigm. Unlike unstable strategies that directly generate entire trees, we introduce a novel progressive construction approach termed "rule generation followed by tree organization." Specifically, LLMs are guided to produce interpretable rules, which are subsequently assembled into a decision tree structure. By integrating prompt engineering with few-shot learning, this method efficiently distills the reasoning capabilities of LLMs into lightweight models. Experimental results demonstrate that the proposed approach significantly improves both classification accuracy and interpretability across multiple real-world datasets while substantially reducing inference overhead.
📝 Abstract
While Large Language Models (LLMs) possess rich world knowledge and impressive generalization capabilities, their direct application to tabular data classification is hindered by high inference costs and limited interpretability. In contrast, decision trees are fast and transparent but often underperform in low-data regimes. In this work, we propose a novel framework that bridges these paradigms by distilling LLM knowledge into interpretable decision trees under a few-shot learning setting. Instead of directly prompting the LLM to generate full trees, which is often unstable and inefficient, we develop a three-stage paradigm that prompts the LLM to generate rules and organize the rules into a tree. Experiments on multiple real-world tabular datasets demonstrate that our method achieves superior accuracy and interpretability with significantly lower prompting overhead compared to existing baselines.
Problem

Research questions and friction points this paper is trying to address.

Tabular Classification
Few-Shot Learning
Large Language Models
Decision Trees
Interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Decision Trees
Few-Shot Learning
Knowledge Distillation
Tabular Classification