π€ AI Summary
This study addresses the friction and tool-circumvention behaviors that arise when AI teaching assistants employ overly restrictive or context-deprived guardrails. To investigate this, we develop large language model-based instructional guardrails and conduct a randomized controlled trial to systematically examine how Socratic versus direct instruction styles, combined with varying levels of context awareness, affect programming learning experiences. Our findings reveal that safety guardrails are not inherently beneficial; notably, the fully context-aware Socratic AI tutor received the lowest student evaluations. This work underscores the necessity of balancing pedagogical guidance with context sensitivity to enhance studentsβ sense of support. Ultimately, it provides critical empirical evidence for optimizing AI teaching assistant design and sustaining student engagement in educational settings.
π Abstract
AI teaching assistants (AI TAs) backed by large language models (LLMs) and pedagogical guardrails are increasingly being integrated into programming courses, providing students with scalable access to hints, conceptual explanations, and code-level feedback. However, guardrails may also create friction. If students feel that the support provided is overly restrictive or poorly contextualized to their current progress, they may bypass approved tools for general-purpose LLMs. To investigate how AI TA design affects students' learning experiences, we conducted a randomized controlled trial with 132 students in an introductory programming course. Students completed three tasks related to code-writing and debugging and were randomly assigned to one of four AI TAs varied across two dimensions: pedagogical guidance style (Socratic vs. Direct instruction) and context awareness (no context vs. full context of the problem and student solution). We examined students' perceptions, interaction behaviors, and evidence of post-task comprehension. Students rated the Socratic AI TA with full context least favorably, reporting significantly lower perceived support for task completion. Descriptively, this condition also showed the highest observed interaction stress, the highest rate of external LLM use, and the lowest proportion of post-task explanations demonstrating full comprehension, though these differences were not statistically significant. These findings suggest that guardrailed AI TAs are not automatically better for learning. Instead, their effectiveness depends on how pedagogical guidance and contextual awareness are balanced in ways that students experience as useful, supportive, and worth continuing to use.