AI Agents 相关度: 8/10

Chatbot-Based Assessment of Code Understanding in Automated Programming Assessment Systems

Eduard Frankford, Erik Cikalleshi, Ruth Breu
arXiv: 2604.07304v1 发布: 2026-04-08 更新: 2026-04-08

AI 摘要

论文提出了一种混合Socratic框架,利用LLM评估学生对代码的理解程度。

主要贡献

  • 对编程教育中会话式评估方法进行了综述
  • 提出了用于自动编程评估系统的混合Socratic框架
  • 讨论了防止LLM生成解释的安全措施

方法论

论文采用了文献综述和框架设计的方法,结合确定性代码分析和双代理会话层。

原文摘要

Large Language Models (LLMs) challenge conventional automated programming assessment because students can now produce functionally correct code without demonstrating corresponding understanding. This paper makes two contributions. First, it reports a saturation-based scoping review of conversational assessment approaches in programming education. The review identifies three dominant architectural families: rule-based or template-driven systems, LLM-based systems, and hybrid systems. Across the literature, conversational agents appear promising for scalable feedback and deeper probing of code understanding, but important limitations remain around hallucinations, over-reliance, privacy, integrity, and deployment constraints. Second, the paper synthesizes these findings into a Hybrid Socratic Framework for integrating conversational verification into Automated Programming Assessment Systems (APASs). The framework combines deterministic code analysis with a dual-agent conversational layer, knowledge tracking, scaffolded questioning, and guardrails that tie prompts to runtime facts. The paper also discusses practical safeguards against LLM-generated explanations, including proctored deployment modes, randomized trace questions, stepwise reasoning tied to concrete execution states, and local-model deployment options for privacy-sensitive settings. Rather than replacing conventional testing, the framework is intended as a complementary layer for verifying whether students understand the code they submit.

标签

LLM 编程教育 自动评估 会话式评估 代码理解

arXiv 分类

cs.SE cs.AI