Researchers at Drexel University attempted to program language models to simulate students with varying levels of algebra proficiency. However, their efforts to have the models "respond like a struggling student" fell short, as the models could not convincingly take on that role. The models, including Gemini 3.1 Flash Lite, Claude Haiku 4.5, and GPT-5.4-mini, were tasked with solving 379 algebra problems while embodying five distinct student profiles. Regardless of whether they were designated as "top students" or "underperformers," their accuracy remained impressively high, fluctuating between 96.8% and 100%. The challenge lies in the fact that these models already possess the correct answers and find it difficult to downplay their capabilities. To address this, the researchers divided the task into two parts: first, a separate algorithm assesses a student's knowledge and potential errors; then, a language model articulates the student's answer. This approach led to more realistic variations in performance, with the "near-expert" student achieving an accuracy of 85.2%, the average student at 57.8%, and the struggling student at 44.1%. The researchers admit that they haven't yet compared these virtual "students" with actual students, indicating that their work primarily serves as a method for teaching AI to make credible mistakes rather than a precise representation of human learning processes.
Informational material. 18+.