THE CRUNCH

Microsoft and the University of Illinois have built a system called StudentSim that creates digital replicas of individual students from limited data. These replicas make realistic mistakes and respond to tutoring, allowing researchers to train AI tutors more quickly than using real learners. The method outperforms the large language model GPT-5.4 in chess, English, and math tests.

The researchers note that training AI tutors with real students is prohibitively expensive and time-consuming. StudentSim addresses this by using a two-stage training process. First, a base model learns shared patterns from pooled student data. Then, it is tailored to an individual student using only a few records of that person's work. This approach prevents the model from overfitting to limited data.

The system uses Alibaba's Qwen3-4B-Instruct language model as its base. In tests across chess, English, and math, StudentSim outperformed GPT-5.4 in all three subjects. In chess, StudentSim correctly predicted a player's next move about twice as often and almost always followed corrective guidance. GPT-5.4 and specialized chess models fell behind in these areas.

The researchers also used a student replica to improve a chess tutor. Professional chess players evaluated three versions of the tutor. The version trained with StudentSim scored highest on all three measures, making the fewest serious factual errors and receiving the highest scores for explanation quality and adaptation to the individual student. The tutor trained with GPT-5.4 scored worse on factual accuracy than the tutor that received no extra training.

WHAT HAPPENS NEXT

The researchers note this is a proof of concept, not a claim to have built the best tutor.