A course-specific assistant still disappointed
A University of Maryland working paper dated September 29 reports a randomized study involving 2,379 undergraduates and 30 instructors during fall 2025. Instructors were assigned within groups of the same or similar courses to offer a study assistant or not. The assistant drew on instructor-selected course materials and could provide citations. It was more closely tied to the class than a general chatbot.
In the tighter comparison—1,353 students taught by 22 instructors in sections of the same courses—offering the assistant reduced final grades by an estimated 4.12 points on a 100-point scale, or 0.37 standard deviations. The broader sample did not provide statistically significant evidence of a final-grade loss. Recorded participation on the existing course platform declined substantially in both samples.
Those results describe the rollout of this particular assistant. Only about 15% of students offered it used it at least once, and students in both conditions could access other AI tools. It would be wrong to say every student who chatted with it lost four points. The researchers also cannot establish that reduced course-platform activity caused the lower grades.
Its configuration deserves attention. Few instructors switched away from the default direct-instruction mode; only two ended the term using the mode that guides students toward an answer. Grounding an assistant in the right readings may help it answer the right question. That alone does not ensure the student gets enough practice.
A smaller trial made students attempt the answer
A separate September working draft from Boston University researchers studied a structured tutor in an online MBA corporate-finance course. It randomized 86 students between tutor access and a comparison group that could still use consumer AI. The five-week tutor sessions asked one question at a time, required an attempt before an explanation, and used hints or smaller steps when a student got stuck.
The authors report an adjusted 6.63-point advantage on their 55-point learning assessment among the 51 students who completed both assessment waves. Their analysis of the instructor's final exam, available for all 86 randomized students, found a smaller positive difference of 2.6 points out of 100.
That is encouraging, with limits. The paired-assessment result leaves out students who did not finish both tests; similar completion rates between groups do not erase that concern. This was one graduate course, and the paper is a working draft. Its score cannot be placed beside Maryland's final-grade estimate as if the two trials were measuring the same thing.
Nor did either trial directly randomize students between the Maryland assistant and the Boston tutor. The contrast suggests a design question worth pursuing, not proof that one prompt rule explains the difference. What does the learner have to produce before the software supplies the next explanation?
A nicer voice did not improve measured mastery
The Boston study also alternated voice and text weeks within the tutor group. Voice produced more conversational turns, and students increasingly preferred it. Yet measured weekly mastery was statistically equivalent within the study's stated bounds, while voice cost 2.8 times as much to deliver in this setup.
Those are provider costs for the research system, not a price comparison between consumer subscriptions. They do separate two buying decisions that can otherwise get muddled: which format makes studying tolerable, and which format helps you learn more.
Someone who finds typing tiring may reasonably prefer voice. A person studying quietly beside a sleeping household may want text. The trial gives no reason to assume the more human-sounding option will teach either person better. Choose the channel you can comfortably use, then check the learning separately.
Try a short session that leaves evidence on paper
For your next study session, choose a concept from the course rather than asking the chatbot to invent a curriculum. Follow the course's AI rules. You can ask: ‘Use this material to ask me one question at a time. Wait for my attempt. If I am wrong, give a small hint before showing a solution.’ That request may improve the shape of a conversation; it does not turn a general chatbot into a validated tutor.
When the explanation finally makes sense, hide it. Try another question from an instructor-provided practice set, or explain the idea in your notebook and compare it with the course material afterward. Leave yourself a brief note about where you got stuck. Revisit that spot in the next session before reopening the chat.
This is a practical check, not a new research finding. It can reveal a familiar disappointment: the answer was clear while it was in front of you, but you still could not begin alone. If that keeps happening, change the session or ask a teacher, classmate or human tutor for help. More chatbot messages are unlikely to settle the question.
A shorter evening of study is welcome. So is a patient explanation. The bargain is poor if the same chapter takes another evening because nothing stayed with you.