The assistant disappeared before the final questions

Researchers recruited 1,222 US-based participants through the online research platform Prolific across three randomized experiments: 354 for the first fraction task, 667 for its replication, and 201 for reading comprehension. After attention and ability checks, the analyzed samples were 307, 585 and 168 respectively. These were brief, paid online sessions, not a semester of lessons or months of workplace observation.

Participants assigned AI access could ask a sidebar assistant for help during an initial set of questions. In the first experiment, the assistant had the fraction problem and its solution in advance. A participant could ask for the answer with very little effort. For the final three questions, the assistant was removed and participants were asked to work without outside help.

While help was available, the AI groups answered more accurately. After it disappeared, they performed worse than the groups that had worked without AI. Researchers also tracked skipping: people could decline a problem without losing payment. That behavior, rather than a direct test of attention span, supplied the study's measure of persistence.

The larger replication deserves its own sentence

The first experiment had a potential selection problem. Participants were excluded partly on performance during a phase when one group could obtain correct answers from AI. The second experiment added an unaided pretest and changed the comparison interface to address that concern.

In this larger replication, the AI group's lower final solve rate remained statistically significant. Its higher overall skip rate did not. The reading-comprehension experiment found both lower solve rates and higher skip rates after AI removal. So the evidence for worse independent performance was more consistent across these experiments than the evidence for increased skipping.

The researchers also divided AI participants by how they said they used the assistant. People who sought direct answers did worse afterward than those who sought hints. That comparison is suggestive, but it was not a randomized trial of answer-giving versus hint-giving. People chose their own approach. It cannot establish that switching a chatbot to hints would prevent the problem.

Ten minutes is an exposure window, not a damage threshold

The paper explicitly leaves the duration of the effect unresolved. Participants were tested immediately after AI removal. The researchers do not know whether the disadvantage survives hours or days, or whether repeated use makes it accumulate or fade. A short experiment can reveal an immediate cost without telling us the shape of a long-term habit.

The authors propose that instant answers may reset expectations about how much effort a task should require. They also point to displaced opportunities for practice. Those are explanations to investigate, not measurements of a changed brain or proof that every AI-assisted task weakens the user.

It would be equally careless to dismiss the finding because the tasks involved fractions and reading passages. Random assignment gives these comparisons weight. The fair reading is narrower: in these settings, easy access to answers improved assisted performance while leaving people worse prepared to continue alone immediately afterward.

Some effort is worth keeping. Some is just work.

This research gives a reason to distinguish a skill you want to retain from a chore you want finished. A new employee learning to interpret a report needs opportunities to make the judgment. Someone turning an already-approved report into a different file format may reasonably want the software to do the whole thing. These are examples of different goals, not tasks tested in this study.

For a team, the useful discussion is which decisions people still need to make when help is unavailable. Protect time for those decisions during normal paid work. Do not turn concern about AI into an extra evening of compulsory practice, or quietly treat every saved minute as room for another assignment.

Nor should an assistant impose a lesson on someone who explicitly asked it to finish a routine chore. A tool should make the choice understandable: are you asking it to take the task away, or help you get better at doing it? The experiment does not settle how best to build that choice. It does show why immediate correctness cannot be the only measure of useful help.

You do not need to swear off a chatbot because of a ten-minute headline. You do need room to notice when the answer arrives but your willingness to attempt the next problem does not.