Picture two high school classrooms, same teacher, same syllabus, same final exam. In one, students have access to an AI chatbot that answers their math questions with the full solution, step by step, the moment they ask. In the other, an almost identical chatbot — same model, same look — refuses to give the answer: it only offers a hint, one step, and waits for the student to try the next one on their own.
Six months later, a group of researchers measured what had happened in each classroom. The results leave little room for ambiguity — and explain, better than almost anything else we've written about AI and learning, why the debate shouldn't be "should we use AI in class, yes or no?" but "who is configuring that AI, and how?"
The Same Model, Two Opposite Results
The study is real: nearly 1,000 high school students in Turkey, during the 2023-24 school year, split into three groups — one with no AI access (control), one with GPT Base (a standard ChatGPT that answers with the direct solution), and one with GPT Tutor (the same model, instructed to guide without revealing the answer) (Wharton/Sungu, via Tech & Learning).
On practice exercises, both AI groups improved massively over the control group: 48% better with GPT Base, 127% better with GPT Tutor. So far, AI looks like an outright success.
But the final exam, taken without any help, tells a different story. The GPT Base group — the one that got direct answers — scored 17% worse than the control group: all that practice-exercise improvement evaporated, and left students worse prepared than if they'd never touched the AI at all. The GPT Tutor group, by contrast, ended up equal to the control — neither better nor worse.
Alp Sungu, a Wharton professor and co-author of the study, puts it bluntly: "students potentially use it as an answer machine, rather than as a tool that facilitates learning." And he adds something worth underlining, because it dismantles easy optimism: "I thought GPT Tutor was going to beat the control group. That wasn't the case." Not even the well-designed version outperformed skipping AI altogether — it simply avoided the harm the poorly-designed version caused.
The Difference Is in the Design, Not the Model
If that finding sounds like "AI with hints doesn't help much," another study — this one from Harvard and MIT, with 186 university physics students — qualifies the conclusion in an important way (via Edutopia).
The AI tutor in this study followed two very specific rules, written literally into its instructions: give only one step at a time, never the full solution at once; and it could confirm whether an answer was correct, but never reveal it before the student attempted it. With just those two rules — nothing more — students who used the AI tutor achieved a learning gain more than double that of students who attended an active lecture, while also finishing the material 10 minutes faster and reporting higher motivation.
The researchers' explanation connects to something pedagogy has known for decades: the reason a one-on-one human tutor usually outperforms a lecture hall is personalized, immediate feedback. A well-designed AI tutor can replicate that part of the mechanism — but, as the Turkish study shows, replicating it badly is worse than not having it at all.
It's the same conclusion, in different words, reached by the OECD's PISA 2025 report, which we already covered on Papitek: the OECD itself sums up the philosophy as AI needing to be scaffolding, not a crutch — something that supports the student's reasoning, not something that replaces it. Two university studies and a report with data from 91 countries, all pointing to exactly the same place.
So Who's Configuring This in the Classroom?
This is where the optimism runs into reality. If the difference between "help" and "crutch" lies in how AI is configured and introduced into a specific task, the logical question is: who is actually doing that design work at your kid's school?
According to a recent Bellwether report, as of late 2024 only about half of U.S. school districts had provided any kind of training to their teaching staff on generative AI (Bellwether, "Productive Struggle"). The other half, in practice, left each teacher to figure it out alone. The report itself doesn't sugarcoat it: "the gap between what students are doing right now and what schools are prepared to address is widening," and adds a line that should unsettle any education system still sitting on its hands: "students can't wait for a district's AI policy and teacher training to slowly catch up."
This isn't a knock on teachers — it's almost the opposite. A teacher with no specific training on how to configure or choose an AI tool has no way of knowing, unless someone explains it, that the difference between "answers directly" and "gives a hint" can be the difference between helping a student and harming them. The report calls for, among other things, redesigning tasks so students have to show their thinking process, not just the final result — exactly the principle that separated GPT Tutor from GPT Base in the Turkish study.
What You Can Do Now, Without Waiting for School
While teacher training catches up — and the Bellwether report itself admits that will take a while — there's real room for a family to apply the same principle at home, with whatever AI tool they're already using:
Ask for explanations, not answers. Before accepting what the AI gives you, ask it "why?" or "walk me through the reasoning step by step" instead of copying the result directly.
Verify before using. Treating an AI's answer as a first draft that needs checking, not as a settled fact, is the habit that shows up most consistently in the studies that do show real learning.
Ask school if there's a clear policy — not just which tools are allowed, but how they're configured. "Does it give direct answers or guide with hints?" is a perfectly reasonable question for a parent-teacher meeting.
Don't confuse "uses a lot of AI" with "is learning a lot." The Turkish study makes it clear: students with GPT Base felt like they were doing better (they improved in practice) and ended up worse prepared. The feeling of progress doesn't always match real progress.
A Well-Built Scaffold Goes Unnoticed
You don't notice a good scaffold from the street once the building is finished — you only notice it if it was badly built, because then the building leans. Something similar happens with AI in learning: when it's well designed, you barely notice it was there, because the student is still the one doing the thinking. When it's badly designed, the result also takes a while to show — until the unassisted exam arrives.
The technology to do this well already exists, and so do the studies that prove it. What's missing, based on everything we've reviewed here, isn't more AI or less AI — it's more people trained to build the scaffold properly before handing it to students. Do you already know what kind of scaffold your kid is using this week?


