← Ayo Osunjuyigbe

· AI

You can write an essay you cannot quote

The MIT Media Lab preprint landed on arXiv on the 10th. The flattening started immediately: ChatGPT fries your brain, MIT proved it, tell everyone. I sat down and read the thing they actually ran.

Nataliya Kosmyna, Pattie Maes, and six co-authors put 54 people — Boston-area, ages 18–39, mean 23, mostly students from MIT, Wellesley, Harvard, Tufts, Northeastern — in a 32-channel EEG net and made them write SAT essays for twenty minutes. Three groups. ChatGPT only. Google with -ai tacked on so the search box would not cheat. Or nothing: no tabs, no model, just what was already in their head. Same people, three sessions. Eighteen of them came back for a fourth, and the researchers swapped the tools.

They call the bill cognitive debt. Not AI makes you stupid. You skip the expensive part now. You pay later when you need the idea without the tool.

I have been paying that bill on code for a year. This is the first time I have seen someone put a headset on the writing version instead of arguing from a vibe.

What they asked

Essay writing is a nasty task on purpose. You hold the argument and the sentence at the same time. Schools already use it as a proxy for can this person think on a clock. The prompts are old SAT topics — happiness, philanthropy, art, courage — three choices per session, nine across the study.

The ChatGPT group got a lab account and was forbidden from opening a browser. The search group could open anything except an LLM. The brain-only group got a blank page. Then an interview: can you quote a sentence from the essay you just finished, without looking? Do you own this text? Are you happy with it?

Those quote questions are the ones I cannot stop thinking about.

Session 1. Two different questions. Can you quote anything from the essay you just finished? Is the line actually in it? Fifteen of eighteen ChatGPT people could not produce a usable quote. None of the eighteen got an accurate one. Search and brain-only: two people each said they had trouble pulling a line; on the accuracy check, three search people and two brain-only people missed. Ownership: sixteen of eighteen brain-only people said the essay was fully theirs. The ChatGPT group split — nine claimed full ownership, three said they owned none of it, the rest split the credit 50/70/90. Search people split credit too, but nobody in that group or the brain-only group said they owned nothing. Satisfaction was high anyway: 17/18 in the ChatGPT group, 15/18 brain-only, 18/18 search. The tool makes a thing that looks done. You feel fine. You cannot remember a line.

They expected session 2 to look different because everyone now knew the quote questions were coming. Ability: brain-only 18/18, ChatGPT and search 2/18 still stuck. Accuracy: four ChatGPT misses versus two in each of the other groups. By session 3 the ChatGPT group was 13/18 I can quote and six misses on the accurate one. Search and brain-only were clean on ability. Two people in each of those groups still missed the accurate line. The ChatGPT people did not catch up. In the later sessions a lot of them stopped fighting the model and just pasted.

The essays still scored. Human teachers and an AI judge did not fail the ChatGPT papers. They noticed the sameness — same named entities, same n-grams, same ontology inside a topic. The model’s default SAT answer has a smell. Teachers closed the loop on structure. They did not need a watermark.

The headset is not a morality tale

EEG here is directed connectivity (dDTF) across bands, not a brain score. The pattern is boring and consistent. Connectivity scales with how much work you still have to do. Brain-only: widest networks. Search: middle. ChatGPT: weakest coupling. Search is not the same as ChatGPT. You still pick a result, reject a tab, hold a claim against a source. The model hands you the synthesis. The effort that actually sticks the idea in your head is what drops.

The paper does not give you a single X percent dumber number. Anyone who is already quoting one made it up. What the authors wrote is that the three groups are statistically different, and that the LLM sessions look like less coordinated effort in the bands they care about for attention and control.

Session 4 is the part they told language models to read first, which is a joke I assume they are already tired of.

Nine people who had written with ChatGPT for three sessions had to write without it, on a prompt they had already used. Nine who had been brain-only finally got the model. The LLM-to-brain people: seven of nine could not quote anything; one in nine produced a correct line. Only three of nine recognized all three of their old prompts. The brain-to-LLM people: nine of nine recognized every prompt; seven of nine quoted correctly. Their connectivity went up when the tool arrived on a topic they had already bled on. They prompted like the search group — information-seeking, not write the essay. The people who had lived in the tool for three sessions did not collapse to a session-1 novice, and they also never grew the session-2/3 consolidation the unaided group had. Above the floor. Below the people who had practiced without it. Intermediate. Stuck.

The authors’ own educational sentence, not mine: do the hard pass first, then let the model in. The other order is cheaper in the moment and thinner when you take the model away.

They say this last bit is preliminary. Eighteen people. See limitations.

What cognitive debt is not

It is not never use the model. I am writing this in a window that will get a pass from one. The paper is an essay study on Boston-area students, on a twenty-minute SAT clock, with ChatGPT only, over four months of scheduling, not four months of daily use. Eighteen people in the swap session. It is a preprint. EEG cannot see the hippocampus. They did not split idea-generation from typing. They did not test code, math, or a job you already know how to do.

If you flatten that into MIT proved ChatGPT causes brain damage, you are doing the thing the LLM group did to the SAT prompt: accepting the convenient summary.

What survives the flattening:

You can produce text you do not remember. Teachers can still score it. You will say you are satisfied. Ownership gets weird. The next time you have to do it cold, the people who never outsourced the first pass are the ones who still have the argument.

That is offload. We already knew the Google version: you remember the place, not the fact. The model is worse because there is no place. There is a paragraph that arrived finished.

When I ban the first draft

I am applying this to the work I actually do, not to a morality lecture about students.

Learning the thing. A codebase I do not own yet. A protocol I have not implemented. A school essay whose point I could not explain in the hallway. The first pass is brain-only, or search with the model’s autocomplete off. Then I let Cursor or ChatGPT attack the draft. Session 4 in this paper is that order, and it is the one where connectivity and recall both look like a person using a tool instead of a person being used by one. If I cannot point at the file and say what the function does without opening the chat, I did not learn the repo. I rented a paragraph about it.

Producing the thing I already know. Boilerplate. A regex I could write drunk. Reformatting a type I have written ten times. The model can go first. I am not trying to encode a new schema. I am trying not to spend Tuesday on glue.

The test I stole from their interview. Close the tab. Quote one sentence you just accepted. If you cannot, you did not write it, and you should not ship it as if you did — not a PR description, not a design doc, not a take-home. The teachers in this study still scored the homogeneous essays. Your reviewer might too. That is not the same as you being able to defend a line in standup.

Search, in their data, is the honest middle. It is slower than ChatGPT and more work than a blank page, and people still owned the text. If your research is one prompt and a paste, you skipped even that.

I am not throwing the model out of school. I am throwing it out of the first twenty minutes when the point of the twenty minutes is that you leave with the argument still in your head. Use it after. Use it to fight the draft. Use it to translate. Do not use it as the draft if you will later be asked, even by yourself, what you said.

Kosmyna’s group wants longitudinal work before anyone calls this a net positive for brains. Fair. The session-1 quote test is enough for how I work tomorrow. The essay is not yours if you cannot say a sentence of it out loud.

References