The error message makes no sense to you. You did not write the function it points to, or you did, three prompts ago, and you cannot say what it does. So the only move left is to paste the error back and hope.
Sometimes it works. Then the project grows, and each fix breaks something else. The ceiling is not your ability. It is that nobody, including you, holds the picture.
AI writes a large share of your code now, and that is fine. What you keep depends on how you run it. The difference is a short loop you run on purpose.
Even if AI makes you faster, code you don't understand is a loan
The obvious objection is that if the AI is faster, you can skip understanding and take the speed. On a throwaway script, you can. On anything you will live with, the speed is a loan, and understanding is how you pay it back.
The next feature. If you know the code, you notice that a new requirement fits an existing queue, or that it will collide with an assumption three files away. If you do not, you get a feature that works alone and fights the rest of the system.
The next prompt. "Fix the bug" gets a guess. "The retry path tells the queue the job is done after the send, so check that ordering" gets a fix. What you ask is limited by your model of the codebase.
The 2 a.m. incident. If you understand the code, you read the error and have a hypothesis in five minutes. If you do not, you are back to paste and hope.
A controlled study found AI users understood their own code less
In January 2026, researchers at Anthropic published a study that quizzed developers on code they had just written, some with an AI assistant and some without.
Here is the study. Shen and Tamkin (2026, Anthropic) ran a randomized controlled trial with 52 mostly junior developers learning Trio, a Python library all of them were unfamiliar with. 26 worked with an AI assistant and 26 coded by hand. Everyone then took a quiz on the code they had just written.
The people who used the assistant scored lower on the quiz, on average. Using AI did not guarantee the lower score. That is the finding most retellings skip, and it is the one that matters. Same tool, same task, different outcomes, and the people who kept more were the ones who worked differently.
The group with the assistant averaged 50% on the quiz. The hand-coding group averaged 67%, and that difference was statistically significant.
The group with the assistant finished about two minutes faster, and that difference was not significant. Speed barely moved. Understanding did.
The participants who retained more used the assistant to build comprehension while they worked: they asked follow-up questions, requested explanations, or posed conceptual questions. The pattern is worth copying, even though the study does not prove it caused the higher scores. The authors measured understanding right after the task, not weeks later, and the subgroups behind the "how they used it" finding are small. What the study does show is a significant gap in understanding between coding with the assistant and coding by hand.
One more line in the paper names the tool many of us use now:
"this setup is different from agentic coding products like Claude Code; we expect that the impacts of such programs on skill development are likely to be more pronounced than the results here."
The assistant in the study sat in a sidebar and wrote code when asked. An agentic tool reads your repository, edits files, and runs commands, so far more of the work can pass by without you. The loop below is built for that setting.
Junior engineers and vibe coders have the most to lose when AI does the struggling
Two kinds of engineers feel this most, and both are doing something reasonable.
If you are a junior engineer, you are being asked to produce like a senior while you are still learning to be one. Today's juniors are the seniors who will one day debug the outage and make the design call.
The time you once spent struggling with a codebase is now the time the AI spends writing it. You can ship, and you still need to learn. Those two goals pull against each other.
If you are a vibe coder, you can already build real things, and that is a genuine skill. The limit shows up when the next change depends on how the parts fit together. A model of the system is what lets you change one part without breaking three others.
One double-sent email, four steps: the AI lays out the trade-off, you decide
If you have ever had an app post something twice after a retry, this is that bug.
A queue worker sends a confirmation email, and when a job retries, the customer gets the email twice. Here is the afternoon with the loop, one step at a time.
Plan. Before you ask the AI anything, write down what you think. Three things: what you would look at first, what you think is wrong, and which parts you will write yourself. The parts you name are the ones you claim.
If you have no plan, write the guess you have: "I would start with the retry code. I do not know yet why it sends twice." A guess on the page is enough, because the next step needs something to compare against.
Here, you write that the worker sends the email and then tells the queue the job is done, so a timeout between those two steps makes the queue deliver the job again. Your fix is a check on a sent marker, keyed by the confirmation, before each send. You claim that check.
Diff and spar. Now ask the AI for its plan and read it beside yours, as a diff: agree where you agree, argue where you differ, before anything is built.
It agrees on the cause. It adds something you did not have: pass an idempotency key to the email provider as well, if it supports one. That is a unique ID the provider uses to ignore a repeated request.
It also pushes on your fix. If you set the marker before the send and the process crashes, the customer never gets the email. If you set it after, a crash in the gap still allows one duplicate.
Which failure is worse, a missing confirmation or a doubled one? The AI can lay out the trade-off, but you own what customers see. You decide a duplicate is the lesser harm and set the marker after the send.
That choice removes the failure you actually had, a double send on every timeout retry. It accepts a rare one, a crash in the gap between send and marker, and the provider key covers most of that.
Build with the parts you claim. You write the marker check. The AI builds the provider key and the test that simulates a retry.
Record. After the commit, write two or three lines in a file in your repo: what you predicted, what you demonstrated (what you got right without help), and what you missed. Yours reads: predicted the redelivery cause, correct; had not considered provider-side idempotency.
That second half is the useful one. It is what you will be glad to have at 2 a.m.
That is the whole loop, and it needs a text file and an AI chat. If you want it to run itself, Scaffold is a free, open-source plugin for Claude Code built around these four steps.
Run the same afternoon with it, and it asks for your prediction before it shows you its plan. The question, word for word from the plugin's instructions, ends:
"…what do you think needs to change, and why?"
Its plan comes back as a diff against yours, and it argues where it genuinely disagrees before it builds. The kind of question it asks is "if we shipped your plan as-is, which of my items would bite first?" The build waits on your plain yes.
After the commit, it writes the record into a ledger of what you demonstrated, plain markdown in the repo and kept apart from what the AI knows about the code. An entry for this afternoon would say that you predicted the redelivery cause correctly and did not predict the provider key. The plugin's template shows the shape of the evidence lines. A real entry carries the date where the placeholder sits:
- YYYY-MM-DD plan: predicted dedup unprompted. (+)
- YYYY-MM-DD commit abc123: missed retry-safety. (−)
Two unaided correct predictions on a concept (three if they look alike) move it to "understanding". When you say you understand something, you have dated evidence in your own repo to point to. It is an honest answer with receipts, one you can show a manager or read back yourself. Install it with:
/plugin marketplace add kayashaolu/systemthinkinglab
/plugin install scaffold@systemthinkinglab
Planning before you prompt costs minutes, so spend them where they pay
The loop costs something.
A few minutes of thinking before you see an answer. A slower start, and sometimes finding out your plan was the weaker one. Typing you could have delegated. A minute at the end of the task.
The answer is a budget. Spend your struggle where the understanding lives: the decision about the marker, the trade-off, the debugging you will do next month. Let the AI have boilerplate, test scaffolding, and plumbing you have written many times before.
Keep a size gate as well. A typo-level fix does not need the loop. Fix it, commit it, move on.
You can tell the difference by asking whether you could explain the change to a teammate if the AI were gone.
One test for AI-written code: could you explain it without the AI?
I use one test for all of this. If I cannot explain why it works, I do not control it. I rent it.
Try the loop on your next bug, by hand with a text file or with the plugin, and see which side of that test you land on.
Frequently Asked Questions
- Does using an AI coding assistant hurt your skills?
- It can. In a randomized controlled trial (Shen and Tamkin 2026, Anthropic) with 52 mostly junior developers learning the Python library Trio, the group using an AI assistant averaged 50% on a quiz about their own code, against 67% for the group coding by hand, a statistically significant gap. Using AI did not guarantee a lower score. Participants who used the assistant to build comprehension, through follow-up questions, explanations, or conceptual questions, retained more, though the study does not claim those habits caused the difference.
- Is understanding the code worth it if AI makes you faster?
- Yes, for work you will live with. The person who understands the code can diagnose a production incident at 2 a.m., make better design decisions because they know what is already there, and write better prompts and plans. In the study, the assistant group was only about two minutes faster, a difference that was not significant, while the quiz gap was.
- Are agentic coding tools riskier for skill development?
- The study's authors expect so. Their footnote says the setup was different from agentic coding products like Claude Code, and that they expect the impacts of such programs on skill development to be more pronounced than the results in the study. The plan-first loop is built for that setting.
- What is the plan-first loop?
- Four steps you run on each task. Plan: before you ask the AI, write what you would look at first, what you think is wrong, and which parts you will write yourself (the parts you claim). If you have no plan, a guess on the page is enough. Diff and spar: read the AI's plan beside yours and argue where you differ before anything is built. Build with the parts you claim: write what you claimed and let the AI build the rest. Record: after the commit, write two or three lines in a file in your repo on what you predicted, what you got right without help, and what you missed.
- What should I write in the record step?
- Two or three lines in a file in your repo: what you predicted, what you demonstrated (what you got right without help), and what you missed. Over time the file shows where your understanding is solid and where it is thin.
- Does running a loop slow you down too much?
- It can on some tasks, which is why the loop comes with a budget. Spend your struggle on the decisions and trade-offs where understanding lives, and let the AI handle boilerplate and plumbing. A typo-level fix skips the loop entirely.