A few months ago I wrote about whether we were turning into mere execution engines, outsourcing the thinking and keeping only the typing. I'd watched a junior developer ship in twenty minutes what would have taken me an afternoon, and I couldn't decide if I was impressed or worried. Since then the question hasn't gone away. It's just gotten more specific.

I was watching the first couple of episodes of The Queen's Gambit recently, and there's a scene that stuck with me. Beth Harmon takes a tranquilliser before a big match because she's convinced she can't visualise the board without it. The show is careful about this. Her talent was real before the pills. What the pills gave her was a ritual she'd come to believe she needed. The dependency wasn't chemical, it was psychological, and that's a very different problem to solve.
I started asking myself the same question about how I use Claude. Not whether I'm relying on it, because clearly I am, every day, across half a dozen side projects. The real question is whether that reliance has quietly turned into belief rather than judgement. Am I still checking the work, or have I just decided I don't need to?
Dependency isn't the right frame
In the original post I argued that senior engineers catch subtle AI mistakes because they built intuition the hard way, line by line, and that newer developers might never build that foundation at all. I still think that's true. But framing the whole thing as dependency versus independence turned out to be the wrong axis. Nobody serious is suggesting we hand-roll memory management again because calculators made arithmetic obsolete as a skill worth guarding. The interesting question was never whether you rely on a tool. It's whether you still have a foothold on the problem when the tool is wrong.
Here's where I landed, and a few people pushed back on this when I said it out loud: I'd rather define the outcome, write the constraints, and trust the result when the tests pass, than insist on reading every function Claude produces. At the volume code gets written now, reading everything isn't just impractical, it's not even the right kind of attention. Understanding has moved up a level, from what does this function do to what does this system guarantee.
The catch is the one that actually matters. A test suite only encodes the failure modes you already thought to check for. If something breaks in a way your tests didn't catch, the real test isn't whether you can read the offending function. It's whether you can reason about why it broke at all, or whether you're stuck prompting the same tool that wrote the bug to go diagnose itself, with no independent foothold of your own. That's the Beth Harmon moment. Not using the tool, but losing the ability to function without checking first whether you still can.
What I actually changed
So the practical question became: how do you build guardrails for the failures you haven't thought of, rather than the ones you have? A test you write by hand only ever covers your own imagination, and that's the gap that matters, not the volume of code being produced. I ended up with a short sequence I now run on anything beyond a trivial fix, roughly in order of how much it costs to set up.
The cheapest one is also the easiest to skip under pressure, which is exactly why I don't let myself skip it. After Claude writes something, I ask it directly what inputs or conditions would break this, and what it had to assume because I didn't specify it. That's a genuinely separate pass from the one that wrote the code, and it surfaces assumptions neither of us said out loud.
The bigger shift is moving from example tests to property tests where it makes sense. An example test says input X gives output Y, which only ever covers cases someone thought to write down. A property test states an invariant, parsing then serialising should be lossless, this should never throw on any string, and a tool like fast-check generates hundreds of cases against it. You're no longer limited by what you personally imagined could go wrong.
For anything touching money, auth, or a public boundary, I've started running mutation testing as well. Tools like Stryker deliberately break the code in small ways and check whether the existing tests actually notice. A green suite that misses every mutation isn't testing anything, it's just reassuring you. And for anything non-trivial, I try to get a second, genuinely fresh pass, a different session with no memory of why the code was written a particular way, looking only at the diff and the stated goal. Having the rationale in your head biases you toward defending the choice rather than questioning it.
None of this is exotic. It's mostly existing testing discipline, pointed deliberately at the blind spot that matters now, which is the gap between what you tested and what you didn't think to.
The middle path, made operational
I ended the last post saying the answer sits somewhere between uncritical embrace and stubborn resistance, and that the people who thrive will keep enough expertise to tell when AI genuinely helps versus when it's quietly misleading them. I still believe that, but I'll admit it wasn't especially useful as a plan. It told me what stance to take without telling me what to actually do on a Tuesday afternoon with a pull request open.
This is the more useful version. Define the outcome, trust the process when you can see it working, but treat your own test suite as a record of what you already know to worry about, not a ceiling on what could go wrong. Build the second layer of checks deliberately rather than assuming the first layer caught everything. And watch for the one honest signal that something's actually wrong, not whether you're using the tool, but whether you can still reason your way to a failure it didn't catch.
Beth Harmon spends the whole story learning she never needed the pills, she just believed she did. I don't think the answer here is to prove you don't need Claude. It's to keep enough of your own judgement in the loop that you'd notice if you did.