| LatchKey.ai | Archive | About | Consulting | Forward |

Issue #23: The Success of Failure
Lead Line:
The push is coming from every direction now. The trade press says this is the year AI graduates from toy to workforce. The vendor at the booth swears his platform could be running half your shop by Labor Day. The startup kids say just use it already, what's the holdup. Even Brad, two barstools and two Caymuses down, who still calls it "the ChatGPT," has a position. Put it to work.
Everyone agrees. For real this time.
Yet under all that noise, the same question that’s kept the whole project idling in the driveway since last year sits unanswered. But it’s getting louder now. Can I trust it?
We keep asking it like it’s a verdict. Guilty or not guilty. And the trial never ends, because the defendant refuses to sit still. It nails the contract summary on Tuesday. It invents a court case on Thursday. Verdict vacated, jury re-seated, start over.
The marketplace offers two answers, but at this point they both feel more like insults. "Always verify everything" isn't trust, it's a second job. "You're just prompting it wrong" isn't advice, it's a scold with a course to sell. So the mandate keeps getting louder while the question stays open, and after enough Tuesdays and Thursdays it curdles into the version nobody says at the conference but everybody mutters in the parking lot:
If I have to babysit it, is it really working for me?
Fair question. Here's a better one.
What would proof actually look like? If this trust is supposed to be earned, what would earning it even mean? More than a good week, surely. A benchmark? An exam? Who writes it? Who grades it? What's on it?
Don’t answer that, because the honest answer to the trust question just happens to be the one sentence nobody with a booth or a course can afford to say.
No one knows.
Not the vendors, not the analysts, not the people writing newsletters (points to self). There is no golden scorecard coming, no threshold where a light turns green or a gavel drops and the industry collectively announces that trust is now officially solved.
The tool is too fluid a thing to pin a verdict like that on. It changes under our feet monthly, sometimes weekly. And then there's the part that just sounds wrong until it finally clicks: the fallibility is a feature, not a bug.
The same flexibility that lets it draft a contract summary it was never specifically built to write is the flexibility that lets it be wrong in ways no spec sheet could fence in. The range and the wrongness are one property. That property is what makes the thing special, and it's also why every trust rating we will ever be handed is insufficient, because it was measured in somebody else's shop, on somebody else's work, against somebody else's bar.
Which means the verdict we've all been waiting for can only be forged in one place.
I spent more than twenty years around NASA engineers, and they had a relationship with failure that would look deranged anywhere else. When they built the parachute meant to land a one-ton rover safely on the surface of Mars, they tested it under the harshest conditions they could muster. There were people whose entire job was orchestrating the breaking. And they rooted for it to fail. Out loud. Because failure is where the learning lives. If the chute just works, you didn't push the system far enough, and a system you haven't pushed to failure is a system you can’t truly rely on. So you test it to the breaking point to learn its true boundaries. Then you fly it well inside them.
That philosophy ports into any shop that would rather measure than guess.
Stop auditioning the machine on live work and just hoping. Design tests meant to break it. Pick a task you could grade blindfolded, work you already know cold, so there's no debate about what right looks like. Then push. Longer documents. Messier inputs. The ugly edge cases you'd never lead with. Keep going until it drops something, and when it does, resist the flinch. Root for the break. The failure point is the prize.
Once you find it, that boundary is worth more than a hundred flawless demos, because now the calibration writes itself: real work rides below the line, with margin. The lessons get codified into guardrails you can actually rely on, because you measured them yourself. "It summarizes; it doesn't calculate." "Fine to draft; I hit send."
And the tasks where it flat misses the bar entirely? Just fire it from those. No drama, no ‘verdict on AI’, no referendum on the future. Just a tool operating within its own constraints. Clear because you defined them. As much a genius in its lane as a disappointment outside of it. Nobody throws a table lamp into the pool to light the water, and nobody calls the lamp a failure for it. This tool, this job, this shop, this month. Dated, and revisable the next time the ground shifts.
The reflex objection: “who has the time to run a test lab on top of a day job?” Ah, but here's the beauty part. The same relentless speed that makes this thing impossible to certify makes it absurdly cheap to interrogate. Design the test once, or better yet, have it help you design the test once, and it will run a hundred variations before Brad can finish telling you about his new gravel bike. More data than we could ever need, faster than we can believe.
It’s the old Background Calculator again. Pricing this kind of diligence like a 2003 video render job: overnight and fingers crossed. But in 2026, with AI, digging the trench isn't the labor our reflexes think it is anymore. The trench practically digs itself. What it's waiting on is somebody to draw where it goes.
And this means that the trust issue is less about trusting the AI and more about trusting yourself. Not the greeting-card aphorism kind, or some buried instinct you're supposed to reach in and reawaken. The engineer's kind. The trust that you mapped your own need case and dug the trench exactly deep enough for the tool to move through without spilling over. Run a couple weeks of honest tests and you’ll come out with a ledger. What earned weight. What got fired. What checking actually costs. Your numbers, from your work, against your bar. It’s the one object in this entire discourse that cannot be sold to you off the shelf. But it can be built with you.
The bonus is that the same tests will also sort the people selling you miracle machines, but that's next month's conversation. What matters first is simpler. The next time the clamoring chorus of trade press, booth barkers, and Brad is insisting it's time to go all in or just sit it out, you will be the only person in the room who isn't guessing.
No one is coming to certify this thing for us. Nothing that moves this fast can hold still long enough for a verdict. But that was never a flaw. It's the deal. So push it. Break it on purpose. Alone or with help. Mark where and then fly under the line. The success of failure is that the trust it builds is real.
Rhythm Section:
Every receipt is what happens when trust is declared instead of measured.
MIT Project NANDA — ~95% of enterprise genAI pilots, no measurable P&L return, $30–40B in. Dry read: nobody found the boundary before the money did.
HBR workslop — workers estimate 15.4% of what they receive is AI workslop; ~2 hrs to fix each; 42% found the sender less trustworthy. Dry read: work shipped from above the failure line; the receiver runs the test the sender skipped.
WalkMe — 3,750 people, 14 countries: 54% bypassed mandated AI tools last month; 33% skipped AI entirely. Dry read: mandated trust vs earned trust. (51-days figure omitted.)
Epoch AI — Claude Opus 4.7, 14 hrs autonomous, 2–17 weeks of engineering, $251 in tokens (via Mollick). Dry read: the upside receipt; why the ledger has a date column.
Goldman Sachs (Schneider) — token use projected 24x to 120 quadrillion/month, 2026→2030. Dry read: the ground will not hold still; certification-by-waiting is structurally impossible.
Bridge:
Flip over the lamp on your desk. Somewhere on the underside, stamped or stickered, there's a small circled UL. It's on the toaster, the power strip, the string of patio lights you trusted all summer without a second thought. You have bought that mark ten thousand times. You have read it approximately never.
It exists because of a fair that kept catching fire. Chicago, 1893: the World's Columbian Exposition wired the Palace of Electricity with a hundred thousand light bulbs to show off the miracle of alternating current, and small fires kept breaking out in the cheap jute draping the structure. The insurers underwriting the spectacle had no way to tell the safe equipment from the fire-in-waiting. So they sent for a young engineer from Boston named William Henry Merrill, who did the one thing nobody selling the miracle wanted done. He tested it. The fair made it through without a major electrical fire, and the next year the underwriters staked Merrill to his own laboratory, where the entire job was pushing equipment until it failed: burning it, overloading it, shaking it until something let go, so the mark on the survivors could mean something.
That is the whole secret of the most trusted symbol in your house. It's not a promise. It's a record of survived abuse. Nobody ever handed down a grand verdict on electricity. Somebody just kept burning toasters until we could trust the boundaries.
Interlude:
No one hands you the boundary. You build the test, root for the break, and stamp the line into the tool yourself.
Liner Notes:
Mollick, "The twilight of the chatbots" (June 30, 2026). The full argument for what comes after the chat window, from a man who measures this stuff for a living. Also where the strangest receipt in this issue surfaced.
Feynman's Challenger appendix. Ten pages he fought to get into the official report. Read the odds management declared, then the odds the engineers measured, and sit with the gap.
Petroski, To Engineer Is Human. An engineer's 1985 classic on how every dependable thing around you got that way. Forty years old and fresh as the today’s headlines.
The dahart trust comment. A stranger on the internet, a hammer, a bag of nails, and a box of screws. The best sixty seconds you'll spend on the trust question all month.
The B-Side:
The Proof is Bang On
For most of firearms history, a new gun barrel was a gamble that paid out in fingers. But that kind of risk tends to slow sales. So the gunmakers built a ritual to replace the gamble. London chartered its proof house in 1637, and the test was as brutal as the promise of the product: load the barrel with a charge far past any sane working load, more than any customer would ever fire, and touch it off from behind a wall. Barrels that burst became scrap. Barrels that held were stamped into the metal with a proof mark, an indentation that couldn't rub off or be quietly forgotten. The mark was the certification that someone already tried to destroy this, and it refused.
Nobody bought a barrel on the reassurance of the salesman. They bought the scar.
Four centuries later, the proof houses still run, because they solved the trust problem the only way it’s ever been solved. Not with proof of what it can do, but a verifiable record of what it won’t.
Reader Signal:
Hit Reply To This Email And Let Me Know Your Thoughts
What's one AI question you're sick of not getting a straight answer to?
What's one task you've already fired your AI from? And one it's genuinely earned?
Thanks for reading,
-Ep
Miss any past issues? Find them here: CTRL-ALT-ADAPT Archive
Know someone with AI trust issues? Forward this. Help them find the line.
Did this newsletter find you? If you liked what you read and want to join the conversation, CTRL-ALT-ADAPT is a weekly newsletter for experienced professionals navigating AI without the hype.
LatchKey AI provides professional AI consulting and services.
For more information or to book a call, head to latchkey.ai


