← All episodes
Translated Strategy · · 7 min read

The Engine Got Smarter. Your Setup Didn't.

You have been improving your AI setup for a year. It got worse. Every rule you added fixed something real. None of them ever asked to leave. Here is the ninety-minute audit to prune what is quietly steering the work.

An ops manager at a service business opens her Claude Project on a Wednesday morning. The custom instructions have been there since February. Eleven pinned prompts. Four uploaded reference PDFs. Three examples from a client who is no longer a client. A guardrail she wrote after a bad output in April that she is pretty sure the model does not need anymore. She has not touched any of it in seven weeks. The Claude Project is still working.

It is also, quietly, doing something she does not remember asking it to do. The sales email drafter she runs on Wednesdays came out formatted as a table this morning. She stares at the table for a minute. She does not remember writing a rule about tables. She reads through her custom instructions and finds it. Line forty-three, added in May, after she got frustrated with a comparison summary that lost its structure. She had typed “format output as a table when comparing two or more items” and moved on with her day. She did not remember she had written it.

The engine got smarter. Her setup did not.

She has been improving that setup for eight months. Every rule she added fixed a real problem the day she added it. None of them ever asked to leave. What she is looking at is not a Claude Project anymore. It is a slow accretion. A system nobody can see, steering the work in ways nobody remembers deciding.

The one you keep polishing gets worse

Here’s the thing. The engine keeps upgrading. The models you use are measurably smarter than they were in February. Better at reasoning. Better at inference. Better at knowing when to ask a clarifying question instead of guessing. The setup around them, the recipe you have been building all year, has not been evaluated once. Not the custom instructions. Not the pinned prompts. Not the uploaded PDFs. Not the never do X unless Y rules you wrote in a rush during a bad Monday six months ago.

Every one of those rules fixed something real. That is the whole trap. If a rule fixed nothing, you would notice. It got kept because it worked. It stayed because there was no ritual for asking whether it still needed to work. And now the setup has thirty-eight rules, eleven of them are still doing real work, and the other twenty-seven are steering the model into corners nobody meant to build.

The output feels a little worse. The team keeps rewriting it. The rewrites feel like quality control. The setup passes every internal check.

“But it worked in February.”

Yeah. In February the model needed the training wheels. It needed the exact steps. It needed a specific output format. It needed the guardrail against a hallucination the current model does not make anymore. The training wheels are still bolted on. The kid learned to ride six months ago.

Two ingredients rot the fastest

Our Professional Recipe names seven things that make an AI job hold. Training, Context, Guardrails, Examples, Format, Escalation, Feedback. Do not memorize the list. Two of them are the ones that bloat quietly when nobody is looking, and those are the two to audit first. Context and Examples.

Context is the situation-specific detail. The customer type, the deal stage, the product line, the last three interactions, the constraint your business runs under this quarter. Context is right and useful the day you write it. Context is also the ingredient most likely to still be describing your business the way it looked six months ago.

Look at the shape of it across three shops.

The home services dispatcher whose routing rules still reference two trucks the shop replaced in April. The insurance renewal workflow whose carrier notes still cover a product the carrier discontinued in Q1. The medical practice manager whose intake script still asks about a billing code the office no longer bills. None of that is broken enough to trigger a fix. All of it is now steering. The model is doing exactly what the setup asked for. What the setup asked for stopped matching the business.

Examples is the “here is what good looks like” set. Three to five real work samples. Examples is a supernova ingredient. It is the highest-leverage thing in the recipe when it is right, and it is the fastest thing to rot in place. The client whose emails you loaded in March is no longer a client. Their voice was the anchor. Their style set the pattern. Your setup is still writing to their style, for your current customers. The dish tastes wrong, and nobody can quite name why.

Here’s the thing about Examples specifically. You did not do anything wrong. You loaded the best material you had at the time. Six months later, the material stopped being the best. The recipe kept using it because nobody wrote a rule that says “revisit the examples in ninety days.”

The corrections were never audited

There is another failure mode underneath all of this. It has a name. The Silent Critic. Every time the team caught a bad output over the last eight months, they fixed it. They rewrote the email. They swapped the number. They corrected the tone. They typed the version they wished the model had produced. Then they moved on.

None of those corrections went back into the recipe. Not the setup. Not the examples. Not the custom instructions. The lessons were real. They literally never turned into ingredients. The team is doing the work of the recipe every week. The recipe is not learning.

That is what a mature setup that has never been pruned actually looks like. Rules from February. Examples from a departed client. Guardrails against a hallucination the current model no longer makes. Twelve months of quiet corrections that live in your team’s heads and Google Docs and never made it back into the system prompt.

The receipts are on the model. The bloat is on you.

Monday move: the ninety-minute audit

Pick one AI setup you have been building for at least three months. One. Not the whole stack. One Claude Project. One Custom GPT. One automation. One Copilot template. One email drafter.

Paste the whole thing into a doc. Every custom instruction. Every pinned prompt. Every reference file description. Every guardrail. Every persona. Every format rule. If it steers the output, it is on the page.

Count the words. Just knowing the number is useful. If your setup is over five thousand words and it has never been audited, you already have your answer about whether it needs pruning.

Sort every line into three piles.

Still needed. The rule fixes a problem you would still notice if it went away tomorrow. It has been used in the last thirty days. You could name the exact scenario in one sentence.

Fixed a problem that is gone. The rule was written for a customer segment you no longer serve, a product you sunset, a model quirk the current model does not have, a client whose examples are stale, a Slack channel that is archived. Retire it. Not revisit later. Retire it now.

Not sure. The middle pile is the tell. If you cannot decide whether a rule is still doing useful work, the odds it is doing useful work are low. Retire the middle pile too. If you retire something that was actually load-bearing, you will find out inside a week when an output goes sideways, and you can re-add it with a fresh diagnosis. The one you re-add is doing real work. The rest were not.

Then test the setup against the original job. Same input. Pruned setup. Does the output hold. Metric: word count before versus after, and did the task quality hold or improve. Most of the time it does. Sometimes it improves. Almost never does it get meaningfully worse.

One guardrail. If the setup has never been audited, do not build a new one this week. You already have a pruning problem. Adding a second overgrown garden does not help.

Then put a recurring block on your calendar. Every ninety days. Ninety-minute audit. Same three piles. Same test at the end. Prune is a rhythm, not a moment.

Does that make sense?

So.

The engine got smarter. Your setup did not.

You have been improving the wrong side of the equation for a year. The models have been upgrading themselves for free. The recipe you built around them has been accumulating. Every rule you added fixed something real, and none of them ever asked to leave.

Ninety minutes on a Wednesday. Three piles. One test at the end.

That is the whole audit.


Source influence: Nate B Jones on harness bloat and pruning. Distilled and operator-translated.

Framework spine: The Professional Recipe, the Context and Examples ingredients, with a callback to Quality Control failure mode #4 (The Silent Critic). Read the Professional Recipe.

~ source material · Source influence: Nate B Jones on harness bloat and pruning. Distilled and operator-translated.

~ keep going up next
~ if you got value here

Reu talks about this stuff on stages too.

Keynotes, panels, workshops. For conferences, operating companies, and trade associations.

Book Reu to speak →