Three things our AI agent couldn't tell us
A morning brief said its budget was gone after spending 2% of it. Three questions an agent gets confidently wrong, and why no fix is a better prompt.
Our cross-project agent (“Director”) prepares a briefing every morning and pushes it to a Telegram account. On August 11 it opened like this:
“Budget’s out mid-read — but I’ve got the live to-do desk plus fresh commits, and that’s enough to call it.”
Reasonable. Candid, even — a system admitting a constraint and shipping what it had.
It had spent twenty-eight cents of a fifteen dollar cap.
Chasing that one wrong sentence turned up two more things the agent couldn’t tell us, each worse than the last. What follows is the whole thread pulled, in the order it came apart.
Why it stopped
It will answer. It will not check.
A system that has spent two percent of its budget will tell you, with some feeling, that the budget is gone.
What actually stopped it was a step limit: eight tool round-trips, then the loop closes. The trace says so plainly — iterations: 8 against a maximum of 8. It ran out of turns, not money.
The cause was one line. When the loop breaks, it hands the model a closing instruction, and ours read:
Budget or step limit reached. Stop reading now and give Dan your best answer with what you have — name any gap you couldn’t close.
A disjunction, with no indication which half fired. The model resolved it the only way available to it: by guessing. There was no bug and no bad data. We had asked a question we already knew the answer to, and then believed the guess.
Most of that sentence was working correctly
Worth separating, because the instinct on reading a wrong error message is to distrust the whole apparatus. The loop noticed it was cut off, stopped cleanly, answered usefully from partial information, and named the gaps it hadn’t closed. That is exactly what you want from a bounded agent. The defect was the label, not the behaviour — and those get conflated all the time, usually by removing the half that worked.
Three limits wearing one name
Our loop can stop for three different reasons, and they point at three different fixes.
| Fires when | What it means | What you’d do |
|---|---|---|
| step limit | 8 tool round-trips used | more turns, or cheaper reads |
| cost cap | $15 spent | genuinely expensive — investigate |
| output cap | 200k chars of tool output | it’s reading huge files |
One message covered all three. And here is the part that turns a mislabel into something structural: an iteration costs about three and a half cents. Against a fifteen dollar cap, that is roughly four hundred iterations of headroom. The step limit binds at about two percent of the cost cap. The generous budget we had configured was unreachable; the loop always died on steps first, at a fiftieth of the spend we had authorised.
So “budget’s out” was wrong in a deeper way than a mislabel. That budget had never once been the thing that stopped it, and never could have been.
Where the number came from
Every constant in a config file is a prediction about the future, made by someone who hadn’t met it yet. Ours was set on June 29, in the same commit that first built the loop, and never touched again. That commit’s own test run used three iterations — so eight was “about two and a half times what we just saw,” a sane opening guess and nothing more.
Its neighbour in the same file, the output cap, carries a comment naming the exact incident it was written to prevent. The step limit has no such story. One was a response to a failure; the other was a number somebody needed.
And it governed two paths with nothing in common: a chat where a human waits — eight round-trips is about seventy seconds, which is the edge of tolerable — and an unattended job at seven in the morning with a ten-minute allowance and nobody watching. The unattended path was being sized by the chat path’s constraint.
Name the limit in the close-out. Tell the model which cap fired and rule out the others, so it never has to infer a cause it was never given.
Record it as a field, not prose. We now write the limit that fired into the decision log, so a forced close is a fact you can query rather than a claim you have to read.
Ask an agent to report what it was told, not what it must infer. It receives a truncation notice — that’s knowledge. It does not receive a reason for the loop ending — that’s a guess about its own machinery. Only ever ask for the first kind.
What it stopped seeing
Nothing failed. That is the part worth being nervous about.
Fixing the label left the better question standing: why did a routine briefing need eight steps at all?
Because it spent eight of its twelve tool calls on one file. Two whole-file reads and six searches, all against the same calendar. The obvious diagnosis is a stuck loop. It isn’t — every one of those calls succeeded, and every one was different.
A join, by hand
The calendar had crossed the read tool’s byte ceiling. The file was 51,100 bytes; the limit was 40,000. A whole-file read returns the first 40,000 bytes and a note saying it stopped.
11,100
bytes past the ceiling — holding the 8 newest to-dos of 58, three of which the briefing went on to cite.
So it went around through search. And search is line-oriented, while a calendar to-do is a record: title, due date and status sit on three separate lines, with long descriptions folded across continuations. No single search returns one whole item. To reconstruct “what’s overdue and what is it called,” you search once per field and zip the results by position. Six searches is about what that costs.
The degradation has a birthday
Had this always been happening, or was it new? Both halves are checkable, which is the whole argument for keeping a decision log. The calendar is version-controlled, so its size over time is a matter of record. Every morning run is logged with its call count.
The file crossed 40,000 bytes on August 3. The grinding starts August 6.
| Before Aug 3 | Aug 6 onward | |
|---|---|---|
| Calls on the calendar, per run | 0–1 | 4, 5, 5, 8, 9, 10 |
| Runs ending at the step limit | 2 in 22 days | 4 of the last 6 |
| Typical cost | $0.10–0.15 | $0.22–0.28 |
The three-day lag is the honest detail. Right after crossing, only about four items fell off the end and it could work without them. As the lost tail grew to eight — and as those newest items became the ones that mattered — it had to go get them. The degradation scaled with the size of the loss, which is why it ramped instead of snapping.
Six weeks of a working capability, then a silent stop, with no code change anywhere near it. The trigger was a to-do list getting longer.
The expensive failure is the merciful one
Everything above is the version that announces itself: grinding, limit hits, a cost bump, eventually a briefing bad enough that someone notices. That failure turns up on the invoice. We were lucky.
The same defect has a silent form. Here is every oversized file our agents actually read, across 45 days.
| File | Size | Visible | whole reads | searches | |
|---|---|---|---|---|---|
| leads.yaml | 371 KB | 10% | 8 | 2 | |
| project-plan.md | 90 KB | 44% | 7 | 16 | |
| calendar.ics | 51 KB | 78% | 23 | 36 | |
| NEWSROOM.md | 42.7 KB | 93% | 9 | 2 |
Read the last two columns, not the sizes. Where searches far outrun whole reads, the agent is fighting — visible, expensive, self-announcing. That’s the calendar and the project plan.
Now the first row. That file was opened eight times, a tenth of it seen, with no follow-up searches, no limit hit, no cost spike and no hedge in a single reply. It didn’t fight. It read a tenth of that file eight times and never once mentioned it.
A confident answer built on a tenth of a file is indistinguishable from a confident answer built on all of it.
The truncation notice was in every one of those eight results. It read the notice and moved on — which is exactly what act one predicts, arriving from a different direction.
One quieter still: a design doc of ours carries a marker on line two of its own front matter, read: full, meaning read this whole, never conclude from a part of it. It has been seven percent past the ceiling since it crossed. The doc says read it all. The substrate has quietly been unable to.
The guard worked perfectly
That 40,000-byte ceiling is layer two of a three-layer guard we built six weeks earlier, after a vague instruction sent a search across scraped HTML with single lines up to 731,000 characters. It exhausted memory, and what came back assembled a 3.77-million-token prompt the API rejected outright.
The guard did its job. That failure has never recurred. It had only ever been asked to protect the process, and nobody had told it the agent was in the room.
This is not an argument against clipping. It is an argument that a clip is a silent lie by omission unless something checks the size.
Two detectors, both computed outside the agent. First: a run that ended at its iteration limit — ours would have fired on August 6, five days early. Second, and the one that matters more: a whole-file read against a file bigger than the ceiling. Pure arithmetic — the trace has the path, the filesystem has the size, severity is one minus the ratio. That second one catches the case that produces no symptom at all.
Make the clip state its own size. Ours now reports bytes shown, bytes total and the percentage. Before, a ten-percent read and a ninety-nine-percent read produced the identical sentence.
Test that documents meant to be read whole still fit. A limit sized once, against data that grows, has an expiry date nobody scheduled. Make the calendar tell you.
What it figured out
It solved the problem. Then the turn ended.
Go back to those twelve calls and read them in order, because they are not the flailing they look like.
First the whole-file read, which came back truncated. Then the field-by-field searches, the workaround. Then — and this is the part I keep returning to — one more whole-file read, as though the front door might behave differently the second time. And then it abandoned the raw file entirely, listed the directory, and opened the assembled view: a page where the joining is already done, generated nightly, sitting right there the whole time.
It found the right answer on the twelfth call. The wall was at the thirteenth.
And the next morning it started cold and walked the same three steps again. It had done so for six consecutive mornings.
This is the failure that outlives the other two, because raising a limit doesn’t touch it. What an agent works out inside a turn does not survive the turn, unless something writes it down.
Persistence is not memory
We thought we had this covered, and the shape of being wrong is worth showing, because it is a common way to be wrong.
Every turn our agent takes is persisted — the reply, the tool trace, the cost, the iteration count. Over a hundred of them. And each new turn replays the last six turns for continuity. That is a real capability and it is not memory:
| What replay gives you | What it doesn’t |
|---|---|
| The last 6 turns, verbatim | The 7th. Ours reached back four days. |
| Continuity inside one channel | Anything from another. The morning job had never read a word of the chat thread, or the reverse. |
| A transcript | A conclusion. Nothing is distilled, so nothing compounds. |
An agent with a rolling transcript can recall what it said. It cannot know what it learned.
The part we didn’t have to build
Here the story turns embarrassing in a useful direction. Asked whether we should build a memory, the honest answer looked like yes — until someone pointed out that another agent on the same box had been keeping one for five weeks.
A dated, append-only file of findings, written by the agent through a helper that refuses to edit prior entries, read at the start of every run. Its entries have titles like “occurrences #2 and #3: it’s a law, not an incident” — an agent promoting a recurrence into a rule, which is what memory is actually for.
Two details from that contract are worth stealing outright.
Newest entries go at the top, because the read tool returns a bounded prefix and an append-at-bottom file silently loses its newest entries first. Someone wrote that defence down a month before the same trap ate our calendar — and nothing propagated it from the one place that knew to the several places that didn’t.
And the rule that makes agent-written memory safe without supervision:
Memory is for pattern recognition — recurrences, trend lines, what a decision cost last time. Never for current state.
That single line answers the obvious objection. A store the agent writes itself could drift, go stale, become confidently wrong — unless it is structurally barred from being cited as current fact. The calendar says what’s due. The commit log says what shipped. The memory says what keeps happening. When they disagree about now, the memory loses.
Give it somewhere to write, and constrain the write. Append-only, dated, newest-first, corrections as new entries rather than edits. A memory that can rewrite its own history is worse than none, because nothing downstream can cite it.
Inject it; don’t make the agent go find it. Ours arrives in the turn state before the first tool call. Anything needed on every turn should never cost a tool call.
Rank it below live observation, explicitly, in the prompt. Memory for patterns, tools for the present. That precedence is the whole safety story.
The same shape, three times
Why it stopped, what it stopped seeing, what it worked out. Three questions, three confident wrong answers, and one shape underneath all of them: in each case the agent’s own account was the unreliable layer, and the fix was something outside it — a recorded field, an arithmetic check, a written file.
None of it needed an exotic setup. A read tool with a byte limit, a file that grows, and a loop with a cap is the entire recipe, and most agent stacks have all three by Tuesday. What made it visible here was a decision log nobody had thought to query that way.
The fixed run does the same job in three iterations instead of eight, costs ten cents instead of twenty-eight, and finishes on its own. On its first outing with a memory it spent the freed turns tracing a dependency nobody had followed through, and found six finished pieces of work that had been frozen for twenty-five days behind a condition that could never be met.
Which is the argument, really. None of this was about byte limits. It was about how much an agent can do once it stops spending its turns rediscovering things.