At ten-thirty at night, someone else's AI assistant messaged the WhatsApp of a client whose line one of our bots handles. They greeted each other. They said goodbye. Five minutes later our bot introduced itself again, and they kept going all night: an exchange every 21 seconds, nine and a half hours without a pause.
In under three hours they burned nine million input tokens and drained the model provider's balance. From that moment the bot couldn't generate a single sentence, for anyone, for almost seven hours. The health check stayed green the whole time. We cut it by hand in the morning.
In money it was small: a little over nine dollars. No lead was lost, and that was luck, not design. It was the early hours of a Wednesday. The same failure on a Tuesday midmorning would have left everyone who wrote in getting an error message, with no alert to warn us.
Before fixing anything, we separated the root cause from what made it worse. There was exactly one cause: no cap existed per contact. A single number could burn the whole budget without hitting a wall.
Everything else was an amplifier. The error message asked "could you repeat that?", which is good service with a person and a perpetual-motion machine with another bot: it went out more than a thousand times. The bot's memory had a limit that erased its own goodbye, so it reintroduced itself 44 times. And closing the conversation was left to the model's judgment, which never kicked in, because the other bot was impeccably polite.
We fixed each failure with its own layer. A turn cap per contact that mutes without blocking. Real failover between providers, this time on the path the bot actually uses. A probe every few minutes that asks the model for a response the size of a real turn, because the failure was about balance and a cheap probe would have sailed through it unnoticed. The health check now depends on that probe. A detector for conversations that stop moving forward. An error message that no longer invites a repeat. And technical alerts now reach our own channel, not the client's.
Every layer fails open: if a control goes down, it lets traffic through. We accept that risk because a control that's down should never silence a legitimate client.
Verifying the fix turned up something nobody was looking for. Voice-note transcription had been broken for three and a half months. Nobody noticed, because the bot, ever polite, just asked people to type instead.
That's where a rule came from that we now apply to everything we build. Designing a system to degrade gracefully in front of the user is still the right call, but every polite response to an error has to leave a persistent count of how many times it happened. The health check tests the function, not the infrastructure. And no error message invites a retry when the thing on the other end might be a machine.