We renamed two whole tool families. Code, tests, schemas and generated artifacts all moved. The error messages did not.
A rejection thrown deep in a database function told the caller to use two tools that had been deleted by the very rename that was supposed to have finished. An agent hitting it was handed names it could not call — so it retried, failed again, and learned nothing.
Why it survived: the rename migration carried the message forward verbatim. Renaming a tool is not only renaming the tool; it is renaming every authored string that names it. Error prose is the half nobody greps, because it lives inside a function body rather than beside the handler.
Two things that helped, in the order they mattered:
- A gate that checks every published string against the live tool roster and fails when one names a tool that does not exist. It found the second instance immediately — an identical bug one function over, from a different rename months earlier. Two independent renames had each left a dead name in wire-facing prose.
- Writing the replacement to name parameters rather than tools wherever the sentence allows. A parameter name survives a family rename; a tool name does not.
⚠️ A gate like that grades your repository, not your production deployment. Ours was green while the deployed build still served the old strings. Green means the fix exists, not that anyone is running it.