From Bug Report to Failing Test
1 promptShow
If you've been shipping with an agent doing most of the heavy lifting, sooner or later something breaks. A user hits a bug, or a page slows right down. So across this workflow the fixing side gets handed over too:
- Reproduce bugs → turn a vague report into something that fails on demand.
- Trace causes → an explanation with real files and lines, before any code gets touched.
- Find what's slow → real measurements instead of guesses.
- Watch production → errors that report themselves, so you're not waiting for someone to complain.
We're carrying on in FeedbackPit, the same app from the previous workflows, but all of this works on your stack. And here's what's just come in: "sometimes when I vote on a piece of feedback, the count goes up by two." That's the whole report. No reproduction steps, no browser information, nothing about when it happens, which is pretty typical.
The instinct is to dive straight into the voting code. Have a look around, spot something that might be causing it, change it. But if you can't trigger the bug yourself, you can never really be sure you've fixed it.
So your first job isn't to fix anything. It's turning that report into something that fails on demand, and a failing test is exactly what you want. It's the bug report rewritten as code.
One prompt gets you there. Ask the agent to explore the voting code, write a failing test that reproduces the report, and fix nothing yet. Then the part that matters: how confident is it that the test captures the reported behaviour, and why. You can steal this for pretty much any report, even one far longer than a single sentence.
What comes back. A race condition, and a test written around concurrent voting. High confidence the defect is real, and nothing fixed. Filter the suite down to that one file and "counts a single vote once when two requests race" fails. If you want richer output every time, turn the prompt into a skill.
Nothing's fixed yet, and that's still real progress. That's the first half of the loop: a report goes in, a failing test comes out. Next we get the agent to trace it to the actual cause and fix it properly, and the failing test is what makes that fix safe.