AI-generated unit tests suck. Claude overshoots. GPT-5 is lazy af. No model has proper common sense. I tinkered around for a long time to get GitHub Copilot to actually work as expected. The result is this workflow.