If my codebase needs to pass a manual code review which finds an endless slew of LLM hallucinations that need correcting, it adds an immense amount of time to the process to rectify them. And these tools create insanely long winded and bloated code bases so the review itself is much longer from the outset. We're not talking about rigorous mathematical proofs, it's much more akin to a smoke test, and even that is a huge burden to bear when using those tools such that I view them as useless from this caveat alone. And it's basic best practice in the industry to perform these reviews, they are vital to quality results.
Even hobbyist software dev work would feel this impact. If you aren't testing the functions an LLM writes you will be putting buggy work with major obvious edge cases and vulnerabilities out there. The tests they write for themselves are often incomplete, happy path, or plainly avoiding actual testing by just asserting true.
From seeing these shortcomings in a field where technical proof and automated testing is a possibility, I can only imagine the level of human review necessary on less quantitative work to ensure any sort of reliability or accuracy. They cannot be trusted as accurate or authoritative at any point in the chain. ML can be useful as statistics for fuzzy and approximate data, but that's really where the usefulness ends