I think that's a false impression brought about by the endless articles using bad API math. They aren't running the APi at cost, they are also rerouting easy task to smaller models and build fine tuned quants in the first week or two that they replace their models with. It's very noticeable, the models are only at their best in the first few weeks than there is noticable degradation.
I don't get how people can readily distrust OpenAI but drink the "there's no profit in it" koolaid. Granted, they aren't in the best of spots if businesses start running their own models but I think it's overblown a bit personally.
I can run a small model that's about as good as SOTA a year ago on my 8 GB card and run it for 15 people at a comfy rate, those giant GPUs they have in datacenters are each serving 100+ users, but their APIs are priced as if they needed one GPU per person.