The fact that they "average" their inputs is why there are comparatively few typos (different sources have different typos, so they average out), not too many emojis or internet lingo (again, different sources use different ones in different places, so they average away), and why they produce such tedious stock output (it's an average of the inputs, so all the little quirks and idioms that make human communucation more vibrant have been blended away).
I'm sure there is some filtering on the inputs to try to remove the worst of it, but ultimately it's still just taking the rest and building it's probability tables from that, which leads to the homogenised outputs we see.
Mind you, having said there are fewer typos, the last time I bothered trying to get one to write some code, it managed to misspell a popular library name in multiple places, which gives some indication of how bad the inputs are, how bad the tokeniser is, or possibly both.