LLMs are notoriously bad at math.
Or rather, they're completely incapable of it. Even the most "advanced" AI systems basically just say "hmm... I think this might be a math problem" and then pass the data into a regular math API to do the calculation.
But a simple summarization bot isn't going to be wired up for calling a bunch of external tools, so... yeah, it just spits out an answer that looks statistically correct without having any real bearing on reality.