you are viewing a single comment's thread
view the rest of the comments
[–] 136 points 21 hours ago (10 children)

This seems to be the mathematical equivalent of using AI to submit 200 unverified pull requests to an open source project and then telling the maintainers it's their job to figure it all out.

The one mathematician saying he's not going to spend hours of his time to verify if a slop report is true, let alone do it for dozens of reports, reminds me of all those projects making rules that unverified AI pull requests will be trashed.

  • source
  • hideshow 10 child comments
  • [–] 1 point 5 hours ago* (last edited 2 hours ago)

    It definitely is in the very same family of rudeness to the system. Like with everything anymore, we have to design every system to account for bad faith actors, because they are too abundant to ignore and will quickly inundate any good faith, mutual trust, or professional courtesy based system with their "flood the field with bullshit" approach to everything.

    I think in this case, as in many others, we need to shift the burden of contribution more onto the contributor so that spurious contributions are more costly to the contributor than the system they're contributing to. Much like with bots now disrespecting robots.txt, we have to add an element of discomfort to that violation of trust, since the violators lack any senses of community, respect, or shame to keep them acting in good faith.

    A system that achieves the same goal as the Nepenthes Web project that punishes and wastes AI resources that disrespect web sites' bot policies may be needed for software and scientific contributions to raise the bar on contributions such that it is not so easy for prompt engineers to copy and paste some LLM output and call themselves mathematicians.

    Excuse me while I go take a shower after saying "prompt engineer".

  • source
  • parent
  • [–] 41 points 17 hours ago (1 child)

    I think OpenAI has retracted some of them already. For OpenAI it’s a really low-risk thing: if it’s wrong, then just retract the paper. If it’s right though, the fame all goes to OpenAI. Meanwhile, people who spent their whole life researching on this topic, needs to confirm this manually for OpenAI.

  • source
  • parent
  • hideshow 1 child comment
  • I mean, this is a step worse actually. We've already seen mathematicians claiming that these systems have actively scooped them. Lots of academics are using these systems regularly.

    At this point, every time I see a paper being "published" by an AI company, I'm wondering who they stole the result from.

  • source
  • parent
  • [–] 12 points 20 hours ago*

    Also makes me wonder if it found and exploited (or got caught out by) some bugs in the Lean programming language...

    (Not saying Lean is buggy, but finding bugs seems more likely to me, as a programmer who knows not much about mathematical proofs since I haven't looked at anything like that since uni, about a decade ago)

  • source
  • parent
  • [+] -8 points 19 hours ago* (last edited 19 hours ago) (4 children)

    At least half of these are formally verified. Although they did retract 3 papers. Out of about 700

  • source
  • parent
  • hideshow 4 child comments
  • [–] 22 points 19 hours ago (3 children)

    No they aren't, about 40% of them claim to have been self-verified. But when actual mathematicians looks at them, it isn't even clear if the "proof" included is proving the thing the paper claims to prove. It all has to be looked at with great scrutiny to find out if any of it even has merit.

    OpenAI just dropped a bunch of busy work on the entire field of mathematics that may or may not turn into anything at all...

  • source
  • parent
  • hideshow 3 child comments
  • [–] 7 points 18 hours ago* (1 child)

    Saying "self-verified" is massively downplaying it. They are formalized and checked in Lean4 1, which is a programming language used by mathematicians today, to mechanically check their proofs to rule out human error. In other words a theory being stated in Lean generally means it's more likely to be correct than one stated in mere human language.

    Now, Lean, like any piece of software, has had bugs. A while back someone exploited a Lean bug to "prove" the Collatz conjecture 2. So it is possible that AI agents found a bug and used it. But the vibes I got from mathematicians in the field is that that's not very likely.

  • source
  • parent
  • hideshow 1 child comment
  • [–] 5 points 12 hours ago

    It’s interesting that the one person who actually seems to know what they’re talking about is the one getting downvoted. AI is bad at many things for many reasons, but that doesn’t mean we should just assume that anything derived from AI is automatically slop.

  • source
  • parent