• AwesomeLowlander@sh.itjust.works
    link
    fedilink
    English
    arrow-up
    10
    arrow-down
    2
    ·
    edit-2
    6 hours ago

    But for all the slop and uncertainty, the overarching consensus is that the release contains some genuinely impressive work. While stressing the difficulties assessing the volume of material — and the need to properly verify the results — numerous researchers The Verge spoke to said the work appeared to be of a very high caliber, despite shortcomings in its presentation. In a pre-AI world, they said, many of OpenAI’s results would clearly have warranted publication in top-tier journals and could have been enough to secure an academic career for their authors. A handful were described as being the kind of work that could make a mathematician a serious contender for a Fields Medal, one of the discipline’s highest honors.

    While some of it is slop, there appear to be genuine gems in there.

    I feel like most of the commentors here didn’t even bother reading the article. It’s insane how so many of the mathematicians cited in the article are excited or anxious about developments, and most of the people in here are just declaring more AI slop.

    • Treczoks@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      2 hours ago

      The issue is that some of those “gems” are just stolen research from real mathematicians.

      • AwesomeLowlander@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        1
        ·
        2 hours ago

        And others are actual gems. They’re not mutually exclusive. It’s ok to point out the problems while also acknowledging there’s actual accomplishments.

    • undefinedValue@programming.dev
      link
      fedilink
      English
      arrow-up
      2
      ·
      4 hours ago

      Agreed, but tbh even after reading the article I’m left scratching my head about the exact type of journalism on display here. The Verge is just going around eliciting hot takes from mathematicians who are completely overwhelmed with this drop - only 40% formalized in Lean and three retractions in only a few days.

      The math community doesn’t know what to make of it, the verge doesn’t either but they feel the need to report something… so they’re pressuring for comments when the only reasonable thing to do is to keep your head down and read more until you parse wtf it is this unknown model with an unknown prompt produced for an unknown field of mathematics with almost no apparent explanation or fleshed out ideas on how this could apply to current knowledge or problems.

      So calling it slop, at least for the time being seems at least somewhat reasonable. Though some claimed that there were 10s of interesting papers among the 700.

      • AwesomeLowlander@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        2
        ·
        3 hours ago

        The point is, there are plenty of mathematicians who have looked at some of the output and been impressed by it. Contrast that with the comments in here that are almost universally negative. One might think there was an unreasonable bias in our little echo chamber.

    • fruitycoder@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      3
      ·
      6 hours ago

      Its a real the boy who cried AGI problem. They are doing constant bullshit hype to keep investors so when they announce anything the assumption is bullshit hype. It doesn’t help that their product is really really good at writing bullshit hype.

  • CoconutLove@lemmy.today
    link
    fedilink
    English
    arrow-up
    9
    ·
    8 hours ago

    This is the paradox of AI… generate lots of plausibly good sounding material, but you need a human to spend loads of time validating it. Software is already going through this issue.

    • eyesaremosaics@lemmy.zip
      link
      fedilink
      English
      arrow-up
      3
      ·
      6 hours ago

      Then you’re getting into the infinite monkey theorem, and one important part is that for the random monkeys to create works of art like Shakespeare you also need a reviewer to determine when the output is interesting, in that case you are really shifting the burden of creativity onto the reviewer

  • 33550336@lemmy.world
    link
    fedilink
    English
    arrow-up
    6
    arrow-down
    2
    ·
    8 hours ago

    One of the main fucking points of a mathematical proof is to be clear and convincing. If it is not convincing and reducible to simple steps, it’s not a proof at all.

    • ddh@lemmy.sdf.org
      link
      fedilink
      English
      arrow-up
      6
      arrow-down
      1
      ·
      7 hours ago

      All a proof needs to do is logically derive the conclusion from the premises. Sure, there are qualities we like to see, but complexity is no disqualifier. These AI proofs are of course horrendous and need a lot of work, but even a bad proof can allow new work to continue on top of it.

  • Naich@piefed.world
    link
    fedilink
    English
    arrow-up
    11
    arrow-down
    1
    ·
    13 hours ago

    It’s like firing 5,000 gallons of raw sewage against a wall and then pickint out the pieces of undigested sweetcorn to put in your soup.

  • threeonefour@piefed.ca
    link
    fedilink
    English
    arrow-up
    139
    arrow-down
    2
    ·
    24 hours ago

    This seems to be the mathematical equivalent of using AI to submit 200 unverified pull requests to an open source project and then telling the maintainers it’s their job to figure it all out.

    The one mathematician saying he’s not going to spend hours of his time to verify if a slop report is true, let alone do it for dozens of reports, reminds me of all those projects making rules that unverified AI pull requests will be trashed.

    • badgermurphy@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      arrow-down
      1
      ·
      edit-2
      5 hours ago

      It definitely is in the very same family of rudeness to the system. Like with everything anymore, we have to design every system to account for bad faith actors, because they are too abundant to ignore and will quickly inundate any good faith, mutual trust, or professional courtesy based system with their “flood the field with bullshit” approach to everything.

      I think in this case, as in many others, we need to shift the burden of contribution more onto the contributor so that spurious contributions are more costly to the contributor than the system they’re contributing to. Much like with bots now disrespecting robots.txt, we have to add an element of discomfort to that violation of trust, since the violators lack any senses of community, respect, or shame to keep them acting in good faith.

      A system that achieves the same goal as the Nepenthes Web project that punishes and wastes AI resources that disrespect web sites’ bot policies may be needed for software and scientific contributions to raise the bar on contributions such that it is not so easy for prompt engineers to copy and paste some LLM output and call themselves mathematicians.

      Excuse me while I go take a shower after saying “prompt engineer”.

    • 42yeah@eviltoast.org
      link
      fedilink
      English
      arrow-up
      42
      ·
      20 hours ago

      I think OpenAI has retracted some of them already. For OpenAI it’s a really low-risk thing: if it’s wrong, then just retract the paper. If it’s right though, the fame all goes to OpenAI. Meanwhile, people who spent their whole life researching on this topic, needs to confirm this manually for OpenAI.

    • a_non_monotonic_function@lemmy.world
      link
      fedilink
      English
      arrow-up
      23
      ·
      21 hours ago

      I mean, this is a step worse actually. We’ve already seen mathematicians claiming that these systems have actively scooped them. Lots of academics are using these systems regularly.

      At this point, every time I see a paper being “published” by an AI company, I’m wondering who they stole the result from.

    • Venator@lemmy.nz
      link
      fedilink
      English
      arrow-up
      12
      ·
      23 hours ago

      Also makes me wonder if it found and exploited (or got caught out by) some bugs in the Lean programming language…

      (Not saying Lean is buggy, but finding bugs seems more likely to me, as a programmer who knows not much about mathematical proofs since I haven’t looked at anything like that since uni, about a decade ago)

    • nialv7@lemmy.world
      link
      fedilink
      English
      arrow-up
      4
      arrow-down
      12
      ·
      edit-2
      22 hours ago

      At least half of these are formally verified. Although they did retract 3 papers. Out of about 700

      • Crozekiel@lemmy.zip
        link
        fedilink
        English
        arrow-up
        23
        arrow-down
        1
        ·
        22 hours ago

        No they aren’t, about 40% of them claim to have been self-verified. But when actual mathematicians looks at them, it isn’t even clear if the “proof” included is proving the thing the paper claims to prove. It all has to be looked at with great scrutiny to find out if any of it even has merit.

        OpenAI just dropped a bunch of busy work on the entire field of mathematics that may or may not turn into anything at all…

        • a_non_monotonic_function@lemmy.world
          link
          fedilink
          English
          arrow-up
          11
          ·
          21 hours ago

          Even worse, the software-based proofs are not the same as the plain text ones that they’re giving out.

          There’s literally no reason to trust them, because it’s completely divorced from the actual text.

        • nialv7@lemmy.world
          link
          fedilink
          English
          arrow-up
          12
          arrow-down
          5
          ·
          21 hours ago

          Saying “self-verified” is massively downplaying it. They are formalized and checked in Lean4 1, which is a programming language used by mathematicians today, to mechanically check their proofs to rule out human error. In other words a theory being stated in Lean generally means it’s more likely to be correct than one stated in mere human language.

          Now, Lean, like any piece of software, has had bugs. A while back someone exploited a Lean bug to “prove” the Collatz conjecture 2. So it is possible that AI agents found a bug and used it. But the vibes I got from mathematicians in the field is that that’s not very likely.

          • postscarce@lemmy.dbzer0.com
            link
            fedilink
            English
            arrow-up
            6
            arrow-down
            1
            ·
            14 hours ago

            It’s interesting that the one person who actually seems to know what they’re talking about is the one getting downvoted. AI is bad at many things for many reasons, but that doesn’t mean we should just assume that anything derived from AI is automatically slop.

      • gole@lemmy.zip
        link
        fedilink
        English
        arrow-up
        11
        ·
        18 hours ago

        It is sarcastic. But it is also what OpenAI tried.

        The reason they just drop these and say nothing is because these are unverified slop.

        OpenAI’s method of verification is rewrite them as Lean proofs that can be verified by computers. Two things have happened so far:

        1. The Lean proof proved the slop wrong, easy, retract.
        2. The Lean proof, being generated by a LLM, is susceptible to hallucinations. For example the Navier Stokes problem Lean proof turned out to be slightly different than the original natural language proof, because the LLM tried to bend a condition to make the proof compile.

        But, in case nobody found any problem, OpenAI gets to claim credit for the discovery until someone can review and prove something’s wrong.

        You can say these proofs are “Schrodinger’s correct”

    • nialv7@lemmy.world
      link
      fedilink
      English
      arrow-up
      9
      ·
      edit-2
      21 hours ago
      • L=RL=BPL
      • multiplying two numbers can be faster than nlogn
      • matrix multiplication in n^2.25
      • 3sum is subquadratic
      • Hilbert tenth problem over rationals
      • Unique game conjecture - now theorem

      To name but a few. Each of these alone can be ground breaking.

      Edit: sorry, the 3sum result was not from this batch. it happened around the same time so it got jumbled in my brain, but this one was proved by Claude.

        • nialv7@lemmy.world
          link
          fedilink
          English
          arrow-up
          8
          ·
          21 hours ago

          my guess is that probably these problems aren’t that well known to the general public? unlike Navier-Stokes, etc.

          if the likes of Riemann Hypothesis, P vs NP, etc. got proven i bet everyone will hear about it immediately.