Executives working on AI at Microsoft and OpenAI admitted what its critics have been saying all along: Large language models are predatory pieces of technology that have been built on what a Microsoft executive called “an astonishing theft of unprecedented proportions,” and the “largest theft of labor in human history.” An internal Microsoft document said generative AI products have created a “doom loop” that is killing “the entire web.”

Those statements and a series of other mask-off moments feature heavily in an unredacted court filing that was unsealed Thursday in the behemoth New York Times vs OpenAI copyright lawsuit that has been winding its way through the court system for years. In a filing asking for summary judgment (basically, a filing with the court asking it to rule), lawyers for the New York Times laid out a series of admissions made by Microsoft and OpenAI executives in documents and depositions that until now had remained either sealed or redacted at the request of Microsoft and OpenAI.

It’s easy to see why the AI companies wanted to hide this from the public. The statements, taken together, are some of the most damning indictments of the ways LLMs were trained, how they worked, and the immediate threat they pose to human labor. It is a reminder that even as AI becomes more powerful and companies try to shift the narrative to the supposed existential risk of “superintelligent” AI, the tools they have already built were created by stealing from human creativity and labor and are by definition existential threats to the human labor market.


  • RunawayFixer@lemmy.world
    link
    fedilink
    English
    arrow-up
    14
    ·
    2 hours ago

    The best succinct description of llm that I’ve read was “plagiarism machine”, because that’s basically what they are: automated plagiarism. At best they create a collage of other works, at worst they copy one work verbatim, but they never create something original because they can’t.

    Copying something is a lot easier and thus cheaper than creating an original work from scratch, so genuine creators cannot hope to compete on cost. Which leads to less people being able to afford to earn a living from creating works, which leads to less original works being created.

    Short term that’s great for the neoliberal company executives: fire the expensive artists, create cheap slop with the plagiarism machine, and thus maximize profits now. That the plagiarism machine becomes stagnant because not enough new original works are being created is a future problem, by which time the current executives will have already jumped ship.

    But while it’s great for the neoliberal business model, for the rest of society it will suck. So now that we have confirmation that the AI executives know how bad their products are, and that they were just publicly lying about it in the style of tobacco executives, will anything be done about this “astonishing theft of unprecedented proportions”? Personally I doubt it, there’s too much regulatory capture.

  • abcdqfr@lemmy.world
    link
    fedilink
    English
    arrow-up
    1
    arrow-down
    1
    ·
    12 minutes ago

    A vast minority of this community interacts with llm ever and it shows

  • [object Object]@lemmy.ca
    link
    fedilink
    English
    arrow-up
    2
    ·
    1 hour ago

    AI is the borg.

    It is inherently evil in its creation and uses.

    It doesn’t have to be, but it largely is.

  • Mrkawfee@lemmy.world
    link
    fedilink
    English
    arrow-up
    32
    ·
    4 hours ago

    Internet search is terrible now. Every website I go to reads like it was generated by an LLM.

    • clif@lemmy.world
      link
      fedilink
      English
      arrow-up
      10
      ·
      2 hours ago

      Because it was generated by a llm.

      I did a search for the torque spec on a castle nut last week and the second result was for a nut (as in, food nuts that you eat) website and the llm had gone all in on nuts and added a page about castle nuts. It even generated an image of a castle for the top because… castle nut.

      It proceeded to provide vague instructions and a torque recommendation of 2x the actual spec.

      Then there was the other one about a small 4 stroke engine where it stated “other sites will say you don’t need to mix oil with gas but you absolutely must!” (You don’t and shouldn’t) … I wonder how many people have fucked up their shit by trusting llm garbage without knowing better.

    • Doom@discuss.online
      link
      fedilink
      English
      arrow-up
      13
      ·
      3 hours ago

      Glad I spent my teens and 20s getting stoned and reading the shit out of wikipedia before this shit happened. Can’t trust anything written online now.

    • XiELEd@piefed.social
      link
      fedilink
      English
      arrow-up
      3
      ·
      2 hours ago

      I once looked for a tutorial on a process in a Minecraft mod and there was a site that gave out incorrect information that was obviously written by an LLM.

    • grrgyle@slrpnk.net
      link
      fedilink
      English
      arrow-up
      2
      ·
      2 hours ago

      Yeah I don’t even use ai so I can’t compare, but even paid search services are noticeably so much more frustrating.

  • grrgyle@slrpnk.net
    link
    fedilink
    English
    arrow-up
    5
    ·
    2 hours ago

    Actually surprised how candid some of these statements are. Maybe points to more internal resistance than I would have assumed… like, they see the problem.

    Their focus on threats to media companies is probably borne out of fear of litigation, so maybe they only pay a lil lip service to how they also hurt “content creators” (people).

    Could also be that they know that when talking to the representatives of runaway financial automatons like multinational corporations, they have to appeal in terms the machine will understand.

  • BeMoreCareful@lemmy.world
    link
    fedilink
    English
    arrow-up
    2
    arrow-down
    1
    ·
    1 hour ago

    I’ve started to think of LLMs like a worm, running around copying and pasting gobbledegook. It’s the ultimate filter, the gibberish curtain of advertising

  • melsaskca@lemmy.ca
    link
    fedilink
    English
    arrow-up
    3
    arrow-down
    1
    ·
    2 hours ago

    One more generation of this shit, two at most, and tech will be working for the people again, unless we blow everything to smithereens before then.

  • k0e3@lemmy.ca
    link
    fedilink
    English
    arrow-up
    17
    arrow-down
    6
    ·
    4 hours ago

    LLMs are NOT destroying the planet. It’s the humans running the companies making LLMs.

    • SaharaMaleikuhm@feddit.org
      link
      fedilink
      English
      arrow-up
      6
      ·
      2 hours ago

      Okay, the let’s destroy the humans running the companies making LLMs before they destroy us. Fetch the guillotines!

    • BilSabab@lemmy.world
      link
      fedilink
      English
      arrow-up
      5
      ·
      4 hours ago

      technically, it’s the hardware that helps humans destroy the planet. and that hardware is astoundingly inefficient in what they’re trying to do with it. if you need to stack a literal data center with GPUs - maybe you need to develop some new tech.

      • k0e3@lemmy.ca
        link
        fedilink
        English
        arrow-up
        2
        arrow-down
        1
        ·
        3 hours ago

        But the LLM and the hardware is not making the decision and approving anything. It’s the humans.

        • BilSabab@lemmy.world
          link
          fedilink
          English
          arrow-up
          4
          ·
          3 hours ago

          humans destroy the planet with gear. that’s what i meant. and it’s the humans who think they can build superintelligent AI with tech and gear so woefully unfit for that it is not even funny.

  • CosmoNova@lemmy.world
    link
    fedilink
    English
    arrow-up
    43
    ·
    7 hours ago

    They knew what they were doing every step of the way. They are criminals that need to be disarmed and locked away. And we need to create a new Internet from scratch somehow thanks to these donkeys.

  • Grandwolf319@sh.itjust.works
    link
    fedilink
    English
    arrow-up
    26
    ·
    7 hours ago

    “Our AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,’” the document said.

    So basically the ouroboros

    • BilSabab@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      4 hours ago

      funny thing is that this ouroboros thing been known ever since LLM concept became a thing. that’s why it took forever to take off beyond r&d experiments. basically, the only remotely adequate way to apply LLM is what NotebookLM does - which literally parsing an uploaded document and reiterating it in some preset format. that’s literally it.

  • kablez@lemmy.world
    link
    fedilink
    English
    arrow-up
    37
    ·
    8 hours ago

    Gonna get real weird soon when they run out of rich new training data and they begin to consume their own shit. When that happens their entire model will collapse and if the bubble hasn’t popped already that may be what causes it.

    • WorldsDumbestMan@lemmy.today
      link
      fedilink
      English
      arrow-up
      1
      ·
      7 minutes ago

      Or they can just keep a database of actual data on their servers, instead of getting new data every single time for some fucking reason.

    • MalReynolds@slrpnk.net
      link
      fedilink
      English
      arrow-up
      7
      ·
      6 hours ago

      Pretty sure they mostly use the pre-AI internet (that they scraped and kept) and synthetic data currently. Probably trying (and failing so far or we’d have heard) to adapt to using video as training material at the moment, but developments there will likely apply to robotics at some point. Here’s hoping the current chuds have crashed and burned before then and that some sanity has taken over from unfettered capitalist oligarchs dreams of computer slavery.

        • MalReynolds@slrpnk.net
          link
          fedilink
          English
          arrow-up
          1
          ·
          1 hour ago

          Flock does pattern recognition, a quite old piece of machine learning. Nothing to do with training a large language model or other ‘AI’ model.

    • RepleteLocum@lemmy.blahaj.zone
      link
      fedilink
      English
      arrow-up
      23
      arrow-down
      1
      ·
      8 hours ago

      They’re already doing it. They call it distilling when they take it from another llm. Pretty sure most content was already stolen in the early days and they now rely on distillation and stealing new content.

        • BilSabab@lemmy.world
          link
          fedilink
          English
          arrow-up
          1
          ·
          2 hours ago

          in a manner of speaking - you always end up there. not by design though. models operate via continuous refinement and you can only optimize a model so much until it is a mess and you need to figure out where to roll back. so you either get shit like semantic drift or variance decay and you can whack a mole it to an extent but then you hit the rlhf wall when the model starts gaming its reinforcement framework and the fat lady sings.

        • DeadDigger@lemmy.zip
          link
          fedilink
          English
          arrow-up
          1
          ·
          2 hours ago

          It’s either agi or model collapse. For every AI system actually. If you have a high enough adoption you start to muss original data so if your AI is not self sufficient in time it will collapse, because it will be trained on its own data, which just is an incentive loop

    • Thorry@feddit.org
      link
      fedilink
      English
      arrow-up
      3
      arrow-down
      1
      ·
      6 hours ago

      Why do you think there has been such an emphasis on hacking with LLMs lately (especially by OpenAI). They figured out all those vulnerability databases were an excellent source for training material. In the past they scraped those, but just for general language training. Now they’ve specifically trained the models on the information within. Some team figured out how to use that data to train a model and have testing scenarios automated so they could write a good reward function. It wasn’t that they figured out the models are good at hacking, they ran out of content and found a new source of good data.

      With all the books they’ve been scanning I wonder if the next thing is going to be a writing assistant or editor or something like that. Even though writing good books is an art form and the actual writing down of the words is the easiest part (still not easy tho).

      These companies are starving for content and they’ve not just poisoned buy absolutely destroyed the content well that is the internet. Given they were already hitting diminishing returns hard, it doesn’t matter too much to them probably. But more compute and storage has also been hitting diminishing returns hard and customers are complaining about the cost. So they are getting a bit desperate on how to improve these things at all.