“How dare AI companies pirate books to train their AIs! They should be paying for their training data!”
…
“Not like that!”
This bit of clickbait outrage has been making the rounds for weeks now. They’re using the standard approach to bulk scanning books. The “rare editions” they’re scanning are rare because nobody is interested in them. It’s stuff like old guidebooks from a particular town fair in 1970 or a book on crochet in early colonial Kansas or whatever, they were sitting in warehouses and would have eventually been pulped if not bought in bulk. This is their one shot at preservation, frankly. And they’re doing this because the law requires them to do it.
And in fairness, they should be able to share their scans - with each other, with the public, and with Archive.org.
But that means relaxing copyright, not getting even more tightassed about it.
Every fucking website has to find a scandalized take about the several companies buying one (1) of every book.
Most of them were equally scandalized when the same companies did the obvious thing and grabbed a shadow library torrent.
Almost none of them recognize that when they demanded things be done properly - that’s what this is.
I don’t recall anybody demanding AI training involve destruction of physical media
They can’t resell it, or their scan isn’t a “backup copy,” it’s just piracy again. Same reason they can’t share their scans of apparently-rare books nobody’s cared enough to put on Archive.org.
Cutting the spine off is the simple way of scanning a mass-produced book. The alternatives are Rube Goldberg devices, or abundant tedious human labor.
What are you asking for, if not this? A private library of obscure forgotten works? A warehouse full of carefully-preserved dead weight? This was always the immediate alternative, and we told you as much, when y’all huffed about simple piracy. Like the average person on Lemmy gives two shits about piracy, outside this context.
I asking for them to go out of business if their business model is destroying art to manufacture slop
When the bubble bursts, this tech isn’t going away.
Last year some blog recreated GPT-2, the first LLM ‘too dangerous to release!,’ for twenty bucks. Corpus to model within the hour. The billions spent on research since then have let newer models of similar size benchmark a hair below the big-iron cutting edge.
I recently found out Chroma, once a leading image model, was made by like two people, for presumably $200,000. Not nothing… but within reach of even small organizations. Recent pressure has been toward making models even smaller and less structured than that. Several updated video and editing models have quietly declined to release their weights, because they’re small enough and powerful enough that most people would just run them locally. That desire, and that potential, will not vanish for wishful thinking.
If models had to train on nothing but public-domain, BSD, and CC0 materials, they’d still work the same way. If the money dries up to-morrow, then what’s possible with hobbyist research alone will continue. If even they hit a brick wall, barely past current capabilities, that’s still a couple gigs of linear algebra that’ll try and do anything you ask.
Meanwhile - people will still make art. People make art even when it’s illegal. Photography didn’t stop anyone from drawing.
It feels like you are hearing what I’m not saying. I’m aware that LLMs as a technology are here to stay. I’m aware that the existence of slop is not an existential threat to the practice of art. Frankly, I wasn’t one of the folks bothered by AI training piracy, as long as the companies involved weren’t hypocritically attempting to protect any of their own IP. Neither of these points are material to my stance that a business model that relies on destruction without producing a more valuable output is not a societal good and that an optimal market (i.e.: one that optimizes for societal good) would not allow such a business to be sustainable or exist at any kind of capital-intensive scale. The fact that the market we have is bending over backwards to allow companies like this to continue to exist is bad. I couldn’t give two shits what tech is involved, destruction of art at the industrial scale in order to produce slop at digital scale is a bad thing
destruction of art at the industrial scale
One copy. One copy of each book. And they’re not doing sexy vault heists for the last original manuscripts, they’re just buying antiquarian junk by the ton, through a service that avoids giving them duplicates. If this process is a threat to the preservation of any particular work, it’s probably because that work is undesired by anyone else in the world. Its eBay listing would end with zero bids. If it wasn’t going from a warehouse to a scanner to the dump, it would likely go straight to the dump.
If you want this process to aid in general preservation, maybe take aim at how Anthropic got fined over a billion dollars for infringement. That’s their motivation for doing this - that’s the kind of money they risk if they share all their super-rare scans. Of garbage.
This is not proper either
Elaborate.


