Obviously fuck Meta, but the difference is the publishing. Scraping and copying for personal or business use is a civil matter. And machine learning from unlicensed material is also fine - as long as the material isn’t “memorized” and an AI model can’t reproduce it.
Aaron bravely published the papers which is a criminal matter. That is why I believe copyright and IP law is the real theft, they take down copies for people to learn from.
There are many many millions of people who read and learned from pirated textbooks and who use those skills to do things. Who watched pirated amines or comics use that to learn how to draw.
Basically we should not be siding with the unethical side of IP law just to oppose AI corporations. They can afford to buy the books and media, and a simple purchase will do. And they can afford the lawsuits. And they will figure out the memorization problem, so that AI models learn but not memorize (which is only happens in like 1% of the cases and only when you specifically ask for something “just like that”).
China (so far) is saying that the AI models they produce should be open weight and be available to all people on earth. So if we have any issues it should be with the monopolization of AI models that concentrate this new developing immense power of AI in the hands of a few plutocrats. Which the IP laws might actually help them with.
He didn’t publish them, as far as I know. The charges brought against him were centered around the allegedly “fraudulent” use of his JSTOR account and the fact that he downloaded the files from the premises of an institution he did not belong to by plugging his laptop in a network switch in a restricted area where he wasn’t allowed to be. The actual deed was minor, likely not even criminal as far as digital rights were concerned, and the charges were famously so out of proportion that even other attorneys and legal scholars questioned them publicly. The publishers wanted to make an example, and it drove Swartz into suicide before a proper trial could be held.
Meta and other companies do essentially the same thing at a much larger scale and with the clear intent to monetize it through publishing AI models built on top of it all. The main difference is the legal climate, which has changed since/due to Swartz, and the lack of clarity that the law has for these new applications. Substantively though, there is a good bit of hypocrisy here, and it makes sense to point this out.
Obviously fuck Meta, but the difference is the publishing. Scraping and copying for personal or business use is a civil matter. And machine learning from unlicensed material is also fine - as long as the material isn’t “memorized” and an AI model can’t reproduce it.
Aaron bravely published the papers which is a criminal matter. That is why I believe copyright and IP law is the real theft, they take down copies for people to learn from.
There are many many millions of people who read and learned from pirated textbooks and who use those skills to do things. Who watched pirated amines or comics use that to learn how to draw.
Basically we should not be siding with the unethical side of IP law just to oppose AI corporations. They can afford to buy the books and media, and a simple purchase will do. And they can afford the lawsuits. And they will figure out the memorization problem, so that AI models learn but not memorize (which is only happens in like 1% of the cases and only when you specifically ask for something “just like that”).
China (so far) is saying that the AI models they produce should be open weight and be available to all people on earth. So if we have any issues it should be with the monopolization of AI models that concentrate this new developing immense power of AI in the hands of a few plutocrats. Which the IP laws might actually help them with.
He didn’t publish them, as far as I know. The charges brought against him were centered around the allegedly “fraudulent” use of his JSTOR account and the fact that he downloaded the files from the premises of an institution he did not belong to by plugging his laptop in a network switch in a restricted area where he wasn’t allowed to be. The actual deed was minor, likely not even criminal as far as digital rights were concerned, and the charges were famously so out of proportion that even other attorneys and legal scholars questioned them publicly. The publishers wanted to make an example, and it drove Swartz into suicide before a proper trial could be held.
Meta and other companies do essentially the same thing at a much larger scale and with the clear intent to monetize it through publishing AI models built on top of it all. The main difference is the legal climate, which has changed since/due to Swartz, and the lack of clarity that the law has for these new applications. Substantively though, there is a good bit of hypocrisy here, and it makes sense to point this out.
I don’t think he did.
I thought that was where a bunch of the initial core of libgen/anna’s archive was from? Maybe I’m wrong