In addition to the possible business threat, forcing OpenAI to identify its use of copyrighted data would expose the company to potential lawsuits. Generative AI systems like ChatGPT and DALL-E are trained using large amounts of data scraped from the web, much of it copyright protected. When companies disclose these data sources it leaves them open to legal challenges. OpenAI rival Stability AI, for example, is currently being sued by stock image maker Getty Images for using its copyrighted data to train its AI image generator.
Aaaaaand there it is. They don’t want to admit how much copyrighted materials they’ve been using.
If I do a book report based on a book that I picked up from the library, am I violating copyright? If I write a movie review for a newspaper that tells the plot of the film, am I violating copyright? Now, if the information that they have used is locked behind paywalls and obtained illegally, then sure, fire ze missiles, but if it is readily accessible and not being reprinted wholesale by the AI, then it doesn’t seem that different from any of the other millions of ways we use data in every day life. Just because a machine learned it instead of a human, I don’t believe that it makes it inherently wrong. I am very open to discussion on this, and if anyone has a counter-argument, I’d love to hear it, because this is a new field of technology that we should all talk about and learn to understand better.
Edit: I asked GPT-4 what it thought about this, and here is what it said:
That’s very cool and all but while we have this debate there are artists getting ripped off.
You aren’t having a debate. You’re blindly claiming that artists are getting ripped off, because maybe they are a bit, or maybe they’re latching onto any reason that lets them still have professional careers in 30 years.
I’m not making blind claims. And I won’t point you to the sources either. I’m not making any homework for anyone today. Dig the subject and post us some information if you are really into the debate thing.
If you can provide some sources with real data from people that have proven a loss of income due to getting “ripped off” by AI, I’d love to look over it. Until then, it’s a witch hunt.
I can provide you with reddit posts from artists who are replaced by AI.
Would you like it served with a cup of tea and some sandwiches?
If you have some that have actual proof in them, sure. That’s exactly what I’m looking for. However, if it amounts to nothing more than hearsay, then no, I don’t think I want them.
Had you spent a minimum of time digging the subject you would know exactly what I’m talking about.
I’m not making your sandwich for you, you will have to make your sandwich yourself.
Burden of proof being what it is, I’ll leave the sandwich making to those with the meat and bread.
I won’t take any burden for your majesty. Do your homework.