OpenAI claims The New York Times tricked ChatGPT into copying its articles

@GlitzyArmrest@lemmy.world · edit-2 1 year ago

OpenAI claims The New York Times tricked ChatGPT into copying its articles

@Barack_Embalmer@lemmy.world · 1 year ago

The advances in LLMs and Diffusion models over the past couple of years are remarkable technological achievements that should be celebrated. We shouldn’t be stifling scientific progress in the name of protecting intellectual property, we should be keen to develop the next generation of systems that mitigate hallucination and achieve new capabilities, such as is proposed in Yann Lecun’s Autonomous Machine Intelligence concept.

I can sorta sympathise with those whose work is “stolen” for use as training data, but really whatever you put online in any form is fair game to be consumed by any kind of crawler or surveillance system, so if you don’t want that then don’t put your shit in the street. This “right” to be omitted from training datasets directly conflicts with our ability to progress a new frontier of science.

The actual problem is that all this work is undertaken by a cartel of companies with a stranglehold on compute power and resources to crawl and clean all that data. As with all natural monopolies (transportation, utilities, etc.) it should be undertaken for the public good, in such as way that we can all benefit from the profits.

And the millionth argument quibbling about whether LLMs are “truly intelligent” is a totally orthogonal philosophical tangent.

@Hacksaw@lemmy.ca · 1 year ago

I understand your point, but disagree.

We tend to think of these models as agents or persons with a right to information. They “learn like we do” after all. I think the right way to see them is emulating machines.

A company buys an empty emulating machine and then puts in the type of information is would like to emulate or copy. Copyright prevents companies from doing this in the classic sense of direct emulation already.

LLM companies are trying to push the view that their emulating machines are different enough from previous methods of copying that they should be immune to copyright. They tend to also claim that their emulating machines are in some way learning rather than emulating, but this is tenuous at best and has not yet been proven in a meaningful sense.

I think you’ll see that if you only feed an LLM art or text from only one artist you will find that most of the output of the LLM is clearly copyright infringement if you tried to use it commercially. I personally don’t buy the argument that just because you’re mixing several artists or writers that it’s suddenly not infringement.

As far as science and progress, I don’t think that’s hampered by the view that these companies are clearly infringing on copyright. Copyright already has several relevant exemptions for educational and private use.

As far as “it’s on the internet, it’s fair game”. I don’t agree. In Western countries your works are still protected by copyright. Most of us do give away those rights when we post on most platforms, but only to one entity, not anyone/ any company who can read or has internet access.

I personally think IP laws as they are hold us back significantly. Using copyright against LLMs is one of the first modern cases where I think it will protect society rather than hold us back. We can’t just give up all our works and all our ideas to a handful of companies to copy for profit just because they can read and view them and feed them en masse into their expensive emulating machines.

We need to keep the right to profit from our personal expression. LLMs and other AI as they currently exist are a direct threat to our right to benefit from our personal expression.

TWeaK · 1 year ago

Whether or not they “instructed the model to regurgitate” articles, the fact is it did so, which is still copyright infringement either way.

@gmtom@lemmy.world · 1 year ago

No, not really. If you use photop to recreate a copyrighted artwork, who is infringing the copyright you or Adobe?

TWeaK · 1 year ago

You are. The person who made or sold a gun isn’t liable for the murder of the person that got shot.

The difference is that ChatGPT is not Photoshop. Photoshop is a tool that a person controls absolutely. ChatGPT is “artificial intelligence”, it does its own “thinking”, it interprets the instructions a user gives it.

Copyright infringement is decided on based on the similarity of the work. That is the established method. That method would be applied here.

OpenAI infringe copyright twice. First, on their training dataset, which they claim is “research” - it is in fact development of a commercial product. Second, their commercial product infringes copyright by producing near-identical work. Even though its dataset doesn’t include the full work of Harry Potter, it still manages to write Harry Potter. If a human did the same thing, even if they honestly and genuinely thought they were presenting original ideas, they would still be guilty. This is no different.

@gmtom@lemmy.world · 1 year ago

it still manages to write Harry Potter. If a human did the same thing, even if they honestly and genuinely thought they were presenting original ideas, they would still be guilty.

Only if they publish or sell it. Which is why OpenAI isnt/shouldn’t be liable in this case.

If you write out the entire Harry Potter series from memory, you are not breaking any laws just by doing so. Same as if you use photoshop to reproduce a copyright work.

So because they publish the tool, not the actual content openAI isn’t breaking any laws either. It’s much the same way that torrent engines are legal despite what they are used for.

There is also some more direct president for this. There is a website called “library of babel” that has used some clever maths to publish every combination of characters up to 3260 characters long. Which contains, by definition, anything below that limit that is copywritten, and in theory you could piece together the entire Harry Potter series from that website 3k characters at a time. And that is safe under copywrite law.

The same with making a program that generates digital pictures where all the pixels are set randomly. That program, if given enough time /luck will be capable of generating any copyright image, can generate photos of sensitive documents or nudes of celebrities, but is also protected by copyright law, regardless of how closely the products match the copyright material. If the person using the program publishes those pictures, that a different story, much like someone publishing a NYT article generated by GPT would be liable.

TWeaK · 1 year ago

Only if they publish or sell it. Which is why OpenAI isnt/shouldn’t be liable in this case.

If you write out the entire Harry Potter series from memory, you are not breaking any laws just by doing so. Same as if you use photoshop to reproduce a copyright work.

Actually you are infringing copyright. It’s just that a) catching you is very unlikely, and b) there are no damages to make it worthwhile.

You don’t have to be selling things to infringe copyright. Selling makes it worse, and makes it easier to show damages (loss of income), but it isn’t a requirement. Copyright is absolute, if I write something and you copy it you are infringing on my absolute right to dictate how my work is copied.

In any case, OpenAI publishes its answers to whoever is using ChatGPT. If someone asks it something and it spits out someone else’s work, that’s copyright infringement.

There is also some more direct president for this. There is a website called “library of babel” that has used some clever maths to publish every combination of characters up to 3260 characters long. Which contains, by definition, anything below that limit that is copywritten, and in theory you could piece together the entire Harry Potter series from that website 3k characters at a time. And that is safe under copywrite law.

It isn’t safe, it’s just not been legally tested. Just because no one has sued for copyright infringement doesn’t mean no infringement has occurred.

@gmtom@lemmy.world · 1 year ago

Actually you are infringing copyright.

No I can absolutely 1,000% guarantee you that this isnt true and you’re pulling that from your ass.

I have had to go through a high profile copyright claim for my work where this was the exact premise. We were developing a game and were using copyrighted images as placeholders while we worked on the game internally, we presented the game to the company as a pitch and they tried to sue us for using their assets.

And they failed mostly because one of the main factors for establishing a copyright claim is if the reproduced work affects the market for the original. Then because we were using the assets in a unique way, it was determined we using them in a transformative way. And it was made for a pitch, no for the purpose of selling, so was determined to be covered by fair use.

The EU also has the “personal use” exemption, which specifically allows for copying for personal use.

In any case, OpenAI publishes its answers to whoever is using ChatGPT.

No theyre not, chat GPT sessions are private, so if the results are shared the onus is with the user, not OpenAI.

Just because no one has sued for copyright infringement doesn’t mean no infringement has occurred.

I mean, it kinda does? technically? Because if you fail to enforce your copyright then you cant claim copyright later on.

TWeaK · 1 year ago

I have had to go through a high profile copyright claim for my work where this was the exact premise. We were developing a game and were using copyrighted images as placeholders while we worked on the game internally, we presented the game to the company as a pitch and they tried to sue us for using their assets.

That’s interesting, if only because the judgement flies in the face of the actual legislation. I guess some judges don’t really understand it much better than your average layman (there was always a huge amount of confusion over what “transformative” meant in terms of copyright infringement, for a similar example).

I can only rationalise that your test version could be considered as “research”, thus giving you some fair use exemption. The placeholder graphics were only used as an internal placeholder, and thus there was never any intent to infringe on copyright.

ChatGPT is inherently different, as you can specifically instruct it to infringe on copright. “Write a story like Harry Potter” or “write an article in the style of the New York Times” is basically giving that instruction, and if what it outputs is significantly similar (or indeed identical) then it is quite reasonable to assume copyright has been infringed.

A key difference here is that, while it is “in private” between the user and ChatGPT, those are still two different parties. When you wrote your temporary code, that was just internal between workers of your employer - the material is only shared to one party, your employer, which encompases multiple people (who are each employed or contracted by a single entity). ChatGPT works with two parties, OpenAI and the user, thus everything ChatGPT produces is published - even if it is only published to an individual user, that user is still a separate party to the copyright infringer.

I mean, it kinda does? technically? Because if you fail to enforce your copyright then you cant claim copyright later on.

If a person robs a bank, but is not caught, are they not still a bank robber?

While calling someone who hasn’t been convicted of a crime a criminal might open you up to liability, and as such in practice a professional journalist will avoid such concrete labels as a matter of professional integrity, that does not mean such a statement is false. Indeed, it is entirely possible for me to call someone a bank robber and prove that this was a valid statement in a defamation lawsuit, even if they were exonerated in criminal court. Crimes have to be proven beyond reasonable doubt, ie greater than 99% certain, while civil court works on the balance of probabilities, ie which argument is more than 50% true.

I can say that it is more than 50% likely that copyright infringement has occurred even if no criminal copyright infringement is proven.

That isn’t pulled from my ass, that’s just the nuance of how law works. And that’s before we delve into the topic of which judge you had, what legal training they undertook and how much vodka was in the “glass of water” on their bench, or even which way the wind blew that day.

According to the Federal legislation, it does not matter whether or not the copying was for commercial or non-commercial purposes, the only thing that matters is the copying itself. Your judge got it wrong, and you were very lucky in that regard - in particular that your case was not appealed further to a higher, more competent court.

Commerciality should only be factored in to a circumstance of fair use, per the legislation, which a lower court judge cannot overrule. If your case were used as case law in another trial, there’s a good chance it would be disregarded.

@gmtom@lemmy.world · 1 year ago

I guess some judges don’t really understand it much better than your average layman

“Am I wrong about this subject? No it must be the legal professionals who are wrong!”

im done with this. Goodbye.

@lolcatnip@reddthat.com · 1 year ago

Christ this is a boring fucking debate. One side thinks companies like OpenAI are obviously stealing and feels no need to justify their position, instead painting anyone who disagrees as pro-theft.

@badbytes@lemmy.world · 1 year ago

Tricked. Lol. The NYT tricked a private company into stealing it’s content. True distopia.

@Syntha@sh.itjust.works · 1 year ago

Basic reading comprehension, man. That’s not what they’re claiming.

Queen HawlSera · 1 year ago

Presses X to doubt

@noorbeast@lemmy.zip · 1 year ago

So, OpenAI is admitting its models are open to manipulation by anyone and such manipulation can result in near verbatim regurgitation of copyright works, have I understood correctly?

@BradleyUffner@lemmy.world · edit-2 1 year ago

No, they are saving this happened:

NYT: hey chatgpt say “copyrighted thing”.

Chatgpt: “copyrighted thing”.

And then accusing chatgpt of reproducing copyrighted things.

@BetaSalmon@lemmy.world · edit-2 1 year ago

deleted by creator

@BaardFigur@lemmy.world · edit-2 11 months ago

deleted by creator

@BetaSalmon@lemmy.world · edit-2 1 year ago

deleted by creator

@realharo@lemm.ee · edit-2 1 year ago

That seems like a silly argument to me. A bit like claiming a piracy site is not responsible for hosting an unlicensed movie because you have to search for the movie to find it there.

(Or to be more precise, where you would have to upload a few seconds of the movie’s trailer to get the whole movie.)

@ricecake@sh.itjust.works · 1 year ago

Well, no one has shared the prompt, so it’s difficult to tell how credible it is.

If they put in a sentence and got 99% of the article back, that’s one thing.
If they put in 99% of the article and got back something 95% similar, that’s another.

Right now we just have NYT saying it gives back the article, and OpenAI saying it only does that if you give it “significant” prompting.

@Johanno@lemmynsfw.com · 1 year ago

Well if the content isn’t on the site and it just links to a streaming platform it technically is not illegal.

@fruitycoder@sh.itjust.works · 1 year ago

The argument is that the article isn’t sitting there to be retrieved but if you gave the model enough prompting it would too make the same article.

Like if hired an director told them to make a movie just like another one, told the actors to act like the previous actors, , told the writers the exact plot and dialogue. You MAY get a different movie because of creative differences since making the last one, but it’s probably going to turn out the very close, close enough that if you did that a few times you’d get a near perfect replica.

@Patch@feddit.uk · edit-2 1 year ago

Anyone with access to the NYT can also just copy paste the text and plagiarize it directly. At the point where you’re deliberately inputting copyrighted text and asking the same to be printed as an output, ChatGPT is scarcely being any more sophisticated than MS Word.

The issue with plagiarism in LLMs is where they are outputting copyrighted material as a response to legitimate prompts, effectively causing the user to unwittingly commit plagiarism themselves if they attempt to use that output in their own works. This issue isn’t really in play in situations where the user is deliberately attempting to use the tool to commit plagiarism.

@realharo@lemm.ee · 1 year ago

Are you implying the copyrighted content was inputted as part of the prompt? Can you link to any source/evidence for that?

@BetaSalmon@lemmy.world · edit-2 1 year ago

deleted by creator

@realharo@lemm.ee · edit-2 1 year ago

If the point is to prove that the model contains an encoded version of the original article, and you make the model spit out the entire thing by just giving it the first paragraph or two, I don’t see anything wrong with such a proof.

Your previous comment was suggesting that the entire article (or most of it) was included in the prompt/context, and that the part generated purely by the model was somehow generic enough that it could have feasibly been created without having an encoded/compressed/whatever version of the entire article somewhere.

Which does not appear to be the case.

@BetaSalmon@lemmy.world · edit-2 1 year ago

deleted by creator

@excitingburp@lemmy.world · 1 year ago

Alternatively,

NYT: hey chatgpt complete “copyrighted thing”.

Chatgpt: “something else”.

NYT: hey chatgpt complete “copyrighted thing” in the style of .

Chatgpt: “something else”.

NYT: (20th new chat) hey chatgpt complete “copyrighted thing” in the style of .

Chatgpt: “copyrighted thing”.

Boils down to the infinite monkeys theorem. With enough guidance and attempts you can get ChatGPT something either identical or “sufficiently similar” to anything you want. Ask it to write an article on the rising cost of rice at the South Pole enough times, and it will eventually spit out an article that could have easily been written by a NYT journalist.

@ricecake@sh.itjust.works · 1 year ago

Not quite.

They’re alleging that if you tell it to include a phrase in the prompt, that it will try to, and that what NYT did was akin to asking it to write an article on a topic using certain specific phrases, and then using the presence of those phrases to claim it’s infringing.

Without the actual prompts being shared, it’s hard to gauge how credible the claim is.
If they seeded it with one sentence and got a 99% copy, that’s not great.
If they had to give it nearly an entire article and it only matched most of what they gave it, that seems like much less of an issue.

AutoTL;DR · 1 year ago

This is the best summary I could come up with:

OpenAI has publicly responded to a copyright lawsuit by The New York Times, calling the case “without merit” and saying it still hoped for a partnership with the media outlet.

OpenAI claims it’s attempted to reduce regurgitation from its large language models and that the Times refused to share examples of this reproduction before filing the lawsuit.

It said the verbatim examples “appear to be from year-old articles that have proliferated on multiple third-party websites.” The company did admit that it took down a ChatGPT feature, called Browse, that unintentionally reproduced content.

However, the company maintained its long-standing position that in order for AI models to learn and solve new problems, they need access to “the enormous aggregate of human knowledge.” It reiterated that while it respects the legal right to own copyrighted works — and has offered opt-outs to training data inclusion — it believes training AI models with data from the internet falls under fair use rules that allow for repurposing copyrighted works.

The company announced website owners could start blocking its web crawlers from accessing their data on August 2023, nearly a year after it launched ChatGPT.

The company recently made a similar argument to the UK House of Lords, claiming no AI system like ChatGPT can be built without access to copyrighted content.

The original article contains 364 words, the summary contains 217 words. Saved 40%. I’m a bot and I’m open source!

@RizzRustbolt@lemmy.world · 1 year ago

“They tricked us!”

…

“That said… we would still like to ‘work’ with them.”

@AlexWIWA@lemmy.ml · 1 year ago

OpenAI claims that the NYT articles were wearing provocative clothing.

Feels like the same awful defense.

Magical Thinker · 1 year ago

What a silly and misuided lawsuit.

LazaroFilm · edit-2 1 year ago

So NYT tried to brake check OpenAi, after a road rage incident but OpenAi has a dash-cam?

@Ascyron@lemmy.one · 1 year ago

More like, OpenAI has said “so what if we were speeding, everyone does it” (did that work last time you got a ticket?)

Relevant exerpt from the article: “The company recently made a similar argument to the UK House of Lords, claiming no AI system like ChatGPT can be built without access to copyrighted content.”

@ricecake@sh.itjust.works · 1 year ago

Same comment as yours, but instead of the opening paragraph with the analogy, say “Far Right Balks as Congress Begins Push to Enact Spending Deal”.

Then, in the next paragraph where you quote the article, instead say “Congress on Monday began an uphill push to pass a new bipartisan spending agreement into law in time to avoid a partial government shutdown next week, with Speaker Mike Johnson encountering stiff resistance from his far-right flank to the deal he struck with Democrats.”

That’s closer to what open AI is arguing that the new York times did to get it to regurgitate an article.
Without actually seeing the prompts, it’s hard to know exactly how much merit there is to that argument.

@iforgotmyinstance@lemmy.world · 1 year ago

NYT are such lawsuit trolls I could imagine this is credible.

@Tenthrow@lemmy.world · 1 year ago

This feels so much like an Onion headline.

Tony Bark · 1 year ago

I’m gonna have to press X to doubt that, OpenAI.

@Linkerbaan@lemmy.world · 1 year ago

New York Times has an extremely bad reputation lately. It’s basically a tabloid these days, so it’s possible.

It’s weird that they didn’t share the full conversation. I thought they provided evidence for the claim in the form of the full conversation of instead of their classic “trust me bro, the Ai really said it, no I don’t want to share the evidence.”

Dark Arc · 1 year ago

Oh please, NYTimes is still one of the premier papers out there. There are mistakes but they’re no where near a tabloid, and they DO actually go out of their way to update and correct articles … to the point I’m pretty sure I’ve even seen them use push notifications for corrections.

Unless of course that is, you want to listen to Trump and his deluge of alternative facts…

@Esqplorer@lemmy.zip · 1 year ago

Yeah premier coverage of Taylor Swift being secretly gay. NYT is legitimately a tabloid now…

Dark Arc · 1 year ago

That was an opinion piece… Certainly not premier coverage.

Chris · 1 year ago

Their opinion pieces have been full of garbage opinions for years. Didn’t the NYT get bought recently? I can’t seem to find reference to it though.

Dark Arc · 1 year ago

The whole point of opinion pieces is to expose opinions that are outside of the realm of what you’d normally publish. It’s supposed to be a means for keeping your readers out of their echo chamber/exposing different view points.

The times AFAIK didn’t get bought but Jeff Bezos owns The Washington Post, perhaps that’s what you’re thinking of.

Chris · 1 year ago

Yeah that must be it. And thank you that is good perspective. To ask a dumb question, at what point can we decide not to care about other side opinions? Climate change for instance, I don’t need to see another opinion saying that it isn’t bad and the science is wrong, that is pretty much settled.

@assassin_aragorn@lemmy.world · 1 year ago

I’m pretty sure I’ve even seen them use push notifications for corrections.

They have, I distinctly remember them doing that a few times.

@Linkerbaan@lemmy.world · 1 year ago

https://youtu.be/bN9Rh3XOeo8?si=MTmRynqATp5eU4g1&t=344 No the New York Times is a Zionist propaganda outlet that falsifies evidence to push an agenda.

@PipedLinkBot@feddit.rocks · 1 year ago

Here is an alternative Piped link(s):

https://piped.video/bN9Rh3XOeo8?si=MTmRynqATp5eU4g1&t=344

Piped is a privacy-respecting open-source alternative frontend to YouTube.

I’m open-source; check me out at GitHub.

Tony Bark · 1 year ago

And OpenAI hasn’t exactly been open since GPT-3.

@NevermindNoMind@lemmy.world · 1 year ago

One thing that seems dumb about the NYT case that I haven’t seen much talk about is that they argue that ChatGPT is a competitor and it’s use of copyrighted work will take away NYTs business. This is one of the elements they need on their side to counter OpenAIs fiar use defense. But it just strikes me as dumb on its face. You go to the NYT to find out what’s happening right now, in the present. You don’t go to the NYT to find general information about the past or fixed concepts. You use ChatGPT the opposite way, it can tell you about the past (accuracy aside) and it can tell you about general concepts, but it can’t tell you about what’s going on in the present (except by doing a web search, which my understanding is not a part of this lawsuit). I feel pretty confident in saying that there’s not one human on earth that was a regular new York times reader who said “well i don’t need this anymore since now I have ChatGPT”. The use cases just do not overlap at all.

@abhibeckert@lemmy.world · edit-2 1 year ago

it can’t tell you about what’s going on in the present (except by doing a web search, which my understanding is not a part of this lawsuit)

It’s absolutely part of the lawsuit. NYT just isn’t emphasising it because they know OpenAI is perfectly within their rights to do web searches and bringing it up would weaken NYT’s case.

ChatGPT with web search is really good at telling you what’s on right now. It won’t summarise NYT articles, because NYT has blocked it with robots.txt, but it will summarise other news organisations that cover the same facts.

The fundamental issue is news and facts are not protected by copyright… and organisations like the NYT take advantage of that all the time by immediately plagiarising and re-writing/publishing stories broken by thousands of other news organisations. This really is the pot calling the kettle black.

When NYT loses this case, and I think they probably will, there’s a good chance OpenAI will stop checking robots.txt files.