35 publishers sue OpenAI, Microsoft for training ChatGPT with their content
7 mins read

35 publishers sue OpenAI, Microsoft for training ChatGPT with their content

Newspaper publishers across the country are suing OpenAI and Microsoft for allegedly scraping their websites to train their flagship artificial intelligence models.

Read more ‘Spraying and praying’ is ruining the Bay Area tech job market

The lawsuit, filed in the U.S. District Court for the Southern District of New York on Wednesday, alleges that by using unlicensed and paywalled content to train versions of ChatGPT, OpenAI and Microsoft have profited from the publishers’ written work. Those actions have caused newsrooms to lose out on advertising and subscription revenues, the complaint alleges, by deterring people from accessing the content directly from the plaintiffs. (Microsoft does not own OpenAI, but it maintains a significant minority stake in the company and owns the computing infrastructure used to power OpenAI’s large language models.) 

Thirty-five newspaper publishing companies, largely independent and locally owned, have signed on as co-plaintiffs. They represent nearly 400 news outlets across 33 states, including three California papers: the Tahoe Daily Tribune in South Lake Tahoe, the Sierra Sun in Truckee and the Needles Desert Star in Needles.

“The Publishers’ journalism was essential to the Defendants’ explosive growth, and unless Defendants are held accountable for stealing, stripping, and misusing the Publishers’ content, the AI boom Defendants orchestrated and benefit from will be a death knell for local journalism — which remains the most trusted news sources in America,” the lawsuit reads.

Drew Pusateri, a spokesman for OpenAI, told SFGATE in a statement Thursday that the company trains its models on “publicly available data” and that its work is “grounded in fair use.”

But Michael Bolden, the dean of UC Berkeley’s School of Journalism, contends that just because the written work may appear on the open internet does not mean the copyright is waived. He highlighted that a news outlets’ path to publication is an “expensive effort,” from compensation for reporters and editors for newsgathering to the actual cost of publication. 

“The notion that this content is freely available and just sort of dropped down from the heavens is not accurate,” Bolden said in an interview on Friday. “There is an intellectual effort that goes into producing work, including journalism, and companies need to be compensated for that.”

SFGATE reached out to the editors of the three California newspapers whose publishers are named plaintiffs in the lawsuit. Laney Griffo, the editor of the Tahoe Daily Tribune and Sierra Sun, declined to comment, while a representative for the papers’ publisher, Ogden Newspapers Inc., did not respond to a request for comment. Microsoft also did respond to a request for comment before the time of publication. 

Matt Platkin, the plaintiffs’ attorney, alleged in a statement to SFGATE on Friday that OpenAI “systematically and willfully stole” copyrighted materials, which in turn harmed local communities. 

“These actions are not only illegal, but are harmful to essential community papers that are already facing economic pressures and challenges,” he said. “Local reporters should not have their work stolen without credit or compensation, and new technology does not come with an exemption from copyright law.”

Read more Costco receipt helps tie Bay Area woman to baby found dead in dumpster

The modern media landscaping is ever-evolving, with popularity in print products at an all-time low, according to a Pew Research Center analysis. Even news websites’ readerships are in decline as social media and alternative media cut into traditional journalism spaces, per a Reuters Institute study. The lawsuit alleges that while OpenAI has become increasingly valuable, plaintiffs have been “robbed” of subscription, advertising and content licensing revenue.

“The Publishers have spent billions of dollars to sustain this work,” the filing reads. “Defendants helped themselves to all of it — without providing a cent of compensation.”

Large language models such as OpenAI’s ChatGPT process an enormous amount of data every second. In order to function at that scale it does and to keep up with the increasing use and demand for accuracy, the tech needs to be fed a lot of information. Companies do this by taking large snapshots of the internet, including links to news articles, and feeding them into the chatbot’s code to fine-tune its responses. Text from this data is broken down into “tokens,” which the models memorize and use to better predict how to respond to questions from people using the tools. 

OpenAI has continued to use larger and larger data sets to train its newer models. According to OpenAI’s transparency reports cited in the lawsuit, ChatGPT-2 was trained on a singular data set containing 45 million links to written works posted on Reddit. But the next version, ChatGPT-3, was trained on multiple of these types of data sets, including Common Crawl, which is referred to in the lawsuit as a “copy of the internet.” The lawsuit alleges that Common Crawl contains hundreds of thousands of tokens consisting of the plaintiffs’ paywalled content.

image

The Bay Area’s best free newsletter.

Stay informed, and entertained.

By signing up, you agree to our Terms Of Use and acknowledge that your information will be used as described in our Privacy Policy.

Several other larger news operations are litigating similar claims against OpenAI and Microsoft, including the New York Times and the Intercept. For smaller newsrooms with limited resources, the fight for compensation against these AI companies has a “David and Goliath” feel, Bolden said, especially as OpenAI is on track to file what will be one of the most valuable initial public offerings in U.S. history. 

Bolden told SFGATE he was hopeful that multiple organizations were banding together in this effort. Beyond publishers being just compensated for the work they produce, Bolden said copyright laws need to be strengthened to protect local journalism and reporters as AI becomes more powerful. He added that AI developers should also have ongoing discussions with publishers about how to create an equitable information environment.

“These models are very ever more powerful, and they’re going to continue to be developed, and we really need to ensure that there are standards being set so that work that is generated and owned by someone is not taken without permission and reused in a way that benefits another company entirely,” Bolden said.

More News

— Youth pastor accused of pushing wife off Zion cliff found dead
— Costco receipt helps tie Bay Area woman to baby’s death
— New homeowners find human remains after buying Calif. property
— Remote workers tried to flock to a Calif. beach. The city shut it down.

Read more Tourist killed by crocodile at Puerto Vallarta beach

Sign up for daily SFGATE breaking news alerts here.

Leave a Reply

Your email address will not be published. Required fields are marked *