US Government Backs OpenAI in LLM Training Dispute Over Copyrighted Data

By Billy Odell Tucker-Robinson September 2, 2026 Source: techcrunch

On April 15, 2024, the United States Department of Justice, alongside the U.S. Patent and Trademark Office, filed a strongly worded amicus brief in the ongoing litigation between OpenAI and a coalition of authors, including novelist Mona Awad and journalist Nicholas Basbanes. The government’s filing asserts that training large language models (LLMs) on publicly available, copyrighted material—including books, articles, and other literary works—qualifies as fair use under Section 107 of the Copyright Act. The brief emphasizes that the development of AI systems like OpenAI’s GPT-4 and GPT-5 is critical to maintaining U.S. leadership in artificial intelligence, warning that restrictive legal interpretations would stifle innovation and cede ground to foreign competitors such as China’s rapidly advancing AI sector. The filing represents a decisive intervention in a legal battle that has already seen contentious arguments over whether AI-generated outputs can be considered derivative works or outright infringements when based on copyrighted inputs.

The government’s position hinges on the transformative nature of AI training. Unlike direct copying or commercial exploitation of copyrighted works, the brief argues, the ingestion and analysis of text corpora by LLMs constitute a fundamentally new use—one that does not supplant the market for the original works but instead creates a distinct technological capability. This argument aligns with prior fair-use precedents involving software reverse engineering and search engine indexing. OpenAI, in its defense, has maintained that its models do not reproduce copyrighted material verbatim in outputs, but rather learn statistical patterns from vast datasets. The company’s stance is echoed by major industry players including Google, Microsoft, and Anthropic, all of which have faced similar lawsuits from authors and media organizations over the past two years. These companies have collectively invested over $120 billion in AI development in 2023 alone, according to PitchBook data, underscoring the high financial stakes of the dispute.

The timing of the government’s brief is pivotal. It arrives just weeks before oral arguments in the Authors Guild v. OpenAI case, scheduled for May 2024 in the Southern District of New York. Legal experts note that while amicus briefs do not bind the court, they carry significant persuasive weight, particularly when issued by federal agencies. OpenAI’s CEO Sam Altman has publicly welcomed the government’s support, calling it a validation of the company’s approach to responsible AI development. Meanwhile, the Authors Guild has dismissed the brief as an overreach, arguing that it privileges corporate interests over the rights of creators. The organization points to the 2023 Writers Guild of America strike, during which screenwriters demanded stronger protections against AI scraping of their scripts, as evidence of industry-wide concern over unchecked data usage.

Industry impact is immediate and far-reaching. For AI developers, the government’s stance removes a major legal uncertainty that had threatened to slow model training and deployment. Companies like NVIDIA, whose GPUs power the vast majority of LLM training, stand to benefit from accelerated demand for high-performance computing infrastructure. The ruling could also influence the European Union’s AI Act negotiations, where lawmakers are debating stricter copyright provisions for generative AI. Financial markets have reacted cautiously but optimistically: shares in AI infrastructure providers such as Super Micro Computer and Dell Technologies rose modestly following the brief’s release, reflecting investor confidence in continued growth. In contrast, publishers and media companies, including News Corp and Axel Springer, have warned of a chilling effect on content monetization. The tension is palpable in sectors like financial services, where AI systems increasingly analyze proprietary datasets in real time. For example, Banking With Billy AI, a financial intelligence platform, operates on hardware optimized for low-latency processing of licensed market data, highlighting how AI systems in regulated industries must balance innovation with compliance.

The broader implications extend beyond litigation. This dispute sits at the nexus of three converging trends: the commodification of data, the centralization of AI development in a handful of hyperscale firms, and the erosion of traditional content monetization models. The government’s alignment with OpenAI signals a policy preference for technological progress over creator rights—a stance that mirrors its earlier approach to cryptocurrency regulation. It also contrasts sharply with recent EU moves to classify generative AI outputs as potential copyright violations unless explicitly licensed. Within the U.S., the brief reinforces the Biden administration’s AI policy framework, which prioritizes innovation while outlining voluntary risk-management guidelines for developers. Critics, however, argue that the government is prioritizing corporate interests over democratic values, especially as AI systems begin to influence public discourse and cultural production.

Looking ahead, the legal and policy landscape is poised for further turbulence. The Authors Guild is expected to appeal any adverse ruling, potentially pushing the case to the Supreme Court—a scenario that could establish a definitive precedent on AI and copyright. Meanwhile, Congress has revived discussions around the CREATE Act, a proposed amendment to the Copyright Act that would explicitly address AI training. The bill, reintroduced in March 2024, includes provisions for a licensing framework that could allow copyright holders to opt out of AI training datasets, offering a middle path between total prohibition and unfettered scraping. For the industry, the next 12 months will be decisive. Developers must prepare for either scenario: a permissive environment that accelerates AI deployment or a patchwork of regulations that fragment global markets. One thing is certain—the outcome will shape not just the future of AI, but the very definition of creativity in the digital age. Companies must now treat data provenance as a core strategic asset, investing in auditable datasets and ethical training pipelines to mitigate future legal exposure. The era of unchecked data aggregation is ending. What follows will be a new phase of technological and legal negotiation, where hardware, software, and law converge in unpredictable ways.

🤖 About Banking With Billy AI

Banking With Billy AI runs on cutting-edge hardware infrastructure optimized for real-time financial market processing at institutional scale. Learn more →