AI training sparks copyright controversy as OpenAI and Microsoft partially win lawsuit

The United States Court of Appeals for the Ninth Circuit in San Francisco upheld a California judge’s decision on Wednesday that OpenAI and Microsoft should not be held responsible for infringing developers’ rights under the U.S. copyright law.

In 2022, two anonymous plaintiffs representing a group of open-source code developers filed a lawsuit on GitHub, bringing Microsoft and OpenAI to court. The plaintiffs alleged that these companies used code from the GitHub repository to train OpenAI’s Codex and Microsoft’s Copilot AI programming tools, a practice that violated the U.S. Digital Millennium Copyright Act (DMCA) and open-source license terms.

The DMCA prohibits removing copyright management information (CMI) from protected works. CMI includes information that identifies the work, such as the author and copyright owner.

Developers claimed that Codex and Copilot failed to attribute the original authors when copying code, thus violating relevant provisions of the DMCA.

This lawsuit is one of the earliest among a wave of lawsuits brought by copyright holders (including authors, news organizations, and record companies) against tech companies for using protected materials in AI training. Unlike most cases focusing on software developers, the core issue in these lawsuits tends to be “copyright infringement.”

Federal District Judge Jon Tigar had previously dismissed the developers’ DMCA-based allegations but allowed them to proceed with the lawsuit against OpenAI and Microsoft for violating open-source license agreements.

The Ninth Circuit Court of Appeals on Wednesday decided to uphold the district judge’s dismissal of the DMCA claims. The court determined that Copilot and Codex did not remove copyright information from existing works but rather generated new works that did not contain such information.

Copyright claims related to the output of AI systems often rely on traditional principles of infringement. The Appeals Court on Wednesday made a distinction between DMCA allegations and general infringement claims—the latter requiring evidence of “substantial similarity” between the output of AI and the copyrighted works in question.

The Court of Appeals stated that if infringement under the DMCA could be established solely based on “substantial similarity,” the law would “substitute traditional copyright protections and expose defendants to the devastating liability of enhanced damages under the DMCA.”

The ruling by the Appeals Court did not determine whether the similarity between AI-generated code and copyrighted code is sufficient to constitute copyright infringement.