Between Copyright and Computer Science: The Law and Ethics of Generative AI

69 Pages Posted: 7 Mar 2024

See all articles by Deven R. Desai

Deven R. Desai

Georgia Institute of Technology - Scheller College of Business

Mark Riedl

Georgia Institute of Technology - College of Computing

Date Written: February 16, 2024

Abstract

Copyright and computer science continue to intersect and clash, but they can coexist. The advent of new technologies such as digitization of visual and aural creations, sharing technologies, search engines, social media offerings, and more challenge copyright-based industries and reopen questions about the reach of copyright law. Breakthroughs in artificial intelligence research, especially Large Language Models that leverage copyrighted material as part of training models, are the latest examples of the ongoing tension between copyright and computer science. The exuberance, rush-to-market, and edge problem cases created by a few misguided companies now raises challenges to core legal doctrines and may shift Open Internet practices for the worse. That result does not have to be, and should not be, the outcome.

This Article shows that, contrary to some scholars’ views, fair use law does not bless all ways that someone can gain access to copyrighted material even when the purpose is fair use. Nonetheless, the scientific need for more data to advance AI research means access to large book corpora and the Open Internet is vital for the future of that research. The copyright industry claims, however, that almost all uses of copyrighted material must be compensated, even for non-expressive uses. The Article’s solution accepts that both sides need to change. It is one that forces the computer science world to discipline its behaviors and, in some cases, pay for copyrighted material. It also requires the copyright industry to abandon its belief that all uses must be compensated or restricted to uses sanctioned by the copyright industry. As part of this re-balancing, the Article addresses a problem that has grown out of this clash and under theorized.

The exuberance, rush-to-market, and edge problem cases created by a few misguided companies now raises challenges to core legal doctrines and may shift Open Internet practices for the worse. Legal doctrine and scholarship have not solved what happens if a company ignores Website code signals such as “robots.txt” and “do not train.” In addition, companies such as the New York Times now use terms of service that assert you cannot use their copyrighted material to train software. Drawing the doctrine of fair access as part of fair use that indicates researchers may have to pay for access to books, we show that same logic indicates such signals and terms should not be held against fair uses of copyrighted material on the Open Internet.

In short, this Article rebalances the equilibrium between copyright and computer science for the age of AI.

Keywords: copyright, artificial intelligence, computer science, generative AI, Google Books, machine learning, transformers, ethics, licensing, gridlock economy

JEL Classification: C6, K00, 03, 033, 032, 038, Z1, Z10

Suggested Citation

Desai, Deven R. and Riedl, Mark, Between Copyright and Computer Science: The Law and Ethics of Generative AI (February 16, 2024). Georgia Tech Scheller College of Business Research Paper No. 4735776, Available at SSRN: https://ssrn.com/abstract=4735776 or http://dx.doi.org/10.2139/ssrn.4735776

Deven R. Desai (Contact Author)

Georgia Institute of Technology - Scheller College of Business ( email )

800 West Peachtree St.
Atlanta, GA 30308
United States

HOME PAGE: http://scheller.gatech.edu/directory/faculty/desai/index.html

Mark Riedl

Georgia Institute of Technology - College of Computing ( email )

Do you have a job opening that you would like to promote on SSRN?

Paper statistics

Downloads
570
Abstract Views
2,058
Rank
97,280
PlumX Metrics