PART I

Identifying the problem


‘Thinking II’. 2020. Watercolor on paper. 
The independent artist and intellectual are among the few remaining personalities equipped to resist the stereotyping and consequent death of genuinely lively things.
— C. Wright Mills, The powerless people: The social rôle of the intellectual, 1945 [1]
IDENTIFYING THE PROBLEM

Lack of consent & broken trust

Creative work is an inherently vulnerable act. There is no guarantee that it will be paid or even recognized. It can have either incredible value or no value at all. Sharing media was the original intent of the web. Sir Tim Berners-Lee designed the URL to be a destination to anything and that permission to create the world at the other end of a link inspired endless digitization of places, things, and people [2] . The only recompense one asks in return is a link back, a citation, an attribution — and on some platforms, like Pinterest, attribution isn’t necessarily required for the maker of the work to harvest the fruits of the pollination of their work. Exposure and recognition can lead to customers.

Over decades, many artists and creators uploaded their work and shared it online trusting that it would be viewed and used by other people in predictable ways: read, remixed, downloaded, traded. What they didn’t expect was that it would be used without permission to create models that, according to the United States Copyright Office [3] directly compete with them in the marketplace for digital creative goods. There were no individual consent requests for the usage of creative works for training for large language models and image generation models, as they were imagined to be given freely as a communal resource. Now that such a breach of trust has occurred, it can directly threaten willingness to make and share creative works online. It also damages the willingness for companies to pay for creative data. This is both a tragedy of the commons and a free-rider problem, rolled into one — but instead of individuals exploiting the communal resource, it’s large businesses. 

To take this, an entire world of people’s identities and creative effort, loosely based on a culture of collective benefit and goodwill, and to use it to make a technology that puts those people out of business, is a massive breach of trust and a particularly painful theft. Trust, defined by Mayer et. al (1995) [4], is “a willingness to be vulnerable with another party”. If this trust is broken and left unrepaired, who will be willing to build the vibrant, thoughtful, and beautiful internet of tomorrow?

A mismatch of scale and value

In the development of generative machine learning models, unfathomable numbers of digital creative works were used as training material. The value of these individual works to their creators does not match their value to the developers of machine learning models. A single book can create years of revenue and royalties for a single author and the cost to read or use that book is paid for by the people reading it. However,a single book in a dataset of millions has, individually, a value approaching zero. The value of the data hasn't disappeared, it has been absorbed into the aggregate [5], where the creators can no longer access it. This mismatch in scale between value to individuals and value as part of scraped training data is at the core of the moral issues of training machine learning models on publicly available creative works. 

These value extraction mechanisms are essentially exploitation, which is also how use of creative works as data is described by top data and media academics [6] and policy institutions [7]. In the Oxford English Dictionary, ‘exploitation’ is defined as “The action or fact of taking advantage of something or someone in an unfair or unethical manner; utilization of something for one's own ends” [8]. In describing the action of getting value from a free communal resource, those who do so have also defined that action as unfair or unethical. 

The transformation of artistic works into training data strips the works of their attribution and ownership frameworks. However, when creative works are used in this way, the harms to the creators are diffuse yet individually extremely painful — the opposite of the concentrated monetary value a company might get from collecting and leveraging the works for training machine learning models.

Institutions are the rules of the game in a society or, more formally, are the humanly devised constraints that shape human interaction. In consequence they structure incentives in human exchange, whether political, social, or economic.
— Douglass North, Institutions, institutional change and economic performance, Cambridge University Press, 1990 [9]


Real harms with unrealistic remuneration

The internet itself  is an institution with actors, expected interactions, and incentives. For such an institution to continue and grow, the breach of trust must be repaired. A path to trust doesn’t require complete forgiveness, but instead, the establishment of new boundaries and engagement expectations. New structures and rules, that, if followed over time, can result in the regrowth of trust. 

If we think of the internet as a community maintained by its users, then a societal trust repair framework is a good starting point. To rebuild trust between groups of people, a path to repair must pass through something resembling what Archbishop Desmond Tutu described in the aftermath of apartheid: a process of truth, acknowledgment, and restorative justice [10]. However, those who maintain the structure of the internet are not the same individuals or groups that committed the breach of trust. Thus, to use this framework, it must be expanded to include those who benefit from the resource without contributing to it.

If the internet is a collective resource, then an ethics based on Ostrom’s Principles for managing the commons could be applied [11]:

Ostrom’s Principles for Managing the Commons

However, very few of those principles have been respected or upheld in the creation of generative machine learning models from CommonCrawl, unlicensed, and even copyrighted data.

If we think of the internet itself as a business, where employees (data contributors) receive benefits from the business owners (machine learning builders and maintainers) we can start to move away from the concept of online media as a collective resource. This reframing opens up helpful frameworks for how such a relationship can be made just. The 2011 United Nations “Guiding principles on business and human rights”establish a framework of ‘Protect, Respect, and Remedy’ for human rights in business. These pillars lean on the state to encourage business compliance with the principles, but without a formal contract between data creators and machine learning developers, these principles are not recognized or enforceable.

The most useful framework is a blend of these three, imagining the internet more loosely as an institutional organization but not one with a specified purpose. Bachmann et al. (2015) provide a useful conceptual framework for repairing trust in institutions [13]. In their paper, they identify 6 mechanisms for trust repair, and underlying mechanisms. Using this framework, it’s possible to identify these repair mechanisms at work as different members of the institution attempt repair. 

Sensemaking

“A shared understanding or accepted account of the trust violation is required for effective trust repair”

The first question in determining harms to creators and artists is whether or not their works were used for training. In the aftermath of the release of StableDiffusion and ChatGPT, a websites titled ‘Have I Been Trained?’ (2022, now Spawning.ai) [14] and ‘Don’t train on me’ (2024, now https://trufo.ai/) [15] attempted to help find that truth. The website ‘Knowing Machines’ [16]  includes visual stories, a reading list,explainers and collections to guide creatives through the processes that led to the large-scale usage of public works for private gain. However, the technical realities of parsing machine learning data sets make it difficult to understand the impact on individuals. 

Relational

“Trust repair requires social rituals and symbolic acts to resolve negative emotions caused by the violation and re-establish the social order in the relationship’“

The second question is whether or not the outputs of the models are competing against the artists whose works were used for training. According to the US office of copyright and IP in the US [3], the outputs of machine learning models can indeed be considered competitive in the creative market. It is also theoretically possible to determine the precise value that each token in the training data provides for the model, as Yoon finds, using Data Valuation using Reinforcement Learning (DVRL) “Organizations that sell data can use it for pricing of each datum” [17]. Thus, we can confirm that economic harm is being done, and it is theoretically possible to calculate the remuneration value to victims. 

Regulation and controls

Trust repair requires formal rules and controls to constrain untrustworthy behaviour and hence prevent a future trust violation”

Over the past 4 years, there have been many new and ongoing court cases focused on intellectual property and copyright, brought forward by large publishers and extremely powerful and popular artists and writers. The choice for large players is practically between litigation (NyTimes/Microsoft) and licensing (Time, Disney, OpenAI), and in some cases, both. 

However, using creative works for training in many cases may qualify legally as fair use [18]. In addition, for a small creator, because of the low individual value of their contribution to the data set, and the financial power needed to seek remuneration, the possibility for remuneration through licensing and litigation remains low. In fact, these are not often possible or desired paths for small creators, and some, such as Re:Create movement, would prefer to continue to use and support the creation of generative machine learning models using their work [19]. 

Often, it is prohibitively expensive to defend copyright, and current policies cannot force model developers to attribute data sources. Even if every large language model or image generation model did indeed publish a record of every image or text source used to train the machine learning model, and a passionate team of lawyers aimed to file a class action lawsuit against each of the firms that used these creative works as training data, a class action lawsuit could take a decade to finalize and settle. This remuneration, which it is difficult to gauge in size and format, is unrealistic to rely on within the working life of the artists who were exploited. 

Transparency

“Transparently sharing relevant information about organizational decision processes and functioning with stakeholders helps restore trust”

Model cards, which many large firms publish to describe their large language modes, do not cite the precise data that is used from each source: OpenAI [20], Mistral [21], Anthropic [22], Google [23], DeepSeek [24], HuggingFace [25] all drescribe data mix, not specific tokens. Even open source models, like the image evaluation models LAION, which does cite every image in the dataset, would take over 700 years for a single person to review [26]. When the dataset was reviewed, it was found to contain images of abuse of children. Producing and publishing such a citation would also attract continuous lawsuits, which does not incentivize model development labs to do so. 

Ethical culture 

“Trust repair requires informal cultural controls to constrain untrustworthy behaviour and promote trustworthy behaviour, and hence prevent a future trust violation”

The backlash against using creative works as training data without consent or compensation has been extremely robust. As described in ‘The Creative Double Bind” in the Harvard Law Review, using generative models is both a shortcut and a harmful one [27]. Many creative groups and culture self-police against the use of AI, e.g. the comic artist employed for this work advertises his work on Instagram as “Zero AI”[28], as does the Authors Guild on their website: “Human Authored”[29], and the podcast network iHeartPodcasts, “Guaranteed Human”[30]. Some of the larger and louder groups pushing for correction of this injustice include Stealing isn’t Innovation [31], and the AI-related terms after the 2023 SAG-AFTRA Strike [32]. However, restoring a sense of justice in this case remains difficult. 

Transference

“Trust repair can be facilitated by transferring trust from a credible party to the discredited party”

The SAG-AFTRA strike delivered a strong sense of public credibility to Hollywood in general. Hollywood remains a beloved institution globally, and a focal point of creative employment and work. A single film employs a gamut of creative work including but not limited to writers, songwriters, animators, actors, computer graphics artists, directors, and producers. In 2023, the strike pitted these creatives against the use of AI, and they effectively negotiated its use and non-use for their work. Since 2024, large AI companies have aimed to gain the trust of this highly-credible group, through incubators [33], funding [34], partnerships [35], product integrations [36], public-facing principles [37] and research [38], with mixed results [39]. 

From the other direction, the non-profit Fairly Trained [40] aims to provide certifications to companies that ‘don’t use any copyrighted work without a license’. This certification incentivizes companies to either buy licenses for the work, or use unlicensed with.  Pragmatically, unlicensed work is work for which the license cannot be enforced, either because the owner is unaware of its use, or does not have the legal means to sue for misuse. 


[1] Mills, C. W. (1945). The powerless people: The social rôle of the intellectual. Bulletin of the American Association of University Professors, 31(2), 231–243. https://doi.org/10.2307/40221218

[2] Berners-Lee, T., & Witt, S. (2025). This is for everyone: The unfinished story of the World Wide Web. Macmillan.

[3] U.S. Copyright Office. (2025, May 9). Copyright and artificial intelligence, part 3: Generative AI training [Pre-publication version]. https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf

[4] Mayer, R. C., Davis, J. H., & Schoorman, F. D. (1995). An Integrative Model of Organizational Trust. The Academy of Management Review, 20(3), 709–734. https://doi.org/10.2307/258792

[5] Coyle, D., &  Manley, A. (2024).  What is the value of data? A review of empirical methods. Journal of Economic Surveys,  38,  1317–1337. https://doi.org/10.1111/joes.12585

[6] Vera Nieto, D., Celona, L., & Fernandez Labrador, C. (2022). Understanding aesthetics with language: A photo critique dataset for aesthetic assessment. Advances in Neural Information Processing Systems, 35. https://www.research-collection.ethz.ch/entities/publication/c7c7e9db-e144-4cf5-b47e-7a2453b1ff36

[7] Directive 2019/1024 of the European Parliament and of the Council. (2019, June 26). Official Journal of the European Union. https://eur-lex.europa.eu/legal-content/EN/TXT/PDF/?uri=CELEX:32019L1024

[8] “Exploitation, N., Sense 1.c.” Oxford English Dictionary, Oxford UP, December 2025, https://doi.org/10.1093/OED/3660582354.

[9] North, D. C. (1990). Institutions, institutional change and economic performance. Cambridge University Press.

[10] Tutu, D. (1999). No future without forgiveness. Doubleday. https://archive.org/details/nofuturewithoutf0000tutu

[11] Ostrom, E. (1990). Governing the commons: The evolution of institutions for collective action. Cambridge University Press. https://www.actu-environnement.com/media/pdf/ostrom_1990.pdf

[12] United Nations Human Rights Office of the High Commissioner. (2011). Guiding principles on business and human rights: Implementing the United Nations "Protect, Respect and Remedy" framework. United Nations. https://www.ohchr.org/sites/default/files/documents/publications/guidingprinciplesbusinesshr_en.pdf

[13] Bachmann, R., Gillespie, N., & Priem, R. (2015). Repairing trust in organizations and institutions: Toward a conceptual framework. Organization Studies, 36(9), 1123–1142. https://doi.org/10.1177/0170840615599334

[14] Spawning. (2022, September 14). Have I Been Trained? [Web archive]. Wayback Machine. https://web.archive.org/web/20220914211816/https://haveibeentrained.com/

[15] Don't Train on Me. (n.d.). Don't Train on Me. Retrieved July 4, 2026, from https://app.dont-train-on-me.org/

[16] Knowing Machines. (n.d.). Knowing Machines: Tracing the histories, practices, and politics of machine learning systemshttps://knowingmachines.org/

[17] Yoon, J., Arik, S., & Pfister, T. (2020). Data valuation using reinforcement learning. Proceedings of the 37th International Conference on Machine Learning, 119, 10842–10851. https://proceedings.mlr.press/v119/yoon20a.html

[18] Knowing Machines. (2025, November 21). Is GenAI allowed to produce outputs "in the style of" a human artist? Knowing Machines. https://knowingmachines.org/knowing-legal-machines/legal-explainer/questions/is-genai-allowed-to-produce-outputs-in%20the-style-of-a-human%20artist

[19]  Re:Create Coalition. (2025, May 28). Breaking down the USCO report on generative AI training and Re:Create's "non-takeaways." https://www.recreatecoalition.org/breaking-down-the-usco-report-on-generative-ai-training-and-recreates-non-takeaways/

[20] OpenAI (GPT-5.5) OpenAI. (2026, April 23). GPT-5.5 System Card. Retrieved July 4, 2026, from https://openai.com/index/gpt-5-5-system-card/

[21] Mistral AI Mistral AI. (n.d.). [System card or legal documentation]. Retrieved July 4, 2026, from https://legal.cms.mistral.ai/assets/1e37fffd-7ea5-469b-822f-05dcfbb43623

22] Anthropic Anthropic. (n.d.). Model system cards. Retrieved July 4, 2026, from https://www.anthropic.com/system-cards

[23] Google DeepMind Google DeepMind. (n.d.). Model cards. Retrieved July 4, 2026, from https://deepmind.google/models/model-cards/[

24] DeepSeek DeepSeek AI. (2026, April 27). DeepSeek V4 technical documentation. Retrieved July 4, 2026, from https://fe-static.deepseek.com/chat/transparency/deepseek-V4-model-card-EN.pdf

[25] CompVis / Stable Diffusion Rombach, R., & Esser, P. (2022). Stable Diffusion [Model card]. Hugging Face. https://huggingface.co/CompVis/stable-diffusion

[26] Buschek, C., & Thorp, J. (n.d.). Models all the way down. Knowing Machines. Retrieved July 4, 2026, from https://knowingmachines.org/models-all-the-way

[27] Harvard Law Review. (2025). Chapter two: Artificial intelligence and the creative double bind. Harvard Law Review, 138, 1585–1608.

[28] Patterson, G. @carpaintings.art. (n.d.). Home [Instagram profile]. Instagram. Retrieved July 4, 2026, from https://www.instagram.com/carpaintings.art/

[30] iHeartMedia. (2026). Guaranteed human at CES 2026: Real voices drive real outcomes. Retrieved July 4, 2026, from https://www.iheartmedia.com/advertise/insights/articles/guaranteed-human‍ ‍

[31]  Stealing Isn't Innovation. (2026). Stealing isn't innovation: America's creative community speaks out. https://www.stealingisntinnovation.com

[32] SAG-AFTRA. (n.d.). Artificial intelligence resources: 2023 TV/theatrical contracts. Retrieved July 4, 2026, from https://www.sagaftra.org/contracts-industry-resources/contracts/2023-tvtheatrical-contracts/artificial-intelligence-resources

[33] Google. (n.d.). We're introducing Flow Sessions and our first filmmaker in residence. Google Blog. Retrieved July 4, 2026, from https://blog.google/innovation-and-ai/models-and-research/google-labs/flow-resident-filmmaker/

34] Google. (n.d.). Building a community-led future for AI in film with Sundance Institute. Google Blog. Retrieved July 4, 2026, from https://blog.google/company-news/outreach-and-initiatives/google-org/sundance-institute-ai-education/

[35] Universal Music Group. (2025, October 29). Universal Music Group and Udio announce Udio's first strategic agreements for new licensed AI music creation platform. Retrieved July 4, 2026, from https://www.universalmusic.com/universal-music-group-and-udio-announce-udios-first-strategic-agreements-for-new-licensed-ai-music-creation-platform/

[36] Anthropic. (n.d.). Claude for creative work. Retrieved July 4, 2026, from https://www.anthropic.com/news/claude-for-creative-work

[37] YouTube. (n.d.). Our principles for partnering with the music industry on AI technology. YouTube Blog. Retrieved July 4, 2026, from https://blog.youtube/inside-youtube/partnering-with-the-music-industry-on-ai/

[38] Google DeepMind. (2024, October). New generative AI tools open the doors of music creation. Retrieved July 4, 2026, from https://deepmind.google/blog/new-generative-ai-tools-open-the-doors-of-music-creation/

39] Heritage, S. (2026, February 2). Requiem for a film-maker: Darren Aronofsky’s AI revolutionary war series is a horror. The Guardianhttps://www.theguardian.com/film/2026/feb/02/darren-aronofsky-ai-revolutionary-war-series-review

[40] Fairly Trained. (n.d.). Fairly Trained. Retrieved July 4, 2026, from https://www.fairlytrained.org/

Previous
Previous

Introduction and philosophy