“The independent artist and intellectual are among the few remaining personalities equipped to resist the stereotyping and consequent death of genuinely lively things.”
IDENTIFYING THE PROBLEMLack of consent & broken trust
Creative work is an inherently vulnerable act. There is no guarantee that it will be paid or even recognized. It can have either incredible value or no value at all. Sharing media was the original intent of the web. Sir Tim Berners-Lee designed the URL to be a destination to anything and that permission to create the world at the other end of a link inspired endless digitization of places, things, and people [2] . The only recompense one asks in return is a link back, a citation, an attribution — and on some platforms, like Pinterest, attribution isn’t necessarily required for the maker of the work to harvest the fruits of the pollination of their work. Exposure and recognition can lead to customers.
Over decades, many artists and creators uploaded their work and shared it online trusting that it would be viewed and used by other people in predictable ways: read, remixed, downloaded, traded. What they didn’t expect was that it would be used without permission to create models that, according to the United States Copyright Office [3] directly compete with them in the marketplace for digital creative goods. There were no individual consent requests for the usage of creative works for training for large language models and image generation models, as they were imagined to be given freely as a communal resource. Now that such a breach of trust has occurred, it can directly threaten willingness to make and share creative works online. It also damages the willingness for companies to pay for creative data. This is both a tragedy of the commons and a free-rider problem, rolled into one — but instead of individuals exploiting the communal resource, it’s large businesses.
To take this, an entire world of people’s identities and creative effort, loosely based on a culture of collective benefit and goodwill, and to use it to make a technology that puts those people out of business, is a massive breach of trust and a particularly painful theft. Trust, defined by Mayer et. al (1995) [4], is “a willingness to be vulnerable with another party”. If this trust is broken and left unrepaired, who will be willing to build the vibrant, thoughtful, and beautiful internet of tomorrow?
A mismatch of scale and value
In the development of generative machine learning models, unfathomable numbers of digital creative works were used as training material. The value of these individual works to their creators does not match their value to the developers of machine learning models. A single book can create years of revenue and royalties for a single author and the cost to read or use that book is paid for by the people reading it. However,a single book in a dataset of millions has, individually, a value approaching zero. The value of the data hasn't disappeared, it has been absorbed into the aggregate [5], where the creators can no longer access it. This mismatch in scale between value to individuals and value as part of scraped training data is at the core of the moral issues of training machine learning models on publicly available creative works.
These value extraction mechanisms are essentially exploitation, which is also how use of creative works as data is described by top data and media academics [6] and policy institutions [7]. In the Oxford English Dictionary, ‘exploitation’ is defined as “The action or fact of taking advantage of something or someone in an unfair or unethical manner; utilization of something for one's own ends” [8]. In describing the action of getting value from a free communal resource, those who do so have also defined that action as unfair or unethical.
The transformation of artistic works into training data strips the works of their attribution and ownership frameworks. However, when creative works are used in this way, the harms to the creators are diffuse yet individually extremely painful — the opposite of the concentrated monetary value a company might get from collecting and leveraging the works for training machine learning models.
“Institutions are the rules of the game in a society or, more formally, are the humanly devised constraints that shape human interaction. In consequence they structure incentives in human exchange, whether political, social, or economic.”
Real harms with unrealistic remuneration
The internet itself is an institution with actors, expected interactions, and incentives. For such an institution to continue and grow, the breach of trust must be repaired. A path to trust doesn’t require complete forgiveness, but instead, the establishment of new boundaries and engagement expectations. New structures and rules, that, if followed over time, can result in the regrowth of trust.
If we think of the internet as a community maintained by its users, then a societal trust repair framework is a good starting point. To rebuild trust between groups of people, a path to repair must pass through something resembling what Archbishop Desmond Tutu described in the aftermath of apartheid: a process of truth, acknowledgment, and restorative justice [10]. However, those who maintain the structure of the internet are not the same individuals or groups that committed the breach of trust. Thus, to use this framework, it must be expanded to include those who benefit from the resource without contributing to it.
If the internet is a collective resource, then an ethics based on Ostrom’s Principles for managing the commons could be applied [11]:
Ostrom’s Principles for Managing the Commons
-
No such legal licensing regime exists.
-
It is possible to track bot traffic and thereby scraping actions, but not how many people made copies of work using generative ML.
-
The legal system as a lever for individual remuneration is inaccessible for many artists.
-
This is an individual calculation.
-
This pricinple is upheld in efforts to identify whther work has been used for training or not. However, users cannot monitor if other people are leveraging their work through the end generative model.
-
Legal settlements and fines are still ongoing, however without proof of abuse, and agreement that abuse took place, this is impossible.
-
In limited cases data owners and creators are consulted in the design and deployment of the machine learning models, but rarely in the initial develpment.
-
In the EU, policy frameworks from the European comission directly address data use, ethics, and compensation.
However, very few of those principles have been respected or upheld in the creation of generative machine learning models from CommonCrawl, unlicensed, and even copyrighted data.
If we think of the internet itself as a business, where employees (data contributors) receive benefits from the business owners (machine learning builders and maintainers) we can start to move away from the concept of online media as a collective resource. This reframing opens up helpful frameworks for how such a relationship can be made just. The 2011 United Nations “Guiding principles on business and human rights”establish a framework of ‘Protect, Respect, and Remedy’ for human rights in business. These pillars lean on the state to encourage business compliance with the principles, but without a formal contract between data creators and machine learning developers, these principles are not recognized or enforceable.
The most useful framework is a blend of these three, imagining the internet more loosely as an institutional organization but not one with a specified purpose. Bachmann et al. (2015) provide a useful conceptual framework for repairing trust in institutions [13]. In their paper, they identify 6 mechanisms for trust repair, and underlying mechanisms. Using this framework, it’s possible to identify these repair mechanisms at work as different members of the institution attempt repair.
Sensemaking
“A shared understanding or accepted account of the trust violation is required for effective trust repair”
The first question in determining harms to creators and artists is whether or not their works were used for training. In the aftermath of the release of StableDiffusion and ChatGPT, a websites titled ‘Have I Been Trained?’ (2022, now Spawning.ai) [14] and ‘Don’t train on me’ (2024, now https://trufo.ai/) [15] attempted to help find that truth. The website ‘Knowing Machines’ [16] includes visual stories, a reading list,explainers and collections to guide creatives through the processes that led to the large-scale usage of public works for private gain. However, the technical realities of parsing machine learning data sets make it difficult to understand the impact on individuals.
Relational
“Trust repair requires social rituals and symbolic acts to resolve negative emotions caused by the violation and re-establish the social order in the relationship’“
The second question is whether or not the outputs of the models are competing against the artists whose works were used for training. According to the US office of copyright and IP in the US [3], the outputs of machine learning models can indeed be considered competitive in the creative market. It is also theoretically possible to determine the precise value that each token in the training data provides for the model, as Yoon finds, using Data Valuation using Reinforcement Learning (DVRL) “Organizations that sell data can use it for pricing of each datum” [17]. Thus, we can confirm that economic harm is being done, and it is theoretically possible to calculate the remuneration value to victims.
Regulation and controls
“Trust repair requires formal rules and controls to constrain untrustworthy behaviour and hence prevent a future trust violation”
Over the past 4 years, there have been many new and ongoing court cases focused on intellectual property and copyright, brought forward by large publishers and extremely powerful and popular artists and writers. The choice for large players is practically between litigation (NyTimes/Microsoft) and licensing (Time, Disney, OpenAI), and in some cases, both.
However, using creative works for training in many cases may qualify legally as fair use [18]. In addition, for a small creator, because of the low individual value of their contribution to the data set, and the financial power needed to seek remuneration, the possibility for remuneration through licensing and litigation remains low. In fact, these are not often possible or desired paths for small creators, and some, such as Re:Create movement, would prefer to continue to use and support the creation of generative machine learning models using their work [19].
Often, it is prohibitively expensive to defend copyright, and current policies cannot force model developers to attribute data sources. Even if every large language model or image generation model did indeed publish a record of every image or text source used to train the machine learning model, and a passionate team of lawyers aimed to file a class action lawsuit against each of the firms that used these creative works as training data, a class action lawsuit could take a decade to finalize and settle. This remuneration, which it is difficult to gauge in size and format, is unrealistic to rely on within the working life of the artists who were exploited.
Transparency
“Transparently sharing relevant information about organizational decision processes and functioning with stakeholders helps restore trust”
Model cards, which many large firms publish to describe their large language modes, do not cite the precise data that is used from each source: OpenAI [20], Mistral [21], Anthropic [22], Google [23], DeepSeek [24], HuggingFace [25] all drescribe data mix, not specific tokens. Even open source models, like the image evaluation models LAION, which does cite every image in the dataset, would take over 700 years for a single person to review [26]. When the dataset was reviewed, it was found to contain images of abuse of children. Producing and publishing such a citation would also attract continuous lawsuits, which does not incentivize model development labs to do so.
Ethical culture
“Trust repair requires informal cultural controls to constrain untrustworthy behaviour and promote trustworthy behaviour, and hence prevent a future trust violation”
The backlash against using creative works as training data without consent or compensation has been extremely robust. As described in ‘The Creative Double Bind” in the Harvard Law Review, using generative models is both a shortcut and a harmful one [27]. Many creative groups and culture self-police against the use of AI, e.g. the comic artist employed for this work advertises his work on Instagram as “Zero AI”[28], as does the Authors Guild on their website: “Human Authored”[29], and the podcast network iHeartPodcasts, “Guaranteed Human”[30]. Some of the larger and louder groups pushing for correction of this injustice include Stealing isn’t Innovation [31], and the AI-related terms after the 2023 SAG-AFTRA Strike [32]. However, restoring a sense of justice in this case remains difficult.
Transference
“Trust repair can be facilitated by transferring trust from a credible party to the discredited party”
The SAG-AFTRA strike delivered a strong sense of public credibility to Hollywood in general. Hollywood remains a beloved institution globally, and a focal point of creative employment and work. A single film employs a gamut of creative work including but not limited to writers, songwriters, animators, actors, computer graphics artists, directors, and producers. In 2023, the strike pitted these creatives against the use of AI, and they effectively negotiated its use and non-use for their work. Since 2024, large AI companies have aimed to gain the trust of this highly-credible group, through incubators [33], funding [34], partnerships [35], product integrations [36], public-facing principles [37] and research [38], with mixed results [39].
From the other direction, the non-profit Fairly Trained [40] aims to provide certifications to companies that ‘don’t use any copyrighted work without a license’. This certification incentivizes companies to either buy licenses for the work, or use unlicensed with. Pragmatically, unlicensed work is work for which the license cannot be enforced, either because the owner is unaware of its use, or does not have the legal means to sue for misuse.

