What happened
Sony Music and Universal Music Group (UMG) have initiated a new legal action against AI music generator Suno, specifically targeting its v6 model. The labels contend that the model remains fundamentally infringing because it incorporates data generated by earlier versions of Suno’s software, which the plaintiffs allege were trained on copyrighted music scraped from platforms like YouTube without authorization.
According to the report by The Verge, Sony and UMG filed the complaint alleging that Suno’s v6 model is the 'fruit of the same poisoned tree' as its predecessors. The labels argue that Suno is engaging in 'model laundering,' where the value of copyrighted expression is passed from original recordings into initial models, then into the outputs of those models, and finally into the training set for v6.
The plaintiffs specifically allege that Suno utilized a process of distillation to train v6, effectively teaching the new model to replicate the results of previous 'teacher' models that were allegedly built on unlicensed data. The complaint asserts that even if v6 was not directly trained on the plaintiffs' recordings, it remains 'informed by, and benefits from' the unauthorized copies retained by Suno.
Suno’s Jack Brody previously stated to The Verge that v6 was 'trained from the ground up, with a new set of data,' which included user-generated content. However, the company has not provided specific details regarding the composition of this dataset or the extent to which previous model outputs were utilized in the training process.
Sony and UMG remain notable holdouts in the AI music space, having refused to sign licensing agreements with Suno. This lawsuit represents a significant escalation in their efforts to hold AI developers accountable for the provenance of their training data.
Source details: theverge.com ↗
Why it matters
This lawsuit introduces the concept of 'model laundering' into the legal discourse surrounding . By arguing that training a new model on the outputs of an infringing predecessor perpetuates copyright violation, the labels are challenging the industry practice of iterative model development. This case could establish a critical legal precedent regarding whether AI companies can 'cleanse' their training data through successive generations of models, potentially forcing a shift in how AI developers curate and document their training sets to avoid liability.
The core of this dispute centers on the legal definition of 'fresh' training data in the context of . If the court accepts the labels' argument, it could invalidate the common industry practice of using or model outputs to refine subsequent versions of AI systems.
This case highlights the ongoing tension between AI developers and the music industry regarding intellectual property rights. For creators and rights holders, the outcome will determine whether they have a viable path to claim infringement when their work is used to train AI models that are several generations removed from the original source material.
The practical implication for the AI industry is a potential requirement for rigorous '' documentation. If companies cannot prove that their models are free from the influence of infringing data, they may face significant legal and financial risks, regardless of how many iterations they perform.
Interactive Mechanism: How It Actually Works
Explore the underlying technology behind this development interactively.
crm_get_transaction(id='4092').Impossibility results in algorithmic fairness (e.g. Kleinberg et al., Chouldechova) show what?
What to watch next
The court's interpretation of 'model laundering' and whether it accepts the labels' argument that distillation and iterative training constitute a continuation of copyright infringement. Observers should monitor whether Suno provides transparency regarding its training data for v6, as the company has thus far declined to elaborate on the specific nature of the 'user data' it claims to have used.
The legal discovery process will be critical, as it may force Suno to disclose the specific datasets and methodologies used to train v6. The extent to which the court allows the labels to probe the 'black box' of Suno's training pipeline will be a key indicator of the case's trajectory.
Whether other major music labels or publishers join the suit or file similar actions against other AI music generators. The industry is currently watching to see if this 'model laundering' theory gains traction in other copyright litigation involving .
The availability of specific technical evidence regarding the distillation process mentioned in the complaint. Currently, the technical details of how v6 was trained remain an unknown, as Suno has not publicly verified the labels' claims.