LONDON, February 29, 2024 – A massive data scrape potentially impacting the future of music and artificial intelligence has come to light: an activist group claims to have copied roughly 86 million songs-nearly 300 terabytes of data-from Spotify.
Archive Claims Nearly Complete Spotify Library Copy
Table of Contents
The group, Anna’s Archive, says its goal is to preserve music for future generations, but experts warn the data could be a boon for AI music progress.
- Anna’s Archive alleges it scraped 86 million music files and associated metadata from Spotify.
- Spotify confirmed an investigation into “unlawful scraping” and has taken action against involved accounts.
- The scraped data could be valuable for training AI systems to compose or mimic musical styles.
- The incident reignites debate over copyright law and its request to AI-driven technologies.
The collective, known for previously linking to pirated books and academic papers, detailed its actions in a blog post, stating its intention to build a “preservation archive” for music. They argue that archiving widely listened-to songs would safeguard cultural heritage against threats like war, natural disasters, or funding cuts. The group intends to distribute the collection using peer-to-peer torrent technology.
The potential impact of this data scrape extends beyond preservation, raising significant concerns about its use in artificial intelligence. Large audio datasets are in high demand for training generative AI systems capable of composing or mimicking music styles.
AI Development and Copyright Concerns
Industry observers are already speculating about the value of the scraped data to AI developers. Yoav Zimmerman, co-founder of AI startup Third Chair, noted on LinkedIn that such a dataset could theoretically allow individuals to “create their own personal free version of Spotify” or enable companies to “train on modern music at scale.” “The only thing stopping them is copyright law and the deterrent of enforcement,” Zimmerman wrote.
This incident echoes recent legal battles involving Meta, where chief executive mark Zuckerberg allegedly approved the use of data from LibGen-a repository of pirated books-for AI training, despite internal warnings about copyright infringement. While Meta successfully defended itself against initial claims, the legal dispute remains ongoing.
Spotify confirmed it is indeed investigating the claims, stating the accessed material did not represent its full catalog.A Spotify spokesperson said the company had already shut down accounts involved in what it described as “unlawful scraping.” “An investigation into unauthorised access identified that a third party scraped public metadata and used illicit tactics to circumvent digital rights management to access some audio files,” the spokesperson added. The company stated it did not believe the music had yet been released publicly.
growing Debate Over AI and Copyright
The episode arrives during a period of intensifying global debate over how copyright law should apply to AI. Creative industries argue their work is being harvested without consent, while AI companies maintain that broad access to data is essential for innovation.
A recent UK government survey on AI revealed overwhelming public support for stronger copyright protections, with a clear preference for protecting existing creative rights. Musicians, authors, and visual artists have cautioned that AI systems trained on existing creative work risk devaluing human creativity while disproportionately benefiting technology firms.
Last month, the science, innovation and technology secretary, Liz kendall, expressed sympathy for artists’ concerns regarding their copyrighted work being used by AI companies without compensation, stating she wanted to “reset” the debate.
Spotify has stated it has introduced additional safeguards since the Anna’s Archive proclamation and is actively monitoring for suspicious activity aimed at bypassing copyright protections.
