Skip to main content

Clarifications and Corrections

1 month 2 weeks ago
An article about the trial of a new Ebola vaccine inaccurately said that the trial would inject people with Ebola. In fact, the vaccine candidate does not contain Ebola virus and no volunteers...

Book Publishers Sue Google For Copyright Infringement Over Gemini AI Training

1 month 2 weeks ago
Major publishers Hachette, Cengage, Elsevier, and author Scott Turow have sued Google, accusing it of using millions of copyrighted books to train Gemini without permission or payment, in "one of the most prolific infringements of copyrighted materials in history." The Guardian reports: The publishers argue that Google repurposed books that had been supplied for limited services such as Google Books, Google Play Books and Google Scholar. Those services allowed Google to use the works in specific ways -- for example, to display searchable snippets or sell ebooks -- but not, the lawsuit claims, to copy them for training commercial AI products. "Desperate to maintain its online dominance, Google abandoned its early motto of 'Don't be evil' and engaged in one of the most prolific infringements of copyrighted materials in history," the suit states (PDF). According to the complaint, the tech company made copies of copyrighted books to train Gemini without permission or payment, despite internal discussions acknowledging the legal risks. The filing claims Google flagged internally that it could face "$10Bs-$100Bs in potential fines" for using texts provided by publishers for Google Play Books. The publishers say Google's actions are harming authors and the wider publishing industry, arguing that AI-generated content could negatively impact book sales. It notes that, for example, Gemini could generate "a 100-page murder mystery set in a quiet seaside town filled with secrets, that substitutes for an original copyrighted murder mystery on which Gemini trained" in 20 minutes for 39 cents. "No publisher or author can compete with that." The lawsuit names a number of specific books that the publishers allege were among the copyrighted works used without permission, including NK Jemisin's The Fifth Season, and Lemony Snicket's Who Could That Be at This Hour?

Read more of this story at Slashdot.

BeauHD

Spotify Is Now an AI Chatbot, Too

1 month 2 weeks ago
Spotify is testing a new "Talk to Spotify" AI feature for Premium subscribers that will let them chat with an AI assistant to explore music, podcasts, and audiobooks. The feature can answer questions about what users are listening to, adjust playback through follow-up prompts, and offer more personalized recommendations. The Verge reports: Amazon Music introduced a similar feature last year when it integrated Alexa Plus into the service. Spotify's chatbot goes a step beyond providing AI-powered recommendations and general trivia, however, because it references your playlists, favorite artists, repeat listens, and listening data when responding to requests. That means you can ask questions about your own listening history to check when you first heard a specific song, or see what genres you've been into lately if you can't hold out for the annual Wrapped insights. The updated AI capabilities are more conversational than older features like Prompted Playlist, which automatically builds playlists based on descriptions. Now, you can ask the Spotify chatbot to "play some songs I haven't heard before," and control what's being played with further instructions like requesting specific artists or asking to make it "more upbeat." Spotify says the new conversational experience aims to make the platform "more personal and useful for every listener," making this one of several ways that the company is trying to address complaints about its algorithm. You can also ask the Spotify AI general questions about whatever you're listening to, making the feature feel similar to using chatbot services like Google's Gemini or OpenAI's ChatGPT. That includes asking for when a song was released, exploring other titles an author has written when listening to one of their audiobooks, or checking if a podcast guest has appeared on other audio shows.

Read more of this story at Slashdot.

BeauHD

Hack Reveals Suno AI Music Generator Scraped YouTube, Deezer, and Genius

1 month 2 weeks ago
A hacker who breached Suno reportedly revealed source code and training-library details showing the AI music generator scraped millions of songs and lyrics from sources including YouTube Music, Deezer, Genius, Pond5, Jamendo, Freesound, and podcast RSS feeds. "The hacked data is a rare look at exactly how AI models and tools are built," reports 404 Media. "Suno is one of the largest AI music generation tools on the internet, and has been the subject of several major lawsuits from the record industry, which accused the company of training on millions of copyrighted songs." Suno maintains that its models were trained on publicly available music files and metadata as fair use. 404 Media reports: The Recording Industry Association of America accused Suno of ripping songs directly from YouTube; the hacked data seen by 404 Media confirms this. The hacked material includes source code that appears to be from 2023 and 2024 that includes scraping instructions and details about the scope of at least some of the scraping. For example, the comments in one file note that they will pull from "genius_hq, youtube_music, freesound, jamendo, imp, deezer, ytm_tagged," and that "non-music will be filtered out." A file called "youtube_music" notes that at the time the file was last updated, it had ingested "2,013,545 music clips." Another file contains comments about different datasets Suno had created, which included "113,879 hours of youtube_music," "17,615 hours of genius_hq," "410 hours of free sound," "19,514 hours of imslp," "3,726 hours of jamendo," "62,117 hours of pond5_music," "12,287 hours of deezer," "152,162 hours of ytm_tagged," and "103 hours of musescore_lyrics." In total, this is at least decades worth of music. Other code the hacker shared with 404 Media appeared to look specifically for vocals by searching specifically for acapella versions of songs on YouTube. The code also suggested that Suno was using proxies to scrape songs from YouTube through a company called Bright Data, which sells scraping tools, infrastructure, and data services. Additional code shows that with the help of an online tool called PodcastIndex, Suno identified 420,000 different podcasts that had at least five, 30-minute episodes and sought to download roughly 1 million hours of podcasts. [...] The hacker, ellie.191, told 404 Media they breached the company by hacking an individual employee using the Shai-Hulud worm, a supply chain attack that allowed hackers to harvest GitHub and cloud service credentials. They said they also accessed Suno's customer list, which included customers' emails and/or phone numbers and Stripe payment details, depending on what they used to login. The hacker provided a sample of some of the customers, some of whom confirmed to 404 Media they had used their phone number to sign up for Suno and said they were never notified of a breach. The hacker told 404 Media they had no specific motivation for hacking Suno and said "I like to hack anything and everything."

Read more of this story at Slashdot.

BeauHD

FCC Plans To Repeal 39% TV Ownership Cap

1 month 2 weeks ago
The FCC plans to vote on repealing local TV ownership limits, including the 39% national audience cap that currently restricts how much of the U.S. market a single broadcast group can reach. Engadget reports: On August 6, commissioners will hold a ballot to repeal Section 303 of the Communications Act, and with it the 39 percent rule. In essence, the rule limits the reach of a local TV network to no more than 39 percent of the U.S.' total audience market. In its place, the FCC would move to a system whereby it would personally approve or reject TV ownership deals on a case-by-case basis. It's not clear if the FCC even has the authority to reject Section 303 without the explicit consent of the legislature. As Lawrence J. Spiwak wrote in the Yale Journal on Regulation back in January, Section 10 of the Communications Act expressly forbids the FCC from bending the rules around Section 303. "Americans no longer trust the legacy national media to report the news fairly or accurately," wrote FCC Chairman Brendan Carr in an op-ed published on Breitbart. "In fact, only eight percent of Americans have a great deal of trust in mass media. That figure is even lower among Republicans -- sitting at a mere three percent." "... Many local broadcast TV stations are getting hollowed out as a result and turning into little more than mouthpieces for programming produced in New York and Hollywood," he alleged. "That is not what Congress or the FCC intended."

Read more of this story at Slashdot.

BeauHD