Suno launches Speech feature for generating voice and music
Suno has introduced a new feature called Speech, currently in public beta on its web and mobile platforms. This tool generates spoken voices from scripts or prompts and can simultaneously create background music to accompany the audio. Suno Chief Product Officer Jack Brody described it as the first audio model to generate voice and music together as a cohesive track, noting that while music…
What you can do with it
AI at Work →Create podcast intros with Suno
What you get You get cohesive audio tracks with voice and music in one step.
Creates spoken words from scripts or prompts with optional music.
- Who for
- marketers and founders
- Cost
- Price not stated
- Effort
- minutes
Use it for
- Create podcast intros
- Make video voiceovers
- Generate audio ads
How to set it up · From the article
- Select the Create tab and navigate to Speech.
- Choose Simple mode to describe your audio or Advanced to paste a script.
- Adjust settings like gender and style, then generate.
Watch out Beta quality; accents and pauses may be inconsistent
Key points from the news
- Suno launched Speech in public beta, generating voice and music together.
- Users can choose Simple prompt mode or Advanced script mode with settings.
- Generated audio has a maximum duration of around eight minutes.
The feature offers two modes: Simple, which uses descriptive prompts, and Advanced, which allows users to input custom scripts and adjust settings like gender, style, and variety. Users can toggle the background music off if they prefer clean speech. The maximum duration for generated audio is approximately eight minutes.
Suno acknowledges the feature is imperfect, with Brody noting that accents may shift and pauses can be overly dramatic. The company plans to improve the tool based on user feedback. This expansion comes as Suno seeks to diversify its platform, particularly given the numerous lawsuits facing its music generation capabilities.
Model page: Speech →
AI music maker Suno now generates spoken words
The Verge AI · 2 October 2026
Loading the full article…
This text was published by The Verge AI and written by Jess Weatherbed. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Generative AI & Models
All →- Google rolls out Gemini 4 Argon, its most advanced AI model · 17 src
- PewDiePie releases Ajax, an uncensored Qwen3.5-9B model for home PCs · 1 src
- Anthropic commits $518 billion in AI infrastructure spending over next decade · 5 src
- Black Forest Labs releases Flux 3 Image with multi-step editing · 1 src
- Perplexity AI releases decision model pplx-decider-v1-27b fine‑tuned from Qwen3.8-27B · 1 src
Comments
via GitHub Discussions