Stable Audio 3.0, Explained Simply
What Stable Audio 3.0 is, what it does well, and what it costs โ released 2026-05-20 by Stability AI.
TLDR
- Stable Audio 3.0 is a new family of AI models from Stability AI that generates music and sound effects.
- You can run some of these models directly on your own phone or laptop, completely offline.
- You can generate full songs up to over six minutes long, and you fully own whatever you create.
- The models are trained on fully licensed data, meaning the company legally paid for the audio it learned from.
Want a heads-up when Stability AI updates?
We email you the moment a new Stability AI version drops โ plus plain-English release notes like these. No spam.
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
Stable Audio 3.0 is a set of AI tools designed for generating audio, music, and sound effects. It is built for musicians, developers, and creative tinkerers who want to experiment with making AI-generated tracks or build new music apps.
What it does well
- Runs on your own devices: Unlike most AI that needs the internet, the "Small" models can create full songs or sound effects right on your phone or laptop.
- Longer tracks: It can generate audio for more than six minutes, letting you create complete songs rather than just short clips.
- Audio editing: You can modify just one part of a track, rewrite a section, or extend a song without having to start all over again.
- Custom training: It supports "LoRa," a method that lets you teach the model your own specific style or library of sounds efficiently.
How it compares
- Compared to older models: The previous "Small" model could only generate 11 seconds of audio, and the older "Open" model maxed out at 47 seconds. The new 3.0 Small model can generate up to two minutes, while the Medium and Large models can go past six minutes.
- Compared to rivals: The makers say this is the only model capable of full music composition directly on a device. Also, unlike other open music models that either block commercial use or risk legal trouble because they learned from unlicensed music, this model was trained on fully licensed data.
What it costs and what it can handle
- Length limits: The Small SFX and Small models generate up to two minutes of audio. The Medium and Large models can generate over six minutes (up to 6:20).
- Pricing and Access: The Small SFX, Small, and Medium models are free to download as "open weights" (meaning the public can download the model's core files). The Large model is available through an API (a service developers pay to access) or for big companies to host on their own servers.
- Ownership: Under the Stability AI Community License, you own your outputs and can sell or distribute them. However, if your organization makes over $1 million a year, you have to get an Enterprise License.
Worth knowing
- Different sizes for different needs: You have to choose the right model for your task. "Small SFX" is just for sound effects, "Small" is for full songs on portable devices, "Medium" has better musical structure and longer tracks, and "Large" is built for high-volume business platforms.
- Where to find them: The free downloadable models are on Hugging Face (a popular website for sharing AI models), while the Large model is accessed through the company's API or partner platforms like ComfyUI.
This is our plain-language summary. Read the complete, official notes on the Stability AI official changelog โ.
