Stable Audio 2.0 icon

Stable Audio 2.0

Open weights
Stability AI

Stable Audio 2.0 is Stability AI's latest music and sound generation model, released in April 2024, capable of producing high-quality stereo audio up to 3 minutes in length at 44.1kHz from text prompts.

Text to Audio

Key Highlights

3-Minute Coherent Music

Generates music up to 3 minutes with coherent song structures including intros, verses, choruses, and outros.

Audio-to-Audio Transformation

Transforms uploaded audio samples into new compositions while preserving structural elements.

Licensed Training Data

Trained on licensed dataset from AudioSparx, providing commercial use safety.

CD-Quality Audio

Professional-grade outputs with stereo audio production at 44.1kHz sample rate.

About

Stable Audio 2.0 is Stability AI's second-generation audio generation model, representing a significant advancement in AI-powered music and sound creation. Released in April 2024, the model builds upon the original Stable Audio by extending output duration from 90 seconds to 3 minutes, introducing audio-to-audio transformation capabilities, and improving overall generation quality with coherent musical structure.

The model generates stereo audio at CD-quality 44.1kHz sample rate, producing professional-grade sound suitable for commercial use. A key technical achievement is the model's ability to generate audio with coherent musical structure — songs feature appropriate intros, verses, choruses, bridges, and outros that follow genre conventions. This structural coherence, which was a major weakness of earlier audio generation models, makes Stable Audio 2.0's outputs much more usable as complete musical pieces.

The audio-to-audio generation capability is a notable innovation. Users can upload existing audio samples and use text prompts to guide transformation of these samples into new compositions. The model can maintain rhythm, melody, or structural elements from the input while applying new instrumentation, genre characteristics, or sonic textures. This enables creative workflows like remixing, style transfer, and sample-based composition.

Training data for Stable Audio 2.0 comes from a licensed dataset provided by AudioSparx, a music licensing library. This licensed training approach, similar to Adobe's strategy with Firefly, provides commercial safety for users generating content for business applications, reducing legal risk associated with AI-generated music.

The model is accessible through the Stable Audio web platform at stableaudio.com, which provides an intuitive interface for text-to-audio and audio-to-audio generation. API access is available for developers integrating audio generation into applications. Stability AI has also released an open-source version of the model under its Community License for non-commercial research.

In the competitive landscape, Stable Audio 2.0 occupies a distinct niche between Suno/Udio (which focus on vocal-rich popular songs) and production-oriented tools like AudioCraft and MusicGen (which focus on instrumental generation). Its combination of licensed training data, audio-to-audio capabilities, and the availability of an open-source variant for research makes it a unique offering in the AI audio space.

Use Cases

1

Background Music Production

Creating custom background music for video content, podcasts, and presentations.

2

Sound Effect Design

Producing custom sound effects and ambient sounds for games, films, and applications.

3

Remix and Style Transfer

Creating creative remixes by transforming existing audio samples into different genres and styles.

4

Research and Prototyping

Audio generation research and prototype application development with the open-source model.

Pros & Cons

Pros

  • Licensed training data provides safety for commercial use
  • Audio-to-audio transformation opens unique creative possibilities
  • Music generation with coherent song structures up to 3 minutes
  • Open source variant available for research and learning

Cons

  • Vocal quality cannot reach Suno or Udio level
  • 3-minute maximum duration insufficient for longer compositions
  • Open source variant licensed for non-commercial research only
  • Genre diversity and production quality behind Suno/Udio

Technical Details

Parameters

undisclosed

License

Stability AI Community License + Commercial

Features

  • Text-to-Audio Generation
  • Audio-to-Audio Transformation
  • 44.1kHz Stereo Output
  • 3-Minute Duration
  • Song Structure Coherence
  • Sound Effect Generation
  • Licensed Training Data
  • Open Source Variant

Benchmark Results

MetricValueCompared ToSource
Max Duration3 minutesSuno: 4 min, Udio: 2 minStability AI
Sample Rate44.1kHz stereoCD qualityStability AI
Training DataLicensed (AudioSparx)Suno/Udio: undisclosedStability AI

Available Platforms

stable audio platform
hugging face
api

News & References

Frequently Asked Questions

Related Models

AudioCraft icon

AudioCraft

Meta|N/A

AudioCraft is Meta AI's comprehensive open-source framework for generative audio research and applications, bringing together three specialized models under a single integrated platform: MusicGen for music generation, AudioGen for sound effect synthesis, and EnCodec for neural audio compression. Released in August 2023 under the MIT license, AudioCraft provides a unified codebase that simplifies working with state-of-the-art audio generation models through consistent APIs and shared infrastructure. The framework is built on a transformer-based architecture where audio signals are first compressed into discrete tokens by EnCodec, then generated autoregressively by task-specific language models. MusicGen handles text-to-music generation with melody conditioning support, while AudioGen specializes in environmental sounds, sound effects, and non-musical audio from text descriptions. EnCodec serves as the neural audio codec backbone, compressing audio at various bitrates while maintaining high perceptual quality. AudioCraft supports multiple model sizes, stereo generation, and provides extensive training and inference utilities. The framework includes pre-trained models for immediate use and tools for training custom models on user-provided datasets. As a Python library installable via pip, AudioCraft integrates seamlessly into existing machine learning and audio processing pipelines. It is widely used by researchers studying audio generation, developers building creative audio tools, content creators needing original music and sound effects, and game studios requiring dynamic audio systems. AudioCraft represents Meta's most significant contribution to open-source audio AI and has become the foundation for numerous community projects and commercial applications in the rapidly growing AI audio generation space.

Open weights
AudioLDM 2 icon

AudioLDM 2

CUHK & Surrey|N/A

AudioLDM 2 is a unified audio generation framework developed by researchers at the Chinese University of Hong Kong and the University of Surrey, capable of producing music, sound effects, and speech from text descriptions within a single model. Building on the original AudioLDM, version 2 introduces a universal audio representation called Language of Audio that bridges the gap between different audio types by encoding them into a shared semantic space. The model combines a GPT-2 language model for understanding text inputs with an AudioMAE encoder for audio conditioning, feeding into a latent diffusion model that generates audio spectrograms which are converted to waveforms. This architecture enables AudioLDM 2 to handle diverse audio generation tasks without requiring separate specialized models for each audio type. The model achieves competitive performance across multiple benchmarks including text-to-music, text-to-sound-effects, and text-to-speech evaluations. AudioLDM 2 generates audio at up to 48 kHz with good perceptual quality for both musical and non-musical content. Released in August 2023 under a research license, the model is open source with code and pre-trained weights available on GitHub and Hugging Face. AudioLDM 2 supports audio inpainting, style transfer, and super-resolution in addition to text-conditioned generation. The model is particularly relevant for researchers studying unified audio generation, content creators needing diverse audio types from a single tool, and developers building comprehensive audio generation systems. Its unified approach to handling speech, music, and environmental sounds makes it a versatile foundation for multi-purpose audio applications.

Open weights
Bark icon

Bark

Suno AI|N/A

Bark is a transformer-based text-to-audio generation model developed by Suno AI that converts text into natural-sounding speech, music, and sound effects. Released as open source under the MIT license in April 2023, Bark goes far beyond traditional text-to-speech systems by generating not only spoken words but also laughter, sighs, music, and ambient sounds from text descriptions. The model uses a GPT-style autoregressive transformer architecture with EnCodec audio tokenizer to generate audio tokens that are then decoded into waveforms. Bark supports multiple languages including English, Chinese, French, German, Hindi, Italian, Japanese, Korean, Polish, Portuguese, Russian, Spanish, and Turkish, making it one of the most multilingual open-source audio generation models available. The model can clone voice characteristics from short audio samples, allowing users to generate speech in specific voices or speaking styles. Bark operates in a zero-shot manner, meaning it can produce diverse outputs without task-specific fine-tuning. Generation includes natural prosody, emotion, and intonation that closely mimics human speech patterns. The model generates audio at 24 kHz sample rate with reasonable quality for most applications. As a fully open-source project with pre-trained weights available on Hugging Face and GitHub, Bark is widely used by developers building voice applications, content creators producing multilingual audio, and researchers exploring generative audio models. The model is particularly valued for its versatility in handling diverse audio types within a single unified architecture and its accessibility for rapid prototyping of audio generation applications.

Open weights
MusicGen icon

MusicGen

Meta|3.3B

MusicGen is a single-stage transformer-based music generation model developed by Meta AI Research as part of the AudioCraft framework. Released in June 2023 under the MIT license, MusicGen uses a single autoregressive language model operating over compressed discrete audio representations from EnCodec, unlike cascading approaches that require multiple models. The model comes in multiple sizes ranging from 300M to 3.3B parameters, allowing users to balance quality against computational requirements. MusicGen generates high-quality mono and stereo music at 32 kHz from text descriptions, supporting a wide range of genres, instruments, moods, and musical styles. Users can describe desired music using natural language prompts specifying genre, tempo, instrumentation, and atmosphere, and the model produces coherent musical compositions that follow the specified characteristics. Beyond text-to-music generation, MusicGen supports melody conditioning where an existing audio clip guides the melodic structure of the generated output, enabling more controlled music creation. The model achieves strong results across both objective metrics and subjective listening evaluations, producing music that sounds natural and musically coherent for durations up to 30 seconds. As a fully open-source model with code and weights available on GitHub and Hugging Face, MusicGen has become one of the most widely adopted AI music generation tools in both research and creative communities. It integrates easily into existing audio production workflows through the Audiocraft Python library and various community-built interfaces. MusicGen is particularly popular among content creators, game developers, and musicians who need royalty-free background music generated on demand.

Open weights

Quick Info

Parametersundisclosed
Typediffusion
LicenseStability AI Community License + Commercial
Released2024-04
CreatorStability AI

Links

Tags

stable-audio
müzik
text-to-audio
ses-efekti
stability-ai
Visit Website
Stable Audio 2.0: open-weight Text to Audio AI model profile | tasarim.ai