Models

ACE-Step 1.5 Brings Local Music Generation to C++

ACE-Step 1.5 is now available in quantized GGUF format alongside a lightweight C++17 runtime, enabling developers to generate 48kHz stereo music locally without needing Python or PyTorch.

AlphaSignal1 day agoModels
Image: AlphaSignal

A community release has brought quantized GGUF weights and a dedicated C++17 runtime to the ACE-Step 1.5 text-to-music model. Powered by the new acestep.cpp project built on GGML, the implementation allows engineers to deploy the audio pipeline locally across CPU, CUDA, Metal, and Vulkan backends. The framework outputs stereo 48kHz audio directly from text captions and lyrics, bypassing the heavy Python environments, CUDA-specific wheels, and full-precision checkpoints traditionally required for local generative audio.

The architecture splits responsibilities across several specialized models. Text captions and optional lyrics are processed by a Qwen3 text encoder before passing to a Qwen3 language model, available in 0.6B, 1.7B, and 4B parameter sizes. This language model converts prompts into song metadata, lyrics, and a 5 Hz audio-code sequence with a 64,000-entry vocabulary, where each code covers 200 milliseconds. Next, a flow-matching diffusion transformer—available in 2B or 4B XL sizes—converts the plan into 25 Hz audio latents that define timbre, transients, and stereo structure, before a BF16 VAE synthesizes the final waveform.

Quantization options span from BF16 down to Q4_K_M, though the 4B language model explicitly skips Q4_K_M because that compression level breaks audio code generation. The full Q8_0 turbo model set requires approximately 7.7 GB of storage. According to the underlying research paper, ACE-Step 1.5 can generate full songs in under 2 seconds on an A100 GPU while utilizing under 4GB of VRAM. Beyond standard generation, the system supports cover generation, audio repainting, vocal-to-BGM conversion, LoRA personalization, and over 50 languages.

For AI practitioners and product teams, this port drastically lowers the system requirements and integration complexity for audio synthesis. By using GGUF storage formats and native C++ execution, developers can embed low-latency music generation into consumer applications or edge devices without managing heavy PyTorch dependencies.

This is our own summary of reporting by AlphaSignal

More in Models