Theme
AI Resources
ACE-Step 1.5
ACE-Step 1.5 is a locally runnable music model for generating complete songs, guiding results with reference audio, editing selected passages, separating tracks, and training lightweight style adapters.
It combines a language model that plans a song with a diffusion model that produces the audio, giving creators more than a single text-to-song prompt box. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.
What it is
A local music generation model
ACE-Step 1.5 can turn a description and lyrics into music from 10 seconds to 10 minutes long. Its official workflow also covers reference-guided generation, covers, repainting, accompaniment, stem separation, and multi-track additions.
Why it stands out
One local workflow from draft to revision
The useful difference is the range of work kept in one project: plan a song, generate it, revise a section, use a reference track, extract musical information, separate stems, or train a LoRA on a small set of songs.
Availability
Code, model weights, demo, and setup guides
The public release includes the original 2B path and newer ACE-Step 1.5 XL 4B DiT variants, alongside a hosted demo, installation guides, hardware guidance, and a research paper. Local setup supports NVIDIA, AMD, Intel, Apple Silicon, and CPU paths, with different limits.
Why it matters
What makes it useful
ACE-Step 1.5 lets creators test more of a music workflow on their own hardware instead of stopping at one generated track. The same project can create a song, take guidance from existing audio, repair a section, build accompaniment around vocals, separate stems, and train a small style adapter.
What to know
Where it fits
Treat it as a local AI music workbench that can feed a wider production setup, not as a replacement for a digital audio workstation. It is most relevant when you want to generate and reshape source audio before arranging, mixing, or finishing it elsewhere.
Notable points
What stands out
The speed and quality comparisons come from the project and its paper, not an independent LifeHubber test. Model choice now matters too: the 2B path keeps the lower-memory baseline, while the XL 4B variants trade higher memory needs for a different generation path.
Before using
What to review
For the 2B path, the installation guide lists Python 3.11-3.12, about 10 GB of disk space for the core models, at least 4 GB VRAM for DiT-only mode, and at least 6 GB for the language-model-plus-DiT path. The XL 4B variants list at least 12 GB VRAM with offload or quantization and at least 20 GB without; CPU inference is supported but described as significantly slower.
The project reports inconsistent results across random seeds and durations, weaker performance in some styles, rough transitions during repainting or extension, coarse vocals, and limited fine-grained control.
Reference audio, cover generation, and style training can involve work you do not own. Before publishing or monetizing a result, check the project terms and make sure you have permission to use any lyrics, recordings, voices, or other source material you provide.
Reader fit
Who may find it relevant
Musicians and creators who want a locally runnable starting point for full-song generation and revision.
Builders comparing GPU requirements, model variants, APIs, and local creative workflows.
Less relevant if you want a polished hosted service with no setup, or a full arranging and mixing environment.
Editorial note
Why LifeHubber lists it
LifeHubber lists ACE-Step 1.5 because it connects local song generation with practical revision tools. It gives creators a concrete way to compare how much of an AI music workflow can stay on their own machine, and what still needs careful listening, editing, and rights checks.
Source links
Source materials
Reader note
Before relying on this entry
LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.
More in Music / Image Gen Models
Keep browsing this category
Explore more media-generation model resources.
CoMoVi
IGL-HKUST/CoMoVi
A framework for co-generating 3D human motion and realistic videos, with a focus on motion-conditioned video generation and training workflows.
Boogu Image
boogu-project/Boogu-Image
A 10B research model family for text-to-image generation and instruction-based image editing, with Base, Turbo, Edit, and Edit-Turbo checkpoints, bilingual Chinese-English text rendering, local inference code, and public demos.
PersonaLive
GVCLab/PersonaLive
A portrait image-animation framework for live-streaming-style video generation research, with offline and online inference, pretrained weights, a Web UI, and acceleration notes.
For project maintainers
Listed here? You can use the badge.
If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.