Theme
AI Visuals
LongLive
LongLive is an NVIDIA Labs infrastructure codebase for long video generation.
LongLive 2.0 provides parallel training and inference for long video, with NVFP4 and FP8 paths, multi-shot and image-to-video support, sequence parallelism, and asynchronous decoding. The project also publishes LongLive-RAG for retrieving long-video context during generation. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.
What it is
Infrastructure for long video generation
LongLive tackles the systems work needed for longer video generation rather than offering a simple consumer editor or one-prompt app.
Why it stands out
Parallelism, lower precision, and long sequences
Sequence parallelism, NVFP4 and FP8 paths, multi-shot support, image-to-video, asynchronous decoding, and retrieval all target the cost and continuity problems that grow with longer clips.
Availability
Repo, docs, models, and papers
The repository includes training and inference code, configuration files, documentation, model links, project pages, papers, and project-reported performance tables for readers comparing the technical direction.
Why it matters
What makes it useful
Long video generation can run into memory, speed, and continuity limits that a short demo hides. LongLive exposes the training, inference, precision, parallelism, decoding, and retrieval choices used to push past those limits.
What to know
Where it fits
It fits researchers and experienced builders comparing long-video training, inference optimisation, and interactive or streaming systems. It is infrastructure, not a no-code video generator.
Notable points
What stands out
The current project includes LongLive 2.0, NVFP4 training and inference, an FP8 inference path, multi-shot and image-to-video support, sequence parallelism, asynchronous decoding, LongLive-RAG, and the earlier LongLive 1.0 real-time interactive work.
Before using
What to review
Whether the goal is LongLive 2.0 infrastructure work or the older LongLive 1.0 branch.
The CUDA, GPU, model-checkpoint, NVFP4, TransformerEngine, FourOverSix, and configuration requirements for the intended setup.
The project-reported FPS, VBench, and model-table claims before using them as settled comparisons across video-generation systems.
Reader fit
Who may find it relevant
Readers tracking long video generation and real-time or interactive video systems.
Builders comparing training and inference infrastructure for diffusion-based video generation.
Less relevant for readers looking for a no-code video generator or casual laptop-friendly creative workflow.
Editorial note
Why LifeHubber lists it
LongLive is included because it exposes the systems work behind longer generated video: parallelism, lower-precision paths, asynchronous decoding, and retrieval. It helps experienced builders decide whether their real bottleneck is the model, the creative workflow, or the infrastructure carrying a long sequence.
Source links
Source materials
Reader note
Before relying on this entry
LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.
What to explore next
Connect long-video infrastructure to models and workflows.
LongLive handles the systems work behind longer generated video.
More in Music / Image Gen Models
Keep browsing this category
Explore more media-generation model resources.
ACE-Step 1.5
ace-step/ACE-Step-1.5
A locally runnable music model with 2B and XL 4B variants for full-song generation, reference-guided creation, repainting, accompaniment, stem separation, and lightweight style training.
Boogu Image
boogu-project/Boogu-Image
A 10B research model family for text-to-image generation and instruction-based image editing, with Base, Turbo, Edit, and Edit-Turbo checkpoints, bilingual Chinese-English text rendering, local inference code, and public demos.
CoMoVi
IGL-HKUST/CoMoVi
A framework for co-generating 3D human motion and realistic videos, with a focus on motion-conditioned video generation and training workflows.
For project maintainers
Listed here? You can use the badge.
If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.