LIFEHUBBER
Choose theme

AI Resources

MOSS-VL

GitHub stars: 747 GitHub forks: 36 Declared license: Apache-2.0: Apache-2.0 Last pushed September 28, 2026: Pushed 3d ago
Stats from GitHub

MOSS-VL is an OpenMOSS family of 11B vision-language models for realtime video streams, offline image and video work, and continued pre-training or fine-tuning.

Realtime, Instruct-0708 and Base-0708 serve different tasks. The project also bundles a specialized backend for several concurrent video streams and a browser demo with optional voice and memory components; these have their own setup and protocols. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.

What it is

Three checkpoints for different video workflows

Realtime handles continuous timestamped streams, Instruct-0708 is the offline instruction-following checkpoint, and Base-0708 is intended for continued pre-training and downstream fine-tuning.

Why it stands out

Watching and answering can happen together

The realtime design accepts timestamped frames while generating text. Questions can arrive before a video finishes, rather than being attached only to a completed clip.

Availability

Model, backend and demo materials

The public materials include checkpoints, reference inference scripts, a separate SGLang-format realtime backend, a browser application and fine-tuning examples. Downloading a checkpoint alone does not install the application.

Why it matters

What makes it useful

When replaying a file as a live stream, sampling and pacing are separate choices. The realtime guide recommends --playback-speed 1 for model inference because accelerating or disabling pacing can change generation behavior; it reserves speed 0 for source-only diagnostics such as --dry-run. --sample-fps chooses how often to sample, while --max-frames caps the replay. A longer recording can be truncated by that cap even when its playback speed is unchanged.

Notable points

What stands out

A client receiving the reference service's output gets raw model text, which can include response, silence and round-start control tokens. The realtime guide allows a client to hide those tokens in its display without changing model decoding. Decide whether a troubleshooting interface should show raw output or a reading interface should hide the markers; a display choice is separate from changing what the model generates.

Before using

What to review

The quantization guide separates stored weight conversion from generation-time KV-cache settings. It notes that weight quantization alone may not be enough for long videos because the cache grows with the stream. Its NF4 recipe is Transformers-only; its SGLang cache example instead uses the engine's own KV dtype. A smaller weight file does not describe the entire runtime memory configuration.

For remote deployment of the reference WebSocket service, its guide says the example has no authentication and recommends an authenticated reverse proxy with TLS and WebSocket upgrade support. The local example's availability is not a claim that it supplies those remote-access controls.

Reader fit

Who may find it relevant

For builders preparing fine-tuning examples, single-turn and multi-turn records place media differently. The prompt/response format automatically prepends media markers; conversations requires explicit image or video markers in the messages. Each video segment consumes a video marker. The guide warns that when there are too few video markers, preprocessing expands them and their final placement may differ from the original expectation.

Editorial note

Why LifeHubber lists it

In the bundled Demo, choose whether a long conversation should be summarized or reset. The docs say both commands require backend session recreation support. Its documentation says /compact keeps session Memory and recent turns but requires Memory and a configured summary service. /clear instead discards model context and session retrieval Memory while preserving chat archives and connection settings. The docs explicitly say these commands do not guarantee correction of visual hallucinations. Clearing model context and deleting a saved conversation are different actions.

Source links

Source materials

Reader note

Before relying on this entry

LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.

What to explore next

Take the fine-tuning workflow beyond the reference example.

For production or domain-specific fine-tuning, the MOSS-VL guide recommends LLaMA Factory or ms-swift rather than its minimal training reference. If that is your next task, continue with the framework's dataset, adapter, evaluation and export workflow.

Advertisements

Advertisements

For project maintainers

Listed here? You can use the badge.

If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.

See what’s moving