LIFEHUBBER
Choose theme

AI Resources

Hy4 preview

Hugging Face likes: 502 Hugging Face downloads, last 30 days: 25.9K Declared license: Apache-2.0: Apache-2.0 Last modified August 28, 2026: Modified 1mo ago
Stats from Hugging Face

Hy4 preview is Tencent Hy Team's early release of a very large Mixture-of-Experts language model for coding, document and analysis work, game development, research, tool use, and long-context tasks.

Tencent lists 770B total parameters, 49B active parameters per token, a 1 million-token context window, public BF16 and FP8 weights, and dedicated deployment paths for vLLM and SGLang. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.

What it is

A 770B-total MoE model

Hy4 preview has 49B active parameters per token across a 770B-parameter backbone. Tencent also builds in a multi-token-prediction layer for speculative decoding.

Why it stands out

One model for very large work contexts

Tencent positions the model around long-horizon software work, files and office artifacts, playable game prototypes, difficult research questions, and tool-driven tasks, with a 1M-token context window for large inputs.

Availability

Public weights, heavy deployment

The official Hugging Face pages host BF16 and FP8 checkpoints. The vLLM recipe uses an eight-GPU FP8 setup; current SGLang recipes vary by hardware and precision, including four-, eight-, and sixteen-GPU configurations through its dedicated image.

Why it matters

What makes it useful

For a coding task spread across many files, Tencent positions Hy4 preview around long software workflows and large work contexts. A team evaluating the public weights can supply project material to a local serving stack rather than reduce the task to one short snippet. File access, tool execution, permissions, and checks still belong to the application around the model.

Notable points

What stands out

The model's 1M-token ceiling is not the default serving budget in every recipe. SGLang documents 131K context on its H200 BF16 setup and 262K on several newer GPU setups; the vLLM AMD recipes document different validated limits too. Choose the input budget from the matching recipe rather than assuming the model headline fits your hardware. Tencent also warns that this preview can reason and verify for too long; its productivity and benchmark results remain project-reported evidence.

Before using

What to review

Start with the exact BF16 or FP8 model page and the matching current vLLM or SGLang recipe; the deployment examples use dedicated images and model-specific parsers.

Budget for model files, GPU memory, cache, runtime overhead, and the parallelism in the exact hardware recipe. The SGLang FP8 path requires Blackwell GPUs; its H200 recipe uses BF16 across two eight-GPU nodes.

Test whether the default high reasoning effort helps the task. Tencent documents a no-think setting for direct responses and separately warns that the preview may reason and verify for too long.

Review tool calls, generated code, documents, calculations, and game or research outputs before they affect a repository, business decision, public surface, or other real system.

The project identifies Apache-2.0. Review the current terms at the main official source for your intended use.

Reader fit

Who may find it relevant

Model and infrastructure teams with multi-GPU systems evaluating flagship-scale public weights.

Builders testing very long coding, document, analysis, research, or tool-driven workflows.

Teams comparing high-reasoning and direct-response modes behind an OpenAI-compatible local API.

Less relevant for readers seeking a small local model, a consumer chatbot, or image and video generation.

Editorial note

Why LifeHubber lists it

Hy4 preview is useful for builders who need to display final answers separately from reasoning in an application. Its documented vLLM and SGLang client examples expose those parts separately, with an explicit field mapping for each runtime. When displaying a response, check which field your chosen serving example uses for the final answer. The vLLM recipe prints thinking from message.reasoning and the answer from message.content; the SGLang cookbook uses message.reasoning_content and message.content. Read the answer from content and handle any separate reasoning field deliberately, rather than assuming both runtimes expose the same field name. These are the documented client examples, not a guarantee for every SDK or configuration.

Source links

Source materials

Reader note

Before relying on this entry

LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.

What to explore next

Compare the preview with a smaller model and the wider model field.

Hy4 preview pushes model size and context much further than Hy3, while also demanding substantially heavier infrastructure. These paths help separate that scale choice from model-to-agent setup and other model options.

Advertisements

Advertisements

For project maintainers

Listed here? You can use the badge.

If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.

See what’s moving