Choose theme
AI Resources
Hy4 preview
Hy4 preview is Tencent Hy Team's early release of a very large Mixture-of-Experts language model for coding, document and analysis work, game development, research, tool use, and long-context tasks.
Tencent lists 770B total parameters, 49B active parameters per token, a 1 million-token context window, public BF16 and FP8 weights, and dedicated deployment paths for vLLM and SGLang. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.
What it is
A 770B-total MoE model
Hy4 preview has 49B active parameters per token across a 770B-parameter backbone. Tencent also builds in a multi-token-prediction layer for speculative decoding.
Why it stands out
One model for very large work contexts
Tencent positions the model around long-horizon software work, files and office artifacts, playable game prototypes, difficult research questions, and tool-driven tasks, with a 1M-token context window for large inputs.
Availability
Public weights, heavy deployment
The official Hugging Face pages host BF16 and FP8 checkpoints. The vLLM recipe uses an eight-GPU FP8 setup; current SGLang recipes vary by hardware and precision, including four-, eight-, and sixteen-GPU configurations through its dedicated image.
Why it matters
What makes it useful
For a coding task spread across many files, Tencent positions Hy4 preview around long software workflows and large work contexts. A team evaluating the public weights can supply project material to a local serving stack rather than reduce the task to one short snippet. File access, tool execution, permissions, and checks still belong to the application around the model.
What to know
Where it fits
Choose the BF16 or FP8 checkpoint and match its current recipe to your GPUs before wiring an application to the local API. The vLLM recipe enables the hy_v4 reasoning and tool-call parsers; SGLang uses its Hunyuan parsers through auto. The Hugging Face repositories contain about 1.56 TB of BF16 files or 814 GB of FP8 files, so plan storage and serving memory separately. These are multi-GPU deployments, not laptop or no-setup chatbots.
Notable points
What stands out
The model's 1M-token ceiling is not the default serving budget in every recipe. SGLang documents 131K context on its H200 BF16 setup and 262K on several newer GPU setups; the vLLM AMD recipes document different validated limits too. Choose the input budget from the matching recipe rather than assuming the model headline fits your hardware. Tencent also warns that this preview can reason and verify for too long; its productivity and benchmark results remain project-reported evidence.
Before using
What to review
Start with the exact BF16 or FP8 model page and the matching current vLLM or SGLang recipe; the deployment examples use dedicated images and model-specific parsers.
Budget for model files, GPU memory, cache, runtime overhead, and the parallelism in the exact hardware recipe. The SGLang FP8 path requires Blackwell GPUs; its H200 recipe uses BF16 across two eight-GPU nodes.
Test whether the default high reasoning effort helps the task. Tencent documents a no-think setting for direct responses and separately warns that the preview may reason and verify for too long.
Review tool calls, generated code, documents, calculations, and game or research outputs before they affect a repository, business decision, public surface, or other real system.
The project identifies Apache-2.0. Review the current terms at the main official source for your intended use.
Reader fit
Who may find it relevant
Model and infrastructure teams with multi-GPU systems evaluating flagship-scale public weights.
Builders testing very long coding, document, analysis, research, or tool-driven workflows.
Teams comparing high-reasoning and direct-response modes behind an OpenAI-compatible local API.
Less relevant for readers seeking a small local model, a consumer chatbot, or image and video generation.
Editorial note
Why LifeHubber lists it
Hy4 preview is useful for builders who need to display final answers separately from reasoning in an application. Its documented vLLM and SGLang client examples expose those parts separately, with an explicit field mapping for each runtime. When displaying a response, check which field your chosen serving example uses for the final answer. The vLLM recipe prints thinking from message.reasoning and the answer from message.content; the SGLang cookbook uses message.reasoning_content and message.content. Read the answer from content and handle any separate reasoning field deliberately, rather than assuming both runtimes expose the same field name. These are the documented client examples, not a guarantee for every SDK or configuration.
Source links
Source materials
Reader note
Before relying on this entry
LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.
What to explore next
Compare the preview with a smaller model and the wider model field.
Hy4 preview pushes model size and context much further than Hy3, while also demanding substantially heavier infrastructure. These paths help separate that scale choice from model-to-agent setup and other model options.
More in AI Models
Keep browsing this category
Explore more AI model resources.
Gemma 4
google/gemma-4
A Google DeepMind Gemma 4 model family collection with public checkpoints including Gemma 4 12B, a dense multimodal model Google describes around local agentic workflows, native audio input, and encoder-free vision/audio handling.
MiniMax H3 Integrations
MiniMax-AI/awesome-minimax-h3-integration
A community-maintained MiniMax H3 integration index that maps checkpoints, hardware and VRAM starting points, runtimes, ComfyUI nodes, prompting tools, acceleration routes, and deployment options.
GLM-5.3-Flash
zai-org/GLM-5.3-Flash
A Z.ai 320B-total, 18B-active multimodal mixture-of-experts model for coding, agent workflows, visual input, and long-context work, with a 1M-token context window, public FP8 and BF16 weights, API access, and several technical serving paths.
For project maintainers
Listed here? You can use the badge.
If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.