Project MonetRequest demo
Home/Blog/Tencent Hy4 Preview: API, Pricing, Benchmarks, 1M Context & Open Weights

AI · Project Monet Briefing

Tencent Hy4 Preview: API Pricing, Benchmarks, 1M Context and Open Weights

Tencent's 770B-total MoE preview combines 49B active parameters, a 1M-token context window, Apache-2.0 weights and official self-hosting recipes.

Published 2026-08-28 · Updated 2026-08-28 · By Project Monet Editorial Team

Project Monet editorial graphic for Tencent Hy4 Preview, a 770B MoE model with 1M-token context

01

What is Hy4 Preview?

Tencent released Hy4 Preview on August 28, 2026 as an open-weight flagship mixture-of-experts model for long-horizon coding, office work, game development and scientific tasks.

Tencent documents 770 billion total backbone parameters, 49 billion active parameters per token, 78 layers, 256 routed experts plus one shared expert, a 1 million-token context window and a native MTP layer for speculative decoding.

02

Open weights, license and local deployment

Tencent publishes official full and FP8 checkpoints under the Apache License 2.0. The official model card and repository document deployment through vLLM and SGLang, including Hy4-specific reasoning and tool-call parsers.

Once deployed, either runtime can expose an OpenAI-compatible endpoint. Tencent also points to AngelSlim for compression and quantization work.

The model is self-hostable, but its scale makes this a multi-GPU infrastructure project rather than a normal desktop-model workflow. See the Hy4 Preview local deployment guide for the official paths and hardware caveats.

03

Hosted API access and current pricing

Hosted access is currently verified through OpenRouter under tencent/hy4-preview. At the August 28 review, OpenRouter listed $0.834 per million input tokens, $2.501 per million output tokens and $0.042 per million cache-read tokens.

OpenRouter lists a 1,048,576-token context window, up to 64,000 completion tokens, tool calling and structured outputs. Actual usable context, latency and cost still depend on the provider and request.

04

Hy4 Preview benchmarks

Tencent reports a blind comparison across 203 engineering tasks evaluated by 163 internal experts. Its published average scores are 2.99 for Hy4 Preview, 2.92 for GLM 5.3 and 2.94 for Kimi K3.

Tencent also reports 46.8% wins, 12.8% ties and 40.4% losses against GLM 5.3, and 51.2% wins, 7.9% ties and 40.9% losses against Kimi K3.

05

Who should use Hy4 Preview?

Hy4 is most relevant when open weights, a very large context window, agentic coding, tool use or infrastructure control matter. Teams prioritising inexpensive hosted inference should compare live price, latency and reliability with GLM, Kimi, Qwen and closed models on their own workloads.

The next useful evidence will be independent benchmarks, mature lower-bit conversions and reproducible serving data. Until then, Tencent's release is best treated as a credible early signal rather than proof of a universal winner.

Sources

Primary and supporting sources

Facts were rechecked against the linked sources immediately before publication. Pricing, product availability and rollout status can change.

Project Monet

Useful signals. Clear decisions. Better digital work.

Project Monet turns relevant shifts in AI, creator tools and the web into practical context—and builds focused websites for businesses ready to grow.

Request a free homepage concept