vLLM
The [Model] Support Hy3 preview pull request was merged on April 23, 2026.
This reference covers the open mixture reasoning and agent checkpoint with a quantized companion. It separates text-backed benchmark claims from the image-only appendix and keeps provider pricing distinct.
Hy3 is the vendor reference for tencent/Hy3, as of Aug 12, 2026. Model scale: 295B total; 21B activated; 3.8B MTP; Context length: 256K, as of Aug 12, 2026. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | tencent/Hy3 | huggingface.coRetrieved |
| Identity | Vendor description | Hy3 is a 295B-parameter Mixture-of-Experts (MoE) model with 21B active parameters and 3.8B MTP layer parameters, developed by the Tencent Hy Team. | huggingface.coRetrieved |
| Identity | Lineage | Following the Hy3 Preview launch in late April, we gathered feedback from 50+ products and scaled up post-training with higher quality data. | huggingface.coRetrieved |
| Identity | OpenRouter release date | Jul 6, 2026 | openrouter.aiRetrieved |
| Architecture | Model scale | 295B total; 21B activated; 3.8B MTP | huggingface.coRetrieved |
| Architecture | Layer configuration | 80 layers excluding MTP; 1 MTP layer | huggingface.coRetrieved |
| Architecture | Expert routing | 192 experts, top-8 activated | huggingface.coRetrieved |
| Architecture | Attention | 64 (GQA, 8 KV heads, head dim 128) | huggingface.coRetrieved |
| Architecture | Hidden and intermediate sizes | hidden 4096; intermediate 13312 | huggingface.coRetrieved |
| Architecture | Vocabulary | 120832 | huggingface.coRetrieved |
| Architecture | Precision | BF16 | huggingface.coRetrieved |
| Architecture | Vendor architecture framing | Built on a hybrid fast-and-slow-thinking Mixture-of-Experts (MoE) architecture | www.tencent.comRetrieved |
| Context | Context length | 256K | huggingface.coRetrieved |
| Modalities | Pipeline tag | text generation | huggingface.coRetrieved |
| License | Weights license | Hy3 is released under the Apache License 2.0. | huggingface.coRetrieved |
| License | License file | Apache License, Version 2.0; Copyright (C) 2026 Tencent. All rights reserved. | raw.githubusercontent.comRetrieved |
| Pricing | Tencent API price: OpenRouter listing | $0.1288 per 1M input tokens / $0.5336 per 1M output tokens | openrouter.aiRetrieved |
| Availability | Repository access | Ungated repository. | huggingface.coRetrieved |
| Availability | Weight distribution | We open-source Hy3 and Hy3-FP8 model weights on Hugging Face, ModelScope, GitCode, and CNB. | huggingface.coRetrieved |
| Availability | FP8 companion | tencent/Hy3-FP8; FP8 quantized instruct model | huggingface.coRetrieved |
| Availability | WorkBuddy access | available free of charge to users worldwide until 31 August 2026 (Pacific Time) | www.tencent.comRetrieved |
| Availability | vLLM recipe floor | vLLM 0.26.0+ | recipes.vllm.aiRetrieved |
| Availability | Vendor serving recommendation | For production serving, we recommend using vLLM or SGLang, both of which provide dedicated recipes for Hy3 | huggingface.coRetrieved |
| Availability | Main weights | tencent/Hy3 | huggingface.coRetrieved |
| Availability | Quantized instruct weights | tencent/Hy3-FP8 | huggingface.coRetrieved |
"Today, we introduce Hy3, which outperforms similar-size models and rivals flagship open-source models with 2-5x parameters."
huggingface.coRetrieved
"Hy3 recorded more than 68 times as many API calls as the previous-generation model and ranked first globally on OpenRouter's global LLM usage leaderboard within one week of launch."
www.tencent.comRetrieved
These results are vendor-claimed, not independently measured by RunInfra.
| Benchmark | Vendor-claimed value | Source and date |
|---|---|---|
| Expert blind evaluation | Hy3 scored 2.67/4, outperforming GLM-5.1 at 2.51/4 | huggingface.coRetrieved |
| Internal hallucination evaluation | dropped from 12.5% to 5.4% in internal evaluations | huggingface.coRetrieved |
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
The [Model] Support Hy3 preview pull request was merged on April 23, 2026.
The Hy3 day-zero cookbook pull request was merged on July 6, 2026.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of Aug 12, 2026, Release identity: tencent/Hy3.
As of Aug 12, 2026, Repository access: Ungated repository.
As of Aug 12, 2026, Repository access: Ungated repository.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs