概要
Tencent Hy3 is a 295B-parameter Mixture-of-Experts (MoE) model with only 21B active parameters per token, released under the commercially permissive Apache 2.0 license on July 6, 2026. Positioned as a cost-effective, production-ready alternative to larger frontier models, Hy3 excels in agentic tasks, tool orchestration, and long-context reasoning while featuring dramatically reduced hallucination rates (5.4%) for enhanced reliability. The model represents a significant improvement over its April preview, incorporating feedback from over 50 internal Tencent product teams and showing substantial gains in reasoning, agent capabilities, and real-world deployment stability.
Hy3's strategic positioning focuses on delivering frontier-adjacent performance at a fraction of the cost of larger models like GLM-5.2 (744B parameters), making it particularly attractive for organizations that prioritize cost efficiency, licensing flexibility, and production reliability over absolute peak performance in coding tasks. The model's design is optimized for deployment on export-compliant hardware (like Nvidia's H20-3e) while also running efficiently on standard Western data center GPUs, and its 256K context window supports complex, multi-step workflows that are essential for modern agent applications.
ベンチマーク&性能
Hy3 delivers competitive benchmark performance across multiple categories, with notable strengths in agentic and reasoning tasks. According to Tencent's published benchmarks and independent reports:
| Benchmark | Hy3 Score | GLM-5.2 Score | Notes |
|-----------|-----------|---------------|-------|
| SWE-Bench Verified | 78.0% | 84.2% | Hy3 trails in coding tasks |
| BrowseComp | 84.2 | 78.0 | Hy3 leads in agentic search |
| MCP-Atlas | 79.1 | 72.0 | Hy3 leads in tool orchestration |
| GPQA Diamond | 90.4% | 88.0% | Hy3 leads in STEM reasoning |
| DeepSWE | 28.0% | 46.2% | Significant coding gap |
| Terminal-Bench 2.1 | 71.7 | 81.0 | Coding benchmark |
| FrontierScience-Olympiad | 74.8 | 72.5 | Scientific reasoning |
Hy3's hallucination rate dropped from 12.5% in the preview version to 5.4% in the release, with commonsense error rates falling from 25.4% to 12.7%. The model shows particular strength in agentic tasks, scoring 84.2 on BrowseComp (competitive with Claude Opus 4.8 at 84.3 and GPT-5.5 at 84.4) and 91.0 on DeepSearchQA. Long-context performance is strong with a 256K token context window and 73.4 on AA-LCR for retrieval tasks.
In blind evaluations with 270 domain experts using real-world tasks, Hy3 scored 2.67/4, outperforming GLM-5.1's 2.51/4, with advantages in frontend development, CI/CD, and data/storage tasks.
詳細比較
**Hy3 vs GLM-5.2:**
- **Architecture & Scale**: Hy3 (295B total, 21B active) vs GLM-5.2 (744B total, ~40B active). GLM-5.2 has nearly double the per-token compute.
- **Coding Performance**: GLM-5.2 dominates all coding benchmarks (SWE-Bench: 84.2% vs 78.0%, DeepSWE: 46.2% vs 28.0%).
- **Agentic Tasks**: Hy3 leads in search and tool orchestration (BrowseComp: 84.2 vs 78.0, MCP-Atlas: 79.1 vs 72.0).
- **Cost**: Hy3 is significantly cheaper ($0.14/$0.56 per million tokens vs GLM-5.2's $1.40/$4.40).
- **Deployment**: Hy3's FP8 footprint is under 300GB (fits 8x H200 node) vs GLM-5.2's ~744GB (requires 8x H200 minimum).
- **License**: Hy3 uses Apache 2.0 (no regional restrictions) vs GLM-5.2's restrictive license excluding EU/UK/South Korea.
**Hy3 vs DeepSeek V4 Flash:**
- **Scale**: Similar total parameters (295B vs 284B), but DeepSeek V4 activates fewer parameters (~13B vs 21B).
- **Coding**: DeepSeek V4 Flash shows better coding benchmark performance in some tests.
- **Cost**: DeepSeek V4 Flash is slightly cheaper ($0.14/$0.28 vs $0.14/$0.56).
- **Deployment**: DeepSeek V4 handles aggressive quantization better, while Hy3 offers more consistent agent scaffold performance.
- **Reliability**: Users report Hy3 stays on track more reliably in complex tasks, while DeepSeek V4 can build incorrect mental models.
**Hy3 vs Frontier Proprietary Models (GPT-5.5, Claude Opus 4.8):**
- **Performance**: Hy3 approaches frontier performance on agentic tasks (BrowseComp parity with Claude Opus 4.8) but trails on complex coding and math benchmarks.
- **Cost**: Dramatically cheaper than proprietary models (71x cheaper than Fable 5 for input tokens).
- **Flexibility**: Apache 2.0 license enables self-hosting and customization vs API-only access.
コミュニティ評価
The developer and researcher community has shown significant interest in Hy3, with reactions focusing on several key aspects:
**Positive Reactions:**
- The Apache 2.0 license change has been widely praised as the "real headline," removing regional restrictions that hampered previous Chinese open models. Researchers on X noted that "if the scores hold up, Tencent has just become one of the leaders of open source."
- Developers appreciate the model's production focus, with one noting: "I've found DS4 Flash to be very temperamental via Claude Code... Hy3 isn't as fast, but so far it seems to stay on track much more reliably."
- The cost-to-performance ratio is highlighted as a major advantage, particularly for organizations with infrastructure constraints.
- Integration with multiple platforms (OpenRouter, Hermes, Kilo, Cline, etc.) is seen as lowering adoption barriers.
**Skepticism & Critiques:**
- Some community members express skepticism about benchmark validity, with comments suggesting potential "benchmaxxing" in Chinese models.
- The coding performance gap vs GLM-5.2 is frequently noted as a significant limitation for developer-focused use cases.
- Practical deployment concerns exist regarding GPU requirements and KV cache limitations ("with Hy3 quantized to FP4, there is only room for 130K tokens of KV cache").
- The fast development cycle from preview to release (about 10 weeks) raises questions about thoroughness.
**Adoption Patterns:**
- Early adoption is strongest among teams prioritizing agentic workflows, cost efficiency, and licensing flexibility.
- The model is being tested in production environments within Tencent's ecosystem (WorkBuddy, CodeBuddy, Yuanbao) with reported latency drops of 54% and success rates above 99.99%.
- Open-source communities on Hugging Face (11,849 downloads last month) and ModelScope are actively engaging with the model.
ユースケース
1. **Agentic Workflows and Tool Orchestration**: Hy3 excels at complex agent tasks requiring search, tool coordination, and multi-step planning. With scores of 84.2 on BrowseComp and 79.1 on MCP-Atlas, it's ideal for building search-heavy agents, customer service bots with tool access, and automated workflow systems. Choose Hy3 over GLM-5.2 when agentic capabilities matter more than raw coding performance.
2. **Cost-Sensitive Production Deployments**: At $0.14/$0.56 per million tokens, Hy3 delivers frontier-adjacent performance at 71x lower cost than models like Fable 5. It's optimal for high-volume applications like chatbots, document processing, and enterprise productivity tools where cost per token is critical. Self-hosting on a single 8-GPU node further reduces operational costs.
3. **Reliability-First Applications**: With a 5.4% hallucination rate and explicit training to avoid fabrication, Hy3 is suitable for applications where accuracy matters more than creativity: legal document analysis, medical information systems, financial modeling, and enterprise knowledge bases. The consistent performance across agent scaffolds (CodeBuddy, Cline, KiloCode) makes it reliable for teams using diverse development environments.
4. **Global Deployments Requiring Permissive Licensing**: The Apache 2.0 license without regional restrictions makes Hy3 the clear choice for multinational enterprises, especially those serving EU/UK/South Korean markets where GLM-5.2's license creates legal barriers. It's also ideal for open-source projects that need commercial-friendly model weights.
**When to Choose Alternatives:**
- Choose GLM-5.2 for repository-scale software engineering, complex coding tasks, and when you have 8x H200 budget available.
- Choose DeepSeek V4 Flash for maximum cost efficiency on simpler tasks and when aggressive quantization is needed.
- Choose proprietary models (GPT-5.5, Claude Opus 4.8) for frontier complex reasoning and when absolute peak performance is required regardless of cost.
最新ニュース
1. **Official Release (July 6, 2026)**: Tencent officially released Hy3, upgrading from the April preview with improved performance, reliability, and production stability. The model incorporates feedback from 50+ internal product teams.
2. **License Change**: Hy3 is now available under Apache 2.0 license, removing regional restrictions that affected previous Chinese open models and enabling global enterprise adoption.
3. **Free Availability**: Hy3 is available for free on OpenRouter and Novita AI through July 21, 2026, allowing developers to test the model without cost commitment.
4. **Platform Integration**: The model is now available on multiple global developer platforms including OpenRouter, Hermes, Kilo, Cline, OpenClaw, OpenCode, and Cherry Studio, in addition to Hugging Face and ModelScope.
5. **Production Deployment**: Hy3 is already deployed across Tencent's products (WorkBuddy, CodeBuddy, Yuanbao, Marvis, ima) with reported metrics: 54% latency drop, 47% end-to-end task duration reduction, and 99.99% success rates in agent workflows.
6. **API Pricing**: Tencent Cloud offers Hy3 at approximately 1 yuan per million input tokens and 4 yuan per million output tokens (~$0.14/$0.56 per million tokens), with cache-hit pricing at 0.25 yuan per million.
7. **Deployment Ecosystem**: Dedicated deployment recipes for vLLM and SGLang are now available, with optimization for Nvidia H20-3e GPUs (export-compliant) and standard Western GPUs (H100, H200, B200).