Zhipu AIオープンソース

GLM 5.1

このモデルを比較

Zhipu AI開発の高性能基盤モデル。中国語対応に優れ、多様なタスクに対応。

シェア:XはてブLINE

パラメータ

754

コンテキスト長

200K

ライセンス

MIT

リリース日

2026-03-27

日本語性能

🌐多言語対応

一般的な多言語対応モデル。基本的な日本語処理は可能だが、特化モデルには劣る。

API料金

入力料金(1Mトークンあたり)

$1.4

出力料金(1Mトークンあたり)

$

課金モード: standard

強み

    弱み

      活用例

        深度分析

        Arena Elo (Text)

        1472

        #16 verified leaderboard; Coding Arena: 1524

        SWE-Bench Pro

        58.4%

        #1 globally — ahead of GPT-5.4 (57.7) & Claude Opus 4.6 (57.3)

        SWE-Bench Verified

        77.8%

        3 pts behind Claude Opus 4.6 (80.8%)

        GPQA Diamond

        86.2%

        vs Claude Opus 4.6: 91.3%, GPT-5.2: 92.4%

        Input Price

        $1.40/1M

        Output: $4.40/1M; cached: $0.26/1M

        Parameters (Active/Total)

        40B / 744B

        MoE: 256 experts, top-8 routing + 1 shared

        Context Window

        200K tokens

        Max output: 128K tokens

        強み

        • Best open-weight model for agentic coding — leads SWE-Bench Pro (58.4%) and CyberGym (68.7%) globally
        • Exceptional long-horizon autonomous execution: demonstrated 8-hour uninterrupted agentic sessions across 1,700 steps
        • Fully open under MIT license with competitive API pricing (~$1.40/$4.40 per 1M tokens), undercutting Claude Opus by 10-17×

        弱み

        • Text-only — no multimodal input (image, audio, video) unlike GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro
        • Lags frontier models on general reasoning: GPQA Diamond 86.2% vs 92.4% (GPT-5.2), HLE 31.0 vs 45.0 (Claude Opus 4.6)
        • Slower generation speed (~40-44 tok/s) and verbose output, increasing latency and effective cost for interactive use

        競合比較

        ModelArenaSWEGPQAPrice
        Claude Opus 4.6 (Anthropic)~150380.8%91.3%$15/$75
        GPT-5.4 (OpenAI)N/AN/AN/A~$12/$60 est.
        DeepSeek V3.2N/A73.1%N/A$0.27/$1.10

        GLM-5.1 is a 744B-parameter Mixture-of-Experts model developed by Zhipu AI (rebranded as Z.AI), released on April 7, 2026. It represents the current pinnacle of open-weight coding and agentic AI, achieving the highest public score on SWE-Bench Pro (58.4%) — edging past both GPT-5.4 and Claude Opus 4.6. The model was trained entirely on 100,000 Huawei Ascend 910B chips with zero NVIDIA hardware, making it a landmark achievement in China's AI self-sufficiency under US export controls. Z.AI completed a Hong Kong IPO in January 2026 at a ~$31.3B valuation, becoming the world's first publicly traded foundation model company.

        Architecture-wise, GLM-5.1 is a post-training refinement of GLM-5, sharing the same MoE backbone (744B total, 40B active per token, 256 routed experts with top-8 selection) but with significantly enhanced coding, agentic, and long-horizon capabilities. Its defining innovation is extended autonomous execution: the model can sustain a plan-execute-analyze-optimize loop for up to 8 hours without human checkpoints, demonstrated by building a complete Linux desktop environment through 655 iterations and optimizing a vector database to 6.9× its original throughput. The model ships under the MIT license with weights available on Hugging Face.

        Where GLM-5.1 falls short is general-purpose reasoning and multimodal tasks. It trails frontier closed-source models by 5-12 points on GPQA Diamond (86.2%), Humanity's Last Exam (31.0 base), and AIME 2026 (95.3%). It is text-only, with no image or audio processing. Generation speed (~40-44 tok/s) is below the competitive median, and the model's verbose output style inflates token usage. For organizations building autonomous coding agents or long-running engineering pipelines, GLM-5.1 is the strongest open-weight option available; for multimodal workflows or maximum reasoning capability, closed-source alternatives remain superior.

        分析生成日: 2026-07-17