DeepSeekオープンソース

DeepSeek V4 Pro

このモデルを比較

DeepSeek開発の高性能MoEコーディングモデル。1.6Tパラメータ(49B活性化)、100万トークンコンテキスト対応。CSA+HCAハイブリッドアーキテクチャ採用。

シェア:XはてブLINE

パラメータ

1.6

コンテキスト長

1M

ライセンス

MIT

リリース日

2026-04-24

日本語性能

高品質日本語

多言語対応モデルのうち、日本語処理に優れた性能を持つモデル。

API料金

入力料金(1Mトークンあたり)

$0.435

出力料金(1Mトークンあたり)

$

課金モード: standard

強み

    弱み

      活用例

        深度分析

        Arena Elo

        1462

        Upper tier among all models (Swfte.com)

        SWE-Bench Verified

        80.6%

        Ties Claude Opus 4.6 (80.8%) and Gemini 3.1 Pro (80.6%)

        LiveCodeBench Pass@1

        93.5%

        #1 among all benchmarked models

        GPQA Diamond

        90.1%

        Behind Gemini 3.1 Pro (94.3%) and GPT-5.4 (93.0%)

        Output Price

        $3.48/1M

        ~7-8x cheaper than Claude Opus 4.7 ($25) and GPT-5.5 ($30)

        Context Window

        1M tokens

        Hybrid CSA+HCA attention; 10% KV cache vs V3.2 at 1M ctx

        Codeforces Rating

        3206

        Leads all models; GPT-5.4 at 3168, Gemini at 3052

        Artificial Analysis Intelligence Index

        52

        #2 open weights model, behind Kimi K2.6 (54)

        強み

        • Best-in-class coding benchmarks among open models: #1 LiveCodeBench (93.5%), #1 Codeforces (3206), tied SWE-Bench Verified (80.6%)
        • Exceptional cost efficiency at $3.48/M output (~7-8x cheaper than closed frontier models with comparable quality on most tasks)
        • MIT-licensed open weights with 1M-token context window, agentic tool-use support (MCPAtlas 73.6%), and dual Thinking/Non-Thinking modes

        弱み

        • Significant factual knowledge gaps: HLE 37.7% (vs Gemini's 44.4%), SimpleQA 57.9% (vs Gemini's 75.6%), and 94% hallucination rate on AA-Omniscience
        • High latency (~28s average per response vs 5.8s for Claude Opus 4.7) and context quality degradation past ~800K tokens
        • Political censorship embedded in training weights; regulatory bans in Italy, Denmark, Australia, South Korea, and multiple US states due to data collection practices

        競合比較

        ModelArenaSWEGPQAPrice
        Claude Opus 4.6N/A80.8%91.3%$5.00/$25.00
        GPT-5.4N/An/r93.0%$2.50/$15.00
        Gemini 3.1 ProN/A80.6%94.3%$2.00/$12.00

        DeepSeek V4 Pro is a 1.6-trillion-parameter Mixture-of-Experts model (49B active per token) released on April 24, 2026, as part of a two-tier open-weight lineup alongside V4 Flash. It introduces a hybrid Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) architecture that dramatically improves long-context efficiency — requiring only 27% of V3.2's inference FLOPs and 10% of its KV cache at the full 1-million-token context window. Pre-trained on 32T+ tokens with a comprehensive post-training pipeline featuring domain-specific expert cultivation followed by unified consolidation, V4 Pro ships under the MIT license with weights on HuggingFace and ModelScope.

        The model's strongest performance is in coding and competitive programming: it leads all public models on LiveCodeBench (93.5% Pass@1), Codeforces (3206 rating), and ties Claude Opus 4.6 and Gemini 3.1 Pro on SWE-Bench Verified (80.6%). On the Artificial Analysis Intelligence Index it scores 52, making it the #2 open-weight reasoning model behind only Kimi K2.6. Its GDPval-AA agentic score of 1554 leads all open-weight models, and its MCPAtlas Public score of 73.6 nearly matches Claude Opus 4.6's 73.8. However, it shows meaningful gaps in factual knowledge (HLE 37.7%, SimpleQA 57.9%) and high hallucination rates (94% on AA-Omniscience), reflecting training emphasis on code and math over broad world knowledge.

        Pricing is the model's most disruptive feature. At $1.74/M input (cache miss) and $3.48/M output — roughly 7-8x cheaper than Claude Opus 4.7 and GPT-5.5 on output tokens — V4 Pro collapses the price floor for frontier-adjacent quality. The companion V4 Flash model ($0.14/$0.28) provides 85-95% of V4 Pro's quality at 12x lower cost for most workloads. Combined with automatic context caching and a 50% off-peak discount, DeepSeek has created the most cost-efficient frontier-adjacent model stack available. The trade-offs are real — higher latency, political censorship in weights, regulatory concerns, and knowledge gaps — but for teams building production coding agents and extraction pipelines, V4 Pro represents the strongest quality-per-dollar proposition in the market as of July 2026.

        分析生成日: 2026-07-17