Moonshot AIプロプライエタリ

Kimi K2.6

このモデルを比較

Moonshot AI開発の最新高性能モデル。K2.5の改良版で更高的性能を実現。

シェア:XはてブLINE

パラメータ

1

コンテキスト長

256K

ライセンス

https://huggingface.co/moonshotai/Kimi-K2-Base/raw/main/LICENSE

リリース日

2026-04-20

日本語性能

🌐多言語対応

一般的な多言語対応モデル。基本的な日本語処理は可能だが、特化モデルには劣る。

API料金

入力料金(1Mトークンあたり)

$0.95

出力料金(1Mトークンあたり)

$

課金モード: standard

強み

    弱み

      活用例

        深度分析

        Arena Elo

        1462

        Coding sub-arena: 1514; BenchLM #24 overall, #11 verified

        SWE-Bench Pro

        58.6%

        #1 — beats GPT-5.4 (57.7%) and Claude Opus 4.6 (53.4%)

        SWE-Bench Verified

        80.2%

        Essentially tied with Claude (80.8%) and Gemini (80.6%)

        DeepSearchQA (F1)

        92.5%

        Leads Claude (91.3%) and GPT-5.4 (78.6%)

        Input Price (Official API)

        $0.60/1M

        ~8x cheaper than Claude Opus 4.6 ($5/1M)

        Architecture

        1T params (32B active)

        MoE; 256K context; open-weight (Modified MIT)

        強み

        • Best-in-class open-weight coding model — leads SWE-Bench Pro over all frontier competitors at $0.60/1M input
        • Exceptional agentic and long-horizon execution — 300 sub-agent swarms, 13-hour autonomous coding demos, strongest DeepSearchQA and HLE-with-tools scores
        • Fully open-source (Modified MIT) with native video/image input via MoonViT, deployable on 8×H100 with INT4 quantization

        弱み

        • Lags GPT-5.4 and Gemini 3.1 Pro on pure reasoning (HLE no-tools: 34.7% vs 44.4%; AIME: 96.4% vs 99.2%)
        • Ecosystem and documentation skew Chinese-first; English-language enterprise support and compliance tooling less mature than OpenAI/Anthropic
        • Verbose output (~160M tokens on AA Intelligence Index eval, 4× median) inflates cost in production and can cause downstream parsing issues

        競合比較

        ModelArenaGPQAPrice
        GPT-5.4 (xhigh)N/A92.8%~$5/~$15
        Claude Opus 4.6 (max effort)N/A91.3%$5/$25
        Gemini 3.1 Pro (thinking high)N/A94.3%N/A

        Kimi K2.6, released April 20, 2026 by Moonshot AI, is the strongest open-weight coding and agentic model available as of mid-2026. Built on a 1-trillion-parameter Mixture-of-Experts architecture that activates only 32B parameters per token, it delivers frontier-tier performance at inference costs closer to a dense 32B model. It leads all compared models — including GPT-5.4 and Claude Opus 4.6 — on SWE-Bench Pro (58.6%), the benchmark most representative of real-world software engineering, while costing roughly 8x less per input token than Claude Opus 4.6 on Moonshot's official API ($0.60/1M vs $5/1M).

        The model's design philosophy centers on practical, long-horizon execution rather than raw reasoning ceiling. It scales horizontally to 300 parallel sub-agents in its Agent Swarm architecture, has demonstrated multi-day autonomous operation in production settings, and shows strong generalization across programming languages including niche ones like Zig. On agentic benchmarks requiring tool use and multi-step search (HLE with tools: 54.0%, DeepSearchQA: 92.5%), it outperforms every proprietary competitor. The 256K context window, native multimodal input via MoonViT, and full open-source release under a Modified MIT license make it uniquely positioned for teams that need control over their infrastructure.

        The tradeoff is clear: on pure reasoning without tools — graduate-level science (GPQA Diamond: 90.5% vs 94.3% for Gemini), competition math (AIME: 96.4% vs 99.2% for GPT-5.4), and the hardest knowledge benchmarks (HLE without tools: 34.7% vs 44.4% for Gemini) — the proprietary models retain a meaningful edge. For teams building production AI systems, K2.6 is best understood as the execution engine in a multi-model stack: route structured coding, agentic workflows, and cost-sensitive batch jobs to K2.6, and reserve frontier proprietary models for high-stakes single-turn reasoning where accuracy is non-negotiable.

        分析生成日: 2026-07-17