アリババプロプライエタリ

Qwen3.6-Max-Preview

このモデルを比較

アリババ開発の高性能MoEモデル。多言語対応と高い推論能力を特徴とする。

シェア:XはてブLINE

パラメータ

1

コンテキスト長

262K

ライセンス

プロプライエタリ

リリース日

2026-04-20

日本語性能

高品質日本語

多言語対応モデルのうち、日本語処理に優れた性能を持つモデル。

API料金

入力料金(1Mトークンあたり)

$1.3

出力料金(1Mトークンあたり)

$

課金モード: standard

強み

    弱み

      活用例

        深度分析

        AA Intelligence Index

        52

        #3 globally, behind GPT-5.4 and Claude Opus 4.7

        SWE-bench Pro

        #1

        ~58.4%, top among all evaluated models

        Terminal-Bench 2.0

        65.4%

        Tied with Claude Opus 4.6

        GPQA Diamond

        88.8%

        Strong but trails Gemini 3.1 Pro (94.3%)

        Input Price

        $1.30/1M

        ~$1.04 via OpenRouter (20% discount)

        Context Window

        262K tokens

        vs Claude Opus 4.7 and GPT-5.4 at 1M

        強み

        • Leads six agentic coding benchmarks simultaneously including SWE-bench Pro, SciCode (+10.8 over Plus), and SkillsBench (+9.9 over Plus)
        • preserve_thinking feature carries reasoning traces across multi-turn agentic workflows, reducing context loss in complex tool-calling chains
        • API compatible with both OpenAI and Anthropic SDK formats — drop-in substitution with a single model string change

        弱み

        • First closed-weight Qwen flagship breaks three years of open-source tradition; no self-hosting or fine-tuning possible
        • Slow output speed (~38–45 tok/s) falls below the median of 62 t/s for comparable reasoning models, creating latency challenges
        • 262K context window is 4× smaller than Claude Opus 4.7 and GPT-5.4 (both 1M), limiting large-codebase and long-document workflows

        競合比較

        ModelArenaSWEGPQAPrice
        Claude Opus 4.6N/A80.8%~90%$15/$75
        GPT-5.4N/A88.7%~92%$2.50/$10
        Qwen3.6-PlusN/A78.8%88.2%$0.29/$1.65

        Qwen3.6-Max-Preview is Alibaba's first proprietary, closed-weight flagship model, released April 20, 2026. Built on a sparse Mixture-of-Experts architecture with an estimated ~1 trillion total parameters, it targets agentic coding, scientific programming, and multi-step tool-calling workflows. The model topped six coding benchmarks at launch — SWE-bench Pro, Terminal-Bench 2.0, SkillsBench, SciCode, QwenClawBench, and QwenWebBench — and ranks #3 globally on the Artificial Analysis Intelligence Index with a score of 52. Its signature innovation is the preserve_thinking feature, which retains the model's chain-of-thought reasoning across conversation turns, addressing a critical failure mode in multi-step agent loops where context coherence degrades over sequential tool calls.

        The closed-weight decision marks a significant strategic pivot for Alibaba, which built its global developer reputation on open-source releases across the entire Qwen family. Two open-weight siblings — Qwen3.6-27B and Qwen3.6-35B-A3B — were released the same week under Apache 2.0, signaling a deliberate tiering strategy: open models for the community, proprietary flagship for monetization. This positions Max-Preview directly against GPT-5.4 and Claude Opus 4.7 at the frontier API tier, while lower-cost options like Qwen3.6-Plus ($0.29/$1.65 per 1M tokens) continue serving cost-sensitive workloads.

        The model arrives as a 'preview' with explicit acknowledgment from Alibaba that development is ongoing. While its agentic coding scores are genuinely competitive — particularly on SWE-bench Pro and SciCode — it trails Claude Opus 4.6 on SWE-bench Verified (80.8% vs ~73%) and GPT-5.4 on composite benchmarks (BenchLM 89 vs 81). The 262K context window, text-only modality, and below-median output speed are practical constraints. Pricing at $1.30/$7.80 per 1M tokens sits between the budget Plus tier and Western frontier pricing, making it a mid-range option for teams optimizing agentic coding cost-per-quality.

        分析生成日: 2026-07-17