Anthropicプロプライエタリ

Claude Opus 4.7

このモデルを比較

Anthropic開発の高性能基盤モデル。幅広いタスクに対応し、バランスの取れた性能を提供。

シェア:XはてブLINE

パラメータ

非公開

コンテキスト長

ライセンス

プロプライエタリ

リリース日

2026-04-16

日本語性能

高品質日本語

多言語対応モデルのうち、日本語処理に優れた性能を持つモデル。

API料金

入力料金(1Mトークンあたり)

$2.5

出力料金(1Mトークンあたり)

$

課金モード: standard

強み

    弱み

      活用例

        深度分析

        LMArena Elo

        1492

        WebDev Elo: 1560; Vision Elo: 1304

        SWE-bench Verified

        87.6%

        Up from 80.8% on Opus 4.6

        SWE-bench Pro

        64.3%

        vs GPT-5.4: 57.7%

        GPQA Diamond

        94.2%

        Near parity with GPT-5.4 (94.4%)

        Input/Output Price

        $5 / $25 per 1M tokens

        Same rate card as Opus 4.6; new tokenizer inflates 10–35%

        Context Window

        1M input / 128K output

        Max image resolution: 3.75MP (2,576px long edge)

        強み

        • Industry-leading agentic coding: 87.6% SWE-bench Verified, 64.3% SWE-bench Pro — best-in-class among broadly available models
        • 3.3× higher image resolution (3.75MP) unlocks dense screenshot analysis, diagram extraction, and computer-use agents
        • Strongest multi-tool orchestration at 77.3% MCP-Atlas (+9 points over GPT-5.4), ideal for complex agent pipelines
        • Self-verification before reporting and literal instruction following improve reliability on long-running autonomous tasks

        弱み

        • Severe long-context retrieval regression: MRCR v2 8-needle at 1M tokens drops from 78.3% (Opus 4.6) to 32.2%
        • New tokenizer inflates real-world costs 10–35% despite unchanged per-token rate card
        • BrowseComp regressed to 79.3% from 84.0%; trails GPT-5.4 (89.3%) and Gemini 3.1 Pro (85.9%) on web research tasks

        競合比較

        ModelArenaSWEGPQAPrice
        Claude Opus 4.7149287.6%94.2%$5/$25 per 1M tokens
        GPT-5.4N/A from sources84.1%94.4%Not reported in sources
        Gemini 3.1 ProN/A from sources80.6%94.3%Not reported in sources

        Claude Opus 4.7, released April 16, 2026, is Anthropic's most capable broadly available model and a direct upgrade to Opus 4.6. It targets agentic software engineering, high-fidelity vision tasks, and long-running autonomous workflows. The model introduces self-verification behavior (it checks its own outputs before reporting back), literal instruction following, higher-resolution image processing (3.75MP, up from 1.15MP), and improved file-system memory for multi-session agent work. Pricing is unchanged at $5/$25 per million input/output tokens, though a new tokenizer inflates effective costs by 10–35% depending on content type.

        Opus 4.7 sets new benchmarks for Anthropic on SWE-bench Verified (87.6%), SWE-bench Pro (64.3%), and MCP-Atlas (77.3%), with particularly strong gains on the hardest coding problems. Partner testimonials from Cursor, Replit, Notion, Vercel, and others confirm real-world production improvements of 10–15% in task success rates and 3× more production task resolution on some benchmarks. The model introduces a new 'xhigh' effort level between 'high' and 'max', and ships with task budgets (public beta) for cost control on long-running agent loops.

        However, the release comes with notable trade-offs. Long-context multi-needle retrieval collapsed on MRCR v2 (78.3% → 32.2% at 1M tokens), BrowseComp regressed nearly 5 points, and the new tokenizer raises actual costs despite unchanged pricing. Anthropic positions Opus 4.7 below its unreleased Claude Mythos Preview, and describes it as the first broadly released model carrying cybersecurity safeguards from Project Glasswing — including automated blocking of high-risk cyber requests. The system card rates it 'largely well-aligned and trustworthy, though not fully ideal.'

        分析生成日: 2026-07-17