Anthropicプロプライエタリ

Claude Sonnet 5

このモデルを比較

Anthropic開発の次世代基盤モデル。最新の訓練データとアーキテクチャで高い性能を実現。

シェア:XはてブLINE

パラメータ

非公開

コンテキスト長

1M

ライセンス

プロプライエタリ

リリース日

2026-06-30

日本語性能

高品質日本語

多言語対応モデルのうち、日本語処理に優れた性能を持つモデル。

API料金

入力料金(1Mトークンあたり)

$2

出力料金(1Mトークンあたり)

$

課金モード: standard

強み

    弱み

      活用例

        深度分析

        Artificial Analysis Intelligence Index

        53

        #5 overall, 2-3 points behind GPT-5.5 (xhigh) and Opus 4.8 (max)

        SWE-bench Pro

        63.2%

        vs Sonnet 4.6: 58.1%, Opus 4.8: 69.2%

        Terminal-Bench 2.1

        80.4%

        Beats Opus 4.8 (74.6%); largest single improvement over 4.6

        OSWorld-Verified

        81.2%

        vs Sonnet 4.6: 78.5%, Opus 4.8: 83.4%

        Standard Input Price

        $3/1M tokens

        Intro $2/1M through Aug 31; vs Opus $5/1M

        Standard Output Price

        $15/1M tokens

        Intro $10/1M through Aug 31; vs Opus $25/1M

        強み

        • Near-Opus agentic performance at roughly half the cost; beats Opus 4.8 on Terminal-Bench 2.1 and ties on knowledge work (GDPval-AA v2)
        • Strongest prompt-injection robustness of any model tested (0.19% attack success rate), with lowest MASK lying rate (3.1%)
        • Best-in-class agentic follow-through: runs tests, reads failures, patches code, and reruns without human prompting

        弱み

        • Updated tokenizer inflates token counts 1.0-1.35x for identical text, making real-world costs higher than headline per-token pricing suggests
        • Opus 4.8 retains clear leads on hardest tasks: +17.2 points on USAMO olympiad math, +6.0 on SWE-bench Pro, +6.6 on HLE no-tools
        • Intentionally low cybersecurity capability; defaults to strong cyber safeguards unsuitable for legitimate security research without exemptions

        競合比較

        ModelArenaPrice
        Claude Opus 4.8N/A$5/$25 per 1M
        GPT-5.5N/A$5/$30 per 1M
        Gemini 3.5 FlashN/APer Google pricing

        Claude Sonnet 5, released June 30, 2026, is Anthropic's most agentic mid-tier model to date, closing much of the performance gap to the flagship Opus 4.8 while costing roughly half as much per token. It scored 63.2% on SWE-bench Pro (up 5.1 points from Sonnet 4.6), 80.4% on Terminal-Bench 2.1 (actually beating Opus 4.8's 74.6%), and tied Opus 4.8 on knowledge work with a 1618 GDPval-AA v2 Elo. The model uses an updated tokenizer that can inflate token counts by up to 1.35x, which Anthropic offset with introductory pricing of $2/$10 per million tokens through August 31, 2026, after which standard pricing of $3/$15 applies.

        The model's defining characteristic is reliable agentic follow-through—early users report it finishing multi-step coding, debugging, and business automation tasks where previous Sonnets would stall or require human nudging. It ships with five effort levels (low through max), adaptive thinking that defaults to high, and a 1-million-token context window. Safety improvements are significant: the model achieved the lowest prompt-injection attack success rate (0.19%) and lowest sycophantic lying rate (3.1%) of any tested model.

        Sonnet 5 is now the default model on Claude's Free and Pro plans and is available across Claude Code, the Claude API, AWS Bedrock, Google Vertex AI, and Microsoft Foundry, plus third-party tools like Cursor, VS Code, and GitHub Copilot. For the vast majority of production agentic workloads, it represents the new best value in the Claude lineup, with Opus 4.8 now reserved for the hardest accuracy-critical tasks and cyber-research use cases.

        分析生成日: 2026-07-17