Google Deep Mindプロプライエタリ

Gemini 3.5 Live Translate

このモデルを比較

Google DeepMind開発の最新マルチモーダルモデル。テキスト・画像・動画を統合処理。

シェア:XはてブLINE

パラメータ

非公開

コンテキスト長

ライセンス

プロプライエタリ

リリース日

2026-06-09

日本語性能

高品質日本語

多言語対応モデルのうち、日本語処理に優れた性能を持つモデル。

API料金

入力料金(1Mトークンあたり)

$3.5

出力料金(1Mトークンあたり)

$

課金モード: standard

強み

    弱み

      活用例

        深度分析

        Languages Supported

        70+

        auto-detected; 2,000+ pair combinations in one session

        First-Audio Latency

        ~2.9 seconds

        median 2,947 ms (LiveLingo benchmark, p10–p90: 2,859–3,104 ms)

        Comprehension Score

        4.93 / 5

        #1 among tested systems (LiveLingo 2026 benchmark, 120 utterances)

        Effective Cost

        ~$0.037/min

        $3.50/1M input tokens, $21.00/1M output tokens

        Input Context

        128K tokens

        audio context window; 64K token output limit

        Base Architecture

        Gemini 3 Pro

        natively multimodal audio model, not a cascaded STT→MT→TTS pipeline

        強み

        • Native audio-to-audio architecture produces the most natural-sounding translated speech, preserving speaker intonation, pacing, and pitch
        • Broadest distribution of any live translation system—rolled into Google Translate, Google Meet, and the Gemini Live API simultaneously
        • Continuous streaming translation (~3s latency) enables real conversational flow instead of turn-by-turn pauses

        弱み

        • Audio-only output with no streaming text mode, no per-speaker attribution, and no ability to edit or revise spoken output mid-utterance
        • Voice inconsistency documented by Google's own model card—voices can shift gender, get stuck, or bleed across speakers in multi-party scenarios
        • Structural blind spot on code-switched audio: when source speech switches into the target language, content silently disappears (~28% loss in benchmark testing)

        競合比較

        ModelArenaSWEGPQAPrice
        OpenAI gpt-realtime-translateN/AN/AN/A$0.06/min est.
        Google Cloud STT v2 + Translate v3N/AN/AN/Avaries per service
        Azure Speech TranslationN/AN/AN/Avaries per tier

        Gemini 3.5 Live Translate, released June 9, 2026, is Google DeepMind's specialized audio-to-audio translation model built on the Gemini 3 Pro architecture. Unlike traditional cascaded pipelines (speech-to-text → machine translation → text-to-speech), it processes audio natively—accepting 16kHz PCM chunks and outputting 24kHz translated speech in near real-time. The model auto-detects 70+ languages and preserves the speaker's vocal characteristics, producing translated audio that sounds substantially more natural than generic TTS reading a translation aloud. It ships simultaneously across Google Translate (consumer), Google Meet (enterprise), and the Gemini Live API (developer), giving Google an unmatched distribution advantage in the live translation space.

        The model represents a fundamental architectural shift in how translation is delivered. Rather than being a feature bolted onto existing products, it positions translation as an ambient capability—a layer that sits inside conversations, meetings, and apps. Google's integration with partners like Grab (10M+ monthly voice calls), Agora, LiveKit, and Pipecat signals that the long-term vision is embedded translation infrastructure, not a standalone translator app. At ~$0.037 per minute of translated conversation, the economics are viable for high-value business use cases while remaining accessible to developers through a free tier in Google AI Studio.

        However, the model card is refreshingly honest about limitations. Voice consistency degrades in multi-speaker sessions, language detection struggles with accents and similar language pairs, and the irreversible audio commitment means late-resolving syntax (common in Mandarin, Japanese) can produce factual inversions. The absence of streaming text output, per-speaker attribution, and mid-utterance revision makes it unsuitable for use cases requiring verbatim records or speaker disambiguation. This is a model optimized for fluid conversational translation—not transcription, not summarization, and not the kind of editable output that enterprise compliance workflows demand.

        分析生成日: 2026-07-17