{"schemaVersion":"1.0","type":"Article","types":["Article"],"slug":"introducing-gemini-3-8-live-and-3-8-live-extended-thinking-kd69g","url":"https://api.zyvop.com/introducing-gemini-3-8-live-and-3-8-live-extended-thinking-kd69g","title":"Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking","subtitle":"Google's most advanced live-dialogue AI yet, with background reasoning, visual grounding, and real-time multilingual voice support.","tldr":"Google launched Gemini 3.8 Live and 3.8 Live Extended Thinking on September 15, 2026, its most advanced voice AI yet. The models reason and speak simultaneously, ground responses in live video, switch between 97 languages, and top several speech benchmarks.","keywords":["Gemini","Google DeepMind","voice AI","Gemini 3.8 Live and 3.8 Live Extended Thinking","AI News"],"entities":["Arpan Singh","Gemini","Google DeepMind","voice AI","Gemini 3.8 Live and 3.8 Live Extended Thinking","AI News","ZyVOP"],"keyTakeaways":["Google announced two new voice-first AI models on September 15, 2026: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most capable live-dialogue models yet.","Both are native speech-to-speech models built on Gemini 3 Pro, meaning audio goes in and audio comes out directly instead of routing through separate speech-recognition, language, and text-to-speech systems.","The pitch is simple: a voice agent that can think, act, and keep talking, all at once."],"headings":["Two models, two jobs","What's actually new","Specs, benchmarks, and limits","Pricing and availability","Ecosystem and safety","The bigger picture"],"outboundLinks":["https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/","https://deepmind.google/models/model-cards/gemini-3-8-audio/","https://artificialanalysis.ai/speech-to-speech","https://ai.google.dev/gemini-api/docs/pricing","https://ai.google.dev/gemini-api/docs/live-api","https://aistudio.google.com/live","https://deepmind.google/models/synthid/"],"contentText":"Google announced two new voice-first AI models on September 15, 2026: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most capable live-dialogue models yet. Both are native speech-to-speech models built on Gemini 3 Pro, meaning audio goes in and audio comes out directly instead of routing through separate speech-recognition, language, and text-to-speech systems. The pitch is simple: a voice agent that can think, act, and keep talking, all at once. Two models, two jobs Google split the release into two models with clearly different jobs: Gemini 3.8 Live Gemini 3.8 Live Extended Thinking Built for Scale and cost efficiency High-complexity, multi-step tasks Strengths Fluid dialogue, visual grounding, fast responses Deeper reasoning, background tool orchestration Best for High-volume, everyday voice interactions Agentic workflows, coding, multi-step bookings What's actually new The headline feature is background reasoning: instead of going silent while it works something out, Extended Thinking acknowledges a request verbally, keeps the conversation moving while it reasons and calls tools in the background, then narrates its progress as the task completes. Live adds its own tricks: near real-time visual grounding from a camera or screen share, automatic language switching across 97 languages mid-conversation, and background tool execution that doesn't interrupt the chat. Google's demos show the models guiding onboarding with live visual context, playing chess by watching a physical board, and turning hand-drawn sketches into working React components through voice feedback. Specs, benchmarks, and limits Per the Gemini 3.8 Audio model card, both models accept audio, image, video, and text input (128,000-token context) and produce audio and text output (64,000-token ceiling). Known limitations: they can hallucinate, occasionally time out, and their knowledge cutoff is January 2025. On Google's own benchmarks, Extended Thinking takes the #1 spot on Artificial Analysis' Speech-to-Speech Quality Index at 82.6, leads agentic task completion at 68.6% on τ-Voice, and scores 97.7% on Big Bench Audio. Standard Live placed second in the Speech Agent Arena, the cheaper, higher-volume sibling of the two. Treat these as a starting point, not the final word. Independent benchmarking will follow, as it usually does after a launch like this. Pricing and availability On the standard Gemini API tier: Per 1M tokens Per minute Audio input $3.00 ~$0.005 Audio output $12.00 ~$0.018 A free tier is available to get started. (Check Google's pricing page for current rates, since introductory-pricing windows shift.) Both models are rolling out now. Developers get them via the Gemini API and Google AI Studio; enterprises get private-preview access through Gemini Enterprise; and everyday users will find Live in Search Live and Extended Thinking in the Gemini Live app, plus Docs, Gmail, and Keep for eligible subscribers. Ecosystem and safety Voice infrastructure providers like LiveKit, LangChain, and Vercel already support the Gemini Live API, and early enterprise partners include Salesforce, Genspark, and Lumeris. Every audio clip carries a SynthID watermark for provenance, and Google's frontier-safety assessment found no new capability thresholds crossed relative to Gemini 3.7 Flash. The bigger picture The real story here isn't the benchmark scores. It's the UX pattern: a model that acknowledges a request out loud and narrates its own background work solves the \"does it actually understand me\" anxiety that's dogged voice assistants for years. That's my read, not Google's: every voice-assistant builder, OpenAI included, has run into the same latency-versus-reasoning wall, and whichever one makes \"thinking out loud\" feel natural instead of gimmicky wins the next round. It's a UX bet worth watching more closely than any single benchmark number. Sources: Google's official announcement and the Gemini 3.8 Audio model card, both published September 15, 2026, plus the Gemini API pricing documentation.","contentHash":"sha256:dbc11d365b8e3c1b6ffd0bf5c6f1c78ef9da7257684a3018a50fca64c50936f3","authorName":"Arpan Singh","authorUrl":"https://api.zyvop.com/author/arpan","authorSameAs":[],"category":"AI News","tags":["Gemini","Google DeepMind","voice AI","Gemini 3.8 Live and 3.8 Live Extended Thinking"],"audience":"Readers and engineers researching AI News","tone":"Practical and evidence-based engineering guidance","readingTimeMinutes":3,"wordCount":616,"faqs":null,"primaryTopic":"AI News","publishedAt":"2026-09-16T06:02:21.398Z","updatedAt":"2026-09-16T06:06:24.324Z","canonicalUrl":"https://api.zyvop.com/introducing-gemini-3-8-live-and-3-8-live-extended-thinking-kd69g"}