Close Menu
Techy101 –

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Is Parking A Gas Car In An EV Charging Spot Illegal? It’s Complicated

    September 24, 2026

    Amazon reduziert mehrere Segway E-Scooter

    September 24, 2026

    Qualcomm goes official with Snapdragon X laptop Linux support

    September 23, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Is Parking A Gas Car In An EV Charging Spot Illegal? It’s Complicated
    • Amazon reduziert mehrere Segway E-Scooter
    • Qualcomm goes official with Snapdragon X laptop Linux support
    • Discord Begins Rolling Out Age Verification–And A Sponsorship With FanDuel
    • This popular video editing app is finally coming to Android
    • iPhone 18 Pro vs. iPhone Duo: Comparing Apple’s New Pro Phone Against Its First Foldable
    • The Xiaomi 18 Pro and Pro Max are Now Official, and Come with Qualcomm’s Powerful New Processors
    • I tried the Beats 360 headphones ahead of launch – here’s what I thought
    Facebook X (Twitter) Instagram Pinterest YouTube LinkedIn TikTok
    Techy101 –Techy101 –
    • Home
    • Laptops
    • Mobiles
    • Gaming
    • Gadgets
    • Apps
    • AI
    • How To
    • Reviews
    Techy101 –
    Home»AI»Google Launches Gemini 3.8 Live and Extended Thinking Voice Models – Unite.AI
    AI

    Google Launches Gemini 3.8 Live and Extended Thinking Voice Models – Unite.AI

    By RepublisherSeptember 16, 2026No Comments5 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    Google Launches Gemini 3.8 Live and Extended Thinking Voice Models – Unite.AI
    Share
    Facebook Twitter LinkedIn Pinterest Email



    Google announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026, a pair of live dialogue models rolling out across the Gemini API, Google AI Studio, Gemini Enterprise, Search Live, Gemini Live, and Google Workspace.

    The models were introduced by Tom Ouyang, a principal engineer, and Malini Jaganathan, a member of technical staff, writing on behalf of the Gemini Audio Team. Google describes the pair as its most advanced live dialogue models yet, built for natural conversation, and says gains in intelligence and parallel reasoning make them more intuitive to collaborate with on complex tasks executed by voice. The company positions Gemini 3.8 Live for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding, while Gemini 3.8 Live Extended Thinking is built for high-complexity tasks that call for increased intelligence and multi-step reasoning. According to Google, the models give developers and enterprises the building blocks for reliable, production-ready voice agents while making conversations with Gemini across the Gemini app, Workspace, and Search more fluid.

    Reported Benchmark Results

    Google reports that Gemini 3.8 Live Extended Thinking captured the top overall position on Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6, and that it leads in agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark. The company also reports a 97.7% score on Big Bench Audio and says Extended Thinking maintains a highly competitive price point relative to other frontier models. Google says Gemini 3.8 Live placed second in Artificial Analysis’ Speech Agent Arena, citing high preference among users, and describes it as highly cost-effective and built for scale. On ServiceNow’s EVA-Bench, a benchmark for evaluating voice agents, Google says its models push the Pareto Frontier for complex workflows by successfully balancing accuracy with conversational quality; the post notes this evaluation was run on the Live API on Gemini Enterprise Agent Platform.

    Real-Time Conversation Capabilities

    According to Google, Gemini 3.8 Live processes visual inputs in near real time, adding context the company says produces more helpful responses. The model automatically detects and transitions between 97 supported languages mid-conversation, and it executes tools and API calls in the background while the conversation continues, acknowledging requests and keeping the dialogue going while tasks finish. For tasks that require deeper reasoning, Google says Extended Thinking reasons and speaks simultaneously, using early verbal cues to acknowledge prompts naturally and live progress narration to walk users through multi-step background tasks as they advance, without breaking the conversational flow.

    Demonstration videos published with the announcement show Gemini 3.8 Live guiding an employee onboarding session in real time using visual context to answer live questions, playing chess in near real time, building complete business plans and custom marketing toolkits through natural speech, and powering step-by-step troubleshooting help inside Search Live. Extended Thinking demonstrations show the model turning raw sketches and near-real-time voice feedback into functional React components, coordinating multi-step restaurant reservations through asynchronous function calls without interrupting the conversation, and working across Docs Live, Gmail Live, and Keep Live in Google Workspace.

    Live API and Developer Ecosystem

    Google names Agora, Fishjam, LiveKit, Pipecat, Vercel, and Vision Agents as developer platforms that use the Gemini Live API to let developers build and deploy voice-driven interfaces while the platforms manage the underlying real-time media streaming infrastructure. The company says it is also partnering with Salesforce, Genspark, and Lumeris on the new models, pointing to their interest in the models’ latency, fluidity, and tool-calling capabilities.

    The official Live API documentation describes the interface as enabling low-latency, real-time voice and vision interactions with Gemini, processing continuous streams of audio, images, and text to deliver immediate spoken responses over a stateful WebSocket connection. Documented inputs are raw 16-bit PCM audio at 16kHz, JPEG images at up to one frame per second, and text, with audio output as raw 16-bit PCM at 24kHz. Developers can choose a server-to-server implementation, in which a backend forwards client stream data to the Live API, or a client-to-server implementation, in which frontend code connects directly over WebSockets. The documentation lists use cases spanning retail shopping assistants and support agents, interactive game characters, voice- and video-enabled experiences in robotics, smart glasses, and vehicles, health companions, financial advisory tools, education mentors, real-time translation, and live transcription and captioning.

    Google states that all audio generated by its AI products is watermarked with SynthID, an imperceptible watermark woven directly into the audio output so that AI-generated content remains detectable to help prevent misinformation. The post points readers to the model card for details on the company’s safety and responsibility approach.

    Availability and Rollout

    Both models began rolling out on September 15. For developers, they are available in the Gemini API and Google AI Studio, and for enterprises they are in private preview in Gemini Enterprise. Gemini 3.8 Live is available to everyone in Search Live, while Extended Thinking is available in Gemini Live, in Docs for Google AI Pro and Ultra subscribers through Workspace, and in Gmail and Keep for all Google AI subscribers. Google says Gemini Enterprise for Customer Experience support for both models is coming soon, and that Extended Thinking is coming soon to Google Workspace business customers.



    Source link

    Extended Gemini Google launches live models Thinking Unite.AI Voice
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleGoogle breathes new life into older Pixel Watches with Gemini Personalization, gesture expansions, and more
    Next Article vivo starts teasing the V80, promises it will deliver “cinematic moments”
    Republisher
    • Website

    Related Posts

    Laptops

    ChatGPT Voice can now check your email, manage your calendar, search Slack, and use GPT-6

    September 23, 2026
    Apps

    Discord launches controversial age-verification system “using the most privacy-preserving approach we could build” months after first delay

    September 23, 2026
    Mobiles

    Gemini just got a big upgrade with new connected apps for work, creativity, and life

    September 23, 2026
    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    Is Parking A Gas Car In An EV Charging Spot Illegal? It’s Complicated

    September 24, 2026

    AMD is apparently gearing up to raise GPU prices right after Nvidia’s steep hike

    August 1, 2026

    LanceDB Vector Database Guide: Features anndPython Demo

    August 1, 2026
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Latest Reviews

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Latest Post

    Is Parking A Gas Car In An EV Charging Spot Illegal? It’s Complicated

    September 24, 2026

    AMD is apparently gearing up to raise GPU prices right after Nvidia’s steep hike

    August 1, 2026

    LanceDB Vector Database Guide: Features anndPython Demo

    August 1, 2026
    Recent Posts
    • Is Parking A Gas Car In An EV Charging Spot Illegal? It’s Complicated
    • Amazon reduziert mehrere Segway E-Scooter
    • Qualcomm goes official with Snapdragon X laptop Linux support
    • Discord Begins Rolling Out Age Verification–And A Sponsorship With FanDuel
    • This popular video editing app is finally coming to Android

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram Pinterest YouTube LinkedIn TikTok
    • About Us
    • Contact Us
    • Privacy Policy
    • Terms & Conditions
    • Disclaimer
    © 2026 techy101. Designed by Pro.

    Type above and press Enter to search. Press Esc to cancel.