🔥 Trending on HN

Gemini 3.8 Live: An AI That Keeps Talking While It Works

3 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
visual grounding(visual grounding)

Using visual input as part of the model's conversation context.

private previews(private previews)

Limited releases that let selected users test a service.

SynthID(SynthID)

A hidden watermark Google says it adds to AI-generated audio.

Google, the company behind Gemini, announced Gemini 3.8 Live, a voice conversation model, and Gemini 3.8 Live Extended Thinking, a version for complex tasks, on September 15, 2026. The announcement is in Google's official blog. Google says the models can reason near real time, understand visual input, and run tools while a conversation continues. The announcement drew 350 points and 224 comments on Hacker News, a technology news site. Those numbers show community attention. They do not prove the claims or the models' quality.

What Google announced

Live is built for scale and cost efficiency. Google says it combines smooth dialogue with conversational intelligence. It can process visual input almost in real time. Google calls this visual grounding: using what the model sees as conversation context. It can detect and switch among 97 supported languages during one conversation. While it calls tools and APIs, it can keep talking with the user.

Extended Thinking is built for high-complexity tasks. It works through multiple reasoning steps while speaking. It can give short cues, such as saying it will check something, and narrate progress. Google shows examples with bookings and asynchronous function calls. These examples show the intended experience, not proof that every task will work reliably.

Background: two ways to handle live voice

The release focuses on more than smarter answers. It focuses on keeping the conversation active while work happens behind the scenes. Live emphasizes speed, fluid dialogue, and efficiency. Extended Thinking emphasizes complex workflows and deeper reasoning. Google also presents both models as building blocks for voice agents.

Why it matters

Voice use combines listening, seeing, speaking, and service calls in one flow. If an AI can acknowledge a request, work in the background, and report progress, the user can keep the conversation moving. That is the promise behind the release. It also makes signaling important: a smooth voice should not make an unfinished task sound complete.

Google reports outside evaluation results for Extended Thinking: 82.6 for speech-to-speech quality, 68.6% for voice-agent task completion, and 97.7% for audio reasoning. It also places Live second in another speech-agent evaluation. These figures come from Google's announcement. The post does not, by itself, provide independent verification or full test conditions.

What is confirmed

For developers, both models are starting in the Gemini API, Google's developer interface, and Google AI Studio, a developer workspace. For companies, both have private previews in Gemini Enterprise. Live is also starting in Search Live, Google's voice search experience. Extended Thinking is rolling out in Gemini Live. Google lists Docs, Gmail, and Keep in Workspace for eligible subscribers.

Google also says all audio generated by its AI products carries an imperceptible SynthID watermark. It is intended to help identify AI-generated audio and prevent misinformation.

What remains unknown

The post does not give a full price table. It does not spell out availability by country, device, or plan. It also does not explain language switching in noisy conditions, what happens after a failed tool call, or what confirmation controls users get. Benchmark conditions and comparison sets are not fully detailed either.

What to watch next

Watch whether conversations stay natural while tools run. Watch task time, API cost, and visual accuracy. Independent repetitions of the benchmarks would make the claims easier to judge. For everyday users, the key signal will be the rollout across Search Live, Gemini Live, and Workspace. Until then, Gemini 3.8 Live shows a clear direction for voice AI, but its impact still depends on testing.

💬 Gemini 3.8 Live: Natural to use, but hard to evaluate cleanly

HN commenters praise Gemini 3.8 Live for speed, readable conversation, and multilingual voice use, while others point to the chess demo, reported hallucinations, lost context, voice fatigue, and the confounding effect of search harnesses. These are user reports and opinions, not independent verification.

  • Some commenters thought the chess demo looked bad because the model appeared to miss a very common checkmate pattern, weakening the impression of an advanced system.
  • Others argued that a live LLM completing a legal game without inventing pieces or using guided decoding is itself a meaningful achievement.
  • Users report that Gemini feels relatively readable, quick, and enjoyable for open-ended discussion and brainstorming. Opinions are more divided on coding, while some commenters value its multilingual speech and translation.
  • One user report says Gemini Live can identify different speakers and translate during multilingual conversations. Another reports that overly expressive voice intonation becomes tiring, showing that naturalness is not universally preferred.
  • A user who says they used Pro/Thinking reported hallucinations on research questions. This suggests that pleasant conversation and reliable fact-finding are separate properties.
  • Another user report describes Gemini suddenly losing earlier context and derailing a conversation, with a harsher failure mode than gradual context compression in another model.
  • A counterargument is that comparisons through Google's app may partly measure the search harness and allocated computation, not just the underlying model.
  • Taken together, the thread supports separating conversational usability from reliability on difficult reasoning, research, coding, and long-running interactions.

mature digest at 327 comments (revision 4). We fetched 300 comments and sampled 120 across the thread. These are HN users’ reports, not independently verified facts.

🔥 Trending on HN

Gemini 3.8 Live lets AI talk, see, and work at once

📰 Full story: Gemini 3.8 Live: An AI That Keeps Talking While It Works

Google has announced two voice models that can keep talking while they handle tasks.

2 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
API(API)

A way for software to use a model or service.

private previews(private previews)

Limited releases for testing.

SynthID(SynthID)

A hidden marker for audio made by AI.

💡 The gist

  • Live is built for quick, smooth voice conversations.
  • Extended Thinking handles harder, multi-step tasks.
  • Both can use visual input and tools during a conversation.

Google, the company behind Gemini, announced the models on September 15, 2026. Read the official announcement. Live can process visual input almost in real time. Google says it can detect and switch among 97 supported languages. It can also run tools and API calls in the background. The conversation can continue while those tasks finish.

Extended Thinking is designed for complex work. It can reason through several steps while speaking. It may give short updates, such as saying it will check something. Google shows examples with bookings and several calls made at different times. This matters because users can keep talking while the AI works. The AI can acknowledge the request and explain its progress.

Google reports strong results on several outside benchmarks. The announcement lists 82.6 for speech-to-speech quality. It lists 68.6% for voice-agent task completion. It lists 97.7% for audio reasoning. Live is listed second in a speech-agent evaluation. These are results reported by Google. They are not independent proof. Test rules and comparisons are not fully explained.

The models are rolling out in stages. Developers can start with the Gemini API, a way for software to use Gemini. They can also use Google AI Studio. Enterprise users can see private previews in Gemini Enterprise. Live is also starting in Search Live, a voice search feature. Extended Thinking is coming to Gemini Live. Google lists Docs, Gmail, and Keep for eligible Workspace subscribers.

Google says its AI-generated audio has a SynthID watermark. The mark is hard to notice. It should help identify AI-made audio.

The announcement drew attention on Hacker News, a technology news site. It received 350 points and 224 comments. Those numbers measure attention. They do not confirm the models' quality.

Important details remain open. The announcement does not give full prices or every country and device. It does not show how language switching works in noisy places. It also does not explain what happens after a failed tool call. The next test is real use. Watch task speed, API cost, visual accuracy, and independent benchmark checks.

💬 Gemini 3.8 Live: Nice to talk to, but not perfect

Some commenters find Gemini 3.8 Live fast, clear, and good with languages. Others report mistakes in chess, research, and long chats. These are personal reports, not official test results.

  • The chess demo looked weak to some people because it seemed to miss an easy checkmate.
  • Others said that playing a full legal game is impressive for a live language model.
  • Users say it is pleasant for brainstorming, and some say it handles multilingual speech and translation well.
  • Other users report made-up research answers and losing track of earlier conversation.
  • Some people also find the voice too expressive and tiring.
  • Comparisons may be affected by Google's search system and computing budget, not only by the model itself.

mature digest at 327 comments (revision 4). We fetched 300 comments and sampled 120 across the thread. These are HN users’ reports, not independently verified facts.

🔥 Trending on HN

A talking Gemini that can keep working

📰 Full story: Gemini 3.8 Live: An AI That Keeps Talking While It Works

Google announced a voice AI that keeps talking while it works.

1 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
Hacker News(Hacker News)

A website that shares technology news.

Gemini API(Gemini API)

A way for software to use Gemini.

SynthID(SynthID)

A hidden mark in AI-made sound.

Google, the company behind Gemini, announced a new voice AI. Read the announcement. Its name is Gemini 3.8 Live, a voice model. It can talk with you. It can look at pictures and other things you show it. Google says it can switch among 97 languages. It can also ask tools to do jobs. The talking can continue while jobs finish.

Google also announced Gemini 3.8 Live Extended Thinking, the version for harder jobs. It works through several steps. It can tell you short updates while it thinks. Google showed an example with bookings.

The news appeared on Hacker News, a technology news site. It got 350 points and 224 comments. These numbers show attention. They do not show that the AI is correct.

Google says the models will arrive in stages. Developers can use the Gemini API, a way for software to use Gemini. Some Google apps will get them too. Google also says AI-made audio has a SynthID watermark, a hidden mark in the sound. This mark may help people spot AI-made sound.

We still do not know every price. We do not know every device. We do not know how well it works everywhere.

💬 What people say about Gemini 3.8 Live

People noticed good things and bad things.

  • It can talk quickly and use different languages, some people say.
  • But it seemed to make an easy chess mistake.
  • Some people say it forgets old parts of a long talk.
  • Some people also say its voice is too dramatic and tiring.

mature digest at 327 comments (revision 4). We fetched 300 comments and sampled 120 across the thread. These are HN users’ reports, not independently verified facts.

Sources