gemini-live-api-dev
Use this skill when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live.
Install / Use
npx skills add google-gemini/gemini-skills --skill gemini-live-api-devInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
OtherSupported Platforms
Our assessment of gemini-live-api-dev
gemini-live-api-dev scores 95/100 on our quality scale, 6th of 115 Other skills we index (top 6%).
Its SKILL.md is 18 KB long, well organised into 45 sections with 15 code examples: a thorough specification that gives an agent plenty to work with.
With 4,205 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 5 days ago, so gemini-live-api-dev is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
gemini-live-api-dev compared with similar skills
All 4 of these similar skills score higher than gemini-live-api-dev; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| gemini-live-api-dev (this skill)by google-gemini | 95 | 4.2k | 5d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 86.0k | 13d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.0k | today | CLAUDE.md |
| rufloby ruvnet | 100 | 73.4k | today | CLAUDE.md |
| crawl4aiby unclecode | 100 | 84.4k | 3d ago | MCP Server |
Frequently asked questions
- How do I install gemini-live-api-dev?
- Run
npx skills add google-gemini/gemini-skills --skill gemini-live-api-dev. The install tabs above show the steps for each supported agent. - Which AI agents does gemini-live-api-dev work with?
- It is written for Gemini CLI, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is gemini-live-api-dev safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is gemini-live-api-dev still maintained?
- The repository was last updated 5 days ago, so gemini-live-api-dev is actively maintained.
Skill content
View source on GitHubname: gemini-live-api-dev description: Use this skill when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live. Covers WebSocket-based audio/video/text streaming, voice activity detection (VAD), background reasoning (extended thinking), asynchronous function calling, session management, ephemeral tokens, live transcription, and live translation. SDKs covered - google-genai (Python), @google/genai (JavaScript/TypeScript).
Gemini Live API Development Skill
Overview
The Live API enables low-latency, real-time voice and video interactions with Gemini over WebSockets. It processes continuous streams of audio, video, or text to deliver immediate, human-like spoken responses and background reasoning.
Key capabilities:
- Bidirectional audio streaming — real-time mic-to-speaker conversations
- Background reasoning (extended thinking) — multi-step background reasoning with spoken conversational fillers
- Live streaming transcription — real-time speech-to-text with interim and finalized streams
- Video streaming — send camera/screen frames alongside audio
- Text input/output — send and receive text within a live session
- Audio transcriptions — get text transcripts of both input and output audio
- Voice Activity Detection (VAD) — automatic server VAD, client-side Hybrid VAD, and manual Push-to-Talk
- Asynchronous function calling — non-blocking tool execution while audio continues streaming
- Full-session client content — inject and update conversation turns mid-stream
- Session management — context compression, session resumption, GoAway signals
- Ephemeral tokens — secure client-side authentication
[!NOTE] The Live API connects directly via WebSockets. For WebRTC support or simplified integration, use a partner integration.
Models
Current Models (Use These)
gemini-3.8-live— Default option for most low-latency voice agent experiences and real-time dialogue without reasoning delays. Supports interleaved reasoning, asynchronous function calling by default (behavior: NON_BLOCKING), and full-session client content updates.gemini-3.8-live-extended-thinking— High-reasoning audio-to-audio model recommended when higher background reasoning is required during live interactions. Processes background reasoning and async tool calls (behavior: NON_BLOCKINGrequired) while streaming continuous spoken conversational fillers; lifecycle managed viainteraction_status(IN_PROGRESSvsIDLE).gemini-3.5-transcribe-live— Real-time streaming speech-to-text with interim hypotheses, finalized transcripts, smart formatting, and Hybrid VAD.gemini-3.5-live-translate-preview— Real-time speech-to-speech streaming translation across 70+ languages.
[!WARNING] Legacy Models (
gemini-3.1-flash-live-preview,gemini-2.5-flash-native-audio-*,gemini-live-2.5-flash-preview,gemini-2.0-flash-live-001): Readreferences/migration.mdfor breaking protocol changes (behavior: "NON_BLOCKING",thinking_level,interaction_status,send_client_content).
SDKs
- Python:
google-genai>=2.3.0—pip install -U google-genai - JavaScript/TypeScript:
@google/genai>=2.3.0—npm install @google/genai
[!WARNING] Legacy SDKs
google-generativeai(Python) and@google/generative-ai(JS) are deprecated. Never use them.
Partner Integrations
To streamline real-time audio/video app development, use a third-party integration supporting the Gemini Live API over WebRTC or WebSockets:
- LiveKit — Use the Gemini Live API with LiveKit Agents.
- Pipecat by Daily — Create a real-time AI chatbot using Gemini Live and Pipecat.
- Fishjam by Software Mansion — Create live video and audio streaming applications with Fishjam.
- Vision Agents by Stream — Build real-time voice and video AI applications with Vision Agents.
- Voximplant — Connect inbound and outbound calls to Live API with Voximplant.
- Firebase AI SDK — Get started with the Gemini Live API using Firebase AI Logic.
Audio Formats
- Input: Raw PCM, little-endian, 16-bit, mono. 16kHz native (will resample others). MIME type:
audio/pcm;rate=16000 - Output: Raw PCM, little-endian, 16-bit, mono. 24kHz sample rate.
[!IMPORTANT] Use
send_realtime_input/sendRealtimeInputfor all real-time streaming user input (audio, video, and text). On Gemini 3.8 models,send_client_content/sendClientContentis supported across the full session lifecycle with explicit roles (userormodel) to inject conversation context (turn_complete=trueunconditionally interrupts active generation).
[!WARNING] Do not use
mediainsendRealtimeInput. Use the specific keys:audiofor audio data,videofor images/video frames, andtextfor text input.
Quick Start
Authentication
Python
from google import genai
client = genai.Client(api_key="YOUR_API_KEY")
JavaScript
import { GoogleGenAI } from '@google/genai';
const ai = new GoogleGenAI({ apiKey: 'YOUR_API_KEY' });
Connecting to the Live API
Python
from google.genai import types
config = types.LiveConnectConfig(
response_modalities=[types.Modality.AUDIO],
system_instruction=types.Content(
parts=[types.Part(text="You are a helpful assistant.")]
)
)
async with client.aio.live.connect(model="gemini-3.8-live", config=config) as session:
pass # Session is active
JavaScript
const session = await ai.live.connect({
model: 'gemini-3.8-live',
config: {
responseModalities: ['audio'],
systemInstruction: { parts: [{ text: 'You are a helpful assistant.' }] }
},
callbacks: {
onopen: () => console.log('Connected'),
onmessage: (response) => console.log('Message:', response),
onerror: (error) => console.error('Error:', error),
onclose: () => console.log('Closed')
}
});
Sending Text
Python
await session.send_realtime_input(text="Hello, how are you?")
JavaScript
session.sendRealtimeInput({ text: 'Hello, how are you?' });
Sending Audio
Python
await session.send_realtime_input(
audio=types.Blob(data=chunk, mime_type="audio/pcm;rate=16000")
)
JavaScript
session.sendRealtimeInput({
audio: { data: chunk.toString('base64'), mimeType: 'audio/pcm;rate=16000' }
});
Sending Video
Python
# frame: raw JPEG-encoded bytes
await session.send_realtime_input(
video=types.Blob(data=frame, mime_type="image/jpeg")
)
JavaScript
session.sendRealtimeInput({
video: { data: frame.toString('base64'), mimeType: 'image/jpeg' }
});
Receiving Audio and Text
[!IMPORTANT] A single server event can contain multiple content parts simultaneously (e.g., audio chunks and transcript). Always process all parts in each event to avoid missing content.
Python
async for response in session.receive():
content = response.server_content
if content:
# Audio — process ALL parts in each event
if content.model_turn:
for part in content.model_turn.parts:
if part.inline_data:
audio_data = part.inline_data.data
# Transcription
if content.input_transcription:
print(f"User: {content.input_transcription.text}")
if content.output_transcription:
print(f"Gemini: {content.output_transcription.text}")
# Interruption
if content.interrupted is True:
pass # Stop playback, clear audio queue
JavaScript
// Inside the onmessage callback
const content = response.serverContent;
if (content?.modelTurn?.parts) {
for (const part of content.modelTurn.parts) {
if (part.inlineData) {
const audioData = part.inlineData.data; // Base64 encoded
}
}
}
if (content?.inputTranscription) console.log('User:', content.inputTranscription.text);
if (content?.outputTranscription) console.log('Gemini:', content.outputTranscription.text);
if (content?.interrupted) { /* Stop playback, clear audio queue */ }
Background Reasoning (Extended Thinking)
Use gemini-3.8-live-extended-thinking when your voice agent must evaluate complex data, plan multiple steps, or handle long-running tools. The model speaks natural conversational fillers (e.g. "Checking flight options now...") while executing asynchronous tools in the background.
Key requirements:
- Thinking config: Set
thinking_config=types.ThinkingConfig(thinking_level="low")("minimal"|"low"|"medium"|"high"). - Non-blocking tools: All function declarations must set
behavior="NON_BLOCKING". Synchronous blocking mode is not supported and returns an error. - Lifecycle tracking (
interaction_status): Do not rely onturn_complete=Truealone to detect turn completion. Monitormessage.interaction_status(Python) /message.interactionStatus(JS):"IN_PROGRESS": Server is reasoning, speaking conversational fillers, or waiting for async tool responses."IDLE": Server has completed all background reasoning and tool calls; session is ready for user input.
See references/migration.md and the Thinking in Live API Guide for complete Python and JavaScript implementation examples.
Live Translation (Gemini Live Translate)
The Live API supports real-time, low-latency streaming translation of speech (audio) across 70+ languages. For full details on options and capabilities, see the Live Translate Guide.
Model
gemini-3.5-live-translate-preview— The recommended translation model for all Live Translate use cases.
Configuration (TranslationConfig)
To enable translation, specify a TranslationConfig object inside your live session setup:
- Python SDK: Configure the connection using
translation_configonLiveConnectConfig:config = types.LiveConnectConfig( response_modalities=[types.Modality.AUDIO], translation_config=types.TranslationConfig( target_language_code="es", # Target language code (e.g. es, fr, pl) echo_target_language=True, ), input_audio_transcription=types.AudioTranscriptionConfig(), output_audio_transcription=types.AudioTranscriptionConfig(), ) - Raw WebSockets: Place
translationConfiginsidegenerationConfig:{ "setup": { "model": "models/gemini-3.5-live-translate-preview", "generationConfig": { "responseModalities": ["AUDIO"], "translationConfig": { "targetLanguageCode": "es", "echoTargetLanguage": true } } } }
Live Streaming Transcription (Gemini Live Transcribe)
The Live API supports real-time streaming speech-to-text over WebSockets with low-latency interim hypotheses, finalized transcripts, and Hybrid VAD. For full details, see the Live Transcription Guide and Colab Cookbook.
Model
gemini-3.5-transcribe-live
Modes
smart: cleans up filler words, resolves inline self-corr
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
86.0kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.0kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
ruflo
73.4k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
crawl4ai
84.4kOpen-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
