Google AI Models

Explore 29 curated AI models from Google and compare context, API pricing, modalities and capabilities.

29models tracked

Model directory

29 models in this view

Compare
Active

A fast multimodal Gemini model listed for responsive agent workflows, coding and multi-step reasoning.

Context
1.05M
Input
$0.375
Output
$1.88
TextImageVideoFile
View model

Google

Chirp 3

Active

Chirp 3 is Google's latest multilingual speech-to-text model. It offers enhanced transcription accuracy across 24 GA languages and 77+ preview languages, with support for automatic language detection, automatic punctuation, and…

Context
0
Input
$16,000
Output
$0.00
Audio
View model

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool…

Context
1.05M
Input
$0.50
Output
$3
TextImageFileAudio
View model
Active

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic…

Context
1.05M
Input
$0.25
Output
$1.5
TextImageVideoFile
View model

Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2.5 Flash performance across…

Context
1.05M
Input
$0.25
Output
$1.5
TextImageVideoFile
View model

Gemini 3.1 Flash TTS Preview is a text-to-speech model from Google, and a substantial generational step up from Gemini 2.5 Flash TTS. It takes text input and produces audio output…

Context
33K
Input
$1
Output
$20
Text
View model

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation…

Context
1.05M
Input
$2
Output
$12
AudioFileImageText
View model

Gemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a general bash tool when more efficient third-party…

Context
1.05M
Input
$2
Output
$12
TextAudioImageVideo
View model
Active

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution…

Context
1.05M
Input
$1.5
Output
$9
TextImageVideoFile
View model
Active

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

Context
1.05M
Input
$0.30
Output
$2.5
TextImageVideoFile
View model
Active

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and…

Context
1.05M
Input
$0.75
Output
$3.75
TextImageVideoFile
View model
Active

gemini-embedding-001 provides a unified cutting edge experience across domains, including science, legal, finance, and coding. This embedding model has consistently held a top spot on the Massive Text Embedding Benchmark…

Context
20K
Input
$0.15
Output
$0.00
Text
View model
Active

Gemini Embedding 2 is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It supports…

Context
8K
Input
$0.20
Output
$0.00
TextImageFileAudio
View model

Gemini Embedding 2 Preview is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It…

Context
8K
Input
$0.20
Output
$0.00
TextImageFileAudio
View model
Active

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.

Context
262K
Input
$0.07
Output
$0.34
ImageTextVideo
View model
Active

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.

Context
262K
Input
$0.10
Output
$0.34
ImageTextVideo
View model

This model always redirects to the latest model in the Google Gemini Flash family.

Context
1.05M
Input
$0.375
Output
$1.88
TextImageVideoFile
View model

This model always redirects to the latest model in the Google Gemini Pro family.

Context
1.05M
Input
$2
Output
$12
AudioFileImageText
View model
Active

30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate…

Context
1.05M
Input
$0.00
Output
$0.00
TextImage
View model
Active

Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate high-quality, 48kHz…

Context
1.05M
Input
$0.00
Output
$0.00
TextImage
View model

Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual understanding. It is capable of image generation,…

Context
33K
Input
$0.30
Output
$2.5
ImageText
View model

Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines…

Context
66K
Input
$0.50
Output
$3
ImageText
View model

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation…

Context
66K
Input
$0.25
Output
$1.5
ImageText
View model

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and…

Context
66K
Input
$2
Output
$12
ImageText
View model

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and…

Context
66K
Input
$2
Output
$12
ImageText
View model

Google

Veo 3.1

Active

Google's state-of-the-art video generation model, built for maximum visual fidelity in final production cuts. Veo 3.1 generates high-quality 1080p video from text or image prompts with native synchronized audio —…

Context
0
Input
$0.00
Output
$0.00
TextImage
View model
Active

Google's mid-tier video generation model balancing speed and quality. Veo 3.1 Fast generates high-quality video from text or image prompts with native synchronized audio, offering faster turnaround than Veo 3.1…

Context
0
Input
$0.00
Output
$0.00
TextImage
View model
Active

Google's most cost-effective video generation model, designed for high-volume applications and rapid iteration. Veo 3.1 Lite generates 720p and 1080p video from text or image prompts with native synchronized audio…

Context
0
Input
$0.00
Output
$0.00
TextImage
View model
OUR METHOD

How AIToolly handles model data

Model-native facts and provider endpoint facts stay separate. Third-party catalogs support discovery and provider snapshots, while official documentation and verified model cards take priority for model facts.

01Sources and verification
02Verified
03Capability