LLM API Pricing

Google Models

Browse all 14 Google models available on TokenHot.

One API key gives you access to 14 Google models with live pricing and no separate provider account.

Gemini
CONTEXT1.05M
Input$0.7874/M
Output$3.9370/M
The flagship Flash model for long-horizon software engineering, autonomous agents, and complex enterprise workflows. It keeps the speed and cost profile of Flash while adding stronger reasoning and tool orchestration. It accepts text, images, video, audio, and PDF input, with about a 1.05M-token context window, search grounding, function calling, code execution, structured outputs, and tunable thinking. Prefer it over Flash-Lite for harder tasks, at higher usage cost.
Input Type:
Output Type:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong ContextCode Execution
Try model
Gemini
CONTEXT1.05M
Input$0.7500/M
Output$3.7500/M
A new fast multimodal model in the Gemini 3 family, released in August 2026, built for low-cost, high-throughput multimodal understanding. It accepts text, image, video, audio and document inputs with a 1M-token context, suited to large-scale processing, batch workloads and real-time interactions.
Input Type:
Output Type:
ReasoningTool UseLong Context
Try model
CONTEXT1.05M
Input$0.3000/M
Output$2.5000/M
The fastest and most cost-effective model in the Gemini 3.5 line, built for high-throughput, low-latency work such as translation, classification, document processing, and lightweight agentic workflows. It supports multimodal inputs and up to a 1M-token context, while demanding reasoning, long-horizon planning, or high-quality code changes are better handled by a higher-tier Flash model.
Input Type:
Output Type:
Long Context
Try model
Gemini
PER SEC$0.2450/s
Built for any-input-to-video multimodal creation, putting stronger emphasis than earlier Gemini video workflows on unified orchestration of text, image, audio, and video materials. Compared with Veo, Kling, and Runway, it is better for integrating complex multimodal assets into one creative workflow.
Input Type:
Output Type:
Multimodal Output
Try model
CONTEXT131K
PER IMG$0.0400/image
Built for image generation and editing, with stronger multimodal understanding, conversational edits, and rapid visual prototyping than earlier Gemini image capabilities. Compared with GPT Image, Qwen Image, and Seedream, it is useful for combining documents, images, and prompts in one visual workflow.
Input Type:
Output Type:
ReasoningMultimodal Output
Try model
CONTEXT1.05M
Input$2.0000/M
Output$12.0000/M
Built for multimodal understanding, coding, and long-context tasks, improving over Gemini 3.0 / 2.5 in throughput, context, and tool capabilities. Compared with GPT, Claude, and Qwen models, it is strong for unified analysis workflows across text, images, audio, video, and documents.
Input Type:
Output Type:
Long Context
Try model
CONTEXT1M
Input$0.5000/M
Output$3.0000/M
Google's high-speed thinking model designed for agentic workflows, multi-turn chat, and coding assistance. Supports text, image, audio, video, and PDF input with a 1M token context window. Features configurable thinking levels, tool use, and structured output. Broad quality improvements over Gemini 2.5 Flash across reasoning, multimodal understanding, and reliability.
Input Type:
Output Type:
ReasoningTool UseFunction CallingStructured OutputLong Context
Try model
CONTEXT66K
PER IMG$0.1340/image
Built for image generation and editing, with stronger multimodal understanding, conversational edits, and rapid visual prototyping than earlier Gemini image capabilities. Compared with GPT Image, Qwen Image, and Seedream, it is useful for combining documents, images, and prompts in one visual workflow.
Input Type:
Output Type:
ReasoningMultimodal Output
Try model
Gemini
PER CALL$0.1100/call
Built for high-fidelity video generation, with a more mature native-audio, camera-control, and reference-input workflow than Veo 2/3. Compared with Kling, Runway, and Seedance, it is better suited to cinematic shorts, ad storyboards, and videos requiring stable physical motion. The Lite tier favors high-throughput, lower-cost drafts over the standard tier.
Input Type:
Output Type:
Multimodal Output
Try model
Gemini
PER CALL$0.2211/call
Built for high-fidelity video generation, with a more mature native-audio, camera-control, and reference-input workflow than Veo 2/3. Compared with Kling, Runway, and Seedance, it is better suited to cinematic shorts, ad storyboards, and videos requiring stable physical motion. The Fast tier is better for rapid concept validation than the standard tier.
Input Type:
Output Type:
Multimodal Output
Try model
Gemini
Google
PER CALL$1.6544/call
Built for high-fidelity video generation, with a more mature native-audio, camera-control, and reference-input workflow than Veo 2/3. Compared with Kling, Runway, and Seedance, it is better suited to cinematic shorts, ad storyboards, and videos requiring stable physical motion.
Input Type:
Output Type:
Multimodal Output
Try model
Gemini
PER IMG$0.0300/image
This is a discounted trial route. Generation may take longer and has roughly a 10% chance of failing. It is best for cost-sensitive image generation trials where queueing and occasional failures are acceptable; for more stable production use, choose a regular route.
Input Type:
Output Type:
Multimodal Output
Gemini
PER CALL$0.0200/call
This is a discounted trial route. Generation may take longer and has roughly a 10% chance of failing. It is best for cost-sensitive image generation trials where queueing and occasional failures are acceptable; for more stable production use, choose a regular route.
Input Type:
Output Type:
Multimodal Output
Gemini
PER IMG$0.0300/image
This is a discounted trial route. Generation may take longer and has roughly a 10% chance of failing. It is best for cost-sensitive image generation trials where queueing and occasional failures are acceptable; for more stable production use, choose a regular route.
Input Type:
Output Type:
Multimodal Output