Google Fast multimodal processing
Gemini 3.5 Flash-Lite
Move high-throughput everyday tasks through a low-latency Gemini route.
Gemini 3.5 Flash-Lite is the lightweight Gemini route in TokenHot's catalog. It accepts text, images, audio, video and documents, with a 1.05M-token context value and text output. Use it for translation, classification, document processing and lightweight agent steps where throughput and predictable output matter more than deep planning.
- Text + image + audio + video
- 1.05M context
- High-throughput text
Flash-Lite is positioned for routine, high-throughput work. It should not be treated as a guarantee for complex reasoning, code changes or long-horizon planning; route those tasks to a stronger model and validate every result.
AI Chat
Send a prompt and see the response here.
gemini-3.5-flash-liteStart a conversation with this model and ask follow-up questions.
USD /1M tokens
Pricing
| Group | Input | Output | Cache Read | Image | Audio input |
|---|---|---|---|---|---|
| default | $0.3 | $2.49999 | $0.03 | $0.3 | $0.3 |
Public list prices are shown in USD. Final charges may vary by account group and usage tier.
Overview
What is Gemini 3.5 Flash-Lite?
A good Flash-Lite request has a narrow job and a known output shape: classify, extract, translate, summarize or route. Keep a stable fixture set so throughput gains remain visible alongside accuracy and repair rate.
The model accepts many input types, but preprocessing still matters. Store document pages, audio transcripts and video timestamps so a reviewer can trace why a result was produced.
Define a small output contract
Keep each call focused and return only the fields the downstream system needs.
Preserve source boundaries
Label pages, frames and audio segments so extraction results remain traceable.
Escalate complexity early
Use a task classifier or policy to send deep reasoning and code changes to a stronger route.
From brief to deliverable
Put Gemini 3.5 Flash-Lite to work
Operations and knowledge teams
Classify and extract a document queue
Turn incoming documents into typed fields, categories and review flags for a downstream workflow.
- Tools you need
- Document parsing, OCR when needed and schema validation.
- Keep in mind
- OCR or parsing errors can propagate into a clean-looking result; retain page evidence and escalate uncertain text.
View a task brief
Extract the requested fields from the supplied documents. Return valid JSON with page references, missing values and a review flag. Do not invent text that cannot be read.
Support and content teams
Route a mixed media queue
Classify text, image, audio and video items into a next action without making a final high-impact decision.
- Tools you need
- Media decoding, transcript extraction and queue management.
- Keep in mind
- A route decision is not a judgment of hidden state or policy compliance; escalate ambiguous cases.
View a task brief
Classify each supplied media item using the allowed categories. Return the evidence used, the next queue and an escalation flag for ambiguity.
Localization teams
Translate a recurring content stream
Translate and normalize short content while preserving names, placeholders, formatting and prohibited terms.
- Tools you need
- Translation memory, glossary checks and a human language reviewer.
- Keep in mind
- Fluency does not prove legal, cultural or product correctness.
View a task brief
Translate the supplied content into the target locale. Preserve placeholders and named terms, then list any ambiguous or culturally sensitive wording for review.
Capabilities in context
Key features
| Feature | What you get | Why it matters |
|---|---|---|
| Input coverage | Text, image, audio, video and document input are listed | Use the input type that preserves evidence and store page, frame or segment references. |
| Context capacity | 1,048,576 tokens shown in the catalog | Large documents still need chunking, retrieval and a clear output budget. |
| Fast positioning | Built for low-latency, high-throughput work | Benchmark the complete queue, including preprocessing, retries and human review. |
| Text output | The catalog lists text output | Use a schema for downstream systems and keep image or video generation on a dedicated route. |
| Dual endpoints | Gemini and OpenAI endpoint types are listed | Pin the endpoint protocol and test request shape, media handling and error semantics separately. |
Modalities and endpoint types come from TokenHot's current catalog. Verify the selected Gemini or OpenAI protocol, media limits and safety behavior before production use.
Audience & fit
Where Gemini 3.5 Flash-Lite fits
Operations teams
Process document and media queues with explicit schemas, references and review flags.
Localization teams
Use narrow translation contracts with glossary and placeholder checks.
Product teams building intake flows
Route mixed inputs quickly, then escalate ambiguous or consequential decisions to people or stronger models.
Cost context
Plan the cost of a useful result
Measure queue throughput
Compare cost per accepted item, including decoding, OCR, retries and review time.
Keep media segments small
A full recording or document can be costly and harder to audit. Send the segments needed for the decision.
Separate translation QA
Language review and legal checks are part of the workflow budget, even when the model call is inexpensive.
Use the current pricing table for available rates and account groups. Model IDs and prices come from TokenHot’s catalog; provider pricing and subscriptions are separate.
API
Code examples
curl --location --request POST 'https://api.tokenhot.ai/v1beta/models/gemini-3.5-flash-lite:generateContent' \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data-raw '{
"contents": [
{
"role": "user",
"parts": [
{
"text": "Hello"
}
}
]
}'FAQ
Frequently asked questions
What is Gemini 3.5 Flash-Lite best suited for?
Gemini 3.5 Flash-Lite is suited to the input, output, and capability types listed on this page. Test production workloads before choosing it for a critical system.
How is Gemini 3.5 Flash-Lite priced?
Pricing is shown in the pricing section above. Actual cost depends on usage volume, account group, and the request payload.
How can I call Gemini 3.5 Flash-Lite?
Use the model ID shown above with one of the supported API protocols. Code examples are provided when usage data is available.
What inputs does Gemini 3.5 Flash-Lite accept?
The catalog lists text, image, audio, video and document input. Confirm file size, duration and endpoint-specific request limits.
Is it intended for complex reasoning?
It is positioned for fast, cost-effective routine work. Route complex planning, deep coding or high-impact ambiguity to a stronger model.
Can it produce images or videos?
The catalog lists text output for this model. Use a dedicated image or video route for those deliverables.
Which API protocol should I use?
The catalog lists Gemini and OpenAI endpoint types. Choose one deliberately and test its media and error behavior before rollout.
Get started
Build with Gemini 3.5 Flash-Lite
API