Skip to main content
9 – 17 UHR +49 8031 3508270 LUITPOLDSTR. 9, 83022 ROSENHEIM
DE / EN
LLM Google USA

Google Gemini

Google Gemini via Vertex AI (Gemini Enterprise Agent Platform). New since July 21, 2026: Gemini 3.6 Flash (GA, $1.50/1M input) and Gemini 3.5 Flash-Lite (GA, 350 tokens/s). Gemini 3.5 Flash Cyber starts as a government-only pilot. Gemini 3.5 Pro further delayed. GDPR EU deployment: Gemini 3.5 Flash via the EU multi-region endpoint. innFactory AI Consulting Germany.

License Proprietary
GDPR Hosting Available
Context Up to 2M Tokens
Modality Text, Image, Audio, Video, PDF, Code → Text, Image, Audio, Code

Versions

Overview of available model variants

ModelReleaseEUStrengthsWeaknessesStatus
Gemini 3.6 Flash (GA)
21 July 2026
New workhorse model: better coding, knowledge work and multimodal performance than 3.5 Flash DeepSWE 49% (Gemini 3.5 Flash: 37%), OSWorld-Verified 83% (computer use) 17% fewer output tokens than 3.5 Flash – at a lower price ($1.50/1M input, $7.50/1M output) 1M token input context (1,048,576), 65,536 output, knowledge cutoff March 2026
EU multi-region endpoint not yet confirmed (as of 24 July 2026) Sampling parameters temperature/top_p/top_k are deprecated on 3.6 Flash – review existing prompt integrations
Current
Gemini 3.5 Flash-Lite (GA)
21 July 2026
Fastest model in the 3.5 series: 350 output tokens/second (measured by Artificial Analysis) Very affordable: $0.30/1M input, $2.50/1M output Ideal for high-volume automation with low latency
Lower capability than 3.6 Flash / 3.5 Flash EU multi-region endpoint not yet confirmed
Current
Gemini 3.5 Flash Cyber
Announced 21 July 2026 (limited-access pilot)
Specialised cybersecurity model, paired with Google's CodeMender agent Focus on vulnerability detection and remediation
Governments and trusted partners only – not commercially available
Preview
Gemini 3.5 Flash (GA) Recommended
19 May 2026 (Google I/O 2026)
Outperforms Gemini 3.1 Pro on Terminal-Bench 2.1, GDPval-AA Elo and MCP Atlas 289 tokens/second – ~4× faster than other frontier models Agentic-first: multi-hour autonomous coding and research pipelines Default model in Gemini app, AI Mode in Search, Antigravity and Gemini Enterprise EU multi-region endpoint for GDPR-compliant routing inside the EU geography
No single-region EU endpoint (Frankfurt / Netherlands) yet — only the EU multi-region endpoint
Current
Gemini 3.5 Pro (Limited Preview)
Limited Preview since May 2026, GA date open
Designed as orchestrator/planner for multi-agent workflows Operates 3.5/3.6 Flash as sub-agents 2M token context window and announced Deep Think reasoning mode
GA repeatedly delayed – originally expected for June 2026, date now open EU multi-region availability not yet confirmed
Preview
Gemini 3.1 Pro (GA)
February 2026
Complex reasoning 2M token context window Multimodal
Surpassed by Gemini 3.5 Flash on key benchmarks
Current
Gemini 3.1 Flash (GA)
January 2026
Strong price-performance ratio 1M token context window
Superseded by Gemini 3.5 Flash
Current
Gemini 3.1 Flash Thinking
February 2026
Extended reasoning Strong on STEM tasks
Higher latency due to thinking process
Current
Gemini 3.1 Pro Deep Research
February 2026
Multi-hop research Long analysis tasks
Specialised, not general-purpose
Current
Gemini 3 Pro (Preview)
January 2026
Reasoning-first Multimodal
Superseded by 3.1 Pro GA
Preview
Gemini 3 Flash (Preview)
January 2026
Fast Strong multimodal performance
Superseded by 3.1 Flash GA
Preview
Gemini 2.5 Pro
2025
Proven track record
Deprecated
Deprecated
Gemini 2.0 Flash
December 2024
Cost-efficient
Deprecated
Deprecated

Use Cases

Typical applications for this model

Video Analysis
Research & Document Analysis
Multimodal Applications
Vibe Coding
Agentic Workflows
Data Analysis
Google Workspace Integration

Technical Details

API, features and capabilities

API & Availability
Availability Public
Requests/Min 1500
Tokens/Min 4000000
Latency (TTFT) ~300ms
Throughput ~200 Tokens/Sec
Features & Capabilities
Tool Use Function Calling Structured Output Vision Reasoning Mode Code Execution Web Browsing File Upload Realtime API
Training & Knowledge
Knowledge Cutoff 2026-03 (Gemini 3.6 Flash)
Fine-Tuning Available (Supervised Fine-Tuning)
Language Support
Best Quality English, German, French, Spanish, Japanese, Korean, Chinese
Supported 100+ Languages
Excellent multilingual capabilities through multimodal training

Hosting & Compliance

GDPR-compliant hosting options and licensing

GDPR-Compliant Hosting Options
Google Cloud
Netherlands (europe-west4)
Vertex AI
License & Hosting
License Proprietary
Security Filters Customizable
Enterprise Support Yes
SLA Available Yes
Cloud Only

Benchmarks

Performance comparison with standardized tests

DeepSWE (Gemini 3.6 Flash)
49
OSWorld-Verified (Gemini 3.6 Flash)
83

innFactory AI Consulting, based in Rosenheim, Germany, supports enterprises across the DACH region (Germany, Austria, Switzerland) and Europe with GDPR-compliant deployment of Google Gemini. On 21 July 2026, Google released three new models: Gemini 3.6 Flash (GA) as the new workhorse model with better coding and 17 percent fewer output tokens than 3.5 Flash, Gemini 3.5 Flash-Lite (GA) as the fastest model in the 3.5 series at 350 output tokens per second, and Gemini 3.5 Flash Cyber as a specialised security model in a limited-access pilot for governments and vetted partners. Gemini 3.5 Pro, by contrast, keeps slipping: Google is still testing the orchestrator model with partners and no longer names a GA date. In parallel, DeepMind confirmed that the pre-training run for Gemini 4 is already underway.

EU availability (as of 24 July 2026): For GDPR-compliant workloads, Gemini 3.5 Flash remains the reference – it is available via the EU multi-region endpoint on Vertex AI / Gemini Enterprise Agent Platform; requests are routed exclusively inside the EU geography and covered by the Vertex AI Data Processing Addendum. For the new Gemini 3.6 Flash, EU multi-region availability is not yet confirmed – though for 3.5 Flash the EU endpoint followed just two days after GA; we are monitoring continuously. A single-region endpoint in europe-west3 (Frankfurt) or europe-west4 (Netherlands) is still missing for the 3.x family; for strict single-region pinning, Gemini 2.5 Pro and Gemini 2.0 Flash in europe-west4 remain the path.

Key Strengths

Gemini 3.6 Flash – The New Workhorse (GA since 21 July 2026)

Gemini 3.6 Flash (gemini-3.6-flash) improves coding, knowledge work, and multimodal processing over 3.5 Flash while being cheaper: $1.50/1M input, $7.50/1M output. On the DeepSWE coding benchmark the model jumps from 37 to 49 percent, and it reaches 83 percent on OSWorld-Verified (computer use). At the same time it produces 17 percent fewer output tokens than 3.5 Flash – Google directly addressing developer feedback on verbosity, which lowers effective per-task cost further. The knowledge cutoff advances to March 2026; the context window stays at 1 million input tokens (65,536 output). The model is available via the Gemini API (AI Studio, Android Studio), Antigravity, the Gemini Enterprise Agent Platform, and the Gemini app. Migration note: the sampling parameters temperature, top_p, and top_k are deprecated on the new models – review existing integrations.

Gemini 3.5 Flash-Lite – Speed for High-Volume Workloads (GA since 21 July 2026)

Gemini 3.5 Flash-Lite (gemini-3.5-flash-lite) is the fastest model in the 3.5 series: 350 output tokens per second as measured by Artificial Analysis, priced at $0.30/1M input and $2.50/1M output. That makes it a fit for high-volume automation, classification, and latency-critical applications – for example as a sub-agent in multi-agent pipelines. Google is additionally rolling the model out in Google Search.

Gemini 3.5 Flash Cyber – Security Model in a Pilot Program

With Gemini 3.5 Flash Cyber, Google announced its first dedicated cybersecurity model, paired with the CodeMender security agent. It remains exclusive for now: access is limited to governments and vetted partners via a limited-access pilot program. The model is therefore not currently relevant for enterprises, but it shows Google’s course of rolling out security-critical capabilities in a controlled way – a pattern we already know from the government clearance processes at OpenAI GPT-5.6 and Anthropic Claude Fable 5.

Gemini 3.5 Flash – Agentic-first

With Gemini 3.5 Flash, Google positions its fastest tier above its own Pro flagship for the first time: 3.5 Flash outperforms Gemini 3.1 Pro on Terminal-Bench 2.1 (coding), GDPval-AA Elo (real-world agentic) and MCP Atlas (tool use). The model is built for multi-hour autonomous workflows – coding pipelines, research projects, even building entire systems – pausing for human input at decision points. An optimised variant is reported to reach up to 12× the speed of other frontier models at equivalent quality.

Gemini 3.5 Pro – Orchestrator (Further Delayed)

Gemini 3.5 Pro remains in partner testing: the GA originally expected for June 2026 has slipped, and in the July 21 release blog post Google only states the model will be “broadly available as soon as it’s ready”. The model is designed as the orchestrator/planner that drives Gemini 3.5/3.6 Flash instances as sub-agents. Google has announced a 2M token context window and a Deep Think reasoning mode for Pro, explicitly positioning the family for multi-agent architectures rather than classic chatbots.

Gemini Spark – Agentic 24/7 Assistant

At Google I/O 2026, Google introduced Gemini Spark, an agentic personal assistant built on Gemini 3.5 and the Antigravity platform. Spark runs continuously on Google’s cloud, can be addressed directly via a dedicated Gmail address, and executes background tasks across Chrome and Google Workspace. Spark supports the Model Context Protocol (MCP) for external integrations and is rolling out first to Google AI Ultra subscribers.

Industry-Leading Context Window

Gemini 3.1 Pro processes up to 2 million tokens in a single context. This enables analysis of extensive contract portfolios, technical documentation, or entire codebases in a single request. The Flash models offer a 1 million token context window at significantly lower cost.

Native Multimodal Processing

The Gemini 3.x family natively processes text, images, audio, video, and PDF documents. This enables use cases such as automated video analysis, document extraction from scanned PDFs, or meeting evaluation combining audio and visual material.

Google Ecosystem

  • Vertex AI: Enterprise deployment with SLA (EU hosting only for older models)
  • Antigravity 2.0: Agentic development platform for Gemini 3.5 and 3.6
  • Gemini Enterprise: Enterprise frontend with access to 3.6 Flash and 3.5 Flash
  • Google Workspace: Integration with Docs, Sheets, Gmail, and more
  • Search Grounding: Access to current web information
  • Google AI Studio: Rapid prototyping and API access

Specialised Variants

  • Gemini 3.1 Flash Thinking: Extended reasoning for complex STEM tasks with a transparent thinking process
  • Gemini 3.1 Pro Deep Research: Specialised for multi-step research and long analysis tasks

EU Availability

As of 24 July 2026: Gemini 3.5 Flash is GA on the EU multi-region endpoint on Vertex AI / Gemini Enterprise Agent Platform. Requests are routed exclusively inside the EU geography and the model is covered by the Vertex AI Data Processing Addendum. For the new models Gemini 3.6 Flash and Gemini 3.5 Flash-Lite (GA since 21 July), EU multi-region availability is not yet confirmed. Gemini 3.5 Pro remains in partner testing without a GA date.

What works from EU projects today

  • Gemini 3.5 Flash via the EU multi-region endpoint — GDPR-compliant routing inside the EU geography, GA on Vertex AI / Gemini Enterprise Agent Platform.
  • Gemini 3.6 Flash and 3.5 Flash-Lite via the global endpoint — available for workloads that don’t require regional binding.
  • Gemini 2.5 Pro and Gemini 2.0 Flash as single-region regional endpoints in europe-west4 (Netherlands), partial coverage in europe-west3 (Frankfurt) — for customers that need strict pinning to a specific EU country.

What is still missing

  • Confirmed EU multi-region availability for Gemini 3.6 Flash and Gemini 3.5 Flash-Lite — for 3.5 Flash the EU endpoint arrived two days after GA (19 → 21 May 2026); we expect a similar pattern and are monitoring continuously
  • Single-region endpoints for the 3.x family in europe-west3 (Frankfurt) or europe-west4 (Netherlands)
  • Gemini 3.5 Pro (partner testing, GA date open) — EU availability to be confirmed at GA launch

We recommend monitoring the Google Cloud regional availability documentation and the Gemini Enterprise Agent Platform release notes regularly.

Integration with CompanyGPT

Gemini models are integrated in CompanyGPT. Gemini 3.5 Flash is wired in via the EU multi-region endpoint for GDPR-compliant frontier reasoning and agentic coding workloads. Gemini 2.5 Pro and Gemini 2.0 Flash remain available as single-region endpoints in europe-west4 for customers that require strict country-level data residency. As soon as Google ships single-region 3.5 endpoints in Frankfurt or the Netherlands we will add them as default options.

Our Recommendation

Since 21 July 2026, Gemini 3.6 Flash is Google’s most capable GA model – but for GDPR-compliant EU workloads, Gemini 3.5 Flash remains the recommendation until the EU endpoint is confirmed:

  • Frontier reasoning and agentic workflows in EU: Gemini 3.5 Flash via the EU multi-region endpoint – switch to 3.6 Flash as soon as Google confirms the EU endpoint (we will update this page)
  • Workloads without EU residency requirements: Gemini 3.6 Flash via the global endpoint – better coding at a lower price than 3.5 Flash
  • High-volume and latency-critical: Gemini 3.5 Flash-Lite ($0.30/1M input, 350 tokens/s) via the global endpoint
  • Strict single-country EU residency: Gemini 2.5 Pro in europe-west4
  • Cross-cloud alternative: Anthropic Claude Opus 4.8 or Claude Sonnet 4.6 via AWS Bedrock (in-region Ireland / Stockholm / Frankfurt) — useful for hybrid Multi-Cloud strategies or when the workload is not bound to Google Cloud

We are closely tracking the EU rollout of Gemini 3.6 Flash and 3.5 Flash-Lite and will update this page once Google ships the endpoints.

Cost estimation for this model

For up-to-date token pricing, model variants and EU availability, see our sister project ai-prices.eu. It helps you compare and estimate the operational cost of leading AI models for your specific use case.

Compare prices on ai-prices.eu

ai-prices.eu is a project by innFactory AI Consulting GmbH and provides transparent cost estimates for leading AI models.

Consultation for this model?

We help you select and integrate the right AI model for your use case.