Skip to main content
9 – 17 UHR +49 8031 3508270 LUITPOLDSTR. 9, 83022 ROSENHEIM
DE / EN
LLM Google USA

Google Gemini

Google Gemini via Vertex AI (Gemini Enterprise Agent Platform). New since 2 September 2026: Gemini 3.8 Flash (GA, gemini-3.8-flash, introductory price $0.75/1M input through 31 Dec 2026) and Gemini 3.8 Flash Cyber (Fairwind Program only). Key point for the EU: per Google's locations docs, 3.8, 3.7 and 3.6 Flash run only in the global region without data residency – Gemini 3.5 Flash remains the latest Gemini with EU data residency (DRZ + MLP). As of 3 September 2026. innFactory AI Consulting Germany.

License Proprietary
GDPR Hosting Available
Context Up to 2M Tokens
Modality Text, Image, Audio, Video, PDF, Code → Text, Image, Audio, Code

Versions

Overview of available model variants

ModelReleaseEUStrengthsWeaknessesStatus
Gemini 3.8 Flash (GA)
2 September 2026
Positioned by Google as its most intelligent Flash model to date; third Flash release within six weeks Per Google's evaluation: DeepSWE v1.1 73.7% (3.7 Flash: 65.3%), Terminal-bench 2.1 89.4% (85.8%), Terminal-bench 4.0 19.1% (11.2%) HLE-Verified 54.9% (53.6%), LVBench 87.8% agentic (85.4%), OSWorld 2.0 59.0% (50.6%), GDPval-AA v2 Elo 1545 (1482) Vals Finance Agent v2 61.4%, Harvey's Legal Agent Benchmark 10.0% (per Google) Introductory price through 31 Dec 2026: $0.75/1M input, $3.75/1M output (caching $0.075) – identical to 3.7 and 3.6 Flash 1M token input context (1,048,576), 65,536 output, knowledge cutoff March 2026, thinking levels low/medium (default)/high Input: text, image, video, audio, PDF; computer use (preview), batch, flex and priority inference
No EU data residency commitment: no DRZ and no MLP in 'eu', global region only Per Google the model 'works harder' and uses more tokens at higher effort levels – Artificial Analysis measures around 30% more output tokens per task than 3.7 Flash, i.e. higher cost per task despite the same list price Text output only, no Live API; agentic video understanding (GA since 1 September 2026) is listed by Google only for 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite thinking_level 'minimal' is rejected; sampling parameters temperature/top_p/top_k/candidate_count deprecated From 1 January 2027 the price doubles to $1.50/1M input, $7.50/1M output
Current
Gemini 3.8 Flash Cyber
Announced 2 September 2026 (Fairwind Program)
Successor to Gemini 3.5 Flash Cyber; per Google CyberGym pass@1 86.2% (3.5 Flash Cyber: 77.5%) Internal vulnerability benchmark across 20 languages 71.0% (3.7 Flash: 58.9%), CWE-Bench pass@1 47.2% (per Google) Focus on vulnerability detection and remediation
Fairwind participants only – not commercially available
Preview
Gemini 3.7 Flash (GA)
13 August 2026
Positioned by Google as a workhorse model for coding, web dev, knowledge work and agentic workflows DeepSWE v1.1 65.3% (3.6 Flash: 49.0%), FrontierCode 1.1 Main 43.6% (34.4%), WebDev Arena Elo 1588 (1538) Terminal-bench 2.1 85.8%, Harvey LAB-AA 90.7%, LVBench 85.4% Agentic video understanding GA since 1 September 2026 (no surcharge, per Google up to 88% fewer tokens on video tasks) Introductory price through 31 Dec 2026: $0.75/1M input, $3.75/1M output (caching $0.075) 1M token input context (1,048,576), 65,536 output, thinking levels low/medium/high, knowledge cutoff March 2026 (per model card)
No EU data residency commitment: no DRZ and no MLP in 'eu', global region only Sampling parameters temperature/top_p/top_k are deprecated – review existing prompt integrations Superseded by Gemini 3.8 Flash (2 September 2026) as the current Flash model; Google lists 3.7 Flash as previous generation From 1 January 2027 the price doubles to $1.50/1M input, $7.50/1M output
Current
Gemini 3.6 Flash (GA)
21 July 2026
Workhorse model: better coding, knowledge work and multimodal performance than 3.5 Flash DeepSWE 49% (Gemini 3.5 Flash: 37%), OSWorld-Verified 83% (computer use) 17% fewer output tokens than 3.5 Flash Introductory price cut to $0.75/1M input, $3.75/1M output through 31 Dec 2026 (launch price: $1.50/$7.50) 1M token input context (1,048,576), 65,536 output, knowledge cutoff March 2026
No EU machine learning processing – usable from the EU only via the global region Sampling parameters temperature/top_p/top_k are deprecated on 3.6 Flash – review existing prompt integrations Superseded by Gemini 3.7 Flash (13 August 2026) and 3.8 Flash (2 September 2026)
Current
Gemini 3.5 Flash-Lite (GA)
21 July 2026
Fastest model in the 3.5 series: 350 output tokens/second (measured by Artificial Analysis) Very affordable: $0.30/1M input, $2.50/1M output Ideal for high-volume automation with low latency
Lower capability than 3.7 / 3.6 Flash No documented EU data residency
Current
Gemini 3.5 Flash Cyber
Announced 21 July 2026 (limited-access pilot)
Specialised cybersecurity model, paired with Google's CodeMender agent Focus on vulnerability detection and remediation
Governments and trusted partners only – not commercially available Superseded by Gemini 3.8 Flash Cyber (Fairwind Program) since 2 September 2026
Preview
Gemini 3.5 Flash (GA) Recommended
19 May 2026 (Google I/O 2026)
Latest Gemini model with EU data residency: DRZ and MLP supported in the EU multi-region Outperforms Gemini 3.1 Pro on Terminal-Bench 2.1, GDPval-AA Elo and MCP Atlas 289 tokens/second – ~4× faster than other frontier models Agentic-first: multi-hour autonomous coding and research pipelines Default model in Gemini app, AI Mode in Search, Antigravity and Gemini Enterprise $1.50/1M input, $9.00/1M output
No single-region EU endpoint (Frankfurt / Netherlands) — only the EU multi-region endpoint Surpassed on coding benchmarks by 3.6, 3.7 and 3.8 Flash; Google itself now calls 3.5 Flash a legacy Flash model (status still Stable, deprecated without a shutdown date)
Current
Gemini 3.5 Pro (Limited Preview)
Limited Preview since May 2026, GA date open
Designed as orchestrator/planner for multi-agent workflows Operates 3.5/3.6/3.7 Flash as sub-agents 2M token context window and announced Deep Think reasoning mode
GA repeatedly delayed – as of 3 September 2026 without a new date EU availability not confirmed
Preview
Gemini 3.1 Pro (Preview)
Preview since 19 February 2026
Pro flagship of the Gemini family Complex reasoning, 2M token context window Multimodal; $2/1M input, $12/1M output at up to 200k context
Surpassed on key benchmarks by the newer Flash models Preview status, no EU data residency
Preview
Gemini 3.1 Flash (GA)
January 2026
Strong price-performance ratio 1M token context window
Superseded by Gemini 3.5 Flash
Current
Gemini 3.1 Flash Thinking
February 2026
Extended reasoning Strong on STEM tasks
Higher latency due to thinking process
Current
Gemini 3.1 Pro Deep Research
February 2026
Multi-hop research Long analysis tasks
Specialised, not general-purpose
Current
Gemini 3 Pro (Preview)
January 2026
Reasoning-first Multimodal
Shut down on 9 March 2026
Deprecated
Gemini 3 Flash (Preview)
January 2026
Fast Strong multimodal performance
Superseded by 3.1 Flash GA
Preview
Gemini 2.5 Pro
2025
Proven track record, EU data residency with DRZ and MLP
Legacy generation, technically surpassed by the 3.x family
Current
Gemini 2.0 Flash
December 2024
Was cost-efficient
Shut down since 1 June 2026
Deprecated

Use Cases

Typical applications for this model

Video Analysis
Research & Document Analysis
Multimodal Applications
Vibe Coding
Agentic Workflows
Data Analysis
Google Workspace Integration

Technical Details

API, features and capabilities

API & Availability
Availability Public
Requests/Min 1500
Tokens/Min 4000000
Latency (TTFT) ~300ms
Throughput ~200 Tokens/Sec
Features & Capabilities
Tool Use Function Calling Structured Output Vision Reasoning Mode Code Execution Web Browsing File Upload Realtime API
Training & Knowledge
Knowledge Cutoff 2026-03 (Gemini 3.8 Flash, 3.7 Flash and 3.6 Flash per model cards)
Fine-Tuning Available (Supervised Fine-Tuning)
Language Support
Best Quality English, German, French, Spanish, Japanese, Korean, Chinese
Supported 100+ Languages
Excellent multilingual capabilities through multimodal training

Hosting & Compliance

GDPR-compliant hosting options and licensing

GDPR-Compliant Hosting Options
Google Cloud
EU multi-region (aiplatform.eu.rep.googleapis.com)
Gemini Enterprise Agent Platform (formerly Vertex AI). Data residency at rest (DRZ) and machine learning processing (MLP) in the EU are documented in Google's locations documentation for Gemini 3.5 Flash and Gemini 2.5 Pro, but not for Gemini 3.6, 3.7 and 3.8 Flash.
License & Hosting
License Proprietary
Security Filters Customizable
Enterprise Support Yes
SLA Available Yes
Cloud Only

Benchmarks

Performance comparison with standardized tests

DeepSWE v1.1 (Gemini 3.8 Flash)
73.7%
Terminal-bench 2.1 (Gemini 3.8 Flash)
89.4%
Terminal-bench 4.0 (Gemini 3.8 Flash)
19.1%
HLE-Verified (Gemini 3.8 Flash)
54.9%
OSWorld 2.0 (Gemini 3.8 Flash)
59.0%
LVBench (Gemini 3.8 Flash)
87.8%
GDPval-AA v2 (Gemini 3.8 Flash)
1545
Artificial Analysis Intelligence Index (Gemini 3.8 Flash)
59
DeepSWE v1.1 (Gemini 3.7 Flash)
65.3
Terminal-bench 2.1 (Gemini 3.7 Flash)
85.8
FrontierCode 1.1 Main (Gemini 3.7 Flash)
43.6
Harvey LAB-AA (Gemini 3.7 Flash)
90.7
LVBench (Gemini 3.7 Flash)
85.4
OSWorld-Verified (Gemini 3.6 Flash)
83

innFactory AI Consulting, based in Rosenheim, Germany, supports enterprises across the DACH region (Germany, Austria, Switzerland) and Europe with GDPR-compliant deployment of Google Gemini. On 2 September 2026, Google released Gemini 3.8 Flash – straight to GA, with no preview phase, together with the access-restricted Gemini 3.8 Flash Cyber. It is the third Flash release within six weeks: on 21 July 2026 Google shipped Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber, on 13 August 2026 Gemini 3.7 Flash. Google describes 3.8 Flash as its most intelligent Flash model to date with “substantial gains from 3.7 Flash”. Gemini 3.5 Pro, by contrast, remains unreleased – as of September 2026 there is still no Pro model above Gemini 3.1 Pro.

EU availability (as of 3 September 2026): With Gemini it pays to read closely, because two Google documents read differently. The Gemini Enterprise release notes of 2 September report Gemini 3.8 Flash as “generally available (GA) in the global, us, and eu regions” – as previously 3.7 Flash (13 August) and 3.6 Flash (18 August, now without an allowlist). The locations and data residency documentation (last updated 2 September 2026), however, lists all three models as “Only available in the global region” – with no data residency at rest and no machine learning processing in us or eu. Both can be reconciled: Gemini 3.8 Flash is selectable in an EU deployment, but per Google’s documentation administrators must acknowledge a warning that traffic is routed to the global endpoint – without regional data residency. There is therefore no EU processing commitment for 3.8, 3.7 or 3.6 Flash. Gemini 3.5 Flash remains the latest Gemini model with full EU data residency (DRZ and MLP), and therefore our recommendation for residency-bound workloads – even though Google itself now calls it a “legacy Flash model”.

Key Strengths

Gemini 3.8 Flash – The Current Flash Model (GA since 2 September 2026)

Gemini 3.8 Flash (gemini-3.8-flash) has been available straight to GA since 2 September 2026 – with no preview and no preview suffix in the model ID. Google positions it as its most intelligent Flash model to date and stresses that it “works harder” on complex tasks: it runs additional reasoning steps and calls tools iteratively. For cost planning that is the most important framing, because the list price is identical to 3.7 Flash while token consumption per task is higher (see below).

Google’s published benchmark figures (model evaluation PDF, as of September 2026, pass@1) compared with Gemini 3.7 Flash:

BenchmarkGemini 3.8 FlashGemini 3.7 Flash
DeepSWE v1.1 (autonomous coding)73.7%65.3%
Terminal-bench 2.189.4%85.8%
Terminal-bench 4.019.1%11.2%
GDPval-AA v2 (Elo)15451482
HLE-Verified54.9%53.6%
OSWorld 2.0 (partial score)59.0%50.6%
LVBench (video understanding, agentic)87.8%85.4%
Vals Finance Agent v261.4%59.0%
Harvey’s Legal Agent Benchmark10.0%8.8%
CharXiv Reasoning86.2%84.5%
LABBench286.2%82.1%

The DeepSWE figure is confirmed by the independent DeepSWE leaderboard (74% ± 1). In the Artificial Analysis Intelligence Index 3.8 Flash scores 59 points (3.7 Flash: 56, 3.6 Flash: 52). The same measurement also shows that 3.8 Flash generates around 30% more output tokens per task than 3.7 Flash, raising cost per task by roughly 40% at the same list price. Anyone migrating from 3.7 to 3.8 Flash should therefore measure cost per task rather than compare token prices.

The context window is 1,048,576 input tokens and 65,536 output tokens; the knowledge cutoff per the model card is March 2026. The model processes text, image, video, audio and PDF; output is text only (no Live API). Reasoning is controlled via the thinking levels low, medium (default) and high; minimal is rejected with an error, and thinking_budget has been replaced by thinking_level. Tooling covers function calling, structured output, code execution, file search, URL context, Google Search grounding, batch, flex and priority inference, and computer use (preview). The agentic video understanding introduced on 1 September 2026 is explicitly listed by Google only for 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite – not for 3.8 Flash.

Pricing (Gemini API): Through 31 December 2026 the same introductory price as for 3.7 and 3.6 Flash applies: $0.75 per 1M input tokens and $3.75 per 1M output tokens (including thinking tokens; caching $0.075; batch $0.375 / $1.875; priority $1.35 / $6.75). From 1 January 2027 list prices rise to $1.50 / $7.50 / $0.15. Per Google the introductory price applies both in AI Studio and on the Gemini Enterprise Agent Platform. There is no context-length tiering; audio input is not priced separately.

Migration note: Besides temperature, top_p and top_k, candidate_count is deprecated as well; FunctionResponse objects require call_id and name. Existing integrations on 3.5 or 3.6 Flash that still send thinking_level: minimal break on the switch.

Gemini 3.7 Flash – Previous Generation (GA since 13 August 2026)

Gemini 3.7 Flash (gemini-3.7-flash) has been generally available since 13 August 2026 – shipped straight to GA with no preview phase. Google positions the model for coding, web development, knowledge work, and agentic workflows, and emphasises that the gains come not from a retraining run but from algorithmic improvements and evaluated developer feedback.

Google’s published benchmark results compared to Gemini 3.6 Flash:

BenchmarkGemini 3.7 FlashGemini 3.6 Flash
DeepSWE v1.1 (autonomous coding)65.3%49.0%
FrontierCode 1.1 Main43.6%34.4%
WebDev Arena (Elo)15881538
Terminal-bench 2.185.8%
Harvey LAB-AA (legal reasoning)90.7%
LVBench (video understanding)85.4%

The context window is 1,048,576 input tokens and 65,536 output tokens. The model processes text, image, video, audio, and PDF; reasoning is controlled via the thinking levels low, medium, and high (there is no “minimal” level). Tooling covers function calling, structured output, code execution, file search, Google Search grounding, and computer use (preview). Relevant for media workloads: since 1 September 2026 agentic video understanding is GA for 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite – at no surcharge, per Google with up to 88% fewer tokens, up to 66% lower cost and up to 7% better accuracy on LongVideoBench.

Pricing (Gemini API): Through 31 December 2026 an introductory price of $0.75 per 1M input tokens and $3.75 per 1M output tokens applies (caching $0.075). From 1 January 2027 list prices rise to $1.50 / $7.50 / $0.15 – anyone building a business case now should factor that doubling in. A limited free tier is available.

The knowledge cutoff per the model card is March 2026. Since 2 September 2026 Google lists 3.7 Flash as previous generation; for new projects without residency requirements 3.8 Flash is the current choice. Migration note: the sampling parameters temperature, top_p, and top_k have been deprecated since 21 July 2026 and this applies to 3.7 Flash as well – review existing integrations.

Gemini 3.6 Flash – Previous-Step Workhorse (GA since 21 July 2026)

Gemini 3.6 Flash (gemini-3.6-flash) improves coding, knowledge work, and multimodal processing over 3.5 Flash. On the DeepSWE coding benchmark the model jumps from 37 to 49 percent, and it reaches 83 percent on OSWorld-Verified (computer use). At the same time it produces 17 percent fewer output tokens than 3.5 Flash – Google directly addressing developer feedback on verbosity, which lowers effective per-task cost further. The knowledge cutoff is March 2026; the context window is 1 million input tokens (65,536 output).

Price correction versus our 24 July status: at launch, 3.6 Flash cost $1.50/1M input and $7.50/1M output. It now sits – like 3.7 Flash – at the introductory price of $0.75/1M input and $3.75/1M output, limited to 31 December 2026. The model is available via the Gemini API (AI Studio, Android Studio), Antigravity, the Gemini Enterprise Agent Platform, and the Gemini app.

Gemini 3.5 Flash-Lite – Speed for High-Volume Workloads (GA since 21 July 2026)

Gemini 3.5 Flash-Lite (gemini-3.5-flash-lite) is the fastest model in the 3.5 series: 350 output tokens per second as measured by Artificial Analysis, priced at $0.30/1M input and $2.50/1M output. That makes it a fit for high-volume automation, classification, and latency-critical applications – for example as a sub-agent in multi-agent pipelines. Google is additionally rolling the model out in Google Search.

Gemini 3.8 Flash Cyber and the Fairwind Program – Security Models with Restricted Access

With Gemini 3.5 Flash Cyber, Google had announced its first dedicated cybersecurity model on 21 July 2026, paired with the CodeMender security agent and exclusive to governments and vetted partners. On 2 September 2026 the successor Gemini 3.8 Flash Cyber followed, provided exclusively through the new Fairwind Program – for governments, critical infrastructure operators and core technology platforms, per Google with over 650 partners at launch. Google reports for 3.8 Flash Cyber, among other figures, 86.2% pass@1 on CyberGym (3.5 Flash Cyber: 77.5%), 71.0% on an internal vulnerability benchmark across 20 programming languages (3.7 Flash: 58.9%) and 47.2% pass@1 on CWE-Bench. For enterprises outside the program the model is therefore not available, but it shows Google’s course of rolling out security-critical capabilities in a controlled way – a pattern we already know from the clearance processes at OpenAI GPT-5.6 and Anthropic Claude Mythos 5.1.

Gemini 3.5 Flash – Agentic-first

With Gemini 3.5 Flash, Google positions its fastest tier above its own Pro flagship for the first time: 3.5 Flash outperforms Gemini 3.1 Pro on Terminal-Bench 2.1 (coding), GDPval-AA Elo (real-world agentic) and MCP Atlas (tool use). The model is built for multi-hour autonomous workflows – coding pipelines, research projects, even building entire systems – pausing for human input at decision points. An optimised variant is reported to reach up to 12× the speed of other frontier models at equivalent quality.

Gemini 3.5 Pro – Orchestrator (Still Delayed)

Gemini 3.5 Pro is still in partner testing in early September 2026, with no new date; the 3.8 Flash launch blog does not mention it. The timeline: announced at Google I/O on 19 May 2026, the June target missed, a reported 17 July target missed as well – on 21 July Google shipped three smaller models instead, followed by two further Flash generations on 13 August and 2 September. Media reports (Bloomberg, Axios, Forbes, 9to5google, all 13 August 2026) attribute the delay to coding performance falling short of internal targets; Google has not confirmed this.

The model is designed as the orchestrator/planner that drives Gemini 3.5/3.6/3.7 Flash instances as sub-agents. Google has announced a 2M token context window and a Deep Think reasoning mode. In practical terms for users: the Pro flagship is still Gemini 3.1 Pro (gemini-3.1-pro-preview, in preview since 19 February 2026, $2/1M input and $12/1M output at up to 200k context). There is no Gemini 3.5, 3.6, 3.7 or 3.8 Pro – the 3.5 to 3.8 generations are Flash-only lines.

Gemini Spark – Agentic 24/7 Assistant

At Google I/O 2026, Google introduced Gemini Spark, an agentic personal assistant; it is now built on Gemini 3.7 Flash. Spark runs continuously on Google’s cloud, can be addressed directly via a dedicated Gmail address, and executes background tasks across Chrome and Google Workspace. Spark supports the Model Context Protocol (MCP) for external integrations.

Relevant for customers in Europe: Google states Spark availability in over 160 countries, but explicitly excludes the EEA, Switzerland, the United Kingdom, and Nigeria. That applies to the end-user product, not the API – Gemini 3.7 Flash itself is usable from the EU via the Gemini API and the Gemini Enterprise Agent Platform.

Very Large Context Window

Gemini 3.1 Pro processes up to 2 million tokens in a single context. This enables analysis of extensive contract portfolios, technical documentation, or entire codebases in a single request. The Flash models offer a 1,048,576 token context window at significantly lower cost.

Native Multimodal Processing

The Gemini 3.x family natively processes text, images, audio, video, and PDF documents. This enables use cases such as automated video analysis, document extraction from scanned PDFs, or meeting evaluation combining audio and visual material.

Google Ecosystem

  • Gemini Enterprise Agent Platform (formerly Vertex AI): Enterprise deployment with SLA and data residency options
  • Antigravity: Agentic development platform for the current Flash models
  • Gemini Enterprise: Enterprise frontend with access to 3.8, 3.7, 3.6, and 3.5 Flash
  • Google Workspace: Integration with Docs, Sheets, Gmail, and more
  • Search Grounding: Access to current web information
  • Google AI Studio: Rapid prototyping and API access

Specialised Variants

  • Gemini 3.1 Flash Thinking: Extended reasoning for complex STEM tasks with a transparent thinking process
  • Gemini 3.1 Pro Deep Research: Specialised for multi-step research and long analysis tasks

The Model Family in September 2026

The Gemini family has become broad. Sorting by maturity helps with selection:

Generally available (GA), with list prices per 1M tokens:

ModelInputOutput
Gemini 3.8 Flash$0.75 (through 31 Dec 2026, then $1.50)$3.75 (then $7.50)
Gemini 3.7 Flash$0.75 (through 31 Dec 2026, then $1.50)$3.75 (then $7.50)
Gemini 3.6 Flash$0.75 (through 31 Dec 2026, then $1.50)$3.75 (then $7.50)
Gemini 3.5 Flash$1.50$9.00
Gemini 3.5 Flash-Lite$0.30$2.50
Gemini 3.1 Flash-Lite$0.25$1.50

The Gemini 2.5 line (Pro, Flash, Flash-Lite) continues as a legacy generation and remains relevant for EU projects because it supports full data residency.

New since late August 2026: Gemini Omni 1.1 Flash (gemini-omni-1.1-flash, GA since 27 August 2026) extends video generation with scene extension in 10-second steps up to 40 seconds, first/last-frame interpolation, a 360p draft mode and 1080p/4K upscaling; billing is by output tokens (5,792 tokens per second of 720p video). The preview gemini-omni-flash-preview shuts down on 30 September 2026. Gemini 3.5 Transcribe (gemini-3.5-transcribe, diarization, word timestamps, vocabulary bias) and Gemini 3.5 Transcribe Live (WebSocket streaming) have been available since 26 August 2026 – the Gemini API changelog calls it GA, Google’s blog a public preview.

In preview: Gemini 3.1 Pro (the Pro flagship), Gemini 3 Flash, Gemini 3.5 Live Translate (over 70 languages), Gemini 3.1 Flash Live, Gemini 3.1 Flash TTS, Deep Research and Deep Research Max, the Antigravity Agent, Gemini Embedding 2, and Gemini Robotics ER 2.

Image models: Nano Banana 2 (gemini-3.1-flash-image), Nano Banana 2 Lite, and Nano Banana Pro (gemini-3-pro-image, up to 4K).

Deep Think: Per the secondary sources available to us, only Gemini 3.1 Deep Think currently exists, with consumer access via AI Ultra. There is no 3.5 Deep Think, because the underlying 3.5 Pro base model does not exist. We are not aware of any verifiable Gemini 3.x Nano line.

Data Residency: Keep DRZ and MLP Apart

The most important point for GDPR projects using Gemini is a distinction that Google’s documentation maintains consistently but that often blurs in marketing material:

  • DRZ – data residency at rest: stored data stays in the chosen region.
  • MLP – machine learning processing: the model’s processing also takes place in the chosen region.

Only both together amount to full EU localisation. A model with only a DRZ commitment may still process prompts outside the EU. This is exactly where the Gemini models differ significantly. The following status comes from Google’s locations documentation:

ModelUS / EU multi-regionIn-country
Gemini 3.7 Flash“Only available in the global region”no
Gemini 3.6 Flashus: DRZ + MLP with allowlist · eu: “Only available in the global region”no
Gemini 3.5 FlashDRZ and MLP supportedIndia, Japan, Singapore, UK; Canada no
Gemini 2.5 ProDRZ and MLP supportedCanada, Japan

Region Model

Google distinguishes three multi-regions: global is the default and offers the best performance and newest features – but no residency guarantee. Alongside it are us and eu. The EU multi-region endpoint is https://aiplatform.eu.rep.googleapis.com. In addition there are in-country regions (GA, each with an allowlist): ca, in, asia-northeast1 (Japan), sg, and europe-west2 (UK).

Zero Data Retention

For the Gemini API / AI Studio, ZDR is available only for “paid services” – there is no ZDR in the free tier. Where ZDR is approved, Google strips prompts, responses, and identifying metadata before logging. Not disableable even with ZDR, however: Google Search grounding (30 days retention) and Google Maps grounding (30 days). Further retention applies by default in the Interactions API (conversation state; requires store: false), the Live API (session state for up to 24 hours), and the File API (until the user deletes the file).

On the Gemini Enterprise Agent Platform, ZDR is achievable per vendor documentation and practitioner reports, but requires active steps: data caching for Google models must be disabled and abuse monitoring logging must be turned off – the latter via opt-out through invoiced billing or as an exception via a request to Google support. For some “advanced AI features”, these sources indicate ZDR is not possible. We treat these statements as secondary-sourced and verify them contractually in each project.

The Gemini API Is Not the Enterprise Platform

The Gemini API / AI Studio is available in over 190 countries and therefore across the entire EU, but carries no EU data residency and no GDPR-specific commitments. Google explicitly points to the Gemini Enterprise Agent Platform for compliance commitments. For production processing of personal data, going directly through AI Studio is therefore usually not the right path.

Structural Lag in EU Regions

New Gemini models arrive first in the global or US regions and reach EU regions with substantial delay – or, as with 3.6, 3.7 and 3.8 Flash, not at all with full MLP. This is not a one-off but a recurring pattern that developers also document in Google’s own developer forum. For architecture planning that means: if you need EU data residency, do not assume you can always use the newest model, and drive model selection through configuration rather than hard-wiring it in code.

Model Lifecycle and Deprecations

Google retires endpoints at a noticeable cadence. For anyone running production deployments, tracking the lifecycle belongs in the operational routine.

Already shut down:

  • gemini-2.0-flash and gemini-2.0-flash-lite – 1 June 2026

  • gemini-3-pro-preview – 9 March 2026

  • gemini-3.1-flash-image-preview and gemini-3-pro-image-preview – 25 June 2026

  • Veo 2.0 and Veo 3.0 endpoints – 30 June 2026

  • imagen-4.0-generate-001 including its Ultra and Fast variants, plus Gemini 3 Image – 17 August 2026

  • Grok 4.1 models on the Agent Platform – 20 August 2026

  • gemini-robotics-er-1.6-preview – 31 August 2026 (successor gemini-robotics-er-2-preview)

Upcoming:

  • gemini-omni-flash-preview30 September 2026 (successor gemini-omni-1.1-flash)
  • gemini-2.5-flash-image2 October 2026
  • Deprecated open model endpoints – 21 October 2026
  • gemini-3.1-flash-lite – deprecated since 7 May 2026, shutdown 7 May 2027 (successor gemini-3.5-flash-lite)
  • Gemini 2.5 Pro / Flash / Flash-Lite: deprecated in the Gemini API, but without a shutdown date. For the Agent Platform, secondary sources point to a retirement in October 2026; we could not verify Google’s lifecycle page directly. Anyone using 2.5 Pro as the EU legacy option should plan the migration now.
  • Gemini 3.5 Flash: listed as deprecated in the Gemini API since 19 May 2026 – without a shutdown date, status still “Stable”.

Parameter deprecation: since 21 July 2026, the sampling parameters temperature, top_p, and top_k have been deprecated (on 3.8 Flash additionally candidate_count). This applies to Gemini 3.7 and 3.8 Flash as well. Older integrations that hard-wire these parameters are a common stumbling block when switching models – we recommend configuring model IDs and request parameters centrally rather than scattering them across application code.

EU Availability

As of 3 September 2026: Gemini 3.5 Flash is the latest Gemini model for which Google documents both data residency at rest and machine learning processing in the EU multi-region (since 31 August 2026 additionally in-country in Canada). Requests are routed inside the EU geography and covered by the Data Processing Addendum. Gemini 3.8 Flash, 3.7 Flash and 3.6 Flash are selectable in EU deployments, but run via the global region – with no residency commitment. Gemini 3.5 Pro remains in partner testing without a date. New: Gemini 3.1 Flash Image (Nano Banana 2) has been GA in the us and eu multi-regions since 31 August 2026.

What works from EU projects today

  • Gemini 3.5 Flash via the EU multi-region endpoint — with DRZ and MLP in the EU, GA on the Gemini Enterprise Agent Platform. Our recommendation for residency-bound workloads.
  • Gemini 2.5 Pro — legacy generation, likewise with DRZ and MLP in the EU multi-region; retirement on the Agent Platform in October 2026 per secondary sources (see lifecycle).
  • Gemini 3.8 Flash, 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite via the global region — for workloads without residency requirements. In an EU deployment the models are selectable; per Google’s documentation administrators acknowledge a warning that traffic goes to the global endpoint.

What is still missing

  • EU machine learning processing for Gemini 3.8 Flash, 3.7 Flash and 3.6 Flash — per the locations documentation (as of 2 September 2026) global region only, even though the release notes list all three models as GA in us and eu
  • Single-region endpoints (europe-west1/3/4) for Gemini 3.8 Flash — not documented by Google
  • Documented EU data residency for Gemini 3.5 Flash-Lite
  • In-country regions inside the EU — for the relevant Gemini models Google lists in-country residency only outside the EU (Canada, India, Japan, Singapore, UK), not in europe-west3 or europe-west4
  • Gemini 3.5 Pro (partner testing, date open) — EU availability to be assessed at GA launch

We recommend monitoring Google’s locations and data residency documentation and the Gemini Enterprise release notes regularly — and, in case of doubt, following the locations documentation, because it carries the residency commitments.

A Note on the EU AI Act

In addition, the enforceable transparency obligations of the EU AI Act have applied since 2 August 2026, with fines of up to 3 percent of global annual turnover. According to the secondary sources available to us, Google has signed the EU AI Act Transparency Code of Practice and marks generated content with SynthID watermarks. For model selection this is an additional checkpoint, but it does not replace the residency assessment.

Integration with CompanyGPT

Gemini models are integrated in CompanyGPT. Gemini 3.5 Flash is wired in via the EU multi-region endpoint for GDPR-compliant reasoning and agentic coding workloads — with data residency at rest and machine learning processing in the EU. Gemini 2.5 Pro is available as a legacy option with the same residency properties. As soon as Google documents EU MLP for a newer Gemini model, we will add it as a default option.

Our Recommendation

Since 2 September 2026, Gemini 3.8 Flash has been Google’s most capable GA model. For model selection in EU projects, however, the deciding factor is not the benchmark table but whether the workload requires EU data residency:

  • With EU residency requirements (personal data, regulated industries): Gemini 3.5 Flash via the EU multi-region endpoint. It is the latest Gemini model with DRZ and MLP in the EU and therefore remains our EU recommendation – even though Google now calls it a legacy Flash model (without a shutdown date).
  • Without residency requirements, focused on coding and agentic workflows: Gemini 3.8 Flash via the global region — per Google clear benchmark gains over 3.7 Flash at the introductory price of $0.75/1M input. Factor two things into your business case: the roughly 30% higher token consumption per task (Artificial Analysis measurement) and the price doubling from 1 January 2027. Gemini 3.7 Flash remains an option when token consumption per task matters more than the last benchmark points.
  • High-volume and latency-critical: Gemini 3.5 Flash-Lite ($0.30/1M input, 350 tokens/s) or Gemini 3.1 Flash-Lite ($0.25/1M input) via the global region.
  • Legacy estate with EU residency: Gemini 2.5 Pro — technically surpassed, but with full EU data residency.
  • Cross-cloud alternative: Anthropic Claude Opus 5 or Claude Sonnet 5 (model page) via AWS Bedrock (in-region Ireland / Stockholm, Frankfurt via EU geo inference) — useful for multi-cloud strategies or when the workload is not bound to Google Cloud. Claude Fable 5.1, released on 1 September 2026, runs with EU data residency via the Europe multi-region of the Agent Platform (global only on Bedrock), but as a covered model is subject to mandatory 30-day retention.

Unsure which option fits your data protection impact assessment? We review the residency and contractual situation in your specific project and track the EU rollout of the Gemini models continuously — we will update this page as soon as Google documents EU MLP for a newer model.

Cost estimation for this model

For up-to-date token pricing, model variants and EU availability, see our sister project ai-prices.eu. It helps you compare and estimate the operational cost of leading AI models for your specific use case.

Compare prices on ai-prices.eu

ai-prices.eu is a project by innFactory AI Consulting GmbH and provides transparent cost estimates for leading AI models.

Frequently Asked Questions

What does Gemini 3.8 Flash cost?

Gemini 3.8 Flash (gemini-3.8-flash, GA since 2 September 2026) costs USD 0.75 per 1 million input tokens and USD 3.75 per 1 million output tokens including thinking tokens in the Gemini API at an introductory rate through 31 December 2026; context caching USD 0.075, batch USD 0.375 / 1.875. From 1 January 2027 the price is USD 1.50 / 7.50. Per Google the introductory price also applies on the Gemini Enterprise Agent Platform.

Can Gemini 3.8 Flash be used with EU data residency?

No, as of 3 September 2026. The Gemini Enterprise release notes list 3.8 Flash as GA in global, us and eu, but Google's locations and data residency documentation lists the model, like 3.7 and 3.6 Flash, as 'Only available in the global region' without data residency at rest and without ML processing in the EU. The latest Gemini model with documented EU data residency (DRZ and MLP) remains Gemini 3.5 Flash.

What is new in Gemini 3.8 Flash compared with 3.7 Flash?

Per Google's evaluation, DeepSWE v1.1 rises from 65.3 to 73.7 percent, Terminal-bench 2.1 from 85.8 to 89.4 percent and HLE-Verified from 53.6 to 54.9 percent. Context window (1,048,576 input, 65,536 output) and list price stay the same. Per Google the model 'works harder' and uses more tokens; Artificial Analysis measures around 30 percent more output tokens per task, i.e. higher cost per task.

Is there a Gemini 3.8 Pro or Gemini 3.5 Pro?

No. Generations 3.5 to 3.8 are Flash-only lines. Gemini 3.5 Pro was announced on 19 May 2026 and is still in partner testing in early September 2026 without a date. The Pro flagship remains Gemini 3.1 Pro (gemini-3.1-pro-preview, USD 2 / 12 per 1M tokens up to 200k context).

Which Gemini model do we recommend for GDPR-bound workloads?

Gemini 3.5 Flash via the EU multi-region endpoint of the Gemini Enterprise Agent Platform: the latest Gemini model for which Google documents data residency at rest and ML processing in the EU (USD 1.50 / 9.00 per 1M tokens). Gemini 2.5 Pro remains a legacy option with EU residency. Gemini 3.8, 3.7 and 3.6 Flash suit workloads without residency requirements via the global region.

What is Gemini 3.8 Flash Cyber?

A cybersecurity model announced on 2 September 2026 that is provided exclusively through Google's Fairwind Program for governments, critical infrastructure operators and core technology platforms (per Google over 650 partners at launch). Google reports, among other figures, 86.2 percent pass@1 on CyberGym. The model is not available to enterprises outside the program.

Consultation for this model?

We help you select and integrate the right AI model for your use case.