innFactory AI Consulting, based in Rosenheim, Germany, supports enterprises across the DACH region (Germany, Austria, Switzerland) and Europe with GDPR-compliant deployment of Google Gemini. On 2 September 2026, Google released Gemini 3.8 Flash – straight to GA, with no preview phase, together with the access-restricted Gemini 3.8 Flash Cyber. It is the third Flash release within six weeks: on 21 July 2026 Google shipped Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber, on 13 August 2026 Gemini 3.7 Flash. Google describes 3.8 Flash as its most intelligent Flash model to date with “substantial gains from 3.7 Flash”. Gemini 3.5 Pro, by contrast, remains unreleased – as of September 2026 there is still no Pro model above Gemini 3.1 Pro.
EU availability (as of 3 September 2026): With Gemini it pays to read closely, because two Google documents read differently. The Gemini Enterprise release notes of 2 September report Gemini 3.8 Flash as “generally available (GA) in the
global,us, andeuregions” – as previously 3.7 Flash (13 August) and 3.6 Flash (18 August, now without an allowlist). The locations and data residency documentation (last updated 2 September 2026), however, lists all three models as “Only available in theglobalregion” – with no data residency at rest and no machine learning processing inusoreu. Both can be reconciled: Gemini 3.8 Flash is selectable in an EU deployment, but per Google’s documentation administrators must acknowledge a warning that traffic is routed to the global endpoint – without regional data residency. There is therefore no EU processing commitment for 3.8, 3.7 or 3.6 Flash. Gemini 3.5 Flash remains the latest Gemini model with full EU data residency (DRZ and MLP), and therefore our recommendation for residency-bound workloads – even though Google itself now calls it a “legacy Flash model”.
Key Strengths
Gemini 3.8 Flash – The Current Flash Model (GA since 2 September 2026)
Gemini 3.8 Flash (gemini-3.8-flash) has been available straight to GA since 2 September 2026 – with no preview and no preview suffix in the model ID. Google positions it as its most intelligent Flash model to date and stresses that it “works harder” on complex tasks: it runs additional reasoning steps and calls tools iteratively. For cost planning that is the most important framing, because the list price is identical to 3.7 Flash while token consumption per task is higher (see below).
Google’s published benchmark figures (model evaluation PDF, as of September 2026, pass@1) compared with Gemini 3.7 Flash:
| Benchmark | Gemini 3.8 Flash | Gemini 3.7 Flash |
|---|---|---|
| DeepSWE v1.1 (autonomous coding) | 73.7% | 65.3% |
| Terminal-bench 2.1 | 89.4% | 85.8% |
| Terminal-bench 4.0 | 19.1% | 11.2% |
| GDPval-AA v2 (Elo) | 1545 | 1482 |
| HLE-Verified | 54.9% | 53.6% |
| OSWorld 2.0 (partial score) | 59.0% | 50.6% |
| LVBench (video understanding, agentic) | 87.8% | 85.4% |
| Vals Finance Agent v2 | 61.4% | 59.0% |
| Harvey’s Legal Agent Benchmark | 10.0% | 8.8% |
| CharXiv Reasoning | 86.2% | 84.5% |
| LABBench2 | 86.2% | 82.1% |
The DeepSWE figure is confirmed by the independent DeepSWE leaderboard (74% ± 1). In the Artificial Analysis Intelligence Index 3.8 Flash scores 59 points (3.7 Flash: 56, 3.6 Flash: 52). The same measurement also shows that 3.8 Flash generates around 30% more output tokens per task than 3.7 Flash, raising cost per task by roughly 40% at the same list price. Anyone migrating from 3.7 to 3.8 Flash should therefore measure cost per task rather than compare token prices.
The context window is 1,048,576 input tokens and 65,536 output tokens; the knowledge cutoff per the model card is March 2026. The model processes text, image, video, audio and PDF; output is text only (no Live API). Reasoning is controlled via the thinking levels low, medium (default) and high; minimal is rejected with an error, and thinking_budget has been replaced by thinking_level. Tooling covers function calling, structured output, code execution, file search, URL context, Google Search grounding, batch, flex and priority inference, and computer use (preview). The agentic video understanding introduced on 1 September 2026 is explicitly listed by Google only for 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite – not for 3.8 Flash.
Pricing (Gemini API): Through 31 December 2026 the same introductory price as for 3.7 and 3.6 Flash applies: $0.75 per 1M input tokens and $3.75 per 1M output tokens (including thinking tokens; caching $0.075; batch $0.375 / $1.875; priority $1.35 / $6.75). From 1 January 2027 list prices rise to $1.50 / $7.50 / $0.15. Per Google the introductory price applies both in AI Studio and on the Gemini Enterprise Agent Platform. There is no context-length tiering; audio input is not priced separately.
Migration note: Besides temperature, top_p and top_k, candidate_count is deprecated as well; FunctionResponse objects require call_id and name. Existing integrations on 3.5 or 3.6 Flash that still send thinking_level: minimal break on the switch.
Gemini 3.7 Flash – Previous Generation (GA since 13 August 2026)
Gemini 3.7 Flash (gemini-3.7-flash) has been generally available since 13 August 2026 – shipped straight to GA with no preview phase. Google positions the model for coding, web development, knowledge work, and agentic workflows, and emphasises that the gains come not from a retraining run but from algorithmic improvements and evaluated developer feedback.
Google’s published benchmark results compared to Gemini 3.6 Flash:
| Benchmark | Gemini 3.7 Flash | Gemini 3.6 Flash |
|---|---|---|
| DeepSWE v1.1 (autonomous coding) | 65.3% | 49.0% |
| FrontierCode 1.1 Main | 43.6% | 34.4% |
| WebDev Arena (Elo) | 1588 | 1538 |
| Terminal-bench 2.1 | 85.8% | — |
| Harvey LAB-AA (legal reasoning) | 90.7% | — |
| LVBench (video understanding) | 85.4% | — |
The context window is 1,048,576 input tokens and 65,536 output tokens. The model processes text, image, video, audio, and PDF; reasoning is controlled via the thinking levels low, medium, and high (there is no “minimal” level). Tooling covers function calling, structured output, code execution, file search, Google Search grounding, and computer use (preview). Relevant for media workloads: since 1 September 2026 agentic video understanding is GA for 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite – at no surcharge, per Google with up to 88% fewer tokens, up to 66% lower cost and up to 7% better accuracy on LongVideoBench.
Pricing (Gemini API): Through 31 December 2026 an introductory price of $0.75 per 1M input tokens and $3.75 per 1M output tokens applies (caching $0.075). From 1 January 2027 list prices rise to $1.50 / $7.50 / $0.15 – anyone building a business case now should factor that doubling in. A limited free tier is available.
The knowledge cutoff per the model card is March 2026. Since 2 September 2026 Google lists 3.7 Flash as previous generation; for new projects without residency requirements 3.8 Flash is the current choice. Migration note: the sampling parameters temperature, top_p, and top_k have been deprecated since 21 July 2026 and this applies to 3.7 Flash as well – review existing integrations.
Gemini 3.6 Flash – Previous-Step Workhorse (GA since 21 July 2026)
Gemini 3.6 Flash (gemini-3.6-flash) improves coding, knowledge work, and multimodal processing over 3.5 Flash. On the DeepSWE coding benchmark the model jumps from 37 to 49 percent, and it reaches 83 percent on OSWorld-Verified (computer use). At the same time it produces 17 percent fewer output tokens than 3.5 Flash – Google directly addressing developer feedback on verbosity, which lowers effective per-task cost further. The knowledge cutoff is March 2026; the context window is 1 million input tokens (65,536 output).
Price correction versus our 24 July status: at launch, 3.6 Flash cost $1.50/1M input and $7.50/1M output. It now sits – like 3.7 Flash – at the introductory price of $0.75/1M input and $3.75/1M output, limited to 31 December 2026. The model is available via the Gemini API (AI Studio, Android Studio), Antigravity, the Gemini Enterprise Agent Platform, and the Gemini app.
Gemini 3.5 Flash-Lite – Speed for High-Volume Workloads (GA since 21 July 2026)
Gemini 3.5 Flash-Lite (gemini-3.5-flash-lite) is the fastest model in the 3.5 series: 350 output tokens per second as measured by Artificial Analysis, priced at $0.30/1M input and $2.50/1M output. That makes it a fit for high-volume automation, classification, and latency-critical applications – for example as a sub-agent in multi-agent pipelines. Google is additionally rolling the model out in Google Search.
Gemini 3.8 Flash Cyber and the Fairwind Program – Security Models with Restricted Access
With Gemini 3.5 Flash Cyber, Google had announced its first dedicated cybersecurity model on 21 July 2026, paired with the CodeMender security agent and exclusive to governments and vetted partners. On 2 September 2026 the successor Gemini 3.8 Flash Cyber followed, provided exclusively through the new Fairwind Program – for governments, critical infrastructure operators and core technology platforms, per Google with over 650 partners at launch. Google reports for 3.8 Flash Cyber, among other figures, 86.2% pass@1 on CyberGym (3.5 Flash Cyber: 77.5%), 71.0% on an internal vulnerability benchmark across 20 programming languages (3.7 Flash: 58.9%) and 47.2% pass@1 on CWE-Bench. For enterprises outside the program the model is therefore not available, but it shows Google’s course of rolling out security-critical capabilities in a controlled way – a pattern we already know from the clearance processes at OpenAI GPT-5.6 and Anthropic Claude Mythos 5.1.
Gemini 3.5 Flash – Agentic-first
With Gemini 3.5 Flash, Google positions its fastest tier above its own Pro flagship for the first time: 3.5 Flash outperforms Gemini 3.1 Pro on Terminal-Bench 2.1 (coding), GDPval-AA Elo (real-world agentic) and MCP Atlas (tool use). The model is built for multi-hour autonomous workflows – coding pipelines, research projects, even building entire systems – pausing for human input at decision points. An optimised variant is reported to reach up to 12× the speed of other frontier models at equivalent quality.
Gemini 3.5 Pro – Orchestrator (Still Delayed)
Gemini 3.5 Pro is still in partner testing in early September 2026, with no new date; the 3.8 Flash launch blog does not mention it. The timeline: announced at Google I/O on 19 May 2026, the June target missed, a reported 17 July target missed as well – on 21 July Google shipped three smaller models instead, followed by two further Flash generations on 13 August and 2 September. Media reports (Bloomberg, Axios, Forbes, 9to5google, all 13 August 2026) attribute the delay to coding performance falling short of internal targets; Google has not confirmed this.
The model is designed as the orchestrator/planner that drives Gemini 3.5/3.6/3.7 Flash instances as sub-agents. Google has announced a 2M token context window and a Deep Think reasoning mode. In practical terms for users: the Pro flagship is still Gemini 3.1 Pro (gemini-3.1-pro-preview, in preview since 19 February 2026, $2/1M input and $12/1M output at up to 200k context). There is no Gemini 3.5, 3.6, 3.7 or 3.8 Pro – the 3.5 to 3.8 generations are Flash-only lines.
Gemini Spark – Agentic 24/7 Assistant
At Google I/O 2026, Google introduced Gemini Spark, an agentic personal assistant; it is now built on Gemini 3.7 Flash. Spark runs continuously on Google’s cloud, can be addressed directly via a dedicated Gmail address, and executes background tasks across Chrome and Google Workspace. Spark supports the Model Context Protocol (MCP) for external integrations.
Relevant for customers in Europe: Google states Spark availability in over 160 countries, but explicitly excludes the EEA, Switzerland, the United Kingdom, and Nigeria. That applies to the end-user product, not the API – Gemini 3.7 Flash itself is usable from the EU via the Gemini API and the Gemini Enterprise Agent Platform.
Very Large Context Window
Gemini 3.1 Pro processes up to 2 million tokens in a single context. This enables analysis of extensive contract portfolios, technical documentation, or entire codebases in a single request. The Flash models offer a 1,048,576 token context window at significantly lower cost.
Native Multimodal Processing
The Gemini 3.x family natively processes text, images, audio, video, and PDF documents. This enables use cases such as automated video analysis, document extraction from scanned PDFs, or meeting evaluation combining audio and visual material.
Google Ecosystem
- Gemini Enterprise Agent Platform (formerly Vertex AI): Enterprise deployment with SLA and data residency options
- Antigravity: Agentic development platform for the current Flash models
- Gemini Enterprise: Enterprise frontend with access to 3.8, 3.7, 3.6, and 3.5 Flash
- Google Workspace: Integration with Docs, Sheets, Gmail, and more
- Search Grounding: Access to current web information
- Google AI Studio: Rapid prototyping and API access
Specialised Variants
- Gemini 3.1 Flash Thinking: Extended reasoning for complex STEM tasks with a transparent thinking process
- Gemini 3.1 Pro Deep Research: Specialised for multi-step research and long analysis tasks
The Model Family in September 2026
The Gemini family has become broad. Sorting by maturity helps with selection:
Generally available (GA), with list prices per 1M tokens:
| Model | Input | Output |
|---|---|---|
| Gemini 3.8 Flash | $0.75 (through 31 Dec 2026, then $1.50) | $3.75 (then $7.50) |
| Gemini 3.7 Flash | $0.75 (through 31 Dec 2026, then $1.50) | $3.75 (then $7.50) |
| Gemini 3.6 Flash | $0.75 (through 31 Dec 2026, then $1.50) | $3.75 (then $7.50) |
| Gemini 3.5 Flash | $1.50 | $9.00 |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 |
The Gemini 2.5 line (Pro, Flash, Flash-Lite) continues as a legacy generation and remains relevant for EU projects because it supports full data residency.
New since late August 2026: Gemini Omni 1.1 Flash (gemini-omni-1.1-flash, GA since 27 August 2026) extends video generation with scene extension in 10-second steps up to 40 seconds, first/last-frame interpolation, a 360p draft mode and 1080p/4K upscaling; billing is by output tokens (5,792 tokens per second of 720p video). The preview gemini-omni-flash-preview shuts down on 30 September 2026. Gemini 3.5 Transcribe (gemini-3.5-transcribe, diarization, word timestamps, vocabulary bias) and Gemini 3.5 Transcribe Live (WebSocket streaming) have been available since 26 August 2026 – the Gemini API changelog calls it GA, Google’s blog a public preview.
In preview: Gemini 3.1 Pro (the Pro flagship), Gemini 3 Flash, Gemini 3.5 Live Translate (over 70 languages), Gemini 3.1 Flash Live, Gemini 3.1 Flash TTS, Deep Research and Deep Research Max, the Antigravity Agent, Gemini Embedding 2, and Gemini Robotics ER 2.
Image models: Nano Banana 2 (gemini-3.1-flash-image), Nano Banana 2 Lite, and Nano Banana Pro (gemini-3-pro-image, up to 4K).
Deep Think: Per the secondary sources available to us, only Gemini 3.1 Deep Think currently exists, with consumer access via AI Ultra. There is no 3.5 Deep Think, because the underlying 3.5 Pro base model does not exist. We are not aware of any verifiable Gemini 3.x Nano line.
Data Residency: Keep DRZ and MLP Apart
The most important point for GDPR projects using Gemini is a distinction that Google’s documentation maintains consistently but that often blurs in marketing material:
- DRZ – data residency at rest: stored data stays in the chosen region.
- MLP – machine learning processing: the model’s processing also takes place in the chosen region.
Only both together amount to full EU localisation. A model with only a DRZ commitment may still process prompts outside the EU. This is exactly where the Gemini models differ significantly. The following status comes from Google’s locations documentation:
| Model | US / EU multi-region | In-country |
|---|---|---|
| Gemini 3.7 Flash | “Only available in the global region” | no |
| Gemini 3.6 Flash | us: DRZ + MLP with allowlist · eu: “Only available in the global region” | no |
| Gemini 3.5 Flash | DRZ and MLP supported | India, Japan, Singapore, UK; Canada no |
| Gemini 2.5 Pro | DRZ and MLP supported | Canada, Japan |
Region Model
Google distinguishes three multi-regions: global is the default and offers the best performance and newest features – but no residency guarantee. Alongside it are us and eu. The EU multi-region endpoint is https://aiplatform.eu.rep.googleapis.com. In addition there are in-country regions (GA, each with an allowlist): ca, in, asia-northeast1 (Japan), sg, and europe-west2 (UK).
Zero Data Retention
For the Gemini API / AI Studio, ZDR is available only for “paid services” – there is no ZDR in the free tier. Where ZDR is approved, Google strips prompts, responses, and identifying metadata before logging. Not disableable even with ZDR, however: Google Search grounding (30 days retention) and Google Maps grounding (30 days). Further retention applies by default in the Interactions API (conversation state; requires store: false), the Live API (session state for up to 24 hours), and the File API (until the user deletes the file).
On the Gemini Enterprise Agent Platform, ZDR is achievable per vendor documentation and practitioner reports, but requires active steps: data caching for Google models must be disabled and abuse monitoring logging must be turned off – the latter via opt-out through invoiced billing or as an exception via a request to Google support. For some “advanced AI features”, these sources indicate ZDR is not possible. We treat these statements as secondary-sourced and verify them contractually in each project.
The Gemini API Is Not the Enterprise Platform
The Gemini API / AI Studio is available in over 190 countries and therefore across the entire EU, but carries no EU data residency and no GDPR-specific commitments. Google explicitly points to the Gemini Enterprise Agent Platform for compliance commitments. For production processing of personal data, going directly through AI Studio is therefore usually not the right path.
Structural Lag in EU Regions
New Gemini models arrive first in the global or US regions and reach EU regions with substantial delay – or, as with 3.6, 3.7 and 3.8 Flash, not at all with full MLP. This is not a one-off but a recurring pattern that developers also document in Google’s own developer forum. For architecture planning that means: if you need EU data residency, do not assume you can always use the newest model, and drive model selection through configuration rather than hard-wiring it in code.
Model Lifecycle and Deprecations
Google retires endpoints at a noticeable cadence. For anyone running production deployments, tracking the lifecycle belongs in the operational routine.
Already shut down:
gemini-2.0-flashandgemini-2.0-flash-lite– 1 June 2026gemini-3-pro-preview– 9 March 2026gemini-3.1-flash-image-previewandgemini-3-pro-image-preview– 25 June 2026Veo 2.0 and Veo 3.0 endpoints – 30 June 2026
imagen-4.0-generate-001including its Ultra and Fast variants, plus Gemini 3 Image – 17 August 2026Grok 4.1 models on the Agent Platform – 20 August 2026
gemini-robotics-er-1.6-preview– 31 August 2026 (successorgemini-robotics-er-2-preview)
Upcoming:
gemini-omni-flash-preview– 30 September 2026 (successorgemini-omni-1.1-flash)gemini-2.5-flash-image– 2 October 2026- Deprecated open model endpoints – 21 October 2026
gemini-3.1-flash-lite– deprecated since 7 May 2026, shutdown 7 May 2027 (successorgemini-3.5-flash-lite)- Gemini 2.5 Pro / Flash / Flash-Lite: deprecated in the Gemini API, but without a shutdown date. For the Agent Platform, secondary sources point to a retirement in October 2026; we could not verify Google’s lifecycle page directly. Anyone using 2.5 Pro as the EU legacy option should plan the migration now.
- Gemini 3.5 Flash: listed as deprecated in the Gemini API since 19 May 2026 – without a shutdown date, status still “Stable”.
Parameter deprecation: since 21 July 2026, the sampling parameters temperature, top_p, and top_k have been deprecated (on 3.8 Flash additionally candidate_count). This applies to Gemini 3.7 and 3.8 Flash as well. Older integrations that hard-wire these parameters are a common stumbling block when switching models – we recommend configuring model IDs and request parameters centrally rather than scattering them across application code.
EU Availability
As of 3 September 2026: Gemini 3.5 Flash is the latest Gemini model for which Google documents both data residency at rest and machine learning processing in the EU multi-region (since 31 August 2026 additionally in-country in Canada). Requests are routed inside the EU geography and covered by the Data Processing Addendum. Gemini 3.8 Flash, 3.7 Flash and 3.6 Flash are selectable in EU deployments, but run via the global region – with no residency commitment. Gemini 3.5 Pro remains in partner testing without a date. New: Gemini 3.1 Flash Image (Nano Banana 2) has been GA in the
usandeumulti-regions since 31 August 2026.
What works from EU projects today
- Gemini 3.5 Flash via the EU multi-region endpoint — with DRZ and MLP in the EU, GA on the Gemini Enterprise Agent Platform. Our recommendation for residency-bound workloads.
- Gemini 2.5 Pro — legacy generation, likewise with DRZ and MLP in the EU multi-region; retirement on the Agent Platform in October 2026 per secondary sources (see lifecycle).
- Gemini 3.8 Flash, 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite via the global region — for workloads without residency requirements. In an EU deployment the models are selectable; per Google’s documentation administrators acknowledge a warning that traffic goes to the global endpoint.
What is still missing
- EU machine learning processing for Gemini 3.8 Flash, 3.7 Flash and 3.6 Flash — per the locations documentation (as of 2 September 2026) global region only, even though the release notes list all three models as GA in
usandeu - Single-region endpoints (europe-west1/3/4) for Gemini 3.8 Flash — not documented by Google
- Documented EU data residency for Gemini 3.5 Flash-Lite
- In-country regions inside the EU — for the relevant Gemini models Google lists in-country residency only outside the EU (Canada, India, Japan, Singapore, UK), not in
europe-west3oreurope-west4 - Gemini 3.5 Pro (partner testing, date open) — EU availability to be assessed at GA launch
We recommend monitoring Google’s locations and data residency documentation and the Gemini Enterprise release notes regularly — and, in case of doubt, following the locations documentation, because it carries the residency commitments.
A Note on the EU AI Act
In addition, the enforceable transparency obligations of the EU AI Act have applied since 2 August 2026, with fines of up to 3 percent of global annual turnover. According to the secondary sources available to us, Google has signed the EU AI Act Transparency Code of Practice and marks generated content with SynthID watermarks. For model selection this is an additional checkpoint, but it does not replace the residency assessment.
Integration with CompanyGPT
Gemini models are integrated in CompanyGPT. Gemini 3.5 Flash is wired in via the EU multi-region endpoint for GDPR-compliant reasoning and agentic coding workloads — with data residency at rest and machine learning processing in the EU. Gemini 2.5 Pro is available as a legacy option with the same residency properties. As soon as Google documents EU MLP for a newer Gemini model, we will add it as a default option.
Our Recommendation
Since 2 September 2026, Gemini 3.8 Flash has been Google’s most capable GA model. For model selection in EU projects, however, the deciding factor is not the benchmark table but whether the workload requires EU data residency:
- With EU residency requirements (personal data, regulated industries): Gemini 3.5 Flash via the EU multi-region endpoint. It is the latest Gemini model with DRZ and MLP in the EU and therefore remains our EU recommendation – even though Google now calls it a legacy Flash model (without a shutdown date).
- Without residency requirements, focused on coding and agentic workflows: Gemini 3.8 Flash via the global region — per Google clear benchmark gains over 3.7 Flash at the introductory price of $0.75/1M input. Factor two things into your business case: the roughly 30% higher token consumption per task (Artificial Analysis measurement) and the price doubling from 1 January 2027. Gemini 3.7 Flash remains an option when token consumption per task matters more than the last benchmark points.
- High-volume and latency-critical: Gemini 3.5 Flash-Lite ($0.30/1M input, 350 tokens/s) or Gemini 3.1 Flash-Lite ($0.25/1M input) via the global region.
- Legacy estate with EU residency: Gemini 2.5 Pro — technically surpassed, but with full EU data residency.
- Cross-cloud alternative: Anthropic Claude Opus 5 or Claude Sonnet 5 (model page) via AWS Bedrock (in-region Ireland / Stockholm, Frankfurt via EU geo inference) — useful for multi-cloud strategies or when the workload is not bound to Google Cloud. Claude Fable 5.1, released on 1 September 2026, runs with EU data residency via the Europe multi-region of the Agent Platform (global only on Bedrock), but as a covered model is subject to mandatory 30-day retention.
Unsure which option fits your data protection impact assessment? We review the residency and contractual situation in your specific project and track the EU rollout of the Gemini models continuously — we will update this page as soon as Google documents EU MLP for a newer model.
