Skip to main content
9 – 17 UHR +49 8031 3508270 LUITPOLDSTR. 9, 83022 ROSENHEIM
DE / EN
LLM DeepSeek China

DeepSeek

DeepSeek V4-Pro GA (13 Aug 2026), V4-Flash-0731 and the vision model: 1M context, peak/off-peak pricing since 16 Aug 2026, EU hosting and our recommendation. As of 3 September 2026.

License MIT (Code), Model Agreement (V3), MIT (R1, V4-Flash repository)
GDPR Hosting Available
Context 128K-1M Tokens
Modality Text, Image, Code → Text, Code

Versions

Overview of available model variants

ModelReleaseEUStrengthsWeaknessesStatus
DeepSeek-V4-Pro-0813 (GA)
13 August 2026
General availability since 13 Aug 2026 in app, web (Expert Mode) and API – model ID unchanged Reasoning effort in three levels (low/high/max) Native support for the Responses API format, adapted for Codex Per DeepSeek: Terminal-Bench 2.1 87.9, NL2Repo 61.5, DeepSWE 62.7, HLE 42.7 / 60.0 (without / with tools) 1M token context, up to 384K output tokens
API prices since 16 Aug 2026 considerably higher than before (see pricing overview) Open weights for the 0813 build not mentioned in the changelog – not confirmed by DeepSeek Extremely high resource requirements for self-hosting
Current
DeepSeek-V4-Flash-0731 Recommended
31 July 2026
Same base as V4-Flash (284B parameters, ~13B active) – per DeepSeek only re-post-trained, with a focus on agent tasks Per DeepSeek: Terminal-Bench 2.1 82.7, NL2Repo 54.2, DeepSWE 54.4, DSBench-Hard 59.6 Native support for the Responses API format, adapted for Codex Reasoning effort low/high/max, 1M token context, up to 384K output tokens Still the lower-priced V4 variant
Price increase as of 16 Aug 2026 (see pricing overview) Open weights for the 0731 build not mentioned in the changelog – not confirmed by DeepSeek
Current
DeepSeek-V4-Flash-Vision-Exp
21 August 2026
Experimental multimodal model: image input via base64, URL or Files API Per DeepSeek on par with V4-Flash for pure text tasks Per DeepSeek: Terminal-Bench 2.1 83.9, NL2Repo 57.7, DeepSWE 59.3, DSBench-Hard 63.6, Chartography 64.3 1M token context, up to 384K output tokens; Chat Completions, Messages and Responses API
Explicitly labelled experimental – no commitments on stability or lifetime Thinking mode not listed for this model in DeepSeek's pricing overview No open weights announced
Preview
DeepSeek-V4-Pro (Preview, April 2026)
24 April 2026
1.6 trillion parameters (~49B active – MoE) 1M token context window (Compressed Sparse Attention + Heavily Compressed Attention) Thinking and non-thinking mode Open weights on HuggingFace Available in Microsoft Foundry since May 2026
Extremely high resource requirements for self-hosting EU region on AWS Bedrock not yet confirmed (US-first rollout) Replaced on the API by the 0813 build since 13 Aug 2026
Current
DeepSeek-V4-Flash (Preview, April 2026)
24 April 2026
284B parameters (~13B active – MoE) 1M token context window Open weights on HuggingFace (MIT License) Cost-efficient alternative to V4-Pro Available in Microsoft Foundry since May 2026
EU region on AWS Bedrock not yet confirmed Replaced on the API by the 0731 build since 31 Jul 2026
Current
DeepSeek-V3.2
December 2025
Current generation Open source (Model Agreement) Now available on AWS, Azure, Vertex AI
Resource intensive
Current
DeepSeek-V3.1
2025
Stable Available on AWS Bedrock EU
Current
DeepSeek-R1
January 2025
Reasoning focus MIT License
Current

Use Cases

Typical applications for this model

Coding & Software Development
Mathematics & Science
Reasoning Tasks
Research & Development
Self-Hosted Deployments
Agentic Workflows

Technical Details

API, features and capabilities

API & Availability
Availability Public
Latency (TTFT) ~800ms
Features & Capabilities
Tool Use Function Calling Structured Output Vision Reasoning Mode File Upload
Training & Knowledge
Knowledge Cutoff 2025 (V4)
Fine-Tuning Available (LoRA, Full, PEFT)
Language Support
Best Quality English, Chinese
Supported 50+ languages
Best quality in English and Chinese, good quality in German

Hosting & Compliance

GDPR-compliant hosting options and licensing

GDPR-Compliant Hosting Options
AWS
Frankfurt (eu-central-1)
Amazon Bedrock - V3.1/V3.2 available
Azure
West Europe
Microsoft Foundry - V3/R1 plus V4-Flash/V4-Pro (since May 2026)
Google Cloud
Frankfurt (europe-west3)
Vertex AI - V3.2/R1 available
Self-Hosted
Own Infrastructure
Open source - full control
License & Hosting
License MIT (Code), Model Agreement (V3), MIT (R1, V4-Flash repository)
Security Filters Customizable
On-Premise

Benchmarks

Performance comparison with standardized tests

Terminal-Bench 2.1 (V4-Pro-0813)
87.9
Terminal-Bench 2.1 (V4-Flash-0731)
82.7
Terminal-Bench 2.1 (V4-Flash-Vision-Exp)
83.9
NL2Repo (V4-Pro-0813)
61.5
DeepSWE (V4-Pro-0813)
62.7
HLE with tools (V4-Pro-0813)
60.0
SWE-Bench (V4-Pro, April 2026)
80.6%

Update September 2026: The summer brought three changes. On 31 July 2026 DeepSeek moved deepseek-v4-flash to the V4-Flash-0731 build (same architecture, re-post-trained for agent tasks). On 13 August 2026 V4-Pro became generally available – in app, web and API, with reasoning effort low/high/max and native Responses API format. Since 21 August 2026 there is also the experimental vision model deepseek-v4-flash-vision-exp. In parallel DeepSeek raised its API prices as of 16 August 2026 and introduced a peak/off-peak model. Details in the “Summer 2026 updates” section. As of 3 September 2026.

Update June 2026: Since May 2026, DeepSeek V4-Flash and V4-Pro are also available in Microsoft Foundry – making the V4 generation usable on a hyperscaler with EU data residency for the first time. On AWS Bedrock, EU regions continue to offer V3.1, V3.2 and R1; V4-Pro typically rolls out to US regions first. innFactory AI Consulting from Germany advises on all deployment options.

Update April 2026: On 24 April 2026, DeepSeek released the V4 generation. V4-Flash (284B) and V4-Pro (1.6T parameters) offer 1M token context via a new hybrid attention mechanism (Compressed Sparse Attention + Heavily Compressed Attention). Both models are available as open weights on HuggingFace.

Summer 2026 updates: V4-Flash-0731, V4-Pro GA and Vision-Exp

V4-Flash-0731 (31 July 2026)

  • Model ID unchanged deepseek-v4-flash; per the DeepSeek documentation updated to the DeepSeek-V4-Flash-0731 build
  • Architecture and size unchanged (284B parameters, ~13B active); per DeepSeek re-post-trained only
  • Agent benchmarks per DeepSeek: Terminal-Bench 2.1 82.7, NL2Repo 54.2, DeepSWE 54.4, DSBench-Hard 59.6
  • Native support for the Responses API format, specifically adapted for Codex
  • DeepSeek’s changelog does not say whether the 0731 weights have been published on Hugging Face – so we do not treat that as given. The V4-Flash repository released in April remains under the MIT License.

V4-Pro GA (13 August 2026)

  • General availability in app, web (“Expert Mode”) and API; model ID stays deepseek-v4-pro, listed in the documentation as DeepSeek-V4-Pro-0813
  • Reasoning effort in three levels: low for simple tasks, high for everyday agent work, max for complex tasks – also applies to V4-Flash
  • Native Responses API format with Codex integration
  • Agent benchmarks per DeepSeek: Terminal-Bench 2.1 87.9, NL2Repo 61.5, DeepSWE 62.7, HLE 42.7 (without tools) / 60.0 (with tools)
  • Open weights for the 0813 build: not confirmed by DeepSeek

V4-Flash-Vision-Exp (21 August 2026)

  • Model ID deepseek-v4-flash-vision-exp, explicitly experimental
  • Image input via base64, URL or Files API; Chat Completions, Messages and Responses API
  • 1M token context, up to 384K output tokens (per pricing overview); a thinking mode is not listed there for this model
  • Per DeepSeek on par with V4-Flash for pure text tasks (agent, reasoning, world knowledge); benchmarks per DeepSeek: Terminal-Bench 2.1 83.9, NL2Repo 57.7, DeepSWE 59.3, DSBench-Hard 63.6, Chartography 64.3
  • Billed at V4-Flash rates; images are converted into tokens depending on their dimensions (max. 384 tokens per image) and charged as input tokens

API pricing since 16 August 2026 (16:00 UTC)

DeepSeek now bills by time of day. Peak hours are Monday to Friday 01:00–04:00 and 06:00–10:00 UTC (CEST: 03:00–06:00 and 08:00–12:00); all other hours are off-peak with a 50% discount.

ModelCache hitCache miss (input)Output
V4-Flash / Vision-Exp – peakUSD 0.014USD 0.44USD 1.32
V4-Flash / Vision-Exp – off-peakUSD 0.007USD 0.22USD 0.66
V4-Pro – peakUSD 0.044USD 1.32USD 3.96
V4-Pro – off-peakUSD 0.022USD 0.66USD 1.98

All figures in USD per 1M tokens, source: DeepSeek pricing overview, as of 3 September 2026. According to secondary sources V4-Flash was previously billed at a flat USD 0.14 input and USD 0.28 output per 1M tokens; DeepSeek itself does not list the old prices in its changelog. For batch and agent workloads it pays to shift into off-peak hours – from a European perspective that means afternoon, evening and night.

DeepSeek V4 - The New Generation (April 2026)

DeepSeek has made a significant leap with the V4 generation:

V4-Flash

  • 284B parameters total, ~13B active (MoE)
  • 1M token context window
  • Thinking and non-thinking mode
  • API: deepseek-v4-flash (0731 build since 31 Jul 2026, see above)
  • Open weights on HuggingFace (MIT License)

V4-Pro

  • 1.6 trillion parameters total, ~49B active (MoE)
  • 1M token context window – only 27% of the FLOPs and 10% of the KV cache compared to V3.2
  • Thinking and non-thinking mode
  • API: deepseek-v4-pro (GA build 0813 since 13 Aug 2026, see above)
  • 80.6% SWE-Bench (per DeepSeek, April 2026)

Note: The previous API names deepseek-chat and deepseek-reasoner were discontinued on 24 July 2026. The pricing overview now lists only deepseek-v4-flash, deepseek-v4-pro and deepseek-v4-flash-vision-exp.

Key Strengths

Open Source & Licensing

DeepSeek offers full transparency:

  • Public Weights: Fully available on GitHub/Hugging Face
  • Licensing: R1 under MIT, V3 under a separate Model Agreement
  • Community: Active development
  • Customizable: Fine-tuning and modifications possible

MoE Architecture

DeepSeek uses innovative Mixture-of-Experts:

  • 671B parameters total, but only 37B active per request
  • Efficient: High performance with reduced resource requirements
  • Multihead Latent Attention: New attention mechanism

Reasoning Capabilities (R1)

DeepSeek-R1 shows transparent thinking processes:

  • Chain-of-thought is made visible
  • Particularly strong in mathematics and logic
  • Comparable to OpenAI o1

EU Availability (as of 3 September 2026)

DeepSeek is available through all three major cloud providers in EU regions. The summer builds 0731 and 0813 initially apply to the direct DeepSeek API; to our knowledge AWS, Microsoft and Google have not confirmed whether or when their hyperscaler deployments will follow.

AWS Bedrock

  • Regions: Frankfurt (eu-central-1), Ireland (eu-west-1)
  • Models: DeepSeek-V3.1, V3.2
  • Advantage: Serverless, immediate availability

Microsoft Foundry (formerly Azure AI Foundry)

  • Regions: West Europe, Sweden Central
  • Models: V3, R1, V4-Flash and V4-Pro (since May 2026)
  • Advantage: Azure ecosystem integration, now with the V4 generation

Google Vertex AI

  • Regions: Frankfurt (europe-west3), Netherlands (europe-west4)
  • Models: V3.2, R1
  • Advantage: Vertex AI Model Garden

Self-Hosting

Still available for maximum control and full GDPR compliance.

Important Notes

Data Privacy Considerations

Update February 2026: With availability on AWS Bedrock, Azure AI, and Google Vertex AI in EU regions, enterprises can now use DeepSeek GDPR-compliant in the cloud!

  • Cloud Hosting (EU): Data remains in EU regions with AWS/Azure/Google
  • Direct API: DeepSeek servers in China (caution with sensitive data)
  • Self-Hosting: Still the option with maximum control

For Enterprises: Cloud providers offer EU data residency with full compliance. Self-hosting remains an alternative for highest security requirements.

Self-Hosting as a Solution

The open-source model can be operated in your own infrastructure:

  • All data remains under your control
  • No dependency on external APIs
  • Full GDPR compliance possible
  • Hardware requirements: Multiple high-end GPUs (A100/H100)

Price-Performance

Even after the price adjustment DeepSeek remains an attractively priced option:

  • API: Peak/off-peak pricing since 16 Aug 2026 (see table above); cache hits and off-peak usage reduce costs considerably
  • Self-Hosting: Free to use (only hardware costs) – based on the open April weights
  • No License Fees: R1 and the V4-Flash repository under MIT, V3 under Model Agreement

Our Recommendation (as of 3 September 2026)

DeepSeek is technically impressive and, according to vendor figures, delivers strong results in reasoning, coding and agent tasks. With EU availability on AWS, Azure and Google, enterprises can use DeepSeek in a GDPR-compliant way.

For most enterprises, we recommend:

  • Cloud option: DeepSeek-V4-Flash (0731 build) via an EU cloud provider or – for non-sensitive data – via the direct API. On the API, shifting batch workloads into off-peak hours halves token costs since 16 Aug 2026.
  • Demanding agent workflows: V4-Pro (GA build 0813) with reasoning effort high or max; prices are roughly three times those of V4-Flash.
  • Self-hosting: DeepSeek-V4-Flash (open April weights, MIT) or V3.2 for maximum control and customizability. DeepSeek has not confirmed whether the summer builds 0731/0813 will be released as weights.
  • Vision-Exp: Suitable for evaluation, not yet for production use given its experimental status.

The choice depends on your requirements for control, compliance, and technical resources.

Cost estimation for this model

For up-to-date token pricing, model variants and EU availability, see our sister project ai-prices.eu. It helps you compare and estimate the operational cost of leading AI models for your specific use case.

Compare prices on ai-prices.eu

ai-prices.eu is a project by innFactory AI Consulting GmbH and provides transparent cost estimates for leading AI models.

Consultation for this model?

We help you select and integrate the right AI model for your use case.