Angebot

DeepSeek V4-Flash API Preise 2026: Peak/Off-Peak Tokenkosten erklärt

KI-Engineering DeepSeek API · Preise · Token
2026-08-22 ~14 Min.

Fazit: Seit 2026-08-16 rechnet DeepSeek UTC Peak/Off-Peak × Cache ab; Off-Peak genau halb.

Wasserscheide ist Output-Anteil × UTC-Fenster. Off-Peak out $0.66/M, Peak $1.32/M.

Fazit zuerst

  1. Seit 2026-08-16 16:00 UTC: deepseek-v4-flash mit UTC Peak/Off-Peak; Off-Peak genau halb.
  2. Off-Peak miss in $0.22/M, out $0.66/M; Peak verdoppelt. Hit Off-Peak in $0.007/M.
  3. Flat $0.14/$0.28 vor 8/16 obsolet.
  4. Wasserscheide: Output-Anteil × UTC-Fenster. Thinking standardmäßig an, gleicher Tarif.
  5. 1M Kontext + OpenAI-kompatibel {API_BASE}. Peak-Planung und Cache sind Hebel.
Nicht die model id. UTC Peak/Off-Peak und Output-Anteil sind die Wasserscheide.
DeepSeek V4-Flash Peak/Off-Peak-Tabelle
Zuerst UTC-Fenster und Cache-Trefferquote, dann Million-Token-Zeilen.

0. Fazit zuerst

2026-08-22 DeepSeek official pricing (effective 2026-08-16 16:00 UTC) splits deepseek-v4-flash (checkpoint DeepSeek-V4-Flash-0731) into peak/off-peak × cache-hit/miss × input/output. Base https://api.deepseek.com, OpenAI-compatible; 1M context; thinking on by default at same rates.

Off-peak Flash: cache-miss in $0.22/M, out $0.66/M; peak doubles to $0.44 / $1.32. Cache-hit off-peak in $0.007/M. Peak UTC 01:00–04:00 and 06:00–10:00. Legacy flat $0.14/$0.28 dead. Gemini compare: Gemini API 2026; stack: AI Agent stack.

1. Warum kein Einzelpreis

Teams budgeted pre-Aug 16 flat rates, launched into UTC peak, and output at $1.32/M blew budgets. Others quoted off-peak input while agent loops emit ≥3× output; thinking billed the same. No cache on 80k system prompts is expensive at peak miss $0.44/M; cache-hit off-peak gap ≈31×. August 2026 adds a time axis—not Gemini Batch, not Claude seats. Claude: Claude Code 2026.

  • Peak is whole-table ×2; sessions can span both windows.
  • Cache-hit vs miss ≈31× off-peak—skipping cache is repeat-input tax.
  • 1M context is a cap, not free quota.
  • Thinking defaults on; same price.
  • MCP retries at peak rebill output. Host: MCP placement.

2. Vier Dimensionen

2.1 Peak vs off-peak (UTC)

Peak 01:00–04:00 and 06:00–10:00 UTC; off-peak half price. Cron batch to off-peak.

2.2 Cache-hit vs miss

Stable prefixes hit at Flash off-peak $0.007/M.

2.3 Input vs output

Off-peak out $0.66 is 3× miss in $0.22; peak out $1.32 is 94× hit in $0.014. Output share is the main dial.

2.4 Flash vs Pro

Pro off-peak out $1.98 ≈3× Flash; router only.

3. Vergleich gleicher Spalten

Ebene / OptionEinstiegAusführungKontextKostenRechtegrenze
Off-peak Flashhttps://api.deepseek.com1M ctxcache tiersin $0.007–0.22, out $0.66Keys backend
Peak FlashUTC peak stampsamesamein $0.014–0.44, out $1.32Time-sliced
Legacy flatpre–8/16nonenonein $0.14, out $0.28obsolete
Off-peak Prodeepseek-v4-prostrongercachein $0.022–0.66, out $1.98hard cases
Cachestable prefixcut misssystem+Schemahit −97%change=remiss
HostCloud Mac/VPSMCP, signrepospeak retry wastealways-on

Asymmetrisch: Peak→Off-Peak spart 50%; Output-Anteil 70%→40% spart ähnlich viel. Beides.

3.1 Offizielle Übersicht

ModelWindowhit inmiss inoutNote
deepseek-v4-flashOff-peak$0.007$0.22$0.660731
deepseek-v4-flashPeak$0.014$0.44$1.32UTC 01–04, 06–10
deepseek-v4-proOff-peak$0.022$0.66$1.98peak ×2
(legacy)Flat$0.0028$0.14$0.28dead

4. Szenarien

Ebene / OptionEinstiegAusführungKontextKostenRechtegrenze
DemoFlash off-peakshort<8k<$5/motest keys
Daily productFlash+cachecap toolsshort ctxout 3× inUTC monitor
Code agentFlash; Pro routethinking samerepo summaryout+retryMCP on
Night batchoff-peak cronchunk 1MJSONLout $0.66queue policy
Long RAGhigh cache hitstatic cachedshort querymiss controlledversion prefix
iOS agentCloud Mac+MCPinfer onlycerts on Macavoid peak retrykeychain stays

5. Empfohlene Stacks

A Personal: Flash off-peak, short prompts, not legacy $0.14/$0.28
B Product: Flash+cache, batch cron off-peak, Pro router only
C Multi-model: DeepSeek + Gemini/Claude split (stack article)
D Apple: DeepSeek infer; MCP/sign on Cloud Mac

6. Fallstricke

  • 1: Pre–Aug 16 flat rates—obsolete.
  • 2: Input-only quotes; output $0.66 off / $1.32 peak.
  • 3: Thinking-off discount—none; same card.
  • 4: No cache on large system prompts.
  • 5: Ignoring UTC 06:00–10:00 peak band.
  • 6: Unbounded retries at peak.
  • 7: Default Pro when Flash suffices.

7. 7 Schritte

  1. Open official page; confirm 2026-08-16 16:00 UTC and peak windows.
  2. Run prod prompt off-peak and peak; log cache hit/miss.
  3. Cache system prompt and tool Schema; target >80% hit.
  4. Cap tool steps and daily USD; queue peak alerts to off-peak.
  5. Cron batch/evals to UTC off-peak.
  6. Move MCP off closing laptops. MCP + Cloud Mac.
  7. Internal rate card: off-peak/peak columns; diff invoices by UTC hour.

8. FAQ

Was kostet DeepSeek V4-Flash jetzt?

Off-peak miss in $0.22/M, out $0.66/M; peak doubles. Hit off-peak in $0.007/M.

Peak vs Off-Peak?

UTC 01:00–04:00 and 06:00–10:00 peak (2×). Billing uses API request UTC timestamp.

Thinking gleich bepreist?

Same price; thinking on by default.

Vergleich mit Gemini?

No single sticker compare. High cache + off-peak batch favors DeepSeek input; peak output agents need $1.32/M. Gemini 2026.

Welche model id?

deepseek-v4-flash (0731). Pro: deepseek-v4-pro. Base https://api.deepseek.com.

Warum Cloud-Mac?

DeepSeek sells inference only. Peak retries rebill output at $1.32/M. Cloud Mac cuts wasted output tokens.

9. Zusammenfassung

Ab Aug 2026 ist DeepSeek V4-Flash API Preis UTC Peak/Off-Peak × Cache × In/Out-Mix. Bei 70 % Output spart Scheduling allein ~35 %. Output, Cache, Peak-Retries zusammen.

Token für Inferenz, nicht Peak-Retries

Stückpreise sind öffentlich. Budget sprengt Timeout plus MCP-Abbruch und Peak-Neuausgabe $1.32/M. Dediziertes M4-Cloud-Mac hält Host, Repo und MCP auf einem Pfad.

Vergleich · Tarife · Start