Fazit zuerst
- Seit 2026-08-16 16:00 UTC: deepseek-v4-flash mit UTC Peak/Off-Peak; Off-Peak genau halb.
- Off-Peak miss in $0.22/M, out $0.66/M; Peak verdoppelt. Hit Off-Peak in $0.007/M.
- Flat $0.14/$0.28 vor 8/16 obsolet.
- Wasserscheide: Output-Anteil × UTC-Fenster. Thinking standardmäßig an, gleicher Tarif.
- 1M Kontext + OpenAI-kompatibel
{API_BASE}. Peak-Planung und Cache sind Hebel.
Nicht die model id. UTC Peak/Off-Peak und Output-Anteil sind die Wasserscheide.
0. Fazit zuerst
2026-08-22 DeepSeek official pricing (effective 2026-08-16 16:00 UTC) splits deepseek-v4-flash (checkpoint DeepSeek-V4-Flash-0731) into peak/off-peak × cache-hit/miss × input/output. Base https://api.deepseek.com, OpenAI-compatible; 1M context; thinking on by default at same rates.
Off-peak Flash: cache-miss in $0.22/M, out $0.66/M; peak doubles to $0.44 / $1.32. Cache-hit off-peak in $0.007/M. Peak UTC 01:00–04:00 and 06:00–10:00. Legacy flat $0.14/$0.28 dead. Gemini compare: Gemini API 2026; stack: AI Agent stack.
1. Warum kein Einzelpreis
Teams budgeted pre-Aug 16 flat rates, launched into UTC peak, and output at $1.32/M blew budgets. Others quoted off-peak input while agent loops emit ≥3× output; thinking billed the same. No cache on 80k system prompts is expensive at peak miss $0.44/M; cache-hit off-peak gap ≈31×. August 2026 adds a time axis—not Gemini Batch, not Claude seats. Claude: Claude Code 2026.
- Peak is whole-table ×2; sessions can span both windows.
- Cache-hit vs miss ≈31× off-peak—skipping cache is repeat-input tax.
- 1M context is a cap, not free quota.
- Thinking defaults on; same price.
- MCP retries at peak rebill output. Host: MCP placement.
2. Vier Dimensionen
2.1 Peak vs off-peak (UTC)
Peak 01:00–04:00 and 06:00–10:00 UTC; off-peak half price. Cron batch to off-peak.
2.2 Cache-hit vs miss
Stable prefixes hit at Flash off-peak $0.007/M.
2.3 Input vs output
Off-peak out $0.66 is 3× miss in $0.22; peak out $1.32 is 94× hit in $0.014. Output share is the main dial.
2.4 Flash vs Pro
Pro off-peak out $1.98 ≈3× Flash; router only.
3. Vergleich gleicher Spalten
| Ebene / Option | Einstieg | Ausführung | Kontext | Kosten | Rechtegrenze |
|---|---|---|---|---|---|
| Off-peak Flash | https://api.deepseek.com | 1M ctx | cache tiers | in $0.007–0.22, out $0.66 | Keys backend |
| Peak Flash | UTC peak stamp | same | same | in $0.014–0.44, out $1.32 | Time-sliced |
| Legacy flat | pre–8/16 | none | none | in $0.14, out $0.28 | obsolete |
| Off-peak Pro | deepseek-v4-pro | stronger | cache | in $0.022–0.66, out $1.98 | hard cases |
| Cache | stable prefix | cut miss | system+Schema | hit −97% | change=remiss |
| Host | Cloud Mac/VPS | MCP, sign | repos | peak retry waste | always-on |
Asymmetrisch: Peak→Off-Peak spart 50%; Output-Anteil 70%→40% spart ähnlich viel. Beides.
3.1 Offizielle Übersicht
| Model | Window | hit in | miss in | out | Note |
|---|---|---|---|---|---|
| deepseek-v4-flash | Off-peak | $0.007 | $0.22 | $0.66 | 0731 |
| deepseek-v4-flash | Peak | $0.014 | $0.44 | $1.32 | UTC 01–04, 06–10 |
| deepseek-v4-pro | Off-peak | $0.022 | $0.66 | $1.98 | peak ×2 |
| (legacy) | Flat | $0.0028 | $0.14 | $0.28 | dead |
4. Szenarien
| Ebene / Option | Einstieg | Ausführung | Kontext | Kosten | Rechtegrenze |
|---|---|---|---|---|---|
| Demo | Flash off-peak | short | <8k | <$5/mo | test keys |
| Daily product | Flash+cache | cap tools | short ctx | out 3× in | UTC monitor |
| Code agent | Flash; Pro route | thinking same | repo summary | out+retry | MCP on |
| Night batch | off-peak cron | chunk 1M | JSONL | out $0.66 | queue policy |
| Long RAG | high cache hit | static cached | short query | miss controlled | version prefix |
| iOS agent | Cloud Mac+MCP | infer only | certs on Mac | avoid peak retry | keychain stays |
5. Empfohlene Stacks
A Personal: Flash off-peak, short prompts, not legacy $0.14/$0.28 B Product: Flash+cache, batch cron off-peak, Pro router only C Multi-model: DeepSeek + Gemini/Claude split (stack article) D Apple: DeepSeek infer; MCP/sign on Cloud Mac
6. Fallstricke
- 1: Pre–Aug 16 flat rates—obsolete.
- 2: Input-only quotes; output $0.66 off / $1.32 peak.
- 3: Thinking-off discount—none; same card.
- 4: No cache on large system prompts.
- 5: Ignoring UTC 06:00–10:00 peak band.
- 6: Unbounded retries at peak.
- 7: Default Pro when Flash suffices.
7. 7 Schritte
- Open official page; confirm 2026-08-16 16:00 UTC and peak windows.
- Run prod prompt off-peak and peak; log cache hit/miss.
- Cache system prompt and tool Schema; target >80% hit.
- Cap tool steps and daily USD; queue peak alerts to off-peak.
- Cron batch/evals to UTC off-peak.
- Move MCP off closing laptops. MCP + Cloud Mac.
- Internal rate card: off-peak/peak columns; diff invoices by UTC hour.
8. FAQ
Was kostet DeepSeek V4-Flash jetzt?
Off-peak miss in $0.22/M, out $0.66/M; peak doubles. Hit off-peak in $0.007/M.
Peak vs Off-Peak?
UTC 01:00–04:00 and 06:00–10:00 peak (2×). Billing uses API request UTC timestamp.
Thinking gleich bepreist?
Same price; thinking on by default.
Vergleich mit Gemini?
No single sticker compare. High cache + off-peak batch favors DeepSeek input; peak output agents need $1.32/M. Gemini 2026.
Welche model id?
deepseek-v4-flash (0731). Pro: deepseek-v4-pro. Base https://api.deepseek.com.
Warum Cloud-Mac?
DeepSeek sells inference only. Peak retries rebill output at $1.32/M. Cloud Mac cuts wasted output tokens.
9. Zusammenfassung
Ab Aug 2026 ist DeepSeek V4-Flash API Preis UTC Peak/Off-Peak × Cache × In/Out-Mix. Bei 70 % Output spart Scheduling allein ~35 %. Output, Cache, Peak-Retries zusammen.