Offre

Prix DeepSeek V4-Flash API 2026 : coût tokens peak/off-peak expliqué

Ingénierie IA DeepSeek API · prix · tokens
2026-08-22 ~14 min

En bref : Depuis 2026-08-16, DeepSeek facture peak/off-peak UTC × cache ; off-peak = moitié.

La ligne de partage est part sortie × fenêtre UTC. Off-peak out $0.66/M, peak $1.32/M.

Conclusion d'abord

  1. Depuis 2026-08-16 16:00 UTC : deepseek-v4-flash en peak/off-peak UTC ; off-peak = moitié.
  2. Off-peak miss in $0.22/M, out $0.66/M ; peak ×2. Hit off-peak in $0.007/M.
  3. Flat $0.14/$0.28 avant 8/16 obsolète.
  4. Ligne de partage : part sortie × fenêtre UTC. Thinking ON par défaut, même tarif.
  5. 1M contexte + compatible OpenAI {API_BASE}. Planification peak et cache = leviers.
Ce n'est pas le model id. Peak/off-peak UTC et part de sortie font la ligne de partage.
Tableau peak/off-peak DeepSeek V4-Flash
Verrouillez fenêtres UTC et taux de cache, puis lisez les lignes au million.

0. Conclusion d'abord

2026-08-22 DeepSeek official pricing (effective 2026-08-16 16:00 UTC) splits deepseek-v4-flash (checkpoint DeepSeek-V4-Flash-0731) into peak/off-peak × cache-hit/miss × input/output. Base https://api.deepseek.com, OpenAI-compatible; 1M context; thinking on by default at same rates.

Off-peak Flash: cache-miss in $0.22/M, out $0.66/M; peak doubles to $0.44 / $1.32. Cache-hit off-peak in $0.007/M. Peak UTC 01:00–04:00 and 06:00–10:00. Legacy flat $0.14/$0.28 dead. Gemini compare: Gemini API 2026; stack: AI Agent stack.

1. Pourquoi pas un prix unique

Teams budgeted pre-Aug 16 flat rates, launched into UTC peak, and output at $1.32/M blew budgets. Others quoted off-peak input while agent loops emit ≥3× output; thinking billed the same. No cache on 80k system prompts is expensive at peak miss $0.44/M; cache-hit off-peak gap ≈31×. August 2026 adds a time axis—not Gemini Batch, not Claude seats. Claude: Claude Code 2026.

  • Peak is whole-table ×2; sessions can span both windows.
  • Cache-hit vs miss ≈31× off-peak—skipping cache is repeat-input tax.
  • 1M context is a cap, not free quota.
  • Thinking defaults on; same price.
  • MCP retries at peak rebill output. Host: MCP placement.

2. Quatre dimensions

2.1 Peak vs off-peak (UTC)

Peak 01:00–04:00 and 06:00–10:00 UTC; off-peak half price. Cron batch to off-peak.

2.2 Cache-hit vs miss

Stable prefixes hit at Flash off-peak $0.007/M.

2.3 Input vs output

Off-peak out $0.66 is 3× miss in $0.22; peak out $1.32 is 94× hit in $0.014. Output share is the main dial.

2.4 Flash vs Pro

Pro off-peak out $1.98 ≈3× Flash; router only.

3. Comparaison à en-têtes identiques

Couche / optionEntréeExécutionContexteCoûtPérimètre
Off-peak Flashhttps://api.deepseek.com1M ctxcache tiersin $0.007–0.22, out $0.66Keys backend
Peak FlashUTC peak stampsamesamein $0.014–0.44, out $1.32Time-sliced
Legacy flatpre–8/16nonenonein $0.14, out $0.28obsolete
Off-peak Prodeepseek-v4-prostrongercachein $0.022–0.66, out $1.98hard cases
Cachestable prefixcut misssystem+Schemahit −97%change=remiss
HostCloud Mac/VPSMCP, signrepospeak retry wastealways-on

Conclusion asymétrique : peak→off-peak −50 % ; part sortie 70 %→40 % autant ou plus. Les deux.

3.1 Aide-mémoire officiel

ModelWindowhit inmiss inoutNote
deepseek-v4-flashOff-peak$0.007$0.22$0.660731
deepseek-v4-flashPeak$0.014$0.44$1.32UTC 01–04, 06–10
deepseek-v4-proOff-peak$0.022$0.66$1.98peak ×2
(legacy)Flat$0.0028$0.14$0.28dead

4. Scénarios

Couche / optionEntréeExécutionContexteCoûtPérimètre
DemoFlash off-peakshort<8k<$5/motest keys
Daily productFlash+cachecap toolsshort ctxout 3× inUTC monitor
Code agentFlash; Pro routethinking samerepo summaryout+retryMCP on
Night batchoff-peak cronchunk 1MJSONLout $0.66queue policy
Long RAGhigh cache hitstatic cachedshort querymiss controlledversion prefix
iOS agentCloud Mac+MCPinfer onlycerts on Macavoid peak retrykeychain stays

5. Stacks recommandés

A Personal: Flash off-peak, short prompts, not legacy $0.14/$0.28
B Product: Flash+cache, batch cron off-peak, Pro router only
C Multi-model: DeepSeek + Gemini/Claude split (stack article)
D Apple: DeepSeek infer; MCP/sign on Cloud Mac

6. Pièges

  • 1: Pre–Aug 16 flat rates—obsolete.
  • 2: Input-only quotes; output $0.66 off / $1.32 peak.
  • 3: Thinking-off discount—none; same card.
  • 4: No cache on large system prompts.
  • 5: Ignoring UTC 06:00–10:00 peak band.
  • 6: Unbounded retries at peak.
  • 7: Default Pro when Flash suffices.

7. 7 étapes

  1. Open official page; confirm 2026-08-16 16:00 UTC and peak windows.
  2. Run prod prompt off-peak and peak; log cache hit/miss.
  3. Cache system prompt and tool Schema; target >80% hit.
  4. Cap tool steps and daily USD; queue peak alerts to off-peak.
  5. Cron batch/evals to UTC off-peak.
  6. Move MCP off closing laptops. MCP + Cloud Mac.
  7. Internal rate card: off-peak/peak columns; diff invoices by UTC hour.

8. FAQ

Prix actuel DeepSeek V4-Flash ?

Off-peak miss in $0.22/M, out $0.66/M; peak doubles. Hit off-peak in $0.007/M.

Définition peak/off-peak ?

UTC 01:00–04:00 and 06:00–10:00 peak (2×). Billing uses API request UTC timestamp.

Thinking même tarif ?

Same price; thinking on by default.

Comparaison Gemini ?

No single sticker compare. High cache + off-peak batch favors DeepSeek input; peak output agents need $1.32/M. Gemini 2026.

Quel model id ?

deepseek-v4-flash (0731). Pro: deepseek-v4-pro. Base https://api.deepseek.com.

Pourquoi Mac cloud ?

DeepSeek sells inference only. Peak retries rebill output at $1.32/M. Cloud Mac cuts wasted output tokens.

9. Synthèse

Après août 2026, le prix DeepSeek V4-Flash API = peak/off-peak UTC × cache × mix in/out. À 70 % de sortie, le seul décalage horaire ~−35 %. Sortie, cache, retries peak ensemble.

Tokens pour l'inférence, pas les retries peak

Les prix unitaires sont publics. Le budget explose au timeout laptop + MCP coupé puis re-sortie peak $1.32/M. Cloud Mac M4 dédié garde Host, dépôt et MCP allumés.

Comparer · Offres · Démarrer