DeepSeek V4.1 Flash API pricing
DeepSeek's current Flash model with native multimodal support and strong coding, reasoning and agentic capabilities.
How much does DeepSeek V4.1 Flash API cost?
DeepSeek V4.1 Flash uses variable TokenGate pricing. The rates below depend on the request conditions listed in each row.
| Condition | Input tokens / 1M | Cached input / 1M | Output tokens / 1M |
|---|---|---|---|
| Cache hit, off-peak | $0.0030 | — | — |
| Cache hit, peak | $0.0060 | — | — |
| Cache miss, off-peak | $0.15 | — | — |
| Cache miss, peak | $0.30 | — | — |
| Output, off-peak | — | — | $0.60 |
| Output, peak | — | — | $1.20 |
- Pricing depends on cache hit or miss and on peak or off-peak hours.
Use DeepSeek V4.1 Flash with the OpenAI SDK
curl https://api.tokengate.com/v1/chat/completions \
-H "Authorization: Bearer $TOKENGATE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"deepseek-flash","messages":[{"role":"user","content":"Hello"}]}'Change the model identifier to use another supported model through the same TokenGate integration.
Model specifications
- Provider
- DeepSeek
- API model identifier
- deepseek-flash
- Generation
- 4.1
- Provider lifecycle
- Active
- Input modalities
- Text, Image
- Output modalities
- Text
- Capabilities
- Streaming, Function calling, Structured outputs, Vision, Reasoning, Tool use
Best for
- Coding
- Agents
- Complex reasoning
- High volume
- Multimodal
Access DeepSeek V4.1 Flash through TokenGate
Use supported models through a single TokenGate integration.
- · One API integration
- · OpenAI-compatible interface
- · Payment in USDT
Frequently asked questions
Related models
Gemini 3.5 Flash
Google Flash model with a large context window for high-volume multimodal workloads.
- Input / 1M
- $1.50
- Output / 1M
- $9.00
- Context
- 1.05M
- Generation
- 3.5
Streaming · Function calling · Structured outputs
Gemini 3.5 Flash API pricingGemini 3.1 Pro Preview
Google's Pro preview model for complex reasoning, coding and multimodal agentic work.
- Input / 1M
- Variable pricing
- Output / 1M
- —
- Context
- 1.05M
- Generation
- 3.1
Streaming · Function calling · Structured outputs
Gemini 3.1 Pro Preview API pricingGPT-5.6 Luna
Cost-focused GPT-5.6 model for high-volume workloads, routine automation and applications where API efficiency matters.
- Input / 1M
- $0.20
- Output / 1M
- $1.20
- Context
- 1.05M
- Generation
- 5.6
Streaming · Function calling · Structured outputs
GPT-5.6 Luna API pricingTokenGate
