deepseek-v4.1-flash
@deepseek/deepseek-v4.1-flash
DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the cost-efficient tier of the V4.1 family.
- All contexts
- Input $0.3Output $1.2Cached $0.006
Providers & technical details 12 for @deepseek/deepseek-v4.1-flash
DeepSeek
deepseek/deepseek-v4.1-flash- Context
- 1,048,576 tokens
- Max output
- 384,000 tokens
- Throughput
- 136 tokens/s
- Latency
- 1.112 s
- Uptime
- 99.9981%
- Data collection
- Prompt With Training
- All contexts
- Input $0.15Output $0.6Cached $0.003
DeepInfra:offline
Disableddeepseek/deepseek-v4.1-flash- Context
- 1,048,576 tokens
- Max output
- 131,072 tokens
- Throughput
- 24 tokens/s
- Latency
- 3.146 s
- Uptime
- 91.9323%
- Data collection
- Zero
- All contexts
- Input $0.2Output $0.6Cached $0.006
Fireworks
deepseek/deepseek-v4.1-flash- Context
- 1,048,576 tokens
- Max output
- 943,718 tokens
- Throughput
- 105 tokens/s
- Latency
- 1.318 s
- Uptime
- 99.5117%
- Data collection
- Zero
- All contexts
- Input $0.22Output $0.66Cached $0.007
Morph
deepseek/deepseek-v4.1-flash- Context
- 1,048,576 tokens
- Max output
- 943,718 tokens
- Throughput
- 27 tokens/s
- Latency
- 1.915 s
- Uptime
- 99.473%
- Data collection
- Zero
- All contexts
- Input $0.225Output $0.9Cached $0.0225
SiliconFlow
deepseek/deepseek-v4.1-flash- Context
- 1,048,576 tokens
- Max output
- 393,216 tokens
- Throughput
- 166 tokens/s
- Latency
- 1.117 s
- Uptime
- 100%
- Data collection
- Zero
- All contexts
- Input $0.3Output $1.2Cached $0.006
Modal
deepseek/deepseek-v4.1-flash- Context
- 1,048,576 tokens
- Max output
- 943,718 tokens
- Throughput
- 76 tokens/s
- Latency
- 1.3945 s
- Uptime
- 95.4147%
- Data collection
- Zero
- All contexts
- Input $0.3Output $1.2Cached $0.03
Wafer
deepseek/deepseek-v4.1-flash- Context
- 1,048,576 tokens
- Max output
- 943,718 tokens
- Throughput
- 37 tokens/s
- Latency
- 0.9055 s
- Uptime
- 99.6663%
- Data collection
- Zero
- All contexts
- Input $0.3Output $1.2Cached $0.006
Parasail
deepseek/deepseek-v4.1-flash- Context
- 1,048,576 tokens
- Max output
- 943,718 tokens
- Throughput
- 92 tokens/s
- Latency
- 1.186 s
- Uptime
- 97.7013%
- Data collection
- Zero
- All contexts
- Input $0.3Output $1.2Cached $0.006
GMICloud
deepseek/deepseek-v4.1-flash- Context
- 1,048,575 tokens
- Max output
- 943,717 tokens
- Throughput
- 93 tokens/s
- Latency
- 3.4645 s
- Uptime
- 99.9774%
- Data collection
- Prompt No Training
- All contexts
- Input $0.3Output $1.2Cached $0.006
Io Net
deepseek/deepseek-v4.1-flash- Context
- 262,124 tokens
- Max output
- 131,072 tokens
- Throughput
- 71 tokens/s
- Latency
- 1.148 s
- Uptime
- 99.8894%
- Data collection
- Zero
- All contexts
- Input $0.3Output $1.2Cached $0.003
Novita
deepseek/deepseek-v4.1-flash- Context
- 1,048,576 tokens
- Max output
- 393,216 tokens
- Throughput
- 130 tokens/s
- Latency
- 2.059 s
- Uptime
- 100%
- Data collection
- Zero
- All contexts
- Input $0.3Output $1.2Cached $0.006
Venice
deepseek/deepseek-v4.1-flash- Context
- 1,000,000 tokens
- Max output
- 131,072 tokens
- Throughput
- 97 tokens/s
- Latency
- 0.979 s
- Uptime
- 99.5052%
- Data collection
- Zero
- All contexts
- Input $0.375Output $1.5Cached $0.0075