a devlog on machines & languages

23 DeepSeek V4.1 Flash

DeepSeek V4.1 Flash was launched yesterday and is already available on OpenCode Go. It now has Vision support. The price change from August 16 has been partially reversed. For now:

TypeOff-peakPeak
Cached Read$0.003/M$0.006/M
Input$0.15/M$0.30/M
Output$0.60/M$1.20/M

It seems that the DeepSeek V4 Flash Vision experiment was a success. One of the most important changes compared to DeepSeek V4 Flash, for those running quantized models at home, is the model size:

MetricDeepSeek V4 FlashDeepSeek V4.1 Flash
ArchitectureMoEMoE
Backbone Params284B552B
Activated Params13B8B/16B
Context Length1M1M
PrecisionFP8 MixedFP8 Mixed

With V4, you could run Q3 quants on a 128GB RAM machine. With V4.1, you cannot.