WTF Is Going On(Weights, Tools & Frameworks)

한국어
deepseek-ai/DeepSeek-V4.1-Flash
image-text-to-text mit FP8

What this is

763B parameters. Takes pictures and text together and answers in text. You can hand it a screenshot and ask about it.

Packaged as FP8. FP8 needs a GPU that supports it — Hopper or newer.

mit — commercial use allowed.

Release chainonly the steps this archive can prove

  1. weights posted2026-09-10
  2. argued about999 points · 570 comments

What they saidthe thread’s own top comments, verbatim

DeepSeek v4.1 Flash — 999 points · 570 comments
kouteiheika
It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers. […
rao-v
As I also said on Twitter - it really amazes me how fearless Deepseek are. Every single model release is packed with new and crazy clever ideas and somehow, they always commit to training them at near frontier scale. I know everybody wants…
k9294
I'm surprised more people aren't talking about the cache hit price: $0.003 per million tokens. I have a feeling that the price of 1 million tokens transmitted over the internet is more expensive than cache hit. Are we close to making the c…

Measured herefrom our own daily capture — upstream publishes today only · one day in this window was never captured

← the board