DeepSeek-V3 Quietly Rewrote the Rules on Training Cost

Released at the very end of December 2024, DeepSeek-V3 was DeepSeek's flagship Mixture-of-Experts base model: 671 billion total parameters, with only a fraction active per token, released under the fully permissive MIT license.

Why it mattered

DeepSeek-V3's headline claim was reaching frontier-competitive benchmark performance at a training cost dramatically lower than what the industry had generally assumed was necessary for models at that capability level. It also laid the architectural and training groundwork for DeepSeek-R1 barely a month later -- the release that would turn DeepSeek from a notable open-weight lab into a genuinely disruptive one.

On the record

See the full hash breakdown on DeepSeek-V3's Explorer page.