LongCut logo

DeepSeek Just Made Closed AI Look Ridiculous

By Two Minute Papers

Summary

Topics Covered

  • Open Weights Force A Price War
  • Same Architecture, Massive Post-Training Gains
  • Ten Specialist Teachers Distill One Student
  • Six Weeks From Paper To Product

Full Transcript

Yes, DeepSeek 4 Pro is here. This time,

the real one. I know it gets confusing.

The previous version was called preview, and this is called 0813.

The numbers say it performs better than the much smaller flash version. When

building a Rubik's Cube, flash did not completely understand the 3D structure of the object. Lots of missing parts,

lots of blackness. But, with the pro, look, much better understanding of structure. Then, I was also surprised by

structure. Then, I was also surprised by this.

Holy mother of papers. Look at that. It

is inching closer and closer to fable quality. Yep, another challenger

quality. Yep, another challenger appeared, and it gets better. They give

all this for us for free, which is absolutely incredible. Okay, but what

absolutely incredible. Okay, but what does it mean for us? The weights are available for free for all of us, but, you know, few have the hardware to host it at home. I'd love to, but I don't

have that kind of hardware. Other

options include Lambda or using it hosted by DeepSeek themselves. But, they

just raised their prices dramatically, about 2 and 1/2 to 5x the previous prices. Now, I bet you can already

prices. Now, I bet you can already imagine the clickbait headlines saying, "It's over."

"It's over." I think what they should also say is that DeepSeek has MIT licensed open weights. What does that mean? Well,

weights. What does that mean? Well,

anyone can run the exact same model at their own price. And, look, they do. A

bunch of hosts available, and they all compete on price. That is amazing for us, and it is very likely to push the Frontier Labs to give us fellow scholars

something even better, and quickly. And

all this improvement comes from the same architecture. But, how is that even

architecture. But, how is that even possible? The model structure is the

possible? The model structure is the same, yet it is massively better than the preview was less than 4 months ago.

So, how? Dear fellow scholars, this is Two Minute Papers with Dr. Károly Zsolnai-Fehér. Once again, a lot of the

Zsolnai-Fehér. Once again, a lot of the magic happens after pre-training. During

post-training, DeepSeek creates several specialist models for mathematics, coding, and agentic work. Now, we have to stop here for a moment. People

confuse these with the experts in mixture of experts. That's not quite the same. Those are little pieces within one

same. Those are little pieces within one neural network. These are not. These are

neural network. These are not. These are

separately trained model checkpoints.

Okay, so what then? Then comes

distillation. Yeah. They take more than 10 of these specialist teachers and train one final model to absorb their abilities. So, the student model says,

abilities. So, the student model says, "This is what I would do." Then the teacher says, "Well, this is what I would have done." Then the student adapts its brain [clears throat] to be

more like its teacher. Do it with 10 teachers and you see that the student indeed improves like crazy. They also

added this part to it. Instead of just predicting one token at a time, it drafts several tokens ahead. It does it much better than previous techniques.

And hold on to your papers, fellow scholars, because DeepSeek reports up to 78% faster generation for V4 Pro. Real,

measurable speed up in real use that you get right now and benefit from it. And

here is something absolutely insane.

This was a research paper, let's see, 6 weeks ago. And now everyone is using it. Let me say it again, a research

it. Let me say it again, a research paper only 6 weeks ago. One of the best papers of the year. And it is coming

alive right in our hands, for free.

Incredible. Full breakdown video in the description. And don't forget, we own

description. And don't forget, we own and can run the weights. No one

downgrades us to a different model if we type the wrong keyword. No games. That

is incredible. Even if I can't run it at home, which I would love to do. But,

there are options. What a time to be alive. This is open science and open

alive. This is open science and open research at its best. And it's important that we talk about it. Why? Because the

future belongs to those who understand it. Use DeepSeek and use DeepSpark. Take

it. Use DeepSeek and use DeepSpark. Take

advantage of them. Oh, and I plan to talk about DeepSeek's incredible no agent harness as well. Novel design,

really powerful. If you're interested, consider subscribing and hitting the bell. I use Lambda to reproduce AI

bell. I use Lambda to reproduce AI research papers often in minutes. It's

also great to train your own models or fine-tune an existing one. Run inference

or text-to-image or video, easy-peasy.

Running a DeepSeek chatbot or agent, super fast, super reliable. Lambda gives

you powerful Nvidia GPUs to run your own experiments. I test ideas from the

experiments. I test ideas from the papers I cover and moments later, results.

Love it. Seriously, try it out now at lambda.ai/papers.

lambda.ai/papers.

Loading...

Loading video analysis...