DeepSeek R1 Zero

DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors.

MATH-500

95.9%

i
AIME 2024

86.7%

i
GPQA

73.3%

i
LiveCodeBench

50.0%

i