Skip to content
Not available in this workspace
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Support
  • Works With OR
  • Data

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Favicon for deepseek

DeepSeek: R1 Distill Llama 8B

deepseek/deepseek-r1-distill-llama-8b

Model weights

DeepSeek R1 Distill Llama 8B is a distilled large language model based on Llama-3.1-8B-Instruct, using outputs from DeepSeek R1. The model combines advanced distillation techniques to achieve high performance across multiple benchmarks, including:

  • AIME 2024 pass@1: 50.4
  • MATH-500 pass@1: 89.1
  • CodeForces Rating: 1205

The model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.

Hugging Face:

  • Llama-3.1-8B(opens in new tab)
  • DeepSeek-R1-Distill-Llama-8B(opens in new tab) |

Modalities

Released

Feb 7, 2025

Knowledge Cutoff

Jul 2024

ActivityFAQ

Activity

Token volume and request traffic to this model over time.

About DeepSeek: R1 Distill Llama 8B

OpenRouter makes DeepSeek: R1 Distill Llama 8B available through a unified, OpenAI-compatible API using the model ID deepseek/deepseek-r1-distill-llama-8b.

DeepSeek: R1 Distill Llama 8B accepts text and returns text.

It was released on February 7, 2025; its knowledge cutoff is July 31, 2024.

More models from DeepSeek

  • DeepSeek V4 Pro 0813
  • DeepSeek V4 Flash 0731
  • DeepSeek V4 Pro 0423

Frequently asked questions

DeepSeek R1 Distill Llama 8B is a distilled large language model based on Llama-3.1-8B-Instruct, using outputs from DeepSeek R1. The model combines advanced distillation techniques to achieve high performance across multiple benchmarks, including: - AIME 2024 pass@1: 50.4 - MATH-500 pass@1: 89.1 - CodeForces Rating: 1205 The model leverages fine-tuning from DeepSeek R1's outputs, enabling...

DeepSeek V4 Pro 0813, DeepSeek V4 Flash 0731, DeepSeek V4 Pro 0423 and 9 more are other text models from DeepSeek.

R1 Distill Llama 8B was released on February 7, 2025. Its knowledge cutoff is July 31, 2024.