Qwen3 Highlights

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features:

Uniquely support of seamless switching between thinking mode (for complex logical reasoning, math, and coding) and non-thinking mode (for efficient, general-purpose dialogue) within single model, ensuring optimal performance across various scenarios. Significantly enhancement in its reasoning capabilities, surpassing previous QwQ (in thinking mode) and Qwen2.5 instruct models (in non-thinking mode) on mathematics, code generation, and commonsense logical reasoning. Superior human preference alignment, excelling in creative writing, role-playing, multi-turn dialogues, and instruction following, to deliver a more natural, engaging, and immersive conversational experience. Expertise in agent capabilities, enabling precise integration with external tools in both thinking and unthinking modes and achieving leading performance among open-source models in complex agent-based tasks. Support of 100+ languages and dialects with strong capabilities for multilingual instruction following and translation.

Model Details

Qwen3-8B-Lite is a lightweight variant of the Qwen3 series of generative AI models with approximately 8 billion parameters designed for efficient language understanding and generation tasks. It is optimized for faster inference and lower compute requirements while maintaining competitive accuracy.

  • Model type: Transformer-based generative language model (e.g., decoder-only or encoder-decoder)
  • Language(s): Primarily English (or specify other supported languages)
  • Finetuned from: Base Qwen3-8B

Model Description

  • Model Type: Causal Language Model (decoder-only transformer)
  • Training Stage: Pretraining and Post-training quantization (FP8)
  • Number of Parameters: Approximately 8.2 billion total parameters
  • Number of Parameters (Non-Embedding): Approximately 6.95 billion
  • Number of Layers: 36 transformer layers
  • Number of Attention Heads (Grouped Query Attention - GQA): 32 heads for Query 8 heads for Key-Value

Direct Use

  • Text completion, generation, and summarization
  • Chatbots and conversational AI
  • Language understanding tasks like classification and question answering

Downstream Use [optional]

  • Fine-tuning on domain-specific data
  • Integration into NLP pipelines and applications
Downloads last month
4
Safetensors
Model size
8B params
Tensor type
F16
F32
U8
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 1 Ask for provider support