DeepThink-Gemma-2-2B-GRPO (Model Weights)
This repository contains the final model weights for Gemma-2-2B-IT, fine-tuned using Group Relative Policy Optimization (GRPO).
Model Specifications
- Base Model: Google Gemma-2-2b-it.
- Size: 4.10 GB.
- Format: JAX / Flax.
- Purpose: Enhanced logical reasoning capabilities.
Directory Structure
- final_model/: Contains the necessary model weights and configuration files for deployment.
Deployment and Usage
These weights are ready for inference in environments that support the JAX framework. For training scripts and detailed implementation, please refer to the associated GitHub repository.