GGUF
conversational

Model Specifications

Architecture

Property Value
Base Model IBM Granite 4.1
Architecture Type Decoder-Only Transformer
Transformer Layers 40
Hidden Size 4096
Attention Heads 32
KV Heads (GQA) 8
Head Dimension 128
Intermediate Size 12800
Vocabulary Size 100,352
Context Length 131,072 Tokens
Activation Function SiLU
RoPE Theta 10,000,000

Fine-Tuning Statistics

Metric Value
Fine-Tuning Method LoRA
Trainable Parameters 98,959,360
Total Parameters 4,494,921,728
Trainable Percentage 2.20%
Base Parameters Frozen 97.80%
Training Framework Unsloth
Optimizer AdamW 8-bit

Quantization Details

Property Value
Output Format GGUF
Quantization Method Q4_K_M
Quantization Type K-Quant Medium
Deployment Size ~5 GB
Runtime Engine Ollama / llama.cpp

Memory Analysis

KV Cache Formula

KV Cache Per Token:

KV Cache = 2 Γ— Layers Γ— KV Heads Γ— Head Dimension Γ— 2 Bytes

Calculation:

2 Γ— 40 Γ— 8 Γ— 128 Γ— 2

= 163,840 Bytes

β‰ˆ 160 KB per Token


Estimated Runtime Memory Usage

Context Length KV Cache Total Runtime Memory
4K Tokens ~655 MB ~6.2 GB
8K Tokens ~1.31 GB ~7.0 GB
16K Tokens ~2.62 GB ~8–9 GB
32K Tokens ~5.24 GB ~11 GB

Hardware Requirements

Training Environment

Component Value
GPU NVIDIA T4
VRAM 16 GB
Quantization 4-bit NF4
Fine-Tuning Method LoRA

Inference Environment

Component Value
GPU RTX 2080
VRAM 8 GB
System RAM 32 GB
Recommended Context 8192 Tokens
Quantization Q4_K_M

Deployment Artifacts

Artifact Purpose
granite_python_lora.zip LoRA Adapter Backup
adapter_model.safetensors Fine-Tuned Weights
granite-4.1-8b.Q4_K_M.gguf Deployable Model
Modelfile Ollama Configuration
granite-python Ollama Model Name

Project Workflow

Dataset β†’ LoRA Fine-Tuning β†’ Adapter Export β†’ Model Merge β†’ GGUF Conversion β†’ Q4_K_M Quantization β†’ Ollama Deployment β†’ VS Code Integration β†’ Custom Python Code Generation AI Model

Downloads last month
41
GGUF
Model size
9B params
Architecture
granite
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support