Instructions to use adeljebali/llama3.1-gec-strict with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use adeljebali/llama3.1-gec-strict with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/Meta-Llama-3.1-8B-bnb-4bit") model = PeftModel.from_pretrained(base_model, "adeljebali/llama3.1-gec-strict") - Transformers
How to use adeljebali/llama3.1-gec-strict with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="adeljebali/llama3.1-gec-strict")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("adeljebali/llama3.1-gec-strict") model = AutoModelForCausalLM.from_pretrained("adeljebali/llama3.1-gec-strict", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use adeljebali/llama3.1-gec-strict with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf adeljebali/llama3.1-gec-strict:Q4_K_M # Run inference directly in the terminal: llama cli -hf adeljebali/llama3.1-gec-strict:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf adeljebali/llama3.1-gec-strict:Q4_K_M # Run inference directly in the terminal: llama cli -hf adeljebali/llama3.1-gec-strict:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf adeljebali/llama3.1-gec-strict:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf adeljebali/llama3.1-gec-strict:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf adeljebali/llama3.1-gec-strict:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf adeljebali/llama3.1-gec-strict:Q4_K_M
Use Docker
docker model run hf.co/adeljebali/llama3.1-gec-strict:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use adeljebali/llama3.1-gec-strict with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "adeljebali/llama3.1-gec-strict" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "adeljebali/llama3.1-gec-strict", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/adeljebali/llama3.1-gec-strict:Q4_K_M
- SGLang
How to use adeljebali/llama3.1-gec-strict with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "adeljebali/llama3.1-gec-strict" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "adeljebali/llama3.1-gec-strict", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "adeljebali/llama3.1-gec-strict" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "adeljebali/llama3.1-gec-strict", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Ollama
How to use adeljebali/llama3.1-gec-strict with Ollama:
ollama run hf.co/adeljebali/llama3.1-gec-strict:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use adeljebali/llama3.1-gec-strict with Docker Model Runner:
docker model run hf.co/adeljebali/llama3.1-gec-strict:Q4_K_M
- Lemonade
How to use adeljebali/llama3.1-gec-strict with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull adeljebali/llama3.1-gec-strict:Q4_K_M
Run and chat with the model
lemonade run user.llama3.1-gec-strict-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Model Card
Model Description
A French grammar correction model designed primarily for learners of French as a second language (FSL/FLE). It corrects grammar, spelling, syntax, punctuation, and stylistic issues while preserving the original meaning and tone as much as possible.
The model is especially effective for:
- French learners and students
- Academic and professional writing
- Language practice and self-correction
- Improving fluency and sentence naturalness
It can also perform high-quality translation from English to French, making it useful both as a grammar corrector and as a lightweight bilingual writing assistant.
Optimized for local inference with LM Studio and compatible with GGUF quantizations for efficient CPU or GPU deployment.
- Developed by: Adel Jebali, Concordia University
- Funded by: SSHRC
- Language(s) (NLP): French
- License: Apache 2.0
- Finetuned: from Llama 3.1-8B
Uses
Correct you French written texts! It is not a chat model.
Bias, Risks, and Limitations
This AI LLM is not 100% bullet-proof. Errors are still possible.
How to use
Works best with LM Studio and is available in three quantization formats:
4-bit β Q4_K_M β 4.92 GB Best choice for low-memory systems and entry-level hardware. Recommended for Macs with only 8 GB of unified memory or older GPUs. Offers the fastest loading times and lowest VRAM usage, with a small trade-off in output quality and coherence.
6-bit β Q6_K β 6.6 GB Excellent balance between quality, speed, and memory consumption. A strong default option for most users with 12β16 GB of RAM or mid-range GPUs. In many cases, it delivers quality close to 8-bit while remaining significantly lighter.
8-bit β Q8_0 β 8.54 GB Highest quality and most faithful outputs among the available quantizations. Recommended if your hardware can handle it, especially with a dedicated GPU or Apple Silicon Mac with sufficient unified memory. Produces more stable generations, better grammatical consistency, and fewer hallucinations.
Recommendations
- 8 GB Macs: use Q4_K_M
- 16 GB systems: use Q6_K for the best balance
- 24 GB+ RAM or modern GPU: use Q8_0 for maximum quality
For optimal performance in LM Studio:
- Enable GPU offloading when available
- Increase the context length only if needed, since larger contexts consume more memory
- On Apple Silicon Macs, Metal acceleration significantly improves inference speed
- Temperature 0 (or 0.1)
- Min p 0
- Top k 0
- Top p 1
If your priority is:
- Maximum speed / lowest memory usage β Q4_K_M
- Best balance β Q6_K
- Best overall quality β Q8_0
Paper
A. Jebali, "Developing a Grammatical Error Correction System for French Second Language Written Texts," 2025 5th International Conference on Electrical, Computer and Energy Technologies (ICECET), Paris, France, 2025, pp. 1-6, doi: 10.1109/ICECET63943.2025.11472110.
- Downloads last month
- 12
4-bit
6-bit
8-bit
Model tree for adeljebali/llama3.1-gec-strict
Base model
meta-llama/Llama-3.1-8B