Instructions to use krish99/saathi-lite-preview with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use krish99/saathi-lite-preview with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf krish99/saathi-lite-preview:Q4_0 # Run inference directly in the terminal: llama cli -hf krish99/saathi-lite-preview:Q4_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf krish99/saathi-lite-preview:Q4_0 # Run inference directly in the terminal: llama cli -hf krish99/saathi-lite-preview:Q4_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf krish99/saathi-lite-preview:Q4_0 # Run inference directly in the terminal: ./llama-cli -hf krish99/saathi-lite-preview:Q4_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf krish99/saathi-lite-preview:Q4_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf krish99/saathi-lite-preview:Q4_0
Use Docker
docker model run hf.co/krish99/saathi-lite-preview:Q4_0
- LM Studio
- Jan
- Ollama
How to use krish99/saathi-lite-preview with Ollama:
ollama run hf.co/krish99/saathi-lite-preview:Q4_0
- Unsloth Desktop
- Docker Model Runner
How to use krish99/saathi-lite-preview with Docker Model Runner:
docker model run hf.co/krish99/saathi-lite-preview:Q4_0
- Lemonade
How to use krish99/saathi-lite-preview with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull krish99/saathi-lite-preview:Q4_0
Run and chat with the model
lemonade run user.saathi-lite-preview-Q4_0
List all available models
lemonade list
- Atomic Chat
YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Saathi Lite (preview) โ Gemma 3 1B, Q4_0
Base model for the Lite tier of Saathi, a private offline journaling companion.
- Converted with llama.cpp (b11192) from google/gemma-3-1b-it-qat-q4_0-unquantized. Google's own q4_0 GGUF has sliding_window 1024 and no freq_base_swa (the config says 512 / 10000), which breaks prompts longer than 512 tokens; this file fixes that.
- File: gemma-3-1b-it-qat-q4_0-saathi.gguf, 720,425,696 bytes, sha256 fe611e51b2e038f26bf77a1eb29bac385cbf1b1256629b6b255d444d2e8f590b
- Perplexity (60 dialogue pairs, n_ctx 256): this file 49.60, Google's q4_0 49.34.
- Reply style adapter: saathi-lite-r1-f16.gguf, 26,116,992 bytes, sha256 13b0cd5c7135f049b3949d9180ce4e08e03d8eb89b2b10b465f303e59d31b316 Reply style adapter trained on 1,021 Saathi examples (LoRA r16) on the QAT-unquantized weights.
Gemma is provided under and subject to the Gemma Terms of Use found at https://ai.google.dev/gemma/terms
- Downloads last month
- 47
Hardware compatibility
Log In to add your hardware
4-bit
16-bit
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support