YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Saathi Lite (preview) โ€“ Gemma 3 1B, Q4_0

Base model for the Lite tier of Saathi, a private offline journaling companion.

  • Converted with llama.cpp (b11192) from google/gemma-3-1b-it-qat-q4_0-unquantized. Google's own q4_0 GGUF has sliding_window 1024 and no freq_base_swa (the config says 512 / 10000), which breaks prompts longer than 512 tokens; this file fixes that.
  • File: gemma-3-1b-it-qat-q4_0-saathi.gguf, 720,425,696 bytes, sha256 fe611e51b2e038f26bf77a1eb29bac385cbf1b1256629b6b255d444d2e8f590b
  • Perplexity (60 dialogue pairs, n_ctx 256): this file 49.60, Google's q4_0 49.34.
  • Reply style adapter: saathi-lite-r1-f16.gguf, 26,116,992 bytes, sha256 13b0cd5c7135f049b3949d9180ce4e08e03d8eb89b2b10b465f303e59d31b316 Reply style adapter trained on 1,021 Saathi examples (LoRA r16) on the QAT-unquantized weights.

Gemma is provided under and subject to the Gemma Terms of Use found at https://ai.google.dev/gemma/terms

Downloads last month
47
GGUF
Model size
1.0B params
Architecture
gemma3
Hardware compatibility
Log In to add your hardware

4-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support