Access PetInst-LLM under the Gemma Terms

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

This is a modified Gemma model derivative. By accessing it, you agree to the Gemma Terms of Use and incorporated Prohibited Use Policy linked in this repository.

Log in or Sign Up to review the conditions and access this model content.

PetInst-LLM 1B v6.18 โ€” Q8_0 GGUF

PetInst-LLM 1B is an independently modified research derivative of Google's google/gemma-3-1b-it, specialized for one silent, schema-constrained virtual-pet function decision. This repository is the Q8_0 GGUF conversion of manfye/PetInst-LLM-1B-MLX.

This project is not affiliated with, sponsored by, or endorsed by Google. Google DeepMind created the upstream Gemma model; the PetInst-LLM project created and evaluated these modifications.

Artifact

  • File: petinst-llm-1b-v6.18-q8_0.gguf
  • Quantization: Q8_0
  • Size: 1,069,306,336 bytes (1,019.77 MiB)
  • SHA-256: 61179bfe75d067f371d9165b17e94fb028c0b0a724e66fdde84b9656d8c513f7
  • Runtime: llama.cpp-compatible GGUF; intended Expo integration uses llama.rn
  • Serialization: petinst.json-call/v1.3

Required decoding contract

Supply three to five complete semantic calls and constrain generation to exactly one declared call. The selected call contains name plus bounded targetId, effect, emotion, intensity, and manner arguments. Trusted application code validates the result and owns every state mutation.

Evaluation

On the answer-free 220-row Control D set with the same v1.3 prompt and declared-call constraint:

Metric Result
Exact/tool accuracy 97.27%
Strict validity 100%
Memory-policy accuracy 95%
Memory read/write 100% / 100%
Negative-expression accuracy 90%
State-precedence accuracy 100%
Safety accuracy 100%
Mean / P95 / max latency 0.797 / 0.889 / 0.932 s
Mean generation rate 69.02 tokens/s

Latency was measured sequentially on an Apple M4 using llama.cpp-server with Metal and a 4,096-token runtime context. It is not physical-phone, llama.rn, peak-RSS, thermal, or airplane-mode proof.

This Q8_0 artifact preserves the selected MLX policy's 97.27% result. It supersedes an earlier rejected affine-Q5 conversion that reached only 85.91%.

Limitations

  • Research-only and release-unapproved.
  • The current Expo app still needs a PetInst JSON-call adapter before this can replace its existing model contract.
  • Generation must be constrained to the declared complete calls; unconstrained behavior is not represented by these metrics.
  • Only English synthetic research data was evaluated.

Terms and modification notice

Access and use are subject to the Gemma Terms of Use and the incorporated Gemma Prohibited Use Policy. See NOTICE and MODIFICATIONS.md.

Downloads last month
-
GGUF
Model size
1.0B params
Architecture
gemma3
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for manfye/PetInst-LLM-1B-GGUF

Quantized
(1)
this model