SAMS bounded language model v3 (prototype)

The Transformers checkpoint is in merged-hf/, the LoRA adapter is in adapter/, and Q4_K_M/Q5_K_M files are in gguf/. Evaluation results and the artifact manifest are included. The model only renders supplied facts and selects a supplied question ID. Deterministic services retain all medical, emergency, robot motion, navigation, power, and hardware decisions. The synthetic data is prototype-unreviewed, so this model is not medically or robotically validated and must not be deployed as a safety controller.

Current release decision: REJECTED for user-facing SAMS

An independently authored 200-case quality audit was run against each GGUF after the original contract test. The final rendered utterance was evaluated as model speech followed by the deterministic approved question selected by next_question_id.

Model Independent quality passes Pass rate 90% gate
Q4_K_M 140 / 200 70.0% FAIL
Q5_K_M 151 / 200 75.5% FAIL

Both models retained 100% strict JSON, contract-field, trusted-number, prompt-injection, and no-unsafe-command performance in this audit. They failed user-facing quality because they frequently omitted trusted measurement values, produced canned check-in language, and did not reliably continue natural climber questions. The current artifacts are preserved for research and debugging but are not approved for the SAMS user-facing release.

The existing-output audit also found internal or robotic language in 405/630 Q4 cases and 411/630 Q5 cases. All 25 low-confidence transcript cases in each original suite reached Qwen and collapsed to one canned response; low-confidence STT must instead bypass the LLM and use deterministic input recovery.

Downloads

Original locked contract evaluation

The complete 630-case locked set was run against each downloadable GGUF with temperature 0, thinking disabled, and strict constrained JSON decoding. This produced 1,260 recorded generations. This test measures the narrow machine contract only; it does not override the failed independent user-facing quality gate above.

Model Full-contract passes Pass rate Gate
Q4_K_M 619 / 630 98.25% PASS
Q5_K_M 609 / 630 96.67% PASS

Both pass the required 90% full-contract threshold. Strict JSON, exact schema, value types, speech length, supplied questions, fact IDs, forbidden advice/commands, required states, required questions, and required fact echoes were 100% for both quantizations. The 11 Q4 and 21 Q5 misses were all numbers_grounded failures. Most claimed an unsupported 100 percent; a few spoke numeric identifier suffixes that the conservative checker did not accept. Production integration should retain the deterministic number-grounding post-check and reject or regenerate every flagged response.

These scores measure the bounded language/output contract on synthetic prototype data. They are not general language-model accuracy, medical validation, field safety validation, or robot-controller validation.

Downloads last month
-
GGUF
Model size
2B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for iteratehack/sam-qwen3-1.7b-sams-v3

Finetuned
Qwen/Qwen3-1.7B
Quantized
(354)
this model