Cam1

Cam1 is a custom decision model based on the pinned Qwen3.5-4B backbone, a LoRA adapter, and a joint-schema decision head. It returns Choice probabilities and Noul decisions directly. This repository contains the model assets and the minimal inference runtime needed to use them. The upstream backbone is downloaded separately at the revision in inference_config.json.

The adapter, decision head and tokenizer files are byte-identical to the evaluated checkpoint. The export removes training-only metadata and optimizer state, so its packaging hash differs. It requires the custom decision-model inference runtime; the adapter alone is not the complete model. The validated serving policy is full FP32, reference kernels, GPU placement and a 16,384-token context limit.

The completed official Decision Index 0.3 public suite scored 30.72 Public points, with complete: true, 140,178 scoreable requests across 43 benchmarks, and 138,071 successful responses. The 2,107 scoreable refusals comprise 107 overlength requests and 2,000 PhishNChips inputs rejected by the evaluated schema validator. No evidence or options were truncated. See the submission results. Public points are chance-corrected, coverage-adjusted index points, not accuracy. This is not a maintainer Full score or leaderboard admission claim.

cam1_engine:Cam1Engine is the default evaluated behavior. The separate cam1_engine:Cam1CompatibilityEngine accepts empty Noul criteria and null Choice descriptions. That generic post-evaluation fix answered all 2,000 Phish inputs with 51.2% verdict accuracy and exactly preserved the 507-request regression set; it has no new complete-suite score. Its results are not merged into the submitted run. Both entrypoints use the same weights, tokenizer, prompts, FP32 numerical policy and temperature.

Reproduce inference

Use Linux, Python 3.12 and one NVIDIA RTX PRO 6000 with 96 GB VRAM. The evaluated system used Python 3.12.15, driver 580.178.04, Torch 2.11.0/CUDA 13.0, Transformers 5.10.2 and PEFT 0.18.1. The original container was python@sha256:edd0b3ec946cc68bd20e39480bd03929016d0fb6254cbf870cd3bc55ea25116d. Follow REPRODUCE.md for the pinned setup and runner command. runtime.zip contains only inference source, its dependency lock and release identity binding; no training scripts, optimizer state, planning notes or evaluation questions are included.

The runtime requires a fresh process per schema policy, full FP32 parameters, GPU-only placement, math SDPA/reference kernels, disabled TF32/autocast, deterministic algorithms, and a 16,384-token limit. Temperature is fixed at 1.333521432163324. It was originally calibrated under BF16; calibration validity under this FP32 runtime is unproven. Cross-runtime equivalence to BF16 is not claimed. Score-primitive quality is not established by the Choice/Noul evaluation.

Measured successful-request latency in the complete run was mean 408.97 ms, median 255.28 ms, p80 398.75 ms and p95 915.28 ms. These public-suite measurements do not certify the maintainers' separate held-out latency test.

The joint-schema head is derived from Cloudflare's Apache-2.0 implementation (revision 17f0b0ad64efb65d273590632833508766b2aae6). Qwen3.5-4B is attributed to the Qwen team. See LICENSE.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for cbuk10001/Cam1

Finetuned
Qwen/Qwen3.5-4B
Adapter
(730)
this model