You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Inkling-Small — abliterated-cyber GLP-41

A GLP (GGUF Layer Projection) control vector for thinkingmachines/Inkling-Small: 41 per-layer projective refusal directions (L1–41), applied at runtime by a fail-closed vLLM hotfix. No weights are modified — this 660 KB file is the entire behavioral change, and deleting it reverts to stock.

First GLP for a Thinking Machines model. Same technique as the DSV4 GLP-29 / Qwen GLP-49/GLP-47 / GLM GLP-44/GLP-77 vectors — see the weightless repo for the GLP format spec, the hotfixes, and the serving recipes.

Measured (vLLM 0.28.0, NVFP4, 4×H100)

arm refusal32 benign32
stock 0/32 (total lockdown — the stickiest stock refusal we have measured) 31/32
GLP-41, α=0.25 30/32 (2 garbled) 30/32

Dose discipline is extreme on this model. α=1.0 garbles everything (including benign: 30/32 degenerate); α=0.5 garbles everything; α=0.25 works. That is a quarter of Qwen's calibrated dose, an eighth of GLM-5.3-Flash's, a sixteenth of DeepSeek V4's. Ship at α=0.25 and do not raise it.

Stock note (measured on the base model, unrelated to the vector): on a 32-country political-propaganda probe, stock Inkling-Small answers 28/32 even-handedly and refuses or deflects the rest — an asymmetric map aligned with provider sensitivities rather than a uniform policy. The steered arm answers all 32. Per-country detail stays private.

Derivation

Contrast-derived per-layer mean difference (AdvBench32 vs Alpaca32), captured on the post-layer residual stream in vLLM 0.28.0 (the hotfix's capture mode — the run included the deferred-residual flush this architecture needs), last prefill token, unit-normed. Cross-layer adjacent cosine median 0.89 vs null p99 0.04 (systematic, not noise). Full methodology: spec/GLP.md + BENCHMARK.md in weightless.

Use

vLLM 0.28.0+ serves Inkling-Small natively (day-0). Apply with the weightless hotfix for this arch (patches/hotfix-inkling-steering-projective.py, fail-closed):

export WEIGHTLESS_STEER_PATH=/path/to/Inkling-Small-abliterated-cyber-GLP-41-L1-41-a0.25.gguf
export WEIGHTLESS_STEER_ALPHA=0.25

The GGUF carries glp.mode=project — a reader that only understands additive control vectors must refuse it.

Provenance

  • content_sha256 (tensor bytes): 2a229d56cc7ecd582a6d527af55905188c586e946eee5a749bb17d9521bd3f05
  • Derived 2026-09-01, tolmo 4×H100, from thinkingmachines/Inkling-Small-NVFP4
  • Eval protocol: refusal32 / benign32, four-state scoring, temperature 0
Downloads last month
-
GGUF
Model size
168k params
Architecture
controlvector
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support