Kambo-v1 SQL + Code, DynQuant 4-bit

VikramPal/kambo-v1-sql-code quantized with DynQuant to 4.25 bits per parameter, scales included: the same byte budget as uniform 4-bit, spent unevenly. 0.836 GiB of weights instead of 3.150 GiB. The other width: 3-bit.

Against the bf16 fine-tune this checkpoint loses 4.85 points on text-to-SQL (49.06% against 53.91%, separated after Holm correction; by source (exploratory rows) Gretel -3.06, WikiSQL -10.76 and Spider dev -0.73, so WikiSQL accounts for 74% of the net loss in items) and is not separated from the fine-tune on HumanEval (+0.61) and MBPP (-0.60). Beside the unquantized base model it scores higher on text-to-SQL (49.06% against 42.87%) and MBPP (28.20% against 25.40%), and the same on HumanEval (31.10%); these are descriptions, not planned tests. At 0.13% fewer bytes than uniform 4-bit, it beats that control on text-to-SQL (+5.30), and is not separated from it on HumanEval (+6.71) and MBPP (+4.40, uncorrected p = 0.0115). On held-out loss (KL to the fine-tune, paired by conversation) it is closer to the fine-tune than uniform 4-bit (|z| = 23.0) and one draw of the permuted-signal null (|z| = 13.0).

What this is

quantized from VikramPal/kambo-v1-sql-code, the bf16 fine-tune of VikramPal/kambo-v1
method DynQuant 0.5.3, per-matrix widths of 2, 3, 4 or 8 bits from the fine-tune's own training signal, asymmetric, groups of 128
size 0.836 GiB of weights (4.2479 bits per parameter, scales and the bf16 remainder included); 898,079,352 bytes of safetensors
memory 0.836 GiB resident on the GPU after loading (weights and buffers, before any KV cache), 100.0% of the map's prediction; 0.836 GiB on disk
loads with transformers + dynquant, trust_remote_code=True, bfloat16 only. The usage snippet below ran against this repo's files with dynquant 0.5.3 under transformers 5.14.1 (torch 2.13.0+cu130) and 5.18.0 (torch 2.14.1+cu130)

Results

arm text-to-SQL (2,454) HumanEval (164) MBPP (500) weights bits/param
Kambo-v1 (base) 42.87% (1052/2454) 31.10% (51/164) 25.40% (127/500) 3.150 GiB 16.0000
fine-tune, bf16 53.91% (1323/2454) 30.49% (50/164) 28.80% (144/500) 3.150 GiB 16.0000
DynQuant 4-bit (this repo) 49.06% (1204/2454) 31.10% (51/164) 28.20% (141/500) 0.836 GiB 4.2479
uniform 4-bit 43.77% (1074/2454) 24.39% (40/164) 23.80% (119/500) 0.837 GiB 4.2535
DynQuant 3-bit 38.75% (951/2454) 19.51% (32/164) 20.00% (100/500) 0.640 GiB 3.2495
uniform 3-bit 24.33% (597/2454) 2.44% (4/164) 6.40% (32/500) 0.641 GiB 3.2538

Accuracy with the correct count. This arm was scored from this repo's packed checkpoint. The uniform arms put every quantized matrix at one width with the same quantizer and byte accounting. Each DynQuant arm's bytes are within 0.13% of its uniform control's.

Text-to-SQL by source:

arm Gretel test (a training source) WikiSQL test (a training source) Spider dev (not a training source; see below)
Kambo-v1 (base) 52.93% (433/818) 51.71% (423/818) 23.96% (196/818)
fine-tune, bf16 60.15% (492/818) 75.92% (621/818) 25.67% (210/818)
DynQuant 4-bit 57.09% (467/818) 65.16% (533/818) 24.94% (204/818)
uniform 4-bit 50.73% (415/818) 61.37% (502/818) 19.19% (157/818)
DynQuant 3-bit 45.48% (372/818) 56.72% (464/818) 14.06% (115/818)
uniform 3-bit 28.24% (231/818) 41.20% (337/818) 3.55% (29/818)

Spider is not one of the three training sources, but sql-create-context, which is, was built partly from Spider. Training rows asking a Spider dev question (after folding case, punctuation and whitespace) were removed; other Spider-derived rows (Spider train questions, for instance) can be in the training mix.

How this arm compares

McNemar exact over the per-item hits: every row pairs two arms on the same problems in the same order, so only the items the two arms disagree on (+ won by the first arm, − by the second) carry information. Delta is the first arm minus the second, in points. The 95% interval is exact and conditional on the number of disagreements (Clopper–Pearson on the first arm's share of them, scaled by their share of the items), so it excludes zero exactly when the unadjusted p is below 0.05. p (Holm) is step-down corrected across the 18 planned tests that could be computed (18 were declared before any fine-tuned arm was scored: 6 comparisons × 3 tasks). separated means Holm p < 0.05; not separated means this test cannot tell the two arms apart, not that they are equal: the interval shows how large a difference remains possible.

comparison task first second delta (pts) 95% CI disagreements p p (Holm) verdict
DynQuant 4-bit vs the bf16 fine-tune text-to-SQL 1204/2454 1323/2454 -4.85 [-6.22, -3.39] +110 / −229 9.48e-11 1.33e-09 separated
DynQuant 4-bit vs the bf16 fine-tune HumanEval 51/164 50/164 +0.61 [-5.70, +6.77] +13 / −12 1.00 1.00 not separated
DynQuant 4-bit vs the bf16 fine-tune MBPP 141/500 144/500 -0.60 [-3.30, +2.18] +21 / −24 0.766 1.00 not separated
DynQuant 4-bit vs uniform 4-bit text-to-SQL 1204/2454 1074/2454 +5.30 [+3.68, +6.84] +273 / −143 1.82e-10 2.36e-09 separated
DynQuant 4-bit vs uniform 4-bit HumanEval 51/164 40/164 +6.71 [-0.29, +12.28] +20 / −9 0.0614 0.249 not separated
DynQuant 4-bit vs uniform 4-bit MBPP 141/500 119/500 +4.40 [+0.95, +7.46] +46 / −24 0.0115 0.0744 not separated

Secondary and exploratory rows, not corrected for multiplicity (the per-source rows are cuts of the text-to-SQL row with the same label, and the pooled-code row is the union of the two code rows with that label; neither is further evidence):

comparison task first second delta (pts) 95% CI disagreements p
DynQuant 4-bit vs the bf16 fine-tune text-to-SQL / gretel 467/818 492/818 -3.06 [-5.05, -0.83] +27 / −52 0.00655
DynQuant 4-bit vs the bf16 fine-tune text-to-SQL / wikisql 533/818 621/818 -10.76 [-13.06, -8.03] +32 / −120 3.53e-13
DynQuant 4-bit vs the bf16 fine-tune text-to-SQL / spider 204/818 210/818 -0.73 [-3.29, +1.86] +51 / −57 0.631
DynQuant 4-bit vs the bf16 fine-tune code (humaneval+mbpp) 192/664 194/664 -0.30 [-2.86, +2.28] +34 / −36 0.905
DynQuant 4-bit vs uniform 4-bit text-to-SQL / gretel 467/818 415/818 +6.36 [+3.44, +9.00] +98 / −46 1.75e-05
DynQuant 4-bit vs uniform 4-bit text-to-SQL / wikisql 533/818 502/818 +3.79 [+0.75, +6.67] +90 / −59 0.0137
DynQuant 4-bit vs uniform 4-bit text-to-SQL / spider 204/818 157/818 +5.75 [+3.05, +8.16] +85 / −38 2.72e-05
DynQuant 4-bit vs uniform 4-bit code (humaneval+mbpp) 192/664 159/664 +4.97 [+1.93, +7.70] +66 / −33 0.00119
4-bit vs 3-bit text-to-SQL 1204/2454 951/2454 +10.31 [+8.77, +11.71] +357 / −104 1.64e-33
4-bit vs 3-bit text-to-SQL / gretel 467/818 372/818 +11.61 [+8.80, +13.99] +129 / −34 3.15e-14
4-bit vs 3-bit text-to-SQL / wikisql 533/818 464/818 +8.44 [+5.54, +10.99] +110 / −41 1.79e-08
4-bit vs 3-bit text-to-SQL / spider 204/818 115/818 +10.88 [+8.23, +13.07] +118 / −29 6.16e-14
4-bit vs 3-bit HumanEval 51/164 32/164 +11.59 [+4.46, +16.51] +26 / −7 0.00132
4-bit vs 3-bit MBPP 141/500 100/500 +8.20 [+4.69, +11.09] +61 / −20 5.66e-06
4-bit vs 3-bit code (humaneval+mbpp) 192/664 132/664 +9.04 [+5.99, +11.60] +87 / −27 1.53e-08

Held-out loss

Teacher-forced over the 999 conversations held out of the training mixture (2% of every stratum, never trained on): 90,517 assistant tokens. KL and argmax agreement compare each arm with the bf16 fine-tune token by token; NLL and token accuracy score each arm against the held-out reference text. This is the fine-tune's own training distribution, so it measures distance from the fine-tune there (for the quantized rows, what quantization did; for the base row, what fine-tuning did), not general ability.

arm NLL (nats/token) KL(fine-tune ‖ arm) argmax agrees with fine-tune token accuracy
fine-tune, bf16 (the reference) 0.1324 0.0000 100.00% 95.84%
Kambo-v1 (base) 0.1851 0.0602 97.24% 94.74%
DynQuant 4.25 map, encoded 0.1485 0.0166 98.15% 95.38%
DynQuant 4-bit, packed 0.1485 0.0166 98.15% 95.38%
uniform 4-bit 0.1648 0.0332 97.25% 94.88%
permuted-signal null, 4.25 (one draw) 0.1536 0.0211 97.81% 95.22%
DynQuant 3.25 map, encoded 0.1912 0.0594 96.19% 94.11%
DynQuant 3-bit, packed 0.1912 0.0594 96.19% 94.11%
uniform 3-bit 0.3476 0.2118 91.64% 90.24%
permuted-signal null, 3.25 (one draw) 0.2048 0.0718 95.69% 93.75%

Paired against the controls at 4.25 bits: each row is this arm minus the control on the same tokens, summed within each conversation, with the standard error clustered by conversation, since the tokens of one conversation are not independent. A negative KL difference means this arm is closer to the fine-tune; a negative NLL difference means it puts more probability on the held-out reference text. Each pair is read on its KL z at |z| > 1.96, uncorrected (every KL result on these cards also clears a Bonferroni bar across all 4 pairs, |z| > 2.50); the NLL z is shown, not read.

against KL difference (this arm − it) z conversations where this arm is closer NLL difference z
uniform 4-bit -0.01659 -23.0 866 of 999 -0.01626 -15.6
permuted-signal null, 4.25 (one draw) -0.00450 -13.0 696 of 999 -0.00509 -7.2

The permuted-signal null is one seeded within-role permutation (--score-null shuffle --null-seed 0): it moved each module's recorded score, and its measured sensitivity where it has one, to another module of the same role (190 of 205 modules moved; the other 15 drew their own place or, like the embedding, are alone in their role). No module gained or lost a measured sensitivity in the move. Its z is conditional on that draw and does not include the spread across shuffles.

Packed and encoded agree exactly on the held-out loss. The same map scored through this repo's packed checkpoint and through bf16 encoding gives identical per-token losses and predictions, so on this loss the uniform arms (scored encoded, here and on the tasks), the null arms (scored encoded) and this arm (scored packed) differ in their bit widths and in nothing else. Task scores were not cross-checked between the two storage paths.

How the bits were allocated

DynQuant gives every quantized module its own width from {2, 3, 4, 8} bits (each layer's 16 routed experts share one width per batched bank), with groups of 128 weights sharing one scale and one zero point (asymmetric). Those cost 0.25 bits per weight, so a uniform 4-bit recipe costs 4.2535 bits per parameter here, counting scales, zero points and the bf16 remainder below, and this map was asked for 4.25. It achieved 4.2479 bits (898,002,432 bytes) against uniform 4-bit's 4.2535 (899,182,080 bytes): -0.13% bytes. The uniform figures are re-priced: DynQuant 0.5.3 charges a uniform map for 448,512 of the 499,456 bf16 remainder parameters (it leaves out the 50,944 norm weights), and here it pays for the whole remainder, as this map does.

Widths come from the fine-tune itself. During training DynQuant's tracker recorded each trained module's gradient-norm variance across optimizer steps and its activation RMS and, for the 133 single-matmul modules, the channel moments from which the loss's sensitivity to quantization is computed; the allocator spends the byte budget where each byte buys the largest drop in priced damage, subject to per-role floors.

width parameters share
2-bit 169,869,312 10.1%
3-bit 509,607,936 30.1%
4-bit 801,243,136 47.4%
8-bit 209,977,344 12.4%

By module family (one row per projection, summed over the layers that have it; the model has 24):

family modules parameters widths (share of the family's parameters)
moe.w1 24 452,984,832 4b 100%
moe.w2 24 452,984,832 2b 21%, 3b 54%, 4b 25%
moe.w3 24 452,984,832 2b 17%, 3b 58%, 4b 25%
model.embed_tokens 1 155,582,464 8b 100%
conv.in_proj 18 56,623,104 4b 72%, 8b 28%
moe.sw1 24 28,311,552 4b 96%, 8b 4%
moe.sw2 24 28,311,552 4b 63%, 8b 37%
moe.sw3 24 28,311,552 4b 79%, 8b 21%
conv.out_proj 18 18,874,368 4b 67%, 8b 33%
self_attn.o_proj 6 6,291,456 4b 17%, 8b 83%
self_attn.q_proj 6 6,291,456 8b 100%
self_attn.k_proj 6 1,572,864 8b 100%
self_attn.v_proj 6 1,572,864 8b 100%

Kept in bf16 and not quantized: 499,456 parameters (0.030% of the model): moe.router (24 tensors, 393,216), conv.conv (18 tensors, 55,296), input_layernorm (24 tensors, 24,576), post_attention_layernorm (24 tensors, 24,576), model.norm (1 tensor, 1,024), self_attn.k_norm (6 tensors, 384), self_attn.q_norm (6 tensors, 384). These are the 24 routers, which pick each token's experts, the short-convolution kernels and the norms, all left in the compute dtype by DynQuant's classification.

72 modules holding 1,358,954,496 parameters (80.4%) were priced by a proxy, not by measurement. DynQuant prices a module from the channel moments its hooks collect, and it collects them only for modules that are a single matmul (the tied embedding is measured through lm_head). Kambo stores each layer's 16 routed experts as three batched tensors (moe.w1, moe.w2, moe.w3) whose forward spans two matmuls and a nonlinearity, so the tracker followed those banks (measure_expert_banks=True; 72 recorded) and recorded their gradient statistics and activation RMS, but no moments. Those banks were priced from the tracker's plasticity score (each bank's within-role rank of log1p of its gradient-norm variance across optimizer steps; DynQuant 0.5.3's default score does not use saliency) times their size times an error curve, scaled to the measured modules' units: the allocator's fallback. The other 133 modules (331,743,232 parameters) were priced from measured moments.

Every module is at or above its role's floor.

Expert banks are grouped along their output axis. A Linear's weight is [out, in] and DynQuant groups along the stored last axis, the input. Kambo stores its expert banks input-major ([experts, in, out]), so the same rule groups each of their 128-weight blocks across 128 output channels of one input. Measured on the base model's layers 0, 12 and 23 at 4 bits, the reconstruction error of the shipped grouping relative to grouping along the input is 0.995 for w1, 1.017 for w2 and 0.999 for w3 (1.000 = no difference, above 1 = the shipped grouping is worse); the banks were not transposed.

Evaluation

All scores come from dynquant eval (dynquant 0.5.3) with the transformers backend, bf16, the chat template, greedy decoding and every decode setting pinned identically across arms:

task items prompt max new tokens scored by
text-to-SQL 2,454 2 solved examples as prior chat turns, then the question 320 execution match: the query runs against the item's database and its result set must equal the reference query's
HumanEval 164 one user turn: complete the function, in a single code block 1024 the item's unit tests, pass@1
MBPP 500 (test split) one user turn: the task and its tests 1024 the item's unit tests, pass@1

Text-to-SQL deals 818 items from each of Gretel's test split, WikiSQL's test split and Spider's 1,034-item dev set (its validation split), in rotation. An item is admitted only if its database holds rows and its reference query returns some, and not a single row of NULLs and zeros, so a wrong query cannot match by also returning nothing; items whose schema and rows exceed 6,000 characters are skipped. Gretel's schemas carry their own INSERTs, WikiSQL's databases are built from its real Wikipedia tables, and Spider's databases, rows included, are inlined from a mirror. MBPP's records say shots: 3, but the chat framing ignores exemplars (DynQuant logs that it does), so every MBPP prompt is the single turn above. Generated code runs in a sandbox (exec/linux/py3.12/rlimits/t=8s/m=4096MB).

Decoding is deterministic. Kambo's experts are summed with a bf16 index_add whose CUDA atomics round in arrival order, and a top-2 router can turn that last bit into a different expert, so two plain runs of one checkpoint disagree on a few items. Every arm was therefore run under torch.use_deterministic_algorithms(True); a repeat of 96 text-to-SQL items reproduced every prediction (EXACT). Launchers recorded across the arms above: dq_det. The base model's first, plain-launcher run is kept as a secondary row on the bf16 card, so the size of the launcher effect is on record.

Greedy is checked, not assumed. The checkpoints were evaluated with the generation defaults inherited from the base model, which sample (below; this repo's own are greedy); the evaluation overrides them, and a check on the fine-tune confirmed that its generations are greedy: 8 of 8 generations were identical under seeds 1 and 2, and 583 of 584 generated tokens are the argmax of a teacher-forced pass over the same text; the one that is not trails it by 0.125 logits, a near-tie inside the check's 0.25-logit tolerance.

Training, data and decontamination are described on the bf16 fine-tune's card.

What is not claimed

  • Storage and resident memory are measured; speed is not. This card makes no claim about decode throughput or latency against bf16.
  • One runtime. Loaded and scored through transformers with DynQuant's quantizer. Not tested with vLLM, llama.cpp, TGI or any other runtime; Kambo's architecture is custom code, so most will not load it at all.
  • bfloat16 only. Loading with dtype=torch.float32 fails at the first layer (expected scalar type Half but found Float): the packed embedding emits its scales' dtype, which the format pins to fp16 under an fp32 model. bf16 loads and runs on GPU and CPU.
  • The allocation is mostly proxy-priced (see above): DynQuant's measured sensitivities cover 19.6% of the quantized parameters.
  • Two skills. Text-to-SQL and Python function writing, plus held-out loss on the training mixture. A quantization that holds these can lose something else.

Usage

pip install dynquant torch transformers accelerate
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

import dynquant

# Before from_pretrained. Without it transformers does not recognise the packed
# format, warns, and returns a model with randomly initialised weights.
dynquant.register_hf_quantizer()

model_id = "VikramPal/kambo-v1-sql-code-DynQuant-4bit"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id, trust_remote_code=True, dtype=torch.bfloat16, device_map="cuda"
)

schema = "CREATE TABLE employees (id INTEGER, name TEXT, dept TEXT, salary INTEGER);"
question = "Which employees in Sales earn more than 50000?"
prompt = (
    "Write a single SQL query that answers the question, using only the tables in the "
    "schema. Return just the query, with no explanation.\n\n"
    f"Schema:\n{schema}\n\nQuestion: {question}"
)
messages = [{"role": "user", "content": prompt}]
inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt", return_dict=True
).to(model.device)
out = model.generate(**inputs, max_new_tokens=320)  # greedy: see generation_config.json
print(tokenizer.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))

This repo's generation_config.json is greedy, which is a change from the base model's. Kambo-v1 ships do_sample: true, temperature 0.7, top_p 0.9 and top_k 2; the top_k is the MoE routing width (top_k: 2 in config.json) carried into the generation defaults, and it restricts every sampled token to the two most likely. Every number on this card was measured greedy, so a plain generate() call here decodes greedily too. To sample, pass do_sample=True with your own temperature, top_p and top_k. transformers 5 still fills the unset top_k from config.json and warns that it "may be ignored"; greedy decoding does ignore it.

Prompt format

The model was trained and evaluated on these wordings, and answers best when asked in them. ChatML, no system message (none is inserted when you supply none, which is how it was trained).

Text-to-SQL
Write a single SQL query that answers the question, using only the tables in the schema. Return just the query, with no explanation.

Schema:
{CREATE TABLE ... statements}

Question: {question}
Python function from a signature and docstring (HumanEval style)
Complete the following Python function. Write the entire function, including the signature, inside a single ```python code block. Do not write tests, examples, or an explanation.

```python
{signature and docstring}```
Python function from a description and tests (MBPP style)
You are an expert Python programmer. Write a Python function for this task:

{description}

Your code must pass these tests:

```python
{assert statements}
```

Return only the function, in a single ```python code block, with no explanation.

License

Released under the Apache License 2.0, as the base model is; see NOTICE. Training data, each under its own license: gretelai/synthetic_text_to_sql (apache-2.0), Salesforce/wikisql (unknown, as the dataset card states it), b-mc2/sql-create-context (cc-by-4.0), nvidia/OpenCodeInstruct (cc-by-4.0). Evaluated on: gretelai/synthetic_text_to_sql (apache-2.0), Salesforce/wikisql (unknown, as the dataset card states it), xlangai/spider (cc-by-sa-4.0), premai-io/spider (no license stated on the dataset card), openai/openai_humaneval (mit), google-research-datasets/mbpp (cc-by-4.0).

Citation

@misc{kambo_v1_2026,
  title  = {Kambo-v1: A 1.7B Hybrid Convolution-Attention Mixture-of-Experts Language Model},
  author = {Kamboj, Vikrampal},
  year   = {2026},
  note   = {Apache-2.0},
  url    = {https://huggingface.co/VikramPal/kambo-v1}
}
Downloads last month
26
Safetensors
Model size
0.2B params
Tensor type
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for VikramPal/kambo-v1-sql-code-DynQuant-4bit

Quantized
(2)
this model

Datasets used to train VikramPal/kambo-v1-sql-code-DynQuant-4bit