Kambo-v1 SQL + Code, DynQuant 3-bit

VikramPal/kambo-v1-sql-code quantized with DynQuant to 3.25 bits per parameter, scales included: the same byte budget as uniform 3-bit, spent unevenly. 0.640 GiB of weights instead of 3.150 GiB. The other width: 4-bit.

This 3-bit checkpoint scores below the base model on all three tasks, so it is published as a measurement of DynQuant at 3 bits, not as a model to use in place of the base model. The 4-bit checkpoint scores above the base model on text-to-SQL and MBPP, and level with it on HumanEval.

At 3.25 bits this checkpoint scores below the unquantized base model on all three tasks (text-to-SQL 38.75% against 42.87%, HumanEval 19.51% against 31.10%, MBPP 20.00% against 25.40%; a description, not a planned test). Against the bf16 fine-tune it loses 15.16 points on text-to-SQL, 10.98 on HumanEval and 8.80 on MBPP, all separated after Holm correction. At 0.13% fewer bytes than uniform 3-bit, it beats that control on text-to-SQL (+14.43), HumanEval (+17.07) and MBPP (+13.60). On held-out loss (KL to the fine-tune, paired by conversation) it is the closest of the 3.25-bit allocations measured: closer to the fine-tune than uniform 3-bit (|z| = 47.9) and one draw of the permuted-signal null (|z| = 15.0). Its KL to the fine-tune is 0.0594 against the unquantized base model's 0.0602, and it picks the fine-tune's top token at 96.19% of positions against the base model's 97.24% (descriptions, not tests). In secondary comparisons, the 4-bit version, 211,058,688 bytes (0.197 GiB) larger, is ahead of this one on text-to-SQL (+10.31, uncorrected p = 1.64e-33), HumanEval (+11.59, uncorrected p = 0.00132) and MBPP (+8.20, uncorrected p = 5.66e-06).

What this is

quantized from VikramPal/kambo-v1-sql-code, the bf16 fine-tune of VikramPal/kambo-v1
method DynQuant 0.5.3, per-matrix widths of 2, 3, 4 or 8 bits from the fine-tune's own training signal, asymmetric, groups of 128
size 0.640 GiB of weights (3.2495 bits per parameter, scales and the bf16 remainder included); 687,020,624 bytes of safetensors
memory 0.640 GiB resident on the GPU after loading (weights and buffers, before any KV cache), 100.0% of the map's prediction; 0.640 GiB on disk
loads with transformers + dynquant, trust_remote_code=True, bfloat16 only. The usage snippet below ran against this repo's files with dynquant 0.5.3 under transformers 5.14.1 (torch 2.13.0+cu130) and 5.18.0 (torch 2.14.1+cu130)

Results

arm text-to-SQL (2,454) HumanEval (164) MBPP (500) weights bits/param
Kambo-v1 (base) 42.87% (1052/2454) 31.10% (51/164) 25.40% (127/500) 3.150 GiB 16.0000
fine-tune, bf16 53.91% (1323/2454) 30.49% (50/164) 28.80% (144/500) 3.150 GiB 16.0000
DynQuant 4-bit 49.06% (1204/2454) 31.10% (51/164) 28.20% (141/500) 0.836 GiB 4.2479
uniform 4-bit 43.77% (1074/2454) 24.39% (40/164) 23.80% (119/500) 0.837 GiB 4.2535
DynQuant 3-bit (this repo) 38.75% (951/2454) 19.51% (32/164) 20.00% (100/500) 0.640 GiB 3.2495
uniform 3-bit 24.33% (597/2454) 2.44% (4/164) 6.40% (32/500) 0.641 GiB 3.2538

Accuracy with the correct count. This arm was scored from this repo's packed checkpoint. The uniform arms put every quantized matrix at one width with the same quantizer and byte accounting. Each DynQuant arm's bytes are within 0.13% of its uniform control's.

Text-to-SQL by source:

arm Gretel test (a training source) WikiSQL test (a training source) Spider dev (not a training source; see below)
Kambo-v1 (base) 52.93% (433/818) 51.71% (423/818) 23.96% (196/818)
fine-tune, bf16 60.15% (492/818) 75.92% (621/818) 25.67% (210/818)
DynQuant 4-bit 57.09% (467/818) 65.16% (533/818) 24.94% (204/818)
uniform 4-bit 50.73% (415/818) 61.37% (502/818) 19.19% (157/818)
DynQuant 3-bit 45.48% (372/818) 56.72% (464/818) 14.06% (115/818)
uniform 3-bit 28.24% (231/818) 41.20% (337/818) 3.55% (29/818)

Spider is not one of the three training sources, but sql-create-context, which is, was built partly from Spider. Training rows asking a Spider dev question (after folding case, punctuation and whitespace) were removed; other Spider-derived rows (Spider train questions, for instance) can be in the training mix.

How this arm compares

McNemar exact over the per-item hits: every row pairs two arms on the same problems in the same order, so only the items the two arms disagree on (+ won by the first arm, − by the second) carry information. Delta is the first arm minus the second, in points. The 95% interval is exact and conditional on the number of disagreements (Clopper–Pearson on the first arm's share of them, scaled by their share of the items), so it excludes zero exactly when the unadjusted p is below 0.05. p (Holm) is step-down corrected across the 18 planned tests that could be computed (18 were declared before any fine-tuned arm was scored: 6 comparisons × 3 tasks). separated means Holm p < 0.05; not separated means this test cannot tell the two arms apart, not that they are equal: the interval shows how large a difference remains possible.

comparison task first second delta (pts) 95% CI disagreements p p (Holm) verdict
DynQuant 3-bit vs the bf16 fine-tune text-to-SQL 951/2454 1323/2454 -15.16 [-16.51, -13.65] +91 / −463 5.66e-61 1.02e-59 separated
DynQuant 3-bit vs the bf16 fine-tune HumanEval 32/164 50/164 -10.98 [-16.28, -3.66] +8 / −26 0.00294 0.0264 separated
DynQuant 3-bit vs the bf16 fine-tune MBPP 100/500 144/500 -8.80 [-11.62, -5.32] +19 / −63 1.15e-06 1.15e-05 separated
DynQuant 3-bit vs uniform 3-bit text-to-SQL 951/2454 597/2454 +14.43 [+12.54, +16.19] +519 / −165 1.72e-43 2.92e-42 separated
DynQuant 3-bit vs uniform 3-bit HumanEval 32/164 4/164 +17.07 [+11.99, +18.26] +29 / −1 5.77e-08 6.93e-07 separated
DynQuant 3-bit vs uniform 3-bit MBPP 100/500 32/500 +13.60 [+10.42, +15.85] +80 / −12 1.72e-13 2.57e-12 separated

Secondary and exploratory rows, not corrected for multiplicity (the per-source rows are cuts of the text-to-SQL row with the same label, and the pooled-code row is the union of the two code rows with that label; neither is further evidence):

comparison task first second delta (pts) 95% CI disagreements p
DynQuant 3-bit vs the bf16 fine-tune text-to-SQL / gretel 372/818 492/818 -14.67 [-16.89, -11.95] +29 / −149 1.20e-20
DynQuant 3-bit vs the bf16 fine-tune text-to-SQL / wikisql 464/818 621/818 -19.19 [-21.51, -16.34] +31 / −188 1.33e-28
DynQuant 3-bit vs the bf16 fine-tune text-to-SQL / spider 115/818 210/818 -11.61 [-13.89, -8.89] +31 / −126 8.67e-15
DynQuant 3-bit vs the bf16 fine-tune code (humaneval+mbpp) 132/664 194/664 -9.34 [-11.90, -6.28] +27 / −89 6.44e-09
DynQuant 3-bit vs uniform 3-bit text-to-SQL / gretel 372/818 231/818 +17.24 [+13.76, +20.31] +197 / −56 1.40e-19
DynQuant 3-bit vs uniform 3-bit text-to-SQL / wikisql 464/818 337/818 +15.53 [+11.31, +19.45] +225 / −98 1.22e-12
DynQuant 3-bit vs uniform 3-bit text-to-SQL / spider 115/818 29/818 +10.51 [+8.58, +11.83] +97 / −11 2.39e-18
DynQuant 3-bit vs uniform 3-bit code (humaneval+mbpp) 132/664 36/664 +14.46 [+11.93, +16.24] +109 / −13 4.68e-20
4-bit vs 3-bit text-to-SQL 1204/2454 951/2454 +10.31 [+8.77, +11.71] +357 / −104 1.64e-33
4-bit vs 3-bit text-to-SQL / gretel 467/818 372/818 +11.61 [+8.80, +13.99] +129 / −34 3.15e-14
4-bit vs 3-bit text-to-SQL / wikisql 533/818 464/818 +8.44 [+5.54, +10.99] +110 / −41 1.79e-08
4-bit vs 3-bit text-to-SQL / spider 204/818 115/818 +10.88 [+8.23, +13.07] +118 / −29 6.16e-14
4-bit vs 3-bit HumanEval 51/164 32/164 +11.59 [+4.46, +16.51] +26 / −7 0.00132
4-bit vs 3-bit MBPP 141/500 100/500 +8.20 [+4.69, +11.09] +61 / −20 5.66e-06
4-bit vs 3-bit code (humaneval+mbpp) 192/664 132/664 +9.04 [+5.99, +11.60] +87 / −27 1.53e-08

Held-out loss

Teacher-forced over the 999 conversations held out of the training mixture (2% of every stratum, never trained on): 90,517 assistant tokens. KL and argmax agreement compare each arm with the bf16 fine-tune token by token; NLL and token accuracy score each arm against the held-out reference text. This is the fine-tune's own training distribution, so it measures distance from the fine-tune there (for the quantized rows, what quantization did; for the base row, what fine-tuning did), not general ability.

arm NLL (nats/token) KL(fine-tune ‖ arm) argmax agrees with fine-tune token accuracy
fine-tune, bf16 (the reference) 0.1324 0.0000 100.00% 95.84%
Kambo-v1 (base) 0.1851 0.0602 97.24% 94.74%
DynQuant 4.25 map, encoded 0.1485 0.0166 98.15% 95.38%
DynQuant 4-bit, packed 0.1485 0.0166 98.15% 95.38%
uniform 4-bit 0.1648 0.0332 97.25% 94.88%
permuted-signal null, 4.25 (one draw) 0.1536 0.0211 97.81% 95.22%
DynQuant 3.25 map, encoded 0.1912 0.0594 96.19% 94.11%
DynQuant 3-bit, packed 0.1912 0.0594 96.19% 94.11%
uniform 3-bit 0.3476 0.2118 91.64% 90.24%
permuted-signal null, 3.25 (one draw) 0.2048 0.0718 95.69% 93.75%

Paired against the controls at 3.25 bits: each row is this arm minus the control on the same tokens, summed within each conversation, with the standard error clustered by conversation, since the tokens of one conversation are not independent. A negative KL difference means this arm is closer to the fine-tune; a negative NLL difference means it puts more probability on the held-out reference text. Each pair is read on its KL z at |z| > 1.96, uncorrected (every KL result on these cards also clears a Bonferroni bar across all 4 pairs, |z| > 2.50); the NLL z is shown, not read.

against KL difference (this arm − it) z conversations where this arm is closer NLL difference z
uniform 3-bit -0.15240 -47.9 990 of 999 -0.15635 -44.5
permuted-signal null, 3.25 (one draw) -0.01235 -15.0 687 of 999 -0.01357 -11.6

The permuted-signal null is one seeded within-role permutation (--score-null shuffle --null-seed 0): it moved each module's recorded score, and its measured sensitivity where it has one, to another module of the same role (190 of 205 modules moved; the other 15 drew their own place or, like the embedding, are alone in their role). No module gained or lost a measured sensitivity in the move. Its z is conditional on that draw and does not include the spread across shuffles.

This map is closer to the fine-tune than every control at 3.25 bits (|z| ≥ 15.0; the permuted-signal null is one draw).

Packed and encoded agree exactly on the held-out loss. The same map scored through this repo's packed checkpoint and through bf16 encoding gives identical per-token losses and predictions, so on this loss the uniform arms (scored encoded, here and on the tasks), the null arms (scored encoded) and this arm (scored packed) differ in their bit widths and in nothing else. Task scores were not cross-checked between the two storage paths.

How the bits were allocated

DynQuant gives every quantized module its own width from {2, 3, 4, 8} bits (each layer's 16 routed experts share one width per batched bank), with groups of 128 weights sharing one scale and one zero point (asymmetric). Those cost 0.25 bits per weight, so a uniform 3-bit recipe costs 3.2538 bits per parameter here, counting scales, zero points and the bf16 remainder below, and this map was asked for 3.25. It achieved 3.2495 bits (686,943,744 bytes) against uniform 3-bit's 3.2538 (687,844,864 bytes): -0.13% bytes. The uniform figures are re-priced: DynQuant 0.5.3 charges a uniform map for 448,512 of the 499,456 bf16 remainder parameters (it leaves out the 50,944 norm weights), and here it pays for the whole remainder, as this map does.

Widths come from the fine-tune itself. During training DynQuant's tracker recorded each trained module's gradient-norm variance across optimizer steps and its activation RMS and, for the 133 single-matmul modules, the channel moments from which the loss's sensitivity to quantization is computed; the allocator spends the byte budget where each byte buys the largest drop in priced damage, subject to per-role floors.

width parameters share
2-bit 660,602,880 39.1%
3-bit 455,344,128 26.9%
4-bit 555,089,920 32.8%
8-bit 19,660,800 1.2%

By module family (one row per projection, summed over the layers that have it; the model has 24):

family modules parameters widths (share of the family's parameters)
moe.w1 24 452,984,832 2b 13%, 3b 33%, 4b 54%
moe.w2 24 452,984,832 2b 67%, 3b 33%
moe.w3 24 452,984,832 2b 67%, 3b 33%
model.embed_tokens 1 155,582,464 4b 100%
conv.in_proj 18 56,623,104 4b 100%
moe.sw1 24 28,311,552 4b 96%, 8b 4%
moe.sw2 24 28,311,552 3b 4%, 4b 79%, 8b 17%
moe.sw3 24 28,311,552 3b 4%, 4b 92%, 8b 4%
conv.out_proj 18 18,874,368 4b 94%, 8b 6%
self_attn.o_proj 6 6,291,456 4b 33%, 8b 67%
self_attn.q_proj 6 6,291,456 4b 33%, 8b 67%
self_attn.k_proj 6 1,572,864 8b 100%
self_attn.v_proj 6 1,572,864 8b 100%

Kept in bf16 and not quantized: 499,456 parameters (0.030% of the model): moe.router (24 tensors, 393,216), conv.conv (18 tensors, 55,296), input_layernorm (24 tensors, 24,576), post_attention_layernorm (24 tensors, 24,576), model.norm (1 tensor, 1,024), self_attn.k_norm (6 tensors, 384), self_attn.q_norm (6 tensors, 384). These are the 24 routers, which pick each token's experts, the short-convolution kernels and the norms, all left in the compute dtype by DynQuant's classification.

72 modules holding 1,358,954,496 parameters (80.4%) were priced by a proxy, not by measurement. DynQuant prices a module from the channel moments its hooks collect, and it collects them only for modules that are a single matmul (the tied embedding is measured through lm_head). Kambo stores each layer's 16 routed experts as three batched tensors (moe.w1, moe.w2, moe.w3) whose forward spans two matmuls and a nonlinearity, so the tracker followed those banks (measure_expert_banks=True; 72 recorded) and recorded their gradient statistics and activation RMS, but no moments. Those banks were priced from the tracker's plasticity score (each bank's within-role rank of log1p of its gradient-norm variance across optimizer steps; DynQuant 0.5.3's default score does not use saliency) times their size times an error curve, scaled to the measured modules' units: the allocator's fallback. The other 133 modules (331,743,232 parameters) were priced from measured moments.

12 modules sit below their role's floor (363,200,512 parameters; at their floors they would take 110,821,376 more bytes, 16.1% of this map). DynQuant's allocator starts every module at its role's floor and goes below floors only when the floors alone cost more than the budget, as they do at 3.25 bits; it then cuts where it prices the damage per byte lowest until the map fits, spends what the last cut overshot on upgrades elsewhere, and records each module left below its floor as a violation rather than refusing the map:

role modules width floor
embedding 1 4-bit 8-bit
moe.expert.gate 3 2-bit 4-bit
moe.expert.gate 8 3-bit 4-bit

The embedding is also the output head (the weights are tied), so its 4-bit table is read twice per token: once to embed and once to score the vocabulary.

Expert banks are grouped along their output axis. A Linear's weight is [out, in] and DynQuant groups along the stored last axis, the input. Kambo stores its expert banks input-major ([experts, in, out]), so the same rule groups each of their 128-weight blocks across 128 output channels of one input. Measured on the base model's layers 0, 12 and 23 at 3 bits, the reconstruction error of the shipped grouping relative to grouping along the input is 0.995 for w1, 1.016 for w2 and 0.999 for w3 (1.000 = no difference, above 1 = the shipped grouping is worse); the banks were not transposed.

Evaluation

All scores come from dynquant eval (dynquant 0.5.3) with the transformers backend, bf16, the chat template, greedy decoding and every decode setting pinned identically across arms:

task items prompt max new tokens scored by
text-to-SQL 2,454 2 solved examples as prior chat turns, then the question 320 execution match: the query runs against the item's database and its result set must equal the reference query's
HumanEval 164 one user turn: complete the function, in a single code block 1024 the item's unit tests, pass@1
MBPP 500 (test split) one user turn: the task and its tests 1024 the item's unit tests, pass@1

Text-to-SQL deals 818 items from each of Gretel's test split, WikiSQL's test split and Spider's 1,034-item dev set (its validation split), in rotation. An item is admitted only if its database holds rows and its reference query returns some, and not a single row of NULLs and zeros, so a wrong query cannot match by also returning nothing; items whose schema and rows exceed 6,000 characters are skipped. Gretel's schemas carry their own INSERTs, WikiSQL's databases are built from its real Wikipedia tables, and Spider's databases, rows included, are inlined from a mirror. MBPP's records say shots: 3, but the chat framing ignores exemplars (DynQuant logs that it does), so every MBPP prompt is the single turn above. Generated code runs in a sandbox (exec/linux/py3.12/rlimits/t=8s/m=4096MB).

Decoding is deterministic. Kambo's experts are summed with a bf16 index_add whose CUDA atomics round in arrival order, and a top-2 router can turn that last bit into a different expert, so two plain runs of one checkpoint disagree on a few items. Every arm was therefore run under torch.use_deterministic_algorithms(True); a repeat of 96 text-to-SQL items reproduced every prediction (EXACT). Launchers recorded across the arms above: dq_det. The base model's first, plain-launcher run is kept as a secondary row on the bf16 card, so the size of the launcher effect is on record.

Greedy is checked, not assumed. The checkpoints were evaluated with the generation defaults inherited from the base model, which sample (below; this repo's own are greedy); the evaluation overrides them, and a check on the fine-tune confirmed that its generations are greedy: 8 of 8 generations were identical under seeds 1 and 2, and 583 of 584 generated tokens are the argmax of a teacher-forced pass over the same text; the one that is not trails it by 0.125 logits, a near-tie inside the check's 0.25-logit tolerance.

Training, data and decontamination are described on the bf16 fine-tune's card.

What is not claimed

  • Storage and resident memory are measured; speed is not. This card makes no claim about decode throughput or latency against bf16.
  • One runtime. Loaded and scored through transformers with DynQuant's quantizer. Not tested with vLLM, llama.cpp, TGI or any other runtime; Kambo's architecture is custom code, so most will not load it at all.
  • bfloat16 only. Loading with dtype=torch.float32 fails at the first layer (expected scalar type Half but found Float): the packed embedding emits its scales' dtype, which the format pins to fp16 under an fp32 model. bf16 loads and runs on GPU and CPU.
  • The allocation is mostly proxy-priced (see above): DynQuant's measured sensitivities cover 19.6% of the quantized parameters.
  • Two skills. Text-to-SQL and Python function writing, plus held-out loss on the training mixture. A quantization that holds these can lose something else.

Usage

pip install dynquant torch transformers accelerate
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

import dynquant

# Before from_pretrained. Without it transformers does not recognise the packed
# format, warns, and returns a model with randomly initialised weights.
dynquant.register_hf_quantizer()

model_id = "VikramPal/kambo-v1-sql-code-DynQuant-3bit"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id, trust_remote_code=True, dtype=torch.bfloat16, device_map="cuda"
)

schema = "CREATE TABLE employees (id INTEGER, name TEXT, dept TEXT, salary INTEGER);"
question = "Which employees in Sales earn more than 50000?"
prompt = (
    "Write a single SQL query that answers the question, using only the tables in the "
    "schema. Return just the query, with no explanation.\n\n"
    f"Schema:\n{schema}\n\nQuestion: {question}"
)
messages = [{"role": "user", "content": prompt}]
inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt", return_dict=True
).to(model.device)
out = model.generate(**inputs, max_new_tokens=320)  # greedy: see generation_config.json
print(tokenizer.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))

This repo's generation_config.json is greedy, which is a change from the base model's. Kambo-v1 ships do_sample: true, temperature 0.7, top_p 0.9 and top_k 2; the top_k is the MoE routing width (top_k: 2 in config.json) carried into the generation defaults, and it restricts every sampled token to the two most likely. Every number on this card was measured greedy, so a plain generate() call here decodes greedily too. To sample, pass do_sample=True with your own temperature, top_p and top_k. transformers 5 still fills the unset top_k from config.json and warns that it "may be ignored"; greedy decoding does ignore it.

Prompt format

The model was trained and evaluated on these wordings, and answers best when asked in them. ChatML, no system message (none is inserted when you supply none, which is how it was trained).

Text-to-SQL
Write a single SQL query that answers the question, using only the tables in the schema. Return just the query, with no explanation.

Schema:
{CREATE TABLE ... statements}

Question: {question}
Python function from a signature and docstring (HumanEval style)
Complete the following Python function. Write the entire function, including the signature, inside a single ```python code block. Do not write tests, examples, or an explanation.

```python
{signature and docstring}```
Python function from a description and tests (MBPP style)
You are an expert Python programmer. Write a Python function for this task:

{description}

Your code must pass these tests:

```python
{assert statements}
```

Return only the function, in a single ```python code block, with no explanation.

License

Released under the Apache License 2.0, as the base model is; see NOTICE. Training data, each under its own license: gretelai/synthetic_text_to_sql (apache-2.0), Salesforce/wikisql (unknown, as the dataset card states it), b-mc2/sql-create-context (cc-by-4.0), nvidia/OpenCodeInstruct (cc-by-4.0). Evaluated on: gretelai/synthetic_text_to_sql (apache-2.0), Salesforce/wikisql (unknown, as the dataset card states it), xlangai/spider (cc-by-sa-4.0), premai-io/spider (no license stated on the dataset card), openai/openai_humaneval (mit), google-research-datasets/mbpp (cc-by-4.0).

Citation

@misc{kambo_v1_2026,
  title  = {Kambo-v1: A 1.7B Hybrid Convolution-Attention Mixture-of-Experts Language Model},
  author = {Kamboj, Vikrampal},
  year   = {2026},
  note   = {Apache-2.0},
  url    = {https://huggingface.co/VikramPal/kambo-v1}
}
Downloads last month
93
Safetensors
Model size
0.2B params
Tensor type
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for VikramPal/kambo-v1-sql-code-DynQuant-3bit

Quantized
(2)
this model

Datasets used to train VikramPal/kambo-v1-sql-code-DynQuant-3bit