Instructions to use litert-community/Kev-4B-LiteRT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT
How to use litert-community/Kev-4B-LiteRT with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Kev-4B for LiteRT
jaredpalmer/kev-4b is a decision model by Jared Palmer. It reads a text, called the state, and typed questions about it: yes/no (noul), multiple choice (choice) and rating (score). For each question it returns an answer with probabilities. It never generates text. A pointer head of two linear layers reads the hidden states of Qwen/Qwen3.5-4B-Base, adapted with a rank-16 LoRA, and turns them into those probabilities. Requests and responses follow the /v1/systemone shape of the author's server (github.com/jaredpalmer/kev).
This repository holds the backbone, with the LoRA folded into the weights, as eight LiteRT files. Six are row graphs, for rows of up to 64, 128, 256, 512, 1,024 and 2,048 tokens; a row is the state plus one question. Two are shared-state pairs, for states of up to 128 and 256 tokens: a pair runs the state once, then each question from it, for questions of up to 64 tokens. The repository also holds the pointer head, the tokenizer files, a Python host, test fixtures and the conversion scripts. The graphs return hidden states. The host builds the input rows, chooses the files and applies the head.
The reference for every check is the author's fp32 PyTorch code. Every file ran on the Apple M4 Max GPU (Metal, float32 precision) with ai-edge-litert 2.2.0, on every test question it can hold; the L1024 file and both pairs also ran on its CPU. A near tie is a question whose two most likely options in the reference are 0.02 or less apart; 9 of the 401 test questions are. Every one of these runs gave the reference's most likely option on every question, near ties included. The largest difference of any option's probability from the reference was 0.0152.
This model was measured on a desktop only. It did not fit a 12 GB Galaxy S26: in two tries, with the earlier release's L1024 file and with this release's L64 file (7.8 GB each), the phone ran out of memory while the graph was compiling. For that phone, use the 0.8B sibling, litert-community/Kev-0.8B-LiteRT, which runs on it.
An invented support ticket with three questions, the state abbreviated:
{
"state": "Ticket #48213, opened by Mara Quellen.\n\nHi, I ordered the Thistlebeam desk lamp (order TB-20931) on September 14 and was charged twice on my card … Could you refund the duplicate charge? I need it sorted before my card statement closes on Friday. …",
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this ticket?",
"criteria": {"billing": "Charges, refunds and invoices", "shipping": "Deliveries, tracking and lost parcels",
"returns": "Exchanges and sending a product back", "technical": "Product faults and setup help"}},
"deadline": {"type": "noul", "instructions": "Does the customer ask for action by a specific deadline?",
"criteria": {"true": "The ticket names a day or date by which something must happen",
"false": "No deadline is stated"}},
"mood": {"type": "score", "instructions": "How upset is the customer?", "criteria": ["Calm", "Annoyed", "Angry"]}
}
}
On the desktop CPU, the host returned this response both through the Ls128 pair, its default route for this request, and through the L256 graph. It is examples/run_example.expected.json, which leaves out latency_ms. That file is unchanged from the earlier release.
{
"model": "kev-latest",
"answers": {
"team": {"type": "choice", "choice": "billing", "confidence": 0.9477,
"probabilities": {"billing": 0.9608, "shipping": 0.0117, "returns": 0.0147, "technical": 0.0128}},
"deadline": {"type": "noul", "noul": 0.8813},
"mood": {"type": "score", "score": 0.3022, "legend": {"0": "Calm", "1": "Annoyed", "2": "Angry"},
"probabilities": {"0": 0.7472, "1": 0.2034, "2": 0.0494}, "confidence": 0.5467}
},
"usage": {"input_tokens": 216, "output_tokens": 190}
}
Other formats:
- ONNX:
midudev/kev-4b-ONNX - GGUF:
ggml-org/Kev-4B-GGUF - MLX:
RoderickQiu/kev-4b-mlx-8bit, and the author's repository has an MLX backend (kev.mlx_model).
We have not checked their contents.
What changed on 2026-10-05
- Files: the earlier release (2026-10-04) had three row graphs, L512, L1024 and L2048. This one has six, from L64 to L2048, and two shared-state pairs, Ls128 and Ls256, for questions of up to 64 tokens. The L512, L1024 and L2048 files keep their names but hold new graphs.
- Kernel: the Gated DeltaNet chunk kernel computes the same values with fewer operators, 11,226 instead of 27,921 in the L512 fp32 graph. In PyTorch fp32, its probabilities stay within 5.3e-6 of the earlier graph's on all 402 test questions, with the same most likely option on each.
- Norms: in Kev-4B, the inputs of two norms are scaled down by a power of two. In fp32 this gives exactly the same values. Under float16 storage, it keeps the final norm's sum of squares below the float16 maximum.
- Speed, on the same Mac in the same timing run, GPU at float32 precision: a 300-token row on L512 takes 396.0 to 396.5 ms instead of 464.4 to 464.7 ms (two runs of each file), and a 1,805-token row on L2048 takes 1,705.4 ms instead of 1,824.2 ms. An 80-token row can now run on L128, in 118.5 ms; the earlier release had nothing below L512 (464.4 to 464.7 ms).
- Host:
host/kev_litert.pyadds the 64-, 128- and 256-token graphs, the pair path and the optionsmode,pair_ratio,handoverandconstant_tensor_sharing. It is the same file as in the 0.8B sibling.
Files
| File | Bytes | Role |
|---|---|---|
kev-4b_rowprefill_L64_fp16fc_i8emb.tflite |
7,785,929,712 | Row graph for rows of up to 64 tokens. |
kev-4b_rowprefill_L128_fp16fc_i8emb.tflite |
7,786,450,960 | Row graph for rows of up to 128 tokens. |
kev-4b_rowprefill_L256_fp16fc_i8emb.tflite |
7,787,432,576 | Row graph for rows of up to 256 tokens. The example uses it with --mode row, or when no pair is present. |
kev-4b_rowprefill_L512_fp16fc_i8emb.tflite |
7,789,789,264 | Row graph for rows of up to 512 tokens. |
kev-4b_rowprefill_L1024_fp16fc_i8emb.tflite |
7,796,075,120 | Row graph for rows of up to 1,024 tokens. |
kev-4b_rowprefill_L2048_fp16fc_i8emb.tflite |
7,814,938,320 | Row graph for rows of up to 2,048 tokens. |
kev-4b_sharedstate_Ls128_Lq64_fp16fc_i8emb.tflite |
7,789,323,680 | Shared-state pair for states of up to 128 tokens and questions of up to 64. The example uses it by default. |
kev-4b_sharedstate_Ls256_Lq64_fp16fc_i8emb.tflite |
7,789,936,224 | Shared-state pair for states of up to 256 tokens and questions of up to 64. |
head/kev_4b_pointer_head.safetensors |
5,245,360 | Pointer head: q.weight [256, 2560], q.bias, k.weight, k.bias, float32. |
head/kev_4b_pointer_head.json |
3,427 | Temperature, delimiter and pad token ids, the head formula and the graph contract. |
tokenizer/tokenizer.json |
19,989,325 | The source repository's tokenizer.json, unchanged. |
tokenizer/tokenizer_config.json |
1,128 | The source repository's tokenizer_config.json, unchanged. |
host/kev_litert.py |
54,212 | Python host: request to rows, routing, graphs, pointer head and response. It also runs from the command line. |
host/requirements-host.txt |
393 | Pinned packages for the Python host. |
examples/ |
run_example.py, the example above, and run_example.expected.json, its expected response. |
|
android/CardSnippet.kt |
7,041 | The Kotlin block below. |
android/measure/ |
Sources of the Android measurement activity behind the phone tries (not a sample app), and of an NPU runner used with Kev-0.8B, with a README. | |
fixtures/ |
Test requests (156 with their text, 221 by reference) and the script that rebuilds the full set; the reference's results for all 402 questions; 12 tokenizer probes; the SemIf license. | |
conversion/ |
Conversion and check scripts, their environment files, and a README with the commands in order. | |
REPRODUCE.md |
16,713 | How to reproduce the files and the checks. |
LICENSE |
11,358 | Apache License 2.0 text. |
NOTICE |
2,003 | Attribution. |
SHA256SUMS |
SHA-256 checksums of the files. |
All eight graph files hold the same weights (7.8 GB each) and differ in their input lengths. A download of any one of them works on its own. Each graph computes all of its positions whatever the row length, so the smallest row graph that holds a row takes the least time. The Python host picks that graph for each question, or a pair for a request's questions. It refuses a row that no file present can hold instead of cutting it.
Each file holds the token embedding as an int8 table, the 32 layers of the backbone and the final RMSNorm, and no LM head. A row graph returns the hidden state at every position of its row. A pair's question step returns the hidden states of the question's 64 positions. The FULLY_CONNECTED operators (296 in the L128 file) read float16 weights through DEQUANTIZE operators (ai-edge-quantizer float casting), the recipe of the earlier release. The embedding table, [248,320 × 2,560], is int8 with one scale per row. Activations, the 24 causal convolutions of the Gated DeltaNet layers and the delta rule stay in float32.
Minimal usage
Python: desktop CPU or GPU
The host needs tokenizers, numpy, safetensors and ai-edge-litert (2.2.0), pinned with their dependencies in host/requirements-host.txt. It does not need PyTorch or the kev package.
hf download litert-community/Kev-4B-LiteRT --exclude "*.tflite" --local-dir Kev-4B-LiteRT
hf download litert-community/Kev-4B-LiteRT kev-4b_rowprefill_L256_fp16fc_i8emb.tflite --local-dir Kev-4B-LiteRT
cd Kev-4B-LiteRT
pip install -r host/requirements-host.txt
python examples/run_example.py --check
These commands download every file except the graphs, plus the L256 graph. examples/run_example.py sends the ticket request above on the CPU and prints the response. --check also compares it with examples/run_example.expected.json. With the Ls128 pair present, the example's default mode sends the three questions through the pair instead, and --mode row keeps them on the L256 graph. On the CPU, the process footprint was 15 GB with one row graph and 31 GB with a pair. To download all eight graph files, leave out --exclude "*.tflite".
In your own code:
import sys
sys.path.insert(0, "host") # run from the repository root
from kev_litert import KevLiteRT
request = {"state": "Order #1182 arrived with a cracked screen.",
"questions": {"refund": {"type": "noul", "instructions": "Should we offer a refund?"}}}
with KevLiteRT.from_dir(".") as kev:
print(kev.decide(request))
from_dir finds the row graphs and pairs that are present, the head and the tokenizer in this repository's layout. It compiles a file only when a request needs it. With the L512 file alone, it serves rows of up to 512 tokens. decide() returns the response: model, answers, usage and latency_ms. The default is the CPU with 4 threads; threads= changes the count. accelerator="gpu" runs the files on the GPU with float32 precision (GpuOptions(enforce_f32=True)). A row that fits no graph raises RowTooLong, and a non-finite hidden state raises NonFiniteOutput.
Several questions on one state, through a pair on the GPU:
request = {"state": "Order #1182 arrived with a cracked screen. The customer asks for a replacement before Friday.",
"questions": {"refund": {"type": "noul", "instructions": "Should we offer a refund?"},
"deadline": {"type": "noul", "instructions": "Does the customer name a deadline?"},
"mood": {"type": "score", "instructions": "How upset is the customer?",
"criteria": ["Calm", "Annoyed", "Angry"]}}}
with KevLiteRT.from_dir(".", accelerator="gpu", mode="pair", constant_tensor_sharing=False) as kev:
print(kev.decide(request)) # the state runs once, then each question runs from it
mode="pair" sends every question through the smallest pair that holds the state. A state or question that does not fit raises PairDoesNotFit. The default, mode="auto", chooses for each request by the rule in the host contract below. It would send this short request's questions to the L64 graph, and the ticket's to the Ls128 pair.
On a Mac with enough memory, constant_tensor_sharing=False is the faster setting for pairs. On the Mac GPU, a 5-question request took 492.6 ms through the Ls128 pair without sharing and 871.8 ms with it, the default. Without sharing, the process footprint was 36.9 GB after compile and 54.9 GB at peak, against 15.9 GB and 17.2 GB with it.
From the command line:
python host/kev_litert.py --graph kev-4b_rowprefill_L512_fp16fc_i8emb.tflite \
--head head/kev_4b_pointer_head.safetensors --tokenizer tokenizer/tokenizer.json --request request.json
--graph takes one or more of the eight files, row graphs and pairs. With only the L512 file, a row longer than 512 tokens is refused. --accel gpu selects the GPU at float32 precision. --mode, --pair-ratio, --handover and --no-constant-tensor-sharing set the matching host options, and --request - reads the request from stdin.
Kotlin: Android GPU with explicit FP32
android/CardSnippet.kt holds the CompiledModel calls for a row graph and for a pair. It is the 0.8B sibling's file, unchanged, and its defaults are for Kev-0.8B. With Kev-4B, pass precision = CompiledModel.GpuOptions.Precision.FP32 to both classes and layers = 32 to KevPairGraph. Kev-4B did not fit the 12 GB phone it was tried on, so for this model the block documents the call shape only: it has not run with Kev-4B.
// SPDX-License-Identifier: Apache-2.0
package com.kev.snippet
import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import com.google.ai.edge.litert.Environment
import com.google.ai.edge.litert.TensorBuffer
import java.io.File
private const val PAD = 248044 // <|endoftext|>, right padding
private const val STATE = 248060
private const val QUESTION = 248061
private const val OPTION_END = 248050
private const val DECIDE = 248062
private fun gpu(precision: CompiledModel.GpuOptions.Precision, constantTensorSharing: Boolean? = null) =
CompiledModel.Options(Accelerator.GPU).apply {
gpuOptions = CompiledModel.GpuOptions(constantTensorSharing = constantTensorSharing, precision = precision)
}
/** The rows the pointer head reads: [decide (the last token), each option's 248050], d floats each. */
private fun readoutRows(hidden: FloatArray, length: Int, last: Int, optionEnds: IntArray): List<FloatArray> {
val d = hidden.size / length
return (listOf(last) + optionEnds.toList()).map { hidden.copyOfRange(it * d, (it + 1) * d) }
}
/**
* One Kev row-prefill graph (`*_rowprefill_L{64,128,256,512,1024,2048}_fp16fc_i8emb.tflite`) on the GPU:
* ids int32 [1, L] + valid float32 [1, L] -> hidden float32 [1, L, d] (d = 1024 for Kev-0.8B, 2560 for Kev-4B).
* precision: FP16_WITH_FP32_ACCUM (float16 storage, float32 accumulation) or FP32. The default precision (plain float16
* activations) moves the probabilities outside the parity tolerance. Kev-4B has not run on a phone: on a 12 GB Galaxy
* S26 both tries, at FP32 and at FP16_WITH_FP32_ACCUM, used up the memory while the graph compiled.
*/
class KevRowGraph(
file: File,
private val length: Int,
env: Environment,
precision: CompiledModel.GpuOptions.Precision = CompiledModel.GpuOptions.Precision.FP16_WITH_FP32_ACCUM,
) : AutoCloseable {
private val model = CompiledModel.create(file.absolutePath, gpu(precision), env)
private val inputs = listOf("ids", "valid").associateWith { model.createInputBuffer(it, "serving_default") }
private val outputs = mapOf("hidden" to model.createOutputBuffer("hidden", "serving_default"))
/**
* row = [248060] + state + [248061] + instructions + for each option ([248049] + option + [248050]) + [248062], with
* the Kev repository's tokenizer.json and no special tokens added; optionEnds = the index of each option's 248050.
*/
fun readout(row: IntArray, optionEnds: IntArray): List<FloatArray> {
require(row.size <= length && row.last() == DECIDE) { "a row ends with 248062 and fits $length tokens" }
require(optionEnds.all { row[it] == OPTION_END }) { "optionEnds must point at 248050 tokens" }
inputs.getValue("ids").writeInt(IntArray(length) { if (it < row.size) row[it] else PAD })
inputs.getValue("valid").writeFloat(FloatArray(length) { if (it < row.size) 1f else 0f })
model.run(inputs, outputs, "serving_default")
return readoutRows(outputs.getValue("hidden").readFloat(), length, row.size - 1, optionEnds)
}
override fun close() {
(inputs.values + outputs.values).forEach { it.close() }
model.close()
}
}
/**
* One Kev shared-state pair (`*_sharedstate_Ls{Ls}_Lq{Lq}_fp16fc_i8emb.tflite`) on the GPU: `state_prefill_<Ls>` runs
* the request's state once, then `question_step_<Ls>_<Lq>` runs each question from it. state_prefill's output buffers go
* to question_step as its inputs (no copy through the app). constantTensorSharing (default true) keeps one copy of the
* weights on the GPU for the two signatures. Without it the pair is faster, but the GPU holds the weights twice: on a
* Galaxy S26 (12 GB) at FP16_WITH_FP32_ACCUM, Kev-0.8B's Ls 128 pair answered a 5-question request in 430.9 ms without
* sharing and 624.8 ms with it (median of the requests timed with the GPU clock not capped), and the phone's smallest
* MemAvailable in the Ls 128 and Ls 256 runs (gate and timing) was 2.7 to 3.1 GB without sharing and 5.6 to 6.1 GB
* with it.
* layers = 24 for Kev-0.8B, 32 for Kev-4B: every fourth layer (3, 7, ...) is an attention layer (state k_<l>, v_<l>),
* the others are Gated DeltaNet layers (gdn_state_<l>, conv_tail_<l>).
*/
class KevPairGraph(
file: File,
private val ls: Int,
private val lq: Int,
env: Environment,
layers: Int = 24,
precision: CompiledModel.GpuOptions.Precision = CompiledModel.GpuOptions.Precision.FP16_WITH_FP32_ACCUM,
constantTensorSharing: Boolean = true,
) : AutoCloseable {
private val model =
CompiledModel.create(file.absolutePath, gpu(precision, if (constantTensorSharing) true else null), env)
private val sigState = "state_prefill_$ls"
private val sigQuestion = "question_step_${ls}_$lq"
private val stateNames =
(0 until layers).flatMap { if (it % 4 == 3) listOf("k_$it", "v_$it") else listOf("gdn_state_$it", "conv_tail_$it") }
private val stateIn = listOf("ids", "valid").associateWith { model.createInputBuffer(it, sigState) }
private val stateOut = stateNames.associateWith { model.createOutputBuffer(it, sigState) }
private val questionIn = listOf("ids", "valid", "state_valid").associateWith { model.createInputBuffer(it, sigQuestion) }
private val questionOut = mapOf("hidden" to model.createOutputBuffer("hidden", sigQuestion))
private val questionInputs: Map<String, TensorBuffer> = questionIn + stateOut
/** state = [248060] + the state's tokens, at most Ls. Run once per request, before its questions. */
fun runState(state: IntArray) {
require(state.size <= ls && state.first() == STATE) { "a state starts with 248060 and fits $ls tokens" }
val valid = FloatArray(ls) { if (it < state.size) 1f else 0f }
stateIn.getValue("ids").writeInt(IntArray(ls) { if (it < state.size) state[it] else PAD })
stateIn.getValue("valid").writeFloat(valid)
model.run(stateIn, stateOut, sigState)
questionIn.getValue("state_valid").writeFloat(valid)
}
/**
* branch = [248061] + instructions + for each option ([248049] + option + [248050]) + [248062], at most Lq tokens: the
* row without the state. optionEnds index the branch. Returns the same rows as KevRowGraph.readout on the whole row,
* up to float rounding.
*/
fun readout(branch: IntArray, optionEnds: IntArray): List<FloatArray> {
require(branch.size <= lq && branch.first() == QUESTION && branch.last() == DECIDE) {
"a question starts with 248061, ends with 248062 and fits $lq tokens"
}
require(optionEnds.all { branch[it] == OPTION_END }) { "optionEnds must point at 248050 tokens" }
questionIn.getValue("ids").writeInt(IntArray(lq) { if (it < branch.size) branch[it] else PAD })
questionIn.getValue("valid").writeFloat(FloatArray(lq) { if (it < branch.size) 1f else 0f })
model.run(questionInputs, questionOut, sigQuestion)
return readoutRows(questionOut.getValue("hidden").readFloat(), lq, branch.size - 1, optionEnds)
}
override fun close() {
(stateIn.values + stateOut.values + questionIn.values + questionOut.values).forEach { it.close() }
model.close()
}
}
Before calling readout, the app builds row and optionEnds: it tokenizes the state, the instructions and each option with this repository's tokenizer.json (no special tokens added, <|name|> rewritten to <¦name¦>) and joins them with the delimiter ids of the host contract below. With a pair, it calls runState once per request, then readout for each question with the question's branch and the option indices counted in it. Afterwards, the app applies the pointer head in float32 to the returned vectors, with the weights in head/kev_4b_pointer_head.safetensors and the temperature in the JSON beside it, and takes the softmax.
Host contract
host/kev_litert.py implements this contract. It ports the request handling of the kev package at tag kev-1.0.
- Render the request. A JSON state becomes
key: valuelines. The options become text:no[: description]andyes[: description]for noul,name[: description]for choice, and the level text for score. - Tokenize each text with
tokenizer/tokenizer.jsonandadd_special_tokens=False. Rewrite<|name|>in caller text to<¦name¦>before tokenizing, so that caller text cannot produce a delimiter. - Build one row per question:
[248060] + state + [248061] + instructions, then[248049] + option + [248050]for each option, then[248062]. The five delimiters are state (248060), question (248061), option start (248049), option end (248050) and decide (248062). The state part of every row is[248060] + state; the rest of the row is the question's branch. - Route the request. With
mode="row", each question takes the smallest row graph whose length L holds its row. Withmode="pair", the state part takes the smallest pair whose length Ls holds it, and each branch takes that pair's question step, which holds 64 tokens. Withmode="auto", the default, the questions whose branches fit the smallest pair that holds the state take that pair if their row graphs would compute more thanpair_ratio(1.5) times as many positions: R > 1.5 × P, where R is the sum of the lengths of those row graphs and P = Ls + 64 × n for n questions. The other questions, and all of them when the condition fails, take their row graphs. A row that fits no graph raisesRowTooLongand is never cut. In modepair, a state or branch that does not fit raisesPairDoesNotFit. - Run a row graph. Pad
idson the right with 248044.validis 1.0 on real tokens and 0.0 on padding. Theserving_defaultsignature takesidsint32[1, L]andvalidfloat32[1, L], and returnshiddenfloat32[1, L, 2560], the hidden states after the final RMSNorm at every position. Positions are the constants 0 to L−1. The graph has no input or output for model state, so each row starts from zero. - Run a pair.
state_prefill_<Ls>takes the state part'sidsandvalid[1, Ls]and returns 64 state tensors: the recurrent state and conv tail of each of the 24 Gated DeltaNet layers, and the keys and values of each of the 8 attention layers. They hold 61,079,552 bytes per request for Ls128 and 69,468,160 for Ls256.question_step_<Ls>_64takes a branch'sidsandvalid[1, 64], the state'svalidasstate_valid[1, Ls]and the state tensors, and returnshiddenfloat32[1, 64, 2560]; its positions continue after the state.handover="direct", the default, passes the state step's output buffers to the question step as its inputs;"host"reads them back and writes them into the question step's input buffers. On the GPU,constant_tensor_sharing=True, the default, keeps one copy of the weights for the two signatures, andFalsekeeps a copy per signature. With and without sharing, the probabilities and read-out hidden states are bit-identical (314 questions on Ls128, 341 on Ls256). - Read out in float32. The decide vector is the hidden state of the row's last real token (248062). Option k's vector is the hidden state of its closing 248050. On the pair path these positions are counted in the branch. Then
z_k = ((h_opt_k · Wkᵀ + bk) · (h_decide · Wqᵀ + bq)) / 16 / T, with T = 2.406050072164233, andp = softmax(z). - Answer. Noul returns
noul= p(yes). Choice returnschoice(the most likely option),probabilitiesandconfidence= (p_max − 1/K) / (1 − 1/K) for K options. Score returnsscore= Σ i·p_i over the levels counted from 0,legend,probabilitiesandconfidence= max(0, 1 − E|level − mode| / D), where D is the mean absolute deviation of a uniform distribution over the levels. Numbers are rounded to 4 decimals, as the author'sround_probdoes.
Use this repository's tokenizer.json, which is the source repository's file, and read it with tokenizers.Tokenizer.from_file. It stores the tokenizer pipeline that transformers builds for the author's code (Qwen2Tokenizer). The base repository's own tokenizer.json, read the same way, gives different ids for text with combining marks, such as Devanagari, and for a few special-looking strings such as <think>. It differs on 4 of the 12 probes in fixtures/tokenizer_probes.json. The host refuses it.
On all 402 test questions, the host's token ids equal those of the author's encode. Run on the published files on Metal at float32 precision, the host's results equal those of the agreement runs below bit for bit: 452 of 452 questions through the row graphs, and 314 (Ls128) and 341 (Ls256) through the pairs with sharing. On 20 rows of the L128 file, the host on the CPU stays within 3.7e-6 of the Metal results.
On the GPU, set float32 precision explicitly: GpuOptions(enforce_f32=True) in Python, CompiledModel.GpuOptions(precision = FP32) in Kotlin. The default precision moves the probabilities outside the tolerance (see GPU precision below).
Measured agreement and speed
Agreement with the author's fp32 code
The reference is the author's code at tag kev-1.0 on the CPU in float32, with the author's lock file (torch 2.8.0, transformers 5.17.0, peft 0.21.0): Checkpoint.load(cpu, dtype=float32), then forward in the row form. Probabilities are compared after the temperature.
The test set has 377 requests with 402 questions. The author's evals/v4/transfer-v4/development.jsonl gives 220 of them: records 1 to 60 (MMLU, 4 options), 20 from each of its 7 other sources and 20 score questions. SemIf authored144 gives 144 items, mapped to 3-option choice questions. The 12 invented records, written for this conversion, hold 37 questions of all three types. One more request is a control, described below. Without the control there are 401 questions: 280 choice, 93 noul and 28 score. The rows of 392 questions have at most 366 tokens. The other 9 questions sit on 3 long invented states, with rows of 1,369 to 1,805 tokens. Only the L2048 file holds them.
"Same most likely option" counts the questions that are not near ties. Max |Δp| is the largest absolute difference of any option's probability, and mean |Δp| is the mean over all options. The tolerance set before the runs was: the same most likely option on every question that is not a near tie, max |Δp| of at most 0.02 and mean |Δp| of at most 0.002. The control is the request of one test question with one word of its instructions changed ("correctly" to "incorrectly"). Its column gives its max |Δp| from that question's reference; a run detects the change when this is above 0.02.
| File | Runtime | Questions | Same most likely option | Near ties | Max |Δp| | Mean |Δp| | Control |
|---|---|---|---|---|---|---|---|
| L64 | Mac GPU (Metal), float32 precision | 72 | 71/71 | 1/1 | 0.0070 | 3.6e-4 | not run |
| L128 | Mac GPU (Metal), float32 precision | 321 | 312/312 | 9/9 | 0.0152 | 4.4e-4 | 0.0528 |
| L256 | Mac GPU (Metal), float32 precision | 385 | 376/376 | 9/9 | 0.0152 | 4.6e-4 | 0.0528 |
| L512 | Mac GPU (Metal), float32 precision | 392 | 383/383 | 9/9 | 0.0152 | 4.5e-4 | 0.0528 |
| L1024 | Mac GPU (Metal), float32 precision | 392 | 383/383 | 9/9 | 0.0152 | 4.5e-4 | 0.0528 |
| L2048 | Mac GPU (Metal), float32 precision | 401 | 392/392 (long rows 9/9, max 0.0010) | 9/9 | 0.0152 | 4.5e-4 | 0.0528 |
| L1024 | Mac CPU, 8 threads | 392 | 383/383 | 9/9 | 0.0152 | 4.5e-4 | 0.0528 |
| Pair Ls128 | Mac CPU; Mac GPU (Metal), float32 precision, with and without sharing | 313 | 308/308 | 5/5 | 0.0152 | 4.1e-4 | 0.0528 |
| Pair Ls256 | Mac CPU; Mac GPU (Metal), float32 precision, with and without sharing | 340 | 335/335 | 5/5 | 0.0152 | 4.4e-4 | 0.0528 |
No question changed its answer on any of these runs, near ties included. The numbers equal those of the earlier release's files (max 0.0152). On the same questions, the GPU results of the two releases differ by at most 8.8e-6. Every run that holds the control's row detected the change. That row has 94 tokens, so the L64 file does not hold it. On Metal, is_fully_accelerated is true for all eight files.
Desktop (Apple M4 Max)
These times come from an Apple M4 Max (128 GB, macOS 27.0) with ai-edge-litert 2.2.0 through the Python CompiledModel API, measured while no other GPU job ran. A time is the wall clock of writing the inputs, running and reading the output back: the median of 20 calls after 5 warm-up calls. Here a call runs one question's row.
| Row | File | GPU (Metal), float32 precision | CPU, 8 threads |
|---|---|---|---|
| 64 tokens | L64 | 70.8 ms | 228.7 ms |
| 80 tokens | L128 | 118.5 ms | 450.0 ms |
| 128 to 142 tokens (one question of a 5-question request) | L256 | 209.8 ms | 815.3 ms |
| The 5-question request (sum of 5 calls, all on L256) | L256 | 1,048.8 ms | 4,083.5 ms |
| 300 tokens | L512 | 396.0 to 396.5 ms (two runs) | 1,497.0 ms |
| 1,000 tokens | L1024 | 801.0 ms | 2,928.4 ms |
| 1,805 tokens | L2048 | 1,705.4 ms | 5,934.9 ms |
The GPU compile took 29 to 49 s. With the GPU at float32 precision and one row graph loaded, the process footprint was 17.0 to 20.8 GB after compile and peaked at 37.0 to 38.1 GB. With the CPU, it was 14.7 to 16.4 GB after compile.
Requests: row graphs and pairs
A request's time is the wall clock of all its calls, timed as above. Through the row graphs, each question runs on the smallest graph that holds its row. Through a pair, the state runs once and each question then runs from it, with the state handed over directly. GPU means Metal at float32 precision; CPU means 8 threads.
| Request | Row graphs, GPU | Pair, GPU, with sharing | Pair, GPU, without sharing | Row graphs, CPU | Pair, CPU |
|---|---|---|---|---|---|
| 1 question; state 30 tokens; row 94 tokens (L128); Ls128 pair | 117.4 ms | 331.2 ms | 201.1 ms | 444.8 ms | 755.5 ms |
| 2 questions; state 109 tokens; rows on L256; Ls128 pair | 419.5 ms | 465.9 ms | 274.0 ms | 1,629.2 ms | 1,019.0 ms |
| 3 questions; state 109 tokens; rows on L256; Ls128 pair | 629.2 ms | 600.6 ms | 347.2 ms | 2,444.0 ms | 1,287.1 ms |
| 5 questions; state 99 tokens; 3 rows on L256, 2 on L128; Ls128 pair | 864.1 ms | 871.8 ms | 492.6 ms | 3,333.0 ms | 1,814.2 ms |
| 2 questions; state 150 tokens; Ls256 pair | 419.5 ms | 571.4 ms | 376.6 ms | 1,624.4 ms | 1,410.8 ms |
| 3 questions; state 167 tokens; Ls256 pair | 628.8 ms | 705.6 ms | 449.9 ms | 2,438.1 ms | 1,677.0 ms |
On the GPU, a pair request splits into one state step and one question step per question:
| Pair, GPU | Without sharing | With sharing |
|---|---|---|
| Ls128 | 128.2 ms + 72.9 ms per question | 196.8 ms + 134.4 ms per question |
| Ls256 | 228.5 ms + 73.9 ms per question | 298.7 ms + 136.0 ms per question |
With a pair on the GPU, the process footprint was 15.9 GB after compile and 17.2 GB at peak with sharing, and 36.9 GB and 54.9 GB without, for either pair. The compile took 13 to 22 s with sharing and 62 to 74 s without. On the CPU, the footprint was 31 GB with a pair and 15 to 16 GB with one row graph.
With sharing, the default, the pairs on the GPU took about as long as the row graphs, or longer. Without sharing, they took less time than the row graphs on every request of two or more questions. On the CPU too, a pair took less time on every request of two or more questions, and more on the 1-question request.
This release and the earlier one
The earlier release's files were timed again in the same timing run as this release's, on the same Mac, with the GPU at float32 precision:
| Row | Earlier release (2026-10-04) | This release |
|---|---|---|
| 300 tokens | L512: GPU 464.4 to 464.7 ms (two runs), CPU 1,967.5 ms | L512: GPU 396.0 to 396.5 ms (two runs), CPU 1,497.0 ms |
| 1,805 tokens | L2048: GPU 1,824.2 ms | L2048: GPU 1,705.4 ms |
| 80 tokens | No graph below L512, which took 464.4 to 464.7 ms on the GPU | L128: GPU 118.5 ms |
The earlier release's card gave 2,320.3 ms on the GPU for the 5-question request (state 99 tokens, five calls on L512), measured in an earlier run. In this release, that request took 864.1 ms through the row graphs, 492.6 ms through the Ls128 pair without sharing and 871.8 ms with sharing.
Galaxy S26
This model did not fit one Galaxy S26 (12 GB). Both tries compiled one graph for the GPU. Both ran out of memory during the compile, before any call ran. The CPU was not tried.
- 2026-10-03, the L1024 file of the earlier release (7.8 GB), FP32 precision: the delegate took all 34,313 nodes, and 9 s later the low-memory killer stopped the process. By then the process had peaked at 5.7 GB of resident memory (VmHWM), and the killer's log line reported 7.2 GB of its memory in swap.
- 2026-10-05, the L64 file of this release (7,785,929,712 bytes), FP16_WITH_FP32_ACCUM precision, with the phone's memory watched: the delegate took all 5,307 nodes in one partition. 6 s later, MemAvailable had fallen from 6.8 GB to 1.2 GB, the process had peaked at 5.9 GB (VmHWM), and the GPU's allocation was still 0.3 GB. The watch then stopped the app. By that time the low-memory killer had started reclaiming other apps.
Not measured: phones with more than 12 GB, the NPU (on Kev-0.8B, compiling a 1.26 GB file on the phone used 5.5 to 6.2 GB of VmHWM), and GPUs on Linux or Windows.
GPU precision
Run the GPU at float32 precision. In the Python host this is precision="fp32", the default for accelerator="gpu" (GpuOptions(enforce_f32=True)); in Kotlin it is Precision.FP32. On the Mac GPU, the default precision (float16 activations; precision="fp16" in the host) gave finite hidden states on every row but moved the probabilities outside the tolerance:
| File | Runtime | Questions | Same most likely option | Near ties | Max |Δp| | Mean |Δp| | Control |
|---|---|---|---|---|---|---|---|
| L128 | Mac GPU (Metal), default precision (float16 activations) | 321 | 312/312 | 8/9 | 0.0456 | 3.8e-3 | 0.0561 |
| L512 | Mac GPU (Metal), default precision (float16 activations) | 392 | 383/383 | 6/9 | 0.0675 | 4.0e-3 | 0.0488 |
At float32 precision, the same two files stay within 0.0152 and keep all 9 near ties. The Kotlin and C APIs also offer FP16_WITH_FP32_ACCUM (float16 storage, float32 accumulation). No Kev-4B output has been checked at that precision: the phone try that used it ran out of memory before any call. The Python host raises NonFiniteOutput instead of answering from a non-finite row.
Limits
- It does not generate text. A row holds at most 2,048 tokens: the state plus one question. A pair holds a state of up to 256 tokens and questions of up to 64. Longer states are not handled.
- The row graphs compute the state again for every question, so a request with Q questions takes Q calls. A pair computes the state once per request.
- The model needs desktop-class memory. On the Mac GPU, the process footprint with one row graph was 17 to 21 GB after compile and peaked at 38 GB; with a pair without sharing, it was 37 GB and peaked at 55 GB. On the CPU it was 15 to 16 GB, and 31 GB with a pair. It did not fit the 12 GB Galaxy S26 (two tries), and no other device was measured.
- Run the GPU at float32 precision. The default precision moved the probabilities outside the tolerance.
- Probabilities differ from the reference by up to 0.0152. For this model, the source of that difference was not measured separately; for the 0.8B sibling, it was the int8 embedding table. No test question changed its answer on the CPU or on the GPU at float32 precision, near ties included.
- The agreement numbers measure how closely the conversion follows the author's fp32 code, not task accuracy. This conversion did not measure accuracy or calibration again; see the source model card. The temperature is the author's fitted value, unchanged.
- Input languages and intended uses follow the source model card.
- Measured on one Mac (Apple M4 Max). Converting one file (the fp32 export) used about 50 GB of memory.
Provenance, conversion and license
- Source:
jaredpalmer/kev-4bat tagv1.0(commit6cfce5c2fa4b4bd64026336ab649c5ca78857d52) by Jared Palmer, Apache-2.0. Base:Qwen/Qwen3.5-4B-Baseat revision1001bb4d826a52d1f399e183466143f4da7b741b, Apache-2.0. - Model: the base with a rank-16 LoRA on 12 projection types and a pointer head of two linear layers with 256 dimensions. The base has 32 layers (24 Gated DeltaNet layers, a form of linear attention, and 8 full-attention layers), hidden size 2560 and a vocabulary of 248,320.
- Weights: the author's
scripts/merge_lora_checkpoint.pyat tagkev-1.0folded the LoRA into 248 weights in float32. Read with the author's loader, the folded checkpoint gives bit-identical probabilities to the adapter checkpoint on all 402 questions. - Graph: the Qwen3.5 text model of transformers 5.14.1, re-authored for export. The Gated DeltaNet chunk kernel is rewritten at rank 4 or less: the tail padding becomes a concat, and the diagonal and triangular masks become constants. The Gated DeltaNet layers of Kev-4B have 16 key heads and 32 value heads; the query and key heads are copied to match by concat, at rank 4 instead of through a rank-5 tensor. Attention copies the GQA keys and values by concat. In the earlier release, the hidden states before and after this rewrite differed by 2.6e-4.
- Kernel of this release: it changes how the chunk kernel computes, not what it computes. It forms the inverse inside each chunk by recursive doubling instead of step-by-step substitution. Softplus, the exponentials, the decays inside a chunk and the gated norm's input are written so that float16 storage does not overflow or lose precision in them. In Kev-4B, each layer's input norm and the final RMSNorm also take their input scaled down by a power of two, with eps scaled to match.
conversion/README.mddescribes each change. - Exported with litert-torch 0.9.4 and quantized with ai-edge-quantizer 0.9.0. The fp32 export used about 50 GB of memory per file.
- Operators: 5,307 in L64, 6,436 in L128, 8,116 in L256, 11,476 in L512, 18,196 in L1024 and 31,636 in L2048. Each pair has 6,777 (Ls128) or 8,457 (Ls256) in
state_prefilland 5,406 inquestion_step. No file has a CUSTOM operator, an int64 tensor or a tensor of rank above 4. - Scripts:
conversion/holds every conversion and check script, with its environment files and a README of the commands. It is the same folder as in the 0.8B sibling, and the Kev-4B commands take--model 4b. Rebuilt from these scripts alone, the L128 file came out with the same SHA-256 as the published one.REPRODUCE.mdlists the sources, the environments, the steps and where each number comes from. - Test data: the requests are the same as in the 0.8B sibling, and
fixtures/oracle_probs.jsonholds this model's reference.fixtures/holds the SemIf records (MIT, from github.com/TheoLeeCJ/SemIf at commitca3ba65f) and the 12 invented records with their text. The 221 transfer-v4 records, the control included, are listed by reference only: file, tag, line, line SHA-256 and_meta.id. Their source datasets carry different licenses: on the Hub cards, tweet_eval is unknown and SciQ is CC BY-NC 3.0.fixtures/rebuild_requests.pyrestores them from the author's GitHub repository at tagkev-1.0. - Training data and evaluation: see the source model card.
License: Kev-4B and Qwen3.5-4B-Base are licensed under Apache 2.0. This repository is released under the same license (LICENSE), except the SemIf fixture records, which are MIT (fixtures/LICENSE-SemIf-MIT.txt). Attribution, including the code adapted from the kev package and from litert-torch, is in NOTICE.
- Downloads last month
- 61