Instructions to use litert-community/d1-3B-LiteRT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT
How to use litert-community/d1-3B-LiteRT with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
d1-3B for LiteRT
This repository is a community conversion of LiquidAI/d1-3B to LiteRT, made with litert-torch 0.9.4 for the CompiledModel API of ai-edge-litert 2.2.0: on an Apple M4 Max with Metal at float32 precision, the source model card's three-question example takes a median 153.9 ms and its photo with one question 389.3 ms (2026-10-09). A Mac sample app runs it on your own text, photo and questions: LiteRT-Models/d1-3b.
LiquidAI/d1-3B is a decision model by Liquid AI, post-trained from LiquidAI/LFM2.5-VL-3B. It reads a state (text, a JSON value, pictures, or a mix) and named questions about it: yes/no (noul), a pick from named options (choice) and an ordered rating (score, 2 to 10 levels). For each question it returns the probability of every option, read from the logits of a few option tokens at the prompt's last position. It never generates text. Requests and responses follow the provider's system_one(): a state and questions in, typed answers and a token count out. What it is for and how it scores: the source model card and the blog post.
This repository holds the model as LiteRT files for desktop runtimes. Six row graphs take rows of up to 128, 256, 512, 1,024, 2,048 and 4,096 tokens. Three shared-state pairs run a state of up to 64, 128 or 256 tokens once, then each question from it. Two more graphs, the picture tower and the projector, turn pictures into tokens. The text graphs take the embeddings of a row's tokens and return the hidden states after the final norm. A Python host builds the rows, writes the embeddings, picks the graphs and reads the answers out. The repository also holds the tokenizer, the tables the host reads, test fixtures and the conversion scripts.
The reference for every check is the provider's own code in float32 on the CPU. On an Apple M4 Max, on the CPU and on the Metal GPU at float32 precision, every file gave the reference's most likely option on every test question it holds, the one near tie included. Its probabilities differed from the reference by at most 1.02e-5 on text, and by at most 3.39e-5 on the three test requests with pictures.
The files are for desktops and laptops. A 12 GB Galaxy S26 did not hold the model: the GPU and CPU delegates accepted an earlier 5.14 GB form of it, but during its compile the phone ran short of memory and a memory guard stopped the app (see Android).
The text example of the source model card, one state and three questions:
{"state": "I was charged twice this month, please refund one of them.",
"questions": {
"refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Charges, refunds, invoices", "technical": "App or site faults",
"fraud": "Suspected unauthorised use"}},
"urgency": {"type": "score", "instructions": "How urgent is this?",
"criteria": ["Can wait", "Today", "Blocking the customer now"]}}}
The host sent it through the 64-token pair: the state once, then one step per question. The answers' probabilities, rounded to 5 decimals:
| Question | Option | Provider, float32 CPU | This repository, CPU | This repository, Metal (float32) |
|---|---|---|---|---|
refund (noul) |
yes | 0.98962 | 0.98962 | 0.98962 |
team (choice) |
billing | 0.98351 | 0.98351 | 0.98351 |
| technical | 0.01065 | 0.01065 | 0.01065 | |
| fraud | 0.00585 | 0.00585 | 0.00585 | |
urgency (score) |
Can wait | 0.33858 | 0.33858 | 0.33858 |
| Today | 0.56744 | 0.56744 | 0.56744 | |
| Blocking the customer now | 0.09398 | 0.09398 | 0.09398 |
The response reports 116 input tokens and no output tokens. Before rounding, both runs differ from the reference by at most 1.3e-6 on these probabilities.
Other formats:
- GGUF:
LiquidAI/d1-3B-GGUF, by Liquid AI - w8a8:
LiquidAI/d1-3B-w8a8, by Liquid AI - ONNX:
onnx-community/d1-3B-ONNX - MLX:
DJLougen/d1-3B-MLX-8bit
We have not checked their contents.
Files
| File | Bytes | Role |
|---|---|---|
d1-3b_rowprefill_embeds_L128_fp16fc.tflite |
4,871,345,568 | Row graph for rows of up to 128 tokens. |
d1-3b_rowprefill_embeds_L256_fp16fc.tflite |
4,871,607,712 | Row graph for rows of up to 256 tokens. |
d1-3b_rowprefill_embeds_L512_fp16fc.tflite |
4,872,525,216 | Row graph for rows of up to 512 tokens. |
d1-3b_rowprefill_embeds_L1024_fp16fc.tflite |
4,875,933,088 | Row graph for rows of up to 1,024 tokens. |
d1-3b_rowprefill_embeds_L2048_fp16fc.tflite |
4,889,040,192 | Row graph for rows of up to 2,048 tokens. |
d1-3b_rowprefill_embeds_L4096_fp16fc.tflite |
4,940,420,512 | Row graph for rows of up to 4,096 tokens. |
d1-3b_sharedstate_embeds_Ls64_Lq64_fp16fc.tflite |
4,871,972,336 | Shared-state pair: a state of up to 64 tokens, then questions of up to 64 tokens each. |
d1-3b_sharedstate_embeds_Ls128_Lq64_fp16fc.tflite |
4,872,108,304 | Shared-state pair: a state of up to 128 tokens, then questions of up to 64 tokens each. |
d1-3b_sharedstate_embeds_Ls256_Lq128_fp16fc.tflite |
4,872,620,816 | Shared-state pair: a state of up to 256 tokens, then questions of up to 128 tokens each. |
d1-3b_vision_tower_fp16fc.tflite |
825,613,520 | Picture tower, one call per tile. |
d1-3b_projector_fp16fc.tflite |
27,281,424 | Picture projector, one call per tile. |
tables/embed_table.safetensors |
524,288,104 | The tied token embedding in bfloat16, 128,000 × 2,048. The host writes its rows into the text graphs. |
tables/readout_table.safetensors |
10,118,944 | Float32 rows of the same table for the 1,234 token ids that answers are read from. |
tables/vision_position_table.safetensors |
1,179,736 | The tower's position table in float32, 16 × 16 × 1,152. |
tokenizer/tokenizer.json |
17,905,750 | The source repository's tokenizer.json, unchanged. |
contract.json |
The host contract as data: token ids, graph inputs and outputs, buckets, pairs, the routing table, the picture path, and each file's checks. | |
host/ |
Python host: d1_litert.py (rows, row graphs, read-out, answers), d1_shared_state.py (pairs and routing), d1_vision.py (pictures), and requirements-host.txt. |
|
examples/ |
run_example.py, the source card's examples, and run_example.expected.json, the provider's answers to them. |
|
fixtures/ |
Test requests: 161 records with their text, 221 by reference to the kev repository with a script that restores them, the four control requests, the provider's reference probabilities and tokenizer probes. | |
conversion/ |
Conversion, check and timing scripts, their environments, summaries of their records, and a README with the commands in order. | |
REPRODUCE.md |
Sources, environments, steps and where each number comes from. | |
LICENSE, NOTICE, SHA256SUMS |
The LFM Open License v1.0 (the source repository's file), attribution, and the SHA-256 of every file. |
Each text graph holds the whole text decoder, from the embeddings to the final norm: 4.87 to 4.94 GB per file, 45.3 GB for the repository. A graph computes every position of its length whatever the row's length, so the host takes the smallest graph that holds a row. load_host() (below) compiles only the files a request needs, and works with the files that are present.
Minimal usage
A desktop app with an editor, built on this host: LiteRT-Models/d1-3b (Mac, Python).
Python: desktop CPU or GPU
hf download litert-community/d1-3B-LiteRT --local-dir d1-3B-LiteRT
cd d1-3B-LiteRT
pip install -r host/requirements-host.txt
python examples/run_example.py --check # CPU (XNNPACK, 8 threads)
python examples/run_example.py --check --accel gpu # GPU at float32 precision (Metal on a Mac)
The host needs tokenizers, numpy, safetensors, pillow and ai-edge-litert (2.2.0), pinned with their dependencies in host/requirements-host.txt. It does not need PyTorch or transformers. It was tested with Python 3.14.6 on macOS 27.0.
examples/run_example.py sends the source card's two examples: the text request above, and its photo with the question "How many cats are there?". Then it sends a state too long for any graph, which the host refuses. The photo (COCO val2017 000000039769.jpg, 640 × 480) is not in this repository: the script downloads it into memory and checks its SHA-256 before use. --check compares the answers with examples/run_example.expected.json, which holds the provider's float32 CPU answers. Every probability must be within 1e-5, and the most likely options, the input token counts, the routes and the refusal must be equal. On the Apple M4 Max the whole run took 14.1 s on the CPU and 61.5 s on Metal, compiles included, with a peak memory footprint of 23.1 GB and 33.5 GB.
In your own code:
import sys
sys.path[:0] = ["host", "examples"] # run from the repository root
from run_example import load_host
request = {
"state": "The parcel came two days late and the box was crushed, but the lamp inside works.",
"questions": {
"late": {"type": "noul", "instructions": "Did the parcel arrive late?"},
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"shipping": "Deliveries and couriers", "returns": "Replacements and refunds"}},
},
}
with load_host(".", accelerator="cpu") as host: # accelerator="gpu": the GPU at float32 precision
print(host.route(request))
print(host.decide(request))
load_host() builds the host from contract.json with the files that are present. It compiles a graph file only when a request needs it, and keeps at most one text graph compiled (keep=): the least recently used one is closed when another is needed. route() tells where a request's questions run: here, the 64-token pair. decide() returns the answers and usage. The default is the CPU with 8 threads (threads=); accelerator="gpu" sets GpuOptions(enforce_f32=True). A request with pictures carries them as "images": PIL images, file paths or the bytes of image files. A row that no graph holds raises RowTooLong.
Host contract
host/d1_litert.py implements this contract with the provider's defaults (its prompt.py, runner.py and api.py at the revision below). contract.json holds the same rules as data.
- Render one row per question:
<|startoftext|><|im_start|>user\n, the state block,\nQUESTION:\n, the question block, then<|im_end|>\n<|im_start|>assistant\n. A string state goes in as it is, any other JSON value asjson.dumps(value, ensure_ascii=False, indent=2), followed by two newlines. A null state leaves out the state block and theQUESTION:line. - The question block.
choice: the instructions,Options:, one line<code> <description>per option, andReply with the option code only..noul: the instructions (withYes: …andNo: …lines when criteria are given) andReply with yes or no only..score: the instructions, one line<i> <level>per level, counted from zero, andReply with a single digit 0-<K-1> only.. Option codes are the names themselves when every name is one letter, elseA,B,C, … up to 26 options, else00,01, …. A code that is not one token, or whose token another option already took, is replaced by the next free single-token alias. - Tokenize with
tokenizer/tokenizer.json(tokenizers.Tokenizer.from_file) and no special tokens: the BOS is part of the text. The special ids are<|startoftext|>124894,<|im_start|>124899,<|im_end|>124900 and<|pad|>124893. - Take the smallest row graph that holds the row, and pad the row on the right with
<|pad|>. A longer row is refused (RowTooLong), never cut. The answer slot is the row's last real token. - Write
embedsfloat32[1, L, 2048]: the rows oftables/embed_table.safetensorsat the row's ids, widened from bfloat16 to float32 (exact), with the pad's row on the padding.validfloat32[1, L]is one on real tokens and zero on padding. - Run the row graph's
serving_defaultsignature:embedsandvalidin,hiddenfloat32[1, L, 2048]out, the hidden states after the final RMSNorm at every position. The positions are constants in the graph, from zero, and no state carries over between calls. - Read out: for each id of an option's group, logit = h · E[id] in float32, with h the hidden state at the answer slot and E the rows of
tables/readout_table.safetensors. An option scores its group's highest logit, and p is the softmax over the options, in float64. A choice option's group is the token of its code, plus the token of the code with a leading space when that is one token. A noul question reads the single-token forms of yes / Yes / YES against those of no / No / NO; a score level i reads the digit i. This equals the provider's read-out, a log-softmax over the whole vocabulary: it shifts every logit by the same amount, which neither the group maximum nor the softmax sees. - Answer:
noulgives P(yes) asnoul;choicegiveschoice(the most likely name),confidence(its probability) andprobabilities;scoregivesscore(the expected level),confidence,probabilitiesandlegend.usagecountsinput_tokensthe provider's way (the row for one question; the state's tokens once plus each question's own tokens for several) andoutput_tokens: 0.
Pictures
A request with images renders one <image> per picture at the head of the user turn. The host shrinks a picture larger than 1024 × 1024 pixels, then cuts a large picture into 512 × 512 tiles (2 to 10, in the grid closest to its aspect ratio) plus a thumbnail; a small picture is one tile. A tile is up to 1,024 patches of 16 × 16 pixels, 768 values each. Its tokens in the row are <|image_start|> (125009), for each tile of a split picture <|img_row_r_col_c|> (124908 + 10(r − 1) + (c − 1)) and 256 <image> tokens (124907), then <|img_thumbnail|> (125008) and the thumbnail's tokens, and <|image_end|> (125010). A one-tile picture has its tokens alone between the start and the end.
The tower graph takes pixels [1, 1024, 768], pos [1, 1024, 1152] (the position table resized to the tile's patch grid by the host) and mask [1, 1024], and returns features [1, 1024, 1152]. The host folds each 2 × 2 block of the tile's patches into one row of 4,608 values. The projector maps soft [1, 256, 4608] to mm [1, 256, 2048], whose leading rows are the tile's tokens. These take the row's <image> positions in order, and the row runs on a row graph like text. host/d1_vision.py ports each step of the provider's processor in transformers 5.14.1, PyTorch's antialiased resizes included; on 18 synthetic pictures it gives the processor's pixel values bit for bit.
Shared-state pairs
A pair is one file with two signatures over one copy of the weights:
state_prefill_<Ls>takesembeds[1, Ls, 2048]andvalid[1, Ls]for the state's tokens, the row up toQUESTION:\n. It returns 38 state tensors:conv_tail_<l>[1, 2048, 2]for the 22 short-convolution layers, the convolution's input at the last two state tokens, andk_<l>,v_<l>[1, 8, Ls, 64]for the 8 attention layers.question_step_<Ls>_<Lq>takes one question's own tokens asembeds[1, Lq, 2048]andvalid[1, Lq], the state'svalidasstate_valid[1, Ls], and the 38 tensors. It returnshidden[1, Lq, 2048]. The question's positions continue after the state inside the graph, and its answer slot is its last token.
The host passes state_prefill's output buffers to question_step as its inputs (handover="direct"), or reads them back and writes them (handover="host"). Both give the same bits on the Mac CPU and Metal. In PyTorch, before export, the pairs gave the row form's probabilities within 3.8e-6 on 906 runs that cover all 415 text questions.
Routing
A request with pictures runs row by row. For any other request, the host compares the expected times of the forms that hold it, from the call times measured on the Mac's Metal GPU at float32 and stored in contract.json. Row graphs cost the sum of their rows' calls: 61.84, 108.71, 207.54, 420.37, 889.29 and 2,049.99 ms for L128 to L4096. A pair costs one state call plus one call per question: 37.80 + 38.72 ms (Ls64), 60.44 + 38.91 ms (Ls128) and 106.04 + 62.93 ms (Ls256). The smaller expected time wins, and a tie goes to the rows. The three-question example takes the Ls64 pair: 153.96 ms expected, against 185.52 ms for three L128 rows. On the 16 fixture requests with several questions, the pair route and the row route answered alike: the same most likely options, probabilities within 2.1e-6, and the same refusal of the one request with a question that has no instructions.
Measured agreement and speed
Agreement with the provider's float32 code
The reference is the provider's code at the revision below, in float32 on the CPU, one row per question. The test set has 382 requests with 417 questions: 415 text questions, the photo question of the source card, and one question that the provider's code refuses because it has no instructions. They are 221 requests from the transfer-v4 development file of github.com/jaredpalmer/kev, 144 SemIf items, the source card's two examples and 15 requests written for these conversions. One question is a near tie: its two most likely options in the reference are 0.02 or less apart. Four control requests each change the state or the instructions of a test request. Each must move some option's probability by more than 0.02, which shows that the checks see the input.
The tolerance, set before the runs: the same most likely option on every question that is not a near tie, max |Δp| of at most 0.02 over all options, and mean |Δp| of at most 0.002. Each file runs the questions whose rows it holds. Every text file passes on the CPU (XNNPACK, 8 threads) and on Metal at float32 precision (GpuOptions(enforce_f32=True)), with every operator on the delegate in one partition.
| File | Questions | Same most likely option (CPU and Metal) | Near tie kept | Max |Δp|, CPU / Metal | Mean |Δp|, CPU / Metal | Controls moved |
|---|---|---|---|---|---|---|
| L128 | 311 | 311/311 | — | 6.5e-6 / 1.0e-5 | 3.7e-7 / 3.6e-7 | 4/4 |
| L256 | 388 | 387/387 | 1/1 | 6.5e-6 / 8.4e-6 | 3.6e-7 / 3.7e-7 | 4/4 |
| L512 | 395 | 394/394 | 1/1 | 6.5e-6 / 7.6e-6 | 3.5e-7 / 3.5e-7 | 4/4 |
| L1024 | 398 | 397/397 | 1/1 | 6.5e-6 / 7.6e-6 | 3.5e-7 / 3.5e-7 | 4/4 |
| L2048 | 411 | 410/410 | 1/1 | 6.5e-6 / 7.6e-6 | 3.5e-7 / 3.4e-7 | 4/4 |
| L4096 | 415 | 414/414 | 1/1 | 6.5e-6 / 7.6e-6 | 3.5e-7 / 3.4e-7 | 4/4 |
| Pair Ls64+Lq64 | 201 | 201/201 | — | 6.5e-6 / 4.4e-6 | 3.7e-7 / 3.2e-7 | 3/3 |
| Pair Ls128+Lq64 | 290 | 290/290 | — | 6.5e-6 / 4.0e-6 | 3.6e-7 / 3.0e-7 | 4/4 |
| Pair Ls256+Lq128 | 390 | 390/390 | — | 7.0e-6 / 4.2e-6 | 3.9e-7 / 3.5e-7 | 4/4 |
The near tie's row has 183 tokens and its question alone 136, so the L128 file and the pairs do not hold it. One control request has a 72-token state, so the Ls64 pair holds 3 of the 4. Each control moved some option by at least 0.83.
Pictures end to end
Three requests with pictures went through the host on the shipped picture graphs and row graphs: the source card's photo (one tile, 271 tokens, L512), a 1280 × 853 picture split into 6 tiles (3 across, 2 down) plus a thumbnail (7 tiles, 1,856 tokens, L2048), and two pictures (2 tiles, 621 tokens, L1024). The token ids and the pixel values equal the provider's processor bit for bit. All three gave the reference's most likely option, with max |Δp| 3.4e-5 and mean 1.4e-5 on the CPU, and 3.1e-5 and 1.4e-5 on Metal at float32.
On-device demo
| Picture | Question | Answer | Probabilities: one / two / more |
|---|---|---|---|
| COCO val2017 000000039769.jpg (two cats on a sofa, 640 × 480) | How many cats are there? (choice: one, two, more) |
two | Metal (float32): 0.00613 / 0.98531 / 0.00855 CPU: 0.00613 / 0.98531 / 0.00855 Provider, float32 CPU: 0.00613 / 0.98532 / 0.00855 |
Run with examples/run_example.py on this repository's files: Apple M4 Max, macOS 27.0, ai-edge-litert 2.2.0 through the Python CompiledModel API, Metal at float32 precision (--accel gpu) and the CPU with 8 threads. The photo is one tile; its row has 271 tokens and runs on the L512 graph.
Desktop speed (Apple M4 Max)
These times come from an Apple M4 Max (128 GB, macOS 27.0) with ai-edge-litert 2.2.0 through the Python CompiledModel API. A time is the median of 20 requests after 5 warm-up requests. A request's time covers every graph call: writing the inputs (the host's embedding rows included), running, and reading the whole output back.
| Request | Route | Metal, float32 | CPU, 8 threads |
|---|---|---|---|
1 question, row of 39 tokens (the card example's refund) |
L128 row graph | 61.8 ms | 213.8 ms |
| 3 questions, state of 19 tokens (the card example) | Ls64 pair | 153.9 ms | 500.6 ms |
| The same 3 questions on L128 row graphs | 3 × L128 | 185.5 ms | 641.7 ms |
| 5 questions, state of 102 tokens | Ls128 pair | 255.1 ms | — |
| The same 5 questions on L256 row graphs | 5 × L256 | 544.0 ms | — |
| 1 question on a state of 3,425 tokens (row of 3,470) | L4096 row graph | 2,050.0 ms | — |
| The card's photo and 1 question (1 tile, 271 tokens) | tower, projector, L512 | 389.3 ms | 1,176.9 ms |
| 1 picture split into 6 tiles plus a thumbnail, 1 question (1,856 tokens) | 7 tower calls, L2048 | 2,183.5 ms | — |
One call of each row graph on Metal at float32: 61.8 ms (L128), 108.7 ms (L256), 207.5 ms (L512), 420.4 ms (L1024), 889.3 ms (L2048) and 2,050.0 ms (L4096). On the CPU: 213.8 ms (L128) and 390.5 ms (L256). The photo's 389.3 ms on Metal are 24.3 ms of preprocessing, 156.2 ms in the tower, 1.1 ms in the projector and 207.3 ms in the L512 graph; on the CPU the tower took 403.5 ms. "—": not measured.
Compiling a graph on Metal at float32 took 20.3 to 28.3 s for a row graph and 39.2 to 45.5 s for a pair, once per process. Each compiled graph holds its own copy of the weights: after compiling, the process footprint was 13.9 to 14.1 GB with one row graph and 23.2 GB with one pair on Metal, and 10.8 GB and 21.0 GB on the CPU. The pairs run without constant tensor sharing, so the GPU holds their weights once per signature.
Android
We could not run the model within the tolerance on a 12 GB Android phone. On one Galaxy S26 (SM-S942Q, Android 16) with LiteRT 2.2.0 through the Kotlin CompiledModel API, an earlier form of the 256-token row graph did not load. That form held float16 FULLY_CONNECTED weights like this repository's files, plus an int8 copy of the embedding table inside the graph: 5.14 GB, against 4.87 GB for this repository's L256 file. A memory guard stopped the app when the phone's MemAvailable fell below 1,500,000 kB (1.54 GB).
- GPU (OpenCL,
FP16_WITH_FP32_ACCUM): the delegate took all 1,694 nodes. 8 s into the compile, the app held 5.15 GB resident and 3.82 GB in swap, both still growing, and MemAvailable had fallen from 6.53 GB to 0.82 GB. The GPU's page allocation stayed at or below 0.30 GB. - CPU (XNNPACK, 4 threads): XNNPACK took 1,692 of the 1,694 nodes. 8 s into the compile, the app held 4.69 GB resident and 4.58 GB in swap, and MemAvailable had fallen from 7.32 GB to 0.81 GB.
- A 2.71 GB form with dynamic int8 FULLY_CONNECTED weights loaded on the CPU and took 415.1 ms per question (median of 10 calls, 4 threads). Its probabilities moved by up to 0.157 from the reference on 100 test questions, outside the 0.02 tolerance. For that file the phone gave the Mac CPU's hidden states bit for bit, on all 105 rows it ran; on the Mac, the same file moves the probabilities by up to 0.25.
The files of this repository were not run on the phone. The phone's NPU was not tried. GB here is bytes / 10⁹; the phone's /proc values are kB × 1,024 / 10⁹.
GPU precision
Run the GPU at float32 precision: the host sets GpuOptions(enforce_f32=True) for accelerator="gpu". At the Metal GPU's default precision (float16 storage), every text file stays finite but moves the probabilities outside the tolerance: max |Δp| from 0.044 to 0.086, and up to 2 questions per file change their most likely option. On the picture requests, the default precision moves the probabilities by up to 0.52 and flips the split picture's answer.
Limits
- The model never generates text. A row holds at most 4,096 tokens: the state plus one question, pictures included. A longer row is refused with
RowTooLong, never cut. The provider's card gives a context length of 32,768 tokens. - A pair holds a state of up to 256 tokens and questions of up to 128 tokens; other requests run row by row.
- A picture takes one tile, or 2 to 10 tiles plus a thumbnail: at most 2,816 picture tokens, which must fit the row with the text.
- Run the GPU at float32 precision. Its default precision is outside the tolerance (see GPU precision).
- Each compiled graph needs its own memory: 13.9 to 23.2 GB of process footprint per graph on Metal (see Desktop speed). Keep one or two text graphs compiled at a time.
- The probabilities differ from the provider's float32 code by up to 1.02e-5 on text and 3.39e-5 with pictures on the Mac. A question whose two most likely options are closer than that can change its answer; the test set's one near tie did not.
- The agreement numbers measure how closely this conversion follows the provider's code, not task accuracy. Accuracy and calibration were not measured again; see the source model card. Input languages and intended uses follow the source model card; the test requests are in English.
- Measured on one Mac (Apple M4 Max) and one Galaxy S26. Other machines and GPUs were not measured.
Provenance, conversion and license
- Source:
LiquidAI/d1-3Bat revisionda1fe36a861f24690f27f622dca1d8688503d113. The later commits on its main branch, up to051bcc46on 2026-10-07, changeREADME.mdonly. Base:LiquidAI/LFM2.5-VL-3B. Both are licensed under the LFM Open License v1.0. - Model: a SigLIP2 picture tower (27 layers, width 1,152) and projector, and the LFM2 hybrid text decoder (30 layers: 22 short-convolution and 8 grouped-query attention with 32 query and 8 key-value heads; hidden size 2,048; a 128,000-token vocabulary; tied embeddings).
- Text graphs: the provider's own LFM2 modules, with the attention re-expressed for export: batched matrix products (BATCH_MATMUL) with an additive mask, the grouped keys and values copied by concat, and the rotary tables as constants. A pair's question step picks its rotary rows with a one-hot FULLY_CONNECTED kept in float32. Exported with litert-torch 0.9.4 from the weights in float32 (transformers 5.14.1, torch 2.13.0), then quantized with ai-edge-quantizer 0.9.0: every other FULLY_CONNECTED weight is stored in float16 and read through a DEQUANTIZE operator. Activations stay in float32. A row graph has 1,691 operators, 166 of them FULLY_CONNECTED; a pair has 1,682 in
state_prefilland 1,717 inquestion_step. None is CUSTOM, no tensor is int64, and no tensor has a rank above 4. - Picture graphs: the SigLIP2 tower up to its final layer norm, without the pooling head, with the position table's resize moved to the host (the
posinput); the projector after the host's 2 × 2 unshuffle. Both store their FULLY_CONNECTED weights in float16; the tower has 1,559 operators. - Tables: the tied embedding's bytes as they are in the checkpoint (bfloat16), its float32 rows for the 1,234 read-out ids, and the tower's position table in float32.
- Scripts:
conversion/holds every conversion, check and timing script, with its environments and a README of the commands.REPRODUCE.mdlists the sources, the steps and where each number comes from. - Test data:
fixtures/holds the requests. The 144 SemIf records (MIT,fixtures/LICENSE-SemIf-MIT.txt), the source card's two examples and the 15 requests written for these conversions are included with their text. The 221 transfer-v4 requests quote public datasets under their own licenses, so they are listed by reference only (file, tag, line and SHA-256 in the development file of github.com/jaredpalmer/kev at tagkev-1.0);fixtures/rebuild_requests.pyrestores the full set from that file. The provider's float32 probabilities for every test question it answers are infixtures/reference_probs.json, which also names the one question it refuses.
License: d1-3B and LFM2.5-VL-3B are licensed under the LFM Open License v1.0 (LICENSE, a copy of the source repository's file). This repository is released under the same license, except the SemIf fixture records (MIT). Note the license's limit on commercial use by organizations with annual revenue of US$10 million or more. Attribution is in NOTICE.
Changes from the original work. The weights were converted from the checkpoint's safetensors (bfloat16) to LiteRT flatbuffers: float16 FULLY_CONNECTED weights, float32 for the rest. The text decoder was re-expressed for export as above, and split into row graphs of fixed lengths and shared-state pairs. The picture tower's position-table resize and the projector's unshuffle moved to the host. The embedding table ships separately and the host writes its rows. The tokenizer is unchanged. The host ports the provider's request handling and picture preprocessing to numpy. This repository is a community conversion and is not affiliated with Liquid AI.
- Downloads last month
- 80