Altitude A2UI β€” browser build (Gemma 4 E2B, int8)

A single 2.14 GiB file that generates A2UI layouts built only from the Altitude design system, running entirely inside a browser tab on WebGPU. No server, no API key, nothing leaves the machine. The same file runs on Android and through the LiteRT-LM command line.

Given a request in plain English β€” "a settings page for a smart thermostat with schedule and away mode" β€” it returns a JSON layout referencing Altitude components by their real tag names (al-card, al-button, al-tabs). Nothing it emits is executable; the host renders the layout with its own component library.

Results

250 held-out requests, never seen in training, scored by the Altitude guardrail. A pass means the output needed no repair at all: valid JSON, every component and property real, and a sound layout tree with no duplicate ids, dangling references or orphaned nodes.

Prompt Right on the first try Valid output Invented components
This file (int8, LiteRT-LM/WebGPU) full catalog 240 / 250 (96.0%) 250 / 250 none
This file compact catalog 70 / 250 (28.0%) 213 / 250 3.2% of answers
Same model, full precision, on a server full catalog 239 / 250 250 / 250 none
Same model, full precision, on a server compact catalog 238 / 250 250 / 250 none
Untrained Gemma 4 E2B full catalog 0 / 250 207 / 250 4.0% of answers

Read the second row before you use this file. The headline number belongs to a specific prompt, and the next section explains why that is not a footnote.

Against the full catalog prompt it has never invented a component that does not exist, at any checkpoint of any training run, and it uses 53 of the 55 components in the catalog.

The ten remaining failures there are bookkeeping rather than misunderstanding: repeated content, a duplicate parent reference, one reused id. It knows the design system; it occasionally loses track of what it has already written.

The prompt is part of the model

This model was trained on two renderings of its catalog, a full one of about 2,750 tokens and a compact one of about 1,550. At full precision both work: 239 and 238 of 250. After compression only one of them does. This file scores 240 on the full rendering and 70 on the compact one. 170 requests that are answered correctly under the full prompt fail under the compact prompt, and not one goes the other way.

The failure is identifier drift. The model names the parts of the page it is about to build, then builds them under different, shorter names:

{"id":"root", "children":["pageHeader","starterStepper","buttonRow"]},
{"id":"Header",  ...}    // should be pageHeader
{"id":"Stepper", ...}    // should be starterStepper
{"id":"button",  ...}    // should be buttonRow

Two thirds of the failures are that. A tenth of the answers come out too damaged to parse as JSON at all.

Two explanations were tested and ruled out. It is not the vocabulary trim: both renderings tokenize about two tokens longer under the trimmed tokenizer than the original, symmetrically and trivially. It is not the compact prompt itself: the full-precision weights handle it at 238. It takes the compression and the shorter prompt together, and each is harmless alone. The mechanism is not established.

What this means if you use this file. The prompt below is not a suggestion, it is part of what was measured. Shorten it, re-order it, translate it or trim the catalog and the 240 no longer applies to you, possibly by a very large margin, and you will get no warning other than output that looks subtly wrong. Compression appears to cost robustness that a longer, more explicit prompt is covering up. If you change the prompt, re-measure on your own held-out set before you trust it.

That generalises past this model. A benchmark score does not transfer across a change of prompt once the weights have been compressed. We learned it the expensive way: a public demo of this file ran on the compact prompt for a day, on the strength of the 238 that was measured on uncompressed weights.

Running it

Needs WebGPU and roughly 3 GB of browser memory. Works on a current laptop; will not work on a phone browser. Median 35 s per request on an M1 with 16 GB β€” budget for that, and stream the output, or the page will look frozen.

prompt is not free text. It is the full catalog rendering wrapped in this model's chat turns, and it is what every number above was measured with. It ships in this repository as prompt-full.txt, with __REQUEST__ where the user's request goes, so you can reproduce the numbers exactly. Templating is applied by hand because the artifact's own chat template uses a construct the runtime's template engine does not implement.

import litert_lm

template = open('prompt-full.txt').read()          # ships with this model; already chat-wrapped
prompt = template.replace('__REQUEST__', request)  # do not shorten it: see the section above

eng = litert_lm.Engine('altitude-a2ui-e2b-v32k-int8.litertlm', backend=litert_lm.Backend.GPU)
s = eng.create_session(apply_prompt_template=False,                  # the text above is already templated
                       sampler_config=litert_lm.SamplerConfig(top_k=1))  # greedy
s.run_prefill([prompt])
print(s.run_decode().texts[0])

In the browser the file is placed in the WASM heap rather than streamed, because the runtime's default path refuses an ordinary .litertlm: it can stream neither the tokenizer section nor a standard prefill/decode model.

How it was made, and what that costs you

Fine-tuned from google/gemma-4-E2B-it (Apache-2.0) on 3,487 examples whose targets were all machine-checked to need zero repair, then merged, vocabulary-trimmed from 262,144 entries to 32,768, and exported at int8.

The trim is what makes it fit: the two vocabulary tables are 54% of the model's 5.1 B parameters, and the browser allows about 3.3 GB. Trimming preserves tokenization exactly for the training and evaluation corpus β€” 16,648 texts re-tokenized with zero mismatches β€” but text far outside that distribution will split into more pieces than the original model would use. In practice that means English prose about interfaces is fine, and unusual languages or heavy Unicode are not what this was built for.

One caveat on that check, since it is the kind of thing that invites misplaced confidence: it covered the prompt bodies, not the chat-wrapped strings actually fed to the runtime, which is why a shipped prompt tokenizes about two tokens longer than the original tokenizer would make it. That difference is real but tiny, it affects both renderings equally, and it is not the cause of the compact-prompt collapse above.

Only int8 survives. Every int4 variant tried destroyed the fine-tune.

Scope, honestly

The claim is about this design system, its 55-component catalog, the full catalog prompt, and requests that look like the 250 tested: plain-English descriptions of ordinary interfaces. It is not a claim about arbitrary interfaces, other design systems, or a model that needs no checker. Roughly one answer in twenty still needs the guardrail.

Attribution

Base model google/gemma-4-E2B-it by Google, Apache-2.0. Fine-tune and A2UI catalog by Southleft. The A2UI protocol is Google's.

Downloads last month
49
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for southleft/altitude-a2ui-e2b-litertlm

Finetuned
(384)
this model