gpt-oss-vl-exp
A vision-language model built by bolting a translator onto a frozen 20B reasoning model β no new architecture, no new tokens, no touching the brain's weights.
What this is
image β SigLIP2-so400m (frozen eye, 428M)
β 2-layer MLP adapter (16.5M, TRAINED) β "the translator"
β 256 vectors in gpt-oss's own embedding space
β gpt-oss-20b (frozen brain, 20.9B MoE) β does the thinking
+ LoRA inserts on attention q/k/v/o (7.96M, TRAINED) β stage 2
The splice is LLaVA-style: render the harmony chat prompt with an
{{IMAGE}} marker, split it, embed both halves, and concatenate the image
vectors in between. No special tokens, no modified attention, nothing inside
gpt-oss was ever re-designed β the transformer is a universal continuer of
vector conversations, so anything placed in its embedding space gets
processed like text. (Proof from our own logs: the same brain fed random
junk vectors produces noise; fed trained-adapter vectors, it reads
diagrams.)
Results (honest version)
- Stage 1 (adapter only): training loss 7.52 β 1.13 in 1,250 steps. Smoke test: the frozen brain opened a harmony analysis channel and reasoned about a parallelogram using point labels that exist only in the image.
- Stage 2 (+LoRA on attention): loss 1.10 β 0.96. Text-channel discipline improved sharply; image descriptions got concretely grounded (it read the name "Sultan" off a classroom whiteboard).
- 24-question trap exam: ~40% VQA accuracy across both stages β above chance, below any other vision LLM from the past 2 years. Honest verdict: a real working VLM prototype that "sees the gist, misses the details miserably."
- Live A/B on a 16GB consumer GPU: same map, same question, greedy, only the LoRA toggled β with inserts: correct "United States of America" with grounded map-reading; without: hallucinated geography collapsing into a repetition loop. The trained inserts are load-bearing for grounding and channel stability (but don't reliably aim at truth β an orange cat comes out "pig"/"rabbit"/"white" depending on toggles and temperature; an orange sunset sky reads "blue": the diagram-diet adapter never learned object color). Full profile in the HF model card.
Interesting tidbits
- 24.5M trained parameters = 0.12% of the stack. The other 99.88% is frozen OpenAI + Google weights doing their day jobs.
- Total compute bill: ~$20 on Modal (A100-80GB, ~7 GPU-hours across everything, including every failed launch) β plus $0.40 on OpenRouter for the coding agent (Pi using GLM5.3-Flash) that wrote the whole thing.
- gpt-oss's native MXFP4 quantization means the brain checkpoint is smaller than a Q5 requant of itself would be. We kept it untouched.
- The eye outputs 256 native-aspect patch tokens at 448px (SigLIP2 NaFlex, patchified), not a fixed grid β variable token counts per image.
- Failure curriculum, lovingly preserved: 6 data-prep attempts, 4 training launch failures (int unpacking, dtype boundary, batch-dim strip, 76GB OOM from full-vocab logits), 1 smoke-test scoping bug, and 1 launcher typo. Every one is documented in the git history.
The two halves
Code and weights are split across two hosts, and each is useless without the other:
- GitHub (https://github.com/OverThereNotHere/gpt-oss-vl) β the runner and the story. Without the weights it's a car manual with no car.
- Hugging Face (you are here) β the three weights pieces. Without the code they're three inert piles of numbers.
Weights
| File | What it is |
|---|---|
gpt-oss-20b/ |
unmodified OpenAI checkpoint (brain) |
siglip2-so400m-naflex/ |
unmodified Google checkpoint (eye) |
stage2_final.pt |
our trained artifacts: {"adapter": 16.5M, "lora": 96 A/B inserts, "step": 1250} |
Credits
Built in one weekend for ~$20.50, mostly by asking "but why does that work?" until the answers got interesting.
Model tree for NotHereNorThere/gpt-oss-vl-exp
Base model
google/siglip2-so400m-patch16-naflex