P-JiT (Pyramid-JiT)

Code: github.com/Linum-AI/pyramid-jit
Blog post: Pyramid-JiT: Low-Res Drafts, High-Res Images

⚠️ Research checkpoint

This is an experimental research artifact, not a production model. P-JiT was trained for 138M samples and has not been post-trained. It is a research preview on the road to our v3 model. We're releasing it to share our preliminary findings with the broader field and to encourage others to explore efficient training methods like ours. Expect rough edges.

P-JiT matches Linum v2's final FD-DINOv2 in 11.3x fewer samples

P-JiT generates better samples than Linum v2, with 4x more aggressive token reduction: the same prompt from Linum v2, JiT-DDT and P-JiT

P-JiT reaches Linum v2's FD-DINOv2 on 11.3x fewer training samples and trains in 4.3x fewer GPU-hours, at 4x the pixels. Both figures are from the blog post, which has the full write-up.

P-JiT is a 2.2B-parameter pixel-space text-to-image diffusion transformer from Linum: a single decoder-only DiT that reads the noisy image as 32x32-pixel patches (256 tokens at 512x512) together with the caption and predicts the clean image directly, with no VAE. During training, extra readout heads predicted the image at 128x128 (after block 10) and 256x256 (after block 16) on the way to the final 512x512 head. Captions are encoded with Qwen3.5-4B (hidden layers 7/15/27 concatenated).

Code, sampler, and the reference training loss: github.com/Linum-AI/pyramid-jit.

Samples

Six P-JiT samples at 512x512: a red-haired woman, an old fisherman, an oil painting of a ship in a storm, a tiger in a river, a fish mosaic, an animated robot

512x512 samples from the blog post's appendix.

Files

File Contents
model.safetensors EMA weights, fp32, 538 tensors, 2,184,091,606 parameters (8.74 GB). Keep fp32: the model runs its fp32 master weights under bf16 autocast, which is what the samples in this card were generated with.
config.json Architecture, text-encoder recipe, sampler defaults, and provenance (training step, EMA half-life).
export_manifest.json Export record: weights SHA256, tensor and parameter counts, training step.

Usage

git clone https://github.com/Linum-AI/pyramid-jit && cd pyramid-jit && pip install -e .
PROMPT="A close-up portrait of a young white woman with vibrant, fiery red hair cascading over \
her shoulders in soft waves, framed from the shoulders up and centered against a softly blurred \
warm-toned background. Her fair, lightly freckled complexion sets off piercing green eyes and a \
subtle, closed-lipped smile. Soft natural light enters from the left of the frame, highlighting \
the texture of her hair and the curve of her cheek while leaving the right side in gentle shadow. \
A shallow depth of field renders the background into smooth, neutral bokeh. Lights dangle out of \
focus on the left side of the frame."
python generate.py --weights Linum-AI/pyramid-jit --qwen_model_path Qwen/Qwen3.5-4B \
    --prompt "$PROMPT" --seeds 42
from pyramid_jit import PyramidJiT, QwenTextEncoder, SamplerConfig, generate

model = PyramidJiT.from_pretrained("Linum-AI/pyramid-jit")    # downloads from the Hub
text_encoder = QwenTextEncoder("Qwen/Qwen3.5-4B")
images = generate(model=model, text_encoder=text_encoder, prompt=prompt, seeds=[42],
                  sampler=SamplerConfig())

This generates the top-left sample above. The model was trained on dense captions (around 80 words); rewrite short prompts into a detailed description first.

Sampling: 50 Euler steps, adaptive projected guidance at scale 15, initial noise scale 2. The initial noise is drawn exactly as torch.randn draws it on an H100 SXM, the GPU the model was trained on, so a seed gives the same image on any NVIDIA GPU (see the code repository's README).

Provenance

Weights are the EMA (half-life 3,907 steps, from step 977) after 134,766 steps at batch size 1,024: 138M image samples, with logit-normal timesteps (0.8, 0.8) for the first 50M samples and (-0.2, 1.0) after.

The two readout heads (2.85M parameters) and the 145M-parameter PixelREPA masked transformer adapter (the representation-alignment loss's projection head) are training-only and not included; reference implementations are in the code repository's loss.py.

Authorship

This model card, and the code repository it points to, were written by Claude (Anthropic's Opus 5.5 model, running in Claude Code) at Linum's request: it extracted the model, inference code and training loss from Linum's internal experiment repository, exported these weights, and verified that they reproduce the internal sample images pixel for pixel. The model, its training, and the review of this release are Linum's.

License

Apache-2.0. Copyright 2026 Linum Inc.

Downloads last month
4
Safetensors
Model size
2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support