MicroCLIP checkpoints

Every trained checkpoint behind the results in github.com/umutonuryasar/microclip: a CLIP-style image–text model built from scratch on a single-GPU budget, used to test whether SigLIP's sigmoid loss keeps its small-batch advantage over softmax (InfoNCE) at small scale. At this scale it does not: softmax matched or beat sigmoid at every batch size from 128 to 512.

Live demo: huggingface.co/spaces/umutonuryasar/microclip

Contents

One best.pt per run: the lowest-validation-loss weights over training. The main comparison is 3 seeds (42/43/44) at batch 128 and 512; batch 256 and all ablations are single-seed. Checkpoints are byte-identical to the training outputs; the MD5 sums below let you verify that.

Run Config in the GitHub repo MD5
abl_init_he configs/ablations/init_he.yml cb1730335670064a520d502c2c1e8bfa
abl_init_xavier configs/ablations/init_xavier.yml 67b3043f79775e0103ad129f95cfdbe7
abl_lr_constant configs/ablations/lr_constant.yml 968a6201effa74ef067ce05ab523ac5b
abl_sgd configs/ablations/optimizer_sgd.yml 3644b925db5b8566bb3e1c5f363d31ac
abl_vit_tiny configs/ablations/vit_tiny.yml f10c641a9b975743cbfbe5ae3d7962cb
base_s42_ep10 configs/base.yml db0f367fc5875576aebd6a55e9694563
sigmoid_b128_s42 configs/ablations/sigmoid_b128.yml 11d4f57fcf4d6ec4bb7e08c0ac9f76b3
sigmoid_b128_s43 configs/ablations/sigmoid_b128.yml 971bee951f74a8db4494fc315551ad3a
sigmoid_b128_s44 configs/ablations/sigmoid_b128.yml 10845415f9c6de6e9d3329abbb179c40
sigmoid_b256 configs/ablations/sigmoid_b256.yml fb7a9f5c0e4d42f3217ebba7e20515e2
sigmoid_b512_s42 configs/sigmoid_b512.yml 41b0daa4dd4a21c0afa9064a895feb15
sigmoid_b512_s43 configs/sigmoid_b512.yml 29d9a7bd82e1cb47d11da404674381ca
sigmoid_b512_s44 configs/sigmoid_b512.yml 3aa749c68e315d949e65ad52f359dd05
softmax_b128_s42 configs/ablations/softmax_b128.yml c68331d17d7b74e82e27775636e6d5a5
softmax_b128_s43 configs/ablations/softmax_b128.yml 96328777032b47bb728fd758d41f4156
softmax_b128_s44 configs/ablations/softmax_b128.yml b7727297db18a5a1ac89630bc5765151
softmax_b256 configs/ablations/softmax_b256.yml 68977bdea7b5b73cf6d78162bb3145a0
softmax_b512_s42 configs/softmax_b512.yml 972068c8032016370e75d191b3fa33a7
softmax_b512_s43 configs/softmax_b512.yml 70bc48b76a997d9966b4d060f1c09366
softmax_b512_s44 configs/softmax_b512.yml b1faafb315d0b4425fc16fe412f56b59

Also included: eval_per_run.csv and eval_summary.csv (identical to results/ in the repo), eval_results.csv (an earlier single-seed evaluation pass) and queue_log.txt (the training queue's timestamps).

Loading

Each file is a dict with a single model key holding the state dict, loadable with weights_only=True. The config defines the architecture; it must be the one the run was trained with, per the table above.

import torch
from microclip.config import load_config
from microclip.data.tokenizer import CaptionTokenizer
from microclip.models.microclip import MicroCLIP

cfg = load_config("configs/softmax_b512.yml")
tokenizer = CaptionTokenizer(cfg["tokenizer"]["path"])
model = MicroCLIP(cfg, vocab_size=tokenizer.vocab_size, max_len=cfg["data"]["max_text_len"])
state = torch.load("runs/softmax_b512_s42/best.pt", map_location="cpu", weights_only=True)
model.load_state_dict(state["model"])

Evaluate with the repo's script, for example COCO 5K retrieval:

python scripts/evaluate.py --config configs/softmax_b512.yml \
    --checkpoint runs/softmax_b512_s42/best.pt --task retrieval

Limits

Trained from scratch on COCO train2017 only, for 30 epochs (the recipe ablations for 10). Absolute retrieval quality is modest by design; see the repo README for the full results and limitations.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train umutonuryasar/microclip-checkpoints