CL Tagger v3 โ Phase 2 backbone
This is the teacher-free, from-scratch FoveaTag visual backbone accepted by the Phase 2 representation-health gate on 2026-10-03. It is an experimental feature extractor, not a trained tagger. No tag head, vocabulary, learned router, or tag predictions are included. Passing this gate authorizes supervised experiments; it does not establish superiority over SigLIP2 or validate downstream tagging accuracy.
Architecture and input strategy
The shared encoder combines anti-aliased convolutional downsampling, depthwise gated blocks, shifted window attention with continuous source-coordinate 2-D RoPE, and a reduced-token global attention stage. Configuration: widths [96, 192, 384, 512], depths [2, 3, 8, 6], window size 8, attention heads [8, 8], MLP ratio 3, and 10 geometry features. The exact parameter count and tensor dtypes are recorded in release_metadata.json.
Pretraining uses a 448ร448 overview and four 256ร256 source-image local canvases. Local fields cover multiple scales (base field levels 2, 1, 0.5; scale jitter 0.8โ1.2). Detail comes from cropping the source, not magnifying an already downscaled overview. Geometry keeps each view anchored to the source image. A caller may encode additional local views in bounded batches, but this release does not decide which regions to visit.
The local map has stride 16 and 384 channels; global tokens have stride 32 and 512 channels. The encoder supports varying canvas dimensions, but unbounded full-image input is not the intended efficient inference path: global attention is quadratic in its reduced token count. Resolution extrapolation and adaptive routing remain downstream evaluation tasks. The training decoder rejects sources above 40 million pixels; this is an operational decode cap, not a claim that every larger source is supported.
Training and provenance
- Random initialization; no SigLIP2, DINO, or other external pretrained weights and no external teacher.
- Self-supervised coordinate correspondence, masked reconstruction, chroma reconstruction, variance, covariance, and centering objectives. No image-level tag supervision in this phase.
- Private/local photographic and Danbooru-style illustration collections, sampled with domain weights 0.15/0.85. Training images, image paths, personal metadata, dataset database, and annotations are not distributed. Collection provenance does not imply ownership of underlying image rights.
- Ten registered source passes; 15,467,980 presented samples, 15,264,198 decoded, 203,782 skipped (1.32%). Most skips occurred in pass 7 during a suspected transient storage outage; the cause is not proven.
- Selected checkpoint: zero-based epoch 9, optimizer step 954,024 (nominal full budget 966,750). Actual compute shortfall is retained, not silently represented as full-budget training.
- Batch size 16, learning rate 1e-4, weight decay 1e-4, warmup 5,000 steps, seed 20260929. Objective and input settings are included in
config.json.
Phase 2 evaluation
Fixed held-out evaluation used 512 photos and 2,048 illustrations, with zero decode skips, compared with same-seed random initialization. These measure representation health, not classification quality.
| Aggregate measure | Photos | Illustrations |
|---|---|---|
| Correspondence loss, pretrained | 0.03093 | 0.02912 |
| Correspondence loss, random initialization | 0.63928 | 0.46538 |
| Effective rank, pretrained local features | 111.61 | 98.14 |
| Effective rank, pretrained overview features | 261.04 | 233.70 |
Final weighted held-out pretraining loss: 0.11306. Feature-health evidence is consistent with avoiding the earlier collapsed solution. It does not prove semantic utility, fine-detail recognition, cross-domain generalization, or production readiness. Phase 3 supervised A/B/C comparison is separate and ongoing at publication.
Use
Install torch, safetensors, numpy, Pillow, and huggingface_hub in your environment. Download the snapshot and explicitly import the small accompanying implementation; this is not a Transformers AutoModel checkpoint and does not require trust_remote_code.
import json
import sys
from pathlib import Path
import torch
from PIL import Image
from safetensors.torch import load_file
from huggingface_hub import snapshot_download
root = Path(snapshot_download("celstk/cl_tagger_v3_backbone"))
sys.path.insert(0, str(root))
from encoder import FoveaEncoder, FoveaEncoderConfig
from views import ViewConfig, build_views, pil_to_tensor
cfg = json.loads((root / "config.json").read_text())
encoder = FoveaEncoder(FoveaEncoderConfig(**cfg["encoder"])).eval()
encoder.load_state_dict(load_file(str(root / "model.safetensors")), strict=True)
image = Image.open("your_image.jpg")
if image.width * image.height > cfg["max_source_pixels"]:
raise ValueError("Source exceeds the published training decode cap")
views = build_views(image, ViewConfig(**cfg["views"]), seed=0, training=False)
images = pil_to_tensor(views.overview).unsqueeze(0)
geometry = torch.tensor([views.overview_spec.normalized_geometry(image.width, image.height)])
with torch.inference_mode():
features = encoder(images, geometry)
print(features.local_map.shape, features.global_summary.shape)
# Encode each source-native local view using its own geometry, or batch them.
for canvas, spec in zip(views.local_views, views.local_specs):
geometry = torch.tensor([spec.normalized_geometry(image.width, image.height)])
with torch.inference_mode():
local_features = encoder(pil_to_tensor(canvas).unsqueeze(0), geometry)
views.py supplies the training-compatible aspect-preserving canvases, RGB normalization, and geometry. Geometry encodes normalized crop center/extent, logarithmic field scale and source pixels per canvas pixel, and normalized content rectangle. Do not substitute zero geometry or an unrelated image processor. Pin a Hub commit revision for reproducibility. The example uses CPU float32 execution; users may move both inputs and model to their device/dtype.
Files and integrity
model.safetensors: exact, unchanged Phase 2 accepted encoder state only.encoder.py,views.py: matching standalone source implementation and view construction.config.json: encoder, view, objective, and training settings without private paths.release_metadata.json: artifact hash, selection, source revision, parameter count, provenance, and aggregate validation.phase2_gate.json: public path-free projection of the accepted decision and evidence hashes; not a substitute for the original private gate verifier.SHA256SUMS: checksums for the published files, excluding itself.
Expected weight SHA-256: 3efe259fc1f4a559dea0906c275c16d0a11b52d055b782e41c1c81258777a4f6.
Optimizer/RNG/resume state, reconstruction heads, dataset manifests, and local filesystem paths are excluded. Source corresponds to SushiUI design-branch commit 8064bc4a50eeaa67d0a3acd8225c39e0a204dac3; this repository includes the required encoder implementation so users do not need that branch installed.
Limitations and license
This is a research-stage RGB backbone. Alpha-aware representation and transparency recognition are not validated. Native detail does not recover information absent from the source, and extra views cost compute. No individual tags or per-tag results are published here.
No model or accompanying-code license has been selected for this release yet. Public availability is not an affirmative grant of reuse, redistribution, or commercial rights. No third-party pretrained-weight license is inherited because no such weights were used; dataset/image rights remain separate and must not be inferred from this statement.
- Downloads last month
- -