You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

CL Tagger v3 โ€” Phase 2 backbone

This is the teacher-free, from-scratch FoveaTag visual backbone accepted by the Phase 2 representation-health gate on 2026-10-03. It is an experimental feature extractor, not a trained tagger. No tag head, vocabulary, learned router, or tag predictions are included. Passing this gate authorizes supervised experiments; it does not establish superiority over SigLIP2 or validate downstream tagging accuracy.

Architecture and input strategy

The shared encoder combines anti-aliased convolutional downsampling, depthwise gated blocks, shifted window attention with continuous source-coordinate 2-D RoPE, and a reduced-token global attention stage. Configuration: widths [96, 192, 384, 512], depths [2, 3, 8, 6], window size 8, attention heads [8, 8], MLP ratio 3, and 10 geometry features. The exact parameter count and tensor dtypes are recorded in release_metadata.json.

Pretraining uses a 448ร—448 overview and four 256ร—256 source-image local canvases. Local fields cover multiple scales (base field levels 2, 1, 0.5; scale jitter 0.8โ€“1.2). Detail comes from cropping the source, not magnifying an already downscaled overview. Geometry keeps each view anchored to the source image. A caller may encode additional local views in bounded batches, but this release does not decide which regions to visit.

The local map has stride 16 and 384 channels; global tokens have stride 32 and 512 channels. The encoder supports varying canvas dimensions, but unbounded full-image input is not the intended efficient inference path: global attention is quadratic in its reduced token count. Resolution extrapolation and adaptive routing remain downstream evaluation tasks. The training decoder rejects sources above 40 million pixels; this is an operational decode cap, not a claim that every larger source is supported.

Training and provenance

  • Random initialization; no SigLIP2, DINO, or other external pretrained weights and no external teacher.
  • Self-supervised coordinate correspondence, masked reconstruction, chroma reconstruction, variance, covariance, and centering objectives. No image-level tag supervision in this phase.
  • Private/local photographic and Danbooru-style illustration collections, sampled with domain weights 0.15/0.85. Training images, image paths, personal metadata, dataset database, and annotations are not distributed. Collection provenance does not imply ownership of underlying image rights.
  • Ten registered source passes; 15,467,980 presented samples, 15,264,198 decoded, 203,782 skipped (1.32%). Most skips occurred in pass 7 during a suspected transient storage outage; the cause is not proven.
  • Selected checkpoint: zero-based epoch 9, optimizer step 954,024 (nominal full budget 966,750). Actual compute shortfall is retained, not silently represented as full-budget training.
  • Batch size 16, learning rate 1e-4, weight decay 1e-4, warmup 5,000 steps, seed 20260929. Objective and input settings are included in config.json.

Phase 2 evaluation

Fixed held-out evaluation used 512 photos and 2,048 illustrations, with zero decode skips, compared with same-seed random initialization. These measure representation health, not classification quality.

Aggregate measure Photos Illustrations
Correspondence loss, pretrained 0.03093 0.02912
Correspondence loss, random initialization 0.63928 0.46538
Effective rank, pretrained local features 111.61 98.14
Effective rank, pretrained overview features 261.04 233.70

Final weighted held-out pretraining loss: 0.11306. Feature-health evidence is consistent with avoiding the earlier collapsed solution. It does not prove semantic utility, fine-detail recognition, cross-domain generalization, or production readiness. Phase 3 supervised A/B/C comparison is separate and ongoing at publication.

Use

Install torch, safetensors, numpy, Pillow, and huggingface_hub in your environment. Download the snapshot and explicitly import the small accompanying implementation; this is not a Transformers AutoModel checkpoint and does not require trust_remote_code.

import json
import sys
from pathlib import Path
import torch
from PIL import Image
from safetensors.torch import load_file
from huggingface_hub import snapshot_download

root = Path(snapshot_download("celstk/cl_tagger_v3_backbone"))
sys.path.insert(0, str(root))
from encoder import FoveaEncoder, FoveaEncoderConfig
from views import ViewConfig, build_views, pil_to_tensor

cfg = json.loads((root / "config.json").read_text())
encoder = FoveaEncoder(FoveaEncoderConfig(**cfg["encoder"])).eval()
encoder.load_state_dict(load_file(str(root / "model.safetensors")), strict=True)
image = Image.open("your_image.jpg")
if image.width * image.height > cfg["max_source_pixels"]:
    raise ValueError("Source exceeds the published training decode cap")
views = build_views(image, ViewConfig(**cfg["views"]), seed=0, training=False)
images = pil_to_tensor(views.overview).unsqueeze(0)
geometry = torch.tensor([views.overview_spec.normalized_geometry(image.width, image.height)])
with torch.inference_mode():
    features = encoder(images, geometry)
print(features.local_map.shape, features.global_summary.shape)
# Encode each source-native local view using its own geometry, or batch them.
for canvas, spec in zip(views.local_views, views.local_specs):
    geometry = torch.tensor([spec.normalized_geometry(image.width, image.height)])
    with torch.inference_mode():
        local_features = encoder(pil_to_tensor(canvas).unsqueeze(0), geometry)

views.py supplies the training-compatible aspect-preserving canvases, RGB normalization, and geometry. Geometry encodes normalized crop center/extent, logarithmic field scale and source pixels per canvas pixel, and normalized content rectangle. Do not substitute zero geometry or an unrelated image processor. Pin a Hub commit revision for reproducibility. The example uses CPU float32 execution; users may move both inputs and model to their device/dtype.

Files and integrity

  • model.safetensors: exact, unchanged Phase 2 accepted encoder state only.
  • encoder.py, views.py: matching standalone source implementation and view construction.
  • config.json: encoder, view, objective, and training settings without private paths.
  • release_metadata.json: artifact hash, selection, source revision, parameter count, provenance, and aggregate validation.
  • phase2_gate.json: public path-free projection of the accepted decision and evidence hashes; not a substitute for the original private gate verifier.
  • SHA256SUMS: checksums for the published files, excluding itself.

Expected weight SHA-256: 3efe259fc1f4a559dea0906c275c16d0a11b52d055b782e41c1c81258777a4f6. Optimizer/RNG/resume state, reconstruction heads, dataset manifests, and local filesystem paths are excluded. Source corresponds to SushiUI design-branch commit 8064bc4a50eeaa67d0a3acd8225c39e0a204dac3; this repository includes the required encoder implementation so users do not need that branch installed.

Limitations and license

This is a research-stage RGB backbone. Alpha-aware representation and transparency recognition are not validated. Native detail does not recover information absent from the source, and extra views cost compute. No individual tags or per-tag results are published here.

No model or accompanying-code license has been selected for this release yet. Public availability is not an affirmative grant of reuse, redistribution, or commercial rights. No third-party pretrained-weight license is inherited because no such weights were used; dataset/image rights remain separate and must not be inferred from this statement.

Downloads last month
-
Safetensors
Model size
37.4M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support