pixel-detector-5-f1: whole-image AI-image detector (research model)

pixel-detector-5-f1.onnx is a CLIP ViT-B/16 vision transformer, fine-tuned end to end to separate camera photos from AI-generated images. It reads one 224 × 224 view of the whole image and returns one logit: positive leans AI, negative leans camera.

It is the learned detector of Slop Scope, a research app built to test JEV. JEV is a typed decision model the app reaches through OpenRouter. It receives only fixed text facts, never pixels or numbers, and it is the app's only classifier. This model's logit becomes at most one of those facts.

Research model, not a verdict.

  • The logit is not a calibrated probability, and no score verifies the origin of a single image.
  • The model misses or leaves undecided about a quarter of our development images.
  • It calls some heavily retouched real portraits AI.
  • Do not use it to make decisions about people.

Files

File Bytes SHA-256
pixel-detector-5-f1.onnx 171,894,049 a3d4b9aaf1b9096116465cad4808b0b4497a16eb16ef0fdb71023751327d27d8

This repository holds only the model and this card: no training images, no manifests.

Model details

  • Architecture. The CLIP ViT-B/16 vision tower (openai/clip-vit-base-patch16, 86 M parameters) with a linear head (1536 → 1). The head reads two vectors joined together:
    • the pooled output;
    • the post-layernorm mean of the patch tokens.
  • Training. End to end:
    • 5 epochs, batch 32;
    • AdamW: backbone learning rate 1e-5, head 1e-3, weight decay 0.05;
    • one-cycle schedule (10% warm-up), fp16 autocast, seed 7.
    • The last epoch is the model; validation was reported, not used to pick an epoch.
    • The loss is binary cross-entropy with logits; the target is 1 for AI.
  • Export. ONNX opset 17 (IR 8, PyTorch 2.10.0 exporter):
    • Weights of 1,024 values or more are stored as float16 and cast back to float32 in the graph, so all computation is float32.
    • Batch size is fixed at 1.
  • Input. view: float32 [1, 3, 224, 224], RGB, values 0–255 (not normalised). CLIP's mean and standard deviation are applied inside the graph.
  • Output. logit: float32 [1].
  • Parity. PyTorch and ONNX Runtime 1.24.1 (CPU) agree within 0.0049 logit (tolerance 0.02) on 6 synthetic and 12 real files. The app runs the file with onnxruntime-web 1.30.0 (WebAssembly).

How to use

The view must be made exactly as in training. A different resizer gives a different view and a shifted logit. This is the training pipeline in Python with Pillow:

import numpy as np
import onnxruntime as ort
from PIL import Image, ImageOps

def decode(path):
    """EXIF orientation applied, transparency on white, RGB; ICC profiles ignored."""
    img = ImageOps.exif_transpose(Image.open(path))
    if img.mode in ('RGBA', 'LA') or (img.mode == 'P' and 'transparency' in img.info):
        rgba = img.convert('RGBA')
        img = Image.alpha_composite(Image.new('RGBA', rgba.size, (255, 255, 255, 255)), rgba)
    return img.convert('RGB')

def view(img):
    """The 224 x 224 view the model was trained on, or None (no score) under 256 px on the shorter side."""
    w, h = img.size
    if min(w, h) < 256:
        return None
    # 1. Long side to 320 px with Lanczos, never upscaled (Python's round: half to even).
    s = 320 / max(w, h) if max(w, h) > 320 else 1.0
    W, H = max(1, round(w * s)), max(1, round(h * s))
    # 2. Centre-crop the shorter side down to a multiple of 16.
    W2, H2 = (W, H // 16 * 16) if W >= H else (W // 16 * 16, H)
    if min(W2, H2) < 16:
        return None
    if (W, H) != img.size:
        img = img.resize((W, H), Image.LANCZOS)
    left, top = (W - W2) // 2, (H - H2) // 2
    img = img.crop((left, top, left + W2, top + H2))
    # 3. Short side to 224 px with bicubic, then a centre crop of 224 x 224.
    s = 224 / min(W2, H2)
    W3, H3 = max(224, round(W2 * s)), max(224, round(H2 * s))
    img = img.resize((W3, H3), Image.BICUBIC)
    left, top = (W3 - 224) // 2, (H3 - 224) // 2
    return np.asarray(img.crop((left, top, left + 224, top + 224)), dtype=np.uint8)

session = ort.InferenceSession('pixel-detector-5-f1.onnx')
v = view(decode('photo.jpg'))
if v is not None:
    x = v.astype(np.float32).transpose(2, 0, 1)[None]          # [1, 3, 224, 224], RGB 0-255
    logit = float(session.run(['logit'], {'view': x})[0][0])   # > 0 leans AI, < 0 leans camera

The Slop Scope app ports Pillow's resampler to JavaScript. Its views are byte-identical to Pillow's on 47 test files.

Decision bands used in Slop Scope

The app turns the logit into at most one fact, using four fixed cutoffs. The cutoffs were set on a calibration population that the model did not train on:

  • public photos and their platform copies;
  • phone photos from held-out devices, with their WhatsApp and Facebook copies;
  • AI images from older and current generators.
Band Rule Budget on the calibration population
strong AI logit > 6.036 at most 0.5% of real rows
AI logit > 2.799 at most 2% of real rows
camera logit ≤ −4.133 at most 2% of AI rows
strong camera logit ≤ −7.895 at most 0.5% of AI rows

Between −4.133 and 2.799 there is no band. These are budgets on one calibration population, not error rates you can expect on your images.

Training data

All sources allow publishing a non-commercial research model trained on them. The terms were checked on 2026-09-27; this is our reading, not legal advice.

Source Use Images in training Licence
OpenFake (arXiv 2509.09495): AI images (75 generators in the main pool); real photos from its Pexels and LAION subsets only train, validation, calibration 4,773 real, 4,748 AI CC BY-NC 4.0
VISION (University of Florence; Shullani et al. 2017): phone photos from 28 devices, originals plus WhatsApp and Facebook copies train, calibration (held-out devices) 8,218 real CC BY-SA 4.0
bitmind/nano-banana and eight Rapidata text-to-image preference sets (listed above; 32 generator labels in total) train, validation, calibration 3,813 AI MIT; CDLA-Permissive-2.0
  • Totals.
    • Training: 21,552 images (12,991 real, 8,561 AI).
    • Validation: 1,663.
    • Calibration: 5,140 real and 2,662 AI rows.
  • Exclusions.
    • Near-duplicates (exact hash, or difference hash within Hamming distance 6) were removed against the calibration and development sets and across the train/validation line.
    • Unsplash Lite was never used: its terms allow training only for internal business purposes.
    • The authors' private development folders were never used for training or calibration.
  • Augmentation, both classes.
    • Two times in three, a shared-copy recipe: messaging app, forwarded, web export, reduced, sharpened or PNG-reduced, with JPEG quality 70–95.
    • A mild edit 15% of the time, and a scan-like pass 5% of the time.
    • A random resampler and a horizontal flip.

JEV's part in training

Each training image was weighted by how JEV fared on it in three earlier judging rounds:

  • 1 if JEV was right;
  • 3 if JEV was uncertain;
  • 5 if JEV was wrong.

Images JEV never judged took their stratum's miss rate, shrunk toward the class rate, and the weights were rescaled to a mean of 1 per class. Only 1,336 of the 21,552 images had a JEV verdict, and 92% of those were uncertain. So f1's weights are almost uniform (SD 0.06).

We make no claim that JEV weighting improved this model. Later experiments retrained f1's recipe with and without JEV weighting, on development data:

  • JEV weighting gave more misses (376 against 316 of 810) but fewer wrong answers (24 against 39).
  • JEV in both training and classification did not differ from no JEV (238 against 231 misses on our round-4 set, p = 0.45).

Evaluation

All numbers below come from development data. Round 4 and the private set informed how this model was built and chosen, so none of them is a held-out accuracy claim. "Decided" means the image fell in a band; the rest are undecided.

Set AUC Real: decided right / wrong AI: decided right / wrong
Validation 0.992
Calibration 0.984
Round 4 (public development set) 0.978 70.8% / 2.9% 87.0% / 1.7%
Private labelled development set (the authors' own photos and AI images) 0.946 37.7% / 16.9% 89.9% / 0.5%
Held-out VISION phones (leave one device out; one device has a same-model twin in training), with WhatsApp and Facebook copies 98–99% / 0%
Newest generators (prompts held out) 93.2% / 1.6%

The "decided" columns use the 2% bands.

Comparison on our 810-image development screen, counting a miss as wrong or undecided:

  • this model missed 215 (27%), 35 of them wrong;
  • the Community Forensics ViT-S detector missed 421 (p < 0.0001).

The project's own bar is not met. It allows at most 1% misclassified, with undecided counted as a miss.

The app's evidence gate. Before a detector fact may reach JEV, it should fire on at most 2% of wrong-class images and source groups.

  • The camera-side bands pass.
  • Both AI-side bands fail on real development photos:
    • strong AI: 7 of 121 real source groups, 5.4% of real variants;
    • AI: 19 of 121 groups, 9.4% of variants.

Limitations and known failure modes

  • Retouched real portraits (studio backdrops, Lightroom-style event edits) are often called AI: about a quarter of one retouched sub-folder in our private set.
  • People as a cue. 54% of the training images with people in them were AI, against 40% overall, so a person in the picture pushes the model toward AI.
  • Low resolution. The view is 224 px: the model judges content and style, not camera sensor texture.
  • Recompression. Real photos re-shared through messaging apps from devices outside the training data are decided less often. The held-out phones' WhatsApp and Facebook copies do well.
  • Small images. Images under 256 px on the shorter side get no score.
  • Generators change. New generators can be missed; at least one current-generator portrait was called real.
  • Not a probability. The logit is not calibrated. Bands calibrated on one population do not carry over to another.

Intended use

  • Intended: research on AI-image detection, and use as one bounded, disclosed input to a system that keeps its own decision layer and uncertainty, as in Slop Scope.
  • Out of scope:
    • commercial use;
    • decisions about individuals;
    • facial recognition or biometric profiling (also excluded by OpenFake's terms);
    • presenting the logit as proof that an image is or is not AI-generated.

Use in Slop Scope

Slop Scope pins this file by commit, size and SHA-256. A visitor's browser downloads it only after the visitor clicks Load model, checks it and keeps it locally. The image never leaves the browser.

The band becomes one of the dossier's fixed text facts, and JEV judges the dossier. JEV never receives the logit, a number or a pixel. Hosting must stay non-commercial.

Licence and credits

Licence. CC BY-NC 4.0 (non-commercial), carried from OpenFake. VISION is CC BY-SA 4.0. Whether its share-alike term reaches model weights is unsettled, so we name it here.

Credits, with licences:

  • OpenFake: ComplexDataLab/OpenFake, arXiv:2509.09495, CC BY-NC 4.0. That includes its Pexels and LAION real subsets: third-party stock and web images that OpenFake redistributes.
  • VISION: D. Shullani, M. Fontani, M. Iuliani, O. Al Shaya and A. Piva, "VISION: a video and image dataset for source identification", EURASIP Journal on Information Security, 2017. CC BY-SA 4.0.
  • bitmind/nano-banana: MIT.
  • Rapidata: the eight preference sets listed above, CDLA-Permissive-2.0.
  • CLIP ViT-B/16: OpenAI, A. Radford et al., "Learning Transferable Visual Models From Natural Language Supervision", 2021. MIT.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for GPTchatly/slope-scope

Quantized
(8)
this model

Datasets used to train GPTchatly/slope-scope

Paper for GPTchatly/slope-scope