SenseNova-U1.5-8B-MoT-pose-8bit

A pose-transfer tier of SenseNova-U1.5-8B-MoT for sensenova-u1-swift on Apple silicon: the RefControl pose adapter merged into the base checkpoint, then quantized to 8-bit (group 64) on the two transformer streams.

What it does

Restages the person in a reference photo into the pose of an OpenPose-style skeleton, keeping their identity, clothing and scene. It is a two-reference edit, and the slot order matters:

Image-1: <image>
Image-2: <image>
apply pose from image 1 with reference from image 2

Image 1 is the skeleton, image 2 the identity/appearance frame. Swap them and the model reproduces the appearance frame's own pose and ignores the skeleton.

sensenova-cli --weights SenseNova-U1.5-8B-MoT-pose-8bit \
  --edit-image skeleton.png --edit-image identity.jpg \
  --prompt $'Image-1: <image>\nImage-2: <image>\napply pose from image 1 with reference from image 2' \
  --width 768 --height 768 --steps 28 --cfg 4.0 --out posed.npy

From Swift, through the MLXEngine wrapper:

let package = SenseNovaU1Package(configuration: .init(variant: .pose8))
try await package.load()
let response = try await package.run(
    IEditRequest(
        images: [skeleton, identity],            // 1 = pose, 2 = appearance
        prompt: SenseNovaVariant.posePrompt,
        width: 768, height: 768, steps: 28, seed: 42))

Performance (M5 Max)

768² two-reference edit, 28 steps, cfg 4 20.2 s (0.72 s/step)
peak memory 22.5 GB
resident after load 19.0 GB
artifact load 2.6 s

Provenance

  • Base: sensenova/SenseNova-U1.5-8B-MoT @ 07d76f61 (Apache-2.0). Hub main has since moved to 19bc874e, a bf16 re-shard of the same weights; this artifact was converted from the pinned revision.
  • Adapter: RefControl pose LoRA v1 @ 3000 steps — rank 32 / alpha 32, neo_hf_lora layout, 294 gen-stream (*_mot_gen) projections, trained on (skeleton, reference, target) triples. Merged in fp32 as W += (alpha/rank)·BA and cast back; all 294 targets matched.
  • Conversion: sensenova-cli --weights <base> --lora <adapter> --quant 8 --convert <out>.
  • The merge is exactly as wide as the adapter claims: against the plain 8-bit artifact, 882 tensors differ — 294 targets × {{weight, scales, biases}} — and nothing else. A zero-B copy of the adapter through the same path reproduces the plain artifact tensor for tensor.

Training data: Wikimedia Commons (CC-BY / CC-BY-SA) and Pexels footage. The source frames are not redistributed; these are merged weights.

Tiers

pose-8bit reproduces the pose-bf16 renders on all five valid held-out fixtures, so it is the recommended tier. pose-bf16 exists for reference-precision work.

This is a base-checkpoint merge, not a distillation — text-to-image, think mode and visual question answering all still run on these weights, though the generation stream is pose-biased. For plain generation use the base tiers.

Licence

Apache-2.0, following the base checkpoint. Port code is MIT.

Downloads last month
27
Safetensors
Model size
18B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/SenseNova-U1.5-8B-MoT-pose-8bit

Finetuned
(12)
this model