YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
NAVA β image-server model staging folder (NVMe)
Joint audio-video diffusion DiT (Baidu ERNIE Research, Wan2.x fork). Staged on
fast NVMe (/home) for Phase 1 load-test iteration. Source bundle (slow HDD):
/storage/Models/robingg1/NAVA/.
Layout consumed by image-server/src/models/video/nava_core/ (vendored) +
build_nava.py (mmgp-driven loader, Phase 1).
NAVA/
βββ dit/
β βββ nava_dit_bf16.safetensors 12.59 GB 1052 tensors, 6.30B params, bf16
β (extracted from NAVA.ckpt via
β image-server/scripts/extract_nava_bf16.py;
β 'backbone.' prefix STRIPPED β bare WanAVModel)
βββ text_encoder/
β βββ models_t5_umt5-xxl-enc-fp8.safetensors 6.73 GB FP8 e4m3 quanto, Wan T5Encoder
β β flat 'blocks.N.' layout β scaled_fp8 handler
β β (REUSED from Wan2.2-S2V-14B-fp8 β same frozen
β β Wan UMT5-XXL NAVA uses; drop-in, no remap)
β βββ google/umt5-xxl/ tokenizer (spiece.model + tokenizer.json +
β config + special_tokens)
βββ vae_video/
β βββ Wan2.2_VAE.safetensors 2.82 GB Wan2.2 TI2V 48-channel video VAE, fp32
β (vid_in/out_dim=48; NOT the 16-ch Wan2.1 VAE.
β Converted from .pth via scripts/pth_to_safetensors.py,
β byte-parity verified; original .pth on /storage bundle)
βββ vae_audio/
β βββ ltx-2.3-22b-dev_audio_vae.safetensors 0.36 GB Lightricks LTX-2.3 audio VAE decoder
βββ config/
β βββ config.json WanAVModel constructor config (== NAVA_6B.json):
β β model_type=ti2v, dim 3072, ffn 14336, heads 24,
β β double 10 / single 20, vid 48 / audio 128,
β β text_len 512, qk_norm, temporal_rope_scale 0.24
β βββ nava.yaml pipeline/inference config (guidance scales,
β β shift=5, unipc, 4-pass CFG flags)
β βββ example_prompts.jsonl 5 sample prompts (prompt / spk_wavs / image_path)
Total ~21 GB.
Notes / provenance
- DiT prefix: keys are bare (
patch_embedding.weight,single_blocks.0...) β thebackbone.prefix from the training checkpoint was stripped, so this loads directly into a bareWanAVModel. Strict-load key-match againstWanAVModel.state_dict()is the Phase 1 gate. - Text encoder: FP8, Wan layout, drop-in. Routine Phase 1 sanity check = diff FP8 prompt
embeddings vs the bf16 reference. bf16 fallback (NOT copied here, lives on HDD):
/storage/Models/robingg1/NAVA/Wan2.2-TI2V-5B/models_t5_umt5-xxl-enc-bf16.pth(11.36 GB). - FP8 DiT (Phase 4): will be produced by
scripts/quantize_nava_fp8.pyand dropped intodit/nava_dit_fp8.safetensorsalongside the bf16. - Not needed: the
model_id: Qwen3-1.7Bin nava.yaml is prompt-rewriting only, not the model. - Full integration plan + recon: image_server_kernels memory
project_nava_kernel_fit.md.
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support