YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

NAVA β€” image-server model staging folder (NVMe)

Joint audio-video diffusion DiT (Baidu ERNIE Research, Wan2.x fork). Staged on fast NVMe (/home) for Phase 1 load-test iteration. Source bundle (slow HDD): /storage/Models/robingg1/NAVA/.

Layout consumed by image-server/src/models/video/nava_core/ (vendored) + build_nava.py (mmgp-driven loader, Phase 1).

NAVA/
β”œβ”€β”€ dit/
β”‚   └── nava_dit_bf16.safetensors          12.59 GB  1052 tensors, 6.30B params, bf16
β”‚                                           (extracted from NAVA.ckpt via
β”‚                                            image-server/scripts/extract_nava_bf16.py;
β”‚                                            'backbone.' prefix STRIPPED β†’ bare WanAVModel)
β”œβ”€β”€ text_encoder/
β”‚   β”œβ”€β”€ models_t5_umt5-xxl-enc-fp8.safetensors   6.73 GB  FP8 e4m3 quanto, Wan T5Encoder
β”‚   β”‚                                            flat 'blocks.N.' layout β†’ scaled_fp8 handler
β”‚   β”‚                                            (REUSED from Wan2.2-S2V-14B-fp8 β€” same frozen
β”‚   β”‚                                             Wan UMT5-XXL NAVA uses; drop-in, no remap)
β”‚   └── google/umt5-xxl/                         tokenizer (spiece.model + tokenizer.json +
β”‚                                                config + special_tokens)
β”œβ”€β”€ vae_video/
β”‚   └── Wan2.2_VAE.safetensors               2.82 GB  Wan2.2 TI2V 48-channel video VAE, fp32
β”‚                                            (vid_in/out_dim=48; NOT the 16-ch Wan2.1 VAE.
β”‚                                             Converted from .pth via scripts/pth_to_safetensors.py,
β”‚                                             byte-parity verified; original .pth on /storage bundle)
β”œβ”€β”€ vae_audio/
β”‚   └── ltx-2.3-22b-dev_audio_vae.safetensors  0.36 GB  Lightricks LTX-2.3 audio VAE decoder
β”œβ”€β”€ config/
β”‚   β”œβ”€β”€ config.json                          WanAVModel constructor config (== NAVA_6B.json):
β”‚   β”‚                                         model_type=ti2v, dim 3072, ffn 14336, heads 24,
β”‚   β”‚                                         double 10 / single 20, vid 48 / audio 128,
β”‚   β”‚                                         text_len 512, qk_norm, temporal_rope_scale 0.24
β”‚   β”œβ”€β”€ nava.yaml                             pipeline/inference config (guidance scales,
β”‚   β”‚                                         shift=5, unipc, 4-pass CFG flags)
β”‚   └── example_prompts.jsonl                 5 sample prompts (prompt / spk_wavs / image_path)

Total ~21 GB.

Notes / provenance

  • DiT prefix: keys are bare (patch_embedding.weight, single_blocks.0...) β€” the backbone. prefix from the training checkpoint was stripped, so this loads directly into a bare WanAVModel. Strict-load key-match against WanAVModel.state_dict() is the Phase 1 gate.
  • Text encoder: FP8, Wan layout, drop-in. Routine Phase 1 sanity check = diff FP8 prompt embeddings vs the bf16 reference. bf16 fallback (NOT copied here, lives on HDD): /storage/Models/robingg1/NAVA/Wan2.2-TI2V-5B/models_t5_umt5-xxl-enc-bf16.pth (11.36 GB).
  • FP8 DiT (Phase 4): will be produced by scripts/quantize_nava_fp8.py and dropped into dit/nava_dit_fp8.safetensors alongside the bf16.
  • Not needed: the model_id: Qwen3-1.7B in nava.yaml is prompt-rewriting only, not the model.
  • Full integration plan + recon: image_server_kernels memory project_nava_kernel_fit.md.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support