Video-ORA-4B-GGUF

Video-ORA-4B is a 4-billion-parameter unified video understanding model built on Qwen3.5-4B and post-trained with OraRL (Annotations as Rollouts) — an annotation-augmented, on-policy reinforcement learning method that equips a single compact model to handle seven task families with direct, task-native answers and no chain-of-thought decoding: temporal grounding, visual tracking, image/video segmentation, spatial grounding, spatial-temporal grounding, video question answering, and spatial intelligence. It retains a 262,144-token native context window from its base model and delivers strong compact-model results in matched seven-family benchmark comparisons against multimodal baselines despite skipping reasoning traces at inference, with evaluations run using direct-answer prompts (enable_thinking=False) to match its reported protocol. The model is served via vLLM (with a qwen3 reasoning parser and configurable video frame sampling) or a Transformers serving endpoint, occupies roughly 8.6 GiB in BF16 weight loading, and is intended for research on structured video/spatial perception, benchmark evaluation, and task-specific adaptation — explicitly out of scope for safety-critical decisions, identity inference, or surveillance deployment. Trained on public dataset splits with evaluation identities and media excluded from the training mixture, it is released under the Apache License 2.0, consistent with its Qwen3.5-4B base, and serves as the smaller, more deployment-friendly sibling to Video-ORA-9B in the OraRL model family.

Model: https://huggingface.co/OraRL/Video-ORA-4B

Model Files

File Name Quant Type File Size File Link
Video-ORA-4B.BF16.gguf BF16 9.7 GB Download
Video-ORA-4B.F16.gguf F16 9.7 GB Download
Video-ORA-4B.Q3_K_L.gguf Q3_K_L 2.69 GB Download
Video-ORA-4B.Q3_K_M.gguf Q3_K_M 2.54 GB Download
Video-ORA-4B.Q3_K_S.gguf Q3_K_S 2.34 GB Download
Video-ORA-4B.Q4_0.gguf Q4_0 2.9 GB Download
Video-ORA-4B.Q4_K_M.gguf Q4_K_M 3.07 GB Download
Video-ORA-4B.Q4_K_S.gguf Q4_K_S 2.92 GB Download
Video-ORA-4B.Q5_0.gguf Q5_0 3.43 GB Download
Video-ORA-4B.Q5_K_M.gguf Q5_K_M 3.51 GB Download
Video-ORA-4B.Q5_K_S.gguf Q5_K_S 3.43 GB Download
Video-ORA-4B.mmproj-bf16.gguf mmproj-bf16 676 MB Download
Video-ORA-4B.mmproj-f16.gguf mmproj-f16 676 MB Download

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Downloads last month
354
GGUF
Model size
5B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/Video-ORA-4B-GGUF

Finetuned
Qwen/Qwen3.5-4B
Quantized
(2)
this model

Collection including prithivMLmods/Video-ORA-4B-GGUF