Omida R1 (v0.1 weights, research package notes)

Omida R1 is an open-weight research model for operator-style technical work: reading a local workspace, running commands, repairing a small program, and stating a conclusion only when the evidence supports it. The R-line is the public research track. Its stated role is to develop data and methods that later private flagship models can build on. This repository is not that private model, and nothing here measures one.

The weights in this repo are the v0.1 merged checkpoint. This documentation update does not replace them.

Lineage

Omida R1 v0.1 is a QLoRA supervised fine-tune of Qwen/Qwen3.5-9B, pinned at commit c202236235762e1c871ad0ccb60c8ee5ba337b9a. That revision was still the model repo head when checked on 2026-10-03. The model was not trained from scratch.

The upstream Qwen card describes Qwen3.5-9B as a post-trained model whose own base is Qwen/Qwen3.5-9B-Base. Omida's published base is the post-trained 9B checkpoint, not that base.

From the original model card, which remains the source for these training facts:

  • The BF16 tree is the Qwen checkpoint with the Omida adapter merged into the text layers.
  • Vision layers were frozen during training.
  • The corpus was 32 Omida-authored synthetic demonstrations: 25 train, 5 dev, 2 internal test.
  • The run used 2 epochs, rank 8, alpha 16, NF4 QLoRA, and learning rate 8e-5.
  • Adapter run id in omida_merge_metadata.json: r1-sft-001.
  • Merge description: BF16 safetensors, LoRA delta B @ A * alpha/r, 248 merged targets, 9,653,104,368 parameters.

config.json in this repo identifies the architecture as Qwen3_5ForConditionalGeneration (qwen3_5). The parameter total reported by the Hugging Face model API matches the merge metadata. The weight files were not re-hashed for this documentation update.

Intended use and demonstrated capability

Intended use is research on operator behavior: software investigation, repair, and evidence-based conclusions, including the choice to say that the evidence is insufficient.

Demonstrated capability is not established here. The original card points at reports/EVAL_REPORT.md for measured results. That file is not in this repository, not in the connected project tree that was available on 2026-10-03, and not in the public GitHub repositories of gh0st359 that were scanned. It is unavailable. No scores were reconstructed. v0.1 was trained on 32 synthetic demonstrations, which is a very small corpus and is not evidence of broad superiority to Qwen.

Files that the original card named but that are not in this repo

These paths were in the first model card and are not in the public file list at revision db78eadc2adb7280ce47c3c635804ecc7a46024d:

  • models/omida-r1-adapter/
  • inference/run_omida.py
  • inference/default_config.json
  • reports/EVAL_REPORT.md

The 32 demonstration texts were not published with the weights. They are not part of the new corpus.

Files that are present and should be used as the lineage record:

Research package

The v0.1 weights in this repository are unchanged. Benchmark v0.3 and corpus v0.3 are new research artifacts. They do not change the checkpoint.

The executable package is in research/ in this same repository: generators, verifiers, corpus builder, JSON Schema, training export, evaluation runner, and tests. Install and commands are in that README. There is no separate public Git remote for this package.

Versioned outputs:

Concise docs:

v0.2 had 22 families, 44 templates, and 132 instances. v0.3 keeps those smoke tasks and adds 12 demanding scenarios, including a family-held-out split. Counts after the CPU run are in the benchmark summary. None of this is a model score, and none of it is a private holdout.

Installation and inference

The examples below are adapted from the Qwen3.5-9B model card retrieved on 2026-10-03. The model id is changed to this merged checkpoint. They were not executed in the research-package environment. The weights were not downloaded or loaded.

Qwen3.5 thinks by default. The card says /think and /nothink are not the supported switch. Non-thinking mode is chat_template_kwargs.enable_thinking=False on an OpenAI-compatible server. For vLLM and SGLang tool calling, that card names the qwen3_coder tool-call parser. This merged repo ships the Qwen3.5 chat template; the parser name was not re-checked against a live server.

Coding-oriented thinking-mode sampling from that card: temperature=0.6, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0. Framework support for those fields varies.

# Not run here. Official Qwen card: current vLLM from the nightly index.
uv pip install vllm --torch-backend=auto --extra-index-url https://wheels.vllm.ai/nightly

vllm serve gh0st359/omida-r1 \
  --port 8000 \
  --tensor-parallel-size 1 \
  --max-model-len 262144 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder

The Qwen card also shows a text-only server flag, --language-model-only, which skips the vision encoder. Omida's card says vision layers were frozen, and the merged tree still contains the vision config. Whether --language-model-only is appropriate for this merged tree was not tested.

pip install -U openai
export OPENAI_BASE_URL="http://127.0.0.1:8000/v1"
export OPENAI_API_KEY="EMPTY"
from openai import OpenAI

client = OpenAI()
response = client.chat.completions.create(
    model="gh0st359/omida-r1",
    messages=[{"role": "user", "content": "Summarize the files in the current directory, then stop."}],
    max_tokens=4096,
    temperature=0.6,
    top_p=0.95,
    presence_penalty=0.0,
    extra_body={"top_k": 20},
)
print(response)

max_tokens=4096 is an evaluation budget chosen for the research runner, not a value from the Qwen card. The card's own examples use much larger limits.

Transformers serving, also from that card and also not run here:

pip install "transformers[serving] @ git+https://github.com/huggingface/transformers.git@main"
transformers serve --force-model gh0st359/omida-r1 --port 8000 --continuous-batching

The published chat template asks for tool calls in this shape, and only before any trailing text:

<tool_call>
<function=example_function_name>
<parameter=example_parameter_1>
value_1
</parameter>
</function>
</tool_call>

The research package parses that markup and was tested with fixture strings and with canned reference tool calls. That test did not run this model.

What this documentation update did not do

  • It did not modify weight shards, the tokenizer, config.json, LICENSE, or omida_merge_metadata.json.
  • It did not run inference, quantization, a merge, or a training job.
  • It did not fill in the missing evaluation report.

The public file list and revision checked before writing these notes: db78eadc2adb7280ce47c3c635804ecc7a46024d (2026-10-03).

Downloads last month
35
Safetensors
Model size
10B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for omida-ai/omida-r1

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(999)
this model
Quantizations
1 model