DAFT Qwen2.5-Coder-7B-Instruct — checkpoint 2000

Intermediate checkpoint from full supervised fine-tuning for NVIDIA GPU assembly to AMD GPU assembly translation. Training is ongoing; this is not the final selected model.

  • Base model: Qwen/Qwen2.5-Coder-7B-Instruct, revision c03e6d358207e414f1eca0bb1891e29f1db0e242.
  • Optimizer step: 2,000 of 10,219 planned steps; one epoch, all 163,495 eligible training pairs.
  • Dataset: ahmedheakl/daft-sm89-rdna-functions, revision a833e9887bb5e3623c52dc0ce57480784fd77ce6.
  • Context: 32,768 total prompt and response tokens; full fine-tuning, BF16, two H200s, ZeRO-3, effective batch 16.
  • Leakage policy: benchmarks_exact_v2, excluding only exact whitespace-normalized source OR target benchmark matches from training.
  • Validation: target cross-entropy on 281 eligible CASS pairs and 337 eligible Rodinia pairs. These suites are used for checkpoint selection, not independent final testing. Functional correctness has not been established.
  • Export includes model weights and tokenizer. DeepSpeed optimizer/RNG state remains in the local training checkpoint and is not included in this model repository.
  • Training code: https://github.com/aasim-m/DAFT-experiment-setup
  • W&B: https://wandb.ai/daft/daft-asplos/runs/9fd9f1100493

Input format

Use the tokenizer's chat template with a user message containing:

Translate the following NVIDIA GPU assembly function into corresponding AMD GPU assembly. Preserve the function's behavior. Return only the translated assembly, without explanations or Markdown fences.
<your NVIDIA assembly function>

Validation losses at step 2000

  • eval_qwen25_cass_validation_loss: 0.08126474171876907
  • eval_qwen25_rodinia_validation_loss: 0.16459216177463531
Downloads last month
259
Safetensors
Model size
333k params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for aasim-m/daft-qwen2.5-coder-7b-instruct-checkpoint-2000

Base model

Qwen/Qwen2.5-7B
Finetuned
(452)
this model