SeqAtom-Coder-1.5B-Instruct

Experimental Seq language adaptation of Qwen/Qwen2.5-Coder-1.5B-Instruct. This folder contains standalone merged FP16 safetensors weights and a tokenizer. It does not require the original LoRA adapter at inference time. Intended upload destination: tomjnet/SeqAtom-Coder-1.5B-Instruct.

Training

Only 10 training examples, 10 validation examples, and 10 test examples. Three epochs, nine optimizer steps, rank 8 LoRA on q_proj/v_proj, alpha 16, dropout 0.05, learning rate 0.0002, batch size 1, accumulation 4. NF4 double quantization and FP32 computation on a GTX 1650 (4 GB). Completion-only training; configured maximum length 512 tokens. Mean training loss: 2.813219. Validation loss by epoch: 2.634003, 2.571821, 2.543699. Validation token accuracy at epoch 3: 0.548294.

Held-out evaluation

Same original test prompts, chat template, greedy decoding, and 256-token generation budget for both models. One NF4 base model with adapter disabled or enabled; FP32 compute; one example at a time. Reference-answer scoring masks prompt tokens, includes the assistant ending markers, and is token-weighted. No test examples were used for additional training or checkpoint selection here.

Metric Base Qwen SeqAtom adapter
Completion loss 2.433137 2.272492
Reference perplexity 11.3946 9.7035
Teacher-forced token accuracy 0.5672 0.5902
Exact reference matches / 10 0 0
Outputs reaching token limit / 10 5 4

These are adapter evaluation results, not a full reevaluation of the merged model. The merge uses the original unquantized base, so outputs can differ from the NF4 evaluation. The merged model was reloaded and smoke-tested on CPU. FP32 merge maximum logit difference on the smoke prompt: 8.4638596e-06.

Observed behavior: both models answer the compiler-command question with seqc, despite scoring zero strict full-text matches. The adapter lowers reference-answer loss but still misses the colon in the code-repair example, uses unrelated syntax for the sales workflow, and gives the same incorrect backend.C explanation as the base model. More training examples and validation against the actual Seq compiler are needed before claiming reliable Seq code.

Limitations

This is a pipeline prototype, not a validated Seq coding assistant. The test set is tiny and contains related concepts and tasks to the training set. Exact prompt overlap audit: {'train': 0, 'validation': 0}. Loss and token accuracy measure fit to reference text, not executable correctness. No generated code was executed or checked with a Seq compiler. Inspect outputs and validate syntax against your actual Seq implementation before using them.

Memory-conscious local inference

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig

path = "./output/seqatom-merged"  # Or the Hub ID after upload
tokenizer = AutoTokenizer.from_pretrained(path)
model = AutoModelForCausalLM.from_pretrained(
    path, device_map={"": 0}, dtype=torch.float32,
    attn_implementation="eager",
    quantization_config=BitsAndBytesConfig(
        load_in_4bit=True, bnb_4bit_quant_type="nf4",
        bnb_4bit_use_double_quant=True, bnb_4bit_compute_dtype=torch.float32,
    ),
).eval()
messages = [
    {"role": "system", "content": "You are SeqAtom, an expert assistant for the Seq programming language."},
    {"role": "user", "content": "Create a Seq program that prints Hello World."},
]
inputs = tokenizer.apply_chat_template(messages, tokenize=True,
    add_generation_prompt=True, return_dict=True, return_tensors="pt").to("cuda")
with torch.inference_mode():
    result = model.generate(**inputs, max_new_tokens=128, do_sample=False,
        pad_token_id=tokenizer.pad_token_id, eos_token_id=tokenizer.eos_token_id)
print(tokenizer.decode(result[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Use one model and one prompt at a time on a 4 GB GPU. The tested environment is recorded in requirements.lock.txt. The FP16 artifact was merged on CPU.

Provenance

Base revision: not recorded; loaded the locally resolved main revision. The training configuration did not pin a base or dataset revision. This run's Transformers configuration did not expose a commit hash, so historical revision identity is unverified. Dataset fingerprints are in evaluation_summary.json. Base model license: Apache 2.0; original LICENSE is included.

Downloads last month
141
Safetensors
Model size
2B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tomjnet/SeqAtom-Coder-1.5B-Instruct

Finetuned
(216)
this model
Quantizations
1 model

Dataset used to train tomjnet/SeqAtom-Coder-1.5B-Instruct