Instructions to use tomjnet/SeqAtom-Coder-1.5B-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use tomjnet/SeqAtom-Coder-1.5B-Instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="tomjnet/SeqAtom-Coder-1.5B-Instruct") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("tomjnet/SeqAtom-Coder-1.5B-Instruct") model = AutoModelForCausalLM.from_pretrained("tomjnet/SeqAtom-Coder-1.5B-Instruct", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use tomjnet/SeqAtom-Coder-1.5B-Instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "tomjnet/SeqAtom-Coder-1.5B-Instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tomjnet/SeqAtom-Coder-1.5B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/tomjnet/SeqAtom-Coder-1.5B-Instruct
- SGLang
How to use tomjnet/SeqAtom-Coder-1.5B-Instruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "tomjnet/SeqAtom-Coder-1.5B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tomjnet/SeqAtom-Coder-1.5B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "tomjnet/SeqAtom-Coder-1.5B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tomjnet/SeqAtom-Coder-1.5B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use tomjnet/SeqAtom-Coder-1.5B-Instruct with Docker Model Runner:
docker model run hf.co/tomjnet/SeqAtom-Coder-1.5B-Instruct
SeqAtom-Coder-1.5B-Instruct
Experimental Seq language adaptation of Qwen/Qwen2.5-Coder-1.5B-Instruct.
This folder contains standalone merged FP16 safetensors weights and a tokenizer.
It does not require the original LoRA adapter at inference time.
Intended upload destination: tomjnet/SeqAtom-Coder-1.5B-Instruct.
Training
Only 10 training examples, 10 validation examples, and 10 test examples. Three epochs, nine optimizer steps, rank 8 LoRA on q_proj/v_proj, alpha 16, dropout 0.05, learning rate 0.0002, batch size 1, accumulation 4. NF4 double quantization and FP32 computation on a GTX 1650 (4 GB). Completion-only training; configured maximum length 512 tokens. Mean training loss: 2.813219. Validation loss by epoch: 2.634003, 2.571821, 2.543699. Validation token accuracy at epoch 3: 0.548294.
Held-out evaluation
Same original test prompts, chat template, greedy decoding, and 256-token generation budget for both models. One NF4 base model with adapter disabled or enabled; FP32 compute; one example at a time. Reference-answer scoring masks prompt tokens, includes the assistant ending markers, and is token-weighted. No test examples were used for additional training or checkpoint selection here.
| Metric | Base Qwen | SeqAtom adapter |
|---|---|---|
| Completion loss | 2.433137 | 2.272492 |
| Reference perplexity | 11.3946 | 9.7035 |
| Teacher-forced token accuracy | 0.5672 | 0.5902 |
| Exact reference matches / 10 | 0 | 0 |
| Outputs reaching token limit / 10 | 5 | 4 |
These are adapter evaluation results, not a full reevaluation of the merged model. The merge uses the original unquantized base, so outputs can differ from the NF4 evaluation. The merged model was reloaded and smoke-tested on CPU. FP32 merge maximum logit difference on the smoke prompt: 8.4638596e-06.
Observed behavior: both models answer the compiler-command question with
seqc, despite scoring zero strict full-text matches. The adapter lowers
reference-answer loss but still misses the colon in the code-repair example,
uses unrelated syntax for the sales workflow, and gives the same incorrect
backend.C explanation as the base model. More training examples and validation
against the actual Seq compiler are needed before claiming reliable Seq code.
Limitations
This is a pipeline prototype, not a validated Seq coding assistant. The test set is tiny and contains related concepts and tasks to the training set. Exact prompt overlap audit: {'train': 0, 'validation': 0}. Loss and token accuracy measure fit to reference text, not executable correctness. No generated code was executed or checked with a Seq compiler. Inspect outputs and validate syntax against your actual Seq implementation before using them.
Memory-conscious local inference
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
path = "./output/seqatom-merged" # Or the Hub ID after upload
tokenizer = AutoTokenizer.from_pretrained(path)
model = AutoModelForCausalLM.from_pretrained(
path, device_map={"": 0}, dtype=torch.float32,
attn_implementation="eager",
quantization_config=BitsAndBytesConfig(
load_in_4bit=True, bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True, bnb_4bit_compute_dtype=torch.float32,
),
).eval()
messages = [
{"role": "system", "content": "You are SeqAtom, an expert assistant for the Seq programming language."},
{"role": "user", "content": "Create a Seq program that prints Hello World."},
]
inputs = tokenizer.apply_chat_template(messages, tokenize=True,
add_generation_prompt=True, return_dict=True, return_tensors="pt").to("cuda")
with torch.inference_mode():
result = model.generate(**inputs, max_new_tokens=128, do_sample=False,
pad_token_id=tokenizer.pad_token_id, eos_token_id=tokenizer.eos_token_id)
print(tokenizer.decode(result[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Use one model and one prompt at a time on a 4 GB GPU. The tested environment is recorded in requirements.lock.txt. The FP16 artifact was merged on CPU.
Provenance
Base revision: not recorded; loaded the locally resolved main revision.
The training configuration did not pin a base or dataset revision. This run's
Transformers configuration did not expose a commit hash, so historical revision
identity is unverified. Dataset fingerprints are in evaluation_summary.json.
Base model license: Apache 2.0; original LICENSE is included.
- Downloads last month
- 141
Model tree for tomjnet/SeqAtom-Coder-1.5B-Instruct
Base model
Qwen/Qwen2.5-1.5B