A128-ours

Whole-epoch backups of the last16 in-target drafter for Qwen3.6-27B, trained with128 uniformly sampled anchors per sequence, rank64, split memory. Available epochs: 1, 2. Planned epochs:1–6.

Each epoch-NN/ contains adapters.pt, its native config/provenance and SHA256 checksums. Epoch2 onward also include trainer.pt with float32 masters, AdamW state, data RNG and torch RNG for training recovery. Epoch1 was imported as weights only: its original optimizer/master state was not uploaded and is unavailable. Continuation starts with a fresh optimizer from the epoch1 weights, then preserves optimizer state through epoch6. Its learning rate decays continuously from approximately1.66e-4 to3e-5.

These are custom specdec.fused_adapter.v1 drafter checkpoints, not a standalone copy of the27B target. Adapter inference uses frozen project code commit41a449da2aa50b674a615c45fb8bc7cda9de8466: https://github.com/ms-choi-126/spec_dec_naacl Training recovery additionally needs the continuous-schedule wrapper and config: https://github.com/ms-choi-126/spec_dec_naacl/tree/main/docs/handoff/a128_continuous6_20261008 Use that handoff's train.py with --resume; the wrapper registers the explicit continuation_lr field while preserving native resume checks. The target weights and training corpus are not included.

Only integer epochs are published. The watcher verifies remote file hashes before marking an epoch backed up; staged hard links protect against pruning.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for CHOI0126/A128-ours

Base model

Qwen/Qwen3.6-27B
Adapter
(554)
this model