Text Generation
Transformers
Safetensors
Japanese
English
looped_gdn
gated-deltanet
linear-attention
looped-transformer
weight-sharing
distillation
custom_code
conversational
Instructions to use summerMC/Looped-Qwen-pre with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use summerMC/Looped-Qwen-pre with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="summerMC/Looped-Qwen-pre", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("summerMC/Looped-Qwen-pre", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use summerMC/Looped-Qwen-pre with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "summerMC/Looped-Qwen-pre" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "summerMC/Looped-Qwen-pre", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/summerMC/Looped-Qwen-pre
- SGLang
How to use summerMC/Looped-Qwen-pre with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "summerMC/Looped-Qwen-pre" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "summerMC/Looped-Qwen-pre", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "summerMC/Looped-Qwen-pre" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "summerMC/Looped-Qwen-pre", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use summerMC/Looped-Qwen-pre with Docker Model Runner:
docker model run hf.co/summerMC/Looped-Qwen-pre
summerMC/Looped-Qwen-pre
j-llm/Qwen3.5-2B-SpeedX2 (全層 Gated DeltaNet) の 24 層を
空間マージして重み共有し、2 重ループにした Looped GDN (反復ステップ毎 LoRA 付き)。教師 Qwen/Qwen3.5-2B から蒸留。
for t in range(6): # 外側ループ
h = RoPE注入(h) # 位置の再注入 + ループ段埋め込み
for k in range(loops[t]): # 内側ループ (ゲート付き) loops = [4, 4, 4, 4, 4, 4]
h = h + g_k(h) * (GDN_t(h) - h) # GDN_t = 元の 4 層を平均マージした共有ブロック
使い方 (transformers)
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "summerMC/Looped-Qwen-pre"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True, dtype=torch.bfloat16).cuda().eval()
ids = tok.apply_chat_template([{"role": "user", "content": "日本の首都はどこ?"}], add_generation_prompt=True,
return_tensors="pt", return_dict=True, enable_thinking=False)["input_ids"].cuda()
out = model.generate(ids, max_new_tokens=128, do_sample=True, temperature=0.7, top_p=0.9, repetition_penalty=1.1)
print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True))
pip install -U transformers flash-linear-attention 推奨 (GDN 高速カーネル)。
ファイル
| ファイル | 内容 |
|---|---|
modeling_looped_gdn.py / configuration_looped_gdn.py |
transformers 用カスタムコード (trust_remote_code) |
looped_gdn.py |
モデル本体・空間マージ変換 (python looped_gdn.py --src j-llm/Qwen3.5-2B-SpeedX2 --groups 6) |
scripts/colab_looped_distill.py |
蒸留スクリプト (Colab に貼るだけ) |
scripts/infer_looped_gdn.py |
推論 / 対話スクリプト |
scripts/colab_ast_transplant.py |
高度な空間移植 (活性考慮 RegMean + DP 可変長分割 + 活性考慮 LoRA) |
scripts/test_looped_gdn.py |
単体テスト |
looped_config.json |
旧形式の設定 (LoopedGDNForCausalLM.load 用) |
学習
- Stage 1: グループ単位 hidden 蒸留 (相対 MSE + コサイン) / Stage 2: 教師 logits KL + hidden 損失
- データ: C4 (日本語・英語)
注意
蒸留による近似で教師より性能は低下します。ライセンスは元モデル (Qwen/Qwen3.5-2B) に従います。 元モデル: Qwen Team / 全 GDN 化: summerMC, j-llm
- Downloads last month
- 350