A100 Qwen API runtime

Reproducible OpenAI-compatible local API for a single NVIDIA A100 40 GB GPU. It uses the pinned llama.cpp b11177 CUDA 12.8 binary and two pinned GGUF models:

API alias Model Quant File size Source
qwen38 Qwen3.8-27B Q6_K 22.56 GB 6block GGUF
qwen36 Qwen3.6-35B-A3B UD-Q5_K_M 26.46 GB Unsloth GGUF

Both files take about 49.02 GB of disk space. The CUDA archives take another 0.76 GB until removed. Only one model can be resident on a 40 GB A100 at a time. Model weights, binaries, logs, benchmark outputs, and API keys are excluded from this repo. The installer uses Hugging Face Hub/Xet for chunked model downloads at pinned revisions and verifies the published SHA-256 digests.

An optional 3.16 GB Qwen3.8 MTP draft can be installed with WITH_MTP=1 bash scripts/install.sh. The two-model router leaves speculative decoding off by default.

Requirements

  • Linux x86-64 with an NVIDIA A100 40 GB and a CUDA 12.8-compatible driver.
  • bash, curl, tar, sha256sum, setsid, nohup, taskset, and Python 3.
  • The pinned prebuilt binary needs libgomp.so.1. This machine provides it in /opt/conda/lib; adjust LD_LIBRARY_PATH in serve.sh on other machines if needed.
  • Approximately 55 GB of free storage and internet access for the initial download. Python venv and pip are needed for the pinned Hub/Xet downloader.

Install and serve

git clone https://huggingface.co/ShoaibSSM/a100-qwen-api-runtime
cd a100-qwen-api-runtime
bash scripts/install.sh
scripts/start.sh

start.sh uses nohup and setsid, so closing Codex or the shell does not stop the API process while the container is running. It generates a random API key in secrets/api_keys.txt (mode 0600) if needed and binds to 127.0.0.1:8080. Both models start unloaded. A request with "model":"qwen38" or "model":"qwen36" loads it on demand; at most one model runs at once. After 15 minutes without inference work, the model sleeps and releases its memory. The next request wakes it. Switching or waking models incurs a cold-load delay. Check readiness with curl http://127.0.0.1:8080/health, logs with tail -f logs/server.log, and stop with scripts/stop.sh. Change the idle timer with LLM_IDLE_SECONDS=600 scripts/start.sh.

The router uses full GPU offload, flash attention, Q8 KV cache, 65,536 total context tokens, four parallel slots, medium reasoning effort, and 2048/512 logical/physical batch sizes. In llama.cpp b11177, the startup log confirms n_slots = 4 and n_ctx_slot = 16384, so each slot gets a 16K context. CPU pools are explicitly bounded: 16 generation threads, 16 batch threads, and 4 HTTP threads; OpenMP is capped at 4 and optional BLAS pools at 1. --threads-http matters here: its b11177 default is automatic and had grown with the 100 CPUs visible inside this container. GPU offload is forced so a model that does not fit fails visibly instead of falling back to CPU offload. Continuous batching is enabled. Set LLM_REASONING_EFFORT=xhigh for harder tasks and allow enough max_tokens for both reasoning and the final answer.

Run scripts/check_threads.sh to inspect the cgroup task limit, process thread counts, llama processes, and thread-related environment.

On this A100, four simultaneous 384-token Qwen3.8 requests completed through slots 0–3. The llama processes used at most 14 and 12 threads during the check; total cgroup tasks peaked at 215/512 (idle baseline 190). Generation logs measured about 21.8 tokens/s per active slot during that run. GPU process memory was about 24.2 GiB. These are observed values for this container, not guarantees under other workloads.

Add GLM-4.7-Flash and GPT-OSS-120B

The router automatically adds these aliases after their files are downloaded and the router is restarted:

API model GGUF Router settings
glm47q8 GLM-4.7-Flash Q8_0 2 slots, 32K total context (16K each), full GPU offload
gptoss120 GPT-OSS-120B native MXFP4 1 slot, 16K context, MoE weights kept in system RAM, remaining layers offloaded to GPU

Install the Hugging Face CLI if needed, then run from this directory. Downloads are pinned to the listed repository revisions and go into models/:

python -m pip install -U 'huggingface_hub[hf_xet]'
hf download ggml-org/GLM-4.7-Flash-GGUF GLM-4.7-Flash-Q8_0.gguf \
  --revision 7559e96b7e324ab405897dc2b91492b0f376ad4a --local-dir models

The GPT-OSS file is two shards; download both into the same directory:

hf download lmstudio-community/gpt-oss-120b-GGUF \
  gpt-oss-120b-MXFP4-00001-of-00002.gguf \
  gpt-oss-120b-MXFP4-00002-of-00002.gguf \
  --revision ffa0c82eff830f6644fa19b14ef2c0e11f7cd1e8 --local-dir models

Both commands resume through Hugging Face Hub/Xet. GLM Q8_0 is 31.84 GB; the router gives it a smaller context/slot allocation than Qwen to leave A100 VRAM headroom. A Q6_K GLM file is about 24.61 GB if you prefer a lower memory footprint; Q8_0 is already a high quality quant, so start with the requested file. GPT-OSS's MXFP4 is its native published quantization; the two shards total about 63.39 GB. The GPT-OSS preset keeps MoE weights on CPU because this model cannot fit wholly in 40 GB VRAM. Expect it to be much slower than the fully GPU-resident models.

After the downloads finish, restart the router so it discovers the files:

scripts/stop.sh
scripts/start.sh

Select them with API model values glm47q8 and gptoss120. GPT-OSS takes longer to load and has only one request slot in this initial configuration.

LLM_CONTEXT=32768 LLM_PARALLEL=2 scripts/start.sh

Use the API

API_KEY="$(cat secrets/api_keys.txt)"
curl -sS http://127.0.0.1:8080/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{"model":"qwen38","messages":[{"role":"user","content":"Explain this algorithm."}],"max_tokens":512}'

With the OpenAI Python client, set base_url="http://127.0.0.1:8080/v1", api_key to the generated key, and model to qwen38 or qwen36 per request.

Reach it through ngrok

From another shell in this directory, install the standalone ngrok Linux agent and add your own ngrok account authtoken. This token is distinct from the LLM API key.

curl -fsSL https://bin.ngrok.com/c/bNyj1mQVY4c/ngrok-v3-stable-linux-amd64.tgz | tar -xz
read -rsp 'ngrok authtoken: ' NGROK_AUTHTOKEN; echo
./ngrok config add-authtoken "$NGROK_AUTHTOKEN"
unset NGROK_AUTHTOKEN
GOMAXPROCS=2 nohup setsid ./ngrok http 8080 --log=stdout > logs/ngrok.log 2>&1 < /dev/null &
echo $! > runtime/ngrok.pid

Read the HTTPS URL with curl -s http://127.0.0.1:4040/api/tunnels | python -c 'import json,sys; print(next(t["public_url"] for t in json.load(sys.stdin)["tunnels"] if t["public_url"].startswith("https://")))'. Use that URL plus /v1/chat/completions from the other machine with the Authorization: Bearer <LLM API key> header. The API and tunnel persist after the shell closes, but both stop if the container stops. The tunnel URL may change on restart. Stop the tunnel with kill "$(cat runtime/ngrok.pid)".

On your laptop, set LLM_URL to that HTTPS URL and LLM_API_KEY to the contents of secrets/api_keys.txt from this server:

export LLM_URL='https://YOUR-TUNNEL.ngrok-free.app'
read -rsp 'LLM API key: ' LLM_API_KEY; echo
export LLM_API_KEY
curl -sS "$LLM_URL/v1/chat/completions" \
  -H "Authorization: Bearer $LLM_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{"model":"qwen38","messages":[{"role":"user","content":"Hello"}],"max_tokens":512}'

For an OpenAI-compatible client: base_url="$LLM_URL/v1", api_key="$LLM_API_KEY", and model="qwen38" or model="qwen36". The first call after sleep, or a model switch, can take longer while weights load; give the client at least a 180-second timeout.

Compare on this hardware

Run the same API timing prompts through the router if you want a local comparison:

python scripts/bench.py qwen38
python scripts/bench.py qwen36

Stop the API server before running raw throughput checks. These measure prompt processing and token generation without chat-template or reasoning-length effects:

scripts/bench-throughput.sh qwen38 > results/qwen38-throughput.json
scripts/bench-throughput.sh qwen36 > results/qwen36-throughput.json

bench.py saves answers and latency to results/. It is a smoke test and timing check, not a SWE-bench evaluation. Qwen3.8's model card reports stronger recent coding results than the older Qwen3.6 card, but the exact tasks, harnesses, and full precision versus quantized weights differ. Compare quality on your own repository and research prompts before selecting the permanent default.

Update this repo without an interactive browser login

Install the Hugging Face CLI with python -m pip install -U huggingface_hub. Create a write token on Hugging Face, then enter it without putting it in shell history:

read -rsp 'HF write token: ' HF_TOKEN; echo
export HF_TOKEN
hf auth login --token "$HF_TOKEN"
hf upload ShoaibSSM/a100-qwen-api-runtime . . --repo-type model \
  --exclude 'models/**' 'runtime/**' 'logs/**' 'results/**' 'secrets/**' '.git/**'
unset HF_TOKEN

The model repositories retain their own licenses and provenance; this repo contains deployment code only.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support