GGUF
conversational

Tmax-9B-GGUF

Direct GGUF Quantizations of Tmax-9B

This repository provides GGUF quantized models for allenai/tmax-9b.

Tmax-9B is a 9 billion parameter terminal-agent model developed by AllenAI and collaborators. Built on top of the Qwen3.5-9B architecture and further trained using reinforcement learning for terminal-based tasks, it is designed to perform complex command-line and software engineering workflows while maintaining strong general-purpose reasoning capabilities. These GGUF versions are optimized for efficient CPU and GPU inference using llama.cpp and compatible tools.

This release includes various quantization levels (e.g., Q2_K, Q3_K_M, Q4_K_M, Q5_K_M, Q6_K, Q8_0) to suit different hardware capabilities and performance requirements.

Table of Contents 📝

  1. Usage
  2. 📃 License
  3. 🙏 Acknowledgements

▶ Usage

1. Download Models

Download models using huggingface-cli:

pip install "huggingface_hub[cli]"
huggingface-cli download samuelchristlie/tmax-9B-gguf --local-dir ./tmax-9B-gguf

You can also download directly from this page

2. Inference

To use these GGUF files, you'll need a compatible inference engine like llama.cpp or clients built on top of it (e.g., Ollama, LM Studio, KoboldCpp, text-generation-webui with llama.cpp backend).

📃 License

This model is a GGUF conversion of the original allenai/tmax-9b model. The original model is licensed under the Apache 2.0 License, and this derivative work adheres to the terms of that license. Please review the original license for full details.

🙏 Acknowledgements

Downloads last month
1,905
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for samuelchristlie/tmax-9b-gguf

Finetuned
Qwen/Qwen3.5-9B
Quantized
(411)
this model