Instructions to use migtissera/Tess-4-27B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use migtissera/Tess-4-27B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="migtissera/Tess-4-27B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("migtissera/Tess-4-27B") model = AutoModelForMultimodalLM.from_pretrained("migtissera/Tess-4-27B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use migtissera/Tess-4-27B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "migtissera/Tess-4-27B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "migtissera/Tess-4-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/migtissera/Tess-4-27B
- SGLang
How to use migtissera/Tess-4-27B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "migtissera/Tess-4-27B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "migtissera/Tess-4-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "migtissera/Tess-4-27B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "migtissera/Tess-4-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use migtissera/Tess-4-27B with Docker Model Runner:
docker model run hf.co/migtissera/Tess-4-27B
Complements To Your Work
Tess, does indeed, exhibit improved depth of knowledge, slightly better adherence to instructions, and frankly better task when doing front end tasks.
I've been running some very long document topic/relationship extraction on it as a head to head against the base 27B and Tess routinely picks topics that are more aligned. It gets fed a set of 600 topics and very long documents, some over 300k tokens (I run 1/2 mil context length), and is outperforming 27B, with thusfar zero incorrectly formatted JSON output.
I frequently task agentic chat with webpage generation as outputs, things like... get me the news of the day return in a web page, get the public sentiment about GPT 5.6 dropping and build me a web page presenting it all, etc. I use it as a test of tool adherence, tool calling, judgement, and front end design preference, by leaving it open ended on purpose, the model's trained behavior biases surface.
Tess does better on these tasks than base 27B, by a longshot. The gap from base 27B to Tess is like the gap from 35B to 27B from what I can find.
Another task, long running many turn, 150+ tool call programming task in the strictest possible environment for code quality, standards etc. Most models cannot complete it at all, getting stuck in file read/think/edit/fail tests/lint loops. 27B can complete it about 1/2 the time. Tess has nailed it each time, and, more importantly, the output it gives is closer to the frontier models in quality.
So, well done, Tess is the first 27B I have found that is not inherently broken in some deep and subtle way making it unusable.
Thanks man, appreciate the feedback! I'm just getting started again (took a two year break to touch some grass, lol) -- expect a lot more coming. Just find me on Twitter.
