Instructions to use InternScience/Agents-A1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use InternScience/Agents-A1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="InternScience/Agents-A1") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("InternScience/Agents-A1") model = AutoModelForMultimodalLM.from_pretrained("InternScience/Agents-A1", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use InternScience/Agents-A1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "InternScience/Agents-A1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "InternScience/Agents-A1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/InternScience/Agents-A1
- SGLang
How to use InternScience/Agents-A1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "InternScience/Agents-A1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "InternScience/Agents-A1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "InternScience/Agents-A1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "InternScience/Agents-A1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use InternScience/Agents-A1 with Docker Model Runner:
docker model run hf.co/InternScience/Agents-A1
Very very good model
Please do whatever you did again once new models drop! you're on the right target to get people on the local side to adopt this model. maybe add in some SWE training as well if it doesn't hurt whatever magic is in here!
Thank you for your suggestion! We’ll optimize for SWE-related scenarios in the next version.
could you look into also making a 122ba10 version? this model is great but that is currently the "underserved" size on local.
Ironically- optimizing for pure SWE isn't gonna make the model better at coding- when coding try to optimize for reasoning as its often a force multiplier for complex tasks, and will lead to better design choices and overall better code output
here check out this model and its training methodology for reasoning, agentic use, and coding: https://huggingface.co/AlexWortega/SIQ-1-35B
Ive used it and compared SIQ to ornith - which is purely fine tuned for coding- and yielded far better results from SIQ
Here are a few other notable models trained for reasoning/coding:
Edit: https://huggingface.co/Jackrong/Qwopus3.6-27B-Coder
https://huggingface.co/Jackrong/Qwopus3.6-35B-A3B-Coder
https://huggingface.co/Jackrong/Qwopus3.6-35B-A3B-v1
https://huggingface.co/FINAL-Bench/Darwin-36B-Opus
Hope this helps! 🙏
@el4 Thanks for the thoughtful suggestion and for sharing the references. I agree that better coding ability is not just about optimizing for pure SWE benchmarks. SWE performance depends on many underlying atomic capabilities, and long-horizon reasoning is certainly one of the most important ones.
This is a key area we are actively building, and we expect to see noticeable improvements in the next version. Really appreciate the input — please stay tuned!