Deploying on AWS documentation

Use AWS-hosted models with agent harnesses

Hugging Face's logo
Join the Hugging Face community

and get access to the augmented documentation experience

to get started

Use AWS-hosted models with agent harnesses

Agent harnesses such as Pi, Hermes Agent, or Tau can use models hosted on Amazon SageMaker AI, SageMaker JumpStart, or Amazon Bedrock. The harness runs locally and sends model requests to an OpenAI-compatible endpoint on AWS.

This guide covers client-side agent harnesses. Amazon Bedrock Agents is a separate managed AWS service for building and orchestrating agents.

The endpoint protocol and the model capability are separate requirements:

  • The endpoint must support streaming OpenAI Chat Completions.
  • The model must reliably support tool calling. A model that can chat but cannot produce valid tool calls will not complete an agent loop.

Before you begin

For SageMaker AI and JumpStart, use a real-time endpoint backed by a container that implements /v1/chat/completions, such as the Hugging Face vLLM or SGLang DLC. The IAM identity invoking the endpoint needs:

  • sagemaker:InvokeEndpoint on the endpoint ARN.
  • sagemaker:CallWithBearerToken on "*".

For Amazon Bedrock, select a model that supports the Chat Completions API and create an Amazon Bedrock API key. Prefer short-term keys for production; long-term keys are intended for exploration.

Connect to SageMaker AI or JumpStart

SageMaker AI and JumpStart deployments use the same runtime URL. JumpStart deploys the selected model as a SageMaker AI endpoint, so the connection settings depend on the resulting endpoint, not how it was created.

The base URL for a single-model endpoint is:

https://runtime.sagemaker.<REGION>.amazonaws.com/endpoints/<ENDPOINT_NAME>/openai/v1

SageMaker AI authenticates OpenAI-compatible clients with short-lived bearer tokens. Generate one with the SageMaker Python SDK:

from datetime import timedelta
from sagemaker.core.token_generator import generate_token

token = generate_token(
    region="us-west-2",
    expiry=timedelta(hours=1),
)

Tokens are valid for up to 12 hours and cannot outlive the AWS credentials used to create them. Generate them at the point of use instead of saving them in a configuration file.

Pi

Pi can resolve an API key by running a shell command. Save this script as ~/.pi/agent/sagemaker-token.py so Pi can request a fresh token when it connects:

from sagemaker.core.token_generator import generate_token

print(generate_token(region="us-west-2"))

Add the provider to ~/.pi/agent/models.json:

{
  "providers": {
    "sagemaker": {
      "baseUrl": "https://runtime.sagemaker.us-west-2.amazonaws.com/endpoints/my-endpoint/openai/v1",
      "api": "openai-completions",
      "apiKey": "!python ~/.pi/agent/sagemaker-token.py",
      "models": [
        {
          "id": "Qwen/Qwen3.8-27B",
          "name": "Qwen3.8 27B on SageMaker",
          "reasoning": false,
          "input": ["text"],
          "contextWindow": 32768,
          "maxTokens": 4096,
          "compat": {
            "supportsDeveloperRole": false,
            "supportsReasoningEffort": false
          }
        }
      ]
    }
  }
}

Replace the Region, endpoint name, model ID, and model limits with your deployment values. Open Pi’s model picker and select the SageMaker model. See Pi custom models for the complete provider schema.

Hermes Agent

Hermes supports OpenAI-compatible endpoints through its openai-api provider. Export a fresh token in the shell that starts Hermes:

export OPENAI_BASE_URL="https://runtime.sagemaker.us-west-2.amazonaws.com/endpoints/my-endpoint/openai/v1"
export OPENAI_API_KEY="$(python -c 'from sagemaker.core.token_generator import generate_token; print(generate_token(region=\"us-west-2\"))')"

Then configure the model in ~/.hermes/config.yaml:

model:
  default: Qwen/Qwen3.8-27B
  provider: openai-api

The environment variable keeps the short-lived token out of the configuration file. Generate a new token when it expires. See Hermes AI providers for other configuration methods.

Tau

Tau supports OpenAI-compatible endpoints through custom providers. Because SageMaker bearer tokens are short-lived, pass the token through an environment variable generated in the shell that starts Tau:

export SAGEMAKER_API_KEY="$(python -c 'from sagemaker.core.token_generator import generate_token; print(generate_token(region=\"us-west-2\"))')"

Add the provider to ~/.tau/catalog.toml:

schema_version = 1

[[providers]]
name = "sagemaker"
display_name = "Amazon SageMaker"
kind = "openai-compatible"
base_url = "https://runtime.sagemaker.us-west-2.amazonaws.com/endpoints/my-endpoint/openai/v1"
api_key_env = "SAGEMAKER_API_KEY"
credential_name = "sagemaker"
models = ["Qwen/Qwen3.8-27B"]
default_model = "Qwen/Qwen3.8-27B"

[providers.context_windows]
"Qwen/Qwen3.8-27B" = 32768

Replace the Region, endpoint name, model ID, and context window with your deployment values, then start Tau with the provider:

tau --provider sagemaker

The environment variable keeps the short-lived token out of the configuration file. Generate a new token when it expires. You can also run /login custom inside Tau for an interactive setup. See Tau providers and models for the complete provider schema.

Connect to Amazon Bedrock

Amazon Bedrock exposes an OpenAI-compatible endpoint at:

https://bedrock-runtime.<REGION>.amazonaws.com/openai/v1

Set your Bedrock API key as an environment variable:

export AWS_BEARER_TOKEN_BEDROCK="<BEDROCK_API_KEY>"

Pi

Add a Bedrock provider to ~/.pi/agent/models.json:

{
  "providers": {
    "bedrock": {
      "baseUrl": "https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1",
      "api": "openai-completions",
      "apiKey": "$AWS_BEARER_TOKEN_BEDROCK",
      "models": [
        {
          "id": "openai.gpt-oss-120b",
          "name": "GPT OSS 120B on Bedrock",
          "reasoning": true,
          "input": ["text"],
          "contextWindow": 131072,
          "maxTokens": 8192
        }
      ]
    }
  }
}

Use a model ID and limits supported by Bedrock in your Region.

Hermes Agent

Point the OpenAI-compatible provider at Bedrock:

export OPENAI_BASE_URL="https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1"
export OPENAI_API_KEY="$AWS_BEARER_TOKEN_BEDROCK"

Configure the selected model in ~/.hermes/config.yaml:

model:
  default: openai.gpt-oss-120b
  provider: openai-api

Tau

Add a Bedrock provider to ~/.tau/catalog.toml:

schema_version = 1

[[providers]]
name = "bedrock"
display_name = "Amazon Bedrock"
kind = "openai-compatible"
base_url = "https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1"
api_key_env = "AWS_BEARER_TOKEN_BEDROCK"
credential_name = "bedrock"
models = ["openai.gpt-oss-120b"]
default_model = "openai.gpt-oss-120b"

[providers.context_windows]
"openai.gpt-oss-120b" = 131072

Then start Tau with the provider:

tau --provider bedrock

Use a model ID and limits supported by Bedrock in your Region.

Use another agent harness

For any harness that supports OpenAI Chat Completions, provide the same three values:

  1. The SageMaker AI or Bedrock base URL.
  2. A SageMaker bearer token or Bedrock API key.
  3. The deployed model ID.

Confirm that the client supports streaming responses and tool calls. OpenAI-compatible text generation alone is not sufficient for an agent workflow.

Troubleshooting

  • 401 or 403 response: refresh an expired token, verify that the token and endpoint use the same Region, and check the invocation permissions.
  • 404 response: use the base URL ending in /openai/v1; do not append /chat/completions when the harness adds it automatically.
  • The model answers but never calls tools: verify that the model supports tool calling and that the serving container has tool calling enabled. For vLLM, configure the appropriate tool-call parser for the model.
  • Malformed tool calls: reduce the tool schema complexity, verify the model’s expected chat template, and test a single tool before running a larger agent workflow.
  • Streaming errors: confirm that the container returns Chat Completions streams as server-sent events.

Examples

Update on GitHub