Deploying on AWS documentation

Use Agent Skills with SageMaker AI

Hugging Face's logo
Join the Hugging Face community

and get access to the augmented documentation experience

to get started

Use Agent Skills with SageMaker AI

Hugging Face Agent Skills give coding agents reusable instructions, scripts, and safeguards for AI and ML workflows. The hf-cloud-* skills guide an agent through a SageMaker AI deployment, from discovering your AWS context to selecting a container, creating an endpoint, validating it, and cleaning it up.

These skills control SageMaker resources. They are independent of the model that powers your agent, so you can use them whether the agent itself runs on a local model, a SageMaker endpoint, Amazon Bedrock, or another provider.

The hf-cli skill covers Hugging Face Hub operations such as finding models and managing repositories. The hf-cloud-* skills cover the AWS deployment lifecycle. Install both when an agent needs to select a model from the Hub and deploy it to SageMaker AI.

Install the skills

Install the Hugging Face CLI if it is not already available:

curl -LsSf https://hf.co/cli/install.sh | bash

Install the SageMaker skills:

hf skills add hf-cloud-aws-context-discovery
hf skills add hf-cloud-python-env-setup
hf skills add hf-cloud-sagemaker-deployment-planner
hf skills add hf-cloud-sagemaker-iam-preflight
hf skills add hf-cloud-sagemaker-production-defaults
hf skills add hf-cloud-serving-image-selection

The CLI installs each skill with its supporting scripts, references, and templates. Run hf skills update to fetch newer versions later.

If your harness does not have a location that the CLI detects automatically, pass --dest with one of its Agent Skills directories. For example, Pi discovers .agents/skills/, Hermes Agent discovers ~/.hermes/skills/, and Tau discovers ~/.tau/skills/ and ~/.agents/skills/.

Available SageMaker skills

How the workflow runs

When you ask an agent to deploy a model, the skills work together:

  1. The deployment planner identifies the model type, traffic pattern, latency needs, and cost constraints.
  2. AWS context discovery resolves the active profile, Region, account, and caller identity.
  3. Python environment setup creates an isolated environment for the AWS tooling.
  4. IAM preflight finds and validates a SageMaker execution role.
  5. Serving image selection chooses the correct Hugging Face DLC and regional image URI.
  6. Production defaults creates the endpoint with autoscaling and alarms, runs a real invocation, checks logs, and provides the teardown command.

The planner asks for confirmation before creating paid resources. It also checks endpoint quotas before recommending an instance type when permissions allow.

Ask your agent

Once the skills are installed, describe the outcome rather than the AWS plumbing. For example:

Deploy Qwen/Qwen3.8-27B to SageMaker for an internal coding agent.
Traffic will be low but interactive, and I want to minimize idle cost.
Deploy BAAI/bge-large-en-v1.5 as a production embedding endpoint.
Use my current AWS profile and pick a cost-effective CPU instance.
Check whether the SageMaker endpoint named my-endpoint is healthy,
run a smoke test, and inspect its recent errors.
Tear down the endpoint my-endpoint and its associated autoscaling
policies and alarms.

You can also name a skill explicitly:

Use hf-cloud-serving-image-selection to choose the right container
for Qwen/Qwen3-Reranker-4B.

Review before the agent acts

An Agent Skill is guidance and executable support code, not an additional AWS permission boundary. The agent can only perform actions allowed by the AWS identity in the active shell.

Before approving a deployment, review:

  • The AWS profile, Region, and account.
  • The endpoint type and instance type.
  • The estimated idle and traffic costs.
  • The IAM execution role.
  • Whether the endpoint needs VPC, KMS, or data-capture settings that are specific to your organization.

Keep IAM permissions narrow and retain the teardown command printed at the end of a deployment. SageMaker real-time endpoints incur charges while they remain in service.

Update on GitHub