Deploying on AWS documentation
Use Agent Skills with SageMaker AI
Use Agent Skills with SageMaker AI
Hugging Face Agent Skills give coding agents reusable instructions, scripts, and safeguards for AI and ML workflows. The hf-cloud-* skills guide an agent through a SageMaker AI deployment, from discovering your AWS context to selecting a container, creating an endpoint, validating it, and cleaning it up.
These skills control SageMaker resources. They are independent of the model that powers your agent, so you can use them whether the agent itself runs on a local model, a SageMaker endpoint, Amazon Bedrock, or another provider.
The hf-cli skill covers Hugging Face Hub operations such as finding models and managing repositories. The hf-cloud-* skills cover the AWS deployment lifecycle. Install both when an agent needs to select a model from the Hub and deploy it to SageMaker AI.
Install the skills
Install the Hugging Face CLI if it is not already available:
curl -LsSf https://hf.co/cli/install.sh | bash
Install the SageMaker skills:
hf skills add hf-cloud-aws-context-discovery hf skills add hf-cloud-python-env-setup hf skills add hf-cloud-sagemaker-deployment-planner hf skills add hf-cloud-sagemaker-iam-preflight hf skills add hf-cloud-sagemaker-production-defaults hf skills add hf-cloud-serving-image-selection
The CLI installs each skill with its supporting scripts, references, and templates. Run hf skills update to fetch newer versions later.
If your harness does not have a location that the CLI detects automatically, pass --dest with one of its Agent Skills directories. For example, Pi discovers .agents/skills/, Hermes Agent discovers ~/.hermes/skills/, and Tau discovers ~/.tau/skills/ and ~/.agents/skills/.
Available SageMaker skills
Chooses real-time, scale-to-zero, async, serverless, batch, or Bedrock based on the model and traffic.
Reads the active AWS profile, Region, account, credentials, and caller identity without guessing.
Creates an isolated Python environment with current AWS dependencies.
Finds and validates an existing SageMaker execution role before attempting to create one.
Matches the model to a current Hugging Face DLC and resolves the correct regional image URI.
Deploys with autoscaling, alarms, tags, smoke tests, and a safe teardown path.
How the workflow runs
When you ask an agent to deploy a model, the skills work together:
- The deployment planner identifies the model type, traffic pattern, latency needs, and cost constraints.
- AWS context discovery resolves the active profile, Region, account, and caller identity.
- Python environment setup creates an isolated environment for the AWS tooling.
- IAM preflight finds and validates a SageMaker execution role.
- Serving image selection chooses the correct Hugging Face DLC and regional image URI.
- Production defaults creates the endpoint with autoscaling and alarms, runs a real invocation, checks logs, and provides the teardown command.
The planner asks for confirmation before creating paid resources. It also checks endpoint quotas before recommending an instance type when permissions allow.
Ask your agent
Once the skills are installed, describe the outcome rather than the AWS plumbing. For example:
Deploy Qwen/Qwen3.8-27B to SageMaker for an internal coding agent. Traffic will be low but interactive, and I want to minimize idle cost.
Deploy BAAI/bge-large-en-v1.5 as a production embedding endpoint. Use my current AWS profile and pick a cost-effective CPU instance.
Check whether the SageMaker endpoint named my-endpoint is healthy, run a smoke test, and inspect its recent errors.
Tear down the endpoint my-endpoint and its associated autoscaling policies and alarms.
You can also name a skill explicitly:
Use hf-cloud-serving-image-selection to choose the right container for Qwen/Qwen3-Reranker-4B.
Review before the agent acts
An Agent Skill is guidance and executable support code, not an additional AWS permission boundary. The agent can only perform actions allowed by the AWS identity in the active shell.
Before approving a deployment, review:
- The AWS profile, Region, and account.
- The endpoint type and instance type.
- The estimated idle and traffic costs.
- The IAM execution role.
- Whether the endpoint needs VPC, KMS, or data-capture settings that are specific to your organization.
Keep IAM permissions narrow and retain the teardown command printed at the end of a deployment. SageMaker real-time endpoints incur charges while they remain in service.
Update on GitHub