SkillAgentSearch skills...

global-capacity-orchestrator-on-aws

Official AWS Solutions Guidance for AI/ML and HPC at planetary scale. One API. Every Accelerator. Any Region. MIT-0 licensed. https://docs.aws.amazon.com/solutions/eks-automode-clusters-with-global-capacity-orchestrator-on-aws/

Install / Use

claude mcp add aws-solutions-library-samples -- npx -y github:aws-solutions-library-samples/global-capacity-orchestrator-on-aws

If the server publishes to npm under a different name, use that package instead — check the repo README.

About this skill
🔌

MCP Server

Model Context Protocol server

Quality Score

78/100

Supported Platforms

Claude Code
Claude Desktop

Our assessment of global-capacity-orchestrator-on-aws

global-capacity-orchestrator-on-aws scores 78/100 on our quality scale, 815th of 963 AI & Machine Learning skills we index.

Its MCP Server is 68 KB long, well organised into 33 sections with 14 code examples: long enough that it reads more like full documentation than a focused instruction file, which agents can find harder to follow.

It has 48 GitHub stars, so there is little community track record yet; judge it on its content.

Substance
21/30
Structure
20/20
Description
15/15
Adoption
7/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 2 days ago, so global-capacity-orchestrator-on-aws is actively maintained.
  • It is released under the MIT-0 license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 97/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

global-capacity-orchestrator-on-aws compared with similar skills

All 4 of these similar skills score higher than global-capacity-orchestrator-on-aws; compare them before choosing.

SkillScoreStarsUpdatedFormat
global-capacity-orchestrator-on-aws (this skill)by aws-solutions-library-samples78482d agoMCP Server
claude-memby thedotmack10097.5ktodayCLAUDE.md
Agent-Reachby Panniantong10093.0k22d agoCLAUDE.md
Understand-Anythingby Egonex-AI10085.5k1d agoCLAUDE.md
headroomby headroomlabs-ai10074.6ktodayCLAUDE.md

Frequently asked questions

How do I install global-capacity-orchestrator-on-aws?
Run claude mcp add aws-solutions-library-samples -- npx -y github:aws-solutions-library-samples/global-capacity-orchestrator-on-aws. The install tabs above show the steps for each supported agent.
Which AI agents does global-capacity-orchestrator-on-aws work with?
It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
Is global-capacity-orchestrator-on-aws safe to use?
It is MIT-0-licensed and scores 97/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is global-capacity-orchestrator-on-aws still maintained?
The repository was last updated 2 days ago, so global-capacity-orchestrator-on-aws is actively maintained.
<div align="center"> <h1>Guidance for EKS AutoMode Clusters with<br><em>Global Capacity Orchestrator</em> on AWS</h1> <p><b><i>One API. Every Accelerator. Any Region.</i></b></p> <p><b>Global Capacity Orchestrator (GCO)</b> runs accelerated workloads — LLM training and inference, batch ML, HPC — on <a href="https://docs.aws.amazon.com/eks/latest/userguide/automode.html">EKS Auto Mode</a> clusters in as many AWS Regions as you configure, behind one <a href="https://aws.amazon.com/iam/">IAM</a>-authenticated <a href="docs/API.md">API</a>, <a href="docs/CLI.md">CLI</a> and <a href="gco_mcp/README.md">MCP server</a>. It finds where NVIDIA GPU, <a href="https://aws.amazon.com/ai/machine-learning/trainium/">Trainium</a>, <a href="https://aws.amazon.com/ai/machine-learning/inferentia/">Inferentia</a> and CPU capacity actually is, places jobs there, and serves inference endpoints with automatic cross-region failover.</p> <!-- BEGIN BADGE TABLE --> <p> <a href="https://github.com/aws-solutions-library-samples/global-capacity-orchestrator-on-aws/actions/workflows/unit-tests.yml"><img src="https://github.com/aws-solutions-library-samples/global-capacity-orchestrator-on-aws/actions/workflows/unit-tests.yml/badge.svg?branch=main" alt="Unit Tests"></a> <a href="https://github.com/aws-solutions-library-samples/global-capacity-orchestrator-on-aws/actions/workflows/integration-tests.yml"><img src="https://github.com/aws-solutions-library-samples/global-capacity-orchestrator-on-aws/actions/workflows/integration-tests.yml/badge.svg?branch=main" alt="Integration Tests"></a> <a href="https://github.com/aws-solutions-library-samples/global-capacity-orchestrator-on-aws/actions/workflows/security.yml"><img src="https://github.com/aws-solutions-library-samples/global-capacity-orchestrator-on-aws/actions/workflows/security.yml/badge.svg?branch=main" alt="Security"></a> <a href="https://github.com/aws-solutions-library-samples/global-capacity-orchestrator-on-aws/actions/workflows/lint.yml"><img src="https://github.com/aws-solutions-library-samples/global-capacity-orchestrator-on-aws/actions/workflows/lint.yml/badge.svg?branch=main" alt="Linting"></a> </p> <p> <a href="https://aws-solutions-library-samples.github.io/global-capacity-orchestrator-on-aws/python-coverage/"><img src="https://img.shields.io/endpoint?url=https%3A%2F%2Faws-solutions-library-samples.github.io%2Fglobal-capacity-orchestrator-on-aws%2Fpython-coverage-badge.json" alt="Python coverage"></a> <a href="https://aws-solutions-library-samples.github.io/global-capacity-orchestrator-on-aws/bash-coverage/"><img src="https://img.shields.io/endpoint?url=https%3A%2F%2Faws-solutions-library-samples.github.io%2Fglobal-capacity-orchestrator-on-aws%2Fbash-coverage-badge.json" alt="Bash coverage"></a> <a href="https://aws-solutions-library-samples.github.io/global-capacity-orchestrator-on-aws/nodejs-coverage/"><img src="https://img.shields.io/endpoint?url=https%3A%2F%2Faws-solutions-library-samples.github.io%2Fglobal-capacity-orchestrator-on-aws%2Fnodejs-coverage-badge.json" alt="Node.js coverage"></a> </p> <p> <a href="https://github.com/aws-solutions-library-samples/global-capacity-orchestrator-on-aws/releases/latest"><img src="https://img.shields.io/github/v/release/aws-solutions-library-samples/global-capacity-orchestrator-on-aws?sort=semver&display_name=tag" alt="Latest Release"></a> <a href="https://aws-solutions-library-samples.github.io/global-capacity-orchestrator-on-aws/"><img src="https://img.shields.io/badge/docs-wiki-blue" alt="Wiki"></a> </p> <p> <a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT--0-blue" alt="License: MIT-0"></a> </p> <!-- END BADGE TABLE --> </div>

Start here. With git and a container runtime (Docker, Finch or Podman) installed, this is the whole journey from nothing to a running deployment:

git clone https://github.com/aws-solutions-library-samples/global-capacity-orchestrator-on-aws.git
cd global-capacity-orchestrator-on-aws
./scripts/setup-dev-alias.sh   # builds the dev container and installs the `gco` shell function
source ~/.zshrc                # or ~/.bashrc — the script prints which file it updated
gco stacks deploy-all -y       # stand up every region in cdk.json; `gco stacks destroy-all -y` tears it down

Or let an agent drive: gco autopilot opens a Claude Code session — gco autopilot --engine codex an OpenAI Codex one, gco autopilot --engine opencode an OpenCode one — on Amazon Bedrock with the GCO MCP server already wired in, and you ask for what you want. Get started has the details and the alternatives, the Quick Start walks through your first job, and the wiki is the short orientation site.

GCO Live Demo

A real gco session, not a mock-up: fleet-wide status with cost and policy agreement, capacity discovery, four schedulers (Volcano, Kueue, YuniKorn, Slurm) plus KEDA running the queue processor, FSx, Valkey, an Aurora Serverless v2 pgvector database, a globally replicated vector store answering a semantic query over GCO's own docs, EFS, and live LLM inference — reproducible from demo/live_demo.sh. Deploy, teardown and all three Autopilot engines are recorded under See it running.

<details> <summary><b>Table of Contents</b></summary> </details>

Why GCO?

Running GPU workloads at scale is hard. You need to find regions with available capacity, provision clusters, handle authentication, deal with failover, and persist outputs after pods terminate. GCO solves all of this with a single deployable platform.

| Challenge | Traditional Approach | With GCO | |-----------|---------------------|--------------| | GPU availability | Manually check each region | Capacity tools and auto-region workflows compare configured regions | | Node provisioning | Pre-provision or wait for scaling | EKS Auto Mode provisions on-demand | | Multi-region ops | Manage clusters separately | One platform across unlimited SDK-known Regions in one partition | | Authentication | Configure per-cluster access | IAM-based, uses existing AWS credentials | | Job outputs | Lost unless persisted | EFS/FSx and per-region S3 available to every job that mounts or writes them | | Inference serving | Deploy and manage per-region | Deploy once across selected Regions; global failover in aws | | Failover | Manual intervention required | Automatic via Global Accelerator in aws; explicit regional selection elsewhere |

When to use GCO:

  • You need to run GPU workloads (training, inference, batch processing)
  • You want to deploy inference endpoints across multiple regions with a single command
  • You want multi-region redundancy without managing multiple clusters
  • You prefer IAM authentication over kubeconfig management
  • You need job outputs to persist after completion

What it does. Spins up EKS Auto Mode clusters across any number of SDK-known CloudFormation Regions in one AWS partition. In commercial aws, Global Accelerator provides latency-aware anycast routing and automatic failover behind the global workload API; other partitions (aws-cn and aws-us-gov) use IAM-authenticated regional workload APIs while retaining the aggregate global API. Capacity tools and auto-region queue/CLI workflows select a target Region, EKS Auto Mode provisions matching nodes from the built-in system and general-purpose NodePools plus project-managed GPU x86, GPU ARM, inference, EFA, Mooncake EFA, Neuron, and CPU NodePools, and shared storage persists workload outputs. Network routing never substitutes for live GPU-capacity placement.

Why it's different. Capacity-aware placement tools and auto-region workflows, partition-aware authenticated routing, full-stack observability (CloudWatch dashboards, alarms, SNS), and a CDK app validated across the full curated configuration matrix in CI. Read Core Concepts for the ideas behind it and the Learning Path if Kubernetes is new to you.

Get started

Prerequisites

Recommended path — the dev container only needs:

  • AWS credentials configured for the AWS CLI (or a ~/.aws directory to mount in)
  • Git and a container runtime: Docker, Finch, or Podman (Colima also works). The container ships Python 3.14, Node.js 24, CDK, kubectl, the AWS CLI, and Docker CLI + Buildx at pinned versions.

Host install path (advanced) additionally needs:

  • Python 3.14+ and Node.js 24 (use .nvmrc)
  • npm 12.0.2 and the repository's locked tooling graph: run npm ci --ignore-scripts --no-audit --no-fund at the repository root; gco prefers its local node_modules/.bin/cdk over a global CLI
  • A clean Python virtual environment or pipx — GCO pins exact versions of many packages, so installing into an existing environment commonly fails with dependency-resolver errors. If you hit ResolutionImpossible, switch to the dev container instead of debugging your local environment.

Run everything from the dev container

GCO pins exact versions of a lot of Python packages (CDK, AWS SDKs, FastAPI, mypy, Ruff, etc.), and installing them on top of an existing Python environment is the most common source of "it doesn't install" reports. The dev container ships a f

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars48
CategoryAI
Updated2d ago
Forks11

Languages

Python

Trust signals

97/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

1 info
global-capacity-orchestrator-on-aws — MCP Server: Install & Safety Check | SkillAgent