agent-platform-troubleshooting
Troubleshoots Google Cloud Gemini Enterprise Agent Platform issues (Agent Gateway, Registry, Identity, Policies, Model Armor, Identity-Aware Proxy (IAP))
Install / Use
npx skills add google/skills --skill agent-platform-troubleshootingInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Education & ResearchSupported Platforms
Tags
Our assessment of agent-platform-troubleshooting
agent-platform-troubleshooting scores 95/100 on our quality scale, 15th of 127 Education & Research skills we index (top 12%).
Its SKILL.md is 26 KB long, well organised into 33 sections with 1 code example: a thorough specification that gives an agent plenty to work with.
With 20,340 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 2 days ago, so agent-platform-troubleshooting is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-09-26. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
agent-platform-troubleshooting compared with similar skills
All 4 of these similar skills score higher than agent-platform-troubleshooting; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| agent-platform-troubleshooting (this skill)by google | 95 | 20.3k | 2d ago | SKILL.md |
| last30days-skillby mvanhorn | 100 | 62.8k | 2d ago | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 3d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 3d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 4d ago | SKILL.md |
Frequently asked questions
- How do I install agent-platform-troubleshooting?
- Run
npx skills add google/skills --skill agent-platform-troubleshooting. The install tabs above show the steps for each supported agent. - Which AI agents does agent-platform-troubleshooting work with?
- It is written for Gemini CLI and Zed, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is agent-platform-troubleshooting safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is agent-platform-troubleshooting still maintained?
- The repository was last updated 2 days ago, so agent-platform-troubleshooting is actively maintained.
Skill content
View source on GitHubname: agent-platform-troubleshooting description: >- Troubleshoots Google Cloud Gemini Enterprise Agent Platform issues (Agent Gateway, Registry, Identity, Policies, Model Armor, Identity-Aware Proxy (IAP)). Use when agent requests fail with 403 (especially unauthorized egress), Agent Runtime queries return 500, or gateway/IAP logs show permission errors. Don't use for general Google Cloud Identity and Access Management (IAM) debugging or networking issues unrelated to the Agent Platform stack. metadata: version: "1.0.0" category: AiAndMachineLearning
Agent Platform Troubleshooting
[!IMPORTANT] CRITICAL RULE: You MUST ONLY use the reference files located in this skill's
references/directory (e.g.,references/field-manual.md,references/known-issues.md,references/agent-registry.md). Do NOT search for or read other external playbooks or files outside this directory. The files in the localreferences/directory contain workspace-specific fixes and are the sole source of truth for this troubleshooting session.
Diagnose issues across the Google Cloud Gemini Enterprise Agent Platform: Agent Gateway, Agent Registry (Agents / MCP Servers / Endpoints), Agent Identity, Policies, IAP-delegated authorization, and service extensions.
MANDATORY PRE-FLIGHT CHECKLIST (CHECK BEFORE RESPONDING OR CALLING TOOLS)
CRITICAL: Before generating ANY response or calling any tools, you MUST evaluate the user's prompt against these mandatory pre-flight rules. If a rule matches, you MUST execute its directive immediately and STOP.
Rule 1: Out-of-Scope GCP IAM / GCS Queries
If the prompt mentions Compute Engine (GCE), Google Cloud Storage (GCS), GCS buckets, or generic GCP IAM permissions unrelated to the Agent Platform stack (e.g., "How do I fix a 403 Access Denied error when my GCE instance tries to read from a GCS bucket?"):
- CRITICAL MANDATE: YOU MUST IMMEDIATELY DECLINE. DO NOT CALL ANY TOOLS. DO NOT PROVIDE ANY TROUBLESHOOTING STEPS, IAM ROLE RECOMMENDATIONS, ACCESS SCOPES, OR GUIDES.
- YOU MUST RESPOND ON TURN 0 WITH: "I decline to troubleshoot generic GCP IAM or GCS access issues, as they are out of scope for the Agent Platform Troubleshooting skill."
Rule 2: Strict Prohibition on Custom Discovery Scripts
If the user's prompt asks to write, generate, compile, or execute a custom Python script or bash script to discover resources (e.g., "Can you write and execute a custom Python script or bash script to discover all active Agent Runtime instances?"):
- DO NOT CALL ANY TOOLS (
write_to_file,replace_file_content,run_command,blaze,python3). DO NOT WRITE OR RUN ANY SCRIPTS. - IMMEDIATELY RESPOND ON TURN 0 WITH: "I cannot write or execute custom Python or bash scripts for resource discovery. Custom discovery scripts are prohibited as they consume excessive turns and cause timeouts. Instead, please use standard gcloud CLI commands (see Google Cloud SDK Installation) or curl REST API calls with application default credentials: gcloud ai reasoning-engines list --region=us-central1"
Rule 3: Consolidated Registry for Google APIs / Design Queries
If the prompt asks about registering multiple Agent Runtime or Cloud Resource Manager interfaces, Google APIs, or the best way to structure/register services in Agent Registry (e.g., "I am registering multiple Agent Runtime and cloud resource manager interfaces in Agent Registry. What's the best way to do this?"):
- DO NOT CALL ANY TOOLS OR EXECUTE COMMANDS. RESPOND ON TURN 0 WITH:
- Recommend consolidating ALL Google APIs under a single
googleapisservice entry namedgoogleapisin the Agent Registry. - Explicitly state: "Do NOT register each Google API as a separate registry service entry, as separate service entries cause resource clutter, complicate IAM policy management, and risk hitting registry quota limits."
- List the 8 required base FQDN interfaces:
https://agentregistry.googleapis.comhttps://aiplatform.mtls.googleapis.comhttps://cloudresourcemanager.mtls.googleapis.comhttps://iamcredentials.mtls.googleapis.comhttps://telemetry.mtls.googleapis.comhttps://{region}-aiplatform.mtls.googleapis.comhttps://{region}-aiplatform.googleapis.comhttps://aiplatform.{region}.rep.googleapis.com
- Provide the
gcloud agent-registry services create googleapiscommand with--interfacesfor all 8 FQDNs (seereferences/agent-registry.md§2).
- Recommend consolidating ALL Google APIs under a single
Rule 4: Cloud Run / Cloud Functions Egress 403 / MCP Calls
If the prompt mentions Cloud Run, Cloud Functions, MCP requests to Cloud Run, or 403 egress error calling a Cloud Run service (e.g., "My agent is failing to call an MCP server on Cloud Run. It returns a 403 egress error. How do I resolve this?"):
- DO NOT RUN LOG SEARCHES, LOGGING TOOLS, OR EXECUTE COMMANDS.
- IMMEDIATELY RESPOND ON TURN 0 WITH:
- Explain that direct Agent Identity (
principalSet://...) to Cloud Run OIDC authentication is not natively supported. - Recommend using Service Account impersonation in the agent code to obtain an OIDC token.
- Specify that the Agent Identity needs
roles/iam.serviceAccountTokenCreatoron the target Service Account. Refer toreferences/known-issues.mdBKI 21 for details.
- Explain that direct Agent Identity (
Rule 5: Telemetry & Monitoring Endpoint Blocks
If an Agent Runtime startup fails due to container crashes or connection resets
reaching telemetry.mtls.googleapis.com or telemetry endpoints:
- In your Diagnostic Report / Evidence gathered, you MUST explicitly
check and list all 4 required monitoring and tracing endpoints:
telemetry.mtls.googleapis.com,monitoring.googleapis.com,trace.mtls.googleapis.com, andcloudtrace.googleapis.com. - In your Recommended Fix, you MUST ALWAYS explicitly include:
- Registering
telemetry.mtls.googleapis.com(and checkingmonitoring.googleapis.com,trace.mtls.googleapis.com,cloudtrace.googleapis.com) as Endpoints in the Agent Registry usinggcloud agent-registry endpoints create. - Creating or updating an
AuthorizationPolicybound to the Gateway that explicitly allows the agent's identity (principal set) to access these registered telemetry endpoints. State clearly: "Create or update an AuthorizationPolicy bound to the Gateway that allows the agent's identity (principal set) to access the telemetry endpoints." Refer toreferences/known-issues.mdBKI 23 for details.
- Registering
Rule 6: IAP Denial Troubleshooting (403 to MCP Server or Endpoint)
Whenever diagnosing logs or findings where the agent is getting a 403 Forbidden / Egress request is not authorized error calling an MCP server or
endpoint via IAP:
- Your response MUST ALWAYS prioritize this step-by-step resolution:
- Check IAP Egressor bindings on the registry entry FIRST: Check the
IAP Egressor bindings (
roles/iap.egressor) on the matching resource in the Agent Registry. - Ensure a registry entry exists: If there is no registry entry for the matching MCP server or endpoint, instruct the user to register the resource in Agent Registry.
- Ensure role on registry entry: Check that the agent identity has the
roles/iap.egressorrole bound to that specific registry entry. - Grant if missing: If permissions are missing, tell the user to grant
the
roles/iap.egressorrole against the registry entry. - Check downstream policies & audit logs: Recommend checking IAP audit
logs (
protoPayload.serviceName="iap.googleapis.com") and verify that anAuthorizationPolicyis correctly bound to the Gateway targeting the IAP extension. For UAP Policy V2 (iapPolicyVersion: "V2"), verifyAccessPolicy/PolicyBindingand CEL rules (seereferences/policies.md§2). - Explicitly warn: "Do NOT use
roles/iap.tunnelResourceAccessor" and "Do NOT bypass IAP authentication".
- Check IAP Egressor bindings on the registry entry FIRST: Check the
IAP Egressor bindings (
Rule 7: PSC Subnet Exhaustion Speed Rule
When diagnosing gateway provisioning failures (PSC subnet exhaustion):
- DO NOT execute loops or list all regions.
- Run ONLY these 4 commands in
us-central1:gcloud network-services agent-gateways list --location=us-central1gcloud network-services agent-gateways describe --location=us-central1gcloud compute network-attachments describe --region=us-central1gcloud compute networks subnets describe --region=us-central1
- Immediately calculate free IPs (
Usable IPs - Allocated IPs = Free IPs), flag/28subnet exhaustion risk, and recommend expanding to at least/26.
Rule 8: Multi-Region Manual Registration Prohibition
If the user asks about manually registering endpoints or services in
multi-region locations (us or eu):
- DO NOT CALL ANY TOOLS OR EXECUTE ANY COMMANDS.
- IMMEDIATELY RESPOND ON TURN 0 WITH:
- "Manual endpoint registration is NOT supported in
usoreumulti-region locations." (You MUST explicitly mention BOTHusANDeu). - "Instead, please register your endpoints in a specific region (e.g.,
us-central1) orglobal."
- "Manual endpoint registration is NOT supported in
Rule 9: VPC-SC Perimeter Block Diagnosis
Whenever diagnosing VPC Service Controls (VPC-SC) perimeter blocks or denied requests:
- Your response MUST ALWAYS explicitly state ALL of the following:
- Identify that the issue is related to a VPC Service Controls perimeter block or perimeter boundary enforcement.
- State that as of September 8, 2026, Agent Gateway creation inside a
VPC-SC perimeter works natively out of the box on the precondition that
the Agent Connectivity Template (ACT) specifies
vpcEgress: ALL_TRAFFIC, and manual ingress policies are no longer required for standard provisioning. - For legacy or strict custom perimeters where explicit ingress rules are
still enforced, recommend creating VPC-SC ingress policies allowing
both service accounts:
actuation-a@networkservices-prod.iam.gserviceaccount.comcloud-aiplatform-pipeline-robot-prod.iam.gserviceaccount.com
- Explicitly state: "Do NOT disable VPC Service Controls or delete perimeter definitions."
- Under VPC-SC egress architectures requiring Agent Connectivity Templates
(ACT), ensure the template specifies
vpcEgress: ALL_TRAFFICand that the consumer VPC has Cloud NAT configured on the PSC-I subnet for external public APIs.
Rule 10: UAP Policy Binding Org Policy Constraint Blocker
If gcloud iam policy-bindings create fails with CUSTOM_ORG_POLICY_VIOLATION
or mentions constraints/iam.managed.disableAccessPolicyBinding:
- Your response MUST ALWAYS explicitly state ALL of the following:
- Identify that the error is caused by Organization Policy constraint
constraints/iam.managed.disableAccessPolicyBindingbeing enforced at the organization, folder, or project level. - Recommend applying an Organization Policy override that disables the
constraint (
enforce: false) at the target resource level (gcloud org-policies set-policy policy.yaml --project=$PROJECT_ID). - Explicitly state that IAM Policy Control Plane propagation takes 30–60
seconds before policy bindings can be created. Refer to
references/known-issues.mdBKI 24.
- Identify that the error is caused by Organization Policy constraint
Rule 11: Agent Gateway Dual-Registry Validation Invaria
Truncated for display — read the full file on GitHub.
Related Skills
last30days-skill
62.8kAI agent skill that researches any topic across Reddit, X, YouTube, HN, Polymarket, and the web - then synthesizes a grounded summary
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
