gke-alert-configuration
Configures alerting policies in Terraform for Google Kubernetes Engine (GKE) clusters, workloads, and services using PromQL and Google Cloud Managed Service for Prometheus
Install / Use
npx skills add google/skills --skill gke-alert-configurationInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
OperationsSupported Platforms
Our assessment of gke-alert-configuration
gke-alert-configuration scores 97/100 on our quality scale, 22nd of 259 Operations skills we index (top 9%).
Its SKILL.md is 21 KB long, well organised into 10 sections with 2 code examples: a thorough specification that gives an agent plenty to work with.
With 20,340 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 2 days ago, so gke-alert-configuration is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-09-26. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
gke-alert-configuration compared with similar skills
All 4 of these similar skills score higher than gke-alert-configuration; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| gke-alert-configuration (this skill)by google | 97 | 20.3k | 2d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 85.4k | 10d ago | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 3d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 3d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 4d ago | SKILL.md |
Frequently asked questions
- How do I install gke-alert-configuration?
- Run
npx skills add google/skills --skill gke-alert-configuration. The install tabs above show the steps for each supported agent. - Which AI agents does gke-alert-configuration work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is gke-alert-configuration safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is gke-alert-configuration still maintained?
- The repository was last updated 2 days ago, so gke-alert-configuration is actively maintained.
Skill content
View source on GitHubname: gke-alert-configuration metadata: version: "1.0.0" category: CloudInfrastructure canonical_source: https://github.com/google/skills/tree/main/skills/cloud/gke-alert-configuration description: >- Configures alerting policies in Terraform for Google Kubernetes Engine (GKE) clusters, workloads, and services using PromQL and Google Cloud Managed Service for Prometheus. Use when writing, analyzing, validating, or deploying Terraform alerting policies to monitor GKE service latency, traffic, error rates using Multi-Window Multi-Burn-Rate SLO alerts, memory saturation, and cluster health such as CrashLoopBackOff and Node NotReady conditions. Don't use for non-GKE compute runtimes such as standalone Compute Engine VMs or standalone Cloud Run services without GKE.
GKE Alert Configuration
This skill provides guidelines and best practices for creating robust, high-signal alerting policies for Google Kubernetes Engine workloads using Google Cloud Managed Service for Prometheus and Terraform. It ensures comprehensive coverage of the 4 Golden Signals and key cluster health metrics while minimizing alert noise.
Critical Rules
-
Negative Triggers and Scope Redirection for Non-GKE Standalone Runtimes:
- This skill is strictly scoped to Google Kubernetes Engine (GKE) workloads, clusters, and services using PromQL and Google Cloud Managed Service for Prometheus.
- Do not use for non-GKE compute runtimes, such as standalone Compute Engine virtual machines or standalone Cloud Run services without GKE.
- STOP AND RESPOND DIRECTLY (Do Not Edit Files): When the user
requests alert configuration for non-GKE compute infrastructure:
- Do not write, create, edit, or validate any Terraform files on disk.
- Immediately stop and respond directly to the user in chat:
- Explicitly Clarify Out-of-Scope: State clearly that standalone Compute Engine virtual machine monitoring or standalone Cloud Run monitoring is out of scope for this GKE-specific PromQL alerting skill, which is designed specifically for GKE workloads using Google Cloud Managed Service for Prometheus and PromQL.
- Do Not Generate GKE PromQL Alerts: Do not create or generate Kubernetes PromQL alert policies or fabricate Kubernetes container, pod, or node resources for non-GKE infrastructure.
- Redirect the User: Guide and redirect the user to standard
Google Cloud Monitoring metrics, such as
compute.googleapis.com/instance/cpu/utilizationorrun.googleapis.com/request_latencies, using standardgoogle_monitoring_alert_policywithcondition_thresholdor MQL, or recommend the relevant specialized Cloud observability skill.
-
Mandatory
kube-state-metrics(KSM) Cost Guardrail:- Deploying open-source
kube-state-metricsin Google Cloud Managed Service for Prometheus incurs billable metric ingestion costs. - STOP AND ASK PERMISSION FIRST (Do Not Edit Files): When a requested
alert rule relies on Tier 2 KSM metrics (such as
kube_cronjob_*,kube_pod_status_phase,kube_persistentvolume_*,kube_deployment_*,kube_statefulset_*,kube_job_*, orkube_daemonset_*), do not write, create, edit, or validate any Terraform files or generate alert policies before obtaining user approval. - Instead, you must immediately stop and respond directly to the user
to:
- Alert the user that the requested alert requires
kube-state-metrics. - Explain the cost impact: Detail that
kube-state-metricsincurs billable sample ingestion costs in Google Cloud Managed Service for Prometheus. - Ask for explicit permission: Ask the user for explicit permission before assuming, enabling, or generating KSM-dependent alert configurations.
- Recommend filtering or allowlisting: Suggest and recommend
filtering or allowlisting only the specific required metrics, such
as using a
PodMonitoringresource withmetricRelabeling(action: keep) or KSM--metric-allowlistto minimize ingestion costs. Provide a concrete allowlist example.
- Alert the user that the requested alert requires
- Always prefer Non-KSM Native Alternatives (Tier 1 cAdvisor or native
GKE metrics documented in
metrics_and_alerts_catalog.md)
whenever possible, such as using
container_memory_working_set_bytesandcontainer_spec_memory_limit_bytesinstead ofkube_pod_container_resource_limits. - Explicit Tier and Cost Surcharge Identification in Response: In
every response where you generate or recommend an alerting policy, you
must explicitly state its classification tier and cost impact:
- Tier 1 native or standard metric (GKE built-in metrics, cAdvisor
container_*, kubelet volume stats, kubelet node conditions, and control-plane metrics; see metrics_and_alerts_catalog.md): State that it is a Tier 1 native or standard metric with zero KSM cost surcharge. - Tier 2 KSM metric: State that it is a Tier 2 KSM-dependent
metric and follow the permission and allowlisting guardrail above.
(Tip: Generally, metrics with the
kube_prefix that represent resource state or metadata belong to Tier 2).
- Tier 1 native or standard metric (GKE built-in metrics, cAdvisor
- Deploying open-source
-
Plan-Validate-Execute Loop for Approved File Edits: When modifying, adding, or merging approved Terraform files on disk in a workspace, follow the three-phase workflow:
- Plan: Draft a structured change plan (
changes.json) containing proposed policy resource names, PromQL expressions, grouping labels, and durations. - Validate: Run the pre-edit validation script (
python3 scripts/validate_config.py --plan changes.json) to verify PromQL grammar, lookback windows, duration rules, and ensure no duplicate signals exist. - Execute: After the plan passes validation, apply or merge changes
in-place into the target Terraform configuration (
alerts.tf). - Note: When answering questions or providing Terraform snippets directly in chat where no disk modification is requested, output the complete, valid Terraform HCL block in your response.
- Plan: Draft a structured change plan (
-
Configure the 4 Golden Signals and Cluster Health: Always ensure the target Kubernetes workload or service has the following alerting coverage:
- Latency (P95 response time)
- Errors (Multi-Window Multi-Burn-Rate SLO alerts, such as Fast Burn 1 hour / 5 minutes with factor 14.4, Slow Burn 6 hours / 30 minutes with factor 6.0; do not use simple static ratios)
- Traffic (Sudden drop or complete metric disappearance using
absent()ordefault 0syntax, or overload spikes) - Saturation (Memory Limit Utilization Only): When describing or
configuring alert policies for a cluster or project, include ONLY
Memory Saturation (
container_memory_working_set_bytes/container_spec_memory_limit_bytes). Do NOT include CPU saturation alerts or listcontainer_cpu_usage_seconds_totalas an alert metric because CPU is compressible and throttled by CFS quotas rather than causing uncompressible fatal termination (OOM). - Cluster Health (Pod CrashLooping, Node NotReady)
-
PromQL Only (Managed Prometheus): You must use
condition_prometheus_query_languagewith PromQL. Do NOT use MQL or standardcondition_thresholdunless explicitly requested. Google Cloud Managed Service for Prometheus is the standard telemetry ingestion path for GKE. -
Terraform Only: Write the generated observability configuration ONLY as Terraform (
.tf) files, such asalerts.tfandvariables.tf. -
Dynamic Multi-Resource Alerting (No Hardcoding): You must not hardcode specific pod names, node names, or service names in alerting conditions unless explicitly requested. Alerting policies must be written to cover resources dynamically:
- Always use grouping aggregations (
by (cluster, namespace, service, pod, container)) instead of filtering to a single instance. This allows a single alert policy to dynamically track each service or pod separately. - Always declare and use Terraform variables for
project_id,cluster_name, andnamespace(var.project_id,var.cluster_name,var.namespace) to make the configuration reusable across environments. Always define these variables invariables.tf(or within the configuration) and reference all three in policies or PromQL label matchers.
- Always use grouping aggregations (
-
No Redundant Duration Windows on Lookbacks:
- When PromQL expressions already use an aggregated lookback window (such
as
increase(...[15m]) > 3or multi-window SLO burn rates), the query time window already smooths out transient spikes. - Adding a Terraform duration on top of a PromQL lookback window increases the Mean Time to Detect (MTTD) without providing additional smoothing benefits.
- In these cases, set Terraform
duration = "0s"(or"60s"). Do not enforceduration = "300s"on top of[15m], which delays critical crashloop alerts by up to 20 minutes total (15 minutes + 5 minutes). - Use
duration = "300s"only on instantaneous gauge conditions, such askube_node_status_condition == 0.
- When PromQL expressions already use an aggregated lookback window (such
as
-
Use SLO Burn Rates Instead of Simple Ratios: For error rate alerting, always generate Multi-Window Multi-Burn-Rate (MWMBR) SLO alerts (such as 14.4x burn rate over 1 hour and 5 minute windows for a 99% SLO) rather than simple error rate ratios (
rate(5xx)/rate(total) > 0.05), which produce excessive false alarms on low traffic. -
Robust Traffic Drop Detection (
absent()/default 0): When monitoring for traffic drops to zero, do not userate(...) == 0alone because Prometheus time series disappear completely when no requests occur (evaluating to an empty vector rather than 0). Usedefault 0syntax, such assum(rate(...[5m])) default 0 == 0, orabsent(...) == 1. -
Notification Channels: By default, never configure any notification channels without user input. If the user explicitly provides a notification channel, configure the alerts to use it. Otherwise, you must prompt the user in your response to ask if they would like to configure one.
-
Consult GKE Metrics and Open-Source Alerts Catalog: When designing or generating evaluation suites or alerting policies, consult metrics_and_alerts_catalog.md for public GKE metrics (
kubernetes.io/) and open-source Kubernetes alerts (awesome-prometheus-alerts). -
Plain English Response: You must include a plain English explanation for what the alerts do in your response. Explain what the alert measures, what the threshold represents, and what a trigger indicates.
-
User Labels: Include a
user_labelsblock in allgoogle_monitoring_alert_policyresources to track policies created by this skill:user_labels = { created-with-google-skill = "g
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
85.4kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
