gke-ai-troubleshooting-node-unresponsive-timeout
Diagnose and mitigate GKE TPU or GPU nodes stuck in NotReady / NodeStatusUnknown ("Kubelet stopped posting node status") due to host kernel panics, hardware lockups, or disabled node auto-repair
Install / Use
npx skills add google/skills --skill gke-ai-troubleshooting-node-unresponsive-timeoutInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Development & EngineeringSupported Platforms
Tags
Our assessment of gke-ai-troubleshooting-node-unresponsive-timeout
gke-ai-troubleshooting-node-unresponsive-timeout scores 97/100 on our quality scale, 161st of 4,570 Development & Engineering skills we index (top 4%).
Its SKILL.md is 12 KB long, well organised into 9 sections with 2 code examples: a thorough specification that gives an agent plenty to work with.
With 20,340 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 14 days ago, so gke-ai-troubleshooting-node-unresponsive-timeout is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
gke-ai-troubleshooting-node-unresponsive-timeout compared with similar skills
All 4 of these similar skills score higher than gke-ai-troubleshooting-node-unresponsive-timeout; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| gke-ai-troubleshooting-node-unresponsive-timeout (this skill)by google | 97 | 20.3k | 14d ago | SKILL.md |
| ai-job-searchby MadsLorentzen | 100 | 45.3k | 2d ago | CLAUDE.md |
| claude-howtoby luongnv89 | 100 | 41.8k | 8d ago | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 15d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 15d ago | SKILL.md |
Frequently asked questions
- How do I install gke-ai-troubleshooting-node-unresponsive-timeout?
- Run
npx skills add google/skills --skill gke-ai-troubleshooting-node-unresponsive-timeout. The install tabs above show the steps for each supported agent. - Which AI agents does gke-ai-troubleshooting-node-unresponsive-timeout work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is gke-ai-troubleshooting-node-unresponsive-timeout safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is gke-ai-troubleshooting-node-unresponsive-timeout still maintained?
- The repository was last updated 14 days ago, so gke-ai-troubleshooting-node-unresponsive-timeout is actively maintained.
Skill content
View source on GitHubname: gke-ai-troubleshooting-node-unresponsive-timeout description: >- Diagnose and mitigate GKE TPU or GPU nodes stuck in NotReady / NodeStatusUnknown ("Kubelet stopped posting node status") due to host kernel panics, hardware lockups, or disabled node auto-repair. Use when nodes stop heartbeating beyond the node auto-repair threshold and pods remain stuck in Terminating. Don't use for healthy nodes, pod-only application crashes, or routine GKE upgrades. metadata: version: "1.0.0" category: Containers
Troubleshoot unresponsive GKE TPU and GPU nodes (NodeStatusUnknown)
When the Compute Engine host of a TPU or GPU node has a fatal hardware error,
kernel panic, or non-maskable interrupt (NMI) lockup, the guest OS stops
responding. The kubelet can no longer send heartbeats, so the node Ready
condition becomes Unknown with Reason: NodeStatusUnknown (Kubelet stopped posting node status.). If node auto-repair is disabled on the node pool, GKE
doesn't repair the node. The node can stay NotReady, and pods on it can stay
in Terminating, which blocks multi-host JobSet workloads from recovering.
Prerequisites
- Tools: Install the
Google Cloud SDK (
gcloud) andkubectl. - Cloud Billing & Project Configuration: Verify an active billing account is
linked (
gcloud billing projects describe {project_id}), authenticate (gcloud auth login), set the target project (gcloud config set project {project_id}), and ensurecontainer.googleapis.com,compute.googleapis.com,logging.googleapis.com, andmonitoring.googleapis.comare enabled. - Required IAM Roles:
- Kubernetes Engine Viewer (
roles/container.viewer) - Compute Viewer (
roles/compute.viewer) - Logs Viewer (
roles/logging.viewer) - Monitoring Viewer (
roles/monitoring.viewer) - For remediation (
[High Risk]steps): Kubernetes Engine Cluster Admin (roles/container.clusterAdmin)
- Kubernetes Engine Viewer (
- Documentation:
- Troubleshoot nodes with the NotReady status in GKE (Sections: "Check the node's status and conditions", "Confirm node preemption", "Verify that the node has recovered")
- Deploy TPU workloads in GKE Standard (Sections: "Monitor health metrics for TPU nodes and node pools", "Configure auto repair for TPU slice nodes")
- Troubleshoot OOM events
- View GKE logs (Sections: "System logs")
- Viewing serial port output
- Troubleshoot Linux VM boot issues due to kernel panic
- Auto-repair nodes (Sections: "Settings for Autopilot and Standard", "Verify node auto-repair is enabled for a Standard node pool", "Get information about recent automated repair events", "Enable auto-repair for an existing Standard node pool", "Repair criteria", "Node repair process", "Node auto repair in TPU slice nodes")
- Troubleshooting VM shutdowns and reboots (Sections: "Querying Cloud Audit Logs", "Reviewing Cloud Audit Logs")
Read-only rule: Run read-only diagnostic commands only. Never drain, delete, or re-create nodes, or run any other command that changes the cluster. Give the user any fix to apply themselves.
When you recommend a fix, link the doc section that describes it.
Diagnostic workflow
Step 0: Collect context and set the investigation window [Low Risk]
Collect the target parameters. By default, query a 60-minute window `[T - 30m, T
- 30m]
around{issue_time}`:
{project_id}: Google Cloud project ID{cluster_name}: GKE cluster name{location}: Cluster region or zone{nodepool_name}: Target TPU or GPU node pool name{node_name}: Unresponsive GKE node name (and its Compute Engine{zone}){issue_time}: Incident timestamp in RFC3339 UTC{start_time}:{issue_time} - 30m{end_time}:{issue_time} + 30m
Step 1: Verify the NodeStatusUnknown heartbeat timeout [Low Risk]
- Check Kubernetes node conditions: To inspect the node status and verify
whether the
Readycondition isUnknownwithReason: NodeStatusUnknown(Kubelet stopped posting node status.), follow the instructions in the section Check the node's status and conditions. - Query Cloud Logging (read-only LQL): Query
k8s_nodeandk8s_clusterlogs across[{start_time}, {end_time}]to confirm when the control plane lost heartbeat contact with{node_name}:
(resource.type="k8s_node" OR resource.type="k8s_cluster")
resource.labels.cluster_name="{cluster_name}"
("{node_name}" AND ("NodeNotReady" OR "NodeStatusUnknown" OR "Kubelet stopped posting node status"))
timestamp >= "{start_time}" AND timestamp <= "{end_time}"
- Query Cloud Monitoring (read-only PromQL): Correlate the duration of the
Unknownstate using the GKE system metrickubernetes.io/node/status_condition(kubernetes_io:node_status_condition, GKE1.32.1-gke.1357001+) documented in Monitor health metrics for TPU nodes and node pools, filtered bycondition="Ready"andstatus="Unknown":
kubernetes_io:node_status_condition{
monitored_resource="k8s_node",
cluster_name="{cluster_name}",
node_name="{node_name}",
condition="Ready",
status="Unknown"
}
- Decision logic:
- If the node
Readycondition isTrueandNodeStatusUnknownis absent, rule out an unresponsive node timeout and pivot to workload-level troubleshooting (for example, Troubleshoot OOM events) rather than repairing or draining the node. - If
ReadyisUnknown(NodeStatusUnknown), proceed to Step 2.
- If the node
Step 2: Inspect serial port output for a kernel panic [Low Risk]
The guest OS can't send logs after a fatal kernel freeze, so check the serial port output of the node's VM:
- Cloud Logging: If serial port logging is enabled (i.e. VM metadata
serial-port-logging-enableis set totrue), GKE system logs in Cloud Logging include the node's serial port output. See the section System logs. Use Cloud Logging when the VM is stopped or has already been replaced by auto-repair, or when you need more than the most recent output. - Running VM: Follow
Viewing serial port output
to retrieve the serial port 1 output (
gcloud compute instances get-serial-port-outputwith--port=1) for{node_name}in{zone}. This method returns only the most recent 1 MB of output per port.
Consult
Troubleshoot Linux VM boot issues due to kernel panic
to identify documented kernel panic and hardware crash patterns (such as Fatal Machine check, hung_task: blocked tasks, or NMI: Not continuing) in the
serial port output.
Step 3: Check node auto-repair status [Low Risk]
Find out why GKE hasn't repaired the unresponsive node:
- Autopilot clusters: Autopilot always repairs nodes, and you can't turn this off (see Settings for Autopilot and Standard). Skip the configuration check and check the repair history.
- Standard clusters: Follow
Verify node auto-repair is enabled for a Standard node pool
to check whether
autoRepairis enabled on{nodepool_name}. - Repair history (both modes): Auto-repair can be enabled and the repair can still fail. Per Configure auto repair for TPU slice nodes, check the repair status, including the failure reason, in the operation history described in Get information about recent automated repair events. If the failure is caused by insufficient quota, tell the user to contact their Google Cloud account representative to increase the quota.
Step 4: Check Compute Engine system events [Low Risk]
To list the system events for {node_name} around {issue_time}, follow the
instructions in the section
Querying Cloud Audit Logs.
Compare the method field with the table in the "Reviewing Cloud Audit Logs"
section of the same document, and look for:
compute.instances.hostError: a hardware or software issue on the physical host caused the VM to crash.compute.instances.preempted: Compute Engine preempted a Spot VM or preemptible VM. For preempted nodes, also see Confirm node preemption.compute.instances.automaticRestart: Compute Engine restarted the VM after ahostErrororterminateOnHostMaintenanceevent.compute.instances.guestTerminate: the VM's operating system initiated the shutdown.
Step 5: Resolution [High Risk]
Guardrails:
- Never force-delete stuck
Terminatingpods on an unresponsive node. Force deletion doesn't wait for the kubelet to confirm that the pod has stopped, so a replacement pod can start while the old one is still running. See Force Delete StatefulSet Pods.- Never delete GKE-managed Compute Engine VM instances directly (
gcloud compute instances delete). Instead, check how long the node has reportedNodeStatusUnknown(Step 1), check the serial console output (Step 2), and rely on node auto-repair, as described in the following steps.
- Enable node auto-repair (Standard clusters only):
- Autopilot clusters always auto-repair nodes, so skip this step.
- Before the user enables auto-repair on a multi-host TPU slice node pool, tell them that GKE re-creates the entire node pool when a node in it needs repair (see step 2).
- If
autoRepairis disabled on{nodepool_name}, link the user to Enable auto-repair for an existing Standard node pool and Configure auto repair for TPU slice nodes so they can apply the change themselves.
- Let GKE repair the node:
- Explain that GKE repairs a node that reports
NotReadyor no status for the documented time threshold by draining and re-creating it, and link the "Repair criteria" and "Node repair process" sections of Auto-repair nodes. GKE waits one hour for the drain to complete. If the drain doesn't
- Explain that GKE repairs a node that reports
Truncated for display — read the full file on GitHub.
Related Skills
ai-job-search
45.3kThe job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
claude-howto
41.8kA visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
