google-cloud-solution-agentic-analytics-spark-knowledge-catalog
Discovers requirements and designs an end-to-end governed agentic analytics solution using Knowledge Catalog and Managed Service for Apache Spark (Lightning Engine)
Install / Use
npx skills add google/skills --skill google-cloud-solution-agentic-analytics-spark-knowledge-catalogInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AutomationSupported Platforms
Our assessment of google-cloud-solution-agentic-analytics-spark-knowledge-catalog
google-cloud-solution-agentic-analytics-spark-knowledge-catalog scores 91/100 on our quality scale, 460th of 1,267 Automation skills we index (top 37%).
Its SKILL.md is 17 KB long, well organised into 13 sections and no code examples: a thorough specification that gives an agent plenty to work with.
With 20,340 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 2 days ago, so google-cloud-solution-agentic-analytics-spark-knowledge-catalog is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-09-26. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
google-cloud-solution-agentic-analytics-spark-knowledge-catalog compared with similar skills
All 4 of these similar skills score higher than google-cloud-solution-agentic-analytics-spark-knowledge-catalog; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| google-cloud-solution-agentic-analytics-spark-knowledge-catalog (this skill)by google | 91 | 20.3k | 2d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 85.4k | 10d ago | CLAUDE.md |
| rufloby ruvnet | 100 | 73.3k | 1d ago | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 83.7k | today | MCP Server |
| algorithmic-artby anthropics | 100 | 177.9k | 3d ago | SKILL.md |
Frequently asked questions
- How do I install google-cloud-solution-agentic-analytics-spark-knowledge-catalog?
- Run
npx skills add google/skills --skill google-cloud-solution-agentic-analytics-spark-knowledge-catalog. The install tabs above show the steps for each supported agent. - Which AI agents does google-cloud-solution-agentic-analytics-spark-knowledge-catalog work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is google-cloud-solution-agentic-analytics-spark-knowledge-catalog safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is google-cloud-solution-agentic-analytics-spark-knowledge-catalog still maintained?
- The repository was last updated 2 days ago, so google-cloud-solution-agentic-analytics-spark-knowledge-catalog is actively maintained.
Skill content
View source on GitHubname: google-cloud-solution-agentic-analytics-spark-knowledge-catalog metadata: version: "1.0.0" category: MultiProductSolutions description: >- Discovers requirements and designs an end-to-end governed agentic analytics solution using Knowledge Catalog and Managed Service for Apache Spark (Lightning Engine). Use when designing data science and analytics workflows across structured and unstructured distributed data (including in S3, Azure Blob, AlloyDB, and Iceberg), establishing metadata governance with Knowledge Catalog aspect types, or grounding agentic IDEs (VS Code, Antigravity) by using the Google Cloud Data Agent Kit. Don't use for provisioning borderless data lakehouse infrastructure (use google-cloud-solution-agentic-ai-borderless-data-lakehouse instead).
Agentic analytics across cloud providers and data types
This skill provides a workflow to design and implement a governed, secure pipeline for agentic analytics solution across structured and unstructured data that's distributed across Google Cloud, on-premises systems, and other cloud providers.
Overview of the workflow
The workflow consists of the following phases:
- Phase 1: Requirements discovery. Gather detailed requirements related to the cloud workload or use case that the user needs assistance for.
- Phase 2: Solution architecture. Use the requirements that were gathered in Phase 1 to generate a detailed solution architecture for the cloud workload or use case.
- Phase 3: Solution validation. Create a plan to validate the generated solution, generate validation instructions and scripts, and run the validation.
- Phase 4: Solution packing and presentation. Consolidate the generated content and present the solution.
Important notes about the workflow:
- Strict phase separation: During Phase 1 (Requirements discovery), when you ask the user clarifying questions, DON'T recommend, propose, or outline any architectural designs, technical decompositions, cloud services, or component mappings.
- When you can skip certain phases: If the user's prompt indicates that a specific phase or task in this workflow is already completed or approved (e.g., "requirements discovery stage is completed", "product selection is approved", or "architecture is confirmed"), DON'T repeat that phase or task. Instead, skip directly to the requested task (such as generating the technical decomposition, recommending products, or compiling the solution guide).
Phase 1: Requirements discovery and analysis
-
Request the user to describe the functional requirements (business processes, activities, and use cases) of their workload. Ask the user the following questions, one question at a time:
- What are your primary inventory data sources? Are they unstructured (e.g., PDF flavor recipes, invoices) or structured (e.g., historical sales in Iceberg)?
- Where are these sources hosted? Are they split across AWS S3, Azure Blob, Google Cloud Storage, or databases like AlloyDB?
- How do you manage and federate metadata across your data sources within Google Cloud and in external locations (such as other cloud providers)?
- What are your analytical and computational requirements to join, clean, and run forecast models over large-scale distributed data?
- What types of natural language prompts do your data scientists or operational agents expect to execute in their agentic IDE (VS Code or Antigravity IDE)?
-
Request the user to describe the non-functional requirements of their workload.
The following are examples of questions you can ask to gather non-functional requirements:
- Security, privacy, and compliance: What data privacy rules, regulatory compliance (e.g., GDPR, HIPAA), or data governance requirements must the system adhere to?
- Reliability: What are your uptime, high-availability, fault-tolerance, and disaster recovery objectives (RTO/RPO)?
- Performance: What target query latencies and SLA expectations does your workload require?
- Operations: What operational monitoring metrics do your data scientists and engineers need?
- Cost & Sustainability: Do you have specific budget constraints and data egress/transfer cost requirements?
-
Ask the user whether the workload currently runs on other cloud providers or on-premises.
- If the user answers "yes", then ask the user to describe the architecture of the current deployment.
- If the user answer "no", then proceed to the next step.
-
Request the user to describe dependencies, if any, on other workloads, products, or tools. The following are examples of questions that you can ask to get information about the dependencies:
- Do you have any upstream or downstream dependencies on external systems (e.g., identity providers, data curation platforms, CI/CD pipelines, or active data catalogs)?
- Are there any requirements for your general data-engineering software delivery lifecycle (e.g., version control, testing, data quality assurance)? Provide the path to a directory or examples of these artifacts.
-
Review the input that the user has provided so far, and check whether there are any ambiguities or contradictions.
If you identify any ambiguities or contradictions in the requirements that the user has provided (e.g., zero-copy vs copying data to a repository), then do the following for each ambiguity or contradiction that you identify:
- Describe the ambiguity or contradiction (e.g., explain why copying data contradicts the zero-copy requirement and also incurs data-transfer costs).
- Ask the user how they wish to resolve the ambiguity or contradiction.
- If the user delegates the choice to you (e.g., the user replies with "do what you think is best" or "you decide"), then provide a clear suggestion to resolve the ambiguity or contradiction (e.g., suggest prioritizing zero-copy remote queries), explain your reasoning (e.g., to eliminate multi-cloud fees and data duplication), and ask the user to approve your suggestion.
Critical: Until all the ambiguities and contradictions that you identify are resolved according to the preceding guidance, you must NOT recommend or generate any architecture design, technical decomposition, or Google Cloud product recommendations.
-
Important: DON'T start this step if there are unresolved contradictions or ambiguities from Step 5.
Generate a technical decomposition of the components of the workload.
- The technical decomposition must break down the solution into logical components.
- The decomposition MUST address role-based security and credentials within the relevant layers.
- The decomposition MUST be organized under the following four layers,
which represent a standard architectural pattern for agentic analytics
solutions, flowing from user interaction through data context and
governance to core data processing:
- User-interaction layer (IDE): e.g., agentic development environment.
- Grounding and trusted data: e.g., foundation model, MCP servers, and data warehouse in the cloud.
- Metadata curation: e.g., metadata scanning.
- Data processing and analytics: e.g., analytics workflows, Spark data processing, and external data stores.
-
Request the user to approve the generated technical decomposition.
-
If the user requests changes, then generate an updated technical decomposition.
-
Repeat steps 5 through 8 until the user approves the generated technical decomposition.
-
After the user approves the technical decomposition, proceed to Phase 2. Important: Don't proceed to the next phase until the user approves the generated technical decomposition of the workload.
Phase 2: Solution architecture
Ground all generated content
For each task in this phase, to ensure that the generated content aligns with the latest and official Google Cloud guidance, ground the generated content by using the following resources:
- Google Developer Knowledge MCP server
- Instructions to connect to the MCP Server: https://developers.google.com/knowledge/mcp.md.txt
- Server: https://developerknowledge.googleapis.com/mcp
- Tools:
developerknowledge:search_documentsdeveloperknowledge:get_documentsdeveloperknowledge:answer_query
- Tools:
- Relevant skills from https://github.com/google/skills
- Official Google Cloud documentation, including the following:
- Reference architecture for agentic cross-cloud analytics workflows across multi-cloud data lakes, structured data warehouses, and unstructured data stores: https://docs.cloud.google.com/architecture/agentic-ai-cross-cloud-analytics.md.txt
- Decision-making guides for the products and topics that are relevant to the workload: https://github.com/google/skills/blob/main/skills/cloud/google-cloud-solution-architecture/references/decision-making-guides.md
- Best-practices guides for the products and topics that are relevant to the workload: https://github.com/google/skills/blob/main/skills/cloud/google-cloud-solution-architecture/references/best-practices-guides.md
Task 2.1: Identify Google Cloud products and features required for the workload.
- For each component in the confirmed technical decomposition, identify the
appropriate Google Cloud products and features, based on the guidance in the
following resources and adjusted suitably based on the approved technical
decomposition:
references/product-selection-guidance.mdhttps://github.com/google/skills/blob/main/skills/cloud/google-cloud-solution-architecture/references/decision-making-guides.md
- Present the generated product recommendations and ask the user to approve the recommendations.
- If the user requests changes, then make the required changes.
- Repeat steps 2 and 3 until the user approves the product recommendations.
- After the user approves the product recommendations, proceed to Task 2.2.
Task 2.2: Generate an architecture diagram.
- Generate an architecture diagram in Mermaid format: https://github.com/mermaid-js/mermaid.
- Present the generated diagram to the user and ask the user to approve the architecture diagram.
- If the user requests changes, then make the required changes.
- Repeat steps 2 and 3 until the user approves the architecture diagram.
- After the user approves the architecture diagram, proceed to Task 2.3.
Task 2.3: Generate an architecture description.
- Generate a description that explains the purpose of each component, the relationships between the components, and the task flow or data flow.
- Present the generated architecture description to the user and ask the user to approve the description.
- If the user requests any changes, then make the required changes.
- Repeat steps 2 and 3 until the user approves the architecture description.
- After the user approves the architecture description, proceed to Task 2.4.
Task 2.4: Generate design recommendations.
-
Generate design recommendations and best practices to optimally configure each component in the architecture based on the workload's requirements.
Important:
- When you generate design recommendations, consider the following:
- Functional
- When you generate design recommendations, consider the following:
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
85.4kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
ruflo
73.3k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
Scrapling
83.7k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
