ai-llm-engineering
Operational skill hub for LLM system architecture, evaluation, deployment, and optimization (modern production standards). Links to specialized skills for prompts, RAG, agents, and safety.
Install / Use
npx skills add Microck/ordinary-claude-skills --skill ai-llm-engineeringInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
SecuritySupported Platforms
Our assessment of ai-llm-engineering
ai-llm-engineering scores 83/100 on our quality scale, 853rd of 1,119 Security skills we index.
Its SKILL.md is 12 KB long, well organised into 24 sections with 1 code example: a thorough specification that gives an agent plenty to work with.
It has 399 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 30 days ago, so ai-llm-engineering is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
ai-llm-engineering compared with similar skills
All 4 of these similar skills score higher than ai-llm-engineering; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| ai-llm-engineering (this skill)by Microck | 83 | 399 | 30d ago | SKILL.md |
| algorithmic-artby anthropics | 100 | 177.9k | 14d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 14d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 15d ago | SKILL.md |
| ui-ux-pro-maxby nextlevelbuilder | 100 | 130.2k | 15d ago | SKILL.md |
Frequently asked questions
- How do I install ai-llm-engineering?
- Run
npx skills add Microck/ordinary-claude-skills --skill ai-llm-engineering. The install tabs above show the steps for each supported agent. - Which AI agents does ai-llm-engineering work with?
- It is written for Zed, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is ai-llm-engineering safe to use?
- It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is ai-llm-engineering still maintained?
- The repository was last updated 30 days ago, so ai-llm-engineering is actively maintained.
Skill content
View source on GitHubname: ai-llm-engineering description: | Operational skill hub for LLM system architecture, evaluation, deployment, and optimization (modern production standards). Links to specialized skills for prompts, RAG, agents, and safety. Integrates recent advances: PEFT/LoRA fine-tuning, hybrid RAG handoff (see dedicated skill), vLLM 24x throughput, multi-layered security (90%+ bypass for single-layer), automated drift detection (18-second response), and CI/CD-aligned evaluation.
LLM Engineering – Operational Skill Hub
A single resource for executing, validating, and scaling LLM systems with modern production standards, while delegating domain depth to specialized skills.
This skill provides quick reference, decision frameworks, and navigation to detailed operational patterns for:
- Data, training, fine-tuning (PEFT/LoRA standard)
- Evaluation (automated testing, metrics, rollout gates)
- Deployment (vLLM 24x throughput, FP8/FP4 quantization)
- LLMOps (automated drift detection, retraining)
- Safety (multi-layered defenses, AI-powered guardrails)
For detailed patterns: See Resources and Templates sections below.
Quick Reference
| Task | Tool/Framework | Command/Pattern | When to Use | |------|----------------|-----------------|-------------| | RAG Pipeline | LlamaIndex, LangChain | Page-level chunking + hybrid retrieval | Dynamic knowledge, 0.648 accuracy | | Agentic Workflow | LangGraph, AutoGen, CrewAI | ReAct, multi-agent orchestration | Complex tasks, tool use required | | Prompt Design | Anthropic, OpenAI guides | CoT, few-shot, structured | Task-specific behavior control | | Evaluation | LangSmith, W&B, RAGAS | Multi-metric (hallucination, bias, cost) | Quality validation, A/B testing | | Production Deploy | vLLM, TensorRT-LLM | FP8/FP4 quantization, 24x throughput | High-throughput serving, cost optimization | | Monitoring | Arize Phoenix, LangFuse | Drift detection, 18-second response | Production LLM systems |
Decision Tree: LLM System Architecture
Building LLM application: [Architecture Selection]
├─ Need current knowledge?
│ ├─ Simple Q&A? → Basic RAG (page-level chunking + hybrid retrieval)
│ └─ Complex retrieval? → Advanced RAG (reranking + contextual retrieval)
│
├─ Need tool use / actions?
│ ├─ Single task? → Simple agent (ReAct pattern)
│ └─ Multi-step workflow? → Multi-agent (LangGraph, CrewAI)
│
├─ Static behavior sufficient?
│ ├─ Quick MVP? → Prompt engineering (CI/CD integrated)
│ └─ Production quality? → Fine-tuning (PEFT/LoRA)
│
└─ Best results?
└─ Hybrid (RAG + Fine-tuning + Agents) → Comprehensive solution
See Decision Matrices for detailed selection criteria.
When to Use This Skill
Claude should invoke this skill when the user asks about:
- LLM preflight/project checklists, production best practices, or data pipelines
- Building or deploying RAG, agentic, or prompt-based LLM apps
- Prompt design, chain-of-thought (CoT), ReAct, or template patterns
- Troubleshooting LLM hallucination, bias, retrieval issues, or production failures
- Evaluating LLMs: benchmarks, multi-metric eval, or rollout/monitoring
- LLMOps: deployment, rollback, scaling, resource optimization
- Technology stack selection (models, vector DBs, frameworks)
- Production deployment strategies and operational patterns
Scope Boundaries (Use These Skills for Depth)
- Prompt design & CI/CD → ai-prompt-engineering
- RAG pipelines & chunking → ai-llm-rag-engineering
- Search tuning (BM25, HNSW, hybrid) → ai-llm-search-retrieval
- Agent architectures & tools → ai-agents-development
- Serving optimization/quantization → ai-llm-ops-inference
- Production deployment/monitoring → ai-ml-ops-production
- Security/guardrails → ai-ml-ops-security
Resources (Best Practices & Operational Patterns)
Comprehensive operational guides with checklists, patterns, and decision frameworks:
Core Operational Patterns
-
Project Planning Patterns - Stack selection, FTI pipeline, performance budgeting
- AI engineering stack selection matrix
- Feature/Training/Inference (FTI) pipeline blueprint
- Performance budgeting and goodput gates
- Progressive complexity (prompt → RAG → fine-tune → hybrid)
-
Production Checklists - Pre-deployment validation and operational checklists
- LLM lifecycle checklist (modern production standards)
- Data & training, RAG pipeline, deployment & serving
- Safety/guardrails, evaluation, agentic systems
- Reliability & data infrastructure (DDIA-grade)
- Weekly production tasks
-
Common Design Patterns - Copy-paste ready implementation examples
- Chain-of-Thought (CoT) prompting
- ReAct (Reason + Act) pattern
- RAG pipeline (minimal to advanced)
- Agentic planning loop
- Self-reflection and multi-agent collaboration
-
Decision Matrices - Quick reference tables for selection
- RAG type decision matrix (naive → advanced → modular)
- Production evaluation table with targets and actions
- Model selection matrix (GPT-4, Claude, Gemini, self-hosted)
- Vector database, embedding model, framework selection
- Deployment strategy matrix
-
Anti-Patterns - Common mistakes and prevention strategies
- Data leakage, prompt dilution, RAG context overload
- Agentic runaway, over-engineering, ignoring evaluation
- Hard-coded prompts, missing observability
- Detection methods and prevention code examples
Domain-Specific Patterns
- LLMOps Best Practices - Operational lifecycle and deployment patterns
- Evaluation Patterns - Testing, metrics, and quality validation
- Prompt Engineering Patterns - Quick reference (canonical skill: ai-prompt-engineering)
- Agentic Patterns - Quick reference (canonical skill: ai-agents-development)
- RAG Best Practices - Quick reference (canonical skill: ai-llm-rag-engineering)
Note: Each resource file includes preflight/validation checklists, copy-paste reference tables, inline templates, anti-patterns, and decision matrices.
Templates (Copy-Paste Ready)
Production templates by use case and technology:
RAG Pipelines
- Basic RAG - Simple retrieval-augmented generation
- Advanced RAG - Hybrid retrieval, reranking, contextual embeddings
Prompt Engineering
- Chain-of-Thought - Step-by-step reasoning pattern
- ReAct - Reason + Act for tool use
Agentic Workflows
- Reflection Agent - Self-critique and improvement
- Multi-Agent - Manager-worker orchestration
Data Pipelines
- Data Quality - Validation, deduplication, PII detection
Deployment
- LLM Deployment - Production deployment with monitoring
Evaluation
- Multi-Metric Evaluation - Comprehensive testing suite
Related Skills
This skill integrates with complementary Claude Code skills:
Core Dependencies
- ai-llm-rag-engineering - Advanced RAG patterns, chunking strategies, hybrid retrieval, reranking
- ai-llm-search-retrieval - Search optimization, BM25 tuning, vector search, ranking pipelines
- ai-prompt-engineering - Systematic prompt design, evaluation, testing, and optimization
- ai-agents-development - Agent architectures, tool use, multi-agent systems, autonomous workflows
Production & Operations
- ai-llm-development - Model training, fine-tuning, dataset creation, instruction tuning
- ai-llm-ops-inference - Production serving, quantization, batching, GPU optimization
- ai-ml-ops-production - Deployment patterns, monitoring, drift detection, API design
- ai-ml-ops-security - Security guardrails, prompt injection defense, privacy protection
External Resources
See data/sources.json for 50+ curated authoritative sources:
- Official LLM platform docs - OpenAI, Anthropic, Gemini, Mistral, Azure OpenAI, AWS Bedrock
- Open-source models and frameworks - HuggingFace Transformers, LLaMA, vLLM, PEFT/LoRA, DeepSpeed
- RAG frameworks and vector DBs - LlamaIndex, LangChain, LangGraph, Haystack, Pinecone, Qdrant, Chroma
- 2025 Agentic frameworks - Anthropic Agent SDK, AutoGen, CrewAI, LangGraph Multi-Agent, Semantic Kernel
- 2025 RAG innovations - Microsoft GraphRAG (knowledge graphs), Pathway (real-time), hybrid retrieval
- Prompt engineering - Anthropic Prompt Library, Prompt Engineering Guide, CoT/ReAct patterns
- Evaluation and monitoring - OpenAI Evals, HELM, Anthropic Evals, LangSmith, W&B, Arize Phoenix
- Production deployment - LiteLLM, Ollama, RunPod, Together AI, vLLM serving
Usage
For New Projects
- Start with Production Checklists - Validate all pre-deployment requirements
- Use Decision Matrices - Select technology stack
- Reference Project Planning Patterns - Design FTI pipeline
- Implement with Common Design Patterns - Copy-paste code examples
- Avoid Anti-Patterns - Learn from common mistakes
For Troubleshooting
- Check Anti-Patterns - Identify failure modes and mitigations
- Use Decision Matrices - Evaluate if architecture fits use case
- Reference Common Design Patterns - Verify implementation correctness
For Ongoing Operations
- Follow Production Checklists - Weekly operational tasks
- Integrate Evaluation Patterns - Continuous quality monitoring
- Apply LLMOps Best Practices - Deployment and rollback procedures
Navigation Summary
Quick Decisions: Decision Matrices Pre-Deployment: Production Checklists Planning: Project Planning Patterns Implementation: Common Design Patterns Troubleshooting: Anti-Patterns
Domain Depth: LLMOps | [Evaluation](resources/ev
Truncated for display — read the full file on GitHub.
Related Skills
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
ui-ux-pro-max
130.2kUI/UX design intelligence for web, mobile, and desktop. This skill should be used when designing, building, reviewing, or fixing interfaces, including pages, components, design systems, accessibility, interaction, responsive layout, typography, color, charts, and stack-specific UI implementation.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
