sgrna-design-guide
Three-tiered sgRNA design guide using validated Addgene sequences, CRISPick pre-computed datasets, or de novo design rules for CRISPR experiments
Install / Use
npx skills add jaechang-hits/SciAgent-Skills --skill sgrna-design-guideInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Development & EngineeringSupported Platforms
Our assessment of sgrna-design-guide
sgrna-design-guide scores 91/100 on our quality scale, 1168th of 4,619 Development & Engineering skills we index (top 26%).
Its SKILL.md is 24 KB long, well organised into 67 sections with 21 code examples: a thorough specification that gives an agent plenty to work with.
It has 367 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 37 days ago, so sgrna-design-guide is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
sgrna-design-guide compared with similar skills
All 4 of these similar skills score higher than sgrna-design-guide; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| sgrna-design-guide (this skill)by jaechang-hits | 91 | 367 | 37d ago | SKILL.md |
| ai-job-searchby MadsLorentzen | 100 | 45.0k | 1d ago | CLAUDE.md |
| claude-howtoby luongnv89 | 100 | 41.7k | 4d ago | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 12d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 12d ago | SKILL.md |
Frequently asked questions
- How do I install sgrna-design-guide?
- Run
npx skills add jaechang-hits/SciAgent-Skills --skill sgrna-design-guide. The install tabs above show the steps for each supported agent. - Which AI agents does sgrna-design-guide work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is sgrna-design-guide safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is sgrna-design-guide still maintained?
- The repository was last updated 37 days ago, so sgrna-design-guide is actively maintained.
Skill content
View source on GitHubname: sgrna-design-guide description: Three-tiered sgRNA design guide using validated Addgene sequences, CRISPick pre-computed datasets, or de novo design rules for CRISPR experiments license: open
sgRNA Design Guide: Three-Tiered Approach
Metadata
Short Description: Comprehensive guide for finding or designing sgRNAs using validated sequences, CRISPick datasets, or de novo design tools.
Authors: Ohagent Team
Version: 1.0
Last Updated: November 2025
License: CC BY 4.0
Commercial Use: Allowed
Citations and Acknowledgments
If you use validated sgRNAs from our database (Option 1):
- Database Source: Addgene (https://www.addgene.org)
- Citation: Always cite the original publication associated with each sgRNA using the PubMed ID provided in the database
- Acknowledgment: "Validated sgRNA sequences obtained from Addgene (https://www.addgene.org/crispr/reference/grna-sequence/)"
If you use CRISPick designs (Option 2):
- Acknowledgment Statement: "Guide designs provided by the CRISPick web tool of the GPP at the Broad Institute"
- Citation for Cas9 designs (SpCas9, SaCas9): Sanson KR, et al. Optimized libraries for CRISPR-Cas9 genetic screens with multiple modalities. Nat Commun. 2018;9(1):5416. PMID: 30575746
- Citation for Cas12a designs (AsCas12a, enAsCas12a): DeWeirdt PC, et al. Optimization of AsCas12a for combinatorial genetic screens in human cells. Nat Biotechnol. 2021;39(1):94-104. PMID: 32661438
- Note: This paper describes enAsCas12a optimization; specify which variant you used in your methods
Overview
This guide provides a three-tiered approach to sgRNA design, prioritizing validated sequences before moving to computational predictions. Always start with Option 1 and proceed to subsequent options only if needed.
Key Concepts
Validated sgRNA Databases (Tier 1)
Validated sgRNAs are guide sequences that have been experimentally tested in published work, with documented cell line, cutting efficiency, and (often) off-target characterization. Addgene curates such sequences alongside the deposited plasmids that carry them, and a downloadable CSV (addgene_grna_sequences.csv) provides a searchable index keyed by gene symbol, target species, and application (cut / activate / RNA targeting). Using a validated sgRNA is preferred because the largest source of CRISPR experimental failure is poor on-target activity that would have been detected during the original validation.
Computational sgRNA Scoring (Tier 2)
When no validated sgRNA exists, pre-computed genome-scale designs from the Broad Institute's CRISPick service are the next best option. CRISPick provides 238 datasets covering multiple genomes, Cas variants, and applications, with each candidate guide ranked by on-target efficiency (how reliably it cuts), off-target specificity (how unlikely it is to cut elsewhere), and a combined rank that balances both. The on-target models behind CRISPick are Sanson 2018 (Cas9) and DeWeirdt 2021 (Cas12a). Combined Rank is the recommended default; use On-Target or Off-Target rank only when the experiment specifically prioritizes one over the other.
Off-Target Stringency and PAM Compatibility
PAM (Protospacer Adjacent Motif) requirements differ by Cas variant: SpCas9 needs NGG (3'), SaCas9 needs NNGRRT (3'), AsCas12a and enAsCas12a need TTTV (5'). Critically, AsCas12a and enAsCas12a are different enzymes: enAsCas12a is an engineered variant with broadened activity, and guides optimized for one will not perform identically on the other. Always match the CRISPick dataset to the exact Cas variant used in the lab. Off-target stringency thresholds (typically Off-Target Rank or a CFD-style score) trade specificity against the size of the candidate pool — tighter thresholds yield fewer but cleaner guides.
De Novo Design Rules (Tier 3)
When neither validated sequences nor pre-computed datasets cover the target (non-model organism, custom locus, etc.), apply rule-based design: 20 bp protospacer for SpCas9/SaCas9 (23–25 bp for Cas12a), GC content 40–60%, avoid TTTT (Pol III terminator) and homopolymer runs >4. Target location matters: early exons (first 50% of coding sequence) for knockout, −200 to +1 from TSS for CRISPRa, −50 to +300 from TSS for CRISPRi.
Decision Framework
sgRNA design decision tree
└── Does Addgene database (Tier 1) contain a validated sgRNA for your gene + species + application?
├── Yes -> Use the validated sgRNA(s); cite the original PubMed reference
└── No -> Run advanced literature search (mandatory before Tier 2)
├── Found in literature -> Use the literature sgRNA; cite the paper
└── Still no match -> Is the target organism + Cas variant covered by a CRISPick dataset (Tier 2)?
├── Yes -> Download dataset, filter by gene, sort by Combined Rank
│ └── Pick top 3-4 sgRNAs (ideally from different exons for redundancy)
└── No -> De novo design (Tier 3) using 20 bp / PAM / GC / avoid-TTTT rules
└── Validate experimentally (Sanger / T7E1 / amplicon-seq)
| Situation | Recommended tier | Rationale |
|-----------|------------------|-----------|
| Common human/mouse gene with prior CRISPR publications | Tier 1 (Addgene + literature) | Validated sequences come with measured cutting efficiency; lowest experimental risk |
| Genome-scale screen of a model organism | Tier 2 (CRISPick) | Pre-computed datasets cover whole genomes with consistent scoring |
| Single human gene knockout, no Addgene hit | Tier 2 (CRISPick GRCh38 SpCas9 CRISPRko), filter by Combined Rank | Best balance of efficiency and specificity for one-off knockouts |
| CRISPRa / CRISPRi experiment | Tier 2 with the matching CRISPRa / CRISPRi dataset | Activation/inhibition models target proximal-promoter windows, not coding exons |
| Cas12a (AsCas12a vs enAsCas12a) | Tier 2 with the exact Cas12a variant dataset | Guides for one variant are not interchangeable with the other |
| Non-model organism not covered by CRISPick | Tier 3 (de novo rules) | No high-throughput training data exists; rule-based design + experimental validation |
| Small custom locus (e.g., engineered cassette) | Tier 3 | Locus is not in any reference genome; design directly against the target sequence |
| Maximum specificity required (e.g., therapeutic application) | Tier 2 sorted by Off-Target Rank | Prioritize guides with the fewest predicted off-target sites |
| Maximum cutting efficiency required (e.g., difficult cell line) | Tier 2 sorted by On-Target Rank | Accept slightly higher off-target risk in exchange for activity |
Best Practices
- Always exhaust Tier 1 before moving to Tier 2. Even if
addgene_grna_sequences.csvreturns no rows, run the literature search step (Method 2). Many validated sgRNAs are published in supplementary materials but never deposited at Addgene. - Match the CRISPick dataset to the exact Cas variant. AsCas12a and enAsCas12a are distinct enzymes; SpCas9 and SaCas9 require different PAMs. Using the wrong dataset wastes experimental resources.
- Use Combined Rank as the default selector. It balances on-target efficiency and off-target specificity; reach for
On-Target RankorOff-Target Rankonly when the experiment has a specific reason to prioritize one. - Order 3–4 sgRNAs per target gene, ideally from different exons. A single guide can fail for reasons unrelated to its predicted score (chromatin context, secondary structure, PCR drop-out). Redundancy is cheap insurance.
- For knockouts, target the first 50% of the coding sequence. N-terminal cuts are most likely to produce a true loss-of-function via NMD; C-terminal cuts can leave partially functional protein.
- For CRISPRa, design within −200 to +1 bp of the TSS; for CRISPRi, within −50 to +300 bp. Effector activity drops sharply outside these windows.
- Validate experimentally. Even high-ranked guides fail ~20–30% of the time. Confirm editing with Sanger sequencing, T7E1, ICE, or amplicon-seq before committing to phenotype assays.
- Cite the original validation paper for Tier 1, and CRISPick + Sanson 2018 / DeWeirdt 2021 for Tier 2. Reproducibility depends on traceable provenance for every guide sequence.
Common Pitfalls
-
Pitfall: Skipping the literature search step (Method 2) when the Addgene CSV returns no hits.
- How to avoid: Treat Method 2 as mandatory; many validated guides only appear in published supplementary materials.
-
Pitfall: Using AsCas12a guides with enAsCas12a (or vice versa) because both share TTTV PAM.
- How to avoid: Match the CRISPick dataset filename's Cas tag (
AsCas12avsenAsCas12a) to the exact enzyme variant used at the bench.
- How to avoid: Match the CRISPick dataset filename's Cas tag (
-
Pitfall: Sorting by On-Target Rank only and ending up with high-activity, low-specificity guides that produce off-target editing.
- How to avoid: Use Combined Rank as the default; only deviate when the experiment explicitly requires extreme efficiency or specificity.
-
Pitfall: Designing a single sgRNA per gene and concluding "the guide doesn't work" when one fails.
- How to avoid: Always order and test 3–4 guides per gene from independent loci/exons.
-
Pitfall: Targeting late exons or 3' UTRs for knockout. Truncated proteins from late-exon edits are often partially functional and confound phenotype interpretation.
- How to avoid: For knockouts, restrict candidates to the first 50% of the coding sequence (or the first 3 exons).
-
Pitfall: Using a CRISPRko dataset for a CRISPRa or CRISPRi experiment. The protospacer windows are different — coding exons for knockout, proximal promoter for activation/inhibition.
- How to avoid: Download the dataset whose application tag (
CRISPRko/CRISPRa/CRISPRi) matches the experiment.
- How to avoid: Download the dataset whose application tag (
-
Pitfall: Designing guides containing
TTTT. This sequence terminates RNA Pol III transcription, so the sgRNA simply will not be expressed from a U6 promoter.- How to avoid: Always filter de novo designs (and double-check Tier 1/2 picks) against
TTTTand homopolymer runs >4.
- How to avoid: Always filter de novo designs (and double-check Tier 1/2 picks) against
-
Pitfall: Failing to cite the original publication for a validated sgRNA.
- How to avoid: Record the
PubMed_IDfrom the Addgene CSV (or the DOI from the literature search) as part of the design record and include it in the methods section.
- How to avoid: Record the
Option 1: Search Validated sgRNA Sequences (Recommended First)
1.1 Search for Validated sgRNAs
IMPORTANT: You MUST complete BOTH Method 1 AND Method 2 before proceeding to Option 2. Do not skip Method 2 even if Method 1 finds no results.
Method 1: Search Our Database (Fastest)
We maintain a curated database of 300+ validated sgRNA sequences from Addgene with experimental evidence.
Location: resource/addgene_grna_sequences.csv (relative to this skill directory)
Search the database:
import pandas as pd
# Load the database
df = pd.read_csv('addgene_grna_sequences.csv')
# Search for your gene
gene_name = "TP53"
results = df[df['Target_Gene'].str.upper() == gene_name.upper()]
# Filter by species and application
results_filtered = results[
(results['Target_Species'] == 'H. sapiens') &
(results['Application'] == 'cut') # or 'activate', 'RNA targeting'
]
# Display results with references
print(results_filtered[['Target_Gene', 'Target_Sequence',
'Plasmid_ID', 'PubMed_ID', 'Depositor']])
Database columns:
Target_Gene: Gene symbolTarget_Species: Organism (H. sapiens, M. musculus, etc.)Target_Sequence: 20bp sgRNA sequence (5' to 3')Application: cut (knockout), activate (CRISPRa), RNA targeting (CRISPRi)Cas9_Species: S. pyogenes, S. aureus, etc.Plasmid_ID: Addgene plasmid numberPlasmid_URL: Direct link to plasmid pagePubMed_ID: Publication reference (cite this in your work)
Truncated for display — read the full file on GitHub.
Related Skills
ai-job-search
45.0kThe job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
claude-howto
41.7kA visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
