SkillAgentSearch skills...

sgrna-design-guide

Three-tiered sgRNA design guide using validated Addgene sequences, CRISPick pre-computed datasets, or de novo design rules for CRISPR experiments

Install / Use

npx skills add jaechang-hits/SciAgent-Skills --skill sgrna-design-guide

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

91/100

Supported Platforms

Universal

Our assessment of sgrna-design-guide

sgrna-design-guide scores 91/100 on our quality scale, 1168th of 4,619 Development & Engineering skills we index (top 26%).

Its SKILL.md is 24 KB long, well organised into 67 sections with 21 code examples: a thorough specification that gives an agent plenty to work with.

It has 367 GitHub stars, a meaningful sign that others use it.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
11/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 37 days ago, so sgrna-design-guide is actively maintained.
  • No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
  • Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

Safety scan

No issues found

Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.

Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.

sgrna-design-guide compared with similar skills

All 4 of these similar skills score higher than sgrna-design-guide; compare them before choosing.

SkillScoreStarsUpdatedFormat
sgrna-design-guide (this skill)by jaechang-hits9136737d agoSKILL.md
ai-job-searchby MadsLorentzen10045.0k1d agoCLAUDE.md
claude-howtoby luongnv8910041.7k4d agoCLAUDE.md
algorithmic-artby anthropics100177.9k12d agoSKILL.md
pptxby anthropics100177.9k12d agoSKILL.md

Frequently asked questions

How do I install sgrna-design-guide?
Run npx skills add jaechang-hits/SciAgent-Skills --skill sgrna-design-guide. The install tabs above show the steps for each supported agent.
Which AI agents does sgrna-design-guide work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is sgrna-design-guide safe to use?
Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is sgrna-design-guide still maintained?
The repository was last updated 37 days ago, so sgrna-design-guide is actively maintained.

name: sgrna-design-guide description: Three-tiered sgRNA design guide using validated Addgene sequences, CRISPick pre-computed datasets, or de novo design rules for CRISPR experiments license: open

sgRNA Design Guide: Three-Tiered Approach


Metadata

Short Description: Comprehensive guide for finding or designing sgRNAs using validated sequences, CRISPick datasets, or de novo design tools.

Authors: Ohagent Team

Version: 1.0

Last Updated: November 2025

License: CC BY 4.0

Commercial Use: Allowed

Citations and Acknowledgments

If you use validated sgRNAs from our database (Option 1):

  • Database Source: Addgene (https://www.addgene.org)
  • Citation: Always cite the original publication associated with each sgRNA using the PubMed ID provided in the database
  • Acknowledgment: "Validated sgRNA sequences obtained from Addgene (https://www.addgene.org/crispr/reference/grna-sequence/)"

If you use CRISPick designs (Option 2):

  • Acknowledgment Statement: "Guide designs provided by the CRISPick web tool of the GPP at the Broad Institute"
  • Citation for Cas9 designs (SpCas9, SaCas9): Sanson KR, et al. Optimized libraries for CRISPR-Cas9 genetic screens with multiple modalities. Nat Commun. 2018;9(1):5416. PMID: 30575746
  • Citation for Cas12a designs (AsCas12a, enAsCas12a): DeWeirdt PC, et al. Optimization of AsCas12a for combinatorial genetic screens in human cells. Nat Biotechnol. 2021;39(1):94-104. PMID: 32661438
    • Note: This paper describes enAsCas12a optimization; specify which variant you used in your methods

Overview

This guide provides a three-tiered approach to sgRNA design, prioritizing validated sequences before moving to computational predictions. Always start with Option 1 and proceed to subsequent options only if needed.

Key Concepts

Validated sgRNA Databases (Tier 1)

Validated sgRNAs are guide sequences that have been experimentally tested in published work, with documented cell line, cutting efficiency, and (often) off-target characterization. Addgene curates such sequences alongside the deposited plasmids that carry them, and a downloadable CSV (addgene_grna_sequences.csv) provides a searchable index keyed by gene symbol, target species, and application (cut / activate / RNA targeting). Using a validated sgRNA is preferred because the largest source of CRISPR experimental failure is poor on-target activity that would have been detected during the original validation.

Computational sgRNA Scoring (Tier 2)

When no validated sgRNA exists, pre-computed genome-scale designs from the Broad Institute's CRISPick service are the next best option. CRISPick provides 238 datasets covering multiple genomes, Cas variants, and applications, with each candidate guide ranked by on-target efficiency (how reliably it cuts), off-target specificity (how unlikely it is to cut elsewhere), and a combined rank that balances both. The on-target models behind CRISPick are Sanson 2018 (Cas9) and DeWeirdt 2021 (Cas12a). Combined Rank is the recommended default; use On-Target or Off-Target rank only when the experiment specifically prioritizes one over the other.

Off-Target Stringency and PAM Compatibility

PAM (Protospacer Adjacent Motif) requirements differ by Cas variant: SpCas9 needs NGG (3'), SaCas9 needs NNGRRT (3'), AsCas12a and enAsCas12a need TTTV (5'). Critically, AsCas12a and enAsCas12a are different enzymes: enAsCas12a is an engineered variant with broadened activity, and guides optimized for one will not perform identically on the other. Always match the CRISPick dataset to the exact Cas variant used in the lab. Off-target stringency thresholds (typically Off-Target Rank or a CFD-style score) trade specificity against the size of the candidate pool — tighter thresholds yield fewer but cleaner guides.

De Novo Design Rules (Tier 3)

When neither validated sequences nor pre-computed datasets cover the target (non-model organism, custom locus, etc.), apply rule-based design: 20 bp protospacer for SpCas9/SaCas9 (23–25 bp for Cas12a), GC content 40–60%, avoid TTTT (Pol III terminator) and homopolymer runs >4. Target location matters: early exons (first 50% of coding sequence) for knockout, −200 to +1 from TSS for CRISPRa, −50 to +300 from TSS for CRISPRi.

Decision Framework

sgRNA design decision tree
└── Does Addgene database (Tier 1) contain a validated sgRNA for your gene + species + application?
    ├── Yes -> Use the validated sgRNA(s); cite the original PubMed reference
    └── No  -> Run advanced literature search (mandatory before Tier 2)
        ├── Found in literature -> Use the literature sgRNA; cite the paper
        └── Still no match -> Is the target organism + Cas variant covered by a CRISPick dataset (Tier 2)?
            ├── Yes -> Download dataset, filter by gene, sort by Combined Rank
            │         └── Pick top 3-4 sgRNAs (ideally from different exons for redundancy)
            └── No  -> De novo design (Tier 3) using 20 bp / PAM / GC / avoid-TTTT rules
                      └── Validate experimentally (Sanger / T7E1 / amplicon-seq)

| Situation | Recommended tier | Rationale | |-----------|------------------|-----------| | Common human/mouse gene with prior CRISPR publications | Tier 1 (Addgene + literature) | Validated sequences come with measured cutting efficiency; lowest experimental risk | | Genome-scale screen of a model organism | Tier 2 (CRISPick) | Pre-computed datasets cover whole genomes with consistent scoring | | Single human gene knockout, no Addgene hit | Tier 2 (CRISPick GRCh38 SpCas9 CRISPRko), filter by Combined Rank | Best balance of efficiency and specificity for one-off knockouts | | CRISPRa / CRISPRi experiment | Tier 2 with the matching CRISPRa / CRISPRi dataset | Activation/inhibition models target proximal-promoter windows, not coding exons | | Cas12a (AsCas12a vs enAsCas12a) | Tier 2 with the exact Cas12a variant dataset | Guides for one variant are not interchangeable with the other | | Non-model organism not covered by CRISPick | Tier 3 (de novo rules) | No high-throughput training data exists; rule-based design + experimental validation | | Small custom locus (e.g., engineered cassette) | Tier 3 | Locus is not in any reference genome; design directly against the target sequence | | Maximum specificity required (e.g., therapeutic application) | Tier 2 sorted by Off-Target Rank | Prioritize guides with the fewest predicted off-target sites | | Maximum cutting efficiency required (e.g., difficult cell line) | Tier 2 sorted by On-Target Rank | Accept slightly higher off-target risk in exchange for activity |

Best Practices

  1. Always exhaust Tier 1 before moving to Tier 2. Even if addgene_grna_sequences.csv returns no rows, run the literature search step (Method 2). Many validated sgRNAs are published in supplementary materials but never deposited at Addgene.
  2. Match the CRISPick dataset to the exact Cas variant. AsCas12a and enAsCas12a are distinct enzymes; SpCas9 and SaCas9 require different PAMs. Using the wrong dataset wastes experimental resources.
  3. Use Combined Rank as the default selector. It balances on-target efficiency and off-target specificity; reach for On-Target Rank or Off-Target Rank only when the experiment has a specific reason to prioritize one.
  4. Order 3–4 sgRNAs per target gene, ideally from different exons. A single guide can fail for reasons unrelated to its predicted score (chromatin context, secondary structure, PCR drop-out). Redundancy is cheap insurance.
  5. For knockouts, target the first 50% of the coding sequence. N-terminal cuts are most likely to produce a true loss-of-function via NMD; C-terminal cuts can leave partially functional protein.
  6. For CRISPRa, design within −200 to +1 bp of the TSS; for CRISPRi, within −50 to +300 bp. Effector activity drops sharply outside these windows.
  7. Validate experimentally. Even high-ranked guides fail ~20–30% of the time. Confirm editing with Sanger sequencing, T7E1, ICE, or amplicon-seq before committing to phenotype assays.
  8. Cite the original validation paper for Tier 1, and CRISPick + Sanson 2018 / DeWeirdt 2021 for Tier 2. Reproducibility depends on traceable provenance for every guide sequence.

Common Pitfalls

  • Pitfall: Skipping the literature search step (Method 2) when the Addgene CSV returns no hits.

    • How to avoid: Treat Method 2 as mandatory; many validated guides only appear in published supplementary materials.
  • Pitfall: Using AsCas12a guides with enAsCas12a (or vice versa) because both share TTTV PAM.

    • How to avoid: Match the CRISPick dataset filename's Cas tag (AsCas12a vs enAsCas12a) to the exact enzyme variant used at the bench.
  • Pitfall: Sorting by On-Target Rank only and ending up with high-activity, low-specificity guides that produce off-target editing.

    • How to avoid: Use Combined Rank as the default; only deviate when the experiment explicitly requires extreme efficiency or specificity.
  • Pitfall: Designing a single sgRNA per gene and concluding "the guide doesn't work" when one fails.

    • How to avoid: Always order and test 3–4 guides per gene from independent loci/exons.
  • Pitfall: Targeting late exons or 3' UTRs for knockout. Truncated proteins from late-exon edits are often partially functional and confound phenotype interpretation.

    • How to avoid: For knockouts, restrict candidates to the first 50% of the coding sequence (or the first 3 exons).
  • Pitfall: Using a CRISPRko dataset for a CRISPRa or CRISPRi experiment. The protospacer windows are different — coding exons for knockout, proximal promoter for activation/inhibition.

    • How to avoid: Download the dataset whose application tag (CRISPRko / CRISPRa / CRISPRi) matches the experiment.
  • Pitfall: Designing guides containing TTTT. This sequence terminates RNA Pol III transcription, so the sgRNA simply will not be expressed from a U6 promoter.

    • How to avoid: Always filter de novo designs (and double-check Tier 1/2 picks) against TTTT and homopolymer runs >4.
  • Pitfall: Failing to cite the original publication for a validated sgRNA.

    • How to avoid: Record the PubMed_ID from the Addgene CSV (or the DOI from the literature search) as part of the design record and include it in the methods section.

Option 1: Search Validated sgRNA Sequences (Recommended First)

1.1 Search for Validated sgRNAs

IMPORTANT: You MUST complete BOTH Method 1 AND Method 2 before proceeding to Option 2. Do not skip Method 2 even if Method 1 finds no results.

Method 1: Search Our Database (Fastest)

We maintain a curated database of 300+ validated sgRNA sequences from Addgene with experimental evidence.

Location: resource/addgene_grna_sequences.csv (relative to this skill directory)

Search the database:

import pandas as pd

# Load the database
df = pd.read_csv('addgene_grna_sequences.csv')

# Search for your gene
gene_name = "TP53"
results = df[df['Target_Gene'].str.upper() == gene_name.upper()]

# Filter by species and application
results_filtered = results[
    (results['Target_Species'] == 'H. sapiens') &
    (results['Application'] == 'cut')  # or 'activate', 'RNA targeting'
]

# Display results with references
print(results_filtered[['Target_Gene', 'Target_Sequence',
                        'Plasmid_ID', 'PubMed_ID', 'Depositor']])

Database columns:

  • Target_Gene: Gene symbol
  • Target_Species: Organism (H. sapiens, M. musculus, etc.)
  • Target_Sequence: 20bp sgRNA sequence (5' to 3')
  • Application: cut (knockout), activate (CRISPRa), RNA targeting (CRISPRi)
  • Cas9_Species: S. pyogenes, S. aureus, etc.
  • Plasmid_ID: Addgene plasmid number
  • Plasmid_URL: Direct link to plasmid page
  • PubMed_ID: Publication reference (cite this in your work)

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars367
CategoryDevelopment
Updated1mo ago
Forks36

Languages

Python

Trust signals

88/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

1 medium