It's of huge value to be able to predict CRISPR target efficiency ahead of time. Xu et al have published an analysis of multiple guide RNA data sets and extracted what they claim is an improved model for target cleavage efficiency prediction. This data is all for the S.pyogenes native Cas9.
Xu et al. Sequence determinants of improved CRISPR sgRNA design.
Genome Res. 2015 Aug;25(8):1147-57. doi: 10.1101/gr.191452.115. Epub 2015 Jun 10.
Their paper is important to me for several reasons. First, they have examined two independently-published "large" guide RNA data sets that had mutagenesis-efficiency data, which allows more confidence that trends of sequence preferences are holding up across labs and platforms. Second, they validated their predictive model on a small (in comparison to genome-wide, but still not bad) data set of new CRISPR targets and corresponding guide RNAs. Third, they did "in silico validation" by turning their model loose on another target/indel data set, and showed improved performance of their predictive model over a previously published model. See ROC curves in Fig. 4b. This allows an ability to weed out "50-60% of the inefficient sgRNAs…at the cost of 10-20% of efficient sgRNAs misclassified." That is, misclassified as inefficient.
For those who are interested in genome-wide knockout screening experiments these sorts of models are very good for increasing efficiency of the screens. Moreover, if you wish to knockout particular genes, it will allow you to test or use fewer targets per gene till you find one that works well.
OK, now the sobering reality for nerds like me is that predictive models, even with great ROC curves, have false positive and false negative rates that will bite you in the behind on a regular basis if you are designing large projects around the function of single CRISPR targets. I'm still facing this issue for precision knock-in projects, for which there are often not many targets to choose from. And with transgenic mice we always want the efficiency as high as possible. For cell lines, hey, that's not as much a problem if you can subclone the edited lines.
But let's get back to the CRISPR target sequence preferences. The bottom line here is that the last three bases of the protospacer seem to have the most influence on cleavage efficiency, with a C preferred at the -3 position (relative to the PAM), and G's at -2 and -1. Also, G's are helpful at the -17 to -14 region, while A's are good at the -12 to -9 region. Finally, a C seems helpful at +1 following the PAM.
Looking back at the Wang et al paper, they also reported a preference for A's at around -10 to -8, and essentially a "GCRR" preference for bases -4 to -1. This makes sense since Xu have based their model partly on the the Wang data. However, Xu et al point out that the apparent G preference at the -20 position is probably an artifact of the Wang sgRNA library in that these may have had increased efficiency due to enhanced transcription, not activity per se.
General GC-richness in the protospacer is known to correlate with CRISPR mutagenesis. Could that just be driven by the GC-rich preferences of the last few bases? Otherwise, GC-richness doesn't clearly emerge from the Xu model, at least to me anyway. I took a crack at this by looking at a data set from Gagnon et al, mostly because I could handle the size of their sgRNA list in an excel spreadsheet without exploding my own brain or my iMac. My impression is that GC richness is still "good" even when the last 4 bases of the protospacer are similar. Here's an example. From Gagnon et al's list of 122 sgRNAs with indel numbers, I ranked them according to how well they matched the "GCRR" of the last four bases. I based this on the Wang et al paper although I think it is very similar to that corresponding part of the Xu model. My "score" ranged from 0 to 7. Then I examined the subset of 30 targets that all had a same "score" of 5. So these targets are all controlled, at least kinda sorta, for their variation in bases -4 to -1 in that they have similar strength of matching to the "GCRR" motif. Finally, I graphed the indel frequencies versus the GC content of the first 16 bases of their protospacers. Here is the data. y axis= indel frequency (in a zebrafish model), x axis= # of GC base pairs in first 16 bases.
This ain't close to something I'd submit for peer review but I do see a trend. GC richness in the first 16 bases of the protospacer correlates with cleavage efficiency, even within a group of targets for which the 3' ends are similar. So for now - I will continue to prefer overall GC-rich targets that also have at least some matching to the "CGG", or "GCRR", motif at the very 3' end.
Also, "CGG" matches the high-efficiency 3' end reported by Farboud and Meyer so there's another corroboration.
So, the answer to my previous post "Are there sequence preferences near the 3' end of the #CRISPR protospacer? …" is, yes. And this holds up for S.pyogenes Cas9 when used across human, mouse, fish and C.elegans models.
A final note - these data all refer to cleavage and/or knockout efficiencies. CRISPRi and CRISPRa screens, which do not lead to or require DNA cleavage, have different sequence preferences which Xu et al also modeled in detail.
Happy CRISPRing.
New developments in CRISPR technology, with a focus on mouse and human cell applications.
Showing posts with label sgRNA. Show all posts
Showing posts with label sgRNA. Show all posts
Monday, August 17, 2015
Wednesday, May 20, 2015
About using DNA or RNA for mouse embryo #CRISPR injections.
I got a question:
Isn't the disadvantage of injecting DNA the threat of integration and more frequent mosaicism than in the case of RNA as Cas is expressed quicker? Do you have some direct experience with that? Thanks!
Um, well yes. Yes. Those are the disadvantages. Also I will add that because the RNA should lead to quicker Cas9 expression, mutagenesis efficiencies will likely be higher than with DNA vectors.
So why use DNA at all? Well, the issues are mostly practical. DNA vectors are easy to customize for CRISPR. Although the issues of efficiency and mosaicism are potentially problematic, I have seen pretty consistent success* in generating simple indel mutations following injections of PX330-style CRISPR-Cas9 DNA plasmids. That is, consistent double digit percentages of founders carrying mutations as assayed by PCR and/or sequencing. In addition, in our core we have obtained HDR-mediated codon editing rates in the 10-20% range using PX330 vectors co-injected with appropriate "donor" oligos. But this is dependent on cooperative CRISPR sites that have a high rate of baseline cleavage.
Another practical consideration is that not everyone can routinely synthesize high-quality RNAs in vitro with consistency. Quality DNA is relatively easy to prepare and QC. RNA is much less so - especially for the 4+ kilobase Cas9 mRNA. OK, so some of you are saying "Come one, my lab makes RNAs all the time - no prob! " . That's awesome, but the empirical observation is that it's not trivial to get proficient at making long mRNAs, and to keep on top of the key reagent issues (RNAses, enzymes going bad, etc.).
Also, CRISPR DNA plasmids are immediately useful for cell culture gene editing experiments. Some labs will be making these anyway so they will have them on hand, ready to go.
What I am also observing - which many others have reported - is that a fraction of CRISPR sites just don't cut very well, even when the sequence characteristics of the site seem OK. (Like, somewhere on the order of 1/3 to 1/4 of CRISPR target sites?) Most of our injections to date have been using DNA plasmids. It's possible that RNAs might save the day for some of these sites.
Isn't the disadvantage of injecting DNA the threat of integration and more frequent mosaicism than in the case of RNA as Cas is expressed quicker? Do you have some direct experience with that? Thanks!
Um, well yes. Yes. Those are the disadvantages. Also I will add that because the RNA should lead to quicker Cas9 expression, mutagenesis efficiencies will likely be higher than with DNA vectors.
So why use DNA at all? Well, the issues are mostly practical. DNA vectors are easy to customize for CRISPR. Although the issues of efficiency and mosaicism are potentially problematic, I have seen pretty consistent success* in generating simple indel mutations following injections of PX330-style CRISPR-Cas9 DNA plasmids. That is, consistent double digit percentages of founders carrying mutations as assayed by PCR and/or sequencing. In addition, in our core we have obtained HDR-mediated codon editing rates in the 10-20% range using PX330 vectors co-injected with appropriate "donor" oligos. But this is dependent on cooperative CRISPR sites that have a high rate of baseline cleavage.
Another practical consideration is that not everyone can routinely synthesize high-quality RNAs in vitro with consistency. Quality DNA is relatively easy to prepare and QC. RNA is much less so - especially for the 4+ kilobase Cas9 mRNA. OK, so some of you are saying "Come one, my lab makes RNAs all the time - no prob! " . That's awesome, but the empirical observation is that it's not trivial to get proficient at making long mRNAs, and to keep on top of the key reagent issues (RNAses, enzymes going bad, etc.).
Also, CRISPR DNA plasmids are immediately useful for cell culture gene editing experiments. Some labs will be making these anyway so they will have them on hand, ready to go.
What I am also observing - which many others have reported - is that a fraction of CRISPR sites just don't cut very well, even when the sequence characteristics of the site seem OK. (Like, somewhere on the order of 1/3 to 1/4 of CRISPR target sites?) Most of our injections to date have been using DNA plasmids. It's possible that RNAs might save the day for some of these sites.
The ability to do precise HDR-mediated editing/insertions, rather than simple indels, is very compelling and is the direction most of our CRISPR ideas are going in terms of new mouse models. But coding modifications usually have extremely narrow CRISPR target choices that are imposed by the science; if you want to change a codon, you'll probably need a target as close as possible - preferably overlapping the codon. There won't be many to choose from. So getting the highest efficiency cleavage rates may be critical for some of these projects - for these, Cas9 mRNA or protein may be needed.
Finally, these issues of target efficiency really call for pre-validation of sites. This can be done by transfecting CRISPR plasmids into cooperative cell lines, e.g. NIH3T3 for mouse targets, followed by PCR and mismatch cleavage assays, which can then be quantified. But then - if you go through the trouble to do that, you will have generated the DNA plasmids and thus have the DNA reagent ready for injection.
Having said all that, although I really like the convenience of plasmids, the RNA problems are all about sourcing them. A few vendors, such as Sigma-Aldrich can provide custom guide RNAs and Cas9 mRNA that work. (FYI I do not receive any compensation from Sigma). The RNA reagent expense is less than the cost of mouse embryo injections. I suppose zebrafish researchers may balk at the cost, as they will have more capacity to inject fish eggs, in their own labs usually, and may be more willing to make RNAs in-house. For mice, you'll be usually working with a transgenic core and spending thousands of bucks per experiment. Vendor-supplied RNAs may be worth the money. Thanks for the question!
*Actually, "consistent" may be misleading… To clarify, about 75% of the NHEJ projects I've been observing have had this level of success. So - more success than not, but then again, not perfectly consistent.
*Actually, "consistent" may be misleading… To clarify, about 75% of the NHEJ projects I've been observing have had this level of success. So - more success than not, but then again, not perfectly consistent.
Labels:
assay,
cleavage,
HDR,
optimization,
plasmid,
pronuclear injection,
PX330,
RNA,
RNA vs DNA,
sgRNA
Monday, October 20, 2014
More protocols for mouse mutagenesis with #CRISPR: newly published in CPHG.
Harms et al have created this detailed set of protocols for conducting CRISPR/Cas9 mutagenesis and editing in mouse embryos. Included are more guidelines and instructions for choosing targets, then on to ligation protocols for cloning protospacers in sgRNA expression vectors, in vitro RNA synthesis, pronuclear injection, and follow-up screening/genotyping of founder animals, as well as HDR donor design and considerations. Also a timeline schematic.
Mouse Genome Editing Using the CRISPR/Cas System.Harms DW, Quadros RM, Seruggia D, Ohtsuka M, Takahashi G, Montoliu L, Gurumurthy CB.Curr Protoc Hum Genet. 2014 Oct 1;83:15.7.1-15.7.27. doi: 10.1002/0471142905.hg1507s83.
Mouse Genome Editing Using the CRISPR/Cas System.Harms DW, Quadros RM, Seruggia D, Ohtsuka M, Takahashi G, Montoliu L, Gurumurthy CB.Curr Protoc Hum Genet. 2014 Oct 1;83:15.7.1-15.7.27. doi: 10.1002/0471142905.hg1507s83.
Labels:
adapters,
protocol,
protospacer,
sgRNA,
target
Wednesday, September 10, 2014
Paper: More details about what sequence features make a CRISPR target highly "cuttable".
This was interesting for several nice reasons. First, my interpretation is that this suggests that while most CRISPR targets have some level of acceptable targeting by Cas9/sgRNA, there is a subset of sites that are *very* susceptible. Therefore, if you want to make a null allele in a gene it's a good idea to score all the possible sites first with this sort of algorithm.
Rational design of highly active sgRNAs for CRISPR-Cas9-mediated gene inactivation. Doench et al, Nat Biotechnol. 2014 Sep 3. doi: 10.1038/nbt.3026. [Epub ahead of print].
Second, they of course have made their scoring algorithm available as a web tool: http://www.broadinstitute.org/rnai/public/analysis-tools/sgrna-design
I took a crack at it with a small DNA sequence from mouse that I know has a decent CRISPR target. Interestingly, my pre-validated target had a lousy score (~0.05 on a scale from 0 to 1) despite my knowledge that it works pretty well in my hands. I take this to mean that of course, no scoring algorithm is perfect, but also that most targets actually fall into the so-so category of activity. This jibes with data in this paper somewhat and I won't jump to conclusions based on my N=1! I'd be interested to hear other people's results from running their targets through this algorithm.
Rational design of highly active sgRNAs for CRISPR-Cas9-mediated gene inactivation. Doench et al, Nat Biotechnol. 2014 Sep 3. doi: 10.1038/nbt.3026. [Epub ahead of print].
Second, they of course have made their scoring algorithm available as a web tool: http://www.broadinstitute.org/rnai/public/analysis-tools/sgrna-design
I took a crack at it with a small DNA sequence from mouse that I know has a decent CRISPR target. Interestingly, my pre-validated target had a lousy score (~0.05 on a scale from 0 to 1) despite my knowledge that it works pretty well in my hands. I take this to mean that of course, no scoring algorithm is perfect, but also that most targets actually fall into the so-so category of activity. This jibes with data in this paper somewhat and I won't jump to conclusions based on my N=1! I'd be interested to hear other people's results from running their targets through this algorithm.
Subscribe to:
Posts (Atom)
