Showing posts with label Cas9. Show all posts
Showing posts with label Cas9. Show all posts

Thursday, January 28, 2016

For #CRISPR HDR, use donor oligos that are complementary to the "gRNA strand". A new paper shows why; see my blog post.

(ERRATUM:  I made on correction to this post on March 11 2016.  In the original version of the post I stated that the paper implied that the donor oligo should have "additional length of homology on the PAM-distal side as compared to the PAM-proximal side".  That was a mistake - it turns out the opposite was true.   The authors found that additional length of homology on the PAM-proximal side was favorable.  I got confused because the PAM in figure 3 is on the "bottom" strand, not the top, so the PAM-proximal side is to the left of the cut site in their oligo schematics.  Figure 3 has an "upside down" Cas9 icon in keeping with this fact.  Thank you Scot for letting me know about the error!).


After a long break from blogging... here's a nice nugget of insight about CRISPR-mediated homologous recombination.    Back in late 2014 I blogged about 
Optimal design of ssODNs (donor oligos) for #CRISPR - length and strandedness data? . In that post I pointed out a curious observation that HDR oligos work "better" when they are designed from the strand that is complementary to the protospacer/gRNA sequence.   This was somewhat counterintuitive to me, as one might think that in a complementary HDR oligo would tend to anneal to the gRNA, reducing its availability or kinetics somewhat and generally interfering with Cas9's job.   But empirically, this was not the case; complementary oligos work better.

Now, Richardson et al seem to have found an explanation. (Richardson et al, Enhancing homology-directed genome editing by catalytically active and inactive CRISPR-Cas9 using asymmetric donor DNA. Nature Biotechnology (2016, Published online 20 January 2016)). Turns out that following double-stranded DNA cleavage by Cas9, the first component of the DNA molecule that it releases is the 3' end of the DNA strand that corresponds to the protospacer/gRNA.  Quote: Hence, although Cas9 globally dissociates from duplex DNA in a symmetric fashion (Fig. 1c), it appears that the enzyme locally releases the PAM-distal nontarget strand after cleavage but before dissociation."  And here is a nice diagram from the supplemental material of the paper that shows how this strand "breathes" after cleavage.  I thank the senior author, Jacob Corn, for graciously allowing me to reproduce the image here:




(OK - before moving forward, let's get clarity on the terms here; when we're comparing the two strands of the DNA containing the CRISPR target, the "target" strand is the DNA strand that directly will anneal to the gRNA.  Thus, the "target strand" is complementary to the gRNA.  The non-target DNA strand encodes the protospacer and the "NGG" PAM sequence.   Got it? )   

Read the above quote again.  The first bit of DNA that is released by Cas9 is the single-stranded 3' end of the non-target strand.  This immediately suggests 2 things:

1.  Cas9 releases the non-target strand before the target strand.  Thus the non-target strand is available sooner than the target strand to potentially engage with a complementary donor molecule and jump-start homologous recombination.      

2.  The 3' end of the non-target strand, which is "PAM-distal", is released first.  This also suggests that the design considerations for homology might be different for the PAM-distal and PAM-proximal sides of the cleavage location.

For this second point, the practical consideration is that commercially available ssDNA oligos are usually limited to 200 bases or less (depending on the vendor's capabilities; 120 base oligos can work well too).  So, we are limited in the length of homology we can actually apply to each side of the cleavage site.   Instead of centering the oligo (e.g. for a 120 base HDR oligo = ~60 bases of homology on each side) it may be better to skew the oligo design to have more homology on one side of the cut site versus the other.  That is what Richardson et al found - at least for one target they investigated in detail  (See Fig. 3c, d, e.)

Here's another thought. The authors note that Cas9 actually stays on the DNA for quite a while after it cleaves both strands - about 5 hours.    This might have something to do with why DNA repair takes significantly longer on Cas9-cleaved breaks than on breaks induced with radiation - Cas9 may just sitting there, sterically hindering the DNA repair proteins from accessing the free DNA ends at the break.  Perhaps, CRISPR mutagenesis efficiencies could be further enhanced by increasing Cas9's intrinsic off-rate?  On this note, another recent paper by Kleinstiver et al (Keith Joung lab) suggests that directed mutations can destabilize Cas9's non-specific interaction with DNA.   The authors note that this reduces off-target cleavage significantly while preserving on-target cleavage .   While they did not see dramatic increases in on-target efficiency, perhaps in some experimental contexts there might be?    Hmm.    

Monday, August 17, 2015

#CRISPR target sequence preferences are being clarified. Xu et al Genome Research paper.

It's of huge value to be able to predict CRISPR target efficiency ahead of time.  Xu et al have published an analysis of multiple guide RNA data sets and extracted what they claim is an improved model for target cleavage efficiency prediction.   This data is all for the S.pyogenes native Cas9. 

Xu et al. Sequence determinants of improved CRISPR sgRNA design.
Genome Res. 2015 Aug;25(8):1147-57. doi: 10.1101/gr.191452.115. Epub 2015 Jun 10.

Their paper is important to me for several reasons.  First, they have examined two independently-published "large" guide RNA data sets that had mutagenesis-efficiency data, which allows more confidence that trends of sequence preferences are holding up across labs and platforms.   Second, they validated their predictive model on a small (in comparison to genome-wide, but still not bad) data set of new CRISPR targets and corresponding guide RNAs.  Third, they did "in silico validation" by turning their model loose on another target/indel data set, and showed improved performance of their predictive model over a previously published model.   See ROC curves in Fig. 4b.    This allows an ability to weed out "50-60% of the inefficient sgRNAs…at the cost of 10-20% of efficient sgRNAs misclassified."  That is, misclassified as inefficient.    

For those who are interested in genome-wide knockout screening experiments these sorts of models are very good for increasing efficiency of the screens.   Moreover, if you wish to knockout particular genes, it will allow you to test or use fewer targets per gene till you find one that works well.

OK, now the sobering reality for nerds like me is that predictive models, even with great ROC curves, have false positive and false negative rates that will bite you in the behind on a regular basis if you are designing large projects around the function of single CRISPR targets.  I'm still facing this issue for precision knock-in projects, for which there are often not many targets to choose from.   And with transgenic mice we always want the efficiency as high as possible.   For cell lines, hey, that's not as much a problem if you can subclone the edited lines.

But let's get back to the CRISPR target sequence preferences.  The bottom line here is that the last three bases of the protospacer seem to have the most influence on cleavage efficiency, with a C preferred at the -3 position (relative to the PAM), and G's at -2 and -1.    Also, G's are helpful at the -17 to -14 region, while A's are good at the -12 to -9 region.  Finally, a C seems helpful at +1 following the PAM.

Looking back at the Wang et al paper, they also reported a preference for A's at around -10 to -8, and essentially a "GCRR" preference for bases -4 to -1.  This makes sense since Xu have based their model partly on the the Wang data.   However, Xu et al point out that the apparent G preference at the -20 position is probably an artifact of the Wang sgRNA library in that these may have had increased efficiency due to enhanced transcription, not activity per se.

General GC-richness in the protospacer is known to correlate with CRISPR mutagenesis.  Could that just be driven by the GC-rich preferences of the last few bases?   Otherwise, GC-richness doesn't clearly emerge from the Xu model, at least to me anyway.  I took a crack at this by looking at a data set from Gagnon et al, mostly because I could handle the size of their sgRNA list in an excel spreadsheet without exploding my own brain or my iMac.  My impression is that GC richness is still "good" even when the last 4 bases of the protospacer are similar.   Here's an example.  From Gagnon et al's list of 122 sgRNAs with indel numbers, I ranked them according to how well they matched the "GCRR" of the last four bases.  I based this on the Wang et al paper although I think it is very similar to that corresponding part of the Xu model.   My  "score" ranged from 0 to 7.  Then I examined the subset of 30 targets that all had a same "score" of 5.   So these targets are all controlled, at least kinda sorta, for their  variation in bases -4 to -1 in that they have similar strength of matching to the "GCRR" motif.  Finally, I graphed the indel frequencies versus the GC content of the first 16 bases of their protospacers.   Here is the data.  y axis= indel frequency (in a zebrafish model), x axis= # of GC base pairs in first 16 bases.

This ain't close to something I'd submit for peer review but I do see a trend.  GC richness in the first 16 bases of the protospacer correlates with cleavage efficiency, even within a group of targets for which the 3' ends are similar.   So for now - I will continue to prefer overall GC-rich targets that also have at least some matching to the "CGG", or "GCRR", motif at the very 3' end.  

Also, "CGG" matches the high-efficiency 3' end reported by Farboud and Meyer so there's another corroboration.

So, the answer to my previous post "Are there sequence preferences near the 3' end of the #CRISPR protospacer? …" is, yes.  And this holds up for S.pyogenes Cas9 when used across human, mouse, fish and C.elegans models.   

A final note - these data all refer to cleavage and/or knockout efficiencies.  CRISPRi and CRISPRa screens, which do not lead to or require DNA cleavage, have different sequence preferences which Xu et al also modeled in detail.   

Happy CRISPRing.

Monday, July 6, 2015

My review of recent Joung lab paper with new PAM specificities! for Cas9: will broaden choice of #CRISPR targets.

It was just a matter of time before someone mutagenized Cas9 to try to change its PAM preferences.  (What's a PAM? Look here if you aren't hip yet).   Although that "NGG" motif is pretty abundant in genomes - heck, it's only 2 bases - it sure would be nice to be able to target even more sequences with high specificity.  For example, the closer the CRISPR site is to the site where precise genome editing is required, the more efficient it will (probably) be;  having more PAM choices will only be helpful in this situation.  But, having more targets isn't very useful unless the properties of CRISPR specificity and sensitivity remain robust.  

So now the Joung lab has led the way with efforts to coax Cas9 into preferring new PAMs, as described in Kleinstiver et al's new paper in Nature.    I really like this paper.   It introduces new Cas9 variants with preferences for NGA, or NGCG PAMs. The new variants are not complicated to engineer, maintain high cleavage activity, work in vivo, and have low off-target effects similar to native Cas9.  

Previously, Anders et al had published a Nature paper describing Cas9-DNA structural interaction.  (Senior author of this paper was Martin Jinek, who was first author of the seminal 2012 Dounda/Charpentier Science paper).  An interesting nugget in that paper was a first attempt to alter Cas9 specificity by mutating the two amino acids (Arg1333 and Arg 1335), which apparently interact directly with the two guanine bases of the PAM motif, to glutamine.  This was inspired by Cas9 variants from non-S.pyogenes species which prefer A-rich PAMs and have glutamines in the homologous positions.   However these changes alone couldn't make S.pyogenes' Cas9 prefer NAA instead of NGG.

Enter Kleinstiver et al.  They began a systematic attempt to engineer new PAM specificities into S.pyogenes Cas9.  First, they used a clever bacterial assay to measure PAM preference in which the CRISPR target is within a toxic gene.  Cas9 variants with different codon changes were then introduced.  In this setup, target cleavage disrupts the gene and allows the bugs to grow, allowing one to sequence the survivors to figure out which Cas9 variants worked.  In this manner they identified combinations of codon changes that allowed Cas9 to recognize a NGA PAM.  2 variant combos, "VQR" and "EQR" , emerged as being best at now preferring NGA over NGG.

Then, they used a different assay to measure preference for all the possible different PAMs for selected Cas9 mutants.  See Figs. 1e, 1f for these data nicely visualized.  For example you can clearly see how wild type Cas9 greatly prefers NGG, but has a little ability to use NAG as had been previously reported by many - in fact many off-target analyses consider NAG PAMs as well as NGGs.   Then, they tested the VQR and EQR variants, which revealed that these now are sensitive to the fourth base in the PAM.  Specifically, Cas9-VQR "likes" NGAG, NGAA, NGAT, NGCG the best.    Interestingly, Cas9-EQR preferred NGAG almost exclusively.  The authors concluded that the T1337R variant is a gain-of-function allowing sensitivity to the fourth base, which is then specified by other codon variants.  Cool.

Next, they found that the "VRER" combination allowed specific preference for a NGCG PAM.  Note that this GCG motif is much less common in mammalian DNA than the other PAMs - after all, it's got a CG in it - but that also means it's off-target potential is lower.   Since I know lots of genes with GC-rich regions I'm betting this PAM will come in handy.

A few more points from the paper:   The new variants work in zebrafish in human cells, and have good activity and low off-target effects.  Additionally, they noticed that the D1135E variant actually increased PAM specificity for the wild-type NGG PAM relative to NAG - see Fig 3a, and furthermore it reduced off-target cleavage on other off-target sites that have mismatches in protospacer but (presumably) the NGG PAM.  In other words D1135E reduces off-target cleavage in general (at least somewhat) without reducing on-target cleavage.   Sounds good to me!

Finally, they examined two Cas9 genes from other bacterial species and showed they could carefully measure their normal spectra of PAM preferences (which are different from the NGG of S.pyogenes).  In other words they are poised to do the same mutagenesis work on these other Cas9 proteins, which will add even more PAM choices to the toolkit.   

i was about to write "these Cas9 variants should be widely available soon", then I thought Hmm, better check Addgene.   Sure enough:   VQR, EQR, and VRER expression plasmids are already available!  Kudos to Keith Joung and his lab for making these available to the world.  Happy CRISPRing with new PAMs!