Showing posts with label cleavage. Show all posts
Showing posts with label cleavage. Show all posts

Thursday, January 28, 2016

For #CRISPR HDR, use donor oligos that are complementary to the "gRNA strand". A new paper shows why; see my blog post.

(ERRATUM:  I made on correction to this post on March 11 2016.  In the original version of the post I stated that the paper implied that the donor oligo should have "additional length of homology on the PAM-distal side as compared to the PAM-proximal side".  That was a mistake - it turns out the opposite was true.   The authors found that additional length of homology on the PAM-proximal side was favorable.  I got confused because the PAM in figure 3 is on the "bottom" strand, not the top, so the PAM-proximal side is to the left of the cut site in their oligo schematics.  Figure 3 has an "upside down" Cas9 icon in keeping with this fact.  Thank you Scot for letting me know about the error!).


After a long break from blogging... here's a nice nugget of insight about CRISPR-mediated homologous recombination.    Back in late 2014 I blogged about 
Optimal design of ssODNs (donor oligos) for #CRISPR - length and strandedness data? . In that post I pointed out a curious observation that HDR oligos work "better" when they are designed from the strand that is complementary to the protospacer/gRNA sequence.   This was somewhat counterintuitive to me, as one might think that in a complementary HDR oligo would tend to anneal to the gRNA, reducing its availability or kinetics somewhat and generally interfering with Cas9's job.   But empirically, this was not the case; complementary oligos work better.

Now, Richardson et al seem to have found an explanation. (Richardson et al, Enhancing homology-directed genome editing by catalytically active and inactive CRISPR-Cas9 using asymmetric donor DNA. Nature Biotechnology (2016, Published online 20 January 2016)). Turns out that following double-stranded DNA cleavage by Cas9, the first component of the DNA molecule that it releases is the 3' end of the DNA strand that corresponds to the protospacer/gRNA.  Quote: Hence, although Cas9 globally dissociates from duplex DNA in a symmetric fashion (Fig. 1c), it appears that the enzyme locally releases the PAM-distal nontarget strand after cleavage but before dissociation."  And here is a nice diagram from the supplemental material of the paper that shows how this strand "breathes" after cleavage.  I thank the senior author, Jacob Corn, for graciously allowing me to reproduce the image here:




(OK - before moving forward, let's get clarity on the terms here; when we're comparing the two strands of the DNA containing the CRISPR target, the "target" strand is the DNA strand that directly will anneal to the gRNA.  Thus, the "target strand" is complementary to the gRNA.  The non-target DNA strand encodes the protospacer and the "NGG" PAM sequence.   Got it? )   

Read the above quote again.  The first bit of DNA that is released by Cas9 is the single-stranded 3' end of the non-target strand.  This immediately suggests 2 things:

1.  Cas9 releases the non-target strand before the target strand.  Thus the non-target strand is available sooner than the target strand to potentially engage with a complementary donor molecule and jump-start homologous recombination.      

2.  The 3' end of the non-target strand, which is "PAM-distal", is released first.  This also suggests that the design considerations for homology might be different for the PAM-distal and PAM-proximal sides of the cleavage location.

For this second point, the practical consideration is that commercially available ssDNA oligos are usually limited to 200 bases or less (depending on the vendor's capabilities; 120 base oligos can work well too).  So, we are limited in the length of homology we can actually apply to each side of the cleavage site.   Instead of centering the oligo (e.g. for a 120 base HDR oligo = ~60 bases of homology on each side) it may be better to skew the oligo design to have more homology on one side of the cut site versus the other.  That is what Richardson et al found - at least for one target they investigated in detail  (See Fig. 3c, d, e.)

Here's another thought. The authors note that Cas9 actually stays on the DNA for quite a while after it cleaves both strands - about 5 hours.    This might have something to do with why DNA repair takes significantly longer on Cas9-cleaved breaks than on breaks induced with radiation - Cas9 may just sitting there, sterically hindering the DNA repair proteins from accessing the free DNA ends at the break.  Perhaps, CRISPR mutagenesis efficiencies could be further enhanced by increasing Cas9's intrinsic off-rate?  On this note, another recent paper by Kleinstiver et al (Keith Joung lab) suggests that directed mutations can destabilize Cas9's non-specific interaction with DNA.   The authors note that this reduces off-target cleavage significantly while preserving on-target cleavage .   While they did not see dramatic increases in on-target efficiency, perhaps in some experimental contexts there might be?    Hmm.    

Monday, September 28, 2015

Move over, Cas9: Cpf1 may be your new #CRISPR competition.

This past weekend at the Pilgrimage music festival in Franklin, TN I had the pleasure of seeing Weezer rock out, then I walked a few hundred yards to another stage to watch Wilco do the same.  These bands each had their own stage, but in the CRISPR world, Cas9 is a superstar that now has to share the stage with a newcomer:  Cpf1.  (Yeah, I know that's a goofy setup but it really was a good festival and it was on my mind.  Now on to the science.)

Last Friday Feng Zhang’s group published a paper in Cell that immediately grabbed a lot of attention, and rightly so.  They reveal that the Cpf1 class of CRISPR effector proteins may be an attractive alternative to Cas9.   Although Cpf1 has many similarities to Cas9, it has some significant differences that are very interesting – and could lead to improved efficiencies for some types of gene editing.  

Cpf1 Is a Single RNA-Guided Endonuclease of a Class 2 CRISPR-Cas System.  Zetsche et al., 2015, Cell 163, 1–13October 22, 2015  (Avail. online Sept. 25 2015).

Here is my summary…First, they reviewed some background on CRISPR systems in bacteria to cover the basis of the study.   There are two major classes of CRISPR systems based mainly on the proteins involved; the cleavage effectors of class 1 are complexes of multiple proteins, while the class 2 effectors are single proteins like Cas9.  Within class 2 there are two subtypes of systems: those with Cas9, and another that has Cpf1.   Cpf1 means “CRISPR from Prevotella and Francisella 1”.  Some bacterial species carry both Cas9-CRISPR and Cpf1-CRISPR loci in their genomes.   Like Cas9, Cpf1 has RuvC-like DNA cleavage domains but it lacks some of the other domain and neighbor-gene features of Cas9, so it’s clearly distinct in its evolution.    It’s a largish protein of ~1300 amino acids, similar in size to Cas9.

They picked the Cpf1-CRISPR gene system of Francisella novicida strain U112 to study first since there were clear homologies of Cpf1-CRISPR spacer sequences to various prophage in this species – further suggesting that Cpf1 is important in bacterial immunity and so it’s well adapted to slice and dice target DNAs.   (Sidebar: what’s F. novicida? A pretty rare human pathogen, originally isolated from the Great Salt Lake in Utah.  It’s related to the better known bug F. tularensis which is one of the most infectious pathogens known.)     

By transferring the F. novicida Cpf1-CRISPR gene locus into E. coli they quickly established that it prefers a  “TTN” PAM motif that is located 5’ to its protospacer target – not 3’, as per Cas9.  So right away it’s distinct in having a PAM that isn’t G-rich and is on the opposite side of the protospacer. 

Like Cas9, Cpf1 binds a crRNA that carries the protospacer sequence for base-pairing the target.  But for me the biggest surprise in the paper is that unlike Cas9, Cpf1 does not require a separate tracrRNA – in fact, there’s no sign of a tracrRNA gene at the Cpf1-CRISPR locus.   Thus, Cpf1 merely needs a cRNA that is about 43 bases long –of which 24 nt is protospacer and 19 nt is the constitutive direct repeat sequence.   This is very different than Cas9 – even by fusing the crRNA and tracrRNA, the single RNA that Cas9 needs is still ~100 nt long.

Furthermore, the Cpf1 crRNA does not have the long stemloop structure that is typical of RNAseIII-mediated processing to cut it out of its primary transcript.   It has a much shorter stemloop that is required for Cpf1 activity, however.   But surprisingly, Cpf1 itself is apparently directly responsible for cleaving the 43-base cRNAs apart from the primary transcript in the first place! This isn’t conclusively proven yet, but is pretty likely based on their experiments.   

Next, two more surprises comes from the cleavage sites on the target DNA.  The cut sites are staggered by about 5 bases.  This should create “sticky overhangs” that might be exploitable to enable gene editing via NHEJ-mediated-ligation of DNA fragments with matching ends.   And, the cut sites are in the 3’ end of the protospacer, distal to the 5’ end where the PAM is.    The cut positions usually follow the 18th base on the protospacer strand and the 23rd base on the complementary strand (the one that pairs to the crRNA).

They tested if they could inactivate the DNA cleavage domains via homologous mutations in codons known to do this in Cas9.  However, the resulting Cpf1 mutants don’t have “nickase” activity – they can’t cut either strand.   So it’s not clear that Cpf1 nickases can be made and in fact the authors suggest that the cleavage might require some sort of dimerization.  I can’t visualize how that would work yet but I’m sure it will be figured out in the near future…

Base substitution experiments then showed that, as per Cas9 CRISPRs, there is a “seed” region close to the PAM in which single base substitutions completely prevent cleavage activity.   Therefore, unlike the Cas9 CRISPR target the cleavage sites and the seed region do not overlap.  This immediately suggests a potential improvement over Cas9 in mammalian HDR-mediated repair efficiency.    This is because any initial cleavage events that might lead to “simple” NHEJ indels might still be substrates for cleavage  - and thus allowing additional chances for HDR-mediated edting to occur.  With Cas9, an indel mutation will almost always disrupt the target seeds and then it’s game over for HDR.     

Finally, they did the important work to screen various Cpf1 proteins from different bacterial species to see if any would actually work in mammalian cells.  This is because that despite codon optimization and attachment of nuclear localization signals, most of these bacterial proteins just don’t work right when you put them inside human or mouse cells.  Therefore they tested 16 different Cpf1 proteins.  Of these, for seven proteins they could identify PAM signatures using their E. coli assay.  They all had similar T-rich PAMs. 

Of these seven proteins, only two worked well in human HEK293 cells – AbCpf1, from an Acidaminococcus, and LbCpf1, which is from a Lachnospiraceae; interestingly these are apparently both anaerobic bacteria sometimes found in mammalian intestines.  Anyway these can both generate indels at specific targets in human cells at a rate similar to Cas9 – typically, 10-20% Surveyor assay numbers were observed, when they tested HEK cells following simple transfections.

Bottom line: Cpf1 may be the real deal as a serious competitor for Cas9.  Is Cas9 suddenly obsolete?  Hardly.  First, we don't yet know if Cpf1 is as specific as Cas9 – though there is every reason to think it may be.   So the off-target effects need to be carefully measured. Second, we don’t know how widespread the targeting efficiencies will be across sites (although the initial tests seem very promising).  Third, although Cpf1 may be better than Cas9 for mediating insertions of DNA, it’s not yet been shown if that is true.   However it may have some nice advantages over Cas9, not the least of which is that its guide RNA is only 43 bases long.  It will thus be feasible to purchase directly synthesized guide RNAs for Cpf1, perhaps with chemical modifications to enhance stability.

Probably, Cpf1 and Cas9 will both be in the spotlight for a long while to come.  Look for more Cpf1 papers to start coming out very soon and for Cpf1 plasmids to appear in Addgene.  Happy CRISPRing, everyone.




Monday, August 17, 2015

#CRISPR target sequence preferences are being clarified. Xu et al Genome Research paper.

It's of huge value to be able to predict CRISPR target efficiency ahead of time.  Xu et al have published an analysis of multiple guide RNA data sets and extracted what they claim is an improved model for target cleavage efficiency prediction.   This data is all for the S.pyogenes native Cas9. 

Xu et al. Sequence determinants of improved CRISPR sgRNA design.
Genome Res. 2015 Aug;25(8):1147-57. doi: 10.1101/gr.191452.115. Epub 2015 Jun 10.

Their paper is important to me for several reasons.  First, they have examined two independently-published "large" guide RNA data sets that had mutagenesis-efficiency data, which allows more confidence that trends of sequence preferences are holding up across labs and platforms.   Second, they validated their predictive model on a small (in comparison to genome-wide, but still not bad) data set of new CRISPR targets and corresponding guide RNAs.  Third, they did "in silico validation" by turning their model loose on another target/indel data set, and showed improved performance of their predictive model over a previously published model.   See ROC curves in Fig. 4b.    This allows an ability to weed out "50-60% of the inefficient sgRNAs…at the cost of 10-20% of efficient sgRNAs misclassified."  That is, misclassified as inefficient.    

For those who are interested in genome-wide knockout screening experiments these sorts of models are very good for increasing efficiency of the screens.   Moreover, if you wish to knockout particular genes, it will allow you to test or use fewer targets per gene till you find one that works well.

OK, now the sobering reality for nerds like me is that predictive models, even with great ROC curves, have false positive and false negative rates that will bite you in the behind on a regular basis if you are designing large projects around the function of single CRISPR targets.  I'm still facing this issue for precision knock-in projects, for which there are often not many targets to choose from.   And with transgenic mice we always want the efficiency as high as possible.   For cell lines, hey, that's not as much a problem if you can subclone the edited lines.

But let's get back to the CRISPR target sequence preferences.  The bottom line here is that the last three bases of the protospacer seem to have the most influence on cleavage efficiency, with a C preferred at the -3 position (relative to the PAM), and G's at -2 and -1.    Also, G's are helpful at the -17 to -14 region, while A's are good at the -12 to -9 region.  Finally, a C seems helpful at +1 following the PAM.

Looking back at the Wang et al paper, they also reported a preference for A's at around -10 to -8, and essentially a "GCRR" preference for bases -4 to -1.  This makes sense since Xu have based their model partly on the the Wang data.   However, Xu et al point out that the apparent G preference at the -20 position is probably an artifact of the Wang sgRNA library in that these may have had increased efficiency due to enhanced transcription, not activity per se.

General GC-richness in the protospacer is known to correlate with CRISPR mutagenesis.  Could that just be driven by the GC-rich preferences of the last few bases?   Otherwise, GC-richness doesn't clearly emerge from the Xu model, at least to me anyway.  I took a crack at this by looking at a data set from Gagnon et al, mostly because I could handle the size of their sgRNA list in an excel spreadsheet without exploding my own brain or my iMac.  My impression is that GC richness is still "good" even when the last 4 bases of the protospacer are similar.   Here's an example.  From Gagnon et al's list of 122 sgRNAs with indel numbers, I ranked them according to how well they matched the "GCRR" of the last four bases.  I based this on the Wang et al paper although I think it is very similar to that corresponding part of the Xu model.   My  "score" ranged from 0 to 7.  Then I examined the subset of 30 targets that all had a same "score" of 5.   So these targets are all controlled, at least kinda sorta, for their  variation in bases -4 to -1 in that they have similar strength of matching to the "GCRR" motif.  Finally, I graphed the indel frequencies versus the GC content of the first 16 bases of their protospacers.   Here is the data.  y axis= indel frequency (in a zebrafish model), x axis= # of GC base pairs in first 16 bases.

This ain't close to something I'd submit for peer review but I do see a trend.  GC richness in the first 16 bases of the protospacer correlates with cleavage efficiency, even within a group of targets for which the 3' ends are similar.   So for now - I will continue to prefer overall GC-rich targets that also have at least some matching to the "CGG", or "GCRR", motif at the very 3' end.  

Also, "CGG" matches the high-efficiency 3' end reported by Farboud and Meyer so there's another corroboration.

So, the answer to my previous post "Are there sequence preferences near the 3' end of the #CRISPR protospacer? …" is, yes.  And this holds up for S.pyogenes Cas9 when used across human, mouse, fish and C.elegans models.   

A final note - these data all refer to cleavage and/or knockout efficiencies.  CRISPRi and CRISPRa screens, which do not lead to or require DNA cleavage, have different sequence preferences which Xu et al also modeled in detail.   

Happy CRISPRing.

Wednesday, May 20, 2015

About using DNA or RNA for mouse embryo #CRISPR injections.

I got a question:

Isn't the disadvantage of injecting DNA the threat of integration and more frequent mosaicism than in the case of RNA as Cas is expressed quicker? Do you have some direct experience with that? Thanks! 

Um, well yes.  Yes.  Those are the disadvantages.  Also I will add that because the RNA should lead to quicker Cas9 expression,  mutagenesis efficiencies will likely be higher than with DNA vectors.

So why use DNA at all?  Well, the issues are mostly practical.  DNA vectors are easy to customize for CRISPR.  Although the issues of efficiency and mosaicism are potentially problematic, I have seen pretty consistent success* in generating simple indel mutations following injections of PX330-style CRISPR-Cas9 DNA plasmids.  That is, consistent double digit percentages of founders carrying mutations as assayed by PCR and/or sequencing.    In addition, in our core we have obtained HDR-mediated codon editing rates in the 10-20% range using PX330 vectors co-injected with appropriate "donor" oligos.  But this is dependent on cooperative CRISPR sites that have a high rate of baseline cleavage.

Another practical consideration is that not everyone can routinely synthesize high-quality RNAs in vitro with consistency.   Quality DNA is relatively easy to prepare and QC.   RNA is much less so - especially for the 4+ kilobase Cas9 mRNA.   OK, so some of you are saying "Come one, my lab makes RNAs all the time - no prob! " .    That's awesome, but the empirical observation is that it's not trivial to get proficient at making long mRNAs, and to keep on top of the key reagent issues (RNAses, enzymes going bad, etc.).

Also, CRISPR DNA plasmids are immediately useful for cell culture gene editing experiments.  Some labs will be making these anyway so they will have them on hand, ready to go.   

What I am also observing - which many others have reported - is that a fraction of CRISPR sites just don't cut very well, even when the sequence characteristics of the site seem OK.  (Like, somewhere on the order of 1/3 to 1/4 of CRISPR target sites?) Most of our injections to date have been using DNA plasmids.   It's possible that RNAs might save the day for some of these sites.    


The ability to do precise HDR-mediated editing/insertions, rather than simple indels, is very compelling and is the direction most of our CRISPR ideas are going in terms of new mouse models.  But coding modifications usually have extremely narrow CRISPR target choices that are imposed by the science;  if you want to change a codon, you'll probably need a target as close as possible - preferably overlapping the codon.  There won't be many to choose from.  So getting the highest efficiency cleavage rates may be critical for some of these projects - for these, Cas9 mRNA or protein may be needed.

Finally, these issues of target efficiency really call for pre-validation of sites.  This can be done by transfecting CRISPR plasmids into cooperative cell lines, e.g. NIH3T3 for mouse targets, followed by PCR and mismatch cleavage assays, which can then be quantified.  But then - if you go through the trouble to do that, you will have generated the DNA plasmids and thus have the DNA reagent ready for injection.   

Having said all that, although I really like the convenience of plasmids, the RNA problems are all about sourcing them.  A few vendors, such as Sigma-Aldrich can provide custom guide RNAs and Cas9 mRNA that work.  (FYI I do not receive any compensation from Sigma).  The RNA reagent expense is less than the cost of mouse embryo injections.  I suppose zebrafish researchers may balk at the cost, as they will have more capacity to inject fish eggs, in their own labs usually, and may be more willing to make RNAs in-house.  For mice, you'll be usually working with a transgenic core and spending thousands of bucks per experiment.  Vendor-supplied RNAs may be worth the money.    Thanks for the question!

*Actually, "consistent" may be misleading… To clarify, about 75% of the NHEJ projects I've been observing have had this level of success.   So - more success than not, but then again, not perfectly consistent.  

Wednesday, April 29, 2015

The reported off-target effects in the recent Liang et al human embryo #CRISPR paper are partly incorrect.


As widely reported last week, a group in China has published results of CRISPR editing experiments in human triponuclear embryos (Liang et al, Protein & Cell 2015).   The news blurb in Nature is worth a read to get the context of the paper, which follows on the heels of a previous statement published in Science by leaders in the CRISPR field and others, in which they discourage CRISPR experiments in human embryos at this time pending further discussion of the implications of such research.    

In this post I won’t get into the ethical implications of the paper (which is more than I can deal with in one post anyway!).   Here I’ll discuss the technical results.   The Liang et al paper does not actually present much data that is very surprising - it is not unexpected that CRISPR can induce targeted mutations in humans, since it works in basically every species in which it’s been tried.    Their target gene was HBB (beta-globin) and they attempted HDR using the familiar approach of coinjecting Cas9 mRNA, guide RNA, and donor oligo ssDNA.

Here’s their 4 main points, paraphrased from the abstract:   
  1. Efficiency of HDR was low.
  2. Edited embryos were mosaic.   
  3. Off-target mutations were evident.
  4. A separate, highly homologous gene (HBD) could serve as donor template for repair, thus introducing sequences inadvertently from the other gene into the target gene.

Of these, points #1 and 2 were not surprising to those who have injected CRISPR reagents into mouse embryos, and not all that different.   The reported HDR efficiency was 14%, which is in line with at least some published mouse experiments (e.g. Singh et al 2014).   I will state that in our mouse core we are apparently seeing HDR results around 10-20% efficiency across several experiments.    Mosaicism has also been previously reported in CRISPR mice (Yen et al 2014).  Point #4 was kind of novel, but in retrospect not completely weird, since the HBD gene (delta-globin) is over 90% identical to HBB.

Ok, so regarding point #3 - off-target mutations...I was very interested in this because the authors reported four distinct off-target (OT) mutations in human embryos associated with the single CRISPR guide RNA they used, and this has already been described in the media as being substantially higher than OT rates in animal embryos.  Meaning, mouse embryos.      

One of these OTs was particularly surprising, as it looked like a very poor match to the target protospacer indeed - although the 3’-most 11 bases (the “seed” region”) matched the target, the 5’ bases only matched 1 out of 9 bases, for a total of 8 mismatches.    Frankly, this really scared me, because if true it means that the current methods used to predict OTs are not nearly broad enough.   But this level of OT mismatch was much greater than any OT I had seen before.

Bottom line: After looking at their data, I now firmly believe that only one of the four OT mutations were actually new mutations caused by CRISPR.  The other three were simply polymorphisms, already present in the  germline, that the authors mistakenly classified as OTs.

Here is how they did their OT analysis and my interpretation of the data.   Tripronuclear human embryos were obtained from a fertility clinic; you can identify these microscopically at the 1-cell zygote stage.  Fertilized by 2 sperm by accident, they are effectively triploid, and are absolutely unable to survive to term as normal pregnancies - but they can survive well enough during short-term CRISPR experiments, in which the embryos are only kept alive for a few days in vitro.   Briefly, 86 tripronuclear embryos were injected; 71 survived the injection; 56 of these were GFP-positive (used to as a reporter to show expression of injected reagents) and used for DNA analyses of on- and/or off-target effects.   28 of these embryos had on-target indel mutations and/or the desired HDR edit and were used for OT analysis.   

Note that they had originally chosen this particular CRISPR target from 3 potential targets they looked at in their gene;  one of these didn’t cut well and was not used further . For each of the other two, 7 sites were identified as the “top” potential OTs by using the MIT tool.  Of the two targets, one was found to have no OT mutations at the 7 potential OT sites when it was tested in 293T cells.  So they decided to work with this CRISPR target.  

2 of the 7 OT’s were found to have mutations in the injected embryos.  These were named G1-OT4 and G1-OT5 and are the first two of the four total OT mutations they claimed to identify (Figure 3A).  T7 mismatch assays were used for this analysis.   293T cell transfections with the guide RNA had already shown a lack of mutations across the 7 OT sites, that is, they were negative by T7 assays.  I’ll come back to these later.

They then did whole-exome sequencing on six of the embryos to identify potentially even more mutated OTs.  From this data, they first called indels and SNVs (single nucleotide variants) and then searched for protospacer similarity “allowing for ≤6 mismatches or perfect match of the last 10 nt 3′ of the gRNA” anywhere within 100 bp of the indels.  (Not sure if they did anything more with the SNVs.)   This identified two apparently new OTs, in the 3’ UTRs of  the C1QC and TTR genes, each found in one embryo (Figure 3B).  These were confirmed by T7 assays.

So - what does the T7 mismatch assay really indicate?  It reveals heterozygosity within the PCR product.   Of course, new mutations can cause this.  But so can plain old polymorphisms.   This is a drawback of using mismatch assays when applied to polymorphic samples.     

The next question is, simply, are there common human polymorphisms in the PCR products used in the OT analysis?  It’s easy to check this using the UCSC genome browser and the 1000 genomes site.   

For OT #1, a.k.a. “G1-OT4”, (Fig. 1C and 3A; PCR, hg19, chr11:132761837-132762356; intron of OPCML) there are no known common polymorphisms within the PCR that are close to the OT.    The closest SNP, rs79549129, has a minor allele frequency (MAF) of 1.2% but zero in asian populations.   There are no other annotated variants near the OT with a significant MAF.  The closest “common” SNP is rs2659601 but it’s about 50 bp from one end of the PCR product.  I don’t think that could produce the band sizes seen in the Fig. S3 T7 assays. From what I can tell from their Fig. S3 & S4, their T7 assays are compatible with new mutations that have been induced by cleavage close to the CRISPR OT site.   Thus, these look like “real” CRISPR OT effects at this site.    6/20 of on-target embryos, or 30%, had mutations at this OT.  So this looks like a real OT effect that replicates across embryos, but not in 293T cells.   

But then, polymorphisms become apparent in the other OTs...

For OT #2, a. k. a. “G1-OT5”, (Fig. 1C and 3A; PCR, hg19 chr22:31000551-31001000; intron of TULP4), there are two common SNPs on either side of the OT:  rs616358 (G/C) and rs628203 (T/C).   Haplotype derivations in Southern Han Chinese suggest haplotype population frequencies of ~56% GT, 35% GC, 9% CT, and zero % CC.  So we would expect to see plenty of heterozygosity in this PCR - easily observable by T7 assays, in the range of 50% or so being positive in a population based sample from this geographic location  - no CRISPR required.  This OT was a false positive.  

UCSC screen grab showing SNPs close to the OT (black bar in middle)

This leaves the two additional OTs they discovered by whole-exome sequencing.  Remember their workflow: they called indels in their data sets, then looked for nearby partial matches to the CRISPR target.   However, they apparently did not filter out known polymorphisms first.

OT #3 was in the 3’ UTR of the TTR gene.   Inspection of the OT sequence location (given in Fig. S6) on the UCSC browser clearly shows that rs143948820 is a known 9-base indel contained completely inside the OT.       Turns out that it’s uncommon outside of Asia but it has a MAF of ~2% in Southern Han Chinese.  With a heterozygosity of ~4% in normal diploids, it’s totally possible that 1 out of 6 triploid embryos would carry this variant.  This OT was very likely a false positive. 
UCSC screen grab; OT is black bar, indel variant is long red bar.


Finally, OT #4 was in the 3’ UTR of the C1QC gene.   And similar to the case above, this OT overlaps with a known 17-base indel, rs142916975, that has a MAF of 38% in Southern Han Chinese.   Heterozygosity should be close to 50%.   In fact, I’m surprised they got a false positive in only 1 of their 6 samples.   This is almost certainly a false positive.
UCSC screen grab; OT in black, indel variant in blue.



In summary, only 1 of the OTs holds up to scrutiny.    Importantly, neither OT found by exome sequencing holds up.  This flips their conclusion on its head: “Our whole-exome sequencing result only covered a fraction of the genome and likely underestimated the off- target effects in human 3PN zygotes.”.   While it’s certainly possible that some more OTs could be found by whole genome sequencing, the exome data was essentially totally negative.   Note that they chose a CRISPR to work with because it had a low apparent OT rate in 293T cells.   In retrospect, it was just by luck that G1-OT5 has a negative T7 assay in 293T cells.  It could have been heterozygous, but it's apparently not.

To their credit, the authors rightly restate that it’s going to be critical moving forward to carefully analyze off-target effects in any human applications of CRISPR.  This is widely agreed upon (see below).  However this paper made some technical mistakes in this regard.  While underestimating off-target effects could certainly have serious negative consequences for future CRISPR-based clinical treatments for genetic disease - and nobody wants that - overestimating them could generate an excess of hesitation to research the feasibility of such treatments within the broader scientific community.  

As stated by Baltimore et al in their Science commentary:

“It is critical to implement appropriate and standardized benchmarking methods to determine the frequency of off-target effects and to assess the physiology of cells and tissues that have undergone genome editing.”

I don’t think I’m blowing smoke here that we all need to get this right, as the media quickly reported on the Liang et al conclusions:  

From Wired:  “But—and this is a big but—using the technique without proper guidance could result in unforeseen consequences. The Chinese researchers, for example, found mutations in many of the embryos in genes other than the ones they’d targeted with CRISPR/Cas9.”

From Time, quoting Carl Zimmer from National Geographic:   “The experiment “came out poorly,” Zimmer says; in some cases, DNA was placed in the wrong spot and “off-target” mutations were discovered in the DNA.”

From the Washington Post:   “And in some of the embryos, the gene editing caused unintended mutations in other genes.”

From USA today (emphasis is mine): “The team also found that the complex used in the procedure was also acting on other parts of the genome, leading to other bits of it mutating. That happened much more than in previous experiments on adult human cells and animal embryos — and could happen yet more if the whole genome were used, as it would be if the embryo were to be implanted.”

            (OK, note from this last article the specific comparison to the very observations that I have blogged about in more detail than most people probably ever wanted to hear about...My point is that, due to the technical problems in the Liang paper, I don’t think we can yet say the off-target effects were “much more than in previous experiments on adult human cells and animal embryos”. )

One final note - my analysis of this paper should not be interpreted to mean that I fully endorse CRISPR experimentation or applications in human embryos.   I also applaud the authors' cautionary tone that the incomplete efficiency of CRISPR editing in humans is a problem that any therapeutic applications need to address.

Whew, this was the longest post yet.   

Tuesday, April 14, 2015

Are there sequence preferences near the 3' end of the #CRISPR protospacer? Paper from the C. elegans field explores this.

When the first word in a paper title is "Dramatic", I certainly wonder if I will agree after reading it…It's worth a blog post at any rate.   This paper by Farboud, B. and Meyer, B.J. is titled "Dramatic Enhancement of Genome Editing byCRISPR/Cas9 Through Improved Guide RNA Design" (Genetics, Vol. 199, 959–971 April 2015).   

As is true for other model organisms, CRISPR is very useful in nematodes for performing mutagenesis.  In this paper the authors were inspired by the previous observation that the Cas9 protein physically associates with the PAM motif (NGG) sort of promiscuously across DNA.  This had also previously led to the discovery that - in vitro - a CRISPR target region that is generally rich in GG dinucleotides will enable higher rates of cleavage at a unique CRISPR target within that region, than if the region is otherwise reduced in GG content.  In other words, general GG density probably "attracts" Cas9 and keeps more of it around, which in turn may enable faster recognition of the actual target.

Using this idea, the authors tested whether simply keeping an extra GG motif nearby the actual PAM NGG motif would enhance CRISPR mutagenesis in worms.  Turns out that if the PAM is followed (3') by another NGG, it doesn't help.  However, if the first 3 bases 5' to the PAM are NGG - that is, the last 2 bases of the protospacer are GG - they saw a consistent, and yes, dramatic improvement in recovery rates of CRISPR mutants.  This was a pretty striking finding and was validated across about 8 to 10 sites.  For comparison they tested "shifted" targets where they just shifted the protospacer 5' by 3 bases, and used those last 3 bases of the first protospacer as the PAM as the control.   These were almost uniformally poor in terms of absolute numbers of mutagenesis, with numbers usually at zero - meaning with their particular system of worm injections the baseline rate is pretty low.   The targets that had the GG at the protospacer 3' end usually had high mutation rates in double digit percentages.

So is this observed in other animals/cells/systems?   Well, based on my own work and from what I see in the literature, in mice we definitely don't need to have a GG at the protospacer 3' end to get high efficiency mutagenesis.   There are still no consistent rules here, but there are trends for sure.  Cas9 does seem to prefer purines in the last two bases of the protospacer as this seems to enhance gRNA loading (Wang et al 2014).     GC richness across the protospacer is definitely good (Gagnon et al 2014) which must correlate with G's in the protospacer 3' end.  Interestingly, Doench et al 2014 observed a preference for purine in the last base of the protospacer, but not much preference for the penultimate base (see their Fig 3a).

On the other hand I can't identify many examples yet of mammalian CRISPR targets that were published, had a GG in the last bases of the protospacer, had hard mutagenesis rates published and enough other targets in the same paper for a good comparison.  

Wednesday, January 21, 2015

TIDE: an online tool for evaluating #CRISPR gene editing in sequence trace files.

TIDE is a neat new web tool that's designed for a specific problem that I've definitely been dealing with. Following a CRISPR experiment, either in cell lines or animals, it's not trivial to quantify how well editing/mutagenesis worked and what sort of mutations were generated.   This is well summarized in the introduction of this paper so I won't repeat that, but I have certainly had these situations:  first, staring at ABI chromatograms following sequencing of PCR products from founder mice, and second, trying to quantify cleavage in pools of transfected cells.     Of course, the target site PCRs are going to usually contain mixtures of molecules with different mutations, and likely some amount of wild-type allele (for sure in pooled cells, often in founder animals).   So direct sequencing is hard to interpret as the actual chromatogram data past the cleavage site is usually a jumble of overlapping staggered sequences.

What TIDE does is actually to quantify the underlying non-wild-type sequence signal in the chromatogram data 3' to the expected cleavage site, then it quantifies the apparent contribution of specific, underlying mutant alleles, based on the pretty good assumption that most of the mutations generated by CRISPR will be short indels.  This seems to be a extension of PolyPeakParser, which I blogged about previously, but it's able to deal with multiple mutant alleles.  

I had a recent data set of sequence files from a mouse CRISPR experiment, so I thought I'd compare our independent analysis of the founder mice to TIDE's interpretation.  The gene is anonymized but I can state that it was a straightforward attempt to create indel mutations in a gene of interest.    Here's what we did:   About 25% of pups were positive for new mutations as revealed by Surveyor assays.  We then sequenced PCR products on 8 founder littermates, of which 6 were known Surveyor-positive and 2 were of unknown status.    

The last 2 (#19, #20) mice had normal, wild-type sequencing data.   The other 6 mice had very jumbled sequences past the cleavage site.   After some serious staring at the chromatograms - which took a while - I made some guesses that some of them had specific indel mutations.  However some of them were just too complex for me to figure out.      

Then I analyzed all of them with TIDE, using the sequence file from wild-type mouse #20 as the control file (which TIDE requires).   Here's the results:


Pup #
Pre-TIDE manual interpretation
TIDE result
2
WT allele and at least 2 different mutant alleles present. Could not interpret mutations at all.
No significant results, but the sequence quality was rather poor to begin with.   
6
WT allele and a 1-bp deletion allele. Germline transmission confirmed.
66.5% WT, 24.8 % 1-bp deletion.  
10
No WT allele; one 3-bp deletion; plus a complex (discontiguous)  4-bp deletion.  Germline transmission confirmed of both alleles at essentially mendelian rates.
10.9% WT, 44.9% 3-bp deletion, 33.7% 4-bp deletion. 
16
WT allele and 2 different mutant alleles present. Could not interpret mutations.
22.5% WT, 58.5 % 1-bp deletion, 9.7% 5-bp insertion.
22
WT allele and a 1-bp deletion allele.  Germline transmission confirmed.
60% WT, 30.2% 1-bp deletion. 
24
No WT allele, but multiple (>3) mutant alleles.
At least 4 different deletions of -2, -12, -28, -29 bp, each at low levels.
19
WT allele predominates.
75% WT;  7.4% 2-bp insertion; 8% 8-bp deletion.
20
WT allele predominates.
(Used #19 as control) 82.8% WT, 10.7% 4-bp deletion.

I was fairly impressed by the TIDE results.  First, it agreed with my specific interpretations for #6, 10 and 22, which were actually confirmed by germline transmission.   Second, it was able to correctly call 2 mutations at the same time in mouse #10.   Third, it made interpretations that made sense for founders #16 and 24, which I had given up on.    

Finally, I didn't really give the algorithm the optimal control sequence.  Instead I used the file for an apparently wild-type founder animal (#20).  However - when the files from #19 and #20 were used as controls to analyze each other, low levels of mutant alleles were detected.  And yes, if you go back to the chromatograms you can see a little underlying signal that may be a bit more than "usual" past the cleavage site - but it's very easy to miss.   This result is actually consistent with the experiment, since the embryos were injected with a PX330 plasmid, which may persist past the 1-cell stage and thus lead to low levels of mosaicism.      

Based on the imperfect controls I used, I would not take the TIDE quantitation of allele fractions literally.   However the qualitative results were pretty good and I wasn't able to find anything manually that TIDE didn't.  Also, this is a very fast analysis if you are performing sequencing on the PCRs anyway.   Moreover, it's easy to apply this analysis to PCRs on transfected pools of cells to measure CRISPR mutagenesis.  I'm looking forward to trying TIDE in this context as well.

Easy quantitative assessment of genome editing by sequence trace decomposition.  Eva K. Brinkman, Tao Chen, Mario Amendola and Bas van Steensel.    Nucleic Acids Research, 2014, Vol. 42, No. 22 e168