Wednesday, January 21, 2015

TIDE: an online tool for evaluating #CRISPR gene editing in sequence trace files.

TIDE is a neat new web tool that's designed for a specific problem that I've definitely been dealing with. Following a CRISPR experiment, either in cell lines or animals, it's not trivial to quantify how well editing/mutagenesis worked and what sort of mutations were generated.   This is well summarized in the introduction of this paper so I won't repeat that, but I have certainly had these situations:  first, staring at ABI chromatograms following sequencing of PCR products from founder mice, and second, trying to quantify cleavage in pools of transfected cells.     Of course, the target site PCRs are going to usually contain mixtures of molecules with different mutations, and likely some amount of wild-type allele (for sure in pooled cells, often in founder animals).   So direct sequencing is hard to interpret as the actual chromatogram data past the cleavage site is usually a jumble of overlapping staggered sequences.

What TIDE does is actually to quantify the underlying non-wild-type sequence signal in the chromatogram data 3' to the expected cleavage site, then it quantifies the apparent contribution of specific, underlying mutant alleles, based on the pretty good assumption that most of the mutations generated by CRISPR will be short indels.  This seems to be a extension of PolyPeakParser, which I blogged about previously, but it's able to deal with multiple mutant alleles.  

I had a recent data set of sequence files from a mouse CRISPR experiment, so I thought I'd compare our independent analysis of the founder mice to TIDE's interpretation.  The gene is anonymized but I can state that it was a straightforward attempt to create indel mutations in a gene of interest.    Here's what we did:   About 25% of pups were positive for new mutations as revealed by Surveyor assays.  We then sequenced PCR products on 8 founder littermates, of which 6 were known Surveyor-positive and 2 were of unknown status.    

The last 2 (#19, #20) mice had normal, wild-type sequencing data.   The other 6 mice had very jumbled sequences past the cleavage site.   After some serious staring at the chromatograms - which took a while - I made some guesses that some of them had specific indel mutations.  However some of them were just too complex for me to figure out.      

Then I analyzed all of them with TIDE, using the sequence file from wild-type mouse #20 as the control file (which TIDE requires).   Here's the results:


Pup #
Pre-TIDE manual interpretation
TIDE result
2
WT allele and at least 2 different mutant alleles present. Could not interpret mutations at all.
No significant results, but the sequence quality was rather poor to begin with.   
6
WT allele and a 1-bp deletion allele. Germline transmission confirmed.
66.5% WT, 24.8 % 1-bp deletion.  
10
No WT allele; one 3-bp deletion; plus a complex (discontiguous)  4-bp deletion.  Germline transmission confirmed of both alleles at essentially mendelian rates.
10.9% WT, 44.9% 3-bp deletion, 33.7% 4-bp deletion. 
16
WT allele and 2 different mutant alleles present. Could not interpret mutations.
22.5% WT, 58.5 % 1-bp deletion, 9.7% 5-bp insertion.
22
WT allele and a 1-bp deletion allele.  Germline transmission confirmed.
60% WT, 30.2% 1-bp deletion. 
24
No WT allele, but multiple (>3) mutant alleles.
At least 4 different deletions of -2, -12, -28, -29 bp, each at low levels.
19
WT allele predominates.
75% WT;  7.4% 2-bp insertion; 8% 8-bp deletion.
20
WT allele predominates.
(Used #19 as control) 82.8% WT, 10.7% 4-bp deletion.

I was fairly impressed by the TIDE results.  First, it agreed with my specific interpretations for #6, 10 and 22, which were actually confirmed by germline transmission.   Second, it was able to correctly call 2 mutations at the same time in mouse #10.   Third, it made interpretations that made sense for founders #16 and 24, which I had given up on.    

Finally, I didn't really give the algorithm the optimal control sequence.  Instead I used the file for an apparently wild-type founder animal (#20).  However - when the files from #19 and #20 were used as controls to analyze each other, low levels of mutant alleles were detected.  And yes, if you go back to the chromatograms you can see a little underlying signal that may be a bit more than "usual" past the cleavage site - but it's very easy to miss.   This result is actually consistent with the experiment, since the embryos were injected with a PX330 plasmid, which may persist past the 1-cell stage and thus lead to low levels of mosaicism.      

Based on the imperfect controls I used, I would not take the TIDE quantitation of allele fractions literally.   However the qualitative results were pretty good and I wasn't able to find anything manually that TIDE didn't.  Also, this is a very fast analysis if you are performing sequencing on the PCRs anyway.   Moreover, it's easy to apply this analysis to PCRs on transfected pools of cells to measure CRISPR mutagenesis.  I'm looking forward to trying TIDE in this context as well.

Easy quantitative assessment of genome editing by sequence trace decomposition.  Eva K. Brinkman, Tao Chen, Mario Amendola and Bas van Steensel.    Nucleic Acids Research, 2014, Vol. 42, No. 22 e168

Friday, January 16, 2015

Watch Dr. Jennifer Doudna's Vanderbilt Flexner Discovery lecture on #CRISPR online: tinyurl.com/nwbeon7

Vanderbilt has set up a nice site for viewing of the Discovery lectures, so check out the link to hear and see Dr. Doudna's talk from last week.   The audio/video was pretty good - this is what we saw in the overflow room (which also overflowed!).  A great talk by Dr. Doudna.

http://mediasite.vanderbilt.edu/Mediasite/Play/b8701c579ec6438abcc0188275efa16a1d

Monday, January 5, 2015

Monday, December 15, 2014

COSMID, a tool to find #CRISPR off-targets including indels (usually not found w/other tools).

From Georgia Tech comes a new web tool, COSMID, for finding CRISPR off-targets (OTs).  But, this one definitely adds something new to the already busy realm of OT-finding tools: the ability to find OTs with small insertions or deletions relative to the target, not just base-pair mismatches.   Like "mismatched" OTs, off-target cleavage can occur at "indel-OTs" as well; see Lin et al, Nucl. Acids Res.42(11): 7473-7485.    

COSMID is described in a new publication by T.J. Cradick, et al.  COSMID: A Web-based Tool for Identifying and Validating CRISPR/Cas Off-target Sites, Molecular Therapy—Nucleic Acids (2014) 3, e214.    I've only used it briefly so far but it has a nice interface that allows the user to search for OTs with a user-specified max number of mismatches in combination with single-base insertions and/or deletions.  It also can perform PCR primer design for the OTs with the primers optimized for Surveyor-style mutation screening.    Thanks go to T.J. and colleagues for the nice addition to the CRISPR off-target screening toolkit.

Alas, this already raises a new caveat about my previous post reviewing off-targets in mice: I'm pretty sure that most or all of those citations did not look at indel OTs, only mismatch OTs.    !

Monday, December 1, 2014

Here's my review of published #CRISPR off-target mutation data from mouse embryo injections.


Here is a literature review of CRISPR off-target (OT) mutation analysis in mouse oocytes.    This review only concerns published experiments using “native” Cas9 that cuts both DNA strands, and not nickase-Cas9 experiments.  Although nickase-Cas9 is much less prone to OT mutation, editing is still generally less efficient than with native Cas9 .  Therefore it’s important to know whether the potential problem of off-target mutation rates with native Cas9 outweigh its utility.   All the data below is pertinent to mouse zygotes.  Other systems such as cell lines may have different OT rates.

Here’s a few pertinent questions to preface this review:  First, how should potential off-targets (OTs) be defined ahead of time?  It’s complicated by the fact that mismatches are less tolerated within the “seed” region of 8-12 bases proximal to the PAM site, and more tolerated in the more distal (5’) region of the protospacer.   So some groups define OTs as having perfect matches to the seed region, while other groups defined them as simply having fewer than a threshold maximum number of mismatches anywhere in the protospacer.   Alternatively, they can be scored for cleavage potential by algorithms such as the MIT CRISPR design tool.

Second, how is CRISPR performed? Some of these groups used RNA or DNA injections;  most used slightly varying injection concentrations.   

Third, how were the OTs screened?  Most of these did direct sequencing on PCR from founders or Surveyor-type assays.   Also, OTs often cut at lower efficiency than the on-target but the results depend on the assay sensitivity and the number of pups screened, which varies across studies.   So the data here is only a general comparison.

I’m not focusing on those differences here, since the overall picture is broadly similar - OT rates were generally low to nil.   

Let’s start with the pair of 2013 Cell papers from the Jaenisch lab.   

1.  Wang et al. (Cell 2013) was the first report of CRISPR-mediated mutagenesis in mouse zygotes.   They only considered OTs with perfect matches to the 12 bases adjacent to the PAM and also the PAM itself (NGG).  For 2 targets, they defined 7 total OTs. (A third gRNA they used had no OTs by this definition). In 7 mutants pups carrying mutations at 2 simultaneously-targeted gRNA targets, they found zero mutations at the 7 OTs.   
Bottom line:  7 OTs screened, 0 mutated.

2.  Yang et al (Cell 2013) screened OTs that were defined at having “up to 3 or 4” mismatches.  (In my experience most mouse CRISPR targets have several-to-many 3-base mismatches, and I’m guessing that most targets will have many 4-base mismatches in mammalian genomes.)  I believe they screened OTs for 5 targets across 4 different genes.   A total of 35 pups and ES cell clones were screened in separate experiments using different gRNAs.  Of 47 OTs, only 3 had mutations. Two of those sites were mutated in multiple mice, indicating fairly high rates at these particular OTs.   However, the mutated OTs all had only 2 or 1 mismatches, and the “high rate” OTs only had 1 mismatch near the 5’ end, distal to the seed region.   
Bottom line:  47 OTs screened, 3 mutated.

3. Li et al.  (Nat. Biotech. 2013)Of 4 targets they used, only two had OTs with fewer than 4 mismatches or perfect seed matches, so they focused on those. 12 founders were screened. In a subset of pups also screened some more OTs that had perfect seed matches but were otherwise totally mismatched. 
Bottom line: more than 13 OTs screened, 0 mutated.

4.  Mashiko et al (Sci. Rep. 2013) were the first to publish on injecting plasmid DNAs into mouse zygotes for transient CRISPR expression.    Similar to Wang et al, they defined OTs as having a perfect match to the 12-13 bases adjacent to the PAM.  For two targets, they defined 7 and 4 OTs respectively; in 16 and 8 mutant pups made with either gRNA, they found one pup with a single OT mutation.
Bottom line:  11 OTs screened, 1 mutated.

5. Fujii et al (NAR 2013) targeted the Rosa26 locus, and reported that OT rates dropped as injected RNA concentrations were lowered.   They inspected 10 OTs with “3 or 4 mismatches” but each of these actually had a mismatch to the “N” of the PAM motif, which doesn’t affect targeting, so these were really “2” and “3” mismatched OTs.  No OT mutations were found for the true 3-mismatch OTs.  However they found mutations in all four 2-mismatch OTs,  particularly when injecting higher RNA concentrations .  They also examined 12 OTs for 2 targets in Hprt and found no OT mutations.
Bottom line:  22 OTs screened; 4 mutated but only in “2-mismatch” OTs.

6.   Wu et al (Cell Stem Cell 2013) targeted the Crygc gene and defined OTs as having no mismatches in a 14 base seed region of the target.  Of 10 OTs, 1 was mutated in 2 out of 12 pups.   
Bottom line:  10 OTs screened, 1 mutated.

7. Inui et al (Sci. Rep. 2014) examined about 10 OTs total for two targets, in Sox9 and Sf-1.
Bottom line:  ~10 OTs screened, 0 mutated.

8. Zhou et al (FEBS J. 2014) retargeted an EGFP cassette separately with two gRNAs.   OTs were defined using the MIT CRISPR design tool, and they analyzed a subset of these (15 OTs per target) by surveyor assay on the founder pups.  For one target, they detected mutations in 4 OTs, but not in any OTs for the second target.
Bottom line: 30 OTs screened, 4 mutated.

9. Mizuno et al (Mamm. Genome 2014) targeted the Tyr gene and screened 5 OTs in founders by sequencing.
Bottom line:  5 OTs screened, 0 mutated.

10.  Han et al (RNA Biol. 2014) used 4 targets and identified OTs with the MIT CRISPR design tool.  They screened 3 founders for the “top 5” potential OTs.
Bottom line:  20 OTs screened, 0 mutated.
  

Here’s a quasi-meta-analysis:

From these 10 studies, 5 (50%) were able to detect some degree of off-target mutation.
But from ~175 OT’s screened, mutations in only 13 (7%) were detected.  Several of these OTs had fewer than 3 mismatches to the target.

In conclusion, the consensus from many studies of CRISPR-mediated mouse engineering demonstrates that native Cas9 has a low rate of off-target effects in mouse zygotes.  Of course, targets should still be pre-screened when possible to avoid those that will have more potential off-targets, particularly those with fewer than 3 mismatches within the protospacer.   

Doug Mortlock 2014.


Bibliography



Monday, November 24, 2014

Optimal design of ssODNs (donor oligos) for #CRISPR - length and strandedness data?

(UPDATE Jan. 29 2016:  Also see my new post "For #CRISPR HDR, use donor oligos that are complementary to the "gRNA strand". A new paper shows why"- this supports the choice of strand to use, but also impacts the placement of the homology arms.)


For this CRISPR question I am going back to this paper:  Yang et al, Optimization of scarless human stem cellgenome editing, Nucleic Acids Research, 2013, 1–13  (from the Church lab).   Although this paper had a lot of TALEN data it had comparative data for CRISPR-mediated editing.  In this case, a 2-bp mismatch was engineered into the CCR5 locus in human iPS cells.    

First of all, from now on I will use the term "ssODN" to refer to single-stranded donor oligonucleotides.  (Hooray, more jargon!)   In Fig. 3d of Yang et al, they presented a nice series of data in which they varied both the length and the strandedness of the ssODN used for editing in conjunction with a single gRNA, which was held constant of course.   The mismatch was within the CRISPR target and was always positioned in the middle of the ssODN.  However, ssODN length varied from 50 to 110 nt.  (Strangely, the top panel has longer ssODNs also drawn schematically but no data for those was shown).  In addition, ssODNs corresponding to either the same strand or the complementary strand to the gRNA were tested.  

Now, if you're like me, you might naively assume that the complementary-stranded ssODN would be worse in mediating editing because it might base-pair with the gRNA itself, preventing the gRNA from functioning properly.    Sounds reasonable?  Turns out, the opposite was true - at least for this individual target.  Maximal efficiency of editing was achieved with the complementary ssODN at about 70 nt length, with an absolute efficiency of ~1.5% - not too shabby considering it's iPS cells and without any selection.  Interestingly, when non-complementary ssODN (i.e., same strand as the gRNA) were used, the efficiency never reached that efficiency, but it did increase with length up to the maximum tested length (110 nt) where it reached about 0.5%.  At this length, it was basically the same whether the complementary or non-complementary ssODN was used.

What should one take from this?  Well, the trouble with these sorts of tests is that when you test certain parameters you have to keep the other parameters fixed.  In this case the gRNA was kept constant.  So the peak in efficiency at 70 nt with the complementary ssODN might be a peculiarity of that particular sequence - perhaps it forms a secondary structure that just happens to inhibit the competing NHEJ pathway, for example?   

Also, there is a decent theme forming that for mouse oocyte injections, ssODNs should have homology arms of about 60 nt.  This puts the minimum length at 120 nt.


P.S. ...another thought added 12/15/14:    Note that for the above observation that max. editing efficiency was achieved with a complementary-strand ssODN of ~70 nt, this means the homology arms were each only ~35 nt long.  In fact, a 50 nt ssODN was as efficient as a 90 nt ssODN, meaning ~25 nt homology arms were also workable with a complementary ssODN.  But the non-complementary ssODN worked better with longer arms.  See Fig. 3 from the Yang paper.

Friday, November 21, 2014

New volume of Methods in Enzymology devoted to #CRISPR, TALENs, ZFNs.

The volume is entitled "The Use of CRISPR/Cas9, ZFNs, and TALENs in Generating Site-Specific Genome Alterations",  Methods in Enzymology, Volume 546, Pages 2-549 (2014).  Edited by Jennifer A. Doudna and Erik J. Sontheimer.    

A variety of great topics are addressed in this volume, including many that I've talked about in this blog, and it features chapters written by many of the authors of key papers that are also discussed in my blog posts.   Too many interesting chapters to list all the pubmed links here!