Wednesday, October 28, 2015

#CRISPR / gene editing featured at Festival of Genomics Nov. 3-5 in San Mateo. #genomicsfest

I will be participating in a CRISPR mouse editing workshop next week (Nov. 3) at The Festival of Genomics, in San Mateo, CA.  The full workshop title is "CRISPR/Cas9 Genome Editing Pipelines for Mice and Rats" and I'll be joined by Thom Saunders (University of Michigan) and Kevin Peterson (The Jackson Laboratory).  

The Festival of Genomics is a fairly new, recurring conference that has a lower registration fee than most traditional scientific-society meetings.  The speaker lineup for the meeting next week is pretty strong, with a strong biotech presence not surprising given the locale, but also strong plenaries from academics as well (Jennifer Doudna, of course; Carlos Bustamante, Manolis Kellis etc).  

Looks like Tues. Nov. 3 is mostly workshops, and Weds & Thurs Nov. 4-5 will feature plenary talks in the morning and then four concurrent tracks of sessions: Genome Editing, Genome Analysis, Genomic Medicine, and Data Analysis.    If you're in the SF bay area, check it out.

Friday, October 23, 2015

Updated guidelines for mouse #CRISPR injections in the Vanderbilttransgenic core.

I have revised the CRISPR guidelines on our Vanderbilt transgenic core's website.

 https://labnodes.vanderbilt.edu/resource/view/id/5265/collection_id/14/community_id/8

Basically it describes the basic information about what to do, and more importantly what to know, to get a CRISPR mouse project started through our core.   This may be most useful for my Vanderbilt peeps but others may find it interesting, as it gives some insight into how our core facility is communicating these sorts of guidelines to our users.  

These guidelines are not heavy on the up-front, nitty-gritty CRISPR design aspects as there are other places to find that stuff - for example, in the archives of this blog, as well as through the links I provide on the right side of this blog page to some helpful tools.  

Thursday, October 1, 2015

UK researchers ask you to submit your opinions about gene editing - link to web survey. #CRISPR

Dr. Lara Marks and Dr. Silvia Camporesi would like you to tell them your opinions about gene editing technology via their online survey.   What do you think? Let them know.  

Dr. Marks edits the hosting website, What is Biotechnology.   Dr. Camporesi is a bioethicist at Kings' College London.

Monday, September 28, 2015

Move over, Cas9: Cpf1 may be your new #CRISPR competition.

This past weekend at the Pilgrimage music festival in Franklin, TN I had the pleasure of seeing Weezer rock out, then I walked a few hundred yards to another stage to watch Wilco do the same.  These bands each had their own stage, but in the CRISPR world, Cas9 is a superstar that now has to share the stage with a newcomer:  Cpf1.  (Yeah, I know that's a goofy setup but it really was a good festival and it was on my mind.  Now on to the science.)

Last Friday Feng Zhang’s group published a paper in Cell that immediately grabbed a lot of attention, and rightly so.  They reveal that the Cpf1 class of CRISPR effector proteins may be an attractive alternative to Cas9.   Although Cpf1 has many similarities to Cas9, it has some significant differences that are very interesting – and could lead to improved efficiencies for some types of gene editing.  

Cpf1 Is a Single RNA-Guided Endonuclease of a Class 2 CRISPR-Cas System.  Zetsche et al., 2015, Cell 163, 1–13October 22, 2015  (Avail. online Sept. 25 2015).

Here is my summary…First, they reviewed some background on CRISPR systems in bacteria to cover the basis of the study.   There are two major classes of CRISPR systems based mainly on the proteins involved; the cleavage effectors of class 1 are complexes of multiple proteins, while the class 2 effectors are single proteins like Cas9.  Within class 2 there are two subtypes of systems: those with Cas9, and another that has Cpf1.   Cpf1 means “CRISPR from Prevotella and Francisella 1”.  Some bacterial species carry both Cas9-CRISPR and Cpf1-CRISPR loci in their genomes.   Like Cas9, Cpf1 has RuvC-like DNA cleavage domains but it lacks some of the other domain and neighbor-gene features of Cas9, so it’s clearly distinct in its evolution.    It’s a largish protein of ~1300 amino acids, similar in size to Cas9.

They picked the Cpf1-CRISPR gene system of Francisella novicida strain U112 to study first since there were clear homologies of Cpf1-CRISPR spacer sequences to various prophage in this species – further suggesting that Cpf1 is important in bacterial immunity and so it’s well adapted to slice and dice target DNAs.   (Sidebar: what’s F. novicida? A pretty rare human pathogen, originally isolated from the Great Salt Lake in Utah.  It’s related to the better known bug F. tularensis which is one of the most infectious pathogens known.)     

By transferring the F. novicida Cpf1-CRISPR gene locus into E. coli they quickly established that it prefers a  “TTN” PAM motif that is located 5’ to its protospacer target – not 3’, as per Cas9.  So right away it’s distinct in having a PAM that isn’t G-rich and is on the opposite side of the protospacer. 

Like Cas9, Cpf1 binds a crRNA that carries the protospacer sequence for base-pairing the target.  But for me the biggest surprise in the paper is that unlike Cas9, Cpf1 does not require a separate tracrRNA – in fact, there’s no sign of a tracrRNA gene at the Cpf1-CRISPR locus.   Thus, Cpf1 merely needs a cRNA that is about 43 bases long –of which 24 nt is protospacer and 19 nt is the constitutive direct repeat sequence.   This is very different than Cas9 – even by fusing the crRNA and tracrRNA, the single RNA that Cas9 needs is still ~100 nt long.

Furthermore, the Cpf1 crRNA does not have the long stemloop structure that is typical of RNAseIII-mediated processing to cut it out of its primary transcript.   It has a much shorter stemloop that is required for Cpf1 activity, however.   But surprisingly, Cpf1 itself is apparently directly responsible for cleaving the 43-base cRNAs apart from the primary transcript in the first place! This isn’t conclusively proven yet, but is pretty likely based on their experiments.   

Next, two more surprises comes from the cleavage sites on the target DNA.  The cut sites are staggered by about 5 bases.  This should create “sticky overhangs” that might be exploitable to enable gene editing via NHEJ-mediated-ligation of DNA fragments with matching ends.   And, the cut sites are in the 3’ end of the protospacer, distal to the 5’ end where the PAM is.    The cut positions usually follow the 18th base on the protospacer strand and the 23rd base on the complementary strand (the one that pairs to the crRNA).

They tested if they could inactivate the DNA cleavage domains via homologous mutations in codons known to do this in Cas9.  However, the resulting Cpf1 mutants don’t have “nickase” activity – they can’t cut either strand.   So it’s not clear that Cpf1 nickases can be made and in fact the authors suggest that the cleavage might require some sort of dimerization.  I can’t visualize how that would work yet but I’m sure it will be figured out in the near future…

Base substitution experiments then showed that, as per Cas9 CRISPRs, there is a “seed” region close to the PAM in which single base substitutions completely prevent cleavage activity.   Therefore, unlike the Cas9 CRISPR target the cleavage sites and the seed region do not overlap.  This immediately suggests a potential improvement over Cas9 in mammalian HDR-mediated repair efficiency.    This is because any initial cleavage events that might lead to “simple” NHEJ indels might still be substrates for cleavage  - and thus allowing additional chances for HDR-mediated edting to occur.  With Cas9, an indel mutation will almost always disrupt the target seeds and then it’s game over for HDR.     

Finally, they did the important work to screen various Cpf1 proteins from different bacterial species to see if any would actually work in mammalian cells.  This is because that despite codon optimization and attachment of nuclear localization signals, most of these bacterial proteins just don’t work right when you put them inside human or mouse cells.  Therefore they tested 16 different Cpf1 proteins.  Of these, for seven proteins they could identify PAM signatures using their E. coli assay.  They all had similar T-rich PAMs. 

Of these seven proteins, only two worked well in human HEK293 cells – AbCpf1, from an Acidaminococcus, and LbCpf1, which is from a Lachnospiraceae; interestingly these are apparently both anaerobic bacteria sometimes found in mammalian intestines.  Anyway these can both generate indels at specific targets in human cells at a rate similar to Cas9 – typically, 10-20% Surveyor assay numbers were observed, when they tested HEK cells following simple transfections.

Bottom line: Cpf1 may be the real deal as a serious competitor for Cas9.  Is Cas9 suddenly obsolete?  Hardly.  First, we don't yet know if Cpf1 is as specific as Cas9 – though there is every reason to think it may be.   So the off-target effects need to be carefully measured. Second, we don’t know how widespread the targeting efficiencies will be across sites (although the initial tests seem very promising).  Third, although Cpf1 may be better than Cas9 for mediating insertions of DNA, it’s not yet been shown if that is true.   However it may have some nice advantages over Cas9, not the least of which is that its guide RNA is only 43 bases long.  It will thus be feasible to purchase directly synthesized guide RNAs for Cpf1, perhaps with chemical modifications to enhance stability.

Probably, Cpf1 and Cas9 will both be in the spotlight for a long while to come.  Look for more Cpf1 papers to start coming out very soon and for Cpf1 plasmids to appear in Addgene.  Happy CRISPRing, everyone.




Tuesday, August 18, 2015

Webtool link for getting Xu et al #CRISPR target scores for your sequence of interest.

A followup to yesterday's post about the Xu et al paper,  Sequence determinants of improved CRISPR sgRNA design:   They have also kindly made a public webtool for generating CRISPR scores with their model.   It's a cut-and-paste that accepts up to 10000 bases.   Simple and quick.    

Of course, their source code is available too in their supplemental material and here.   

Monday, August 17, 2015

#CRISPR target sequence preferences are being clarified. Xu et al Genome Research paper.

It's of huge value to be able to predict CRISPR target efficiency ahead of time.  Xu et al have published an analysis of multiple guide RNA data sets and extracted what they claim is an improved model for target cleavage efficiency prediction.   This data is all for the S.pyogenes native Cas9. 

Xu et al. Sequence determinants of improved CRISPR sgRNA design.
Genome Res. 2015 Aug;25(8):1147-57. doi: 10.1101/gr.191452.115. Epub 2015 Jun 10.

Their paper is important to me for several reasons.  First, they have examined two independently-published "large" guide RNA data sets that had mutagenesis-efficiency data, which allows more confidence that trends of sequence preferences are holding up across labs and platforms.   Second, they validated their predictive model on a small (in comparison to genome-wide, but still not bad) data set of new CRISPR targets and corresponding guide RNAs.  Third, they did "in silico validation" by turning their model loose on another target/indel data set, and showed improved performance of their predictive model over a previously published model.   See ROC curves in Fig. 4b.    This allows an ability to weed out "50-60% of the inefficient sgRNAs…at the cost of 10-20% of efficient sgRNAs misclassified."  That is, misclassified as inefficient.    

For those who are interested in genome-wide knockout screening experiments these sorts of models are very good for increasing efficiency of the screens.   Moreover, if you wish to knockout particular genes, it will allow you to test or use fewer targets per gene till you find one that works well.

OK, now the sobering reality for nerds like me is that predictive models, even with great ROC curves, have false positive and false negative rates that will bite you in the behind on a regular basis if you are designing large projects around the function of single CRISPR targets.  I'm still facing this issue for precision knock-in projects, for which there are often not many targets to choose from.   And with transgenic mice we always want the efficiency as high as possible.   For cell lines, hey, that's not as much a problem if you can subclone the edited lines.

But let's get back to the CRISPR target sequence preferences.  The bottom line here is that the last three bases of the protospacer seem to have the most influence on cleavage efficiency, with a C preferred at the -3 position (relative to the PAM), and G's at -2 and -1.    Also, G's are helpful at the -17 to -14 region, while A's are good at the -12 to -9 region.  Finally, a C seems helpful at +1 following the PAM.

Looking back at the Wang et al paper, they also reported a preference for A's at around -10 to -8, and essentially a "GCRR" preference for bases -4 to -1.  This makes sense since Xu have based their model partly on the the Wang data.   However, Xu et al point out that the apparent G preference at the -20 position is probably an artifact of the Wang sgRNA library in that these may have had increased efficiency due to enhanced transcription, not activity per se.

General GC-richness in the protospacer is known to correlate with CRISPR mutagenesis.  Could that just be driven by the GC-rich preferences of the last few bases?   Otherwise, GC-richness doesn't clearly emerge from the Xu model, at least to me anyway.  I took a crack at this by looking at a data set from Gagnon et al, mostly because I could handle the size of their sgRNA list in an excel spreadsheet without exploding my own brain or my iMac.  My impression is that GC richness is still "good" even when the last 4 bases of the protospacer are similar.   Here's an example.  From Gagnon et al's list of 122 sgRNAs with indel numbers, I ranked them according to how well they matched the "GCRR" of the last four bases.  I based this on the Wang et al paper although I think it is very similar to that corresponding part of the Xu model.   My  "score" ranged from 0 to 7.  Then I examined the subset of 30 targets that all had a same "score" of 5.   So these targets are all controlled, at least kinda sorta, for their  variation in bases -4 to -1 in that they have similar strength of matching to the "GCRR" motif.  Finally, I graphed the indel frequencies versus the GC content of the first 16 bases of their protospacers.   Here is the data.  y axis= indel frequency (in a zebrafish model), x axis= # of GC base pairs in first 16 bases.

This ain't close to something I'd submit for peer review but I do see a trend.  GC richness in the first 16 bases of the protospacer correlates with cleavage efficiency, even within a group of targets for which the 3' ends are similar.   So for now - I will continue to prefer overall GC-rich targets that also have at least some matching to the "CGG", or "GCRR", motif at the very 3' end.  

Also, "CGG" matches the high-efficiency 3' end reported by Farboud and Meyer so there's another corroboration.

So, the answer to my previous post "Are there sequence preferences near the 3' end of the #CRISPR protospacer? …" is, yes.  And this holds up for S.pyogenes Cas9 when used across human, mouse, fish and C.elegans models.   

A final note - these data all refer to cleavage and/or knockout efficiencies.  CRISPRi and CRISPRa screens, which do not lead to or require DNA cleavage, have different sequence preferences which Xu et al also modeled in detail.   

Happy CRISPRing.