# Welcome to the CLIP Forum

The purpose of this wiki is to provide a real-time, community discussion of crosslinking and immunoprecipitation (CLIP)-based studies of protein-RNA interactions and RNA methylation. Both the experimental [(Garzia et al., 2017; Lee and Ule, 2018; Ule et al., 2018; Wheeler et al., 2018)](https://paperpile.com/c/3c2FI2/rPrX+jqZg+DBLj+A8KI) and computational [(Chakrabarti et al., 2018; Chen et al., 2019; De and Gorospe, 2017; Moore and AC’t Hoen, 2019)](https://paperpile.com/c/3c2FI2/2SQ6+Mj9Z+dLnN+YsbD) aspects have been reviewed, and also summarised more recently in the primer on [CLIP and complementary methods](https://www.nature.com/articles/s43586-021-00018-1).&#x20;

Our most recent version of improved iCLIP protocol is available from [bioRxiv](https://www.biorxiv.org/content/10.1101/2021.08.27.457890v1.full.pdf), and we established [iMaps](https://imaps.goodwright.com/), a free platform for password-protected data analysis, which contains many computational tools tailored to CLIP data.

CLIP methods continue evolving and additional questions can arise in specific applications, so this forum serves to enable further discussion and to address the newly emerging questions.

**If you would like to ask a question that is not answered already,** please ask your question on our [Slack workspace](https://join.slack.com/t/imapsgroup/shared_invite/zt-r24y3591-Xbhnym2t38u_urU~I0K0lQ) where we have a dedicated channel: #general-questions-for-clip-forum.

This forum summarises questions that arose in the past either through e-mail or Slack discussions - most questions are answered by Jernej Ule, but where answers are contributed by others this is usually indicated.


# Cell lysis

**Are there protocols for nuclear vs cytoplasmic CLIP for shuttling RBPs?**

Several publications, including from Ule lab, have done this. I add the relevant part from an old protocol below, this is meant to be a quick fractionation to minimise RNase activity, it may need to be optimised for your types of cells/tissues:

*Cytoplasmic Lysis buffer (CLB):*

* 50 mM Tris-HCl, pH 7.4
* 10mM NaCl
* 0.5% Igepal CA-630 (Sigma I8896)
* 0.25% Triton X-100
* 1mM EDTA

On the day of the experiment, add complete protease inhibitor cocktail (Roche, 25x) to the amount of buffer required for lysis.

*CBP-adjusting buffer:*

* 50 mM Tris-HCl, pH 7.4
* 100 mM NaCl
* 4% Igepal CA-630 (Sigma I8896)
* 0.4% SDS
* 2% sodium deoxycholate

*Nucleus Lysis Buffer (NLB)*

* 50 mM Tris-HCl, pH 7.4
* 100 mM NaCl
* 1% Igepal CA-630 (Sigma I8896)
* 0.1% SDS
* 0.5% sodium deoxycholate

On the day of the experiment, add cOmplete protease inhibitor cocktail (Roche, 25x) to the amount of buffer required for lysis.

2.3 A. Resuspend cell pellet (from step 1, each pellet is usually enough for 1 IP) in 1.5 ml of cold CLB. Make sure to fully resuspend the pellet by pipetting up and down with a 1ml pipette tip, and then rotate in cold room for 5 minutes. Afterwards spin at 4C, 1000 x g for 3 minutes. Collect the supernatant for cytoplasmic iCLIP and add 0.5 ml of CBP-adjusting buffer.

2.3 B. Resuspend the pellet in 1ml CLB, transfer to a 1.5 mL tube and spin again at 4C, 1000 x g for 3 minutes. Discard the supernatant, and add 1ml of NLB to the pellet (supplemented with protease inhibitors). Then use Bioruptor for 10 cycles with alternating 30 secs on/ off at low intensity. Six samples can be sonicated at the same time.&#x20;

2.3 C. Measure lysate concentration with Bradford assay and dilute all samples with NLB to 1 mg/ml.

## **RNase**

**We were not able to get a clean result after end-labeling of the RNA (Step-8) like you showed in Figure 2. In our trials, lane 1 and 2 (no antibody controls) were blank, just like what you showed in the figure. But our lane 3 and 4 both have a smear in the entire lane (even below the size of our protein), and the signal in lane 3 (the high RNase treated sample) was only slightly lower than lane 4 (low RNase). There was a band in lane 3 (high RNase) but not lane 4 (low RNase) at the size of our protein, but it was accompanied by significant amount of smear. Do you know why there was so much signal below the size of our protein, and how we can get rid of the smear in the high RNase treated sample?**

This is a phenomenon that also occurred to us in the past. Most often, it results from immunoprecipitation conditions that are not stringent enough for the IPed protein, which allow co-purification of other associated RNA-binding proteins that migrate below the size of the protein. For this purpose, it would help to try a few different IP conditions - sometimes it is sufficient just to increase the volume of lysis buffer added to the pellets and incubate the beads 5min rotating in cold room in the high-salt wash buffer during washes. An alternative is to increase the concentration of RNAse in the high-RNAse condition. It’s likely that the concentration isn’t high enough for your conditions.

## **RNase inhibitors**

**How do RNasin (Promega) and anti-RNase (Ambion) compare?**

They are similar RNAses, we use them in the lysis buffer when we wish to selectively inhibit RNase A, but not RNase I. Otherwise, SUPERaseIn (Life Technologies, AM2696) can be used if you want to also inhibit RNase I.

## **Sonication & getting rid of the DNA**&#x20;

**After RNase/Turbo DNase treatment of the lysate and centrifugation of the lysate, intact chromosomes were not completely precipitated and took up most of lysate volume as viscous goo. I could not pipet out the sup since pipet tips caught the goo very easily. I wonder Turbo DNase did not work well enough. A short sonication would help to chop off DNAs, but it would also shear RNAs. Would it make sense to increase Turbo DNase (amount and/or length), or to pretreat the lysate with Turbo DNase for some time before adding RNase? I could spin the pellet with higher than 20,000xG by using ultracentrifuge if it is an easier solution.**

Sonication is the only solution that works for us. Shearing of RNA is not a problem, since in any case we want to partially digest the RNA.\ <br>


# SDS-Page

## **Benefits of the PAGE**

**We would love to omit this step, and go directly from IP elution and PK treatment to library construction. Do we need to include the extraction of protein-RNA complexes from PAGE+transfer? Shouldn't the IP step remove RNA fragments that are non-specifically bound? And anyway, if it doesn't, wouldn't these non-specific RNA be also  labeled and present on the gel?**&#x20;

If IPing endogenous proteins, antibodies most often recognise a native protein, and therefore one is not able to include a denaturation step. Current iCLIP conditions are optimised for use of such antibodies. While the washing is as stringent as possible, it will not disrupt very stable RNP complexes. Therefore, there is a chance that you will co-IP other RBPs, and RNAs that stick to the beads or are non-specifically bound to your RBP.

The purpose of SDS-PAGE is therefore three-fold:

1. Increase the stringency of purification, because free non-crosslinked RNAs will migrate lower on the gel, and will not stick to the membrane as well after transfer. Also, if you have co-IPed another RBP of a different MW, this RBP will migrate at a different size on the gel, so it is possible to avoid it by cutting your band in a way that includes only the RBP-RNA complex of interest.
2. Control the quality of purification. Even if you are using a purification protocol that is generally clean, it’s reassuring to confirm this with SDS-PAGE before proceeding. The whole iCLIP sequencing and data analysis can take >month, and SDS-PAGE+transfer takes less than a day. If one can notice at this stage that something didn’t go right, it will save a lot of effort down the road.&#x20;
3. Monitor the quality of RNase fragmentation. Appropriate RNase conditions are crucial for iCLIP, and we find that they are sensitive to the batch of RNase, the concentration of extract, the type of cell you use, and the RBP that you study. Visualising the shift in the size of RBP-RNA complexes is the best way to ensure that the fragmentation is appropriate. This has major impact on the resulting data, as described in [(Haberman et al., 2017)](https://paperpile.com/c/3c2FI2/2HpP).

**What do you think about the option of labeling and running only a sample of the IP – just for analytical purposes, and do the rest of the protocol by digesting the RNP directly off the beads? This will spare the hassle of going through PAGE, yet will give indications about RNA fragmentation and IP quality.**

If you wish to skip the PAGE, then I’d recommend stringent denaturing purification conditions, which lead to clean complexes that shouldn’t require SDS-PAGE for further purification. For instance, you could use affinity purification-based methods that include a 6M urea wash, as described by the CLAP method (see <http://www.sciencedirect.com/science/article/pii/S1672022914000230> for a summary of basic options). Or you could use the two-step urea-iCLIP that uses 3xFlag and 6M urea: <https://www.ncbi.nlm.nih.gov/pubmed/24184352>. As you say, even if skipping the PAGE for most samples, it is useful to run some representative samples on PAGE for quality control of RNase and IP conditions. This protocol will work with most affinity tags, but for endogenous proteins you will need to be lucky, so that your antibody recognises a denatured protein.

## **Loading the gel**

**Our protein molecular weight is 72kDa, do you think I should add reducing agent in the 1xNupage Loading buffer during the elution step?**

Yes, in reducing conditions light and heavy chains will migrate at 25kDa and 50kDa, respectively, and will thus not interfere with migration of your protein-RNA complexes.

## **Transfer**

**Prior to transfer, would we be able to detect RNA on the get by ethidium bromide staining or are the RNA levels too low.**

Signal would be too low, and there would be more background. Transfer is necessary to remove free RNA that passes through the membrane.

**Do we need to use the NuPAGE gels from Invitrogen, and a specific brand of  nitrocellulose?**

We recommend that you initially follow the protocol exactly, because we haven’t tested it with other reagents. Once the protocol works in your hands, you can then try comparing results with different products. You need to use a gel with neutral pH to prevent alkaline hydrolysis, and a brand of 100% pure nitrocellulose.

**Transfer efficiency of crosslinked complexes is low (more than 50% stays in the gel even after overnight transfer). How can one improve this?**

Hi Svetlana, this is an interesting observation. In the past we have monitored efficiency with radioactivity, and it has been very high in our hands, but it may depend on the condition. Do you find it to be RBP-dependent, or general across RBPs? Is the RBP of high MW? How did you determine the efficiency, perhaps you can paste the result into this document?&#x20;

**Hi Jernej! To answer: 1 - Now I see it for one RBP \~100kDa, but I remember seeing this before with another (FUS). 2 - I measure pixel intensity on the 16bit IP scan with ImageJ to get an idea. Here is an example of 4-12% NuPAGE and standard transfer (20% MeOH, \~1h), \~⅔ left in the gel. Using 3-8% gel and overnight transfer with 10% MeOH helped a bit (but still a lot is left in the gel). I should add that this is 365nm crosslink with 4SU, but looks similar for 254nm, in my hands.**

![](https://lh6.googleusercontent.com/AwDcNAfd0EshVlyfHdoVGJUEF2chLqF93xpSqHmYHUZnM2QpYplgo4_SBIT8PjArH9KlQwpuBCyZDbzqLP3NChx2DTbSm2DlvK7Qn_eOncLxe9_qtcz3j5kVhwq3jtDviscfPuMD)

Hi Svetlana, we have now tested this using an infrared adaptor. We observed negligible signal remaining in the gel (<5%), I attach the image below. The gel is on top (signal is hardly detectable here), and membrane on bottom (where you can see strong signal). So I’m unsure why transfer didn’t work well for you - let’s discuss at the upcoming EMBO conference.

#### ![](https://lh5.googleusercontent.com/_Mx9uz3e-YlXiI22yvcfjyUSN5-Ix8RRQy5iKtzOwN_kH9QIzv3AvXCEwHISbp9UQc5aai2oCDB7GuMUT2YbvTfLMhpB3yihE_D5sD4PYkK1e5cyfzyOSD6hhO_EBe0wwnl-mGUw)

## **Analysis of X-ray film results**

**In our first experiments, we obtained additional radioactive bands that don't seem to respond to RNase treatments. Have you ever seen these kind of bands in your experiments? What do you think these bands could be?**

It is important to analyse results from material that was not crosslinked to evaluate this. If bands are present, then the signal is most likely coming from  direct labelling of proteins (which might be due to a contaminating kinase from lysate, or non-specific PNK activity). If bands require crosslinking, then they might represent proteins crosslinked to microRNAs.

**We see an absence of signal in the low RNase sample around 130kDa, creating a gap in the otherwise nicely diffuse signal corresponding to the shifted protein-RNA complex.  Have you seen such gaps before?**

This is due to migration of antibody at this size on non-reducing gel, which pushes off other proteins. Since your protein had MW>50kDa, you need to use reducing gel to avoid this problem.

**We thought that we could see whether our protein crosslinks to RNA by looking at its shift on the SDS page gel after crosslinking, without labelling RNA. But we never could detect such shift.**

Correct, we also cannot see a shift on Western blot after crosslinking. This is due to low crosslinking efficiency, and also due to the fact that crosslinked complexes migrate as a more diffuse band (depending on the amount of RNAse use, of course).

**After end-labeling RNAs and running the samples in SDS-PAGE/autoradiography. There are clear differences in patterns between low and high RNase treatment, but “high” treatment did not give an accumulated signal at the size of the target protein as expected.**

The high RNAse treatment should give a band slightly above the MW of the protein - it can be as much as 5kDa higher. The RNAse concentration needs to be adjusted to each cell or tissue type, so try a range of concentrations to see which one gives the sharpest band. If concentration is too high, the signal may be lost, because in our hands PNK seems not to phosphorylate RNA well when only one crosslinked nucleotide is left.

## **Phase Lock tubes**

**Is it necessary to use Phase Lock Gel Heavy tube for phase separation?**

It is not necessary for phase separation. The reason for using these tubes is to ensure that no phenol is carried over into the next step. But the protocol can be performed without the Phase Lock Gel Heavy tube, as long as the user is very careful when collecting the aqueous phase.<br>


# Crosslinking

**Can I increase UV intensity like 600 mJ/cm2 or multiple times (5 or 6 times using 400 mJ/cm2 setting)?**

Increases in efficiency with crosslinking longer than recommended in the protocol can be seen, depending on the protein studied. However, we advise to follow the protocol described in Ule et al, Methods 2005, with 150 mL/cm2 setting for monolayer cells. Prolonged cross-linking could initiate a DNA-damage response in the cells, and can lead to crosslinking of multiple proteins to nearby sites on RNAs.  This could decrease the resolution of the method, and cause co-purification of non-specific proteins and other non-specific RNAs that crosslink to these proteins.

**Can frozen tissue be effectively cross-linked?**

Yes, the frozen tissue pieces can be ground with mortar and pestle on liquid nitrogen (see <https://www.youtube.com/watch?v=inpoDyAuyXk>) and then the frozen powder crosslinked. This is helpful especially for tissues that are not soft enough to dissociate using the standard pipetting trituration method described in our protocol. Alternatively, if the frozen tissue is soft enough for trituration after thawing, it can be crosslinked after thawing and dissociation as described in [(Ule et al., 2005)](https://paperpile.com/c/3c2FI2/XHHA).

**To increase the UV crosslinking efficiency,  I wonder if I can first homogenize the tissue with Dounce homogenizer, then do UV crosslinking with the supernatant after spinning, if the homogenization will not affect the complex.**

Homogenisation will certainly affect the protein-RNA interactions, and could lead to non-physiologic interactions with RNAs that normally don’t co-localise with the complex in cells. I don’t think that the  potential increase in cross-link efficiency is worth taking the risk of getting non-informative data.

**What is the exact nature of the covalent bonds between the UV-crosslinked protein and RNA?**

I assume you are referring to the biophysics behind the bond formation? I find the explanation from this paper quite nice [(Erika C Urdanetabenedikt, 2020)](https://paperpile.com/c/3c2FI2/F16B).


# On-bead biochemistry

## **Dephosphorylation**

**For dephosphorylation of RNA 3'ends, pH 6.5 PNK buffer is used, rather than the pH 7.6 buffer, provided by NEB. Have you compared these two conditions internally? Is 5x PNK pH 6.5 buffer available commercially?**

This buffer needs to be prepared by the user. We haven’t compared conditions, but increased phophatase activity of PNK at lower pH has been reported in literature, you can read more in [(Wang and Shuman, 2002)](https://paperpile.com/c/3c2FI2/IR33).

## **RNA ligation**

**For the RNA ligation step, do you think PEG4000 works as well as PEG400?**

PEG4000 works fine in solution, but when used on-bead, it in our hands interfered with immunoprecipitation efficiency. We are not sure why, but it might have a greater tendency to stick to the beads.

**Is it possible to use a 3' linker with a phosphorylated 5' end instead of a pre-adenylated 5' end and adding some ATP during the 3' linker ligation step?**

Yes, just follow the protocol as described in [(König et al., 2010)](https://paperpile.com/c/3c2FI2/mqax). It is important in this case to also dephosphorylate using a phosphatase (as described in original protocol), rather than PNK, because contaminating PNK in the ligation reaction containing ATP would lead to phosphorylation of 5’ ends of RNA, and their self-circularisation.

**Why do you use final 10mM DTT in the ligation buffer, even though NEB uses 1mM final DTT? And why do you make 4x stock of buffer?**

Several companies use 10mM final concentration (including Ambion, Takara and Promega), and this concentration was tested in our hands. DTT also inhibits RNases, therefore we decided to use the higher concentration. Furthermore, DTT is unstable, therefore higher concentration ensures that enough reducing activity is present during overnight incubation. We make 4x stock of buffer to avoid precipitation, which can occur with 10x stock.

## **32P gamma ATP labelling**

**You say to use 0.8 uL of P32 gamma ATP. what concentration of mCi/unit should I use?**

We get stocks of P32 gamma ATP at activity of 3000 Ci/mmol (concentration of 10mCi/ml). We order from here:[ http://www.perkinelmer.com/Catalog/Product/ID/BLU002A100UC](http://www.perkinelmer.com/Catalog/Product/ID/BLU002A100UC)<br>


# Adapters and Primers

## **L3 linker**

**L3 linker DNA does not have a 3'-puromycin**

It has ddC (dideoxycytosine), which works equally well to block ligation.

## **RT primer and random barcode (UMI)**

**Why does the RT primer contain only nine "linker L3 specific" bases for the reverse transcription step and not the maximum (20)? If we don't have the barcode in the RT primer, do you think we should add twenty "linker L3 specific" bases for the reverse transcription instead of nine?**

It's important that the RT primer doesn’t have all 20 of the L3-specific bases, otherwise circularised RT primer could serve as a template for PCR after linearization. Even though we run a cDNA gel after RT, the concentration of primer is so high that it can contaminate the purified cDNAs. We have made great effort in the early stages of iCLIP development, tried RT primers with different numbers of L3-specific bases, and found that having nine bases gave best results.

**Do you somehow test the RT primers with different barcodes to ensure that they all have efficient priming?**

We tested the oligos by splitting a single experiment before RT, and then comparing PCR products at the end. This can performed whenever ordering a new batch of primers. We initially found some variations between primer qualities when ordering non-purified DNA, and this variation was not related to their sequence (it depended on the synthesis batch, and might have to do with residual salt). However, now we order them as HPLC purified, and we don’t find much variation anymore.

**All Rclip 1-16 primers have two consecutive wobbles (NN) at the 5'end immediately followed by a 4bp bar code and another 3bp-long wobble (NNN). What is the purpose of this?**

The wobbles (NNN) are part of unique molecular identifier (UMI), also referred to as random barcode. Their analysis allows to distinguish individual cDNAs from PCR amplicons (see chapter ‘[Use of random barcode in data analysis](https://docs.google.com/document/d/1qSMMjwvFtytVcmyAgcLrktgE2KjMuLK55f1oCcexaWw/edit?hl=en_US#heading=h.1gbpzde00a4h)’). The three wobbles will be positioned at the start of sequencing read, which ensures efficient cluster identification by the Illumina software, which uses the first 4 cycles, and thus high sequence diversity is required in the first nucleotides. The two more wobbles at the 5'-end of the primer are added to avoid potential effects of defined barcodes on the efficiency of cDNA circularization.

We order the primer from IDT, and specify ‘N’ at the corresponding positions in the primer to define these wobbles.&#x20;

**The cut-oligo and Rclip RT primers do not have 5'-amino (6 carbon) group ( you mentioned that they are important in a Methods, CLIP paper, 2009).**

They were important to block ligation of 5’ end of adapter at the time, but here the protocol is now different, and we do not need to block ligation.

**We have recently had a problem with using rt5clip. In two independent experiments, rt5clip samples were taking over 90% of the reads, while the libraries using rt5clip were not necessarily looking different from the other libraries. I also noticed that rt5clip sequence was dropped from the Huppertz et al. 2014 Methods paper. What is the reason for this, did you observe a similar problem?**

Yes, I believe we avoided rt5clip at the time, because we tested each ordered primer and saw this one being an outlier. That said, we interpreted it as variation in quality of synthesis rather than primer sequence, as no primer would reproducibly be an outlier across stocks. In your case, it seems that it is best to exclude your current rt5clip primer stock from future experiments too. Thank you so much for your fast answer, this answers all our worries!

## **Cut oligo**

**What is the purpose of the cut oligo?**

The cut oligo was used in the original variant of iCLIP, and was used to cut the linearized cDNA with BamH1.

**Why does the Cut Primer contain "aaaa" at the end ?**

To prevent it from being a template for PCR<br>

## **PCR primers**

**What sort of primer I should use if I want to start by cloning my insert into TOPO vector instead of doing nextGen sequencing?**

Same primer can be used as described in the protocol. TOPO cloning doesn’t require any specific primer.

**Do P5 and P3 solexa primers bind the library through the same sequence? Doesn't this cause PCR artefacts?**

The last 14 nucleotides of P5 and P3 solexa primers are identical. Interestingly, that doesn’t create any artefacts in our hands.<br>


# Alternative cDNA library prep protocols

We’ve described a comparison of the variant options for the library preparation protocols in our [2018 review](https://authors.elsevier.com/a/1WUmG3vVUP5-cM).

## **Kits as an alternative for library prep**

**What is the feasibility of doing RNA library prep using a commercially available kit (NEB Ultra, or Illumina Tru-seq), after RNA extraction following proteinase digestion? In this way, the isolated RNA will be treated as a normal RNA sample and random hexamer primers would be used for reverse-transcription.**

The pro is that you can use an established kit, so possibly less optimisation needed. The cons are several. Some are described in our review (greater specificity with on-bead ligation, the capacity for non-radioactive visualisation in irCLIP, etc). In addition, I’m not aware of a kit that would ligate the adapter to cDNA to allow amplification of truncated cDNAs. So when using kits, data will be similar to HITS-CLIP, and restricted to readthrough cDNAs. Also, CLIP RNAs are often too short to allow priming with random hexamers. So all in all, use of kits may work in some cases, but could lead to biased data with limited resolution.<br>


# cDNA purification

## **Precipitation**

**At the beginning of step 10 (and 11-13), after spinning for 20 mins, the only pellet I saw was a small blue fleck.  this stayed in tact through the EtOH wash and then dissolved in water.  Is there more of a pellet at the bottom of the tube that I just can't see or is this blue fleck the only thing I have to worry about not removing? I'm afraid there was RNA/cDNA at the bottom of my tube that I pipetted out in my attempt to remove all the alcohol (while being careful to leave the small blue fleck in the tube).**

The blue fleck contains the RNA - so it’s the only thing to worry about ;-)

**When the RNA or cDNA was precipitated and washed after extraction, some samples had relatively large pellets, which might not be completely soluble in the small amount of buffer (7.25uL) used to dissolve the pellet prior to RT.  Obviously it was not made of pure RNA. I would think the pellets are mostly made up with salt or some component of nitrocellulose membrane. Would it be OK to just ignore it and go ahead?**

Yes, the additional precipitate is most likely salt or other contamination. This may inhibit other biochemical steps, so I advice against proceeding. See Huppertz et al, Method 2014 manuscript, which discusses ways how to ensure that precipitations are clean.

## **Size marker**

**We're trying to troubleshoot iCLIP experiments following your Jove protocol... what ladder did you use in Figure 3? We've had some issues where the amplified library size isn't what we expect given the size we're selecting for at the cDNA (pre-circular ligase) step, and we're trying to figure out what the cause is.**

I agree this can be a problem of size selection, because linear cDNA runs more similar to the size of single-stranded RNA than double-stranded DNA marker, which are normally used.  Therefore we suggest to load 6 µl RNA century size marker (Invitrogen AM7140; diluted 1/30, and stored as aliquots at -20) for the cDNA gel.

## **Gel**

**After reverse transcription, cDNA products are resolved in TBE-urea gels and a DNA low molecular weight marker is used (double strand!).  Does this marker run correctly? As single strand? Or as a mixture of complementary bands?.**

If denatured correctly this marker runs as single strands. Denaturing time can be prolonged to 5 min to ensure full denaturation. And don't use too much marker, because when overloaded it doesn't denature fully.

## **Visualisation**

**How did you see the ladder in 11.4?**

You cut off the part of the gel that contains the ladder and stain it with Sybr green II.

## **Purification**

**During purification of the cDNA from the gel, you separately excise the "Low", Medium" and "High"  RNA fractions. Do you sequence them separately as well? If yes, how do you barcode them (i.e. different barcodes for each fractions or the same barcodes for the same fraction from  a different biological replica?) If you pool them together, at what stage do you do it?**

We keep it separate during PCR. We don’t barcode them, but we do mix them after PCR if they all look good. Sometimes the shortest cDNAs contain primer dimmer - in this case we don’t sequence them. We can’t figure out after sequencing which sequence came from which band.<br>


# Enzymes

**Is it necessary to use same companies for enzymes in the protocol, such as Fermentas for BamHI or Invitrogen for polymerase (Accuprime Supermix 1)?**

NEB BamHI would work just as well. Otherwise, we would recommend some testing before changing companies.

**Can the Invitrogen platinum Hot start PCR mix be used over the Invitrogen AccuPrime SuperMix I for the final PCR amplification stage?**&#x20;

Certainly. We ourselves have lately switched to Phusion HF Master mix, which is more efficient (it brings down PCR by several cycles, without much change in library quality).<br>


# PCR Products

**When analysing PCR products, I see a band corresponding to the size of primer dimers, especially in the sample that was cut low from cDNA gel.**

Yes, it is common to see this band in the sample that was cut low from cDNA gel, and sometimes also in other samples. This is due to contamination from short cDNAs that only contain the sequence of RT primer. If this primer dimer is the dominant product on gel, we advise against sequencing the corresponding sample.

**At the end of the procedure, I didn't understand if you purify the final PCR product or if the PCR product is ready for sequencing?**

PCR products are ready. The gel is just for quality control. Alternative methods (bioanalyser) are also possible.

**In the protocol, is the final PCR product quantified before submitting?**

Yes, the PCR product needs to be quantified. We use both qPCR and bioanalyser.

**Can I gel purify PCR products and re-PCR using the same primers?**

Normally, the products of the first PCR should look clean on the gel, otherwise it is a sign of a library that is of low complexity, and is unlikely to generate informative data. Therefore we advise against re-PCR, but it can be done as the last resort.

**After final amplification of the library and gel extraction, there are several distinct bands, and A260/280 indicated there were good amount of DNAs. However, QC results (capillary electrophoresis) provided by the sequencing company indicated there were no peaks in some (if not all) samples (a few samples gave a very sharp peak at \~120bp, but they say “NA” in the section of sample concentration). Since they charge by lanes, not by number of samples, for NGS analysis, I wonder if I should include these samples for sequencing while trying to obtain better samples (quality/quantity) in the second try.**

I advise against proceeding with the poor samples. 120bp peak is primer dimer. If you overamplify your library, there may be additional peaks on the gel at higher MW, but these would be just additional artefacts.  If you mix with good samples, they will just produce useless reads that won’t map. Make sure  you follow the quality control steps described in Huppertz et al, Methods 2014 - especially comparing the PCR products obtained after cutting from different sizes of cDNA gel - the sizes of these PCR products should agree with where cutting was done.

**We wonder if the iCLIP protocol could give problems with small RNAs, as truncated fragments may be too short to harbour useful information.**

We find that if a protein crosslinks to small RNAs, these will be the dominant sequence species in the resulting data. In such case, there will be enough reads for these.

**I have been trying a variation of iCLIP, the irCLIP protocol by Zarnegar et al., and came across your google document with questions and answers. First of all, it is very useful, thank you very much. Secondly, I have tried the protocol with one of the hnRNP proteins and after the library preparation (which I think follows Flynn et al., Fast-iCLIP), I see a strong peak at 148 bp and a fainter smear ranging from \~120-200 bp. The empty PCR products are reported to be 137 bp. I was wondering if this pattern is normal to be seen by bioanalyzer and if not, have you ever experienced a discreet stronger band in your iCLIP final libraries.**

You are right, due to the longer barcodes, the primer dimer band migrates around 137 bp in the irCLIP protocol. I’m unsure what the 148 bp band would be, but it may also be a type of an artefact - usually strong peaks are artefacts. We tend to have a problem with primer artefacts when using the irCLIP protocol, so gel purification is a must. In your case, I guess purification of signal >155bp would probably be best.

**I have been using the iCLIP protocol modified with the irCLIP protocol. After the amplification with the solexa primers I have run my cDNA library on a TBE gel and cut between 155bp and 400bp. When I took the samples for analysis of the tape station what I am seeing is in the picture below. I am not sure if the samples are good enough to send for sequencing or if I have some primer dimers in this. Could I please have some advice on how to proceed please?**

![](https://lh4.googleusercontent.com/tUUZkn5CSCmBu_m2N9n9_7AGSlCOajbtmKk7aFEJRISlzfDH0vtAbXmjZ69qr9_fOkywtVbrNkuOd8GXb-UV4iRM2lLPYJ6a74W4ayE_NawhUrq_EMtwlTxB-Zize_8zHv4xa310)

![](https://lh3.googleusercontent.com/6pRctCsxaEMwazpvvm73AEJEID3_LmhjbvORA3Ah8YNpJbmkQ_9yIYxV1kbDCtPRglKcBL31JISu8NH58BwQkwpJe5qaqQ6pXD1Eix3WnxkVGxwQt8kVcIgcKQUgj_Txt_75sAJO)

The sizes of the first library look a bit unusual (very long products with peak at 272, but are you sure this size is right? - it seems to me that this ‘272’ peak is same as the 139 peak from the 2nd library), but probably fine, but in the 2nd library the peak at 139 are indeed primer artefacts (these we find particularly problematic with irCLIP). So a 2nd gel purification for the 2nd library would be good where you again cut above 155nt (it seems that with the 1st purification you didn’t manage to get rid of the artefact well). If the two libraries have different barcodes, you can possibly mix them before you load the gel again to repurify.<br>


# qPCR

**What primer/probe sequences did you use to quantify the final PCR library with qPCR before submitting to sequencing?**

We use Invitrogen Platinum qPCR Supermix w/ROX and these primers:

DLP: \[6FAM] CCCTACACGACGCTCTTCCGATCT \[TAMRA]

Primer 1: AATGATACGGCGACCACCGAGATC

Primer 2: CAAGCAGAAGACGGCATACGAGATC<br>


# Illumina Sequencing

**Are the solexa primers that are given in the paper designed for a paired-end sequencing?**

Yes, but we use them for single-end sequencing.

**My library is ready, so I’d like to know whether need customized sequencing primer or just the standard primer for Illumina sequencing.**

The standard sequencing primer is used.

**What kind of sequencing was employed? Does a standard Illumina sequencing flow need to be customized?**

We use 75 or 100 cycles single-end sequencing with HiSeq. Standard protocol.

**How many reads do you normally aim for per sample?**

10 million if the quality measures were good (low PCR cycle number, expected size distribution of amplicons). After initial round of sequencing, we often then resequence the library if we see the ratio of cDNA/read count is close to 1 (i.e., most reads have unique UMIs).

**Do you use CLIP-only samples in a flow cell or do you mix with other samples? Our facility doesn't like mixing "unusual" samples together.**

We had designed our custom indexes in such a way that they are different from indexes of all other methods and kits used by the facility at the Crick institute. So we let our facility to decide if they wish to mix with other samples, and in most cases they do. They haven't seen any problems when mixing our 'unusual' iCLIP libraries with others, it  doesn't affect the quality of other samples.<br>


# miCLIP (methylation iCLIP)

**Is cDNA truncation the same in the miCLIP protocol for analyzing m6A positions with Abcam and Sysy antibodies to m6A?**

Yes, [miCLIP can  be used to map the methyl-5-cytosine (m5C) or methyl-6-adenosine (m6A) modifications in RNA by exploiting cDNA truncation](https://www.dropbox.com/s/71dn55vr82gc4sv/2017_miCLIP%20review.pdf?dl=0). In case of antibodies against m6A, a polypeptide is left at the cross-linked nucleotide after proteinase K digestion of the antibody, because the antibody is covalently crosslinked to the RNA in the same way as an RNA-binding protein would be crosslinked in the case of iCLIP. Apart from crosslinking, the rest of the protocol of miCLIP is very similar to iCLIP.


# CLIP-rtPCR (site-specific analysis)

**CLIP-rtPCR is sometimes used to demonstrate RBP interaction with a specific RNA or a site on RNA. Since iCLIP analysis is based on the frequent RT stalling at crosslinking sites, how does that affect the design and interpretation of CLIP-rtPCR? Presumably both forward and reverse PCR primers should be 3’ to crosslinking sites, so that cDNAs can be efficiently produced from the purified RNA fragments? How frequent is the stalling?**

Stalling (i.e., cDNA truncation) depends on many factors, but especially the RT conditions (see <https://pubmed.ncbi.nlm.nih.gov/28790018/>). Various studies tend to estimate it between 3-30% depending on RBP, RT condition, dataset, so quite a wide range. So CLIP-rtPCR may work to some extent even if primers are not just downstream of the crosslinking region. That said, there is also another issue - if applications of CLIP-rtPCR don’t use the SDS-PAGE purification step (which helps remove the non-crosslinked RNA), it’s possible that non-crosslinked RNA will be present, which could be overrepresented in the readout (because there’s no signal loss due to crosslink-induced truncation). It’s also likely that the protein crosslinks to many sites on the RNA, so RNA fragments will exist where crosslink doesn’t overlap with the amplified region (especially if low RNAse is used, so that RNA fragments are long).<br>


# iiCLIP protocol

*Questions pertaining to* [*iiCLIP protocol*](https://www.biorxiv.org/content/10.1101/2021.08.27.457890v1.full) *answered by Flora Lee.*

**I have a couple of questions for you regarding the 3' adapter ligation step (without barcodes, for the moment). I would be very grateful if you could help me understand it better.**

<figure><img src="https://1565883286-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-MbHJ7HdNsgyJurmXKPb%2Fuploads%2F6hpUB9H8NW4eP7QDOUtT%2Fimage.png?alt=media&amp;token=241a7f89-7a04-40bf-b847-ceb51a1dbea4" alt=""><figcaption><p><em>irCLIP adapter substrate: /5Phos/AG ATC GGA AGA GCG GTT CAG AAA AAA AAA AAA /iAzideN/AA AAA AAA AAA A/3Bio/ (same as Zarnegar et al, Nat methods, 2016).</em></p></figcaption></figure>

<figure><img src="https://1565883286-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-MbHJ7HdNsgyJurmXKPb%2Fuploads%2FHiyvaiBgLt7LhbGPWEc1%2Fimage.png?alt=media&amp;token=8af74159-c2dd-4b0f-85f1-2ec2f7096a41" alt=""><figcaption><p><em>Construction schematic for generation of irCLIP DNA adapter (no need for phosphorylation step if I buy the oligo with 5’ phosphate from IDT)</em></p></figcaption></figure>

**Q1)  I understand that the pre-adenylation step is necessary to adenylate the 5' end of the adaptor so that it can covalently bind to the 3' dephosphorylated end of the RNA fragments. IDT will custom adenylate oligonucleotides, why don't you recommend buying the sequence already pre-adenylated?**\
\
This is for cost reasons. It is much more expensive to order adenylated oligos and not guaranteed yield. Whereas the 5’phosphate modification is inexpensive.&#x20;

**Q2) I see from your bioRxiv 2021 preprint that the introduction of enzymatic RecJ adapter removal after 3′ adapter ligation is meant to improve the efficiency of downstream steps by minimising artefacts that can be caused by adapter carry-over. However, from a chemical point of view, I am not clear which part of the adapter is removed in this step.**\
\
RecJ is DNA specific from the 5’ end, so will degrade the adapter from the 5’ end. If the adapter is ligated to RNA, the 5’ end will be RNA, hence ligated products are protected.

&#x20;**Q3) what do you need the biotin at the end of the IR L3 adapter for? to capture and purify final cDNAs with streptavidin beads? Or for chemiluminescence RNA detection as in** [**seCLIP 2022**](https://eur03.safelinks.protection.outlook.com/?url=https%3A%2F%2Fwww.nature.com%2Farticles%2Fs41596-022-00680-z\&data=05%7C02%7Ccharlotte.capitanchik%40kcl.ac.uk%7C8a76993657d34dacb8f808dc36bf0456%7C8370cf1416f34c16b83c724071654356%7C0%7C0%7C638445441248809756%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C0%7C%7C%7C\&sdata=lVMIUOPZ5%2B0kM%2B9I0GXNvhBOxrFiqGglZbwRlcy%2FK54%3D\&reserved=0)**?**

For our iiCLIP protocol, you don’t need the biotin. You can replace it with something like /3ddC/ or any other modification that will block ligation (to minimise adapter concatamers).


# Demultiplexing

**Are the indices in line with read 1? They are also shorter than what our sequencing facility normally uses, are they hard to demultiplex?**

We developed a tool for demultiplexing libraries with complex barcoding (including in-line barcodes), as is the case of iCLIP. The tool is called Ultraplex, it's fast and easy to use, please find here: <https://wellcomeopenresearch.org/articles/6-141>. See Fig1 for an example read from a typical iCLIP library, and the demultiplexing workflow. This tool is also implemented within <https://imaps.goodwright.com/>, a free platform for password-protected data analysis dedicated especially to CLIP.


# Use of random barcode in data analysis

**I am interested in the evaluation of random barcodes, which I can't completely understand. The barcode marks individual cDNA, but how can the barcode solve the problem of PCR artefacts? For example, if there are two barcode at a particular position, one barcode having 100 reads, another barcode having 200 reads, then the total reads for both barcodes should be the same if PCR efficiency is the same for all cDNAs, so you should choose the minimal number of reads, i.e., 100 reads. This is my guessing, I don't know if it's correct.**

A good example of random barcode analysis is the Fig 1C in “iCLIP Predicts the Dual Splicing Effects of TIA-RNA Interactions by Wang et al, PLOS biology, 2010”. This shows you the random barcode for each sequence, and the number of sequences that had the same barcode is shown in the brackets. If multiple sequences mapping to the same position in the genome have the same random barcode, then they are all counted as 1. In your example, you have only two different random barcodes or sequences mapping to the same position, so the cDNA count = 2. Such analysis can properly correct for PCR artefacts.


# Mapping to repetitive elements/RNAs

**How do I ensure that I accurately represent repetitive RNAs (rRNA, tRNA, snRNA, repetitive elements etc) in my analysis?**

If you perform a standard genome mapping and only take singly mapping reads, it is likely you will misrepresent the ncRNA portion of your CLIP library. This is because RNA species such as rRNA, tRNA and snRNA have many copies in the genome, that often have a high sequence identity. This means that it is possible to map short CLIP reads to multiple gene copies, and if you exclude multi-mapping reads, then these reads will be excluded. Another thing that makes this difficult is that in the case of rRNA and tRNA, genome annotations are likely incomplete, and reads that you think might be pre-mRNA, might actually come from an unannotated intronic tRNA for example (see (Schwartz, 2018) for an example of where this can become contentious).&#x20;

One approach to this issue is to randomly assign multi-mapping reads to one of the genomic locations where they map. The issues with this approach arise when it comes to gene by gene quantification - a gene with multiple high identity copies in the genome will be penalised in terms of counts because multi-mapping reads will be spread amongst all copies. In addition, the assignment of a read to certain groups can become complicated, for example if a read maps between tRNA, introns, intergenic spaces, for example, then you will be forced to come up with a hierarchy.&#x20;

Other approaches that look to solve this issue involve some kind of pre-mapping. This means mapping reads to certain groups of RNA species before mapping to the whole genome. This effectively reduces the sequence space available for reads to map to and so there are several considerations in doing this: if I am too lenient in terms of mismatches .etc then I may map a read to (say) tRNA, that could map much better without mismatches to somewhere else in the genome, however if I am too stringent, then I will fail to assign a true tRNA read as “tRNA”. Pre-mapping involves some assumption making, in that it is typical to map to the most abundant ncRNA first (rRNA, tRNA). Because they are the most abundant it makes sense that a read that could map to these RNAs probably does originate from these species - but this might not always be the case. Even pre-mapping will leave you with problems when it comes to individual gene quantification. How best, for example, to categorise reads mapping to some combination of the 193 annotated U1 snRNA genes in the human genome? We can use some knowledge of biology to help us make sensible groupings.&#x20;

In the case of the FAST-iCLIP pipeline, the authors decide to quantify tRNAs at the level of anticodon groups for example, rather than individual genes. A further, CLIP-specific issue, is assigning the location of crosslinks to reads that multimap within ncRNA. One option here is to make metaprofiles over groups of ncRNA. Another solution is to map to single representative copies of specific species, for example to map to one copy of the 45s rDNA cluster in the case of rRNA. Some software tools that aim to help in this area are:&#x20;

* [**FAST-iCLIP**](https://github.com/ChangLab/FAST-iCLIP)&#x20;
* [**CLAM**](https://github.com/Xinglab/CLAM)

*written by Charlotte Capitanchik*


# Differential CLIP binding analysis

#### **How many replicates do I need?**

One measure to address this question is to calculate correlation between replicates, which helps to determine the number required to make a statistical analysis meaningful i.e. in the case that your replicates are very variable, then you will need more replicates.

*written by Charlotte Capitanchik*&#x20;

#### **How can I control for changes in gene expression between my conditions?**

Within DeSeq2 and EdgeR it is possible to include a covariate in your analysis, enabling you to include an expression dataset alongside the CLIP data. Additionally some tools that were originally developed for MeRIP-Seq appear to also work in the case of CLIP and specifically have the capacity to include an expression dataset (for example [QNB](https://bmcbioinformatics.biomedcentral.com/articles/10.1186/s12859-017-1808-4)). The various approaches tested for MeRIP-Seq analysis [(McIntyre et al., 2019)](https://paperpile.com/c/3c2FI2/8eQz) are thus inspiring for CLIP analysis too. In my own hands, I find that the DeSeq2 approach is the most forgiving.

*written by Charlotte Capitanchik*&#x20;

**If you compare RNA binding targets of a protein between different tissues, does the data need to be adjusted for differential gene expression across tissues?**

Often, a change in binding targets between tissues will be due to differential gene expression. To try and distinguish whether this is the case, you might want to correlate the log2 fold change in gene expression between two tissues with the log2 fold change in CLIP peak signal. To perform an analysis that adjusts for differential gene expression see above, but this is probably more relevant in a treatment vs. control situation than in a tissue vs. tissue comparison.

*written by Charlotte Capitanchik*<br>


# To be Answered

* **How to accurately compare two CLIP libraries of the same RBP with \~10-fold difference of CLIPped RBP abundance in the cell? How can we account for RBP/target RNA ratio?**
* **Is it necessary to use spike ins, and if so, what kind of spike ins and how should I analyse this?**
* **How do I choose/remove redundant annotations from genome annotation files?**&#x20;
* **How can I compare multiple RBP baits with different IP efficiencies?**
* **What are the best ways to access and analyze published CLIP data?**
* **What available pipelines are out there for analysing different CLIP datasets (iCLIP & PAR-CLIP)? Are there advantages/disadvantages to different established pipelines?**
* **Which statistical methods are best at identifying cross-link sites?**<br>


# References

[**Chakrabarti, A.M., Haberman, N., Praznik, A., Luscombe, N.M., and Ule, J. (2018). Data Science Issues in Studying Protein–RNA Interactions with CLIP Technologies. Annu. Rev. Biomed. Data Sci. 1, 235–261.**](http://paperpile.com/b/3c2FI2/2SQ6)

[**Chen, X., Castro, S.A., Liu, Q., Hu, W., and Zhang, S. (2019). Practical considerations on performing and analyzing CLIP-seq experiments to identify transcriptomic-wide RNA-protein interactions. Methods 155, 49–57.**](http://paperpile.com/b/3c2FI2/dLnN)

[**De, S., and Gorospe, M. (2017). Bioinformatic tools for analysis of CLIP ribonucleoprotein data. Wiley Interdiscip. Rev. RNA 8.**](http://paperpile.com/b/3c2FI2/YsbD)

[**Erika C Urdanetabenedikt (2020). Fast and unbiased purification of RNA-protein complexes after UV cross-linking. Methods 178, 72–82.**](http://paperpile.com/b/3c2FI2/F16B)

[**Garzia, A., Meyer, C., Morozov, P., Sajek, M., and Tuschl, T. (2017). Optimization of PAR-CLIP for transcriptome-wide identification of binding sites of RNA-binding proteins. Methods 118-119, 24–40.**](http://paperpile.com/b/3c2FI2/A8KI)

[**Haberman, N., Huppertz, I., Attig, J., König, J., Wang, Z., Hauer, C., Hentze, M.W., Kulozik, A.E., Le Hir, H., Curk, T., et al. (2017). Insights into the design and interpretation of iCLIP experiments. Genome Biol. 18.**](http://paperpile.com/b/3c2FI2/2HpP)

[**König, J., Zarnack, K., Rot, G., Curk, T., Kayikci, M., Zupan, B., Turner, D.J., Luscombe, N.M., and Ule, J. (2010). iCLIP reveals the function of hnRNP particles in splicing at individual nucleotide resolution. Nat. Struct. Mol. Biol. 17, 909–915.**](http://paperpile.com/b/3c2FI2/mqax)

[**Lee, F.C.Y., and Ule, J. (2018). Advances in CLIP Technologies for Studies of Protein-RNA Interactions. Mol. Cell 69, 354–369.**](http://paperpile.com/b/3c2FI2/jqZg)

[**McIntyre, A.B.R., Gokhale, N.S., Cerchietti, L., and Jaffrey, S.R. (2019). Limits in the detection of m6A changes using MeRIP/m6A-seq. BioRxiv.**](http://paperpile.com/b/3c2FI2/8eQz)

[**Moore, K.S., and AC’t Hoen, P. (2019). Computational approaches for the analysis of RNA–protein interactions: A primer for biologists. J. Biol. Chem.**](http://paperpile.com/b/3c2FI2/Mj9Z)

[**Schwartz, S. (2018). m1A within cytoplasmic mRNAs at single nucleotide resolution: a reconciled transcriptome-wide map. RNA 24, 1427–1436.**](http://paperpile.com/b/3c2FI2/lWgc)

[**Smith, T., Heger, A., and Sudbery, I. (2017). UMI-tools: modeling sequencing errors in Unique Molecular Identifiers to improve quantification accuracy. Genome Res. 27, 491–499.**](http://paperpile.com/b/3c2FI2/s7kA)

[**Ule, J., Jensen, K., Mele, A., and Darnell, R.B. (2005). CLIP: a method for identifying protein-RNA interaction sites in living cells. Methods 37, 376–386.**](http://paperpile.com/b/3c2FI2/XHHA)

[**Ule, J., Hwang, H.-W., and Darnell, R.B. (2018). The Future of Cross-Linking and Immunoprecipitation (CLIP). Cold Spring Harb. Perspect. Biol. 10.**](http://paperpile.com/b/3c2FI2/rPrX)

[**Wang, L.K., and Shuman, S. (2002). Mutational analysis defines the 5’-kinase and 3'-phosphatase active sites of T4 polynucleotide kinase. Nucleic Acids Res. 30, 1073–1080.**](http://paperpile.com/b/3c2FI2/IR33)[**Wheeler, E.C., Van Nostrand, E.L., and Yeo, G.W. (2018). Advances and challenges in the detection of transcriptome-wide protein-RNA interactions. Wiley Interdiscip. Rev. RNA 9.**](http://paperpile.com/b/3c2FI2/DBLj)


