Showing posts with label Molecular biology. Show all posts
Showing posts with label Molecular biology. Show all posts

Friday, July 16, 2010

Tangled lovers: the true tale of siRNA…

Recap:

So, last time, I mentioned how DNA was transcribed into RNA that itself was translated into proteins. We also spend some time on how complementary nucleotide sequences bonded each others (if you are not familiar with the subject, go check it out, we will wait).



Role in gene regulation:

So, if you look at the second step of the central dogma, you can see the recently translated mRNA floating around, patiently waiting for the ribosomal complex to bind it and start transcribing it into protein…

Except that, at the last minute, a short RNA sweeps in and bind the mRNA instead. This block the binding site for the ribosomes and prevent the translation (that is called “RNA interference”).

Worse, viruses are the only organisms whose genome can sometime be found in the form of double stranded RNA so these structures are quickly targeted by dicer proteins that cut them into pieces.

This constitute a somewhat crude but a efficient mechanism to control the level of expression of a gene referred to as “posttranscriptional gene silencing”.

Bacterial CRISPRs:

A very cool application of RNA interference might be immunological.

There are sequences, present on most bacterial genomes, called (Clustered Regularly Interspaced Short Palindromic Repeats) that, when they were first discovered, nobody really understood what they did. Then somebody noticed that, at their center, was a short sequence that looked really like that of a bacteriophage (the term for any virus infecting bacteria).

At which point it was a small step to think that these CRISPRs might be involved in RNA interference with their matching sequence on the viral genome and, sure enough, it was quickly demonstrated that the presence of these sequence correlated with resistance against virus invasion. Moreover, knocking down these sequences rendered the previously resistant bacteria susceptible to infection.

It was a pretty cool find all by itself. But then, somebody decided to take a susceptible bacterial culture, infect it with a bacteriophage and look for survivors. Sure enough, a few bacteria did survive infection. Furthermore, when challenged again with the same virus, all the descendants from these survivors appeared resistant to the virus. At this point, looking back at the genome, the researcher noticed that the bacteria now harbored a shining new CRISPR region it didn’t before the infection, and this region did correspond to the sequence of the virus. In short, they had evidences suggesting that the bacteria, somehow, had been able to acquire part of the virus sequence and use it for protection. Technically, you could describe that as ‘acquired immunity’, if the term was not already taken.

In addition to this complementary DNA sequence, the CRISPRs are surrounded by a variety of sequences, termed Cas (for CRISPRs associated genes). These genes are extremely heterogeneous and code for a variety of proteins, more than 40 families of Cas have been described, several of which appear to be DICERS, molecules that specialize in cutting nucleotide sequences. Interestingly, Cas are very well distributed among the bacteria suggesting that a lot of horizontal transmission, gene passing between bacteria, is taking place.

The means through which these viral sequences are then integrate to the bacterial genome are not as of yet known. It is likely that some Cas are responsible, it might, for example, be an additional function of some of the DICERs…

Immunological role:

Interestingly, a mechanism, called RISC, had been discovered in eukaryote before, and it’s really cool:

So, here you have your strand of viral RNA, it’s either double stranded or it is single stranded, but will become, if briefly, at some point in the process of copying itself.

At any rate there is this protein floating around in the cytoplasm called a DICER (cool name, right? It’d make a great name for a super-hero. Or maybe one of these knives sold on TV infomercial). Anyway, the DICER’s job is to find and bind such a double stranded RNA and then, as its name suggest, it cut if off into pieces (ok, it would be an old school super-hero, like the Spectre but maybe he could have a team-up with the Punisher, or something). But the DICER’s job doesn’t stop there, it also pick up a small portion of the RNA strand that he just made a mess of.

Then the dicer bound by a molecule termed TRBP (for human immunodeficiency virus Transactivating Response RNA-Binding Protein) that apparently functions as a matchmaker, recruiting the Argonaute2 protein.

This second protein loads the RNA sequence into its groove. Then it destroys one of the strands and starts floating around in the cytoplasm, still carrying one strand of the RNA, now termed the “guide strand” the DICER-TRBP helped loading. This way, when it encounters a RNA sequence which is complementary from the guide strand, this complementary sequence (the “passenger strand”) will be bound and then cleaved and, in effect, the RISC acts as a molecular targeting system for the argonaute2 protein.

This mechanism presents some similarities with the one recently described for the CRISPRs. Furthermore, it is very well spread all over the eukaryotic kingdom, it is very common and im)portant in plants, but has also been described in humans. This suggests that it involved in a common ancestor, far away don’t the evolutionary tree. In fact, it is even quite possible that this ancestor was a bacteria and that the RISC system is but an evolutionary refinement of the good ol’ CRIPRs…

A few stuff about molecular biology.

I wanted to talk about a few of the really cool stuffs that one can find out by studying molecular biology.

These can be really neat and often are living witnesses of our most distant evolutionary past, carried deep in the heart our own cells.

But, yeah, when I started about thinking how to write up the coolest ones, I realized that I’d probably do a small recap of the bases, so here it comes.



Recap, the central dogma of molecular biology:

What we refer to as ‘the central dogma of molecular biology’ was first proposed by one Francis Crick of doubly helical fame. In short, it describes how genetic information is contained into DNA, how this DNA is transcribed into RNA and how this DNA is then translated into proteins.

One could, with a little smirk, describe it as the product of a simple time, because quite a few exceptions to this dogma have since been found, as always, Nature is a messy place. Still, it works well enough.



Now, the four DNA bases tend to bind to each other in pair: cytosine tends to bind guanine and thymine tends to bind adenine (in RNA, thymine is absent and instead another closely similar molecule, uracil, binds adenine). That means that complementary sequences will naturally tend to bind each other: CGTACGTA will naturally tend to bind GCATGCAT sequences, in fact, this is the combined forces of all these tiny base-to-base liaisons that bind together the two branches of the double helix (in fact, two complementary sequences of DNA).


Here you see the two helices bound together by the meeting of the nucleotide pairs

(image ruthlessly pilferaged from Genomes -third Edition- by T. A. Brown; University of Manchester)



Another point worth mentioning is that both lesion are not equal, cytosine that interacts with guanine through three hydrogen bonds will form stronger bonds than adenine that only interacts with thymine or uracil through two hydrogen bonds.



Here are the formula of the five nucleotide bases, also showing the interaction between the DNA base pairs, weaker between adenine and thymine or uracil than between cytosine and guanine.



The transcription and transcription termination in prokaryotes.

Now, imagine, if you will, the single lonely chromosome of a simple bacterium. Getting closer, we can discern the long flowing chain of this double stranded chain of DNA. Floating by is the bulky shape of a RNA polymerase, a complex of 5 different sub-units. Another molecule, the Sigma factor 70, pass by and start interacting with the RNA polymerase. These interaction slowly nudge the RNA polymerase toward the chromosome, and more precisely toward the -35 box (a sequence located 35 nucleotide upstream from the gene of interest). Then, the RNA polymerase finally bind the DNA, forming the ‘promoter complex’. It then starts pulling the DNA toward it, in the process, separating the two branches downstream and accumulating the energy it will need (imagine unwinding two rubber bands, tightly wound one on to the other). This lead to the formation of a ‘open promoter complex’ where, just uphead from the transcription complex, one may observe a short region where the two strings of DNA are separated (we call it the transcription bubble), a bit like the zipper of an overloaded bag, where the two half of the zipper start to break apart.

Inside the bubble, base pairs are added on after the other that match the ones on the DNA strand and, slowly, the RNA messenger is constructed that form a sequence complementary to the DNA. The two strands interact with each other, forming an ephemeral liaison that help stabilizing the transcription complex.



Here is a nice schema to illustrate this (taken from Principle of Biochemistry, Lehninger et al., 2000: get it, it’s a standard).



But now, it’s when it becomes really cool. At this point, the RNA arrives at the end of the gene. How does the transcription complex ‘know’ where to stop?

Well, there are two big ways in bacteria and the coolest part is common in both. You see, by the end of the gene there is what we call a ‘palyndromic sequence’. In English, a palindrome is a word or a sentence that can be read from both left to right and right to left, like the word, ‘radar’ for example. In molecular biologist English, a palindrome is a sequence that is identical when read in one direction to the one on the complementary strand when read in the opposite direction: for example CGTTAACG is palindromic, its complementary sequence would be GCAATTGC which, read from right to left, become CGTTAACG.

The importance of this palyndromic sequence is that it is complementary with itself. Look at it, the first pair (C) will bind the last one (G), the second pair (G) will bind the second to last (C) and the two Ts will bind the two As

Now, if we have a long palindromic sequence, the left and right portion of the sequence will tend to bind with each other, forming what is called a ‘hairpin loop’.






This hairpin loop fold onto itself, instead of interacting with the DNA sequence, which would stabilize the transcription complex. This destabilize the whole complex and slow down the transcription process.

Then, two things can happen. In about half the genes, this sequence is followed by a long sequence of adenine. As we mentioned, the adenine-thymine and adenine-uracil pairs are weaker than the cytosine guanine ones. This further weaken the binding of the mRNA on the DNA strain to the point where the two strands break up, interrupting the transcription process.

The other relies on what is call a Rho factor which is another helicase (it breaks bond between nucleotide pair). Essentially, he Rho binds the transcribed DNA sequence a bit upstream from the RNA polymerase and start moving in the same direction at roughly the same speed. Normally, the RNA polymerase own migration along the DNA strand allows it to stay ahead. But, as we mentioned, the RNA polymerase is slowed down at the hairpin level, this allows the Rho factor to catch up and start breaking up the pair lesions at the level of the RNA polymerase. This, in turn, leads to the breaking up of the transcription complex and the end of the transcription…


Conclusion

So, there you go, it’s certainly not complete, but reading that should give you the basis for understanding the really cool stuff I’d like to talk about next (starting with small interfering RNAs).