Chapter Notes

Molecular Basis of Inheritance
25 min read

The DNA

Deoxyribonucleic acid, or DNA, is a long molecule made up of repeating units called deoxyribonucleotides. It serves as the primary genetic material for most living organisms, carrying the instructions for development, functioning, growth, and reproduction. The length of a DNA molecule, measured in the number of nucleotides or base pairs (bp), is a unique characteristic of each organism.

Example
  • Bacteriophage ϕ×174\phi \times 174 has 5386 nucleotides.
  • Bacteriophage lambda has 48502 base pairs (bp).
  • Escherichia coli has 4.6×1064.6 \times 10^6 bp.
  • The haploid content of human DNA is 3.3×1093.3 \times 10^9 bp.

Structure of a Polynucleotide Chain

A polynucleotide chain is a polymer formed by linking many nucleotides together. Each nucleotide has three core components:

  1. A pentose sugar (deoxyribose in DNA, ribose in RNA).
  2. A nitrogenous base.
  3. A phosphate group.

The nitrogenous bases are of two types:

  • Purines: Adenine (A) and Guanine (G).
  • Pyrimidines: Cytosine (C), Thymine (T), and Uracil (U).
Note
In DNA, the bases are A, G, C, and T. In RNA, Thymine (T) is replaced by Uracil (U).

A nucleoside is formed when a nitrogenous base attaches to the 1' carbon of the pentose sugar via an N-glycosidic linkage. When a phosphate group attaches to the 5' carbon of this nucleoside via a phosphoester linkage, a nucleotide is formed.

To build the chain, two nucleotides are linked together by a 3'-5' phosphodiester linkage. This process repeats, forming a long polynucleotide chain with a distinct polarity.

  • The 5'-end has a free phosphate group at the 5'-carbon of the sugar.
  • The 3'-end has a free hydroxyl (-OH) group at the 3'-carbon of the sugar. The sugar and phosphate groups form the backbone of the chain, with the nitrogenous bases projecting outwards from it (Figure 5.1).

The Double Helix Structure of DNA

In 1953, James Watson and Francis Crick proposed the famous Double Helix model for DNA structure, building upon the X-ray diffraction data from Maurice Wilkins and Rosalind Franklin, and Erwin Chargaff's observations on base pairing. Chargaff noted that in double-stranded DNA, the ratio of Adenine to Thymine and Guanine to Cytosine is always one.

Salient features of the Double-helix structure of DNA:

  • It consists of two polynucleotide chains. The backbone is made of sugar and phosphate, and the nitrogenous bases project inward.
  • The two chains have anti-parallel polarity, meaning if one strand runs in the 535' \rightarrow 3' direction, the other runs in the 353' \rightarrow 5' direction.
  • The bases on the two strands are paired through hydrogen bonds (H-bonds).
    • Adenine (A) forms two hydrogen bonds with Thymine (T).
    • Guanine (G) forms three hydrogen bonds with Cytosine (C).
  • This specific pairing means a purine always pairs with a pyrimidine, which keeps the distance between the two strands uniform.
  • The two chains are coiled in a right-handed fashion. The pitch of the helix is 3.4 nm, with roughly 10 bp in each turn. The distance between adjacent base pairs is approximately 0.34 nm.
  • The stacking of one base pair over another, in addition to the H-bonds, provides stability to the helical structure.

This structure's key feature is complementarity. Because of the base pairing rules, if you know the sequence of one strand, you can predict the sequence of the other. This immediately suggested how DNA could be copied, a crucial aspect for a genetic material.

Central Dogma of Molecular Biology

Proposed by Francis Crick, the Central Dogma states that genetic information flows in one direction: DNATranscriptionRNATranslationProtein\text{DNA} \xrightarrow{\text{Transcription}} \text{RNA} \xrightarrow{\text{Translation}} \text{Protein}

In some viruses, this flow can be reversed (from RNA to DNA), a process known as reverse transcription.

Packaging of DNA Helix

A typical mammalian cell contains about 2.2 meters of DNA, which must fit inside a nucleus that is only about 10610^{-6} m in diameter. This requires incredible compaction.

In Prokaryotes:

  • Prokaryotes like E. coli lack a defined nucleus. Their DNA is found in a region called the nucleoid.
  • The negatively charged DNA is held together with some positively charged proteins, organized into large loops.

In Eukaryotes:

  • The organization is far more complex. It involves a set of positively charged proteins called histones, which are rich in basic amino acids like lysine and arginine.
  • Histones are organized into a unit of eight molecules called a histone octamer.
  • The negatively charged DNA wraps around the positively charged histone octamer to form a structure called a nucleosome. A typical nucleosome contains about 200 bp of DNA.
  • These nucleosomes are the repeating units of chromatin, which under an electron microscope looks like a "beads-on-string" structure (Figure 5.4b).
  • This chromatin fiber is further coiled and condensed to form chromosomes, a process that requires an additional set of proteins called Non-histone Chromosomal (NHC) proteins.

Chromatin exists in two forms in the nucleus:

  • Euchromatin: Loosely packed, stains light, and is transcriptionally active.
  • Heterochromatin: Densely packed, stains dark, and is transcriptionally inactive.

The Search for Genetic Material

For a long time, scientists debated whether protein or DNA was the genetic material. Several key experiments settled this question.

Transforming Principle

In 1928, Frederick Griffith's experiment with Streptococcus pneumoniae bacteria provided the first clue. This bacterium has two strains:

  • S strain (smooth): Has a protective mucous coat and is virulent (causes pneumonia and death in mice).
  • R strain (rough): Lacks the coat and is non-virulent (mice live).

Griffith's observations:

  1. Injecting live S strain into mice \rightarrow Mice die.
  2. Injecting live R strain into mice \rightarrow Mice live.
  3. Injecting heat-killed S strain into mice \rightarrow Mice live.
  4. Injecting a mixture of heat-killed S strain and live R strain into mice \rightarrow Mice die.

From the dead mice in the fourth step, Griffith recovered living S strain bacteria. He concluded that some 'transforming principle' from the heat-killed S bacteria had been transferred to the R bacteria, enabling them to make a smooth coat and become virulent. He suspected this was the genetic material, but couldn't identify it.

Biochemical Characterisation of Transforming Principle

From 1933-44, Oswald Avery, Colin MacLeod, and Maclyn McCarty worked to identify Griffith's 'transforming principle'. They purified biochemicals (proteins, DNA, RNA) from heat-killed S cells.

  • They discovered that only DNA from the S bacteria could transform R bacteria into S bacteria.
  • Treating the mixture with protein-digesting enzymes (proteases) or RNA-digesting enzymes (RNases) did not stop the transformation.
  • However, treating the mixture with DNA-digesting enzymes (DNase) did inhibit transformation.

They concluded that DNA is the hereditary material. However, not all biologists were convinced.

The Genetic Material is DNA

The definitive proof came from the 1952 experiment by Alfred Hershey and Martha Chase, who worked with bacteriophages (viruses that infect bacteria). They wanted to find out if protein or DNA from the virus entered the bacterial cell to direct the production of new viruses.

Experimental Setup:

  1. Batch 1: They grew viruses on a medium with radioactive sulfur (35S^{35}\text{S}), which gets incorporated into proteins (as DNA does not contain sulfur).
  2. Batch 2: They grew viruses on a medium with radioactive phosphorus (32P^{32}\text{P}), which gets incorporated into DNA (as proteins do not contain phosphorus).

Procedure:

  • The radioactive phages were allowed to infect E. coli bacteria.
  • The viral coats were separated from the bacteria by agitating them in a blender.
  • The mixture was centrifuged to separate the heavier bacteria from the lighter virus particles.

Results (Figure 5.5):

  • Bacteria infected with viruses having radioactive protein (35S^{35}\text{S}) were not radioactive. The radioactivity was found in the supernatant with the virus particles.
  • Bacteria infected with viruses having radioactive DNA (32P^{32}\text{P}) were radioactive.

Conclusion: DNA, not protein, is the genetic material that is passed from the virus to the bacteria.

Properties of Genetic Material (DNA versus RNA)

To serve as a genetic material, a molecule must fulfill four criteria:

  1. It should be able to replicate. Both DNA and RNA can do this due to their complementary base pairing.
  2. It should be chemically and structurally stable. DNA is more stable than RNA. The 2'-OH group on every nucleotide in RNA is reactive and makes RNA easily degradable. The presence of Thymine instead of Uracil also adds to DNA's stability.
  3. It should provide scope for slow changes (mutation) required for evolution. Both DNA and RNA can mutate. However, RNA, being unstable, mutates at a faster rate.
  4. It should be able to express itself in the form of 'Mendelian Characters'. RNA can directly code for protein synthesis. DNA is dependent on RNA for this process.
Note
While both can function as genetic material, DNA is better suited for storing genetic information due to its stability. RNA is better for the transmission of genetic information and performing dynamic functions.

RNA World

The evidence suggests that RNA was the first genetic material. Essential life processes like metabolism, splicing, and translation evolved around RNA. RNA could act as both a genetic material and a catalyst (in the form of ribozymes). However, because it was a catalyst, it was reactive and unstable. Therefore, DNA evolved from RNA through chemical modifications that made it more stable, making it a better molecule for storing genetic information for the long term.

Replication

Watson and Crick's model immediately suggested a mechanism for DNA copying, which they termed semiconservative DNA replication.

The Scheme: The two strands of the parental DNA separate, and each strand acts as a template for the synthesis of a new, complementary strand. After replication, each of the two daughter DNA molecules contains one original (parental) strand and one newly synthesized strand.

The Experimental Proof

This model was proven in 1958 by Matthew Meselson and Franklin Stahl using E. coli.

Experiment:

  1. They grew E. coli for many generations in a medium containing a heavy isotope of nitrogen, 15N^{15}\text{N}. The DNA of these bacteria became uniformly "heavy".
  2. They then transferred these bacteria to a medium with the normal, lighter isotope, 14N^{14}\text{N}.
  3. They collected samples after 20 minutes (one generation) and 40 minutes (two generations) and separated the DNA using cesium chloride (CsCl) density gradient centrifugation.

Results (Figure 5.7):

  • After 20 minutes: The extracted DNA had a hybrid or intermediate density, halfway between heavy 15N^{15}\text{N}-DNA and light 14N^{14}\text{N}-DNA.
  • After 40 minutes: The extracted DNA consisted of equal amounts of hybrid DNA and light DNA.

This experiment confirmed that DNA replication is semiconservative. A similar experiment by Taylor and colleagues on Vicia faba (faba beans) showed that DNA in chromosomes also replicates semiconservatively.

The Machinery and the Enzymes

DNA replication is a complex process requiring a host of enzymes.

  • The main enzyme is DNA-dependent DNA polymerase, which uses a DNA template to synthesize a new DNA strand. It polymerizes nucleotides very rapidly (around 2000 bp per second in E. coli) and with high accuracy.
  • Deoxyribonucleoside triphosphates serve a dual purpose: they act as substrates and provide the energy for polymerization.
  • Replication occurs at a replication fork, a Y-shaped structure where the DNA helix is unwound.
  • DNA polymerase can only synthesize in the 535' \rightarrow 3' direction. This leads to a complication:
    • On the template strand with 353' \rightarrow 5' polarity, synthesis is continuous.
    • On the template strand with 535' \rightarrow 3' polarity, synthesis is discontinuous, occurring in short fragments (Okazaki fragments). These fragments are later joined by the enzyme DNA ligase.
  • Replication does not start randomly but at a specific site called the origin of replication (ori).

In eukaryotes, DNA replication takes place during the S-phase of the cell cycle and must be tightly coordinated with cell division.

Transcription

Transcription is the process of copying genetic information from one strand of DNA into RNA. It is governed by the principle of complementarity, with one key difference: Adenine in the DNA template pairs with Uracil (U) in the RNA, not Thymine.

Unlike replication, where the entire DNA is copied, transcription only copies a segment of DNA, and only one of the two strands.

Transcription Unit

A transcription unit in DNA has three main regions:

  1. A Promoter: A DNA sequence located upstream (at the 5'-end of the coding strand) that provides the binding site for RNA polymerase. It defines which strand will be the template.
  2. The Structural gene: The segment of DNA that is actually transcribed into RNA.
  3. A Terminator: A sequence located downstream (at the 3'-end of the coding strand) that signals the end of transcription.

Within the structural gene, the two DNA strands are named:

  • Template Strand: The strand with 353' \rightarrow 5' polarity that is used as the template for RNA synthesis.
  • Coding Strand: The strand with 535' \rightarrow 3' polarity. It is not used as a template, but its sequence is identical to the synthesized RNA (with T instead of U). All references are made with respect to this strand.

Transcription Unit and the Gene

A gene is the functional unit of inheritance. In terms of DNA sequence, a cistron is a segment of DNA that codes for a polypeptide.

  • Prokaryotes often have polycistronic genes, where multiple genes are part of one transcription unit and are regulated together.
  • Eukaryotes typically have monocistronic genes. These genes are often "split."
    • Exons are the coding sequences that appear in the final, mature RNA.
    • Introns are the non-coding, intervening sequences that are removed during RNA processing.

Types of RNA and the process of Transcription

In Bacteria:

  • There are three major types of RNA: mRNA (messenger RNA), tRNA (transfer RNA), and rRNA (ribosomal RNA).
  • A single enzyme, DNA-dependent RNA polymerase, transcribes all three types.
  • The process involves three steps:
    1. Initiation: RNA polymerase binds to the promoter with the help of an initiation-factor (σ\sigma).
    2. Elongation: The polymerase moves along the DNA, synthesizing RNA.
    3. Termination: When the polymerase reaches the terminator region, a termination-factor (ρ\rho) helps release the newly made RNA.
  • In bacteria, transcription and translation can be coupled (occur at the same time) because there is no nucleus and the mRNA does not require processing.

In Eukaryotes: The process is more complex:

  1. There are at least three different RNA polymerases:
    • RNA polymerase I transcribes rRNAs.
    • RNA polymerase II transcribes the precursor of mRNA, called heterogeneous nuclear RNA (hnRNA).
    • RNA polymerase III transcribes tRNA, 5srRNA, and snRNAs.
  2. The primary transcript (hnRNA) is non-functional and must be processed. This involves:
    • Splicing: The introns are removed, and the exons are joined together.
    • Capping: An unusual nucleotide (methyl guanosine triphosphate) is added to the 5'-end.
    • Tailing: A tail of 200-300 adenylate residues (poly-A tail) is added to the 3'-end.
  • After processing, the mature mRNA is transported out of the nucleus to the cytoplasm for translation.

Genetic Code

Translation is the process of converting the genetic information from a polymer of nucleotides (mRNA) into a polymer of amino acids (protein). The set of rules that governs this conversion is the genetic code.

George Gamow proposed that a triplet code (a combination of 3 bases) would be necessary to specify all 20 amino acids (43=644^3 = 64 possible codons). This was later confirmed by the work of Har Gobind Khorana, Marshall Nirenberg, and Severo Ochoa.

Salient features of the genetic code:

  • The codon is triplet. 61 codons code for amino acids, and 3 codons (UAA, UAG, UGA) are stop codons that terminate translation.
  • The code is degenerate. A single amino acid can be coded by more than one codon.
  • The code is read in a contiguous fashion. There are no pauses or punctuations between codons.
  • The code is nearly universal. For example, the codon UUU codes for Phenylalanine (phe) in almost all organisms, from bacteria to humans.
  • AUG has a dual function. It codes for Methionine (met) and also serves as the initiator codon.

Mutations and Genetic Code

The relationship between genes and proteins is clearly seen through mutations.

  • A point mutation, a change in a single base pair, can cause a change in an amino acid. A classic example is sickle cell anemia, where a single base change leads to the substitution of glutamate with valine in the beta-globin chain.
  • Frameshift mutations are caused by the insertion or deletion of one or two bases in the DNA. This shifts the "reading frame" of the codons, altering every amino acid from the point of the mutation onwards.
Example
Consider the sentence: RAM HAS RED CAP
  • If we insert 'B', it becomes: RAM HAS BRE DCA P (The reading frame is shifted).
  • If we insert 'BIG' (three letters), it becomes: RAM HAS BIG RED CAP (The reading frame is restored after the insertion). The insertion or deletion of three (or a multiple of three) bases adds or removes one or more amino acids but leaves the rest of the reading frame intact.

tRNA- the Adapter Molecule

Francis Crick proposed that an adapter molecule must exist that could read the genetic code on the mRNA and also bind to a specific amino acid. This molecule is transfer RNA (tRNA).

  • tRNA has an anticodon loop with a three-base sequence that is complementary to an mRNA codon.
  • It has an amino acid acceptor end where a specific amino acid attaches.
  • There are specific tRNAs for each amino acid.
  • An initiator tRNA recognizes the start codon (AUG). There are no tRNAs for stop codons.
  • The 2D structure of tRNA resembles a clover-leaf, while its actual 3D structure is a compact, inverted 'L' shape (Figure 5.12).

Translation

Translation is the process of protein synthesis, where the sequence of codons on an mRNA molecule is used to create a specific sequence of amino acids in a polypeptide chain.

Key steps and components:

  1. Charging of tRNA (Aminoacylation): This is the first step. An amino acid is activated using energy from ATP and then linked to its corresponding tRNA.
  2. Ribosomes: These are the cellular factories for protein synthesis. A ribosome is made of rRNA and proteins and consists of a large and a small subunit. It provides the site for translation and also acts as a catalyst (the 23S rRNA in bacteria is a ribozyme) for forming the peptide bond that links amino acids together.
  3. mRNA: Provides the template with the codons. A translational unit on an mRNA is the region flanked by a start codon (AUG) and a stop codon. It also contains untranslated regions (UTRs) at the 5' and 3' ends, which are required for efficient translation.

The Process of Translation (Figure 5.13):

  1. Initiation: The small ribosomal subunit binds to the mRNA. The initiator tRNA, carrying methionine, binds to the start codon (AUG). The large subunit then joins to form the complete ribosome.
  2. Elongation: The ribosome moves along the mRNA, one codon at a time. For each codon, the corresponding charged tRNA binds, and the amino acid it carries is added to the growing polypeptide chain via a peptide bond.
  3. Termination: When the ribosome reaches a stop codon (UAA, UAG, or UGA), a release factor binds to it. This terminates translation and releases the completed polypeptide from the ribosome.

Regulation of Gene Expression

Gene expression is the process by which information from a gene is used to synthesize a functional product, like a protein. This process is tightly regulated to ensure that genes are only expressed when and where their products are needed.

In eukaryotes, regulation can occur at multiple levels:

  • Transcriptional level (controlling which genes are transcribed).
  • Processing level (regulating splicing).
  • Transport of mRNA from the nucleus to the cytoplasm.
  • Translational level (controlling which mRNAs are translated).

In prokaryotes, the primary site of control is the initiation of transcription. This is often achieved through operons.

The Lac operon

The lac operon, first described by Francois Jacob and Jacque Monod, is a classic example of gene regulation in E. coli. An operon is a cluster of genes that are transcribed together and regulated by a single promoter and operator.

The lac operon is involved in the metabolism of lactose. It consists of:

  • One regulatory gene (i gene): Codes for a repressor protein.
  • Three structural genes:
    • z gene: Codes for beta-galactosidase (β\beta-gal), which breaks lactose into glucose and galactose.
    • y gene: Codes for permease, which transports lactose into the cell.
    • a gene: Codes for transacetylase.
  • A promoter (p): The binding site for RNA polymerase.
  • An operator (o): The binding site for the repressor protein.

Regulation of the Lac Operon (Figure 5.14): This is an example of negative regulation, where the operon is normally "off" but can be turned "on" by an inducer (lactose).

  • In the absence of lactose (Inducer):
    • The i gene constitutively produces the repressor protein.
    • The repressor binds to the operator region (o).
    • This binding blocks RNA polymerase from accessing the promoter, so the structural genes (z, y, a) are not transcribed. The operon is OFF.
  • In the presence of lactose (Inducer):
    • Lactose (or its isomer, allolactose) enters the cell and binds to the repressor protein.
    • This binding changes the shape of the repressor, causing it to detach from the operator.
    • With the operator now free, RNA polymerase can bind to the promoter and transcribe the structural genes.
    • The enzymes for lactose metabolism are produced. The operon is ON.

When the lactose is used up, the repressor becomes free again, binds back to the operator, and shuts the operon off.

Human Genome Project

The Human Genome Project (HGP) was a massive international research effort launched in 1990 with the goal of sequencing the entire human genome. It was completed in 2003.

Why was it a "Mega Project"?

  • The human genome has approximately 3×1093 \times 10^9 bp.
  • The estimated cost was about 9 billion US dollars.
  • The vast amount of data generated required the development of a new field, Bioinformatics, to store, retrieve, and analyze the data.

Goals of HGP:

  • Identify all approximately 20,000-25,000 genes in human DNA.
  • Determine the sequences of the 3 billion chemical base pairs.
  • Store this information in databases.
  • Improve tools for data analysis.
  • Address the ethical, legal, and social issues (ELSI) arising from the project.

Methodologies:

  • DNA was isolated, broken into smaller fragments, and cloned into vectors like BAC (bacterial artificial chromosomes) and YAC (yeast artificial chromosomes) to be amplified.
  • The fragments were sequenced using automated DNA sequencers based on a method developed by Frederick Sanger.
  • Computer programs were used to align the overlapping fragments to reconstruct the full genome sequence.
  • Two main approaches were used:
    1. Expressed Sequence Tags (ESTs): Identifying and sequencing only the genes that are expressed as RNA.
    2. Sequence Annotation: Sequencing the entire genome (both coding and non-coding regions) and then assigning functions to different regions.

Salient Features of Human Genome

  • The human genome contains 3164.7 million base pairs.
  • The total number of genes is estimated at around 30,000.
  • The average gene consists of 3000 bases. The largest known human gene is dystrophin, at 2.4 million bases.
  • Less than 2% of the genome codes for proteins.
  • A very large portion of the genome is made of repetitive sequences, stretches of DNA that are repeated many times.
  • 99.9% of the nucleotide sequence is exactly the same in all people.
  • Chromosome 1 has the most genes (2968), and the Y chromosome has the fewest (231).
  • About 1.4 million locations of single nucleotide polymorphisms (SNPs), single-base differences between individuals, have been identified.

Applications and Future Challenges

The HGP has revolutionized biological research. It allows scientists to study all genes in a genome at once (genomics), leading to a better understanding of biological systems, human health, and disease. This knowledge promises new ways to diagnose, treat, and prevent thousands of disorders.

DNA Fingerprinting

Since 99.9% of human DNA is identical, DNA fingerprinting focuses on the 0.1% that varies between individuals. It is a quick technique to compare the DNA of two or more people by analyzing specific regions of repetitive DNA.

These repetitive sequences are separated from the bulk DNA during density gradient centrifugation and are called satellite DNA. They show a high degree of polymorphism (variation at the genetic level), which forms the basis of DNA fingerprinting.

The technique was developed by Alec Jeffreys, who used a type of satellite DNA called Variable Number of Tandem Repeats (VNTR). The number of repeats in a VNTR is highly variable among individuals.

The Process of DNA Fingerprinting (Figure 5.16): The original technique involved Southern blot hybridization.

  1. Isolation of DNA from a sample (e.g., blood, saliva, hair).
  2. Digestion of DNA with restriction enzymes.
  3. Separation of DNA fragments by size using gel electrophoresis.
  4. Transferring (blotting) the separated DNA fragments to a synthetic membrane (like nylon).
  5. Hybridization of the membrane with a radiolabeled VNTR probe.
  6. Detection of the hybridized fragments by autoradiography, which reveals a unique pattern of bands for each individual.

The sensitivity of this technique has been greatly increased by the use of Polymerase Chain Reaction (PCR), which can amplify tiny amounts of DNA.

Applications:

  • Forensic science: To identify criminals by matching DNA from a crime scene with suspects.
  • Paternity testing: To resolve parentage disputes, as a child inherits their polymorphisms from their parents.
  • Evolutionary biology: To study genetic diversity in populations.

Way to go! You've finished this chapter 🎉

That's real dedication — you read through every section. Keep up this momentum, revisit anything that felt tricky, and you'll be exam-ready in no time. Explore more from this chapter below.