The RNA Bricks database (http://iimcb.genesilico.pl/rnabricks), stores information about recurrent RNA 3D motifs and their interactions, found in experimentally determined RNA structures and in RNA–protein complexes. In contrast to other similar tools (RNA 3D Motif Atlas, RNA Frabase, Rloom) RNA motifs, i.e. ‘RNA bricks’ are presented in the molecular environment, in which they were determined, including RNA, protein, metal ions, water molecules and ligands. All nucleotide residues in RNA bricks are annotated with structural quality scores that describe real-space correlation coefficients with the electron density data (if available), backbone geometry and possible steric conflicts, which can be used to identify poorly modeled residues. The database is also equipped with an algorithm for 3D motif search and comparison. The algorithm compares spatial positions of backbone atoms of the user-provided query structure and of stored RNA motifs, without relying on sequence or secondary structure information. This enables the identification of local structural similarities among evolutionarily related and unrelated RNA molecules. Besides, the search utility enables searching ‘RNA bricks’ according to sequence similarity, and makes it possible to identify motifs with modified ribonucleotide residues at specific positions.
Creating useful software is a major activity of many scientists, including bioinformaticians. Nevertheless, software development in an academic setting is often unsystematic, which can lead to problems associated with maintenance and long-term availibility. Unfortunately, well-documented software development methodology is difficult to adopt, and technical measures that directly improve bioinformatic programming have not been described comprehensively. We have examined 22 software projects and have identified a set of practices for software development in an academic environment. We found them useful to plan a project, support the involvement of experts (e.g. experimentalists), and to promote higher quality and maintainability of the resulting programs. This article describes 12 techniques that facilitate a quick start into software engineering. We describe 3 of the 22 projects in detail and give many examples to illustrate the usage of particular techniques. We expect this toolbox to be useful for many bioinformatics programming projects and to the training of scientific programmers.
software development; programming; project management; software quality
R.MwoI is a Type II restriction endonucleases enzyme (REase), which specifically recognizes a palindromic interrupted DNA sequence 5′-GCNNNNNNNGC-3′ (where N indicates any nucleotide), and hydrolyzes the phosphodiester bond in the DNA between the 7th and 8th base in both strands. R.MwoI exhibits remote sequence similarity to R.BglI, a REase with known structure, which recognizes an interrupted palindromic target 5′-GCCNNNNNGGC-3′. A homology model of R.MwoI in complex with DNA was constructed and used to predict functionally important amino acid residues that were subsequently targeted by mutagenesis. The model, together with the supporting experimental data, revealed regions important for recognition of the common bases in DNA sequences recognized by R.BglI and R.MwoI. Based on the bioinformatics analysis, we designed substitutions of the S310 residue in R.MwoI to arginine or glutamic acid, which led to enzyme variants with altered sequence selectivity compared with the wild-type enzyme. The S310R variant of R.MwoI preferred the 5′-GCCNNNNNGGC-3′ sequence as a target, similarly to R.BglI, whereas the S310E variant preferentially cleaved a subset of the MwoI sites, depending on the identity of the 3rd and 9th nucleotide residues. Our results represent a case study of a REase sequence specificity alteration by a single amino acid substitution, based on a theoretical model in the absence of a crystal structure.
DNA is continuously exposed to many different damaging agents such as environmental chemicals, UV light, ionizing radiation, and reactive cellular metabolites. DNA lesions can result in different phenotypical consequences ranging from a number of diseases, including cancer, to cellular malfunction, cell death, or aging. To counteract the deleterious effects of DNA damage, cells have developed various repair systems, including biochemical pathways responsible for the removal of single-strand lesions such as base excision repair (BER) and nucleotide excision repair (NER) or specialized polymerases temporarily taking over lesion-arrested DNA polymerases during the S phase in translesion synthesis (TLS). There are also other mechanisms of DNA repair such as homologous recombination repair (HRR), nonhomologous end-joining repair (NHEJ), or DNA damage response system (DDR). This paper reviews bioinformatics resources specialized in disseminating information about DNA repair pathways, proteins involved in repair mechanisms, damaging agents, and DNA lesions.
NpmA, a methyltransferase that confers resistance to aminoglycosides was identified in an Escherichia coli clinical isolate. It belongs to the kanamycin–apramycin methyltransferase (Kam) family and specifically methylates the 16S rRNA at the N1 position of A1408. We determined the structures of apo-NpmA and its complexes with S-adenosylmethionine (AdoMet) and S-adenosylhomocysteine (AdoHcy) at 2.4, 2.7 and 1.68 Å, respectively. We generated a number of NpmA variants with alanine substitutions and studied their ability to bind the cofactor, to methylate A1408 in the 30S subunit, and to confer resistance to kanamycin in vivo. Residues D30, W107 and W197 were found to be essential. We have also analyzed the interactions between NpmA and the 30S subunit by footprinting experiments and computational docking. Helices 24, 42 and 44 were found to be the main NpmA-binding site. Both experimental and theoretical analyses suggest that NpmA flips out the target nucleotide A1408 to carry out the methylation. NpmA is plasmid-encoded and can be transferred between pathogenic bacteria; therefore it poses a threat to the successful use of aminoglycosides in clinical practice. The results presented here will assist in the development of specific NpmA inhibitors that could restore the potential of aminoglycoside antibiotics.
Sgm (Sisomicin-gentamicin methyltransferase) from antibiotic-producing bacterium Micromonospora zionensis is an enzyme that confers resistance to aminoglycosides like gentamicin and sisomicin by specifically methylating G1405 in bacterial 16S rRNA. Sgm belongs to the aminoglycoside resistance methyltransferase (Arm) family of enzymes that have been recently found to spread by horizontal gene transfer among disease-causing bacteria. Structural characterization of Arm enzymes is the key to understand their mechanism of action and to develop inhibitors that would block their activity. Here we report the structure of Sgm in complex with cofactors S-adenosylmethionine (AdoMet) and S-adenosylhomocysteine (AdoHcy) at 2.0 and 2.1 Å resolution, respectively, and results of mutagenesis and rRNA footprinting, and protein-substrate docking. We propose the mechanism of methylation of G1405 by Sgm and compare it with other m7G methyltransferases, revealing a surprising diversity of active sites and binding modes for the same basic reaction of RNA modification. This analysis can serve as a stepping stone towards developing drugs that would specifically block the activity of Arm methyltransferases and thereby re-sensitize pathogenic bacteria to aminoglycoside antibiotics.
Bacterial Dsb enzymes are involved in the oxidative folding of many proteins, through the formation of disulfide bonds between their cysteine residues. The Dsb protein network has been well characterized in cells of the model microorganism Escherichia coli. To gain insight into the functioning of the Dsb system in epsilon-Proteobacteria, where it plays an important role in the colonization process, we studied two homologs of the main Escherichia coli Dsb oxidase (EcDsbA) that are present in the cells of the enteric pathogen Campylobacter jejuni, the most frequently reported bacterial cause of human enteritis in the world.
Methods and Results
Phylogenetic analysis suggests the horizontal transfer of the epsilon-Proteobacterial DsbAs from a common ancestor to gamma-Proteobacteria, which then gave rise to the DsbL lineage. Phenotype and enzymatic assays suggest that the two C. jejuni DsbAs play different roles in bacterial cells and have divergent substrate spectra. CjDsbA1 is essential for the motility and autoagglutination phenotypes, while CjDsbA2 has no impact on those processes. CjDsbA1 plays a critical role in the oxidative folding that ensures the activity of alkaline phosphatase CjPhoX, whereas CjDsbA2 is crucial for the activity of arylsulfotransferase CjAstA, encoded within the dsbA2-dsbB-astA operon.
Our results show that CjDsbA1 is the primary thiol-oxidoreductase affecting life processes associated with bacterial spread and host colonization, as well as ensuring the oxidative folding of particular protein substrates. In contrast, CjDsbA2 activity does not affect the same processes and so far its oxidative folding activity has been demonstrated for one substrate, arylsulfotransferase CjAstA. The results suggest the cooperation between CjDsbA2 and CjDsbB. In the case of the CjDsbA1, this cooperation is not exclusive and there is probably another protein to be identified in C. jejuni cells that acts to re-oxidize CjDsbA1. Altogether the data presented here constitute the considerable insight to the Epsilonproteobacterial Dsb systems, which have been poorly understood so far.
MODOMICS, a database devoted to the systems biology of RNA modification, has been subjected to substantial improvements. It provides comprehensive information on the chemical structure of modified nucleosides, pathways of their biosynthesis, sequences of RNAs containing these modifications and RNA-modifying enzymes. MODOMICS also provides cross-references to other databases and to literature. In addition to the previously available manually curated tRNA sequences from a few model organisms, we have now included additional tRNAs and rRNAs, and all RNAs with 3D structures in the Nucleic Acid Database, in which modified nucleosides are present. In total, 3460 modified bases in RNA sequences of different organisms have been annotated. New RNA-modifying enzymes have been also added. The current collection of enzymes includes mainly proteins for the model organisms Escherichia coli and Saccharomyces cerevisiae, and is currently being expanded to include proteins from other organisms, in particular Archaea and Homo sapiens. For enzymes with known structures, links are provided to the corresponding Protein Data Bank entries, while for many others homology models have been created. Many new options for database searching and querying have been included. MODOMICS can be accessed at http://genesilico.pl/modomics.
The understanding of folding and function of RNA molecules depends on the identification and classification of interactions between ribonucleotide residues. We developed a new method named ClaRNA for computational classification of contacts in RNA 3D structures. Unique features of the program are the ability to identify imperfect contacts and to process coarse-grained models. Each doublet of spatially close ribonucleotide residues in a query structure is compared to clusters of reference doublets obtained by analysis of a large number of experimentally determined RNA structures, and assigned a score that describes its similarity to one or more known types of contacts, including pairing, stacking, base–phosphate and base–ribose interactions. The accuracy of ClaRNA is 0.997 for canonical base pairs, 0.983 for non-canonical pairs and 0.961 for stacking interactions. The generalized squared correlation coefficient (GC2) for ClaRNA is 0.969 for canonical base pairs, 0.638 for non-canonical pairs and 0.824 for stacking interactions. The classifier can be easily extended to include new types of spatial relationships between pairs or larger assemblies of nucleotide residues. ClaRNA is freely available via a web server that includes an extensive set of tools for processing and visualizing structural information about RNA molecules.
In this study, we present the discovery and characterization of a highly thermostable endolysin from bacteriophage Ph2119 infecting Thermus strain MAT2119 isolated from geothermal areas in Iceland. Nucleotide sequence analysis of the 16S rRNA gene affiliated the strain with the species Thermus scotoductus. Bioinformatics analysis has allowed identification in the genome of phage 2119 of an open reading frame (468 bp in length) coding for a 155-amino-acid basic protein with an Mr of 17,555. Ph2119 endolysin does not resemble any known thermophilic phage lytic enzymes. Instead, it has conserved amino acid residues (His30, Tyr58, His132, and Cys140) that form a Zn2+ binding site characteristic of T3 and T7 lysozymes, as well as eukaryotic peptidoglycan recognition proteins, which directly bind to, but also may destroy, bacterial peptidoglycan. The purified enzyme shows high lytic activity toward thermophiles, i.e., T. scotoductus (100%), Thermus thermophilus (100%), and Thermus flavus (99%), and also, to a lesser extent, toward mesophilic Gram-negative bacteria, i.e., Escherichia coli (34%), Serratia marcescens (28%), Pseudomonas fluorescens (13%), and Salmonella enterica serovar Panama (10%). The enzyme has shown no activity against a number of Gram-positive bacteria analyzed, with the exception of Deinococcus radiodurans (25%) and Bacillus cereus (15%). Ph2119 endolysin was found to be highly thermostable: it retains approximately 87% of its lytic activity after 6 h of incubation at 95°C. The optimum temperature range for the enzyme activity is 50°C to 78°C. The enzyme exhibits lytic activity in the pH range of 6 to 10 (maximum at pH 7.5 to 8.0) and is also active in the presence of up to 500 mM NaCl.
The structure of Bacillus subtilis TrmB (BsTrmB), the tRNA (m7G46) methyltransferase, was determined at a resolution of 2.1 Å. This is the first structure of a member of the TrmB family to be determined by X-ray crystallography. It reveals a unique variant of the Rossmann-fold methyltransferase (RFM) structure, with the N-terminal helix folded on the opposite site of the catalytic domain. The architecture of the active site and a computational docking model of BsTrmB in complex with the methyl group donor S-adenosyl-l-methionine and the tRNA substrate provide an explanation for results from mutagenesis studies of an orthologous enzyme from Escherichia coli (EcTrmB). However, unlike EcTrmB, BsTrmB is shown here to be dimeric both in the crystal and in solution. The dimer interface has a hydrophobic core and buries a potassium ion and five water molecules. The evolutionary analysis of the putative interface residues in the TrmB family suggests that homodimerization may be a specific feature of TrmBs from Bacilli, which may represent an early stage of evolution to an obligatory dimer.
R.DpnI consists of N-terminal catalytic and C-terminal winged helix domains that are separately specific for the Gm6ATC sequences in Dam-methylated DNA. Here we present a crystal structure of R.DpnI with oligoduplexes bound to the catalytic and winged helix domains and identify the catalytic domain residues that are involved in interactions with the substrate methyl groups. We show that these methyl groups in the Gm6ATC target sequence are positioned very close to each other. We further show that the presence of the two methyl groups requires a deviation from B-DNA conformation to avoid steric conflict. The methylation compatible DNA conformation is complementary with binding sites of both R.DpnI domains. This indirect readout of methylation adds to the specificity mediated by direct favorable interactions with the methyl groups and solvation/desolvation effects. We also present hydrogen/deuterium exchange data that support ‘crosstalk’ between the two domains in the identification of methylated DNA, which should further enhance R.DpnI methylation specificity.
•AZIN2, unlike ornithine decarboxylase, exists as a monomer in solution.•Conserved residues among AZIN2 orthologs are critical for the binding to antizymes (AZs).•Substitution of the conserved residues affects the ability of AZIN2 to modulate polyamine levels.•AZIN2 and AZs are extremely labile proteins, which mutually stabilize each other.•Other proteolytic systems, besides the 26S proteasome, might be involved in AZIN2 degradation.
Ornithine decarboxylase (ODC) is the key enzyme in the polyamine biosynthetic pathway. ODC levels are controlled by polyamines through the induction of antizymes (AZs), small proteins that inhibit ODC and target it to proteasomal degradation without ubiquitination. Antizyme inhibitors (AZIN1 and AZIN2) are proteins homologous to ODC that bind to AZs and counteract their negative effect on ODC. Whereas ODC and AZIN1 are well-characterized proteins, little is known on the structure and stability of AZIN2, the lastly discovered member of this regulatory circuit. In this work we first analyzed structural aspects of AZIN2 by combining biochemical and computational approaches. We demonstrated that AZIN2, in contrast to ODC, does not form homodimers, although the predicted tertiary structure of the AZIN2 monomer was similar to that of ODC. Furthermore, we identified conserved residues in the antizyme-binding element, whose substitution drastically affected the capacity of AZIN2 to bind AZ1. On the other hand, we also found that AZIN2 is much more labile than ODC, but it is highly stabilized by its binding to AZs. Interestingly, the administration of the proteasome inhibitor MG132 caused differential effects on the three AZ-binding proteins, having no effect on ODC, preventing the degradation of AZIN1, but unexpectedly increasing the degradation of AZIN2. Inhibitors of the lysosomal function partially prevented the effect of MG132 on AZIN2. These results suggest that the degradation of AZIN2 could be also mediated by an alternative route to that of proteasome. These findings provide new relevant information on this unique regulatory mechanism of polyamine metabolism.
AZ, antizyme; AZBE, antizyme-binding element; AZIN, antizyme inhibitor; ERGIC, endoplasmic reticulum-Golgi intermediate compartment; ODC, ornithine decarboxylase; GDT_TS, global distance test total score; HA, hemagglutinin; HEK, human embryonic kidney; PAGE, polyacrylamide gel electrophoresis; RMSD, root-mean-square deviation; TGN, trans-Golgi network; Antizyme; Antizyme-binding element; Homology modeling; Polyamines; Protein degradation; Proteasome inhibitors
Ribonuclease H-like (RNHL) superfamily, also called the retroviral integrase superfamily, groups together numerous enzymes involved in nucleic acid metabolism and implicated in many biological processes, including replication, homologous recombination, DNA repair, transposition and RNA interference. The RNHL superfamily proteins show extensive divergence of sequences and structures. We conducted database searches to identify members of the RNHL superfamily (including those previously unknown), yielding >60 000 unique domain sequences. Our analysis led to the identification of new RNHL superfamily members, such as RRXRR (PF14239), DUF460 (PF04312, COG2433), DUF3010 (PF11215), DUF429 (PF04250 and COG2410, COG4328, COG4923), DUF1092 (PF06485), COG5558, OrfB_IS605 (PF01385, COG0675) and Peptidase_A17 (PF05380). Based on the clustering analysis we grouped all identified RNHL domain sequences into 152 families. Phylogenetic studies revealed relationships between these families, and suggested a possible history of the evolution of RNHL fold and its active site. Our results revealed clear division of the RNHL superfamily into exonucleases and endonucleases. Structural analyses of features characteristic for particular groups revealed a correlation between the orientation of the C-terminal helix with the exonuclease/endonuclease function and the architecture of the active site. Our analysis provides a comprehensive picture of sequence-structure-function relationships in the RNHL superfamily that may guide functional studies of the previously uncharacterized protein families.
Risk alleles for complex diseases are widely spread throughout human populations. However, little is known about the geographic distribution and frequencies of risk alleles, which may contribute to differences in disease susceptibility and prevalence among populations. Here, we focus on Crohn's disease (CD) as a model for the evolutionary study of complex disease alleles. Recent genome-wide association studies and classical linkage analyses have identified more than 70 susceptible genomic regions for CD in Europeans, but only a few have been confirmed in non-European populations. Our analysis of eight European-specific susceptibility genes using HapMap data shows that at the NOD2 locus the CD-risk alleles are linked with a haplotype specific to CEU at a frequency that is significantly higher compared with the entire genome. We subsequently examined nine global populations and found that the CD-risk alleles spread through hitchhiking with a high-frequency haplotype (H1) exclusive to Europeans. To examine the neutrality of NOD2, we performed phylogenetic network analyses, coalescent simulation, protein structural prediction, characterization of mutation patterns, and estimations of population growth and time to most recent common ancestor (TMRCA). We found that while H1 was significantly prevalent in European populations, the H1 TMRCA predated human migration out of Africa. H1 is likely to have undergone negative selection because 1) the root of H1 genealogy is defined by a preexisting amino acid substitution that causes serious conformational changes to the NOD2 protein, 2) the haplotype has almost become extinct in Africa, and 3) the haplotype has not been affected by the recent European expansion reflected in the other haplotypes. Nevertheless, H1 has survived in European populations, suggesting that the haplotype is advantageous to this group. We propose that several CD-risk alleles, which destabilize and disrupt the NOD2 protein, have been maintained by natural selection on standing variation because the deleterious haplotype of NOD2 is advantageous in diploid individuals due to heterozygote advantage and/or intergenic interactions.
Crohn's disease; NOD2; hitchhiking effect; natural selection; standing variation; mildly deleterious mutation
QA-RecombineIt provides a web interface to assess the quality of protein 3D structure models and to improve the accuracy of models by merging fragments of multiple input models. QA-RecombineIt has been developed for protein modelers who are working on difficult problems, have a set of different homology models and/or de novo models (from methods such as I-TASSER or ROSETTA) and would like to obtain one consensus model that incorporates the best parts into one structure that is internally coherent. An advanced mode is also available, in which one can modify the operation of the fragment recombination algorithm by manually identifying individual fragments or entire models to recombine. Our method produces up to 100 models that are expected to be on the average more accurate than the starting models. Therefore, our server may be useful for crystallographic protein structure determination, where protein models are used for Molecular Replacement to solve the phase problem. To address the latter possibility, a special feature was added to the QA-RecombineIt server. The QA-RecombineIt server can be freely accessed at http://iimcb.genesilico.pl/qarecombineit/.
The continuously increasing amount of RNA sequence and experimentally determined 3D structure data drives the development of computational methods supporting exploration of these data. Contemporary functional analysis of RNA molecules, such as ribozymes or riboswitches, covers various issues, among which tertiary structure modeling becomes more and more important. A growing number of tools to model and predict RNA structure calls for an evaluation of these tools and the quality of outcomes their produce. Thus, the development of reliable methods designed to meet this need is relevant in the context of RNA tertiary structure analysis and can highly influence the quality and usefulness of RNA tertiary structure prediction in the nearest future. Here, we present RNAlyzer—a computational method for comparison of RNA 3D models with the reference structure and for discrimination between the correct and incorrect models. Our approach is based on the idea of local neighborhood, defined as a set of atoms included in the sphere centered around a user-defined atom. A unique feature of the RNAlyzer is the simultaneous visualization of the model-reference structure distance at different levels of detail, from the individual residues to the entire molecules.
We present a continuous benchmarking approach for the assessment of RNA secondary structure prediction methods implemented in the CompaRNA web server. As of 3 October 2012, the performance of 28 single-sequence and 13 comparative methods has been evaluated on RNA sequences/structures released weekly by the Protein Data Bank. We also provide a static benchmark generated on RNA 2D structures derived from the RNAstrand database. Benchmarks on both data sets offer insight into the relative performance of RNA secondary structure prediction methods on RNAs of different size and with respect to different types of structure. According to our tests, on the average, the most accurate predictions obtained by a comparative approach are generated by CentroidAlifold, MXScarna, RNAalifold and TurboFold. On the average, the most accurate predictions obtained by single-sequence analyses are generated by CentroidFold, ContextFold and IPknot. The best comparative methods typically outperform the best single-sequence methods if an alignment of homologous RNA sequences is available. This article presents the results of our benchmarks as of 3 October 2012, whereas the rankings presented online are continuously updated. We will gladly include new prediction methods and new measures of accuracy in the new editions of CompaRNA benchmarks.
A key step in proliferation of retroviruses is the conversion of their RNA genome to double-stranded DNA, a process catalysed by multifunctional reverse transcriptases (RTs). Dimeric and monomeric RTs have been described, the latter exemplified by the enzyme of Moloney murine leukaemia virus. However, structural information is lacking that describes the substrate binding mechanism for a monomeric RT. We report here the first crystal structure of a complex between an RNA/DNA hybrid substrate and polymerase-connection fragment of the single-subunit RT from xenotropic murine leukaemia virus-related virus, a close relative of Moloney murine leukaemia virus. A comparison with p66/p51 human immunodeficiency virus-1 RT shows that substrate binding around the polymerase active site is conserved but differs in the thumb and connection subdomains. Small-angle X-ray scattering was used to model full-length xenotropic murine leukaemia virus-related virus RT, demonstrating that its mobile RNase H domain becomes ordered in the presence of a substrate—a key difference between monomeric and dimeric RTs.
MODOMICS is a database of RNA modifications that provides comprehensive information concerning the chemical structures of modified ribonucleosides, their biosynthetic pathways, RNA-modifying enzymes and location of modified residues in RNA sequences. In the current database version, accessible at http://modomics.genesilico.pl, we included new features: a census of human and yeast snoRNAs involved in RNA-guided RNA modification, a new section covering the 5′-end capping process, and a catalogue of ‘building blocks’ for chemical synthesis of a large variety of modified nucleosides. The MODOMICS collections of RNA modifications, RNA-modifying enzymes and modified RNAs have been also updated. A number of newly identified modified ribonucleosides and more than one hundred functionally and structurally characterized proteins from various organisms have been added. In the RNA sequences section, snRNAs and snoRNAs with experimentally mapped modified nucleosides have been added and the current collection of rRNA and tRNA sequences has been substantially enlarged. To facilitate literature searches, each record in MODOMICS has been cross-referenced to other databases and to selected key publications. New options for database searching and querying have been implemented, including a BLAST search of protein sequences and a PARALIGN search of the collected nucleic acid sequences.
Ribonucleases (RNases) are valuable tools applied in the analysis of RNA sequence, structure and function. Their substrate specificity is limited to recognition of single bases or distinct secondary structures in the substrate. Currently, there are no RNases available for purely sequence-dependent fragmentation of RNA. Here, we report the development of a new enzyme that cleaves the RNA strand in DNA–RNA hybrids 5 nt from a nonanucleotide recognition sequence. The enzyme was constructed by fusing two functionally independent domains, a RNase HI, that hydrolyzes RNA in DNA–RNA hybrids in processive and sequence-independent manner, and a zinc finger that recognizes a sequence in DNA–RNA hybrids. The optimization of the fusion enzyme’s specificity was guided by a structural model of the protein-substrate complex and involved a number of steps, including site-directed mutagenesis of the RNase moiety and optimization of the interdomain linker length. Methods for engineering zinc finger domains with new sequence specificities are readily available, making it feasible to acquire a library of RNases that recognize and cleave a variety of sequences, much like the commercially available assortment of restriction enzymes. Potentially, zinc finger-RNase HI fusions may, in addition to in vitro applications, be used in vivo for targeted RNA degradation.
The spliceosome is a molecular machine that performs the excision of introns from eukaryotic pre-mRNAs. This macromolecular complex comprises in human cells five RNAs and over one hundred proteins. In recent years, many spliceosomal proteins have been found to exhibit intrinsic disorder, that is to lack stable native three-dimensional structure in solution. Building on the previous body of proteomic, structural and functional data, we have carried out a systematic bioinformatics analysis of intrinsic disorder in the proteome of the human spliceosome. We discovered that almost a half of the combined sequence of proteins abundant in the spliceosome is predicted to be intrinsically disordered, at least when the individual proteins are considered in isolation. The distribution of intrinsic order and disorder throughout the spliceosome is uneven, and is related to the various functions performed by the intrinsic disorder of the spliceosomal proteins in the complex. In particular, proteins involved in the secondary functions of the spliceosome, such as mRNA recognition, intron/exon definition and spliceosomal assembly and dynamics, are more disordered than proteins directly involved in assisting splicing catalysis. Conserved disordered regions in spliceosomal proteins are evolutionarily younger and less widespread than ordered domains of essential spliceosomal proteins at the core of the spliceosome, suggesting that disordered regions were added to a preexistent ordered functional core. Finally, the spliceosomal proteome contains a much higher amount of intrinsic disorder predicted to lack secondary structure than the proteome of the ribosome, another large RNP machine. This result agrees with the currently recognized different functions of proteins in these two complexes.
In eukaryotic cells, introns are spliced out of proteincoding mRNAs by a highly dynamic and extraordinarily plastic molecular machine called the spliceosome. In recent years, multiple regions of intrinsic structural disorder were found in spliceosomal proteins. Intrinsically disordered regions lack stable native three-dimensional structure in solutions, which makes them structurally flexible and/or able to switch between different conformations. Hence, intrinsically disordered regions are the ideal candidate responsible for the spliceosome's plasticity. Intrinsically disordered regions are also frequently the sites of post-translational modifications, which were also proven to be important in spliceosome dynamics. In this article, we describe the results of a structural bioinformatics analysis focused on intrinsic disorder in the spliceosomal proteome. We systematically analyzed all known human spliceosomal proteins with regards to the presence and type of intrinsic disorder. Almost a half of the combined sequence of these spliceosomal proteins is predicted to be intrinsically disordered, and the type of intrinsic disorder in a protein varies with its function and its location in the spliceosome. The parts of the spliceosome that act earlier in the process are more disordered, which corresponds to their role in establishing a network of interactions, while the parts that act later are more ordered.
Exonuclease VII (ExoVII) is a bacterial nuclease involved in DNA repair and recombination that hydrolyses single-stranded DNA. ExoVII is composed of two subunits: large XseA and small XseB. Thus far, little was known about the molecular structure of ExoVII, the interactions between XseA and XseB, the architecture of the nuclease active site or its mechanism of action. We used bioinformatics methods to predict the structure of XseA, which revealed four domains: an N-terminal OB-fold domain, a middle putatively catalytic domain, a coiled-coil domain and a short C-terminal segment. By series of deletion and site-directed mutagenesis experiments on XseA from Escherichia coli, we determined that the OB-fold domain is responsible for DNA binding, the coiled-coil domain is involved in binding multiple copies of the XseB subunit and residues D155, R205, H238 and D241 of the middle domain are important for the catalytic activity but not for DNA binding. Altogether, we propose a model of sequence–structure–function relationships in ExoVII.