In summary: (i) Trinity obtained the highest number of genes associated to all three categories, (ii) SPAdes obtained the overall highest representation of GO-terms in all three categories, and (iii) SOAPdenovo-Trans and SPAdes obtained the highest number of uniquely represented GO-terms in all categories (Table 3). and 27.88% were uniquely assembled by Trinity, while 27.65% were uniquely assembled by SPAdes. The non-redundant merging of all three assemblies output permitted the annotation of 9232 sequences, which was 23% more when compared to each software and 28% more when compared to the previous annotation; moreover, the description of 65 novel theraphotoxins was possible. In the generation of data for non-model organisms, as well as in the search for novel peptides with biotechnological interest, it is highly recommended to employ at least two different transcriptome assemblers. are around JC-1 180 times more potent than ziconotide [11,12,13]. Moreover, DBPs have shown potential in the treatment of other pathologies, such as in models of ischemic stroke as shown by the peptide PcTx1 (also shows activity against Gram-positive and Gram-negative bacteria as well as fungi, yeast, and parasites of the genus or displays activity against methicillin-resistant (MRSA), and the peptide CIT1a from displays activity against the yeast [16]. Furthermore, the peptide lycosin-I (from and and is also able to induce the apoptosis of prostate cancer cells, while the peptide latarcin-3a from is able to generate pores in the envelope of the HIV virus [12]. This broad range of activities shows the enormous biotechnological potential of spider venoms; yet, less than 1% of the hypothetical potential novel molecules have been reported in the literature [17]. Proteomic studies from spider venoms have been limited in the amount of data that they can provide, since the collection of the sample is technically complex [18]. Therefore, transcriptomics appears as a preferable tool for the prospection of bioactive molecules from arachnids, as the collection of the sample is greatly simplified [19,20]. Next-generation-sequencing (NGS) tools allow the collection of vast amounts of data from tissues such as the venom gland. Identification of sequences from NGS data is usually carried by pipelines that rely on alignment methodologies, e.g., BLAST [19,21,22,23]. However, alignment methodologies specialize in the detection of homology between closely related sequences. Thus, for non-model organisms, methodologies for the Rabbit Polyclonal to PSMD2 detection of distantly related sequences, which rely instead on the identification of conserved motifs by statistical models, e.g., hidden Markov models (HMMs) [24,25], must also be implemented, as these could show higher specificity and sensibility. For organisms such as those found in the order, due to the lack of reported data at the genome and transcriptome level, the assembly and the annotation of genomes or transcriptomes become important endeavors in the characterization process; thus, the retrieval of the nucleotide information is paramount. Venoms are highly complex substances: a high splicing activity [26,27], as well as the presence of several duplicated genes within the venom gland [28], accounts for the presence of highly homologous JC-1 and paralogous sequences that can distort the quantity and quality of the information recovered. Therefore, the use of various assemblers may also be recommended since taxa-specific as well as tissue-specific biases and issues may arise [29,30]. The main objective of this study is to improve the available transcriptomic resources of the spider (order, we have improved the amount of information obtained from venom gland transcriptome by re-assembling the reads with three different free de novo assembly algorithms. The new results allowed us to annotate and report 9232 genes and unveiled new bioactive peptides with potential interest in biomedicine. 2. Results 2.1. Quality of the Transcripts Assembled with Trinity and SPAdes Are Similar, While Outperforming That of SOAPdenovo-Trans The 24 million 101 bp pair-ended raw reads of the venom gland transcriptome of were cleaned from low-quality and Illumina adaptor sequences with the software TrimGalore (v0.6.3 available at https://github.com/FelixKrueger/TrimGalore, accessed on 31 July 2020) and assembled using three free access software: Trinity (v2.1.1 available at https://github.com/trinityrnaseq/trinityrnaseq, accessed on 31 July 2020), SPAdes (v3.13.1 available at https://github.com/ablab/spades, accessed on 31 July 2020), and SOAPdenovo-Trans (v1.0.4 available at https://github.com/aquaskyline/SOAPdenovo-Trans, accessed on 31 July 2020). On average, 36% and 85% of the obtained contigs from SPAdes and SOAPdenovo-Trans had less than 200 bp, respectively (Table 1). Removal of these sequences showed that Trinity assembled 21% more contigs than SPAdes and 52% more contigs than SOAPdenovo-Trans on average; however, the N50 JC-1 metric was higher on average for those contigs assembled with SPAdes (802 bp), followed then by Trinity (784 bp), and finally by SOAPdenovo-Trans (486 bp); the GC content statistic showed no major differences (Table 1). Raw read representation on the assembled contigs was also calculated with the software Bowtie2 (v.2.2.5 available at https://github.com/BenLangmead/bowtie2, accessed JC-1 31 July 2020). All assemblies obtained.