The Role of GenSAS, BLAST, and UniProt in Bioinformatics for Gene Research

Main Article Content

Rumo Bian

Keywords

bioinformatics, gene function annotation, GenSAS, BLAST, bioinformatics proteomics

Abstract

The rapid expansion of genomic and proteomic data has intensified the need for robust bioinformatics tools capable of transforming raw sequences into meaningful biological insights. This review highlights the roles of three foundational platforms—GenSAS, BLAST, and UniProt—in gene research workflows. GenSAS offers comprehensive genome annotation capabilities, particularly for non -model organisms; BLAST enables efficient sequence alignment and homology detection; and UniProt serves as a rich repository of cu rated protein information. This research analyzes their individual functions, integration potential, and applications across structural annotation, functional prediction, and comparative genomics. Case studies demonstrate their use in diverse research cont exts, including plant genome annotation, evolutionary analysis, and protein function characterization. We also discuss future directions, emphasizing the importance of machine learning integration, multi -omics interoperability, and enhanced user accessibil ity. Together, these tools provide a synergistic framework that continues to drive advances in gene discovery and bioinformatics.

Abstract 36 | PDF Downloads 22

References

  • [1] Humann JL, Lee T, Ficklin S, Main D. Structural and Functional Annotation of Eukaryotic Genomes with GenSAS. Methods Mol Biol. 2019; 1962:29-51. doi: 10.1007/978-1-4939-9173-0_3.
  • [2] Ladunga, I. Finding Homologs in Amino Acid Sequences Using Network BLAST Searches. Curr Protoc Bioinformatics. 2017; 59:3.4.1-3.4.24. doi: 10.1002/cpbi.34.
  • [3] Singh H, Raghava GP. BLAST -based structural annotation of protein residues using the Protein Data Bank. Biol Direct. 2016;11(1):4. doi: 10.1186/s13062-016-0106-9.
  • [4] Bateman et al. (2023). UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Research, 51(D1), D523-D531.
  • [5] Altschul SF, Gish W, Miller W, Myers EW, Lipman DJ. Basic local alignment search tool. J Mol Biol. 1990;215(3):403-10.
  • [6] Camacho C, Coulouris G, Avagyan V, Ma N, Papadopoulos J, Bealer K, Madden TL. BLAST+: architecture and applications. BMC Bioinformatics. 2009; 10:421.
  • [7] Boratyn GM, Schä ffer AA, Agarwala R, Altschul SF, Lipman DJ, Madden TL. Domain-enhanced lookup time accelerated BLAST. Biol Direct. 2012; 7:12.
  • [8] Madden TL. The BLAST Sequence Analysis Tool. In: McEntyre J, Ostell J, editors. The NCBI Handbook [Internet]. 2nd edition. Bethesda (MD): National Center for Biotechnology Information (US); 2013.
  • [9] Buchfink B, Xie C, Huson DH. Fast and sensitive protein alignment using DIAMOND. Nat Methods. 2015;12(1):59-60. doi: 10.1038/nmeth.3176.
  • [10] Geer LY, Marchler -Bauer A, Geer RC, Han L, He J, He S, Liu C, Shi W, Bryant SH. The NCBI BioSystems database. Nucleic Acids Res. 2010;38(Database issue): D492-6. doi: 10.1093/nar/gkp858.
  • [11] Ashburner M, Ball CA, Blake JA, et al. Gene ontology: a tool for the unification of biology. The Gene Ontology Consortium. Nat Genet. 2000;25(1):25-29.
  • [12] Finn RD, Clements J, Eddy SR. HMMER web server: interactive sequence similarity searching. Nucleic Acids Res. 2011;39(Web Server issue): W29-W37.
  • [13] Eddy SR. Profile hidden Markov models. Bioinformatics. 1998;14(9):755-763.
  • [14] Punta M, Coggill PC, Eberhardt RY, et al. The Pfam protein families database. Nucleic Acids Res. 2012;40(Database issue): D290-D301.
  • [15] Jones DT. Protein secondary structure prediction based on position-specific scoring matrices. J Mol Biol. 1999;292(2):195-202.
  • [16] Hunter S, Jones P, Mitchell A, et al. InterPro in 2011: new developments in the family and domain prediction database. Nucleic Acids Res. 2012;40 (Database issue): D306-D312.
  • [17] UniProt Consortium. UniProt: a worldwide hub of protein knowledge. Nucleic Acids Res. 2019;47(D1): D506-D515.
  • [18] Steinegger M, Sö ding J. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat Biotechnol. 2017;35(11):1026-1028.
  • [19] Li W, Jaroszewski L, Godzik A. Clustering of highly homologous sequences to reduce the size of large protein databases. Bioinformatics. 2001;17(3):282-283.
  • [20] Lee T, Main D. GenSAS: a web -based integrated genome sequence annotation pipeline. Methods Mol Biol. 2019; 1962:33-45. doi: 10.1007/978-1-4939-9173-0_4.
  • [21] Mount DW. Using BLOSUM in sequence alignment and evolutionary distance estimation. Methods Mol Biol. 2008; 484:575-580.
  • [22] Lemoine F, Domelevo Entfellner JB, Wilkinson E, et al. Renewing Felsenstein’s phylogenetic bootstrap in the era of big data. Nature. 2018;556(7702):452-456.
  • [23] Sievers F, Wilm A, Dineen D, et al. Fast, scalable generation of high -quality protein multiple sequence alignments using Clustal Omega. Mol Syst Biol. 2011; 7:539. doi: 10.1038/msb.2011.75.
  • [24] Sayers EW, Barrett T, Benson DA, et al. Database resources of the National Center for Biotechnology Information. Nucleic Acids Res. 2011;39(Database issue): D38-D51.
  • [25] Thimm O, Blä sing O, Gibon Y, et al. MAPMAN: a user-driven tool to display genomics data sets onto diagrams of metabolic pathways and other biological processes. Plant J. 2004;37(6):914-939.
  • [26] Quinlan AR, Hall IM. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 2010;26(6):841-842.
  • [27] Szklarczyk D, Gable AL, Lyon D, et al. STRING v11: protein-protein association networks with increased coverage, supporting functional discovery in genome -wide experimental datasets. Nucleic Acids Res. 2019;47(D1): D607-D613.
  • [28] Katoh K, Standley DM. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol Biol Evol. 2013;30(4):772-780.
  • [29] Bailey TL, Johnson J, Grant CE, Noble WS. The MEME Suite. Nucleic Acids Res. 2015;43(W1): W39 - W49.
  • [30] Xie C, Mao X, Huang J, et al. KOBAS 2.0: a web server for annotation and identification of enriched pathways and diseases. Nucleic Acids Res. 2011;39(Web Server issue): W316-W322.