जैव सूचना विज्ञान
Genetic algorithm July 2026 vanc
Bioinformatics (IPA: /ˌbaɪ.oʊˌɪnfɚˈmætɪks/) is an interdisciplinary field of science that develops computational methods and software tools for understanding biological data, especially when the data sets are large and complex. Bioinformatics integrates principles from biology, chemistry, physics, computer science, data science, computer programming, information engineering, mathematics, and statistics to analyze and interpret biological data.[1] This process can sometimes be referred to as computational biology; however, the distinction between the two terms is often disputed. The term computational biology can refer to building and using models of biological systems.
Some of the main sub-branches of bioinformatics are computational genomics, computational epigenetics, computational immunology, and computational metabolomics.
Computational, statistical, and computer programming techniques have been used for computer simulation analyses of biological queries. They include reused specific analysis "pipelines", particularly in the field of genomics, such as by the identification of genes and single nucleotide polymorphisms (SNPs). These pipelines are used to better understand the genetic basis of disease, unique adaptations, desirable properties (especially in agricultural species), or differences between populations. Bioinformatics also includes proteomics, which aims to understand the organizational principles within nucleic acid and protein sequences.[2]
Image and signal processing allow the extraction of useful results from large amounts of raw data. It aids in sequencing and annotating genomes and their observed mutations. Bioinformatics includes text mining of biological literature and the development of biological and gene ontologies to organize and query biological data. It also plays a role in the analysis of gene and protein expression and regulation. Bioinformatic tools aid in comparing, analyzing, and interpreting genetic and genomic data and in the understanding of evolutionary aspects of molecular biology. At a more integrative level, it helps analyze and catalogue the biological pathways and networks that are an important part of systems biology. In structural biology, it aids in the simulation and modeling of DNA,[3] RNA,[3][1] proteins[4] as well as biomolecular interactions.[5][6][7][8]
History
The first definition of the term bioinformatics was coined by Paulien Hogeweg and Ben Hesper in 1970, to refer to the study of information processes in biotic systems.[9][10][11][12][13] This definition placed bioinformatics as a field parallel to biochemistry (the study of chemical processes in biological systems).[10]
Bioinformatics and computational biology involved the analysis of biological data, particularly DNA, RNA, and protein sequences. The field of bioinformatics experienced explosive growth starting in the mid-1990s, driven largely by the Human Genome Project and by rapid advances in DNA sequencing technology.[14]
Analyzing biological data to produce meaningful information involves writing and running software programs that use algorithms from graph theory, artificial intelligence, soft computing, data mining, image processing, and computer simulation. The algorithms in turn depend on theoretical foundations such as discrete mathematics, control theory, system theory, information theory, and statistics.
Sequences
There has been a tremendous advance in speed and cost reduction since the completion of the Human Genome Project, with some labs able to sequence over 100,000 billion bases each year, and a full genome can be sequenced for $1,000 or less.[15]
Computers became essential in molecular biology when protein sequences became available after Frederick Sanger determined the sequence of insulin in the early 1950s.[16][17] Comparing multiple sequences manually turned out to be impractical. Margaret Oakley Dayhoff, a pioneer in the field,[18] compiled one of the first protein sequence databases, initially published as books[19] as well as methods of sequence alignment and molecular evolution.[20] Another early contributor to bioinformatics was Elvin A. Kabat, who pioneered biological sequence analysis in 1970 with his comprehensive volumes of antibody sequences released online with Tai Te Wu between 1980 and 1991.[21]
In the 1970s, new techniques for sequencing DNA were applied to bacteriophage MS2 and øX174, and the extended nucleotide sequences were then parsed with informational and statistical algorithms. These studies showed that well-known features, such as coding segments and the triplet code, could be revealed in straightforward statistical analyses and were proof of the concept that bioinformatics would be insightful.[22][23]
Goals
Bioinformatics focuses on the analysis, interpretation of various types of data combined to form a comprehensive picture of cell physiology. This includes nucleotide and amino acid sequences, protein domains, and protein structures.[24]
Important sub-disciplines within bioinformatics and computational biology include:
- Development and implementation of computer programs to efficiently access, manage, and use various types of information.
- Development of new mathematical algorithms and statistical measures to assess relationships among members of large data sets. For example, there are methods to locate a gene within a sequence, to predict protein structure and/or function, and to cluster protein sequences into families of related sequences.
The primary goal of bioinformatics is to increase the understanding of biological processes. What sets it apart from other approaches is its focus on developing and applying computationally intensive techniques to achieve this goal. Examples include: pattern recognition, data mining, machine learning algorithms, and visualization. Major research efforts in the field include sequence alignment, gene finding, genome assembly, drug design, drug discovery, protein structure alignment, protein structure prediction, prediction of gene expression and protein–protein interactions, genome-wide association studies, the modeling of evolution and cell division/mitosis.
Bioinformatics entails the creation and advancement of databases, algorithms, computational and statistical techniques, and theory to solve formal and practical problems arising from the management and analysis of biological data.
Over the past few decades, rapid developments in genomic and other molecular research technologies and developments in information technologies have combined to produce a tremendous amount of information related to molecular biology. Bioinformatics is the name given to these mathematical and computing approaches used to glean understanding of biological processes.
Common activities in bioinformatics include mapping and analyzing DNA and protein sequences, aligning DNA and protein sequences to compare them, and creating and viewing 3-D models of protein structures.
Sequence analysis
মূল নিবন্ধ: Sequence alignment, Sequence database, Alignment-free sequence analysis Since the bacteriophage Phage Φ-X174 was sequenced in 1977,[25] the DNA sequences of thousands of organisms have been decoded and stored in databases. This sequence information is analyzed to determine genes that encode proteins, RNA genes, regulatory sequences, structural motifs, and repetitive sequences. A comparison of genes within a species or between different species can show similarities between protein functions, or relations between species (the use of molecular systematics to construct phylogenetic trees). With the growing amount of data, it long ago became impractical to analyze DNA sequences manually. Computer programs such as BLAST are used routinely to search sequences—as of 2008, from more than 260,000 organisms, containing over 190 billion nucleotides.[26]
DNA sequencing
মূল নিবন্ধ: DNA sequencing
Before sequences can be analyzed, they are obtained from a data storage bank, such as GenBank. DNA sequencing is still a non-trivial problem as the raw data may be noisy or affected by weak signals. Algorithms have been developed for base calling for the various experimental approaches to DNA sequencing.
Sequence assembly
মূল নিবন্ধ: Sequence assembly Most DNA sequencing techniques produce short fragments of sequence that need to be assembled to obtain complete gene or genome sequences. The shotgun sequencing technique (used by The Institute for Genomic Research (TIGR) to sequence the first bacterial genome, Haemophilus influenzae)[27] generates the sequences of many thousands of small DNA fragments (ranging from 35 to 900 nucleotides long, depending on the sequencing technology). The ends of these fragments overlap and, when aligned properly by a genome assembly program, can be used to reconstruct the complete genome. Shotgun sequencing yields sequence data quickly, but the task of assembling the fragments can be quite complicated for larger genomes. For a genome as large as the human genome, it may take many days of CPU time on large-memory, multiprocessor computers to assemble the fragments, and the resulting assembly usually contains numerous gaps that must be filled in later. Shotgun sequencing is the method of choice for virtually all genomes sequenced (rather than chain-termination or chemical degradation methods), and genome assembly algorithms are a critical area of bioinformatics research.
আরও দেখুন: sequence analysis, sequence mining, sequence profiling tool, sequence motif
Genome annotation
মূল নিবন্ধ: Gene prediction
In genomics, annotation refers to the process of marking the stop and start regions of genes and other biological features in a sequenced DNA sequence. Many genomes are too large to be annotated by hand. As the rate of sequencing exceeds the rate of genome annotation, genome annotation has become the new bottleneck in bioinformatics.
Genome annotation can be classified into three levels: the nucleotide, protein, and process levels.
Gene finding is a chief aspect of nucleotide-level annotation. For complex genomes, a combination of ab initio gene prediction and sequence comparison with expressed sequence databases and other organisms can be successful. Nucleotide-level annotation also allows the integration of genome sequence with other genetic and physical maps of the genome.
The principal aim of protein-level annotation is to assign function to the protein products of the genome. Databases of protein sequences and functional domains and motifs are used for this type of annotation. About half of the predicted proteins in a new genome sequence tend to have no obvious function.
Understanding the function of genes and their products in the context of cellular and organismal physiology is the goal of process-level annotation. An obstacle of process-level annotation has been the inconsistency of terms used by different model systems. The Gene Ontology Consortium is helping to solve this problem.[28]
The first description of a comprehensive annotation system was published in 1995[27] by The Institute for Genomic Research, which performed the first complete sequencing and analysis of the genome of a free-living (non-symbiotic) organism, the bacterium Haemophilus influenzae.[27] The system identifies the genes encoding all proteins, transfer RNAs, ribosomal RNAs, in order to make initial functional assignments. The GeneMark program trained to find protein-coding genes in Haemophilus influenzae is constantly changing and improving.
Following the goals that the Human Genome Project left to achieve after its closure in 2003, the ENCODE project was developed by the National Human Genome Research Institute. This project is a collaborative data collection of the functional elements of the human genome that uses next-generation DNA-sequencing technologies and genomic tiling arrays, technologies able to automatically generate large amounts of data at a dramatically reduced per-base cost but with the same accuracy (base call error) and fidelity (assembly error).
Gene function prediction
While genome annotation is primarily based on sequence similarity (and thus homology), other properties of sequences can be used to predict the function of genes. In fact, most gene function prediction methods focus on protein sequences as they are more informative and more feature-rich. For instance, the distribution of hydrophobic amino acids predicts transmembrane segments in proteins. However, protein function prediction can also use external information such as gene (or protein) expression data, protein structure, or protein–protein interactions.[29]
Computational evolutionary biology
আরও দেখুন: Computational phylogenetics
विकासवादी जीवविज्ञान (Evolutionary biology), प्रजातियों (species) की उत्पत्ति और वंश परंपरा के साथ-साथ समय के साथ उनमें होने वाले परिवर्तनों का अध्ययन है। सूचना विज्ञान (Informatics) ने शोधकर्ताओं को निम्नलिखित कार्य करने में सक्षम बनाकर विकासवादी जीवविज्ञानियों की सहायता की है:
- शारीरिक वर्गीकरण या केवल शारीरिकीय प्रेक्षणों के बजाय, जीवों के डीएनए (DNA) में होने वाले परिवर्तनों को मापकर बड़ी संख्या में जीवों के विकास का पता लगाना,
- संपूर्ण जीनोम (genomes) की तुलना करना, जो अधिक जटिल विकासवादी घटनाओं के अध्ययन की अनुमति देता है, जैसे कि जीन डुप्लीकेशन (gene duplication), हॉरिजॉन्टल जीन ट्रांसफर (horizontal gene transfer), और बैक्टीरियल स्पीसिएशन (speciation) में महत्वपूर्ण कारकों की भविष्यवाणी,
- समय के साथ सिस्टम के परिणामों की भविष्यवाणी करने के लिए जटिल कम्प्यूटेशनल पॉपुलेशन जेनेटिक्स (population genetics) मॉडल बनाना[30]
- प्रजातियों और जीवों की लगातार बढ़ती संख्या पर जानकारी को ट्रैक करना और साझा करना
तुलनात्मक जीनोमिक्स (Comparative genomics)
মূল নিবন্ধ: Comparative genomics
तुलनात्मक जीनोम विश्लेषण का मूल विभिन्न जीवों में जीनों (ऑर्थोलॉजी विश्लेषण) या अन्य जीनोमिक विशेषताओं के बीच पत्राचार (correspondence) की स्थापना करना है। दो जीनोम के विचलन (divergence) के लिए जिम्मेदार विकासवादी प्रक्रियाओं का पता लगाने के लिए इंटरजीनोमिक मानचित्र बनाए जाते हैं। विभिन्न संगठनात्मक स्तरों पर काम करने वाली कई विकासवादी घटनाएँ जीनोम विकास को आकार देती हैं। सबसे निचले स्तर पर, बिंदु उत्परिवर्तन (point mutations) व्यक्तिगत न्यूक्लियोटाइड को प्रभावित करते हैं। उच्च स्तर पर, बड़े गुणसूपीय खंड (chromosomal segments) डुप्लीकेशन, पार्श्व स्थानांतरण (lateral transfer), व्युत्क्रमण (inversion), ट्रांसपोज़िशन, विलोपन (deletion) और प्रविष्टि (insertion) से गुजरते हैं।[31] संकरण (hybridization), पॉलीप्लॉइडाइजेशन और एंडोसियोबियोसिस (endosymbiosis) की प्रक्रियाओं में संपूर्ण जीनोम शामिल होते हैं जो तीव्र प्रजातीकरण (rapid speciation) की ओर ले जाते हैं। जीनोम विकास की जटिलता गणितीय मॉडल और एल्गोरिदम के डेवलपर्स के सामने कई रोमांचक चुनौתियां पेश करती है, जो पार्सिमनी मॉडल पर आधारित समस्याओं के लिए सटीक, ह्यूरिस्टिक्स (heuristics), निश्चित पैरामीटर और सन्निकटन एल्गोरिदम (approximation algorithms) से लेकर, प्रायिकता मॉडल पर आधारित समस्याओं के बेजियन विश्लेषण (Bayesian analysis) के लिए मार्कोव चेन मोंटे कार्लो (Markov chain Monte Carlo) एल्गोरिदम तक, एल्गोरिथम, सांख्यिकीय और गणितीय तकनीकों के स्पेक्ट्रम का सहारा लेते हैं।
इनमें से कई अध्ययन अनुक्रमों को प्रोटीन परिवारों (protein families) में असाइन करने के लिए अनुक्रम समरूपता (sequence homology) की पहचान पर आधारित हैं।[32]
पैन जीनोमिक्स (Pan genomics)
মূল নিবন্ধ: Pan-genome
पैन जीनोमिक्स 2005 में टेटेलिन (Tettelin) और मेडिनी (Medini) द्वारा प्रस्तुत एक अवधारणा है। पैन जीनोम किसी विशेष मोनोफायलेटिक (monophyletic) वर्गीकरण समूह की संपूर्ण जीन शस्त्रागार (gene repertoire) है। हालांकि शुरुआत में इसे किसी प्रजाति के निकट संबंधित उपभेदों (strains) पर लागू किया गया था, इसे वंश (genus), संघ (phylum) आदि जैसे बड़े संदर्भ में भी लागू किया जा सकता है। इसे दो भागों में विभाजित किया गया है: कोर जीनोम (Core genome), अध्ययन के तहत सभी जीनोम के लिए सामान्य जीनों का एक सेट (अक्सर जीवित रहने के लिए महत्वपूर्ण हाउसकीपिंग जीन), और डिस्पेंसेबल/फ्लेक्सिबल जीनोम (Dispensable/Flexible genome): जीनों का एक सेट जो अध्ययन के तहत सभी में नहीं बल्कि एक या कुछ जीनोम में मौजूद होता है। बैक्टीरियल प्रजातियों के पैन जीनोम की विशेषता बताने के लिए बायोइन्फॉर्मेटिक्स टूल BPGA का उपयोग किया जा सकता है।[33]
रोग की आनुवंशिकी (Genetics of disease)
মূল নিবন্ধ: Genome-wide association studies
2013 तक, कुशल उच्च-थ्रूपुट अगली पीढ़ी की अनुक्रमण (next-generation sequencing) तकनीक के अस्तित्व ने कई अलग-अलग मानव विकारों के कारणों की पहचान की अनुमति दी है। सरल मेंडेलियन वंशागति (Mendelian inheritance) को 3,000 से अधिक विकारों के लिए देखा गया है जिनकी पहचान ऑनलाइन मेंडेलियन इनहेरिटेंस इन मैन (Online Mendelian Inheritance in Man) डेटाबेस में की गई है, लेकिन जटिल बीमारियाँ अधिक कठिन हैं। एसोसिएशन अध्ययनों ने कई व्यक्तिगत आनुवंशिक क्षेत्रों को पाया है जो व्यक्तिगत रूप से जटिल बीमारियों (जैसे बांझपन (infertility),[34] स्तन कैंसर (breast cancer)[35] और अ Alzheimer रोग (Alzheimer's disease)[36]) से कमजोर रूप से जुड़े हुए हैं, बजाए किसी एकल कारण के।[37][38] वर्तमान में निदान और उपचार के लिए जीनों का उपयोग करने में कई चुनौतियाँ हैं, जैसे कि हम नहीं जानते कि कौन से जीन महत्वपूर्ण हैं, या एल्गोरिदम द्वारा प्रदान किए जाने वाले विकल्प कितने स्थिर हैं।[39]
जीनोम-व्यापी सहसंबंध अध्ययन (Genome-wide association studies) ने जटिल रोगों और लक्षणों के लिए हज़ारों सामान्य आनुवंशिक प्रकारों (genetic variants) की सफलतापूर्वक पहचान की है; हालांकि, ये सामान्य प्रकार वंशागतिकता (heritability) के केवल एक छोटे से हिस्से को स्पष्ट करते हैं।[40] दुर्लभ प्रकार (Rare variants) इस लापता वंशागतिकता (missing heritability) के कुछ हिस्से के लिए जिम्मेदार हो सकते हैं।[41] बड़े पैमाने पर किए जाने वाले संपूर्ण जीनोम अनुक्रमण (whole genome sequencing) अध्ययनों ने तेज़ी से लाखों संपूर्ण जीनोमों का अनुक्रमण किया है, और ऐसे अध्ययनों ने करोड़ों दुर्लभ प्रकारों की पहचान की है।[42] कार्यात्मक एनोटेशन (Functional annotations) किसी आनुवंशिक प्रकार के प्रभाव या कार्य की भविष्यवाणी करती हैं और दुर्लभ कार्यात्मक प्रकारों को प्राथमिकता देने में मदद करती हैं, तथा इन एनोटेशन को शामिल करने से संपूर्ण जीनोम अनुक्रमण अध्ययनों के दुर्लभ प्रकार के आनुवंशिक सहसंबंध विश्लेषण की शक्ति को प्रभावी ढंग से बढ़ाया जा सकता है।[43] संपूर्ण जीनोम अनुक्रमण (whole-genome sequencing) डेटा के लिए ऑल-इन-वन (all-in-one) दुर्लभ वेरिएंट एसोसिएशन विश्लेषण प्रदान करने के वास्ते कुछ उपकरण विकसित किए गए हैं, जिनमें जीनोटाइप डेटा और उनके कार्यात्मक एनोटेशन का एकीकरण, एसोसिएशन विश्लेषण, परिणाम सारांश और विज़ुअलाइज़ेशन शामिल हैं।[44][45] संपूर्ण जीनोम अनुक्रमण अध्ययनों का मेटा-विश्लेषण (Meta-analysis), जटिल फीनोटाइप्स (phenotypes) से जुड़े दुर्लभ वेरिएंट की खोज के लिए बड़े नमूना आकार (sample sizes) को एकत्र करने की समस्या का एक आकर्षक समाधान प्रदान करता है।[46]
कैंसर में उत्परिवर्तन का विश्लेषण
মূল নিবন্ধ: ओंकोजेनोमिक्स
कैंसर में, प्रभावित कोशिकाओं के जीनोम जटिल या अप्रत्याशित तरीकों से पुनर्व्यवस्थित होते हैं। कैंसर का कारण बनने वाले पॉइंट म्यूटेशन (बिंदु उत्परिवर्तन) की पहचान करने वाली सिंगल-न्यूक्लियोटाइड पॉलिमॉर्फिज्म एरे के अलावा, गुणसूत्रों की बढ़त और हानि (जिसे तुलनात्मक जीनोमिक हाइब्रिडाइजेशन कहा जाता है) की पहचान करने के लिए ओलिगोसैकराइड (oligonucleotide) माइक्रोक्रे का उपयोग किया जा सकता है। ये पहचान विधियाँ प्रति प्रयोग टेराबाइट डेटा उत्पन्न करती हैं।[47] इस डेटा में काफी परिवर्तनशीलता, या शोर (noise) पाया जाता है, और इसलिए वास्तविक कॉपी नंबर परिवर्तनों का अनुमान लगाने के लिए हिडन मार्कोव मॉडल और चेंज-पॉइंट विश्लेषण विधियाँ विकसित की जा रही हैं।[48]
एक्सोम में उत्परिवर्तन द्वारा कैंसर की पहचान करने के लिए दो महत्वपूर्ण सिद्धांतों का उपयोग किया जा सकता है। पहला, कैंसर जीनों में संचित दैहिक (सोमैटिक) उत्परिवर्तन का एक रोग है। दूसरा, कैंसर में ड्राइवर उत्परिवर्तन होते हैं जिन्हें पैसेंजर (यात्री) उत्परिवर्तन से अलग करने की आवश्यकता होती है।[49]
बायोइनफॉर्मेटिक्स में और सुधार से जीनोम में कैंसर को प्रेरित करने वाले उत्परिवर्तन के विश्लेषण द्वारा कैंसर के प्रकारों को वर्गीकृत करना संभव हो सकता है। इसके अलावा, रोग की प्रगति के दौरान कैंसर के नमूनों के अनुक्रम के साथ रोगियों को ट्रैक करना भविष्य में संभव हो सकता है। डेटा का एक अन्य प्रकार जिसके लिए नए इंफॉर्मेटिक्स विकास की आवश्यकता है, वह कई ट्यूमर के बीच आवर्ती पाए जाने वाले लेजन (क्षतियों) का विश्लेषण है।[50]
जीन और प्रोटीन अभिव्यक्ति
जीन अभिव्यक्ति का विश्लेषण
कई जीनों की अभिव्यक्ति (expression) को माइक्रोक्रे, एक्सप्रेस्ड सीडीएनए सीक्वेंस टैग (EST) अनुक्रमण, सीरियल एनालिसिस ऑफ जीन एक्सप्रेशन (SAGE) टैग अनुक्रमण, मासिवली पैरेलल सिग्नेचर सीक्वेंसिंग (MPSS), RNA-Seq (जिसे "होल ट्रांसक्रिप्टोम शॉटगन सीक्वेंसिंग" या WTSS भी कहा जाता है) सहित कई तकनीकों के साथ mRNA स्तरों को मापकर निर्धारित किया जा सकता है, या मल्टीप्लेक्स इन-सीटू हाइब्रिडाइजेशन के विभिन्न अनुप्रयोगों द्वारा निर्धारित किया जा सकता है। ये सभी तकनीकें अत्यधिक शोर-प्रवण (noise-prone) हैं और/या जैविक माप में पूर्वाग्रह के अधीन हैं, और कम्प्यूटेशनल जीव विज्ञान में एक प्रमुख अनुसंधान क्षेत्र में हाई-थ्रूपुट जीन अभिव्यक्ति अध्ययनों में शोर (noise) से सिग्नल को अलग करने के लिए सांख्यिकीय उपकरण विकसित करना शामिल है।[51] ऐसे अध्ययनों का उपयोग अक्सर किसी विकार में शामिल जीनों को निर्धारित करने के लिए किया जाता है: कैंसर कोशिकाओं की एक विशेष आबादी में अप-रेगुलेटेड और डाउन-रेगुलेटेड ट्रांसक्रिप्ट को निर्धारित करने के लिए कोई कैंसरयुक्त उपकला (epithelial) कोशिकाओं के माइक्रोक्रे डेटा की गैर-कैंसरयुक्त कोशिकाओं के डेटा से तुलना कर सकता है।
प्रोटीन अभिव्यक्ति का विश्लेषण
प्रोटीन माइक्रोक्रे और हाई-थ्रूपुट (HT) मास स्पेक्ट्रोमेट्री (MS) एक जैविक नमूने में मौजूद प्रोटीन का एक स्नैपशॉट प्रदान कर सकते हैं। पूर्व दृष्टिकोण को mRNA को लक्षित करने वाले माइक्रोक्रे जैसी ही समस्याओं का सामना करना पड़ता है, बाद वाले में प्रोटीन अनुक्रम डेटाबेस से अनुमानित द्रव्यमान के मुकाबले बड़ी मात्रा में द्रव्यमान डेटा का मिलान करने की समस्या शामिल है, और जब प्रत्येक प्रोटीन से कई अधूरे पेप्टाइड्स का पता लगाया जाता है तो नमूनों का जटिल सांख्यिकीय विश्लेषण होता है। ऊतक के संदर्भ में सेलुलर प्रोटीन स्थानीयकरण को इम्यूनोहिस्टोकैमिस्ट्री और टिश्यू माइक्रोक्रे के आधार पर स्थानिक डेटा के रूप में प्रदर्शित आत्मीयता प्रोटिओमिक्स के माध्यम से प्राप्त किया जा सकता है।[52]
विनियमन का विश्लेषण
जीन विनियमन एक जटिल प्रक्रिया है जहाँ एक संकेत, जैसे कि हार्मोन जैसे बाह्य कोशिका संकेत, अंततः एक या एक से अधिक प्रोटीन की गतिविधि में वृद्धि या कमी की ओर ले जाता है। इस प्रक्रिया में विभिन्न चरणों का पता लगाने के लिए बायोइनफॉर्मेटिक्स तकनीकों को लागू किया गया है।
उदाहरण के लिए, जीन अभिव्यक्ति को जीनोम में आसन्न तत्वों द्वारा विनियमित किया जा सकता है। प्रमोटर विश्लेषण में किसी जीन के प्रोटीन-कोडिंग क्षेत्र के आसपास के डीएनए में अनुक्रम रूपांकनों (sequence motifs) की पहचान और अध्ययन शामिल है। ये रूपांकने इस बात को प्रभावित करते हैं कि उस क्षेत्र को किस हद तक mRNA में प्रतिलेखित (transcribe) किया जाता है। प्रमोटर से बहुत दूर स्थित एन्हांसर तत्व त्रि-आयामी लूपिंग इंटरैक्शन के माध्यम से जीन अभिव्यक्ति को भी विनियमित कर सकते हैं। इन इंटरैक्शन को क्रोमोसोम कॉन्फ़िगरेशन कैप्चर प्रयोगों के बायोइनफॉर्मेटिक्स विश्लेषण द्वारा निर्धारित किया जा सकता है।
जीन विनियमन का अनुमान लगाने के लिए अभिव्यक्ति डेटा का उपयोग किया जा सकता है: प्रत्येक स्थिति में शामिल जीनों के बारे में परिकल्पना बनाने के लिए कोई जीव की विभिन्न प्रकार की अवस्थाओं से माइक्रोक्रे डेटा की तुलना कर सकता है। एक एकल-कोशिका जीव में, कोई विभिन्न तनाव स्थितियों (हीट शॉक, भुखमरी, आदि) के साथ-साथ कोशिका चक्र के चरणों की तुलना कर सकता है। यह निर्धारित करने के लिए कि कौन से जीन सह-अभिव्यक्त (co-expressed) हैं, अभिव्यक्ति डेटा पर क्लस्टरिंग एल्गोरिदम लागू किए जा सकते हैं। उदाहरण के लिए, सह-अभिव्यक्त जीनों के अपस्ट्रीम क्षेत्रों (प्रमोटरों) को अति-प्रतिनिधि नियामक तत्वों के लिए खोजा जा सकता है। जीन क्लस्टरिंग में लागू क्लस्टरिंग एल्गोरिदम के उदाहरणों में के-मीन्स क्लस्टरिंग, सेल्फ-ऑर्गेनाइजिंग मैप (SOMs), श्रेणीबद्ध क्लस्टरिंग (hierarchical clustering), और सहमति क्लस्टरिंग (consensus clustering) विधियाँ शामिल हैं।
कोशिका संगठन का विश्लेषण
कोशिकाओं के भीतर ऑर्गेनेल, जीन, प्रोटीन और अन्य घटकों के स्थान का विश्लेषण करने के लिए कई दृष्टिकोण विकसित किए गए हैं। कई जैविक डेटाबेस में उपकोशिकीय स्थानीयकरण को कैप्चर करने के लिए एक जीन ऑन्टोलॉजी श्रेणी, सेलुलर घटक, तैयार की गई है।
माइक्रोस्कोपी और छवि विश्लेषण
सूक्ष्मदर्शी चित्र ऑर्गेनेल के साथ-साथ अणुओं के स्थान की अनुमति देते हैं, जो रोगों में असामान्यताओं का स्रोत हो सकते हैं।
प्रोटीन स्थानीयकरण
प्रोटीन का स्थान खोजने से हमें यह अनुमान लगाने की अनुमति मिलती है कि वे क्या करते हैं। इसे प्रोटीन कार्य भविष्यवाणी कहा जाता है। उदाहरण के लिए, यदि कोई प्रोटीन केंद्रक में पाया जाता है तो यह जीन विनियमन या स्प्लिसिंग में शामिल हो सकता है। इसके विपरीत, यदि कोई प्रोटीन माइटोकॉन्ड्रिया में पाया जाता है, तो यह श्वसन या अन्य उपापचय प्रक्रियाओं में शामिल हो सकता है। प्रोटीन उपकोशिकीय स्थान डेटाबेस और भविष्यवाणी टूल सहित अच्छी तरह से विकसित प्रोटीन उपकोशिकीय स्थानीयकरण भविष्यवाणी संसाधन उपलब्ध हैं।[53][54]
क्रोमैटिन का नाभिकीय संगठन
মূল নিবন্ধ: Nuclear organization हाई-थ्रूपुट गुणसूत्र संसूचन कैप्चर प्रयोगों जैसे कि Hi-C (experiment) और ChIA-PET से प्राप्त डेटा, त्रि-आयामी संरचना और क्रोमैटिन के नाभिकीय संगठन पर जानकारी प्रदान कर सकता है। इस क्षेत्र में बायोइनफॉर्मेटिक्स की चुनौतियों में जीनोम को डोमेन में विभाजित करना शामिल है, जैसे कि टोपोलॉजिकल एसोसिएटिंग डोमेन (TADs), जो त्रि-आयामी अंतरिक्ष में एक साथ व्यवस्थित होते हैं।[55]
संरचनात्मक बायोइनफॉर्मेटिक्स
মূল নিবন্ধ: Structural bioinformatics, Protein structure prediction আরও দেখুন: Structural motif, Structural domain
प्रोटीन की संरचना का पता लगाना बायोइनफॉर्मेटिक्स का एक महत्वपूर्ण अनुप्रयोग है। प्रोटीन संरचना भविष्यवाणी का गंभीर मूल्यांकन (CASP) एक खुली प्रतियोगिता है जहाँ दुनिया भर के शोध समूह अज्ञात प्रोटीन मॉडल का मूल्यांकन करने के लिए प्रोटीन मॉडल जमा करते हैं।[56][57]
अमीनो अम्ल अनुक्रम
प्रोटीन के रैखिक अमीनो अम्ल अनुक्रम को प्राथमिक संरचना कहा जाता है। प्राथमिक संरचना को इसके लिए कोड करने वाले डीएनए जीन पर कोडन के अनुक्रम से आसानी से निर्धारित किया जा सकता है। अधिकांश प्रोटीनों में, प्राथमिक संरचना अपने मूल वातावरण में किसी प्रोटीन की त्रि-आयामी संरचना को विशिष्ट रूप से निर्धारित करती है। इसका एक अपवाद बोवाइन स्पॉंजीफॉर्म एन्सेफैलोपैथी में शामिल गलत तरीके से मुड़ा हुआ प्रियन प्रोटीन है। यह संरचना प्रोटीन के कार्य से जुड़ी हुई है। अतिरिक्त संरचनात्मक जानकारी में द्वितीयक, तृतीयक और चतुर्धातुक संरचना शामिल हैं। प्रोटीन के कार्य की भविष्यवाणी के लिए एक व्यावहारिक सामान्य समाधान एक खुली समस्या बनी हुई है। अधिकांश प्रयास अब तक उन ह्यूरिस्टिक्स (heuristics) की ओर निर्देशित किए गए हैं जो अधिकांश समय काम करते हैं।
समजातीयता (Homology)
बायोइनफॉर्मेटिक्स की जीनोमिक शाखा में, समजातीयता का उपयोग किसी जीन के कार्य की भविष्यवाणी करने के लिए किया जाता है: यदि जीन A का अनुक्रम, जिसका कार्य ज्ञात है, जीन B के अनुक्रम के समजात (homologous) है, जिसका कार्य अज्ञात है, तो कोई यह अनुमान लगा सकता है कि B, A के कार्य को साझा कर सकता है। संरचनात्मक बायोइनफॉर्मेटिक्स में, समजातीयता का उपयोग यह निर्धारित करने के लिए किया जाता है कि प्रोटीन के कौन से हिस्से संरचना गठन और अन्य प्रोटीनों के साथ संपर्क में महत्वपूर्ण हैं। मौजूदा समजात प्रोटीनों से किसी अज्ञात प्रोटीन की संरचना की भविष्यवाणी करने के लिए समजात मॉडलिंग (Homology modeling) का उपयोग किया जाता है।
इसका एक उदाहरण मनुष्यों में हीमोग्लोबिन और फलदार पौधों (दलहनों) में हीमोग्लोबिन (लेहीमोग्लोबिन) है, जो एक ही प्रोटीन सुपरपरिवार के दूर के रिश्तेदार हैं। दोनों जीव में ऑक्सीजन के परिवहन के समान उद्देश्य को पूरा करते हैं। हालांकि इन दोनों प्रोटीनों में बहुत अलग अमीनो अम्ल अनुक्रम हैं, फिर भी उनकी प्रोटीन संरचनाएं बहुत समान हैं, जो उनके साझा कार्य और साझा पूर्वज को दर्शाती हैं।[58]
प्रोटीन संरचना की भविष्यवाणी के लिए अन्य तकनीकों में प्रोटीन थ्रेडिंग और डी नोवो (शुरुआत से) भौतिकी-आधारित मॉडलिंग शामिल हैं।
संरचनात्मक बायोइनफॉर्मेटिक्स के एक अन्य पहलू में वर्चुअल स्क्रीनिंग मॉडल जैसे मात्रात्मक संरचना-गतिविधि संबंध मॉडल और प्रोटिओकेमेट्रिक मॉडल (PCM) के लिए प्रोटीन संरचनाओं का उपयोग शामिल है। इसके अलावा, उदाहरण के लिए लिगैंड-बाइंडिंग अध्ययन और इन सिलिको म्यूटोजेनेसिस अध्ययन के अनुकरण में प्रोटीन की क्रिस्टल संरचना का उपयोग किया जा सकता है।
Google के DeepMind द्वारा विकसित AlphaFold नामक 2021 के एक डीप-लर्निंग एल्गोरिदम-आधारित सॉफ्टवेयर ने अल्फाफोल्ड प्रोटीन संरचना डेटाबेस में सैकड़ों मिलियन प्रोटीनों के लिए अनुमानित संरचनाएं जारी कीं।[59]
नेटवर्क और सिस्टम जीव विज्ञान
মূল নিবন্ধ: Computational systems biology, Biological network, Interactome
नेटवर्क विश्लेषण जैविक नेटवर्क जैसे उपापचय या प्रोटीन-प्रोटीन इंटरैक्शन नेटवर्क के भीतर के संबंधों को समझने का प्रयास करता है। हालांकि जैविक नेटवर्क को अणु या इकाई के एक ही प्रकार (जैसे जीन) से बनाया जा सकता है, नेटवर्क जीव विज्ञान अक्सर कई अलग-अलग डेटा प्रकारों, जैसे प्रोटीन, छोटे अणु, जीन अभिव्यक्ति डेटा, और अन्य को एकीकृत करने का प्रयास करता है, जो सभी भौतिक रूप से, कार्यात्मक रूप से, या दोनों रूप से जुड़े हुए हैं।
सिस्टम जीव विज्ञान में इन सेलुलर प्रक्रियाओं के जटिल कनेक्शन का विश्लेषण और विज़ुअलाइज़ेशन दोनों करने के लिए कोशिका उप-प्रणालियों (जैसे मेटाबोलाइट्स और एंजाइमों के नेटवर्क जो उपापचय, सिग्नल ट्रांसडक्शन मार्ग और जीन नियामक नेटवर्क बनाते हैं) के कंप्यूटर सिमुलेशन का उपयोग शामिल है। कृत्रिम जीवन या आभासी विकास सरल (कृत्रिम) जीवन रूपों के कंप्यूटर सिमुलेशन के माध्यम से विकासात्मक प्रक्रियाओं को समझने का प्रयास करता है।
आणविक अंतःक्रिया नेटवर्क (Molecular interaction networks)
মূল নিবন্ধ: Protein–protein interaction prediction, interactome [60]
एक्स-रे क्रिस्टलोग्राफी और प्रोटीन परमाणु चुंबकीय अनुनाद स्पेक्ट्रोस्कोपी (प्रोटीन NMR) द्वारा हजारों त्रि-आयामी प्रोटीन संरचनाओं निर्धारित की गई हैं और संरचनात्मक बायोइनफॉर्मेटिक्स में एक केंद्रीय प्रश्न यह है कि क्या प्रोटीन-प्रोटीन इंटरैक्शन प्रयोगों को किए बिना, केवल इन 3D आकारों के आधार पर संभावित प्रोटीन-प्रोटीन इंटरैक्शन की भविष्यवाणी करना व्यावहारिक है। प्रोटीन-प्रोटीन डॉकिंग समस्या से निपटने के लिए विभिन्न प्रकार की विधियाँ विकसित की गई हैं, हालांकि ऐसा प्रतीत होता है कि इस क्षेत्र में अभी बहुत काम किया जाना बाकी है।
क्षेत्र में आने वाली अन्य अंतःक्रियाओं में प्रोटीन-लिगैंड (दवा सहित) और प्रोटीन-पेप्टाइड शामिल हैं। घूर्णन योग्य बांडों के बारे में परमाणुओं की गति का आणविक गतिशील सिमुलेशन आणविक अंतःक्रियाओं का अध्ययन करने के लिए डॉकिंग एल्गोरिदम नामक कम्प्यूटेशनल एल्गोरिदम के पीछे का मूल सिद्धांत है।
जैवविविधता बायोइनफॉर्मेटिक्स
মূল নিবন্ধ: Biodiversity informatics जैवविविधता बायोइनफॉर्मेटिक्स जैवविविधता डेटा के संग्रह और विश्लेषण से संबंधित है, जैसे वर्गीकरण डेटाबेस (taxonomic databases), या माइक्रोबायोम डेटा। ऐसे विश्लेषणों के उदाहरणों में फाइलोजीनेटिक्स, स्थान मॉडलिंग (niche modelling), प्रजाति समृद्धि मैपिंग, डीएनए बारकोडिंग, या प्रजाति पहचान उपकरण शामिल हैं। एक बढ़ता हुआ क्षेत्र मैक्रोइकोलॉजी भी है, यानी यह अध्ययन कि जैवविविधता पारिस्थितिकी और मानवीय प्रभाव से कैसे जुड़ी है, जैसे कि जलवायु परिवर्तन।
अन्य
साहित्य विश्लेषण
মূল নিবন্ধ: Text mining, Biomedical text mining
प्रकाशित साहित्य की भारी संख्या व्यक्तियों के लिए हर पेपर को पढ़ना व्यावहारिक रूप से असंभव बना देती है, जिसके परिणामस्वरूप अनुसंधान के असंबद्ध उप-क्षेत्र होते हैं। साहित्य विश्लेषण का उद्देश्य पाठ संसाधनों के इस बढ़ते पुस्तकालय को खंगालने के लिए कम्प्यूटेशनल और सांख्यिकीय भाषाविज्ञान को नियोजित करना है। उदाहरण के लिए:
- संक्षिप्त नाम पहचान – जैविक शब्दों के लंबे रूप और संक्षिप्त नाम की पहचान करना
- नामित-इकाई पहचान (Named-entity recognition) – जीन नामों जैसे जैविक शब्दों को पहचानना
- प्रोटीन-प्रोटीन अंतःक्रिया – पाठ से पहचानें कि कौन से प्रोटीन किस प्रोटीन के साथ अंतःक्रिया करते हैं
अनुसंधान का यह क्षेत्र सांख्यिकी और कम्प्यूटेशनल भाषाविज्ञान से लिया गया है।
हाई-थ्रूपुट छवि विश्लेषण
बड़ी मात्रा में उच्च-सूचना-सामग्री वाले बायोमेडिकल इमेजरी के प्रसंस्करण, मात्रा निर्धारण (quantification) और विश्लेषण को स्वचालित करने के लिए कम्प्यूटेशनल तकनीकों का उपयोग किया जाता है। आधुनिक छवि विश्लेषण प्रणालियाँ किसी प्रेक्षक की सटीकता, वस्तुनिष्ठता, या गति में सुधार कर सकती हैं। छवि विश्लेषण निदान (diagnostics) और अनुसंधान दोनों के लिए महत्वपूर्ण है। कुछ उदाहरण हैं:
- हाई-थ्रूपुट और उच्च-निष्ठा मात्रा निर्धारण और उप-कोशिकीय स्थानीयकरण (हाई-कंट्रोल स्क्रीनिंग, साइटोहिस्टोपैथोलॉजी, बायोइमेज इंफॉर्मेटिक्स)
- मॉर्फोमेट्रिक्स
- नैदानिक छवि विश्लेषण और विज़ुअलाइज़ेशन
- जीवित जानवरों के सांस लेने वाले फेफड़ों में वास्तविक समय के वायु-प्रवाह पैटर्न का निर्धारण करना
- धमनियों की चोट के विकास के दौरान और ठीक होने के दौरान वास्तविक समय के दृश्यों में ओक्लुजन (रुकावट) के आकार को मापना
- प्रयोगशाला जानवरों की विस्तारित वीडियो रिकॉर्डिंग से व्यवहार संबंधी अवलोकन करना
- चयापचय गतिविधि के निर्धारण के लिए इन्फ्रारेड माप
- डीएनए मैपिंग में क्लोन ओवरलैप का अनुमान लगाना, उदाहरण के लिए सल्स्टन स्कोर
हाई-थ्रूपुट एकल कोशिका डेटा विश्लेषण
মূল নিবন্ধ: Flow cytometry bioinformatics
हाई-थ्रूपुट, कम-माप वाले एकल कोशिका डेटा का विश्लेषण करने के लिए कम्प्यूटेशनल तकनीकों का उपयोग किया जाता है, जैसे कि फ्लो साइटोमेट्री से प्राप्त डेटा। इन विधियों में आम तौर पर कोशिकाओं की आबादी को ढूंढना शामिल होता है जो किसी विशेष रोग अवस्था या प्रायोगिक स्थिति से प्रासंगिक हैं।
ऑन्टोलॉजी और डेटा एकीकरण
जैविक ऑन्टोलॉजी नियंत्रित शब्दावली के निर्देशित एसाइक्लिक ग्राफ हैं। वे जैविक अवधारणाओं और विवरणों के लिए श्रेणियां बनाते हैं ताकि उनका कंप्यूटरों के साथ आसानी से विश्लेषण किया जा सके। जब इस तरह से वर्गीकृत किया जाता है, तो समग्र और एकीकृत विश्लेषण से अतिरिक्त मूल्य प्राप्त करना संभव हो जाता है।[61]
OBO फाउंड्री कुछ ऑन्टोलॉजी को मानकीकृत करने का एक प्रयास था। सबसे व्यापक में से एक जीन ऑन्टोलॉजी है जो जीन के कार्य का वर्णन करती है। ऐसी ऑन्टोलॉजी भी हैं जो फेनोटाइप का वर्णन करती हैं।
डेटाबेस
মূল নিবন্ধ: List of biological databases, Biological database
बायोइनफॉर्मेटिक्स अनुसंधान और अनुप्रयोगों के लिए डेटाबेस आवश्यक हैं। डीएनए और प्रोटीन अनुक्रमों, आणविक संरचनाओं, फेनोटाइप और जैवविविधता सहित कई अलग-अलग सूचना प्रकारों के लिए डेटाबेस मौजूद हैं। डेटाबेस में अनुभवजन्य डेटा (सीधे प्रयोगों से प्राप्त) और अनुमानित डेटा (मौजूदा डेटा के विश्लेषण से प्राप्त) दोनों हो सकते हैं। वे किसी विशेष जीव, मार्ग या रुचि के अणु के लिए विशिष्ट हो सकते हैं। वैकल्पिक रूप से, वे कई अन्य डेटाबेस से संकलित डेटा को शामिल कर सकते हैं। डेटाबेस में अलग-अलग प्रारूप, पहुंच तंत्र हो सकते हैं, और वे सार्वजनिक या निजी हो सकते हैं।
सबसे अधिक उपयोग किए जाने वाले डेटाबेस में से कुछ नीचे सूचीबद्ध हैं:
- जैविक अनुक्रम विश्लेषण में प्रयुक्त: Genbank, UniProt
- संरचना विश्लेषण में प्रयुक्त: प्रोटीन डेटा बैंक (PDB)
- प्रोटीन परिवारों और मोटिफ की खोज में प्रयुक्त: InterPro, Pfam
- अगली पीढ़ी के अनुक्रमण के लिए प्रयुक्त: Sequence Read Archive
- नेटवर्क विश्लेषण में प्रयुक्त: चयापचय मार्ग डेटाबेस (KEGG, BioCyc), इंटरैक्शन विश्लेषण डेटाबेस, कार्यात्मक नेटवर्क
- सिंथेटिक आनुवंशिक सर्किट के डिजाइन में प्रयुक्त: GenoCAD
सॉफ्टवेयर और उपकरण
बायोइनफॉर्मेटिक्स के लिए सॉफ्टवेयर उपकरण में सरल कमांड-लाइन टूल, अधिक जटिल ग्राफिकल प्रोग्राम और स्टैंडअलोन वेब-सेवाएं शामिल हैं। वे बायोइनफॉर्मेटिक्स कंपनियों द्वारा या सार्वजनिक संस्थानों द्वारा बनाए जाते हैं।
ओपन-सोर्स बायोइनफॉर्मेटिक्स सॉफ्टवेयर
মূল নিবন্ধ: List of open-source bioinformatics software আরও দেখুন: List of bioinformatics software
1980 के दशक से कई मुक्त और ओपन-सोर्स सॉफ्टवेयर टूल मौजूद रहे हैं और बढ़ते रहे हैं।[62] उभरते प्रकार के जैविक रीडआउट के विश्लेषण के लिए नए एल्गोरिदम की निरंतर आवश्यकता, अभिनव इन सिलिको प्रयोगों की क्षमता, और स्वतंत्र रूप से उपलब्ध ओपन कोड बेस के संयोजन ने अनुसंधान समूहों के लिए फंडिंग की परवाह किए बिना बायोइनफॉर्मेटिक्स दोनों में योगदान करने के अवसर पैदा किए हैं। ओपन सोर्स टूल अक्सर विचारों के इनक्यूबेटर, या वाणिज्यिक अनुप्रयोगों में समुदाय-समर्थित प्लग-इन के रूप में कार्य करते हैं। वे जैवसूचना एकीकरण की चुनौती में सहायता के लिए डी फैक्टो मानक और साझा वस्तु मॉडल भी प्रदान कर सकते हैं।
ओपन-सोर्स बायोइनफॉर्मेटिक्स सॉफ्टवेयर में Bioconductor, BioPerl, Biopython, BioJava, BioJS, BioRuby, Bioclipse, EMBOSS, .NET Bio, अपने बायोइनफॉर्मेटिक्स एड-ऑन के साथ Orange, Apache Taverna, UGENE और GenoCAD शामिल हैं।
गैर-लाभकारी ओपन बायोइनफॉर्मेटिक्स फाउंडेशन[62] और वार्षिक बायोइनफॉर्मेटिक्स ओपन सोर्स कॉन्फ्रेंस ओपन-सोर्स बायोइनफॉर्मेटिक्स सॉफ्टवेयर को बढ़ावा देते हैं।[63]
बायोइनफॉरमैटिक्स में वेब सेवाएँ
SOAP- और REST-आधारित इंटरफेस विकसित किए गए हैं ताकि क्लाइंट कंप्यूटर दुनिया के अन्य हिस्सों के सर्वर से एल्गोरिदम, डेटा और कंप्यूटिंग संसाधनों का उपयोग कर सकें। इसका मुख्य लाभ यह है कि अंत-उपयोगकर्ताओं (end users) को सॉफ्टवेयर और डेटाबेस रखरखाव के ओवरहेड्स से नहीं जूझना पड़ता है।
बुनियादी बायोइनफॉरमैटिक्स सेवाओं को EBI द्वारा तीन श्रेणियों में वर्गीकृत किया गया है: SSS (सीक्वेंस सर्च सर्विसेज), MSA (मल्टीपल सीक्वेंस एलाइनमेंट), और BSA (बायोलॉजिकल सीक्वेंस एनालिसिस)।[64] इन सेवा-उन्मुख (service-oriented) बायोइनफॉरमैटिक्स संसाधनों की उपलब्धता वेब-आधारित बायोइनफॉरमैटिक्स समाधानों की प्रयोज्यता को प्रदर्शित करती है, और ये एक ही वेब-आधारित इंटरफेस के तहत सामान्य डेटा प्रारूप वाले स्टैंडअलोन टूल के संग्रह से लेकर, एकीकृत, वितरित और विस्तार योग्य बायोइनफॉरमैटिक्स वर्कफ़्लो प्रबंधन प्रणालियों तक फैली हुई हैं।
बायोइनफॉरमैटिक्स वर्कफ़्लो प्रबंधन प्रणालियाँ
মূল নিবন্ধ: Bioinformatics workflow management systems
एक बायोइनफॉरमैटिक्स वर्कफ़्लो प्रबंधन प्रणाली, वर्कफ़्लो प्रबंधन प्रणाली का एक विशिष्ट रूप है जिसे विशेष रूप से बायोइनफॉरमैटिक्स अनुप्रयोग में संगणकीय (computational) या डेटा हेरफेर चरणों की एक श्रृंखला, या एक वर्कफ़्लो को संयोजित और निष्पादित करने के लिए डिज़ाइन किया गया है। ऐसी प्रणालियाँ निम्नलिखित कार्य करने के लिए डिज़ाइन की गई हैं:
- व्यक्तिगत अनुप्रयोग वैज्ञानिकों के लिए अपने स्वयं के वर्कफ़्लो बनाने हेतु उपयोग में आसान वातावरण प्रदान करना,
- वैज्ञानिकों के लिए ऐसे संवादात्मक (interactive) उपकरण प्रदान करना जो उन्हें अपने वर्कफ़्लो को निष्पादित करने और वास्तविक समय में अपने परिणामों को देखने में सक्षम बनाते हों,
- वैज्ञानिकों के बीच वर्कफ़्लो को साझा करने और पुनः उपयोग करने की प्रक्रिया को सरल बनाना, और
- वैज्ञानिकों को वर्कफ़्लो निष्पादन परिणामों और वर्कफ़्लो निर्माण चरणों के उद्भव (provenance) को ट्रैक करने में सक्षम बनाना।
यह सेवा प्रदान करने वाले कुछ प्लेटफॉर्म हैं: Galaxy, Kepler, Taverna, UGENE, Anduril, HIVE।
बायोकंप्यूट और बायोकंप्यूट ऑब्जेक्ट्स
2014 में, यूएस फूड एंड ड्रग एडमिनिस्ट्रेशन ने बायोइनफॉरमैटिक्स में पुनरुत्पादकता (reproducibility) पर चर्चा करने के लिए नेशनल इंस्टीट्यूट ऑफ हेल्थ बेथेस्डा कैंपस में आयोजित एक सम्मेलन प्रायोजित किया।[65] अगले तीन वर्षों में, हितधारकों के एक संघ (consortium) ने इस बात पर चर्चा करने के लिए नियमित रूप से मुलाकात की कि आगे चलकर बायोकंप्यूट प्रतिमान (paradigm) क्या बनेगा।[66] इन हितधारकों में सरकार, उद्योग और शैक्षणिक संस्थाओं के प्रतिनिधि शामिल थे। सत्र के नेताओं ने FDA और NIH संस्थानों और केंद्रों की कई शाखाओं, ह्यूमन वैरीओम प्रोजेक्ट और यूरोपीय फेडरेशन फॉर मेडिकल इंफॉर्मेटिक्स सहित गैर-लाभकारी संस्थाओं, और स्टैनफोर्ड, न्यू यॉर्क जीनोम सेंटर, और जॉर्ज वाशिंगटन यूनिवर्सिटी सहित अनुसंधान संस्थानों का प्रतिनिधित्व किया।
यह तय किया गया कि बायोकंप्यूट प्रतिमान डिजिटल 'लैब नोटबुक' के रूप में होगा जो बायोइनफॉरमैटिक्स प्रोटोकॉल की पुनरुत्पादकता (reproducibility), प्रतिकृति (replication), समीक्षा (review) और पुन: उपयोग (reuse) की अनुमति देता है। यह प्रस्ताव इसलिए दिया गया था ताकि सामान्य कार्मिक प्रवाह के दौरान एक शोध समूह के भीतर अधिक निरंतरता सक्षम हो सके और साथ ही समूहों के बीच विचारों के आदान-प्रदान को बढ़ावा मिल सके। अमेरिकी FDA ने इस कार्य को वित्त पोषित किया ताकि पाइपलाइनों पर जानकारी उनके नियामक कर्मचारियों के लिए अधिक पारदर्शी और सुलभ हो सके।[67]
2016 में, समूह ने बेथेस्डा में NIH में फिर से मुलाकात की और एक बायोकंप्यूट ऑब्जेक्ट की संभावना पर चर्चा की, जो बायोकंप्यूट प्रतिमान का एक उदाहरण है। इस कार्य को "मानक परीक्षण उपयोग" (standard trial use) दस्तावेज़ और bioRxiv पर अपलोड किए गए एक प्रीप्रिंट पेपर दोनों के रूप में कॉपी किया गया था। बायोकंप्यूट ऑब्जेक्ट JSON-कृत रिकॉर्ड को कर्मचारियों, सहयोगियों और नियामकों के बीच साझा किए जाने की अनुमति देता है।[68][69]
शिक्षा प्लेटफॉर्म
जबकि कई विश्वविद्यालयों में बायोइनफॉरमैटिक्स को व्यक्तिगत रूप से मास्टर डिग्री के रूप में पढ़ाया जाता है, इस विषय में सीखने और प्रमाणन प्राप्त करने के लिए कई अन्य पद्धतियाँ और प्रौद्योगिकियाँ उपलब्ध हैं। बायोइनफॉरमैटिक्स की संगणकीय प्रकृति इसे कंप्यूटर-सहायता प्राप्त और ऑनलाइन शिक्षण के अनुकूल बनाती है।[70][71] बायोइनफॉरमैटिक्स की अवधारणाओं और तरीकों को सिखाने के लिए डिज़ाइन किए गए सॉफ्टवेयर प्लेटफॉर्म में रोज़ालैंड और स्विस इंस्टीट्यूट ऑफ बायोइनफॉरमैटिक्स ट्रेनिंग पोर्टल के माध्यम से पेश किए जाने वाले ऑनलाइन पाठ्यक्रम शामिल हैं। Canadian Bioinformatics Workshops अपनी वेबसाइट पर क्रिएटिव कॉमन्स लाइसेंस के तहत प्रशिक्षण कार्यशालाओं के वीडियो और स्लाइड प्रदान करता है। 4273π प्रोजेक्ट या 4273pi प्रोजेक्ट[72] मुफ्त में ओपन-सोर्स शैक्षणिक सामग्री भी प्रदान करता है। यह पाठ्यक्रम कम लागत वाले रास्पबेरी पाई कंप्यूटरों पर चलता है और इसका उपयोग वयस्कों और स्कूली छात्रों को पढ़ाने के लिए किया गया है।[73][74] 4273 को शिक्षाविदों और शोध कर्मचारियों के एक संघ द्वारा सक्रिय रूप से विकसित किया गया है जिन्होंने रास्पबेरी पाई कंप्यूटर और 4273π ऑपरेटिंग सिस्टम का उपयोग करके अनुसंधान-स्तरीय बायोइनफॉरमैटिक्स चलाया है।[75][76]
MOOC प्लेटफॉर्म बायोइनफॉरमैटिक्स और संबंधित विषयों में ऑनलाइन प्रमाणन भी प्रदान करते हैं, जिनमें कैलिफोर्निया विश्वविद्यालय, सैन डिएगो की कौरसेरा बायोइनफॉरमैटिक्स स्पेशलाइजेशन, हॉपकिंस विश्वविद्यालय की जीनोमिक डेटा साइंस स्पेशलाइजेशन, और हार्वर्ड विश्वविद्यालय की एडएक्स डेटा एनालिसिस फॉर लाइफ साइंसेज एक्ससीरीज़ शामिल हैं।
सम्मेलन
कई बड़े सम्मेलन हैं जो बायोइनफॉरमैटिक्स से संबंधित हैं। कुछ सबसे उल्लेखनीय उदाहरण European Conference on Computational Biology (ECCB), Intelligent Systems for Molecular Biology (ISMB), पैसिफिक सिंपोजियम ऑन बायोकंप्यूटिंग (PSB), और Research in Computational Molecular Biology (RECOMB) हैं।
यह भी देखें
संदर्भ
- ↑ 1.0 1.1 (2014). "Bioinformatics Curriculum Guidelines: Toward a Definition of Core Competencies". PLOS Comput Biol. 10 (3). doi:10.1371/journal.pcbi.1003496.
- ↑ (26 July 2013). Bioinformatics. Encyclopaedia Britannica.
- ↑ 3.0 3.1 (June 2012). "Modeling nucleic acids". Current Opinion in Structural Biology. 22 (3) 273–8. doi:10.1016/j.sbi.2012.03.012.
- ↑ (July 2016). "Coarse-Grained Protein Models and Their Applications". Chemical Reviews. 116 (14) 7898–936. doi:10.1021/acs.chemrev.6b00163.
- ↑ (2016). "Computational Biology and Bioinformatics: Gene Regulation". CRC Press/Taylor & Francis Group. ISBN 978-1-4987-2497-5.
- ↑ (January 2015). "Structure-based modeling of protein: DNA specificity". Briefings in Functional Genomics. 14 (1) 39–49. doi:10.1093/bfgp/elu044.
- ↑ (2014). Biomolecular Modelling and Simulations. 96 77–111. Academic Press. ISBN 978-0-12-800013-7. doi:10.1016/bs.apcsb.2014.06.008.
- ↑ (August 2018). "Protein-peptide docking: opportunities and challenges". Drug Discovery Today. 23 (8) 1530–1537. doi:10.1016/j.drudis.2018.05.006.
- ↑ (2003). "Early bioinformatics: the birth of a discipline—a personal view". Bioinformatics. 19 (17) 2176–2190. doi:10.1093/bioinformatics/btg309.
- ↑ 10.0 10.1 (2011). "The Roots of Bioinformatics in Theoretical Biology". PLOS Computational Biology. 7 (3). doi:10.1371/journal.pcbi.1002021.
- ↑ (1970). "BIO-INFORMATICA: een werkconcept". Het Kameleon. 1 (6) 28–29.
- ↑ (2021). "Bio-informatics: a working concept. A translation of "Bio-informatica: een werkconcept" by B. Hesper and P. Hogeweg".
- ↑ (1978). "Simulating the growth of cellular forms". Simulation. 31 (3) 90–96. doi:10.1177/003754977803100305.
- ↑ (27 November 2019). "A brief history of bioinformatics". Briefings in Bioinformatics. 20 (6) 1981–1996. doi:10.1093/bib/bby063.
- ↑ (2022). Whole Genome Sequencing Cost. Sequencing.com.
- ↑ (1951). "The Amino-acid Sequence in the Phenylalanyl Chain of Insulin. I. The identification of lower peptides from partial hydrolysates". Biochemical Journal. 49 (4) 463–81. doi:10.1042/bj0490463.
- ↑ (1953). "The Amino-acid Sequence in the Glycyl Chain of Insulin. I. The identification of lower peptides from partial hydrolysates". Biochemical Journal. 53 (3) 353–66. doi:10.1042/bj0530353.
- ↑ (2004). Digital Code of Life: How Bioinformatics is Revolutionizing Science, Medicine, and Business. John Wiley & Sons. ISBN 978-0-471-32788-2.
- ↑ (1965). ATLAS of PROTEIN SEQUENCE and STRUCTURE. National Biomedical Research Foundation.
- ↑ (April 1966). "Evolution of the Structure of Ferredoxin Based on Living Relics of Primitive Amino Acid Sequences". Science. 152 (3720) 363–6. doi:10.1126/science.152.3720.363.
- ↑ (January 2000). "Kabat database and its applications: 30 years after the first variability plot". Nucleic Acids Research. 28 (1) 214–8. doi:10.1093/nar/28.1.214.
- ↑ (1979). "A Search for Patterns in the Nucleotide Sequence of the MS2 Genome". Journal of Mathematical Biology. 7 (3) 219–230. doi:10.1007/BF00275725.
- ↑ (February 1981). "The coding function of nucleotide sequences can be discerned by statistical analysis". Journal of Theoretical Biology. 88 (3) 409–20. doi:10.1016/0022-5193(81)90274-5.
- ↑ (2006). Essential Bioinformatics. 4. Cambridge University Press. ISBN 978-0-511-16815-4.
- ↑ (February 1977). "Nucleotide sequence of bacteriophage phi X174 DNA". Nature. 265 (5596) 687–95. doi:10.1038/265687a0.
- ↑ (January 2008). "GenBank". Nucleic Acids Research. 36 (Database issue) D25-30. doi:10.1093/nar/gkm929.
- ↑ 27.0 27.1 27.2 (July 1995). "Whole-genome random sequencing and assembly of Haemophilus influenzae Rd". Science. 269 (5223) 496–512. doi:10.1126/science.7542800.
- ↑ (2001). "Genome annotation: from sequence to biology". Nature. 2 (7) 493–503. doi:10.1038/35080529.
- ↑ (April 2011). "Protein function prediction: towards integration of similarity metrics". Current Opinion in Structural Biology. 21 (2) 180–8. doi:10.1016/j.sbi.2011.02.001.
- ↑ (March 2010). "Simulation of genes and genomes forward in time". Current Genomics. 11 (1) 58–61. doi:10.2174/138920210790218007.
- ↑ (2002). "Genomes". Oxford.
- ↑ (October 2002). "Comparative analysis of comparative genomic hybridization microarray technologies: report of a workshop sponsored by the Wellcome Trust". Cytometry. 49 (2) 43–8. doi:10.1002/cyto.10153.
- ↑ (April 2016). "BPGA- an ultra-fast pan-genome analysis pipeline". Scientific Reports. 6. doi:10.1038/srep24373.
- ↑ (May 2014). "Genetic susceptibility to male infertility: news from genome-wide association studies". Andrology. 2 (3) 315–21. doi:10.1111/j.2047-2927.2014.00188.x.
- ↑ (2014). "Genome-wide association studies and the clinic: a focus on breast cancer". Biomarkers in Medicine. 8 (2) 287–96. doi:10.2217/bmm.13.121.
- ↑ (October 2013). "Genome-wide association studies in Alzheimer's disease: a review". Current Neurology and Neuroscience Reports. 13 (10). doi:10.1007/s11910-013-0381-0.
- ↑ (2013). "Pharmacogenomics". 1015 127–46. ISBN 978-1-62703-434-0. doi:10.1007/978-1-62703-435-7_8.
- ↑ (June 2009). "Potential etiologic and functional implications of genome-wide association loci for human diseases and traits". Proceedings of the National Academy of Sciences of the United States of America. 106 (23) 9362–7. doi:10.1073/pnas.0903103106.
- ↑ (2010). "2010 International Conference on System Science and Engineering". 1–2. ISBN 978-1-4244-6472-2. doi:10.1109/ICSSE.2010.5551766.
- ↑ (October 2009). "Finding the missing heritability of complex diseases". Nature. 461 (7265) 747–753. doi:10.1038/nature08494.
- ↑ (March 2022). "Assessing the contribution of rare variants to complex trait heritability from whole-genome sequence data". Nature Genetics. 54 (3) 263–273. doi:10.1038/s41588-021-00997-7.
- ↑ (February 2021). "Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program". Nature. 590 (7845) 290–299. doi:10.1038/s41586-021-03205-y.
- ↑ (September 2020). "Dynamic incorporation of multiple in silico functional annotations empowers rare variant association analysis of large whole-genome sequencing studies at scale". Nature Genetics. 52 (9) 969–983. doi:10.1038/s41588-020-0676-4.
- ↑ (December 2022). "A framework for detecting noncoding rare-variant associations of large-scale whole-genome sequencing studies". Nature Methods. 19 (12) 1599–1611. doi:10.1038/s41592-022-01640-x.
- ↑ (December 2022). "STAARpipeline: an all-in-one rare-variant tool for biobank-scale whole-genome sequencing data". Nature Methods. 19 (12) 1532–1533. doi:10.1038/s41592-022-01641-w.
- ↑ (January 2023). "Powerful, scalable and resource-efficient meta-analysis of rare variant associations in large whole genome sequencing studies". Nature Genetics. 55 (1) 154–164. doi:10.1038/s41588-022-01225-6.
- ↑ (2010). "CGHTRIMMER: Discretizing noisy Array CGH Data".
- ↑ (2010). "A statistical approach for detecting genomic aberrations in heterogeneous tumor samples from single nucleotide polymorphism genotyping data". Genome Biology. 11 (9) R92. doi:10.1186/gb-2010-11-9-r92.
- ↑ (2012-12-27). "Chapter 14: Cancer genome analysis". PLOS Computational Biology. 8 (12). doi:10.1371/journal.pcbi.1002824.
- ↑ (2014). "Cancer Genomics". 13–30. Academic Press. ISBN 978-0-12-396967-5. doi:10.1016/B978-0-12-396967-5.00002-5.
- ↑ (July 2006). "VOMBAT: prediction of transcription factor binding sites using variable order Bayesian trees". Nucleic Acids Research. 34 (Web Server issue) W529-33. doi:10.1093/nar/gkl212.
- ↑ The Human Protein Atlas. www.proteinatlas.org.
- ↑ The human cell. www.proteinatlas.org.
- ↑ (May 2017). "A subcellular map of the human proteome". Science. 356 (6340). doi:10.1126/science.aal3321.
- ↑ (September 2015). "Analysis methods for studying the 3D architecture of the genome". Genome Biology. 16 (1). doi:10.1186/s13059-015-0745-7.
- ↑ (2019). "Critical Assessment of Methods of Protein Structure Prediction (CASP) – Round XIII". Proteins. 87 (12) 1011–1020. doi:10.1002/prot.25823.
- ↑ Home - CASP14. predictioncenter.org.
- ↑ (August 2007). "Plant hemoglobins: a molecular fossil record for the evolution of oxygen transport". Journal of Molecular Biology. 371 (1) 168–79. doi:10.1016/j.jmb.2007.05.029.
- ↑ AlphaFold Protein Structure Database. alphafold.ebi.ac.uk.
- ↑ (May 2008). "The binary protein interactome of Treponema pallidum--the syphilis spirochete". PLOS ONE. 3 (5). doi:10.1371/journal.pone.0002292.
- ↑ Groß, Anika; Pruski, Cédric; Rahm, Erhard. (2016-01-01). Evolution of biomedical ontologies and mappings: Overview of recent approaches. Computational and Structural Biotechnology Journal. 14 333–340. doi:10.1016/j.csbj.2016.08.002.
- ↑ 62.0 62.1 Open Bioinformatics Foundation: About us. Official website.
- ↑ Open Bioinformatics Foundation: BOSC. Official website.
- ↑ (2009). Handbook of Statistical Analysis and Data Mining Applications. 328. Academic Press. ISBN 978-0-08-091203-5.
- ↑ Office of the Commissioner. Advancing Regulatory Science – Sept. 24–25, 2014 Public Workshop: Next Generation Sequencing Standards. www.fda.gov.
- ↑ (2017). "Biocompute Objects-A Step towards Evaluation and Validation of Biomedical Scientific Computations". PDA Journal of Pharmaceutical Science and Technology. 71 (2) 136–146. doi:10.5731/pdajpst.2016.006734.
- ↑ Office of the Commissioner. Advancing Regulatory Science – Community-based development of HTS standards for validating data and computation and encouraging interoperability. www.fda.gov.
- ↑ (December 2018). "Enabling precision medicine via standard communication of HTS provenance, analysis, and results". PLOS Biology. 16 (12). doi:10.1371/journal.pbio.3000099.
- ↑ (2017-09-03). BioCompute Object (BCO) project is a collaborative and community-driven framework to standardize HTS computational data. 1. BCO Specification Document: user manual for understanding and creating B.. biocompute-objects.
- ↑ Campbell, A. Malcolm. (2003-06-01). "Public Access for Teaching Genomics, Proteomics, and Bioinformatics". Cell Biology Education. 2 (2) 98–111. doi:10.1187/cbe.03-02-0007.
- ↑ Arenas, Miguel. (September 2021). "General considerations for online teaching practices in bioinformatics in the time of COVID -19". Biochemistry and Molecular Biology Education. 49 (5) 683–684. doi:10.1002/bmb.21558.
- ↑ (August 2013). "4273π: bioinformatics education on low cost ARM hardware". BMC Bioinformatics. 13 522. doi:10.1186/1471-2105-14-243.
- ↑ (2015). "University-level practical activities in bioinformatics benefit voluntary groups of pupils in the last 2 years of school". International Journal of STEM Education. 2 (17). doi:10.1186/s40594-015-0030-z.
- ↑ (2016). "Bringing computational science to the public". SpringerPlus. 5 (259). doi:10.1186/s40064-016-1856-7.
- ↑ (October 2015). "Comparison of the protein-coding gene content of Chlamydia trachomatis and Protochlamydia amoebophila using a Raspberry Pi computer". BMC Research Notes. 8 (561). doi:10.1186/s13104-015-1476-2.
- ↑ (October 2015). "A comparison of the protein-coding genomes of two green sulphur bacteria, Chlorobium tepidum TLS and Pelodictyon phaeoclathratiforme BU-1". BMC Research Notes. 8 (565). doi:10.1186/s13104-015-1535-8.
अग्रिम पठन
35em
- Sehgal et al. : Structural, phylogenetic and docking studies of D-amino acid oxidase activator (DAOA), a candidate schizophrenia gene. Theoretical Biology and Medical Modelling 2013 10 :3.
- Achuthsankar S Nair Computational Biology & Bioinformatics – A gentle Overview সংরক্ষিত সংস্করণ, Communications of Computer Society of India, January 2007
- Aluru, Srinivas, ed. Handbook of Computational Molecular Biology. Chapman & Hall/Crc, 2006. 1-58488-406-1 (Chapman & Hall/Crc Computer and Information Science Series)
- Baldi, P and Brunak, S, Bioinformatics: The Machine Learning Approach, 2nd edition. MIT Press, 2001. 0-262-02506-X
- Barnes, M.R. and Gray, I.C., eds., Bioinformatics for Geneticists, first edition. Wiley, 2003. 0-470-84394-2
- Baxevanis, A.D. and Ouellette, B.F.F., eds., Bioinformatics: A Practical Guide to the Analysis of Genes and Proteins, third edition. Wiley, 2005. 0-471-47878-4
- Baxevanis, A.D., Petsko, G.A., Stein, L.D., and Stormo, G.D., eds., Current Protocols in Bioinformatics. Wiley, 2007. 0-471-25093-7
- Cristianini, N. and Hahn, M. Introduction to Computational Genomics সংরক্ষিত সংস্করণ, Cambridge University Press, 2006. (9780521671910 |0-521-67191-4)
- Durbin, R., S. Eddy, A. Krogh and G. Mitchison, Biological sequence analysis. Cambridge University Press, 1998. 0-521-62971-3
- (September 2004). "Bioinformatics software resources". Briefings in Bioinformatics. 5 (3) 300–4. doi:10.1093/bib/5.3.300.
- Keedwell, E., Intelligent Bioinformatics: The Application of Artificial Intelligence Techniques to Bioinformatics Problems. Wiley, 2005. 0-470-02175-6
- Kohane, et al. Microarrays for an Integrative Genomics. The MIT Press, 2002. 0-262-11271-X
- Lund, O. et al. Immunological Bioinformatics. The MIT Press, 2005. 0-262-12280-4
- Pachter, Lior and Sturmfels, Bernd. "Algebraic Statistics for Computational Biology" Cambridge University Press, 2005. 0-521-85700-7
- Pevzner, Pavel A. Computational Molecular Biology: An Algorithmic Approach The MIT Press, 2000. 0-262-16197-4
- Soinov, L. Bioinformatics and Pattern Recognition Come Together সংরক্ষিত সংস্করণ Journal of Pattern Recognition Research (JPRR সংরক্ষিত সংস্করণ), Vol 1 (1) 2006 p. 37–41
- Stevens, Hallam, Life Out of Sequence: A Data-Driven History of Bioinformatics, Chicago: The University of Chicago Press, 2013, 9780226080208
- Tisdall, James. "Beginning Perl for Bioinformatics" O'Reilly, 2001. 0-596-00080-4
- Catalyzing Inquiry at the Interface of Computing and Biology (2005) CSTB report সংরক্ষিত সংস্করণ
- Calculating the Secrets of Life: Contributions of the Mathematical Sciences and computing to Molecular Biology (1995) সংরক্ষিত সংস্করণ
- Foundations of Computational and Systems Biology MIT Course
- Computational Biology: Genomes, Networks, Evolution Free MIT Course সংরক্ষিত সংস্করণ
बाहरी कड़ियाँ
En-Bioinformatics.ogg bioinformatics
स्रोत: अंग्रेज़ी विकिपीडिया के “Bioinformatics” लेख का अनुवाद। मूल लेख: https://en.wikipedia.org/wiki/Bioinformatics यह अनुवाद स्वचालित रूप से तैयार किया गया है।
