WO2021085702A1 - 유전자 가위 효과를 분석하는 방법 및 장치 - Google Patents
유전자 가위 효과를 분석하는 방법 및 장치 Download PDFInfo
- Publication number
- WO2021085702A1 WO2021085702A1 PCT/KR2019/014785 KR2019014785W WO2021085702A1 WO 2021085702 A1 WO2021085702 A1 WO 2021085702A1 KR 2019014785 W KR2019014785 W KR 2019014785W WO 2021085702 A1 WO2021085702 A1 WO 2021085702A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- sequence
- learning model
- effect
- information
- scissors
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B25/00—ICT specially adapted for hybridisation; ICT specially adapted for gene or protein expression
- G16B25/10—Gene or protein expression profiling; Expression-ratio estimation or normalisation
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
- C12N15/09—Recombinant DNA-technology
- C12N15/10—Processes for the isolation, preparation or purification of DNA or RNA
- C12N15/102—Mutagenizing nucleic acids
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
- C12N15/09—Recombinant DNA-technology
- C12N15/11—DNA or RNA fragments; Modified forms thereof; Non-coding nucleic acids having a biological activity
- C12N15/113—Non-coding nucleic acids modulating the expression of genes, e.g. antisense oligonucleotides; Antisense DNA or RNA; Triplex- forming oligonucleotides; Catalytic nucleic acids, e.g. ribozymes; Nucleic acids used in co-suppression or gene silencing
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N9/00—Enzymes; Proenzymes; Compositions thereof; Processes for preparing, activating, inhibiting, separating or purifying enzymes
- C12N9/14—Hydrolases (3)
- C12N9/16—Hydrolases (3) acting on ester bonds (3.1)
- C12N9/22—Ribonucleases [RNase]; Deoxyribonucleases [DNase]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B15/00—ICT specially adapted for analysing two-dimensional [2D] or three-dimensional [3D] molecular structures, e.g. structural or functional relations or structure alignment
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
- G16B20/20—Allele or variant detection, e.g. single nucleotide polymorphism [SNP] detection
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B35/00—ICT specially adapted for in silico combinatorial libraries of nucleic acids, proteins or peptides
- G16B35/10—Design of libraries
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B35/00—ICT specially adapted for in silico combinatorial libraries of nucleic acids, proteins or peptides
- G16B35/20—Screening of libraries
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B40/00—ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
- G16B40/20—Supervised data analysis
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N2310/00—Structure or type of the nucleic acid
- C12N2310/10—Type of nucleic acid
- C12N2310/20—Type of nucleic acid involving clustered regularly interspaced short palindromic repeats [CRISPR]
Definitions
- the technology to be described below relates to a technique for analyzing the effect of the shearing gene based on the sequence of the shearing gene.
- Programmable nucleases are artificial enzymes that cut off genetically trait sites in cells and individuals. Genetic scissors also refers to a gene editing technique that removes a target gene with an artificial enzyme.
- the research goal in the field of genetic scissors is to increase on-target efficiency and at the same time reduce off-target activity on other untargeted sequences.
- Genetic shears can also work when the sequence of the guide RNA and the target sequence are not complementary. In this case, the genetic scissors may cause DNA mutation by conjugating to a sequence other than the target sequence.
- the technique described below attempts to detect effective genetic scissors using a learning model.
- the method of analyzing the effect of the scissors by using a learning model is the step of receiving sequence data of the scissors by an analysis device, and the structure of the structure constituted by the target specific sequence of the scissors by the analysis device using the sequence data. Generating information, the analysis device inputting the structural information into the learning model, and the analysis device evaluating the effect of the genetic scissors based on the information output from the learning model.
- the analysis device for predicting the effect of the genetic scissors includes an input device that receives sequence data of the genetic scissors, a program that generates structural information about the structure of the sequence based on the nucleotide sequence, and the structural information of the genetic scissors.
- a program that generates structural information about the structure of the sequence based on the nucleotide sequence
- the structural information of the genetic scissors Using the program and a storage device for storing a learning model that analyzes the effect, structural information on the target specific sequence of the genetic scissors is generated from the sequence data, and the generated structural information is input to the learning model to It includes a computing device that evaluates the effect.
- the system for predicting the effect of the genetic scissors stores a structure prediction server that generates structural information about the structure of the sequence based on the nucleotide sequence and a learning model that analyzes the effect of the genetic scissors based on the structural information of the genetic scissors. And an evaluation server for receiving structural information on a target specific sequence of the genetic scissors from the sequence data of the genetic scissors from the structure prediction server, and inputting the structural information into the learning model to evaluate the effect of the genetic scissors.
- the technique described below quickly predicts whether or not it is effective for a specific target sequence based on the sequence data of the genetic scissors.
- the technique described below accurately predicts the sequence of the genetic scissors effective for the target sequence using a machine learning model trained in advance.
- 5 is another example of a process for predicting the effect of the shearing gene.
- 6 is an example of an artificial neural network.
- 9 is an example of an analysis device for predicting the effect of the scissors gene.
- 11 is another example of the evaluation of the model performance predicting the effect of the genetic scissors.
- Terms such as 1st, 2nd, A, B, etc. may be used to describe various components, but the components are not limited by the above terms, and only for the purpose of distinguishing one component from other components. Is only used.
- a first component may be referred to as a second component, and similarly, a second component may be referred to as a first component without departing from the scope of the rights of the technology described below.
- the term and/or includes a combination of a plurality of related listed items or any of a plurality of related listed items.
- each of the processes constituting the method may occur differently from the specified order unless a specific order is clearly stated in the context. That is, each of the processes may occur in the same order as the specified order, may be performed substantially simultaneously, or may be performed in the reverse order.
- Subjects include cells, tissues or organisms. Entity basically means including humans, animals, plants, microorganisms, and the like.
- Genetic shears are artificial enzymes that cut sites of genetic properties in cells and individuals. Genetic scissors also refers to a gene editing technique that removes a target gene with an artificial enzyme. For example, it is possible to treat a disease by removing damaged DNA using genetic scissors and inserting normal DNA.
- CRISPR Zinc Finger Nuclease
- ZFN Zinc Finger Nuclease
- TALENs Transcriptor Activator-Like Effector Nucleases
- CRISPR corresponds to the third generation of genetic scissors, and is currently the most attention-grabbing genetic scissors. Hereinafter, for convenience of explanation, it will be described based on CRISPR.
- the target gene or the target sequence means a sequence to be edited by the genetic scissors.
- CRISPR gene scissors are composed of RNA (Crisper RNA) that has a complementary base of DNA that is subject to gene editing and an enzyme that removes the target gene.
- CRISPR Cas9 is composed of CRISPR, which plays a role in finding a target gene, and Cas9, a protein that removes a target gene.
- gene scissors using proteins other than Cas 9 eg, Cpf1 are also being studied.
- CRISPR Cas9 which is most studied.
- CRISPR Cas9 consists of a guide RNA and Cas9 to find the target sequence.
- Guide RNA is composed of sgRNA (single guide RNA).
- Guide RNA plays a key role in finding the target sequence.
- Guide RNA can be synthesized by the researcher to have a specific sequence. There are various techniques for synthesizing and screening guide RNAs.
- the technique described below is for detecting gene scissors that have good effect on the target sequence and less side effects on sequences other than the target sequence.
- Machine learning is a field of artificial intelligence, which refers to the field in which algorithms are developed so that computers can learn.
- a learning model is a model that is trained to identify and recognize a specific type of pattern for a data set, and the result is a file format that stores program code.
- learning models such as artificial neural networks and decision trees, depending on the approach method.
- the technique described below detects a genetic scissors effective for a target sequence based on the sequence of the genetic scissors.
- the researcher verifies the effect on the candidate genetic scissors using the technique described below.
- researchers can design effective genetic scissors based on the techniques described below.
- CRISPR Cas9 is an example of the structure of CRISPR Cas9.
- the CRISPR gene shear has a sequence that is complementary to the target sequence.
- the constant region contains a loop structure, which is the site to which the Cas 9 protein binds.
- Sequences specific to the target sequence are sometimes referred to as crRNA (CRISPR RNA).
- the site that binds to Cas 9 is sometimes referred to as tracrRNA.
- Guide RNA can be said to be a complex composed of crRNA and tracrRNA. Part of the crRNA and part of the tracrRNA combine to form a loop structure.
- the target-specific sequence includes a 20 nt (nucleotide) long spacer.
- the conventional research method predicted the effect of gene scissors by measuring the binding force to the target while changing nt of the sequence of the spacer region one by one.
- the technique to be described below aims to detect an effective genetic scissor structure for the target sequence, centering on the crRNA specific for the target sequence.
- the technology described below is intended to detect a target sequence-specific and designable sequence.
- the designable sequence may be 20 nt in sgRNA. Alternatively, the designable sequence may be the entire crRNA.
- Target specific sequences include sequences that directly bind to a target.
- the target-specific sequence may include a sequence that indirectly affects binding of the target sequence.
- the target-specific sequence may include all or part of the crRNA.
- the analysis device is the subject of predicting the effect of the genetic scissors.
- the analysis device analyzes whether the corresponding genetic scissors are effective for a specific target sequence based on the sequence of the genetic scissors.
- the analysis device is illustrated in the form of analysis servers 110 and 210 and a computer terminal 300.
- 2A is an example of a system 100 including an analysis server 110 and a structure prediction server 120.
- the user terminal 10 transmits the genetic scissors sequence to the analysis server 110.
- Genetic scissor sequences are a form of digital data.
- the analysis server 110 transmits the genetic scissors sequence to the structure prediction server 120.
- the structure prediction server 120 predicts a secondary structure or a tertiary structure constituted by the sequence based on the nucleotide sequence.
- the structure prediction server 120 may predict the secondary structure or the tertiary structure of the target-specific sequence among the gene scissor sequences.
- the structure prediction server 120 may predict a secondary structure or a tertiary structure for the received sequence using such a commercial program or its own program.
- the structure prediction server 120 transmits information on the secondary structure or the tertiary structure to the analysis server 110.
- Structure information is a secondary structure image of RNA, a tertiary structure image, a matrix expressing the secondary structure, a matrix expressing the tertiary structure, a distance map generated based on the second/third structure, and the second order. It may be any one of various forms such as a contact map generated based on a third order structure.
- Structural information corresponds to input data input to the learning model.
- the analysis server 110 holds a previously trained learning model.
- the learning model is a model that outputs information on the effect of genetic scissors based on structural information.
- the analysis server 110 inputs the received structure information into the learning model.
- the analysis server 110 predicts or verifies the effect of the genetic scissors based on information output from the learning model.
- the analysis server 110 may transmit the analysis result to the user terminal 10.
- the analysis server 110 may generate a distance map, an energy value, and the like by processing information on a secondary structure or a tertiary structure for a target-specific sequence received from the structure prediction server 120.
- FIG. 2B is an example of a system 200 including an analysis server 210.
- the user terminal 20 transmits the genetic scissors sequence to the analysis server 210.
- Genetic scissor sequences are a form of digital data.
- the analysis server 210 predicts the effect of the genetic scissors.
- Analysis server 210 holds a program or model for predicting a secondary structure or a tertiary structure constituted by the sequence based on the nucleotide sequence.
- the analysis server 210 predicts the secondary structure or the tertiary structure of the target-specific sequence among the gene scissor sequences.
- the analysis server 210 generates the above-described structure information.
- the analysis server 210 holds a previously trained learning model.
- the analysis server 210 inputs the received structure information into the learning model.
- the analysis server 210 predicts or verifies the effect of the genetic scissors based on information output from the learning model.
- the analysis server 210 may transmit the analysis result to the user terminal 20.
- the 2(C) is an example of an analysis device in the form of a computer terminal 310.
- the computer terminal 310 holds a program or model for predicting a secondary structure or a tertiary structure constituted by the sequence based on the nucleotide sequence.
- the computer terminal 310 holds a learning model that outputs information on the effect of the genetic scissors based on the structure information.
- the computer terminal 310 predicts the secondary structure or the tertiary structure of the target-specific sequence among the genetic scissors sequence based on the genetic scissors sequence input from the user 30.
- the computer terminal 310 generates structure information including information about the predicted secondary structure or tertiary structure or related information.
- the computer terminal 310 inputs the generated structure information into the learning model.
- the computer terminal 310 predicts or verifies the effect of the genetic scissors based on the information output from the learning model.
- the user 30 can check the analysis result.
- 3(A) corresponds to the secondary structure image of RNA. This is an example where the input data (structure information) is a general video. 3(A) shows an example of a secondary structure composed of a relatively large number of sequences. In the case of genetic scissors, the secondary structure constituted by the guide RNA or the target-specific sequence may be simpler than that of FIG. 3(A).
- 3(B) is an example of a distance map generated based on a secondary structure constituted by an RNA sequence.
- the distance map is a two-dimensional matrix with a horizontal axis and a vertical axis.
- the horizontal axis and the vertical axis are the nucleotide sequences that are labeled according to a certain order.
- the distance map represents the distance from one nucleotide to another nucleotide in the secondary structure.
- 3(B) is an example in which the distance is displayed in a certain color. For example, a brighter color may indicate a closer distance.
- the input data is a 2D distance map image. Furthermore, since the elements of the horizontal axis and the vertical axis are the same, the distance map of FIG. 3B is symmetrical with respect to the diagonal. Accordingly, input data may use only a triangular area excluding an overlapping area.
- 3(A) and 3(B) are examples of the above-described structure information.
- the learning model may use other types of input data in addition to structure information.
- Other types of input data are also nucleotide sequence specific information. For example, a researcher can use a commercial program to calculate the energy required to break a double bond from a nucleotide sequence.
- the input data can be a nucleotide sequence specific energy value.
- energy values (DNA-DNA, DNA-RNA hybrid duplexes energy), which are one of the gene scissors information and one of the variables that can be used in the learning model, are expressed.
- Energy value I is the value required to break the DNA strand, that is, the DNA-DNA bond
- the energy value I is the energy required to separate the RNA strand after the bond is formed in the loosened DNA part.
- binding energy the energy required to release the double bond
- the input data may be in the form of a vector encoding a nucleotide sequence.
- 3(D) is an example of a one-hot encoding procedure for a portion of the genetic scissors sequence.
- the result data of one-hot encoding is called a one-hot vector. If the nucleotide in the sequence is composed of A, C, G, and U, the corresponding position per nucleotide is expressed as 1 and the rest are expressed as 0.
- the learning model can use various input data types as shown in FIG. 3.
- the analysis device 110, 220, or 310 may predict the effect of the specific gene scissors through the prediction process 400 described in FIG. 4.
- the analysis device receives the genetic scissors sequence data (410).
- Sequence data is typically in the form of digital data.
- the sequence data may be any one of a sequence of the entire genetic scissors, a sequence of a guide RNA (crRNA+tracrRNA), a partial sequence of a sequence of a guide RNA, crRNA+tracrRNA, and crRNA.
- the analysis device may generate structural information based on all of the received scissors sequence data or part of the scissors sequence data (420). For example, the analysis device may generate structural information based on a target-specific sequence.
- An example of using a distance map is shown in the lower part of FIG. 4.
- the distance map of FIG. 4 corresponds to image data representing the secondary structure of RNA.
- the analysis device inputs structural information into the learning model and analyzes the effect on the input sequence data (specific gene scissors) (430).
- the lower part of FIG. 4 shows a learning model in the form of an artificial neural network.
- the analysis device determines the effect of the genetic scissors to be analyzed based on the result output from the learning model.
- the learning model may be a binary classification model that performs only two classifications, such as a value of 0 or 1. In this case, the learning model outputs information as to whether the input genetic scissors sequence is effective (1) or not (0) for a specific target sequence. Furthermore, the learning model can be classified into any one of multiple classes. In this case, the learning model may classify whether or not the input scissor sequence is effective for a specific target sequence into one of an effective class, a candidate class, and an ineffective class.
- 5 is another example of a process 500 for predicting the effect of the shearing gene.
- 5 is an example of a process of predicting the effect of the genetic scissors by using additional information in addition to structural information.
- the additional information may include at least one of a sequence-specific binding energy value and a one-hot vector.
- the analysis device receives the genetic scissors sequence data (510).
- Sequence data is typically in the form of digital data.
- the sequence data may be any one of a sequence of the entire genetic scissors, a sequence of a guide RNA (crRNA+tracrRNA), a partial sequence of a sequence of a guide RNA, crRNA+tracrRNA, and crRNA.
- the analysis device may generate structural information and additional information based on all of the received scissors sequence data or part of the scissors sequence data (520).
- the analysis device may generate structural information based on a target-specific sequence.
- 5 shows an example of using a distance map as structure information.
- the analysis device predicts the secondary structure of the target-specific sequence using a tool for predicting the secondary structure (521).
- the analysis device generates a distance map based on the secondary structure (522).
- the analysis device may generate additional information based on the target-specific sequence.
- the analysis device may calculate the binding energy value of the target-specific sequence using the avidity prediction tool (523).
- the analysis device may generate a one-hot vector by one-hot encoding the target-specific sequence (524).
- the analysis device inputs structural information and additional information into the learning model and analyzes the effect on the input sequence data (specific gene scissors) (530).
- the analysis device determines the effect of the genetic scissors to be analyzed based on the result output from the learning model.
- the learning model may be a binary classification model that performs only two classifications, such as a value of 0 or 1. In this case, the learning model outputs information as to whether the input genetic scissors sequence is effective (1) or not (0) for a specific target sequence. Furthermore, the learning model can be classified into any one of multiple classes. In this case, the learning model may classify whether or not the input scissor sequence is effective for a specific target sequence into one of an effective class, a candidate class, and an ineffective class.
- CNN convolutional neural networks
- RNN recurrent neural networks
- FFNN feedforward neural networks
- CNN convolutional neural networks
- the CNN model will be mainly described.
- the CNN model is mainly used in the field of computer vision. Recently, the CNN model can receive and process not only images, but also natural language processing, a matrix composed of vectors, and the like.
- 6 is an example of an artificial neural network that is a learning model. 6 is an example of a CNN model. 6 corresponds to a model that receives structural information (image information) specific to a target sequence.
- the CNN includes a convolution layer (Conv), a pooling layer (Pool), and a fully connected layer. In addition, multiple layers may be repeatedly arranged.
- the upper CNN of FIG. 6 may have a structure of 5 convolutional layers, 2 pooling layers, and 2 fully connected layers.
- the convolution layer outputs a feature map through a convolution operation on an input image.
- a filter that performs a convolution operation is also called a kernel.
- the size of the filter is called the filter size or kernel size.
- An operation parameter constituting the kernel is called a kernel parameter, a filter parameter, or a weight.
- the convolutional layer performs convolution and nonlinear operations.
- the convolution operation is performed on a window of a certain size.
- the window can be moved one by one from the upper left to the lower right of the image, and the size of the movement can be adjusted at a time.
- the size of the movement is called a stride.
- the convolutional layer performs a convolution operation on all areas of the input image while moving the window in the input image.
- the convolution layer can maintain the dimension of the input image after the convolution operation by padding the edge of the image.
- the nonlinear operation layer is a layer that determines output values from neurons (nodes).
- the nonlinear operation layer uses a transfer function. Transfer functions include Relu and sigmoid functions.
- the pooling layer sub-samples the feature map obtained as a result of the operation in the convolutional layer.
- Pooling operations include max pooling and average pooling. Maximum pooling selects the largest sample value within the window. Average pooling is sampled as the average value of the values included in the window.
- the all-connected layer finally classifies the input image.
- the all-connected layer receives all the values output from the previous convolutional layer and performs a final classification.
- the all-connected layer outputs a classification result using a softmax function.
- FIG. 7 is another example of an artificial neural network that is a learning model. 7 corresponds to a learning model that receives and analyzes structural information and additional information specific to a target sequence.
- the learning model receives and processes a distance map (A), a binding energy (B), and a one-hot vector (C).
- the convolutional neural network eg, CNN
- CNN may be composed of a convolutional layer, a pooling layer, and a full connection layer.
- the distance map is input to the initial convolutional layer of the CNN.
- the distance map is converted into a feature map through a convolutional layer and a pooling layer.
- the feature map is input to the entire connection layer.
- the binding energy may be input to the entire connection layer.
- a one-hot vector may also be input to the full connection layer.
- the all-connected layer can be finally classified based on the feature map, the binding energy, and the one-hot vector.
- the all-connected layer may convert a two-dimensional feature map into a one-dimensional matrix and then perform computational processing, convert the combined energy and one-hot vector information into a one-dimensional matrix, and then connect behind the feature map to perform final classification. That is, the all-connected layer may convert the feature map and additional information (at least one of a combination energy and a one-hot vector) into one one-dimensional matrix, and perform final classification based on the converted one-dimensional matrix.
- 8 is an example of a process of training a learning model. 8 shows the configuration of a system 600 for training a learning model.
- the training data is a data set consisting of ⁇ input data, classification values for input data ⁇ .
- x i is the i-th input data
- y i is the classification value for the input data x i.
- the computer terminal 50 receives or generates sequence data of the genetic scissors.
- the sequence data is expressed as s.
- s i is the i-th sequence data.
- the structure prediction apparatus 610 is a device similar to the structure prediction apparatus 120 of FIG. 2.
- the structure prediction device 610 predicts the secondary structure or the tertiary structure of the corresponding nucleotide sequence based on the input nucleotide sequence.
- the structure prediction apparatus 610 may generate an image corresponding to a secondary structure or a tertiary structure as structure information.
- the structure prediction apparatus 610 may generate related information as structure information based on a secondary structure or a tertiary structure.
- the structure prediction apparatus 610 may generate structure information in a form such as an image including a structure, a distance map, and a contact map.
- the structure prediction apparatus 610 may generate a sequence-specific binding energy, a one-hot vector, and the like (additional information).
- Structure prediction unit 610 outputs x i receives the i s.
- the structure prediction apparatus 610 may store x i together with an identifier for i in a database (DB, 630).
- the scoring device 620 calculates the effect of the scissor on a specific target as a specific score based on the scissor sequence.
- the score representing the genetic shear effect is called an effect score.
- the effect score can be calculated using a known solution. For example, a solution such as GenomeCRISPR_full05112017 (https://www.dkfz.de/signaling/crispr-downloads/GENOMECRISPR/) can be used.
- the effect score is a quantitative evaluation of the effect of a specific gene clip on a specific target.
- Score calculator 620 outputs y i receives the i s.
- the score calculating device 620 may store y i together with an identifier for i in the database (DB, 630).
- the structure prediction device 610 and the score calculation device 620 are computing devices capable of processing certain data.
- the structure prediction device 610 and the score calculation device 620 may be devices such as a server, a PC, and a smart device.
- the database 630 stores training data sets for n sequences.
- the training device 640 trains a learning model using training data.
- the training device 640 iteratively trains the learning model so that the learning model outputs a value y i with respect to the input data x i.
- the learning process corresponds to the process of optimizing the weights used in the CNN model. For example, weight optimization may use a gradient descent method.
- the learning model When training is properly completed, the learning model outputs the value y i for the structure information x i.
- the effect score y i may be either 0 or 1.
- the effect score y i may be an integer or a real value within a certain range.
- FIG. 8 separately shows an apparatus 610 for predicting a structure based on the scissors sequence and an apparatus 620 for calculating an effect score based on the scissors sequence.
- one device such as the user terminal 50, may generate structural information based on the genetic scissors sequence and generate training data by calculating an effect score for the corresponding genetic scissors sequence.
- the user terminal 50 uses a web solution (eg, RNAfoldWebServer, http://rna.tbi.univie.ac.at/cgi-bin/RNAWebSuite/RNAfold.cgi) that predicts the RNA structure, or an installed program. Can be used to predict the secondary structure / tertiary structure for a gene scissor sequence or a target-specific sequence.
- the user terminal 50 may generate structure information such as a distance map based on the predicted structure.
- the user terminal 50 may calculate the effect score using a solution such as GenomeCRISPR_full05112017.
- FIG 8 shows the user terminal 50 and the training device 640 as separate objects.
- the object for preparing the training data and the object for training the learning model may be the same.
- the learning model evaluates the effect or suitability of the scissor sequence for a specific target sequence. Therefore, it is preferable that the learning model is learned in advance with a different model for each target sequence.
- the analysis device 700 is a device corresponding to the analysis device 110, 210, or 310 of FIG. 2.
- the analysis device 700 predicts the effect of the genetic scissors having a specific sequence using the above-described learning model.
- the analysis device 700 may be physically implemented in various forms.
- the analysis device 700 may have a form such as a computer device such as a PC, a server of a network, or a chipset dedicated to image processing.
- the computer device may include a mobile device such as a smart device.
- the analysis device 700 includes a storage device 710, a memory 720, an operation device 730, an interface device 740, a communication device 750, and an output device 760.
- the storage device 710 stores a neural network model that predicts the effect of the genetic scissors.
- the neural network model must be trained in advance.
- the storage device 710 may store a program for predicting a secondary structure or a tertiary structure based on an RNA sequence.
- the storage device 710 may store a program for generating structure information based on a secondary structure or a tertiary structure.
- the storage device 710 may store other programs or source codes required for data processing.
- the storage device 710 may store the inputted gene scissors sequence and analysis results (gene scissors effect).
- the memory 720 may store data and information generated in the process of analyzing the data received by the analysis device 700.
- the interface device 740 is a device that receives certain commands and data from the outside.
- the interface device 740 may receive a genetic scissors sequence or a guide RNA sequence from an input device physically connected or an external storage device.
- the interface device 740 may receive a learning model for data analysis.
- the interface device 740 may receive training data, information, and parameter values for training a learning model.
- the communication device 750 refers to a configuration for receiving and transmitting certain information through a wired or wireless network.
- the communication device 750 may receive a genetic scissors sequence or a guide RNA sequence from an external object.
- the communication device 750 may also receive data for model training.
- the communication device 750 may transmit an analysis result determined for the input sequence to an external object.
- the communication device 750 to the interface device 740 are devices that receive certain data or commands from the outside.
- the communication device 750 to the interface device 740 may be referred to as an input device.
- the input device may input or receive structural information on the genetic scissors to be analyzed.
- the input device may receive structure information from an external server or DB.
- the input device may receive secondary structure or tertiary structure information about the genetic scissors to be analyzed.
- the output device 760 is a device that outputs certain information.
- the output device 760 may output an interface required for a data processing process, an analysis result, and the like.
- the computing device 730 may predict a secondary structure or a tertiary structure for the genetic scissors sequence using a program stored in the storage device 710.
- the computing device 730 may generate an image of an RNA secondary structure or a tertiary structure.
- the computing device 730 may generate structure information such as a distance map based on the RNA secondary structure or the tertiary structure. That is, the computing device 730 may generate input data to which the learning model can be input.
- the computing device 730 may pre-process data if necessary.
- the computing device 730 may input the generated input data (structural information) into the learning model to predict an effect on the gene scissor sequence to be analyzed.
- the computing device 730 may predict an effect on the gene scissor sequence based on a result of binary classification or multiple classification for the effect.
- the computing device 730 may train a learning model used to predict the effect of the genetic scissors sequence by using the given training data.
- the computing device 730 may be a device such as a processor, an AP, or a chip in which a program is embedded that processes data and processes certain operations.
- Training data is GenomeCRISPR_full05112017 data (hereinafter, data set).
- the data set consisted of 38 million genetic scissor sequences.
- the scissor sequences in the data set are not unique scissor sequences, but genetic scissor sequences that have been duplicated in multiple cell lines.
- One gene shear has different effects from cell line to cell line.
- the data set was prepared as a data set that is easy to learn through the following statistical processing.
- the investigator excluded sequences tested in less than 30 cell lines from the initial data set. This is the filtering of some data to predict the presence or absence of truncation while maintaining the uniqueness of the genetic scissors sequence.
- the researcher defines a gene scissor sequence with an effect score of 5 or more in 80% or more cell lines as good (cut off), and a gene scissor sequence with an effect score of -5 or less in 80% or more cell lines is less effective (not truncated). It was defined as.
- the filtered data set consisted of 25,593 gene scissor sequences.
- the filtered data set consisted of 12,481 data labeled with good effect and 13,112 data labeled with low effect.
- the learning model implemented the model as shown in FIG. 7 described above.
- the learning model is a model that predicts effects according to binary classification.
- the learning model implemented four learning models with different input data. The four learning models are shown in Table 1 below.
- Model classification Input data 1st model 2D street map 2nd model (i) two-dimensional distance map, (ii) binding energy 3rd model (i) 2D distance map, (ii) One-hot vector 4th model (i) two-dimensional distance map, (ii) binding energy, (iii) one-hot vector
- All four learning models performed k-fold cross validation. k was set to 5. The training data was randomly mixed and divided into 5 pieces. Of the five, four were used for learning, and the other one was used for verification data.
- the data for verification was changed to data different from that of the previous fold, and the training data was used except for one data selected as the data for verification.
- the learning model was trained using different combinations of training data, and verified using different data. Furthermore, the 5-fold cross-validation was repeated 30 times.
- the performance evaluation measures of the learning model used the accuracy and the area under the curve (AUC) of ROC (Receiver Operating Characteristics). Performance was analyzed based on the distribution of the average performance of 5 data sets generated by 5-fold cross-validation per experiment iterations.
- the x-axis is the false positive rate (FPR) of the classification model
- the y-axis is the true positive rate (TPR).
- FPR false positive rate
- TPR true positive rate
- AUC means the probability that the predicted value for the positive object of the classifier is higher than the predicted value for the negative object.
- AUC can be estimated by finding the area under the ROC curve.
- Accuracy compares the model predicted value with the data y value, and indicates the rate at which the classification matches for all data.
- Accuracy (ACC) can be expressed by the following equation. TP is the number of true positives, TN is the number of true negatives, FP is the number of false positives, and FN is the number of false negatives.
- 10 is an example of the model performance evaluation predicting the effect of the genetic scissors.
- 10 shows the accuracy per epoch for the learning model.
- Fig. 10 shows the average of the results of performing the 5-fold cross-validation 30 times (overall training is 150 times).
- Fig. 10(A) is a result showing the accuracy per epoch using the training data
- Fig. 10(B) is a result showing the accuracy per epoch using the verification data.
- the fourth model distance map + binding energy + one-hot vector
- the fourth model has the highest accuracy among the four learning models.
- the order of accuracy is: the fourth model> the second model (distance map + combined energy)> the third model (distance map + one-hot vector)> the first model (distance map).
- the learning model showed relatively high performance when energy was used together with a distance map as input data.
- 11 is another example of the evaluation of the model performance predicting the effect of the genetic scissors.
- 11 shows a distribution diagram obtained by repeatedly obtaining the average AUC value of the 5-fold 30 times as a box plot.
- 11(A) is a result of analysis using training data.
- 11(B) is a result of analysis using verification data. Looking at the median value of AUC based on the verification data, the learning model including energy is about 0.93, and the learning model without energy is about 0.922. Therefore, it can be said that the performance of the model including energy is better than the model that does not include energy.
- 12 is another example for evaluating the model performance predicting the effect of the shearing gene.
- 12 is a box plot showing the performance of TPR and TNR criteria for a learning model.
- 12(A) is a result of evaluating the TPR performance based on a threshold value of 0.5
- FIG. 12(B) is a result of evaluating the TNR performance based on a threshold value of 0.5. Referring to FIG. 12, it can be seen that the model including energy as input data has higher performance than the model not including energy.
- the learning model training method and the genetic scissors effect analysis method using the learning model as described above may be implemented as a program (or application) including an executable algorithm that can be executed on a computer.
- the program may be provided by being stored in a non-transitory computer readable medium.
- the non-transitory readable medium refers to a medium that stores data semi-permanently and can be read by a device, rather than a medium that stores data for a short moment, such as a register, cache, and memory.
- a non-transitory readable medium such as a CD, DVD, hard disk, Blu-ray disk, USB, memory card, ROM, or the like.
Landscapes
- Engineering & Computer Science (AREA)
- Life Sciences & Earth Sciences (AREA)
- Health & Medical Sciences (AREA)
- Physics & Mathematics (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Chemical & Material Sciences (AREA)
- Genetics & Genomics (AREA)
- Biotechnology (AREA)
- General Health & Medical Sciences (AREA)
- Molecular Biology (AREA)
- Theoretical Computer Science (AREA)
- Biophysics (AREA)
- Medical Informatics (AREA)
- Biomedical Technology (AREA)
- Spectroscopy & Molecular Physics (AREA)
- General Engineering & Computer Science (AREA)
- Evolutionary Biology (AREA)
- Organic Chemistry (AREA)
- Zoology (AREA)
- Bioinformatics & Computational Biology (AREA)
- Wood Science & Technology (AREA)
- Biochemistry (AREA)
- Library & Information Science (AREA)
- Data Mining & Analysis (AREA)
- Software Systems (AREA)
- Artificial Intelligence (AREA)
- Evolutionary Computation (AREA)
- Microbiology (AREA)
- Mathematical Physics (AREA)
- General Physics & Mathematics (AREA)
- Crystallography & Structural Chemistry (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Computing Systems (AREA)
- Plant Pathology (AREA)
- Analytical Chemistry (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Computational Linguistics (AREA)
- Medicinal Chemistry (AREA)
- Bioethics (AREA)
- Databases & Information Systems (AREA)
Abstract
학습모델을 이용하여 유전자 가위를 설계하는 방법은 분석장치가 유전자 가위의 서열 데이터를 입력받는 단계, 상기 분석장치가 상기 서열 데이터를 이용하여 유전자 가위의 표적 특이적 서열이 구성하는 구조에 대한 구조 정보를 생성하는 단계, 상기 분석장치가 상기 구조 정보를 학습모델에 입력하는 단계 및 상기 분석장치가 상기 학습모델이 출력하는 정보를 기준으로 상기 유전자 가위의 효과를 평가하는 단계를 포함한다.
Description
이하 설명하는 기술은 유전자 가위 서열을 기준으로 유전자 가위 효과를 분석하는 기법에 관한 것이다.
유전자 가위(programmable nuclease)는 세포 및 개체의 유전자 특성 부위를 자르는 인공 효소이다. 유전자 가위는 인공 효소로 표적 유전자를 제거하는 유전자 편집 (genome editing) 기술을 의미하기도 한다.
유전자 가위 분야의 연구 목표는 표적 서열에 대한 효율(on-target efficiency)을 높이고, 동시에 목표하지 않은 다른 서열에 대한 작용(off-target activity)을 줄이는 것이다. 유전자 가위는 가이드 RNA의 서열과 표적 서열이 일부 상호보완이 아닌 경우 경우에도 작용할 수 있다. 이 경우, 유전자 가위가 표적 서열 외에 다른 서열에 접합하여 DNA 돌연변이를 유발할 수 있다.
종래 연구는 유전자 가위에서 표적 서열에 특이적인 서열을 변경해가면서 표적 서열에 효과적인지 테스트를 수행하는 방식을 사용한다. 이와 같은 종래 방법은 유전자 서열 중 하나 이상의 뉴클레오타이드를 변경하면서, 표적 서열에 대한 유전자 가위의 효과를 검증하는 실험을 반복한다. 결국, 종래 접근법은 효과적인 유전자 가위를 탐색하는데 매우 많은 시간과 비용이 소요된다는 한계가 있다.
이하 설명하는 기술은 학습모델을 사용하여 효과적인 유전자 가위를 검출하고자 한다.
학습모델을 이용하여 유전자 가위 효과를 분석하는 방법은 분석장치가 유전자 가위의 서열 데이터를 입력받는 단계, 상기 분석장치가 상기 서열 데이터를 이용하여 유전자 가위의 표적 특이적 서열이 구성하는 구조에 대한 구조 정보를 생성하는 단계, 상기 분석장치가 상기 구조 정보를 학습모델에 입력하는 단계 및 상기 분석장치가 상기 학습모델이 출력하는 정보를 기준으로 상기 유전자 가위의 효과를 평가하는 단계를 포함한다.
유전자 가위 효과를 예측하는 분석장치는 유전자 가위의 서열 데이터를 입력받는 입력장치, 뉴클레오타이드 서열 기준으로 해당 서열이 구성하는 구조에 대한 구조 정보를 생성하는 프로그램 및 유전자 가위의 구조 정보를 기준으로 유전자 가위의 효과를 분석하는 학습모델을 저장하는 저장장치 및 상기 프로그램을 이용하여 상기 서열 데이터에서 유전자 가위의 표적 특이적 서열에 대한 구조 정보를 생성하고, 생성된 구조 정보를 상기 학습모델에 입력하여 유전자 가위의 효과를 평가하는 연산장치를 포함한다.
유전자 가위 효과를 예측하는 시스템은 뉴클레오타이드 서열 기준으로 해당 서열이 구성하는 구조에 대한 구조 정보를 생성하는 구조 예측 서버 및 유전자 가위의 구조 정보를 기준으로 유전자 가위의 효과를 분석하는 학습모델을 저장하고, 상기 구조 예측 서버로부터 유전자 가위의 서열 데이터에서 유전자 가위의 표적 특이적 서열에 대한 구조 정보를 수신하고, 상기 구조 정보를 상기 학습모델에 입력하여 유전자 가위의 효과를 평가하는 평가 서버를 포함한다.
이하 설명하는 기술은 유전자 가위의 서열 데이터를 기준으로 특정한 표적 서열에 효과적인지 빠르게 예측한다. 이하 설명하는 기술은 사전에 훈련된 기계학습모델을 사용하여 표적 서열에 효과적인 유전자 가위의 서열을 정확하게 예측한다.
도 1은 CRISPR Cas9 구조에 대한 예이다.
도 2는 유전자 가위 효과를 예측하는 시스템에 대한 예이다.
도 3은 입력 데이터에 대한 예이다.
도 4는 유전자 가위 효과를 예측하는 과정에 대한 예이다.
도 5는 유전자 가위 효과를 예측하는 과정에 대한 다른 예이다.
도 6은 인공 신경망에 대한 예이다.
도 7은 인공 신경망에 대한 다른 예이다.
도 8은 학습모델을 훈련하는 과정에 대한 예이다.
도 9는 유전자 가위 효과를 예측하는 분석장치에 대한 예이다.
도 10은 유전자 가위 효과 예측한 모델 성능 평가에 대한 예이다.
도 11은 유전자 가위 효과 예측한 모델 성능 평가에 대한 다른 예이다.
도 12는 유전자 가위 효과 예측한 모델 성능 평가에 대한 또 다른 예이다.
이하 설명하는 기술은 다양한 변경을 가할 수 있고 여러 가지 실시례를 가질 수 있는 바, 특정 실시례들을 도면에 예시하고 상세하게 설명하고자 한다. 그러나, 이는 이하 설명하는 기술을 특정한 실시 형태에 대해 한정하려는 것이 아니며, 이하 설명하는 기술의 사상 및 기술 범위에 포함되는 모든 변경, 균등물 내지 대체물을 포함하는 것으로 이해되어야 한다.
제1, 제2, A, B 등의 용어는 다양한 구성요소들을 설명하는데 사용될 수 있지만, 해당 구성요소들은 상기 용어들에 의해 한정되지는 않으며, 단지 하나의 구성요소를 다른 구성요소로부터 구별하는 목적으로만 사용된다. 예를 들어, 이하 설명하는 기술의 권리 범위를 벗어나지 않으면서 제1 구성요소는 제2 구성요소로 명명될 수 있고, 유사하게 제2 구성요소도 제1 구성요소로 명명될 수 있다. 및/또는 이라는 용어는 복수의 관련된 기재된 항목들의 조합 또는 복수의 관련된 기재된 항목들 중의 어느 항목을 포함한다.
본 명세서에서 사용되는 용어에서 단수의 표현은 문맥상 명백하게 다르게 해석되지 않는 한 복수의 표현을 포함하는 것으로 이해되어야 하고, "포함한다" 등의 용어는 설시된 특징, 개수, 단계, 동작, 구성요소, 부분품 또는 이들을 조합한 것이 존재함을 의미하는 것이지, 하나 또는 그 이상의 다른 특징들이나 개수, 단계 동작 구성요소, 부분품 또는 이들을 조합한 것들의 존재 또는 부가 가능성을 배제하지 않는 것으로 이해되어야 한다.
또, 방법 또는 동작 방법을 수행함에 있어서, 상기 방법을 이루는 각 과정들은 문맥상 명백하게 특정 순서를 기재하지 않은 이상 명기된 순서와 다르게 일어날 수 있다. 즉, 각 과정들은 명기된 순서와 동일하게 일어날 수도 있고 실질적으로 동시에 수행될 수도 있으며 반대의 순서대로 수행될 수도 있다.
이하 명세서에서 사용하는 기술 내지 용어에 대하여 먼저 설명한다.
개체(subject)는 세포, 조직 또는 유기체를 포함한다. 개체는 기본적으로 인간, 동물, 식물, 미생물 등을 포함하는 의미이다.
유전자 가위는 세포 및 개체의 유전자 특성 부위를 자르는 인공 효소이다. 유전자 가위는 인공 효소로 표적 유전자를 제거하는 유전자 편집 (genome editing) 기술을 의미하기도 한다. 예컨대, 유전자 가위를 이용하여 손상된 DNA를 제거하고, 정상 DNA를 삽입하여 질병을 치료할 수 있다.
대표적인 유전자 가위는 징크핑거 뉴클레이즈(Zinc Finger Nuclease, ZFN), 탈렌(TALENs·Transcriptor Activator-Like Effector Nucleases), 크리스퍼(CRISPR)가 있다. 크리스퍼는 3세대 유전자 가위에 해당하며, 현재 가장 주목을 받는 유전자 가위이다. 이하 설명의 편의를 위하여 크리스퍼를 기준으로 설명한다.
표적 유전자 내지 표적 서열은 유전자 가위의 편집 대상이 되는 서열을 의미한다.
크리스퍼 유전자 가위는 유전자 편집의 대상이 되는 DNA의 상보적 염기를 지니는 RNA(크리스퍼 RNA) 및 표적 유전자를 제거하는 효소로 구성된다. 크리스퍼 캐스9(CRISPR Cas9)은 표적 유전자를 찾는 역할을 수행하는 크리스퍼와 표적 유전자를 제거하는 단백질인 캐스9으로 구성된다. 한편, 크리스퍼 유전자 가위 관련한 연구가 진행되면서, 캐스 9이 아닌 다른 단백질(예컨대, Cpf1)을 이용하는 유전자 가위도 연구되고 있다. 이하 설명의 편의를 위하여 가장 많이 연구되는 CRISPR Cas9을 기준으로 설명한다.
CRISPR Cas9은 표적 서열은 찾기 위한 가이드 RNA 및 Cas9으로 구성된다. 가이드 RNA는 sgRNA(single guide RNA)로 구성된다. 가이드 RNA가 표적 서열을 찾는데 핵심적 역할을 한다. 가이드 RNA는 연구자가 특정 서열을 갖도록 합성할 수 있다. 가이드 RNA를 합성하는 기술 및 스크리닝하는 기술은 다양하다.
이하 설명하는 기술은 표적 서열에 대한 효과가 좋고, 표적 서열이 아닌 다른 서열에 대한 부작용이 적은 유전자 가위를 검출하기 위한 것이다.
기계 학습(machine learning)은 인공 지능의 한 분야로, 컴퓨터가 학습할 수 있도록 알고리즘을 개발하는 분야를 의미한다. 학습모델은 데이터 집합에 대해 특정 유형의 패턴을 파악하고 인식하도록 학습된 모델이고, 결과물은 프로그램 코드를 저장하는 파일 형태이다. 학습모델은 접근 방법에 따라 인공신경망, 결정 트리 등과 같은 다양한 유형의 모델이 있다.
이하 설명하는 기술은 유전자 가위의 서열을 기준으로 표적 서열에 효과적인 유전자 가위를 검출한다. 연구자는 이하 설명하는 기술을 이용하여 유전자 가위 후보에 대한 효과를 검증한다. 연구자는 이하 설명하는 기술을 기반으로 효과적인 유전자 가위를 설계할 수 있다.
도 1은 CRISPR Cas9 구조에 대한 예이다. 크리스퍼 유전자 가위는 목표 서열에 상보성을 갖는 서열을 갖는다. 고유 영역(constant region)은 루프 구조를 포함하는데, 캐스 9 단백질이 결합하는 부위이다. 목표 서열에 특이적인 서열은 crRNA(CRISPR RNA)라고 명명하기도 한다. 캐스 9과 결합하는 부위를 tracrRNA라고 명명하기도 한다. 가이드 RNA는 crRNA와 tracrRNA로 구성되는 복합체라고 할 수 있다. crRNA 일부와 tracrRNA 일부가 결합하여 루프 구조를 형성하기도 한다. 표적 특이적 서열은 20nt(nucleotide) 길이의 스페이서(spacer)를 포함한다.
종래 연구 방법은 스페이서 영역의 서열 중 nt를 하나씩 변경해가면서 표적에 대한 결합력을 측정하여 유전자 가위의 효과를 예측하였다.
이하 설명하는 기술은 표적 서열에 특이적인 crRNA를 중심으로 표적 서열에 효과적인 유전자 가위 구조를 검출하고자 한다. 이하 설명하는 기술은 표적 서열 특이적이고, 설계 가능한 서열을 검출하고자 한다. 설계 가능한 서열은 sgRNA 중 20nt일 수 있다. 또는 설계 가능한 서열은 crRNA 전체일 수도 있다.
이하 설명하는 기술은 학습모델을 이용하여 sgRNA 전체 또는 일부 서열을 분석한다. 이하 학습모델이 입력받는 서열을 표적 특이적 서열이라고 명명한다. 표적 특이적 서열은 표적에 직접 결합하는 서열을 포함한다. 나아가, 표적 특이적 서열은 표적 서열 결합에 간접적으로 영향을 주는 서열을 포함할 수 있다. 따라서, 표적 특이적 서열은 crRNA 전체 또는 일부를 포함할 수도 있다.
도 2는 유전자 가위 효과를 예측하는 시스템에 대한 예이다. 도 2는 3가지 유형의 시스템 내지 장치를 도시한다. 분석 장치가 유전자 가위의 효과를 예측하는 주체이다. 분석 장치는 유전자 가위의 서열을 기준으로 해당 유전자 가위가 특정 표적 서열에 효과적인지 여부를 분석한다. 도 2에서 분석 장치는 분석 서버(110, 210) 및 컴퓨터 단말(300)의 형태로 도시하였다.
도 2(A)는 분석 서버(110) 및 구조 예측 서버(120)를 포함하는 시스템(100)에 대한 예이다. 사용자 단말(10)은 유전자 가위 서열을 분석 서버(110)에 전달한다. 유전자 가위 서열은 디지털 데이터의 형태이다. 분석 서버(110)는 유전자 가위 서열을 구조 예측 서버(120)에 전달한다. 구조 예측 서버(120)는 뉴클레오타이드 서열을 기준으로 해당 서열이 구성하는 2차 구조 또는 3차 구조를 예측한다. 구조 예측 서버(120)는 유전자 가위 서열 중 표적 특이적 서열의 2차 구조 또는 3차 구조를 예측할 수 있다.
RNA 서열을 기준으로 해당 RNA 서열이 구성하는 2차 구조 또는 3차 구조를 예측하는 상용툴도 있다. 예컨대, RNAdraw, RNAfold, RNAstructure 등이 있다. 따라서, 구조 예측 서버(120)는 이와 같은 상용 프로그램 또는 자체 프로그램을 사용하여 수신하는 서열에 대한 2차 구조 또는 3차 구조를 예측할 수 있다. 구조 예측 서버(120)는 2차 구조 또는 3차 구조에 대한 정보를 분석 서버(110)에 전달한다.
이하, 표적 특이적 서열에 대한 2차/3차 구조에 대한 정보 및 2차/3차 구조를 기준으로 생성된 정보를 모두 구조 정보라고 명명한다. 구조 정보는 RNA의 2차 구조 영상, 3차 구조 영상, 2차 구조를 표현하는 매트릭스, 3차 구조를 표현하는 매트릭스, 2차/3차 구조 기준으로 생성되는 거리 맵(distance map), 2차/3차 구조 기준으로 생성되는 접촉 맵(contact map) 등 다양한 형태 중 어느 하나일 수 있다. 구조 정보는 학습모델에 입력되는 입력 데이터에 해당한다.
분석 서버(110)는 사전에 훈련된 학습모델을 보유한다. 학습모델은 구조 정보를 기준으로 유전자 가위에 대한 효과에 대한 정보를 출력하는 모델이다. 분석 서버(110)는 수신한 구조 정보를 학습모델에 입력한다. 분석 서버(110)는 학습모델이 출력하는 정보를 기준으로 유전자 가위에 대한 효과를 예측 내지 검증한다. 분석 서버(110)는 분석한 결과를 사용자 단말(10)에 전송할 수 있다.
분석 서버(110)는 구조 예측 서버(120)로부터 수신한 표적 특이적 서열에 대한 2차 구조 또는 3차 구조에 대한 정보를 가공하여 거리 맵, 에너지 값 등을 생성할 수도 있다.
도 2(B)는 분석 서버(210)를 포함하는 시스템(200)에 대한 예이다. 사용자 단말(20)은 유전자 가위 서열을 분석 서버(210)에 전달한다. 유전자 가위 서열은 디지털 데이터의 형태이다.
분석 서버(210)는 유전자 가위의 효과를 예측한다. 분석 서버(210)는 뉴클레오타이드 서열을 기준으로 해당 서열이 구성하는 2차 구조 또는 3차 구조를 예측하는 프로그램 내지 모델을 보유한다. 분석 서버(210)는 유전자 가위 서열 중 표적 특이적 서열의 2차 구조 또는 3차 구조를 예측한다. 분석 서버(210)는 전술한 구조 정보를 생성한다.
분석 서버(210)는 사전에 훈련된 학습모델을 보유한다. 분석 서버(210)는 수신한 구조 정보를 학습모델에 입력한다. 분석 서버(210)는 학습모델이 출력하는 정보를 기준으로 유전자 가위에 대한 효과를 예측 내지 검증한다. 분석 서버(210)는 분석한 결과를 사용자 단말(20)에 전송할 수 있다.
도 2(C)는 컴퓨터 단말(310) 형태의 분석 장치에 대한 예이다. 컴퓨터 단말(310)은 뉴클레오타이드 서열을 기준으로 해당 서열이 구성하는 2차 구조 또는 3차 구조를 예측하는 프로그램 내지 모델을 보유한다. 또한, 컴퓨터 단말(310)은 구조 정보를 기준으로 유전자 가위에 대한 효과에 대한 정보를 출력하는 학습모델을 보유한다.
컴퓨터 단말(310)은 사용자(30)로부터 입력받은 유전자 가위 서열을 기준으로 유전자 가위 서열 중 표적 특이적 서열의 2차 구조 또는 3차 구조를 예측한다. 컴퓨터 단말(310)은 예측되는 2차 구조 또는 3차 구조에 대한 정보 내지 연관 정보를 포함하는 구조 정보를 생성한다. 컴퓨터 단말(310)은 생성한 구조 정보를 학습모델에 입력한다. 컴퓨터 단말(310)은 학습모델이 출력하는 정보를 기준으로 유전자 가위에 대한 효과를 예측 내지 검증한다. 사용자(30)는 분석 결과를 확인할 수 있다.
도 3은 입력 데이터에 대한 예이다.
도 3(A)는 RNA의 2차 구조 영상에 해당한다. 입력 데이터(구조 정보)가 일반적인 영상인 예이다. 도 3(A)는 비교적 많은 서열로 구성된 2차 구조를 예로 도시하였다. 유전자 가위 경우, 가이드 RNA 또는 표적 특이적 서열이 구성하는 2차 구조는 도 3(A)보다 단순할 수도 있다.
도 3(B)는 RNA 서열이 구성하는 2차 구조를 기준으로 생성한 거리 맵에 대한 예이다. 거리 맵은 가로축과 세로축을 갖는 2차원 매트릭스이다. 가로축 및 세로축은 일정한 순서에 따라 라벨링되는 뉴클레오타이드 서열이다. 거리 맵은 2차 구조에서 하나의 뉴클레오타이드를 기준으로 다른 뉴클레오타이드와의 거리를 나타낸다. 도 3(B)는 거리를 일정한 색상으로 표시한 예이다. 예컨대, 더 밝은 색일수록 가까운 거리를 나타낼 수 있다.
도 3(B)에서 입력 데이터는 2차원 거리 맵 영상이다. 나아가, 도 3(B)의 거리 맵은 가로축과 세로축의 요소가 동일하기 때문에 대각선을 기준으로 대칭된다. 따라서, 입력 데이터는 중복되는 영역을 제외한 삼각형 영역만을 사용할 수도 있다.
도 3(A) 및 도 3(B)는 전술한 구조 정보에 대한 예이다. 한편, 학습모델은 구조 정보 외에 다른 형태의 입력 데이터를 사용할 수도 있다. 다른 형태의 입력 데이터도 뉴클레오타이드 서열 특이적인 정보이다. 예컨대, 연구자는 상용 프로그램을 이용하여 뉴클레오타이드 서열로부터 이중 결합을 푸는데 소요되는 에너지 값을 연산할 수 있다.
입력 데이터는 뉴클레오타이드 서열 특이적인 에너지 값이 될 수 있다. 도 3(C)에는 유전자 가위 정보 중 하나, 그리고 학습모델에서 사용할 수 있는 변수 중 하나인 에너지 값 (DNA-DNA, DNA-RNA hybrid duplexes energy)을 표현한다. 에너지 값 는 DNA strand, 즉 DNA-DNA 결합을 푸는 것에 필요한 수치이며, 에너지 값 는 풀어진 DNA 부분에 RNA strand와 결합 형성후 결합 분리할 시 필요한 에너지 수치이다. 이하, 이중 결합을 푸는데 소요되는 에너지(hybrid duplexes energy)를 결합 에너지라고 명명한다.
나아가, 입력 데이터는 뉴클레오타이드 서열을 인코딩한 벡터 형태일 수 있다. 도 3(D)는 유전자 가위 서열 일부분을 원-핫 인코딩 절차 예시다. 원-핫 인코딩한 결과 데이터를 원-핫 벡터라고 명명한다. 서열 내 뉴클레오타이드가 A, C, G, U로 구성이 되어있으면 뉴클레오타이드 당 해당하는 위치에 1로 표현하고 나머지는 0으로 표현한다. 학습모델은 도 3과 같은 다양한 입력 데이터 유형을 이용할 수 있다.
도 4는 유전자 가위 효과를 예측하는 과정(400)에 대한 예이다. 분석장치(110, 220 또는 310)는 도 4에서 설명하는 예측 과정(400)을 통해 특정 유전자 가위의 효과를 예측할 수 있다.
도 4는 구조 정보를 이용하여 유전자 가위 효과를 예측하는 과정에 대한 예이다. 분석 장치는 유전자 가위 서열 데이터를 입력받는다(410). 서열 데이터는 통상적으로 디지털 데이터 형태이다. 서열 데이터는 유전자 가위 전체의 서열, 가이드 RNA의 서열(crRNA+tracrRNA), 가이드 RNA의 서열의 일부 서열, crRNA + tracrRNA 및 crRNA 중 어느 하나일 수 있다.
분석 장치는 입력받은 유전자 가위 서열 데이터 전체 또는 유전자 가위 서열 데이터 일부를 기준으로 구조 정보를 생성할 수 있다(420). 예컨대, 분석 장치는 표적 특이적 서열을 기준으로 구조 정보를 생성할 수 있다. 도 4의 하단에서 거리맵을 사용한 예를 도시하였다. 도 4의 거리 맵은 RNA의 2차 구조를 표현하는 영상 데이터에 해당한다.
분석장치는 구조 정보를 학습모델에 입력하여 입력된 서열 데이터(특정 유전자 가위)에 대한 효과를 분석한다(430). 도 4의 하단에는 인공신경망 형태의 학습모델을 도시하였다. 분석장치는 학습모델이 출력하는 결과를 기준으로 분석 대상인 유전자 가위의 효과를 판단한다.
학습모델은 0 또는 1의 값과 같이 2개의 분류만을 하는 이진 분류 모델일 수 있다. 이 경우 학습모델은 입력된 유전자 가위 서열이 특정 표적 서열에 효과적인지(1) 또는 효과가 없는지(0)에 대한 정보를 출력한다. 나아가 학습모델은 다중 클래스 중 어느 하나로 분류할 수 있다. 이 경우 학습모델은 입력된 유전자 가위 서열이 특정 표적 서열에 효과적인지 여부를 효과적인 클래스, 후보 클래스 및 효과없는 클래스 중 어느 하나로 분류할 수도 있다.
도 5는 유전자 가위 효과를 예측하는 과정(500)에 대한 다른 예이다. 도 5는 구조 정보 외에 부가 정보를 활용하여 이용하여 유전자 가위 효과를 예측하는 과정에 대한 예이다. 부가 정보는 서열 특이적인 결합 에너지 값 및 원-핫 벡터 중 적어도 하나를 포함할 수 있다.
분석 장치는 유전자 가위 서열 데이터를 입력받는다(510). 서열 데이터는 통상적으로 디지털 데이터 형태이다. 서열 데이터는 유전자 가위 전체의 서열, 가이드 RNA의 서열(crRNA+tracrRNA), 가이드 RNA의 서열의 일부 서열, crRNA + tracrRNA 및 crRNA 중 어느 하나일 수 있다.
분석 장치는 입력받은 유전자 가위 서열 데이터 전체 또는 유전자 가위 서열 데이터 일부를 기준으로 구조 정보 및 부가 정보를 생성할 수 있다(520).
분석 장치는 표적 특이적 서열을 기준으로 구조 정보를 생성할 수 있다. 도 5는 구조 정보로 거리 맵을 사용하는 예를 도시한다. 분석 장치는 2차 구조를 예측하는 툴(tool)을 사용하여 표적 특이적 서열의 2차 구조를 예측한다(521). 분석 장치는 2차 구조를 기준으로 거리 맵을 생성한다(522).
분석 장치는 표적 특이적 서열을 기준으로 부가 정보를 생성할 수 있다. 분석 장치는 결합력 예측 툴을 사용하여 표적 특이적 서열의 결합 에너지 값을 연산할 수 있다(523). 또한, 분석 장치는 표적 특이적 서열을 원-핫 인코딩하여 원-핫 벡터를 생성할 수도 있다(524).
분석장치는 구조 정보 및 부가 정보를 학습모델에 입력하여 입력된 서열 데이터(특정 유전자 가위)에 대한 효과를 분석한다(530). 분석장치는 학습모델이 출력하는 결과를 기준으로 분석 대상인 유전자 가위의 효과를 판단한다.
학습모델은 0 또는 1의 값과 같이 2개의 분류만을 하는 이진 분류 모델일 수 있다. 이 경우 학습모델은 입력된 유전자 가위 서열이 특정 표적 서열에 효과적인지(1) 또는 효과가 없는지(0)에 대한 정보를 출력한다. 나아가 학습모델은 다중 클래스 중 어느 하나로 분류할 수 있다. 이 경우 학습모델은 입력된 유전자 가위 서열이 특정 표적 서열에 효과적인지 여부를 효과적인 클래스, 후보 클래스 및 효과없는 클래스 중 어느 하나로 분류할 수도 있다.
신경망 모델은 RNN(Recurrent Neural Networks), FFNN(feedforward neural network), CNN(convolutional neural network) 등 다양한 모델이 사용될 수 있다. 이하 설명의 편의를 위하여 CNN 모델을 중심으로 설명한다. CNN 모델은 주로 컴퓨터 비전 분야에 사용된다. 최근 CNN 모델은 영상뿐만 아니라, 자연어처리, 벡터로 구성된 매트릭스 등을 입력받아 처리할 수 있다.
도 6은 학습모델인 인공 신경망에 대한 예이다. 도 6은 CNN 모델에 대한 예이다. 도 6은 표적 서열 특이적인 구조 정보(영상 정보)를 입력받는 모델에 해당한다.
CNN은 컨볼루션 계층 (convolution layer, Conv), 풀링 계층 (pooling layer, Pool) 및 전연결 계층(fully connected layer)을 포함한다. 또한, 계층들은 반복적으로 다수가 배치될 수 있다. 도 6의 상단 CNN은 5개의 컨볼루션 계층, 2개의 풀링 계층, 2개의 전연결 계층(Fully connected layer) 구조를 가질 수 있다.
컨볼루션 계층은 입력 이미지에 대한 컨볼루션 연산을 통해 특징맵(feature map)을 출력한다. 이때 컨볼루션 연산을 수행하는 필터(filter)를 커널(kernel) 이라고도 부른다. 필터의 크기를 필터 크기 또는 커널 크기라고 한다. 커널을 구성하는 연산 파라미터(parameter)를 커널 파라미터(kernel parameter), 필터 파라미터(filter parameter), 또는 가중치(weight)라고 한다.
컨볼루션 계층은 컨볼루션 연산과 비선형 연산을 수행한다.
컨볼루션 연산은 일정한 크기의 윈도우에서 수행된다. 윈도우는 영상의 좌측 상단에서 우측 하단까지 한 칸씩 이동할 수 있고, 한 번에 이동하는 이동 크기를 조절할 수 있다. 이동 크기를 스트라이드(stride)라고 한다. 컨볼루션 계층은 입력이미지에서 윈도우를 이동하면서 입력이미지의 모든 영역에 대하여 컨볼루션 연산을 수행한다. 컨볼루션 계층은 영상의 가장 자리에 패딩(padding)을 하여 컨볼루션 연산 후 입력 영상의 차원을 유지할 수 있다.
비선형 연산 계층(nonlinear operation layer)은 뉴런(노드)에서 출력값을 결정하는 계층이다. 비선형 연산 계층은 전달 함수(transfer function)를 사용한다. 전달 함수는 Relu, sigmoid 함수 등이 있다.
풀링 계층(pooling layer)은 컨볼루션 계층에서의 연산 결과로 얻은 특징맵을 서브 샘플링(sub sampling)한다. 풀링 연산은 최대 풀링(max pooling)과 평균 풀링(average pooling) 등이 있다. 최대 풀링은 윈도우 내에서 가장 큰 샘플 값을 선택한다. 평균 풀링은 윈도우에 포함된 값의 평균 값으로 샘플링한다.
전연결 계층은 최종적으로 입력 영상을 분류한다. 전연결 계층은 이전 컨볼루션 계층에서 출력하는 값을 모두 입력받아 최종적인 분류를 한다. 도 4에서 전연결 계층은 소프트맥스(softmax) 함수를 사용하여 분류 결과를 출력한다.
도 7은 학습모델인 인공신경망에 대한 다른 예이다. 도 7은 표적 서열 특이적인 구조 정보와 부가 정보를 입력받아 분석하는 학습모델에 해당한다. 도 7에서 학습모델은 거리 맵(A), 결합 에너지(B) 및 원-핫 벡터(C)를 입력받아 처리한다. 합성곱 신경망(예컨대, CNN)은 도 6과 유사하게 컨볼루션 계층, 풀링 계층 및 전연결 계층으로 구성될 수 있다.
거리 맵은 CNN의 최초 컨볼루션 계층에 입력된다. 거리 맵은 컨볼루션 계층과 풀링 계층을 통하여 특징맵으로 변환된다. 특징맵은 전연결 계층에 입력된다. 결합 에너지는 전연결 계층에 입력될 수 있다. 또한, 원-핫 벡터도 전연결 계층에 입력될 수 있다.
전연결 계층은 특징맵, 결합 에너지 및 원-핫 벡터를 기준으로 최종적인 분류를 할 수 있다. 예컨대, 전연결 계층은 2차원 특징 맵을 1차원 매트릭스로 변환 후 연산 처리를 하고, 결합 에너지 및 원-핫 벡터 정보를 1차원 매트릭스로 변환 후 특징 맵 뒤에 연결하며 최종적인 분류를 할 수 있다. 즉, 전연결 계층은 특징맵과 부가 정보(결합 에너지 및 원-핫 벡터 중 적어도 하나)를 하나의 1차원 매트릭스로 변환하고, 변환된 1차원 매트릭스 기반하여 최종 분류를 할 수 있다.
도 8은 학습모델을 훈련하는 과정에 대한 예이다. 도 8은 학습모델을 훈련하는 시스템(600) 구성을 도시한다.
훈련데이터를 마련하는 과정을 먼저 설명한다. 훈련데이터는 {입력데이터, 입력데이터에 대한 분류값}으로 구성된 데이터 세트이다. 훈련 데이터를 G라고 하자. 훈련 데이터 세트가 전체 n개인 경우, 훈련 데이터 세트 G =(xi, yi)라고 표현할 수 있다. 이때, 0≤i< n이다. xi는 i 번째 입력데이터이고, yi는 입력데이터 xi에 대한 분류값이다.
컴퓨터 단말(50)은 유전자 가위의 서열 데이터를 입력받거나 생성한다. 서열 데이터를 s라고 표현한다. si는 i 번째 서열 데이터이다.
구조 예측 장치(610)는 도 2의 구조 예측 장치(120)와 유사한 장치이다. 구조 예측 장치(610)는 입력되는 뉴클레오타이드 서열을 기준으로 해당 뉴클레오타이드 서열의 2차 구조 또는 3차 구조를 예측한다. 구조 예측 장치(610)는 2차 구조 또는 3차 구조에 해당하는 영상을 구조 정보로 생성할 수 있다. 또는, 구조 예측 장치(610)는 2차 구조 또는 3차 구조를 기준으로 연관된 정보를 구조 정보로 생성할 수도 있다. 예컨대, 구조 예측 장치(610)는 구조를 포함하는 영상, 거리 맵, 접촉 맵 등과 같은 형태의 구조 정보를 생성할 수 있다. 또한, 구조 예측 장치(610)는 서열 특이적인 결합 에너지, 원-핫 벡터 등(부가 정보)을 생성할 수도 있다. 구조 예측 장치(610)는 si를 입력받아 xi를 출력한다. 구조 예측 장치(610)는 xi를 i에 대한 식별자와 함께 데이터베이스(DB, 630)에 저장할 수 있다.
점수 연산 장치(620)는 유전자 가위 서열을 기준으로 특정 표적에 대한 유전자 가위의 효과를 특정 점수로 산출한다. 유전자 가위 효과를 나타내는 점수를 효과 점수(effect score)라고 명명한다. 효과 점수는 공지된 솔루션을 이용하여 연산할 수 있다. 예컨대, GenomeCRISPR_full05112017 (https://www.dkfz.de/signaling/crispr-downloads/GENOMECRISPR/)과 같은 솔루션을 이용할 수 있다. 또는 유전자 가위 관련 실험 정보를 관리하는 DB를 활용하여 특정 표적에 대한 특정 서열의 유전자 가위 효과를 추정할 수도 있다. 효과 점수는 특정 표적에 대한 특정 유전자 가위의 효과를 정량적으로 평가한 값이다. 효과 점수는 동일한 표적이라도, 유전자 가위를 구성하는 서열에 따라 달라진다. 점수 연산 장치(620)는 si를 입력받아 yi를 출력한다. 점수 연산 장치(620)는 yi를 i에 대한 식별자와 함께 데이터베이스(DB, 630)에 저장할 수 있다.
구조 예측 장치(610) 및 점수 연산 장치(620)는 일정한 데이터 처리가 가능한 연산장치이다. 예컨대, 구조 예측 장치(610) 및 점수 연산 장치(620)는 서버, PC, 스마트 기기 등과 같은 장치일 수 있다.
데이터베이스(630)는 n개의 서열에 대한 훈련데이터 세트를 저장한다고 가정한다.
훈련 장치(640)는 훈련데이터를 이용하여 학습모델을 훈련한다. 훈련 장치(640)는 입력데이터 xi에 대하여 학습모델이 yi라는 값을 출력하도록 학습모델을 반복적으로 훈련한다. CNN 경우 목적 함수를 최소화하는 방향으로 학습된다. 학습 과정은 CNN 모델에서 사용하는 가중치를 최적화하는 과정에 해당한다. 예컨대, 가중치 최적화는 경사 하강법(gradient descent method)을 이용할 수 있다.
훈련이 제대로 끝나면, 학습모델은 xi라는 구조 정보에 대하여 yi라는 값을 출력하게 된다. 효과 점수 yi는 0 또는 1 중 어느 하나의 값일 수 있다. 또는 효과 점수 yi는 일정한 범위 내의 정수 내지 실수 값일 수도 있다.
도 8은 유전자 가위 서열을 기준으로 구조를 예측하는 장치(610) 및 유전자 가위 서열을 기준으로 효과 점수를 연산하는 장치(620)를 별도로 도시하였다. 그러나, 사용자 단말(50)과 같은 하나의 장치가 유전자 가위 서열을 기준으로 구조 정보를 생성하고, 해당 유전자 가위 서열에 대한 효과 점수를 연산하여 훈련 데이터를 생성할 수 있다.
예컨대, 사용자 단말(50)은 RNA 구조를 예측하는 웹 솔루션(예컨대, RNAfoldWebServer,http://rna.tbi.univie.ac.at/cgi-bin/RNAWebSuite/RNAfold.cgi)을 이용하거나, 설치된 프로그램을 이용하여 유전자 가위 서열 또는 표적 특이적 서열에 대한 2차 구조/3차 구조를 예측할 수 있다. 사용자 단말(50)은 예측된 구조를 기준으로 거리 맵과 같은 구조 정보를 생성할 수 있다. 또한, 사용자 단말(50)은 GenomeCRISPR_full05112017과 같은 솔루션을 이용하여 효과 점수를 연산할 수도 있다.
도 8은 사용자 단말(50)과 훈련 장치(640)를 별도의 객체로 도시하였다. 다만, 훈련 데이터를 마련하는 객체와 학습모델을 훈련하는 객체가 동일할 수도 있다.
한편, 학습모델은 특정한 표적 서열에 대한 유전자 가위 서열의 효과 내지 적합도를 평가한다. 따라서, 학습모델은 표적 서열마다 서로 다른 모델로 사전에 학습되는 것이 바람직하다.
도 9는 유전자 가위 효과를 예측하는 분석장치(700)에 대한 예이다. 분석장치(700)는 도 2의 분석 장치(110, 210 또는 310)에 해당하는 장치이다.
분석장치(700)는 전술한 학습모델을 이용하여 특정 서열을 갖는 유전자 가위의 효과를 예측한다. 분석장치(700)는 물리적으로 다양한 형태로 구현될 수 있다. 예컨대, 분석장치(700)는 PC와 같은 컴퓨터 장치, 네트워크의 서버, 영상 처리 전용 칩셋 등의 형태를 가질 수 있다. 컴퓨터 장치는 스마트 기기 등과 같은 모바일 기기를 포함할 수 있다.
분석장치(700)는 저장장치(710), 메모리(720), 연산장치(730), 인터페이스 장치(740), 통신장치(750) 및 출력장치(760)를 포함한다.
저장장치(710)는 유전자 가위의 효과를 예측하는 신경망 모델을 저장한다. 신경망 모델는 사전에 학습되어야 한다. 저장장치(710)는 RNA 서열을 기준으로 2차 구조 또는 3차 구조를 예측하는 프로그램을 저장할 수 있다. 저장장치(710)는 2차 구조 또는 3차 구조를 기준으로 구조 정보를 생성하는 프로그램을 저장할 수도 있다. 나아가 저장장치(710)는 데이터 처리에 필요한 다른 프로그램 내지 소스 코드 등을 저장할 수 있다. 저장장치(710)는 입력되는 유전자 가위 서열 및 분석 결과(유전자 가위 효과)를 저장할 수 있다.
메모리(720)는 분석장치(700)가 수신한 데이터를 분석하는 과정에서 생성되는 데이터 및 정보 등을 저장할 수 있다.
인터페이스 장치(740)는 외부로부터 일정한 명령 및 데이터를 입력받는 장치이다. 인터페이스 장치(740)는 물리적으로 연결된 입력 장치 또는 외부 저장장치로부터 유전자 가위 서열 또는 가이드 RNA 서열을 입력받을 수 있다. 인터페이스 장치(740)는 데이터 분석을 위한 학습모델을 입력받을 수 있다. 인터페이스 장치(740)는 학습모델 훈련을 위한 학습데이터, 정보 및 파라미터값을 입력받을 수도 있다.
통신장치(750)는 유선 또는 무선 네트워크를 통해 일정한 정보를 수신하고 전송하는 구성을 의미한다. 통신장치(750)는 외부 객체로부터 유전자 가위 서열 또는 가이드 RNA 서열을 수신할 수 있다. 통신장치(750)는 모델 학습을 위한 데이터도 수신할 수 있다. 통신장치(750)는 입력된 서열에 대하여 결정된 분석 결과를 외부 객체로 송신할 수 있다.
통신장치(750) 내지 인터페이스 장치(740)는 외부로부터 일정한 데이터 내지 명령을 전달받는 장치이다. 통신장치(750) 내지 인터페이스 장치(740)를 입력장치라고 명명할 수 있다.
입력 장치는 분석 대상인 유전자 가위에 대한 구조 정보를 입력 내지 수신받을 수 있다. 예컨대, 입력 장치는 외부 서버나 DB로부터 구조 정보를 수신할 수 있다. 입력 장치는 분석 대상인 유전자 가위에 대한 2차 구조 또는 3차 구조 정보를 수신할 수도 있다.
출력장치(760)는 일정한 정보를 출력하는 장치이다. 출력장치(760)는 데이터 처리 과정에 필요한 인터페이스, 분석 결과 등을 출력할 수 있다.
연산 장치(730)는 저장장치(710)에 저장된 프로그램을 이용하여 유전자 가위 서열에 대한 2차 구조 또는 3차 구조를 예측할 수 있다. 연산 장치(730)는 RNA 2차 구조 또는 3차 구조 영상을 생성할 수 있다. 또한, 연산 장치(730)는 RNA 2차 구조 또는 3차 구조를 기준으로 거리 맵과 같은 구조 정보를 생성할 수도 있다. 즉, 연산 장치(730)는 학습모델이 입력받을 수 있는 입력 데이터를 생성할 수 있다. 연산 장치(730)는 필요한 경우 데이터 전처리를 할 수도 있다.
연산 장치(730)는 생성한 입력 데이터(구조 정보)를 학습모델에 입력하여, 분석 대상인 유전자 가위 서열에 대한 효과를 예측할 수 있다. 연산 장치(730)는 효과에 대한 이진 분류 또는 다중 분류 결과를 기준으로 유전자 가위 서열에 대한 효과를 예측할 수 있다.
한편, 연산 장치(730)는 주어진 훈련 데이터를 이용하여 유전자 가위 서열의 효과 예측에 사용되는 학습모델을 훈련할 수도 있다.
연산 장치(730)는 데이터를 처리하고, 일정한 연산을 처리하는 프로세서, AP, 프로그램이 임베디드된 칩과 같은 장치일 수 있다.
이하 전술한 학습모델에 대한 성능을 검증하는 실험 결과를 설명한다.
먼저, 실제 학습모델 훈련에 사용한 데이터에 대하여 설명한다. 훈련 데이터는 GenomeCRISPR_full05112017 데이터 (이하, 데이터 셋)이다. 상기 데이터 셋은 3,800 만개의 유전자 가위 서열들로 구성되어 있다. 데이터 셋 내 유전자 가위 서열들은 고유한 유전자 가위 서열이 아니고, 다수의 세포주에 중복 실험된 유전자 가위 서열들이다. 하나의 유전자 가위는 세포주마다 다른 효과를 보인다.
데이터 셋은 다음과 같은 통계 처리를 통하여 학습이 용이한 데이터 셋으로 준비하였다. 연구자는 최초 데이터 셋에서 30개 미만의 세포주에 실험된 서열을 제외하였다. 이는 유전자 가위 서열의 고유성을 유지하며 잘림 유무를 예측하고자 일부 데이터를 필터링한 것이다. 또한, 연구자는 80% 이상의 세포주에서 효과 점수 5점 이상인 유전자 가위 서열을 효과 좋음(잘 잘림)으로 정의하고, 80% 이상의 세포주에서 효과 점수가 -5점 이하인 유전자 가위 서열을 효과 낮음(안잘림)으로 정의하였다. 필터링된 데이터 셋은 25,593개의 유전자 가위 서열로 구성되었다. 필터링된 데이터 셋은 효과 좋음으로 라벨링된 데이터들 12,481개 및 효과 낮음으로 라벨딩된 데이터들 13,112개로 구성되었다.
학습모델은 전술한 도 7과 같은 모델을 구현하였다. 학습모델은 이진 분류에 따른 효과 예측을 하는 모델이다. 학습모델은 입력 데이터가 서로 다른 4가지 학습모델을 구현하였다. 4가지 학습모델은 아래의 표 1과 같다.
| 모델 구분 | 입력 데이터 |
| 제1 모델 | 2차원 거리 맵 |
| 제2 모델 | (i) 2차원 거리 맵, (ii) 결합 에너지 |
| 제3 모델 | (i) 2차원 거리 맵, (ii) 원-핫 벡터 |
| 제4 모델 | (i) 2차원 거리 맵, (ii) 결합 에너지, (iii) 원-핫 벡터 |
4개의 학습모델은 모두 k-폴드 교차검증(k-fold cross validation)을 수행하였다. k는 5로 설정하였다. 학습 데이터는 랜덤으로 섞어서 5개로 나누었다. 5개 중 4개는 학습용으로 사용하였고, 나머지 1개는 검증용 데이터로 사용하였다.
다음 폴드에서 검증용 데이터는 이전 폴드와 다른 데이터로 변경하였고, 학습 데이터는 검증용 데이터로 선택된 하나의 데이터를 제외한 나머지 데이터를 사용하였다. 결국, 5번의 검증 과정에서, 학습모델은 서로 다른 조합의 학습 데이터들을 사용하여 학습되었고, 서로 다른 데이터를 사용하여 검증되었다. 나아가, 5-폴드 교차 검증을 30번 반복 수행하였다.
학습모델의 성능 평가 측도는 정확도 (accuracy), ROC (Receiver Operating Characteristics)의 곡선 아래 면적 (Area Under the Curve, AUC) 등을 사용하였다. 성능은 실험 횟수 (iterations) 당 5-폴드 교차검증에 의해 생성된 5개의 데이터 셋의 평균 성능의 분포를 기준으로 분석하였다.
ROC 곡선을 나타내는 그래프는 x축이 분류 모형의 위양성률 (false positive rate, FPR)이고, y축이 진양성률 (true positive rate, TPR)이다. TPR은 민감도(sensitivity)를 나타내고, FPR은 '1-특이도(specificity)'를 나타낸다.
이진 분류(binary classification) 문제에 대해, AUC는 분류기의 양성 데이터(positive object)에 대한 예측 값이 음성데이터(negative object)에 대한 예측 값보다 높을 확률을 의미한다. AUC는 ROC 곡선의 아래 면적을 구함으로써 추정 가능하다.
정확도는 모델 예측 값과 데이터 y값을 비교하며 전체 데이터에 대하여 분류가 일치하는 비율을 나타낸다. 정확도(ACC)는 아래 수학식으로 표현할 수 있다. TP는 True positive 개수, TN은 True negative 개수, FP는 False positive 개수이고, FN은 False negative 개수이다.
도 10은 유전자 가위 효과 예측한 모델 성능 평가에 대한 예이다. 도 10은 학습모델에 대하여 에폭(epoch) 당 정확도를 나타낸다. 도 10은 5-폴드 교차 검증을 30번 반복(전체 학습을 150번)한 결과의 평균을 나타낸다. 도 10(A)는 학습 데이터를 이용하여 에폭 당 정확도를 나타낸 결과이고, 도 10(B)는 검증 데이터를 이용하여 에폭 당 정확도를 나타낸 결과이다. 도 10의 결과를 살펴보면, 제4 모델(거리맵 + 결합 에너지 + 원-핫 벡터)이 4개의 학습모델 중 가장 정확도가 높았다. 정확도 순서는 제4 모델 > 제2 모델(거리맵 + 결합 에너지) > 제3 모델(거리맵 + 원-핫 벡터) > 제1 모델(거리맵)이다. 전체적으로 학습모델은 입력데이터로 거리맵과 함께 에너지가 사용될 때 비교적 높은 성능을 보였다.
도 11은 유전자 가위 효과 예측한 모델 성능 평가에 대한 다른 예이다. 도 11은 5-폴드의 평균 AUC 값을 30번 반복적으로 구한 분포도를 박스 플롯(box plot)으로 나타낸다. 도 11(A)는 훈련 데이터를 사용하여 분석한 결과이다. 도 11(B)는 검증 데이터를 사용하여 분석한 결과이다. 검증 데이터 기준으로 AUC의 중간 값을 살펴보면, 에너지를 포함하는 학습모델은 약 0.93, 포함하지 않은 학습모델은 약 0.922이다. 따라서, 전체적으로 에너지를 포함한 모델 성능이 포함하지 않은 모델보다 성능이 좋다고 할 수 있다.
도 12는 유전자 가위 효과 예측한 모델 성능 평가에 대한 또 다른 예이다. 도 12는 학습모델에 대한 TPR과 TNR 기준의 성능을 나타내는 박스 플롯이다. 도 12(A)는 임계값 0.5 기준으로 TPR의 성능을 평가한 결과이고, 도 12(B)는 임계값 0.5 기준으로 TNR의 성능을 평가한 결과이다. 도 12를 살펴보면, 입력 데이터로 에너지를 포함한 모델이 에너지를 포함하지 않은 모델보다 성능이 더 높다는 것을 알 수 있다.
또한, 상술한 바와 같은 학습모델 훈련 방법 및 학습모델을 이용한 유전자 가위 효과 분석 방법은 컴퓨터에서 실행될 수 있는 실행가능한 알고리즘을 포함하는 프로그램(또는 어플리케이션)으로 구현될 수 있다. 상기 프로그램은 비일시적 판독 가능 매체(non-transitory computer readable medium)에 저장되어 제공될 수 있다.
비일시적 판독 가능 매체란 레지스터, 캐쉬, 메모리 등과 같이 짧은 순간 동안 데이터를 저장하는 매체가 아니라 반영구적으로 데이터를 저장하며, 기기에 의해 판독(reading)이 가능한 매체를 의미한다. 구체적으로는, 상술한 다양한 어플리케이션 또는 프로그램들은 CD, DVD, 하드 디스크, 블루레이 디스크, USB, 메모리카드, ROM 등과 같은 비일시적 판독 가능 매체에 저장되어 제공될 수 있다.
본 실시례 및 본 명세서에 첨부된 도면은 전술한 기술에 포함되는 기술적 사상의 일부를 명확하게 나타내고 있는 것에 불과하며, 전술한 기술의 명세서 및 도면에 포함된 기술적 사상의 범위 내에서 당업자가 용이하게 유추할 수 있는 변형 예와 구체적인 실시례는 모두 전술한 기술의 권리범위에 포함되는 것이 자명하다고 할 것이다.
Claims (17)
- 분석장치가 유전자 가위의 서열 데이터를 입력받는 단계;상기 분석장치가 상기 서열 데이터를 이용하여 유전자 가위의 표적 특이적 서열이 구성하는 구조에 대한 구조 정보를 생성하는 단계;상기 분석장치가 상기 구조 정보를 학습모델에 입력하는 단계; 및상기 분석장치가 상기 학습모델이 출력하는 정보를 기준으로 상기 유전자 가위의 효과를 평가하는 단계를 포함하는 학습모델을 이용하여 유전자 가위 효과를 분석하는 방법.
- 제1항에 있어서,상기 표적 특이적 서열은 상기 유전자 가위를 구성하는 가이드 RNA 전체 또는 일부 서열인 학습모델을 이용하여 유전자 가위 효과를 분석하는 방법.
- 제1항에 있어서,상기 분석장치는 뉴클레오타이드 서열을 기준으로 구조를 예측하는 프로그램을 이용하여 상기 표적 특이적 서열에 대한 2차 구조 또는 3차 구조를 예측하고, 상기 2차 구조 또는 3차 구조에 대한 상기 구조 정보를 생성하는 학습모델을 이용하여 유전자 가위 효과를 분석하는 방법.
- 제1항에 있어서,상기 구조 정보는 상기 표적 특이적 서열의 2차 구조 또는 3차 구조에 대한 영상인 학습모델을 이용하여 유전자 가위 효과를 분석하는 방법.
- 제1항에 있어서,상기 구조 정보는 상기 표적 특이적 서열의 2차 구조 또는 3차 구조에서 한 쌍의 뉴클레오타이드 사이의 거리 정보를 시각적으로 표현한 거리 맵(distance map)인 학습모델을 이용하여 유전자 가위 효과를 분석하는 방법.
- 제1항에 있어서,상기 분석장치는 상기 서열 데이터를 이용하여 결합 에너지(hybrid duplexes energy) 및 원-핫 벡터(one-hot vector) 중 적어도 하나를 포함하는 부가 정보를 생성하고, 상기 학습모델에 상기 부가 정보를 더 입력하는 학습모델을 이용하여 유전자 가위 효과를 분석하는 방법.
- 제6항에 있어서,상기 학습모델은 컨볼루션 계층, 풀링 계층 및 전연결 계층을 포함하는 인공신경망 모델이고,상기 컨볼루션 계층과 상기 풀링 계층은 상기 구조 정보를 특징 맵으로 변환하고,상기 전연결 계층은 상기 특징 맵과 상기 부가 정보를 입력받아 상기 표적 특이적 서열을 포함하는 유전자 가위에 대한 효과 정보를 출력하는 학습모델을 이용하여 유전자 가위 효과를 분석하는 방법.
- 유전자 가위의 서열 데이터를 입력받는 입력장치;뉴클레오타이드 서열 기준으로 해당 서열이 구성하는 구조에 대한 구조 정보를 생성하는 프로그램 및 유전자 가위의 구조 정보를 기준으로 유전자 가위의 효과를 분석하는 학습모델을 저장하는 저장장치; 및상기 프로그램을 이용하여 상기 서열 데이터에서 유전자 가위의 표적 특이적 서열에 대한 구조 정보를 생성하고, 생성된 구조 정보를 상기 학습모델에 입력하여 유전자 가위의 효과를 평가하는 연산장치를 포함하는 유전자 가위 효과를 예측하는 분석장치.
- .제8항에 있어서,상기 표적 특이적 서열은 상기 유전자 가위를 구성하는 가이드 RNA 전체 또는 일부 서열인 유전자 가위 효과를 예측하는 분석장치.
- 제8항에 있어서,상기 구조 정보는 상기 표적 특이적 서열의 2차 구조 또는 3차 구조에서 한 쌍의 뉴클레오타이드 사이의 거리 정보를 시각적으로 표현한 거리 맵(distance map)인 유전자 가위 효과를 예측하는 분석장치.
- 제8항에 있어서,상기 저장장치는 뉴클레오타이드 서열에 대한 결합 에너지(hybrid duplexes energy)를 연산하는 프로그램을 더 저장하고,상기 연산장치는 상기 서열 데이터를 이용하여 결합 에너지를 연산하고, 상기 결합 에너지를 상기 학습모델에 더 입력하는 유전자 가위 효과를 예측하는 분석장치.
- 제8항에 있어서,상기 연산장치는 상기 서열 데이터에 대한 원-핫 인코딩(one-hot encoding)을 하여 원-핫 벡터를 생성하고, 상기 원-핫 벡터를 상기 학습모델에 더 입력하는 유전자 가위 효과를 예측하는 분석장치.
- 제8항에 있어서,상기 연산장치는 상기 서열 데이터를 이용하여 결합 에너지(hybrid duplexes energy) 및 원-핫 벡터(one-hot vector) 중 적어도 하나를 포함하는 부가 정보를 생성하고, 상기 학습모델에 상기 부가 정보를 더 입력하는 유전자 가위 효과를 예측하는 분석장치.
- 제13항에 있어서,상기 학습모델은 컨볼루션 계층, 풀링 계층 및 전연결 계층을 포함하는 인공신경망 모델이고,상기 컨볼루션 계층과 상기 풀링 계층은 상기 구조 정보를 특징 맵으로 변환하고,상기 전연결 계층은 상기 특징 맵과 상기 부가 정보를 입력받아 상기 표적 특이적 서열을 포함하는 유전자 가위에 대한 효과 정보를 출력하는 유전자 가위 효과를 예측하는 분석장치.
- 컴퓨터에서 제1항 내지 제7항 중 어느 하나의 항에 학습모델을 이용하여 유전자 가위를 설계하는 방법을 실행하기 위한 프로그램을 기록한 컴퓨터로 읽을 수 있는 기록 매체.
- 뉴클레오타이드 서열 기준으로 해당 서열이 구성하는 구조에 대한 구조 정보를 생성하는 구조 예측 서버; 및유전자 가위의 구조 정보를 기준으로 유전자 가위의 효과를 분석하는 학습모델을 저장하고, 상기 구조 예측 서버로부터 유전자 가위의 서열 데이터에서 유전자 가위의 표적 특이적 서열에 대한 구조 정보를 수신하고, 상기 구조 정보를 상기 학습모델에 입력하여 유전자 가위의 효과를 평가하는 평가 서버를 포함하는 유전자 가위 효과를 예측하는 시스템.
- 제16항에 있어서,상기 구조 정보는 상기 표적 특이적 서열의 2차 구조 또는 3차 구조에 대한 영상 또는 상기 표적 특이적 서열의 2차 구조 또는 3차 구조에서 한 쌍의 뉴클레오타이드 사이의 거리 정보를 시각적으로 표현한 거리 맵(distance map)인 유전자 가위 효과를 예측하는 시스템.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| KR10-2019-0138120 | 2019-10-31 | ||
| KR1020190138120A KR102166070B1 (ko) | 2019-10-31 | 2019-10-31 | 유전자 가위 효과를 분석하는 방법 및 장치 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2021085702A1 true WO2021085702A1 (ko) | 2021-05-06 |
Family
ID=73035141
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/KR2019/014785 Ceased WO2021085702A1 (ko) | 2019-10-31 | 2019-11-04 | 유전자 가위 효과를 분석하는 방법 및 장치 |
Country Status (2)
| Country | Link |
|---|---|
| KR (1) | KR102166070B1 (ko) |
| WO (1) | WO2021085702A1 (ko) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114649052A (zh) * | 2022-03-22 | 2022-06-21 | 山东省计算中心(国家超级计算济南中心) | 基于超算互联网的rna结构预测方法及系统 |
Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR20190048926A (ko) * | 2017-10-31 | 2019-05-09 | 연세대학교 산학협력단 | 딥러닝을 이용한 rna-가이드 뉴클레아제의 활성 예측 시스템 |
-
2019
- 2019-10-31 KR KR1020190138120A patent/KR102166070B1/ko active Active
- 2019-11-04 WO PCT/KR2019/014785 patent/WO2021085702A1/ko not_active Ceased
Patent Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR20190048926A (ko) * | 2017-10-31 | 2019-05-09 | 연세대학교 산학협력단 | 딥러닝을 이용한 rna-가이드 뉴클레아제의 활성 예측 시스템 |
Non-Patent Citations (5)
| Title |
|---|
| "Bioinformatics", 28 November 2012, INTECH , ISBN: 9789535108788, article SUMAN GHOSAL, SHAOLI DAS, JAYPROKAS CHAKRABARTI: "Computational Approaches for Designing Efficient and Specific siRNAs", pages: 261 - 276, XP055210577, DOI: 10.5772/50125 * |
| FUSI, N. ET AL.: "In silico predictive modeling of CRISP/Cas9 guide efficiency", THE PREPRAINT SERVER FOR BIOLOGY, 26 June 2015 (2015-06-26), pages 1 - 31, XP055892489, DOI: 10.1101/021568 * |
| GUOHUI CHUAI, HANHUI MA, JIFANG YAN, MING CHEN, NANFANG HONG, DONGYU XUE, CHI ZHOU, CHENYU ZHU, KE CHEN, BIN DUAN, FENG GU, SHENG : "DeepCRISPR: optimized CRISPR guide RNA design by deep learning", GENOME BIOLOGY, vol. 19, no. 1, 80, 1 December 2018 (2018-12-01), XP055716006, DOI: 10.1186/s13059-018-1459-4 * |
| MICHAL J PIETAL;NATALIA SZOSTAK;KRISTIAN M ROTHER;JANUSZ M BUJNICKI: "RNAmap2D – calculation, visualization and analysis of contact and distance maps for RNA and protein-RNA complex struct", BMC BIOINFORMATICS, BIOMED CENTRAL , LONDON, GB, vol. 13, no. 1, 21 December 2012 (2012-12-21), GB , pages 333, XP021134544, ISSN: 1471-2105, DOI: 10.1186/1471-2105-13-333 * |
| WANG LEI, ZHANG JUHUA: "Prediction of sgRNA on-target activity in bacteria by deep learning", BMC BIOINFORMATICS, vol. 20, no. 1, 1 December 2019 (2019-12-01), pages 517, XP055932349, DOI: 10.1186/s12859-019-3151-4 * |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114649052A (zh) * | 2022-03-22 | 2022-06-21 | 山东省计算中心(国家超级计算济南中心) | 基于超算互联网的rna结构预测方法及系统 |
| CN114649052B (zh) * | 2022-03-22 | 2024-10-25 | 山东省计算中心(国家超级计算济南中心) | 基于超算互联网的rna结构预测方法及系统 |
Also Published As
| Publication number | Publication date |
|---|---|
| KR102166070B1 (ko) | 2020-10-16 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Tampuu et al. | ViraMiner: Deep learning on raw DNA sequences for identifying viral genomes in human samples | |
| CN111489324B (zh) | 一种融合多模态先验病理深度特征的宫颈图像分类方法 | |
| WO2019107614A1 (ko) | 제조 공정에서 딥러닝을 활용한 머신 비전 기반 품질검사 방법 및 시스템 | |
| Atikuzzaman et al. | Human activity recognition system from different poses with cnn | |
| WO2019235828A1 (ko) | 투 페이스 질병 진단 시스템 및 그 방법 | |
| WO2021107676A1 (ko) | 인공지능 기반 염색체 이상 검출 방법 | |
| WO2019178291A1 (en) | Methods for data segmentation and identification | |
| WO2020045848A1 (ko) | 세그멘테이션을 수행하는 뉴럴 네트워크를 이용한 질병 진단 시스템 및 방법 | |
| CN111564179B (zh) | 一种基于三元组神经网络的物种生物学分类方法及系统 | |
| WO2022146050A1 (ko) | 우울증 진단을 위한 인공지능 연합학습 방법 및 시스템 | |
| WO2021075742A1 (ko) | 딥러닝 기반의 가치 평가 방법 및 그 장치 | |
| WO2021010671A2 (ko) | 뉴럴 네트워크 및 비국소적 블록을 이용하여 세그멘테이션을 수행하는 질병 진단 시스템 및 방법 | |
| CN113221929A (zh) | 一种图像处理方法以及相关设备 | |
| Alsharif et al. | Machine learning technology to recognize American sign language alphabet | |
| WO2019045147A1 (ko) | 딥러닝을 pc에 적용하기 위한 메모리 최적화 방법 | |
| WO2023195564A1 (ko) | 공간전사체정보 분석장치 및 이를 이용한 분석방법 | |
| CN117611901B (zh) | 一种基于全局和局部对比学习的小样本图像分类方法 | |
| WO2020032561A2 (ko) | 다중 색 모델 및 뉴럴 네트워크를 이용한 질병 진단 시스템 및 방법 | |
| WO2021085702A1 (ko) | 유전자 가위 효과를 분석하는 방법 및 장치 | |
| WO2021177532A1 (ko) | 인공지능을 이용하여 정렬된 염색체 이미지의 분석을 통한 염색체 이상 판단 방법, 장치 및 컴퓨터프로그램 | |
| CN116439663A (zh) | 基于自监督学习和多视图学习的睡眠分期系统 | |
| Sharma et al. | Detecting human embryo cleavage stages using YOLO V5 object detection algorithm | |
| JP2021072106A (ja) | イメージ処理システム | |
| Bai et al. | A unified deep learning model for protein structure prediction | |
| WO2023075402A1 (ko) | 메틸화된 무세포 핵산을 이용한 암 진단 및 암 종 예측방법 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19950392 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19950392 Country of ref document: EP Kind code of ref document: A1 |