WO2024070589A1 - オフターゲットリスク解析方法、オフターゲットリスク解析システム、プログラム、記録媒体 - Google Patents
オフターゲットリスク解析方法、オフターゲットリスク解析システム、プログラム、記録媒体 Download PDFInfo
- Publication number
- WO2024070589A1 WO2024070589A1 PCT/JP2023/032841 JP2023032841W WO2024070589A1 WO 2024070589 A1 WO2024070589 A1 WO 2024070589A1 JP 2023032841 W JP2023032841 W JP 2023032841W WO 2024070589 A1 WO2024070589 A1 WO 2024070589A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- sequence
- target
- virtual
- score
- target sequence
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B30/00—ICT specially adapted for sequence analysis involving nucleotides or amino acids
Definitions
- the present disclosure relates to an off-target risk analysis method and an off-target prediction system that analyze the risk of off-target effects occurring in genome editing and its derived technologies.
- Genome editing is a technology that recognizes the DNA sequence of a target region on a genome sequence (hereafter referred to as the target sequence) and introduces mutations such as deletions, substitutions, and insertions into the sequence of any target gene using a DNA-binding tool that can cleave the target region.
- DNA-binding tools include zinc-finger nucleases (ZFNs), TALE nucleases (Transcription Activator-Like Effector Nucleases), and CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats)/Cas (CRISPR-associated protein).
- off-target effects a phenomenon in which unexpected mutations are introduced into non-target genome sequences
- the frequency of off-target effects depends on the DNA sequence recognized by the DNA-binding tool, so even if the same DNA-binding tool (e.g., CRISPR/Cas9) is used, the frequency of off-target effects varies if the target sequence is different.
- the risk of introducing unexpected mutations into other genome sequences that are not the target genome sequence in genome editing is also called off-target risk.
- DNA-binding tools e.g., transcriptional regulation, epigenetic editing, chromosome imaging, etc.
- technologies that apply DNA-binding tools e.g., transcriptional regulation, epigenetic editing, chromosome imaging, etc.
- zinc fingers e.g., TALEs, and CRISPR/Cas with inactivated nuclease activity, so there is a risk of off-target effects depending on the target sequence.
- the off-target risk prediction methods shown in (1) and (2) above can only be used for biological species whose entire genome information has been decoded, or for specific individual organisms, tissues, cell clones, varieties, bacterial strains, and virus strains, etc. Therefore, it has been difficult to apply these methods to predicting off-target risks when using DNA-binding tools for biological species (including industrial organisms) whose genome information has not been decoded well or whose genomes are difficult to decode because they contain difficult-to-read sequence elements such as repeat structures.
- genomic information is not definitive because spontaneous mutations may occur in each individual and cell. Even by referring to the sequences of reference genomes and unique genomes sequenced in advance, it is not possible to predict potential off-target effects of DNA-binding tools that may be manifested in the presence of spontaneous mutations.
- DNA-binding tools for example, medical treatments that apply DNA-binding tools
- the potential off-target risks of DNA-binding tools can be a major risk, but until now there has been no method that can predict these risks in advance.
- the off-target risk analysis method includes a virtual sequence generation step of generating a plurality of virtual sequences including a sequence identical to a target sequence and a sequence in which at least one mutation has been introduced into the target sequence, a score calculation step of calculating, for each of the plurality of virtual sequences, a score related to the probability that a DNA-binding tool that recognizes the target sequence will act, and a prediction result output step of outputting, based on the calculated score, a prediction result indicating the possibility that the DNA-binding tool will act on a sequence different from the target sequence.
- the off-target risk analysis system includes a virtual sequence generation unit that generates a plurality of virtual sequences including a sequence identical to a target sequence and a sequence in which at least one mutation has been introduced into the target sequence, a score calculation unit that calculates a score related to the probability that a DNA-binding tool that recognizes the target sequence will act for each of the plurality of virtual sequences, and a prediction result output unit that outputs a prediction result indicating the possibility that the DNA-binding tool will act on a sequence different from the target sequence based on the calculated score.
- the off-target risk analysis system may be realized by a computer.
- the control program of the off-target risk analysis system that realizes the off-target risk analysis system on a computer by causing the computer to operate as each part (software element) of the off-target risk analysis system, and the computer-readable recording medium on which it is recorded, also fall within the scope of the present disclosure.
- FIG. 1 is a block diagram showing an example of a schematic configuration of an off-target risk analysis system.
- FIG. 13 is a diagram illustrating an example of a virtual array.
- FIG. 13 is a diagram illustrating an example of a virtual array.
- FIG. 13 is a diagram illustrating an example of a virtual array.
- 13 is a flowchart showing an example of a processing flow executed by the off-target risk analysis device.
- FIG. 1 is a diagram illustrating an example of a schematic configuration of an off-target risk analysis system.
- FIG. 1 is a block diagram showing an example of a schematic configuration of an off-target risk analysis device.
- FIG. 4 is a diagram illustrating an example of a data structure of user information.
- FIG. 13 illustrates an example of a data structure of an analysis result log.
- 1 is a graph showing the correlation between the prediction results output using the off-target risk analysis method according to the present disclosure and the off-target incidence rate obtained as a result of analyzing the off-target effect in actual cells.
- 1 is a graph showing the correlation between the prediction results output using the off-target risk analysis method according to the present disclosure and the off-target incidence rate obtained as a result of analyzing the off-target effect in actual cells.
- 1 is a graph showing the correlation between the prediction results output using the off-target risk analysis method according to the present disclosure and the off-target incidence rate obtained as a result of analyzing the off-target effect in actual cells.
- 13 is a scatter plot showing the correlation between the specificity prediction score calculated using the MOFF score and the “Off-on-ratio” score in the “TTISS” group.
- the off-target risk analysis system 100 is a system that outputs a prediction result indicating the possibility that a DNA binding tool that recognizes a target sequence may act on a sequence different from the target sequence.
- a DNA-binding tool is a tool that specifically binds to a specific genomic region and is capable of cutting, modifying, editing, etc., the genomic region.
- the genomic region that is cut, modified, edited, etc. by the DNA-binding tool may be a specific region on the genome of a cell, and the DNA-binding tool may also be referred to as a genome editing tool.
- off-target effect When a DNA-binding tool that recognizes a target sequence acts on a sequence other than the target sequence, this is called an "off-target effect.”
- off-target effects on genes in the genome of a cell can result in the activation of cancer genes and the inactivation of tumor suppressor genes in that cell.
- the effects of off-target effects have a permanent effect on cells. Therefore, it is desirable to accurately estimate the risk of off-target effects (hereinafter referred to as off-target risk) before actually using a DNA-binding tool that recognizes a target sequence.
- the off-target risk analysis system 100 executes the following processes (a) to (c).
- a "DNA binding tool” may be any of (i) to (v) below.
- a TALE Transcription Activator-Like Effector
- Pentatricopeptide repeats (PPRs) or fusion polypeptides combining pentatricopeptide repeats with a functional domain.
- PPRs Pentatricopeptide repeats
- PPRs Pentatricopeptide repeats
- a wild-type CRISPR/Cas CRISPR-associated protein
- a fusion polypeptide-nucleic acid complex combining a wild-type CRISPR/Cas with a functional domain.
- a modified CRISPR/Cas CRISPR-associated protein
- a fusion polypeptide-nucleic acid complex combining a modified CRISPR/Cas with a functional domain.
- the target sequence is a base sequence recognized by a domain having a zinc finger protein motif.
- the target sequence is a base sequence recognized by a region to which a TAL module (Transcription Activator-Like Module) is bound.
- TAL module Transcription Activator-Like Module
- the target sequence is a base sequence recognized by a region of consecutive PPR motifs.
- the target sequence is a base sequence complementary to the guide RNA (gRNA) that forms a complex with Cas, and a protospacer adjacent motif (PAM) sequence recognized by Cas.
- gRNA guide RNA
- PAM protospacer adjacent motif
- a DNA-binding tool that falls under any of the above (i) to (v) may misidentify unintended DNA that has a base sequence similar to the target sequence.
- Non-Patent Documents 1 to 3 All of the analytical methods described in Non-Patent Documents 1 to 3 are capable of analyzing off-target risks when DNA-binding tools are used on biological species whose entire genome information has been decoded, or on specific individual organisms, tissues, cell clones, varieties, bacterial strains, and virus strains. However, it is difficult to apply the analytical methods described in Non-Patent Documents 1 to 3 to predict off-target risks when DNA-binding tools are used on biological species (including industrial organisms) whose genome information has not been decoded or whose genome is difficult to decode.
- the off-target risk analysis system 100 can output a prediction result indicating the off-target risk of a DNA-binding tool that recognizes a target sequence, provided that the target sequence is given.
- the off-target risk analysis system 100 can evaluate the risk of causing an off-target effect that a DNA-binding tool that recognizes a target sequence potentially has, without referring to the genomic sequence of the target to which the DNA-binding tool is applied.
- Fig. 1 is a block diagram showing an example of a schematic configuration of the off-target risk analysis system 100.
- the off-target risk analysis system 100 may include an off-target risk analysis device 1 and a display device 4.
- FIG. 1 shows an off-target risk analysis system 100 that includes one off-target risk analysis device 1 and one display device 4.
- the configuration of the off-target risk analysis system 100 is not limited to this.
- the number of display devices 4 in the off-target risk analysis system 100 may be zero or more than one.
- the off-target risk analysis device 1 and the display device 4 are connected to each other so that they can communicate with each other.
- the off-target risk analysis device 1 and the display device 4 may be directly connected by wire or wirelessly, or may be connected via a communication network.
- the form of the communication network is not limited, and may be a local area network (LAN) or the Internet.
- the off-target risk analysis device 1 is a device that uses target sequence data to output a prediction result indicating the off-target risk of a DNA-binding tool.
- the output prediction result may be transmitted from the off-target risk analysis device 1 to a display device 4.
- the display device 4 may typically be a computer, smartphone, tablet terminal, etc. used by a user of the off-target risk analysis system 100.
- FIG. 1 shows an off-target risk analysis system 100 in which the display device 4 is separate from the off-target risk analysis device 1.
- the configuration of the off-target risk analysis system 100 is not limited to this.
- the display device 4 may be a device integrated with the off-target risk analysis device 1, in which case the display device 4 may be a display unit (display, etc.) provided in the off-target risk analysis device 1.
- the off-target risk analysis device 1 includes a control unit 10, a storage unit 20, and an input unit 30.
- control unit 10 may be a CPU (Central Processing Unit).
- the control unit 10 reads a control program, which is software stored in the memory unit 20, and expands it into a memory such as a RAM (Random Access Memory) to execute various functions. Note that, in the memory unit 20 shown in Figure 1, the control program is not shown in order to simplify the explanation.
- the control unit 10 includes a target sequence receiving unit 11, a virtual sequence generating unit 12, a score calculating unit 13, and a prediction result output unit 14.
- the target sequence receiving unit 11 receives target sequence data indicating a target sequence input using the input unit 30.
- the target sequence receiving unit 11 may store the received target sequence data in the memory unit 20.
- the virtual sequence generating unit 12 generates a plurality of virtual sequences from the target sequence data, including a sequence identical to the target sequence and a sequence in which at least one mutation has been introduced into the target sequence.
- the virtual sequence generating unit 12 virtually generates a variety of virtual sequence data including sequences that may be misrecognized by a DNA binding tool that recognizes the target sequence.
- the virtual sequence generating unit 12 may generate virtual sequences based on virtual sequence generation rules 21 stored in the memory unit 20.
- the mutation that the virtual sequence generation unit 12 introduces into the target sequence may be any one of substitutions, deletions, and insertions.
- Figs. 2 to 4 are diagrams showing examples of virtual sequences.
- the virtual sequences shown in Figs. 2 to 4 are generated from a target sequence derived from the human ⁇ -globin gene, which consists of 23 bases.
- the virtual sequence may be shorter or longer than 23 bases.
- the target sequence is not limited to that derived from the human ⁇ -globin gene.
- the virtual sequence does not have to be a continuous sequence, and may be, for example, a discontinuous sequence that contains a single or multiple arbitrary bases (N) at one or multiple locations.
- FIG. 2 shows an example of a virtual sequence in which a substitution has been introduced into a target sequence.
- the virtual sequence generation unit 12 may generate a plurality of virtual sequences including a sequence identical to the target sequence and a sequence in which a substitution has been introduced into the target sequence.
- sequence M1 (SEQ ID NO: 1) is the same sequence as the target sequence.
- Sequences M2 to M7 show examples of sequences in which a substitution has been introduced into one of the nucleotides contained in sequence M1.
- sequence M2 (SEQ ID NO: 2) is a sequence in which the "A" at the 5' end of sequence M1 has been replaced with a "T”
- sequence M3 (SEQ ID NO: 3) is a sequence in which it has been replaced with a "G”
- sequence M4 (SEQ ID NO: 4) is a sequence in which it has been replaced with a "C”.
- sequence M5 (SEQ ID NO: 5) is a sequence in which the second "G” from the 5' end of sequence M1 has been replaced with an "A”
- sequence M6 (SEQ ID NO: 6) is a sequence in which it has been replaced with a "T”
- sequence M7 (SEQ ID NO: 7) is a sequence in which it has been replaced with a "C”.
- sequences M2 to M7 are shown as examples of virtual sequences in which a substitution has been introduced into one nucleotide of the target sequence, but the virtual sequences generated by the virtual sequence generation unit 12 are not limited to these.
- the virtual sequence generation unit 12 may comprehensively generate sequences in which a single base substitution has been introduced into the target sequence.
- the virtual sequence generation unit 12 may generate virtual sequences in which substitutions have been introduced into multiple nucleotides of the target sequence.
- the virtual sequence generation unit 12 may comprehensively generate sequences in which two-base substitutions, three-base substitutions, and four-base substitutions have been introduced into the target sequence.
- the multiple virtual sequences generated by the virtual sequence generating unit 12 may include the following sequences: A sequence in which at least one adenine (A) in the target sequence is replaced with at least one of thymine (T), cytosine (C), and guanine (G), and/or A sequence in which at least one thymine (T) in the target sequence is replaced with at least one of adenine (A), cytosine (C), and guanine (G), and/or A sequence in which at least one cytosine (C) in the target sequence is replaced with at least one of adenine (A), thymine (T), and guanine (G), and/or A sequence in which at least one guanine (G) in the target sequence is replaced with at least one of adenine (A), thymine (T), and cytosine (C).
- FIG. 3 shows an example of a virtual sequence in which a deletion has been introduced into a target sequence.
- the virtual sequence generation unit 12 may generate a plurality of virtual sequences including a sequence identical to the target sequence and a sequence in which a deletion has been introduced into the target sequence.
- sequence M1 (SEQ ID NO: 1) is the same sequence as the target sequence.
- Sequences M8 to M10 show examples of sequences in which a deletion has been introduced into one of the nucleotides contained in sequence M1.
- sequence M8 (SEQ ID NO: 8) is a sequence in which the "A" at the 5' end of sequence M1 has been deleted
- sequence M9 (SEQ ID NO: 9) is a sequence in which the second "G” from the 5' end of sequence M1 has been deleted
- sequence M10 (SEQ ID NO: 10) is a sequence in which the third "C" from the 5' end of sequence M1 has been deleted.
- sequences M8 to M10 are shown as examples of virtual sequences in which a deletion has been introduced into one nucleotide of the target sequence, but the virtual sequences generated by the virtual sequence generation unit 12 are not limited to these.
- the virtual sequence generation unit 12 may comprehensively generate sequences in which a sequence in which a single base deletion has been introduced into the target sequence.
- the virtual sequence generation unit 12 may generate virtual sequences in which deletions have been introduced into multiple nucleotides of the target sequence.
- the virtual sequence generation unit 12 may comprehensively generate sequences in which a two-base deletion, a three-base deletion, and a four-base deletion have been introduced into the target sequence.
- FIG. 4 shows an example of a virtual sequence in which an insertion has been introduced into a target sequence.
- the virtual sequence generating unit 12 may generate a plurality of virtual sequences including a sequence identical to the target sequence and a sequence in which an insertion has been introduced into the target sequence.
- sequence M1 (SEQ ID NO: 1) is the same sequence as the target sequence.
- Sequences M11 to M18 show examples of sequences in which one base has been inserted into sequence M1.
- sequence M11 (SEQ ID NO: 11) is a sequence in which "A” has been inserted between the 5'-end and second nucleotide "AG” of sequence M1
- sequence M12 (SEQ ID NO: 12) is a sequence in which "T” has been inserted
- sequence M13 (SEQ ID NO: 13) is a sequence in which "G” has been inserted
- sequence M14 (SEQ ID NO: 14) is a sequence in which "C” has been inserted.
- Sequence M15 is a sequence in which "A” has been inserted between the second and third nucleotides "GC” from the 5'-end of sequence M1
- sequence M16 (SEQ ID NO: 6) is a sequence in which "T” has been inserted
- sequence M17 (SEQ ID NO: 17) is a sequence in which "G” has been inserted
- sequence M18 (SEQ ID NO: 18) is a sequence in which "C" has been inserted.
- sequences M11 to M18 are shown as examples of virtual sequences in which an insertion has been introduced at one position in the target sequence, but the virtual sequences generated by the virtual sequence generation unit 12 are not limited to these.
- the virtual sequence generation unit 12 may generate a sequence in which a two-base insertion has been introduced at one position in the target sequence.
- the virtual sequence generation unit 12 may also comprehensively generate sequences in which an insertion has been introduced at one position in the target sequence.
- the virtual sequence generation unit 12 may generate virtual sequences in which insertions have been introduced at multiple positions in the target sequence.
- the virtual sequence generation unit 12 may comprehensively generate sequences in which insertions have been introduced at two, three, and four positions in the target sequence.
- the virtual sequence generating unit 12 may generate a virtual sequence in which two or more types of mutations are introduced into the target sequence. For example, the virtual sequence generating unit 12 may generate a virtual sequence in which a substitution is introduced into the target sequence, a virtual sequence in which a deletion is introduced into the target sequence, and a virtual sequence in which an insertion is introduced into the target sequence.
- the DNA-binding tool is less likely to act on sequences that have low homology to the target sequence recognized by the DNA-binding tool. Therefore, each of the multiple virtual sequences only needs to differ from the target sequence by four or fewer nucleotides, and there is little need to generate virtual sequences that introduce more than this number of mutations. This ensures the accuracy of the prediction results while reducing the burden on computational resources.
- the off-target risk analysis device 1 may be configured not to generate a virtual sequence that satisfies a specific condition and not to calculate a score, as described below.
- a virtual sequence that satisfies a specific condition may be, for example, a virtual sequence that is predicted from known knowledge to be unlikely to contribute to the occurrence of off-target effects.
- the score calculation unit 13 calculates a score for each of the multiple virtual sequences that is related to the probability that the DNA binding tool that recognizes the target sequence will act.
- the score calculation unit 13 may calculate the score using a value indicating the stability when the DNA binding tool is bound to each of the multiple virtual sequences.
- the virtual sequence generation unit 12 may calculate the score based on score calculation rules 22 stored in the memory unit 20.
- the score calculation rules 22 may be publicly known score calculation rules that apply an in silico analysis method developed according to the type of DNA binding tool. Furthermore, the score calculation rules 22 may be a combination of multiple rules, including publicly known calculation rules.
- CRISPR-Net is a different scoring tool from the MOFF score (Jiecong Lin et al., "CRISPR-Net: A Recurrent Convolutional Network Quantifies CRISPR Off-Target Activities with Mismatches and Indels", Advanced Science, Vol 7, 1903562, 2020) (https://doi.org/10.1002/advs.201903562).
- the DNA-binding tool is any of zinc finger nucleases (ZFNs), TALE nucleases (TALENs), and pentatricopeptide repeat (PPR) nucleases, it is possible to perform similar scoring by, for example, alignment taking into account mismatches and gaps and Tm value calculations, which can be performed by Biophython, etc.
- ZFNs zinc finger nucleases
- TALENs TALE nucleases
- PPR pentatricopeptide repeat
- the prediction result output unit 14 outputs a prediction result indicating the possibility that the DNA binding tool acts on a sequence different from the target sequence based on the calculated score.
- the prediction result output unit 14 may output an evaluation value calculated using all the scores calculated for each of the multiple virtual sequences as the prediction result.
- This evaluation value may be a value obtained by adding up all the scores calculated for each of the multiple virtual sequences.
- this value is not limited to the sum of all the scores, but refers to a calculation that can be expressed as an n-variable function f(s1, s2, ..., sn) for n virtual sequences.
- "s" indicates the score of each virtual sequence.
- This n-variable function f may include not only linear transformations but also nonlinear transformations based on a model generated by machine learning, and may also include a term indicating an error by the computer.
- the n-variable function f does not have to be a single-valued function, and may be a multi-valued function.
- the prediction result may be calculated so that the score calculated for a virtual sequence in which a mutation has been introduced so that the sequence is unlikely to be misrecognized by the DNA binding tool (i.e., a sequence that is unlikely to contribute to the occurrence of off-target effects) does not unduly affect the prediction result.
- the prediction result output unit 14 may output as the prediction result a value that is the sum of only the scores calculated for virtual sequences in which a mutation has been introduced so that the sequence is likely to contribute to the occurrence of off-target effects, among multiple virtual sequences.
- Fig. 5 is a flowchart showing an example of a process flow executed by the off-target risk analysis device 1.
- Fig. 5 also shows a process flow executed by an off-target risk analysis system 100 including the off-target risk analysis device 1.
- the target sequence receiving unit 11 receives a target sequence input by a user (step S1).
- the target sequence data indicating the target sequence may be, for example, text data.
- the virtual sequence generation unit 12 generates multiple virtual sequences including a sequence identical to the target sequence and a sequence in which at least one mutation has been introduced into the target sequence (step S2: virtual sequence generation step).
- the score calculation unit 13 calculates a score related to the probability that the DNA binding tool that recognizes the target sequence will act for each of the multiple virtual sequences generated in step S2 (step S3: score calculation step).
- the prediction result output unit 14 outputs a prediction result indicating the possibility that the DNA binding tool acts on a sequence different from the target sequence based on the score calculated in step S3 (step S4: prediction result output step).
- the off-target risk analysis device 1 generates multiple virtual sequences from a target sequence recognized by the DNA-binding tool, and calculates a score related to the probability that the DNA-binding tool will act for each of the multiple virtual sequences generated. The off-target risk analysis device 1 then uses the calculated score to output a prediction result indicating the possibility that the DNA-binding tool will act on a sequence different from the target sequence. This prediction result is, so to speak, information indicating the potential off-target risk of the DNA-binding tool.
- the off-target risk analysis device 1 can predict the potential off-target risks of a DNA-binding tool even if the genomic information of the target on which the DNA-binding tool is to be acted upon is unknown or uncertain.
- the off-target risk analysis device 1 is provided with an input unit 30 that accepts input of a target sequence by a user, and is configured to output a prediction result to a display device 4, but is not limited to this.
- an off-target risk analysis device 1a may be provided that is communicatively connected to communication devices 5a and 5b used by each user via a communication network 9.
- the off-target risk analysis device 1a receives target sequence data indicating the target sequence from each of the communication devices 5a and 5b. The off-target risk analysis device 1a then transmits to the communication device 5a a prediction result corresponding to the target sequence received from the communication device 5a, and transmits to the communication device 5b a prediction result corresponding to the target sequence received from the communication device 5b.
- FIG. 6 shows the off-target risk analysis system 100a including the communication devices 5a and 5b and the off-target risk analysis device 1a, the system is not limited to this. In the off-target risk analysis system 100a, the off-target risk analysis device 1a may be capable of communicating with two or more communication devices.
- Fig. 7 is a functional block diagram showing a configuration example of an off-target risk analysis system 100a according to one aspect of the present disclosure.
- the same reference numerals are attached to members having the same functions as those described in the above embodiment, and the description thereof will not be repeated.
- the off-target risk analysis device 1a includes a communication unit 16 that functions as a communication interface with the communication devices 5a and 5b.
- the target sequence receiving unit 11 receives target sequence data via the communication unit 16.
- the prediction result output unit 14 transmits the prediction results to each of the communication devices 5a and 5b via the communication unit 16.
- the off-target risk analysis device 1a may generate a web page showing the prediction results for each received target sequence for each target sequence, and provide information for accessing the web page to the user who sent the target sequence.
- communication device 5a and communication device 5b may be communication devices used by users who are pre-registered as users of off-target risk analysis system 100a.
- memory unit 20a may store user information 23 including information about users who are pre-registered as users of off-target risk analysis system 100a.
- FIG. 8 is a diagram showing an example of the data structure of user information 23.
- a user ID assigned to each user may be associated with the user's name, affiliation, and contact information.
- a user assigned user ID "U001” has the name "AA AA”, belongs to “XX University School of Medicine”, and has contact information (e.g., email address) "AAAA@xxx.xx.xx”.
- a user assigned user ID "U002” has the name "BB BB”, belongs to "YY Research Institute”, and has contact information "BBBB@yyy.yy.yyy".
- a prediction result regarding a target sequence received from a user with user ID "U001" is sent to "AAAA@xxx.xx.xx".
- the off-target risk analysis device 1a may store the prediction results for each received target sequence in the analysis result log 24 of the storage unit 20a.
- FIG. 9 is a diagram showing an example of the data structure of the analysis result log 24.
- the user ID assigned to each user may be associated with the target sequence data received from each user, the reception date and time, and the prediction results.
- the target sequence data received from user ID "U001" at "PM 1:50” on “2022/9/1" and the prediction results for the target sequence are stored in association with each other.
- the off-target risk analysis device 1a can provide prediction results obtained by analyzing target sequences received from each of multiple users to each user who is the sender of each target sequence data. For example, an administrator who manages the off-target risk analysis device 1a may charge each user (or each user's affiliated institution) a specified fee as compensation for the service of providing prediction results regarding the received target sequences.
- the functions of the off-target risk analysis device 1, 1a can be realized by a program for causing a computer to function as the device, and a program for causing a computer to function as each control block of the device (particularly each part included in the control unit 10, 10a).
- the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., a memory) as hardware for executing the program.
- control device e.g., a processor
- storage device e.g., a memory
- the program may be recorded on one or more computer-readable recording media, not on a temporary basis.
- the recording media may or may not be included in the device. In the latter case, the program may be supplied to the device via any wired or wireless transmission medium.
- each of the above control blocks can be realized by a logic circuit.
- the scope of this disclosure also includes an integrated circuit in which a logic circuit that functions as each of the above control blocks is formed.
- each process described in each of the above embodiments may be executed by AI (Artificial Intelligence).
- AI Artificial Intelligence
- the AI may run on the control device, or may run on another device (such as an edge computer or a cloud server).
- the off-target risk analysis method includes a virtual sequence generation step of generating a plurality of virtual sequences including a sequence identical to a target sequence and a sequence in which at least one mutation has been introduced into the target sequence; a score calculation step of calculating, for each of the plurality of virtual sequences, a score related to the probability that a DNA-binding tool that recognizes the target sequence will act; and a prediction result output step of outputting, based on the calculated score, a prediction result indicating the possibility that the DNA-binding tool will act on a sequence different from the target sequence.
- multiple virtual sequences are generated, including a sequence identical to the target sequence and a sequence in which at least one mutation has been introduced into the target sequence, and a score related to the probability that a DNA-binding tool that recognizes the target sequence will act is calculated for each of the multiple virtual sequences. Then, based on the calculated score, a prediction result is output that indicates the possibility that the DNA-binding tool will act on a sequence different from the target sequence.
- this off-target risk analysis method makes it possible to evaluate the risk of causing an off-target effect that a DNA-binding tool that recognizes a target sequence potentially has, without referring to the genomic sequence of the subject to which it is applied.
- the at least one mutation may be introduced into the entire target sequence or a part of the target sequence.
- the risk of off-target effects occurring in a portion of the target sequence is higher than the risk of off-target effects occurring in other portions.
- the risk of off-target effects may be assessed by focusing on that portion with a high risk of off-target effects.
- the virtual sequence generation step at least one mutation is introduced into the entire target sequence or into a part of the target sequence. This makes it possible to efficiently evaluate the risk of off-target effects.
- the mutation may be any one of a substitution, a deletion, and an insertion.
- the multiple virtual sequences may include (1) a sequence in which at least one adenine (A) in the target sequence is replaced with at least one of thymine (T), cytosine (C), and guanine (G), and/or (2) a sequence in which at least one thymine (T) in the target sequence is replaced with at least one of adenine (A), cytosine (C), and guanine (G), and/or (3) a sequence in which at least one cytosine (C) in the target sequence is replaced with at least one of adenine (A), thymine (T), and guanine (G), and/or (4) a sequence in which at least one guanine (G) in the target sequence is replaced with at least one of adenine (A), thymine (T), and cytosine (C).
- the above configuration makes it possible to comprehensively generate multiple virtual sequences, including sequences identical to the target sequence and sequences in which at least one substitution has been introduced into the target sequence. This allows for an unbiased evaluation of the potential risk of off-target effects that a DNA-binding tool that recognizes a target sequence has.
- each of the multiple virtual sequences may differ from the target sequence by four or less nucleotides.
- the DNA binding tool is less likely to act on sequences that have low homology to the target sequence recognized by the DNA binding tool. With the above configuration, it is possible to reduce the burden on computational resources while ensuring the accuracy of the prediction results.
- the score in any one of aspects 1 to 5 above, in the score calculation step, the score may be calculated using a value indicating the stability when the DNA-binding tool is bound to each of the multiple virtual sequences.
- the above configuration allows the score to be calculated with high accuracy for each of multiple virtual arrays.
- an evaluation value calculated using all of the scores calculated for each of the multiple virtual sequences may be output as the prediction result.
- the above configuration makes it possible to easily evaluate the risk of off-target effects that a DNA-binding tool that recognizes a target sequence potentially has, using the scores calculated for each of the multiple virtual sequences.
- the off-target risk analysis method in any one of aspects 1 to 7 above, may be such that the DNA-binding tool is (1) a zinc finger or a fusion polypeptide combining a zinc finger with a functional domain, (2) a TALE (Transcription Activator-Like Effector) or a fusion polypeptide combining a TALE with a functional domain, (3) a pentatricopeptide repeat or a fusion polypeptide combining a pentatricopeptide repeat with a functional domain, (4) a wild-type CRISPR/Cas (CRISPR-associated protein) or a fusion polypeptide-nucleic acid complex combining a wild-type CRISPR/Cas with a functional domain, or (5) a modified CRISPR/Cas (CRISPR-associated protein) or a fusion polypeptide-nucleic acid complex combining a modified CRISPR/Cas with a functional domain.
- CRISPR/Cas CRISPR-associated protein
- the off-target risk analysis system includes a virtual sequence generation unit that generates a plurality of virtual sequences including a sequence identical to a target sequence and a sequence in which at least one mutation has been introduced into the target sequence, a score calculation unit that calculates a score related to the probability that a DNA-binding tool that recognizes the target sequence will act for each of the plurality of virtual sequences, and a prediction result output unit that outputs a prediction result indicating the possibility that the DNA-binding tool will act on a sequence different from the target sequence based on the calculated score.
- the above configuration provides the same effect as aspect 1.
- the program according to aspect 10 of the present disclosure causes a computer to execute a virtual sequence generation step of generating a plurality of virtual sequences including a sequence identical to a target sequence and a sequence in which at least one mutation has been introduced into the target sequence, a score calculation step of calculating, for each of the plurality of virtual sequences, a score related to the probability that a DNA-binding tool that recognizes the target sequence will act, and a prediction result output step of outputting, based on the calculated score, a prediction result indicating the possibility that the DNA-binding tool will act on a sequence different from the target sequence.
- the recording medium according to aspect 11 of the present disclosure is a computer-readable recording medium having the program described in aspect 10 recorded thereon.
- Example 1 the first embodiment of the present disclosure will be described with reference to FIGS.
- the correlation between the prediction results output by the off-target risk analysis method according to one embodiment of the present disclosure and the off-target incidence rate calculated as described above was examined for 14 types of target sequences.
- virtual sequences in which four nucleotide substitutions were introduced into the target sequences were comprehensively generated.
- the MOFF score described in Non-Patent Document 3 was used as the score calculation tool used to output the prediction results.
- 10 to 12 are graphs showing the correlation between the prediction results output using the off-target risk analysis method according to the present disclosure and the off-target incidence rate obtained as a result of analyzing the off-target action in actual cells.
- the prediction results shown in FIG. 10 use prediction results output using scores calculated for a virtual sequence in which mutations are comprehensively introduced into the entire target sequence.
- the correlation between the prediction results and the off-target incidence rate is high ( R2 is 0.5675), demonstrating that the prediction accuracy of the prediction results output by the off-target risk analysis method according to one embodiment of the present disclosure is high.
- the prediction results shown in Figure 11 are based on prediction results output using only the scores calculated for the virtual sequence in which a mutation has been introduced into the non-seed region (the 8 nucleotides immediately preceding the PAM) with low specificity in the target sequence. As shown in Figure 11, it was found that the prediction accuracy was improved by outputting the prediction results using only the scores calculated for the virtual sequence in which a mutation has been introduced into the non-seed region with low specificity ( R2 is 0.57465).
- the prediction results shown in Figure 12 are obtained by comprehensively generating virtual sequences in which two nucleotide substitutions are introduced into the target sequence, calculating a score for each virtual sequence, and using the calculated scores to output the prediction results. As shown in Figure 12, even when the number of substitutions introduced to generate the virtual sequence was reduced from 4 to 2, the correlation between the prediction results and the off-target incidence rate decreased ( R2 was 0.5351), but the prediction accuracy of the prediction results remained high.
- Example 2 A second embodiment of the present disclosure will be described below with reference to FIGS.
- the virtual sequence used in this example is a base sequence in which a mismatch of up to one base pair is introduced into the entire target sequence.
- the base sequence used in the off-target analysis experiment in Non-Patent Document 3 and evaluated with the scoring tool CRISPR-Net was used as the target sequence.
- gRNA sequences used in the off-target analysis experiment “CHANGE-seq”
- 59 types used in the off-target analysis experiment “TTISS”
- 10 types used in the off-target analysis experiment “GUIDE-seq” in Non-Patent Document 3 were used.
- guide RNA (gRNA) sequences that satisfy the following condition I were selected and used.
- Genomic DNA sequences including PAM sequences can be extracted using the program "ExtendSeq.py” published at https://github.com/KazukiNakamae/Frame_Editor_sgRNA_selection.
- the selected gRNA sequences were 102 used in the off-target analysis experiment “CHANGE-seq”, 54 used in the off-target analysis experiment “TTISS”, and 8 used in the off-target analysis experiment “GUIDE-seq”.
- a virtual sequence was generated by introducing a mismatch of up to one base pair to the target sequence, which is a base sequence complementary to the selected gRNA sequence.
- a prediction result indicating the possibility that the DNA-binding tool will act on a sequence different from the target sequence was calculated by multiplying the sum of the CRISPR-Net scores by -1. That is, the forecast result is (-1) x ⁇ (CRISPR-Net score).
- a prediction result indicating the possibility that the DNA-binding tool will act on a sequence different from the sequence that matches the target sequence was calculated by multiplying the sum of the MOFF scores by -1. That is, the forecast result is (-1) x ⁇ (MOFF score).
- Figures 13 and 16 are scatter plots for the "CHANGE-seq” group
- Figures 14 and 17 are scatter plots for the "TTISS” group
- Figures 15 and 18 are scatter plots for the "GUIDE-seq” group.
- FIG. 13 is a scatter plot showing the correlation between the specificity prediction score calculated using the CRISPR-Net score and the "Off-on-ratio" score in the "CHANGE-seq” group.
- the Spearman correlation for the correlation between the specificity prediction score calculated using the CRISPR-Net score and the "Off-on-ratio” score in the "CHANGE-seq” group was -0.231839.
- Figure 14 is a scatter plot showing the correlation between the specificity prediction score calculated using the CRISPR-Net score and the "Off-on-ratio" score in the "TTISS" group.
- the Spearman correlation for the correlation between the specificity prediction score calculated using the CRISPR-Net score and the "Off-on-ratio” score in the "TTISS” group was -0.495973.
- Figure 15 is a scatter plot showing the correlation between the specificity prediction score calculated using the CRISPR-Net score and the "Off-on-ratio" score in the "GUIDE-seq” group. As shown in Figure 15, the Spearman correlation for the correlation between the specificity prediction score calculated using the CRISPR-Net score and the "Off-on-ratio" score in the "GUIDE-seq” group was -0.380952.
- [Correlation between specificity prediction score calculated using MOFF score and "Off-on-ratio” score] 16 is a scatter plot showing the correlation between the specificity prediction score calculated using the MOFF score and the "Off-on-ratio" score in the "CHANGE-seq” group. As shown in FIG. 16, the Spearman correlation for the correlation between the specificity prediction score calculated using the MOFF score and the "Off-on-ratio" score in the "CHANGE-seq” group was -0.566788.
- Figure 17 is a scatter plot showing the correlation between the specificity prediction score calculated using the MOFF score and the "Off-on-ratio" score in the "TTISS" group. As shown in Figure 17, the Spearman correlation for the correlation between the specificity prediction score calculated using the MOFF score and the "Off-on-ratio" score in the "TTISS" group was -0.655498.
- Figure 18 is a scatter plot showing the correlation between the specificity prediction score calculated using the MOFF score and the "Off-on-ratio" score in the "GUIDE-seq” group. As shown in Figure 18, the Spearman correlation for the correlation between the specificity prediction score calculated using the MOFF score and the "Off-on-ratio" score in the "GUIDE-seq” group was -0.761904.
- Target sequence receiving unit 12 Virtual sequence generating unit 13
- Score calculation unit 14 Prediction result output unit 100, 100a Off-target risk analysis system S2 Virtual sequence generating step S3 Score calculation step S4 Prediction result output step
Landscapes
- Physics & Mathematics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Chemical & Material Sciences (AREA)
- Analytical Chemistry (AREA)
- Biophysics (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Health & Medical Sciences (AREA)
- Engineering & Computer Science (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Biotechnology (AREA)
- Evolutionary Biology (AREA)
- General Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Theoretical Computer Science (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Abstract
ゲノム情報が未知あるいは不確定であっても、DNA結合性ツールの潜在的なオフターゲットリスクを予測する。オフターゲットリスク解析方法は、標的配列と同一の配列、および該標的配列に少なくとも1つの変異を導入した配列を含む複数の仮想配列を生成する仮想配列生成ステップ(S2)と、標的配列を認識するDNA結合性ツールが作用する確率に関連するスコアを、複数の仮想配列の各々について算出するスコア算出ステップ(S3)と、算出された前記スコアに基づいて、前記DNA結合性ツールが前記標的配列とは異なる配列に作用する可能性を示す予測結果を出力する予測結果出力ステップ(S4)と、を含む。
Description
本開示は、ゲノム編集およびその派生技術においてオフターゲット作用が生じるリスクを解析するオフターゲットリスク解析方法、オフターゲット予測システム等に関する。
ゲノム編集は、ゲノム配列上の標的領域のDNA配列(以下、標的配列)を認識して、当該標的領域を切断可能なDNA結合性ツールを用いて、任意の標的遺伝子の配列に欠失、置換、挿入等の変異を導入する技術である。DNA結合性ツールとしては、例えば、ジンクフィンガーヌクレアーゼ(ZFN:Zinc-Finger Nuclease)、TALEヌクレアーゼ(Transcription Activator-Like Effector Nuclease)、CRISPR(Clustered Regularly Interspaced Short Palindromic Repeats)/Cas(CRISPR-associated protein)等が知られている。
ゲノム編集によって改変された生物および細胞においては、ゲノム領域の標的ゲノム配列への変異導入に加えて、標的ではないゲノム配列に予期せぬ変異が導入される現象(所謂、オフターゲット作用)が生じることが懸念されている。オフターゲット作用が生じる頻度は、DNA結合性ツールによって認識されるDNA配列に依存するため、同じDNA結合性ツール(例えばCRISPR/Cas9)を用いた場合であっても、標的配列が異なれば、オフターゲット作用が生じる頻度は異なる。ゲノム編集において標的ゲノム配列ではない他のゲノム配列に予期せぬ変異が導入されてしまうリスクは、オフターゲットリスクとも呼称される。また、DNA結合性ツールを応用した様々な技術(例えば、転写調節、エピゲノム編集、染色体イメージング等)においても、同様にジンクフィンガー、TALE、およびヌクレアーゼ活性を不活化させたCRISPR/Cas等が用いられることから、標的配列に応じたオフターゲット作用のリスクが存在する。
非特許文献1~3に記載されているように、実験に用いるDNA結合性ツールの設計および選択を支援するために、様々なオフターゲットリスクの解析方法が開発されている。
Pawel Sledzinski et al., "Computational Tools and Resources Supporting CRISPR-Cas Experiments", Cells, Vol. 9, 1288, 2020.
X. Robert Bao et al., "Tools for experimental and computational analyses of off-target editing by programmable nucleases", Nature Protocol, Vol. 16, pp10-26, 2021.
Rongjie Fu et al., "Systematic decomposition of sequence determinants governing CRISPR/CAS9 specificity", Nature Communications, Vol.13, 474, 2022.
従来、オフターゲットリスクを、コンピュータを用いて予測する手法としては、例えば、以下の(1)および(2)が知られている。
(1)ゲノム編集等を行う前に、公共データベース上に記録された生物種、ウイルス株、および人種等の集団を代表する全ゲノム配列(リファレンスゲノム)を参照した相同性検索等によって、オフターゲット作用が生じる可能性のある候補配列を予測する方法。
(2)ゲノム編集等を行う前に、現実にゲノム編集等を受ける生物個体、組織、細胞クローン、品種、菌株、およびウイルス株の全ゲノムシーケンシングによって得た全ゲノム配列(以後、ユニークゲノムと呼称)を参照することで、オフターゲット作用が生じる可能性のある候補配列を予測する方法。
上記の(1)および(2)に示すオフターゲットリスクの予測方法はいずれも、全ゲノム情報が解読された生物種、あるいは特定の生物個体、組織、細胞クローン、品種、菌株、およびウイルス株等に限って利用できる方法である。それゆえ、ゲノム情報の解読が進んでいない、あるいはリピート構造等の難読な配列要素を含むためにゲノム解読が困難な生物種(産業生物も含まれる)にDNA結合性ツールを用いる場合のオフターゲットリスクの予測に適用することは困難であった。
また、全ゲノムが解読された生物であっても、そのゲノム情報は確定情報ではない。なぜならば、個体および細胞の各々において自然突然変異が生じている可能性があるからである。リファレンスゲノム、および事前に解読されたユニークゲノムの配列を参照しても、自然突然変異が生じている場合に表出する、DNA結合性ツールの潜在的なオフターゲット作用について予測することはできない。
このように、DNA結合性ツールを応用した様々な技術(例えば、DNA結合性ツールを応用した医療)においては、DNA結合性ツールの潜在的なオフターゲットリスクは大きなリスクとなり得るが、当該リスクを事前に予測できる手法はこれまでなかった。
本開示の一態様に係るオフターゲットリスク解析方法は、標的配列と同一の配列、および該標的配列に少なくとも1つの変異を導入した配列を含む複数の仮想配列を生成する仮想配列生成ステップと、前記標的配列を認識するDNA結合性ツールが作用する確率に関連するスコアを、前記複数の仮想配列の各々について算出するスコア算出ステップと、算出された前記スコアに基づいて、前記DNA結合性ツールが前記標的配列とは異なる配列に作用する可能性を示す予測結果を出力する予測結果出力ステップと、を含む。
本開示の一態様に係るオフターゲットリスク解析システムは、標的配列と同一の配列、および該標的配列に少なくとも1つの変異を導入した配列を含む複数の仮想配列を生成する仮想配列生成部と、前記標的配列を認識するDNA結合性ツールが作用する確率に関連するスコアを、前記複数の仮想配列の各々について算出するスコア算出部と、算出された前記スコアに基づいて、前記DNA結合性ツールが前記標的配列とは異なる配列に作用する可能性を示す予測結果を出力する予測結果出力部と、を備える。
本開示の各態様に係るオフターゲットリスク解析システムは、コンピュータによって実現してもよく、この場合には、コンピュータを前記オフターゲットリスク解析システムが備える各部(ソフトウェア要素)として動作させることにより前記オフターゲットリスク解析システムをコンピュータにて実現させるオフターゲットリスク解析システムの制御プログラム、およびそれを記録したコンピュータ読み取り可能な記録媒体も、本開示の範疇に入る。
ゲノム情報が未知あるいは不確定であっても、DNA結合性ツールの潜在的なオフターゲットリスクを予測可能なオフターゲットリスク解析方法、オフターゲットリスク解析システム等を提供する。
〔実施形態1〕
以下、本開示の一実施形態について、詳細に説明する。
以下、本開示の一実施形態について、詳細に説明する。
(オフターゲットリスク解析システム100の特徴)
本開示の実施形態1に係るオフターゲットリスク解析システム100は、標的配列を認識するDNA結合性ツールが、標的配列とは異なる配列に作用する可能性を示す予測結果を出力するシステムである。
本開示の実施形態1に係るオフターゲットリスク解析システム100は、標的配列を認識するDNA結合性ツールが、標的配列とは異なる配列に作用する可能性を示す予測結果を出力するシステムである。
DNA結合性ツールは、特定のゲノム領域に特異的に結合し、当該ゲノム領域の切断、改変、および編集等が可能なツールである。DNA結合性ツールによる切断、改変、および編集等を受けるゲノム領域は細胞のゲノム上の特定領域であってもよく、DNA結合性ツールはゲノム編集ツールとも呼称され得る。
標的配列を認識するDNA結合性ツールが、標的配列とは異なる配列に作用することは、「オフターゲット作用」と呼ばれる。例えば、細胞のゲノム上の遺伝子に対してオフターゲット作用が生じた結果、当該細胞において、がん遺伝子の活性化、およびがん抑制遺伝子の不活化等が起こる可能性がある。また、オフターゲット作用の影響は、細胞に対して永続的な効果をもたらす。それゆえ、標的配列を認識するDNA結合性ツールを実際に使用する前に、オフターゲット作用が生じるリスク(以下、オフターゲットリスク)を精度良く見積もることが望ましい。
オフターゲットリスク解析システム100は、下記の(a)~(c)の処理を実行する。
(a)DNA結合性ツールが認識する標的配列と同一の配列、および該標的配列に少なくとも1つの変異を導入した配列を含む複数の仮想配列を生成する。
(b)DNA結合性ツールが作用する確率に関連するスコアを、複数の仮想配列の各々について算出する。
(c)上記(b)において算出されたスコアに基づいて、DNA結合性ツールのオフターゲットリスクを示す予測結果を出力する。
(a)DNA結合性ツールが認識する標的配列と同一の配列、および該標的配列に少なくとも1つの変異を導入した配列を含む複数の仮想配列を生成する。
(b)DNA結合性ツールが作用する確率に関連するスコアを、複数の仮想配列の各々について算出する。
(c)上記(b)において算出されたスコアに基づいて、DNA結合性ツールのオフターゲットリスクを示す予測結果を出力する。
本開示において、「DNA結合性ツール」は、下記の(i)~(v)のいずれかであってもよい。
(i)ジンクフィンガー、またはジンクフィンガーに機能性ドメインを組み合わせた融合ポリペプチド。
(ii)TALE(Transcription Activator-Like Effector)、またはTALEに機能性ドメインを組み合わせた融合ポリペプチド。
(iii)ペンタトリコペプチドリピート(PPR:pentatricopeptide repeat)、またはペンタトリコペプチドリピートに機能性ドメインを組み合わせた融合ポリペプチド。
(iv)野生型のCRISPR/Cas(CRISPR-associated protein)、または野生型のCRISPR/Casに機能性ドメインを組み合わせた融合ポリペプチド-核酸複合体。
(v)改変型のCRISPR/Cas(CRISPR-associated protein)、または改変型のCRISPR/Casに機能性ドメインを組み合わせた融合ポリペプチド-核酸複合体。
(i)ジンクフィンガー、またはジンクフィンガーに機能性ドメインを組み合わせた融合ポリペプチド。
(ii)TALE(Transcription Activator-Like Effector)、またはTALEに機能性ドメインを組み合わせた融合ポリペプチド。
(iii)ペンタトリコペプチドリピート(PPR:pentatricopeptide repeat)、またはペンタトリコペプチドリピートに機能性ドメインを組み合わせた融合ポリペプチド。
(iv)野生型のCRISPR/Cas(CRISPR-associated protein)、または野生型のCRISPR/Casに機能性ドメインを組み合わせた融合ポリペプチド-核酸複合体。
(v)改変型のCRISPR/Cas(CRISPR-associated protein)、または改変型のCRISPR/Casに機能性ドメインを組み合わせた融合ポリペプチド-核酸複合体。
DNA結合性ツールが上記(i)である場合、標的配列は、ジンクフィンガータンパク質モチーフを有するドメインによって認識される塩基配列である。
DNA結合性ツールが上記(ii)である場合、標的配列は、TALモジュール(Transcription Activator-Like Module)を結合させた領域によって認識される塩基配列である。
DNA結合性ツールが上記(iii)である場合、標的配列は、PPRモチーフが連続する領域によって認識される塩基配列である。
DNA結合性ツールが上記(iv)および(v)のいずれかである場合、標的配列は、Casと複合体を形成しているガイドRNA(gRNA:guide RNA)と相補的な塩基配列、およびCasが認識するPAM(protospacer adjacent motif)配列である。
上記(i)~(v)のいずれかであるDNA結合性ツールは、標的配列と類似の塩基配列を有する目的外のDNAを誤認識する可能性がある。
非特許文献1~3に記載の解析方法はいずれも、全ゲノム情報が解読された生物種、あるいは特定の生物個体、組織、細胞クローン、品種、菌株、およびウイルス株などにDNA結合性ツールを用いる場合のオフターゲットリスクを解析可能である。しかし、非特許文献1~3に記載されている解析方法を、ゲノム情報の解読が進んでいない、あるいはゲノム解読が困難な生物種(産業生物も含まれる)にDNA結合性ツールを用いる場合のオフターゲットリスクの予測に適用することは困難である。
これに対し、本開示に係るオフターゲットリスク解析システム100は、標的配列さえ与えられれば、当該標的配列を認識するDNA結合性ツールのオフターゲットリスクを示す予測結果を出力することができる。すなわち、オフターゲットリスク解析システム100は、DNA結合性ツールを適用する対象のゲノム配列を参照することなく、標的配列を認識するDNA結合性ツールが潜在的に有する、オフターゲット作用を生じさせるリスクを評価することができる。
(オフターゲットリスク解析システム100の概略構成)
以下、オフターゲットリスク解析システム100の概略構成について、図1を用いて説明する。図1は、オフターゲットリスク解析システム100の概略構成の一例を示すブロック図である。
以下、オフターゲットリスク解析システム100の概略構成について、図1を用いて説明する。図1は、オフターゲットリスク解析システム100の概略構成の一例を示すブロック図である。
図1に示すように、オフターゲットリスク解析システム100は、オフターゲットリスク解析装置1および表示装置4を備えていてもよい。図1には、オフターゲットリスク解析装置1および表示装置4をそれぞれ1つずつ備えるオフターゲットリスク解析システム100を示している。しかし、オフターゲットリスク解析システム100の構成は、これに限定されない。例えば、オフターゲットリスク解析システム100の表示装置4の数は、0であってもよいし、複数であってもよい。
オフターゲットリスク解析システム100において、オフターゲットリスク解析装置1および表示装置4は、互いに通信可能に接続されている。オフターゲットリスク解析装置1および表示装置4は、直接、有線または無線で接続されていてもよいし、通信ネットワークを介して接続されていてもよい。通信ネットワークの形態は限定されるものではなく、ローカルエリアネットワーク(LAN)でもよいし、インターネットでもよい。
オフターゲットリスク解析装置1は、標的配列データを用いて、DNA結合性ツールのオフターゲットリスクを示す予測結果を出力する装置である。出力された予測結果は、オフターゲットリスク解析装置1から表示装置4へ送信されてもよい。
表示装置4は、典型的には、オフターゲットリスク解析システム100を利用するユーザが使用するコンピュータ、スマートフォン、タブレット端末等であってもよい。なお、図1には、表示装置4がオフターゲットリスク解析装置1と別体であるオフターゲットリスク解析システム100を示している。しかし、オフターゲットリスク解析システム100の構成は、これに限定されない。例えば、表示装置4は、オフターゲットリスク解析装置1と一体の装置であってもよく、この場合、表示装置4は、オフターゲットリスク解析装置1が備える表示部(ディスプレイ等)であってもよい。
(オフターゲットリスク解析装置1の構成)
次に、オフターゲットリスク解析装置1の構成について説明する。オフターゲットリスク解析装置1は、制御部10、記憶部20、および入力部30、を備えている。
次に、オフターゲットリスク解析装置1の構成について説明する。オフターゲットリスク解析装置1は、制御部10、記憶部20、および入力部30、を備えている。
制御部10は、一例において、CPU(Central Processing Unit)であってもよい。制御部10は、記憶部20に記憶されているソフトウェアである制御プログラムを読み取ってRAM(Random Access Memory)等のメモリに展開して各種機能を実行する。なお、図1に示す記憶部20では、説明の簡略化のために、制御プログラムの図示を省略している。
制御部10は、標的配列受付部11、仮想配列生成部12、スコア算出部13、および予測結果出力部14を備えている。
標的配列受付部11は、入力部30を用いて入力された標的配列を示す標的配列データを受け付ける。標的配列受付部11は、受け付けた標的配列データを記憶部20に格納してもよい。
仮想配列生成部12は、標的配列データから、標的配列と同一の配列、および該標的配列に少なくとも1つの変異を導入した配列を含む複数の仮想配列を生成する。すなわち、仮想配列生成部12は、標的配列を認識するDNA結合性ツールが、誤認識する可能性のある配列を含む多様な仮想配列データを仮想的に生成する。仮想配列生成部12は、記憶部20に格納されている仮想配列生成ルール21に基づいて、仮想配列を生成してもよい。
ここで、仮想配列生成部12が標的配列に導入する変異は、置換、欠失、および挿入のうちのいずれかであってもよい。
仮想配列生成部12によって生成される仮想配列について、図2~図4を用いて説明する。図2~図4は、仮想配列の例を示す図である。なお、図2~図4に示す仮想配列は、23塩基から成るヒトβグロビン遺伝子由来の標的配列から生成されている。ただし、本開示において、仮想配列は、23塩基よりも短くても長くてもよい。また、本開示において、標的配列は、ヒトβグロビン遺伝子由来のものに限定されない。また、本開示において、仮想配列は、連続した配列でなくてもよく、例えば、単一または複数の任意の塩基(N)を一か所または複数か所に挟む不連続な配列であってもよい。
まず、標的配列に置換が導入された仮想配列について図2を用いて説明する。図2には、標的配列に置換が導入された仮想配列の例を示している。仮想配列生成部12は、図2に示すように、標的配列と同一の配列、および標的配列に置換を導入した配列を含む複数の仮想配列を生成してもよい。
図2において、配列M1(配列番号1)は、標的配列と同一の配列である。配列M2~M7は、配列M1に含まれるヌクレオチドのうちの1つに置換を導入した配列の例を示している。例えば、配列M2(配列番号2)は、配列M1の5’末端の「A」が「T」に置換された配列であり、配列M3(配列番号3)は「G」に置換された配列であり、配列M4(配列番号4)は、「C」に置換された配列である。また、配列M5(配列番号5)は、配列M1の5’末端から2番目の「G」が「A」に置換された配列であり、配列M6(配列番号6)は、「T」に置換された配列であり、配列M7(配列番号7)は、「C」に置換された配列である。
図2には、標的配列の1つのヌクレオチドに対して置換を導入した仮想配列の例として配列M2~M7が示されているが、仮想配列生成部12によって生成される仮想配列は、これらに限定されない。仮想配列生成部12は、標的配列に1塩基置換を導入した配列を網羅的に生成してもよい。さらに、仮想配列生成部12は、標的配列の複数のヌクレオチドに置換を導入した仮想配列を生成してもよい。例えば、仮想配列生成部12は、標的配列に2塩基置換、3塩基置換、および4塩基置換を導入した配列を網羅的に生成してもよい。
すなわち、変異が置換である場合、仮想配列生成部12によって生成される複数の仮想配列は、下記の配列を含んでいてもよい。
・標的配列の少なくとも1つのアデニン(A)を、チミン(T)、シトシン(C)、およびグアニン(G)の少なくともいずれか1つに置換した配列、および/または、
・標的配列の少なくとも1つのチミン(T)を、アデニン(A)、シトシン(C)、およびグアニン(G)の少なくともいずれか1つに置換した配列、および/または、
・標的配列の少なくとも1つのシトシン(C)を、アデニン(A)、チミン(T)、およびグアニン(G)の少なくともいずれか1つに置換した配列、および/または、
・標的配列の少なくとも1つのグアニン(G)を、アデニン(A)、チミン(T)、およびシトシン(C)の少なくともいずれか1つに置換した配列。
・標的配列の少なくとも1つのアデニン(A)を、チミン(T)、シトシン(C)、およびグアニン(G)の少なくともいずれか1つに置換した配列、および/または、
・標的配列の少なくとも1つのチミン(T)を、アデニン(A)、シトシン(C)、およびグアニン(G)の少なくともいずれか1つに置換した配列、および/または、
・標的配列の少なくとも1つのシトシン(C)を、アデニン(A)、チミン(T)、およびグアニン(G)の少なくともいずれか1つに置換した配列、および/または、
・標的配列の少なくとも1つのグアニン(G)を、アデニン(A)、チミン(T)、およびシトシン(C)の少なくともいずれか1つに置換した配列。
次に、標的配列に欠失が導入された仮想配列について図3を用いて説明する。図3には、標的配列に欠失が導入された仮想配列の例を示している。仮想配列生成部12は、図3に示すように、標的配列と同一の配列、および標的配列に欠失を導入した配列を含む複数の仮想配列を生成してもよい。
図3において、配列M1(配列番号1)は、標的配列と同一の配列である。配列M8~M10は、配列M1に含まれるヌクレオチドのうちの1つに欠失を導入した配列の例を示している。例えば、配列M8(配列番号8)は、配列M1の5’末端の「A」が欠失した配列であり、配列M9(配列番号9)は、配列M1の5’末端から2番目の「G」が欠失した配列であり、配列M10(配列番号10)は、配列M1の5’末端から3番目の「C」が欠失した配列である。
図3には、標的配列の1つのヌクレオチドに欠失を導入した仮想配列の例として配列M8~M10が示されているが、仮想配列生成部12によって生成される仮想配列は、これらに限定されない。仮想配列生成部12は、標的配列に1塩基欠失を導入した配列を導入した配列を網羅的に生成してもよい。さらに、仮想配列生成部12は、標的配列の複数のヌクレオチドに欠失を導入した仮想配列を生成してもよい。例えば、仮想配列生成部12は、標的配列に2塩基欠失、3塩基欠失、および4塩基欠失を導入した配列を網羅的に生成してもよい。
続いて、標的配列に挿入が導入された仮想配列について図4を用いて説明する。図4には、標的配列に挿入が導入された仮想配列の例を示している。仮想配列生成部12は、図4に示すように、標的配列と同一の配列、および標的配列に挿入を導入した配列を含む複数の仮想配列を生成してもよい。
図4において、配列M1(配列番号1)は、標的配列と同一の配列である。配列M11~M18は、配列M1に1塩基挿入を導入した配列の例を示している。例えば、配列M11(配列番号11)は、配列M1の5’末端および2番目のヌクレオチド「AG」の間に「A」が挿入された配列であり、配列M12(配列番号12)は「T」が挿入された配列であり、配列M13(配列番号13)は、「G」が挿入された配列であり、配列M14(配列番号14)は、「C」が挿入された配列である。また、配列M15(配列番号15)は、配列M1の5’末端から2番目および3番目のヌクレオチド「GC」の間に「A」が挿入された配列であり、配列M16(配列番号6)は、「T」が挿入された配列であり、配列M17(配列番号17)は、「G」が挿入された配列であり、配列M18(配列番号18)は、「C」が挿入された配列である。
図4には、標的配列中の1か所に挿入を導入した仮想配列の例として配列M11~M18が示されているが、仮想配列生成部12によって生成される仮想配列は、これらに限定されない。例えば、仮想配列生成部12は、標的配列中の1か所に2塩基挿入を導入した配列を生成してもよい。また、仮想配列生成部12は、標的配列中の1か所に挿入を導入した配列を網羅的に生成してもよい。さらに、仮想配列生成部12は、標的配列の複数の箇所に挿入を導入した仮想配列を生成してもよい。例えば、仮想配列生成部12は、標的配列中の2か所、3か所、および4か所に挿入を導入した配列を網羅的に生成してもよい。
仮想配列生成部12は、標的配列に2種類以上の変異を導入した仮想配列を生成してもよい。例えば、仮想配列生成部12は、標的配列に置換を導入した仮想配列と、標的配列に欠失を導入した仮想配列と、標的配列に挿入を導入した仮想配列とを生成してもよい。
DNA結合性ツールによって認識される標的配列との相同性が低い配列に対しては、当該DNA結合性ツールが作用する可能性は低い傾向がある。それゆえ、複数の仮想配列の各々は、標的配列と4つ以下のヌクレオチドが異なっていればよく、それ以上の変異を導入した仮想配列を生成する必要性は低い。これにより、予測結果の精度を担保しつつ、計算資源への負担を軽減させることができる。
オフターゲットリスク解析装置1は、特定の条件を満たす仮想配列の生成と、後述するスコアの算出とを行わない構成であってもよい。ここで、特定の条件を満たす仮想配列とは、例えば、オフターゲット作用の発生に寄与する可能性が低いことが既知の知見から予期される仮想配列であってもよい。このような構成を採用すれば、オフターゲット作用の発生に寄与する可能性が低い仮想配列に関するスコアが、予測結果に過度に影響することを防ぎ、出力される予測結果の精度を向上させることができる。
図1に戻り、スコア算出部13は、標的配列を認識するDNA結合性ツールが作用する確率に関連するスコアを、複数の仮想配列の各々について算出する。スコア算出部13は、複数の仮想配列の各々にDNA結合性ツールを結合させた場合の安定性を示す値を用いてスコアを算出してもよい。仮想配列生成部12は、記憶部20に格納されているスコア算出ルール22に基づいて、スコアを算出してもよい。スコア算出ルール22は、DNA結合性ツールの種類に応じて開発されたin silico解析手法を適用した公知のスコア算出ルールであってもよい。また、スコア算出ルール22は公知の算出ルールを含めた複数を組み合わせであってもよい。
DNA結合性ツールがCRISPR/Cas9である場合、例えば、MOFFスコア(非特許文献3を参照)、CRISPR-Netスコア、およびCFDスコア等のスコアを算出するためのスコアリングツールと同様のスコア算出ルール22が適用可能である。例えば、CRISPR-Netは、MOFFスコアとは別種のスコアリングツールである(Jiecong Lin et al., “CRISPR-Net: A Recurrent Convolutional Network Quantifies CRISPR Off-Target Activities with Mismatches and Indels”, Advanced Science, Vol 7, 1903562, 2020)(https://doi.org/10.1002/advs.201903562)。DNA結合性ツールがジンクフィンガーヌクレアーゼ(ZFN)、TALEヌクレアーゼ(TALEN)、およびペンタトリコペプチドリピート(PPR)ヌクレアーゼのうちのいずれかである場合、例えば、Biоphythоn等で実施可能なミスマッチ・ギャップを考慮したアライメントとTm値計算等により同様にスコアリングを実施することが可能である。
予測結果出力部14は、算出されたスコアに基づいて、DNA結合性ツールが標的配列とは異なる配列に作用する可能性を示す予測結果を出力する。予測結果出力部14は、複数の仮想配列の各々について算出されたすべてのスコアを用いて算出される評価値を予測結果として出力してもよい。この評価値は、複数の仮想配列の各々について算出されたすべてのスコアを合算した値であってもよい。また、この値は、すべてのスコアの総和に限定されず、n個の仮想配列に対してn変数関数f(s1,s2,・・・,sn)として表現できる計算を指す。ここで、「s」は、各仮想配列のスコアを示す。このn変数関数fは、線型変換だけでなく、機械学習によって生成されるモデルによる非線型変換を含んでもよく、また計算機による誤差を示す項を含んでもよい。また、n変数関数fは一価関数でなくてもよく、多価関数であってもよい。
DNA結合性ツールが誤認識を起こしにくい配列(すなわち、オフターゲット作用の発生に寄与する可能性の低い配列)となるように変異を導入した仮想配列について算出されたスコアが予測結果に過度に影響しないように構成してもよい。この場合、予測結果出力部14は、複数の仮想配列のうち、オフターゲット作用の発生に寄与する可能性の高い配列となるように変異を導入した仮想配列について算出されたスコアのみを合算した値を予測結果として出力してもよい。
(オフターゲットリスク解析装置が実行する処理の流れ)
次に、オフターゲットリスク解析装置が実行する処理の流れについて、図5を用いて説明する。図5は、オフターゲットリスク解析装置1が実行する処理の流れの一例を示すフローチャートである。図5は、オフターゲットリスク解析装置1を備えるオフターゲットリスク解析システム100が実行する処理の流れでもある。
次に、オフターゲットリスク解析装置が実行する処理の流れについて、図5を用いて説明する。図5は、オフターゲットリスク解析装置1が実行する処理の流れの一例を示すフローチャートである。図5は、オフターゲットリスク解析装置1を備えるオフターゲットリスク解析システム100が実行する処理の流れでもある。
まず、標的配列受付部11は、ユーザによる標的配列の入力を受け付ける(ステップS1)。標的配列を示す標的配列データは、例えば、テキストデータであってもよい。
次に、仮想配列生成部12は、標的配列と同一の配列、および該標的配列に少なくとも1つの変異を導入した配列を含む複数の仮想配列を生成する(ステップS2:仮想配列生成ステップ)。
続いて、スコア算出部13は、標的配列を認識するDNA結合性ツールが作用する確率に関連するスコアを、ステップS2において生成された複数の仮想配列の各々について算出する(ステップS3:スコア算出ステップ)。
予測結果出力部14は、ステップS3において算出されたスコアに基づいて、DNA結合性ツールが標的配列とは異なる配列に作用する可能性を示す予測結果を出力する(ステップS4:予測結果出力ステップ)。
上記の構成によれば、オフターゲットリスク解析装置1は、複数の仮想配列を、DNA結合性ツールが認識する標的配列から生成し、生成した複数の仮想配列の各々について、DNA結合性ツールが作用する確率に関連するスコアを算出する。そして、オフターゲットリスク解析装置1は、算出したスコアを用いて、当該DNA結合性ツールが、標的配列とは異なる配列に作用する可能性を示す予測結果を出力する。この予測結果は、いわば、当該DNA結合性ツールの潜在的なオフターゲットリスクを示す情報である。
このように、オフターゲットリスク解析装置1は、DNA結合性ツールを作用させる対象のゲノム情報が未知あるいは不確定であっても、DNA結合性ツールの潜在的なオフターゲットリスクを予測することができる。
〔実施形態2〕
本開示の他の実施形態について、以下に説明する。
本開示の他の実施形態について、以下に説明する。
図1に示すオフターゲットリスク解析システム100は、オフターゲットリスク解析装置1が、ユーザによる標的配列の入力を受け付ける入力部30を備えており、予測結果を表示装置4に出力する構成であったが、これに限定されない。例えば、図6に示すオフターゲットリスク解析システム100aのように、通信ネットワーク9を介して各ユーザが使用する通信装置5a、5bと通信可能に接続されているオフターゲットリスク解析装置1aを備えていてもよい。
図6に示すオフターゲットリスク解析システム100aでは、オフターゲットリスク解析装置1aは、通信装置5a、5bのそれぞれから、標的配列を示す標的配列データを受信する。そして、オフターゲットリスク解析装置1aは、通信装置5aから受け付けた標的配列に対応する予測結果を通信装置5aに送信し、通信装置5bから受け付けた標的配列に対応する予測結果を通信装置5bに送信する。なお、図6は、通信装置5a、5bとオフターゲットリスク解析装置1aとを含むオフターゲットリスク解析システム100aを示しているが、これに限定されない。オフターゲットリスク解析システム100aにおいて、オフターゲットリスク解析装置1aは2以上の通信装置と通信可能であってもよい。
(オフターゲットリスク解析装置1aの構成)
オフターゲットリスク解析装置1aの構成について、図7を用いて説明する。図7は、本開示の一態様に係るオフターゲットリスク解析システム100aの構成例を示す機能ブロック図である。なお、説明の便宜上、上記実施形態にて説明した部材と同じ機能を有する部材については、同じ符号を付記し、その説明を繰り返さない。
オフターゲットリスク解析装置1aの構成について、図7を用いて説明する。図7は、本開示の一態様に係るオフターゲットリスク解析システム100aの構成例を示す機能ブロック図である。なお、説明の便宜上、上記実施形態にて説明した部材と同じ機能を有する部材については、同じ符号を付記し、その説明を繰り返さない。
図7に示すように、オフターゲットリスク解析装置1aは、通信装置5a、5bとの通信インターフェースとして機能する通信部16を備えている。標的配列受付部11は、通信部16を介して、標的配列データを受け付ける。予測結果出力部14は、通信部16を介して、予測結果を通信装置5a、5bのそれぞれに送信する。なお、オフターゲットリスク解析装置1aは、受け付けた標的配列に関する予測結果を示したウェブページを標的配列毎に生成し、当該標的配列の送信元であるユーザに当該ウェブページにアクセスするための情報を提供してもよい。
ここで、通信装置5aおよび通信装置5bは、オフターゲットリスク解析システム100aの利用者として予め登録されているユーザが使用する通信装置であってもよい。この場合、記憶部20aには、オフターゲットリスク解析システム100aの利用者として予め登録されているユーザに関する情報を含むユーザ情報23が格納されていてもよい。
図8は、ユーザ情報23のデータ構造の一例を示す図である。ユーザ情報23には、各ユーザに付与されたユーザIDと、ユーザの氏名、所属、および連絡先とが、対応付けられていてもよい。図8において、ユーザID「U001」が付与されたユーザは、氏名が「AA AA」であり、「XX大学医学部」に所属しており、連絡先(例えば、メールアドレス)は「AAAA@xxx.xx.xx」である。ユーザID「U002」が付与されたユーザは、氏名が「BB BB」であり、「YY研究所」に所属しており、連絡先は「BBBB@yyy.yy.yy」である。例えば、ユーザID「U001」のユーザから受け付けた標的配列に関する予測結果は、「AAAA@xxx.xx.xx」に送信される。
オフターゲットリスク解析装置1aは、受け付けた標的配列毎の予測結果を、記憶部20aの解析結果ログ24に格納してもよい。図9は、解析結果ログ24のデータ構造の一例を示す図である。解析結果ログ24には、各ユーザに付与されたユーザIDと、各ユーザから受け付けた標的配列データ、受付日時、および予測結果とが、対応付けられていてもよい。図9において、ユーザID「U001」から「2022/9/1」の「PM1:50」に受け付けた標的配列データ、および当該標的配列に関する予測結果が対応付けて格納されている。
このような構成を採用すれば、オフターゲットリスク解析装置1aは、複数のユーザのそれぞれから受け付けた標的配列を解析して得た予測結果を、各標的配列データの送信元である各ユーザに提供することができる。例えば、オフターゲットリスク解析装置1aを管理する管理者は、受け付けた標的配列に関する予測結果を提供するサービスに対する対価として、各ユーザ(もしくは各ユーザの所属機関)に対して所定の料金を請求してもよい。
〔ソフトウェアによる実現例〕
オフターゲットリスク解析装置1、1a(以下、「装置」と呼ぶ)の機能は、当該装置としてコンピュータを機能させるためのプログラムであって、当該装置の各制御ブロック(特に制御部10、10aに含まれる各部)としてコンピュータを機能させるためのプログラムにより実現することができる。
オフターゲットリスク解析装置1、1a(以下、「装置」と呼ぶ)の機能は、当該装置としてコンピュータを機能させるためのプログラムであって、当該装置の各制御ブロック(特に制御部10、10aに含まれる各部)としてコンピュータを機能させるためのプログラムにより実現することができる。
この場合、上記装置は、上記プログラムを実行するためのハードウェアとして、少なくとも1つの制御装置(例えばプロセッサ)と少なくとも1つの記憶装置(例えばメモリ)を有するコンピュータを備えている。この制御装置と記憶装置により上記プログラムを実行することにより、上記各実施形態で説明した各機能が実現される。
上記プログラムは、一時的ではなく、コンピュータ読み取り可能な、1または複数の記録媒体に記録されていてもよい。この記録媒体は、上記装置が備えていてもよいし、備えていなくてもよい。後者の場合、上記プログラムは、有線または無線の任意の伝送媒体を介して上記装置に供給されてもよい。
また、上記各制御ブロックの機能の一部または全部は、論理回路により実現することも可能である。例えば、上記各制御ブロックとして機能する論理回路が形成された集積回路も本開示の範疇に含まれる。この他にも、例えば量子コンピュータにより上記各制御ブロックの機能を実現することも可能である。
また、上記各実施形態で説明した各処理は、AI(Artificial Intelligence:人工知能)に実行させてもよい。この場合、AIは上記制御装置で動作するものであってもよいし、他の装置(例えばエッジコンピュータまたはクラウドサーバ等)で動作するものであってもよい。
本開示は上述した各実施形態に限定されるものではなく、請求項に示した範囲で種々の変更が可能であり、異なる実施形態にそれぞれ開示された技術的手段を適宜組み合わせて得られる実施形態についても本開示の技術的範囲に含まれる。
〔まとめ〕
本開示の態様1に係るオフターゲットリスク解析方法は、標的配列と同一の配列、および該標的配列に少なくとも1つの変異を導入した配列を含む複数の仮想配列を生成する仮想配列生成ステップと、前記標的配列を認識するDNA結合性ツールが作用する確率に関連するスコアを、前記複数の仮想配列の各々について算出するスコア算出ステップと、算出された前記スコアに基づいて、前記DNA結合性ツールが前記標的配列とは異なる配列に作用する可能性を示す予測結果を出力する予測結果出力ステップと、を含む。
本開示の態様1に係るオフターゲットリスク解析方法は、標的配列と同一の配列、および該標的配列に少なくとも1つの変異を導入した配列を含む複数の仮想配列を生成する仮想配列生成ステップと、前記標的配列を認識するDNA結合性ツールが作用する確率に関連するスコアを、前記複数の仮想配列の各々について算出するスコア算出ステップと、算出された前記スコアに基づいて、前記DNA結合性ツールが前記標的配列とは異なる配列に作用する可能性を示す予測結果を出力する予測結果出力ステップと、を含む。
上記のオフターゲットリスク解析方法では、標的配列と同一の配列、および標的配列に少なくとも1つの変異を導入した配列を含む複数の仮想配列を生成し、標的配列を認識するDNA結合性ツールが作用する確率に関連するスコアを複数の仮想配列の各々について算出する。そして、算出されたスコアに基づいて、DNA結合性ツールが標的配列とは異なる配列に作用する可能性を示す予測結果を出力する。
上記のオフターゲットリスク解析方法を用いれば、標的配列が与えられれば、当該標的配列を認識するDNA結合性ツールが、標的配列とは異なる配列に作用する可能性を示す予測結果を出力することができる。すなわち、このオフターゲットリスク解析方法によれば、適用する対象のゲノム配列を参照することなく、標的配列を認識するDNA結合性ツールが潜在的に有する、オフターゲット作用を生じさせるリスクを評価することができる。
本開示の態様2に係るオフターゲットリスク解析方法は、上記態様1において、前記仮想配列生成ステップにおいて、前記少なくとも1つの変異は、前記標的配列の全体、または該標的配列の一部分に対して導入されてもよい。
DNA結合性ツールの中には、標的配列の一部分にオフターゲット作用が生じるリスクが、他の部分にオフターゲット作用が生じるリスクよりも高いことが知られている場合がある。この場合、オフターゲット作用が生じるリスクが高い当該一部分に着目して、オフターゲット作用が生じるリスクを評価してもよい。
上記の構成によれば、仮想配列生成ステップにおいて、少なくとも1つの変異は、標的配列の全体、または該標的配列の一部分に対して導入される。これにより、オフターゲット作用が生じるリスクを効率的に評価することができる。
本開示の態様3に係るオフターゲットリスク解析方法は、上記態様1または2において、前記変異は、置換、欠失、および挿入のうちのいずれかであってもよい。
本開示の態様4に係るオフターゲットリスク解析方法は、上記態様3において、前記変異が置換である場合、前記複数の仮想配列は、(1)前記標的配列の少なくとも1つのアデニン(A)を、チミン(T)、シトシン(C)、およびグアニン(G)の少なくともいずれか1つに置換した配列、および/または、(2)前記標的配列の少なくとも1つのチミン(T)を、アデニン(A)、シトシン(C)、およびグアニン(G)の少なくともいずれか1つに置換した配列、および/または、(3)前記標的配列の少なくとも1つのシトシン(C)を、アデニン(A)、チミン(T)、およびグアニン(G)の少なくともいずれか1つに置換した配列、および/または、(4)前記標的配列の少なくとも1つのグアニン(G)を、アデニン(A)、チミン(T)、およびシトシン(C)の少なくともいずれか1つに置換した配列、を含んでいてもよい。
上記の構成によれば、標的配列と同一の配列、および標的配列に少なくとも1つの置換を導入した配列を含む複数の仮想配列を網羅的に生成し得る。これにより、標的配列を認識するDNA結合性ツールが潜在的に有する、オフターゲット作用を生じさせるリスクを偏り無く評価することができる。
本開示の態様5に係るオフターゲットリスク解析方法は、上記態様1~4の何れかにおいて、前記仮想配列生成ステップにおいて、前記複数の仮想配列の各々は、前記標的配列と4つ以下のヌクレオチドが異なっていてもよい。
DNA結合性ツールによって認識される標的配列との相同性が低い配列に対しては、当該DNA結合性ツールが作用する可能性は低い傾向がある。上記の構成によれば、予測結果の精度を担保しつつ、計算資源への負担を軽減させることができる。
本開示の態様6に係るオフターゲットリスク解析方法は、上記態様1~5の何れかにおいて、前記スコア算出ステップにおいて、前記スコアは、前記複数の仮想配列の各々に前記DNA結合性ツールを結合させた場合の安定性を示す値を用いて算出されてもよい。
上記の構成によれば、スコアを複数の仮想配列の各々について精度良く算出することができる。
本開示の態様7に係るオフターゲットリスク解析方法は、上記態様1~6の何れかにおいて、前記予測結果出力ステップにおいて、前記複数の仮想配列の各々について算出されたすべての前記スコアを用いて算出される評価値を前記予測結果として出力してもよい。
上記の構成によれば、複数の仮想配列の各々について算出されたスコアを用いて、標的配列を認識するDNA結合性ツールが潜在的に有する、オフターゲット作用を生じさせるリスクを簡易に評価することができる。
本開示の態様8に係るオフターゲットリスク解析方法は、上記態様1~7の何れかにおいて、前記DNA結合性ツールは、(1)ジンクフィンガー、またはジンクフィンガーに機能性ドメインを組み合わせた融合ポリペプチド、(2)TALE(Transcription Activator-Like Effector)、またはTALEに機能性ドメインを組み合わせた融合ポリペプチド、(3)ペンタトリコペプチドリピート(pentatricopeptide repeat)、またはペンタトリコペプチドリピートに機能性ドメインを組み合わせた融合ポリペプチド、(4)野生型のCRISPR/Cas(CRISPR-associated protein)、または野生型のCRISPR/Casに機能性ドメインを組み合わせた融合ポリペプチド-核酸複合体、または、(5)改変型のCRISPR/Cas(CRISPR-associated protein)、または改変型のCRISPR/Casに機能性ドメインを組み合わせた融合ポリペプチド-核酸複合体、であってもよい。
本開示の態様9に係るオフターゲットリスク解析システムは、標的配列と同一の配列、および該標的配列に少なくとも1つの変異を導入した配列を含む複数の仮想配列を生成する仮想配列生成部と、前記標的配列を認識するDNA結合性ツールが作用する確率に関連するスコアを、前記複数の仮想配列の各々について算出するスコア算出部と、算出された前記スコアに基づいて、前記DNA結合性ツールが前記標的配列とは異なる配列に作用する可能性を示す予測結果を出力する予測結果出力部と、を備える。上記構成によれば、態様1と同様の効果を奏する。
本開示の態様10に係るプログラムは、コンピュータに、標的配列と同一の配列、および該標的配列に少なくとも1つの変異を導入した配列を含む複数の仮想配列を生成する仮想配列生成ステップと、前記標的配列を認識するDNA結合性ツールが作用する確率に関連するスコアを、前記複数の仮想配列の各々について算出するスコア算出ステップと、算出された前記スコアに基づいて、前記DNA結合性ツールが前記標的配列とは異なる配列に作用する可能性を示す予測結果を出力する予測結果出力ステップと、を実行させる。
本開示の態様11に係る記録媒体は、上記態様10に記載のプログラムを記録したコンピュータ読取り可能な記録媒体である。
(実施例1)
以下、本開示の実施例1を図10~図12を用いて説明する。
以下、本開示の実施例1を図10~図12を用いて説明する。
実験的手法を用いて、ヒトゲノム由来の配列を標的配列とするCRISPR/Cas9のオフターゲット候補配列をヒトゲノム全体に亘り検索した結果がいくつか報告されている。(例えば、Lamsfus-Celle,A. et al., Scientific Reports, 10, 10133.2020)。実験的手法を用いた検索の結果を利用して、CRISPR/Cas9のオフターゲット発生率を、総リード数におけるオフターゲットリード数が占める割合として算出可能である。
そこで、本実施例では、本開示の一態様に係るオフターゲットリスク解析方法によって出力された予測結果と、上記のように算出したオフターゲット発生率との相関関係を、14種の標的配列について調べた。なお、予測結果の出力するために、標的配列に4ヌクレオチドの置換を導入した仮想配列を網羅的に生成した。また、予測結果を出力するために用いるスコア算出ツールとして、非特許文献3に記載のMOFFスコアを用いた。
図10~図12は、本開示に係るオフターゲットリスク解析方法を用いて出力された予測結果と、実際の細胞内でのオフターゲット作用を解析した結果得られたオフターゲット発生率との相関関係を示すグラフである。図10に示す予測結果は、標的配列の全体に網羅的に変異が導入された仮想配列について算出されたスコアを用いて出力された予測結果を用いている。図10に示すように、予測結果とオフターゲット発生率との相関関係は高く(R2は0.5675)、本開示の一態様に係るオフターゲットリスク解析方法によって出力された予測結果の予測精度が高いことが実証できた。
図11に示す予測結果は、標的配列の中で、特異性の低い非シード領域(PAMの直前の8ヌクレオチド)に変異が導入された仮想配列について算出されたスコアのみを用いて出力した予測結果を用いている。図11に示すように、特異性の低い非シード領域に変異が導入された仮想配列について算出されたスコアのみを用いて予測結果を出力することで、予測精度が向上することが分かった(R2は0.57465)。
図12に示す予測結果は、予測結果の出力するために、標的配列に2ヌクレオチドの置換を導入した仮想配列を網羅的に生成し、各仮想配列についてスコアを算出し、算出されたスコアを用いて出力された予測結果を用いている。図12に示すように、仮想配列を生成するために導入する置換の数を4から2に減らしても、予測結果とオフターゲット発生率との相関関係の低下は見られるものの(R2は0.5351)、予測結果の予測精度は高いままであった。
(実施例2)
以下、本開示の実施例2を図13~図18を用いて説明する。
以下、本開示の実施例2を図13~図18を用いて説明する。
[標的配列]
本実施例で用いた仮想配列は、標的配列の全体に対して最大1塩基対までのミスマッチを導入した塩基配列である。非特許文献3においてオフターゲット解析実験に利用され、かつスコアリングツールCRISPR-Netでの評価実績のある塩基配列を標的配列とした。
本実施例で用いた仮想配列は、標的配列の全体に対して最大1塩基対までのミスマッチを導入した塩基配列である。非特許文献3においてオフターゲット解析実験に利用され、かつスコアリングツールCRISPR-Netでの評価実績のある塩基配列を標的配列とした。
具体的には、非特許文献3において、オフターゲット解析実験「CHANGE-seq」に使われている108種、オフターゲット解析実験「TTISS」に使われている59種、および、オフターゲット解析実験「GUIDE-seq」に使われている10種のgRNA配列を用いた。具体的には、これらのgRNA配列から、下記の条件Iを満たすガイドRNA(gRNA)配列を選定して用いた。
条件I:https://github.com/KazukiNakamae/Frame_Editor_sgRNA_selectionにて公開されているプログラムである「ExtendSeq.py」を用いてPAM配列を含むゲノムDNA配列が抽出可能。
[仮想配列]
選定されたgRNA配列は、オフターゲット解析実験「CHANGE-seq」に使われた102種、オフターゲット解析実験「TTISS」に使われた54種、オフターゲット解析実験「GUIDE-seq」に使われた8種であった。選定されたgRNA配列と相補的な塩基配列である標的配列に対して、最大1塩基対までのミスマッチを導入し仮想配列を生成した。
選定されたgRNA配列は、オフターゲット解析実験「CHANGE-seq」に使われた102種、オフターゲット解析実験「TTISS」に使われた54種、オフターゲット解析実験「GUIDE-seq」に使われた8種であった。選定されたgRNA配列と相補的な塩基配列である標的配列に対して、最大1塩基対までのミスマッチを導入し仮想配列を生成した。
[スコアリングツールCRISPR-Netによるスコア算出]
スコアリングツールCRISPR-Netによるスコア算出は、データ解析プラットフォーム「Code Ocean」(https://codeocean.com)上で実施した。具体的には、CRISPR-Netの実行スペース(https://codeocean.com/capsule/9553651/tree/v1)のスクリプトのコピーを作成し、生成した仮想配列に対するCRISPR-Netスコアを自動算出できるように実行スクリプト「run.sh」および「CRISPR_Net.py」に書き加えた。
スコアリングツールCRISPR-Netによるスコア算出は、データ解析プラットフォーム「Code Ocean」(https://codeocean.com)上で実施した。具体的には、CRISPR-Netの実行スペース(https://codeocean.com/capsule/9553651/tree/v1)のスクリプトのコピーを作成し、生成した仮想配列に対するCRISPR-Netスコアを自動算出できるように実行スクリプト「run.sh」および「CRISPR_Net.py」に書き加えた。
[CRISPR-Netスコアを用いた特異性予測スコア(予測結果)]
上記のように生成された仮想配列の生成に用いられた標的配列が、非特許文献3におけるオフターゲット解析実験「CHANGE-seq」、「TTISS」、および「GUIDE-seq」のいずれで仕様されたものかに基づいて、仮想配列を3グループに分類した。すなわち、各仮想配列は、「CHANGE-seq」グループ、「TTISS」グループ、および「GUIDE-seq」グループに分類された。
上記のように生成された仮想配列の生成に用いられた標的配列が、非特許文献3におけるオフターゲット解析実験「CHANGE-seq」、「TTISS」、および「GUIDE-seq」のいずれで仕様されたものかに基づいて、仮想配列を3グループに分類した。すなわち、各仮想配列は、「CHANGE-seq」グループ、「TTISS」グループ、および「GUIDE-seq」グループに分類された。
次に、各グループの仮想配列の各々に対して、DNA結合性ツールが標的配列とは異なる配列に作用する可能性を示す予測結果を、CRISPR-Netスコアの総和に-1を積算することによって算出した。すなわち、予告結果は、(-1)×Σ(CRISPR-Netスコア)である。また、各グループの仮想配列の各々に対して、DNA結合性ツールが標的配列と一致する配列と異なる配列に作用する可能性を示す予測結果を、MOFFスコアの総和に-1を積算することによって算出した。すなわち、予告結果は、(-1)×Σ(MOFFスコア)である。以下、本開示に係るオフターゲットリスク解析方法を用いて出力されたこれらの予測結果を「特異性予測スコア」と呼称する。
[特異性予測スコアと実験データとの比較]
上記のように算出された特異性予測スコアを、実際のオフターゲット発生率と比較した。ここで、実際のオフターゲット発生率として、非特許文献3に記載されているオフターゲット解析実験「CHANGE-seq」、「TTISS」、および「GUIDE-seq」の結果を用いた。具体的には、本実施例では、非特許文献3に記載された「CHANGE-seq」グループ、「TTISS」グループ、および「GUIDE-seq」グループの各々についての「Off-on-ratio」スコアをとして用いた。
上記のように算出された特異性予測スコアを、実際のオフターゲット発生率と比較した。ここで、実際のオフターゲット発生率として、非特許文献3に記載されているオフターゲット解析実験「CHANGE-seq」、「TTISS」、および「GUIDE-seq」の結果を用いた。具体的には、本実施例では、非特許文献3に記載された「CHANGE-seq」グループ、「TTISS」グループ、および「GUIDE-seq」グループの各々についての「Off-on-ratio」スコアをとして用いた。
「CHANGE-seq」グループ、「TTISS」グループ、および「GUIDE-seq」グループの各々について、「Off-on-ratio」スコアおよび特異性予測スコアについての散布図を作成した。「Off-on-ratio」スコアと特異性予測スコアとの相関関係は、スピアマンの相関係数(Spearman correlation)によって評価した。
図13および図16は、「CHANGE-seq」グループについての散布図であり、図14および図17は、「TTISS」グループについての散布図であり、図15および図18は、「GUIDE-seq」グループについての散布図である。
<結果>
[CRISPR-Netスコアを用いて算出した特異性予測スコアと「Off-on-ratio」スコアとの相関関係]
図13は、「CHANGE-seq」グループにおける、CRISPR-Netスコアを用いて算出した特異性予測スコアと「Off-on-ratio」スコアとの相関関係を示す散布図である。図13に示すように、「CHANGE-seq」グループにおける、CRISPR-Netスコアを用いて算出した特異性予測スコアと「Off-on-ratio」スコアとの相関関係に関するSpearman correlationは-0.231839であった。「Off-on-ratio」スコアとCRISPR-Netスコアを用いた特異性予測スコアとの間に有意な相関関係(Spearman順位相関係数の無相関検定、p-value<0.05)が認められた。
[CRISPR-Netスコアを用いて算出した特異性予測スコアと「Off-on-ratio」スコアとの相関関係]
図13は、「CHANGE-seq」グループにおける、CRISPR-Netスコアを用いて算出した特異性予測スコアと「Off-on-ratio」スコアとの相関関係を示す散布図である。図13に示すように、「CHANGE-seq」グループにおける、CRISPR-Netスコアを用いて算出した特異性予測スコアと「Off-on-ratio」スコアとの相関関係に関するSpearman correlationは-0.231839であった。「Off-on-ratio」スコアとCRISPR-Netスコアを用いた特異性予測スコアとの間に有意な相関関係(Spearman順位相関係数の無相関検定、p-value<0.05)が認められた。
図14は、「TTISS」グループにおける、CRISPR-Netスコアを用いて算出した特異性予測スコアと「Off-on-ratio」スコアとの相関関係を示す散布図である。図14に示すように、「TTISS」グループにおける、CRISPR-Netスコアを用いて算出した特異性予測スコアと「Off-on-ratio」スコアとの相関関係に関するSpearman correlationは-0.495973であった。「Off-on-ratio」スコアとCRISPR-Netスコアを用いた特異性予測スコアとの間に有意な相関関係(Spearman順位相関係数の無相関検定、p-value<0.05)が認められた。
図15は、「GUIDE-seq」グループにおける、CRISPR-Netスコアを用いて算出した特異性予測スコアと「Off-on-ratio」スコアとの相関関係を示す散布図である。図15に示すように、「GUIDE-seq」グループにおける、CRISPR-Netスコアを用いて算出した特異性予測スコアと「Off-on-ratio」スコアとの相関関係に関するSpearman correlationは-0.380952であった。
[MOFFスコアを用いて算出した特異性予測スコアと「Off-on-ratio」スコアとの相関関係]
図16は、「CHANGE-seq」グループにおける、MOFFスコアを用いて算出した特異性予測スコアと「Off-on-ratio」スコアとの相関関係を示す散布図である。図16に示すように、「CHANGE-seq」グループにおける、MOFFスコアを用いて算出した特異性予測スコアと「Off-on-ratio」スコアとの相関関係に関するSpearman correlationは-0.566788であった。
図16は、「CHANGE-seq」グループにおける、MOFFスコアを用いて算出した特異性予測スコアと「Off-on-ratio」スコアとの相関関係を示す散布図である。図16に示すように、「CHANGE-seq」グループにおける、MOFFスコアを用いて算出した特異性予測スコアと「Off-on-ratio」スコアとの相関関係に関するSpearman correlationは-0.566788であった。
図17は、「TTISS」グループにおける、MOFFスコアを用いて算出した特異性予測スコアと「Off-on-ratio」スコアとの相関関係を示す散布図である。図17に示すように、「TTISS」グループにおける、MOFFスコアを用いて算出した特異性予測スコアと「Off-on-ratio」スコアとの相関関係に関するSpearman correlationは-0.655498であった。
図18は、「GUIDE-seq」グループにおける、MOFFスコアを用いて算出した特異性予測スコアと「Off-on-ratio」スコアとの相関関係を示す散布図である。図18に示すように、「GUIDE-seq」グループにおける、MOFFスコアを用いて算出した特異性予測スコアと「Off-on-ratio」スコアとの相関関係に関するSpearman correlationは-0.761904であった。
[MOFFスコアを用いて算出した特異性予測スコアと「Off-on-ratio」スコアとの相関関係]
「CHANGE-seq」グループ、「TTISS」グループ、および「GUIDE-seq」グループのいずれにおいても、「Off-on-ratio」スコアとMOFFスコアを用いた特異性予測スコアとの間に有意な相関関係(Spearman順位相関係数の無相関検定、p-value<0.05)が認められた。
「CHANGE-seq」グループ、「TTISS」グループ、および「GUIDE-seq」グループのいずれにおいても、「Off-on-ratio」スコアとMOFFスコアを用いた特異性予測スコアとの間に有意な相関関係(Spearman順位相関係数の無相関検定、p-value<0.05)が認められた。
また、「CHANGE-seq」グループ、「TTISS」グループ、および「GUIDE-seq」グループのうち、オフターゲット解析実験に使われたgRNA配列が少ない「GUIDE-seq」を除く「CHANGE-seq」グループおよび「TTISS」グループにおいて、「Off-on-ratio」スコアとCRISPR-Netスコアを用いた特異性予測スコアとの間に有意な相関関係が認められた。
これらの結果から、本開示に係るオフターゲットリスク解析方法におけるスコア算出ルール22として、MOFFスコアおよびCRISPR-Netスコアが適用可能であることが強く示唆された。
11 標的配列受付部
12 仮想配列生成部
13 スコア算出部
14 予測結果出力部
100、100a オフターゲットリスク解析システム
S2 仮想配列生成ステップ
S3 スコア算出ステップ
S4 予測結果出力ステップ
12 仮想配列生成部
13 スコア算出部
14 予測結果出力部
100、100a オフターゲットリスク解析システム
S2 仮想配列生成ステップ
S3 スコア算出ステップ
S4 予測結果出力ステップ
Claims (11)
- 標的配列と同一の配列、および該標的配列に少なくとも1つの変異を導入した配列を含む複数の仮想配列を生成する仮想配列生成ステップと、
前記標的配列を認識するDNA結合性ツールが作用する確率に関連するスコアを、前記複数の仮想配列の各々について算出するスコア算出ステップと、
算出された前記スコアに基づいて、前記DNA結合性ツールが前記標的配列とは異なる配列に作用する可能性を示す予測結果を出力する予測結果出力ステップと、
を含む、オフターゲットリスク解析方法。 - 前記仮想配列生成ステップにおいて、前記少なくとも1つの変異は、前記標的配列の全体、または該標的配列の一部分に対して導入される、
請求項1に記載のオフターゲットリスク解析方法。 - 前記変異は、置換、欠失、および挿入のうちのいずれかである、
請求項1に記載のオフターゲットリスク解析方法。 - 前記変異が置換である場合、前記複数の仮想配列は、
前記標的配列の少なくとも1つのアデニン(A)を、チミン(T)、シトシン(C)、およびグアニン(G)の少なくともいずれか1つに置換した配列、および/または、
前記標的配列の少なくとも1つのチミン(T)を、アデニン(A)、シトシン(C)、およびグアニン(G)の少なくともいずれか1つに置換した配列、および/または、
前記標的配列の少なくとも1つのシトシン(C)を、アデニン(A)、チミン(T)、およびグアニン(G)の少なくともいずれか1つに置換した配列、および/または、
前記標的配列の少なくとも1つのグアニン(G)を、アデニン(A)、チミン(T)、およびシトシン(C)の少なくともいずれか1つに置換した配列、
を含む、
請求項3に記載のオフターゲットリスク解析方法。 - 前記仮想配列生成ステップにおいて、前記複数の仮想配列の各々は、前記標的配列と4つ以下のヌクレオチドが異なっている、
請求項1に記載のオフターゲットリスク解析方法。 - 前記スコア算出ステップにおいて、前記スコアは、前記複数の仮想配列の各々に前記DNA結合性ツールを結合させた場合の安定性を示す値を用いて算出される、
請求項1に記載のオフターゲットリスク解析方法。 - 前記予測結果出力ステップにおいて、前記複数の仮想配列の各々について算出されたすべての前記スコアを用いて算出される評価値を前記予測結果として出力する、
請求項1に記載のオフターゲットリスク解析方法。 - 前記DNA結合性ツールは、
ジンクフィンガー、またはジンクフィンガーに機能性ドメインを組み合わせた融合ポリペプチド、
TALE(Transcription Activator-Like Effector)、またはTALEに機能性ドメインを組み合わせた融合ポリペプチド、
ペンタトリコペプチドリピート(pentatricopeptide repeat)、またはペンタトリコペプチドリピートに機能性ドメインを組み合わせた融合ポリペプチド、
野生型のCRISPR/Cas(CRISPR-associated protein)、または野生型のCRISPR/Casに機能性ドメインを組み合わせた融合ポリペプチド-核酸複合体、または、
改変型のCRISPR/Cas(CRISPR-associated protein)、または改変型のCRISPR/Casに機能性ドメインを組み合わせた融合ポリペプチド-核酸複合体、である、
請求項1から7の何れか1項に記載のオフターゲットリスク解析方法。 - 標的配列と同一の配列、および該標的配列に少なくとも1つの変異を導入した配列を含む複数の仮想配列を生成する仮想配列生成部と、
前記標的配列を認識するDNA結合性ツールが作用する確率に関連するスコアを、前記複数の仮想配列の各々について算出するスコア算出部と、
算出された前記スコアに基づいて、前記DNA結合性ツールが前記標的配列とは異なる配列に作用する可能性を示す予測結果を出力する予測結果出力部と、
を備える、オフターゲットリスク解析システム。 - コンピュータに、
標的配列と同一の配列、および該標的配列に少なくとも1つの変異を導入した配列を含む複数の仮想配列を生成する仮想配列生成ステップと、
前記標的配列を認識するDNA結合性ツールが作用する確率に関連するスコアを、前記複数の仮想配列の各々について算出するスコア算出ステップと、
算出された前記スコアに基づいて、前記DNA結合性ツールが前記標的配列とは異なる配列に作用する可能性を示す予測結果を出力する予測結果出力ステップと、
を実行させるためのプログラム。 - 請求項10に記載のプログラムを記録したコンピュータ読取り可能な記録媒体。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2024549981A JPWO2024070589A1 (ja) | 2022-09-30 | 2023-09-08 |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2022158915 | 2022-09-30 | ||
| JP2022-158915 | 2022-09-30 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024070589A1 true WO2024070589A1 (ja) | 2024-04-04 |
Family
ID=90477418
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2023/032841 Ceased WO2024070589A1 (ja) | 2022-09-30 | 2023-09-08 | オフターゲットリスク解析方法、オフターゲットリスク解析システム、プログラム、記録媒体 |
Country Status (2)
| Country | Link |
|---|---|
| JP (1) | JPWO2024070589A1 (ja) |
| WO (1) | WO2024070589A1 (ja) |
Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2021100731A1 (ja) * | 2019-11-19 | 2021-05-27 | 国立大学法人 長崎大学 | Cas9ヌクレアーゼを用いて相同組み換えを誘導する方法 |
-
2023
- 2023-09-08 JP JP2024549981A patent/JPWO2024070589A1/ja active Pending
- 2023-09-08 WO PCT/JP2023/032841 patent/WO2024070589A1/ja not_active Ceased
Patent Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2021100731A1 (ja) * | 2019-11-19 | 2021-05-27 | 国立大学法人 長崎大学 | Cas9ヌクレアーゼを用いて相同組み換えを誘導する方法 |
Non-Patent Citations (2)
| Title |
|---|
| TAKASHI YAMAMOTO: ""CRISPRdirect" Utilizing Genome Editing Technology: Designing Guide RNA with Little Off-Target Action", 23 November 2021 (2021-11-23), XP093152028, Retrieved from the Internet <URL:https://web.archive.org/web/20211123033113/https://biosciencedbc.jp/about-us/files/leaflet_crisprdirect.pdf> * |
| TAMANO, SHIYU: "Proposal of Evaluation Off-target Effects Method in Gapmer ASO", IPSJ SIG TECHNICAL REPORT, vol. 2022-BIO-69, no. 7, 10 March 2022 (2022-03-10), pages 1 - 7, XP009554159, ISSN: 2188-8590 * |
Also Published As
| Publication number | Publication date |
|---|---|
| JPWO2024070589A1 (ja) | 2024-04-04 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Farrell et al. | BiSulfite Bolt: A bisulfite sequencing analysis platform | |
| Clum et al. | DOE JGI metagenome workflow | |
| Marchant et al. | The C-Fern (Ceratopteris richardii) genome: insights into plant genome evolution with the first partial homosporous fern genome assembly | |
| Yan et al. | Species tree inference methods intended to deal with incomplete lineage sorting are robust to the presence of paralogs | |
| Cariou et al. | Is RAD‐seq suitable for phylogenetic inference? An in silico assessment and optimization | |
| Park et al. | Comparative analyses of DNA methylation and sequence evolution using Nasonia genomes | |
| McCormack et al. | Maximum likelihood estimates of species trees: how accuracy of phylogenetic inference depends upon the divergence history and sampling design | |
| Giarla et al. | The challenges of resolving a rapid, recent radiation: empirical and simulated phylogenomics of Philippine shrews | |
| Huerta-Cepas et al. | The human phylome | |
| Tirosh et al. | Comparative analysis indicates regulatory neofunctionalization of yeast duplicates | |
| Klopfstein et al. | More on the best evolutionary rate for phylogenetic analysis | |
| Elemento et al. | Reconstructing the duplication history of tandemly repeated genes | |
| Woo et al. | A quantitative quasispecies theory-based model of virus escape mutation under immune selection | |
| Cechova et al. | Dynamic evolution of great ape Y chromosomes | |
| Sveinsson et al. | Phylogenetic pinpointing of a paleopolyploidy event within the flax genus (Linum) using transcriptomics | |
| Xu et al. | An efficient pipeline for ancient DNA mapping and recovery of endogenous ancient DNA from whole‐genome sequencing data | |
| Lin et al. | Probing the genomic limits of de-extinction in the Christmas Island rat | |
| Burleigh et al. | Supertree bootstrapping methods for assessing phylogenetic variation among genes in genome-scale data sets | |
| Raza et al. | Resolving the phylogeny of Thladiantha (Cucurbitaceae) with three different target capture pipelines | |
| Yang et al. | Overcoming CRISPR-Cas9 off-target prediction hurdles: A novel approach with ESB rebalancing strategy and CRISPR-MCA model | |
| Baker et al. | Evolution of Alu subfamily structure in the Saimiri lineage of new world monkeys | |
| Byrnes et al. | Reorganization of adjacent gene relationships in yeast genomes by whole-genome duplication and gene deletion | |
| Souza et al. | Detecting clustered independent rare variant associations using genetic algorithms | |
| Trivedi et al. | Balanced training sets improve deep learning-based prediction of CRISPR sgRNA activity | |
| Dussex et al. | Biomolecular analyses reveal the age, sex and species identity of a near-intact Pleistocene bird carcass |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23871826 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2024549981 Country of ref document: JP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 23871826 Country of ref document: EP Kind code of ref document: A1 |