WO2025255780A1 - 纳米孔蛋白复合物、其构建方法及应用 - Google Patents
纳米孔蛋白复合物、其构建方法及应用Info
- Publication number
- WO2025255780A1 WO2025255780A1 PCT/CN2024/099017 CN2024099017W WO2025255780A1 WO 2025255780 A1 WO2025255780 A1 WO 2025255780A1 CN 2024099017 W CN2024099017 W CN 2024099017W WO 2025255780 A1 WO2025255780 A1 WO 2025255780A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- seq
- protein
- nanopore
- mutates
- complex
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C07—ORGANIC CHEMISTRY
- C07K—PEPTIDES
- C07K14/00—Peptides having more than 20 amino acids; Gastrins; Somatostatins; Melanotropins; Derivatives thereof
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
- C12N15/09—Recombinant DNA-technology
- C12N15/63—Introduction of foreign genetic material using vectors; Vectors; Use of hosts therefor; Regulation of expression
- C12N15/64—General methods for preparing the vector, for introducing it into the cell or for selecting the vector-containing host
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6869—Methods for sequencing
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N33/00—Investigating or analysing materials by specific methods not covered by groups G01N1/00 - G01N31/00
- G01N33/48—Biological material, e.g. blood, urine; Haemocytometers
- G01N33/50—Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing
- G01N33/68—Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing involving proteins, peptides or amino acids
Definitions
- This invention relates to the field of nanopore sequencing technology, and more specifically, to a nanopore protein complex, its construction method, and its application.
- Nanopore sequencing technology as an emerging single-molecule sequencing technology, has brought about a disruptive change to the gene sequencing industry with its unique advantages such as high throughput, long read length, speed, in situ detection and label-free operation. It has wide applications in basic theoretical research and biomedical clinical practice in many fields such as molecular biology, medicine, epidemiology and ecology.
- Nanopore sequencing technology is an electrical signal-based sequencing technique that can simultaneously test the sequences of nucleotides, amino acids, or glycans, as well as base, amino acid, or glycan modifications (such as methylation and acylation, phosphorylation, hydroxylation, oxidation, reduction, glycosylation, decarboxylation, and deamination).
- Its core component, the nanoporous protein is intercalated into a membrane, acting as a signal sensor. It separates two electrolytic chambers containing electrolyte. When a voltage is applied between the chambers, a stable current is generated. When the analyte enters the nanopore, it impedes the flow of ions, causing fluctuations in the current signal.
- the sequence information of the analyte can be sequenced in real time.
- nanopore sequencing The theoretical basis of nanopore sequencing is that, under the control of motor proteins, the bases of nucleic acid molecules sequentially pass through the central channel of nanopore proteins embedded in the membrane. The resulting current signal is amplified by the underlying circuitry and converted into the final sequence information by a base recognition algorithm. This approach has been thoroughly demonstrated and is feasible.
- commercially available nanopore sequencers mainly include a series of nanopore sequencers developed by Oxford Nanopore Technologies in the UK, such as MinION, GridION, and PromethION, and the QNome-3841 nanopore sequencer developed by QiCarbon Technology.
- these nanopore sequencers still have significant shortcomings in sequencing accuracy, throughput, and chip stability, failing to meet the ultimate needs of molecular biology research. Therefore, this field requires single-molecule sequencers with high accuracy, high integration, and high stability.
- nanoporin One of the main reasons for the low accuracy of nanopore sequencing is the singularity and low resolution of the signal sensor of its core component, the nanoporin.
- the narrowest contracted region inside the nanoporin is the most discriminative part of the analyte current signal change characteristics and can be used as the internal sensing region of the current signal. This region needs to be sharp enough to have high spatial resolution in both the lateral and longitudinal directions.
- pore proteins Among the pore proteins studied, three main categories are potentially suitable for sequencing: transporter proteins that act as channels for transporting various biomolecules and small molecules inside and outside the cell; pore-forming toxins produced by bacteria or other organisms that disrupt cell membrane permeability; and viral connectors that provide genome transport channels for viral infection of the host.
- transporter proteins that act as channels for transporting various biomolecules and small molecules inside and outside the cell
- pore-forming toxins produced by bacteria or other organisms that disrupt cell membrane permeability
- viral connectors that provide genome transport channels for viral infection of the host.
- MspA Mycobacterium smegmatis pore protein A
- CsgG curli-specific transport channel
- the difficulty in sequencing natural proteins stems from factors such as the challenges posed by recombinant protein in vitro expression and purification systems to protein stability, the stability and symmetry of protein aggregates, and the shape and size of the protein's luminal contraction zone.
- nanopore sequencers on the market have only one contraction region, allowing them to contact only a small portion of the analyte sequence.
- sequences interact with each other, resulting in highly complex conductivity maps down to the underlying sequences. This is especially problematic for homopolymer sequences, which current algorithms struggle to accurately distinguish. This leads to the relatively poor resolution of current nanopore sequencers, resulting in a relatively high error rate.
- One of the main reasons for the slow adoption of nanopore sequencing is that existing literature has demonstrated that having more than one sensing region within a single nanopore can help obtain additional sequence information, providing more opportunities to resolve homopolymer regions and overcoming the disadvantage of low sequencing accuracy.
- the main objective of this invention is to provide a nanoporin complex to solve the problem of low sequencing accuracy when nanoporins with only a single contractile region are applied to nanopore sequencing in the prior art.
- a nanoporin complex comprising: a nanoporin and an accessory protein, wherein the nanoporin is polymerized from multiple nanoporin monomers, and the polymerization of the multiple nanoporin monomers forms a hollow nanoporous cavity; the accessory protein is polymerized from multiple accessory protein monomers, at least a portion of the accessory protein is located within the nanoporous cavity, and the accessory protein and the nanoporous cavity together form a continuous channel; wherein, according to the direction of movement of the analyte through the continuous channel, the continuous channel includes a first sensing region and a second sensing region connected in sequence, the first sensing region being formed by at least a portion of the nanoporin, and the second sensing region being formed by at least a portion of the accessory protein; the nanoporin monomers are selected from any of the following proteins: 1) a protein having the amino acid sequence shown in SEQ ID NO: 1; 2) a protein having at
- the protein has 50% identity and the ability to polymerize to form nanoporous proteins; or 3) a protein that has been substituted, deleted, or added one or more amino acids based on SEQ ID NO: 1 and has the ability to polymerize to form nanoporous proteins;
- the accessory protein monomer is selected from any of the following proteins: i) a protein having the amino acid sequence shown in SEQ ID NO: 3; ii) a protein having at least 50% identity with the amino acid sequence shown in SEQ ID NO: 3 and having the following function: binding to the porin monomer and polymerizing together with the porin monomer to form nanoporous proteins to form accessory proteins; or iii) a protein that has been substituted, deleted, or added one or more amino acids to the amino acid sequence shown in SEQ ID NO: 3 and has the following function: binding to the porin monomer and polymerizing together with the porin monomer to form nanoporous proteins to form accessory proteins.
- the nanoporin and the accessory protein originate from the same species; optionally, all or part of the N-terminal portion of the accessory protein is located within the nanoporous cavity of the nanoporin; preferably, the accessory protein is attached to the nanoporous cavity of the nanoporin through covalent or non-covalent interactions; more preferably, the nanoporin and the accessory protein exist in the form of a nonamericomer.
- the accessory protein monomer is a mutant with a truncated amino acid sequence as shown in SEQ ID NO: 3.
- the truncated mutant is selected from any of the following mutants: retaining only amino acids 23-67 at the N-terminus, retaining only amino acids 23-57 at the N-terminus, retaining only amino acids 23-52 at the N-terminus, retaining only amino acids 23-54 at the N-terminus, retaining only amino acids 23-51 at the N-terminus, retaining only amino acids 23-50 at the N-terminus, retaining only amino acids 23-49 at the N-terminus, retaining only amino acids 23-48 at the N-terminus, or retaining only amino acids 23-47 at the N-terminus.
- the length of the accessory protein monomer is 24 to 45 amino acids; preferably, the amino acid sequence of the accessory protein monomer comes from the following residue position intervals of SEQ ID NO: 3 or its mutants: positions 23 to 48, 23 to 49, 23 to 50, 23 to 51, 23 to 54 or 23 to 57.
- the porin monomer is selected from proteins with mutations at at least one amino acid site in any one or more of the following groups in SEQ ID NO: 1: 1) S71, N74, G75, and F76; preferably, S71 is mutated to G, A, or T; N74 is mutated to G, A, or T;
- the mutations are as follows: 1) G, A, S, or T; G75 mutates to A, S, T, or Q; F76 mutates to A, S, T, N, or Q; 2) E162, R196, S200, and S216; preferably, E162 mutates to A, G, V, L, I, Y, F, or W; R196 mutates to A, G, V, L, I, Y, F, or W; S200 mutates to A, G, V, L, I, Y, F, or W; S216 mutates to A, G, V, L, I, Y, F, or...
- the mutations are: A, G, S, T, N, or Q; R165 mutations are: A, G, S, T, N, or Q; D199 mutations are: A, G, S, T, N, or Q; R207 mutations are: A, G, S, T, N, Q, D, or E; K209 mutations are: A, G, S, T, N, Q, D, or E; K210 mutations are: A, G, S, T, N, Q, D, or E; E213 mutations are: A, G, S, T, N, or Q; E215 mutations are: A, G, S, T, N, or Q.
- the accessory protein monomer is selected from the amino acid sequence shown in SEQ ID NO: 3, in which at least one of the following amino acids is substituted: A39, T42 or N46; preferably, A39 is mutated to S, T, N, G, V, L, I or Q; T42 is mutated to A, G or S; and N46 is mutated to A, G or S.
- first sensing region and the second sensing region have the same or different minimum aperture diameters; preferably, the minimum aperture diameter is 1-3 nm.
- the accessory protein monomer is selected from proteins whose amino acid sequence shown in SEQ ID NO: 3 is mutated to cysteine or substituted with a non-natural amino acid: S23, S24, L25, T28, K30, N31, S33, F34, N47, A50, Q51, N52 or Q53; and/or the porin monomer is selected from proteins whose amino acid sequence shown in SEQ ID NO: 1 is mutated to cysteine or substituted with a non-natural amino acid: N145, T148, K154, L156, L160, S161, R165, S195, Q197, D199, F203, Y205, K209, K210, L211, E213, E215, G217, S219 or N221.
- the porin monomer and accessory protein monomer exhibit mutations in at least one of the following combinations, where the corresponding amino acid at each combination is mutated to cysteine or substituted with a non-natural amino acid: 1) R165 on SEQ ID NO: 1 and S23 on SEQ ID NO: 3; 2) S195 on SEQ ID NO: 1 and S24 on SEQ ID NO: 3; 3) S219 on SEQ ID NO: 1 and L25 on SEQ ID NO: 3; 4) N221 on SEQ ID NO: 1 and L25 on SEQ ID NO: 3; 5) N145C on SEQ ID NO: 1 and V26C on SEQ ID NO: 3; 6) D199 on SEQ ID NO: 1 and T28 on SEQ ID NO: 3; 7) D199 on SEQ ID NO: 1 and K30 on SEQ ID NO: 3; 8) SEQ ID NO: 1 and S23 on SEQ ID NO: 3; 9) S219 on SEQ ID NO: 1 and S23 on SEQ ID NO: 3; 10) S2
- the amino acid at the corresponding site is mutated to cysteine.
- the amino acid is replaced with a non-natural amino acid through modification, which can be direct or indirect modification; preferably, direct modification includes modification through spontaneous reaction or oxidation of the side chain groups of the amino acid; preferably, indirect modification...
- the modification includes chemical modification by attaching small chemical molecules; preferably, the small chemical molecules include chemical crosslinking agents containing functional groups, and the functional groups are resistant to dithiothreitol.
- disulfide bonds are formed between at least some of the sites in the accessory protein monomers that are mutated to cysteine and at least some of the sites in the porin monomers that are mutated to cysteine.
- a nanopore sensor comprising a membrane and any of the above-described nanopore protein complexes, wherein the nanopore proteins in the nanopore protein complex are positioned on the membrane, and the nanopore cavities of the nanopore proteins and their accessory proteins together form a continuous transmembrane channel.
- the membrane includes an amphiphilic molecular layer; preferably, the membrane is a phospholipid bilayer, a diblock copolymer, or a triblock copolymer composed of diacylphosphatidylcholine.
- a nanopore sequencing system including any of the above-described nanopore sensors, the nanopore sequencing system further including: a conductive solution, positive and negative electrodes providing voltage and potential across the membrane, and a measuring device for measuring electrical signals passing through a continuous channel; wherein the nanopore sensor is located in the conductive solution, and the conductive solution is divided into a first chamber and a second chamber.
- the nanopore sequencing system also includes a test molecule, which is selected from polynucleotides, peptides and polysaccharides.
- the test molecule is selected from polynucleotides containing homopolymers.
- the test molecule can be transiently located in a continuous channel, with one end of the test molecule located in the first chamber and the other end located in the second chamber.
- a method for nanopore sequencing comprising: contacting a nanopore sequencing system with a molecule to be tested; applying a potential across a membrane to allow the molecule to enter a continuous channel; and measuring, once or multiple times, the electrical signal generated as the molecule moves relative to the continuous channel, thereby obtaining the sequence of the molecule to be tested.
- the electrical signal includes the current hysteresis signal strength, the current hysteresis duration, or the interval between current hysteresis events.
- the analyte is a polynucleotide
- the nucleotides in the polynucleotide interact with a first contraction region and a second contraction region within a continuous channel.
- Each contraction region in the first and second contraction regions is capable of distinguishing different nucleotides, such that the total current through the continuous channel is influenced by the interaction between each contraction region in the first and second contraction regions and the nucleotides located in each region.
- nucleic acid-binding proteins are used to control the movement of polynucleotides relative to the continuous channel pores.
- a method for preparing the above-mentioned nanoporin complex comprising: constructing expression vectors for porin monomers and accessory protein monomers, respectively; obtaining the nanoporin complex by co-expressing the expression vectors for porin monomers and accessory protein monomers in competent cells and then separating and purifying them; or constructing the nanoporin complex by in vitro recombination.
- an accessory protein is embedded within the nanopore cavity of a nanoporin, thereby introducing a second contraction region on top of the single contraction region inherent in the nanoporin.
- This forms a nanopore complex with a continuous channel containing two contraction regions (or sensing regions).
- nanopore sequencing using such a complex, as the nucleic acid molecule passes through the first and second contraction regions sequentially, nucleotides at different spatial locations of the nucleic acid molecule interact differently with the two contraction regions, generating more electrical signal (e.g., current) change information. Therefore, compared to sequencing with only a single contraction region, this method provides more comprehensive information.
- Nanoporous proteins, using the nanoporous protein complex of this application for nanopore sequencing can more accurately determine homopolymer sequence information, especially for longer homopolymer sequence fragments.
- Figure 1 shows a schematic diagram of the predicted structure of the multimer formed by (signal peptide) BCP34-AP34 in Embodiment 1 of the present invention, wherein A is a side view and B is a top view.
- Figure 2 shows a schematic diagram of the predicted structure of the multimer formed by mature (signal peptide-free) BCP34-AP34 in Example 1 of the present invention.
- A is a side view and B is a top view.
- Figure 3 shows a schematic diagram of the nanopore shrinkage region in the predicted structure of the polymer formed by BCP34-AP34 in Example 1 of the present invention.
- Figure 4 shows a schematic diagram of the side chains in the nanopore shrinkage region of the predicted structure of the polymer formed by BCP34-AP34 in Example 1 of the present invention.
- Figure 5 shows the purification results of the nanoporous complex constructed by the co-expression method in Example 3 of the present invention, where A is the result of purification by Strep column and B is the result of purification by Ni column after purification by Strep column.
- Figure 6 shows the purification results of the nanoporous complex constructed by the co-expression method in Example 3 of the present invention after TEV enzyme digestion, where A is the result of the nanoporous complex before denaturation and B is the result of the nanoporous complex after denaturation.
- Figure 7 shows the purification results of the nanoporous complex constructed by the in vitro recombination method in Example 4 of the present invention, wherein A is the electrophoresis diagram of the nanoporous complex before denaturation, and B is the electrophoresis diagram of the nanoporous complex after denaturation.
- Figure 8 shows a schematic diagram of the interaction region between BCP34 and AP34 in Embodiment 5 of the present invention.
- Figure 9 shows the purification diagram of the cross-linked mutant of BCP34 in Example 5 of the present invention.
- A is the diagram before denaturation
- B is the diagram after denaturation.
- Figure 10 shows the SDS-PAGE electrophoresis results of the nanoporous complex constructed by crosslinking modification in Example 5 of the present invention
- A shows the reaction efficiency of BCP34-2 and AP34-P2 forming a nanoporous complex
- B shows the reaction efficiency of BCP34-3 and AP34-P3 forming a nanoporous complex.
- Figure 11 shows the SDS-PAGE electrophoresis results of nanoporous complexes constructed from AP34 of different lengths in Example 5 of the present invention
- A shows the reaction efficiency of AP34-P2, AP34-P5, AP34-P6, and AP34-P7 forming nanoporous complexes with BCP34-2, respectively
- B shows the reaction efficiency of AP34-P4 forming nanoporous complexes with BCP34-2.
- Figure 12 shows a schematic diagram of a sequencing library containing helicase in Example 7 of the present invention and a schematic diagram of the pairing structure of the sequencing library and cholesterol-containing single-stranded DNA.
- A is a schematic diagram of a sequencing library containing helicase
- B is a schematic diagram of the pairing structure of the sequencing library and cholesterol-containing single-stranded DNA; a: upper strand; b: lower strand; c: double-stranded target fragment; d: helicase; e: cholesterol-labeled single-stranded DNA.
- Figure 13 shows the sequencing current signal of BCP34-1 nanopore protein in Example 7 of the present invention.
- Figure 14 shows the sequencing current signal of the BCP34-1-AP34-P1 nanopore complex in Example 7 of the present invention.
- Figure 15 shows the sequencing current signal of the BCP34-1-AP34-P8 nanopore complex in Example 7 of the present invention.
- Figure 16 shows the sequencing current signal of the BCP34-1-AP34-P9 nanopore complex in Example 7 of the present invention.
- Figure 17 shows the sequencing current signal of the BCP34-2-AP34-P2 nanopore complex in Example 7 of the present invention.
- Figure 21 shows the sequencing current signal of the BCP34-2-AP34-P7 nanopore complex in Example 7 of the present invention.
- Contraction zone or sensing zone refers to the narrowest inner diameter region of the transmembrane channel formed by the nanoporin or the nanoporin lumen and its accessory proteins.
- a transmembrane channel formed by a nanoporin alone has a contraction zone, while the transmembrane channel material formed by the nanoporin and its accessory proteins together constitutes other contraction zones.
- nanopore proteins for nanopore sequencing only have one contraction region, which is insufficient to meet the sequencing resolution requirements for homopolymers (oligonucleotide molecules containing multiple consecutive repeating nucleotides).
- this application utilizes techniques such as porin database mining, structure prediction and analysis, protein assembly, protein modification, nanopore sequencing performance testing and characterization to screen or construct a novel nanopore protein with multiple sensing regions and its mutants.
- This nanopore protein complex is constructed and polymerized from the porin monomer BCP34 and its accessory protein monomer AP34 to form a transmembrane protein with multiple contractile regions.
- This type of nanopore protein complex can be obtained through in vivo co-expression or constructed through in vitro recombinant expression, forming a highly efficient and stable complex.
- nanopore sensor was constructed using this complex, a nanopore sequencing system was built, and the results were validated on homopolymer nucleic acid molecules, demonstrating that the improved nanopore protein complex containing two or more contractile regions provided in this application helps improve the accuracy of distinguishing homopolymer nucleic acid molecules. Therefore, when applied to nanopore sequencing, it can output stable sequencing signals and has high resolution for homopolymer sequences, showing great potential in improving sequencing accuracy and addressing the shortcomings of nanopore sequencing.
- a nanopore complex comprising: nanopore proteins and accessory proteins, wherein the nanopores...
- the protein is polymerized from multiple porous protein monomers, and the polymerization of multiple porous protein monomers forms a hollow nanopore cavity;
- the accessory protein is polymerized from multiple accessory protein monomers, and at least a portion of the accessory protein is embedded in the nanopore cavity, and the accessory protein and the nanopore cavity together form a continuous channel; wherein, according to the direction of movement of the analyte through the continuous channel, the continuous channel includes a first sensing region and a second sensing region connected in sequence, the first sensing region is formed by a portion of the nanoporous protein, and the second sensing region is formed by a portion or all of the accessory protein;
- the nanoporin monomer is selected from any of the following proteins: 1) a protein having the amino acid sequence shown in SEQ ID NO: 1; 2) a protein having at least 50% identity with the amino acid sequence shown in SEQ ID NO: 1 and having the function of polymerizing to form nanoporins; or 3) a protein that has been substituted, deleted or added to SEQ ID NO: 1 by one or more amino acids and has the function of polymerizing to form nanoporins.
- the accessory protein monomer is selected from any of the following proteins: i) a protein having the amino acid sequence shown in SEQ ID NO: 3; ii) a protein having at least 50% identity with the amino acid sequence shown in SEQ ID NO: 3 and having the following function: binding to the porin monomer and polymerizing together with the porin monomer to form the above-mentioned nanoporin to form an accessory protein; or iii) a protein having one or more amino acids substituted, deleted or added to the amino acid sequence shown in SEQ ID NO: 3 and having the following function: binding to the porin monomer and polymerizing together with the porin monomer to form the above-mentioned nanoporin to form an accessory protein.
- the contraction region of a nanopore can distinguish different nucleotides (i.e., those with different bases) that pass through the pore in the nucleic acid to be tested. Each nucleotide interacts with the contraction region, generating a change in electrical current, which can be converted into corresponding sequence information.
- the signals measured from existing single contraction regions may not be sufficiently distinguishable to differentiate and resolve single-base current changes, thus making it impossible to accurately determine the length of the homopolymer solely based on the magnitude of the measured signal.
- the novel nanopore protein complex provided in this application helps improve the distinguishability of sequencing signals for nucleotide homopolymers.
- the nanoporous complex described in this application utilizes an accessory protein embedded within the nanoporous lumen of the nanoporous protein. This introduces a second contractile region on top of the nanoporous protein's existing contractile region, forming a nanoporous complex with a continuous channel containing two contractile regions (or sensing regions).
- an accessory protein embedded within the nanoporous lumen of the nanoporous protein. This introduces a second contractile region on top of the nanoporous protein's existing contractile region, forming a nanoporous complex with a continuous channel containing two contractile regions (or sensing regions).
- nanoporous sequencing using the nanoporous protein complex of this application can more accurately determine homopolymer sequence information, especially for longer homopolymer sequence fragments.
- this nanoporous protein complex can be applied to nanoporous sequencing, outputting stable sequencing signals and possessing high resolution for homopolymer sequences, with a high potential to improve sequencing accuracy and overcome the disadvantage of insufficient accuracy in nanoporous sequencing.
- the nanoporin and the accessory protein are preferably derived from the same species, which helps the two proteins to bind more specifically.
- accessory proteins in this application are a type of protein without pore structure. However, their monomers can non-covalently bind to the main pore protein monomers, and they also polymerize together during the self-assembly of the main pore protein monomers to form a polymer complex of 9 main pore protein monomers + 9 accessory protein monomers.
- the nanoporin in the above nanoporous complex is derived from the previously provided nanoporin monomer BCP34 and its mutants.
- the nanoporin is a nonamer, and its monomer sequence was mined from a metagenomic database at a depth of 11,000 meters in the Mariana Trench.
- the wild-type full-length protein sequence is SEQ ID NO: 1.
- the mature porin monomer BCP34 lacks a signal peptide region, and its amino acid sequence is SEQ ID NO: 2:
- sequence identity refers to the "sequence identity" between two amino acid sequences, that is, the percentage of identical amino acids between the sequences.
- Methods for assessing the degree of sequence identity between amino acids or nucleotides are known to those skilled in the art.
- amino acid sequence identity is typically measured using sequence analysis software. For instance, it can be determined using the BLAST program in the NCBI database. For information on the determination of sequence identity, see, for example: Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987 and Primers for Sequence Analysis, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991.
- proteins that have 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more than 99% e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98.5%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, or even more than 99.9%
- accessory protein monomers shown in SEQ ID NO: 3 and have the function of binding to the porin monomers and polymerizing together with the porin monomers to form the aforementioned nanoporous proteins to form accessory proteins, have active sites, active pockets, active mechanisms, protein structures, etc. that are highly likely to be the same as those of the AP34 monomer.
- Amino acid residues can be represented by three-letter or one-letter amino acid codes according to standards known and agreed upon in the art.
- the abbreviations for amino acid residues are as follows: alanine (Ala; A), asparagine (Asn; N), aspartic acid (Asp; D), arginine (Arg; R), cysteine (Cys; C), glutamic acid (Glu; E), glutamine (Gln; Q), glycine (Gly; G), histidine (His; H), isoleucine (Ile; I), leucine (Leu; L), lysine (Lys; K), methionine (Met; M), phenylalanine (Phe; F), proline (Pro; P), serine (Ser; S), threonine (Thr; T), tryptophan (Trp; W), tyrosine (Tyr; Y), and valine (Val; V).
- a preferred conservative amino acid substitution is the replacement of one amino acid residue from the following groups (1)-(5) with another amino acid from the same group: (1) smaller aliphatic nonpolar or weakly polar residues: Ala, Ser, Thr, Pro, and Gly; (2) negatively charged polar residues and their (uncharged) amides: Asp, Asn, Glu, and Gln; (3) positively charged polar residues: His, Arg, and Lys; (4) larger aliphatic nonpolar residues: Met, Leu, Ile, Val, and Cys; and (5) aromatic residues: Phe, Tyr, and Trp.
- Particularly preferred conservative amino acid substitutions are as follows: Ala is replaced by Gly or Ser; Arg is replaced by Lys; Asn is replaced by Gln or His; Asp is replaced by Glu. Cys is replaced by Ser; Gln is replaced by Asn; Glu is replaced by Asp; Gly is replaced by Ala or Pro; His is replaced by Asn or Gln; Ile is replaced by Leu or Val; Leu is replaced by Ile or Val; Lys is replaced by Arg, Gln, or Glu; Met is replaced by Leu, Tyr, or Ile; Phe is replaced by Met, Leu, or Tyr; Ser is replaced by Thr; Thr is replaced by Ser; Trp is replaced by Tyr; Tyr is replaced by Trp or Phe; and Val is replaced by Ile or Leu.
- Mutants of the BCP34 protein include, but are not limited to, variants with amino acid mutations at one or more of the following positions in SEQ ID NO: 1: S71, N74, G75, and F76; preferably, the mutation at position S71 includes a mutation to G, A, or T; the mutation at position N74 includes a mutation to G, A, S, or T; the mutation at position G75 includes a mutation to A, S, T, N, or Q; the mutation at position F76 includes a mutation to G, A, S, T, N, or Q; and one or more of the following amino acid mutations are present at positions E162, R196, S200, and S216; preferably, the mutation at position E162 includes a mutation to A, G, V, L, I, Y, F, or W; and the mutation at position R196 includes a mutation to A, G, V, L, I, Y, F, or W.
- the mutation includes a mutation to A, G, V, L, I, Y, F, or W;
- the mutation at position S200 includes a mutation to A, G, V, L, I, Y, F, or W;
- the mutation at position S216 includes a mutation to A, G, V, L, I, Y, F, or W; and one or more of the following positions have amino acid mutations: R103, E104, E112, R113, K114, R117, R119, D120, K122, D124, K154, R165, D199, R207, K209, K210, E213, and E215;
- the mutation at position R103 includes a mutation to A, G, S, T, N, or Q;
- the mutation at position E104 includes a mutation to K, R, A, or G.
- the aforementioned accessory protein monomer AP34 and its mutants are proteins from the same species as the porin monomer BCP34.
- the sequence of this protein was mined from a metagenomic database of the Mariana Trench. It can form a centrosymmetric nonamer upon interaction with BCP34, with its N-terminus (e.g., a portion of the N-terminus) or entirely inserted into the nanopores formed by the polymerization of the porin monomer BCP34.
- the full-length wild-type sequence of the accessory protein monomer, SEQ ID NO: 3 is as follows:
- the aforementioned accessory protein monomer AP34 and its mutants can be truncated to alter the intercalation ability and/or membrane adaptability of the pore complex. Truncating modifications include, but are not limited to, retaining only amino acids 23-67 of the N-terminus (i.e., lacking the N-terminus).
- the mutant AP34 fragment can be 24–45 amino acids long, specifically 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 amino acids.
- the contractile region of the porin monomer BCP34 consists of amino acids S71, N74, G75, and F76.
- the mutation directions of these amino acid residues in the contractile region are as follows: mutations at position S71 include mutations to G, A, or T; mutations at position N74 include mutations to G, A, S, or T; mutations at position G75 include mutations to A, S, T, or Q; and mutations at position F76 include mutations to A, S, T, N, or Q.
- the most important amino acid residue in the contractile region of the accessory protein monomer AP34 is A39. Additionally, T42 and N46 may also affect the sequencing signal.
- mutation directions of this contractile region include, but are not limited to, mutations at one or more positions of A39, T42, and N46 in SEQ ID NO: 3.
- mutations at position A39 include mutations to S, T, N, G, V, L, I, or Q
- mutations at position T42 include mutations to A, G, or S
- mutations at position N46 include mutations to A, G, or S.
- the aforementioned mutations in the relevant amino acids of the BCP34 and AP34 contraction regions help increase the signal-to-noise ratio of the current signal, thereby improving the resolution of base recognition and ultimately enhancing the accuracy of sequence determination of the analyte.
- All or part of the AP34 of the aforementioned accessory protein monomer can be localized within the nanopore cavity formed by the polymerization of the porin monomer BCP34.
- the central cavity or pore of the nonamer formed by the polymerization of the accessory protein monomer AP34 is aligned with the nanopore cavity formed by the polymerization of the nanoporin monomer BCP34 to form a continuous channel, such that each interacting chemical group of the analyte translocating through the continuous channel first interacts with the contractile region of the nanoporin, and then with the contractile region of the accessory protein.
- the two contraction regions within the channels of the aforementioned nanoporous composite can have the same or different minimum pore diameters (i.e., the minimum pore diameters of the first and second sensing regions are parallel; the minimum pore diameter of the first sensing region can be larger or smaller than that of the second).
- the pore diameter is included, but not limited to, the range of 1-3 nm.
- dual contraction regions can provide more discriminative current signals between observed currents, thus achieving higher resolution for analytes in homopolymer form.
- the accessory proteins are attached to the nanoporous cavity of the nanoporin via covalent or non-covalent interactions; more preferably, the nanoporin and accessory proteins exist in the form of a nonamericomer.
- attaching the accessory proteins to the nanoporous cavity of the nanoporin via covalent interactions is relatively more stable and helps to form stable continuous channels with at least two contraction zones.
- the AP34 contraction region is stably in contact with the BCP34 lumen primarily at the following residues in SEQ ID NO: 3: 23, 24, 25, 26, 31, 33, 34, 40, 43, and 45.
- the accessory protein monomer AP34 is attached to the nanoporin monomer through one or more covalent bonds and/or one or more non-covalent interactions.
- a relatively stable complex is formed on and with BCP34.
- suitable non-covalent interactions include salt bridges, electrostatic interactions, and ⁇ - ⁇ interactions.
- the aforementioned positions on AP34 include, but are not limited to, S23, S24, L25, T28, K30, N31, S33, F34, N47, A50, Q51, N52, and Q53.
- These positions on BCP34 include, but are not limited to, N145, T148, K154, L156, L160, S161, R165, S195, Q197, D199, F203, Y205, K209, K210, L211, E213, E215, G217, S219, and N221.
- the combinations of interaction sites between nanoporous proteins and accessory proteins include, but are not limited to, one or more of the following: R165 on BCP34 and S23 on AP34; S195 on BCP34 and S24 on AP34; S219 on BCP34 and L25 on AP34; N221 on BCP34 and L25 on AP34; N145C on BCP34 and V26C on AP34; D199 on BCP34 and T28 on AP34; D199 on BCP34 and K30 on AP34; B...
- a method for nanopore sequencing comprising: contacting the nanopore sequencing system with a nucleic acid molecule to be tested; applying a potential across the membrane to allow the nucleic acid molecule to enter a continuous channel; and measuring, once or multiple times, the electrical signal generated as the nucleic acid molecule moves relative to the continuous channel, thereby obtaining the sequence of the nucleic acid molecule to be tested.
- the nucleic acid molecule to be tested is an oligonucleotide molecule.
- the nucleotides in the oligonucleotide molecule interact with a first contraction region and a second contraction region in a continuous channel.
- Each contraction region in the first and second contraction regions can distinguish different nucleotides, such that the total current through the continuous channel is affected by the interaction between each contraction region in the first and second contraction regions and the nucleotides located in the corresponding regions.
- the oligonucleotide molecules move through a series of channels and undergo transmembrane translocation.
- nucleic acid-binding proteins are used to control the movement of the oligonucleotide molecules relative to the pores of the series of channels.
- the oligonucleotide molecule is an oligonucleotide molecule containing a homopolymer, and the above method can determine the nucleotide sequence of the oligonucleotide molecule containing the homopolymer.
- a method for preparing a nanoporous complex comprising: constructing an expression vector for a nanoporous protein monomer and an expression vector for an accessory protein, respectively; co-expressing the expression vectors for the nanoporous protein monomer and the accessory protein monomer in competent cells, and obtaining the nanoporous complex by separation and purification; or constructing the nanoporous complex by in vitro recombination.
- the nanopore complex is obtained by co-expressing an expression vector for a protein monomer and an expression vector for an accessory protein monomer in competent cells, followed by separation and purification. This includes: co-expressing an expression vector for a pore protein monomer and an expression vector for an accessory protein monomer in competent cells, culturing the competent cells to a predetermined OD value, and collecting the co-expressed cells.
- Cells wherein cellular co-expressing porin monomers and accessory protein monomers are spontaneously polymerized to form a nanoporous complex; the co-expressing cells are sequentially subjected to cell disruption, cell membrane dissolution, and purification to obtain the nanoporous complex; preferably, the predetermined OD value is an OD600 value of 0.6-0.8; preferably, the purification includes sequential strep-tagged purification and his-tagged purification; preferably, the molar ratio of porin monomers to accessory protein monomers in the purified nanoporous complex is 1:1.
- the construction of the nanoporous complex via in vitro recombinant method includes: expressing expression vectors for porin monomers and accessory protein monomers in competent cells, and separating and purifying the cells after expression to obtain porin monomers and accessory protein monomers; mixing and incubating the porin monomers and accessory protein monomers at a molar ratio of 1:30-50 to obtain the nanoporous complex; preferably, mixing and incubating the porin monomers and accessory protein monomers in a buffer solution; preferably, the buffer solution includes 20 mM Tris-HCl, 150 mM NaCl, 0.05% Tween 20 and 15% glycerol, pH 8.0; preferably, mixing and incubating are performed at 4°C.
- a protein purification tag can be added to the C-terminus of the porin monomer BCP34 and/or the accessory protein monomer AP34.
- the protein purification tag is well-known in the art and includes, but is not limited to, histidine tags (Poly His), Strep tag II (non-biotinylated avidin), Biotin Avitag (biotinylated short peptide), calmodulin-binding peptide (CBP), or arginine tags (Poly Arg).
- the tag insertion position includes, but is not limited to, residues 48, 49, 50, 51, 54, 57, 62, or 67 of SEQ ID NO: 1 and its mutants, and SEQ ID NO: 3 and its mutants.
- the AP34 protein purification tag can insert a short protease cleavage peptide sequence, which helps remove redundant residues that may interfere with protein complex formation or protein complex pores.
- protease cleavage sequences are well known in the art and include, but are not limited to, thrombin (recognition sequence: SEQ ID NO: 30: LVPRG ⁇ S), Factor Xa (recognition sequence: SEQ ID NO: 31: IE/DG ⁇ R), TEV protease (recognition sequence: SEQ ID NO: 32: ENLYFQ ⁇ G), and HRV 3C protease (recognition sequence: SEQ ID NO: 33: LEVLFQ ⁇ GP), where the arrows indicate the sites of protease action.
- the aforementioned nanoporous complex can also be obtained by co-expressing the porin monomer BCP34 and its mutants, and the accessory protein monomer AP34 and its mutants.
- Co-expression refers to the simultaneous expression of the porin monomer and the accessory protein monomer in suitable host cells, allowing the nanoporous complex to form in vivo.
- host cells include, but are not limited to, BL21(DE3), BL21Star(DE3)pLyss, Rossata(DE3), and Lemo21(DE3).
- At least one gene encoding a porin monomer and a gene encoding an accessory protein monomer in one vector, or at least one accessory protein monomer in a second vector can be co-transformed to express the protein and prepare the nanoporous complex in the transformed cells.
- These two vectors can be controlled by a single promoter or by two independent promoters, where the two independent promoters can be the same or different. This process can be performed in vivo in host cells or in a cell-free expression system.
- Preferred expression vectors include, but are not limited to, vectors with T7 as the promoter, such as PET.28a(+), PET.21a(+), PET.32a(+), etc.
- the nanoporous complex can be purified using the corresponding protein purification tags for BCP34 and AP34 proteins.
- the purification steps are methods known in the art (e.g., ion exchange, gel filtration, hydrophobic interaction column chromatography, etc.), which can be used alone or in different combinations to purify the components of the nanoporous complex.
- the protein purification tags for BCP34 and AP34 can be the same or different; using two different tags can improve the purity of the nanoporous complex to some extent.
- the aforementioned nanopore complexes can also be constructed via in vitro recombination.
- the porin monomer BP34 and its mutants can be encoded in suitable vectors and transformed into suitable host cells for expression.
- the final product is then purified using appropriate protein purification tags, employing methods known in the art (e.g., ion exchange, gel filtration, hydrophobic interactions). Column chromatography (and other methods) can be used alone or in different combinations to purify components of the pore complex.
- the accessory protein monomer AP34 and its mutants can be obtained by expression in the same host cells as the pore protein monomer BP34 and purification in vitro, or directly through solid-phase peptide synthesis.
- the protein sequence is SEQ ID NO: 3 or its mutants and truncated forms.
- the N-terminal signal peptide can guide its secretion onto membrane proteins.
- the C-terminal protein sequence unrelated to pore formation can be cleaved by the protease corresponding to the added sequence.
- the peptide sequence will be an amino acid sequence that directly does not contain a C-terminal signal peptide.
- Methods for peptide synthesis include, but are not limited to, solid-phase synthesis, liquid-phase fractional synthesis, Staudinger ligation, natural chemical ligation, photosensitive auxiliary group ligation, removable auxiliary group ligation, chemical regioselective ligation, carboxylic anhydride (NCA) method, combinatorial chemistry, enzymatic digestion, genetic engineering, and fermentation.
- NCA carboxylic anhydride
- FMOC or BOC solid-phase peptide synthesis is used.
- a nanoporous complex can be formed through in vitro incubation.
- the ratio of BCP34 to AP34 during incubation is included, but is not limited to, 500:1 to 1:1.
- Incubation temperatures include, but are not limited to, 4°C, 16°C, 20°C, 25°C, and 37°C.
- Incubation durations include, but are not limited to, 30 min, 1 h, 2 h, 3 h, 4 h, 5 h, 8 h, 16 h, and 24 h.
- Salt concentrations in the incubation solution include, but are not limited to, 50 mM, 100 mM, 150 mM, 200 mM, 250 mM, 300 mM, 400 mM, and 500 mM.
- the pH values of the incubation solution include, but are not limited to, 7.0, 7.5, 8.0, and 8.5. Oxidizing agents or chemical cross-linking agents that promote the stability of the nanoporous complex can be added during the reaction.
- the incubation method can involve simultaneously mixing the porin monomer BCP34 and the accessory protein monomer AP34 in the reaction solution, or inserting the porin monomer into the membrane and then adding the accessory protein monomer, so that the nanoporous complex can be formed in situ.
- the nanoporous protein complex with dual contraction zones constructed by the method described above can be used to characterize various analytes, including but not limited to various biological or synthetic macromolecules and polymers such as polynucleotides, peptides, and polysaccharides.
- the polynucleotides include DNA and/or RNA and their modifications.
- This invention provides a method for determining the presence or absence of one or more features of a target analyte, the method comprising:
- a. Contact the target analyte with a nanopore having a first contraction region and a second contraction region on a membrane of a nanopore sensor containing the above-mentioned nanopore protein complex BCP34-AP34 or its mutant, so that the target analyte moves relative to the nanopore.
- the target analyte interacts with the nanoporin complex, thereby enabling the target analyte to move relative to the nanopores in the continuous channel.
- the nanoporous protein complex provided in this application can be used to develop and explore more single-molecule sequencers with high accuracy, high integration, and high stability. Furthermore, as a biosensor, this protein complex has extremely high potential for identifying and analyzing the composition and modification information of different organic and inorganic substances, as well as applications in pharmacokinetics or drug screening, and also has broad application prospects in nucleic acid drug delivery and biosensing. It can also be combined with genomics, proteomics, metabolomics, etc., to build a universal measurement platform that meets the needs of full-omics analysis, helping us to understand the laws of life and the mechanisms of disease more deeply.
- Example 1 Predicted structure of AlphaFold2-Multimer for wild-type BCP34-AP34
- FIG. 1 A is a side view of the predicted BCP34-AP34 multimer structure, and B is a top view.
- A is a side view of the predicted mature BCP34-AP34 (without signal peptide) multimer structure, and B is a top view.
- the structure shows that BCP34:AP34 forms a nonamer with a 9:9 stoichiometric ratio through non-covalent interactions, exhibiting C9 symmetry.
- the complex has two contraction regions that play a decisive role in the generation of the current signal: the first contraction region is formed by the polymerization of the porin monomer BCP34, and the second contraction region is formed by the polymerization of the accessory protein monomer AP34, located below the first contraction region. There is a certain distance between the two contraction regions, as shown in Figure 3.
- Figure 4 shows the side chain structures of the important amino acids in each contraction region of the predicted structure of the complex.
- the first contraction region shows that the four amino acids in the side chain are S71, N74, G75, and F76 on BCP34.
- the second contraction region shows that the amino acid in the side chain is A39 on AP34, which is the narrowest point inside AP34. Additionally, T42 and N46 are also relatively narrow, which have a certain influence on the current signal.
- Example 2 Construction of expression vectors for wild-type BCP34 monomers and mutants and AP34 monomers and mutants
- the DNA sequence of the porin monomer BCP34 (SEQ ID NO: 5) was inserted into the multiple cloning region of the vector pET24a after digestion with NdeI and XhoI via in-fusion. StrepII amino acids were added to the C-terminus of the wild-type BCP34 amino acid sequence (SEQ ID NO: 1) as a purification tag, with kanamycin as the selection tag.
- the constructed vector was named pET24a-BCP34.
- other mutants such as BCP34-1 (SEQ ID NO: 6), were constructed using the expression vector of the porin monomer BCP34 as a template.
- the DNA sequence (SEQ ID NO: 7) of the accessory protein monomer AP34-1 was inserted into the multiple cloning region of the vector pET21a after digestion with NdeI and XhoI using an in-fusion method.
- AP34-1 incorporates a TEV restriction site at position 57 of the wild-type AP34 amino acid sequence to facilitate in vitro protein expression and purification.
- Six histidine residues were added to the C-terminus of the AP34-1 amino acid sequence (SEQ ID NO: 8) as a purification tag, with ampicillin as the selection tag.
- the constructed vector was named pET21a-AP34-1. Using the Agilent site-directed mutagenesis kit, other mutants were constructed using the AP34-1 expression vector as a template.
- amino acid sequence of AP34 (SEQ ID NO: 3):
- amino acid sequence of AP34-1 (with a TEV restriction site added at position 57 based on AP34, SEQ ID NO: 8):
- two protein monomers can be co-expressed in a suitable Gram-negative host (such as Escherichia coli) and extracted and purified from the outer membrane to form the complex.
- a suitable Gram-negative host such as Escherichia coli
- the bacterial culture was then evenly spread on plates containing 50 ⁇ g/mL kanamycin and 100 ⁇ g/mL ampicillin and incubated overnight at 37°C. The next day, single colonies were picked and incubated overnight at 37°C and 200 rpm in 5 mL LB medium containing 50 ⁇ g/mL kanamycin and 100 ⁇ g/mL ampicillin.
- the resulting bacterial culture was then inoculated at a volume ratio of 1:100 into 50 mL LB liquid medium containing 50 ⁇ g/mL kanamycin and 100 ⁇ g/mL ampicillin and incubated at 37°C and 200 rpm for 4 h.
- the expanded bacterial culture was inoculated at a volume ratio of 1:100 into 2 L LB liquid medium containing 50 ⁇ g/mL kanamycin and 100 ⁇ g/mL ampicillin, and cultured at 37°C and 200 rpm.
- IPTG was added to a final concentration of 0.5 mM, and the culture was incubated at 16°C and 200 rpm for approximately 16-18 h.
- the bacterial culture was collected by centrifugation at 8000 rpm, and the bacterial cells were frozen and stored at -20°C for later use.
- Buffer solution A 20 mM Tris-HCl, 150 mM NaCl, pH 8.0
- Buffer solution B 20 mM Tris-HCl, 150 mM NaCl, 1% DDM, pH 8.0
- Buffer solution C 20 mM Tris-HCl, 150 mM NaCl, 0.05% Tween 20, 15% gly, pH 8.0
- Buffer solution D 20mM Tris-HCl, 150mM NaCl, 0.05% Tween 20, 15% gly, 5mM desulfurized biotin, pH 8.0 (This buffer solution should be prepared fresh before use).
- Buffer solution E 20 mM Tris-HCl, 150 mM NaCl, 0.05% Tween 20, 15% gly, 300 mM imidazole, pH 8.0
- the bacterial cells were resuspended thoroughly at a ratio of 1g of bacterial cells to 10mL of buffer solution A, and the cells were homogenized using an autoclave until the bacterial solution was clear. Then, the cells were centrifuged at high speed (40,000 rpm, 4°C) for 1 hour using a low-temperature ultracentrifuge (Beckman, Optima TM XPN-90). The supernatant was removed, and the precipitate was resuspended using buffer solution B. The precipitate was then rotated overnight at 4°C. The next day, the cells were centrifuged at 18,000 rpm for 1 hour at 4°C. The supernatant was collected, filtered through a 0.22 ⁇ m filter, and stored at 4°C for later use.
- FIG. 5A shows the protein characterization results of the wild-type complex after purification using the Strep column.
- the denaturation condition was heating the sample at 60 °C for 15 minutes. The results showed that the target protein was in an aggregate state before denaturation and in a monomer state after denaturation.
- the monomers contained mature BCP34-1 and AP34-1 proteins, respectively, proving that the complex was formed on the membrane and purified through the Strep tag of BCP34-1. At this time, there was still an excess of BCP34-1 and some other proteins.
- FIG. 5B shows the protein characterization results of the wild-type complex after purification by the HisTrap column. The results show that the target protein is in an aggregate state before denaturation and in a monomer state after denaturation.
- the monomers contain mature BCP34-1 and AP34-1 proteins, respectively, proving that the complex forms on the membrane and is further purified by the HisTrap tag of AP34-1. At this point, the protein complex is in a 1:1 state.
- TEV enzyme was added to the protein elution solution to remove the redundant C-terminal protein sequence of AP34-1 (SEQ ID NO: 9SDAIRKDKTPIEEFNDRLQRSLLSRITSTISRSIIGIDGAVNPGSFETTDFLIDVTDLGGGQMSITTTDKVTGDQTSIVIETGL), while retaining the key sequence AP34-2 (SEQ ID NO: 10MSKVIFGFFAVLLMCFVVTASASSLVYTPKNPSFGGPAAYGTYLLNNANAQNQFKDPENLYFQ).
- the protein solution with added enzyme was placed on a gyroscope and rotated overnight at 4°C. The obtained target protein was then subjected to SDS-PAGE electrophoresis.
- Figure 6A shows the changes in the complex protein before and after the addition of TEV before denaturation
- Figure 6B shows the changes in the complex protein after the addition of TEV after denaturation. It can be concluded that the C-terminal redundant sequence of AP34-1 was successfully removed by the TEV enzyme, and our target protein was finally obtained.
- the obtained protein was concentrated to 1 mL, passed through Superdex 6increase 10/300GL (Cytiva) equilibrated with buffer solution C, and the target protein was collected and then stored at -80°C.
- Example 4 Construction of multi-shrinkage zone composite wells using in vitro recombination method
- the complex can also be constructed via in vitro recombination.
- a Gram-negative host such as Escherichia coli
- the key fragment of AP34 was constructed by peptide synthesis. Since peptide synthesis is performed in vitro, it does not require the N-terminal signal peptide and the C-terminal sequence that may interfere with the intercalation of the complex pore membrane.
- a peptide synthesized by solid-phase synthesis from Genscript to synthesize the peptide sequence AP34-P1 (SEQ ID NO: 11: SSLVYTPKNPSFGGPAAYGTYLLNNANAQNQFKDP).
- BCP34-1 The expression and purification steps for BCP34-1 were the same as in Example 3, except that the second purification step (histidine) and the TEV digestion step were omitted.
- the protein purified by the Strep-Tactin beads (IBA Lifesciences) chromatography column was directly eluted and diluted to an appropriate concentration for later use. Simultaneously, 2 mg of lyophilized AP34-P1 obtained from Genscript was dissolved in 1 mL of buffer solution C from the example to obtain a 2 mg/mL sample. The sample was vortexed until no remaining peptide powder was visible. Because BCP34-1 contains a small amount of contaminating proteins, accurate concentration measurement is difficult.
- a rough estimate of the sample concentration can be obtained using the intensity of the protein bands on SDS-PAGE relative to known labeled proteins.
- BCP34-1 and AP34-P1 were mixed at a molar ratio of approximately 1:50 and incubated at 4°C and 700 rpm for 4 h, followed by centrifugation at 13,000 rpm for 2 min.
- the obtained target protein was subjected to SDS-PAGE electrophoresis, as shown in Figure 7.
- Figure 7A shows the gel image of the complex before denaturation
- Figure 7B shows the gel image of the complex after denaturation. The results show that the target protein is in an aggregate state before denaturation and in a monomer state after denaturation.
- the gel image only indicates the state and stability of the protein.
- Example 5 Further stabilization of multi-shrinkage nanoporous complexes through covalent cross-linking
- Example 4 Although the complex pores are unstable, the two protein monomers constituting the complex pores have strong non-covalent interactions and can assemble into a polymer state before the proteins denature. In order to further reduce the potential risks in actual nanopore testing applications, in this example, we will modify the protein and introduce cysteine at the interaction interface to further improve the stability of the complex through disulfide bonds.
- Example 1 Based on the structure of Example 1, we further analyzed it. As shown in Figure 8, it can be seen that the interaction regions of BCP34 and AP34 have many adjacent amino acids with closely aligned side chains. The side chains of these amino acids are marked, and potential interacting amino acid pairs are listed in Table 1.
- BCP34-2 SEQ ID NO: 12, R165C on BCP34-1
- AP34-P2 SEQ ID NO: 13, S23C on AP34
- BCP34-3 SEQ ID NO: 14, N145C on BCP34-1
- AP34-P3 SEQ ID NO: 15, V26C on AP34.
- BPC34-2 (based on BCP34-1 with the addition of the R165C mutation, SEQ ID NO: 12):
- AP34-P2 (AP34 with the addition of the S23C mutation, SEQ ID NO: 13):
- BCP34-3 (BCP34 with the addition of the N145C mutation, SEQ ID NO: 14):
- AP34-P3 (AP34 with V26C mutation added, SEQ ID NO: 15):
- BCP34-2 and BCP34-3 were obtained using the *E. coli* expression and membrane purification steps described in Example 4.
- the obtained target proteins were subjected to SDS-PAGE electrophoresis, as shown in Figure 9.
- the concentrations of the two proteins were roughly estimated by using the intensity of the protein bands on the SDS-PAGE gel compared to known labeled proteins.
- Figures 9A and B show the gel images of the changes in BCP34-2 and BCP34-3 before and after denaturation. The results show that the target protein BCP34-2 was in an aggregate state before denaturation and became a monomer state after denaturation; BCP34-3 was in an aggregate state before denaturation and mostly became a monomer state after denaturation.
- AP34-P2 and AP34-P1 were synthesized by GenScript using the solid-phase synthesis method described in Example 4, and the dry powders were dissolved in solution.
- BCP34-2 and AP34-P2, and BCP34-3 and AP34-P3, were mixed at a molar ratio of approximately 1:50 and incubated at 4°C and 700 rpm for 4 h, followed by centrifugation at 13,000 rpm for 2 min.
- the obtained target proteins were characterized by SDS-PAGE electrophoresis, as shown in Figure 10.
- Figure 10A shows the reaction effect of BCP34-2 with AP34-P2; under these reaction conditions, approximately 90% of BCP34-2 and AP34-P2 formed a stable complex BCP34-2-AP34-P2.
- Figure 10B shows the reaction effect of BCP34-3 with AP34-P3; under these reaction conditions, only approximately 20% of BCP34-3 and AP34-P3 formed a stable complex BCP34-3-AP34-P3. Furthermore, the complex was verified and characterized by DTT, showing that it is stable by disulfide bonds, and that these disulfide bonds can be broken in the presence of the reducing agent DTT.
- the obtained target proteins were characterized by SDS-PAGE electrophoresis, as shown in Figure 11, which characterized the complexes BCP34-2-AP34-P4, BCP34-2-AP34-P5, BCP34-2-AP34-P6, and BCP34-2-AP34-P7.
- Figure 11A characterizes the reaction efficiency of AP34-P2, AP34-P5, AP34-P6, AP34-P7, and BCP34-2. It can be seen that their reaction efficiencies are all relatively high, with AP34-P7 and BCP34-2 exhibiting the highest efficiency, approaching complete reaction. Furthermore, the resulting complex can be broken down in the presence of DTT.
- Figure 11B characterizes the reaction efficiency of AP34-P4 and BCP34-2 forming the complex BCP34-2-AP34-P4, providing multi-dimensional verification that complex formation does not disrupt aggregate formation and that complex formation is broken down by DTT.
- AP34-P4 (SEQ ID NO: 16, truncated and retaining 29aa.): CSLVYTPKNPSFGGPAAYGTYLLNNANAQ.
- AP34-P5 (SEQ ID NO: 17, truncated and retaining 28aa.): CSLVYTPKNPSFGGPAAYGTYLLNNANA.
- AP34-P6 (SEQ ID NO: 18, truncated and retaining 27aa.): CSLVYTPKNPSFGGPAAYGTYLLNNAN.
- AP34-P7 (SEQ ID NO: 19, truncated and retaining 26aa.): CSLVYTPKNPSFGGPAAYGTYLLNNA.
- Example 6 Construction of a nanoporous biosensor using a multi-contraction nanoporous composite.
- Single-channel nanopore current measurement was performed using an amplifier based on a digital device; here, a patch-clamp amplifier was used to acquire the current signal.
- Ag/AgCl electrodes were immersed in sequencing buffer (components: 0.47M KCl, 25mM HEPES, 1mM EDTA, 5mM ATP, 25mM MgCl2 , pH 7.6), with the electrodes located in the cis and trans regions of the electrolytic cell, respectively. Sequencing libraries and nanopore reagents were added to the cis region.
- the nanopore protein or nanopore complex was diluted a certain factor using 1xPBS buffer (this protein is generally used at a concentration of 0.1 mg/ml, diluted 100-fold or 10-fold with PBS for embedding), and then subjected to an external electric field.
- a single nanopore or a nanopore composite is inserted into a phospholipid bilayer composed of diacylphosphatidylcholine (DPhPC, 1,2-diphytanoyl-sn-glycero-3-phosphocholine) to form a nanopore biosensor.
- DPhPC diacylphosphatidylcholine
- An applied voltage is then used to obtain the current amplitude value of a single pore protein.
- Example 7 Using the nanopore sensor from Example 6 for DNA sequencing
- Nucleic acid sequence preparation The artificially synthesized sequence SEQ ID NO: 20 was inserted into the multiple cloning site of the PUC57 plasmid. Using the primer combination of SEQ ID NO: 21 and SEQ ID NO: 22, the 3.5kb test sequence SEQ ID NO: 20 was prepared by PCR amplification.
- Sequencing library preparation Construct the sequencing library from the sequence to be tested, SEQ ID NO: 20.
- this library further binds to a cholesterol-containing single-stranded DNA (SEQ ID NO: 26, cholesterol attached to the 5' end of the DNA), forming the structure shown in Figure 12B.
- Sequencing libraries were mixed with sequencing buffer (sequencing buffer: 0.47M KCl, 25mM HEPES, 1mM EDTA, 5mM ATP, 25mM MgCl2, pH 7.6) and added to the nanopore biosensor. After applying an external voltage of 0.14V or 0.18V, DNA was observed to be captured by the nanopore, generating characteristic retardation current amplitude values. Furthermore, the current amplitude value changed as the DNA moved through the nanopore. Different DNA sequences produced different retardation current amplitude values. Single-stranded DNA containing cholesterol can bind to the phospholipid bilayer, which helps the nanopore capture the sequencing library and reduces the loading volume.
- SEQ ID NO: 21 gccatcagattgtgtttgttagt.
- SEQ ID NO: 22 gcttacggttcactactcacga.
- SEQ ID NO: 29 ggttgtttctgttggtgctgatattgct.
- SEQ ID NO: 24 gcaatatcagcaccaacagaaacaacctttgaggcgagcggtcaa.
- SEQ ID NO: 26 cholesterol-ttgaccgctcgcctc.
- Figure 13 shows the current changes and local details of the library DNA passing through the nanoporin BCP34-1 under an applied voltage of 0.18V.
- Figure 14 shows the current changes and local details of the library DNA passing through the dual-detector nanoporin BCP34-1-AP34-P1 under an applied voltage of 0.18V. It can be seen that under an applied voltage of 0.18V, the opening current of nanoporin BCP34 is 230-250 pA, with a sequencing amplitude of approximately 40 pA; the opening current of the dual-detector nanoporin is 130-150 pA, with a sequencing amplitude of approximately 20 pA.
- the current signal read in the DNA homopolymer region differs from that of the mutant protein, exhibiting more stepwise signals (this phenomenon is also observed in different dual-detector nanoporin mutants, with varying degrees of differentiation; a comparison can be seen in Figure 22). It is evident that, under the same voltage, the opening current of the dual-detector porin BCP34-1-AP34-P1 is lower than that of the original single-detector protein BCP34-1, indicating that a second detector is formed outside the original detector of the porin. A dual-detector nanoporin was successfully constructed. Furthermore, the step-like resolution of the homopolymer region indicates that the nanoporin complex provides more information, thus suggesting that this dual-detector nanoporin has the potential to improve the resolution of nanopore sequencing.
- BCP34-1-AP34-P8 is based on AP34 with the main contraction region modified to A39S
- BCP34-1-AP34-P9 is based on AP34 with the main contraction region modified to A39N.
- AP34-P8 (SEQ ID NO: 27, based on AP34 with modifications to the main contraction region A39S):
- AP34-P9 (SEQ ID NO: 28, based on AP34 with modifications to the main contraction region A39N)
- FIG 15 shows the current changes and local details of the library DNA passing through the nanoporin BCP34-1-AP34-P8 under an applied voltage of 0.18V.
- Figure 16 shows the current changes and local details of the library DNA passing through the dual-detector nanoporin BCP34-1-AP34-P9 under an applied voltage of 0.18V. It can be seen that the modification of the main amino acids in the contraction region will cause changes in its ability to resolve characteristic sequences.
- Figure 17 shows the current changes and local details generated when library DNA passes through nanoporous protein BCP34-2-AP34-P2 under an applied voltage of 0.18V.
- Figure 18 shows the current changes and local details generated when library DNA passes through the dual-detector nanoporous protein BCP34-3-AP34-P3 under an applied voltage of 0.18V.
- the opening current of the dual-detector nanoporous protein is 130-150 pA, with a sequencing amplitude of approximately 20 pA.
- the DNA homopolymer recognition current signal differs from that of the single-pore protein, as shown in Figure 22 (where A compares the current signals of BCP34-1 with a single contraction zone and BCP34-1-AP34-1 with a double contraction zone; B compares the current signals of different AP34 main signal regions; C compares the current signals of modified crosslinking sites; and D compares the current signals of AP34 peptides of different lengths).
- Figure 19 shows the current changes and local details generated when library DNA passes through the nanopore complex BCP34-2-AP34-P4 under an applied voltage of 0.18V.
- Figure 20 shows the current changes and local details generated when library DNA passes through the nanopore complex BCP34-2-AP34-P5 under an applied voltage of 0.18V.
- Figure 21 shows the current changes and local details generated when library DNA passes through the nanopore complex BCP34-2-AP34-P7 under an applied voltage of 0.18V.
- Figure 22 shows the characteristics of the current signals of the DNA homopolymer region of different nanopores in Example 7;
- A is a comparison of the current signals of the single contraction zone pore BCP34-1 and the double contraction zone pore BCP34-1-AP34-1;
- B is a comparison of the current signals of the main signal regions of different AP34;
- C is a comparison of the current signals of the modified crosslinking sites;
- D is a comparison of the current signals of AP34 peptides of different lengths.
- the vertical axis represents current (pA) and the horizontal axis represents time (min).
- Each figure shows the current signal within a certain time scale (the specific values can be ignored because: 1. All the protein mutations or truncations involved in this embodiment are mainly to change the sequencing signal and improve the stability of the complex, but do not change the sequencing speed. Therefore, we do not focus on the sequencing speed, that is, we do not need to focus on the time (horizontal axis difference) used to sequence a complete read; 2.
- the sequencing outputs are not generated at completely the same time.
- the sequencing signal reads extracted here are randomly selected and have a sequencing map of representative complete reads of the current sequencing output signal. Therefore, the sequencing time (i.e., the horizontal axis) here has no reference significance and can be ignored).
- This invention develops novel nanopore complexes with sequencing capabilities to overcome the limitations of pore protein types in the field of nanopore sequencing, and to develop and explore more single-molecule sequencers with high accuracy, high integration and high stability.
- This invention is based on the novel porin monomer BCP34 and its mutants.
- AP3434 which belongs to the same genus as the protein
- truncating and modifying the accessory protein monomer AP3434 a nanopore complex is constructed together with the porin monomer BCP34 to create a nanopore protein complex with including but not limited to two contraction regions, and it is applied in the field of nanopore sequencing.
Landscapes
- Life Sciences & Earth Sciences (AREA)
- Health & Medical Sciences (AREA)
- Chemical & Material Sciences (AREA)
- Engineering & Computer Science (AREA)
- Organic Chemistry (AREA)
- Genetics & Genomics (AREA)
- Molecular Biology (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Biomedical Technology (AREA)
- General Health & Medical Sciences (AREA)
- Biotechnology (AREA)
- Biochemistry (AREA)
- Wood Science & Technology (AREA)
- Zoology (AREA)
- Biophysics (AREA)
- Immunology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Microbiology (AREA)
- General Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Hematology (AREA)
- Analytical Chemistry (AREA)
- Cell Biology (AREA)
- Urology & Nephrology (AREA)
- Medicinal Chemistry (AREA)
- Plant Pathology (AREA)
- Food Science & Technology (AREA)
- General Physics & Mathematics (AREA)
- Pathology (AREA)
- Gastroenterology & Hepatology (AREA)
- Peptides Or Proteins (AREA)
Abstract
提供了一种纳米孔蛋白复合物、其构建方法及应用。其中,纳米孔复合物包括:纳米孔蛋白和附属蛋白,纳米孔蛋白包括纳米孔腔,纳米孔腔由多个孔蛋白单体聚合而成;附属蛋白由多个附属蛋白单体聚合而成,附属蛋白的N端嵌入纳米孔腔内,并与纳米孔腔共同形成连续通道;其中,按照穿过连续通道的分析物的移动方向,连续通道包括顺次连通的第一传感区和第二传感区,第一传感区由纳米孔蛋白的一部分形成,第二传感区由附属蛋白的部分或全部形成;纳米孔蛋白单体选自SEQ ID NO:1所示氨基酸序列的蛋白及其变体,附属蛋白单体选自SEQ ID NO:3所示氨基酸序的蛋白及其变体。
Description
本发明涉及纳米孔测序技术领域,具体而言,涉及一种纳米孔蛋白复合物、其构建方法及应用。
纳米孔测序技术作为新兴起的单分子测序技术,其凭借着高通量、长读长、快速度、原位检测和无标记操作等独特优势,给基因测序行业带来了颠覆性的改变,在分子生物学、医学、流行病学和生态学等许多领域的基础理论研究以及生物医学临床实践中具有广泛的应用。
纳米孔测序技术是基于电信号的测序技术,该技术可同时应用于测试核苷酸,氨基酸或聚糖的序列等以及碱基、氨基酸或聚糖修饰(如甲基化和酰化、磷酸化、羟基化、氧化、还原、糖基化、脱羧、脱氨等)等。其核心元件纳米孔蛋白插在膜上,发挥信号传感器作用,它将两个装有电解液的电解室分开,当电压施加在电解室之间时,会产生稳定电流,当待测分析物进入纳米孔时会对离子的流动造成阻碍从而导致电流信号波动。通过实时记录待测物逐一通过纳米孔蛋白时产生的连续阻滞的电信号,并借助机器学习分析并解码电流信号,从而对待测物的序列信息进行实时测序。
纳米孔测序的理论基础为核酸分子在马达蛋白的控制下碱基依次通过嵌于膜中的纳米孔蛋白中央通道,产生的电流信号经底层电路放大后由碱基的识别算法转化为最终序列信息,该路线目前已充分论证,具有可行性,现市场上商用的纳米孔测序仪主要为英国牛津纳米孔技术公司已研制MinION、GridION和PromethION等一系列纳米孔测序仪,齐碳科技也已研制QNome-3841纳米孔测序仪等,然而这些纳米孔测序仪在测序准确度、通量以及芯片稳定性等方面仍存在较大不足,无法满足分子生物学研究的终极需求。因此,本领域中需要高准确度、高集成度以及高稳定性的单分子测序仪。
造成纳米孔测序准确率低的主要原因之一是其核心元件纳米孔蛋白的信号传感器的单一性和低分辨率。纳米孔蛋白内部最窄的收缩区域,是对分析物电流信号变化特征最具区分力的部分,可以作为电流信号的内部传感(sensor)区域,该区域需要足够锐利,以在横向与纵向上均有高的空间分辨能力。
在研究的孔道蛋白中,潜在可用于测序的蛋白主要有三类:作为细胞内外各种生物大分子与小分子物质运输通道的转运(transporter)蛋白;细菌或其它机体产生的破坏细胞膜通透性的成孔毒素(pore-forming toxin)蛋白;为病毒侵染宿主时提供基因组输运通道的病毒连接体(viral connector)。目前,仅有耻垢分枝杆菌孔蛋白A(MspA)、curli特异性转运通道(CsgG)等少数天然蛋白符合测序要求。造成天然蛋白难以测序的原因有重组蛋白体外表达纯化系统对蛋白稳定性的考验,蛋白聚体的稳定性和对称性,蛋白内腔收缩区的形状和孔大小等。
市场上绝大多数孔蛋白收缩区只有一处,只能与分析物序列的少数部分接触。在测序过程中,序列之间相互影响,产生到底层序列的电导图十分复杂,特别对于均聚物序列,当前算法难以准确分辨,这也是导致了目前纳米孔测序仪分辨率较差,从而导致错误率相对较高,
推广较慢的主要原因之一。已有文献证明,当单个纳米孔中具有一个以上的传感区时,可以帮助获得额外的序列信息,提供更多解析均聚物区域的机会,克服纳米孔测序准确率不高的劣势。目前已有极少数具有两个传感器的孔蛋白,但它们的传感器间隔也太短,也不足以提供足够的空间分辨能力。尤其在测试DNA均聚物上,当前纳米孔测得的信号常不能准确判定该区域重复碱基的个数,从而降低了纳米孔测序准确率。
因此,开发具有更多的具有两个或两个以上收缩区域的纳米孔蛋白且实现高精度测序的任务迫在眉睫。
发明内容
本发明的主要目的在于提供一种纳米孔蛋白复合物,以解决现有技术中大多仅具有单一收缩区的纳米孔蛋白应用于纳米孔测序时存在的测序准确率低的问题。
为了实现上述目的,根据本发明的一个方面,提供了一种纳米孔蛋白复合物,该纳米孔蛋白复合物包括:纳米孔蛋白以及附属蛋白,纳米孔蛋白由多个孔蛋白单体聚合而成,且多个孔蛋白单体聚合形成中空的纳米孔腔;附属蛋白由多个附属蛋白单体聚合而成,附属蛋白的至少部分位于纳米孔腔内,附属蛋白与纳米孔腔共同形成连续通道;其中,按照穿过连续通道的分析物的移动方向,连续通道包括顺次连通的第一传感区和第二传感区,第一传感区由纳米孔蛋白的至少部分形成,第二传感区由附属蛋白的至少部分形成;孔蛋白单体选自如下任意一种蛋白:1)具有SEQ ID NO:1所示氨基酸序列的蛋白;2)与SEQ ID NO:1所示氨基酸序列至少具有50%同一性且具有聚合形成纳米孔蛋白功能的蛋白;或3)在SEQ ID NO:1的基础上经过取代、缺失或添加一个或几个氨基酸,且具有聚合形成纳米孔蛋白功能的蛋白;附属蛋白单体选自如下任意一种蛋白:i)具有SEQ ID NO:3所示氨基酸序列的蛋白;ii)与SEQ ID NO:3所示氨基酸序列至少具有50%同一性且具有如下功能的蛋白:与孔蛋白单体结合,并随孔蛋白单体聚合形成纳米孔蛋白而一起聚合形成附属蛋白;或iii)在SEQ ID NO:3所示氨基酸序列上经过取代、缺失或添加一个或几个氨基酸,且具有如下功能的蛋白:与孔蛋白单体结合,并随孔蛋白单体聚合形成纳米孔蛋白而一起聚合形成附属蛋白。
进一步地,纳米孔蛋白和附属蛋白来源于同一个种属;任选地,附属蛋白的全部或附属蛋白的N端部分位于纳米孔蛋白的纳米孔腔内;优选地,附属蛋白通过共价或非共价作用附接于纳米孔蛋白的纳米孔腔内;更优选地,纳米孔蛋白与附属蛋白以九聚体的形式存在。
进一步地,附属蛋白单体为SEQ ID NO:3所示的氨基酸序列截短的突变体,截短的突变体选自如下任意一种突变体:仅保留N端23-67位氨基酸、仅保留N端23-57位氨基酸、仅保留N端23-52位氨基酸、仅保留N端23-54位氨基酸、仅保留N端23-51位氨基酸、仅保留N端23-50位氨基酸、仅保留N端23-49位氨基酸、仅保留N端23-48位氨基酸或仅保留N端23-47位氨基酸。
进一步地,附属蛋白单体的长度为24~45个氨基酸;优选地,附属蛋白单体的氨基酸序列来自SEQ ID NO:3的如下残基位置区间或其突变体:第23~48位、第23~49位、第23~50位、第23~51位、第23~54位或第23~57位。
进一步地,孔蛋白单体选自在SEQ ID NO:1的如下任意一组或多组中的至少一个氨基酸位点发生突变的蛋白:1)S71、N74、G75和F76;优选地,S71突变为G、A或T;N74突
变为G、A、S或T;G75突变为A、S、T或Q;F76突变为A、S、T、N或Q;2)E162、R196、S200和S216;优选地,E162突变为A、G、V、L、I、Y、F或W;R196突变为A、G、V、L、I、Y、F或W;S200突变为A、G、V、L、I、Y、F或W;S216突变为A、G、V、L、I、Y、F或W;3)R103、E104、E112、R113、K114、R117、R119、D120、K122、D124、K154、R165、D199、R207、K209、K210、E213或E215;优选地,R103突变为A、G、S、T、N或Q;E104突变为K、R、A、G、S、T、N或Q;E112突变为K、R、A、G、S、T、N或Q;R113突变为A、G、S、T、N或Q;K114突变为A、G、S、T、N或Q;R117突变为A、G、S、T、N或Q;R119突变为A、G、S、T、N或Q;D120突变为K、R、A、G、S、T、N或Q;K122突变为A、G、S、T、N或Q;D124突变为K、R、A、G、S、T、N或Q;K154突变为A、G、S、T、N或Q;R165突变为A、G、S、T、N或Q;D199突变为A、G、S、T、N或Q;R207突变为A、G、S、T、N、Q、D或E;K209突变为A、G、S、T、N、Q、D或E;K210突变为A、G、S、T、N、Q、D或E;E213突变为A、G、S、T、N或Q;E215突变为A、G、S、T、N或Q。
进一步地,附属蛋白单体选自在SEQ ID NO:3所示氨基酸序列的如下至少一个氨基酸发生取代:A39、T42或N46;优选地,A39突变为S、T、N、G、V、L、I或Q;T42突变为A、G或S;N46突变为A、G或S。
进一步地,第一传感区和第二传感区具有相同或不同的最小孔直径;优选地,最小孔直径为1-3nm。
进一步地,附属蛋白单体选自在SEQ ID NO:3所示氨基酸序列上的如下至少一个氨基酸突变为半胱氨酸或被非天然氨基酸取代的蛋白:S23、S24、L25、T28、K30、N31、S33、F34、N47、A50、Q51、N52或Q53;和/或孔蛋白单体选自在SEQ ID NO:1的基础上的如下至少一个氨基酸突变为半胱氨酸或被非天然氨基酸取代的蛋白:N145、T148、K154、L156、L160、S161、R165、S195、Q197、D199、F203、Y205、K209、K210、L211、E213、E215、G217、S219或N221。
进一步地,孔蛋白单体和附属蛋白单体存在如下至少一个位点组合形式的突变,且各位点组合中相应位点的氨基酸突变为半胱氨酸或被非天然氨基酸取代:1)SEQ ID NO:1上的R165和SEQ ID NO:3上的S23;2)SEQ ID NO:1上的S195和SEQ ID NO:3上的S24;3)SEQ ID NO:1上的S219和SEQ ID NO:3上的L25;4)SEQ ID NO:1上的N221和SEQ ID NO:3上的L25;5)SEQ ID NO:1上的N145C和SEQ ID NO:3上的V26C;6)SEQ ID NO:1上的D199和SEQ ID NO:3上的T28;7)SEQ ID NO:1上的D199和SEQ ID NO:3上的K30;8)SEQ ID NO:1上的E215和SEQ ID NO:3上的K30:9)SEQ ID NO:1上的E213和SEQ ID NO:3上的N31:10)SEQ ID NO:1上的K154和SEQ ID NO:3上的S33:11)SEQ ID NO:1上的E213和SEQ ID NO:3上的S33;12)SEQ ID NO:1上的E213和SEQ ID NO:3上的N47;13)SEQ ID NO:1上的L211和SEQ ID NO:3上的A50;14)SEQ ID NO:1上的L211和SEQ ID NO:3上的Q51;15)SEQ ID NO:1上的K209和SEQ ID NO:3上的N52;16)SEQ ID NO:1上的K210和SEQ ID NO:3上的Q53。
进一步地,任一位点组合中相应位点的氨基酸均突变为半胱氨酸。
进一步地,通过修饰的方式使氨基酸取代为非天然氨基酸,修饰为直接修饰或间接修饰;优选地,直接修饰包括通过氨基酸的侧链基团自发反应或氧化反应进行修饰;优选地,间接
修饰包括通过附接化学小分子进行化学修饰;优选地,化学小分子包括含官能团的化学交联剂,且官能团对二硫苏糖醇有抗性。
进一步地,至少部分附属蛋白单体中突变为半胱氨酸的位点与至少部分孔蛋白单体中突变为半胱氨酸的位点之间形成二硫键。
为了实现上述目的,根据本发明的第二个方面,提供了一种纳米孔传感器,该纳米孔传感器包括膜和上述任一种纳米孔蛋白复合物,其中,纳米孔蛋白复合物中的纳米孔蛋白定位在膜上,纳米孔蛋白的纳米孔腔和附属蛋白共同形成跨膜的连续通道。
进一步地,膜包括两亲分子层;优选地,膜是由二脂酰磷脂酰胆碱组成的磷脂双分子层、两嵌段共聚物或三嵌段共聚物。
根据本发明的第三个方面,提供了一种纳米孔测序系统,纳米孔测序系统包括上述任一种纳米孔传感器,该纳米孔测序系统还包括:导电溶液、跨膜提供电压电势的正负电极以及用于测量通过连续通道的电信号的测量设备;其中,纳米孔传感器位于导电溶液中,且将导电溶液分割为第一腔室和第二腔室。
进一步地,纳米孔测序系统还包括待测分子,待测分子选自多核苷酸、多肽和多糖,优选地,待测分子选自含有均聚物的多核苷酸;优选地,待测分子可瞬时性地位于连续通道内,且待测分子的一端位于第一腔室,另一端位于第二腔室。
根据本发明的第四个方面,提供了一种纳米孔测序的方法,该方法包括:使纳米孔测序系统与待测分子接触;跨膜施加电势,使得待测分子进入连续通道;以及一次或多次测量待测分子相对于连续通道移动时产生的电信号,由此获得待测分子的序列。
进一步地,电信号包括电流阻滞信号强度、电流阻滞持续时间或电流阻滞事件发生的间隔时间。
进一步地,待测分子为多核苷酸,多核苷酸中的核苷酸与连续通道内的第一收缩区域和第二收缩区域相互作用,并且其中第一收缩区域和第二收缩区域中的每个收缩区域能够区分不同的核苷酸,使得通过连续通道的总电流受到第一收缩区域和第二收缩区中的每个收缩区域与定位在区域中的每个区域处的核苷酸之间的相互作用的影响。
进一步地,使用核酸结合蛋白来控制多核苷酸相对于连续通道孔的移动。
根据本发明的第四个方面,提供了上述纳米孔蛋白复合物的制备方法,该制备方法包括:分别构建孔蛋白单体的表达载体和附属蛋白单体的表达载体;通过在感受态细胞中共表达孔蛋白单体的表达载体和附属蛋白单体的表达载体,并经过分离纯化来获得纳米孔蛋白复合物;或通过体外重组的方式构建纳米孔蛋白复合物。
应用本发明的技术方案,利用附属蛋白嵌入纳米孔蛋白的纳米孔腔内,进而在纳米孔蛋白自身具有的一个收缩区的基础上,引入了第二个收缩区,形成了具有两个收缩区(或传感区)的连续通道的纳米孔复合物,利用这样的复合物进行纳米孔测序时,待测核酸分子在依次通过第一收缩区和第二收缩区时,核酸分子不同空间部位的核苷酸分别与两个收缩区发生不同的相互作用,进而产生更多的电信号(比如电流)变化信息,因而相比只具有单收缩区
的纳米孔蛋白,利用本申请的纳米孔蛋白复合物进行纳米孔测序,能够更准确地确定均聚物序列的信息,特别是针对较长的均聚物序列片段。
构成本申请的一部分的说明书附图用来提供对本发明的进一步理解,本发明的示意性实施例及其说明用于解释本发明,并不构成对本发明的不当限定。在附图中:
图1示出了根据本发明的实施例1中的(含信号肽)BCP34-AP34形成的多聚体的预测结构示意图,其中,A为侧视图(sideview),B为俯视图(topview)。
图2示出了本发明的实施例1中的成熟的(无信号肽)BCP34-AP34形成的多聚体的预测结构示意图,A为侧视图(sideview),B为俯视图(topview)。
图3示出了本发明的实施例1中的BCP34-AP34形成的多聚体的预测结构中纳米孔收缩区示意图。
图4示出了本发明的实施例1中的BCP34-AP34形成的多聚体的预测结构中纳米孔收缩区侧链示意图。
图5示出了本发明的实施例3中的共表达法构建的纳米孔复合物的纯化结果图,其中,A是Strep柱纯化的结果图,B是Strep柱纯化后又通过Ni柱纯化的结果图。
图6示出了本发明的实施例3中的共表达法构建的纳米孔复合物TEV酶切的纯化结果图,其中,A是纳米孔复合物未变性的结果图,B是纳米孔复合物变性后的结果图。
图7示出了本发明的实施例4中的体外重组法构建的纳米孔复合物的纯化结果图,其中,A是未变性前的纳米孔复合物的电泳图,B是变性后的纳米孔复合物的电泳图。
图8示出了本发明的实施例5中的BCP34与AP34相互作用区域的结构示意图。
图9示出了本发明的实施例5中BCP34的交联改造突变体纯化图;A是变性前的图,B是变性后的图。
图10示出了本发明的实施例5中交联改造构建的纳米孔复合物的SDS-PAGE电泳结果;A显示的是BCP34-2与AP34-P2形成纳米孔复合物的反应效率,B显示的是BCP34-3与AP34-P3形成纳米孔复合物的反应效率。
图11示出了本发明的实施例5中不同长度的AP34构建的纳米孔复合物的SDS-PAGE电泳结果;A显示的是AP34-P2、AP34-P5、AP34-P6、AP34-P7分别与BCP34-2形成纳米孔复合物的反应效率;B显示的是AP34-P4与BCP34-2形成纳米孔复合物的反应效率。
图12示出了本发明的实施例7中含有解旋酶测序文库示意图以及测序文库与含有胆固醇的单链DNA配对结构示意图,其中,A为含有解旋酶测序文库示意图,B为测序文库与含有胆固醇的单链DNA配对结构示意图;a:上链;b:下链;c:双链目的片段;d:解旋酶;e:胆固醇标记的单链DNA。
图13示出了本发明的实施例7中的BCP34-1纳米孔蛋白测序电流信号图。
图14示出了本发明的实施例7中的BCP34-1-AP34-P1纳米孔复合物测序电流信号图。
图15示出了本发明的实施例7中的BCP34-1-AP34-P8纳米孔复合物测序电流信号图。
图16示出了本发明的实施例7中的BCP34-1-AP34-P9纳米孔复合物测序电流信号图。
图17示出了本发明的实施例7中的BCP34-2-AP34-P2纳米孔复合物测序电流信号图。
图18示出了本发明的实施例7中的BCP34-3-AP34-P3纳米孔复合物测序电流信号图。
图19示出了本发明的实施例7中的BCP34-2-AP34-P4纳米孔复合物测序电流信号图。
图20示出了本发明的实施例7中的BCP34-2-AP34-P5纳米孔复合物测序电流信号图。
图21示出了本发明的实施例7中的BCP34-2-AP34-P7纳米孔复合物测序电流信号图。
图22示出了本发明的实施例7中的不同纳米孔对DNA均聚物区电流信号的特征;A为单收缩区孔BCP34-1与双收缩区BCP34-1-AP34-1电流信号对比;B为不同AP34的主信号区的电流信号对比;C为不同交联位点的改造电流信号对比;D为不同长短的AP34多肽的电流信号对比。
需要说明的是,在不冲突的情况下,本申请中的实施例及实施例中的特征可以相互组合。下面将结合实施例来详细说明本发明。
术语解释:
孔蛋白,在本申请中也叫纳米孔蛋白,是能够用于进行纳米孔测序的孔蛋白。
收缩区或传感区:本申请中是指纳米孔蛋白或者纳米孔蛋白内腔与附属蛋白共同构成的跨膜通道中最窄的内径区域。其中,纳米孔蛋白单独形成的跨膜通道有一个收缩区,而纳米孔蛋白与附属蛋白共同形成纳米孔复合物的跨膜通道物有构成了其它收缩区。
如背景技术部分所提到的,现有的用于纳米孔测序的纳米孔蛋白仅具有一个收缩区,难以满足对含有均聚物(含有多个连续重复核苷酸的寡聚核苷酸分子)的测序区分度的需求。
为了改善这一现状,本申请通过孔蛋白数据库挖掘、结构预测与解析、蛋白组装,蛋白改造,纳米孔测序性能测试和表征等技术手段筛选或构建出了一种新型的具有多个传感区域的纳米孔蛋白及其突变体。该纳米孔蛋白复合物由孔蛋白单体BCP34和其附属蛋白单体AP34共同构建并聚合形成具有多个收缩区的跨膜蛋白。该类纳米孔蛋白复合物可以在体内共表达获得,也可以在体外重组表达构建,形成高效率的稳定复合物。进一步通过利用该复合物构建了纳米孔传感器,构建纳米孔测序体系,并对多聚均聚物核酸分子进行了检测验证,证明本申请提供的改进的含有两个或两个以上收缩区的纳米孔蛋白复合物有助于提升对均聚核酸分子的区分准确度。因而应用于纳米孔测序中,能输出稳定测序信号,并对均聚物序列具备高的分辨率,在提高测序的准确率,改善纳米孔测序准确率不足方面有极高潜力。
在上述研究结果的基础上,申请人提出了本申请的一系列技术方案。在一种典型的实施方式中,提供了一种纳米孔复合物,该纳米孔复合物包括:纳米孔蛋白和附属蛋白,纳米孔
蛋白由多个孔蛋白单体聚合而成,且多个孔蛋白单体聚合形成中空的纳米孔腔;附属蛋白由多个附属蛋白单体聚合而成,附属蛋白的至少部分嵌入纳米孔腔内,附属蛋白与纳米孔腔共同形成连续通道;其中,按照穿过连续通道的分析物的移动方向,连续通道包括顺次连通的第一传感区和第二传感区,第一传感区由纳米孔蛋白的一部分形成,第二传感区由附属蛋白的部分或全部形成;
其中,纳米孔蛋白单体选自如下任意一种蛋白:1)具有SEQ ID NO:1所示氨基酸序列的蛋白;2)与SEQ ID NO:1所示氨基酸序列至少具有50%同一性且具有聚合形成纳米孔蛋白功能的蛋白;或3)在SEQ ID NO:1的基础上经过取代、缺失或添加一个或几个氨基酸,且具有聚合形成纳米孔蛋白功能的蛋白;
附属蛋白单体选自如下任意一种蛋白:i)具有SEQ ID NO:3所示氨基酸序列的蛋白;ii)与SEQ ID NO:3所示氨基酸序列至少具有50%同一性且且具有如下功能的蛋白:与孔蛋白单体结合,并随孔蛋白单体聚合形成上述纳米孔蛋白而一起聚合形成附属蛋白;或iii)在SEQ ID NO:3所示氨基酸序列上经过取代、缺失或添加一个或几个氨基酸,且具有如下功能的蛋白:与孔蛋白单体结合,并随孔蛋白单体聚合形成上述纳米孔蛋白而一起聚合形成附属蛋白。
纳米孔的收缩区能够区分待测核酸穿过孔的不同核苷酸(即具有不同的碱基),每个核苷酸与收缩部相互作用时产生电流变化,可以转换成对应的序列信息。但是对DNA均聚物测序而言,现有单个收缩区测得的信号可能区分度不足以区分和分辨单碱基电流变化,进而不能仅根据测得的信号的大小来准确确定均聚物的长度。而本申请所提供的上述新型的纳米孔蛋白复合物有助于改善对核苷酸均聚物的测序信号的区分度。
本申请的上述纳米孔复合物,利用附属蛋白嵌入纳米孔蛋白的纳米孔腔内,进而在纳米孔蛋白自身具有的一个收缩区的基础上,引入了第二个收缩区,形成了具有两个收缩区(或传感区)的连续通道的纳米孔复合物,利用这样的复合物进行纳米孔测序时,待测核酸分子在依次通过第一收缩区和第二收缩区时,核酸分子不同空间部位的核苷酸分别与两个收缩区发生不同的相互作用,进而产生更多的电信号(比如电流)变化信息,因而相比只具有单收缩区的纳米孔蛋白,利用本申请的纳米孔蛋白复合物进行纳米孔测序,能够更准确地确定均聚物序列的信息,特别是针对较长的均聚物序列片段。也即,该纳米孔蛋白复合物能够应用于纳米孔测序中,输出稳定测序信号,并对均聚物序列具备高的分辨率,有极高潜力提高测序的准确率,改善纳米孔测序准确率不足的劣势。
上述纳米孔蛋白复合物中,纳米孔蛋白和附属蛋白优选来源于同一个种属,这样有助于两种蛋白能更好的特异性结合。
需要说明的是,本申请中的附属蛋白是没有孔结构的一类蛋白。但是其单体能与主孔蛋白单体非共价结合,并在主孔蛋白单体的自组装过程中也一起聚合形成9个主孔蛋白单体+9个附属蛋白单体的聚体复合物。
需要说明的是,上述纳米孔复合物中,纳米孔蛋白是基于我们以往提供的孔蛋白单体BCP34及其突变体改造而来。纳米孔蛋白是一个九聚体,其单体序列是从11000米深的马里亚纳海沟的宏基因组数据库挖掘而来,野生型全长蛋白序列为SEQ ID NO:1:
MKRFIAFAVSMTLVGCASFSPPKGQTSIRDRAQPLSATVTRNALTKLPPPLAP34IPAAVYNIKDQTGQYKPSPSNGFSTAVMQGATSVLVKALLDSRWFIPLEREGLQNLLTERKIIRARDSK
KDLSNLAAASVIIEGSIIAYDSNVRTGGAGAKYLGIGLSEQYREDQVTVNLRAINVNNGRILQSVTSTKMIFSRQLDSGAFGYIRFKKLLEIESGYSYNEPAQLCVIDAIESALIQLIYEGVVGGTWKLKNPADIDSPIFTHYAGQNGPEGQAIM,其中,前15个氨基酸(下划线部分)是该孔蛋白单体BCP34的信号肽区。
成熟的孔蛋白单体BCP34无信号肽区,其氨基酸序列为SEQ ID NO:2:
需要说明的是,本申请中的同源性指两个氨基酸序列之间的“序列同一性(identity)”,即序列之间相同氨基酸的百分比。评估氨基酸或核苷酸之间序列同一性程度的方法是本领域技术人员已知的。例如,氨基酸序列同一度通常用序列分析软件测量。例如,可以通过NCBI数据库的BLAST程序来确定。关于序列同一性的测定,参见例如:分子生物学中的序列分析,von Heinje,G.,Academic Press,1987和序列分析引物,Gribskov,M.和Devereux,J.,eds.,M Stockton Press,New York,1991。
上述与SEQ ID NO:1所示的BCP34纳米孔蛋白单体具有50%、55%、60%、65%、70%、75%、80%、85%、90%、95%、99%以上(比如85%、86%、87%、88%、89%、90%、91%、92%、93%、94%、95%、96%、97%、98%、98.5%、99%、99.5%、99.6%、99.7%、99.8%以上,甚至99.9%以上)同源性且具有聚合形成纳米孔蛋白功能的蛋白质,其活性位点、活性口袋、活性机制、蛋白结构等均与BCP34纳米孔蛋白单体大概率相同。
类似地,与SEQ ID NO:3所示的附属蛋白单体具有50%、55%、60%、65%、70%、75%、80%、85%、90%、95%、99%以上(比如85%、86%、87%、88%、89%、90%、91%、92%、93%、94%、95%、96%、97%、98%、98.5%、99%、99.5%、99.6%、99.7%、99.8%以上,甚至99.9%以上)同源性且具有与孔蛋白单体结合,并随孔蛋白单体聚合形成上述纳米孔蛋白而一起聚合形成附属蛋白的功能的蛋白质,其活性位点、活性口袋、活性机制、蛋白结构等均与AP34单体大概率相同。
氨基酸残基可以根据本领域公知和约定的标准三个字母或一个字母的氨基酸编码来表示。本文中,氨基酸残基缩写如下:丙氨酸(Ala;A)、天冬酰胺(Asn;N)、天冬氨酸(Asp;D)、精氨酸(Arg;R)、半胱氨酸(Cys;C)、谷氨酸(Glu;E)、谷氨酰胺(Gln;Q)、甘氨酸(Gly;G)、组氨酸(His;H)、异亮氨酸(Ile;I)、亮氨酸(Leu;L)、赖氨酸(Lys;K)、蛋氨酸(Met;M)、苯丙氨酸(Phe;F)、脯氨酸(Pro;P),丝氨酸(Ser;S)、苏氨酸(Thr;T)、色氨酸(Trp;W)、酪氨酸(Tyr;Y)和缬氨酸(Val;V)。
保守氨基酸取代或替换在本领域是公知的,例如,保守的氨基酸取代优选如下组(1)-(5)中的一个氨基酸残基被同一组中的另一个氨基酸取代:(1)较小的脂族非极性或弱极性残基:Ala、Ser、Thr、Pro和Gly;(2)带极性负电荷的残基及其(不带电荷的)酰胺:Asp、Asn、Glu和Gln;(3)带极性正电荷的残基:His、Arg和Lys;(4)较大的脂族非极性残基:Met、Leu、Ile、Val和Cys;和(5)芳族残基:Phe、Tyr和Trp。特别优选的保守氨基酸取代如下:Ala被Gly或Ser取代;Arg被Lys取代;Asn被Gln或His取代;Asp被Glu取代;
Cys被Ser取代;Gln被Asn取代;Glu被Asp取代;Gly被Ala或Pro取代;His被Asn或Gln取代;Ile被Leu或Val取代;Leu被Ile或Val取代;Lys被Arg、Gln或Glu取代;Met被Leu、Tyr或Ile取代;Phe被Met、Leu或Tyr取代;Ser被Thr取代;Thr被Ser取代;Trp被Tyr取代;Tyr被Trp或Phe取代;并且Val被Ile或Leu取代。
本领域技术人员也可以根据现有技术中的“blosum62评分矩阵”等本领域技术人员熟知的氨基酸替换规则对氨基酸进行保守替换。
BCP34蛋白的突变体的包括但不限于在SEQ ID NO:1的S71、N74、G75和F76中的一个或多个位置处发生氨基酸突变的变体;优选地,S71位置处的突变包括突变为G、A或T;N74位置处的突变包括突变为G、A、S或T;G75位置处的突变包括突变为A、S、T、N或Q;F76位置处的突变包括突变为G、A、S、T、N或Q;E162、R196、S200和S216的一个或多个位置处具有氨基酸的突变;优选地,E162位置处的突变包括突变为A、G、V、L、I、Y、F或W;R196位置处的突变包括突变为A、G、V、L、I、Y、F或W;S200位置处的突变包括突变为A、G、V、L、I、Y、F或W;S216位置处的突变包括突变为A、G、V、L、I、Y、F或W;R103、E104、E112、R113、K114、R117、R119、D120、K122、D124、K154、R165、D199、R207、K209、K210、E213和E215的一个或多个位置处具有氨基酸的突变;优选地,R103位置处的突变包括突变为A、G、S、T、N或Q;E104位置处的突变包括突变为K、R、A、G、S、T、N或Q;E112位置处的突变包括突变为K、R、A、G、S、T、N或Q;R113位置处的突变包括突变为A、G、S、T、N或Q;K114位置处的突变包括突变为A、G、S、T、N或Q;R117位置处的突变包括突变为A、G、S、T、N或Q;R119位置处的突变包括突变为A、G、S、T、N或Q;D120位置处的突变包括突变为K、R、A、G、S、T、N或Q;K122位置处的突变包括突变为A、G、S、T、N或Q;D124位置处的突变包括突变为K、R、A、G、S、T、N或Q;K154位置处的突变包括突变为A、G、S、T、N或Q;R165位置处的突变包括突变为A、G、S、T、N或Q;D199位置处的突变包括突变为A、G、S、T、N或Q;R207位置处的突变包括突变为A、G、S、T、N、Q、D或E;K209位置处的突变包括突变为A、G、S、T、N、Q、D或E;K210位置处的突变包括突变为A、G、S、T、N、Q、D或E;E213位置处的突变包括突变为A、G、S、T、N或Q;E215位置处的突变包括突变为A、G、S、T、N或Q(具体可参见PCT/CN2022/143298)。
上述的附属蛋白单体AP34及其突变体,是与孔蛋白单体BCP34来源于同一个种属的蛋白,该蛋白的序列是从马里亚纳海沟的宏基因组数据库挖掘而来,它能够在与BCP34的相互作用下形成中心对称的九聚体,其N端(比如,N端的一部分)或全部插入孔蛋白单体BCP34聚合而成的纳米孔腔内。附属蛋白单体的野生序列全长序列SEQ ID NO:3如下:
MSKVIFGFFAVLLMCFVVTASASSLVYTPKNPSFGGPAAYGTYLLNNANAQNQFKDPDAIRKDKTPIEEFNDRLQRSLLSRITSTISRSIIGIDGAVNPGSFETTDFLIDVTDLGGGQMSITTTDKVTGDQTSIVIETGL,其中前22个氨基酸(下划线部分)是该附属蛋白单体AP34的信号肽区,成熟的附属蛋白单体AP34的蛋白片段无信号肽区,序列如SEQ ID NO:4所示:
上述的附属蛋白单体AP34及其突变体,可以进行截短改造,用于改变孔复合物的嵌孔能力或/和与膜的适配性。截短的改造包括但不限于:仅保留N端23-67位氨基酸(即缺少N端
1-22位的信号肽)、仅保留N端23-57位氨基酸、仅保留N端23-52位氨基酸、仅保留N端23-54位氨基酸、仅保留N端23-51位氨基酸、仅保留N端23-50位氨基酸、仅保留N端23-49位氨基酸、仅保留N端23-48位氨基酸和仅保留N端23-47位氨基酸。优选地,突变体AP34片段的长度可以为24~45个氨基酸,具体地可以是24、25、26、27、28、29、30、31、32、33、34、35、36、37、38、39、40、41、42、43、44或45个氨基酸。更优选地,其氨基酸序列来自SEQ ID NO:3的残基第23~48位、第23~49位、第23~50位、第23~51位、第23~54位或者第23~57位对应残基或其任一种变体。
上述纳米孔蛋白复合物中,孔蛋白单体BCP34的收缩区组成氨基酸是S71、N74、G75和F76,优选地,此收缩区氨基酸残基的突变方向是:S71位置处的突变包括突变为G、A或T;N74位置处的突变包括突变为G、A、S或T;G75位置处的突变包括突变为A、S、T或Q;F76位置处的突变包括突变为A、S、T、N或Q。附属蛋白单体AP34的收缩区组成最主要的氨基酸残基为A39,另外T42和N46可能也会对测序信号有影响。此收缩区的突变方向包括但不限于SEQ ID NO:3的A39、T42和N46处的一个或多个位置处具有氨基酸的突变。优选地,A39位置处的突变包括突变为S、T、N、G、V、L、I或Q;T42位置处的突变包括突变为A、G或S;N46位置处的突变包括突变为A、G或S。将BCP34收缩区及AP34的收缩区的相关氨基酸进行上述突变,有助于增加电流信号的信噪比,从而提高碱基识别的分辨率,最终提高对待测物序列的判定的准确性。
上述附属蛋白单体的AP34的全部或部分(比如,N端的一部分)可以定位在孔蛋白单体BCP34聚合而成的纳米孔腔内。附属蛋白单体的AP34聚合而成的九聚体的中心腔或孔隙与纳米孔蛋白单体BCP34聚合而成的纳米孔腔对齐以形成连续通道,使得穿过连续通道易位的分析物的每个相互作用的化学基团首先与纳米孔蛋白的收缩区相互作用,然后与附属蛋白的收缩区相互作用。
上述纳米孔复合物的通道中的两个收缩区域可以具有相同或不同的最小孔直径(即第一传感区的最小孔直径和第二传感区的最小孔直径两者属于平行关系,可以第一传感区的最小孔直径大于第二传感区的最小孔直径,也可以第一传感区的最小孔直径小于第二传感区的最小孔直径)。孔直径的大小包括但不限制在1-3nm范围内。与单一收缩区产生的电流信号相比,双收缩区能够在观察到的电流之间提供更具区分性的电流信号,从而对均聚物形式的待测分析物的分辨率更高。
上述的纳米孔蛋白复合物中,附属蛋白通过共价或非共价作用附接于纳米孔蛋白的纳米孔腔内;更优选地,纳米孔蛋白与附属蛋白以九聚体的形式存在。相较于非共价作用的附接方式,通过共价作用将附属蛋白附接于纳米孔蛋白的纳米腔中的方式相对更稳定,有助于形成稳定的至少具有两个收缩区的连续通道。
为进一步提高附属蛋白在纳米孔蛋白上附接的稳定性,本申请中优选对两者发生相互作用区域的部分氨基酸残基进行突变或修饰。上述附属蛋白单体AP34与孔蛋白单体BCP34相互作用结合区域通常包括AP34的残基23到30和/或51到57,并且这些残基可以包含一种或多种修饰。在孔中形成收缩区域通常包括AP34的残基30到48,并且这些残基可以包含一种或多种修饰。AP34的残基8到29形成α-螺旋。AP34收缩区域主要在SEQ ID NO:3的如下残基处与BCP34内腔稳定接触:23、24、25、26、31、33、34、40、43和45。上述附属蛋白单体AP34通过一个或多个共价键和/或一种或多种非共价的相互作用附接在孔蛋白单体
BCP34上并与之形成相对稳定的复合物。优选地,合适的非共价相互作用包含盐桥、静电相互作用和π-π相互作用。共价作用指可以通过将分子附接到一个或多个半胱氨酸(半胱氨酸连接)、将分子附接到一个或多个赖氨酸、将分子附接到一个或多个非天然氨基酸、表位的酶修饰或末端的修饰对突变体或经修饰的单体进行化学修饰。用于进行此类修饰的合适方法在本领域是众所周知的。修饰方式包括但不限于在这些位置中的任何一个或多个位置处引入半胱氨酸、带电荷的氨基酸、非天然反应性氨基酸或光反应性氨基酸。
上述的这些位置在AP34上包括但不限于S23、S24、L25、T28、K30、N31、S33、F34、N47、A50、Q51、N52和Q53。这些位置在BCP34上包括但不限于N145、T148、K154、L156、L160、S161、R165、S195、Q197、D199、F203、Y205、K209、K210、L211、E213、E215、G217、S219、N221。
纳米孔蛋白和附属蛋白两者之间的相互作用的位点组合包括但不限于以下的一个或多个:BCP34上的R165和AP34上的S23、BCP34上的S195和AP34上的S24、BCP34上的S219和AP34上的L25、BCP34上的N221和AP34上的L25、BCP34上的N145C和AP34上的V26C、BCP34上的D199和AP34上的T28、BCP34上的D199和AP34上的K30、BCP34上的E215和AP34上的K30、BCP34上的E213和AP34上的N31、BCP34上的K154和AP34上的S33、BCP34上的E213和AP34上的S33、BCP34上的E213和AP34上的N47、BCP34上的L211和AP34上的A50、BCP34上的L211和AP34上的Q51、BCP34上的K209和AP34上的N52、BCP34上的K210和AP34上的Q53。具体的氨基酸组合对如表1所示。
表1
为了进一步增强附属蛋白附接在纳米孔蛋白上的稳定性,可以对上述位点组合中的任意一组或多组的位点进行突变(比如,均突变为半胱氨酸,进而在任意一组内的组合位点之间形成二硫键)或修饰。具体地,经修饰的孔蛋白单体BCP34或附属蛋白单体AP34的突变体可以直接通过氨基酸的侧链基团自发反应或氧化反应,也可以通过附接化学基团进行化学修饰。化学基团,比如酰胺基团,具体包括但不限于碘乙酰胺或马来酰亚胺。
根据本申请的第二个方面,提供了一种纳米孔传感器,该纳米孔传感器包括膜和上述纳米孔蛋白复合物,其中,纳米孔蛋白复合物中的纳米孔蛋白定位在膜上,纳米孔蛋白的纳米孔腔和附属蛋白共同形成跨膜的连续通道。
上述纳米孔传感器上的膜包括两亲分子层;优选地,膜是由二脂酰磷脂酰胆碱组成的磷脂双分子层、两嵌段共聚物或三嵌段共聚物。
含有上述两个收缩区的纳米孔蛋白复合物形成的纳米孔传感器,在应用于对均聚物核酸分子进行测序时,能够提高对此类分子的区分度和测序准确度。
根据本申请的第三个方面,提供了一种纳米孔测序系统,该纳米孔测序系统包括上述纳米孔传感器、导电溶液、跨膜提供电压电势的正负电极,以及用于测量通过连续通道的电信号的测量设备,其中,纳米孔传感器位于导电溶液中且将导电溶液分割为第一腔室和第二腔室。
在一些优选的实施例中,该纳米孔测序系统还包括待测分子,其中,待测分子选自含有核酸分子,更优选为含有均聚物区域的核酸分子;优选地,待测分子可瞬时性地位于连续通道内,且待测分子的一端位于第一腔室,另一端位于第二腔室。
根据本申请的第四个方面,提供了一种纳米孔测序的方法,该方法包括:使上述纳米孔测序系统与待测核酸分子接触;跨膜施加电势,使得待测核酸分子进入连续通道;以及一次或多次测量待测核酸分子相对于连续通道移动时产生的电信号,由此获得待测核酸分子的序列。
在本申请中,测量待测核酸分子通过第一收缩区和第二收缩区时产生的具体的电信号的类型有多种,包括但不仅限于电流阻滞信号强度、电流阻滞持续时间或电流阻滞事件发生的间隔时间等。
上述待测核酸分子为寡核苷酸分子,该寡核苷酸分子中的核苷酸与连续通道内的第一收缩区域和第二收缩区域相互作用,并且其中第一收缩区域和第二收缩区域中的每个收缩区域能够区分不同的核苷酸,使得通过连续通道的总电流受到第一收缩区域和第二收缩区中的每个收缩区域与定位在相应区域中的核苷酸之间的相互作用的影响。
具体地,上述寡核苷酸分子移动通过连续通道并实现跨膜易位。在一些优选的实施例中,使用核酸结合蛋白来控制寡核苷酸分子相对于连续通道孔的移动。
在一些优选的实施例中,上述寡核苷酸分子为含有均聚物的寡核苷酸分子,上述方法可以实现对含有均聚物的寡核苷酸分子的核苷酸序列进行测定。
根据本申请的第五个方面,还提供了一种纳米孔复合物的制备方法,该制备方法包括:分别构建纳米孔蛋白单体的表达载体和附属蛋白的表达载体;通过在感受态细胞中共表达孔蛋白单体的表达载体和附属蛋白单体的表达载体,并经过分离纯化来获得纳米孔复合物;或通过体外重组的方式构建纳米孔复合物。
在一些优选的实施例中,通过在感受态细胞中共表达蛋白单体的表达载体和附属蛋白单体的表达载体,并经过分离纯化来获得纳米孔复合物包括:在感受态细胞中共表达孔蛋白单体的表达载体和附属蛋白单体的表达载体,并培养感受态细胞达到预定OD值,收集共表达细
胞,其中共表达细胞内共表达孔蛋白单体和附属蛋白单体,且孔蛋白单体和附属蛋白单体自发聚合形成纳米孔复合物;对共表达细胞依次进行细胞破碎、细胞膜溶解及纯化,获得纳米孔复合物;优选地,预定OD值为OD600值达0.6-0.8;优选地,进行纯化包括依次进行strep标签纯化和his标签纯化;优选地,分离纯化获得的纳米孔复合物中孔蛋白单体和附属蛋白单体的摩尔比为1:1。
在一些优选的实施例中,通过体外重组的方式构建纳米孔复合物包括:在感受态细胞中分别表达孔蛋白单体的表达载体和附属蛋白单体的表达载体,并对各自表达后的细胞分别进行分离纯化,获得孔蛋白单体和附属蛋白单体;将孔蛋白单体和附属蛋白单体按1:30~50的摩尔比混合孵育,获得纳米孔复合物;优选地,将孔蛋白单体和附属蛋白单体置于缓冲溶液中进行混合孵育;优选地,缓冲溶液包括20mM Tris-HCl、150mM NaCl、0.05% Tween20及15%甘油,pH 8.0;优选地,在4℃下进行混合孵育。
为了通过体外表达获得上述纳米孔蛋白复合物,可以在孔蛋白单体BCP34和/或附属蛋白单体AP34的C端处增加蛋白纯化标签。蛋白纯化标签是本领域众所周知的、包括但不限于组氨酸标签(Poly His)、Strep tag II(非生物素化亲和素)、Biotin Avitag(可生物素化的短肽)、钙调蛋白结合肽(CBP)或精氨酸标签(Poly Arg)。标签插入的位置包括但不限于在SEQ ID NO:1及其突变体和SEQ ID NO:3及其突变体的残基48、49、50、51、54、57、62或67位置处。
AP34的蛋白纯化标签可以插入一段蛋白酶切短肽序列,有助于移除可能会干扰蛋白复合物形成或是蛋白复合物插孔的多余残基。这些蛋白酶切序列也是本领域众所周知的,包括但不限于凝血酶(识别序列为SEQ ID NO:30:LVPRG↓S)、Factor Xa(识别序列为SEQ ID NO:31:IE/DG↓R)、TEV蛋白酶(识别序列为SEQ ID NO:32:ENLYFQ↓G)、HRV 3C蛋白酶(识别序列为:SEQ ID NO:33:LEVLFQ↓GP),其中箭头表示蛋白酶作用的位点。
上述纳米孔复合物还可以通过共表达孔蛋白单体BCP34及其突变体和附属蛋白单体AP34及其突变体得到。共表达的方法指在合适的宿主细胞中同时表达孔蛋白单体和附属蛋白单体并允许纳米孔复合物在体内形成。优选地,宿主细胞包括但不限于BL21(DE3)、BL21Star(DE3)pLyss、Rossata(DE3)、Lemo21(DE3)等。具体地,一个载体中编码孔蛋白单体的至少一个基因和编码附属蛋白单体的基因、或第二载体中的编码至少一个附属蛋白单体可以一起转化、以表达蛋白质并在经转化的细胞中制备纳米孔复合物。这两个载体可以在单个启动子的控制下或在两个独立启动子的控制下,将编码孔蛋白单体和附属蛋白单体的两个基因放置在一个载体中,其中两个独立启动子可以相同或不同,该过程可以在宿主细胞体内进行或者在无细胞表达体系中进行。优先地,表达载体包括但不限于以T7为启动子的载体、如PET.28a(+)、PET.21a(+)、PET.32a(+)等。
当纳米孔复合物在体内构建好后,可以通过BCP34和AP34蛋白相应的蛋白纯化标签进行纯化得到。纯化的步骤是本领域已知的方法(例如离子交换、凝胶过滤、疏水相互作用柱色谱法等),可以单独使用,也可以以不同的组合使用以纯化纳米孔复合物的组分。需要说明的是,BCP34和AP34的蛋白纯化标签可以是相同也可以是不同,采用两种不同的标签能够在一定程度上提高纳米孔复合物的纯度。
此外,上述纳米孔复合物还可以采用体外重组的方式构建得到。其中孔蛋白单体BP34及其突变体可以在合适的载体中编码并转化到合适的宿主细胞中表达。然后通过相应的蛋白纯化标签纯化得到,纯化的步骤是本领域已知的方法(例如离子交换、凝胶过滤、疏水相互作用
柱色谱法等),可以单独使用或以不同组合使用以纯化孔复合物的组分。附属蛋白单体AP34及其突变体可以采用与孔蛋白单体BP34相同的宿主细胞中表达并在体外纯化的方式得到,也可以直接通过多肽固相合成的方式得到。
若采用体内表达的方式得到,则蛋白序列取SEQ ID NO:3或其突变体和截短体,在体内N端的信号肽可以引导其分泌在膜蛋白上,在通过纯化标签纯化得到后,对添加蛋白酶裂解位点的AP34,可以采用与添加序列对应的蛋白酶切除与成孔无关的C端蛋白序列。
若采用多肽合成的方式得到,则多肽序列为直接不含C端信号肽的氨基酸序列。多肽合成的方法包括但不限于固相合成、液相分段合成法、施陶丁格连接、天然化学连接、光敏感辅助基连接、可去除辅助基连接、化学区域选择连接、氨基酸的羧内酸酐(NCA)法、组合化学法、酶解法、基因工程法和发酵法。优选地、采用FMOC或BOC固相多肽合成法。
当BCP34和AP34都得到后,可采用体外孵育的方式形成纳米孔复合物。孵育的BCP34和AP34的比例包括但不限于500:1到1:1范围内。孵育的温度包括但不限于4℃、16℃、20℃、25℃和37℃。孵育的时长包括但不限于30min、1h、2h、3h、4h、5h、8h、16h和24h。孵育的反应溶液盐浓度包括但不限于50mM、100mM、150mM、200mM、250mM、300mM、400mM、500mM。孵育的反应溶液的pH值包括但不限于7.0、7.5、8.0、8.5。反应的同时可以加入促进纳米孔复合物稳定性的氧化剂或化学交联剂。孵育方式可以同时将孔蛋白单体BCP34以及附属蛋白单体AP34在反应溶液中混在一起,也可以将孔蛋白单体插入到膜中,然后添加附属蛋白单体,使得纳米孔复合物可以原位形成。
本发明上述方法构建的具有双收缩区的纳米孔蛋白复合物,可用于表征不同的分析物,包括但不限于多核苷酸、多肽和多糖的各种生物或合成大分子和聚合物。优选地,用于表征靶,多核苷酸包括DNA和/或RNA及其修饰物。
本发明提供了一种用于确定靶分析物存在或不存在一个或多个特征的方法,该方法包括:
a.使靶分析物与含有上述纳米孔蛋白复合物BCP34-AP34或其突变体的纳米孔传感器上的膜上的具有第一收缩区和第二收缩区的纳米孔接触,使得该靶分析物相对于纳米孔移动;
b.在靶分析物相对于纳米孔移动时获取一个或多个测量值,从而确定该靶分析物存在或不存在一个或多个特征。
进一步地,上述靶分析物与纳米孔蛋白复合物相互作用从而使得上述靶分析物相对于连续通道上的纳米孔移动。
需要说明的是,本申请提供的纳米孔蛋白复合物可用以开发探索更多高准确度、高集成度以及高稳定性的单分子测序仪。另外,作为一种生物传感器,该蛋白复合物有极高潜力用以鉴别和分析不同有机和无机物的组成和修饰信息,以及药物动力学或药物筛选中的应用,并在核酸药物递送和生物感受上也有广阔的应用前景。还可以结合基因组学、蛋白组学、代谢组学等,搭建一套满足全组学分析需求的通用型测量平台,帮助我们更深刻去理解生命规律与疾病发生机制。
下面将结合具体的实施例来进一步说明本申请的有益效果。
实施例一:野生型BCP34-AP34的AlphaFold2-Multimer的预测结构
我们使用AlphaFold-Multimer-v3进行目标蛋白质序列的多聚体结构预测,AlphaFold-Multimer-v3提供的不同模型参数进行了一系列预测,并选取了多聚体置信度最高的多聚体结构作为最终的预测结果。预测结果如图1和图2所示。其中图1中A为BCP34-AP34多聚体预测结构的侧视图(sideview),图1中B为BCP34-AP34多聚体预测结构的俯视图(topview),图2中A为成熟的BCP34-AP34(无信号肽)多聚体预测结构的侧视图(sideview),图2中B为成熟的BCP34-AP34(无信号肽)多聚体预测结构的俯视图(topview)。结构示出BCP34:AP34是以9:9的化学计量比通过两者之间的非共价作用形成的九聚体,具有C9对称性。
复合物具有两个收缩区,对于电流信号的产生起到决定性作用:第一个收缩区是孔蛋白单体BCP34聚合形成,第二个收缩区是由附属蛋白单体AP34聚合形成,位于第一个收缩区的下方,两个收缩区之间具有一定的距离,如图3所示。图4展示复合物预测结构的各自收缩区的重要氨基酸的侧链结构,第一个收缩区显示出氨基酸侧链的四个氨基酸分别为BCP34上的S71、N74、G75、F76,第二个收缩区显示出氨基酸侧链氨基酸为AP34上的A39,它是AP34内部最窄处,另外还有两处T42和N46也较窄,对电流信号有一定的影响。
实施例二:野生型BCP34单体及突变体和AP34单体及突变体表达载体的构建
通过In-fusion方法,利用NdeI和XhoI酶切后,将孔蛋白单体BCP34的DNA序列(SEQ ID NO:5插入到载体pET24a的多克隆区。在野生型BCP34的氨基酸序列(SEQ ID NO:1)的C端添加StrepII氨基酸作为纯化标签,其中筛选标签为卡那霉素,将构建好的载体命名为pET24a-BCP34。通过定点突变的方法,使用Agilent定点突变试剂盒,以孔蛋白单体BCP34的表达载体为模板,构建相应的其它突变体,如BCP34-1(SEQ ID NO:6)。
通过In-fusion方法,利用NdeI和XhoI酶切后,将附属蛋白单体AP34-1的DNA序列(SEQ ID NO:7)插入到载体pET21a的多克隆区,其中AP34-1是在野生型AP34氨基酸序列的第57位处加入了TEV酶切识别位点,便于蛋白体外的表达和纯化。在AP34-1的氨基酸序列(SEQ ID NO:8)的C端添加6个组氨酸作为纯化标签,其中筛选标签为氨苄青霉素,将构建好的载体命名为pET21a-AP34-1。通过定点突变的方法,使用Agilent定点突变试剂盒,以附属蛋白单体的AP34-1表达载体为模板,构建相应的其它突变体。
BCP34的氨基酸序列(SEQ ID NO:1):
BCP34的DNA序列(SEQ ID NO:5)
BCP34-1突变体的氨基酸序列(N74S、G75N和F76A SEQ ID NO:6):
AP34的氨基酸序列(SEQ ID NO:3):
AP34-1的DNA序列(SEQ ID NO:7):
AP34-1的氨基酸序列(AP34基础上57位处增加TEV酶切识别位点,SEQ ID NO:8):
实施例三:共表达法构建多收缩区复合孔
为了产生纳米孔复合物,可以在合适的革兰氏阴性宿主(如大肠杆菌)中共表达两种蛋白单体,并从外膜中提取并纯化为复合物。
在该实施例中,我们按照实施例二的方法构建了质粒pET24a-BCP34-1和pET21a-AP34-1,并将它们通过热激法共同转入大肠杆菌表达菌株E.coli BL21(DE3)中,然后将菌液均匀涂抹在含50μg/mL卡那霉素和100μg/mL氨苄青霉素的平板上,37℃过夜培养。次日挑取单菌落于含50μg/mL卡那霉素和100μg/mL氨苄青霉素的5mL LB培养基中,37℃,200rpm,过夜培养。将上述所得菌液,按体积比1:100接种于含有50μg/mL卡那霉素和100μg/mL氨苄青霉素的50mL LB液体培养基中,37℃,200rpm,培养4h。将扩大培养的菌液,按体积比1:100接种于含有50μg/mL卡那霉素和100μg/mL氨苄青霉素的2L LB液体培养基中培养,37℃,200rpm。待OD600值达0.6-0.8左右,加入终浓度为0.5mM的IPTG,16℃,200rpm,培养约16-18h。将菌液于8000rpm离心收集,菌体冻存于-20℃待用。
为在体外获得该共表达的蛋白复合物,纯化提取步骤如下:
(1)缓冲溶液配制
缓冲溶液A:20mM Tris-HCl,150mM NaCl pH 8.0
缓冲溶液B:20mM Tris-HCl,150mM NaCl 1% DDM,pH 8.0
缓冲溶液C:20mM Tris-HCl,150mM NaCl,0.05% Tween20,15%gly pH 8.0
缓冲溶液D:20mM Tris-HCl,150mM NaCl,0.05% Tween20,15%gly,5mM脱硫生物素,pH 8.0(此缓冲溶液现配现用)
缓冲溶液E:20mM Tris-HCl,150mM NaCl,0.05% Tween20,15%gly,300mM咪唑,pH 8.0
(1)细胞破碎和细胞膜溶解
按1g菌体加10mL缓冲溶液A的比例充分重悬菌体,用高压均质机破碎细胞至菌体溶液澄清。然后采用低温超高速离心机(Beckman,OptimaTMXPN-90)高速离心,转速40000rpm,温度4℃,离心1h,去除上清溶液,并采用缓冲溶液B将沉淀悬浮起来,然后置于旋转仪上4℃旋转过夜。次日18000rpm 4℃离心1h,取上清,0.22μm滤膜过滤后于4℃待用。
(2)纯化步骤
用AKTA pure层析仪将Strep-Tactin beads(IBA Lifesciences)层析柱利用缓冲溶液C平衡5柱体积(CV)后,2mL/min上样。上样完成后,使用缓冲溶液C冲洗20CV,使用缓冲溶液D洗脱,收集目的蛋白。将纯化后获得的目的蛋白进行SDS-PAGE电泳,图5中A展示了野生型复合物经过strep柱纯化后的蛋白表征结果。变性的条件是将样品在60℃下加热15分钟。结果显示目的蛋白在未变性的情况下为聚体状态,变性后为单体状态,单体分别含有成熟的BCP34-1和AP34-1蛋白,证明复合物在膜上形成,并通过BCP34-1的strep标签纯化出来,此时还有过量的BCP34-1和部分杂蛋白。
接着,将洗脱蛋白溶液加载到在缓冲溶液C中平衡的5mL HisTrap柱上。用>10CVs 5%缓冲液E离子缓冲液A洗涤柱,并用60mL以上的5-100%梯度的缓冲液B洗脱,收集目的蛋白。将纯化后获得的目的蛋白进行SDS-PAGE电泳,图5中B展示了野生型复合物经过his柱纯化后的蛋白表征结果。结果显示目的蛋白在未变性的情况下为聚体状态,变性后为单体状态,单体分别含有成熟的BCP34-1和AP34-1蛋白,证明复合物在膜上形成,并通过AP34-1的his标签进一步纯化出来,此时蛋白复合物为1:1状态。
然后将蛋白洗脱溶液中加入适量的TEV酶用于去除AP34-1的C端冗余蛋白序列(SEQ ID NO:9SDAIRKDKTPIEEFNDRLQRSLLSRITSTISRSIIGIDGAVNPGSFETTDFLIDVTDLGGGQMSITTTDKVTGDQTSIVIETGL),保留关键序列AP34-2(SEQ ID NO:10MSKVIFGFFAVLLMCFVVTASASSLVYTPKNPSFGGPAAYGTYLLNNANAQNQFKDPENLYFQ),将加入酶的蛋白溶液置于旋转仪上4℃旋转过夜,并将获得的目的蛋白进行SDS-PAGE电泳,蛋白跑胶结果如图6展示。图6中A展示的是为变性前的复合物蛋白在加入TEV前后蛋白的变化,图6中B展示的是为变性后的复合物蛋白在加入TEV前后蛋白的变化,可以得出AP34-1的C端冗余序列被TEV酶成功移除,最终获取我们的目标蛋白,将得到的蛋白浓缩至1mL,过经缓冲溶液C平衡的Superdex 6increase 10/300GL(Cytiva),收集目的蛋白,随后储存于-80℃。
实施例四:体外重组法构建多收缩区复合孔
复合物还可以通过体外重组的方式构建,在本实施例中,我们将突变体BCP34-1采用革兰氏阴性宿主(如大肠杆菌)中表达,并提膜纯化,而AP34的关键片段采用多肽合成的方法构建,由于多肽合成在体外进行,因此不需要N端的信号肽以及C端可能会干扰复合孔插膜的序列,在这里我们采用金斯瑞公司(Genscript)通过固相合成的多肽,合成多肽序列AP34-P1(SEQ ID NO:11:SSLVYTPKNPSFGGPAAYGTYLLNNANAQNQFKDP)。
BCP34-1表达和纯化步骤如实施例三中一样,只需去除组氨酸第二步精纯步骤以及TEV酶切步骤,直接将Strep-Tactin beads(IBA Lifesciences)层析柱纯化的蛋白洗脱并稀释到合适的浓度备用。同时分别将从金斯瑞公司(Genscript)获得冻干的2mg的AP34-P1溶解在1mL实施例中的缓冲溶液C中,以获得2mg/mL样品。对样品进行涡流处理直到没有剩余的肽粉可见。由于BCP34-1含有少许杂蛋白,因此很难准确测量浓度。可以使用蛋白质带在SDS-PAGE上相对于已知标记的蛋白强度获得样品的粗略估计。将BCP34-1和AP34-P1以大约1:50的摩尔比混合,并在4℃下以700rpm孵育4h,并以13,000rpm离心2分钟。获得的目的蛋白进行SDS-PAGE电泳,如图7展示。图7中A显示在未变性前复合物的胶图,图7中B显示在变性后复合物的胶图,结果显示目的蛋白在未变性的情况下为聚体状态,变性后为单体状态,由于BCP34-1和AP34-P之间无共价作用连接,且BCP34-1聚体和BCP34-AP34-P1复合物大小差异不大,无法从该胶图直观显示,因此胶图只表明了蛋白的状态及稳定性。
实施例五:通过共价交联进一步稳定多收缩区纳米孔复合物
从实施例四可以看出,尽管复合物孔不稳定,但构成复合物孔的两个蛋白单体具有很强的非共价作用,能在蛋白未变性前组装成聚体状态,为进一步降低其在实际纳米孔测试应用中存在的潜在风险,因此在本实施例中,我们将通过改造蛋白,在相互作用界面上引入半胱氨酸,通过二硫键进一步提高复合物的稳定性。
根据实施例一的结构,我们进一步分析,如图8所示,可以看出BCP34和AP34的相互作用区有很多氨基酸相邻,并且侧链方向靠近,这些氨基酸的侧链被标出,有潜力的相互作用氨基酸对已经列在表1中。在本实施例中,我们主要呈现BCP34-2(SEQ ID NO:12,BCP34-1上R165C)和AP34-P2(SEQ ID NO:13,AP34上S23C)以及BCP34-3(SEQ ID NO:14,BCP34-1上N145C)和AP34-P3(SEQ ID NO:15,AP34上V26C)的结果,这四条序列如下:
BPC34-2(BCP34-1基础上增加R165C突变,SEQ ID NO:12):
AP34-P2(AP34基础上增加S23C突变,SEQ ID NO:13):
BCP34-3(BCP34基础上增加N145C突变,SEQ ID NO:14):
AP34-P3(AP34基础上增加V26C突变,SEQ ID NO:15):
其中BCP34-2和BCP34-3是采用实施例四中的大肠杆菌表达和提膜纯化步骤得到,获得的目的蛋白进行SDS-PAGE电泳,如图9显示,并利用蛋白质带在SDS-PAGE上相对于已知标记的蛋白强度获得样品的粗略估计两个蛋白的浓度。图9中A和B显示在变性前后BCP34-2和BCP34-3的变化胶图,结果显示目的蛋白BCP34-2在未变性的情况下为聚体状态,变性后为单体状态;BCP34-3蛋白在未变性的情况下为聚体状态,变性后大部分为单体状态。
AP34-P2和AP34-P1采用实施例四的固相合成方法由金斯瑞公司合成,并将干粉用溶液溶解。分别将BCP34-2与AP34-P2,BCP34-3与AP34-P3以大约1:50的摩尔比混合,并在4℃下以700rpm孵育4h,并以13,000rpm离心2分钟。获得的目的蛋白进行SDS-PAGE电泳表征,如图10所示。图10中A显示BCP34-2与AP34-P2的反应效果,在该反应条件下,约90%的BCP34-2与AP34-P2形成了稳定的复合物BCP34-2-AP34-P2。图10中B显示BCP34-3与AP34-P3的反应效果,在该反应条件下,只有约20%的BCP34-3与AP34-P3形成了稳定的复合物BCP34-3-AP34-P3。并且通过DTT验证表征,复合物是通过二硫键稳定,且该二硫键在还原剂DTT的存在下可以被破坏。
由于AP34上的S23与BCP34上的R165C的交联效率高,我们还尝试了采用AP34-P2的截短体AP34-P4,AP34-P5,AP34-P6,和AP34-P7(序列如下所示),并分别与蛋白BCP34-2反应,反应条件保持统一,即蛋白与多肽以大约1:50的摩尔比混合,并在4℃下以700rpm孵育4h,并以13000rpm离心2分钟。获得的目的蛋白进行SDS-PAGE电泳表征,如图11所示,表征了复合物BCP34-2-AP34-P4,BCP34-2-AP34-P5,BCP34-2-AP34-P6,BCP34-2-AP34-P7。图11中A表征了AP34-P2,AP34-P5,AP34-P6,AP34-P7与BCP34-2的反应效率,可以看出它们的反应效率都较高,其中以AP34-P7与BCP34-2的反应效率最高,接近完全反应,且形成的复合物可以在DTT的存在下被打破。图11中B表征了AP34-P4与BCP34-2形成复合物BCP34-2-AP34-P4的反应效率,多维度的验证了复合物的形成是不会破坏聚体形成,且复合物形成会被DTT打破。
AP34-P4(SEQ ID NO:16,截短保留29aa.):CSLVYTPKNPSFGGPAAYGTYLLNNANAQ。
AP34-P5(SEQ ID NO:17,截短保留28aa.):CSLVYTPKNPSFGGPAAYGTYLLNNANA。
AP34-P6(SEQ ID NO:18,截短保留27aa.):CSLVYTPKNPSFGGPAAYGTYLLNNAN。
AP34-P7(SEQ ID NO:19,截短保留26aa.):CSLVYTPKNPSFGGPAAYGTYLLNNA。
实施例六:利用多收缩区纳米孔复合物构建纳米孔生物传感器
单通道纳米孔电流测量基于数字化装置的放大器,这里采用膜片钳放大器采集电流信号。Ag/AgCl电极浸润在测序缓冲液(成分包括:0.47M KCl、25mM HEPES、1mM EDTA、5mM ATP、25mM MgCl2、pH7.6)中,电极分别位于电解槽cis和trans区域。测序文库和纳米孔等试剂加入到cis区域中。使用1xPBS缓冲液将纳米孔蛋白或纳米孔复合物稀释一定的倍数后(此蛋白一般用0.1mg/ml的蛋白浓度,用PBS稀释100倍或者10倍嵌孔),在外加电场力作
用下将单个纳米孔或纳米孔复合孔插入由二脂酰磷脂酰胆碱(DPhPC,1,2-diphytanoyl-sn-glycero-3-phosphocholine)组成的磷脂双分子层中,形成纳米孔生物传感器。施加外加电压,获得单个孔蛋白的电流振幅值。
实施例七:将实施例六中的纳米孔传感器用于DNA测序
a)纳米孔BCP34-1以及BCP34-AP34-P1的DNA测序
制备核酸序列:在PUC57质粒的多克隆位点插入人工合成的序列SEQ ID NO:20,利用SEQ ID NO:21、SEQ ID NO:22引物组合,通过PCR扩增制备3.5kb的待测序列SEQ ID NO:20。
制备测序文库:将待测序列SEQ ID NO:20构建为测序文库。
将两条部分区域互补的DNA链的正义链(SEQ ID NO:23-(iSP18)4-SEQ ID NO:29,其中iSP18为间隔臂)和反义链(SEQ ID NO:24)退火后形成接头,与待测双链目的片段PUC57(SEQ ID NO:20)利用T4DNA连接酶在室温下连接并纯化,制备测序文库。然后该测序文库与解旋酶BCH105(SEQ ID NO:25)在25℃孵育1h(摩尔浓度比1:8),形成含有BCH105马达蛋白的如图12A所示结构的测序文库。在测序时,该测序文库能够进一步与带有胆固醇的单链DNA(SEQ ID NO:26,胆固醇连接在DNA的5'端)互补配对结合,形成图12B所示结构。测序文库与测序缓冲液(测序缓冲液:0.47M KCl、25mM HEPES、1mM EDTA、5mM ATP、25mM MgCl2、pH7.6)混合并加入纳米孔生物传感器中;施加外加电压0.14V或0.18V后,观察到DNA被纳米孔捕获,产生特征的阻滞电流振幅值。并且随着DNA通过纳米孔移动,电流振幅值改变。不同的DNA序列产生不同的阻滞电流振幅值。带有胆固醇的单链DNA可以与磷脂双分子层进行结合,有助于纳米孔捕获测序文库,降低测序文库的上样量。
SEQ ID NO:20:
SEQ ID NO:21:gccatcagattgtgtttgttagt。
SEQ ID NO:22:gcttacggttcactactcacga。
SEQ ID NO:23:tttttttttttttttttttttttttttttttttttttttt。
SEQ ID NO:29:ggttgtttctgttggtgctgatattgct。
SEQ ID NO:24:gcaatatcagcaccaacagaaacaacctttgaggcgagcggtcaa。
SEQ ID NO:25:
SEQ ID NO:26:cholesterol-ttgaccgctcgcctc。
图13为在外加电压0.18V作用下,文库DNA穿过纳米孔蛋白BCP34-1时产生的电流变化及局部细节。图14为在外加电压0.18V作用下,文库DNA穿过双检测器纳米孔蛋白BCP34-1-AP34-P1时产生的电流变化及局部细节。可见在0.18V外加电压下,纳米孔蛋白BCP34的开孔电流为230-250pA,测序幅度约为40pA;双检测器纳米孔蛋白的开孔电流为130-150pA,测序幅度约为20pA,在DNA均聚物区读取的电流信号区别于突变体蛋白,呈现更多的台阶化信号(该现象在不同的双检测器纳米孔蛋白突变体也可见,它们区分度有所不同,比较图可见图22)。可见,在相同电压下,双检测器孔蛋白BCP34-1-AP34-P1相比于原单检测器蛋白BCP34-1开孔电流有所降低,表明在孔蛋白的检测器外形成了第二个检测器,
成功构建了双检测器纳米孔蛋白。且对均聚物区的区分度台阶化,说明纳米孔蛋白复合物提供的信息更多,因此推测该双检测器纳米孔蛋白具有提升纳米孔测序分辨率的潜力。
b)不同AP34主信号区纳米孔复合物对DNA的测序表征
按照实施例四的方法,我们构建了额外两个纳米孔复合物BCP34-1-AP34-P8和BCP34-1-AP34-P9。其中AP34-P8是基于AP34基础上对主收缩区改造A39S,AP34-P9是基于AP34基础上对主收缩区改造A39N,并在体外通过多肽合成法获取,序列如下。
AP34-P8(SEQ ID NO:27,基于AP34基础上对主收缩区改造A39S):
AP34-P9(SEQ ID NO:28,基于AP34基础上对主收缩区改造A39N)
按照本实施例的测试方法对其复合物进行了DNA测序。图15为在外加电压0.18V作用下,文库DNA穿过纳米孔蛋白BCP34-1-AP34-P8时产生的电流变化及局部细节。图16为在外加电压0.18V作用下,文库DNA穿过双检测器纳米孔蛋白BCP34-1-AP34-P9时产生的电流变化及局部细节。可以看出对收缩区主要氨基酸的改造,会造成其对特征序列的分辨能力发生变化。
c)不同交联位点位置改造的纳米孔复合物对DNA的测序表征
将实施例五中构建的不同位点交联改造的复合物BCP34-2-AP34-P2和BCP34-3-AP34-P3采用本实施例的测试方法对其复合物进行了DNA测序。
图17为在外加电压0.18V作用下,文库DNA穿过纳米孔蛋白BCP34-2-AP34-P2时产生的电流变化及局部细节。
图18为在外加电压0.18V作用下,文库DNA穿过双检测器纳米孔蛋白BCP34-3-AP34-P3时产生的电流变化及局部细节。
与BCP34-1-AP34-P1复合物对比,这两个复合物嵌孔更加容易,复合物更稳定,不会出现后期开孔电流又增加的现象,输出的高分辨测序电流信号也更多。双检测器纳米孔蛋白的开孔电流为130-150pA,测序幅度约为20pA,DNA均聚物识别电流信号异于单孔蛋白,如图22所示(其中A为单收缩区孔BCP34-1与双收缩区BCP34-1-AP34-1电流信号对比;B为不同AP34的主信号区的电流信号对比;C为不同交联位点的改造电流信号对比;D为不同长短的AP34多肽的电流信号对比)。
d)不同长短的截短AP34与BCP34-2形成的纳米孔复合物对DNA的测序表征
将实施例五种构建的不同长短的多肽与BCP34-2构建的纳米孔复合物BCP34-2-AP34-P4、BCP34-2-AP34-P5、BCP34-2-AP34-P6和BCP34-2-AP34-P7采用本实施例的测试方法对其复合物进行了DNA测序,除了BCP34-2-AP34-P6未获得特征性的双收缩区测序电流信号,其余均获取到了,可能是由于该多肽的折叠不好导致。
图19为在外加电压0.18V作用下,文库DNA穿过纳米孔复合物BCP34-2-AP34-P4时产生的电流变化及局部细节。
图20为在外加电压0.18V作用下,文库DNA穿过纳米孔复合物BCP34-2-AP34-P5时产生的电流变化及局部细节。
图21为在外加电压0.18V作用下,文库DNA穿过纳米孔复合物BCP34-2-AP34-P7时产生的电流变化及局部细节。
图22为本实施例7中的不同纳米孔对DNA均聚物区电流信号的特征;A为单收缩区孔BCP34-1与双收缩区BCP34-1-AP34-1电流信号对比;B为不同AP34的主信号区的电流信号对比;C为不同交联位点的改造电流信号对比;D为不同长短的AP34多肽的电流信号对比。
图13至图22中纵坐标均为电流(pA),横坐标均为时间(min),各图示为某一时间尺度内的电流信号(具体数值可以忽略,原因在于:1.本实施例涉及的所有对蛋白的突变或截短等改造都主要是为了改变测序信号及提高复合物的稳定性,但不改变测序速度,因此我们不关注测序速度,即也不用关注测序一条完整reads所用的时间(横坐标差);2.测序输出非完全同一时间生成,此处截取的测序信号reads是随机选取的,且具有当次测序输出信号的代表完整reads测序图,因此这里的测序时间(即横坐标)并无参考意义,所以可以忽略)。
从上述各图可以看出,不同AP34的长度对测序信号是有影响的,所获取的信号信噪比以及对均聚物序列的分辨率是不同的。
从以上的描述中,可以看出,本发明上述的实施例实现了如下技术效果:
a)当前市场上应用于纳米孔测序仪的孔蛋白绝大多数只有一个收缩区,对于均聚物的解析能力差,导致了纳米孔测序的准确率偏低。本发明通过开发具有两个或两个以上收缩区的新一代纳米孔蛋白,进一步提高多个核苷酸之间的电流差异,将有机会增强碱基识别能力,具有有效提高纳米孔测序准确率的潜力。
b)本发明开发具有测序能力的新型纳米孔复合物,用以突破纳米孔测序领域孔蛋白类型限制,用以开发探索更多高准确度、高集成度以及高稳定性的单分子测序仪。
c)本发明是基于新型孔蛋白单体BCP34及其突变体上,通过挖掘与其蛋白同一种属的附属蛋白AP3434,并对该附属蛋白单体AP3434进行截短和改造,与孔蛋白单体BCP34一起构建纳米孔复合物,创建有包括但不限于两个收缩区的纳米孔蛋白复合物,并将其应用在纳米孔测序领域。
以上所述仅为本发明的优选实施例而已,并不用于限制本发明,对于本领域的技术人员来说,本发明可以有各种更改和变化。凡在本发明的精神和原则之内,所作的任何修改、等同替换、改进等,均应包含在本发明的保护范围之内。
Claims (20)
- 一种纳米孔蛋白复合物,其特征在于,所述纳米孔蛋白复合物包括:纳米孔蛋白,所述纳米孔蛋白由多个孔蛋白单体聚合而成,且多个所述孔蛋白单体聚合形成中空的纳米孔腔;附属蛋白,所述附属蛋白由多个附属蛋白单体聚合而成,所述附属蛋白的至少部分位于所述纳米孔腔内,所述附属蛋白与所述纳米孔腔共同形成连续通道;其中,按照穿过所述连续通道的分析物的移动方向,所述连续通道包括顺次连通的第一传感区和第二传感区,所述第一传感区由所述纳米孔蛋白的至少部分形成,所述第二传感区由所述附属蛋白的至少部分形成;所述孔蛋白单体选自如下任意一种蛋白:1)具有SEQ ID NO:1所示氨基酸序列的蛋白;2)与SEQ ID NO:1所示氨基酸序列至少具有50%同一性且具有聚合形成纳米孔蛋白功能的蛋白;或3)在SEQ ID NO:1的基础上经过取代、缺失或添加一个或几个氨基酸,且具有聚合形成纳米孔蛋白功能的蛋白;所述附属蛋白单体选自如下任意一种蛋白:i)具有SEQ ID NO:3所示氨基酸序列的蛋白;ii)与SEQ ID NO:3所示氨基酸序列至少具有50%同一性且具有如下功能的蛋白:与所述孔蛋白单体结合,并随所述孔蛋白单体聚合形成所述纳米孔蛋白而一起聚合形成所述附属蛋白;或iii)在SEQ ID NO:3所示氨基酸序列上经过取代、缺失或添加一个或几个氨基酸,且具有如下功能的蛋白:与所述孔蛋白单体结合,并随所述孔蛋白单体聚合形成所述纳米孔蛋白而一起聚合形成所述附属蛋白。
- 根据权利要求1所述的纳米孔蛋白复合物,其特征在于,所述纳米孔蛋白和所述附属蛋白来源于同一个种属;任选地,所述附属蛋白全部或所述附属蛋白的至少N端的一部分位于所述纳米孔蛋白的所述纳米孔腔内;优选地,所述附属蛋白通过共价或非共价作用附接于所述纳米孔蛋白的所述纳米孔腔内;更优选地,所述纳米孔蛋白与所述附属蛋白以九聚体的形式存在。
- 根据权利要求1所述的纳米孔蛋白复合物,其特征在于,所述附属蛋白单体为SEQ ID NO:3所示的氨基酸序列截短的突变体,所述截短的突变体选自如下任意一种突变体:仅保留N端23-67位氨基酸、仅保留N端23-57位氨基酸、仅保留N端23-52位氨基酸、仅保留N端23-54位氨基酸、仅保留N端23-51位氨基酸、仅保留N端23-50位氨基酸、仅保留N端23-49位氨基酸、仅保留N端23-48位氨基酸或仅保留N端23-47位氨基酸。
- 根据权利要求1所述的纳米孔蛋白复合物,其特征在于,所述附属蛋白单体的长度为24~45个氨基酸;优选地,所述附属蛋白单体的氨基酸序列来自SEQ ID NO:3的如下残基位置区间或其突变体:第23~48位、第23~49位、第23~50位、第23~51位、第23~54位或第23~57位。
- 根据权利要求1所述的纳米孔蛋白复合物,其特征在于,所述孔蛋白单体选自在SEQ ID NO:1的如下任意一组或多组中的至少一个氨基酸位点发生突变的蛋白:1)S71、N74、G75和F76;优选地,S71突变为G、A或T;N74突变为G、A、S或T;G75突变为A、S、T或Q;F76突变为A、S、T、N或Q;2)E162、R196、S200和S216;优选地,E162突变为A、G、V、L、I、Y、F或W;R196突变为A、G、V、L、I、Y、F或W;S200突变为A、G、V、L、I、Y、F或W;S216突变为A、G、V、L、I、Y、F或W;3)R103、E104、E112、R113、K114、R117、R119、D120、K122、D124、K154、R165、D199、R207、K209、K210、E213或E215;优选地,R103突变为A、G、S、T、N或Q;E104突变为K、R、A、G、S、T、N或Q;E112突变为K、R、A、G、S、T、N或Q;R113突变为A、G、S、T、N或Q;K114突变为A、G、S、T、N或Q;R117突变为A、G、S、T、N或Q;R119突变为A、G、S、T、N或Q;D120突变为K、R、A、G、S、T、N或Q;K122突变为A、G、S、T、N或Q;D124突变为K、R、A、G、S、T、N或Q;K154突变为A、G、S、T、N或Q;R165突变为A、G、S、T、N或Q;D199突变为A、G、S、T、N或Q;R207突变为A、G、S、T、N、Q、D或E;K209突变为A、G、S、T、N、Q、D或E;K210突变为A、G、S、T、N、Q、D或E;E213突变为A、G、S、T、N或Q;E215突变为A、G、S、T、N或Q。
- 根据权利要求1所述的纳米孔蛋白复合物,其特征在于,所述附属蛋白单体选自在SEQ ID NO:3所示氨基酸序列的如下至少一个氨基酸发生取代:A39、T42或N46;优选地,A39突变为S、T、N、G、V、L、I或Q;T42突变为A、G或S;N46突变为A、 G或S。
- 根据权利要求1所述的纳米孔蛋白复合物,其特征在于,所述第一传感区和所述第二传感区具有相同或不同的最小孔直径;优选地,所述最小孔直径为1-3nm。
- 根据权利要求1-7中任一项所述的纳米孔蛋白复合物,其特征在于,所述附属蛋白单体选自在SEQ ID NO:3所示氨基酸序列上的如下至少一个氨基酸突变为半胱氨酸或被非天然氨基酸取代的蛋白:S23、S24、L25、T28、K30、N31、S33、F34、N47、A50、Q51、N52或Q53;和/或所述孔蛋白单体选自在SEQ ID NO:1的基础上的如下至少一个氨基酸突变为半胱氨酸或被非天然氨基酸取代的蛋白:N145、T148、K154、L156、L160、S161、R165、S195、Q197、D199、F203、Y205、K209、K210、L211、E213、E215、G217、S219或N221。
- 根据权利要求8所述的纳米孔蛋白复合物,其特征在于,所述孔蛋白单体和所述附属蛋白单体存在如下至少一个位点组合形式的突变,且各所述位点组合中相应位点的氨基酸突变为半胱氨酸或被非天然氨基酸取代:1)SEQ ID NO:1上的R165和SEQ ID NO:3上的S23;2)SEQ ID NO:1上的S195和SEQ ID NO:3上的S24;3)SEQ ID NO:1上的S219和SEQ ID NO:3上的L25;4)SEQ ID NO:1上的N221和SEQ ID NO:3上的L25;5)SEQ ID NO:1上的N145C和SEQ ID NO:3上的V26C;6)SEQ ID NO:1上的D199和SEQ ID NO:3上的T28;7)SEQ ID NO:1上的D199和SEQ ID NO:3上的K30;8)SEQ ID NO:1上的E215和SEQ ID NO:3上的K30:9)SEQ ID NO:1上的E213和SEQ ID NO:3上的N31:10)SEQ ID NO:1上的K154和SEQ ID NO:3上的S33:11)SEQ ID NO:1上的E213和SEQ ID NO:3上的S33:12)SEQ ID NO:1上的E213和SEQ ID NO:3上的N47;13)SEQ ID NO:1上的L211和SEQ ID NO:3上的A50;14)SEQ ID NO:1上的L211和SEQ ID NO:3上的Q51;15)SEQ ID NO:1上的K209和SEQ ID NO:3上的N52;16)SEQ ID NO:1上的K210和SEQ ID NO:3上的Q53。
- 根据权利要求9所述的纳米孔蛋白复合物,其特征在于,任一所述位点组合中相应位点的所述氨基酸均突变为半胱氨酸。
- 根据权利要求8-10中任一项所述的纳米孔蛋白复合物,其特征在于,通过修饰的方式使所述氨基酸被非天然氨基酸取代,所述修饰为直接修饰或间接修饰;优选地,所述直接修饰包括通过所述氨基酸的侧链基团自发反应或氧化反应进行修饰;优选地,所述间接修饰包括通过附接化学小分子进行化学修饰;优选地,所述化学小分子包括含官能团的化学交联剂,且所述官能团对二硫苏糖醇有抗性。优选地,至少部分所述附属蛋白单体中突变为半胱氨酸的位点与至少部分所述孔蛋白单体中突变为半胱氨酸的位点之间形成二硫键。
- 一种纳米孔传感器,其特征在于,所述纳米孔传感器包括膜和权利要求1-11种任一项所述的纳米孔蛋白复合物,其中,所述纳米孔蛋白复合物中的所述纳米孔蛋白定位在所述膜上,所述纳米孔蛋白的所述纳米孔腔和所述附属蛋白共同形成跨所述膜的连续通道。
- 根据权利要求12所述的纳米孔传感器,其特征在于,所述膜包括两亲分子层;优选地,所述膜是由二脂酰磷脂酰胆碱组成的磷脂双分子层、两嵌段共聚物或三嵌段共聚物。
- 一种纳米孔测序系统,其特征在于,所述纳米孔测序系统包括权利要求12或13所述的纳米孔传感器,所述纳米孔测序系统还包括:导电溶液;跨所述膜提供电压电势的正负电极,以及用于测量通过所述连续通道的电信号的测量设备;其中,所述纳米孔传感器位于所述导电溶液中,且将所述导电溶液分割为第一腔室和第二腔室。
- 根据权利要求14所述的纳米孔测序系统,其特征在于,所述纳米孔测序系统还包括待测分子,所述待测分子选自多核苷酸、多肽和多糖,优选地,所述待测分子选自含有均聚物的多核苷酸;优选地,所述待测分子可瞬时性地位于所述连续通道内,且所述待测分子的一端位 于所述第一腔室,另一端位于所述第二腔室。
- 一种纳米孔测序的方法,其特征在于,所述方法包括:使权利要求14所述的纳米孔测序系统与权利要求15所述的待测分子接触;跨所述膜施加电势,使得所述待测分子进入所述连续通道;以及一次或多次测量所述待测分子相对于所述连续通道移动时产生的电信号,由此获得所述待测分子的序列。
- 根据权利要求16所述的方法,其特征在于,所述电信号包括电流阻滞信号强度、电流阻滞持续时间或电流阻滞事件发生的间隔时间。
- 根据权利要求16所述的方法,其特征在于,所述待测分子为多核苷酸,所述多核苷酸中的核苷酸与所述连续通道内的第一收缩区域和第二收缩区域相互作用,并且其中所述第一收缩区域和所述第二收缩区域中的每个收缩区域能够区分不同的核苷酸,使得通过所述连续通道的总电流受到所述第一收缩区域和所述第二收缩区中的每个收缩区域与定位在所述区域中的每个区域处的核苷酸之间的相互作用的影响。
- 根据权利要求18所述的方法,其特征在于,使用核酸结合蛋白来控制所述多核苷酸相对于所述连续通道孔的移动。
- 权利要求1-11中任一项所述的纳米孔蛋白复合物的制备方法,其特征在于,所述制备方法包括:分别构建孔蛋白单体的表达载体和附属蛋白单体的表达载体;通过在感受态细胞中共表达所述孔蛋白单体的表达载体和所述附属蛋白单体的表达载体,并经过分离纯化来获得所述纳米孔蛋白复合物;或通过体外重组的方式构建所述纳米孔蛋白复合物。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2024/099017 WO2025255780A1 (zh) | 2024-06-13 | 2024-06-13 | 纳米孔蛋白复合物、其构建方法及应用 |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2024/099017 WO2025255780A1 (zh) | 2024-06-13 | 2024-06-13 | 纳米孔蛋白复合物、其构建方法及应用 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025255780A1 true WO2025255780A1 (zh) | 2025-12-18 |
Family
ID=98049937
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2024/099017 Pending WO2025255780A1 (zh) | 2024-06-13 | 2024-06-13 | 纳米孔蛋白复合物、其构建方法及应用 |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025255780A1 (zh) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN107735686A (zh) * | 2015-04-14 | 2018-02-23 | 鲁汶天主教大学 | 具有内部蛋白质衔接子的纳米孔 |
| CN113195736A (zh) * | 2018-11-08 | 2021-07-30 | 牛津纳米孔科技公司 | 孔 |
| CN115974984A (zh) * | 2023-01-17 | 2023-04-18 | 南方科技大学 | 双门孔道蛋白、孔道蛋白突变体、核苷酸序列及其应用 |
| CN117106037A (zh) * | 2017-06-30 | 2023-11-24 | 弗拉芒区生物技术研究所 | 新颖蛋白孔 |
| WO2024138512A1 (zh) * | 2022-12-29 | 2024-07-04 | 深圳华大生命科学研究院 | 新型孔蛋白bcp34、其突变体及其应用 |
-
2024
- 2024-06-13 WO PCT/CN2024/099017 patent/WO2025255780A1/zh active Pending
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN107735686A (zh) * | 2015-04-14 | 2018-02-23 | 鲁汶天主教大学 | 具有内部蛋白质衔接子的纳米孔 |
| CN117106037A (zh) * | 2017-06-30 | 2023-11-24 | 弗拉芒区生物技术研究所 | 新颖蛋白孔 |
| CN113195736A (zh) * | 2018-11-08 | 2021-07-30 | 牛津纳米孔科技公司 | 孔 |
| US20220056517A1 (en) * | 2018-11-08 | 2022-02-24 | Oxford Nanopore Technologies Limited | Pore |
| WO2024138512A1 (zh) * | 2022-12-29 | 2024-07-04 | 深圳华大生命科学研究院 | 新型孔蛋白bcp34、其突变体及其应用 |
| CN115974984A (zh) * | 2023-01-17 | 2023-04-18 | 南方科技大学 | 双门孔道蛋白、孔道蛋白突变体、核苷酸序列及其应用 |
Non-Patent Citations (1)
| Title |
|---|
| ZHANG JIA-YUAN, ZHANG YUNING, WANG LELE, GUO FEI, YUN QUANXIN, ZENG TAO, YAN XU, YU LEI, CHENG LEI, WU WEI, SHI XIAO, CHEN JUNYI, : "A single-molecule nanopore sequencing platform", BIORXIV, CN, 20 August 2024 (2024-08-20), CN, pages 1 - 21, XP093376076, Retrieved from the Internet <URL:https://www.biorxiv.org/content/10.1101/2024.08.19.608720v1.full.pdf> DOI: 10.1101/2024.08.19.608720 * |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12275760B2 (en) | Amino acid-specific binder and selectively identifying an amino acid | |
| JP2024133465A (ja) | 細孔 | |
| JP7282697B2 (ja) | 新規タンパク質細孔 | |
| US20240288416A1 (en) | Artificial nanopores and uses and methods relating thereto | |
| JP7027334B2 (ja) | アルファ溶血素バリアントおよびその使用 | |
| CA3000561C (en) | Alpha-hemolysin variants | |
| JP2019525911A (ja) | 長寿命アルファ溶血素ナノポア | |
| CN109627344B (zh) | cAMP荧光探针及其应用 | |
| WO2024138512A1 (zh) | 新型孔蛋白bcp34、其突变体及其应用 | |
| CN111164096A (zh) | 一种Mmup单体变体及其应用 | |
| JP2025500472A (ja) | 細孔 | |
| WO2024138425A1 (zh) | 一种新型纳米孔蛋白及其应用 | |
| WO2024138424A1 (zh) | 纳米孔蛋白及其应用 | |
| WO2024138565A1 (zh) | 纳米孔蛋白及其突变体和应用 | |
| CN109748970B (zh) | α-酮戊二酸光学探针及其制备方法和应用 | |
| WO2025255780A1 (zh) | 纳米孔蛋白复合物、其构建方法及应用 | |
| WO2024138473A1 (zh) | 孔蛋白单体、孔蛋白及其突变体和其应用 | |
| WO2026055978A1 (zh) | 纳米孔蛋白复合物及其应用 | |
| WO2025236239A1 (zh) | 一种双收缩区纳米孔蛋白复合物及其应用 | |
| WO2025236237A1 (zh) | 一种双限制区纳米孔蛋白复合物及其应用 | |
| WO2025236238A1 (zh) | 一种纳米孔蛋白单体及其应用 | |
| HK40129071A (zh) | 一种双限制区纳米孔蛋白复合物及其应用 | |
| HK40129471A (zh) | 一种双收缩区纳米孔蛋白复合物及其应用 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24942953 Country of ref document: EP Kind code of ref document: A1 |