WO2024258814A1 - Compositions and methods for engineering male sterility in plants - Google Patents
Compositions and methods for engineering male sterility in plants Download PDFInfo
- Publication number
- WO2024258814A1 WO2024258814A1 PCT/US2024/033340 US2024033340W WO2024258814A1 WO 2024258814 A1 WO2024258814 A1 WO 2024258814A1 US 2024033340 W US2024033340 W US 2024033340W WO 2024258814 A1 WO2024258814 A1 WO 2024258814A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- plant
- nucleotides
- seq
- seed
- modification
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C07—ORGANIC CHEMISTRY
- C07K—PEPTIDES
- C07K14/00—Peptides having more than 20 amino acids; Gastrins; Somatostatins; Melanotropins; Derivatives thereof
- C07K14/415—Peptides having more than 20 amino acids; Gastrins; Somatostatins; Melanotropins; Derivatives thereof from plants
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
- C12N15/09—Recombinant DNA-technology
- C12N15/11—DNA or RNA fragments; Modified forms thereof; Non-coding nucleic acids having a biological activity
- C12N15/113—Non-coding nucleic acids modulating the expression of genes, e.g. antisense oligonucleotides; Antisense DNA or RNA; Triplex- forming oligonucleotides; Catalytic nucleic acids, e.g. ribozymes; Nucleic acids used in co-suppression or gene silencing
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
- C12N15/09—Recombinant DNA-technology
- C12N15/63—Introduction of foreign genetic material using vectors; Vectors; Use of hosts therefor; Regulation of expression
- C12N15/79—Vectors or expression systems specially adapted for eukaryotic hosts
- C12N15/82—Vectors or expression systems specially adapted for eukaryotic hosts for plant cells, e.g. plant artificial chromosomes (PACs)
- C12N15/8241—Phenotypically and genetically modified plants via recombinant DNA technology
- C12N15/8261—Phenotypically and genetically modified plants via recombinant DNA technology with agronomic (input) traits, e.g. crop yield
- C12N15/8287—Phenotypically and genetically modified plants via recombinant DNA technology with agronomic (input) traits, e.g. crop yield for fertility modification, e.g. apomixis
- C12N15/8289—Male sterility
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N9/00—Enzymes; Proenzymes; Compositions thereof; Processes for preparing, activating, inhibiting, separating or purifying enzymes
- C12N9/14—Hydrolases (3)
- C12N9/16—Hydrolases (3) acting on ester bonds (3.1)
- C12N9/22—Ribonucleases [RNase]; Deoxyribonucleases [DNase]
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N2310/00—Structure or type of the nucleic acid
- C12N2310/10—Type of nucleic acid
- C12N2310/20—Type of nucleic acid involving clustered regularly interspaced short palindromic repeats [CRISPR]
Definitions
- the present disclosure relates to the field of agricultural biotechnology, and more specifically to methods and compositions for genome editing in plants.
- Hybridization is an important aspect in the breeding of domesticated plants as it enables the introduction of transient hybrid vigor, desirable variation among different germplasms, transgenic trait integration, and generation of novel phenotypes.
- Plant breeders use hybridization or controlled cross-pollination as the starting point of a breeding cycle in different crops.
- Conventional methods for cross-pollination of many crop species, especially soybean typically involves manual pollination and emasculation. This process is complex, time consuming, and labor-intensive due to the anatomy and size of the flowers.
- Typical commercial breeding programs require thousands or even millions of crosses in workflows such as, development crosses, backcrosses, and trait integration.
- the present invention provides a modified plant, plant seed, plant part, or plant cell comprising a genomic modification that reduces expression or activity of TDF1, or a homolog thereof, as compared to the expression or activity of TDF1 or a homolog thereof in an otherwise identical plant, plant seed, plant part, or plant cell that lacks the modification.
- the modification may be present in at least one allele of an endogenous TDF1 gene or homolog thereof.
- the TDF1 gene or homolog thereof may encode a protein comprising a conserved SANT/Myb domain having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the conserved SANT/Myb domain corresponding to position 9 to position 112 of SEQ ID NO:5.
- the modification may disrupt the function of a wild-type allele product of the endogenous TDF1 gene or homolog thereof.
- the modification may be located within a coding region, non-coding region, or any combination thereof, of the endogenous TDF1 gene or homolog thereof.
- the modification may be located in an exon region of said TDF1 gene or homolog thereof. In yet further embodiments, the modification may be located in an intron region of said TDF1 gene or homolog thereof. In even further embodiments, the modification may be located in an exon region and an intron region of said TDF1 gene or homolog thereof.
- the plant, plant seed, plant part, or plant cell may be heterozygous for the modification. In other embodiments, the plant, plant seed, plant part, or plant cell may be homozygous for the modification.
- the plant, plant seed, plant part, or plant cell may comprise a first modification in a first allele of the TDF1 gene and a second modification in a second allele of the TDF1 gene, the first modification and the second modification being different from one another.
- the modification may comprise a deletion, an insertion, a substitution, an inversion, a duplication, or a combination of any thereof.
- the modification may be a deletion.
- the deletion may comprise between 1 nucleotide and 1500 nucleotides.
- the deletion may comprise a region of at least 1 nucleotide, at least 2 nucleotides, at least 5 nucleotides, at least 10 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, at least 500 nucleotides, at least 600 nucleotides, at least 700 nucleotides, at least 800 nucleotides, at least 900 nucleotides, at least 1000 nucleotides, or at least 1500 nucleotides.
- the modification may be an inversion.
- the inversion may be a region of at least 5 consecutive nucleotides, at least 10 consecutive nucleotides, at least 20 consecutive nucleotides, at least 30 consecutive nucleotides, at least 40 consecutive nucleotides, at least 50 consecutive nucleotides, at least 100 consecutive nucleotides, at least 200 consecutive nucleotides, at least 300 consecutive nucleotides, at least 400 consecutive nucleotides, at least 500 consecutive nucleotides, at least 600 consecutive nucleotides, at least 700 consecutive nucleotides, at least 800 consecutive nucleotides, at least 900 consecutive nucleotides, at least 1000 consecutive nucleotides, or at least 1100 consecutive nucleotides.
- the present invention provides a modified plant, plant seed, plant part, or plant cell comprising a genomic modification that reduces expression or activity of TDF1, or a homolog thereof, as compared to the expression or activity of TDF1 or a homolog thereof in an otherwise identical plant, plant seed, plant part, or plant cell that lacks the modification, wherein the modification is present in at least one allele of an endogenous TDF1 gene or homolog thereof.
- the plant may be a leguminous plant, or wherein the plant seed, plant part, or plant cell is a plant seed, plant part, or plant cell of a leguminous plant.
- the leguminous plant may be a soybean plant, a bean plant, a pea plant, a chickpea plant, an alfalfa plant, a peanut plant, a carob plant, a lentil plant, or a licorice plant.
- the leguminous plant may be a soybean plant.
- the leguminous plant may be a soybean plant and the TDF1 gene is GmTDFl -1.
- the plant, plant seed, plant part, or plant cell may comprise a modification in at least one allele of the GmTDFl-1 gene, wherein the modification is selected from the group consisting of: an 8 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:26; a 10 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:27; a 1 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:28; a 6 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO: 29; a 9 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:30; an 8 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:31; an 11 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:32; a 7 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:33
- the plant, plant seed, plant part, or plant cell may comprise at least one polynucleotide sequence selected from the group consisting of SEQ ID NOs:26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, and 53.
- the plant, plant seed, plant part, or plant cell may comprise at least one polynucleotide sequence selected from the group consisting of SEQ ID NOs:26, 27, 28, 29, 30, 31, 32, and 33.
- the plant, plant seed, plant part, or plant cell may comprise at least one polynucleotide sequence selected from the group consisting of SEQ ID NOs:34, 35, 36,
- the modification may result in a male sterility phenotype in the plant, as compared to the male fertile phenotype of an otherwise identical plant that lacks the modification.
- the present invention provides a polynucleotide comprising a sequence selected from the group consisting of SEQ ID NOs:26, 27, 28, 29, 30, 31 , 32, 33, 34, 35, 36, 37,
- the present invention also provides a guide RNA comprising a polynucleotide sequence selected from the group consisting of SEQ ID NOs:23, 24, and 25.
- the polynucleotide sequence may be a spacer sequence.
- the present invention provides a method for producing a modified plant having a male sterility phenotype, the method comprising: a) introducing a modification into at least one target site in an endogenous TDF1 gene or a homolog thereof of a plant cell; b) identifying and selecting one or more plant cells of step a) comprising said modification in said TDF1 gene or homolog thereof; and c) regenerating at least one plant from at least one or more cells selected in step b).
- the target site may be located in a coding region, a non-coding region, or any combination thereof, of an endogenous TDF1 gene or homolog thereof.
- the modification may be introduced by a site-specific genome modification enzyme selected from the group consisting of: an RNA-guided nuclease, a zinc-finger nuclease, a meganuclease, a TALE-nuclease, a recombinase, a transposase, and combinations of any thereof.
- the site-specific genome modification enzyme may be an RNA-guided nuclease comprising a Cas nuclease, a Cpf 1 nuclease, or a variant of either thereof.
- the site-specific genome modification enzyme may be an RNA-guided nuclease comprising a Cpf 1 nuclease. In some embodiments, the site-specific genome modification enzyme may create at least one strand break at the target site. In other embodiments, the modification may be selected from the group consisting of a substitution, an insertion, an inversion, a deletion, a duplication, and a combination thereof. In further embodiments, the modification may be a deletion.
- the deletion may comprise a region of at least 1 nucleotide, at least 2 nucleotides, at least 5 nucleotides, at least 10 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, at least 500 nucleotides, at least 600 nucleotides, at least 700 nucleotides, at least 800 nucleotides, at least 900 nucleotides, at least 1000 nucleotides, or at least 1500 nucleotides.
- the modification is an inversion.
- the inversion may be a region of at least 5 consecutive nucleotides, at least 10 consecutive nucleotides, at least 20 consecutive nucleotides, at least 30 consecutive nucleotides, at least 40 consecutive nucleotides, at least 50 consecutive nucleotides, at least 100 consecutive nucleotides, at least 200 consecutive nucleotides, at least 300 consecutive nucleotides, at least 400 consecutive nucleotides, at least 500 consecutive nucleotides, at least 600 consecutive nucleotides, at least 700 consecutive nucleotides, at least 800 consecutive nucleotides, at least 900 consecutive nucleotides, at least 1000 consecutive nucleotides, or at least 1100 consecutive nucleotides.
- the invention also provides male sterile soybean plants produced by the methods described herein.
- the present invention provides a method for producing an Fi hybrid soybean plant comprising crossing a male-sterile soybean plant comprising a modified GmTDFl- 1 gene with a second, non-isogenic, male-fertile soybean plant to produce an Fi hybrid soybean plant.
- the invention also provides Fi hybrid soybean plants produced by the methods described herein.
- the present invention provides a method for producing an Fi hybrid soybean seed, comprising crossing a modified plant, plant seed, plant part, or plant cell comprising a genomic modification that reduces expression or activity of TDF1, or a homolog thereof, as compared to the expression or activity of TDF1 or a homolog thereof in an otherwise identical plant, plant seed, plant pail, or plant cell that lacks the modification with a second, non-isogenic, male-fertile soybean plant and harvesting the resultant Fi hybrid soybean seed wherein an Fi hybrid soybean seed is produced.
- the invention also provides Fi hybrid soybean seeds produced by the methods described herein.
- the present invention provides a method for producing an Fi hybrid soybean seed comprising crossing a male-sterile soybean plant comprising a modified GmTDFl- 1 gene with a second, non-isogenic, male-fertile soybean plant to produce an Fi hybrid soybean seed.
- the invention also provides Fi hybrid soybean seed produced by the methods described herein.
- the present invention also provides modified plants, plant seed, plant parts, or plant cells, wherein the plants, plant seed, plant parts, or plant cells comprise at least one polynucleotide sequence selected from the group consisting of SEQ ID NOs:26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, and 53.
- the modified plant, plant seed, plant part, or plant cell may be a transgenic soybean plant, transgenic soybean plant seed, transgenic soybean plant part, or transgenic soybean plant cell.
- FIGS. 1A and IB show a schematic illustration of the GmTDFl-1 gene and the GmTDFl- 1 amino acid sequence.
- FIG. 1A shows the exon-intron structure of the GmTDFl-1 gene based on SEQ ID NO:4, where the exons are denoted by grey arrows and all regions lacking arrows denote introns and/or untranslated regions. The approximate site of the T -> A point mutation is also shown.
- FIG. IB shows an alignment of the amino acid sequences of Exon 2 encoded by the wildtype (Exon2_WT) and the ms6 mutant allele (Exon2_ms6). The amino acid sequence of Exon2_WT is based on SEQ ID NO:5. The denotes the site of the Leu to His amino acid change that results from the T -> A point mutation in the nucleic acid sequence.
- FIG. 2 shows a schematic illustration of the sites targeted for editing by the guide RNAs (gRNAs) in the GmTDFl-1 gene.
- the top illustration shows the location of the three exons where the exons are denoted by grey arrows and all regions lacking arrows denote introns and/or untranslated regions. The approximate site of the T -> A point mutation is also shown.
- the bottom illustration denotes the locations and binding directions of the three gRNAs, gRNA_-103C, gRNA_+221C, and gRNA_-l 113G, where each gRNA and the directionality of the strand it binds is denoted by a grey arrowhead.
- FIGS. 3A and 3B show illustrations of alignments of the monoallelic Edits 1-8 (based on SEQ ID NOs:26-33), the GmTDFl-1 coding sequence (based on SEQ ID NOG), and the GmTDFl-1 gene (based on SEQ ID NO:4).
- FIG. 3A shows an alignment of the sequences of each of the Edits 1-8, the GmTDFl-1 coding sequence, and the GmTDFl-1 gene.
- the gaps shown in the GmTDFl-1 coding sequence are representative of introns and/or untranslated regions whereas the gaps shown in the edited sequences are representative of deletions and/or 5 ’and 3’ untranslated regions.
- Edit 3 contains a 1 bp deletion which is not visualized in FIG. 3 A, but it is visualized in FIG. 3B.
- FIG. 3B shows sequence alignments and the specific deletions at a nucleotide level. A indicates that the sequence does not have a base alignment at that position, which represents a deletion in the edited sequences or an intron in the coding sequence.
- FIG. 4 shows an illustration of an alignment of Edits 9-24 (based on SEQ ID NOs:34, 36- 44, 46-50, and 52), the GmTDFl-1 coding sequence (based on SEQ ID NO:3), and the GmTDFl- 1 gene (based on SEQ ID NO:4).
- the gaps shown in the GmTDFl-1 coding sequence are representative of introns and/or untranslated regions whereas the gaps shown in the edited sequences are representative of deletions and/or 5’ and 3’ untranslated regions.
- FIGS. 5A, 5B, 5C, 5D, 5E, 5F, 5G, and 5H show an illustration of an alignment of Edits 25-28 (based on SEQ ID NOs:35, 45, 51, and 53), which comprise inversions, and the GmTDFl- 1 gene (based on SEQ ID NO:4).
- a shown in the GmTDFl-1 gene indicates that the sequence does not have a base alignment at that position, likely caused by the algorithm making the best alignment with the inverted sequences.
- a in SEQ ID NOs:35, 45, 51 , and 53 may represent an absence of a base alignment (as in the GmTDFl-1 gene) or may represent a deletion in the edited sequences, including the loss of nucleotides flanking inversions as a result of DNA repair mechanisms at the cut and ligation sites during the gene editing process, and/or 5’ or 3’ untranslated regions.
- FIG. 6 shows images of a phenotypic comparison of representative pod growth in wild-type plants and plants comprising the heteroallelic combination of mutant alleles at the GmTDFl-1 locus (Edit 9/Edit 25).
- the left panel shows normal pod growth in the wild-type plant and the right panel shows the failed pod growth in the Edit 9/Edit 25 plant.
- FIG. 7 shows images of a plant comprising the heteroallelic combination of mutant alleles at the GmTDFl-1 locus (Edit 9/Edit 25) before and after fertilization in a pod growth rescue experiment.
- the left panel shows the failed pod growth in an Edit 9/Edit 25 plant before pollination and the right panel shows developing pods in an Edit 9/Edit 25 plant after the plant was pollinated with donor pollen from wild- type plants.
- FIG. 8 shows images of a comparison of anther development in wild-type and GmTDFl-1 mutant plants in dissected flowers.
- the left panel shows representative wild-type anther development with pollen shattered.
- the right panel shows representative edited GmTDFl-1 mutant (male sterile) abnormal anther development.
- FIG. 9 shows images of the micro anatomical dissection of anthers from Arabiclopsis thaliana and Glycine max plants, specifically from AtTDFl mutant, GmTDFl-1 mutant, and wildtype plants.
- the leftmost top and bottom panels show wild-type Arabiclopsis anther sections (top) and homozygous AtTDFl mutant Arabidopsis anther sections (bottom).
- “E” designates epidermis
- “En” designates cndothccium
- T designates tapetum
- MSp designates microsporcs
- ML designates middle layer.
- the middle and rightmost top panels show wild-type and GmTDFl-1 heterozygous mutant soybean anther sections.
- the middle and rightmost bottom panels show GmTDFl-1 homozygous mutant soybean anther sections with arrows highlighting phenotypic characteristics that are different from the wild-type.
- FIG. 10 shows a diagrammatic representation of the alleles of Edits 1-8 shown in relation to the GmTDFl-1 Intron 1-Exon 2 border and T- ⁇ A mutation.
- the leftmost panel provides the edit ID.
- the middle panel shows the size (Abp) and relative location of the deletion in that edit.
- the rightmost panel describes the type of mutation caused by the edit.
- the thin, grey, vertical line represents the splice junction between Intron 1 and Exon 2.
- the wide, grey, vertical box represents the location of the classical ms6 T->A mutation.
- FIGS. 11A, 11B, 11C, 11D, and HE show an alignment of TDFl-like amino acid sequences across numerous plant species. Alignment of AtTDFl, GmTDFl-1, and TDFl-like proteins from both dicot and monocot plant species is provided with consensus sequence for the amino acid found in the majority of sequences. The region of highest conservation (the SANT/Myb functional domain) is located in the N-terminal region, while the serine rich domain highly conserved in dicots begins around position 262 of the amino acid sequences.
- SEQ ID NO:1 is a 954-nucleotide sequence representing the coding region of the wild-type Arabidopsis thaliana TDF1 (AtTDFl) gene.
- SEQ ID NO:2 is a 317-amino acid sequence representing the Arabidopsis thaliana TDF1 (AtTDFl) protein sequence.
- SEQ ID NO:2 is the sequence of the polypeptide encoded by SEQ ID NO:1.
- SEQ ID NO:3 is a 1038-nucleotide sequence representing the coding region of the wild-type Glycine max TDF1-1 (GmTDFl-1) gene.
- SEQ ID NO:4 is a 2094-nucleotide sequence representing the genomic DNA sequence of the wild-type GmTDFl-1 gene. The sequence includes a 5’ untranslated region (UTR), the GmTDFl-1 gene including introns and exons, and a 3’ UTR.
- SEQ TD NO:5 is a 345-amino acid sequence representing the Glycine max TDF1-1 (GmTDFl-1) protein sequence.
- SEQ ID NO:5 is the sequence of the polypeptide encoded by the nucleic acid sequence of SEQ ID NO:3.
- SEQ ID NO:6 is a 33-nucleotide sequence corresponding to a forward primer referred to as “NGMAX009625172 - FP” used to genotype seeds and/or plants for GmTDFl-1.
- SEQ ID NO:7 is a 25-nucleotide sequence corresponding to a reverse primer referred to as “NGMAX009625172 - RP” used to genotype seeds and/or plants for GmTDFl-1.
- SEQ ID NO:8 is a 16-nucleotide sequence corresponding to a probe referred to as “NGMAX009625172 - F” used for marker assisted selection of seeds and/or plants for GmTDFl- 1.
- SEQ ID NO:9 is a 15 -nucleotide sequence corresponding to a probe referred to as “NGMAX009625172 - V” used for marker assisted selection of seeds and/or plants for GmTDFl- 1.
- SEQ ID NO: 10 is a 21-nucleotide sequence corresponding to a reverse primer referred to as “ms6 - RP1” used to genotype seeds and/or plants for GmTDFl-1.
- SEQ ID NO: 11 is a 30-nucleotide sequence corresponding to a reverse primer referred to as “ms6 - RP2” used to genotype seeds and/or plants for GmTDFl-1.
- SEQ ID NO: 12 is a 25-nucleotide sequence corresponding to a forward primer referred to as “ms6 - FP1” used to genotype seeds and/or plants for GmTDFl-1.
- SEQ ID NO: 14 is a 15-nucleotide sequence corresponding to a probe referred to as “GMMS6S22427963_2_M” used for marker assisted selection of seeds and/or plants for GmTDFl-1.
- SEQ ID NO: 18 is a 20-nucleotide sequence representing a common scaffold compatible with a Cpfl gene from Lachnospiraceae bacterium ND2006.
- SEQ ID NO: 19 is a 365-nucleotide sequence representing a promoter from Dahlia mosaic virus.
- SEQ ID NO:20 is a 3681 -nucleotide sequence representing a Cpfl RNA-guided endonuclease enzyme codon-optimized for plants from Lachnospiraceae bacterium ND2006.
- SEQ ID NO:21 is a 30-nucleotide sequence representing a nuclear localization signal from Solatium lycopersicum.
- SEQ ID NO:22 is a 427-nucleotide sequence representing a Pol3 promoter from Glycine max.
- SEQ ID NO:23 is a 23-nucleotide guide RNA spacer sequence that corresponds to positions 328-350 on SEQ ID NO:4 in a forward orientation.
- SEQ ID NO:24 is a 23-nucleotide guide RNA spacer sequence that corresponds to positions 188-210 on SEQ ID NO:4 in a reverse orientation.
- SEQ ID NO:25 is a 23-nucleotide guide RNA spacer sequence that corresponds to positions 1198-1220 on SEQ ID NO:4 in a reverse orientation.
- SEQ ID NO:26 is a 1716-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 1.”
- SEQ ID NO:27 is a 1714-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 2.”
- SEQ ID NO:28 is a 1723-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 3.”
- SEQ ID NO:29 is a 1718-nuclcotidc sequence representing an edited GmTDFl-1 allele corresponding to “Edit 4.”
- SEQ ID NQ:30 is a 1715-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 5.”
- SEQ ID NO:31 is a 1716-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 6.”
- SEQ ID NO:32 is a 1713-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 7.”
- SEQ ID NO:33 is a 1717-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 8.”
- SEQ ID NO: 34 is a 1704-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 9.
- SEQ ID NO:35 is a 1653-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 25.”
- SEQ ID NO:36 is a 1709-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 10.
- SEQ ID NO:37 is a 1705 -nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 11.”
- SEQ ID NO:38 is a 1617-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 12.”
- SEQ ID NO:39 is a 1697-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 13.
- SEQ ID NO:40 is a 1715-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 14.”
- SEQ ID NO:41 is a 1705 -nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 15.
- SEQ ID NO:42 is a 699-nuclcotidc sequence representing an edited GmTDFl-1 allele corresponding to “Edit 16.
- SEQ ID NO:43 is a 1713-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 17.
- SEQ ID NO:44 is a 1706-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 18.
- SEQ ID NO:45 is a 1706-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 26.”
- SEQ ID NO:46 is a 669-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 19.
- SEQ ID NO:47 is a 1713-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 20.
- SEQ ID NO:48 is a 1692-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 21.”
- SEQ ID NO:49 is a 1602-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 22.
- SEQ ID NO:50 is a 1659-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 23.”
- SEQ ID NO:51 is a 1707 -nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 27.”
- SEQ ID NO:52 is a 1684-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 24.
- SEQ ID NO: 53 is a 1018-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 28.”
- SEQ ID NO:55 is a 321-amino acid sequence representing the Arachis hypogaea TDFl-like protein sequence.
- SEQ ID NO:56 is a 339-amino acid sequence representing the Cicer arietinum TDFl-likc protein sequence.
- SEQ ID NO:57 is a 333-amino acid sequence representing the Vitis vinifera TDFl-like protein sequence.
- SEQ ID NO:58 is a 268-amino acid sequence representing the Cucumis sativus TDFl-like protein sequence.
- SEQ ID NO:59 is a 320-amino acid sequence representing the Brassica napus TDFl-like protein sequence.
- SEQ ID NO:60 is a 387-amino acid sequence representing the Gossypium hirsutum TDFl- like protein sequence.
- SEQ ID NO:61 is a 263-amino acid sequence representing the Solarium lycopersicum TDF1 - like protein sequence.
- SEQ ID NO:62 is a 319-amino acid sequence representing the Zea mays TDFl-like protein sequence.
- SEQ ID NO:63 is a 328-amino acid sequence representing the Sorghum bicolor TDFl-like protein sequence.
- SEQ ID NO:64 is a 306-amino acid sequence representing the Oryza sativa subsp. japonica TDFl-like protein sequence.
- SEQ ID NO:65 is a 60-nucleotide sequence representing the 5’ junction of an inverted sequence and accompanying deletion(s) or modification(s). It comprises at least 20 bp on either side of the junction point.
- SEQ ID NO:66 is a 60-nucleotide sequence representing the 3’ junction of an inverted sequence and accompanying deletion(s) or modification(s). It comprises at least 20 bp on either side of the junction point.
- SEQ ID NO:67 is a 60-nucleotide sequence representing the 5’ junction of an inverted sequence and accompanying deletion(s) or modification(s). It comprises at least 20 bp on either side of the junction point.
- SEQ ID NO:68 is a 60-nucleotide sequence representing the 3’ junction of an inverted sequence and accompanying deletion(s) or modification(s). It comprises at least 20 bp on either side of the junction point.
- SEQ ID NO:69 is a 60-nucleotide sequence representing the 5’ junction of an inverted sequence and accompanying deletion(s) or modification(s). It comprises at least 20 bp on either side of the junction point.
- SEQ ID NO:70 is a 60-nucleotide sequence representing the 3’ junction of an inverted sequence and accompanying deletion(s) or modification(s). It comprises at least 20 bp on either side of the junction point.
- SEQ ID NO:71 is a 60-nucleotide sequence representing the 5’ junction of an inverted sequence and accompanying deletion(s) or modification(s). It comprises at least 20 bp on either side of the junction point.
- SEQ ID NO:72 is a 60-nucleotide sequence representing the 3’ junction of an inverted sequence and accompanying dclction(s) or modification(s). It comprises at least 20 bp on cither side of the junction point.
- the goal of crop breeding is to produce agricultural plants with desirable traits, such as increased grain yield, improved nutrition quality, and enhanced adaptation to environmental changes.
- desirable traits such as increased grain yield, improved nutrition quality, and enhanced adaptation to environmental changes.
- One of the ways breeders have produced such plants is through hybrid breeding (also called hybridization) using inbred parent varieties derived from diverse gene pools to generate hybrid plants. Because hybrids have a full set of genes from each unique parent, they have much more genetic diversity than cither inbred parent alone. Hybrids also have a combination of traits from each parent that enables the hybrids to outperform either of their inbred parents.
- hybrids are produced by controlled crosspollination, which requires hand-crossing selected lines to produce the desired Fi seed while also preventing unwanted self-pollination through use of a pollination control system.
- Manual emasculation namely the physical removal of either the male floral structure from hermaphrodite flowers or removal of whole male flowers from monoecious plants, remains the predominant method of pollination control in hybrid seed production. This process is time consuming and labor- intensive for most crop species but is especially difficult in soybean (Glycine max). Soybean is autogamous and is thus a self-pollinating species.
- male sterility in plants refers to the inability of a plant to produce functional anthers, pollens, and/or male gametes during their reproductive stages. When used in breeding, it eliminates the need for manual emasculation and, in many cases, pollination by hand. Male sterility can be induced through either chemical or genetic manipulation.
- male sterility systems in many crops have been developed through genetic manipulation of genes involved in male fertility. Plant male sterility can be caused either by nuclear genes alone (genic/genetic male sterility) or through mitochondrial genes that directly or indirectly affect nuclear gene function (cytoplasmic male sterility).
- Cytoplasmic male sterility is used in commercial crop hybrid seed production in a three-line system consisting of male-sterile lines, maintainer lines, and restorer lines, although it often suffers from poor genetic diversity, increased disease susceptibility, and unstable restoration of CMS lines.
- Genic/genetic male sterility is controlled by nuclear Male sterility (Ms) genes without the influence of cytoplasmic sequences controlled by nuclear genes alone.
- GMS can be genetically stable or environment-sensitive (EGMS).
- GMS is a common spontaneous occurrence in flowering plants and occurs in nearly all species. These abnormal plants are usually homozygous recessive (msms) in the self-pollinated progenies of heterozygous (Msms) individuals. Generally, such mutants appear in low frequency and are often lost in the selfpollinated crops. However, in cross-pollinated crops, the recessive mutant allele is protected by its dominant male fertile counterparts as hetero zygotes.
- Soybean is believed to have strong heterosis and utilizing this has been a strategy to increase soybean yield.
- the bottleneck in commercial application of soybean heterosis is a lack of suitable male-sterile lines.
- a three-line system based on CMS has been developed in soybean, very few CMS lines have been identified and restoration of these lines is unreliable. This restricts full use of heterosis in a commercial hybrid seed production system.
- GMS can overcome these disadvantages, but large-scale production of seeds is greatly hindered because half of the progenies in a homozygous male sterile line will restore fertility when male sterile plants are crossed to heterozygous male fertile plants during seed propagation of the male sterile line.
- ms6 is one that has a recessive morphological trait marker that can be used for phenotypic selection.
- the ms6 mutant was identified as an environmentally stable nuclear male sterile mutation at the Ms6 locus that results in a non-pollen phenotype.
- Plants carrying at least one copy of the wild-type Ms6 allele are fertile, whereas ms6ms6 plants are female fertile and completely male sterile due to abnormalities in tapetum development in anthers.
- the Ms6 locus is closely linked the flower color locus W1 which is commonly used for selection of male sterile lines in hybrid seed production.
- W1 is the developmental stage at which they appear.
- Ms6_Wl_ plants (Ms6ms6Wlwl or Ms6Ms6WlWl) have purple hypocotyls and flowers and are fertile whereas those of the genotype ms6ms6wlwl have green hypocotyls and white flowers and are male sterile.
- the hypocotyl color trait is visible shortly after emergence and at that stage purple hypocotyl seedlings must be removed manually from the field.
- not all white flowered plants are male sterile, as white flowered fertile plants may arise from recombination between the W1 and Ms6 loci.
- the plants must therefore be inspected at the flowering stage for anther morphology and plants with normal anthers must be removed. This process is laborious and time consuming, especially due to the structure of soybean flowers.
- ms6 allele for male sterility, especially for large-scale soybean breeding.
- some high yielding and male fertile elite soybean lines have white flowers, and accordingly, the white flower color phenotype could not be used a marker for the presence of ms 6 in homozygous form.
- the ms6 allele also differs from other male sterile, female fertile genes in soybean due to an observed pleiotropic effect on smaller flower size.
- ms6ms6w 1 w 1 Male- sterile plants (ms6ms6w 1 w 1) had flowers of smaller size when compared to fertile, purple flowered Ms6_WlWl plants, however it is not clear if the smaller flower is due to homozygosity of ms6 or wl .
- Flower size is known to play a role in attracting pollinators as bumble bees prefer large floral displays. As such, a small flower phenotype may have a negative impact on pollination in the field.
- the male sterile phenotype conferred by the ms6 allele results from a point mutation that alters a single amino acid.
- Point mutations have the potential to revert back to wild-type by random mutagenesis, and are therefore not considered to be highly genetically stable.
- Successful utilization of male sterility in hybrid breeding requires a male sterility phenotype that is genetically stable and does not incur any yield penalty in various genetic backgrounds and environments, as unstable male sterility leads to low purity of hybrid seeds and reduced grain production.
- Arabidopsis In many plants, male sterility may result from dysfunction of either the tapetum or the reproductive cells, or both.
- Arabidopsis In Arabidopsis, several genes have been reported to be important for tapetai differentiation and function. Specifically, the Arabidopsis DEFECTIVE in TAPETAL DEVELOPMENT and FUNCTION 1 (AtTDFl) gene has been identified as being important in regulation of tapetai cell differentiation and function.
- the tapetum plays important roles in microspore development as it provides nutrients for the developing microspores and pollen, it produces sporopollenin which is deposited on the outer walls of the released microspores, and it produces the enzyme callase which degrades the callose walls of the microspore tetrads that are the products of meiosis.
- AtTDFl mutants were found to have aborted pollen development due to defects in tapetai function. This phenotype is similar to the soybean ms6 phenotype, where ms6ms6 plants are completely male sterile due to abnormalities in tapetum.
- AtTDFl-1 a homolog of AtTDFl on chromosome 13
- GmTDFl-2 a second soybean TDF1 gene was also identified on chromosome 19 (designated GmTDFl-2), it is believed that this locus may be redundant due to its low level of expression and its inability to compensate for the deficiency in ms6ms6 (i.e., male sterile) plants.
- the present disclosure represents a significant advance in the art in that it provides engineered alleles that confer stable male sterility phenotypes in soybeans and other crops, as well as methods for the production thereof, thereby offering improvements in genetic male sterility systems that can lead to reduced breeding times and increased yield and quality.
- the edited alleles described herein cause male sterility in soybean plants without having any pleiotropic effects.
- the methods and compositions disclosed herein offer the opportunity to create diversity that cannot be selected from conventional plant breeding or random mutagenesis.
- the present disclosure provides, in certain embodiments, methods and compositions for the creation of novel alleles at the Ms6 locus via editing of the GmTDFl-1 gene. Therefore, the present disclosure represents a significant advance in the art in that it permits the production of novel engineered alleles in soybean and other crops that confer male sterility phenotypes with the potential to accelerate crop breeding cycles and improved yield gains and quality through capturing heterosis.
- Genome editing can be used to make one or more edit(s) or mutation(s) at a desired target site in the genome of a plant, such as to change expression and/or activity of one or more genes, or to integrate an insertion sequence or transgene at a desired location in a plant genome. Any site or locus within the genome of a plant may potentially be chosen for making a genomic edit (or gene edit) or site-directed integration of a transgene, construct, or transcribable DNA sequence.
- a “target site” or “a target sequence” for genome editing refers to the location of a polynucleotide sequence within a plant genome that is bound and cleaved by a sitespecific nuclease to introduce a strand break into the nucleic acid backbone of the polynucleotide sequence and/or its complementary DNA strand within the plant genome.
- a target site may comprise, for example, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27 , at least 28, at least 29, or at least 30 consecutive nucleotides.
- a “target site” for an RNA-guided nuclease may comprise the sequence of either complementary strand of a double-stranded nucleic acid (DNA) molecule or chromosome at the target site.
- An RNA-guidcd nuclease may bind to a target site through, for example, a non-coding guide RNA (e.g., without being limiting, a CRISPR RNA (crRNA) or a single-guide RNA (sgRNA), as described further herein).
- a non-coding guide RNA e.g., without being limiting, a CRISPR RNA (crRNA) or a single-guide RNA (sgRNA), as described further herein).
- a non-coding guide RNA provided herein may be complementary to a target site (e.g., complementary to either strand of a double-stranded nucleic acid molecule or chromosome at the target site).
- a non-coding guide RNA may not be required for a non-coding guide RNA to bind or hybridize to a target site. For example, at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8 mismatches (or more) between a target site and a non-coding RNA may be tolerated.
- a “target region” or a “targeted region” refers to a polynucleotide sequence or region that is flanked by two or more target sites.
- a target region may be subjected to a mutation, insertion, deletion, substitution, duplication, or inversion.
- “flanked” when used to describe a target region of a polynucleotide sequence or molecule refers to two or more target sites of the polynucleotide sequence or molecule surrounding the target region, with one target site on each side of the target region.
- a “complement,” a “complementary sequence,” and a “reverse complement” are used interchangeably. All three terms refer to the inversely complementary sequence of a nucleotide sequence, i.e., to a sequence complementary to a given sequence in reverse order of the nucleotides.
- antisense refers to DNA or RNA sequences that are complementary to a specific DNA or RNA sequence. Antisense RNA molecules are singlestranded nucleic acids which can combine with a sense RNA strand or sequence or mRNA to form duplexes due to complementarity of the sequences.
- the term “antisense strand” refers to a nucleic acid strand that is complementary to the “sense” strand.
- the “sense strand” of a gene or locus is the strand of DNA or RNA that has the same sequence as an RNA molecule transcribed from the gene or locus (with the exception of uracil in RNA and thymine in DNA).
- a “targeted editing technique” refers to any method, protocol, or technique that allows the precise and/or targeted editing of a specific location in a genome of a plant (i.e., the editing is largely or completely non-random) using a site-specific nuclease, such as a meganuclease, a zine-finger nuclease (ZFN), an RNA-guided endonuclease (e.g., the CRISPR/Cas9 system or the CRISPR/Cpfl system), a TALE (transcription activator-like effector)-endonuclease (TALEN), a recombinase, or a transposase.
- a site-specific nuclease such as a meganuclease, a zine-finger nuclease (ZFN), an RNA-guided endonuclease (e.g., the CRISPR/Cas9 system
- editing refers to generating a targeted mutation, insertion, deletion, substitution, duplication, or inversion, of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 75, at least 100, at least 250, at least 500, at least 1000, at least 2500, at least 5000, at least 10,000, or at least 25,000 nucleotides of an endogenous plant genome nucleic acid sequence.
- an “edit” or “genomic edit” in the singular refers to one such targeted mutation, insertion, deletion, substitution, duplication, or inversion, whereas “edits” or “genomic edits” refers to two or more targeted mutation(s), insertion(s), deletion(s), substitution(s), duplication(s), and/or inversion(s), with each “edit” being introduced via a targeted editing technique.
- Genome editing or targeted editing can be facilitated via the use of one or more site-specific nucleases.
- a site-specific nuclease provided herein may be selected from the group consisting of zinc-finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), meganucleases, RNA-guided endonucleases (e.g., Cas9 and Cpfl), recombinases (without being limiting, for example, a serine recombinase attached to a DNA recognition motif, a tyrosine recombinase attached to a DNA recognition motif), transposases (without being limiting, for example, a DNA transposase attached to a DNA binding domain), or any combination thereof. See, e.g., Khandagale et al. (Plant Biotechnol Rep 10:327-343, 2016); and Gaj et al. (Trends Biotechnol. 31
- the targeted genome editing described herein may comprise the use of an RNA-guided endonuclease.
- an “RNA-guided endonuclease” refers to an RNA- guided DNA endonuclease associated with the CRISPR system.
- an RNA-guided endonuclease may be selected from the group consisting of Casl, CaslB, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csxl2), CaslO, Csyl, Csy2, Csy3, Cse1 , Cse2, Csc 1 , Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl , Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, CsxlO, Csxl6, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4, Cpfl (also known as Casl2a), CasX, CasY
- the CRISPR system in its native context, provides bacteria and archaea with immunity to invading foreign nucleic acids and relies on an RNA-guided endonuclease to cleave the invading DNA or RNA into short sequence fragments and incorporating them into the bacterial CRISPR genomic locus.
- the incorporated short sequences referred to as “protospacers”, and flanking direct repeats are transcribed and processed into CRISPR RNAs (crRNAs).
- crRNAs hybridize with /ra -activating crRNAs (tracrRNAs) to activate the RNA-guided Cas endonuclease to form a ribonucleoprotein (RNP) complex that is guided to a target site.
- tracrRNAs /ra -activating crRNAs
- RNP ribonucleoprotein
- a “protospacer adjacent motif’ (PAM) herein refers to a short nucleotide sequence adjacent to a target sequence (protospacer) that is recognized (targeted) by a guide polynucleotide/Cas endonuclease system described herein.
- a PAM sequence may be present in the genome immediately adjacent and upstream (z'.e. the 5’ end) of the genomic target site sequence complementary to the targeting sequence of the guide RNA - i.e., immediately downstream (3’) to the sense (+) strand of the genomic target site (relative to the targeting sequence of the guide RNA) as known in the art. See, e.g., Wu et al. (Quant Biol. 2(2):59-70, 2014).
- the Cas endonuclease may not successfully recognize a target DNA sequence if the target DNA sequence is not followed by a PAM sequence.
- the sequence and length of a PAM sequence herein can differ depending on the Cas endonuclease used.
- the PAM sequence can be of any length but is typically 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides long.
- CRISPR/Cas9 which is the CRISPR system from Streptococcus pyogenes, was adapted for use in eukaryotes and has been widely used for gene editing in plants.
- the CRISPR/Cas9 system requires both crRNA and tracrRNA to guide the Cas9 protein to recognize and cleave the target DNA double helix.
- Cas9 recognizes the genomic PAM sequence 5’-NGG-3’ (where N is any nucleotide) and, when located on the sense strand adjacent to the target site, will create a blunt- end DSB at the target site, specifically the 5 '-end of the PAM site.
- Cas9 has been observed to recognize other PAM sequences, such as 5’-NAG-3’and 5’-NGA-3,’ which may result in cleavage of non-specific DNA sequences.
- CRISPR/Cpfl is also known as Casl2a
- Casl2a CRISPR/Cpfl
- CRISPR/Cpfl functions in a manner similar to CRISPR/Cas9, it is an even simpler system than CRISPR/Cas9.
- CRISPR/Cpfl requires only one crRNA molecule and no tracrRNA to cleave DNA.
- Cpfl recognizes the genomic PAM sequence 5’-TTTV-3’ (where V is A, G, or C) or 5’-TTN-3’, depending on the Cpfl ortholog. See e.g., Alok et al. Front. Plant Sci. 11 :264, 2020).
- Cpfl recognizes the genomic PAM located on the sense strand adjacent to the target site, it will generate a staggered DSB with a 4- or 5-nt 5' overhang at the target site, specifically the 3'-end of the PAM site.
- a guide RNA molecule may be further provided to direct the endonuclease to a target site in the genome of the plant via base-pairing or hybridization to cause a DSB or nick at or near the target site.
- the guide RNA may be transformed or introduced into a plant cell or tissue as a gRNA molecule, or as a recombinant DNA molecule, construct or vector comprising a transcribable DNA sequence encoding the guide RNA operably linked to a promoter.
- a guide RNA may comprise, for example, a CRISPR RNA (crRNA), a single-chain guide RNA (sgRNA), or any other RNA molecule that may guide or direct an endonuclease to a specific target site in the genome.
- crRNA CRISPR RNA
- sgRNA single-chain guide RNA
- a “single-chain guide RNA” is a RNA molecule comprising a crRNA covalently linked a tracrRNA by a linker sequence, which may be expressed as a single RNA transcript or molecule.
- the guide RNA comprises a guide or targeting sequence that is identical or complementary to a target site within the plant genome, such as at or near a TDF1 gene, such as the GmTDFl-1 gene in soybean.
- Nucleic acid molecules provided herein can combine a crRNA and a tracrRNA into one nucleic acid molecule in what is herein referred to as a “single guide RNA (sgRNA).
- the guide RNA is typically a non-coding RNA molecule that does not encode a protein.
- the guide sequence of the guide RNA may be at least 10 nucleotides in length, such as 12-40 nucleotides, 12-30 nucleotides, 12-20 nucleotides, 12-35 nucleotides, 12-30 nucleotides, 15-30 nucleotides, 17-30 nucleotides, or 17-25 nucleotides in length, or about 12, 13, 1 , 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 or more nucleotides in length.
- the guide sequence may be at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical or complementary to at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, or at least 25 or more consecutive nucleotides of a DNA sequence at the genomic target site.
- RNA-guided endonucleases may be delivered as a protein with or without a guide RNA, or the guide RNA may be complexed with the RNA-guided endonuclease and delivered as a ribonucleoprotein (RNP).
- RNP ribonucleoprotein
- a target gene for genome editing may be any of the DEFECTIVE in TAPETAL DEVELOPMENT and FUNCTION 1 -like genes described herein, including but not limited to one of the two soybean TDF1 genes (hereafter referred to as GmTDF 1 -1 ).
- a guide RNA comprising a guide sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical or complementary to at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, or more consecutive nucleotides of SEQ ID NO:4 or a sequence complementary thereto (e.g., 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or more consecutive nucleotides of SEQ ID NO:4 or a sequence complementary thereto).
- the term "consecutive" in reference to a polynucleotide or protein sequence means without deletions or gaps in the sequence.
- a guide RNA may further comprise one or more other structural or scaffold sequence(s), which may bind or interact with an RNA-guided endonuclease.
- Such scaffold or structural sequences may further interact with other RNA molecules (e.g., tracrRNA).
- a DSB or nick may be used to introduce one or more mutation in a target region in the genome of a plant.
- a “mutation” refers to the permanent alteration of the nucleic acid sequence or amino acid sequence in an organism, as compared to a naturally occurring reference nucleic acid sequence or amino acid sequence from the same organism. It will be appreciated that, when identifying a mutation, the reference sequence should be from the same nucleic acid (c.g., gene or non-coding RNA) or amino acid (c.g., protein).
- a difference between two sequences comprises a mutation
- the comparison should not be made between homologous sequences of two different species or between homologous sequences of two different varieties of a single species.
- the comparison should be made between the edited (e.g., mutated) sequence and the endogenous, nonedited (e.g., “wild-type”) sequence of the same organism.
- the mutation may comprise an insertion, a deletion, a substitution, a duplication, or an inversion.
- mutations such as insertions, deletions, substitutions, duplications, and/or inversions may be introduced at a target site via imperfect repair of the DSB or nick to produce a knock-out or knock-down of a gene.
- a “knock-out” of a gene may be achieved by inducing a DSB or nick at or near the endogenous locus of the gene that results in non-expression of the protein or expression of a non-functional protein, whereas a “knock-down” of a gene may be achieved in a similar manner by inducing a DSB or nick at or near the endogenous locus of the gene that is repaired imperfectly at a site that does not affect the coding sequence of the gene in a manner that would eliminate the function of the encoded protein.
- the site of the DSB or nick within the endogenous locus may be in the upstream or 5’ region of the gene (e.g., a promoter and/or enhancer sequence) to affect or reduce its level of expression.
- Insertions in the coding region of a gene refers to the addition of one or more extra nucleotides into the DNA. Insertions in the coding region of a gene may alter splicing of the mRNA (splice site mutation) or cause a shift in the reading frame (frameshift), both of which can significantly alter the gene product.
- deletion refers to the removal of one or more nucleotides from the DNA. Like insertion mutations, these mutations can alter splicing of the mRNA (splice site mutation) or cause a shift in the reading frame of the gene.
- substitution refers to an exchange of a single nucleotide for another.
- inversion refers to reversing the orientation of a chromosomal segment.
- An inversion can be accompanied by a loss of nucleotides flanking either one or both sites of the inversion due to DNA repair mechanisms occurring at the cut and ligation sites during the formation of an inversion.
- replication refers to the creation of multiple copies of chromosomal regions, increasing the dosage of the genes located within them.
- a “missense mutation” refers to a single nucleotide change that results in a codon that codes for a different amino acid.
- the codon “CGU” encodes an arginine amino acid. If a missense mutation changes the G to a U, producing a “CUU” codon, the codon now encodes a leucine amino acid. Missense mutations can be caused by an insertion, deletion, substitution, duplication, or inversion. The frameshift, missense, or nonsense mutations described herein lead to loss of function or expression of a targeted gene, such as a GmTDFl-1 gene.
- a “loss-of-function mutation” is a mutation in the coding sequence of a gene, which causes the function of the gene product, usually a protein, to be either reduced or completely absent.
- a loss- of-function mutation can, for instance, be caused by the truncation of the gene product.
- a phenotype associated with an allele with a loss-of-function mutation can be either recessive or dominant.
- the terms “natural mutation,” “naturally-occurring mutation,” or “native mutation,” refers to a mutation that arises spontaneously in nature without any involvement of laboratory or experimental procedures or under the exposure to mutagens. Without being bound by scientific theory, a naturally-occurring mutation can arise from a variety of sources, including errors in DNA replication, spontaneous lesion, and transposable elements (or transposon). As used herein, the terms “non-natural mutation,” “non-naturally-occurring mutation,” or “synthetic mutation” refers to non-spontaneous mutation that occurs as a result of experimental procedures, such as exposure to a mutagen or by a site-specific genome modification enzyme.
- the present disclosure provides a modified soybean plant, or plant part thereof, comprising a mutant allele of the GmTDFl -1 gene, wherein the mutant allele comprises at least one genome modification involving at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 110, at least 125, at least 150, at least 175, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, or at least 1100 consecutive nucleotides of a coding region, a non-coding region, or any combination thereof, of the endogenous GmTDFl-1 gene.
- the mutant allele comprises at least one genome modification involving at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least
- Such genome modifications may include: a 8 base pair deletion (bp) wherein the resulting nucleotide sequence is SEQ ID NO:26; a 10 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:27; a 1 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:28; a 6 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:29; a 9 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:30; a 8 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:31; a l l base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:32; a 7 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:33; a 12 base pair deletion and a 8 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:34; two inversions wherein the resulting
- SEQ ID NOs:2 and 55-64 are each the TDFl-like protein homologs found in Arabidopsis thaliana (SEQ ID NO:2), Arachis hypogaea (SEQ ID NO:55), Cicer arietinum (SEQ ID NO:56), Vitis vinifera (SEQ ID NO:57), Cucumis sativus (SEQ ID NO:58), Brassica napus (SEQ ID NO:59), Gossypium hirsutum (SEQ ID NO:60), Solarium lycopersicum (SEQ ID N0:61), Zea mays (SEQ ID NO:62), Sorghum bicolor (SEQ ID NO:63), and Oryza sativa subsp. japonica (SEQ ID NO:64).
- the present disclosure provides a modified soybean plant, or plant part thereof, comprising a mutant allele of the GmTDFl-1 gene, wherein the mutant allele comprises one or more junction sequences, wherein the junction sequences are at least 30, at least 60, or at least 100 nucleotides at the junction site.
- a “junction” or “junction site” is the connection point between the nucleotide sequences at the site of an insertion, deletion, substitution, duplication, or inversion. In the case of a deletion, the junction is the connection point at the site of the deletion of the sequences that previously flanked the deletion.
- the junction would be between nucleotide 1538 and nucleotide 1569.
- the junction is the connection point between the inserted, inverted, or substituted sequence and the flanking DNA sequences.
- one junction is found at the 5’ end of the insertion, substitution or inversion, and another junction is found at the 3’ end of the insertion, substitution, or inversion.
- a “junction sequence” refers to a DNA sequence of any length that spans a junction.
- a junction sequence can comprise at least 10 nucleotides, at least, 15 nucleotides, at least 20 nucleotides, at least 25 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, or more.
- Recombinant DNA constructs and vectors comprising a polynucleotide sequence encoding a site-specific nuclease, such as a zinc-finger nuclease (ZFN), a meganuclease, an RNA-guided endonuclease, a TALE-endonuclease (TALEN), a recombinase, or a transposase, wherein the coding sequence is operably linked to a plant expressible promoter.
- ZFN zinc-finger nuclease
- TALEN TALE-endonuclease
- RNA-guided endonucleases recombinant DNA constructs and vectors are further provided comprising a polynucleotide sequence encoding a guide RNA, wherein the guide RNA comprises a guide sequence of sufficient length having a percent identity or complementarity to a target site within the genome of a plant, such as at or near a targeted GmTDFl-1 gene.
- a polynucleotide sequence of a recombinant DNA construct and vector that encodes a site-specific nuclease or a guide RNA may be operably linked to a plant expressible promoter, such as an inducible promoter, a constitutive promoter, a tissue- specific promoter, etc.
- a recombinant DNA construct or vector may comprise a first polynucleotide sequence encoding a site- specific nuclease and a second polynucleotide sequence encoding a guide RNA that may be introduced into a plant cell together via plant transformation techniques.
- two recombinant DNA constructs or vectors may be provided including a first recombinant DNA construct or vector and a second DNA construct or vector that may be introduced into a plant cell together or sequentially via plant transformation techniques, wherein the first recombinant DNA construct or vector comprises a polynucleotide sequence encoding a site-specific nuclease and the second recombinant DNA construct or vector comprises a polynucleotide sequence encoding a guide RNA.
- a recombinant DNA construct or vector comprising a polynucleotide sequence encoding a sitespecific nuclease may be introduced via plant transformation techniques into a plant cell that already comprises (or is transformed with) a recombinant DNA construct or vector comprising a polynucleotide sequence encoding a guide RNA.
- a recombinant DNA construct or vector comprising a polynucleotide sequence encoding a guide RNA may be introduced via plant transformation techniques into a plant cell that already comprises (or is transformed with) a recombinant DNA construct or vector comprising a polynucleotide sequence encoding a sitespecific nuclease.
- a first plant comprising (or transformed with) a recombinant DNA construct or vector comprising a polynucleotide sequence encoding a site-specific nuclease may be crossed with a second plant comprising (or transformed with) a recombinant DNA construct or vector comprising a polynucleotide sequence encoding a guide RNA.
- a second plant comprising (or transformed with) a recombinant DNA construct or vector comprising a polynucleotide sequence encoding a guide RNA.
- Such recombinant DNA constructs or vectors may be transiently transformed into a plant cell or stably transformed or integrated into the genome of a plant cell.
- vectors comprising polynucleotides encoding a site-specific nuclease, and optionally one or more, two or more, three or more, or four or more gRNAs are provided to a plant cell by transformation methods known in the art (e.g., without being limiting, particle bombardment, PEG-mediated protoplast transfection or Agro ⁇ acteriuzn-mediated transformation).
- vectors comprising polynucleotides encoding a Cpfl nuclease, and optionally one or more, two or more, three or more, or four or more gRNAs arc provided to a plant cell by transformation methods known in the art (e.g., without being limiting, particle bombardment, PEG-mediated protoplast transfection or Agrobacterium-mediated transformation).
- vectors comprising polynucleotides encoding a Cpfl and, optionally one or more, two or more, three or more, or four or more crRNAs are provided to a cell by transformation methods known in the art (e.g., without being limiting, viral transfection, particle bombardment, PEG- mediated protoplast transfection or Agrobacterium-mediated transformation).
- site-specific nucleases such as zine-finger nucleases, TALENs, meganucleases, and recombinases, are not RNA-guided and instead rely on their protein structure to determine their target site for causing the DSB or nick, or they are fused, tethered or attached to a DNA- binding protein domain or motif.
- the protein structure of the site-specific nuclease (or the fused/attached/tethered DNA binding domain) may target the site-specific nuclease to the target site.
- non-RNA-guided site-specific nucleases such as zine-finger nucleases, TALENs, meganucleases, and recombinases
- a target site at or near the genomic locus of an endogenous gene of a plant, such as the GmTDFl-1 gene in soybean, to create a DSB or nick at such genomic locus to knockout or knockdown expression of the GmTDFl-1 gene via repair of the DSB or nick.
- an engineered site-specific nuclease such as a recombinase, zinc finger nuclease (ZFN), meganuclease, or TALEN, may be designed to target and bind to a target site within the genome of a plant corresponding to a sequence within SEQ ID NO:4, or its complementary sequence, to create a DSB or nick at the genomic locus for the GmTDFl-1 gene, which may then lead to the creation of a mutation or insertion of a sequence at the site of the DSB or nick, through cellular repair mechanisms, which may be guided by a donor molecule or template.
- a recombinase zinc finger nuclease (ZFN), meganuclease, or TALEN
- the site-specific nuclease may be a zine-finger nuclease.
- Zinefinger nucleases are synthetic proteins consisting of an engineered zinc finger DNA-binding domain fused to a cleavage domain (or a cleavage half-domain), which may be derived from a restriction endonuclease (e.g., FokI).
- the DNA binding domain may be canonical (C2H2) or non- canonical (e.g., C3H or C4).
- the DNA-binding domain can comprise one or more zinc fingers (e.g., 2, 3, 4, 5, 6, 7, 8, 9 or more zinc fingers) depending on the target site but may typically be composed of 3-4 (or more) zinc-fingers. Multiple zinc fingers in a DNA-binding domain may be separated by linker scqucncc(s).
- ZFNs can be designed to cleave almost any stretch of doublestranded DNA by modification of the zinc finger DNA-binding domain. ZFNs form dimers from monomers composed of a non-specific DNA cleavage domain (e.g., derived from the FokI nuclease) fused to a DNA-binding domain comprising a zinc finger array engineered to bind a target site DNA sequence.
- amino acids at positions -1, +2, +3, and +6 relative to the start of the zinc finger a-helix, which contribute to site-specific binding to the target site, can be changed and customized to fit specific target sequences.
- the other amino acids may form a consensus backbone to generate ZFNs with different sequence specificities.
- a ZFN as used herein, is broad and includes a monomeric ZFN that can cleave double stranded DNA without assistance from another ZFN.
- the term ZFN may also be used to refer to one or both members of a pair of ZFNs that are engineered to work together to cleave DNA at the same site. Because the DNA-binding specificities of zinc finger domains can be re-engineered using one of various methods, customized ZFNs can theoretically be constructed to target nearly any target sequence (e.g., at or near a gene in a plant genome).
- Publicly available methods for engineering zinc finger domains include Context-dependent Assembly (CoDA), Oligomerized Pool Engineering (OPEN), and Modular Assembly.
- the site-specific nuclease may be a TALEN.
- TALENs are artificial restriction enzymes generated by fusing the TALE DNA binding domain to a nuclease domain.
- the nuclease is selected from a group consisting of PvuII, MutH, TevI, Fokl, Alwl, Mlyl, Sbfl, Sdal, StsI, CleDORF, Clo051, and Pept071.
- the Fokl monomers dimerize and cause a doublestranded DNA break at the target site.
- TALEN as used herein, is broad and includes a monomeric TALEN that can cleave double- stranded DNA without assistance from another TALEN.
- the term TALEN also refers to one or both members of a pair of TALENs that work together to cleave DNA at the same site.
- variants of the FokI cleavage domain with mutations have been designed to improve cleavage specificity and cleavage activity.
- the Fokl domain functions as a dimer, requiring two constructs with unique DNA binding domains for sites in the target genome with proper orientation and spacing.
- PvuII, MutH, and TevI cleavage domains are useful alternatives to Fokl and Fokl variants for use with TALEs.
- PvuII functions as a highly specific cleavage domain when coupled to a TALE (see Yank et al., PLoS One 8:e82539, 2013).
- MutH is capable of introducing strand- specific nicks in DNA (see Gabsalilow et al., Nucleic Acids Research. 41:e83, 2013).
- TevI introduces doublestranded breaks in DNA at targeted sites (see Beurdeley et al., Nature Communications 4:1762, 2013).
- TALEs can be engineered to bind practically any DNA sequence, such as at or near the genomic locus of a gene in a plant.
- TALE has a central DNA-binding domain composed of 13-28 repeat monomers of 33-34 amino acids.
- the amino acids of each monomer are highly conserved, except for hypervariable amino acid residues at positions 12 and 13.
- the two variable amino acids are called repeat- variable diresidues (RVDs).
- RVDs repeat- variable diresidues
- the amino acid pairs NI, NG, HD, and NN of RVDs preferentially recognize adenine, thymine, cytosine, and guanine/adenine, respectively, and modulation of RVDs can recognize consecutive DNA bases. This simple relationship between amino acid sequence and DNA recognition has allowed for the engineering of specific DNA binding domains by selecting a combination of repeat segments containing the appropriate RVDs.
- the site-specific nuclease may be a meganuclease.
- Meganucleases which are commonly identified in microbes, such as the LAGLIDADG family of homing endonucleases, are unique enzymes with high activity and long recognition sequences (> 14 bp) resulting in site-specific digestion of target DNA.
- Engineered versions of naturally occurring meganucleases typically have extended DNA recognition sequences (for example, 14 to 40 bp).
- the engineering of mcganuclcascs can be more challenging than ZFNs and TALENs because the DNA recognition and cleavage functions of meganucleases are intertwined in a single domain.
- the site-specific nuclease may be a recombinase.
- recombinases that may be used include a serine recombinase attached to a DNA recognition motif, a tyrosine recombinase attached to a DNA recognition motif, or any recombinase enzyme known in the art attached to a DNA recognition motif.
- the site-specific nuclease is a recombinase or transposase, which may be a DNA transposase or recombinase attached or fused to a DNA binding domain.
- recombinases include a tyrosine recombinase selected from the group consisting of a Cre recombinase, a Gin recombinase, a Flp recombinase, and a Tnpl recombinase attached to a DNA recognition motif provided herein.
- a Cre recombinase or a Gin recombinase provided herein is tethered to a zine-finger DNA-binding domain, a TALE DNA-binding domain, or a Cas9 nuclease.
- a serine recombinase selected from the group consisting of a PhiC31 integrase, an R4 integrase, and a TP- 901 integrase may be attached to a DNA recognition motif provided herein.
- a DNA transposase selected from the group consisting of a TALE-piggyBac and TALE-Mutator may be attached to a DNA binding domain provided herein.
- a “gene” refers to a nucleic acid sequence forming a genetic and functional unit and coding for one or more sequence-related RNA and/or polypeptide molecules.
- a gene generally contains a coding region operably linked to appropriate regulatory sequences that regulate the expression of a gene product (e.g., a polypeptide or a functional RNA).
- a gene can have various sequence elements, including, but not limited to, a promoter, an untranslated region(s) (UTR), exons, introns, and other upstream or downstream regulatory sequences.
- locus is a chromosomal locus or region where a polymorphic nucleic acid, trait determinant, gene, or marker is located.
- locus can be shared by two homologous chromosomes to refer to their corresponding locus or region.
- allele refers to an alternative nucleic acid sequence of a gene or at a particular locus (e.g., a nucleic acid sequence of a gene or locus that is different than other alleles for the same gene or locus).
- Such an allele can be considered (i) wild-type or (ii) mutant if one or more mutations or edits arc present in the nucleic acid sequence of the mutant allele relative to the wild-type allele.
- a mutant allele for a gene may have a reduced or eliminated activity or expression level for the gene relative to the wild- type allele.
- a first allele can occur on one chromosome, and a second allele can occur at the same locus on a second homologous chromosome.
- homozygous refers to a genotype comprising two identical alleles at a given locus in a diploid genome.
- CRISPR- mediated gene editing can result in biallelic edits (that is, different edits are made to the same locus on corresponding homologous chromosomes) resulting in a genotype comprising two nonidentical mutant alleles at a given locus in a diploid genome in Ro plants.
- the non-identical mutant alleles may differ, for example, based on having different additions, deletions, and/or substitutions of one or more nucleotides relative to each other, with the additions, deletions, and/or substitutions of each being sufficiently severe to cause a loss of function of the TDF1 protein encoded by each.
- plants comprising such genotypes may also be referred to as comprising a heteroallelic combination or transheterozygous edits.
- “complementation” or “complement” refers to the ability of two non-identical mutations to restore the wild-type phenotype when present in the same plant.
- Complementation can be full (wild-type phenotype) or partial (milder mutant phenotype).
- “complementation” or “complement” describes when non-identical, recessive mutant alleles act in a complementary manner to produce the recessive phenotype.
- said recessive phenotype is male sterility.
- a monoallelic edit includes events in which an edit is made to only one allele of a locus (z'.e., a modification to the endogenous TDF1 gene in only one of the two homologous chromosomes) and can result in a cell that is heterozygous for the targeted TDF1 modification.
- heterozygous describes a genotype comprising a mutant/edited allele and a wild-type allele at a given locus in a diploid genome.
- a wild-type gene or “wild-type allele” refers to a gene or allele having a sequence or genotype that is most common in a particular plant species, or another sequence or genotype having only natural variations, polymorphisms, or other silent mutations relative to the most common sequence or genotype that do not significantly impact the expression and activity of the gene or allele.
- a “wild-type” gene or allele contains no variation, polymorphism, or any other type of mutation that substantially affects the normal function, activity, expression, or phenotypic consequence of the gene or allele relative to the most common sequence or genotype.
- a “native copy” of a gene refers to a gene that originates from within a given organism, cell, tissue, genome, or chromosome that was not previously modified by human action.
- a “native protein” refers to a protein encoded by a native gene.
- pleiotropy refers to the phenomenon where a single gene (i.e. pleiotropic gene) or locus affects two or more phenotypic traits.
- a “pleiotropic effect” refers to the influence of a mutation in a pleiotropic gene on the traits associated with said gene.
- variant refers to molecules with some differences, generated synthetically or naturally, in their nucleotide or amino acid sequences as compared to a reference (native) polynucleotides or polypeptides, respectively. These differences include substitutions, insertions, deletions or any desired combinations of such changes in a native polynucleotide or amino acid sequence.
- the term “expression” refers to the biosynthesis of a gene product, and typically the transcription and/or translation of a nucleotide sequence, such as an endogenous gene, a heterologous gene, a transgene or an RNA and/or protein coding sequence, in a cell, tissue, organ, or organism, such as a plant, plant part or plant cell, tissue or organ.
- polynucleotide (DNA or RNA) molecule, protein, construct, vector, etc. refers to a polynucleotide or protein molecule or sequence that is man-made and not normally found in nature, and/or is present in a context in which it is not normally found in nature, including a polynucleotide (DNA or RNA) molecule, protein, construct, etc., comprising a combination of two or more polynucleotide or protein sequences that would not naturally occur together in the same manner without human intervention, such as a polynucleotide molecule, protein, construct, etc., comprising at least two polynucleotide or protein sequences that are operably linked but heterologous with respect to each other.
- the term “recombinant” can refer to any combination of two or more DNA or protein sequences in the same molecule (e.g., a plasmid, construct, vector, chromosome, protein, etc.) where such a combination is man-made and not normally found in nature.
- a plasmid, construct, vector, chromosome, protein, etc. e.g., a plasmid, construct, vector, chromosome, protein, etc.
- a recombinant polynucleotide or protein molecule, construct, etc. can comprise polynucleotide or protein sequence(s) that is/are (i) separated from other polynucleotide or protein sequence(s) that exist in proximity to each other in nature, and/or (ii) adjacent to (or contiguous with) other polynucleotide or protein scqucncc(s) that are not naturally in proximity with each other.
- Such a recombinant polynucleotide molecule, protein, construct, etc. can also refer to a polynucleotide or protein molecule or sequence that has been genetically engineered and/or constructed outside of a cell.
- a recombinant DNA molecule can comprise any engineered or man-made plasmid, vector, etc., and can include a linear or circular DNA molecule.
- plasmids, vectors, etc. can contain various maintenance elements including a prokaryotic origin of replication and selectable marker, as well as one or more transgenes or expression cassettes perhaps in addition to a plant selectable marker gene, etc.
- operably linked refers to a functional linkage between a promoter or other regulatory element and an associated transcribable DNA sequence or coding sequence of a gene (or transgene), such that the promoter, etc., operates or functions to initiate, assist, affect, cause, and/or promote the transcription and expression of the associated transcribable DNA sequence or coding sequence, at least in certain cell(s), tissue(s), developmental stage(s), and/or condition(s).
- references in this application to an “isolated DNA molecule” or an “isolated polynucleotide,” or an equivalent term or phrase, is intended to mean that the DNA molecule or polynucleotide is one that is present alone or in combination with other compositions, but not within its natural environment.
- nucleic acid elements such as a coding sequence, intron sequence, untranslated leader sequence, promoter sequence, transcriptional termination sequence, and the like, that are naturally found within the DNA of the genome of an organism are not considered to be “isolated” so long as the element is within the genome of the organism and at the location within the genome in which it is naturally found.
- each of these elements, and subparts of these elements would be “isolated” within the scope of this disclosure so long as the element is not within the genome of the organism and at the location within the genome in which it is naturally found.
- a nucleotide sequence encoding a protein or any naturally occurring variant of that protein would be an isolated nucleotide sequence so long as the nucleotide sequence was not within the DNA of the organism in which the sequence encoding the protein is naturally found.
- a synthetic nucleotide sequence encoding the amino acid sequence of the naturally occurring protein would be considered to be isolated for the purposes of this disclosure.
- any transgenic nucleotide sequence i.e., the nucleotide sequence of the DNA inserted into the genome of the cells of a plant or bacterium, or present in an extrachromosomal vector, would be considered to be an isolated nucleotide sequence whether it is present within the plasmid or similar structure used to transform the cells, within the genome of the plant or bacterium, or present in detectable amounts in tissues, progeny, biological samples or commodity products derived from the plant or bacterium.
- promoter can generally refer to a DNA sequence that contains an RNA polymerase binding site, transcription start site, and/or TATA box and assists or promotes the transcription and expression of an associated transcribable polynucleotide sequence and/or gene (or transgene).
- a promoter can be synthetically produced, varied, or derived from a known or naturally occurring promoter sequence or other promoter sequence.
- a promoter can also include a chimeric promoter comprising a combination of two or more heterologous sequences.
- a promoter of the present disclosure can thus include variants or fragments of promoter sequences that are similar in composition, but not identical to, other promoter sequence(s) known or provided herein.
- a promoter provided herein, or variant or fragment thereof, may comprise a “minimal promoter” which provides a basal level of transcription and is comprised of a TATA box or equivalent DNA sequence for recognition and binding of the RNA polymerase II complex for initiation of transcription.
- a promoter can be classified according to a variety of criteria relating to the pattern of expression of an associated coding or transcribable sequence or gene (including a transgene) operably linked to the promoter, such as constitutive, developmental, tissue-specific, inducible, etc. Promoters that drive expression in all or most tissues of the plant are referred to as “constitutive” promoters. Promoters that drive expression during certain periods or stages of development are referred to as “developmental” promoters.
- tissue-enhanced or “tissue-preferred” promoters.
- tissue-preferred causes relatively higher or preferential expression in a specific tissue(s) of the plant, but with lower levels of expression in other tissue(s) of the plant.
- Promoters that express within a specific tissue(s) of the plant, with little or no expression in other plant tissues are referred to as “tissue-specific” promoters.
- An “inducible” promoter is a promoter that initiates transcription in response to an environmental stimulus such as cold, drought or light, or other stimuli, such as wounding or chemical application.
- a promoter can also be classified in terms of its origin, such as being heterologous, homologous, chimeric, synthetic, etc.
- a “plant-expressible promoter” refers to a promoter that can initiate, assist, affect, cause, and/or promote the transcription and expression of its associated transcribablc DNA sequence, coding sequence or gene in a plant cell or tissue.
- heterologous in reference to a promoter or other regulatory sequence in relation to an associated polynucleotide sequence (e.g., a transcribable DNA sequence or coding sequence or gene) is a promoter or regulatory sequence that is not operably linked to such associated polynucleotide sequence in nature without human introduction - e.g., the promoter or regulatory sequence has a different origin relative to the associated polynucleotide sequence and/or the promoter or regulatory sequence is not naturally occurring in a plant species to be transformed with the promoter or regulatory sequence.
- an “endogenous gene” or an “endogenous locus” refers to a gene or locus at its natural and original chromosomal location.
- the “endogenous TDF1 gene” refers to the TDF1 genic locus at its original chromosomal location.
- an “exon” refers to a segment of a DNA or RNA molecule containing information coding for a protein or polypeptide sequence.
- an “intron” of a gene refers to a segment of a DNA or RNA molecule, which does not contain information coding for a protein or polypeptide, and which is first transcribed into an RNA sequence but then spliced out from a mature RNA molecule.
- an “untranslated region (UTR)” of a gene refers to a segment of an RNA molecule or sequence (e.g., a mRNA molecule) expressed from a gene (or transgene) but excluding the exon and intron sequences of the RNA molecule.
- An “untranslated region (UTR)” also refers to a DNA segment or sequence encoding such a UTR segment of an RNA molecule.
- An untranslated region can be a 5'-UTR or a 3'-UTR depending on whether it is located at the 5' or 3' end of a DNA or RNA molecule or sequence relative to a coding region of the DNA or RNA molecule or sequence (/. ⁇ ?., upstream (5') or downstream (3') of the exon and intron sequences, respectively).
- upstream refers to a nucleic acid sequence that is positioned before the 5' end of a linked nucleic acid sequence.
- downstream refers to a nucleic acid sequence is positioned after the 3' end of a linked nucleic acid sequence.
- 5' refers to the start of a coding DNA sequence or the beginning of an RNA molecule.
- 3' refers to the end of a coding DNA sequence or the end of an RNA molecule. It will be appreciated that an “inversion” refers to reversing the orientation of a given polynucleotide sequence.
- a “transcription termination sequence” refers to a nucleic acid sequence containing a signal that triggers the release of a newly synthesized transcript RNA molecule from an RNA polymerase complex and marks the end of transcription of a gene or locus.
- a "homolog” or “homologues” means a protein in a group of proteins that perform the same biological function, for example, proteins that belong to the same TDFl-like protein family and that provide a common enhanced trait in modified plants of this disclosure.
- Homologs are expressed by homologous genes.
- homologs include orthologs, for example, genes expressed in different species that evolved from common ancestral genes by speciation and encode proteins retain the same function, but do not include paralogs, z.e., genes that arc related by duplication but have evolved to encode proteins with different functions.
- Homologous genes include naturally occurring alleles and artificially-created variants.
- Homologs are inferred from sequence similarity, by comparison of protein sequences, for example, manually or by use of a computer-based tool.
- various pair-wise or multiple sequence alignment algorithms and programs are known in the ail, such as ClustalW or Basic Local Alignment Search Tool® (BLAST), etc., that can be used to compare the sequence identity or similarity between two or more nucleotide or protein sequences.
- BLAST can also be used, for example to search query protein sequences of a base organism against a database of protein sequences of various organisms, to find similar sequences.
- the generated summary Expectation value (E-value) can be used to measure the level of sequence similarity.
- a reciprocal query is used to filter hit sequences with significant E-values for ortholog identification.
- the reciprocal query entails search of the significant hits against a database of protein sequences of the base organism.
- a hit can be identified as an ortholog, when the reciprocal query's best hit is the query protein itself or a paralog of the query protein.
- orthologs are further differentiated from paralogs among all the homologs, which allows for the inference of functional equivalence of genes.
- percent identity As used herein in reference to two or more nucleotide or protein sequences is calculated by (i) comparing two optimally aligned sequences (nucleotide or protein) over a window of comparison, (ii) determining the number of positions at which the identical nucleic acid base (for nucleotide sequences) or amino acid residue (for proteins) occurs in both sequences to yield the number of matched positions, (iii) dividing the number of matched positions by the total number of positions in the window of comparison, and then (iv) multiplying this quotient by 100% to yield the percent identity.
- percent identity or “percent sequence identity” is being calculated in relation to a reference sequence without a particular comparison window being specified, then the percent identity is determined by dividing the number of matched positions over the region of alignment by the total length of the reference sequence. Accordingly, for purposes of the present application, when two sequences (query and subject) are optimally aligned (with allowance for gaps in their alignment), the “percent identity” for the query sequence is equal to the number of identical positions between the two sequences divided by the total number of positions in the query sequence over its length (or a comparison window), which is then multiplied by 100%.
- sequence similarity When percentage of sequence identity is used in reference to proteins it is recognized that residue positions which are not identical often differ by conservative amino acid substitutions, where amino acid residues are substituted for other amino acid residues with similar chemical properties (e.g., charge or hydrophobicity) and therefore do not change the functional properties of the molecule.
- sequences differ in conservative substitutions the percent sequence identity can be adjusted upwards to correct for the conservative nature of the substitution. Sequences that differ by such conservative substitutions are said to have “sequence similarity” or “similarity.” Sequences having a percent identity to a base sequence may exhibit the activity of the base sequence.
- Degeneracy of the genetic code provides the possibility to substitute at least one base of the protein encoding sequence of a gene with a different base without causing the amino acid sequence of the polypeptide produced from the gene to be changed.
- homolog proteins, or their corresponding nucleotide sequences have typically at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or even at least about 99.5% identity over the full length of a protein or its corresponding nucleotide sequence identified as being associated with imparting a male fertility phenotype when expressed in plant cells.
- a TDF1 gene or homolog thereof encodes a protein that has a functional domain with at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the functional domain of SEQ ID NO:5.
- Examples of homologs of the GmTDFl-1 protein (SEQ ID NO:5) include, but are not limited to, the sequences of SEQ ID NOs:55-64. Further homologs and paralogs may readily be identified using the BLAST procedure described above.
- alignment refers to a number of nucleotide bases or amino acid residue sequences aligned by lengthwise comparison so that components in common (i.e. , nucleotide bases or amino acid residues at corresponding positions) may be visually and readily identified. The fraction or percentage of components in common is related to the homology or identity between the sequences.
- An alignment such as that shown in FIGS. 11 A- HE, may be used to identify conserved domains and relatedness within these domains. An alignment may suitably be determined by means of computer programs known in the ail. Specialist databases also exist for the identification of domains, for example, SMART (Schultz et al., Proc. Natl. Acad. Sci.
- homologs of the proteins described herein are also identifiable by the presence of a conserved functional domain(s).
- domain refers to a set of amino acids conserved at specific positions along an alignment of sequences of evolutionarily related proteins. While amino acids at other positions can vary between homologs, amino acids that are highly conserved at specific positions indicate amino acids that are essential in the structure, stability, or activity of a protein. Identified by their high degree of conservation in aligned sequences of a family of protein homologs, they can be used as identifiers to determine if any polypeptide in question belongs to a previously identified polypeptide family (in this case, the proteins useful in the methods of the invention and nucleic acids encoding the same as defined herein).
- TDF1 is believed to be a member of the MYB transcription factor family which is defined by a highly conserved MYB DNA-binding domain near the N-terminus of the protein. Outside this domain however MYB protein sequences exhibit a low degree of sequence conservation, which leads to an overall lower sequence identity across homologs.
- a conserved functional domain with respect to presently disclosed polypeptides refers to a domain within a polypeptide family that exhibits a higher degree of sequence homology, such as at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to a conserved domain of a polypeptide of the invention (e.g., the SANT/Myb domain of SEQ ID NO:5). Sequences that possess or encode for conserved domains that meet these criteria of percentage identity, and that have comparable biological activity to the present polypeptide sequences, thus being members of the MYB transcription factor family, are encompassed by the invention.
- percent complementarity or “percent complementary,” as used herein in reference to two nucleotide sequences, is similar to the concept of percent identity but refers to the percentage of nucleotides of a query sequence that optimally base-pair or hybridize to nucleotides of a subject sequence when the query and subject sequences are linearly arranged and optimally base paired without secondary folding structures, such as loops, stems or hairpins. Such a percent complementarity may be between two DNA strands, two RNA strands, or a DNA strand and an RNA strand.
- the “percent complementarity” is calculated by (i) optimally base-pairing or hybridizing the two nucleotide sequences in a linear and fully extended arrangement (z'.e., without folding or secondary structures) over a window of comparison, (ii) determining the number of positions that base-pair between the two sequences over the window of comparison to yield the number of complementary positions, (iii) dividing the number of complementary positions by the total number of positions in the window of comparison, and (iv) multiplying this quotient by 100% to yield the percent complementarity of the two sequences.
- Optimal base pairing of two sequences may be determined based on the known pairings of nucleotide bases, such as G-C, A-T, and A-U, through hydrogen bonding.
- the percent identity is determined by dividing the number of complementary positions between the two linear sequences by the total length of the reference sequence.
- the “percent complementarity” for the query sequence is equal to the number of base-paired positions between the two sequences divided by the total number of positions in the query sequence over its length (or by the number of positions in the query sequence over a comparison window), which is then multiplied by 100%.
- a “fragment” of a polynucleotide refers to a sequence comprising at least about 50, at least about 75, at least about 95, at least about 100, at least about 125, at least about 150, at least about 175, at least about 200, at least about 225, at least about 250, at least about 275, at least about 300, at least about 500, at least about 600, at least about 700, at least about 750, at least about 800, at least about 900, or at least about 1000 contiguous nucleotides, or longer, of a DNA molecule or protein as disclosed herein. Methods for producing such fragments from a starting promoter molecule are well known in the art. Fragments of a DNA molecule or protein may exhibit the activity of the DNA molecule or protein from which they are derived.
- the present disclosure provides methods for altering a phenotype, such as inducing male sterility in a plant comprising:(a) modifying the genome of a plant cell by: (i) identifying an endogenous gene of the plant corresponding to a TDF1 gene, such as the GmTDFl-1 gene described herein, and its homologs, and (ii) modifying the sequence of the endogenous gene in the plant cell via targeted mutagenesis to modify the expression level of the endogenous gene; and (b) regenerating or developing a plant from the plant cell.
- a modifying the genome of a plant cell by: (i) identifying an endogenous gene of the plant corresponding to a TDF1 gene, such as the GmTDFl-1 gene described herein, and its homologs, and (ii) modifying the sequence of the endogenous gene in the plant cell via targeted mutagenesis to modify the expression level of the endogenous gene; and (b) regenerating or developing a plant from the plant cell.
- TDF1 genes and proteins from different plant species may be identified and considered TDF1 homologs or orthologs for use in the present disclosure if they have a similar nucleic acid and/or protein sequence and/or share conserved amino acids and/or structural domain(s) with at least one known TDF1 gene or protein.
- a plant selectable marker transgene in a transformation vector or construct of the present disclosure may be used to assist in the selection of transformed cells or tissue due to the presence of a selection agent, such as an antibiotic or herbicide, wherein the plant selectable marker transgene provides tolerance or resistance to the selection agent.
- a selection agent such as an antibiotic or herbicide
- the selection agent may bias or favor the survival, development, growth, proliferation, etc., of transformed cells expressing the plant selectable marker gene, such as to increase the proportion of transformed cells or tissues in the Ro plant.
- Commonly used plant selectable marker genes include, for example, those conferring tolerance or resistance to antibiotics, such as kanamycin and paromomycin (nptll), hygromycin B (aph IV), streptomycin or spectinomycin (aadA) and gentamycin (aac3 and aacC4), or those conferring tolerance or resistance to herbicides such as glufosinate (bar or pat), dicamba (DMO) and glyphosate (proA or EPSPS).
- antibiotics such as kanamycin and paromomycin (nptll), hygromycin B (aph IV), streptomycin or spectinomycin (aadA) and gentamycin (aac3 and aacC4)
- kanamycin and paromomycin kanamycin and paromomycin (nptll)
- aph IV hygromycin B
- streptomycin or spectinomycin a
- Plant screenable marker genes may also be used, which provide an ability to visually screen for transformants, such as luciferase or green fluorescent protein (GFP), or a gene expressing a beta glucuronidase or uidA gene (GUS) for which various chromogenic substrates are known. Plant transformation may also be carried out in the absence of selection during one or more steps or stages of culturing, developing or regenerating transformed explants, tissues, plants and/or plant parts.
- transformants such as luciferase or green fluorescent protein (GFP), or a gene expressing a beta glucuronidase or uidA gene (GUS) for which various chromogenic substrates are known.
- GFP green fluorescent protein
- GUS beta glucuronidase or uidA gene
- Methods and compositions are provided for transforming a plant cell, tissue or explant with a recombinant DNA molecule or construct encoding one or more molecules required for targeted genome editing (e.g., guide RNA(s) and/or site-directed nuclease(s)).
- Suitable methods for transformation of host plant cells include virtually any method by which DNA or RNA can be introduced into a cell (for example, where a recombinant DNA construct is stably integrated into a plant chromosome or where a recombinant DNA construct or an RNA is transiently provided to a plant cell) and are well known in the art.
- Two effective methods for cell transformation are bacterially-mediated transformation, such as Agro&acterzuzvr-mediated or R/zzzo&zwm-mediated transformation, and microprojectile or particle bombardment-mediated transformation.
- Microprojectile bombardment methods are illustrated, for example, in U.S. Patent Nos. 5,550,318; 5,538,880; 6,160,208; and 6,399,861.
- Agrobacterium-mcAa cd transformation methods are described, for example in U.S. Patent No. 5,591,616.
- Other methods for plant transformation such as microinjection, electroporation, vacuum infiltration, pressure, sonication, silicon carbide fiber agitation, PEG-mediated transformation, etc., are also known in the art.
- Transformation of plant material is practiced in tissue culture on nutrient media, for example a mixture of nutrients that allow cells to grow in vitro.
- Recipient cell targets include, but are not limited to, meristem cells, shoot tips, hypocotyls, calli, immature or mature embryos, and gametic cells such as microspores and pollen.
- Callus can be initiated from tissue sources including, but not limited to, immature or mature embryos, hypocotyls, seedling apical meristems, microsporcs and the like.
- Cells containing a transgenic nucleus arc grown into transgenic plants, also referred to as Ro plants.
- Ro plant refers to an initial regenerated transformant.
- Ri seed refers to seed produced from selfing Ro plants.
- Ri plant refers to a plant grown from Ri seed.
- R2 seed refers to seed produced from selfing Ri plants.
- R? plant refers to a plant grown from R2 seed.
- Any suitable method or technique for transformation of a plant cell known in the art may be used according to present methods.
- DNA is typically introduced into only a small percentage of target plant cells in any one transformation experiment.
- Marker genes are used to provide an efficient system for identification of those cells that are stably transformed by receiving and integrating a recombinant DNA molecule into their genomes.
- the terms “regeneration” and “regenerating” refer to a process of growing or developing a plant from one or more plant cells through one or more culturing steps. Transformed or edited cells, tissues or explants containing a DNA sequence insertion or edit may be grown, developed or regenerated into transgenic plants in culture, plugs, or soil according to methods known in the ail. Certain embodiments of the disclosure therefore relate to methods and constructs for regenerating a plant from a cell with modified genomic DNA resulting from genome editing. The regenerated plant can then be used to propagate additional plants.
- regenerated plants or a progeny plant, plant part or seed thereof can be screened or selected based on a marker, trait, or phenotype produced by the edit or mutation, or by the site-directed integration of an insertion sequence, transgene, etc., in the developed or regenerated plant, or a progeny plant, plant part or seed thereof. If a given mutation, edit, trait or phenotype is recessive, one or more generations or crosses (e.g., selfing) from the initial Ro plant may be necessary to produce a plant homozygous for the edit or mutation so the trait or phenotype can be observed.
- Progeny plants such as plants grown from Ri seed or in subsequent generations, can be tested for zygosity using any known zygosity assay, such as by using a single nucleotide polymorphism (SNP) assay, DNA sequencing, thermal amplification, or polymerase chain reaction (PCR), and/or Southern blotting that allows for the distinction between heterozygote, homozygote and wild-type plants.
- SNP single nucleotide polymorphism
- PCR polymerase chain reaction
- Methods and techniques are provided for screening for, and/or identifying, cells or plants, etc., for the presence of targeted edits or transgcncs, and selecting cells or plants comprising targeted edits or transgenes, which may be based on one or more phenotypes or traits, or on the presence or absence of a molecular marker or polynucleotide or protein sequence in the cells or plants.
- a “molecular technique” refers to any method known in the fields of molecular biology, biochemistry, genetics, plant biology, or biophysics that involves the use, manipulation, or analysis of a nucleic acid, a protein, or a lipid.
- molecular techniques useful for detecting the presence of a modified sequence in a genome include phenotypic screening; molecular marker technologies such as SNP analysis by TaqMan® or Illumina/Infinium technology; Southern blot; PCR (including amplicon sequencing which consists of the generation of one or more unique PCR products across the genomic region of interest for further sequencing analysis, such as using Next-Gen Sequencing techniques known in the art. Sequence data from each sample is then mapped to a reference sequence to identify consensus differences); enzyme-linked immunosorbent assay (ELISA); and sequencing (e.g., Sanger, Illumina®, 454, Pae-Bio, Ion TorrentTM).
- a method of detection provided herein comprises phenotypic screening.
- a method of detection provided herein comprises SNP analysis. In a further aspect, a method of detection provided herein comprises a Southern blot. In a further aspect, a method of detection provided herein comprises PCR. In a further aspect, a method of detection provided herein comprises amplicon sequencing. In an aspect, a method of detection provided herein comprises ELISA. In a further aspect, a method of detection provided herein comprises determining the sequence of a nucleic acid or a protein. Without being limiting, nucleic acids can be detected using hybridization. Hybridization between nucleic acids is discussed in detail in Sambrook et al. (1989, Molecular Cloning: A Laboratory Manual, 2nd Ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY).
- nucleic acids can be isolated using techniques routine in the ail.
- nucleic acids can be isolated using any method including, without limitation, recombinant nucleic acid technology, and/or PCR.
- General PCR techniques are described, for example in PCR Primer: A Laboratory Manual, Dieffenbach & Dveksler, Eds., Cold Spring Harbor Laboratory Press, 1995.
- Recombinant nucleic acid techniques include, for example, restriction enzyme digestion and ligation, which can be used to isolate a nucleic acid.
- Isolated nucleic acids also can be chemically synthesized, either as a single nucleic acid molecule or as a series of oligonucleotides.
- Detection can be accomplished using detectable labels that may be attached or associated with a hybridization probe or antibody.
- label is intended to encompass the use of direct labels as well as indirect labels.
- Detectable labels include enzymes, prosthetic groups, fluorescent materials, luminescent materials, bioluminescent materials, and radioactive materials.
- the screening and selection of modified (e.g., edited) plants or plant cells can be through any methodologies known to those skilled in the art of molecular biology.
- screening and selection methodologies include, but are not limited to, Southern analysis, PCR amplification for detection of a polynucleotide (including amplicon sequencing), Northern blots, RNase protection, primerextension, RT-PCR amplification for detecting RNA transcripts, Sanger sequencing, Next Generation sequencing technologies (e.g., Illumina®, PacBio®, Ion TorrentTM, etc.) enzymatic assays for detecting enzyme or ribozyme activity of polypeptides and polynucleotides, and protein gel electrophoresis, Western blots, immunoprecipitation, and enzyme-linked immunoassays to detect polypeptides.
- Other techniques such as in situ hybridization, enzyme staining, and immuno staining also can be used to detect the presence or expression of polypeptides and/or polynucleotides. Methods for performing all of the referenced techniques are known in the art.
- polypeptide refers to a chain of at least two covalently linked amino acids.
- Polypeptides can be encoded by polynucleotides provided herein.
- An example of a polypeptide is a protein.
- Proteins provided herein can be encoded by nucleic acid molecules provided herein.
- Polypeptides can be purified from natural sources (e.g., a biological sample) by known methods such as DEAE ion exchange, gel filtration, and hydroxyapatite chromatography.
- a polypeptide also can be purified, for example, by expressing a nucleic acid in an expression vector.
- a purified polypeptide can be obtained by chemical synthesis. The extent of purity of a polypeptide can be measured using any appropriate method, e.g., column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis.
- Polypeptides can be detected using antibodies. Techniques for detecting polypeptides using antibodies include enzyme linked immunosorbent assays (ELISAs), Western blots, immunoprecipitations and immunofluorescence.
- An antibody provided herein can be a polyclonal antibody or a monoclonal antibody.
- An antibody having specific binding affinity for a polypeptide provided herein can be generated using methods well known in the art.
- An antibody provided herein can be attached to a solid support such as a microtiter plate using methods known in the art.
- a plant that may be transformed with a recombinant DNA molecule or transformation vector comprising a guide RNA may include a variety of flowering plants or angiosperms, which may be further defined as including various dicotyledonous (dicot) plant species or monocotyledonous (monocot) plant species.
- a dicot plant could be members of the Fabaceae family (such as legumes), sunflower (Helianthus annuus), safflower (Carthamus tinctorius), sesame (Sesamum spp.), tobacco (Nicotiana tabacum), potato (Solarium tuberosum), cotton (Gossypium barbadense, Gossypium hirsutum), sweet potato (Ipomoea batatas), cassava (Manihot esculenta), coffee (Coffea spp.), tea (Camellia spp.), fruit trees, such as apple (Malus spp.), Primus spp., such as plum, apricot, peach, cherry, etc., pear (Pyrus spp.), fig (Ficus carica), etc., citrus trees (Citrus spp.), cocoa (Theobroma cacao), avocado (Persea americana), olive (Olea europaea),
- Legumes and leguminous plants include peas (Pisum sativum) alfalfa (Medicago sativa), barrel clover (Medicago truncatula), pigeon pea (Cajanus cajan) guar (Cyamopsis tetragonoloba), carob (Ceratonia siliqua), fenugreek (Trigonella foenum-graecum), soybean (Glycine max), common bean (Phaseolus vulgaris), cowpea (Vigna unguiculata), mung bean (Vigna radiata), lima bean (Phaseolus lunatus), fava bean (Vicia faba), lentil (Lens culinaris or Lens esculenta), peanut (Arachis hypogaea), licorice (Glycyrrh
- a monocot plant could be oil palm (Elaeis spp.), coconut (Cocos spp.), banana (Musa spp.), and cereals such as corn (Zea mays), barley (Hordeum vulgare), sorghum (Sorghum bicolor), rice (Oryza sativa), and wheat (Triticum aestivum).
- the present disclosure further applies to other botanical structures analogous to pods of leguminous plants, such as bolls, siliques, fruits, nuts, tubers, etc.
- modified in the context of a plant, plant seed, plant part, plant cell, and/or plant genome, refers to a plant, plant seed, plant part, plant cell, and/or plant genome comprising an engineered change in the expression level and/or endogenous sequence of one or more genes of interest relative to a wild-type or control plant, plant seed, plant part, plant cell, and/or plant genome.
- modified may further refer to a plant, plant seed, plant part, plant cell, and/or plant genome having one or more deletions affecting expression of an endogenous TDF1 gene or homolog introduced through chemical mutagenesis, transposon insertion or excision, or any other known mutagenesis technique, or introduced through genome editing.
- a modified plant, plant seed, plant part, plant cell, and/or plant genome can comprise one or more transgenes.
- a modified plant, plant seed, plant part, plant cell, and/or plant genome includes a mutated, edited and/or transgenic plant, plant seed, plant part, plant cell, and/or plant genome having a modified expression level, expression pattern, and/or sequence of a TDF1 gene or homolog relative to a wild-type or control plant, plant seed, plant part, plant cell, and/or plant genome.
- Modified plants, plant parts, seeds, etc. may have been subjected to mutagenesis, genome editing or site-directed integration, genetic transformation, or a combination thereof.
- Such “modified” plants, plant seeds, plant parts, and plant cells include plants, plant seeds, plant parts, and plant cells that are offspring or derived from “modified” plants, plant seeds, plant parts, and plant cells that retain the molecular change (e.g., change in expression level and/or activity) to the endogenous TDF1 gene, such as the GmTDFl-1 gene.
- a modified seed provided herein may give rise to a modified plant provided herein.
- a modified plant, plant seed, plant part, plant cell, or plant genome provided herein may comprise a recombinant DNA construct or vector or genome edit as provided herein.
- a “modified plant product” may be any product made from a modified plant, plant part, plant cell, or plant chromosome provided herein, or any portion or component thereof.
- Modified plants may be further crossed to themselves or other plants to produce modified plant seeds and progeny.
- a modified plant may also be prepared by crossing a first plant comprising a DNA sequence or construct or an edit (e.g., a genomic deletion) with a second plant lacking the DNA sequence or construct or edit.
- a DNA sequence or inversion may be introduced into a first plant line that is amenable to transformation or editing, which may then be crossed with a second plant line to introgress the DNA sequence or edit (e.g., deletion) into the second plant line.
- Progeny of these crosses can be further backcrossed into the desirable line multiple times, such as through 6 to 8 generations or back crosses, to produce a progeny plant with substantially the same genotype as the original parental line, but for the introduction of the DNA sequence or edit.
- a modified plant, plant cell, or seed provided herein may be a hybrid plant, plant cell, or seed.
- a “hybrid” is created by crossing two plants from different varieties, lines, inbrcds, or species, such that the progeny comprises genetic material from each parent. Skilled artisans recognize that higher order hybrids can be generated as well.
- a modified plant, plant part, plant cell, or seed provided herein may be of an elite variety or an elite line.
- An “elite variety” or an “elite line” refers to a variety that has resulted from breeding and selection for superior agronomic performance.
- control plant refers to a plant (or plant seed, plant part, plant cell, and/or plant genome) that is used for comparison to a modified plant (or modified plant seed, plant part, plant cell, and/or plant genome) and has the same or similar genetic background (e.g., same parental lines, hybrid cross, inbred line, testers, etc.) as the modified plant (or plant seed, plant part, plant cell, and/or plant genome), except for genome edit(s) (e.g., a deletion) affecting a TDF1 gene.
- genetic background e.g., same parental lines, hybrid cross, inbred line, testers, etc.
- a control plant may be an inbred line that is the same as the inbred line used to make the modified plant, or a control plant may be the product of the same hybrid cross of inbred parental lines as the modified plant, except for the absence in the control plant of any transgenic events or genome edit(s) affecting a TDF1 gene.
- an “unmodified control plant” refers to a plant that shares a substantially similar or essentially identical genetic background as a modified plant, but without the one or more engineered changes to the genome (e.g., mutation or edit) of the modified plant.
- a “wild-type plant” refers to a non-transgenic and non-genome edited control plant, plant seed, plant part, plant cell, and/or plant genome.
- a “control” plant, plant seed, plant part, plant cell, and/or plant genome may also be a plant, plant seed, plant part, plant cell, and/or plant genome having a similar (but not the same or identical) genetic background to a modified plant, plant seed, plant part, plant cell, and/or plant genome, if deemed sufficiently similar’ for comparison of the characteristics or traits to be analyzed.
- the terms “suppress,” “suppression,” “inhibit,” “inhibition,” “inhibiting,” “knockout,” “knockdown,” and “downregulation” refer to a lowering, reduction, or elimination of the expression level of an mRNA and/or protein encoded by a target gene in a plant, plant cell, or plant tissue at one or more stage(s) of plant development, as compared to the expression level of such target mRNA and/or protein in a wild-type or control plant, cell, or tissue at the same stage(s) of plant development.
- a modified plant having a GmTDFl-1 gene expression level that is reduced in at least one plant tissue by at least 5%, at least 10%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, or 100%, as compared to a control plant.
- a modified plant having a GmTDFl-1 gene expression level that is reduced in at least one plant tissue by 5%-20%, 5%-25%, 5%-30%, 5%-40%, 5%-50%, 5%- 60%, 5%-70%, 5%-75%, 5%-80%, 5%-90%, 5%-100%, 75%-100%, 50%-100%, 50%-90%, 50%- 75%, 25%- 75%, 30%-80%, or 10%-75%, as compared to a control plant.
- a modified plant having a GmTDFl-1 mRNA level that is reduced in at least one plant tissue by at least 5%, at least 10%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, or 100%, as compared to a control plant.
- a modified plant having a GmTDFl-1 mRNA expression level that is reduced in at least one plant tissue by 5%-20%, 5%-25%, 5%- 30%, 5%-40%, 5%-50%, 5%-60%, 5%-70%, 5%- 75%, 5%-80%, 5%-90%, 5%-100%, 75%-100%, 50%-100%, 50%-90%, 50%-75%, 25%-75%, 30%-80%, or 10%-75%, as compared to a control plant.
- a modified plant having a GmTDFl-1 protein expression level that is reduced in at least one plant tissue by at least 5%, at least 10%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, or 100%, as compared to a control plant.
- a modified plant having a GmTDFl-1 protein expression level that is reduced in at least one plant tissue by 5%-20%, 5%- 25%, 5%-30%, 5%-40%, 5%-50%, 5%-60%, 5%-70%, 5%-75%, 5%-8O%, 5%-90%, 5%-100%, 75%-100%, 50%-100%, 50%-90%, 50%-75%, 25%-75%, 30%-80%, or 10%-75%, as compared to a control plant.
- Modified plants comprising or derived from plant cells that are transformed with a recombinant DNA of this disclosure can be further enhanced with stacked traits, for example, a modified crop plant having an enhanced trait resulting from expression of DNA disclosed herein in combination with one or more genes of agronomic interest that provide a beneficial agronomic trait (such as herbicide and/or pest resistance traits) to crop plants.
- a beneficial agronomic trait such as herbicide and/or pest resistance traits
- the traits conferred by the recombinant DNA constructs of the current disclosure can be stacked with other traits of agronomic interest, such as a trait providing insect resistance such as using a gene from Bacillus thuringensis to provide resistance against lepidoptcran, colcoptcran, homoptcran, hemiopteran, and other insects, or improved quality traits such as improved nutritional value.
- a trait providing insect resistance such as using a gene from Bacillus thuringensis to provide resistance against lepidoptcran, colcoptcran, homoptcran, hemiopteran, and other insects
- improved quality traits such as improved nutritional value.
- Molecules and methods for imparting insect/nematode/virus resistance are disclosed in U.S. Patent Nos. 5,250,515; 5,880,275; 6,506,599; 5,986,175; and U.S. Patent Application Publication No. 2003/0150017 Al.
- Herbicides for which transgenic plant tolerance has been demonstrated and the methods and compositions of the present disclosure can be applied include, but are not limited to, glyphosate, dicamba, glufosinate, sulfonylurea, bromoxynil, norflurazon, 2,4-D (2,4- dichlorophenoxy) acetic acid, aryloxyphenoxy propionates, p-hydroxyphenyl pyruvate dioxygenase inhibitors (HPPD), and protoporphyrinogen oxidase inhibitors (PPO) herbicides.
- Polynucleotide molecules encoding proteins involved in herbicide tolerance include, but are not limited to, a polynucleotide molecule encoding 5-enolpyruvylshikimate-3- phosphate synthase (EPSPS) disclosed in U.S. Patent Nos. 5,094,945; 5,627,061; 5,633,435 and 6,040,497 for imparting glyphosate tolerance; polynucleotide molecules encoding a glyphosate oxidoreductase (GOX) disclosed in U.S. Patent No. 5,463,175 and a glyphosate-N-acetyl transferase (GAT) disclosed in U.S. Patent No.
- EPSPS 5-enolpyruvylshikimate-3- phosphate synthase
- AHAS acetohydroxyacid synthase
- bar genes disclosed in DeBlock et al. (EMBO J. 6:2513-2519, 1987) for imparting glufosinate and bialaphos tolerance
- Patent Application Publication 2003/010609 Al for imparting N-amino methyl phosphonic acid tolerance
- polynucleotide molecules disclosed in U.S. Patent No. 6,107,549 for imparting pyridine herbicide resistance
- molecules and methods for imparting tolerance to multiple herbicides such as glyphosate, atrazine, ALS inhibitors, isoxoflutole and glufosinate herbicides are disclosed in U.S. Patent No. 6,376,754 and U.S. Patent Application Publication 2002/0112260.
- any method that “comprises,” “has” or “includes” one or more steps is not limited to possessing only those one or more steps and can also cover other unlisted steps.
- any composition or device that “comprises,” “has,” or “includes” one or more features is not limited to possessing only those one or more features and can cover other unlisted features.
- a “plant” includes a whole plant, explant, plant part, seedling, or plantlet at any stage of regeneration or development.
- a “plant part” can refer to any organ or intact tissue of a plant, such as a meristem, shoot organ/structure (e.g., leaf, stem or node), root, flower or floral organ/structure (e.g., bract, sepal, petal, stamen, carpel, anther and ovule), seed, embryo, endosperm, seed coat, fruit, the mature ovary, propagule, or other plant tissues (e.g., vascular tissue, dermal tissue, ground tissue, and the like), or any portion thereof.
- Plant parts of the present disclosure can be viable, nonviable, regenerable, and/or non-regenerable.
- a “propagule” can include any plant part that can grow into an entire plant.
- An “embryo” is a part of a plant seed, consisting of precursor tissues (e.g., meristematic tissue) that can develop into all or part of an adult plant.
- An “embryo” may further include a portion of a plant embryo.
- a “meristem” or “meristematic tissue” comprises undifferentiated cells or meristematic cells, which are able to differentiate to produce one or more types of plant parts, tissues or structures, such as all or part of a shoot, stem, root, leaf, seed, etc.
- the “vegetative phase” of plant development is the period of growth between germination and flowering.
- the stages in the vegetative phase of soybean are as follows: VE (emergence), VC (cotyledon stage), V 1 (first trifoliolate leaf), V2 (second trifoliolate leaf), V3 (third trifoliolate leaf), V(n) (nth trifoliolate leaf), and V6 (flowering will soon start).
- the “reproductive phase” of plant development is the period between flowering and the end of harvest.
- the stages in the reproductive phase of soybean are as follows R1 (beginning bloom, first flower); R2 (full bloom, flower in top 2 nodes); R3 (beginning pod, 3/16" pod in top 4 nodes); R4 (full pod, 3/4" pod in top 4 nodes); R5 (1/8" seed in top 4 nodes); R6 (full size seed in top 4 nodes); R7 (beginning maturity, one mature pod); and, R8 (full maturity, 95% of pods on the plant have reached mature color).
- Soybean vegetative and reproductive stages are well known to those of skill in the art and numerous publications describing these stages can be found on the world wide web and elsewhere, such as North Dakota State University publication A- 1174, June 1999, Reviewed and Reprinted August 2004.
- Embodiment l is a modified plant, plant seed, plant part, or plant cell comprising a genomic modification that reduces expression or activity of TDF1, or a homolog thereof, as compared to the expression or activity of TDF1 or a homolog thereof in an otherwise identical plant, plant seed, plant part, or plant cell that lacks the modification.
- Embodiment 2 is the modified plant, plant seed, plant part, or plant cell of embodiment 1 , wherein the modification is present in at least one allele of an endogenous TDF1 gene or homolog thereof.
- Embodiment 3 is the modified plant, plant seed, plant part, or plant cell of embodiment 2, wherein the TDF1 gene or homolog thereof encodes a protein comprising a conserved SANT/Myb domain having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the conserved SANT/Myb domain corresponding to position 9 to position 112 of SEQ ID NO:5.
- Embodiment 4 is the modified plant, plant seed, plant part, or plant cell of embodiment 2 or 3, wherein the modification disrupts the function of a wild-type allele product of the endogenous TDF1 gene or homolog thereof.
- Embodiment 5 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 2-4, wherein the modification is located within a coding region, non-coding region, or any combination thereof, of the endogenous TDF1 gene or homolog thereof.
- Embodiment 6 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 2-5, wherein the modification is located in an exon region of said TDF1 gene or homolog thereof.
- Embodiment 7 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 2-5, wherein the modification is located in an intron region of said TDF1 gene or homolog thereof.
- Embodiment 8 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 2-5, wherein the modification is located in an exon region and an intron region of said TDF1 gene or homolog thereof.
- Embodiment 9 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 1-8, wherein the plant is a leguminous plant, or wherein the plant seed, plant part, or plant cell is a plant seed, plant part, or plant cell of a leguminous plant.
- Embodiment 10 is the modified plant, plant seed, plant part, or plant cell of embodiment 9, wherein the leguminous plant is a soybean plant, a bean plant, a pea plant, a chickpea plant, an alfalfa plant, a peanut plant, a carob plant, a lentil plant, or a licorice plant.
- the leguminous plant is a soybean plant, a bean plant, a pea plant, a chickpea plant, an alfalfa plant, a peanut plant, a carob plant, a lentil plant, or a licorice plant.
- Embodiment 11 is the modified plant, plant seed, plant part, or plant cell of embodiment
- leguminous plant is a soybean plant.
- Embodiment 12 is the modified plant, plant seed, plant part, or plant cell of embodiment
- TDF1 gene is GmTDFl-1.
- Embodiment 13 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 1-12, wherein the plant, plant seed, plant part, or plant cell is heterozygous for the modification.
- Embodiment 14 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 1-12, wherein the plant, plant seed, plant part, or plant cell is homozygous for the modification.
- Embodiment 15 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 1-10, wherein the plant, plant seed, plant part, or plant cell comprises a first modification in a first allele of the TDF1 gene and a second modification in a second allele of the TDF1 gene, the first modification and the second modification being different from one another.
- Embodiment 16 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 1-15, wherein the modification comprises a deletion, an insertion, a substitution, an inversion, a duplication, or a combination of any thereof.
- Embodiment 17 is the modified plant, plant seed, plant part, or plant cell of embodiment
- Embodiment 18 is the modified plant, plant seed, plant part, or plant cell of embodiment
- deletion comprises between 1 nucleotide and 1500 nucleotides.
- Embodiment 19 is the modified plant, plant seed, plant part, or plant cell of embodiment
- deletion comprises a region of at least 1 nucleotide, at least 2 nucleotides, at least 5 nucleotides, at least 10 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, at least 500 nucleotides, at least 600 nucleotides, at least 700 nucleotides, at least 800 nucleotides, at least 900 nucleotides, at least 1000 nucleotides, or at least 1500 nucleotides.
- Embodiment 20 is the modified plant, plant seed, plant part, or plant cell of embodiment 16, wherein the modification is an inversion.
- Embodiment 21 is the modified plant, plant seed, plant part, or plant cell of embodiment 20, wherein the inversion is a region of at least 5 consecutive nucleotides, at least 10 consecutive nucleotides, at least 20 consecutive nucleotides, at least 30 consecutive nucleotides, at least 40 consecutive nucleotides, at least 50 consecutive nucleotides, at least 100 consecutive nucleotides, at least 200 consecutive nucleotides, at least 300 consecutive nucleotides, at least 400 consecutive nucleotides, at least 500 consecutive nucleotides, at least 600 consecutive nucleotides, at least 700 consecutive nucleotides, at least 800 consecutive nucleotides, at least 900 consecutive nucleotides, at least 1000 consecutive nucleotides, or at least 1100 consecutive nucleotides.
- Embodiment 22 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 12-21, wherein the plant, plant seed, plant part, or plant cell comprises a modification in at least one allele of the GmTDFl-1 gene, wherein the modification is selected from the group consisting of: an 8 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:26; a 10 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:27; a 1 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:28; a 6 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:29; a 9 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:30; an 8 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:31; an 11 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:26;
- Embodiment 23 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 12-22, wherein the plant, plant seed, plant part, or plant cell comprises at least one polynucleotide sequence selected from the group consisting of SEQ ID NOs:26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, and 53.
- Embodiment 24 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 12-23, wherein the plant, plant seed, plant part, or plant cell comprises at least one polynucleotide sequence selected from the group consisting of SEQ ID NOs:26, 27, 28, 29, 30, 31, 32, and 33.
- Embodiment 25 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 12-23, wherein the plant, plant seed, plant part, or plant cell comprises at least one polynucleotide sequence selected from the group consisting of SEQ ID NOs:34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, and 53.
- Embodiment 26 is the modified plant of any one of embodiments 1-25, wherein the modification results in a male sterility phenotype in the plant, as compared to the male fertile phenotype of an otherwise identical plant that lacks the modification.
- Embodiment 27 is a polynucleotide comprising a sequence selected from the group consisting of SEQ ID NOs:26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, and 53.
- Embodiment 28 is a guide RNA comprising a polynucleotide sequence selected from the group consisting of SEQ ID NOs:23, 24, and 25.
- Embodiment 29 is the guide RNA of embodiment 28, wherein the polynucleotide sequence is a spacer sequence.
- Embodiment 30 is a method for producing a modified plant having a male sterility phenotype, the method comprising: a) introducing a modification into at least one target site in an endogenous TDF1 gene or a homolog thereof of a plant cell; b) identifying and selecting one or more plant cells of step a) comprising said modification in said TDF1 gene or homolog thereof; and c) regenerating at least one plant from at least one or more cells selected in step b).
- Embodiment 31 is the method of embodiment 30, wherein the target site is located in a coding region, a non-coding region, or any combination thereof, of an endogenous TDF1 gene or homolog thereof.
- Embodiment 32 is the method of embodiment 30, wherein the modification is introduced by a site-specific genome modification enzyme selected from the group consisting of: an RNA- guided nuclease, a zinc-finger nuclease, a meganuclease, a TALE-nuclease, a recombinase, a transposase, and combinations of any thereof.
- a site-specific genome modification enzyme selected from the group consisting of: an RNA- guided nuclease, a zinc-finger nuclease, a meganuclease, a TALE-nuclease, a recombinase, a transposase, and combinations of any thereof.
- Embodiment 33 is the method of embodiment 32, wherein the site-specific genome modification enzyme is an RNA-guided nuclease comprising a Cas nuclease, a Cpfl nuclease, or a variant of either thereof.
- the site-specific genome modification enzyme is an RNA-guided nuclease comprising a Cas nuclease, a Cpfl nuclease, or a variant of either thereof.
- Embodiment 34 is the method of embodiment 33, wherein the site-specific genome modification enzyme is an RNA-guided nuclease comprising a Cpfl nuclease.
- Embodiment 35 is the method of embodiment 30, wherein the modification is introduced by a site-specific genome modification enzyme that creates at least one strand break at the target site.
- Embodiment 36 is the method of embodiment 30, wherein the modification is selected from the group consisting of a substitution, an insertion, an inversion, a deletion, a duplication, and a combination thereof.
- Embodiment 37 is the method of embodiment 36, wherein the modification is a deletion.
- Embodiment 38 is the method of embodiment 37, wherein the deletion comprises a region of at least 1 nucleotide, at least 2 nucleotides, at least 5 nucleotides, at least 10 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, at least 500 nucleotides, at least 600 nucleotides, at least 700 nucleotides, at least 800 nucleotides, at least 900 nucleotides, or at least 1000 nucleotides, or at least 1500 nucleotides.
- Embodiment 39 is the method of embodiment 36, wherein the modification is an inversion.
- Embodiment 40 is the method of embodiment 39, wherein the inversion is a region of at least 5 consecutive nucleotides, at least 10 consecutive nucleotides, at least 20 consecutive nucleotides, at least 30 consecutive nucleotides, at least 40 consecutive nucleotides, at least 50 consecutive nucleotides, at least 100 consecutive nucleotides, at least 200 consecutive nucleotides, at least 300 consecutive nucleotides, at least 400 consecutive nucleotides, at least 500 consecutive nucleotides, at least 600 consecutive nucleotides, at least 700 consecutive nucleotides, at least 800 consecutive nucleotides, at least 900 consecutive nucleotides, at least 1000 consecutive nucleotides, or at least 1100 consecutive nucleotides.
- Embodiment 41 is a male sterile soybean plant produced by any of the methods of embodiments 30-40.
- Embodiment 42 is a method for producing an Fi hybrid soybean plant comprising crossing a male-sterile soybean plant comprising a modified GmTDFl-1 gene with a second, non-isogenic, male-fertile soybean plant to produce an Fi hybrid soybean plant.
- Embodiment 43 is an Fi hybrid soybean plant produced by the method of embodiment 42.
- Embodiment 44 is a method for producing an Fi hybrid soybean seed, comprising crossing the plant of embodiment 1 with a second, non-isogenic, male-fertile soybean plant and harvesting the resultant Fi hybrid soybean seed, wherein an Fi hybrid soybean seed is produced.
- Embodiment 45 is an Fi hybrid soybean seed produced by the method of embodiment 44.
- Embodiment 46 is a modified plant, plant seed, plant part, or plant cell, wherein the plant, plant seed, plant part, or plant cell comprises at least one polynucleotide sequence selected from the group consisting of SEQ ID NOs:26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, and 53.
- Embodiment 47 is the modified plant of embodiment 12, wherein the modification results in a male sterility phenotype in the plant, as compared to the male fertile phenotype of an otherwise identical plant that lacks the modification.
- Embodiment 48 is the modified plant of embodiment 23, wherein the modification results in a male sterility phenotype in the plant, as compared to the male fertile phenotype of an otherwise identical plant that lacks the modification.
- Embodiment 49 is the method of any one of embodiments 30-40, wherein the plant cell of step a) is a soybean plant cell and wherein the produced modified plant having a male sterility phenotype is a soybean plant.
- Embodiment 50 is a method for producing an Fi hybrid soybean seed, comprising crossing the plant of any one of embodiments 1-26 with a second, non-isogenic, male-fertile soybean plant and harvesting the resultant Fi hybrid soybean seed, wherein an Fi hybrid soybean seed is produced.
- Embodiment 51 is an Fi hybrid soybean seed produced by the method of embodiment 50.
- Embodiment 52 is the modified plant, plant seed, plant part, or plant cell of embodiment 46, wherein the plant, plant seed, plant part, or plant cell is a transgenic soybean plant, transgenic soybean plant seed, transgenic soybean plant part, or transgenic soybean plant cell.
- Embodiment 52 is the modified plant, plant seed, plant part, or plant cell of embodiment 12, wherein the plant, plant seed, plant part, or plant cell comprises a first modification in a first allele of the TDF1-1 gene and a second modification in a second allele of the TDF1-1 gene, the first modification and the second modification being different from one another.
- Example 1 Identification of the Candidate Gene Associated with Male Sterility.
- the Ms6 locus in soybean has been identified as being associated with male fertility in soybean.
- a naturally occurring male sterility allele in soybean is caused by the recessive ms6 allele.
- This Ms6 locus is publicly documented in the literature to reside approximately 10 centimorgan (cM) from the flower color locus (Wl) on soybean chromosome 13.
- cM centimorgan
- Wl flower color locus
- an ms6 mutant soybean line was whole genome sequenced and mapped using fine mapping techniques to improve markers for quality seed identification as well as to identify a high- confidence gene candidate for the Ms6 locus.
- the approximate gene locus region of the ms6 trait on chromosome 13 was narrowed via joint linkage mapping.
- Joint linkage mapping was performed for the ms6 mutant line with 5 populations heterozygous for the flower color marker using 27 markers between 15 cM to 44.8 cM on chromosome 13 to identify the most polymorphic markers among the populations.
- This joint linkage mapping narrowed the region coding for the ms6 trait to be between 26 cM and 32 cM on chromosome 13 of the soybean cultivar Williams 82 (Wm82) germplasm.
- eSNPl was used as the candidate eSNP associated with the ms6 allele.
- the variation at eSNPl is a T->A nucleotide change at 28 cM (14,995,614 bp) on chromosome 13 (Table 1).
- Marker eSNPl is associated with the putative region comprising the Ms6 locus on chromosome 13.
- Identification of the candidate gene could allow for gene editing approaches to reproduce the male sterile phenotype and male sterile lines.
- mapping data generated above Using the mapping data generated above, a bioinformatics approach was used to identify the gene associated with eSNPl locus.
- the T- A mutation identified at eSNPl was found in a genic region containing the A7v/>- fa ily of genes.
- This gene (Glymal3.g066600) has high functional domain similarity to the Tapetum Development & Function 1 (AtTDFF) gene in Arabidopsis thaliana.
- This gene which is also known as AtMYB35 (SEQ ID NO:1), encodes a member of the R2R3 MYB transcription factor superfamily that is essential for early tapetum development.
- the AtTDFl protein (SEQ ID NO:2) is understood to play a role in male fertility.
- the candidate gene mutated in the ms6 line is the soybean homolog of AtTDFl and is designated herein as GmTDFl-1.
- the R2R3-MYB protein consists of two major functional parts: a DNA-binding domain (MYB domain) located at the N- terminus and a regulatory region (non-MYB region) located at the C-terminus.
- MYB domain DNA-binding domain
- non-MYB region regulatory region located at the C-terminus.
- the MYB domain is a signature, highly conserved feature in the gene family, whereas the non-MYB regions have diverged among plant species.
- Glycine max TDF1-1 plays an important role in plant growth and development.
- the soybean GmTDFl-1 gene is 2,094 bp in length and comprises three exons and two introns.
- the coding and genomic sequences of the GmTDFl-1 gene are set forth in SEQ ID NOs;3 and 4, respectively.
- This gene has two in-frame translational start codons that encode two potential proteins, a longer GmTDFl-1 protein that is 375 amino acids in length and a shorter GmTDFl-1 protein that is 345 amino acids in length, alternatively translated from the second start codon of the mRNA.
- the first in-frame transcriptional start codon is at nucleotide position 18, and the second in-frame transcriptional start codon is at nucleotide position 108 in SEQ ID NO:4, respectively.
- the protein sequence of the shorter GmTDFl-1 protein is set forth in SEQ ID NO:5.
- the additional 30 amino acids in the extended C-terminal of the longer GmTDFl-1 protein sequence (not shown) identifies only soybean consensus sequences in the GenBank database and suggests that this additional sequence is either soybean-specific, or not translated. Both the shorter and longer GmTDFl-1 sequences and annotations can be found in the SoyBase database (www.soybase.org) depending on the reference genome used.
- Bioinformatic analysis and annotation of the ms6 line utilized for the experiments in these Examples supports use of the shorter GmTDFl-1 protein sequence as the reference and, accordingly, use of the second in-frame start codon of the GmTDFl-1 nucleotide sequence as the start codon. Accordingly, the 90 nucleotides from the first in-frame start codon until the second in-frame start codon are counted as a pail of the 5’ UTR of the GmTDFl-1 gene.
- the CDS utilized as SEQ ID NO:3 represents the shorter version of the nucleotide sequence based on the second in-frame transcriptional start codon.
- the eSNPl was identified at position 347 in SEQ ID NO:4 and occurs at the 5’ end of Exon 2, which spans nucleotides 344 to 473 in SEQ ID NO:4. This is a missense mutation which changes the amino acid residue from leucine to histidine in the GmTDFl-1 polypeptide sequence (SEQ ID NO:5). This missense mutation is associated with the ms6 phenotype (FIGS. 1A and IB).
- a marker was converted from eSNPl and analyzed for high association to the male sterility trait and genetic accuracy of the marker was confirmed. Assay probes and primers were developed and customized for accurate genetic identification of seeds for hybrid production (SEQ ID NOs:6- 12 and 14).
- the two gene editing constructs each contained two functional regions or cassettes relevant to gene editing and the creation of the DSBs in the edited GmTDFl-1 gene region: expression of a Cpfl protein and expression of one or two guide RNAs targeting the GmTDFl-1 gene.
- Each gRNA unit contains a common scaffold compatible with the Cpfl gene (SEQ ID NO: 18), and a unique spacer/targeting sequence complementary to an intended target site as listed in Table 3.
- the coordinates provided in Table 3 are indicative of the nucleotide position(s) in the sequence of SEQ ID NO:4 (z.e., the genomic GmTDFl-1 sequence), in the 5’ to 3’ direction.
- Both guide RNA spacers of construct pM765 are oriented in the antisense direction and thus, the reverse complement sequences of the spacer sequences match the target site given in Table 3. Table 3.
- the Cpfl expression cassette of both editing constructs comprised a Dahlia mosaic virus FLT promoter (SEQ ID NO: 19) operably linked to a sequence encoding a Lachnospiraceae bacterium Cpfl RNA-guided endonuclease enzyme (SEQ ID NO:20) that was codon-optimized for expression in plants, flanked on each side by one copy of a nuclear localization signal (SEQ ID NO:21). See, e.g., Gao et al. (Nature Biotechnol. 35(8):789-792, 2017).
- One type of gRNA expression cassette present in construct pM764, comprised a sequence encoding one GmTDFl-1 gRNA (guide RNA spacer gRNA_+221C) operably linked to a soybean RNA polymerase III (Pol3) promoter (SEQ ID NO:22).
- the other gRNA expression cassette present in construct pM765 (guide RNA spacers gRNA_-103C and gRNA_- 1113G), comprised a sequence encoding two GmTDFl -1 guide RNAs operably linked to the same soybean RNA polymerase III (Pol3) promoter (SEQ ID NO:22).
- the sequences of SEQ ID NOs:24 and 25, listed in Table 3, are spacer sequences that target alternative DSB sites, one near the 3’ end of Exon 1 and one near the 5’ end of Exon 3 in GmTDFl-1, aiming to result in deletion of most of the coding sequence of the GmTDFl-1 gene and ensure loss of function.
- the target sites of the three gRNAs are visualized on the GmTDFl-1 gene.
- the gRNAs were designed to avoid targeting GmTDFl-2, a close paralog of GmTDFl-1 on chromosome 19.
- Example 3 Generation of Novel GmTDFl-1 Alleles Produced Through Gene Editing.
- An inbred wild-type soybean line was transformed using Agrobacterium-mediated transformation with the pM764 or pM765 vector described in Example 2 above.
- the transformed plant tissue was grown to produce mature Ro plants.
- a total of 168 Ro transformants were regenerated per construct and assayed by long or short amplicon sequencing for detection of edits in the endogenous GmTDFl-1 gene. Sequence data from each sample was mapped to a reference sequence to identify consensus differences. Plants with unique deletions were selected to provide diverse coverage of gene mutations in the targeted genomic region.
- Most mutations generated by the CRISPR/Cpfl system can comprise insertions, deletions, and/or inversions depending on the guide RNA(s) used.
- Cpfl -derived edits are located downstream of the PAM site.
- the in the “Ro Genotype” column indicates the status of allele 1/allele 2 wherein “WT” indicates that one allele is unedited and “d” indicates that a deletion had occurred during the editing process.
- the nucleotide positions of each of the deletions described in Table 4 are based on the nucleotide position(s) counted from the first nucleotide of the sequence of SEQ ID NO:4 being “position 1.”.
- the nucleotide positions in Table 4 are based on an alignment created using the BLAST computational alignment method. Generally, visibly misaligned bases at either end of a modification, an occasional byproduct of multiple sequence alignment methods, can be corrected manually so that the coordinates reflect the correct positions of the modifications with reference to SEQ ID NO:4.
- Edit 1 through Edit 8 (SEQ ID NOs:26-33) consisted of deletions from one base pair to 11 base pairs near the site of the ms6 mutation. Edit 1, Edit 4, Edit 5, and Edit 7 resulted in the mutation of a splice site between Intron 1 and Exon 2, whereas Edit 2, Edit 3, Edit 6, and Edit 8 resulted in frameshift mutations. All of the resultant plants were fertile in the Ro generation, meaning that pods were present on plants at representative reproductive stages.
- Edit 10 comprises a 15 bp deletion resulting in a frameshift mutation in the GmTDFl-1 gene (SEQ ID NO:36).
- the coordinates presented in Table 5 are based on an alignment created using the BLAST computational alignment method. Generally, visibly misaligned bases at either end of a modification, an occasional byproduct of multiple sequence alignment methods, can be corrected manually so that the coordinates reflect the correct positions of the modifications with reference to SEQ ID NO:4.
- multiple sets of “Position of Edits (relative to SEQ ID NO:4)” coordinates are possible as already described above for Table 4. However, regardless of the alignment method utilized, all edits are unambiguously defined by their individual sequence presented in Table 5 via the identifier “SEQ ID NO:”.
- Edit 25, Edit 26, Edit 27, and Edit 28 each consisted of an inversion on the GmTDFl-1 gene.
- inversions can be accompanied by loss of nucleotides surrounding the inversion, resulting in deleted sequences on either end of the inverted sequence.
- Cpfl- mediated gene editing the Cpfl endonuclease produces staggered cuts and, as such, overhanging DNA ends are generated. These are subject to endogenous DNA repair mechanisms, such as error- prone non-homologous end joining (NHEJ) which can lead to sequence variation at the cut and ligation sites.
- NHEJ error- prone non-homologous end joining
- Plants heterozygous for the edited alleles displayed phenotypes similar to the wild-type.
- the homozygous mutant soybean phenotype resembles the phenotype observed in homozygous AtTDFl mutant anthers in Arabiclopsis thaliana.
- Similar to the AtTDFl loss-of-function mutant ms6 edited mutations in the GmTDFl-1 gene exhibited disrupted tapetum development.
- Edit 2 and Edit 4 showed typical recessive segregation patterns (near 1:2:1).
- Edit 7 showed roughly a typical recessive segregation pattern (near 1:2:1) but had lower germination than other alleles.
- Edit 3 showed a segregation bias resulting in no homozygotes, likely indicating an undetected second-site mutation.
- Nuclease- ncgativc, edit-positive individuals were identified from 3 of 4 edited lines (Edit 2, Edit 3, and Edit 4) and were used to produce R2 seed (data not shown).
- the male sterility phenotype in publicly available ms6 soybean lines is the result of a non- synonymous single nucleotide polymorphism at the Ms6 locus that introduces only a single amino acid change in the encoded protein sequence.
- the present study demonstrates that the edited alleles described herein comprise nucleotide deletions in a wide range of lengths, many of which would substantially disrupt expression of the GmTDFl-1 gene, yet they all confer to soybean plants the same male sterility phenotype seen in the ms6 soybean lines without having an impact on normal plant development. Accordingly, these novel edited alleles of the GmTDFl-1 gene have valuable future use in hybrid development and production.
- Gene editing technology is available for applications in many plant species.
- the use of gene editing technology here to modify the GmTDFl-1 gene to produce male sterility in soybeans for enablement of hybrid soy production indicates possible applications of gene editing to other crop and plant species with highly homologous TDF/-likc genes for production of male sterility and enablement of additional hybrid crops.
- the gene edits made in the GmTDFl-1 gene produce a phenotype that is consistent with the mutant AtTDFl allele in Arabidopsis thaliana, on both a microanatomical level and through an overall male sterile reproductive phenotype (lack of pod production in soybeans).
- the high level of functional domain sequence homology between the AtTDFl and GmTDFl-1 proteins combined with the phenotypic results indicate potential for success in applying editing of TDF1 -like genes in other plant species showing high functional domain homology.
- TDF1 homologs were identified in other green plants from both dicot and monocot plant species. Sequence identity is conserved primarily within the SANT/Myb functional domain for DNA-binding located near the N-terminus of the protein and, particularly in dicots, a serine rich domain in the middle of each protein.
- An alignment of the amino acid sequences of AtTDFl (SEQ ID NO:2), GmTDFl-1 (SEQ ID NO:5), and other TDFl-like proteins from both dicot and monocot plant species (SEQ ID NOs:55-64) is shown in FIGS. 11 A- 1 IE.
- the SANT/Myb domain of AtTDFl is annotated as spanning from position 9 to position 116 of the sequence of SEQ ID NO:2.
- the SANT/Myb domain of GmTDFl-1 is annotated as spanning from position 9 to position 112 of sequence of SEQ ID NO:5.
- the amino acid sequences of the other TDFl-like proteins shown in FIGS. 11 A-l IE are from species with a high degree of sequence homology to the SANT/Myb functional domain of AtTDFl and GmTDFl-1.
Landscapes
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Genetics & Genomics (AREA)
- Chemical & Material Sciences (AREA)
- Engineering & Computer Science (AREA)
- Organic Chemistry (AREA)
- Molecular Biology (AREA)
- Zoology (AREA)
- Biomedical Technology (AREA)
- Wood Science & Technology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- General Engineering & Computer Science (AREA)
- Biotechnology (AREA)
- Biochemistry (AREA)
- General Health & Medical Sciences (AREA)
- Biophysics (AREA)
- Microbiology (AREA)
- Medicinal Chemistry (AREA)
- Physics & Mathematics (AREA)
- Plant Pathology (AREA)
- Cell Biology (AREA)
- Botany (AREA)
- Gastroenterology & Hepatology (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Breeding Of Plants And Reproduction By Means Of Culturing (AREA)
- Micro-Organisms Or Cultivation Processes Thereof (AREA)
Abstract
Provided are compositions and methods for altering TDF1-1 levels in soybean plants. Methods and compositions are also provided for altering the expression of genes related to male fertility through editing of the soybean TDF1-1 gene. Modified plant cells and plants having a suppression element or mutation reducing the expression or activity of a TDF1 gene are further provided comprising reduced TDF1 levels and a desired trait, such as male sterility.
Description
COMPOSITIONS AND METHODS FOR ENGINEERING MALE STERILITY IN PLANTS
REFERENCE TO RELATED APPLICATIONS
[001] This application claims the benefit of United States provisional application No. 63/507,916, filed June 13, 2023, herein incorporated by reference in its entirety.
INCORPORATION OF SEQUENCE LISTING
[002] The sequence listing that is contained in the file named “MONS536WO_ST26.xml,” which is 115 kilobytes as measured in Microsoft Windows operating system and was created on June 10, 2024, is filed electronically herewith and incorporated herein by reference.
FIELD OF THE INVENTION
[003] The present disclosure relates to the field of agricultural biotechnology, and more specifically to methods and compositions for genome editing in plants.
BACKGROUND OF THE INVENTION
[004] Hybridization is an important aspect in the breeding of domesticated plants as it enables the introduction of transient hybrid vigor, desirable variation among different germplasms, transgenic trait integration, and generation of novel phenotypes. Plant breeders use hybridization or controlled cross-pollination as the starting point of a breeding cycle in different crops. Conventional methods for cross-pollination of many crop species, especially soybean, typically involves manual pollination and emasculation. This process is complex, time consuming, and labor-intensive due to the anatomy and size of the flowers. Typical commercial breeding programs require thousands or even millions of crosses in workflows such as, development crosses, backcrosses, and trait integration. As breeders look to accelerate crop variety development and reduce labor needs, it is critical to develop tools that facilitate a higher throughput in breeding cycles and improve hybridization efficiency. This includes development of a stable male sterility system in soybean. There exists a need for new stable male-sterile lines in a diverse genetic background in order to accelerate hybrid soybean breeding.
SUMMARY
[005] In one aspect, the present invention provides a modified plant, plant seed, plant part, or plant cell comprising a genomic modification that reduces expression or activity of TDF1, or a homolog thereof, as compared to the expression or activity of TDF1 or a homolog thereof in an otherwise identical plant, plant seed, plant part, or plant cell that lacks the modification. In some embodiments, the modification may be present in at least one allele of an endogenous TDF1 gene or homolog thereof. In other embodiments, the TDF1 gene or homolog thereof may encode a protein comprising a conserved SANT/Myb domain having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the conserved SANT/Myb domain corresponding to position 9 to position 112 of SEQ ID NO:5. In some embodiments, the modification may disrupt the function of a wild-type allele product of the endogenous TDF1 gene or homolog thereof. In other embodiments, the modification may be located within a coding region, non-coding region, or any combination thereof, of the endogenous TDF1 gene or homolog thereof. In further embodiments, the modification may be located in an exon region of said TDF1 gene or homolog thereof. In yet further embodiments, the modification may be located in an intron region of said TDF1 gene or homolog thereof. In even further embodiments, the modification may be located in an exon region and an intron region of said TDF1 gene or homolog thereof. In some embodiments, the plant, plant seed, plant part, or plant cell may be heterozygous for the modification. In other embodiments, the plant, plant seed, plant part, or plant cell may be homozygous for the modification. In some embodiments, the plant, plant seed, plant part, or plant cell may comprise a first modification in a first allele of the TDF1 gene and a second modification in a second allele of the TDF1 gene, the first modification and the second modification being different from one another. In other embodiments, the modification may comprise a deletion, an insertion, a substitution, an inversion, a duplication, or a combination of any thereof. In further embodiments, the modification may be a deletion. In even further embodiments, the deletion may comprise between 1 nucleotide and 1500 nucleotides. In yet further embodiments, the deletion may comprise a region of at least 1 nucleotide, at least 2 nucleotides, at least 5 nucleotides, at least 10 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, at least 500 nucleotides, at least 600 nucleotides, at least 700 nucleotides, at least 800
nucleotides, at least 900 nucleotides, at least 1000 nucleotides, or at least 1500 nucleotides. In other embodiments, the modification may be an inversion. In further embodiments, the inversion may be a region of at least 5 consecutive nucleotides, at least 10 consecutive nucleotides, at least 20 consecutive nucleotides, at least 30 consecutive nucleotides, at least 40 consecutive nucleotides, at least 50 consecutive nucleotides, at least 100 consecutive nucleotides, at least 200 consecutive nucleotides, at least 300 consecutive nucleotides, at least 400 consecutive nucleotides, at least 500 consecutive nucleotides, at least 600 consecutive nucleotides, at least 700 consecutive nucleotides, at least 800 consecutive nucleotides, at least 900 consecutive nucleotides, at least 1000 consecutive nucleotides, or at least 1100 consecutive nucleotides.
[006] In another aspect, the present invention provides a modified plant, plant seed, plant part, or plant cell comprising a genomic modification that reduces expression or activity of TDF1, or a homolog thereof, as compared to the expression or activity of TDF1 or a homolog thereof in an otherwise identical plant, plant seed, plant part, or plant cell that lacks the modification, wherein the modification is present in at least one allele of an endogenous TDF1 gene or homolog thereof. In some embodiments, the plant may be a leguminous plant, or wherein the plant seed, plant part, or plant cell is a plant seed, plant part, or plant cell of a leguminous plant. In other embodiments, the leguminous plant may be a soybean plant, a bean plant, a pea plant, a chickpea plant, an alfalfa plant, a peanut plant, a carob plant, a lentil plant, or a licorice plant. In some embodiments, the leguminous plant may be a soybean plant. In other embodiments, the leguminous plant may be a soybean plant and the TDF1 gene is GmTDFl -1. In further embodiments, the plant, plant seed, plant part, or plant cell may comprise a modification in at least one allele of the GmTDFl-1 gene, wherein the modification is selected from the group consisting of: an 8 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:26; a 10 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:27; a 1 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:28; a 6 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO: 29; a 9 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:30; an 8 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:31; an 11 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:32; a 7 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:33; a 12 base pair deletion and an 8 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:34; two inversions, wherein the resulting nucleotide sequence is SEQ ID NO:35; a 15 base pair deletion,
wherein the resulting nucleotide sequence is SEQ TD NO:36; a 19 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:37; a 107 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:38; a 27 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:39; a 9 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:40; a 9 base pair deletion and a 10 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:41; a 1,025 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:42; an 11 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:43; a 7 base pair deletion and an 11 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:44; an inversion, wherein the resulting nucleotide sequence is SEQ ID NO:45; a 1,055 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:46; an 11 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:47; a 16 base pair indel and a 9 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:48; a 122 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:49; a 65 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:50; an inversion, wherein the resulting nucleotide sequence is SEQ ID NO:51; a 4 base pair deletion and a 36 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:52; two inversions, wherein the resulting nucleotide sequence is SEQ ID NO:53; and any combinations of any thereof. In yet further embodiments, the plant, plant seed, plant part, or plant cell may comprise at least one polynucleotide sequence selected from the group consisting of SEQ ID NOs:26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, and 53. In other embodiments, the plant, plant seed, plant part, or plant cell may comprise at least one polynucleotide sequence selected from the group consisting of SEQ ID NOs:26, 27, 28, 29, 30, 31, 32, and 33. In further embodiments, the plant, plant seed, plant part, or plant cell may comprise at least one polynucleotide sequence selected from the group consisting of SEQ ID NOs:34, 35, 36,
37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, and 53. In many embodiments, the modification may result in a male sterility phenotype in the plant, as compared to the male fertile phenotype of an otherwise identical plant that lacks the modification.
[007] In another aspect, the present invention provides a polynucleotide comprising a sequence selected from the group consisting of SEQ ID NOs:26, 27, 28, 29, 30, 31 , 32, 33, 34, 35, 36, 37,
38, 39, 40, 41, 42, 3, 44, 45, 46, 47, 48, 49, 50, 51, 52, and 53. The present invention also provides a guide RNA comprising a polynucleotide sequence selected from the group consisting of SEQ ID
NOs:23, 24, and 25. In some embodiments, the polynucleotide sequence may be a spacer sequence.
[008] In one aspect, the present invention provides a method for producing a modified plant having a male sterility phenotype, the method comprising: a) introducing a modification into at least one target site in an endogenous TDF1 gene or a homolog thereof of a plant cell; b) identifying and selecting one or more plant cells of step a) comprising said modification in said TDF1 gene or homolog thereof; and c) regenerating at least one plant from at least one or more cells selected in step b). In some embodiments, the target site may be located in a coding region, a non-coding region, or any combination thereof, of an endogenous TDF1 gene or homolog thereof. In other embodiments, the modification may be introduced by a site-specific genome modification enzyme selected from the group consisting of: an RNA-guided nuclease, a zinc-finger nuclease, a meganuclease, a TALE-nuclease, a recombinase, a transposase, and combinations of any thereof. In further embodiments, the site-specific genome modification enzyme may be an RNA-guided nuclease comprising a Cas nuclease, a Cpf 1 nuclease, or a variant of either thereof. In yet further embodiments, the site-specific genome modification enzyme may be an RNA-guided nuclease comprising a Cpf 1 nuclease. In some embodiments, the site-specific genome modification enzyme may create at least one strand break at the target site. In other embodiments, the modification may be selected from the group consisting of a substitution, an insertion, an inversion, a deletion, a duplication, and a combination thereof. In further embodiments, the modification may be a deletion. In yet further embodiments, the deletion may comprise a region of at least 1 nucleotide, at least 2 nucleotides, at least 5 nucleotides, at least 10 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, at least 500 nucleotides, at least 600 nucleotides, at least 700 nucleotides, at least 800 nucleotides, at least 900 nucleotides, at least 1000 nucleotides, or at least 1500 nucleotides. In other embodiments, the modification is an inversion. In further embodiments, the inversion may be a region of at least 5 consecutive nucleotides, at least 10 consecutive nucleotides, at least 20 consecutive nucleotides, at least 30 consecutive nucleotides, at least 40 consecutive nucleotides, at least 50 consecutive nucleotides, at least 100 consecutive nucleotides, at least 200 consecutive nucleotides, at least 300 consecutive nucleotides, at least 400 consecutive nucleotides, at least 500 consecutive nucleotides, at least 600 consecutive nucleotides, at least 700 consecutive nucleotides, at least 800 consecutive nucleotides, at least 900
consecutive nucleotides, at least 1000 consecutive nucleotides, or at least 1100 consecutive nucleotides. The invention also provides male sterile soybean plants produced by the methods described herein.
[009] In another aspect, the present invention provides a method for producing an Fi hybrid soybean plant comprising crossing a male-sterile soybean plant comprising a modified GmTDFl- 1 gene with a second, non-isogenic, male-fertile soybean plant to produce an Fi hybrid soybean plant. The invention also provides Fi hybrid soybean plants produced by the methods described herein.
[010] In yet another aspect, the present invention provides a method for producing an Fi hybrid soybean seed, comprising crossing a modified plant, plant seed, plant part, or plant cell comprising a genomic modification that reduces expression or activity of TDF1, or a homolog thereof, as compared to the expression or activity of TDF1 or a homolog thereof in an otherwise identical plant, plant seed, plant pail, or plant cell that lacks the modification with a second, non-isogenic, male-fertile soybean plant and harvesting the resultant Fi hybrid soybean seed wherein an Fi hybrid soybean seed is produced. The invention also provides Fi hybrid soybean seeds produced by the methods described herein.
[011] In yet another aspect, the present invention provides a method for producing an Fi hybrid soybean seed comprising crossing a male-sterile soybean plant comprising a modified GmTDFl- 1 gene with a second, non-isogenic, male-fertile soybean plant to produce an Fi hybrid soybean seed. The invention also provides Fi hybrid soybean seed produced by the methods described herein.
[012] The present invention also provides modified plants, plant seed, plant parts, or plant cells, wherein the plants, plant seed, plant parts, or plant cells comprise at least one polynucleotide sequence selected from the group consisting of SEQ ID NOs:26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, and 53. In some embodiments, the modified plant, plant seed, plant part, or plant cell may be a transgenic soybean plant, transgenic soybean plant seed, transgenic soybean plant part, or transgenic soybean plant cell.
BRIEF DESCRIPTION OF THE DRAWINGS
[013] The following drawings form part of the present specification and are included to further demonstrate certain aspects of the present disclosure. The disclosure may be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein.
[014] FIGS. 1A and IB show a schematic illustration of the GmTDFl-1 gene and the GmTDFl- 1 amino acid sequence. FIG. 1A shows the exon-intron structure of the GmTDFl-1 gene based on SEQ ID NO:4, where the exons are denoted by grey arrows and all regions lacking arrows denote introns and/or untranslated regions. The approximate site of the T -> A point mutation is also shown. FIG. IB shows an alignment of the amino acid sequences of Exon 2 encoded by the wildtype (Exon2_WT) and the ms6 mutant allele (Exon2_ms6). The amino acid sequence of Exon2_WT is based on SEQ ID NO:5. The
denotes the site of the Leu to His amino acid change that results from the T -> A point mutation in the nucleic acid sequence.
[015] FIG. 2 shows a schematic illustration of the sites targeted for editing by the guide RNAs (gRNAs) in the GmTDFl-1 gene. The top illustration shows the location of the three exons where the exons are denoted by grey arrows and all regions lacking arrows denote introns and/or untranslated regions. The approximate site of the T -> A point mutation is also shown. The bottom illustration denotes the locations and binding directions of the three gRNAs, gRNA_-103C, gRNA_+221C, and gRNA_-l 113G, where each gRNA and the directionality of the strand it binds is denoted by a grey arrowhead.
[016] FIGS. 3A and 3B show illustrations of alignments of the monoallelic Edits 1-8 (based on SEQ ID NOs:26-33), the GmTDFl-1 coding sequence (based on SEQ ID NOG), and the GmTDFl-1 gene (based on SEQ ID NO:4). FIG. 3A shows an alignment of the sequences of each of the Edits 1-8, the GmTDFl-1 coding sequence, and the GmTDFl-1 gene. The gaps shown in the GmTDFl-1 coding sequence are representative of introns and/or untranslated regions whereas the gaps shown in the edited sequences are representative of deletions and/or 5 ’and 3’ untranslated regions. It should be noted that Edit 3 contains a 1 bp deletion which is not visualized in FIG. 3 A, but it is visualized in FIG. 3B. FIG. 3B shows sequence alignments and the specific deletions at a nucleotide level. A
indicates that the sequence does not have a base alignment at that position, which represents a deletion in the edited sequences or an intron in the coding sequence.
[017] FIG. 4 shows an illustration of an alignment of Edits 9-24 (based on SEQ ID NOs:34, 36- 44, 46-50, and 52), the GmTDFl-1 coding sequence (based on SEQ ID NO:3), and the GmTDFl- 1 gene (based on SEQ ID NO:4). The gaps shown in the GmTDFl-1 coding sequence are representative of introns and/or untranslated regions whereas the gaps shown in the edited sequences are representative of deletions and/or 5’ and 3’ untranslated regions.
[018] FIGS. 5A, 5B, 5C, 5D, 5E, 5F, 5G, and 5H show an illustration of an alignment of Edits 25-28 (based on SEQ ID NOs:35, 45, 51, and 53), which comprise inversions, and the GmTDFl- 1 gene (based on SEQ ID NO:4). A shown in the GmTDFl-1 gene indicates that the sequence does not have a base alignment at that position, likely caused by the algorithm making the best alignment with the inverted sequences. A in SEQ ID NOs:35, 45, 51 , and 53 may represent an absence of a base alignment (as in the GmTDFl-1 gene) or may represent a deletion in the edited sequences, including the loss of nucleotides flanking inversions as a result of DNA repair mechanisms at the cut and ligation sites during the gene editing process, and/or 5’ or 3’ untranslated regions.
[019] FIG. 6 shows images of a phenotypic comparison of representative pod growth in wild-type plants and plants comprising the heteroallelic combination of mutant alleles at the GmTDFl-1 locus (Edit 9/Edit 25). The left panel shows normal pod growth in the wild-type plant and the right panel shows the failed pod growth in the Edit 9/Edit 25 plant.
[020] FIG. 7 shows images of a plant comprising the heteroallelic combination of mutant alleles at the GmTDFl-1 locus (Edit 9/Edit 25) before and after fertilization in a pod growth rescue experiment. The left panel shows the failed pod growth in an Edit 9/Edit 25 plant before pollination and the right panel shows developing pods in an Edit 9/Edit 25 plant after the plant was pollinated with donor pollen from wild- type plants.
[021] FIG. 8 shows images of a comparison of anther development in wild-type and GmTDFl-1 mutant plants in dissected flowers. The left panel shows representative wild-type anther development with pollen shattered. The right panel shows representative edited GmTDFl-1 mutant (male sterile) abnormal anther development.
[022] FIG. 9 shows images of the micro anatomical dissection of anthers from Arabiclopsis thaliana and Glycine max plants, specifically from AtTDFl mutant, GmTDFl-1 mutant, and wildtype plants. The leftmost top and bottom panels show wild-type Arabiclopsis anther sections (top)
and homozygous AtTDFl mutant Arabidopsis anther sections (bottom). “E” designates epidermis, “En” designates cndothccium, “T” designates tapetum, “MSp” designates microsporcs, and “ML” designates middle layer. The middle and rightmost top panels show wild-type and GmTDFl-1 heterozygous mutant soybean anther sections. The middle and rightmost bottom panels show GmTDFl-1 homozygous mutant soybean anther sections with arrows highlighting phenotypic characteristics that are different from the wild-type.
[023] FIG. 10 shows a diagrammatic representation of the alleles of Edits 1-8 shown in relation to the GmTDFl-1 Intron 1-Exon 2 border and T-^A mutation. The leftmost panel provides the edit ID. The middle panel shows the size (Abp) and relative location of the deletion in that edit. The rightmost panel describes the type of mutation caused by the edit. The thin, grey, vertical line represents the splice junction between Intron 1 and Exon 2. The wide, grey, vertical box represents the location of the classical ms6 T->A mutation.
[024] FIGS. 11A, 11B, 11C, 11D, and HE show an alignment of TDFl-like amino acid sequences across numerous plant species. Alignment of AtTDFl, GmTDFl-1, and TDFl-like proteins from both dicot and monocot plant species is provided with consensus sequence for the amino acid found in the majority of sequences. The region of highest conservation (the SANT/Myb functional domain) is located in the N-terminal region, while the serine rich domain highly conserved in dicots begins around position 262 of the amino acid sequences.
BRIEF DESCRIPTION OF THE SEQUENCES
[025] SEQ ID NO:1 is a 954-nucleotide sequence representing the coding region of the wild-type Arabidopsis thaliana TDF1 (AtTDFl) gene.
[026] SEQ ID NO:2 is a 317-amino acid sequence representing the Arabidopsis thaliana TDF1 (AtTDFl) protein sequence. SEQ ID NO:2 is the sequence of the polypeptide encoded by SEQ ID NO:1.
[027] SEQ ID NO:3 is a 1038-nucleotide sequence representing the coding region of the wild-type Glycine max TDF1-1 (GmTDFl-1) gene.
[028] SEQ ID NO:4 is a 2094-nucleotide sequence representing the genomic DNA sequence of the wild-type GmTDFl-1 gene. The sequence includes a 5’ untranslated region (UTR), the GmTDFl-1 gene including introns and exons, and a 3’ UTR.
[029] SEQ TD NO:5 is a 345-amino acid sequence representing the Glycine max TDF1-1 (GmTDFl-1) protein sequence. SEQ ID NO:5 is the sequence of the polypeptide encoded by the nucleic acid sequence of SEQ ID NO:3.
[030] SEQ ID NO:6 is a 33-nucleotide sequence corresponding to a forward primer referred to as “NGMAX009625172 - FP” used to genotype seeds and/or plants for GmTDFl-1.
[031] SEQ ID NO:7 is a 25-nucleotide sequence corresponding to a reverse primer referred to as “NGMAX009625172 - RP” used to genotype seeds and/or plants for GmTDFl-1.
[032] SEQ ID NO:8 is a 16-nucleotide sequence corresponding to a probe referred to as “NGMAX009625172 - F” used for marker assisted selection of seeds and/or plants for GmTDFl- 1.
[033] SEQ ID NO:9 is a 15 -nucleotide sequence corresponding to a probe referred to as “NGMAX009625172 - V” used for marker assisted selection of seeds and/or plants for GmTDFl- 1.
[034] SEQ ID NO: 10 is a 21-nucleotide sequence corresponding to a reverse primer referred to as “ms6 - RP1” used to genotype seeds and/or plants for GmTDFl-1.
[035] SEQ ID NO: 11 is a 30-nucleotide sequence corresponding to a reverse primer referred to as “ms6 - RP2” used to genotype seeds and/or plants for GmTDFl-1.
[036] SEQ ID NO: 12 is a 25-nucleotide sequence corresponding to a forward primer referred to as “ms6 - FP1” used to genotype seeds and/or plants for GmTDFl-1.
[037] SEQ ID NO: 14 is a 15-nucleotide sequence corresponding to a probe referred to as “GMMS6S22427963_2_M” used for marker assisted selection of seeds and/or plants for GmTDFl-1.
[038] SEQ ID NO: 18 is a 20-nucleotide sequence representing a common scaffold compatible with a Cpfl gene from Lachnospiraceae bacterium ND2006.
[039] SEQ ID NO: 19 is a 365-nucleotide sequence representing a promoter from Dahlia mosaic virus.
[040] SEQ ID NO:20 is a 3681 -nucleotide sequence representing a Cpfl RNA-guided endonuclease enzyme codon-optimized for plants from Lachnospiraceae bacterium ND2006.
[041] SEQ ID NO:21 is a 30-nucleotide sequence representing a nuclear localization signal from Solatium lycopersicum.
[042] SEQ ID NO:22 is a 427-nucleotide sequence representing a Pol3 promoter from Glycine max.
[043] SEQ ID NO:23 is a 23-nucleotide guide RNA spacer sequence that corresponds to positions 328-350 on SEQ ID NO:4 in a forward orientation.
[044] SEQ ID NO:24 is a 23-nucleotide guide RNA spacer sequence that corresponds to positions 188-210 on SEQ ID NO:4 in a reverse orientation.
[045] SEQ ID NO:25 is a 23-nucleotide guide RNA spacer sequence that corresponds to positions 1198-1220 on SEQ ID NO:4 in a reverse orientation.
[046] SEQ ID NO:26 is a 1716-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 1.”
[047] SEQ ID NO:27 is a 1714-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 2.”
[048] SEQ ID NO:28 is a 1723-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 3.”
[049] SEQ ID NO:29 is a 1718-nuclcotidc sequence representing an edited GmTDFl-1 allele corresponding to “Edit 4.”
[050] SEQ ID NQ:30 is a 1715-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 5.”
[051] SEQ ID NO:31 is a 1716-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 6.”
[052] SEQ ID NO:32 is a 1713-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 7.”
[053] SEQ ID NO:33 is a 1717-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 8.”
[054] SEQ ID NO: 34 is a 1704-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 9.
[055] SEQ ID NO:35 is a 1653-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 25.”
[056] SEQ ID NO:36 is a 1709-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 10.
[057] SEQ ID NO:37 is a 1705 -nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 11.”
[058] SEQ ID NO:38 is a 1617-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 12.”
[059] SEQ ID NO:39 is a 1697-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 13.
[060] SEQ ID NO:40 is a 1715-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 14.”
[061] SEQ ID NO:41 is a 1705 -nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 15.
[062] SEQ ID NO:42 is a 699-nuclcotidc sequence representing an edited GmTDFl-1 allele corresponding to “Edit 16.
[063] SEQ ID NO:43 is a 1713-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 17.
[064] SEQ ID NO:44 is a 1706-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 18.
[065] SEQ ID NO:45 is a 1706-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 26.”
[066] SEQ ID NO:46 is a 669-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 19.
[067] SEQ ID NO:47 is a 1713-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 20.
[068] SEQ ID NO:48 is a 1692-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 21.”
[069] SEQ ID NO:49 is a 1602-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 22.
[070] SEQ ID NO:50 is a 1659-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 23.”
[071] SEQ ID NO:51 is a 1707 -nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 27.”
[072] SEQ ID NO:52 is a 1684-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 24.
[073] SEQ ID NO: 53 is a 1018-nucleotide sequence representing an edited GmTDFl-1 allele corresponding to “Edit 28.”
[074] SEQ ID NO:55 is a 321-amino acid sequence representing the Arachis hypogaea TDFl-like protein sequence.
[075] SEQ ID NO:56 is a 339-amino acid sequence representing the Cicer arietinum TDFl-likc protein sequence.
[076] SEQ ID NO:57 is a 333-amino acid sequence representing the Vitis vinifera TDFl-like protein sequence.
[077] SEQ ID NO:58 is a 268-amino acid sequence representing the Cucumis sativus TDFl-like protein sequence.
[078] SEQ ID NO:59 is a 320-amino acid sequence representing the Brassica napus TDFl-like protein sequence.
[079] SEQ ID NO:60 is a 387-amino acid sequence representing the Gossypium hirsutum TDFl- like protein sequence.
[080] SEQ ID NO:61 is a 263-amino acid sequence representing the Solarium lycopersicum TDF1 - like protein sequence.
[081] SEQ ID NO:62 is a 319-amino acid sequence representing the Zea mays TDFl-like protein sequence.
[082] SEQ ID NO:63 is a 328-amino acid sequence representing the Sorghum bicolor TDFl-like protein sequence.
[083] SEQ ID NO:64 is a 306-amino acid sequence representing the Oryza sativa subsp. japonica TDFl-like protein sequence.
[084] SEQ ID NO:65 is a 60-nucleotide sequence representing the 5’ junction of an inverted sequence and accompanying deletion(s) or modification(s). It comprises at least 20 bp on either side of the junction point.
[085] SEQ ID NO:66 is a 60-nucleotide sequence representing the 3’ junction of an inverted sequence and accompanying deletion(s) or modification(s). It comprises at least 20 bp on either side of the junction point.
[086] SEQ ID NO:67 is a 60-nucleotide sequence representing the 5’ junction of an inverted sequence and accompanying deletion(s) or modification(s). It comprises at least 20 bp on either side of the junction point.
[087] SEQ ID NO:68 is a 60-nucleotide sequence representing the 3’ junction of an inverted sequence and accompanying deletion(s) or modification(s). It comprises at least 20 bp on either side of the junction point.
[088] SEQ ID NO:69 is a 60-nucleotide sequence representing the 5’ junction of an inverted sequence and accompanying deletion(s) or modification(s). It comprises at least 20 bp on either side of the junction point.
[089] SEQ ID NO:70 is a 60-nucleotide sequence representing the 3’ junction of an inverted sequence and accompanying deletion(s) or modification(s). It comprises at least 20 bp on either side of the junction point.
[090] SEQ ID NO:71 is a 60-nucleotide sequence representing the 5’ junction of an inverted sequence and accompanying deletion(s) or modification(s). It comprises at least 20 bp on either side of the junction point.
[091] SEQ ID NO:72 is a 60-nucleotide sequence representing the 3’ junction of an inverted sequence and accompanying dclction(s) or modification(s). It comprises at least 20 bp on cither side of the junction point.
DETAILED DESCRIPTION
[092] The goal of crop breeding is to produce agricultural plants with desirable traits, such as increased grain yield, improved nutrition quality, and enhanced adaptation to environmental changes. One of the ways breeders have produced such plants is through hybrid breeding (also called hybridization) using inbred parent varieties derived from diverse gene pools to generate hybrid plants. Because hybrids have a full set of genes from each unique parent, they have much more genetic diversity than cither inbred parent alone. Hybrids also have a combination of traits from each parent that enables the hybrids to outperform either of their inbred parents. The combination of genetic diversity between the two inbred parents is the source of the phenomenon of heterosis or “hybrid vigor.” In many crops, hybrids are produced by controlled crosspollination, which requires hand-crossing selected lines to produce the desired Fi seed while also preventing unwanted self-pollination through use of a pollination control system. Manual emasculation, namely the physical removal of either the male floral structure from hermaphrodite flowers or removal of whole male flowers from monoecious plants, remains the predominant method of pollination control in hybrid seed production. This process is time consuming and labor- intensive for most crop species but is especially difficult in soybean (Glycine max). Soybean is autogamous and is thus a self-pollinating species. This is largely due to the structure of the flower itself, which is small in size, fragile, and contains both the pistil and the stamens situated in a manner that facilitates the anthers to shed mature pollen directly onto the stigma of the same flower. Further, because pollen is shed shortly before or immediately after the flower opens, self- pollination typically occurs before the flowers are fully open. As a result, natural outcrossing occurs at rates below 1%. Because manual emasculation is labor, time, and cost intensive in soybean, utilizing hybrid breeding for heterosis in this crop has continued to be a challenge.
[093] The increasing demand for hybrid varieties in many crops has driven efforts to seek alternative breeding methods that are both more efficient and less costly. The advances in plant cell biology and plant biotechnology have enabled development of genetic methods of pollination control, specifically male sterility. Male sterility in plants refers to the inability of a plant to
produce functional anthers, pollens, and/or male gametes during their reproductive stages. When used in breeding, it eliminates the need for manual emasculation and, in many cases, pollination by hand. Male sterility can be induced through either chemical or genetic manipulation. A number of chemicals (e.g., auxins, anti-auxins, halogenated aliphatic acids, gibberellins, Ethephon, etc.) have been identified as male gametocides but these agents have largely not been adopted due to the extensive specificity required for successful application. Instead, male sterility systems in many crops have been developed through genetic manipulation of genes involved in male fertility. Plant male sterility can be caused either by nuclear genes alone (genic/genetic male sterility) or through mitochondrial genes that directly or indirectly affect nuclear gene function (cytoplasmic male sterility). Cytoplasmic male sterility (CMS) is used in commercial crop hybrid seed production in a three-line system consisting of male-sterile lines, maintainer lines, and restorer lines, although it often suffers from poor genetic diversity, increased disease susceptibility, and unstable restoration of CMS lines. Genic/genetic male sterility (GMS) is controlled by nuclear Male sterility (Ms) genes without the influence of cytoplasmic sequences controlled by nuclear genes alone. GMS can be genetically stable or environment-sensitive (EGMS). GMS is a common spontaneous occurrence in flowering plants and occurs in nearly all species. These abnormal plants are usually homozygous recessive (msms) in the self-pollinated progenies of heterozygous (Msms) individuals. Generally, such mutants appear in low frequency and are often lost in the selfpollinated crops. However, in cross-pollinated crops, the recessive mutant allele is protected by its dominant male fertile counterparts as hetero zygotes.
[094] Soybean is believed to have strong heterosis and utilizing this has been a strategy to increase soybean yield. The bottleneck in commercial application of soybean heterosis is a lack of suitable male-sterile lines. Although a three-line system based on CMS has been developed in soybean, very few CMS lines have been identified and restoration of these lines is unreliable. This restricts full use of heterosis in a commercial hybrid seed production system. GMS can overcome these disadvantages, but large-scale production of seeds is greatly hindered because half of the progenies in a homozygous male sterile line will restore fertility when male sterile plants are crossed to heterozygous male fertile plants during seed propagation of the male sterile line. These male fertile plants must be identified and eliminated during hybrid seed production, which increases labor costs. The application of male sterility in soybean hybrid seed production has typically lagged other crops.
[095] GMS in soybean is a heritable trait and more than 20 ms mutations have been reported (Zhao et al., Front Plant Sci. 10(94): 1-9, 2019). Most of these mutations have been mapped to soybean linkage groups, however very few have been physically mapped and characterized at the molecular level. Mutations causing male sterility often have pleiotropic effects, some of which lead to developmental defects such has reduced plant height and slower vegetative growth, and/or defects in female fertility. Furthermore, given that the phenotype of GMS is only observable after flowering and the recessive nature of ms genes, it is necessary to have a marker that is associated with GMS to facilitate selection of the male sterile seeds or plants as early as possible. Due to the lack of molecular characterization regarding many of ms mutations in soybean, efforts have been made to identify recessive morphological traits associated with male sterility. Among the identified mutants, ms6 is one that has a recessive morphological trait marker that can be used for phenotypic selection. The ms6 mutant was identified as an environmentally stable nuclear male sterile mutation at the Ms6 locus that results in a non-pollen phenotype. Plants carrying at least one copy of the wild-type Ms6 allele are fertile, whereas ms6ms6 plants are female fertile and completely male sterile due to abnormalities in tapetum development in anthers. The Ms6 locus is closely linked the flower color locus W1 which is commonly used for selection of male sterile lines in hybrid seed production. One of the main limitations of morphological markers such as W1 is the developmental stage at which they appear. For example, in the case of the W1 locus, Ms6_Wl_ plants (Ms6ms6Wlwl or Ms6Ms6WlWl) have purple hypocotyls and flowers and are fertile whereas those of the genotype ms6ms6wlwl have green hypocotyls and white flowers and are male sterile. The hypocotyl color trait is visible shortly after emergence and at that stage purple hypocotyl seedlings must be removed manually from the field. At flowering however, not all white flowered plants are male sterile, as white flowered fertile plants may arise from recombination between the W1 and Ms6 loci. The plants must therefore be inspected at the flowering stage for anther morphology and plants with normal anthers must be removed. This process is laborious and time consuming, especially due to the structure of soybean flowers.
[096] There are further disadvantages to use of the ms6 allele for male sterility, especially for large-scale soybean breeding. For example, some high yielding and male fertile elite soybean lines have white flowers, and accordingly, the white flower color phenotype could not be used a marker for the presence of ms 6 in homozygous form. The ms6 allele also differs from other male sterile, female fertile genes in soybean due to an observed pleiotropic effect on smaller flower size. Male-
sterile plants (ms6ms6w 1 w 1) had flowers of smaller size when compared to fertile, purple flowered Ms6_WlWl plants, however it is not clear if the smaller flower is due to homozygosity of ms6 or wl . Flower size is known to play a role in attracting pollinators as bumble bees prefer large floral displays. As such, a small flower phenotype may have a negative impact on pollination in the field. Finally, the male sterile phenotype conferred by the ms6 allele results from a point mutation that alters a single amino acid. Point mutations have the potential to revert back to wild-type by random mutagenesis, and are therefore not considered to be highly genetically stable. Successful utilization of male sterility in hybrid breeding requires a male sterility phenotype that is genetically stable and does not incur any yield penalty in various genetic backgrounds and environments, as unstable male sterility leads to low purity of hybrid seeds and reduced grain production. Furthermore, it is desirable to identify male sterility genes with restricted expression patterns and, ideally, no pleiotropic functions. A need exists for an alternative genetic source of male sterility that is stable and can be used in any background for hybrid seed production in soybean.
[097] In many plants, male sterility may result from dysfunction of either the tapetum or the reproductive cells, or both. In Arabidopsis, several genes have been reported to be important for tapetai differentiation and function. Specifically, the Arabidopsis DEFECTIVE in TAPETAL DEVELOPMENT and FUNCTION 1 (AtTDFl) gene has been identified as being important in regulation of tapetai cell differentiation and function. The tapetum plays important roles in microspore development as it provides nutrients for the developing microspores and pollen, it produces sporopollenin which is deposited on the outer walls of the released microspores, and it produces the enzyme callase which degrades the callose walls of the microspore tetrads that are the products of meiosis. AtTDFl mutants were found to have aborted pollen development due to defects in tapetai function. This phenotype is similar to the soybean ms6 phenotype, where ms6ms6 plants are completely male sterile due to abnormalities in tapetum. A study of soybean ms6 mutants by Yu et al. Theor. Appl. Genet. 134:3661-3674, 2021) identified a homolog of AtTDFl on chromosome 13, designated GmTDFl-1. While a second soybean TDF1 gene was also identified on chromosome 19 (designated GmTDFl-2), it is believed that this locus may be redundant due to its low level of expression and its inability to compensate for the deficiency in ms6ms6 (i.e., male sterile) plants.
[098] The present disclosure represents a significant advance in the art in that it provides engineered alleles that confer stable male sterility phenotypes in soybeans and other crops, as well
as methods for the production thereof, thereby offering improvements in genetic male sterility systems that can lead to reduced breeding times and increased yield and quality. The edited alleles described herein cause male sterility in soybean plants without having any pleiotropic effects. The methods and compositions disclosed herein offer the opportunity to create diversity that cannot be selected from conventional plant breeding or random mutagenesis. Accordingly, provided herein are methods and compositions for generating novel male sterility alleles in soybean that may be used to achieve such benefits, including, for example, rapid introduction of recessive and penetrant mutant GmTDFl-1 alleles into elite soybean lines. To produce soybean and other plants having a male-sterile phenotype, the present disclosure provides, in certain embodiments, methods and compositions for the creation of novel alleles at the Ms6 locus via editing of the GmTDFl-1 gene. Therefore, the present disclosure represents a significant advance in the art in that it permits the production of novel engineered alleles in soybean and other crops that confer male sterility phenotypes with the potential to accelerate crop breeding cycles and improved yield gains and quality through capturing heterosis.
I. Genome Editing
[099] The present disclosure provides, in certain embodiments, plants, plant parts, plant cells, and seeds produced through genome modification using site-specific integration or genome editing. Genome editing can be used to make one or more edit(s) or mutation(s) at a desired target site in the genome of a plant, such as to change expression and/or activity of one or more genes, or to integrate an insertion sequence or transgene at a desired location in a plant genome. Any site or locus within the genome of a plant may potentially be chosen for making a genomic edit (or gene edit) or site-directed integration of a transgene, construct, or transcribable DNA sequence.
[0100] As used herein, a “target site” or “a target sequence” for genome editing refers to the location of a polynucleotide sequence within a plant genome that is bound and cleaved by a sitespecific nuclease to introduce a strand break into the nucleic acid backbone of the polynucleotide sequence and/or its complementary DNA strand within the plant genome. A target site may comprise, for example, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27 , at least 28, at least 29, or at least 30 consecutive nucleotides. A “target site” for an RNA-guided nuclease may comprise the sequence of either complementary
strand of a double-stranded nucleic acid (DNA) molecule or chromosome at the target site. An RNA-guidcd nuclease may bind to a target site through, for example, a non-coding guide RNA (e.g., without being limiting, a CRISPR RNA (crRNA) or a single-guide RNA (sgRNA), as described further herein). A non-coding guide RNA provided herein may be complementary to a target site (e.g., complementary to either strand of a double-stranded nucleic acid molecule or chromosome at the target site). It will be appreciated that perfect identity or complementarity may not be required for a non-coding guide RNA to bind or hybridize to a target site. For example, at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8 mismatches (or more) between a target site and a non-coding RNA may be tolerated. A “target site” for any other site- specific nuclease that may not be guided by a non-coding RNA molecule, such as a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), a meganuclease, etc., also refers to the location of a polynucleotide sequence within a plant genome that is bound and cleaved. As used herein, a “target region” or a “targeted region” refers to a polynucleotide sequence or region that is flanked by two or more target sites. Without being limiting, in some embodiments a target region may be subjected to a mutation, insertion, deletion, substitution, duplication, or inversion. As used herein, “flanked” when used to describe a target region of a polynucleotide sequence or molecule, refers to two or more target sites of the polynucleotide sequence or molecule surrounding the target region, with one target site on each side of the target region.
[0101] As used herein, with respect to a given sequence, a “complement,” a “complementary sequence,” and a “reverse complement” are used interchangeably. All three terms refer to the inversely complementary sequence of a nucleotide sequence, i.e., to a sequence complementary to a given sequence in reverse order of the nucleotides.
[0102] As used herein, the term “antisense” refers to DNA or RNA sequences that are complementary to a specific DNA or RNA sequence. Antisense RNA molecules are singlestranded nucleic acids which can combine with a sense RNA strand or sequence or mRNA to form duplexes due to complementarity of the sequences. The term “antisense strand” refers to a nucleic acid strand that is complementary to the “sense” strand. The “sense strand” of a gene or locus is the strand of DNA or RNA that has the same sequence as an RNA molecule transcribed from the gene or locus (with the exception of uracil in RNA and thymine in DNA).
[0103] As used herein, a “targeted editing technique” refers to any method, protocol, or technique that allows the precise and/or targeted editing of a specific location in a genome of a plant (i.e., the editing is largely or completely non-random) using a site-specific nuclease, such as a meganuclease, a zine-finger nuclease (ZFN), an RNA-guided endonuclease (e.g., the CRISPR/Cas9 system or the CRISPR/Cpfl system), a TALE (transcription activator-like effector)-endonuclease (TALEN), a recombinase, or a transposase.
[0104] As used herein, “editing” or “genome editing” refers to generating a targeted mutation, insertion, deletion, substitution, duplication, or inversion, of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 75, at least 100, at least 250, at least 500, at least 1000, at least 2500, at least 5000, at least 10,000, or at least 25,000 nucleotides of an endogenous plant genome nucleic acid sequence. An “edit” or “genomic edit” in the singular refers to one such targeted mutation, insertion, deletion, substitution, duplication, or inversion, whereas “edits” or “genomic edits” refers to two or more targeted mutation(s), insertion(s), deletion(s), substitution(s), duplication(s), and/or inversion(s), with each “edit” being introduced via a targeted editing technique.
[0105] Genome editing or targeted editing can be facilitated via the use of one or more site-specific nucleases. A site-specific nuclease provided herein may be selected from the group consisting of zinc-finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), meganucleases, RNA-guided endonucleases (e.g., Cas9 and Cpfl), recombinases (without being limiting, for example, a serine recombinase attached to a DNA recognition motif, a tyrosine recombinase attached to a DNA recognition motif), transposases (without being limiting, for example, a DNA transposase attached to a DNA binding domain), or any combination thereof. See, e.g., Khandagale et al. (Plant Biotechnol Rep 10:327-343, 2016); and Gaj et al. (Trends Biotechnol. 31(7):397-405, 2013).
[0106] In an aspect, the targeted genome editing described herein may comprise the use of an RNA-guided endonuclease. As used herein, an “RNA-guided endonuclease” refers to an RNA- guided DNA endonuclease associated with the CRISPR system. According to some embodiments, an RNA-guided endonuclease may be selected from the group consisting of Casl, CaslB, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csxl2), CaslO, Csyl, Csy2,
Csy3, Cse1 , Cse2, Csc 1 , Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl , Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, CsxlO, Csxl6, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4, Cpfl (also known as Casl2a), CasX, CasY, and homologs or modified versions of any thereof, as well as Argonaute (non-limiting examples of Argonaute proteins include Thermus thermophilus Argonaute (TtAgo), Pyrococcus furiosus Argonaute (PfAgo), and Natronobacterium gregoryi Argonaute (NgAgo)), and homologs or modified versions of any thereof). According to some embodiments, an RNA-guided endonuclease is a Cas9 or Cpfl enzyme. According to some embodiments, an RNA-guided endonuclease is a Cpfl enzyme.
[0107] The CRISPR system, in its native context, provides bacteria and archaea with immunity to invading foreign nucleic acids and relies on an RNA-guided endonuclease to cleave the invading DNA or RNA into short sequence fragments and incorporating them into the bacterial CRISPR genomic locus. The incorporated short sequences, referred to as “protospacers”, and flanking direct repeats are transcribed and processed into CRISPR RNAs (crRNAs). These crRNAs hybridize with /ra -activating crRNAs (tracrRNAs) to activate the RNA-guided Cas endonuclease to form a ribonucleoprotein (RNP) complex that is guided to a target site. A prerequisite for cleavage of the target site, however, is the presence of a conserved genomic protospacer-adjacent motif sequence recognized by the Cas endonuclease. A “protospacer adjacent motif’ (PAM) herein refers to a short nucleotide sequence adjacent to a target sequence (protospacer) that is recognized (targeted) by a guide polynucleotide/Cas endonuclease system described herein. A PAM sequence may be present in the genome immediately adjacent and upstream (z'.e. the 5’ end) of the genomic target site sequence complementary to the targeting sequence of the guide RNA - i.e., immediately downstream (3’) to the sense (+) strand of the genomic target site (relative to the targeting sequence of the guide RNA) as known in the art. See, e.g., Wu et al. (Quant Biol. 2(2):59-70, 2014). The Cas endonuclease may not successfully recognize a target DNA sequence if the target DNA sequence is not followed by a PAM sequence. The sequence and length of a PAM sequence herein can differ depending on the Cas endonuclease used. The PAM sequence can be of any length but is typically 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides long.
[0108] CRISPR/Cas9, which is the CRISPR system from Streptococcus pyogenes, was adapted for use in eukaryotes and has been widely used for gene editing in plants. The CRISPR/Cas9 system requires both crRNA and tracrRNA to guide the Cas9 protein to recognize and cleave the
target DNA double helix. Cas9 recognizes the genomic PAM sequence 5’-NGG-3’ (where N is any nucleotide) and, when located on the sense strand adjacent to the target site, will create a blunt- end DSB at the target site, specifically the 5 '-end of the PAM site. Cas9 has been observed to recognize other PAM sequences, such as 5’-NAG-3’and 5’-NGA-3,’ which may result in cleavage of non-specific DNA sequences.
[0109] Recently, the CRISPR/Cpfl system (Cpfl is also known as Casl2a) was discovered as an alternative to the CR1SPR/Cas9 system for genome editing. While CRISPR/Cpfl functions in a manner similar to CRISPR/Cas9, it is an even simpler system than CRISPR/Cas9. CRISPR/Cpfl requires only one crRNA molecule and no tracrRNA to cleave DNA. Cpfl recognizes the genomic PAM sequence 5’-TTTV-3’ (where V is A, G, or C) or 5’-TTN-3’, depending on the Cpfl ortholog. See e.g., Alok et al. Front. Plant Sci. 11 :264, 2020). When Cpfl recognizes the genomic PAM located on the sense strand adjacent to the target site, it will generate a staggered DSB with a 4- or 5-nt 5' overhang at the target site, specifically the 3'-end of the PAM site.
[0110] For RNA-guided endonucleases, a guide RNA molecule may be further provided to direct the endonuclease to a target site in the genome of the plant via base-pairing or hybridization to cause a DSB or nick at or near the target site. The guide RNA may be transformed or introduced into a plant cell or tissue as a gRNA molecule, or as a recombinant DNA molecule, construct or vector comprising a transcribable DNA sequence encoding the guide RNA operably linked to a promoter. As understood in the art, a guide RNA may comprise, for example, a CRISPR RNA (crRNA), a single-chain guide RNA (sgRNA), or any other RNA molecule that may guide or direct an endonuclease to a specific target site in the genome. A “single-chain guide RNA” (or “sgRNA”) is a RNA molecule comprising a crRNA covalently linked a tracrRNA by a linker sequence, which may be expressed as a single RNA transcript or molecule. The guide RNA comprises a guide or targeting sequence that is identical or complementary to a target site within the plant genome, such as at or near a TDF1 gene, such as the GmTDFl-1 gene in soybean. Nucleic acid molecules provided herein can combine a crRNA and a tracrRNA into one nucleic acid molecule in what is herein referred to as a “single guide RNA (sgRNA).”
[0111] The guide RNA is typically a non-coding RNA molecule that does not encode a protein. The guide sequence of the guide RNA may be at least 10 nucleotides in length, such as 12-40 nucleotides, 12-30 nucleotides, 12-20 nucleotides, 12-35 nucleotides, 12-30 nucleotides, 15-30
nucleotides, 17-30 nucleotides, or 17-25 nucleotides in length, or about 12, 13, 1 , 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 or more nucleotides in length. The guide sequence may be at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical or complementary to at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, or at least 25 or more consecutive nucleotides of a DNA sequence at the genomic target site. RNA-guided endonucleases may be delivered as a protein with or without a guide RNA, or the guide RNA may be complexed with the RNA-guided endonuclease and delivered as a ribonucleoprotein (RNP).
[0112] As mentioned above, a target gene for genome editing may be any of the DEFECTIVE in TAPETAL DEVELOPMENT and FUNCTION 1 -like genes described herein, including but not limited to one of the two soybean TDF1 genes (hereafter referred to as GmTDF 1 -1 ). For genome editing at or near the GmTDF 1-1 gene with an RNA-guided endonuclease, a guide RNA may be used comprising a guide sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical or complementary to at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, or more consecutive nucleotides of SEQ ID NO:4 or a sequence complementary thereto (e.g., 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or more consecutive nucleotides of SEQ ID NO:4 or a sequence complementary thereto). As used herein, the term "consecutive" in reference to a polynucleotide or protein sequence means without deletions or gaps in the sequence.
[0113] In addition to the guide sequence, a guide RNA may further comprise one or more other structural or scaffold sequence(s), which may bind or interact with an RNA-guided endonuclease. Such scaffold or structural sequences may further interact with other RNA molecules (e.g., tracrRNA). Methods and techniques for designing targeting constructs and guide RNAs for genome editing and site-directed integration at a target site within the genome of a plant using an RNA-guided endonuclease are known in the art.
[0114] The introduction of a DSB or nick may be used to introduce one or more mutation in a target region in the genome of a plant. As used herein, a “mutation” refers to the permanent alteration of the nucleic acid sequence or amino acid sequence in an organism, as compared to a naturally occurring reference nucleic acid sequence or amino acid sequence from the same
organism. It will be appreciated that, when identifying a mutation, the reference sequence should be from the same nucleic acid (c.g., gene or non-coding RNA) or amino acid (c.g., protein). In determining if a difference between two sequences comprises a mutation, it will be appreciated in the art that the comparison should not be made between homologous sequences of two different species or between homologous sequences of two different varieties of a single species. The comparison should be made between the edited (e.g., mutated) sequence and the endogenous, nonedited (e.g., “wild-type”) sequence of the same organism.
[0115] Without being limiting, in some embodiments, the mutation may comprise an insertion, a deletion, a substitution, a duplication, or an inversion. According to this approach, mutations such as insertions, deletions, substitutions, duplications, and/or inversions may be introduced at a target site via imperfect repair of the DSB or nick to produce a knock-out or knock-down of a gene. A “knock-out” of a gene may be achieved by inducing a DSB or nick at or near the endogenous locus of the gene that results in non-expression of the protein or expression of a non-functional protein, whereas a “knock-down” of a gene may be achieved in a similar manner by inducing a DSB or nick at or near the endogenous locus of the gene that is repaired imperfectly at a site that does not affect the coding sequence of the gene in a manner that would eliminate the function of the encoded protein. For example, the site of the DSB or nick within the endogenous locus may be in the upstream or 5’ region of the gene (e.g., a promoter and/or enhancer sequence) to affect or reduce its level of expression.
[0116] As used herein, the term “insertion” as it relates to a mutation, refers to the addition of one or more extra nucleotides into the DNA. Insertions in the coding region of a gene may alter splicing of the mRNA (splice site mutation) or cause a shift in the reading frame (frameshift), both of which can significantly alter the gene product.
[0117] As used herein, the term “deletion” as it relates to a mutation refers to the removal of one or more nucleotides from the DNA. Like insertion mutations, these mutations can alter splicing of the mRNA (splice site mutation) or cause a shift in the reading frame of the gene.
[0118] As used herein, the term “substitution” as it relates to a mutation refers to an exchange of a single nucleotide for another.
[0119] As used herein, the term “inversion” refers to reversing the orientation of a chromosomal segment. An inversion can be accompanied by a loss of nucleotides flanking either one or both
sites of the inversion due to DNA repair mechanisms occurring at the cut and ligation sites during the formation of an inversion.
[0120] As used herein, the term “duplication” refers to the creation of multiple copies of chromosomal regions, increasing the dosage of the genes located within them.
[0121] As used herein, a “missense mutation” refers to a single nucleotide change that results in a codon that codes for a different amino acid. For example, the codon “CGU” encodes an arginine amino acid. If a missense mutation changes the G to a U, producing a “CUU” codon, the codon now encodes a leucine amino acid. Missense mutations can be caused by an insertion, deletion, substitution, duplication, or inversion. The frameshift, missense, or nonsense mutations described herein lead to loss of function or expression of a targeted gene, such as a GmTDFl-1 gene. A “loss-of-function mutation” is a mutation in the coding sequence of a gene, which causes the function of the gene product, usually a protein, to be either reduced or completely absent. A loss- of-function mutation can, for instance, be caused by the truncation of the gene product. A phenotype associated with an allele with a loss-of-function mutation can be either recessive or dominant.
[0122] As used herein, the terms “natural mutation,” “naturally-occurring mutation,” or “native mutation,” refers to a mutation that arises spontaneously in nature without any involvement of laboratory or experimental procedures or under the exposure to mutagens. Without being bound by scientific theory, a naturally-occurring mutation can arise from a variety of sources, including errors in DNA replication, spontaneous lesion, and transposable elements (or transposon). As used herein, the terms “non-natural mutation,” “non-naturally-occurring mutation,” or “synthetic mutation” refers to non-spontaneous mutation that occurs as a result of experimental procedures, such as exposure to a mutagen or by a site-specific genome modification enzyme.
[0123] In an aspect, the present disclosure provides a modified soybean plant, or plant part thereof, comprising a mutant allele of the GmTDFl -1 gene, wherein the mutant allele comprises at least one genome modification involving at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 110, at least 125, at least 150, at least 175, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, or at least 1100 consecutive nucleotides of a
coding region, a non-coding region, or any combination thereof, of the endogenous GmTDFl-1 gene. Such genome modifications may include: a 8 base pair deletion (bp) wherein the resulting nucleotide sequence is SEQ ID NO:26; a 10 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:27; a 1 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:28; a 6 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:29; a 9 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:30; a 8 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:31; a l l base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:32; a 7 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:33; a 12 base pair deletion and a 8 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:34; two inversions wherein the resulting nucleotide sequence is SEQ ID NO:35; a 15 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:36; a 19 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:37; a 107 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:38; a 27 base pair wherein the resulting nucleotide sequence is SEQ ID NO:39; a 9 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:40; a 9 base pair deletion and a 10 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:41; a 1,025 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:42; a l l base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:43; a 7 base pair deletion and a l l base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:44; an inversion wherein the resulting nucleotide sequence is SEQ ID NO:45; a 1,055 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:46; a l l base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:47; a 16 base pair indel and a 9 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:48; a 122 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:49; a 65 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:50; an inversion wherein the resulting nucleotide sequence is SEQ ID NO:51; or a 4 base pair deletion and a 36 base pair deletion wherein the resulting nucleotide sequence is SEQ ID NO:52; and two inversions wherein the resulting nucleotide sequence is SEQ ID NO:53.
[0124] Other targeted modifications may be made to generate novel alleles in the soybean TDF1- 1 gene and homologs thereof. SEQ ID NOs:2 and 55-64 are each the TDFl-like protein homologs found in Arabidopsis thaliana (SEQ ID NO:2), Arachis hypogaea (SEQ ID NO:55), Cicer
arietinum (SEQ ID NO:56), Vitis vinifera (SEQ ID NO:57), Cucumis sativus (SEQ ID NO:58), Brassica napus (SEQ ID NO:59), Gossypium hirsutum (SEQ ID NO:60), Solarium lycopersicum (SEQ ID N0:61), Zea mays (SEQ ID NO:62), Sorghum bicolor (SEQ ID NO:63), and Oryza sativa subsp. japonica (SEQ ID NO:64).
[0125] In a further aspect, the present disclosure provides a modified soybean plant, or plant part thereof, comprising a mutant allele of the GmTDFl-1 gene, wherein the mutant allele comprises one or more junction sequences, wherein the junction sequences are at least 30, at least 60, or at least 100 nucleotides at the junction site. As used herein, a “junction” or “junction site” is the connection point between the nucleotide sequences at the site of an insertion, deletion, substitution, duplication, or inversion. In the case of a deletion, the junction is the connection point at the site of the deletion of the sequences that previously flanked the deletion. For example, in the case of the 30 base pair deletion from nucleotide 1539 to nucleotide 1568, as compared to reference sequence SEQ ID NO:4 described herein, the junction would be between nucleotide 1538 and nucleotide 1569. In the case of an insertion, substitution, or inversion, the junction is the connection point between the inserted, inverted, or substituted sequence and the flanking DNA sequences. In the case of an insertion, substitution, or inversion, one junction is found at the 5’ end of the insertion, substitution or inversion, and another junction is found at the 3’ end of the insertion, substitution, or inversion. A “junction sequence” refers to a DNA sequence of any length that spans a junction. A junction sequence can comprise at least 10 nucleotides, at least, 15 nucleotides, at least 20 nucleotides, at least 25 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, or more.
II. Constructs for Genome Editing
[0126] Recombinant DNA constructs and vectors are provided comprising a polynucleotide sequence encoding a site-specific nuclease, such as a zinc-finger nuclease (ZFN), a meganuclease, an RNA-guided endonuclease, a TALE-endonuclease (TALEN), a recombinase, or a transposase, wherein the coding sequence is operably linked to a plant expressible promoter. For RNA-guided endonucleases, recombinant DNA constructs and vectors are further provided comprising a polynucleotide sequence encoding a guide RNA, wherein the guide RNA comprises a guide
sequence of sufficient length having a percent identity or complementarity to a target site within the genome of a plant, such as at or near a targeted GmTDFl-1 gene. A polynucleotide sequence of a recombinant DNA construct and vector that encodes a site-specific nuclease or a guide RNA may be operably linked to a plant expressible promoter, such as an inducible promoter, a constitutive promoter, a tissue- specific promoter, etc.
[0127] According to some embodiments, a recombinant DNA construct or vector may comprise a first polynucleotide sequence encoding a site- specific nuclease and a second polynucleotide sequence encoding a guide RNA that may be introduced into a plant cell together via plant transformation techniques. Alternatively, two recombinant DNA constructs or vectors may be provided including a first recombinant DNA construct or vector and a second DNA construct or vector that may be introduced into a plant cell together or sequentially via plant transformation techniques, wherein the first recombinant DNA construct or vector comprises a polynucleotide sequence encoding a site-specific nuclease and the second recombinant DNA construct or vector comprises a polynucleotide sequence encoding a guide RNA. According to some embodiments, a recombinant DNA construct or vector comprising a polynucleotide sequence encoding a sitespecific nuclease may be introduced via plant transformation techniques into a plant cell that already comprises (or is transformed with) a recombinant DNA construct or vector comprising a polynucleotide sequence encoding a guide RNA. Alternatively, a recombinant DNA construct or vector comprising a polynucleotide sequence encoding a guide RNA may be introduced via plant transformation techniques into a plant cell that already comprises (or is transformed with) a recombinant DNA construct or vector comprising a polynucleotide sequence encoding a sitespecific nuclease. According to yet further embodiments, a first plant comprising (or transformed with) a recombinant DNA construct or vector comprising a polynucleotide sequence encoding a site-specific nuclease may be crossed with a second plant comprising (or transformed with) a recombinant DNA construct or vector comprising a polynucleotide sequence encoding a guide RNA. Such recombinant DNA constructs or vectors may be transiently transformed into a plant cell or stably transformed or integrated into the genome of a plant cell.
[0128] In an aspect, vectors comprising polynucleotides encoding a site-specific nuclease, and optionally one or more, two or more, three or more, or four or more gRNAs are provided to a plant cell by transformation methods known in the art (e.g., without being limiting, particle bombardment, PEG-mediated protoplast transfection or Agro^acteriuzn-mediated transformation).
In an aspect, vectors comprising polynucleotides encoding a Cpfl nuclease, and optionally one or more, two or more, three or more, or four or more gRNAs arc provided to a plant cell by transformation methods known in the art (e.g., without being limiting, particle bombardment, PEG-mediated protoplast transfection or Agrobacterium-mediated transformation). In another aspect, vectors comprising polynucleotides encoding a Cpfl and, optionally one or more, two or more, three or more, or four or more crRNAs are provided to a cell by transformation methods known in the art (e.g., without being limiting, viral transfection, particle bombardment, PEG- mediated protoplast transfection or Agrobacterium-mediated transformation).
[0129] Several site-specific nucleases, such as zine-finger nucleases, TALENs, meganucleases, and recombinases, are not RNA-guided and instead rely on their protein structure to determine their target site for causing the DSB or nick, or they are fused, tethered or attached to a DNA- binding protein domain or motif. The protein structure of the site-specific nuclease (or the fused/attached/tethered DNA binding domain) may target the site-specific nuclease to the target site. According to many of these embodiments, non-RNA-guided site-specific nucleases, such as zine-finger nucleases, TALENs, meganucleases, and recombinases, may be designed, engineered and constructed according to known methods to target and bind to a target site at or near the genomic locus of an endogenous gene of a plant, such as the GmTDFl-1 gene in soybean, to create a DSB or nick at such genomic locus to knockout or knockdown expression of the GmTDFl-1 gene via repair of the DSB or nick. For example, an engineered site-specific nuclease, such as a recombinase, zinc finger nuclease (ZFN), meganuclease, or TALEN, may be designed to target and bind to a target site within the genome of a plant corresponding to a sequence within SEQ ID NO:4, or its complementary sequence, to create a DSB or nick at the genomic locus for the GmTDFl-1 gene, which may then lead to the creation of a mutation or insertion of a sequence at the site of the DSB or nick, through cellular repair mechanisms, which may be guided by a donor molecule or template.
[0130] In some embodiments, the site-specific nuclease may be a zine-finger nuclease. Zinefinger nucleases (ZFN) are synthetic proteins consisting of an engineered zinc finger DNA-binding domain fused to a cleavage domain (or a cleavage half-domain), which may be derived from a restriction endonuclease (e.g., FokI). The DNA binding domain may be canonical (C2H2) or non- canonical (e.g., C3H or C4). The DNA-binding domain can comprise one or more zinc fingers (e.g., 2, 3, 4, 5, 6, 7, 8, 9 or more zinc fingers) depending on the target site but may typically be
composed of 3-4 (or more) zinc-fingers. Multiple zinc fingers in a DNA-binding domain may be separated by linker scqucncc(s). ZFNs can be designed to cleave almost any stretch of doublestranded DNA by modification of the zinc finger DNA-binding domain. ZFNs form dimers from monomers composed of a non-specific DNA cleavage domain (e.g., derived from the FokI nuclease) fused to a DNA-binding domain comprising a zinc finger array engineered to bind a target site DNA sequence. The amino acids at positions -1, +2, +3, and +6 relative to the start of the zinc finger a-helix, which contribute to site-specific binding to the target site, can be changed and customized to fit specific target sequences. The other amino acids may form a consensus backbone to generate ZFNs with different sequence specificities.
[0131] Methods and rules for designing ZFNs for targeting and binding to specific target sequences are known in the art. See, e.g., U.S. Patent App. Pub. Nos. 2005/0064474, 2009/0117617, and 2012/0142062. The Fokl nuclease domain may require dimerization to cleave DNA and therefore two ZFNs with their C-terminal regions are needed to bind opposite DNA strands of the cleavage site (separated by 5-7 bp). The ZFN monomer can cut the target site if the two-ZF-binding sites are palindromic. A ZFN, as used herein, is broad and includes a monomeric ZFN that can cleave double stranded DNA without assistance from another ZFN. The term ZFN may also be used to refer to one or both members of a pair of ZFNs that are engineered to work together to cleave DNA at the same site. Because the DNA-binding specificities of zinc finger domains can be re-engineered using one of various methods, customized ZFNs can theoretically be constructed to target nearly any target sequence (e.g., at or near a gene in a plant genome). Publicly available methods for engineering zinc finger domains include Context-dependent Assembly (CoDA), Oligomerized Pool Engineering (OPEN), and Modular Assembly.
[0132] In some embodiments, the site-specific nuclease may be a TALEN. TALENs are artificial restriction enzymes generated by fusing the TALE DNA binding domain to a nuclease domain. In some aspects, the nuclease is selected from a group consisting of PvuII, MutH, TevI, Fokl, Alwl, Mlyl, Sbfl, Sdal, StsI, CleDORF, Clo051, and Pept071. When each member of a TALEN pair binds to the DNA sites flanking a target site, the Fokl monomers dimerize and cause a doublestranded DNA break at the target site. The term TALEN, as used herein, is broad and includes a monomeric TALEN that can cleave double- stranded DNA without assistance from another TALEN. The term TALEN also refers to one or both members of a pair of TALENs that work together to cleave DNA at the same site.
[0133] Besides the wild-type Fokl cleavage domain, variants of the FokI cleavage domain with mutations have been designed to improve cleavage specificity and cleavage activity. The Fokl domain functions as a dimer, requiring two constructs with unique DNA binding domains for sites in the target genome with proper orientation and spacing. Both the number of amino acid residues between the TALEN DNA binding domain and the Fokl cleavage domain and the number of bases between the two individual TALEN binding sites are parameters for achieving high levels of activity. PvuII, MutH, and TevI cleavage domains are useful alternatives to Fokl and Fokl variants for use with TALEs. PvuII functions as a highly specific cleavage domain when coupled to a TALE (see Yank et al., PLoS One 8:e82539, 2013). MutH is capable of introducing strand- specific nicks in DNA (see Gabsalilow et al., Nucleic Acids Research. 41:e83, 2013). TevI introduces doublestranded breaks in DNA at targeted sites (see Beurdeley et al., Nature Communications 4:1762, 2013).
[0134] TALEs can be engineered to bind practically any DNA sequence, such as at or near the genomic locus of a gene in a plant. TALE has a central DNA-binding domain composed of 13-28 repeat monomers of 33-34 amino acids. The amino acids of each monomer are highly conserved, except for hypervariable amino acid residues at positions 12 and 13. The two variable amino acids are called repeat- variable diresidues (RVDs). The amino acid pairs NI, NG, HD, and NN of RVDs preferentially recognize adenine, thymine, cytosine, and guanine/adenine, respectively, and modulation of RVDs can recognize consecutive DNA bases. This simple relationship between amino acid sequence and DNA recognition has allowed for the engineering of specific DNA binding domains by selecting a combination of repeat segments containing the appropriate RVDs.
[0135] The relationship between amino acid sequence and DNA recognition of the TALE binding domain allows for designable proteins. Software programs such as DNAWorks can be used to design TALE constructs. Other methods of designing TALE constructs are known to those of skill in the art. See Doyle et al. (Nucleic Acids Research 40:W 117-122, 2012); Cermak et al. (Nucleic Acids Research 39:e82, 2011); and tale-nt.cac.comell.edu/about.
[0136] In some embodiments, the site-specific nuclease may be a meganuclease. Meganucleases, which are commonly identified in microbes, such as the LAGLIDADG family of homing endonucleases, are unique enzymes with high activity and long recognition sequences (> 14 bp) resulting in site-specific digestion of target DNA. Engineered versions of naturally occurring
meganucleases typically have extended DNA recognition sequences (for example, 14 to 40 bp). The engineering of mcganuclcascs can be more challenging than ZFNs and TALENs because the DNA recognition and cleavage functions of meganucleases are intertwined in a single domain. Specialized methods of mutagenesis and high-throughput screening have been used to create novel meganuclease variants that recognize unique sequences and possess improved nuclease activity. In some embodiments, the site-specific nuclease may be a recombinase. Non-limiting examples of recombinases that may be used include a serine recombinase attached to a DNA recognition motif, a tyrosine recombinase attached to a DNA recognition motif, or any recombinase enzyme known in the art attached to a DNA recognition motif.
[0137] In certain embodiments, the site-specific nuclease is a recombinase or transposase, which may be a DNA transposase or recombinase attached or fused to a DNA binding domain. Nonlimiting examples of recombinases include a tyrosine recombinase selected from the group consisting of a Cre recombinase, a Gin recombinase, a Flp recombinase, and a Tnpl recombinase attached to a DNA recognition motif provided herein. In one aspect of the present disclosure, a Cre recombinase or a Gin recombinase provided herein is tethered to a zine-finger DNA-binding domain, a TALE DNA-binding domain, or a Cas9 nuclease. In another aspect, a serine recombinase selected from the group consisting of a PhiC31 integrase, an R4 integrase, and a TP- 901 integrase may be attached to a DNA recognition motif provided herein. In yet another aspect, a DNA transposase selected from the group consisting of a TALE-piggyBac and TALE-Mutator may be attached to a DNA binding domain provided herein.
[0138] As used herein, a “gene” refers to a nucleic acid sequence forming a genetic and functional unit and coding for one or more sequence-related RNA and/or polypeptide molecules. A gene generally contains a coding region operably linked to appropriate regulatory sequences that regulate the expression of a gene product (e.g., a polypeptide or a functional RNA). A gene can have various sequence elements, including, but not limited to, a promoter, an untranslated region(s) (UTR), exons, introns, and other upstream or downstream regulatory sequences.
[0139] As used herein, “locus” is a chromosomal locus or region where a polymorphic nucleic acid, trait determinant, gene, or marker is located. A “locus” can be shared by two homologous chromosomes to refer to their corresponding locus or region. As used herein, “allele” refers to an alternative nucleic acid sequence of a gene or at a particular locus (e.g., a nucleic acid sequence of
a gene or locus that is different than other alleles for the same gene or locus). Such an allele can be considered (i) wild-type or (ii) mutant if one or more mutations or edits arc present in the nucleic acid sequence of the mutant allele relative to the wild-type allele. A mutant allele for a gene may have a reduced or eliminated activity or expression level for the gene relative to the wild- type allele. For diploid organisms such as soybean, a first allele can occur on one chromosome, and a second allele can occur at the same locus on a second homologous chromosome.
[0140] As used herein, the term “homozygous” refers to a genotype comprising two identical alleles at a given locus in a diploid genome. Given that soybean is a diploid organism, CRISPR- mediated gene editing can result in biallelic edits (that is, different edits are made to the same locus on corresponding homologous chromosomes) resulting in a genotype comprising two nonidentical mutant alleles at a given locus in a diploid genome in Ro plants. The non-identical mutant alleles may differ, for example, based on having different additions, deletions, and/or substitutions of one or more nucleotides relative to each other, with the additions, deletions, and/or substitutions of each being sufficiently severe to cause a loss of function of the TDF1 protein encoded by each. When used in the context of edited alleles, plants comprising such genotypes may also be referred to as comprising a heteroallelic combination or transheterozygous edits. Generally, “complementation” or “complement” refers to the ability of two non-identical mutations to restore the wild-type phenotype when present in the same plant. Complementation can be full (wild-type phenotype) or partial (milder mutant phenotype). However, when used in the context of biallelic edits, “complementation” or “complement” describes when non-identical, recessive mutant alleles act in a complementary manner to produce the recessive phenotype. In some embodiments, said recessive phenotype is male sterility. Conversely, a monoallelic edit includes events in which an edit is made to only one allele of a locus (z'.e., a modification to the endogenous TDF1 gene in only one of the two homologous chromosomes) and can result in a cell that is heterozygous for the targeted TDF1 modification. As used herein, “heterozygous” describes a genotype comprising a mutant/edited allele and a wild-type allele at a given locus in a diploid genome.
[0141] As used herein, a “wild-type gene” or “wild-type allele” refers to a gene or allele having a sequence or genotype that is most common in a particular plant species, or another sequence or genotype having only natural variations, polymorphisms, or other silent mutations relative to the most common sequence or genotype that do not significantly impact the expression and activity of the gene or allele. Indeed, a “wild-type” gene or allele contains no variation, polymorphism, or
any other type of mutation that substantially affects the normal function, activity, expression, or phenotypic consequence of the gene or allele relative to the most common sequence or genotype. As used herein, a “native copy” of a gene refers to a gene that originates from within a given organism, cell, tissue, genome, or chromosome that was not previously modified by human action. Similarly, a “native protein” refers to a protein encoded by a native gene.
[0142] As used herein “pleiotropy” refers to the phenomenon where a single gene (i.e. pleiotropic gene) or locus affects two or more phenotypic traits. As used herein, a “pleiotropic effect” refers to the influence of a mutation in a pleiotropic gene on the traits associated with said gene.
[0143] In general, the term "variant" refers to molecules with some differences, generated synthetically or naturally, in their nucleotide or amino acid sequences as compared to a reference (native) polynucleotides or polypeptides, respectively. These differences include substitutions, insertions, deletions or any desired combinations of such changes in a native polynucleotide or amino acid sequence.
[0144] As used herein, the term “expression” refers to the biosynthesis of a gene product, and typically the transcription and/or translation of a nucleotide sequence, such as an endogenous gene, a heterologous gene, a transgene or an RNA and/or protein coding sequence, in a cell, tissue, organ, or organism, such as a plant, plant part or plant cell, tissue or organ.
[0145] The term “recombinant” in reference to a polynucleotide (DNA or RNA) molecule, protein, construct, vector, etc., refers to a polynucleotide or protein molecule or sequence that is man-made and not normally found in nature, and/or is present in a context in which it is not normally found in nature, including a polynucleotide (DNA or RNA) molecule, protein, construct, etc., comprising a combination of two or more polynucleotide or protein sequences that would not naturally occur together in the same manner without human intervention, such as a polynucleotide molecule, protein, construct, etc., comprising at least two polynucleotide or protein sequences that are operably linked but heterologous with respect to each other. For example, the term “recombinant” can refer to any combination of two or more DNA or protein sequences in the same molecule (e.g., a plasmid, construct, vector, chromosome, protein, etc.) where such a combination is man-made and not normally found in nature. As used in this definition, the phrase “not normally found in nature” means not found in nature without human introduction. A recombinant polynucleotide or protein molecule, construct, etc., can comprise polynucleotide or protein sequence(s) that is/are (i)
separated from other polynucleotide or protein sequence(s) that exist in proximity to each other in nature, and/or (ii) adjacent to (or contiguous with) other polynucleotide or protein scqucncc(s) that are not naturally in proximity with each other. Such a recombinant polynucleotide molecule, protein, construct, etc., can also refer to a polynucleotide or protein molecule or sequence that has been genetically engineered and/or constructed outside of a cell. For example, a recombinant DNA molecule can comprise any engineered or man-made plasmid, vector, etc., and can include a linear or circular DNA molecule. Such plasmids, vectors, etc., can contain various maintenance elements including a prokaryotic origin of replication and selectable marker, as well as one or more transgenes or expression cassettes perhaps in addition to a plant selectable marker gene, etc. The term “operably linked” refers to a functional linkage between a promoter or other regulatory element and an associated transcribable DNA sequence or coding sequence of a gene (or transgene), such that the promoter, etc., operates or functions to initiate, assist, affect, cause, and/or promote the transcription and expression of the associated transcribable DNA sequence or coding sequence, at least in certain cell(s), tissue(s), developmental stage(s), and/or condition(s).
[0146] Reference in this application to an “isolated DNA molecule” or an “isolated polynucleotide,” or an equivalent term or phrase, is intended to mean that the DNA molecule or polynucleotide is one that is present alone or in combination with other compositions, but not within its natural environment. For example, nucleic acid elements such as a coding sequence, intron sequence, untranslated leader sequence, promoter sequence, transcriptional termination sequence, and the like, that are naturally found within the DNA of the genome of an organism are not considered to be “isolated” so long as the element is within the genome of the organism and at the location within the genome in which it is naturally found. However, each of these elements, and subparts of these elements, would be “isolated” within the scope of this disclosure so long as the element is not within the genome of the organism and at the location within the genome in which it is naturally found. Similarly, a nucleotide sequence encoding a protein or any naturally occurring variant of that protein would be an isolated nucleotide sequence so long as the nucleotide sequence was not within the DNA of the organism in which the sequence encoding the protein is naturally found. A synthetic nucleotide sequence encoding the amino acid sequence of the naturally occurring protein would be considered to be isolated for the purposes of this disclosure. For the purposes of this disclosure, any transgenic nucleotide sequence, i.e., the nucleotide sequence of the DNA inserted into the genome of the cells of a plant or bacterium, or present in
an extrachromosomal vector, would be considered to be an isolated nucleotide sequence whether it is present within the plasmid or similar structure used to transform the cells, within the genome of the plant or bacterium, or present in detectable amounts in tissues, progeny, biological samples or commodity products derived from the plant or bacterium.
[0147] As commonly understood in the art, the term “promoter” can generally refer to a DNA sequence that contains an RNA polymerase binding site, transcription start site, and/or TATA box and assists or promotes the transcription and expression of an associated transcribable polynucleotide sequence and/or gene (or transgene). A promoter can be synthetically produced, varied, or derived from a known or naturally occurring promoter sequence or other promoter sequence. A promoter can also include a chimeric promoter comprising a combination of two or more heterologous sequences. A promoter of the present disclosure can thus include variants or fragments of promoter sequences that are similar in composition, but not identical to, other promoter sequence(s) known or provided herein. A promoter provided herein, or variant or fragment thereof, may comprise a “minimal promoter” which provides a basal level of transcription and is comprised of a TATA box or equivalent DNA sequence for recognition and binding of the RNA polymerase II complex for initiation of transcription. A promoter can be classified according to a variety of criteria relating to the pattern of expression of an associated coding or transcribable sequence or gene (including a transgene) operably linked to the promoter, such as constitutive, developmental, tissue-specific, inducible, etc. Promoters that drive expression in all or most tissues of the plant are referred to as “constitutive” promoters. Promoters that drive expression during certain periods or stages of development are referred to as “developmental” promoters. Promoters that drive enhanced expression in certain tissues of the plant relative to other plant tissues are referred to as “tissue-enhanced” or “tissue-preferred” promoters. Thus, a “tissue-preferred” promoter causes relatively higher or preferential expression in a specific tissue(s) of the plant, but with lower levels of expression in other tissue(s) of the plant. Promoters that express within a specific tissue(s) of the plant, with little or no expression in other plant tissues, are referred to as “tissue-specific” promoters. An “inducible” promoter is a promoter that initiates transcription in response to an environmental stimulus such as cold, drought or light, or other stimuli, such as wounding or chemical application. A promoter can also be classified in terms of its origin, such as being heterologous, homologous, chimeric, synthetic, etc.
[0148] As used herein, a “plant-expressible promoter” refers to a promoter that can initiate, assist, affect, cause, and/or promote the transcription and expression of its associated transcribablc DNA sequence, coding sequence or gene in a plant cell or tissue.
[0149] The term “heterologous” in reference to a promoter or other regulatory sequence in relation to an associated polynucleotide sequence (e.g., a transcribable DNA sequence or coding sequence or gene) is a promoter or regulatory sequence that is not operably linked to such associated polynucleotide sequence in nature without human introduction - e.g., the promoter or regulatory sequence has a different origin relative to the associated polynucleotide sequence and/or the promoter or regulatory sequence is not naturally occurring in a plant species to be transformed with the promoter or regulatory sequence.
[0150] As used herein, an “endogenous gene” or an “endogenous locus” refers to a gene or locus at its natural and original chromosomal location. As used herein, the “endogenous TDF1 gene” refers to the TDF1 genic locus at its original chromosomal location.
[0151] As used herein, in the context of a protein-coding gene, an “exon” refers to a segment of a DNA or RNA molecule containing information coding for a protein or polypeptide sequence.
[0152] As used herein, an “intron” of a gene refers to a segment of a DNA or RNA molecule, which does not contain information coding for a protein or polypeptide, and which is first transcribed into an RNA sequence but then spliced out from a mature RNA molecule.
[0153] As used herein, an “untranslated region (UTR)” of a gene refers to a segment of an RNA molecule or sequence (e.g., a mRNA molecule) expressed from a gene (or transgene) but excluding the exon and intron sequences of the RNA molecule. An “untranslated region (UTR)” also refers to a DNA segment or sequence encoding such a UTR segment of an RNA molecule. An untranslated region can be a 5'-UTR or a 3'-UTR depending on whether it is located at the 5' or 3' end of a DNA or RNA molecule or sequence relative to a coding region of the DNA or RNA molecule or sequence (/.<?., upstream (5') or downstream (3') of the exon and intron sequences, respectively).
[0154] As used herein, “upstream” refers to a nucleic acid sequence that is positioned before the 5' end of a linked nucleic acid sequence. As used herein, “downstream” refers to a nucleic acid sequence is positioned after the 3' end of a linked nucleic acid sequence. As used herein, “5'” refers
to the start of a coding DNA sequence or the beginning of an RNA molecule. As used herein, “3'” refers to the end of a coding DNA sequence or the end of an RNA molecule. It will be appreciated that an “inversion” refers to reversing the orientation of a given polynucleotide sequence.
[0155] As used herein, a “transcription termination sequence” refers to a nucleic acid sequence containing a signal that triggers the release of a newly synthesized transcript RNA molecule from an RNA polymerase complex and marks the end of transcription of a gene or locus.
[0156] As used herein, a "homolog" or "homologues" means a protein in a group of proteins that perform the same biological function, for example, proteins that belong to the same TDFl-like protein family and that provide a common enhanced trait in modified plants of this disclosure. Homologs are expressed by homologous genes. With reference to homologous genes, homologs include orthologs, for example, genes expressed in different species that evolved from common ancestral genes by speciation and encode proteins retain the same function, but do not include paralogs, z.e., genes that arc related by duplication but have evolved to encode proteins with different functions. Homologous genes include naturally occurring alleles and artificially-created variants.
[0157] Homologs are inferred from sequence similarity, by comparison of protein sequences, for example, manually or by use of a computer-based tool. For optimal alignment of sequences to calculate their percent identity, various pair-wise or multiple sequence alignment algorithms and programs are known in the ail, such as ClustalW or Basic Local Alignment Search Tool® (BLAST), etc., that can be used to compare the sequence identity or similarity between two or more nucleotide or protein sequences. BLAST, can also be used, for example to search query protein sequences of a base organism against a database of protein sequences of various organisms, to find similar sequences. The generated summary Expectation value (E-value) can be used to measure the level of sequence similarity. Because a protein hit with the lowest E-value for a particular organism may not necessarily be an ortholog or be the only ortholog, a reciprocal query is used to filter hit sequences with significant E-values for ortholog identification. The reciprocal query entails search of the significant hits against a database of protein sequences of the base organism. A hit can be identified as an ortholog, when the reciprocal query's best hit is the query protein itself or a paralog of the query protein. With the reciprocal query process orthologs are
further differentiated from paralogs among all the homologs, which allows for the inference of functional equivalence of genes.
[0158] The terms “percent identity,” “% identity” or “percent identical” or “percent sequence identity” as used herein in reference to two or more nucleotide or protein sequences is calculated by (i) comparing two optimally aligned sequences (nucleotide or protein) over a window of comparison, (ii) determining the number of positions at which the identical nucleic acid base (for nucleotide sequences) or amino acid residue (for proteins) occurs in both sequences to yield the number of matched positions, (iii) dividing the number of matched positions by the total number of positions in the window of comparison, and then (iv) multiplying this quotient by 100% to yield the percent identity. If the “percent identity” or “percent sequence identity” is being calculated in relation to a reference sequence without a particular comparison window being specified, then the percent identity is determined by dividing the number of matched positions over the region of alignment by the total length of the reference sequence. Accordingly, for purposes of the present application, when two sequences (query and subject) are optimally aligned (with allowance for gaps in their alignment), the “percent identity” for the query sequence is equal to the number of identical positions between the two sequences divided by the total number of positions in the query sequence over its length (or a comparison window), which is then multiplied by 100%. When percentage of sequence identity is used in reference to proteins it is recognized that residue positions which are not identical often differ by conservative amino acid substitutions, where amino acid residues are substituted for other amino acid residues with similar chemical properties (e.g., charge or hydrophobicity) and therefore do not change the functional properties of the molecule. When sequences differ in conservative substitutions, the percent sequence identity can be adjusted upwards to correct for the conservative nature of the substitution. Sequences that differ by such conservative substitutions are said to have “sequence similarity” or “similarity.” Sequences having a percent identity to a base sequence may exhibit the activity of the base sequence.
[0159] Degeneracy of the genetic code provides the possibility to substitute at least one base of the protein encoding sequence of a gene with a different base without causing the amino acid sequence of the polypeptide produced from the gene to be changed. When optimally aligned, homolog proteins, or their corresponding nucleotide sequences, have typically at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 92%,
at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or even at least about 99.5% identity over the full length of a protein or its corresponding nucleotide sequence identified as being associated with imparting a male fertility phenotype when expressed in plant cells. According to embodiments of the present invention, a TDF1 gene or homolog thereof encodes a protein that has a functional domain with at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the functional domain of SEQ ID NO:5. Examples of homologs of the GmTDFl-1 protein (SEQ ID NO:5) include, but are not limited to, the sequences of SEQ ID NOs:55-64. Further homologs and paralogs may readily be identified using the BLAST procedure described above.
[0160] As used herein, “alignment” refers to a number of nucleotide bases or amino acid residue sequences aligned by lengthwise comparison so that components in common (i.e. , nucleotide bases or amino acid residues at corresponding positions) may be visually and readily identified. The fraction or percentage of components in common is related to the homology or identity between the sequences. An alignment, such as that shown in FIGS. 11 A- HE, may be used to identify conserved domains and relatedness within these domains. An alignment may suitably be determined by means of computer programs known in the ail. Specialist databases also exist for the identification of domains, for example, SMART (Schultz et al., Proc. Natl. Acad. Sci. USA 95:5857-5864, 1998; Letunic et al., Nucleic Acids Res. 30: 242-244, 2002), InterPro (Mulder et al., Nucleic Acids Res. 31:315-318, 2002), PROSITE (Bairoch and Bucher, Nucleic Acids Res., 22:3583-3589, 1994; Hofmann et al., Nucleic Acids Res. 27:215-219, 1999), or Pfam (Bateman et al., Nucleic Acids Res. 30(l):276-280, 2002). A set of tools for in silico analysis of protein sequences is available on the ExPASY proteomics server (hosted by the Swiss Institute of Bioinformatics (Gasteiger et al., Nucleic Acids Res. 31:3784-3788, 2003).
[0161] Homologs of the proteins described herein are also identifiable by the presence of a conserved functional domain(s). The term “domain” refers to a set of amino acids conserved at specific positions along an alignment of sequences of evolutionarily related proteins. While amino acids at other positions can vary between homologs, amino acids that are highly conserved at specific positions indicate amino acids that are essential in the structure, stability, or activity of a protein. Identified by their high degree of conservation in aligned sequences of a family of protein homologs, they can be used as identifiers to determine if any polypeptide in question belongs to a
previously identified polypeptide family (in this case, the proteins useful in the methods of the invention and nucleic acids encoding the same as defined herein). TDF1 is believed to be a member of the MYB transcription factor family which is defined by a highly conserved MYB DNA-binding domain near the N-terminus of the protein. Outside this domain however MYB protein sequences exhibit a low degree of sequence conservation, which leads to an overall lower sequence identity across homologs. A conserved functional domain with respect to presently disclosed polypeptides refers to a domain within a polypeptide family that exhibits a higher degree of sequence homology, such as at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to a conserved domain of a polypeptide of the invention (e.g., the SANT/Myb domain of SEQ ID NO:5). Sequences that possess or encode for conserved domains that meet these criteria of percentage identity, and that have comparable biological activity to the present polypeptide sequences, thus being members of the MYB transcription factor family, are encompassed by the invention.
[0162] The terms “percent complementarity” or “percent complementary,” as used herein in reference to two nucleotide sequences, is similar to the concept of percent identity but refers to the percentage of nucleotides of a query sequence that optimally base-pair or hybridize to nucleotides of a subject sequence when the query and subject sequences are linearly arranged and optimally base paired without secondary folding structures, such as loops, stems or hairpins. Such a percent complementarity may be between two DNA strands, two RNA strands, or a DNA strand and an RNA strand. The “percent complementarity” is calculated by (i) optimally base-pairing or hybridizing the two nucleotide sequences in a linear and fully extended arrangement (z'.e., without folding or secondary structures) over a window of comparison, (ii) determining the number of positions that base-pair between the two sequences over the window of comparison to yield the number of complementary positions, (iii) dividing the number of complementary positions by the total number of positions in the window of comparison, and (iv) multiplying this quotient by 100% to yield the percent complementarity of the two sequences. Optimal base pairing of two sequences may be determined based on the known pairings of nucleotide bases, such as G-C, A-T, and A-U, through hydrogen bonding. If the “percent complementarity” is being calculated in relation to a reference sequence without specifying a particular comparison window, then the percent identity is determined by dividing the number of complementary positions between the two linear
sequences by the total length of the reference sequence. Thus, for purposes of the present disclosure, when two sequences (query and subject) arc optimally base-paired (with allowance for mismatches or non-base-paired nucleotides but without folding or secondary structures), the “percent complementarity” for the query sequence is equal to the number of base-paired positions between the two sequences divided by the total number of positions in the query sequence over its length (or by the number of positions in the query sequence over a comparison window), which is then multiplied by 100%.
[0163] As used herein, a “fragment” of a polynucleotide refers to a sequence comprising at least about 50, at least about 75, at least about 95, at least about 100, at least about 125, at least about 150, at least about 175, at least about 200, at least about 225, at least about 250, at least about 275, at least about 300, at least about 500, at least about 600, at least about 700, at least about 750, at least about 800, at least about 900, or at least about 1000 contiguous nucleotides, or longer, of a DNA molecule or protein as disclosed herein. Methods for producing such fragments from a starting promoter molecule are well known in the art. Fragments of a DNA molecule or protein may exhibit the activity of the DNA molecule or protein from which they are derived.
[0164] According to another aspect, the present disclosure provides methods for altering a phenotype, such as inducing male sterility in a plant comprising:(a) modifying the genome of a plant cell by: (i) identifying an endogenous gene of the plant corresponding to a TDF1 gene, such as the GmTDFl-1 gene described herein, and its homologs, and (ii) modifying the sequence of the endogenous gene in the plant cell via targeted mutagenesis to modify the expression level of the endogenous gene; and (b) regenerating or developing a plant from the plant cell. Various TDF1 genes and proteins from different plant species may be identified and considered TDF1 homologs or orthologs for use in the present disclosure if they have a similar nucleic acid and/or protein sequence and/or share conserved amino acids and/or structural domain(s) with at least one known TDF1 gene or protein.
[0165] A plant selectable marker transgene in a transformation vector or construct of the present disclosure may be used to assist in the selection of transformed cells or tissue due to the presence of a selection agent, such as an antibiotic or herbicide, wherein the plant selectable marker transgene provides tolerance or resistance to the selection agent. Thus, the selection agent may bias or favor the survival, development, growth, proliferation, etc., of transformed cells expressing
the plant selectable marker gene, such as to increase the proportion of transformed cells or tissues in the Ro plant. Commonly used plant selectable marker genes include, for example, those conferring tolerance or resistance to antibiotics, such as kanamycin and paromomycin (nptll), hygromycin B (aph IV), streptomycin or spectinomycin (aadA) and gentamycin (aac3 and aacC4), or those conferring tolerance or resistance to herbicides such as glufosinate (bar or pat), dicamba (DMO) and glyphosate (proA or EPSPS). Plant screenable marker genes may also be used, which provide an ability to visually screen for transformants, such as luciferase or green fluorescent protein (GFP), or a gene expressing a beta glucuronidase or uidA gene (GUS) for which various chromogenic substrates are known. Plant transformation may also be carried out in the absence of selection during one or more steps or stages of culturing, developing or regenerating transformed explants, tissues, plants and/or plant parts.
III. Transformation Methods
[01661 Methods and compositions are provided for transforming a plant cell, tissue or explant with a recombinant DNA molecule or construct encoding one or more molecules required for targeted genome editing (e.g., guide RNA(s) and/or site-directed nuclease(s)). Suitable methods for transformation of host plant cells include virtually any method by which DNA or RNA can be introduced into a cell (for example, where a recombinant DNA construct is stably integrated into a plant chromosome or where a recombinant DNA construct or an RNA is transiently provided to a plant cell) and are well known in the art. Two effective methods for cell transformation are bacterially-mediated transformation, such as Agro&acterzuzvr-mediated or R/zzzo&zwm-mediated transformation, and microprojectile or particle bombardment-mediated transformation. Microprojectile bombardment methods are illustrated, for example, in U.S. Patent Nos. 5,550,318; 5,538,880; 6,160,208; and 6,399,861. Agrobacterium-mcAa cd transformation methods are described, for example in U.S. Patent No. 5,591,616. Other methods for plant transformation, such as microinjection, electroporation, vacuum infiltration, pressure, sonication, silicon carbide fiber agitation, PEG-mediated transformation, etc., are also known in the art.
[0167] Transformation of plant material is practiced in tissue culture on nutrient media, for example a mixture of nutrients that allow cells to grow in vitro. Recipient cell targets include, but are not limited to, meristem cells, shoot tips, hypocotyls, calli, immature or mature embryos, and gametic cells such as microspores and pollen. Callus can be initiated from tissue sources including,
but not limited to, immature or mature embryos, hypocotyls, seedling apical meristems, microsporcs and the like. Cells containing a transgenic nucleus arc grown into transgenic plants, also referred to as Ro plants. As used herein, “Ro plant” refers to an initial regenerated transformant. As used herein, “Ri seed” refers to seed produced from selfing Ro plants. As used herein, “Ri plant” refers to a plant grown from Ri seed. As used herein, “R2 seed” refers to seed produced from selfing Ri plants. As used herein, “R? plant” refers to a plant grown from R2 seed.
[0168] Any suitable method or technique for transformation of a plant cell known in the art may be used according to present methods. In transformation, DNA is typically introduced into only a small percentage of target plant cells in any one transformation experiment. Marker genes are used to provide an efficient system for identification of those cells that are stably transformed by receiving and integrating a recombinant DNA molecule into their genomes.
[0169] As used herein, the terms “regeneration” and “regenerating” refer to a process of growing or developing a plant from one or more plant cells through one or more culturing steps. Transformed or edited cells, tissues or explants containing a DNA sequence insertion or edit may be grown, developed or regenerated into transgenic plants in culture, plugs, or soil according to methods known in the ail. Certain embodiments of the disclosure therefore relate to methods and constructs for regenerating a plant from a cell with modified genomic DNA resulting from genome editing. The regenerated plant can then be used to propagate additional plants.
[0170] According to an aspect of the present disclosure, regenerated plants or a progeny plant, plant part or seed thereof can be screened or selected based on a marker, trait, or phenotype produced by the edit or mutation, or by the site-directed integration of an insertion sequence, transgene, etc., in the developed or regenerated plant, or a progeny plant, plant part or seed thereof. If a given mutation, edit, trait or phenotype is recessive, one or more generations or crosses (e.g., selfing) from the initial Ro plant may be necessary to produce a plant homozygous for the edit or mutation so the trait or phenotype can be observed. Progeny plants, such as plants grown from Ri seed or in subsequent generations, can be tested for zygosity using any known zygosity assay, such as by using a single nucleotide polymorphism (SNP) assay, DNA sequencing, thermal amplification, or polymerase chain reaction (PCR), and/or Southern blotting that allows for the distinction between heterozygote, homozygote and wild-type plants.
[0171] Methods and techniques are provided for screening for, and/or identifying, cells or plants, etc., for the presence of targeted edits or transgcncs, and selecting cells or plants comprising targeted edits or transgenes, which may be based on one or more phenotypes or traits, or on the presence or absence of a molecular marker or polynucleotide or protein sequence in the cells or plants. As used herein, a “molecular technique” refers to any method known in the fields of molecular biology, biochemistry, genetics, plant biology, or biophysics that involves the use, manipulation, or analysis of a nucleic acid, a protein, or a lipid. Without being limiting, molecular techniques useful for detecting the presence of a modified sequence in a genome include phenotypic screening; molecular marker technologies such as SNP analysis by TaqMan® or Illumina/Infinium technology; Southern blot; PCR (including amplicon sequencing which consists of the generation of one or more unique PCR products across the genomic region of interest for further sequencing analysis, such as using Next-Gen Sequencing techniques known in the art. Sequence data from each sample is then mapped to a reference sequence to identify consensus differences); enzyme-linked immunosorbent assay (ELISA); and sequencing (e.g., Sanger, Illumina®, 454, Pae-Bio, Ion Torrent™). In one aspect, a method of detection provided herein comprises phenotypic screening. In another aspect, a method of detection provided herein comprises SNP analysis. In a further aspect, a method of detection provided herein comprises a Southern blot. In a further aspect, a method of detection provided herein comprises PCR. In a further aspect, a method of detection provided herein comprises amplicon sequencing. In an aspect, a method of detection provided herein comprises ELISA. In a further aspect, a method of detection provided herein comprises determining the sequence of a nucleic acid or a protein. Without being limiting, nucleic acids can be detected using hybridization. Hybridization between nucleic acids is discussed in detail in Sambrook et al. (1989, Molecular Cloning: A Laboratory Manual, 2nd Ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY).
[0172] Nucleic acids can be isolated using techniques routine in the ail. For example, nucleic acids can be isolated using any method including, without limitation, recombinant nucleic acid technology, and/or PCR. General PCR techniques are described, for example in PCR Primer: A Laboratory Manual, Dieffenbach & Dveksler, Eds., Cold Spring Harbor Laboratory Press, 1995. Recombinant nucleic acid techniques include, for example, restriction enzyme digestion and ligation, which can be used to isolate a nucleic acid. Isolated nucleic acids also can be chemically synthesized, either as a single nucleic acid molecule or as a series of oligonucleotides.
[0173] Detection (e.g., of an amplification product, of a hybridization complex, of a polypeptide) can be accomplished using detectable labels that may be attached or associated with a hybridization probe or antibody. The term “label” is intended to encompass the use of direct labels as well as indirect labels. Detectable labels include enzymes, prosthetic groups, fluorescent materials, luminescent materials, bioluminescent materials, and radioactive materials. The screening and selection of modified (e.g., edited) plants or plant cells can be through any methodologies known to those skilled in the art of molecular biology. Examples of screening and selection methodologies include, but are not limited to, Southern analysis, PCR amplification for detection of a polynucleotide (including amplicon sequencing), Northern blots, RNase protection, primerextension, RT-PCR amplification for detecting RNA transcripts, Sanger sequencing, Next Generation sequencing technologies (e.g., Illumina®, PacBio®, Ion Torrent™, etc.) enzymatic assays for detecting enzyme or ribozyme activity of polypeptides and polynucleotides, and protein gel electrophoresis, Western blots, immunoprecipitation, and enzyme-linked immunoassays to detect polypeptides. Other techniques such as in situ hybridization, enzyme staining, and immuno staining also can be used to detect the presence or expression of polypeptides and/or polynucleotides. Methods for performing all of the referenced techniques are known in the art.
[0174] As used herein, the term “polypeptide” refers to a chain of at least two covalently linked amino acids. Polypeptides can be encoded by polynucleotides provided herein. An example of a polypeptide is a protein. Proteins provided herein can be encoded by nucleic acid molecules provided herein. Polypeptides can be purified from natural sources (e.g., a biological sample) by known methods such as DEAE ion exchange, gel filtration, and hydroxyapatite chromatography. A polypeptide also can be purified, for example, by expressing a nucleic acid in an expression vector. In addition, a purified polypeptide can be obtained by chemical synthesis. The extent of purity of a polypeptide can be measured using any appropriate method, e.g., column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis.
[0175] Polypeptides can be detected using antibodies. Techniques for detecting polypeptides using antibodies include enzyme linked immunosorbent assays (ELISAs), Western blots, immunoprecipitations and immunofluorescence. An antibody provided herein can be a polyclonal antibody or a monoclonal antibody. An antibody having specific binding affinity for a polypeptide provided herein can be generated using methods well known in the art. An antibody provided herein can be attached to a solid support such as a microtiter plate using methods known in the art.
[0176] A plant that may be transformed with a recombinant DNA molecule or transformation vector comprising a guide RNA may include a variety of flowering plants or angiosperms, which may be further defined as including various dicotyledonous (dicot) plant species or monocotyledonous (monocot) plant species. A dicot plant could be members of the Fabaceae family (such as legumes), sunflower (Helianthus annuus), safflower (Carthamus tinctorius), sesame (Sesamum spp.), tobacco (Nicotiana tabacum), potato (Solarium tuberosum), cotton (Gossypium barbadense, Gossypium hirsutum), sweet potato (Ipomoea batatas), cassava (Manihot esculenta), coffee (Coffea spp.), tea (Camellia spp.), fruit trees, such as apple (Malus spp.), Primus spp., such as plum, apricot, peach, cherry, etc., pear (Pyrus spp.), fig (Ficus carica), etc., citrus trees (Citrus spp.), cocoa (Theobroma cacao), avocado (Persea americana), olive (Olea europaea), almond (Prunus amygdalus), walnut (Juglans spp.), strawberry (Fragaria spp.), watermelon (Citrullus lanatus), pepper (Capsicum spp.), beet (Beta vulgaris), grape (Vitis, Muscadinia), tomato (Lycopersicon esculentum, Solarium lycopersicum), cucumber (Cucumis sativus), and members of the Brassicaceae family, such as thale cress (Arabidopsis thaliana) and Brassica sp. (e.g., B. napus, B. rapa, B. juncea), particularly those Brassica species useful as sources of seed oil. Legumes and leguminous plants include peas (Pisum sativum) alfalfa (Medicago sativa), barrel clover (Medicago truncatula), pigeon pea (Cajanus cajan) guar (Cyamopsis tetragonoloba), carob (Ceratonia siliqua), fenugreek (Trigonella foenum-graecum), soybean (Glycine max), common bean (Phaseolus vulgaris), cowpea (Vigna unguiculata), mung bean (Vigna radiata), lima bean (Phaseolus lunatus), fava bean (Vicia faba), lentil (Lens culinaris or Lens esculenta), peanut (Arachis hypogaea), licorice (Glycyrrhiza glabra), and chickpea (Cicer arietinum). A monocot plant could be oil palm (Elaeis spp.), coconut (Cocos spp.), banana (Musa spp.), and cereals such as corn (Zea mays), barley (Hordeum vulgare), sorghum (Sorghum bicolor), rice (Oryza sativa), and wheat (Triticum aestivum). Given that the present disclosure may apply to a broad range of plant species, the present disclosure further applies to other botanical structures analogous to pods of leguminous plants, such as bolls, siliques, fruits, nuts, tubers, etc.
IV. Genome Modified Plants
[0177] As used herein, “modified” in the context of a plant, plant seed, plant part, plant cell, and/or plant genome, refers to a plant, plant seed, plant part, plant cell, and/or plant genome comprising an engineered change in the expression level and/or endogenous sequence of one or more genes
of interest relative to a wild-type or control plant, plant seed, plant part, plant cell, and/or plant genome. Indeed, the term “modified” may further refer to a plant, plant seed, plant part, plant cell, and/or plant genome having one or more deletions affecting expression of an endogenous TDF1 gene or homolog introduced through chemical mutagenesis, transposon insertion or excision, or any other known mutagenesis technique, or introduced through genome editing. In an aspect, a modified plant, plant seed, plant part, plant cell, and/or plant genome can comprise one or more transgenes. For clarity, therefore, a modified plant, plant seed, plant part, plant cell, and/or plant genome includes a mutated, edited and/or transgenic plant, plant seed, plant part, plant cell, and/or plant genome having a modified expression level, expression pattern, and/or sequence of a TDF1 gene or homolog relative to a wild-type or control plant, plant seed, plant part, plant cell, and/or plant genome.
[0178] Modified plants, plant parts, seeds, etc., may have been subjected to mutagenesis, genome editing or site-directed integration, genetic transformation, or a combination thereof. Such “modified” plants, plant seeds, plant parts, and plant cells include plants, plant seeds, plant parts, and plant cells that are offspring or derived from “modified” plants, plant seeds, plant parts, and plant cells that retain the molecular change (e.g., change in expression level and/or activity) to the endogenous TDF1 gene, such as the GmTDFl-1 gene. A modified seed provided herein may give rise to a modified plant provided herein. A modified plant, plant seed, plant part, plant cell, or plant genome provided herein may comprise a recombinant DNA construct or vector or genome edit as provided herein. A “modified plant product” may be any product made from a modified plant, plant part, plant cell, or plant chromosome provided herein, or any portion or component thereof.
[0179] Modified plants may be further crossed to themselves or other plants to produce modified plant seeds and progeny. A modified plant may also be prepared by crossing a first plant comprising a DNA sequence or construct or an edit (e.g., a genomic deletion) with a second plant lacking the DNA sequence or construct or edit. For example, a DNA sequence or inversion may be introduced into a first plant line that is amenable to transformation or editing, which may then be crossed with a second plant line to introgress the DNA sequence or edit (e.g., deletion) into the second plant line. Progeny of these crosses can be further backcrossed into the desirable line multiple times, such as through 6 to 8 generations or back crosses, to produce a progeny plant with substantially the same genotype as the original parental line, but for the introduction of the DNA sequence or edit. A modified plant, plant cell, or seed provided herein may be a hybrid plant, plant
cell, or seed. As used herein, a “hybrid” is created by crossing two plants from different varieties, lines, inbrcds, or species, such that the progeny comprises genetic material from each parent. Skilled artisans recognize that higher order hybrids can be generated as well.
[0180] A modified plant, plant part, plant cell, or seed provided herein may be of an elite variety or an elite line. An “elite variety” or an “elite line” refers to a variety that has resulted from breeding and selection for superior agronomic performance.
[0181] As used herein, the term “control plant” (or likewise a “control” plant seed, plant part, plant cell, and/or plant genome) refers to a plant (or plant seed, plant part, plant cell, and/or plant genome) that is used for comparison to a modified plant (or modified plant seed, plant part, plant cell, and/or plant genome) and has the same or similar genetic background (e.g., same parental lines, hybrid cross, inbred line, testers, etc.) as the modified plant (or plant seed, plant part, plant cell, and/or plant genome), except for genome edit(s) (e.g., a deletion) affecting a TDF1 gene. For example, a control plant may be an inbred line that is the same as the inbred line used to make the modified plant, or a control plant may be the product of the same hybrid cross of inbred parental lines as the modified plant, except for the absence in the control plant of any transgenic events or genome edit(s) affecting a TDF1 gene. Similarly, an “unmodified control plant” refers to a plant that shares a substantially similar or essentially identical genetic background as a modified plant, but without the one or more engineered changes to the genome (e.g., mutation or edit) of the modified plant. For purposes of comparison to a modified plant, plant seed, plant part, plant cell, and/or plant genome, a “wild-type plant” (or likewise a “wild-type” plant seed, plant part, plant cell, and/or plant genome) refers to a non-transgenic and non-genome edited control plant, plant seed, plant part, plant cell, and/or plant genome. As used herein, a “control” plant, plant seed, plant part, plant cell, and/or plant genome may also be a plant, plant seed, plant part, plant cell, and/or plant genome having a similar (but not the same or identical) genetic background to a modified plant, plant seed, plant part, plant cell, and/or plant genome, if deemed sufficiently similar’ for comparison of the characteristics or traits to be analyzed.
[0182] As used herein, the terms “suppress,” “suppression,” “inhibit,” “inhibition,” “inhibiting,” “knockout,” “knockdown,” and “downregulation” refer to a lowering, reduction, or elimination of the expression level of an mRNA and/or protein encoded by a target gene in a plant, plant cell, or plant tissue at one or more stage(s) of plant development, as compared to the expression level of
such target mRNA and/or protein in a wild-type or control plant, cell, or tissue at the same stage(s) of plant development. According to some embodiments, a modified plant is provided having a GmTDFl-1 gene expression level that is reduced in at least one plant tissue by at least 5%, at least 10%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, or 100%, as compared to a control plant. According to further embodiments, a modified plant is provided having a GmTDFl-1 gene expression level that is reduced in at least one plant tissue by 5%-20%, 5%-25%, 5%-30%, 5%-40%, 5%-50%, 5%- 60%, 5%-70%, 5%-75%, 5%-80%, 5%-90%, 5%-100%, 75%-100%, 50%-100%, 50%-90%, 50%- 75%, 25%- 75%, 30%-80%, or 10%-75%, as compared to a control plant.
[0183] According to some embodiments, a modified plant is provided having a GmTDFl-1 mRNA level that is reduced in at least one plant tissue by at least 5%, at least 10%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, or 100%, as compared to a control plant. According to some embodiments, a modified plant is provided having a GmTDFl-1 mRNA expression level that is reduced in at least one plant tissue by 5%-20%, 5%-25%, 5%- 30%, 5%-40%, 5%-50%, 5%-60%, 5%-70%, 5%- 75%, 5%-80%, 5%-90%, 5%-100%, 75%-100%, 50%-100%, 50%-90%, 50%-75%, 25%-75%, 30%-80%, or 10%-75%, as compared to a control plant. According to some embodiments, a modified plant is provided having a GmTDFl-1 protein expression level that is reduced in at least one plant tissue by at least 5%, at least 10%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, or 100%, as compared to a control plant. According to some embodiments, a modified plant is provided having a GmTDFl-1 protein expression level that is reduced in at least one plant tissue by 5%-20%, 5%- 25%, 5%-30%, 5%-40%, 5%-50%, 5%-60%, 5%-70%, 5%-75%, 5%-8O%, 5%-90%, 5%-100%, 75%-100%, 50%-100%, 50%-90%, 50%-75%, 25%-75%, 30%-80%, or 10%-75%, as compared to a control plant.
[0184] Modified plants comprising or derived from plant cells that are transformed with a recombinant DNA of this disclosure can be further enhanced with stacked traits, for example, a modified crop plant having an enhanced trait resulting from expression of DNA disclosed herein in combination with one or more genes of agronomic interest that provide a beneficial agronomic trait (such as herbicide and/or pest resistance traits) to crop plants. For example, the traits conferred by the recombinant DNA constructs of the current disclosure can be stacked with other
traits of agronomic interest, such as a trait providing insect resistance such as using a gene from Bacillus thuringensis to provide resistance against lepidoptcran, colcoptcran, homoptcran, hemiopteran, and other insects, or improved quality traits such as improved nutritional value. Molecules and methods for imparting insect/nematode/virus resistance are disclosed in U.S. Patent Nos. 5,250,515; 5,880,275; 6,506,599; 5,986,175; and U.S. Patent Application Publication No. 2003/0150017 Al.
[0185] Herbicides for which transgenic plant tolerance has been demonstrated and the methods and compositions of the present disclosure can be applied include, but are not limited to, glyphosate, dicamba, glufosinate, sulfonylurea, bromoxynil, norflurazon, 2,4-D (2,4- dichlorophenoxy) acetic acid, aryloxyphenoxy propionates, p-hydroxyphenyl pyruvate dioxygenase inhibitors (HPPD), and protoporphyrinogen oxidase inhibitors (PPO) herbicides. Polynucleotide molecules encoding proteins involved in herbicide tolerance known in the art and include, but are not limited to, a polynucleotide molecule encoding 5-enolpyruvylshikimate-3- phosphate synthase (EPSPS) disclosed in U.S. Patent Nos. 5,094,945; 5,627,061; 5,633,435 and 6,040,497 for imparting glyphosate tolerance; polynucleotide molecules encoding a glyphosate oxidoreductase (GOX) disclosed in U.S. Patent No. 5,463,175 and a glyphosate-N-acetyl transferase (GAT) disclosed in U.S. Patent No. Application Publication No. 2003/0083480 Al also for imparting glyphosate tolerance; dicamba monooxygenase disclosed in U.S. Patent Application Publication No. 2003/0135879 Al for imparting dicamba tolerance; a polynucleotide molecule encoding bromoxynil nitrilase (Bxn) disclosed in U.S. Patent No. 4,810,648 for imparting bromoxynil tolerance; a polynucleotide molecule encoding phytoene desaturase (crtl) described in Misawa et al. (Plant J. 4:833-840, 1993) and in Misawa et al. (Plant J. 6:481-489, 1994) for norflurazon tolerance; a polynucleotide molecule encoding acetohydroxyacid synthase (AHAS, also known as ALS) described in Sathasivan et al. (Nucl. Acids Res. 18:2188-2193, 1990) for imparting tolerance to sulfonylurea herbicides; polynucleotide molecules known as bar genes disclosed in DeBlock et al. (EMBO J. 6:2513-2519, 1987) for imparting glufosinate and bialaphos tolerance; polynucleotide molecules disclosed in U.S. Patent Application Publication 2003/010609 Al for imparting N-amino methyl phosphonic acid tolerance; polynucleotide molecules disclosed in U.S. Patent No. 6,107,549 for imparting pyridine herbicide resistance; molecules and methods for imparting tolerance to multiple herbicides such as glyphosate, atrazine, ALS inhibitors,
isoxoflutole and glufosinate herbicides are disclosed in U.S. Patent No. 6,376,754 and U.S. Patent Application Publication 2002/0112260.
[0186] Genetic elements, methods, and transgenes that confer fungal disease resistance may also be used with the present disclosure (U.S. Pat. Nos. 6,653,280; 6,573,361; 6,506,962; 6,316,407; 6,215,048; 5,516,671; 5,773,696; 6,121,436; 6,316,407; 6,506,962). Soybean diseases caused by fungi include, but are not limited to, Phakopsora pachyrhizi, Phakopsora meibomiae (Asian Soybean Rust), Colletotrichum truncatum, Colletotrichum dematium var. truncatum, Glomerella glycines (Soybean Anthracnose), Phytophthorci sojae (Phytophthora root and stem rot), Sclerotinia sclerotiorum (Sclerotinia stem rot), Fusarium solani f. sp. glycines (sudden death syndrome), Fusarium spp. (Fusarium root rot), Macrophomina phaseolina (charcoal rot), Septaria glycines, (Brown Spot), Pythium aphanidermatum, Pythium debaryanum, Pythium irregulare, Pythium ultimum, Pythium myriotylum, Pythium torulosum (Pythium seed decay), Diaporthe phaseolorum var. sojae (Pod blight), Phomopsis longicola (Stem blight), Phomopsis spp. (Phomopsis seed decay), Peronospora manshurica (Downy Mildew), Rhizoctonia solani (Rhizoctonia root and stem rot, Rhizoctonia aerial blight), Phialophora gregata (Brown Stem Rot), Diaporthe phaseolorum var. caulivora (Stem Canker), Cercospora kikuchii (Purple Seed Stain), Alternaria sp. (Target Spot), Cercospora sojina (Frogeye Leafspot), Sclerotium rolfsii (Southern blight), Arkoola nigra (Black leaf blight), Thielaviopsis basicola, (Black root rot), Choanephora infundibulifera, Choanephora trispora Choanephora leaf blight), Leptosphaerulina trifolii (Leptosphaerulina leaf spot), Mycoleptodiscus terrestris (Mycoleptodiscus root rot), Neocosmospora vasinfecta (Neocosmospora stem rot), Phyllosticta sojicola (Phyllosticta leaf spot), Pyrenochaeta glycines (Pyrenochaeta leaf spot), Cylindrocladium crotalariae (Red crown rot), Dactuliochaeta glycines (Red leaf blotch), Spaceloma glycines (Scab), Stemphylium botryosum Stemphylium leaf blight), Corynespora cassiicola (Target spot), Nematospora coryli (Yeast spot), and Phymatotrichum omnivorum (Cotton Root Rot).
V. Definitions
[0187] The following definitions are provided to define and clarify the meaning of these terms in reference to the relevant embodiments of the present disclosure as used herein and to guide those of ordinary skill in the art in understanding the present disclosure. Unless otherwise noted, terms
are to be understood according to their conventional meaning and usage in the relevant art, particularly in the field of molecular biology and plant transformation.
[0188] When introducing elements of the present disclosure or the embodiment(s) thereof, the articles “a,” “an,” “the,” and “said” are intended to mean that there are one or more of the elements. [0189] The term “and/or,” when used in a list of two or more items, means any one of the items, any combination of the items, or all of the items with which this term is associated.
[0190] The terms “comprising,” “including,” and “having” are intended to be inclusive and mean that there may be additional elements other than the listed elements. For example, any method that “comprises,” “has” or “includes” one or more steps is not limited to possessing only those one or more steps and can also cover other unlisted steps. Similarly, any composition or device that “comprises,” “has,” or “includes” one or more features is not limited to possessing only those one or more features and can cover other unlisted features.
[0191] As used herein, a “plant” includes a whole plant, explant, plant part, seedling, or plantlet at any stage of regeneration or development.
[0192] As used herein, a “plant part” can refer to any organ or intact tissue of a plant, such as a meristem, shoot organ/structure (e.g., leaf, stem or node), root, flower or floral organ/structure (e.g., bract, sepal, petal, stamen, carpel, anther and ovule), seed, embryo, endosperm, seed coat, fruit, the mature ovary, propagule, or other plant tissues (e.g., vascular tissue, dermal tissue, ground tissue, and the like), or any portion thereof. Plant parts of the present disclosure can be viable, nonviable, regenerable, and/or non-regenerable. A “propagule” can include any plant part that can grow into an entire plant.
[0193] An “embryo” is a part of a plant seed, consisting of precursor tissues (e.g., meristematic tissue) that can develop into all or part of an adult plant. An "embryo" may further include a portion of a plant embryo.
[0194] A “meristem” or “meristematic tissue” comprises undifferentiated cells or meristematic cells, which are able to differentiate to produce one or more types of plant parts, tissues or structures, such as all or part of a shoot, stem, root, leaf, seed, etc.
[0195] As used herein, the “vegetative phase” of plant development is the period of growth between germination and flowering. The stages in the vegetative phase of soybean are as follows:
VE (emergence), VC (cotyledon stage), V 1 (first trifoliolate leaf), V2 (second trifoliolate leaf), V3 (third trifoliolate leaf), V(n) (nth trifoliolate leaf), and V6 (flowering will soon start). As used herein, the “reproductive phase” of plant development is the period between flowering and the end of harvest. The stages in the reproductive phase of soybean are as follows R1 (beginning bloom, first flower); R2 (full bloom, flower in top 2 nodes); R3 (beginning pod, 3/16" pod in top 4 nodes); R4 (full pod, 3/4" pod in top 4 nodes); R5 (1/8" seed in top 4 nodes); R6 (full size seed in top 4 nodes); R7 (beginning maturity, one mature pod); and, R8 (full maturity, 95% of pods on the plant have reached mature color). Soybean vegetative and reproductive stages are well known to those of skill in the art and numerous publications describing these stages can be found on the world wide web and elsewhere, such as North Dakota State University publication A- 1174, June 1999, Reviewed and Reprinted August 2004.
[0196] All methods described herein can be performed in any suitable order unless otherwise indicated herein or clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided with respect to certain embodiments herein is intended merely to illuminate the present disclosure and does not pose a limitation on the scope of the present disclosure otherwise claimed.
EMBODIMENTS
[0197] For further illustration, additional exemplary, non-limiting embodiments of the present disclosure are set forth below.
[0198] Embodiment l is a modified plant, plant seed, plant part, or plant cell comprising a genomic modification that reduces expression or activity of TDF1, or a homolog thereof, as compared to the expression or activity of TDF1 or a homolog thereof in an otherwise identical plant, plant seed, plant part, or plant cell that lacks the modification.
[0199] Embodiment 2 is the modified plant, plant seed, plant part, or plant cell of embodiment 1 , wherein the modification is present in at least one allele of an endogenous TDF1 gene or homolog thereof.
[0200] Embodiment 3 is the modified plant, plant seed, plant part, or plant cell of embodiment 2, wherein the TDF1 gene or homolog thereof encodes a protein comprising a conserved SANT/Myb domain having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at
least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the conserved SANT/Myb domain corresponding to position 9 to position 112 of SEQ ID NO:5. [0201] Embodiment 4 is the modified plant, plant seed, plant part, or plant cell of embodiment 2 or 3, wherein the modification disrupts the function of a wild-type allele product of the endogenous TDF1 gene or homolog thereof.
[0202] Embodiment 5 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 2-4, wherein the modification is located within a coding region, non-coding region, or any combination thereof, of the endogenous TDF1 gene or homolog thereof.
[0203] Embodiment 6 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 2-5, wherein the modification is located in an exon region of said TDF1 gene or homolog thereof.
[0204] Embodiment 7 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 2-5, wherein the modification is located in an intron region of said TDF1 gene or homolog thereof.
[0205] Embodiment 8 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 2-5, wherein the modification is located in an exon region and an intron region of said TDF1 gene or homolog thereof.
[0206] Embodiment 9 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 1-8, wherein the plant is a leguminous plant, or wherein the plant seed, plant part, or plant cell is a plant seed, plant part, or plant cell of a leguminous plant.
[0207] Embodiment 10 is the modified plant, plant seed, plant part, or plant cell of embodiment 9, wherein the leguminous plant is a soybean plant, a bean plant, a pea plant, a chickpea plant, an alfalfa plant, a peanut plant, a carob plant, a lentil plant, or a licorice plant.
[0208] Embodiment 11 is the modified plant, plant seed, plant part, or plant cell of embodiment
10, wherein the leguminous plant is a soybean plant.
[0209] Embodiment 12 is the modified plant, plant seed, plant part, or plant cell of embodiment
11, wherein the TDF1 gene is GmTDFl-1.
[0210] Embodiment 13 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 1-12, wherein the plant, plant seed, plant part, or plant cell is heterozygous for the modification.
[0211] Embodiment 14 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 1-12, wherein the plant, plant seed, plant part, or plant cell is homozygous for the modification.
[0212] Embodiment 15 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 1-10, wherein the plant, plant seed, plant part, or plant cell comprises a first modification in a first allele of the TDF1 gene and a second modification in a second allele of the TDF1 gene, the first modification and the second modification being different from one another.
[0213] Embodiment 16 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 1-15, wherein the modification comprises a deletion, an insertion, a substitution, an inversion, a duplication, or a combination of any thereof.
[0214] Embodiment 17 is the modified plant, plant seed, plant part, or plant cell of embodiment
16, wherein the modification is a deletion.
[0215] Embodiment 18 is the modified plant, plant seed, plant part, or plant cell of embodiment
17, wherein the deletion comprises between 1 nucleotide and 1500 nucleotides.
[0216] Embodiment 19 is the modified plant, plant seed, plant part, or plant cell of embodiment
18, wherein the deletion comprises a region of at least 1 nucleotide, at least 2 nucleotides, at least 5 nucleotides, at least 10 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, at least 500 nucleotides, at least 600 nucleotides, at least 700 nucleotides, at least 800 nucleotides, at least 900 nucleotides, at least 1000 nucleotides, or at least 1500 nucleotides.
[0217] Embodiment 20 is the modified plant, plant seed, plant part, or plant cell of embodiment 16, wherein the modification is an inversion.
[0218] Embodiment 21 is the modified plant, plant seed, plant part, or plant cell of embodiment 20, wherein the inversion is a region of at least 5 consecutive nucleotides, at least 10 consecutive nucleotides, at least 20 consecutive nucleotides, at least 30 consecutive nucleotides, at least 40 consecutive nucleotides, at least 50 consecutive nucleotides, at least 100 consecutive nucleotides, at least 200 consecutive nucleotides, at least 300 consecutive nucleotides, at least 400 consecutive nucleotides, at least 500 consecutive nucleotides, at least 600 consecutive nucleotides, at least 700 consecutive nucleotides, at least 800 consecutive nucleotides, at least 900 consecutive nucleotides, at least 1000 consecutive nucleotides, or at least 1100 consecutive nucleotides.
[0219] Embodiment 22 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 12-21, wherein the plant, plant seed, plant part, or plant cell comprises a modification in at least one allele of the GmTDFl-1 gene, wherein the modification is selected from the group consisting of: an 8 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:26; a 10 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:27; a 1 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:28; a 6 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:29; a 9 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:30; an 8 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:31; an 11 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:32; a 7 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:33; a 12 base pair deletion and an 8 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:34; two inversions, wherein the resulting nucleotide sequence is SEQ ID NO:35; a 15 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:36; a 19 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:37; a 107 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:38; a 27 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:39; a 9 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:40; a 9 base pair deletion and a 10 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:41; a 1,025 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:42; an 11 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:43; a 7 base pair deletion and an 11 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:44; an inversion, wherein the resulting nucleotide sequence is SEQ ID NO:45; a 1,055 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:46; an 11 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:47; a 16 base pair indel and a 9 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:48; a 122 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:49; a 65 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:50; an inversion, wherein the resulting nucleotide sequence is SEQ ID NO:51 ; a 4 base pair deletion and a 36 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:52; two inversions, wherein the resulting nucleotide sequence is SEQ ID NO:53; and any combinations of any thereof.
[0220] Embodiment 23 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 12-22, wherein the plant, plant seed, plant part, or plant cell comprises at least one polynucleotide sequence selected from the group consisting of SEQ ID NOs:26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, and 53.
[0221] Embodiment 24 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 12-23, wherein the plant, plant seed, plant part, or plant cell comprises at least one polynucleotide sequence selected from the group consisting of SEQ ID NOs:26, 27, 28, 29, 30, 31, 32, and 33.
[0222] Embodiment 25 is the modified plant, plant seed, plant part, or plant cell of any one of embodiments 12-23, wherein the plant, plant seed, plant part, or plant cell comprises at least one polynucleotide sequence selected from the group consisting of SEQ ID NOs:34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, and 53.
[0223] Embodiment 26 is the modified plant of any one of embodiments 1-25, wherein the modification results in a male sterility phenotype in the plant, as compared to the male fertile phenotype of an otherwise identical plant that lacks the modification.
[0224] Embodiment 27 is a polynucleotide comprising a sequence selected from the group consisting of SEQ ID NOs:26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, and 53.
[0225] Embodiment 28 is a guide RNA comprising a polynucleotide sequence selected from the group consisting of SEQ ID NOs:23, 24, and 25.
[0226] Embodiment 29 is the guide RNA of embodiment 28, wherein the polynucleotide sequence is a spacer sequence.
[0227] Embodiment 30 is a method for producing a modified plant having a male sterility phenotype, the method comprising: a) introducing a modification into at least one target site in an endogenous TDF1 gene or a homolog thereof of a plant cell; b) identifying and selecting one or more plant cells of step a) comprising said modification in said TDF1 gene or homolog thereof; and c) regenerating at least one plant from at least one or more cells selected in step b).
[0228] Embodiment 31 is the method of embodiment 30, wherein the target site is located in a coding region, a non-coding region, or any combination thereof, of an endogenous TDF1 gene or homolog thereof.
[0229] Embodiment 32 is the method of embodiment 30, wherein the modification is introduced by a site-specific genome modification enzyme selected from the group consisting of: an RNA- guided nuclease, a zinc-finger nuclease, a meganuclease, a TALE-nuclease, a recombinase, a transposase, and combinations of any thereof.
[0230] Embodiment 33 is the method of embodiment 32, wherein the site-specific genome modification enzyme is an RNA-guided nuclease comprising a Cas nuclease, a Cpfl nuclease, or a variant of either thereof.
[0231] Embodiment 34 is the method of embodiment 33, wherein the site-specific genome modification enzyme is an RNA-guided nuclease comprising a Cpfl nuclease.
[0232] Embodiment 35 is the method of embodiment 30, wherein the modification is introduced by a site- specific genome modification enzyme that creates at least one strand break at the target site.
[0233] Embodiment 36 is the method of embodiment 30, wherein the modification is selected from the group consisting of a substitution, an insertion, an inversion, a deletion, a duplication, and a combination thereof.
[0234] Embodiment 37 is the method of embodiment 36, wherein the modification is a deletion.
[0235] Embodiment 38 is the method of embodiment 37, wherein the deletion comprises a region of at least 1 nucleotide, at least 2 nucleotides, at least 5 nucleotides, at least 10 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, at least 500 nucleotides, at least 600 nucleotides, at least 700 nucleotides, at least 800 nucleotides, at least 900 nucleotides, or at least 1000 nucleotides, or at least 1500 nucleotides.
[0236] Embodiment 39 is the method of embodiment 36, wherein the modification is an inversion. [0237] Embodiment 40 is the method of embodiment 39, wherein the inversion is a region of at least 5 consecutive nucleotides, at least 10 consecutive nucleotides, at least 20 consecutive nucleotides, at least 30 consecutive nucleotides, at least 40 consecutive nucleotides, at least 50 consecutive nucleotides, at least 100 consecutive nucleotides, at least 200 consecutive nucleotides, at least 300 consecutive nucleotides, at least 400 consecutive nucleotides, at least 500 consecutive
nucleotides, at least 600 consecutive nucleotides, at least 700 consecutive nucleotides, at least 800 consecutive nucleotides, at least 900 consecutive nucleotides, at least 1000 consecutive nucleotides, or at least 1100 consecutive nucleotides.
[0238] Embodiment 41 is a male sterile soybean plant produced by any of the methods of embodiments 30-40.
[0239] Embodiment 42 is a method for producing an Fi hybrid soybean plant comprising crossing a male-sterile soybean plant comprising a modified GmTDFl-1 gene with a second, non-isogenic, male-fertile soybean plant to produce an Fi hybrid soybean plant.
[0240] Embodiment 43 is an Fi hybrid soybean plant produced by the method of embodiment 42. [0241] Embodiment 44 is a method for producing an Fi hybrid soybean seed, comprising crossing the plant of embodiment 1 with a second, non-isogenic, male-fertile soybean plant and harvesting the resultant Fi hybrid soybean seed, wherein an Fi hybrid soybean seed is produced.
[0242] Embodiment 45 is an Fi hybrid soybean seed produced by the method of embodiment 44. [0243] Embodiment 46 is a modified plant, plant seed, plant part, or plant cell, wherein the plant, plant seed, plant part, or plant cell comprises at least one polynucleotide sequence selected from the group consisting of SEQ ID NOs:26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, and 53.
[0244] Embodiment 47 is the modified plant of embodiment 12, wherein the modification results in a male sterility phenotype in the plant, as compared to the male fertile phenotype of an otherwise identical plant that lacks the modification.
[0245] Embodiment 48 is the modified plant of embodiment 23, wherein the modification results in a male sterility phenotype in the plant, as compared to the male fertile phenotype of an otherwise identical plant that lacks the modification.
[0246] Embodiment 49 is the method of any one of embodiments 30-40, wherein the plant cell of step a) is a soybean plant cell and wherein the produced modified plant having a male sterility phenotype is a soybean plant.
[0247] Embodiment 50 is a method for producing an Fi hybrid soybean seed, comprising crossing the plant of any one of embodiments 1-26 with a second, non-isogenic, male-fertile soybean plant and harvesting the resultant Fi hybrid soybean seed, wherein an Fi hybrid soybean seed is produced.
[0248] Embodiment 51 is an Fi hybrid soybean seed produced by the method of embodiment 50.
Embodiment 52 is the modified plant, plant seed, plant part, or plant cell of embodiment 46, wherein the plant, plant seed, plant part, or plant cell is a transgenic soybean plant, transgenic soybean plant seed, transgenic soybean plant part, or transgenic soybean plant cell.
[0249] Embodiment 52 is the modified plant, plant seed, plant part, or plant cell of embodiment 12, wherein the plant, plant seed, plant part, or plant cell comprises a first modification in a first allele of the TDF1-1 gene and a second modification in a second allele of the TDF1-1 gene, the first modification and the second modification being different from one another.
[0250] The embodiments described herein may be more readily understood through reference to the following examples, which are provided by way of illustration, and are not intended to be limiting, unless specified.
EXAMPLES
Example 1. Identification of the Candidate Gene Associated with Male Sterility.
[0251] The Ms6 locus in soybean has been identified as being associated with male fertility in soybean. A naturally occurring male sterility allele in soybean is caused by the recessive ms6 allele. This Ms6 locus is publicly documented in the literature to reside approximately 10 centimorgan (cM) from the flower color locus (Wl) on soybean chromosome 13. To more efficiently identify homozygous ms6 mutant seeds for use as female parental lines in soybean hybrid production, an ms6 mutant soybean line was whole genome sequenced and mapped using fine mapping techniques to improve markers for quality seed identification as well as to identify a high- confidence gene candidate for the Ms6 locus. The approximate gene locus region of the ms6 trait on chromosome 13 was narrowed via joint linkage mapping. Joint linkage mapping was performed for the ms6 mutant line with 5 populations heterozygous for the flower color marker using 27 markers between 15 cM to 44.8 cM on chromosome 13 to identify the most polymorphic markers among the populations. This joint linkage mapping narrowed the region coding for the ms6 trait to be between 26 cM and 32 cM on chromosome 13 of the soybean cultivar Williams 82 (Wm82) germplasm. The region between 15 cM to 35 cM, corresponding to 12,742,089 bp and 16,750,332 bp on chromosome 13, was further investigated using fine mapping.
[0252] For the fine mapping, whole genome sequencing was performed on the ms6 line and the region spanning from 15 cM to 35 cM on chromosome 13 was compared to the corresponding regions in a panel of 37 reference lines for which whole genome sequences were available.
Expressed single nucleotide polymorphism (eSNP) calls were extracted from this region in the ms6 line. The allele frequency for each eSNP was calculated and five low allele frequency variants were identified to be unique to the ms6 line and advanced for further evaluation. Of these five variants, two were determined to have a low likelihood for being causative of the male sterility trait due to their locations relative to the putative ms6 region (spanning from 26 cM to 32 cM on chromosome 13) predicted by joint linkage mapping. Of the remaining three variants, two were considered unsuitable for marker design due to their surrounding sequences. As such, the remaining variant, designated eSNPl, was used as the candidate eSNP associated with the ms6 allele. The variation at eSNPl is a T->A nucleotide change at 28 cM (14,995,614 bp) on chromosome 13 (Table 1).
Table 1. Marker eSNPl is associated with the putative region comprising the Ms6 locus on chromosome 13.
[0253] Identification of the candidate gene could allow for gene editing approaches to reproduce the male sterile phenotype and male sterile lines. Using the mapping data generated above, a bioinformatics approach was used to identify the gene associated with eSNPl locus. The T- A mutation identified at eSNPl was found in a genic region containing the A7v/>- fa ily of genes. This gene (Glymal3.g066600) has high functional domain similarity to the Tapetum Development & Function 1 (AtTDFF) gene in Arabidopsis thaliana. This gene, which is also known as AtMYB35 (SEQ ID NO:1), encodes a member of the R2R3 MYB transcription factor superfamily that is essential for early tapetum development. The AtTDFl protein (SEQ ID NO:2) is understood to play a role in male fertility. Based on sequence homology and male fertility-related functional similarity, the candidate gene mutated in the ms6 line is the soybean homolog of AtTDFl and is designated herein as GmTDFl-1. These findings are supported by GmTDFl-1 (SEQ ID NO:5) and AtTDFl (SEQ ID NO:2) share 53% sequence identity over the full length of their amino acid sequences and 83.3% sequence identity in the R2R3 MYB domain. The R2R3-MYB protein consists of two major functional parts: a DNA-binding domain (MYB domain) located at the N- terminus and a regulatory region (non-MYB region) located at the C-terminus. The MYB domain
is a signature, highly conserved feature in the gene family, whereas the non-MYB regions have diverged among plant species.
[0254] Glycine max TDF1-1 (GmTDFI-i) plays an important role in plant growth and development. The soybean GmTDFl-1 gene is 2,094 bp in length and comprises three exons and two introns. The coding and genomic sequences of the GmTDFl-1 gene are set forth in SEQ ID NOs;3 and 4, respectively. This gene has two in-frame translational start codons that encode two potential proteins, a longer GmTDFl-1 protein that is 375 amino acids in length and a shorter GmTDFl-1 protein that is 345 amino acids in length, alternatively translated from the second start codon of the mRNA. The first in-frame transcriptional start codon is at nucleotide position 18, and the second in-frame transcriptional start codon is at nucleotide position 108 in SEQ ID NO:4, respectively. The protein sequence of the shorter GmTDFl-1 protein is set forth in SEQ ID NO:5. The additional 30 amino acids in the extended C-terminal of the longer GmTDFl-1 protein sequence (not shown) identifies only soybean consensus sequences in the GenBank database and suggests that this additional sequence is either soybean-specific, or not translated. Both the shorter and longer GmTDFl-1 sequences and annotations can be found in the SoyBase database (www.soybase.org) depending on the reference genome used. Bioinformatic analysis and annotation of the ms6 line utilized for the experiments in these Examples supports use of the shorter GmTDFl-1 protein sequence as the reference and, accordingly, use of the second in-frame start codon of the GmTDFl-1 nucleotide sequence as the start codon. Accordingly, the 90 nucleotides from the first in-frame start codon until the second in-frame start codon are counted as a pail of the 5’ UTR of the GmTDFl-1 gene. As such, the CDS utilized as SEQ ID NO:3 represents the shorter version of the nucleotide sequence based on the second in-frame transcriptional start codon.
[0255] The eSNPl was identified at position 347 in SEQ ID NO:4 and occurs at the 5’ end of Exon 2, which spans nucleotides 344 to 473 in SEQ ID NO:4. This is a missense mutation which changes the amino acid residue from leucine to histidine in the GmTDFl-1 polypeptide sequence (SEQ ID NO:5). This missense mutation is associated with the ms6 phenotype (FIGS. 1A and IB). The coordinates of each of the exons within the GmTDFl-1 genomic DNA sequence (SEQ ID NO:4) shown in Table 2 are indicative of the nucleotide position(s) counted from the first nucleotide of SEQ ID NO:4 in the 5’ to 3’ direction.
Table 2. Coordinates of exons in the GmTDFl-1 gene sequence.
[0256] A marker was converted from eSNPl and analyzed for high association to the male sterility trait and genetic accuracy of the marker was confirmed. Assay probes and primers were developed and customized for accurate genetic identification of seeds for hybrid production (SEQ ID NOs:6- 12 and 14).
Example 2. Design of Gene Editing Constructs.
[0257] Using gene editing technologies to confirm causality of GmTDFl-1 for male sterility associated with the ms6 trait in soybean plants and produce novel alleles of the candidate gene for the ms6 mutation could allow for production of more environmentally stable male sterile plants for soybean hybrid production and further development of the trait for agronomic use.
[0258] In this example, the two gene editing constructs (pM764 and pM765) each contained two functional regions or cassettes relevant to gene editing and the creation of the DSBs in the edited GmTDFl-1 gene region: expression of a Cpfl protein and expression of one or two guide RNAs targeting the GmTDFl-1 gene. Each gRNA unit contains a common scaffold compatible with the Cpfl gene (SEQ ID NO: 18), and a unique spacer/targeting sequence complementary to an intended target site as listed in Table 3. The coordinates provided in Table 3 are indicative of the nucleotide position(s) in the sequence of SEQ ID NO:4 (z.e., the genomic GmTDFl-1 sequence), in the 5’ to 3’ direction. Both guide RNA spacers of construct pM765 (gRNA_-103C and gRNA_-1113G) are oriented in the antisense direction and thus, the reverse complement sequences of the spacer sequences match the target site given in Table 3.
Table 3. Example guide RNAs used for editing the endogenous GmTDFl-1 gene.
[0259] The Cpfl expression cassette of both editing constructs comprised a Dahlia mosaic virus FLT promoter (SEQ ID NO: 19) operably linked to a sequence encoding a Lachnospiraceae bacterium Cpfl RNA-guided endonuclease enzyme (SEQ ID NO:20) that was codon-optimized for expression in plants, flanked on each side by one copy of a nuclear localization signal (SEQ ID NO:21). See, e.g., Gao et al. (Nature Biotechnol. 35(8):789-792, 2017).
[0260] One type of gRNA expression cassette, present in construct pM764, comprised a sequence encoding one GmTDFl-1 gRNA (guide RNA spacer gRNA_+221C) operably linked to a soybean RNA polymerase III (Pol3) promoter (SEQ ID NO:22). The sequence of SEQ ID NO:23, listed in Table 3, is a spacer sequence targeted near the candidate site of the ms6 point mutation (/.<?., near mutation T347A in the GmTDFl-1 gene), with the intent to produce lesions near the natural mutation locus and cause loss of function of the gene through codon deletion, frameshift, or loss of a splice site, thus resulting in a precise recapitulation of the mutation. The other gRNA expression cassette, present in construct pM765 (guide RNA spacers gRNA_-103C and gRNA_- 1113G), comprised a sequence encoding two GmTDFl -1 guide RNAs operably linked to the same soybean RNA polymerase III (Pol3) promoter (SEQ ID NO:22). The sequences of SEQ ID NOs:24 and 25, listed in Table 3, are spacer sequences that target alternative DSB sites, one near the 3’ end of Exon 1 and one near the 5’ end of Exon 3 in GmTDFl-1, aiming to result in deletion of most of the coding sequence of the GmTDFl-1 gene and ensure loss of function. In FIG. 2, the target sites of the three gRNAs are visualized on the GmTDFl-1 gene. The gRNAs were designed to avoid targeting GmTDFl-2, a close paralog of GmTDFl-1 on chromosome 19.
Example 3. Generation of Novel GmTDFl-1 Alleles Produced Through Gene Editing.
[0261] An inbred wild-type soybean line was transformed using Agrobacterium-mediated transformation with the pM764 or pM765 vector described in Example 2 above. The transformed plant tissue was grown to produce mature Ro plants. A total of 168 Ro transformants were regenerated per construct and assayed by long or short amplicon sequencing for detection of edits in the endogenous GmTDFl-1 gene. Sequence data from each sample was mapped to a reference sequence to identify consensus differences. Plants with unique deletions were selected to provide diverse coverage of gene mutations in the targeted genomic region. Most mutations generated by the CRISPR/Cpfl system can comprise insertions, deletions, and/or inversions depending on the guide RNA(s) used. Cpfl -derived edits are located downstream of the PAM site.
[0262] Plants transformed with construct pM764 were evaluated for small deletions using a short amplicon that extended from Intron 1 to Intron 2. A total of eight monoallelic edits were recovered from construct pM764, which had an editing frequency of approximately 5%. All eight mutants were advanced for continued growth observation and analysis. The resultant edited alleles were sequenced to determine the genotype and plants were characterized for male sterile phenotypes. An additional two regenerants from pM765 (transformation performed on identical genetic background as pM764 transformations) that were identified as having no edits were selected for continued growth and analysis as presumptive wild-type controls. A summary of the monoallelic edits and their characteristics can be found in Table 4. The
in the “Ro Genotype” column indicates the status of allele 1/allele 2 wherein “WT” indicates that one allele is unedited and “d” indicates that a deletion had occurred during the editing process. The nucleotide positions of each of the deletions described in Table 4 are based on the nucleotide position(s) counted from the first nucleotide of the sequence of SEQ ID NO:4 being “position 1.”. The nucleotide positions in Table 4 are based on an alignment created using the BLAST computational alignment method. Generally, visibly misaligned bases at either end of a modification, an occasional byproduct of multiple sequence alignment methods, can be corrected manually so that the coordinates reflect the correct positions of the modifications with reference to SEQ ID NO:4. In some edits, multiple sets of “Nucleotide Positions of Deletion (relative to SEQ ID NO:4)” coordinates are possible as depending on the alignment method utilized, the coordinates may differ due to presence of identical base(s) at either end of a modification. However, regardless of the alignment method
utilized, all edits are unambiguously defined by their individual sequence presented in Table 4 via the identifier “SEQ ID NO:”.
[0263] Edit 1 through Edit 8 (SEQ ID NOs:26-33) consisted of deletions from one base pair to 11 base pairs near the site of the ms6 mutation. Edit 1, Edit 4, Edit 5, and Edit 7 resulted in the mutation of a splice site between Intron 1 and Exon 2, whereas Edit 2, Edit 3, Edit 6, and Edit 8 resulted in frameshift mutations. All of the resultant plants were fertile in the Ro generation, meaning that pods were present on plants at representative reproductive stages.
[0264] Plants transformed with construct pM765 were evaluated for large deletions using a long amplicon that extended between Exon 1 and the 3’ end of Exon 3, which had an editing frequency of approximately 87%. A total of 146 diverse mutants were recovered and the resultant edited
alleles were sequenced to determine the genotype. A subset of the edited alleles produced from pM765 and their characteristics can be found in Tables 5 and 6. The coordinates provided in Table 5 are indicative of the nucleotide position(s) counted from the first nucleotide of the sequence of SEQ ID NO:4 in the 5’ to 3’ direction. In the “Description of Edit” column, the modification of GmTDFl-1 gene induced by the edit is specified. For example, Edit 10 comprises a 15 bp deletion resulting in a frameshift mutation in the GmTDFl-1 gene (SEQ ID NO:36). The coordinates presented in Table 5 are based on an alignment created using the BLAST computational alignment method. Generally, visibly misaligned bases at either end of a modification, an occasional byproduct of multiple sequence alignment methods, can be corrected manually so that the coordinates reflect the correct positions of the modifications with reference to SEQ ID NO:4. In some edits, multiple sets of “Position of Edits (relative to SEQ ID NO:4)” coordinates are possible as already described above for Table 4. However, regardless of the alignment method utilized, all edits are unambiguously defined by their individual sequence presented in Table 5 via the identifier “SEQ ID NO:”.
[0265] As shown in Table 5 above, the editing events resulted in deletions ranging from four base pairs to 1055 base pairs, in the GmTDFl-1 gene. These edits resulted in frameshift mutations and/or the deletion of codons or the GmTDFl-1 gene itself. An alignment of the GmTDFl-1 coding sequence, genomic sequence, and each of Edits 9-24 is shown in FIG. 4.
*possible modification(s) at edges of inversion
[0266] Edit 25, Edit 26, Edit 27, and Edit 28 each consisted of an inversion on the GmTDFl-1 gene. Generally, inversions can be accompanied by loss of nucleotides surrounding the inversion, resulting in deleted sequences on either end of the inverted sequence. In the process of Cpfl- mediated gene editing, the Cpfl endonuclease produces staggered cuts and, as such, overhanging DNA ends are generated. These are subject to endogenous DNA repair mechanisms, such as error- prone non-homologous end joining (NHEJ) which can lead to sequence variation at the cut and ligation sites. An alignment of the GmTDFl -1 genomic sequence and each of Edits 25-28 is shown in FIGS. 5A-5H.
[0267] Of the various edits generated, 10 biallelic mutant plants were selected for continued growth observations and analysis, based on containing multiple lesions, large deletions, or inversions. The mutations generated from pM765 were intended to confirm the male sterility phenotype. An additional two regenerants from pM765 that were identified as having no edits were selected for continued growth and analysis as presumptive wild-type controls. The results are shown in Table 7 below.
Table 7. Phenotypic characteristics of soybean plants comprising various combinations of edited GmTDFl-1 mutant alleles.
[0268] All plants comprising a combination of edits were male sterile, which was demonstrated by a failure to develop pods through development. It is expected that plants carrying any combination of Edits 1-28 will have the same male sterility phenotype. Overall, while the edits showed the predicted variation in fertility, other phenotypic changes, such as plant height and development time were considered to fall within naturally expected variation for the populations (data not shown).
Example 4. Phenotypic Confirmation of Causal Gene Through Gene Editing in Soybean Plants.
[0269] The mutant plants from pM765 were observed through pod maturation and dry down to document late-stage reproductive phenotypes. The lack of pods on these biallelic mutants suggests that GmTDFl-1 on chromosome 13 is the causal gene controlling the ms6 male sterility trait (FIG. 6). To confirm that the detected male sterility phenotype (failure to produce pods) results from male sterility in self-crosses, the selected wild-type control regenerants were used as pollen donors for hand-fertilization of flowers on nodes of podless mutant plants to rescue the male sterile
phenotype. Such fertilizations often resulted in normal pod formation and development (FIG. 7). Observations from these artificial fertilization of flowering nodes included: 1) no viable pods were produced on nodes of edited biallelic mutant plants, except those cross -pollinated with wild-type pollen by hand; 2) pod development on hand-pollinated nodes of mutants was more rapid than development on fertile controls; and 3) viable pods developed on hand-pollinated nodes that attained normal size and quantities per pod, indicating effective fertilization of male sterile edited plants with wild-type pollen. As wild-type pollen rescued the sterile phenotypes, the conclusion can be made that sterility in GmTDFl-1 edited plants is from male, not female, sterility.
[0270] Preliminary dissection and examination of the flowers from wild-type (fertile with pod development) and edited (male-sterile with failure of pod development) plants from those selected from pM765 show clear differences in the development of male (abnormal anthers), but not female, reproductive parts (FIG. 8). This observation further suggests a causal role for GmTDFl-1 in male-specific fertility development.
[0271] Further phenotypic dissection was performed on the microanatomy of reproductive organs, particularly of anther tapetum cells in edited GmTDFl-1 mutants. Buds (optimally in stage 10 of anther development) from selfed mutant (Edit 2 or Edit 4 from pM764) and wild-type plants were fixed, embedded, and sectioned prior to staining with toluidine blue. The stained sections were analyzed for mutant phenotypes (FIG. 9). Analysis found that the anthers of homozygous GmTDFl-1 mutants displayed disrupted tapetum development, with evidence of collapsed microspores and enlarged endothecium compared to wild-type soybean plants. Plants heterozygous for the edited alleles displayed phenotypes similar to the wild-type. The homozygous mutant soybean phenotype resembles the phenotype observed in homozygous AtTDFl mutant anthers in Arabiclopsis thaliana. Similar to the AtTDFl loss-of-function mutant, ms6 edited mutations in the GmTDFl-1 gene exhibited disrupted tapetum development. This microanatomical similarity is consistent with downstream failure to produce pollen, resulting in the male sterility seen in both natural Arabidopsis TDF1 mutants and natural soybean ms6 mutants, indicating a successful anatomic and functional phenotypic recapitulation of the causative ms6 mutation via the edited GmTDFl-1 alleles described here.
Example 5. Characterization of Heritability of Novel GmTDFl-1 Edited Alleles.
[0272] The monoallelic edits generated from pM764, Edit 1 through Edit 8 (SEQ ID NOs:26-33), were used for segregation analysis and seed production. Each of these edited alleles was expected to result in male sterility when present in the homozygous state. A representation of each edited allele, located at the Intron 1/Exon 2 border in the GmTDFl-1 gene, is shown in FIG. 10.
[0273] Seed produced from self-fertilization of Ro plants from each of Edit 2, Edit 3, Edit 4, and Edit 7, as well as an unedited event from pM765 from the same transformation window (WT) were used for planting. A total of 24 Ri seed were sown for each event and grown to generate Ri plants. The Ri plants were subsequently genotyped for segregation analysis. The results of the genotyping are shown in Table 8. Relating to the edit, “HOMO” represents the number of RI plants being homozygous, “HETERO” represents the number of Ri plants being heterozygous, and “NULL” represents the number of unedited Ri plants.
[0274] The segregation analysis showed that Edit 2 and Edit 4 showed typical recessive segregation patterns (near 1:2:1). Edit 7 showed roughly a typical recessive segregation pattern (near 1:2:1) but had lower germination than other alleles. Edit 3 showed a segregation bias resulting in no homozygotes, likely indicating an undetected second-site mutation. Nuclease- ncgativc, edit-positive individuals were identified from 3 of 4 edited lines (Edit 2, Edit 3, and Edit 4) and were used to produce R2 seed (data not shown).
[0275] The male sterility phenotype in publicly available ms6 soybean lines is the result of a non- synonymous single nucleotide polymorphism at the Ms6 locus that introduces only a single amino acid change in the encoded protein sequence. The present study demonstrates that the edited alleles described herein comprise nucleotide deletions in a wide range of lengths, many of which would substantially disrupt expression of the GmTDFl-1 gene, yet they all confer to soybean plants the same male sterility phenotype seen in the ms6 soybean lines without having an impact on normal plant development. Accordingly, these novel edited alleles of the GmTDFl-1 gene have valuable future use in hybrid development and production.
Example 6. Gene Editing of TDF1 Homologs.
[0276] Gene editing technology is available for applications in many plant species. The use of gene editing technology here to modify the GmTDFl-1 gene to produce male sterility in soybeans for enablement of hybrid soy production indicates possible applications of gene editing to other crop and plant species with highly homologous TDF/-likc genes for production of male sterility and enablement of additional hybrid crops.
[0277] Based on the results of the Examples above, the gene edits made in the GmTDFl-1 gene produce a phenotype that is consistent with the mutant AtTDFl allele in Arabidopsis thaliana, on both a microanatomical level and through an overall male sterile reproductive phenotype (lack of pod production in soybeans). The high level of functional domain sequence homology between the AtTDFl and GmTDFl-1 proteins combined with the phenotypic results indicate potential for success in applying editing of TDF1 -like genes in other plant species showing high functional domain homology.
[0278] Other TDF1 homologs were identified in other green plants from both dicot and monocot plant species. Sequence identity is conserved primarily within the SANT/Myb functional domain for DNA-binding located near the N-terminus of the protein and, particularly in dicots, a serine rich domain in the middle of each protein. An alignment of the amino acid sequences of AtTDFl (SEQ ID NO:2), GmTDFl-1 (SEQ ID NO:5), and other TDFl-like proteins from both dicot and monocot plant species (SEQ ID NOs:55-64) is shown in FIGS. 11 A- 1 IE. The SANT/Myb domain of AtTDFl is annotated as spanning from position 9 to position 116 of the sequence of SEQ ID NO:2. Similarly, the SANT/Myb domain of GmTDFl-1 is annotated as spanning from position 9 to position 112 of sequence of SEQ ID NO:5. The amino acid sequences of the other TDFl-like
proteins shown in FIGS. 11 A-l IE are from species with a high degree of sequence homology to the SANT/Myb functional domain of AtTDFl and GmTDFl-1.
[0279] Having described the present disclosure in detail, it will be apparent that modifications, variations, and equivalent embodiments are possible without departing from the spirit and scope of the present disclosure as described herein and in the appended claims. Furthermore, it should be appreciated that all examples in the present disclosure are provided as non-limiting examples.
Claims
1. A modified plant, plant seed, plant part, or plant cell comprising a genomic modification that reduces expression or activity of TDF1, or a homolog thereof, as compared to the expression or activity of TDF1 or a homolog thereof in an otherwise identical plant, plant seed, plant part, or plant cell that lacks the modification.
2. The modified plant, plant seed, plant part, or plant cell of claim 1 , wherein the modification is present in at least one allele of an endogenous TDF1 gene or homolog thereof.
3. The modified plant, plant seed, plant part, or plant cell of claim 2, wherein the TDF1 gene or homolog thereof encodes a protein comprising a conserved SANT/Myb domain having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the conserved SANT/Myb domain corresponding to position 9 to position 112 of SEQ ID NO:5.
4. The modified plant, plant seed, plant part, or plant cell of claim 2, wherein the modification disrupts the function of a wild-type allele product of the endogenous TDF1 gene or homolog thereof.
5. The modified plant, plant seed, plant part, or plant cell of claim 2, wherein the modification is located within a coding region, non-coding region, or any combination thereof, of the endogenous TDF1 gene or homolog thereof.
6. The modified plant, plant seed, plant part, or plant cell of claim 2, wherein the modification is located in an exon region of said TDF1 gene or homolog thereof.
7. The modified plant, plant seed, plant part, or plant cell of claim 2, wherein the modification is located in an intron region of said TDF1 gene or homolog thereof.
8. The modified plant, plant seed, plant part, or plant cell of claim 2, wherein the modification is located in an exon region and an intron region of said TDF1 gene or homolog thereof.
9. The modified plant, plant seed, plant part, or plant cell of claim 2, wherein the plant is a leguminous plant, or wherein the plant seed, plant part, or plant cell is a plant seed, plant part, or plant cell of a leguminous plant.
10. The modified plant, plant seed, plant part, or plant cell of claim 9, wherein the leguminous plant is a soybean plant, a bean plant, a pea plant, a chickpea plant, an alfalfa plant, a peanut plant, a carob plant, a lentil plant, or a licorice plant.
11. The modified plant, plant seed, plant part, or plant cell of claim 10, wherein the leguminous plant is a soybean plant.
12. The modified plant, plant seed, plant part, or plant cell of claim 11, wherein the TDF1 gene is GmTDFl-1.
13. The modified plant, plant seed, plant part, or plant cell of claim 2, wherein the plant, plant seed, plant part, or plant cell is heterozygous for the modification.
14. The modified plant, plant seed, plant part, or plant cell of claim 2, wherein the plant, plant seed, plant part, or plant cell is homozygous for the modification.
15. The modified plant, plant seed, plant part, or plant cell of claim 2, wherein the plant, plant seed, plant part, or plant cell comprises a first modification in a first allele of the TDF1 gene and a second modification in a second allele of the TDF1 gene, the first modification and the second modification being different from one another.
16. The modified plant, plant seed, plant part, or plant cell of claim 2, wherein the modification comprises a deletion, an insertion, a substitution, an inversion, a duplication, or a combination of any thereof.
17. The modified plant, plant seed, plant part, or plant cell of claim 16, wherein the modification is a deletion.
18. The modified plant, plant seed, plant part, or plant cell of claim 17, wherein the deletion comprises between 1 nucleotide and 1500 nucleotides.
19. The modified plant, plant seed, plant part, or plant cell of claim 18, wherein the deletion comprises a region of at least 1 nucleotide, at least 2 nucleotides, at least 5 nucleotides, at least 10 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, at least 500 nucleotides, at least 600 nucleotides, at least 700 nucleotides, at least 800 nucleotides, at least 900 nucleotides, at least 1000 nucleotides, or at least 1500 nucleotides.
20. The modified plant, plant seed, plant part, or plant cell of claim 16, wherein the modification is an inversion.
21. The modified plant, plant seed, plant part, or plant cell of claim 20, wherein the inversion is a region of at least 5 consecutive nucleotides, at least 10 consecutive nucleotides, at least 20 consecutive nucleotides, at least 30 consecutive nucleotides, at least 40 consecutive nucleotides, at least 50 consecutive nucleotides, at least 100 consecutive nucleotides, at least 200 consecutive nucleotides, at least 300 consecutive nucleotides, at least 400 consecutive nucleotides, at least 500 consecutive nucleotides, at least 600 consecutive nucleotides, at least 700 consecutive nucleotides, at least 800 consecutive nucleotides, at least 900 consecutive nucleotides, at least 1000 consecutive nucleotides, or at least 1 100 consecutive nucleotides.
22. The modified plant, plant seed, plant part, or plant cell of claim 12, wherein the plant, plant seed, plant part, or plant cell comprises a modification in at least one allele of the GmTDFl -1 gene, wherein the modification is selected from the group consisting of: an 8 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:26; a 10 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:27; a 1 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:28; a 6 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:29; a 9 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:30; an 8 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:31; an 11 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:32; a 7 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:33; a 12 base pair deletion and an 8 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO: 34; two inversions, wherein the resulting nucleotide sequence is SEQ ID NO:35; a 15 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:36; a 19 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:37; a 107 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:38;
a 27 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:39; a 9 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:40; a 9 base pair deletion and a 10 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:41; a 1,025 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:42; an 11 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:43; a 7 base pair deletion and an 11 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO: 44; an inversion, wherein the resulting nucleotide sequence is SEQ ID NO:45; a 1,055 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:46; an 11 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:47; a 16 base pair indel and a 9 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:48; a 122 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:49; a 65 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:50; an inversion, wherein the resulting nucleotide sequence is SEQ ID NO:51; a 4 base pair deletion and a 36 base pair deletion, wherein the resulting nucleotide sequence is SEQ ID NO:52; two inversions, wherein the resulting nucleotide sequence is SEQ ID NO:53; and any combinations of any thereof.
23. The modified plant, plant seed, plant part, or plant cell of claim 12, wherein the plant, plant seed, plant part, or plant cell comprises at least one polynucleotide sequence selected from the group consisting of SEQ ID NOs:26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, and 53.
24. The modified plant, plant seed, plant part, or plant cell of claim 12, wherein the plant, plant seed, plant part, or plant cell comprises at least one polynucleotide sequence selected from the group consisting of SEQ ID NOs:26, 27, 28, 29, 30, 31, 32, and 33.
25. The modified plant, plant seed, plant part, or plant cell of claim 12, wherein the plant, plant seed, plant part, or plant cell comprises at least one polynucleotide sequence selected from the group consisting of SEQ ID NOs:34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, and 53.
26. The modified plant of claim 1, wherein the modification results in a male sterility phenotype in the plant, as compared to the male fertile phenotype of an otherwise identical plant that lacks the modification.
27. The modified plant of claim 12, wherein the modification results in a male sterility phenotype in the plant, as compared to the male fertile phenotype of an otherwise identical plant that lacks the modification.
28. The modified plant of claim 23, wherein the modification results in a male sterility phenotype in the plant, as compared to the male fertile phenotype of an otherwise identical plant that lacks the modification.
29. A polynucleotide comprising a sequence selected from the group consisting of SEQ ID NOs:26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, and 53.
30. A guide RNA comprising a polynucleotide sequence selected from the group consisting of SEQ ID NOs:23, 24, and 25.
31. A method for producing a modified plant having a male sterility phenotype, the method comprising: a) introducing a modification into at least one target site in an endogenous TDF1 gene or a homolog thereof of a plant cell; b) identifying and selecting one or more plant cells of step a) comprising said modification in said TDF1 gene or homolog thereof; and c) regenerating at least one plant from at least one or more cells selected in step b).
32. The method of claim 31 , wherein the target site is located in a coding region, a non-coding region, or any combination thereof, of an endogenous TDF1 gene or homolog thereof.
33. The method of claim 31, wherein the modification is introduced by a site-specific genome modification enzyme selected from the group consisting of: an RNA-guided nuclease, azinc-finger nuclease, a meganuclease, a TALE-nuclease, a recombinase, a transposase, and combinations of any thereof.
34. The method of claim 33, wherein the site-specific genome modification enzyme is an RNA-guided nuclease comprising a Cas nuclease, a Cpfl nuclease, or a variant of either thereof.
35. The method of claim 34, wherein the site-specific genome modification enzyme is an RNA-guided nuclease comprising a Cpfl nuclease.
36. The method of claim 31, wherein the modification is introduced by a site-specific genome modification enzyme that creates at least one strand break at the target site.
37. The method of claim 31, wherein the modification is selected from the group consisting of a substitution, an insertion, an inversion, a deletion, a duplication, and a combination thereof.
38. The method of claim 37, wherein the modification is a deletion.
39. The method of claim 38, wherein the deletion comprises a region of at least 1 nucleotide, at least 2 nucleotides, at least 5 nucleotides, at least 10 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, at least 500 nucleotides, at least 600 nucleotides, at least 700 nucleotides, at least 800 nucleotides, at least 900 nucleotides, or at least 1000 nucleotides, or at least 1500 nucleotides.
40. The method of claim 37, wherein the modification is an inversion.
41. The method of claim 40, wherein the inversion is a region of at least 5 consecutive nucleotides, at least 10 consecutive nucleotides, at least 20 consecutive nucleotides, at least 30 consecutive nucleotides, at least 40 consecutive nucleotides, at least 50 consecutive nucleotides, at least 100 consecutive nucleotides, at least 200 consecutive nucleotides, at least 300 consecutive nucleotides, at least 400 consecutive nucleotides, at least 500 consecutive nucleotides, at least 600 consecutive nucleotides, at least 700 consecutive nucleotides, at least 800 consecutive nucleotides,
at least 900 consecutive nucleotides, at least 1000 consecutive nucleotides, or at least 1100 consecutive nucleotides.
42. The method of claim 31, wherein the plant cell of step a) is a soybean plant cell and the produced modified plant having a male sterility phenotype is a soybean plant.
43. A male sterile soybean plant produced by the method of claim 42.
44. A method for producing an Fi hybrid soybean plant comprising crossing a male-sterile soybean plant comprising a modified GmTDFl-1 gene with a second, non-isogenic, male-fertile soybean plant to produce an Fi hybrid soybean plant.
45. An Fi hybrid soybean plant produced by the method of claim 44.
46. A method for producing an Fi hybrid soybean seed, comprising crossing the plant of claim 1 with a second, non-isogenic, male-fertile soybean plant and harvesting the resultant Fi hybrid soybean seed, wherein an Fi hybrid soybean seed is produced.
47. An Fi hybrid soybean seed produced by the method of claim 46.
48. A method for producing an Fi hybrid soybean seed, comprising crossing the plant of claim 12 with a second, non-isogenic, male-fertile soybean plant and harvesting the resultant Fi hybrid soybean seed, wherein an Fi hybrid soybean seed is produced.
49. An Fi hybrid soybean seed produced by the method of claim 48.
50. A modified plant, plant seed, plant part, or plant cell, wherein the plant, plant seed, plant part, or plant cell comprises at least one polynucleotide sequence selected from the group consisting of SEQ ID NOs:26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, and 53.
51. The modified plant, plant seed, plant part, or plant cell of claim 50, wherein the plant, plant seed, plant part, or plant cell is a transgenic soybean plant, soybean plant seed, soybean plant part, or soybean plant cell.
52. The modified plant, plant seed, plant part, or plant cell of claim 12, wherein the plant, plant seed, plant part, or plant cell comprises a first modification in a first allele of the TDF1-1 gene and a second modification in a second allele of the TDF1-1 gene, the first modification and the second modification being different from one another.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202363507916P | 2023-06-13 | 2023-06-13 | |
| US63/507,916 | 2023-06-13 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024258814A1 true WO2024258814A1 (en) | 2024-12-19 |
Family
ID=93852623
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2024/033340 Ceased WO2024258814A1 (en) | 2023-06-13 | 2024-06-11 | Compositions and methods for engineering male sterility in plants |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2024258814A1 (en) |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20130333061A1 (en) * | 2008-02-05 | 2013-12-12 | Wei Wu | Isolated novel nucleic acid and protein molecules from soy and methods of using those molecules to generate transgenic plants with enhanced agronomic traits |
| US20160017365A1 (en) * | 2013-03-12 | 2016-01-21 | Pioneer Hi-Bred International, Inc. | Manipulation of dominant male sterility |
| US20190024102A1 (en) * | 2012-03-13 | 2019-01-24 | Pioneer Hi-Bred International, Inc. | Genetic reduction of male fertility in plants |
| JP2022076870A (en) * | 2020-11-10 | 2022-05-20 | 国立大学法人 新潟大学 | Polynucleotides involved in pollen formation, their use, and methods for determining male sterility using this nucleotide sequence. |
-
2024
- 2024-06-11 WO PCT/US2024/033340 patent/WO2024258814A1/en not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20130333061A1 (en) * | 2008-02-05 | 2013-12-12 | Wei Wu | Isolated novel nucleic acid and protein molecules from soy and methods of using those molecules to generate transgenic plants with enhanced agronomic traits |
| US20190024102A1 (en) * | 2012-03-13 | 2019-01-24 | Pioneer Hi-Bred International, Inc. | Genetic reduction of male fertility in plants |
| US20160017365A1 (en) * | 2013-03-12 | 2016-01-21 | Pioneer Hi-Bred International, Inc. | Manipulation of dominant male sterility |
| JP2022076870A (en) * | 2020-11-10 | 2022-05-20 | 国立大学法人 新潟大学 | Polynucleotides involved in pollen formation, their use, and methods for determining male sterility using this nucleotide sequence. |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12234467B2 (en) | Diplospory gene | |
| AU2021286555B2 (en) | Heterozygous CENH3 monocots and methods of use thereof for haploid induction and simultaneous genome editing | |
| US12024711B2 (en) | Methods and compositions for generating dominant short stature alleles using genome editing | |
| EP3490365A1 (en) | Wheat | |
| CN116529376A (en) | Fertility-related gene and application thereof in cross breeding | |
| EP4376593A1 (en) | Methods and compositions relating to maintainer lines for male-sterility | |
| US20230183725A1 (en) | Method for obtaining mutant plants by targeted mutagenesis | |
| CN113754746A (en) | Rice male fertility regulating gene, its application and method for regulating rice fertility using CRISPR-Cas9 | |
| EP4499674A2 (en) | Compositions and methods for enhancing corn traits and yield using genome editing | |
| US20240376489A1 (en) | Compositions and methods for altering plant determinacy | |
| US20230235350A1 (en) | Compositions and methods for altering plant determinacy | |
| WO2024129512A2 (en) | Compositions and methods for site-directed integration | |
| US20230313216A1 (en) | Compositions and methods for enhancing corn traits and yield using genome editing | |
| US20230340517A1 (en) | Compositions and methods for enhancing corn traits and yield using genome editing | |
| WO2025072302A1 (en) | Methods for regulating self-incompatibility for generating hybrids | |
| EP4453219A1 (en) | Regulatory nucleic acid molecules for modifying gene expression in cereal plants |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24823985 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |










