EP3373726A1 - Methods and systems for trait introgression - Google Patents

Methods and systems for trait introgression

Info

Publication number
EP3373726A1
EP3373726A1 EP16864760.0A EP16864760A EP3373726A1 EP 3373726 A1 EP3373726 A1 EP 3373726A1 EP 16864760 A EP16864760 A EP 16864760A EP 3373726 A1 EP3373726 A1 EP 3373726A1
Authority
EP
European Patent Office
Prior art keywords
snp
genome
plants
plant
generation
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Withdrawn
Application number
EP16864760.0A
Other languages
German (de)
French (fr)
Inventor
Cherie OCHSENFELD
Tyler Mansfield
Clive EVANS
Jia YI
Pradeep MARRI
Jennifer Changhong TANG
Jennifer L. Hamilton
Jenelle Meyer
Steve ROUNSLEY
James W. Bing
Rebecca AUS
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Corteva Agriscience LLC
Original Assignee
Dow AgroSciences LLC
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Dow AgroSciences LLC filed Critical Dow AgroSciences LLC
Publication of EP3373726A1 publication Critical patent/EP3373726A1/en
Withdrawn legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B20/00ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
    • G16B20/20Allele or variant detection, e.g. single nucleotide polymorphism [SNP] detection
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B20/00ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B30/00ICT specially adapted for sequence analysis involving nucleotides or amino acids
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B5/00ICT specially adapted for modelling or simulations in systems biology, e.g. gene-regulatory networks, protein interaction networks or metabolic networks

Definitions

  • Quantitative traits are continuously varying due to genetic and environmental influences. Quantitative traits may be distinguished from “qualitative” or “discrete” traits on the basis of two factors: environmental influences on gene expression that produce a continuous distribution of phenotypes; and the complex segregation pattern produced by multi-genic inheritance. The identification of one or more regions of the genome linked to the expression of a quantitative trait led to the discovery of Quantitative Trait Loci (QTL).
  • QTL Quantitative Trait Loci
  • RAPD random-amplified polymorphic DNA
  • RFLP restriction fragment length polymorphism
  • This invention is related to methods and systems for trait introgression, and/or predicting/evaluating numbers of non-rogue plants and rogue plants after back-crossing, where genetic markers are used without evaluating the phenotypes of physical plants.
  • a computerized method for trait introgression, predicting/evaluating numbers of non- rogue plants and rogue plants after back-crossing, and/or selecting plants with the most advantageous genetic profiles comprises:
  • step (b) obtaining sequences from at least one back-crossed plant sample; (c) comparing the sequences obtained in step (b) with the genetic markers database of step (a) for imputation of genetic markers;
  • the method provided further comprises the step of selecting a pre-determined number of plants with the most advantageous genetic profiles for crossing.
  • the pre-determined number is between 5 and 50; between 10 and 100; between 10 and 20; between 10 and 15; or between 5 and 15.
  • the genetic markers comprise single nucleotide polymorphisms (SNPs).
  • the genetic markers database is a SNP database or SNP library.
  • the SNP database or SNP library comprises a genome- wide SNP collection having between 50 and 500; between 100 and 500; between 20 and 200; or between 25 and 300 SNPs.
  • the SNP database or SNP library comprises a genome-wide SNP collection of at least 1,000, 5,000, 10,000, 25,000, 50,000, or 100,000 SNPs. In another embodiment, the SNP database or SNP library comprises a genome- wide SNP collection having between 1,000 and 5,000; between 3,000 and 20,000; between 5,000 and 50,000; or between 10,000 and 100,000 SNPs.
  • the plant is selected from soybean, maize, canola, cotton, wheat, sunflower, and rice.
  • the back-crossed plant sample is from a first generation, second generation, third generation, fourth generation, fifth generation back-crossing plant, or combinations thereof.
  • the comparing step of step (c) uses a Genotyping Imputation Module.
  • the predicting step of step (d) uses a Marker Study Manager module.
  • the Marker Study Manager module provides visualization output for predicting/evaluating numbers of non-rogue plants and rogue plants.
  • the Marker Study Manager module provides visualization output for selecting a pre-determined number of plants with the most advantageous genetic profiles for crossing.
  • the computerized system for trait introgression, and/or predicting/evaluating numbers of non-rogue plants and rogue plants after back-crossing.
  • the computerized system comprises:
  • the Marker Study Manager module provides visualization output for predicting/evaluating numbers of non-rogue plants and rogue plants. In another embodiment, the Marker Study Manager module provides visualization output for selecting a pre-determined number of plants with the most advantageous genetic profiles for crossing. In a further embodiment, the pre-determined number is between 5 and 50; between 10 and 100; between 10 and 20; between 10 and 15; or between 5 and 15.
  • the genetic markers database is a SNP database or SNP library.
  • the SNP database or SNP library comprises a genome-wide SNP collection having between 50 and 500; between 100 and 500; between 20 and 200; or between 25 and 300 SNPs.
  • the SNP database or SNP library comprises a genome- wide SNP collection of at least 1,000, 5,000, 10,000, 25,000, 50,000, or 100,000 SNPs. In another embodiment, the SNP database or SNP library comprises a genome-wide SNP collection having between 1,000 and 5,000; between 3,000 and 20,000; between 5,000 and 50,000; or between 10,000 and 100,000 SNPs.
  • the genetic markers imputation module is a SNP Imputation Module.
  • the plant is selected from soybean, maize, canola, cotton, wheat, sunflower, and rice.
  • the back-crossed plant sample is from a first generation, second generation, third generation, fourth generation, fifth generation back-crossing plant, or combinations thereof.
  • a process for use in a computerized system for trait introgression, predicting/evaluating numbers of non-rogue plants and rogue plants after back- crossing, and/or selecting plants with the most advantageous genetic profiles comprises:
  • a computerized method for trait introgression, predicting/evaluating numbers of non-rogue plants and rogue plants after back-crossing, and/or selecting plants with the most advantageous genetic profiles comprises:
  • step (c) comparing the sequences obtained in step (b) with the genome-wide genetic markers information from parent plants of step (a) for imputation of genetic markers;
  • the method provided further comprises the step of selecting a pre-determined number of plants with the most advantageous genetic profiles for crossing.
  • the pre-determined number is between 5 and 50; between 10 and 100; between 10 and 20; between 10 and 15; or between 5 and 15.
  • the genetic markers comprise single nucleotide polymorphisms (SNPs).
  • the genome-wide genetic markers information from parent plants comprises a genome- wide SNP collection having between 50 and 500; between 100 and 500; between 20 and 200; or between 25 and 300 SNPs.
  • the genome-wide genetic markers information from parent plants comprises a genome- wide SNP collection of at least 1,000, 5,000, 10,000, 25,000, 50,000, or 100,000 SNPs.
  • the genome-wide genetic markers information from parent plants comprises a genome- wide SNP collection having between 1,000 and 5,000; between 3,000 and 20,000; between 5,000 and 50,000; or between 10,000 and 100,000 SNPs.
  • the plant is selected from soybean, maize, canola, cotton, wheat, sunflower, and rice.
  • the back-crossed plant sample is from a first generation, second generation, third generation, fourth generation, fifth generation back-crossing plant, or combinations thereof.
  • the comparing step of step (c) uses a Genotyping Imputation Module.
  • the predicting step of step (d) uses a Marker Study Manager module.
  • the Marker Study Manager module provides visualization output for predicting/evaluating numbers of non-rogue plants and rogue plants.
  • the Marker Study Manager module provides visualization output for selecting a pre-determined number of plants with the most advantageous genetic profiles for crossing.
  • FIG. 1 provides an illustration of the flow chart for the Genotyping Imputation Module provided herein.
  • the methods provided comprise obtaining a SNP database, analyzing sequences against such SNP database, and predicting the outcome between non-rogue plants and rogue plants.
  • the systems provided comprise a SNP database, a Genotype Imputation Module, and a Marker Study Manager module. The methods and/or systems provided will allow users to calculate a more realistic statistical measure to ensure high quality conversions.
  • SNPs single nucleotide polymorphisms
  • SNPs are preferred because technologies are available for automated, high-throughput screening of SNP markers, which can decrease the time to select for and introgress desired trait(s) in plants.
  • SNP markers are ideal because the likelihood that a particular SNP allele is derived from independent origins in the extant population of a particular species is relatively low. Thus, SNP markers are useful for tracking and assisting introgression of alleles associated with desired trait(s).
  • SNP single nucleotide polymorphism
  • NGS Genetics Translation
  • marker assisted trait introgression heavily relies on availability of polymorphic markers at regular intervals across the entire genome to evaluate genetic patterns of relatedness to the recurrent elite parent for high quality conversions.
  • trait introgression involves selection of narrow cross combinations between a recurrent parent and a donor, aims to decrease the amount of linkage drag that could potentially come from a genetically distant donor lines.
  • narrow crosses may limit the number of available polymorphic markers to be used in trait conversions due to genetic similarities between the parents.
  • high-density SNP database/library is provided for such narrow crosses.
  • the SNP database/library provided herein comprises a genome- wide SNP collection of at least 1,000, 5,000, 10,000, 25,000, 50,000, or 100,000 SNPs.
  • the SNP database/library comprises a genome- wide SNP collection of at least 1,000, 5,000, 10,000, 25,000, 50,000, or 100,000 SNPs.
  • the SNP database/library comprises a genome- wide SNP collection of at least 1,000, 5,000, 10,000, 25,000, 50,000, or 100,000 S
  • database/library provided herein comprises a genome- wide SNP collection between 1,000 and 5,000; between 3,000 and 20,000; between 5,000 and 50,000; or between 10,000 and 100,000.
  • the phrase "vector" refers to a piece of DNA, typically double- stranded, which can have inserted into it a piece of foreign DNA.
  • the vector can be for example, of plasmid or viral origin, which typically encodes a selectable or screenable marker or transgenes.
  • the vector is used to transport the foreign or heterologous DNA into a suitable host cell. Once in the host cell, the vector can replicate independently of or coincidental with the host chromosomal DNA. Alternatively, the vector can target insertion of the foreign or heterologous DNA into a host chromosome.
  • transgene vector refers to a vector that contains an inserted segment of DNA, the "transgene” that is transcribed into mRNA or replicated as a RNA within a host cell.
  • transgene refers not only to that portion of inserted DNA that is converted into RNA, but also those portions of the vector that are necessary for the transcription or replication of the RNA.
  • a transgene typically comprises a gene-of-interest but needs not necessarily comprise a polynucleotide sequence that contains an open reading frame capable of producing a protein.
  • transformant or “transgenic” refers to plant cells, plants, and the like that have been transformed or have undergone a transformation procedure.
  • the introduced DNA is usually in the form of a vector containing an inserted piece of DNA.
  • transgenic plant refers to a plant whose genome has been altered by the stable integration of recombinant DNA.
  • a transgenic plant includes a plant regenerated from an originally-transformed plant cell and progeny transgenic plants from later generations or crosses of a transformed plant.
  • recombinant DNA refers to DNA which has been genetically engineered and constructed outside of a cell including DNA containing naturally occurring DNA or cDNA or synthetic DNA.
  • locus refers to a short sequence that is usually unique and usually found at one particular location in the genome by a point of reference; for example a short DNA sequence that is a gene, or part of a gene or intergenic region.
  • a locus can be a unique PCR product at a particular location in the genome.
  • the loci may comprise one or more polymorphisms.
  • genetic locus refers to a location on a chromosome.
  • genomic locus refers to a location within the entire set of chromosomes of an organism.
  • the phrase "marker” refers to a locus on a chromosome that serves to identify a unique position on the chromosome.
  • a genotype may be defined by use of one or a plurality of markers.
  • a “marker” is a polymorphic nucleic acid sequence or nucleic acid feature.
  • a “polymorphism” is a variation among individuals in sequence, particularly in DNA sequence, or feature, such as a transcriptional profile or methylation pattern.
  • Useful polymorphisms include single nucleotide polymorphisms (SNPs), insertions or deletions in DNA sequence (Indels), simple sequence repeats of DNA sequence (SSRs) a restriction fragment length polymorphism, a haplotype, and a tag SNP.
  • SNPs single nucleotide polymorphisms
  • Indels DNA sequence
  • SSRs simple sequence repeats of DNA sequence
  • a genetic marker, a gene, a DNA-derived sequence, a RNA-derived sequence, a promoter, a 5' untranslated region of a gene, a 3' untranslated region of a gene, microRNA, siRNA, a QTL, a satellite marker, a transgene, mRNA, ds mRNA, a transcriptional profile, and a methylation pattern may comprise polymorphisms.
  • a "marker” can be a detectable characteristic that can be used to discriminate between heritable differences between organisms.
  • characteristics may include genetic markers, protein composition, protein levels, oil composition, oil levels, carbohydrate composition, carbohydrate levels, fatty acid composition, fatty acid levels, amino acid composition, amino acid levels, biopolymers, pharmaceuticals, starch composition, starch levels, fermentable starch, fermentation yield, fermentation efficiency, energy yield, secondary compounds, metabolites, morphological characteristics, and agronomic characteristics.
  • the phrase "marker assay” refers to a method for detecting a polymorphism at a particular locus using a particular method, including measurement of at least one phenotype (for example seed color, flower color, or other visually detectable trait), genotyping, restriction fragment length polymorphism (RFLP), single base extension, electrophoresis, sequence alignment, allelic specific oligonucleotide hybridization (ASO), random amplified polymorphic DNA (RAPD), microarray-based technologies, and nucleic acid sequencing technologies, etc.
  • phenotype for example seed color, flower color, or other visually detectable trait
  • genotyping for example seed color, flower color, or other visually detectable trait
  • RFLP restriction fragment length polymorphism
  • ASO allelic specific oligonucleotide hybridization
  • RAPD random amplified polymorphic DNA
  • microarray-based technologies and nucleic acid sequencing technologies, etc.
  • allele refers to an alternative sequence at a particular locus; the length of an allele can be as small as 1 nucleotide base, but is typically larger. Allelic sequence can be amino acid sequence or nucleic acid sequence.
  • single nucleotide polymorphism refers to a polymorphism at a single site wherein the polymorphism constitutes a single base pair change, an insertion of one or more base pairs, or a deletion of one or more base pairs.
  • the phrase "genotype" means the genetic component of the phenotype and it can be indirectly characterized using markers or directly characterized by nucleic acid sequencing. Suitable markers include a phenotypic character, a metabolic profile, a genetic marker, or some other type of marker. A genotype may constitute an allele for at least one genetic marker locus or a haplotype for at least one haplotype window. In some
  • a genotype may represent a single locus and in others it may represent a genome - wide set of loci.
  • the genotype can reflect the sequence of a portion of a chromosome, an entire chromosome, a portion of the genome, and the entire genome.
  • phenotype refers to the detectable characteristics of a cell or organism which are a manifestation of gene expression.
  • linkage refers to relative frequency at which types of gametes are produced in a cross.
  • linkage disequilibrium refers to a statistical association between two loci or between a trait and a marker.
  • Quantitative trait locus or “QTL” refers to a locus that controls to some degree numerically representable traits that are usually continuously distributed.
  • the phrase "allelic state” refers to the nucleic acid sequence that is present in a nucleic acid molecule that contains a genomic polymorphism.
  • the nucleic acid sequence of a DNA molecule that contains a single nucleotide polymorphism may comprise an A, C, G, or T residue at the polymorphic position such that the allelic state is defined by which residue is present at the polymorphic position.
  • the nucleic acid sequence of an RNA molecule that contains a single nucleotide polymorphism may comprise an A, C, G, or U residue at the polymorphic position such that the allelic state is defined by which residue is present at the polymorphic position.
  • nucleic acid sequence of a nucleic acid molecule that contains an Indel may comprise an insertion or deletion of nucleic acid sequences at the polymorphic position such that the allelic state is defined by the presence or absence of the insertion or deletion at the polymorphic position.
  • association when used in reference to a polymorphism and a phenotypic trait or trait index, refers to any statistically significant correlation between the presence of a given allele of a polymorphic locus and the phenotypic trait or trait index value, wherein the value may be qualitative or quantitative.
  • the phrase "typing” refers to any method whereby the specific allelic form of a given soybean genomic polymorphism is determined.
  • a single nucleotide polymorphism SNP
  • Indels Insertion/deletions
  • Indels can be typed by a variety of assays including, but not limited to, marker assays.
  • an elite line refers to any line that has resulted from breeding and selection for superior agronomic performance.
  • An elite plant is any plant from an elite line.
  • the phrase "plant” includes dicotyledons plants and monocotyledons plants.
  • dicotyledons plants include tobacco, Arabidopsis, soybean, tomato, papaya, canola, sunflower, cotton, alfalfa, potato, grapevine, pigeon pea, pea, Brassica, chickpea, sugar beet, rapeseed, watermelon, melon, pepper, peanut, pumpkin, radish, spinach, squash, broccoli, cabbage, carrot, cauliflower, celery, Chinese cabbage, cucumber, eggplant, and lettuce.
  • Examples of monocotyledons plants include corn, rice, wheat, sugarcane, barley, rye, sorghum, orchids, bamboo, banana, cattails, lilies, oat, onion, millet, and triticale.
  • plant also includes a whole plant and any descendant, cell, tissue, or part of a plant.
  • plant parts include any part(s) of a plant, including, for example and without limitation: seed (including mature seed and immature seed); a plant cutting; a plant cell; a plant cell culture; a plant organ (e.g., pollen, embryos, flowers, fruits, shoots, leaves, roots, stems, and explants).
  • a plant tissue or plant organ may be a seed, callus, or any other group of plant cells that is organized into a structural or functional unit.
  • a plant cell or tissue culture may be capable of regenerating a plant having the physiological and morphological characteristics of the plant from which the cell or tissue was obtained, and of regenerating a plant having substantially the same genotype as the plant. In contrast, some plant cells are not capable of being regenerated to produce plants.
  • Regenerable cells in a plant cell or tissue culture may be embryos, protoplasts, meristematic cells, callus, pollen, leaves, anthers, roots, root tips, silk, flowers, kernels, ears, cobs, husks, or stalks.
  • Plant parts include harvestable parts and parts useful for propagation of progeny plants.
  • Plant parts useful for propagation include, for example and without limitation: seed; fruit; a cutting; a seedling; a tuber; and a rootstock.
  • a harvestable part of a plant may be any useful part of a plant, including, for example and without limitation: flower; pollen; seedling; tuber; leaf; stem; fruit; seed; and root.
  • plant cell described or “transformed plant cell” refers to a plant cell that is transformed with stably- integrated, non-natural, recombinant DNA, e.g., by Agrobacterium-mediated transformation or by bombardment using microparticles coated with recombinant DNA or other means.
  • a plant cell of this invention can be an originally- transformed plant cell that exists as a microorganism or as a progeny plant cell that is
  • regenerated into differentiated tissue e.g., into a transgenic plant with stably-integrated, non- natural recombinant DNA, or seed or pollen derived from a progeny transgenic plant.
  • the phrase "trait” refers to a physiological, morphological, biochemical, or physical characteristic of a plant or particular plant material or cell. In some instances, this characteristic is visible to the human eye, including seed or plant size, or can be measured by biochemical techniques, including detecting the protein, starch, or oil content of seed or leaves, or by observation of a metabolic or physiological process (e.g., by measuring uptake of carbon dioxide), or by the observation of the expression level of a gene or genes (e.g., by employing Northern analysis, RT-PCR, microarray gene expression assays), or reporter gene expression systems, or by agricultural observations including stress tolerance, yield, or pathogen tolerance.
  • biochemical techniques including detecting the protein, starch, or oil content of seed or leaves, or by observation of a metabolic or physiological process (e.g., by measuring uptake of carbon dioxide), or by the observation of the expression level of a gene or genes (e.g., by employing Northern analysis, RT-PCR, microarray gene expression assays), or reporter gene expression
  • a “nongenic sequence” or “nongenic genomic sequence” is a native DNA sequence found in the nuclear genome of a plant, having a length of at least 1 Kb, and devoid of any open reading frames, gene sequences, or gene regulatory sequences. Furthermore, the nongenic sequence does not comprise any intron sequence (i.e., introns are excluded from the definition of nongenic). The nongenic sequence cannot be transcribed or translated into protein. Many plant genomes contain nongenic regions, where as much as 95% of the genome can be nongenic, and these regions may be comprised of mainly repetitive DNA.
  • a "genie region” is defined as a polynucleotide sequence that comprises an open reading frame encoding an RNA and/or polypeptide.
  • the genie region may also encompass any identifiable adjacent 5' and 3' non-coding nucleotide sequences involved in the regulation of expression of the open reading frame up to about 2 Kb upstream of the coding region and 1 Kb downstream of the coding region, but possibly further upstream or downstream.
  • a genie region further includes any introns that may be present in the genie region.
  • the genie region may comprise a single gene sequence, or multiple gene sequences interspersed with short spans (less than 1 Kb) of nongenic sequences.
  • the first set has relatively low genome coverage between the two parents XJA40 x 4XP811XT - similar parents where the genome coverage at BC2 is 62% at RPP4; and is 33% at RPP6).
  • the second set has relatively high genome coverage between the two parents LDS51 X MV8735XT - dissimilar parents where the genome coverage at BC2 is 99% at RPP4; and 85% at RPP6).
  • the SNP database provides reference genomes to be compared with individual sequences, and in this example generated 15,636 genome-wide markers for BC2, 14,824 genome-wide markers for BC3, and 15,402 genome-wide markers for BC4 (for the LDS51 X MV8735XT set).
  • the genotyping data (generated using methods known in the art) from the two sets of BC2, BC3, and BC4 generations are compared with the SNP database for the imputation of genotypic data.
  • the flow chart for the Genotyping Imputation Module is illustrated in FIG. 1, where such algorithm is executed in a computer system.
  • the resulting genotyping information is used by the Marker Study Manager module to generate user-friendly visualization of the output for predicting/evaluating numbers of non-rogue plants and rogue plants.
  • a Marker Study Manager (MSM) module is developed using algorithm with a visualization tool.
  • the genotyping data after imputation is channeled through the MSM module to build a chromosome table with conversion metrics. Visualization of such chromosome table enables a user-friendly way to make selections, where the MSM module can generate
  • chromosome tables having 3500 to 15500 SNP markers in each sample in this example.
  • Table 1 summarizes the different populations and the population size of each generation used in this example.
  • the output in Table 1 is solely based on analysis of molecular markers instead of observation from the phenotypes of the physical plants.
  • the genotyping imputation module successfully differentiates rogue plants from non- rogue plants.
  • the rogue plants can be visualized using the Marker Study Manager module with the larger genome view.
  • identification of rogue alleles and/or large number of SNP markers can significantly improve imputation and prediction.
  • the method and system used in this example can process genetic data from thousands of SNP markers which increases the genome coverage required for informed selection decisions especially in narrow crosses.
  • the method and system provided are cost effective because there is no need to design new markers for each project where multiplexing capability is provided.

Landscapes

  • Physics & Mathematics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Evolutionary Biology (AREA)
  • Biophysics (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Medical Informatics (AREA)
  • General Health & Medical Sciences (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Biotechnology (AREA)
  • Analytical Chemistry (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Molecular Biology (AREA)
  • Chemical & Material Sciences (AREA)
  • Genetics & Genomics (AREA)
  • Physiology (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
  • Management, Administration, Business Operations System, And Electronic Commerce (AREA)

Abstract

Provided are methods and/or systems having advantages of cost effective, time saving, and informative user-friendly characteristics to accomplish conversion of an elite inbred into a traited plant without losing agronomic performance. The methods provided comprise obtaining a SNP database, analyzing sequences against such SNP database, and predicting the outcome between non-rogue plants and rogue plants. The systems provided comprise a SNP database, a Genotype Imputation Module, and a Marker Study Manager module. The methods and/or systems provided will allow users to calculate a more realistic statistical measure to ensure high quality conversions.

Description

METHODS AND SYSTEMS FOR TRAIT INTROGRESSION
CROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Patent Application Serial No. 62/253,347, filed November 10, 2015.
BACKGROUND OF THE INVENTION
[0002] Traits that are continuously varying due to genetic and environmental influences are commonly referred to as "quantitative traits." Quantitative traits may be distinguished from "qualitative" or "discrete" traits on the basis of two factors: environmental influences on gene expression that produce a continuous distribution of phenotypes; and the complex segregation pattern produced by multi-genic inheritance. The identification of one or more regions of the genome linked to the expression of a quantitative trait led to the discovery of Quantitative Trait Loci (QTL).
[0003] Different types of molecular markers such as RAPD (random-amplified polymorphic DNA) markers, RFLP (restriction fragment length polymorphism) markers, and SCAR
(sequence-characterized amplified region) markers have been identified. However, most of these markers are low-throughput markers which are not suitable for large-scale screening through automation
[0004] Therefore, there is the need for inventions that are useful to provide high-throughput, high-capacity, and/or high-density approaches for plant breeding and/or trait introgression.
SUMMARY OF THE INVENTION
[0005] This invention is related to methods and systems for trait introgression, and/or predicting/evaluating numbers of non-rogue plants and rogue plants after back-crossing, where genetic markers are used without evaluating the phenotypes of physical plants. In one aspect, provided is a computerized method for trait introgression, predicting/evaluating numbers of non- rogue plants and rogue plants after back-crossing, and/or selecting plants with the most advantageous genetic profiles. The computerized method comprises:
(a) generating a genetic markers database by collecting genome-wide genetic markers
information in a plant;
(b) obtaining sequences from at least one back-crossed plant sample; (c) comparing the sequences obtained in step (b) with the genetic markers database of step (a) for imputation of genetic markers; and
(d) predicting whether the back-crossed plant sample is a rogue or non-rogue plant.
[0006] In one embodiment, the method provided further comprises the step of selecting a pre-determined number of plants with the most advantageous genetic profiles for crossing. In a further embodiment, the pre-determined number is between 5 and 50; between 10 and 100; between 10 and 20; between 10 and 15; or between 5 and 15. In another embodiment, the genetic markers comprise single nucleotide polymorphisms (SNPs). In another embodiment, the genetic markers database is a SNP database or SNP library. In another embodiment, the SNP database or SNP library comprises a genome- wide SNP collection having between 50 and 500; between 100 and 500; between 20 and 200; or between 25 and 300 SNPs. In another embodiment, the SNP database or SNP library comprises a genome-wide SNP collection of at least 1,000, 5,000, 10,000, 25,000, 50,000, or 100,000 SNPs. In another embodiment, the SNP database or SNP library comprises a genome- wide SNP collection having between 1,000 and 5,000; between 3,000 and 20,000; between 5,000 and 50,000; or between 10,000 and 100,000 SNPs. In another embodiment, the plant is selected from soybean, maize, canola, cotton, wheat, sunflower, and rice. In another embodiment, the back-crossed plant sample is from a first generation, second generation, third generation, fourth generation, fifth generation back-crossing plant, or combinations thereof.
[0007] In one embodiment, the comparing step of step (c) uses a Genotyping Imputation Module. In another embodiment, the predicting step of step (d) uses a Marker Study Manager module. In a further embodiment, the Marker Study Manager module provides visualization output for predicting/evaluating numbers of non-rogue plants and rogue plants. In another further embodiment, the Marker Study Manager module provides visualization output for selecting a pre-determined number of plants with the most advantageous genetic profiles for crossing.
[0008] In another aspect, provided is computerized system for trait introgression, and/or predicting/evaluating numbers of non-rogue plants and rogue plants after back-crossing. The computerized system comprises:
(a) a genetic markers database;
(b) a genetic markers imputation module accepting inputs of sequences from at least one back-crossed plant sample; and
(c) a Marker Study Manager module providing visualization output.
[0009] In one embodiment, the Marker Study Manager module provides visualization output for predicting/evaluating numbers of non-rogue plants and rogue plants. In another embodiment, the Marker Study Manager module provides visualization output for selecting a pre-determined number of plants with the most advantageous genetic profiles for crossing. In a further embodiment, the pre-determined number is between 5 and 50; between 10 and 100; between 10 and 20; between 10 and 15; or between 5 and 15. In another embodiment, the genetic markers database is a SNP database or SNP library. In another embodiment, the SNP database or SNP library comprises a genome-wide SNP collection having between 50 and 500; between 100 and 500; between 20 and 200; or between 25 and 300 SNPs. In another embodiment, the SNP database or SNP library comprises a genome- wide SNP collection of at least 1,000, 5,000, 10,000, 25,000, 50,000, or 100,000 SNPs. In another embodiment, the SNP database or SNP library comprises a genome-wide SNP collection having between 1,000 and 5,000; between 3,000 and 20,000; between 5,000 and 50,000; or between 10,000 and 100,000 SNPs. In another embodiment, the genetic markers imputation module is a SNP Imputation Module. In another embodiment, the plant is selected from soybean, maize, canola, cotton, wheat, sunflower, and rice. In another embodiment, the back-crossed plant sample is from a first generation, second generation, third generation, fourth generation, fifth generation back-crossing plant, or combinations thereof.
[0010] In another aspect, provided is a process for use in a computerized system for trait introgression, predicting/evaluating numbers of non-rogue plants and rogue plants after back- crossing, and/or selecting plants with the most advantageous genetic profiles. The process comprises:
(a) inputting sequences of at least one back-crossed plant sample into the system provided herein by an user; and
(b) receiving output from the system provided herein for predicting/evaluating numbers of non-rogue plants and rogue plants.
[0011] In another aspect, provided is a computerized method for trait introgression, predicting/evaluating numbers of non-rogue plants and rogue plants after back-crossing, and/or selecting plants with the most advantageous genetic profiles. The computerized method comprises:
(a) generating genome-wide genetic markers information from parent plants;
(b) obtaining sequences from at least one back-crossed plant sample;
(c) comparing the sequences obtained in step (b) with the genome-wide genetic markers information from parent plants of step (a) for imputation of genetic markers; and
(d) predicting whether the back-crossed plant sample is a rogue or non-rogue plant.
[0012] In one embodiment, the method provided further comprises the step of selecting a pre-determined number of plants with the most advantageous genetic profiles for crossing. In a further embodiment, the pre-determined number is between 5 and 50; between 10 and 100; between 10 and 20; between 10 and 15; or between 5 and 15. In another embodiment, the genetic markers comprise single nucleotide polymorphisms (SNPs). In another embodiment, the genome-wide genetic markers information from parent plants comprises a genome- wide SNP collection having between 50 and 500; between 100 and 500; between 20 and 200; or between 25 and 300 SNPs. In another embodiment, the genome-wide genetic markers information from parent plants comprises a genome- wide SNP collection of at least 1,000, 5,000, 10,000, 25,000, 50,000, or 100,000 SNPs. In another embodiment, the genome-wide genetic markers information from parent plants comprises a genome- wide SNP collection having between 1,000 and 5,000; between 3,000 and 20,000; between 5,000 and 50,000; or between 10,000 and 100,000 SNPs. In another embodiment, the plant is selected from soybean, maize, canola, cotton, wheat, sunflower, and rice. In another embodiment, the back-crossed plant sample is from a first generation, second generation, third generation, fourth generation, fifth generation back-crossing plant, or combinations thereof.
[0013] In one embodiment, the comparing step of step (c) uses a Genotyping Imputation Module. In another embodiment, the predicting step of step (d) uses a Marker Study Manager module. In a further embodiment, the Marker Study Manager module provides visualization output for predicting/evaluating numbers of non-rogue plants and rogue plants. In another further embodiment, the Marker Study Manager module provides visualization output for selecting a pre-determined number of plants with the most advantageous genetic profiles for crossing.
BRIEF DESCRIPTION OF THE DRAWINGS
[0014] FIG. 1 provides an illustration of the flow chart for the Genotyping Imputation Module provided herein.
DETAILED DESCRIPTION OF THE INVENTION
[0015] Provided are methods and/or systems having advantages of cost effective, time saving, and informative user-friendly characteristics to accomplish conversion of an elite inbred into a traited plant without losing agronomic performance. The methods provided comprise obtaining a SNP database, analyzing sequences against such SNP database, and predicting the outcome between non-rogue plants and rogue plants. The systems provided comprise a SNP database, a Genotype Imputation Module, and a Marker Study Manager module. The methods and/or systems provided will allow users to calculate a more realistic statistical measure to ensure high quality conversions.
[0016] Trait introgression and breeding in general can be greatly facilitated by the use of marker-assisted selection. Of the classes of genetic markers, single nucleotide polymorphisms (SNPs) have characteristics which make them preferential to other genetic markers in detecting, selecting for, and introgressing desired trait(s) in plants. SNPs are preferred because technologies are available for automated, high-throughput screening of SNP markers, which can decrease the time to select for and introgress desired trait(s) in plants. Further, SNP markers are ideal because the likelihood that a particular SNP allele is derived from independent origins in the extant population of a particular species is relatively low. Thus, SNP markers are useful for tracking and assisting introgression of alleles associated with desired trait(s).
[0017] Discovery of single nucleotide polymorphism (SNP) markers in plants genome and generation of SNP database can facilitate visualization of the genomes. However, if the SNP database or library does not provide genome-wide coverage, such SNP database or library will not be useful for narrow crosses with high levels of genetic similarity. Also, we are constrained by time and effort to identify and convert polymorphic markers. As Next Generation
Sequencing (NGS) has proven to become an increasingly cost effective way to discover polymorphisms, it can be used as a new tool for generating genome-wide SNP coverage. NGS has also been used for genotyping application as disclosed in Elshire et al., "A robust, simple genotyping-by- sequencing (GBS) approach for high density species" (2011) PLoS One 6(5) el9379 and Sonah et al., "An improved genotyping by sequencing (GBS) approach offering increased versatility and efficiency of SNP discovery and genotyping" (2013) PLoS One 8(1) e54603, the contents of which are hereby incorporated by reference in their entireties. [0018] In addition, marker assisted trait introgression heavily relies on availability of polymorphic markers at regular intervals across the entire genome to evaluate genetic patterns of relatedness to the recurrent elite parent for high quality conversions. Typically trait introgression involves selection of narrow cross combinations between a recurrent parent and a donor, aims to decrease the amount of linkage drag that could potentially come from a genetically distant donor lines. However, such narrow crosses may limit the number of available polymorphic markers to be used in trait conversions due to genetic similarities between the parents. Thus, high-density SNP database/library is provided for such narrow crosses. In some embodiments, the SNP database/library provided herein comprises a genome- wide SNP collection of at least 1,000, 5,000, 10,000, 25,000, 50,000, or 100,000 SNPs. In other embodiments, the SNP
database/library provided herein comprises a genome- wide SNP collection between 1,000 and 5,000; between 3,000 and 20,000; between 5,000 and 50,000; or between 10,000 and 100,000.
[0019] Unless otherwise indicated, all technical and scientific terms used herein have the same meaning as they would to one skilled in the art of the present invention. Practitioners are particularly directed to Sambrook et al. Molecular Cloning: A Laboratory Manual (Second Edition), Cold Spring Harbor Press, Plainview, N.Y., 1989, and Ausubel FM et al. Current Protocols in Molecular Biology, John Wiley & Sons, New York, N.Y., 1993, for definitions and terms of the art. It is to be understood that this invention is not limited to the particular methodology, protocols, and reagents described, as these may vary.
[0020] As used herein, the phrase "about" refers to greater or lesser than the value or range of values stated by 10 percent, but is not intended to designate any value or range of values to only this broader definition. Each value or range of values preceded by the term "about" is also intended to encompass the embodiment of the stated absolute value or range of values
[0021] As used herein, the phrase "vector" refers to a piece of DNA, typically double- stranded, which can have inserted into it a piece of foreign DNA. The vector can be for example, of plasmid or viral origin, which typically encodes a selectable or screenable marker or transgenes. The vector is used to transport the foreign or heterologous DNA into a suitable host cell. Once in the host cell, the vector can replicate independently of or coincidental with the host chromosomal DNA. Alternatively, the vector can target insertion of the foreign or heterologous DNA into a host chromosome.
[0022] As used herein, the phrase "transgene vector" refers to a vector that contains an inserted segment of DNA, the "transgene" that is transcribed into mRNA or replicated as a RNA within a host cell. The phrase "transgene" refers not only to that portion of inserted DNA that is converted into RNA, but also those portions of the vector that are necessary for the transcription or replication of the RNA. A transgene typically comprises a gene-of-interest but needs not necessarily comprise a polynucleotide sequence that contains an open reading frame capable of producing a protein.
[0023] As used herein, the phrase "transformed" or "transformation" refers to the
introduction of DNA into a cell. The phrases "transformant" or "transgenic" refers to plant cells, plants, and the like that have been transformed or have undergone a transformation procedure. The introduced DNA is usually in the form of a vector containing an inserted piece of DNA.
[0024] As used herein, the phrase "transgenic plant" refers to a plant whose genome has been altered by the stable integration of recombinant DNA. A transgenic plant includes a plant regenerated from an originally-transformed plant cell and progeny transgenic plants from later generations or crosses of a transformed plant.
[0025] As used herein, the phrase "recombinant DNA" refers to DNA which has been genetically engineered and constructed outside of a cell including DNA containing naturally occurring DNA or cDNA or synthetic DNA.
[0026] As used herein, the phrase "locus" refers to a short sequence that is usually unique and usually found at one particular location in the genome by a point of reference; for example a short DNA sequence that is a gene, or part of a gene or intergenic region. A locus can be a unique PCR product at a particular location in the genome. The loci may comprise one or more polymorphisms.
[0027] As used herein, the phrase "genetic locus" refers to a location on a chromosome.
[0028] As used herein, the phrase "genomic locus" refers to a location within the entire set of chromosomes of an organism.
[0029] As used herein, the phrase "marker" refers to a locus on a chromosome that serves to identify a unique position on the chromosome. A genotype may be defined by use of one or a plurality of markers. A "marker" is a polymorphic nucleic acid sequence or nucleic acid feature. A "polymorphism" is a variation among individuals in sequence, particularly in DNA sequence, or feature, such as a transcriptional profile or methylation pattern. Useful polymorphisms include single nucleotide polymorphisms (SNPs), insertions or deletions in DNA sequence (Indels), simple sequence repeats of DNA sequence (SSRs) a restriction fragment length polymorphism, a haplotype, and a tag SNP. A genetic marker, a gene, a DNA-derived sequence, a RNA-derived sequence, a promoter, a 5' untranslated region of a gene, a 3' untranslated region of a gene, microRNA, siRNA, a QTL, a satellite marker, a transgene, mRNA, ds mRNA, a transcriptional profile, and a methylation pattern may comprise polymorphisms. In a broader aspect, a "marker" can be a detectable characteristic that can be used to discriminate between heritable differences between organisms. Examples of such characteristics may include genetic markers, protein composition, protein levels, oil composition, oil levels, carbohydrate composition, carbohydrate levels, fatty acid composition, fatty acid levels, amino acid composition, amino acid levels, biopolymers, pharmaceuticals, starch composition, starch levels, fermentable starch, fermentation yield, fermentation efficiency, energy yield, secondary compounds, metabolites, morphological characteristics, and agronomic characteristics.
[0030] As used herein, the phrase "marker assay" refers to a method for detecting a polymorphism at a particular locus using a particular method, including measurement of at least one phenotype (for example seed color, flower color, or other visually detectable trait), genotyping, restriction fragment length polymorphism (RFLP), single base extension, electrophoresis, sequence alignment, allelic specific oligonucleotide hybridization (ASO), random amplified polymorphic DNA (RAPD), microarray-based technologies, and nucleic acid sequencing technologies, etc.
[0031] As used herein, the phrase "allele" refers to an alternative sequence at a particular locus; the length of an allele can be as small as 1 nucleotide base, but is typically larger. Allelic sequence can be amino acid sequence or nucleic acid sequence.
[0032] As used herein, the phrase "single nucleotide polymorphism," or "SNP" refers to a polymorphism at a single site wherein the polymorphism constitutes a single base pair change, an insertion of one or more base pairs, or a deletion of one or more base pairs.
[0033] As used herein, the phrase "genotype"" means the genetic component of the phenotype and it can be indirectly characterized using markers or directly characterized by nucleic acid sequencing. Suitable markers include a phenotypic character, a metabolic profile, a genetic marker, or some other type of marker. A genotype may constitute an allele for at least one genetic marker locus or a haplotype for at least one haplotype window. In some
embodiments, a genotype may represent a single locus and in others it may represent a genome - wide set of loci. In another embodiment, the genotype can reflect the sequence of a portion of a chromosome, an entire chromosome, a portion of the genome, and the entire genome.
[0034] As used herein, the phrase "phenotype" refers to the detectable characteristics of a cell or organism which are a manifestation of gene expression.
[0035] As used herein, the phrase "linkage" refers to relative frequency at which types of gametes are produced in a cross. The phrase "linkage disequilibrium" refers to a statistical association between two loci or between a trait and a marker.
[0036] As used herein, the phrase "quantitative trait locus" or "QTL" refers to a locus that controls to some degree numerically representable traits that are usually continuously distributed.
[0037] As used herein, the phrase "allelic state" refers to the nucleic acid sequence that is present in a nucleic acid molecule that contains a genomic polymorphism. For example, the nucleic acid sequence of a DNA molecule that contains a single nucleotide polymorphism may comprise an A, C, G, or T residue at the polymorphic position such that the allelic state is defined by which residue is present at the polymorphic position. For another example, the nucleic acid sequence of an RNA molecule that contains a single nucleotide polymorphism may comprise an A, C, G, or U residue at the polymorphic position such that the allelic state is defined by which residue is present at the polymorphic position. Similarly, the nucleic acid sequence of a nucleic acid molecule that contains an Indel may comprise an insertion or deletion of nucleic acid sequences at the polymorphic position such that the allelic state is defined by the presence or absence of the insertion or deletion at the polymorphic position.
[0038] As used herein, the phrase "association" when used in reference to a polymorphism and a phenotypic trait or trait index, refers to any statistically significant correlation between the presence of a given allele of a polymorphic locus and the phenotypic trait or trait index value, wherein the value may be qualitative or quantitative.
[0039] As used herein, the phrase "typing" refers to any method whereby the specific allelic form of a given soybean genomic polymorphism is determined. For example, a single nucleotide polymorphism (SNP) is typed by determining which nucleotide is present (i.e. an A, G, T, or C). Insertion/deletions (Indels) are determined by determining if the Indel is present. Indels can be typed by a variety of assays including, but not limited to, marker assays.
[0040] As used herein, the phrase "elite line" refers to any line that has resulted from breeding and selection for superior agronomic performance. An elite plant is any plant from an elite line.
[0041] As used herein, the phrase "plant" includes dicotyledons plants and monocotyledons plants. Examples of dicotyledons plants include tobacco, Arabidopsis, soybean, tomato, papaya, canola, sunflower, cotton, alfalfa, potato, grapevine, pigeon pea, pea, Brassica, chickpea, sugar beet, rapeseed, watermelon, melon, pepper, peanut, pumpkin, radish, spinach, squash, broccoli, cabbage, carrot, cauliflower, celery, Chinese cabbage, cucumber, eggplant, and lettuce.
Examples of monocotyledons plants include corn, rice, wheat, sugarcane, barley, rye, sorghum, orchids, bamboo, banana, cattails, lilies, oat, onion, millet, and triticale.
[0042] As used herein, the term "plant" also includes a whole plant and any descendant, cell, tissue, or part of a plant. The term "plant parts" include any part(s) of a plant, including, for example and without limitation: seed (including mature seed and immature seed); a plant cutting; a plant cell; a plant cell culture; a plant organ (e.g., pollen, embryos, flowers, fruits, shoots, leaves, roots, stems, and explants). A plant tissue or plant organ may be a seed, callus, or any other group of plant cells that is organized into a structural or functional unit. A plant cell or tissue culture may be capable of regenerating a plant having the physiological and morphological characteristics of the plant from which the cell or tissue was obtained, and of regenerating a plant having substantially the same genotype as the plant. In contrast, some plant cells are not capable of being regenerated to produce plants. Regenerable cells in a plant cell or tissue culture may be embryos, protoplasts, meristematic cells, callus, pollen, leaves, anthers, roots, root tips, silk, flowers, kernels, ears, cobs, husks, or stalks.
[0043] Plant parts include harvestable parts and parts useful for propagation of progeny plants. Plant parts useful for propagation include, for example and without limitation: seed; fruit; a cutting; a seedling; a tuber; and a rootstock. A harvestable part of a plant may be any useful part of a plant, including, for example and without limitation: flower; pollen; seedling; tuber; leaf; stem; fruit; seed; and root.
[0044] As used herein, the phrase "plant cell described" or "transformed plant cell" refers to a plant cell that is transformed with stably- integrated, non-natural, recombinant DNA, e.g., by Agrobacterium-mediated transformation or by bombardment using microparticles coated with recombinant DNA or other means. A plant cell of this invention can be an originally- transformed plant cell that exists as a microorganism or as a progeny plant cell that is
regenerated into differentiated tissue, e.g., into a transgenic plant with stably-integrated, non- natural recombinant DNA, or seed or pollen derived from a progeny transgenic plant.
[0045] As used herein, the phrase "trait" refers to a physiological, morphological, biochemical, or physical characteristic of a plant or particular plant material or cell. In some instances, this characteristic is visible to the human eye, including seed or plant size, or can be measured by biochemical techniques, including detecting the protein, starch, or oil content of seed or leaves, or by observation of a metabolic or physiological process (e.g., by measuring uptake of carbon dioxide), or by the observation of the expression level of a gene or genes (e.g., by employing Northern analysis, RT-PCR, microarray gene expression assays), or reporter gene expression systems, or by agricultural observations including stress tolerance, yield, or pathogen tolerance.
[0046] As used herein, a "nongenic sequence" or "nongenic genomic sequence" is a native DNA sequence found in the nuclear genome of a plant, having a length of at least 1 Kb, and devoid of any open reading frames, gene sequences, or gene regulatory sequences. Furthermore, the nongenic sequence does not comprise any intron sequence (i.e., introns are excluded from the definition of nongenic). The nongenic sequence cannot be transcribed or translated into protein. Many plant genomes contain nongenic regions, where as much as 95% of the genome can be nongenic, and these regions may be comprised of mainly repetitive DNA.
[0047] As used herein, a "genie region" is defined as a polynucleotide sequence that comprises an open reading frame encoding an RNA and/or polypeptide. The genie region may also encompass any identifiable adjacent 5' and 3' non-coding nucleotide sequences involved in the regulation of expression of the open reading frame up to about 2 Kb upstream of the coding region and 1 Kb downstream of the coding region, but possibly further upstream or downstream. A genie region further includes any introns that may be present in the genie region. Further, the genie region may comprise a single gene sequence, or multiple gene sequences interspersed with short spans (less than 1 Kb) of nongenic sequences.
[0048] While the invention has been described with reference to specific methods and embodiments, it will be appreciated that various modifications and changes may be made without departing from the invention. All publications cited herein are expressly incorporated herein by reference for the purpose of describing and disclosing compositions and methodologies that might be used in connection with the invention. All cited patents, patent applications, and sequence information in referenced websites and public databases are also incorporated by reference.
EXAMPLES
Example 1
[0049] Two sets of three backcross generations of maize plants (BC2, BC3 and BC4) are selected for evaluation in this example. The first set has relatively low genome coverage between the two parents XJA40 x 4XP811XT - similar parents where the genome coverage at BC2 is 62% at RPP4; and is 33% at RPP6).
[0050] The second set has relatively high genome coverage between the two parents LDS51 X MV8735XT - dissimilar parents where the genome coverage at BC2 is 99% at RPP4; and 85% at RPP6).
[0051] The SNP database provides reference genomes to be compared with individual sequences, and in this example generated 15,636 genome-wide markers for BC2, 14,824 genome-wide markers for BC3, and 15,402 genome-wide markers for BC4 (for the LDS51 X MV8735XT set).
[0052] The genotyping data (generated using methods known in the art) from the two sets of BC2, BC3, and BC4 generations are compared with the SNP database for the imputation of genotypic data. The flow chart for the Genotyping Imputation Module is illustrated in FIG. 1, where such algorithm is executed in a computer system. The resulting genotyping information is used by the Marker Study Manager module to generate user-friendly visualization of the output for predicting/evaluating numbers of non-rogue plants and rogue plants.
[0053] A Marker Study Manager (MSM) module is developed using algorithm with a visualization tool. The genotyping data after imputation is channeled through the MSM module to build a chromosome table with conversion metrics. Visualization of such chromosome table enables a user-friendly way to make selections, where the MSM module can generate
chromosome tables having 3500 to 15500 SNP markers in each sample in this example.
[0054] Table 1 summarizes the different populations and the population size of each generation used in this example. The output in Table 1 is solely based on analysis of molecular markers instead of observation from the phenotypes of the physical plants.
[0055] The genotyping imputation module successfully differentiates rogue plants from non- rogue plants. The rogue plants can be visualized using the Marker Study Manager module with the larger genome view. In addition, identification of rogue alleles and/or large number of SNP markers can significantly improve imputation and prediction.
[0056] The method and system used in this example can process genetic data from thousands of SNP markers which increases the genome coverage required for informed selection decisions especially in narrow crosses. The method and system provided are cost effective because there is no need to design new markers for each project where multiplexing capability is provided.

Claims

A computerized method for trait introgression, comprising,
(a) generating a genetic markers database by collecting genome-wide genetic markers information in a plant;
(b) obtaining sequences from at least one back-crossed plant sample;
(c) comparing the sequences obtained in step (b) with the genetic markers database of step (a) for imputation of genetic markers; and
(d) predicting whether the back-crossed plant sample is a rogue or non-rogue plant.
The method of claim 1, further comprising the step of selecting a pre-determined number of plants with the most advantageous genetic profiles for crossing.
The method of claim 2, wherein the pre-determined number is between 5 and 50.
The method of claim 1, wherein the genetic markers comprise single nucleotide polymorphisms (SNPs).
The method of claim 1, wherein the genetic markers database is a SNP database or SNP library.
The method of claim 5, wherein the SNP database or SNP library comprises a genome- wide SNP collection having between 100 and 500 SNPs.
The method of claim 5, wherein the SNP database or SNP library comprises a genome- wide SNP collection of at least 1,000 SNPs.
The method of claim 5, wherein the SNP database or SNP library comprises a genome- wide SNP collection having between 1,000 and 100,000 SNPs.
The method of claim 1, wherein the plant is selected from soybean, maize, canola, cotton, wheat, sunflower, and rice.
10. The method of claim 1, wherein the back-crossed plant sample is from a first generation, second generation, third generation, fourth generation, fifth generation back-crossing plant, or combinations thereof.
11. The method of claim 5, wherein the comparing step of step (c) uses a Genotyping
Imputation Module.
12. The method of claim 1, wherein the predicting step of step (d) uses a Marker Study
Manager module.
13. The method of claim 12, wherein the Marker Study Manager module provides
visualization output for predicting/evaluating numbers of non-rogue plants and rogue plants.
14. A computerized system for trait introgression, comprising,
(a) a genetic markers database;
(b) a genetic markers imputation module accepting inputs of sequences from at least one back-crossed plant sample; and
(c) a Marker Study Manager module providing visualization output.
15. The system of claim 14, wherein the Marker Study Manager module provides
visualization output for predicting/evaluating numbers of non-rogue plants and rogue plants.
16. The system of claim 14, wherein the Marker Study Manager module provides
visualization output for selecting a pre-determined number of plants with the most advantageous genetic profiles for crossing.
17. The system of claim 16, wherein the pre-determined number is between 5 and 50.
18. The system of claim 14, wherein the genetic markers database is a SNP database or SNP library.
19. The system of claim 18, wherein the SNP database or SNP library comprises a genome- wide SNP collection having between 100 and 500 SNPs.
20. The system of claim 18, wherein the SNP database or SNP library comprises a genome- wide SNP collection of at least 1,000 SNPs.
21. The system of claim 18, wherein the SNP database or SNP library comprises a genome- wide SNP collection having between 1,000 and 100,000 SNPs.
22. The system of claim 14, wherein the genetic markers imputation module is a SNP
Imputation Module.
23. The system of claim 14, wherein the plant is selected from soybean, maize, canola, cotton, wheat, sunflower, and rice.
24. The system of claim 14, wherein the back-crossed plant samples are from a first
generation, second generation, third generation, fourth generation, fifth generation back- crossing plant, or combinations thereof.
25. A process for use in a computerized system for trait introgression, comprising,
(a) inputting sequences of at least one back-crossed plant sample into the system of claim 11 by an user; and
(b) receiving output from the system of claim 14 for predicting/evaluating numbers of non-rogue plants and rogue plants.
26. A computerized method for trait introgression, comprising,
(a) generating genome-wide genetic markers information from parent plants;
(b) obtaining sequences from at least one back-crossed plant sample;
(c) comparing the sequences obtained in step (b) with the genome-wide genetic markers information from parent plants of step (a) for imputation of genetic markers; and
(d) predicting whether the back-crossed plant sample is a rogue or non-rogue plant.
27. The method of claim 26, further comprising the step of selecting a pre-determined
number of plants with the most advantageous genetic profiles for crossing.
28. The method of claim 27, wherein the pre-determined number is between 5 and 50.
29. The method of claim 26, wherein the genetic markers comprise single nucleotide
polymorphisms (SNPs).
30. The method of claim 29, wherein the genome-wide genetic markers information from parent plants comprises a genome-wide SNP collection having between 100 and 500 SNPs.
31. The method of claim 29, wherein the genome- wide genetic markers information from parent plants comprises a genome- wide SNP collection of at least 1,000 SNPs.
32. The method of claim 29, wherein the genome-wide genetic markers information from parent plants comprises a genome- wide SNP collection having between 1,000 and 100,000 SNPs.
33. The method of claim 26, wherein the plant is selected from soybean, maize, canola, cotton, wheat, sunflower, and rice. The method of claim 26, wherein the back-crossed plant sample is from a first generation, second generation, third generation, fourth generation, fifth generation back-crossing plant, or combinations thereof.
The method of claim 29, wherein the comparing step of step (c) uses a Genotyping Imputation Module.
The method of claim 26, wherein the predicting step of step (d) uses a Marker Study Manager module.
The method of claim 36, wherein the Marker Study Manager module provides
visualization output for predicting/evaluating numbers of non-rogue plants and plants.
EP16864760.0A 2015-11-10 2016-10-25 Methods and systems for trait introgression Withdrawn EP3373726A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US201562253347P 2015-11-10 2015-11-10
PCT/US2016/058575 WO2017083091A1 (en) 2015-11-10 2016-10-25 Methods and systems for trait introgression

Publications (1)

Publication Number Publication Date
EP3373726A1 true EP3373726A1 (en) 2018-09-19

Family

ID=58695924

Family Applications (1)

Application Number Title Priority Date Filing Date
EP16864760.0A Withdrawn EP3373726A1 (en) 2015-11-10 2016-10-25 Methods and systems for trait introgression

Country Status (5)

Country Link
EP (1) EP3373726A1 (en)
CN (1) CN108135144A (en)
AR (1) AR107215A1 (en)
BR (1) BR102016025738A2 (en)
WO (1) WO2017083091A1 (en)

Families Citing this family (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US7844609B2 (en) 2007-03-16 2010-11-30 Expanse Networks, Inc. Attribute combination discovery
US10777302B2 (en) * 2012-06-04 2020-09-15 23Andme, Inc. Identifying variants of interest by imputation
CN112514790B (en) * 2020-11-27 2022-04-01 上海师范大学 Rice molecular navigation breeding method and application

Also Published As

Publication number Publication date
AR107215A1 (en) 2018-04-11
CN108135144A (en) 2018-06-08
WO2017083091A1 (en) 2017-05-18
BR102016025738A2 (en) 2017-08-08

Similar Documents

Publication Publication Date Title
Kumar et al. Advances in genomic tools for plant breeding: harnessing DNA molecular markers, genomic selection, and genome editing
Tuberosa et al. Genomics-based approaches to improve drought tolerance of crops
Yano et al. Genome-wide association study using whole-genome sequencing rapidly identifies new genes influencing agronomic traits in rice
US8039686B2 (en) QTL “mapping as-you-go”
Ranc et al. Genome-wide association mapping in tomato (Solanum lycopersicum) is possible using genome admixture of Solanum lycopersicum var. cerasiforme
Kover et al. A multiparent advanced generation inter-cross to fine-map quantitative traits in Arabidopsis thaliana
Chen et al. The development of 7E chromosome-specific molecular markers for Thinopyrum elongatum based on SLAF-seq technology
Siol et al. Patterns of genetic structure and linkage disequilibrium in a large collection of pea germplasm
Li et al. Construction of high-density genetic map and mapping quantitative trait loci for growth habit-related traits of peanut (Arachis hypogaea L.)
Ogawa et al. Haplotype-based allele mining in the Japan-MAGIC rice population
BRPI0812744B1 (en) METHODS FOR SEQUENCE-TARGETED MOLECULAR IMPROVEMENT
Cook et al. Genetic analysis of stay‐green, yield, and agronomic traits in spring wheat
US20170022574A1 (en) Molecular markers associated with haploid induction in zea mays
Francis et al. Molecular characterization and SNP identification using genotyping-by-sequencing in high-yielding mutants of proso millet
CN117210596B (en) Melon SNP site marker combination, SNP site marker detection probe combination, liquid phase chip and application
Lu et al. Genetic basis of maize kernel protein content revealed by high-density bin mapping using recombinant inbred lines
Kumari et al. Association mapping reveals novel genes and genomic regions controlling grain size architecture in mini core accessions of Indian National Genebank wheat germplasm collection
Li et al. EST-SSR primer development and genetic structure analysis of Psathyrostachys juncea Nevski
WO2017083091A1 (en) Methods and systems for trait introgression
Zhang et al. Studies of new EST-SSRs derived from Gossypium barbadense
Sun et al. Evaluation of efficiency of controlled pollination based parentage analysis in a Larix gmelinii var. principis-rupprechtii Mayr. seed orchard
Shavrukov et al. Plant genotyping: from traditional markers to modern technologies
Yu et al. Development of a 50K SNP array for whole-genome analysis and its application in the genetic localization of eggplant (Solanum melongena L.) fruit shape
Wang et al. Identification of molecular markers and candidate regions associated with grain number per spike in Pubing3228 using SLAF-BSA
Park et al. Development of genome-wide single nucleotide polymorphism markers for variety identification of F1 hybrids in cucumber (Cucumis sativus L.)

Legal Events

Date Code Title Description
PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

17P Request for examination filed

Effective date: 20180328

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO RS SE SI SK SM TR

AX Request for extension of the european patent

Extension state: BA ME

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE APPLICATION HAS BEEN WITHDRAWN

18W Application withdrawn

Effective date: 20181116