EP1143787A2 - Molecular profiling for heterosis selection - Google Patents

Molecular profiling for heterosis selection

Info

Publication number
EP1143787A2
EP1143787A2 EP00904457A EP00904457A EP1143787A2 EP 1143787 A2 EP1143787 A2 EP 1143787A2 EP 00904457 A EP00904457 A EP 00904457A EP 00904457 A EP00904457 A EP 00904457A EP 1143787 A2 EP1143787 A2 EP 1143787A2
Authority
EP
European Patent Office
Prior art keywords
plant
expression
progeny
plants
dominant
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Withdrawn
Application number
EP00904457A
Other languages
German (de)
French (fr)
Inventor
Ben Bowen
Mei Guo
Oscar Smith
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Pioneer Hi Bred International Inc
Original Assignee
Pioneer Hi Bred International Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Pioneer Hi Bred International Inc filed Critical Pioneer Hi Bred International Inc
Publication of EP1143787A2 publication Critical patent/EP1143787A2/en
Withdrawn legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • C12Q1/6876Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
    • C12Q1/6888Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for detection or identification of organisms
    • C12Q1/6895Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for detection or identification of organisms for plants, fungi or algae
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • C12Q1/6809Methods for determination or identification of nucleic acids involving differential detection
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q2600/00Oligonucleotides characterized by their use
    • C12Q2600/156Polymorphic or mutational markers
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q2600/00Oligonucleotides characterized by their use
    • C12Q2600/158Expression markers

Definitions

  • Hyb ⁇ d offspring often outperform their parents by a variety of different measures, including yield, adaptability to environmental changes, disease resistance, pest resistance, and the like
  • the improved properties for the hyb ⁇ d as compared to the parents are collectively referred to as "hybrid vigor," or “heterosis”
  • Hyb ⁇ dization between parents of dissimilar genetic stock has been used in animal husbandry and especially for improving major plant crops, such as corn, sugarbeet and sunflower
  • the development of a maize hyb ⁇ d typically involves three steps (1) the selection of plants from va ⁇ ous germplasm pools for initial breeding crosses, (2) the self g of the selected plants from the breeding crosses for several generations to produce a se ⁇ es of inbred lines, which, although different from each other, breed true and are highly uniform, and (3) crossing the selected mbred lines with different mbred lines to produce hybrid progeny (sometimes referred to as "FI" hyb ⁇ ds) Du ⁇ ng the inbreeding process in maize, the vigor of the lines decreases Vigor is restored when two different inbred lines are crossed to produce hyb ⁇ d progeny A consequence of the homozygosity and homogeneity of the mbred lines is that hyb ⁇ ds produced by crossing a defined pair of mbreds are uniform and predictable.
  • heterosis is the result of one or a few general genetic mechanisms, or whether it is the result of many simultaneously interacting processes.
  • heterosis Because of the lack of understanding of the molecular basis for heterosis, crop development has relied upon empi ⁇ cal observations of heterosis for hybrids which result from crossing selected mbred crop strains (or resulting from second order crosses, e.g., in which two mbreds are crossed to produce a hyb ⁇ d which is then crossed with an inbred or hybrid strain to produce a subsequent 3-4 way heterotic hyb ⁇ d) This laborious process has been conducted on a large scale, resulting in increases in desirable measures of heterosis, such as yield, of several percent per year.
  • Molecular methods have been used to a limited extent to supplement crop breeding programs to select desirable inbreds and hyb ⁇ ds. In general, these procedures have been used to identify genetic markers corresponding to desirable or undesirable loci (e.g., "quantitative trait loci" or QTLs) m plants under analysis. Genetic markers represent (mark the location of) specific loci in the genome of a species or closely related species, and sampling of different genotypes at these marker loci reveals genetic va ⁇ ation.
  • QTLs quantitative trait loci
  • the genetic va ⁇ ation at marker loci can then be desc ⁇ bed and applied to genetic studies, commercial breeding, diagnostics, cladistic analysis of va ⁇ ance, or genotyping of samples Because molecular methods are amenable to high throughput analysis and because they do not require yield testing, they can be used to speed the process of crop development. However, although these techniques are of considerable use, and can and do enhance the efficiency of crop breeding programs, they are not currently used, or useful, as a predictor for the more general phenomenon of heterosis.
  • the present invention provides a number of fundamental discove ⁇ es which make it possible to correlate molecular methods and the phenomenon of heterosis, as well as a variety of additional aspects which will be apparent upon complete review.
  • a heterologous nucleic acid that results in expression of expression products from silenced genes is introduced into a target plant.
  • Examples of approp ⁇ ate heterologous nucleic acids include one or more of: a transc ⁇ ption factor which activates a promoter from a silenced gene, a nucleic acid encoded by the silenced gene under the control of a heterologous promoter, and a nucleic acid homologous to the silenced gene with at least one region of difference with the silenced gene, which homologous nucleic acid can recombine with the silenced gene to produce a modified gene. Any of these nucleic acids can be cloned under the control of heterologous promoters and placed into target plants to increase heterosis of the target plants.
  • integrated systems comprising computer databases having expression profile information can be used to select which parental crosses are most likely to result in an increase in the number of expression products (or an optimization of expression products of a selected class, i.e., dominant, under- dominant, over-dommant, additive, or the like) in offsp ⁇ ng
  • consideration of expression profile information provides not only a basis for selecting hybrids from crosses, but, using the methods herein, also identifies desirable crosses to be made.
  • Production and automated consideration of expression profile databases also provides a mechanism for identifying the genetic source of particular expression products, thereby indicating the likely parentage of given hyb ⁇ ds
  • the invention additionally provides methods of cloning and transducing target plants or animals with dominant, additive, under-dominant and over-dommant genes identified by comparative examination of expression profiles BRIEF DESCRIPTION OF THE FIGURES
  • Figure 1 is a scatter plot showing the correlation between the degree of heterosis and % relationship.
  • Figure 2 is a set of bar graphs showing classification of gene expression patterns in Hybrid vs. inbred parents.
  • Figure 3 is a line graph showing the co ⁇ elation between the pattern of gene expression and heterosis.
  • Figure 4 is a set of bar graphs showing dominant, additive and over-/under- dominant RNA expression.
  • Figure 5 is a scatter graph showing the correlation between parental effects on gene expression and heterosis.
  • Figure 6a-c is a set of schematic illustrations showing polymorphic dominant products and their sequences.
  • An "expression profile” is the result of detecting a representative sample of expression products from a cell, tissue or whole organism, or a representation (picture, graph, data table, database, etc.) thereof. For example, many RNA expression products or a cell or tissue can simultaneously be detected on a nucleic acid array, or by the technique of differential display or modification thereof such as Curagen's "GeneCallingTM” technology. Similarly, protein expression products can be tested by various protein detection methods, such as hybridization to peptide or antibody arrays, or by screening phage display libraries.
  • a “portion” or “subportion” of an expression profile, or a “partial profile” is a subset of the data provided by the complete profile, such as the information provided by a subset of the total number of detected expression products.
  • An “expression product” is any product transcribed in a cell from a DNA (e.g., from a gene) or translated from an RNA (e.g., a protein).
  • Example expression products include mRNAs and proteins.
  • a "representative sample" of expression products e.g., from a particular cell, tissue, or whole organism is a sufficiently large number of expression products that statistical comparison of the actual number and/or type of expression products between different cells, tissues, or whole organisms can be made. Ideally, at least about 50%, and typically 60%, 70%, 80%, 90%, 95% or 100% of the total expression products which are detectable by a given technique constitute the "representative sample.”
  • the representative sample will typically include a large number of expression products, as cells, tissues and organisms typically produce a fairly large number of expression products.
  • a typical representative sample of expression products includes between about 100 and 20,000 or more expression products, e.g., about 100-500, 1,000, 1,500, 2,000, 2,500, 3,000, 3,500, 4,000, 4,500, 5,000, 5,500, 6,000, 6,500, 7,000, 7,500, 8,000, 8,500, 9,000, 9,500, 10,000, 10,500, 11,000, 11,500, 12,000, 12,500, 13,000, 13,500, 14,000, 14,500, 15,000, 15,500, 16,000, 16,500, 17,000, 17,500, 18,000, 18,500, 19,000, 19,500, 20,000, or 30,000 expression products, or the like.
  • correlation unless indicated otherwise, is used herein to indicate that a “statistical association” exists between, e.g., an expression product and the degree of heterosis.
  • Dominant expression for an expression product refers to the situation where expression of the product in a progeny differs from one parent, and not the other for the expression product
  • Additional expression for an expression product refers to the situation where expression of the product in a progeny falls within the range of the two parents (and may or may not differ from both parents).
  • “Over-dominant” or “under-dommant” expression for an expression product refers to the situation where expression of an expression product in a progeny differs from both parents and falls outside of the range of the two parents, either over the higher parent value, or under the lower parent value, respectively ( Figure 2).
  • a “biological sample” is a portion of mate ⁇ al isolated from a biological source such as a plant, isolated plant tissue, or plant cell, or a portion of mate ⁇ al made from such a source, such as a cell extract or the like
  • a “promoter” is an array of nucleic acid control sequences which direct transc ⁇ ption of a nucleic acid
  • a promoter includes necessary nucleic acid sequences near the start site of transc ⁇ ption, such as, in the case of a polymerase II type promoter, a TATA element.
  • a promoter also optionally includes distal enhancer or repressor elements which can be located as much as several thousand base pairs from the start site of transcription.
  • a “constitutive” promoter is a promoter which is active in a selected organism under most environmental and developmental conditions.
  • An “inducible” promoter is a promoter which is under environmental or developmental regulation in a selected organism. -
  • hybrid plants refers to plants which result from a cross between genetically different individuals.
  • tester parent refers to a parent that is genetically different from a set of lines to which it is crossed. The cross is for purposes of evaluating differences among the lines in topcross combination. Using a tester parent in a sexual cross allows one of skill to determine the genetic differences between the tested lines on the phenotypic trait with expression of quantitative trait loci in a hybrid combination.
  • topcross combination and “hybrid combination” refer to the processes of crossing a single tester parent to multiple lines. The purposes of producing such crosses is to evaluate the ability of the lines to produce desirable phenotypes in hybrid progeny derived from the line by the tester cross.
  • transgenic plant refers to a plant into which exogenous polynucleotides have been introduced by any process other than sexual cross or selfing. Examples of processes by which this can be accomplished are described below, and include Agrob ⁇ cte ⁇ ' wm-mediated transformation, biolistic methods, electroporation, in planta techniques, and the like. Such a plant containing the exogenous polynucleotides is referred to here as an R l generation transgenic plant. Transgenic plants may also arise from sexual cross or by selfing of transgenic plants into which exogenous polynucleotides have been introduced. DETAILED DESCRIPTION OVERVIEW OF SELECTION FOR HETEROSIS
  • Crop improvement relies extensively on the phenomenon of heterosis.
  • Inbreds and/or hyb ⁇ ds are crossed to produce heterotic hyb ⁇ ds with desirable traits such as high yield, disease resistance, resistance to heat, cold, salinity, insects, fungi, herbicides, pesticides, etc.
  • Secondary desirable traits such as a particular size or shape of ears, solids content, sugar content, oil content, water content, etc., can also be affected by heterosis.
  • the present invention establishes several correlations between the expression of gene products and heterosis, e.g., with respect to yield.
  • genes are silenced du ⁇ ng inbreeding in plants. These correlations provide new methods of selecting heterotic hyb ⁇ ds, without the necessity of field testing every hyb ⁇ d to monitor heterotic traits.
  • expression of a first representative sample of first expression products e.g., RNAs or proteins
  • a first progeny plant e.g., a hyb ⁇ d from resulting from crossing two or more parental lines.
  • the expression products produced in the first progeny plant are quantified and/or monitored for the type of expression product (additive, dominant, under- dominant, over-dominant, etc ).
  • the number of first expression products produced in the first progeny plant is statistically associated with a measure of heterosis in the first progeny plant, as is the number of dominant, additive, under-dommant or over-dom ant, or silenced expression products.
  • the plant is then selected (e.g., against similar measures for a second progeny plant, or a population of progeny plants, or against the parental stock) for further testing based upon the number or type of expression products detected.
  • the plant can be selected for one or more characte ⁇ stic, including: a selected number of expression products, a selected number of dominant expression products, a selected ratio of dominant expression products to total expression products, a desired number of over- or under-dommant expression products, a selected ratio of over- or under- dominant expression products to total expression products, a selected number of additive expression products, and a selected ratio of additive expression products to total expression products.
  • the first progeny plant is selected to maximize the number of dominant expression products and/or to maximize the number of additive expression products, and/or to minimize the number of over- or under-dominant expression products.
  • Crosses can also be selected to minimize silencing in the progeny plant.
  • the parental plants used to produce the first progeny can also be profiled. Resulting parental expression profiles serve any of a va ⁇ ety of purposes.
  • the parental expression profiles can be compared to the first progeny profile to aid in determining whether the progeny show an increase in the number of expression products as compared to parental stocks (thereby indicating that the progeny is likely to be heterotic).
  • compa ⁇ son between the parental expression profiles and the progeny profile is used to determine whether the individual expression products represented m the profile are dominant, additive, under- dominant, over-dominant, or the like
  • the parental expression profiles can also be placed into a database to aid in determining which crosses are most likely to produce heterotic hyb ⁇ ds.
  • Cropsenchymy plants are selected by identifying plants likely to produce progeny plants with a selected number of expression products which are dominant, over-dommant, under-dommant or additive.
  • parents are selected to produce the first progeny plant by selecting for complementary expression of dominant or additive expression products between the parents, or by selecting against expression of over-dominant or under-dommant expression products in the parents
  • An additional statistical association relates to the relationship between parental and progeny plants. It is discovered that plants which exhibit an expression profile that is more similar to the maternal plant than to the paternal plant may be more heterotic. Accordingly, compa ⁇ son of the maternal, paternal and progeny expression profiles can be used to monitor this relationship. In addition, multiple crosses to a single female type can be made (or the results predicted by compa ⁇ son in a database) and the progeny screened (or predicted) for simila ⁇ ty to the female type.
  • silencing was determined to play a significant role in the loss of heterosis due to inbreeding. Accordingly, by compa ⁇ ng parental and progeny plants it is possible to determine which genes are silenced These genes can be rescued, e.g., by cloning the silenced genes and placing them under the control of heterologous promoters, or other strategies noted herein, and transducing the genes back into target plants (e.g., the parental lines, the hyb ⁇ ds, or any other plant). In addition, by compiling database information for which genes are silenced m mbreds, it is possible to decrease silencing in hyb ⁇ ds by selecting crosses where parents have complementary patterns. It is also possible to use these methods to increase the performance (e.g., gram yield, standabihty, etc.) of the inbred lines themselves.
  • the first progeny plant selected by any of the methods herein, or a subsequent progeny plant, or a transgenic plant as desc ⁇ bed above can be subjected to any of the field tests appropriate for monito ⁇ ng one or more desired traits
  • the first progeny plant, or a subsequent progeny plant thereof can be tested for a desired phenotypic trait.
  • the phenotypic trait can be compared between the first progeny plant, or a subsequent progeny plant, and a selected hybrid or inbred plant.
  • the expression profile of the selected hyb ⁇ d or mbred plant can be compared to an expression profile of the first progeny plant, or the subsequent progeny plant.
  • Nucleic acids differentially expressed between the selected hyb ⁇ d or mbred plant and the first progeny plant, or the subsequent progeny plant are identified as targets for cloning.
  • genes that are expressed high yielding hyb ⁇ ds that are not expressed in low yielding hybrids can be determined by compa ⁇ sons of the expression profiles for the high and low yielding hyb ⁇ ds
  • Nucleic acids from (or corresponding to) the differentially expressed genes are cloned for introduction into target nucleic acids After identifying which expression products from the representative sample show an additive, dominant, underdominant, or overdommant expression pattern for at least a portion of the representative sample, or a nucleic acid corresponding to the expression product, can be cloned.
  • the cloned nucleic acid can then be transduced into target plants to test whether the nucleic acid encodes a useful trait, or to improve traits in the target plant. Further details on expression profiling, cloning of nucleic acids, selection of hyb ⁇ ds, integrated systems, screening methods and the like are set forth below.
  • a va ⁇ ety of tissues can be profiled, with immature tissues being preferentially profiled. Immature tissues are prefe ⁇ ed, because it increases the rate at which crops can be screened, as a plant does not have to be grown to matu ⁇ ty However, essentially any tissue, or whole plant, can be profiled.
  • a va ⁇ ety of profiling methods are available, including hybridization of expressed or amplified nucleic acids to a nucleic acid array, hybridization of expressed polypeptides to a protein array, hybridization of peptides or nucleic acids to an antibody array, subtractive hybridization, differential display and others.
  • CROPS TO BE PROFILED The parental or progeny plants can be inbreds or hybrids.
  • the progeny plant is a hybrid, produced by crossing two different inbred lines, or crossing an inbred line and a hybrid line, or crossing two hybrid lines (which are the result of crossing inbred or hybrid lines), or crossing of more than two lines (e.g., to generate polyploid or recombinant plants) in a single cross.
  • a desirable heterotic hybrid Once a desirable heterotic hybrid is identified, it can be treated as such hybrids typically are in breeding schemes, e.g., it can produced in quantity as seed; it can be top crossed to inbred lines to produce a 3-way hybrid plant; it can be selfed to produce more inbred lines, or the like.
  • Monocots such as plants in the grass family (Gramineae), such as plants in the sub families Fetucoideae and Poacoideae, which together include several hundred genera including plants in the genera
  • Agrostis Phleum, Dactylis, Sorgum, Setaria, Zea (e.g., corn), Oryza (e.g., rice), Triticum (e.g., wheat), Secale (e.g., rye), Avena (e.g., oats), Hordeum (e.g., barley), Saccharum, Poa, Festuca, Stenotaphrum, Cynodon, Coix, the Olyreae, Phareae and many others. Plants in the family Gramineae are a particularly preferred target plants for the methods of the invention.
  • Additional preferred targets include other commercially important crops, e.g., from the families Compositae (the largest family of vascular plants, including at least 1,000 genera, including important commercial crops such as sunflower), and Leguminosae or "pea family,” which includes several hundred genera, including many commercially valuable crops such as pea, beans, lentil, peanut, yam bean, cowpeas, velvet beans, soybean, clover, alfalfa, lupine, vetch, lotus, sweet clover, wisteria, and sweetpea.
  • Common crops applicable to the methods of the invention include Zea mays, rice, soybean, sorghum, wheat, oats, barley, millet, sunflower, and canola.
  • RNA PROFILING In one preferred embodiment, the expression products which are detected in the methods of the invention are RNAs, e.g., mRNAs expressed from genes within a cell of the plant or tissue profiled.
  • RNA detection A number of techniques are available for detecting RNAs. For example, northern blot hybridization is widely used for RNA detection, and is generally taught in a variety of standard texts on molecular biology, including: Berger and Kimmel, Guide to
  • RNA can be converted into a double stranded DNA using a reverse transcriptase enzyme and a polymerase. See, Ausubel, Sambrook and Berger, id.
  • detection of mRNAs can be performed by converting, e.g., mRNAs into DNAs, which are subsequently detected in, e.g., a standard "Southern blot" format.
  • DNAs can be amplified to aid in the detection of rare molecules by any of a number of well known techniques, including: the polymerase chain reaction (PCR), the ligase chain reaction (LCR), Q ⁇ -rephcase amplification and other RNA polymerase mediated techniques (e g., NASBA) Examples of these techniques are found m Berger, Sambrook, and Ausubel, id., as well as in Mulhs et al, (1987) U.S. Patent No
  • PCR amphcons of up to 40kb are generated.
  • RNA can be converted into a double stranded DNA suitable for rest ⁇ ction digestion, PCR expansion and sequencing using reverse transcnptase and a polymerase. See, Ausubel, Sambrook and Berger, all supra These general methods can be used for expression profiling. For example, arrays of probes can be spotted onto a surface and expression products (or in vitro amplified nucleic acids corresponding to expression products) can be labeled and hyb ⁇ dized with the array For convenience, it may be helpful to use several arrays simultaneously. It is expected that one of skill is familiar with nucleic acid hyb ⁇ dization. General methods of hyb ⁇ dization are found in Berger, Sambrook and Ausubel, supra, and further in Tijssen (1993) Laboratory
  • solid phase arrays are adapted for the rapid and specific detection of multiple polymorphic nucleotides.
  • a nucleic acid probe is chemically linked to a solid support and a target nucleic acid (e.g., an RNA or corresponding amplified DNA) is hybridized to the probe.
  • a target nucleic acid e.g., an RNA or corresponding amplified DNA
  • hybridization is detected by detecting bound fluorescence.
  • hybridization is typically detected by quenching of the label by the bound nucleic acid.
  • detection of hybridization is typically performed by monitoring a - signal shift such as a change in color, fluorescent quenching, or the like, resulting from proximity of the two bound labels.
  • an array of probes are synthesized on a solid support.
  • chip masking technologies and photoprotective chemistry it is possible to generate ordered arrays of nucleic acid probes with large numbers of probes.
  • These arrays which are known, e.g., as "DNA chips,” or as very large scale immobilized polymer arrays (“VLSIPS”TM arrays) can include millions of defined probe regions on a substrate having an area of about 1cm 2 to several cm 2 .
  • arrays of chemicals, nucleic acids, proteins or the like can also be printed on a solid substrate using printing technologies.
  • these procedures provide a method of producing 4 n different oligonucleotide probes on an array using only 4n synthetic steps.
  • Light-directed combinatorial synthesis of oligonucleotide arrays on a glass surface is performed with automated phosphoramidite chemistry and chip masking techniques similar to photo resist technologies in the computer chip industry.
  • a glass surface is derivatized with a silane reagent containing a functional group, e.g., a hydroxyl (for nucleic acid arrays) or amine group (for peptide or peptide nucleic acid arrays) blocked by a photolabile protecting group.
  • Photolysis through a photolithogaphic mask is used selectively to expose functional groups which are then ready to react with incoming 5'-photoprotected nucleoside phosphoramidites.
  • the phosphoramidites react only with those sites which are illuminated (and thus exposed by removal of the photolabile blocking group).
  • the phosphoramidites only add to those areas selectively exposed from the preceding step. These steps are repeated until the desired array of sequences have been synthesized on the solid surface.
  • Combinatorial synthesis of different oligonucleotide analogues at different locations- on the array is determined by the pattern of illumination during synthesis and the order of addition of coupling reagents.
  • Monitoring of hybridization of target nucleic acids to the array is typically performed with fluorescence microscopes or laser scanning microscopes.
  • one of skill is also able to order custom-made arrays and array-reading devices from manufacturers specializing in a ⁇ ay manufacture. For example, Affymetrix Corp. in Santa Clara, CA manufactures nucleic acid arrays.
  • probe design is influenced by the intended application. For example, where several allele-specific probe-target interactions are to be detected in a single assay, e.g., on a single nucleic acid chip, it is desirable to have similar melting temperatures for all of the probes. Accordingly, the length of the probes are adjusted so that the melting temperatures for all of the probes on the array are closely similar (it will be appreciated that different lengths for different probes may be needed to achieve a particular T m where different probes have different GC contents). Although melting temperature is a primary consideration in probe design, other factors are also optionally used to further adjust probe construction, such as elimination of self-complementarity in the probe (which can inhibit hybridization of a target nucleotide).
  • a restriction site or amplification template for a second primer is incorporated, the primers are optionally longer than those described above by the length of the restriction site, or amplification template site.
  • Standard restriction enzyme sites include 4 base sites, 5 base sites, 6 base sites, 7 base sites, and 8 base sites.
  • An amplification template site for a second primer can be of essentially any length, for example, the site can be about 15-25 nucleotides in length.
  • the amplified products are optionally labeled and are typically resolved by electrophoresis on a polyacrylamide gel; the location(s) where label is present are excised and the labeled product species is/are recovered from the gel portion, typically by elution.
  • the resultant recovered product species can be subcloned into a replicable vector with or without attachment of linkers, amplified further, and/or detected, or even sequenced directly. Sequencing methods are described in Berger, Sambrook and Ausubel, supra.
  • differential display for expression profiling.
  • CuraGen Corp. New Haven CT
  • detected proteins can be derived from one of at least two sources.
  • the proteins which are detected can be either directly isolated from a cell or tissue to be profiled, providing direct detection (and, optionally, quantification) of proteins present in a cell.
  • mRNAs can be translated into cDNA sequences, cloned and expressed. This increases the ability to detect rare RNAs, and makes it possible to immediately associate a detected protein with its coding sequence.
  • nucleic acids it is not necessary even to express nucleic acids in the proper reading frame, as it is typically the presence or absence of an expression product that is, initially, at issue. Even an out of frame peptide is an indicator for the presence of a corresponding RNA.
  • hybridization techniques including western blotting, ELISA assays, and the like are available for detection of specific proteins. See, Ausubel, Sambrook and Berger, supra. See also, Antibodies: A Laboratory Manual, (1988) E. Harlow and D. Lane, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY. Non-hybridization based techniques such as two-dimensional electrophoresis can also be used to simultaneously and specifically detect large numbers of proteins.
  • One typical technology for detecting specific proteins involves making antibodies to the proteins. By specifically detecting binding of an antibody and a given protein, the presence of the protein can be detected.
  • one of skill can easily make antibodies using existing techniques, or modify those antibodies which are commercially or publicly available.
  • general methods of producing polyclonal and monoclonal antibodies are known to those of skill in the art. See, e.g., Paul (ed) (1998) Fundamental Immunology, Fourth Edition Raven Press, Ltd., New York Coligan (1991) Current Protocols in Immunology Wiley/Greene, NY; Harlow and Lane (1989) Antibodies: A Laboratory Manual Cold Spring Harbor Press, NY; Stites et al. (eds.) Basic and Clinical Immunology (4th ed.) Lange Medical Publications, Los-
  • Specific monoclonal and polyclonal antibodies and antisera will usually bind with a K D of at least about .1 ⁇ M, preferably at least about .01 ⁇ M or better, and most typically and preferably, .001 ⁇ M or better.
  • an “antibody” refers to a protein consisting of one or more polypeptide substantially or partially encoded by immunoglobulin genes or fragments of immunoglobulin genes.
  • the recognized immunoglobulin genes include the kappa, lambda, alpha, gamma, delta, epsilon and mu constant region genes, as well as myriad immunoglobulin variable region genes.
  • Light chains are classified as either kappa or lambda.
  • Heavy chains are classified as gamma, mu, alpha, delta, or epsilon, which in turn define the immunoglobulin classes, IgG, IgM, IgA, IgD and IgE, respectively.
  • a typical immunoglobulin (antibody) structural unit is known to comprise a tetramer.
  • Each tetramer is composed of two identical pairs of polypeptide chains, each pair having one "light” (about 25 kD) and one "heavy” chain (about 50-70 kD).
  • the N-terminus of each chain defines a variable region of about 100 to 110 or more amino acids primarily responsible for antigen recognition.
  • the terms variable light chain (V L ) and variable heavy chain (V H ) refer to these light and heavy chains respectively.
  • Antibodies exist as intact immunoglobulins or as a number of well characterized fragments produced by digestion with various peptidases.
  • pepsin digests an antibody below the disulfide linkages in the hinge region to produce F(ab)' 2 a dimer of Fab which itself is a light chain joined to V H -C H 1 by a disulfide bond.
  • the F(ab)' 2 may be reduced under mild conditions to break the disulfide linkage in the hinge region thereby converting the (Fab') 2 dimer into an Fab' monomer.
  • the Fab' monomer is essentially an Fab with part of the hinge region (see, Fundamental Immunology. W.E. Paul, ed., Raven Press, N.Y. (1993), for a more detailed description of other antibody fragments).
  • Antibodies include single chain antibodies, including single chain Fv (sFv) antibodies in which a variable heavy and a variable light chain are joined together (directly or through a peptide linker) to form a continuous polypeptide.
  • sFv single chain Fv
  • antibodies or antibody fragments can be arrayed, e.g., by coupling to an amine moiety fixed to a solid phase array, in a manner similar to that described above for construction of nucleic acid arrays.
  • the antibodies can be labeled, or proteins corresponding to expression products can be labeled. In this manner, it is possible to couple hundreds, or even thousands, of different antibodies to an array.
  • a bacteriophage antibody display library is screened with a polypeptide encoded by a cell, or obtained by expression of mRNAs, differential display, subtractive hybridization or the like.
  • Combinatorial libraries of antibodies have been generated in bacteriophage lambda expression systems which are screened as bacteriophage plaques or as colonies of lysogens (Huse et al. (1989) Science 246: 1275; Caton and Koprowski (1990) Proc. Natl. Acad. Sci. (U.S.A.) 87:6450; Mullinax et al (1990) Proc. Natl. Acad. Sci. (U.S.A.) 87:8095; Persson et al.
  • the patterns of hybridization which are detected provide an indication of the presence or absence of protein sequences. As long as the library or array against which a population of proteins are to be screened can be correlated from one experiment to the next -
  • peptide and nucleic acid hybridization to arrays or libraries can be treated in a manner analogous to a bar code label.
  • Any diverse library or array can be used to screen for the presence or absence of complementary molecules, whether RNA, DNA, protein, or a combination thereof.
  • mass spectrometry is in use for identification of large sets of proteins in samples, and is suitable for identification of many proteins in a sequential or parallel fashion.
  • Hutchens et al. U.S. Pat. 5,719,060 describe methods and apparatus for desorption and ionization of analytes for subsequent analysis by mass spectroscopy and/or biosensors.
  • Sample presenting means with probe elements with "Surfaces Enhanced for Laser Desorption/Ionization" (SELDI) described in the '060 patent is particularly useful in the context of the present invention; however, other approaches described in the '060 are also generally applicable to the present invention.
  • Multi-dimensional gel technology is well-known and described e.g., in Ausubel, supra, Volume 2, Chapter 10.
  • Image analysis of multi-dimensional protein separation gels provides an indication of the proteins that are expressed e.g., in a cell or tissue type. It is worth noting that identification of particular proteins is not necessary; instead, positional and pattern information e.g , of protein staining or fluorescmg patterns is sufficient to identify sets of protein expression products.
  • metabolites can be monitored by any of currently available method, including chromatography, urn or multi dimensional gel separations, hyb ⁇ dization to complementary molecules, or the like
  • the invention provides methods of identifying plant crosses with an increase in probability for heterosis progeny plants. For example, in a preferred method, the expression profiles for a plurality of plants are compared, and the expression profiles are considered by pair-wise comparison. Desirable crosses produce progeny with a selected or optimal number of expression products, or progeny with a selected number or type of expression products that display a dominant, additive, over-dommant or under-dommant expression pattern. Desirably, these compa ⁇ sons are performed in an integrated system which includes a computer The generation and use of databases of expression profile information for performing a va ⁇ ety of comparisons is a feature of the invention.
  • a va ⁇ ety of comparative methods can be performed in an integrated system, e.g., to determine the heterosis (or likely heterosis) of a cross.
  • one simple measure that can be compared across different actual or potential crosses to determine the desirability of a particular cross is to determine the sum of the expressed gene products that differ from a progeny plant in each of a first and second parental plant and the number of expressed gene products that differ between the first and second parental plant. The larger this sum, typically, the more desirable the cross.
  • matrices of possible expression profile combinations for plants are generated.
  • the expression profiles are compiled in a database in a computer and a matrix of possible pair-wise expression profile combinations for the plants is generated and queried using an integrated system comprising a computer with software for generating and comparing matrices.
  • Subsets of potential crosses from all of the possible pair-wise comparisons which exhibit a maximal number of expression profile differences represent one preferred cross.
  • Useful software aids in determining how many genes are expressed, or whether expressed genes are additive, dominant, over-dominant or under-dominant.
  • matrix information can be limited to possible pair-wise crosses for plants from different heterotic groups, or from the same heterotic group.
  • the fidelity of predicted expression profile information increasingly varies as subsequent cross information is considered, and of course, the number of possible crosses increases. Accordingly, typically only one or a few rounds of potential crosses are considered at one time. In any case, selection of a subset of potential crosses from all of the possible pair-wise comparisons which exhibit a maximal number of expression profile differences is desirable. A variety of rules for performing the basic comparisons can be used.
  • crosses are identified in which the sum of: (i) expression products produced in a first plant from a first heterotic group (A,) which are not expressed in a second plant from the first heterotic group (A) to which the first plant is crossed (A.), and which are not expressed in a selected third plant from a second heterotic group (B), plus (ii) the expression products produced in A. which are not produced A, and which are not produced in B, is optimized.
  • This optimization results in crosses which achieve elevated numbers of expression products expressed in heterotic hybrid progeny, and also in an optimization of the number of dominant products expressed.
  • optimization is achieved by determining all possible pair-wise combinations from the first heterotic group and identifying the cross which results in the largest sum of expression products, or by determining all possible pair-wise combinations from the first heterotic group and identifying crosses which result in a hybrid progeny (A, x A.) with a maximal number of differences as compared to B, or by determining all possible pair-wise combinations from the first heterotic group and identifying crosses which result in the hybrid progeny (A, x A..) having a greater number of differences with B than the number of differences between B and A, or B and A..
  • this optimization results in crosses which achieve elevated numbers of expression products expressed in heterotic hybrid progeny, and also in an optimization of the number of dominant products expressed.
  • Such implementations can also be used to improve selection methods per se. For example, in one method, self- or back-crossed progeny derived from the A, x A hybrid are selected which either retain a set of expression products defined by the sum of expression products expressed in A, (but not A. or B) and A. (but not A, or B), or which show a larger number of expression products expressed in a topcross with B than does either A, or A. when topcrossed with B.
  • One approach for comparing profiles is a nested analysis in which expression profiles are successively grouped together, and the many gene expression differences seen in individual pair-wise compa ⁇ sons can be ranked hierarchically in a filtering process. This method is useful for identifying genes expressed in one set of genotypes vs. another, e.g. hybrids vs. inbreds or bulked segregants from the two ends of a quantitative phenotypic distribution.
  • the methods of the invention can include inputing an expression profile for progeny or parental plants into a database of expression profiles. This can be performed manually, but is more typically performed in an automated system.
  • Computer databases of expression profile information can be quite large, with from a few up to several thousand profiles in the database.
  • the database will have expression product profiles of a representative sample of expression products for hybrid progeny plants resulting from at least 10 separate inbred plant crosses, or at least 10 inbred plant expression product profiles.
  • computer system or "integrated system” in the context of this invention refers to a system in which data entering a computer corresponds to physical objects or processes external to the computer, e.g., nucleic acid hybridization or protein binding data and a process that, within a computer, causes a physical transformation of the input signals to different output signals.
  • the input data e.g., hybridization of expression products on a specific array
  • output data e.g., the identification or counting of the sequence hybridized, comparison to similar a ⁇ ays with different test materials, counting and categorization of expression products or the like.
  • the process within the computer is a program by which positive (or negative) hybridization signals are recognized by the computer system and attributed to a region of an array, or other expression profile format (e.g., simple counting of array signals).
  • the program determines which region of the array the hybridized expression products are located on and, optionally, the specific corresponding sequences which the probe is based on (as noted above, no sequence information is required for making or assessing expression profiles).
  • the invention provides integrated systems for plant or plant cell manipulation and hybridization analysis. Typical systems include a digital computer with high-throughput liquid control software, image analysis software, and data interpretation software.
  • a robotic liquid control armature for transferring solutions (e.g., plant cell extracts) from a source to a destination, is typically operably linked to the digital computer.
  • An input device for entering data to the digital computer to control high throughput liquid transfer by the robotic liquid control armature and, optionally, to control transfer by the pinning armature to the solid support is commonly a feature of the integrated system, as is an image scanner for digitizing label signals from labeled probe hyb ⁇ dized to the DNA on the solid support operably linked to the digital computer
  • the image scanner interfaces with the image analysis software to provide a measurement of probe label intensity, where the probe label intensity measurement is interpreted by the data interpretation software to show whether, and to what degree, the labeled probe hybridizes to a label.
  • High throughput screening systems are commercially available (see, e.g., Zymark Corp., Hopkmton, MA; Air Technical Indust ⁇ es, Mentor, OH; Beckman Instruments, Inc Fullerton, CA, Precision Systems, Inc., Natick, MA, etc.). These systems typically automate entire procedures including all sample and reagent pipetting, liquid dispensing, timed incubations, and final readings of the microplate in detector(s) approp ⁇ ate for the assay. These configurable systems provide high throughput and rapid start up as well as a high degree of flexibility and customization For example, the currently available commercial software package, BioWorks® 1 4®, provided by Beckman Instruments, Inc.
  • Optical images viewed (and, optionally, recorded) by a camera or other recording device are optionally further processed- m any of the embodiments herein, e.g., by digitizing the image and/or sto ⁇ ng and analyzing the image on a computer.
  • a va ⁇ ety of commercially available pe ⁇ pheral equipment and software is available for digitizing, sto ⁇ ng and analyzing a digitized video or digitized optical image, e.g., using PC (Intel x86 or pentium chip- compatible DOSTM, OS2TM WINDOWSTM, WINDOWS NTTM or WINDOWS95TM based machines), MACINTOSHTM, or UNIX based
  • a CCD camera includes an array of picture elements (pixels). The light from the specimen is imaged on the CCD. Particular pixels corresponding to regions of the specimen (e g., individual hyb ⁇ dization sites on an array of biological polymers) are sampled to obtain light intensity readings for each position Multiple pixels are processed in parallel to increase speed.
  • Integrated systems for hybridization analysis of the present invention typically include a digital computer with high-throughput liquid control software, image analysis software, data interpretation software, a robotic liquid control armature for transfer ⁇ ng solutions from a source to a destination operably linked to the digital computer, an input device (e.g., a computer keyboard) for ente ⁇ ng data to the digital computer to control high throughput liquid transfer by the robotic liquid control armature and, optionally, an image scanner for digitizing label signals from labeled probe hybridized to expression products, e.g , on a solid support operably linked to the digital computer
  • the image scanner interfaces with the image analysis software to provide a measurement of probe label intensity Typically, the probe label intensity measurement is interpreted by the data interpretation software to show whether the labeled probe hybridizes to the DNA on the solid support.
  • Software to support sample processing can be divided into 4 functional catego ⁇ es: 1) liquid transfer control software, 2) image analysis software
  • applications can share information through data files which the applications can read and create.
  • files can be formatted as simple text files and/or in Microsoft Excel® or other worksheet format. This allows viewing and editing of the files through the use of commercially available software such as Microsoft Excel®.
  • Microsoft Windows® a Microsoft Windows® user interface can be developed for most applications using Microsoft Visual Basic 4.O®. Most applications can be developed for a 32-bit environment to run under Microsoft Windows 95® or 98®. 16-bit applications such as image analysis software developed by Optimas Corporation, Optimas 5.0, can also be useful components of the integrated system.
  • nucleic acid encoding an expression product identified as being of interest by the expression profiling techniques noted herein can be cloned. It is expected that many such nucleic acids, particularly dominant and additive nucleic acids will be encoded by loci responsible for desirable quantitative traits ("QTL” see, Edwards, et al., (1987) in Genetics 115:113). QTL include genes that control, to some degree, nume ⁇ cally quantifiable phenotypic traits such as disease resistance, crop yield, resistance to environmental extremes, etc. In addition to the methods herein, other expe ⁇ mental paradigms can be used to identify, analyze and select for QTL.
  • One paradigm involves crossing two mbred lines and genotyping multiple marker loci and evaluating one to several quantitative phenotypic traits among the progeny of the cross. QTL are then identified and ultimately selected for based on significant statistical associations between the genotypic values determined by genetic marker technology and the phenotypic va ⁇ ability among the segregating progeny. As applied to the present invention, the identification of particular nucleic acids which encode dominant, additive or under or over dominant expression products, or which encode silenced expression products, are potential products of QTLs or other genes or loci of interest.
  • nucleic acids which are genetically linked to DNAs encoding these expression products for transduction into cells (e.g., coding sequences for expression products, or genetically linked coding or non-coding sequences), especially to make transgenic plants.
  • the cloned sequences are also useful as molecular tags- for selected plant strains, e.g., to identify parentage, and are further useful for encoding expression products, including nucleic acids and polypeptides.
  • expression products which are differentially expressed between heterotic and non-heterotic plants are encoded by QTL and are responsible for the phenotypic effects of the QTL.
  • a DNA linked to a locus encoding an expression product is introduced into plant cells, either in culture or in organs of a plant, e.g., leaves, stems, fruit, seed, etc.
  • the expression of natural or synthetic nucleic acids encoded by nucleic acids linked to expression product coding nucleic acids can be achieved by operably linking a cloned nucleic acid of interest, such as an expression product or a genetically linked nucleic acid, to a promoter, incorporating the construct into an expression vector and introducing the vector into a suitable host cell.
  • an endogenous promoter linked to the nucleic acids can be used.
  • Bacterial cells are often used to amplify increase the number of plasmids containing DNA constructs of this invention.
  • the bacteria are grown to log phase and the plasmids within the bacteria can be isolated by a variety of methods known in the art (see, for instance, Sambrook).
  • kits are commercially available for the purification of plasmids from bacteria.
  • Agrobacterium tumefaciens related vectors to infect plants contain transc ⁇ ption and translation terminators, transc ⁇ ption and translation initiation sequences, and promoters useful for regulation of the expression of the particular nucleic acid.
  • the vectors optionally comp ⁇ se gene ⁇ c expression cassettes containing at least one independent terminator sequence, sequences permitting replication of the cassette in eukaryotes, or prokaryotes, or both, (e.g., shuttle vectors) and selection markers for both prokaryotic and eukaryotic systems
  • Vectors are suitable for replication and integration in prokaryotes, eukaryotes, or preferably both.
  • the nucleic acid constructs of the invention are introduced into plant cells, either m culture or in the organs of a plant by a va ⁇ ety of conventional techniques.
  • the DNA construct can be introduced directly into the genomic DNA of the plant cell using techniques such as electroporation and microinjection of plant cell protoplasts, or the DNA constructs can be introduced directly to plant cells using ballistic methods, such as DNA particle bombardment.
  • the DNA constructs are combined with suitable T-DNA flanking regions and introduced into a conventional Agrobacterium tumefaciens host vector.
  • the virulence functions of the Agrobacterium tumefaciens host directs the insertion of the construct and adjacent marker into the plant cell DNA when the cell is infected by the- bacteria.
  • Microinjection techniques are known in the art and well described in the scientific and patent literature.
  • the introduction of DNA constructs using polyethylene glycol precipitation is described in Paszkowski, et al, EMBO J. 3:2717 (1984).
  • Electroporation techniques are described in Fromm, et al, Proc. Nat'l. Acad. Sci. USA
  • Agrobacterium tumefaciens-medi ⁇ ed transformation techniques including disarming and use of binary vectors, are also well described in the scientific literature. See, for example Horsch, et al, Science 233:496-498 (1984), and Fraley, et al., Proc. Nat'l. Acad.
  • Agrobacterium-mediated transformation is a prefe ⁇ ed method of transformation of dicots.
  • recombinant DNA vectors suitable for transformation of plant cells are prepared.
  • a DNA sequence coding for the desired mRNA, polypeptide, or non-expressed sequence is transduced into the plant.
  • the sequence is optionally combined with transcriptional and translational initiation regulatory sequences which will direct the transcription of the sequence from the gene in the intended tissues of the transformed plant.
  • Promoters in nucleic acids linked to loci identified by detecting expression products, are identified, e.g., by analyzing the 5' sequences upstream of a coding sequence in linkage disequilibrium with the loci.
  • promoters will be associated with a QTL. Sequences characteristic of promoter sequences can be used to identify the promoter. Sequences controlling eukaryotic gene expression have been extensively studied. For instance, promoter sequence elements include the TATA box consensus sequence
  • TATAAT which is usually 20 to 30 base pairs upstream of a transcription start site. In most instances the TATA box aids in accurate transcription initiation. In plants, further upstream from the TATA box, at positions -80 to -100, there is typically a promoter element with a series of adenines su ⁇ ounding the trinucleotide G (or T) N G. See, e.g., J. Messing, et al, in Genetic Engineering in Plants, pp. 221-227 (Kosage, Meredith and Hollaender, eds. (1983)). A number of methods are known to those of skill in the art for identifying and characterizing promoter regions in plant genomic DNA.
  • a plant promoter fragment is optionally employed which directs expression of a nucleic acid in any or all tissues of a regenerated plant.
  • constitutive promoters include the cauliflower mosaic virus (CaMV) 35S transcription initiation region, the 1'- or 2'- promoter derived from T-DNA of Agrobacterium tumafaciens, and other transcription initiation regions from various plant genes known to those of skill.
  • the plant promoter may direct expression of the polynucleotide of the invention in a specific tissue (tissue-specific promoters) or may be otherwise under more precise environmental control (inducible promoters).
  • tissue-specific promoters under developmental control include promoters that initiate transcription only in certain tissues, such as fruit, seeds, or flowers. Any of a number of promoters which direct transcription in plant cells can be suitable.
  • the promoter can be either constitutive or inducible.
  • promoters of bacterial origin which operate in plants include the octopine synthase promoter, the nopaline synthase promoter and other promoters derived from native Ti plasmids. See, Herrara-Estrella et al. (1983), Nature. 303:209-213. Viral promoters include the 35S and 19S RNA promoters of cauliflower mosaic virus. See, Odell et al
  • plant promoters include the ribulose-l,3-bisphosphate carboxylase small subunit promoter and the phaseolin promoter.
  • the promoter sequence from the E8 gene and other genes may also be used. The isolation and sequence of the E8 promoter is described in detail in Deikman and Fischer, (1988) EMBO J. 7:3315- 3327.
  • a polyadenylation region at the 3'-end of the coding region is typically included. The polyadenylation region can be de ⁇ ved from the natural gene, from a va ⁇ ety of other plant genes, or from T-DNA.
  • the vector compnsing the sequences from genes encoding expression products of the invention will typically comp ⁇ se a nucleic acid - subsequence which confers a selectable phenotype on plant cells.
  • the vector comp ⁇ smg the sequence will typically comprise a marker gene which confers a selectable phenotype on plant cells.
  • the marker may encode biocide tolerance, particularly antibiotic tolerance, such as tolerance to kanamycin, G418, bleomycm, hygromycin, or herbicide tolerance, such as tolerance to chlorosluforon, or phosph oth ⁇ cm (the active ingredient in the herbicides bialaphos and Basta).
  • crop selectivity to specific herbicides can be conferred by engmee ⁇ ng genes into crops which encode approp ⁇ ate herbicide metabolizing enzymes from other organisms, such as microbes.
  • crops which encode approp ⁇ ate herbicide metabolizing enzymes from other organisms, such as microbes.
  • Padgette et al. (1996) "New weed control opportunities: Development of soybeans with a Round UP ReadyTM gene” In: Herbicide-Resistant Crops (Duke, ed.), pp 53-84, CRC Lewis Publishers, Boca
  • genes that confer tolerance to herbicides include: a gene encoding a chime ⁇ c protein of rat cytochrome P4507A1 and yeast NADPH- cytochrome P450 oxidoreductase (Shiota, et al. (1994) Plant Physiol. 106(1)17, genes for glutathione reductase and superoxide dismutase (Aono, et al. (1995) Plant Cell Physiol.
  • nucleic acids which can be cloned and introduced into plants to modify or complement expression of a gene, including a silenced gene, a dominant gene, and additive gene or the like, can be any of a variety of constructs, depending on the particular application.
  • a nucleic acid encoding a cDNA expressed from an identified gene can be expressed in a plant under the control of a heterologous promoter.
  • a nucleic acid - encoding a transc ⁇ ption factor that regulates a target identified by the methods herein, or that encodes any other moiety affecting transc ⁇ ption can be cloned and transduced into a plant Methods of identifying such factors are replete throughout the literature. For a basic introduction to genetic regulation, see, Lewin (1995) Genes V Oxford University Press Inc , NY (Lewm), and the references cited therein.
  • Transformed plant cells which are de ⁇ ved by any of the above transformation techniques can be cultured to regenerate a whole plant which possesses the transformed genotype and thus the desired phenotype.
  • Such regeneration techniques rely on manipulation of certain phytohormones in a tissue culture growth medium, typically relying on a biocide and/or herbicide marker which has been introduced together with the desired nucleotide sequences.
  • Plant regeneration from cultured protoplasts is desc ⁇ bed in Evans, et al., Protoplasts Isolation and Culture, Handbook of Plant Cell Culture, pp 124-176, Macmilhan Publishing Company, New York, (1983), and Binding, Regeneration of Plants. Plant Protoplasts, pp. 21-73, CRC Press, Boca Raton, (1985).
  • Regeneration can also be obtained from plant callus, explants, somatic embryos (Dandekar, et al., J. Tissue Cult. Meth. 12: 145 (1989); McGranahan, et al., Plant Cell Rep 8:512 (1990)), organs, or parts thereof.
  • Such regeneration techniques are desc ⁇ bed generally in Klee, et al, Ann. Rev, of Plant Phvs 38:467-486 (1987).
  • One of skill will recognize that after the expression cassette is stably incorporated in transgenic plants and confirmed to be operable, it can be introduced into other plants by sexual crossing. Any of a number of standard breeding techniques can be used, depending upon the species to be crossed.
  • GENE SILENCING AND HETEROSIS It is discovered that gene silencing and epigenetic effects play a role in inbreeding depression. As demonstrated herein, the number of genes in hyb ⁇ ds with a dominant pattern of gene expression is correlated with hyb ⁇ d yield, a component of which is found to be relief from inbreeding depression. An other way of conside ⁇ ng genes in this class is to classify them as genes that are expressed at lower levels in one inbred parent than the other.
  • the number of genes in the dominant class were considered as a function of the number of hyb ⁇ ds that share those genes, and the frequency dist ⁇ bution indicated that the overlap between sets of genes cont ⁇ buting to dominant patterns of gene expression in hyb ⁇ ds is essentially random. This suggests that, du ⁇ ng the process of inbreeding, expression of a subset of genes may always be altered (and usually reduced), and that the expression of different random subsets of genes are silenced in different mbreds.
  • allelic (and non-allelic) effects have been described where expression in heterozygotes is normal, but in homozygotes trans-inactivation (or silencing) of both alleles occurs.
  • These effects are mediated by cis-acting regulatory sequences that need to be present at more than one copy (e.g. on different chromosome homologs) to mediate the cooperative assembly of multimeric protein complexes responsible for gene silencing (e.g., Polycomb proteins in Drosophila or SIR proteins in yeast).
  • sequences responsible for these effects most likely occur in intergenic regions outside of the chromatin loops flanked by MARs that contain genes.
  • the present invention provides methods of identifying unique expression products and/or unique profiles (or partial profiles). This ability to identify unique expression products provides one way of ascertaining parentage, which, in turn, provides the ability to determine whether a hybrid comprises proprietary material.
  • a source or the sources of a test plant such as a hybrid can be identified.
  • a representative sample of expression products from the test plant is profiled and the resulting test expression profile is compared to a database of known expression profiles for plants from known inbred or hybrid strains (methods of making such databases are described above).
  • the expression profiles for a selected tissue can be entered into a database for any or every proprietary plant (or clone, or any other source of germ plasm) that a corporation owns.
  • profiling a number of plants it is possible to detect unique expression products and/or expression patterns within the expression profile of specific plants. It is also possible to generate likely expression profiles for hybrid products of members of the database. Any of these expression profiles can be compared to an actual expression profile for a test plant suspected of being derived from a one or more proprietary plant. For example, a matrix of pair-wise comparisons for potential progeny from the expression profiles in the database can be compared to the test expression profile. Either the entire expression profile or a sub portion of the expression profile (i.e., a plurality of characters corresponding to expression products found in the overall profile) comprising at least one unique expression marker can be evaluated.
  • EXAMPLE 1 DIFFERENCES IN RNA EXPRESSION PROFILES CORRELATE WITH HETEROSIS Heterosis is a term used to describe the increased vigor of hybrid progeny in - comparison to their parents. Although heterosis has been widely used in plant breeding for many decades, the molecular mechanisms underlying the phenomenon were previously unknown. In this example, heterosis was studied as a phenotype using CuraGen (CuraGen Corp., New Haven CT) RNA profiling technology to examine differences in RNA expression between hybrids and their inhybrid parents. Using this approach, it was possible to sort out cDNA fragments into different categories, depending on their relative levels of expression in a given hybrid and its two parents.
  • CuraGen CuraGen Corp., New Haven CT
  • the degree of heterosis varies tremendously among hybrids from different parental combinations. In cu ⁇ ent breeding practice, selection for parent combinations which give a high degree of heterosis depends on top-cross yield tests.
  • new methods of monitoring heterosis by identifying genes and gene expression patterns associated with heterosis expression are provided. Specific gene expression patterns associated with heterosis are identified prior to yield testing. This allows screening of larger numbers of top- crosses without having to yield test all combinations.
  • non-optimally expressed genes in existing commercial hybrids can be identified and improved by transgenic manipulation or gene-expression profile assisted selection.
  • PAR poly(ethylene glycol) names
  • Figure 1 graphically represents the correlation between degree of heterosis and % relationship: % relationship is designated on the X axis; Fl-MP heterosis in bu/LCR is given on the Y axis. Data was obtained from 4 locations in JH97.
  • RNAs in each F hybrid were expressed at the same levels as in both parental inbreds. Genetically distantly related inbreds, e.g., the parents of commercial hybrids, had less than 6% of the mRNAs differentially expressed. The number of differentially expressed RNA bands between two inbred parents was positively correlated with the corresponding hybrid yield, demonstrating that either gene expression differences and/or DNA sequence polymorphism between inbred parents are important for heterosis.
  • RNA expression in the hybrid can differ from one inbred parent or the other (dominant), or both (additive or over-/under-dominant).
  • Figure 2 depicts the classification of gene expression patterns in FI hybrids relative to the inbred parents. RNA levels are provided on the vertical axis. Bands in each class exhibited the following expression patterns: (A) Over/under-dominant class: the level of expression in FI hybrid is at least two folds higher or lower than both parents, which have either equal or different levels of expression. In the additive. The majority of RNA expression level differences in both tissues of all hybrids analyzed were in the (B) additive and (C) dominant classes, the mRNA levels of the inbred parents are different.
  • Additive class Fl's expression level falls within the range of the two parents.
  • Dominant class the level of expression in FI hybrid is equal to one parent but different from the other. Two-thirds of the differences observed exhibited additive expression, and the rest of the differences demonstrated a dominant expression pattern.
  • RNA fragments correlated with the degree of heterosis.
  • Hybrid yield in bu/LCR is given on the X axis, while % of bands in each expression class is given on the Y axis (% of bands different: dotted line; % of additive bands: dashed line; and % of dominant bands: solid line).
  • Table 2 The number of genes exhibiting over-/under-dominant. additive or dominant expression patterns in heterotic and non heterotic hybri Expression data derived from 1 replicate/sample. [Note: Discrepancies between table 2 and Table 3 are likely due to different number of sam used.]
  • EXAMPLE 2 PREDICTING HETEROSIS FROM ANALYSIS OF SHARED ADDITIVE BANDS; IDENTIFICATION OF GENES INVOLVED IN HETEROSIS Immature ear mRNA was profiled from 10 hybrids and their respective inbred parents .
  • the genotypes profiled included a number of commercial hybrids and a set from the "PAR 27 series, " in which PAR 27 was used as a common female with a series of males that differed in percent relationship.
  • FI tends to have an expression level closer to the parents with higher expression. This is especially true with the RNA bands that are similar between an FI hybrid and its male parent.
  • Hybrids derived from two inbreds that have optimal complementation to each other to give rise to an heterozygosity condition for most of these regulatory elements had a maximal number of genes "re-activated” and were therefore, heterotic.
  • Crosses of closely related inbreds or inbred lines that did not have such "optimal complementation” had fewer genes re-activated and produced low heterotic hybrids.
  • RNA profile data described in Example 1 are based on the expression patterns of FI hybrids relative to their inbred parents, such as additive vs. non additive classifications and the differences of these catego ⁇ es between heterotic and non-heterotic hyb ⁇ ds. While the results so far were informative, another way of analyzing this data set by comparing the levels of RNA expression of poor hyb ⁇ ds with heterotic hyb ⁇ ds without any involvement of their parents. In compa ⁇ ng all 10 hyb ⁇ ds, which include 3 breeding crosses and 7 commercial hybrids, a list of bands that have similar expression level among heterotic hyb ⁇ ds but different from the non-heterotic hyb ⁇ ds
  • EXAMPLE 5 EXPRESSION PROFILING USING DIFFERENT TISSUES FROM HYBRIDS AND PARENTS
  • RNA profiling data from hyb ⁇ d sets were obtained in maize. Five other sets utilized kernel tissue at 13 days after pollination
  • DAP seedling tissue
  • the 14 hybrid sets analyzed included seven from the PAR 27 series, which covers a spectrum of heterosis levels ranging from commercial hyb ⁇ ds to low heterotic hybrids of sibling crosses; four commercial hyb ⁇ ds from diversified genetic backgrounds other than PAR 27 series and three crosses between inbreds of the same heterotic group, typical of those that would be useful for breeding new mbreds.
  • PAR,, vs. PAR,, vs. PAR,, vs. PAR,, vs. PAR,, vs. PAR,, vs. PAR 27 /PAR 25 PAR,, vs. PAR,, vs. PAR,, vs. Band ID PAR 12 PAR except PAR, vs. PAR, 5 PAR permitting/PAR, precisely PAR-j PAR,, (PAR 2 , cross) PAR 27 /PAR 4 , PAR 27 /PAR 44 PAR 27 /
  • RNA expression of poor hybrids and heterotic hybrids are compared without any involvement of their parents.
  • This approach examines whether the absolute level of expression of a subset of genes are important for heterosis, in addition to the additive vs. non-additive expression patterns we already found.
  • the FI hybrids tend to have the same expression levels as the higher parent, i.e. showing overall an up-regulation of gene expression (Table 10).
  • 34 bands that have a similar expression level among heterotic hybrids but different from the non-heterotic hybrid were identified (Table 9; the last three columns are non-heterotic hybrids). For these 34 bands, the 3 poor hybrids show either higher or lower expression than PAR 19 whereas all other hybrids, which are heterotic, show no or little differences in the expression relative to PAR I9 .

Landscapes

  • Chemical & Material Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Organic Chemistry (AREA)
  • Analytical Chemistry (AREA)
  • Engineering & Computer Science (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Health & Medical Sciences (AREA)
  • Wood Science & Technology (AREA)
  • Zoology (AREA)
  • Biotechnology (AREA)
  • Biochemistry (AREA)
  • General Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Microbiology (AREA)
  • Immunology (AREA)
  • Biophysics (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Molecular Biology (AREA)
  • General Health & Medical Sciences (AREA)
  • Genetics & Genomics (AREA)
  • Botany (AREA)
  • Mycology (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
  • Breeding Of Plants And Reproduction By Means Of Culturing (AREA)

Abstract

Methods of correlating molecular profile information and heterosis are provided. Selection for dominant, additive, or under/overdominant markers provides for improved heterosis. Selection for the number of expression products in an expression profile provides for improved heterosis. Methods of identifying and cloning nucleic acids linked to heterotic traits are provided. Methods of identifying parentage by consideration of expression profiles are provided.

Description

PATENT
MOLECULAR PROFILING FOR HETEROSIS SELECTION
FIELD OF THE INVENTION The invention relates to new methods of improving crop selection and selecting for heterosis using molecular and computer modeling techniques
CROSS-REFERENCE TO RELATED APPLICATIONS
This application is a non-provisional filing of and claims pπoπty to "MOLECULAR PROFILING FOR HETEROSIS" by Ben Bowen et al , USSN 60/116,617 filed January 21, 1999 and "MOLECULAR PROFILING FOR HETEROSIS" by Ben Bowen et al , USSN
60/166,368 filed November 17, 1999
BACKGROUND OF THE INVENTION Hybπd offspring often outperform their parents by a variety of different measures, including yield, adaptability to environmental changes, disease resistance, pest resistance, and the like The improved properties for the hybπd as compared to the parents are collectively referred to as "hybrid vigor," or "heterosis " Hybπdization between parents of dissimilar genetic stock has been used in animal husbandry and especially for improving major plant crops, such as corn, sugarbeet and sunflower
Indeed, for some crops, such as corn (Zea mays), most of the crop which is grown is hybπd offspπng Because crossing these hybrid offspπng results in a loss of vigoi and lack of uniformity, the production of seed of these crops for planting is complex, utilizing mbred strains that are crossed to produce hybπd seed with uniform characteπstics
For example, the development of a maize hybπd typically involves three steps (1) the selection of plants from vaπous germplasm pools for initial breeding crosses, (2) the self g of the selected plants from the breeding crosses for several generations to produce a seπes of inbred lines, which, although different from each other, breed true and are highly uniform, and (3) crossing the selected mbred lines with different mbred lines to produce hybrid progeny (sometimes referred to as "FI" hybπds) Duπng the inbreeding process in maize, the vigor of the lines decreases Vigor is restored when two different inbred lines are crossed to produce hybπd progeny A consequence of the homozygosity and homogeneity of the mbred lines is that hybπds produced by crossing a defined pair of mbreds are uniform and predictable. Once the mbreds that give a supeπor hybπd have been identified, the hybπd seed can be reproduced for as long as the homogeneity of the mbred parents is maintained.
Despite many years of research and the considerable commercial importance of generating hybπds with desirable traits, the molecular basis for heterosis is still essentially unknown. In a few cases, the loss of vigor due to inbreeding can be traced directly to a combination of undesirable genes (e.g., lethal or sublethal recessives). However, the simple genetic combination of such genes is not at all sufficient to explain the phenomenon of heterosis. Even when crosses are optimized to eliminate such problematic genes, the resulting offspπng still show a decrease m vigor when inbred. Furthermore, many phenotypic traits, such as yield, are the result of several interacting genes and it is unclear why combining parents with different genetic backgrounds results in an increase in yield. Indeed, it is not even clear whether heterosis is the result of one or a few general genetic mechanisms, or whether it is the result of many simultaneously interacting processes.
Because of the lack of understanding of the molecular basis for heterosis, crop development has relied upon empiπcal observations of heterosis for hybrids which result from crossing selected mbred crop strains (or resulting from second order crosses, e.g., in which two mbreds are crossed to produce a hybπd which is then crossed with an inbred or hybrid strain to produce a subsequent 3-4 way heterotic hybπd) This laborious process has been conducted on a large scale, resulting in increases in desirable measures of heterosis, such as yield, of several percent per year.
Empiπcal methods based on quantitative genetics theory have resulted in a tripling of hybπd com yield over the last 70 years This has been essential for food secuπty and a major contπbution to the U.S. and world economy. By 2020, the world bank and other groups predict that it will be necessary to double maize production and increase πce and wheat production by 50% to support projected population growth. Such an increase can not be accomplished by increasing acreage in production (there is not enough additional acreage available). It is doubtful that simple empiπcal approaches will be sufficient to increase yield fast enough to meet projected demand.
Molecular methods have been used to a limited extent to supplement crop breeding programs to select desirable inbreds and hybπds. In general, these procedures have been used to identify genetic markers corresponding to desirable or undesirable loci (e.g., "quantitative trait loci" or QTLs) m plants under analysis. Genetic markers represent (mark the location of) specific loci in the genome of a species or closely related species, and sampling of different genotypes at these marker loci reveals genetic vaπation. The genetic vaπation at marker loci can then be descπbed and applied to genetic studies, commercial breeding, diagnostics, cladistic analysis of vaπance, or genotyping of samples Because molecular methods are amenable to high throughput analysis and because they do not require yield testing, they can be used to speed the process of crop development. However, although these techniques are of considerable use, and can and do enhance the efficiency of crop breeding programs, they are not currently used, or useful, as a predictor for the more general phenomenon of heterosis.
Accordingly, there is a need in the art to determine how molecular, or other high-throughput methods, or models, can be applied to predict heterosis in individual organisms and in populations. The present invention provides a number of fundamental discoveπes which make it possible to correlate molecular methods and the phenomenon of heterosis, as well as a variety of additional aspects which will be apparent upon complete review.
SUMMARY OF THE INVENTION It is discovered that the number of gene products expressed at optimum levels in an organism such as a plant correlates with the degree of heterosis the organism displays. Thus, by profiling the expression of RNA or protein m a tissue of a plant, it is possible to predict the level of heterosis the plant will display if tested for a heterotic trait such as yield Use of this correlation permits initial selection of organisms, such as commercial crops, without actual field testing Because of the high throughput nature of molecular methods which can be used to profile expression, this initial selection dramatically speeds the process of increasing desirable traits (and decreasing undesirable traits), resulting in an increase in the rate, e.g , of crop improvement.
It is additionally discovered that there is a correlation between the number of dominant and additive expression products and the heterosis an organism such as a plant displays As above, determination of the number (and/or ratio) of dominant and or additive expression products permits selection of plants for heterosis without field testing. In all cases, profiling methods are used to determine the number, and/or relative ratio of any or all of additive, dominant, or under- or over-dommant expression products, thereby providing methods of selecting plants for increased heterosis based upon observed expression profiles In addition, modeling methods for predicting which crosses from a panel of potential crosses are most likely to result in increases in the number of expressed genes, or the number or ratio of additive or dominant genes, or which minimize the ratio of under- or over-dominant genes- are provided. New selection methods for obtaining desirable plants, and plants obtained by these methods are provided
It is additionally discovered that gene silencing plays a role in heterosis Thus, by monitoπng silencing of genes, it is possible to identify which genes are responsible for heterosis Thus, in one aspect, a heterologous nucleic acid that results in expression of expression products from silenced genes (e g., dominant or additive products) is introduced into a target plant. Examples of appropπate heterologous nucleic acids include one or more of: a transcπption factor which activates a promoter from a silenced gene, a nucleic acid encoded by the silenced gene under the control of a heterologous promoter, and a nucleic acid homologous to the silenced gene with at least one region of difference with the silenced gene, which homologous nucleic acid can recombine with the silenced gene to produce a modified gene. Any of these nucleic acids can be cloned under the control of heterologous promoters and placed into target plants to increase heterosis of the target plants.
In desirable implementations of the methods herein, integrated systems comprising computer databases having expression profile information can be used to select which parental crosses are most likely to result in an increase in the number of expression products (or an optimization of expression products of a selected class, i.e., dominant, under- dominant, over-dommant, additive, or the like) in offspπng Thus, consideration of expression profile information provides not only a basis for selecting hybrids from crosses, but, using the methods herein, also identifies desirable crosses to be made. Production and automated consideration of expression profile databases also provides a mechanism for identifying the genetic source of particular expression products, thereby indicating the likely parentage of given hybπds
The invention additionally provides methods of cloning and transducing target plants or animals with dominant, additive, under-dominant and over-dommant genes identified by comparative examination of expression profiles BRIEF DESCRIPTION OF THE FIGURES
Figure 1 is a scatter plot showing the correlation between the degree of heterosis and % relationship.
Figure 2 is a set of bar graphs showing classification of gene expression patterns in Hybrid vs. inbred parents.
Figure 3 is a line graph showing the coπelation between the pattern of gene expression and heterosis.
Figure 4 is a set of bar graphs showing dominant, additive and over-/under- dominant RNA expression. Figure 5 is a scatter graph showing the correlation between parental effects on gene expression and heterosis.
Figure 6a-c is a set of schematic illustrations showing polymorphic dominant products and their sequences.
DEFINITIONS An "expression profile" is the result of detecting a representative sample of expression products from a cell, tissue or whole organism, or a representation (picture, graph, data table, database, etc.) thereof. For example, many RNA expression products or a cell or tissue can simultaneously be detected on a nucleic acid array, or by the technique of differential display or modification thereof such as Curagen's "GeneCalling™" technology. Similarly, protein expression products can be tested by various protein detection methods, such as hybridization to peptide or antibody arrays, or by screening phage display libraries. A "portion" or "subportion" of an expression profile, or a "partial profile" is a subset of the data provided by the complete profile, such as the information provided by a subset of the total number of detected expression products. An "expression product" is any product transcribed in a cell from a DNA (e.g., from a gene) or translated from an RNA (e.g., a protein). Example expression products include mRNAs and proteins.
A "representative sample" of expression products, e.g., from a particular cell, tissue, or whole organism is a sufficiently large number of expression products that statistical comparison of the actual number and/or type of expression products between different cells, tissues, or whole organisms can be made. Ideally, at least about 50%, and typically 60%, 70%, 80%, 90%, 95% or 100% of the total expression products which are detectable by a given technique constitute the "representative sample." The representative sample will typically include a large number of expression products, as cells, tissues and organisms typically produce a fairly large number of expression products. For example, a typical representative sample of expression products includes between about 100 and 20,000 or more expression products, e.g., about 100-500, 1,000, 1,500, 2,000, 2,500, 3,000, 3,500, 4,000, 4,500, 5,000, 5,500, 6,000, 6,500, 7,000, 7,500, 8,000, 8,500, 9,000, 9,500, 10,000, 10,500, 11,000, 11,500, 12,000, 12,500, 13,000, 13,500, 14,000, 14,500, 15,000, 15,500, 16,000, 16,500, 17,000, 17,500, 18,000, 18,500, 19,000, 19,500, 20,000, or 30,000 expression products, or the like.
The term "correlation" unless indicated otherwise, is used herein to indicate that a "statistical association" exists between, e.g., an expression product and the degree of heterosis.
"Dominant" expression for an expression product refers to the situation where expression of the product in a progeny differs from one parent, and not the other for the expression product "Additive" expression for an expression product refers to the situation where expression of the product in a progeny falls within the range of the two parents (and may or may not differ from both parents). "Over-dominant" or "under-dommant" expression for an expression product refers to the situation where expression of an expression product in a progeny differs from both parents and falls outside of the range of the two parents, either over the higher parent value, or under the lower parent value, respectively (Figure 2). Further, the term "differ" when referπng to values is dependent on the technologies being utilized For example, when using Curagen's "GeneCal ng™" technology, any differences in value less than approximately 1.5 to 2.0 fold different from a given parent is considered not to differ
A "biological sample" is a portion of mateπal isolated from a biological source such as a plant, isolated plant tissue, or plant cell, or a portion of mateπal made from such a source, such as a cell extract or the like
A "promoter" is an array of nucleic acid control sequences which direct transcπption of a nucleic acid As used herein, a promoter includes necessary nucleic acid sequences near the start site of transcπption, such as, in the case of a polymerase II type promoter, a TATA element. A promoter also optionally includes distal enhancer or repressor elements which can be located as much as several thousand base pairs from the start site of transcription. A "constitutive" promoter is a promoter which is active in a selected organism under most environmental and developmental conditions. An "inducible" promoter is a promoter which is under environmental or developmental regulation in a selected organism. -
The phrase "hybrid plants" refers to plants which result from a cross between genetically different individuals.
The phrase "sexually crossed" or sexual reproduction" in the context of seed crop plants refers to the fusion of gametes to produce, e.g., seed by pollination. A "sexual cross" is pollination of one plant by another. "Selfing" is the production of, e.g., seed by self- pollination, i.e., where the pollen and the ovule are from the same plant.
The phrase "tester parent" refers to a parent that is genetically different from a set of lines to which it is crossed. The cross is for purposes of evaluating differences among the lines in topcross combination. Using a tester parent in a sexual cross allows one of skill to determine the genetic differences between the tested lines on the phenotypic trait with expression of quantitative trait loci in a hybrid combination.
The phrases "topcross combination" and "hybrid combination" refer to the processes of crossing a single tester parent to multiple lines. The purposes of producing such crosses is to evaluate the ability of the lines to produce desirable phenotypes in hybrid progeny derived from the line by the tester cross.
The phrase "transgenic plant" refers to a plant into which exogenous polynucleotides have been introduced by any process other than sexual cross or selfing. Examples of processes by which this can be accomplished are described below, and include Agrobαcteπ'wm-mediated transformation, biolistic methods, electroporation, in planta techniques, and the like. Such a plant containing the exogenous polynucleotides is referred to here as an Rl generation transgenic plant. Transgenic plants may also arise from sexual cross or by selfing of transgenic plants into which exogenous polynucleotides have been introduced. DETAILED DESCRIPTION OVERVIEW OF SELECTION FOR HETEROSIS
Crop improvement relies extensively on the phenomenon of heterosis. Inbreds and/or hybπds are crossed to produce heterotic hybπds with desirable traits such as high yield, disease resistance, resistance to heat, cold, salinity, insects, fungi, herbicides, pesticides, etc. Secondary desirable traits such as a particular size or shape of ears, solids content, sugar content, oil content, water content, etc., can also be affected by heterosis. The present invention establishes several correlations between the expression of gene products and heterosis, e.g., with respect to yield. These include a statistical association between the number of gene products and the degree of heterosis displayed; a statistical association between the number of gene products with a dominant expression pattern and the degree of heterosis displayed and a statistical association with the number of gene products with an additive expression pattern and the degree of heterosis displayed In addition, it is discovered that genes are silenced duπng inbreeding in plants. These correlations provide new methods of selecting heterotic hybπds, without the necessity of field testing every hybπd to monitor heterotic traits. In the methods, expression of a first representative sample of first expression products (e.g., RNAs or proteins) is profiled from a first progeny plant (e.g., a hybπd from resulting from crossing two or more parental lines). The expression products produced in the first progeny plant are quantified and/or monitored for the type of expression product (additive, dominant, under- dominant, over-dominant, etc ). As noted above, the number of first expression products produced in the first progeny plant is statistically associated with a measure of heterosis in the first progeny plant, as is the number of dominant, additive, under-dommant or over-dom ant, or silenced expression products. The plant is then selected (e.g., against similar measures for a second progeny plant, or a population of progeny plants, or against the parental stock) for further testing based upon the number or type of expression products detected. Thus, the plant can be selected for one or more characteπstic, including: a selected number of expression products, a selected number of dominant expression products, a selected ratio of dominant expression products to total expression products, a desired number of over- or under-dommant expression products, a selected ratio of over- or under- dominant expression products to total expression products, a selected number of additive expression products, and a selected ratio of additive expression products to total expression products. Typically, the first progeny plant is selected to maximize the number of dominant expression products and/or to maximize the number of additive expression products, and/or to minimize the number of over- or under-dominant expression products. Crosses can also be selected to minimize silencing in the progeny plant.
The parental plants used to produce the first progeny can also be profiled. Resulting parental expression profiles serve any of a vaπety of purposes. The parental expression profiles can be compared to the first progeny profile to aid in determining whether the progeny show an increase in the number of expression products as compared to parental stocks (thereby indicating that the progeny is likely to be heterotic). In addition, compaπson between the parental expression profiles and the progeny profile is used to determine whether the individual expression products represented m the profile are dominant, additive, under- dominant, over-dominant, or the like The parental expression profiles can also be placed into a database to aid in determining which crosses are most likely to produce heterotic hybπds. Potentially desirable crosses among members of the database are selected by identifying plants likely to produce progeny plants with a selected number of expression products which are dominant, over-dommant, under-dommant or additive. For example, parents are selected to produce the first progeny plant by selecting for complementary expression of dominant or additive expression products between the parents, or by selecting against expression of over-dominant or under-dommant expression products in the parents
An additional statistical association relates to the relationship between parental and progeny plants. It is discovered that plants which exhibit an expression profile that is more similar to the maternal plant than to the paternal plant may be more heterotic. Accordingly, compaπson of the maternal, paternal and progeny expression profiles can be used to monitor this relationship. In addition, multiple crosses to a single female type can be made (or the results predicted by compaπson in a database) and the progeny screened (or predicted) for similaπty to the female type.
As noted above, silencing was determined to play a significant role in the loss of heterosis due to inbreeding. Accordingly, by compaπng parental and progeny plants it is possible to determine which genes are silenced These genes can be rescued, e.g., by cloning the silenced genes and placing them under the control of heterologous promoters, or other strategies noted herein, and transducing the genes back into target plants (e.g., the parental lines, the hybπds, or any other plant). In addition, by compiling database information for which genes are silenced m mbreds, it is possible to decrease silencing in hybπds by selecting crosses where parents have complementary patterns. It is also possible to use these methods to increase the performance (e.g., gram yield, standabihty, etc.) of the inbred lines themselves.
The first progeny plant selected by any of the methods herein, or a subsequent progeny plant, or a transgenic plant as descπbed above can be subjected to any of the field tests appropriate for monitoπng one or more desired traits Thus, the first progeny plant, or a subsequent progeny plant thereof, can be tested for a desired phenotypic trait. The phenotypic trait can be compared between the first progeny plant, or a subsequent progeny plant, and a selected hybrid or inbred plant. The expression profile of the selected hybπd or mbred plant can be compared to an expression profile of the first progeny plant, or the subsequent progeny plant. Nucleic acids differentially expressed between the selected hybπd or mbred plant and the first progeny plant, or the subsequent progeny plant are identified as targets for cloning. Similarly, genes that are expressed high yielding hybπds that are not expressed in low yielding hybrids can be determined by compaπsons of the expression profiles for the high and low yielding hybπds Nucleic acids from (or corresponding to) the differentially expressed genes are cloned for introduction into target nucleic acids After identifying which expression products from the representative sample show an additive, dominant, underdominant, or overdommant expression pattern for at least a portion of the representative sample, or a nucleic acid corresponding to the expression product, can be cloned. The cloned nucleic acid can then be transduced into target plants to test whether the nucleic acid encodes a useful trait, or to improve traits in the target plant. Further details on expression profiling, cloning of nucleic acids, selection of hybπds, integrated systems, screening methods and the like are set forth below. EXPRESSION PROFILING
As set forth below, a vaπety of tissues can be profiled, with immature tissues being preferentially profiled. Immature tissues are prefeπed, because it increases the rate at which crops can be screened, as a plant does not have to be grown to matuπty However, essentially any tissue, or whole plant, can be profiled. A vaπety of profiling methods are available, including hybridization of expressed or amplified nucleic acids to a nucleic acid array, hybridization of expressed polypeptides to a protein array, hybridization of peptides or nucleic acids to an antibody array, subtractive hybridization, differential display and others. CROPS TO BE PROFILED The parental or progeny plants can be inbreds or hybrids. Most commonly, the progeny plant is a hybrid, produced by crossing two different inbred lines, or crossing an inbred line and a hybrid line, or crossing two hybrid lines (which are the result of crossing inbred or hybrid lines), or crossing of more than two lines (e.g., to generate polyploid or recombinant plants) in a single cross. Once a desirable heterotic hybrid is identified, it can be treated as such hybrids typically are in breeding schemes, e.g., it can produced in quantity as seed; it can be top crossed to inbred lines to produce a 3-way hybrid plant; it can be selfed to produce more inbred lines, or the like.
Most, if not all, plants and animals show hybrid vigor. Much of the discussion herein relates to commercially valuable crops, as these are an important target of the methods of the invention. However, the methods are general and can be applied to non-commercial crop plants, fungi, and to the production of animals, including poultry, cattle, sheep, pigs, and the like.
Important commercial crops include both monocots and dicots. Monocots such as plants in the grass family (Gramineae), such as plants in the sub families Fetucoideae and Poacoideae, which together include several hundred genera including plants in the genera
Agrostis, Phleum, Dactylis, Sorgum, Setaria, Zea (e.g., corn), Oryza (e.g., rice), Triticum (e.g., wheat), Secale (e.g., rye), Avena (e.g., oats), Hordeum (e.g., barley), Saccharum, Poa, Festuca, Stenotaphrum, Cynodon, Coix, the Olyreae, Phareae and many others. Plants in the family Gramineae are a particularly preferred target plants for the methods of the invention. Additional preferred targets include other commercially important crops, e.g., from the families Compositae (the largest family of vascular plants, including at least 1,000 genera, including important commercial crops such as sunflower), and Leguminosae or "pea family," which includes several hundred genera, including many commercially valuable crops such as pea, beans, lentil, peanut, yam bean, cowpeas, velvet beans, soybean, clover, alfalfa, lupine, vetch, lotus, sweet clover, wisteria, and sweetpea. Common crops applicable to the methods of the invention include Zea mays, rice, soybean, sorghum, wheat, oats, barley, millet, sunflower, and canola.
TISSUES TO BE PROFILED
As noted above, one advantage of the present invention is that the methods can be performed without the necessity of field testing progeny (field testing can, of course, be - used as a part of, or an adjunct to the other methods herein). An extension of this advantage is that immature tissues can be profiled from a test plant, which speeds the testing process. Thus, although expression profiles can be performed from any tissue or whole organism, in one preferred embodiment, the representative samples are from immature tissues or immature plants. For example, an immature ear of the plant, or a whole seedling plant (or any tissue thereof), can be profiled. It will be appreciated that when comparisons are performed, they are typically performed between expression profiles obtained from the same tissue and developmental stage (and environmental conditions) for the plants which are compared. RNA PROFILING In one preferred embodiment, the expression products which are detected in the methods of the invention are RNAs, e.g., mRNAs expressed from genes within a cell of the plant or tissue profiled.
A number of techniques are available for detecting RNAs. For example, northern blot hybridization is widely used for RNA detection, and is generally taught in a variety of standard texts on molecular biology, including: Berger and Kimmel, Guide to
Molecular Cloning Techniques. Methods in Enzymology volume 152 Academic Press, Inc., San Diego, CA (Berger); Sambrook et al., Molecular Cloning - A Laboratory Manual (2nd Ed.), Vol. 1-3, Cold Spring Harbor Laboratory, Cold Spring Harbor, New York, 1989 ("Sambrook") and Current Protocols in Molecular Biology, F.M. Ausubel et al., eds., Current Protocols, a joint venture between Greene Publishing Associates, Inc. and John Wiley &
Sons, Inc., (supplemented through 1998) ("Ausubel")).
Furthermore, one of skill will appreciate that essentially any RNA can be converted into a double stranded DNA using a reverse transcriptase enzyme and a polymerase. See, Ausubel, Sambrook and Berger, id. Thus, detection of mRNAs can be performed by converting, e.g., mRNAs into DNAs, which are subsequently detected in, e.g., a standard "Southern blot" format. Furthermore, DNAs can be amplified to aid in the detection of rare molecules by any of a number of well known techniques, including: the polymerase chain reaction (PCR), the ligase chain reaction (LCR), Qβ-rephcase amplification and other RNA polymerase mediated techniques (e g., NASBA) Examples of these techniques are found m Berger, Sambrook, and Ausubel, id., as well as in Mulhs et al, (1987) U.S. Patent No
4,683,202; PCR Protocols A Guide to Methods and Applications (Innis et al. eds) Academic Press Inc. San Diego, CA (1990) (Innis), Arnheim & Levinson (October 1, 1990) C&EN 36- 47; The Journal Of NIH Research (1991) 3, 81-94, Kwoh et al. (1989) Proc. Natl. Acad. Sci USA 86, 1173; Guatelli et al. (1990) Proc Natl. Acad Sci USA 87. 1874; Lomell et al. (1989) J. Chn. Chem 35, 1826, Landegren et al , (1988) Science 241, 1077-1080; Van Brunt
(1990) Biotechnology 8, 291-294; Wu and Wallace, (1989) Gene 4, 560; Barπnger et al. (1990) Gene 89, 117, and Sooknanan and Malek (1995) Biotechnology 13: 563-564. Improved methods of cloning in vitro amplified nucleic acids are descπbed in Wallace et al., U.S. Pat. No. 5,426,039. Improved methods of amplifying large nucleic acids by PCR are summaπzed in Cheng et al. (1994) Nature 369: 684-685 and the references therein, in which
PCR amphcons of up to 40kb are generated. One of skill will appreciate that essentially any RNA can be converted into a double stranded DNA suitable for restπction digestion, PCR expansion and sequencing using reverse transcnptase and a polymerase. See, Ausubel, Sambrook and Berger, all supra These general methods can be used for expression profiling. For example, arrays of probes can be spotted onto a surface and expression products (or in vitro amplified nucleic acids corresponding to expression products) can be labeled and hybπdized with the array For convenience, it may be helpful to use several arrays simultaneously. It is expected that one of skill is familiar with nucleic acid hybπdization. General methods of hybπdization are found in Berger, Sambrook and Ausubel, supra, and further in Tijssen (1993) Laboratory
Techniques in Biochemistry and Molecular Biology— Hybπdization with Nucleic Acid Probes, e.g., part I chapter 2 "Overview of principles of hybπdization and the strategy of nucleic acid probe assays," Elsevier, New York
In one useful vaπation of these methods, solid phase arrays are adapted for the rapid and specific detection of multiple polymorphic nucleotides. Typically, a nucleic acid probe is chemically linked to a solid support and a target nucleic acid (e.g., an RNA or corresponding amplified DNA) is hybridized to the probe. Either the probe, or the target, or both, can be labeled, typically with a fluorophore. Where the target is labeled, hybridization is detected by detecting bound fluorescence. Where the probe is labeled, hybridization is typically detected by quenching of the label by the bound nucleic acid. Where both the probe and the target are labeled, detection of hybridization is typically performed by monitoring a - signal shift such as a change in color, fluorescent quenching, or the like, resulting from proximity of the two bound labels.
In one embodiment of this concept, an array of probes are synthesized on a solid support. Using chip masking technologies and photoprotective chemistry, it is possible to generate ordered arrays of nucleic acid probes with large numbers of probes. These arrays, which are known, e.g., as "DNA chips," or as very large scale immobilized polymer arrays ("VLSIPS"™ arrays) can include millions of defined probe regions on a substrate having an area of about 1cm2 to several cm2. In addition to photomasking technologies, arrays of chemicals, nucleic acids, proteins or the like can also be printed on a solid substrate using printing technologies.
The construction and use of solid phase nucleic acid arrays to detect target nucleic acids is well described in the literature. See, Fodor, et al. Science 251:767 (1991); Sheldon, et al. Clin. Chem. 39(4):718 (1993); Kozal, et al. Nature Medicine 2(7):753 (1996) and Hubbell, U.S. Pat. No. 5,571,639. In brief, a combinatorial strategy allows for the synthesis of arrays containing a large number of probes using a minimal number of synthetic steps. For instance, it is possible to synthesize and attach all possible DNA 8-mer oligonucleotides (48, or 65,536 possible combinations) using only 32 chemical synthetic steps. In general, these procedures provide a method of producing 4n different oligonucleotide probes on an array using only 4n synthetic steps. Light-directed combinatorial synthesis of oligonucleotide arrays on a glass surface is performed with automated phosphoramidite chemistry and chip masking techniques similar to photo resist technologies in the computer chip industry. Typically, a glass surface is derivatized with a silane reagent containing a functional group, e.g., a hydroxyl (for nucleic acid arrays) or amine group (for peptide or peptide nucleic acid arrays) blocked by a photolabile protecting group. Photolysis through a photolithogaphic mask is used selectively to expose functional groups which are then ready to react with incoming 5'-photoprotected nucleoside phosphoramidites. The phosphoramidites react only with those sites which are illuminated (and thus exposed by removal of the photolabile blocking group). Thus, the phosphoramidites only add to those areas selectively exposed from the preceding step. These steps are repeated until the desired array of sequences have been synthesized on the solid surface. Combinatorial synthesis of different oligonucleotide analogues at different locations- on the array is determined by the pattern of illumination during synthesis and the order of addition of coupling reagents. Monitoring of hybridization of target nucleic acids to the array is typically performed with fluorescence microscopes or laser scanning microscopes.
In addition to being able to design, build and use probe arrays using available techniques, one of skill is also able to order custom-made arrays and array-reading devices from manufacturers specializing in aπay manufacture. For example, Affymetrix Corp. in Santa Clara, CA manufactures nucleic acid arrays.
It will be appreciated that probe design is influenced by the intended application. For example, where several allele-specific probe-target interactions are to be detected in a single assay, e.g., on a single nucleic acid chip, it is desirable to have similar melting temperatures for all of the probes. Accordingly, the length of the probes are adjusted so that the melting temperatures for all of the probes on the array are closely similar (it will be appreciated that different lengths for different probes may be needed to achieve a particular Tm where different probes have different GC contents). Although melting temperature is a primary consideration in probe design, other factors are also optionally used to further adjust probe construction, such as elimination of self-complementarity in the probe (which can inhibit hybridization of a target nucleotide). Techniques for designing and using sets of probes for screening many nucleic acids, such as expression products, simultaneously, and for monitoring expression on nucleic acid arrays are described in EP 0799 897 Al. One way to compare expression products between two cell populations is to identify mRNA species which are differentially expressed between the cell populations (i.e., present at different abundances between the cell populations). In addition to the array techniques noted above, another prefeπed method is to use subtractive hybridization (Lee et al. (1991) Proc. Natl. Acad. Sci. (U.S.A.) 88:2825) or differential display employing arbitrary primer polymerase chain reaction (PCR) (Liang and Pardee (1992) Science 257:967). Each of these methods has been used by various investigators to identify differentially expressed mRNA species. See, Salesiotis et al. (1995) Cancer Lett. 91:47; Jiang et al. (1995) Oncogene 10: 1855; Blok et al. (1995) Prostate 26:213; Shinoura et al. (1995) Cancer Lett. 89:215; Murphy et al. (1993) Cell Growth Differ 4:715: Austruv et al. (1993) Cancer Res. 53:2888: Zhang et al. (1993) Mol. Carcinog. 8:123: and Liang et al. (1992) Cancer Res. 52:6966). The methods have also been used to identify mRNA species which are induced or repressed, e.g.,- by drugs or certain nutrients (Fisicaro et al. (1995) Mol. Immunol. 32:565; Chapman et al. (1995) Mol. Cell. Endocrinol. 108: 108; Douglass et al. (1995) J. Neurosci. 15:2471; Aiello et al. (1994) Proc. Na . Acad. Sci. (U.S.A.) 91 :6231 ; Ace et al. (1994) Endocrinology 134:1305. For the technique of differential display, Liang and Pardee (1992), supra provide theoretical calculations for the selection of 5' and 3' arbitrary primers. Correlation of observed results to the theory is also provided. In practice, 5' primers of less than about 9 nucleotides may not provide adequate specificity (slightly shorter primers of about 8 to 10 nucleotides have been used in PCR methods for analysis of DNA polymorphisms. See also, Williams et al. (1991) Nucleic Acids Research 18: 6531). The primer(s) optionally comprise
5'-terminal sequences which serve to anchor other PCR primers (distal primers) and/or which comprise a restriction site or half-site or other ligatable end. Where a restriction site or amplification template for a second primer is incorporated, the primers are optionally longer than those described above by the length of the restriction site, or amplification template site. Standard restriction enzyme sites include 4 base sites, 5 base sites, 6 base sites, 7 base sites, and 8 base sites. An amplification template site for a second primer can be of essentially any length, for example, the site can be about 15-25 nucleotides in length.
The amplified products are optionally labeled and are typically resolved by electrophoresis on a polyacrylamide gel; the location(s) where label is present are excised and the labeled product species is/are recovered from the gel portion, typically by elution. The resultant recovered product species can be subcloned into a replicable vector with or without attachment of linkers, amplified further, and/or detected, or even sequenced directly. Sequencing methods are described in Berger, Sambrook and Ausubel, supra. Direct sequencing of PCR generated amplicons by selectively incorporating boronated nuclease resistant nucleotides into the amplicons during PCR and digestion of the amplicons with a nuclease to produce sized template fragments has also been proposed (Porter et al. (1997) Nucleic Acids Research 25(8): 1611).
It is expected that one of skill can use, e.g., differential display for expression profiling. In addition, companies such as CuraGen Corp. (New Haven CT) provide robust expression profiling based upon modified differential display techniques. See, e.g., WO
97/15690 by Rothberg et al. Accordingly, one of skill can have expression profiling performed by companies which specialize in such techniques.
PROTEIN PROFILING
In addition to profiling RNAs (or corresponding cDNAs) as described above, it is also possible to profile proteins. In particular, various strategies are available for detecting many proteins simultaneously. As applied to the present invention, detected proteins, corresponding to expression products, can be derived from one of at least two sources. First, the proteins which are detected can be either directly isolated from a cell or tissue to be profiled, providing direct detection (and, optionally, quantification) of proteins present in a cell. Second, mRNAs can be translated into cDNA sequences, cloned and expressed. This increases the ability to detect rare RNAs, and makes it possible to immediately associate a detected protein with its coding sequence. For purposes of the present invention, it is not necessary even to express nucleic acids in the proper reading frame, as it is typically the presence or absence of an expression product that is, initially, at issue. Even an out of frame peptide is an indicator for the presence of a corresponding RNA.
A variety of hybridization techniques, including western blotting, ELISA assays, and the like are available for detection of specific proteins. See, Ausubel, Sambrook and Berger, supra. See also, Antibodies: A Laboratory Manual, (1988) E. Harlow and D. Lane, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY. Non-hybridization based techniques such as two-dimensional electrophoresis can also be used to simultaneously and specifically detect large numbers of proteins.
One typical technology for detecting specific proteins involves making antibodies to the proteins. By specifically detecting binding of an antibody and a given protein, the presence of the protein can be detected. In addition to available antibodies, one of skill can easily make antibodies using existing techniques, or modify those antibodies which are commercially or publicly available. In addition to the art referenced above, general methods of producing polyclonal and monoclonal antibodies are known to those of skill in the art. See, e.g., Paul (ed) (1998) Fundamental Immunology, Fourth Edition Raven Press, Ltd., New York Coligan (1991) Current Protocols in Immunology Wiley/Greene, NY; Harlow and Lane (1989) Antibodies: A Laboratory Manual Cold Spring Harbor Press, NY; Stites et al. (eds.) Basic and Clinical Immunology (4th ed.) Lange Medical Publications, Los-
Altos, CA, and references cited therein; Goding (1986) Monoclonal Antibodies: Principles and Practice (2d ed.) Academic Press, New York, NY; and Kohler and Milstein (1975) Nature 256:495-497. Other suitable techniques for antibody preparation include selection of libraries of recombinant antibodies in phage or similar vectors. See, Huse et al. (1989) Science 246:1275-1281; and Ward et al. (1989) Nature 341:544-546. Specific monoclonal and polyclonal antibodies and antisera will usually bind with a KD of at least about .1 μM, preferably at least about .01 μM or better, and most typically and preferably, .001 μM or better.
As used herein, an "antibody" refers to a protein consisting of one or more polypeptide substantially or partially encoded by immunoglobulin genes or fragments of immunoglobulin genes. The recognized immunoglobulin genes include the kappa, lambda, alpha, gamma, delta, epsilon and mu constant region genes, as well as myriad immunoglobulin variable region genes. Light chains are classified as either kappa or lambda. Heavy chains are classified as gamma, mu, alpha, delta, or epsilon, which in turn define the immunoglobulin classes, IgG, IgM, IgA, IgD and IgE, respectively. A typical immunoglobulin (antibody) structural unit is known to comprise a tetramer. Each tetramer is composed of two identical pairs of polypeptide chains, each pair having one "light" (about 25 kD) and one "heavy" chain (about 50-70 kD). The N-terminus of each chain defines a variable region of about 100 to 110 or more amino acids primarily responsible for antigen recognition. The terms variable light chain (VL) and variable heavy chain (VH) refer to these light and heavy chains respectively. Antibodies exist as intact immunoglobulins or as a number of well characterized fragments produced by digestion with various peptidases. Thus, for example, pepsin digests an antibody below the disulfide linkages in the hinge region to produce F(ab)'2 a dimer of Fab which itself is a light chain joined to VH-CH1 by a disulfide bond. The F(ab)'2 may be reduced under mild conditions to break the disulfide linkage in the hinge region thereby converting the (Fab')2 dimer into an Fab' monomer. The Fab' monomer is essentially an Fab with part of the hinge region (see, Fundamental Immunology. W.E. Paul, ed., Raven Press, N.Y. (1993), for a more detailed description of other antibody fragments). While various antibody fragments are defined in terms of the digestion of an intact antibody, one of skill will appreciate that such Fab' fragments may be synthesized de novo either chemically or by utilizing recombinant DNA methodology. Thus, the term antibody, as used herein also includes antibody fragments either produced by the modification of whole antibodies or synthesized de novo using recombinant DNA methodologies. Antibodies include single chain antibodies, including single chain Fv (sFv) antibodies in which a variable heavy and a variable light chain are joined together (directly or through a peptide linker) to form a continuous polypeptide.
For purposes of the present invention, antibodies or antibody fragments can be arrayed, e.g., by coupling to an amine moiety fixed to a solid phase array, in a manner similar to that described above for construction of nucleic acid arrays. As above for nucleic acid probes, the antibodies can be labeled, or proteins corresponding to expression products can be labeled. In this manner, it is possible to couple hundreds, or even thousands, of different antibodies to an array.
In one embodiment, a bacteriophage antibody display library is screened with a polypeptide encoded by a cell, or obtained by expression of mRNAs, differential display, subtractive hybridization or the like. Combinatorial libraries of antibodies have been generated in bacteriophage lambda expression systems which are screened as bacteriophage plaques or as colonies of lysogens (Huse et al. (1989) Science 246: 1275; Caton and Koprowski (1990) Proc. Natl. Acad. Sci. (U.S.A.) 87:6450; Mullinax et al (1990) Proc. Natl. Acad. Sci. (U.S.A.) 87:8095; Persson et al. (1991) Proc. Natl. Acad. Sci. (U.S.A.) 88:2432). Various embodiments of bacteriophage antibody display libraries and lambda phage expression libraries have been described (Kang et al. (1991) Proc. Natl. Acad. Sci. (U.S.A.)
88:4363; Clackson et al. (1991) Nature 352:624; McCafferty et al. (1990) Nature 348:552; Burton et al. (1991) Proc. Natl. Acad. Sci. (U.S.A.) 88:10134; Hoogenboom et al. (1991) Nucleic Acids Res. 19:4133; Chang et al. (1991) J. Immunol. 147:3610; Breitling et al. (1991) Gene 104: 147; Marks et al. (1991) J. Mol. Biol. 222:581; Barbas et al. (1992) Proc. Natl. Acad. Sci. (U.S.A.) 89:4457; Hawkins and Winter (1992) J. Immunol. 22:867; Marks et al. (1992) Biotechnology 10:779; Marks et al. (1992) J. Biol. Chem. 267:16007; Low an et al (1991) Biochemistry 30: 10832; Lerner et al. (1992) Science 258:1313.
The patterns of hybridization which are detected provide an indication of the presence or absence of protein sequences. As long as the library or array against which a population of proteins are to be screened can be correlated from one experiment to the next -
(e.g., by noting the x-y coordinates of the library or array member), no sequence information is required to compare expression profiles from one representative sample to another. In particular, the mere presence or absence (or degree) of label provides the ability to determine differences. One advantage of using libraries of antibodies for protein detection is that the individual libraries can be uncharacterized. As long as library members have a set spatial relationship, e.g., gridded on a plate, duplicate plates can be made and label patterns to the set spatial relationship determined.
More generally, peptide and nucleic acid hybridization to arrays or libraries (or even simple two dimensional gels) can be treated in a manner analogous to a bar code label. Any diverse library or array can be used to screen for the presence or absence of complementary molecules, whether RNA, DNA, protein, or a combination thereof. By measuring corresponding signal information between different sources of test material (e.g., different hybrid or inbred plants, or different tissues, or the like), it is possible to determine differences in expression products for the different source materials. As set forth below, this process is facilitated by various high throughput integrated systems set forth below.
In addition to array based approaches, mass spectrometry is in use for identification of large sets of proteins in samples, and is suitable for identification of many proteins in a sequential or parallel fashion. For example, Hutchens et al. U.S. Pat. 5,719,060, describe methods and apparatus for desorption and ionization of analytes for subsequent analysis by mass spectroscopy and/or biosensors. Sample presenting means with probe elements with "Surfaces Enhanced for Laser Desorption/Ionization" (SELDI) described in the '060 patent is particularly useful in the context of the present invention; however, other approaches described in the '060 are also generally applicable to the present invention. Two and three dimensional gel based approaches can also be used for the specific and simultaneous identification and quantification of large numbers of proteins from biological samples. Multi-dimensional gel technology is well-known and described e.g., in Ausubel, supra, Volume 2, Chapter 10. Image analysis of multi-dimensional protein separation gels provides an indication of the proteins that are expressed e.g., in a cell or tissue type. It is worth noting that identification of particular proteins is not necessary; instead, positional and pattern information e.g , of protein staining or fluorescmg patterns is sufficient to identify sets of protein expression products.
In addition to identifying expression products, such as proteins or RNA, it is also possible to screen for large numbers of metabolites in cell or tissue samples. The presence, absence or level of a metabolite can be treated as a character for compaπson purposes in the same way that nucleic acids or proteins are discussed herein. Metabolites can be monitored by any of currently available method, including chromatography, urn or multi dimensional gel separations, hybπdization to complementary molecules, or the like
The invention provides methods of identifying plant crosses with an increase in probability for heterosis progeny plants. For example, in a preferred method, the expression profiles for a plurality of plants are compared, and the expression profiles are considered by pair-wise comparison. Desirable crosses produce progeny with a selected or optimal number of expression products, or progeny with a selected number or type of expression products that display a dominant, additive, over-dommant or under-dommant expression pattern. Desirably, these compaπsons are performed in an integrated system which includes a computer The generation and use of databases of expression profile information for performing a vaπety of comparisons is a feature of the invention. Because of the large number of compaπsons between expression profiles (which, as noted above, compπse e.g., detection information from about 1,000 to about 20,000 or more expression products), the most practical way of performing the comparisons is by enteπng the information into one or more database and using a computer to make the comparisons.
A vaπety of comparative methods can be performed in an integrated system, e.g., to determine the heterosis (or likely heterosis) of a cross. For example, one simple measure that can be compared across different actual or potential crosses to determine the desirability of a particular cross is to determine the sum of the expressed gene products that differ from a progeny plant in each of a first and second parental plant and the number of expressed gene products that differ between the first and second parental plant. The larger this sum, typically, the more desirable the cross.
In the integrated systems herein, it is also possible to predict the likely outcomes of crosses between parental plants. In these methods, matrices of possible expression profile combinations for plants are generated. For example, the expression profiles are compiled in a database in a computer and a matrix of possible pair-wise expression profile combinations for the plants is generated and queried using an integrated system comprising a computer with software for generating and comparing matrices. Subsets of potential crosses from all of the possible pair-wise comparisons which exhibit a maximal number of expression profile differences represent one preferred cross. Useful software aids in determining how many genes are expressed, or whether expressed genes are additive, dominant, over-dominant or under-dominant.
Which plants to select as possible crosses is up to the discretion of the user. It is possible simply to test all possible first order crosses in a database. However, it is not possible to test all possible subsequent crosses, as the set size for such a procedure is theoretically infinite. That is, after generating a progeny matrix of expression products for all possible pair-wise parental crosses, the progeny matrix can be used to generate a possible theoretical set of crosses between the hypothesized progeny represented by the progeny matrix and/or the original database of parental expression profiles. A resulting expression profile matrix can be generated for hypothesized subsequent progeny, which can again be compared to any of the preceding expression profile information. In theory, this process can be repeated ad infinitum.
More practically, certain rules can be implemented to reduce the total amount of calculations to be performed. For example, matrix information can be limited to possible pair-wise crosses for plants from different heterotic groups, or from the same heterotic group.
In addition, the fidelity of predicted expression profile information increasingly varies as subsequent cross information is considered, and of course, the number of possible crosses increases. Accordingly, typically only one or a few rounds of potential crosses are considered at one time. In any case, selection of a subset of potential crosses from all of the possible pair-wise comparisons which exhibit a maximal number of expression profile differences is desirable. A variety of rules for performing the basic comparisons can be used. In one desirable implementation, crosses are identified in which the sum of: (i) expression products produced in a first plant from a first heterotic group (A,) which are not expressed in a second plant from the first heterotic group (A) to which the first plant is crossed (A.), and which are not expressed in a selected third plant from a second heterotic group (B), plus (ii) the expression products produced in A. which are not produced A, and which are not produced in B, is optimized. This optimization results in crosses which achieve elevated numbers of expression products expressed in heterotic hybrid progeny, and also in an optimization of the number of dominant products expressed. In another optimization protocol, optimization is achieved by determining all possible pair-wise combinations from the first heterotic group and identifying the cross which results in the largest sum of expression products, or by determining all possible pair-wise combinations from the first heterotic group and identifying crosses which result in a hybrid progeny (A, x A.) with a maximal number of differences as compared to B, or by determining all possible pair-wise combinations from the first heterotic group and identifying crosses which result in the hybrid progeny (A, x A..) having a greater number of differences with B than the number of differences between B and A, or B and A.. As above, this optimization results in crosses which achieve elevated numbers of expression products expressed in heterotic hybrid progeny, and also in an optimization of the number of dominant products expressed.
Such implementations can also be used to improve selection methods per se. For example, in one method, self- or back-crossed progeny derived from the A, x A hybrid are selected which either retain a set of expression products defined by the sum of expression products expressed in A, (but not A. or B) and A. (but not A, or B), or which show a larger number of expression products expressed in a topcross with B than does either A, or A. when topcrossed with B.
One approach for comparing profiles is a nested analysis in which expression profiles are successively grouped together, and the many gene expression differences seen in individual pair-wise compaπsons can be ranked hierarchically in a filtering process. This method is useful for identifying genes expressed in one set of genotypes vs. another, e.g. hybrids vs. inbreds or bulked segregants from the two ends of a quantitative phenotypic distribution.
In any case, the methods of the invention can include inputing an expression profile for progeny or parental plants into a database of expression profiles. This can be performed manually, but is more typically performed in an automated system.
Computer databases of expression profile information can be quite large, with from a few up to several thousand profiles in the database. Typically, the database will have expression product profiles of a representative sample of expression products for hybrid progeny plants resulting from at least 10 separate inbred plant crosses, or at least 10 inbred plant expression product profiles.
The phrase "computer system" or "integrated system" in the context of this invention refers to a system in which data entering a computer corresponds to physical objects or processes external to the computer, e.g., nucleic acid hybridization or protein binding data and a process that, within a computer, causes a physical transformation of the input signals to different output signals. In other words, the input data, e.g., hybridization of expression products on a specific array, is transformed to output data, e.g., the identification or counting of the sequence hybridized, comparison to similar aπays with different test materials, counting and categorization of expression products or the like. The process within the computer is a program by which positive (or negative) hybridization signals are recognized by the computer system and attributed to a region of an array, or other expression profile format (e.g., simple counting of array signals). The program then determines which region of the array the hybridized expression products are located on and, optionally, the specific corresponding sequences which the probe is based on (as noted above, no sequence information is required for making or assessing expression profiles). The invention provides integrated systems for plant or plant cell manipulation and hybridization analysis. Typical systems include a digital computer with high-throughput liquid control software, image analysis software, and data interpretation software. A robotic liquid control armature for transferring solutions (e.g., plant cell extracts) from a source to a destination, is typically operably linked to the digital computer. An input device for entering data to the digital computer to control high throughput liquid transfer by the robotic liquid control armature and, optionally, to control transfer by the pinning armature to the solid support is commonly a feature of the integrated system, as is an image scanner for digitizing label signals from labeled probe hybπdized to the DNA on the solid support operably linked to the digital computer The image scanner interfaces with the image analysis software to provide a measurement of probe label intensity, where the probe label intensity measurement is interpreted by the data interpretation software to show whether, and to what degree, the labeled probe hybridizes to a label.
A number of well known robotic systems have also been developed for solution phase chemistπes. These systems include automated workstations like the automated synthesis apparatus developed by Takeda Chemical Industπes, LTD. (Osaka, Japan) and many robotic systems utilizing robotic arms (Zymate π, Zymark Corporation, Hopkmton,
Mass.; Orca, Hewlett-Packard, Palo Alto, Calif.) which mimic the manual synthetic operations performed by a scientist Any of the above devices are suitable for use with the present invention. The nature and implementation of modifications to these devices (if any) so that they can operate as discussed herein with reference to the integrated system will be apparent to persons skilled in the relevant art
High throughput screening systems are commercially available (see, e.g., Zymark Corp., Hopkmton, MA; Air Technical Industπes, Mentor, OH; Beckman Instruments, Inc Fullerton, CA, Precision Systems, Inc., Natick, MA, etc.). These systems typically automate entire procedures including all sample and reagent pipetting, liquid dispensing, timed incubations, and final readings of the microplate in detector(s) appropπate for the assay These configurable systems provide high throughput and rapid start up as well as a high degree of flexibility and customization For example, the currently available commercial software package, BioWorks® 1 4®, provided by Beckman Instruments, Inc. to control and operate their Biomek® 2000 robotics liquid handler supports a scπpting capability based on the publicly available Tool Command Language (TCL). Beckman has incorporated a TCL interpreter into the Biomek® 2000 and has included TCL extensions (Bioscπpt®) to allow direct motor control and other instrument functionality. A 16-bit (to run under Microsoft Windows 3.1® and Microsoft Windows 95®) application to generate the TCL/Bioscnpt code can be created, e.g., in Microsoft Visual Basic 4.O®. The manufacturers of such systems provide detailed protocols the vaπous high throughput. Thus, for example, Zymark Corp. provides technical bulletins descπbmg screening systems for detecting the modulation of gene transcπption, gand binding, and the like. More recently, microfluidic approaches to reagent manipulation have been developed, e.g., by Cahper Technologies (Palo Alto, CA)
Optical images viewed (and, optionally, recorded) by a camera or other recording device (e.g., a photodiode and data storage device) are optionally further processed- m any of the embodiments herein, e.g., by digitizing the image and/or stoπng and analyzing the image on a computer. A vaπety of commercially available peπpheral equipment and software is available for digitizing, stoπng and analyzing a digitized video or digitized optical image, e.g., using PC (Intel x86 or pentium chip- compatible DOS™, OS2™ WINDOWS™, WINDOWS NT™ or WINDOWS95™ based machines), MACINTOSH™, or UNIX based
(e.g., SUN™ work station) computers
One conventional system carπes light from the specimen field to a cooled charge-coupled device (CCD) camera, in common use in the art. A CCD camera includes an array of picture elements (pixels). The light from the specimen is imaged on the CCD. Particular pixels corresponding to regions of the specimen (e g., individual hybπdization sites on an array of biological polymers) are sampled to obtain light intensity readings for each position Multiple pixels are processed in parallel to increase speed. The apparatus and methods of the invention are easily used for viewing any sample, e.g., by fluorescent or dark field microscopic techniques Integrated systems for hybridization analysis of the present invention typically include a digital computer with high-throughput liquid control software, image analysis software, data interpretation software, a robotic liquid control armature for transferπng solutions from a source to a destination operably linked to the digital computer, an input device (e.g., a computer keyboard) for enteπng data to the digital computer to control high throughput liquid transfer by the robotic liquid control armature and, optionally, an image scanner for digitizing label signals from labeled probe hybridized to expression products, e.g , on a solid support operably linked to the digital computer The image scanner interfaces with the image analysis software to provide a measurement of probe label intensity Typically, the probe label intensity measurement is interpreted by the data interpretation software to show whether the labeled probe hybridizes to the DNA on the solid support. Software to support sample processing can be divided into 4 functional categoπes: 1) liquid transfer control software, 2) image analysis software, 3) data management software, and 4) data interpretation software
Conveniently, applications can share information through data files which the applications can read and create. For flexibility and ease of use, files can be formatted as simple text files and/or in Microsoft Excel® or other worksheet format. This allows viewing and editing of the files through the use of commercially available software such as Microsoft Excel®. Those of skill in the art will recognize that this approach is only one possible set of systems that could be used in the support and facilitation of the process of the present invention. Other systems can easily designed to fit the particular needs of the user in the practice of the invention. By way of example, and not limitation, a Microsoft Windows® user interface can be developed for most applications using Microsoft Visual Basic 4.O®. Most applications can be developed for a 32-bit environment to run under Microsoft Windows 95® or 98®. 16-bit applications such as image analysis software developed by Optimas Corporation, Optimas 5.0, can also be useful components of the integrated system.
CLONING OF EXPRESSION PRODUCTS
Any nucleic acid encoding an expression product identified as being of interest by the expression profiling techniques noted herein, including dominant, additive and over or under dominant expression products can be cloned. It is expected that many such nucleic acids, particularly dominant and additive nucleic acids will be encoded by loci responsible for desirable quantitative traits ("QTL" see, Edwards, et al., (1987) in Genetics 115:113). QTL include genes that control, to some degree, numeπcally quantifiable phenotypic traits such as disease resistance, crop yield, resistance to environmental extremes, etc. In addition to the methods herein, other expeπmental paradigms can be used to identify, analyze and select for QTL. One paradigm involves crossing two mbred lines and genotyping multiple marker loci and evaluating one to several quantitative phenotypic traits among the progeny of the cross. QTL are then identified and ultimately selected for based on significant statistical associations between the genotypic values determined by genetic marker technology and the phenotypic vaπability among the segregating progeny. As applied to the present invention, the identification of particular nucleic acids which encode dominant, additive or under or over dominant expression products, or which encode silenced expression products, are potential products of QTLs or other genes or loci of interest. Accordingly, it is desirable to clone nucleic acids which are genetically linked to DNAs encoding these expression products for transduction into cells (e.g., coding sequences for expression products, or genetically linked coding or non-coding sequences), especially to make transgenic plants. The cloned sequences are also useful as molecular tags- for selected plant strains, e.g., to identify parentage, and are further useful for encoding expression products, including nucleic acids and polypeptides. Often, expression products which are differentially expressed between heterotic and non-heterotic plants are encoded by QTL and are responsible for the phenotypic effects of the QTL. A DNA linked to a locus encoding an expression product is introduced into plant cells, either in culture or in organs of a plant, e.g., leaves, stems, fruit, seed, etc. The expression of natural or synthetic nucleic acids encoded by nucleic acids linked to expression product coding nucleic acids can be achieved by operably linking a cloned nucleic acid of interest, such as an expression product or a genetically linked nucleic acid, to a promoter, incorporating the construct into an expression vector and introducing the vector into a suitable host cell. Alternatively, an endogenous promoter linked to the nucleic acids can be used.
Cloning of Expression Product Sequences into Bacterial Hosts
There are several well-known methods of introducing expression product nucleic acids into bacterial cells, any of which may be used in the present invention. These include: fusion of the recipient cells with bacterial protoplasts containing the DNA, electroporation, projectile bombardment, and infection with viral vectors, etc. Bacterial cells are often used to amplify increase the number of plasmids containing DNA constructs of this invention. The bacteria are grown to log phase and the plasmids within the bacteria can be isolated by a variety of methods known in the art (see, for instance, Sambrook). In addition, a plethora of kits are commercially available for the purification of plasmids from bacteria. For their proper use, follow the manufacturer's instructions (see, for example, EasyPrep™, FlexiPrep™, both from Pharmacia Biotech; StrataClean™, from Stratagene; and, QIAexpress Expression System™ from Qiagen). The isolated and purified plasmids are then further manipulated to produce other plasmids, used to transfect plant cells or incorporated into
Agrobacterium tumefaciens related vectors to infect plants. Typical vectors contain transcπption and translation terminators, transcπption and translation initiation sequences, and promoters useful for regulation of the expression of the particular nucleic acid. The vectors optionally compπse geneπc expression cassettes containing at least one independent terminator sequence, sequences permitting replication of the cassette in eukaryotes, or prokaryotes, or both, (e.g., shuttle vectors) and selection markers for both prokaryotic and eukaryotic systems Vectors are suitable for replication and integration in prokaryotes, eukaryotes, or preferably both. See, Giliman & Smith, Gene 8:81 (1979); Roberts, et al, Nature. 328:731 (1987); Schneider, B., et al., Protein Expr. Punf. 6435: 10 (1995); Berger, Sambrook, Ausubel (all supra). A catalogue of Bacteπa and Bacteπophages useful for cloning is provided, e.g., by the ATCC, e.g., The ATCC Catalogue of Bactena and
Bacteriophage (1992) Ghema et al. (eds) published by the ATCC. Additional basic procedures for sequencing, cloning and other aspects of molecular biology and underlying theoretical considerations are also found in Watson et al (1992) Recombinant DNA. Second Edition Scientific Amencan Books, NY. Transfecting and Manipulating Plant Cells
Methods of transducing plant cells with nucleic acids are generally available. In addition to Berger, Ausubel and Sambrook, useful general references for plant cell cloning, culture and regeneration include Payne et al. (1992) Plant Cell and Tissue Culture in Liquid Systems John Wiley & Sons, Inc. New York, NY (Payne); and Gamborg and Phillips (eds) (1995) Plant Cell, Tissue and Organ Culture; Fundamental Methods Spπnger Lab Manual,
Spπnger-Verlag (Berlin Heidelberg New York) (Gamborg). A vaπety of Cell culture media are descπbed Atlas and Parks (eds) The Handbook of Microbiological Media (1993) CRC Press, Boca Raton, FL (Atlas) Additional information for plant cell culture is found in available commercial literature such as the Life Science Research Cell Culture Catalogue (1998) from Sigma- Aldπch, Inc (St Louis, MO) (Sigma-LSRCCC) and, e.g., the Plant
Culture Catalogue and supplement (1997) also from Sigma-Aldπch, Inc (St Louis, MO) (Sigma-PCCS)
The nucleic acid constructs of the invention are introduced into plant cells, either m culture or in the organs of a plant by a vaπety of conventional techniques. For example, the DNA construct can be introduced directly into the genomic DNA of the plant cell using techniques such as electroporation and microinjection of plant cell protoplasts, or the DNA constructs can be introduced directly to plant cells using ballistic methods, such as DNA particle bombardment. Alternatively, the DNA constructs are combined with suitable T-DNA flanking regions and introduced into a conventional Agrobacterium tumefaciens host vector. The virulence functions of the Agrobacterium tumefaciens host directs the insertion of the construct and adjacent marker into the plant cell DNA when the cell is infected by the- bacteria.
Microinjection techniques are known in the art and well described in the scientific and patent literature. The introduction of DNA constructs using polyethylene glycol precipitation is described in Paszkowski, et al, EMBO J. 3:2717 (1984). Electroporation techniques are described in Fromm, et al, Proc. Nat'l. Acad. Sci. USA
82:5824 (1985). Ballistic transformation techniques are described in Klein, et al, Nature 327:70-73 (1987).
Agrobacterium tumefaciens-medi∑Λed transformation techniques, including disarming and use of binary vectors, are also well described in the scientific literature. See, for example Horsch, et al, Science 233:496-498 (1984), and Fraley, et al., Proc. Nat'l. Acad.
Sci. USA 80:4803 (1983). Agrobacterium-mediated transformation is a prefeπed method of transformation of dicots.
To use isolated sequences corresponding to or linked to expression products in the above techniques, recombinant DNA vectors suitable for transformation of plant cells are prepared. A DNA sequence coding for the desired mRNA, polypeptide, or non-expressed sequence is transduced into the plant. Where the sequence is expressed, the sequence is optionally combined with transcriptional and translational initiation regulatory sequences which will direct the transcription of the sequence from the gene in the intended tissues of the transformed plant. Promoters, in nucleic acids linked to loci identified by detecting expression products, are identified, e.g., by analyzing the 5' sequences upstream of a coding sequence in linkage disequilibrium with the loci. Optionally, such promoters will be associated with a QTL. Sequences characteristic of promoter sequences can be used to identify the promoter. Sequences controlling eukaryotic gene expression have been extensively studied. For instance, promoter sequence elements include the TATA box consensus sequence
(TATAAT), which is usually 20 to 30 base pairs upstream of a transcription start site. In most instances the TATA box aids in accurate transcription initiation. In plants, further upstream from the TATA box, at positions -80 to -100, there is typically a promoter element with a series of adenines suπounding the trinucleotide G (or T) N G. See, e.g., J. Messing, et al, in Genetic Engineering in Plants, pp. 221-227 (Kosage, Meredith and Hollaender, eds. (1983)). A number of methods are known to those of skill in the art for identifying and characterizing promoter regions in plant genomic DNA. See, e.g., Jordano, et al, Plant Cell 1:855-866 (1989); Bustos. et al. Plant Cell 1:839-854 (1989); Green, et al. EMBO J. 7:4035-4044 (1988); Meier, et al. Plant Cell 3:309-316 (1991); and Zhang, et al, Plant Physiology 110:1069-1079 (1996). In construction of recombinant expression cassettes of the invention, a plant promoter fragment is optionally employed which directs expression of a nucleic acid in any or all tissues of a regenerated plant. Examples of constitutive promoters include the cauliflower mosaic virus (CaMV) 35S transcription initiation region, the 1'- or 2'- promoter derived from T-DNA of Agrobacterium tumafaciens, and other transcription initiation regions from various plant genes known to those of skill. Alternatively, the plant promoter may direct expression of the polynucleotide of the invention in a specific tissue (tissue-specific promoters) or may be otherwise under more precise environmental control (inducible promoters). Examples of tissue-specific promoters under developmental control include promoters that initiate transcription only in certain tissues, such as fruit, seeds, or flowers. Any of a number of promoters which direct transcription in plant cells can be suitable. The promoter can be either constitutive or inducible. In addition to the promoters noted above, promoters of bacterial origin which operate in plants include the octopine synthase promoter, the nopaline synthase promoter and other promoters derived from native Ti plasmids. See, Herrara-Estrella et al. (1983), Nature. 303:209-213. Viral promoters include the 35S and 19S RNA promoters of cauliflower mosaic virus. See, Odell et al
(1985) Nature, 313:810-812. Other plant promoters include the ribulose-l,3-bisphosphate carboxylase small subunit promoter and the phaseolin promoter. The promoter sequence from the E8 gene and other genes may also be used. The isolation and sequence of the E8 promoter is described in detail in Deikman and Fischer, (1988) EMBO J. 7:3315- 3327. If polypeptide expression is desired, a polyadenylation region at the 3'-end of the coding region is typically included. The polyadenylation region can be deπved from the natural gene, from a vaπety of other plant genes, or from T-DNA.
The vector compnsing the sequences (e.g., promoters or coding regions) from genes encoding expression products of the invention will typically compπse a nucleic acid - subsequence which confers a selectable phenotype on plant cells. The vector compπsmg the sequence will typically comprise a marker gene which confers a selectable phenotype on plant cells. For example, the marker may encode biocide tolerance, particularly antibiotic tolerance, such as tolerance to kanamycin, G418, bleomycm, hygromycin, or herbicide tolerance, such as tolerance to chlorosluforon, or phosph othπcm (the active ingredient in the herbicides bialaphos and Basta). For example, crop selectivity to specific herbicides can be conferred by engmeeπng genes into crops which encode appropπate herbicide metabolizing enzymes from other organisms, such as microbes. See, Padgette et al. (1996) "New weed control opportunities: Development of soybeans with a Round UP Ready™ gene" In: Herbicide-Resistant Crops (Duke, ed.), pp 53-84, CRC Lewis Publishers, Boca
Raton ("Padgette, 1996"), and Vasil (1996) "Phosphmothπcm-resistant crops" In: Herbicide- Resistant Crops (Duke, ed.), pp 85-91, CRC Lewis Publishers, Boca Raton) (Vasil, 1996). Transgenic plants have been engineered to express a vaπety of herbicide tolerance/metabolizing genes, from a vaπety of organisms. For example, acetohydroxy acid synthase, which has been found to make plants which express this enzyme resistant to multiple types of herbicides, has been cloned into a vaπety of plants (see, e.g., Hattoπ, J., et al. (1995) Mol. Gen. Genet. 246(4):419). Other genes that confer tolerance to herbicides include: a gene encoding a chimeπc protein of rat cytochrome P4507A1 and yeast NADPH- cytochrome P450 oxidoreductase (Shiota, et al. (1994) Plant Physiol. 106(1)17, genes for glutathione reductase and superoxide dismutase (Aono, et al. (1995) Plant Cell Physiol.
36(8): 1687, and genes for vaπous phosphotransferases (Datta, et al. (1992) Plant Mol. Biol. 20(4):619. Similarly, crop selectivity can be conferred by alteπng the gene coding for an herbicide target site so that the altered protein is no longer inhibited by the herbicide (Padgette, 1996). Several such crops have been engineered with specific microbial enzymes for confer selectivity to specific herbicides (Vasil, 1996) Further, nucleic acids which can be cloned and introduced into plants to modify or complement expression of a gene, including a silenced gene, a dominant gene, and additive gene or the like, can be any of a variety of constructs, depending on the particular application. Thus, a nucleic acid encoding a cDNA expressed from an identified gene can be expressed in a plant under the control of a heterologous promoter. Similarly, a nucleic acid - encoding a transcπption factor that regulates a target identified by the methods herein, or that encodes any other moiety affecting transcπption, can be cloned and transduced into a plant Methods of identifying such factors are replete throughout the literature. For a basic introduction to genetic regulation, see, Lewin (1995) Genes V Oxford University Press Inc , NY (Lewm), and the references cited therein.
Regeneration of Transgenic Plants
Transformed plant cells which are deπved by any of the above transformation techniques can be cultured to regenerate a whole plant which possesses the transformed genotype and thus the desired phenotype. Such regeneration techniques rely on manipulation of certain phytohormones in a tissue culture growth medium, typically relying on a biocide and/or herbicide marker which has been introduced together with the desired nucleotide sequences. Plant regeneration from cultured protoplasts is descπbed in Evans, et al., Protoplasts Isolation and Culture, Handbook of Plant Cell Culture, pp 124-176, Macmilhan Publishing Company, New York, (1983), and Binding, Regeneration of Plants. Plant Protoplasts, pp. 21-73, CRC Press, Boca Raton, (1985). Regeneration can also be obtained from plant callus, explants, somatic embryos (Dandekar, et al., J. Tissue Cult. Meth. 12: 145 (1989); McGranahan, et al., Plant Cell Rep 8:512 (1990)), organs, or parts thereof. Such regeneration techniques are descπbed generally in Klee, et al, Ann. Rev, of Plant Phvs 38:467-486 (1987). One of skill will recognize that after the expression cassette is stably incorporated in transgenic plants and confirmed to be operable, it can be introduced into other plants by sexual crossing. Any of a number of standard breeding techniques can be used, depending upon the species to be crossed. GENE SILENCING AND HETEROSIS It is discovered that gene silencing and epigenetic effects play a role in inbreeding depression. As demonstrated herein, the number of genes in hybπds with a dominant pattern of gene expression is correlated with hybπd yield, a component of which is found to be relief from inbreeding depression. An other way of consideπng genes in this class is to classify them as genes that are expressed at lower levels in one inbred parent than the other. When one copy of a gene that is expressed at low levels in one inbred is combined with a copy from another mbred, a frequent outcome in the hybπd is an equivalent level of - expression to that seen with two copies of the gene in one or other of the parental inbreds (most often the more highly expressing parent)
The number of genes in the dominant class were considered as a function of the number of hybπds that share those genes, and the frequency distπbution indicated that the overlap between sets of genes contπbuting to dominant patterns of gene expression in hybπds is essentially random. This suggests that, duπng the process of inbreeding, expression of a subset of genes may always be altered (and usually reduced), and that the expression of different random subsets of genes are silenced in different mbreds.
These results agree well with the classical complementation concepts of metabolic balance and physiological bottlenecks (Hageman et al. 1967 "A biochemical approach to corn breeding" Advan. Agron. 19:45; Schrader, L.E. 1985 "Selection for metabolic balance in maize" pp79-89 in Exploitation of physiological and genetic vaπabihty to enhance crop productivity. Harper J.E.(ed) Waverly Press, Baltimore, and Manglesdorf, A.J. 1952 "Gene interaction in Heterosis, pp321-329 in Heterosis, Gowen, J. (ed) Iowa State College Press, Ames) to explain heterosis. This hypothesis proposes that maize mbred lines have unbalanced metabolic systems with some enzymes at optimum level and some at rate limiting levels, or bottlenecks Hybπds from inbred lines that have different rate limiting systems can overcome the bottlenecks by complementation. Depending on the gene product, a favorable allele can become an unfavorable allele a different developmental stage; and vice-versa. Complementation, therefore results not only from quantitative aspects, i.e , vaπation in the level of expression, but also from qualitative aspects, e.g. vaπation in function due to sequence polymorphisms.
Closely related crosses are less heterotic because, firstly, there are fewer band differences, either in level of expression or in sequence polymorphism, therefore fewer heterozygous loci providing potential opportunities for complementation. Secondly, loci from closely related crosses are more susceptible to gene silencing. In more distantly related crosses, the inbred parents have a higher number of differential bands, and the resulting hybrid tends to express both alleles providing better complementation of unfavorable parental alleles. Such complementation allows for better responses to differing environments or during different developmental stages. Without being bound to a particular theory, epigenetics provide a simple and - elegant explanation for these effects. In Drosophila and other organisms, allelic (and non-allelic) effects have been described where expression in heterozygotes is normal, but in homozygotes trans-inactivation (or silencing) of both alleles occurs. These effects are mediated by cis-acting regulatory sequences that need to be present at more than one copy (e.g. on different chromosome homologs) to mediate the cooperative assembly of multimeric protein complexes responsible for gene silencing (e.g., Polycomb proteins in Drosophila or SIR proteins in yeast). In maize, sequences responsible for these effects most likely occur in intergenic regions outside of the chromatin loops flanked by MARs that contain genes. About 80% of the sequences in these regions are derived from retroelements that may be transcriptionally silenced through natural selection. However, the intergenic regions are also where maize exhibits most DNA sequence polymorphism. Thus, homozygosity of certain intergenic regions in inbreds could lead to adjacent gene silencing, whereas in hybrids fewer intergenic regions will be homozygous for sites that can assemble silencing complexes, so more genes will be derepressed. As new inbreds are created from hybrid crosses, recombination randomizes the intergenic regions across the genome, thereby resulting in a new subset of genes that are silenced when those regions that can assemble silencing complexes are made homozygous. This model explains why inbreds express fewer genes than hybrids (which accounts for their lower yield) and why the number of genes that exhibit a dominant pattern of gene expression in hybrids increases as the percent relationship between inbreds decreases. It also can easily accommodate potential explanations for the existence of heterotic pools, and the higher level of heterosis seen in maize as compared to other cereals (e.g. rice), which have a very different genome organization and level of sequence polymorphism. Finally, it is also possible that in maize, where natural inbreeding occurs infrequently because of its floral characteristics, natural selection may not have acted to eliminate gene silencing at the same rate as in self- fertilizing species. MOLECULAR SECURITY; IDENTIFICATION OF PARENTAL SOURCES BY COMPARISON OF EXPRESSION PROFILES
One general concern in the agricultural industry is that proprietary plant stocks or other sources of germ plasm can sometimes be inadvertently, or even deliberately, misappropriated. Because the germ plasm may be recombined with other sources of germ plasm before producing a product such as a hybrid seed, it is not always possible to tell that the product is improperly derived from proprietary parental plants, clones, or the like.
The present invention provides methods of identifying unique expression products and/or unique profiles (or partial profiles). This ability to identify unique expression products provides one way of ascertaining parentage, which, in turn, provides the ability to determine whether a hybrid comprises proprietary material.
In the methods, a source or the sources of a test plant such as a hybrid can be identified. In the methods, a representative sample of expression products from the test plant is profiled and the resulting test expression profile is compared to a database of known expression profiles for plants from known inbred or hybrid strains (methods of making such databases are described above). For example, the expression profiles for a selected tissue can be entered into a database for any or every proprietary plant (or clone, or any other source of germ plasm) that a corporation owns.
By profiling a number of plants, it is possible to detect unique expression products and/or expression patterns within the expression profile of specific plants. It is also possible to generate likely expression profiles for hybrid products of members of the database. Any of these expression profiles can be compared to an actual expression profile for a test plant suspected of being derived from a one or more proprietary plant. For example, a matrix of pair-wise comparisons for potential progeny from the expression profiles in the database can be compared to the test expression profile. Either the entire expression profile or a sub portion of the expression profile (i.e., a plurality of characters corresponding to expression products found in the overall profile) comprising at least one unique expression marker can be evaluated.
EXAMPLES The following examples are offered by way of illustration, and are not intended to be limiting. One of skill will immediately recognize a variety of alternate procedures, compositions, reagents and the like which can be substituted for those exemplified below.
EXAMPLE 1: DIFFERENCES IN RNA EXPRESSION PROFILES CORRELATE WITH HETEROSIS Heterosis is a term used to describe the increased vigor of hybrid progeny in - comparison to their parents. Although heterosis has been widely used in plant breeding for many decades, the molecular mechanisms underlying the phenomenon were previously unknown. In this example, heterosis was studied as a phenotype using CuraGen (CuraGen Corp., New Haven CT) RNA profiling technology to examine differences in RNA expression between hybrids and their inhybrid parents. Using this approach, it was possible to sort out cDNA fragments into different categories, depending on their relative levels of expression in a given hybrid and its two parents. Data indicated a difference in the number of genes in each category (dominant, under-dominant, over-dominant, additive) between heterotic and non- heterotic hybrids. The results also suggested the ability of this approach to explain the molecular basis of heterosis and the application of the information obtained to plant breeding methods.
The degree of heterosis varies tremendously among hybrids from different parental combinations. In cuπent breeding practice, selection for parent combinations which give a high degree of heterosis depends on top-cross yield tests. In this disclosure, new methods of monitoring heterosis by identifying genes and gene expression patterns associated with heterosis expression are provided. Specific gene expression patterns associated with heterosis are identified prior to yield testing. This allows screening of larger numbers of top- crosses without having to yield test all combinations. Similarly, non-optimally expressed genes in existing commercial hybrids can be identified and improved by transgenic manipulation or gene-expression profile assisted selection.
"PAR" names herein are arbitrary predesignations of commercial and proprietary strain names. Because the invention is applicable to any crop strain, the particular strains used are not critical, or even relevant, to the claimed invention. Accordingly, actual crop strain names are not provided. The PAR, series of hybrids used for RNA profiling are listed in Table 1. In Table 1 , these hybrids range from a highly heterotic commercial hybrid (PAR19 = PAR,/PAR2) to sibling crosses (e.g. PAR,/PAR17) which have much less heterosis. Each hybrid is derived from the same female parent (PAR,) and a male parent with a different percentage pedigree relationship. The coπelation between heterosis and pedigree relationship is given in Table 1. Figure 1 graphically represents the correlation between degree of heterosis and % relationship: % relationship is designated on the X axis; Fl-MP heterosis in bu/LCR is given on the Y axis. Data was obtained from 4 locations in JH97.
Table 1. Hybrids and inbred parents selected for mRNA profiling analysis
In seedlings and immature ears, 90-95% of RNAs in each F, hybrid were expressed at the same levels as in both parental inbreds. Genetically distantly related inbreds, e.g., the parents of commercial hybrids, had less than 6% of the mRNAs differentially expressed. The number of differentially expressed RNA bands between two inbred parents was positively correlated with the corresponding hybrid yield, demonstrating that either gene expression differences and/or DNA sequence polymorphism between inbred parents are important for heterosis.
The level of RNA expression in the hybrid can differ from one inbred parent or the other (dominant), or both (additive or over-/under-dominant). Figure 2 depicts the classification of gene expression patterns in FI hybrids relative to the inbred parents. RNA levels are provided on the vertical axis. Bands in each class exhibited the following expression patterns: (A) Over/under-dominant class: the level of expression in FI hybrid is at least two folds higher or lower than both parents, which have either equal or different levels of expression. In the additive. The majority of RNA expression level differences in both tissues of all hybrids analyzed were in the (B) additive and (C) dominant classes, the mRNA levels of the inbred parents are different. Additive class: Fl's expression level falls within the range of the two parents. Dominant class: the level of expression in FI hybrid is equal to one parent but different from the other. Two-thirds of the differences observed exhibited additive expression, and the rest of the differences demonstrated a dominant expression pattern.
Furthermore, the number of dominant and additive RNA fragments correlated with the degree of heterosis. Initial studies of both seedlings and immature ears, demonstrated correlations between the number of RNA fragments in the over-/under-dominant class and the % relationship between the inbred parents. Five commercial hybrids (PAR19-PAR23) selected for high yield all had high numbers of dominant and additive bands and a lower number of under/over dominant bands.
A new metric that measures the genetic distance between the two parents and the frequency of non-additively expressed RNA's in the over-/under-dominant class was developed. This was defined as the ratio between the sum of the RNA fragment numbers that differ from the hybrid in each of the two parents and the number of RNA's that differ between the two parents, [i.e., (A-F1)+(B-F1)/(A-B)]. High yielding commercial hybrids between distantly related parents and with fewer over-/under-dominant RNA's give a lower ratio close to 1.0 (Tables 2, 3 & 4). Figure 3 illustrates the correlation of gene expression patterns with hybrid yield. Hybrid yield in bu/LCR is given on the X axis, while % of bands in each expression class is given on the Y axis (% of bands different: dotted line; % of additive bands: dashed line; and % of dominant bands: solid line). Table 2. The number of genes exhibiting over-/under-dominant. additive or dominant expression patterns in heterotic and non heterotic hybri Expression data derived from 1 replicate/sample. [Note: Discrepancies between table 2 and Table 3 are likely due to different number of sam used.]
ND= not determined
The number of RNA fragments in the additive class was higher in all heterotic hybπds which include five commercial hybπds (PAR,/PAR2 [+PAR,9], PAR20, PAR21, PAR22, PAR2„ PAR PAR^) The same trend was also found in seedling tissues of selected hybπds analyzed. There was also a strong correlation between the number of dominant RNA bands and % of yield heterosis (Figure 3).
Table 3. Gene expression patterns of hybrids in relation to heterosis Total no. of bands assayed is approximately 14,000 for all genotypes. 1 % is about 140 bands. % of bands different (A-B) : % of bands differentially expressed when comparing the inbred parents of corresponding hybrids. % of bands additive or dominant: % of bands where FI had an additive or dominant expression pattern, respectively. Expression data based on 3 sample replicates.
One way to interpret these data is to assume that for every gene, there is an optimal level of expression. Different mbreds may have subsets of genes that are expressed either below or above the optimum, thus contπbuting to their poor vigor. In hybrids, many genes expressed in parent A, but not in parent B, may be expressed at the same level as in parent A and vice versa. Thus, hybπds will have more genes expressed at an optimum level than either parent A or parent B, and the genes expressed at optimum level in A and B will complement those expressed at sub- or supra-optimal level in the other inbred This aπangement is represented graphically in Figure 4. (Panel A illustrates the dominant class, panels B and C illustrate the additive and over/under dominant classes, respectively «--*: "optimum" level of mRNA expression) Thus, high heterosis is associated with an increase in the dominant and additive classes and a decrease m the over-/under- dominant class. In crosses between related inbreds, the additive class may disappear as more of these genes are likely to be iso-allelic.
Table 4. The number of RNA fragments that differ between parents and between parent and hybrid
The poor correlation between the number of genes in the over-/under- dominant class and the degree of heterosis is surprising. The data suggest that when breeders select for highly heterotic hybrids, they may also be selecting against genes that fall into this class. The logical extension to this argument would be that if derivatives, e.g., of PAR19 are selected or engineered that have fewer or no genes that fall into this class, they will have a higher yield than PAR19 itself.
EXAMPLE 2: PREDICTING HETEROSIS FROM ANALYSIS OF SHARED ADDITIVE BANDS; IDENTIFICATION OF GENES INVOLVED IN HETEROSIS Immature ear mRNA was profiled from 10 hybrids and their respective inbred parents . The genotypes profiled included a number of commercial hybrids and a set from the "PAR27 series, " in which PAR27 was used as a common female with a series of males that differed in percent relationship. Differentially expressed bands among hybrids and inbred parents were categorized according to whether they were additive, non-additive [= over-/under-dominant] or dominant. Analysis of this set of data from profiles of all 10 hybrids showed the following.
First, there was an inverse correlation between heterosis and the number of non-additively expressed sequences. Second, the number of RNA fragments in the additive class was higher in all heterotic hybrids analyzed, which include five commercial hybrids (PAR]9, PAR20, PAR2] , PAR22, PAR23) and PAR24/PAR25 (PAR26 cross). Third, the data also indicated a strong correlation between the number of dominant RNA bands and the degree of heterosis.
Table 5: # of Additive Bands
# of Additive Bands # of Hybrids sharing 26 5 or more
94 4 or more
262 3 or more
612 2 or more
1635 1 or more
Identifying and cloning genes in common to the additive and dominant classes amongst a series of highly heterotic hybrids that share little relationship to each other by pedigree is of value. In comparing the additive class of all 10 hybrids, the additive bands occurring in one or more of the 10 hybrids were considered. The results are shown in Table 5.
The maximum number of hybrids an additive band occurred in was seven out of 10. By analyzing the frequency of bands that occurred in each group mentioned above, in all groups a constant pattern in the bands shared between heterotic hybrids (See Table 6) was detected. Heterosis prediction based on this data gave the following rank: 10>8>2>7>6>1>3>9>4>5. The corresponding hybrids are: PAR23> PAR22> PAR27/PAR25> PAR21> PAR20> PAR19> PAR27/PAR43> PAR24/PAR25> PAR27/PAR37> PAR27/PAR44. Comparing with the actual yield heterosis data in Table 6, the ranking is very close to the yield. Table 6: Heterosis/ Corresponding Hybrid Information
Hybrid # Corresponding Heterosis Frequency Frequency Frequency Frequency hybrid (% of Fl) (26 bands) (94 bands) (262 bands) (612 bands)
1 PAR-7/PAR,4 59 4 15 52 103 167
(PAR,,)
2 PAR,,/PAR,- 544 20 57 137 212
3 PAR27/PAR4, 48 5 13 36 63 94
4 PAR27/PAR„ 25 1 6 8 12 15
5 PAR,,/PAR44 12 8 1 2 4 6
6 PAR , >50 15 48 107 214
7 PAR,, >50 17 52 124 252
8 PAR-, >50 21 62 142 253
10 PAR,, Commercial hybπd 24 67 139 253
The expression patterns of the two SS x SS hybrids, PAR46/PAR48 and PAR46/PAR47, were also informative. The pedigree relationship of the two hybrids are similar (23% and 27%, respectively ); however, the heterosis levels are different significantly, 2.3% for the former and 36.6% for the latter. The difference in the number of additive bands is striking between the two (1 and 60). EXAMPLE 3: ANALYSIS OF DOMINANT GENE EXPRESSION CLASS
Additional observation from further data analysis was that there is a difference in the number of dominant bands contributed by the male vs. female parent, i.e., whether the expression level in FI is the same as that of male or female parent. The number of dominant bands contributed by the male parent are consistently higher across all hybrids analyzed, regardless of the degree of heterosis. However, there is a better correlation between the yield heterosis and the number of the dominant bands contributed by female parents than male parents, especially with the PAR27 series.
When the dominant bands were grouped according to whether they are up- or down-regulated in the hybrid, that is, whether the hybrid is the same as either the higher or lower parent, FI tends to have an expression level closer to the parents with higher expression. This is especially true with the RNA bands that are similar between an FI hybrid and its male parent.
The consistent association of higher numbers of dominant and additive RNA bands with heterotic hybrids, regardless of genetic backgrounds or developmental stages, and- the tendency of up-regulated gene expression of dominant bands in hybrids, suggested that genes in hybrids are in a more active phase than in inbreds. Secondly, more genes are in such an active condition in heterotic hybrids than in poor hybrids. Thus, genes are mostly silenced or inactivated when their regulatory elements are in a homozygous condition, e.g., in inbreds, but re-activated when in a heterozygous condition. Hybrids derived from two inbreds that have optimal complementation to each other to give rise to an heterozygosity condition for most of these regulatory elements had a maximal number of genes "re-activated" and were therefore, heterotic. Crosses of closely related inbreds or inbred lines that did not have such "optimal complementation" had fewer genes re-activated and produced low heterotic hybrids.
Table 7. Difference in the numbers of dominant bands contributed by male vs. female parent.
Genotypes % REL Heterosis Total No. Total No No. of No. of Ratio of
(% of Fl) of bands of domnt do nt domnt male to different bands bands by bands by female
(A-B) male female parent parent
PAR,7-PAR,4 0 08 59 4 1309 605 346 259 1 34 1
PAR,7-PAR,5 0 1 1 54 4 1404 646 419 227 1 85 1
PAR,7-PAR4, 71 48 4 952 462 299 163 1 83 1
PAR,7-PAR,7 86 37 4 505 241 133 108 1 23 1
PAR^-PAR^ 45 18 6 359 180 134 46 2 91 1
PAR,,-PAR,,j 0 04 com 1487 739 415 324 1 28 1
PAR-0-PAR„ 0 04 com 1565 751 417 333 1 25 1
PAR,,-PAR„ 0 01 com 1403 650 370 280 1 32 1
PAR,4-PAR,5 0 21 PAR,, 1484 681 444 237 1 87 1
PAR,4-PAR,5 0 03 com 1297 568 364 204 1 78 1
Another coπelation (statistical association) from this data set is that there is a difference in the number of dominant bands contributed by the male vs. female parent, i.e., the expression level in FI is the same as that of male or female parent. Figure % illustrates parental effects on gene expression of heterotic and non-heterotic hybrids. Total number of dominant bands were calculated for each hybrid as 100%. (Fl=male or female: dominant bands where FI hybrid has equal level of expression as the male or female parent, respectively). The number of dominant bands contributed by the male parent are constantly - higher across all hybrids analyzed, regardless of whether heterotic or non-heterotic (Table 7).
Also, there is a better correlation between the yield heterosis and the number of the dominant bands by female parents than male parents, especially with the PAR27 series (Table 7). For example, the least heterotic hybrid PAR1/PAR17. a sib cross, had 96% of male dominant bands and 4% female dominant bands. Whereas the hybrid exhibiting the highest degree of heterosis PAR1/PAR2 had 60% and 40% male and female dominant bands, respectively. When these dominant bands were grouped according to whether they are up- or down-regulated in the FI, that is, where the FI is the same as either the higher or lower parent, FI tends to have an expression level closer to the parents with higher expression. This is especially true with the RNA bands dominant by the male parents (Table 8).
Table 8. Number of RNA bands where FI is up or down regulated.
Male Parent Female Parent
Hybπds Total Up Regulated Down Regulated Ratio Total (FI Up Regulated Down Regulated Ratio
(Fl = (F = = higher (FI = lower (up down) = female) (FI = higher (FI = lower (up down) male) parent) parent) parent) parent) P PAARR„,7//PPAARR„,4 3 34444 2 25588 8 866 3 1 259 157 102 1 5 1
PAR,7/PAR,, 419 304 1 15 2 6 1 227 154 73 2 1 1
PAR27/PAR4, 299 206 93 2 2 1 163 96 67 1 4 I
PAR27/PAR,7 133 97 36 2 7 1 108 51 57 0 9 1
PAR-,/PAR„ 134 106 28 3.8 1 46 18 28 0 6 1 P PAARR„,,//PPAARR„,,,, 4 41155 2 27744 1 14411 1 9 1 324 152 172 0 9 1
PAR20/PAR„ 417 285 132 2 6 1 333 189 144 1 3 1
PAR3,/PARj, 370 251 119 2 1 1 280 195 85 2 3 1
PAR-VPAR,, 444 270 174 1 6 1 237 165 72 2 3 1
PAR14/PAR„ 363 196 167 1 2 1 204 144 60 2 4 1 EXAMPLE 4: GENES SPECIFICALLY EXPRESSED IN HIGH YIELDING COMMERCIAL HYBRIDS
Most of the analyses so far with the RNA profile data described in Example 1 are based on the expression patterns of FI hybrids relative to their inbred parents, such as additive vs. non additive classifications and the differences of these categoπes between heterotic and non-heterotic hybπds. While the results so far were informative, another way of analyzing this data set by comparing the levels of RNA expression of poor hybπds with heterotic hybπds without any involvement of their parents. In compaπng all 10 hybπds, which include 3 breeding crosses and 7 commercial hybrids, a list of bands that have similar expression level among heterotic hybπds but different from the non-heterotic hybπds
(breeding crosses) was determined.
EXAMPLE 5: EXPRESSION PROFILING USING DIFFERENT TISSUES FROM HYBRIDS AND PARENTS
RNA profiling data from hybπd sets (hybπds and their respective parents) were obtained in maize. Five other sets utilized kernel tissue at 13 days after pollination
("DAP"). A total of 14 hybπd sets for the immature ear (V19), five for the kernel (R2) and three for the seedling tissue (V3) were profiled.
For immature ear tissue, the 14 hybrid sets analyzed included seven from the PAR27 series, which covers a spectrum of heterosis levels ranging from commercial hybπds to low heterotic hybrids of sibling crosses; four commercial hybπds from diversified genetic backgrounds other than PAR27 series and three crosses between inbreds of the same heterotic group, typical of those that would be useful for breeding new mbreds.
The five hybπd sets where kernel tissue was analyzed and the three hybπd sets from the PAR27 series where seedling tissue was analyzed were from the PAR27 series Profiling data of all these hybrids from all three tissues analyzed gave similar expression patterns. However, the immature ear tissue was more informative than seedling tissue and less complicated than the kernel tissue, which is compounded with other effects due to pollen Table 9. RNA bands identified (34) that are differentially expressed in heterotic vs. non- heterotic hybrids (the numbers are the N-fold differences in the expression of each hybrids from PARin).
PAR,, vs. PAR,, vs. PAR,, vs. PAR,, vs. PAR,, vs. PAR,, vs. PAR^/PAR^ PAR,, vs. PAR,, vs PART,, vs.
Band ID PAR12 PAR„ P AR, PAR,5 PAR„ PAR,0 PARjVPAR^ PAR--/PAR-, (PARα cross) PAR2,/PAR43 P R^/PAR^ PAR-,/
PAR,, d010-165 5 0 0 0 0 0 -2 66 -2 19 -6 49 -4 31 dOvO-123 2 0 0 -2 66 -2 06 0 0 3 54 7 09 5 14 d0v0-172 4 7 8 0 0 0 0 8 91 18 48 18 44 29 45 dOvO-104 4 0 0 0 0 0 0 2 36 2 55 2 45 gOmO-359 8 0 0 0 0 0 0 4 97 2 82 2 09 gln0-389 3 0 0 0 2 69 -4 03 0 -2 73 -2 57 -2 98 h0c0-173 1 0 0 0 0 0 0 2 43 2 85 2 16 hOcO-285 4 2 07 2 64 0 2 04 0 0 -2 47 -2 1 -2 88 hOrO-131 4 0 0 0 0 0 3 85 4 83 4 81 5 16 ιOaO-45 6 0 0 0 0 0 0 -2 49 -2 26 -2 53 ιOaO-237 8 0 0 0 0 0 0 2 59 3 3 2 81 ιOaO-242 8 0 0 0 0 0 0 -2 84 -4 76 -2 87 ιOaO-252 1 0 0 0 -2 74 2 72 0 69 21 29 62 69 23 ιOcO-95 6 0 0 -2 51 0 0 0 -3 02 -2 2 -2 76 ιOcO-203 4 0 0 0 0 0 0 -2 26 -2 87 -2 04 ιOcO-312 1 0 0 0 0 -8 74 0 6 1 1 7 1 1 7 9 ιOmO-271 5 0 0 0 0 0 0 -2 92 2 39 2 63
10n0-140 4 0 0 0 0 -2 03 0 2 55 2 1 2 45
10n0-210 5 -2 13 0 -2 21 0 0 12 09 13 43 8 82 16 27 mla0-89 3 0 0 0 0 0 0 3 41 3 74 2 4 mlaO-239 6 -2 54 0 0 0 0 0 -2 8 -4 05 -4 94 mlaO-241 6 0 0 0 0 0 0 -4 -3 14 -2 64 mla0-425 1 0 0 -2 34 0 0 0 4 59 6 64 6 1 rOkO-190 7 0 0 0 0 2 45 0 -2 85 -4 4 -9 01 w9c0-128 3 0 0 4 54 0 0 0 2 53 3 63 3 14 wOcO-230 2 0 0 6 83 0 0 0 -2 72 -3 14 -3 13 wOcO-267 2 -15 18 0 3 09 0 -5 12 2 62 6 62 6 55 10 73 w0c0-381 3 0 0 0 0 0 0 17 45 39 08 44 17 wOhO-251 2 0 0 0 0 0 0 2 64 2 56 2 8 wOhO-406 6 0 0 0 0 0 0 7 34 4 21 7 PAR,, vs. PAR,, vs. PAR,, vs. PAR,, vs. PAR,, vs. PAR,, vs. PAR27/PAR25 PAR,, vs. PAR,, vs PAR,, vs. Band ID PAR12 PAR„ PAR,./PAR,5 PAR„/PAR,„ PAR-j PAR,, (PAR2, cross) PAR27/PAR4, PAR27/PAR44 PAR27/
PAR„ w0ι0-154 4 0 0 -2 04 0 0 0 3 73 3 65 3 8 wOιO-265 3 0 0 0 0 0 0 -2 51 -273 -2 28 yOιO-118 1 0 0 0 0 -2 15 0 4 19 6 57 6 65 y0ι0-254 4 0 0 0 0 0 4 21 2 6 6 95 5 32
sources, such as xenia, maternal effects, etc. Profile analysis of all the samples consistently showed similar correlations between profile information and heterosis to those described in Examples 1, 2 and 3. Dominant bands from all 14 hybrids shared by number of hybrids is presented in Figure 2. The number of dominant bands shared by one or more hybrids is normalized to 100%. The data show that the dominant bands shared by two or more hybrids range from 60- 80%; bands shared by three or more hybrids is about 40-50% and so on. Although the total number of dominant bands was important to make a heterotic hybrid, the dominant bands shared by a higher number of hybrids may not necessarily contribute to the heterosis expression.
Since seedlings show the same trend as immature ears, albeit with different genes involved, it is possible to select at the seedling stage individual hybrid combinations that express the highest number of genes with a dominant expression pattern and that have fewest genes in the over-/under-dominant class. Thus, much larger numbers of F2 top- crosses are screened using this procedure as a first cut, than could be screened by multi- location yield tests alone.
In addition to the analyses above, another way of analyzing the profile data can be used. In this approach, the levels of RNA expression of poor hybrids and heterotic hybrids are compared without any involvement of their parents. This approach examines whether the absolute level of expression of a subset of genes are important for heterosis, in addition to the additive vs. non-additive expression patterns we already found. In the dominantly expressed bands, the FI hybrids tend to have the same expression levels as the higher parent, i.e. showing overall an up-regulation of gene expression (Table 10). In comparing all hybrids with PAR19, 34 bands that have a similar expression level among heterotic hybrids but different from the non-heterotic hybrid were identified (Table 9; the last three columns are non-heterotic hybrids). For these 34 bands, the 3 poor hybrids show either higher or lower expression than PAR19 whereas all other hybrids, which are heterotic, show no or little differences in the expression relative to PARI9.
Table 10. Predominance of up-regulated bands in the hybrids vs. their parents.
Dominant Bands Genotype Total Up- Dn-
PAR7- 656 472 184
PAR20- 672 447 225
PAR30- 629 446 183
PAR24- 670 435 235
PAR,7- 621 431 190
PAR^- 711 422 289
PAR34- 558 342 216
PAR27- 588 319 269
PAR46- 441 313 128
PAR7- 549 311 238
PAR9 " 7- 459 304 155
PAR 6- 346 249 97
PAR27- 269 175 94
PAR97- 191 137 54
EXAMPLE 6: CORRELATIONS TO MALE VERSUS FEMALE PARENTS
As indicated previously, a preponderance of male dominant bands was observed when immature ear mRNA was profiled from hybrids and their respective inbred parents (Figure 5). Selected male dominant bands were screened for allelic sequence polymorphism between inbred parents such that male and female alleles were identified.
Several bands exhibited an allelic polymorphism between the two parental alleles, and these were further tested for mono- or bi-allelic expression in the FI hybrids. PCR primers were designed based on the sequence information and used to amplify cDNAs derived from mRNAs of FI hybrids. More than 20 cDNA clones derived from FI mRNA derived from a single locus were randomly picked and sequenced. All cDNAs expressed in the FI were identical to the allele expressed in the male parent and none were identical to that expressed in the female parent. These results are illustrated in Figure 6a-c which show an allelic 51 expression test of a male dominant band (wOhO) cloned from CuraGen. (A) schematic representation of polymorphic amplification products; B) sequences of 9 random cDNAs from 50% PAR, + 50% PAR2 mRNA used as a control for allelic discrimination in PCR cloning; C) sequences of 10 random cDNAs from PAR,/PAR2 mRNA are all the same as PAR2 allele). This result is consistent with expression of only the male-derived allele and silencing of the female-derived allele. To insure that preferential amplification did not explain the differential amplification results, equal amounts of mRNA from each parent genotype was mixed and amplified by PCR. Of nine cDNA clones sequenced from the control reaction, five were from the male parental allele, and four were from the female parental allele, demonstrating that no discrimination between the alleles occurred during amplification.
Accordingly, the disclosures and descriptions herein are intended to be illustrative, but not limiting, of the scope of the invention which is set forth in the following claims. One of skill will recognize many modifications which fall within the scope of the following claims. For example, all of the methods and compositions herein may be used in different combinations to achieve results selected by one of skill. All publications and patent applications cited herein are incorporated by reference in their entirety for all purposes, as if each were specifically indicated to be incorporated by reference.

Claims

WHAT IS CLAIMED IS. 1. A method of screening for heterosis in plants, compπsmg: (I) profiling expression of a first representative sample of first expression products from a first progeny plant to quantify the expression products produced in the first progeny plant, wherein the number of first expression products produced m the first progeny plant is - correlated with a measure of heterosis in the first progeny plant; or, (n) profiling expression of a second representative sample of second expression products from the first progeny plant to quantify or identify the dominant expression products in the second representative sample, wherein the number of dominant expression products is correlated with a measure of heterosis in the progeny plant.
2. The method of claim 1, further compπsmg' selecting the progeny plant profiled in (l) or (n), based upon the number of first expression products in the first representative sample, or based upon the number of second expression products in the second representative sample that exhibit a dominant expression pattern.
3. A plant selected by the method of claim 2.
4. The method of claim 1, wherein the first and second expression products are independently selected from: mRNAs and proteins.
5. The method of claim 1, wherein the first or second representative sample corresponds to between about 1,000 and about 20,000 gene products
6. The method of claim 1 , wherein expression of at least about 50% of the first or second expression products produced in a selected tissue are detected.
7. The method of claim 1, wherein expression is profiled in step (I) or step (n) using one or more technique selected from: hybπdization of expressed or amplified nucleic acids to a nucleic acid array, hybπdization to a protein array, hybπdization to an antibody aπay, subtractive hybπdization, and differential display
8. The method of claim 1, further compπsing. selecting the first progeny plant for one or more characteristics selected from: a selected number of dominant expression products, a selected ratio of dominant expression products to total expression products, expression of a dominant expression product exhibiting an allelic sequence polymorphism, a desired number of over- or under-dominant expression products, a selected ratio of over- or under-dominant expression products to total expression products, a selected number of additive expression products, and a selected ratio of additive expression products to total expression products
9. The method of claim 1, further comprising- identifying which expression products from the first or second representative sample show a dominant, additive, under-dommant, or over-dom ant expression pattern for at least a portion of the representative sample.
10. The method of claim 1, further comprising: selecting the first progeny plant to maximize the number of dominant expression products or to maximize the number of additive expression products, or to express a dominant expression product exhibiting an allelic sequence polymorphism, or to minimize the number of over- or under-dominant expression products.
11. The method of claim 1, further comprising: cloning at least one nucleic acid encoding an expression product selected from: an additive gene product, a dominant gene product, which dominant gene product optionally has an allelic sequence polymorphism, an over-dominant gene product, an under-dommant gene product, and the product of a transgene derived from the first or second parental plant or the first progeny plant.
12. The method of claim 11, further comprising transducing the at least one nucleic acid into a target plant, resulting in an increase in the number of additive or dominant gene products expressed in the target plant.
13. The method of claim 1, further comprising: crossing a first parent plant with a second parent plant to produce the first progeny plant.
14. The method of claim 13, further comprising: profiling parental expression products from either the first or second parent plant.
15. The method of claim 13, further comprising: profiling expression of parental representative samples of gene products from the first and second parent plant; and, comparing the resulting parental expression profiles of the first and second parent plants with an expression profile of the first progeny plant.
16. The method of claim 13, wherein the first parent plant is a female plant and the second plant is a male plant, the method further comprising: crossing at least a third male plant to the first female plant to produce at least a second progeny plant; comparing an expression profile of the first progeny plant and an expression profile of the second progeny plant to an expression profile of the first female plant; and, selecting the first or second progeny plant based upon similarity to the expression profile of the first female plant.
17. The method of claim 13, further comprising: identifying genes which are silenced in the first parent plant, the second parent plant, or the first progeny plant.
18. The method of claim 13, further compπsing: cloning a nucleic acid encoded by a gene silenced in the first parent plant, the second parent plant or in the first progeny plant.
19. The method of claim 13, further compπsing: introducing a heterologous - nucleic acid into the first parent plant, the second parent plant, the first progeny plant, or a subsequent progeny plant deπved from one or more of: the first parent plant, the second parent plant, or the first progeny plant, which heterologous nucleic acid results in increased expression of an expression product from a silenced gene
20. The method of claim 13, further compπsing: determining a ratio between the sum of expressed gene pioducts that differ from the progeny plant in each of the first and second parent plants and the number of expressed gene products that differ between the first and the second parent plant
21. The method of claim 13, further compπsing: crossing 1 or more additional plants with the first or second parent plant to produce at least one additional progeny plant.
22. The method of claim 13, further compπsing: crossing 1 or more additional plants with the first or second parent plant to produce one or more additional progeny plant; profiling expression of a representative sample of gene products from the one or more additional progeny plant; and, comparing the resulting expression profile of the one or more additional progeny plant with an expression profile of the first progeny plant
23. The method of claim 22, further compπsing: selecting a heterotic progeny plant from a group of progeny plants compπsing the first progeny plant and the one or more additional progeny plant.
24. The method of claim 23, wherein the heterotic progeny plant is selected based upon one or more selectable property selected from: an elevated number of expressed RNAs relative to one or more parental plant; an elevated number of expressed RNAs relative to other progeny plants in the group of progeny plants; an elevated number of RNAs showing a dominant expression pattern relative to one or more parental plant; an elevated number of - gene products showing a dominant expression pattern relative to other progeny plants in the group of progeny plants; an RNA showing an allelic sequence polymorphism relative to one or more parental plant, an RNA showing an allelic sequence polymorphism relative to other progeny plants in the group of progeny plants, a decreased number of gene products showing an over or underdominant gene expression pattern as compared to one or more parental plant; and, a decreased number of gene products showing an over or underdominant expression pattern as compared to other progeny plants in the group of progeny plants.
25. The method of claim 13, wherein the first parent plant, second parent plant, and first progeny plant are independently selected from: an inbred plant, and a hybrid plant.
26. The method of claim 13, wherein the first parent plant is a first inbred plant, the second parent plant is a second inbred plant and the progeny plant is a hybrid plant.
27. The method of claim 13, wherein the first or second parent plant is an inbred or hybrid plant, and the progeny plant is a hybrid plant, the method further comprising crossing a plurality of first additional plants of the same strain as the first parent plant with a plurality second additional plants of the same strain as the second parent plant, to produce a plurality of progeny hybrid plants.
28. The method of claim 27, further comprising topcrossing at least one of the plurality of progeny hybrid plants with a plurality of inbred plants to provide a plurality of topcross plants.
29. The method of claim 28, further compπsing topcrossing the topcross plants to an mbred plant to produce a topcross progeny plant, and, optionally, profiling expression of a representative sample of RNA from the topcross plant or from the topcross progeny plant.
30. The method of claim 28, further compπsing: selfing a test plant selected from: the first parent plant, the second parent plant, the first progeny plant, one of the plurality of progeny hybπd plants, one of the plurality of topcross plants, and one of the plurality of topcross progeny plants; or crossing one or more test plants selected from- the first parent plant, the second parent plant, the first progeny plant, one of the plurality of progeny hybπd plants, one of the plurality of topcross plants, and one of the plurality of topcross progeny plants.
31. The method of claim 30, further compπsing: profiling expression of the test plant.
32. The method of claim 30, further compπsing: profiling expression of an immature tissue from the test plant.
33. The method of claim 28, the method further compπsing: profiling expression of a representative number of expression products from one or more of the plurality of hybrid progeny plants, or progeny thereof; and additionally performing at least one of: (I) determining the number of expression products in the representative sample from the plurality of hybrid progeny plants, or progeny thereof, wherein the number of expression products in the plurality of hybπd progeny plants, or progeny thereof, is correlated with a measure of heterosis in the plurality of hybrid progeny plants, or progeny thereof, (n) determining the number of expression products in the representative sample of expression products from the plurality of hybrid progeny plants, or progeny thereof, wherein the number of expression products exhibiting a dominant expression pattern in the plurality of hybrid progeny plants, or progeny thereof is correlated with a measure of heterosis in the plurality of hybrid progeny plants, or progeny thereof; and, (iii) selecting the plurality of hybrid progeny plants, or progeny thereof for plants which display a selected number of expression products, or a selected number of dominant expression products, thereby selecting for an increase in a measure of heterosis.
34. The method of claim 13, wherein the first and second parent plants are monocots.
35. The method of claim 13, wherein the first and second parent plant are selected from the families Gramineae, Compositae, and Leguminosae.
36. The method of claim 13, wherein the first and second parent plant are selected from: Zea mays, rice, soybean, sorghum, wheat, oats, barley, millet, sunflower, and canola.
37. The method of claim 13, further comprising selecting the first and second parent plant to produce the first progeny plant with a selected number of expression products which are dominant, over-dominant, under-dominant or additive.
38. The method of claim 37, wherein the parents are selected to produce the first progeny plant by selecting for complementary expression of dominant or additive expression products between the parents.
39. The method of claim 1, further comprising: (iii) comparing a set of first expression products in the first progeny plant to a set of second plant expression products from a second plant; or, (iv) comparing a set of expression products exhibiting a dominant expression pattern in the first progeny plant to a set of expression products exhibiting a dominant expression pattern in a second plant.
40. The method of claim 39, wherein step (in) or step (IV) is performed using a computer
41. The method of claim 39, wherein step (iv) or step (v) is performed using a computer, wherein the second number or expressed gene products or the second number of - gene products exhibiting a dominant expression pattern is present in a database in the computer
42. The method of claim 1, wherein the steps of profiling expression are performed in an integrated system compπsing a microprocessor with software for determining one or more of: how many genes are expressed; whether expressed genes are dominant, whether expressed genes are additive; whether expressed genes are over-dominant, and, whether expressed genes are under-dominant
43. The method of claim 1, further compπsing mput g a resulting expression profile for the first progeny plant into a database of expression profiles
44. The method of claim 43, wherein the database is m an integrated system compπsing a computer
45. A database produced by the method of claim 43
46. The database of claim 45, wherein the database is present in a computer
47. The computer database of claim 46, wherein the database compπses expression product profiles of a representative sample of expression products for hybπd progeny plants resulting from at least 10 separate inbred plant crosses
48. The method of claim 43, further compπsing selecting an expression profile from the database, which profile provides a unique subset of expression products.
49. The method of claim 48, further comprising: cloning a nucleic acid which expresses at least one expression product in the unique subset of expression products; or, cloning a nucleic acid which expresses at least one expression product in the unique subset of expression products and transducing the nucleic acid into a heterologous plant; or, - crossing a first selected plant which expresses the unique subset of expression products with a second selected plant which does not express the unique subset of expression products.
50. The method of claim 1, further comprising: selfing the first progeny plant.
51. The method of claim 1, further comprising: selfing the first progeny plant and detecting silencing of dominant expression products in subsequent progeny plants which are derived from selfing the first progeny plant.
52. The method of claim 51, further comprising: cloning a silenced nucleic acid encoding a dominant expression product.
53. The method of claim 51, further comprising: introducing a heterologous nucleic acid that results in expression of dominant expression products from silenced genes.
54. The method of claim 53, wherein the heterologous nucleic acid encodes one or more of: a transcription factor which activates a promoter from a silenced gene; a nucleic acid encoded by the silenced gene under the control of a heterologous promoter; and, a nucleic acid homologous to the silenced gene with at least one [region of difference] with the silenced gene, which homologous nucleic acid can recombine with the silenced gene to produce a modified gene.
55. The method of claim 1, further comprising: testing the first progeny plant or a subsequent progeny plant thereof for a desired trait.
56. The method of claim 1, further compπsing testing the first progeny plant, or a subsequent progeny plant thereof, for a desired phenotypic trait, compaπng the phenotypic trait between the first progeny plant, or the subsequent progeny plant, to a selected hybπd plant; compaπng an expression profile of the selected hybπd plant to an expression profile of the first progeny plant, or the subsequent progeny plant; and, cloning at least one nucleic acid which is differentially expressed between the selected hybπd plant and the first progeny plant, or the subsequent progeny plant
57. The method of claim 56, further compπsing transducing the at least one nucleic acid into a selected plant to produce a transgenic plant.
58. The method of claim 1, wherein the first and second representative samples are from an immature tissue of first progeny plant.
59. The method of claim 58, wherein the immature tissue is an immature ear of the plant, or a seedling plant.
60. A method of identifying plant crosses with an increase m probability for heterosis in progeny plants, compπsing: (I) compaπng expression profiles for a plurality of plants; and (n) determining, by pair-wise compaπsons of the expression profiles, which crosses will produce at least one of the following. (a) progeny with a selected or optimal number of expression products; or, (b) progeny with a selected number or type of expression products that display a dominant, additive, overdommant or underdominant expression pattern.
61. The method of claim 60, further comprising making identified plant crosses to produce progeny plants.
62. The method of claim 60, further compπsing making identified plant crosses to produce progeny plants, which progeny plants are tested for one or more desired trait.
63. The method of claim 60, wherein crosses are identified which maximize - the number of expression products in potential progeny, or which maximize the number of dominant expression products in potential progeny, or which maximize the number of additive expression products m potential progeny, or which minimize the number of over- dominant expression products m potential progeny, or which minimize the number of under- dominant expression products in potential progeny.
64. The method of claim 60, wherein: the plants are inbred plants, hybrid plants, or transgenic plants; and, the plants are selected from: plants in the families Gramineae, Compositae, and Leguminosae; or, the plants are selected from: Zea mays, rice, soybean, sorghum, wheat, oats, barley, millet, sunflower, and canola
65. The method of claim 60, wherein the expression profiles are compiled m a database.
66. The method of claim 60, wherein a matrix of possible pair- wise expression profile combinations for the plants is generated.
67. The method of claim 60, wherein the expression profiles are compiled in a database in a computer and a matrix of possible pair-wise expression profile combinations for the plants is considered using an integrated system compπsing a computer.
68. The method of claim 60, further compπsing: selecting a subset of potential crosses from all of the possible pair-wise compaπsons which exhibit a maximal number of expression profile differences
69. The method of claim 60, further comprising: selecting a subset of potential crosses from all of the possible pair-wise comparisons which exhibit a maximal number of expression profile differences, wherein at least a plurality of the possible pair-wise comparisons are for plants from different heterotic groups.
70. The method of claim 60, wherein the pair-wise comparisons are considered to identify crosses from the same heterotic group.
71. The method of claim 60, further comprising: (iii) identifying crosses where: the sum of: (a) expression products produced in a first plant from a first heterotic group (A,) which are not expressed in a second plant from the first heterotic group (A) to which the first plant is crossed (Ak), and which are not expressed in a selected third plant from a second heterotic group (B); plus (b) the expression products produced in Ak which are not produced A. and which are not produced in B ; is optimized.
72. The method of claim 71, further comprising making a cross identified in (iii).
73. The method of claim 71, wherein optimization is made by: determining all possible pair-wise combinations from the first heterotic group and identifying the cross which results in the largest sum of expression products; or determining all possible pair-wise combinations from the first heterotic group and identifying crosses which result in a hybrid progeny (A, x A,) with a maximal number of differences as compared to B; or, determining all possible pair-wise combinations from the first heterotic group and identifying crosses which result in the hybrid progeny (A, x A.) having a greater number of differences with B than the number of differences between B and A, or B and A..
74. The method of claim 73, further comprising selecting self- or back- crossed progeny derived from the A, x A. hybrid that: retain a set of expression products defined by the sum of expression products expressed in A, (but not A. or B) and A. (but not A, or B); or which show a larger number of expression products expressed in a topcross with B than does either A, or A., when topcrossed with B.
75. A method of identifying a source of a test plant, comprising: profiling expression of a representative sample of expression products from the test plant; and, comparing the resulting test expression profile to a database of known expression profiles for plants from known inbred or hybrid strains.
76. The method of claim 75, wherein the expression profile is for a selected tissue and the database of expression profiles comprises expression profiles for the same tissue from the known inbred or hybrid strains.
77. The method of claim 75, wherein the database of expression profiles is used to provide a matrix of pair-wise comparisons for potential progeny from the expression profiles in the database, which matrix of pair-wise comparisons is compared to the test expression profile .
78. The method of claim 75, wherein the source identified is a sub-portion of the total expression profile, which subportion corresponds to a unique marker for a specific parental strain.
EP00904457A 1999-01-21 2000-01-19 Molecular profiling for heterosis selection Withdrawn EP1143787A2 (en)

Applications Claiming Priority (5)

Application Number Priority Date Filing Date Title
US11661799P 1999-01-21 1999-01-21
US116617P 1999-01-21
US16636899P 1999-11-17 1999-11-17
US166368P 1999-11-17
PCT/US2000/001422 WO2000042838A2 (en) 1999-01-21 2000-01-19 Molecular profiling for heterosis selection

Publications (1)

Publication Number Publication Date
EP1143787A2 true EP1143787A2 (en) 2001-10-17

Family

ID=26814421

Family Applications (1)

Application Number Title Priority Date Filing Date
EP00904457A Withdrawn EP1143787A2 (en) 1999-01-21 2000-01-19 Molecular profiling for heterosis selection

Country Status (6)

Country Link
EP (1) EP1143787A2 (en)
AU (1) AU2621300A (en)
CA (1) CA2358509A1 (en)
HU (1) HUP0200319A3 (en)
MX (1) MXPA01007325A (en)
WO (1) WO2000042838A2 (en)

Families Citing this family (11)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
AU773329B2 (en) * 1999-05-14 2004-05-20 Proteomedica Ab Materials and methods relating to disease diagnosis
AU2002364153A1 (en) * 2001-12-11 2003-06-23 Lynx Therapeutics, Inc. Genetic analysis of gene expression in heterosis
CA2555965A1 (en) * 2004-02-03 2005-08-18 Hybrid Biosciences Pty Ltd Method of identifying genes which promote hybrid vigour and hybrid debility and uses thereof
WO2007012138A1 (en) * 2005-07-29 2007-02-01 Hybrid Biosciences Pty Ltd Identification of genes and their products which promote hybrid vigour or hybrid debility and uses thereof
GB2436564A (en) * 2006-03-31 2007-10-03 Plant Bioscience Ltd Prediction of heterosis and other traits by transcriptome analysis
EP2005193A1 (en) * 2006-04-06 2008-12-24 Monsanto Technology, LLC Method of predicting a trait of interest
US20080083042A1 (en) * 2006-08-14 2008-04-03 David Butruille Maize polymorphisms and methods of genotyping
ES2381457T3 (en) 2007-12-28 2012-05-28 Pioneer Hi-Bred International Inc. Use of a structural variation to analyze genomic differences for the prediction of heterosis
BRPI0920872B1 (en) * 2008-10-06 2018-06-19 Yissum Research Development Company Of The Hebrew University Of Jerusalem Ltd. ISOLATED POLYNUCLEOTIDE THAT ENCODES A MUTANT SFT PROTEIN AND METHOD TO PRODUCE A HYBRID PLANT
US9842252B2 (en) 2009-05-29 2017-12-12 Monsanto Technology Llc Systems and methods for use in characterizing agricultural products
GB201110888D0 (en) * 2011-06-28 2011-08-10 Vib Vzw Means and methods for the determination of prediction models associated with a phenotype

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO1984004758A1 (en) * 1983-05-26 1984-12-06 Plant Resources Inst Process for genetic mapping and cross-breeding thereon for plants

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
See references of WO0042838A3 *

Also Published As

Publication number Publication date
WO2000042838A3 (en) 2001-03-29
CA2358509A1 (en) 2000-07-27
MXPA01007325A (en) 2002-06-04
HUP0200319A3 (en) 2003-12-29
WO2000042838A2 (en) 2000-07-27
AU2621300A (en) 2000-08-07
HUP0200319A2 (en) 2002-05-29

Similar Documents

Publication Publication Date Title
US8039686B2 (en) QTL “mapping as-you-go”
Yu et al. A whole‐genome SNP array (RICE 6 K) for genomic breeding in rice
AU2004303836C1 (en) High lysine maize compositions and methods for detection thereof
Barua et al. Identification of RAPD markers linked to a Rhynchosporium secalis resistance locus in barley using near-isogenic lines and bulked segregant analysis
EP2158336A2 (en) Methods for sequence-directed molecular breeding
EP1143787A2 (en) Molecular profiling for heterosis selection
AU2019312799B2 (en) Method for the quality control of seed lots
Lu et al. Genetic basis of maize kernel protein content revealed by high-density bin mapping using recombinant inbred lines
JP2006345855A (en) Method for identification and/or quantification of nucleotide sequence element specific to genetically modified plant on array
US9617605B2 (en) Molecular markers associated with yellow flash in glyphosate tolerant soybeans
Koebner et al. Actual and potential contributions of biotechnology to wheat breeding
US20070192909A1 (en) Methods for screening for gene specific hybridization polymorphisms (GSHPs) and their use in genetic mapping ane marker development
US5332408A (en) Methods and reagents for backcross breeding of plants
CN114015701A (en) A molecular marker for detecting barley grain shrinkage and its application
US20070048768A1 (en) Methods for screening for gene specific hybridization polymorphisms (GSHPs) and their use in genetic mapping and marker development
Chen et al. A genotyping platform assembled with high-throughput DNA extraction, codominant functional markers, and automated CE system to accelerate marker-assisted improvement of rice
KR100981042B1 (en) Complementary recessive genes and screening markers involved in rice hybrid degeneration
Shimizu et al. Development of a KASP marker set for high-throughput genotyping in Japanese barley breeding programs with various end-use purposes
US20050250205A1 (en) Use of associations between at least one nucleic sequence polymorphism of the sh2 gene and at least one seed quality characteristic in plant selection methods
CN119120767A (en) A fl4InDel1 molecular marker associated with maize floury4 genotype and its application
Oh Tagging downy mildew resistance gene (Sdm) and head smut resistance gen (Shs) in sorghum using RFLP and RAPD markers
MXPA06006574A (en) High lysine maize compositions and methods for detection thereof
HK1159681B (en) High lysine maize compositions and methods for detection thereof

Legal Events

Date Code Title Description
PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

17P Request for examination filed

Effective date: 20010810

AK Designated contracting states

Kind code of ref document: A2

Designated state(s): AT BE CH CY DE DK ES FI FR GB GR IE IT LI LU MC NL PT SE

AX Request for extension of the european patent

Free format text: AL;LT;LV;MK;RO;SI

17Q First examination report despatched

Effective date: 20040216

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWN

18D Application deemed to be withdrawn

Effective date: 20040629