EP1786926A2 - Disease risk estimating method using sequence polymorphisms in a specific region of chromosome 19 - Google Patents

Disease risk estimating method using sequence polymorphisms in a specific region of chromosome 19

Info

Publication number
EP1786926A2
EP1786926A2 EP05771208A EP05771208A EP1786926A2 EP 1786926 A2 EP1786926 A2 EP 1786926A2 EP 05771208 A EP05771208 A EP 05771208A EP 05771208 A EP05771208 A EP 05771208A EP 1786926 A2 EP1786926 A2 EP 1786926A2
Authority
EP
European Patent Office
Prior art keywords
sequence
polymorphism
cancer
rai
seq
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Withdrawn
Application number
EP05771208A
Other languages
German (de)
French (fr)
Inventor
Bjørn Andersen NEXØ
Ulla Birgitte Vogel
Anders BØRGLUM
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Det Nationale Forskningscenter For Arbejdsmiljo
Aarhus Universitet
Original Assignee
Det Nationale Forskningscenter For Arbejdsmiljo
Aarhus Universitet
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Det Nationale Forskningscenter For Arbejdsmiljo, Aarhus Universitet filed Critical Det Nationale Forskningscenter For Arbejdsmiljo
Publication of EP1786926A2 publication Critical patent/EP1786926A2/en
Withdrawn legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • C12Q1/6876Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
    • C12Q1/6883Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material
    • C12Q1/6886Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material for cancer
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q2600/00Oligonucleotides characterized by their use
    • C12Q2600/106Pharmacogenomics, i.e. genetic variability in individual responses to drugs and drug metabolism
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q2600/00Oligonucleotides characterized by their use
    • C12Q2600/118Prognosis of disease development
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q2600/00Oligonucleotides characterized by their use
    • C12Q2600/172Haplotypes

Definitions

  • the present invention provides methods and compositions for identifying human subjects with an increased risk of having or developing disease.
  • this invention relates to the identification and characterization of polymorphisms in the human chromosome 19q, the region r located approximately 19q 13.2-3 correlated with increased risk of developing disease, in particular cancer and the responsive ⁇ ness of a subject to various treatments for cancer.
  • DNA polymorphisms provide an efficient way to study the association of genes and diseases by analysis of linkage and linkage disequilibrum. With the sequencing of the human genome a myriad of hitherto unknown genetic polymorphisms among people have been detected. Most common among these are the single nucleotide polymorphisms, also called SNPs, of which several millions are known. Other ex ⁇ amples are variable number of tandem repeat polymorphisms, insertions, deletions and block modifications. Tandem repeats often have multiple different alleles (vari- ants), whereas the other groups of polymorphisms usually just have two alleles.
  • Some of these genetic polymorphisms probably play a direct role in the biology of the individuals, including their risk of developing disease, but the virtue of the major ⁇ ity is that they can serve as markers for the surrounding DNA, and thus serve as leads during as search for a causative gene polymorphism, as substitutes in the evaluation of its role in health and disease, and as substitutes in the evaluation of the genetic constitution of individuals.
  • Linkage arises because large parts of chromosomes are passed unchanged from parents to off ⁇ spring, so that minor regions of a chromosome tend to flow unchanged from one generation to the next and also to be similar in different branches of the same fam ⁇ ily. Linkage is gradually eroded by recombination occurring in the cells of the germ- line, but typically operates over multiple generations and distances of a number of million bases in the DNA.
  • Linkage disequilibrium deals with whole populations and has its origin in the (distant) forefather in whose DNA a new sequence polymorphism arose.
  • the immediate sur ⁇ roundings in the DNA of the forefather will tend to stay with the new allele for many generations.
  • Recombination and changes in the composition of the population will again erode the association, but the new allele and the alleles of any other polymor ⁇ phism nearby will often be partly associated among unrelated humans even today.
  • a crude estimate suggests that alleles of sequence polymorphisms with distances less than 10000 bases in the DNA will have tended to stay together since modern man arose.
  • Linkage disequilibrium is the results of many stochastic events and as such subject to statistical variation occasionally resulting in discontinuities, lack of a monotonic relationship between association and distance and differences between people of different ethnicity. Therefore, it is often advantageous to study more that one se ⁇ quence polymorphism in a given region. This also allows for further definition of the genetic surroundings of the biologically relevant polymorphism by combining the associated alleles of the different markers into a socalled haplotype.
  • genotypes i.e. the combined analysis of both chromosomes at a given sequence polymorphism.
  • the resulting genotypes of a person, analysed for instance on DNA from peripheral blood leukocytes, are inherently very stable over time. Therefore, this type of analysis can be performed any time in the life of a person and will be applicable to this person for his or her entire life.
  • genetic analyses are ideally suited to predict future risks of disease.
  • a variety of investigations suggest that many diseases in part are determined by the genetic constitution of the individual.
  • One group of genes in particular has been as ⁇ sociated with rare genetic predispositions to cancer.
  • DNA repair genes these are the genes involved in maintaining the integrity of a person's DNA, the so-called DNA repair genes.
  • DNA repair genes One set of such genes are the XP genes which participate in nucleotide excision repair, and, when mutated, give rise to a 1000 fold increased risk of getting skin cancer.
  • XPDe ⁇ one allele of the sequence polymorphism called XPDe ⁇ was associated with a moderately increased risk of getting basal cell carci ⁇ noma, the most common form of skin cancer.
  • XPDe ⁇ one allele of the sequence polymorphism called XPDe ⁇ was associated with a moderately increased risk of getting basal cell carci ⁇ noma, the most common form of skin cancer.
  • Later other groups have studied the association between sequence polymorphisms in this and other DNA repair genes and various forms of cancer. Some have reported positive results.
  • the present invention relates in a first aspect to a group of nucleic acid sequences found to be associated with disease, in particular cancer.
  • the invention further re ⁇ lates to transcriptional and translational products of said sequence.
  • An allele in the r region can be identified as correlated with an increased risk of developing disease, in particular cancer, the prognosis of developed disease, in particular cancer, and responsiveness to disease treatment, in particular cancer treatment on the basis of statistical analyses of the incidence of a particular allele in individuals diagnosed with disease, in particular cancer.
  • the invention relates to a method for estimating the disease risk of an individual comprising
  • the estimation of the disease risk of an individual can involve the comparison of the number and/or kind of polymorphic sequences identified with a predetermined dis ⁇ ease risk profile.
  • a profile can be based on statistical data obtained for a rele ⁇ vant reference group of individuals.
  • the disease is a proliferative dis- ease, such as cancer.
  • sequence of the r region is set forth as SEQ ID NO 1 , originating from the clon ⁇ ing of human chromosome 19q published as part of the contig NT_011109 in the database of human sequences established by National Center for Biotechnology Information and located on the internet at http://www.ncbi.nlm.nih.gov/- qenome/quide/human/ .
  • the presence of an allele is determined by determining the nucleic acid sequence of all or part of the region according to standard molecular biology protocols well known in the art as described for example in Sambrook et al. (1989) and as set forth in the Examples provided herein or products of the nucleic acid sequences.
  • the nucleic acid molecules of the present invention represent in a first aspect nucleic acid sequences forming part of the region r corresponding to position 1522-37752 of SEQ ID NO: 1 , and preferably to certain nucleic acid sequences within the gene referred to herein as RAI.
  • RAI nucleic acid sequences forming part of the region r corresponding to position 1522-37752 of SEQ ID NO: 1
  • the RAI gene is in particular associated with human cancer diseases.
  • the invention relates to a method for estimating the disease prognosis of an individual comprising - in a sample from said individual, assessing in the genetic material a sequence polymorphism
  • the estimation of the disease prognosis of an individual can involve the comparison of the number and/or kind of polymorphic sequences identified with a predetermined disease prognosis profile. Such a profile can be based on statistical data obtained for a relevant reference group of individuals.
  • a method of identifying a human subject as having an in ⁇ creased likelihood of responding to a treatment comprising a) correlating the pres- ence of an r region allele genotype with an increased likelihood of responding to treatment; and b) determining the r region allele genotype of the subject, whereby a subject having an r region allele genotype correlated with an increased likelihood of responding to treatment is identified as having an increased likelihood of responding to treatment.
  • the present invention also relates to method for estimating a treatment re ⁇ sponse of an individual suffering from disease to a disease treatment, comprising
  • the estimation of the individual's response to disease treatment can involve the comparison of the number and/or kind of polymorphic sequences identified with a predetermined cancer treatment response profile.
  • a predetermined cancer treatment response profile can be based on statistical data obtained for a relevant reference group of individuals.
  • the disease is a proliferative disease, such as cancer.
  • the invention also comprises primers or probes for use in the invention, as well as kits including these.
  • the primers and/or probes are preferably capable of hybridising to SEQ ID NO:1, or a part thereor, in particularly the regions relevant to this inven ⁇ tion, or a part thereof, under stringent conditions, as well as to a sequence comple- mentary thereto.
  • the invention also relates to cloning vectors and expression vectors containing the nucleic acid molecules of the invention, as well as hosts which have been transformed with such nucleic acid molecules, including cells genetically engi- neered to contain the nucleic acid molecules of the invention, and/or cells geneti ⁇ cally engineered to express the nucleic acid molecules of the invention.
  • the nucleic acids are preferably isolated from the r region and preferably contain one or more sequence polymorphisms as described herein below in more detail.
  • hosts also include transgenic non-human animals (or prog- eny thereof).
  • the present invention is based on the discovery of the correlation with single nucleotide polymorphisms (SNPs), deletion polymorphisms, insertion poly ⁇ morphisms, dinucleotide polymorphisms and/or tandem repeats in the regions and disease.
  • SNPs single nucleotide polymorphisms
  • deletion polymorphisms e.g., deletion polymorphisms
  • insertion poly ⁇ morphisms e.g., tyl-like polymorphisms
  • dinucleotide polymorphisms e.g., tandem repeats in the regions and disease.
  • the term human includes both a human having or suspected of having a disease and an a-symptomatic human who may be tested for predisposition or susceptibility to disease. At each position the human may be homozygous for an allele or the hu- man may be a heterozygote.
  • Fig. 2 shows a pair-wise linkage disequilibrium between all markers in breast cancer controls.
  • Fig. 3 shows an association of single polymorphisms with early breast cancer.
  • Fig. 4 shows an association of sets of neighboring SNPs with breast cancer.
  • Fig. 5 shows a maximal OR for breast cancer for haplotypes formed by two SNPs.
  • Fig. 6 shows a Lazzeroni estimation of the position of the causative variant using young breast cancer.
  • Fig. 7 shows the overall distribution of p-values for sets of markers plotted against the position on the chromosome for basal cell carcinoma and lung cancer.
  • posi ⁇ tion on the abscissa we used the median marker position in a given cluster of mark ⁇ ers.
  • the ordinate values are the negative logarithms to the overall p-values for a difference between cases and controls associated with a given set of markers.
  • Each curve corresponds to a given size of marker sets, i.e. the length of the haplotypes.
  • Fig. 8 shows odds ratio for cancer versus control between the two homozygotes of each SNP in relation to location on chromosome 19.
  • Fig. 9 shows event free survival for patients who are wild-type carriers of ASE-1 (0) or homozygous or heterozygous carriers of the variant allele (1 ).
  • Fig. 11 shows event free survival for women subdivided by ASE-1 genotype.
  • 0 homozygous carrier of the wild-type allele of ASE-1
  • 1 carrier of the variant allele of ASE-1.
  • Fig. 12 shows overall survival for women subdivided by ASE-1 genotype.
  • 0 homozygous carrier of the wild-type allele of ASE-1
  • 1 carrier of the variant allele of ASE-1
  • Fig. 13 shows event free survival for men subdivided by ASE-1 genotype.
  • 0 homozygous carrier of the wild-type allele of ASE-1
  • 1 carrier of the variant allele of ASE-1.
  • Fig. 15 shows Kaplan Meier plot of survival of lung cancer patients in relation to highrisk haplotype
  • Fig. 16 shows Kaplan Meier plot of survival in relation to XPD K751Q among those lung cancer patients homozygous for the high risk haplotype.
  • Fig. 17 shows Kaplan Meier plot of survival in relation to XPD K751Q among those lung cancer patients not homozygous for the highrisk haplotype.
  • the present invention relates to a characterization of a person's present and/or fu ⁇ ture risk of getting certain forms of disease, in particular a proliferative disease, such as cancer.
  • the characterization is based on the analysis of sequence polymor ⁇ phisms in a region of chromosome 19q in the person.
  • a number of polymorphisms in the chromosomal region 19q 13.2-3 have been identi ⁇ fied and characterised. Surprisingly, the sequence polymorphisms with strongest association to disease appeared to be located outside the gene XPD. More specifi ⁇ cally, the sequences were located in a sub-region harboring the gene RAI and im ⁇ mediately downstream of the RAI gene. The nature of the association between hap- lotype and disease was examined together with the p-values associated with the individual haplotypes. The odds ratios for each haplotype of each set of three neighboring SNPs were also determined. The odds ratio for test of homozygotes of individual markers against cancer status was likewise detemined. The most likely location of a single causative gene variant for individuals of the breast cancer cases was evaluated.
  • the region of chromosome 19q is depicted in Figure 1 as it is presently known together with the presently known or suspected genes.
  • the arrows indicate the directions of transcription of the genes.
  • the absolute chromosome positions shown are from the particular build of NCBPs map of chromsome 19, and will proba- bly change with time. The position of markers used throughout the experiments is indicated.
  • the intron-exon structure of the RAI gene is shown together with the posi ⁇ tion of the inter gene region between RAI and XPD genes.
  • the region r stretches from the beginning of, but not including the XPD gene, to ap- proximately the end of ERCC1 and includes the genes RAI, LOC162978, and ASE- 1. More specifically r is bounded by and includes the following two sequences: AGAACCCCCG CCCCTCCACC TCGTCTCAAA and TCCCTCCCCA GA- GACTGCAC CAGCGCAGCC, and is defined by SEQ ID NO: 1.
  • the region r means SEQ ID NO: 1 and complementary se ⁇ quence as well as transcriptional products and translational products thereof.
  • the gene RAI is defined in the claims as including transcribed sequences of the gene plus a 1500 base upstream promoter region. More specifically RAI is bounded by and includes the following sequences: CATAACCACA ATGATGAGCA TGTATTGAGT and ATGTTGTCCA GGCTGGTCTT GAACTCCTGA. In the present context this section of the region relates to SEQ ID NO: 1 bases 7761-22885 and complementary sequence as well as transcriptional products and translational products thereof.
  • one preferred section of the region stretches approximately from the the beginning of, but not including the XPD gene, to approximately the end of the RAI gene.
  • the region means SEQ ID NO: 1 bases 1 to 25550 and complementary sequence as well as transcriptional products and transla ⁇ tional products thereof.
  • one preferred section of the region stretches approximately from the the beginning of, but not including the XPD gene, to approximately within the RAI gene.
  • the region means SEQ ID NO: 1 bases 1 to 15698 and complementary sequence as well as transcriptional products and translational products thereof.
  • one preferred section of the region r stretches ap ⁇ proximately from outside the XPD gene, to approximately within the RAI gene.
  • the region means SEQ ID NO: 1 bases 4528 to 15698 and comple- mentary sequence as well as transcriptional products and translational products thereof.
  • one preferred section of the region r stretches approximately from the beginning of, but not including the XPD gene, into the inter gene region between the RAI and XPD gene.
  • the region means SEQ ID NO: 1 bases 1 to 1510 and complementary sequence as well as transcriptional products and translational products thereof.
  • one preferred section of the region r stretches approxi- mately from the the beginning of, but not including the XPD gene, throughout the inter gene region between the RAI and XPD gene and into the 3' part of the RAI gene.
  • the region means SEQ ID NO: 1 bases 1710 to 8685 and complementary sequence as well as transcriptional products and translational products thereof.
  • one preferred section of the region r stretches ap ⁇ proximately over the central part of the RAI gene.
  • the region means SEQ ID NO: 1 bases 8987 to 12090 and complementary sequence as well as transcriptional products and translational products thereof.
  • one preferred section of the region r stretches approximately over the middle and 5' part of the RAI gene.
  • the region means SEQ ID NO: 1 bases 15898 to 25550 and complementary sequence as well as tran ⁇ scriptional products and translational products thereof.
  • Fragments or parts of the region r as used herein relates to any fragment of at least 5 nucleic acid redues in length, or multiples of 5 nucleic acid residues in length start ⁇ ing from SEQ ID NO: 1 position 1 , 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100.
  • At least 21 such as at least 22, for example at least 23, such as at least 24, for example at least 26, such as at least 27, for ex- ample at least 28, such as at least 29, for example at least 31 , such as at least 32, for example at least 33, such as at least 34, for example at least 36, such as at least 37, for example at least 38, such as at least 39, for example at least 41 , such as at least 42, for example at least 43, such as at least 44, for example at least 46, such as at least 47, for example at least 48, or at least 100 nucleic acid redues in length, or mutiples of 100 nucleic acid residues in length, starting from SEQ ID NO: 1 posi ⁇ tion 1, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2600, 2700, 2800, 2900
  • Multiples are preferably multiples of e.g. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 , 42, 43, 44, 45, 46, 47, 48, 49 and 50.
  • the length of said fragments will thus be e.g. 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500,
  • the length of said fragments will thus be e.g. 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500,
  • the length of said fragments will thus be e.g. 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500,
  • nucleic acid sequences according to the present invention make it possible to estimate cancer risk in an individual by using sequence polymorphisms originating from a specific region of chromosome 19.
  • Lung cancer affects approxi- mately 10-15 percent of smokers and thus approximately 5 percent of the popula ⁇ tion, somewhat varying from country to country.
  • Malignant melanoma a sun- induced, often lethal form of skin cancer, affects approximately 700 persons a year in Denmark or about 1 percent of the Danish population.
  • sequence polymorphism is understood any single nucleotide, tandem repeat, insertion, deletion or block polymorphism, which varies among humans, whether it is of known biological importance or not.
  • one or more single nucleotide polymorphism(s) at a pre ⁇ determined position in the region are identified and used for e.g. cancer risk profiling and/or cancer treatment response profiling.
  • Presently preferred polymorphism(s) are listed in Tables 1a, 1b and 1c, more preferably at least two polymorphism(s) are selected, most preferably at least three polymorphism(s) are selected.
  • the present invention relates to any polymorphism in the region. Table 1a
  • ERCC1-3' C/T rs762562 NT 011109 18180561 57424 37267 rs numbers were derived from the NCBI's database dbSNP.
  • TCTTTCTTTCTT rs4572514 AGAACCTGTTCAGGCTGGCGGCTCA[C/T]TTGGATGAAC 28691
  • one or more single nucleotide polymorphism(s) at a pre ⁇ determined position in the region are identified and used for e.g. cancer risk profiling and/or cancer treatment response profiling.
  • Presently preferred polymorphism(s) are listed in Tables 2a and 2b, more preferably at least two poly- morphism(s) are selected, most preferably at least three polymorphism(s) are se ⁇ lected.
  • the present invention relates to any polymorphism in the region.
  • RAI intron 1-3 A/C AC092309 25115 39833 19677 rs numbers were derived from the NCBI's database dbSNP.
  • one or more single nucleotide polymorphism(s) at a pre ⁇ determined position in the region are identified and used for e.g. cancer risk profiling and/or cancer treatment response profiling.
  • Presently preferred polymorphism(s) are listed in Tables 3a, 3b and 3c, more preferably at least two polymorphism(s) are selected, most preferably at least three polymorphism(s) are selected.
  • the present invention relates to any polymorphism in the region.
  • RAI intron 1-3 A/C AC092309 25115 39833 19677 rs numbers were derived from the NCBI's database dbSNP.
  • polymorphism(s) are listed below in tables 3c and 3d, more pref ⁇ erably at least two polymorphism(s) are selected, most preferably at least three polymorphism(s) are selected from SEQ ID NO:1 below Table 3c
  • RAI intron 3 A/G rs2017104 NT 011109 18155483 32346 12190 rs numbers were derived from the NCBI's database dbSNP.
  • GAGGCTGCAGTGAGCTGT gactgtgcca ctgcactcca rs2097215 TGACAGTAGA CATCCTGTCA T (A/G) ATAAGTCttt ttttttt 1610 rs 11878644 CATCCCCATA CCAAcccacc (c/t) tactgctctg atctcctcct 2790 rs7252567 tttagtagagacatggttccgcca (C/T) gttgcccaggctggtcttgaactc 4428 rs3047560 ACTAAAAATAAAAAATAAAAAAAA(-/AA) ATAGCCGAG- 4797
  • polymorphism(s) are those listed below in tables 3e and 3f, more preferably at least two polymorphism(s) are selected, most preferably at least three polymorphism(s) are selected from SEQ ID NO: 1 below:
  • RAI intron 3 A/G rs2017104 NT 011109 18155483 32346 12190 rs numbers were derived from the NCBI's database dbSNP.
  • polymorphism(s) are those listed in tables 4a and 4b below, more preferably at least two polymorphism(s) are selected, most preferably at least three polymorphism(s) are selected from the polymorphisms shown below:
  • RAI-3'3 AAA rs3047560 NT 011109 24055 RAI exon 6 A/T rs6966 NT 011109 28043 RAI intron 3 A/G rs2017104 NT 011109 32346 rs numbers were derived from the NCBI's database dbSNP.
  • rs6966 ATTAAGTGCCTTCACACAGC (A/T) CTGGTTTAAT GTTTATAA 8785 rs2017104 gggaggctcg aggcgggc (AJG) gattgcatga gctcaggatt 12190
  • one or more single nucleotide polymor ⁇ phism ⁇ ) at a predetermined position in the region are identified and used for e.g. cancer risk profiling and/or cancer treatment response profiling.
  • the preferred polymorphism(s) are listed in Tables 5a, more preferably at least two polymorphism(s) are selected, most preferably at least three polymor- phism(s) are selected.
  • the present invention relates to any polymorphism in the region.
  • RAI intron 1-3 A/C AC092309 25115 39833 19677 rs numbers were derived from the NCBI's database dbSNP.
  • polymorphism(s) are those listed in table 5b below, more preferably at least two polymorphism(s) are selected, most preferably at least three polymor ⁇ phism ⁇ ) are selected from the polymorphisms shown below:
  • polymorphism(s) are those listed in table 5c below, more preferably at least two polymorphism(s) are selected, most preferably at least three polymorphism(s) are selected from the polymorphisms shown below:
  • RAI-i ⁇ -2 A/T rs8112723 NT 011109 18153497 30360 10204 rs numbers were derived from the NCBI's database dbSNP.
  • polymorphism(s) are those listed in table 5d below, more preferably at least two polymorphism(s) are selected, most preferably at least three polymorphism(s) are selected from the polymorphisms shown below:
  • RAI-3'8 A/G rs8101662 NT 011109 18150911 27774 8516 rs numbers were derived from the NCBI's database dbSNP.
  • At least one of the following combinations of polymor ⁇ phisms is included in the methods:
  • the method described herein is one in which the tandem repeat is at a position as described in Table 6: Table 6
  • the method for diagnosis described herein is preferably one in which the sequence polymorphism is in region r. Testing for the presence of the RAI gene allele is especially preferred because, without wishing to be bound by theoretical considerations, of its association with increased risk of can- cer (as explained herein).
  • one or more polymorphism(s) at a predetermined position in the region r are identified and used for e.g. cancer risk profiling and/or cancer treatment response profiling.
  • Presently preferred polymorphism(s) are the dinucleotide polymorphism RAI-3'3 in which AA or a deletion is present, the sin ⁇ gle nucleotide polymorphism RAI exon 6 in which A or T is present, and RAI intron 3 in which A or G is present.
  • the present invention relates to any polymor ⁇ phism and SNP in the r region.
  • the sequence polymorphism of the invention comprises at least one base differ ⁇ ence, such as at least two base differences, such as at least three base differences, such as at least four base differences, such as eighty one base pair differences.
  • the sequence polymorphism(s) comprises at least one polymor- phism, such as at least two polymorphisms, such as at least three polymorphisms, such as at least four polymorphisms.
  • the sequence polymorphism comprises at least one polymorphism, such as at least two tandem repeat polymorphisms.
  • sequence polymorphism may be a combination of single nucleotide poly- morphism and dinucleotide polymorphism, such as one single nucleotide polymor ⁇ phism and one dinucleotide polymorphism.
  • the status of the individual may be determined by reference to allelic variation at one, two, three, four or more of the above loci.
  • the cell sample used in the present invention may be any suitable cell sample ca ⁇ pable of providing the genetic material for use in the method.
  • the cell sample is a blood sample, a tissue sample, a sample of secretion, semen, ovum, a washing of a body surface (e.g. a buccal swap), a clipping of a body surface (hairs, or nails), such as wherein the cell is selected from white blood cells and tumour tissue.
  • test sample may equally be a nucleic acid sequence corresponding to the sequence in the test sample, that is to say that all or a part of the region in the sample nucleic acid may firstly be amplified using any convenient technique e.g. PCR, before use in the analysis of variation in the region.
  • Detection may be conducted on the sequence of SEQ ID NO: 1 or a complementary sequence as well as on translational (mRNA) and transcriptional products (polypep- tides, proteins) therefrom.
  • Muta- tions or polymorphisms within or flanking the r region can be detected by utilizing a number of techniques. Nucleic acid from any nucleated cell can be used as the starting point for such assay techniques, and may be isolated according to standard nucleic acid preparation procedures that are well known to those of skill in the art. In general, the detection of allelic variation requires a mutation discrimination tech ⁇ nique, optionally an amplification reaction and a signal generation system. Table 7 lists a number of mutation detection techniques, some based on the PCR.
  • Table 8 illustrates various mutation detection techniques capable of being used for SNP detection.
  • Fluorescence Fluorescence: FRET, Fluorescence quenching, Fluorescence polarisation-United Kingdom Patent No. 2228998 (Zeneca Limited)
  • Table 9 illustrates examples of further amplification techniques.
  • Preferred mutation detection techniques include ARMS, ALEX, COPS, Taqman, Molecular Beacons, RFLP, and restriction site based PCR and FRET techniques.
  • Particularly preferred methods include FRET; taqman, ARMS and RFLP based methods.
  • mutations or polymorphisms can be detected by using a microassay of nucleic acid sequences immobilized to a substrate or "gene chip” (see, e.g. Cronin, et a!., 1996, Human Mutation 7:244-255).
  • Caskey et al. (U.S. Pat. No. 5,364,759) describe a DNA profiling assay for detecting short tri and tetra nucleotide repeat sequences. The process includes ex ⁇ tracting the DNA of interest, such as the RAI gene, amplifying the extracted DNA, and labelling the repeat sequences to form a genotypic map of the individual's DNA.
  • the level of RAI gene expression can also be assayed.
  • RNA from a cell type or tissue known, or suspected, to express the RAI gene may be isolated and tested utilizing hybridization or PCR techniques such as are described, above.
  • the isolated cells can be derived from cell culture or from a patient.
  • the analysis of cells taken from culture may be a necessary step in the assessment of cells to be used as part of a cell-based gene therapy technique or, alternatively, to test the ef ⁇ fect of compounds on the expression of the RAI gene.
  • Such analyses may reveal both quantitative and qualitative aspects of the expression pattern of the RAI gene, including activation or inactivation of RAI gene expression.
  • a cDNA molecule is synthesized from an RNA molecule of interest (e.g., by reverse transcription of the RNA mole ⁇ cule into cDNA).
  • a sequence within the cDNA is then used as the template for a nucleic acid amplification reaction, such as a PCR amplification reaction, or the like.
  • the nucleic acid reagents used as synthesis initiation reagents (e.g., primers) in the reverse transcription and nucleic acid amplification steps of this method are chosen from among the RAI gene nucleic acid reagents described above.
  • the preferred lengths of such nucleic acid reagents are at least 9-30 nucleotides.
  • the nucleic acid amplification may be performed using radio- actively or non-radioactively labeled nucleotides.
  • enough amplified product may be made such that the product may be visualized by standard ethidium bromide staining or by utilizing any other suitable nucleic acid staining method.
  • RAI gene expression assays "in situ", i.e., directly upon tissue sections (fixed and/or frozen) of patient tissue obtained from biopsies or resections, such that no nucleic acid purification is necessary.
  • Nucleic acid reagents such as those described above may be used as probes and/or prim ⁇ ers for such in situ procedures (see, for example, Nuovo, G. J., 1992, “PCR In Situ Hybridization: Protocols And Applications", Raven Press, NY).
  • stan ⁇ dard Northern analysis can be performed to determine the level of mRNA expres ⁇ sion of the RAI gene.
  • Another method for detecting sequence polymorphism is by analysing the activity of gene products resulting from the sequences. Accordingly, in one embodiment the detection uses the activity of the RAI gene product as compared to a reference in the method. In particular if the activity of the genes are decreased or increased by at least or about 50 %, such as at least or about 40%, for example at least or about 30%, such as at least or about 20%, for example at least or about 10%, such as at least or about 10%, for example at least or about 5%, such as at least or about 2%, it indicates a sequence polymorphism in the gene.
  • the present invention may combine the result of sequence polymorphism within the region r with sequence polymorphism outside the region in order to increase the probability of the correlation.
  • the primer nucleotide sequences of the invention further include: (a) any nucleotide sequence that hybridizes to a nucleic acid molecule of the region r or its comple ⁇ mentary sequence or RNA products under stringent conditions, e.g., hybridization to filter-bound DNA in 6x sodium chloride/sodium citrate (SSC) at about 45°C followed by one or more washes in 0.2x SSC/0.1% SDS at about 50-65 0 C, or (b) under highly stringent conditions, e.g., hybridization to filter-bound nucleic acid in 6x SSC at about 45°C followed by one or more washes in 0.1 x SSC/0.2% SDS at about 68°C, or under other hybridization conditions which are apparent to those of skill in the art (see, for example, Ausubel F.M.
  • stringent conditions e.g., hybridization to filter-bound DNA in 6x sodium chloride/sodium citrate (SSC) at about 45°C
  • nucleic acid molecule that hybrid- izes to the nucleotide sequence of (a) and (b), above is one that comprises the complement of a nucleic acid molecule of the region s or r or a complementary se ⁇ quence or RNA product thereof.
  • nucleic acid molecules comprising the nucleotide sequences of (a) and (b) comprises nucleic acid mole ⁇ cule of RAI or a complementary sequence or RNA product thereof.
  • oli- gos deoxyoligonucleotides
  • TM melting temperature
  • Exemplary highly stringent conditions may refer, e.g., to washing in 6x SSC/0.05% sodium pyrophosphate at 37 0 C (for about 14-base oligos), 48°C (for about 17-base oligos), 55°C (for about 20-base oligos), and 6O 0 C (for about 23-base oligos).
  • the invention further provides nucleotide primers or probes which de- tect the r region polymorphisms of the invention.
  • the assessment may be conducted by means of at least one nucleic acid primer or probe, such as a primer or probe of DNA, RNA or a nucleic acid analogue such as peptide nucleic acid (PNA) or locked nucleic acid (LNA).
  • the nucleotide primer or probe is preferably capable of hybridis ⁇ ing to a subsequence of the region corresponding to SEQ ID NO: 1, or a part thereof, or a region complementary to SEQ ID NO: 1.
  • an allele-specific oligonucleotide probe capable of detecting a r region polymorphism at one or more of positions in the r region as defined by the positions in SEQ ID NO: 1.
  • the allele-specific oligonucleotide probe is preferably 5-50 nucleotides, more pref ⁇ erably about 5-35 nucleotides, more preferably about 5-30 nucleotides, more pref ⁇ erably at least 9 nucleotides.
  • Such probes will be apparent to the molecular biologist of ordinary skill.
  • Such probes are of any convenient length such as up to 50 bases, up to 40 bases, more conveniently up to 30 bases in length, such as for example 8-25 or 8- 15 bases in length.
  • such probes will comprise base sequences entirely complementary to the corresponding wild type or variant locus in the region. How- ever, if required one or more mismatches may be introduced, provided that the dis ⁇ criminatory power of the oligonucleotide probe is not unduly affected.
  • the probes of the invention may carry one or more labels to facilitate detection.
  • the primers and/or probes are capable of hybridizing to and/or amplifying a subsequence hybridizing to a single nucleotide polymorphism contain ⁇ ing the sequence shown herein selected from the group of subsequences below or a sequence complementary thereto, wherein the polymorphism is denoted as for ex ⁇ ample T/C:
  • gctgcagtga gctgt (-/ACACCTGTGGTCCCAGCTACTCTGG AAGCTGAGGTGGGAGGATCGCTTGAGCCCAAGAGGTGGAGGCTG-
  • CAGTGAGCTGT gactgtgcca ctgcactcca
  • the primers and/or probes are capable of hybridizing to and/or amplifying a subsequence hybridizing to a single nucleotide polymorphism containing the sequence shown herein selected from the group of subsequences below or a sequence complementary thereto, wherein the polymorphism is denoted as for example T/C:
  • the primers and/or probes are capable of hybridizing to and/or amplifying a subsequence hybridizing to a single nucleotide polymorphism containing the sequence shown herein selected from the group of subsequences below or a sequence complementary thereto, wherein the polymorphism is denoted as for example TYC:
  • gctgcagtga gctgt (-/ACACCTGTGGTCCCAGCTACTCTGG AAGCTGAGGTGGGAGGATCGCTTGAGCCCAAGAGGTGGAGGCTG-
  • CAGTGAGCTGT gactgtgcca ctgcactcca
  • the primers and/or probes are capable of hybridizing to a subsequence selected from the group of subsequences below:
  • the primers and/or probes are capable of hybridizing to a subsequence selected from the group of subsequences below:
  • TGTTGTCCAA GCTGGCAGAG (A/G) TTTTTGTTTG TTTGTTTGAG 5.
  • CTCGGGAGGCTGAGGCAGGAGAATC (A/G) CTTGAACTCAGGCA-
  • primers and/or probes capable of hybridizing to a subse ⁇ quence selected from the group of subsequences below:
  • gggaggctcg aggcgggc (A/G) gattgcatga gctcaggatt are the primers and/or probes capable of hybridizing to a subsequence selected from the group of subsequences below:
  • the primers and/or probes are capable of hybridizing to a subsequence selected from the group of subsequences below:
  • primers and/or probes are capable of hybridizing to a subsequence selected from the group of subsequences below:
  • primers and/or probes are capable of hybridizing to a subsequence selected from the group of subsequences below:
  • CTTGCTACAGAATTACAGGCA GCGCCACCGCTCCGGGCTAA 2.
  • CTAAAGACTACA VA
  • tgc ctt cac aca get ctg gtt taa tg 20.
  • egg get aca ggg tta cct gag 40.
  • age tgc cag age tgc ctg ggc One or more primers able to detect subsequences by hybridisation as described may in a particularly preferred embodiment of the method of the invention be se ⁇ lected from
  • a diagnostic nucleic acid primer capable of detecting a r region polymorphism at one or more of positions in the r region as defined by the in SEQ ID NO: 1.
  • the primer or probe may be a diagnostic nucleic acid primer defined as an allele specific primer, used, generally together with a constant primer, in an amplification reaction such as a PCR reaction, which provides the discrimination between alleles through selective amplification of one allele at a particular sequence position.
  • the diagnostic primer is preferably 5-50 nucleotides, more preferably about 5-35 nucleo ⁇ tides, more preferably about 5-30 nucleotides, more preferably at least 9 nucleo ⁇ tides.
  • diagnostic primers compris- ing the sequences set out below as well as derivatives thereof wherein about 6-8 of the nucleotides at the 3' terminus are identical to the sequences given below and wherein up to 10, such as up to 8, 6, 4, 2, or 1 of the remaining nucleotides may be varied without significantly affecting the properties of the diagnostic primer.
  • up to 10 such as up to 8, 6, 4, 2, or 1 of the remaining nucleotides may be varied without significantly affecting the properties of the diagnostic primer.
  • At least two sets of primer(s) and/or probe(s), such as at least three sets of primer(s) and/or probe(s) may be combined in the method thereby increasing the correlation probability.
  • This second or other set of primer(s) and/or probe(s) may be a nucleotide or nucleotide analogues hybridising to a region within the region r or to a sequence different from the region r. Said sequence differ- ent from the region r is preferably a region in chromosome 19, preferably in chromo ⁇ some 19q.
  • such second or other primer or probe may be selected from one or more of the sequences below, or the complementary strands:
  • the primers and probes may be manufactured using any convenient method of syn ⁇ thesis. Examples of such methods may be found in standard textbooks, for example "Protocols for Oligonucleotides and Analogues; Synthesis and Properties," Methods in Molecular Biology Series; Volume 20; Ed. Sudhir Agrawal, Humana ISBN: 0- 89603-247-7; 1993; Lsup.st Edition. If required the primer(s) and probe(s) may be labelled to facilitate detection.
  • a diagnostic kit comprising at least one diagnostic primer of the invention and/or at least one al- lele-specific oligonucleotide primer of the invention.
  • kits may comprise appropriate packaging and instructions for use in the methods of the invention.
  • Such kits may further comprise appropriate buffer(s) and polymerase(s) such as thermostable polymerases, for example taq polymerase.
  • kits can comprise means for amplifying the relevant sequence such as primers, polymerase, deoxynucleotides, buffer, metal ions; and/or means for dis ⁇ criminating the polymorphism, such as one or a set of probes hybridising to the poly ⁇ morphic site, a sequence reaction covering the polymorphic site, an enzyme or an antibody; and/or a secondary amplification system, such as enzyme-conjugated antibodies, or fluorescent antibodies.
  • the kit-of-parts preferably also comprises a detection system, such as a fluorometer, a film, an enzyme reagent or another highly sensitive detection device.
  • kits for detect ⁇ ing the presence of a polypeptide or nucleic acid of the invention in a biological sample i.e., a test sample.
  • a biological sample i.e., a test sample.
  • kits for detect ⁇ ing the presence of a polypeptide or nucleic acid of the invention in a biological sample i.e., a test sample.
  • kits for detect ⁇ ing the presence of a polypeptide or nucleic acid of the invention in a biological sample i.e., a test sample.
  • kits can be used, e.g., to determine if a subject is suffering from or is at increased risk of developing a disorder associated with a dis ⁇ order-causing allele, or aberrant expression or activity of a polypeptide of the inven- tion.
  • the kit can comprise a labeled compound or agent capable of detecting the polypeptide or mRNA or DNA or RAI gene sequences, e.g., encoding the polypeptide in a biological sample.
  • the kit can further comprise a means for de ⁇ termining the amount of the polypeptide or mRNA in the sample (e.g., an antibody which binds the polypeptide or an oligonucleotide probe which binds to DNA or mRNA encoding the polypeptide).
  • Kits can also include instructions for observing that the tested subject is suffering from or is at risk of developing a disorder associ ⁇ ated with aberrant expression of the polypeptide if the amount of the polypeptide or mRNA encoding the polypeptide is above or below a normal level, or if the DNA correlates with presence of an RAI allele that causes a disorder.
  • the kit can comprise, for example: (1) a first antibody (e.g., attached to a solid support) which binds to a polypeptide of the invention; and, op ⁇ tionally, (2) a second, different antibody which binds to either the polypeptide or to the first antibody and is conjugated to a detectable agent.
  • a first antibody e.g., attached to a solid support
  • a second, different antibody which binds to either the polypeptide or to the first antibody and is conjugated to a detectable agent.
  • An allele in the r region can be identified as correlated with an increased risk of de ⁇ veloping cancer on the basis of statistical analyses of the incidence of a particular allele in two groups of individuals with and without cancer, respectively, according to the ⁇ 2 test, which is well known in the art. Furthermore, an allele in the region can be identified as an allele correlated with prognosis of cancer on the basis of statistical analyses of the incidence of a particular allele in individuals demonstrating different prognostic characteristics. Identification of humans having increased likelihood of responding to treat ⁇ ment
  • the present invention provides a method for identifying a human subject as having an increased likelihood of responding positively to a cancer treatment, comprising determining the presence in the subject of a s or r re ⁇ gion allele genotype correlated with an increased likelihood of positive response to treatment, whereby the presence of the genotype identifies the subject as having an increased likelihood of responding to cancer treatment.
  • the treatment mentioned herein may be any cancer treatment, such as conventional cancer treatment, for example X-ray, chemotherapeutics, surgical excision or com ⁇ binations thereof.
  • Gene products of the region r or peptide fragments thereof can be prepared for a variety of uses.
  • such gene products, or peptide fragments thereof can be used for the generation of antibodies, in diagnostic assays.
  • the gene products of the invention include, but are not limited to, human RAI gene products. In the following the invention is described in relation to RAI gene product.
  • Gene product sometimes referred to herein as an "protein” or “polypeptide”, in- eludes those gene products encoded by the RAI gene sequences shown as position 7760-22885 in SEQ ID NO: 1.
  • gene product variants are gene products comprising amino acid residues encoded by the polymorphisms.
  • Such gene product variants also include a variant of the RAI gene product.
  • RAI gene products may include proteins that represent functionally equi ⁇ valent gene products.
  • functionally equivalent RAI gene products are naturally occurring gene products.
  • Functionally equivalent RAI gene products also include gene products that retain at least one of the biological activities of the RAI gene products described above, and/or which are recognized by and bind to antibodies (polyclonal or monoclonal) directed against RAI gene prod ⁇ ucts.
  • the terms "spe- cifically bind” and “specifically recognize” refer to antibodies that bind to RAI gene product epitopes at a higher affinity than they bind to non-RAI (e.g., random) epi ⁇ topes.
  • Such antibodies may include, but are not limited to, polyclonal antibodies, mono- clonal antibodies (mAbs), humanized or chimeric antibodies, single chain antibodies, Fab fragments, F(ab') 2 fragments, fragments produced by a Fab expression library, anti-idiotypic (anti-Id) antibodies, and epitope-binding fragments of any of the above, including the polyclonal and monoclonal antibodies described below.
  • mAbs mono- clonal antibodies
  • Such antibod ⁇ ies may be used, for example, in the detection of a gene product in a biological sample and may, therefore, be utilized as part of a diagnostic or prognostic tech ⁇ nique whereby patients may be tested for abnormal levels of gene products, and/or for the presence of abnormal forms of such gene products.
  • Such antibodies may also be utilized in conjunction with, for example, compound screening schemes, as described, below, for the evaluation of the effect of test compounds on gene product levels and/or activity.
  • various host animals may be immunized by injection with a RAI gene product, or a portion thereof.
  • Such host animals may include, but are not limited to rabbits, mice, and rats, to name but a few.
  • Various adjuvants may be used to increase the immunological response, de ⁇ pending on the host species, including but not limited to Freund's (complete and in ⁇ complete), mineral gels such as aluminum hydroxide, surface active substances such as lysolecithin, pluronic polyols, polyanions, peptides, oil emulsions, keyhole limpet hemocyanin, dinitrophenol, and potentially useful human adjuvants such as BCG (bacille Calmette-Guerin) and Corynebacterium parvum.
  • BCG Bacille Calmette-Guerin
  • Polyclonal antibodies are heterogeneous populations of antibody molecules derived from the sera of animals immunized with an antigen, such as a gene product, or an antigenic functional derivative thereof.
  • an antigen such as a gene product, or an antigenic functional derivative thereof.
  • host animals such as those described above, may be immunized by injection with gene product supplemented with adjuvants as also described above.
  • Monoclonal antibodies which are homogeneous populations of antibodies to a par ⁇ ticular antigen, may be obtained by any technique that provides for the production of antibody molecules by continuous cell lines in culture. These include, but are not limited to, the hybridoma technique of Kohler and Milstein (1975, Nature 256:495- 497; and U.S. Pat. No. 4,376,110), the human B-cell hybridoma technique (Kosbor et al., 1983, Immunology Today 4:72; Cole et al., 1983, Proc. Natl. Acad. Sci. U.S.A.
  • Such antibodies may be of any immunoglobulin class including IgG, IgM, IgE, IgA, IgD and any subclass thereof.
  • the hybridoma producing the mAb of this invention may be cultivated in vitro or in vivo. Production of high titers of mAbs in vivo makes this the presently preferred method of production.
  • chimeric antibodies In addition, techniques developed for the production of "chimeric antibodies" (Morri ⁇ son, et al., 1984, Proc. Natl. Acad. Sci., 81 :6851-6855; Neuberger, et al., 1984, Na ⁇ ture 312:604-608; Takeda, et al., 1985, Nature, 314:452-454) by splicing the genes from a mouse antibody molecule of appropriate antigen specificity together with genes from a human antibody molecule of appropriate biological activity can be used.
  • a chimeric antibody is a molecule in which different portions are derived from different animal species, such as those having a variable region derived from a mur ⁇ ine mAb and a human immunoglobulin constant region.
  • An immunoglobulin light or heavy chain variable region consists of a "framework" region interrupted by three hypervariable regions, referred to as complementarity determining regions (CDRs).
  • CDRs complementarity determining regions
  • the extent of the framework region and CDRs have been precisely defined (see, "Sequences of Proteins of Im ⁇ munological Interest", Kabat, E. et al., U.S. Department of Health and Human Ser ⁇ vices (1983) ).
  • humanized antibodies are antibody molecules from non- human species having one or more CDRs from the non-human species and a framework region from a human immunoglobulin molecule.
  • Antibody fragments that recognize specific epitopes may be generated by known techniques.
  • such fragments include but are not limited to: the F(ab') 2 fragments, which can be produced by pepsin digestion of the antibody molecule and the Fab fragments, which can be generated by reducing the disulfide bridges of the F(ab') 2 fragments.
  • Fab expression libraries may be constructed (Huse, et al., 1989, Science 246:1275-1281) to allow rapid and easy identification of mono ⁇ clonal Fab fragments with the desired specificity.
  • Immunoassays for gene products, conserved variants, or peptide fragments thereof will typically comprise incubating a sample, such as a biological fluid, a tissue ex- tract, freshly harvested cells, or lysates of cells in the presence of a detectably la ⁇ beled antibody capable of identifying gene product, conserved variants or peptide fragments thereof, and detecting the bound antibody by any of a number of tech ⁇ niques well-known in the art.
  • the biological sample may be brought in contact with and immobilized onto a solid phase support or carrier, such as nitrocellulose, that is capable of immobilizing cells, cell particles or soluble proteins.
  • a solid phase support or carrier such as nitrocellulose, that is capable of immobilizing cells, cell particles or soluble proteins.
  • the support may then be washed with suitable buffers followed by treatment with the detectably labeled gene product specific anti ⁇ body.
  • the solid phase support may then be washed with the buffer a second time to remove unbound antibody.
  • the amount of bound label on the solid support may then be detected by conventional means.
  • solid phase support or carrier any support capable of binding an antigen or an antibody.
  • supports or carriers include glass, polystyrene, polypropylene, polyethylene, dextran, nylon, amylases, natural and modified cellu ⁇ loses, polyacrylamides, gabbros, and magnetite.
  • the nature of the carrier can be either soluble to some extent or insoluble for the purposes of the present invention.
  • the support material may have virtually any possible structural configuration so long as the coupled molecule is capable of binding to an antigen or antibody.
  • the support configuration may be spherical, as in a bead, or cylindrical, as in the inside surface of a test tube, or the external surface of a rod.
  • the surface may be flat such as a sheet, test strip, etc.
  • Preferred supports include polystyrene beads. Those skilled in the art will know many other suitable carriers for binding antibody or antigen, or will be able to ascertain the same by use of routine experimentation.
  • RAI gene product-specific antibody can be detectably labeled is by linking the same to an enzyme, malate dehydrogenase, staphylococcal nuclease, delta-5-steroid isomerase, yeast alcohol dehydrogenase, ⁇ -glycero- phosphate, dehydrogenase, triose phosphate isomerase, horseradish peroxidase, alkaline phosphatase, asparaginase, glucose oxidase, ⁇ -galactosidase, ribonucle- ase, urease, catalase, glucose-6-phosphate dehydrogenase, glucoamylase and acetylcholinesterase.
  • the detection can be accomplished by colorimetric methods that employ a chromogenic substrate for the enzyme. Detection may also be ac- complished by visual comparison of the extent of enzymatic reaction of a substrate in comparison with similarly prepared standards.
  • Detection may also be accomplished using any of a variety of other immunoassays. For example, by radioactively labeling the antibodies or antibody fragments, by Ia- beling the antibody with a fluorescent compound.
  • fluorescent labeling compounds are fluorescein isothiocyanate, rhodamine, phyco- erythrin, phycocyanin, allophycocyanin, o-phthaldehyde and fluorescamine.
  • the antibody can also be detectably labeled using fluorescence emitting metals such as 152 Eu, or others of the lanthanide series or by coupling it to a chemilumines- cent compound.
  • Described herein are various applications of gene sequences, gene products, in ⁇ cluding peptide fragments and fusion proteins thereof, and of antibodies directed against gene products and peptide fragments thereof.
  • Such applications include, for example, prognostic and diagnostic evaluation of a disease, such as cancer, and the identification of subjects with a predisposition to such disorders, as described above.
  • the method according to the invention may be used in relation to any cancer form, such as, but not limited to, skin carcinoma including malignant melanoma, breast cancer, lung cancer, colon cancer and other cancers in the gastro-intestinal tract, prostate cancer, lymphoma, leukemia, multiple myeloma, pancreas cancer, head and neck cancer, ovary cancer and other gynecological cancers.
  • the method is relevant for skin cancer, lung cancer, colon cancer, multiple myeloma, and breast cancer, such as skin cancer, breast cancer, multiple myeloma, and lung cancer, such as skin cancer, breast cancer and lung cancer, such as skin cancer and breast cancer, preferably wherein the skin cancer is basal cell carcinoma, such as lung cancer.
  • the cancer is multiple myeloma
  • the cancer is breast cancer.
  • the cancer may also be skin cancer, preferably basal cell carcinoma, for example early age basal cell carcinoma.
  • the method is relevant for both early age cancer and later age cancer, such as early age breast cancer, and such as later age breast cancer.
  • the method is of particular relevance for lung cancer, such as in patients with XPD exon 23 M .
  • the method is also in particular relevant for early age skin cancer, such as early age basal cell carcinoma.
  • Gene nucleic acid sequences can be utilized for transferring re ⁇ combinant nucleic acid sequences to cells and expressing said sequences in recipi- ent cells. Such techniques can be used, for example, in marking cells or for the treatment of cancer. Such treatment can be in the form of gene replacement ther ⁇ apy. Specifically, one or more copies of a normal RAI gene or a portion of the RAI gene that directs the production of an RAI gene product exhibiting normal RAI gene function, may be inserted into the appropriate cells within a patient, using vectors that include, but are not limited to, adenovirus, adeno-associated virus, and retrovi ⁇ rus vectors, in addition to other particles that introduce DNA into cells, such as lipo ⁇ somes.
  • the invention may be used in relation to inflammatory dis- eases, such as, but not limited thereto, rheumatoid arthritis, colitis ulcerosa, Crohn's disease, thyroiditis, neural inflammation as in Alzheimer's disease, and Guillain- Barre syndrome.
  • inflammatory dis- eases such as, but not limited thereto, rheumatoid arthritis, colitis ulcerosa, Crohn's disease, thyroiditis, neural inflammation as in Alzheimer's disease, and Guillain- Barre syndrome.
  • the primary effectors i.e. the mutations causing cancer according to the present invention may be found within the region including and flanked by the marker RAI- 37 and the polymorphism having the sequence
  • the region and markers herein represents the region around position 24000 which is shown to be important in relation to can ⁇ cer according to the present invention.
  • the primary effector may be selected from the group consisting of
  • the primary effectors i.e. the mutations causing cancer according to the present invention may be be selected from the group consisting of
  • the primary effectors i.e. the mutations causing cancer according to the present invention may be be selected from the group consisting of
  • the primary effectors i.e. the mutations causing cancer ac ⁇ cording to the present invention may further be be selected from the group consist ⁇ ing of
  • the primary effectors i.e. the mutations causing cancer according to the present invention may further be be selected from the group con ⁇ sisting of
  • the primary effectors i.e. the mutations causing cancer ac ⁇ cording to the present invention may further be be selected from the group consist ⁇ ing of
  • the primary effectors i.e. the mutations causing cancer according to the present invention may further be be selected from the group consisting of
  • the primary effectors i.e. the mutations causing cancer according to the present invention may further be be selected from the group consisting of
  • the primary effectors i.e. the mutations causing cancer according to the present invention may further be be selected from the group consisting of
  • the primary effector may be the mutation corresponding to the sequence TTTTAG- TAGAGACATGGTTCCGCCA[C/ ⁇ GTTGCCCAGGCTGGTCTTGAACTCC posi ⁇ tioned at 18146823 of Contig nt_011109.
  • the primary effec ⁇ tor may be the mutation corresponding to the sequence ctggggaggctgaggcagga- gaatc[A/G]cttgaaaccgggaggcggaggttgt positioned at 18147126 of Contig nt_011109.
  • the primary effector may be the mutation corresponding to the sequence GATTGTCATGT[G/T]ACATCAGCCAATACT posi ⁇ tioned at 18146233 of Contig nt_011109.
  • the primary effector may be the mutation corresponding to the sequence caggcggatca- caaggtcaggagtt[C/T]gagaccagcctggccaacacagtga positioned at 18148193 of Contig nt_011109.
  • the primary effector may be the mutation corre ⁇ sponding to the sequence CACAGTGAAAC[C/T]CCATCTCTACTAAA positioned at 18149120 of Contig nt_011109.
  • the primary effec- tor may be the mutation corresponding to the sequence AGCCTGGCCAACATG[CZG]TGAAACCCCGTCTCT positioned at 18150815 of Contig nt_011109.
  • the primary effector may be the muta ⁇ tion corresponding to the sequence ctcgggaggctgaggcagga- gaatc[A/G]cttgaactcaggaggcagaggttgc positioned at 18150911 of Contig nt_011109.
  • the primary effector may be the mutation corresponding to the sequence AAGTTTCTCTATT[GZT]TGTTTATAAACA positioned at 18151158 of Contig nt_011109.
  • the primary effector may be the mutation corresponding to the sequence CCCTATGTTGTCCAAGCTGGCAGAG[AZG]TTTTT-GTTTGTTTGTTTGAGAGGGA positioned at 18150199 of Contig nt_011109.
  • the primary effector may be the mutation corresponding to the sequence ac- taaaataaaaaataaaaaaaaaaaa[-ZAA]atagccgagcatggtggtgggtgcc positioned at 18147192 of Contig nt_011109.
  • the primary effector may be the muta- tion corresponding to the sequence AAAAAACTAAAGTGGGGTTTGCGGG[GZT]- AGTGGGAGGGCCCTTCCTGCTAGGT positioned at 18147886 of Contig nt_011109 .
  • the primary effector may be the mutation cor ⁇ responding to the sequence AAAATTAGCCGG[AZG]CGCCATGGCGGGAG posi ⁇ tioned at 18149154 of Contig nt_011109.
  • the primary effector may be the mutation corresponding to the sequence GGTTTAT[ATTTT]Ntgagatggatttt positioned at 18147012 of Contig nt_011109.
  • the primary effectors may however be any combination of the polymorphisms.
  • polymorphisms employed herein for providing a method andZor composi- tions for identifying human subjects with an increased risk of having or developing disease
  • a number of polymorphisms are novel and have been identified by the pre ⁇ sent inventors.
  • the position according to Contig nt_011109 and the nucleotide sequence of the identified polymorphism are provided, whereas the novel polymorphisms cannot be assigned trivial names or identification numbers according to the dbSNP database.
  • the primary effectors i.e. the mutations causing cancer according to the present invention may be found within the region around position 39000 which is shown to be important in relation to cancer according to the present invention.
  • the region includes and is flanked by the marker Rai intron 1 and the polymorphism having the sequence TCCAGCCTGGGCAAGAA[C/G]AGTGAAACTCCAGCTT corresponding to position 18165052 of contig nt_011109.
  • the primary effector may be selected from the group consisting of
  • the primary effectors i.e. the mutations caus ⁇ ing cancer according to the present invention may be be selected from the group consisting of
  • the primary effectors i.e. the mutations causing cancer according to the present invention may be be selected from the group consisting of
  • the primary effectors i.e. the mutations causing cancer according to the present invention may be be selected from the group consisting of
  • the primary effectors i.e. the mutations caus ⁇ ing cancer according to the present invention may be be selected from the group consisting of
  • the primary effectors i.e. the mutations causing cancer according to the present invention may be be selected from the group consisting of
  • the primary effector may be selected from the group con ⁇ sisting of
  • the primary effector may be the muta ⁇ tion corresponding to the sequence ggagcttgcagtgagctga- gatcgc[A/G]ccactgcactccagcctgggcgaca positioned at 18159091 of Contig nt_011109.
  • the primary effector may be the mutation corre ⁇ sponding to the sequence TTCTCCTGACCTC[A/G]TGATCCGCCCACCTCGG posi ⁇ tioned at 18159263 of Contig nt_011109.
  • the primary effector may be the mutation corresponding to the sequence GGGATTACAGG- CATGC[A/G]CCACCAGGCCCAGCTAATTTTTGT positioned at 18160363 of Contig nt_011109.
  • the primary effector may be the mutation corre ⁇ sponding to the sequence TCCAATGGTGACA[A/C]CAGTAAGAGCAGTTAACAG positioned at 18160936 of Contig nt_011109.
  • the primary effector may be the mutation corresponding to the sequence TCCAATGGTGACAA[C/G]AGTAAGAGCAGTTAACAG positioned at 18160937 of Contig nt_011109.
  • the primary effector may be the mutation corresponding to the sequence tacaggcgcccgccac- cacccccag[A/C]taatttttgtatttttagtagagac positioned at 18161433 of Contig nt_011109.
  • the primary effector may be the mutation corre ⁇ sponding to the sequence TTGCCTCAGCCTCCTGA[G/T]TAGCTGGGATTGGAATGAGA positioned at 18161694 of Contig nt_011109.
  • the primary effec- tor may be the mutation corresponding to the sequence TACGA- TAAATAGCTAGA[C/GACCTTGGCGCCACCATCT ⁇ positioned at 18161841 of Contig nt_011109.
  • the primary effector may be the mutation corresponding to the sequence AAAATAATAATAATAATAT- TAA[C/T]CCTGACCTTGGCGCCACCATCT positioned at 18161896 of Contig nt_011109.
  • the primary effector may be the muta ⁇ tion corresponding to the sequence tcgtcctgctacagaatta- caggca[C/T]gcgccaccgctccgggctaattttt positioned at 18162206 of Contig nt_011109.
  • the primary effector may be the mutation corresponding to the sequence CCTCATGAGCCACCCAC[CZT]TCGGCCTCCCAAAGTGCT posi- tioned at 18162309 of Contig nt_011109.
  • the primary effector may be the mutation corresponding to the sequence TGAGCCACCGCGCCC[A/G]GCCGAGACTCACTATTT positioned at 18162356 of Contig nt_011109.
  • the primary effector may be the mu ⁇ tation corresponding to the sequence taaagcgggag- gatggcttgaacct[A ⁇ 3]ggaggcggaggttgcagtgagccga positioned at 18162599 of Contig nt_011109. In a further embodiment.
  • the primary effec ⁇ tor may be the mutation corresponding to the sequence GGAGAGAAGGAGCAGA- GAAC[A/C]TCTCTATGTGGCCA positioned at 18162903 of Contig nt_011109.
  • the primary effector may be the mutation corre- sponding to the sequence ATCCTAAAGACTAC[A/C]TTTCCCAGCATCCCA posi ⁇ tioned at 18162970 of Contig nt_011109.
  • the primary effector may be the mutation corresponding to the sequence TTTCCCAG- CATCCCA[C/T]TGCAATGAGGCTCCTGGCC positioned at 18162986 of Contig nt_011109.
  • the primary effector may be the mutation corresponding to the sequence TCCTGACTCCAGTG[A/C]GGTGCCTACAGTCCTG positioned at 18163200 of Contig nt_011109. In yet a preferred embodiment the primary effector may be the mutation corresponding to the sequence TTCCAGCCTGGGCAAGAA[C/G]AGTGAAACTCCAGCTT positioned at 18165052 of Contig nt_011109. In a further preferred embodiment the primary effectors may however be any combination of the polymorphisms. RAI function
  • the primary effector influences activity or expression of RAI or XPD.
  • the effector could modify the sequence of RAI protein, through amino acid substitution, splicing or termination.
  • one likely position of the effector is in the 3' portion of the gene and that the effector modifies mRNA expression and stability (22).
  • the effector was a modification of an enhancer situated between RAI and XPD, again modulating the expression of RAI, and possibly modulating other nearby genes as well.
  • the primary effector may be located in the promoter or first intron of RAI and influence the transcriptional activity of the gene.
  • RAI protein is an inhibitor of ReIA, a subunit of the transcription factor NF- ⁇ B (23, 24). NF- ⁇ B has long been implicated in both cell proliferation and apoptosis. Modulation of NF- ⁇ B may well be part of a "crunch-time scenario", in ⁇ voked when the cell has to muster its forces and make life-and-death decisions. A recent scientific paper suggests that the choice between the cell survival and death is regulated by the relative activity of the two subunits encoded by ReIA and c-rel, respectively (25). By neutralising the ReIA product, RAI protein would presumably shift this balance towards apoptosis. RAI protein may also influence p53 activity (26).
  • the present invention relates to method of estimating the disease risk of an individ ⁇ ual comprising further a predictor of RAI action.
  • the length of the RAI transcript may be used as a predictor.
  • the quantity of RAI transcripts may be used as a predictor.
  • the present invention comprises the combined RAI transcript characteristics described above as a predictor of RAI action and the estimation of the disease risk of an individual. Examples
  • HRT hormone replacement therapy
  • menopausal status were also recorded.
  • blood, urine, fat tissue and other biological material was sam ⁇ pled and stored in a biobank at -15O 0 C.
  • Cohort members were identified by a unique identification number, which is allocated to every Danish citizen by the Central Population Registry. Cohort members were linked to The Central Population Regis ⁇ try for information on vital status and immigration. Information on cancer occurrence among cohort members was obtained through record linkage to the Danish Cancer Registry, which collects information on all inhabitants in Denmark who develop can ⁇ cer (7). Linkage was performed by use of the personal identification number.
  • Basal cell carcinoma BCC basal cell carcinoma
  • the groups of Caucasian Americans with and without BCC have been described previously (Athas et al, Cancer Res. 51 :5786-5793, 1991; Wei et al, Proc. Natl. Acad. Sci USA, 90: 1614-8, 1994). Briefly, the study was a clinic based case control study at the Johns Hopkins Hospital, which serves multiple participating dermatologists in Maryland. Cases were histo-pathologically confirmed primary BCCs and were diagnosed be ⁇ tween 1987 and 1990. The controls were patients from the same physician practices and had a diagnosis of mild skin disorders. All participants were Caucasians living near Baltimore and were between 20 and 60 years of age.
  • the controls were fre- quency matched to the cases by age and sex. Cases and controls with any other forms of cancer were excluded.
  • the study subjects were asked if they had any blood relatives with skin cancer, and were asked to specify the type of cancer. Study subjects with relatives with basal cell carcinoma and squamous cell carcinoma and 'skin cancer' were included in the group of subjects with a family of skin cancer. Subjects with relatives with melanoma were not included.
  • the subjects gave informed consent, were examined by dermatologists, com ⁇ pleted a structured questionnaire and provided blood. Available frozen lymphocytes were genotyped. Initially, 71 cases and 118 controls were included in this study. However, the number of persons varied between analyses, as the supply of DNAs gradually was depleted. In case of the SNP RAI intron 1 only 133 persons could be genotyped reliably.
  • the examples relate to prediction from sequence polymorphisms in and around the region r to cancer, see Fig. 1 for an overview of a subregion of chromosome 19q.
  • 27 markers were investigated within this 69 kb stretch for associa ⁇ tion with breast, lung, skin cancer, and multiple myeloma, using linkage disequilib ⁇ rium mapping based on single markers, as well as, on haplotypes combining several neighboring markers.
  • haplotypes with maximal association to all three cancers are centered around the gene RAI.
  • Table 10 lists the polymorphisms used in this study, their nature, the numbers in the NCBI database dbSNP, the sequence defining the SNPs, and their position therein, the currently estimated chromsome positions and their relative position within the region of interest (which can readily be calculated based on the information provided in the table).
  • Table 11 lists the primers used for the PCR reactions.
  • Table 12 lists the probes used for detection and typing.
  • Table 13 lists the PCR regimen for the poly ⁇ morphisms using Lightcycler, Taqman and ABI 2720.
  • rs numbers were derived from the NCBI's database dbSNP; Nucleotide sequence of polymorphisms identified in the present invention with no trivial name as yet:
  • Haplotype trend regression is basically a two-stage procedure. First, genotype results in combination with population assumptions such as Hardy-Weinberg equilibrium are used to con ⁇ struct all haplotype probabilities corresponding to a given set of markers for each individual. Secondly, the disease state (1 for cases, 0 for controls) is regressed on the haplotype probabilities of all individuals, resulting in a p-value for the overall as- sociation of the set of markers with disease, and parameters for association for each specific haplotype with disease.
  • population assumptions such as Hardy-Weinberg equilibrium are used to con ⁇ struct all haplotype probabilities corresponding to a given set of markers for each individual.
  • the disease state (1 for cases, 0 for controls) is regressed on the haplotype probabilities of all individuals, resulting in a p-value for the overall as- sociation of the set of markers with disease, and parameters for association for each specific haplotype with disease.
  • the HelixTree program will use sets of markers of defined size, typically 2 to 4 neighboring markers at a time, to scan the entire region. It can then be used to derive frequencies of the individual haplotypes, to calculate an overall p-value for the distribution of haplotypes covering a given set of markers among cases and controls, and to calculate p-values for the distribution of each hap ⁇ lotype derived from a given set of markers.
  • the program RASCAL for performing gene localization according to Lazzeroni (16) was implemented in Delphi (B.A. Nex ⁇ , unpublished). We used a bootstrap set of 10000 sets of haplotypes selected from the original set with replacement, and to avoid a Q-form with negative values we used a relaxation value of 0.25. A 95% con- fidence interval for the location of the causative gene variant was derived from the Q-form. Places, where it took on a value less that 3.85 above the minimum, were considered inside the confidence interval.
  • primers and probes can be designed based on the information provided herein, in order to detect polymorphism by other means than sequencing, for example by PCR ampliation methods as also used herein.
  • Primers were purchased from DNA Technology (Aarhus, Denmark), probes for the lightcycler were obtained from MolBiol (Tvhofer Weg 11-12, Berlin) and the probes for Taqman were supplied by Applied Biosystem.
  • 5'-cta gag get cag tgt taa tct gtt cct ERCC1-3' 5' gga cag atg gca atg atg g g
  • 5'-cta gag get cag tgt taa tct gtt cct ERCC1-3' 5' gga cag atg gca atg atg g g
  • R is Roche buffer
  • H homebrew buffer
  • M mastermix
  • T denotes Tagman technique
  • L Lightcycler
  • A ABI2720 se ⁇ quencing, respectively.
  • R is Roche buffer: 10 pmole primers
  • H is Homebrew: 10 pmole primers 1 pmole probes
  • Polymorphisms by real-time PCR using Taqman probes. Polymor ⁇ phisms analyzed on Taqman were scored on the basis of the relative reaction of the 2 probes. Some polymorphisms were analysed using the ABI Prism 7700 sequence detection system (Applied Biosystems, Foster City, Ca, USA). PCR Primers and Taqman probes were designed using Primer Express v 1.0 (Applied Biosystems).
  • the reactions were performed in MicroAmp optical tubes sealed with MicroAmp op ⁇ tical caps (Applied Biosystems) containing a 10 ⁇ l reaction volume (Mastermix): 1x Taqman buffer A, 2.5mM MgCI 2 , 200 ⁇ M each of dATP dCTP, dGTP, 400 ⁇ M dUTP, 80OnM each primer, 200nm each probe, 0,01 U/ ⁇ L AmpErase UNG, 0,025 U/ ⁇ L AmpliTaq Gold Polymerase. Tubes were incubated at 50 0 C for 2 min followed by 10 min at 95 0 C. The incubation was succeeded by 45 cycles with thermal cycler condi ⁇ tions as shown in Table 13.
  • XPD exon 10 PCR was performed by rapid-cycling in a reaction vol ⁇ ume of 20 ⁇ l with 0.5 ⁇ M of each primer, 0.045 ⁇ M of anchor and sensor probe, 3.5 mM MgCI 2, approximately 7 - 25 ng genomic DNA, and 2 ⁇ l LightCycler DNA Master Hybridization probe buffer (Roche Molecular Biochemicals, Cat. No 2158 825).
  • This buffer contains Taq DNA polymerase, dNTP mix, and 10 mM MgC ⁇ and 5% DMSO.
  • the temperature cycling consisted of denaturation at 95 0 C for 2 sec, followed by 46 cycles consisting of 10 sec at 95 0 C, 15 sec at 53°C, and 30 sec at 72°C, see table 4a for PCR regimen for the polymorphisms using LightCyclerTM.
  • the last annealing pe- riod at 72°C was extended to 120 sec followed by denaturation for 10 sec at 95°C.
  • the melting profile was determined by a temperature ramp from 50 0 C to 95°C with a rate of 0.1 degree/sec.
  • the region consists of 2 major haplotype blocks with a few interspersed markers in between.
  • One haplotype block spans roughly 16 markers and 26 kb from XPD exon6 to RAI intron1-3, while the other haplotype block spans 6 markers and 9 kb from RAI-5' to ERCC1-3'3.
  • Fig. 9 depicts the relative risks and the lower confidence limits of the relative risks for postmenopausal breast cancer below age 55 as a function of the location of the markers. The values were initially calculated with the wild-type allele as reference, however, if the calculation gave a RR values less than 1 we used the reciprocal value of the RR and the reciprocal value of the high confidence limit in- stead.
  • haplotype trend regression is basically a two-stage procedure. First, genotype results in combination with population as- sumptions such as Hardy-Weinberg equilibrium are used to construct all haplotype probabilities corresponding to a given set of markers for each individual.
  • the disease state (1 for cases, 0 for controls) is regressed on the haplotype prob ⁇ abilities of all individuals, resulting in a p-value for the overall association of the set of markers with disease, and parameters for association for each specific haplotype with disease.
  • the HelixTree program will use sets of markers of defined size, typi ⁇ cally 2 to 4 neighboring markers at a time, to scan the entire region. It can then be used to derive frequencies of the individual haplotypes, to calculate an overall p- value for the distribution of haplotypes covering a given set of markers among cases and controls, and to calculate p-values for the distribution of each haplotype derived from a given set of markers.
  • Fig. 4 shows the overall distribution of p-values for sets of markers plotted against the position on the chromosome for breast cancer.
  • position on the abscissa we used the median marker position in a given set of markers.
  • the ordinate values are the negative logarithms to the overall p-values for a difference between cases and controls associated with a given set of markers.
  • Each curve corresponds to a given size of marker sets, i.e. the number of markers in the haplotypes.
  • Fig. 4 The results in Fig. 4 suggest that the association with breast cancer has two peaks. One was located at roughly 24 kb and one at 39 kb. When the cases were broken down into early and late breast cancer the peak at 24 kb was present in both groups, while the peak at 39 kb was only present in the older group. The peaks were clearly present in curves for haplotypes of 2, 3 and 4 neighboring markers. Larger sets were less informative, presumably due to increased degrees of freedom, but they essentially corroborated the results (results not shown).
  • this association at 24 kb was present for breast cancer both in young and in older female cases. This is the first time an association to RAI has been re ⁇ ported for cancers in older populations.
  • the position of the optimal set of markers corresponds to the 3' part of the gene RAI and the inter-gene region between RAI and XPD and was identical for young and older cancer cases.
  • the peak at 39 kb which is only present in the older breast cancers correspond to the 5' part of RAI.
  • haplotypes for the set of 2 markers, located at 39 kb. Again the deviation of young and old breast cancer cases from the controls had similarity (RAI intron1-2 9 RAI intron1-3 a up; RAI intron1-2 g RAI intron1-3° down). The same pattern was present in the lung cancer data. This is remarkable, as neither the young breast cancers nor the lung cancers showed a peak in this position, and sug- gests an underlying similarity of haplotypes in the different groups.
  • DNA was isolated from lymphocytes from humans from the American cohort of pa ⁇ tients with basal cell carcinoma and controls, described in Materials and Methods, was typed with respect to a number of sequence polymorphisms located in and around the claimed region r. The salient features of the Americant cohort are shown in Table 16.
  • Figure 7 shows the same kind of result for early basal cell carcinoma. Again we found a strong association with the region around RAI, with the curve for sets of 3 markers possibly shifted slightly to the right, and the curve for 4 markers showing suggestions of an extra association at higher positions.
  • DNA from humans from the Danish cohort of patients with lung cancer and controls was typed with respect to a number of sequence polymorphisms located in and around the claimed region r.
  • Fig. 7 shows that also this disease is associated with sets of markers span ⁇ ning the distal end of RAI and the inter-gene region in the population with the XPD exon23 AA genotype. This is the group of patients free of influence of the XPD gene. No association was evident in the two other groups (results not shown).
  • the ASE-1 genotype influences relapse-free survival and survival in Multiple Mye ⁇ loma 391 patients diagnosed with multiple myeloma were found eligible for transplantation in Denmark in the period 1993-2004.
  • the multiple myeloma patients were treated with high dose alkylating chemotherapy followed by autologous bone marrow transplantation.
  • the overall survival was sig- nificantly shorter for patients who experienced relapse than for patients who did not.
  • ASE-1 genotype was found to strongly influence the event-free survival (EFS) and overall survival in patients with multiple myeloma who were auto-transplanted.
  • EFS event-free survival
  • variant allele carries had 1.5 year longer event-free survival, see Fig. 9.
  • the EFS was simi ⁇ lar for men and women, although the difference was only statistically significant for women (Table 18). This means that the ASE-1 genotype can predict who will benefit most from the present treatment regimen.
  • Fig. 11 illustrates, female carriers of the variant allele of ASE-1 polymorphism had a longer relapse-free period than homozygous carriers of the wild-type allele (median 1479 days vs. 714 days), while Figure 12 illustrates that they lived longer (median 3015 vs. 1897). Similar analyses for the men revealed the same tenden ⁇ cies but no statistically significant differences (median 963 vs 736 days for the re ⁇ lapse-free period; mean 2747 vs. 2030 days for the survival), Fig. 13 and Fig. 14, respectively.
  • High-risk haplotype carriers (defined as ERCC1 exon 4 M , ASE1 exon1 GG , RAI In- tron 1 M ) had a mean survival of 359 days compared to a mean survival of 244 for non-carriers (Table 19 and Fig. 15).
  • the genotype XPD K751Q was also deter ⁇ mined.
  • the p-value is for a cox regression.

Landscapes

  • Chemical & Material Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Organic Chemistry (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Engineering & Computer Science (AREA)
  • Immunology (AREA)
  • Pathology (AREA)
  • Analytical Chemistry (AREA)
  • Zoology (AREA)
  • Genetics & Genomics (AREA)
  • Wood Science & Technology (AREA)
  • Physics & Mathematics (AREA)
  • Biotechnology (AREA)
  • Microbiology (AREA)
  • Molecular Biology (AREA)
  • Hospice & Palliative Care (AREA)
  • Biophysics (AREA)
  • Oncology (AREA)
  • Biochemistry (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Health & Medical Sciences (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Abstract

The present invention provides methods and composition for identifying human subjects with an increased risk of having or developing cancercancer. In particular, this invention relates to the identification and characterization of polymorphisms in the human chromosome 19q, the region r located approximately 19q, the region r located approximately 19q13.2-3 correlated with increased risk of developing cancer cancer and the responsiveness of a subject to various treatments for cancer. An allele in the r region can be identified as correlated with an increased risk of developing cancercancer, the prognosis of developed cancercancer, and responsiveness to cancer treatment, on the basis of the statistical analyses of the incidence of a particular allele in individuals diagnosed with cancercancer. The invention further relates to probes and kits comprising the probes useful in the diagnostic.

Description

Disease risk estimating method using sequence polymorphisms in a specific region of chromosome 19
The present invention provides methods and compositions for identifying human subjects with an increased risk of having or developing disease. In particular, this invention relates to the identification and characterization of polymorphisms in the human chromosome 19q, the region r located approximately 19q 13.2-3 correlated with increased risk of developing disease, in particular cancer and the responsive¬ ness of a subject to various treatments for cancer.
Background
DNA polymorphisms provide an efficient way to study the association of genes and diseases by analysis of linkage and linkage disequilibrum. With the sequencing of the human genome a myriad of hitherto unknown genetic polymorphisms among people have been detected. Most common among these are the single nucleotide polymorphisms, also called SNPs, of which several millions are known. Other ex¬ amples are variable number of tandem repeat polymorphisms, insertions, deletions and block modifications. Tandem repeats often have multiple different alleles (vari- ants), whereas the other groups of polymorphisms usually just have two alleles. Some of these genetic polymorphisms probably play a direct role in the biology of the individuals, including their risk of developing disease, but the virtue of the major¬ ity is that they can serve as markers for the surrounding DNA, and thus serve as leads during as search for a causative gene polymorphism, as substitutes in the evaluation of its role in health and disease, and as substitutes in the evaluation of the genetic constitution of individuals.
The association of an allele of one sequence polymorphism with particular alleles of other sequence polymorphisms in the surrounding DNA has two origins, known in the genetic field as linkage and linkage disequilibrium, respectively. Linkage arises because large parts of chromosomes are passed unchanged from parents to off¬ spring, so that minor regions of a chromosome tend to flow unchanged from one generation to the next and also to be similar in different branches of the same fam¬ ily. Linkage is gradually eroded by recombination occurring in the cells of the germ- line, but typically operates over multiple generations and distances of a number of million bases in the DNA.
Linkage disequilibrium deals with whole populations and has its origin in the (distant) forefather in whose DNA a new sequence polymorphism arose. The immediate sur¬ roundings in the DNA of the forefather will tend to stay with the new allele for many generations. Recombination and changes in the composition of the population will again erode the association, but the new allele and the alleles of any other polymor¬ phism nearby will often be partly associated among unrelated humans even today. A crude estimate suggests that alleles of sequence polymorphisms with distances less than 10000 bases in the DNA will have tended to stay together since modern man arose. Linkage disequilbrium in limited populations, for instance Europeans, often extends over longer distances. This can be the result of newer mutations, but can also be a consequence of one or more "bottlenecks" with small population sizes and considerable inbreeding in the history of the current population. Two obvious possi¬ bilities for "bottlenecks" in Europeans are the exodus from Africa and the repopula- tion of Europe after the last ice age.
Linkage disequilibrium is the results of many stochastic events and as such subject to statistical variation occasionally resulting in discontinuities, lack of a monotonic relationship between association and distance and differences between people of different ethnicity. Therefore, it is often advantageous to study more that one se¬ quence polymorphism in a given region. This also allows for further definition of the genetic surroundings of the biologically relevant polymorphism by combining the associated alleles of the different markers into a socalled haplotype.
Humans in general carry two copies of each human chromosome in each cell. There are exceptions to this rule, not relevant to this application. We therefore speak about genotypes i.e. the combined analysis of both chromosomes at a given sequence polymorphism. The resulting genotypes of a person, analysed for instance on DNA from peripheral blood leukocytes, are inherently very stable over time. Therefore, this type of analysis can be performed any time in the life of a person and will be applicable to this person for his or her entire life. By the same token such genetic analyses are ideally suited to predict future risks of disease. A variety of investigations suggest that many diseases in part are determined by the genetic constitution of the individual. One group of genes in particular has been as¬ sociated with rare genetic predispositions to cancer. These are the genes involved in maintaining the integrity of a person's DNA, the so-called DNA repair genes. One set of such genes are the XP genes which participate in nucleotide excision repair, and, when mutated, give rise to a 1000 fold increased risk of getting skin cancer. For this reason we have previously investigated single nucleotide polymorphisms in one DNA repair gene XPD for association with risk of skin cancer in a cohort of Cauca¬ sian Americans, and found that one allele of the sequence polymorphism called XPDeδ was associated with a moderately increased risk of getting basal cell carci¬ noma, the most common form of skin cancer. Later other groups have studied the association between sequence polymorphisms in this and other DNA repair genes and various forms of cancer. Some have reported positive results.
Very little is known about the function of the gene RAI. It was cloned because its protein product binds to and inhibits ReIA of the transcription regulator NF-kappaB. Other studies suggest that it may interact with the product p53 of a tumor suppres¬ sor gene.
Summary of the invention
The present invention relates in a first aspect to a group of nucleic acid sequences found to be associated with disease, in particular cancer. The invention further re¬ lates to transcriptional and translational products of said sequence. An allele in the r region can be identified as correlated with an increased risk of developing disease, in particular cancer, the prognosis of developed disease, in particular cancer, and responsiveness to disease treatment, in particular cancer treatment on the basis of statistical analyses of the incidence of a particular allele in individuals diagnosed with disease, in particular cancer.
Thus, in a first aspect the invention relates to a method for estimating the disease risk of an individual comprising
- in a sample from said individual assessing in the genetic material a se¬ quence polymorphism - in a region corresponding to SEQ ID NO: 1, or a part thereof, or - in a region complementary to SEQ ID NO: 1 , or a part thereof, or
- in a transcription product from a sequence in a region corresponding to SEQ ID NO: 1 , or a part thereof, or
- or translation product from a sequence in a region corresponding to SEQ ID NO: 1 , or a part thereof,
- obtaining a sequence polymorphism response,
- estimating the disease risk of said individual based on the sequence polymor¬ phism response.
The estimation of the disease risk of an individual can involve the comparison of the number and/or kind of polymorphic sequences identified with a predetermined dis¬ ease risk profile. Such a profile can be based on statistical data obtained for a rele¬ vant reference group of individuals. In particular the disease is a proliferative dis- ease, such as cancer.
The sequence of the r region is set forth as SEQ ID NO 1 , originating from the clon¬ ing of human chromosome 19q published as part of the contig NT_011109 in the database of human sequences established by National Center for Biotechnology Information and located on the internet at http://www.ncbi.nlm.nih.gov/- qenome/quide/human/ .
The presence of an allele is determined by determining the nucleic acid sequence of all or part of the region according to standard molecular biology protocols well known in the art as described for example in Sambrook et al. (1989) and as set forth in the Examples provided herein or products of the nucleic acid sequences.
In particular, the nucleic acid molecules of the present invention represent in a first aspect nucleic acid sequences forming part of the region r corresponding to position 1522-37752 of SEQ ID NO: 1 , and preferably to certain nucleic acid sequences within the gene referred to herein as RAI. As demonstrated in the Examples pre¬ sented below, the RAI gene is in particular associated with human cancer diseases.
Furthermore, the invention relates to a method for estimating the disease prognosis of an individual comprising - in a sample from said individual, assessing in the genetic material a sequence polymorphism
- in a region corresponding to SEQ ID NO: 1 , or a part thereof, or
- in a region complementary to SEQ ID NO: 1 , or a part thereof, or - in a transcription product from a sequence in a region corresponding to SEQ ID
NO: 1 , or a part thereof, or
- or translation product from a sequence in a region corresponding to SEQ ID NO: 1 , or a part thereof,
- obtaining a sequence polymorphism response, - estimating the disease prognosis of said individual based on the sequence polymorphism response.
The estimation of the disease prognosis of an individual can involve the comparison of the number and/or kind of polymorphic sequences identified with a predetermined disease prognosis profile. Such a profile can be based on statistical data obtained for a relevant reference group of individuals.
Additionally provided is a method of identifying a human subject as having an in¬ creased likelihood of responding to a treatment, comprising a) correlating the pres- ence of an r region allele genotype with an increased likelihood of responding to treatment; and b) determining the r region allele genotype of the subject, whereby a subject having an r region allele genotype correlated with an increased likelihood of responding to treatment is identified as having an increased likelihood of responding to treatment.
Thus, the present invention also relates to method for estimating a treatment re¬ sponse of an individual suffering from disease to a disease treatment, comprising
- in a sample from said individual, assessing in the genetic material a sequence polymorphism - in a region corresponding to SEQ ID NO: 1 , or a part thereof, or
- in a region complementary to SEQ ID NO: 1 , or a part thereof, or
- in a transcription product from a sequence in a region corresponding to SEQ ID NO: 1 , or a part thereof, or
- or translation product from a sequence in a region corresponding to SEQ ID NO: 1 , or a part thereof, - obtaining a sequence polymorphism response,
- estimating the individual's response to the disease treatment based on the sequence polymorphism response.
The estimation of the individual's response to disease treatment can involve the comparison of the number and/or kind of polymorphic sequences identified with a predetermined cancer treatment response profile. Such a profile can be based on statistical data obtained for a relevant reference group of individuals. In particular the disease is a proliferative disease, such as cancer.
The invention also comprises primers or probes for use in the invention, as well as kits including these. The primers and/or probes are preferably capable of hybridising to SEQ ID NO:1, or a part thereor, in particularly the regions relevant to this inven¬ tion, or a part thereof, under stringent conditions, as well as to a sequence comple- mentary thereto.
Furthermore, the invention also relates to cloning vectors and expression vectors containing the nucleic acid molecules of the invention, as well as hosts which have been transformed with such nucleic acid molecules, including cells genetically engi- neered to contain the nucleic acid molecules of the invention, and/or cells geneti¬ cally engineered to express the nucleic acid molecules of the invention. The nucleic acids are preferably isolated from the r region and preferably contain one or more sequence polymorphisms as described herein below in more detail. In addition to host cells and cell lines, hosts also include transgenic non-human animals (or prog- eny thereof).
In particular, the present invention is based on the discovery of the correlation with single nucleotide polymorphisms (SNPs), deletion polymorphisms, insertion poly¬ morphisms, dinucleotide polymorphisms and/or tandem repeats in the regions and disease. Thus, polymorphisms have been found in the r region as shown in table 1. However, the present invention is not limited to the polymorphisms shown in table 1 , but does include any polymorphism in the region. Deletions and insertions have been found in the r region as shown in table 1. However, the present invention is not limited to the deletions shown in table 1 , but does include any deletion and insertion in the region. The term human includes both a human having or suspected of having a disease and an a-symptomatic human who may be tested for predisposition or susceptibility to disease. At each position the human may be homozygous for an allele or the hu- man may be a heterozygote.
Drawings
Fig. 1 shows an overview of a subregion of chromosome 19q in which the exon- intron structure of RAI, XPD and ASE-1 is shown. Letters correspond to the specific marker position. The position in NT_011109 is shown. Marker abbreviations are as follows: a=XPDexon10, b=XPD exonδ, c=XPD-4bp, d=XPD-81bp, e=XPD-5'2, f=XPD5'3, g=RAI-3'7, h=RAI-3'3, i=RAI-3'6, j=RAI-3'5, k=RAI-3'4, I=RAI-3'8, m=RAI exonβ, n=RAI intron5-2, O=RAI intron3, p=RAI-intron1, q=RAI intron1-2, r=RAI in- tron1-3, s=RAI-5'2, t=RAI-5'3, u=RAI-5', V=ASEI -5'2, w= ASE1 exoni , X=ERCCI -3'
Fig. 2 shows a pair-wise linkage disequilibrium between all markers in breast cancer controls.
Fig. 3 shows an association of single polymorphisms with early breast cancer.
Fig. 4 shows an association of sets of neighboring SNPs with breast cancer.
Fig. 5 shows a maximal OR for breast cancer for haplotypes formed by two SNPs.
Fig. 6 shows a Lazzeroni estimation of the position of the causative variant using young breast cancer.
Fig. 7 shows the overall distribution of p-values for sets of markers plotted against the position on the chromosome for basal cell carcinoma and lung cancer. As posi¬ tion on the abscissa we used the median marker position in a given cluster of mark¬ ers. The ordinate values are the negative logarithms to the overall p-values for a difference between cases and controls associated with a given set of markers. Each curve corresponds to a given size of marker sets, i.e. the length of the haplotypes. Fig. 8 shows odds ratio for cancer versus control between the two homozygotes of each SNP in relation to location on chromosome 19.
Fig. 9 shows event free survival for patients who are wild-type carriers of ASE-1 (0) or homozygous or heterozygous carriers of the variant allele (1 ).
Fig. 10 shows overall survival subdivided by ASE-1 genotype. 0=homozygous car¬ rier of the wild-type allele of ASE-1 , 1 = carrier of the variant allele of ASE-1.
Fig. 11 shows event free survival for women subdivided by ASE-1 genotype.
0=homozygous carrier of the wild-type allele of ASE-1 , 1= carrier of the variant allele of ASE-1.
Fig. 12 shows overall survival for women subdivided by ASE-1 genotype. 0=homozygous carrier of the wild-type allele of ASE-1 , 1 = carrier of the variant allele of ASE-1
Fig. 13 shows event free survival for men subdivided by ASE-1 genotype. 0=homozygous carrier of the wild-type allele of ASE-1 , 1 = carrier of the variant allele of ASE-1.
Fig. 14 shows overall survival for men subdivided by ASE-1 genotype. 0=homozygous carrier of the wild-type allele of ASE-1 , 1= carrier of the variant allele of ASE-1.
Fig. 15 shows Kaplan Meier plot of survival of lung cancer patients in relation to highrisk haplotype
Fig. 16 shows Kaplan Meier plot of survival in relation to XPD K751Q among those lung cancer patients homozygous for the high risk haplotype.
Fig. 17 shows Kaplan Meier plot of survival in relation to XPD K751Q among those lung cancer patients not homozygous for the highrisk haplotype. Detailed description of the invention
The present invention relates to a characterization of a person's present and/or fu¬ ture risk of getting certain forms of disease, in particular a proliferative disease, such as cancer. The characterization is based on the analysis of sequence polymor¬ phisms in a region of chromosome 19q in the person.
A number of polymorphisms in the chromosomal region 19q 13.2-3 have been identi¬ fied and characterised. Surprisingly, the sequence polymorphisms with strongest association to disease appeared to be located outside the gene XPD. More specifi¬ cally, the sequences were located in a sub-region harboring the gene RAI and im¬ mediately downstream of the RAI gene. The nature of the association between hap- lotype and disease was examined together with the p-values associated with the individual haplotypes. The odds ratios for each haplotype of each set of three neighboring SNPs were also determined. The odds ratio for test of homozygotes of individual markers against cancer status was likewise detemined. The most likely location of a single causative gene variant for individuals of the breast cancer cases was evaluated.
The region of chromosome 19q, more precisely the region located in 19q13.2-3, with which the present invention is concerned, is depicted in Figure 1 as it is presently known together with the presently known or suspected genes. The arrows indicate the directions of transcription of the genes. The absolute chromosome positions shown are from the particular build of NCBPs map of chromsome 19, and will proba- bly change with time. The position of markers used throughout the experiments is indicated. The intron-exon structure of the RAI gene is shown together with the posi¬ tion of the inter gene region between RAI and XPD genes.
The region r stretches from the beginning of, but not including the XPD gene, to ap- proximately the end of ERCC1 and includes the genes RAI, LOC162978, and ASE- 1. More specifically r is bounded by and includes the following two sequences: AGAACCCCCG CCCCTCCACC TCGTCTCAAA and TCCCTCCCCA GA- GACTGCAC CAGCGCAGCC, and is defined by SEQ ID NO: 1.
In the present context the region r means SEQ ID NO: 1 and complementary se¬ quence as well as transcriptional products and translational products thereof. In one preferred embodiment, the gene RAI is defined in the claims as including transcribed sequences of the gene plus a 1500 base upstream promoter region. More specifically RAI is bounded by and includes the following sequences: CATAACCACA ATGATGAGCA TGTATTGAGT and ATGTTGTCCA GGCTGGTCTT GAACTCCTGA. In the present context this section of the region relates to SEQ ID NO: 1 bases 7761-22885 and complementary sequence as well as transcriptional products and translational products thereof.
In another embodiment, one preferred section of the region stretches approximately from the the beginning of, but not including the XPD gene, to approximately the end of the RAI gene. In the present context the region means SEQ ID NO: 1 bases 1 to 25550 and complementary sequence as well as transcriptional products and transla¬ tional products thereof.
In an even more preferred embodiment, one preferred section of the region stretches approximately from the the beginning of, but not including the XPD gene, to approximately within the RAI gene. In the present context the region means SEQ ID NO: 1 bases 1 to 15698 and complementary sequence as well as transcriptional products and translational products thereof.
In a most preferred embodiment, one preferred section of the region r stretches ap¬ proximately from outside the XPD gene, to approximately within the RAI gene. In the present context the region means SEQ ID NO: 1 bases 4528 to 15698 and comple- mentary sequence as well as transcriptional products and translational products thereof.
In a third embodiment, one preferred section of the region r stretches approximately from the beginning of, but not including the XPD gene, into the inter gene region between the RAI and XPD gene. In the present context the region means SEQ ID NO: 1 bases 1 to 1510 and complementary sequence as well as transcriptional products and translational products thereof.
In a further embodiment, one preferred section of the region r stretches approxi- mately from the the beginning of, but not including the XPD gene, throughout the inter gene region between the RAI and XPD gene and into the 3' part of the RAI gene. In the present context the region means SEQ ID NO: 1 bases 1710 to 8685 and complementary sequence as well as transcriptional products and translational products thereof.
In yet a further embodiment, one preferred section of the region r stretches ap¬ proximately over the central part of the RAI gene. In the present context the region means SEQ ID NO: 1 bases 8987 to 12090 and complementary sequence as well as transcriptional products and translational products thereof.
In a final embodiment, one preferred section of the region r stretches approximately over the middle and 5' part of the RAI gene. In the present context the region means SEQ ID NO: 1 bases 15898 to 25550 and complementary sequence as well as tran¬ scriptional products and translational products thereof.
Modifications to the human genome map are known to occur from time to time. It is therefore possible that the defining sequences quoted above will change slightly in future maps.
Fragments or parts of the region r as used herein relates to any fragment of at least 5 nucleic acid redues in length, or multiples of 5 nucleic acid residues in length start¬ ing from SEQ ID NO: 1 position 1 , 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100. For example at least 21 , such as at least 22, for example at least 23, such as at least 24, for example at least 26, such as at least 27, for ex- ample at least 28, such as at least 29, for example at least 31 , such as at least 32, for example at least 33, such as at least 34, for example at least 36, such as at least 37, for example at least 38, such as at least 39, for example at least 41 , such as at least 42, for example at least 43, such as at least 44, for example at least 46, such as at least 47, for example at least 48, or at least 100 nucleic acid redues in length, or mutiples of 100 nucleic acid residues in length, starting from SEQ ID NO: 1 posi¬ tion 1, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2600, 2700, 2800, 2900, 3000, and so forth, each fragment starting position having an increment of 100 nucleic acid residues. Multiples are preferably multiples of e.g. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 , 42, 43, 44, 45, 46, 47, 48, 49 and 50.
For fragments starting at position 1, the length of said fragments will thus be e.g. 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500,
1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2600, 2700, 2800, 2900, 3000, and so forth, using suitable multiplicators as listed herein above.
For fragments starting at position 100, the length of said fragments will thus be e.g. 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500,
1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2600, 2700, 2800, 2900, 3000, and so forth, using suitable multiplicators as listed herein above.
For fragments starting at position 7700, the length of said fragments will thus be e.g. 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500,
1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2600, 2700, 2800, 2900, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10000, 10500, 11000, 11500, 12000, 12500, 13000, 13500, 14000, 14500, 15000, and so forth, using suitable multiplicators such as e.g. the ones listed herein above.
The nucleic acid sequences according to the present invention make it possible to estimate cancer risk in an individual by using sequence polymorphisms originating from a specific region of chromosome 19.
Estimation of disease risks has a number of important applications, which in the following is exemplified with respect to cancer, but also apply to other disease, as described herein:
(1) Individuals with reasons to suspect that they are at risk for getting cancer would be able to clarify their situation and, if possible, take protective action. Alternatively, anti-cancer campaigns, companies, hospitals or other institutions could offer a ser¬ vice to help people clarify their situation. It would for instance be possible to test persons, when they got their first basal cell carcinoma, which is often recurrent and also is a moderate predictor for other cancers. If the persons were in a high-risk group, one could then advice them about, or they could of their own accord choose, risk-reducing behaviour, such as avoidance of excessive sun-exposure, abstaining from smoking etc. About 5 percent of the Danish population will at some point in their life get a basal cell carcinoma.
(2) Anti-cancer campaigns, companies, hospitals or other institutions would be able to define relevant target subpopulations and focus information on risk-reducing be¬ haviour on these persons. They might perhaps also be in a position to inform the remainder of the population that they need not worry. Lung cancer affects approxi- mately 10-15 percent of smokers and thus approximately 5 percent of the popula¬ tion, somewhat varying from country to country. Malignant melanoma, a sun- induced, often lethal form of skin cancer, affects approximately 700 persons a year in Denmark or about 1 percent of the Danish population.
(3) The drugs used in cancer treatment are often carcinogenic themselves and indi¬ vidual responses to them vary considerably, both with respect to tolerance to the treatment and with respect to efficacy of the treatment. It is an obvious possibility that the region of chromosome 19 here dealt with, which contains DNA repair genes known to modulate carcinogen responses, also modulates response to anti-cancer agents. Hence, analysis of the region may facilitate better choices of treatment for cancer, and/or help predict the future course of disease.
By sequence polymorphism is understood any single nucleotide, tandem repeat, insertion, deletion or block polymorphism, which varies among humans, whether it is of known biological importance or not.
Position of sequence polymorphism in the region r
In one embodiment of the methods of the invention, preferably the method for diag- nosis as described herein, one or more single nucleotide polymorphism(s) at a pre¬ determined position in the region (SEQ ID NO:1) are identified and used for e.g. cancer risk profiling and/or cancer treatment response profiling. Presently preferred polymorphism(s) are listed in Tables 1a, 1b and 1c, more preferably at least two polymorphism(s) are selected, most preferably at least three polymorphism(s) are selected. However, the present invention relates to any polymorphism in the region. Table 1a
Trivial name Kind dbSNP # Sequence Position Relative Position in in Position SEQ ID sequence NO:1
XPD-81bp 81bp/- rs3916787 NT 011109 18143027 19890 632
XPD-5'2 AJG rs2097215 NT 011109 18144005 20868 1610
XPD-5'3 C/T rs11878644 NT 011109 18145185 22048 2790
RAl-3'7 C/T rs7252567 NT 011109 18146823 23686 4428
RAI-3'3 AA/- rs3047560 NT 011109 18147192 24055 4797
RAI-3'6 AJC 10422489 NT 011109 18147886 24749 5491
RAI-3'5 C/T 10426701 NT 011109 18148193 25056 5798
RAI-3'4 A/G rs4544343 NT 011109 18150199 27062 7804
RAI-3'8 AJG rs8101662 NT 011109 18150911 27774 8516
RAI exon6 A/T rs6966 NT 011109 18151180 28043 8785
RAI-iδ-2 A/T rs8112723 NT 011109 18153497 30360 10204
RAI intron 3 A/G rs2017104 NT 011109 18155483 32346 12190
RAI intron 1 A/G rs1970764 NT 011109 18159091 35954 15798
RAI intron 1-2 A/G AC092309 24351 39069 18913
RAI intron 1-3 A/C AC092309 25115 39833 19677
RAI-5'2 C/T rs4803814 NT 011109 18168944 45807 25650
RAI-5'3 C/T rs4803815 NT 011109 18168948 45811 25564
RAI-5' C/T rs4572514 NT 011109 18171984 48847 28691
ASE1-5'2 G/T rs2226949 NT 011109 18175327 52190 32035
ASE1 exon 1 A/G rs967591 NT 011109 18178152 55015 34854
ERCC1-3' C/T rs762562 NT 011109 18180561 57424 37267 rs numbers were derived from the NCBI's database dbSNP.
Table 1b in SEQ ID NO: 1 rs3916787 gctgcagtga gctgt (-/ACACCTGTGGTCCCAGCTACTCTGG 632 AAGCTGAGGTGGGAGGATCGCTTGAGCCCAAGAGGTG- GAGGCTGCAGTGAGCTGT) gactgtgcca ctgcactcca rs2097215 TG ACAGTAG A CATCCTGTCA T (A/G) ATAAGTCttt ttttttt 1610 rs11878644 CATCCCCATA CCAAcccacc (c/t) tactgctctg atctcctcct 2790 rs7252567 ttttagtagagacatggttccgcca (C/T) gttgcccaggctggtcttgaactc 4428 rs3047560 ACTAAAAATAAAAAATAAAAAAAA(-/AA) ATAGCCGAG- 4797
CATGGTGGTGGGGTGC rs 10422489 AAAAAAC- 5491
TAAAGTGGGGTTTGCGGG(G/T)AGTGGGAGGGCCCTTC
CTGCTAGG rs10426701 ggatcacaag gtcaggagtt (c/t) gagaccagcc tggccaacac 5798 rs4544343 TGTTGTCCAA GCTGGCAGAG (A/G) TTTTTGTTTG 7804
TTTGTTTGAG rs8101662 CTCGGGAGGCTGAGGCAGGAGAATC 8516
(A/G)CTTGAACTCAGGCAGAGGTTG rs6966 ATTAAGTGCCTTCACACAGC (A/T) CTGGTTTAAT GTTTATAA 8785 rs8112723 ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa 10204 rs2017104 gggaggctcg aggcgggc (A/G) gattgcatga gctcaggatt 12190 rs1970764 tgcagtgagc tgagatcgc (A/G) ccactgcact ccagcctggg 15798
AC092309 CTTGCTACAGAATTACAGGCA (T/C) 18913
GCGCCACCGCTCCGGGCTAA AC092309 CTAAAGACTACA (-/A) TTTCCCAGCATCCCATTG 19677 rs4803814 CCCTCCCTGCTTGCTTGCTTTCTCTtC/ηTCTCTCTTTCTT 25650
TCTTTCTTTCTT rs4803815 CCCTGCTTGCTTGCTTTCTCTCTCT[C/ηTCTTTCTTTCTT 25564
TCTTTCTTTCTT rs4572514 AGAACCTGTTCAGGCTGGCGGCTCA[C/T]TTGGATGAAC 28691
AGGGAGTGTGTGAC rs2226949 CCCCCTTCTTAGGACG- 32035
CATGGGGGTfG/ηGAGAGAACGGGGAGATAGACAGAG rs967591 TGCGAGCAGCCCGGGCTA- 34854
CAGGGTT[A/G]CCTGAGGTGTGGGTCCCAGGATGG rs762562 GGCGCCTCAACAGCCAGAAG- 37267
GAGCG[A/G]AGCCTCAGGCCCAGGCAGCTCTGG
Table 1c
Trivial name Kind dbSNP # Sequence Position Relative in sequence position
XPDexon23 A/C Rs1052559 NT 011109 18123137 0
XPDexon 10 A/G Rs1799793 NT 011109 18135477 12340
XPDexon6 A/C Rs238406 NT 011109 18136527 13390
XPDintron3 A/G Rs1799783 NT 011109 18140254 17117
XPD-4bp GACAJ- Rs3916791 NT 011109 18142345 19208
XPD-81bp 81bp/- Rs3916787 NT 011109 18143027 19890
XPD-5'2 A/G Rs2097215 NT 011109 18144005 20868
XPD-5'3 C/T Rs11878644 NT 011109 18145185 22048
RAI-3'7 C/T Rs7252567 NT 011109 18146823 23686
RAI-3'3 AA/- RS3047560 NT 011109 18147192 24055
RAI-3'4 A/G Rs4544343 NT 011109 18150199 27062
RAIexon6 A/T Rs6966 NT 011109 18151180 28043
RAIintron5-3 C/T Rs10417235 NT 011109 18152171 29034
RAIintron5-2 A/T Rs8112723 NT 011109 18153497 30360
RAIintron3 A/G RS2017104 NT 011109 18155483 32346
RAIintron 1 A/G Rs1970764 NT 011109 18159091 35954
RAIintron 1-2 C/T RS6509210 NT 011109 18162206 39069
RAIintron1-3 A/C NT 011109 18162970 39833
RAI-5'2 C/T RS4803814 NT 011109 18168944 45807
RAI-5'3 C/T RS4803815 NT 011109 18168948 45811
RAIS' C/T RS4572514 NT 011109 18171984 48847
ASE1-5'2 G/T RS2226949 NT 011109 18175327 52190
ASE1 exon 1 A/G RS967591 NT 011109 18178152 55015
ERCC1-3'2 A/G RS735482 NT 011109 18180220 57083
ERCC1-3' C/T RS762562 NT 011109 18180561 57424
ERCC1-3'3 A/G RS2336219 NT 011109 18180624 57487
ERCC1 exon4 C/T Rs3177700 NT 011109 18191871 68734
In one embodiment of the methods of the invention, preferably the method for diag¬ nosis as described herein, one or more single nucleotide polymorphism(s) at a pre¬ determined position in the region (SEQ ID NO:1) are identified and used for e.g. cancer risk profiling and/or cancer treatment response profiling. Presently preferred polymorphism(s) are listed in Tables 2a and 2b, more preferably at least two poly- morphism(s) are selected, most preferably at least three polymorphism(s) are se¬ lected. However, the present invention relates to any polymorphism in the region.
Table 2a
Trivial name Kind dbSNP # Sequence Position Relative Position in in Position SEQ ID sequence NO:1
RAI-3'4 A/G rs4544343 NT 011109 18150199 27062 7804
RAI-3'8 A/G rs8101662 NT 011109 18150911 27774 8516
RAI exon 6 A/T rs6966 NT 011109 18151180 28043 8785
RAI-I5-2 A/T rs8112723 NT 011109 18153497 30360 10204
RAI intron 3 A/G rs2017104 NT 011109 18155483 32346 12190
RAI intron 1 A/G rs 1970764 NT 011109 18159091 35954 15798
RAI intron 1-2 A/G AC092309 24351 39069 18913
RAI intron 1-3 A/C AC092309 25115 39833 19677 rs numbers were derived from the NCBI's database dbSNP.
Table 2b
NO: 1 rs4544343 TGTTGTCCAA GCTGGCAGAG (A/G) TmTGTTTG 7804
TTTGTTTGAG rs8101662 CTCGGGAGGCTGAGGCAGGAGAATC 8516
(A/G)CTTGAACTCAGGCAGAGGTTG rs6966 ATTAAGTGCCTTCACACAGC (A/T) CTGGTTTAAT GTTTATAA 8785 rs8112723 ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa 10204 rs2017104 gggaggctcg aggcgggc (A/G) gattgcatga gctcaggatt 12190 rs1970764 tgcagtgagc tgagatcgc (A/G) ccactgcact ccagcctggg 15798
AC092309 CTTGCTACAGMTTACAGGCA (T/C) 18913
GCGCCACCGCTCCGGGCTAA
AC092309 CTAAAGACTACA (-/A) TTTCCCAGCATCCCATTG 19677
In one embodiment of the methods of the invention, preferably the method for diag- nosis as described herein, one or more single nucleotide polymorphism(s) at a pre¬ determined position in the region (SEQ ID NO:1) are identified and used for e.g. cancer risk profiling and/or cancer treatment response profiling. Presently preferred polymorphism(s) are listed in Tables 3a, 3b and 3c, more preferably at least two polymorphism(s) are selected, most preferably at least three polymorphism(s) are selected. However, the present invention relates to any polymorphism in the region.
Table 3a
Trivial name Kind dbSNP # Sequence Position Relative Position in in Position SEQ ID sequence NO:1
XPD-81bp 81 bp/- Rs3916787 NT 011109 18143027 19890 632 XPD-5'2 A/G Rs2097215 NT 011109 18144005 20868 1610
XPD-5'3 C/T Rs1187864 NT_011109 18145185 22048 2790
4
RAI-3'7 C/T RS7252567 NT 011109 18146823 23686 4428
RAl-3'3 AA/- RS3047560 NT 011109 18147192 24055 4797
RAI-3'6 A/C 10422489 NT 011109 18147886 24749 5491
RAI-3'5 C/T 10426701 NT 011109 18148193 25056 5798
RA!-3'4 A/G RS4544343 NT 011109 18150199 27062 7804
RAI-3'8 A/G Rs8101662 NT 011109 18150911 27774 8516
RAI exon 6 A/T Rs6966 NT 011109 18151180 28043 8785
RAN5-2 A/T Rs8112723 NT 011109 18153497 30360 10204
RAI intron 3 A/G Rs2017104 NT 011109 18155483 32346 12190
RAI intron 1 A/G Rs1970764 NT 011109 18159091 35954 15798
RAI intron 1-2 A/G AC092309 24351 39069 18913
RAI intron 1-3 A/C AC092309 25115 39833 19677 rs numbers were derived from the NCBI's database dbSNP.
Table 3b in SEQ ID NO: 1 rs3916787 gctgcagtga gctgt (-/ACACCTGTGGTCCCAGCTACTCTGG 632 AAGCTGAGGTGGGAGGATCGCTTGAGCCCAAGAGGTG- GAGGCTGCAGTGAGCTGT) gactgtgcca ctgcactcca rs2097215 TGACAGTAGA CATCCTGTCA T (A/G) ATAAGTCttt ttttttt 1610 rs 11878644 CATCCCCATA CCAAcccacc (c/t) tactgctctg atctcctcct 2790 rs7252567 ttttagtagagacatggttccgcca (C/T) gttgcccaggctggtcttgaactc 4428 rs3047560 ACTAAAAATAAAAAATAAAAAAAA(-/AA) ATAGCCGAG- 4797
CATGGTGGTGGGGTGC rs10422489 AAAAAAC- 5491
TAAAGTGGGGTTTGCGGG(G/T)AGTGGGAGGGCCCTTC
CTGCTAGG rs 10426701 ggatcacaag gtcaggagtt (c/t) gagaccagcc tggccaacac 5798 rs4544343 TGTTGTCCAA GCTGGCAGAG (A/G) TTTTTGTTTG 7804
TTTGTTTGAG rs8101662 CTCGGGAGGCTGAGGCAGGAGAATC 8516
(A/G)CTTGAACTCAGGCAGAGGTTG rs6966 ATTAAGTGCCTTCACACAGC (A/T) CTGGTTTAAT GTTTATAA 8785 rs8112723 ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa 10204 rs2017104 gggaggctcg aggcgggc (A/G) gattgcatga gctcaggatt 12190 rs1970764 tgcagtgagc tgagatcgc (A/G) ccactgcact ccagcctggg 15798
AC092309 CTTGCTACAGAATTACAGGCA (T/C) 18913
GCGCCACCGCTCCGGGCTAA
AC092309 CTAAAGACTACA (-/A) TTTCCCAGCATCCCATTG 19677
More preferably polymorphism(s) are listed below in tables 3c and 3d, more pref¬ erably at least two polymorphism(s) are selected, most preferably at least three polymorphism(s) are selected from SEQ ID NO:1 below Table 3c
Trivial name Kind dbSNP # Sequence Position Relative Position in in Position SEQ ID sequence NO:1
XPD-81 bp 81 bp/- Rs3916787 NT 011109 18143027 19890 632
XPD-5'2 A/G Rs2097215 NT 011109 18144005 20868 1610
XPD-5'3 C/T Rs 1187864 NT_011109 18145185 22048 2790
4
RAI-3'7 C/T rs7252567 NT 011109 18146823 23686 4428
RAI-3'3 AA/- rs3047560 NT 011109 18147192 24055 4797
RAI-3'6 A/C 10422489 NT 011109 18147886 24749 5491
RAI-3'5 C/T 10426701 NT 011109 18148193 25056 5798
RAI-3'4 A/G rs4544343 NT 011109 18150199 27062 7804
RAI-3'8 A/G rs8101662 NT 011109 18150911 27774 8516
RAI exon 6 A/T rs6966 NT 011109 18151180 28043 8785
RAl-iδ-2 A/T rs8112723 Nf~011109 18153497 30360 10204
RAI intron 3 A/G rs2017104 NT 011109 18155483 32346 12190 rs numbers were derived from the NCBI's database dbSNP.
Table 3d
GAGGCTGCAGTGAGCTGT) gactgtgcca ctgcactcca rs2097215 TGACAGTAGA CATCCTGTCA T (A/G) ATAAGTCttt ttttttt 1610 rs 11878644 CATCCCCATA CCAAcccacc (c/t) tactgctctg atctcctcct 2790 rs7252567 ttttagtagagacatggttccgcca (C/T) gttgcccaggctggtcttgaactc 4428 rs3047560 ACTAAAAATAAAAAATAAAAAAAA(-/AA) ATAGCCGAG- 4797
CATGGTGGTGGGGTGC rs 10422489 AAAAAAC- 5491
TAAAGTGGGGTTTGCGGG(G/T)AGTGGGAGGGCCCTTC
CTGCTAGG rs 10426701 ggatcacaag gtcaggagtt (c/t) gagaccagcc tggccaacac 5798 rs4544343 TGTTGTCCAA GCTGGCAGAG (A/G) I I I I I GTTTG 7804
TTTGTTTGAG rs8101662 CTCGGGAGGCTGAGGCAGGAGAATC 8516
(A/G)CTTGAACTCAGGCAGAGGTTG rs6966 ATTAAGTGCCTTCACACAGC (ATT) CTGGTTTAAT GTTTATAA 8785 rs8112723 ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa 10204 rs2017104 gggaggctcg aggcgggc (A/G) gattgcatga gctcaggatt 12190
Even more preferably polymorphism(s) are those listed below in tables 3e and 3f, more preferably at least two polymorphism(s) are selected, most preferably at least three polymorphism(s) are selected from SEQ ID NO: 1 below:
Table 3e
Trivial name Kind dbSNP # Sequence Position Relative Position in in Position SEQ ID sequence NO:1
RAI-3'3 AA/- rs3047560 NT 011109 18147192 24055 4797
RAI-3'6 A/C 10422489 NT 011109 18147886 24749 5491
RAI-3'5 CfT 10426701 NT 011109 18148193 25056 5798
RAI-3'4 A/G rs4544343 NT 011109 18150199 27062 7804
RAI-3'8 A/G rs8101662 NT 011109 18150911 27774 8516
RAI exon 6 A/T rs6966 NT 011109 18151180 28043 8785
RAI-15-2 A/T rsδ 112723 NT 011109 18153497 30360 10204
RAI intron 3 A/G rs2017104 NT 011109 18155483 32346 12190 rs numbers were derived from the NCBI's database dbSNP.
Table 3f
rs10422489 AAAAAAC- 5491
TAAAGTGGGGTTTGCGGG(G/T)AGTGGGAGGGCCCTTC
CTGCTAGG rs10426701 ggatcacaag gtcaggagtt (c/t) gagaccagcc tggccaacac 5798 rs4544343 TGTTGTCCAA GCTGGCAGAG (A/G) TTTTTGTTTG 7804
TTTGTTTGAG rs8101662 CTCGGGAGGCTGAGGCAGGAGAATC 8516
(A/G)CTTGAACTCAGGCAGAGGTTG rs6966 ATTAAGTGCCTTCACACAGC (A/T) CTGGTTTAAT GTTTATAA 8785 rs8112723 ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa 10204 rs2017104 gggaggctcg aggcgggc (A/G) gattgcatga gctcaggatt 12190
Most preferably polymorphism(s) are those listed in tables 4a and 4b below, more preferably at least two polymorphism(s) are selected, most preferably at least three polymorphism(s) are selected from the polymorphisms shown below:
Table 4a
Trivial name Kind dbSNP # Sequence Position Relative in SEQ ID 1 Position
RAI-3'3 AAA rs3047560 NT 011109 24055 RAI exon 6 A/T rs6966 NT 011109 28043 RAI intron 3 A/G rs2017104 NT 011109 32346 rs numbers were derived from the NCBI's database dbSNP.
Table 4b
rs6966 ATTAAGTGCCTTCACACAGC (A/T) CTGGTTTAAT GTTTATAA 8785 rs2017104 gggaggctcg aggcgggc (AJG) gattgcatga gctcaggatt 12190
In yet another embodiment according to the methods of the invention, preferably the method for diagnosis as described herein, one or more single nucleotide polymor¬ phism^) at a predetermined position in the region (SEQ ID NO:1) are identified and used for e.g. cancer risk profiling and/or cancer treatment response profiling. In this embodiment the preferred polymorphism(s) are listed in Tables 5a, more preferably at least two polymorphism(s) are selected, most preferably at least three polymor- phism(s) are selected. However, the present invention relates to any polymorphism in the region.
Table 5a
Trivial name Kind dbSNP# Sequence Position Relative Position in in Position SEQID sequence NO:1
XPD-81bp 81bp/- rs3916787 NT 011109 18143027 19890 632
XPD-5'3 C/T rs11878644 NT 011109 18145185 22048 2790
RAI-37 C/T rs7252567 NT 011109 18146823 23686 4428
RAI-3'3 AA/- rs3047560 NT 011109 18147192 24055 4797
RAI-3'6 A/C 10422489 NT 011109 18147886 24749 5491
RAI-3'5 C/T 10426701 NT 011109 18148193 25056 5798
RAI-3'4 A/G rs4544343 NT 011109 18150199 27062 7804
RAI-3'8 A/G rs8101662 NT 011109 18150911 27774 8516
RAI-iδ-2 A/T rs8112723 NT 011109 18153497 30360 10204
RAI intron 1-2 A/G AC092309 24351 39069 18913
RAI intron 1-3 A/C AC092309 25115 39833 19677 rs numbers were derived from the NCBI's database dbSNP.
More preferably, polymorphism(s) are those listed in table 5b below, more preferably at least two polymorphism(s) are selected, most preferably at least three polymor¬ phism^) are selected from the polymorphisms shown below:
Table 5b
Trivial name Kind dbSNP# Sequence Position Relative Position in in Position SEQID sequence NO:1
XPD-5'3 C/T rs11878644 NT 011109 18145185 22048 2790
RAI-3'7 C/T rs7252567 NT 011109 18146823 23686 4428
RAI-3'3 AA/- rs3047560 NT 011109 18147192 24055 4797
RAI-3'6 A/C 10422489 NT 011109 18147886 24749 5491
RAI-3'5 C/T 10426701 NT 011109 18148193 25056 5798
RAI-3'4 A/G rs4544343 NT 011109 18150199 27062 7804
RAI-3'8 A/G rs8101662 NT 011109 18150911 27774 8516
RAI-i5-2 A/T rs8112723 NT 011109 18153497 30360 10204
RAIintron 1-2 A/G AC092309 24351 39069 18913 RAl intron 1-3 A/C AC092309 25115 39833 19677 rs numbers were derived from the NCBI's database dbSNP.
In an even more preferred embodiment polymorphism(s) are those listed in table 5c below, more preferably at least two polymorphism(s) are selected, most preferably at least three polymorphism(s) are selected from the polymorphisms shown below:
Table 5c
Trivial name Kind dbSNP# Sequence Position Relative Position in in Position SEQlD sequence NO:1
XPD-5'3 C/T rs11878644 NT 011109 18145185 22048 2790
RAI-3'7 C/T rs7252567 NT 011109 18146823 23686 4428
RAI-3'3 AA/- rs3047560 NT 011109 18147192 24055 4797
RAI-3'6 A/C 10422489 NT 011109 18147886 24749 5491
RAI-3'5 C/T 10426701 NT 011109 18148193 25056 5798
RAI-3'4 A/G rs4544343 NT 011109 18150199 27062 7804
RAI-3'8 A/G rs8101662 NT 011109 18150911 27774 8516
RAI-iδ-2 A/T rs8112723 NT 011109 18153497 30360 10204 rs numbers were derived from the NCBI's database dbSNP.
In a most preferred embodiment polymorphism(s) are those listed in table 5d below, more preferably at least two polymorphism(s) are selected, most preferably at least three polymorphism(s) are selected from the polymorphisms shown below:
Table 5d
Trivial name Kind dbSNP# Sequence Position Relative Position in in position SEQID sequence NO:1
XPD-5'3 C/T rs11878644 NT 011109 18145185 22048 2790
RAI-3'7 C/T rs7252567 NT 011109 18146823 23686 4428
RAI-3'3 AA/- rs3047560 NT 011109 18147192 24055 4797
RAI-3'6 A/C 10422489 NT 011109 18147886 24749 5491
RAI-3'5 C/T 10426701 NT 011109 18148193 25056 5798
RAI-3'4 A/G rs4544343 NT 011109 18150199 27062 7804
RAI-3'8 A/G rs8101662 NT 011109 18150911 27774 8516 rs numbers were derived from the NCBI's database dbSNP.
In a preferred embodiment at least one of the following combinations of polymor¬ phisms is included in the methods:
In another embodiment of the invention preferably the method described herein is one in which the tandem repeat is at a position as described in Table 6: Table 6
Identification in uniSTS2
D19S908 STS-W67936
D19S543
D19S393
STS-R48186
GDB:181915 RH47033
GDB:190019
2 UniSTS is a database of unique sequence tag sites established by National Center for Biotechnology Information and located on the internet at http://www.ncbi.nlm.nih.αov/entrez/querv.fcαi?db=unists
In another embodiment of the invention, the method for diagnosis described herein is preferably one in which the sequence polymorphism is in region r. Testing for the presence of the RAI gene allele is especially preferred because, without wishing to be bound by theoretical considerations, of its association with increased risk of can- cer (as explained herein).
In one embodiment of the methods of the invention, preferably the method for diag¬ nosis as described herein, one or more polymorphism(s) at a predetermined position in the region r (SEQ ID NO:1) are identified and used for e.g. cancer risk profiling and/or cancer treatment response profiling. Presently preferred polymorphism(s) are the dinucleotide polymorphism RAI-3'3 in which AA or a deletion is present, the sin¬ gle nucleotide polymorphism RAI exon 6 in which A or T is present, and RAI intron 3 in which A or G is present. However, the present invention relates to any polymor¬ phism and SNP in the r region.
The sequence polymorphism of the invention comprises at least one base differ¬ ence, such as at least two base differences, such as at least three base differences, such as at least four base differences, such as eighty one base pair differences. As described above the sequence polymorphism(s) comprises at least one polymor- phism, such as at least two polymorphisms, such as at least three polymorphisms, such as at least four polymorphisms. Also, the sequence polymorphism comprises at least one polymorphism, such as at least two tandem repeat polymorphisms.
Also, the sequence polymorphism may be a combination of single nucleotide poly- morphism and dinucleotide polymorphism, such as one single nucleotide polymor¬ phism and one dinucleotide polymorphism.
The status of the individual may be determined by reference to allelic variation at one, two, three, four or more of the above loci.
Cell sample
The cell sample used in the present invention may be any suitable cell sample ca¬ pable of providing the genetic material for use in the method. In a preferred em- bodiment, the cell sample is a blood sample, a tissue sample, a sample of secretion, semen, ovum, a washing of a body surface (e.g. a buccal swap), a clipping of a body surface (hairs, or nails), such as wherein the cell is selected from white blood cells and tumour tissue.
It will be appreciated that the test sample may equally be a nucleic acid sequence corresponding to the sequence in the test sample, that is to say that all or a part of the region in the sample nucleic acid may firstly be amplified using any convenient technique e.g. PCR, before use in the analysis of variation in the region.
Detection methods
Detection may be conducted on the sequence of SEQ ID NO: 1 or a complementary sequence as well as on translational (mRNA) and transcriptional products (polypep- tides, proteins) therefrom.
It will be apparent to the person skilled in the art that there are a large number of analytical procedures which may be used to detect the presence or absence of vari¬ ant nucleotides at one or more of positions mentioned herein in the r region. Muta- tions or polymorphisms within or flanking the r region can be detected by utilizing a number of techniques. Nucleic acid from any nucleated cell can be used as the starting point for such assay techniques, and may be isolated according to standard nucleic acid preparation procedures that are well known to those of skill in the art. In general, the detection of allelic variation requires a mutation discrimination tech¬ nique, optionally an amplification reaction and a signal generation system. Table 7 lists a number of mutation detection techniques, some based on the PCR. These may be used in combination with a number of signal generation systems, a selection of which is listed in Table 8. Further amplification techniques are listed in Table 9. Many current methods for the detection of allelic variation are reviewed by Nollau et al., Clin. Chem. 43, 1114-1120, 1997; and in standard textbooks, for example "Labo¬ ratory Protocols for Mutation Detection", Ed. by U. Landegren, Oxford University Press, 1996 and "PCR", 2nd Edition by Newton & Graham, BIOS Scientific Publish¬ ers Limited, 1997.
Table 7
Abbreviations:
ALEX Amplification refractory mutation system linear extension
APEX Arrayed primer extension ARMS Amplification refractory mutation system b-DNA Branched DNA
CMC Chemical mismatch cleavage bp base pair
COPS Competitive oligonucleotide priming system DGGE Denaturing gradient gel electrophoresis
FRET Fluorescence resonance energy transfer
LCR Ligase chain reaction
MASDA Multiple allele specific diagnostic assay
NASBA Nucleic acid sequence based amplification OLA Oligonucleotide ligation assay
PCR Polymerase chain reaction
PTT Protein truncation test
RFLP Restriction fragment length polymorphism
SDA Strand displacement amplification SNP Single nucleotide polymorphism SSCP Single-strand conformation polymorphism analysis
SSR Self sustained replication
TGGE Temperature gradient gel electrophoresis
Table 8 illustrates various mutation detection techniques capable of being used for SNP detection.
Table 8
General techniques: DNA sequencing, Sequencing by hybridisation, SNAPshot.
Scanning techniques: PJT*, SSCP, DOGE, TGGE, Cleavase, Heteroduplex analy¬ sis, CMC, Enzymatic mismatch cleavage
Hybridisation Based techniques
Solid phase bybridisation: Dot blots, MASDA, Reverse dot blots, Oligonucleotide arrays (DNA Chips)
Solution phase hybridisation: Taqman -U.S. Pat. No. 5,210,015 & 5,487,972 (Hoff¬ mann-La Roche), Molecular Beacons — Tyagi et al (1996), Nature Biotechnology, 14, 303; WO 95/13399 (Public Health Inst., New York), Lightcycler, optionally in combination with FRET.
Extension Based: ARMS, ALEX - European Patent No. EP 332435 B1 (Zeneca Limited), COPS - Gibbs et al (1989), Nucleic Acids Research, 17, 2347.
Incorporation Based: Mini-sequencing, APEX
Restriction Enzyme Based: RFLP, Restriction site generating PCR
Ligation Based: OLA
Other: Invader assay Various Signal Generation or Detection Systems is listed below:
Fluorescence: FRET, Fluorescence quenching, Fluorescence polarisation-United Kingdom Patent No. 2228998 (Zeneca Limited)
Other: Chemiluminescence, Electrochemiluminescence, Raman, Radioactivity, CoI- orimetric, Hybridisation protection assay, Mass spectrometry
Table 9 illustrates examples of further amplification techniques.
Table 9
SSR, NASBA, LCR, SDA, b-DNA
Preferred mutation detection techniques include ARMS, ALEX, COPS, Taqman, Molecular Beacons, RFLP, and restriction site based PCR and FRET techniques.
Particularly preferred methods include FRET; taqman, ARMS and RFLP based methods.
In a preferred embodiment, mutations or polymorphisms can be detected by using a microassay of nucleic acid sequences immobilized to a substrate or "gene chip" (see, e.g. Cronin, et a!., 1996, Human Mutation 7:244-255).
Further, improved methods for analyzing DNA polymorphisms, which can be utilized for the identification of region r specific mutations, have been described that capital¬ ize on the presence of variable numbers of short, tandemly repeated DNA sequen¬ ces between the restriction enzyme sites. For example, Weber (U.S. Pat. No. 5,075,217) describes a DNA marker based on length polymorphisms in blocks of (dC-dA)n-(dG-dT)n short tandem repeats. The average separation of (dC-dA)n-(dG- dT)n blocks is estimated to be 30,000-60,000 bp. Markers that are so closely spaced exhibit a high frequency co-inheritance, and are extremely useful in the iden¬ tification of genetic mutations, such as, for example, mutations within the RAI gene, and the diagnosis of diseases and disorders related to RAI mutations. Also, Caskey et al. (U.S. Pat. No. 5,364,759) describe a DNA profiling assay for detecting short tri and tetra nucleotide repeat sequences. The process includes ex¬ tracting the DNA of interest, such as the RAI gene, amplifying the extracted DNA, and labelling the repeat sequences to form a genotypic map of the individual's DNA.
The level of RAI gene expression can also be assayed. For example, RNA from a cell type or tissue known, or suspected, to express the RAI gene may be isolated and tested utilizing hybridization or PCR techniques such as are described, above. The isolated cells can be derived from cell culture or from a patient. The analysis of cells taken from culture may be a necessary step in the assessment of cells to be used as part of a cell-based gene therapy technique or, alternatively, to test the ef¬ fect of compounds on the expression of the RAI gene. Such analyses may reveal both quantitative and qualitative aspects of the expression pattern of the RAI gene, including activation or inactivation of RAI gene expression.
In one embodiment of such a detection scheme, a cDNA molecule is synthesized from an RNA molecule of interest (e.g., by reverse transcription of the RNA mole¬ cule into cDNA). A sequence within the cDNA is then used as the template for a nucleic acid amplification reaction, such as a PCR amplification reaction, or the like. The nucleic acid reagents used as synthesis initiation reagents (e.g., primers) in the reverse transcription and nucleic acid amplification steps of this method are chosen from among the RAI gene nucleic acid reagents described above. The preferred lengths of such nucleic acid reagents are at least 9-30 nucleotides. For detection of the amplified product, the nucleic acid amplification may be performed using radio- actively or non-radioactively labeled nucleotides. Alternatively, enough amplified product may be made such that the product may be visualized by standard ethidium bromide staining or by utilizing any other suitable nucleic acid staining method.
Additionally, it is possible to perform such RAI gene expression assays "in situ", i.e., directly upon tissue sections (fixed and/or frozen) of patient tissue obtained from biopsies or resections, such that no nucleic acid purification is necessary. Nucleic acid reagents such as those described above may be used as probes and/or prim¬ ers for such in situ procedures (see, for example, Nuovo, G. J., 1992, "PCR In Situ Hybridization: Protocols And Applications", Raven Press, NY). Alternatively, if a sufficient quantity of the appropriate cells can be obtained, stan¬ dard Northern analysis can be performed to determine the level of mRNA expres¬ sion of the RAI gene.
Activity of the gene
Another method for detecting sequence polymorphism is by analysing the activity of gene products resulting from the sequences. Accordingly, in one embodiment the detection uses the activity of the RAI gene product as compared to a reference in the method. In particular if the activity of the genes are decreased or increased by at least or about 50 %, such as at least or about 40%, for example at least or about 30%, such as at least or about 20%, for example at least or about 10%, such as at least or about 10%, for example at least or about 5%, such as at least or about 2%, it indicates a sequence polymorphism in the gene.
Mutations outside the region
The present invention may combine the result of sequence polymorphism within the region r with sequence polymorphism outside the region in order to increase the probability of the correlation.
Primers
The primer nucleotide sequences of the invention further include: (a) any nucleotide sequence that hybridizes to a nucleic acid molecule of the region r or its comple¬ mentary sequence or RNA products under stringent conditions, e.g., hybridization to filter-bound DNA in 6x sodium chloride/sodium citrate (SSC) at about 45°C followed by one or more washes in 0.2x SSC/0.1% SDS at about 50-650C, or (b) under highly stringent conditions, e.g., hybridization to filter-bound nucleic acid in 6x SSC at about 45°C followed by one or more washes in 0.1 x SSC/0.2% SDS at about 68°C, or under other hybridization conditions which are apparent to those of skill in the art (see, for example, Ausubel F.M. et al., eds., 1989, Current Protocols in Molecular Biology, Vol. I, Green Publishing Associates, Inc., and John Wiley & sons, Inc., New York, at pp. 6.3.1-6.3.6 and 2.10.3). Preferably the nucleic acid molecule that hybrid- izes to the nucleotide sequence of (a) and (b), above, is one that comprises the complement of a nucleic acid molecule of the region s or r or a complementary se¬ quence or RNA product thereof. In a preferred embodiment, nucleic acid molecules comprising the nucleotide sequences of (a) and (b), comprises nucleic acid mole¬ cule of RAI or a complementary sequence or RNA product thereof.
Among the nucleic acid molecules of the invention are deoxyoligonucleotides ("oli- gos") which hybridize under highly stringent or stringent conditions to the nucleic acid molecules described above. In general, for probes between 14 and 70 nucleo¬ tides in length the melting temperature (TM) is calculated using the formula:
Tm(°C)=81.5+16.6(log [monovalent cations (molar)])+0.41 (% G+C)-(500/N)
where N is the length of the probe. If the hybridization is carried out in a solution containing formamide, the melting temperature is calculated using the equation Tm(°C)=81.5+16.6(log[monovalent cations (molar)])+0.41 (% G+C)-(0.61% forma- mide)-(500/N) where N is the length of the probe. In general, hybridization is carried out at about 20-25 degrees below Tm (for DNA-DNA hybrids) or 10-15 degrees be¬ low Tm (for RNA-DNA hybrids).
Exemplary highly stringent conditions may refer, e.g., to washing in 6x SSC/0.05% sodium pyrophosphate at 370C (for about 14-base oligos), 48°C (for about 17-base oligos), 55°C (for about 20-base oligos), and 6O0C (for about 23-base oligos).
Accordingly, the invention further provides nucleotide primers or probes which de- tect the r region polymorphisms of the invention. The assessment may be conducted by means of at least one nucleic acid primer or probe, such as a primer or probe of DNA, RNA or a nucleic acid analogue such as peptide nucleic acid (PNA) or locked nucleic acid (LNA). The nucleotide primer or probe is preferably capable of hybridis¬ ing to a subsequence of the region corresponding to SEQ ID NO: 1, or a part thereof, or a region complementary to SEQ ID NO: 1.
According to one aspect of the present invention there is provided an allele-specific oligonucleotide probe capable of detecting a r region polymorphism at one or more of positions in the r region as defined by the positions in SEQ ID NO: 1. The allele-specific oligonucleotide probe is preferably 5-50 nucleotides, more pref¬ erably about 5-35 nucleotides, more preferably about 5-30 nucleotides, more pref¬ erably at least 9 nucleotides.
The design of such probes will be apparent to the molecular biologist of ordinary skill. Such probes are of any convenient length such as up to 50 bases, up to 40 bases, more conveniently up to 30 bases in length, such as for example 8-25 or 8- 15 bases in length. In general such probes will comprise base sequences entirely complementary to the corresponding wild type or variant locus in the region. How- ever, if required one or more mismatches may be introduced, provided that the dis¬ criminatory power of the oligonucleotide probe is not unduly affected. The probes of the invention may carry one or more labels to facilitate detection.
In one embodiment, the primers and/or probes are capable of hybridizing to and/or amplifying a subsequence hybridizing to a single nucleotide polymorphism contain¬ ing the sequence shown herein selected from the group of subsequences below or a sequence complementary thereto, wherein the polymorphism is denoted as for ex¬ ample T/C:
1. gctgcagtga gctgt (-/ACACCTGTGGTCCCAGCTACTCTGG AAGCTGAGGTGGGAGGATCGCTTGAGCCCAAGAGGTGGAGGCTG-
CAGTGAGCTGT) gactgtgcca ctgcactcca
2. TGACAGTAGA CATCCTGTCA T (A/G) ATAAGTCttt ttttttt
3. CATCCCCATA CCAAcccacc (c/t) tactgctctg atctcctcct
4. ttttagtagagacatggttccgcca (C/T) gttgcccaggctggtcttgaactc 5. ACTAAAAATAAAAAATAAAAAAAA(-/AA) ATAGCCGAG-
CATGGTGGTGGGGTGC
6. AAAAAACTAAAGTGGGGTTTGCGGG (G/T) AGTGGGAGGGCCCTTCCTGCTAGG
7. ggatcacaag gtcaggagtt (c/t) gagaccagcc tggccaacac 8. TGTTGTCCAA GCTGGCAGAG (A/G) TTTTTGTTTG TTTGTTTGAG
9. CTCGGGAGGCTGAGGCAGGAGAATC (A/G) CTTGAACTCAGGCA- GAGGTTG
10. ATTAAGTGCCTTCACACAGC (A/T) CTGGTTTAAT GTTTATAA
11. ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa 12. gggaggctcg aggcgggc (A/G) gattgcatga gctcaggatt 13. tgcagtgagc tgagatcgc (A/G) ccactgcact ccagcctggg
14. CTTGCTACAGAATTACAGGCA (T/C) GCGCCACCGCTCCGGGCTAA
15. CTAAAGACTACA (-/A) TTTCCCAGCATCCCATTG
16. CCCTCCCTGCTTGCTTGCTTTCTCT [C/η TCTCTCTTTCTTTCTTTCTTTCTT
17. CCCTGCTTGCTTGCTTTCTCTCTCT [CfT] TCTTTCTTTCTTTCTTTCTTTCTT
18. AGAACCTGTTCAGGCTGGCGGCTCA [C/η TTGGAT- GAACAGGGAGTGTGTGAC 19. CCCCCTTCTTAGGACGCATGGGGGT [GfT] GAGAGAACGGGGAGATA-
GACAGAG
20. TGCGAGCAGCCCGGGCTACAGGGTT [AfG] CCTGAGGTGTGGGTCCCAGGATGG
21. GGCGCCTCAACAGCCAGAAGGAGCG [A/G] AGCCTCAGGCCCAGG- CAGCTCTGG
In another embodiment of the methods of the invention the primers and/or probes are capable of hybridizing to and/or amplifying a subsequence hybridizing to a single nucleotide polymorphism containing the sequence shown herein selected from the group of subsequences below or a sequence complementary thereto, wherein the polymorphism is denoted as for example T/C:
1. TGTTGTCCAA GCTGGCAGAG (A/G) TTTTTGTTTG TTTGTTTGAG
2. CTCGGGAGGCTGAGGCAGGAGAATC (A/G) CTTGAACTCAGGCA- GAGGTTG
3. ATTAAGTGCCTTCACACAGC (ATT) CTGGTTTAAT GTTTATAA
4. ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa
5. gggaggctcg aggcgggc (A/G) gattgcatga gctcaggatt
6. tgcagtgagc tgagatcgc (A/G) ccactgcact ccagcctggg 7. CTTGCTACAGAATTACAGGCA (T/C) GCGCCACCGCTCCGGGCTAA
8. CTAAAGACTACA (-IA) TTTCCCAGCATCCCATTG
In a preferred embodiment of the methods of the invention the the primers and/or probes are capable of hybridizing to and/or amplifying a subsequence hybridizing to a single nucleotide polymorphism containing the sequence shown herein selected from the group of subsequences below or a sequence complementary thereto, wherein the polymorphism is denoted as for example TYC:
1. gctgcagtga gctgt (-/ACACCTGTGGTCCCAGCTACTCTGG AAGCTGAGGTGGGAGGATCGCTTGAGCCCAAGAGGTGGAGGCTG-
CAGTGAGCTGT) gactgtgcca ctgcactcca
2. TGACAGTAGA CATCCTGTCA T (A/G) ATAAGTCttt ttttttt
3. CATCCCCATA CCAAcccacc (c/t) tactgctctg atctcctcct
4. ttttagtagagacatggttccgcca (CA") gttgcccaggctggtcttgaactc 5. ACTAAAAATAAAAAATAAAAAAAA(-/AA) ATAGCCGAG-
CATGGTGGTGGGGTGC
6. AAAAAACTAAAGTGGGGTTTGCGGG (G/T) AGTGGGAGGGCCCTTCCTGCTAGG
7. ggatcacaag gtcaggagtt (c/t) gagaccagcc tggccaacac 8. TGTTGTCCAA GCTGGCAGAG (AJG) TTTTTGTTTG TTTGTTTGAG
9. CTCGGGAGGCTGAGGCAGGAGAATC (A/G) CTTGAACTCAGGCA- GAGGTTG
10. ATTAAGTGCCTTCACACAGC (AIT) CTGGTTTAAT GTTTATAA
11. ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa 12. gggaggctcg aggcgggc (A/G) gattgcatga gctcaggatt
13. tgcagtgagc tgagatcgc (AJG) ccactgcact ccagcctggg
14. CTTGCTACAGAATTACAGGCA (T/C) GCGCCACCGCTCCGGGCTAA
15. CTAAAGACTACA (-/A) TTTCCCAGCATCCCATTG
In a preferred embodiment, the primers and/or probes are capable of hybridizing to a subsequence selected from the group of subsequences below:
1. gctgcagtga gctgt (-/ACACCTGTGGTCCCAGCTACTCTGG AAGCTGAGGTGGGAGGATCGCTTGAGCCCAAGAGGTGGAGGCTG- CAGTGAGCTGT) gactgtgcca ctgcactcca
2. TGACAGTAGA CATCCTGTCA T (A/G) ATAAGTCttt ttttttt
3. CATCCCCATA CCAAcccacc (c/t) tactgctctg atctcctcct
4. ttttagtagagacatggttccgcca (C/T) gttgcccaggctggtcttgaactc
5. ACTAAAAATAAAAAATAAAAAAAA(-/AA) ATAGCCGAG- CATGGTGGTGGGGTGC 6. AAAAAACTAAAGTGGGGTTTGCGGG (G/T) AGTGGGAGGGCCCTTCCTGCTAGG
7. ggatcacaag gtcaggagtt (c/t) gagaccagcc tggccaacac
8. TGTTGTCCAA GCTGGCAGAG (A/G) TTTTTGTTTG TTTGTTTGAG 9. CTCGGGAGGCTGAGGCAGGAGAATC (A/G) CTTGAACTCAGGCA-
GAGGTTG
10. ATTAAGTGCCTTCACACAGC (A/T) CTGGTTTAAT GTTTATAA
11. ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa
12. gggaggctcg aggcgggc (A/G) gattgcatga gctcaggatt
In an even more preferred embodiment, the primers and/or probes are capable of hybridizing to a subsequence selected from the group of subsequences below:
1. ACTAAAAATAAAAAATAAAAAAAA(-/AA) ATAGCCGAG- CATGGTGGTGGGGTGC
2. AAAAAACTAAAGTGGGGTTTGCGGG (G/T) AGTGGGAGGGCCCTTCCTGCTAGG
3. ggatcacaag gtcaggagtt (c/t) gagaccagcc tggccaacac
4. TGTTGTCCAA GCTGGCAGAG (A/G) TTTTTGTTTG TTTGTTTGAG 5. CTCGGGAGGCTGAGGCAGGAGAATC (A/G) CTTGAACTCAGGCA-
GAGGTTG
6. ATTAAGTGCCTTCACACAGC (A/T) CTGGTTTAAT GTTTATAA
7. ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa
8. gggaggctcg aggcgggc (A/G) gattgcatga gctcaggatt
Most preferred are the primers and/or probes capable of hybridizing to a subse¬ quence selected from the group of subsequences below:
1. ACTAAAAATAAAAAATAAAAAAAA(-/AA) ATAGCCGAG- CATGGTGGTGGGGTGC
2. ATTAAGTGCCTTCACACAGC (A/T) CTGGTTTAAT GTTTATAA
3. gggaggctcg aggcgggc (A/G) gattgcatga gctcaggatt In a second embodiment according to the methods of the invention are the primers and/or probes capable of hybridizing to a subsequence selected from the group of subsequences below:
1. gctgcagtga gctgt (-/ACACCTGTGGTCCCAGCTACTCTGG
AAGCTGAGGTGGGAGGATCGCTTGAGCCCAAGAGGTGGAGGCTG- CAGTGAGCTGT) gactgtgcca ctgcactcca
2. CATCCCCATA CCAAcccacc (c/t) tactgctctg atctcctcct
3. ttttagtagagacatggttccgcca (C/T) gttgcccaggctggtcttgaactc 4. ACTAAAAATAAAAAATAAAAAAAA(-/AA) ATAGCCGAG-
CATGGTGGTGGGGTGC
5. AAAAAACTAAAGTGGGGTTTGCGGG(G/T) AGTGGGAGGGCCCTTCCTGCTAGG
6. ggatcacaag gtcaggagtt (c/t) gagaccagcc tggccaacac 7. TGTTGTCCAA GCTGGCAGAG (A/G) TTTTTGTTTG TTTGTTTGAG
8. CTCGGGAGGCTGAGGCAGGAGAATC (A/G) CTTGAACTCAGGCA- GAGGTTG
9. ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa
10. CTTGCTACAGAATTACAGGCA (T/C) GCGCCACCGCTCCGGGCTAA 11. CTAAAGACTACA (-/A) TTTCCCAGCATCCCATTG
In a preferred embodiment, the primers and/or probes are capable of hybridizing to a subsequence selected from the group of subsequences below:
1. CATCCCCATA CCAAcccacc (c/t) tactgctctg atctcctcct
2. ttttagtagagacatggttccgcca (C/T) gttgcccaggctggtcttgaactc
3. ACTAAAAATAAAAAATAAAAAAAA(-/AA) ATAGCCGAG- CATGGTGGTGGGGTGC
4. AAAAAACTAAAGTGGGGTTTGCGGG (G/T) AGTGGGAGGGCCCTTCCTGCTAGG
5. ggatcacaag gtcaggagtt (c/t) gagaccagcc tggccaacac
6. TGTTGTCCAA GCTGGCAGAG (A/G) TTTTTGTTTG TTTGTTTGAG
7. CTCGGGAGGCTGAGGCAGGAGAATC (A/G) CTTGAACTCAGGCA- GAGGTTG 8. ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa In an even more preferred embodiment the primers and/or probes are capable of hybridizing to a subsequence selected from the group of subsequences below:
1. CATCCCCATA CCAAcccacc (c/t) tactgctctg atctcctcct
2. ttttagtagagacatggttccgcca (C/T) gttgcccaggctggtcttgaactc
3. ACTAAAAATAAAAAATAAAAAAAA(-/AA) ATAGCCGAG- CATGGTGGTGGGGTGC
4. AAAAAACTAAAGTGGGGTTTGCGGG (GU) AGTGGGAGGGCCCTTCCTGCTAGG
5. ggatcacaag gtcaggagtt (c/t) gagaccagcc tggccaacac
6. TGTTGTCCAA GCTGGCAGAG (A/G) TTTTTGTTTG TTTGTTTGAG
7. CTCGGGAGGCTGAGGCAGGAGAATC (A/G) CTTGAACTCAGGCA- GAGGTTG
Even more preferred are the primers and/or probes are capable of hybridizing to a subsequence selected from the group of subsequences below:
1. CTTGCTACAGAATTACAGGCA (T/C) GCGCCACCGCTCCGGGCTAA 2. CTAAAGACTACA (VA) TTTCCCAGCATCCCATTG
3. ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa
In one embodiment the primers or probes able to detect subsequences as described above are selected from one or more of the following:
1. agt cac age tea ctg cag cct c
2. ace tct tgg get caa gcg ate etc
3. aaa aaa aga ctt ate atg aca gga tgt ct
4. gca aga etc cgt ccc aga aaa aga aaa 5. tcctctctctcccccagctcattttg
6. ACCCCACCCTACTGCTCTGATCTC
7. agg ctg gtc ttg aac tec tgg get taa g
8. ggt tec gee acg ttg cc
9. teg get att ttt ttt ttt att ttt tta tt 10. att aca ggc ace cac cac cat g 11. ace cca ctt tag ttt ttt ttt cct eta gtg ate gcc
12. gcc etc cca eta ccc gca
13. aga egg ggt ttc act gtg ttg gc
14. agg ctg gtc tea aac tec tga c 15. CAAACAAACAAAAACCTCTGCCA
16. CTTGGACAACATAGGGAGACCCTGTGT
17. tgc etc age etc ccg agt age t
18. cct cct gag ttc aag cga ttc tc
19. tgc ctt cac aca get ctg gtt taa tg 20. tgc ctt cac aca gca ctg gtt taa tg
21. ACCATGTTGGCCAGGCTGGTTTT
22. ATCTACTGACCTCAAATGATCCACCT
23. tgc aat ccg ccc gcc
24. cca ggc tgg ttt gga aat cct gag etc 25. ctg aga teg cac cac tgc ac
26. ggg agg egg age ttg cag tga
27. gcg cat gcc tgt aat tct gta
28. cag gac gag cca cag aca aaa etc c
29. tgc aat gag get cct ggc c 30. act aca att tec cag cat ccc a
31. cct ccc tec etc cct gc
32. tgc ttg ctt tct tct tct
33. tec ctg ctt get tgx ttt etc t
34. tct etc ttt ctt tct ttc ttt c 35. tgt tea tec aaa tga gcc gc
36. age ctg aac agg ttc tgt tec ttc gac tt
37. caa get get ate teg ace gat ctt
38. ggg tga cca ccc tgc cag cc
39. egg get aca ggg tta cct gag 40. tct gca ace tgg tgc gag cag c
41. tga ggc tec get cct tct gg
42. age tgc cag age tgc ctg ggc One or more primers able to detect subsequences by hybridisation as described may in a particularly preferred embodiment of the method of the invention be se¬ lected from
1. teg get att ttt ttt ttt att ttt tta tt
2. att aca ggc ace cac cac cat g
3. tgc ctt cac aca get ctg gtt taa tg
4. tgc ctt cac aca gca ctg gtt taa tg
5. tgc aat ccg cec gcc 6. cca ggc tgg ttt gga aat ect gag etc
According to another aspect of the present invention there is provided a diagnostic nucleic acid primer capable of detecting a r region polymorphism at one or more of positions in the r region as defined by the in SEQ ID NO: 1.
The primer or probe may be a diagnostic nucleic acid primer defined as an allele specific primer, used, generally together with a constant primer, in an amplification reaction such as a PCR reaction, which provides the discrimination between alleles through selective amplification of one allele at a particular sequence position. The diagnostic primer is preferably 5-50 nucleotides, more preferably about 5-35 nucleo¬ tides, more preferably about 5-30 nucleotides, more preferably at least 9 nucleo¬ tides.
In accordance with the present invention diagnostic primers are provided, compris- ing the sequences set out below as well as derivatives thereof wherein about 6-8 of the nucleotides at the 3' terminus are identical to the sequences given below and wherein up to 10, such as up to 8, 6, 4, 2, or 1 of the remaining nucleotides may be varied without significantly affecting the properties of the diagnostic primer. Conven¬ iently, the sequence of the diagnostic primer is as written below.
Furthermore, as described above at least two sets of primer(s) and/or probe(s), such as at least three sets of primer(s) and/or probe(s) may be combined in the method thereby increasing the correlation probability. This second or other set of primer(s) and/or probe(s) may be a nucleotide or nucleotide analogues hybridising to a region within the region r or to a sequence different from the region r. Said sequence differ- ent from the region r is preferably a region in chromosome 19, preferably in chromo¬ some 19q. In particular such second or other primer or probe may be selected from one or more of the sequences below, or the complementary strands:
1. tgc etc ace cct gta ate c
2. get tgt aat ccc age tac teg
3. caa cac tea cac ccc aca g
4. aga tea cgc cac tgc act c
5. ttg aca att gag caa aga gc 6. ttg gat tac aga cgt gag c
7. agt gca gee tea act tec
8. cca gtc caa aca ata tga tec
9. cat gat tea ctg cac cca ace
10. ttt cac tct tgt tgc cca age 11. cac cat gcc tgg etc caa tgt
12. age cca gga att caa gg
13. aga aca ttg gag cca gg
14. aga ate act gga ate cag g
15. ttt tea cac agg tec aat cc 16. act gca ace tec ate tec
17. caa tta agt gcc ttc aca cag ca
18. ggt caa gag ttc aag ace age
19. ccc tgc ccc ace tct cc
20. agt caa ttt ctg tgc aaa eta ctt tta ttt 21. gag gca aca gga aca aac c
22. cat tgg aat gag cag aaa cc
23. taa cat aaa gaa tea gga gga ggc
24. agt tgg etc ate tgc etc tt
25. tgg eta aca egg tga aac c 26. gga ate caa aga ttc tat gat gg
27. act cct gac ttc aaa tga tec
28. tag ccc cca gtc acg ttc c
29. aga agt cca aga gtt tgc age
30. ttc tea gtc cca gaa tga ace 31. cca ctt agg taa aca cct ctt 32. ctg caa tga gcc gag ata gaa
33. cca ctt agg taa aca cct ctt
34. ctg caa tga gcc gag ata gaa
35. atg ttg ggg aga ctg agg 36. ccg cat eta act tat tct gg
37. aac tac etc tgc aaa ccc age
38. ttg gaa tgg agg gat tct ace
39. ggt ttt ctg etc tgc aca eg
40. cct ttc tec ttc cac caa eg 41. gga cag atg gca atg atg g
42. tct tct tct tgg tgg atg tgg
The primers and probes may be manufactured using any convenient method of syn¬ thesis. Examples of such methods may be found in standard textbooks, for example "Protocols for Oligonucleotides and Analogues; Synthesis and Properties," Methods in Molecular Biology Series; Volume 20; Ed. Sudhir Agrawal, Humana ISBN: 0- 89603-247-7; 1993; Lsup.st Edition. If required the primer(s) and probe(s) may be labelled to facilitate detection.
Kit
According to another aspect of the present invention, there is provided a diagnostic kit comprising at least one diagnostic primer of the invention and/or at least one al- lele-specific oligonucleotide primer of the invention.
The diagnostic kits may comprise appropriate packaging and instructions for use in the methods of the invention. Such kits may further comprise appropriate buffer(s) and polymerase(s) such as thermostable polymerases, for example taq polymerase.
Preferred kits can comprise means for amplifying the relevant sequence such as primers, polymerase, deoxynucleotides, buffer, metal ions; and/or means for dis¬ criminating the polymorphism, such as one or a set of probes hybridising to the poly¬ morphic site, a sequence reaction covering the polymorphic site, an enzyme or an antibody; and/or a secondary amplification system, such as enzyme-conjugated antibodies, or fluorescent antibodies. The kit-of-parts preferably also comprises a detection system, such as a fluorometer, a film, an enzyme reagent or another highly sensitive detection device.
The methods described herein may be performed, for example, by utilizing pre- packaged diagnostic kits. The invention therefore also encompasses kits for detect¬ ing the presence of a polypeptide or nucleic acid of the invention in a biological sample (i.e., a test sample). Such kits can be used, e.g., to determine if a subject is suffering from or is at increased risk of developing a disorder associated with a dis¬ order-causing allele, or aberrant expression or activity of a polypeptide of the inven- tion. For example, the kit can comprise a labeled compound or agent capable of detecting the polypeptide or mRNA or DNA or RAI gene sequences, e.g., encoding the polypeptide in a biological sample. The kit can further comprise a means for de¬ termining the amount of the polypeptide or mRNA in the sample (e.g., an antibody which binds the polypeptide or an oligonucleotide probe which binds to DNA or mRNA encoding the polypeptide). Kits can also include instructions for observing that the tested subject is suffering from or is at risk of developing a disorder associ¬ ated with aberrant expression of the polypeptide if the amount of the polypeptide or mRNA encoding the polypeptide is above or below a normal level, or if the DNA correlates with presence of an RAI allele that causes a disorder.
For antibody-based kits, the kit can comprise, for example: (1) a first antibody (e.g., attached to a solid support) which binds to a polypeptide of the invention; and, op¬ tionally, (2) a second, different antibody which binds to either the polypeptide or to the first antibody and is conjugated to a detectable agent.
Identification of an allele as having implication for risk of cancer
An allele in the r region can be identified as correlated with an increased risk of de¬ veloping cancer on the basis of statistical analyses of the incidence of a particular allele in two groups of individuals with and without cancer, respectively, according to the χ2test, which is well known in the art. Furthermore, an allele in the region can be identified as an allele correlated with prognosis of cancer on the basis of statistical analyses of the incidence of a particular allele in individuals demonstrating different prognostic characteristics. Identification of humans having increased likelihood of responding to treat¬ ment
It is further contemplated that the present invention provides a method for identifying a human subject as having an increased likelihood of responding positively to a cancer treatment, comprising determining the presence in the subject of a s or r re¬ gion allele genotype correlated with an increased likelihood of positive response to treatment, whereby the presence of the genotype identifies the subject as having an increased likelihood of responding to cancer treatment.
The treatment mentioned herein may be any cancer treatment, such as conventional cancer treatment, for example X-ray, chemotherapeutics, surgical excision or com¬ binations thereof.
Protein Products of the Gene(s)
Gene products of the region r or peptide fragments thereof, can be prepared for a variety of uses. For example, such gene products, or peptide fragments thereof, can be used for the generation of antibodies, in diagnostic assays.
The gene products of the invention include, but are not limited to, human RAI gene products. In the following the invention is described in relation to RAI gene product.
Gene product, sometimes referred to herein as an "protein" or "polypeptide", in- eludes those gene products encoded by the RAI gene sequences shown as position 7760-22885 in SEQ ID NO: 1. Among gene product variants are gene products comprising amino acid residues encoded by the polymorphisms. Such gene product variants also include a variant of the RAI gene product.
In addition, RAI gene products may include proteins that represent functionally equi¬ valent gene products. In preferred embodiments, such functionally equivalent RAI gene products are naturally occurring gene products. Functionally equivalent RAI gene products also include gene products that retain at least one of the biological activities of the RAI gene products described above, and/or which are recognized by and bind to antibodies (polyclonal or monoclonal) directed against RAI gene prod¬ ucts.
Antibodies to Gene Products
Described herein are methods for the production of antibodies capable of specifi¬ cally recognizing one or more gene product epitopes or epitopes of conserved vari¬ ants or peptide fragments of the gene products. Furthermore, antibodies that spe¬ cifically recognize mutant forms are encompassed by the invention. The terms "spe- cifically bind" and "specifically recognize" refer to antibodies that bind to RAI gene product epitopes at a higher affinity than they bind to non-RAI (e.g., random) epi¬ topes.
Such antibodies may include, but are not limited to, polyclonal antibodies, mono- clonal antibodies (mAbs), humanized or chimeric antibodies, single chain antibodies, Fab fragments, F(ab')2 fragments, fragments produced by a Fab expression library, anti-idiotypic (anti-Id) antibodies, and epitope-binding fragments of any of the above, including the polyclonal and monoclonal antibodies described below. Such antibod¬ ies may be used, for example, in the detection of a gene product in a biological sample and may, therefore, be utilized as part of a diagnostic or prognostic tech¬ nique whereby patients may be tested for abnormal levels of gene products, and/or for the presence of abnormal forms of such gene products. Such antibodies may also be utilized in conjunction with, for example, compound screening schemes, as described, below, for the evaluation of the effect of test compounds on gene product levels and/or activity.
For the production of antibodies against a gene product, various host animals may be immunized by injection with a RAI gene product, or a portion thereof. Such host animals may include, but are not limited to rabbits, mice, and rats, to name but a few. Various adjuvants may be used to increase the immunological response, de¬ pending on the host species, including but not limited to Freund's (complete and in¬ complete), mineral gels such as aluminum hydroxide, surface active substances such as lysolecithin, pluronic polyols, polyanions, peptides, oil emulsions, keyhole limpet hemocyanin, dinitrophenol, and potentially useful human adjuvants such as BCG (bacille Calmette-Guerin) and Corynebacterium parvum. Polyclonal antibodies are heterogeneous populations of antibody molecules derived from the sera of animals immunized with an antigen, such as a gene product, or an antigenic functional derivative thereof. For the production of polyclonal antibodies, host animals such as those described above, may be immunized by injection with gene product supplemented with adjuvants as also described above.
Monoclonal antibodies, which are homogeneous populations of antibodies to a par¬ ticular antigen, may be obtained by any technique that provides for the production of antibody molecules by continuous cell lines in culture. These include, but are not limited to, the hybridoma technique of Kohler and Milstein (1975, Nature 256:495- 497; and U.S. Pat. No. 4,376,110), the human B-cell hybridoma technique (Kosbor et al., 1983, Immunology Today 4:72; Cole et al., 1983, Proc. Natl. Acad. Sci. U.S.A. 80:2026-2030), and the EBV-hybridoma technique (Cole et al., 1985, Monoclonal Antibodies And Cancer Therapy, Alan R. Liss, Inc., pp. 77-96). Such antibodies may be of any immunoglobulin class including IgG, IgM, IgE, IgA, IgD and any subclass thereof. The hybridoma producing the mAb of this invention may be cultivated in vitro or in vivo. Production of high titers of mAbs in vivo makes this the presently preferred method of production.
In addition, techniques developed for the production of "chimeric antibodies" (Morri¬ son, et al., 1984, Proc. Natl. Acad. Sci., 81 :6851-6855; Neuberger, et al., 1984, Na¬ ture 312:604-608; Takeda, et al., 1985, Nature, 314:452-454) by splicing the genes from a mouse antibody molecule of appropriate antigen specificity together with genes from a human antibody molecule of appropriate biological activity can be used. A chimeric antibody is a molecule in which different portions are derived from different animal species, such as those having a variable region derived from a mur¬ ine mAb and a human immunoglobulin constant region. (See, e.g., Cabilly et al., U.S. Pat. No. 4,816,567; and Boss et al., U.S. Pat. No. 4,816397, which are incorpo- rated herein by reference in their entirety.)
In addition, techniques have been developed for the production of humanized anti¬ bodies. (See, e.g., Queen, U.S. Pat. No. 5,585,089, which is incorporated herein by reference in its entirety.) An immunoglobulin light or heavy chain variable region consists of a "framework" region interrupted by three hypervariable regions, referred to as complementarity determining regions (CDRs). The extent of the framework region and CDRs have been precisely defined (see, "Sequences of Proteins of Im¬ munological Interest", Kabat, E. et al., U.S. Department of Health and Human Ser¬ vices (1983) ). Briefly, humanized antibodies are antibody molecules from non- human species having one or more CDRs from the non-human species and a framework region from a human immunoglobulin molecule.
Alternatively, techniques described for the production of single chain antibodies (U.S. Pat. No. 4,946,778; Bird, 1988, Science 242:423-426; Huston, et al., 1988, Proc. Natl. Acad. Sci. U.S.A. 85:5879-5883; and Ward, et al., 1989, Nature 334:544- 546) can be adapted to produce single chain antibodies against gene products. Sin¬ gle chain antibodies are formed by linking the heavy and light chain fragments of the Fv region via an amino acid bridge, resulting in a single chain polypeptide.
Antibody fragments that recognize specific epitopes may be generated by known techniques. For example, such fragments include but are not limited to: the F(ab')2 fragments, which can be produced by pepsin digestion of the antibody molecule and the Fab fragments, which can be generated by reducing the disulfide bridges of the F(ab')2 fragments. Alternatively, Fab expression libraries may be constructed (Huse, et al., 1989, Science 246:1275-1281) to allow rapid and easy identification of mono¬ clonal Fab fragments with the desired specificity.
Immunoassays for gene products, conserved variants, or peptide fragments thereof will typically comprise incubating a sample, such as a biological fluid, a tissue ex- tract, freshly harvested cells, or lysates of cells in the presence of a detectably la¬ beled antibody capable of identifying gene product, conserved variants or peptide fragments thereof, and detecting the bound antibody by any of a number of tech¬ niques well-known in the art.
The biological sample may be brought in contact with and immobilized onto a solid phase support or carrier, such as nitrocellulose, that is capable of immobilizing cells, cell particles or soluble proteins. The support may then be washed with suitable buffers followed by treatment with the detectably labeled gene product specific anti¬ body. The solid phase support may then be washed with the buffer a second time to remove unbound antibody. The amount of bound label on the solid support may then be detected by conventional means.
By "solid phase support or carrier" is intended any support capable of binding an antigen or an antibody. Well-known supports or carriers include glass, polystyrene, polypropylene, polyethylene, dextran, nylon, amylases, natural and modified cellu¬ loses, polyacrylamides, gabbros, and magnetite. The nature of the carrier can be either soluble to some extent or insoluble for the purposes of the present invention. The support material may have virtually any possible structural configuration so long as the coupled molecule is capable of binding to an antigen or antibody. Thus, the support configuration may be spherical, as in a bead, or cylindrical, as in the inside surface of a test tube, or the external surface of a rod. Alternatively, the surface may be flat such as a sheet, test strip, etc. Preferred supports include polystyrene beads. Those skilled in the art will know many other suitable carriers for binding antibody or antigen, or will be able to ascertain the same by use of routine experimentation.
One of the ways in which the RAI gene product-specific antibody can be detectably labeled is by linking the same to an enzyme, malate dehydrogenase, staphylococcal nuclease, delta-5-steroid isomerase, yeast alcohol dehydrogenase, α-glycero- phosphate, dehydrogenase, triose phosphate isomerase, horseradish peroxidase, alkaline phosphatase, asparaginase, glucose oxidase, β-galactosidase, ribonucle- ase, urease, catalase, glucose-6-phosphate dehydrogenase, glucoamylase and acetylcholinesterase. The detection can be accomplished by colorimetric methods that employ a chromogenic substrate for the enzyme. Detection may also be ac- complished by visual comparison of the extent of enzymatic reaction of a substrate in comparison with similarly prepared standards.
Detection may also be accomplished using any of a variety of other immunoassays. For example, by radioactively labeling the antibodies or antibody fragments, by Ia- beling the antibody with a fluorescent compound. Among the most commonly used fluorescent labeling compounds are fluorescein isothiocyanate, rhodamine, phyco- erythrin, phycocyanin, allophycocyanin, o-phthaldehyde and fluorescamine. The antibody can also be detectably labeled using fluorescence emitting metals such as 152Eu, or others of the lanthanide series or by coupling it to a chemilumines- cent compound.
Diseases
Described herein are various applications of gene sequences, gene products, in¬ cluding peptide fragments and fusion proteins thereof, and of antibodies directed against gene products and peptide fragments thereof. Such applications include, for example, prognostic and diagnostic evaluation of a disease, such as cancer, and the identification of subjects with a predisposition to such disorders, as described above.
The method according to the invention may be used in relation to any cancer form, such as, but not limited to, skin carcinoma including malignant melanoma, breast cancer, lung cancer, colon cancer and other cancers in the gastro-intestinal tract, prostate cancer, lymphoma, leukemia, multiple myeloma, pancreas cancer, head and neck cancer, ovary cancer and other gynecological cancers. In particular the method is relevant for skin cancer, lung cancer, colon cancer, multiple myeloma, and breast cancer, such as skin cancer, breast cancer, multiple myeloma, and lung cancer, such as skin cancer, breast cancer and lung cancer, such as skin cancer and breast cancer, preferably wherein the skin cancer is basal cell carcinoma, such as lung cancer. For example the cancer is multiple myeloma, In another embodi¬ ment the cancer is breast cancer. The cancer may also be skin cancer, preferably basal cell carcinoma, for example early age basal cell carcinoma.
In particular, the method is relevant for both early age cancer and later age cancer, such as early age breast cancer, and such as later age breast cancer.
The method is of particular relevance for lung cancer, such as in patients with XPD exon 23 M . The method is also in particular relevant for early age skin cancer, such as early age basal cell carcinoma.
Gene nucleic acid sequences, described above, can be utilized for transferring re¬ combinant nucleic acid sequences to cells and expressing said sequences in recipi- ent cells. Such techniques can be used, for example, in marking cells or for the treatment of cancer. Such treatment can be in the form of gene replacement ther¬ apy. Specifically, one or more copies of a normal RAI gene or a portion of the RAI gene that directs the production of an RAI gene product exhibiting normal RAI gene function, may be inserted into the appropriate cells within a patient, using vectors that include, but are not limited to, adenovirus, adeno-associated virus, and retrovi¬ rus vectors, in addition to other particles that introduce DNA into cells, such as lipo¬ somes.
In another embodiment, the invention may be used in relation to inflammatory dis- eases, such as, but not limited thereto, rheumatoid arthritis, colitis ulcerosa, Crohn's disease, thyroiditis, neural inflammation as in Alzheimer's disease, and Guillain- Barre syndrome.
Primary effectors The primary effectors i.e. the mutations causing cancer according to the present invention may be found within the region including and flanked by the marker RAI- 37 and the polymorphism having the sequence
AAGTTTCTCTATT[G/T]TGTTTATAAACA corresponding to position 18151158of contig nt_011109 and 18150815, respectively. The region and markers herein represents the region around position 24000 which is shown to be important in relation to can¬ cer according to the present invention. The primary effector may be selected from the group consisting of
novel polymorphisms identified in the present invention
In yet another embodiment of the present invention the primary effectors i.e. the mutations causing cancer according to the present invention may be be selected from the group consisting of
In a preferred embodiment the primary effectors i.e. the mutations causing cancer according to the present invention may be be selected from the group consisting of
In another embodiment the primary effectors i.e. the mutations causing cancer ac¬ cording to the present invention may further be be selected from the group consist¬ ing of
In yet another embodiment the primary effectors i.e. the mutations causing cancer according to the present invention may further be be selected from the group con¬ sisting of
In another embodiment the primary effectors i.e. the mutations causing cancer ac¬ cording to the present invention may further be be selected from the group consist¬ ing of
In another preferred embodiment the primary effectors i.e. the mutations causing cancer according to the present invention may further be be selected from the group consisting of
In yet another preferred embodiment the primary effectors i.e. the mutations causing cancer according to the present invention may further be be selected from the group consisting of
In yet another preferred embodiment the primary effectors i.e. the mutations causing cancer according to the present invention may further be be selected from the group consisting of
The primary effector may be the mutation corresponding to the sequence TTTTAG- TAGAGACATGGTTCCGCCA[C/ηGTTGCCCAGGCTGGTCTTGAACTCC posi¬ tioned at 18146823 of Contig nt_011109. In another embodiment the primary effec¬ tor may be the mutation corresponding to the sequence ctggggaggctgaggcagga- gaatc[A/G]cttgaaaccgggaggcggaggttgt positioned at 18147126 of Contig nt_011109. However, in another preferred embodiment the primary effector may be the mutation corresponding to the sequence GATTGTCATGT[G/T]ACATCAGCCAATACT posi¬ tioned at 18146233 of Contig nt_011109. In yet another embodiment the primary effector may be the mutation corresponding to the sequence caggcggatca- caaggtcaggagtt[C/T]gagaccagcctggccaacacagtga positioned at 18148193 of Contig nt_011109. In a further embodiment the primary effector may be the mutation corre¬ sponding to the sequence CACAGTGAAAC[C/T]CCATCTCTACTAAA positioned at 18149120 of Contig nt_011109. In another preferred embodiment the primary effec- tor may be the mutation corresponding to the sequence AGCCTGGCCAACATG[CZG]TGAAACCCCGTCTCT positioned at 18150815 of Contig nt_011109. In yet another embodiment the primary effector may be the muta¬ tion corresponding to the sequence ctcgggaggctgaggcagga- gaatc[A/G]cttgaactcaggaggcagaggttgc positioned at 18150911 of Contig nt_011109. In a further embodiment the primary effector may be the mutation corresponding to the sequence AAGTTTCTCTATT[GZT]TGTTTATAAACA positioned at 18151158 of Contig nt_011109. However, in another preferred embodiment the primary effector may be the mutation corresponding to the sequence CCCTATGTTGTCCAAGCTGGCAGAG[AZG]TTTTT-GTTTGTTTGTTTGAGAGGGA positioned at 18150199 of Contig nt_011109. Similarly in yet a further embodiment, the primary effector may be the mutation corresponding to the sequence ac- taaaaataaaaaaataaaaaaaa[-ZAA]atagccgagcatggtggtgggtgcc positioned at 18147192 of Contig nt_011109. In another embodiment the primary effector may be the muta- tion corresponding to the sequence AAAAAACTAAAGTGGGGTTTGCGGG[GZT]- AGTGGGAGGGCCCTTCCTGCTAGGT positioned at 18147886 of Contig nt_011109 . In a further embodiment the primary effector may be the mutation cor¬ responding to the sequence AAAATTAGCCGG[AZG]CGCCATGGCGGGAG posi¬ tioned at 18149154 of Contig nt_011109. In an especially preferred embodiment the primary effector may be the mutation corresponding to the sequence GGTTTAT[ATTTT]Ntgagatggatttt positioned at 18147012 of Contig nt_011109. The primary effectors may however be any combination of the polymorphisms.
Among the polymorphisms employed herein for providing a method andZor composi- tions for identifying human subjects with an increased risk of having or developing disease, a number of polymorphisms are novel and have been identified by the pre¬ sent inventors. For such polymorphisms the position according to Contig nt_011109 and the nucleotide sequence of the identified polymorphism are provided, whereas the novel polymorphisms cannot be assigned trivial names or identification numbers according to the dbSNP database.
The primary effectors i.e. the mutations causing cancer according to the present invention may be found within the region around position 39000 which is shown to be important in relation to cancer according to the present invention. The region includes and is flanked by the marker Rai intron 1 and the polymorphism having the sequence TCCAGCCTGGGCAAGAA[C/G]AGTGAAACTCCAGCTT corresponding to position 18165052 of contig nt_011109. The primary effector may be selected from the group consisting of
In another embodiment of the invention the primary effectors i.e. the mutations caus¬ ing cancer according to the present invention may be be selected from the group consisting of
In yet another embodiment of the invention the primary effectors i.e. the mutations causing cancer according to the present invention may be be selected from the group consisting of
In a further embodiment of the invention the primary effectors i.e. the mutations causing cancer according to the present invention may be be selected from the group consisting of
In another embodiment of the invention the primary effectors i.e. the mutations caus¬ ing cancer according to the present invention may be be selected from the group consisting of
In yet a further embodiment of the invention the primary effectors i.e. the mutations causing cancer according to the present invention may be be selected from the group consisting of
In a further embodiment the primary effector may be selected from the group con¬ sisting of
In one embodiment of the present invention the primary effector may be the muta¬ tion corresponding to the sequence ggagcttgcagtgagctga- gatcgc[A/G]ccactgcactccagcctgggcgaca positioned at 18159091 of Contig nt_011109. In another embodiment the primary effector may be the mutation corre¬ sponding to the sequence TTCTCCTGACCTC[A/G]TGATCCGCCCACCTCGG posi¬ tioned at 18159263 of Contig nt_011109. In yet another embodiment the primary effector may be the mutation corresponding to the sequence GGGATTACAGG- CATGC[A/G]CCACCAGGCCCAGCTAATTTTTGT positioned at 18160363 of Contig nt_011109. In a further embodiment the primary effector may be the mutation corre¬ sponding to the sequence TCCAATGGTGACA[A/C]CAGTAAGAGCAGTTAACAG positioned at 18160936 of Contig nt_011109. In a preferred embodiment the primary effector may be the mutation corresponding to the sequence TCCAATGGTGACAA[C/G]AGTAAGAGCAGTTAACAG positioned at 18160937 of Contig nt_011109. In a further preferred embodiment the primary effector may be the mutation corresponding to the sequence tacaggcgcccgccac- cacccccag[A/C]taatttttgtatttttagtagagac positioned at 18161433 of Contig nt_011109. In a further preferred embodiment the primary effector may be the mutation corre¬ sponding to the sequence TTGCCTCAGCCTCCTGA[G/T]TAGCTGGGATTGGAATGAGA positioned at 18161694 of Contig nt_011109. However, in another embodiment the primary effec- tor may be the mutation corresponding to the sequence TACGA- TAAATAGCTAGA[C/GACCTTGGCGCCACCATCTη positioned at 18161841 of Contig nt_011109. In an especially preferred embodiment the primary effector may be the mutation corresponding to the sequence AAAATAATAATAATAATAT- TAA[C/T]CCTGACCTTGGCGCCACCATCT positioned at 18161896 of Contig nt_011109. In further preferred embodiment the primary effector may be the muta¬ tion corresponding to the sequence tcgtcctgctacagaatta- caggca[C/T]gcgccaccgctccgggctaattttt positioned at 18162206 of Contig nt_011109. In another embodiment the primary effector may be the mutation corresponding to the sequence CCTCATGAGCCACCCAC[CZT]TCGGCCTCCCAAAGTGCT posi- tioned at 18162309 of Contig nt_011109. In yet another embodiment the primary effector may be the mutation corresponding to the sequence TGAGCCACCGCGCCC[A/G]GCCGAGACTCACTATTT positioned at 18162356 of Contig nt_011109. In yet a further embodiment the primary effector may be the mu¬ tation corresponding to the sequence taaagcgggag- gatggcttgaacct[AΛ3]ggaggcggaggttgcagtgagccga positioned at 18162599 of Contig nt_011109. In a further embodiment. In an additional embodiment the primary effec¬ tor may be the mutation corresponding to the sequence GGAGAGAAGGAGCAGA- GAAC[A/C]TCTCTATGTGGCCA positioned at 18162903 of Contig nt_011109. However, in another embodiment the primary effector may be the mutation corre- sponding to the sequence ATCCTAAAGACTAC[A/C]TTTCCCAGCATCCCA posi¬ tioned at 18162970 of Contig nt_011109. Furthermore, in a further embodiment the primary effector may be the mutation corresponding to the sequence TTTCCCAG- CATCCCA[C/T]TGCAATGAGGCTCCTGGCC positioned at 18162986 of Contig nt_011109. In yet a further embodiment the primary effector may be the mutation corresponding to the sequence TCCTGACTCCAGTG[A/C]GGTGCCTACAGTCCTG positioned at 18163200 of Contig nt_011109. In yet a preferred embodiment the primary effector may be the mutation corresponding to the sequence TTCCAGCCTGGGCAAGAA[C/G]AGTGAAACTCCAGCTT positioned at 18165052 of Contig nt_011109. In a further preferred embodiment the primary effectors may however be any combination of the polymorphisms. RAI function
Presumably, the primary effector influences activity or expression of RAI or XPD. Possibly, one of the markers analyzed is the effector itself. However, there are sev¬ eral alternative possibilities. First, the effector could modify the sequence of RAI protein, through amino acid substitution, splicing or termination. Secondly, one likely position of the effector is in the 3' portion of the gene and that the effector modifies mRNA expression and stability (22). Thirdly, one could imagine that the effector was a modification of an enhancer situated between RAI and XPD, again modulating the expression of RAI, and possibly modulating other nearby genes as well. Finally, the primary effector may be located in the promoter or first intron of RAI and influence the transcriptional activity of the gene. If a primary result of the effector is modified RAI action, then one result may be modified apoptosis. RAI protein is an inhibitor of ReIA, a subunit of the transcription factor NF-κB (23, 24). NF-κB has long been implicated in both cell proliferation and apoptosis. Modulation of NF-κB may well be part of a "crunch-time scenario", in¬ voked when the cell has to muster its forces and make life-and-death decisions. A recent scientific paper suggests that the choice between the cell survival and death is regulated by the relative activity of the two subunits encoded by ReIA and c-rel, respectively (25). By neutralising the ReIA product, RAI protein would presumably shift this balance towards apoptosis. RAI protein may also influence p53 activity (26).
Predictor of RAI action
The present invention relates to method of estimating the disease risk of an individ¬ ual comprising further a predictor of RAI action. In one embodiment the length of the RAI transcript may be used as a predictor. In another embodiment the quantity of RAI transcripts may be used as a predictor. The present invention comprises the combined RAI transcript characteristics described above as a predictor of RAI action and the estimation of the disease risk of an individual. Examples
Materials and methods
Study groups. Diet, Cancer and Health (DCH) is a Danish prospective follow-up study. Individuals eligible for invitation were born in Denmark, were living in the Co¬ penhagen or Aarhus areas, and were not at the time of invitation registered with a previous diagnosis of cancer (including non-melanoma skin cancer) in the Danish Cancer Registry. Invited to participate were 160,725 individuals aged 50-64 years of which 57.053 individuals were recruited (22). Among these 542 were later notified to the Danish Cancer Register with a cancer diagnosed before the date of enrolment and were therefore excluded. At enrolment (1993-1997), detailed information on diet, smoking habits, lifestyle, weight, height, reproduction, medical treatment, and other socio-economic characteristics and environmental exposures were collected. For women hormone replacement therapy (HRT) and menopausal status were also recorded. Moreover, blood, urine, fat tissue and other biological material was sam¬ pled and stored in a biobank at -15O0C. Cohort members were identified by a unique identification number, which is allocated to every Danish citizen by the Central Population Registry. Cohort members were linked to The Central Population Regis¬ try for information on vital status and immigration. Information on cancer occurrence among cohort members was obtained through record linkage to the Danish Cancer Registry, which collects information on all inhabitants in Denmark who develop can¬ cer (7). Linkage was performed by use of the personal identification number.
Study groups for breast cancer. In "Diet Health and Cancer" 79,729 women aged 50-64 years were invited to participate in the study and 29,875 accepted the invita¬ tion. Of these a total of 326 women that later were reported to the Danish Cancer Registry with a cancer before the visit to the study clinic were excluded from the study. In addition, 8 women were excluded from the study because they did not fill in the lifestyle questionnaire. Because the present analysis aimed at the subgroup of women, who were postmenopausal at study entry, we further excluded 4,844 sup¬ posedly pre-menopausal women, including 4,798 women who had reported at least one menstruation no more than 12 months prior to entry and no use of hormone replacement therapy (HRT), nine women who gave a lifetime history of no men¬ struations and 37 women who did not answer the questions about current or previ- ous use of HRT, leaving 24,697 postmenopausal women. Each cohort member was followed-up for breast cancer occurrence from date of entry, i.e. date of visit to the study centre until the date of diagnosis of any cancer (except for non-melanoma skin cancer), date of death, date of emigration, or 31 De- cember 2000, whichever came first. A total of 434 women were diagnosed with inci¬ dent breast cancer during the follow up period (Vogel et al. In press), and a similar number was selected as controls by individual matching for age, HRT use and menopausal status.
Study groups for lung cancer. These were also recruited from "Diet, Health and Cancer". Among the 56.511 individuals with no previous cancer diagnosis, 2291 with incomplete, inconsistent or missing information on smoking habits were excluded. Among the remainder 54.220 cohort members included in this study, 265 cases of lung cancer, diagnosed between 1994 and 2001 , were identified in the files of the Danish Cancer Registry (23). Also from among the cohort members, a sub-cohort of 272 persons (including 4 of the cases) were selected at random and weighted in numbers similarly to cases within strata defined by sex, year of birth (5-year inter¬ vals) and duration of smoking (10 years intervals). Blood samples were available in the biobank for 521 of the 533 selected individuals (Vogel et al. In press).
Study groups for basal cell carcinoma (basocellular carcinoma BCC). The groups of Caucasian Americans with and without BCC have been described previously (Athas et al, Cancer Res. 51 :5786-5793, 1991; Wei et al, Proc. Natl. Acad. Sci USA, 90: 1614-8, 1994). Briefly, the study was a clinic based case control study at the Johns Hopkins Hospital, which serves multiple participating dermatologists in Maryland. Cases were histo-pathologically confirmed primary BCCs and were diagnosed be¬ tween 1987 and 1990. The controls were patients from the same physician practices and had a diagnosis of mild skin disorders. All participants were Caucasians living near Baltimore and were between 20 and 60 years of age. The controls were fre- quency matched to the cases by age and sex. Cases and controls with any other forms of cancer were excluded. In the questionnaire, the study subjects were asked if they had any blood relatives with skin cancer, and were asked to specify the type of cancer. Study subjects with relatives with basal cell carcinoma and squamous cell carcinoma and 'skin cancer' were included in the group of subjects with a family of skin cancer. Subjects with relatives with melanoma were not included. At the clinic visit the subjects gave informed consent, were examined by dermatologists, com¬ pleted a structured questionnaire and provided blood. Available frozen lymphocytes were genotyped. Initially, 71 cases and 118 controls were included in this study. However, the number of persons varied between analyses, as the supply of DNAs gradually was depleted. In case of the SNP RAI intron 1 only 133 persons could be genotyped reliably.
The examples relate to prediction from sequence polymorphisms in and around the region r to cancer, see Fig. 1 for an overview of a subregion of chromosome 19q. In the present study 27 markers were investigated within this 69 kb stretch for associa¬ tion with breast, lung, skin cancer, and multiple myeloma, using linkage disequilib¬ rium mapping based on single markers, as well as, on haplotypes combining several neighboring markers. We report that haplotypes with maximal association to all three cancers are centered around the gene RAI.
Typing of SNPs
Table 10 lists the polymorphisms used in this study, their nature, the numbers in the NCBI database dbSNP, the sequence defining the SNPs, and their position therein, the currently estimated chromsome positions and their relative position within the region of interest (which can readily be calculated based on the information provided in the table). Table 11 lists the primers used for the PCR reactions. Table 12 lists the probes used for detection and typing. Table 13 lists the PCR regimen for the poly¬ morphisms using Lightcycler, Taqman and ABI 2720.
The particular sequence polymorphisms analysed in these examples are listed in Table 10, together with their sources of information and their definition as se¬ quences.
Table 10
Trivial name Kind dbSNP # Sequence Position Relative in se¬ Position quence
XPD exon 23 A/C Rs1052559 NT_011109 18123137 0
XPD exon 10 A/G Rs1799793 NT_011109 18135477 12340
XPD exon 6 A/C RS238406 NT_011109 18136527 13390 XPD intron 3 A/G Rs 1799783 NT_011109 18140254 17117
XPD-4bp GACA/- RS3916791 NT_011109 18142345 19208
XPD-81 bp 81 bp/- RS3916787 NT_011109 18143027 19890
XPD-5'2 A/G Rs2097215 NT_011109 18144005 20868
XPD-5'3 C/T Rs11878644 NT_011109 18145185 22048
RAI-3'7 C/T RS7252567 NT_011109 18146823 23686
None 1) ATTTT/- 18147012 23875
None 2) A/G RS2377329 18147126 23989
RAI-3'3 AA/- Rs3047560 NT_011109 18147192 24055
None 3) G/T 18146233
None 4) G/T rs10422489 18147886
None 5) C/T rs10426701 18148193
None 6) C/T 18149120
None 7) A/G 18149154
RAI-3'4 A/G RS4544343 NT_011109 18150199 27062
None 8) C/G 18150815
None 9) A/G RS8101662 18150911
None 10) G/T 18151158
RAI exon 6 A/T RS6966 NT_011109 18151180 28043
RAI intron5-3 C/T Rs10417235 NT_011109 18152171 29034
RyA/ intron5-2 A/T Rs8112723 NT_011109 18153497 30360
RAI intron 3 A/G Rs2017104 NT_011109 18155483 32346
RAI intron 1 A/G Rs 1970764 NT_011109 18159091 35954
None 11) A/G 18159263
None 12) A/G 18160363
None 13) A/C 18160936
None 14) C/G 18160937
None 15) A/C rs10419090 18161433
None 16) G/T 18161694
None 17) C/GACCTTGG 18161841
CGCCAC-
CATCTT
None 18) C/T 18161896
RAI intron 1-2 C/T AC092309 26729 39249 NT_011109 18162206
None 19) C/T 18162309
None 20) A/G 18162356
None 21) A/G rs12986272 18162599
None 22) A/C 18162903
R/\/ intron1-3 A/C AC092309 27493 39833
NT_011109 18162970
None 23) C/T 18162986
None 24) A/C 18163200
None 25) C/G 18165052
RAI-5'2 C/T RS4803814 NT_011109 18168944 45807
RAI-5'3 C/T RS4803815 NT_011109 18168948 45811
RAI-5' C/T Rs4572514 NT_011109 18171984 48847
ASE1-5'2 G/T Rs2226949 NT_011109 18175327 52190
ASE1 exon 1 A/G Rs967591 NT_011109 18178152 55015
ERCC1-3'2 A/G Rs735482 NT_011109 18180220 57083
ERCC1-3' C/T Rs762562 NT_011109 18180561 57424
ERCC1-3'3 A/G RS2336219 NT_011109 18180624 57487
ERCC1 exon 4 C/T Rs3177700 NT_011109 18191871 68734
rs numbers were derived from the NCBI's database dbSNP; Nucleotide sequence of polymorphisms identified in the present invention with no trivial name as yet:
I) GGTTTAT[ATTTT]Ntgagatggatttt
2) ctggggaggctgaggcaggagaatc[A/G]cttgaaaccgggaggcggaggttgt 3) GATTGTCATGT[G/T]ACATCAGCCAATACT 4) AAAAAACTAAAGTGGGGTTTGCGGG[G/ηAGTGGGAGGGCCCTTCCTGCTAGGT
5) caggcggatcacaaggtcaggagtt[C/T]gagaccagcctggccaacacagtga
6) CACAGTGAAAC[C/ηCCATCTCTACTAAA
7) AAAATTAGCCGG[A/G]CGCCATGGCGGGAG 8) AGCCTGGCCAACATG[C/G]TGAAACCCCGTCTCT 9) ctcgggaggctgaggcaggagaatc[A/G]cttgaactcaggaggcagaggttgc 10) AAGTTTCTCTATT[G/ηTGTTTATAAACA
I 1 ) TTCTCCTGACCTC[A/G]TGATCCGCCCACCTCGG
12) GGGATTACAGGCATGCCA/GJCCACCAGGCCCAGCTAAI I I I I GT
13) TCCAATGGTGACA[A/C]CAGTAAGAGCAGTTAACAG 14) TCCAATGGTGACAA[C/G]AGTAAGAGCAGTTAACAG
15) tacaggcgcccgccaccacccccag[A/C]taatttttgtatttttagtagagac
16) TTGCCTCAGCCTCCTGA[G/T]TAGCTGGGATTGGAATGAGA
17) TACGATAAATAGCTAGAtC/GACCTTGGCGCCACCATCTT] 18) AAAATAATAATAATAATATTAA[C/ηCCTGACCTTGGCGCCACCATCT
19) CCTCATGAGCCACCCAC[C/T]TCGGCCTCCCAAAGTGCT 20) TGAGCCACCGCGCCC[A/G]GCCGAGACTCACTATTT 21 ) taaagcgggaggatggcttgaacct[A/G]ggaggcggaggttgcagtgagccga 22) GGAGAGAAGGAGCAGAGAAC[A/C]TCTCTATGTGGCCA 23) TTTCCCAGCATCCCA[C/ηTGCAATGAGGCTCCTGGCC
24) TCCTGACTCCAGTG[A/C]GGTGCCTACAGTCCTG 25) TTCCAGCCTGGGCAAGAA[C/G]AGTGAAACTCCAGCTT
Statistics and haplotype assignment Data recording and calculations and tests of allele frequencies were performed in
SPSS and Excel. Calculation of the relative risk and confidence intervals for the sin¬ gle polymorphisms was performed in SAS (SAS Institute, Gary, NC, USA).
Simultaneous analysis of multiple SNPs employing haplotype trend regression (17) was performed with HelixTree (GoldenTree, Bozeman, MT, USA). Haplotype trend regression is basically a two-stage procedure. First, genotype results in combination with population assumptions such as Hardy-Weinberg equilibrium are used to con¬ struct all haplotype probabilities corresponding to a given set of markers for each individual. Secondly, the disease state (1 for cases, 0 for controls) is regressed on the haplotype probabilities of all individuals, resulting in a p-value for the overall as- sociation of the set of markers with disease, and parameters for association for each specific haplotype with disease. The HelixTree program will use sets of markers of defined size, typically 2 to 4 neighboring markers at a time, to scan the entire region. It can then be used to derive frequencies of the individual haplotypes, to calculate an overall p-value for the distribution of haplotypes covering a given set of markers among cases and controls, and to calculate p-values for the distribution of each hap¬ lotype derived from a given set of markers.
The program RASCAL for performing gene localization according to Lazzeroni (16) was implemented in Delphi (B.A. Nexø, unpublished). We used a bootstrap set of 10000 sets of haplotypes selected from the original set with replacement, and to avoid a Q-form with negative values we used a relaxation value of 0.25. A 95% con- fidence interval for the location of the causative gene variant was derived from the Q-form. Places, where it took on a value less that 3.85 above the minimum, were considered inside the confidence interval.
Three programs were used for assigning haplotypes to individuals on the bases of genotype data: HelixTree, Arlequin (18) and Phase (19). Arlequin (18) like HelixTree is a maximum likelihood algorithm, while Phase in addition includes a penalty for each new haplotype that is brought into play. Furthermore, the program Arlequin includes missing values in its table of frequencies. To compensate for the latter, we normalised the values corresponding to fully defined haplotypes before including them in the analysis. All three were run under Windows 2000. The figure of the link¬ age disequilibrium among the controls (Fig. 2) was also derived with the program HelixTree.
Primers and probes The primers used for typing the polymorphisms are listed in Table 11. Table 12 lists the probes used for the polymorphisms.
Polymorphisms listed in Table 10 as having no trivial name have been deduced by sequencing in the present study.
A person skilled in the art will appreciate that primers and probes can be designed based on the information provided herein, in order to detect polymorphism by other means than sequencing, for example by PCR ampliation methods as also used herein.
Primers were purchased from DNA Technology (Aarhus, Denmark), probes for the lightcycler were obtained from MolBiol (Tempelhofer Weg 11-12, Berlin) and the probes for Taqman were supplied by Applied Biosystem.
Table 11 Primers for the polymorphisms
Trivial name Primers
XPD exon 23 5' atg cac cag gaa ccg ttt atg g
5' tct gtt etc tgc agg agg ate XPD exon 10 5' gat caa aga gac aga cga gc
51 gaa gcc cag gaa atg c XPD exon 6 5' gta cca gca tga cac cag cct
5' tec etc cct gag ccc tg XPD intron 3 5' aag gca gac aaa gga agg
5' gca agg aga agg aac agg XPD-4bp 5' caa tea aaa aga aaa cat gg
5' tga gac gag gtg gag g XPD-81bp 5' tgc etc ace cct gta ate c 5' get tgt aat cec age tac teg
XPD-5'2 5' caa cac tea cac cec aca g
5' aga tea cgc cac tgc act c XPD-5'3 5' ttg aca att gag caa aga gc
5' tgg gat tac aga cgt gag c RAI-3'7 5' cca gtc caa aca ata tga tec
5' agt gca gcc tea act tec RAI-3'3 5' cat gat tea ctg cac cca ace
5' ttt cac tct tgt tgc cca age RAI-3'4 5' ttt tea cac aag tec aat cc
5' act gca ace tec ate tec RAI exon 6 5' cec tgc cec ace tct cc
5' agt caa ttt ctg tgc aaa eta ctt tta ttt RAI intron 5-3 5' cat gac gag ace ctg tct eta eta aa
5' cac etc ccg gat tea agt ga RAI intron 5-2 5' gag gca aca gga aca aac c
5' cat tgg att gag cag aaa cc RAI intron 3 5' taa cat aaa gaa tea gga gga ggc
5' agt tgg etc ate tgc etc tt RAI intron 1 5' tgg eta aca egg tga aac c
5' gga ate caa aga ttc tat gat gg RAI intron 1-2 5' act cct gac ttc aaa tga tec
5' tag cec cca gtc acg ttc c RAI intron 1-3 5' aga agt cca aga gtt tgc age
5' ttc tea gtc cca gaa tga ace RAI-5'2 5' cca ctt agg taa aca cct ctt
5' ctg caa tga gcc gag ata gaa RAI-5'3 5' cca ctt agg taa aca cct ctt
5' ctg caa tga gcc gag ata gaa RAI-5' 5' atg ttg ggg aga ctg agg
5' ccg cat eta act tat tct gg ASE1-5'2 5' aac tac etc tgc aaa cec age
5' ttg gaa tgg agg gat tct ace ASE1 exon 1 5' ggt ttt ctg etc tgc aca eg
5' cct ttc tec ttc cac caa eg ERCC1-3'2 5'-aca gag cec aca gtg gag aca
5'-cta gag get cag tgt taa tct gtt cct ERCC1-3' 5' gga cag atg gca atg atg g
5' tct tct tct tgg tgg atg tgg ERCCI-3'3 5'-acc atg gcg cct caa ca
5'-gaa ttg get cag tea ctg tgt ga ERCC1 exon 4 5' ggc cct gtg gtt ate aag g
5' tct cat aga aca gtc cag aac act g RAI-3'3 nested 1 5' cat gat tea ctg cac cca ace
5' ttt cac tct tgt tgc cca age RAI-3'3 nested 2 5' 6FAM-ctt gca cag tgg etc atg c
5' tct tgt tgc cca age tgg
Trivial name Primers ASE 1 exon 1 5' ggt ttt ctg etc tgc aca eg
5' cct ttc tec ttc cac caa eg
ERCCV-3'2 5'-aca gag cec aca gtg gag aca
5'-cta gag get cag tgt taa tct gtt cct ERCC1-3' 5' gga cag atg gca atg atg g
5' tct tct tct tgg tgg atg tgg ERCC1-3'3 5'-acc atg gcg cct caa ca
5'-gaa ttg get cag tea ctg tgt ga ERCC1 exon 4 5' ggc cct gtg gtt ate aag g
5' tct cat aga aca gtc cag aac act g RAI-3'3 nested 1 5' cat gat tea ctg cac cca ace 5' ttt cac tct tgt tgc cca age RAI-3'3 nested 2 5' 6FAM-ctt gca cag tgg etc atg c 5' tct tgt tgc cca age tgg
Table 12 Probes for the polymorphisms
Trivial name Probes
XPD exon 23 5' FAM-ctc tat cct ctg cag cg-MGB
5' VIC-tat cct ctt gag cgt ct-MGB
XPD exon 10 5' LC Red640- cgt get gcc caa cga agt g -p
5' gga cgc cca cct ggc caa cc -fluoresceine
XPD exon 6 5' FAM-ccc cac tgc cgc ttc tat gag gt-TAMRA
5' VIC-ccc cac tgc cga ttc tat gag gtt-TAMRA
XPD intron 3 5' LC Red640- ccc tgc ccc cca act ttg ga -p
5' gcc tec aat gaa cac aag etc -fluoresceine
XPD-Λbp 5' LC Red640- cct ggg ttc gat caa tac tea gac a -p
5' etc get ate ttg etc aag ctg ate teg aac -fluoresceine
XPD-81bp 5' agt cac age tea ctg cag cct c -fluoresceine
51 LC Red640- ace tct tgg get caa gcg ate etc -p
XPD-5'2 5' LC Red640- aaa aaa aga ctt ate atg aca gga tgt ct -p
5' gca aga etc cgt ccc aga aaa aga aaa -fluoresceine
XPD-5'3 5' LC Red640- tec tct etc tec ccc age tea ttt tg -p
5' aac cca ccc tac tgc tct gat etc -fluoresceine
RAI-3'7 5' LC Red640- agg ctg gtc ttg aac tec tgg get taa g -p
5' ggt tec gcc acg ttg cc -fluoresceine
RAI-3'3 5' LC Red640- teg get att ttt ttt ttt att ttt tta tt -p
5' att aca ggc ace cac cac cat g -fluoresceine
RAI-3'4 5' LC Red640- cct gga caa cat agg gag ace ctg tgt -p
5' caa aca aac aaa aac etc tgc ca -fluoresceine
RAI exon 6 5' 6-FAM-tgc ctt cac aca get ctg gtt taa tg - TAMRA
5' VIC - tgc ctt cac aca gca ctg gtt taa tg - TAMRA
RAI intron 5-3 5' FAM-tgg tgg tgc atg cct gta ate cc-BHQ
5' Yakima Yellow-tgg tgg tgc atg ccc gta atc-BHQ
RAI intron 5-2 5' ace atg ttg gcc agg ctg gtt tt -fluoresceine
5' LC Red640- ate tac tga cct caa atg ate cac ct -p
RAI intron 3 5' LC Red640- tgc aat ccg ccc gcc -p
5' cca ggc tgg ttt gga aat cct gag etc -fluoresceine
RAI intron 1 5' LC Red640- ctg aga teg cac cac tgc ac -p
5' ggg agg egg age ttg cag tga -fluoresceine
RAI intron 1-2 5' gcg cat gcc tgt aat tct gta -fluoresceine
5' LC Red640- cag gac gag cca cag aca aaa etc c -p
RAI intron 1-3 51 LC Red640- tgc aat gag get cct ggc c -p
5' act aca ttt ccc age ate cca -fluoresceine
RAI- 5'2 5' cct ccc tec etc cct gc -fluoresceine
5' LC Red640- tgc ttg ctt tct etc tct -p
R>4/-5'3 5' tec ctg ctt get tgc ttt etc t -fluoresceine
5' LCRed 640- tct etc ttt ctt tct ttc ttt c -p
R/A/-51 5' tgt tea tec aaa tga gcc gc -fluoresceine
51 LC Red640- age ctg aac agg ttc tgt tec ttc gac tt -p
ΛSE7-5'2 5' LC Red640- caa get get ate teg ace gat ctt -p
5' ggg tga cca ccc tgc cag cc -fluoresceine
ASE1 exon 1 5' LC Red640- egg get aca ggg tta cct gag -p
5' tct gca ace tgg tgc gag cag c -fluoresceine
ERCC?-3'2 5' FAM-aaa ggg aaa gaa ace t-MGB
5' VIC- cca aag gga cag aaa-MGB
ERCC7-3' 5' LC Red640- tga ggc tec get cct tct gg -p
5' age tgc cag age tgc ctg ggc -fluoresceine ERCC1-3'3 5' VIC -aca gca aaa tgc cac agt-MGB
5' FAM-aca gca aga tgc cac ag-MGB
ERCC1 exon 4 5' cgc aac gtg ccc tgg gaa t -fluorescein
5' LC Red640- tgq cqa cgt aat tec cga eta tgt get g -D
Table 13. PCR regimens used
Trivial Tech Buffer Buffer modifi- Denaturation Annealing Elongation name cation
XPD exon T M - 15 sec 94 0C 60 sec 60
23 0C
XPD exon L R 5% DMSO 10 sec 95 0C 15 sec 53 30 sec 72
10 0C 0C
XPD exon T M 15 sec 94 0C 60 sec 63
6 0C
XPD In- L H - 2 sec 95 0C 15 sec 30 sec 72 tron 3 55°C 0C
XPD-4bp L H 2x Buffer 10 sec 95 0C 15 sec 60 30 sec 72
0C 0C
XPD- L H 5% DMSO 10 sec 95 0C 15 sec 66 30 sec 72
81 bp 0C 0C
XPD-5'2 L H 2x dNTP + 10 sec 95 0C 15 sec 64 30 sec 72
2x Buffer 0C 0C
XPD-5'3 L H _ 2 sec 95 0C 15 sec 71 30 sec 72
0C 0C
RA/-37 L H 2 sec 95 0C 15 sec 67 30 sec 72
0C °c
RAI-3'3 L H 2x dNTP + 2 sec 95 0C 15 sec 70 30 sec 72
2x Buffer 0C 0C
RAI-3'4 L H 2 sec 95 0C 15 sec 67 30 sec 72
0C 0C
RAI exon T M 15 sec 94 C 60 sec 60
6 0C
RAI intron T M _ 15 sec 94 C 60 sec 63
5-3 0C f?<4/ intron L H 2 sec 95 0C 15 sec 60 40 sec 72
5-2 "C 0C
R/A/ intron L H 2x dNTP 2 sec 95 0C 15 sec 63 30 sec 72
3 0C 0C
R/A/ intron L H 2x Buffer 0 sec 95 0C 10 sec 57 15 sec 72
1 0C 0C
RΛ/ in¬ L H 2 sec 95 0C 15 sec 66 30 sec 72 tron 1-2 0C 0C
RyA/ intron L H - 2 sec 95 0C 15 sec 65 30 sec 72
1-3 0C 0C
R>4/-5'2 L R 2 sec 95 0C 15 sec 61 30 sec 72
0C 0C
RAI-5'3 L R 2 sec 95 0C 15 sec 61 30 sec 72
0C 0C
RAI-5' L H 1M Betain 2 sec 95 0C 15 sec 63 30 sec 72
0C 0C
ASE1-5'2 L H 5% DMSO + 2 sec 95 0C 15 sec 63 30 sec 72
1 VzX Primer 0C 0C
ASE1 L H - 2 sec 95 0C 15 sec 60 30 sec 72 exon 1 0C 0C
ERCC1- T M 15 sec 94 0C 60 sec 60
3'2 0C ERCC1-3' L H - 2 sec 95 ° C 15 sec 62 30 sec 72
0C 0C
ERCC1- T M 15 sec 94 0C 60 sec 62
3'3 0C
ERCC1 L H 5% DMSO 0 sec 95 ° C 15 sec 57 25 sec 72 exon 4 0C 0C
RAI-3'3 A H - 30 sec 94 0C 30 sec 64 30 sec 72 nested 1 0C 0C
RAI-3'3 A H - 30 sec 94 30 sec 62 30 sec 72 nested 2 0C 0C
Where R is Roche buffer, H is homebrew buffer and M is mastermix, respectively, and where T denotes Tagman technique, L is Lightcycler and A is ABI2720 se¬ quencing, respectively.
R is Roche buffer: 10 pmole primers
1 pmoie probes 1.3 nmole dNTP
1 ul Hybridization Mix for Lightcycler (Roche) 1 ul MgCI2 for Lightcycler (Roche)
1 ul DNA water to 10 ul
H is Homebrew: 10 pmole primers 1 pmole probes
1.3 nmole dNTP
0.25 ul Titanium Taq (BD Biosciences, Palo Alto, CA)
1 ul Titanium Taq Buffer (BD Biosciences, Palo AIo, CA)
1 ul DNA water to 10 ul
Determination of polymorphisms by real-time PCR using Taqman probes. Polymor¬ phisms analyzed on Taqman were scored on the basis of the relative reaction of the 2 probes. Some polymorphisms were analysed using the ABI Prism 7700 sequence detection system (Applied Biosystems, Foster City, Ca, USA). PCR Primers and Taqman probes were designed using Primer Express v 1.0 (Applied Biosystems). The reactions were performed in MicroAmp optical tubes sealed with MicroAmp op¬ tical caps (Applied Biosystems) containing a 10 μl reaction volume (Mastermix): 1x Taqman buffer A, 2.5mM MgCI2, 200 μM each of dATP dCTP, dGTP, 400μM dUTP, 80OnM each primer, 200nm each probe, 0,01 U/μL AmpErase UNG, 0,025 U/μL AmpliTaq Gold Polymerase. Tubes were incubated at 500C for 2 min followed by 10 min at 950C. The incubation was succeeded by 45 cycles with thermal cycler condi¬ tions as shown in Table 13.
Determination of polymorphisms by Lightcycler. Genotypes of the American per- sons and the Danish persons for polymorphisms in XPD exon 10, XPD-4bp, XPD- 81 bp, XPD-5'2, RAI-3'3, RAI-37, RAI intron 3, RAI intron 1 , RAI intron 1-2, RAI-5', RAI-5'2, RAI-5'3 ASE1-5'2, ASE1 exon 1 , ERCC1-3' and ERCC1 exon 4 were de¬ tected using LightCycler™ (Roche Molecular Biochemicals, Mannheim, Germany), scored on the bases on the temperature dependency ("melting profile") of the fluo- rescence. For XPD exon 10 PCR was performed by rapid-cycling in a reaction vol¬ ume of 20 μl with 0.5 μM of each primer, 0.045 μM of anchor and sensor probe, 3.5 mM MgCI2, approximately 7 - 25 ng genomic DNA, and 2 μl LightCycler DNA Master Hybridization probe buffer (Roche Molecular Biochemicals, Cat. No 2158 825). This buffer contains Taq DNA polymerase, dNTP mix, and 10 mM MgC^ and 5% DMSO. The temperature cycling consisted of denaturation at 950C for 2 sec, followed by 46 cycles consisting of 10 sec at 950C, 15 sec at 53°C, and 30 sec at 72°C, see table 4a for PCR regimen for the polymorphisms using LightCycler™. For detection of the polymorphisms XPD-4bp, XPD-81bp, XPD-5'2, RAI-3'3, RAI intron 3, RAI intron 1, RAl intron 1-2, RAI-5', ASE1-5'2, ASE1-5'2, ASE1 exon 1 , ERCC1-3' and ERCC1 exon 4 using LightCycler™ PCR was performed using Homebrew buffer, the com¬ position of which is listed in table 13a. XPD exon 10, RAI-5'2 and RAI-5'3 were run in Roche buffer, the composition of which is also listed in table 13. Buffer modifica¬ tions for the individual assay are shown in table 13 below. The temperature cycling parameters are shown for each of the reactions in table 13. The last annealing pe- riod at 72°C was extended to 120 sec followed by denaturation for 10 sec at 95°C. The melting profile was determined by a temperature ramp from 500C to 95°C with a rate of 0.1 degree/sec.
EXAMPLE 1 Pair-wise linkage disequilibrium of all markers of controls
To get an initial overview over the many markers the pair-wise linkage disequilibrium of all markers in the controls was calculated, see Fig. 2. It is clear that the region consists of 2 major haplotype blocks with a few interspersed markers in between. One haplotype block spans roughly 16 markers and 26 kb from XPD exon6 to RAI intron1-3, while the other haplotype block spans 6 markers and 9 kb from RAI-5' to ERCC1-3'3.
EXAMPLE 2 Relative risk of breast cancer
Blood samples were collected and frozen from a large number of Danish postmeno¬ pausal women who had suffered from breast cancer. An age limit of 55 years was used to separate early from late cases. The cut-off was forced by a previous deci- sion to use 50 years of age as an entrance criteria for inclusion in the cohort. The salient features of this cohort are shown in Table 14. DNAs were purified from the blood samples of these persons and the polymorphisms as shown in Table 10 were typed.
Table 14 Features of the investigated breast cancer cases and controls. Breast cancer
Background population Danes
Recruitment Population based
Epidemiological design Cohort
No. of cases 428
No. of controls 433
Age at inclusion 50-64 yrs of age
Cutoff for young cases 55 yrs of age
No. of young cases 62
The relative risk of breast cancer was calculated for the two homozygous forms of all the markers as well as the confidence intervals. This relatively simple analysis has the advantage that no a priori assumptions about the distribution of genotypes are needed. Fig. 9 depicts the relative risks and the lower confidence limits of the relative risks for postmenopausal breast cancer below age 55 as a function of the location of the markers. The values were initially calculated with the wild-type allele as reference, however, if the calculation gave a RR values less than 1 we used the reciprocal value of the RR and the reciprocal value of the high confidence limit in- stead. It is obvious that a significant association of markers with cancer exists in the region from 13 000 to 39 000 bases, corresponding to the first mentioned haplo¬ type block and covering most of the gene RAI and the 5' region of the gene XPD. The corresponding curves for the other age-brackets did not show significant re¬ gions of association (results not shown). EXAMPLE 3
Distribution of p-values for sets of markers in relation to breast cancer
As previous experience (Vogel et al. In press; Nexø et al. In press) had indicated to us that combined analyses involving multiple markers were superior to analysis of individual markers, we used the analytical technique known as haplotype trend re¬ gression (Zaykin et al. 2002), more specifically the program HelixTree, to associate sets of markers with the individual diseases. Haplotype trend regression is basically a two-stage procedure. First, genotype results in combination with population as- sumptions such as Hardy-Weinberg equilibrium are used to construct all haplotype probabilities corresponding to a given set of markers for each individual. Secondly, the disease state (1 for cases, 0 for controls) is regressed on the haplotype prob¬ abilities of all individuals, resulting in a p-value for the overall association of the set of markers with disease, and parameters for association for each specific haplotype with disease. The HelixTree program will use sets of markers of defined size, typi¬ cally 2 to 4 neighboring markers at a time, to scan the entire region. It can then be used to derive frequencies of the individual haplotypes, to calculate an overall p- value for the distribution of haplotypes covering a given set of markers among cases and controls, and to calculate p-values for the distribution of each haplotype derived from a given set of markers.
The entire sets of controls were used for the analyses. This broke the matching of the data sets, but matching on other criteria than ethnicity and age may well be ir¬ relevant for genetic association studies (Hemminki & Forsti 2001) (21).
Fig. 4 shows the overall distribution of p-values for sets of markers plotted against the position on the chromosome for breast cancer. As position on the abscissa we used the median marker position in a given set of markers. The ordinate values are the negative logarithms to the overall p-values for a difference between cases and controls associated with a given set of markers. Each curve corresponds to a given size of marker sets, i.e. the number of markers in the haplotypes.
The results in Fig. 4 suggest that the association with breast cancer has two peaks. One was located at roughly 24 kb and one at 39 kb. When the cases were broken down into early and late breast cancer the peak at 24 kb was present in both groups, while the peak at 39 kb was only present in the older group. The peaks were clearly present in curves for haplotypes of 2, 3 and 4 neighboring markers. Larger sets were less informative, presumably due to increased degrees of freedom, but they essentially corroborated the results (results not shown).
Remarkably, this association at 24 kb was present for breast cancer both in young and in older female cases. This is the first time an association to RAI has been re¬ ported for cancers in older populations. The position of the optimal set of markers corresponds to the 3' part of the gene RAI and the inter-gene region between RAI and XPD and was identical for young and older cancer cases. The peak at 39 kb which is only present in the older breast cancers correspond to the 5' part of RAI.
EXAMPLE 4
Frequency of individual haplotypes made up from 2 neighbouring markers in relation to breast cancer.
To understand better the nature of the association, we tabulated the frequencies of individual haplotypes made up of 2 neighboring markers centring on the position at 24 kb with maximal association, in various groups of cases and controls (RAI-3'3 RAI-3'4; Table 15). We also tabulated p-values associated with the individual haplo- types. It was the same haplotypes that differed whin young and older breast cancers were compared to controls (RAI-3'3S RAI-3'4C down; RAI-3'31 RAI-3'4* up). A simi¬ larity between patterns in the different cancer forms was not obvious, however, the frequencies of the haplotypes in the control groups were similar.
Similarly, we calculated haplotypes for the set of 2 markers, located at 39 kb. Again the deviation of young and old breast cancer cases from the controls had similarity (RAI intron1-29 RAI intron1-3a up; RAI intron1-2g RAI intron1-3° down). The same pattern was present in the lung cancer data. This is remarkable, as neither the young breast cancers nor the lung cancers showed a peak in this position, and sug- gests an underlying similarity of haplotypes in the different groups.
Table 15. Frequency of 2-haplotypes of cancer cases and controls and their asso¬ ciation with cancer p-values in parentheses relative to breast cancer controls, all ages
p-values in parentheses relative to skin cancer controls, all ages
3) p-value could not be calculated
The assignment of frequencies to individual haplotypes also made it possible to de¬ rive a measure of association which was stochastically independent of group sizes. Under the assumption that a putative causative variant must be at least equally as associated with disease as any surrogate haplotype we calculated the odds ratios for each haplotype of each set of two neighbouring markers versus the three other haplotypes of the set and chose the maximum value. This value we then plotted against the median position of the set (Fig. 5). The resulting curves indicate a maxi¬ mal association of the set RAI-3'3 [RAI-3'4 at 25 kb with a value of approximately 4.4 for the young women, with rapid fall-off at either side. The curve for older women as expected had two maxima, one at the same place as the young women with a maximum of 7.3, and one associated with the set RAI intron1-2 RAI intron1-3 with a maximum of the same height.
To make sure that the derived haplotype frequencies did not reflect a peculiarity of the particular program used, we also calculated the frequencies of individual haplo¬ types in cases from breast cancer and basal cell carcinoma combined and in the corresponding controls, using 2 other computer programs, Arlequin and Phase. Ar- lequin (18) like HelixTree is a maximum likelihood algorithm, while Phase in addition includes a penalty for each new haplotype that is brought into play. Furthermore, the program Arlequin includes missing values in its table of frequencies. To compensate for this, we normalised the values corresponding to fully defined haplotypes before including them in the analysis. It was evident that results from all 3 programs were similar. In particular, HTR and Arlequin produced essentially identical results (0 to 2 percent difference, except when the values were essentially zero). The results ob¬ tained with Phase occasionally diverged 5 to 10 percent of the value, presumably a reflection of the different optimization criterion, but the change appeared to affect both cases and controls. Therefore, we are confident that our biological conclusions are independent of which computer program was used (results not shown).
We also assigned long haplotypes to all individuals of the breast cancer cases and controls using the program Phase (Stephens et al. 2001 ). To avoid overloading this program and to remove SNPs of little importance we reduced the set to 22 SNPs, by eliminating those with lowest heterogeneity. The resulting haplotypes were used as input to a program implementing Lazzeroni's technique (Lazzeroni 1998) (16) for locating gene variation (RASCAL, BA Nexø, unpublished). We used a bootstrap set of 10000 sets of haplotypes selected from the original set with replacement, and to avoid a Q-form with negative values we used a relaxation value of 0.25. With these values we found the most likely location to be at 20 kb very close to the marker %PD_81 bp for the early breast cancers, (Fig. 6). Changing the relaxation factor to 0.5 moved the most likely location to 23 kb bases. A 95% confidence interval for the location of the causative gene variant was derived from the quadratic form for the residual variance. Places, where it took on a value less that 3.85 above the mini¬ mum were considered inside the confidence interval. The curve suggested that the Cl stretches from approximately 10 kb to approximately 57 kb. Using the same pa¬ rameters, we found an almost identical position for the gene variant among the late breast cancers, however in this case no part of the region could be statistically ex- eluded as a possible location for the variant.
EXAMPLE 5
Distribution of p-values for sets of markers in relation to early basal cell carcinoma
DNA was isolated from lymphocytes from humans from the American cohort of pa¬ tients with basal cell carcinoma and controls, described in Materials and Methods, was typed with respect to a number of sequence polymorphisms located in and around the claimed region r. The salient features of the Americant cohort are shown in Table 16.
Table 16. Features of the investigated basal cell carcinoma cases and con¬ trols.
Basal cell carcinoma
Background population Caucasian Americans
Recruitment Hospital based
Epidemiological design Case-Control
No. of cases 71
No. of controls 118
Age at inclusion 29 - 60
Cutoff for young cases 50 yrs of age
No. of young cases 45
Figure 7 shows the same kind of result for early basal cell carcinoma. Again we found a strong association with the region around RAI, with the curve for sets of 3 markers possibly shifted slightly to the right, and the curve for 4 markers showing suggestions of an extra association at higher positions.
EXAMPLE 6
Distribution of p-values for sets of markers in relation to lung cancer
DNA from humans from the Danish cohort of patients with lung cancer and controls was typed with respect to a number of sequence polymorphisms located in and around the claimed region r.
Evidence exists that a separate effector is associated with the marker XPD exon 23 which influences the lung cancer risk and thus may interfere with accurate mapping.
Therefore the Danish men and women who had suffered from lung cancer were stratified according to the value of the marker XPD exon 23. The salient features of the Danish cohort are shown in Table 17. Table 17. Features of the investigated lung cancer cases and controls.
Lung cancer Lung cancer w/ XPD exon23AA
Background population Danes Danes
Recruitment Population based Population based
Epidemiological design Cohort Subcohort
No. of cases 249 83
No. of controls 259 115
Age at inclusion 50 - 64 yrs of age 50 - 64 yrs of age
Cutoff for young cases 56 yrs of age 56 yrs of age
No. of young cases 33» 9
To minimize the interference of this second effector, we stratified the young popula¬ tion for XPD exon 23, and calculated overall p -values for each value of XPD exon23. Fig. 7 shows that also this disease is associated with sets of markers span¬ ning the distal end of RAI and the inter-gene region in the population with the XPD exon23AA genotype. This is the group of patients free of influence of the XPD gene. No association was evident in the two other groups (results not shown).
EXAMPLE 7
The results obtained in examples 1 , and 3 to 6 rely on the use of the HelixTree pro- gram. In an attempt to control for this dependency the odds ratios for test of the two repective homozygotes of each individual marker against cancer status were plotted against the marker position on the chromosome 19, see Fig. 8. Where one of the four homozygotic groups was empty heterozygotes were included in the analysis. This relatively simple analysis has the advantage that no assumptions about the phase of different markers are necessary.
The data for early cancers show a sharp maximum in the marker RAI introni, and breast cancer in addition a minor peak in XPD_4bp. Among the late cancers there was no striking association with disease, but mapping with single SNPs is also not expected to be particular sensitive (results not shown).
EXAMPLE 8
The ASE-1 genotype influences relapse-free survival and survival in Multiple Mye¬ loma 391 patients diagnosed with multiple myeloma were found eligible for transplantation in Denmark in the period 1993-2004.
304 patients were genotyped for the polymorphisms ERCC1 exon4, RAI introni, ASE-1 e1 , XPD exon23, XPD exon10. Clinical parameters were available for 360 patients.
The multiple myeloma patients were treated with high dose alkylating chemotherapy followed by autologous bone marrow transplantation. The overall survival was sig- nificantly shorter for patients who experienced relapse than for patients who did not.
ASE-1 genotype was found to strongly influence the event-free survival (EFS) and overall survival in patients with multiple myeloma who were auto-transplanted. Thus, homozygous carriers of the wild-type allele had mean EFS of 1160 days whereas carriers of the variant allele had a mean EFS of 1718 days (p= 0.0018). Thus variant allele carries had 1.5 year longer event-free survival, see Fig. 9. The EFS was simi¬ lar for men and women, although the difference was only statistically significant for women (Table 18). This means that the ASE-1 genotype can predict who will benefit most from the present treatment regimen.
The overall survival was also found to be influenced by the ASE-1 genotype. Thus, homozygous carriers of the wild-type allele had an overall survival of 2117 days whereas carriers of the variant allele lived 2727 days, which is 1.67 years longer (p=0.029), see Fig. 10. There was no gender effect but the difference was only sta- tistically significant for women.
This means that the ASE-1 genotype can predict who will benefit the most from the present treatment regimen.
Table 18
Event-free survival and overall survival for patients with different ASE-1 genotypes
Group Median event-free sur- p Median overall sur- P vival, EFS, (5-95% Cl) vival (5-95 % Cl)
_
GG 736 (634, 838) 0.0018 1902 (1613, 2191) 0.0293 AA+AG 1397 (901, 1893) 3015 (2093, 3937)
Men
GG 736 (616,856) 0.106 2030 (1741, 2404) 0.2297
AA+AG 963 (693,1233) 2747 (2377, 3077)
Women
GG 714 (426, 1002) 0.0075 1897 (1384, 2410) 0.0368
AA+AG 1479 (1019, 2939) 3015 (2261 , 3764)
As Fig. 11 illustrates, female carriers of the variant allele of ASE-1 polymorphism had a longer relapse-free period than homozygous carriers of the wild-type allele (median 1479 days vs. 714 days), while Figure 12 illustrates that they lived longer (median 3015 vs. 1897). Similar analyses for the men revealed the same tenden¬ cies but no statistically significant differences (median 963 vs 736 days for the re¬ lapse-free period; mean 2747 vs. 2030 days for the survival), Fig. 13 and Fig. 14, respectively.
EXAMPLE 9
Survival from time of diagnosis and genotypes of relevant genes were determined for 432 lung cancer patients who had been enrolled in the prospective study 'Diet, Cancer and Health'.
High-risk haplotype carriers (defined as ERCC1 exon 4 M, ASE1 exon1GG, RAI In- tron 1 M) had a mean survival of 359 days compared to a mean survival of 244 for non-carriers (Table 19 and Fig. 15). The genotype XPD K751Q was also deter¬ mined.
Table 19.
Genotype N Survival (days)
High Risk carrier.
Yes 105 359.3 +/- 37.1
No 319 244.5 +/- 17.5 0.002 Missing 8
XPD K751Q
AA 146 330.8 +/- 33.8 0.053
AC 218 248.9 +/- 20.7
CC 64 255.8 +/- 31.6
Missing 4 a) p for cox regression
When the two genotypes were combined, there were significant differences in sur¬ vival between carriers of different genotype combinations, see Table 20. Thus, ho¬ mozygous carriers of both favorable genotypes had a mean survival of 415 days whereas carrier of an unfavorable had a lower mean survival time, see Fig.16 and Fig. 17, respectively.
Table 20.
High-risk carrier XPD K751Q N Survival (days) P
No AA 101 296 +/- 37 0.001 a
No AC 164 222 +/- 24
No CC 51 217 +/- 29
Yes AA 44 415 +/- 74
Yes AC 51 318 +/- 42
Yes CC 10 346 +/- 83
Missing 11
a) the p-value is for a cox regression. P for high-risk is 0.004 and for XPD K751Q p=0.035.
EXAMPLE 10 Odds ratios for alleles in young cases vs controls
In analogy with the methods employed and results obtained in examples 1-9 the alleles of the markers listed below in table 20 have been found to have the indi¬ cated odds ratios for breast cancer in young cases versus controls when determined by sequencing of the region in approximately 10 cases and 10 controls. Further- more, to make sure that an odds ratio could be calculated in all instances (i.e. that the divisor was never 0) we added a value of 0.5 to each of the four alleles occur¬ rences. Table 20

Claims

Claims
1. A method for estimating the disease risk of an individual comprising
- in a sample from said individual assessing in the genetic material a se¬ quence polymorphism
- in a region corresponding to SEQ ID NO: 1, or a part thereof, or
- in a region complementary to SEQ ID NO: 1 , or a part thereof, or
- in a transcription product from a sequence in a region corresponding to SEQ ID NO: 1 , or a part thereof, or
- or translation product from a sequence in a region corresponding to SEQ ID NO: 1 , or a part thereof,
- obtaining a sequence polymorphism response,
- estimating the disease risk of said individual based on the sequence polymor¬ phism response.
2. The method according to claim 1 , wherein the cell sample is a blood sample, a tissue sample, a sample of secretion, semen, ovum, a washing of a body sur- face, such as a buccal swap, a clipping of a body surface, including hairs and nails.
3. The method according to any of the preceding claims, wherein the cell is se¬ lected from white blood cells and tumor tissue.
4. The method according to any of the preceding claims, wherein the sequence polymorphism comprises at least one mutation base change.
5. The method according to any of the preceding claims, wherein the sequence polymorphism comprises at least two base changes.
6. The method according to any of the preceding claims, wherein the sequence polymorphism comprises at least one single nucleotide polymorphism.
7. The method according to any of the preceding claims, wherein the sequence polymorphism comprises at least two single nucleotide polymorphisms.
8. The method according to any of the preceding claims, wherein the sequence polymorphism comprises at least one tandem repeat polymorphism.
9. The method according to any of the preceding claims, wherein the sequence polymorphism comprises at least two tandem repeat polymorphisms.
10. The method according to any of the preceding claims, wherein the sequence polymorphism comprises at least one deletion polymorphism.
11. The method according to any of the preceding claims, wherein the sequence polymorphism comprises at least one insertion polymorphism.
12. The method according to any of the preceding claims, wherein the sequence polymorphism comprises at least one dinucleotide polymorphism.
13. The method according to any of the preceding claims, wherein the cancer is se- lected from skin carcinoma including malignant melanoma, breast cancer, lung cancer, colon cancer and other cancers in the gastro-intestinal tract, prostate cancer, lymphoma, leukemia, multiple myeloma, pancreas cancer, head and neck cancer, ovary cancer and other gynecological cancers.
14. The method according to any of the preceding claims, wherein the cancer is se¬ lected from skin cancer, lung cancer, colon cancer, breast cancer, and multiple myeloma.
15. The method according to any of the preceding claims, wherein the cancer is se- lected from skin cancer, lung cancer, colon cancer and breast cancer.
16. The method according to any of the preceding claims, wherein the cancer is se¬ lected from skin cancer, lung cancer, and breast cancer.
17. The method according to any of the preceding claims, wherein the cancer is lung cancer.
18. The method according to any of the preceding claims, wherein the cancer is lung cancer in an individual with marker XPD exon 23 M.
19. The method according to any of the preceding claims, wherein the cancer is se¬ lected from skin cancer and breast cancer.
20. The method according to any of the preceding claims 13-16 and 19, wherein the skin cancer is basal cell carcinoma.
21. The method according to any of the preceding claims, wherein the cancer is multiple myeloma.
22. The method according to any of the preceding claims, wherein the assessment is conducted by means of at least one nucleic acid primer or probe, such as a primer or probe of DNA, RNA or a nucleic acid analogue such as peptide nucleic acid (PNA) or locked nucleic acid (LNA).
23. The method according to claim 22, wherein the nucleotide primer or probe is capable of hybridising to a subsequence of the region corresponding to SEQ ID NO: 1 , or a part thereof, or a region complementary to SEQ ID NO:1.
24. The method according to claim 22 or 23, wherein the primer or probe has a length of at least 9 nucleotide or peptide monomers.
25. The method according to any of the preceding claims 22-24, wherein at least one nucleotide primer or probe is capable of hybridising to a subsequence se- lected from the group of subsequences
1. gctgcagtga gctgt (-/ACACCTGTGGTCCCAGCTACTCTGG AAGCTGAGGTGGGAGGATCGCTTGAGCCCAAGAGGTG- GAGGCTGCAGTGAGCTGT) gactgtgcca ctgcactcca 2. TGACAGTAGA CATCCTGTCA T (A/G) ATAAGTCttt ttttttt 3. CATCCCCATA CCAAcccacc (c/t) tactgctctg atctcctcct
4. ttttagtagagacatggttccgcca (C/T) gttgcccaggctggtcttgaactc
5. ACTAAAAATAAAAAATAAAAAAAA(-/AA) ATAGCCGAG- CATGGTGGTGGGGTGC 6. AAAAAACTAAAGTGGGGTTTGCGGG (G/T)
AGTGGGAGGGCCCTTCCTGCTAGG
7. ggatcacaag gtcaggagtt (c/t) gagaccagcc tggccaacac
8. TGTTGTCCAA GCTGGCAGAG (A/G) TTTTTGTTTG TTTGTTTGAG 9. CTCGGGAGGCTGAGGCAGGAGAATC (AZG) CTTGAACTCAGG-
CAGAGGTTG
10. ATTAAGTGCCTTCACACAGC (AAT) CTGGTTTAAT GTTTATAA
11. ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa
12. gggaggctcg aggcgggc (A/G) gattgcatga gctcaggatt 13. tgcagtgagc tgagatcgc (A/G) ccactgcact ccagcctggg
14. CTTGCTACAGAATTACAGGCA (T/C) GCGCCACCGCTCCGGGCTAA
15. CTAAAGACTACA (-/A) TTTCCCAGCATCCCATTG
16. CCCTCCCTGCTTGCTTGCTTTCTCT [C/η TCTCTCTTTCTTTCTTTCTTTCTT
17. CCCTGCTTGCTTGCTTTCTCTCTCT [C/T] TCTTTCTTTCTTTCTTTCTTTCTT
18. AGAACCTGTTCAGGCTGGCGGCTCA [C/η TTGGAT- GAACAGGGAGTGTGTGAC 19. CCCCCTTCTTAGGACGCATGGGGGT [GAT] GAGA-
GAACGGGGAGATAGACAGAG
20. TGCGAGCAGCCCGGGCTACAGGGTT [A/G] CCTGAGGTGTGGGTCCCAGGATGG
21. GGCGCCTCAACAGCCAGAAGGAGCG [A/G] AGCCTCAGGCCCAGGCAGCTCTGG
or to a sequence complementary to any of the subsequences.
26. The method according to claim 25, wherein at least one nucleotide primer or probe is capable of hybridising to a subsequence selected from the group of subsequences
1. TGTTGTCCAA GCTGGCAGAG (A/G) TTTTTGTTTG TTTGTTTGAG 2. CTCGGGAGGCTGAGGCAGGAGAATC (A/G) CTTGAACTCAGGCA
GAGGTTG
3. ATTAAGTGCCTTCACACAGC (A/T) CTGGTTTAAT GTTTATAA
4. ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa
5. gggaggctcg aggcgggc (A/G) gattgcatga gctcaggatt 6. tgcagtgagc tgagatcgc (A/G) ccactgcact ccagcctggg
7. CTTGCTACAGAATTACAGGCA (T/C) GCGCCACCGCTCCGGGCTAA
8. CTAAAGACTACA (-/A) TTTCCCAGCATCCCATTG
or to a sequence complementary to any of the subsequences.
27. The method according to claim 25, wherein at least one nucleotide primer or probe is capable of hybridising to a subsequence selected from the group of subsequences
lgctgcagtga gctgt (-/ACACCTGTGGTCCCAGCTACTCTGG.
AAGCTGAGGTGGGAGGATCGCTTGAGCCCAAGAGGTGGAGGCTG-
CAGTGAGCTGT) gactgtgcca ctgcactcca
2.TGACAGTAGA CATCCTGTCA T (A/G) ATAAGTCttt ttttttt 3. CATCCCCATA CCAAcccacc (c/t) tactgctctg atctcctcct
4. ttttagtagagacatggttccgcca (C/T) gttgcccaggctggtcttgaactc
5. ACTAAAAATAAAAAATAAAAAAAA(-/AA) ATAGCCGAG- CATGGTGGTGGGGTGC
6. AAAAAACTAAAGTGGGGTTTGCGGG (G/T) AGTGGGAGGGCCCTTCCTGCTAGG
7. ggatcacaag gtcaggagtt (c/t) gagaccagcc tggccaacac
8. TGTTGTCCAA GCTGGCAGAG (A/G) TTTTTGTTTG TTTGTTTGAG
9. CTCGGGAGGCTGAGGCAGGAGAATC (A/G) CTTGAACTCAGGCA- GAGGTTG 10. ATTAAGTGCCTTCACACAGC (A/T) CTGGTTTAAT GTTTATAA 11. ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa
12. gggaggctcg aggcgggc (A/G) gattgcatga gctcaggatt
13. tgcagtgagc tgagatcgc (A/G) ccactgcact ccagcctggg
14. CTTGCTACAGAATTACAGGCA (T/C) GCGCCACCGCTCCGGGCTAA 15. CTAAAGACTACA (-/A) TTTCCCAGCATCCCATTG
or to a sequence complementary to any of the subsequences.
28. The method according to claim 25, wherein at least one nucleotide primer or probe is capable of hybridising to a subsequence selected from the group of subse¬ quences
1. gctgcagtga gctgt (-/ACACCTGTGGTCCCAGCTACTCTGG AAGCTGAGGTGGGAGGATCGCTTGAGCCCAAGAGGTGGAGGCTG- CAGTGAGCTGT) gactgtgcca ctgcactcca
2. TGACAGTAGA CATCCTGTCA T (A/G) ATAAGTCttt ttttttt
3. CATCCCCATA CCAAcccacc (c/t) tactgctctg atctcctcct
4. ttttagtagagacatggttccgcca (C/T) gttgcccaggctggtcttgaactc
5. ACTAAAAATAAAAAATAAAAAAAA(-/AA) ATAGCCGAG- CATGGTGGTGGGGTGC
6. AAAAAACTAAAGTGGGGTTTGCGGG (GfT) AGTGGGAGGGCCCTTCCTGCTAGG
7. ggatcacaag gtcaggagtt (c/t) gagaccagcc tggccaacac
8. TGTTGTCCAA GCTGGCAGAG (A/G) TTTTTGTTTG TTTGTTTGAG 9. CTCGGGAGGCTGAGGCAGGAGAATC (A/G) CTTGAACTCAGGCA-
GAGGTTG
10. ATTAAGTGCCTTCACACAGC (AfT) CTGGTTTAAT GTTTATAA
11. ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa
12. gggaggctcg aggcgggc (A/G) gattgcatga gctcaggatt
or to a sequence complementary to any of the subsequences.
29. The method according to claim 25, wherein at least one nucleotide primer or probe is capable of hybridising to a subsequence selected from the group of subse- quences 1. ACTAAAAATAAAAAATAAAAAAAA(-/AA) ATAGCCGAG- CATGGTGGTGGGGTGC
2. AAAAAACTAAAGTGGGGTTTGCGGG (G/T) AGTGGGAGGGCCCTTCCTGCTAGG 3. ggatcacaag gtcaggagtt (c/t) gagaccagcc tggccaacac
4. TGTTGTCCAA GCTGGCAGAG (A/G) TTTTTGTTTG TTTGTTTGAG
5. CTCGGGAGGCTGAGGCAGGAGAATC (A/G) CTTGAACTCAGG- CAGAGGTTG
6. ATTAAGTGCCTTCACACAGC (A/T) CTGGTTTAAT GTTTATAA 7. ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa
8. gggaggctcg aggcgggc (A/G) gattgcatga gctcaggatt
or to a sequence complementary to any of the subsequences.
30. The method according to claim 25, wherein at least one nucleotide primer or probe is capable of hybridising to a subsequence selected from the group of subse¬ quences
1. ACTAAAAATAAAAAATAAAAAAAA(-/AA) ATAGCCGAG- CATGGTGGTGGGGTGC
2. ATTAAGTGCCTTCACACAGC (A/T) CTGGTTTAAT GTTTATAA
3. gggaggctcg aggcgggc (A/G) gattgcatga gctcaggatt
or to a sequence complementary to any of the subsequences.
31. The method according to claim 25, wherein at least one nucleotide primer or probe is capable of hybridising to a subsequence selected form the roup of subse¬ quences
1. gctgcagtga gctgt (-/ACACCTGTGGTCCCAGCTACTCTGG AAGCTGAGGTGGGAGGATCGCTTGAGCCCAAGAGGTGGAGG-
CTGCAGTGAGCTGT) gactgtgcca ctgcactcca
2. CATCCCCATA CCAAcccacc (c/t) tactgctctg atctcctcct
3. ttttagtagagacatggttccgcca (C/T) gttgcccaggctggtcttgaactc
4. ACTAAAAATAAAAAATAAAAAAAA(-/AA) ATAGCCGAG- CATGGTGGTGGGGTGC 5. AAAAAACTAAAGTGGGGTTTGCGGG (GfT) AGTGGGAGGGCCCTTCCTGCTAGG
6. ggatcacaag gtcaggagtt (c/t) gagaccagcc tggccaacac
7. TGTTGTCCAA GCTGGCAGAG (A/G) TTTTTGTTTG TTTGTTTGAG
8. CTCGGGAGGCTGAGGCAGGAGAATC (A/G) CTTGAACTCAGG- CAGAGGTTG
9. ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa
10. CTTGCTACAGAATTACAGGCA (T/C) GCGCCACCGCTCCGGGCTAA
11. CTAAAGACTACA (-/A) TTTCCCAGCATCCCATTG
or to a sequence complementary to any of the subsequences.
32. The method according to claim 25, wherein at least one nucleotide primer or probe is capable of hybridising to a subsequence selected form the group of subse¬ quences
1. CATCCCCATA CCAAcccacc (c/t) tactgctctg atctcctcct 2. ttttagtagagacatggttccgcca (C/T) gttgcccaggctggtcttgaactc
3. ACTAAAAATAAAAAATAAAAAAAA(-/AA) ATAGCCGAG- CATGGTGGTGGGGTGC
4. AAAAAACTAAAGTGGGGTTTGCGGG (G/T) AGTGGGAGGGCCCTTCCTGCTAGG 5. ggatcacaag gtcaggagtt (c/t) gagaccagcc tggccaacac
6. TGTTGTCCAA GCTGGCAGAG (A/G) TTTTTGTTTG TTTGTTTGAG
7. CTCGGGAGGCTGAGGCAGGAGAATC (A/G) CTTGAACTCAGG- CAGAGGTTG
8. ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa
or to a sequence complementary to any of the subsequences.
33. The method according to claim 25, wherein at least one nucleotide primer or probe is capable of hybridising to a subsequence selected form the roup of subse- quences 1. CATCCCCATA CCAAcccacc (c/t) tactgctctg atctcctcct
2. ttttagtagagacatggttccgcca (C/T) gttgcccaggctggtcttgaactc
3. ACTAAAAATAAAAAATAAAAAAAA(-/AA) ATAGCCGAG- CATGGTGGTGGGGTGC
4. AAAAAACTAAAGTGGGGTTTGCGGG (GfT) AGTGGGAGGGCCCTTCCTGCTAGG
5. ggatcacaag gtcaggagtt (c/t) gagaccagcc tggccaacac
6. TGTTGTCCAA GCTGGCAGAG (A/G) TTTTTGTTTG TTTGTTTGAG 7. CTCGGGAGGCTGAGGCAGGAGAATC (AZG) CTTGAACTCAGG-
CAGAGGTTG
or to a sequence complementary to any of the subsequences.
34. The method according to claim 25, wherein at least one nucleotide primer or probe is capable of hybridising to a subsequence selected form the roup of subse¬ quences
1. CTTGCTACAGAATTACAGGCA (T/C) GCGCCACCGCTCCGGGCTAA
2. CTAAAGACTACA (-/A) TTTCCCAGCATCCCATTG
3. ttgggagacc aaggcaggtg gate (a/t) tttgaggtca gtagatcaaa
or to a sequence complementary to any of the subsequences.
35. The method according to any of the preceding claims, wherein at least one se¬ quence polymorphism is assessed in a region corresponding to SEQ ID NO: 1 posi¬ tion 7760-22885 (RAI).
36. The method according to any of the preceding claims, wherein at least one se¬ quence polymorphism is assessed in a region corresponding to SEQ ID NO: 1 posi¬ tion 1 -15698.
37. The method according to any of the preceding claims, wherein at least one se¬ quence polymorphism is assessed in a region corresponding to SEQ ID NO: 1 posi¬ tion 4528 -15698.
38. The method according to any of the preceding claims, wherein at least one se¬ quence polymorphism is assessed in a region corresponding to SEQ ID NO: 1 posi¬ tion 34391- 37752.
39. The method according to any of the preceding claims, wherein at least one se- quence polymorphism is assessed in a region corresponding to SEQ ID NO: 1 posi¬ tion 4897-12090.
40. The method according to any of the preceding claims, wherein at least one se¬ quence polymorphism is assessed in a region corresponding to SEQ ID NO: 1 posi- tion 1-1510.
41. The method according to any of the preceding claims, wherein at least one se¬ quence polymorphism is assessed in a region corresponding to SEQ ID NO: 1 posi¬ tion 1710-8685.
42. The method according to any of the preceding claims, wherein at least one se¬ quence polymorphism is assessed in a region corresponding to SEQ ID NO: 1 posi¬ tion 8887-12090.
43. The method according to any of the preceding claims, wherein at least one se¬ quence polymorphism is assessed in a region corresponding to SEQ ID NO: 1 posi¬ tion 15898-25550.
44. The method according to any of the preceding claims, wherein at least one se- quence polymorphism is assessed in a region of SEQ ID NO:1 flanked by and in¬ cluding the markers RAI 3'7 and RAI 3'4.
45. The method according to any of the preceding claims, wherein at least one se¬ quence polymorphism is assessed in a region of SEQ ID NO:1 flanked by and in- eluding the markers RAI 3'3 and the marker sequence GGTTTATfATTTηNtgagatggatttt.
46. The method according to any of the preceding claims, wherein at least one se¬ quence polymorphism is the marker sequence GGTTTAT[ATTTT]Ntgagatggatttt.
47. The method according to any of the preceding claims, wherein at least one se¬ quence polymorphism is RAI 3'3.
48. The method according to any of the preceding claims, wherein at least one se¬ quence polymorphism is assessed in a region of SEQ ID NO:1 flanked by and in¬ cluding the marker RAI intron 1 and the marker sequence: TTCCAGCCTGGGCA- GAA[C/G]AGTGAAACTCCAGCTT.
49. The method according to any of the preceding claims, wherein at least one se¬ quence polymorphism is assessed in a region of SEQ ID NO:1 flanked by and in¬ cluding the markers RAI intron 1 and RAI intron 1-3.
50. The method according to any of the preceding claims, wherein at least one se- quence polymorphism is assessed in a region of SEQ ID NO:1 flanked by and in¬ cluding the marker sequence: GGGATTACAGGCATGC[A/G]CCACCAGG- CCCAGCTAATTTTTGT and the marker rs10419090.
51. The method according to any of the preceding claims, wherein at least one se- quence polymorphism is the marker sequence: GGGATTACAGG-
CATGC[A/G]CCACCAGGCCCAGCTAATTTTTGT
52. The method according to any of the preceding claims, wherein at least one se¬ quence polymorphism is the marker sequence: taaagcgggag- gatggcttgaacct[A/G]ggaggcggaggttgcagtgagccga
53. The method according to any of the preceding claims, wherein at least two diffe¬ rent probes are used, one or more probes being selected from the probes as de¬ fined in any of claims 25-34, and one or more probes being capable of hybridising to a sequence different from SEQ ID NO: 1 , or a part thereof, or to a sequence com¬ plementary to a region different from SEQ ID NO: 1, or a part thereof.
54. The method according to claim 1 , wherein the translational product from a se- quence in a region corresponding to SEQ ID NO: 1, or a part thereof, is an antibody, such as a monoclonal or polyclonal antibody.
55. A method for estimating the disease prognosis of an individual comprising
- in a sample from said individual, assessing in the genetic material a sequence polymorphism
- in a region corresponding to SEQ ID NO: 1 , or a part thereof, or
- in a region complementary to SEQ ID NO: 1 , or a part thereof, or
- in a transcription product from a sequence in a region corresponding to SEQ ID NO: 1 , or a part thereof, or
- or translation product from a sequence in a region corresponding to SEQ ID NO: 1 , or a part thereof,
- obtaining a sequence polymorphism response,
- estimating the disease prognosis of said individual based on the sequence polymorphism response.
56. The method according to claim 55, wherein the method has any of the features as defined in any of the claims 2-54.
57. A method for estimating a treatment response of an individual suffering from cancer to a disease treatment, comprising
- in a sample from said individual, assessing in the genetic material a sequence polymorphism - in a region corresponding to SEQ ID NO: 1 , or a part thereof, or
- in a region complementary to SEQ ID NO: 1 , or a part thereof, or
- in a transcription product from a sequence in a region corresponding to SEQ ID NO: 1 , or a part thereof, or
- or translation product from a sequence in a region corresponding to SEQ ID NO: 1 , or a part thereof, - obtaining a sequence polymorphism response,
- estimating the individual's response to the disease treatment based on the sequence polymorphism response.
58. The method according to claim 57, wherein the method has any of the features as defined in any of the claims 2-54.
59. A primer or probe for use in a method as defined in any of the claims above, said primer or probe being selected from the group of nucleotides
1. agt cac age tea ctg cag cct c
2. ace tct tgg get caa gcg ate etc
3. aaa aaa aga ctt ate atg aca gga tgt ct
4. gca aga etc cgt ccc aga aaa aga aaa 5. tcctctctctcccccagctcattttg
6. ACCCCACCCTACTGCTCTGATCTC
7. agg ctg gtc ttg aac tec tgg get taa g
8. ggt tec gcc acg ttg cc
9. teg get att ttt ttt ttt att ttt tta tt 10. att aca ggc ace cac cac cat g
11. ace cca ctt tag ttt ttt ttt cct eta gtg ate gcc
12. gcc etc cca eta ccc gca
13. aga egg ggt ttc act gtg ttg gc
14. agg ctg gtc tea aac tec tga c 15. CAAACAAACAAAAACCTCTGCCA
16. CTTGGACAACATAGGGAGACCCTGTGT
17. tgc etc age etc ccg agt age t
18. cct cct gag ttc aag cga ttc tc
19. tgc ctt cac aca get ctg gtt taa tg 20. tgc ctt cac aca gca ctg gtt taa tg
21. ACCATGTTGGCCAGGCTGGTTTT
22. ATCTACTGACCTCAAATGATCCACCT
23. tgc aat ccg ccc gcc
24. cca ggc tgg ttt gga aat cct gag etc 25. ctg aga teg cac cac tgc ac 26. ggg agg egg age ttg cag tga
27. gcg cat gee tgt aat tct gta
28. cag gac gag cca cag aca aaa etc c
29. tgc aat gag get cct ggc c 30. act aca att tec cag cat ccc a
31. cct ccc tec etc cct gc
32. tgc ttg ctt tct tct tct
33. tec ctg ctt get tgx ttt etc t
34. tct etc ttt ctt tct ttc ttt c 35. tgt tea tec aaa tga gee gc
36. age ctg aac agg ttc tgt tec ttc gac tt
37. caa get get ate teg ace gat ctt
38. ggg tga cca ccc tgc cag cc
39. egg get aca ggg tta cct gag 40. tct gca ace tgg tgc gag cag c
41. tga ggc tec get cct tct gg
42. age tgc cag age tgc ctg ggc
60. A primer or probe for use in a method as defined in any of the claims above as the other probe said primer or probe being selected from
1. tgc etc ace cct gta ate c
2. get tgt aat ccc age tac teg
3. caa cac tea cac ccc aca g 4. aga tea cgc cac tgc act c
5. ttg aca att gag caa aga gc
6. ttg gat tac aga cgt gag c
7. agt gca gee tea act tec
8. cca gtc caa aca ata tga tec 9. cat gat tea ctg cac cca ace
10. ttt cac tct tgt tgc cca age
11. cac cat gee tgg etc caa tgt
12. age cca gga att caa gg
13. aga aca ttg gag cca gg 14. aga ate act gga ate cag g 15. ttt tea cac agg tec aat cc
16. act gca ace tec ate tec
17. caa tta agt gcc ttc aca cag ca
18. ggt caa gag ttc aag ace age 19. ccc tgc ccc ace tct cc
20. agt caa ttt ctg tgc aaa eta ctt tta ttt
21. gag gca aca gga aca aac c
22. cat tgg aat gag cag aaa cc
23. taa cat aaa gaa tea gga gga ggc 24. agt tgg etc ate tgc etc tt
25. tgg eta aca egg tga aac c
26. gga ate caa aga ttc tat gat gg
27. act cct gac ttc aaa tga tec
28. tag ccc cca gtc acg ttc c 29. aga agt cca aga gtt tgc age
30. ttc tea gtc cca gaa tga ace
31. cca ctt agg taa aca cet ctt
32. ctg caa tga gcc gag ata gaa
33. cca ctt agg taa aca cct ctt 34. ctg caa tga gcc gag ata gaa
35. atg ttg ggg aga ctg agg
36. ccg cat eta act tat tct gg
37. aac tac etc tgc aaa ccc age
38. ttg gaa tgg agg gat tct ace 39. ggt ttt ctg etc tgc aca eg
40. cct ttc tec ttc cac caa eg
41. gga cag atg gca atg atg g
42. tct tct tct tgg tgg atg tgg
61. The primer or probe according to any of claims 59 or 60 or as defined in any claims 22-34, wherein the probe is operably linked to at least one label, such as operably linked to two different labels.
62. The primer or probe according to claim 61 , wherein the label is selected from TEX1 TET, TAM, ROX, R6G, ORG, HEX, FLU, FAM, DABSYL, Cy7, Cy5, Cy3, BOFL, BOF, BO-X, BO-TRX, BO-TMR, JOE, 6JOE, VlC, 6FAM, LCRed640, LCRed705, TAMRA, Biotin, Digoxigenin, DuO-family, Daq-family.
63. The primer or probe according to claim 61 , wherein the primer or probe is oper- ably linked to a surface.
64. The primer or probe according to claim 63, wherein the surface is the surface of microbeads or a DNA chip.
65. An antibody directed to an epitope of a RAI gene product.
EP05771208A 2004-08-18 2005-08-17 Disease risk estimating method using sequence polymorphisms in a specific region of chromosome 19 Withdrawn EP1786926A2 (en)

Applications Claiming Priority (4)

Application Number Priority Date Filing Date Title
DKPA200401249 2004-08-18
DKPA200500274 2005-02-23
DKPA200500918 2005-06-22
PCT/DK2005/000529 WO2006018023A2 (en) 2004-08-18 2005-08-17 Disease risk estimating method using sequence polymorphisms in a specific region of chromosome 19

Publications (1)

Publication Number Publication Date
EP1786926A2 true EP1786926A2 (en) 2007-05-23

Family

ID=35457796

Family Applications (1)

Application Number Title Priority Date Filing Date
EP05771208A Withdrawn EP1786926A2 (en) 2004-08-18 2005-08-17 Disease risk estimating method using sequence polymorphisms in a specific region of chromosome 19

Country Status (2)

Country Link
EP (1) EP1786926A2 (en)
WO (1) WO2006018023A2 (en)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2006136170A2 (en) * 2005-06-22 2006-12-28 Aarhus Universitet Disease risk estimating method using sequence polymorphisms in a specific region of chromosome 19

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20060147915A1 (en) * 2002-06-27 2006-07-06 Aarhus Universitet Disease risk estimating method fusing sequence polymorphisms in a specific region of chromosome 19

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
See references of WO2006018023A2 *

Also Published As

Publication number Publication date
WO2006018023A2 (en) 2006-02-23
WO2006018023A3 (en) 2006-12-14

Similar Documents

Publication Publication Date Title
US10407738B2 (en) Markers for breast cancer
US20030108910A1 (en) STK15 (STK6) gene polymorphism and methods of determining cancer risk
US20100092959A1 (en) Single nucleotide polymorphisms as genetic markers for childhood leukemia
US20060147915A1 (en) Disease risk estimating method fusing sequence polymorphisms in a specific region of chromosome 19
WO2006136170A2 (en) Disease risk estimating method using sequence polymorphisms in a specific region of chromosome 19
EP1786926A2 (en) Disease risk estimating method using sequence polymorphisms in a specific region of chromosome 19
US20110104669A1 (en) Method to predict iris color
AU2013203435B2 (en) Markers for breast cancer
AU2012202265B2 (en) Markers for breast cancer
HK1172063A (en) Markers for breast cancer
HK1171477A (en) Markers for breast cancer
HK1172062A (en) Markers for breast cancer
HK1172062B (en) Markers for breast cancer
HK1121499B (en) Markers for breast cancer
HK1171477B (en) Markers for breast cancer

Legal Events

Date Code Title Description
PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

AK Designated contracting states

Kind code of ref document: A2

Designated state(s): AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IS IT LI LT LU LV MC NL PL PT RO SE SI SK TR

AX Request for extension of the european patent

Extension state: AL BA HR MK YU

17P Request for examination filed

Effective date: 20070614

RAX Requested extension states of the european patent have changed

Extension state: YU

Payment date: 20070614

Extension state: MK

Payment date: 20070614

Extension state: HR

Payment date: 20070614

Extension state: BA

Payment date: 20070614

Extension state: AL

Payment date: 20070614

RBV Designated contracting states (corrected)

Designated state(s): AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IS IT LI LT LU LV MC NL PL PT RO SE SI SK TR

17Q First examination report despatched

Effective date: 20081021

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWN

18D Application deemed to be withdrawn

Effective date: 20090303