WO2014134970A1 - New biomarker for type 2 diabetes - Google Patents

New biomarker for type 2 diabetes Download PDF

Info

Publication number
WO2014134970A1
WO2014134970A1 PCT/CN2014/000226 CN2014000226W WO2014134970A1 WO 2014134970 A1 WO2014134970 A1 WO 2014134970A1 CN 2014000226 W CN2014000226 W CN 2014000226W WO 2014134970 A1 WO2014134970 A1 WO 2014134970A1
Authority
WO
WIPO (PCT)
Prior art keywords
diabetes
sequence
type
subject
risk
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2014/000226
Other languages
French (fr)
Inventor
Ronald Ching-Wan MA
Wing-Yee So
Juliana Chung-Ngor CHAN
Cheng Hu
Rong Zhang
Weiping JIA
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
SHANGHAI JIAOTONG UNIVERSITY AFFILIATED SIXTH PEOPLE'S HOSPITAL
Hospital Authority
Chinese University of Hong Kong CUHK
Original Assignee
SHANGHAI JIAOTONG UNIVERSITY AFFILIATED SIXTH PEOPLE'S HOSPITAL
Hospital Authority
Chinese University of Hong Kong CUHK
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by SHANGHAI JIAOTONG UNIVERSITY AFFILIATED SIXTH PEOPLE'S HOSPITAL, Hospital Authority, Chinese University of Hong Kong CUHK filed Critical SHANGHAI JIAOTONG UNIVERSITY AFFILIATED SIXTH PEOPLE'S HOSPITAL
Publication of WO2014134970A1 publication Critical patent/WO2014134970A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • C12Q1/6876Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
    • C12Q1/6883Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q2600/00Oligonucleotides characterized by their use
    • C12Q2600/156Polymorphic or mutational markers

Definitions

  • Type 2 diabetes is a common complex disease characterised by deficient insulin secretion and decreased insulin sensitivity.
  • T2D Type 2 diabetes
  • 285 million people worldwide were affected by type 2 diabetes (Shaw et al, (2010) Diabetes Res Clin Pract 87: 4-14), with 60% of them located in Asia (Chan et al. (2009) JAMA 301 : 2129-2140; Ramachandran et al. (2010) Lancet 375: 408-418).
  • the present invention provides a method for assessing the presence or risk of type 2 diabetes (T2D) or cardiovascular disease in a subject.
  • the method includes these steps: (a) performing an assay that determines nucleotide sequence of at least a portion of the PAX4-SND1 genomic sequence that is present in a biological sample taken from the subject, and (b) comparing the sequence determined in step (a) with a standard sequence of the corresponding PAX4-SND1 genomic sequence, wherein a variation in the sequence determined in step (a) when compared with the standard sequence indicates the presence or risk of type 2 diabetes or cardiovascular disease in the subject.
  • the sample is a blood or saliva sample.
  • the subject is of Asian descent, such as a Chinese, Korean, Japanese, especially a Han Chinese.
  • the subject has a family history of type 2 diabetes but has not been diagnosed of type 2 diabetes, while in other cases, the subject has no family history and has not been diagnosed of type 2 diabetes.
  • the method is particular effective in detecting or assessing the risk of developing cardiovascular disease in a subject, who has already been diagnosed with type 2 diabetes. After performing steps (a) and (b) as described above, a sequence variation indicates the presence or risk of developing cardiovascular disease in the subject.
  • the assay in step (a) may comprise an amplification reaction, such as a polymerase chain reaction (PCR); or the assay in step (a) may comprise mass spectrometry.
  • amplification reaction such as a polymerase chain reaction (PCR)
  • mass spectrometry a polymerase chain reaction
  • sequence variants include a polymorphism rs7801 1 1 , rs806187, rsl 40971 , rs806179, rs806178, rs7781 189, or rs806176.
  • the method of this invention is not limited to be used with just one sequence variation. In some cases, especially after a sequence variation is detected in the portion of PAX4-SND1 genomic sequence after steps (a) and (b), a further step may be taken to detect a second sequence variation in a second genomic sequence, e.g. , one that is different from the PAX4-SND1 genomic sequence, or one is in another portion of the PAX4-SND1 genomic sequence.
  • Two or more such additional sequence variants can be used for this purpose.
  • the diagnosis of presence or risk of type 2 diabetes or cardiovascular disease is further supported.
  • One example of the second sequence variation is a polymorphism rs2737250, located near TRPS1. Additional examples can be found in Tables 2, 6, 7, and 8.
  • one or more treatment steps should be taken.
  • a physician may prescribe administering to the subject a cholesterol lowering drug or a blood glucose lowering drug.
  • the subject once indicated as at risk of developing type 2 diabetes or cardiovascular disease according to the methods described above, may receive one or more further steps of monitoring for any of these conditions on a regular basis, utilizing physical examination tools, laboratory tests and application of various scanning and/or scoping technologies to image high risk anatomical areas. Preventive steps may also be taken such as changing dietary habits, increasing physical activity level, etc.
  • the present invention provides a kit for assessing the presence or risk of T2D or cardiovascular disease in a subject.
  • the kit includes two oligonucleotide primers for specifically amplifying: (1) at least a segment of the PAX4-SND1 genomic sequence; or (2) complement of (1), in an amplification reaction.
  • Such an amplification reaction may be a polymerase chain reaction (PCR), such as RT-PCR.
  • the kit may include an agent that can differentially indicate a sequence variation within the genomic sequence following its amplification, e.g., an oligonucleotide probe that specifically binds to one version of the genomic sequence but not to other versions.
  • the kit typically further includes an instruction manual. BRIEF DESCRIPTION OF THE DRAWINGS
  • Figure 1 Summary of study design. CHB, Han Chinese in Beijing, China; JPT, Japanese in Tokyo, Japan.
  • Figure 2 Manhattan plot of combined genome-wide association results from the Hong Kong 1, Hong Kong 2 and Shanghai studies based on the random effect models.
  • the j-axis represents the -logio p value
  • the x-axis represents the 2,925,090 analysed SNPs.
  • the dotted line indicates the threshold of significance p ⁇ x 10 5 .
  • There are 44 points with p ⁇ x 10 ⁇ 5 and the arrows and labels localise the susceptibility loci to type 2 diabetes uncovered in the present study.
  • Figure 3 Regional plots for the identified variant rsl0229583, including results for both genotyped and imputed SNPs in the Chinese population.
  • the top positioned circle and diamond represent the sentinel SNP in meta-analysis of three GWAS in the stage 1 and the East Asian meta-analysis in stages 1+2+3, respectively.
  • Other SNPs are coloured according to their level of LD, which is measured by r 2 , with the sentinel SNP.
  • the recombination rates estimated from the 1000 Genomes project JPT+CHB data are shown. CHB, Han Chinese in Beijing, China; JPT, Japanese in Tokyo, Japan.
  • Figure 4 Forest plot for meta-analysis of the association between type 2 diabetes and rs 10229583 for all populations in the present study. ORs and 95% CIs were reported with respect to the type 2 diabetes-related risk alleles (G).
  • FIG. 6 Multidimensional scaling analysis (MDS) plot showing the first two principal components, based on genotype data of 1 1 populations from HapMap (African ancestry in Southwest USA (ASW), Utah residents with Northern and Western European ancestry from the CEPH collection (CEU), Han Chinese in Beijing, China (CHB), Chinese in Metropolitan Denver, Colorado (CHD), kanni Indians in Houston, Texas (GIH), Japanese in Tokyo, Japan (JPT), Luhya in Webuye, Kenya (LWK), Mexican ancestry in Los Angeles, California (MEX), Maasai in Kinyawa, Kenya (MKK), Samsung in Italy (TSI) and Yoruban in Ibadan, Nigeria (YRI)), as well as the 3 case-controls cohorts (Hong Kong GWAS 1 (HK1), Hong Kong GWAS 2 (HK2) and Shanghai GWAS (SH)) in the stage 1 genome scan of the present study.
  • FIG. 7 Multidimensional scaling analysis (MDS) plot shows the first two principal components, based on genotype data of 3 case-controls cohorts (Hong Kong GWAS 1 , Hong Kong GWAS 2 and Shanghai GWAS) in the stage 1 genome scan of the present study without HapMap scaling.
  • MDS Multidimensional scaling analysis
  • Figure 8 Q-Q plot for combined genome-wide association results in a total of 684 T2D patients and 955 controls based on the 2,925,090 analyzed SNPs.
  • the curvy lines above and below the straight diagonal lines represent the upper and lower boundaries of the 95% confidence bands.
  • Figure 9 Distribution of ENCODE open chromatin sites around the identified gene region (NCBI Build 36.1/hgl8 CHR7:127033000-127084685) annotated in the UCSC human genome browser on human (website: genome.ucsc.edu/).
  • the distribution of open chromatin across the identified gene region in pancreatic islets (depicted as peaks) is highlighted inside the box (labeled Panlsle FAIRE FD and highlighted by red arrow).
  • the position of rs 10229583 and other tagging SNPs in high LD to this lead SNP are marked by the black arrows at the top of the figure.
  • Figure 10 Comparison of varLD scores within 100Kb region centred on index SNP rsl0229583 between pairs of populations using HapMap phase III CHB, JPT, CEU and YRI data.
  • Figure 11 Linkage disequilibrium for SNPs within the region near PAX4 on chromosome 7 between 126.95 Mb and 127.06 Mb (Build 36). Pairwise r 2 among SNPs for HapMap CEU and CHB are indicated in upper and lower block, respectively. Shades of grey represent the strength of pairwise r 2 . Rsl0229583 and rs6467136 are the SNPs showing significant association with T2D in the present study and the study conducted by the East Asian Consortium, respectively. DEFINITIONS
  • Type 2 diabetes refers to a metabolic disorder that is characterized by high blood glucose in the context of varying combinations of insulin resistance and insulin deficiency.
  • Type 2 diabetes may be caused by a combination of lifestyle and genetic factors. Diabetes can be caused by distinct clinical entities such as endocrine disorders (e.g., Cushing's syndrome) and chronic pancreatitis.
  • type 2 diabetes often include polyuria (frequent urination), polydipsia (increased thirst), polyphagia (increased hunger), fatigue, and weight loss.
  • the abnormal neurohormonal and metabolic milieu characterized by hyperglycemia, dyslipidemia and low grade inflammation can trigger a cascade of signaling pathways, which can lead to cell death and dysregulated cell growth, giving rise to multiple morbidities including heart disease, strokes, limb amputation, visual loss, kidney failure, cancers, and cognitive impairment.
  • cardiovascular disease refers to a broad class of diseases that involve the heart or blood vessels (arteries and veins) and affect the cardiovascular system, such as conditions related to atherosclerosis (arterial disease). These include but not limited to stroke, coronary heart disease and peripheral vascular disease.
  • Known risk factors for cardiovascular diseases include unhealthy eating, lack of exercise, obesity, suboptimally managed diabetes, abnormal blood lipids, high blood pressure, excessive consumption of alcohol, use of tobacco, as well as genetic background.
  • diabetic cardiovascular disease specifically refers to a cardiovascular disease that is associated with or secondary to diabetes.
  • a BMI of 20 to 25 kg/m 2 is considered optimal weight; a BMI lower than 20 kg/m 2 suggests the person is underweight whereas a BMI above 25 kg/m 2 may indicate the person is overweight; a BMI above 30 kg/m suggests the person is obese; and a BMI over 40 kg/m indicates the person to be morbidly obese.
  • Asians have more body fat for the same degree of BMI and waist circumference.
  • Asians are defined as ⁇ 23 kg/m and >25 kg/m respectively. While high BMI may predict risk for diabetes or prediabetes, people with low BMI, which correlates with beta cell function, are also at high risk, especially if these subjects develop central obesity, which tends to be associated with insulin resistance or reduced insulin sensitivity.
  • biological sample includes any section of tissue or bodily fluid taken from a test subject such as a biopsy and autopsy sample, and frozen section taken for histologic purposes, or processed forms of any of such samples.
  • Biological samples include blood and blood fractions or products (e.g., serum, plasma, platelets, white blood cells, red blood cells, and the like), sputum or saliva, lymph and tongue tissue, cultured cells, e.g. , primary cultures, explants, and transformed cells, stool, urine, stomach biopsy tissue etc.
  • a biological sample is typically obtained from a eukaryotic organism, which may be a mammal, may be a primate and may be a human subject.
  • biopsy refers to the process of removing a tissue sample for diagnostic or prognostic evaluation, and to the tissue specimen itself. Any biopsy technique known in the art can be applied to the methods of the present invention. The biopsy technique applied will depend on the tissue type to be evaluated (e.g., tongue, colon, prostate, kidney, bladder, lymph node, liver, bone marrow, blood cell, stomach tissue, etc.) among other factors. Representative biopsy techniques include, but are not limited to, excisional biopsy, incisional biopsy, needle biopsy, surgical biopsy, and bone marrow biopsy and may comprise endoscopy such as colonoscopy. A wide range of biopsy techniques are well known to those skilled in the art who will choose between them and implement them with minimal experimentation.
  • isolated nucleic acid molecule means a nucleic acid molecule that is separated from other nucleic acid molecules that are usually associated with the isolated nucleic acid molecule.
  • an "isolated" nucleic acid molecule includes, without limitation, a nucleic acid molecule that is free of nucleotide sequences that naturally flank one or both ends of the nucleic acid in the genome of the organism from which the isolated nucleic acid is derived (e.g., a cDNA or genomic DNA fragment produced by a polymerase chain reaction or restriction endonuclease digestion).
  • an isolated nucleic acid molecule is generally introduced into a vector (e.g., a cloning vector or an expression vector) for convenience of manipulation or to generate a fusion nucleic acid molecule.
  • an isolated nucleic acid molecule can include an engineered nucleic acid molecule such as a recombinant or a synthetic nucleic acid molecule.
  • nucleic acid or “polynucleotide” refers to deoxyribonucleic acids (DNA) or ribonucleic acids (RNA) and polymers thereof in either single- or double-stranded form.
  • nucleic acids containing known analogs of natural nucleotides that have similar binding properties as the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides.
  • a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions), alleles, orthologs, single nucleotide
  • SNPs polymorphisms
  • complementary sequences as well as the sequence explicitly indicated.
  • degenerate codon substitutions may be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and/or deoxyinosine residues (Batzer et ah, Nucleic Acid Res. 19:5081 (1991);
  • nucleic acid is used interchangeably with gene, cDNA, and mRNA encoded by a gene.
  • gene means the segment of DNA involved in producing a polypeptide chain; it includes regions preceding and following the coding region (leader and trailer) involved in the transcription and/or translation of the gene product and the regulation of the transcription and/or translation, as well as intervening sequences (introns) between individual coding segments (exons).
  • polypeptide polypeptide
  • peptide protein
  • protein protein
  • amino acid polymers in which one or more amino acid residue is an artificial chemical mimetic of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers and non-naturally occurring amino acid polymers.
  • the terms encompass amino acid chains of any length, including full-length proteins (i.e., antigens), wherein the amino acid residues are linked by covalent peptide bonds.
  • amino acid refers to naturally occurring and synthetic amino acids, as well as amino acid analogs and amino acid mimetics that function in a manner similar to the naturally occurring amino acids.
  • Naturally occurring amino acids are those encoded by the genetic code, as well as those amino acids that are later modified, e.g., hydroxyproline, ⁇ -carboxyglutamate, and O-phosphoserine.
  • amino acid analogs refers to compounds that have the same basic chemical structure as a naturally occurring amino acid, i.e., a carbon that is bound to a hydrogen, a carboxyl group, an amino group, and an R group, e.g., homoserine, norleucine, methionine sulfoxide, methionine methyl sulfonium. Such analogs have modified R groups (e.g., norleucine) or modified peptide backbones, but retain the same basic chemical structure as a naturally occurring amino acid.
  • amino acid mimetics refer to chemical compounds that have a structure that is different from the general chemical structure of an amino acid, but that functions in a manner similar to a naturally occurring amino acid.
  • Amino acids may include those having non-naturally occurring D-chirality, as disclosed in WO01/12654, which may improve the stability (e.g., half-life), bioavailability, and other characteristics of a polypeptide comprising one or more of such D-amino acids. In some cases, one or more, and potentially all of the amino acids of a therapeutic polypeptide have D-chirality.
  • Amino acids may be referred to herein by either the commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical
  • immunoglobulin or "antibody” (used interchangeably herein) refers to an antigen-binding protein having a basic four-polypeptide chain structure consisting of two heavy and two light chains, said chains being stabilized, for example, by interchain disulfide bonds, which has the ability to specifically bind antigen. Both heavy and light chains are folded into domains.
  • antibody also refers to antigen- and epitope-binding fragments of antibodies, e.g., Fab fragments, that can be used in immunological affinity assays.
  • Fab fragments antigen- and epitope-binding fragments of antibodies
  • pepsin digests an antibody C-terminal to the disulfide linkages in the hinge region to produce F(ab)' 2 , a dimer of Fab which itself is a light chain joined to V R -C R I by a disulfide bond.
  • the F(ab)' 2 can be reduced under mild conditions to break the disulfide linkage in the hinge region thereby converting the (Fab') 2 dimer into an Fab' monomer.
  • the Fab' monomer is essentially a Fab with part of the hinge region (see, e.g., Fundamental Immunology, Paul, ed., Raven Press, N.Y. (1993), for a more detailed description of other antibody fragments). While various antibody fragments are defined in terms of the digestion of an intact antibody, one of skill will appreciate that fragments can be synthesized de novo either chemically or by utilizing recombinant DNA methodology. Thus, the term antibody also includes antibody fragments either produced by the modification of whole antibodies or synthesized using recombinant DNA methodologies.
  • the specified binding agent e.g., an antibody
  • Specific binding of an antibody under such conditions may require an antibody that is selected for its specificity for a particular protein or a protein but not its similar "sister" proteins.
  • immunoassay formats may be used to select antibodies specifically immunoreactive with a particular protein or in a particular form.
  • solid-phase ELISA immunoassays are routinely used to select antibodies specifically immunoreactive with a protein (see, e.g., Harlow & Lane, Antibodies, A Laboratory Manual (1988) for a description of immunoassay formats and conditions that can be used to determine specific immuno reactivity).
  • a specific or selective binding reaction will be at least twice background signal or noise and more typically more than 10 to 100 times background.
  • the term “specifically bind” when used in the context of referring to a polynucleotide sequence forming a double-stranded complex with another polynucleotide sequence describes "polynucleotide hybridization” based on the Watson-Crick base-pairing, as provided in the definition for the term “polynucleotide hybridization method.”
  • a "polynucleotide hybridization method" as used herein refers to a method for detecting the presence and/or quantity of a pre-determined polynucleotide sequence based on its ability to form Watson-Crick base-pairing, under appropriate hybridization conditions, with a polynucleotide probe of a known sequence. Examples of such hybridization methods include Southern blot, Northern blot, and in situ hybridization.
  • Primers refer to oligonucleotides that can be used in an amplification method, such as a polymerase chain reaction (PCR), to amplify a nucleotide sequence based on the polynucleotide sequence corresponding to a gene of interest, e.g., the cDNA or human genomic sequence PAX4-SND1 or a portion thereof.
  • PCR polymerase chain reaction
  • at least one of the PCR primers for amplification of a polynucleotide sequence is sequence-specific for that polynucleotide sequence. The exact length of the primer will depend upon many factors, including temperature, source of the primer, and the method used.
  • the oligonucleotide primer typically contains at least 10, or 15, or 20, or 25 or more nucleotides, although it may contain fewer nucleotides or more nucleotides.
  • the factors involved in determining the appropriate length of primer are readily known to one of ordinary skill in the art.
  • primer pair means a pair of primers that hybridize to opposite strands a target DNA molecule or to regions of the target DNA which flank a nucleotide sequence to be amplified.
  • the term "primer site” means the area of the target DNA or other nucleic acid to which a primer hybridizes.
  • a "label,” “detectable label,” or “detectable moiety” is a composition detectable by spectroscopic, photochemical, biochemical, immunochemical, chemical, or other physical means.
  • useful labels include 32 P, fluorescent dyes, electron-dense reagents, enzymes ⁇ e.g., as commonly used in an ELISA), biotin, digoxigenin, or haptens and proteins that can be made detectable, e.g., by incorporating a radioactive component into the peptide or used to detect antibodies specifically reactive with the peptide.
  • a detectable label is attached to a probe or a molecule with defined binding characteristics ⁇ e.g., a polypeptide with a known binding specificity or a polynucleotide), so as to allow the presence of the probe (and therefore its binding target) to be readily detectable.
  • binding characteristics e.g., a polypeptide with a known binding specificity or a polynucleotide
  • a "standard sequence” as used herein refers to the polynucleotide sequence of a predetermined genomic DNA segment, e.g., a defined portion or the entire length of a human genomic sequence of a given range and location, such as the human PAX4-SND1 genomic sequence, including 2 kb upstream and 2 kb downstream flanking sequences, that is present in a publically accessible database, e.g., the University of California Santa Cruz database (hgl 8), as the standard human genomic sequence for these particular genes.
  • a genomic DNA sequence determined from a test sample is compared with a "standard sequence,” the test sequence is aligned with the "standard sequence” at the corresponding nucleotide bases of the genomic sequence to reveal any sequence variation.
  • the standard genomic sequences for human PAX4 and SND1 genes are provided as below:
  • PAX4 (paired box 4, chr7: 127250346-127255982, hgl9) is the closest gene to the SNP rsl0229583 (chr7: 127246903, hgl9).
  • the PAX4 gene encodes 4 protein-coding isoforms (Tablel). Because the PAX4 gene is at the reverse (-) strand, rsl0229583 is about 3.4 kb downstream to the most 3 ' exon of the PAX4 gene.
  • PAX4-001 ENST00000341640. 2010 ENSP00000339906, 343 9 127250346 127255780
  • PAX4-002 ENST00000463946 ⁇ 613 ENSP00000451923 341 8 127250992 127255724
  • PAX4-004 ENST00000378740 I 08S ENSP00000368014 34S ⁇ 127250865 127255982
  • the term "amount” as used in this application refers to the quantity of a polynucleotide of interest or a polypeptide of interest present in a sample. Such quantity may be expressed in the absolute terms, i.e. , the total quantity of the polynucleotide or polypeptide in the sample, or in the relative terms, i.e. , the concentration of the polynucleotide or polypeptide in the sample.
  • an effective amount of a cholesterol lowering drug or a blood glucose lowering drug is the amount of said drug to achieve a decreased level of cholesterol or blood glucose, respectively, in a patient who has been given the drug for therapeutic purposes.
  • An amount adequate to accomplish this is defined as the "therapeutically effective dose.”
  • the dosing range varies with the nature of the therapeutic agent being administered and other factors such as the route of administration and the severity of a patient's condition.
  • subject or “subject in need of treatment,” as used herein, includes individuals who seek medical attention due to risk of, or actual suffering from type 2 diabetes or cardiovascular/renal disease associated with diabetes. Subjects also include individuals currently undergoing therapy that seek manipulation of the therapeutic regimen. Subjects or individuals in need of treatment include those that demonstrate symptoms of type 2 diabetes or related cardiovascular/renal disease, or are at risk of suffering from type 2 diabetes or diabetic cardiovascular/renal disease or related symptoms.
  • a subject in need of treatment includes individuals with a genetic predisposition or family history for type 2 diabetes or diabetic cardiovascular/renal disease, those who have suffered relevant symptoms in the past, those who have been exposed to a triggering substance or event, as well as those suffering from chronic or acute symptoms of the condition.
  • a "subject in need of treatment” may be at any age of life.
  • the present inventors performed studies to identify new type 2 diabetes susceptibility loci in Southern Han Chinese individuals. A meta-analysis was performed of three GWAS comprising 684 patients with type 2 diabetes and 955 controls, and analysed 2.9 million
  • SNPs single-nucleotide polymorphisms
  • Putatively associated SNPs (p ⁇ l > ⁇ 10 ⁇ 5 ) were genotyped de novo in two independent Southern Han Chinese cohorts (10,383 cases and 6,974 controls), and SNPs reaching a genome- wide significance of /? ⁇ 5x 10 " were replicated in silico in five East Asian and three non-East Asian populations for a total of 31,541 cases and 60,344 controls.
  • the inventors discovered for the first time the correlation between genomic sequence variation in the human PAX4-SND1 genomic sequence and medical conditions such as type 2 diabetes and diabetic cardiovascular and renal diseases in human subjects.
  • This discovery allows medical professionals to identify subjects at risk cardiovascular or renal disease in a patient with type 2 diabetes or assess the risk of developing type 2 diabetes and diabetic cardiovascular and/or renal disease in a subject at risk by studying the subject's PAX4-SND1 genomic sequence and then comparing the subject's sequence with a standard PAX4-SND1 genomic sequence that has been determined as a part of the standard human genome. Detection of such sequence variation(s) indicates the presence or elevated risk of developing type 2 diabetes or diabetic cardiovascular and/or renal disease in the subject, as well as the early onset of these conditions.
  • the detection of pertinent genomic sequence variation(s) can further guide physicians to devise or modify treatment plans for a subject in both prevention and therapeutic measures.
  • type 2 diabetes patients have an increased risk of coronary heart disease (CHD) when they possess the genomic sequence variation in the human PAX4-SND1 genomic sequence at 7q32.
  • CHD coronary heart disease
  • a recent genome -wide association study in the Chinese populations identified association between a novel variant at 7q32 near paired box 4 (PAX4) and T2D, which was confirmed in other East Asian populations. This study aimed to investigate the association of this novel 7q32 variant and CHD risk in an 8- year prospective cohort of Chinese patients with T2D.
  • nucleic acids sizes are given in either kilobases (kb) or base pairs (bp). These are estimates derived from agarose or acrylamide gel electrophoresis, from sequenced nucleic acids, or from published DNA sequences.
  • kb kilobases
  • proteins sizes are given in kilodaltons (kDa) or amino acid residue numbers. Protein sizes are estimated from gel electrophoresis, from sequenced proteins, from derived amino acid sequences, or from published protein sequences.
  • Oligonucleotides that are not commercially available can be chemically synthesized, e.g. , according to the solid phase phosphoramidite triester method first described by Beaucage and Caruthers, Tetrahedron Lett.
  • oligonucleotides are synthesized using any art-recognized strategy, e.g., native acrylamide gel electrophoresis or anion-exchange high performance liquid chromatography (HPLC) as described in Pearson and Reanier, J. Chrom. 255: 137-149 (1983).
  • HPLC high performance liquid chromatography
  • sequence of interest used in this invention e.g., the polynucleotide sequence of the human PAX4 and SND1 genes, and synthetic oligonucleotides (e.g., primers) can be verified using, e.g., the chain termination method for sequencing double-stranded templates of Wallace et al., Gene 16: 21-26 (1981).
  • sequence of interest used in this invention e.g., the polynucleotide sequence of the human PAX4 and SND1 genes, and synthetic oligonucleotides (e.g., primers) can be verified using, e.g., the chain termination method for sequencing double-stranded templates of Wallace et al., Gene 16: 21-26 (1981).
  • the present invention relates to determining at least a portion of the genomic sequence of a pertinent region, such as the human PAX4-SND1 segment and/or its transcript(s), found in a biological sample taken from a person being tested, as a means to detect the presence and/or to assess the risk of developing type 2 diabetes, or cardiovascular/renal diseases in that person.
  • the first steps of practicing this invention are to obtain a biological sample (e.g., tissue or bodily fluid sample) from a test subject and extract genomic DNA or RNA from the sample.
  • a biological sample is obtained from a person to be tested or assessed for risk of developing type 2 diabetes or associated cardiovascular or renal disease using a method of the present invention. Collection of a tissue or fluid sample from an individual is performed in accordance with the standard protocol laboratories, hospitals or clinics generally follow, such as during a biopsy, blood drawing, saliva collection, or oral swab. An appropriate amount of sample is collected and may be stored according to standard procedures prior to further preparation.
  • genomic DNA found in a subject's sample may be performed using essentially any tissue or bodily fluid, so long as genomic DNA is expected to be present in such sample.
  • the methods for preparing tissue or fluid samples for nucleic acid extraction are well known among those of skill in the art. For example, a subject's epithelial tissue sample should be first treated to disrupt cellular membrane so as to release nucleic acids contained within the cells.
  • Possible sequence variation within a segment of a pertinent genomic sequence is investigated to provide indication as to whether a test subject is suffering from type 2 diabetes and associated cardiovascular or renal disease, or whether the subject is at risk of developing type 2 diabetes and associated complications including cardiovascular or renal disease in the future.
  • a segment of the genomic sequence of an appropriate length is selected for sequencing analysis.
  • the segment may be chosen from the genomic sequence of a pertinent gene or genes defined by the same boundaries defining the gene's genomic sequence, plus about 2,000 base pairs upstream and downstream from the boundaries.
  • the human PAX4- SND1 genomic sequence will encompass the upstream boundary of the upstream gene, PAX4, to the downstream boundary of the downstream gene, SNDl, plus 2000 bp upstream and downstream from the upstream and downstream boundaries, respectively.
  • the PAX4-SND1 genomic sequence may be as long as: 2000 bp genomic sequence immediately upstream from the PAX4 genomic sequence + the PAX4 genomic sequence + the genomic sequence between the PAX4 and SNDl genomic sequences + the SNDl genomic sequence + 2000 bp genomic sequence immediately downstream from the SNDl genomic sequence.
  • the length of the genomic sequence being analyzed may be a segment of the above and is usually at least 15 or 20 contiguous nucleotides, and may be longer with at least 25, 30, 50, 100, 200, 300, 400, or more contiguous nucleotides.
  • RNA contamination should be eliminated to avoid interference with DNA analysis.
  • other components such as proteins and lipids may be removed from the biological sample prior to further analysis of the genomic DNA.
  • RNA RNA sequence-based analysis such that the genomic sequence of one or more of the pertinent genes, or one or more of its transcripts, found in a test subject may be determined and then compared with a standard sequence to detect any possible sequence variation.
  • An amplification reaction is optional prior to the sequence analysis.
  • a variety of polynucleotide amplification methods are well established and frequently used in research. For instance, the general methods of polymerase chain reaction (PCR) for polynucleotide sequence amplification are well known in the art and are thus not described in detail herein.
  • PCR reagents and protocols are also available from commercial vendors, such as Roche Molecular Systems.
  • PCR amplification is typically used in practicing the present invention, one of skill in the art will recognize that amplification of the relevant genomic sequence may be accomplished by any known method, such as the ligase chain reaction (LCR), transcription- mediated amplification, and self-sustained sequence replication or nucleic acid sequence-based amplification ( ASBA), each of which provides sufficient amplification.
  • LCR ligase chain reaction
  • ASBA nucleic acid sequence-based amplification
  • Additional means suitable for determining the polynucleotide sequence of a genomic DNA include but are not limited to mass spectrometry, primer extension, polynucleotide hybridization, real-time PCR, melting curve analysis, high resolution melting analysis, heteroduplex analysis, pyrosequencing, and electrophoresis.
  • genomic DNA sequence variations may also be detected by way of analyzing RNA sequences transcribed from the pertinent DNA sequences, which may include portion of the coding sequence or non-coding sequence of a genomic locus of interest (e.g., the PAX4-SND1 genomic sequence).
  • Methods for RNA extraction from a biological sample, sequence analysis of RNA or DNA molecules, optionally involving amplification techniques such as reverse transcription based amplification processes, e.g., RT-PCR, are well known in the art.
  • Suitable samples for RNA sequence analysis may include peripheral blood monocytes (PBMC) and specific tissue samples such as fat and muscles.
  • PBMC peripheral blood monocytes
  • specific tissue samples such as fat and muscles.
  • the standard genomic sequence(s) for one or more pertinent genes such as the human PAX4 and SND1 genes and their isoforms, will be chosen before the comparison with a test subject's genomic sequence of the corresponding gene at the corresponding location may be performed.
  • genomic sequence variation in one or more of the specific genes named above and the presence or heightened risk of developing type 2 diabetes or cardiovascular and renal diseases among subjects having such variation, especially those fitting certain profiles, such as those of Asian descent, in particular Han Chinese
  • the present inventors have provided a valuable tool for clinicians to determine, often in combination with other information and diagnostic or predictive or screening test results, how a subject having certain genomic sequence variation(s) should be monitored and/or treated for type 2 diabetes and diabetic cardiovascular or renal disease such that the symptoms of these conditions may be prevented, eliminated, ameliorated, reduced in severity and/or frequency, or delayed in their onset.
  • a physician may arrange for regular monitoring of various symptoms of type 2 diabetes or diabetic cardiovascular and renal diseases in a subject who has been deemed by the method of the present invention to have an elevated risk of developing type 2 diabetes.
  • the physician may also prescribe both pharmacological and non-pharmacological treatments such as lifestyle modification ⁇ e.g., reduce body weight by 5%, high fiber diet, walking for at least 150 minutes weekly) and medicines known to reduce risk of onset of diabetes (e.g., metformin, alpha glucosidase inhibitors, lipase inhibitors) to a subject who has been deemed by the method of the present invention to have an elevated risk of developing type 2 diabetes.
  • lifestyle modification e.g., reduce body weight by 5%, high fiber diet, walking for at least 150 minutes weekly
  • medicines known to reduce risk of onset of diabetes e.g., metformin, alpha glucosidase inhibitors, lipase inhibitors
  • the attending physician may prescribe medications to control risk factors such as high levels of blood cholesterol and triglycerol (e.g., statins and fibrates) and reduce angiotensin II activity (e.g., Angiotensin converting enzyme inhibitor (ACEI) and angiotensin II receptor blocker (ARB)), as well as place the subject under regular testing and monitoring of coronary artery condition and kidney function.
  • risk factors such as high levels of blood cholesterol and triglycerol (e.g., statins and fibrates) and reduce angiotensin II activity (e.g., Angiotensin converting enzyme inhibitor (ACEI) and angiotensin II receptor blocker (ARB)
  • ACEI Angiotensin converting enzyme inhibitor
  • ARB angiotensin II receptor blocker
  • the present invention provides compositions and kits for practicing the methods described herein to detect possible genomic sequence variation of certain gene(s) and the transcripts thereof in a subject, which can be used for various purposes such as detecting or diagnosing the presence of type 2 diabetes and diabetic cardiovascular or renal disease in a subject, determining the risk of developing type 2 diabetes and diabetic cardiovascular or renal disease in a subject, and guiding the treatment plan for these conditions in the subject.
  • Kits for carrying out assays for determining the nucleotide sequence of a relevant genomic sequence typically include at least one oligonucleotide useful for specific hybridization with a predetermined segment of a pertinent genomic sequence (e.g., human PAX4-SND1 genomic sequence).
  • this oligonucleotide is labeled with a detectable moiety.
  • the oligonucleotide specifically hybridizes with the standard sequence only but not with any of the variant sequences.
  • the oligonucleotide specifically hybridizes with one particular version of the variant sequence but not with other versions, nor with the standard sequence.
  • kits may include at least two oligonucleotide primers that can be used in the amplification of at least one segment of a pertient genomic sequence (such as the human PAX4-SND1 genomic sequence) or transcripts thereof by PCR.
  • at least one of the oligonucleotide primers is designed to anneal only to the standard sequence or only to a particular version of the variant sequences, for example, the G allele of rs 10229583.
  • kits of this invention may provide instruction manuals ⁇ e.g., internet- based decision support tools) to guide users in analyzing test samples and assessing the presence or future risk of type 2 diabetes and diabetic cardiovascular or renal disease in a test subject.
  • instruction manuals e.g., internet- based decision support tools
  • the present invention can also be embodied in a device or a system comprising one or more such devices, which is capable of carrying out all or some of the method steps described herein.
  • the device or system performs the following steps upon receiving a biological sample taken from a subject being tested for detecting type 2 diabetes or diabetic cardiovascular or renal disease, assessing the risk of developing type 2 diabetes or diabetic cardiovascular or renal disease, or guiding treatment of a subject having or at risk of developing any one of these conditions: (a) determining in the sample the nucleotide sequence of a pertinent genomic DNA segment or its transcript; (b) comparing the sequence determined from the sample with a corresponding standard sequence; and (c) providing an output indicating whether type 2 diabetes or diabetic cardiovascular/renal disease is present in the subject or whether the subject is at risk of developing type 2 diabetes or diabetic
  • the device or system of the invention performs the task of steps (b) and (c), after step (a) has been performed and the genomic sequence determined from (a) has been entered into the device.
  • the device or system is partially or fully automated.
  • GWAS Genome -wide association studies
  • the present inventors performed a meta-analysis of three GWAS comprising 684 patients with type 2 diabetes and 955 controls of Southern Han Chinese descent.
  • the inventors followed up the top signals in two independent Southern Han Chinese cohorts (totalling 10,383 cases and 6,974 controls) and performed in silico replication in multiple populations. They identified CDKN2A/B and four novel type 2 diabetes association signals with p ⁇ 1 x 10 "5 from the meta-analysis. Thirteen loci within these four loci were followed up in two independent Chinese cohorts, and rs 10229583 at 7q32 was found to be associated with type 2 diabetes in a combined analysis of 1 1,067 cases and 7,929 controls l0 "8; OR [95% CI] 1.18 [1.11 , 1.25]). In silico replication revealed consistent associations across multiethnic groups, including five East
  • the rs 10229583 risk variant was associated with elevated fasting plasma glucose, impaired beta cell function in controls, and an earlier age at diagnosis for the cases.
  • the novel variant lies within an islet-selective cluster of open regulatory elements. There was significant heterogeneity of effect between Han Chinese and individuals of European descent, Malaysians and Indians. rs 10229583 near PAX4 is identified as a novel locus for type 2 diabetes in Chinese and other populations and provides new insights into the pathogenesis of type 2 diabetes.
  • stage 1 Participants In the first-stage discovery cohort (stage 1), genome-wide scanning was performed in three different case-control samples: 198 Hong Kong Chinese individuals (99 patients with type 2 diabetes and 99 healthy controls) in Hong Kong GWAS 1 , 1 ,047 Hong Kong Chinese individuals (388 with type 2 diabetes and 659 controls) in Hong Kong GWAS 2 and 394 Shanghai Chinese (197 patients with type 2 diabetes and 197 normal controls) in the Shanghai GWAS.
  • Individuals included in the stage 2 replication included 5,366 with type 2 diabetes and 2,474 controls from Hong Kong, and 4,035 cases and 3,964 controls from Shanghai. 325 cases and 368 controls from 178 Hong Kong families were also included, as well as 657 cases and 168 controls from 248 Shanghai families.
  • Table 4 shows the quality control for the participants in stage 1.
  • SNPs were excluded from further analysis if: (1) /? ⁇ l x l0 ⁇ 4 for HWE; (2) minor allele frequency (MAF) was ⁇ 1%; (3) call rate was ⁇ 95%; in particular, SNPs with MAF>1 % but ⁇ 5% were excluded if their call rate was ⁇ 99%; or (4) the SNPs showed a significant difference in MAF (p ⁇ x 1(T 4 ) between the Hong Kong control cohorts with other conditions (450 with epilepsy cases, 1 10 with eczema and 99 non-hypertensive individuals). Only SNPs that passed the quality control criteria for both cases and controls were used for further analysis. Table 5 shows the quality control of the genotyping results in stage 1. Genotypes were imputed for autosomal SNPs according to the 1000 Genomes reference panel. See the ESM Methods for further details.
  • MassARRAY platform (Sequenom; San Diego, CA, USA). Family samples were genotyped using TaqMan SNP Genotyping Assays (Applied Biosystems, Foster City, CA, USA) or by direct sequencing.
  • GWAMA software website: well.ox.ac.uk/gwama/
  • Magi and Morris BMC Bioinformatics 2010, 1 1 :288 was used to calculate the combined estimates of the ORs (95% CIs) from multiple groups by weighting the natural log-transformed ORs of each study using the inverse of their variance under the random effect model (DerSimonian and Laird (1986) Control Clin Trials 7: 177-188).
  • the random effect model SNPs with some degree of heterogeneity between studies were excluded, which helped to attenuate the number of false-positive findings in this study.
  • ALR alternating logistic regressions
  • bioinformatics and czs-expression quantitative trait loci (eQTL) analysis was performed for functional implication of the identified SNP. See the ESM Methods for additional information on methods, including adjustment for genomic control and the gene network analysis.
  • Meta-analysis of patients with Chinese ancestry A summary of the study design and the clinical characteristics of the participants in all stages are shown in Fig. 1 and Table 1.
  • stage 1 684 patients with type 2 diabetes and 955 controls were genotyped. No population stratification were detected between case and control individuals in multidimensional scaling analysis for all GWAS (Fig. 7).
  • Meta-analysis was implemented to combine the individual association results for 2,925,090 imputed and genotyped SNPs (under additive genetic models) available in all three GWAS using the inverse-variance approach for random effect models.
  • stage 1 meta-analysis of three Chinese GWAS 44 SNPs within five loci were prioritised for follow-up (Fig. 2 and Table 8). No substantial change was observed in the stage 1 results after adjusting either for ⁇ . s (1.01— 1.04 in individual cohorts) or the first principle component in the meta-analysis, reflecting that the results were not likely to be due to population stratification (Fig. 8 and Table 9).
  • CDKN2A/B has previously been reported to be strongly associated with type 2 diabetes.
  • two SNPs in CDKN2A/B showing strong signals for type 2 diabetes in the present study were in high LD (r 2 ⁇ 0.8) with rs 1081 1661 , which is well-replicated in most populations.
  • 13 top and proxy SNPs among the remaining 42 SNPs in four regions were taked forward to stage 2, de novo replication, in two independent Chinese case-control cohorts (Table 6).
  • Genotypes were successfully obtained for 1 1 SNPs in Hong Kong replication 1 cohort with 5,366 cases and 2,474 controls, and Shanghai replication 1 cohort with 4,035 cases and 3,964 controls to proceed for subsequent analysis (Table 7).
  • rsl0229583 and rs2737250 located on chromosomes 7 and 8, respectively, gave p ⁇ 4.5x l0 " (threshold of significance after Bonferroni correction) with the same directions of association as the original signals (Table 2).
  • These two SNPs were genotyped in 1 ,518 additional samples from 426 families of Han Chinese descent (325 cases and 368 controls from 178 Hong Kong families, and 657 cases and 168 controls from 248 Shanghai families).
  • FPG fasting plasma glucose
  • variant and its tagging SNPs lie within an area near PAX4 and SND1 , which is enriched with DNase I hypersenstitive sites, histone H3 lysine modifications and CCCTC factor binding in human islets (Fig. 9) (Stitzel et al. (2010) Cell Metab 12: 443-455).
  • eQTL data were only available for PAX4 in adipose tissue, but not LCLs or skin, for which no expression data were available from MuTHER. There was a nominal association (p ⁇ 0.05) between the variant and expression of C7orf54 and ARF5 in LCLs, and C7orf68 in adipose tissue.
  • the r 2 between the GWAS SNP and the peak eQTL SNPs ranged between 0.56 and 1.
  • Type 2 diabetes in Asians is characterised by an earlier AAD, strong family history and evidence of impaired beta cell function (Chan et al. (2009) JAMA 301 : 2129-2140; and
  • the novel locus for type 2 diabetes identified, rsl 0229583, is located downstream of the ARF5 and PAX4 genes in 7q32, and upstream of SND1.
  • PAX4 a member of the paired box family of transcription factors, plays a critical role in pancreatic beta cell formation during fetal development (Bran et al. (2004) J Cell Biol 167: 1 123-1 135; Li et al. (2006) Leuk Res 30: 1547- 1553) and is therefore a very strong candidate for the implicated gene.
  • PAX4 is expressed in early pancreatic endocrine cells, but expression is later restricted to beta cells and it is not expressed in mature pancreas (Habener et al. (2005) Endocrinology 146: 1025-1034). In pancreatic endocrine cells, PAX4 represses ghrelin and glucagon expression, and can induce the expression of PDX1, a key transcription factor for islet development. Targeted disruption of PAX4 in mice was found to lead to reduced beta cell mass at birth (Wang et al. (2004) Dev Biol 266: 178-189).
  • a risk variant at HNF4a has been found to be associated with increased risk of type 2 diabetes, and carriers of the risk allele have impaired beta cell function (Silander et al. (2004) Diabetes 53: 1 141-1 149).
  • the MAF of the R121W PAX4 mutation was 1 % in Asians, and the mutation is in low LD with rs 10229583. It is possible that both rare mutations and common variation within the same gene confer risk towards type 2 diabetes independently.
  • the common variant here identified, rsl0229583 may be associated with altered gene expression, while the other rare non-synonymous mutations lead to impaired gene function.
  • the recent East Asian meta-analysis comprising eight type 2 diabetes GWAS identified a locus on chromosome 7 near GRIP and GCC1-PAX4 to be associated with type 2 diabetes.
  • the protein encoded by GCC1 may play a role in transmembrane transport (Luke et al. (2005) Biochem J 388: 835-841).
  • the variant identified from the East Asian study, rs6467136, appears to be independent of our signal, with r 2 0.044 in our Chinese samples (Fig. 1 1).
  • ARF5 belongs to a family of guanine nucleotide -binding proteins that have been shown to play a role in vesicular trafficking and as activators of phospholipase D (Lebeda and Haun (1999) Gene 237: 209-214). Islet expression of ARF5 was found to be induced threefold in rats receiving a high-carbohydrate diet (Song et al. (2001) Diabetes 50: 2053-2060). The nearby SND1 gene, also known as the plOO transcription co- activator, is a member of the micronuclease family and plays a key role in transcription and splicing. The pi 00 transcriptional co-activator is present in endocrine cells and tissues, including the pancreas of cattle (Broadhurst et al. (2005) Biochim Biophys Acta 1681 : 126-133).
  • PAX4 mutations were first identified in Asian MODY probands (Shimajiri et al. (2001) Diabetes 50: 2864-2869; Tokuyama et al. (2006) Metabolism 55: 213- 216), but seldom found in those of European descent (Dupont et al. (1999) Diabetologia 42: 480- 484; Dusatkova et al. (2010) Diabet Med 27: 1459-1460). This suggests that PAX4, like KCNQl, may be particularly relevant for the pathogenesis of type 2 diabetes in East Asians individuals.
  • rs 10229583 is also in strong LD with a region spanning the neighbouring SND1 gene (Fig. 3). Further resequencing and transethnic mapping should help to identify the causal gene variant for type 2 diabetes within this region.
  • rs 10229583 near PAX4 has been identified as a novel locus for type 2 diabetes in Chinese and other populations, providing new insights into the pathogenesis of type 2 diabetes.
  • Example 2 Sequence Variation Near PAX4 Linked to Early Onset of Coronary Heart Disease Among Type 2 Diabetes Patients
  • Type 2 diabetes (T2D) patients have a 2-4 fold increased risk of coronary heart disease (CHD) compared with the general population.
  • CHD coronary heart disease
  • a recent genome-wide association study conducted by the present inventors in the Chinese populations identified association between a sequence variant at 7q32 near paired box 4 (PAX4) and T2D, which was confirmed in other East Asian populations. See details in Example 1.
  • Type 2 diabetes was diagnosed according to the 1998 World Health Organization (WHO) criteria. Patients with classic type 1 diabetes with acute ketotic presentation or continuous requirement of insulin within 1 year of diagnosis were excluded. Written informed consent was obtained from all participants. This study was approved by the Clinical Research Ethics Committee of the Chinese University of Hong Kong.
  • WHO World Health Organization
  • stage 1 genome -wide scanning was performed in 202 Hong Kong Chinese individuals (102 type 2 diabetes patients and 100 healthy controls) (Hong Kong GWAS 1 cohort).
  • 102 type 2 diabetes cases were selected with young-onset diabetes diagnosed at age ⁇ 40 years, positive family history and overweight, and 100 controls were selected using the criteria of 1) no past diagnostic history of type 2 diabetes, impaired fasting glucose (IFG) or impaired glucose tolerance (IGT); 2) without family history of type 2 diabetes; and 3) with BMI ⁇ 25 kg/m 2 and waist circumference ⁇ 90 cm and 80 cm for men and women, respectively.
  • IGF impaired fasting glucose
  • ITT impaired glucose tolerance
  • Hong Kong Chinese individuals were genome-scanned (400 type 2 diabetes patients from the Hong Kong Diabetes Registry and 668 non-diabetic controls) (Hong Kong GWAS 2 cohort).
  • the 668 diseased controls were individuals aged > 16 years old with diseases other than type 2 diabetes that included 457 epilepsy cases, 11 1 eczema cases and 100 healthy individuals without hypertension (recruited from the control arm of a hypertension study).
  • the genome -wide scan was performed in 394 samples, including 197 type 2 diabetes patients and 197 normal glucose regulation controls.
  • the type 2 diabetes patients were probands of diabetic pedigrees with fasting plasma glucose > 7.0 mmol/L and/or 2-h post plasma glucose > 1 1.1 mmol/L who were diagnosed before 40 years old.
  • Type 1 diabetes and mitochondrial diabetes were excluded based on clinical, immunological and genetic criteria.
  • the controls were individuals with normal glucose regulation with fasting plasma glucose ⁇ 6.1 mmol/L and 2-h plasma glucose ⁇ 7.8 mmol/L as assessed by standard 75g OGTTs, negative diabetic family history, aged over 50 years old and with a BMI below 23kg/m 2 .
  • HOMA-IR homeostasis model assessment of beta-cell function
  • ⁇ - ⁇ homeostasis model assessment of beta-cell function
  • Stumvoll indices for beta-cell function were calculated for Shanghai controls which underwent OGTT with measurement of insulin levels (Stumvoll et al. (2000) Diabetes Care 23: 295-301).
  • the case cohort consisted of 5,366 unrelated type 2 diabetes patients (mean age 56.7 ⁇ 13.4 years, 45.1% male, mean duration of T2D 6.6 ⁇ 6.9 years) selected from the Hong Kong Diabetes Registry (HKDR).
  • the control cohort consisted of 2474 individuals ascertained from 3 sources: a) 985 adolescents from a community-based school survey of cardiovascular risk factors (mean age 15.5 ⁇ 1.9 years, 44.2% male) (Ng et al. (2010) J Clin Endocrinol Metab 95: 2418- 2425); b) 513 hospital staff and adult volunteers participating in a community-based health screening program (mean age 42.0 ⁇ 10.4 years, 47% male) (Ng et al.
  • SHDS I and II Shanghai Diabetes Study (SHDS) I and II, which are community-based surveys of diabetes performed in 1998-2001 (SHDS I) and 2007-2008 (SHDS II).
  • the controls had fasting plasma glucose ⁇ 6.1 mmol/L and 2-h plasma glucose ⁇ 7.8 mmol/L as assessed by standard 75g OGTTs, and had no family history of diabetes mellitus.
  • Type 2 diabetes cases were selected from individuals registered as having type 2 diabetes. Diabetes was originally diagnosed according to the World Health Organization (WHO) criteria, type 2 diabetes was clinically defined as disease with a gradual adult onset. Individuals who tested positive for antibodies to glutamic acid decarboxylase (GAD) and those diagnosed with a mitochondrial disease or MODY were not included in the case group.
  • WHO World Health Organization
  • Controls were individuals registered as individuals not having type 2 diabetes but with diseases other than type 2 diabetes, comprised of 13 distinct diseases, or healthy volunteers. Individuals who had been analyzed in the previous report (Yamauchi et al., 2010 Nat Genet 42(10):864-868) were excluded from the present study. Altogether, 4,878 individuals with type 2 diabetes (case 1 , age, 65.8 ⁇ 10.0 years; BMI, 24.1 ⁇ 3.8 kg/m2; (all values are expressed as mean ⁇ s.d.)) and
  • 3,345controls (control 1 , age, 52.5 ⁇ 15.2 years; BMI, 22.5 ⁇ 3.8 kg/m2; (all values are expressed as mean ⁇ s.d.)) were genotyped. A total of 7,541 individuals belonging to the Hondo cluster (4,470 cases and 3,071 controls) were selected. Samples were directly genotyped using Illurnina HurnanHap610-Quad (type 2 diabetes patients) and 550K BeadChip (controls).
  • nondiabetic control individuals were as follows: (1) no history of diabetes and (2) fasting plasma glucose ⁇ 5.6 mmol/L and plasma glucose 2-h after ingestion of 75gm oral glucose load ⁇ 7.8 mmol/L at both baseline and follow up studies.
  • the Singapore case-control study contained individuals from three sources: 1) 1998 Singapore National Health Survey (NHS98); 2) Singapore Malay Eye Study (SiMES); and 3) Singapore Diabetes Cohorts Study (SDCS) (Tan et ah (2010) J Clin Endocrinol Metab 95: 390- 397).
  • NHS98 United States National Health Survey
  • SiMES Singapore Malay Eye Study
  • SDCS Singapore Diabetes Cohorts Study
  • FPG fasting plasma glucose
  • 2HPG 2 hour post-challenge glucose
  • Type 2 diabetic cases were identified as those with previously diagnosed type 2 diabetes and current use of antidiabetic treatment or who meet the following criteria: 1) 30 ⁇ age ⁇ 70, 2) fasting plasma glucose >7.0 mmol/1, 3) 2-h postprandial plasma glucose >1 1.1 mmol/1 in a standard 75 g oral glucose tolerance test (OGTT) or plasma HbAlc >6.5%.
  • the nondiabetic controls were selected according to the following criteria: 1) age >30, 2) no past history of diagnosis of diabetes and no family history of diabetes, 3) fasting glucose ⁇ 5.6 mmol/1, 4) 2-h OGTT ⁇ 7.8 mmol/1 and/or HbAl c content ⁇ 5.6%.
  • the studies were approved by local ethnic committees of each participating institution, and written informed consents were obtained from all participants.
  • the DNA samples were genotyped using the Illumina Human660W-Quad BeadChip (Illumina, Inc., San Diego, CA, USA), and the genotypes were called using the Illumina GenCall algorithm. Some of samples were excluded if their genotype call rates ⁇ 97%, excessive heterozygosity, gender mismatches between the reported and genetically inferred gender or duplicates among other samples. Principle component analysis was used to assess population structure of the samples and detected outliers along the first two eigenvectors which were excluded from further analyses. SNPs with genotype call rate ⁇ 95%, MAF ⁇ 0.5% or deviation from Hardy- Weinberg equilibrium (p ⁇ 10 "6 ) in control groups were also excluded.
  • DIAGRAM+ study comprised 8,130 type 2 diabetes cases and 38,987 controls from eight type 2 diabetes GWAS of European descent, including the Wellcome Trust Case Control Consortium (WTCCC), Diabetes Genetics Initiative (DGI) and Finland-US Investigation of NIDDM genetics (FUSION) scans (the individuals of a previous joint analysis), with those from scans performed by deCODE genetics, the Diabetes Gene Discovery Group, the Cooperative Health Research in the Region of Augsburg group (KORAgen), the Rotterdam study and the European Special Population Research Network (EUROSPAN) (for details of sample characteristics, see Supplementary Table 2 in Voight et ah (2010) Nat Genet 42: 579-589). Additional information on methods
  • the SNP ID (rs number) was standardized according to dbSNP build 129, and their physical positions were standardized according to build 36. SNPs were further excluded sequentially if: 1) their polymorphisms were A T or C/G; 2) absent from dbSNP build 129; 3) genotyped in only case or only control cohorts; 4) absent from the 1000 Genomes reference panel for CHB+JPT (March 2010 release of pilot project 1). For each sample set in stage 1 , all SNPs were aligned to the positive strand and imputed (via the MLE approach) using the MACH 1.0 software (Li et al. (2010) Genet Epidemiol 34: 816-834). Genotypes were imputed for autosomal SNPs that were present in the March 2010 release of phased 1000
  • Genomes genotype data from 60 CHB+JPT founders (Nature 467: 1061-1073, 2010), but were not present in the genome -wide chip or did not pass direct genotyping QC. Cases and controls were merged into a single cohort for imputation based on 440,194, 435,953 and 274,752 quality autosomal SNPs in Hong Kong GWAS 1, Hong Kong GWAS 2 and Shanghai GWAS case- control cohorts, respectively.
  • For the Hong Kong GWAS 1 cohort one-step imputation was applied.
  • two-step imputation was used to improve imputation efficiency, by randomly selecting 100 cases and 100 controls for model parameter estimation first before imputation.
  • Genomic control was applied to correct for relatedness of the individuals and adjust for potential population stratification (Devlin and Roeder (1999) Biometrics 55: 997- 1004).
  • the inflation factor ⁇ was estimated by taking the median of the distribution of the ⁇ 2 statistic from all quality SNPs in association test, and then divide by the median of the expected ⁇ 2 distribution.
  • the p values were calculated corrected for genomic control by dividing the observed ⁇ statistic by ⁇ . In this study, they were adjusted for GC in two levels. Firstly, each individual study was corrected for ⁇ separately in directly genotyped and imputed SNPs. Then they were further adjusted for GC on the meta-analysis results. eQTL analysis
  • MuTHER resource website: muther.ac.uk
  • Log 2 transformed expression signals were normalized separately per tissue as follows: quantile normalization was performed across technical replicates of each individual followed by quantile normalization across all individuals.
  • Genotyping was done with a combination of Illumina arrays (HumanHap300, HumanHap610Q, lMDuo and 1.2MDuo). Untyped HapMap2 SNPs were imputed using the IMPUTE software package (v2). The number of adipose samples with genotypes and expression values is per tissue was 778 for LCLs, 667 in skin and 776 in adipose.
  • the eQTL data were also examined for SNPs which showed nominal association (p ⁇ 0.05) with the expression of nearby genes in the different tissues from the MuTHER dataset.
  • the LD between the MuTHER eQTL peaks within the dataset and rsl 0229583 were then examined.
  • GenCord project is the study of association analysis for eQTL to nearby SNPs in three cell types (primary fibroblasts, lymphoblastoid cells and T cells) from the umbilical cords of 75 individuals (Dimas et al. (2009) Science 325: 1246- 1250).
  • the Java based application is the study of association analysis for eQTL to nearby SNPs in three cell types (primary fibroblasts, lymphoblastoid cells and T cells) from the umbilical cords of 75 individuals (Dimas et al. (2009) Science 325: 1246- 1250).
  • Genevar website: sanger.ac.uk/resources/software/genevar/
  • Genevar was used to retrieve the relevant association results of our variant and expression of genes within 1MB of the SNP from the GenCord project.
  • the LD between the variant and the eQTL peak within the dataset was also examined.
  • SNAP website: broadinstitute.org/mpg/snap/ldsearch.php. Selection criteria were as follows: (1) 1000 Genome Pilot 1 ; (2) r 2 limit > 0.8; (3) Population Panel: CHBJPT; and (4) Distance
  • GeneMania (Warde-Farley et ah (2010) Nucleic Acids Res 38: W214-220). Data was updated by GeneMania as of Feb 2012. The weight on each edge, as represented by the thickness of the edge, was computed by GeneMania and reflects the degree of confidence of the relationships between any gene pair within the network. Genes that are identified from the MuTHER eQTL analysis above were entered as prior knowledge, and were used to guide the gene-network building. Pathway information was obtained from Pathway Commons (Cerami et al. (201 1) Nucleic Acids Res 39: D685-690) and genetic interactions were obtained from BioGrid (Stark et al. (201 1) Nucleic Acids Res 39: D698-704).
  • T2D patient 99 (40.4) 40.6 ⁇ 8.8 31.8 ⁇ 7.7 8.0 ⁇ 8.3 30.9 ⁇ 4.4 -
  • HK2 Diseased control 659 (48.7) 37.1 ⁇ 17.0 - - 23.3 ⁇ 3.7 -
  • T2D patient 388 (49.5) 60.6 ⁇ 10.8 51.1 ⁇ 12.1 9.5 ⁇ 7.0 25.0 ⁇ 3.8 -
  • T2D patient 197 (57.9) 41.6 ⁇ 10.4 34.5 ⁇ 4.8 7.3 ⁇ 8.5 23.8 ⁇ 4.1 -
  • T2D patient 4035 (52.0) 61.2 ⁇ 12.1 54.2 ⁇ 11.3 7.2 ⁇ 6.9 24.5 ⁇ 3.5 -
  • T2D patient 325 (40.6) 48.0 ⁇ 14.4 41.7 ⁇ 13.1 6.3 ⁇ 7.6 25.9 ⁇ 4.4 -
  • T2D patient 657 (43.7) 54.6 ⁇ 15.6 50.0 ⁇ 14.2 4.9 ⁇ 7.3 23.9 ⁇ 3.5 -
  • T2D patient 4465 (68.0) 65.8 ⁇ 10.0 56.5 ⁇ 11.4 9.4 ⁇ 8.4 24.1 ⁇ 3.8
  • T2D patient 1042 (51.7) 56.4 ⁇ 8.6 25.5 ⁇ 3.3 7.0 ⁇ 2.6 Korean 2 Control 1305 (54.5) 65.2 ⁇ 2.6 23.9 ⁇ 3.0 5.0 ⁇ 0.5
  • T2D patient 1082 (37.2) 65.1 ⁇ 9.7 55.7 ⁇ 12.0 - 25.3 ⁇ 3.9 -
  • T2D patient 928 (64.9) 63.7 ⁇ 10.8 52.2 ⁇ 14.4 25.4 ⁇ 3.8
  • T2D patient 794 (51.0) 62.3 ⁇ 9.90 54.4 ⁇ 11.2 - 27.8 ⁇ 4.9 -
  • T2D type 2 diabetes
  • P, Pmeta and p het represent p values from logistic regression without any adjustment under the additive genetic model, meta-analysis under a fixed effect model
  • Asians Korean replication None 1,042 2,943 0.894 0.878 1.17 (0.99, 1.38) 0.0577
  • Phet refers to the p value obtained from the heterogeneity test
  • Step 1 Monomorphic SNPs 58,440 57,680 0 53,422 194 51 ,746 25,397 24,237
  • Step 2 SNPs with 100% missing genotype 78 78 0 112 0 0 2 0
  • Step 3 SNPs with MAF ⁇ 5% and their call rate
  • Step 4 SNPs with MAF > 5% and their call rate
  • Step 5 SNPs with overall MAF ⁇ 1% 9,656 9,967 81 26,447 1 ,311 27,618 236 214
  • Step 6 SNPs without HWE (P ⁇ 1 10 ⁇ 4 ) 463 458 54 598 19 575 12,281 12,486
  • Step 7 SNPs with different MAF between
  • HKl 99 T2D vs 99 controls
  • HK2 388 T2D vs 659 controls
  • SHGWA 197 T2D vs 197 controls
  • HK1 99 T2D vs 99 controls
  • HK2 388 T2D vs 659 controls
  • SHGWA 197 T2D vs 197 controls
  • Non- gene(s) risk (95% CI) (95% CI) (95% CI) (95% CI) (uncorrect)
  • P unadjusted and P adjusted represent P values calculated from linear regression with and without adjustmen for sex, BMI and Hb a i c (where appropriate) under the additive genetic model.
  • P values were obtained from met analysis of two Chinese T2D case cohorts (Hong Kong and Shanghai) under a fixed effect model.
  • Q test P and I 2 refer to the statistical significance and quantified index of heterogeneity test of OR between Chinese and other populations, respectively.
  • Monte Carlo and ⁇ test P refer to the P values testing for the regional LD variation and allelic difference between CHB and other populations, respectively.
  • CHB Han Chinese in Beijing, China
  • JPT Japanese in Tokyo, Japan
  • CEU Utah residents with Northern and Western European ancestry from the CEPH collection
  • YRI Yoruban in Ibadan, Nigeria.

Landscapes

  • Chemical & Material Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Health & Medical Sciences (AREA)
  • Organic Chemistry (AREA)
  • Wood Science & Technology (AREA)
  • Analytical Chemistry (AREA)
  • Zoology (AREA)
  • Genetics & Genomics (AREA)
  • Engineering & Computer Science (AREA)
  • Pathology (AREA)
  • Immunology (AREA)
  • Microbiology (AREA)
  • Molecular Biology (AREA)
  • Biotechnology (AREA)
  • Biophysics (AREA)
  • Physics & Mathematics (AREA)
  • Biochemistry (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Health & Medical Sciences (AREA)
  • Medicines That Contain Protein Lipid Enzymes And Other Medicines (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Abstract

The present invention provides a method for assessing the presence and risk of developing type 2 diabetes and cardiovascular disease in a subject by detecting sequence variation in the genomic sequence located between the PAX4 and SND1 genes in 7q32. A kit and device useful for such a method are also provided. In addition, the present invention provides a method for treating type 2 diabetes in patients who have been tested and shown to have the pertinent genetic variation(s).

Description

NEW BIOMARKER FOR TYPE 2 DIABETES
RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Patent Application No. 61/775,103, filed March 8, 2013, the contents of which are hereby incorporated by reference in the entirety. BACKGROUND OF THE INVENTION
[0002] Type 2 diabetes (T2D) is a common complex disease characterised by deficient insulin secretion and decreased insulin sensitivity. In 2010, 285 million people worldwide were affected by type 2 diabetes (Shaw et al, (2010) Diabetes Res Clin Pract 87: 4-14), with 60% of them located in Asia (Chan et al. (2009) JAMA 301 : 2129-2140; Ramachandran et al. (2010) Lancet 375: 408-418). China now has the largest number of patients with diabetes in the world, with an estimated 92 million affected individuals, and an additional 150 million with impaired glucose tolerance (Yang et al. (2010) N Engl J Med 362: 1090-1 101).
[0003] To identify common type 2 diabetes susceptibility variants, large-scale genome -wide association studies (GWAS) have been conducted in white individuals, yielding more than 60 genetic loci to date (Zeggini et al. (2008) Nat Genet 40: 638-645; Voight et al. (2010) Nat Genet 42: 579-589). Although many of these regions have been successfully replicated in Asian populations (Hu et al. (2009) Diabetologia 52: 1322-1325; Hu et al. (2010) PLoS One 5: el 5542; Hu et al. (2009) PLoS One 4: e7643; Ng et al. (2008) Diabetes 57: 2226-2233; Tarn et al. (2010) PLoS One 5: el 1428), discrepancies in allelic frequencies and effect sizes have demonstrated that interethnic differences exist. GWAS conducted in Japanese individuals (Yasuda et al. (2008) Nat Genet 40: 1092-1097; Unoki et al. (2008) Nat Genet 40: 1098-1 102), as well as metaanalyses of GWAS in South Asian (Kooner et al. (201 1) Nat Genet 43: 984-989) and East Asian (Cho et al. (201 1) Nat Genet 44:67-72) groups, have revealed additional variants not detected in GWAS with white individuals, with several signals, including KCNQ1, later replicated in many populations. Previous GWAS in Chinese suggested several loci but lacked large-scale replication (Cui et al. (2010) PLoS One 6: e22353; Tsai et al. (2010) PLoS Genet 6: el 000847; Shu et al. (2010) PLoS Genet 6: el OOl 127; Cho et al. (2009) Nat Genet 41 : 527-534).
[0004] Because of the enormous social and economic impact type 2 diabetes and associated complications such as cardiovascular and renal diseases, there exist clear and immediate needs to develop new and effective means for accurate diagnosis or early assessment a patient's risk of developing these diseases in the future, such that early intervention may be performed to minimize the harmful effects associated with these diseases and/or the risk of developing the diseases. The present invention fulfills this and other related needs. BRIEF SUMMARY OF THE INVENTION
[0005] In a first aspect, the present invention provides a method for assessing the presence or risk of type 2 diabetes (T2D) or cardiovascular disease in a subject. The method includes these steps: (a) performing an assay that determines nucleotide sequence of at least a portion of the PAX4-SND1 genomic sequence that is present in a biological sample taken from the subject, and (b) comparing the sequence determined in step (a) with a standard sequence of the corresponding PAX4-SND1 genomic sequence, wherein a variation in the sequence determined in step (a) when compared with the standard sequence indicates the presence or risk of type 2 diabetes or cardiovascular disease in the subject.
[0006] In some cases, the sample is a blood or saliva sample. In some cases, the subject is of Asian descent, such as a Chinese, Korean, Japanese, especially a Han Chinese. In some cases, the subject has a family history of type 2 diabetes but has not been diagnosed of type 2 diabetes, while in other cases, the subject has no family history and has not been diagnosed of type 2 diabetes.
[0007] In one particular use, the method is particular effective in detecting or assessing the risk of developing cardiovascular disease in a subject, who has already been diagnosed with type 2 diabetes. After performing steps (a) and (b) as described above, a sequence variation indicates the presence or risk of developing cardiovascular disease in the subject.
[0008] In any embodiment of the above described method, the assay in step (a) may comprise an amplification reaction, such as a polymerase chain reaction (PCR); or the assay in step (a) may comprise mass spectrometry.
[0009] In the application of the method described above, when the subject is indicated as having or at risk of developing type 2 diabetes, or having or at risk of developing cardiovascular disease, further therapeutic or prophylactic measure may be taken, such as taking the step of administering to the subject a cholesterol lowering drug or a blood glucose lowering drug, as deemed appropriate by the attending physician. [0010] A variety of genomic sequence variations can be used in practicing the method of this invention. One example is a polymorphism of rs 10229583, for instance, the G allele of polymorphism rsl0229583. Other sequence variants include a polymorphism rs7801 1 1 , rs806187, rsl 40971 , rs806179, rs806178, rs7781 189, or rs806176. [0011] The method of this invention is not limited to be used with just one sequence variation. In some cases, especially after a sequence variation is detected in the portion of PAX4-SND1 genomic sequence after steps (a) and (b), a further step may be taken to detect a second sequence variation in a second genomic sequence, e.g. , one that is different from the PAX4-SND1 genomic sequence, or one is in another portion of the PAX4-SND1 genomic sequence. Two or more such additional sequence variants can be used for this purpose. When at least one additional variant is detected the diagnosis of presence or risk of type 2 diabetes or cardiovascular disease is further supported. One example of the second sequence variation is a polymorphism rs2737250, located near TRPS1. Additional examples can be found in Tables 2, 6, 7, and 8.
[0012] In some cases, after the subject is indicated as having developed or at increased risk of type 2 diabetes or cardiovascular disease, one or more treatment steps should be taken. For example, a physician may prescribe administering to the subject a cholesterol lowering drug or a blood glucose lowering drug. On the other hand, the subject, once indicated as at risk of developing type 2 diabetes or cardiovascular disease according to the methods described above, may receive one or more further steps of monitoring for any of these conditions on a regular basis, utilizing physical examination tools, laboratory tests and application of various scanning and/or scoping technologies to image high risk anatomical areas. Preventive steps may also be taken such as changing dietary habits, increasing physical activity level, etc.
[0013] In a second aspect, the present invention provides a kit for assessing the presence or risk of T2D or cardiovascular disease in a subject. The kit includes two oligonucleotide primers for specifically amplifying: (1) at least a segment of the PAX4-SND1 genomic sequence; or (2) complement of (1), in an amplification reaction. Such an amplification reaction may be a polymerase chain reaction (PCR), such as RT-PCR. Optionally, the kit may include an agent that can differentially indicate a sequence variation within the genomic sequence following its amplification, e.g., an oligonucleotide probe that specifically binds to one version of the genomic sequence but not to other versions. The kit typically further includes an instruction manual. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1: Summary of study design. CHB, Han Chinese in Beijing, China; JPT, Japanese in Tokyo, Japan.
[0015] Figure 2: Manhattan plot of combined genome-wide association results from the Hong Kong 1, Hong Kong 2 and Shanghai studies based on the random effect models. The j-axis represents the -logio p value, and the x-axis represents the 2,925,090 analysed SNPs. The dotted line indicates the threshold of significance p<\ x 10 5. There are 44 points with p<\ x 10~5, and the arrows and labels localise the susceptibility loci to type 2 diabetes uncovered in the present study.
[0016] Figure 3: Regional plots for the identified variant rsl0229583, including results for both genotyped and imputed SNPs in the Chinese population. The top positioned circle and diamond represent the sentinel SNP in meta-analysis of three GWAS in the stage 1 and the East Asian meta-analysis in stages 1+2+3, respectively. Other SNPs are coloured according to their level of LD, which is measured by r2, with the sentinel SNP. The recombination rates estimated from the 1000 Genomes project JPT+CHB data are shown. CHB, Han Chinese in Beijing, China; JPT, Japanese in Tokyo, Japan.
[0017] Figure 4: Forest plot for meta-analysis of the association between type 2 diabetes and rs 10229583 for all populations in the present study. ORs and 95% CIs were reported with respect to the type 2 diabetes-related risk alleles (G).
[0018] Figure 5: Associations of the risk variant (G allele) of rsl0229583 with measures of insulin secretion in Chinese controls, (a) Association with reduced ΗΟΜΑ-β in a Hong Kong Chinese adolescent cohort (p=0.0221). (b) Association with a reduced Stumvoll Index of beta cell function in healthy Shanghai controls (p=0.0303). (c) Association of the risk variant with higher FPG in healthy Shanghai controls (p=0.0460). Data are expressed as mean (for FPG) or geometric mean (for ΗΟΜΑ-β and Stumvoll Index). SDs or 95% CIs are expressed as error bars. The number of individuals analysed for each genotype is shown in parentheses under each column.
[0019] Figure 6: Multidimensional scaling analysis (MDS) plot showing the first two principal components, based on genotype data of 1 1 populations from HapMap (African ancestry in Southwest USA (ASW), Utah residents with Northern and Western European ancestry from the CEPH collection (CEU), Han Chinese in Beijing, China (CHB), Chinese in Metropolitan Denver, Colorado (CHD), Gujarati Indians in Houston, Texas (GIH), Japanese in Tokyo, Japan (JPT), Luhya in Webuye, Kenya (LWK), Mexican ancestry in Los Angeles, California (MEX), Maasai in Kinyawa, Kenya (MKK), Tuscan in Italy (TSI) and Yoruban in Ibadan, Nigeria (YRI)), as well as the 3 case-controls cohorts (Hong Kong GWAS 1 (HK1), Hong Kong GWAS 2 (HK2) and Shanghai GWAS (SH)) in the stage 1 genome scan of the present study.
[0020] Figure 7: Multidimensional scaling analysis (MDS) plot shows the first two principal components, based on genotype data of 3 case-controls cohorts (Hong Kong GWAS 1 , Hong Kong GWAS 2 and Shanghai GWAS) in the stage 1 genome scan of the present study without HapMap scaling.
[0021] Figure 8: Q-Q plot for combined genome-wide association results in a total of 684 T2D patients and 955 controls based on the 2,925,090 analyzed SNPs. The curvy lines above and below the straight diagonal lines represent the upper and lower boundaries of the 95% confidence bands.
[0022] Figure 9: Distribution of ENCODE open chromatin sites around the identified gene region (NCBI Build 36.1/hgl8 CHR7:127033000-127084685) annotated in the UCSC human genome browser on human (website: genome.ucsc.edu/). The distribution of open chromatin across the identified gene region in pancreatic islets (depicted as peaks) is highlighted inside the box (labeled Panlsle FAIRE FD and highlighted by red arrow). The position of rs 10229583 and other tagging SNPs in high LD to this lead SNP are marked by the black arrows at the top of the figure.
[0023] Figure 10: Comparison of varLD scores within 100Kb region centred on index SNP rsl0229583 between pairs of populations using HapMap phase III CHB, JPT, CEU and YRI data.
[0024] Figure 11: Linkage disequilibrium for SNPs within the region near PAX4 on chromosome 7 between 126.95 Mb and 127.06 Mb (Build 36). Pairwise r2 among SNPs for HapMap CEU and CHB are indicated in upper and lower block, respectively. Shades of grey represent the strength of pairwise r2. Rsl0229583 and rs6467136 are the SNPs showing significant association with T2D in the present study and the study conducted by the East Asian Consortium, respectively. DEFINITIONS
[0025] The term "type 2 diabetes" (T2D) refers to a metabolic disorder that is characterized by high blood glucose in the context of varying combinations of insulin resistance and insulin deficiency. Type 2 diabetes may be caused by a combination of lifestyle and genetic factors. Diabetes can be caused by distinct clinical entities such as endocrine disorders (e.g., Cushing's syndrome) and chronic pancreatitis. However, the majority of people with diabetes have risk factors including but not limited to obesity, hypertension, high blood cholesterol, metabolic syndrome (high triglyceride, low HDL-C, high blood glucose, high blood pressure, large waist), which may share common metabolic pathways, further amplified by aging, energy dense diets (e.g., high- fat and high glucose), sedentary lifestyle and use of certain drugs (e.g., beta blockers, steroids). On the other hand, having relatives (especially first degree) with type 2 diabetes increases risks of developing type 2 diabetes substantially. Symptoms of type 2 diabetes often include polyuria (frequent urination), polydipsia (increased thirst), polyphagia (increased hunger), fatigue, and weight loss. The abnormal neurohormonal and metabolic milieu characterized by hyperglycemia, dyslipidemia and low grade inflammation can trigger a cascade of signaling pathways, which can lead to cell death and dysregulated cell growth, giving rise to multiple morbidities including heart disease, strokes, limb amputation, visual loss, kidney failure, cancers, and cognitive impairment.
[0026] The term "cardiovascular disease" refers to a broad class of diseases that involve the heart or blood vessels (arteries and veins) and affect the cardiovascular system, such as conditions related to atherosclerosis (arterial disease). These include but not limited to stroke, coronary heart disease and peripheral vascular disease. Known risk factors for cardiovascular diseases include unhealthy eating, lack of exercise, obesity, suboptimally managed diabetes, abnormal blood lipids, high blood pressure, excessive consumption of alcohol, use of tobacco, as well as genetic background. In this application, the term "diabetic cardiovascular disease" specifically refers to a cardiovascular disease that is associated with or secondary to diabetes.
[0027] As used herein, the term "body mass index" or "BMI" refers to a number calculated from a person's weight and height to reflect the "fatness" or "thinness" of a person. More specifically, BMI = mass (kg) / (height (m))2 or mass (lb) x 703 / (height (in))2. Typically, in Caucasian populations, a BMI of 20 to 25 kg/m2 is considered optimal weight; a BMI lower than 20 kg/m2 suggests the person is underweight whereas a BMI above 25 kg/m2 may indicate the person is overweight; a BMI above 30 kg/m suggests the person is obese; and a BMI over 40 kg/m indicates the person to be morbidly obese. Compared to Caucasians, Asians have more body fat for the same degree of BMI and waist circumference. Thus, normal weight and obesity
2 2
in Asians are defined as <23 kg/m and >25 kg/m respectively. While high BMI may predict risk for diabetes or prediabetes, people with low BMI, which correlates with beta cell function, are also at high risk, especially if these subjects develop central obesity, which tends to be associated with insulin resistance or reduced insulin sensitivity.
[0028] In this disclosure, the term "biological sample" or "sample" includes any section of tissue or bodily fluid taken from a test subject such as a biopsy and autopsy sample, and frozen section taken for histologic purposes, or processed forms of any of such samples. Biological samples include blood and blood fractions or products (e.g., serum, plasma, platelets, white blood cells, red blood cells, and the like), sputum or saliva, lymph and tongue tissue, cultured cells, e.g. , primary cultures, explants, and transformed cells, stool, urine, stomach biopsy tissue etc. A biological sample is typically obtained from a eukaryotic organism, which may be a mammal, may be a primate and may be a human subject. [0029] In this disclosure, the term "biopsy" refers to the process of removing a tissue sample for diagnostic or prognostic evaluation, and to the tissue specimen itself. Any biopsy technique known in the art can be applied to the methods of the present invention. The biopsy technique applied will depend on the tissue type to be evaluated (e.g., tongue, colon, prostate, kidney, bladder, lymph node, liver, bone marrow, blood cell, stomach tissue, etc.) among other factors. Representative biopsy techniques include, but are not limited to, excisional biopsy, incisional biopsy, needle biopsy, surgical biopsy, and bone marrow biopsy and may comprise endoscopy such as colonoscopy. A wide range of biopsy techniques are well known to those skilled in the art who will choose between them and implement them with minimal experimentation.
[0030] In this disclosure, the term "isolated" nucleic acid molecule means a nucleic acid molecule that is separated from other nucleic acid molecules that are usually associated with the isolated nucleic acid molecule. Thus, an "isolated" nucleic acid molecule includes, without limitation, a nucleic acid molecule that is free of nucleotide sequences that naturally flank one or both ends of the nucleic acid in the genome of the organism from which the isolated nucleic acid is derived (e.g., a cDNA or genomic DNA fragment produced by a polymerase chain reaction or restriction endonuclease digestion). Such an isolated nucleic acid molecule is generally introduced into a vector (e.g., a cloning vector or an expression vector) for convenience of manipulation or to generate a fusion nucleic acid molecule. In addition, an isolated nucleic acid molecule can include an engineered nucleic acid molecule such as a recombinant or a synthetic nucleic acid molecule. A nucleic acid molecule existing among hundreds to millions of other nucleic acid molecules within, for example, a nucleic acid library (e.g., a cDNA or genomic library) or a gel (e.g., agarose, or polyacrylamine) containing restriction-digested genomic DNA, is not an "isolated" nucleic acid.
[0031] The term "nucleic acid" or "polynucleotide" refers to deoxyribonucleic acids (DNA) or ribonucleic acids (RNA) and polymers thereof in either single- or double-stranded form.
Unless specifically limited, the term encompasses nucleic acids containing known analogs of natural nucleotides that have similar binding properties as the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions), alleles, orthologs, single nucleotide
polymorphisms (SNPs), and complementary sequences as well as the sequence explicitly indicated. Specifically, degenerate codon substitutions may be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and/or deoxyinosine residues (Batzer et ah, Nucleic Acid Res. 19:5081 (1991);
Ohtsuka et ah, J. Biol. Chem. 260:2605-2608 (1985); and Rossolini et ah, Mol. Cell. Probes 8:91-98 (1994)). The term nucleic acid is used interchangeably with gene, cDNA, and mRNA encoded by a gene.
[0032] The term "gene" means the segment of DNA involved in producing a polypeptide chain; it includes regions preceding and following the coding region (leader and trailer) involved in the transcription and/or translation of the gene product and the regulation of the transcription and/or translation, as well as intervening sequences (introns) between individual coding segments (exons).
[0033] In this application, the terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to a polymer of amino acid residues. The terms apply to amino acid polymers in which one or more amino acid residue is an artificial chemical mimetic of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers and non-naturally occurring amino acid polymers. As used herein, the terms encompass amino acid chains of any length, including full-length proteins (i.e., antigens), wherein the amino acid residues are linked by covalent peptide bonds.
[0034] The term "amino acid" refers to naturally occurring and synthetic amino acids, as well as amino acid analogs and amino acid mimetics that function in a manner similar to the naturally occurring amino acids. Naturally occurring amino acids are those encoded by the genetic code, as well as those amino acids that are later modified, e.g., hydroxyproline, γ-carboxyglutamate, and O-phosphoserine. For the purposes of this application, amino acid analogs refers to compounds that have the same basic chemical structure as a naturally occurring amino acid, i.e., a carbon that is bound to a hydrogen, a carboxyl group, an amino group, and an R group, e.g., homoserine, norleucine, methionine sulfoxide, methionine methyl sulfonium. Such analogs have modified R groups (e.g., norleucine) or modified peptide backbones, but retain the same basic chemical structure as a naturally occurring amino acid. For the purposes of this application, amino acid mimetics refer to chemical compounds that have a structure that is different from the general chemical structure of an amino acid, but that functions in a manner similar to a naturally occurring amino acid.
[0035] Amino acids may include those having non-naturally occurring D-chirality, as disclosed in WO01/12654, which may improve the stability (e.g., half-life), bioavailability, and other characteristics of a polypeptide comprising one or more of such D-amino acids. In some cases, one or more, and potentially all of the amino acids of a therapeutic polypeptide have D-chirality.
[0036] Amino acids may be referred to herein by either the commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical
Nomenclature Commission. Nucleotides, likewise, may be referred to by their commonly accepted single-letter codes.
[0037] The term "immunoglobulin" or "antibody" (used interchangeably herein) refers to an antigen-binding protein having a basic four-polypeptide chain structure consisting of two heavy and two light chains, said chains being stabilized, for example, by interchain disulfide bonds, which has the ability to specifically bind antigen. Both heavy and light chains are folded into domains.
[0038] The term "antibody" also refers to antigen- and epitope-binding fragments of antibodies, e.g., Fab fragments, that can be used in immunological affinity assays. There are a number of well characterized antibody fragments. Thus, for example, pepsin digests an antibody C-terminal to the disulfide linkages in the hinge region to produce F(ab)'2, a dimer of Fab which itself is a light chain joined to VR-CRI by a disulfide bond. The F(ab)'2 can be reduced under mild conditions to break the disulfide linkage in the hinge region thereby converting the (Fab')2 dimer into an Fab' monomer. The Fab' monomer is essentially a Fab with part of the hinge region (see, e.g., Fundamental Immunology, Paul, ed., Raven Press, N.Y. (1993), for a more detailed description of other antibody fragments). While various antibody fragments are defined in terms of the digestion of an intact antibody, one of skill will appreciate that fragments can be synthesized de novo either chemically or by utilizing recombinant DNA methodology. Thus, the term antibody also includes antibody fragments either produced by the modification of whole antibodies or synthesized using recombinant DNA methodologies.
[0039] The phrase "specifically binds," when used in the context of describing a binding relationship of a particular molecule to a protein or peptide, refers to a binding reaction that is determinative of the presence of the protein in a heterogeneous population of proteins and other biologies. Thus, under designated binding assay conditions, the specified binding agent (e.g., an antibody) binds to a particular protein at least two times the background and does not substantially bind in a significant amount to other proteins present in the sample. Specific binding of an antibody under such conditions may require an antibody that is selected for its specificity for a particular protein or a protein but not its similar "sister" proteins. A variety of immunoassay formats may be used to select antibodies specifically immunoreactive with a particular protein or in a particular form. For example, solid-phase ELISA immunoassays are routinely used to select antibodies specifically immunoreactive with a protein (see, e.g., Harlow & Lane, Antibodies, A Laboratory Manual (1988) for a description of immunoassay formats and conditions that can be used to determine specific immuno reactivity). Typically a specific or selective binding reaction will be at least twice background signal or noise and more typically more than 10 to 100 times background. On the other hand, the term "specifically bind" when used in the context of referring to a polynucleotide sequence forming a double-stranded complex with another polynucleotide sequence describes "polynucleotide hybridization" based on the Watson-Crick base-pairing, as provided in the definition for the term "polynucleotide hybridization method."
[0040] A "polynucleotide hybridization method" as used herein refers to a method for detecting the presence and/or quantity of a pre-determined polynucleotide sequence based on its ability to form Watson-Crick base-pairing, under appropriate hybridization conditions, with a polynucleotide probe of a known sequence. Examples of such hybridization methods include Southern blot, Northern blot, and in situ hybridization.
[0041] "Primers" as used herein refer to oligonucleotides that can be used in an amplification method, such as a polymerase chain reaction (PCR), to amplify a nucleotide sequence based on the polynucleotide sequence corresponding to a gene of interest, e.g., the cDNA or human genomic sequence PAX4-SND1 or a portion thereof. Typically, at least one of the PCR primers for amplification of a polynucleotide sequence is sequence-specific for that polynucleotide sequence. The exact length of the primer will depend upon many factors, including temperature, source of the primer, and the method used. For example, for diagnostic and prognostic applications, depending on the complexity of the target sequence, the oligonucleotide primer typically contains at least 10, or 15, or 20, or 25 or more nucleotides, although it may contain fewer nucleotides or more nucleotides. The factors involved in determining the appropriate length of primer are readily known to one of ordinary skill in the art. In this disclosure the term "primer pair" means a pair of primers that hybridize to opposite strands a target DNA molecule or to regions of the target DNA which flank a nucleotide sequence to be amplified. In this disclosure, the term "primer site" means the area of the target DNA or other nucleic acid to which a primer hybridizes.
[0042] A "label," "detectable label," or "detectable moiety" is a composition detectable by spectroscopic, photochemical, biochemical, immunochemical, chemical, or other physical means. For example, useful labels include 32P, fluorescent dyes, electron-dense reagents, enzymes {e.g., as commonly used in an ELISA), biotin, digoxigenin, or haptens and proteins that can be made detectable, e.g., by incorporating a radioactive component into the peptide or used to detect antibodies specifically reactive with the peptide. Typically a detectable label is attached to a probe or a molecule with defined binding characteristics {e.g., a polypeptide with a known binding specificity or a polynucleotide), so as to allow the presence of the probe (and therefore its binding target) to be readily detectable.
[0043] A "standard sequence" as used herein refers to the polynucleotide sequence of a predetermined genomic DNA segment, e.g., a defined portion or the entire length of a human genomic sequence of a given range and location, such as the human PAX4-SND1 genomic sequence, including 2 kb upstream and 2 kb downstream flanking sequences, that is present in a publically accessible database, e.g., the University of California Santa Cruz database (hgl 8), as the standard human genomic sequence for these particular genes. When a genomic DNA sequence determined from a test sample is compared with a "standard sequence," the test sequence is aligned with the "standard sequence" at the corresponding nucleotide bases of the genomic sequence to reveal any sequence variation. For this particular application, the standard genomic sequences for human PAX4 and SND1 genes (including some isoforms) are provided as below:
PAX4 (paired box 4, chr7: 127250346-127255982, hgl9) is the closest gene to the SNP rsl0229583 (chr7: 127246903, hgl9). The PAX4 gene encodes 4 protein-coding isoforms (Tablel). Because the PAX4 gene is at the reverse (-) strand, rsl0229583 is about 3.4 kb downstream to the most 3 ' exon of the PAX4 gene.
Table 1 Summary of PAX4 isoforms
i me Transcript ID Length Protein ID Length Exon Start End
(bp) (dd)
PAX4-001 ENST00000341640. 2010 ENSP00000339906, 343 9 127250346 127255780
NM 006193.2 NP 006184.2
PAX4-002 ENST00000463946 613 ENSP00000451923 341 8 127250992 127255724
PAX4-004 ENST00000378740 I 08S ENSP00000368014 34S ίθ 127250865 127255982
PAX4-201 ENST00000338516 ENSP00000344297 273 9 127250346 127255982 [0044] The term "amount" as used in this application refers to the quantity of a polynucleotide of interest or a polypeptide of interest present in a sample. Such quantity may be expressed in the absolute terms, i.e. , the total quantity of the polynucleotide or polypeptide in the sample, or in the relative terms, i.e. , the concentration of the polynucleotide or polypeptide in the sample.
[0045] The term "effective amount" as used herein refers to an amount of a given substance that is sufficient in quantity to produce a desired effect. For example, an effective amount of a cholesterol lowering drug or a blood glucose lowering drug is the amount of said drug to achieve a decreased level of cholesterol or blood glucose, respectively, in a patient who has been given the drug for therapeutic purposes. An amount adequate to accomplish this is defined as the "therapeutically effective dose." The dosing range varies with the nature of the therapeutic agent being administered and other factors such as the route of administration and the severity of a patient's condition.
[0046] The term "subject" or "subject in need of treatment," as used herein, includes individuals who seek medical attention due to risk of, or actual suffering from type 2 diabetes or cardiovascular/renal disease associated with diabetes. Subjects also include individuals currently undergoing therapy that seek manipulation of the therapeutic regimen. Subjects or individuals in need of treatment include those that demonstrate symptoms of type 2 diabetes or related cardiovascular/renal disease, or are at risk of suffering from type 2 diabetes or diabetic cardiovascular/renal disease or related symptoms. For example, a subject in need of treatment includes individuals with a genetic predisposition or family history for type 2 diabetes or diabetic cardiovascular/renal disease, those who have suffered relevant symptoms in the past, those who have been exposed to a triggering substance or event, as well as those suffering from chronic or acute symptoms of the condition. A "subject in need of treatment" may be at any age of life.
DETAILED DESCRIPTION OF THE INVENTION
I. Introduction
[0047] The present inventors performed studies to identify new type 2 diabetes susceptibility loci in Southern Han Chinese individuals. A meta-analysis was performed of three GWAS comprising 684 patients with type 2 diabetes and 955 controls, and analysed 2.9 million
(genotyped and imputed) single-nucleotide polymorphisms (SNPs) in an additive model.
Putatively associated SNPs (p<l >< 10~5) were genotyped de novo in two independent Southern Han Chinese cohorts (10,383 cases and 6,974 controls), and SNPs reaching a genome- wide significance of /?<5x 10" were replicated in silico in five East Asian and three non-East Asian populations for a total of 31,541 cases and 60,344 controls.
[0048] The inventors discovered for the first time the correlation between genomic sequence variation in the human PAX4-SND1 genomic sequence and medical conditions such as type 2 diabetes and diabetic cardiovascular and renal diseases in human subjects. This discovery allows medical professionals to identify subjects at risk cardiovascular or renal disease in a patient with type 2 diabetes or assess the risk of developing type 2 diabetes and diabetic cardiovascular and/or renal disease in a subject at risk by studying the subject's PAX4-SND1 genomic sequence and then comparing the subject's sequence with a standard PAX4-SND1 genomic sequence that has been determined as a part of the standard human genome. Detection of such sequence variation(s) indicates the presence or elevated risk of developing type 2 diabetes or diabetic cardiovascular and/or renal disease in the subject, as well as the early onset of these conditions. The detection of pertinent genomic sequence variation(s) can further guide physicians to devise or modify treatment plans for a subject in both prevention and therapeutic measures.
[0049] It is of particular interest that the present inventors revealed that type 2 diabetes patients have an increased risk of coronary heart disease (CHD) when they possess the genomic sequence variation in the human PAX4-SND1 genomic sequence at 7q32. A recent genome -wide association study in the Chinese populations identified association between a novel variant at 7q32 near paired box 4 (PAX4) and T2D, which was confirmed in other East Asian populations. This study aimed to investigate the association of this novel 7q32 variant and CHD risk in an 8- year prospective cohort of Chinese patients with T2D. The 7q32 variant was genotyped in 5264 T2D patients [age mean ± SD = 56.3 ± 13.3 years, % males = 45.0] without history of CHD at baseline. Associations of the variant under multiplicative, dominant and recessive genetic models with new onset of CHD were examined by Cox proportional hazard models with adjustment of conventional risk factors at baseline. During the mean follow-up period of 6.9 ± 6.7 years, 395 patients (7.5%) developed incident CHD. Subjects who carried the common, T2D risk allele (G) of the 7q32 variant have higher risk for CHD (hazard ratio [HR] (95% CI) =
I .43 (1.15 - 1.79), P = 1.4 x 10"3 under additive model; HR (95% CI) = 1.56 (1.22 - 2.00), P = 3.7 x 10"4 under recessive model) after adjusting for sex, age, duration of diabetes, smoking status and HbAic. Further adjustments of conventional risk factors including body mass index, lipid profiles, blood pressure, history of retinopathy and neuropathy, urinary albumin excretion rate, glomerular filtration rate and use of medications did not change the association. In summary, these findings suggest that the novel T2D-associated 7q32 variant near PAX4 was not only associated with T2D, but also identifies T2D subjects with increased risk of CHD.
II. General Methodology
[0050] Practicing this invention utilizes routine techniques in the field of molecular biology.
Basic texts disclosing the general methods of use in this invention include Sambrook and Russell, Molecular Cloning, A Laboratory Manual (3rd ed. 2001); Kriegler, Gene Transfer and
Expression: A Laboratory Manual (1990); and Current Protocols in Molecular Biology (Ausubel
Figure imgf000016_0001
[0051] For nucleic acids, sizes are given in either kilobases (kb) or base pairs (bp). These are estimates derived from agarose or acrylamide gel electrophoresis, from sequenced nucleic acids, or from published DNA sequences. For proteins, sizes are given in kilodaltons (kDa) or amino acid residue numbers. Protein sizes are estimated from gel electrophoresis, from sequenced proteins, from derived amino acid sequences, or from published protein sequences. [0052] Oligonucleotides that are not commercially available can be chemically synthesized, e.g. , according to the solid phase phosphoramidite triester method first described by Beaucage and Caruthers, Tetrahedron Lett. 22: 1859-1862 (1981), using an automated synthesizer, as described in Van Devanter et. ah, Nucleic Acids Res. 12:6159-6168 (1984). Purification of oligonucleotides is performed using any art-recognized strategy, e.g., native acrylamide gel electrophoresis or anion-exchange high performance liquid chromatography (HPLC) as described in Pearson and Reanier, J. Chrom. 255: 137-149 (1983).
[0053] The sequence of interest used in this invention, e.g., the polynucleotide sequence of the human PAX4 and SND1 genes, and synthetic oligonucleotides (e.g., primers) can be verified using, e.g., the chain termination method for sequencing double-stranded templates of Wallace et al., Gene 16: 21-26 (1981). III. Acquisition of Biological Samples and Analysis of Genomic DNA Sequence
[0054] The present invention relates to determining at least a portion of the genomic sequence of a pertinent region, such as the human PAX4-SND1 segment and/or its transcript(s), found in a biological sample taken from a person being tested, as a means to detect the presence and/or to assess the risk of developing type 2 diabetes, or cardiovascular/renal diseases in that person. Thus, the first steps of practicing this invention are to obtain a biological sample (e.g., tissue or bodily fluid sample) from a test subject and extract genomic DNA or RNA from the sample.
A. Acquisition and Preparation of Samples
[0055] A biological sample is obtained from a person to be tested or assessed for risk of developing type 2 diabetes or associated cardiovascular or renal disease using a method of the present invention. Collection of a tissue or fluid sample from an individual is performed in accordance with the standard protocol laboratories, hospitals or clinics generally follow, such as during a biopsy, blood drawing, saliva collection, or oral swab. An appropriate amount of sample is collected and may be stored according to standard procedures prior to further preparation.
[0056] The analysis of genomic DNA found in a subject's sample according to the present invention may be performed using essentially any tissue or bodily fluid, so long as genomic DNA is expected to be present in such sample. The methods for preparing tissue or fluid samples for nucleic acid extraction are well known among those of skill in the art. For example, a subject's epithelial tissue sample should be first treated to disrupt cellular membrane so as to release nucleic acids contained within the cells. B. Determination of Genomic Sequence
[0057] Possible sequence variation within a segment of a pertinent genomic sequence (such as the human PAX4-SND1 sequence), or one or more of its transcripts, is investigated to provide indication as to whether a test subject is suffering from type 2 diabetes and associated cardiovascular or renal disease, or whether the subject is at risk of developing type 2 diabetes and associated complications including cardiovascular or renal disease in the future.
[0058] Typically a segment of the genomic sequence of an appropriate length is selected for sequencing analysis. The segment may be chosen from the genomic sequence of a pertinent gene or genes defined by the same boundaries defining the gene's genomic sequence, plus about 2,000 base pairs upstream and downstream from the boundaries. For instance, the human PAX4- SND1 genomic sequence will encompass the upstream boundary of the upstream gene, PAX4, to the downstream boundary of the downstream gene, SNDl, plus 2000 bp upstream and downstream from the upstream and downstream boundaries, respectively. In other words, the PAX4-SND1 genomic sequence may be as long as: 2000 bp genomic sequence immediately upstream from the PAX4 genomic sequence + the PAX4 genomic sequence + the genomic sequence between the PAX4 and SNDl genomic sequences + the SNDl genomic sequence + 2000 bp genomic sequence immediately downstream from the SNDl genomic sequence. The length of the genomic sequence being analyzed may be a segment of the above and is usually at least 15 or 20 contiguous nucleotides, and may be longer with at least 25, 30, 50, 100, 200, 300, 400, or more contiguous nucleotides.
1. DNA Extraction and Treatment
[0059] Methods for extracting DNA from a biological sample are well known and routinely practiced in the art of molecular biology, see, e.g., Sambrook and Russell, supra. RNA contamination should be eliminated to avoid interference with DNA analysis. Optionally, other components (such as proteins and lipids) may be removed from the biological sample prior to further analysis of the genomic DNA.
2. Optional Amplification and Sequence Analysis
[0060] Following the desired processing of DNA RNA in a biological sample, the DNA/RNA is then subjected to sequence-based analysis, such that the genomic sequence of one or more of the pertinent genes, or one or more of its transcripts, found in a test subject may be determined and then compared with a standard sequence to detect any possible sequence variation. An amplification reaction is optional prior to the sequence analysis. A variety of polynucleotide amplification methods are well established and frequently used in research. For instance, the general methods of polymerase chain reaction (PCR) for polynucleotide sequence amplification are well known in the art and are thus not described in detail herein. For a review of PCR methods, protocols, and principles in designing primers, see, e.g. , Innis, et ah, PCR Protocols: A Guide to Methods and Applications, Academic Press, Inc. N.Y., 1990. PCR reagents and protocols are also available from commercial vendors, such as Roche Molecular Systems.
[0061] Although PCR amplification is typically used in practicing the present invention, one of skill in the art will recognize that amplification of the relevant genomic sequence may be accomplished by any known method, such as the ligase chain reaction (LCR), transcription- mediated amplification, and self-sustained sequence replication or nucleic acid sequence-based amplification ( ASBA), each of which provides sufficient amplification.
[0062] Techniques for polynucleotide sequence determination are also well established and widely practiced in the relevant research field. For instance, the basic principles and general techniques for polynucleotide sequencing are described in various research reports and treatises on molecular biology and recombinant genetics, such as Wallace et ah, supra; Sambrook and Russell, supra, and Ausubel et ah, supra. DNA sequencing methods routinely practiced in research laboratories, either manual or automated, can be used for practicing the present invention. Additional means suitable for determining the polynucleotide sequence of a genomic DNA for practicing the methods of the present invention include but are not limited to mass spectrometry, primer extension, polynucleotide hybridization, real-time PCR, melting curve analysis, high resolution melting analysis, heteroduplex analysis, pyrosequencing, and electrophoresis.
3. Determining Genomic DNA Sequence Variation Based on RNA Sequence Variation
[0063] As an alternative, genomic DNA sequence variations may also be detected by way of analyzing RNA sequences transcribed from the pertinent DNA sequences, which may include portion of the coding sequence or non-coding sequence of a genomic locus of interest (e.g., the PAX4-SND1 genomic sequence). Methods for RNA extraction from a biological sample, sequence analysis of RNA or DNA molecules, optionally involving amplification techniques such as reverse transcription based amplification processes, e.g., RT-PCR, are well known in the art. Suitable samples for RNA sequence analysis may include peripheral blood monocytes (PBMC) and specific tissue samples such as fat and muscles.
IV. Corresponding Standard Sequences
[0064] In order to practice the method of this invention, the standard genomic sequence(s) for one or more pertinent genes, such as the human PAX4 and SND1 genes and their isoforms, will be chosen before the comparison with a test subject's genomic sequence of the corresponding gene at the corresponding location may be performed.
V. Therapeutic and Preventive Measures
[0065] By illustrating the correlation between genomic sequence variation in one or more of the specific genes named above and the presence or heightened risk of developing type 2 diabetes or cardiovascular and renal diseases among subjects having such variation, especially those fitting certain profiles, such as those of Asian descent, in particular Han Chinese, the present inventors have provided a valuable tool for clinicians to determine, often in combination with other information and diagnostic or predictive or screening test results, how a subject having certain genomic sequence variation(s) should be monitored and/or treated for type 2 diabetes and diabetic cardiovascular or renal disease such that the symptoms of these conditions may be prevented, eliminated, ameliorated, reduced in severity and/or frequency, or delayed in their onset. For example, a physician may arrange for regular monitoring of various symptoms of type 2 diabetes or diabetic cardiovascular and renal diseases in a subject who has been deemed by the method of the present invention to have an elevated risk of developing type 2 diabetes. The physician may also prescribe both pharmacological and non-pharmacological treatments such as lifestyle modification {e.g., reduce body weight by 5%, high fiber diet, walking for at least 150 minutes weekly) and medicines known to reduce risk of onset of diabetes (e.g., metformin, alpha glucosidase inhibitors, lipase inhibitors) to a subject who has been deemed by the method of the present invention to have an elevated risk of developing type 2 diabetes. For a subject who has been deemed by the method of the present invention to suffer from or at risk of developing diabetic cardiovascular or renal disease, the attending physician may prescribe medications to control risk factors such as high levels of blood cholesterol and triglycerol (e.g., statins and fibrates) and reduce angiotensin II activity (e.g., Angiotensin converting enzyme inhibitor (ACEI) and angiotensin II receptor blocker (ARB)), as well as place the subject under regular testing and monitoring of coronary artery condition and kidney function. VI. KITS AND DEVICES
[0066] The present invention provides compositions and kits for practicing the methods described herein to detect possible genomic sequence variation of certain gene(s) and the transcripts thereof in a subject, which can be used for various purposes such as detecting or diagnosing the presence of type 2 diabetes and diabetic cardiovascular or renal disease in a subject, determining the risk of developing type 2 diabetes and diabetic cardiovascular or renal disease in a subject, and guiding the treatment plan for these conditions in the subject.
[0067] Kits for carrying out assays for determining the nucleotide sequence of a relevant genomic sequence typically include at least one oligonucleotide useful for specific hybridization with a predetermined segment of a pertinent genomic sequence (e.g., human PAX4-SND1 genomic sequence). Optionally, this oligonucleotide is labeled with a detectable moiety. In some cases, the oligonucleotide specifically hybridizes with the standard sequence only but not with any of the variant sequences. In other cases, the oligonucleotide specifically hybridizes with one particular version of the variant sequence but not with other versions, nor with the standard sequence.
[0068] In some cases, the kits may include at least two oligonucleotide primers that can be used in the amplification of at least one segment of a pertient genomic sequence (such as the human PAX4-SND1 genomic sequence) or transcripts thereof by PCR. In some examples, at least one of the oligonucleotide primers is designed to anneal only to the standard sequence or only to a particular version of the variant sequences, for example, the G allele of rs 10229583.
[0069] In addition, the kits of this invention may provide instruction manuals {e.g., internet- based decision support tools) to guide users in analyzing test samples and assessing the presence or future risk of type 2 diabetes and diabetic cardiovascular or renal disease in a test subject.
[0070] Furthermore, the present invention can also be embodied in a device or a system comprising one or more such devices, which is capable of carrying out all or some of the method steps described herein. For instance, in some cases, the device or system performs the following steps upon receiving a biological sample taken from a subject being tested for detecting type 2 diabetes or diabetic cardiovascular or renal disease, assessing the risk of developing type 2 diabetes or diabetic cardiovascular or renal disease, or guiding treatment of a subject having or at risk of developing any one of these conditions: (a) determining in the sample the nucleotide sequence of a pertinent genomic DNA segment or its transcript; (b) comparing the sequence determined from the sample with a corresponding standard sequence; and (c) providing an output indicating whether type 2 diabetes or diabetic cardiovascular/renal disease is present in the subject or whether the subject is at risk of developing type 2 diabetes or diabetic
cardiovascular/renal disease. In other cases, the device or system of the invention performs the task of steps (b) and (c), after step (a) has been performed and the genomic sequence determined from (a) has been entered into the device. Preferably, the device or system is partially or fully automated.
EXAMPLES
[0071] The following examples are provided by way of illustration only and not by way of limitation. Those of skill in the art will readily recognize a variety of non-critical parameters that could be changed or modified to yield essentially the same or similar results.
Example 1: Genomic Sequence Variation at 7q32 Linked to Type 2 Diabetes
INTRODUCTION
[0072] Most genetic variants identified for type 2 diabetes have been discovered in European populations. Genome -wide association studies (GWAS) was performed in a Chinese population with the aim of identifying novel variants for type 2 diabetes in Asians.
[0073] The present inventors performed a meta-analysis of three GWAS comprising 684 patients with type 2 diabetes and 955 controls of Southern Han Chinese descent. The inventors followed up the top signals in two independent Southern Han Chinese cohorts (totalling 10,383 cases and 6,974 controls) and performed in silico replication in multiple populations. They identified CDKN2A/B and four novel type 2 diabetes association signals with p< 1 x 10"5 from the meta-analysis. Thirteen loci within these four loci were followed up in two independent Chinese cohorts, and rs 10229583 at 7q32 was found to be associated with type 2 diabetes in a combined analysis of 1 1,067 cases and 7,929 controls
Figure imgf000022_0001
l0"8; OR [95% CI] 1.18 [1.11 , 1.25]). In silico replication revealed consistent associations across multiethnic groups, including five East
Asian populations
Figure imgf000022_0002
). The rs 10229583 risk variant was associated with elevated fasting plasma glucose, impaired beta cell function in controls, and an earlier age at diagnosis for the cases. The novel variant lies within an islet-selective cluster of open regulatory elements. There was significant heterogeneity of effect between Han Chinese and individuals of European descent, Malaysians and Indians. rs 10229583 near PAX4 is identified as a novel locus for type 2 diabetes in Chinese and other populations and provides new insights into the pathogenesis of type 2 diabetes.
METHODS
[0074] Participants In the first-stage discovery cohort (stage 1), genome-wide scanning was performed in three different case-control samples: 198 Hong Kong Chinese individuals (99 patients with type 2 diabetes and 99 healthy controls) in Hong Kong GWAS 1 , 1 ,047 Hong Kong Chinese individuals (388 with type 2 diabetes and 659 controls) in Hong Kong GWAS 2 and 394 Shanghai Chinese (197 patients with type 2 diabetes and 197 normal controls) in the Shanghai GWAS. Individuals included in the stage 2 replication included 5,366 with type 2 diabetes and 2,474 controls from Hong Kong, and 4,035 cases and 3,964 controls from Shanghai. 325 cases and 368 controls from 178 Hong Kong families were also included, as well as 657 cases and 168 controls from 248 Shanghai families.
[0075] Case-control samples for in silico replication in stage 3 were taken from several published type 2 diabetes GWAS in East Asian individuals. These included the Korea
Association Resource Study (Cho et al. (2009) Nat Genet 41 : 527-534), the Singapore Chinese from the Singapore Diabetes Cohort Study and the Singapore Prospective Study Program (Tan et al. 2010 J Clin Endocrinol Metab 95: 390-397), the BioBank Japan Study (Unoki et al. (2008) Nat Genet 40: 1098-1 102) and a Han Chinese Study (Li et al. 2013 Diabetes doi: 10.2337/dbl2- 0454). For stage 4 in silico replication in other populations, Malaysian participants from the Singapore Malay Eye Study, Indian participants from the Singapore Indian Eye Study (Tan et al. 2010 J Clin Endocrinol Metab 95: 390-397) and participants of European descent in the Diabetes Genetics Replication And Meta-analysis (DIAGRAM) Consortium (Voight et al. (2010) Nat Genet 42: 579-589) were included.
[0076] The study design, type 2 diabetes diagnostic criteria and clinical evaluation used in each study are described in the electronic supplementary material (ESM) Methods. The clinical characteristics of the study individuals are described in Table 1. Each study obtained approval from the appropriate institutional review boards of the respective institutions, and written informed consent was obtained from all participants. The overall study design is depicted in Fig. 1. [0077] Quality control on the samples for the GWAS In this study, individuals were excluded from further analysis if: (1) duplicate samples existed; (2) the sex identified from the X chromosome was discordant with the sex obtained from the medical records; (3) the genotype call rate yield was <98%. Possible familial relationship was detected using estimates of identity by descent derived from pair-wise analyses of independence (r2~0) and quality SNPs.
Individuals with evidence for relatedness were excluded (·β>0.05). Table 4 shows the quality control for the participants in stage 1.
[0078] To discriminate individuals from different geographical origins, multidimensional scaling analysis was conducted using the genotype data obtained from unrelated individuals in the present study and the other 1 1 populations studied by the HapMap project (Fig. 6).
Individuals were excluded from subsequent analyses if they lay between clusters. [0079] Genotyping and quality control on the SNP data Individuals for the stage 1 , 3 and 4 analyses were genotyped using high-density SNP typing arrays that covered the entire genome. Only autosomal SNPs were included. Quality checks for SNPs were performed in the case and control samples separately, although the same criteria were applied to each. SNPs were excluded from further analysis if: (1) /?<l x l0~4 for HWE; (2) minor allele frequency (MAF) was <1%; (3) call rate was <95%; in particular, SNPs with MAF>1 % but <5% were excluded if their call rate was <99%; or (4) the SNPs showed a significant difference in MAF (p<\ x 1(T4) between the Hong Kong control cohorts with other conditions (450 with epilepsy cases, 1 10 with eczema and 99 non-hypertensive individuals). Only SNPs that passed the quality control criteria for both cases and controls were used for further analysis. Table 5 shows the quality control of the genotyping results in stage 1. Genotypes were imputed for autosomal SNPs according to the 1000 Genomes reference panel. See the ESM Methods for further details.
[0080] For de novo replication in stage 2, all selected SNPs were genotyped in the Hong Kong and Shanghai case-control samples by a primer extension of multiplex products with detection by Matrix-assisted laser desorption ionisation-time of flight mass spectroscopy using a
MassARRAY platform (Sequenom; San Diego, CA, USA). Family samples were genotyped using TaqMan SNP Genotyping Assays (Applied Biosystems, Foster City, CA, USA) or by direct sequencing.
[0081] Statistical analysis All statistical analyses were performed using PLINK version 1.07 (website: pngu.mgh.harvard.edu/~purcell/plink/) (Purcell et al (2007) Am J Hum Genet 81 :559— 575), SAS version 9.1 (SAS Institute, Cary, NC, USA) or SPSS for Windows version 18 (SPSS, Chicago, IL, USA), unless specified otherwise. Haploview version 4.1 was used to generate pair- wise linkage disequilibrium (LD) measures (r2).
[0082] To test for an association with type 2 diabetes, logistic regression was applied under an additive genetic model using the MACH2DAT software (Li et al. (2009) Annu Rev Genomics Hum Genet 10: 387-406) adjusted for sex and age according to situations in the individual studies.
[0083] To combine the type 2 diabetes association results in stage 1 , GWAMA software (website: well.ox.ac.uk/gwama/) (Magi and Morris BMC Bioinformatics 2010, 1 1 :288) was used to calculate the combined estimates of the ORs (95% CIs) from multiple groups by weighting the natural log-transformed ORs of each study using the inverse of their variance under the random effect model (DerSimonian and Laird (1986) Control Clin Trials 7: 177-188). By using the random effect model, SNPs with some degree of heterogeneity between studies were excluded, which helped to attenuate the number of false-positive findings in this study. Cochran's Q statistic (p<0.05) and / index were used to assess the heterogeneity of ORs between studies. [0084] The most strongly associated SNPs were prioritised for follow-up in stage 2 based on the meta-analysis results from stage 1. SNPs located within a previously reported type 2 diabetes locus were excluded. 13 top and proxy SNPs from four distinct loci available in all three GWAS were finally considered with (1) a meta-analysis p<\ x 10"5; (2) a heterogeneity test /?>0.05; (3) the same direction of risk allele across all three GWAS; (4) a common allele frequency
(MAF>0.1). For SNPs imputed across all three studies, the most significant SNP was selected associated with type 2 diabetes. Tables 6 and 7 provide the details of the selected SNPs and the quality control for the genotyping results in stage 2.
[0085] In the replication stage, genotype frequencies were compared between cases and controls using logistic regression under additive genetic model. In the family studies, alternating logistic regressions (ALR) implemented in the SAS procedure GENMOD was used to test for the association between type 2 diabetes and SNPs under an additive genetic model adjusted for age and sex. ALR is one type of generalised estimating equation applicable to binary outcomes that can handle correlated data (e.g., familial correlation). ORs (95% CIs) are presented in both analyses. Meta-analyses and heterogeneity tests were conducted as described previously to combine estimates of the ORs (95% CIs) from multiple case-control and family groups under the fixed effect model. Multiple testing in the combined analysis of the case-control study were controlled by Bonferroni correction, and p<4.5 x 10" (0.05 divided by 1 1 SNPs in the stage 2 replication studies) was used as the threshold for filtering SNPs genotyped in the family studies.
[0086] Continuous data are presented as mean±SD or geometric mean (95% CI). Traits were logio-transformed due to skewed distributions. Associations between genotypes and quantitative traits were tested by linear regression (adjusted for sex, age and/or BMI) in each healthy control cohort, as were associations for age at diagnosis (AAD) among patients with type 2 diabetes (adjusted for sex, BMI and/or HbAic). Meta-analyses implemented by GWAMA were applied to combine effect size (fi±SE) from multiple groups under the fixed effect model.
[0087] bioinformatics and czs-expression quantitative trait loci (eQTL) analysis was performed for functional implication of the identified SNP. See the ESM Methods for additional information on methods, including adjustment for genomic control and the gene network analysis.
RESULTS
[0088] Meta-analysis of patients with Chinese ancestry A summary of the study design and the clinical characteristics of the participants in all stages are shown in Fig. 1 and Table 1. In stage 1 , 684 patients with type 2 diabetes and 955 controls were genotyped. No population stratification were detected between case and control individuals in multidimensional scaling analysis for all GWAS (Fig. 7). Meta-analysis was implemented to combine the individual association results for 2,925,090 imputed and genotyped SNPs (under additive genetic models) available in all three GWAS using the inverse-variance approach for random effect models.
[0089] In the stage 1 meta-analysis of three Chinese GWAS, 44 SNPs within five loci were prioritised for follow-up (Fig. 2 and Table 8). No substantial change was observed in the stage 1 results after adjusting either for λ. s (1.01— 1.04 in individual cohorts) or the first principle component in the meta-analysis, reflecting that the results were not likely to be due to population stratification (Fig. 8 and Table 9).
[0090] Of the five loci identified in stage 1 , CDKN2A/B has previously been reported to be strongly associated with type 2 diabetes. In line with the inventors' previous findings, two SNPs in CDKN2A/B showing strong signals for type 2 diabetes in the present study were in high LD (r2~0.8) with rs 1081 1661 , which is well-replicated in most populations. After eliminating the signal of CDKN2A/B and redundant markers, 13 top and proxy SNPs among the remaining 42 SNPs in four regions were taked forward to stage 2, de novo replication, in two independent Chinese case-control cohorts (Table 6). Genotypes were successfully obtained for 1 1 SNPs in Hong Kong replication 1 cohort with 5,366 cases and 2,474 controls, and Shanghai replication 1 cohort with 4,035 cases and 3,964 controls to proceed for subsequent analysis (Table 7). Of these, rsl0229583 and rs2737250, located on chromosomes 7 and 8, respectively, gave p <4.5x l0" (threshold of significance after Bonferroni correction) with the same directions of association as the original signals (Table 2). These two SNPs were genotyped in 1 ,518 additional samples from 426 families of Han Chinese descent (325 cases and 368 controls from 178 Hong Kong families, and 657 cases and 168 controls from 248 Shanghai families).
Although no significant association was detected in either family study using ALR, all were in the concordant direction for rs 10229583 (Table 10). Taken together, the overall observed association for type 2 diabetes with rs 10229583 by combining all studies from Chinese ancestry in stages 1 and 2 yielded an OR (95% CI) of 1.18 (1.1 1 , 1.25) with a corresponding p=2.6x 10"8 (Table 3). For another variant taken to genotyping in family samples, rs2737250, meta-analysis of GWAS and de novo genotyping in the Hong Kong and Shanghai case-control samples revealed OR 1.10 (1.05, 1.15) with a corresponding /?=7.05x l0~5 using a fixed effect model (p for heterogeneity test=0.0012, =0.852), with OR 1.16 (1.01 , 1.33), /?=0.0299 by random effect model. However, genotyping of the variant in the Hong Kong and Shanghai family samples suggested an association in the opposite direction (Table 10). [0091] Meta-analysis in East Asian and other populations To further validate the association of rsl0229583 with type 2 diabetes, in an silico replication of rsl0229583 was conducted in five East Asian GWAS (one Japanese, two Korean, one Singapore Chinese and one Han Chinese study), and three non-East Asian GWAS (Singapore Indian, Singapore Malaysian and the DIAGRAM Consortium). Meta-analysis for the East Asian populations (£>=2.3x l0~10) gave an OR (95% CI) of 1.14 (1.09, 1.19). Among non-East Asian populations, replication of the association was observed in participants of European descent from the DIAGRAM Consortium (p=8.6x l0 3), with OR 1.06 and 95% CI (1.02, 1.12) (Table 3 and Figs 3 and 4).
[0092] Impact of rs 10229583 on clinical traits and course of disease The inventors next investigated the associations of rsl0229583 with the AAD of type 2 diabetes and quantitative metabolic traits related to type 2 diabetes. Among all the patients with type 2 diabetes, individuals who carried the common, type 2 diabetes risk allele (G) were concordantly and significantly younger at the time of diagnosis in both Hong Kong and Shanghai, and the meta- analysis showed that presence of the risk variant had a significant association with younger AAD
Figure imgf000028_0001
which remained unchanged following adjustment for sex and BMI (Table 1 1). A nominal association of the G-alleles of rsl0229583 with beta cell function was also observed as assessed by HOMA- healthy Hong Kong adolescents, a reduced Stumvo
and increased fasting plasma glucose (FPG) level (
Figure imgf000028_0002
healthy Shanghai adults (Fig. 5).
[0093] Functional implication of the identified locus rs 10229583 In order to evaluate the functional implication of our identified variant, an extensive bioinformatics analysis was performed. Consistent with its observed effect on pancreatic beta cell function, the gene region of our locus has been identified as one of the islet-selective clusters of open regulatory elements using a formaldehyde-assisted isolation of regulatory elements coupled with high-throughput sequencing in human pancreatic islets (Gaulton et al. (2010) Nat Genet 42: 255-259). In addition, the variant and its tagging SNPs lie within an area near PAX4 and SND1 , which is enriched with DNase I hypersenstitive sites, histone H3 lysine modifications and CCCTC factor binding in human islets (Fig. 9) (Stitzel et al. (2010) Cell Metab 12: 443-455).
[0094] The relationship of rs 10229583 with eQTLs in adipose tissue and other tissues was next investigated in available datasets. The variant rs 1440971 , a proxy of our associated SNP
(MAF-0.1 , r =0.8 and D'=l , to rsl0229583), was significantly associated with the level of expression of GRM8, ARF5 and PAX4 in lymphoblastoid cells in the GenCord Project, although this did not correlate with the eQTL peak (Dimas et al. (2009) Science 325: 1246-1250).
Analysis of all eQTLs associated with rsl0229583, or its close proxy, rsl440971 , was performed using data from the Multiple Tissue Human Expression Resource (MuTHER) Consortium (Nica et al. (201 1) PLoS Genet 7: el002003). Of note, eQTL data were only available for PAX4 in adipose tissue, but not LCLs or skin, for which no expression data were available from MuTHER. There was a nominal association (p<0.05) between the variant and expression of C7orf54 and ARF5 in LCLs, and C7orf68 in adipose tissue. The r2 between the GWAS SNP and the peak eQTL SNPs ranged between 0.56 and 1.
[0095] Complex diseases such as type 2 diabetes are caused by a combination of alterations, and each genomic perturbation or alteration can potentially impact on thousands of genes (Kim et al. (201 1) PLoS Comput Biol 7: el 001095; Barrenas et al. (2012) Genome Biol 13: R46). Nevertheless, functionally important genes often organise into the same pathway of functional grouping. Therefore, one can overlay the alterations on a gene network that was built using highly confident gene-gene relationships (Lee et al. (201 1) Genome Res 21 : 1 109-1 121).
Interactions between these genes were identified, with additional interaction with other key pancreatic transcription factors such as NEUGR03. Taken together, it is speculated that ARF5, GCC1, SND1 and PAX4 may function together with NEUGR03 in the same network for pancreatic islet development.
[0096] Heterogeneity of effect in Chinese vs other ethnic groups To investigate why the novel loci identified in the present study had not been detected in previous GWAS performed in other populations, the heterogeneity of effect between East Asians and Europeans was examined. There was no evidence of heterogeneity of effect between Chinese, Korean and Japanese populations, but significant heterogeneity of effect was seen between Han Chinese and individuals of European descent in the DIAGRAM Consortium, as well as between Chinese, Malaysians and Indians (Table 12). [0097] To test for the variation of LD structure between Chinese and other populations, the targeted varLD approach was implemented to examine the pattern of r2 between every pair of SNPs within the lOOkb region centred on our index SNP rsl0229583. (website:
statgen.nus.edu.sg/~SGVP/software/varld.html) (Ong and Teo Bioinformatics 2010; 26(9): 1269-70). This region shows highly significant evidence of LD variation between Chinese, European (Monte Carlo [MC] /?=0.0018), and African (MC p=0.0003) individuals, but nominal evidence of variations between Chinese and Japanese (MC /?=0.0107) (Table 13 and Figs 10 and 1 1). Discrepancies were also observed in allele frequency of rs 10229583 between East Asians, Europeans and Africans (Table 13).
DISCUSSION
[0098] In this study the inventors report a meta-analysis of GWAS for type 2 diabetes in a Chinese population, and have identified a novel diabetes-associated locus. Furthermore, the inventors replicated the association in additional East Asian samples and found an association in samples of European descent. In addition to the multiethnic samples used, this study also benefits from a detailed phenotyping of the Chinese samples, which allowed additional analyses of the effect of the risk variant on clinical traits and the course of disease to be carried out. [0099] Type 2 diabetes in Asians is characterised by an earlier AAD, strong family history and evidence of impaired beta cell function (Chan et al. (2009) JAMA 301 : 2129-2140; and
Ramachandran et al. (2010) Lancet 375: 408-418). In a recent nationwide study conducted in China, the prevalence of diabetes was 3.2% among persons aged 20-39 years, and 1 1.5% among adults aged 40-59 (Yang et al. (2010) N Engl J Med 362: 1090-1 101). The risk variant identified here, rsl0229583, was associated with earlier AAD in both the Hong Kong and Shanghai samples, highlighting its potential contribution to young-onset diabetes in the Chinese population. Healthy adults and adolescents who carry the risk variant were found to have elevated fasting glucose and impaired beta cell function, respectively. [0100] The novel locus for type 2 diabetes identified, rsl 0229583, is located downstream of the ARF5 and PAX4 genes in 7q32, and upstream of SND1. PAX4, a member of the paired box family of transcription factors, plays a critical role in pancreatic beta cell formation during fetal development (Bran et al. (2004) J Cell Biol 167: 1 123-1 135; Li et al. (2006) Leuk Res 30: 1547- 1553) and is therefore a very strong candidate for the implicated gene. The gene region lies within an area of islet-specific cluster of open chromatin sites and may therefore act in cis with local chromatin and regulatory changes (DerSimonian and Laird (1986) Control Clin Trials 7: 177-188). PAX4 is expressed in early pancreatic endocrine cells, but expression is later restricted to beta cells and it is not expressed in mature pancreas (Habener et al. (2005) Endocrinology 146: 1025-1034). In pancreatic endocrine cells, PAX4 represses ghrelin and glucagon expression, and can induce the expression of PDX1, a key transcription factor for islet development. Targeted disruption of PAX4 in mice was found to lead to reduced beta cell mass at birth (Wang et al. (2004) Dev Biol 266: 178-189).
[0101] Several human studies have implicated PAX4 in the pathogenesis of diabetes (Mauvais- Jarvis et al. (2004) Hum Mol Genet 13: 3151-3159; Shimajiri et al. (2001) Diabetes 50: 2864- 2869). In one report, a missense mutation (R121 W) was identified in six heterozygous patients and one homozygous patient out of 200 unrelated Japanese patients with type 2 diabetes
(Shimajiri et al. (2001) Diabetes 50: 2864-2869). For example, Japanese patients carrying PAX4 mutations have severe defects in first-phase insulin secretion (Tokuyama et al. (2006)
Metabolism 55: 213-216). Mutations in PAX4 may lead to rare monogenic forms of young-onset diabetes (Plengvidhya et al. (2007) J Clin Endocrinol Metab 92: 2821-2826). Common variants in several other MODY genes, namely, HNF4a, HNFla and TCF2, have been identified as susceptibility loci for type 2 diabetes (McCarthy MI (2010) N Engl J Med 363: 2339-2350). [0102] The finding of this study is consistent with other studies that have highlighted the important role of genes implicated in pancreatic development in the pathogenesis of type 2 diabetes. In a previous study, a risk variant at HNF4a has been found to be associated with increased risk of type 2 diabetes, and carriers of the risk allele have impaired beta cell function (Silander et al. (2004) Diabetes 53: 1 141-1 149). The MAF of the R121W PAX4 mutation was 1 % in Asians, and the mutation is in low LD with rs 10229583. It is possible that both rare mutations and common variation within the same gene confer risk towards type 2 diabetes independently. The common variant here identified, rsl0229583, may be associated with altered gene expression, while the other rare non-synonymous mutations lead to impaired gene function. For example, while common non-coding variants in MTNR1B increase type 2 diabetes risk with a modest effect, large-scale resequencing has identified rare loss-of-function MTNR1B variants that significantly contribute towards type 2 diabetes risk (Bonne fond et al. (2012) Nat Genet 44: 297-301). Some regulatory elements harbouring type 2 diabetes-associated loci have recently been found to exhibit allele-specific differences in activity, providing evidence supporting the functional role of non-coding common variants identified through GWAS (DerSimonian and Laird (1986) Control Clin Trials 7: 177-188; Gaulton et al. (2010) Nat Genet 42: 255-259).
[0103] The recent East Asian meta-analysis comprising eight type 2 diabetes GWAS identified a locus on chromosome 7 near GRIP and GCC1-PAX4 to be associated with type 2 diabetes. The protein encoded by GCC1 may play a role in transmembrane transport (Luke et al. (2005) Biochem J 388: 835-841). The variant identified from the East Asian study, rs6467136, appears to be independent of our signal, with r2=0.044 in our Chinese samples (Fig. 1 1). Furthermore, no change was found in the effect size of rsl0229583 after conditioning on rs6467136 (OR [95% CI]=1.20[ 1.1 1 , 1.29], p=4.6x 10"6 vs OR [95% CI]= 1.19 [ 1.10, 1.29), p=\ .6* 10"5, before and after the conditional analysis in 9,886 Chinese samples). Likewise, rs6467136 had little change in effect after conditioning on rsl0229583 (OR [95% CI]=1.09 [1.02, 1.17], p=0.0125 before; OR [95% CI]=1.07 [0.99, 1.15], p=0.0729 after). In the recent analysis from the DIAGRAM Consortium, rs231362 near KCNQl was identified to be associated with type 2 diabetes. This signal is independent of the original signal identified in the Japanese population as revealed by conditional analysis. Consistent with the evidence observed for KCNQl, this finding highlighted that multiple common genetic variations within the same gene region may independently contribute to disease risk (Voight et al. (2010) Nat Genet 42: 579-589; Yasuda et al. (2008) Nat Genet 40: 1092-1097). It will be worthwhile undertaking a further investigation of this region to search for population-specific and/or disease causal variants in different ethnic groups by fine- mapping as well as transethnic mapping.
[0104] The other genes in the region of this newly identified sequence variant are also potential candidate genes for diabetes. ARF5 belongs to a family of guanine nucleotide -binding proteins that have been shown to play a role in vesicular trafficking and as activators of phospholipase D (Lebeda and Haun (1999) Gene 237: 209-214). Islet expression of ARF5 was found to be induced threefold in rats receiving a high-carbohydrate diet (Song et al. (2001) Diabetes 50: 2053-2060). The nearby SND1 gene, also known as the plOO transcription co- activator, is a member of the micronuclease family and plays a key role in transcription and splicing. The pi 00 transcriptional co-activator is present in endocrine cells and tissues, including the pancreas of cattle (Broadhurst et al. (2005) Biochim Biophys Acta 1681 : 126-133).
[0105] Among the type 2 diabetes loci first identified in non-European populations other than KCNQl, few have consistently been found to show a significant association in studies of individuals of European descent (Voight et al. (2010) Nat Genet 42: 579-589; Kooner et al.
(201 1) Nat Genet 43: 984-989; Cho et al. (201 1) Nat Genet 44:67-72; McCarthy MI (2010) N Engl J Med 363: 2339-2350). The diabetes gene variant here identified, rs 10229583, also showed a significant association in Europeans in the DIAGRAM Consortium, with a smaller effect size compared with East Asian individuals (p=0.0024 by Cochran's Q statistics,
7 =0.8913). Interestingly, rare PAX4 mutations were first identified in Asian MODY probands (Shimajiri et al. (2001) Diabetes 50: 2864-2869; Tokuyama et al. (2006) Metabolism 55: 213- 216), but seldom found in those of European descent (Dupont et al. (1999) Diabetologia 42: 480- 484; Dusatkova et al. (2010) Diabet Med 27: 1459-1460). This suggests that PAX4, like KCNQl, may be particularly relevant for the pathogenesis of type 2 diabetes in East Asians individuals. Interestingly, rs 10229583 is also in strong LD with a region spanning the neighbouring SND1 gene (Fig. 3). Further resequencing and transethnic mapping should help to identify the causal gene variant for type 2 diabetes within this region.
[0106] The novel locus the present inventors identified in Chinese individuals with type 2 diabetes has not been detected in previous GWAS performed in mainly individuals of European descent. A highly significant LD variation was noted between Chinese and European individuals in the region surrounding our identified variant. There is also significant variation in allele frequencies in Chinese compared with Europeans, as well as between Chinese and African individuals. This ethnic difference in LD pattern and risk allele frequency may lead to a differential impact in different populations and warrants further investigation by resequencing.
[0107] This study has several limitations. The sample size of the GWAS was modest, resulting in limited power to identify genetic variants with small effect sizes. The study was limited to Southern Han Chinese, although the consistent replication seen in other East Asian population suggests that the findings may be applicable to other populations of Chinese descent.
[0108] In summary, rs 10229583 near PAX4 has been identified as a novel locus for type 2 diabetes in Chinese and other populations, providing new insights into the pathogenesis of type 2 diabetes.
Example 2: Sequence Variation Near PAX4 Linked to Early Onset of Coronary Heart Disease Among Type 2 Diabetes Patients
[0109] Type 2 diabetes (T2D) patients have a 2-4 fold increased risk of coronary heart disease (CHD) compared with the general population. A recent genome-wide association study conducted by the present inventors in the Chinese populations identified association between a sequence variant at 7q32 near paired box 4 (PAX4) and T2D, which was confirmed in other East Asian populations. See details in Example 1.
[0110] This study aimed to investigate the association of this novel 7q32 variant and CHD risk in an 8 -year prospective cohort of Chinese patients with T2D. The 7q32 variant was geno typed in 5264 T2D patients [age mean ± SD = 56.3 ± 13.3 years, % males = 45.0] without history of CHD at baseline. Associations of the variant under multiplicative, dominant and recessive genetic models with new onset of CHD were examined by Cox proportional hazard models with adjustment of conventional risk factors at baseline. During the mean follow-up period of 6.9 ± 6.7 years, 395 patients (7.5%) developed incident CHD. Subjects who carried the common, T2D risk allele (G) of the 7q32 variant have a higher risk for CHD (hazard ratio [HR] (95% CI) = 1.43 (1.15 - 1.79), P = 1.4 10"3 under additive model; HR (95% CI) = 1.56 (1.22 - 2.00), P = 3.7 x 10"4 under recessive model) after adjusting for sex, age, duration of diabetes, smoking status and HbAic. Further adjustments of conventional risk factors including body mass index, lipid profiles, blood pressure, history of retinopathy and neuropathy, urinary albumin excretion rate, glomerular filtration rate and use of medications did not change the association. [0111] In summary, the findings of this study indicate that the novel T2D-associated 7q32 variant near PAX4 is not only associated with T2DM, but also identifies T2DM subjects with increased risk of CHD.
ELECTRONIC SUPPLEMENTARY MATERIAL (ESM)
STUDY SAMPLES AND TYPE 2 DIABETES DIAGNOSIS
Stage 1: discovery samples
Chinese University of Hong Kong
[0112] The study design, ascertainment, inclusion criteria and pheno typing procedures of participants included in this study were previously described (Ng et ah (2008) Diabetes 57:
2226-2233). All participants were of southern Han Chinese ancestry residing in Hong Kong.
Type 2 diabetes was diagnosed according to the 1998 World Health Organization (WHO) criteria. Patients with classic type 1 diabetes with acute ketotic presentation or continuous requirement of insulin within 1 year of diagnosis were excluded. Written informed consent was obtained from all participants. This study was approved by the Clinical Research Ethics Committee of the Chinese University of Hong Kong.
[0113] In the first stage discovery cohort (stage 1), genome -wide scanning was performed in 202 Hong Kong Chinese individuals (102 type 2 diabetes patients and 100 healthy controls) (Hong Kong GWAS 1 cohort). 102 type 2 diabetes cases were selected with young-onset diabetes diagnosed at age <40 years, positive family history and overweight, and 100 controls were selected using the criteria of 1) no past diagnostic history of type 2 diabetes, impaired fasting glucose (IFG) or impaired glucose tolerance (IGT); 2) without family history of type 2 diabetes; and 3) with BMI≤ 25 kg/m2 and waist circumference≤ 90 cm and 80 cm for men and women, respectively.
[0114] In addition, 1 ,068 Hong Kong Chinese individuals were genome-scanned (400 type 2 diabetes patients from the Hong Kong Diabetes Registry and 668 non-diabetic controls) (Hong Kong GWAS 2 cohort). The 668 diseased controls were individuals aged > 16 years old with diseases other than type 2 diabetes that included 457 epilepsy cases, 11 1 eczema cases and 100 healthy individuals without hypertension (recruited from the control arm of a hypertension study).
Shanghai Jiao Tong University Diabetes Study (SJTUDS)
[0115] All samples were recruited from Shanghai Diabetes Institute of Shanghai Jiao Tong
University. The genome -wide scan was performed in 394 samples, including 197 type 2 diabetes patients and 197 normal glucose regulation controls. The type 2 diabetes patients were probands of diabetic pedigrees with fasting plasma glucose > 7.0 mmol/L and/or 2-h post plasma glucose > 1 1.1 mmol/L who were diagnosed before 40 years old. Type 1 diabetes and mitochondrial diabetes were excluded based on clinical, immunological and genetic criteria. The controls were individuals with normal glucose regulation with fasting plasma glucose < 6.1 mmol/L and 2-h plasma glucose < 7.8 mmol/L as assessed by standard 75g OGTTs, negative diabetic family history, aged over 50 years old and with a BMI below 23kg/m2.
Clinical studies
[0116] All Hong Kong and Shanghai Chinese individuals underwent detailed clinical investigation. For the Hong Kong study, fasting blood samples of all studied participants were collected for fasting plasma glucose (FPG) and fasting plasma insulin (FPI). For the Shanghai study, fasting blood samples were obtained from cases, whilst controls had blood collected at baseline, and 120 min during a 75g oral glucose tolerance test (OGTT). Detailed clinical information, including age at diagnosis and presence of diabetic complications, were
documented in all cases as described. Homeostasis model assessment of insulin resistance
(HOMA-IR) was calculated as (FPIxFPG) ÷ 22.5, and homeostasis model assessment of beta-cell function (ΗΟΜΑ-β) was calculated as FPIx20 ÷ (FPG-3.5) (Matthews et al. (1985) Diabetologia 28: 412-419). Stumvoll indices for beta-cell function were calculated for Shanghai controls which underwent OGTT with measurement of insulin levels (Stumvoll et al. (2000) Diabetes Care 23: 295-301).
Stage 2: de novo replication samples
Hong Kong replication 1 (case-control cohort)
[0117] The case cohort consisted of 5,366 unrelated type 2 diabetes patients (mean age 56.7 ± 13.4 years, 45.1% male, mean duration of T2D 6.6 ± 6.9 years) selected from the Hong Kong Diabetes Registry (HKDR). The control cohort consisted of 2474 individuals ascertained from 3 sources: a) 985 adolescents from a community-based school survey of cardiovascular risk factors (mean age 15.5 ± 1.9 years, 44.2% male) (Ng et al. (2010) J Clin Endocrinol Metab 95: 2418- 2425); b) 513 hospital staff and adult volunteers participating in a community-based health screening program (mean age 42.0 ± 10.4 years, 47% male) (Ng et al. (2010) J Clin Endocrinol Metab 95: 2418-2425) and c) 976 healthy elderly individuals free of diabetes selected from 4,000 elderly individuals recruited from the community (mean age 72.3 ± 5.3 years, 51.4% male) (Tang et al. (2010) Bone 46: 543-550).
Shanghai replication 1 (case-control cohort)
[0118] All participants were of Chinese Han ancestry and resided in Shanghai and the nearby area. The type 2 diabetes patients (n=4,036) were selected from the Shanghai Diabetes Institute Inpatient Database (SHDIID), which recruited participants from inpatients in the Department of Endocrinology and Metabolism, Shanghai Jiao Tong University Affiliated Sixth People's Hospital beginning in 2001. Type 1 diabetes and mitochondrial diabetes were excluded based on clinical, immunological and genetic criteria. The controls (n=3,964) were selected from
Shanghai Diabetes Study (SHDS) I and II, which are community-based surveys of diabetes performed in 1998-2001 (SHDS I) and 2007-2008 (SHDS II). The controls had fasting plasma glucose <6.1 mmol/L and 2-h plasma glucose < 7.8 mmol/L as assessed by standard 75g OGTTs, and had no family history of diabetes mellitus.
Hong Kong replication 2 (family cohort)
[0119] The study design, ascertainment, inclusion criteria and phenotyping of the Hong Kong Family Diabetes Study (HKFDS) have been described elsewhere (Ng et al. (2004) Diabetes 53: 1609-1613). Briefly, 325 individuals with type 2 diabetes (mean age 48.0 ± 14.4, 40.6% male) and 368 control subjects (mean age 37.0 ± 13.6, 41.0% male) were selected from 178 families consisting of siblings, parents, spouses, and offspring (> 16 years) ascertained through a proband with type 2 diabetes. Patients with clinical or autoimmune type 1 diabetes and families with known maturity-onset diabetes of the young or mitochondrial DNA nucleotide 3243 A>G mutations were excluded.
Shanghai replication 2 (family cohort)
[0120] 248 type 2 diabetes pedigrees were recruited with 657 type 2 diabetes patients and 168 individuals with normal glucose regulation. Pedigrees with any type 1 diabetic patient or mitochondria diabetic patient were excluded.
Stages 3 and 4: in silico replication samples
BioBank Japan (BBJ)
[0121] The individuals were recruited from several medical institutes in Japan, including Fukujuji Hospital, Iizuka Hospital, Iwate Medical University School of Medicine, National Hospital Organization Osaka National Hospital, Nihon University, Nippon Medical School, Osaka Medical Center for Cancer and Cardiovascular Diseases, The Cancer Institute Hospital of Japanese Foundation for Cancer Research, Tokushukai Hospitals and Tokyo Metropolitan Geriatric Hospital. Type 2 diabetes cases were selected from individuals registered as having type 2 diabetes. Diabetes was originally diagnosed according to the World Health Organization (WHO) criteria, type 2 diabetes was clinically defined as disease with a gradual adult onset. Individuals who tested positive for antibodies to glutamic acid decarboxylase (GAD) and those diagnosed with a mitochondrial disease or MODY were not included in the case group. Controls were individuals registered as individuals not having type 2 diabetes but with diseases other than type 2 diabetes, comprised of 13 distinct diseases, or healthy volunteers. Individuals who had been analyzed in the previous report (Yamauchi et al., 2010 Nat Genet 42(10):864-868) were excluded from the present study. Altogether, 4,878 individuals with type 2 diabetes (case 1 , age, 65.8 ±10.0 years; BMI, 24.1 ± 3.8 kg/m2; (all values are expressed as mean ± s.d.)) and
3,345controls (control 1 , age, 52.5 ± 15.2 years; BMI, 22.5 ± 3.8 kg/m2; (all values are expressed as mean ± s.d.)) were genotyped. A total of 7,541 individuals belonging to the Hondo cluster (4,470 cases and 3,071 controls) were selected. Samples were directly genotyped using Illurnina HurnanHap610-Quad (type 2 diabetes patients) and 550K BeadChip (controls).
Korea Association Resource Study (KARE)
[0122] Two KARE study cohorts were established as part of the Korean Genome
Epidemiology Study (KoGES) in 2001 (Cho et al. (2009) Nat Genet 41 : 527-534). The sampling base for both cohorts was in the Kyung Gi-Do province, close to Seoul, the capital of the Republic of Korea. Both cohorts were designed to allow longitudinal prospective study and adopted the same investigational strategy. Participants have been examined every two years since baseline (2001). More than 260 traits have been extensively examined through
epidemiological surveys, physical examinations, and laboratory tests applied to cohort members. A total of 10,038 individuals from KARE study cohorts were genotyped with Affymetrix Genome- Wide Human SNP array 5.0. to undertake a large-scale GWA analysis for type 2 diabetes and numerous complex quantitative traits. Of them, 1 ,042 individuals were included as type 2 diabetes cases according to the following criteria: (1) treatment of type 2 diabetes, (2) fasting plasma glucose > 7 mmol/L or plasma glucose 2-h after ingestion of 75gm oral glucose load > 1 1.1 mmol/L and (3) age of disease onset > 40 years. The inclusion criteria of nondiabetic control individuals (n = 2,943) were as follows: (1) no history of diabetes and (2) fasting plasma glucose < 5.6 mmol/L and plasma glucose 2-h after ingestion of 75gm oral glucose load < 7.8 mmol/L at both baseline and follow up studies.
Singapore Chinese. Malaysian and Indian populations (case-control cohorts)
[0123] The Singapore case-control study contained individuals from three sources: 1) 1998 Singapore National Health Survey (NHS98); 2) Singapore Malay Eye Study (SiMES); and 3) Singapore Diabetes Cohorts Study (SDCS) (Tan et ah (2010) J Clin Endocrinol Metab 95: 390- 397). In the NHS98 cohort, individuals with fasting plasma glucose (FPG) <6.0 mmol/L and 2 hour post-challenge glucose (2HPG) <7.0 mmol/L were defined as NGT. Individuals with FPG >6.0 and <7.0 mmol/L, and 2HPG > 7.0 and <7.8 mmol/L, were defined as having IFG.
Individuals with FPG >7.0 mmol/L, and 2HPG >7.8 and <1 1.1 mmol/L, were defined as IGT. A total of 838 IFG/IGT individuals were excluded, leaving 3,032 NGT controls (2,196 Chinese, 472 Malays, and 364 Indians) available for selection. Individuals from the NHS98 and SDCS cohorts with: 1) a reported history of type 2 diabetes; 2) FPG >7.0 mmol/L; or 3) 2HPG>1 1.1 mmol/L were defined as cases. 453 NHS98 cases (224 Chinese, 1 13 Malays, and 1 16 Indians) and 1 ,703 SDCS cases (1 ,317 Chinese, 256 Malays, and 130 Indians) were available for selection. In the SiMES cohort, individuals with non-fasting PG< 1 1.1 mmol/L and HbAic <6.1 % (2 SD above the mean for the nondiabetic population) were defined as controls (N= 1 ,785).
Individuals with a reported history of type 2 diabetes or non-fasting PG level >1 1.1 mmol/L were defined as cases (N =707). From these three sources, the following was included: 1) 2,010 type 2 diabetes cases and 1 ,945 NGT controls of Chinese ancestry, 2) 794 type 2 diabetes cases and 1 ,240 NGT controls of Malaysian ancestry, 3) and 977 type 2 diabetes cases and 1 ,169 NGT controls of Indian ancestry, for analysis. Chinese Hans (case-control cohort)
[0124] This study included 1999 type 2 diabetic cases and 1976 nondiabetic controls drawn from the Nutrition and Health of Aging Population in China (NHAPC) study (312 cases and 815 controls), the Gut Microbiota and Obesity Study (GTOS) (82 cases and 163 controls), the Fudan- Huashan Study (FHS) (807 cases and 339 controls), and the Beijing Diabetes Survey (798 cases and 659 controls). Details on the study have been published previously (Li et al., 2013 Diabetes 62(l):291 -298). All participants were unrelated Chinese Hans from Beijing and Shanghai. Type 2 diabetic cases were identified as those with previously diagnosed type 2 diabetes and current use of antidiabetic treatment or who meet the following criteria: 1) 30< age <70, 2) fasting plasma glucose >7.0 mmol/1, 3) 2-h postprandial plasma glucose >1 1.1 mmol/1 in a standard 75 g oral glucose tolerance test (OGTT) or plasma HbAlc >6.5%. The nondiabetic controls were selected according to the following criteria: 1) age >30, 2) no past history of diagnosis of diabetes and no family history of diabetes, 3) fasting glucose <5.6 mmol/1, 4) 2-h OGTT <7.8 mmol/1 and/or HbAl c content <5.6%. The studies were approved by local ethnic committees of each participating institution, and written informed consents were obtained from all participants.
[0125] The DNA samples were genotyped using the Illumina Human660W-Quad BeadChip (Illumina, Inc., San Diego, CA, USA), and the genotypes were called using the Illumina GenCall algorithm. Some of samples were excluded if their genotype call rates < 97%, excessive heterozygosity, gender mismatches between the reported and genetically inferred gender or duplicates among other samples. Principle component analysis was used to assess population structure of the samples and detected outliers along the first two eigenvectors which were excluded from further analyses. SNPs with genotype call rate < 95%, MAF < 0.5% or deviation from Hardy- Weinberg equilibrium (p < 10"6) in control groups were also excluded. After all these quality control processes, 495,686 SNPs and 3,712 samples, including 1 ,873 type 2 diabetic cases (861 male and 1012 female) and 1 ,839 controls (803 male and 1036 female), remained for association analyses. DIAGRAM Consortium (case-control cohort)
[0126] DIAGRAM+ study comprised 8,130 type 2 diabetes cases and 38,987 controls from eight type 2 diabetes GWAS of European descent, including the Wellcome Trust Case Control Consortium (WTCCC), Diabetes Genetics Initiative (DGI) and Finland-US Investigation of NIDDM genetics (FUSION) scans (the individuals of a previous joint analysis), with those from scans performed by deCODE genetics, the Diabetes Gene Discovery Group, the Cooperative Health Research in the Region of Augsburg group (KORAgen), the Rotterdam study and the European Special Population Research Network (EUROSPAN) (for details of sample characteristics, see Supplementary Table 2 in Voight et ah (2010) Nat Genet 42: 579-589). Additional information on methods
Imputation
[0127] Before imputation, the SNP ID (rs number) was standardized according to dbSNP build 129, and their physical positions were standardized according to build 36. SNPs were further excluded sequentially if: 1) their polymorphisms were A T or C/G; 2) absent from dbSNP build 129; 3) genotyped in only case or only control cohorts; 4) absent from the 1000 Genomes reference panel for CHB+JPT (March 2010 release of pilot project 1). For each sample set in stage 1 , all SNPs were aligned to the positive strand and imputed (via the MLE approach) using the MACH 1.0 software (Li et al. (2010) Genet Epidemiol 34: 816-834). Genotypes were imputed for autosomal SNPs that were present in the March 2010 release of phased 1000
Genomes genotype data from 60 CHB+JPT founders (Nature 467: 1061-1073, 2010), but were not present in the genome -wide chip or did not pass direct genotyping QC. Cases and controls were merged into a single cohort for imputation based on 440,194, 435,953 and 274,752 quality autosomal SNPs in Hong Kong GWAS 1, Hong Kong GWAS 2 and Shanghai GWAS case- control cohorts, respectively. For the Hong Kong GWAS 1 cohort, one-step imputation was applied. For the Hong Kong GWAS 2 and Shanghai GWAS cohorts, two-step imputation was used to improve imputation efficiency, by randomly selecting 100 cases and 100 controls for model parameter estimation first before imputation.
[0128] A total of 3,356,999 SNPs in Hong Kong GWAS 1 , 3,355,668 SNPs in Hong Kong GWAS 2 and 3,087,246 SNPs in Shanghai GWAS that passed the quality control filters (with predicted Rsq>0.5 and and MAF>0.1) after imputation were included in the association analyses for individual cohorts. Rsq is an imputation accuracy measure to estimate the squared correlation between the true and imputed genotypes.
Genomic control
[0129] Genomic control (GC) was applied to correct for relatedness of the individuals and adjust for potential population stratification (Devlin and Roeder (1999) Biometrics 55: 997- 1004). The inflation factor λ was estimated by taking the median of the distribution of the χ2 statistic from all quality SNPs in association test, and then divide by the median of the expected χ2 distribution. The p values were calculated corrected for genomic control by dividing the observed χ statistic by λ. In this study, they were adjusted for GC in two levels. Firstly, each individual study was corrected for λ separately in directly genotyped and imputed SNPs. Then they were further adjusted for GC on the meta-analysis results. eQTL analysis
[0130] Gene expression analysis was performed using data available from the MuTHER (Multiple Tissue Human Expression Resource) Consortium and GenCord Projects. The
MuTHER resource (website: muther.ac.uk) includes lymphoblastoid cell lines (LCLs), skin and adipose tissue derived simultaneously from a subset of well-phenotyped healthy female twins from the TwinsUK adult registry ( ica et al. (201 1) PLoS Genet 7: el002003). Whole-genome expression profiling of the samples, each with either two or three technical replicates, were performed using the Illumina Human HT-12 V3 BeadChips (Illumina Inc) according to the protocol supplied by the manufacturer. Log2 transformed expression signals were normalized separately per tissue as follows: quantile normalization was performed across technical replicates of each individual followed by quantile normalization across all individuals. Genotyping was done with a combination of Illumina arrays (HumanHap300, HumanHap610Q, lMDuo and 1.2MDuo). Untyped HapMap2 SNPs were imputed using the IMPUTE software package (v2). The number of adipose samples with genotypes and expression values is per tissue was 778 for LCLs, 667 in skin and 776 in adipose. Association between rsl 0229583 (MAF > 5%, IMPUTE info > 0.8) and the normalized mRNA expression values of genes within 1MB of the SNP were performed with the GenABEL/ProbABEL packages using the polygenic linear model incorporating a kinship matrix in GenABEL followed by the ProbABEL mmscore test with imputed genotypes. Age and experimental batch were included as cofactors. A multiple-testing correction was applied to the cz's-association results. Genome-wide FDR of 1 % for multiple testing corresponds to ap value threshold ofp=5.1 x l0~5 in adipose tissue, 7.8x l0~5 in LCLs and 3.81 x 10 in skin. The eQTL data were also examined for SNPs which showed nominal association (p<0.05) with the expression of nearby genes in the different tissues from the MuTHER dataset. The LD between the MuTHER eQTL peaks within the dataset and rsl 0229583 were then examined.
[0131] The GenCord project is the study of association analysis for eQTL to nearby SNPs in three cell types (primary fibroblasts, lymphoblastoid cells and T cells) from the umbilical cords of 75 individuals (Dimas et al. (2009) Science 325: 1246- 1250). The Java based application
Genevar (website: sanger.ac.uk/resources/software/genevar/) (Yang et al. (2010) Bioinformatics 26: 2474-2476) was used to retrieve the relevant association results of our variant and expression of genes within 1MB of the SNP from the GenCord project. The LD between the variant and the eQTL peak within the dataset was also examined.
Gene network analysis
[0132] All SNPs in strong LD with rsl 0229583 were identified using data extracted from
SNAP (website: broadinstitute.org/mpg/snap/ldsearch.php). Selection criteria were as follows: (1) 1000 Genome Pilot 1 ; (2) r2 limit > 0.8; (3) Population Panel: CHBJPT; and (4) Distance
Maximum 500kb. Twenty- four tagSNPs in strong LD to rsl 0229583 with r >0.8 were extracted. The maximum genomic distance to rsl 0229583 was 34kb. Experimentally validated gene-gene interactions (pathways, genetic interactions and physical interactions) were obtained from
GeneMania (Warde-Farley et ah (2010) Nucleic Acids Res 38: W214-220). Data was updated by GeneMania as of Feb 2012. The weight on each edge, as represented by the thickness of the edge, was computed by GeneMania and reflects the degree of confidence of the relationships between any gene pair within the network. Genes that are identified from the MuTHER eQTL analysis above were entered as prior knowledge, and were used to guide the gene-network building. Pathway information was obtained from Pathway Commons (Cerami et al. (201 1) Nucleic Acids Res 39: D685-690) and genetic interactions were obtained from BioGrid (Stark et al. (201 1) Nucleic Acids Res 39: D698-704).
[0133] All patents, patent applications, and other publications, including GenBank Accession Numbers, cited in this application are incorporated by reference in the entirety for all purposes.
Table 1. Clinical characteristics of the participants
N Age AAD Diabetes duration BMI FPG
Study Cohort
(Male %) (years) (year) (years) (kg/m2) (mmol/1)
;e 1 (genome scan)
HK1 Control 99 (36.4) 37.3 ± 10.2 - - 20.8 ±2.0 4.7 ±0.4
T2D patient 99 (40.4) 40.6 ±8.8 31.8 ±7.7 8.0 ±8.3 30.9 ±4.4 -
HK2 Diseased control 659 (48.7) 37.1 ± 17.0 - - 23.3 ±3.7 -
T2D patient 388 (49.5) 60.6 ± 10.8 51.1 ± 12.1 9.5 ±7.0 25.0 ±3.8 -
SH Control 197 (50.8) 66.4 ± 10.1 - - 20.6 ± 1.7 4.8 ±0.4
T2D patient 197 (57.9) 41.6 ± 10.4 34.5 ±4.8 7.3 ±8.5 23.8 ±4.1 -
;e 2 (de novo replication in Chinese)
HK1 Adolescent control 985 (44.2) 15.5 ±1.9 - - 22.7 ±5.4 4.9 ±0.4
Adults control 513 (47.0) 42.0 ± 10.4 - - 19.9 ±3.5 4.7 ±0.3
Elderly control 976 (51.4) 72.3 ±5.3 - - 23.2 ±3.2 -
T2D patient 5366 (45.1) 56.7 ± 13.4 48.8 ± 14.9 6.6 ±6.9 24.6 ±5.3 -
SHI Control 3964 (37.6) 51.3 ± 13.5 - - 23.6 ±3.2 5.0 ±0.5
T2D patient 4035 (52.0) 61.2 ± 12.1 54.2 ± 11.3 7.2 ±6.9 24.5 ±3.5 -
H Family 2 Control 368 (41.0) 37.0 ± 13.6 - - 24.0 ±4.1 4.9 ±0.4
T2D patient 325 (40.6) 48.0 ± 14.4 41.7 ± 13.1 6.3 ±7.6 25.9 ±4.4 -
SH Family 2 Control 168 (51.2) 62.8 ± 11.2 - - 23.7 ±3.5 4.8 ±0.6
T2D patient 657 (43.7) 54.6 ± 15.6 50.0 ± 14.2 4.9 ±7.3 23.9 ±3.5 -
Stage 3 (in silico replication in East Asians)
N Age AAD Diabetes duration BMI FPG
Study Cohort
(Male %) (years) (year) (years) (kg/m2) (mmol/1)
Japanese Control 3023 (54.5) 51.9 ± 15.2 22.4 ±3.7
T2D patient 4465 (68.0) 65.8 ± 10.0 56.5 ± 11.4 9.4 ±8.4 24.1 ±3.8
Korean 1 Control 2943 (46.0) 51.1 ±8.6 24.1 ±3.0 4.5 + 0.4
T2D patient 1042 (51.7) 56.4 ±8.6 25.5 ±3.3 7.0 ±2.6 Korean 2 Control 1305 (54.5) 65.2 ±2.6 23.9 ±3.0 5.0 ±0.5
T2D patient 1183 (46.5) 58.6 ±7.1 25.2 ±3.4 7.4 ±2.7
Singapore
Control 1006 (21.6) 47.7 ±11.1 - - 22.3 ±3.7 4.7 ±0.4 Chinese 1
T2D patient 1082 (37.2) 65.1 ±9.7 55.7 ± 12.0 - 25.3 ±3.9 -
Singapore
Control 939 (63.8) 46.7 ± 10.2 - - 22.8 ±3.4 4.7 ± 0.5 Chinese 2
T2D patient 928 (64.9) 63.7 ± 10.8 52.2 ± 14.4 25.4 ±3.8
Chinese Control 1839 (43.7) 54.1 ±9.2 - 24.00 ±3.18 5.04 ±0.35
T2D patient 1873 (46.0) 58.6 ±8.4 25.00 ±3.24 8.43 ±2.90
Stage 4 (in silico replication in non- East Asians)
Singapore
Control 1240 (52.0) 56.9 ± 11.4 - - 25.1 ±4.8 - Malaysian
T2D patient 794 (51.0) 62.3 ±9.90 54.4 ± 11.2 - 27.8 ±4.9 -
Singapore
Control 1169 (48.4) 55.7 ±9.7 - - 25.3 ±4.4 - Indians
T2D patient 977 (54.4) 60.7 ±9.9 51.4 ± 10.6 27.1 ±5.1
DIAGRAM+ Control 38987 (-) - - T2D patient 8130 (-)
Data are shown as N, percentage or mean±SD
T2D, type 2 diabetes
Table 2. Association results for type 2 diabetes (T2D) with 11 top and proxy SNPs in de novo replication stage in Chinese populations
Hong Kong replication 1 Shanghai replication 1
C„ombined
(5366 T2D vs 2474 controls) (4035 T2D vs 3964 controls)
Chrom Nearest Position ^mor^ Case Control OR Case Control OR OR / a
osome gene(s) (B36) ™J°r MAF MAF (95% CI) PsMitive MAF MAF (95% CI) ΛωΜνε (95% CI) (uncorrected)
allele
127034 1 14 1.16 3.7x10- 1.15
rsl0229583 7 PAX4 A G 0.847 0.83 0.0077 0.846 0.825 l.OxlO-5 0.6406 0.000
139 ( 1.03, 1.23) ( 1.08, 1.27) ( 1.08, 1.22)
116725 1.08 1.06
rs2721960 8 TRPS1 T/C 0.657 0.644 L05 0.1566 0.655 0.638 0.0277 0.0095 0.7067 0.000
904 ( 0.98, 1.14) ( 1.01, 1.15) ( 1.02, 1.12)
116731 1.09 1.08
rs2737250 8 TRPS1 G/A 0.631 0.62 L05 0.1807 0.641 0.621 0.0090 0.0045 0.4582 0.000
048 ( 0.98, 1.12) ( 1.02, 1.16) ( 1.02, 1.12)
COL13 713100 1.03 1.01
rs3858158 10 C/T 0.516 0.521 °-98 0.6000 0.569 0.561 0.3211 0.506
Al 56 ( 0.92, 1.05) ( 0.97, 1.10) ( 0.96, 1.05) 0.7408 0.3026
COL13 713102 0 99 1.04 1.02
rs2395272 10 A G 0.531 0.534 0.7680 0.594 0.584 0.2312 0.364
Al 61 ( 0.93, 1.06) ( 0.97, 1.11) ( 0.97, 1.06) 0.4589 0.3027
COL13 713110 1.05 1.01
rs57703465 10 T/C 0.654 0.662 °-96 0.3463 0.667 0.656 0.1502 0.6467 0.0976 0.765
Al 74 ( 0.89, 1.04) ( 0.98, 1.12) ( 0.96, 1.06)
120045 0.97 0.99
rs 11065441 12 P2RX7 C/T 0.728 0.724 l m 0.6224 0.728 0.733 0.4312 0.7756 0.3748 0.000
354 ( 0.94, 1.11) ( 0.91, 1.04) ( 0.94, 1.05)
120054 0.98 1.00
rs684201 12 P2RX7 A/G 0.73 0.726 02 0.5916 0.735 0.739 0.5609 0.9462 0.4332 0.000
726 ( 0.94, 1.10) ( 0.91, 1.05) ( 0.94, 1.05)
120064 0 97 0.98 0.98
rsll065450 12 P2RX7 A/C 0.682 0.688 0.4995 0.702 0.707 0.5520 0.3699 0.9111 0.000
040 ( 0.90, 1.05) ( 0.92, 1.05) ( 0.93, 1.03)
120078 0.61 0.99 1.00
rs208290 12 P2RX7 T/C 0.609 1 -01 0.7086 0.643 0.645 0.7950 0.9605 0.6472 0.000
439 2 ( 0.94, 1.09) ( 0.93, 1.06) ( 0.95, 1.05)
120081 0.72 1 03 0.98 1.00
rs!0849851 12 P2RX7 G/A 0.72 0.4079 0.737 0.741 0.5237 0.9308 0.3002 0.068
027 7 ( 0.95, 1.12) ( 0.91, 1.05) ( 0.95, 1.05)
Nearest Entrez genes within 250 kb
P, Pmeta and phet represent p values from logistic regression without any adjustment under the additive genetic model, meta-analysis under a fixed effect model
(uncorrected for multiple testing) and test of heterogeneity, respectively
5 ORs are reported with respect to the minor allele
Table 3. Association results for rsl 0229583 and type 2 diabetes (T2D)
Risk allele
N frequencies
Figure imgf000046_0001
(uncorrecte
Stage Cohort Adjustment T2D Control T2D Control OR (95% CI) d GC)
1. Discovery Hong Kong GWAS 1 Sex and age 99 99 0.879 0.818 1.48 (0.85, 2.59) 0.1645
Hong Kong GWAS 2 Sex and age 388 659 0.857 0.820 1.56 (1.14, 2.13) 0.0055
Shanghai GWAS None 197 197 0.873 0.777 1.92 (1.32, 2.79) 5.0x l0-4
Meta-analysis of GWAS 684 955 1.66 (1.33, 2.07) 7.7x l0-6 0.6455 0.000
2. de novo Hong Kong replication 1 None 5,366 2,474 0.847 0.831 1.13 (1.03, 1.24) 7.7x l0-3
replications in Hong
Kong and Shanghai Shanghai replication 1 None 4,035 3,964 0.846 0.825 1.17 (1.07, 1.27) 3.7x l0-4
Hong Kong family
replication 2 Sex and age 325 368 0.872 0.856 1.22 (0.85, 1.74) 0.2817
Shanghai family
replication 2 Sex and age 657 168 0.824 0.813 1.09 (0.80, 1.49) 0.5757
Replication in Chinese 10,383 6,974 1.15 ( 1.08, 1.22) l .Ox lO-5 0.6406 0.000 Meta-analysis of
Chinese 11,067 7,929 1.18 (1.1 1, 1.25) 2.6x l0-8 0.0839 0.596
3. In silico Japanese replication None 4,465 3,023 0.892 0.881 1.11 (1.01, 1.23) 0.0379
replications in East
Asians Korean replication 1 None 1,042 2,943 0.894 0.878 1.17 (0.99, 1.38) 0.0577
Korean replication 2 None 1,183 1,305 0.841 0.844 0.98 (0.84, 1.15) 0.8101
Singapore Chinese
1.07 (0.95, 1.20) 0.2728
replication 1 None 1,082 1,006 0.832 0.819
Risk allele
N frequencies
Figure imgf000047_0001
(uncorrecte
Stage Cohort Adjustment T2D Control T2D Control OR (95% CI) d GC)
Singapore Chinese
replication 2 None 928 939 0.833 0.816
1,873 1,839 0.8396 0.8167 0.01091
Chinese replication First 2 PCs 1.17 (1.04, 1.32)
Replication in other East
6.0x 10
Asian 10,573 11,055 1.10 (1.04, 1.17) 0.6767 0.000
Meta-analysis of East
Asian 21,640 18,984 1.14 (1.09, 1.19) 2.3x l0_lu 0.5939 0.000
4. In silico Singapore Malaysian
replications in South replication None 794 1,204 0.798 0.804 0.97 (0.83, 1.14) 0.7185
Asians and Singapore Indian
Europeans replication None 977 1,169 0.647 0.682 0.86 (0.76, 0.98) 0.0276
DIAGRAM None 8,130 38,987 1.06 (1.02, 1.12) 8.6x l0"3
Replication in non-East
Asian 9,901 41,360 1.03 (0.99, 1.08) 0.1156 0.0042 0.878
ORs and 95% CIs were reported with respect to the T2D-related risk alleles (G)
Phet refers to the p value obtained from the heterogeneity test
GC, genomic control; PC, Principal components
Table 4. Quality control of the participants in stage 1
HK 1 HK 2 SH
Controls Cases Controls Cases Controls Cases
Non-
Epilepsy Eczema
hypertension
Number of subject before QC 102 100 457 111 100 400 200 200
Exclusion criteria:
duplicate 0 0 0 0 0 0 0 0
Relatedness 2 0 0 0 0 2 0 0 gender problem 1 1 7 1 0 8 3 3 overall call rate <0.98 0 0 0 0 0 0 2 1 problem of population stratification 0 0 0 0 1 2 0 0 missing phenotype data 0 0 0 0 0 0 0 0
Number of subject after QC 99 99 450 110 99 388 197 197
Table 5. Quality control of the genotyping results in stage 1
H 1 HK 2 SH
Controls Cases Diseased Controls Cases Controls Cases
Non-
Epilepsy Eczema
hypertension
Number of SNPs on chromosomes 1 - 22 before
541 ,224 541 ,224 460,992 576,857 492,147 576,498 334,157 334,157 QC
Stepwise exclusion criteria:
Step 1 : Monomorphic SNPs 58,440 57,680 0 53,422 194 51 ,746 25,397 24,237 Step 2: SNPs with 100% missing genotype 78 78 0 112 0 0 2 0 Step 3: SNPs with MAF < 5% and their call rate
3,050 3,132 1 18 894 250 2,476 1 ,388 3,468 is < 99%
Step 4: SNPs with MAF > 5% and their call rate
2,000 2,591 0 1 ,546 70 2,037 630 2,926 is < 95%
Step 5: SNPs with overall MAF < 1% 9,656 9,967 81 26,447 1 ,311 27,618 236 214 Step 6: SNPs without HWE (P < 1 10~4) 463 458 54 598 19 575 12,281 12,486 Step 7: SNPs with different MAF between
47 47 47
diseased control cohorts (P < 1 x 10 4)
Number of SNPs on chromosome 1 - 22 after QC 467,537 467,318 460,692 493,791 490,256 492,046 294,223 290,826
Table 6. Selection rationales of SNPs for stage 2 replication
SNP Chr Gene Reason for selection
Unconditional analysis
rs 10229583 7 PAX4 genotyped SNP
7 PAX4 proxy of genotyped SNP (MAF ~ 0.1, r = 0.77, D' = 1 , to rsl 0229583) rs2737250 8 TRPS1 genotyped SNP (r = 0.9, D' = 0.1 , to top imputed SNP, MAF = 0.3)
rs2721960 8 TRPS1 2nd top imputed SNP (MAF = = 0.3)
rs3858158 10 COL13A1 top imputed SNP
rs2395272 10 COL13A1 2nd top imputed SNP
rs57703465 10 COL13A1 3rd top imputed SNP
rs 1 1065453 12 P2RX7 top imputed SNP
rsl0849851 12 P2RX7 genotyped SNP
rs684201 12 P2RX7 proxy of genotyped SNP (r2 = 0.94, D' = 1 to rsl 0849851 , MAF = 0.24)
rs 11065450 12 P2RX7 proxy of genotyped SNP (r2 = 0.65, D' = 1 to rsl 0849851 , MAF ~ 0.3)
rs 11065441 12 P2RX7 proxy of genotyped SNP (r2 = 0.9, D' = 0.96 to rsl0849851, MAF ~ 0.3)
rs208290 12 P2RX7 proxy of genotyped SNP (r2 = 0.56, D' = 1 to rsl 0849851 , MAF ~ 0.35)
Remark: SNPs in CDKN2A/B were not selected for replication due to their high r with the reported SNP rsl 0811661 (r = 0.8). SNPs failed in genotyping were highlighted in red colour.
Table 7. Quality control of the genotyping results in stage 2
Frequencies of
Hong Kong replication 1 Shanghai replication 1 g Kong replication 2 Shanghai replication 2 minor allele for
Case-control study Case-control study Family study Family study T2D in Chinese
MAF
Minor / SNP SNP
HapMap HapMa in HWE SNP MAF in HWE SNP HWE Mendel HWE Mendel
Chr dbSNP ID Gene Major call call
CEU p CHB contro p call rate control p call rate p error p error allele rate rate
1
7 rsl0229583 PAX4 A/G 0.990 0.0279 1 0.985 0.8633 0
0.261 0.202 0.996 0.170 0.3501 0.972 0.175 0.5036
8 rs2721 60 TRPS1 T/C 0.347 0.356 0.899 0.356 0.0656 0.965 0.362 0.4052
8 rs2737250 TRPS1 G/A 0.985 0.8557 1 0.845 0.3024 0
0.308 0.339 0.943 0.380 0.0808 0.976 0.379 0.4339
10 rs3858158 COL13A1 C/T 0.965 0.479 0.3148 0.975 0.439 0.8454
10 rs2395272 COL13A1 A/G 0.959 0.466 0.0521 0.942 0.416 0.7150
10 rs57703465 COL13A1 T/C 0.912 0.338 0.1079 0.974 0.344 0.9153
12 rsll065441 P2RX7 C/T 0.062 0.244 0.893 0.276 0.9092 0.971 0.267 0.3252
12 rs684201 P2RX7 A/G 0.070 0.244 0.944 0.274 0.2899 0.974 0.261 0.3832
12 rsl 1065450 P2RX7 A/C 0.239 0.262 0.938 0.312 0.3192 0.970 0.293 0.6150
12 i-s208290 P2RX7 T/C 0.336 0.321 0.905 0.391 0.9262 0.971 0.355 0.7000
12 rsl0849851 P2RX7 G/A 0.062 0.244 0.943 0.280 0.6369 0.977 0.259 1.0000
Figure imgf000052_0001
50
HKl (99 T2D vs 99 controls) HK2 (388 T2D vs 659 controls) SHGWA (197 T2D vs 197 controls) Combined
Nearest RA/ RA OR RA OR RA OR OR Pmeta
SNP Chr ,ΙΜΡ Rsq Podded IMP Rsq P IMP Rsq 1'
gene(s) NRA Freq (95% CI) Freq (95% CI) ' Freq (95% CI) _ (95% CI) (uncorrect)
(0.97, 2.31) (1.19, 1.95) (1.19,2.17) 1.30, 1.84)
1.53 1.61 1.55
rs2737244 8 TRPS1 A/T 0.670 ^50 0.0634 1 0.9721 0.670 6.9x10" 1 0.9638 0.639 0.0016 0.9775 8.5x10" 0.9533 0.000
(0.97, 2.31) (1.19, 1.95) (1.19,2.17) 1.30, 1.84)
1.46 1.44 1.63 1.50
rs2721962 8 TRPS1 T/A 0.652 0.0838 1 0.9839 0.644 0.0031 1 0.9727 0.618 9.5x10"· 0.991 3.2x10" 0.8047 0.000
(0.95, 2.24) (1.13, 1.83) (1.21,2.19) 1.27, 1.79)
1.43 1.45 1.63 1.51
rs2178951 8 TRPS1 A/G 0.655 0.1022 1 0.9763 0.646 0.0026 1 0.966 0.622 0.0010 0.9854 3.4x10" 0.8118 0.000
(0.93,2.21) (1.14, 1.85) (1.21.2.19) 1.27, 1.79)
1.45 1.43 1.63 1.50
rs2205380 8 TRPS1 G/A 0.651 0.0850 1 0.994 0.642 0.0037 1 0.9828 0.619 9.1x10"· 0.9957 3.9x10 0.7780 0.000
(0.95, 2.23) (1.12.1.81) (1.22.2.19) 1.26, 1.78)
1.41 1.63 1.49
rs2737247* 8 TRPS1 A/G 0.651 0.0852 0 1 0.640 0.0048 0 1 0.619 8.9x10"· 1 4.9 10 0.7400 0.000
(0.95, 2.23) (1.11, 1.79) (1.22, 2.18) 1.25, 1.76)
145 1.41 1.63 1.49
rs2737248 8 TRPS1 A/G 0.651 0.0848 1 0.9947 0.640 0.0048 1 0.999 0.619 8.9x10"· 0.9984 4.9x10 0.7400 0.000
(0.95, 2.23) (1.11, 1.79) (1.22,2.18) 1.25, 1.76)
1.46 1.40 1.64 1.48
rs2737250* 8 TRPS1 A/G 0.654 0.0812 0 1 0.641 0.0058 0 1 0.617 8.3x10"· 0.9985 5.6x10 0.7112 0.000
(0.95, 2.23) (1.10, 1.77) (1.22,2.19) 1.25, 1.76)
213 1.44 1.42 1.51
rsl0965243 9 CDKN2A/B A/G 0.596 0.0017 1 0.8185 0.623 0.0019 0 1 0.621 0.0135 1 8.6x10 0.3268 0.106
(1.31,3.45) (1.14, 1.82) (1.07, 1.89) 1.26, 1.81)
212 1.44 1.42 1.51
rsl0965245 9 CDKN2A/B G/A 0.596 0.0017 1 0.8244 0.623 0.0019 0.9843 0.621 0.0138 0.9859 7.3x10- 0.3310 0.096
(1.31,3.44) (1.14, 1.82) 1 (1.07, 1.89) 1.26, 1.81)
1.71 1.33 1.55
rs3858158* 10 COL13A1 T/C 0.586 L55 0.0628 1 0.8185 0.569 2.4 x10 s 0.8 .604 0.0696 0.8264 2.1x10 0.4686 0.000
(0.97, 2.46) (1.33.2.20) 1 299 0
(0.98.1.81) 1.29, 1.85)
1.67 1. 1 1.51
rs2395272* 10 COL13A1 G/A 0.583 L5° 0.0728 1 0.8875 0.566 2.5x10 s 0.8927 0.605 0.0794 0.8875 3.1x10 0.4497 0.000
(0.96, 2.33) (1.31, 2.13) 1 (0.97, 1.76) .1.27, 1.80)
1.57 1.95 1.45 1.72
rs57703465* 10 COL13A1 G/A 0.717 0.1363 1 0.5826 0.704 4.9x10 s 0.5924 0.745 0.0901 0.5391 9.1x10 0.5446 0.000
(0.86, 2.86) (1.40,2.70) 1 (0.94, 2.24) ,1.35,2.19)
129 1.60 1.67 1.57
rs3861798 12 P2RX7 TIC 0.762 0.3212 1 0.979 0.751 4.9x10"· 0.9687 0.712 0.0022 0.9758 4.0x10 0.7063 0.000
(0.78,2.15) (1.22,2.09) 1 (1.20,2.32) ;i.30, 1.91)
chrl2:1200498 1.40 , 1.63 1.73 1.63
12 P2RX7 C/T 0.744 0.2345 1 0.7762 0.762 9.8x10-" 0.8267 0.732 0.0026 0.8319 6.7x10 0.8136 0.000
16 (0.80, 2.43) (1.21,2.20) 1 (1.20,2.50) ;i.32, 2.01)
142 1.52 1.65 1.55
rsll065445 12 P2RX7 C/G 0.733 ' 0.1844 1 0.8786 0.742 0.0016 0.9711 0.717 0.0029 0.9815 8.4x10- 0.8818 0.000
(0.84, 2.39) (1.17, 1.99) 1 (1.18,2.30) ;i.28, 1.88)
1.42 1.52 1.64 1.55
rs684201* 12 P2RX7 G/A 0.733 0.1834 1 0.8791 0.742 0.0016 0.9725 0.717 0.0029 0.9823 8.6x10 0.8864 0.000
(0.84, 2.40) (1.17, 1.99) 1 (1.18.2.30) ;i.28, 1.88)
138 1.52 1.66 1.54
rs3900976 12 P2RX7 C/T 0.733 ' 0.2208 1 0.882 0.741 0.0018 0.9735 0.717 0.0025 0.9845 9.8x10 0.8360 0.000
(0.82, 2.33) (1.16, 1.98) (1.19,2.31) [1.27, 1.87)
HK1 (99 T2D vs 99 controls) HK2 (388 T2D vs 659 controls) SHGWA (197 T2D vs 197 controls) Combined
Nearest RA RA OR RA OR RA OR OR
SNP Chr IMP IMP Rsq
gene(s) NRA Freq (95% CI) ' Freq (95% CI) ' Freq (95% CI) (95% CI) (uncorrect)
1.39
rs79362551 12 P2RX7 G/A 0.733 0.2109 1 0.8771 0.742 ' 0.0017 1 0.9725 0.717 ' 0.0027 1 0.9798 ' 9.7x l0"6 0.8567 0.000
(0.83, 2.35) (1.17, 1.98) (1.18, 2.30) (1 27, 1.87)
1.44 1.53 1.64 1.55
rsl 1065452 12 P2RX7 G/T 0.734 0.1703 1 0.8721 0.742 0.0015 1 0.9696 0.717 0.0032 1 0.9747 8.5x lff6 0.9069 0.000
(0.85, 2.43) (1.17, 2.00) (1.17, 2.29) (1 28, 1.88)
1 38 1.62 1.66 1.60
rsl 1065453* 12 P2RX7 C/T 0.772 ' 0.2355 1 0.9307 0.750 4.2 xl0~4 1 0.9578 0.721 0.0025 1 0.9669 2.9x l0~6 0.8362 0.000
(0.81, 2.34) (1.23, 2.12) (1.19, 2.32) (1 31, 1.94)
1 52 1.63 . 1.56 1.59
rsl0849849 12 P2RX7 A/G 0.772 ' 0.1237 1 0.9089 0.755 4.8xl0~4 1 0.918 0.732 0.0099 1 0.9141 5.7X 10-6 0.9704 0.000
(0.89, 2.62) (1.23, 2.15) (1.11 , 2.21) (1 30, 1.95)
1 58 1.63 . 1.50 1.58
rsl0849850 12 P2RX7 A/G 0.769 ' 0.0813 1 0.9764 0.749 2.8xl0~4 1 0.9845 0.725 0.0134 1 0.9908 3.5X 10-6 0.9286 0.000
(0.94, 2.67) (1.25, 2.14) (1.09, 2.08) (1 30, 1.92)
1.64 , 1.51 1.58
rs208289 12 P2RX7 A/G 0.721 L55 0.1064 1 0.8194 0.709 3.8xlff4 1 0.8477 0.692 0.0154 1 0.858 6.4x 10 6 0.9312 0.000
(0.91, 2.63) (1.24, 2.16) (1.08, 2.11) (1 29, 1.92)
1.64 1.47 1.56
rs208292 12 P2RX7 A/G 0.765 ^50 0.1176 1 0.9991 0.746 2.2 xl0~4 1 0.9973 0.724 0.0184 1 0.9981 5.2χ 10~6 0.8717 0.000
(0.90, 2.51) (1.26, 2.14) (1.06, 2.04) (1 29, 1.89)
1.50 1.64 . 1.47 1.56
rsl0849851* 12 P2RX7 A/G 0.765 0.1179 0 1 0.746 2.2 xl0~4 0 1 0.724 0.0185 0 1 5.1 x l0 6 0.8704 0.000
(0.90, 2.51) (1.26, 2.14) (1.07, 2.04) (1 29, 1.89)
1.64 . 1.47 1.56
rs6489794 12 P2RX7 G/A 0.765 ^50 0.1174 1 0.9978 0.746 2.2 xl0~4 1 0.996 0.724 0.0180 1 0.9968 4.9x 10 6 0.8706 0.000
(0.90, 2.51) (1.26, 2.14) (1.07, 2.04) (1 29, 1.89)
1.64 , 1.48 1.57
rsl 1065458 12 P2RX7 C/T 0.767 L56 0.0872 1 0.9878 0.747 2.6 x Iff4 1 0.9834 0.724 0.0170 1 0.9905 4.4x lff6 0.8894 0.000
(0.93, 2.62) (1.25, 2.15) (1.07, 2.05) (1 30, 1.91)
"*" refers to the SNP selected for stage 2 replication. Nearest Entrez genes within 250 kb. Pad)uae P> Pmaa and Pha represent P values from logistic regression with/without adjustment for age and sex under the additiv meta-analysis under a random effects models (uncorrected for genomic control) and test of heterogeneity, respectively. Risk allele refers to the allele with a higher frequency in T2D patients than in controls in stage OR, odds ratio are reported with respect to the risk allele. RA, risk allele; NRA, non-risk allele; IMP, 1 = imputed, 2 = genotyped; Rsq, indicates the imputation quality provided from MACH.
Table 9. Logistic regression results of T2D with the selected top SNPs adjusted for the first principle component in stage 1
HKl (99 T2D vs 99 HK2 (388 T2D vs 659 SHGWA (197 T2D vs 197
controls) controls) controls) Combined
Nearest Risk/ OR OR OR OR p
Non- gene(s) risk (95% CI) (95% CI) (95% CI) (95% CI) (uncorrect)
SNP Chr allele p p p
1.41 1.55 1.91 1.64
rs 10229583 7 PAX4 G/A (0.8, 2.48) 0.236 (1.13, 2.13) 0.0061 (1.31 , 2.78) 7.9 x 10"4 (1.31, 2.05) 1.3 x 10"5 0.5996
1.47 1.4 1.65 1.49
rs2737250 8 TRPS1 A/G (0.96, 2.26) 0.0775 (1.1, 1.78) 0.006 (1.23, 2.21) 8.4 x 10-4 (1.26, 1.77) 4.5 x 10-6 0.6965
1.56 1.64 1.65 1.63
rs2721960 8 TRPS1 G/A (0.96, 2.54) 0.0711 (1.25, 2.16) 3.3 x W4 (1.19, 2.28) 0.0027 (1.35, 1.98) 6.0 x 10-7 0.9808
1.65 1.72 1.32 1.56
rs3858158 10 COL13A1 T/C (1.03, 2.65) 0.0389 (1.33, 2.21) 3.0 x 10-5 (0.97, 1.81) 0.0762 (1.3, 1.87) 1.8 x 10"6 0.4195
1.58 1.68 1.3 1.53
rs2395272 10 COL13A1 G/A (1.01, 2.49) 0.0462 (1.31, 2.14) 3.1 x 10"5 (0.96, 1.75) 0.0864 (1.28, 1.82) 2.9 x 10-6 0.4337
1.68 1.95 1.44 1.74
rs57703465 10 COL13A1 G/A (0.91, 3.08) 0.0955 (1.41, 2.71) 6.2 x 10-5 (0.94, 2.23) 0.0972 (1.37, 2.2) 5.7 x 10-6 0.5372
1.43 1.61 1.64 1.59
rsl 1065453 12 P2RX7 C/T (0.84, 2.44) 0.1906 (1.23, 2.11) 5.5 x W4 (1.17, 2.3) 0.0038 (1.31, 1.94) 3.1 x 10"6 0.9082
1.56 1.64 1.45 1.56
rsl0849851 12 P2RX7 A/G (0.93, 2.62) 0.0906 (1.25, 2.14) 3.1 x W4 (1.04, 2.02) 0.0266 (1.28, 1.9) 7.6 x 10"6 0.8538
1.5 1.52 1.63 1.55
rs684201 12 P2RX7 G/A (0.89, 2.54) 0.1314 (1.17, 1.98) 0.002 (1.17, 2.27) 0.0043 (1.28, 1.88) 6.5 x 10-6 0.9396
Table 10. Alternating logistic regression results of T2D for PAX4 rsl0229583 and TRPSl rs2737250 SNPs in de novo replication of family study in Chinese populations
N PAX4 rsl0229583 (G) TRPSl rs2737250 (A)
Contr Contro
Contro Case OR Case OR
Study Case ol
1 RAF (95% CI) r"ddmve Phet 1
RAF (95% CI) addilive Phe
RAF RAF
Hong Kong 1.22 0.86
325 368 0.872 0.856 0.2817 0.682 0.704 0.2026
replication 2 (0.85, 1.74) (0.68, 1.09)
Shanghai replication 1.09 0.82
657 168 0.824 0.813 0.5757 0.618 0.646 0.0218
2 (0.80, 1.49) (0.70, 0.97)
Meta-analysis of
1.15 0.83
family studies in 982 536 0.2590 0.6567 0.000 0.0091 0.7859 0.000
(0.91, 1.45) (0.73, 0.96)
Chinese
Meta-analysis of
1.18 1.07
studies in stage 1 and 11067 7929 2.6x l0 8 0.0839 0.596 0.0034 1.5x l0-5 0.855
(1.11, 1.25) (1.02, 1.12)
2 in Chinese
Nearest Entrez genes within 250 kb. additive and Phet represents P values from alternating logistic regression with adjustment of sex and age under additive genetic model or meta-analysis under fixed effect models and test of heterogeneity, respectively. Risk allele of genetic variant is indicated within the parentheses. OR, odds ratios are reported with respect to the risk allele in stage 1 ; RAF, risk allele frequency.
Table 11. Associations of rsl0229583 with age at diagnosis in Chinese T2D patients
Age at diagnosis
Study Genotypes N Mean ± SD P unadjusted Padjusted isex and BMI) adjured (sex, BMI and Hbau)
Hong Kong T2D case AA 129 51.1 ± 11.7 0.0043 0.0095 0.0101
AG 1479 50.4 ± 13.1
GG 4129 49.4 ± 12.8
Shanghai T2D case AA 103 54.2 ± 11.5 0.0193 0.0185 0.3173
AG 1 110 53.3 ± 11.8
GG 3103 52.4 ± 12.2
Combined 2.3 x 10 4 4.6 x 10 4 0.0095
Test of heterogeneity Phet 0.8202 0.9850 0.3429
I2 0.000 0.000 0.000
Data are expressed as n and mean ± SD. P unadjusted and P adjusted represent P values calculated from linear regression with and without adjustmen for sex, BMI and Hbaic (where appropriate) under the additive genetic model. In the combined analysis, P values were obtained from met analysis of two Chinese T2D case cohorts (Hong Kong and Shanghai) under a fixed effect model.
Table 12. Heterogeneity test of OR of rsl0229583 among different populations in the present study
OR Heterogeneity of OR
Population (95% CI) Q test P I2 chinese (1.1 U .25)
1 09
Japanese + Korean ( Λ Λ 1 , , ΟΛ 0.1254 0.5742
(1 .U1 , 1.1 ο)
DIAGRAM+ 02 °^ 12) 0.0088 0.8544
Malaysian + Indian (0 82 00) 1.1 x 10"5 0.9485
Q test P and I2 refer to the statistical significance and quantified index of heterogeneity test of OR between Chinese and other populations, respectively.
Table 13. Comparisons of regional LD and allelic discrepancies of rsl0229583 among different populations using the HapMap phase 3 data
Genotype frequencies Allele frequencies Regional LD variation Allelic difference
Population G/G A/G A/A Risk allele (G) Non-risk allele (A) Monte Carlo P X2 test
CHB 0.664 0.307 0.029 0.818 0.182 -- ~
JPT 0.777 0.223 0.000 0.888 0.112 0.0107 0.0278
CEU 0.566 0.345 0.088 0.739 0.261 0.0018 0.0342
YRI 0.517 0.415 0.068 0.724 0.276 0.0003 0.0085
Monte Carlo and χ test P refer to the P values testing for the regional LD variation and allelic difference between CHB and other populations, respectively. CHB: Han Chinese in Beijing, China; JPT: Japanese in Tokyo, Japan; CEU: Utah residents with Northern and Western European ancestry from the CEPH collection; YRI: Yoruban in Ibadan, Nigeria.

Claims

WHAT IS CLAIMED IS:
1. A method for assessing the presence or risk of type 2 diabetes or cardiovascular disease in a subject, comprising the steps of:
(a) performing an assay that determines nucleotide sequence of at least a portion of PAX4-SND1 genomic sequence present in a biological sample taken from the subject, and
(b) comparing the sequence determined in step (a) with a standard sequence of the corresponding genomic sequence, wherein a variation in the sequence determined in step (a) when compared with the standard sequence indicates the presence or risk of type 2 diabetes or cardiovascular disease in the subject.
2. The method of claim 1, wherein the sample is a blood or saliva sample.
3. The method of claim 1, wherein the subject is of Asian descent.
4. The method of claim 1, wherein the subject is a Chinese.
5. The method of claim 1, wherein the subject is a Han Chinese.
6. The method of claim 1, wherein the subject has a family history of type 2 diabetes but has not been diagnosed of type 2 diabetes.
7. The method of claim 1, wherein the subject has been diagnosed with type 2 diabetes, and wherein a variation in the sequence indicates the presence or risk of developing cardiovascular disease in the subject.
8. The method of claim 1, wherein the assay in step (a) comprises an amplification reaction.
9. The method of claim 8, wherein the amplification reaction is a polymerase chain reaction (PCR).
10. The method of claim 1, wherein the assay in step (a) comprises mass spectrometry.
11. The method of claim 1, wherein, when the subject is indicated as having or at risk of developing type 2 diabetes, further comprising the step of administering to the subject a cholesterol lowering drug or a blood glucose lowering drug.
12. The method of claim 1, wherein the sequence variation is a polymorphism rs 10229583.
13. The method of claim 1, wherein the sequence variation is the G allele of polymorphism rs 10229583.
14. The method of claim 1, wherein the sequence variation is a polymorphism rs7801 1 1 , rs806187, rsl440971 , rs806179, rs806178, rs7781 189, or rs806176.
15. The method of claim 1, further comprising a step of detecting a second sequence variation in a second genomic sequence, after a sequence variation is detected in the portion of PAX4-SND1 genomic sequence after steps (a) and (b).
16. The method of claim 15, wherein the second sequence variation is a polymorphism rs2737250.
17. The method of claim 1, wherein the second sequence variation is one of the polymorphisms listed in Tables 2, 6, 7, and 8.
18. The method of claim 1, further comprising the steps of detecting a second, third, or more sequence variations in a second, third, or more genomic sequences, after a sequence variation is detected in the portion of PAX4-SND1 genomic sequence after steps (a) and (b).
19. A kit for assessing the presence or risk of type 2 diabetes in a subject, comprising two oligonucleotide primers for specifically amplifying: (1) at least a segment of the PAX4-SND1 genomic sequence; or (2) complement of (1), in an amplification reaction.
20. The kit of claim 19, wherein the amplification reaction is a polymerase chain reaction (PCR).
21. The kit of claim 19, further comprising an oligonucleotide probe that specifically binds to product of the amplification reaction.
22. The kit of claim 19, further comprising an instruction manual.
PCT/CN2014/000226 2013-03-08 2014-03-10 New biomarker for type 2 diabetes Ceased WO2014134970A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US201361775103P 2013-03-08 2013-03-08
US61/775,103 2013-03-08

Publications (1)

Publication Number Publication Date
WO2014134970A1 true WO2014134970A1 (en) 2014-09-12

Family

ID=51490610

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2014/000226 Ceased WO2014134970A1 (en) 2013-03-08 2014-03-10 New biomarker for type 2 diabetes

Country Status (1)

Country Link
WO (1) WO2014134970A1 (en)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2019237209A1 (en) * 2018-06-15 2019-12-19 Opti-Thera Inc. Polygenic risk scores for predicting disease complications and/or response to therapy
CN111584080A (en) * 2020-04-18 2020-08-25 中国医学科学院北京协和医院 Diagnostic model for distinguishing MODY from T1D and T2D

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101960019A (en) * 2007-04-03 2011-01-26 国家科学研究中心 FTO gene polymorphisms associated to obesity and/or type II diabetes
WO2012097903A1 (en) * 2011-01-20 2012-07-26 Université Libre de Bruxelles Methylation patterns of type 2 diabetes patients

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101960019A (en) * 2007-04-03 2011-01-26 国家科学研究中心 FTO gene polymorphisms associated to obesity and/or type II diabetes
WO2012097903A1 (en) * 2011-01-20 2012-07-26 Université Libre de Bruxelles Methylation patterns of type 2 diabetes patients

Non-Patent Citations (3)

* Cited by examiner, † Cited by third party
Title
KOOPTIWUT, S. ET AL.: "Defective PAX4 R192H transcriptional repressor activities associated with maturity onset diabetes of the young and early onset-age of type 2 diabetes.", JOURNAL OF DIABETES AND ITS COMPLICATION., vol. 26, 20 April 2012 (2012-04-20), pages 343 - 347 *
MA, R.C.W. ET AL.: "Genome-wide association study in a Chinese population identifies a susceptibility locus for type 2 diabetes at 7q32 near PAX4", DIABETOLOGIA, vol. 56, 27 March 2013 (2013-03-27), pages 1291 - 1305 *
YUAN, WENLI ET AL.: "Research progress on association of PAX4 with diabetes.", INTERNATIONAL JOURNAL OF LABORATORY MEDICINE., vol. 33, no. 16, 31 August 2012 (2012-08-31), pages 1985 - 1987 *

Cited By (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2019237209A1 (en) * 2018-06-15 2019-12-19 Opti-Thera Inc. Polygenic risk scores for predicting disease complications and/or response to therapy
CN112513295A (en) * 2018-06-15 2021-03-16 Opti 西拉公司 Multi-gene risk score for predicting disease complications and/or response to treatment
EP3807883A4 (en) * 2018-06-15 2022-03-23 Opti-Thera Inc. POLYGENIC RISK SCORES TO PREDICT DISEASE COMPLICATIONS AND/OR RESPONSE TO THERAPY
US12467091B2 (en) 2018-06-15 2025-11-11 Opti-Thera Inc. Polygenic risk scores for predicting disease complications and/or response to therapy
CN111584080A (en) * 2020-04-18 2020-08-25 中国医学科学院北京协和医院 Diagnostic model for distinguishing MODY from T1D and T2D
CN111584080B (en) * 2020-04-18 2022-07-26 中国医学科学院北京协和医院 Diagnostic kit and diagnostic system for distinguishing MODY from T1D and T2D

Similar Documents

Publication Publication Date Title
EP1978107A1 (en) Fto gene polymorphisms associated to obesity and/or type II diabetes
US20110129820A1 (en) Human diabetes susceptibility tnfrsf10b gene
Turki et al. Transcription factor-7-like 2 gene variants are strongly associated with type 2 diabetes in Tunisian Arab subjects
US10689702B2 (en) Biomarkers for diabetes
US9518298B2 (en) DACH1 as a biomarker for diabetes
CN103525899B (en) Diabetes B susceptibility loci and detection method and test kit
WO2008144940A1 (en) Biomarker for hypertriglyceridemia
WO2014134970A1 (en) New biomarker for type 2 diabetes
Makni et al. Association of glucose transporter 1 polymorphisms with type 2 diabetes in the Tunisian population
US8236497B2 (en) Methods of diagnosing cardiovascular disease
CN104411824A (en) Snp markers associated with polycystic ovary syndrome
US20080194419A1 (en) Genetic Association of Polymorphisms in the Atf6-Alpha Gene with Insulin Resistance Phenotypes
US20100105057A1 (en) Human diabetes susceptibility tnfrsf10d gene
US20100184839A1 (en) Allelic polymorphism associated with diabetes
WO2010033825A2 (en) Genetic variants associated with abdominal aortic aneurysms
US20100285459A1 (en) Human Diabetes Susceptibility TNFRSF10A gene
US20100151462A1 (en) Human diabetes susceptibility shank2 gene
WO2008122671A1 (en) Human diabetes susceptibility tnfrsf10c gene
HK1202899B (en) Biomarkers for diabetes
WO2006136791A1 (en) Polymorphisms and haplotypes in p2x7 gene and their use in determining susceptibility for atherosclerosis-mediated diseases

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 14759776

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 14759776

Country of ref document: EP

Kind code of ref document: A1