Novel Proteins and Nucleic Acids Encoding Same
Field of the Invention
The present invention relates to novel polypeptides that are targets of small molecule drugs and that have properties related to stimulation of biochemical or physiological responses in a cell, a tissue, an organ or an organism. More particularly, the novel polypeptides are gene products of novel genes, or are specified biologically active fragments or derivatives thereof. Methods of use encompass diagnostic and prognostic assay procedures as well as methods of treating diverse pathological conditions. The present invention discloses novel associations of proteins and polypeptides and the nucleic acids that encode them with various diseases or pathologies. The proteins and related proteins that are similar to them, are encoded by a cDNA and/or by genomic DNA. The proteins, polypeptides and their cognate nucleic acids were identified by Curagen Corporation in certain cases. The XYZase-encoded protein and any variants, thereof, are suitable as diagnostic markers, targets for an antibody therapeutic and targets for small molecule drugs. As such the current invention embodies the use of recombinantly expressed and/or endogenously expressed protein in various screens to identify such therapeutic antibodies and/or therapeutic small molecules.
Background
Eukaryotic cells are characterized by biochemical and physiological processes which under normal conditions are exquisitely balanced to achieve the preservation and propagation of the cells. When such cells are components of multicellular organisms such as vertebrates, or more particularly organisms such as mammals, the regulation of the biochemical and physiological processes involves intricate signaling pathways. Frequently, such signaling pathways are constituted of extracellular signaling proteins, cellular receptors that bind the signaling proteins and signal transducing components located within the cells.
Signaling proteins may be classified as endocrine effectors, paracrine effectors or autocrine effectors. Endocrine effectors are signaling molecules secreted by a given organ into the circulatory system, which are then transported to a distant target organ or tissue. The target cells include the receptors for the endocrine effector, and when the endocrine effector binds, a signaling cascade is induced. Paracrine effectors involve secreting cells and receptor cells in close proximity to each other, for example two different classes of cells in the same tissue or organ. One class of cells secretes the paracrine effector, which then reaches the second class of cells, for example by diffusion through the extracellular fluid. The second class of cells contains the receptors for the paracrine effector; binding of the effector results in induction of the signaling cascade that elicits the corresponding biochemical or physiological effect. Autocrine effectors are highly analogous to paracrine effectors, except that the same cell type that secretes the autocrine effector also contains the receptor. Thus the autocrine effector binds to receptors on the same cell, or on identical neighboring cells. The binding process then elicits the characteristic biochemical or physiological effect.
Signaling processes may elicit a variety of effects on cells and tissues including by way of nonlimiting example induction of cell or tissue proliferation, suppression of growth or proliferation, induction of differentiation or maturation of a cell or tissue, and suppression of differentiation or maturation of a cell or tissue.
Many pathological conditions involve dysregulation of expression of important effector proteins. In certain classes of pathologies the dysregulation is manifested as diminished or suppressed level of synthesis and secretion protein effectors. In a clinical setting a subject may be suspected of suffering from a condition brought on by diminished or suppressed levels of a protein effector of interest. Therefore there is a need to be able to assay for the level of the protein effector of interest in a biological sample from such a subject, and to compare the level with that characteristic of a nonpatho logical condition. There further is a
need to provide the protein effector as a product of manufacture. Administration of the effector to a subject in need thereof is useful in treatment of the pathological condition, or the protein effector deficiency or suppression may be favorably acted upon by the administration of another small molecule drug product. Accordingly, there is a need for a method of treatment of a pathological condition brought on by a diminished or suppressed levels of the protein effector of interest.
Small molecule targets have been implicated in various disease states or pathologies. These targets may be proteins, and particularly enzymatic proteins, which are acted upon by small molecule drugs for the purpose of altering target function and achieving a desired result. Cellular, animal and clinical studies can be performed to elucidate the genetic contribution to the etiology and pathogenesis of conditions in which small molecule targets are implicated in a variety of physiologic, pharmaco logic or native states. These studies utilize the core technologies at CuraGen Corporation to look at differential gene expression, protein-protein interactions, large-scale sequencing of expressed genes and the association of genetic variations such as, but not limited to, single nucleotide polymoφhisms (SNPs) or splice variants in and between biological samples from experimental and control groups. The goal of such studies is to identify potential avenues for therapeutic intervention in order to prevent, treat the consequences or cure the conditions.
In order to treat diseases, pathologies and other abnormal states or conditions in which a mammalian organism has been diagnosed as being, or as being at risk for becoming, other than in a normal state or condition, it is important to identify new therapeutic agents. Such a procedure includes at least the steps of identifying a target component within an affected tissue or organ, and identifying a candidate therapeutic agent that modulates the functional attributes of the target. The target component may be any biological macromolecule implicated in the disease or pathology. Commonly the target is a polypeptide or protein with specific functional attributes. Other classes of macromolecule may be a nucleic acid, a polysaccharide, a lipid such as a complex lipid or a glycolipid; in addition a target may be a sub-cellular structure or extra-cellular structure that is comprised of more than one of these classes of macromolecule. Once such a target has been identified, it may be employed in a screening assay in order to identify favorable candidate therapeutic agents from among a large population of substances or compounds.
In many cases the objective of such screening assays is to identify small molecule candidates; this is commonly approached by the use of combinatorial methodologies to develop the population of substances to be tested. The implementation of high throughput
screening methodologies is advantageous when working with large, combinatorial libraries of compounds.
It is an objective of this invention to provide at least one target biopolymer that is intended to serve as the macromolecular component in a screening assay for identifying candidate pharmaceutical agents.
It is another objective of the present invention to provide screening assays that positively identify candidate pharmaceutical agents from among a combinatorial library of low molecular weight substances or compounds.
It is still a further objective of this invention to employ the candidate pharmaceutical agents in any of a variety of in vitro, ex vivo and in vivo assays in order to identify pharmaceutical agents with advantageous therapeutic applications in the treatment of a disease, pathology, or abnormal state or condition in a mammal.
Summary Of The Invention
The invention is based in part upon the discovery of nucleic acid sequences encoding novel polypeptides. These nucleic acids and polypeptides, as well as derivatives, homologs, analogs and fragments thereof, will hereinafter be collectively designated as "NOVX" nucleic acid, which represents the nucleotide sequence selected from the group consisting of SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178, or polypeptide sequences, which represents the group consisting of SEQ ID NO: 2n, wherein n is an integer between 1 and 178.
In one aspect, the invention provides an isolated polypeptide comprising a mature form of a NOVX amino acid. One example is a variant of a mature form of a NOVX amino acid sequence, wherein any amino acid in the mature form is changed to a different amino acid, provided that no more than 15% of the amino acid residues in the sequence of the mature form are so changed. The amino acid can be, for example, a NOVX amino acid sequence or a variant of a NOVX amino acid sequence, wherein any amino acid specified in the chosen sequence is changed to a different amino acid, provided that no more than 15%) of the amino acid residues in the sequence are so changed. The invention also includes fragments of any of these. In another aspect, the invention also includes an isolated nucleic acid that encodes a NONX polypeptide, or a fragment, homolog, analog or derivative thereof.
Also included in the invention is a ΝOVX polypeptide that is a naturally occurring allelic variant of a ΝOVX sequence. In one embodiment, the allelic variant includes an amino acid sequence that is the translation of a nucleic acid sequence differing by a single nucleotide
from a NOVX nucleic acid sequence. In another embodiment, the NOVX polypeptide is a variant polypeptide described therein, wherein any amino acid specified in the chosen sequence is changed to provide a conservative substitution. In one embodiment, the invention discloses a method for determining the presence or amount of the NOVX polypeptide in a sample. The method involves the steps of: providing a sample; introducing the sample to an antibody that binds immunospecifically to the polypeptide; and determining the presence or amount of antibody bound to the NOVX polypeptide, thereby determining the presence or amount of the NOVX polypeptide in the sample. In another embodiment, the invention provides a method for determining the presence of or predisposition to a disease associated with altered levels of a NOVX polypeptide in a mammalian subject. This method involves the steps of: measuring the level of expression of the polypeptide in a sample from the first mammalian subject; and comparing the amount of the polypeptide in the sample of the first step to the amount of the polypeptide present in a control sample from a second mammalian subject known not to have, or not to be predisposed to, the disease, wherein an alteration in the expression level of the polypeptide in the first subject as compared to the control sample indicates the presence of or predisposition to the disease.
In a further embodiment, the invention includes a method of identifying an agent that binds to a NOVX polypeptide. This method involves the steps of: introducing the polypeptide to the agent; and determining whether the agent binds to the polypeptide. In various embodiments, the agent is a cellular receptor or a downstream effector.
In another aspect, the invention provides a method for identifying a potential therapeutic agent for use in treatment of a pathology, wherein the pathology is related to aberrant expression or aberrant physiological interactions of a NOVX polypeptide. The method involves the steps of: providing a cell expressing the NOVX polypeptide and having a property or function ascribable to the polypeptide; contacting the cell with a composition comprising a candidate substance; and determining whether the substance alters the property or function ascribable to the polypeptide; whereby, if an alteration observed in the presence of the substance is not observed when the cell is contacted with a composition devoid of the substance, the substance is identified as a potential therapeutic agent. In another aspect, the invention describes a method for screening for a modulator of activity or of latency or predisposition to a pathology associated with the NOVX polypeptide. This method involves the following steps: administering a test compound to a test animal at increased risk for a pathology associated with the NOVX polypeptide, wherein the test animal recombinantly expresses the NOVX polypeptide. This method involves the steps of measuring the activity of
the NOVX polypeptide in the test animal after administering the compound of step; and comparing the activity of the protein in the test animal with the activity of the NOVX polypeptide in a control animal not administered the polypeptide, wherein a change in the activity of the NOVX polypeptide in the test animal relative to the control animal indicates the test compound is a modulator of latency of, or predisposition to, a pathology associated with the NOVX polypeptide. In one embodiment, the test animal is a recombinant test animal that expresses a test protein transgene or expresses the transgene under the control of a promoter at an increased level relative to a wild-type test animal, and wherein the promoter is not the native gene promoter of the transgene. In another aspect, the invention includes a method for modulating the activity of the NOVX polypeptide, the method comprising introducing a cell sample expressing the NOVX polypeptide with a compound that binds to the polypeptide in an amount sufficient to modulate the activity of the polypeptide.
The invention also includes an isolated nucleic acid that encodes a NOVX polypeptide, or a fragment, homolog, analog or derivative thereof. In a preferred embodiment, the nucleic acid molecule comprises the nucleotide sequence of a naturally occurring allelic nucleic acid variant. In another embodiment, the nucleic acid encodes a variant polypeptide, wherein the variant polypeptide has the polypeptide sequence of a naturally occurring polypeptide variant. In another embodiment, the nucleic acid molecule differs by a single nucleotide from a NOVX nucleic acid sequence. In one embodiment, the NOVX nucleic acid molecule hybridizes under stringent conditions to the nucleotide sequence selected from the group consisting of SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178, or a complement of the nucleotide sequence. In another aspect, the invention provides a vector or a cell expressing a NOVX nucleotide sequence.
In one embodiment, the invention discloses a method for modulating the activity of a NOVX polypeptide. The method includes the steps of: introducing a cell sample expressing the NOVX polypeptide with a compound that binds to the polypeptide in an amount sufficient to modulate the activity of the polypeptide. In another embodiment, the invention includes an isolated NOVX nucleic acid molecule comprising a nucleic acid sequence encoding a polypeptide comprising a NOVX amino acid sequence or a variant of a mature form of the NOVX amino acid sequence, wherein any amino acid in the mature form of the chosen sequence is changed to a different amino acid, provided that no more than 15% of the amino acid residues in the sequence of the mature form are so changed. In another embodiment, the invention includes an amino acid sequence that is a variant of the NOVX amino acid sequence, in which any amino acid specified in the chosen sequence is changed to a different
amino acid, provided that no more than 15%> of the amino acid residues in the sequence are so changed.
In one embodiment, the invention discloses a NOVX nucleic acid fragment encoding at least a portion of a NOVX polypeptide or any variant of the polypeptide, wherein any amino acid of the chosen sequence is changed to a different amino acid, provided that no more than 10%) of the amino acid residues in the sequence are so changed. In another embodiment, the invention includes the complement of any of the NOVX nucleic acid molecules or a naturally occurring allelic nucleic acid variant. In another embodiment, the invention discloses a NOVX nucleic acid molecule that encodes a variant polypeptide, wherein the variant polypeptide has the polypeptide sequence of a naturally occurring polypeptide variant. In another embodiment, the invention discloses a NOVX nucleic acid, wherein the nucleic acid molecule differs by a single nucleotide from a NOVX nucleic acid sequence.
In another aspect, the invention includes a NOVX nucleic acid, wherein one or more nucleotides in the NOVX nucleotide sequence is changed to a different nucleotide provided that no more than 15%> of the nucleotides are so changed. In one embodiment, the invention discloses a nucleic acid fragment of the NOVX nucleotide sequence and a nucleic acid fragment wherein one or more nucleotides in the NOVX nucleotide sequence is changed from that selected from the group consisting of the chosen sequence to a different nucleotide provided that no more than 15%> of the nucleotides are so changed. In another embodiment, the invention includes a nucleic acid molecule wherein the nucleic acid molecule hybridizes under stringent conditions to a NOVX nucleotide sequence or a complement of the NOVX nucleotide sequence. In one embodiment, the invention includes a nucleic acid molecule, wherein the sequence is changed such that no more than 15% of the nucleotides in the coding sequence differ from the NOVX nucleotide sequence or a fragment thereof.
In a further aspect, the invention includes a method for determining the presence or amount of the NOVX nucleic acid in a sample. The method involves the steps of: providing the sample; introducing the sample to a probe that binds to the nucleic acid molecule; and determining the presence or amount of the probe bound to the NOVX nucleic acid molecule, thereby determining the presence or amount of the NOVX nucleic acid molecule in the sample. In one embodiment, the presence or amount of the nucleic acid molecule is used as a marker for cell or tissue type.
In another aspect, the invention discloses a method for determining the presence of or predisposition to a disease associated with altered levels of the NOVX nucleic acid molecule of in a first mammalian subject. The method involves the steps of: measuring the amount of
NOVX nucleic acid in a sample from the first mammalian subject; and comparing the amount of the nucleic acid in the sample of step (a) to the amount of NOVX nucleic acid present in a control sample from a second mammalian subject known not to have or not be predisposed to, the disease; wherein an alteration in the level of the nucleic acid in the first subject as compared to the control sample indicates the presence of or predisposition to the disease.
Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, suitable methods and materials are described below. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. In the case of conflict, the present specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting.
Other features and advantages of the invention will be apparent from the following detailed description and claims.
Detailed Description Of The Invention
The present invention provides novel nucleotides and polypeptides encoded thereby. Included in the invention are the novel nucleic acid sequences, their encoded polypeptides, antibodies, and other related compounds. The sequences are collectively referred to herein as "NOVX nucleic acids" or "NOVX polynucleotides" and the corresponding encoded polypeptides are referred to as "NOVX polypeptides" or "NOVX proteins." Unless indicated otherwise, "NOVX" is meant to refer to any of the novel sequences disclosed herein. Table 1 provides a summary of the NOVX nucleic acids and their encoded polypeptides.
TABLE 1. Sequences and Corresponding SEQ ID Numbers
Table 1 indicates homology of NOVX nucleic acids to known protein families. Thus, the nucleic acids and polypeptides, antibodies and related compounds according to the invention corresponding to a NOVX as identified in column 1 of Table 1 will be useful in therapeutic and diagnostic applications implicated in, for example, pathologies and disorders associated with the known protein families identified in column 5 of Table 1.
NOVX nucleic acids and their encoded polypeptides are useful in a variety of applications and contexts. The various NOVX nucleic acids and polypeptides according to the invention are useful as novel members of the protein families according to the presence of domains and sequence relatedness to previously described proteins. Additionally, NOVX nucleic .acids and polypeptides can also be used to identify proteins that are members of the family to which the NOVX polypeptides belong.
Consistent with other known members of the family of proteins, identified in column 5 of Table 1, the NOVX polypeptides of the present invention show homology to, and contain domains that are characteristic of, other members of such protein families. Details of the sequence relatedness and domain analysis for each NOVX are presented in Example A.
The NOVX nucleic acids and polypeptides can also be used to screen for molecules, which inhibit or enhance NOVX activity or function. Specifically, the nucleic acids and polypeptides according to the invention may be used as targets for the identification of small molecules that modulate or inhibit diseases associated with the protein families listed in Table 1.
The NOVX nucleic acids and polypeptides are also useful for detecting specific cell types. Details of the expression analysis for each NOVX are presented in Example C. Accordingly, the NOVX nucleic acids, polypeptides, antibodies and related compounds according to the invention will have diagnostic and therapeutic applications in the detection of a variety of diseases with differential expression in normal vs. diseased tissues, e.g.a variety of cancers.
Additional utilities for NOVX nucleic acids and polypeptides according to the invention are disclosed herein.
The present invention is based on the identification of biological macromolecules differentially modulated in a pathologic state, disease, or an abnormal condition or state. Among the pathologies or diseases of present interest include metabolic diseases including those related to endocrinologic disorders, cancers, various tumors and neoplasias, inflammatory disorders, central nervous system disorders, and similar abnormal conditions or
states. In very significant embodiments of the present invention, the biological macromolecules implicated in the pathologies and conditions are proteins and polypeptides, and in such cases the present invention is related as well to the nucleic acids that encode them. Methods that may be employed to identify relevant biological macromolecules include any procedures that detect differential expression of nucleic acids encoding proteins and polypeptides associated with the disorder, as well as procedures that detect the respective proteins and polypeptides themselves. Significant methods that have been employed by the present inventors, include GeneCalling ® technology and SeqCalling TM technology, disclosed respectively, in U. S. Patent No. 5,871,697, and in U. S. Ser. No. 09/417,386, filed Oct. 13, 1999, each of which is incoφorated herein by reference in its entirety. GeneCalling ® is also described in Shimkets, et al., "Gene expression analysis by transcript profiling coupled to a gene database query" Nature Biotechnology 17:198-803 (1999).
The invention provides polypeptides and nucleotides encoded thereby that have been identified as having novel associations with a disease or pathology, or an abnormal state or condition, in a mammal. The present invention further identifies a set of proteins and polypeptides, including naturally occurring polypeptides, precursor forms or proproteins, or mature forms of the polypeptides or proteins, which are implicated as targets for therapeutic agents in the treatment of various diseases, pathologies, abnormal states and conditions. A target may be employed in any of a variety of screening methodologies in order to identify candidate therapeutic agents which interact with the target and in so doing exert a desired or favorable effect. The candidate therapeutic agent is identified by screening a large collection of substances or compounds in an important embodiment of the invention. Such a collection may comprise a combinatorial library of substances or compounds in which, in at least one subset of substances or compounds, the individual members are related to each other by simple structural variations based on a particular canonical or basic chemical structure. The variations may include, by way of nonlimiting example, changes in length or identity of a basic framework of bonded atoms; changes in number, composition and disposition of ringed structures, bridge structures, alicyclic rings, and aromatic rings; and changes in pendent or substituents atoms or groups that are bonded at particular positions to the basic framework of bonded atoms or to the ringed structures, the bridge structures, the alicyclic structures, or the aromatic structures.
A polypeptide or protein described herein, and that serves as a target in the screening procedure, includes the product of a naturally occurring polypeptide or precursor form or proprotein. The naturally occurring polypeptide, precursor or proprotein includes, e.g., the
full-length gene product, encoded by the corresponding gene. The naturally occurring polypeptide also includes the polypeptide, precursor or proprotein encoded by an open reading frame described herein. A "mature" form of a polypeptide or protein arises as a result of one or more naturally occurring processing steps as they may occur within the cell, including a host cell. The processing steps occur as the gene product arises, e.g., via cleavage of the amino-terminal methionine residue encoded by the initiation codon of an open reading frame, or the proteolytic cleavage of a signal peptide or leader sequence. Thus, a mature form arising from a precursor polypeptide or protein that has residues 1 to N, where residue 1 is the N- terminal methionine, would have residues 2 through N remaining. Alternatively, a mature form arising from a precursor polypeptide or protein having residues 1 to N, in which an amino-terminal signal sequence from residue 1 to residue M is cleaved, includes the residues from residue M+l to residue N remaining. A "mature" form of a polypeptide or protein may also arise from non-proteolytic post-translational modification. Such non-proteolytic processes include, e.g., glycosylation, myristylation or phosphorylation. In general, a mature polypeptide or protein may result from the operation of only one of these processes, or the combination of any of them.
As used herein, "identical" residues correspond to those residues in a comparison between two sequences where the equivalent nucleotide base or amino acid residue in an alignment of two sequences is the same residue. Residues are alternatively described as "similar" or "positive" when the comparisons between two sequences in an alignment show that residues in an equivalent position in a comparison are either the same amino acid or a conserved amino acid as defined below.
As used herein, a "chemical composition" relates to a composition including at least one compound that is either synthesized or extracted from a natural source. A chemical compound may be the product of a defined synthetic procedure. Such a synthesized compound is understood herein to have defined properties in terms of molecular formula, molecular structure relating the association of bonded atoms to each other, physical properties such as chromatographic or spectroscopic characterizations, and the like. A compound extracted from a natural source is advantageously analyzed by chemical and physical methods in order to provide a representation of its defined properties, including its molecular formula, molecular structure relating the association of bonded atoms to each other, physical properties such as chromatographic or spectroscopic characterizations, and the like.
As used herein, a "candidate therapeutic agent" is a chemical compound that includes at least one substance shown to bind to a target biopolymer. In important embodiments of the
invention, the target biopolymer is a protein or polypeptide, a nucleic acid, a polysaccharide or proteoglycan, or a lipid such as a complex lipid. The method of identifying compounds that bind to the target effectively eliminates compounds with little or no binding affinity, thereby increasing the potential that the identified chemical compound may have beneficial therapeutic applications. In cases where the "candidate therapeutic agent" is a mixture of more than one chemical compound, subsequent screening procedures may be carried out to identify the particular substance in the mixture that is the binding compound, and that is to be identified as a candidate therapeutic agent.
As used herein, a "pharmaceutical agent" is provided by screening a candidate therapeutic agent using models for a disease state or pathology in order to identify a candidate exerting a desired or beneficial therapeutic effect with relation to the disease or pathology. Such a candidate that successfully provides such an effect is termed a pharmaceutical agent herein. Nonlimiting examples of model systems that may be used in such screens include particular cell lines, cultured cells, tissue preparations, whole tissues, organ preparations, intact organs, and nonhuman mammals. Screens employing at least one system, and preferably more than one system, may be employed in order to identify a pharmaceutical agent. Any pharmaceutical agent so identified may be pursued in further investigation using human subjects.
NOVX Nucleic Acids and Polypeptides
NOVX clones
NOVX nucleic acids and their encoded polypeptides are useful in a variety of applications and contexts. The various NOVX nucleic acids and polypeptides according to the invention are useful as novel members of the protein families according to the presence of domains and sequence relatedness to previously described proteins. Additionally, NOVX nucleic acids and polypeptides can also be used to identify proteins that are members of the family to which the NOVX polypeptides belong.
The NOVX genes and their corresponding encoded proteins are useful for preventing, treating or ameliorating medical conditions, e.g., by protein or gene therapy. Pathological conditions can be diagnosed by determining the amount of the new protein in a sample or by determining the presence of mutations in the new genes. Specific uses are described for each of the NOVX genes, based on the tissues in which they are most highly expressed. Uses include developing products for the diagnosis or treatment of a variety of diseases and disorders.
The NOVX nucleic acids and proteins of the invention are useful in potential diagnostic and therapeutic applications and as a research tool. These include serving as a specific or selective nucleic acid or protein diagnostic and/or prognostic marker, wherein the presence or amount of the nucleic acid or the protein are to be assessed, as well as potential therapeutic applications such as the following: (i) a protein therapeutic, (ii) a small molecule drug target, (iii) an antibody target (therapeutic, diagnostic, drug targeting/cytotoxic antibody), (iv) a nucleic acid useful in gene therapy (gene delivery/gene ablation), and (v) a composition promoting tissue regeneration in vitro and in vivo (vi) biological defense weapon.
In one specific embodiment, the invention includes an isolated polypeptide comprising an amino acid sequence selected from the group consisting of: (a) a mature form of the amino acid sequence selected from the group consisting of SEQ ID NO: 2n, wherein n is an integer between 1 and 178; (b) a variant of a mature form of the amino acid sequence selected from the group consisting of SEQ ID NO: 2n, wherein n is an integer between 1 and 178, wherein any amino acid in the mature form is changed to a different amino acid, provided that no more than 15%) of the amino acid residues in the sequence of the mature form are so changed; (c) an amino acid sequence selected from the group consisting of SEQ ID NO: 2n, wherein n is an integer between 1 and 178; (d) a variant of the amino acid sequence selected from the group consisting of SEQ ID NO:2n, wherein n is an integer between 1 and 178 wherein any amino acid specified in the chosen sequence is changed to a different amino acid, provided that no more than 15%> of the amino acid residues in the sequence are so changed; and (e) a fragment of any of (a) through (d).
In another specific embodiment, the invention includes an isolated nucleic acid molecule comprising a nucleic acid sequence encoding a polypeptide comprising an amino acid sequence selected from the group consisting of: (a) a mature form of the amino acid sequence given SEQ ID NO: 2n, wherein n is an integer between 1 and 178; (b) a variant of a mature form of the amino acid sequence selected from the group consisting of SEQ ID NO: 2n, wherein n is an integer between 1 and 178 wherein any amino acid in the mature form of the chosen sequence is changed to a different amino acid, provided that no more than 15%> of the amino acid residues in the sequence of the mature form are so changed; (c) the amino acid sequence selected from the group consisting of SEQ ID NO: 2n, wherein n is an integer between 1 and 178; (d) a variant of the amino acid sequence selected from the group consisting of SEQ ID NO: 2n, wherein n is an integer between 1 and 178, in which any amino acid specified in the chosen sequence is changed to a different amino acid, provided that no more than 15%> of the amino acid residues in the sequence are so changed; (e) a nucleic acid
fragment encoding at least a portion of a polypeptide comprising the amino acid sequence selected from the group consisting of SEQ ID NO: 2n, wherein n is an integer between 1 and 178 or any variant of said polypeptide wherein any amino acid of the chosen sequence is changed to a different amino acid, provided that no more than 10%> of the amino acid residues in the sequence are so changed; and (f) the complement of any of said nucleic acid molecules.
In yet another specific embodiment, the invention includes an isolated nucleic acid molecule, wherein said nucleic acid molecule comprises a nucleotide sequence selected from the group consisting of: (a) the nucleotide sequence selected from the group consisting of SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178; (b) a nucleotide sequence wherein one or more nucleotides in the nucleotide sequence selected from the group consisting of SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178 is changed from that selected from the group consisting of the chosen sequence to a different nucleotide provided that no more than 15%> of the nucleotides are so changed; (c) a nucleic acid fragment of the sequence selected from the group consisting of SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178; and (d) a nucleic acid fragment wherein one or more nucleotides in the nucleotide sequence selected from the group consisting of SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178 is changed from that selected from the group consisting of the chosen sequence to a different nucleotide provided that no more than 15% of the nucleotides are so changed.
One aspect of the invention pertains to isolated nucleic acid molecules that encode NOVX polypeptides or biologically active portions thereof. Also included in the invention are nucleic acid fragments sufficient for use as hybridization probes to identify NOVX-encoding nucleic acids (e.g., NOVX mRNAs) and fragments for use as PCR primers for the amplification and/or mutation of NOVX nucleic acid molecules. As used herein, the term "nucleic acid molecule" is intended to include DNA molecules (e.g., cDNA or genomic DNA), RNA molecules (e.g., mRNA), analogs of the DNA or RNA generated using nucleotide analogs, and derivatives, fragments and homologs thereof. The nucleic acid molecule may be single-stranded or double-stranded, but preferably is comprised double- stranded DNA.
An NOVX nucleic acid can encode a mature NOVX polypeptide. As used herein, a "mature" form of a polypeptide or protein disclosed in the present invention is the product of a naturally occurring polypeptide or precursor form or proprotein. The naturally occurring polypeptide, precursor or proprotein includes, by way of nonlimiting example, the full-length gene product, encoded by the corresponding gene. Alternatively, it may be defined as the
polypeptide, precursor or proprotein encoded by an ORF described herein. The product "mature" form arises, again by way of nonlimiting example, as a result of one or more naturally occurring processing steps as they may take place within the cell, or host cell, in which the gene product arises. Examples of such processing steps leading to a "mature" form of a polypeptide or protein include the cleavage of the N-terminal methionine residue encoded by the initiation codon of an ORF, or the proteolytic cleavage of a signal peptide or leader sequence. Thus a mature form arising from a precursor polypeptide or protein that has residues 1 to N, where residue 1 is the N-terminal methionine, would have residues 2 through N remaining after removal of the N-terminal methionine. Alternatively, a mature form arising from a precursor polypeptide or protein having residues 1 to N, in which an N-terminal signal sequence from residue 1 to residue M is cleaved, would have the residues from residue M+1 to residue N remaining. Further as used herein, a "mature" form of a polypeptide or protein may arise from a step of post-translational modification other than a proteolytic cleavage event. Such additional processes include, by way of non- limiting example, glycosylation, myristoylation or phosphorylation. In general, a mature polypeptide or protein may result from the operation of only one of these processes, or a combination of any of them.
The term "probes", as utilized herein, refers to nucleic acid sequences of variable length, preferably between at least about 10 nucleotides (nt), 100 nt, or as many as approximately, e.g., 6,000 nt, depending upon the specific use. Probes are used in the detection of identical, similar, or complementary nucleic acid sequences. Longer length probes are generally obtained from a natural or recombinant source, are highly specific, and much slower to hybridize than shorter-length oligomer probes. Probes may be single- or double-stranded and designed to have specificity in PCR, membrane-based hybridization technologies, or ELIS A-like technologies.
The term "isolated" nucleic acid molecule, as utilized herein, is one, which is separated from other nucleic acid molecules which are present in the natural source of the nucleic acid. Preferably, an "isolated" nucleic acid is free of sequences which naturally flank the nucleic acid (i.e., sequences located at the 5'- and 3'-termini of the nucleic acid) in the genomic DNA of the organism from which the nucleic acid is derived. For example, in various embodiments, the isolated NOVX nucleic acid molecules can contain less than about 5 kb, 4 kb, 3 kb, 2 kb, 1 kb, 0.5 kb or 0.1 kb of nucleotide sequences which naturally flank the nucleic acid molecule in genomic DNA of the cell/tissue from which the nucleic acid is derived (e.g., brain, heart, liver, spleen, etc.). Moreover, an "isolated" nucleic acid molecule, such as a cDNA molecule, can be substantially free of other cellular material or culture medium when produced by
recombinant techniques, or of chemical precursors or other chemicals when chemically synthesized.
A nucleic acid molecule of the invention, e.g., a nucleic acid molecule having the nucleotide sequence SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178, or a complement of this aforementioned nucleotide sequence, can be isolated using standard molecular biology techniques and the sequence information provided herein. Using all or a portion of the nucleic acid sequence of SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178 as a hybridization probe, NOVX molecules can be isolated using standard hybridization and cloning techniques (e.g., as described in Sambrook, et αl, (eds.), MOLECULAR CLONING: A LABORATORY MANUAL 2nd Ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989; and Ausubel, et αl., (eds.), CURRENT PROTOCOLS IN MOLECULAR BIOLOGY, John Wiley & Sons, New York, NY, 1993.)
A nucleic acid of the invention can be amplified using cDNA, mRNA or alternatively, genomic DNA, as a template and appropriate oligonucleotide primers according to standard PCR amplification techniques. The nucleic acid so amplified can be cloned into an appropriate vector and characterized by DNA sequence analysis. Furthermore, oligonucleotides corresponding to NOVX nucleotide sequences can be prepared by standard synthetic techniques, e.g. , using an automated DNA synthesizer.
As used herein, the term "oligonucleotide" refers to a series of linked nucleotide residues, which oligonucleotide has a sufficient number of nucleotide bases to be used in a PCR reaction. A short oligonucleotide sequence may be based on, or designed from, a genomic or cDNA sequence and is used to amplify, confirm, or reveal the presence of an identical, similar or complementary DNA or RNA in a particular cell or tissue. Oligonucleotides comprise portions of a nucleic acid sequence having about 10 nt, 50 nt, or 100 nt in length, preferably about 15 nt to 30 nt in length. In one embodiment of the invention, an oligonucleotide comprising a nucleic acid molecule less than 100 nt in length would further comprise at least 6 contiguous nucleotides SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178, or a complement thereof. Oligonucleotides may be chemically synthesized and may also be used as probes.
In another embodiment, an isolated nucleic acid molecule of the invention comprises a nucleic acid molecule that is a complement of the nucleotide from the group consisting of SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178, or a portion of this nucleotide sequence (e.g., a fragment that can be used as a probe or primer or a fragment encoding a biologically-active portion of an NOVX polypeptide). A nucleic acid molecule that is
complementary to the nucleotide sequence from the group consisting of SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178 is one that is sufficiently complementary to the nucleotide sequence from the group consisting of SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178 that it can hydrogen bond with little or no mismatches to the nucleotide sequence from the group consisting of SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178, thereby forming a stable duplex.
As used herein, the term "complementary" refers to Watson-Crick or Hoogsteen base pairing between nucleotides units of a nucleic acid molecule, and the term "binding" means the physical or chemical interaction between two polypeptides or compounds or associated polypeptides or compounds or combinations thereof. Binding includes ionic, non-ionic, van der Waals, hydrophobic interactions, and the like. A physical interaction can be either direct or indirect. Indirect interactions may be through or due to the effects of another polypeptide or compound. Direct binding refers to interactions that do not take place through, or due to, the effect of another polypeptide or compound, but instead are without other substantial chemical intermediates.
Fragments provided herein are defined as sequences of at least 6 (contiguous) nucleic acids or at least 4 (contiguous) amino acids, a length sufficient to allow for specific hybridization in the case of nucleic acids or for specific recognition of an epitope in the case of amino acids, respectively, and are at most some portion less than a full length sequence. Fragments may be derived from any contiguous portion of a nucleic acid or amino acid sequence of choice. Derivatives are nucleic acid sequences or amino acid sequences formed from the native compounds either directly or by modification or partial substitution. Analogs are nucleic acid sequences or amino acid sequences that have a structure similar to, but not identical to, the native compound but differs from it in respect to certain components or side chains. Analogs may be synthetic or from a different evolutionary origin and may have a similar or opposite metabolic activity compared to wild type. Homologs are nucleic acid sequences or amino acid sequences of a particular gene that are derived from different species.
A full-length NOVX clone is identified as containing an ATG translation start codon and an in-frame stop codon. Any disclosed NOVX nucleotide sequence lacking an ATG start codon therefore encodes a truncated C-terminal fragment of the respective NOVX polypeptide, and requires that the corresponding full-length cDNA extend in the 5' direction of the disclosed sequence. Any disclosed NOVX nucleotide sequence lacking an in- frame stop codon similarly encodes a truncated N-terminal fragment of the respective NOVX
polypeptide, and requires that the corresponding full-length cDNA extend in the 3' direction of the disclosed sequence.
Derivatives and analogs may be full length or other than full length, if the derivative or analog contains a modified nucleic acid or amino acid, as described below. Derivatives or analogs of the nucleic acids or proteins of the invention include, but are not limited to, molecules comprising regions that are substantially homologous to the nucleic acids or proteins of the invention, in various embodiments, by at least about 70%>, 80%>, or 95%> identity (with a preferred identity of 80-95%) over a nucleic acid or amino acid sequence of identical size or when compared to an aligned sequence in which the alignment is done by a computer homology program known in the art, or whose encoding nucleic acid is capable of hybridizing to the complement of a sequence encoding the aforementioned proteins under stringent, moderately stringent, or low stringent conditions. See e.g. Ausubel, et al, CURRENT PROTOCOLS IN MOLECULAR BIOLOGY, John Wiley & Sons, New York, NY, 1993, and below.
A "homologous nucleic acid sequence" or "homologous amino acid sequence," or variations thereof, refer to sequences characterized by a homology at the nucleotide level or amino acid level as discussed above. Homologous nucleotide sequences encode those sequences coding for isoforms of NOVX polypeptides. Isoforms can be expressed in different tissues of the same organism as a result of, for example, alternative splicing of RNA. Alternatively, isoforms can be encoded by different genes. In the invention, homologous nucleotide sequences include nucleotide sequences encoding for an NOVX polypeptide of species other than humans, including, but not limited to: vertebrates, and thus can include, e.g., frog, mouse, rat, rabbit, dog, cat cow, horse, and other organisms. Homologous nucleotide sequences also include, but are not limited to, naturally occurring allelic variations and mutations of the nucleotide sequences set forth herein. A homologous nucleotide sequence does not, however, include the exact nucleotide sequence encoding human NOVX protein. Homologous nucleic acid sequences include those nucleic acid sequences that encode conservative amino acid substitutions (see below) in SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178, as well as a polypeptide possessing NOVX biological activity. Various biological activities of the NOVX proteins are described below.
An NOVX polypeptide is encoded by the open reading frame ("ORF") of an NOVX nucleic acid. An ORF corresponds to a nucleotide sequence that could potentially be translated into a polypeptide. A stretch of nucleic acids comprising an ORF is uninterrupted by a stop codon. An ORF that represents the coding sequence for a full protein begins with an ATG "start" codon and terminates with one of the three "stop" codons, namely, TAA, TAG, or
TGA. For the puφoses of this invention, an ORF may be any part of a coding sequence, with or without a start codon, a stop codon, or both. For an ORF to be considered as a good candidate for coding for a bonafide cellular protein, a minimum size requirement is often set, e.g., a stretch of DNA that would encode a protein of 50 amino acids or more.
The nucleotide sequences determined from the cloning of the human NOVX genes allows for the generation of probes and primers designed for use in identifying and/or cloning NOVX homologues in other cell types, e.g. from other tissues, as well as NOVX homologues from other vertebrates. The probe/primer typically comprises substantially purified oligonucleotide. The oligonucleotide typically comprises a region of nucleotide sequence that hybridizes under stringent conditions to at least about 12, 25, 50, 100, 150, 200, 250, 300, 350 or 400 consecutive sense strand nucleotide sequence SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178; or an anti-sense strand nucleotide sequence of SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178.
Probes based on the human NOVX nucleotide sequences can be used to detect transcripts or genomic sequences encoding the same or homologous proteins. In various embodiments, the probe further comprises a label group attached thereto, e.g. the label group can be a radioisotope, a fluorescent compound, an enzyme, or an enzyme co-factor. Such probes can be used as a part of a diagnostic test kit for identifying cells or tissues which mis- express an NOVX protein, such as by measuring a level of an NOVX-encoding nucleic acid in a sample of cells from a subject e.g., detecting NOVX mRNA levels or determining whether a genomic NOVX gene has been mutated or deleted.
"A polypeptide having a biologically-active portion of an NOVX polypeptide" refers to polypeptides exhibiting activity similar, but not necessarily identical to, an activity of a polypeptide of the invention, including mature forms, as measured in a particular biological assay, with or without dose dependency. A nucleic acid fragment encoding a "biologically- active portion of NOVX" can be prepared by isolating a portion SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178, that encodes a polypeptide having an NOVX biological activity (the biological activities of the NOVX proteins are described below), expressing the encoded portion of NOVX protein (e.g., by recombinant expression in vitro) and assessing the activity of the encoded portion of NOVX.
NOVX Nucleic Acid and Polypeptide Variants
The invention further encompasses nucleic acid molecules that differ from the nucleotide sequences shown in SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178
due to degeneracy of the genetic code and thus encode the same NOVX proteins as that encoded by the nucleotide sequences shown in SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178. In another embodiment, an isolated nucleic acid molecule of the invention has a nucleotide sequence encoding a protein having an amino acid sequence shown in SEQ ID NO: 2n, wherein n is an integer between 1 and 178.
In addition to the human NOVX nucleotide sequences shown in SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178, it will be appreciated by those skilled in the art that DNA sequence polymoφhisms that lead to changes in the amino acid sequences of the NOVX polypeptides may exist within a population (e.g., the human population). Such genetic polymoφhism in the NOVX genes may exist among individuals within a population due to natural allelic variation. As used herein, the terms "gene" and "recombinant gene" refer to nucleic acid molecules comprising an open reading frame (ORF) encoding an NOVX protein, preferably a vertebrate NOVX protein. Such natural allelic variations can typically result in l-5%o variance in the nucleotide sequence of the NOVX genes. Any and all such nucleotide variations and resulting amino acid polymoφhisms in the NOVX polypeptides, which are the result of natural allelic variation and that do not alter the functional activity of the NOVX polypeptides, are intended to be within the scope of the invention.
Moreover, nucleic acid molecules encoding NOVX proteins from other species, and thus that have a nucleotide sequence that differs from the human SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178 are intended to be within the scope of the invention. Nucleic acid molecules corresponding to natural allelic variants and homologues of the NOVX cDNAs of the invention can be isolated based on their homology to the human NOVX nucleic acids disclosed herein using the human cDNAs, or a portion thereof, as a hybridization probe according to standard hybridization techniques under stringent hybridization conditions.
Accordingly, in another embodiment, an isolated nucleic acid molecule of the invention is at least 6 nucleotides in length and hybridizes under stringent conditions to the nucleic acid molecule comprising the nucleotide sequence of SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178. In another embodiment, the nucleic acid is at least 10, 25, 50, 100, 250, 500, 750, 1000, 1500, or 2000 or more nucleotides in length. In yet another embodiment, an isolated nucleic acid molecule of the invention hybridizes to the coding region. As used herein, the term "hybridizes under stringent conditions" is intended to describe conditions for hybridization and washing under which nucleotide sequences at least 60%o homologous to each other typically remain hybridized to each other.
Homologs (i.e., nucleic acids encoding NOVX proteins derived from species other than human) or other related sequences (e.g., paralogs) can be obtained by low, moderate or high stringency hybridization with all or a portion of the particular human sequence as a probe using methods well known in the art for nucleic acid hybridization and cloning.
As used herein, the phrase "stringent hybridization conditions" refers to conditions under which a probe, primer or oligonucleotide will hybridize to its target sequence, but to no other sequences. Stringent conditions are sequence-dependent and will be different in different circumstances. Longer sequences hybridize specifically at higher temperatures than shorter sequences. Generally, stringent conditions are selected to be about 5 °C lower than the thermal melting point (Tm) for the specific sequence at a defined ionic strength and pH. The Tm is the temperature (under defined ionic strength, pH and nucleic acid concentration) at which 50% of the probes complementary to the target sequence hybridize to the target sequence at equilibrium. Since the target sequences are generally present at excess, at Tm, 50%) of the probes are occupied at equilibrium. Typically, stringent conditions will be those in which the salt concentration is less than about 1.0 M sodium ion, typically about 0.01 to 1.0 M sodium ion (or other salts) at pH 7.0 to 8.3 and the temperature is at least about 30°C for short probes, primers or oligonucleotides (e.g., 10 nt to 50 nt) and at least about 60°C for longer probes, primers and oligonucleotides. Stringent conditions may also be achieved with the addition of destabilizing agents, such as formamide.
Stringent conditions are known to those skilled in the art and can be found in Ausubel, et al, (eds.), CURRENT PROTOCOLS IN MOLECULAR BIOLOGY, John Wiley & Sons, N.Y. (1989), 6.3.1-6.3.6. Preferably, the conditions are such that sequences at least about 65%>, 70%, 75%o, 85%, 90%o, 95%., 98%, or 99%. homologous to each other typically remain hybridized to each other. A non-limiting example of stringent hybridization conditions are hybridization in a high salt buffer comprising 6X SSC, 50 mM Tris-HCl (pH 7.5), 1 mM EDTA, 0.02%) PVP, 0.02% Ficoll, 0.02%. BSA, and 500 mg/ml denatured salmon sperm DNA at 65°C, followed by one or more washes in 0.2X SSC, 0.01% BSA at 50°C. An isolated nucleic acid molecule of the invention that hybridizes under stringent conditions to the sequences SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178, corresponds to a naturally-occurring nucleic acid molecule. As used herein, a "naturally-occurring" nucleic acid molecule refers to an RNA or DNA molecule having a nucleotide sequence that occurs in nature (e.g., encodes a natural protein).
In a second embodiment, a nucleic acid sequence that is hybridizable to the nucleic acid molecule comprising the nucleotide sequence of SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178, or fragments, analogs or derivatives thereof, under conditions of moderate stringency is provided. A non-limiting example of moderate stringency hybridization conditions are hybridization in 6X SSC, 5X Denhardt's solution, 0.5%. SDS and 100 mg/ml denatured salmon sperm DNA at 55°C, followed by one or more washes in IX SSC, 0.1% SDS at 37°C. Other conditions of moderate stringency that may be used are well-known within the art. See, e.g., Ausubel, et al. (eds.), 1993, CURRENT PROTOCOLS IN MOLECULAR BIOLOGY, John Wiley & Sons, NY, and Kriegler, 1990; GENE TRANSFER AND EXPRESSION, A LABORATORY MANUAL, Stockton Press, NY.
In a third embodiment, a nucleic acid that is hybridizable to the nucleic acid molecule comprising the nucleotide sequences SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178, or fragments, analogs or derivatives thereof, under conditions of low stringency, is provided. A non-limiting example of low stringency hybridization conditions are hybridization in 35% formamide, 5X SSC, 50 mM Tris-HCl (pH 7.5), 5 mM EDTA, 0.02% PVP, 0.02% Ficoll, 0.2% BSA, 100 mg/ml denatured salmon sperm DNA, 10% (wt/vol) dextran sulfate at 40°C, followed by one or more washes in 2X SSC, 25 mM Tris-HCl (pH 7.4), 5 mM EDTA, and 0.1%> SDS at 50°C. Other conditions of low stringency that may be used are well known in the art (e.g., as employed for cross-species hybridizations). See, e.g., Ausubel, et al. (eds.), 1993, CURRENT PROTOCOLS IN MOLECULAR BIOLOGY, John Wiley & Sons, NY, and Kriegler, 1990, GENE TRANSFER AND EXPRESSION, A LABORATORY MANUAL, Stockton Press, NY; Shilo and Weinberg, 1981. Proc Natl Acad Sci USA 78: 6789-6792.
Conservative Mutations
In addition to naturally-occurring allelic variants of NOVX sequences that may exist in the population, the skilled artisan will further appreciate that changes can be introduced by mutation into the nucleotide sequences SEQ ED NO: 2n-l, wherein n is an integer between 1 and 178, thereby leading to changes in the amino acid sequences of the encoded NOVX proteins, without altering the functional ability of said NOVX proteins. For example, nucleotide substitutions leading to amino acid substitutions at "non-essential" amino acid residues can be made in the sequence SEQ ID NO: 2n, wherein n is an integer between 1 and 178. A "non-essential" amino acid residue is a residue that can be altered from the wild-type sequences of the NOVX proteins without altering their biological activity, whereas an
"essential" amino acid residue is required for such biological activity. For example, amino acid residues that are conserved among the NOVX proteins of the invention are predicted to be particularly non-amenable to alteration. Amino acids for which conservative substitutions can be made are well-known within the art.
Another aspect of the invention pertains to nucleic acid molecules encoding NOVX proteins that contain changes in amino acid residues that are not essential for activity. Such NOVX proteins differ in amino acid sequence from SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178 yet retain biological activity. In one embodiment, the isolated nucleic acid molecule comprises a nucleotide sequence encoding a protein, wherein the protein comprises an amino acid sequence at least about 45% homologous to the amino acid sequences SEQ ID NO: 2n, wherein n is an integer between 1 and 178. Preferably, the protein encoded by the nucleic acid molecule is at least about 60% homologous to SEQ ID NO: 2n, wherein n is an integer between 1 and 178; more preferably at least about 70%> homologous SEQ ID NO: 2n, wherein n is an integer between 1 and 178; still more preferably at least about 80%> homologous to SEQ ID NO: 2n, wherein n is an integer between 1 and 178; even more preferably at least about 90%. homologous to SEQ ID NO: 2n, wherein n is an integer between 1 and 178; and most preferably at least about 95% homologous to SEQ ID NO: 2n, wherein n is an integer between 1 and 178.
An isolated nucleic acid molecule encoding an NOVX protein homologous to the protein of SEQ ID NO: 2n, wherein n is an integer between 1 and 178 can be created by introducing one or more nucleotide substitutions, additions or deletions into the nucleotide sequence of SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178, such that one or more amino acid substitutions, additions or deletions are introduced into the encoded protein.
Mutations can be introduced into SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178 standard techniques, such as site-directed mutagenesis and PCR-mediated mutagenesis. Preferably, conservative amino acid substitutions are made at one or more predicted, non-essential amino acid residues. A "conservative amino acid substitution" is one in which the amino acid residue is replaced with an amino acid residue having a similar side chain. Families of amino acid residues having similar side chains have been defined within the art. These families include amino acids with basic side chains (e.g., lysine, arginine, histidine), acidic side chains (e.g., aspartic acid, glutamic acid), uncharged polar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, cysteine), nonpolar side chains (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan), beta-branched side chains (e.g., threonine, valine, isoleucine) and aromatic side
chains (e.g., tyrosine, phenylalanine, tryptophan, histidine). Thus, a predicted non-essential amino acid residue in the NOVX protein is replaced with another amino acid residue from the same side chain family. Alternatively, in another embodiment, mutations can be introduced randomly along all or part of an NOVX coding sequence, such as by saturation mutagenesis, and the resultant mutants can be screened for NOVX biological activity to identify mutants that retain activity. Following mutagenesis SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178, the encoded protein can be expressed by any recombinant technology known in the art and the activity of the protein can be determined.
The relatedness of amino acid families may also be determined based on side chain interactions. Substituted amino acids may be fully conserved "strong" residues or fully conserved "weak" residues. The "strong" group of conserved amino acid residues may be any one of the following groups: STA, NEQK, NHQK, NDEQ, QHRK, MILV, MILF, HY, FYW, wherein the single letter amino acid codes are grouped by those amino acids that may be substituted for each other. Likewise, the "weak" group of conserved residues may be any one of the following: CSA, ATV, SAG, STNK, STPA, SGND, SNDEQK, NDEQHK, NEQHRK, HFY, wherein the letters within each group represent the single letter amino acid code.
In one embodiment, a mutant NOVX protein can be assayed for ( ) the ability to form proteimprotein interactions with other NOVX proteins, other cell-surface proteins, or biologically-active portions thereof, (ii) complex formation between a mutant NOVX protein and an NOVX ligand; or (iii) the ability of a mutant NOVX protein to bind to an intracellular target protein or biologically-active portion thereof; (e.g. avidin proteins).
In yet another embodiment, a mutant NOVX protein can be assayed for the ability to regulate a specific biological function (e.g., regulation of insulin release).
Antisense Nucleic Acids
Another aspect of the invention pertains to isolated antisense nucleic acid molecules that are hybridizable to or complementary to the nucleic acid molecule comprising the nucleotide sequence of SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178, or fragments, analogs or derivatives thereof. An "antisense" nucleic acid comprises a nucleotide sequence that is complementary to a "sense" nucleic acid encoding a protein (e.g., complementary to the coding strand of a double-stranded cDNA molecule or complementary to an mRNA sequence). In specific aspects, antisense nucleic acid molecules are provided that comprise a sequence complementary to at least about 10, 25, 50, 100, 250 or 500 nucleotides or an entire NOVX coding strand, or to only a portion thereof. Nucleic acid molecules
encoding fragments, homologs, derivatives and analogs of an NOVX protein of SEQ ID NO: 2n, wherein n is an integer between 1 and 178, or antisense nucleic acids complementary to an NOVX nucleic acid sequence of SEQ ED NO: 2n-l, wherein n is an integer between 1 and 178, are additionally provided.
In one embodiment, an antisense nucleic acid molecule is antisense to a "coding region" of the coding strand of a nucleotide sequence encoding an NOVX protein. The term "coding region" refers to the region of the nucleotide sequence comprising codons which are translated into amino acid residues. In another embodiment, the antisense nucleic acid molecule is antisense to a "noncoding region" of the coding strand of a nucleotide sequence encoding the NOVX protein. The term "noncoding region" refers to 5' and 3' sequences which flank the coding region that are not translated into amino acids (i.e., also referred to as 5' and 3' untranslated regions).
Given the coding strand sequences encoding the NOVX protein disclosed herein, antisense nucleic acids of the invention can be designed according to the rules of Watson and Crick or Hoogsteen base pairing. The antisense nucleic acid molecule can be complementary to the entire coding region of NOVX mRNA, but more preferably is an oligonucleotide that is antisense to only a portion of the coding or noncoding region of NOVX mRNA. For example, the antisense oligonucleotide can be complementary to the region surrounding the translation start site of NOVX mRNA. An antisense oligonucleotide can be, for example, about 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50 nucleotides in length. An antisense nucleic acid of the invention can be constructed using chemical synthesis or enzymatic ligation reactions using procedures known in the art. For example, an antisense nucleic acid (e.g., an antisense oligonucleotide) can be chemically synthesized using naturally-occurring nucleotides or variously modified nucleotides designed to increase the biological stability of the molecules or to increase the physical stability of the duplex formed between the antisense and sense nucleic acids (e.g., phosphorothioate derivatives and acridine substituted nucleotides can be used).
Examples of modified nucleotides that can be used to generate the antisense nucleic acid include: 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxylmethyl) uracil, 5-carboxymethylaminomethyl- 2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, beta-D-galactosylqueosine, inosine, N6-isopentenyladenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-mefhylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, beta-D-marmosylqueosine, 5'-methoxycarboxyrnethyluracil, 5-mefhoxyuracil,
2-methylthio-N6-isopentenyladenine, uracil-5-oxyacetic acid (v), wybutoxosine, pseudouracil, queosine, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetic acid methylester, uracil-5-oxyacetic acid (v), 5-mefhyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl) uracil, (acp3)w, and 2,6-diaminopurine. Alternatively, the antisense nucleic acid can be produced biologically using an expression vector into which a nucleic acid has been subcloned in an antisense orientation (i.e., RNA transcribed from the inserted nucleic acid will be of an antisense orientation to a target nucleic acid of interest, described further in the following subsection).
The antisense nucleic acid molecules of the invention are typically administered to a subject or generated in situ such that they hybridize with or bind to cellular mRNA and/or genomic DNA encoding an NOVX protein to thereby inhibit expression of the protein (e.g., by inhibiting transcription and/or translation). The hybridization can be by conventional nucleotide complementarity to form a stable duplex, or, for example, in the case of an antisense nucleic acid molecule that binds to DNA duplexes, through specific interactions in the major groove of the double helix. An example of a route of administration of antisense nucleic acid molecules of the invention includes direct injection at a tissue site. Alternatively, antisense nucleic acid molecules can be modified to target selected cells and then administered systemically. For example, for systemic administration, antisense molecules can be modified such that they specifically bind to receptors or antigens expressed on a selected cell surface (e.g., by linking the antisense nucleic acid molecules to peptides or antibodies that bind to cell surface receptors or antigens). The antisense nucleic acid molecules can also be delivered to cells using the vectors described herein. To achieve sufficient nucleic acid molecules, vector constructs in which the antisense nucleic acid molecule is placed under the control of a strong pol II or pol III promoter are preferred.
In yet another embodiment, the antisense nucleic acid molecule of the invention is an α-anomeric nucleic acid molecule. An α-anomeric nucleic acid molecule forms specific double-stranded hybrids with complementary RNA in which, contrary to the usual β-units, the strands run parallel to each other. See, e.g., Gaultier, et al, 1987. Nucl. Acids Res. 15: 6625-6641. The antisense nucleic acid molecule can also comprise a 2'-o-methylribonucleotide (See, e.g., Inoue, et al. 1987. Nucl. Acids Res. 15: 6131-6148) or a chimeric RNA-DNA analogue (See, e.g., Inoue, et al, 1987. FEBS Lett. 215: 327-330.
Ribozymes and PNA Moieties
Nucleic acid modifications include, by way of non-limiting example, modified bases, and nucleic acids whose sugar phosphate backbones are modified or derivatized. These modifications are carried out at least in part to enhance the chemical stability of the modified nucleic acid, such that they may be used, for example, as antisense binding nucleic acids in therapeutic applications in a subject.
In one embodiment, an antisense nucleic acid of the invention is a ribozyme. Ribozymes are catalytic RNA molecules with ribonuclease activity that are capable of cleaving a single-stranded nucleic acid, such as an mRNA, to which they have a complementary region. Thus, ribozymes (e.g., hammerhead ribozymes as described in Haselhoff and Gerlach 1988. Nature 334: 585-591) can be used to catalytically cleave NOVX mRNA transcripts to thereby inhibit translation of NOVX mRNA. A ribozyme having specificity for an NOVX-encoding nucleic acid can be designed based upon the nucleotide sequence of an NOVX cDNA disclosed herein (i.e., SEQ ED NO: 2n-l, wherein n is an integer between 1 and 178). For example, a derivative of a Tetrahymena L-19 IVS RNA can be constructed in which the nucleotide sequence of the active site is complementary to the nucleotide sequence to be cleaved in an NOVX-encoding mRNA. See, e.g., U.S. Patent 4,987,071 to Cech, et al. and U.S. Patent 5,116,742 to Cech, et al. NOVX mRNA can also be used to select a catalytic RNA having a specific ribonuclease activity from a pool of RNA molecules. See, e.g., Barrel et al, (1993) Science 261:1411-1418.
Alternatively, NOVX gene expression can be inhibited by targeting nucleotide sequences complementary to the regulatory region of the NOVX nucleic acid (e.g., the NOVX promoter and/or enhancers) to form triple helical structures that prevent transcription of the NOVX gene in target cells. See, e.g., Helene, 1991. Anticancer Drug Des. 6: 569-84; Helene, et al. 1992. Ann. NY. Acad. Sci. 660: 27-36; Maher, 1992. Bioassays 14: 807-15.
In various embodiments, the NOVX nucleic acids can be modified at the base moiety, sugar moiety or phosphate backbone to improve, e.g., the stability, hybridization, or solubility of the molecule. For example, the deoxyribose phosphate backbone of the nucleic acids can be modified to generate peptide nucleic acids. See, e.g., Hyrup, et al, 1996. BioorgMed Chem 4: 5-23. As used herein, the terms "peptide nucleic acids" or "PNAs" refer to nucleic acid mimics (e.g., DNA mimics) in which the deoxyribose phosphate backbone is replaced by a pseudopeptide backbone and only the four natural nucleobases are retained. The neutral backbone of PNAs has been shown to allow for specific hybridization to DNA and RNA under conditions of low ionic strength. The synthesis of PNA oligomers can be performed using
standard solid phase peptide synthesis protocols as described in Hyrup, et al, 1996. supra; Perry-O'Keefe, et al, 1996. Proc. Natl. Acad. Sci. USA 93: 14670-14675.
PNAs of NOVX can be used in therapeutic and diagnostic applications. For example, PNAs can be used as antisense or anti gene agents for sequence-specific modulation of gene expression by, e.g. , inducing transcription or translation arrest or inhibiting replication. PNAs of NOVX can also be used, for example, in the analysis of single base pair mutations in a gene (e.g., PNA directed PCR clamping; as artificial restriction enzymes when used in combination with other enzymes, e.g., Si nucleases (See, Hyrup, et al, 1996.supra); or as probes or primers for DNA sequence and hybridization (See, Hyrup, et al, 1996, supra; Perry-O'Keefe, et al, 1996. supra).
In another embodiment, PNAs of NOVX can be modified, e.g., to enhance their stability or cellular uptake, by attaching lipophilic or other helper groups to PNA, by the formation of PNA-DNA chimeras, or by the use of liposomes or other techniques of drug delivery known in the art. For example, PNA-DNA chimeras of NOVX can be generated that may combine the advantageous properties of PNA and DNA. Such chimeras allow DNA recognition enzymes (e.g., RNase H and DNA polymerases) to interact with the DNA portion while the PNA portion would provide high binding affinity and specificity. PNA-DNA chimeras can be linked using linkers of appropriate lengths selected in terms of base stacking, number of bonds between the nucleobases, and orientation (see, Hyrup, et al., 1996. supra). The synthesis of PNA-DNA chimeras can be performed as described in Hyrup, et al, 1996. supra and Finn, et al, 1996. Nucl Acids Res 24: 3357-3363. For example, a DNA chain can be synthesized on a solid support using standard phosphoramidite coupling chemistry, and modified nucleoside analogs, e.g., 5'-(4-methoxytrityl)amino-5'-deoxy-thymidine phosphoramidite, can be used between the PNA and the 5' end of DNA. See, e.g., Mag, et al, 1989. Nucl Acid Res 17: 5973-5988. PNA monomers are then coupled in a stepwise manner to produce a chimeric molecule with a 5' PNA segment and a 3' DNA segment. See, e.g., Finn, et al, 1996. supra. Alternatively, chimeric molecules can be synthesized with a 5' DNA segment and a 3' PNA segment. See, e.g., Petersen, et al, 1975. Bioorg. Med. Chem. Lett. 5: 1119-11124.
In other embodiments, the oligonucleotide may include other appended groups such as peptides (e.g. , for targeting host cell receptors in vivo), or agents facilitating transport across the cell membrane (see, e.g., Letsinger, et al, 1989. Proc. Natl. Acad. Sci. U.S.A. 86: 6553-6556; Lemaitre, et al, 1987. Proc. Natl. Acad. Sci. 84: 648-652; PCT Publication No. WO88/09810) or the blood-brain barrier (see, e.g., PCT Publication No. WO 89/10134). In
addition, oligonucleotides can be modified with hybridization triggered cleavage agents (see, e.g., Krol, et al., 1988. BioTechniques 6:958-976) or intercalating agents (see, e.g., Zon, 1988. Pharm. Res. 5: 539-549). To this end, the oligonucleotide may be conjugated to another molecule, e.g., a peptide, a hybridization triggered cross-linking agent, a transport agent, a hybridization-triggered cleavage agent, and the like.
NOVX Polypeptides
A polypeptide according to the invention includes a polypeptide including the amino acid sequence of NOVX polypeptides whose sequences are provided in SEQ TD NO: 2n, wherein n is an integer between 1 and 178. The invention also includes a mutant or variant protein any of whose residues may be changed from the corresponding residues shown in SEQ ED NO: 2n, wherein n is an integer between 1 and 178 while still encoding a protein that maintains its NOVX activities and physiological functions, or a functional fragment thereof.
In general, an NOVX variant that preserves NOVX-like function includes any variant in which residues at a particular position in the sequence have been substituted by other amino acids, and further include the possibility of inserting an additional residue or residues between two residues of the parent protein as well as the possibility of deleting one or more residues from the parent sequence. Any amino acid substitution, insertion, or deletion is encompassed by the invention. In favorable circumstances, the substitution is a conservative substitution as defined above.
One aspect of the invention pertains to isolated NOVX proteins, and biologically- active portions thereof, or derivatives, fragments, analogs or homologs thereof. Also provided are polypeptide fragments suitable for use as immunogens to raise anti-NOVX antibodies. In one embodiment, native NONX proteins can be isolated from cells or tissue sources by an appropriate purification scheme using standard protein purification techniques. In another embodiment, ΝOVX proteins are produced by recombinant DΝA techniques. Alternative to recombinant expression, an ΝOVX protein or polypeptide can be synthesized chemically using standard peptide synthesis techniques.
An "isolated" or "purified" polypeptide or protein or biologically-active portion thereof is substantially free of cellular material or other contaminating proteins from the cell or tissue source from which the ΝOVX protein is derived, or substantially free from chemical precursors or other chemicals when chemically synthesized. The language "substantially free of cellular material" includes preparations of ΝOVX proteins in which the protein is separated from cellular components of the cells from which it is isolated or recombinantly-produced. In
one embodiment, the language "substantially free of cellular material" includes preparations of NOVX proteins having less than about 30% (by dry weight) of non-NOVX proteins (also referred to herein as a "contaminating protein"), more preferably less than about 20%> of non-NOVX proteins, still more preferably less than about 10%. of non-NOVX proteins, and most preferably less than about 5%. of non-NOVX proteins. When the NOVX protein or biologically-active portion thereof is recombinantly-produced, it is also preferably substantially free of culture medium, i.e., culture medium represents less than about 20%, more preferably less than about 10%, and most preferably less than about 5%> of the volume of the NOVX protein preparation.
The language "substantially free of chemical precursors or other chemicals" includes preparations of NOVX proteins in which the protein is separated from chemical precursors or other chemicals that are involved in the synthesis of the protein. In one embodiment, the language "substantially free of chemical precursors or other chemicals" includes preparations of NOVX proteins having less than about 30%. (by dry weight) of chemical precursors or non-NOVX chemicals, more preferably less than about 20%. chemical precursors or non-NOVX chemicals, still more preferably less than about 10%. chemical precursors or non-NOVX chemicals, and most preferably less than about 5% chemical precursors or non-NOVX chemicals.
Biologically-active portions of NOVX proteins include peptides comprising amino acid sequences sufficiently homologous to or derived from the amino acid sequences of the NOVX proteins (e.g., the amino acid sequence shown in SEQ ED NO: 2n, wherein n is an integer between 1 and 178) that include fewer amino acids than the full-length NOVX proteins, and exhibit at least one activity of an NOVX protein. Typically, biologically-active portions comprise a domain or motif with at least one activity of the NOVX protein. A biologically-active portion of an NOVX protein can be a polypeptide which is, for example, 10, 25, 50, 100 or more amino acid residues in length.
Moreover, other biologically-active portions, in which other regions of the protein are deleted, can be prepared by recombinant techniques and evaluated for one or more of the functional activities of a native NOVX protein.
In an embodiment, the NOVX protein has an amino acid sequence shown SEQ ED NO: 2n, wherein n is an integer between 1 and 178. In other embodiments, the NOVX protein is substantially homologous to SEQ ED NO: 2n, wherein n is an integer between 1 and 178, and retains the functional activity of the protein of SEQ TD NO: 2n, wherein n is an integer between 1 and 178, yet differs in amino acid sequence due to natural allelic variation or
mutagenesis, as described in detail, below. Accordingly, in another embodiment, the NOVX protein is a protein that comprises an amino acid sequence at least about 45 %> homologous to the amino acid sequence SEQ ID NO: 2n, wherein n is an integer between 1 and 178, and retains the functional activity of the NOVX proteins of SEQ TD NO: 2n, wherein n is an integer between 1 and 178.
Determining Homology Between Two or More Sequences
To determine the percent homology of two amino acid sequences or of two nucleic acids, the sequences are aligned for optimal comparison puφoses (e.g., gaps can be introduced in the sequence of a first amino acid or nucleic acid sequence for optimal alignment with a second amino or nucleic acid sequence). The amino acid residues or nucleotides at corresponding amino acid positions or nucleotide positions are then compared. When a position in the first sequence is occupied by the same amino acid residue or nucleotide as the corresponding position in the second sequence, then the molecules are homologous at that position (i.e., as used herein amino acid or nucleic acid "homology" is equivalent to amino acid or nucleic acid "identity").
The nucleic acid sequence homology may be determined as the degree of identity between two sequences. The homology may be determined using computer programs known in the art, such as GAP software provided in the GCG program package. See, Needleman and Wunsch, 1970. JMol Biol 48: 443-453. Using GCG GAP software with the following settings for nucleic acid sequence comparison: GAP creation penalty of 5.0 and GAP extension penalty of 0.3, the coding region of the analogous nucleic acid sequences referred to above exhibits a degree of identity preferably of at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99%, with the CDS (encoding) part of the DNA from the group consisting of SEQ TD NO: 2n-l, wherein n is an integer between 1 and 178.
The term "sequence identity" refers to the degree to which two polynucleotide or polypeptide sequences are identical on a residue-by-residue basis over a particular region of comparison. The term "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over that region of comparison, determining the number of positions at which the identical nucleic acid base (e.g., A, T, C, G, U, or I, in the case of nucleic acids) occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the region of comparison (i.e., the window size), and multiplying the result by 100 to yield the percentage of sequence identity. The term "substantial identity" as used herein denotes a characteristic of a
polynucleotide sequence, wherein the polynucleotide comprises a sequence that has at least 80 percent sequence identity, preferably at least 85 percent identity and often 90 to 95 percent sequence identity, more usually at least 99 percent sequence identity as compared to a reference sequence over a comparison region.
Chimeric and Fusion Proteins
The invention also provides NOVX chimeric or fusion proteins. As used herein, an NOVX "chimeric protein" or "fusion protein" comprises an NOVX polypeptide operatively- linked to a non-NOVX polypeptide. An "NOVX polypeptide" refers to a polypeptide having an amino acid sequence corresponding to an NOVX protein SEQ ED NO: 2n, wherein n is an integer between 1 and 178, whereas a "non-NOVX polypeptide" refers to a polypeptide having an amino acid sequence corresponding to a protein that is not substantially homologous to the NOVX protein, e.g., a protein that is different from the NOVX protein and that is derived from the same or a different organism. Within an NOVX fusion protein the NOVX polypeptide can correspond to all or a portion of an NOVX protein. In one embodiment, an NOVX fusion protein comprises at least one biologically-active portion of an NOVX protein. In another embodiment, an NOVX fusion protein comprises at least two biologically-active portions of an NOVX protein. In yet another embodiment, an NOVX fusion protein comprises at least three biologically-active portions of an NOVX protein. Within the fusion protein, the term "operatively-linked" is intended to indicate that the NOVX polypeptide and the non-NOVX polypeptide are fused in- frame with one another. The non-NOVX polypeptide can be fused to the N- terminus or C-terminus of the NOVX polypeptide.
In one embodiment, the fusion protein is a GST-NOVX fusion protein in which the NOVX sequences are fused to the C-terminus of the GST (glutathione S-transferase) sequences. Such fusion proteins can facilitate the purification of recombinant NOVX polypeptides.
In another embodiment, the fusion protein is an NOVX protein containing a heterologous signal sequence at its N-terminus. In certain host cells (e.g., mammalian host cells), expression and or secretion of NOVX can be increased through use of a heterologous signal sequence.
In yet another embodiment, the fusion protein is an NOVX-immunoglobulin fusion protein in which the NOVX sequences are fused to sequences derived from a member of the immunoglobulin protein family. The NOVX-immunoglobulin fusion proteins of the invention can be incoφorated into pharmaceutical compositions and administered to a subject to inhibit
an interaction between an NOVX ligand and an NOVX protein on the surface of a cell, to thereby suppress NOVX-mediated signal transduction in vivo. The NOVX-immunoglobulin fusion proteins can be used to affect the bioavailability of an NOVX cognate ligand. Inhibition of the NOVX ligand/NOVX interaction may be useful therapeutically for both the treatment of proliferative and differentiative disorders, as well as modulating (e.g. promoting or inhibiting) cell survival. Moreover, the NOVX-immunoglobulin fusion proteins of the invention can be used as immunogens to produce anti-NOVX antibodies in a subject, to purify NOVX ligands, and in screening assays to identify molecules that inhibit the interaction of NOVX with an NOVX ligand.
An NOVX chimeric or fusion protein of the invention can be produced by standard recombinant DNA techniques. For example, DNA fragments coding for the different polypeptide sequences are ligated together in-frame in accordance with conventional techniques, e.g., by employing blunt-ended or stagger-ended termini for ligation, restriction enzyme digestion to provide for appropriate termini, filling-in of cohesive ends as appropriate, alkaline phosphatase treatment to avoid undesirable joining, and enzymatic ligation. In another embodiment, the fusion gene can be synthesized by conventional techniques including automated DNA synthesizers. Alternatively, PCR amplification of gene fragments can be carried out using anchor primers that give rise to complementary overhangs between two consecutive gene fragments that can subsequently be annealed and reamplified to generate a chimeric gene sequence (see, e.g., Ausubel, et al. (eds.) CURRENT PROTOCOLS IN MOLECULAR BIOLOGY, John Wiley & Sons, 1992). Moreover, many expression vectors are commercially available that already encode a fusion moiety (e.g., a GST polypeptide). An NOVX-encoding nucleic acid can be cloned into such an expression vector such that the fusion moiety is linked in- frame to the NOVX protein.
NOVX Agonists and Antagonists
The invention also pertains to variants of the NOVX proteins that function as either
NOVX agonists (i.e., mimetics) or as NOVX antagonists. Variants of the NOVX protein can be generated by mutagenesis (e.g., discrete point mutation or truncation of the NOVX protein).
An agonist of the NOVX protein can retain substantially the same, or a subset of, the biological activities of the naturally occurring form of the NOVX protein. An antagonist of the NOVX protein can inhibit one or more of the activities of the naturally occurring form of the NOVX protein by, for example, competitively binding to a downstream or upstream member of a cellular signaling cascade which includes the NOVX protein. Thus, specific
biological effects can be elicited by treatment with a variant of limited function. In one embodiment, treatment of a subject with a variant having a subset of the biological activities of the naturally occurring form of the protein has fewer side effects in a subject relative to treatment with the naturally occurring form of the NOVX proteins.
Variants of the NOVX proteins that function as either NOVX agonists (i.e., mimetics) or as NOVX antagonists can be identified by screening combinatorial libraries of mutants (e.g., truncation mutants) of the NOVX proteins for NOVX protein agonist or antagonist activity. In one embodiment, a variegated library of NOVX variants is generated by combinatorial mutagenesis at the nucleic acid level and is encoded by a variegated gene library. A variegated library of NOVX variants can be produced by, for example, enzymatically ligating a mixture of synthetic oligonucleotides into gene sequences such that a degenerate set of potential NOVX sequences is expressible as individual polypeptides, or alternatively, as a set of larger fusion proteins (e.g., for phage display) containing the set of NOVX sequences therein. There are a variety of methods which can be used to produce libraries of potential NOVX variants from a degenerate oligonucleotide sequence. Chemical synthesis of a degenerate gene sequence can be performed in an automatic DNA synthesizer, and the synthetic gene then ligated into an appropriate expression vector. Use of a degenerate set of genes allows for the provision, in one mixture, of all of the sequences encoding the desired set of potential NOVX sequences. Methods for synthesizing degenerate oligonucleotides are well-known within the art. See, e.g., Narang, 1983. Tetrahedron 39: 3; Itakura, et al, 1984. Annu. Rev. Biochem. 53: 323; Itakura, et al, 1984. Science 198: 1056; Eke, et al, 1983. Nucl. Acids Res. 11: 477.
Polypeptide Libraries
In addition, libraries of fragments of the NOVX protein coding sequences can be used to generate a variegated population of NOVX fragments for screening and subsequent selection of variants of an NOVX protein. In one embodiment, a library of coding sequence fragments can be generated by treating a double stranded PCR fragment of an NOVX coding sequence with a nuclease under conditions wherein nicking occurs only about once per molecule, denaturing the double stranded DNA, renaturing the DNA to form double-stranded DNA that can include sense/antisense pairs from different nicked products, removing single stranded portions from reformed duplexes by treatment with Si nuclease, and ligating the resulting fragment library into an expression vector. By this method, expression libraries can
be derived which encodes N-terminal and internal fragments of various sizes of the NOVX proteins.
Various techniques are known in the art for screening gene products of combinatorial libraries made by point mutations or truncation, and for screening cDNA libraries for gene products having a selected property. Such techniques are adaptable for rapid screening of the gene libraries generated by the combinatorial mutagenesis of NOVX proteins. The most widely used techniques, which are amenable to high throughput analysis, for screening large gene libraries typically include cloning the gene library into replicable expression vectors, transforming appropriate cells with the resulting library of vectors, and expressing the combinatorial genes under conditions in which detection of a desired activity facilitates isolation of the vector encoding the gene whose product was detected. Recursive ensemble mutagenesis (REM), a new technique that enhances the frequency of functional mutants in the libraries, can be used in combination with the screening assays to identify NOVX variants. See, e.g., Arkin and Yourvan, 1992. Proc. Natl. Acad. Sci. USA 89: 7811-7815; Delgrave, et al, 1993. Protein Engineering 6:327-331.
NOVX Antibodies
The term "antibody" as used herein refers to immunoglobulin molecules and immunologically active portions of immunoglobulin (Ig) molecules, i.e., molecules that contain an antigen binding site that specifically binds (immunoreacts with) an antigen. Such antibodies include, but are not limited to, polyclonal, monoclonal, chimeric, single chain, Fab, Fab' and F(ab')2 fragments, and an Fa expression library. In general, antibody molecules obtained from humans relates to any of the classes IgG, IgM, IgA, IgE and IgD, which differ from one another by the nature of the heavy chain present in the molecule. Certain classes have subclasses as well, such as IgGi, IgG , and others. Furthermore, in humans, the light chain may be a kappa chain or a lambda chain. Reference herein to antibodies includes a reference to all such classes, subclasses and types of human antibody species.
An isolated protein of the invention intended to serve as an antigen, or a portion or fragment thereof, can be used as an immunogen to generate antibodies that immunospecifically bind the antigen, using standard techniques for polyclonal and monoclonal antibody preparation. The full-length protein can be used or, alternatively, the invention provides antigenic peptide fragments of the antigen for use as immunogens. An antigenic peptide fragment comprises at least 6 amino acid residues of the amino acid sequence of the full length protein, such as an amino acid sequence shown in SEQ ED NO: 2n, wherein n is an
integer between 1 and 178, and encompasses an epitope thereof such that an antibody raised against the peptide forms a specific immune complex with the full length protein or with any fragment that contains the epitope. Preferably, the antigenic peptide comprises at least 10 amino acid residues, or at least 15 amino acid residues, or at least 20 amino acid residues, or at least 30 amino acid residues. Preferred epitopes encompassed by the antigenic peptide are regions of the protein that are located on its surface; commonly these are hydrophilic regions.
In certain embodiments of the invention, at least one epitope encompassed by the antigenic peptide is a region of NOVX that is located on the surface of the protein, e.g., a hydrophilic region. A hydrophobicity analysis of the human NOVX protein sequence will indicate which regions of a NOVX polypeptide are particularly hydrophilic and, therefore, are likely to encode surface residues useful for targeting antibody production. As a means for targeting antibody production, hydropathy plots showing regions of hydrophilicity and hydrophobicity may be generated by any method well known in the art, including, for example, the Kyte Doolittle or the Hopp Woods methods, either with or without Fourier transformation. See, e.g., Hopp and Woods, 1981, Proc. Nat. Acad. Sci. USA 78: 3824-3828; Kyte and Doolittle 1982, J. Mol. Biol. 157: 105-142, each incoφorated herein by reference in their entirety. Antibodies that are specific for one or more domains within an antigenic protein, or derivatives, fragments, analogs or homologs thereof, are also provided herein.
A protein of the invention, or a derivative, fragment, analog, homolog or ortholog thereof, may be utilized as an immunogen in the generation of antibodies that immunospecifically bind these protein components.
Various procedures known within the art may be used for the production of polyclonal or monoclonal antibodies directed against a protein of the invention, or against derivatives, fragments, analogs homologs or orthologs thereof (see, for example, Antibodies: A Laboratory Manual, Harlow E, and Lane D, 1988, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, incoφorated herein by reference). Some of these antibodies are discussed below.
Polyclonal Antibodies
For the production of polyclonal antibodies, various suitable host animals (e.g., rabbit, goat, mouse or other mammal) may be immunized by one or more injections with the native protein, a synthetic variant thereof, or a derivative of the foregoing. An appropriate immunogenic preparation can contain, for example, the naturally occurring immunogenic protein, a chemically synthesized polypeptide representing the immunogenic protein, or a
recombinantly expressed immunogenic protein. Furthermore, the protein may be conjugated to a second protein known to be immunogenic in the mammal being immunized. Examples of such immunogenic proteins include but are not limited to keyhole limpet hemocyanin, serum albumin, bovine thyroglobulin, and soybean trypsin inhibitor. The preparation can further include an adjuvant. Various adjuvants used to increase the immunological response include, but are not limited to, Freund's (complete and incomplete), mineral gels (e.g., aluminum hydroxide), surface active substances (e.g., lysolecithin, pluronic polyols, polyanions, peptides, oil emulsions, dinitrophenol, etc.), adjuvants usable in humans such as Bacille Calmette-Guerin and Corynebacterium parvum, or similar immunostimulatory agents. Additional examples of adjuvants which can be employed include MPL-TDM adjuvant (monophosphoryl Lipid A, synthetic trehalose dicorynomycolate).
The polyclonal antibody molecules directed against the immunogenic protein can be isolated from the mammal (e.g., from the blood) and further purified by well known techniques, such as affimty chromatography using protein A or protein G, which provide primarily the IgG fraction of immune serum. Subsequently, or alternatively, the specific antigen which is the target of the immunoglobulin sought, or an epitope thereof, may be immobilized on a column to purify the immune specific antibody by immunoaffinity chromatography. Purification of immunoglobulins is discussed, for example, by D. Wilkinson (The Scientist, published by The Scientist, Inc., Philadelphia PA, Vol. 14, No. 8 (April 17, 2000), pp. 25-28).
Monoclonal Antibodies
The term "monoclonal antibody" (MAb) or "monoclonal antibody composition", as used herein, refers to a population of antibody molecules that contain only one molecular species of antibody molecule consisting of a unique light chain gene product and a unique heavy chain gene product. In particular, the complementarity determining regions (CDRs) of the monoclonal antibody are identical in all the molecules of the population. MAbs thus contain an antigen binding site capable of immunoreacting with a particular epitope of the antigen characterized by a unique binding affinity for it.
Monoclonal antibodies can be prepared using hybridoma methods, such as those described by Kohler and Milstein, Nature. 256:495 (1975). In a hybridoma method, a mouse, hamster, or other appropriate host animal, is typically immunized with an immunizing agent to
elicit lymphocytes that produce or are capable of producing antibodies that will specifically bind to the immunizing agent. Alternatively, the lymphocytes can be immunized in vitro.
The immunizing agent will typically include the protein antigen, a fragment thereof or a fusion protein thereof. Generally, either peripheral blood lymphocytes are used if cells of human origin are desired, or spleen cells or lymph node cells are used if non-human mammalian sources are desired. The lymphocytes are then fused with an immortalized cell line using a suitable fusing agent, such as polyethylene giycol, to form a hybridoma cell [Goding, Monoclonal Antibodies: Principles and Practice, Academic Press, (1986) pp. 59- 103]. Immortalized cell lines are usually transformed mammalian cells, particularly myeloma cells of rodent, bovine and human origin. Usually, rat or mouse myeloma cell lines are employed. The hybridoma cells can be cultured in a suitable culture medium that preferably contains one or more substances that inhibit the growth or survival of the unfused, immortalized cells. For example, if the parental cells lack the enzyme hypoxanthine guanine phosphoribosyl transferase (HGPRT or HPRT), the culture medium for the hybridomas typically will include hypoxanthine, aminopterin, and thymidine ("HAT medium"), which substances prevent the growth of HGPRT-deficient cells.
Preferred immortalized cell lines are those that fuse efficiently, support stable high level expression of antibody by the selected antibody-producing cells, and are sensitive to a medium such as HAT medium. More prefened immortalized cell lines are murine myeloma lines, which can be obtained, for instance, from the Salk Institute Cell Distribution Center, San Diego, California and the American Type Culture Collection, Manassas, Virginia. Human myeloma and mouse-human heteromyeloma cell lines also have been described for the production of human monoclonal antibodies [Kozbor, J. Immunol., 133:3001 (1984); Brodeur et al., Monoclonal Antibody Production Techniques and Applications, Marcel Dekker, Inc., New York, (1987) pp. 51-63].
The culture medium in which the hybridoma cells are cultured can then be assayed for the presence of monoclonal antibodies directed against the antigen. Preferably, the binding specificity of monoclonal antibodies produced by the hybridoma cells is determined by immunoprecipitation or by an in vitro binding assay, such as radioimmunoassay (RIA) or enzyme-linked immunoabsorbent assay (ELISA). Such techniques and assays are known in the art. The binding affinity of the monoclonal antibody can, for example, be determined by the Scatchard analysis of Munson and Pollard, Anal. Biochem., 107:220 (1980). It is an objective, especially important in therapeutic applications of monoclonal antibodies, to
identify antibodies having a high degree of specificity and a high binding affinity for the target antigen.
After the desired hybridoma cells are identified, the clones can be subcloned by limiting dilution procedures and grown by standard methods (Goding,1986). Suitable culture media for this pvupose include, for example, Dulbecco's Modified Eagle's Medium and RPMI- 1640 medium. Alternatively, the hybridoma cells can be grown in vivo as ascites in a mammal.
The monoclonal antibodies secreted by the subclones can be isolated or purified from the culture medium or ascites fluid by conventional immunoglobulin purification procedures such as, for example, protein A-Sepharose, hydroxylapatite chromatography, gel electrophoresis, dialysis, or affinity chromatography.
The monoclonal antibodies can also be made by recombinant DNA methods, such as those described in U.S. Patent No. 4,816,567. DNA encoding the monoclonal antibodies of the invention can be readily isolated and sequenced using conventional procedures (e.g., by using oligonucleotide probes that are capable of binding specifically to genes encoding the heavy and light chains of murine antibodies). The hybridoma cells of the invention serve as a preferred source of such DNA. Once isolated, the DNA can be placed into expression vectors, which are then transfected into host cells such as simian COS cells, Chinese hamster ovary (CHO) cells, or myeloma cells that do not otherwise produce immunoglobulin protein, to obtain the synthesis of monoclonal antibodies in the recombinant host cells. The DNA also can be modified, for example, by substituting the coding sequence for human heavy and light chain constant domains in place of the homologous murine sequences (U.S. Patent No. 4,816,567; Morrison, Nature 368, 812-13 (1994)) or by covalently joining to the immunoglobulin coding sequence all or part of the coding sequence for a non-immunoglobulin polypeptide. Such a non-immunoglobulin polypeptide can be substituted for the constant domains of an antibody of the invention, or can be substituted for the variable domains of one antigen-combining site of an antibody of the invention to create a chimeric bivalent antibody.
Humanized Antibodies
The antibodies directed against the protein antigens of the invention can further comprise humanized antibodies or human antibodies. These antibodies are suitable for administration to humans without engendering an immune response by the human against the administered immunoglobulin. Humanized forms of antibodies are chimeric immunoglobulins, immunoglobulin chains or fragments thereof (such as Fv, Fab, Fab', F(ab')2 or other antigen-
binding subsequences of antibodies) that are principally comprised of the sequence of a human immunoglobulin, and contain minimal sequence derived from a non-human immunoglobulin. Humanization can be performed following the method of Winter and co-workers (Jones et al., Nature, 321:522-525 (1986); Riechmann et al., Nature. 332:323-327 (1988); Verhoeyen et al., Science. 239:1534-1536 (1988)), by substituting rodent CDRs or CDR sequences for the conesponding sequences of a human antibody. (See also U.S. Patent No. 5,225,539.) In some instances, Fv framework residues of the human immunoglobulin are replaced by conesponding non-human residues. Humanized antibodies can also comprise residues which are found neither in the recipient antibody nor in the imported CDR or framework sequences. In general, the humanized antibody will comprise substantially all of at least one, and typically two, variable domains, in which all or substantially all of the CDR regions conespond to those of a non-human immunoglobulin and all or substantially all of the framework regions are those of a human immunoglobulin consensus sequence. The humanized antibody optimally also will comprise at least a portion of an immunoglobulin constant region (Fc), typically that of a human immunoglobulin (Jones et al., 1986; Riechmann et al., 1988; and Presta, Cun. Op. Struct. Biol.. 2:593-596 (1992)).
Human Antibodies
Fully human antibodies essentially relate to antibody molecules in which the entire sequence of both the light chain and the heavy chain, including the CDRs, arise from human genes. Such antibodies are termed "human antibodies", or "fully human antibodies" herein. Human monoclonal antibodies can be prepared by the trioma technique; the human B-cell hybridoma technique (see Kozbor, et al., 1983 Immunol Today 4: 72) and the EBV hybridoma technique to produce human monoclonal antibodies (see Cole, et al., 1985 In: MONOCLONAL ANTIBODIES AND CANCER THERAPY, Alan R. Liss, Inc., pp. 77-96). Human monoclonal antibodies may be utilized in the practice of the present invention and may be produced by using human hybridomas (see Cote, et al., 1983. Proc Natl Acad Sci USA 80: 2026-2030) or by transforming human B-cells with Epstein Ban Virus in vitro (see Cole, et al., 1985 In: MONOCLONAL ANTIBODIES AND CANCER THERAPY, Alan R. Liss, Inc., pp. 77-96).
In addition, human antibodies can also be produced using additional techniques, including phage display libraries (Hoogenboom and Winter, J. Mol. Biol, 227:381 (1991);
Marks et al., J. Mol. Biol., 222:581 (1991)). Similarly, human antibodies can be made by introducing human immunoglobulin loci into transgenic animals, e.g., mice in which the endogenous immunoglobulin genes have been partially or completely inactivated. Upon
challenge, human antibody production is observed, which closely resembles that seen in humans in all respects, including gene reanangement, assembly, and antibody repertoire. This approach is described, for example, in U.S. Patent Nos. 5,545,807; 5,545,806; 5,569,825; 5,625,126; 5,633,425; 5,661,016, and in Marks et al. (Bio/Technology 10. 779-783 (1992)); Lonberg et al. (Nature 368 856-859 (1994)); Morrison ( Nature 368, 812-13 (1994)); Fishwild et al,( Nature Biotechnology 14, 845-51 (1996)); Neuberger (Nature Biotechnology 14, 826 (1996)); and Lonberg and Huszar (Intern. Rev. Immunol. 13 65-93 (1995)).
Human antibodies may additionally be produced using transgenic nonhuman animals which are modified so as to produce fully human antibodies rather than the animal's endogenous antibodies in response to challenge by an antigen. (See PCT publication WO94/02602). The endogenous genes encoding the heavy and light immunoglobulin chains in the nonhuman host have been incapacitated, and active loci encoding human heavy and light chain immunoglobulins are inserted into the host's genome. The human genes are incoφorated, for example, using yeast artificial chromosomes containing the requisite human DNA segments. An animal which provides all the desired modifications is then obtained as progeny by crossbreeding intermediate transgenic animals containing fewer than the full complement of the modifications. The prefeπed embodiment of such a nonhuman animal is a mouse, and is termed the Xenomouse™ as disclosed in PCT publications WO 96/33735 and WO 96/34096. This animal produces B cells which secrete fully human immunoglobulins. The antibodies can be obtained directly from the animal after immunization with an immunogen of interest, as, for example, a preparation of a polyclonal antibody, or alternatively from immortalized B cells derived from the animal, such as hybridomas producing monoclonal antibodies. Additionally, the genes encoding the immunoglobulins with human variable regions can be recovered and expressed to obtain the antibodies directly, or can be further modified to obtain analogs of antibodies such as, for example, single chain Fv molecules.
An example of a method of producing a nonhuman host, exemplified as a mouse, lacking expression of an endogenous immunoglobulin heavy chain is disclosed in U.S. Patent No. 5,939,598. It can be obtained by a method including deleting the J segment genes from at least one endogenous heavy chain locus in an embryonic stem cell to prevent reanangement of the locus and to prevent formation of a transcript of a reananged immunoglobulin heavy chain locus, the deletion being effected by a targeting vector containing a gene encoding a selectable marker; and producing from the embryonic stem cell a transgenic mouse whose somatic and germ cells contain the gene encoding the selectable marker.
A method for producing an antibody of interest, such as a human antibody, is disclosed in U.S. Patent No. 5,916,771. It includes introducing an expression vector that contains a nucleotide sequence encoding a heavy chain into one mammalian host cell in culture, introducing an expression vector containing a nucleotide sequence encoding a light chain into another mammalian host cell, and fusing the two cells to form a hybrid cell. The hybrid cell expresses an antibody containing the heavy chain and the light chain.
In a further improvement on this procedure, a method for identifying a clinically relevant epitope on an immunogen, and a conelative method for selecting an antibody that binds immunospecifically to the relevant epitope with high affimty, are disclosed in PCT publication WO 99/53049.
Fa Fragments and Single Chain Antibodies
According to the invention, techniques can be adapted for the production of single-chain antibodies specific to an antigenic protein of the invention (see e.g., U.S. Patent No. 4,946,778). In addition, methods can be adapted for the construction of Fab expression libraries (see e.g., Huse, et al., 1989 Science 246: 1275-1281) to allow rapid and effective identification of monoclonal Fab fragments with the desired specificity for a protein or derivatives, fragments, analogs or homologs thereof. Antibody fragments that contain the idiotypes to a protein antigen may be produced by techniques known in the art including, but not limited to: (i) an F(ab')2 fragment produced by pepsin digestion of an antibody molecule; (ii) an Fa fragment generated by reducing the disulfide bridges of an F(ab')2 fragment; (iii) an Fab fragment generated by the treatment of the antibody molecule with papain and a reducing agent and (iv) Fv fragments.
Bispecific Antibodies
Bispecific antibodies are monoclonal, preferably human or humanized, antibodies that have binding specificities for at least two different antigens. In the present case, one of the binding specificities is for an antigenic protein of the invention. The second binding target is any other antigen, and advantageously is a cell-surface protein or receptor or receptor subunit. Methods for making bispecific antibodies are known in the art. Traditionally, the recombinant production of bispecific antibodies is based on the co-expression of two immunoglobulin heavy-chain/light-chain pairs, where the two heavy chains have different specificities (Milstein and Cuello, Nature, 305:537-539 (1983)). Because of the random assortment of
immunoglobulin heavy and light chains, these hybridomas (quadromas) produce a potential mixture often different antibody molecules, of which only one has the correct bispecific structure. The purification of the conect molecule is usually accomplished by affinity chromatography steps. Similar procedures are disclosed in WO 93/08829, published 13 May 1993, and in Traunecker et al., EMBO J., 10:3655-3659 (1991).
Antibody variable domains with the desired binding specificities (antibody-antigen combining sites) can be fused to immunoglobulin constant domain sequences. The fusion preferably is with an immunoglobulin heavy-chain constant domain, comprising at least part of the hinge, CH2, and CH3 regions. It is prefened to have the first heavy-chain constant region (CHI) containing the site necessary for light-chain binding present in at least one of the fusions. DNAs encoding the immunoglobulin heavy-chain fusions and, if desired, the immunoglobulin light chain, are inserted into separate expression vectors, and are cotransfected into a suitable host organism. For further details of generating bispecific antibodies see, for example, Suresh et al., Methods in Enzymology, 121:210 (1986).
According to another approach described in WO 96/27011, the interface between a pair of antibody molecules can be engineered to maximize the percentage of heterodimers which are recovered from recombinant cell culture. The prefened interface comprises at least a part of the CH3 region of an antibody constant domain. In this method, one or more small amino acid side chains from the interface of the first antibody molecule are replaced with larger side chains (e.g. tyrosine or tryptophan). Compensatory "cavities" of identical or similar size to the large side chain(s) are created on the interface of the second antibody molecule by replacing large amino acid side chains with smaller ones (e.g. alanine or threonine). This provides a mechanism for increasing the yield of the heterodimer over other unwanted end-products such as homodimers.
Bispecific antibodies can be prepared as full length antibodies or antibody fragments (e.g. F(ab')2 bispecific antibodies). Techniques for generating bispecific antibodies from antibody fragments have been described in the literature. For example, bispecific antibodies can be prepared using chemical linkage. Brennan et al., Science 229:81 (1985) describe a procedure wherein intact antibodies are proteolytically cleaved to generate F(ab')2 fragments. These fragments are reduced in the presence of the dithiol complexing agent sodium arsenite to stabilize vicinal dithiols and prevent intermolecular disulfide formation. The Fab' fragments generated are then converted to thionitrobenzoate (TNB) derivatives. One of the Fab'-TNB derivatives is then reconverted to the Fab'-thiol by reduction with mercaptoefhylamine and is mixed with an equimolar amount of the other Fab'-TNB derivative to form the bispecific
antibody. The bispecific antibodies produced can be used as agents for the selective immobilization of enzymes.
Additionally, Fab' fragments can be directly recovered from E. coli and chemically coupled to form bispecific antibodies. Shalaby et al., J. Exp. Med. 175:217-225 (1992) describe the production of a fully humanized bispecific antibody F(ab')2 molecule. Each Fab' fragment was separately secreted from E. coli and subjected to directed chemical coupling in vitro to form the bispecific antibody. The bispecific antibody thus formed was able to bind to cells overexpressing the ErbB2 receptor and normal human T cells, as well as trigger the lytic activity of human cytotoxic lymphocytes against human breast tumor targets.
Various techniques for making and isolating bispecific antibody fragments directly from recombinant cell culture have also been described. For example, bispecific antibodies have been produced using leucine zippers. Kostelny et al., J. Immunol. 148(5): 1547-1553 (1992). The leucine zipper peptides from the Fos and Jun proteins were linked to the Fab' portions of two different antibodies by gene fusion. The antibody homodimers were reduced at the hinge region to form monomers and then re-oxidized to form the antibody heterodimers. This method can also be utilized for the production of antibody homodimers. The "diabody" technology described by Hollinger et al., Proc. Natl. Acad. Sci. USA 90:6444-6448 (1993) has provided an alternative mechanism for making bispecific antibody fragments. The fragments comprise a heavy-chain variable domain (VH) connected to a light-chain variable domain (VL) by a linker which is too short to allow pairing between the two domains on the same chain. Accordingly, the VH and V domains of one fragment are forced to pair with the complementary VL and VH domains of another fragment, thereby forming two antigen-binding sites. Another strategy for making bispecific antibody fragments by the use of single-chain Fv (sFv) dimers has also been reported. See, Gruber et al., J. Immunol. 152:5368 (1994). Antibodies with more than two valencies are contemplated. For example, trispecific antibodies can be prepared. Tutt et al., J. Immunol. 147:60 (1991).
Exemplary bispecific antibodies can bind to two different epitopes, at least one of which originates in the protein antigen of the invention. Alternatively, an anti-antigenic arm of an immunoglobulin molecule can be combined with an arm which binds to a triggering molecule on a leukocyte such as a T-cell receptor molecule (e.g. CD2, CD3, CD28, or B7), or Fc receptors for IgG (FcγR), such as FcγRI (CD64), FcγRII (CD32) and FcγRIII (CD 16) so as to focus cellular defense mechanisms to the cell expressing the particular antigen. Bispecific antibodies can also be used to direct cytotoxic agents to cells which express a particular antigen. These antibodies possess an antigen-binding arm and an arm which binds a cytotoxic
agent or a radionuclide chelator, such as EOTUBE, DPTA, DOT A, or TETA. Another bispecific antibody of interest binds the protein antigen described herein and further binds tissue factor (TF).
Heteroconjugate Antibodies
Heteroconjugate antibodies are also within the scope of the present invention. Heteroconjugate antibodies are composed of two covalently joined antibodies. Such antibodies have, for example, been proposed to target immune system cells to unwanted cells (U.S. Patent No. 4,676,980), and for treatment of HIV infection (WO 91/00360; WO 92/200373; EP 03089). It is contemplated that the antibodies can be prepared in vitro using known methods in synthetic protein chemistry, including those involving crosslinking agents. For example, immunotoxins can be constructed using a disulfide exchange reaction or by forming a thioether bond. Examples of suitable reagents for this puφose include iminothiolate and mefhyl-4-mercaptobutyrimidate and those disclosed, for example, in U.S. Patent No. 4,676,980.
Effector Function Engineering
It can be desirable to modify the antibody of the invention with respect to effector function, so as to enhance, e.g., the effectiveness of the antibody in treating cancer. For example, cysteine residue(s) can be introduced into the Fc region, thereby allowing interchain disulfide bond formation in this region. The homodimeric antibody thus generated can have improved internalization capability and/or increased complement-mediated cell killing and antibody-dependent cellular cytotoxicity (ADCC). See Caron et al., J. Exp Med., 176: 1191- 1195 (1992) and Shopes, J. Immunol., 148: 2918-2922 (1992). Homodimeric antibodies with enhanced anti-tumor activity can also be prepared using heterobifunctional cross-linkers as described in Wolff et al. Cancer Research, 53: 2560-2565 (1993). Alternatively, an antibody can be engineered that has dual Fc regions and can thereby have enhanced complement lysis and ADCC capabilities. See Stevenson et al., Anti-Cancer Drug Design. 3: 219-230 (1989).
Immunoconjugates
The invention also pertains to immunoconjugates comprising an antibody conjugated to a cytotoxic agent such as a chemotherapeutic agent, toxin (e.g., an enzymatically active
toxin of bacterial, fungal, plant, or animal origin, or fragments thereof), or a radioactive isotope (i.e., a radioconjugate).
Chemotherapeutic agents useful in the generation of such immunoconjugates have been described above. Enzymatically active toxins and fragments thereof that can be used include diphtheria A chain, nonbinding active fragments of diphtheria toxin, exotoxin A chain (from Pseudomonas aeruginosa), ricin A chain, abrin A chain, modeccin A chain, alpha-sarcin, Aleurites fordii proteins, dianthin proteins, Phytolaca americana proteins (PAPI, PAPII, and PAP-S), momordica charantia inhibitor, curcin, crotin, sapaonaria officinalis inhibitor, gelonin, mitogellin, restrictocin, phenomycin, enomycin, and the tricothecenes. A variety of radionuclides are available for the production of radioconjugated antibodies. Examples include 2,2Bi, 131I, 131In, 90Y, and 186Re.
Conjugates of the antibody and cytotoxic agent are made using a variety of bifunctional protein-coupling agents such as N-succinimidyl-3-(2-pyridyldithiol) propionate (SPDP), iminothiolane (IT), bifunctional derivatives of imidoesters (such as dimethyl adipimidate HCL), active esters (such as disuccinimidyl suberate), aldehydes (such as glutareldehyde), bis-azido compounds (such as bis (p-azidobenzoyl) hexanediamine), bis- diazonium derivatives (such as bis-(p-diazoniumbenzoyl)-ethylenediamine), diisocyanates (such as tolyene 2,6-diisocyanate), and bis-active fluorine compounds (such as 1,5-difluoro- 2,4-dinitrobenzene). For example, a ricin immunotoxin can be prepared as described in Vitetta et al., Science. 238: 1098 (1987). Carbon- 14-labeled l-isothiocyanatobenzyl-3- methyldiethylene triaminepentaacetic acid (MX-DTPA) is an exemplary chelating agent for conjugation of radionucleotide to the antibody. See WO94/11026.
In another embodiment, the antibody can be conjugated to a "receptor" (such streptavidin) for utilization in tumor pretargeting wherein the antibody-receptor conjugate is administered to the patient, followed by removal of unbound conjugate from the circulation using a clearing agent and then administration of a "ligand" (e.g., avidin) that is in turn conjugated to a cytotoxic agent.
Immunoliposomes
The antibodies disclosed herein can also be formulated as immunoliposomes. Liposomes containing the antibody are prepared by methods known in the art, such as described in Epstein et al., Proc. Natl. Acad. Sci. USA. 82: 3688 (1985); Hwang et al., Proc. Natl Acad. Sci. USA. 77: 4030 (1980); and U.S. Pat. Nos. 4,485,045 and 4,544,545. Liposomes with enhanced circulation time are disclosed in U.S. Patent No. 5,013,556.
Particularly useful liposomes can be generated by the reverse-phase evaporation method with a lipid composition comprising phosphatidylcholine, cholesterol, and PEG- derivatized phosphatidylethanolamine (PEG-PE). Liposomes are extruded through filters of defined pore size to yield liposomes with the desired diameter. Fab' fragments of the antibody of the present invention can be conjugated to the liposomes as described in Martin et al .,_J. Biol. Chem., 257: 286-288 (1982) via a disulfide-interchange reaction. A chemotherapeutic agent (such as Doxorubicin) is optionally contained within the liposome. See Gabizon et al., J. National Cancer Inst.. 81(19): 1484 (1989).
Diagnostic Applications of Antibodies Directed Against the Proteins of the Invention
Antibodies directed against a protein of the invention may be used in methods known within the art relating to the localization and/or quantitation of the protein (e.g., for use in measuring levels of the protein within appropriate physiological samples, for use in diagnostic methods, for use in imaging the protein, and the like). In a given embodiment, antibodies against the proteins, or derivatives, fragments, analogs or homologs thereof, that contain the antigen binding domain, are utilized as pharmacologically-active compounds (see below).
An antibody specific for a protein of the invention can be used to isolate the protein by standard techniques, such as immunoaffinity chromatography or immunoprecipitation. Such an antibody can facilitate the purification of the natural protein antigen from cells and of recombinantly produced antigen expressed in host cells. Moreover, such an antibody can be used to detect the antigenic protein (e.g., in a cellular lysate or cell supernatant) in order to evaluate the abundance and pattem of expression of the antigenic protein. Antibodies directed against the protein can be used diagnostically to monitor protein levels in tissue as part of a clinical testing procedure, e.g., to, for example, determine the efficacy of a given treatment regimen. Detection can be facilitated by coupling (i.e., physically linking) the antibody to a detectable substance. Examples of detectable substances include various enzymes, prosthetic groups, fluorescent materials, luminescent materials, bioluminescent materials, and radioactive materials. Examples of suitable enzymes include horseradish peroxidase, alkaline phosphatase, β-galactosidase, or acetylcholinesterase; examples of suitable prosthetic group complexes include streptavidin/biotin and avidin/biotin; examples of suitable fluorescent materials include umbelliferone, fluorescein, fluorescein isothiocyanate, rhodamine, dichlorotriazinylamine fluorescein, dansyl chloride or phycoerythrin; an example of a luminescent material includes luminol; examples of bioluminescent materials include
luciferase, luciferin, and aequorin, and examples of suitable radioactive material include 125I
Antibody Therapeutics
Antibodies of the invention, including polyclonal, monoclonal, humanized and fully human antibodies, may used as therapeutic agents. Such agents will generally be employed to treat or prevent a disease or pathology in a subject. An antibody preparation, preferably one having high specificity and high affinity for its target antigen, is administered to the subject and will generally have an effect due to its binding with the target. Such an effect may be one of two kinds, depending on the specific nature of the interaction between the given antibody molecule and the target antigen in question. In the first instance, administration of the antibody may abrogate or inhibit the binding of the target with an endogenous ligand to which it naturally binds. In this case, the antibody binds to the target and masks a binding site of the naturally occurring ligand, wherein the ligand serves as an effector molecule. Thus the receptor mediates a signal transduction pathway for which ligand is responsible.
Alternatively, the effect may be one in which the antibody elicits a physiological result by virtue of binding to an effector binding site on the target molecule. In this case the target, a receptor having an endogenous ligand which may be absent or defective in the disease or pathology, binds the antibody as a surrogate effector ligand, initiating a receptor-based signal transduction event by the receptor.
A therapeutically effective amount of an antibody of the invention relates generally to the amount needed to achieve a therapeutic objective. As noted above, this may be a binding interaction between the antibody and its target antigen that, in certain cases, interferes with the functioning of the target, and in other cases, promotes a physiological response. The amount required to be administered will furthermore depend on the binding affinity of the antibody for its specific antigen, and will also depend on the rate at which an administered antibody is depleted from the free volume other subject to which it is administered. Common ranges for therapeutically effective dosing of an antibody or antibody fragment of the invention may be, by way of nonlimiting example, from about 0.1 mg/kg body weight to about 50 mg/kg body weight. Common dosing frequencies may range, for example, from twice daily to once a week.
Pharmaceutical Compositions of Antibodies
Antibodies specifically binding a protein of the invention, as well as other molecules identified by the screening assays disclosed herein, can be administered for the treatment of various disorders in the form of pharmaceutical compositions. Principles and considerations involved in preparing such compositions, as well as guidance in the choice of components are provided, for example, in Remington : The Science And Practice Of Pharmacy 19th ed. (Alfonso R. Gennaro, et al., editors) Mack Pub. Co., Easton, Pa. : 1995; Drug Absoφtion Enhancement : Concepts, Possibilities, Limitations, And Trends, Harwood Academic Publishers, Langhorne, Pa., 1994; and Peptide And Protein Drug Delivery (Advances In Parenteral Sciences, Vol. 4), 1991, M. Dekker, New York.
If the antigenic protein is intracellular and whole antibodies are used as inhibitors, internalizing antibodies are prefened. However, liposomes can also be used to deliver the antibody, or an antibody fragment, into cells. Where antibody fragments are used, the smallest inhibitory fragment that specifically binds to the binding domain of the target protein is prefened. For example, based upon the variable-region sequences of an antibody, peptide molecules can be designed that retain the ability to bind the target protein sequence. Such peptides can be synthesized chemically and/or produced by recombinant DNA technology. See, e.g., Marasco et al., Proc. Natl. Acad. Sci. USA, 90: 7889-7893 (1993). The formulation herein can also contain more than one active compound as necessary for the particular indication being treated, preferably those with complementary activities that do not adversely affect each other. Alternatively, or in addition, the composition can comprise an agent that enhances its function, such as, for example, a cytotoxic agent, cytokine, chemotherapeutic agent, or growth-inhibitory agent. Such molecules are suitably present in combination in amounts that are effective for the puφose intended.
The active ingredients can also be entrapped in microcapsules prepared, for example, by coacervation techniques or by interfacial polymerization, for example, hydroxymethylcellulose or gelatin-microcapsules and poly-(methylmethacrylate) microcapsules, respectively, in colloidal drug delivery systems (for example, liposomes, albumin microspheres, microemulsions, nano-particles, and nanocapsules) or in macroemulsions.
The formulations to be used for in vivo administration must be sterile. This is readily accomplished by filtration through sterile filtration membranes.
Sustained-release preparations can be prepared. Suitable examples of sustained-release preparations include semipermeable matrices of solid hydrophobic polymers containing the antibody, which matrices are in the form of shaped articles, e.g., films, or microcapsules.
Examples of sustained-release matrices include polyesters, hydrogels (for example, poly(2- hydroxyethyl-methacrylate), or poly(vinylalcohol)), polylactides (U.S. Pat. No. 3,773,919), copolymers of L-glutamic acid and γ ethyl-L-glutamate, non-degradable ethylene- vinyl acetate, degradable lactic acid-glycolic acid copolymers such as the LTJPRON DEPOT ™ (injectable microspheres composed of lactic acid-glycolic acid copolymer and leuprolide acetate), and poly-D-(-)-3-hydroxybutyric acid. While polymers such as ethylene- vinyl acetate and lactic acid-glycolic acid enable release of molecules for over 100 days, certain hydrogels release proteins for shorter time periods.
ELISA Assay
An agent for detecting an analyte protein is an antibody capable of binding to an analyte protein, preferably an antibody with a detectable label. Antibodies can be polyclonal, or more preferably, monoclonal. An intact antibody, or a fragment thereof (e.g., Fab or F(a )2) can be used. The term "labeled", with regard to the probe or antibody, is intended to encompass direct labeling of the probe or antibody by coupling (i.e., physically linking) a detectable substance to the probe or antibody, as well as indirect labeling of the probe or antibody by reactivity with another reagent that is directly labeled. Examples of indirect labeling include detection of a primary antibody using a fluorescently-labeled secondary antibody and end-labeling of a DNA probe with biotin such that it can be detected with fluorescently-labeled streptavidin. The term "biological sample" is intended to include tissues, cells and biological fluids isolated from a subject, as well as tissues, cells and fluids present within a subject. Included within the usage of the term "biological sample", therefore, is blood and a fraction or component of blood including blood serum, blood plasma, or lymph. That is, the detection method of the invention can be used to detect an analyte mRNA, protein, or genomic DNA in a biological sample in vitro as well as in vivo. For example, in vitro techniques for detection of an analyte mRNA include Northern hybridizations and in situ hybridizations. In vitro techniques for detection of an analyte protein include enzyme linked immunosorbent assays (ELISAs), Western blots, immunoprecipitations, and immunofluorescence. In vitro techniques for detection of an analyte genomic DNA include Southem hybridizations. Procedures for conducting immunoassays are described, for example in "ELISA: Theory and Practice: Methods in Molecular Biology", Vol. 42, J. R. Crowther (Ed.) Human Press, Totowa, NJ, 1995; "Immunoassay", E. Diamandis and T. Christopoulus, Academic Press, Inc., San Diego, CA, 1996; and "Practice and Thory of Enzyme Immunoassays", P. Tijssen, Elsevier Science Publishers, Amsterdam, 1985. Furthermore, in
vivo techniques for detection of an analyte protein include introducing into a subject a labeled anti-an analyte protein antibody. For example, the antibody can be labeled with a radioactive marker whose presence and location in a subject can be detected by standard imaging techniques.
NOVX Recombinant Expression Vectors and Host Cells
Another aspect of the invention pertains to vectors, preferably expression vectors, containing a nucleic acid encoding an NOVX protein, or derivatives, fragments, analogs or homologs thereof. As used herein, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. One type of vector is a "plasmid", which refers to a circular double stranded DNA loop into which additional DNA segments can be ligated. Another type of vector is a viral vector, wherein additional DNA segments can be ligated into the viral genome. Certain vectors are capable of autonomous replication in a host cell into which they are introduced (e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of a host cell upon introduction into the host cell, and thereby are replicated along with the host genome. Moreover, certain vectors are capable of directing the expression of genes to which they are operatively-linked. Such vectors are refened to herein as "expression vectors". In general, expression vectors of utility in recombinant DNA techniques are often in the form of plasmids. In the present specification, "plasmid" and "vector" can be used interchangeably as the plasmid is the most commonly used form of vector. However, the invention is intended to include such other forms of expression vectors, such as viral vectors (e.g., replication defective retroviruses, adenoviruses and adeno-associated viruses), which serve equivalent functions.
The recombinant expression vectors of the invention comprise a nucleic acid of the invention in a form suitable for expression of the nucleic acid in a host cell, which means that the recombinant expression vectors include one or more regulatory sequences, selected on the basis of the host cells to be used for expression, that is operatively-linked to the nucleic acid sequence to be expressed. Within a recombinant expression vector, "operably-linked" is intended to mean that the nucleotide sequence of interest is linked to the regulatory sequence(s) in a manner that allows for expression of the nucleotide sequence (e.g., in an in vitro transcription/translation system or in a host cell when the vector is introduced into the host cell).
The term "regulatory sequence" is intended to includes promoters, enhancers and other expression control elements (e.g., polyadenylation signals). Such regulatory sequences are described, for example, in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990). Regulatory sequences include those that direct constitutive expression of a nucleotide sequence in many types of host cell and those that direct expression of the nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). It will be appreciated by those skilled in the art that the design of the expression vector can depend on such factors as the choice of the host cell to be transformed, the level of expression of protein desired, etc. The expression vectors of the invention can be introduced into host cells to thereby produce proteins or peptides, including fusion proteins or peptides, encoded by nucleic acids as described herein (e.g., NOVX proteins, mutant forms of NOVX proteins, fusion proteins, etc.).
The recombinant expression vectors of the invention can be designed for expression of NOVX proteins in prokaryotic or eukaryotic cells. For example, NOVX proteins can be expressed in bacterial cells such as Escherichia coli, insect cells (using baculovirus expression vectors) yeast cells or mammalian cells. Suitable host cells are discussed further in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990). Alternatively, the recombinant expression vector can be transcribed and translated in vitro, for example using T7 promoter regulatory sequences and T7 polymerase. Expression of proteins in prokaryotes is most often carried out in Escherichia coli with vectors containing constitutive or inducible promoters directing the expression of either fusion or non- fusion proteins. Fusion vectors add a number of amino acids to a protein encoded therein, usually to the amino terminus of the recombinant protein. Such fusion vectors typically serve three puφoses: ( ) to increase expression of recombinant protein; (ii) to increase the solubility of the recombinant protein; and (iii) to aid in the purification of the recombinant protein by acting as a ligand in affinity purification. Often, in fusion expression vectors, a proteolytic cleavage site is introduced at the junction of the fusion moiety and the recombinant protein to enable separation of the recombinant protein from the fusion moiety subsequent to purification of the fusion protein. Such enzymes, and their cognate recognition sequences, include Factor Xa, thrombin and enterokinase. Typical fusion expression vectors include pGEX (Pharmacia Biotech Inc; Smith and Johnson, 1988. Gene 67: 31-40), pMAL (New England Biolabs, Beverly, Mass.) and pRIT5 (Pharmacia, Piscataway, N.J.) that fuse glutathione S-transferase (GST), maltose E binding protein, or protein A, respectively, to the target recombinant protein.
Examples of suitable inducible non-fusion E. coli expression vectors include pTrc (A rann et al, (1988) Gene 69:301-315) and pET 1 Id (Srudier et al, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990) 60-89).
One strategy to maximize recombinant protein expression in E. coli is to express the protein in a host bacteria with an impaired capacity to proteolytically cleave the recombinant protein. See, e.g., Gottesman, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990) 119-128. Another strategy is to alter the nucleic acid sequence of the nucleic acid to be inserted into an expression vector so that the individual codons for each amino acid are those preferentially utilized in E. coli (see, e.g., Wada, et al, 1992. Nucl. Acids Res. 20: 2111-2118). Such alteration of nucleic acid sequences of the invention can be carried out by standard DNA synthesis techniques.
In another embodiment, the NOVX expression vector is a yeast expression vector. Examples of vectors for expression in yeast Saccharomyces cerivisae include pYepSecl (Baldari, et al, 1987. EMBOJ. 6: 229-234), pMFa (Kurjan and Herskowitz, 1982. Cell 30: 933-943), pJRY88 (Schultz et al, 1987. Gene 54: 113-123), pYES2 (Invitrogen Coφoration, San Diego, Calif), and picZ (InVitrogen Coφ, San Diego, Calif).
Alternatively, NOVX can be expressed in insect cells using baculovirus expression vectors. Baculovirus vectors available for expression of proteins in cultured insect cells (e.g., SF9 cells) include the pAc series (Smith, et al, 1983. Mol. Cell. Biol. 3: 2156-2165) and the pVL series (Lucklow and Summers, 1989. Virology 170: 31-39).
In yet another embodiment, a nucleic acid of the invention is expressed in mammalian cells using a mammalian expression vector. Examples of mammalian expression vectors include pCDM8 (Seed, 1987. Nature 329: 840) and pMT2PC (Kaufman, et al, 1987. EMBO J. 6: 187-195). When used in mammalian cells, the expression vector's control functions are often provided by viral regulatory elements. For example, commonly used promoters are derived from polyoma, adenovirus 2, cytomegalovirus, and simian virus 40. For other suitable expression systems for both prokaryotic and eukaryotic cells see, e.g., Chapters 16 and 17 of Sambrook, et al, MOLECULAR CLONING: A LABORATORY MANUAL. 2nd ed., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., 1989.
In another embodiment, the recombinant mammalian expression vector is capable of directing expression of the nucleic acid preferentially in a particular cell type (e.g., tissue-specific regulatory elements are used to express the nucleic acid). Tissue-specific regulatory elements are known in the art. Non-limiting examples of suitable tissue-specific promoters include the albumin promoter (liver-specific; Pinkert, et al., 1987. Genes Dev. 1:
268-277), lymphoid-specific promoters (Calame and Eaton, 1988. Adv. Immunol. 43: 235-275), in particular promoters of T cell receptors (Winoto and Baltimore, 1989. EMBO J. 8: 729-733) and immunoglobulins (Banerji, et al, 1983. Cell 33: 729-740; Queen and Baltimore, 1983. Cell 33: 741-748), neuron-specific promoters (e.g., the neurofilament promoter; Byrne and Ruddle, 1989. Proc. Natl. Acad. Sci. USA 86: 5473-5477), pancreas-specific promoters (Edlund, et al, 1985. Science 230: 912-916), and mammary gland-specific promoters (e.g., milk whey promoter; U.S. Pat. No. 4,873,316 and European Application Publication No. 264,166). Developmentally-regulated promoters are also encompassed, e.g., the murine hox promoters (Kessel and Grass, 1990. Science 249: 374-379) and the α-fetoprotein promoter (Campes and Tilghman, 1989. Genes Dev. 3: 537-546).
The invention further provides a recombinant expression vector comprising a DNA molecule of the invention cloned into the expression vector in an antisense orientation. That is, the DNA molecule is operatively-linked to a regulatory sequence in a manner that allows for expression (by transcription of the DNA molecule) of an RNA molecule that is antisense to NOVX mRNA. Regulatory sequences operatively linked to a nucleic acid cloned in the antisense orientation can be chosen that direct the continuous expression of the antisense RNA molecule in a variety of cell types, for instance viral promoters and/or enhancers, or regulatory sequences can be chosen that direct constitutive, tissue specific or cell type specific expression of antisense RNA. The antisense expression vector can be in the form of a recombinant plasmid, phagemid or attenuated virus in which antisense nucleic acids are produced under the control of a high efficiency regulatory region, the activity of which can be determined by the cell type into which the vector is introduced. For a discussion of the regulation of gene expression using antisense genes see, e.g., Weintraub, et al, "Antisense RNA as a molecular tool for genetic analysis," Reviews-Trends in Genetics, Vol. 1(1) 1986.
Another aspect of the invention pertains to host cells into which a recombinant expression vector of the invention has been introduced. The terms "host cell" and "recombinant host cell" are used interchangeably herein. It is understood that such terms refer not only to the particular subject cell but also to the progeny or potential progeny of such a cell. Because certain modifications may occur in succeeding generations due to either mutation or environmental influences, such progeny may not, in fact, be identical to the parent cell, but are still included within the scope of the term as used herein.
A host cell can be any prokaryotic or eukaryotic cell. For example, NOVX protein can be expressed in bacterial cells such as E. coli, insect cells, yeast or mammalian cells (such as
Chinese hamster ovary cells (CHO) or COS cells). Other suitable host cells are known to those skilled in the art.
Vector DNA can be introduced into prokaryotic or eukaryotic cells via conventional transformation or transfection techniques. As used herein, the terms "transformation" and "transfection" are intended to refer to a variety of art-recognized techniques for introducing foreign nucleic acid (e.g., DNA) into a host cell, including calcium phosphate or calcium chloride co-precipitation, DEAE-dextran-mediated transfection, lipofection, or elecfroporation. Suitable methods for transforming or transfecting host cells can be found in Sambrook, et al. (MOLECULAR CLONING: A LABORATORY MANUAL. 2nd ed., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., 1989), and other laboratory manuals.
For stable transfection of mammalian cells, it is known that, depending upon the expression vector and transfection technique used, only a small fraction of cells may integrate the foreign DNA into their genome. In order to identify and select these integrants, a gene that encodes a selectable marker (e.g., resistance to antibiotics) is generally introduced into the host cells along with the gene of interest. Various selectable markers include those that confer resistance to drugs, such as G418, hygromycin and methotrexate. Nucleic acid encoding a selectable marker can be introduced into a host cell on the same vector as that encoding NOVX or can be introduced on a separate vector. Cells stably transfected with the introduced nucleic acid can be identified by drug selection (e.g., cells that have incoφorated the selectable marker gene will survive, while the other cells die).
A host cell of the invention, such as a prokaryotic or eukaryotic host cell in culture, can be used to produce (i.e., express) NOVX protein. Accordingly, the invention further provides methods for producing NOVX protein using the host cells of the invention. In one embodiment, the method comprises culturing the host cell of invention (into which a recombinant expression vector encoding NOVX protein has been introduced) in a suitable medium such that NOVX protein is produced. In another embodiment, the method further comprises isolating NOVX protein from the medium or the host cell.
Transgenic NOVX Animals
The host cells of the invention can also be used to produce non-human transgenic animals. For example, in one embodiment, a host cell of the invention is a fertilized oocyte or an embryonic stem cell into which NOVX protein-coding sequences have been introduced.
Such host cells can then be used to create non-human transgenic animals in which exogenous
NOVX sequences have been introduced into their genome or homologous recombinant animals in which endogenous NOVX sequences have been altered. Such animals are useful for studying the function and or activity of NOVX protein and for identifying and/or evaluating modulators of NOVX protein activity. As used herein, a "transgenic animal" is a non-human animal, preferably a mammal, more preferably a rodent such as a rat or mouse, in which one or more of the cells of the animal includes a fransgene. Other examples of transgenic animals include non-human primates, sheep, dogs, cows, goats, chickens, amphibians, etc. A transgene is exogenous DNA that is integrated into the genome of a cell from which a transgenic animal develops and that remains in the genome of the mature animal, thereby directing the expression of an encoded gene product in one or more cell types or tissues of the transgenic animal. As used herein, a "homologous recombinant animal" is a non-human animal, preferably a mammal, more preferably a mouse, in which an endogenous NOVX gene has been altered by homologous recombination between the endogenous gene and an exogenous DNA molecule introduced into a cell of the animal, e.g., an embryonic cell of the animal, prior to development of the animal.
A transgenic animal of the invention can be created by introducing NOVX-encoding nucleic acid into the male pronuclei of a fertilized oocyte (e.g., by microinjection, retroviral infection) and allowing the oocyte to develop in a pseudopregnant female foster animal. The human NOVX cDNA sequences SEQ ID NO: 2n-l, wherein n is an integer between 1 and 178 can be introduced as a fransgene into the genome of a non-human animal. Alternatively, a non-human homologue of the human NOVX gene, such as a mouse NOVX gene, can be isolated based on hybridization to the human NOVX cDNA (described further supra) and used as a transgene. Intronic sequences and polyadenylation signals can also be included in the transgene to increase the efficiency of expression of the transgene. A tissue-specific regulatory sequence(s) can be operably-linked to the NOVX transgene to direct expression of NOVX protein to particular cells. Methods for generating transgenic animals via embryo manipulation and microinjection, particularly animals such as mice, have become conventional in the art and are described, for example, in U.S. Patent Nos. 4,736,866; 4,870,009; and 4,873,191; and Hogan, 1986. In: MANIPULATING THE MOUSE EMBRYO, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N. Y. Similar methods are used for production of other transgenic animals. A transgenic founder animal can be identified based upon the presence of the NOVX transgene in its genome and/or expression of NOVX mRNA in tissues or cells of the animals. A transgenic founder animal can then be used to breed additional animals canying the fransgene. Moreover, transgenic animals carrying a fransgene-
encoding NOVX protein can further be bred to other transgenic animals carrying other transgenes.
To create a homologous recombinant animal, a vector is prepared which contains at least a portion of an NOVX gene into which a deletion, addition or substitution has been introduced to thereby alter, e.g., functionally disrupt, the NOVX gene. The NOVX gene can be a human gene (e.g., the cDNA of SEQ TD NO: 2n-l, wherein n is an integer between 1 and 178), but more preferably, is a non-human homologue of a human NOVX gene. For example, a mouse homologue of human NOVX gene of SEQ TD NO: 2n-l, wherein n is an integer between 1 and 178 can be used to construct a homologous recombination vector suitable for altering an endogenous NOVX gene in the mouse genome. In one embodiment, the vector is designed such that, upon homologous recombination, the endogenous NOVX gene is functionally disrapted (i.e., no longer encodes a functional protein; also refened to as a "knock out" vector).
Alternatively, the vector can be designed such that, upon homologous recombination, the endogenous NOVX gene is mutated or otherwise altered but still encodes functional protein (e.g., the upstream regulatory region can be altered to thereby alter the expression of the endogenous NOVX protein). In the homologous recombination vector, the altered portion of the NOVX gene is flanked at its 5'- and 3'-termini by additional nucleic acid of the NOVX gene to allow for homologous recombination to occur between the exogenous NOVX gene carried by the vector and an endogenous NOVX gene in an embryonic stem cell. The additional flanking NOVX nucleic acid is of sufficient length for successful homologous recombination with the endogenous gene. Typically, several kilobases of flanking DNA (both at the 5'- and 3'-termini) are included in the vector. See, e.g., Thomas, et al, 1987. Cell 51 : 503 for a description of homologous recombination vectors. The vector is ten introduced into an embryonic stem cell line (e.g., by elecfroporation) and cells in which the introduced NOVX gene has homologously-recombined with the endogenous NOVX gene are selected. See, e.g., Li, et al, 1992. Cell 69: 915.
The selected cells are then injected into a blastocyst of an animal (e.g., a mouse) to form aggregation chimeras. See, e.g., Bradley, 1987. In: TERATOCARCINOMAS AND EMBRYONIC STEM CELLS: A PRACTICAL APPROACH, Robertson, ed. ERL, Oxford, pp. 113-152. A chimeric embryo can then be implanted into a suitable pseudopregnant female foster animal and the embryo brought to term. Progeny harboring the homologously-recombined DNA in their germ cells can be used to breed animals in which all cells of the animal contain the homologously-recombined DNA by germline transmission of the fransgene. Methods for
constructing homologous recombination vectors and homologous recombinant animals are described further in Bradley, 1991. Curr. Opin. Biotechnol. 2: 823-829; PCT International Publication Nos.: WO 90/11354; WO 91/01140; WO 92/0968; and WO 93/04169.
In another embodiment, transgenic non-humans animals can be produced that contain selected systems that allow for regulated expression of the transgene. One example of such a system is the cre/loxP recombinase system of bacteriophage PI. For a description of the cre/loxP recombinase system, See, e.g., Lakso, et al, 1992. Proc. Natl. Acad. Sci. USA 89: 6232-6236. Another example of a recombinase system is the FLP recombinase system of Saccharomyces cerevisiae. See, O'Gorman, et al, 1991. Science 251:1351-1355. If a cre/loxP recombinase system is used to regulate expression of the transgene, animals containing transgenes encoding both the Cre recombinase and a selected protein are required. Such animals can be provided through the construction of "double" transgenic animals, e.g., by mating two transgenic animals, one containing a transgene encoding a selected protein and the other containing a transgene encoding a recombinase.
Clones of the non-human transgenic animals described herein can also be produced according to the methods described in Wilmut, et al, 1997. Nature 385: 810-813. In brief, a cell (e.g., a somatic cell) from the transgenic animal can be isolated and induced to exit the growth cycle and enter G0 phase. The quiescent cell can then be fused, e.g. , through the use of electrical pulses, to an enucleated oocyte from an animal of the same species from which the quiescent cell is isolated. The reconstructed oocyte is then cultured such that it develops to morula or blastocyte and then transfened to pseudopregnant female foster animal. The offspring borne of this female foster animal will be a clone of the animal from which the cell (e.g., the somatic cell) is isolated.
Pharmaceutical Compositions
The NOVX nucleic acid molecules, NOVX proteins, and anti-NOVX antibodies (also refened to herein as "active compounds") of the invention, and derivatives, fragments, analogs and homologs thereof, can be incoφorated into pharmaceutical compositions suitable for administration. Such compositions typically comprise the nucleic acid molecule, protein, or antibody and a pharmaceutically acceptable carrier. As used herein, "pharmaceutically acceptable carrier" is intended to include any and all solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic and absoφtion delaying agents, and the like, compatible with pharmaceutical administration. Suitable carriers are described in the most recent edition of Remington's Pharmaceutical Sciences, a standard reference text in the field,
which is incoφorated herein by reference. Prefened examples of such carriers or diluents include, but are not limited to, water, saline, finger's solutions, dextrose solution, and 5%. human serum albumin. Liposomes and non-aqueous vehicles such as fixed oils may also be used. The use of such media and agents for pharmaceutically active substances is well known in the art. Except insofar as any conventional media or agent is incompatible with the active compound, use thereof in the compositions is contemplated. Supplementary active compounds can also be incoφorated into the compositions.
A pharmaceutical composition of the invention is formulated to be compatible with its intended route of administration. Examples of routes of administration include parenteral, e.g., intravenous, intradermal, subcutaneous, oral (e.g., inhalation), transdermal (i.e., topical), transmucosal, and rectal administration. Solutions or suspensions used for parenteral, intradermal, or subcutaneous application can include the following components: a sterile diluent such as water for injection, saline solution, fixed oils, polyethylene glycols, glycerine, propylene giycol or other synthetic solvents; antibacterial agents such as benzyl alcohol or methyl parabens; antioxidants such as ascorbic acid or sodium bisulfite; chelating agents such as ethylenediaminetetraacetic acid (EDTA); buffers such as acetates, citrates or phosphates, and agents for the adjustment of tonicity such as sodium chloride or dextrose. The pH can be adjusted with acids or bases, such as hydrochloric acid or sodium hydroxide. The parenteral preparation can be enclosed in ampoules, disposable syringes or multiple dose vials made of glass or plastic.
Phaπnaceutical compositions suitable for injectable use include sterile aqueous solutions (where water soluble) or dispersions and sterile powders for the extemporaneous preparation of sterile injectable solutions or dispersion. For intravenous administration, suitable carriers include physiological saline, bacteriostatic water, Cremophor EL (BASF, Parsippany, N.J.) or phosphate buffered saline (PBS). In all cases, the composition must be sterile and should be fluid to the extent that easy syringeability exists. It must be stable under the conditions of manufacture and storage and must be preserved against the contaminating action of microorganisms such as bacteria and fungi. The carrier can be a solvent or dispersion medium containing, for example, water, ethanol, polyol (for example, glycerol, propylene giycol, and liquid polyethylene giycol, and the like), and suitable mixtures thereof. The proper fluidity can be maintained, for example, by the use of a coating such as lecithin, by the maintenance of the required particle size in the case of dispersion and by the use of surfactants. Prevention of the action of microorganisms can be achieved by various antibacterial and antifungal agents, for example, parabens, chlorobutanol, phenol, ascorbic
acid, thimerosal, and the like. In many cases, it will be preferable to include isotonic agents, for example, sugars, polyalcohols such as manitol, sorbitol, sodium chloride in the composition. Prolonged absoφtion of the injectable compositions can be brought about by including in the composition an agent which delays absoφtion, for example, aluminum monostearate and gelatin.
Sterile injectable solutions can be prepared by incoφorating the active compound (e.g., an NOVX protein or anti-NOVX antibody) in the required amount in an appropriate solvent with one or a combination of ingredients enumerated above, as required, followed by filtered sterilization. Generally, dispersions are prepared by incoφorating the active compound into a sterile vehicle that contains a basic dispersion medium and the required other ingredients from those enumerated above. In the case of sterile powders for the preparation of sterile injectable solutions, methods of preparation are vacuum drying and freeze-drying that yields a powder of the active ingredient plus any additional desired ingredient from a previously sterile-filtered solution thereof.
Oral compositions generally include an inert diluent or an edible carrier. They can be enclosed in gelatin capsules or compressed into tablets. For the puφose of oral therapeutic administration, the active compound can be incoφorated with excipients and used in the form of tablets, troches, or capsules. Oral compositions can also be prepared using a fluid carrier for use as a mouthwash, wherein the compound in the fluid carrier is applied orally and swished and expectorated or swallowed. Pharmaceutically compatible binding agents, and/or adjuvant materials can be included as part of the composition. The tablets, pills, capsules, troches and the like can contain any of the following ingredients, or compounds of a similar nature: a binder such as microcrystalline cellulose, gum tragacanth or gelatin; an excipient such as starch or lactose, a disintegrating agent such as alginic acid, Primogel, or corn starch; a lubricant such as magnesium stearate or Sterotes; a glidant such as colloidal silicon dioxide; a sweetening agent such as sucrose or saccharin; or a flavoring agent such as peppermint, methyl salicylate, or orange flavoring.
For administration by inhalation, the compounds are delivered in the form of an aerosol spray from pressured container or dispenser which contains a suitable propellant, e.g., a gas such as carbon dioxide, or a nebulizer.
Systemic administration can also be by transmucosal or transdermal means. For transmucosal or transdermal administration, penetrants appropriate to the barrier to be permeated are used in the formulation. Such penetrants are generally known in the art, and include, for example, for transmucosal administration, detergents, bile salts, and fusidic acid
derivatives. Transmucosal administration can be accomplished through the use of nasal sprays or suppositories. For transdermal administration, the active compounds are formulated into ointments, salves, gels, or creams as generally known in the art.
The compounds can also be prepared in the form of suppositories (e.g., with conventional suppository bases such as cocoa butter and other glycerides) or retention enemas for rectal delivery.
In one embodiment, the active compounds are prepared with carriers that will protect the compound against rapid elimination from the body, such as a controlled release formulation, including implants and microencapsulated delivery systems. Biodegradable, biocompatible polymers can be used, such as ethylene vinyl acetate, polyanhydrides, polyglycolic acid, collagen, polyorthoesters, and polylactic acid. Methods for preparation of such formulations will be apparent to those skilled in the art. The materials can also be obtained commercially from Alza Coφoration and Nova Pharmaceuticals, Inc. Liposomal suspensions (including liposomes targeted to infected cells with monoclonal antibodies to viral antigens) can also be used as pharmaceutically acceptable carriers. These can be prepared according to methods known to those skilled in the art, for example, as described in U.S. Patent No. 4,522,811.
It is especially advantageous to formulate oral or parenteral compositions in dosage unit form for ease of administration and uniformity of dosage. Dosage unit form as used herein refers to physically discrete units suited as unitary dosages for the subject to be treated; each unit containing a predetermined quantity of active compound calculated to produce the desired therapeutic effect in association with the required pharmaceutical carrier. The specification for the dosage unit forms of the invention are dictated by and directly dependent on the unique characteristics of the active compound and the particular therapeutic effect to be achieved, and the limitations inherent in the art of compounding such an active compound for the treatment of individuals.
The nucleic acid molecules of the invention can be inserted into vectors and used as gene therapy vectors. Gene therapy vectors can be delivered to a subject by, for example, intravenous injection, local administration (see, e.g., U.S. Patent No. 5,328,470) or by stereotactic injection (see, e.g., Chen, et al, 1994. Proc. Natl. Acad. Sci. USA 91: 3054-3057). The pharmaceutical preparation of the gene therapy vector can include the gene therapy vector in an acceptable diluent, or can comprise a slow release matrix in which the gene delivery vehicle is imbedded. Alternatively, where the complete gene delivery vector can be produced
intact from recombinant cells, e.g., retroviral vectors, the pharmaceutical preparation can include one or more cells that produce the gene delivery system.
The pharmaceutical compositions can be included in a container, pack, or dispenser together with instructions for administration.
Screening and Detection Methods
The isolated nucleic acid molecules of the invention can be used to express NOVX protein (e.g., via a recombinant expression vector in a host cell in gene therapy applications), to detect NOVX mRNA (e.g., in a biological sample) or a genetic lesion in an NOVX gene, and to modulate NOVX activity, as described further, below. In addition, the NOVX proteins can be used to screen drugs or compounds that modulate the NOVX protein activity or expression as well as to treat disorders characterized by insufficient or excessive production of NOVX protein or production of NOVX protein forms that have decreased or abenant activity compared to NOVX wild-type protein (e.g.; diabetes (regulates insulin release); obesity (binds and fransport lipids); metabolic disturbances associated with obesity, the metabolic syndrome X as well as anorexia and wasting disorders associated with chronic diseases and various cancers, and infectious disease(possesses anti-microbial activity) and the various dyslipidemias. In addition, the anti-NOVX antibodies of the invention can be used to detect and isolate NOVX proteins and modulate NOVX activity. In yet a further aspect, the invention can be used in methods to influence appetite, absoφtion of nutrients and the disposition of metabolic substrates in both a positive and negative fashion.
The invention further pertains to novel agents identified by the screening assays described herein and uses thereof for treatments as described, supra.
Screening Assays
The invention provides a method (also refened to herein as a "screening assay") for identifying modulators, i.e., candidate or test compounds or agents (e.g., peptides, peptidomimetics, small molecules or other drugs) that bind to NOVX proteins or have a stimulatory or inhibitory effect on, e.g., NOVX protein expression or NOVX protein activity. The invention also includes compounds identified in the screening assays described herein. In one embodiment, the invention provides assays for screening candidate or test compounds which bind to or modulate the activity of the membrane-bound form of an NOVX protein or polypeptide or biologically-active portion thereof. The test compounds of the invention can be
obtained using any of the numerous approaches in combinatorial library methods known in the art, including: biological libraries; spatially addressable parallel solid phase or solution phase libraries; synthetic library methods requiring deconvolution; the "one-bead one-compound" library method; and synthetic library methods using affinity chromatography selection. The biological library approach is limited to peptide libraries, while the other four approaches are applicable to peptide, non-peptide oligomer or small molecule libraries of compounds. See, e.g., Lam, 1997 '. Anticancer Drug Design 12: 145.
A "small molecule" as used herein, is meant to refer to a composition that has a molecular weight of less than about 5 kD and most preferably less than about 4 kD. Small molecules can be, e.g., nucleic acids, peptides, polypeptides, peptidomimetics, carbohydrates, lipids or other organic or inorganic molecules. Libraries of chemical and/or biological mixtures, such as fungal, bacterial, or algal extracts, are known in the art and can be screened with any of the assays of the invention.
Examples of methods for the synthesis of molecular libraries can be found in the art, for example in: DeWitt, et al, 1993. Proc. Natl. Acad. Sci. U.S.A. 90: 6909; Erb, et al, 1994. Proc. Natl. Acad. Sci. U.S.A. 91: 11422; Zuckermann, et al, 1994. J. Med. Chem. 37: 2678; Cho, et al, 1993. Science 261 : 1303; Canell, et al, 1994. Angew. Chem. Int. Ed. Engl 33: 2059; Carell, et al, 1994. Angew. Chem. Int. Ed. Engl. 33: 2061; and Gallop, et al, 1994. J. Med. Chem. 37: 1233.
Libraries of compounds may be presented in solution (e.g., Houghten, 1992. Biotechniques 13: 412-421), or on beads (Lam, 1991. Nature 354: 82-84), on chips (Fodor, 1993. Nature 364: 555-556), bacteria (Ladner, U.S. Patent No. 5,223,409), spores (Ladner, U.S. Patent 5,233,409), plasmids (Cull, et al, 1992. Proc. Natl. Acad. Sci. USA 89: 1865-1869) or on phage (Scott and Smith, 1990. Science 249: 386-390; Devlin, 1990. Science 249: 404-406; Cwirla, et al, 1990. Proc. Natl. Acad. Sci. U.S.A. 87: 6378-6382; Felici, 1991. J. Mol. Biol. 222: 301-310; Ladner, U.S. Patent No. 5,233,409.).
In one embodiment, an assay is a cell-based assay in which a cell which expresses a membrane-bound form of NOVX protein, or a biologically-active portion thereof, on the cell surface is contacted with a test compound and the ability of the test compound to bind to an NOVX protein determined. The cell, for example, can of mammalian origin or a yeast cell. Determining the ability of the test compound to bind to the NOVX protein can be accomplished, for example, by coupling the test compound with a radioisotope or enzymatic label such that binding of the test compound to the NOVX protein or biologically-active portion thereof can be determined by detecting the labeled compound in a complex. For
example, test compounds can be labeled with 1251, 35S, 14C, or 3H, either directly or indirectly, and the radioisotope detected by direct counting of radioemission or by scintillation counting. Alternatively, test compounds can be enzymatically-labeled with, for example, horseradish peroxidase, alkaline phosphatase, or luciferase, and the enzymatic label detected by determination of conversion of an appropriate substrate to product. In one embodiment, the assay comprises contacting a cell which expresses a membrane-bound form of NOVX protein, or a biologically-active portion thereof, on the cell surface with a known compound which binds NOVX to form an assay mixture, contacting the assay mixture with a test compound, and determining the ability of the test compound to interact with an NOVX protein, wherein determining the ability of the test compound to interact with an NOVX protein comprises determining the ability of the test compound to preferentially bind to NOVX protein or a biologically-active portion thereof as compared to the known compound.
In another embodiment, an assay is a cell-based assay comprising contacting a cell expressing a membrane-bound form of NOVX protein, or a biologically-active portion thereof, on the cell surface with a test compound and determining the ability of the test compound to modulate (e.g., stimulate or inhibit) the activity of the NOVX protein or biologically-active portion thereof. Determining the ability of the test compound to modulate the activity of NOVX or a bio logically- active portion thereof can be accomplished, for example, by determining the ability of the NOVX protein to bind to or interact with an NOVX target molecule. As used herein, a "target molecule" is a molecule with which an NOVX protein binds or interacts in nature, for example, a molecule on the surface of a cell which expresses an NOVX interacting protein, a molecule on the surface of a second cell, a molecule in the extracellular milieu, a molecule associated with the internal surface of a cell membrane or a cytoplasmic molecule. An NOVX target molecule can be a non-NOVX molecule or an NOVX protein or polypeptide of the invention. In one embodiment, an NOVX target molecule is a component of a signal transduction pathway that facilitates transduction of an extracellular signal (e.g. a signal generated by binding of a compound to a membrane-bound NOVX molecule) through the cell membrane and into the cell. The target, for example, can be a second intercellular protein that has catalytic activity or a protein that facilitates the association of downstream signaling molecules with NOVX.
Determining the ability of the NOVX protein to bind to or interact with an NOVX target molecule can be accomplished by one of the methods described above for determining direct binding. In one embodiment, determining the ability of the NOVX protein to bind to or interact with an NOVX target molecule can be accomplished by determining the activity of the
target molecule. For example, the activity of the target molecule can be determined by detecting induction of a cellular second messenger of the target (i.e. intracellular Ca2+, diacylglycerol, EP3, etc.), detecting catalytic/enzymatic activity of the target an appropriate substrate, detecting the induction of a reporter gene (comprising an NOVX-responsive regulatory element operatively linked to a nucleic acid encoding a detectable marker, e.g. , luciferase), or detecting a cellular response, for example, cell survival, cellular differentiation, or cell proliferation.
In yet another embodiment, an assay of the invention is a cell-free assay comprising contacting an NOVX protein or biologically-active portion thereof with a test compound and determining the ability of the test compound to bind to the NOVX protein or biologically- active portion thereof. Binding of the test compound to the NOVX protein can be determined either directly or indirectly as described above. In one such embodiment, the assay comprises contacting the NOVX protein or biologically-active portion thereof with a known compound which binds NOVX to form an assay mixture, contacting the assay mixture with a test compound, and determining the ability of the test compound to interact with an NOVX protein, wherein determining the ability of the test compound to interact with an NOVX protein comprises determining the ability of the test compound to preferentially bind to NOVX or biologically-active portion thereof as compared to the known compound.
In still another embodiment, an assay is a cell-free assay comprising contacting NOVX protein or biologically-active portion thereof with a test compound and determining the ability of the test compound to modulate (e.g. stimulate or inhibit) the activity of the NOVX protein or biologically-active portion thereof. Determining the ability of the test compound to modulate the activity of NOVX can be accomplished, for example, by determining the ability of the NOVX protein to bind to an NOVX target molecule by one of the methods described above for determining direct binding. In an alternative embodiment, determining the ability of the test compound to modulate the activity of NOVX protein can be accomplished by determining the ability of the NOVX protein further modulate an NOVX target molecule. For example, the catalytic/enzymatic activity of the target molecule on an appropriate substrate can be determined as described, supra.
In yet another embodiment, the cell-free assay comprises contacting the NOVX protein or biologically-active portion thereof with a known compound which binds NOVX protein to form an assay mixture, contacting the assay mixture with a test compound, and determining the ability of the test compound to interact with an NOVX protein, wherein determining the ability of the test compound to interact with an NOVX protein comprises determining the
ability of the NOVX protein to preferentially bind to or modulate the activity of an NOVX target molecule.
The cell- free assays of the invention are amenable to use of both the soluble form or the membrane-bound form of NOVX protein. In the case of cell-free assays comprising the membrane-bound form of NOVX protein, it may be desirable to utilize a solubilizing agent such that the membrane-bound form of NOVX protein is maintained in solution. Examples of such solubilizing agents include non-ionic detergents such as n-octylglucoside, n-dodecylglucoside, n-dodecylmaltoside, octanoyl-N-methylglucamide, decanoyl-N-methylglucamide, Triton® X-100, Triton® X-l 14, Thesit®, Isotridecypoly(ethylene giycol ether)n, N-dodecyl~N,N-dimethyl-3-ammonio-l -propane sulfonate, 3-(3-cholamidopropyl) dimethylamminiol-1 -propane sulfonate (CHAPS), or 3-(3-cholamidopropyl)dimethylamminiol-2-hydroxy- 1 -propane sulfonate (CHAPSO).
In more than one embodiment of the above assay methods of the invention, it may be desirable to immobilize either NOVX protein or its target molecule to facilitate separation of complexed from uncomplexed forms of one or both of the proteins, as well as to accommodate automation of the assay. Binding of a test compound to NOVX protein, or interaction of NOVX protein with a target molecule in the presence and absence of a candidate compound, can be accomplished in any vessel suitable for containing the reactants. Examples of such vessels include microtiter plates, test tubes, and micro-centrifuge tubes. In one embodiment, a fusion protein can be provided that adds a domain that allows one or both of the proteins to be bound to a matrix. For example, GST-NOVX fusion proteins or GST-target fusion proteins can be adsorbed onto glutathione sepharose beads (Sigma Chemical, St. Louis, MO) or glutathione derivatized microtiter plates, that are then combined with the test compound or the test compound and either the non-adsorbed target protein or NOVX protein, and the mixture is incubated under conditions conducive to complex formation (e.g., at physiological conditions for salt and pH). Following incubation, the beads or microtiter plate wells are washed to remove any unbound components, the matrix immobilized in the case of beads, complex determined either directly or indirectly, for example, as described, supra. Alternatively, the complexes can be dissociated from the matrix, and the level of NOVX protein binding or activity determined using standard techniques.
Other techniques for immobilizing proteins on matrices can also be used in the screening assays of the invention. For example, either the NOVX protein or its target molecule can be immobilized utilizing conjugation of biotin and streptavidin. Biotinylated NOVX protein or target molecules can be prepared from biotin-NHS
(N-hydroxy-succinimide) using techniques well-known within the art (e.g. , biotinylation kit, Pierce Chemicals, Rockford, 111.), and immobilized in the wells of streptavidin-coated 96 well plates (Pierce Chemical). Alternatively, antibodies reactive with NOVX protein or target molecules, but which do not interfere with binding of the NOVX protein to its target molecule, can be derivatized to the wells of the plate, and unbound target or NOVX protein trapped in the wells by antibody conjugation. Methods for detecting such complexes, in addition to those described above for the GST-immobilized complexes, include immunodetection of complexes using antibodies reactive with the NOVX protein or target molecule, as well as enzyme-linked assays that rely on detecting an enzymatic activity associated with the NOVX protein or target molecule.
In another embodiment, modulators of NOVX protein expression are identified in a method wherein a cell is contacted with a candidate compound and the expression of NOVX mRNA or protein in the cell is determined. The level of expression of NOVX mRNA or protein in the presence of the candidate compound is compared to the level of expression of NOVX mRNA or protein in the absence of the candidate compound. The candidate compound can then be identified as a modulator of NOVX mRNA or protein expression based upon this comparison. For example, when expression of NOVX mRNA or protein is greater (i.e., statistically significantly greater) in the presence of the candidate compound than in its absence, the candidate compound is identified as a stimulator of NOVX mRNA or protein expression. Alternatively, when expression of NOVX mRNA or protein is less (statistically significantly less) in the presence of the candidate compound than in its absence, the candidate compound is identified as an inhibitor of NOVX mRNA or protein expression. The level of NOVX mRNA or protein expression in the cells can be determined by methods described herein for detecting NOVX mRNA or protein.
In yet another aspect of the invention, the NOVX proteins can be used as "bait proteins" in a two-hybrid assay or three hybrid assay (see, e.g., U.S. Patent No. 5,283,317; Zervos, et al, 1993. Ce/772: 223-232; Madura, et al, 1993. J. Biol. Chem. 268: 12046-12054; Barrel, et al, 1993. Biotechniques 14: 920-924; Iwabuchi, et al, 1993. Oncogene 8: 1693-1696; and Brent WO 94/10300), to identify other proteins that bind to or interact with NOVX ("NOVX-binding proteins" or "NOVX-bp") and modulate NOVX activity. Such NONX-binding proteins are also likely to be involved in the propagation of signals by the ΝOVX proteins as, for example, upstream or downstream elements of the ΝOVX pathway.
The two-hybrid system is based on the modular nature of most transcription factors, which consist of separable DΝA-binding and activation domains. Briefly, the assay utilizes
two different DNA constracts. In one construct, the gene that codes for NOVX is fused to a gene encoding the DNA binding domain of a known franscription factor (e.g., GAL-4). In the other construct, a DNA sequence, from a library of DNA sequences, that encodes an unidentified protein ("prey" or "sample") is fused to a gene that codes for the activation domain of the known transcription factor. If the "bait" and the "prey" proteins are able to interact, in vivo, forming an NOVX-dependent complex, the DNA-binding and activation domains of the transcription factor are brought into close proximity. This proximity allows transcription of a reporter gene (e.g., LacZ) that is operably linked to a transcriptional regulatory site responsive to the transcription factor. Expression of the reporter gene can be detected and cell colonies containing the functional transcription factor can be isolated and used to obtain the cloned gene that encodes the protein which interacts with NOVX.
The invention further pertains to novel agents identified by the aforementioned screening assays and uses thereof for treatments as described herein.
Detection Assays
Portions or fragments of the cDNA sequences identified herein (and the conesponding complete gene sequences) can be used in numerous ways as polynucleotide reagents. By way of example, and not of limitation, these sequences can be used to: (i) map their respective genes on a chromosome; and, thus, locate gene regions associated with genetic disease; (ii) identify an individual from a minute biological sample (tissue typing); and (iii) aid in forensic identification of a biological sample. Some of these applications are described in the subsections, below.
Chromosome Mapping
Once the sequence (or a portion of the sequence) of a gene has been isolated, this sequence can be used to map the location of the gene on a chromosome. This process is called chromosome mapping. Accordingly, portions or fragments of the NOVX sequences, SEQ TD NO: 2n-l, wherein n is an integer between 1 and 178, or fragments or derivatives thereof, can be used to map the location of the NOVX genes, respectively, on a chromosome. The mapping of the NOVX sequences to chromosomes is an important first step in conelating these sequences with genes associated with disease.
Briefly, NOVX genes can be mapped to chromosomes by preparing PCR primers (preferably 15-25 bp in length) from the NOVX sequences. Computer analysis of the NOVX, sequences can be used to rapidly select primers that do not span more than one exon in the
genomic DNA, thus complicating the amplification process. These primers can then be used for PCR screening of somatic cell hybrids containing individual human chromosomes. Only those hybrids containing the human gene conesponding to the NOVX sequences will yield an amplified fragment.
Somatic cell hybrids are prepared by fusing somatic cells from different mammals (e.g., human and mouse cells). As hybrids of human and mouse cells grow and divide, they gradually lose human chromosomes in random order, but retain the mouse chromosomes. By using media in which mouse cells cannot grow, because they lack a particular enzyme, but in which human cells can, the one human chromosome that contains the gene encoding the needed enzyme will be retained. By using various media, panels of hybrid cell lines can be established. Each cell line in a panel contains either a single human chromosome or a small number of human chromosomes, and a full set of mouse chromosomes, allowing easy mapping of individual genes to specific human chromosomes. See, e.g., D'Eustachio, et al, 1983. Science 220: 919-924. Somatic cell hybrids containing only fragments of human chromosomes can also be produced by using human chromosomes with translocations and deletions.
PCR mapping of somatic cell hybrids is a rapid procedure for assigning a particular sequence to a particular chromosome. Three or more sequences can be assigned per day using a single thermal cycler. Using the NOVX sequences to design oligonucleotide primers, sub- localization can be achieved with panels of fragments from specific chromosomes. Fluorescence in situ hybridization (FISH) of a DNA sequence to a metaphase chromosomal spread can further be used to provide a precise chromosomal location in one step. Chromosome spreads can be made using cells whose division has been blocked in metaphase by a chemical like colcemid that disrapts the mitotic spindle. The chromosomes can be treated briefly with trypsin, and then stained with Giemsa. A pattern of light and dark bands develops on each chromosome, so that the chromosomes can be identified individually. The FISH technique can be used with a DNA sequence as short as 500 or 600 bases. However, clones larger than 1,000 bases have a higher likelihood of binding to a unique chromosomal location with sufficient signal intensity for simple detection. Preferably 1 ,000 bases, and more preferably 2,000 bases, will suffice to get good results at a reasonable amount of time. For a review of this technique, see, Verma, et al, HUMAN CHROMOSOMES: A MANUAL OF BASIC TECHNIQUES (Pergamon Press, New York 1988).
Reagents for chromosome mapping can be used individually to mark a single chromosome or a single site on that chromosome, or panels of reagents can be used for
marking multiple sites and/or multiple chromosomes. Reagents conesponding to noncoding regions of the genes actually are prefened for mapping puφoses. Coding sequences are more likely to be conserved within gene families, thus increasing the chance of cross hybridizations during chromosomal mapping.
Once a sequence has been mapped to a precise chromosomal location, the physical position of the sequence on the chromosome can be conelated with genetic map data. Such data are found, e.g., in McKusick, MENDELIAN INHERITANCE IN MAN, available on-line through Johns Hopkins University Welch Medical Library). The relationship between genes and disease, mapped to the same chromosomal region, can then be identified through linkage analysis (co-inheritance of physically adjacent genes), described in, e.g., Egeland, et al, 1987. Nature, 325: 783-787.
Moreover, differences in the DNA sequences between individuals affected and unaffected with a disease associated with the NOVX gene, can be determined. If a mutation is observed in some or all of the affected individuals but not in any unaffected individuals, then the mutation is likely to be the causative agent of the particular disease. Comparison of affected and unaffected individuals generally involves first looking for structural alterations in the chromosomes, such as deletions or translocations that are visible from chromosome spreads or detectable using PCR based on that DNA sequence. Ultimately, complete sequencing of genes from several individuals can be performed to confirm the presence of a mutation and to distinguish mutations from polymoφhisms.
Tissue Typing
The NOVX sequences of the invention can also be used to identify individuals from minute biological samples. In this technique, an individual's genomic DNA is digested with one or more restriction enzymes, and probed on a Southern blot to yield unique bands for identification. The sequences of the invention are useful as additional DNA markers for RFLP ("restriction fragment length polymoφhisms," described in U.S. Patent No. 5,272,057). Furthermore, the sequences of the invention can be used to provide an alternative technique that determines the actual base-by-base DNA sequence of selected portions of an individual's genome. Thus, the NOVX sequences described herein can be used to prepare two PCR primers from the 5'- and 3'-termini of the sequences. These primers can then be used to amplify an individual's DNA and subsequently sequence it.
Panels of conesponding DNA sequences from individuals, prepared in this manner, can provide unique individual identifications, as each individual will have a unique set of such
DNA sequences due to allelic differences. The sequences of the invention can be used to obtain such identification sequences from individuals and from tissue. The NOVX sequences of the invention uniquely represent portions of the human genome. Allelic variation occurs to some degree in the coding regions of these sequences, and to a greater degree in the noncoding regions. It is estimated that allelic variation between individual humans occurs with a frequency of about once per each 500 bases. Much of the allelic variation is due to single nucleotide polymoφhisms (SNPs), which include restriction fragment length polymoφhisms (RFLPs).
Each of the sequences described herein can, to some degree, be used as a standard against which DNA from an individual can be compared for identification puφoses. Because greater numbers of polymoφhisms occur in the noncoding regions, fewer sequences are necessary to differentiate individuals. The noncoding sequences can comfortably provide positive individual identification with a panel of perhaps 10 to 1,000 primers that each yield a noncoding amplified sequence of 100 bases. If predicted coding sequences, such as those in SEQ TD NO: 2n-l, wherein n is an integer between 1 and 178 are used, a more appropriate number of primers for positive individual identification would be 500-2,000.
Predictive Medicine
The invention also pertains to the field of predictive medicine in which diagnostic assays, prognostic assays, pharmacogenomics, and monitoring clinical trials are used for prognostic (predictive) puφoses to thereby treat an individual prophylactically. Accordingly, one aspect of the invention relates to diagnostic assays for determining NOVX protein and/or nucleic acid expression as well as NOVX activity, in the context of a biological sample (e.g., blood, serum, cells, tissue) to thereby determine whether an individual is afflicted with a disease or disorder, or is at risk of developing a disorder, associated with abenant NOVX expression or activity. The disorders include metabolic disorders, diabetes, obesity, infectious disease, anorexia, cancer-associated cachexia, cancer, neurodegenerative disorders, Alzheimer's Disease, Parkinson's Disorder, immune disorders, and hematopoietic disorders, and the various dyslipidemias, metabolic disturbances associated with obesity, the metabolic syndrome X and wasting disorders associated with chronic diseases and various cancers. The invention also provides for prognostic (or predictive) assays for determining whether an individual is at risk of developing a disorder associated with NOVX protein, nucleic acid expression or activity. For example, mutations in an NOVX gene can be assayed in a biological sample. Such assays can be used for prognostic or predictive puφose to thereby
prophylactically treat an individual prior to the onset of a disorder characterized by or associated with NOVX protein, nucleic acid expression, or biological activity.
Another aspect of the invention provides methods for determining NOVX protein, nucleic acid expression or activity in an individual to thereby select appropriate therapeutic or prophylactic agents for that individual (refened to herein as "pharmacogenomics"). Pharmacogenomics allows for the selection of agents (e.g., drugs) for therapeutic or prophylactic treatment of an individual based on the genotype of the individual (e.g., the genotype of the individual examined to determine the ability of the individual to respond to a particular agent.)
Yet another aspect of the invention pertains to monitoring the influence of agents (e.g., drags, compounds) on the expression or activity of NOVX in clinical trials.
These and other agents are described in further detail in the following sections.
Diagnostic Assays
An exemplary method for detecting the presence or absence of NOVX in a biological sample involves obtaining a biological sample from a test subject and contacting the biological sample with a compound or an agent capable of detecting NOVX protein or nucleic acid (e.g., mRNA, genomic DNA) that encodes NOVX protein such that the presence of NOVX is detected in the biological sample. An agent for detecting NOVX mRNA or genomic DNA is a labeled nucleic acid probe capable of hybridizing to NOVX mRNA or genomic DNA. The nucleic acid probe can be, for example, a full-length NOVX nucleic acid, such as the nucleic acid of SEQ ED NO: 2n-l, wherein n is an integer between 1 and 178, or a portion thereof, such as an oligonucleotide of at least 15, 30, 50, 100, 250 or 500 nucleotides in length and sufficient to specifically hybridize under stringent conditions to NOVX mRNA or genomic DNA. Other suitable probes for use in the diagnostic assays of the invention are described herein.
An agent for detecting NOVX protein is an antibody capable of binding to NOVX protein, preferably an antibody with a detectable label. Antibodies can be polyclonal, or more preferably, monoclonal. An intact antibody, or a fragment thereof (e.g., Fab or F(ab')2) can be used. The term "labeled", with regard to the probe or antibody, is intended to encompass direct labeling of the probe or antibody by coupling (i.e., physically linking) a detectable substance to the probe or antibody, as well as indirect labeling of the probe or antibody by reactivity with another reagent that is directly labeled. Examples of indirect labeling include detection of a primary antibody using a fluorescently-labeled secondary antibody and
end-labeling of a DNA probe with biotin such that it can be detected with fluorescently- labeled streptavidin. The term "biological sample" is intended to include tissues, cells and biological fluids isolated from a subject, as well as tissues, cells and fluids present within a subject. That is, the detection method of the invention can be used to detect NOVX mRNA, protein, or genomic DNA in a biological sample in vitro as well as in vivo. For example, in vitro techniques for detection of NOVX mRNA include Northern hybridizations and in situ hybridizations. In vitro techniques for detection of NOVX protein include enzyme linked immunosorbent assays (ELIS As), Western blots, immunoprecipitations, and immunofluorescence. In vitro techniques for detection of NOVX genomic DNA include Southem hybridizations. Furthermore, in vivo techniques for detection of NOVX protein include introducing into a subject a labeled anti-NOVX antibody. For example, the antibody can be labeled with a radioactive marker whose presence and location in a subject can be detected by standard imaging techniques.
In one embodiment, the biological sample contains protein molecules from the test subject. Alternatively, the biological sample can contain mRNA molecules from the test subject or genomic DNA molecules from the test subject. A prefened biological sample is a peripheral blood leukocyte sample isolated by conventional means from a subject.
In another embodiment, the methods further involve obtaining a control biological sample from a control subject, contacting the control sample with a compound or agent capable of detecting NOVX protein, mRNA, or genomic DNA, such that the presence of NOVX protein, mRNA or genomic DNA is detected in the biological sample, and comparing the presence of NOVX protein, mRNA or genomic DNA in the control sample with the presence of NOVX protein, mRNA or genomic DNA in the test sample.
The invention also encompasses kits for detecting the presence of NOVX in a biological sample. For example, the kit can comprise: a labeled compound or agent capable of detecting NOVX protein or mRNA in a biological sample; means for determining the amount of NOVX in the sample; and means for comparing the amount of NOVX in the sample with a standard. The compound or agent can be packaged in a suitable container. The kit can further comprise instructions for using the kit to detect NOVX protein or nucleic acid.
Prognostic Assays
The diagnostic methods described herein can furthermore be utilized to identify subjects having or at risk of developing a disease or disorder associated with abenant NOVX expression or activity. For example, the assays described herein, such as the preceding
diagnostic assays or the following assays, can be utilized to identify a subject having or at risk of developing a disorder associated with NOVX protein, nucleic acid expression or activity. Alternatively, the prognostic assays can be utilized to identify a subject having or at risk for developing a disease or disorder. Thus, the invention provides a method for identifying a disease or disorder associated with abenant NOVX expression or activity in which a test sample is obtained from a subject and NOVX protein or nucleic acid (e.g., mRNA, genomic DNA) is detected, wherein the presence of NOVX protein or nucleic acid is diagnostic for a subject having or at risk of developing a disease or disorder associated with abenant NOVX expression or activity. As used herein, a "test sample" refers to a biological sample obtained from a subject of interest. For example, a test sample can be a biological fluid (e.g., serum), cell sample, or tissue.
Furthermore, the prognostic assays described herein can be used to determine whether a subject can be administered an agent (e.g., an agonist, antagonist, peptidomimetic, protein, peptide, nucleic acid, small molecule, or other drag candidate) to treat a disease or disorder associated with abenant NOVX expression or activity. For example, such methods can be used to determine whether a subject can be effectively treated with an agent for a disorder. Thus, the invention provides methods for determining whether a subject can be effectively treated with an agent for a disorder associated with abenant NOVX expression or activity in which a test sample is obtained and NOVX protein or nucleic acid is detected (e.g., wherein the presence of NOVX protein or nucleic acid is diagnostic for a subject that can be administered the agent to treat a disorder associated with abenant NOVX expression or activity).
The methods of the invention can also be used to detect genetic lesions in an NOVX gene, thereby determining if a subject with the lesioned gene is at risk for a disorder characterized by abenant cell proliferation and/or differentiation. In various embodiments, the methods include detecting, in a sample of cells from the subject, the presence or absence of a genetic lesion characterized by at least one of an alteration affecting the integrity of a gene encoding an NOVX-protein, or the misexpression of the NOVX gene. For example, such genetic lesions can be detected by ascertaining the existence of at least one of: (i) a deletion of one or more nucleotides from an NOVX gene; (ii) an addition of one or more nucleotides to an NOVX gene; (iii) a substitution of one or more nucleotides of an NOVX gene, (iv) a chromosomal reanangement of an NOVX gene; (v) an alteration in the level of a messenger RNA transcript of an NOVX gene, (vi) abenant modification of an NOVX gene, such as of the methylation pattern of the genomic DNA, (VH) the presence of a non- wild-type splicing pattern
of a messenger RNA transcript of an NOVX gene, (viii) a non- wild-type level of an NOVX protein, (ix) allelic loss of an NOVX gene, and (x) inappropriate post-translational modification of an NOVX protein. As described herein, there are a large number of assay techniques known in the art which can be used for detecting lesions in an NOVX gene. A prefened biological sample is a peripheral blood leukocyte sample isolated by conventional means from a subject. However, any biological sample containing nucleated cells may be used, including, for example, buccal mucosal cells.
In certain embodiments, detection of the lesion involves the use of a probe/primer in a polymerase chain reaction (PCR) (see, e.g., U.S. Patent Nos. 4,683,195 and 4,683,202), such as anchor PCR or RACE PCR, or, alternatively, in a ligation chain reaction (LCR) (see, e.g., Landegran, et αl., 1988. Science 241: 1077-1080; and Nakazawa, et αl., 1994. Proc. Nαtl Acαd. Sci. USA 91 : 360-364), the latter of which can be particularly useful for detecting point mutations in the NOVX-gene (see, Abravaya, et αl., 1995. Nucl. Acids Res. 23: 675-682). This method can include the steps of collecting a sample of cells from a patient, isolating nucleic acid (e.g., genomic, mRNA or both) from the cells of the sample, contacting the nucleic acid sample with one or more primers that specifically hybridize to an NOVX gene under conditions such that hybridization and amplification of the NOVX gene (if present) occurs, and detecting the presence or absence of an amplification product, or detecting the size of the amplification product and comparing the length to a control sample. It is anticipated that PCR and or LCR may be desirable to use as a preliminary amplification step in conjunction with any of the techniques used for detecting mutations described herein.
Alternative amplification methods include: self sustained sequence replication (see, Guatelli, et αl, 1990. Proc. Nαtl. Acαd. Sci. USA 87: 1874-1878), transcriptional amplification system (see, Kwoh, et αl., 1989. Proc. Nαtl. Acαd. Sci. USA 86: 1173-1177); Qβ Replicase (see, Lizardi, et αl, 1988. BioTechnology 6: 1197), or any other nucleic acid amplification method, followed by the detection of the amplified molecules using techniques well known to those of skill in the art. These detection schemes are especially useful for the detection of nucleic acid molecules if such molecules are present in very low numbers.
In an alternative embodiment, mutations in an NOVX gene from a sample cell can be identified by alterations in restriction enzyme cleavage patterns. For example, sample and control DNA is isolated, amplified (optionally), digested with one or more restriction endonucleases, and fragment length sizes are determined by gel electrophoresis and compared. Differences in fragment length sizes between sample and control DNA indicates mutations in the sample DNA. Moreover, the use of sequence specific ribozymes (see, e.g., U.S. Patent
No. 5,493,531) can be used to score for the presence of specific mutations by development or loss of a ribozyme cleavage site.
In other embodiments, genetic mutations in NOVX can be identified by hybridizing a sample and control nucleic acids, e.g., DNA or RNA, to high-density anays containing hundreds or thousands of oligonucleotides probes. See, e.g., Cronin, et al, 1996. Human Mutation 1: 244-255; Kozal, et al., 1996. Nat. Med. 2: 753-759. For example, genetic mutations in NOVX can be identified in two dimensional anays containing light-generated DNA probes as described in Cronin, et al, supra. Briefly, a first hybridization anay of probes can be used to scan through long stretches of DNA in a sample and control to identify base changes between the sequences by making linear anays of sequential overlapping probes. This step allows the identification of point mutations. This is followed by a second hybridization anay that allows the characterization of specific mutations by using smaller, specialized probe anays complementary to all variants or mutations detected. Each mutation anay is composed of parallel probe sets, one complementary to the wild-type gene and the other complementary to the mutant gene.
In yet another embodiment, any of a variety of sequencing reactions known in the art can be used to directly sequence the NOVX gene and detect mutations by comparing the sequence of the sample NOVX with the conesponding wild-type (confrol) sequence. Examples of sequencing reactions include those based on techniques developed by Maxim and Gilbert, 1977. Proc. Natl. Acad. Sci. USA 74: 560 or Sanger, 1977. Proc. Natl. Acad. Sci. USA 74: 5463. It is also contemplated that any of a variety of automated sequencing procedures can be utilized when performing the diagnostic assays (see, e.g., Naeve, et al, 1995. Biotechniques 19: 448), including sequencing by mass spectrometry (see, e.g., PCT International Publication No. WO 94/16101; Cohen, et al, 1996. Adv. Chromatography 36: 127-162; and Griffin, et al, 1993. Appl Biochem. Biotechnol. 38: 147-159).
Other methods for detecting mutations in the NOVX gene include methods in which protection from cleavage agents is used to detect mismatched bases in RNA/RNA or RNA/DNA heteroduplexes. See, e.g., Myers, et al, 1985. Science 230: 1242. In general, the art technique of "mismatch cleavage" starts by providing heteroduplexes of formed by hybridizing (labeled) RNA or DNA containing the wild-type NOVX sequence with potentially mutant RNA or DNA obtained from a tissue sample. The double-stranded duplexes are treated with an agent that cleaves single-stranded regions of the duplex such as which will exist due to basepair mismatches between the confrol and sample strands. For instance, RNA/DNA duplexes can be treated with RNase and DNA/DNA hybrids treated with Si
nuclease to enzymatically digesting the mismatched regions. In other embodiments, either DNA/DNA or RNA/DNA duplexes can be treated with hydroxylamine or osmium tetroxide and with piperidine in order to digest mismatched regions. After digestion of the mismatched regions, the resulting material is then separated by size on denaturing polyacrylamide gels to determine the site of mutation. See, e.g., Cotton, et al, 1988. Proc. Natl. Acad. Sci. USA 85: 4397; Saleeba, et al, 1992. Methods Enzymol. 217: 286-295. In an embodiment, the control DNA or RNA can be labeled for detection.
In still another embodiment, the mismatch cleavage reaction employs one or more proteins that recognize mismatched base pairs in double-stranded DNA (so called "DNA mismatch repair" enzymes) in defined systems for detecting and mapping point mutations in NOVX cDNAs obtained from samples of cells. For example, the mutY enzyme of E. coli cleaves A at G/A mismatches and the fhymidine DNA glycosylase from HeLa cells cleaves T at G/T mismatches. See, e.g., Hsu, et al, 1994. Carcinogenesis 15: 1657-1662. According to an exemplary embodiment, a probe based on an NOVX sequence, e.g. , a wild-type NOVX sequence, is hybridized to a cDNA or other DNA product from a test cell(s). The duplex is treated with a DNA mismatch repair enzyme, and the cleavage products, if any, can be detected from electrophoresis protocols or the like. See, e.g., U.S. Patent No. 5,459,039.
In other embodiments, alterations in electrophoretic mobility will be used to identify mutations in NOVX genes. For example, single strand conformation polymoφhism (SSCP) may be used to detect differences in electrophoretic mobility between mutant and wild type nucleic acids. See, e.g., Orita, et al, 1989. Proc. Natl. Acad. Sci. USA: 86: 2766; Cotton, 1993. Mutat. Res. 285: 125-144; Hayashi, 1992. Genet. Anal. Tech. Appl. 9: 73-79. Single-stranded DNA fragments of sample and control NOVX nucleic acids will be denatured and allowed to renature. The secondary structure of single-stranded nucleic acids varies according to sequence, the resulting alteration in electrophoretic mobility enables the detection of even a single base change. The DNA fragments may be labeled or detected with labeled probes. The sensitivity of the assay may be enhanced by using RNA (rather than DNA), in which the secondary structure is more sensitive to a change in sequence. In one embodiment, the subject method utilizes heteroduplex analysis to separate double stranded heteroduplex molecules on the basis of changes in electrophoretic mobility. See, e.g., Keen, et al, 1991. Trends Genet. 7: 5.
In yet another embodiment, the movement of mutant or wild-type fragments in polyacrylamide gels containing a gradient of denaturant is assayed using denaturing gradient gel electrophoresis (DGGΕ). See, e.g., Myers, et al, 1985. Nature 313: 495. When DGGΕ is
used as the method of analysis, DNA will be modified to insure that it does not completely denature, for example by adding a GC clamp of approximately 40 bp of high-melting GC-rich DNA by PCR. In a further embodiment, a temperature gradient is used in place of a denaturing gradient to identify differences in the mobility of confrol and sample DNA. See, e.g., Rosenbaum and Reissner, 19S1. Biophys. Chem. 265: 12753.
Examples of other techniques for detecting point mutations include, but are not limited to, selective oligonucleotide hybridization, selective amplification, or selective primer extension. For example, oligonucleotide primers may be prepared in which the known mutation is placed centrally and then hybridized to target DNA under conditions that permit hybridization only if a perfect match is found. See, e.g., Saiki, et al, 1986. Nature 324: 163; Saiki, et al, 1989. Proc. Natl. Acad. Sci. USA 86: 6230. Such allele specific oligonucleotides are hybridized to PCR amplified target DNA or a number of different mutations when the oligonucleotides are attached to the hybridizing membrane and hybridized with labeled target DNA.
Alternatively, allele specific amplification technology that depends on selective PCR amplification may be used in conjunction with the instant invention. Oligonucleotides used as primers for specific amplification may cany the mutation of interest in the center of the molecule (so that amplification depends on differential hybridization; see, e.g., Gibbs, et al, 1989. Nucl. Acids Res. 17: 2437-2448) or at the extreme 3'-terminus of one primer where, under appropriate conditions, mismatch can prevent, or reduce polymerase extension (see, e.g., Prossner, 1993. Tibtech. 11 : 238). In addition it may be desirable to introduce a novel restriction site in the region of the mutation to create cleavage-based detection. See, e.g., Gasparini, et al, 1992. Mol Cell Probes 6: 1. It is anticipated that in certain embodiments amplification may also be performed using Taq ligase for amplification. See, e.g., Barany, 1991. Proc. Natl. Acad. Sci. USA 88: 189. In such cases, ligation will occur only if there is a perfect match at the 3 '-terminus of the 5' sequence, making it possible to detect the presence of a known mutation at a specific site by looking for the presence or absence of amplification.
The methods described herein may be performed, for example, by utilizing pre-packaged diagnostic kits comprising at least one probe nucleic acid or antibody reagent described herein, which may be conveniently used, e.g. , in clinical settings to diagnose patients exhibiting symptoms or family history of a disease or illness involving an NOVX gene.
Furthermore, any cell type or tissue, preferably peripheral blood leukocytes, in which NOVX is expressed may be utilized in the prognostic assays described herein. However, any
biological sample containing nucleated cells may be used, including, for example, buccal mucosal cells.
Pharmacogenomics
Agents, or modulators that have a stimulatory or inhibitory effect on NOVX activity (e.g., NOVX gene expression), as identified by a screening assay described herein can be administered to individuals to treat (prophylactically or therapeutically) disorders (The disorders include metabolic disorders, diabetes, obesity, infectious disease, anorexia, cancer- associated cachexia, cancer, neurodegenerative disorders, Alzheimer's Disease, Parkinson's Disorder, immune disorders, and hematopoietic disorders, and the various dyslipidemias, metabolic disturbances associated with obesity, the metabolic syndrome X and wasting disorders associated with chronic diseases and various cancers.) In conjunction with such freatment, the pharmacogenomics (i.e., the study of the relationship between an individual's genotype and that individual's response to a foreign compound or drug) of the individual may be considered. Differences in metabolism of therapeutics can lead to severe toxicity or therapeutic failure by altering the relation between dose and blood concentration of the pharmacologically active drag. Thus, the pharmacogenomics of the individual permits the selection of effective agents (e.g., drugs) for prophylactic or therapeutic freatments based on a consideration of the individual's genotype. Such pharmacogenomics can further be used to determine appropriate dosages and therapeutic regimens. Accordingly, the activity of NOVX protein, expression of NOVX nucleic acid, or mutation content of NOVX genes in an individual can be determined to thereby select appropriate agent(s) for therapeutic or prophylactic treatment of the individual.
Pharmacogenomics deals with clinically significant hereditary variations in the response to drags due to altered drag disposition and abnormal action in affected persons. See e.g., Eichelbaum, 1996. Clin. Exp. Pharmacol. Physiol, 23: 983-985; Linder, 1997. Clin.
Chem., 43: 254-266. In general, two types of pharmacogenetic conditions can be differentiated. Genetic conditions transmitted as a single factor altering the way drugs act on the body (altered drag action) or genetic conditions transmitted as single factors altering the way the body acts on drugs (altered drag metabolism). These pharmacogenetic conditions can occur either as rare defects or as polymoφhisms. For example, glucose-6-phosphate dehydrogenase (G6PD) deficiency is a common inherited enzymopathy in which the main clinical complication is hemolysis after ingestion of oxidant drags (anti-malarials, sulfonamides, analgesics, nitrofurans) and consumption of fava beans.
As an illustrative embodiment, the activity of drug metabolizing enzymes is a major determinant of both the intensity and duration of drag action. The discovery of genetic polymoφhisms of drag metabolizing enzymes (e.g., N-acetyltransferase 2 (NAT 2) and cytochrome PREGNANCY ZONE PROTEIN PRECURSOR enzymes CYP2D6 and CYP2C19) has provided an explanation as to why some patients do not obtain the expected drug effects or show exaggerated drug response and serious toxicity after taking the standard and safe dose of a drug. These polymoφhisms are expressed in two phenotypes in the population, the extensive metabolizer (EM) and poor metabolizer (PM). The prevalence of PM is different among different populations. For example, the gene coding for CYP2D6 is highly polymoφhic and several mutations have been identified in PM, which all lead to the absence of functional CYP2D6. Poor metabolizers of CYP2D6 and CYP2C19 quite frequently experience exaggerated drug response and side effects when they receive standard doses. If a metabolite is the active therapeutic moiety, PM show no therapeutic response, as demonstrated for the analgesic effect of codeine mediated by its CYP2D6-formed metabolite moφhine. At the other extreme are the so called ultra-rapid metabolizers who do not respond to standard doses. Recently, the molecular basis of ultra-rapid metabolism has been identified to be due to CYP2D6 gene amplification.
Thus, the activity of NOVX protein, expression of NOVX nucleic acid, or mutation content of NOVX genes in an individual can be determined to thereby select appropriate agent(s) for therapeutic or prophylactic treatment of the individual. In addition, pharmacogenetic studies can be used to apply genotyping of polymoφhic alleles encoding drug-metabolizing enzymes to the identification of an individual's drag responsiveness phenotype. This knowledge, when applied to dosing or drug selection, can avoid adverse reactions or therapeutic failure and thus enhance therapeutic or prophylactic efficiency when treating a subject with an NOVX modulator, such as a modulator identified by one of the exemplary screening assays described herein.
Monitoring of Effects During Clinical Trials
Monitoring the influence of agents (e.g., drags, compounds) on the expression or activity of NOVX (e.g., the ability to modulate abenant cell proliferation and/or differentiation) can be applied not only in basic drug screening, but also in clinical trials. For example, the effectiveness of an agent determined by a screening assay as described herein to increase NOVX gene expression, protein levels, or upregulate NOVX activity, can be monitored in clinical trails of subjects exhibiting decreased NOVX gene expression, protein
levels, or downregulated NOVX activity. Alternatively, the effectiveness of an agent determined by a screening assay to decrease NOVX gene expression, protein levels, or downregulate NOVX activity, can be monitored in clinical trails of subjects exhibiting increased NOVX gene expression, protein levels, or upregulated NOVX activity. In such clinical trials, the expression or activity of NOVX and, preferably, other genes that have been implicated in, for example, a cellular proliferation or immune disorder can be used as a "read out" or markers of the immune responsiveness of a particular cell.
By way of example, and not of limitation, genes, including NOVX, that are modulated in cells by treatment with an agent (e.g. , compound, drug or small molecule) that modulates NOVX activity (e.g. , identified in a screening assay as described herein) can be identified. Thus, to study the effect of agents on cellular proliferation disorders, for example, in a clinical trial, cells can be isolated and RNA prepared and analyzed for the levels of expression of NOVX and other genes implicated in the disorder. The levels of gene expression (i.e., a gene expression pattern) can be quantified by Northern blot analysis or RT-PCR, as described herein, or alternatively by measuring the amount of protein produced, by one of the methods as described herein, or by measuring the levels of activity of NOVX or other genes. In this manner, the gene expression pattern can serve as a marker, indicative of the physiological response of the cells to the agent. Accordingly, this response state may be determined before, and at various points during, treatment of the individual with the agent.
In one embodiment, the invention provides a method for monitoring the effectiveness of treatment of a subject with an agent (e.g., an agonist, antagonist, protein, peptide, peptidomimetic, nucleic acid, small molecule, or other drug candidate identified by the screening assays described herein) comprising the steps of (i) obtaining a pre-administration sample from a subject prior to administration of the agent; (ii) detecting the level of expression of an NOVX protein, mRNA, or genomic DNA in the preadministration sample; (iii) obtaining one or more post-administration samples from the subject; (iv) detecting the level of expression or activity of the NOVX protein, mRNA, or genomic DNA in the post-administration samples; (v) comparing the level of expression or activity of the NOVX protein, mRNA, or genomic DNA in the pre-administration sample with the NOVX protein, mRNA, or genomic DNA in the post administration sample or samples; and (vi) altering the administration of the agent to the subject accordingly. For example, increased administration of the agent may be desirable to increase the expression or activity of NOVX to higher levels than detected, i.e., to increase the effectiveness of the agent. Alternatively, decreased
administration of the agent may be desirable to decrease expression or activity of NOVX to lower levels than detected, i.e., to decrease the effectiveness of the agent.
Methods of Treatment
The invention provides for both prophylactic and therapeutic methods of treating a subject at risk of (or susceptible to) a disorder or having a disorder associated with abenant NOVX expression or activity. The disorders include cardiomyopathy, atherosclerosis, hypertension, congenital heart defects, aortic stenosis, atrial septal defect (ASD), atrioventricular (A-V) canal defect, ductus arteriosus, pulmonary stenosis, subaortic stenosis, ventricular septal defect (VSD), valve diseases, tuberous sclerosis, scleroderma, obesity, transplantation, adrenoleukodystrophy, congenital adrenal hypeφlasia, prostate cancer, neoplasm; adenocarcinoma, lymphoma, uteras cancer, fertility, hemophilia, hypercoagulation, idiopathic thrombocytopenic puφura, immunodeficiencies, graft versus host disease, AIDS, bronchial asthma, Crohn's disease; multiple sclerosis, treatment of Albright Hereditary Ostoeodystrophy, and other diseases, disorders and conditions of the like.
These methods of freatment will be discussed more fully, below.
Disease and Disorders
Diseases and disorders that are characterized by increased (relative to a subject not suffering from the disease or disorder) levels or biological activity may be treated with Therapeutics that antagonize (i.e., reduce or inhibit) activity. Therapeutics that antagonize activity may be administered in a therapeutic or prophylactic manner. Therapeutics that may be utilized include, but are not limited to: ( ) an aforementioned peptide, or analogs, derivatives, fragments or homologs thereof; (ii) antibodies to an aforementioned peptide; (iii) nucleic acids encoding an aforementioned peptide; (iv) administration of antisense nucleic acid and nucleic acids that are "dysfunctional" (i.e., due to a heterologous insertion within the coding sequences of coding sequences to an aforementioned peptide) that are utilized to "knockout" endogenous function of an aforementioned peptide by homologous recombination (see, e.g., Capecchi, 1989. Science 244: 1288-1292); or (v) modulators ( i.e., inhibitors, agonists and antagonists, including additional peptide mimetic of the invention or antibodies specific to a peptide of the invention) that alter the interaction between an aforementioned peptide and its binding partner.
Diseases and disorders that are characterized by decreased (relative to a subject not suffering from the disease or disorder) levels or biological activity may be treated with Therapeutics that increase (i.e., are agonists to) activity. Therapeutics that upregulate activity may be administered in a therapeutic or prophylactic manner. Therapeutics that may be utilized include, but are not limited to, an aforementioned peptide, or analogs, derivatives, fragments or homologs thereof; or an agonist that increases bioavailability.
Increased or decreased levels can be readily detected by quantifying peptide and/or RNA, by obtaining a patient tissue sample (e.g., from biopsy tissue) and assaying it in vitro for RNA or peptide levels, structure and/or activity of the expressed peptides (or mRNAs of an aforementioned peptide). Methods that are well-known within the art include, but are not limited to, immunoassays (e.g., by Western blot analysis, immunoprecipitation followed by sodium dodecyl sulfate (SDS) polyacrylamide gel electrophoresis, immunocytochemistry, etc.) and/or hybridization assays to detect expression of mRNAs (e.g., Northern assays, dot blots, in situ hybridization, and the like).
Prophylactic Methods
In one aspect, the invention provides a method for preventing, in a subject, a disease or condition associated with an abenant NOVX expression or activity, by administering to the subject an agent that modulates NOVX expression or at least one NOVX activity. Subjects at risk for a disease that is caused or contributed to by abenant NOVX expression or activity can be identified by, for example, any or a combination of diagnostic or prognostic assays as described herein. Administration of a prophylactic agent can occur prior to the manifestation of symptoms characteristic of the NOVX abenancy, such that a disease or disorder is prevented or, alternatively, delayed in its progression. Depending upon the type of NOVX abenancy, for example, an NOVX agonist or NOVX antagonist agent can be used for treating the subject. The appropriate agent can be determined based on screening assays described herein. The prophylactic methods of the invention are further discussed in the following subsections. Therapeutic Methods
Another aspect of the invention pertains to methods of modulating NOVX expression or activity for therapeutic puφoses. The modulatory method of the invention involves contacting a cell with an agent that modulates one or more of the activities of NOVX protein activity associated with the cell. An agent that modulates NOVX protein activity can be an agent as described herein, such as a nucleic acid or a protein, a naturally-occurring cognate
ligand of an NOVX protein, a peptide, an NOVX peptidomimetic, or other small molecule. In one embodiment, the agent stimulates one or more NOVX protein activity. Examples of such stimulatory agents include active NOVX protein and a nucleic acid molecule encoding NOVX that has been introduced into the cell. In another embodiment, the agent inhibits one or more NOVX protein activity. Examples of such inhibitory agents include antisense NOVX nucleic acid molecules and anti-NOVX antibodies. These modulatory methods can be performed in vitro (e.g., by culturing the cell with the agent) or, alternatively, in vivo (e.g., by administering the agent to a subject). As such, the invention provides methods of treating an individual afflicted with a disease or disorder characterized by abenant expression or activity of an NOVX protein or nucleic acid molecule. In one embodiment, the method involves administering an agent (e.g., an agent identified by a screening assay described herein), or combination of agents that modulates (e.g., up-regulates or down-regulates) NOVX expression or activity. In another embodiment, the method involves administering an NOVX protein or nucleic acid molecule as therapy to compensate for reduced or abenant NOVX expression or activity.
Stimulation of NOVX activity is desirable in situations in which NOVX is abnormally downregulated and/or in which increased NOVX activity is likely to have a beneficial effect. One example of such a situation is where a subject has a disorder characterized by abenant cell proliferation and/or differentiation (e.g., cancer or immune associated disorders). Another example of such a situation is where the subject has a gestational disease (e.g., preclampsia).
Determination of the Biological Effect of the Therapeutic
In various embodiments of the invention, suitable in vitro or in vivo assays are performed to determine the effect of a specific Therapeutic and whether its administration is indicated for freatment of the affected tissue.
In various specific embodiments, in vitro assays may be performed with representative cells of the type(s) involved in the patient's disorder, to determine if a given Therapeutic exerts the desired effect upon the cell type(s). Compounds for use in therapy may be tested in suitable animal model systems including, but not limited to rats, mice, chicken, cows, monkeys, rabbits, and the like, prior to testing in human subjects. Similarly, for in vivo testing, any of the animal model system known in the art may be used prior to administration to human subjects.
Prophylactic and Therapeutic Uses of the Compositions of the Invention
The NOVX nucleic acids and proteins of the invention are useful in potential prophylactic and therapeutic applications implicated in a variety of disorders including, but not limited to: metabolic disorders, diabetes, obesity, infectious disease, anorexia, cancer- associated cancer, neurodegenerative disorders, Alzheimer's Disease, Parkinson's Disorder, immune disorders, hematopoietic disorders, and the various dyslipidemias, metabolic disturbances associated with obesity, the metabolic syndrome X and wasting disorders associated with chronic diseases and various cancers.
As an example, a cDNA encoding the NOVX protein of the invention may be useful in gene therapy, and the protein may be useful when administered to a subject in need thereof. By way of non-limiting example, the compositions of the invention will have efficacy for treatment of patients suffering from: metabolic disorders, diabetes, obesity, infectious disease, anorexia, cancer-associated cachexia, cancer, neurodegenerative disorders, Alzheimer's Disease, Parkinson's Disorder, immune disorders, hematopoietic disorders, and the various dyslipidemias.
Both the novel nucleic acid encoding the NOVX protein, and the NOVX protein of the invention, or fragments thereof, may also be useful in diagnostic applications, wherein the presence or amount of the nucleic acid or the protein are to be assessed. A further use could be as an anti-bacterial molecule (i.e., some peptides have been found to possess anti-bacterial properties). These materials are further useful in the generation of antibodies, which immunospecifically-bind to the novel substances of the invention for use in therapeutic or diagnostic methods.
Sequence Analyses
The sequence of NOVX was derived by laboratory cloning of cDNA fragments, by in silico prediction of the sequence. cDNA fragments covering either the full length of the DNA sequence, or part of the sequence, or both, were cloned. In silico prediction was based on sequences available in CuraGen's proprietary sequence databases or in the public human sequence databases, and provided either the full length DNA sequence, or some portion thereof.
The laboratory cloning was performed using one or more of the methods summarized below:
SeqCaIling™Technology: cDNA was derived from various human samples representing multiple tissue types, normal and diseased states, physiological states, and developmental states from different donors. Samples were obtained as whole tissue, primary cells or tissue cultured primary cells or cell lines. Cells and cell lines may have been treated with biological or chemical agents that regulate gene expression, for example, growth factors, chemokines or steroids. The cDNA thus derived was then sequenced using CuraGen Coφoration's SeqCalling technology which is disclosed in full in U. S. Ser. Nos. 09/417,386 filed Oct. 13, 1999, and 09/614,505 filed July 11, 2000. Sequence traces were evaluated manually and edited for conections if appropriate. cDNA sequences from all samples were assembled together, sometimes including public human sequences, using bioinformatics programs to produce a consensus sequence for each assembly. Each assembly is included in CuraGen Coφoration's database. Sequences were included as components for assembly when the extent of identity with another component was at least 95% over 50 bp. Each assembly represents a gene or portion thereof and includes information on variants, such as splice forms single nucleotide polymoφhisms (SNPs), insertions, deletions and other sequence variations.
Variant sequences are also included in this application. A variant sequence can include a single nucleotide polymoφhism (SNP). A SNP can, in some instances, be refened to as a "cSNP" to denote that the nucleotide sequence containing the SNP originates as a cDNA. A SNP can arise in several ways. For example, a SNP may be due to a substitution of one nucleotide for another at the polymoφhic site. Such a substitution can be either a transition or a transversion. A SNP can also arise from a deletion of a nucleotide or an insertion of a nucleotide, relative to a reference allele. In this case, the polymoφhic site is a site at which one allele bears a gap with respect to a particular nucleotide in another allele. SNPs occurring within genes may result in an alteration of the amino acid encoded by the gene at the position of the SNP. Infragenic SNPs may also be silent, when a codon including a SNP encodes the same amino acid as a result of the redundancy of the genetic code. SNPs occurring outside the region of a gene, or in an intron within a gene, do not result in changes in any amino acid sequence of a protein but may result in altered regulation of the expression pattern. Examples include alteration in temporal expression, physiological response regulation, cell type expression regulation, intensity of expression, and stability of transcribed message.
Presented information includes that associated with genomic clones, public genes and ESTs sharing sequence identity with the disclosed sequence and CuraGen Coφoration's Electronic Northern bioinformatic tool.
Examples Example A: Sequence related information
The NO VI clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 1 A.
Further analysis of the NOVla protein yielded the following properties shown in Table
IB.
Table IB. Protein Sequence Properties NOVla
Psort 0.6500 probability located in cytoplasm; 0.2340 probability located analysis: in lysosome (lumen); 0.1000 probability located in mitochondrial matrix space; 0.0000 probability located in endoplasmic reticulum (membrane)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVla protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table lC.
In a BLAST search of public sequence databases, the NOVla protein was found to have homology to the proteins shown in the BLASTP data in Table ID.
PFam analysis predicts that the NOVla protein contains the domains shown in the Table IE.
Table IE. Domain Analysis of NOVla
Example 2.
The NOV2 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 2A.
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 2B.
Further analysis of the NOV2a protein yielded the following properties shown in Table
2C.
Table 2C. Protein Sequence Properties NOV2a
Psort 0.6400 probability located in plasma membrane; 0.4600 probability located analysis: in Golgi body; 0.3700 probability located in endoplasmic reticulum (membrane); 0.1000 probability located in endoplasmic reticulum (lumen)
SignalP Likely cleavage site between residues 38 and 39 analysis:
A search of the NOV2a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 2D.
In a BLAST search of public sequence databases, the NOV2a protein was found to have homology to the proteins shown in the BLASTP data in Table 2E.
PFam analysis predicts that the NOV2a protein contains the domains shown in the Table 2F.
Example 3.
The NOV3 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 3 A.
Further analysis of the NOV3a protein yielded the following properties shown in Table
3B.
Table 3B. Protein Sequence Properties NOV3a
PSort 0.6850 probability located in endoplasmic reticulum (membrane); analysis: 0.6400 probability located in plasma membrane; 0.4600 probability located in Golgi body; 0.2400 probability located in nucleus
SignalP Likely cleavage site between residues 25 and 26 analysis:
A search of the NOV3a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 3C.
In a BLAST search of public sequence databases, the NOV3a protein was found to have homology to the proteins shown in the BLASTP data in Table 3D.
PFam analysis predicts that the NOV3a protein contains the domains shown in the Table 3E.
Table 3E. Domain Analysis of NOV3a
Example 4.
The NOV4 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 4A.
Further analysis of the NOV4a protein yielded the following properties shown in Table
4B.
Table 4B. Protein Sequence Properties NOV4a
Psort 0.9600 probability located in nucleus; 0.4776 probability located in analysis: mitochondrial matrix space; 0.3000 probability located in microbody (peroxisome); 0.1837 probability located in mitochondrial inner membrane
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV4a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 4C.
In a BLAST search of public sequence databases, the NOV4a protein was found to have homology to the proteins shown in the BLASTP data in Table 4D.
PFam analysis predicts that the NOV4a protein contains the domains shown in the Table 4E.
Example 5.
The NOV5 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 5A.
Table 5A. NOV5 Sequence Analysis —
Further analysis of the NOV5a protein yielded the following properties shown in Table
5B.
Table 5B. Protein Sequence Properties NOV5a
Psort 0.4500 probability located in cytoplasm; 0.3000 probability located in analysis: microbody (peroxisome); 0.1897 probability located in lysosome (lumen); 0.1000 probability located in mitochondrial matrix space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV5a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 5C.
In a BLAST search of public sequence databases, the NOV5a protein was found to have homology to the proteins shown in the BLASTP data in Table 5D.
PFam analysis predicts that the NOV5a protein contains the domains shown in the Table 5E.
Table 5E. Domain Analysis of NOV5a
Example 6.
The NOV6 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 6A.
Further analysis of the NOV6a protein yielded the following properties shown in Table
6B.
Table 6B. Protein Sequence Properties NOV6a
PSort 0.4500 probability located in cytoplasm; 0.3490 probability located in analysis:
matrix space; 0.1000 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV6a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 6C.
In a BLAST search of public sequence databases, the NOV6a protein was found to have homology to the proteins shown in the BLASTP data in Table 6D.
Table 6D. Public BLASTP Results for NOV6a
PFam analysis predicts that the NOV6a protein contains the domains shown in the Table 6E.
Example 7.
The NOV7 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 7A.
Further analysis of the NOV7a protein yielded the following properties shown in Table
7B.
Table 7B. Protein Sequence Properties NOV7a
PSort 0.9800 probability located in nucleus; 0.1000 probability located in analysis: mitochondrial matrix space; 0.1000 probability located in lysosome (lumen); 0.0000 probability located in endoplasmic reticulum (membrane)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV7a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 7C.
In a BLAST search of public sequence databases, the NOV7a protein was found to have homology to the proteins shown in the BLASTP data in Table 7D.
PFam analysis predicts that the NOV7a protein contains the domains shown in the Table 7E.
Example 8.
The NOV8 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 8A.
Table 8A. NOV8 Sequence Analysis
SEQ ED NO: 19 2296 bp
NOV8a, CGGCGGCGGCGGCAGTAGAAATGATGGAAGAATTGCATAGCCTGGACCCACGACGGCA
GAAATTATTGGAGGCCAGGTTTACTGGAGTAGGTGTTAGTAAGGGACCACTTAATAGT
CG57871-01 DNA Sequence GAGTCTTCCAACCAGAGCTTGTGCAGCGTCGGATCCTTGAGTGATAAAGAAGTAGAGA CTCCCAAGAAAAAGCAGAATGACCAGCGAAATCGGAAAAGAAAAGCTGAACCATATGA AAGTAGCCAAGGGAAAGGCACTCCTAGGGGACATAAAATTAGTGATTACTTTGAGTTT GCTGGGGGAAGCGGGCCGGGAACCAGCCCTGGCAGAAGTGTTCCACCAGTTGCACGAT CCTCACTGCAACATTCTTTATCCAATCCCTTACCGCGACGAGTAGAACAGCCCCTCTA TGGTTTAGATGGCAGTGCTGCAAAGGAGGCAACGGAGGAGCAGTCTGCTCTGCCAACC CTCATGTCAGTGATGCTAGCAAAACCTCGGCTTGACACAGAGCAGCTGGCGCAAAGGG GAGCTGGCCTCTGCTTCACTTTTGTTTCAGCTCAGCAAAACAGTCCCTCATCTACGGG ATCTGGCAACACAGAGCATTCCTGCAGCTCCCAAAAACAGATCTCCATCCAGCACAGA CAGACCCAGTCCGACCTCACAATAGAAAAAATATCTGCACTAGAAAACAGTAAGAATT CTGACTTAGAGAAGAAGGAGGGAAGAATAGATGATTTATTAAGAGCCATCTGTGATTT GAGACGGCAGATTGATGAACAGCAAAAGATGCTAGAGAAATACAAGGAACGATTAAAT AGATGTGTGACAATGAGCAAGAAACTCCTTATAGAAAAGTCAAAACAAGAGAAGATGG CGTGTAGAGATAAGAGCATGCAAGACCGCTTGAGACTGGGCCACTTTACTACGTCTGA CCACGGAGCCAAATTTACTGAGCAGTGGACAGATGGTTATGCTTTTCAGAATCTTATC AAGCAACAGGAAAGGATAAATTCACAGAGGGAAGAGATAGAAAGACAACGGAAAATGT TAGCAAAGCGGAAACCTCCTGCCATGGGTCAGGCCCCTCCTGCAACCAATGAGCAGAA ACAGTGGAAAAGCAAGACCAATGGAGCTGAAAATGAAACGTTAACGTTAAAAGAATAC CATGAACAAGAAGAAATCTTCAAACTCAGATTAGGTCATCTTAAAAAGGAGGAAGCAG AGATCCAGGCAGAGCTGGAGAGGCTAGAAAGGGTTAGAAAACTACATATCAGGGAAGT AAAAAGGATACATAATGAAGATAATTCACAATTTAAATATCATCCAACGCTAAATGAC AGATATTTGTTGTTACATCTTTTGGGTAGAGGAGGTTTCAGTGAAGTTTACAAGGCAT TTGATCTAACAGAGCAAAGATACGTAGCTGTGAAAATTCACCAGTTAAATAAAAACTG GAGAGATGAGAAAAAGGAGAATTACCACAAGCATGCATGTAGGGAATACCGGATTCAT AAAGAGCTGGACCATCCCAGAATAGTTAAGCTGTATGATTACTTTTCACTGGATACTG ACTCGTTTTGTACAGTATTAGAATACTGTGAGGGAAATGATCTGGACTTCTACCTGAA ACAGCACAAATTAATGTCAGAGAAAGAGGCCCGGTCCATTATCATGCAGATTGTGAAT GCTTTAAAGTACTTAAATGAAATAAAACCTCCCATCATACACTATGACCTCAAACCAG GTAATATTCTTTTAGAAAATGGTACAGCGTGTGGAGAGATAAAAATTACAGATTTTGG TCTTTCGAAGATCATGGATGATGATAGCTACAATTCAGTGGATGGCATGGAGCTAACA TCACAAGGTGCTGGTACTTATTGGTATTTACCACCAGAGTGTTTTGTGGTTGGGAAAG AACCACCAAAGATCTCAAATAAAGTTGATGTGTGGTCGGTGGGTGTGATCTTCTATCA GTGTCTTTATGGAAGGAAGCCTTTTGGCCATAACCAGTCTCAGCAAGACATCCTACAA GAGAATACGATTCTTAAAGCTACTGAAGTGCAGTTCCCGCCAAAGCCGGTAGTAACAC CTGAAGCAAAGGCGTTGATTCGACGATGCTTGGCCTACCGAAAGGAGGACCGCATTGA TGTCCAGCAGCTGGCCTGTGATCCCTACTTGTTGCCTCACATCCGAAAGTCAGTCTCT ACGAGTAGCCCTGCTGGAGCTGCTATTGCATCAACCTCTGGGGCGTCCAATAACAGTT CTTCTAATTGAGACTGACTCCAAGGCCACAAACT
ORF Start: ATG at 24 ORF Stop: TGA at 2271
SEQ ED NO: 20 749 aa MW at 85415.8kD
NOV8a, MEELHSLDPRRQKLLEARFTGVGVSKGPLNSESSNQSLCSVGSLSDKEVETPKKKQND QR RKRKAEPYESSQGKGTPRGHKISDYFEFAGGSGPGTSPGRSVPPVARSSLQHSLS
CG57871-01 Protein Sequence NP PRRVEQPLYGLDGSAAKEATEEQSALPTLMSVMLAKPRLDTEQ AQRGAGLCFTF VSAQQNSPSSTGSGNTEHSCSSQKQISIQHRQTQSDLTIEKISALENSK SDLEKKEG RIDD LRAICDLRRQIDEQQKM EKYKER NRCVTMSKKLLIEKSKQEKMACRDKSMQ DRLRLGHFTTSDHGAKFTEQ TDGYAFQNLIKQQERINSQREEIERQRKMLAKRKPPA MGQAPPATNEQKQWKSKTNGAENET TLKEYHEQEEIFK RLGHLKKEEAEIQAELER LERVRKLHIREVKRIHNEDNSQFKYHPTL DRY LLH GRGGFSEVYKAFDLTEQRY VAVKIHQLNKNWRDEKKENYHKHACREYRIHKELDHPRIVKLYDYFSLDTDSFCTVLE YCEGNDLDFYLKQHK MSEKEARSIIMQIVNA KYL EIKPPIIHYDLKPGNILLENG TACGEIKITDFGLSKIMDDDSYNSVDGME TSQGAGTYWYLPPECFWGKEPPKISNK VDV SVGVIFYQC YGRKPFGHNQSQQDILQENTI KATEVQFPPKPWTPEAKALIR RCLAYRKEDRIDVQQLACDPYLLPHIRKSVSTSSPAGAAIASTSGASNNSSSN
Further analysis of the NOV8a protein yielded the following properties shown in Table
8B.
Table 8B. Protein Sequence Properties NOV8a
PSort 0.9600 probability located in nucleus; 0.1000 probability located in mitochondrial analysis: matrix space; 0.1000 probability located in lysosome (lumen); 0.0000 probability located in endoplasmic reticulum (membrane)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV8a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 8C.
In a BLAST search of public sequence databases, the NOV8a protein was found to have homology to the proteins shown in the BLASTP data in Table 8D.
PFam analysis predicts that the NOV8a protein contains the domains shown in the Table 8E.
The NOV9 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 9A.
CACAGTACACATGAACAAGGCCAGTCCTCCATTTCCTCTTATCTCCAACGCACAAGAT CTTGCTCAAGAGGTACAAACTGTTTTGAAGCCAGTTCATCATAAGGAAGGACAAGAAC TAACTGCTTTGCTGAATACTCCACATATTCAGGCACTTTTACTGGCCCACGATAAGGT TGCTGAGCAGGAAATGCAGCTAGAGCCCATTACAGATGAGAGAGTTTATGAAAGTATT GGCCAGTATGGAGGAGAAACTGTAAAAATAGTTCGTATAGAAAAGGCTCGTGATATTC CGTTGGGTGCTACAGTTCGTAATGAAATGGACTCTGTCATCATTAGCCGGATAGTAAA AGGGGGTGCTGCAGAGAAAAGTGGTCTGTTGCATGAAGGAGATGAAGTTCTAGAGATT AATGGCATTGAAATTCGGGGGAAAGATGTCAATGAGGTTTTTGACTTGTTGTCTGATA TGCATGGTACTTTGACTTTTGTCCTGATTCCCAGTCAACAGATCAAGCCGCCTCCTGC CAAGGAAACAGTAATCCATGTAAAAGCTCATTTTGACTATGACCCCTCAGATGACCCT TATGTTCCATGTCGAGAGTTAGGTCTGTCTTTTCAAAAAGGTGATATACTTCATGTGA TCAGTCAAGAAGATCCAAACTGGTGGCAGGCCTACAGGGAAGGGGACGAAGATAATCA ACCTCTAGCCGGGCTTGTTCCAGGGAAAAGCTTTCAGCAGCAAAGGGAAGCCATGAAA CAAACCATAGAAGAAGATAAGGAGCCAGAAAAATCAGGTAAACTGTGGTGTGCAAAGA AGAATAAAAAGAAGAGGAAAAAGGTTTTATATAATGCCAATAAAAATGATGATTATGA CAACGAGGAGATCTTAACCTATGAGGAAATGTCACTTTATCATCAGCCAGCAAATAGG AAGAGACCTATCATCTTGATTGGTCCACAGAACTGTGGCCAGAATGAATTGCGTCAGA GGCTCATGAACAAAGAAAAGGACCGCTTTGCATCTGCAGTTCCTCGTACAACCCGGAG TAGGCGAGACCAAGAAGTAGCCGGTAGAGATTACCACTTTGTTTCGCGGCAAGCATTC GAGGCAGACATAGCAGCTGGAAAGTTCATTGAGCATGGTGAATTTGAGAAGAATTTGT ATGGAACTAGCATAGATTCTGTACGGCAAGTGATCAACTCTGGCAAAATATGTCTTTT AAGTCTTCGTACACAGTCATTGAAGACTCTCCGGAATTCAGATTTGAAACCATATATT ATCTTCATTGCACCCCCTTCACAAGAAAGACTTCGGGCATTATTGGCCAAAGAAGGCA AGAATCCAAAGCCTGAAGAGTTGAGAGAAATCATTGAGAAGACAAGAGAGATGGAGCA GAACAATGGCCACTACTTTGATACGGCAATTGTGAATTCCGATCTTGATAAAGCCTAT CAGGAATTGCTTAGGTTAATTAACAAACTTGATACTGAACCTCAGTGGGTACCATCCA CTTGGCTGAGGTGAAAGAAACATCCATTCT
ORF Start: ATG at 17 ORF Stop: TGA at 2042
SEQ ED NO: 22 675 aa MW at 77311.8kD
NOV9a, MTTSHMNGHVTEESDSEVKNVDLASPEEHQKHREMAVDCPGDLGTRMMPIRRSAQLER IRQQQEDMRRRREEEGKKQELDLNSSMRLKKLAQIPPKTGIDNPMFDTEEGIVLESPH
CG58590-01 Protein Sequence YAVKILEIEDLFSSLKHIQHTLVDSQSQEDISLLLQLVQNKDFQNAFKIHNAITVHMN KASPPFPLISNAQDLAQEVQTVLKPVHHKEGQELTALLNTPHIQA LLAHDKVAEQEM Q EPITDERVYESIGQYGGETVKIVRIEKARDIPLGATVRNEMDSVIISRIVKGGAAE KSGLLHEGDEVLEINGIEIRGKDVNEVFDL SD HGTLTFVLIPSQQIKPPPAKETVI HVKAHFDYDPSDDPYVPCRELGLΞFQKGDI HVISQEDPNWWQAYREGDEDNQPLAG VPGKSFQQQREAMKQTIEEDKEPEKSGKL CAKKNKKKRKKVLYNANKNDDYDNEEIL TYEEMSLYHQPANRKRPIILIGPQNCGQNELRQRLMNKEKDRFASAVPRTTRSRRDQE VAGRDYHFVSRQAFEADIAAGKFIEHGEFEKN YGTSIDSVRQVINSGKICLLSLRTQ S KTLRNSDLKPYIIFIAPPSQER RALLAKEG NPKPEE REIIEKTREMEQN GHY FDTAIV SDLDKAYQELLRLINKLDTEPQWVPSTWLR
SEQ ID NO: 23 2030 bp
NOV9b, CCATGACAACATCCCATATGAATGGGCATGTTACAGAGGAATCAGACAGCGAAGTAAA AAATGTTGATCTTGCATCACCAGAGGAACATCAGAAGCACCGAGAGATGGCTGTTGAC
CG58590-02 DNA Sequence TGCCCTGGAGATTTGGGCACCAGGATGATGCCAATACGTCGAAGTGCACAGTTGGAGC GTATTCGGCAACAACAGGAGGACATGAGGCGTAGGAGAGAGGAAGAAGGGAAAAAGCA AGAACTTGACCTTAATTCTTCCATGAGACTTAAGAAACTAGCCCAAATTCCTCCAAAG ACCGGAATAGATAACCCTATGTTTGATACAGAGGAAGGAATTGTCTTAGAAAGTCCTC ATTATGCTGTGAAAATATTAGAAATAGAAGACTTGTTTTCTTCACTTAAACATATCCA ACATACTTTGGTAGATTCTCAGAGCCAGGAGGATATTTCACTGCTTTTACAACTTGTT CAAAATAAGGATTTCCAGAATGCATTTAAGATACACAATGCCATCACAGTACATATGA ACAAGGCCAGTCCTCCATTTCCTCTTATCTCCAACGCACAAGATCTTGCTCAAGAGGT ACAAACTGTTTTGAAGCCAGTTCATCATAAGGAAGGACAAGAACTAACTGCTTTGCTG AATACTCCACATATTCAGGCACTTTTACTGGCCCACGATAAGGTTGCTGAGCAGGAAA TGCAGCTAGAGCCCATTACAGATGAGAGAGTTTATGAAAGTATTGGCCAGTATGGAGG AGAAACTGTAAAAATAGTTCGTATAGAAAAGGCTCGTGATATTCCGTTGGGTGCTACA GTTCGTAATGAAATGGACTCTGTCATCATTAGCCGGATAGTAAAAGGGGGTGCTGCAG AGAAAAGTGGTCTGTTGCATGAAGGAGATGAAGTTCTAGAGATTAATGGCATTGAAAT TCGGGGGAAAGATGTCAATGAGGTTTTTGACCTGTTGTCTGATATGCATGGTACTTTG ACTTTTGTCCTGATTCCCAGTCAACAGATCAAGCCGCCTCCTGCCAAGGAAACAGTAA TCCATGTAAAAGCTCATTTTGACTATGACCCCTCAGATGACCCTTATGTTCCATGTCG AGAGTTAGGTCTGTCTTTTCAAAAAGGTGATATACTTCATGTGATCAGTCAAGAAGAT CCAAACTGGTGGCAGGCCTACAGGGAAGGGGACGAAGATAATCAACCTCTAGCCGGGC TTGTTCCAGGGAAAAGCTTTCAGCAGCAAAGGGAAGCCATGAAACAAACCATAGAAGA AGATAAGGAGCCAGAAAAATCAGGAAAACTGTGGTGTGCAAAGAAGAATAAAAAGAAG AGGAAAAAGGTTTTATATAATGCCAATAAAAATGATGATTATGACAACGAGGAGATCT TAACCTATGAGGAAATGTCACTTTATCATCAGCCAGCAAATAGGAAGAGACCTATCAT CTTGATTGGTCCACAGAACTGTGGCCAGAATGAATTGCGTCAGAGGCTCATGAACAAA GAAAAGGACCGCTTTGCATCTGCAGTTCCTCATACAACCCGGAGTAGGCGAGACCAAG AAGTAGCCGGTAGAGATTACCACTTTGTTTCGCGGCAAGCATTCGAGGCAGACATAGC AGCTGGAAAGTTCATTGAGCATGGTGAATTTGAGAAGAATTTGTATGGAACTAGCATA GATTCTGTACGGCAAGTGATCAACTCTGGCAAAATATGTCTTTTAAGTCTTCGTACAC AGTCATTGAAGACTCTCCGGAATTCAGATTTGAAACCATATATTATCTTCATTGCACC
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 9B.
Further analysis of the NOV9a protein yielded the following properties shown in Table
9C.
Table 9C. Protein Sequence Properties NOV9a
PSort 0.7000 probability located in nucleus; 0.1000 probability located in analysis: mitochondrial matrix space; 0.1000 probability located in lysosome (lumen); 0.0000 probability located in endoplasmic reticulum (membrane)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV9a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 9D.
In a BLAST search of public sequence databases, the NOV9a protein was found to have homology to the proteins shown in the BLASTP data in Table 9E.
PFam analysis predicts that the NOV9a protein contains the domains shown in the Table 9F.
Example 10.
The NOV10 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 10A.
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 10B.
Table 10B. Comparison of NOVlOa against NOVlOb.
Protein NOVlOa Residues/ Identities/ Sequence Match Residues Similarities for the Matched Region
NOVlOb 1..184 163/184 (88%) 1..184 164/184 (88%)
Further analysis of the NOVlOa protein yielded the following properties shown in Table IOC.
Table IOC. Protein Sequence Properties NOVlOa
PSort 0.4500 probability located in cytoplasm; 0.1206 probability located in analysis: microbody (peroxisome); 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVlOa protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 10D.
Table 10D. Geneseq Results for NOVlOa
NOVlOa Identities/
Geneseq Protein/Organism/Length [Patent #, Expect Residues/ Similarities for Identifier Date] Value
In a BLAST search of public sequence databases, the NOVlOa protein was found to have homology to the proteins shown in the BLASTP data in Table 10E.
PFam analysis predicts that the NOVlOa protein contains the domains shown in the Table 10F.
Example 11.
The NO VI 1 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 11 A.
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 1 IB.
Further analysis of the NOVl la protein yielded the following properties shown in Table l lC.
Table 11C. Protein Sequence Properties NOVlla
PSort 0.4698 probability located in microbody (peroxisome); 0.4500 probability analysis: located in cytoplasm; 0.1958 probability located in lysosome (lumen); 0.1000 probability located in mitochondrial matrix space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVl la protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 1 ID.
In a BLAST search of public sequence databases, the NOVl la protein was found to have homology to the proteins shown in the BLASTP data in Table 1 IE.
PFam analysis predicts that the NOVl la protein contains the domains shown in the Table 1 IF.
Example 12.
The NOVl 2 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 12 A.
Table 12 A. NOVl 2 Sequence Analysis
SEQ TD NO: 37 3696 bp
NOV12a, GTGTAAAAATACTGTCCATTTAATGTTTTCTGGGACTTTAGGTAAGAATATGAAAACT
CAACCACCCTTGAGCAGGATGAACCGGGAGGAATTGGAGGACAGTTTCTTTCGACTTC
CG57819-01 DNA Sequence GCGAAGATCACATGTTGGTGAAGGAGCTTTCTTGGAAGCAACAGGATGAGATCAAAAG GCTGAGGACCACCTTGCTGCGGTTGACCGCTGCTGGCCGGGACCTGCGGGTCGCGGAG GAGGCGGCGCCGCTCTCGGAGACCGCAAGGCGCGGGCAGAAGGCGGGATGGCGGCAGC GCCTCTCCATGCACCAGCGCCCCCAGATGCACCGACTGCAAGGGCATTTCCACTGCGT CGGCCCTGCCAGCCCCCGCCGCGCCCAGCCTCGCGTCCAAGTGGGACACAGACAGCTC CACACAGCCGGTGCACCGGTGCCGGAGAAACCCAAGAGGGGTAGGGACAGGCTGAGCT ACACAGCCCCTCCATCGTTTAAGGAGCATGCGACAAATGAAAACAGAGGTGAAGTAGC CAGTAAACCCAGTGAACTGGCCCACATCATGGCCAGCAATACCATGCAAGTGGAAGAG CCACCCAAGTCTCCTGAGAAAATGTGGCCTAAAGATGAAAATTTTGAACAGAGAAGCT CATTGGAGTGTGCTCAGAAGGCTGCAGAGCTTCGGGCTTCCATTAAAGAGAAGGTAGA GCTGATTCGACTTAAGAAGCTCTTACATGAAAGAAATGCTTCATTGGTTATGACAAAA GCACAATTAACAGAAGTTCAAGAGGTGAGTTGCCATCTTTTGACCCAGAATCAGGGAA TCCTGAGTGCAGCCCATGAGGCCCTCCTCAAGCAAGTGAATGAGCTCAGGGCAGAGCT GAAGGAAGAAAGCAAGAAGGCTGTGAGCTTGAAGAGCCAACTGGAAGATGTGTCTATC TTGCAGATGACTCTGAAGGAGTTTCAGGAGAGAGTTGAAGATTTGGAAAAAGAACGAA AATTGCTGAATGACAATTATGACAAACTCTTAGAAAGCAGTGACAGCTCCAGTCAGCC CCACTGGAGCAACGAGCTCATAGCGGAACAGCTACAGCAGCAAGTCTCTCAGCTGCAG GATCAGCTGGATGCTGAGCTGGAGGACAAGAGAAAAGTTTTACTTGAGCTGTCCAGGG AGAAAGCCCAAAATGAGGATCTGAAGCTTGAAGTCACCAACATACTTCAGAAGCATAA ACAGGAAGTAGAGCTCCTCCAAAATGCAGCCACAATTTCCCAACCTCCTGACAGGCAA TCTGAACCAGCCACTCACCCAGCTGTATTGCAAGAGAACACTCAGATCCAGCCAAGTG AACCCAAAAACCAAGAAGAAAAGAAACTGTCCCAGGTGCTAAATGAGTTGCAAGTATC ACACGCAGAGACCACATTGGAACTAGAAAAGACCAGGGACATGCTTATTCTGCAGCGC AAAATCAACGTGTGTTATCAGGAGGAACTGGAGGCAATGATGACAAAAGCTGACAATG ATAATAGAGATCACAAAGAAAAGCTGGAGAGGTTGACTCGACTACTAGACCTCAAGAA TAACCGTATCAAGCAGCTGGAAGAACAGCTCAAAGATGTTGCTTATGGCACCCGACCG TTGTCGTTATGTTTGGAAACACTGCCAGCCCATGGAGATGAGGATAAAGTGGATATTT CTCTGCTGCATCAGGGTGAGAATCTTTTTGAACTGCACATCCACCAGGCCTTCCTGAC ATCTGCCGCCCTAGCTCAGGCTGGAGATACCCAACCTACCACTTTCTGCACCTATTCC TTCTATGACTTTGAAACCCACTGTACCCCATTATCTGTGGGGCCACAGCCCCTCTATG ACTTCACCTCCCAGTATGTGATGGAGACAGATTCGCTTTTCTTACACTACCTTCAAGA GGCTTCAGCCCGGCTTGACATACACCAGGCCATGGCCAGTGAACACAGCACTCTTGCT GCAGGATGGATTTGCTTTGACAGGGTGCTAGAGACTGTGGAGAAAGTCCATGGCTTGG CCACACTGATTGGTGCTGGTGGAGAAGAGTTCGGGGTTCTAGAGTACTGGATGAGGCT
Further analysis of the NOVl 2a protein yielded the following properties shown in Table 12B.
Table 12B. Protein Sequence Properties NOVl 2a
PSort 0.9600 probability located in nucleus; 0.3000 probability located in analysis: microbody (peroxisome); 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
No Known Signal Sequence Predicted
analysis:
A search of the NOVl 2a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 12C.
In a BLAST search of public sequence databases, the NOVl 2a protein was found to have homology to the proteins shown in the BLASTP data in Table 12D.
PFam analysis predicts that the NOVl 2a protein contains the domains shown in the Table 12E.
Example 13.
The NOVl 3 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 13 A.
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 13B.
Further analysis of the NOVl 3a protein yielded the following properties shown in Table 13C.
Table 13C. Protein Sequence Properties NOV13a
PSort analysis: 0.6500 probability located in plasma membrane; 0.5064 probability located in mitochondrial matrix space; 0.3844 probability located in microbody (peroxisome); 0.2556 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVl 3a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 13D.
In a BLAST search of public sequence databases, the NOVl 3a protein was found to have homology to the proteins shown in the BLASTP data in Table 13E.
PFam analysis predicts that the NOVl 3a protein contains the domains shown in the Table 13F.
Example 14.
The NOVl 4 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 14 A.
Table 14A. NOV14 Sequence Analysis
SEQ ID NO: 43 1790 bp
NOV14a, TCTCCCTCCCGCGCGATGGCCTCGGCGCTGAGCTATGTCTCCAAGTTCAAGTCCTTCG
TGATCTTGTTCGTCACCCCGCTCCTGCTGCTGCCACTCGTCATTCTGATGCCCGCCAA
CG57758-01 DNA Sequence GGTCAGTTGTGCCTACGTCATCATCCTCATGGCCATTTACTGGTGCACAGAAGTCATC CCTCTGGCTGTCACCTCTCTCATGCCTGTCTTGCTTTTCCCACTCTTCCAGATTCTGG ACTCCAGGCAGGTGTGTGTCCAGTACATGAAGGACACCAACATGCTGTTCCTGGGCGG CCTCATCGTGGCCGTGGCTGTGGAGCGCTGGAACCTGCACAAGAGGATCGCCCTGCGC ACGCTCCTCTGGGTGGGGGCCAAGCCTGCACGGCTGATGCTGGGCTTCATGGGCGTCA CAGCCCTCCTGTCCATGTGGATCAGTAACACGGCAACCACGGCCATGATGGTGCCCAT CGTGGAGGCCATATTGCAGCAGATGGAAGCCACAAGCGCAGCCACCGAGGCCGGCCTG GAGCTGGTGGACAAGGGCAAGGCCAAGGAGCTGCCAGGGAGTCAAGTGATTTTTGAAG GCCCCACTCTGGGGCAGCAGGAAGACCAAGAGCGGAAGAGGTTGTGTAAGGCCATGAC CCTGTGCATCTGCTACGCGGCCAGCATCGGGGGCACCGCCACCCTGACCGGGACGGGA CCCAACGTGGTGCTCCTGGGCCAGATGAACGAGTTGTTTCCTGACAGCAAGGACCTCG TGAACTTTGCTTCCTGGTTTGCATTTGCCTTTCCCAACATGCTGGTGATGCTGCTGTT CGCCTGGCTGTGGCTCCAGTTTGTTTACATGTTCTCCAGTTTTAAAAAGTCCTGGGGC TGCGGGCTAGAGAGCAAGAAAAACGAGAAGGCTGCCCTCAAGGTGCTGCAGGAGGAGT ACCGGAAGCTGGGGCCCTTGTCCTTCGCGGAGATCAACGTGCTGATCTGCTTCTTCCT GCTGGTCATCCTGTGGTTCTCCCGAGACCCCGGCTTCATGCCCGGCTGGCTGACTGTT GCCTGGGTGGAGGGTGAGACAAAGTATGTCTCCGATGCCACTGTGGCCATCTTTGTGG CCACCCTGCTATTCATTGTGCCTTCACAGAAGCCCAAGTTTAACTTCCGCAGCCAGAC TGAGGAAGGTAAGTCTCCTGTTCTGATCGCCCCCCCTCCCCTGCTGGATTGGAAGGTA ACCCAGGAGAAAGTGCCCTGGGGCATCGTGCTGCTACTAGGGGGCGGATTTGCTCTGG CTAAAGGATCCGAGGCCTCGGGGCTGTCCGTGTGGATGGGGAAGCAGATGGAGCCCTT GCACGCAGTGCCCCCGGCAGCCATCACCTTGATCTTGTCCTTGCTCGTTGCCGTGTTC ACTGAGTGCACAAGCAACGTGGCCACCACCACCTTGTTCCTGCCCATCTTTGCCTCCA TGTCTCGCTCCATCGGCCTCAATCCGCTGTACATCATGCTGCCCTGTACCCTGAGTGC CTCCTTTGCCTTCATGTTGCCTGTGGCCACCCCTCCAAATGCCATCGTGTTCACCTAT GGGCACCTCAAGGTTGCTGACATGGTGAAAACAGGAGTCATAATGAACATAATTGGAG TCTTCTGTGTGTTTTTGGCTGTCAACACCTGGGGACGGGCCATATTTGACTTGGATCA TTTCCCTGACTGGGCTAATGTGACACATATTGAGACTTAGGAAGAGCCACAAGACCAC
ACACACAGCCCTTACCCTCCTCAGGACTACCGAACCTTCTGGCACACCTT
ORF Start: ATG at 16 ORF Stop: TAG at 1720
SEQ ED NO: 44 568 aa MW at 62592.9kD
NOVl 4a, MASALSYVSKFKSFVILFVTPLLLLPLVILMPAKVSCAYVIILMAIYWCTEVIPLAVT SLMPVLLFPLFQILDSRQVCVQYMKDTNMLFLGGLIVAVAVERWNLHKRIALRTLL V
CG57758-01 Protein Sequence GAKPARLMLGFMGVTALLSMWISNTATTAMMVPIVEAILQQMEATSAATEAGLELVDK GKAKELPGSQVIFEGPTLGQQEDQERKRLCKAMTLCICYAASIGGTATLTGTGPNWL LGQMNELFPDSKDLVNFAS FAFAFPNMLVMLLFA LWLQFVYMFSSFKKS GCGLES KKNEKAALKVLQEEYRKLGPLSFAEINVLICFFLLVIL FSRDPGFMPG LTVA VEG ETKYVSDATVAIFVATLLFIVPSQKPKFNFRSQTEEGKSPVLIAPPPLLD KVTQEKV P GIVLLLGGGFALAKGSEASGLSV MGKQMEPLHAVPPAAITLILSLLVAVFTECTS NVATTTLFLPIFASMSRSIGLNPLYIMLPCTLSASFAFMLPVATPPNAIVFTYGHLKV ADMVKTGVIMNIIGVFCVFLAVNT GRAIFDLDHFPDWANVTHIET
SEQ ED NO: 45 1899 bp
NOV14b, CGTCTCGCCCGCCAGTCTCCCTCCCGCGCGATGGCCTCGGCGCTGAGCTATGTCTCCA
AGTTCAAGTCCTTCGTGATCTTGTTCGTCACCCCGCTCCTGCTGCTGCCACTCGTCAT
CG57758-02 DNA Sequence TCTGATGCCCGCCAAGGTCAGTTGCTGTGCCTACGTCATCATCCTCATGGCCATTTAC TGGTGCACAGAAGTCATCCCTCTGGCTGTCACCTCTCTCATGCCTGTCTTGCTTTTCC CACTCTTCCAGATTCTGGACTCCAGGCAGGTGTGTGTCCAGTACATGAAGGACACCAA CATGCTGTTCCTGGGCGGCCTCATCGTGGCCGTGGCTGTGGAGCGCTGGAACCTGCAC AAGAGGATCGCCCTGCGCACGCTCCTCTGGGTGGGGGCCAAGCCTGCACGGCTGATGC TGGGCTTCATGGGCGTCACAGCCCTCCTGTCCATGTGGATCAGTAACACGGCAACCAC GGCCATGATGGTGCCCATCGTGGAGGCCATATTGCAGCAGATGGAAGCCACAAGCGCA GCCACCGAGGCCGGCCTGGAGGGACAAGGTACCACAATAAACAACCTGAATGCACTGG AGGATGATACAGTGAAAGCAGTACTAGGAGGAAAGTGTGTAGCTATAATAAGCACTTA CGTCAAAAAAGTAGAAAAACTTCAAATAAACAATCTAATGACACCTCTTAAAAAACTA GAAAAGCAAGAGCAACAGGACCTAGGGCCTGGCATCAGGCCTCAGGACTCTGCCCAGT GCCAGGAAGACCAAGAGCGGAAGAGGTTGTGTAAGGCCATGACCCTGTGCATCTGCTA CGCGGCCAGCATCGGGGGCACCGCCACCCTGACCGGGACGGGACCCAACGTGGTGCTC CTGGGCCAGATGAACGAGTTGTTTCCTGACAGCAAGGACCTCGTGAACTTTGCTTCCT GGTTTGCATTTGCCTTTCCCAACATGCTGGTGATGCTGCTGTTCGCCTGGCTGTGGCT CCAGTTTGTTTACATGTTCTCCAGTTTTAAAAAGTCCTGGGGCTGCGGGCTAGAGAGC AAGAAAAACGAGAAGGCTGCCCTCAAGGTGCTGCAGGAGGAGTACCGGAAGCTGGGGC CCTTGTCCTTCGCGGAGATCAACGTGCTGATCTGCTTCTTCCTGCTGGTCATCCTGTG GTTCTCCCGAGACCCCGGCTTCATGCCCGGCTGGCTGACTGTTGCCTGGGTGGAGGGT GAGACAAAGTCAGTCTCCGATGCCACTGTGGCCATCTTTGTGGCCACCCTGCTATTCA TTGTGCCTTCACAGAAGCCCAAGTTTAACTTCCGCAGCCAGACTGAGGAAGGTAAGTC TCCTGTTCTGATCGCCCCCCCTCCCCTGCTGGATTGGAAGGTAACCCAGGAGAAAGTG CCCTGGGGCATCGTGCTGCTACTAGGGGGCGGATTTGCTCTGGCTAAAGGATCCGAGG
CCTCGGGGCTGTCCGTGTGGATGGGGAAGCAGATGGAGCCCTTGCACGCAGTGCCCCC GGCAGCCATCACCTTGATCTTGTCCTTGCTCGTTGCCGTGTTCACTGAGTGCACAAGC AACGTGGCCACCACCACCTTGTTCCTGCCCATCTTTGCCTCCATGTCTCGCTCCATCG GCCTCAATCCGCTGTACATCATGCTGCCCTGTACCCTGAGTGCCTCCTTTGCCTTCAT GTTGCCTGTGGCCACCCCTCCAAATGCCATCGTGTTCACCTATGGGCACCTCAAGGTT GCTGACATGGTAAAAACAGGAGTCATAATGAACATAATTGGAGTCTTCTGTGTGTTTT TGGCTGTCAACACCTGGGGACGGGCCATATTTGACTTGGATCATTTCCCTGACTGGGC TAATGTGACACATATTGAGACTTAGGAAGAGCCACAAGACCAC
ORF Start: ATG at 31 ORF Stop: TAG at 1879
SEQ ID NO: 46 616 aa MW at 67816.9kD
NOV14b, MASALSYVSKFKSFVILFVTPLLLLPLVILMPAKVSCCAYVIILMAIYWCTEVIPLAV TSLMPVLLFPLFQILDSRQVCVQYMKDTNMLFLGGLIVAVAVER NLHKRIALRTLLW
CG57758-02 Protein Sequence VGAKPARLMLGFMGVTALLSM ISNTATTAMMVPIVEAILQQMEATSAATEAGLEGQG TTINNLNALEDDTVKAVLGGKCVAIISTYVKKVEKLQINNLMTPLKKLEKQEQQDLGP GIRPQDSAQCQEDQERKRLCKAMTLCICYAASIGGTATLTGTGPNWLLGQMNELFPD SKDLVNFAS FAFAFPNMLVMLLFAWL LQFVYMFSSFKKSWGCGLESKKNEKAALKV LQEEYRKLGPLSFAEINVLICFFLLVIL FSRDPGFMPG LTVA VEGETKSVSDATV AIFVATLLFIVPSQKPKFNFRSQTEEGKSPVLIAPPPLLDWKVTQEKVP GIVLLLGG GFALAKGSEASGLSV MGKQMEPLHAVPPAAITLILSLLVAVFTECTSNVATTTLFLP IFASMSRSIGLNPLYIMLPCTLSASFAFMLPVATPPNAIVFTYGHLKVADMVKTGVIM NIIGVFCVFLAVNT GRAIFDLDHFPD ANVTHIET
SEQ ID NO: 47 1899 bp
NOVl 4c, CGTCTCGCCCGCCAGTCTCCCTCCCGCGCGATGGCCTCGGCGCTGAGCTATGTCTCCA
AGTTCAAGTCCTTCGTGATCTTGTTCGTCACCCCGCTCCTGCTGCTGCCACTCGTCAT
CG57758-03 DNA Sequence TCTGATGCCCGCCAAGGTCAGTTGCTGTGCCTACGTCATCATCCTCATGGCCATTTAC TGGTGCACAGAAGTCATCCCTCTGGCTGTCACCTCTCTCATGCCTGTCTTGCTTTTCC CACTCTTCCAGATTCTGGACTCCAGGCAGGTGTGTGTCCAGTACATGAAGGACACCAA CATGCTGTTCCTGGGCGGCCTCATCGTGGCCGTGGCTGTGGAGCGCTGGAACCTGCAC AAGAGGATCGCCCTGCGCACGCTCCTCTGGGTGGGGGCCAAGCCTGCACGGCTGATGC TGGGCTTCATGGGCGTCACAGCCCTCCTGTCCATGTGGATCAGTAACACGGCAACCAC GGCCATGATGGTGCCCATCGTGGAGGCCATATTGCAGCAGATGGAAGCCACAAGCGCA GCCACCGAGGCCGGCCTGGAGGGACAAGGTACCACAATAAACAACCTGAATGCACTGG AGGATGATACAGTGAAAGCAGTACTAGGAGGAAAGTGTGTAGCTATAATAAGCACTTA CGTCAAAAAAGTAGAAAAACTTCAAATAAACAATCTAATGACACCTCTTAAAAAACTA GAAAAGCAAGAGCAACAGGACCTAGGGCCTGGCATCAGGCCTCAGGACTCTGCCCAGT GCCAGGAAGACCAAGAGCGGAAGAGGTTGTGTAAGGCCATGACCCTGTGCATCTGCTA CGCGGCCAGCATCGGGGGCACCGCCACCCTGACCGGGACGGGACCCAACGTGGTGCTC CTGGGCCAGATGAACGAGTTGTTTCCTGACAGCAAGGACCTCGTGAACTTTGCTTCCT GGTTTGCATTTGCCTTTCCCAACATGCTGGTGATGCTGCTGTTCGCCTGGCTGTGGCT CCAGTTTGTTTACATGTTCTCCAGTTTTAAAAAGTCCTGGGGCTGCGGGCTAGAGAGC AAGAAAAACGAGAAGGCTGCCCTCAAGGTGCTGCAGGAGGAGTACCGGAAGCTGGGGC CCTTGTCCTTCGCGGAGATCAACGTGCTGATCTGCTTCTTCCTGCTGGTCATCCTGTG GTTCTCCCGAGACCCCGGCTTCATGCCCGGCTGGCTGACTGTTGCCTGGGTGGAGGGT GAGACAAAGTCAGTCTCCGATGCCACTGTGGCCATCTTTGTGGCCACCCTGCTATTCA TTGTGCCTTCACAGAAGCCCAAGTTTAACTTCCGCAGCCAGACTGAGGAAGGTAAGTC TCCTGTTCTGATCGCCCCCCCTCCCCTGCTGGATTGGAAGGTAACCCAGGAGAAAGTG CCCTGGGGCATCGTGCTGCTACTAGGGGGCGGATTTGCTCTGGCTAAAGGATCCGAGG CCTCGGGGCTGTCCGTGTGGATGGGGAAGCAGATGGAGCCCTTGCACGCAGTGCCCCC GGCAGCCATCACCTTGATCTTGTCCTTGCTCGTTGCCGTGTTCACTGAGTGCACAAGC AACGTGGCCACCACCACCTTGTTCCTGCCCATCTTTGCCTCCATGTCTCGCTCCATCG GCCTCAATCCGCTGTACATCATGCTGCCCTGTACCCTGAGTGCCTCCTTTGCCTTCAT GTTGCCTGTGGCCACCCCTCCAAATGCCATCGTGTTCACCTATGGGCACCTCAAGGTT GCTGACATGGTAAAAACAGGAGTCATAATGAACATAATTGGAGTCTTCTGTGTGTTTT TGGCTGTCAACACCTGGGGACGGGCCATATTTGACTTGGATCATTTCCCTGACTGGGC TAATGTGACACATATTGAGACTTAGGAAGAGCCACAAGACCAC
ORF Start: ATG at 31 ORF Stop: TAG at 1879
SEQ TD NO: 48 616 aa MW at 67816.9kD
NOV14c, MASALSYVSKFKSFVILFVTPLLLLPLVILMPAKVSCCAYVIILMAIY CTEVIPLAV TSLMPVLLFPLFQILDSRQVCVQYMKDTNMLFLGGLIVAVAVER NLHKRIALRTLLW
CG57758-03 Protein Sequence VGAKPARLMLGFMGVTALLSMWISNTATTAMMVPIVEAILQQMEATSAATEAGLEGQG TTINNLNALEDDTVKAVLGGKCVAIISTYVKKVEKLQINNLMTPLKKLEKQEQQDLGP GIRPQDSAQCQEDQERKRLCKAMTLCICYAASIGGTATLTGTGPNWLLGQMNELFPD SKDLVNFAS FAFAFPNMLVMLLFAWLWLQFVYMFSSFKKS GCGLESKKNEKAALKV LQEEYRKLGPLSFAEINVLICFFLLVILWFSRDPGFMPG LTVAWVEGETKSVSDATV AIFVATLLFIVPSQKPKFNFRSQTEEGKSPVLIAPPPLLD KVTQEKVP GIVLLLGG GFALAKGSEASGLSV MGKQMEPLHAVPPAAITLILSLLVAVFTECTSNVATTTLFLP IFASMSRSIGLNPLYIMLPCTLSASFAFMLPVATPPNAIVFTYGHLKVADMVKTGVIM NlIGVFCVFLAVNT GRAIFDLDHFPDWANVTHIET
SEQ ID NO: 49 1606 bp
NOV14d, GATGGCCTCGGCGCTGAGCTATGTCTCCAAGTTCAAGTCCTTCGTGATCTTGTTCGTC ACCCCGCTCCTGCTGCTGCCACTCGTCATTCTGATGCCCGCCAAGTTTGTCAGGTGTG
CG57758-04 DNA Sequence CCTACGTCATCATCCTCATGGCCATTTACTGGTGCACAGAAGTCATCCCTCTGGCTGT CACCTCTCTCATGCCTGTCTTGCTTTTCCCACTCTTCCAGATTCTGGACTCCAGGCAG GTGTGTGTCCAGTACATGAAGGACACCAACATGCTGTTCCTGGGCGGCCTCATCGTGG CCGTGGCTGTGGAGCGCTGGAACCTGCACAAGAGGATCGCCCTGCGCACGCTCCTCTG GGTGGGGGCCAAGCCTGCACGGCTGATGCTGGGCTTCATGGGCGTCACAGCCCTCCTG TCCATGTGGATCAGTAACACGGCAACCACGGCCATGATGGTGCCCATCGTGGAGGCCA TATTGCAGCAGATGGAAGCCACAAGCGCAGCCACCGAGGCCGGCCTGGAGCTGGTGGA CAAGGGCAAGGCCAAGGAGCTGCCAGGGAGTCAAGTGATTTTTGAAGGCCCCACTCTG GGGCAGCAGGAAGACCAAGAGCGGAAGAGGTTGTGTAAGGCCATGACCCTGTGCATCT GCTACGCGGCCAGCATCGGGGGCACCGCCACCCTGACCGGGACGGGACCCAACGTGGT GCTCCTGGGCCAGATGAACGAGTTGTTTCCTGACAGCAAGGACCTCGTGAACTTTGCT TCCTGGTTTGCATTTGCCTTTCCCAACATGCTGGTGATGCTGCTGTTCGCCTGGCTGT GGCTCCAGTTTGTTTACATGAGATTCAATTTTAAAAAGTCCTGGGGCTGCGGGCTAGA GAGCAAGAAAAACGAGAAGGCTGCCCTCAAGGTGCTGCAGGAGGAGTACCGGAAGTTG GGGCCCTTGTCCTTCGCGGAGATCAACGTGCTGATCTGCTTCTTCCTGCTGGTCATCC TGTGGTTCTCCCGAGACCCCGGCTTCATGCCCGGCTGGCTGACTGTTGCCTGGGTGGA GGGTGAGACAAAGTATGTCTCCGATGCCACTGTGGCCATCTTTGTGGCCACCCTGCTA TTCATTGTGCCTTCACAGAAGCCCAAGTTTAACTTCCGCAGCCAGACTGAGGAAGAAA GGAAAACTCCATTTTATCCCCCTCCCCTGCTGGATTGGAAGGTAACCCAGGAGAAAGT GCCCTGGGGCATCGTGCTGCTACTAGGGGGCGGATTTGCTCTGGCTAAAGGATCCGAG GCCTCGGGGCTGTCCGTGTGGATGGGGAAGCAGATGGAGCCCTTGCACGCAGTGCCCC CGGCAGCCATCACCTTGATCTTGTCCTTGCTCGTTGCCGTGTTCACTGAGTGCACAAG CAACGTGGCCACCACCACCTTGTTCCTGCCCATCTTTGCCTCCATGGTGAAAACAGGA GTCATAATGAACATAATTGGAGTCTTCTGTGTGTTTTTGGCTGTCAACACCTGGGGAC GGGCCATATTTGACTTGGATCATTTCCCTGACTGGGCTAATGTGACACATATTGAGAC TTAGGAAGAGCCACAAGACCACACACATAGCCCTTACCCT
ORF Start: ATG at 2 ORF Stop: TAG at 1568
SEQ ED NO: 50 522 aa MW at 58109.6kD
NOV14d, MASALSYVSKFKSFVI FVTPLLLLPLVILMPAKFVRCAYVIILMAIY CTEVIPLAV TSLMPVLLFPLFQILDSRQVCVQYMKDTNMLFLGGLIVAVAVER NLHKRIALRTLL
CG57758-04 Protein Sequence VGAKPARLMLGFMGVTALLSMWISNTATTAMMVPIVEAILQQMEATSAATEAGLELVD KGKAKELPGSQVIFEGPTLGQQEDQERKRLCKAMTLCICYAASIGGTATLTGTGPNW LLGQMNELFPDSKDLVNFAS FAFAFPNMLVMLLFAWLWLQFVYMRFNFKKS GCGLE SKKNEKAALKVLQEEYRKLGPLSFAEINVLICFFLLVIL FSRDPGFMPGWLTVAWVE GETKYVSDATVAIFVATLLFIVPSQKPKFNFRSQTEEERKTPFYPPPLLD KVTQEKV P GIVLLLGGGFALAKGSEASGLSVWMGKQMEPLHAVPPAAITLILSLLVAVFTECTS NVATTTLFLPIFASMVKTGVIMNIIGVFCVFLAVNTWGRAIFDLDHFPD ANVTHIET
SEQ ED NO: 51 1781 bp
NOV14e, GATGGCCTCGGCGCTGAGCTATGTCTCCAAGTTCAAGTCCTTCGTGATCTTGTTCGTC ACCCCGCTCCTGCTGCTGCCACTCGTCATTCTGATGCCCGCCAAGTTTGTCAGGTGTG
CG57758-05 DNA Sequence CCTACGTCATCATCCTCATGGCCATTTACTGGTGCACAGAAGTCATCCCTCTGGCTGT CACCTCTCTCATGCCTGTCTTGCTTTTCCCACTCTTCCAGATTCTGGACTCCAGGCAG GTGTGTGTCCAGTACATGAAGGACACCAACATGCTGTTCCTGGGCGGCCTCATCGTGG CCGTGGCTGTGGAGCGCTGGAACCTGCACAAGAGGATCGCCCTGCGCACGCTCCTCTG GGTGGGGGCCAAGCCTGCACGGCTGATGCTGGGCTTCATGGGCGTCACAGCCCTCCTG TCCATGTGGATCAGTAACACGGCAACCACGGCCATGATGGTGCCCATCGTGGAGGCCA TATTGCAGCAGATGGAAGCCACAAGCGCAGCCACCGAGGCCGGCCTGGAGCTGGTGGA CAAGGGCAAGGCCAAGGAGCTGCCAGGGAGTCAAGTGATTTTTGAAGGCCCCACTCTG GGGCAGCAGGAAGACCAAGAGCGGAAGAGGTTGTGTAAGGCCATGACCCTGTGCATCT GCTACGCGGCCAGCATCGGGGGCACCGCCACCCTGACCGGGACGGGACCCAACGTGGT GCTCCTGGGCCAGATGAACGAGTTGTTTCCTGACAGCAAGGACCTCGTGAACTTTGCT TCCTGGTTTGCATTTGCCTTTCCCAACATGCTGGTGATGCTGCTGTTCGCCTGGCTGT GGCTCCAGTTTGTTTACATGAGATTCAATTTTAAAAAGTCCTGGGGCTGCGGGCTAGA GAGCAAGAAAAACGAGAAGGCTGCCCTCAAGGTGCTGCAGGAGGAGTACCGGAAGTTG GGGCCCTTGTCCTTCGCGGAGATCAACGTGCTGATCTGCTTCTTCCTGCTGGTCATCC TGTGGTTCTCCCGAGACCCCGGCTTCATGCCCGGCTGGCTGACTGTTGCCTGGGTGGA GGGTGAGACAAAGTATGTCTCCGATGCCACTGTGGCCATCTTTGTGGCCACCCTGCTA TTCATTGTGCCTTCACAGAAGCCCAAGTTTAACTTCCGCAGCCAGACTGAGGAAGAAA GGAAAACTCCATTTTATCCCCCTCCCCTGCTGGATTGGAAGGTAACCCAGGAGAAAGT GCCCTGGGGCATCGTGCTGCTACTAGGGGGCGGATTTGCTCTGGCTAAAGGATCCGAG GCCTCGGGGCTGTCCGTGTGGATGGGGAAGCAGATGGAGCCCTTGCACGCAGTGCCCC CGGCAGCCATCACCTTGATCTTGTCCTTGCTCGTTGCCGTGTTCACTGAGTGCACAAG CAACGTGGCCACCACCACCTTGTTCCTGCCCATCTTTGCCTCCATGAATCACGTCCCC AAGAGCTTCTGTGTTCTGTACGGTGATGTTGCAGTGCTGTCTTTCCGCAGTCTCGCTC CATCGGCCTCAATCCGCTGTACATCATGCTGCCCTGTACCCTGAGTGCCTCCTTTGCC TTCATGTTGCCTGTGGCCACCCCTCCAAATGCCATCGTGTTCACCTATGGGCACCTCA
AGGTTGCTGACATGGTGAAAACAGGAGTCATAATGAACATAATTGGAGTCTTCTGTGT
GTTTTTGGCTGTCAACACCTGGGGACGGGCCATATTTGACTTGGATCATTTCCCTGAC
TGGGCTAATGTGACACATATTGAGACTTAGGAAGAGCCACA
ORF Start: ATG at 2 jORF Stop: TGA at 1550
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 14B.
Further analysis of the NOVl 4a protein yielded the following properties shown in Table 14C.
Table 14C. Protein Sequence Properties NOV14a
PSort analysis: 0.6400 probability located in plasma membrane; 0.4600 probability located in Golgi body; 0.3700 probability located in endoplasmic reticulum (membrane); 0.1000 probability located in endoplasmic reticulum (lumen)
SignalP Likely cleavage site between residues 38 and 39 analysis:
A search of the NOVl 4a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 14D.
In a BLAST search of public sequence databases, the NOVl 4a protein was found to have homology to the proteins shown in the BLASTP data in Table 14E.
PFam analysis predicts that the NOVl 4a protein contains the domains shown in the Table 14F.
Example 15.
The NOVl 5 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 15 A.
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 15B.
Further analysis of the NOVl 5a protein yielded the following properties shown in Table 15C.
Table 15C. Protein Sequence Properties NOV15a
PSort 0.7600 probability located in nucleus; 0.3000 probability located in analysis: microbody (peroxisome); 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV15a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several . homologous proteins shown in Table 15D.
In a BLAST search of public sequence databases, the NOVl 5a protein was found to have homology to the proteins shown in the BLASTP data in Table 15E.
PFam analysis predicts that the NOVl 5a protein contains the domains shown in the Table 15F.
Example 16.
The NOVl 6 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 16A.
Further analysis of the NOVl 6a protein yielded the following properties shown in Table 16B.
Table 16B. Protein Sequence Properties NOVl 6a
PSort 0.9081 probability located in mitochondrial matrix space; 0.6000 probability analysis: located in mitochondrial inner membrane; 0.6000 probability located in mitochondrial intermembrane space; 0.6000 probability located in mitochondrial outer membrane
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVl 6a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 16C.
In a BLAST search of public sequence databases, the NOVl 6a protein was found to have homology to the proteins shown in the BLASTP data in Table 16D.
PFam analysis predicts that the NOVl 6a protein contains the domains shown in the Table 16E.
Table 16E. Domain Analysis of NOVl 6a
Identities/
Pfam Domain NOVl 6a Match Region Similarities Expect Value for the Matched Region
No Significant Matches Found
Example 17.
The NOVl 7 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 17A.
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 17B.
Further analysis of the NOVl 7a protein yielded the following properties shown in Table 17C.
Table 17C. Protein Sequence Properties NOVl 7a
PSort 0.4500 probability located in cytoplasm; 0.3000 probability located in microbody analysis: (peroxisome); 0.1682 probability located in lysosome (lumen); 0.1000 probability located in mitochondrial matrix space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVl 7a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 17D.
In a BLAST search of public sequence databases, the NOVl 7a protein was found to have homology to the proteins shown in the BLASTP data in Table 17E.
PFam analysis predicts that the NOVl 7a protein contains the domains shown in the Table 17F.
Example 18.
The NOVl 8 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 18 A.
Table 18A. NOV18 Sequence Analysis
SEQ ID NO: 69 2109 bp
NOVl 8a, GGGTCCGGCGGGCATCGGCAAGACCATGGCGGCCAAAAATATCCTGTACGACTGGGCG
GCGGGCAAGCTGTACCAGGGCCAGGTGGACTTCGCCTTCTTCATGCCCTGCGGCGAGC
CG58553-01 DNA Sequence TGCTGGAGAGGCCGGGCACGCGCAGCCTGGCTGACCTGATCCTGGACCAGTGCCCCGA CCGCGGCGCGCCGGTGCCGCAGATGCTGGCCCAGCCGCAGCGGCTGCTCTTCATCCTG GACGGCGCGGACGAGCTGCCGGCGCTGGGGGGCCCCGAGGCCGCGCCCTGCACAGACC CCTTCGAGGCGGCGAGCGGCGCGCGGGTGCTAGGCGGGCTGCTGAGTAAGGCGCTGCT GCCCACGGCCCTCCTGCTGGTGACCACGCGCGCCGCCGCCCCCGGGAGGCTGCAGGGC CGCCTGTGTTCCCCGCAGTGCGCCGAGGTGCGCGGCTTCTCCGACAAGGACAAGAAGA AGTATTTCTACAAGTTCTTCCGGGATGAGAGGAGGGCCGAGCGCGCCTACCGCTTCGT GAAGGAGAACGAGACGCTGTTCGCGCTGTGCTTCGTGCCCTTCGTGTGCTGGATCGTG TGCACCGTGCTGCGCCAGCAGCTGGAGCTCGGTCGGGACCTGTCGCGCACGTCCAAGA CCACCACGTCAGTGTACCTGCTTTTCATCACCAGCGTTCTGAGCTCGGCTCCGGTAGC CGACGGGCCCCGGTTGCAGGGCGACCTGCGCAATCTGTGCCGCCTGGCCCGCGAGGGC GTCCTCGGACGCAGGGCGCAGTTTGCCGAGAAGGAACTGGAGCAACTGGAGCTTCGTG GCTCCAAAGTGCAGACGCTGTTTCTCAGCAAAAAGGAGCTGCCGGGCGTGCTGGAGAC AGAGGTCACCTACCAGTTCATCGACCAGAGCTTCCAGGAGTCCTTCGCGGCACTGTCC TACCTGCTGGAGGACGGCGGGGTGCCCAGGACCGCGGCTGGCGGCGTTGGGACACTCC TGCGTGGGGACGCCCAGCCGCACAGCCACTTGGTGCTCACCACGCGCTTCCTCTTCGG ACTGCTGAGCGCGGAGCGGATGCGCGACATCGAGCGCCACTTCGGCTGCATGGTTTCA GAGCGTGTGAAGCAGGAGGCCCTGCGGTGGGTGCAGGGACAGGGACAGGGCTGCCCCG GAGTGGCACCAGAGGTGACCGAGGGGGCCAAAGGGCTCGAGGACACCGAAGAGCCAGA GGAGGAGGAGGAGGGAGAGGAGCCCAACTACCCACTGGAGTTGCTGTACTGCCTGTAC GAGACGCAGGAGGACGCGTTTGTGCGCCAAGCCCTGGGCCGGTTCCCGGAGCTGGCGC TGCAGCGAGTGCGCTTCTGCCGCATGGACGTGGCTGTTCTGAGCTACTGCGTGAGGTG CTGCCCTGCTGCACAGGCACTGCGGCTGATCAGCTGCAGATTGGTTGCTGCGCAGGAG AAGAAGAAGAAGAGCCTGGGGAAGCGGCTCCAGGCCAGCCTGGGCACCACAAAACAAC TGCCAGCCTCCCTTCTTCATCCACTCTTTCAGGCAATGACTGACCCACTGTGCCATCT GAGCAGCCTCACGCTGTCCCACTGCAAACTCCCTGACGCGGTCTGCCGAGACCTTTCT GAGGCCCTGAGGGCAGCCCCCGCACTGACGGAGCTGGGCCTCCTCCACAACAGGCTCA GTGAGGCAGGACTGCGTATGCTGAGTGAGGGCCTAGCCTGGCCGCAGTGCAGGGTGCA GACGGTCAGGGTACAGCTGCCTGACCCCCAGCGAGGGCTCCAGTACCTGGTGGGTATG CTTCGGCAGAGCCCTGCCCTGACCACCCTGGATCTCAGCGGCTGCCAACTGCCCGCCC CCATGGTGACCTACCTGTGTGCAGTCCTGCAGCACCAGGGATGCGGCCTGCAGACCCT CAGTCTGGCCTCTGTGGAGCTGAGCGAGCAGTCACTACAGGAGCTTCAGGCTGTGAAG AGAGCAAAGCCGGATCTGGTCATCACACACCCAGCGCTGGACGGCCACCCACAACCTC CCAAGGAACTCATCTCGACCTTCTGAGGCTCTGGTGGCCAGAGCAGGGTGGAAGACCC TAGTCAAAGTCCCTGTGGAGA
ORF Start: ATG at 26 ORF Stop: TGA at 2054
SEQ TD NO: 70 676 aa MW at 74650.3kD
NOVl 8a, MAAKNILYDWAAGKLYQGQVDFAFFMPCGELLERPGTRSLADLILDQCPDRGAPVPQM LAQPQRLLFILDGADELPALGGPEAAPCTDPFEAASGARVLGGLLSKALLPTALLLVT
CG58553-01 Protein Sequence TRAAAPGRLQGRLCSPQCAEVRGFSDKDKKKYFYKFFRDERRAERAYRFVKENETLFA LCFVPFVCWIVCTVLRQQLELGRDLSRTSKTTTSVYLLFITSVLSSAPVADGPRLQGD LRNLCRLAREGVLGRRAQFAEKELEQLELRGSKVQTLFLSKKELPGVLETEVTYQFID QSFQESFAALSYLLEDGGVPRTAAGGVGTLLRGDAQPHSHLVLTTRFLFGLLSAERMR DIERHFGCMVSERVKQEALR VQGQGQGCPGVAPEVTEGAKGLEDTEEPEEEEEGEEP NYPLELLYCLYETQEDAFVRQALGRFPELALQRVRFCRMDVAVLSYCVRCCPAAQALR LISCRLVAAQEKKKKSLGKRLQASLGTTKQLPASLLHPLFQAMTDPLCHLSSLTLSHC KLPDAVCRDLSEALRAAPALTELGLLHNRLSEAGLR LSEGLA PQCRVQTVRVQLPD PQRGLQYLVGMLRQSPALTTLDLSGCQLPAP VTYLCAVLQHQGCGLQTLSLASVELS EQSLQELQAVKRAKPDLVITHPALDGHPQPPKELISTF
Further analysis of the NOVl 8a protein yielded the following properties shown in Table 18B.
Table 18B. Protein Sequence Properties NOVl 8a
Psort 0.7400 probability located in nucleus; 0.6000 probability located in endoplasmic analysis: reticulum (membrane); 0.3000 probability located in microbody (peroxisome); 0.1000 probability located in mitochondrial inner membrane
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVl 8a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 18C.
In a BLAST search of public sequence databases, the NOVl 8a protein was found to have homology to the proteins shown in the BLASTP data in Table 18D.
PFam analysis predicts that the NOVl 8a protein contains the domains shown in the Table 18E.
Table 18E. Domain Analysis of NOV18a
Identities/
Pfam Domain NOVl 8a Match Region Similarities Expect Value for the Matched Region
No Significant Matches Found
Example 19.
The NOVl 9 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 19 A.
AGGTGGATGTGACCCAAGGAGAGTGCTACCCGGTGTACTGGAACCGTGCTGATAAAAT ACCAGTAATGCGTGGACAGTGGTTTATTGACGGCACTTGGCAGCCTCTAGAAGAGGAA GAAAGTAATTTAATTGAGCAAGAACATCTCAATTGTTTTAGGGGCCAGCAGATGCAGG AAAATTTCGATATTGAAGTGTCAAAATCCATAGATGGAAAAGATGCTGTTCATAGTTT CAAGTTGAGTCGAAACCATGTGGACTGGCACAGTGTGGATGAAGTATATCTTTATAGT GATGCAACAACATCTAAAATTGCAAGAACAGTTACCCAAAAACTGGGATTTTCTAAAG CATCAAGTAGTGGTACCAGACTTCATAGAGGTTATGTAGAAGAAGCCACATTAGAAGA CAAGCCATCACAGACTACCCATATTGTATTTGTTGTGCATGGCATTGGGCAGAAAATG GACCAAGGAAGAATTATCAAAAATACAGCTATGATGAGAGAAGCTGCAAGAAAAATAG AAGAAAGGCATTTTTCCAACCATGCAACACATGTTGAATTTCTGCCTGTTGAGTGGCG GTCAAAACTTACTCTTGATGGAGACACTGTTGATTCCATTACTCCTGACAAAGTACGA GGTTTAAGGGATATGCTGAACAGCAGTGCAATGGACATAATGTATTATACTAGTCCAC TTTATAGAGATGAACTAGTTAAAGGCCTTCAGCAAGAGCTGAATCGATTGTATTCCCT TTTCTGTTCTCGGAATCCAGACTTTGAAGAAAAAGGGGGTAAAGTCTCAATAGTATCA CATTCCTTGGGATGTGTAATTACTTATGACATAATGACTGGCTGGAATCCAGTTCGGC TGTATGAACAGTTGCTGCAAAAGGAAGAAGAGTTGCCTGATGAACGATGGATGAGCTA TGAAGAACGACATCTTCTTGATGAACTCTATATAACTAAACGACGGCTGAAGGAAATA GAAGAACGGCTTCACGGATTGAAAGCATCATCTATGACACAAACACCTGCCTTAAAAT TTAAGGTAGAGAATTTCTTCTGTATGGGATCCCCATTAGCAGTTTTCTTGGCGTTGCG TGGCATCCGCCCAGGAAATACTGGAAGTCAAGACCATATTTTGCCTAGAGAGATTTGT AACCGGTTACTAAATATTTTTCATCCTACAGATCCAGTGGCTTATAGATTAGAACCAT TAATACTGAAACACTACAGCAACATTTCACCTGTCCAGATCCACTGGTACAATACTTC AAATCCTTTACCTTATGAACATATGAAGCCAAGCTTTCTCAACCCAGCTAAAGAACCT ACCTCAGTTTCAGAGAATGAAGGCATTTCAACCATACCAAGCCCTGTGACCTCACCAG TTTTGTCCCGCCGACACTATGGAGAATCTATAACAAATATAGGCAAAGCAAGCATATT AGGTGCTGCTAGCATTGGAAAGGGACTTGGAGGAATGTTGTTCTCAAGATTTGGACGT TCATCTACAACACAGTCATCTGAAACATCAAAAGACTCAATGGAAGATGAGAAGAAGC CAGTTGCCTCACCTTCTGCTACCACCGTAGGGACACAGACCCTTCCACATAGCAGTTC TGGCTTCCTCGATTCTGCAGTGGAGTTGGATCACAGGATTGATTTTGAACTCAGAGAA GGCCTTGTGGAGAGCCGCTATTGGTCAGCTGTCACGTCGCATACTGCCTATTGGTCAT CCTTGGATGTTGCCCTTTTTCTTTTAACCTTCATGTATAAACATGAGCACGATGATGA TGCAAAACCCAATTTAGATCCAATCTGAACTCTCTTGAAGGACATGAATGGCCTAAAA CTGATTTTTTTTTTTTCC
ORF Start: ATG at 20 ORF Stop: TGA at 2636
SEQ ID NO: 72 872 aa MW at 97063.4kD
NOVl 9a, MNYPGRGSPRSPEHNGRGGGGGAWELGSDARPAFGGGVCCFEHLPGGDPDDGDVPLAL LRGEPGLHLAPGTDDHNHHLALDPCLSDENYDFSSAESGSSLRYYSEGESGGGGSSLS
CG58626-01 Protein Sequence LHPPQQPPLVPTNSGGGGATGGSPGERKRTRLGGPAARHRYEWTELGPEEVR FYKE DKKT KPFIGYDSLRIELAFRTLLQTTGARPQGGDRDGDHVCSPTGPASSSGEDDDED RACGFCQSTTGHEPEMVELVNIEPVCVRGGLYEVDVTQGECYPVYWNRADKIPVMRGQ WFIDGTWQPLEEEESNLIEQEHLNCFRGQQMQENFDIEVSKSIDGKDAVHSFKLSRNH VD HSVDEVYLYSDATTSKIARTVTQKLGFSKASSSGTRLHRGYVEEATLEDKPSQTT HIVFWHGIGQKMDQGRIIKNTAMMREAARKIEERHFSNHATHVEFLPVE RSKLTLD GDTVDSITPDKVRGLRDMLNSSAMDIMYYTSPLYRDELVKGLQQELNRLYSLFCSRNP DFEEKGGKVSIVSHSLGCVITYDIMTG NPVRLYEQLLQKEEELPDER MSYEERHLL DELYITKRRLKEIEERLHGLKASSMTQTPALKFKVENFFCMGSPLAVFLALRGIRPGN TGSQDHILPREICNRLLNIFHPTDPVAYRLEPLILKHYSNISPVQIHWYNTSNPLPYE HMKPSFLNPAKEPTSVSENEGISTIPSPVTSPVLSRRHYGESITNIGKASILGAASIG KGLGGMLFSRFGRSSTTQSSETSKDS EDEKKPVASPSATTVGTQTLPHSSSGFLDSA VELDHRIDFELREGLVESRYWSAVTSHTAYWSSLDVALFLLTFMYKHEHDDDAKPNLD PI
Further analysis of the NOVl 9a protein yielded the following properties shown in Table 19B.
Table 19B. Protein Sequence Properties NOVl 9a
PSort 0.4555 probability located in microbody (peroxisome); 0.4500 probability analysis: located in cytoplasm; 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVl 9a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 19C.
In a BLAST search of public sequence databases, the NOVl 9a protein was found to have homology to the proteins shown in the BLASTP data in Table 19D.
PFam analysis predicts that the NOVl 9a protein contains the domains shown in the Table 19E.
Example 20.
The NOV20 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 20A.
Table 20A. NOV20 S __equence Analysis
Further analysis of the NOV20a protein yielded the following properties shown in Table 20B.
Table 20B. Protein Sequence Properties NOV20a
PSort 0.3000 probability located in nucleus; 0.1000 probability located in analysis: mitochondrial matrix space; 0.1000 probability located in lysosome (lumen); 0.0000 probability located in endoplasmic reticulum (membrane)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV20a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 20C.
In a BLAST search of public sequence databases, the NOV20a protein was found to have homology to the proteins shown in the BLASTP data in Table 20D.
PFam analysis predicts that the NOV20a protein contains the domains shown in the Table 20E.
Table 20E. Domain Analysis of NOV20a
Identities/
Pfam Domain NOV20a Match Region Similarities Expect Value for the Matched Region
No Significant Matches Found
Example 21.
The NOV21 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 21 A.
Table 21A. NOV21 Sequence Analysis
SEQ TD NO: 75 7741 bp
NOV21a, TTGTCTCTTTGTGTTTTCCAGACATTCTAAGTGAGACTGTCCACATCATCTAGGAAAA
TGGTGGCCCTGTCCTTAAAGATTTGTGTGCGCCACTGCAACGTGGTGAAGACCATGCA
CG57804-01 DNA Sequence GTTTGAACCATCTACAGCTGTGTACGATGCGTGTCGAGTCATTCGGGAACGGGTGCCT GAGGCACAAACTGGGCAAGCTTCTGACTATGGACTCTTTCTTTCGGATGAAGACCCGA GGAAAGGGATTTGGCTGGAAGCGGGCAGAACACTGGATTACTACATGTTGCGGAATGG GGATATTTTGGAATATAAAAAGAAACAGAGACCTCAGAAAATCCGGATGCTGGATGGA TCTGTGAAGACAGTGATGGTGGATGATTCCAAGACTGTGGGGGAGCTCCTGGTCACTA TTTGTAGCAGAATAGGAATAACAAATTATGAAGAATACTCCTTAATCCAAGAAACTAT TGAAGAAAAGAAAGAGGAAGGAACGGGCACACTCAAAAAAGACAGGACACTGTTACGA GATGAGAGGAAAATGGAGAAGTTGAAGGCCAAGCTGCACACAGATGATGACCTAAATT GGCTGGATCACAGCCGAACATTCAGAGAACAAGGAGTAGATGAAAACGAAACGTTGCT GCTTAGACGGAAGTTCTTTTACTCTGATCAGAATGTAGATTCGAGAGACCCCGTGCAG CTGAACTTGCTTTATGTTCAGGCACGGGATGACATCCTGAATGGCTCTCACCCTGTCT CCTTCGAGAAAGCTTGTGAGTTTGGTGGATTTCAAGCCCAGATACAATTTGGACCTCA TGTGGAACATAAACACAAACCTGGATTTTTAGATCTGAAGGAATTCCTGCCCAAAGAA TATATCAAGCAGAGAGGAGCTGAAAAGAGGATCTTTCAGGAGCATAAGAACTGCGGAG AGATGAGTGAGATAGAAGCCAAGGTCAAGTACGTCAAACTCGCACGGTCCCTCCGCAC ATATGGCGTGTCCTTCTTCCTGGTGAAGGAGAAGATGAAAGGCAAGAACAAGCTGGTG CCTCGCCTGCTGGGGATCACCAAAGACTCGGTGATGCGCGTGGATGAGAAGACCAAGG AAGTGCTGCAGGAGTGGCCCCTCACCACCGTCAAGCGCTGGGCAGCCTCACCCAAGAG CTTCACACTGGATTTTGGGGAGTATCAGGAAAGCTACTATTCAGTACAAACCACCGAG GGAGAGCAGATATCCCAGCTGATTGCAGGCTACATTGACATCATCCTGAAAAAGGGAA CATACGTGACATCTGTGGGGTCTCCTCATTGCACTCCACATGGCTGGTGTTCTCTCAG TGACCAAACCACTTTTCCCGGCAGGTCCACCATCTTGCAGCAGCAGTTCAACCGGACC GGGAAGGCAGAGCACGGCTCAGTGGCGCTGCCGGCCGTGATGCGCTCGGGCTCCAGCG GGCCTGAGACCTTCAACGTTGGCAGCATGCCCTCGCCACAGCAGCAGGTCATGGTTGG GCAGATGCACCGAGGCCACATGCCGCCACTGACCTCAGCCCAGCAGGCCCTGATGGGG ACCATCAACACAAGCATGCACGCCGTCCAGCAGGCCCAGGATGATCTCAGTGAGCTCG ACTCGCTGCCACCTCTCGGCCAGGATATGGCATCTAGGGTATGGGTTCAGAACAAAGT CGACGAATCCAAACACGAAATCCATTCTCAAGTTGATGCTATCACGGCCGGAACGGCT TCAGTTGTTAACCTCACAGCTGGTGACCCTGCAGACACTGACTACACAGCTGTGGGAT GTGCGATCACCACTATTTCTTCCAACCTGACGGAGATGTCCAAGGGTGTGAAGCTATT GGCCGCCCTCATGGATGATGAGGTGGGCAGCGGGGAGGACTTGCTCAGAGCTGCCAGG ACCCTCGCTGGGGCGGTGTCAGACTTGCTGAAAGCTGTGCAGCCTACTTCTGGAGAGC CTCGACAGACAGTTTTGACTGCTGCTGGCAGCATCGGACAAGCCAGTGGGGATCTTCT GAGACAGATTGGAGAGAATGAGACTGATGAGCGATTCCAGGATGTTTTAATGAGTTTG GCCAAAGCTGTTGCCAATGCAGCTGCCATGTTGGTACTAAAGGCAAAGAATGTTGCCC AAGTGGCCGAAGACACTGTCCTACAGAACAGGGTAATTGCTGCTGCCACCCAGTGTGC CCTCTCCACCTCCCAGCTTGTGGCATGTGCCAAGGTTGTGAGCCCCACTATTAGCTCC CCTGTGTGCCAGGAGCAGCTGATTGAAGCAGGGAAGCTGGTGGACCGCTCGGTGGAGA ACTGTGTCCGTGCCTGCCAGGCGGCCACTACCGATAGTGAGCTCCTGAAGCAGGTCAG CGCAGCGGCCAGCGTGGTCAGCCAGGCCCTCCATGATCTCCTGCAGCATGTGCGGCAG TTTGCCAGCCGAGGCGAGCCCATCGGCCGCTACGACCAGGCTACTGACACCATCATGT GTGTCACCGAGAGCATCTTCAGCTCCATGGGTGACGCTGGTGAAATGGTGCGCCAGGC GCGGGTTCTGGCCCAAGCCACATCAGACCTCGTCAATGCCATGAGGTCAGATGCAGAA GCCGAAATCGACATGGAGAATTCAAAGAAGCTCCTGGCAGCAGCAAAACTCTTAGCTG ACTCCACTGCTCGCATGGTGGAAGCTGCAAAGGGGGCTGCAGCCAACCCAGAGAATGA GGACCAGCAGCAAAGGCTGAGAGAAGCTGCAGAAGGCCTCCGGGTAGCAACCAACGCA GCTGCCCAGAATGCTATTAAGAAAAAAATTGTCAACCGACTGGAGGTTGCAGCCAAGC AGGCCGCAGCGGCAGCCACACAGACCATCGCCGCCTCCCAGAATGCAGCTGTTTCCAA CAAGAACCCTGCGGCCCAGCAGCAGCTGGTCCAGAGTTGCAAGGCAGTGGCTGATCAC ATCCCTCAGCTGGTCCAGGGAGTGAGGGGGAGCCAAGCTCAAGCTGAAGACCTGAGTG CCCAGCTGGCTCTCATCATCTCCAGCCAGAACTTCCTCCAGCCTGGAAGCAAGATGGT GTCCTCTGCCAAAGCCGCAGTGCCCACCGTGAGTGACCAGGCCGCAGCCATGCAGCTG AGCCAGTGTGCCAAGAACCTGGCCACCAGCTTGGCGGAGCTGCGTACCGCCTCGCAGA AGGCCCATGAAGCTTGTGGTCCGATGGAAATCGATTCAGCTCTGAATACGGTGCAGAC GCTTAAGAATGAACTGCAGGATGCCAAGATGGCAGCCGTGGAGAGCCAGCTGAAGCCA CTTCCAGGGGAAACGCTGGAAAAATGTGCTCAGGACCTGGGAAGCACATCCAAGGCGG TGGGCTCCTCCATGGCACAGCTGCTGACCTGTGCTGCTCAAGGCAACGAACACTACAC AGGGGTGGCTGCTAGAGAGACGGCCCAAGCTCTGAAAACACTGGCCCAGGCCGCCCGT GGAGTGGCTGCATCGACAACCGACCCCGCGGCCGCCCATGCCATGTTAGATTCTGCTC GAGACGTGATGGAGGGCTCCGCCATGCTCATTCAAGAGGCCAAGCAGGCCCTGATTGC ACCTGGAGATGCAGAGCGTCAACAAAGACTGGCTCAGGTGGCTAAAGCCGTCTCACAC TCCTTGAATAACTGCGTAAATTGCCTCCCTGGGCAGAAGGATGTGGACGTGGCCTTGA
AGAGCATCGGGGAGTCCAGCAAGAAGCTGCTTGTGGATTCGCTACCTCCAAGCACGAA GCCTTTCCAGGAAGCCCAGAGTGAACTGAACCAGGCAGCAGCTGATCTGAACCAGTCT GCTGGGGAAGTGGTCCATGCCACCCGGGGCCAGAGTGGAGAGTTGGCTGCAGCCTCTG GAAAGTTCAGTGATGATTTTGGTGAATTCCTCGATGCTGGCATTGAGATGGCTGGCCA AGCTCAGACAAAAGAAGACCAGATCCAAGTGATAGGGAACCTCAAGAATATCTCGATG GCATCCAGCAAGCTGCTGTTAGCTGCCAAGTCTCTCTCTGTAGATCCAGGAGCTCCCA ATGCGAAAAATCTCCTGGCTGCAGCTGCAAGAGCTGTGACAGAGAGCATCAATCAACT CATCACTCTGTGTACCCAACAAGCTCCGGGCCAGAAAGAGTGCGATAATGCCCTGCGG GAGCTCGAGACTGTGAAGGGGATGTTGGACAATCCTAATGAACCTGTTAGTGACCTCT CTTACTTTGACTGCATTGAGAGTGTGATGGAAAACTCCAAGGTTCTGGGTGAATCGAT GGCAGGGATTTCACAGAATGCCAAGACCGGAGACCTCCCTGCCTTTGGGGAATGTGTG GGGATTGCATCCAAGGCTCTCTGTGGGCTGACAGAGGCTGCAGCCCAGGCTGCATACT TGGTTGGCATCTCTGATCCAAACAGCCAGGCAGGCCACCAGGGCCTGGTGGACCCCAT CCAGTTTGCCAGGGCTAACCAGGCCATCCAGATGGCATGCCAGAACTTGGTGGACCCT GGCAGCAGCCCATCACAGGTCCTGTCAGCCGCCACAATTGTTGCCAAGCACACGTCAG CCTTGTGCAATGCCTGCCGCATCGCCTCATCCAAGACGGCCAACCCAGTAGCCAAGAG GCACTTCGTCCAGTCAGCCAAGGAAGTCGCCAACAGCACTGCCAACCTGGTGAAGACC ATCAAGGCCCTGGATGGGGATTTCTCTGAAGACAACCGCAATAAGTGTCGCATCGCCA CCGCACCCTTGATTGAAGCTGTGGAGAACCTGACAGCGTTCGCCTCAAACCCTGAGTT TGTCAGCATTCCTGCCCAGATCAGCTCCGAGGGTTCCCAGGCACAGGAACCAATCCTG GTCTCAGCCAAGACCATGCTGGAGAGTTCATCGTACCTCATTCGCACTGCACGCTCTC TGGCCATCAACCCCAAAGACCCACCCACCTGGTCTGTACTGGCTGGACATTCCCATAC AGTGTCCGACTCCATCAAGAGTCTCATCACTTCTATCAGGGACAAGGCCCCTGGACAG AGGGAGTGTGATTACTCCATCGATGGCATCAACCGGTGCATCCGGGACATCGAGCAGG CCTCGCTGGCCGCCGTCAGCCAGAGCCTGGCCACGAGGGACGACATCTCTGTGGAGGC CCTGCAGGAGCAGCTGACTTCGGTGGTCCAGGAAATCGGACACCTTATCGATCCCATC GCCACAGCGGCTCGGGGAGAAGCAGCTCAGCTGGGACATAAGGTGACACAACTGGCAA GCTATTTTGAGCCCTTGATCTTAGCCGCAGTTGGTGTGGCCTCCAAGATTCTTGATCA TCAGCAGCAGATGACGGTGCTGGACCAGACCAAGACTCTCGCAGAGTCTGCCTTGCAG ATGTTGTATGCAGCCAAAGAAGGTGGCGGAAACCCCAAGGCACAACACACCCATGACG CCATCACAGAGGCCGCCCAGTTGATGAAGGAAGCCGTGGATGACATCATGGTGACGCT GAACGAAGCTGCCAGTGAAGTGGGGCTGGTTGGGGGCATGGTGGACGCCATTGCAGAA GCCATGAGCAAGCTGGATGAAGGCACTCCTCCAGAACCAAAGGGAACATTTGTCGACT ATCAGACGACTGTGGTTAAATACTCCAAAGCCATTGCGGTGACAGCTCAGGAAATGAT GACTAAGTCGGTTACTAACCCGGAGGAGTTGGGAGGACTGGCTTCACAAATGACCAGT GACTATGGGCACCTGGCTTTCCAGGGCCAGATGGCAGCAGCCACGGCGGAACCAGAGG AGATCGGATTCCAGATTCGCACTCGTGTGCAGGACCTGGGCCACGGCTGTATCTTCCT GGTGCAGAAGGCAGGGGCCCTCCAGGTCTGCCCCACAGACAGCTACACCAAGAGGGAG CTGATCGAATGCGCCCGTGCCGTCACGGAAAAGGTCTCCTTGGTGCTCTCGGCTCTCC AGGCCGGGAACAAAGGAACCCAGGCATGCATTACAGCCGCCACCGCTGTGTCTGGGAT CATTGCCGACCTGGACACCACCATTATGTTTGCAACAGCGGGGACGCTGAATGCAGAG AACAGTGAGACCTTCGCAGACCACAGGGAGAACATTCTCAAGACGGCCAAGGCCTTGG TAGAAGACACGAAACTACTTGTGTCAGGAGCTGCGTCCACTCCTGACAAGCTGGCCCA GGCGGCCCAGTCCTCAGCAGCCACCATCACCCAGCTCGCAGAAGTGGTCAAGCTGGGG GCAGCCAGCCTGGGCTCCGACGACCCCGAGACCCAGGTGGATTTGATCAATGCCATCA AAGATGTGGCCAAGGCCCTTTCTGATCTCATCAGTGCTACCAAGGGAGCTGCCAGCAA GCCAGTGGACGACCCTTCCATGTACCAGCTCAAGGGGGCTGCCAAGGTGATGGTGACC AATGTCACCTCGCTCCTCAAGACTGTAAAGGCAGTGGAGGATGAGGCCACCCGGGGCA CCAGGGCGCTTGAGGCCACAATTGAATGCATAAAGCAGGAGCTTACGGTGTTCCAGTC AAAAGACGTACCTGAAAAGACATCATCACCTGAAGAATCCATAAGGATGACGAAAGGC ATCACCATGGCAACAGCCAAAGCCGTGGCAGCTGGGAACTCATGTAGACAGGAGGACG TGATTGCTACTGCCAACCTGAGCCGGAAAGCCGTGTCAGATATGTTGACGGCTTGCAA GCAAGCATCCTTCCACCCCGATGTCAGTGACGAGGTGAGAACCAGAGCCTTGCGTTTC GGGACGGAGTGCACCCTTGGCTACTTGGACCTCCTGGAGCACGTCTTGGTGATTCTTC AGAAACCAACCCCAGAATTCAAGCAGCAGCTGGCCGCTTTCTCCAAGCGAGTCGCCGG CGCTGTGACAGAGCTCATCCAGGCGGCGGAAGCCATGAAAGGAACAGAGTGGGTGGAT CCAGAAGACCCAACTGTCATTGCAGAAACAGAGTTACTGGGGGCTGCAGCATCCATCG AAGCTGCTGCTAAGAAGTTAGAGCAACTGAAGCCAAGAGCAAAACCAAAACAAGCGGA TGAGACCCTGGACTTTGAGGAACAGATCTTGGAAGCTGCTAAATCCATTGCTGCTGCC ACAAGCGCCCTGGTCAAATCGGCCTCAGCAGCCCAGAGGGAGCTGGTGGCCCAAGGAA AGGTGGGCTCCATCCCTGCCAATGCTGCAGACGACGGACAGTGGTCACAGGGGCTGAT TTCTGCTGCCCGGATGGTGGCGGCTGCGACCAGCAGTCTCTGTGAGGCGGCCAATGCC TCCGTTCAGGGACACGCCAGCGAGGAGAAGCTCATCTCATCTGCCAAGCAGGTCGCCG CTTCCACGGCTCAGCTGCTGGTGGCCTGCAAGGTGAAGGCCGACCAGGATTCAGAGGC CATGAGGCGGCTACAGGCGGCAGGAAATGCTGTGAAAAGAGCCTCAGACAATCTTGTC CGTGCAGCCCAGAAGGCAGCTTTTGGCAAAGCTGATGACGACGATGTTGTAGTGGAAA CCAAGTTTGTGGGGGGCATTGCTCAGATCATCGCCGCCCAGGAAGAAATGCTAAAGAA AGAGCGAGAACTGGAAGAAGCAAGGAAAAAACTGGCCCAAATCCGCCAGCAGCAGTAT AAGTTTTTACCCACCGAGCTGAGGGAAGATGAGGGCTAAAGGTGCGAGCCCAGATGGC GAGCCCCAGGGGATGGCCCTGGCTGAA
ORF Start: ATG at 58 ORF Stop: TAA at 7693
SEQ ED NO: 76 2545 aa MW at 271692.8kD
NOV21a, MVALSLKICVRHCNWKTMQFEPSTAVYDACRVIRERVPEAQTGQASDYGLFLSDEDP RKGI LEAGRTLDYYMLRNGDILEYKKKQRPQKIRMLDGSVKTVMVDDSKTVGELLVT
CG57804-01 Protein Sequence ICSRIGITNYEEYSLIQETIEEKKEEGTGTLKKDRTLLRDERKMEKLKAKLHTDDDLN WLDHSRTFREQGVDENETLLLRRKFFYSDQNVDSRDPVQLNLLYVQARDDILNGSHPV SFEKACEFGGFQAQIQFGPHVEHKHKPGFLDLKEFLPKEYIKQRGAEKRIFQEHKNCG EMSEIEAKVKYVKLARSLRTYGVSFFLVKEKMKGKNKLVPRLLGITKDSVMRVDEKTK EVLQEWPLTTVKR AASPKSFTLDFGEYQESYYSVQTTEGEQISQLIAGYIDIILKKG TYVTSVGSPHCTPHG CSLSDQTTFPGRSTILQQQFNRTGKAEHGSVALPAVMRSGSS GPETFNVGSMPSPQQQVMVGQMHRGHMPPLTSAQQALMGTINTSMHAVQQAQDDLSEL DSLPPLGQDMASRVWVQNKVDESKHEIHSQVDAITAGTASWNLTAGDPADTDYTAVG CAITTISSNLTEMSKGVKLLAALMDDEVGSGEDLLRAARTLAGAVSDLLKAVQPTSGE PRQTVLTAAGSIGQASGDLLRQIGENETDERFQDVLMSLAKAVANAAAMLVLKAKNVA QVAEDTVLQNRVIAAATQCALSTSQLVACAKWSPTISSPVCQEQLIEAGKLVDRSVE NCVRACQAATTDSELLKQVSAAASWSQALHDLLQHVRQFASRGEPIGRYDQATDTIM CVTESIFSSMGDAGEMVRQARVLAQATSDLVNAMRSDAEAEIDMENSKKLLAAAKLLA DSTARMVEAAKGAAANPENEDQQQRLREAAEGLRVATNAAAQNAIKKKIVNRLEVAAK QAAAAATQTIAASQNAAVSNKNPAAQQQLVQSCKAVADHIPQLVQGVRGSQAQAEDLS AQLALIISSQNFLQPGSKMVSSAKAAVPTVSDQAAAMQLSQCAKNLATSLAELRTASQ KAHEACGPMEIDSALNTVQTLKNELQDAKMAAVESQLKPLPGETLEKCAQDLGSTSKA VGSSMAQLLTCAAQGNEHYTGVAARETAQALKTLAQAARGVAASTTDPAAAHAMLDSA RDVMEGSAMLIQEAKQALIAPGDAERQQRLAQVAKAVSHSLNNCVNCLPGQKDVDVAL KSIGESSKKLLVDSLPPSTKPFQEAQSELNQAAADLNQSAGEWHATRGQSGELAAAS GKFSDDFGEFLDAGIEMAGQAQTKEDQIQVIGNLKNISMASSKLLLAAKSLSVDPGAP NAKNLLAAAARAVTESINQLITLCTQQAPGQKECDNALRELETVKGMLDNPNEPVSDL SYFDCIESVMENSKVLGESMAGISQNAKTGDLPAFGECVGIASKALCGLTEAAAQAAY LVGISDPNSQAGHQGLVDPIQFARANQAIQMACQNLVDPGSSPSQVLSAATIVAKHTS ALCNACRIASSKTANPVAKRHFVQSAKEVANSTANLVKTIKALDGDFSEDNRNKCRIA TAPLIEAVENLTAFASNPEFVSIPAQISSEGSQAQEPILVSAKTMLESSSYLIRTARS LAINPKDPPTWSVLAGHSHTVSDSIKSLITSIRDKAPGQRECDYSIDGINRCIRDIEQ ASLAAVSQSLATRDDISVEALQEQLTSWQEIGHLIDPIATAARGEAAQLGHKVTQLA SYFEPLILAAVGVASKILDHQQQMTVLDQTKTLAESALQMLYAAKEGGGNPKAQHTHD AITEAAQLMKEAVDDIMVTLNEAASEVGLVGGMVDAIAEAMSKLDEGTPPEPKGTFVD YQTTWKYSKAIAVTAQEMMTKSVTNPEELGGLASQMTSDYGHLAFQGQMAAATAEPE EIGFQIRTRVQDLGHGCIFLVQKAGALQVCPTDSYTKRELIECARAVTEKVSLVLSAL QAGNKGTQACITAATAVSGIIADLDTTIMFATAGTLNAENSETFADHRENILKTAKAL VEDTKLLVSGAASTPDKLAQAAQSSAATITQLAEWKLGAASLGSDDPETQVDLINAI KDVAKALSDLISATKGAASKPVDDPSMYQLKGAAKVMVTNVTSLLKTVKAVEDEATRG TRALEATIECIKQELTVFQSKDVPEKTSSPEESIRMTKGITMATAKAVAAGNSCRQED VIATANLSRKAVSDMLTACKQASFHPDVSDEVRTRALRFGTECTLGYLDLLEHVLVIL QKPTPEFKQQLAAFSKRVAGAVTELIQAAEAMKGTE VDPEDPTVIAETELLGAAASI EAAAKKLEQLKPRAKPKQADETLDFEEQILEAAKSIAAATSALVKSASAAQRELVAQG KVGSIPANAADDGQ SQGLISAARMVAAATSSLCEAANASVQGHASEEKLISSAKQVA ASTAQLLVACKVKADQDSEAMRRLQAAGNAVKRASDNLVRAAQKAAFGKADDDDVWE TKFVGGIAQIIAAQEEMLKKERELEEARKKLAQIRQQQYKFLPTELREDEG
Further analysis of the NOV21a protein yielded the following properties shown in Table 21B.
Table 21B. Protein Sequence Properties NOV21a
PSort 0.5964 probability located in mitochondrial matrix space; 0.3037 probability analysis: located in mitochondrial inner membrane; 0.3037 probability located in mitochondrial intermembrane space; 0.3037 probability located in mitochondrial outer membrane
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV2 la protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 21C.
In a BLAST search of public sequence databases, the NOV2 la protein was found to have homology to the proteins shown in the BLASTP data in Table 2 ID.
PFam analysis predicts that the NOV21a protein contains the domains shown in the Table 2 IE.
Example 22.
The NOV22 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 22A.
Table 22A. NOV22 Sequence Analysis
SEQ ED NO: 77 2214 bp
NOV22a, ATTCCTCCCTGCCCCTCGTGCAGCCGCTGCCATGGCCCAGACACTGCAGATGGAGATC
CCGAACTTCGGCAACAGCATCCTGGAGTGCCTCAATGAACAGCGGCTGCAGGGCCTGT
CG57551-01 DNA Sequence ACTGTGACGTGTCAGTGGTGGTCAAGGGCCATGCCTTCAAGGCCCACCGGGCCGTGCT TGCTGCCAGCAGCTCCTACTTCCGGGACCTGTTCAACAACAGCCGCAGCGCCGTGGTG GAGCTGCCGGCGGCTGTGCAGCCCCAGTCTTTCCAGCAGATCCTCAGCTTCTGCTACA CGGGCCGGCTGAGCATGAACGTGGGCGACCAGTTCCTGCTCATGTACACGGCTGGCTT CCTGCAGATCCAGGAGATCATGGAGAAGGGCACCGAGTTCTTCCTCAAGGTGAGCTCC CCGAGCTGCGACTCCCAGGGCCTGCATGCGGAGGAGGCCCCATCGTCGGAGCCCCAGA GCCCCGTGGCGCAGACATCGGGCTGGCCAGCCTGTAGCACCCCGCTGCCCCTCGTGTC GCGGGTGAAGACGGAGCAGCAGGAGTCGGACTCCGTGCAGTGCATGCCCGTGGCCAAG CGGCTGTGGGACAGTGGCCAGAAGGAGGCTGGGGGCGGCGGCAATGGCAGCCGCAAGA TGGCCAAGTTCTCCACGCCGGACCTGGCTGCCAACCGGCCTCACCAGCCCCCGCCACC CCAACAGGCTCCGGTGGTGGCAGCAGCCCAGCCCGCCGTGGCTGCGGGAGCAGGGCAG CCAGCCGGTGGGGTGGCAGCAGCAGGGGGTGTGGTGAGTGGGCCCAGCACGTCGGAGC GGACCAGCCCAGGCACCTCAAGCGCCTACACCAGCGACAGCCCTGGCTCCTACCACAA TGAGGAGGACGAGGAGGAGGATGGTGGCGAGGAGGGCATGGATGAGCAGTACCGGCAG ATCTGCAACATGTACACCATGTACAGCATGATGAACGTCGGCCAGACAGCCGAGAAGG TGGAGGCCCTCCCGGAGCAGGTAGCCCCCGAGTCCCGAAATCGCATCCGGGTTCGGCA AGACCTGGCGTCTCTCCCGGCTGAACTTATCAACCAGATTGGGAACCGCTGCCACCCC AAGCTCTACGACGAGGGCGACCCCTCTGAGAAGCTGGAGCTGGTGACAGGCACCAACG TGTACATCACAAGGGCGCAGCTGATGAACTGCCACGTCAGCGCAGGCACGCGGCACAA GGTCCTACTGCGGCGGCTCCTGGCCTCCTTCTTTGACCGGAACACGCTGGCCAACAGC TGCGGCACCGGCATCCGCTCTTCTACCAACGATCCCCGTCGGAAGCCCCTGGACAGCC GCGTGCTCCACGCTGTCAAGTACTACTGCCAGAACTTCGCCCCCAACTTCAAGGAGAG CGAGATGAATGCCATCGCGGCCGACATGTGCACCAACGCCCGCCGCGTCGTGCGCAAG AGCTGGATGCCCAAGGTCAAGGTGCTCAAGGCTGAGGATGACGCCTACACCACCTTCA TCAGTGAAACGGGCAAGATCGAGCCGGACATGATGGGTGTGGAGCATGGCTTCGAGAC CGCCAGCCACGAGGGCGAGGCGGGTCCCATCGCTGAAGCCCTGCAGTAACCCGCCCAG
CCTCCCGCGGGGCCGCACACTTCCCCTCCCAACACACACACACACCTGCCATCTTGGT
CATGAGCTACTGTCTGTCCCTCCCCAGGACCCGCGGTGGGTGCTGCATGTTCCCGGCC
CTCTGCCCCTCCTGTCCTACCCCCTTTCCCCACCGAGAGCTGGGCCGGGAGAGGACCG
CAGGGCAGGTGGCGTGAGGTCCGTGTTGCCTTCTTTAACACACACTCGTGCAGTGGGG
GAGTTCTGGCTCCCCAACCTAACCCCTAGCCGTCATCTCCACACTCACCAGGCCCACC
AGGGGAGGGGGCTGGCCTGGGGGTCTTGGGAAGGCCCCTCCCCAGGCCTTAGGCCACC
TCGCGGAAGCCTTCAGCCTCCGCCCCTCACTGCAGCCCCTTGGGACTTGAGGGGGGCC
CCAGGGGTTCTCAGGACCCCTCCCACCACCTCCCAGTGCTTCCACGTCTCCAAAAGCG
CCTTCCTGTCACCCTCGTCTATCCCTGCGCCTGGGGGCTGGGGTAGGCGAGGCCGTGG
GGACTACCCATTTTATAGCTGGGGAAACAGGCTCCGAGAAATTGCACAACCGACCTCA
GGTGGCCGGC
ORF Start: ATG at 32[θRF Stop: TAA at 1613
SEQ ED NO: 78 527 aa MW at 57283.8kD
NOV22a, MAQTLQMEIPNFGNSILECLNEQRLQGLYCDVSVWKGHAFKAHRAVLAASSSYFRDL FNNSRSAWELPAAVQPQSFQQILSFCYTGRLSMNVGDQFLLMYTAGFLQIQEIMEKG
CG57551-01 Protein Sequence TEFFLKVSSPSCDSQGLHAEEAPSSEPQSPVAQTSGWPACSTPLPLVSRVKTEQQESD SVQCMPVAKRL DSGQKEAGGGGNGSRKMAKFSTPDLAANRPHQPPPPQQAPWAAAQ PAVAAGAGQPAGGVAAAGGWSGPSTSERTSPGTSSAYTSDSPGSYHNEEDEEEDGGE EGMDEQYRQICNMYTMYSMMNVGQTAEKVEALPEQVAPESRNRIRVRQDLASLPAELI NQIGNRCHPKLYDEGDPSEKLELVTGTNVYITRAQLMNCHVSAGTRHKVLLRRLLASF FDRNTLANSCGTGIRSSTNDPRRKPLDSRVLHAVKYYCQNFAPNFKESEMNAIAADMC
Further analysis of the NOV22a protein yielded the following properties shown in Table 22B.
Table 22B. Protein Sequence Properties NOV22a
PSort analysis: 0.6000 probability located in nucleus; 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen); 0.0000 probability located in endoplasmic reticulum (membrane)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV22a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 22C.
In a BLAST search of public sequence databases, the NOV22a protein was found to have homology to the proteins shown in the BLASTP data in Table 22D.
PFam analysis predicts that the NOV22a protein contains the domains shown in the Table 22E.
Example 23.
The NOV23 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 23A.
Further analysis of the NOV23a protein yielded the following properties shown in Table 23B.
Table 23B. Protein Sequence Properties NOV23a
PSort 0.6500 probability located in cytoplasm; 0.2271 probability located in lysosome analysis: (lumen); 0.1000 probability located in mitochondrial matrix space; 0.0000 probability located in endoplasmic reticulum (membrane)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV23a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 23C.
In a BLAST search of public sequence databases, the NOV23a protein was found to have homology to the proteins shown in the BLASTP data in Table 23D.
PFam analysis predicts that the NOV23a protein contains the domains shown in the Table 23E.
Example 24.
The NOV24 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 24A.
CTGGCAGACATCCTCCGGGAATTCAACCCTTCCCTGAAGGGCTTCTCTGTTGGCACTG GGAAAGAAACCAGTCCTAATGCCTTCTTAAACCAGGCTGTGGCAGGAGGCCGAGCTGA GCAGGCCAGGAGGCTGGTGGACCTGATGAAGAATGACACGAGGATACACTTTCAGGAA GACTGGAAGATAATAACCCTGTTTATAGGCGGCAATGACCTCTGTGATTTCTGCAATG ATCTGGTACACTATTCTCCCCAGAACTTCACAGACAACATTGGAAAGGCCCTGGACAT CCTCCATGCTGAGTCTCAGGTTCCTCGGGCATTTGTGAACCTGGTGACGGTGCTTGAG ATCGTCAACCTGAGGGAGCTGTACCAGGAGAAAAAAGTCTACTGCCCAAGGATGATCC TCAGGTCACTGTGTCCCTGTGTCCTGAAGTTTGATGATAACTCAACAGAACTTGCTAC CCTCATCGAATTCAACAAGAAGTTTCAGGAGAAGACCCACCAACTGATTGAGAGTGGG CGATATGACACAAGGGAAGATTTTACTGTGGTTGTGCAGCCGTTCTTTGAAAACGTGG ACATGCCAAAGACCCAGGAAGGATTGCCTGACAACTCTTTCTTCGCTCCTGACTGTTT CCACTTCAGCAGCAAGTCTCACTCCCGAGCAGCCAGTGCTCTCTGGAACAATATGCTG GAGCCTGTTGGCCAGAAGACGACTCGTCATAAGTTTGAAAACAAGATCAATATCACAT GTCCGTCACAGGTCCAGCCGTTTCTGAGGACCTACAAGAACAGCATGCAGGGTCATGG GACCTGGCTGCCATGCAGGGACAGAGCCCCTTCTGCCTTGCACCCTACCTCAGTGCAT GCCCTGAGACCTGCAGACATCCAAGTTGTGGCTGCTCTGGGGGATTCTCTGACCGCTG GCAATGGAATTGGCTCCAAACCAGACGACCTCCCCGATGTCACCACACAGTATCGGGG ACTGTCATACAGTGCAGGAGGGGACGGCTCCCTGGAGAATGTGACCACCTTACCTAGT TCTATCCTTCGGGAGTTTAACAGAAACCTCACAGGCTACGCCGTGGGCACGGGTGATG CCAATGACACGAATGCATTCCTCAATCAAGCTGTTCCCGGAGCAAAGGCTAGGGATCT TATGAGCCAAGTCCAAACTCTGATGCAGAAGATGAAAGATGATCATAGAGTAAATTTC CATGAAGACTGGAAGGTCATCACAGTGCTGATCGGAGGCAGCGATTTATGTGACTACT GCACAGATTCGAATCTGTATTCTGCAGCCAACTTTGTTCACCATCTCCGCAATGCCTT GGACGTCCTGCATAGAGAGGTGCCCAGAGTCCTGGTCAACCTCGTGGACTTCCTGAAC CCCACTATCATGCGGCAGGTGTTCCTGGGAAACCCAGACAAGTGCCCAGTGCAGCAGG CCAGCGTTTTGTGTAACTGCGTTCTGACCCTGCGGGAGAACTCCCAAGAGCTAGCCAG GCTGGAGGCCTTCAGCCGAGCCTACCAGAGCAGCATGCGCGAGCTGGTGGGGTCAGGC CGCTATGACACGCAGGAGGACTTCTCTGTGGTGCTGCAGCCCTTCTTCCAGAACATCC AGCTCCCTGTCCTGCAGGATGGGCTCCCAGATACGTCCTTCTTTGCCCCAGACTGCAT CCACCCAAATCAGAAATTCCACTCCCAGCTGGCCAGAGCCCTTTGGACCAATATGCTT GAACCACTTGGAAGCAAAACAGAGACCCTGGACCTGAGAGCAGAGATGCCCATCACCT GTCCCACTCAGAATGAGCCCTTCCTGAGAACCCCTCGGAATAGTAACTACACGTACCC CATCAAGCCAGCCATTGAGAACTGGGGCAGTGACTTCCTGTGTACAGAGTGGAAGGCT TCCAATAGTGTTCCAACCTCTGTCCACCAGCTCCGACCAGCAGACATCAAAGTGGTGG CCGCCCTGGGTGACTCTCTGACTACAGCAGTGGGAGCTCGACCAAACAACTCCAGTGA CCTACCCACATCTTGGAGGGGACTCTCTTGGAGCATTGGAGGGGATGGGAACTTGGAG ACTCACACCACACTGCCCAGTATTCTGAAGAAGTTCAACCCTTACCTCCTTGGCTTCT CTACCAGCACCTGGGAGGGGACAGCAGGACTAAATGTGGCAGCGGAAGGGGCCAGAGC TAGGAGGGACATGCCAGCCCAGGCCTGGGACCTGGTAGAGCGAATGAAAAACAGCCCC ATACACTTTCAGGAAGACTGGAAGATAATAACCCTGTTTATAGGCGGCAATGACCTCT GTGATTTCTGCAATGATCTGGTAGGTGAATATGTTCAGCACATCCAACAGGCCCTGGA CATCCTCTCTGAGGAGCTCCCAAGGGCTTTCGTCAACGTGGTGGAGGTCATGGAGCTG GCTAGCCTGTACCAGGGCCAAGGCGGGAAATGTGCCATGCTGGCAGCTCAGAACAACT GCACTTGCCTCAGACACTCGCAAAGCTCCCTGGAGAAGCAAGAACTGAAGAAAGTGAA CTGGAACCTCCAGCATGGCATCTCCAGTTTCTCCTACTGGCACCAATACACACAGCGT GAGGACTTTGCGGTTGTGGTGCAGCCTTTCTTCCAAAACACACTCACCCCACTGAACA GAGGGGACACTGACCTCACCTTCTTCTCCGAGGACTGTTTTCACTTCTCAGACCGCGG GCATGCCGAGATGGCCATCGCACTCTGGAACAACATGCTGGAACCAGTGGGCCGCAAG ACTACCTCCAACAACTTCACCCACAGCCGAGCCAAACTCAAGTGCCCCTCTCCTGTGA GTCCTTACCTCTACACCCTGCGGAACAGCCGATTGCTCCCAGACCAGGCTGAAGAAGC CCCCGAGGTGCTCTACTGGGCTGTCCCAGTGGCAGCGGGAGTCGGCCTTGTGGTGGGC ATCATCGGGACAGTGGTCTGGAGGTGCAGGAGAGGTGGCCGGAGGGAAGATCCTCCAA TGAGCCTGCGCACTGTGGCCCTCTAGGCCCGGGG
ORF Start: ATG at 1 ORF Stop: TAG at 4258
SEQ ID NO: 82 1419 aa MW at 158435. lkD
NOV24a, MT DTAL TSVFLIGLLPTLGFANCILQTSGKMCTLRGRYPQPPQPPLCLSPLVHQLR PADIKWAALGNDETFQESGAGQLSEPDPRQ S PQACLPGVKKEMQDWGERTPSRR
CG57399-01 Protein Sequence RSLRRREALVPAAGKESLCRQDIFISLLEIIKHFPPSPQDINLEKDWKLVTLFIGVND LCHYCPLVQGPVIDLGGMDTLHSLQLPRAFVNWEVMELASLYQGQGGKCAMLAAQEA NSLLASSRYSEQESFTWFQPFFYETTPSDPRLQDSTTLAWHL NRMMEPAGEKDEP LSVKHGRPMKCPSQESPYLFSYRNSNYLTRLQKPQDKLVREGAEIRCPDKDPSDTVPT SVHRLKPADI VIGALGDSLTAGNGAGSTPGNVLDVLTQYRGLS SVGGDENIGTVTT LADILREFNPSLKGFSVGTGKETSPNAFLNQAVAGGRAEQARRLVDLMKNDTRIHFQE D KIITLFIGGNDLCDFCNDLVHYSPQNFTDNIGKALDILHAESQVPRAFVNLVTVLE IVNLRELYQEKKVYCPRMILRSLCPCVLKFDDNSTELATLIEFNKKFQEKTHQLIESG RYDTREDFTVWQPFFENVDMPKTQEGLPDNSFFAPDCFHFSSKSHSRAASAL NNML EPVGQKTTRHKFENKINITCPSQVQPFLRTYKNSMQGHGT LPCRDRAPSALHPTSVH ALRPADIQWAALGDSLTAGNGIGSKPDDLPDVTTQYRGLSYSAGGDGSLENVTTLPS SILREFNRNLTGYAVGTGDANDTNAFLNQAVPGAKARDLMSQVQTLMQKMKDDHRVNF HED KVITVLIGGSDLCDYCTDSNLYSAANFVHHLRNALDVLHREVPRVLVNLVDFLN PTIMRQVFLGNPDKCPVQQASVLCNCVLTLRENSQELARLEAFSRAYQSSMRELVGSG RYDTQEDFSWLQPFFQNIQLPVLQDGLPDTSFFAPDCIHPNQKFHSQLARALWTNML EPLGSKTETLDLRAEMPITCPTQNEPFLRTPRNSNYTYPIKPAIENWGSDFLCTEWKA SNSVPTSVHQLRPADIKWAALGDSLTTAVGARPNNSSDLPTSWRGLSWSIGGDGNLE
THTTLPSILKKFNPYLLGFSTSTWEGTAGLNVAAEGARARRDMPAQAWDLVERMKNSP IHFQEDWKIITLFIGGNDLCDFCNDLVGEYVQHIQQALDILSEELPRAFVNWEVMEL ASLYQGQGGKCA LAAQNNCTCLRHSQSSLEKQELKKVNWNLQHGISSFSY HQYTQR EDFAVWQPFFQNTLTPLNRGDTDLTFFSEDCFHFSDRGHAEMAIALWNNMLEPVGRK TTSNNFTHSRAKLKCPSPVSPYLYTLRNSRLLPDQAEEAPEVLY AVPVAAGVGLWG I IGTWWRCRRGGRREDPPMSLRTVAL
SEQ ED NO: 83 1624 bp
NOV24b, GCCGGCTGACATCAATGTAATTGGAGCCCTGGGTGACTCTCTCACGGCAGGCAATGGG
GCCGGGTCCACACCTGGGAACGTCTTGGACGTCTTGACTCAGTACCGAGGCCTGTCCT
CG57399-02 DNA Sequence GGAGCGTCGGCGGAGATGAGAACATCGGCACCGTTACCACCCTGGCGAACATCCTCCG
GGAATTCAACCCTTCCCTGAAGGGCTTCTCTGTTGGCACTGGGAAAGAAACCAGTCCT
AATGCCTTCTTAAACCAGGCTGTGGCAGGAGGCCGAGCTGAGGATCTACCTGTCCAGG
CCAGGAGGCTGGTGGACCTGATGAAGAATGACACGAGGATACACTTTCAGGAAGACTG
GAAGATAATAACCCTGTTTATAGGCGGCAATGACCTCTGTGATTTCTGCAATGATCTG GTCCACTATTCTCCCCAGAACTTCACAGACAACATTGGAAAGGCCCTGGACATCCTCC ATGCTGAGGTTCCTCGGGCATTTGTGAACCTGGTGACGGTGCTTGAGATCGTCAACCT GAGGGAGCTGTACCAGGAGAAAAAAGTCTACTGCCCAAGGATGATCCTCAGGTCTCTG TGTCCCTGTGTCCTGAAGTTTGATGATAACTCAACAGAACTTGCTACCCTCATCGAAT TCAACAAGAAGTTTCAGGAGAAGACCCACCAACTGATTGAGAGTGGGCGATATGACAC AAGGGAAGATTTTACTGTGGTTGTGCAGCCGTTCTTTGAAAACGTGGACATGCCAAAG ACCTCGGAAGGATTGCCTGACAACTCTTTCTTCGCTCCTGACTGTTTCCACTTCAGCA GCAAGTCTCACTCCCGAGCAGCCAGTGCTCTCTGGAACAATATGCTGGAGCCTGTTGG CCAGAAGACGACTCGTCATAAGTTTGAAAACAAGATCAATATCACATGTCCGAACCAG GTCCAGCCGTTTCTGAGGACCTACAAGAACAGCATGCAGGGTCATGGGACCTGGCTGC CATGCAGGGACAGAGCCCCTTCTGCCTTGCACCCTACCTCAGTGCATGCCCTGAGACC TGCAGACATCCAAGTTGTGGCTGCTCTGGGGGATTCTCTGACCGCTGGCAATGGAATT GGCTCCAAACCAGACGACCTCCCCGATGTCACCACACAGTATCGGGGACTGTCATACA GAGAAAGTAAACCAGGGTTCTTATCAGACTCCTGGGTCAGCAAATCCAACAGGAAATG CACCAGAAAAGCACCAAATCCCTGAATCTTCACCTCCCCGCTTGCATGTATACGTGTA CACGTGGTGTTCCTACGTCTCTGTTTACTGTCTTTATGTGTTTATTCATGTTGTCTTG
TAGTCACACAGCTGCCTTTACATATATGTACACATCTGCACAGAAAACCTCTGAAACC
CATCGCACACTTCGAGAGGCCATAACCAAGACACAATCACAATCAGCCATGTCTTGAA
AGATTAGCAATTCGACAAGAGGAAAGGGTGAGAAAGGGCATCCCGAACACGGAAGTGG
AGAAGCTCAGGGTGTGTCAGGCGAGCGGTTGCGTGTAGATATTCTCAAGTTTCTTTCT
CTCCTAATAAAGTTCTCATTCCTGTAGGCTTCAAAGTAAGTGGCGAGTAGCTCAGAAT
ORF Start: ATG at 311 ORF Stop: TGA at 1241
SEQ ID NO: 84 310 aa MW at 35240.6kD
NOV24b, MKNDTRIHFQEDWKIITLFIGGNDLCDFCNDLVHYSPQNFTDNIGKALDILHAEVPRA FVNLVTVLEIVNLRELYQEKKVYCPRMILRSLCPCVLKFDDNSTELATLIEFNKKFQE
CG57399-02 Protein Sequence KTHQLIESGRYDTREDFTVWQPFFENVDMPKTSEGLPDNSFFAPDCFHFSSKSHSRA ASAL NNMLEPVGQKTTRHKFENKINITCPNQVQPFLRTYKNS QGHGTWLPCRDRAP SALHPTSVHALRPADIQWAALGDSLTAGNGIGSKPDDLPDVTTQYRGLSYRESKPGF LSDSWVSKSNRKCTRKAPNP
SEQ ED NO: 85 4425 bp
NOV24c, CTGGAGCATTCTGGCATGGGGCTGCGGCCAGGCATTTTCCTCCTGGAGCTGCTGCTGC
TTCTGGGGCAAGGTACCCCTCAGATCCATACCTCTCCTAGAAAGAGTACATTGGAAGG
CG57399-03 DNA Sequence GCAGCTATGGCCAGAGACAGTTCACTCTCTGAAGCCTTCTGATATTAAATTTGTGGCA GCCATTGGCAATCTGGAAATTGTGCCAGACCCAGGGACGGGCGATCTGGAGAAGCAAG ACGAAAGGCCACAGCAGGTGTGCATGGGAGTGATGACAGTCCTTTCAGACATCATCAG ATATTTCAGTCCTTCTGTTCCAATGCCTGTGTGCCACACTGGAAAGAGAGTCATACCC CACGATGGTGCTGAGGACTTGTGGATTCAGGCTCAAGAACTGGTGAGAAACATGAAAG AGAACCAACTTGACTTTCAATTTGACTGGAAGCTCATCAATGTGTTCTTCAGTAATGC AAGCCAGTGTTACCTGTGCCCCTCTGCTCAACAGAATGGGCTTGCGGCGGGCGGCGTG GATGAGCTGATGGGGGTGCTGGACTACCTGCAGCAGGAGGTGCCCAGAGCATTTGTAA ACCTGGTGGACCTCTCTGAGGTTGCAGAGGTCTCTCGTCAGTATCACGGCACTTGGCT CAGCCCTGCACCAGAGCCCTGTAATTGCTCAGAGGAGACCACCCGGCTGGCCAAGGTG GTGATGCAGTGGTCTTATCAGGAAGCCTGGAACAGCCTCCTGGCCTCCAGCAGGTACA GTGAGCAGGAGTCCTTCACCGTGGTTTTCCAGCCTTTCTTCTATGAGACCACCCCATC TGACCCCCGACTCCAGGATTCTACCACGCTGGCCTGGCATCTCTGGAATAGGATGATG GAGCCAGCAGGAGAGAAAGATGAGCCATTGAGTGTAAAACACGGGAGGCCAATGAAGT GTCCCTCTCAGGAGAGCCCCTATCTGTTCAGCTACAGAAACAGCAACTACCTGACCAG ACTGCAGAAACCCCAAGACAAGCTTGAGGTAAGAGAAGGAGCGGAAATCAGATGTCCT GACAAAGACCCCTCCGATACGGTTCCCACCTCAGTTCATAGGCTGAAGCCGGCTGACA TCAACGTAATTGGAGCCCTGGGTGACTCTCTCACGGCAGGCAATGGGGCCGGGTCCAC ACCTGGGAACGTCTTGGACGTCTTGACTCAGTACCGAGGCCTGTCCTGGAGCGTCGGC GGAGATGAGAACATCGGCACCGTTACCACCCTGGCGGACATCCTCCGGGAATTCAACC CTTCCCTGAAGGGCTTCTCTGTTGGCACTGGGAAAGAAACCAGTCCTAATGCCTTCTT AAACCAGGCTGTGGCAGGAGGCCGAGCTGAGCAGGCCAGGAGGCTGGTGGACCTGATG AAGAATGACACGAGGATACACTTTCAGGAAGACTGGAAGATAATAACCCTGTTTATAG GCGGCAATGACCTCTGTGATTTCTGCAATGATCTGGTACACTATTCTCCCCAGAACTT CACAGACAACATTGGAAAGGCCCTGGACATCCTCCATGCTGAGGTTCCTCGGGCATTT
GTGAACCTGGTGACGGTGCTTGAGATCGTCAACCTGAGGGAGCTGTACCAGGAGAAAA AAGTCTACTGCCCAAGGATGATCCTCAGGTCACTGTGTCCCTGTGTCCTGAAGTTTGA TGATAACTCAACAGAACTTGCTACCCTCATCGAATTCAACAAGAAGTTTCAGGAGAAG ACCCACCAACTGATTGAGAGTGGGCGATATGACACAAGGGAAGATTTTACTGTGGTTG TGCAGCCGTTCTTTGAAAACGTGGACATGCCAAAGACCCAGGAAGGATTGCCTGACAA CTCTTTCTTCGCTCCTGACTGTTTCCACTTCAGCAGCAAGTCTCACTCCCGAGCAGCC AGTGCTCTCTGGAACAATATGCTGGAGCCTGTTGGCCAGAAGACGACTCGTCATAAGT TTGAAAACAAGATCAATATCACATGTCCGAACCAGGTAGAGTGGCCGTTTCTGAGGAC CTACAAGAACAGCATGCAGGGTCATGGGACCTGGCTGCCATGCAGGGACAGAGCCCCT TCTGCCTTGCACCCTACCTCAGTGCATGCCCTGAGACCTGCAGACATCCAAGTTGTGG CTGCTCTGGGGGATTCTCTGACCGCTGGCAATGGAATTGGCTCCAAACCAGACGACCT CCCCGATGTCACCACACAGTATCGGGGACTGTCATACAGTGCAGGAGGGGACGGCTCC CTGGAGAATGTGACCACCTTACCTGATATCCTTCGGGAGTTTAACAGAAACCTCACAG GCTACGCCGTGGGCACGGGTGATGCCAATGACACGAATGCATTCCTCAATCAAGCTGT TCCCGGAGCAAAGGCTAGGGATCTTATGAGCCAAGTCCAAACTCTGATGCAGAAGATG AAAGATGATCATAGAGTAAATTTCCATGAAGACTGGAAGGTCATCACAGTGCTGATCG GAGGCAGCGATTTATGTGACTACTGCACAGATTCGAATCTGTATTCTGCAGCCAACTT TGTTCACCATCTCCGCAATGCCTTGGACGTCCTGCATAGAGAGGTGCCCAGAGTCCTG GTCAACCTCGTGGACTTCCTGAACCCCACTATCATGCGGCAGGTGTTCCTGGGAAACC CAGACAAGTGCCCAGTGCAGCAGGCCAGCGTTTTGTGTAACTGCGTTCTGACCCTGCG GGAGAACTCCCAAGAGCTAGCCAGGCTGGAGGCCTTCAGCCGAGCCTACCAGAGCAGC ATGCGCGAGCTGGTGGGGTCAGGCCGCTATGACACGCAGGAGGACTTCTCTGTGGTGC TGCAGCCCTTCTTCCAGAACATCCAGCTCCCTGTCCTGCAGGATGGGCTCCCAGATAC GTCCTTCTTTGCCCCAGACTGCATCCACCCAAATCAGAAATTCCACTCCCAGCTGGCC AGAGCCCTTTGGACCAATATGCTTGAACCACTTGGAAGCAAAACAGAGACCCTGGACC TGAGAGCAGAGATGCCCATCACCTGTCCCACTCAGAATGAGCCCTTCCTGAGAACCCC TCGGAATAGTAACTACACGTACCCCATCAAGCCAGCCATTGAGAACTGGGGCAGTGAC TTCCTGTGTACAGAGTGGAAGGCTTCCAATAGTGTTCCAACCTCTGTCCACCAGCTCC GACCAGCAGACATCAAAGTGGTGGCCGCCCTGGGTGACTCTCTGACTGTGGCAGTGGG AGCTCGACCAAACAACTCCAGTGACCTACCCACATCTTGGAGGGGACTCTCTTGGAGC ATTGGAGGGGATGGGAACTTGGAGACTCACACCACACTGCCCGACATTCTGAAGAAGT TCAACCCTTACCTCCTTGGCTTCTCTACCAGCACCTGGGAGGGGACAGCAGGACTAAA TGTGGCAGCGGAAGGGGCCAGAGCTAGGGACATGCCAGCCCAGGCCTGGGACCTGGTA GAGCGAATGAAAAACAGCCCCCAGGACATCAACCTGGAGAAAGACTGGAAGCTGGTCA CACTCTTCATTGGGGTCAACGACTTGTGTCATTACTGTGAGAATCCGGTAGGCGAATA TGTTCAGCACATCCAACAGGCCCTGGACATCCTCTCTGAGGAGCTCCCAAGGGCTTTC GTCAACGTGGTGGAGGTCATGGAGCTGGCTAGCCTGTACCAGGGCCAAGGCGGGAAAT GTGCCATGCTGGCAGCTCAGAACAACTGCACTTGCCTCAGACACTCGCAAAGCTCCCT GGAGAAGCAAGAACTGAAGAAAGTGAACTGGAACCTCCAGCATGGCATCTCCAGTTTC TCCTACTGGCACCAATACACACAGCGTGAGGACTTTGCGGTTGTGGTGCAGCCTTTCT TCCAAAACACACTCACCCCACTGAACAGAGGGGACACTGACCTCACCTTCTTCTCCGA GGACTGTTTTCACTTCTCAGACCGCGGGCATGCCGAGATGGCCATCGCACTCTGGAAC AACATGCTGGAACCAGTGGGCCGCAAGACTACCTCCAACAACTTCACCCACAGCCGAG CCAAACTCAAGTGCCCCTCTCCTGAGAGCCCTTACCTCTACACCCTGCGGAACAGCCG ATTGCTCCCAGACCAGGCTGAAGAAGCCCCCGAGGTGCTCTACTGGGCTGTCCCAGTG GCAGCGGGAGTCGGCCTTGTGGTGGGCATCATCGGGACAGTGGTCTGGAGGTGCAGGA GAGGTGGCCGGAGGGAAGATCCTCCAATGAGCCTGCGCACTGTGGCCCTCTAGGCCCG GGGGTGGGTCCTCACCCTAAACTCCCTATAGCCACTCTCTTCACCGCCCTCTGCCCCA
GCCACTCCCGGCCACCAGGACATGCTTCAATGCCTGGTGCCATAGGAAGCCCAGGGGA
CAGTCACAACTTCTTGG
ORF Start: ATG at 16 ORF Stop: TAG at 4285
SEQ TD NO: 86 1423 aa MW at l59352.7kD
NOV24c, GLRPGIFLLELLLLLGQGTPQIHTSPRKSTLEGQL PETVHSLKPSDIKFVAAIGNL EIVPDPGTGDLEKQDERPQQVCMGVMTVLSDIIRYFSPSVPMPVCHTGKRVIPHDGAE
CG57399-03 Protein Sequence DL IQAQELVRNMKENQLDFQFD KLINVFFSNASQCYLCPSAQQNGLAAGGVDELMG VLDYLQQEVPRAFVNLVDLSEVAEVSRQYHGTWLSPAPEPCNCSEETTRLAKWMQ S YQEAWNSLLASSRYSEQESFTWFQPFFYETTPSDPRLQDSTTLA HLWNRMMEPAGE KDEPLSVKHGRPMKCPSQESPYLFSYRNSNYLTRLQKPQDKLEVREGAEIRCPDKDPS DTVPTSVHRLKPADINVIGALGDSLTAGNGAGSTPGNVLDVLTQYRGLS SVGGDENI GTVTTLADILREFNPSLKGFSVGTGKETSPNAFLNQAVAGGRAEQARRLVDLMKNDTR IHFQED KIITLFIGGNDLCDFCNDLVHYSPQNFTDNIGKALDILHAEVPRAFVNLVT VLEIVNLRELYQEKKVYCPRMILRSLCPCVLKFDDNSTELATLIEFNKKFQEKTHQLI ESGRYDTREDFTVWQPFFENVDMPKTQEGLPDNSFFAPDCFHFSSKSHSRAASAL N NMLEPVGQKTTRHKFENKINITCPNQVEWPFLRTYKNSMQGHGTWLPCRDRAPSALHP TSVHALRPADIQWAALGDSLTAGNGIGSKPDDLPDVTTQYRGLSYSAGGDGSLENVT TLPDILREFNRNLTGYAVGTGDANDTNAFLNQAVPGAKARDLMSQVQTLMQKMKDDHR VNFHED KVITVLIGGSDLCDYCTDSNLYSAANFVHHLRNALDVLHREVPRVLVNLVD FLNPTIMRQVFLGNPDKCPVQQASVLCNCVLTLRENSQELARLEAFSRAYQSSMRELV GSGRYDTQEDFSWLQPFFQNIQLPVLQDGLPDTSFFAPDCIHPNQKFHSQLARALWT NMLEPLGSKTETLDLRAEMPITCPTQNEPFLRTPRNSNYTYPIKPAIEN GSDFLCTE KASNSVPTSVHQLRPADIKWAALGDSLTVAVGARPNNSSDLPTSWRGLS SIGGDG NLETHTTLPDILKKFNPYLLGFSTST EGTAGLNVAAEGARARDMPAQA DLVERMKN SPQDINLEKDWKLVTLFIGVNDLCHYCENPVGEYVQHIQQALDILSEELPRAFVNWE VMELASLYQGQGGKCAMLAAQNNCTCLRHSQSSLEKQELKKVNWNLQHGISSFSY HQ
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 24B.
Further analysis of the NOV24a protein yielded the following properties shown in Table 24C.
Table 24C. Protein Sequence Properties NOV24a
PSort 0.6850 probability located in endoplasmic reticulum (membrane); 0.6400 analysis: probability located in plasma membrane; 0.4600 probability located in Golgi body; 0.1080 probability located in nucleus
SignalP Likely cleavage site between residues 24 and 25 analysis:
A search of the NOV24a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 24D.
n a BLAST search of public sequence databases, the NOV24a protein was found to have homology to the proteins shown in the BLASTP data in Table 24E.
PFam analysis predicts that the NOV24a protein contains the domains shown in the Table 24F.
Example 25.
The NOV25 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 25 A.
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 25B.
Table 25B. Comparison of NOV25a against NOV25b through NOV25c.
Protein Sequence
Further analysis of the NOV25a protein yielded the following properties shown in Table 25C.
Table 25C. Protein Sequence Properties NOV25a
PSort 0.4500 probability located in cytoplasm; 0.3630 probability located in microbody analysis: (peroxisome); 0.1958 probability located in lysosome (lumen); 0.1000 probability located in mitochondrial matrix space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV25a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 25D.
In a BLAST search of public sequence databases, the NOV25a protein was found to have homology to the proteins shown in the BLASTP data in Table 25E.
PFam analysis predicts that the NOV25a protein contains the domains shown in the Table 25F.
Table 25F. Domain Analysis of NOV25a
Identities/
Pfam Domain NOV25a Match Region Similarities Expect Value for the Matched Region
No Significant Matches Found
Example 26.
The NOV26 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 26A.
Further analysis of the NOV26a protein yielded the following properties shown in Table 26B.
Table 26B. Protein Sequence Properties NOV26a
PSort 0.4500 probability located in cytoplasm; 0.2585 probability located in lysosome analysis: (lumen); 0.1940 probability located in microbody (peroxisome); 0.1000 probability located in mitochondrial matrix space
A search of the NOV26a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 26C.
In a BLAST search of public sequence databases, the NOV26a protein was found to have homology to the proteins shown in the BLASTP data in Table 26D.
PFam analysis predicts that the NOV26a protein contains the domains shown in the Table 26E.
Example 27.
The NOV27 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 27A.
Table 27A. NOV27 Sequence Analysis
SEQ ED NO: 95 1333 bp
NOV27a, CCTGGCCCCCAAGCTCCCCACTCTGGTGCCCCGAGCAGCCCTGTGGGCAAGCAGCCGC
CGCCATGGCCGAGCACCTGGAGCTGCTGGCAGAGATGCCCATGGTGGGCAGGATGAGC
CG57364-01 DNA Sequence ACACAGGAGCGGCTGAAGCATGCCCAGAAGCGGCGCGCCCAGCAGGTGAAGATGTGGG CCCAGGCTGAGAAGGAGGCCCAGGGCAAGAAGGGTCCTGGGGAGCGTCCCCGGAAGGA GGCAGCCAGCCAAGGGCTCCTGAAGCAGGTCCTCTTCCCTCCCAGTGTTGTCCTTCTG GAGGCCGCTGCCCGAAATGACCTGGAAGAAGTCCGCCAGTTCCTTGGGAGTGGGGTCA GCCCTGACTTGGCCAACGAGGACGGCCTGACGGCCCTGCACCAGTGCTGCATTGATGA TTTCCGAGAGATGGTGCAGCAGCTCCTGGAGGCTGGGGCCAACATCAATGCCTGTGAC AGTGAGTGCTGGACGCCTCTGCATGCTGCGGCCACCTGCGGCCACCTGCACCTGGTGG AGCTGCTCATCGCCAGTGGCGCCAATCTCCTGGCGGTCAACACCGACGGGAACATGCC CTATGACCTGTGTGATGATGAGCAGACGCTGGACTGCCTGGAGACTGCCATGGCCGAC CGTGGCATCACCCAGGACAGCATCGAGGCCGCCCGGGCCGTGCCAGAACTGCGCATGC TGGACGACATCCGGAGCCGGCTGCAGGCCGGGGCAGACCTCCATGCCCCCCTGGACCA CGGGGCCACGCTGCTGCACGTCGCAGCCGCCAACGGGTTCAGCGAGGCGGCTGCCCTG CTGCTGGAACACCGAGCCAGCCTGAGCGCTAAGGACCAAGACGGCTGGGAGCCGCTGC ACGCCGCGGCCTACTGGGGCCAGGTGCCCCTGGTGGAGCTGCTCGTGGCGCACGGGGC CGACCTGAACGCAAAGTCCCTGATGGACGAGACGCCCCTTGATGTGTGCGGGGACGAG GAGGTGCGGGCCAAGCTGCTGGAGCTGAAGCACAAGCACGACGCCCTCCTGCGCGCCC AGAGCCGCCAGCGCTCCTTGCTGCGCCGCCGCACCTCCAGCGCCGGCAGCCGCGGGAA GGTGGTGAGGCGGGATGAGCCTAACCCAGCGCAGCGGCTGACGCATGTCCCAGAAGCG GCGCGCCCAGCAGGTGAAGATGTGGGCCCAGGCTGAGAAGGAGGCCCAGGGCAAGAAG GGTCCTGGGGAGCGTCCCCGGAAGGAGGCAGCCAGCCAAGGGCTCCTGAAGCAGGTCC
TCTTCCCTCCCAGTGTTGTCCTTCTGGAGGCCGCTGCCCGAAATGACCTGGAAGAAG
ORF Start: ATG at 63 ORF Stop: TGA at 1194
SEQ ID NO: 96 377 aa MW at 41019.9kD
NOV27a, MAEHLELLAEMPMVGRMSTQERLKHAQKRRAQQVK WAQAEKEAQGKKGPGERPRKEA ASQGLLKQVLFPPSWLLEAAARNDLEEVRQFLGSGVSPDLANEDGLTALHQCCIDDF
CG57364-01 Protein Sequence REMVQQLLEAGANINACDSECWTPLHAAATCGHLHLVELLIASGANLLAVNTDGN PY DLCDDEQTLDCLETAMADRGITQDSIEAARAVPELRMLDDIRSRLQAGADLHAPLDHG ATLLHVAAANGFSEAAALLLEHRASLSAKDQDGWEPLHAAAY GQVPLVELLVAHGAD LNAKSLMDETPLDVCGDEEVRAKLLELKHKHDALLRAQSRQRSLLRRRTSSAGSRGKV VRRDEPNPAQRLTHVPEAARPAGEDVGPG
Further analysis of the NOV27a protein yielded the following properties shown in Table 27B.
Table 27B. Protein Sequence Properties NOV27a
PSort 0.3000 probability located in microbody (peroxisome); 0.3000 probability analysis: located in nucleus; 0.1547 probability located in lysosome (lumen); 0.1000 probability located in mitochondrial matrix space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV27a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 27C.
In a BLAST search of public sequence databases, the NOV27a protein was found to have homology to the proteins shown in the BLASTP data in Table 27D.
PFam analysis predicts that the NOV27a protein contains the domains shown in the Table 27E.
Example 28.
The NOV28 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 28A.
GGGTCCAAGTCGCAGAGCCGCTCCCGGAGCAGGAGTGACTCCCCACCGAGACAGGCCC CCCGCAGCGCTCCCTACAAAGGCTCTGAGATTCGGGGCTCCCGGAAGTCCAAGGACTG CAAGTACCCCCAGAAGCCACACAAGTCTCGGAGCCGGAGTTCTTCCCGTTCTCGAAGC AGGTCACGGGAGCGGGCGGATAATCCGGGAAAATACAAGAAGAAAAGTCATTACTACA GAGATCAGCGACGAGAGCGCTCGAGGTCGTATGAACGCACAGGCCGTCGCTATGAGCG GGACCACCCTGGGCACAGCAGGCATCGGAGGTGACACGTGCTTCAGACCGGTCTGGGG TGCGGCGCACACCTGGGCCCGTGCAGGGCTCAGCTCGGCAGCAGCTCTGAGGGCAGCT
CAATGAAAAAGTGAATGCACACGCCCTTGTTGGCGTG
ORF Start: ATG at 44 ORF Stop: TGA at 1598
SEQ ID NO: 98 518 aa MW at 58034.5kD
NOV28a, MTAAAAGAAGSAAPAAAAGAPGSGGAPSGSQGVLIGDRLYSGVLITLENCLLPDDKLR FTPSMSSGLDTDTETDLRWGCELIQAAGILLRLPQVAMATGQVLFQRFFYTKSFVKH
CG59348-01 Protein Sequence SMEHVSMACVHLASKIEEAPRRIRDVINVFHRLRQLRDKKKPVPLLLDQDYVNLKNQI IKAERRVLKELGFCVHVKHPHKIIVMYLQVLECERNQHLVQTS NYMNDSLRTDVFVR FQPESIACACIYLAARTLEIPLPNRPH FLLFGATEEEIQEICLKILQLYARKKVDLT HLEGEVEKRKHAIEEAKAQARGLLPGGTQVLDGTSGFSPAPKLVESPKEGKGSKPSPL SVKNTKRRLEGAKKAKADSPVNGLPKGRESRSRSRSREQSYSRSPSRSASPKRRKSDS GSTSGGSKSQSRSRSRSDSPPRQAPRSAPYKGSEIRGSRKSKDCKYPQKPHKSRSRSS SRSRSRSRERADNPGKYKKKSHYYRDQRRERSRSYERTGRRYERDHPGHSRHRR
Further analysis of the NOV28a protein yielded the following properties shown in Table 28B.
Table 28B. Protein Sequence Properties NOV28a
Psort 0.5500 probability located in endoplasmic reticulum (membrane); 0.2400 analysis: probability located in nucleus; 0.1900 probability located in lysosome (lumen); 0.1000 probability located in endoplasmic reticulum (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV28a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 28C.
In a BLAST search of public sequence databases, the NOV28a protein was found to have homology to the proteins shown in the BLASTP data in Table 28D.
PFam analysis predicts that the NOV28a protein contains the domains shown in the Table 28E.
Table 28E. Domain Analysis of NOV28a
Identities/
NOV28a Match Similarities Expect
Pfam Domain Region for the Matched Value
Region
Example 29.
The NOV29 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 29 A.
TGGGGCTGCTGGGCCCCCTGGACTGGCTGGGCCACCCCCCTCAGATCAGCCTCTTCTA CATTTTCAATTTCCTCAAGTACACCCTCTGGCCATGCCCAGTCCTGGCCCTCGTGCCC TGGGCAGTGCACATGTTCAGTGCCCAGGAAGCACCGCCCATCCACTCTTCCTGACTTC
TTGTGTGCCTCCCTTTCCTTTCCCTCCCACAAAGCCAACACTCTGTGACCACCACACT
CCAGGAGGCAGCCCCATCCCCTTCCAGCCCCTAAGTAGGCCCTCCCCTCCCTAAATCT
GCTTCCGCACCACCTGGTCTTAGCCCCAAAGATGGGCCTTCTCTCTCCCAGATAAGTT
GGTCCTCCCTCTGCCTTTCCTCTCAAGCCCCCAAAGAGCAAAGGCAACAGCAAGACCA
GCGGGTTCTTGCAACACTGTGAGGGGCAGCCAGGGCGGAAAGTACAGACTCA
ORF Start: ATG at 58 ORF Stop: TGA at 1096
SEQ ED NO: 102 346 aa MW at 38718.0kD
NOV29b, MESTLGAGIVIAEALQNQLA LENV LWITFLGDPKILFLFYFPAAYYASRRVGIAVL WISLITEWLNLIFKWFLFGDRPFWWVHESGYYSQAPAQVHQFPSSCETGPGSPSGHCM
CG59245-02 Protein Sequence ITGAAL PI TALSSQVATRARSR VRVMPSLAYCTFLLAVGLSRIFILAHFPHQVLA GLITGAVLG LMTPRVPMERELSFYGLTALALMLGTSLIYWTLFTLGLDLSWSISLAF K CERPEWIHVDSRPFASLSRDSGAALGLGIALHSPCYAQVRRAQLGNGQKIACLVLA MGLLGPLD LGHPPQISLFYIFNFLKYTL PCPVLALVPWAVHMFSAQEAPPIHSS
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 29B.
Table 29B. Comparison of NOV29a against NOV29b.
NOV29a Residues/ Identities/
Protein Sequence Match Residues Similarities for the Matched Region
NOV29b 1..337 335/347 (96%) 1..346 335/347 (96%)
Further analysis of the NOV29a protein yielded the following properties shown in Table 29C.
Table 29C. Protein Sequence Properties NOV29a
PSort 0.6850 probability located in endoplasmic reticulum (membrane); 0.6400 analysis: probability located in plasma membrane; 0.4600 probability located in Golgi body; 0.1000 probability located in endoplasmic reticulum (lumen)
SignalP Likely cleavage site between residues 41 and 42 analysis:
A search of the NOV29a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 29D.
In a BLAST search of public sequence databases, the NOV29a protein was found to have homology to the proteins shown in the BLASTP data in Table 29E.
PFam analysis predicts that the NOV29a protein contains the domains shown in the Table 29F.
Example 30.
The NOV30 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 30 A.
CG59241-01 Protein Sequence LDESDDPGVPLAPPGPEAFSGEPFNLHRFYNRSCHRLEDMLLYCSYQGGPCGPHNFSV VFTRYGKCYTFNSGRDGRPRLKTMKGGTGNGLEIMLDIQQDEYLPVWGETDETSFEAG IKVQIHSQDEPPFIDQLGFGVAPGFQTFVACQEQRIYLPPP GTCKAVTMDSDFFDSY SITACRIDCETRYLVENCNCRMVHMPGDAPYCTPEQYKECADPALDFLVEKDQEYCVC EMPCNLTRYGKELSMVKIPSKASAKYLAKKFNKSEQYIGENILVLDIFFEVLNYETIE QKKAYEIAGLLGDIGGQMGLFIGASILTVLELFDYAYEWIKHKLCRRGKCQKEAKRS SADKGVALSLDDVKRHNPCESLRGHPAGMTYAANILPHHPARGTFEDFTC
Further analysis of the NOV30a protein yielded the following properties shown in Table 30B.
Table 30B. Protein Sequence Properties NOV30a
PSort 0.7900 probability located in plasma membrane; 0.3000 probability located in analysis: Golgi body; 0.2000 probability located in endoplasmic reticulum (membrane); 0.1000 probability located in mitochondrial inner membrane
SignalP Likely cleavage site between residues 60 and 61 analysis:
A search of the NOV30a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 30C.
In a BLAST search of public sequence databases, the NOV30a protein was found to have homology to the proteins shown in the BLASTP data in Table 30D.
PFam analysis predicts that the NOV30a protein contains the domains shown in the Table 30E.
Example 31.
The NOV31 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 31 A.
Table 31A. NOV31 Sequence Analysis
SEQ ED NO: 105 1949 bp
NOV31a, TGCCTGGCTATGGCCCGACTGCTCAGGTCTGCAACCTGGGAGCTGTTCCCCTGGAGGG
GCTACTGCTCCCAGTCCCTGCAGGGAGAGCTCTGCAGGGACTTCGTAGAGGCTCTGAA
CG58602-01 DNA Sequence GGCCGTGGTGGGCGGCTCCCACGTGTCCACTGCCGCGGTGGTCCGAGAGCAGCACGGG CGCGATGAGTCGGTGCACAGGTGCGAACCTCCTGATGCTGTGGTGTGGCCCCAGAACG TGGAGCAGGTCAGCCGGCTGGCAGCCCTGTGCTATCGCCAAGGTGTGCCCATCATCCC ATTCGGCACCGGCACCGGGCTTGAGGGTGGCGTCTGTGCTGTGCAGGGCGGCGTCTGC GTTAACCTGACGCATATGGACCGAATCCTGGAGCTGAACCAGGAGGACTTCTCTGTGG TGGTGGAGCCAGGTGTCACCCGCAAAGCCCTCAACGCCCACCTGCGGGACAGCGGCCT CTGGTTTCCTCCAGACCCAGGCGCGGACGCCTCTCTCTGTGGCATGGCGGCCACCGGG GCGTCGGGGACCAACGCGGTCCGCTACGGCACCATGCGGGACAACGTGCTCAACCTGG AGGTGGTGCTGCCCGACGGGCGGCTGCTGCACACGGCGGGCCGAGGCCTCATCACAGA TTCCACTGCTGCATTCCCCCACATCAGCCCCACTGAGTGCTTTTCCCAGGGGCCAGGG CCTCATGTCAATTCTCCTCACCCTGCCCCTGAGGCCACAGTGGCCGCCACGTGTGCGT TCCCCAGTGTCCAGGCTGCTGTGGACAGCACTGTACACATCCTCCAGGCTGCAGTGCC CGTAGCCCGCATTGAGTTCCTGGATGAAGTCATGATGGATGCCTGCAACAGGTACAGC AAGCTGAATTGCTTAGTGGCGCCCACACTCTTCCTGGAGTTCCATGGCTCCCAGCAGG CACTGGAGGAGCAGCTGCAGCGCACAGAGGAGATAGTCCAGCAGAACGGAGCCTCTGA CTTCTCCTGGGCCAAGGAGGCCGAGGAGCGCAGCCGGCTTTGGACAGCACGGCACAAT GCCTGGTACGCAGCCCTGGCCACGCGGCCAGGCTGCAAGGGCTACTCCACGGATGTGT GTGTGCCCATCTCCCGGCTGCCGGAGATCGTGGTGCAGACCAAGGAGGATCTGAATGC CTCAGGACTCACAGGAAGCATTGTCGGGCATGTGGGTGACGGCAACTTCCACTGCATC CTGCTGGTCAACCCTGATGACGCCGAGGAACTGGGCAGGGTCAAGGCTTTTGCAGAAC AGCTGGGCAGGCGGGCACTGGCTCTCCACGGAACGTGCACGGGGGAGCATGGCATCGG AATGGGCAAGCGGCAGCTGCTGCAGGAGGAGGTGGGCGCCGTGGGCGTGGAGACCATG CGGCAGCTCAAGGCCGTGCTAGACCCCCAAGGCCTCATGAATCCAGGCAAAGTGCTGT QAAGGGGGTCTGAGCACTTAGCCCACAAGTTCCCTGACTACGGAGCCGGTTCTGGAAC TTTTCTTCATGCCACGGCCCCTGCAAGGAAATAGATGCTGAGGCAGTCTTCCTGCCAG
CGAGCCCACTGTATCTGGGCCCAAGGCCAGAGGGCCCAGAGAGAAGCCTGAGCACCGT
GTTACCTCCCTGGCCCTCTGGCTGGCCCCAGGAGCCTTTGGTTCAGTAAACGACCCAG
GGTGGTTCCCAGCAAAGCTGCTTCCTCTCTGCTCCTACGCATCCTGTCCTGGCGGGAA
GAGAGCGTCTGGGTCCATTCAAGACTCTGATGACACCCCTCCCCGAGGCCTCCCACTG
CCGGGGTCCCAGGACCCTTCCCCCTTCACCTGGTGACAGGAACACTCCTTTCCTGGTA
TGGAACGTGAGCTCCCGTGACATGATGATAGGTCTTCTCCTTGGGGCCTCCCCCAATA
AATCTGTAATAAACCTGAAACCCACCTACAGCTAA
ORF Start: ATG at 10 ORF Stop: TGA at 1450
SEQ ED NO: 106 480 aa MW at 51629.1kD
NOV31a, MARLLRSAT ELFP RGYCSQSLQGELCRDFVEALKAWGGSHVSTAAWREQHGRDE SVHRCEPPDAW PQNVEQVSRLAALCYRQGVPIIPFGTGTGLEGGVCAVQGGVCVNL
CG58602-01 Protein Sequence THMDRILELNQEDFSVWEPGVTRKALNAHLRDSGLWFPPDPGADASLCGMAATGASG TNAVRYGTMRDNVLNLEWLPDGRLLHTAGRGLITDSTAAFPHISPTECFSQGPGPHV NSPHPAPEATVAATCAFPSVQAAVDSTVHILQAAVPVARIEFLDEVMMDACNRYSKLN CLVAPTLFLEFHGSQQALEEQLQRTEEIVQQNGASDFSWAKEAEERSRL TARHNAWY
Ϊ9Ϊ
Further analysis of the NOV3 la protein yielded the following properties shown in Table 3 IB.
Table 31B. Protein Sequence Properties NOV31a
PSort 0.6574 probability located in mitochondrial matrix space; 0.3502 probability analysis: located in mitochondrial inner membrane; 0.3502 probability located in mitochondrial intermembrane space; 0.3502 probability located in mitochondrial outer membrane
SignalP Likely cleavage site between residues 20 and 21 analysis:
A search of the NOV31a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 31C.
In a BLAST search of public sequence databases, the NOV3 la protein was found to have homology to the proteins shown in the BLASTP data in Table 3 ID.
PFam analysis predicts that the NOV3 la protein contains the domains shown in the Table 3 IE.
Example 32.
The NOV32 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 32A.
Further analysis of the NOV32a protein yielded the following properties shown in Table 32B.
Table 32B. Protein Sequence Properties NOV32a
PSort 0.5500 probability located in endoplasmic reticulum (membrane); 0.3200 analysis: probability located in microbody (peroxisome); 0.2368 probability located in lysosome (lumen); 0.1000 probability located in endoplasmic reticulum (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV32a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 32C.
AAR29922 CRP - Homo sapiens, 225 aa. 14..224 100/218 (45%) 2e-43 [WO9221364-A, 10-DEC-1992] 11..224 132/218 (59%)
AAR74769 Female hamster protein, lfhp 24-222 95/206 (46%) 6e-43 Cricetus cricetus, 210 aa. 1..199 132/206 (63%) [WO9505394-A, 23-FEB-1995]
AAY76844 Human C reactive protein (CRP) 24-224 98/208 (47%) le-42 sequence - Homo sapiens, 206 aa. 2..205 128/208 (61%) [JP2000014388-A, 18-JAN-2000]
In a BLAST search of public sequence databases, the NOV32a protein was found to have homology to the proteins shown in the BLASTP data in Table 32D.
PFam analysis predicts that the NOV32a protein contains the domains shown in the Table 32E.
Example 33.
The NOV33 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 33 A.
Table 33A. NOV33 Sequence Analysis
SEQ ED NO: 109 3350 bp
NOV33a, TAATGAGGAGACTGAGTTTGTGGTGGCTGCTGAGCAGGGTCTGTCTGCTGTTGCCGCC GCCCTGCGCACTGGTGCTGGCCGGGGTGCCCAGCTCCTCCTCGCACCCGCAGCCCTGC
CG58183-01 DNA Sequence CAGATCCTCAAGCGCATCGGGCACGCGGTGAGGGTGGGCGCGGTGCACTTGCAGCCCT GGACCACCGCCCCCCGCGCGGCCAGCCGCGCTCCGGACGACAGCCGAGCAGGAGCCCA GAGGGATGAGCCGGAGCCAGGGACTAGGCGGTCCCCGGCGCCCTCGCCGGGCGCACGC TGGTTGGGGAGCACCCTGCATGGCCGGGGGCCGCCGGGCTCCCGTAAGCCCGGGGAGG GCGCCAGGGCGGAGGCCCTGTGGCCACGGGACGCCCTCCTATTTGCCGTGGACAACCT GAACCGCGTGGAAGGGCTGCTACCCTACAACCTGTCTTTGGAAGTAGTGATGGCCATC GAGGCAGGCCTGGGCGATCTGCCACTTTTGCCCTTCTCCTCCCCTAGTTCGCCATGGA GCAGTGACCCTTTCTCCTTCCTGCAAAGTGTGTGCCATACCGTGGTGGTGCAAGGGGT GTCGGCGCTGCTCGCCTTCCCCCAGAGCCAGGGCGAAATGATGGAGCTCGACTTGGTC AGCTTAGTCCTGCACATTCCAGTGATCAGCATCGTGCGCCACGAGTTTCCACGGGAGA GTCAGAATCCCCTTCACCTACAACTGAGTTTAGAAAATTCATTAAGTTCTGATGCTGA TGTCACTGTCTCAATCCTGACCATGAACAACTGGTACAATTTTAGCTTGTTGCTGTGC CAGGAAGACTGGAACATCACCGACTTCCTCCTCCTTACCCAGAATAATTCCAAGTTCC ACCTTGGTTCTATCATCAACATCACCGCTAACCTCCCCTCCACCCAGGACCTCTTGAG CTTCCTACAGATCCAGCTTGAGAGTATTAAGAACAGCACACCCACAGTGGTGATGTTT GGCTGCGACATGGAAAGTATCCGGCGGATTTTCGAAATTACAACCCAGTTTGGGGTCA TGCCCCCTGAACTTCGTTGGGTGCTGGGAGATTCCCAGAATGTGGAGGAACTGAGGAC AGAGGGTCTGCCCTTAGGGCTCATTGCTCATGGAAAAACAACACAGTCTGTCTTTGAG CACTACGTACAAGATGCTATGGAGCTGGTCGCAAGAGCTGTAGCCACAGCCACCATGA TCCAACCAGAACTTGCTCTCATTCCCAGCACGATGAACTGCATGGAGGTGGAAACTAC AAATCTCACTTCAGGACAATATTTATCAAGGTTTCTAGCCAATACCACTTTCAGAGGC CTCAGTGGTTCCATCAGAGTAAAAGGTTCCACCATCGTCAGCTCAGAAAACAACTTTT TCATCTGGAATCTTCAACATGACCCCATGGGAAAGCCAATGTGGACCCGCTTGGGCAG CTGGCAGGGGGGAAAGATTGTCATGGACTATGGAATATGGCCAGAGCAGGCCCAGAGA CACAAAACCCACTTCCAACATCCAAGTAAGCTACACTTGAGAGTGGTTACCCTGATTG AGCATCCTTTTGTCTTCACAAGGGAGGTAGATGATGAAGGCTTGTGCCCTGCTGGCCA ACTCTGTCTAGACCCCATGACTAATGACTCTTCCACATTGGACAGCCTTTTTAGCAGC CTCCATAGCAGTAATGATACAGTGCCCATTAAATTCAAGAAGTGCTGCTATGGATATT GCATTGATCTGCTGGAAAAGATAGCAGAAGACATGAACTTTGACTTCGACCTCTATAT TGTAGGGGATGGAAAGTATGGAGCATGGAAAAATGGGCACTGGACTGGGCTAGTGGGT GATCTCCTGAGAGGGACTGCCCACATGGCAGTCACTTCCTTTAGCATCAATACTGCAC GGAGCCAGGTGATAGATTTCACCAGCCCTTTCTTCTCCACCAGCTTGGGCATCTTAGT GAGGACCCGAGATACAGCAGCTCCCATTGGAGCCTTCATGTGGCCACTCCACTGGACA ATGTGGCTGGGGATTTTTGTGGCTCTGCACATCACTGCCGTCTTCCTCACTCTGTATG AATGGAAGAGTCCATTTGGTTTGACTTCCAAGGGGCGAAATAGAAGTAAAGTCTTCTC CTTTTCTTCAGCCTTGAACATCTGTTATGCCCTCTTGTTTGGCAGAACAGTGGCCATC AAACCTCCAAAATGTTGGACTGGAAGGTTTCTAATGAACCTTTGGGCCATTTTCTGTA TGTTTTGCCTTTCCACATACACGGCAAACTTGGCTGCTGTCATGGTAGGTGAGAAGAT CTATGAAGAGCTTTCTGGAATACATGACCCCAAGTTACATCATCCTTCCCAAGGATTC CGCTTTGGAACTGTCCGAGAAAGCAGTGCTGAAGATTATGTGAGACAAAGTTTCCCAG AGATGCATGAATATATGAGAAGGTACAATGTTCCAGCCACCCCTGATGGAGTGGAGTA TCTGAAGAATGATCCAGAGAAACTAGACGCCTTCATCATGGACAAAGCCCTTCTGGAT TATGAAGTGTCAATAGATGCTGACTGCAAACTTCTCACTGTGGGGAAGCCATTTGCCA TAGAAGGTTACGGCATTGGCCTCCCACCCAACTCTCCATTGACCGCCAACATATCCGA GCTAATCAGTCAATACAAGTCACATGGGTTTATGGATATGCTCCATGACAAGTGGTAC AGGGTGGTTCCCTGTGGCAAGAGAAGTTTTGCTGTCACGGAGACTTTGCAAATGGGCA
Further analysis of the NOV33a protein yielded the following properties shown in Table 33B.
Table 33B. Protein Sequence Properties NOV33a
PSort 0.6400 probability located in plasma membrane; 0.4600 probability located in analysis: Golgi body; 0.3700 probability located in endoplasmic reticulum (membrane); 0.1000 probability located in endoplasmic reticulum (lumen)
SignalP Likely cleavage site between residues 34 and 35 analysis:
A search of the NOV33a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 33C.
In a BLAST search of public sequence databases, the NOV33a protein was found to have homology to the proteins shown in the BLASTP data in Table 33D.
PFam analysis predicts that the NOV33a protein contains the domains shown in the Table 33E.
Example 34.
The NOV34 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 34A.
NOV34a, MGEWAFLGSLLDAVQLQSPLVGRLWLWMLIFRILVLATVGGAVFEDEQEEFVCNTLQ PGCRQTCYDRAFPVSHYRF LFHILLLSAPPVLFWYSMHRAGKEAGGAEAAAQCAPG
CG59315-01 Protein Sequence LPEAQCAPCALRARRARRCYLLSVALRLLAELTFLGGQALLYGFRVAPHFACAGPPCP HTVDCFVSRPTEKTVFVLFYFAVGLLSALLSVAELGHLL KGRPRAGERDNRCNRAHE EAQKLLPPPPPPPPPPALPSRRPGPEPCAPPAYAHPAPASLRECGSGRGRNAPMAPRC GRHRLTPYPPAALPQGPSSLSPANSRELCPGENQPRTGVSASPPLVPTDTSQPRSYLS SFLEAGGEGS TQECKHACTQLHCLPSPPADAARVPLPRSPSWQGGRRRALHSGFPTP PSRSQART
Further analysis of the NOV34a protein yielded the following properties shown in Table 34B.
Table 34B. Protein Sequence Properties NOV34a
PSort 0.6000 probability located in plasma membrane; 0.4000 probability located in analysis: Golgi body; 0.3000 probability located in endoplasmic reticulum (membrane); 0.0300 probability located in mitochondrial inner membrane
SignalP Likely cleavage site between residues 39 and 40 analysis:
A search of the NOV34a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 34C.
In a BLAST search of public sequence databases, the NOV34a protein was found to have homology to the proteins shown in the BLASTP data in Table 34D.
PFam analysis predicts that the NOV34a protein contains the domains shown in the Table 34E.
Example 35.
The NOV35 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 35 A.
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 35B.
Table 35B. Comparison of NOV35a against NOV35b.
NOV35a Residues/ Identities/
Protein Sequence Match Residues Similarities for the Matched Region
NOV35b 50..148 97/99 (97%) 1..99 98/99 (98%)
Further analysis of the NOV35a protein yielded the following properties shown in Table 35C.
Table 35C. Protein Sequence Properties NOV35a
Psort 0.3700 probability located in outside; 0.1697 probability located in microbody analysis: (peroxisome); 0.1000 probability located in endoplasmic reticulum (membrane); 0.1000 probability located in endoplasmic reticulum (lumen)
SignalP Likely cleavage site between residues 20 and 21 analysis:
A search of the NOV35a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 35D.
In a BLAST search of public sequence databases, the NOV35a protein was found to have homology to the proteins shown in the BLASTP data in Table 35E.
PFam analysis predicts that the NOV35a protein contains the domains shown in the Table 35F.
Table 35F. Domain Analysis of NOV35a
Identities/
Pfam Domain NOV35a Match Region Similarities Expect Value for the Matched Region lys: domain 1 of 1 20-145 68/129 (53%) 8e-58 107/129 (83%)
Example 36.
The NOV36 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 36A.
Table 36A. NOV36 Sequence Analysis
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 36B.
Further analysis of the NOV36a protein yielded the following properties shown in Table 36C.
Table 36C. Protein Sequence Properties NOV36a
PSort 0.5666 probability located in microbody (peroxisome); 0.4500 probability analysis: located in cytoplasm; 0.1562 probability located in lysosome (lumen); 0.1000 probability located in mitochondrial matrix space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV36a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 36D.
In a BLAST search of public sequence databases, the NOV36a protein was found to have homology to the proteins shown in the BLASTP data in Table 36E.
PFam analysis predicts that the NOV36a protein contains the domains shown in the Table 36F.
Table 36F. Domain Analysis of NOV36a
Identities/
Pfam Domain NOV36a Match Region Similarities Expect Value for the Matched Region
No Significant Matches Found
Example 37.
The NOV37 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 37A.
Further analysis of the NOV37a protein yielded the following properties shown in Table 37B.
Table 37B. Protein Sequence Properties NOV37a
PSort 0.6400 probability located in microbody (peroxisome); 0.4500 probability analysis: located in cytoplasm; 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV37a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 37C.
Ln a BLAST search of public sequence databases, the NOV37a protein was found to have homology to the proteins shown in the BLASTP data in Table 37D.
PFam analysis predicts that the NOV37a protein contains the domains shown in the Table 37E.
Example 38.
The NOV38 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 38 A.
Further analysis of the NOV38a protein yielded the following properties shown in Table 38B.
Table 38B. Protein Sequence Properties NOV38a
PSort 0.4404 probability located in mitochondrial matrix space; 0.3000 probability analysis: located in microbody (peroxisome); 0.1257 probability located in mitochondrial inner membrane; 0.1257 probability located in mitochondrial intermembrane space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV38a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 38C.
In a BLAST search of public sequence databases, the NOV38a protein was found to have homology to the proteins shown in the BLASTP data in Table 38D.
Table 38D. Public BLASTP Results for NOV38a
NOV38a
Protein Identities/ Residues/ Expect
Accession Protein/Organism/Length Similarities for the
Match Value
Number Matched Portion Residues
No Significant Matches Found
PFam analysis predicts that the NOV38a protein contains the domains shown in the Table 38E.
Example 39.
The NOV39 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 39 A.
Table 39A. NOV39 Sequence Analysis
SEQ ED NO: 125 1421 bp
NOV39a, ACCATTTCAGAGATGTCTTCCAGAAGTACCAAAGATTTAATTAAAAGTAAGTGGGGAT
CGAAGCCTAGTAACTCCAAATCCGAAACTACATTAGAAAAATTAAAGGGAGAAATTGC
CG59371-01 DNA Sequence ACACTTAAAGACATCAGTGGATGAAATCACAAGTGGGAAAGGAAAGCTGACTGATAAA GAGAGACACAGACTTTTGGAGAAAATTCGAGTCCTTGAGGCTGAGAAGGAGAAGAATG CTTATCAACTCACAGAGAAGGACAAAGAAATACAGCGACTGAGAGACCAACTGAAGGC CAGATATAGTACTACCACATTGCTTGAACAGCTGGAAGAGACAACGAGAGAAGGAGAA AGGAGGGAGCAGGTGTTGAAAGCCTTATCTGAAGAGAAAGACGTATTGAAACAACAGT TGTCTGCTGCAACCTCACGAATTGCTGAACTTGAAAGCAAAACCAATACACTCCGTTT ATCACAGACTGTGGCTCCAAACTGCTTCAACTCATCAATAAATAATATTCATGAAATG GAAATACAGCTGAAAGATGCTCTGGAGAAAAATCAGCAGTGGCTCGTGTATGATCAGC AGCGGGAAGTCTATGTAAAAGGACTTTTAGCAAAGATCTTTGAGTTGGAAAAGAAAAC GGAAACAGCTGCTCATTCACTCCCACAGCAGACAAAAAAGCCTGAATCAGAAGGTTAT CTTCAAGAAGAGAAGCAGAAATGTTACAACGATCTCTTGGCAAGTGCAAAAAAAGATC TTGAGGTTGAACGACAAACCATAACTCAGCTGAGTTTTGAACTGAGTGAATTTCGAAG AAAATATGAAGAAACCCAAAAAGAAGTTCACAATTTAAATCAGCTGTTGTATTCACAA AGAAGGGCAGATGTGCAACATCTGGAAGATGATAGGCATAAAACAGAGAAGATACAAA AACTCAGGGAAGAGAATGATATTGCTAGGGGAAAACTTGAAGAAGAGAAGAAGAGATC CGAAGAGCTCTTATCTCAGGTCCAGTCTCTTTACACATCTCTGCTAAAGCAGCAAGAA GAACAAACAAGGGTAGCTCTGTTGGAACAACAGATGCAGGCATGTACTTTAGACTTTG AAAATGAAAAACTCGACCGTCAACATGTGCAGCATCAATTGCATGTAATTCTTAAGGA GCTCCGAAAAGCAAGAAAAAATATAACACAGTTGGAATCCTTGAAACAGCTTCATGAG TTTGCCATCACAGAGCCATTAGTCACTTTCCAAGGAGAGACTGAAAACAGAGAAAAAG TTGCCGCCTCACCAAAAAGTCCCACTGCTGCACTCAATGGAAGCCTGGTGGAATGTCC CAAGTGCAATATACAGTATCCAGCCACTGAGCATCGCGATCTGCTTGTCCATGTGGAA TACTGTTCAAAGTAGCAAAATAAGTATTT
ORF Start: ATG at 13 ORF Stop: TAG at 1405
SEQ ED NO: 126 464 aa MW at 54045.6kD
NOV39a, MSSRSTKDLIKSK GSKPSNSKSETTLEKLKGEIAHLKTSVDEITSGKGKLTDKERHR LLEKIRVLEAEKEKNAYQLTEKDKEIQRLRDQLKARYSTTTLLEQLEETTREGERREQ
CG59371-01 Protein Sequence VLKALSEEKDVLKQQLSAATSRIAELESKTNTLRLSQTVAPNCFNSSINNIHEMEIQL KDALEKNQQ LVYDQQREVYVKGLLAKIFELEKKTETAAHSLPQQTKKPESEGYLQEE KQKCYNDLLASAKKDLEVERQTITQLSFELSEFRRKYEETQKEVHNL QLLYSQRRAD VQHLEDDRHKTEKIQKLREENDIARGKLEEEKKRSEELLSQVQSLYTSLLKQQEEQTR VALLEQQMQACTLDFENEKLDRQHVQHQLHVILKELRKARKNITQLESLKQLHEFAIT EPLVTFQGETENREKVAASPKSPTAALNGSLVECPKCNIQYPATEHRDLLVHVEYCSK
Further analysis of the NOV39a protein yielded the following properties shown in Table 39B.
Table 39B. Protein Sequence Properties NOV39a
PSort 0.4500 probability located in cytoplasm; 0.3000 probability located in microbody analysis: (peroxisome); 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV39a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 39C.
In a BLAST search of public sequence databases, the NOV39a protein was found to have homology to the proteins shown in the BLASTP data in Table 39D.
PFam analysis predicts that the NOV39a protein contains the domains shown in the Table 39E.
Table 39E. Domain Analysis of NOV39a
Identities/
Pfam Domain NOV39a Match Region Similarities Expect Value for the Matched Region
No Significant Matches Found
Example 40.
The NOV40 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 40A.
TGGCACTCTCCGCAAGGGACCGAGCCATGAAGGAGTCTCAACAGGGACCCAAAGGGGA GGCCCCCAAGGCCGACCTCAACAAACCTCTTTACATTGATACCAAAATGCGGCCCAGC CTGGATGCCGGCTTCCCTACGGTCACCAGGCAGAACACCCGGGGACCCCTGAGGCGGC AGGAGACGGAGAACAAGTACGAGACCGACCTGGGCCGAGACCGGAAAGGCGATGACAA GAAGAACATGCTGATCGACATCATGGACACGTCCCAGCAGAAGTCGGCTGGCCTGCTG ATGGTGCACACCGTGGACGCCACTAAGCTGGACAACGCCCTGCAGGAAGAGGACGAGA AGGCAGAGGTGGAGATGAAGCCAGACAGCTCGCCGTCCGAGGTGCCAGAAGGTGTTTC CGAAACCGAAGGTGCTTTACAGATCTCCGCTGCCCCCGAGCCCACCACCGTGCCCGGC AGAACCATCGTCGCGGTGGGCTCCATGGAAGAGGCGGTGATTTTGCCATTCCGCATCC CTCCTCCCCCTCTGGCATCCGTGGACTTGGATGAGGATTTTATTTTTACAGAGCCATT GCCTCCTCCCCTGGAATTTGCAAATAGTTTTGATATCCCCGATGACCGGGCAGCTTCT GTCCCGGCTCTCTCAGACTTAGTGAAGCAGAAGAAAAGCGACACCCCTCAGTCCCCTT- CGTTGAACTCCAGCCAACCAACCAACTCTGCAGACAGCAAGAAGCCAGCCAGTCTTTC AAACTGTCTGCCTGCCTCATTCCTGCCACCCCCTGAAAGCTTTGACGCCGTCGCCGAC TCTGGGATCGAGGAGGTGGACAGCCGGAGTAGCAGCGACCACCACCTCGAGACGACCA GCACTATCTCCACCGTGTCTAGCATCTCCACCCTGTCTTCCGAAGGTGGAGAGAATGT GGACACCTGCACAGTCTATGCAGATGGGCAAGCATTTATGGTTGACAAACCCCCAGTA CCTCCTAAGCCAAAAATGAAGCCCATCATTCACAAAAGCAATGCACTTTATCAAGACG CGCTCGTGGAAGAAGATGTAGATAGCTTTGTTATCCCCCCGCCCGCTCCCCCGCCCCC GCCGGGCAGTGCCCAGCCTGGGATGGCCAAGGTTCTCCAGCCAAGGACCTCCAAGTTG TGGGGCGACGTCACAGAGATCAAAAGCCCGATTCTCTCAGGCCCAAAGGCAAACGTTA TTAGTGAATTGAACTCTATCCTACAGCAAATGAACCGAGAGAAATTGGCAAAGCCGGG GGAAGGACTGGATTCACCAATGGGAGCCAAGTCCGCCAGCCTCGCTCCAAGAAGCCCG GAGATCATGAGCACCATCTCAGGTACACGGAGCACGACGGTCACCTTCACTGTTCGCC CCGGCACCTCCCAGCCCATCACCCTGCAGAGCCGGCCCCCCGACTATGAAAGCAGGAC CTCAGGAACAAGACGTGCCCCAAGCCCTGTGGTCTCGCCAACAGAGATGAACAAAGAG ACCCTGCCCGCCCCCCTGTCTGCTGCCACCGCCTCTCCTTCTCCCGCTCTCTCAGATG TCTTTAGCCTTCCAAGCCAGCCCCCTTCTGGGGATCTATTTGGCTTGAACCCAGCGGG ACGCAGTAGGTCGCCATCCCCCTCGATACTGCAACAGCCAATCTCAAATAAGCCTTTT ACAACTAAACCTGTCCACCTGTGGACTAAACCAGATGTGGCCGATTGGCTGGAAAGTC TAAACTTGGGTGAACATAAAGAGGCCTTCATGGACAATGAGATCGATGGCAGTCACTT ACCAAACCTGCAGAAGGAGGACCTCATCGATCTTGGGGTAACTCGAGTCGGGCACAGA ATGAACATAGAAAGGGCTTTGAAACAGCTGCTGGACAGATAAGGACGGCTGCTCTCCA CCTCGCAGACTGCTCTTGTTATAAGTAGAGATGGGCTCGTGCTGAAACATCTGAATGC
CAAGCGAAGTC
ORF Start: ATG at 67 ORF Stop: TAA at 3868
SEQ ED NO: 128 1267 aa MW at l36108.7kD
NOV40a, MMMNVPGGGAAAVMMTGYNNGRCPRNSLYSDCIIEEKTWLQKKDNEGFGFVLRGAKA DTPIEEFTPTPAFPALQYLESVDEGGVA QAGLRTGDFLIEVNNENWKVGHRQWNM
CG59346-01 Protein Sequence IRQGGNHLVLKWTVTRNLDPDDTARKKAPPPPKRAPTTALTLRSKSMTSELEELDKP EEIVPASKPSRAAENMAVEPRVATIKQRPSSRCFPAGSDMNVSGRTLGPRGRGPTVPP RLSGLQSVYERQGIAVMTPTVPGSPKAPFLGIPRGTMRRQKSIGITEEERQFLAPPML KFTRSLSMPDTSEDIPPPPQSVPPSPPPPSPTTYNCPKSPTPRVYGTIKPAFNQNSAA KVSPATRSDTVATMMREKGMYFRRELDRYSLDSEDLYSRNAGPQANFRNKRGQMPENP YSEVGKIASKAVYVPAKPARRKGMLVKQSNVEDSPEKTCSIPIPTIIVKEPSTSSSGK SSQGSSMEIDPQAPEPPSQLRPDESLTVSSPFAAAIAGAVRDREKRLEARRNSPAFLS TDLGDEDVGLGPPAPRTRPSMFPEEGDFADEDSAEQLSSPMPSATPREPENHFVGGAE ASAPGEAGRPLNSTSKAQGPESSPAVPSASSGTAGPGNYVHPLTGRLLDPSSPLALAL SARDRAMKESQQGPKGEAPKADLNKPLYIDTKMRPSLDAGFPTVTRQNTRGPLRRQET ENKYETDLGRDRKGDDKKNMLIDIMDTSQQKSAGLL VHTVDATKLDNALQEEDEKAE VEMKPDSSPSEVPEGVSETEGALQISAAPEPTTVPGRTIVAVGSMEEAVILPFRIPPP PLASVDLDEDFIFTEPLPPPLEFANSFDIPDDRAASVPALSDLVKQKKSDTPQSPSLN SSQPTNSADSKKPASLSNCLPASFLPPPESFDAVADSGIEEVDSRSSSDHHLETTSTI STVSSISTLSSEGGENVDTCTVYADGQAFMVDKPPVPPKPKMKPIIHKSNALYQDALV EEDVDSFVIPPPAPPPPPGSAQPGMAKVLQPRTSKL GDVTEIKSPILSGPKANVISE LNSILQQMNREKLAKPGEGLDSPMGAKSASLAPRSPEIMSTISGTRSTTVTFTVRPGT SQPITLQSRPPDYESRTSGTRRAPSPWSPTEMNKETLPAPLSAATASPSPALSDVFS LPSQPPSGDLFGLNPAGRSRSPSPSILQQPISNKPFTTKPVHL TKPDVAD LESLNL GEHKEAFMDNEIDGSHLPNLQKEDLIDLGVTRVGHRMNIERALKQLLDR
Further analysis of the NOV40a protein yielded the following properties shown in Table 40B.
Table 40B. Protein Sequence Properties NOV40a
PSort 0.4500 probability located in cytoplasm; 0.3000 probability located in microbody analysis: (peroxisome); 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV40a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 40C.
In a BLAST search of public sequence databases, the NOV40a protein was found to have homology to the proteins shown in the BLASTP data in Table 40D.
PFam analysis predicts that the NOV40a protein contains the domains shown in the Table 40E.
Example 41.
The NOV41 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 41 A.
Table 41A. NOV41 Sequence Analysis
SEQ ED NO: 129 2069 bp
NOV41a, GGACACTGACATGGACTGAAGGAGTAGAAAGCACTATAAATGTCTTTCCTTATCTGTG
TGTACTCTTATCTCACTGTTCTATTTTTTCTCCTCATTTATATTAACTCTTTCTTACC
CG57814-01 DNA Sequence TTTTTTTCTGAACTTCTAGGCCTTCTCTTTCCAGAACTGGTGGAAGACAAATGAAACG
GCCAAGATGGTAAGAAACAAGCCGCATTTCTCCTTGGGGAGACTGATAATTTAAAAGG
TTTGTTGTGTCAGAAACATTCCCAGCTTCATCACCAACCCTTTCCTTCCACCTCTGCC
CACTGGAGACCACTTACATCCCGAAGCGGACGCGGCAGCTGAAGTCAGGAAACCATGC
ATCACATTAGCAGGAGCCAACTGCAGACTTTAAACTCCGTTCAACATGTGGATGCGGC
AGAGAAATGACCTGTCCAGACAAGCCGGGGCAGCTCATAAACTGGTTCATCTGCTCCC
TGTGCGTCCCGCGGGTGCGTAAGCTCTGGAGCAGCCGGCGTCCAAGGACCCGGAGAAA CCTTCTGCTGGGCACTGCGTGTGCCATCTACTTGGGCTTCCTGGTGAGCCAGGTGGGG AGGGCCTCTCTCCAGCATGGACAGGCGGCTGAGAAGGGGCCACATCGCAGCCGCGACA CCGCCGAGCCATCCTTCCCTGAGATACCCCTGGATGGTACCCTGGCCCCTCCAGAGTC CCAGGGCAATGGGTCCACTCTGCAGCCCAATGTGGTGTACATTACCCTACGCTCCGAG CGCAGCAAGCCGGCCAATATCCGTGGCACCGTGAAGCCCAAGCGCAGGAAAAAGCATG CAGTGGCATCGGCTGCCCCAGGGCAGGAGGCTTTGGTCGGACCATCCCTTCAGCCGCA GGAAGCGGCAAGGGAAGCTGATGCTGTAGCACCTGGGTACGCTCAGGGAGCAAACCTG GTTAAGATTGGAGAGCGACCCTGGAGGTTGGTGCGGGGTCCGGGAGTGCGAGCCGGGG GCCCAGACTTCCTGCAGCCCAGCTCCAGGGAGAGCAACATTAGGATCTACAGCGAGAG CGCCCCCTCCTGGCTGAGCAAAGATGACATCCGAAGAATGCGACTCTTGGCGGACAGC GCAGTGGCAGGGCTCCGGCCTGTGTCCTCTAGGAGCGGAGCCCGTTTGCTGGTGCTGG AGGGGGGCGCACCTGGCGCTGTGCTCCGCTGTGGCCCTAGCCCCTGTGGGCTTCTCAA GCAGCCCTTGGACATGAGTGAGGTGTTTGCCTTCCACCTAGACAGGATCCTGGGGCTC AACAGGACCCTGCCGTCTGTGAGCAGGAAAGCAGAGTTCATCCAAGATGGCCGCCCAT GCCCCATCATTCTTTGGGATGCATCTTTATCTTCAGCAAGTAATGACACCCATTCTTC TGTTAAGCTCACCTGGGGAACTTATCAGCAGTTGCTGAAACAGAAATGCTGGCAGAAT GGCCGAGTACCCAAGCCTGAATCAGGTTGTACTGAAATACATCATCATGAGTGGTCCA AGATGGCACTCTTTGATTTTTTGTTACAGATTTATAATCGCTTAGATACAAATTGCTG TGGATTCAGACCTCGCAAGGAAGATGCCTGTGTACAGAATGGATTGAGGCCAAAATGT GATGACCAAGGTTCTGCGGCTCTAGCACACATTATCCAGCGAAAGCATGACCCAAGGC ATTTGGTTTTTATAGACAACAAGGGTTTCTTTGACAGGAGTGAAGATAACTTAAACTT CAAATTGTTAGAAGGCATCAAAGAGTTTCCAGCTTCTGCAGTTTCTGTTTTGAAGAGC CAGCACTTACGGCAGAAACTTCTTCAGTCTCTGTTTCTTGATAAAGTGTATTGGGAAA GTCAAGGAGGTAGACAAGGAATTGAAAAGCTTATCGATGTAATAGAACACAGAGCCAA AATTCTTATCACCTATATCAATGCACACGGGGTCAAAGTATTACCTATGAATGAATGA CAAAAGAATCTTCTGGCTAGGGTGTTAGATATATTTATGCATTTTTGGTTTTGTTTTT
AAATCAAGCACATCAACCTCAAGCCCGTTTAGCAATGAG
ORF Start: ATG at 413 ORF Stop: TGA at 1970
SEQ ED NO: 130 519 aa MW at 57552.4kD
NOV41a, MTCPDKPGQLINWFICSLCVPRVRKLWSSRRPRTRRNLLLGTACAIYLGFLVSQVGRA SLQHGQAAEKGPHRSRDTAEPSFPEIPLDGTLAPPESQGNGSTLQP WYITLRSERS
CG57814-01 Protein Sequence KPANIRGTVKPKRRKKHAVASAAPGQEALVGPSLQPQEAAREADAVAPGYAQGANLVK IGERPWRLVRGPGVRAGGPDFLQPSSRESNIRIYSESAPSWLSKDDIRRMRLLADSAV AGLRPVSSRSGARLLVLEGGAPGAVLRCGPSPCGLLKQPLDMSEVFAFHLDRILGLNR TLPSVSRKAEFIQDGRPCPIILWDASLSSASNDTHSSVKLT GTYQQLLKQKC QNGR VPKPESGCTEIHHHE SKMALFDFLLQIYNRLDTNCCGFRPRKEDACVQNGLRPKCDD QGSAALAHIIQRKHDPRHLVFIDNKGFFDRSEDNLNFKLLEGIKEFPASAVSVLKSQH LRQKLLQSLFLDKVYWESQGGRQGIEKLIDVIEHRAKILITYINAHGVKVLPMNE
SEQ ED NO: 131 1740 bp
NOV41b, GGCAGCTGAAGTCAGGAAACCATGCATCACATTAGCAGGAGCCAACTGCAGACTTTAA
ACTCCGTTCAACATGTGGATGCGGCAGAGAAATGACCTGTCCAGACAAGCCGGGGCAG
CG57814-02 DNA Sequence CTCATAAACTGGTTCATCTGCTCCCTGTGCGTCCCGCGGGTGCGTAAGCTCTGGAGCA GCCGGCGTCCAAGGACCCGGAGAAACCTTCTGCTGGGCACTGCGTGTGCCATCTACTT GGGCTTCCTGGTGAGCCAGGTGGGGAGGGCCTCTCTCCAGCATGGACAGGCGGCTGAG AAGGGGCCACATCGCAGCCGCGACACCGCCGAGCCATCCTTCCCTGAGATACCCCTGG ATGGTACCCTGGCCCCTCCAGAGTCCCAGGGCAATGGGTCCACTCTGCAGCCCAATGT GGTGTACATTACCCTACGCTCCAAGCGCAGCAAGCCGGCCAATATCCGTGGCACCGTG AAGCCCAAGCGCAGGAAAAAGCATGCAGTGGCATCGGCTGCCCAAGGGCAGGAGGCTT TGGTCGGACCATCCCTTCAGCCGCAAGAAGCGGCAAGGGAAGCTGATGCTGTAGCACT GGGTACGCTCAGGAGCAAACTGGTTAAGATGGAGAGCGACCCTGAAGGTGGTGCGGGG TCGGGAGTGCGAGCCGGGGGCCCAGACTTCCTGCAGCCCAGCTCCAGGGAGAGCAACA TTAGGATCTACAGCGAGAGCGCCCCCTCCTGGCTGAGCAAAGATGACATCCGAAGAAT GCGACTCTTGGCGGACAGCGCAGTGGCAGGGCTCCGGCCTGTGTCCTCTAGGAGCGGA GCCCGTTTGCTGGTGCTGGAGGGGGGCGCACCTGGCGCTGTGCTCCGCTGTGGCCCTA GCCCCTGTGGGCTTCTCAAGCAGCCCTTGGACATGAGTGAGGTGTTTGCCTTCCACCT AGACAGGATCCTGGGGCTCAACAGGACCCTGCCGTCTGTGAGCAGGAAAGCAGAGTTC
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 41B.
Further analysis of the NOV4 la protein yielded the following properties shown in Table 4 IC.
Table 41 C. Protein Sequence Properties NOV41a
PSort 0.5500 probability located in endoplasmic reticulum (membrane); 0.2404 analysis: probability located in lysosome (lumen); 0.1000 probability located in endoplasmic reticulum (lumen); 0.1000 probability located in outside
SignalP Likely cleavage site between residues 59 and 60 analysis:
A search of the NOV41a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 41D.
In a BLAST search of public sequence databases, the NOV4 la protein was found to have homology to the proteins shown in the BLASTP data in Table 4 IE.
PFam analysis predicts that the NOV4 la protein contains the domains shown in the Table 41F.
Example 42.
The NOV42 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 42A.
Further analysis of the NOV42a protein yielded the following properties shown in Table 42B.
Table 42B. Protein Sequence Properties NOV42a
analysis: probability located in plasma membrane; 0.4600 probability located in Golgi body; 0.1000 probability located in endoplasmic reticulum (lumen)
SignalP Likely cleavage site between residues 32 and 33 analysis:
A search of the NOV42a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 42C.
In a BLAST search of public sequence databases, the NOV42a protein was found to have homology to the proteins shown in the BLASTP data in Table 42D.
PFam analysis predicts that the NOV42a protein contains the domains shown in the Table 42E.
Example 43.
The NOV43 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 43 A.
Further analysis of the NOV43a protein yielded the following properties shown in Table 43B.
Table 43B. Protein Sequence Properties NOV43a
PSort 0.6500 probability located in cytoplasm; 0.1000 probability located in analysis: mitochondrial matrix space; 0.1000 probability located in lysosome (lumen); 0.0053 probability located in microbody (peroxisome)
A search of the NOV43a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 43C.
In a BLAST search of public sequence databases, the NOV43a protein was found to have homology to the proteins shown in the BLASTP data in Table 43D.
PFam analysis predicts that the NOV43a protein contains the domains shown in the Table 43E.
Example 44.
The NOV44 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 44A.
Table 44A. NOV44 Sequence Analysis
SEQ ED NO: 137 1561 bp
NOV44a, AAGAATTGTAGCTCTCCACTGAATTGCAGGGGTTCTTGAATGTTGTCAACATTTGGAG
GCAGTTGGAGGAGGGAGCTCTATTGATGAAAAATGGCTACATATTCAAAATTTCAGTG
CG59432-01 DNA Sequence TATACCAGGAAGATAATTCAATTCAATCTCTGGCTTACCCAAAGAATCTTGGAGTTAC
TGCCAATGAGGAAATCCCCAGGGTCTAATAAAAATATCTTTAGGAGTGAAGGAGTTAA
CTGAGTGTGTAAGCTTTATCTTCTGTCCAATGGACTTGTGGTTTGCTTATAAAACTCT
CCAGTAAATAATTGTTAGAGACCTGTCATTGATAGCAGTTGCTAGTTGCTGCCTTTTA
AGAGCTCGTTGATTCCTCTGCAAGGTGGTGCAGCATCCTCTGTCCCTTCATTCATTTC
AGATCTACTCAGGTCTCCCTGTAAACAGATCTCTCGGATCAATAAGCATGAATGACGA
AGACTACAGCACCATCTATGACACAATCCAAAATGAGAGGACGTATGAGGTTCCAGAC CAGCCAGAAGAAAATGAAAGTCCCCATTATGATGATGTCCATGAGTACTTAAGGCCAG AAAATGATTTATATGCCACTCAGCTGAATACCCATGAGTATGATTTTGTGTCAGTCTA TACCATTAAGGGTGAAGAGACCAGCTTGGCCTCTGTCCAGTCAGAAGACAGAGGCTAC
CTCCTGCCTGATGAGATATACTCTGAACTCCAGGAGGCTCATCCAGGTGAGCCCCAGG AGGACAGGGGCATCTCAATGGAAGGGTTATATTCATCAACCCAGGACCAGCAACTCTG CGCAGCAGAACTCCAGGAGAATGGGAGTGTGATGAAGGAAGATCTGCCTTCTCCTTCA AGCTTCACCATTCAGCACAGTAAGGCCTTCTCTACCACCAAGTATTCCTGCTATTCTG ATGCTGAAGGTTTGGAAGAAAAGGAGGGAGCTCACATGAACCCTGAGATTTACCTCTT TGTGAAGGTAAGGTCTGCCTCTGACAGGCATACCCTGTTCATGCAGATATTATGGCTG GTGTTTTATTTTGCTCTGAATGACCAGGGAAAGATTCATAATGCCATGGTCCTTGGAT CTCAATACATATTCAGGAGTCGGAGGGACTAAATCAGTCATTAGAGTGTACTCAGCTC TTCACAAAATTAGAGGAATTGGAAGGTGCATTTAAAGCACGTATTTAATCACTGACTT
TTACATACCATGGGCAAAGTATTTTTCAAAACGGTTCACATAAGTGAGCCATAACTGC
TGCCCAAATCCTTGCCATTGTGGCTGACATTAAGTACATTTTTCTGTCTGGTTAAATT
TCCTTTGTCGACATGTTTAAAAGTGAAACCAAAGCTTGTGAAAGAAAGACCTTCTTGT
GCTTCTAAGGTCACAGATTTGTCAGATAGGTGGTCAATAAAGGCTATCTCTGTCACTA
GCTTGCCCCTTTGGCACCAATATAACTAAAAATTTGATGAAGTCAAATGATTTCAGTA
GTAGTAAGACACTACCAGTGTTAATGTTTAATACTTACGATATCTAAACAGAA
ORF Start: ATG at 454 ORF Stop: TAA at 1132
SEQ ED NO: 138 226 aa MW at 26132.2kD
NOV44a, MNDEDYSTIYDTIQNERTYEVPDQPEENESPHYDDVHEYLRPENDLYATQLNTHEYDF VSVYTIKGEETSLASVQSEDRGYLLPDEIYSELQEAHPGEPQEDRGIS EGLYSSTQD
CG59432-01 Protein Sequence QQLCAAELQENGSVMKEDLPSPSSFTIQHSKAFSTTKYSCYSDAEGLEEKEGAHMNPE IYLFVKVRSASDRHTLFMQIL LVFYFALNDQGKIHNAMVLGSQYIFRSRRD
SEQ ID NO: 139 809 bp
NOV44b, ATCCTCTGTCCCTTCATTCATTTCAGATCTACTCAGGTCTCCCTGTAAACAGATCTCT
CGGATCAATAAGCATGAATGACGAAGACTACGGCACCATCTATGACACAATCCAAAAT
CG59432-02 DNA Sequence GAGAGGACGTATGAGGTTCCAGACCAGCCAGAAGAAAATGAAAGTCCCCATTATGATG ATGTCCATGAGTACTTAAGGCCAGAAAATGATTTATATGCCACTCAGCTGAATACCCA TGAGTATGATTTTGTGTCAGTCTATACCATTAAGGGTGAAGAGACCAGCTTGGCCTCT GTCCAGTCAGAAGACAGAGGCTACCTCCTGCCTGATGAGATATACTCTGAACTCCAGG AGGCTCATCCAGGTGAGCCCCAGGAGGACAGGGGCATCTCAATGGAAGGGTTATATTC ATCAACCCAGGACCAGCAACTCTGCGCAGCAGAACTCCAGGAGAATGGGAGTGTGATG AAGGAAGATCTGCCTTCTCCTTCAAGCTTCACCATTCAGCACAGTAAGGCCTTCTCTA CCACCAAGTATTCCTGCTATTCTGATGCTGAAGGTTTGGAAGAAAAGGAGGGAGCTCA CATGAACCCTGAGATTTACCTCTTTGTGAAGGTAAGGTCTGCCTCTGACAGGCATACC CTGTTCATGCAGATATTATGGCTGGTGTTTTATTTTGCTCTGAATGACCAGGGAAAGA TTCATAATGCCATGGTCCTTGGATCTCAATACATATTCAGGAGTCGGAGGGACTAAAT CAGTCATTAGAGTGTACTCAGCTCTTCACAAAATTAGAGGAATTGGAAGGTGCAT
ORF Start: ATG at 72 ORF Stop: TAA at 750
SEQ ID NO: 140 226 aa MW at 26102.2kD
NOV44b, MNDEDYGTIYDTIQNERTYEVPDQPEENESPHYDDVHEYLRPENDLYATQLNTHEYDF VSVYTIKGEETSLASVQSEDRGYLLPDEIYSELQEAHPGEPQEDRGISMEGLYSSTQD
CG59432-02 Protein Sequence QQLCAAELQENGSVMKEDLPSPSSFTIQHSKAFSTTKYSCYSDAEGLEEKEGAHMNPE IYLFVKVRSASDRHTLFMQIL LVFYFALNDQGKIHNAMVLGSQYIFRSRRD
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 44B.
Table 44B. Comparison of NOV44a against NOV44b.
NOV44a Residues/ Identities/
Protein Sequence Match Residues Similarities for the Matched Region
NOV44b 1..226 225/226 (99%) 1..226 225/226 (99%)
Further analysis of the NOV44a protein yielded the following properties shown in Table 44C.
Table 44C. Protein Sequence Properties NOV44a
PSort 0.6500 probability located in cytoplasm; 0.1000 probability located in analysis: mitochondrial matrix space; 0.1000 probability located in lysosome (lumen); 0.0000 probability located in endoplasmic reticulum (membrane)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV44a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 44D.
Table 44D. Geneseq Results for NOV44a
NOV44a
Identities/
Geneseq Protein/Organism/Length Residues/ Expect
Similarities for the Identifier [Patent #, Date] Match Value
Matched Region
Residues
No Significant Matches Found
In a BLAST search of public sequence databases, the NOV44a protein was found to have homology to the proteins shown in the BLASTP data in Table 44E.
PFam analysis predicts that the NOV44a protein contains the domains shown in the Table 44F.
Table 44F. Domain Analysis of NOV44a
Identities/
Pfam Domain NOV44a Match Region Similarities Expect Value for the Matched Region
No Significant Matches Found
Example 45.
The NOV45 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 45 A.
Further analysis of the NOV45a protein yielded the following properties shown in Table 45B.
Table 45B. Protein Sequence Properties NOV45a
PSort 0.6400 probability located in plasma membrane; 0.4600 probability located in analysis: Golgi body; 0.3700 probability located in endoplasmic reticulum (membrane); 0.1000 probability located in endoplasmic reticulum (lumen)
SignalP Likely cleavage site between residues 42 and 43 analysis:
A search of the NOV45a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 45 C.
In a BLAST search of public sequence databases, the NOV45a protein was found to have homology to the proteins shown in the BLASTP data in Table 45D.
PFam analysis predicts that the NOV45a protein contains the domains shown in the Table 45E.
Example 46.
The NOV46 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 46A.
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 46B.
Further analysis of the NOV46a protein yielded the following properties shown in Table 46C.
Table 46C. Protein Sequence Properties NOV46a
PSort 0.4500 probability located in cytoplasm; 0.3000 probability located in microbody analysis: (peroxisome); 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV46a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 46D.
In a BLAST search of public sequence databases, the NOV46a protein was found to have homology to the proteins shown in the BLASTP data in Table 46E.
PFam analysis predicts that the NOV46a protein contains the domains shown in the Table 46F.
Example 47.
The NOV47 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 47A.
Further analysis of the NOV47a protein yielded the following properties shown in Table 47B.
Table 47B. Protein Sequence Properties NOV47a
PSort 0.8500 probability located in endoplasmic reticulum (membrane); 0.4400 analysis: probability located in plasma membrane; 0.4244 probability located in microbody (peroxisome); 0.1000 probability located in mitochondrial inner membrane
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV47a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 47C.
In a BLAST search of public sequence databases, the NOV47a protein was found to have homology to the proteins shown in the BLASTP data in Table 47D.
PFam analysis predicts that the NOV47a protein contains the domains shown in the Table 47E.
Table 47E. Domain Analysis of NOV47a
Identities/
Pfam Domain NOV47a Match Region Similarities Expect Value for the Matched Region
No Significant Matches Found
Example 48.
The NOV48 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 48A.
Further analysis of the NOV48a protein yielded the following properties shown in Table 48B.
Table 48B. Protein Sequence Properties NOV48a
PSort 0.6400 probability located in microbody (peroxisome); 0.4500 probability analysis: located in cytoplasm; 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV48a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 48C.
In a BLAST search of public sequence databases, the NOV48a protein was found to have homology to the proteins shown in the BLASTP data in Table 48D.
PFam analysis predicts that the NOV48a protein contains the domains shown in the Table 48E.
Example 49.
The NOV49 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 49 A.
Table 49 A. NOV49 Sequence Analysis
SEQ ED NO: 151 1934 bp
NOV49a, CTTGATTACGGAGACTGAACCTTCATAGGGTGCGCACTTACCAAGGACAGGAAGGTTT
CTCTGTTTGAAGGGCTTTAAACTTATAACAAAGAAAATAAAAATGACGACTTCGTCTA
CG59377-01 DNA Sequence TCAGACGGCAGATGAAAAACATCGTGAACAATTACTCAGAGGCAGAAATCAAAGTCCG GGAAGCCACCTCCAATGACCCGTGGGGCCCGTCCAGTTCTCTGATGACCGAGATTGCC GACCTGACCTACAACGTGGTGGCCTTCTCGGAGATCATGAGCATGGTGTGGAAGCGGC TGAATGACCATGGCAAGAACTGGCGGCATGTGTACAAGGCGCTGACCCTGCTGGACTA CCTCATCAAGACAGGCTCCGAACGTGTGGCCCAGCAGTGCCGGGAGAACATCTTCGCC ATCCAGACCCTGAAGGACTTCCAGTACATTGACCGAGATGGCAAGGACCAGGGCATCA ATGTGCGTGAGAAGTCAAAGCAACTGGTGGCTCTCCTCAAGGACGAGGAACGGTTGAA GGCTGAGAGGGCCCAGGCTCTCAAAACCAAAGAGCGCATGGCCCAGGTTGCCACTGGC ATGGGCAGCAACCAGATCACCTTTGGGCGAGGCTCCAGCCAGCCCAACCTCTCCACCA GCCACTCGGAGCAGGAGTATGGCAAGGCCGGGGGCTCCCCGGCCTCCTACCATGGCTC CACCTCCCCGCGAGTGTCCTCCGAGCTGGAGCAAGCCCGGCCCCAGACTAGTGGAGAA GAGGAGCTTCAGCTGCAGCTGGCACTTGCCATGAGCAGAGAAGTGGCTGAGCAGGAAG AACGCCTCAGGCGGGGTGATGACCTCAGATTACAGATGGCCCTGGAAGAAAGCCGAAG GGACACAGTTAAAATTCCAAAAAAGAAAGAGCAGACTACGCTGTTGGATTTAATGGAT GCTCTCCCCAGCTCGGGCCCCGCGGCCCAGAAAGCAGAGCCCTGGGGCCCGTCAGCCT CCACTAACCAGACCAACCCCTGGGGCGGGCCAGCGGCTCCTGCGAGTACTTCAGACCC CTGGCCATCGTTTGGTACCAAGCCAGCTGCCTCCATTGACCCATGGGGGGTGCCCACT GGAGCCACCGCACAATCTGTCCCCAAGAACTCGGACCCCTGGGCAGCTTCACAGCAGC CTGCCTCCAGTGCTGGGAAAAGAGCTTCTGACGCGTGGGGCGCAGTCTCCACCACCAA GCCCGTGTCTGTCTCTGGGTCCTTTGAGCTCTTCAGTAATCTGAATGGTACAATTAAA
GATGACTTTTCTGAATTTGACAACCTTCGGACTTCAAAAAAAACAGCCGAATCTGTGA CCTCTCTGCCATCCCAAAACAATGGAACTACCAGCCCTGACCCCTTTGAGTCTCAACC CCTGACTGTCGCCTCAAGCAAGCCCAGCAGTGCCCGGAAAACACCTGAGTCCTTCCTG GGCCCCAACGCGGCCCTGGTGAACCTGGACTCACTGGTGACCAGGCCTGCCCCACCAG CCCAGTCCCTCAACCCTTTCCTGGCACCAGGTGCTCCCGCCACCTCGGCCCCTGTTAA CCCTTTCCAGGTGAACCAGCCCCAGCCGCTGACACTGAACCAGCTTCGGGGGAGCCCA GTCCTGGGGACCAGCACATCCTTTGGGCCTGGCCCAGGAGTGGAGTCCATGGCTGTGG CCTCGATGACCTCCGCGGCCCCACAGCCAGCTCTGGGGGCCACTGGTTCCTCTCTGAC ACCACTGGGCCCTGCAATGATGAACATGGTGGGCAGTGTGGGTATACCCCCATCAGCA GCCCAGGCCACTGGCACAACCAACCCTTTCCTTCTCTAGTGCCTGGGCCTGGGACCCA CCCAGAGCACCTGTGCTGGAGGATGCCGAGCAGGGACTCTCGTCTGTGGGACGGGATC
CAAGAGTTTGGGGATTAGGG
ORF Start: ATG at 101 ORF Stop: TAG at 1835
SEQ ED NO: 152 578 aa MW at 61651.2kD
NOV49a, MTTSSIRRQMKNIVNNYSEAEIKVREATSNDP GPSSSLMTEIADLTYNWAFSEI S MVWKRLNDHGKN RHVYKALTLLDYLIKTGSERVAQQCRENIFAIQTLKDFQYIDRDG
CG59377-01 Protein Sequence KDQGINVREKSKQLVALLKDEERLKAERAQALKTKERMAQVATGMGSNQITFGRGSSQ PNLSTSHSEQEYGKAGGSPASYHGSTSPRVSSELEQARPQTSGEEELQLQLALAMSRE VAEQEERLRRGDDLRLQMALEESRRDTVKIPKKKEQTTLLDLMDALPSSGPAAQKAEP WGPSASTNQTNPWGGPAAPASTSDPWPSFGTKPAASIDPWGVPTGATAQSVPKNSDPW AASQQPASSAGKRASDAWGAVSTTKPVSVSGSFELFSNLNGTIKDDFSEFDNLRTSKK TAESVTSLPSQNNGTTSPDPFESQPLTVASSKPSSARKTPESFLGPNAALVNLDSLVT RPAPPAQSLNPFLAPGAPATSAPVNPFQVNQPQPLTLNQLRGSPVLGTSTSFGPGPGV ESMAVASMTSAAPQPALGATGSSLTPLGPAMMNMVGSVGIPPSAAQATGTTNPFLL
Further analysis of the NOV49a protein yielded the following properties shown in Table 49B.
Table 49B. Protein Sequence Properties NOV49a
PSort 0.4936 probability located in mitochondrial matrix space; 0.3000 probability analysis: located in nucleus; 0.2087 probability located in mitochondrial inner membrane; 0.2087 probability located in mitochondrial intermembrane space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV49a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 49C.
In a BLAST search of public sequence databases, the NOV49a protein was found to have homology to the proteins shown in the BLASTP data in Table 49D.
PFam analysis predicts that the NOV49a protein contains the domains shown in the Table 49E.
Table 49E. Domain Analysis of NOV49a
Pfam Domain NOV49a Match Region Expect Value
Example 50.
The NOV50 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 50A.
Table 50A. NOV50 Sequence Analysis
SEQ ED NO: 153 2580 bp
NOV50a, ATGCTGCTGGCCCCCTTTTATTGCTGGGTGTGTGCCCATGCTGCTGGCCCCCTTTTAT TGCTGGGCAGTGACAAACTGTACCATCAGTGGCTCTCCACTGTCCGGAAAGGAAGTGG
CG59258-01 DNA Sequence AGCAATTCTGAATACTGTAAAGACCAAAGCAAATCCGGCCATGAAGACTGTCTACAAG TTCGACATTGCCGAGAATGGCTGCGCCCCCACCCCAGAAGAGCAGCTGCCAAAGACTG CACCGTCCCCACTGGTGGAGGCCAAGGACCCCAAGCTCCGAGAAGACCGGCGGCCAAT CACAGTCCACTTTGGACAGGACCAGTCTGAGATGTCTTTCAGCTCAGCACTCACTCAC GGCAAAGAGAGTGCCCGGACCCAGCCGGAGAGAGTCGTTGACAGGACTGGCGAGCCCC TGAATCCTGAGCGCGCTCTCTCCGGAGATCATCTCTGGCCTGTTACGCACTTGCTCTG GGCAACCCTGGGCAAGTCCTTGCTTGCCCTCATCTGTGAAATGGGTAGCAGCCCTCGT TCCCTGCAGAGGAGCCTTGCGCTGCTGGGGACACCCCAGCTTATTTGGGAAACTGCAA CCACCATGGCCGATGGCCCCACCACGCCCTGTCTAGGAAGCAGAGGCCTCCCCAGCAG CGTGTCCACTGTGCCCCTGGCCCTGCGTGAAGTGCCATCAGATGCCCCGCATCCCTGC AGCAGGGCCCTCGTGACTGGCCTCACAGATGAGGACACAGAGGCCCAGGGAAGTCACT TGCTTGCCAAAGTCACTCAGCAAACCATGTCTGTCTGGCTCTCAGAAAATGGGAAAGA AGCCTGGGCATTCAGCCATGAGGGAGCCACGGCTGTAGCCAGTGGAATGACGTACCCT CAGTCCAGGATGTGCACCCGGGCAGCCAGGTCCCACAGCCACTACTTTCTTGCCCCCA CCACTGCTCCCACAGTTCCCAGAACTCAGTCTCCAGATCTGGGCTCCAGGATGCAGAG GCTGTCCTCAGGCCTGGTAAAGCCCTTGCGACACTATGCGGTCTTCCTCTCCGAAGAC TCCTCTGATGATGAATGCCAGCGGGAAGAGGGCCCGAGCTCTGGCTTCACCGAGAGCT TTTTCTTCTCCGCTCCCTTTGAATGGCCGCAGCCGTATCGGACACTCAGGGAGTCAGA CAGCGCGGAAGGCGACGAGGCAGAGAGTCCAGAGCAGCAAGTGCGGAAGTCCACAGGC CCTGTCCCAGCTCCCCCTGACCGGGCTGCCAGCATCGACCTTCTGGAAGACGTCTTCA GCAACCTGGACATGGAGGCCGCACTGCAGCCACTGGGCCAGGCCAAGAGCTTAGAGGA CCTTCGTGCCCCCAAAGACCTGAGGGAGCAGCCAGGGACCTTTGACTATCAGAGGCTG GATCTGGGCGGGAGTGAGAGGAGCCGCGGGGTGACAGTGGCCTTGAAGCTTACCCACC CGTACAACAAGCTCTGGAGCCTGGGCCAGGACGACATGGCCATCCCCAGCAAGCCCCC AGCTGCCTCCCCTGAGAAGCCCTCAGCCCTGCTCGGAAACTCCCTGGCCCTGCCTCGA AGGCCCCAGAACCGGGACAGCATCCTGAACCCCAGTGACAAGGAGGAGGTGCCCACCC CTACTCTGGGCAGCATCACCATCCCCCGGCCCCAAGGCAGGAAGACCCCAGAGCTGGG CATCGTGCCTCCACCGCCCATTCCCCGCCCGGCCAAGCTCCAGGCTGCCGGCGCCGCA CTTGGTGACGTCTCAGAGCGGCTGCAGACGGATCGGGACAGGCGAGCTGCCCTGAGTC CAGGGCTCCTGCCTGGTGTTGTCCCCCAAGGCCCCACTGAACTGCTCCAGCCGCTCAG CCCTGGCCCCGGGGCTGCAGGCACGAGCAGTGACGCCCTGCTCGCCCTCCTGGACCCG CTCAGCACAGCCTGGTCAGGCAGCACCCTCCCGTCACGCCCCGCCACCCCGAATGTAG CCACCCCATTCACCCCCCAATTCAGCTTCCCCCCTGCAGGGACACCCACCCCATTCCC ACAGCCACCACTCAACCCCTTTGTCCCATCCATGCCAGCAGCCCCACCCACCCTGCCC CTGGTCTCCACACCAGCCGGGCCTTTCGGGGCCCCTCCAGCTTCCCTGGGGCCGGCTT TTGCGTCCGGCCTCCTGCTGTCCAGTGCTGGCTTCTGTGCCCCTCACAGGTCTCAGCC CAACCTCTCCGCCCTCTCCATGCCCAACCTCTTTGGCCAGATGCCCATGGGCACCCAC ACGAGCCCCCTACAGCCGCTGGGTCCCCCAGCAGTTGCCCCGTCGAGGATCCGAACGT TGCCCCTGGCCCGCTCAAGTGCCAGGGCTGCTGAGACCAAGCAGGGGCTGGCCCTGAG GCCTGGAGACCCCCCGCTTCTGCCTCCCAGGCCCCCTCAAGGCCTGGAGCCAACACTG CAGCCCTCTGCTCCTCAACAGGCCAGAGACCCCTTTGAGGATTTGTTACAGAAAACCA AGCAAGACGTGAGCCCGAGTCCGGCCCTGGCCCCGGCCCCAGACTCGGTGGAGCAGCT
Further analysis of the NOV50a protein yielded the following properties shown in Table 50B.
Table 50B. Protein Sequence Properties NOV50a
PSort 0.4500 probability located in cytoplasm; 0.3000 probability located in microbody analysis: (peroxisome); 0.1940 probability located in lysosome (lumen); 0.1000 probability located in mitochondrial matrix space
SignalP Likely cleavage site between residues 15 and 16 analysis:
A search of the NOV50a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 50C.
In a BLAST search of public sequence databases, the NOV50a protein was found to have homology to the proteins shown in the BLASTP data in Table 50D.
PFam analysis predicts that the NOV50a protein contains the domains shown in the Table 50E.
Table 50E. Domain Analysis of NOV50a
Pfam Domain NOV50a Match Region Expect Value
Similarities for the Matched Region
No Significant Matches Found
Example 51.
The NOV51 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 51 A.
Table 51A. NOV51 Sequence Analysis
SEQ ED NO: 155 1394 bp
NOV51a, GTGGCCTGCTCCTGCAGCAATCCCAGGACCCCCTGCTCATGGGGCTGTTTCCTACTAA
CCCCAAAGAGAAGACCCAGGAGGAACCCCCTGGCCAGAGCAGGGCCCCTGTGTTGACC
CG59492-01 DNA Sequence GTGGTGTCCAAGTTCAAGGCCTCACTGGAGCAGCTTCTGCAGGTCCTACACAGCACCA CGCCCCACTACATTCGCTGCATCAAGCCCAACAGCCAGGGCCAGGCGCAGACCTTTCT CCAAGAGGAGGTCCTGAGCCAGCTGGAGGCCTGTGGCCTCGTGGAGACCATCCATATC AGTGCTGCTGGCTTCCCCATCCGGGTCTCTCACCGAAACTTTGTAGAACGATACAAGT TACTAAGAAGGCTTCATCCTTGCACATCCTCTGGCCCCGACAGCCCATATCCTGCCAA AGGGCTCCCTGAATGGTGTCCACACAGCGAGGAAGCCACGCTTGAACCTCTCATCCAG GACATTCTCCACACTCTGCCGGTCCTAACTCAGGCAGCAGCCATAACTGGTGACTCGG CTGAGGCCATGCCAGCCCCCATGCACTGTGGCAGGACCAAGGTGTTCATGACTGACTC TATGCTGGAGCTTCTGGAATGTGGGCGTGCCCGGGTGCTGGAGCAGTGTGCCCGCTGC ATCCAGGGTGGCTGGAGGCGACACCGGCACCGAGAGCAGGAGCGGCAGTGGCGGGCCG TCATGCTCATCCAGGCAGCCATTCGTTCCTGGTTAACTCGGAAACACATCCAGAGGCT GCATGCAGCTGCCACAGTCATCAAGCGTGCATGGCAGAAGTGGAGAATCAGAATGGCC TGCCTTGCTGCTAAAGAGCTGGATGGTGTGGAAGAAAAACACTTCTCTCAAGCTCCCT GTTCCCTGAGCACCTCGCCGCTGCAGACCAGGCTCCTGGAGGCAATAATCCGCTTCTG GCCCCTGGGACTGGTCCTGGCCAATACGGCTATGGGTGTAGGCAGCTTTCAGAGGAAA TTAGTGGTCTGGGCTTGCCTCCAGCTCCCCAGGGGCAGCCCCAGTAGCTACACTGTCC AGACAGCACAAGACCAGGCTGGTGTCACGTCCATCCGAGCGCTGCCTCAGGGATCGAT AAAGTTTCACTGCAGAAAGTCTCCACTGCGGTATGCTGACATCTGCCCTGAACCTTCA CCCTACAGCATTGCAGGCTTTAATCAGATTCTGCTGGAAAGACACAGGCTGATCCACG TGACCTCTTCTGCCTTCACTGGGCTGGGGTGATCCTTGGTGCCTTTGTTTCCACAAGG CCTTTTCCTGCCCCCTGCCTTGCCAAAGACATTTAATCAGCACACAGCTGCCAGACTA
TTCCCACAGTGCTCCAAATGCACATGAACAACAGTGACGGCTCCAGCCTTCGACCCAG
AG
ORF Start: ATG at 39 ORF Stop: TGA at 1248
SEQ ED NO: 156 403 aa MW at 45142.8kD
NOV51a, MGLFPTNPKEKTQEEPPGQSRAPVLTWSKFKASLEQLLQVLHSTTPHYIRCIKPNSQ GQAQTFLQEEVLSQLEACGLVETIHISAAGFPIRVSHRNFVERYKLLRRLHPCTSSGP
CG59492-01 Protein Sequence DSPYPAKGLPE CPHSEEATLEPLIQDILHTLPVLTQAAAITGDSAEAMPAPMHCGRT KVFMTDSMLELLECGRARVLEQCARCIQGG RRHRHREQERQ RAVMLIQAAIRS LT RKHIQRLHAAATVIKRAWQK RIR ACLAAKELDGVEEKHFSQAPCSLSTSPLQTRLL EAIIRF PLGLVLANTAMGVGSFQRKLWWACLQLPRGSPSSYTVQTAQDQAGVTSIR ALPQGSIKFHCRKSPLRYADICPEPSPYSIAGFNQILLERHRLIHVTSSAFTGLG
Further analysis of the NOV51a protein yielded the following properties shown in Table 5 IB.
Table 51B. Protein Sequence Properties NOV51a
PSort 0.3000 probability located in nucleus; 0.2029 probability located in lysosome analysis: (lumen); 0.1000 probability located in mitochondrial matrix space; 0.0320 probability located in microbody (peroxisome)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV5 la protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 51C.
In a BLAST search of public sequence databases, the NOV5 la protein was found to have homology to the proteins shown in the BLASTP data in Table 5 ID.
PFam analysis predicts that the NOV5 la protein contains the domains shown in the Table 5 IE.
Example 52.
The NOV52 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 52A.
Table 52A. NOV52 Sequence Analysis
SEQ ID NO: 157 1380 bp
NOV52a, TAGAATTCCAGCGGCCGCTGAAATCCTCACTCGGTCAGTTCCTCGGGCGAGTTACGGG
GACGACCTGCGGGAGCACGCGGGCAGTGGCCGGACGCTGAAGCCCAGGAGAGCGATGG
CG59564-01 DNA Sequence AGACGTATGCGGAGGTTGGGAAGGAGGGCAAGCCTTCCTGTGCATCGGTGGATCTGCA GGGAGACAGCTCCTTACAGGTGGAGATTTCTGACGCAGTGAGTGAGCGGGACAAGGTG AAATTCACTGTTCAAACAAAGAGCTGCCTCCCTCACTTCGCCCAGACCGAGTTCTCAG TCGTGCGGCAGCACGAGGAGTTCATCTGGCTGCATGATGCCTACGTGGAGAATGAGGA GTACGCCGGCCTCATCATCCCCCCAGCCCCTCCGAGGCCAGACTTTGAGGCTTCGAGG GAAAAGCTACAGAAATTGGGCGAGGGGGACAGCTCTGTCACTCGGGAAGAGTTTGCCA AGATGAAGCAGGAGCTGGAAGCGGAGTACCTGGCCATCTTTAAGAAGACAGTTGCGAT GCACGAAGTCTTTCTGCAGCGCCTGGCGGCCCACCCCACCCTGCGTCGAGACCACAAC TTCTTTGTGTTTTTGGAATATGGACAGGATCTGAGTGTCCGGGGGAAGAACAGGAAGG AGCTCCTCGGAGGGTTTCTGAGGAATATTGTGAAGTCCGCGGATGAAGCCCTCATCAC GGGCATGTCAGGGCTCAAGGAGGTGGATGACTTCTTTGAGCATGAGAGGACCTTCCTG TTGGAGTATCACACCCGTATCCGAGATGCCTGCCTGCGGGCCGACCGCGTCATGCGCG CCCACAAGTGCCTGGCAGACGATTATATCCCTATCTCAGCTGCGCTGAGCAGTCTGGG AACACAGGAAGTCAACCAGCTAAGGACGAGCTTCCTCAAATTGGCAGAGCTCTTTGAC CGGCTGAGGAAGCTGGAGGGCCGGGTGGCTTCCGATGAGGACCTGAAGCTGTCAGACA TGCTGAGGTACTACATGCGTGACTCACAGGCAGCCAAGGACCTGCTGTACCGGCGGCT GCGGGCACTGGCCGACTACGAGAATGCCAACAAGGCGCTGGACAAGGCGCGCACCAGG AACCGGGAGGTGCGGCCCGCCGAGAGCCACCAGCAGCTGTGCTGCCAACGCTTCGAGC GCCTCTCCGACTCCGCCAAGCAAGAGCTCATGGACTTCAAGTCCCGCCGGGTCTCCTC TTTTCGAAAGAATCTCATTGAGCTGGCAGAGCTGGAGCTCAAACACGCCAAGGCCAGC ACCCTGATTCTCCGGAACACCCTTGTTGCCCTAAAGGGGGAGCCTTAGAGTAGCCAGA GCTCAGCCAGACCCTAATCTGGGATCTCCAGTGACCAGGGTATCCC
ORF Start: ATG at 113 ORF Stop: TAG at 1322
SEQ ID NO: 158 403 aa MW at 46384.2kD
NOV52a, METYAEVGKEGKPSCASVDLQGDSSLQVEISDAVSERDKVKFTVQTKSCLPHFAQTEF SWRQHEEFI LHDAYVENEEYAGLIIPPAPPRPDFEASREKLQKLGEGDSSVTREEF
CG59564-01 Protein Sequence AKMKQELEAEYLAIFKKTVAMHEVFLQRLAAHPTLRRDHNFFVFLEYGQDLSVRGKNR KELLGGFLRNIVKSADEALITGMSGLKEVDDFFEHERTFLLEYHTRIRDACLRADRVM RAHKCLADDYIPISAALSSLGTQEVNQLRTSFLKLAELFDRLRKLEGRVASDEDLKLS DMLRYYMRDSQAAKDLLYRRLRALADYENANKALDKARTRNREVRPAESHQQLCCQRF ERLSDSAKQELMDFKSRRVSSFRKNLIELAELELKHAKASTLILRNTLVALKGEP
Further analysis of the NOV52a protein yielded the following properties shown in Table 52B.
Table 52B. Protein Sequence Properties NOV52a
PSort 0.6500 probability located in cytoplasm; 0.1000 probability located in analysis: mitochondrial matrix space; 0.1000 probability located in lysosome (lumen); 0.0000 probability located in endoplasmic reticulum (membrane)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV52a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 52C.
In a BLAST search of public sequence databases, the NOV52a protein was found to have homology to the proteins shown in the BLASTP data in Table 52D.
PFam analysis predicts that the NOV52a protein contains the domains shown in the Table 52E.
Example 53.
The NOV53 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 53 A.
Table 53A. NOV53 Sequence Analysis
SEQ ED NO: 159 3056 bp
NOV53a, CTCCTGCGGGGTCAAATACAGAATTTACGCACCCTTGGCTTCCTTGGAGCCTAGCGGC
TCTCCCCGCGTCCAAGATGGCGGCAGAAGCAGCTGGTGGGAAATACAGAAGCACAGTC
CG59553-01 DNA Sequence AGCAAAAGCAAAGACCCCTCGGGGCTGCTCATCTCTGTGATCAGGACTCTGTCTACTA GTGACGATGTCGAAGACAGGGAAAATGAAAAGGGTCGCCTTGAAGAAGCCTACGAGAA ATGTGACCGTGACCTGGATGAATTGATTGTACAGCACTACACAGAATTGACGACAGCC ATTCGCACATACCAGAGCATCACAGAGCGCATCACTAACTCCCGAAATAAAATAAAGC AGGTAAAAGAGAACCTGCTTTCATGCAAGATGCTGCTGCACTGCAAACGGGATGAGCT TCGGAAACTGTGGATTGAAGGAATTGAGCATAAGCATGTCCTGAACTTGTTGGATGAA ATTGAGAATATCAAGCAAGTGCCTCAAAAGCTGGAACAGTGCATGGCCAGCAAGCACT ATCTCAGTGCCACTGACATGTTGGTGTCAGCAGTTGAGTCTTTGGAGGGCCCCCTGCT CCAGGTGGAAGGACTGAGTGACCTTCGACTAGAGCTTCACAGCAAGAAGATGAACCTT CACTTGGTTCTCATAGATGAACTACACCGGCACCTGTACATCAAATCGACTAGCCGAG TTGTGCAGCGTAACAAGGAAAAAGGGAAAATCAGCTCCCTCGTGAAAGATGCTTCTGT TCCTCTGATTGATGTTACAAACCTCCCTACTCCTCGAAAATTCCTTGATACCTCTCAC TATTCTACTGCTGGAAGCTCAAGTGTGAGGGAGATAAATCTGCAGGACATCAAGGAAG ATTTAGAATTGGATCCAGAGGAAAACAGCACCCTGTTTATGGGTATCCTCATTAAGGG CTTGGCGAAACTGAAGAAGATCCCAGAAACAGTTAAGGCAATCATAGAGCGCTTGGAG CAGGAGTTGAAGCAAATTGTGAAGAGGTCTACAACCCAGGTGGCAGACAGTGGCTATC AGCGGGGGGAGAACGTTACTGTGGAGAACCAACCAAGGTTGCTTCTAGAACTGCTGGA GTTACTGTTTGACAAGTTTAATGCTGTAGCCGCTGCACACTCTGTGGTCCTGGGATAC CTGCAGGACACTGTAGTGACTCCACTGACTCAGCAGGAAGATATCAAACTGTATGATA TGGCAGATGTATGGGTGAAGATCCAAGATGTTCTACAGATGCTATTAACTGAGTACTT GGATATGAAAAATACTCGTACGGCCTCTGAACCATCAGCTCAACTAAGCTATGCCAGC ACTGGACGAGAGTTTGCAGCCTTTTTTGCCAAGAAGAAACCTCAAAGGCCAAAAAATT CTCTTTTCAAGTTCGAATCGTCCTCCCATGCCATCAGTATGAGCGCCTATCTGCGAGA ACAGAGAAGGGAGCTCTATAGTCGGAGTGGAGAACTGCAAGGGGGTCCTGATGACAAC TTAATTGAAGGTGGAGGAACAAAATTTGTCTGCAAACCTGGAGCCAGAAACATTACCG TCATATTCCACCCATTACTAAGATTTATTCAGGAGATTGAGCATGCTCTGGGTCTTGG CCCAGCCAAACAGTGTCCTCTTCGAGAGTTTCTCACCGTGTACATCAAAAACATCTTT CTCAATCAAGTCTTGGCTGAGATCAACAAGGAGATTGAAGGAGTCACTAAAACATCTG
ACCCTTTGAAGATTCTGGCCAACGCAGACACCATGAAGGTGCTGGGAGTGCAGCGGCC TCTCCTACAGAGCACAATCATTGTGGAGAAGACAGTTCAAGACCTCCTGAACCTGATG CATGACTTGAGTGCATATTCAGATCAATTCCTCAACATGGTGTGCGTGAAGCTCCAGG AGTACAAGGACACCTGCACTGCAGCTTACAGGGGTATTGTCCAGTCAGAAGAAAAACT TGTCATCAGTGCATCCTGGGCAAAAGATGATGATATCAGCAGACTCTTGAAATCTCTA CCAAACTGGATGAATATGGCTCAACCCAAACAGCTGAGGCCAAAAAGAGAGGAGGAAG AAGATTTCATAAGGGCAGCTTTTGGCAAGGAGTCTGAAGTTCTTATTGGGAACCTGGG TGATAAATTAATCCCTCCACAAGACATCCTTCGTGACGTCAGTGACCTCAAAGCCTTG GCCAACATGCATGAAAGCCTGGAATGGTTGGCAAGTCGAACAAAGTCAGCTTTCTCCA ATCTTTCTACATCCCAGATGCTTTCTCCTGCTCAAGACAGCCACACGAACACGGATCT CCCCCCAGTGTCAGAGCAGATCATGCAGACTCTCAGTGAACTTGCCAAATCGTTCCAG GATATGGCTGACCGCTGCTTGCTTGTCTTACATCTGGAAGTGAGGGTTCACTGTTTCC ACTATCTTATCCCTCTTGCAAAGGAGGGGAACTATGCCATTGTGGCTAATGTGGAAAG TATGGATTATGACCCCCTGGTGGTCAAGCTCAACAAAGATATCAGCGCCATTGAAGAG GCCATGAGCGCCAGCCTTCAGCAGCACAAGTTCCAGTATATCTTCGAAGGCCTGGGCC ACCTGATCTCCTGCATCCTCATTAATGGTGCCCAGTACTTCAGGCGCATCAGTGAGTC TGGCATCAAGAAAATGTGTAGGAACATTTTTGTTCTTCAGCAGAATTTGACCAACATC ACCATGTCGCGGGAGGCAGACCTGGACTTTGCAAGGCAGTACTACGAGATGCTTTACA ACACAGCTGACGAGCTCCTGAACCTGGTGGTGGACCAGGGTGTGAAGTACACGGAGCT GGAGTACATCCACGCTCTGACCCTGCTGCACCGCAGCCAGACTGGGGTGGGGGAACTG ACCACCCAGAACACGAGCTGCAGAGGAGGCTCAAAGAGATCATCTGCGAGCAGGCTGC CATCAAGCAAGCCACCAAGGACAAGAAGATAACTACCGTTTAGCAGGGCGTACTGCGG TTGGTGACGGGGGTCCCCTCAGTCACACTCACTTTTTTCC
ORF Start: ATG at 75 ORF Stop: TAA at 2988
SEQ ED NO: 160 971 aa MW at l09984.9kD
NOV53a, MAAEAAGGKYRSTVSKSKDPSGLLISVIRTLSTSDDVEDRENEKGRLEEAYEKCDRDL DELIVQHYTELTTAIRTYQSITERITNSRNKIKQVKENLLSCKMLLHCKRDELRKL I
CG59553-01 Protein Sequence EGIEHKHVLNLLDEIENIKQVPQKLEQCMASKHYLSATDMLVSAVESLEGPLLQVEGL SDLRLELHSKKMNLHLVLIDELHRHLYIKSTSRWQRNKEKGKISSLVKDASVPLIDV TNLPTPRKFLDTSHYSTAGSSSVREINLQDIKEDLELDPEENSTLFMGILIKGLAKLK KIPETVKAIIERLEQELKQIVKRSTTQVADSGYQRGENVTVENQPRLLLELLELLFDK FNAVAAAHSWLGYLQDTWTPLTQQEDIKLYDMADV VKIQDVLQMLLTEYLDMKNT RTASEPSAQLSYASTGREFAAFFAKKKPQRPKNSLFKFESSSHAISMSAYLREQRREL YSRSGELQGGPDDNLIEGGGTKFVCKPGARNITVIFHPLLRFIQEIEHALGLGPAKQC PLREFLTVYIKNIFLNQVLAEINKEIEGVTKTSDPLKILANADTMKVLGVQRPLLQST IIVEKTVQDLLNLMHDLSAYSDQFLNMVCVKLQEYKDTCTAAYRGIVQSEEKLVISAS AKDDDISRLLKSLP MNMAQPKQLRPKREEEEDFIRAAFGKESEVLIGNLGDKLIP PQDILRDVSDLKALANMHESLE LASRTKSAFSNLSTSQMLSPAQDSHTNTDLPPVSE QIMQTLSELAKSFQDMADRCLLVLHLEVRVHCFHYLIPLAKEGNYAIVANVESMDYDP LWKLNKDISAIEEAMSASLQQHKFQYIFEGLGHLISCILINGAQYFRRISESGIKKM CRNIFVLQQNLTNITMSREADLDFARQYYEMLYNTADELLNLWDQGVKYTELEYIHA LTLLHRSQTGVGELTTQNTSCRGGSKRSSASRLPSSKPPRTRR
Further analysis of the NOV53a protein yielded the following properties shown in Table 53B.
Table 53B. Protein Sequence Properties NOV53a
PSort 0.5500 probability located in endoplasmic reticulum (membrane); 0.1900 analysis: probability located in lysosome (lumen); 0.1000 probability located in endoplasmic reticulum (lumen); 0.1000 probability located in outside
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV53a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 53C.
In a BLAST search of public sequence databases, the NOV53a protein was found to have homology to the proteins shown in the BLASTP data in Table 53D.
PFam analysis predicts that the NOV53a protein contains the domains shown in the Table 53E.
Table 53E. Domain Analysis of NOV53a
Identities/
Pfam Domain NOV53a Match Region Similarities Expect Value for the Matched Region
No Significant Matches Found
Example 54.
The NOV54 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 54A.
Further analysis of the NOV54a protein yielded the following properties shown in Table 54B.
Table 54B. Protein Sequence Properties NOV54a
PSort 0.5500 probability located in endoplasmic reticulum (membrane); 0.1900 analysis: probability located in lysosome (lumen); 0.1000 probability located in endoplasmic reticulum (lumen); 0.1000 probability located in outside
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV54a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 54C.
In a BLAST search of public sequence databases, the NOV54a protein was found to have homology to the proteins shown in the BLASTP data in Table 54D.
PFam analysis predicts that the NOV54a protein contains the domains shown in the Table 54E.
Example 55.
The NOV55 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 55 A.
ACTGGGAAAAGTAGTTTAGGTGACATGTTCTCACCTATCAGAGATGATGCTGTAGTTA ACAAGGGAAGTGATGAGTCCATAGGCAAAGGAGATGGCTTTGACTTTCTACCGCAGTT GAACTCAGTGTTTCCTCCAAGAAAAAATCCAGTAACTTCAAGTACTTCAGTATTGCAT TCTAGTCCTCTTAATGTTTTTATGGGATCTCCAGGGAAAGAGGAAAATGAAAACCGTG ATCTAACAGCTGAGTCTAAGAAAATATATATGGGAAAACAGGAATCTAAAGACTCCTT CAAACAGTTAGCAAAGTTGGTCACATCTGGTGCTGAAAGTGGAAATCTAAATACCTCT CCATCATCTAACCAAACAAGAAATTCTGAGAAATTTGAAAAGCCAGAGAATGAAATTG AAGCCCAGTTGATATGTGAACCCCCAATCAATGGATCCTCAACTCCAAATCCAAAGAT AGCATCTTCTGTCACTGCTGGAGTTGCCAGTTCACTCTCAGAAAAAATAGCCGACAGC ATTGGAAATAACCGGCAAAATGCACCATTGACTTCCATTCAAATTCGTTTTATTCAGA ACATGATACAGGAAACGTTGGATGACTTTAGAGAAGCATGCCATAGGGACATTGTGAA TTTGCAAGTGGAGATGATTAAACAGTTTCATATGCAACTGAATGAAATGCATTCTTTG CTGGAAAGATACTCAGTGAATGAAGGTTTAGTGGCTGAAATTGAAAGACTACGAGAAG AAAACAAAAGATTACGGGCCCACTTTTGAAATTTCAGTGAATACCTTAATGTTCTGTA
ATTTGGGAAGTTTCTGGCAACACAGAACTACATAGAATCAT
ORF Start: ATG at 22 ORF Stop: TGA at 1999
SEQ ED NO: 164 659 aa MW at 71851.2kD
NOV55a, MQENLRFASSGDDIKI DASSMTLVDKFNPHTSPHGISSICWSSNSNFLVTASSSGDK IWSSCKCKPVPLLELAEGQKQTCVNLNSTSMYLVSGGLNNTVNIWDLKSKRVHRSLK
CG59435-01 Protein Sequence DHKDQVTCVTYNNDCYIASGSLSGEIILHSVTTNLSSTPFGHGSNQVRHLKYSLFKK SLLGSVSDNGIVTL DVNSQSPYHNFDSVHKAPASGICFSPVNELLFVTIGLDKRIIL YDTSSKKLVKTLVADTPLTAVDFMPDGATLAIGSSRGKIYQYDLRMLKSPVKTISAHK TSVQCIAFQYSTVLTKSSLNKGCSNKPTTVNKRSVNVNAASGGVQNSGIVREAPATSI ATVLPQPMTSA GKGTVAVQEKAGLPRSINTDTLSKETDSGKNQDFSSFDDTGKSSLG DMFSPIRDDAWNKGSDESIGKGDGFDFLPQLNSVFPPRKNPVTSSTSVLHSSPLNVF MGSPGKEENENRDLTAESKKIYMGKQESKDSFKQLAKLVTSGAESGNLNTSPSSNQTR NSEKFEKPENEIEAQLICEPPINGSSTPNPKIASSVTAGVASSLSEKIADSIGNNRQN APLTSIQIRFIQNMIQETLDDFREACHRDIVNLQVEMIKQFHMQLNEMHSLLERYSVN EGLVAEIERLREENKRLRAHF
SEQ ED NO: 165 2009 bp
NOV55b, AAACTATTTGTAGGCGCAGTCATGCAGGAAAACCTCAGATTTGCTTCATCAGGAGATG
ATATTAAAATATGGGATGCTTCATCTATGACATTGGTGGATAAATTCAACCCACACAC
CG59435-02 DNA Sequence ATCACCACATGGAATCAGCTCAATATGTTGGAGCAGCAATAATAACTTTTTAGTAACA GCATCTTCCAGTGGCGACAAAATAGTTGTCTCAAGTTGCAAATGTAAACCTGTTCCAC TTTTAGAGCTTGCTGAAGGGCAAAAGCAGACATGTGTCAATTTAAATTCTACATCTAT GTATTTGGTAAGCGGAGGCCTAAATAACACTGTTAATATTTGGGATTTAAAATCAAAA AGAGTTCATCGATCTCTTAAGGATCATAAAGATCAAGTAACTTGTGTAACATACAATT GGAATGATTGCTACATTGCTTCTGGATCTCTTAGTGGTGAAATTATTTTACACAGTGT AACCACTAATTTATCTAGTACTCCTTTTGGCCATGGTAGTAACCAGTCTGTTCGGCAC TTGAAGTACTCCTTGTTTAAGAAATCACTACTGGGCAGTGTTTCGGATAATGGAATAG TAACTCTCTGGGATGTAAATAGTCAGAGTCCATACCATAACTTTGACAGTGTACACAA AGCTCCAGCGTCAGGCATCTGTTTTTCTCCTGTCAATGAATTGCTCTTTGTAACCATA GGCTTGGATAAAAGAATCATCCTCTATGACACTTCAAGTAAGAAGCTAGTGAAAACTT TAGTGGCTGACACTCCTCTAACTGCGGTAGATTTCATGCCTGATGGAGCCACTTTGGC TATTGGATCTTCCCGGGGGAAAATATATCAATATGATTTAAGAATGTTGAAATCACCA GTTAAGACCATCAGTGCTCACAAGACATCTGTGCAGTGTATAGCATTTCAGTACTCCA CTGTTCTTACTAAGTCAAGTTTAAATAAAGGCTGTTCAAATAAGCCCACAACAGTGAA CAAACGAAGTGTTAATGTGAATGCTGCTAGTGGAGGAGTTCAGAATTCCGGAATTGTC AGAGAAGCACCTGCCACGTCCATTGCCACAGTTCTACCACAACCTATGACATCAGCTA TGGGGAAAGGAACAGTTGCTGTTCAAGAAAAAGCAGGTTTGCCTCGAAGCATAAACAC AGACACTTTATCTAAGGAAACAGACAGTGGAAAAAATCAGGATTTCTCCAGCTTTGAT GATACTGGGAAAAGTAGTTTAGGTGACATGTTCTCACCTATCAGAGATGATGCTGTAG TTAACAAGGGAAGTGATGAGTCCATAGGCAAAGGAGATGGCTTTGACTTTCTACCGCA GTTGAACTCAGTGTTTCCTCCAAGAAAAAATCCAGTAACTTCAAGTACTTCAGTATTG CATTCTAGTCCTCTTAATGTTTTTATGGGATCTCCAGGGAAAGAGGAAAATGAAAACC GTGATCTAACAGCTGAGTCTAAGAAAATATATATGGGAAAACAGGAATCTAAAGACTC CTTCAAACAGTTAGCAAAGTTGGTCACATCTGGTGCTGAAAGTGGAAATCTAAATACC TCTCCATCATCTAACCAAACAAGAAATTCTGAGAAATTTGAAAAGCCAGAGAATGAAA TTGAAGCCCAGTTGATATGTGAACCCCCAATCAATGGATCCTCAACTCCAAATCCAAA GATAGCATCTTCTGTCACTGCTGGAGTTGCCAGTTCACTCTCAGAAAAAATAGCCGAC AGCATTGGAAATAACCGGCAAAATGCACCATTGACTTCCATTCAAATTCGTTTTATTC AGAACATGATACAGGAAACGTTGGATGACTTTAGAGAAGCATGCCATAGGGACATTGT GAATTTGCAAGTGGAGATGATTAAACAGTTTCATATGCAACTGAATGAAATGCATTCT TTGCTGGAAAGATACTCAGTGAATGAAGGTTTAGTGGCTGAAATTGAAAGACTACGAG AAGAAAACAAAAGATTACGGGCCCACTTTTGAAATTT
ORF Start: ATG at 22 ORF Stop: TGA at 2002
SEQ ED NO: 166 660 aa MW at 71965.3kD
NOV55b, MQENLRFASSGDDIKIWDASSMTLVDKFNPHTSPHGISSICWSSNNNFLVTASSSGDK IWSSCKCKPVPLLELAEGQKQTCVNLNSTSMYLVSGGLNNTVNIWDLKSKRVHRSLK
CG59435-02 Protein Sequence DHKDQVTCVTYN NDCYIASGSLSGEI ILHSVTTNLSSTPFGHGSNQSVRHLKYSLFK
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 55B.
Further analysis of the NOV55a protein yielded the following properties shown in Table 55C.
Table 55C. Protein Sequence Properties NOV55a
PSort 0.4500 probability located in cytoplasm; 0.3000 probability located in microbody analysis: (peroxisome); 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV55a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 55D.
In a BLAST search of public sequence databases, the NOV55a protein was found to have homology to the proteins shown in the BLASTP data in Table 55E.
PFam analysis predicts that the NOV55a protein contains the domains shown in the Table 55F.
Example 56.
The NOV56 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 56A.
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 56B.
Table 56B. Comparison of NOV56a against NOV56b.
NOV56a Residues/ Identities/
Protein Sequence Match Residues Similarities for the Matched Region
NOV56b 1..577 543/577 (94%) 1..543 543/577 (94%)
Further analysis of the NOV56a protein yielded the following properties shown in Table 56C.
Table 56C. Protein Sequence Properties NOV56a
PSort 0.6400 probability located in microbody (peroxisome); 0.4712 probability located analysis: in mitochondrial matrix space; 0.1737 probability located in mitochondrial inner membrane; 0.1737 probability located in mitochondrial intermembrane space
SignalP Likely cleavage site between residues 23 and 24 analysis:
A search of the NOV56a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 56D.
In a BLAST search of public sequence databases, the NOV56a protein was found to have homology to the proteins shown in the BLASTP data in Table 56E.
PFam analysis predicts that the NOV56a protein contains the domains shown in the Table 56F.
Table 56F. Domain Analysis of NOV56a
Identities/
NOV56a Match Expect
Pfam Domain Similarities Region Value
Example 57.
The NOV57 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 57A.
Table 57A. NOV57 Sequence Analysis
SEQ ED NO: 171 2501 bp
NOV57a, ACACCATGACCACCCTTGATGATAAGTTGCTGGGGGAGAAACTGCAGTACTACTATAG
CAGCAGTGAGGATGAGGACAGTGACCACGAGGACAAGGACCGAGGCAGATGTGCCCCA
CG59354-01 DNA Sequence GCCAGCAGTTCTGTGCCTGCAGAGGCTGAGCTGGCAGGCGAAGGCATCTCAGTTAACA CAATGACTCTGAAGGAGTTTGCCATAATGAATGAGGACCAAGATGATGAAGAGTTTCT GCAGCAGTACCGGAAGCAGCGAATGGAAGAGATGCGGCAGCAGCTTCACAAGGGGCCC CAATTCAAGCAGGTTTTTGAGATCTCCAGTGGAGAAGGGTTTTTAGACATGATTGATA AAGAACAGAAAAGCATTGTCATCATGGTTCATATTTATGAGGATGGCATTCCAGGGAC CGAAGCCATGAATGGTTGCATGATCTGCCTTGCCGCAGAGTACCCAGCTGTCAAGTTC TGCAAGGTGAAGAGCTCAGTTATTGGCGCCAGCAGTCAGTTCACCAGGAATGCCCTTC CTGCCCTGCTGATCTATAAGGGGGGTGAATTGATCGGCAATTTTGTTCGTGTTACTGA CCAGCTGGGGGATGATTTCTTTGCTGTGGACCTTGAAGCTTTTCTCCAGGAATTTGGA TTACTCCCAGAAAAGGAAGTCTTGGTGCTGACATCTGTGCGTAACTCTGCCACGTGTC ACAGTGAGGATAGCGACCTGGAAATAGATTGAACTGATAGTCTAGTTGCATAGATTTC
TCATTGTTTGGGTTGGAATACACGTCATTGTTTATTTTTGTTCCTTTGTCTTCTGGCT
TTTCAGCTGTTCTTTGTAGTCCCTTTTATTATGCATAAAATAAAGAAATTCTTAGATT
AAATCAGAATGCTGAATAACCTTGTAGCTAGCAATAAGGTGACTTACAATTGTATAAA
CAGGAAGCCAGGCTTTTGAACTGTTTACTTAAGATTCTGTGGTGTGACATCTCTGTTA
TTGTTTCCAGTCAATATTTACAAAGCATCCTAAAGACAGGGTCTTGGAAATTGTCTTC
AGATGATCTTAGAGGTCTCTGCCAAGTCTGAGAGTATAATTCTGTAGGTATTGTGTTA
TTTGCAACGTAAATAGTGCATTTTCTTAATCAAATGATTGTAAATTATATTTACTTGT
AATCAGTTCCATAGCTTTAGACGGTGGTTAGATTTTTTTTTTCCCCACCAGGGTCTTG
TTTAAAGGGGTGAGCCACCGCACCCAGTCCTGAGGGGTGGCCTCTGCTGCTGGATTTC
ATGTCTTCCTCCAGCATGACTAAGTCTGGAACAGCAGGAAGGGTTGATGCTTACTGAC
CTGGTGATGTTAGAAGACAAGTAGTTTATGGATTTAAACATTAGAGCTGGAGTGGGGC
TGGAAATCTTTGTAAAGGAAGTTCTTTCAGTAAGATGCCCCTGCTTGTCTTTGTCTCT
TTTTTGTTTAACAAGGTAACTTTTTGTTTAACAAGGTAACTTTTTGTTTAACCTAGAT
TTTTTTTAAAACTTTTTTTTTTTTTCATATTGGAAAAGTAATTCATATTCAGTAGAGG
AAAACTGACCAAAACAGAAGCAAAAATAAGAAAATTAAAATAATCTCTAATCCTACTA
CCTAGAATAAAACACTATTAATATTTTGGTCTGTTTCCTGCCAAGGTGTTTTCTGTGT
ATACATGGATATTTTGTTTGTTTTTAAACAAAACGATGGGATCATTCTGAACATACTG
TTCTATAGTATGGTCAGCTAATAATATATCAGACCTTTTTTTTATATTATTAAATATT
CTACAACTTTTTAAAAATGTCTATTAATATTCCATCGTATAGATGTGATATAATTTGC
TTGATGGTTGTCTCTTAAAAAGAAAGATAGCAAATACTTTTTTTAAATTACAAAAGTG
ATAGATGTTCATTGTAGAAAATGTAATAAACACTGTTAAGACTTAAAAGCCATATAAT
TCCACCAACCAAAATTAATCCCTTTTGTCATATTTCTAGTCATTTTTATAGCCTTTTT
TTTCTATGTATTTATAATAATTATCATTTGCGTTTTTTTCCTTTTTTTAACTTTAAAA
ATGTATATTCTAGGGTCAGGGGAAATGTAATCTGGAATTAAATATTAGCCTTAAAATT
CACAATTTTGATTTTCCTGGCTTTTCAGGAATTGACTAACTGTAAAAGAGTCTTGAAA
GTATTTAGTCAACAAACAGAGTGCATTTTTTTTTTTTTGACTAAGAAAGCTCGTTGTA
GTAGAAAGGGTGGAATGTATTGAAAATTATTAGAAGCAGGGAAGTATTGTTAGTCTAG
CTTATTTCCTTTCAGTCTTTTTTCAATATTTTTATAAACATTGAGTACTTACTGAATT
TAGTTCTGTGCTCTTCCTTATTTAGTGTTGTATCATAAATACTTTGATGTTTCAAACA
TTCTAAATAAATAATTTTCAGTGGCTTCATAATAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAA
ORF Start: ATG at 6 ORF Stop: TGA at 726
SEQ ED NO: 172 240 aa MW at 26866.9kD
NOV57a, MTTLDDKLLGEKLQYYYSSSEDEDSDHEDKDRGRCAPASSSVPAEAELAGEGISVNTM TLKEFAIMNEDQDDEEFLQQYRKQRMEEMRQQLHKGPQFKQVFEISSGEGFLDMIDKE
CG59354-01 Protein Sequence QKSIVIMVHIYEDGIPGTEAMNGCMICLAAEYPAVKFCKVKSSVIGASSQFTRNALPA LLIYKGGELIGNFVRVTDQLGDDFFAVDLEAFLQEFGLLPEKEVLVLTSVRNSATCHS EDSDLEID
SEQ ED NO: 173 893 bp
NOV57b, CACCATGACCACCCTGTATGATAAGTTGCTGGGGGAGAAACTGCAGTACTACTATAGC
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 57B.
Further analysis of the NOV57a protein yielded the following properties shown in Table 57C.
Table 57C. Protein Sequence Properties NOV57a
PSort 0.6500 probability located in cytoplasm; 0.1000 probability located in analysis: mitochondrial matrix space; 0.1000 probability located in lysosome (lumen); 0.0000 probability located in endoplasmic reticulum (membrane)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV57a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 57D.
In a BLAST search of public sequence databases, the NOV57a protein was found to have homology to the proteins shown in the BLASTP data in Table 57E.
PFam analysis predicts that the NOV57a protein contains the domains shown in the Table 57F.
Example 58.
The NOV58 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 58 A.
NOV58b, MDPNEDTE NDILRDFGILPPKEESKDEIEEMVLRLQKEAMVKPFEKMTLAQLKEAED EFDEEDMQAVETYRKKRLQEWKALKKKQKFGELREISGNQYVNEVTNAEEDVWVIIHL
CG59319-02 Protein Sequence YRSSIPMCLLVNQHLSLLARKFPETKFVKAIVNSCIQHYHDNCLPTIFVYKNGQIEAK FIGIIECGGINLKLEELE KLAEVGAIQTDLEENPRKDMVDMMVSSIRNTSIHDDSDS SNSDNDTK
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 58B.
Table 58B. Comparison of NOV58a against NOV58b.
NOV58a Residues/ Identities/
Protein Sequence Match Residues Similarities for the Matched Region
NOV58b 1..239 216/239 (90%) 2..240 216/239 (90%)
Further analysis of the NOV58a protein yielded the following properties shown in Table 58C.
Table 58C. Protein Sequence Properties NOV58a
PSort 0.8800 probability located in nucleus; 0.1000 probability located in analysis: mitochondrial matrix space; 0.1000 probability located in lysosome (lumen); 0.0000 probability located in endoplasmic reticulum (membrane)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV58a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 58D.
In a BLAST search of public sequence databases, the NOV58a protein was found to have homology to the proteins shown in the BLASTP data in Table 58E.
PFam analysis predicts that the NOV58a protein contains the domains shown in the Table 58F.
Example 59.
The NOV59 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 59A.
Further analysis of the NOV59a protein yielded the following properties shown in Table 59B.
Table 59B. Protein Sequence Properties NOV59a
PSort 0.6400 probability located in plasma membrane; 0.4600 probability located in analysis: Golgi body; 0.3700 probability located in endoplasmic reticulum (membrane); 0.1000 probability located in endoplasmic reticulum (lumen)
SignalP Likely cleavage site between residues 25 and 26 analysis:
A search of the NOV59a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 59C.
In a BLAST search of public sequence databases, the NOV59a protein was found to have homology to the proteins shown in the BLASTP data in Table 59D.
PFam analysis predicts that the NOV59a protein contains the domains shown in the Table 59E.
Example 60.
The NOV60 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 60A.
Table 60A. NOV60 Sequence Analysis
SEQ ED NO: 183 1201 bp
NOV60a, AGGATAACTTTATATGTTGCAAAATGACTCACATAGTATATTTTATTTAACCAGCCTA
ATTTCAAGGCTGTTTAGTTGCTTGAAAAGAAGGTTTTTATTTGTTCTTTGCATGTACT
CG59557-01 DNA Sequence TAGAATGCTGACTGTGTTTTATGAGCCAACAAGTGAAACCGCTGAAAATATGGATCCA
GAGAATCAGACAATGGTGACTGAGTTTTATTTCTCTGATTTTCCTCAATCTAAGAATG GCAGCCTCTTATTCTTCATTCCTATGCTCTTTATTTATATATTCATTCTTGTTGGAAA TTTCATGATTTTCTTTGCTGTCCGACCGGACCCCCATCTCCATAATCCTATGTACAGT TTTATCAGTGTCTTCTCCTTCCTGGAGATTTGGTACACCACCGTGACTATCCCCAAGA TGCTCTCCAACCTTCTCAGTGAACAGAAAACCATCTCTTTCATAGGTTGCCTCCTGCA GATGTACTTCTTCCACTCACTCGGGGTCACAGAAGCCCTAGTCCTCACAGTGATGGCC ATTGACAGGTGTGTAGCCATCTGCAACCCCCTTCGCTATGCAATCACTATGTCCCCTA GACTGTGCATCCAGCTCTCCACTGGCTCTTGCATTTTTGGCTTCCTCATGTTACTGCC AGAGATTGTGTGCATTTCCACTCTTCCATTCTGTGGCGCCAACCAAATTCATCAACTC TTTTGTGACTTTGAACCTGTGCTGCAGTTAGCCTGCACAGATACGTACATAATTCTGG TTGAAGATGTGATCCGTGCTATTTCCATTCTGACCTCTGTCTCTGTCATCACCCTTTT CTATTTAAGAATCATCACGGTGATCCTGAGGATTCCCTCTGGTGAGAGTCGTCAGAAG GCTTTCTTCACATGTGCAGCCCACATTGCTATTTTCTTGCTGTTTTTTGGCAGTGTGT CACTCATGTATCTGCGCTTCTCTGTCACATTCCCACCATTACTGGACAAGGCCATTGC ACTGATGTTTGCTGTCCTTGCCCTACTTTTCAACCCAGTAATCTATAGTCTGAGGAAC AAAGATATGAAAAACGCCACCAAGAAAATCCTCTGTTCTCAAAAGATGTTCAATGCCT CTGGGAGCTAATGGAGTTCACACACACCTCTTCAAAGAAATCTCATCATCTCCTTAAG TTTAAAATGCTAACAAATCAGTTTTTTTAAATTACCATGCA
ORF Start: ATG at 121 ORF Stop: TAA at l l ll
SEQ ED NO: 184 330 aa MW at 37439. lkD
NOV60a, MLTVFYEPTSETAENMDPENQTMVTEFYFSDFPQSKNGSLLFFIPMLFIYIFILVGNF MIFFAVRPDPHLHNP YSFISVFSFLEIWYTTVTIPKMLSNLLSEQKTISFIGCLLQM
CG59557-01 Protein Sequence YFFHSLGVTEALVLTVMAIDRCVAICNPLRYAITMSPRLCIQLSTGSCIFGFLMLLPE
IVCISTLPFCGANQIHQLFCDFEPVLQLACTDTYIILVEDVIRAISILTSVSVITLFY LRIITVILRIPSGESRQKAFFTCAAHIAIFLLFFGSVSLMYLRFSVTFPPLLDKAIAL MFAVLALLFNPVIYSLRNKDMKNATKKILCSQKMFNASGS
Further analysis of the NOV60a protein yielded the following properties shown in Table 60B.
Table 60B. Protein Sequence Properties NOV60a
PSort 0.6000 probability located in plasma membrane; 0.4000 probability located in analysis: Golgi body; 0.3000 probability located in endoplasmic reticulum (membrane); 0.0300 probability located in mitochondrial inner membrane
SignalP Likely cleavage site between residues 67 and 68 analysis:
A search of the NOV60a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 60C.
In a BLAST search of public sequence databases, the NOV60a protein was found to have homology to the proteins shown in the BLASTP data in Table 60D.
PFam analysis predicts that the NOV60a protein contains the domains shown in the Table 60E.
Example 61.
The NOV61 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 61 A.
SEQ ID NO: 185 | l06l"bp"
NOV61a, CAATCTGGTCCTAAGTGATCTTTTTCTTTTTCACAGGGAAATGGGGGAAAATCAGACA
ATGGTCACAGAGTTCCTCCTACTGGGATTTCTCCTGGGCCCAAGGATTCAGATGCTCC
CG59555-01 DNA Sequence TCTTTGGGCTCTTCTCCCTGTTCTATATCTTCACCCTGCTGGGGAACGGGGCCATCCT GGGGCTCATCTCACTGGACTCCAGACTCCACACCCCCATGTACTTCTTCCTCTCACAC CTGGCTGTCGTCGACATCGCCTACACCCGCAACACGGTGCCCCAGATGCTGGCGAACC TCCTGCATCCAGCCAAGCCCATCTCCTTTGCTGGCTGCATGACGCAGACCTTTCTCTG TTTGAGTTTTGGACACAGCGAATGTCTCCTGCTGGTGCTGATGTCCTACGATCGTTAC GTGGCCATCTGCCACCCTCTCCGATACTCCGTCATCATGACCTGGAGAGTCTGCATCA CCCTGGCCGTCACTTCCTGGACGTGTGGCTCCCTCCTGGCTCTGGCCCATGTGGTTCT CATCCTAAGACTGCCCTTCTCTGGGCCTCATGAAATCAACCACTTCTTCTGTGAAATC CTGTCTGTCCTCAGGCTGGCCTGTGCTGACACCTGGCTCAACCAGGTGGTCATCTTTG CAGCCTGCGTGTTCTTCCTGGTGGGGCCACCCAGCCTGGTGCTTGTCTCCTACTCGCA CATCCTGGCGGCCATCCTGAGGATCCAGTCTGGGGAGGGCCGCAGAAAGGCCTTCTCC ACCTGCTCCTCCCACCTCTGCGTGGTGGGACTCTTCTTTGGCAGTGCCATCATCATGT ACATGGCCCCCAAGTCCCGCCATCCTGAGGAGCAGCAAAAGGTCTTTTTTCTATTTTA CAGTTTTTTCAACCCAACACTTAACCCCCTGATTTACAGCCTGAGGAACGGAGAGGTC AAGGGTGCCCTGAGGAGAGCACTGGGCAAGGAAAGTCATTCCTAACTGGTGTGACATT
TGACTCTCCCTCCTCAGTCATCTCCTGGAATCTTGGTACCAAATACCACCTAAGTTCA
CTACTCTCTTTATATCA
ORF Start: ATG at 41 ORF Stop: TAA at 971
SEQ ED NO: 186 310 aa MW at 34713.8kD
NOV61a, MGENQTMVTEFLLLGFLLGPRIQMLLFGLFSLFYIFTLLGNGAILGLISLDSRLHTPM YFFLSHLAWDIAYTRNTVPQMLANLLHPAKPISFAGCMTQTFLCLSFGHSECLLLVL
CG59555-01 Protein Sequence SYDRYVAICHPLRYSVIMT RVCITLAVTS TCGSLLALAHWLILRLPFSGPHEIN HFFCEILSVLRLACADT LNQWIFAACVFFLVGPPSLVLVSYSHILAAILRIQSGEG RRKAFSTCSSHLCWGLFFGSAIIMYMAPKSRHPEEQQKVFFLFYSFFNPTLNPLIYS LRNGEVKGALRRALGKESHS
Further analysis of the NOV6 la protein yielded the following properties shown in Table 61B.
Table 61B. Protein Sequence Properties NOV61a
PSort 0.6400 probability located in plasma membrane; 0.4600 probability located in analysis: Golgi body; 0.3700 probability located in endoplasmic reticulum (membrane); 0.1000 probability located in endoplasmic reticulum (lumen)
SignalP Likely cleavage site between residues 43 and 44 analysis:
A search of the NOV6 la protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 61C.
Ln a BLAST search of public sequence databases, the NOV61a protein was found to have homology to the proteins shown in the BLASTP data in Table 6 ID.
PFam analysis predicts that the NOV6 la protein contains the domains shown in the Table 6 IE.
Example 62.
The NOV62 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 62A.
NPLIYTLRNRDVKAAITKIMSQDPGCDRSI
Further analysis of the NOV62a protein yielded the following properties shown in Table 62B.
Table 62B. Protein Sequence Properties NOV62a
PSort 0.6000 probability located in plasma membrane; 0.4000 probability located in analysis: Golgi body; 0.3000 probability located in endoplasmic reticulum (membrane); 0.3000 probability located in microbody (peroxisome)
SignalP Likely cleavage site between residues 57 and 58 analysis:
A search of the NOV62a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 62C.
In a BLAST search of public sequence databases, the NOV62a protein was found to have homology to the proteins shown in the BLASTP data in Table 62D.
PFam analysis predicts that the NOV62a protein contains the domains shown in the Table 62E.
Example 63.
The NOV63 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 63A.
NOV63a, GACCTTTCATCACACTCTGGTCATTTACAAACTGTTATTAAGGAATGGGGGACAAGCA
GCCCTGGGTCACAGAATTCATCCTGGTTGGATTCCAGCTCAGTGCAGAGATGGAGATC
CG59540-01 DNA Sequence TTTCTCTCTTGCATCTTCTCCCTGTTATATCTCTTCAGTCTACTGAGGAATGGCATGA ACATGGGACTCATCTGTCTGGATCCCAGACTACACACCCCCATATACTTCTTCCTGTC ACACTTGGCCGTCATTGACATATACTATGCTTCCAACAATTTGCTCAACATGCTGGAA AACCTAGTGAAACACAAAAAAACTATCTCGTTCATCTCTTGCATTATGCAGATGGCTT TGTATTTGACTTTTGCTGCTGCAGTGTGCATGATTTTGGTGGTGATGTCCTATGACAG ATTTGTGGCGATCTGCCATCCCCTGCATTACACTGTCATCATGAACTGGAGAGTGTGC ACAGTACTGGCTATTACTTCCTGGGCATGTGGATTTTCCCTGGCCCTCATAAATCTAA TTCTCCTTCTAAGGCTGCCCTTCTGTGGGCCCCAGGAGGTGAACCACTTCTTCGGTGA AATTCTGTCTGTCCTCAAACTGGCCTGTGCAGACACCTGGATTAATGAAATTTTTGTC TTTGCTGGTGGTGTGTTTGTCTTAGTCGGGCCCCTTTCCTTGATGCTGATCTCCTACA TGCGCATCCTCTTGGCCATCCTGAAGATCCAGTCAGGCGAGGGCCACAGAAAGGACTT CTCTACCTGCTCCTCCCACCTCTGTGTGGTGGGGTTCTTCTTTGCCAACGCCATTGTC ATGTACATGGCCCCCAAGTCCCGCCATCCCGAGGAGCAGCAGAAGGTCCTTTCCCTGT TTTGCAGCCTTTGGAATCAGGTGCTGAACCCCCCTCTGATCTACAGCTTGAGGAATGC AGAGGTCAAGAGTGCCCCACAAGAGGGCCACTGAAGAAGGAGAGGCTGATGTTACAAT
CTCAAAGGCACCACGAGGAGAGGGCCTGCTCCGACAAATGGGGAAGTTGGCTTTTT
ORF Start: ATG at 45 ORF Stop: TGA at 960
SEQ ED NO: 190 305 aa MW at 34554.8kD
NOV63a, MGDKQPWVTEFILVGFQLSAEMEIFLSCIFSLLYLFSLLRNGMNMGLICLDPRLHTPI YFFLSHLAVIDIYYASNNLLNMLENLVKHKKTISFISCIMQMALYLTFAAAVCMILW
CG59540-01 Protein Sequence MSYDRFVAICHPLHYTVIMN RVCTVLAITS ACGFSLALINLILLLRLPFCGPQEVN HFFGEILSVLKLACADT INEIFVFAGGVFVLVGPLSLMLISYMRILLAILKIQSGEG HRKDFSTCSSHLCWGFFFANAIVMYMAPKSRHPEEQQKVLSLFCSLWNQVLNPPLIY SLRNAEVKSAPQEGH
Further analysis of the NOV63a protein yielded the following properties shown in Table 63B.
Table 63B. Protein Sequence Properties NOV63a
PSort 0.6000 probability located in plasma membrane; 0.4000 probability located in analysis: Golgi body; 0.3000 probability located in endoplasmic reticulum (membrane); 0.3000 probability located in microbody (peroxisome)
SignalP Likely cleavage site between residues 43 and 44 analysis:
A search of the NOV63a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 63C.
In a BLAST search of public sequence databases, the NOV63a protein was found to have homology to the proteins shown in the BLASTP data in Table 63D.
PFam analysis predicts that the NOV63a protein contains the domains shown in the Table 63E.
Example 64.
The NOV64 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 64A.
NOV64b, FFIPLFVIYIFIVIGNLIVFFAVRVDTRLHNPMYNFISIFSFLEI YTTATIPKMLSI LISRQRTISMVGCLLQMYFFHSLGNSEGILLTTMAIDRYVAICNPLRYPTIMTPGLCV
CG59280-02 Protein Sequence QLSVGSCIFGFLVLLPEIA ISTLPFCGPNQIHQIFCDFEPVLRLACTDTSMILIEDV IHAVAIVFSVLIIAFSYIRIITVILRIPSVEGRQKAFSTCAAHLSVFLMFYGSVSLMY LRFSATFPPILDTAVALMFAVLAPFFNPIIYSFRNKDMKIAIKKLFCPQKMVNLSVD
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 64B.
Table 64B. Comparison of NOV64a against NOV64b.
NOV64a Residues/ Identities/
Protein Sequence Match Residues Similarities for the Matched Region
NOV64b 27..315 289/289 (100%) 1..289 289/289 (100%)
Further analysis of the NOV64a protein yielded the following properties shown in Table 64C.
Table 64C. Protein Sequence Properties NOV64a
PSort 0.6000 probability located in plasma membrane; 0.4000 probability located in analysis: Golgi body; 0.3000 probability located in endoplasmic reticulum (membrane); 0.3000 probability located in microbody (peroxisome)
SignalP Likely cleavage site between residues 54 and 55 analysis:
A search of the NOV64a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 64D.
In a BLAST search of public sequence databases, the NOV64a protein was found to have homology to the proteins shown in the BLASTP data in Table 64E.
PFam analysis predicts that the NOV64a protein contains the domains shown in the Table 64F.
Table 64F. Domain Analysis of NOV64a
Identities/
Pfam Domain NOV64a Match Region Similarities Expect Value for the Matched Region
Example 65.
The NOV65 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 65 A.
Further analysis of the NOV65a protein yielded the following properties shown in Table 65B.
Table 65B. Protein Sequence Properties NO 65a
PSort 0.6000 probability located in plasma membrane; 0.4000 probability located in analysis: Golgi body; 0.3888 probability located in mitochondrial inner membrane; 0.3030 probability located in mitochondrial intermembrane space
SignalP Likely cleavage site between residues 45 and 46 analysis:
A search of the NOV65a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 65C.
In a BLAST search of public sequence databases, the NOV65a protein was found to have homology to the proteins shown in the BLASTP data in Table 65D.
PFam analysis predicts that the NOV65a protein contains the domains shown in the Table 65E.
Example 66.
The NOV66 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 66A.
Further analysis of the NOV66a protein yielded the following properties shown in Table 66B.
Table 66B. Protein Sequence Properties NOV66a
PSort 0.6000 probability located in plasma membrane; 0.4000 probability located in analysis: Golgi body; 0.3000 probability located in endoplasmic reticulum (membrane); 0.2007 probability located in mitochondrial inner membrane
SignalP Likely cleavage site between residues 50 and 51 analysis:
A search of the NOV66a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 66C.
In a BLAST search of public sequence databases, the NOV66a protein was found to have homology to the proteins shown in the BLASTP data in Table 66D.
PFam analysis predicts that the NOV66a protein contains the domains shown in the Table 66E.
Example 67.
The NOV67 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 67 A.
Further analysis of the NOV67a protein yielded the following properties shown in Table 67B.
Table 67B. Protein Sequence Properties NOV67a
PSort 0.6000 probability located in plasma membrane; 0.4047 probability located in analysis: mitochondrial inner membrane; 0.4000 probability located in Golgi body; 0.3480 probability located in mitochondrial intermembrane space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV67a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 67C.
In a BLAST search of public sequence databases, the NOV67a protein was found to have homology to the proteins shown in the BLASTP data in Table 67D.
PFam analysis predicts that the NOV67a protein contains the domains shown in the Table 67E.
Example 68.
The NOV68 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 68A.
Further analysis of the NOV68a protein yielded the following properties shown in Table 68B.
Table 68B. Protein Sequence Properties NOV68a
PSort 0.6400 probability located in plasma membrane; 0.4600 probability located in analysis: Golgi body; 0.3700 probability located in endoplasmic reticulum (membrane); 0.1000 probability located in endoplasmic reticulum (lumen)
SignalP Likely cleavage site between residues 50 and 51 analysis:
A search of the NOV68a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 68C.
In a BLAST search of public sequence databases, the NOV68a protein was found to have homology to the proteins shown in the BLASTP data in Table 68D.
PFam analysis predicts that the NOV68a protein contains the domains shown in the Table 68E.
Example 69.
The NOV69 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 69 A.
Further analysis of the NOV69a protein yielded the following properties shown in Table 69B.
Table 69B. Protein Sequence Properties NOV69a
PSort 0.6000 probability located in plasma membrane; 0.4000 probability located in analysis: Golgi body; 0.3000 probability located in endoplasmic reticulum (membrane); 0.0300 probability located in mitochondrial inner membrane
SignalP Likely cleavage site between residues 40 and 41 analysis:
A search of the NOV69a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 69C.
In a BLAST search of public sequence databases, the NOV69a protein was found to have homology to the proteins shown in the BLASTP data in Table 69D.
PFam analysis predicts that the NOV69a protein contains the domains shown in the Table 69E.
Example 70.
The NOV70 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 70A.
Further analysis of the NOV70a protein yielded the following properties shown in Table 70B.
Table 70B. Protein Sequence Properties NOV70a
PSort 0.6000 probability located in plasma membrane; 0.4000 probability located in analysis: Golgi body; 0.3000 probability located in endoplasmic reticulum (membrane); 0.2007 probability located in mitochondrial inner membrane
SignalP Likely cleavage site between residues 50 and 51 analysis:
A search of the NOV70a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 70C.
In a BLAST search of public sequence databases, the NOV70a protein was found to have homology to the proteins shown in the BLASTP data in Table 70D.
PFam analysis predicts that the NOV70a protein contains the domains shown in the Table 70E.
Example 71.
The NOV71 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 71 A.
Further analysis of the NOV71a protein yielded the following properties shown in Table 71B.
Table 71 B. Protein Sequence Properties NO 71a
PSort 0.6000 probability located in plasma membrane; 0.4047 probability located in analysis: mitochondrial inner membrane; 0.4000 probability located in Golgi body; 0.3480 probability located in mitochondrial intermembrane space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV7 la protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 71C.
In a BLAST search of public sequence databases, the NOV7 la protein was found to have homology to the proteins shown in the BLASTP data in Table 7 ID.
PFam analysis predicts that the NOV7 la protein contains the domains shown in the Table 7 IE.
Example 72.
The NOV72 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 72A.
Further analysis of the NOV72a protein yielded the following properties shown in Table 72B.
Table 72B. Protein Sequence Properties NOV72a
PSort 0.6000 probability located in plasma membrane; 0.4000 probability located in analysis: Golgi body; 0.3000 probability located in endoplasmic reticulum (membrane); 0.0300 probability located in mitochondrial inner membrane
SignalP Likely cleavage site between residues 44 and 45 analysis:
A search of the NOV72a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 72C.
In a BLAST search of public sequence databases, the NOV72a protein was found to have homology to the proteins shown in the BLASTP data in Table 72D.
PFam analysis predicts that the NOV72a protein contains the domains shown in the Table 72E.
Example 73.
The NOV73 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 73 A.
Further analysis of the NOV73a protein yielded the following properties shown in Table 73B.
Table 73B. Protein Sequence Properties NOV73a
Psort 0.8110 probability located in plasma membrane; 0.6400 probability located in analysis: endoplasmic reticulum (membrane); 0.3700 probability located in Golgi body; 0.1839 probability located in microbody (peroxisome)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV73a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 73C.
In a BLAST search of public sequence databases, the NOV73a protein was found to have homology to the proteins shown in the BLASTP data in Table 73D.
PFam analysis predicts that the NOV73a protein contains the domains shown in the Table 73E.
Example 74.
The NOV74 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 74A.
Further analysis of the NOV74a protein yielded the following properties shown in Table 74B.
Table 74B. Protein Sequence Properties NOV74a
PSort 0.4328 probability located in mitochondrial matrix space; 0.3000 probability analysis: located in microbody (peroxisome); 0.1137 probability located in mitochondrial inner membrane; 0.1137 probability located in mitochondrial intermembrane space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV74a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 74C.
In a BLAST search of public sequence databases, the NOV74a protein was found to have homology to the proteins shown in the BLASTP data in Table 74D.
PFam analysis predicts that the NOV74a protein contains the domains shown in the Table 74E.
Example 75.
The NOV75 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 75 A.
Further analysis of the NOV75a protein yielded the following properties shown in Table 75B.
Table 75B. Protein Sequence Properties NOV75a
PSort 0.6500 probability located in cytoplasm; 0.1000 probability located in analysis: mitochondrial matrix space; 0.1000 probability located in lysosome (lumen); 0.0442 probability located in microbody (peroxisome)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV75a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 75C.
In a BLAST search of public sequence databases, the NOV75a protein was found to have homology to the proteins shown in the BLASTP data in Table 75D.
PFam analysis predicts that the NOV75a protein contains the domains shown in the Table 75E.
Example 76.
The NOV76 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 76A.
Table 76A. NOV76 Sequence Analysis
SEQ ED NO: 217 |7497 bp~
NOV76a, ATGGTCTTGCTTCTTTGTCTATCTTGTCTGATTTTCTCCTGTCTGACCTTTTCCTGGT TAAAAATCTGGGGGAAAATGACGGACTCCAAGCCGATCACCAAGAGTAAATCAGAAGC
CG59641-01 DNA Sequence AAACCTCATCCCGAGCCAGGAGCCCTTTCCAGCCTCTGATAACTCAGGGGAGACACCG CAGAGAAATGGGGAGGGCCACACTCTGCCCAAGACACCCAGCCAGGCCGAGCCAGCCT CCCACAAAGGCCCCAAAGATGCCGGTCGGCGGAGAAACTCCCTACCACCCTCCCACCA GAAGCCCCCAAGAAACCCCCTTTCTTCCAGTGACGCAGCACCCTCCCCAGAGCTTCAA GCCAACGGGACTGGGACACAAGGTCTGGAGGCCACAGATACCAATGGCCTGTCCTCCT CAGCCAGGCCCCAGGGCCAGCAAGCTGGCTCCCCCTCCAAAGAAGACAAGAAGCAGGC AAACATCAAGAGGCAGCTGATGACCAACTTCATCCTGGGCTCTTTTGATGACTACTCC TCCGACGAGGACTCTGTTGCTGGCTCATCTCGTGAGTCTACCCGGAAGGGCAGCCGGG CCAGCTTGGGGGCCCTGTCCCTGGAGGCTTATCTGACCACAAGGCCGAGCATGTCGGG ACTCCACCTGGTGAAGAGGGGACGGGAACACAAGAAGCTGGACCTGCACAGAGACTTT ACCGTGGCTTCTCCCGCTGAGTTTGTCACACGCTTTGGGGGGGATCGGGTCATCGAGA AGGTGCTTATTGCCAACAACGGGATTGCCGCCGTGAAGTGCATGCGCTCCATCCGCAG GTGGGCCTATGAGATGTTCCGCAACGAGCGGGCCATCCGGTTTGTTGTGATGGTGACC CCCGAGGACCTTAAGGCCAACGCAGAGTACATCAAGATGGCGGATCATTACGTCCCCG TCCCAGGAGGGCCCAATAACAACAACTATGCCAACGTGGAGCTGATTGTGGACATTGC CAAGAGAATCCCCGTGCAGGCGGTGTGGGCTGGCTGGGGCCATGCTTCAGAAAACCCT AAACTTCCGGAGCTGCTGTGCAAGAATGGAGTTGCTTTCTTAGGCCCTCCCAGTGAGG CCATGTGGGCCTTAGGAGATAAGATCGCCTCCACCGTTGTCGCCCAGACGCTACAGGT CCCAACCCTGCCCTGGAGTGGAAGCGGCCTGACAGTGGAGTGGACAGAAGATGATCTG CAGCAGGGAAAAAGAATCAGTGTCCCAGAAGATGTTTATGACAAGGGTTGCGTGAAAG ACGTAGATGAGGGCTTGGAGGCAGCAGAAAGAATTGGTTTTCCATTGATGATCAAAGC TTCTGAAGGTGGCGGAGGGAAGGGAATCCGGAAGGCTGAGAGTGCGGAGGACTTCCCG ATCCTTTTCAGACAAGTACAGAGTGAGATCCCAGGCTCGCCCATCTTTCTCATGAAGC TGGCCCAGCACGCCCGTCACCTGGAAGTTCAGATCCTCGCTGACCAGTATGGGAATGC TGTGTCTCTGTTTGGTCGCGACTGCTCCATCCAGCGGCGGCATCAGAAGATCGTTGAG GAAGCACCGGCCACCATCGCCCCGCTGGCCATATTCGAGTTCATGGAGCAGTGTGCCA TCCGCCTGGCCAAGACCGTGGGCTATGTGAGTGCAGGGACAGTGGAATACCTCTATAG TCAGGATGGCAGCTTCCACTTCTTGGAGCTGAATCCTCGCTTGCAGGTGGAACATCCC TGCACAGAAATGATTGCTGATGTTAATCTGCCGGCCGCCCAGCTACAGATCGCCATGG GCGTGCCACTGCACCGGCTGAAGGATATCCGGCTTCTGTATGGAGAGTCACCATGGGG AGTGACTCCCATTTCTTTTGAAACCCCCTCAAACCCTCCCCTCGCCCGAGGCCACGTC ATTGCCGCCAGAATCACCAGCGAAAACCCAGACGAGGGTTTTAAGCCGAGCTCCGGGA CTGTCCAGGAACTGAATTTCCGGAGCAGCAAGAACGTGTGGGGTTACTTCAGCGTGGC CGCTACTGGAGGCCTGCACGAGTTTGCGGATTCCCAATTTGGGCACTGCTTCTCCTGG GGAGAGAACCGGGAAGAGGCCATTTCGAACATGGTGGTGGCTTTGAAGGAACTGTCCA TCCGAGGCGACTTTAGGACTACCGTGGAATACCTCATTAACCTCCTGGAGACCGAGAG CTTCCAGAACAACGACATCGACACCGGGTGGTTGGACTACCTCATTGCTGAGAAAGTG CAGGCGGAGAAACCGGATATCATGCTTGGGGTGGTATGCGGGGCCTTGAACGTGGCCG ATGCGATGTTCAGAACGTGCATGACAGATTTCTTACACTCCCTGGAAAGGGGCCAGGT CCTCCCAGCGGATTCACTACTGAACCTCGTAGATGTGGAATTAATTTACGGAGGTGTT AAGTACATTCTCAAGGTGGCCCGGCAGTCTCTGACCATGTTCGTTCTCATCATGAATG GCTGCCACATCGAGATTGATGCCCACCGGCTGAATGATGGGGGGCTCCTGCTCTCCTA CAATGGGAACAGCTACACCACCTACATGAAGGAAGAGGTTGACAGTTACCGAATTACC ATCGGCAATAAGACGTGTGTGTTTGAGAAGGAGAACGATCCTACAGTCCTGAGATCCC CCTCGGCTGGGAAGCTGACACAGTACACAGTGGAGGATGGGGGCCACGTTGAGGCTGG GAGCAGCTACGCTGAGATGGAGGTGATGAAGATGATCATGACCCTGAACGTTCAGGAA AGAGGCCGGGTGAAGTACATCAAGCGTCCAGGTGCCGTGCTGGAAGCAGGCTGCGTGG TGGCCAGGCTGGAGCTCGATGACCCTTCTAAAGTCCACCCGGCTGAACCGTTCACAGG AGAACTCCCTGCCCAGCAGACACTGCCCATCCTCGGAGAGAAACTGCACCAGGTCTTC CACAGCGTCCTGGAAAACCTCACCAACGTCATGAGTGGCTTTTGTCTGCCAGAGCCCG TTTTTAGCATAAAGCTGAAGGAGTGGGTGCAGAAGCTCATGATGACCCTCCGGCACCC GTCACTGCCGCTGCTGGAGCTGCAGGAGATCATGACCAGCGTGGCAGGCCGCATCCCC GCCCCTGTGGAGAAGTCTGTCCGCAGGGTGATGGCCCAGTATGCCAGCAACATCACCT CGGTGCTGTGCCAGTTCCCCAGCCAGCAGATAGCCACCATCCTGGACTGCCATGCAGC CACCCTGCAGCGGAAGGCTGATCGAGAGGTCTTCTTCATCAACACCCAGAGCATCGTG CAGTTGGTCCAGAGATACCGCAGCGGGATCCGCGGCTATATGAAAACAGTGGTGTTGG ATCTCCTGAGAAGATACTTGCGTGTTGAGAGCAAGGCAAGAGATGCTGATGCCAACAC CAGTGGGATGGTGGGGGGCGTGAGGAGCCTGAGCTTTACCTCTGTGTGGTGTTTTGTC TCCCCCGAATCCCACTACGACAAGTGTGTGATAAACCTCAGGGAGCAGTTCAAGCCAG ACATGTCCCAGGTGCTGGACTGCATCTTCTCCCACGCACAGGTGGCCAAGAAGAACCA GCTGGTGATCATGTTGATCGATGAGCTGTGTGGCCCAGACCCTTCCCTGTCGGACGAG CTGATCTCCATCCTCAACGAGCTCACTCAGCTGAGCAAAAGCGAGCACTGCAAAGTGG
CCCTCAGAGCCCGGCAGATCCTGATTGCCTCCCACCTCCCCTCCTACGAGCTGCGGCA TAACCAGGTGGAGTCCATTTTCCTGTCTGCCATTGACATGTACGGCCACCAGTTCTGC CCCGAGAACCTCAAGAAATTAATACTTTCGGAAACAACCATCTTCGACGTCCTGCCTA CTTTCTTCTATCACGCAAACAAAGTCGTGTGCATGGCGTCCTTGGAGGTTTACGTGCG GAGGGGCTACATCGCCTATGAGTTAAACAGCCTGCAGCACCGGCAGCTCCCGGACGGC ACCTGCGTGGTAGAATTCCAGTTCATGCTGCCGTCCTCCCACCCAAACCGGATGACCG TGCCCATCAGCATCACCAACCCTGACCTGCTGAGGCACAGCACAGAGCTCTTCATGGA CAGCGGCTTCTCCCCACTGTGCCAGCGCATGGGAGCCATGGTAGCCTTCAGGAGATTC GAGGACTTCACCAGAAATTTTGATGAAGTCATCTCTTGCTTCGCCAACGTGCCCAAAG ACACCCCCCTCTTCAGCGAGGCCCGCACCTCCCTATACTCCGAGGATGACTGCAAGAG CCTCAGAGAAGAGCCCATCCACATTCTGAATGTGTCCATCCAGTGTGCAGACCACCTG GAGGATGAGGCACTGGTGCCGATTTTACGGACATTCGTACAGTCCAAGAAAAATATCC TTGTGGATTATGGACTCCGACGAATCACATTCTTGATTGCCCAAGAGTTTGCAGAAGA TCGCATTTACCGTCACTTGGAACCTGCCCTGGCCTTCCAGCTGGAACTTAACCGGATG CGTAACTTCGATCTGACCGCCGTGCCCTGTGCCAACCACAAGATGCACCTTTACCTGG GTGCTGCCAAGGTGAAGGAAGGTGTGGAAGTGACGGACCATAGGTTCTTCATCCGCGC CATCATCAGGCACTCTGACCTGATCACAAAGGAAGCCTCCTTCGAATACCTGCAGAAC GAGGGTGAGCGGCTGCTCCTGGAGGCCATGGACGAGCTGGAGGTGGCGTTCAATAACA CCAGCGTGCGCACCGACTGCAACCACATCTTCCTCAACTTCGTGCCCACTGTCATCAT GGACCCCTTCAAGATCGAGGAGTCCGTGCGCTACATGGTTATGCGCTACGGCAGCCGG CTGTGGAAACTCCGTGTGCTACAGGCTGAGGTCAAGATCAACATCCGCCAGACCACCA CCGGCAGTGCCGTTCCCATCCGCCTGTTCATCACCAATGAGTCGGGCTACTACCTGGA CATCAGCCTCTACAAAGAAGTGACTGACTCCAGATCTGGAAATATCATGTTTCACTCC TTCGGCAACAAGCAAGGGCCCCAGCACGGGATGCTGATCAATACTCCCTACGTCACCA AGGATCTGCTCCAGGCCAAGCGATTCCAGGCCCAGACCCTGGGAACCACCTACATCTA TGACTTCCCGGAAATGTTCAGGCAGGCAAGTCCGGCGGCTCAGACGCGGGTACATGTG CACAATGTGCAGGCTCTCTTTAAACTGTGGGGCTCCCCAGACAAGTATCCCAAAGACA TCCTGACATACACTGAATTAGTGTTGGACTCTCAGGGCCAGCTGGTGGAGATGAACCG ACTTCCTGGTGGAAATGAGGTGGGCATGGTGGCCTTCAAAATGAGGTTTAAGACCCAG GAGTACCCGGAAGGACGGGATGTGATCGTCATCGGCAATGACATCACCTTTCGCATTG GATCCTTTGGCCCTGGAGAGGACCTTCTGTACCTGCGGGCATCCGAGATGGCCCGGGC AGAGGGCATTCCCAAAATTTACGTGGCAGCCAACAGTGGCGCCCGTATTGGCATGGCA GAGGAGATCAAACACATGTTCCACGTGGCTTGGGTGGACCCAGAAGACCCCCACAAAA AAAAAAAAACAGTGGCTTTCAGTGCAGGGAACTGGATTCGTAGCCTCACTAAAGTATT TTTTAAGGGATTTAAATACCTGTACCTGACTCCCCAAGACTACACCAGAATCAGCTCC CTGAACTCCGTCCACTGTAAACACATCGAGGAAGGAGGAGAGTCCAGATACATGATCA CGGATATCATCGGGAAGGATGATGGCTTGGGCGTGGAGAATCTGAGGGGCTCAGGCAT GATTGCTGGGGAGTCCTCTCTGGCTTACGAAGAGATCGTCACCATTAGCTTGGTGACC TGCCGAGCCATTGGGATTGGGGCCTACTTGGTGAGGCTGGGCCAGCGAGTGATCCAGG TGGAGAATTCCCACATCATCCTCACAGGAGCAAGTGCTCTCAACAAGGTCCTGGGAAG AGAGGTCTACACATCCAACAACCAGCTGGGTGGCGTTCAGATCATGCATTACAATGGT GTCTCCCACATCACCGTGCCAGATGACTTTGAGGGGGTTTATACCATCCTGGAGTGGC TGTCCTATATGCCAAAGGATAATCACAGCCCTGTCCCTATCATCACACCCACTGACCC CATTGACAGAGAAATTGAATTCCTCCCATCCAGAGCTCCCTACGACCCCCGGTGGATG CTTGCAGGAAGGCCTCACCCAACTCTGAAGGGAACGTGGCAGAGCGGATTCTTTGACC ACGGCAGTTTCAAGGAAATCATGGCACCCTGGGCGCAGACCGTGGTGACAGGACGAGC AAGGCTTGGGGGGATTCCCGTGGGAGTGATTGCTGTGGAGACACGGACTGTGGAGGTG GCAGTCCCTGCAGACCCTGCCAACCTGGATTCTGAGGCCAAGATAATTCAGCAGGCAG GACAGGTGTGGTTCCCAGACTCAGCCTACAAAACCGCCCAGGCCGTCAAGGACTTCAA CCGGGAGAAGTTGCCCCTGATGATCTTTGCCAACTGGAGGGGGTTCTCCGGTGGCATG AAAGACATGTATGACCAGGTGCTGAAGTTTGGAGCCTACATCGTGGACGGCCTTAGAC AATACAAACAGCCCATCCTGATCTATATCCCGCCCTATGCGGAGCTCCGGGGAGGCTC CTGGGTGGTCATAGATGCCACCATCAACCCGCTGTGCATAGAAATGTATGCAGACAAA GAGAGCAGGGGTGGTGTTCTGGAACCAGAGGGGACAGTGGAGATTAAGTTCCGAAAGA AAGATCTGATAAAGTCCATGAGAAGGATCGATCCAGCTTACAAGAAGCTCATGGAACA GCTAGGGGAACCTGATCTCTCCGACAAGGACCGAAAGGACCTGGAGGGCCGGCTAAAG GCTCGCGAGGACCTGCTGCTCCCCATCTACCACCAGGTGGCGGTGCAGTTCGCCGACT TCCATGACACACCCGGCCGGATGCTGGAGAAGGGCGTCATATCTGACATCCTGGAGTG GAAGACCGCACGCACCTTCCTGTATTGGCGTCTGCGCCGCCTCCTCCTGGAGGACCAG GTCAAGCAGGAGATCCTGCAGGCCAGCGGGGAGCTGAGTCACGTGCATATCCAGTCCA TGCTGCGTCGCTGGTTCGTGGAGACGGAGGGGGCTGTCAAGGCCTACTTGTGGGACAA CAACCAGGTGGTTGTGCAGTGGCTGGAACAGCACTGGCAGGCAGGGGATGGCCCGCGC TCCACCATCCGTGAGAACATCACGTACCTGAAGCACGACTCTGTCCTCAAGACCATCC GAGGCCTGGTTGAAGAAAACCCCGAGGTGGCCGTGGACTGTGTGATATACCTGAGCCA GCACATCAGCCCAGCTGAGCGGGCGCAGGTCGTTCACCTGCTGTCTACCATGGACAGC CCGGCCTCCACCTGA
ORF Start: ATG at 1 ORF Stop: TGA at 7495
SEQ ID NO: 218 2498 aa MW at 280484.4kD
NOV76a, MVLLLCLSCLIFSCLTFS LKIWGKMTDSKPITKSKSEANLIPSQEPFPASDNSGETP QRNGEGHTLPKTPSQAEPASHKGPKDAGRRRNSLPPSHQKPPRNPLSSSDAAPSPELQ
CG59641-01 Protein Sequence ANGTGTQGLEATDTNGLSSSARPQGQQAGSPSKEDKKQANIKRQLMTNFILGSFDDYS SDEDSVAGSSRESTRKGSRASLGALSLEAYLTTRPSMSGLHLVKRGREHKKLDLHRDF TVASPAEFVTRFGGDRVIEKVLIANNGIAAVKCMRSIRR AYEMFRNERAIRFWMVT PEDLKANAEYIKMADHYVPVPGGPNNNNYANVELIVDIAKRIPVQAVWAGWGHASENP
KLPELLCKNGVAFLGPPSEAM ALGDKIASTWAQTLQVPTLP SGSGLTVEWTEDDL QQGKRISVPEDVYDKGCVKDVDEGLEAAERIGFPLMIKASEGGGGKGIRKAESAEDFP ILFRQVQSEIPGSPIFLMKLAQHARHLEVQILADQYGNAVSLFGRDCSIQRRHQKIVE EAPATIAPLAIFEFMEQCAIRLAKTVGYVSAGTVEYLYSQDGSFHFLELNPRLQVEHP CTEMIADVNLPAAQLQIAMGVPLHRLKDIRLLYGESPWGVTPISFETPSNPPLARGHV IAARITSENPDEGFKPSSGTVQELNFRSSKNVWGYFSVAATGGLHEFADSQFGHCFS GENREEAISNMWALKELSIRGDFRTTVEYLINLLETESFQNNDIDTG LDYLIAEKV QAEKPDIMLGWCGALNVADAMFRTCMTDFLHSLERGQVLPADSLLNLVDVELIYGGV KYILKVARQSLTMFVLIMNGCHIEIDAHRLNDGGLLLSYNGNSYTTYMKEEVDSYRIT IGNKTCVFEKENDPTVLRSPSAGKLTQYTVEDGGHVEAGSSYAE EVMKMIMTLNVQE RGRVKYIKRPGAVLEAGCWARLELDDPSKVHPAEPFTGELPAQQTLPILGEKLHQVF HSVLENLTNVMSGFCLPEPVFSIKLKE VQKLMMTLRHPSLPLLELQEIMTSVAGRIP APVEKSVRRVMAQYASNITSVLCQFPSQQIATILDCHAATLQRKADREVFFINTQSIV QLVQRYRSGIRGYMKTWLDLLRRYLRVESKARDADANTSGMVGGVRSLSFTSV CFV SPESHYDKCVINLREQFKPDMSQVLDCIFSHAQVAKKNQLVIMLIDELCGPDPSLSDE LISILNELTQLSKSEHCKVALRARQILIASHLPSYELRHNQVESIFLSAIDMYGHQFC PENLKKLILSETTIFDVLPTFFYHANKWCMASLEVYVRRGYIAYELNSLQHRQLPDG TCWEFQFMLPSSHPNRMTVPISITNPDLLRHSTELFMDSGFSPLCQRMGAMVAFRRF EDFTRNFDEVISCFANVPKDTPLFSEARTSLYSEDDCKSLREEPIHILNVSIQCADHL EDEALVPILRTFVQSKKNILVDYGLRRITFLIAQEFAEDRIYRHLEPALAFQLELNRM RNFDLTAVPCANHKMHLYLGAAKVKEGVEVTDHRFFIRAIIRHSDLITKEASFEYLQN EGERLLLEAMDELEVAFNNTSVRTDCNHIFLNFVPTVIMDPFKIEESVRYMVMRYGSR L KLRVLQAEVKINIRQTTTGSAVPIRLFITNESGYYLDISLYKEVTDSRSGNIMFHS FGNKQGPQHGMLINTPYVTKDLLQAKRFQAQTLGTTYIYDFPEMFRQASPAAQTRVHV HNVQALFKLWGSPDKYPKDILTYTELVLDSQGQLVEMNRLPGGNEVGMVAFKMRFKTQ EYPEGRDVIVIGNDITFRIGSFGPGEDLLYLRASE ARAEGIPKIYVAANSGARIGMA EEIKH FHVA VDPEDPHKKKKTVAFSAGNWIRSLTKVFFKGFKYLYLTPQDYTRISS LNSVHCKHIEEGGESRYMITDIIGKDDGLGVENLRGSGMIAGESSLAYEEIVTISLVT CRAIGIGAYLVRLGQRVIQVENSHIILTGASALNKVLGREVYTSNNQLGGVQIMHYNG VSHITVPDDFEGVYTILE LSYMPKDNHSPVPIITPTDPIDREIEFLPSRAPYDPRWM LAGRPHPTLKGT QSGFFDHGSFKEIMAPWAQTWTGRARLGGIPVGVIAVETRTVEV AVPADPANLDSEAKIIQQAGQVWFPDSAYKTAQAVKDFNREKLPLMIFAN RGFSGGM KDMYDQVLKFGAYIVDGLRQYKQPILIYIPPYAELRGGSWWIDATINPLCIEMYADK ESRGGVLEPEGTVEIKFRKKDLIKSMRRIDPAYKKL EQLGEPDLSDKDRKDLEGRLK AREDLLLPIYHQVAVQFADFHDTPGRMLEKGVISDILEWKTARTFLYWRLRRLLLEDQ VKQEILQASGELSHVHIQSMLRRWFVETEGAVKAYLWDNNQWVQWLEQHWQAGDGPR STIRENITYLKHDSVLKTIRGLVEENPEVAVDCVIYLSQHISPAERAQWHLLSTMDS PAST
Further analysis of the NOV76a protein yielded the following properties shown in Table 76B.
Table 76B. Protein Sequence Properties NOV76a
PSort 0.6850 probability located in endoplasmic reticulum (membrane); 0.6400 analysis: probability located in plasma membrane; 0.4600 probability located in Golgi body; 0.1000 probability located in endoplasmic reticulum (lumen)
SignalP Likely cleavage site between residues 25 and 26 analysis:
A search of the NOV76a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 76C.
In a BLAST search of public sequence databases, the NOV76a protein was found to have homology to the proteins shown in the BLASTP data in Table 76D.
PFam analysis predicts that the NOV76a protein contains the domains shown in the Table 76E.
Example 77.
The NOV77 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 77A.
Further analysis of the NOV77a protein yielded the following properties shown in Table 77B.
Table 77B. Protein Sequence Properties NOV77a
PSort 0.3000 probability located in microbody (peroxisome); 0.3000 probability analysis: located in nucleus; 0.1526 probability located in lysosome (lumen); 0.1000 probability located in mitochondrial matrix space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV77a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 77C.
In a BLAST search of public sequence databases, the NOV77a protein was found to have homology to the proteins shown in the BLASTP data in Table 77D.
PFam analysis predicts that the NOV77a protein contains the domains shown in the Table 77E.
Example 78.
The NOV78 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 78A.
Further analysis of the NOV78a protein yielded the following properties shown in Table 78B.
Table 78B. Protein Sequence Properties NOV78a
PSort 0.8000 probability located in microbody (peroxisome); 0.1000 probability located analysis: in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen); 0.0000 probability located in endoplasmic reticulum (membrane)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV78a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 78C.
In a BLAST search of public sequence databases, the NOV78a protein was found to have homology to the proteins shown in the BLASTP data in Table 78D.
Table 78D. Public BLASTP Results for NOV78a
PFam analysis predicts that the NOV78a protein contains the domains shown in the Table 78E.
Example 79.
The NOV79 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 79A.
Table 79 A. NOV79 Sequence Analysis
SEQ ED NO: 223 4203 bp
NOV79a, AATGTGATGGGATCACTAGCATGTCTGCGGAGAGCGGCCCTGGGACGAGATTGAGAAA
TCTGCCAGTAATGGGGGATGGACTAGAAACTTCCCAAATGTCTACAACACAGGCCCAG
CG59452-01 DNA Sequence GCCCAACCCCAGCCAGCCAACGCAGCCAGCACCAACCCCCCGCCCCCAGAGACCTCCA ACCCTAACAAGCCCAAGAGGCAGACCAACCAACTGCAATACCTGCTCAGAGTGGTGCT CAAGACACTATGGAAACACCAGTTTGCATGGCCTTTCCAGCAGCCTGTGGATGCCGTC AAGCTGAACCTCCCTGATTACTATAAGATCATTAAAACGCCTATGGATATGGGAACAA TAAAGAAGCGCTTGGAAAACAACTATTACTGGAATGCTCAGGAATGTATCCAGGACTT CAACACTATGTTTACAAATTGTTACATCTACAACAAGCCTGGAGATGACATAGTCTTA ATGGCAGAAGCTCTGGAAAAGCTCTTCTTGCAAAAAATAAATGAGCTACCCACAGAAG AAACCGAGATCATGATAGTCCAGGCAAAAGGAAGAGGACGTGGGAGGAAAGAAACAGG TACAGCAAAACCTGGCGTTTCCACGGTACCAAACACAACTCAAGCATCGACTCCTCCG CAGACCCAGACCCCTCAGCCGAATCCTCCTCCTGTGCAGGCCACGCCTCACCCCTTCC CTGCCGTCACCCCGGACCTCATCGTCCAGACCCCTGTCATGACAGTGGTGCCTCCCCA GCCACTGCAGACGCCCCCGCCAGTGCCCCCCCAGCCACAACCCCCACCCGCTCCAGCT CCCCAGCCCGTACAGAGCCACCCACCCATCATCGCGGCCACCCCACAGCCTGTGAAGA CAAAGAAGGGAGTGAAGAGGAAAGCAGACACCACCACCCCCACCACCATTGACCCCAT TCACGAGCCACCCTCGCTGCCCCCGGAGCCCAAGACCACCAAGCTGGGCCAGCGGCGG GAGAGCAGCCGGCCTGTGAAACCTCCAAAGAAGGACGTGCCCGACTCTCAGCAGCACC CAGCACCAGAGAAGAGCAGCAAGGTCTCGGAGCAGCTCAAGTGCTGCAGCGGCATCCT CAAGGAGATGTTTGCCAAGAAGCACGCCGCCTACGCCTGGCCCTTCTACAAGCCTGTG GACGTGGAGGCACTGGGCCTACACGACTACTGTGACATCATCAAGCACCCCATGGACA TGAGCACAATCAAGTCTAAACTGGAGGCCCGTGAGTACCGTGATGCTCAGGAGTTTGG TGCTGACGTCCGATTGATGTTCTCCAACTGCTATAAGTACAACCCTCCTGACCATGAG GTGGTGGCCATGGCCCGCAAGCTCCAGGATGTGTTCGAAATGCGCTTTGCCAAGATGC CGGACGAGCCTGAGGAGCCAGTGGTGGCCGTGTCCTCCCCGGCAGTGCCCCCTCCCAC CAAGGTTGTGGCCCCGCCCTCATCCAGCGACAGCAGCAGCGATAGCTCCTCGGACAGT GACAGTTCGACTGATGACTCTGAGGAGGAGCGAGCCCAGCGGCTGGCTGAGCTCCAGG AGCAGCTCAAAGCCGTGCACGAGCAGCTTGCAGCCCTCTCTCAGCCCCAGCAGAACAA ACCAAAGAAAAAGGAGAAAGACAAGAAGGAAAAGAAAAAAGAAAAGCACAAAAGGAAA GAGGAAGTGGAAGAGAATAAAAAAAGCAAAGCCAAGGAACCTCCTCCTAAAAAGACGA AGAAAAATAATAGCAGCAACAGCAATGTGAGCAAGAAGGAGCCAGCGCCCATGAAGAG CAAGCCCCCTCCCACGTATGAGTCGGAGGAAGAGGACAAGTGCAAGCCTATGTCCTAT GAGGAGAAGCGGCAGCTCAGCTTGGACATCAACAAGCTCCCCGGCGAGAAGCTGGGCC GCGTGGTGCACATCATCCAGTCACGGGAGCCCTCCCTGAAGAATTCCAACCCCGACGA GATTGAAATCGACTTTGAGACCCTGAAGCCGTCCACACTGCGTGAGCTGGAGCGCTAT GTCACCTCCTGTTTGCGGAAGAAAAGGAAACCTCAAGCTGAGAAAGTTGATGTGATTG CCGGCTCCTCCAAGATGAAGGGCTTCTCGTCCTCAGAGTCGGAGAGCTCCAGTGAGTC CAGCTCCTCTGACAGCGAAGACTCCGAAACAGAGATGGCTCCGAAGTCAAAAAAGAAG GGGCACCCCGGGAGGGAGCAGAAGCAGCACCATCATCACCACCATCAGCAGATGCAGC AGGCCCCGGCTCCTGTGCCCCAGCAGCCGCCCCCGCCTCCCCAGCAGCCCCCACCGCC TCCACCTCCGCAGCAGCAACAGCAGCCGCCACCCCCGCCTCCCCCACCCTCCATGCCG CAGCAGGCAGCCCCGGCGATGAAGTCCTCGCCCCCACCCTTCATTGCCACCCAGGTGC CCGTCCTGGAGCCCCAGCTCCCAGGCAGCGTCTTTGACCCCATCGGCCACTTCACCCA GCCCATCCTGCACCTGCCGCAGCCTGAGCTGCCCCCTCACCTGCCCCAGCCGCCTGAG CACAGCACTCCACCCCATCTCAACCAGCACGCAGTGGTCTCTCCTCCAGCTTTGCACA ACGCACTACCCCAGCAGCCATCACGGCCCAGCAACCGAGCCGCTGCCCTGCCTCCCAA GCCCGCCCGGCCCCCAGCCGTGTCACCAGCCTTGACCCAAACACCCCTGCTCCCACAG CCCCCCATGGCCCAACCCCCCCAAGTGCTGCTGGAGGATGAAGAGCCACCTGCCCCAC CCCTCACCTCCATGCAGATGCAGCTGTACCTGCAGCAGCTGCAGAAGGTGCAGCCCCC TACGCCGCTACTCCCTTCCGTGAAGGTGCAGTCCCAGCCCCCACCCCCCCTGCCGCCC CCACCCCACCCCTCTGTGCAGCAGCAGCTGCAGCAGCAGCCGCCACCACCCCCACCAC CCCAGCCCCAGCCTCCACCCCAGCAGCAGCATCAGCCCCCTCCACGGCCCGTGCACTT GCAGCCCATGCAGTTTTCCACCCACATCCAACAGCCCCCGCCACCCCAGGGCCAGCAG CCCCCCCATCCGCCCCCAGGCCAGCAGCCACCCCCGCCGCAGCCTGCCAAGCCTCAGC AAGTCATCCAGCACCACCATTCACCCCGGCACCACAAGTCGGACCCCTACTCAACCGG TCACCTCCGCGAAGCCCCCTCCCCGCTTATGATACATTCCCCCCAGATGTCACAGTTC CAGAGCCTGACCCACCAGTCTCCACCCCAGCAAAACGTCCAGCCTAAGAAACAGGTAA CTGGCAGGGCTGGGCCAAGTCCTGTGGGCCAGGGCCGGGGGTGCCTGCCCACCTCACC GGCCGCTGTGCCTGTGCCATCCCAGGAGCTGCGTGCTGCCTCCGTGGTCCAGCCCCAG CCCCTCGTGGTGGTGAAGGAGGAGAAGATCCACTCACCCATCATCCGCAGCGAGCCCT TCAGCCCCTCGCTGCGGCCGGAGCCCCCCAAGCACCCGGAGAGCATCAAGGCCCCCGT TTATGTTCCAGGGCCGGAAATGAAGCCTGTGGATGTCGGGAGGCCTGTGATCCGGCCC CCAGAGCAGAACGCACCGCCACCAGGGGCCCCTGACAAGGACAAACAGAAACAGGAGC CGAAGACTCCAGTTGCGCCCAAAAAGGACCTGAAAATCAAGAACATGGGCTCCTGGGC CAGCCTAGTGCAGAAGCATCCGACCACCCCCTCCTCCACAGCCAAGTCATCCAGCGAC AGCTTCGAGCAGTTCCGCCGCGCCGCTCGGGAGAAAGAGGAGCGTGAGAAGGCCCTGA AGGCTCAGGCCGAGCACGCTGAGAAGGAGAAGGAGCGGCTGCGGCAGGAGCGCATGAG GAGCCGAGAGGACGAGGATGCGCTGGAGCAGGCCCGGCGGGCCCATGAGGAGGCACGT CGGCGCCAGGAGCAGCAGCAGCAGCAGCGCCAGGAGCAACAGCAGCAGCAGCAACAGC AAGCAGCTGCGGTGGCTGCCGCCGCCACCCCACAGGCCCAGAGCTCCCAGCCCCAGTC CATGCTGGACCAGCAGAGGGAGTTGGCCCGGAAGCGGGAGCAGGAGCGAAGACGCCGG GAAGCCATGGCAGCTACCATTGACATGAATTTCCAGAGTGATCTATTGTCAATATTTG AAGAAAATCTTTTCTGAGCGCACCTAG
Further analysis of the NOV79a protein yielded the following properties shown in Table 79B.
Table 79B. Protein Sequence Properties NOV79a
PSort 0.9800 probability located in nucleus; 0.3000 probability located in microbody analysis: (peroxisome); 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV79a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 79C.
In a BLAST search of public sequence databases, the NOV79a protein was found to have homology to the proteins shown in the BLASTP data in Table 79D.
PFam analysis predicts that the NOV79a protein contains the domains shown in the Table 79E.
Table 79E. Domain Analysis of NOV79a
Example 80.
The NOV80 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 80A.
Table 80A. NOV80 Sequence Analysis
SEQ ID NO: 225 1776 bp
NOV80a, TGGTTCGTTTATTCCTGGGGTTGTCATATCATGGCTTATAATGACACAGACAGAAACC
AGACTGAGAAGCTCCTAAAAAGAGTACGAGAACTGGAGCAAGAGGTGCAAAGACTTAA
CG59572-01 DNA Sequence AAAGGAACAGGCCAAAAATAAGGAGGACTCAAACATTAGAGAAAATTCAGCAGGAGCT GGAAAAACTAAGCGTGCATTTGATTTCAGTGCTCATGGCCGAAGACACGTAGCCCTAA GAATAGCCTATATGGGCTGGGGATACCAGGGCTTTGCTAGTCAGGAAAACACAAATAA TACCATTGAAGAGAAACTGTTTGAAGCTCTAACCAAGACTCGACTAGTAGAAAGCAGA CAGACATCCAACTATCACCGATGTGGGAGAACAGATAAAGGAGTTAGTGCCTTTGGAC AGGTGATCTCACTTGACCTTCGCTCTCAGTTTCCAAGGGGCAGGGATTCCGAGGACTT TAATGTAAAAGAGGAGGCTAATGCTGCTGCTGAAGAGATCCGTTATACCCACATTCTC AATCGGGTACTCCCTCCAGACATCCGTATATTGGCCTGGGCCCCTGTAGAACCAAGCT TCAGTGCTAGGTTCAGCTGCCTTGAGCGGACTTACCGCTATTTTTTCCCTCGTGCTGA TTTAGATATTGTAACCATGGATTATGCAGCTCAGAAGTATGTTGGCACCCATGATTTC AGGAACTTGTGTAAAATGGATGTAGCCAACGGTGTGATTAATTTTCAGAGGACTATTC TATCTGCTCAAGTACAGCTAGTGGGCCAGAGCCCAGGTGAGGGGAGATGGCAAGAACC TTTCCAGTTATGTCAGTTTGAAGTGACTGGCCAGGCATTCCTTTATCATCAAGTCCGA TGTATGATGGCTATCCTCTTTCTGATTGGCCAAGGAATGGAGAAGCCAGAGATTATTG ATGAGCTGCTGAATATAGAGAAAAATCCCCAAAAGCCTCAATATAGTATGGCTGTAGA ATTTCCTCTAGTCTTATATGACTGTAAGTTTGAAAATGTCAAGTGGATCTATGACCAG GAGGCTCAGGAGTTCAATATTACCCACCTACAACAACTGTGGGCTAATCATGCTGTCA AAACTCACATGTTGTATAGTATGCTACAAGGACTGGACACTGTTCCAGTACCCTGTGG AATAGGACCAAAGATGGATGGAATGACAGAATGGGGAAATGTTAAGCCCTCTGTCATA AAGCAGACCAGTGCCTTTGTAGAAGGAGTGAAGATGCGCACATATAAGCCCCTCATGG ACCGTCCTAAATGCCAAGGACTGGAATCCCGGATCCAGCATTTTGTACGTAGGGGACG AATTGAGCACCCACATTTATTCCATGAGGAAGAAACAAAAGCCAAAAGGGACTGTAAT GACACACTAGAGGAAGAGAATACTAATTTGGAGACACCAACGAAGAGGGTCTGTGTTG ACACAGAAATTAAAAGCATCATTTAACCATAGACAATTTGCCAGGATCTAGGAACCAC CTAATGGTAGGTGGACAGAAAAGGAAAAAAAAAAAAATTTACTTGCAAGTACTAGGAA
TTCAGATGATCAGCTCTTAAAAAAAAAAAAAAAGCAAAAAGACTAAAGCCCTATTAAG
GAAGTTATTGCTTTAATAAGAAATTTCAAATATTCTCTTATCCCGGTCCAAAAGGATT
AAGCGATTAAAGAACGTAAAATGGAGATGTATTTACATACACCTGGAAACCTGTGCCT
TGTATTCAAATTCATTAAAGCCTAATCCTGCAAGAA
ORF Start: ATG at 31 ORF Stop: TAA at 1474
SEQ ID NO: 226 481 aa MW at 55646.8kD
NOV80a, MAYNDTDRNQTEKLLKRVRELEQEVQRLKKEQAKNKEDSNIRENSAGAGKTKRAFDFS AHGRRHVALRIAYMG GYQGFASQEMTNNTIEEKLFEALTKTRLVESRQTSNYHRCGR
CG59572-01 Protein Sequence TDKGVSAFGQVISLDLRSQFPRGRDSEDFNVKEEANAAAEEIRYTHILNRVLPPDIRI LA APVEPSFSARFSCLERTYRYFFPRADLDIVTMDYAAQKYVGTHDFRNLCKMDVAN GVINFQRTILSAQVQLVGQSPGEGR QEPFQLCQFEVTGQAFLYHQVRCMMAILFLIG QGMEKPEIIDELLNIEKNPQKPQYSMAVEFPLVLYDCKFENVK IYDQEAQEFNITHL QQL ANHAVKTHMLYSMLQGLDTVPVPCGIGPKMDGMTEWGNVKPSVIKQTSAFVEGV KMRTYKPLMDRPKCQGLESRIQHFVRRGRIEHPHLFHEEETKAKRDCNDTLEEENTNL ETPTKRVCVDTEIKSII
SEQ ED NO: 227 1508 bp
NOV80b, CATGGCTTATAATGACACAGACAGAAACCAGACTGAGAAGCTCCTAAAAAGAGTACGA GAACTGGAGCAAGAGGTGCAAAGACTTAAAAAGGAACAGGCCAAAAATAAGGAGGACT
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 80B.
Further analysis of the NOV80a protein yielded the following properties shown in Table 80C.
Table 80C. Protein Sequence Properties NOV80a
PSort 0.6500 probability located in cytoplasm; 0.1000 probability located in analysis: mitochondrial matrix space; 0.1000 probability located in lysosome (lumen); 0.0142 probability located in microbody (peroxisome)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV80a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 80D.
In a BLAST search of public sequence databases, the NOV80a protein was found to have homology to the proteins shown in the BLASTP data in Table 80E.
PFam analysis predicts that the NOV80a protein contains the domains shown in the Table 80F.
Example 81.
The NOV81 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 81 A.
TCCAGGGACACCTGCATCGTCATCTCAGGGGAGAGTGGGGCAGGGAAGACAGAAGCCA GTAAGCACATCATGCAGTACATCGCTGCTGTCACCAATCCAAGCCAGAGGGCTGAGGT GGAGAGGGTCAAGGACGTGCTGCTCAAGTCCACCTGTGTGCTGGAGGCCTTTGGCAAT GCCCGCACCAACCGCAATCACAACTCCAGCCGCTTTGGCAAGTACATGGACATCAACT TTGACTTCAAGGGGGACCCGATCGGAGGACACATCCACAGCTACCTACTGGAGAAGTC TCGGGTCCTCAAGCAGCACGTGGGTGAAAGAAACTTCCACGCCTTCTACCAATTGCTG AGAGGCAGTGAGGACAAGCAGCTGCATGAACTGCACTTGGAGAGAAACCCTGCTGTAT ACAATTTCACACACCAGGGAGCAGGACTCAACATGACTGTGAGTGATGAGCAGAGCCA CCAGGCAGTGACCGAGGCCATGAGGGTCATCGGCTTCAGTCCTGAAGAGGTGGAGTCT GTGCATCGCATCCTGGCTGCCATATTGCACCTGGGAAACATCGAGTTTGTGGAGACGG AGGAGGGTGGGCTGCAGAAGGAGGGCCTGGCAGTGGCCGAGGAGGCACTGGTGGACCA TGTGGCTGAGCTGACGGCCACACCCCGGGACCTCGTGCTCCGCTCCCTGCTGGCTCGC ACAGTTGCCTCGGGAGGCAGGGAACTCATAGAGAAGGGCCACACTGCAGCTGAGGCCA GCTATGCCCGGGATGCCTGTGCCAAGGCAGTGTACCAGCGGCTGTTTGAGTGGGTGGT GAACAGGATCAACAGTGTCATGGAACCCCGGGGCCGGGATCCTCGGCGTGATGGCAAG GACACAGTCATTGGCGTGCTGGACATCTATGGCTTCGAGGTGTTTCCCGTCAACAGTT TCGAGCAGTTCTGCATCAACTACTGCAACGAGAAGCTGCAGCAGCTATTCATCCAGCT CATCCTGAAGCAGGAACAGGAAGAGTACGAGCGCGAGGGCATCACCTGGCAGAGCGTT GAGTATTTCAACAACGCCACCATTGTGGATCTGGTGGAGCGGCCCCACCGTGGCATCC TGGCCGTGCTGGACGAGGCCTGCAGCTCTGCTGGCACCATCACTGACCGAATCTTCCT GCAGACCCTGGACATGCACCACCGCCATCACCTACACTACACCAGCCGCCAGCTCTGC CCCACAGACAAGACCATGGAGTTTGGCCGAGACTTCCGGATCAAGCACTATGCAGGGG ACGTCACGTACTCCGTGGAAGGCTTCATCGACAAGAACAGAGATTTCCTCTTCCAGGA CTTCAAGCGGCTGCTGTACAACAGCACGGACCCCACTCTACGGGCCATGTGGCCGGAC GGGCAGCAGGACATCACAGAGGTGACCAAGCGCCCCCTGACGGCTGGCACACTCTTCA AGAACTCCATGGTGGCCCTGGTGGAGAACCTTGCCTCCAAGGAGCCCTTCTACGTCCG CTGCATCAAGCCCAATGAGGACAAGGTAGCTGGGAAGCTGGATGAGAACCACTGTCGC CACCAGGTCGCATACCTGGGGCTGCTGGAGAATGTGAGGGTCCGCAGGGCTGGCTTCG CTTCCCGCCAGCCCTACTCTCGATTCCTGCTCAGGTACAAGATGACCTGTGAATACAC ATGGCCCAACCACCTGCTGGGCTCCGACAAGGCAGCCGTGAGCGCTCTCCTGGAGCAG CACGGGCTGCAGGGGGACGTGGCCTTTGGCCACAGCAAGCTGTTCATCCGCTCACCCC GGACACTGGTCACACTGGAGCAGAGCCGAGCCCGCCTCATCCCCATCATTGTGCTGCT ATTGCAGAAGGCATGGCGGGGCACCTTGGCGAGGTGGCGCTGCCGGAGGCTGAGGGCT ATCTACACCATCATGCGCTGGTTCCGGAGACACAAGGTGCGGGCTCACCTGGCTGAGC TGCAGCGGCGATTCCAGGCTGCAAGGCAGCCGCCACTCTACGGGCGTGACCTTGTGTG GCCGCTGCCCCCTGCTGTGCTGCAGCCCTTCCAGGACACCTGCCACGCACTCTTCTGC AGGTGGCGGGCCCGGCAGCTGGTGAAGAACATCCCCCCTTCAGACATGCCCCAGATCA AGGCCAAGGTGGCCGCCATGGGGGCCCTGCAAGGGCTTCGTCAGGACTGGGGCTGCCG ACGGGCCTGGGCCCGAGACTACCTGTCCTCTGCCACTGACAATCCCACAGCATCAAGC CTGTTTGCTCAGCGACTAAAGACACTTCAGGACAAAGATGGCTTCGGGGCTGTGCTCT TTTCAAGCCATGTCCGCAAGGTGAACCGCTTCCACAAGATCCGGAACCGGGCCCTCCT GCTCACAGACCAGCACCTCTACAAGCTGGACCCTGACCGGCAGTACCGGGTGATGCGG GCCGTGCCCCTTGAGGCGGTGACGGGGCTGAGCGTGACCAGCGGAGGAGACCAGCTGG TGGTGCTGCACGCCCGCGGCCAGGACGACCTCGTGGTGTGCCTGCACCGCTCCCGGCC GCCATTGGACAACCGCGTTGGGGAGCTGGTGGGCGTGCTGGCCGCACACTGCCGCAGG GAGGGCCGCACCCTGGAGGTTCGCGTCTCCGACTGCATCCCACTAAGCCATCGCGGGG TCCGGCGCCTCATCTCCGTGGAGCCCAGGCCGGAGCAGCCAGAGCCCGATTTCCGCTG CGCTCGCGGCTCCTTCACCCTGCTCTGGCCCAGCCGCTGAGCGCCCGCACCCGCCGCA CCCCGA
ORF Start: ATG at 15 ORF Stop: TGA at 3054
SEQ ED NO: 230 1013 aa MW at 116044.5kD
NOV81a, MEDEEGPEYGKPDFVLLDQVTMEDFMRNLQLRFEKGRIYTYIGEVLVSVNPYQELPLY GPEAIARYQGRELYERPPHLYAVANAAYKAMKHRSRDTCIVISGESGAGKTEASKHIM
CG59522-01 Protein Sequence QYIAAVTNPSQRAEVERVKDVLLKSTCVLEAFGNARTNRNHNSSRFGKYMDINFDFKG DPIGGHIHSYLLEKSRVLKQHVGERNFHAFYQLLRGSEDKQLHELHLERNPAVYNFTH QGAGLNMTVSDEQSHQAVTEAMRVIGFSPEEVESVHRILAAILHLGNIEFVETEEGGL QKEGLAVAEEALVDHVAELTATPRDLVLRSLLARTVASGGRELIEKGHTAAEASYARD ACAKAVYQRLFEWWNRINSVMEPRGRDPRRDGKDTVIGVLDIYGFEVFPVNSFEQFC INYCNEKLQQLFIQLILKQEQEEYEREGIT QSVEYFNNATIVDLVERPHRGILAVLD EACSSAGTITDRIFLQTLDMHHRHHLHYTSRQLCPTDKT EFGRDFRIKHYAGDVTYS VEGFIDKNRDFLFQDFKRLLYNSTDPTLRAM PDGQQDITEVTKRPLTAGTLFKNSMV ALVENLASKEPFYVRCIKPNEDKVAGKLDENHCRHQVAYLGLLENVRVRRAGFASRQP YSRFLLRYKMTCEYT PNHLLGSDKAAVSALLEQHGLQGDVAFGHSKLFIRSPRTLVT LEQSRARLIPIIVLLLQKA RGTLARWRCRRLRAIYTIMR FRRHKVRAHLAELQRRF QAARQPPLYGRDLV PLPPAVLQPFQDTCHALFCR RARQLVKNIPPSDMPQIKAKVA AMGALQGLRQD GCRRA ARDYLSSATDNPTASSLFAQRLKTLQDKDGFGAVLFSSHV RKVNRFHKIRNRALLLTDQHLYKLDPDRQYRVMRAVPLEAVTGLSVTSGGDQLWLHA RGQDDLWCLHRSRPPLDNRVGELVGVLAAHCRREGRTLEVRVSDCIPLSHRGVRRLI SVEPRPEQPEPDFRCARGSFTLL PSR
Further analysis of the NOV8 la protein yielded the following properties shown in
Table 8 IB.
Table 81B. Protein Sequence Properties NOV81a
PSort 0.8800 probability located in nucleus; 0.3902 probability located in microbody analysis: (peroxisome); 0.2210 probability located in lysosome (lumen); 0.1000 probability located in mitochondrial matrix space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV8 la protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 81C.
In a BLAST search of public sequence databases, the NOV8 la protein was found to have homology to the proteins shown in the BLASTP data in Table 8 ID.
Table 81D. Public BLASTP Results for NOV81a
Protein/Organism/Length
PFam analysis predicts that the NOV8 la protein contains the domains shown in the Table 8 IE.
Example 82.
The NOV82 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 82A.
Further analysis of the NOV82a protein yielded the following properties shown in Table 82B.
Table 82B. Protein Sequence Properties NOV82a
PSort 0.4066 probability located in microbody (peroxisome); 0.3000 probability analysis: located in nucleus; 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV82a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 82C.
In a BLAST search of public sequence databases, the NOV82a protein was found to have homology to the proteins shown in the BLASTP data in Table 82D.
PFam analysis predicts that the NOV82a protein contains the domains shown in the Table 82E.
Example 83.
The NOV83 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 83A.
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 83B.
Further analysis of the NOV83a protein yielded the following properties shown in Table 83C.
Table 83C. Protein Sequence Properties NOV83a
PSort 0.6500 probability located in cytoplasm; 0.1000 probability located in analysis: mitochondrial matrix space; 0.1000 probability located in lysosome (lumen); 0.0000 probability located in endoplasmic reticulum (membrane)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV83a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 83D.
Table 83D. Geneseq Results for NOV83a
NOV83a Identities/
Geneseq Protein/Organism/Length [Patent Residues/ Similarities for Expect Identifier #, Date] Match the Matched Value Residues Region
AAM79976 Human protein SEQ TD NO 3622 - 1..101 100/101 (99%) le-52 Homo sapiens, 125 aa. 25..125 100/101 (99%) [WO200157190-A2, 09-AUG-2001]
AAM78992 Human protein SEQ ED NO 1654 - 1..101 100/101 (99%) le-52 Homo sapiens, 101 aa. 1..101 100/101 (99%) [WO200157190-A2, 09-AUG-2001]
AAY49967 Human sentrin protein sequence - 1..101 89/101 (88%) 2e-45 Homo sapiens, 101 aa. [US5985664- 1..101 94/101 (92%) A, 16-NOV-1999]
AAW87984 Ubiquitin-like domain of the protein 1..101 89/101 (88%) 2e-45 SUMO1 - Mammalia, 101 aa. 1..101 94/101 (92%) [WO9857978-A1, 23-DEC-1998]
AAW60079 Homo sapiens sentrin- 1 polypeptide 1..101 89/101 (88%) 2e-45 - Homo sapiens, 101 aa. 1..101 94/101 (92%) [WO9820038-A1, 14-MAY-1998]
In a BLAST search of public sequence databases, the NOV83a protein was found to have homology to the proteins shown in the BLASTP data in Table 83E.
PFam analysis predicts that the NOV83a protein contains the domains shown in the Table 83F.
Example 84.
The NOV84 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 84A.
Table 84A. NOV84 Sequence Analysis
SEQ ID NO: 237 912 bp
NOV84a, ACTCACTAATGGGCTCGAGCGGCTGCCTGTGTTTCAGCGGCTCGGGGAAATCCACCGT
GGGCGCCCTGCTGGCATCTGAGCTGGGATGGAAATTCTATGATGCTGATGATTATCAC
CG59586-01 DNA Sequence CCGGAGGAAAATCGAAGGAAGATGGGAAAAGGCATACCGCTCAATGACCAGGACCGGA TTCCATGGCTCTGTAACTTGCATGACATTTTACTAAGAGATGTAGCCTCGGGACAGCG TGTGGTTCTAGCCTGTTCAGCCCTGAAGAAAACGTACAGAGACATATTAACACAAGGA AAAGATGGTGTAGCTCTGAAGTGTGAGGAGTCGGGAAAGGAAGCAAAGCAGGCTGAGA TGCAGCTCCTGGTGGTCCATCTGAGCGGGTCGTTTGAGGTCATCTCTGGACGCTTACT CAAAAGAGAGGGACATTTTATGCCCCCTGAATTATTGCAGTCCCAGTTTGAGACTCTG GAGCCCCCAGCAGCTCCAGAAAACTTTATCCAAATAAGTGTGGACAAAAATGTTTCAG AGATAATTGCTACAATTATGGAAACCCTAAAAATGAAATGACAATGATTTTGTATCAG
TGGTCCAAACAGAACTAAGCATAAATCATTGTGCCATCCCAAACCTCGTTCCAGCCGC
CTTGCCCATACTAGATTCTAAATGTTTCTAAAGGCAAACCCCAATGTGTCAAGACAGA
CTTGTTTAGGTGTAATTTTAGGAATTATGCTGGTTCATCAGGAAGCAGAGGGGGAGTT
TTAAAAGTCAAGCTTAAATTGAAGTTTAAATTCATCTATAACCAAATCAAATGATCAG
AGGAAATTCTGTAATCAATGCTGGAAATCGTTACATTGTTTAGAACATTCTTGCTCAT
GCCTGTATTTGCACAAATAAATGAAACTTCGCTGTAAAAAAA
ORF Start: ATG at 9 ORF Stop: TGA at 561
SEQ ID NO: 238 184 aa MW at 20352.2kD
NOV84a, MGSSGCLCFSGSGKSTVGALLASELGWKFYDADDYHPEENRRKMGKGI PLNDQDRI PW LCNLHDILLRDVASGQRWLACSALKKTYRDILTQGKDGVALKCEESGKEAKQAEMQL
CG59586-01 Protein Sequence LWHLSGSFEVISGRLLKREGHFMPPELLQSQFETLEPPAAPENFIQISVDKNVSEII ATIMETLKMK
Further analysis of the NOV84a protein yielded the following properties shown in Table 84B.
Table 84B. Protein Sequence Properties NOV84a
PSort 0.6500 probability located in cytoplasm; 0.1000 probability located in analysis: mitochondrial matrix space; 0.1000 probability located in lysosome (lumen); 0.1000 probability located in plasma membrane
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV84a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 84C.
In a BLAST search of public sequence databases, the NOV84a protein was found to have homology to the proteins shown in the BLASTP data in Table 84D.
PFam analysis predicts that the NOV84a protein contains the domains shown in the Table 84E.
Example 85.
The NOV85 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 85A.
TCTCTGATTGCCGAGATGGAGGTGTTCACGATTGAGCTGTCGAAGAACTTCGGTGTCA AGGAATGGCACGAGAGCCTCGCGAAGTTGCTGCTCGAGTGTGGCAAGGACGAGAAGAA GCGGACGTTTCTCTTCGCCGACACCCAGCTGGCGCATCCGACGTTTCTGGAGGATGTG GCGGGCCTGCTCACATCGGGTGATGTGCCGAACCTCTTTGAGGACCAAGATATCGAGC TCATCAACGACAAGTTTCGCGGCGTCTGCCTAAGCGAGAACCTGCCAACGACGAAGGT GTCGGTGTACGCGCGCTTTGTGAAGGAGGCGCGAGCCAACCTGCACCTTGTGCTCGCC TTCTCTCCCATCGGAGAGGCGTTTCGCAGCCGCCTGCGTATGTTCCCATCGCTCATTG CGTGCTGCACAATCGACTGGTTTGCTGAGTGGCCATCCGAGGCGCTACTGTCGGTAGC CGCAGTGCAGCTGAACGCCGGCGACGTTACTGACGTCATGGGGGCGGCAAGCCATGCC GACTTGCCGGGCTGCTTCCAGGCAGTGCACCGCGCGGCGGCGGAGGTGACGGAGCGCT TCTTCACGGAAACGCGTCGTCGCTCGTACGTGACGCCGACGTCCTATCTGTCGCTCCT CTCCAACTTCAAAGTGATGGCGGCGGCAAAACGCCGCTTCGTTCGCGAGCAGCGCGGC CGCCTCGAGAAGGGGCTGGAGAAGCTGCGGCACACCGAGGTGCAAGTGGCGGAGCTGG AGGCCCAGCTCAAGGCGCAGCAGCCGGTTCTGGTGCAGAAAAAGGCAGAGATTCAGTC GATGATGGAGCGGCTGACGGTGGACCGAAAGGAGGCGGCGGTGAAGGAGGCGGACGCG CGCAGGGAGGCCCAGCTTCCCGGTGGCCGTGCTGCATACGGCGGTGAAGATGACGAAT GAGCCGCCGATGGGGCTGCGGGCGAACGTGATGCGCTCCTACTACGGCTTCACTCCCG AGGACCTCGAGCAGGAGGAGAAGCCCGCCGAGTTCAAAAAGATGTTGATGGCATCCGC
ATGCCTGGTCCCATACCCGAGCACTGAAGAGCAGGGTCTCTGGAGCCTGGCATCGTGG
GGTGGCCCTCAGCTTCCCCACTCACTGTGGGAAGTTTCCTTAGTGTCTCTGAGCCTGT
TTCCTCATCCGTTGCCTGAGGATAAACCTGCTTCAGGATTGTTGGTGAAAAGACTTCC
CTCACCTAGCTTCTGTAACGCCACTGCATGCCACCACTGCTGAGTACTGTTTGTTTGC
TAGGTTGGTGTCATTCTCATTTTACCAGAAAGTGAAGCTC
ORF Start: ATG at 41 ORF Stop: TGA at 3944
SEQ ID NO: 240 1301 aa MW at 146115.7kD
NOV85a, MNNYVLNDEIGQGAFSTIYKGRYRTTTEFYAIASIDKKRRERWNCVQLLRSMHHSNV IEFHNWYETNNHLWIITEYCTGGDMSTILRSNINLTΕQAVQAFGRDVAMGLMYIHSKG
CG59704-01 Protein Sequence WYNDLQTRNLLMDSAAMLRFHDFSLACLFQDAATRPLVGTPLYMAPELFMADRPLYS MASDL SFGCVLHELATGKPPFAASDLETLLGDILTSPTPAVPGAPESFQTLLCGLLE KDPLKRYAWVDWRSEF DEPLPLPSNGFPSQVAWEDYKRSRSGRGASQYNWTDSDVR VAVAHAVGAAKSNASTHNVEERERAAATLNVAKELDFTASAAMLLERLPERTQERAAH ATGHVATAHGSLVHGCPSTASAATSPRRSRTRRRCSRL KRSKPLSRASSRGCPSTSL RHPGMRERHWTGLSQKLGMKLVPGDTLMLLEDCEPLLAHRDTIISYCEVAAKESQIEM TLKD RAKWETKCFIIEAYKETGTYILKDTSEWELLDEHLNWQQLQFSPFKGYFEE SITDWERSLNLISDILEQWLECQRAWRYLEPILNSEDIAMQLPRLSTLFEKVDRTWRR VMGNAHAQPNALEYCIGTNKLLDHLREANRLLEVLQHLMAQKVNVAAVGPTGTGKSIS LARLVLGGGMPANFLGLNFTFSAQTKCTVLQNSLMAKFDKRRSHVYGAPAGKHFLIFI DDANLPQPEKYGAQPPVELLRQMLAQGGFYNFTGGIKWSSIIDCSLALAMGPPGGGRS RVSNRF RYFNYLAFPEMSDMSKRTILQAILVGGLAQSGLADRLANVASAWDSTLRV FRKCTQVFLPTPAHVHYSFNMRDVMRVFPLLYTADKSVLQSEESIVRLMHEMQRVFY DRLVDATDKGLFIEYLNAELPSMGVDKSYNEWKADRLIFADVLSDKGVYEQITDMNA LTTRMNELLEAYNDENEVKMNLVLFLDAIEHVCRISRVLRLPNGHCLLLGVGGSGRKS LTRLACSLIAE EVFTIELSKNFGVKEWHESLAKLLLECGKDEKKRTFLFADTQLAHP TFLEDVAGLLTSGDVPNLFEDQDIELINDKFRGVCLSENLPTTKVSVYARFVKEARAN LHLVLAFSPIGEAFRSRLRMFPSLIACCTIDWFAEWPSEALLSVAAVQLNAGDVTDVM GAASHADLPGCFQAVHRAAAEVTERFFTETRRRSYVTPTSYLSLLΞNFKVMAAAKRRF VREQRGRLEKGLEKLRHTEVQVAELEAQLKAQQPVLVQKKAEIQSMMERLTVDRKEAA VKEADARREAQLPGGRAAYGGEDDE
Further analysis of the NOV85a protein yielded the following properties shown in Table 85B.
Table 85B. Protein Sequence Properties NOV85a
PSort 0.8800 probability located in nucleus; 0.3562 probability located in microbody analysis: (peroxisome); 0.1671 probability located in lysosome (lumen); 0.1000 probability located in mitochondrial matrix space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV85a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 85C.
In a BLAST search of public sequence databases, the NOV85a protein was found to have homology to the proteins shown in the BLASTP data in Table 85D.
PFam analysis predicts that the NOV85a protein contains the domains shown in the Table 85E.
Example 86.
The NOV86 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 86A.
Table 86A. NOV86 Sequence Analysis
SEQ ID NO: 241 1420 bp
NOV86a, GTCCAGCTTTAGCTCTCTGCTCGCCGCCGCCGCTGTCGCCGCCACCTCCTCTGATCTA
CGAAAGTCATGTTACCCAACACCGGGAGGCTGGCAGGATGTACAGTTTTTATCACAGG
CG59628-01 DNA Sequence TGCAAGCCGTGGCATTGGCAAAGCTATTGCATTGAAAGCAGCAAAGGATGGAGCAAAT ATTGTTATTGCTGCAAAGACCGCCCAGCCACATCCAAAACTTCTAGGCACAATCTATA CTGCTGCTGAAGAAATTGAAGCAGTTGGAGGAAAGGCCTTGCCATGTATTGTTGATGT GAGAGATGAACAGCAGATCAGTGCTGCAGTGGAGAAAGCCATCAAGAAATTTGGAGGA ATTGATATTCTGGTAAATAATGCCAGTGCCATTAGTTTGACCAATACATTGGACACAC CTACCAAGAGATTGGATCTGATGATGAACGTGAACACCAGAGGCACCTACCTTGCATC TAAAGCATGTATTCCTTATTTGAAAAAGAGCAAAGTTGCTCATATCCTCAATATCAGT CCACCACTGAACCTAAATCCAGTTTGGTTCAAACAGCACTGTGCTTATACCATTGCTA AGTATGGTATGTCTATGTATGTGCTTGGAATGGCAGAAGAATTTAAAGGTGAAATTGC AGTCAATGCATTATGGCCTAAAACAGCCATACACACTGCTGCTATGGATATGCTGGGA GGACCTGGTATCGAAAGCCAGTGTAGAAAAGTTGATATCATTGCAGATGCAGCATATT CCATTTTCCAAAAGCCAAAAAGTTTTACTGGCAACTTTGTCATTGATGAAAATATCTT AAAAGAAGAAGGAATAGAAAATTTTGACGTTTATGCAATTAAACCAGGTCATCCTTTG CAACCAGATTTCTTCTTAGATGAATACCCAGAAGCAGTTAGCAAGAAAGTGGAATCAA CTGGTGCTGTTCCAGAATTCAAAGAAGAGAAACTGCAGCTGCAACCAAAACCACGTTC TGGAGCTGTGGAAGAAACATTTAGAATTGTTAAGGACTCTCTCAGTGATGATGTTGTT AAAGCCACTCAAGCAATCTATCTGTTTGAACTCTCCGGTGAAGATGGTGGCACGTGGT TTCTTGATCTGAAAAGCAAGGGTGGGAATGTCGGATATGGAGAGCCTTCTGATCAGGC AGATGTGGTGATGAGTATGACTACTGATGACTTTGTAAAAATGTTTTCAGGTAAACTA AAACCAACAATGGCATTCATGTCAGGGAAATTGAAGATTAAAGGTAACATGGCCCTAG CAATCAAATTGGAGAAGCTAATGAATCAGATGAATGCCAGACTGTGAAGGAAAATATA
AAAAAAAAGTCGACTGCTATGCTCAAAAAGTAAAAAAAGCTCAACAGTTAAAATCTAA
TGTTTGTTTTCTTTCCTGTTATATTATA
ORF Start: ATG at 67 ORF Stop: TGA at 1321
SEQ ID NO: 242 418 aa MW at 45394.2kD
NOV86a, MLPNTGRLAGCTVFITGASRGIGKAIALKAAKDGANIVIAAKTAQPHPKLLGTIYTAA EEIEAVGGKALPCIVDVRDEQQISAAVEKAIKKFGGIDILVNNASAISLTNTLDTPTK
CG59628-01 Protein Sequence RLDLMMNVNTRGTYLASKACIPYLKKSKVAHILNISPPLNLNPV FKQHCAYTIAKYG MSMYVLGMAEEFKGEIAVNAL PKTAIHTAAMDMLGGPGIESQCRKVDIIADAAYSIF QKPKSFTGNFVIDENILKEEGIENFDVYAIKPGHPLQPDFFLDEYPEAVSKKVESTGA VPEFKEEKLQLQPKPRSGAVEETFRIVKDSLSDDWKATQAIYLFELSGEDGGT FLD LKSKGGNVGYGEPSDQADWMSMTTDDFVKMFSGKLKPTMAFMSGKLKIKGNMALAIK LEKLMNQMNARL
Further analysis of the NOV86a protein yielded the following properties shown in Table 86B.
Table 86B. Protein Sequence Properties NOV86a
PSort 0.5500 probability located in endoplasmic reticulum (membrane); 0.5000 analysis: probability located in microbody (peroxisome); 0.1900 probability located in lysosome (lumen); 0.1000 probability located in endoplasmic reticulum (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV86a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 86C.
In a BLAST search of public sequence databases, the NOV86a protein was found to have homology to the proteins shown in the BLASTP data in Table 86D.
PFam analysis predicts that the NOV86a protein contains the domains shown in the Table 86E.
Example 87.
The NOV87 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 87A.
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 87B.
Table 87B. Comparison of NOV87a against NOV87b.
NOV87a Residues/ Identities/
Protein Sequence Match Residues Similarities for the Matched Region
NOV87b 1..288 288/288 (100%) 1..288 288/288 (100%)
Further analysis of the NOV87a protein yielded the following properties shown in Table 87C.
Table 87C. Protein Sequence Properties NOV87a
PSort 0.4500 probability located in cytoplasm; 0.3000 probability located in microbody analysis: (peroxisome); 0.2110 probability located in lysosome (lumen); 0.1000 probability located in mitochondrial matrix space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV87a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 87D.
In a BLAST search of public sequence databases, the NOV87a protein was found to have homology to the proteins shown in the BLASTP data in Table 87E.
PFam analysis predicts that the NOV87a protein contains the domains shown in the Table 87F.
Example 88.
The NOV88 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 88 A.
GTGACTCAGATGTGGTGGTGCAGGAGCTCAAGTCCATGGTGGCCACCAAGATCGCCAA ATATGCTGTGCCTGATGAGATCCTGGTGGTGAAACGTCTTCCAAAAACCAGGTCTGGG AAGGTCATGCGGCGGCTCCTGAGGAAGATCATCACTAGTGAGGCCCAGGAGCTGGGAG ACACTACCACCTTGGAGGACCCCAGCATCATCGCAGAGATCCTGAGTGTCTACCAGAA GTGCAAGGACAAGCAGGCTGCTGCTAAGTGAGCTGGCACCTTGTGGGGCTCTTGGGAT
GGGCGGGCACCCAAGCCCTGGCTTGTCCTTCCCAGAAGGTACCCCTGAGGTTGGCGTC
TTCCTACGT
ORF Start: ATG at 50 ORF Stop: TGA at 2117
SEQ ID NO: 248 689 aa MW at 74855.9kD
NOV88a, MAARTLGRGVGRLLGSLRGLSGQPARPPCGVSAPRRAASGPSGSAPAVAAAAAQPGSY PALSAQAAREPAAF GPLARDTLVWDTPYHTVWDCDFSTGKIGWFLGGQLNVSVNCLD
CG59671-02 Protein Sequence QHVRKSPESVALI ERDEPGTEVRITYRELLETTCRLANTLKRHGVHRGDRVAIYMPV SPLAVAAMLACARIGAVHTVIFAGFSAESLAGRINDAKCKWITFNQGLRGGRWELK KIVDEAVKHCPTVQHVLVAHRTDNKVHMGDLDVPLEQEMAKEDPVCAPESMGSEDMLF MLYTSGSTGMPKGIVHTQAGYLLYAALTHKLVFDHQPGDIFGCVADIG ITGHSYWY GPLCNGATSVLFESTPVYPNAGRYWETVERLKINQFYGAPTAVRLLLKYGDAWVKKYD RSSLRTLGSVGEPINCEA E LHRWGDSRCTLVDT QTETGGICIAPRPSEEGAEI LPAMAMRPFFGIVPVLMDEKGSWEGSNVSGALCISQA PGMARTIYGDHQRFVDAYF KAYPGYYFTGDGAYRTEGGYYQITGRMDDVINISGHRLGTAEIEDAIADHPAVPESAV IGYPHDIKGEAAFAFIWKDSAGDSDWVQELKS VATKIAKYAVPDEILWKRLPKT RSGKVMRRLLRKIITSEAQELGDTTTLEDPSIIAEILSVYQKCKDKQAAAK
Further analysis of the NOV88a protein yielded the following properties shown in Table 88B.
Table 88B. Protein Sequence Properties NOV88a
PSort 0.6500 probability located in plasma membrane; 0.6000 probability located in analysis: nucleus; 0.4340 probability located in mitochondrial inner membrane; 0.3000 probability located in Golgi body
SignalP Likely cleavage site between residues 23 and 24 analysis:
A search of the NOV88a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 88C.
Table 88C. Geneseq Results for NOV88a
NOV88a Identities/
Geneseq Protein/Organism/Length [Patent Residues/ Similarities for Expect Identifier #, Date] Match the Matched Value Residues Region
In a BLAST search of public sequence databases, the NOV88a protein was found to have homology to the proteins shown in the BLASTP data in Table 88D.
PFam analysis predicts that the NOV88a protein contains the domains shown in the Table 88E.
Example 89.
The NOV89 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 89A.
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 89B.
Further analysis of the NOV89a protein yielded the following properties shown in Table 89C.
Table 89C. Protein Sequence Properties NOV89a
PSort 0.4500 probability located in cytoplasm; 0.3000 probability located in microbody analysis: (peroxisome); 0.1685 probability located in lysosome (lumen); 0.1000 probability located in mitochondrial matrix space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV89a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 89D.
In a BLAST search of public sequence databases, the NOV89a protein was found to have homology to the proteins shown in the BLASTP data in Table 89E.
PFam analysis predicts that the NOV89a protein contains the domains shown in the Table 89F.
142/239 (59%)
Ndr: domain 1 of 1 22-346 210/340 (62%) 3.7e-211 311/340 (91%)
Example 90.
The NOV90 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 90A.
Further analysis of the NOV90a protein yielded the following properties shown in Table 90B.
Table 90B. Protein Sequence Properties NOV90a
PSort 0.4500 probability located in cytoplasm; 0.1400 probability located in microbody analysis: (peroxisome); 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV90a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 90C.
Table 90C. Geneseq Results for NOV90a
Geneseq Protein/Organism/Length [Patent #, NOV90a Identities/ Expect Identifier Date] Value
In a BLAST search of public sequence databases, the NOV90a protein was found to have homology to the proteins shown in the BLASTP data in Table 90D.
PFam analysis predicts that the NOV90a protein contains the domains shown in the Table 90E.
Example 91.
The NOV91 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 91 A.
Further analysis of the NOV91a protein yielded the following properties shown in Table 9 IB.
Table 91B. Protein Sequence Properties NOV91a
PSort 0.6500 probability located in cytoplasm; 0.1000 probability located in analysis: mitochondrial matrix space; 0.1000 probability located in lysosome (lumen); 0.0000 probability located in endoplasmic reticulum (membrane)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV9 la protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 91C.
In a BLAST search of public sequence databases, the NOV91a protein was found to have homology to the proteins shown in the BLASTP data in Table 91D.
PFam analysis predicts that the NOV9 la protein contains the domains shown in the Table 9 IE.
Table 91E. Domain Analysis of NOV91a
Identities/
Pfam Domain NOV91a Match Region Similarities Expect Value for the Matched Region
No Significant Matches Found
Example 92.
The NOV92 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 92A.
ATCAGCGACTTGCGGACCGAGGACAGCGGCACCTACATTTGTGAGGTCACCAACACCT TCGGTTCGGCAGAGGCCACAGGCATCCTCATGGTCATTGATCCCCTTCATGTGACCCT GACACCAAAGAAGCTGAAGACCGGCATTGGCAGCACGGTCATCCTCTCCTGTGCCCTG ACGGGCTCCCCAGAGTTCACCATCCGCTGGTATCGCAACACGGAGCTGGTGCTGCCTG ACGAGGCCATCTCCATCCGCGGGCTCAGCAACGAGACGCTGCTCATCACCTCGGCCCA GAAGAGCCATTCCGGGGCCTACCAGTGCTTCGCTACCCGCAAGGCCCAGACCGCCCAG GACTTTGCCATCATTGCACTTGAGGATGGCACGCCCCGCATCGTCTCGTCCTTCAGCG AGAAGGTGGTCAACCCCGGGGAGCAGTTCTCACTGATGTGTGCGGCCAAGGGCGCCCC GCCCCCCACGGTCACCTGGGCCCTCGACGATGAGCCCATCGTGCGGGATGGCAGCCAC CGCACCAACCAGTACACCATGTCGGACGGCACCACCATCAGCCACATGAACGTCACAG GCCCCCAGATCCGCGACGGGGGCGTGTACCGGTGCACAGCGCGGAACTTGGTGGGCAG TGCTGAATATCAGGCGCGAATAAACGTAAGAGGCCCACCCAGCATCCGGGCTATGCGG AACATCACAGCAGTCGCCGGGCGGGACACCCTTATCAACTGCAGGGTCATCGGCTATC CCTACTACTCCATCAAGTGGTACAAGGATGCCTTGCTGCTGCCAGACAACCACCGCCA GGTGGTGTTTGAGAATGGGACCCTCAAGCTGACTGACGTGCAGAAGGGCATGGATGAG GGGGAGTACCTGTGCAGTGTCCTCATCCAGCCCCAGCTCTCCATCAGCCAGAGCGTTC ACGTAGCCGTCAAAGTGCCCCCTCTGATCCAGCCCTTCGAATTCCCACCCGCCTCCAT CGGCCAGCTGCTCTACATTCCCTGTGTGGTGTCCTCGGGGGACATGCCCATCCGTATC ACCTGGAGGAAGGACGGACAGGTGATCATCTCAGGCTCGGGCGTGACCATCGAGAGCA AGGAATTCATGAGCTCCCTGCAGATCTCTAGCGTCTCCCTCAAGCACAACGGCAACTA TACATGCATCGCCAGCAACGCAGCCGCCACCGTGAGCCGGGAGCGTCAGCTCATCGTG CGTGTGCCCCCTCGATTTGTGGTGCAACCCAACAACCAGGATGGCATCTACGGCAAAG CTGGTGTGCTCAACTGCTCGGTGGACGGCTACCCCCCACCCAAGGTCATGTGGAAGCA TGCCAAGGGGAGCGGGAACCCCCAGCAGTACCACCCTGTGCCCCTCACTGGCCGCATC CAGATCCTGCCCAACAGCTCGCTGCTGATCCGCCACGTCCTAGAAGAGGACATCGGCT ACTACCTCTGCCAGGCCAGCAACGGCGTAGGCACCGACATCAGCAAGTCCATGTTCCT CACAGTCAAGATCCCGGCCATGATCACTTCCCACCCCAACACCACCATCGCCATCAAG GGCCATGCGAAGGAGCTAAACTGCACGGCACGGGGTGAGCGGCCCATCATCATCCGCT GGGAGAAGGGGGACACAGTCATCGACCCTGACCGCGTCATGCGGTATGCCATCGCCAC CAAGGACAACGGCGACGAGGTCGTCTCCACACTGAAGCTCAAGCCCGCTGACCGTGGG GACTCTGTGTTCTTCAGCTGCCATGCCATCAACTCGTATGGGGAGGACCGGGGCTTGA TCCAACTCACTGTGCAAGAGCCCCCCGACCCCCCAGAGCTGGAGATCCGGGAGGTGAA GGCCCGGAGCATGAACCTGCGCTGGACCCAGCGATTCGACGGGAACAGCATCATCACG GGCTTCGACATTGAATACAAGAACAAATCAGATTCCTGGGACTTCAAGCAGTCCACAC GCAACATCTCCCCCACCATCAACCAGGCCAACATTGTGGACTTGCACCCGGCATCTGT GTACAGCATCCGCATGTACTCTTTCAACAAGATTGGCCGCAGTGAACCAAGCAAGGAG CTCACCATCAGCACTGAGGAGGCCGCTCCCGATGGGCCCCCCATGGATGTTACCTTGC AGCCAGTGACCTCACAGAGCATCCAGGTGACCTGGAAGGCACCCAAGAAGGAGCTGCA GAACGGTGTCATCCGGGGCTACCAGATTGGCTACAGAGAGAACAGCCCCGGCAGCAAC GGGCAGTACAGCATCGTGGAGATGAAGGCCACGGGGGACAGCGAGGTCTACACCCTGG ACAACCTCAAGAAGTTCGCCCAGTATGGGGTGGTGGTCCAAGCCTTCAATCGGGCTGG CACGGGGCCCTCTTCCAGCGAGATCAATGCCACCACTCTGGAGGATGTGCCCAGCCAG CCCCCTGAGAACGTCCGGGCCCTGTCCATCACTTCTGACGTGGCCGTCATCTCCTGGT CAGAGCCCCCGCGCAGCACCCTCAATGGCGTCCTCAAAGGCTATCGGGTCATCTTCTG GTCCCTCTATGTTGATGGGGAGTGGGGCGAGATGCAGAACATCACCACCACGCGGGAG CGGGTGGAGCTGCGGGGCATGGAGAAGTTCACCAACTACAGCGTCCAGGTGCTGGCCT ACACCCAGGCTGGGGACGGCGTACGCAGCAGTGTGCTCTACATCCAGACCAAGGAGGA CGTTCCAGGTCCCCCTGCTGGCATCAAAGCTGTCCCTTCATCAGCTAGCAGTGTGGTT GTGTCTTGGCTCCCCCCTACCAAGCCCAACGGGGTGATCCGCAAGTACACCATCTTCT GTTCCAGCCCCGGGTCTGGCCAGCCGGCTCCCAGCGAGTACGAGACGAGTCCAGAGCA GCTCTTCTACCGGATCGCCCACCTAAACCGCGGTCAGCAGTATCTGCTGTGGGTGGCC GCCGTCACCTCTGCCGGCCGGGGCAACAGCAGCGAGAAGGTGACCATCGAGCCTGCTG GCAAGGCCCCAGCAAAGATCATCTCCTTTGGGGGCACCGTGACAACACCTTGGATGAA AGATGTTCGGCTGCCTTGCAATTCAGTGGGAGATCCAGCCCCTGCTGTGAAGTGGACC AAGGACAGTGAAGACTCGGCCATTCCAGTGTCCATGGATGGGCACCGGCTCATCCACA CCAATGGCACACTGCTGCTGCGTGCAGTGAAGGCTGAGGACTCTGGCTACTACACGTG CACGGCCACCAACACTGGTGGCTTTGACACCATCATCGTCAACCTTCTGGTGCAAGTT CCCCCGGACCAGCCCCGCCTCACTGTCTCCAAAACCTCAGCTTCGTCCATCACCCTGA CCTGGATTCCAGGTGACAATGGGGGCAGCTCCATCCGAGGCTTCGTGCTACAGTACTC GGTGGACAACAGCGAGGAGTGGAAGGATGTGTTCATCAGCTCCAGCGAGCGCTCCTTC AAGCTGGACAGCCTCAAGTGTGGCACGTGGTACAAGGTGAAGCTGGCAGCCAAGAACA GCGTGGGCTCTGGGCGCATCAGCGAGATCATCGAGGCCAAGACCCACGGGCGGGAGCC CTCCTTCAGCAAAGACCAACACCTCTTCACCCACATCAACTCCACGCATGCTCGGCTT AACCTGCAGGGCTGGAACAATGGGGGCTGCCCTATCACAGCCATCGTTCTGGAGTACC GGCCCAAGGGGACCTGGGCCTGGCAGGGCCTCCGGGCCAACAGCTCCGGGGAGGTGTT TCTGACGGAACTGCGAGAGGCCACGTGGTACGAGCTGCGCATGAGGGCTTGCAACAGT GCGGGCTGCGGCAATGAAACAGCCCAGTTCGCCACCCTGGACTACGATGGCAGCACCA TTCCACCCATCAAGTCTGCTCAAGGTGAAGGGGATGATGTGAAGAAGCTGTTCACCAT CGGCTGCCCTGTCATCCTGGCCACACTGGGGGTGGCACTGCTCTTCATCGTACGCAAG AAGAGGAAGGAGAAACGGCTGAAGCGACTCCGAGATGCAAAGAGTTTGGCAGAAATGT TGATAAGCAAGAACAATAGAAGCTTTGACACCCCTGTGAAAGGGCCACCCCAGGGCCC ACGGCTACACATTGACATCCCCAGGGTCCAGCTGCTCATCGAGGACAAAGAAGGCATC AAGCAACTGGGAGATGACAAGGCCACCATCCCTGTGACAGATGCTGAGTTCAGCCAAG CTGTCAACCCACAGAGCTTCTGTACTGGCGTCTCCTTGCACCACCCAACCCTCATCCA GAGCACAGGACCCCTCATCGACATGTCTGACATCCGGCCAGGAACCAATCCAGTGTCC AGGAAGAATGTGAAGTCAGCCCACAGCACCCGGAACCGGTACTCAAGCCAGTGGACCC TGACCAAGTGCCAGGCCTCCACACCTGCCCGCACCCTCACCTCCGACTGGCGCACCGT
GGGCTCCCAGCATGGTGTCACGGTCACTGAGAGTGACAGCTACAGTGCCAGCCTGTCC CAGGACACAGACAAAGGAAGGAACAGCATGGTGTCCACTGAGAGTGCCTCTTCCACCT ACGAGGAGCTGGCCCGGGCCTATGAGCATGCCAAGCTGGAGGAGCAGCTGCAGCACGC CAAGTTTGAGATCACCGAGTGCTTCATCTCTGACAGTTCCTCTGACCAGATGACCACA GGCACCAACGAGAACGCCGACAGCATGACATCCATGAGCACACCCTCAGAGCCTGGCA TCTGCCGCTTTACCGCCTCACCACCCAAGCCCCAGGATGCGGACCGGGGCAAAAACGT GGCTGTGCCCATCCCTCACCGGGCCAACAAGAGTGACTACTGCAACCTGCCCCTGTAT GCCAAGTCAGAGGCCTTCTTTCGAAAGGCAGATGGACGTGAGCCCTGCCCCGTGGTCC CACCCCGTGAGGCCTCCATCCGGAACCTGGCTCGAACCTACCACACCCAGGCTCGCCA CCTGACCCTGGACCCTGCCAGCAAGTCCTTGGGCCTTCCCCACCCAGGGGCCCCCGCT GCCGCCTCCACAGCCACCTTACCTCAGAGGACTCTGGCCATGCCAGCCCCCCCAGCCG GCACAGCCCCCCCAGCCCCCGGCCCCACCCCTGCTGAGCCACCCACCGCCCCCAGCGC TGCCCCTCCGGCCCCCAGCACCGAGCCTCCACGAGCCGGGGGCCCACACACCAAAATG GGGGGCTCCAGGGACTCGCTTCTCGAGATGAGCACATCGGGGGTAGGGAGGTCTCAGA AGCAGGGGGCCGGGGCCTACTCCAAATCCTACACCCTGGTGTAGGGCCGGCAGGAAGA GCAGCCACGCCTGGGCCGCGCCGCGCCGCAGCCCCACACGCCAGCTCGGCTGTTTTTC
TGCATTATTTATATTCAACTGACAGACAAAAACCAACCAACGACAAAACAAAAACCCC
CAATCATGAACGCCTGTACATAGAACTCTTTTGTACAAATGAAACTATTTTCTTCTTC
TCCATGAAGCCAGGGCACAAAGAATTTGACAGTACAAGTCAAATCCCCCACCCCACAA
AATATGTGTGGAGATATATATACATATATAGACAGACAGGAACGCCTCCACGAGCTAT
ATATCTATATATTTCTCTCACCCTATTTTGAGACAGAGGCACAAAGACTCAGCAATTT
TTTTCCCTCCTCCTCACCTTCCCCCCAGTCTAGGTGGTTTTGACAAAGACCAAAATCC
CAACTCAGAGACACTGCATGCGATTTTACTGTTCCAAGAAAACCAGGAGTTGCTTCAA
TTTGCAGATGCTTATGTGTTAATACCTTTTTCTATGAAAAAAGACCCAGCGCCGTGTG
CAATAAAGGTTATGTTTCCAAAAAAAAGCTT
ORF Start: ATG at 129 ORF Stop: TAG at 5958
SEQ ED NO: 264 1943 aa MW at 211904.3kD
NOV92a, MPIRIT RKDGQVIISGSGVTIESKEFMSSLQISSVSLKHNGNYTCIASNAAATVSIV SPEHRFFITYHGGLYISDVQKEDALSTYRCITKHKYSGETRQSNGARLSVTDPAESIP
CG59754-02 Protein Sequence TILDGFHSQEV AGHTVELPCTASGYPIPAIR LKDGRPLPADSR TKRITGLTISDL RTEDSGTYICEVTNTFGSAEATGILMVIDPLHVTLTPKKLKTGIGSTVILSCALTGSP EFTIR YRNTELVLPDEAISIRGLSNETLLITSAQKSHSGAYQCFATRKAQTAQDFAI IALEDGTPRIVSSFSEKWNPGEQFSLMCAAKGAPPPTVT ALDDEPIVRDGSHRTNQ YTMSDGTTISHMNVTGPQIRDGGVYRCTARNLVGSAEYQARINVRGPPSIRAMRNITA VAGRDTLINCRVIGYPYYSIK YKDALLLPDNHRQWFENGTLKLTDVQKGMDEGEYL CSVLIQPQLSISQSVHVAVKVPPLIQPFEFPPASIGQLLYIPCWSSGDMPIRIT RK DGQVIISGSGVTIESKEFMSSLQISSVSLKHNGNYTCIASNAAATVSRERQLIVRVPP RFWQPNNQDGIYGKAGVLNCSVDGYPPPKVMWKHAKGSGNPQQYHPVPLTGRIQILP NSSLLIRHVLEEDIGYYLCQASNGVGTDISKSMFLTVKIPAMITSHPNTTIAIKGHAK ELNCTARGERPIIIR EKGDTVIDPDRVMRYAIATKDNGDEWSTLKLKPADRGDSVF FSCHAINSYGEDRGLIQLTVQEPPDPPELEIREVKARSMNLRWTQRFDGNSIITGFDI EYKNKSDS DFKQSTRNISPTINQANIVDLHPASVYSIRMYSFNKIGRSEPSKELTIS TEEAAPDGPPMDVTLQPVTSQSIQVTWKAPKKELQNGVIRGYQIGYRENSPGSNGQYS IVEMKATGDSEVYTLDNLKKFAQYGVWQAFNRAGTGPSSSEINATTLEDVPSQPPEN VRALSITSDVAVIS SEPPRSTLNGVLKGYRVIF SLYVDGE GEMQNITTTRERVEL RG EKFTNYSVQVLAYTQAGDGVRSSVLYIQTKEDVPGPPAGIKAVPSSASSVWSWL PPTKPNGVIRKYTIFCSSPGSGQPAPSEYETSPEQLFYRIAHLNRGQQYLLWVAAVTS AGRGNSSEKVTIEPAGKAPAKIISFGGTVTTPWMKDVRLPCNSVGDPAPAVK TKDSE DSAI VSMDGHRLIHTNGTLLLRAVKAEDSGYYTCTATNTGGFDTIIVNLLVQVPPDQ PRLTVSKTSASSITLTWIPGDNGGSSIRGFVLQYSVDNSEE KDVFISSSERSFKLDS LKCGT YKVKLAAKNSVGSGRISEIIEAKTHGREPSFSKDQHLFTHINSTHARLNLQG NNGGCPITAIVLEYRPKGT AWQGLRANSSGEVFLTELREAT YELRMRACNSAGCG NETAQFATLDYDGS IPPIKSAQGEGDDVKKLFTIGCPVILATLGVALLFIVRKKRKE KRLKRLRDAKSLAEMLISKNNRSFDTPVKGPPQGPRLHIDIPRVQLLIEDKEGIKQLG DDKATIPVTDAEFSQAVNPQSFCTGVSLHHPTLIQSTGPLIDMSDIRPGTNPVSRKNV KSAHSTRNRYSSQWTLTKCQASTPARTLTSD RTVGSQHGVTVTESDSYSASLSQDTD KGRNSMVSTESASSTYEELARAYEHAKLEEQLQHAKFEITECFISDSSSDQ TTGTNE NADSMTSMSTPSEPGICRFTASPPKPQDADRGKNVAVPIPHRANKSDYCNLPLYAKSE AFFRKADGREPCPWPPREASIRNLARTYHTQARHLTLDPASKSLGLPHPGAPAAAST ATLPQRTLAMPAPPAGTAPPAPGPTPAEPPTAPSAAPPAPSTEPPRAGGPHTKMGGSR DSLLEMSTSGVGRSQKQGAGAYSKSYTLV
SEQ ID NO: 265 6049 bp
NOV92b, CCACAGAGGGGAAATGCCAGCTTCCCTCTCCCTGGGGCTCCGTGCCCCCTCTGATCCA
GCCCTTCGAATTCCCACCCGCCTCCATCGGCCAGCTGCTCTACATTCCCTGTGTGGTG
CG59754-01 DNA Sequence TCCTCGGGGGACATGCCCATCCGTATCACCTGGAGGAAGGACGGACAGGTGATCATCT
CAGGCTCGGGCGTGACCATCGAGAGCAAGGAATTCATGAGCTCCCTGCAGATCTCTAG CGTCTCCCTCAAGCACAACGGCAACTATACATGCATCGCCAGCAACGCAGCCGCCACC GTGAGCATTGTGTCTCCAGAACACAGGTTTTTTATTACCTACCACGGCGGGCTGTACA TCTCTGACGTACAGAAGGAGGACGCCCTCTCCACCTATCGCTGCATCACCAAGCACAA GTATAGCGGGGAGACCCGGCAGAGCAATGGGGCACGCCTCTCTGTGACAGACCCTGCT GAGTCGATCCCCACCATCCTGGATGGCTTCCACTCCCAGGAAGTGTGGGCCGGCCACA CCGTGGAGCTGCCCTGCACCGCCTCGGGCTACCCTATCCCCGCCATCCGCTGGCTCAA GGATGGCCGGCCCCTCCCGGCTGACAGCCGCTGGACCAAGCGCATCACAGGGCTGACC
ATCAGCGACTTGCGGACCGAGGACAGCGGCACCTACATTTGTGAGGTCACCAACACCT TCGGTTCGGCAGAGGCCACAGGCATCCTCATGGTCATTGATCCCCTTCATGTGACCCT GACACCAAAGAAGCTGAAGACCGGCATTGGCAGCACGGTCATCCTCTCCTGTGCCCTG ACGGGCTCCCCAGAGTTCACCATCCGCTGGTATCGCAACACGGAGCTGGTGCTGCCTG ACGAGGCCATCTCCATCCGCGGGCTCAGCAACGAGACGCTGCTCATCACCTCGGCCCA GAAGAGCCATTCCGGGGCCTACCAGTGCTTCGCTACCCGCAAGGCCCAGACCGCCCAG GACTTTGCCATCATTGCACTTGAGGATGGCACGCCCCGCATCGTCTCGTCCTTCAGCG AGAAGGTGGTCAACCCCGGGGAGCAGTTCTCACTGATGTGTGCGGCCAAGGGCGCCCC GCCCCCCACAGTCACCTGGGCCCTCGACGATGAGCCCATCGTGCGGGATGGCAGCCAC CGCACCAACCAGTACACCATGTCGGACGGCACCACCATCAGCCACATGAACGTCACAG GCCCCCAGATCCGCGACGGGGGCGTGTACCGGTGCACAGCGCGGAACTTGGTGGGCAG TGCTGAATATCAGGCGCGAATAAACGTAAGAGGCCCACCCAGCATCCGGGCTATGCGG AACATCACAGCAGTCGCCGGGCGGGACACCCTTATCAACTGCAGGGTCATCGGCTATC CCTACTACTCCATCAAGTGGTACAAGGATGCCTTGCTGCTGCCAGACAACCACCGCCA GGTGGTGTTTGAGAATGGGACCCTCAAGCTGACTGACGTGCAGAAGGGCATGGATGAG GGGGAGTACCTGTGCAGTGTCCTCATCCAGCCCCAGCTCTCCATCAGCCAGAGCGTTC ACGTAGCCGTCAAAGTGCCCCCTCTGATCCAGCCCTTCGAATTCCCACCCGCCTCCAT CGGCCAGCTGCTCTACATTCCCTGTGTGGTGTCCTCGGGGGACATGCCCATCCGTATC ACCTGGAGGAAGGACGGACAGGTGATCATCTCAGGCTCGGGCGTGACCATCGAGAGCA AGGAATTCATGAGCTCCCTGCAGATCTCTAGCGTCTCCCTCAAGCACAACGGCAACTA TACATGCATCGCCAGCAACGCAGCCGCCACCGTGAGCCGGGAGCGTCAGCTCATCGTG CGTGTGCCCCCTCGATTTGTGGTGCAACCCAACAACCAGGATGGCATCTACGGCAAAG CTGGTGTGCTCAACTGCTCGGTGGACGGCTACCCCCCACCCAAGGTCATGTGGAAGCA TGCCAAGGGTAGCGGGAACCCCCAGCAGTACCACCCTGTGCCCCTCACTGGCCGCATC CAGATCCTGCCCAACAGCTCGCTGCTGATCCGCCACGTCCTAGAAGAGGACATCGGCT ACTACCTCTGCCAGGCCAGCAACGGCGTAGGCACCGACATCAGCAAGTCCATGTTCCT CACAGTCAAGATCCCCACCATCCTGGATGGCTTCCACTCCCAGGAAGTGTGGGCCGGC CACACCGTGGAGCTGCCCTGCACCGCCTCGGGCTACCCTATCCCCGCCATCCGCTGGC TCAAGGATGGCCGGCCCCTCCCGGCTGACAGCCGCTGGACCAAGCGCATCACAGGGCT GACCATCAGCGACTTGCGGACCGAGGACAGCGGCACCTACATTTGTGAGGTCACCAAC ACCTTCGGTGAGGCCACAGGCATCCTCATGGTCATTGGTGAGGAGCCCCCCGACCCCC CAGAGCTGGAGATCCGGGAGGTGAAGGCCCGGAGCATGAACCTGCGCTGGACCCAGCG ATTCGACGGGAACAGCATCATCACGGGCTTCGACATTGAATACAAGAACAAATCAGAT TCCTGGGACTTCAAGCAGTCCACACGCAACATCTCCCCCACCATCAACCAGGCCAACA TTGTGGACTTGCACCCGGCATCTGTGTACAGCATCCGCATGTACTCTTTCAACAAGAT TGGCCGCAGTGAACCAAGCAAGGAGCTCACCATCAGCACTGAGGAGGCCTCAGCTCCC GATGGGCCCCCCATGGATGTTACCTTGCAGCCAGTGACCTCACAGAGCATCCAGGTGA CCTGGAAGCAGGCACCCAAGAAGGAGCTGCAGAACGGTGTCATCCGGGGCTACCAGAT TGGCTACAGAGAGAACAGCCCCGGCAGCAACGGGCAGTACAGCATCGTGGAGATGAAG GCCACGGGGGACAGCGAGGTCTACACCCTGGACAACCTCAAGAAGTTCGCCCAGTATG GGGTGGTGGTCCAGGCCTTCAATCGGGCTGGCACGGGGCCCTCTTCCAGCGAGATCAA TGCCACCACTCTGGAGGATGTGCCCAGCCAGCCCCCTGAGAACGTCCGGGCCCTGTCC ATCACTTCTGACGTGGCCGTCATCTCCTGGTCAGAGCCCCCGCGCAGCACCCTCAATG GCGTCCTCAAAGGCTATCGGGTCATCTTCTGGTCCCTCTATGTTGATGGGGAGTGGGG CGAGATGCAGAACATCACCACCACGCGGGAGCGGGTGGAGCTGCGGGGCATGGAGAAG TTCACCAACTACAGCGTCCAGGTGCTGGCCTACACCCAGGCTGGGGACGGCGTACGCA GCAGTGTGCTCTACATCCAGACCAAGGAGGACGTTCCAGGTCCCCCTGCTGGCATCAA AGCTGTCCCTTCATCAGCTAGCAGTGTGGTTGTGTCTTGGCTCCCCCCTACCAAGCCC AACGGGGTGATCCGCAAGTACACCATCTTCTGTTCCAGCCCCGCCCCGCAGGCTCCCA GCGAGTACGAGACGAGTCCAGAGCAGCTCTTCTACCGGATCGCCCACCTAAACCGCGG TCAGCAGTATCTGCTGTGGGTGGCCGCCGTCACCTCTGCCGGCCGGGGCAACAGCAGC GAGAAGGTGACCATCGAGCCTGCTGGCAAGGCCCCAGCAAAGATCATCTCCTTTGGGG GCACCGTGACAACACCTTGGATGAAAGATGTTCGGCTGCCTTGCAATTCAGTGGGAGA TCCAGCCCCTGCTGTGAAGTGGACCAAGGACAGTGAAGACTCGGCCATTCCAGTGTCC ATGGATGGGCACCGGCTCATCCACACCAATGGCACACTGCTGCTGCGTGCAGTGAAGG CTGAGGACTCTGGCTACTACACGTGCACGGCCACCAACACTGGTGGCTTTGACACCAT CATCGTCAACCTTCTGGTGCAAGTTCCCCCGGACCAGCCCCGCCTCACTGTCTCCAAA ACCTCAGCTTCGTCCATCACCCTGACCTGGATTCCAGGTGACAATGGGGGCAGCTCCA TCCGAGGTTTTGTGCTACAGTACTCGGTGGACAACAGCGAGGAGTGGAAGGATGTGTT CATCAGCTCCAGCGAGCGCTCCTTCAAGCTGGACAGCCTCAAGTGTGGCACGTGGTAC AAGGTGAAGCTGGCAGCCAAGAACAGCGTGGGCTCTGGGCGCATCAGCGAGATCATCG AGGCCAAGACCCACGGGCGGGAGCCCTCCTTCAGCAAAGACCAACACCTCTTCACCCA CATCAACTCCACGCATGCTCGGCTTAACCTGCAGGGCTGGAACAATGGGGGCTGCCCT ATCACAGCCATCGTTCTGGAGTACCGGCCCAAGGGGACCTGGGCCTGGCAGGGCCTCC GGGCCAACAGCTCCGGGGAGGTGTTTCTGACGGAACTGCGAGAGGCCACGTGGTACGA GCTGCGCATGAGGGCTTGCAACAGTGCGGGCTGCGGCAATGAAACAGCCCAGTTCGCC ACCCTGGACTACGATGGCAGTACCATTCCACCCATCAAGTCTGCTCAAGGTGAAGGGG ATGATGTGAAGAAGCTGTTCACCATCGGCTGCCCTGTCATCCTGGCCACACTGGGGGT GGCACTGCTCTTCATCGTACGCAAGAAGAGGAAGGAGAAACGGCTGAAGCGACTCCGA GATGCAAAGAGTTTGGCAGAAATGTTGATAAGCAAGAACAATAGAAGCTTTGACACCC CTGTGAAAGGGCCACCCCAGGGCCCACGGCTACACATTGACATCCCCAGGGTCCAGCT GCTCATCGAGGACAAAGAAGGCATCAAGCAACTGGGTGAGGACAAGGCCACCATCCCT GTGACAGATGCTGAGTTCAGCCAAGCTGTCAACCCACAGAGCTTCTGTACTGGCGTCT CCTTGCACCACCCAACCCTCATCCAGAGCACAGGACCCCTCATCGACATGTCTGACAT CCGGCCAGGAACCGATCCAGTGTCCAGGAAGAATGTGAAGTCAGCCCACAGCACCCGG AACCGGTACTCAAGCCAGTGGACCCTGACCAAGTGCCAGGCCTCCACACCTGCCCGCA CCCTCACCTCCGACTGGCGCACCGTGGGCTCCCAGCATGGTGTCACGGTCACTGAGAG
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 92B.
Table 92B. Comparison of NOV92a against NOV92b.
NOV92a Residues/ Identities/
Protein Sequence Match Residues Similarities for the Matched Region
Further analysis of the NOV92a protein yielded the following properties shown in Table 92C.
Table 92C. Protein Sequence Properties NOV92a
PSort 0.7000 probability located in plasma membrane; 0.3000 probability located in analysis: microbody (peroxisome); 0.3000 probability located in nucleus; 0.2000 probability located in endoplasmic reticulum (membrane)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV92a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 92D.
In a BLAST search of public sequence databases, the NOV92a protein was found to have homology to the proteins shown in the BLASTP data in Table 92E.
PFam analysis predicts that the NOV92a protein contains the domains shown in the Table 92F.
Example 93.
The NOV93 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 93 A.
Further analysis of the NOV93a protein yielded the following properties shown in Table 93B.
Table 93B. Protein Sequence Properties NOV93a
PSort 0.6000 probability located in plasma membrane; 0.4000 probability located in analysis: Golgi body; 0.3000 probability located in endoplasmic reticulum (membrane); 0.1000 probability located in mitochondrial inner membrane
SignalP Likely cleavage site between residues 7 and 8 analysis:
A search of the NOV93a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 93 C.
Table 93C. Geneseq Results for NOV93a
NOV93a Identities/
Geneseq Protein/Organism/Length [Patent Expect Residues/ Similarities for Identifier #, Date] Value
In a BLAST search of public sequence databases, the NOV93a protein was found to have homology to the proteins shown in the BLASTP data in Table 93D.
PFam analysis predicts that the NOV93a protein contains the domains shown in the Table 93E.
Table 93E. Domain Analysis of NOV93a
Identities/
Pfam Domain NOV93a Match Region Similarities Expect Value for the Matched Region
No Significant Matches Found
Example 94.
The NOV94 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 94A.
Table 94A. NOV94 Sequence Analysis
SEQ TD NO: 269 2949 bp
NOV94a, GTCCGCCTCCGGGCCGCCGAGCCGCAGCCGCCGAGATGGGGGCCGCCCCGGGCCGCGC
CCCCGCCGGGTCCCGCCCGCCGCGCTGCCGCTGAGCGCATGGGCCCGGACCGCGCCGC
CG59761-01 DNA Sequence GCCGCTCCGGGAGCCGGGCCCGGGGTCCCGCCACCACCGCGCGCGGGACAGATTGATT CACTTTGGAGCTGTAAGTACTGATGTATTAGGGTGCAGCGCTCATTGTTCATTGACGC AGAGTCCCAAAATGAATATCCAAGAGCAGGGTTTCCCCTTGGACCTCGGAGCAAGTTT CACCGAAGATGCTCCCCGACCCCCAGTGCCTGGTGAGGAGGGAGAACTGGTGTCCACA GACCCGAGGCCCGCCAGCTACAGTTTCTGCTCCGGGAAAGGTGTTGGCATTAAAGGTG AGACTTCGACGGCCACTCCGAGGCGCTCGGATCTGGACCTGGGGTATGAGCCTGAGGG CAGTGCCTCCCCCACCCCACCATACTTGAAGTGGGCTGAGTCACTGCATTCCCTGCTG GATGACCAAGATGGGATAAGCCTGTTCAGGACTTTCCTGAAGCAGGAGGGCTGTGCCG ACTTGCTGGACTTCTGGTTTGCCTGCACTGGCTTCAGGAAGCTGGAGCCCTGTGACTC GAACGAGGAGAAGAGGCTGAAGCTGGCGAGAGCCATCTACCGAAAGTACATTCTTGAT AACAATGGCATCGTGTCCCGGCAGACCAAGCCAGCCACCAAGAGCTTCATAAAGGGCT GCATCATGAAGCAGCTGATCGATCCTGCCATGTTTGACCAGGCCCAGACCGAAATCCA GGCCACTATGGAGGAAAACACCTATCCCTCCTTCCTTAAGTCTGATATTTATTTGGAA TATACGAGGACAGGCTCGGAGAGCCCCAAAGTCTGTAGTGACCAGAGCTCTGGGTCAG GGACAGGGAAGGGCATATCTGGATACCTGCCGACCTTAAATGAAGATGAGGAATGGAA GTGTGACCAGGACATGGATGAGGACGATGGCAGAGACGCTGCTCCCCCCGGAAGACTC CCTCAGAAGCTGCTCCTGGAGACAGCTGCCCCGAGGGTCTCCTCCAGTAGACGGTACA GCGAAGGCAGAGAGTTCAGGTATGGATCCTGGCGGGAGCCAGTCAACCCCTATTATGT CAATGCCGGCTATGCCCTGGCCCCAGCCACCAGTGCCAACGACAGCGAGCAGCAGAGC CTGTCCAGCGATGCAGACACCCTGTCCCTCACGGACAGCAGCGTGGATGGGATCCCCC CATACAGGATCCGTAAGCAGCACCGCAGGGAGATGCAGGAGAGCGTGCAGGTCAATGG GCGGGTGCCCCTACCTCACATTCCCCGCACGTACCGGGTGCCGAAGGAGGTCCGCGTG GAGCCTCAGAAGTTCGCGGAGGAGCTCATCCACCGCCTGGAGGCTGTGCAGCGCACGC GGGAGGCCGAGGAGAAGCTGGAGGAGCGGCTGAAGCGCGTGCGCATGGAGGAGGAAGG TGAGGACGGCGATCCATCATCAGGGCCCCCAGGGCCGTGTCACAAGCTGCCTCCCGCC CCCGCTTGGCACCACTTCCCGCCCCGCCTGTGTTGGACATGGGCTTGTGCCGGGCTCC GGGATGCACACGAGGAGAACCCTGAGAGCATCCTGGACGAGCACGTACAGCGTGTGCT GAGGACACCTGGCCGCCAGTCGCCTGGGCCTGGCCATCGCTCCCCGGACAGTGGGCAC GTGGCCAAGATGCCAGTGGCACTGGGGGGTGCCGCCTCGGGGCACGGGAAGCACGTAC CCAAGTCAGGGGCGAAGCTGGACGCGGCCGGCCTGCACCACCACCGACACGTCCACCA CCACGTCCACCACAGCACAGCCCGGCCCAAGGAGCAGGTGGAGGCCGAGGCCACCCGC AGGGCCCAGAGCAGCTTCGCCTGGGGCCTGGAACCACACAGCCATGGGGCAAGGTCCC GAGGCTACTCAGAGAGTGTTGGCGCTGCCCCCAACGCCAGTGATGGCCTCGCCCACAG TGGGAAGGTGGGCGTGGCGTGCAAAAGAAATGCCAAGAAGGCCGAGTCGGGGAAGAGC GCCAGCACCGAGGTGCCAGGTGCCTCGGAGGATGCGGAGAAGAACCAGAAAATCATGC
AGTGGATCATTGAGGGGGAAAAGGAGATCAGCAGGCACCGCAGGACCGGCCACGGGTC TTCGGGGACGAGGAAGCCACAGCCCCATGAGAACTCCAGACCCTTGTCCCTTGAGCAC CCCTGGGCCGGCCCTCAGCTCCGGACCTCCGTGCAGCCCTCCCACCTCTTCATCCAAG ACCCCACCATGCCACCCCACCCAGCTCCCAACCCCCTAACCCAGCTGGAGGAGGCGCG CCGACGTCTGGAGGAGGAAGAAAAGAGAGCCAGCCGAGCACCCTCCAAGCAGAGGTAT GTGCAGGAGGTTATGCGGCGGGGACGCGCCTGCGTCAGGCCAGCGTGCGCGCCGGTGC TGCACGTGGTACCAGCCGTGTCGGACATGGAGCTCTCCGAGACAGAGACAAGATCGCA GAGGAAGGTGGGCGGCGGGAGTGCCCAGCCGTGTGACAGCATCGTTGTGGCGTACTAC TTCTGCGGGGAACCCATCCCCTACCGCACCCTGGTGAGGGGCCGCGCTGTCACCCTGG GCCAGTTCAAGGAGCTGCTGACCAAAAAGGGCAGCTACAGATACTACTTCAAGAAAGT GAGCGACGAGTTTGACTGTGGGGTGGTGTTTGAGGAGGTTCGAGAGGACGAGGCCGTC CTGCCCGTCTTTGAGGAGAAGATCATCGGCAAAGTGGAGAAGGTGGACTGATAGGCTG GTGGGCTGGCCGCTGTGCCAGGCGAGGCCCTTGGCGGGCACGGGTGTCACGGCCAGGC
AGATGACCTCGTACTCAGGAGCCCGATGGGGAACAGTGTTGGGTGTACC
ORF Start: ATG at 97 ORF Stop: TGA at 2833
SEQ ID NO: 270 912 aa MW at 101118.1kD
NOV94a, MGPDRAAPLREPGPGSRHHRARDRLIHFGAVSTDVLGCSAHCSLTQSPKMNIQEQGFP LDLGASFTEDAPRPPVPGEEGELVSTDPRPASYSFCSGKGVGIKGETSTATPRRSDLD
CG59761-01 Protein Sequence LGYEPEGSASPTPPYLKWAESLHSLLDDQDGISLFRTFLKQEGCADLLDF FACTGFR KLEPCDSNEEKRLKLARAIYRKYILDNNGIVSRQTKPATKSFIKGCIMKQLIDPAMFD QAQTEIQATMEENTYPSFLKSDIYLEYTRTGSESPKVCSDQSSGSGTGKGISGYLPTL NEDEEWKCDQDMDEDDGRDAAPPGRLPQKLLLETAAPRVSSSRRYSEGREFRYGS RE PVNPYYVNAGYALAPATSANDSEQQSLSSDADTLSLTDSSVDGIPPYRIRKQHRREMQ ESVQVNGRVPLPHIPRTYRVPKEVRVEPQKFAEELIHRLEAVQRTREAEEKLEERLKR VRMEEEGEDGDPSSGPPGPCHKLPPAPA HHFPPRLCWT ACAGLRDAHEENPESILD EHVQRVLRTPGRQSPGPGHRSPDSGHVAKMPVALGGAASGHGKHVPKSGAKLDAAGLH HHRHVHHHVHHSTARPKEQVEAEATRRAQSSFA GLEPHSHGARSRGYSESVGAAPNA SDGLAHSGKVGVACKRNAKKAESGKSASTEVPGASEDAEKNQKIMQ IIEGEKEISRH RRTGHGSSGTRKPQPHENSRPLSLEHP AGPQLRTSVQPSHLFIQDPTMPPHPAPNPL TQLEEARRRLEEEEKRASRAPSKQRYVQEVMRRGRACVRPACAPVLHWPAVSDMELS ETETRSQRKVGGGSAQPCDSIWAYYFCGEPIPYRTLVRGRAVTLGQFKELLTKKGSY RYYFKKVSDEFDCGWFEEVREDEAVLPVFEEKIIGKVEKVD
Further analysis of the NOV94a protein yielded the following properties shown in Table 94B.
Table 94B. Protein Sequence Properties NOV94a
PSort 0.6000 probability located in nucleus; 0.3000 probability located in microbody analysis: (peroxisome); 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV94a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 94C.
Table 94C. Geneseq Results for NOV94a
NOV94a Identities/
Geneseq Protein/Organism/Length [Patent Residues/ Similarities for Expect Identifier #, Date] Match the Matched Value
Residues Region
In a BLAST search of public sequence databases, the NOV94a protein was found to have homology to the proteins shown in the BLASTP data in Table 94D.
PFam analysis predicts that the NOV94a protein contains the domains shown in the Table 94E.
Table 94E. Domain Analysis of NOV94a
Pfam Domain NOV94a Match Region Expect Value
Example 95.
The NOV95 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 95 A.
Table 95A. NOV95 Sequence Analysis
SEQ TD NO: 271 2223 bp
NOV95a, TTGCAGGCATCACCCACGCCCTCTGCACCCACGCTGGAGGACGGGGAGGTTGTCAGGG
GCTATGATGAGATGAGTGGGGGCCGCTTCGACTTTGATGATGGAGGGGCGTACTGCGG
CG59756-01 DNA Sequence GGGCTGGGAGGGGGGAAAGGCCCATGGGCATGGACTGTGCACAGGCCCCAAGGGCCAG GGCGAATACTCTGGCTCCTGGAACTTTGGCTTTGAGGTGGCAGGTGTCTACACCTGGC CCAGCGGAAACACCTTTGAGGGATACTGGAGCCAGGGCAAACGGCATGGGCTGGGCAT AGAGACCAAGGGGCGCTGGCTCTACAAGGGCGAGTGGACACATGGCTTCAAGGGACGC TACGGAATCCGGCAGAGCTCAAGCAGCGGTGCCAAGTATGAGGGCACCTGGAACAATG GCCTGCAAGACGGCTATGGCACCGAGACCTATGCTGATGGAGGGACGTACCAAGGCCA GTTCACCAACGGCATGCGCCATGGCTACGGAGTACGCCAGAGCGTGCCCTACGGGATG GCCGTGGTGGTGCGCTCGCCGCTGCGCACGTCGCTGTCGTCCCTGCGCAGCGAGCACA GCAACGGCACGGTGGCCCCGGACTCTCCCGCCTCGCCGGCCTCCGACGGCCCCGCGCT GCCCTCGCCCGCCATCCCGCGTGGCGGCTTCGCGCTCAGCCTCCTGGCCAATGCCGAG GCGGCCGCGCGGGCGCCCAAGGGCGGCGGCCTCTTCCAGCGGGGCGCGCTGCTGGGCA AGCTGCGGCGCGCAGAGTCGCGCACGTCCGTGGGTAGCCAGCGCAGCCGTGTCAGCTT CCTTAAGAGCGACCTCAGCTCGGGCGCCAGCGACGCCGCGTCCACCGCCAGCCTGGGA GAGGCCGCCGAGGGCGCCGACGAGGCCGCACCCTTCGAGGCCGATATCGACGCCACCA CCACCGAGACCTACATGGGCGAGTGGAAGAACGACAAACGCTCGGGCTTCGGCGTGAG CGAACGCTCCAGTGGCCTCCGCTACGAGGGCGAGTGGCTGGACAACCTGCGCCACGGC TATGGCTGCACCACGCTGCCCGACGGCCACCGCGAGGAGGGCAAGTACCGCCACAACG TGCTGGTCAAGGACACCAAGCGCCGCATGCTGCAGCTCAAGAGCAACAAGGTCCGCCA GAAAGTGGAGCACAGTGTGGAGGGTGCCCAGCGCGCCGCTGCTATCGCGCGCCAGAAG GCCGAGATTGCCGCCTCCAGGACAAGCCACGCCAAGGCCAAAGCTGAGGCAGCGGAAC AGGCCGCCCTGGCTGCCAACCAGGAGTCCAACATTGCTCGCACTTTGGCCAGGGAGCT GGCTCCGGACTTCTACCAGCCAGGTCCGGAATATCAGAAGCGCCGGCTGCTGCAGGAG ATCCTGGAGAACTCGGAGAGCCTGCTGGAGCCCCCCGACCGGGGCGCCGGCGCAGCGG GCCTCCCACAGCCGCCCCGCGAGAGCCCGCAGCTGCACGAGCGTGAGACCCCTCGGCC CGAGGGTGGCTCCCCGTCACCGGCCGGGACGCCCCCGCAGCCCAAGCGGCCCAGGCCC GGGGTGTCCAAGGACGGCCTGCTGAGCCCAGGCGCCTGGAACGGCGAGCCCAGCGGTG AGGGCAGCCGGTCAGTCACTCCGTCCGAGGGCGCGGGCCGCCGCAGCCCCGCGCGTCC AGCCACCGAGCGCATGGCCATCGAGGCTCTGCAGGCACCGCCTGCGCCGTCGCGGGAG CCGGAGGTGGCGCTTTACCAGGGCTACCACAGCTATGCTGTGCGCACCACGCCGCCCG AGCCCCCACCCTTTGAGGACCAGCCCGAGCCCGAGGTCTCCGGGTCCGAGTCCGCGCC CTCGTCCCCGGCCACCGCCCCGCTGCAGGCCCCCACGCTCCGAGGCCCCGAGCCTGCA CGCGAGACCCCCGCCAAGCTGGAGCCCAAGCCCATCATCCCCAAAGCCGAGCCCAGGG CCAAGGCCCGCAAGACTGAGGCTCGAGGGCTGACCAAGGCGGGGGCCAAGAAGAAGGC GCGGAAGGAGGCCGCACTGGCGGCAGAGGCGGAGGTGGAGGTGGAAGAGGTCCCCAAC ACCATCCTCATCTGCATGGTGATCCTGCTGAACATCGGCCTGGCCATCCTCTTTGTTC ACCTCCTGACCTGACCGTCGCTTACCAGGTGCAGCCAGCTGGCTGGAGGAGGGGTTGG GGGGCAGGAGCCCCTGGGG
ORF Start: ATG at 70 ORF Stop: TGA at 2158
SEQ ED NO: 272 |696 aa |MW at 74220.7kD
NOV95a, MSGGRFDFDDGGAYCGG EGGKAHGHGLCTGPKGQGEYSGS NFGFEVAGVYT PSGN TFEGYWSQGKRHGLGIETKGR LYKGE THGFKGRYGIRQSSSSGAKYEGT NNGLQD
CG59756-01 Protein Sequence GYGTETYADGGTYQGQFTNG RHGYGVRQSVPYGMAVWRSPLRTSLSSLRSEHSNGT VAPDSPASPASDGPALPSPAIPRGGFALSLLANAEAAARAPKGGGLFQRGALLGKLRR AESRTSVGSQRSRVSFLKSDLSSGASDAASTASLGEAAEGADEAAPFEADIDATTTET YMGEWKNDKRSGFGVSERSSGLRYEGEWLDNLRHGYGCTTLPDGHREEGKYRHNVLVK DTKRRMLQLKSNKVRQKVEHSVEGAQRAAAIARQKAEIAASRTSHAKAKAEAAEQAAL AANQESNIARTLARELAPDFYQPGPEYQKRRLLQEILENSESLLEPPDRGAGAAGLPQ PPRESPQLHERETPRPEGGSPSPAGTPPQPKRPRPGVSKDGLLSPGANGEPSGEGSR SVTPSEGAGRRSPARPATERMAIEALQAPPAPSREPEVALYQGYHSYAVRTTPPEPPP FEDQPEPEVSGSESAPSSPATAPLQAPTLRGPEPARETPAKLEPKPIIPKAEPRAKAR KTEARGLTKAGAKKKARKEAALAAEAEVEVEEVPNTILICMVILLNIGLAILFVHLLT
Further analysis of the NOV95a protein yielded the following properties shown in Table 95B.
Table 95B. Protein Sequence Properties NOV95a
PSort 0.8000 probability located in nucleus; 0.7000 probability located in plasma analysis: membrane; 0.3133 probability located in microbody (peroxisome); 0.2000 probability located in endoplasmic reticulum (membrane)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV95a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 95C.
In a BLAST search of public sequence databases, the NOV95a protein was found to have homology to the proteins shown in the BLASTP data in Table 95D.
PFam analysis predicts that the NOV95a protein contains the domains shown in the Table 95E.
Example 96.
The NOV96 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 96A.
TTCTCGGTCTTCCATGGAAATGCCTTCACAGCCAGCTCCACGAACAGTCACAGATGAG GAGATAAATTTTGTTAAGACCTGTCTTCAGAGATGGAGGAGTGAGATTGAACAAGATA TACAAGATTTAAAGACTTGTATTGCAAGTACTACTCAGACTATTGAACAGATGTACTG CGATCCTCTCCTTCGTCAGGTGCCTTATCGCTTGCATGCAGTTCTTGTTCATGAAGGA CAAGCAAATGCTGGACACTATTGGGCCTATATCTATAATCAACCCCGACAGAGCTGGC TCAAGTACAATGACATCTCTGTTACTGAATCTTCCTGGGAAGAAGTTGAAAGAGATTC CTATGGAGGCCTGAGAAATGTTAGTGCTTACTGTCTGATGTACATTAATGCCAAACTA CCCTACTTCAATGCAGAGGCAGCCCCAACTGAATCAGATCAAATGTCAGAAGTGGAAG CCCTATCTGTGGAACTCAAGCATTACATTCAGGAGGATAACTGGCGGTTTGAGCAGGA AGTAGAGGAGTGGGAAGAAGAGCAGTCTTGCAAAATCCCTCAAATGGAGTCCTCCCCC AACTCCTCATCACAGGGCTACTCTACATCACAAGAGCCTTCAGTAGCCTCTTCTCATG GGGTTCGCTGCTTGTCATCTGAGCATGCTGTGATTGTAAAGGAGCAAACTGCCCAGGC TATTGCAAACACAGCCCGTGCCTATGAGAAGAGCGGTGTAGAAGCGGCACTGAGTGAG GCATTCCATGAAGAATACTCCAGGCTCTATCAGCTTGCCAAAGAGACCCCCACCTCTC ACAGTGATCCTCGACTTCAGCATGTCCTTGTCTACTTTTTCCAAAATGAAGCACCCAA AAGGGTAGTAGAACGAACCCTTCTGGAACAGTTTGCAGATAAAAATCTTAGCTATGAT GAAAGATCAATCAGCATTATGAAGGTGGCTCAAGCGAAACTGAAGGAAATTGGTCCAG ATGACATGAATATGGAAGAGTACAAGAGGTGGCATGAAGATTATAGTTTGTTCCGAAA AGTGTCTGTGTATCTCCTAACAGGCCTAGAACTCTATCAAAAAGGAAAGTACCAAGAG GCACTTTCCTACCTGGTATATGCCTACCAGAGCAATGCTGCCCTGCTGATGAAGGGGC CCCGCCGGGGGGTCAAAGAATCCGTGATTGCTTTATACCGAAGAAAATGCCTTCTGGA GCTGAATGCCAAAGCAGCTTCTCTTTTTGAAACAAATGATGATCACTCCGTAACTGAG GGCATTAATGTGATGAATGAACTGATCATCCCCTGCATTCACCTTATCATTAATAATG ACATTTCCAAGGATGATCTGGATGCCATTGAGGTCATGAGAAACCATTGGTGCTCTTA CCTTGGGCAAGATATTGCAGAAAATCTGCAGCTGTGCCTAGGGGAGTTTCTACCCAGA CTTCTAGATCCTTCTGCAGAAATCATCGTCTTGAAAGAGCCTCCAACTATTCGACCCA ATTCTCCCTATGACCTATGTAGCCGATTTGCAGCTGTCATGGAGTCAATTCAGGGAGT TTCAACTGTGACAGTGAAATAAGCTCCCACATGTTCAAGGCCCATTCTGGTTCCTGGC TGCCTGCCTCTTGCACAGAAGTTCGTTGTCATAGTGCTCACCTTGGGAAAAGGATTAG
GTGGGCACA
ORF Start: ATG at 17 ORF Stop: TAA at 3152
SEQ ED NO: 274 1045 aa MW at 119041.7kD
NOV96a, MTAELQQDDAAGAADGHGSSCQMLLNQLREITGIQDPSFLHEALKASNGDITQAVSLL TDERVKEPSQDTVATEPSEVEGSAANKEVLAKVIDLTHDNKDDLQAAIALSLLESPKI
CG59708-01 Protein Sequence QADGRDLNRMHEATSAETKRSKRKRCEVWGENPNPNDWRRVDG PVGLKNVGNTC FS AVIQSLFQLPEFRRLVLSYSLPQNVLENCRSHTEKRNIMFMQELQYLFALMMGSNRKF VDPSAALDLLKGAFRSSEEQQQDVSEFTHKLLD LEDAFQLAVNVNSPRNKSENPMVQ LFYGTFLTEGVREGKPFCNNETFGQYPLQVNGYRNLDECLEGAMVEGDVELLPSDHSV KYGQERWFTKLPPVLTFELSRFEFNQSLGQPEKIHNKLEFPQIIYMDRYMYRSKELIR NKRECIRKLKEEIKILQQKLERYVKYGSGPARFPLPDMLKYVIEFASTKPASESCPPE SDTHMTLPLSSVHCSVSDQTSKESTSTESSSQDVESTFSSPEDSLPKSKPLTSSRSSM EMPSQPAPRTVTDEEINFVKTCLQRWRSEIEQDIQDLKTCIASTTQTIEQMYCDPLLR QVPYRLHAVLVHEGQANAGHY AYIYNQPRQSWLKYNDISVTESSWEEVERDSYGGLR NVSAYCLMYINAKLPYFNAEAAPTESDQMSEVEALSVELKHYIQEDNWRFEQEVEE E EEQSCKIPQMESSPNSSSQGYSTSQEPSVASSHGVRCLSSEHAVIVKEQTAQAIANTA RAYEKSGVEAALSEAFHEEYSRLYQLAKETPTSHSDPRLQHVLVYFFQNEAPKRWER TLLEQFADKNLSYDERSISIMKVAQAKLKEIGPDDMNMEEYKRWHEDYSLFRKVSVYL LTGLELYQKGKYQEALSYLVYAYQSNAALLMKGPRRGVKESVIALYRRKCLLELNAKA ASLFETNDDHSVTEGINVMNELIIPCIHLIINNDISKDDLDAIEVMRNH CSYLGQDI AENLQLCLGEFLPRLLDPSAEIIVLKEPPTIRPNSPYDLCSRFAAVMESIQGVSTVTV K
SEQ ED NO: 275 3044 bp
NOV96b, CGTAGGCGCTTCGGCCATGACTGCGGAGCTGCAGCAGGACGACGCGGCCGGCGCGGCA
GACGGCCACGGCTCGAGCTGCCAAATGCTGTTAAATCAACTGAGAGAAATCACAGGCA
CG59708-02 DNA Sequence TTCAGGACCCTTCCTTTCTCCATGAAGCTCTGAAGGCCAGTAATGGTGACATTACTCA GGCAGTCAGCCTTCTCACTGATGAGAGAGTTAAGGAGCCCAGTCAAGACACTGTTGCT ACAGAACCATCTGAAGTAGAGGGGAGTGCTGCCAACAAGGAAGTATTAGCAAAAGTTA TAGACCTTACTCATGATAACAAAGATGATCTTCAGGCTGCCATTGCTTTGAGTCTACT GGAGTCTCCCAAAATTCAAGCTGATGGAAGAGATCTTAACAGGATGCATGAAGCAACC TCTGCAGAAACTAAACGCTCAAAGAGAAATATCATGTTTATGCAAGAGCTTCAGTATT TGTTTGCTCTAATGATGGGATCAAATAGAAAATTTGTAGACCCGTCTGCAGCCCTGGA TCTATTAAAGGGAGCATTCCGATCATCTGAGGAACAGCAGCAAGATGTGAGTGAATTC ACACACAAGCTCCTGGATTGGCTAGAGGACGCATTCCAGCTAGCTGTTAATGTTAACA GTCCCAGGAACAAATCTGAAAATCCAATGGTGCAGCTGTTCTATGGTACTTTCCTGAC TGAAGGGGTTCGTGAAGGAAAACCCTTTTGTAACAATGAGACCTTCGGCCAGTATCCT CTTCAGGTAAACGGTTATCGCAACTTAGACGAGTGTTTGGAAGGGGCCATGGTGGAGG GTGATGTTGAGCTTCTTCCCTCCGATCACTCGGTGAAGTATGGACAAGAGCGTTGGTT TACAAAGCTACCTCCAGTGTTGACCTTTGAACTCTCAAGATTTGAGTTTAATCAGTCC CTTGGGCAGCCAGAGAAAATTCACAATAAGCTGGAATTTCCTCAGATTATTTATATGG ACAGGTACATGTACAGGAGCAAGGAGCTTATTCGAAATAAGAGAGAGTGTATTCGAAA GTTGAAGGAGGAAATAAAAATTCTGCAGCAAAAATTGGAAAGGTATGTGAAATATGGC TCAGGCCCAGCTCGGTTCCCGCTCCCGGACATGCTGAAATATGTTATTGAATTTGCTA GTACAAAACCTGCCTCAGAAAGCTGTCCACCTGAAAGTGACACACATATGACATTACC
ACTTTCTTCAGTGCACTGCTCGGTTTCTGACCAGACATCCAAGGAAAGTACAAGTACA GAAAGCTCTTCTCAGGATGTTGAAAGTACCTTTTCTTCTCCTGAAGATTCTTTACCCA AGTCTAAACCACTGACATCTTCTCGGTCTTCCATGGAAATGCCTTCACAGCCAGCTCC ACGAACAGTCACAGATGAGGAGATAAATTTTGTTAAGACCTGTCTTCAGAGATGGAGG AGTGAGATTGAACAAGATATACAAGATTTAAAGACTTGTATTGCAAGTACTACTCAGA CTATTGAACAGATGTACTGCGATCCTCTCCTTCGTCAGGTGCCTTATCGCTTGCATGC AGTTCTTGTTCATGAAGGACAAGCAAATGCTGGACACTATTGGGCCTATATCTATAAT CAACCCCGACAGAGCTGGCTCAAGTACAATGACATCTCTGTTACTGAATCTTCCTGGG AAGAAGTTGAAAGAGATTCCTATGGAGGCCTGAGAAATGTTAGTGCTTACTGTCTGAT GTACATTAATGCCAAACTACCCTACTTCAATGCAGAGGCAGCCCCAACTGAATCAGAT CAAATGTCAGAAGTGGAAGCCCTATCTGTGGAACTCAAGCATTACATTCAGGAGGATA ACTGGCGGTTTGAGCAGGAAGTAGAGGAGTGGGAAGAAGAGCAGTCTTGCAAAATCCC TCAAATGGAGTCCTCCCCCAACTCCTCATCACAGGGCTACTCTACATCACAAGAGCCT TCAGTAGCCTCTTCTCATGGGGTTCGCTGCTTGTCATCTGAGCATGCTGTGATTGTAA AGGAGCAAACTGCCCAGGCTATTGCAAACACAGCCCGTGCCTATGAGAAGAGCGGTGT AGAAGCGGCACTGAGTGAGGCATTCCATGAAGAATACTCCAGGCTCTATCAGCTTGCC AAAGAGACCCCCACCTCTCACAGTGATCCTCGACTTCAGCATGTCCTTGTCTACTTTT TCCAAAATGAAGCACCCAAAAGGGTAGTAGAACGAACCCTTCTGGAACAGTTTGCAGA TAAAAATCTTAGCTATGATGAAAGATCAATCAGCATTATGAAGGTGGCTCAAGCGAAA CTGAAGGAAATTGGTCCAGATGACATGAATATGGAAGAGTACAAGAGGTGGCATGAAG ATTATAGTTTGTTCCGAAAAGTGTCTGTGTATCTCCTAACAGGCCTAGAACTCTATCA AAAAGGAAAGTACCAAGAGGCACTTTCCTACCTGGTATATGCCTACCAGAGCAATGCT GCCCTGCTGATGAAGGGGCCCCGCCGGGGGGTCAAAGAATCCGTGATTGCTTTATACC GAAGAAAATGCCTTCTGGAGCTGAATGCCAAAGCAGCTTCTCTTTTTGAAACAAATGA TGATCACTCCGTAACTGAGGGCATTAATGTGATGAATGAACTGATCATCCCCTGCATT CACCTTATCATTAATAATGACATTTCCAAGGATGATCTGGATGCCATTGAGGTCATGA GAAACCATTGGTGCTCTTACCTTGGGCAAGATATTGCAGAAAATCTGCAGCTGTGCCT AGGGGAGTTTCTACCCAGACTTCTAGATCCTTCTGCAGAAATCATCGTCTTGAAAGAG CCTCCAACTATTCGACCCAATTCTCCCTATGACCTATGTAGCCGATTTGCAGCTGTCA TGGAGTCAATTCAGGGAGTTTCAACTGTGACAGTGAAATAAGCTCCCACATGTTCAAG GCCCATTCTGGTTCCTGGCTGCCTGCCTCTTGCACAGAAGTTCGTTGTCATAGTGCTC
ACCTTGGGAAAAGGATTAGGTGGGCACA
ORF Start: ATG at 17 ORF Stop: TAA at 2939
SEQ ED NO: 276 974 aa MW at l l0687.3kD
NOV96b, MTAELQQDDAAGAADGHGSSCQMLLNQLREITGIQDPSFLHEALKASNGDITQAVSLL TDERVKEPSQDTVATEPSEVEGSAANKEVLAKVIDLTHDNKDDLQAAIALSLLESPKI
CG59708-02 Protein Sequence QADGRDLNRMHEATSAETKRSKRNIMFMQELQYLFALMMGSNRKFVDPSAALDLLKGA FRSSEEQQQDVSEFTHKLLD LEDAFQLAVNVNSPRNKSENPMVQLFYGTFLTEGVRE GKPFCNNETFGQYPLQVNGYRNLDECLEGAMVEGDVELLPSDHSVKYGQER FTKLPP VLTFELSRFEFNQSLGQPEKIHNKLEFPQIIYMDRYMYRSKELIRNKRECIRKLKEEI KILQQKLERYVKYGSGPARFPLPD LKYVIEFASTKPASESCPPESDTHMTLPLSSVH CSVSDQTSKESTSTESSSQDVESTFSSPEDSLPKSKPLTSSRSS E PSQPAPRTVTD EEINFVKTCLQR RSEIEQDIQDLKTCIASTTQTIEQMYCDPLLRQVPYRLHAVLVHE GQANAGHY AYIYNQPRQS LKYNDISVTESSWEEVERDSYGGLRNVSAYCLMYINAK LPYFNAEAAPTESDQMSEVEALSVELKHYIQEDNWRFEQEVEE EEEQSCKIPQMESS PNSSSQGYSTSQEPSVASSHGVRCLSSEHAVIVKEQTAQAIANTARAYEKSGVEAALS EAFHEEYSRLYQLAKETPTSHSDPRLQHVLVYFFQNEAPKRWERTLLEQFADKNLSY DERSISIMKVAQAKLKEIGPDDMNMEEYKRWHEDYSLFRKVSVYLLTGLELYQKGKYQ EALSYLVYAYQSNAALLMKGPRRGVKESVIALYRRKCLLELNAKAASLFETNDDHSVT EGINVMNELIIPCIHLIINNDISKDDLDAIEVMRNHWCSYLGQDIAENLQLCLGEFLP RLLDPSAEIIVLKEPPTIRPNSPYDLCSRFAAVMESIQGVSTVTVK
SEQ ED NO: 277 3231 bp
NOV96c, GCGCTTCGGCCATGACTGCGGAGCTGCAGCAGGACGACGCGGCCGGCGCGGCAGACGG
CCACGGCTCGAGCTGCCAAATGCTGTTAAATCAACTGAGAGAAATCACAGGCATTCAG
CG59708-03 DNA Sequence GACCCTTCCTTTCTCCATGAAGCTCTGAGGGCCAGTAATGGTGACATTACTCAGGCAG TCAGCCTTCTCACTGATGAGAGAGTTAAGGAGCCCAGTCAAGACACTGTTGCTACAGA ACCATCTGAAGTAGAGGGGAGTGCTGCCAACAAGGAAGTATTAGCAAAAGTTATAGAC CTTACTCATGATAACAAAGATGATCTTCAGGCTGCCATTGCTTTGAGTCTACTGGAGT CTCCCAAAATTCAAGCTGATGGAAGAGATCTTAACAGGATGCATGAAGCAACCTCTGC AGAAACTAAACGCTCAAAGAGAAAACGCTGTGAAGTCTGGGGAGAAAACCCCAATCCC AATGACTGGAGGAGAGTTGATGGTTGGCCAGTTGGGCTGAAAAATGTTGGCAATACAT GTTGGTTTAGTGCTGTTATTCAGTCTCTCTTTCAATTGCCTGAATTTCGAAGACTTGT TCTCAGTTATAGTCTGCCACAAAATGTACTTGAAAATTGTCGAAGTCATACAGAAAAG AGAAATATCATGTTTATGCAAGAGCTTCAGTATTTGTTTGCTCTAATGATGGGATCAA ATAGAAAATTTGTAGACCCGTCTGCAGCCCTGGATCTATTAAAGGGAGCATTCCGATC ATCTGAGGAACAGCAGCAAGATGTGAGTGAATTCACACACAAGCTCCTGGATTGGCTA GAGGACGCATTCCAGCTAGCTGTTAATGTTAACAGTCCCAGGAACAAATTTGAAAATC CAATGGTGCAGCTGTTCTATGGTACTTTCCTGACTGAAGGGGTTCGTGAAGGAAAACC CTTTTGTAACAATGAGACCTTCGGCCAGTATCCTCTTCAGGTAAACGGTTATCGCAAC TTAGACGAGTGTTTGGAAGGGGCCATGGTGGAGGGTGATGTTGAGCTTCTTCCCTCCG ATCACTCGGTGAAGTATGGACAAGAGCGTTGGTTTACAAAGCTACCTCCAGTGTTGAC CTTTGAACTCTCAAGATTTGAGTTTAATCAGTCCCTTGGGCAGCCAGAGAAAATTCAC AATAAGCTGGAATTTCCTCAGATTATTTATATGGACAGGTACATGTACAGGAGCAAGG
AGCTTATTCGAAATAAGAGAGAGTGTATTCGAAAGTTGAAGGAGGAAATAAAAATTCT GCAGCAAAAATTGGAAGGGTATGTGAAATATGGCTCAGGCCCAGCTCGGTTCCCGCTC CCGGACATGCTGAAATATGTTATTGAATTTGCTAGTACAAAACCTGCCTCAGAAAGCT GTCCACCTGAAAGTGACACACATATGACATTACCACTTTCTTCAGTGCACTGCTCGGT TTCTAACCAGACATCCAAGGAAAGTACAAGTACAGAAAGCTCTTCTCAGGATGTTGAA AGTACCTTTTCTTCTCCTGAAGATTCTTTACCCAAGTCTAAACCACTGACATCTTCTC GGTCTTCCATGGAAATGCCTTCACAGCCAGCTCCACGAACAGTCACAGATGAGGAGAT AAATTTTGTTAAGACCTGTCTTCAGAGATGGAGGAGTGAGATTGAACAAGATATACAA GATTTAAAGACTTGTATTGCAAGTACTACTCAGACTATTGAACAGATGTACTGCGATC CTCTCCTTCGTCAGGTGCCTTATCGCTTGCATGCAGTTCTTGTTCATGAAGGACAAGC AAATGCTGGACACTATTGGGCCTATATCTATAATCAACCCCGACAGAGCTGGCTCAAG TACAATGACATCTCTGTTACTGAATCTTCCTGGGAAGAAGTTGAAAGAGATTCCTATG GAGGCCTGAGAAATGTTAGTGCTTACTGTCTGATGTACATTAACGACAAACTACCCTA CTTCAATGCAGAGGCAGCCCCAACTGAATCAGATCAAATGTCAGAAGTGGAAGCCCTA TCTGTGGAACTCAAGCATTACATTCAGGAGGATAACTGGCGGTTTGAGCAGGAAGTAG AGGAGTGGGAAGAAGAGCAGTCTTGCAAAATCCCTCAAATGGAGTCCTCCACCAACTC CTCATCACAGGACTACTCTACATCACAAGAGCCTTCAGTAGCCTCTTCTCATGGGGTT CGCTGCTTGTCATCTGAGCATGCTGTGATTGTAAAGGAGCAAACTGCCCAGGCTATTG CAAACACAGCCCGTGCCTATGAGAAGAGCGGTGTAGAAGCGGCACTGAGTGAGGCATT CCATGAAGAATACTCCAGGCTCTATCAGCTTGCCAAAGAGACCCCCACCTCTCACAGT GATCCTCGACTTCAGCATGTCCTTGTCTACTTTTTCCAAAATGAAGCACCCAAAAGGG TAGTAGAACGAACCCTTCTGGAACAGTTTGCAGATAAAAATCTTAGCTATGATGAAAG ATCAATCAGCATTATGAAGGTGGCTCAAGCGAAACTGAAGGAAATTGGTCCAGATGAC ATGAATATGGAAGAGTACAAGAAGTGGCATGAAGATTATAGTTTGTTCCGAAAAGTGT CTGTGTATCTCCTAACAGGCCTAGAACTCTATCAAAAAGGAAAGTACCAAGAGGCACT TTCCTACCTGGTATATGCCTACCAGAGCAATGCTGCCCTGCTGATGAAGGGGCCCCGC CGGGGGGTCAAAGAATCCGTGATTGCTTTATACCGAAGAAAATGCCTTCTGGAGCTGA ATGCCAAAGCAGCTTCTCTTTTTGAAACAAATGATGATCACTCCGTAACTGAGGGCAT TAATGTGATGAATGAACTGATCATCCCCTGCATTCACCTTATCATTAATAATGACATT TCCAAGGATGATCTGGATGCCATTGAGGTCATGAGAAACCATTGGTGCTCTTACCTTG GGCAAGATATTGCAGAAAATCTGCAGCTGTGCCTAGGGGAGTTTCTACCCAGACTTCT AGATCCTTCTGCAGAAATCATCGTCTTGAAAGAGCCTCCAACTATTCGACCCAATTCT CCCTATGACCTATGTAGCCGATTTGCAGCTGTCATGGAGTCAATTCAGGGAGTTTCAA CTGTGACAGTGAAATAAGCTCCCACATGTTCAAGGCCCATTCTGGTTCCTGGCTGCCT GCCTCTTGCACAGAAGTTCGTTGTCATAGTGCTCACCTTGG
ORF Start: ATG at 12|θRF Stop: TAA at 3147
SEQ ED NO: 278 1045 aa MW at l l9107.7kD
NOV96c, MTAELQQDDAAGAADGHGSSCQ LLNQLREITGIQDPSFLHEALRASNGDITQAVSLL TDERVKEPSQDTVATEPSEVEGSAANKEVLAKVIDLTHDNKDDLQAAIALSLLESPKI
CG59708-03 Protein Sequence QADGRDLNRMHEATSAETKRSKRKRCEVWGENPNPND RRVDG PVGLKNVGNTCWFS AVIQSLFQLPEFRRLVLSYSLPQNVLENCRSHTEKRNIMFMQELQYLFALMMGSNRKF VDPSAALDLLKGAFRSSEEQQQDVSEFTHKLLD LEDAFQLAVNVNSPRNKFENP VQ LFYGTFLTEGVREGKPFCNNETFGQYPLQVNGYRNLDECLEGAMVEGDVELLPSDHSV KYGQER FTKLPPVLTFELSRFEFNQSLGQPEKIHNKLEFPQIIYMDRYMYRSKELIR NKRECIRKLKEEIKILQQKLEGYVKYGSGPARFPLPDMLKYVIEFASTKPASESCPPE SDTHMTLPLSSVHCSVSNQTSKESTSTESSSQDVESTFSSPEDSLPKSKPLTSSRSSM EMPSQPAPRTVTDEEINFVKTCLQRWRSEIEQDIQDLKTCIASTTQTIEQMYCDPLLR QVPYRLHAVLVHEGQANAGHYWAYIYNQPRQSWLKYNDISVTESS EEVERDSYGGLR NVSAYCLMYINDKLPYFNAEAAPTESDQMSEVEALSVELKHYIQEDNWRFEQEVEEWE EEQSCKIPQMESSTNSSSQDYSTSQEPSVASSHGVRCLSSEHAVIVKEQTAQAIANTA RAYEKSGVEAALSEAFHEEYSRLYQLAKETPTSHSDPRLQHVLVYFFQNEAPKRWER TLLEQFADKNLSYDERSISIMKVAQAKLKEIGPDDMNMEEYKKWHEDYSLFRKVSVYL LTGLELYQKGKYQEALSYLVYAYQSNAALLMKGPRRGVKESVIALYRRKCLLELNAKA ASLFETNDDHSVTEGINVMNELIIPCIHLIINNDISKDDLDAIEVMRNH CSYLGQDI AENLQLCLGEFLPRLLDPSAEIIVLKEPPTIRPNSPYDLCSRFAAVMESIQGVSTVTV K
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 96B.
Table 96B. Comparison of NOV96a against NOV96b through NOV96c.
NOV96a Residues/ Identities/
Protein Sequence Match Residues Similarities for the Matched Region
NOV96b 1 209-1045 j 805/837 (96%) 138..974 j 805/837 (96%)
NOV96c 1..1045 979/1045 (93%) 1..1045 j 981/1045 (93%)
Further analysis of the NOV96a protein yielded the following properties shown in Table 96C.
Table 96C. Protein Sequence Properties NOV96a
PSort 0.8800 probability located in nucleus; 0.3000 probability located in microbody analysis: (peroxisome); 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV96a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 96D.
In a BLAST search of public sequence databases, the NOV96a protein was found to have homology to the proteins shown in the BLASTP data in Table 96E.
PFam analysis predicts that the NOV96a protein contains the domains shown in the Table 96F.
Table 96F. Domain Analysis of NOV96a
Identities/
Pfam Domain NOV96a Match Region Similarities Expect Value for the Matched Region
Example 97.
The NOV97 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 97A.
Table 97A. NOV97 Sequence Analysis
SEQ ED NO: 279 1601 bp
NOV97a, AGGGCAGAGGCCACAGCGCCATCCCCTTCCCCATGGTCTCCCTACCCCCAACCTGCAC
TGGGCGCTCCGCCCAGAGGTGAGTCCCTCCCAGCCCTTCTCTCCTTCTGTCCTAGCCA
CG59559-01 DNA Sequence TCCGCAGAGCCATCCTGTGCAAAGGAAGGAGCTAGGCTGTGCGCCCTGGGCGTCATGA
TCCTTCTGCGGGCCTCCGAAGTGCGGCAGCTGCTTCACAATAAGTTCGTGGTCATCCT GGGGGACTCTGTGCATAGGGCAGTATACAAGGACCTGGTGCTTCTGCTGCAGAAGGAC CGCCTGCTCACTCCCGGGCAGCTTAGAGCAAGGGGGGAGCTGAACTTCGAACAAGATG AGCTGGTGGACGGAGGCCAGCGGGGCCACATGCACAACGGCCTTAACTACCGTGAGGT CCGCGAGTTCCGCTCCGACCACCATCTGGTACGTTTTTACTTCCTCACCCGCGTGTAC TCCGATTACCTCCAGACCATCTTGAAAGAGCTGCAGTCGGGCGAGCACGCCCCCGACC TGGTCATCATGAATTCCTGCCTCTGGGACATCTCCAGGTATGGTCCGAACTCCTGGAG AAGCTACCTGGAGAACCTGGAGAACCTGTTCCAGTGCCTGGGCCAGGTGCTGCCCGAG TCTTGCCTCCTGGTGTGGAACACGGCCATGCCTGTGGGCGAGGAAGTCACCGGGGGTT TTCTTCCGCCCAAGCTCCGGCGGCAGAAGGCCACCTTCCTGAAAAACGAAGTGGTCAA AGCCAACTTCCACAGCGCCACCGAGGCACGTAAACATAACTTCGATGTACTGGACTTG CATTTCCACTTCCGCCACGCGAGGGAGAACCTGCACTGGGACGGGGTGCACTGGAATG GACGTGTGCACCGCTGCCTCTCCCAGCTGCTGCTGGCCCACGTGGCCGACGCCTGGGG TGTGGAGCTGCCCCACCGCCACCCCGTGGGCGAGTGGATCAAGAAGAAAAAACCTGGC CCGAGAGTCGAAGGGCCGCCCCAGGCCAACAGAAATCACCCGGCCTTACCTCTGTCCC CACCCTTACCTTCCCCCACATACCGCCCCCTGCTTGGGTTCCCACCCCAGCGCTTGCC GCTGCTCCCGCTCCTGTCCCCACAGCCTCCTCCTCCCATTCTCCATCACCAGGGAATG CCCCGGTTCCCACAGGGTCCCCCAGATGCCTGTTTTTCCTCAGACCATACTTTCCAGT CGGATCAATTCTATTGCCATTCAGATGTCCCCTCATCAGCCCATGCAGGTTTCTTCGT CGAAGACAATTTTATGGTTGGTCCTCAGCTGCCTATGCCCTTCTTCCCCACACCCCGT TATCAGCGGCCTGCCCCAGTGGTACATAGGGGTTTTGGCAGGTATCGTCCCCGTGGCC CCTATACGCCCTGGGGACAGCGGCCTCGACCTTCAAAGAGAAGGGCCCCAGCCAATCC TGAGCCAAGGCCTCAATAGACGGACCTAGGCCTTATTTCCTCTTTATGAACATGGATT GGACAGATCTGACACTTCCTTTCCATTGCTTGGCCTGAACAGACTGACCTTGTTAACT
TAAGCCTGGAGTCCATGCCTCGTCTTCCTTTTGTT
ORF Start: ATG at 171 [ORF Stop: TAG at 1467
SEQ ED NO: 280 432 aa MW at 49726.6kD
NOV97a, MILLRASEVRQLLHNKFWILGDSVHRAVYKDLVLLLQKDRLLTPGQLRARGELNFEQ DELVDGGQRGHMHNGLNYREVREFRSDHHLVRFYFLTRVYSDYLQTILKELQSGEHAP
CG59559-01 Protein Sequence DLVIMNSCL DISRYGPNS RSYLENLENLFQCLGQVLPESCLLVWNTAMPVGEEVTG GFLPPKLRRQKATFLKNEWKANFHSATEARKHNFDVLDLHFHFRHARENLH DGVHW NGRVHRCLSQLLLAHVADA GVELPHRHPVGE IKKKKPGPRVEGPPQANRNHPALPL SPPLPSPTYRPLLGFPPQRLPLLPLLSPQPPPPILHHQGMPRFPQGPPDACFSSDHTF QSDQFYCHSDVPSSAHAGFFVEDNF VGPQLPMPFFPTPRYQRPAPWHRGFGRYRPR GPYTP GQRPRPSKRRAPANPEPRPQ
Further analysis of the NOV97a protein yielded the following properties shown in Table 97B.
Table 97B. Protein Sequence Properties NOV97a
PSort 0.5937 probability located in mitochondrial matrix space; 0.5103 probability analysis: located in microbody (peroxisome); 0.4900 probability located in nucleus; 0.3252 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV97a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 97C.
In a BLAST search of public sequence databases, the NOV97a protein was found to have homology to the proteins shown in the BLASTP data in Table 97D.
PFam analysis predicts that the NOV97a protein contains the domains shown in the Table 97E.
Table 97E. Domain Analysis of NOV97a
Identities/
Pfam Domain NOV97a Match Region Similarities Expect Value for the Matched Region
No Significant Matches Found
Example 98.
The NOV98 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 98A.
Further analysis of the NOV98a protein yielded the following properties shown in Table 98B.
Table 98B. Protein Sequence Properties NOV98a
PSort 0.4766 probability located in mitochondrial matrix space; 0.4500 probability analysis: located in cytoplasm; 0.1822 probability located in mitochondrial inner membrane; 0.1822 probability located in mitochondrial intermembrane space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV98a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 98C.
In a BLAST search of public sequence databases, the NOV98a protein was found to have homology to the proteins shown in the BLASTP data in Table 98D.
PFam analysis predicts that the NOV98a protein contains the domains shown in the Table 98E.
Example 99.
The NOV99 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 99A.
Further analysis of the NOV99a protein yielded the following properties shown in Table 99B.
Table 99B. Protein Sequence Properties NOV99a
PSort 0.6000 probability located in plasma membrane; 0.4000 probability located in analysis: Golgi body; 0.3000 probability located in endoplasmic reticulum (membrane); 0.1000 probability located in mitochondrial inner membrane
SignalP Likely cleavage site between residues 55 and 56 analysis:
A search of the NOV99a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 99C.
In a BLAST search of public sequence databases, the NOV99a protein was found to have homology to the proteins shown in the BLASTP data in Table 99D.
PFam analysis predicts that the NOV99a protein contains the domains shown in the Table 99E.
Example 100.
The NOVl 00 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 100A.
Further analysis of the NOVlOOa protein yielded the following properties shown in Table 100B.
Table 100B. Protein Sequence Properties NOVlOOa
PSort 0.3600 probability located in mitochondrial matrix space; 0.3000 probability analysis: located in microbody (peroxisome); 0.1808 probability located in lysosome (lumen); 0.0000 probability located in endoplasmic reticulum (membrane)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVlOOa protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table lOOC.
SEP-2000]
In a BLAST search of public sequence databases, the NOVlOOa protein was found to have homology to the proteins shown in the BLASTP data in Table 100D.
PFam analysis predicts that the NOVlOOa protein contains the domains shown in the Table 100E.
Example 101.
The NOVl 01 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 101 A.
Further analysis of the NOVlOla protein yielded the following properties shown in Table 101B.
Table 101B. Protein Sequence Properties NOVlOla
PSort 0.5708 probability located in mitochondrial matrix space; 0.4996 probability analysis: located in mitochondrial intermembrane space; 0.2852 probability located in mitochondrial inner membrane; 0.2852 probability located in mitochondrial outer membrane
SignalP Likely cleavage site between residues 23 and 24 analysis:
A search of the NOVlOla protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table lOlC.
In a BLAST search of public sequence databases, the NOVlOla protein was found to have homology to the proteins shown in the BLASTP data in Table 101D.
PFam analysis predicts that the NOVlOla protein contains the domains shown in the Table 101E.
Example 102.
The NOVl 02 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 102A.
Further analysis of the NOVl 02a protein yielded the following properties shown in Table 102B.
Table 102B. Protein Sequence Properties NOV102a
PSort 0.6400 probability located in microbody (peroxisome); 0.4500 probability analysis: located in cytoplasm; 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
No Known Signal Sequence Predicted
analysis:
A search of the NOVl 02a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 102C.
In a BLAST search of public sequence databases, the NOVl 02a protein was found to have homology to the proteins shown in the BLASTP data in Table 102D.
PFam analysis predicts that the NOVl 02a protein contains the domains shown in the Table 102E.
Example 103.
The NOVl 03 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 103 A.
Table 103 A. NOVl 03 Sequence Analysis
SEQ ED NO: 291 8860 bp
NOV103a, GGATCCTTGAGGGCACTGGTGCGACTTTCAGGTGAGGTCTTAGCAGATGAAAGCGGCT
GGCTGTGGCCCGCGCCAGTAGTGCTTTCTGCTCCGCACTCGCCGTGAGCCAGGTGTGC
CG59773-01 DNA Sequence AACCGGATTTGGGGCGAGGGTCGCGCTGGCTACCTCGCATGCGCAGAGCCGGAAGCCC
GCTGACCGGACTACAGCTCCCAGAAGAGCCTTGTGGAGGCCGCAGACGCGAAGCCGCT
GGCGCCATCTTGAAATCTGATCCTCCATCCCCGAGGCTTTGCGTCTGCGCGGCCGGCC
GCTGCTGCTCCGGGAGCCCAGTCTGCTAAAAGGGGAGGACGTTGAGGACGCGGCGGCT
GGCGGGAGAGACAGCTGGGGAGAGACATGGCAGGGTCGGAGCGCGGCCTGCGCCTCTG
TCACTCAGCATCCTCTTAGGCGTTTCCACGCCCGCCCCCTGCCCGAGGGGCGGGGCTG
ACGGCTCTGGTACCCGGAGTCGGCGCGCGGGGCAGGGGCGCGCCCCTGCAGAGTGGGG
ACCCCACTGGGCTGTGCCATGCTGACCGGAGACCACCGAGGCGGGAGACAGAGCGCGG
CGAAGAGCCATTGAGTGGTCACCCAGTAGCCGCCGCCGCCGCCGCCTCGGGAAGCTTG
CCACCCGCTAGGAGGGAAGATGAAGGAGATTTGCAGGATCTGTGCCCGAGAGCTGTGT
GGAAACCAGCGGCGCTGGATCTTCCACACGGCGTCCAAGCTCAATCTCCAGGTTCTGC TTTCGCACGTCTTGGGCAAGGATGTCCCCCGCGATGGCAAAGCCGAGTTCGCTTGCAG CAAGTGTGCTTTCATGCTTGATCGAATCTATCGATTCGACACAGTTATTGCCCGGATT GAAGCGCTTTCTATTGAGCGCTTGCAAAAGCTGCTACTGGAGAAGGATCGCCTCAAGT TCTGCATTGCCAGTATGTATCGGAAGAATAACGATGACTCTGGCGCGGAGATCAAGGC GGGGAATGGGACGGTTGACATGTCCGTCTTACCCGATGCGAGATACTCTGCACTGCTC
CAGGAGGACTTCGCCTATTCAGGGTTTGAGTGCTGGGTGGAGAATGAGGATCAGATCC AGGAGCCACACAGCTGCCATGGTTCAGAAGGCCCTGGAAACCGACCCAGGAGATGCCG TGGTTGTGCCGCTTTGCGGGTTGCTGATTCTGACTATGAAGCCATTTGTAAGGTACCT CGAAAGGTGGCCAGAAGTATCTCCTGCGGCCCTTCTAGCAGGTGGTCGACCAGCATTT GCACTGAAGAACCAGCGTTGTCTGAGGTTGGGCCACCCGACTTAGCAAGCACAAAGGT ACCCCCAGATGGAGAAAGCATGGAGGAAGAGACGCCTGGTTCCTCTGTGGAATCTTTG GATGCAAGCGTCCAGGCTAGCCCTCCACAACAGAAAGATGAGGAGACTGAGAGAAGTG CAAAGGAACTTGGAAAGTGTGACTGTTGTTCAGATGATCAGGCTCCGCAGCATGGGTG TAATCACAAGCTGGAATTAGCTCTTAGCATGATTAAAGGTCTTGATTATAAGCCCATC CAGAGCCCCCGAGGGAGCAGGCTTCCGATTCCAGTGAAATCCAGCCTACCTGGAGCCA AGCCTGGCCCTAGCATGACAGATGGAGTTAGTTCCGGTTTCCTTAACAGGTCTTTGAA ACCCCTTTACAAGACACCTGTGAGTTATCCCTTGGAGCTTTCAGACCTGCAGGAGCTG TGGGATGATCTCTGTGAAGATTATTTGCCGCTCCGGGTCCAGCCCATGACTGAAGAGT TGCTGAAACAACAAAAGCTGAATTCACATGAGACCACTATAACTCAGCAGTCTGTATC TGATTCCCACTTGGCAGAACTCCAGGAAAAAATCCAGCAAACAGAGGCCACCAACAAG ATTCTTCAAGAGAAACTTAATGAAATGAGCTATGAACTAAAGTGTGCTCAGGAGTCGT CTCAAAAGCAAGATGGTACAATTCAGAACCTCAAGGAAACTCTGAAAAGCAGGGAACG TGAGACTGAGGAGTTGTACCAGGTAATTGAAGGTCAAAATGACACAATGGCAAAGCTT CGAGAAATGCTGCACCAAAGCCAGCTTGGACAACTTCACAGCTCAGAGGGTACTTCTC CAGCTCAGCAACAGGTAGCTCTGCTTGATCTTCAGAGTGCTTTATTCTGCAGCCAACT TGAAATACAGAAGCTCCAGAGGGTGGTACGACAGAAAGAGCGCCAACTGGCTGATGCC AAACAATGTGTGCAATTTGTAGAGGCTGCAGCACACGAGAGTGAACAGCAGAAAGAGG CTTCTTGGAAACATAACCAGGAATTGCGAAAAGCCTTGCAGCAGCTACAAGAAGAATT GCAGAATAAGAGCCAACAGCTTCGTGCCTGGGAGGCTGAAAAATACAATGAGATTCGA ACCCAGGAACAAAACATCCAGCACCTAAACCATAGTCTGAGTCACAAGGAGCAGTTGC TTCAGGAATTTCGGGAGCTCCTACAGTATCGAGATAACTCAGACAAAACCCTTGAAGC AAATGAAATGTTGCTTGAGAAACTTCGCCAGCGAATACATGATAAAGCTGTTGCTCTG GAGCGGGCTATAGATGAAAAATTCTCTGCTCTAGAAGAGAAAGAAAAAGAACTGCGCC AGCTTCGTCTTGCTGTGAGAGAGCGAGATCATGACTTAGAGAGACTGCGCGATGTCCT CTCCTCCAATGAAGCTACTATGCAAAGTATGGAGAGTCTCCTGAGGGCCAAAGGCCTG GAAGTGGAACAGTTATCTACTACCTGTCAAAACCTCCAGTGGCTGAAAGAAGAAATGG AAACCAAATTTAGCCGTTGGCAGAAGGAACAAGAGAGTATCATTCAGCAGTTACAGAC GTCTCTTCATGATAGGAACAAAGAAGTGGAGGATCTTAGTGCAACACTGCTCTGCAAA CTTGGACCAGGGCAGAGTGAGATAGCAGAGGAGCTGTGCCAGCGTCTACAGCGAAAGG AAAGGATGCTGCAGGACCTTCTAAGTGATCGAAATAAACAAGTGCTGGAACATGAAAT GGAGATTCAAGGCCTGCTTCAGTCTGTGAGCACCAGGGAGCAGGAAAGCCAAGCTGCT GCAGAGAAGTTGGTGCAAGCCTTAATGGAAAGAAATTCAGAATTACAGGCCCTGCGCC AATATTTAGGAGGGAGAGACTCCCTGATGTCCCAAGCACCCATCTCTAACCAACAAGC TGAAGTTACCCCCACTGGCCGTCTTGGAAAACAGACTGATCAAGGTTCAATGCAGATA CCTTCCAGAGATGATAGCACTTCATTGACTGCCAAAGAGGATGTCAGCATACCCAGAT CCACATTAGGAGACTTGGACACAGTTGCAGGGCTGGAAAAAGAACTGAGTAATGCCAA AGAGGAACTTGAACTCATGGCTAAAAAAGAAAGAGAAAGTCAGATGGAACTTTCTGCT CTACAGTCCATGATGGCTGTGCAGGAAGAAGAGCTGCAGGTGCAGGCTGCTGATATGG AGTCTCTGACCAGGAACATACAGATTAAAGAAGATCTCATAAAGGACCTGCAAATGCA ACTGGTTGATCCTGAAGACATACCAGCTATGGAACGCCTGACCCAGGAAGTCTTACTT CTTCGGGAAAAAGTTGCTTCAGTAGAATCCCAGGGTCAAGAAATTTCAGGAAACCGAA GACAACAGTTGCTGCTGATGCTAGAAGGACTAGTAGATGAACGGAGTCGGCTCAATGA GGCCTTACAAGCAGAGAGACAGCTCTATAGCAGTCTGGTGAAGTTCCATGCCCATCCA GAGAGCTCTGAGAGAGACCGAACTCTGCAGGTGGAACTGGAAGGGGCTCAGGTGTTAC GCAGTCGGCTAGAAGAAGTTCTTGGAAGAAGCTTGGAGCGCTTAAACAGGCTGGAGAC CCTGGCCGCCATTGGAGGTGCAGCTGCAGGGGATGACACCGAAGATACAAGCACTGAG TTCACTGACAGTATTGAGGAGGAGGCTGCACACCATAGTCACCAGCAACTTGTCAAGG TGGCTTTGGAGAAAAGTCTGGCAACTGTGGAGACCCAGAACCCATCTTTTTCCCCTCC TTCTCCGATGGGAGGGGACAGTAACAGGTGTCTTCAGGAAGAAATGCTCCACCTGAGG GCTGAGTTCCACCAGCACTTAGAAGAGAAGAGGAAAGCTGAGGAGGAACTGAAGGAGC TAAAGGCTCAAATTGAGGAAGCAGGATTCTCCTCAGTGTCCCACATCAGGAACACCAT GCTGAGCCTTTGCCTTGAGAATGCGGAGCTGAAAGAGCAGATGGGAGAAGCAATGTCT GATGGATGGGAGATCGAGGAAGACAAGGAGAAGGGCGAGGTGATGGTTGAGACTGTGG TAACCAAAGAGGGTCTGAGTGAGAGTAGCCTTCAGGCTGAGTTCAGAAAGCTCCAGGG AAAACTGAAGAATGCCCACAATATCATCAACCTCCTCAAAGAACAACTTGTGCTGAGT AGCAAGGAAGGGAATAGTAAACTTACTCCAGAGCTCCTTGTGCATCTGACCAGCACCA TTGAAAGAATAAACACAGAACTGGTTGGTTCCCCTGGGAAGCACCAACACCAAGAGGA GGGGAATGTGACTGTGAGGCCTTTCCCCAGACCCCAGAGCCTTGACCTTGGGGCTACC TTCACAGTGGATGCCCACCAATTGGATAACCAGTCCCAGCCTCGTGACCCTGGGCCTC AGTCAGCGTTTAGCCTACCAGGGTCCACCCAGCACCTGCGCTCCCAGCTGTCACAATG CAAACAACGCTATCAAGATCTCCAGGAGAAGCTGCTGCTATCAGAAGCCACTGTCTTT GCTCAGGCTAACGAGCTGGAGAAATACAGAGTTATGCTTACAGGTGAATCCTTGGTGA AGCAGGACAGCAAGCAGATCCAGGTGGACCTCCAGGACCTGGGCTATGAGACTTGTGG CCGAAGCGAGAATGAGGCTGAACGGGAGGAAACCACCAGTCCTGAGTGTGAGGAGCAC AACAGCCTCAAGGAAATGGTCCTGATGGAGGGGCTGTGCTCTGAGCAGGGACGCCGGG GCTCAACACTGGCTAGTTCCTCTGAGAGGAAGCCCTTGGAGAACCAGCTAGGGAAGCA GGAAGAGTTCCGGGTATATGGAAAGTCAGAAAACATCTTGGTCCTACGAAAGGACATC AAAGATCTGAAGGCCCAGCTGCAGAATGCCAACAAGGTCATTCAAAACCTCAAGAGCC GGGTCCGGTCCCTCTCAGTTACAAGTGATTATTCGTCTAGTCTGGAAAGACCCCGGAA GCTGAGAGCTGTTGGCACCTTGGAGGGGTCTTCACCTCATAGTGTCCCTGATGAGGAT GAGGGGTGGCTGTCTGATGGCACTGGGGCTTTCTACTCTCCAGGGCT CAGGCCAAAA AGGACCTGGAGAGTCTCATCCAGAGAGTATCCCAGCTGGAGGCCCAGCTCCCAAAAAA
TGGACTAGAAGAGAAGCTGGCTGAGGAGCTGAGATCAGCCTCGTGGCCTGGGAAATAT GATTCCCTGATTCAGGATCAGGCCCGGGAACTGTCTTACCTACGGCAAAAAATACGAG AAGGGAGAGGTATTTGTTATCTTATCACCCGGCATGCAAAAGATACAGTAAAATCTTT TGAGGATCTCCTAAGGAGCAATGACATTGACTACTACCTGGGACAGAGCTTCCGGGAG CAACTCGCCCAGGGAAGCCAGCTGACAGAGAGGCTCACCAGCAAACTCAGCACCAAGG ATCATAAAAGTGAGAAAGATCAAGCTGGACTTGAGCCACTGGCCCTCAGGCTCAGCAG GGAGCTGCAGGAGAAGGAGAAAGTGATTGAAGTCCTGCAGGCCAAGCTGGATGCTCGG TCCCTCACACCCTCCAGCAGCCATGCCTTGTCTGACTCCCACCGCTCTCCCAGCAGCA CCTCTTTCCTGTCTGATGAACTGGAAGCCTGCTCTGACATGGACATAGTCAGCGAGTA CACACACTATGAAGAGAAGAAAGCTTCTCCCAGTCACTCAGATTCCATCCATCATTCG AGTCATTCTGCTGTGTTGTCTTCTAAACCATCATCAACCAGTGCATCTCAGGGGGCTA AGGCCGAATCCAACAGCAACCCCATCAGCTTGCCAACTCCCCAGAATACCCCCAAGGA GGCCAACCAGGCCCATTCAGGCTTTCATTTTCACTCCATACCCAAGCTGGCTAGCCTT CCTCAGGCACCATTGCCCTCAGCTCCATCCAGCTTCCTGCCTTTCAGCCCCACTGGCC CTCTCCTCCTTGGCTGCTGTGAGACACCAGTGGTCTCCTTGGCTGAGGCTCAGCAGGA GCTACAGATGCTGCAGAAGCAGTTGGGAGAAAGTGCCAGCACTGTTCCTCCTGCTTCC ACAGCTACATTGCTGAGCAACGACTTGGAAGCCGACTCTTCCTACTACCTCAACTCTG CCCAGCCTCACTCTCCTCCAAGGGGCACCATAGAACTGGGAAGAATCCTAGAGCCTGG GTACCTGGGCAGCAGTGGCAAGTGGGATGTGATGAGGCCTCAGAAAGGGAGTGTATCT GGGGACCTATCCTCAGGCTCCTCTGTGTACCAGCTTAACTCCAAACCCACAGGGGCTG ACCTGCTGGAAGAGCATCTTGGTGAAATCCGGAACCTGCGCCAGCGCCTGGAGGAGTC CATCTGCATCAATGACCGCCTACGGGAGCAACTGGAACACCGGCTGACCTCTACTGCT CGTGGAAGGGGATCCACTTCTAACTTCTACAGTCAGGGCCTGGAGTCCATACCTCAGC TCTGCAATGAGAACAGAGTCCTCAGGGAAGACAATCGAAGACTTCAGGCTCAACTGAG TCATGTTTCCAGAGAGCACTCCCAGGAAACAGAAAGCCTGAGGGAGGCTCTGCTGTCC TCTCGATCCCACCTTCAAGAGCTGGAAAAGGAGCTGGAGCACCAGAAGGTGGAAAGGC AGCAGCTTTTGGAAGACTTGAGGGAGAAGCAGCAAGAGGTCTTGCATTTCAGGGAGGA ACGTCTTTCCCTCCAGGAAAACGACTCCAGACTGCAGCACAAGCTGGTTCTCCTGCAG CAACAGTGTGAAGAGAAACAGCAGCTCTTTGAGTCCCTCCAGTCAGAGCTACAAATCT ACGAGGCACTTTATGGCAATTCCAAGAAGGGGCTGAAAGGCTTGGGTTTGGATACTTC TCCAGTAATGAAGACCCCTCCCAAGCTAGAGGGTGATGCTACTGATGGCTCCTTTGCC AATAAGCATGGCCGCCATGTCATTGGCCACATTGATGACTACAGTGCCCTAAGACAGC AGATTGCGGAGGGCAAGCTGCTGGTCAAAAAGATAGTGTCTCTTGTGAGATCAGCGTG CAGCTTCCCTGGCCTTGAAGCCCAAGGCACAGAGGTGCTAGGCAGCAAAGGTATTCAT GAGCTTCGGAGCAGCACCAGTGCCCTGCACCATGCCCTAGAGGAGTCGGCTTCCCTCC TCACCATGTTCTGGAGAGCAGCCCTGCCAAGCACCCACATCCCTGTGCTGCCTGGCAA AGTGGGAGAATCAACAGAAAGGGAACTTCTGGAACTGAGAACCAAAGTATCCAAACAG GAGCGGCTCCTTCAGAGCACAACTGAGCATCTGAAGAACGCCAACCAGCAGAAGGAGA GCATGGAGCAGTTCATCGTCAGCCAGCTAACCAGAACACATGATGTTTTAAAGAAGGC AAGGACTAACTTAGAGGTGAAATCCCTAAGGGCTCTGCCATGTACTCCAGCCTTGTGA CCCTTGCCTTCCAGGAACCATGCAAGAAGCGCAGCCACCAGAAGTCCTTAAAACAGCA
GGAAAGGTGGGCCTGTCCCCCTTTTGTGCAGCTACCTATCTGCTGAGGAGCATCTGGG
CCTCATTCCTCCAAGTCCACGGGAGGGTCCAGAAGAGGGAGTCAGAGATGTATCCTGG
TGGAGCTGGGAGAAAGGCAGAAAGCCTTTCTGACAGCTATGGAATACGATTAGCCAAG
GTCCACTTGGCCCAGCACTAAGAAAAAGATGCGTAGTTTGCACAGAAGGTTTTGTGAT
CCTGCCTCTCAACAGCCCCAGCAGCTTGGGAACTAGCAAGAGCACATTTCTTGCCTCA
TCAGCTGTCCTGAGATGGAAAACTCAGTGGATATAGGACCCTGATTCCGATGAAAGGG
GCACGTGGTCCCAATGCTGGAGCTCCTCTGGCAGGTTCTAAAAGCACACTACTGAGCA
GCGGTGCCCTGCCGGACACTGCTGGCGGGGGCTCAGTGAGCACTACTCACAGATCCAC
ACCTGACCCTGTTGGGTCGAGTCAGGCTGGGCCTTGGTCTGCACTGTAGCACCTGTGT
TCTTTGAGTTCACATCATGAATGTGGTGACTTCCCAGATACCATCTCAGGCTTAACCT
AGCACATCCTATTTCTTTTCTTCTATGATATCCAAATTGGACTGACCTCACTTCAAAG
TTGCTGTCCCATTTTGTCACCCTATCTTATCTCGGGGAAATTGCAGACTGATGGCCAG
ACCAACTCTGTTGAAATTCTTGCATAGAGCAAACCTGTGCTCATTTTTAAGTGGCATG
GGAGAGGCCCCCAGCCTAGTAAAGCCTAGTCTGTGTCTTCACAGTGCTGGTAGAATGT
GTTTGTGTGTATAAATATATGATATAGATTTATATATGTTGCTAACGCCATATATTGA
AGGCCAACATAACTGGTGGACAGGGTGGGTGACAGAAAATGAAAGCCTTTTTGGTGAT
TGTTAAAGCAAGATGTGTATAAAGAAATAAATAGTTTTTCTTTC
ORF Start: ATG at 658 ORF Stop: TGA at 7828
SEQ ID NO: 292 2390 aa MW at 268843.7kD
NOV103a, MKEICRICARELCGNQRRWIFHTASKLNLQVLLSHVLGKDVPRDGKAEFACSKCAFML DRIYRFDTVIARIEALSIERLQKLLLEKDRLKFCIASMYRKNNDDSGAEIKAGNGTVD
CG59773-01 Protein Sequence MSVLPDARYSALLQEDFAYSGFEC VENEDQIQEPHSCHGSEGPGNRPRRCRGCAALR VADSDYEAICKVPRKVARSISCGPSSRWSTSICTEEPALSEVGPPDLASTKVPPDGES MEEETPGSSVESLDASVQASPPQQKDEETERSAKELGKCDCCSDDQAPQHGCNHKLEL ALSMIKGLDYKPIQSPRGSRLPIPVKSSLPGAKPGPSMTDGVSSGFLNRSLKPLYKTP VSYPLELSDLQELWDDLCEDYLPLRVQPMTEELLKQQKLNSHETTITQQSVSDSHLAE LQEKIQQTEATNKILQEKLNEMSYELKCAQESSQKQDGTIQNLKETLKSRERETEELY QVIEGQNDTMAKLREMLHQSQLGQLHSSEGTSPAQQQVALLDLQSALFCSQLEIQKLQ RWRQKERQLADAKQCVQFVEAAAHESEQQKEASWKHNQELRKALQQLQEELQNKSQQ LRA EAEKYNEIRTQEQNIQHLNHSLSHKEQLLQEFRELLQYRDNSDKTLEANEMLLE KLRQRIHDKAVALERAIDEKFSALEEKEKELRQLRLAVRERDHDLERLRDVLSSNEAT MQSMESLLRAKGLEVEQLSTTCQNLQ LKEEMETKFSRWQKEQESIIQQLQTSLHDRN KEVEDLSATLLCKLGPGQSEIAEELCQRLQRKERMLQDLLSDRNKQVLEHEMEIQGLL
QSVSTREQESQAAAEKLVQALMERNSELQALRQYLGGRDSLMSQAPISNQQAEVTPTG RLGKQTDQGSMQIPSRDDSTSLTAKEDVSIPRSTLGDLDTVAGLEKELSNAKEELELM AKKERESQMELSALQSM AVQEEELQVQAADMESLTRNIQIKEDLIKDLQMQLVDPED IPAMERLTQEVLLLREKVASVESQGQEISGNRRQQLLLMLEGLVDERSRLNEALQAER QLYSSLVKFHAHPESSERDRTLQVELEGAQVLRSRLEEVLGRSLERLNRLETLAAIGG AAAGDDTEDTSTEFTDSIEEEAAHHSHQQLVKVALEKSLATVETQNPSFSPPSPMGGD SNRCLQEEMLHLRAEFHQHLEEKRKAEEELKELKAQIEEAGFSSVSHIRNTMLSLCLE NAELKEQMGEAMSDGWEIEEDKEKGEVMVETWTKEGLSESSLQAEFRKLQGKLKNAH NIINLLKEQLVLSSKEGNSKLTPELLVHLTSTIERINTELVGSPGKHQHQEEGNVTVR PFPRPQSLDLGATFTVDAHQLDNQSQPRDPGPQSAFSLPGSTQHLRSQLSQCKQRYQD LQEKLLLSEATVFAQANELEKYRV LTGESLVKQDSKQIQVDLQDLGYETCGRSENEA EREETTSPECEEHNSLKEMVLMEGLCSEQGRRGSTLASSSERKPLENQLGKQEEFRVY GKSENILVLRKDIKDLKAQLQNANKVIQNLKSRVRSLSVTSDYSSSLERPRKLRAVGT LEGSSPHSVPDEDEGWLSDGTGAFYSPGLQAKKDLESLIQRVSQLEAQLPKNGLEEKL AEELRSASWPGKYDSLIQDQARELSYLRQKIREGRGICYLITRHAKDTVKSFEDLLRS NDIDYYLGQSFREQLAQGSQLTERLTSKLSTKDHKSEKDQAGLEPLALRLSRELQEKE KVIEVLQAKLDARSLTPSSSHALSDSHRSPSSTSFLSDELEACSDMDIVSEYTHYEEK KASPSHSDSIHHSSHSAVLSSKPSSTSASQGAKAESNSNPISLPTPQNTPKEANQAHS GFHFHSIPKLASLPQAPLPSAPSSFLPFSPTGPLLLGCCETPWSLAEAQQELQMLQK QLGESASTVPPASTATLLSNDLEADSSYYLNSAQPHSPPRGTIELGRILEPGYLGSSG K DVMRPQKGSVSGDLSSGSSVYQLNSKPTGADLLEEHLGEIRNLRQRLEESICINDR LREQLEHRLTSTARGRGSTSNFYSQGLESIPQLCNENRVLREDNRRLQAQLSHVSREH SQETESLREALLSSRSHLQELEKELEHQKVERQQLLEDLREKQQEVLHFREERLSLQE NDSRLQHKLVLLQQQCEEKQQLFESLQSELQIYEALYGNSKKGLKGLGLDTSPVMKTP PKLEGDATDGSFANKHGRHVIGHIDDYSALRQQIAEGKLLVKKIVSLVRSACSFPGLE AQGTEVLGSKGIHELRSSTSALHHALEESASLLTMF RAALPSTHIPVLPGKVGESTE RELLELRTKVSKQERLLQSTTEHLKNANQQKES EQFIVSQLTRTHDVLKKARTNLEV KSLRALPCTPAL
SEQ ID NO: 293 7161 bp
NOVl 03b, GTTGAGGGGGCAATCGGGCACGCTCCTCCCCATGGGTTGCCCATCATGTCTAATGGAT
ATCGCACTCTGTCCCAGCACCTCAATGACCTGAAGAAGGAGAACTTCAGCCTCAAGCT
CG59773-02 DNA Sequence GCGCATCTACTTCCTGGAGGAGCGCATGCAACAGAAGTATGAGGCCAGCCGGGAGGAC ATCTACAAGCGGAACATTGAGCTGAAGGTTGAAGTGGAGAGCTTGAAACGAGAACTCC AGGACAAGAAACAGCATCTGGATAAAACATGGGCTGATGTGGAGAATCTCAACAGTCA GAATGAAGCTGAGCTCCGACGCCAGTTTGAGGAGCGACAGCAGGAGACGGAGCATGTT TATGAGCTCTTGGAGAATAAGATCCAGCTTCTGCAGGAGGAATCCAGGCTAGCAAAGA ATGAAGCTGCGCGGATGGCAGCTCTGGTGGAAGCAGAGAAGGAGTGTAACCTGGAGCT CTCAGAGAAACTGAAGGGAGTCACCAAAAACTGGGAAGATGTACCAGGAGACCAGGTC AAGCCCGACCAATACACTGAGGCCCTGGCCCAGAGGGACAGGAGAATTGAAGAACTGA ATCAGAGCCTGGCTGCCCAGGAGAGGCTTGTAGAACAGCTATCTCGGGAGAAACAACA ACTGCTACATCTGTTGGAGGAGCCAACTAGCATGGAAGTGCAGCCCATGACTGAAGAG TTGCTGAAACAACAAAAGCTGAATTCACATGAGACCACTATAACTCAGCAGTCTGTAT CTGATTCCCACTTGGCAGAACTCCAGGAAAAAATCCAGCAAACAGAGGCCACCAACAA GATTCTTCAAGAGAAACTTAATGAAATGAGCTATGAACTAAAGTGTGCTCAGGAGTCG TCTCAAAAGCAAGATGGTACAATTCAGAACCTCAAGGAAACTCTGAAAAGCAGGGAAC GTGAGACTGAGGAGTTGTACCAGGTAATTGAAGGTCAAAATGACACAATGGCAAAGCT TCGAGAAATGCTGCACCAAAGCCAGCTTGGACAACTTCAGAGCTCAGAGGGTACTTCT CCAGCTCAGCAACAGGTAGCTCTGCTTGATCTTCAGAGTGCTTTATTCTGCAGCCAAC TTGAAATACAGAAGCTCCAGAGGGTGGTACGACAGAAAGAGCGCCAACTGGCTGATGC CAAACAATGTGTGCAATTTGTAGAGGCTGCAGCACACGAGAGTGAACAGCAGAAAGAG GCTTCTTGGAAACATAACCAGGAATTGCGAAAAGCCTTGCAGCAGCTACAAGAAGAAT TGCAGAATAAGAGCCAACAGCTTCGTGCCTGGGAGGCTGAAAAATACAATGAGATTCG AACCCAGGAACAAAACATCCAGCACCTAAACCATAGTCTGAGTCACAAGGAGCAGTTG CTTCAGGAATTTCGGGAGCTCCTACAGTATCGAGATAACTCAGACAAAACCCTTGAAG CAAATGAAATGTTGCTTGAGAAACTTCGCCAGCGAATACATGATAAAGCTGTTGCTCT GGAGCGGGCTATAGATGAAAAATTCTCTGCTCTAGAAGAGAAAGAAAAAGAACTGCGC CAGCTTCGTCTTGCTGTGAGAGAGCGAGATCATGACTTAGAGAGACTGCGCGATGTCC TCTCCTCCAATGAAGCTACTATGCAAAGTATGGAGAGTCTCCTGAGGGCCAAAGGCCT GGAAGTGGAACAGTTATCTACTACCTGTCAAAACCTCCAGTGGCTGAAAGAAGAAATG GAAACCAAATTTAGCCGTTGGCAGAAGGAACAAGAGAGTATCATTCAGCAGTTACAGA CGTCTCTTCATGATAGGAACAAAGAAGTGGAGGATCTTAGTGCAACACTGCTCTGCAA ACTTGGACCAGGGCAGAGTGAGATAGCAGAGGAGCTGTGCCAGCGTCTACAGCGAAAG GAAAGGATGCTGCAGGACCTTCTAAGTGATCGAAATAAACAAGTGCTGGAACATGAAA TGGAGATTCAAGGCCTGCTTCAGTCTGTGAGCACCAGGGAGCAGGAAAGCCAAGCTGC TGCAGAGAAGTTGGTGCAAGCCTTAATGGAAAGAAATTCAGAATTACAGGCCCTGCGC CAATATTTAGGAGGGAGAGACTCCCTGATGTCCCAAGCACCCATCTCTAACCAACAAG CTGAAGTTACCCCCACTGGCCGTCTTGGAAAACAGACTGATCAAGGTTCAATGCAGAT ACCTTCCAGAGATGATAGCACTTCATTGACTGCCAAAGAGGATGTCAGCATACCCAGA TCCACATTAGGAGATTTGGACACAGTTGCAGGGCTGGAAAAAGAACTGAGTAATGCCA AAGAGGAACTTGAACTCATGGCTAAAAAAGAAAGAGAATCACAGATGGAACTTTCTGC TCTACAGTCCATGATGGCTGTGCAGGAAGAAGAGCTGCAGGTGCAGGCTGCTGATATG GAGTCTCTGACCAGGAACATACAGATTAAAGAAGATCTCATAAAGGACCTGCAAATGC AACTGGTTGATCCTGAAGACATACCAGCTATGGAACGCCTGACCCAGGAAGTCTTACT TCTTCGGGAAAAAGTTGCTTCAGTAGAATCCCAGGGTCAAGAAATTTCAGGAAACCGA AGACAACAGCAGTTGCTGCTGATGCTAGAAGGACTAGTAGATGAACGGAGTCGGCTCA
ATGAGGCCTTACAAGCAGAGAGACAGCTCTATAGCAGTCTGGTGAAGTTCCATGCCCA TCCAGAGAGCTCTGAGAGAGACCGAACTCTGCAGGTGGAACTGGAAGGGGCTCAGGTG TTACGCAGTCGGCTAGAAGAAGTTCTTGGAAGAAGCTTGGAGCGCTTAAACAGGCTGG AGACCCTGGCCGCCATTGGAGGTGCAGCTGCAGGGGATGACACCGAAGATACAAGCAC TGAGTTCACTGACAGTATTGAGGAGGAGGCTGCACACCATAGTCACCAGCAACTTGTC AAGGTGGCTTTGGAGAAAAGTCTGGCAACTGTGGAGACCCAGAACCCATCTTTTTCCC CTCCTTCTCCGATGGGAGGGGACAGTAACAGGTGTCTTCAGGAAGAAATGCTCCACCT GAGGGCTGAGATCCACCAGCACTTAGAAGAGAAGAGGAAAGCTGAGGAGGAACTGAAG GAGCTAAAGGCTCAAATTGAGGAAGCAGGATTCTCCTCAGTGTCCCACATCAGGAACA CCATGCTGAGCCTTTGCCTTGAGAATGCGGAGCTGAAAGAGCAGATGGGAGAAACAAT GTCTGATGGATGGGAGATCGAGGAAGACAAGGAGAAGGGCGAGGTGATGGTTGAGACT GTGGTAACCAAAGAGGGTCTGAGTGAGAGTAGCCTTCAGGCTGAGTTCAGAAAGCTCC AGGGAAAACTGAAGAATGCCCACAATATCATCAACCTCCTCAAAGAACAACTTGTGCT GAGTAGCAAGGAAGGGAATAGTAAACTTACTCCAGAGCTCCTTGTGCATCTGACCAGC ACCATCGAAAGAATAAACACAGAACTGGTTGGTTCCCCTGGGAAGCACCAACACCAAG AGGAGGGGAATGTGACTGTGAGGCCTTTCCCCAGACCCCAGAGCCTTGACCTTGGGGC TACCTTCACAGTGGATGCCCACCAACAGTTGGATAACCAGTCCCAGCCTCGTGACCCT GGGCCTCAGCCAGCGTTTAGCCTACCAGGGTCCACCCAGCACCTGCGCTCCCAGCTGT CACAATGCAAACAACGCTATCAAGATCTCCAGGAGAAGCTGCTGCTATCAGAAGCCAC TGTCTTTGCTCAGGCTAACGAGCTGGAGAAATACAGAGTTATGCTTAGTGAATCCTTG GTGAAGCAGGACAGCAAGCAGATCCAGGTGGACTTCCAGGACCTGGGCTATGAGACTT GTGGCCGAAGCGAGAATGAGGCTGAACGGGAGGAAACCACCAGTCCTGAGTGTGAGGA GCACAACAGCCTCAAGGAAATGGTCCTGATGGAGGGGCTGTGCTCTGAGCAGGGACGC CGGGGCTCAACACTGGCTAGTTCCTCTGAGAGGAAGCCCTTGGAGAACCAGCTAGGGA AGCAGGAAGAGTTCCGGGTATATGGAAAGTCAGAAAACATCTTGGTCCTACGAAAGGA CATCGAAGATCTGAAGGCCCAGCTGCAGAATGCCAACAAGGTCATTCAAAACCTCAAG AGCCGGGTCCGGTCCCTCTCAGTTACAAGTGATTATTCGTCTAGTCTGGAAAGACCCC GGAAGCTGAGAGCTGTTGGCACCTTGGAGGGGTCTTCACCTCATAGTGTCCCTGATGA GGATGAGGGGTGGCTGTCTGATGGCACTGGGGCTTTCTACTCTCCAGGGCTTCAGGCC AAAAAGGACCTGGAGAGTCTCATCCAGAGAGTATCCCAGCTGGAGGCCCAGCTCCCAG AAAATGGACTAGAAGAGAAGCTGGCTGAGGAGCTGAGATCAGCCTCGTGGCCTGGGAA ATATGATTCCCTGATTCAGGATCAGGCCCGGGAACTGTCTTACCTACGGCAAAAAATA CGAGAAGGGAGAGGTATTTGTTATCTTATCACCCAGCATGCAAAAGATACAGTAAAAT CTTTTGAGGATCTCCTAAGGAGCAATGACATTGACTACTACCTGGGACAGAGCTTCCG GGAGCAACTCGCCCAGGGAAGCCAGCTGACAGAGAGGCTCACCAGCAAACTCAGCACA GAGGATCATAAAAGTGAGAAAGATCAAGCTGGACTTGAGCCACTGGCCCTCAGGCTCA GCAGGGAGCTGCAGGAGAAGGAGAAAGTGATTGAAGTCCTGCAGGCCAAGCTGGATGC TCGGTCCCTCACACCCTCCAGCAGCCGTGCCTTGTCTGACTCCCACCGCTCTCCCAGC AGCACCTCTTTCCTGTCTGATGAGCTGGAAGCCTGCTCTGACATGGACATAGTCAGCG AGTACACACACTATGAAGAGAAGAAAGCTTCTCCCAGTCACTCAGGTAGCAGTGCATC TCAGGGGGCTAAGGCCGAATCCAACAGCAACCCCATCAGCTTGCCAACTCCCCAGAAT ACCCCCAAGGAGGCCAACCAAGCCCATTCAGGCTTTCATTTTCACTCCATACCCAAGC TGGCTAGCCTTCCTCAGGCACCATTGCCCTCAGCTCCATCCAGCTTCCTGCCTTTCAG CCCCACTGGCCCTCCCCTCCTTGGCTGCTGTGAGACACCAGAGGTCTCCTTGGCTGAG TCTCAGCAGGAGCTACAGATGCTGCAGAAGCAGTTGGGAGAAAGTAGCACTGTTCCTC CTGCTTCCACAGCTACATTGCTGAGCAACGACTTGGAAGCCGACTCTTCCTACTACCT CAACTCTGCCCAGCCTCACTCTCCTCCAAGGGGCACCATAGAACTGGGAAGAATCCTA GAGCCTGGGTACCTGGGCAGCAGTGGCAAGTGGGATGTGATGAGGCCTCAGAAAGGGA GTGTATCTGGGGACCTATCCTCAGGCTCCTCTGTGTACCAGCTTAACTCCAAACCCAC AGGGGCTGACCTGCTGGAAGAGCATCTTGGTGAAATCTGGAACCTGCGCCAGCGCCTG GAGGAGTCCATCTGCATCAATGACTGCCTACGGGAGCAACTGGAACACCGGCTGACCT CTACTGCTCGTGGAAGGGGATCCACTTCTAACTTCTACAGTCAGGGCCTGGAGTCCAT ACCTCAGCTCTGCAATGAGAACAGAGTCCTCAGGGAAGAAAATCGAAGACTTCAGGCT CAACTGAGTCATGTTTCCAGAGGTCACTCCCAGGAAACAGAAAGCCTGAGGGAGGCTC TGCTGTCCTCTCGATCCCACCTTCAAGAGCTGGAAAAGGAGCTGGAGCACCAGAAGGT GGAAAGGCAGCAGCTTTTGGAAGACTTGAGGGAGAAGCAGCAAGAGGTCTTGCATTTC AGGGAGGAACGTCTTTCCCTCCAGGAAAACGACTCCAGACTGCAGCACAAGCTGGTTC TCCTGCAGCAACAGTGTGAAGAGAAACAGCAGCTCTTTGAGTCCCTCCAGTCAGAGCT ACAAATCTACGAGGCACTTTATGGCAATTCCAAGAAGGGGCTGAAAGCTTACAGCCTG GATGCCTGTCACCAAATCCCTTTGAGCAGTGACCTGAGCCACCTGGTGGCAGAGGTAC AAGCTCTGAGAGGGCAGCTGGAGCAGAGCATTCAGGGGAACAATTGTCTGCGACTGCA GCTGCAACAGCAGCTGGAGAGCGGTGCTGGCAAAGCCAGCCTCAGCCCCTCCTCCATT AACCAGAACTTCCCAGCCAGCACTGACCCTGGAAACAAGCAGCTGCTCCTCCAAGGTT CAGCTGTGTCCCCTCCAGTCCGGGATGTTGGTATGAATTCCCCAGCTCTGGTCTTCCC CAGCTCTGCTTCCTCTACTCCTGGCTCAGATTCAGTTGTGTTGTCATTTTCTTTTTCA GGCTTGGGTTTGGATACTTCTCCAGTAATGAAGACCCCTCCCAAGCTAGAGGGTGATG CTACTGATGGCTCCTTTGCCAATAAGCATGGCCGCCATGTCATTGGCCACATTGATGA CTACAGTGCCCTAAGACAGCAGATTGCGGAGGGCAAGCTGCTGGTCAAAAAGATAGTG TCTCTTGTGAGATCAGCGTGCAGCTTCCCTGGCCTTGAAGCCCAAGGCACAGAGGGCA GCAAAGGCATTCATGAGCTTCGGAGCAGCACCAGTGCCCTGCACCATGCCCTAGAGGA GTCGGCTTCCCTCCTCACCATGTTCTGGAGAGCGGCCCTGCCAAGCACCCACATCCCT GTGCTGCCTGGCAAACAGGGAGAATCAACAGAAAGGGAACTTCTGGAACTGAGAACCA AAGTATCCAAACAGGAGCAGCTCCTTCAGAGCACAACTGAGCATCTGAAGAACGCCAA CCAGCAGAAGGAGAGCATGGAACAGTTCATTGTCAGCGTAACCAGAACACATGATGTT TTAAAGAAGGCAAGGACTAACTTAGAGGTGAAATCCCTAAGGGCTCTGCCGTGTACTC CAGCCTTGTGACCCTTGCCTTCCAGGAACCATGCAAGAAGCGCAGCCACCAGAAGTCC TTAAAACAGCAGGAAAGGTGAGCCTGTCCCCCTTTTGTGCAGCTACCTATCTGCTGAG
GAGCATCTGGGCCTCATTCCTCCAAGT
ORF Start: ATG at 46 ORF Stop: TGA at 7027
SEQ TD NO: 294 2327 aa MW at 263034.6kD
NOVl 03b, MSNGYRTLSQHLNDLKKENFSLKLRIYFLEERMQQKYEASREDIYKRNIELKVEVESL KRELQDKKQHLDKTWADVENLNSQNEAELRRQFEERQQETEHVYELLENKIQLLQEES
CG59773-02 Protein Sequence RLAKNEAAR AALVEAEKECNLELSEKLKGVTKNWEDVPGDQVKPDQYTEALAQRDRR IEELNQSLAAQERLVEQLSREKQQLLHLLEEPTSMEVQPMTEELLKQQKLNSHETTIT QQSVSDSHLAELQEKIQQTEATNKILQEKLNEMSYELKCAQESSQKQDGTIQNLKETL KSRERETEELYQVIEGQNDTMAKLREMLHQSQLGQLQSSEGTSPAQQQVALLDLQSAL FCSQLEIQKLQRWRQKERQLADAKQCVQFVEAAAHESEQQKEAS KHNQELRKALQQ LQEELQNKSQQLRAWEAEKYNEIRTQEQNIQHLNHSLSHKEQLLQEFRELLQYRDNSD KTLEANEMLLEKLRQRIHDKAVALERAIDEKFSALEEKEKELRQLRLAVRERDHDLER LRDVLSSNEATMQSMESLLRAKGLEVEQLSTTCQNLQWLKEEMETKFSRWQKEQESII QQLQTSLHDRNKEVEDLSATLLCKLGPGQSEIAEELCQRLQRKERMLQDLLSDRNKQV LEHEMEIQGLLQSVSTREQESQAAAEKLVQALMERNSELQALRQYLGGRDSLMSQAPI SNQQAEVTPTGRLGKQTDQGSMQIPSRDDSTSLTAKEDVSI RSTLGDLDTVAGLEKE LSNAKEELEL AKKERESQMELSALQSMMAVQEEELQVQAADMESLTRNIQIKEDLIK DLQMQLVDPEDIPAMERLTQEVLLLREKVASVESQGQEISGNRRQQQLLLMLEGLVDE RSRLNEALQAERQLYSSLVKFHAHPESSERDRTLQVELEGAQVLRSRLEEVLGRSLER LNRLETLAAIGGAAAGDDTEDTSTEFTDSIEEEAAHHSHQQLVKVALEKSLATVETQN PSFSPPSPMGGDSNRCLQEE LHLRAEIHQHLEEKRKAEEELKELKAQIEEAGFSSVS HIRNTMLSLCLENAELKEQMGETMSDGWEIEEDKEKGEVMVETWTKEGLSESSLQAE FRKLQGKLKNAHNIINLLKEQLVLSSKEGNSKLTPELLVHLTSTIERINTELVGSPGK HQHQEEGNVTVRPFPRPQSLDLGATFTVDAHQQLDNQSQPRDPGPQPAFSLPGSTQHL RSQLSQCKQRYQDLQEKLLLSEATVFAQANELEKYRVMLSESLVKQDSKQIQVDFQDL GYETCGRSENEAEREETTSPECEEHNSLKEMVLMEGLCSEQGRRGSTLASSSERKPLE NQLGKQEEFRVYGKSENILVLRKDIEDLKAQLQNANKVIQNLKSRVRSLSVTSDYSSS LERPRKLRAVGTLEGSSPHSVPDEDEGWLSDGTGAFYSPGLQAKKDLESLIQRVSQLE AQLPENGLEEKLAEELRSASWPGKYDSLIQDQARELSYLRQKIREGRGICYLITQHAK DTVKSFEDLLRSNDIDYYLGQSFREQLAQGSQLTERLTSKLSTEDHKSEKDQAGLEPL ALRLSRELQEKEKVIEVLQAKLDARSLTPSSSRALSDSHRSPSSTSFLSDELEACSDM DIVSEYTHYEEKKASPSHSGSSASQGAKAESNSNPISLPTPQNTPKEANQAHSGFHFH SIPKLASLPQAPLPSAPSSFLPFSPTGPPLLGCCETPEVSLAESQQELQMLQKQLGES STVPPASTATLLSNDLEADSSYYLNSAQPHSPPRGTIELGRILEPGYLGSSGK DVMR PQKGSVSGDLSSGSSVYQLNSKPTGADLLEEHLGEIWNLRQRLEESICINDCLREQLE HRLTSTARGRGSTSNFYSQGLESIPQLCNENRVLREENRRLQAQLSHVSRGHSQETES LREALLSSRSHLQELEKELEHQKVERQQLLEDLREKQQEVLHFREERLSLQENDSRLQ HKLVLLQQQCEEKQQLFESLQSELQIYEALYGNSKKGLKAYSLDACHQIPLSSDLSHL VAEVQALRGQLEQSIQGNNCLRLQLQQQLESGAGKASLSPSSINQNFPASTDPGNKQL LLQGSAVSPPVRDVGMNSPALVFPSSASSTPGSDSWLSFSFSGLGLDTSPVMKTPPK LEGDATDGSFANKHGRHVIGHIDDYSALRQQIAEGKLLVKKIVSLVRSACSFPGLEAQ GTEGSKGIHELRSSTSALHHALEESASLLTMF RAALPSTHIPVLPGKQGESTERELL ELRTKVSKQEQLLQSTTEHLKNANQQKESMEQFIVSVTRTHDVLKKARTNLEVKSLRA LPCTPAL
SEQ ED NO: 295 7084 bp
NOVl 03c, GTTGAGGGGGCAATCGGGCACGCTCCTCCCCATGGGTTGCCCATCATGTCTAATGGAT
ATCGCACTCTGTCCCAGCACCTCAATGACCTGAAGAAGGAGAACTTCAGCCTCAAGCT
CG59773-03 DNA Sequence GCTCATCTACTTCCTGGAGGAGCGCATGCAACAGAAGTATGAGGCCAGCCGGGAGGAC
ATCTACAAGCGGGGGTGATGTGGAGAATCTCAACAGTCAGAATGAAGCTGAGCTCCGA CGCCAGTTTGAGGAGCGACAGCAGGAGACGGAGCATGTTTATGAGCTCTTGGAGAATA AGATCCAGCTTCTGCAGGAGGAATCCAGGCTAGCAAAGAATGAAGCTGCGCGGATGGC AGCTCTGGTGGAAGCAGAGAAGGAGTGTAACCTGGAGCTCTCAGAGAAACTGAAGGGA GTCACCAAAAACTGGGAAGATGTACCAGGAGACCAGGTCAAGCCCGACCAATACACTG AGACCCTGGCCCAGAGGGACAAGAGAATTGAAGAACTGAATCAGAGCCTGGCTGCCCA GGAGAGGCTTGTAGAACAGCTATCTCGGGAGAAACAACAACTGCTACATCTGTTGGAG GAGCCAACTAGCATGGAAGTGCAGCCCATGACTGAAGAGTTGCTGAAACAACAAAAGC TGAATTCACATGAGACCACTATAACTCAGCAGTCTGTATCTGATTCCCACTTGGCAGA ACTCCAGGAAAAAATCCAGCAAACAGAGGCCACCAACAAGATTCTTCAAGAGAAACTT AATGAAATGAGCTATGAACTAAAGTGTGCTCAGGAGTCGTCTCAAAAGCAAGATGGTA CAATTCAGAACCTCAAGGAAACTCTGAAAAGCAGGGAACGTGAGACTGAGGAGTTGTA CCAGGTAATTGAAGGTCAAAATGACACAATGGCAAAGCTTCGAGAAATGCTGCACCAA AGCCAGCTTGGACAACTTCAGAGCTCAGAGGGTACTTCTCCAGCTCAGCAACAGGTAG CTCTGCTTGATCTTCAGAGTGCTTTATTCTGCAGCCAACTTGAAATACAGAAGCTCCA GAGGGTGGTACGACAGAAAGAGCGCCAACTGGCTGATGCCAAACAATGTGTGCAATTT GTAGAGGCTGCAGCACACGAGAGTGAACAGCAGAAAGAGGCTTCTTGGAAACATAACC AGGAATTGCGAAAAGCCTTGCAGCAGCTACAAGAAGAATTGCAGAATAAGAGCCAACA GCTTCGTGCCTGGGAGGCTGAAAAATACAATGAGATTCGAACCCAGGAACAAAACATC CAGCACCTAAACCATAGTCTGAGTCACAAGGAGCAGTTGCTTCAGGAATTTCGGGAGC TCCTACAGTATCGAGATAACTCAGACAAAACCCTTGAAGCAAATGAAATGTTGCTTGA GAAACTTCGCCAGCGAATACATGATAAAGCTGTTGCTCTGGAGCGGGCTATAGATGAA AAATTCTCTGCTCTAGAAGAGAAAGAAAAAGAACTGCGCCAGCTTCGTCTTGCTGTGA GAGAGCGAGATCATGACTTAGAGAGACTGCGCGATGTCCTCTCCTCCAATGAAGCTAC TATGCAAAGTATGGAGAGTCTCCTGAGGGCCAAAGGCCTGGAAGTGGAACAGTTATCT
ACTACCTGTCAAAACCTCCAGTGGCTGAAAGAAGAAATGGAAACCAAATTTAGCCGTT GGCAGAAGGAACAAGAGAGTATCATTCAGCAGTTACAGACGTCTCTTCATGATAGGAA CAAAGAAGTGGAGGATCTTAGTGCAACACTGCTCTGCAAACTTGGACCAGGGCAGAGT GAGATAGCAGAGGAGCTGTGCCAGCGTCTACAGCGAAAGGAAAGGATGCTGCAGGACC TTCTAAGTGATCGAAATAAACAAGTGCTGGAACATGAAATGGAGATTCAAGGCCTGCT TCAGTCTGTGAGCACCAGGGAGCAGGAAAGCCAAGCTGCTGCAGAGAAGTTGGTGCAA GCCTTAATGGAAAGAAATTCAGAATTACAGGCCCTGCGCCAATATTTAGGAGGGAGAG ACTCCCTGATGTCCCAAGCACCCATCTCTAACCAACAAGCTGAAGTTACCCCCACTGG CCGTCTTGGAAAACAGACTGATCAAGGTTCAATGCAGATACCTTCCAGAGATGATAGC ACTTCATTGACTGCCAAAGAGGATGTCAGCATACCCAGATCCACATTAGGAGATTTGG ACACAGTTGCAGGGCTGGAAAAAGAACTGAGTAATGCCAAAGAGGAACTTGAACTCAT GGCTAAAAAAGAAAGAGAATCACAGATGGAACTTTCTGCTCTACAGTCCATGATGGCT GTGCAGGAAGAAGAGCTGCAGGTGCAGGCTGCTGATATGGAGTCTCTGACCAGGAACA TACAGATTAAAGAAGATCTCATAAAGGACCTGCAAATGCAACTGGTTGATCCTGAAGA CATACCAGCTATGGAACGCCTGACCCAGGAAGTCTTACTTCTTCGGGAAAAAGTTGCT TCAGTAGAATCCCAGGGTCAAGAAATTTCAGGAAACCGAAGACAACAGCAGTTGCTGC TGATGCTAGAAGGACTAGTAGATGAACGGAGTCGGCTCAATGAGGCCTTACAAGCAGA GAGACAGCTCTATAGCAGTCTGGTGAAGTTCCATGCCCATCCAGAGAGCTCTGAGAGA GACCGAACTCTGCAGGTGGAACTGGAAGGGGCTCAGGTGTTACGCAGTCGGCTAGAAG AAGTTCTTGGAAGAAGCTTGGAGCGCTTAAACAGGCTGGAGACCCTGGCCGCCATTGG AGGTGCAGCTGCAGGGGATGACACCGAAGATACAAGCACTGAGTTCACTGACAGTATT GAGGAGGAGGCTGCACACCATAGTCACCAGCAACTTGTCAAGGTGGCTTTGGAGAAAA GTCTGGCAACTGTGGAGACCCAGAACCCATCTTTTTCCCCTCCTTCTCCGATGGGAGG GGACAGTAACAGGTGTCTTCAGGAAGAAATGCTCCACCTGAGGGCTGAGATCCACCAG CACTTAGAAGAGAAGAGGAAAGCTGAGGAGGAACTGAAGGAGCTAAAGGCTCAAATTG AGGAAGCAGGATTCTCCTCAGTGTCCCACATCAGGAACACCATGCTGAGCCTTTGCCT TGAGAATGCGGAGCTGAAAGAGCAGATGGGAGAAACAATGTCTGATGGATGGGAGATC GAGGAAGACAAGGAGAAGGGCGAGGTGATGGTTGAGACTGTGGTAACCAAAGAGGGTC TGAGTGAGAGTAGCCTTCAGGCTGAGTTCAGAAAGCTCCAGGGAAAACTGAAGAATGC CCACAATATCATCAACCTCCTCAAAGAACAACTTGTGCTGAGTAGCAAGGAAGGGAAT AGTAAACTTACTCCAGAGCTCCTTGTGCATCTGACCAGCACCATCGAAAGAATAAACA CAGAACTGGTTGGTTCCCCTGGGAAGCACCAACACCAAGAGGAGGGGAATGTGACTGT GAGGCCTTTCCCCAGACCCCAGAGCCTTGACCTTGGGGCTACCTTCACAGTGGATGCC CACCAACAGTTGGATAACCAGTCCCAGCCTCGTGACCCTGGGCCTCAGCCAGCGTTTA GCCTACCAGGGTCCACCCAGCACCTGCGCTCCCAGCTGTCACAATGCAAACAACGCTA TCAAGATCTCCAGGAGAAGCTGCTGCTATCAGAAGCCACTGTCTTTGCTCAGGCTAAC GAGCTGGAGAAATACAGAGTTATGCTTAGTGAATCCTTGGTGAAGCAGGACAGCAAGC AGATCCAGGTGGACTTCCAGGACCTGGGCTATGAGACTTGTGGCCGAAGCGAGAATGA GGCTGAACGGGAGGAAACCACCAGTCCTGAGTGTGAGGAGCACAACAGCCTCAAGGAA ATGGTCCTGATGGAGGGGCTGTGCTCTGAGCAGGGACGCCGGGGCTCAACACTGGCTA GTTCCTCTGAGAGGAAGCCCTTGGAGAACCAGCTAGGGAAGCAGGAAGAGTTCCGGGT ATATGGAAAGTCAGAAAACATCTTGGTCCTACGAAAGGACATCGAAGATCTGAAGGCC CAGCTGCAGAATGCCAACAAGGTCATTCAAAACCTCAAGAGCCGGGTCCGGTCCCTCT CAGTTACAAGTGATTATTCGTCTAGTCTGGAAAGACCCCGGAAGCTGAGAGCTGTTGG CACCTTGGAGGGGTCTTCACCTCATAGTGTCCCTGATGAGGATGAGGGGTGGCTGTCT GATGGCACTGGGGCTTTCTACTCTCCAGGGCTTCAGGCCAAAAAGGACCTGGAGAGTC TCATCCAGAGAGTATCCCAGCTGGAGGCCCAGCTCCCAGAAAATGGACTAGAAGAGAA GCTGGCTGAGGAGCTGAGATCAGCCTCGTGGCCTGGGAAATATGATTCCCTGATTCAG GATCAGGCCCGGGAACTGTCTTACCTACGGCAAAAAATACGAGAAGGGAGAGGTATTT GTTATCTTATCACCCAGCATGCAAAAGATACAGTAAAATCTTTTGAGGATCTCCTAAG GAGCAATGACATTGACTACTACCTGGGACAGAGCTTCCGGGAGCAACTCGCCCAGGGA AGCCAGCTGACAGAGAGGCTCACCAGCAAACTCAGCACAGAGGATCATAAAAGTGAGA AAGATCAAGCTGGACTTGAGCCACTGGCCCTCAGGCTCAGCAGGGAGCTGCAGGAGAA GGAGAAAGTGATTGAAGTCCTGCAGGCCAAGCTGGATGCTCGGTCCCTCACACCCTCC AGCAGCCGTGCCTTGTCTGACTCCCACCGCTCTCCCAGCAGCACCTCTTTCCTGTCTG ATGAGCTGGAAGCCTGCTCTGACATGGACATAGTCAGCGAGTACACACACTATGAAGA GAAGAAAGCTTCTCCCAGTCACTCAGGTAGCAGTGCATCTCAGGGGGCTAAGGCCGAA TCCAACAGCAACCCCATCAGCTTGCCAACTCCCCAGAATACCCCCAAGGAGGCCAACC AAGCCCATTCAGGCTTTCATTTTCACTCCATACCCAAGCTGGCTAGCCTTCCTCAGGC ACCATTGCCCTCAGCTCCATCCAGCTTCCTGCCTTTCAGCCCCACTGGCCCTCCCCTC CTTGGCTGCTGTGAGACACCAGAGGTCTCCTTGGCTGAGTCTCAGCAGGAGCTACAGA TGCTGCAGAAGCAGTTGGGAGAAAGTAGCACTGTTCCTCCTGCTTCCACAGCTACATT GCTGAGCAACGACTTGGAAGCCGACTCTTCCTACTACCTCAACTCTGCCCAGCCTCAC TCTCCTCCAAGGGGCACCATAGAACTGGGAAGAATCCTAGAGCCTGGGTACCTGGGCA GCAGTGGCAAGTGGGATGTGATGAGGCCTCAGAAAGGGAGTGTATCTGGGGACCTATC CTCAGGCTCCTCTGTGTACCAGCTTAACTCCAAACCCACAGGGGCTGACCTGCTGGAA GAGCATCTTGGTGAAATCTGGAACCTGCGCCAGCGCCTGGAGGAGTCCATCTGCATCA ATGACTGCCTACGGGAGCAACTGGAACACCGGCTGACCTCTACTGCTCGTGGAAGGGG ATCCACTTCTAACTTCTACAGTCAGGGCCTGGAGTCCATACCTCAGCTCTGCAATGAG AACAGAGTCCTCAGGGAAGAAAATCGAAGACTTCAGGCTCAACTGAGTCATGTTTCCA GAGGTCACTCCCAGGAAACAGAAAGCCTGAGGGAGGCTCTGCTGTCCTCTCGATCCCA CCTTCAAGAGCTGGAAAAGGAGCTGGAGCACCAGAAGGTGGAAAGGCAGCAGCTTTTG GAAGACTTGAGGGAGAAGCAGCAAGAGGTCTTGCATTTCAGGGAGGAACGTCTTTCCC TCCAGGAAAACGACTCCAGACTGCAGCACAAGCTGGTTCTCCTGCAGCAACAGTGTGA AGAGAAACAGCAGCTCTTTGAGTCCCTCCAGTCAGAGCTACAAATCTACGAGGCACTT TATGGCAATTCCAAGAAGGGGCTGAAAGCTTACAGCCTGGATGCCTGTCACCAAATCC CTTTGAGCAGTGACCTGAGCCACCTGGTGGCAGAGGTACAAGCTCTGAGAGGGCAGCT
GGAGCAGAGCATTCAGGGGAACAATTGTCTGCGACTGCAGCTGCAACAGCAGCTGGAG AGCGGTGCTGGCAAAGCCAGCCTCAGCCCCTCCTCCATTAACCAGAACTTCCCAGCCA GCACTGACCCTGGAAACAAGCAGCTGCTCCTCCAAGGTTCAGCTGTGTCCCCTCCAGT CCGGGATGTTGGTATGAATTCCCCAGCTCTGGTCTTCCCCAGCTCTGCTTCCTCTACT CCTGGCTCAGATTCAGTTGTGTTGTCATTTTCTTTTTCAGGCTTGGGTTTGGATACTT CTCCAGTAATGAAGACCCCTCCCAAGCTAGAGGGTGATGCTACTGATGGCTCCTTTGC CAATAAGCATGGCCGCCATGTCATTGGCCACATTGATGACTACAGTGCCCTAAGACAG CAGATTGCGGAGGGCAAGCTGCTGGTCAAAAAGATAGTGTCTCTTGTGAGATCAGCGT GCAGCTTCCCTGGCCTTGAAGCCCAAGGCACAGAGGGCAGCAAAGGCATTCATGAGCT TCGGAGCAGCACCAGTGCCCTGCACCATGCCCTAGAGGAGTCGGCTTCCCTCCTCACC ATGTTCTGGAGAGCGGCCCTGCCAAGCACCCACATCCCTGTGCTGCCTGGCAAACAGG GAGAATCAACAGAAAGGGAACTTCTGGAACTGAGAACCAAAGTATCCAAACAGGAGCA GCTCCTTCAGAGCACAACTGAGCATCTGAAGAACGCCAACCAGCAGAAGGAGAGCATG GAACAGTTCATTGTCAGCGTAACCAGAACACATGATGTTTTAAAGAAGGCAAGGACTA ACTTAGAGGTGAAATCCCTAAGGGCTCTGCCGTGTACTCCAGCCTTGTGACCCTTGCC TTCCAGGAACCATGCAAGAAGCGCAGCCACCAGAAGTCCTTAAAACAGCAGGAAAGGT
GAGCCTGTCCCCCTTTTGTGCAGCTACCTATCTGCTGAGGAGCATCTGGGCCTCATTC
CTCCAAGT
ORF Start: ATG at 155 ORF Stop: TGA at 6950
SEQ ID NO: 296 2265 aa MW at 255081.5kD
NOV103c, MRPAGRTSTSGGDVENLNSQNEAELRRQFEERQQETEHVYELLENKIQLLQEESRLAK NEAARMAALVEAEKECNLELSEKLKGVTKN EDVPGDQVKPDQYTETLAQRDKRIEEL
CG59773-03 Protein Sequence NQSLAAQERLVEQLSREKQQLLHLLEEPTSMEVQPMTEELLKQQKLNSHETTITQQSV SDSHLAELQEKIQQTEATNKILQEKLNEMSYELKCAQESSQKQDGTIQNLKETLKSRE RETEELYQVIEGQNDTMAKLREMLHQSQLGQLQSSEGTSPAQQQVALLDLQSALFCSQ LEIQKLQRWRQKERQLADAKQCVQFVEAAAHESEQQKEAS KHNQELRKALQQLQEE LQNKSQQLRA EAEKYNEIRTQEQNIQHLNHSLSHKEQLLQEFRELLQYRDNSDKTLE ANEMLLEKLRQRIHDKAVALERAIDEKFSALEEKEKELRQLRLAVRERDHDLERLRDV LSSNEATMQSMESLLRAKGLEVEQLSTTCQNLQ LKEEMETKFSR QKEQESIIQQLQ TSLHDRNKEVEDLSATLLCKLGPGQSEIAEELCQRLQRKERMLQDLLSDRNKQVLEHE MEIQGLLQSVSTREQESQAAAEKLVQALMERNSELQALRQYLGGRDSLMSQAPISNQQ AEVTPTGRLGKQTDQGS QIPSRDDSTSLTAKEDVSIPRSTLGDLDTVAGLEKELSNA KEELELMAKKERESQMELSALQS MAVQEEELQVQAADMESLTRNIQIKEDLIKDLQM QLVDPEDIPAMERLTQEVLLLREKVASVESQGQEISGNRRQQQLLLMLEGLVDERSRL NEALQAERQLYSSLVKFHAHPESSERDRTLQVELEGAQVLRSRLEEVLGRSLERLNRL ETLAAIGGAAAGDDTEDTSTEFTDSIEEEAAHHSHQQLVKVALEKSLATVETQNPSFS PPSPMGGDSNRCLQEEMLHLRAEIHQHLEEKRKAEEELKELKAQIEEAGFSSVSHIRN TMLSLCLENAELKEQMGETMSDG EIEEDKEKGEVMVETWTKEGLSESSLQAEFRKL QGKLKNAHNIINLLKEQLVLSSKEGNSKLTPELLVHLTSTIERINTELVGSPGKHQHQ EEGNVTVRPFPRPQSLDLGATFTVDAHQQLDNQSQPRDPGPQPAFSLPGSTQHLRSQL SQCKQRYQDLQEKLLLSEATVFAQANELEKYRVMLSESLVKQDSKQIQVDFQDLGYET CGRSENEAEREETTSPECEEHNSLKEMVLMEGLCSEQGRRGSTLASSSERKPLENQLG KQEEFRVYGKSENILVLRKDIEDLKAQLQNANKVIQNLKSRVRSLSVTSDYSSSLERP RKLRAVGTLEGSSPHSVPDEDEGWLSDGTGAFYSPGLQAKKDLESLIQRVSQLEAQLP ENGLEEKLAEELRSAS PGKYDSLIQDQARELSYLRQKIREGRGICYLITQHAKDTVK SFEDLLRSNDIDYYLGQSFREQLAQGSQLTERLTSKLSTEDHKSEKDQAGLEPLALRL SRELQEKEKVIEVLQAKLDARSLTPSSSRALSDSHRSPSSTSFLSDELEACSDMDIVS EYTHYEEKKASPSHSGSSASQGAKAESNSNPISLPTPQNTPKEANQAHSGFHFHSIPK LASLPQAPLPSAPSSFLPFSPTGPPLLGCCETPEVSLAESQQELQ LQKQLGESSTVP PASTATLLSNDLEADSSYYLNSAQPHSPPRGTIELGRILEPGYLGSSGK DVMRPQKG SVSGDLSSGSSVYQLNSKPTGADLLEEHLGEI NLRQRLEESICINDCLREQLEHRLT STARGRGSTSNFYSQGLESIPQLCNENRVLREENRRLQAQLSHVSRGHSQETESLREA LLSSRSHLQELEKELEHQKVERQQLLEDLREKQQEVLHFREERLSLQENDSRLQHKLV LLQQQCEEKQQLFESLQSELQIYEALYGNSKKGLKAYSLDACHQIPLSSDLSHLVAEV QALRGQLEQSIQGNNCLRLQLQQQLESGAGKASLSPSSINQNFPASTDPGNKQLLLQG SAVSPPVRDVGMNSPALVFPSSASSTPGSDSWLSFSFSGLGLDTSPVMKTPPKLEGD ATDGSFANKHGRHVIGHIDDYSALRQQIAEGKLLVKKIVSLVRSACSFPGLEAQGTEG SKGIHELRSSTSALHHALEESASLLTMF RAALPSTHIPVLPGKQGESTERELLELRT KVSKQEQLLQSTTEHLKNANQQKESMEQFIVSVTRTHDVLKKARTNLEVKSLRALPCT PAL
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 103B.
Table 103B. Comparison of NOVl 03a against NOV103b through NOV103c.
NOVl 03a Residues/ Identities/
Protein Sequence Match Residues Similarities for the Matched Region
Further analysis of the NOVl 03 a protein yielded the following properties shown in Table 103C.
Table 103C. Protein Sequence Properties NOV103a
PSort 0.5855 probability located in mitochondrial matrix space; 0.4200 probability analysis: located in nucleus; 0.3000 probability located in microbody (peroxisome); 0.2957 probability located in mitochondrial inner membrane
SignalP Likely cleavage site between residues 39 and 40 analysis:
A search of the NOVl 03 a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 103D.
In a BLAST search of public sequence databases, the NOV103a protein was found to have homology to the proteins shown in the BLASTP data in Table 103E.
PFam analysis predicts that the NOV103a protein contains the domains shown in the Table 103F.
Example 104.
The NOVl 04 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 104A.
Further analysis of the NOVl 04a protein yielded the following properties shown in Table 104B.
Table 104B. Protein Sequence Properties NOV104a
PSort 0.6400 probability located in plasma membrane; 0.4600 probability located in analysis: Golgi body; 0.3700 probability located in endoplasmic reticulum (membrane); 0.1000 probability located in endoplasmic reticulum (lumen)
SignalP Likely cleavage site between residues 64 and 65 analysis:
A search of the NOVl 04a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 104C.
In a BLAST search of public sequence databases, the NOVl 04a protein was found to have homology to the proteins shown in the BLASTP data in Table 104D.
Table 104D. Public BLASTP Results for NOVl 04a
Protein/Organism/Length
PFam analysis predicts that the NOVl 04a protein contains the domains shown in the Table 104E.
Example 105.
The NOVl 05 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 105 A.
PFam analysis predicts that the NOVl 05a protein contains the domains shown in the Table 105E.
Table 105E. Domain Analysis of NOVl 05a
Identities/
Pfam Domain NOVl 05a Match Region Similarities Expect Value for the Matched Region
No Significant Matches Found
Example 106.
The NOVl 06 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 106 A.
411
Further analysis of the NOV105a protein yielded the following properties shown in Table 105B.
Table 105B. Protein Sequence Properties NOVl 05a
PSort 0.6760 probability located in plasma membrane; 0.1000 probability located in analysis: endoplasmic reticulum (membrane); 0.1000 probability located in endoplasmic reticulum (lumen); 0.1000 probability located in outside
SignalP Likely cleavage site between residues 29 and 30 analysis:
A search of the NOV105a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 105C.
In a BLAST search of public sequence databases, the NOVl 05a protein was found to have homology to the proteins shown in the BLASTP data in Table 105D.
410
Further analysis of the NOVl 06a protein yielded the following properties shown in Table 106B.
Table 106B. Protein Sequence Properties NOV106a
PSort 0.6400 probability located in microbody (peroxisome); 0.4500 probability analysis: located in cytoplasm; 0.3122 probability located in lysosome (lumen); 0.1000 probability located in mitochondrial matrix space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVl 06a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 106C.
In a BLAST search of public sequence databases, the NOVl 06a protein was found to have homology to the proteins shown in the BLASTP data in Table 106D.
PFam analysis predicts that the NOVl 06a protein contains the domains shown in the Table 106E.
Example 107.
The NOVl 07 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 107 A.
Table 107A. NOVl 07 Sequence Analysis
SEQ ED NO: 303 4091 bp
NOV107a, AAGCAAGAGGCTGAGATGGATCTTGAGGCGGCAAAGAACGGAACAGCCTGGCGCCCCA
CGAGCGCGGAGGGCGACTTTGAACTGGGCATCAGCAGCAAACAAAAAAGGAAAAAAAC
CG57468-01 DNA Sequence GAAGACAGTGAAAATGATTGGAGTATTAACATTGTTTCGATACTCCGATTGGCAGGAT AAATTGTTTATGTCGCTGGGTACCATCATGGCCATAGCTCACGGATCAGGTCTCCCCC TCATGATGATAGTATTTGGAGAGATGACTGACAAATTTGTTGATACTGCAGGAAACTT CTCCTTTCCAGTGAACTTTTCCTTGTCGCTGCTAAATCCAGGCAAAATTCTGGAAGAA GAAATGACTAGATATGCATATTACTACTCAGGATTGGGTGCTGGAGTTCTTGTTGCTG CCTATATACAAGTTTCATTTTGGACTTTGGCAGCTGGTCGACAGATCAGGAAAATTAG GCAGAAGTTTTTTCATGCTATTCTACGACAGGAAATAGGATGGTTTGACATCAATGAC ACCACTGAACTCAATACGCGGCTAACAGATGACATCTCCAAAATCAGTGAAGGAATTG GTGACAAGGTTGGAATGTTCTTTCAAGCAGTAGCCACGTTTTTTGCAGGATTCATAGT GGGATTCATCAGAGGATGGAAGCTCACCCTTGTGATAATGGCCATCAGCCCTATTCTA GGACTCTCTGCAGCCGTTTGGGCAAAGATACTCTCGGCATTTAGTGACAAAGAACTAG CTGCTTATGCAAAAGCAGGCGCCGTGGCAGAAGAGGCTCTGGGGGCCATCAGGACTGT GATAGCTTTCGGGGGCCAGAACAAAGAGCTGGAAAGGTATCAGAAACATTTAGAAAAT GCCAAAGAGATTGGAATTAAAAAAGCTATTTCAGCAAACATTTCCATGGGTATTGCCT TCCTGTTAATATATGCATCATATGCACTGGCCTTCTGGTATGGATCCACTCTAGTCAT ATCAAAAGAATATACTATTGGAAATGCAATGACAGTTTTTTTTTCAATCCTAATTGGA GCTATGGCCATCGGAGAAACGCTCGTTTTGGCTCCTGAATATTCCAAAGCCAAATCGG GGGCTGCGCATCTGTTTGCCTTGTTGGAAAAGAAACCAAATATAGACAGCCGCAGTCA AGAAGGGAAAAAGCCAGTAAGCGACACATGTGAAGGGAATTTAGAGTTTCGAGAAGTC TCTTTCTTCTATCCATGTCGCCCAGATGTTTTCATCCTCCGTGGCTTATCCCTCAGTA TTGAGCGAGGAAAGACAGTAGCATTTGTGGGGAGCAGCGGCTGTGGGAAAAGCACTTC TGTTCAACTTCTGCAGAGACTTTATGACCCCGTGCAAGGACAAGTGGATGGTGTGGAT GCAAAAGAATTGAATGTACAGTGGCTCCGTTCCCAAATAGCAATCGTTCCTCAAGAGC CTGTGCTCTTCAACTGCAGCATTGCTGAGAACATCGCCTATGGTGACAACAGCCGTGT GGTGCCATTAGATGAGATCAAAGAAGCCGCAAATGCAGCAAATATCCATTCTTTTATT GAAGGTCTCCCTGAGAAATACAACACACAAGTTGGACTGAAAGGAGCACAGCTTTCTG GCGGCCAGAAACAAAGACTAGCTATTGCAAGGGCTCTTCTCCAAAAACCCAAAATTTT ATTGTTGGATGAGGCCACTTCAGCCCTCGATAATGACAGTGAGTGGCAGGTGGTTCAG CATGCCCTTGATAAAGCCAGGACGGGAAGGACATGCCTAGTGGTCACTCACAGGCTCT CTGCAATTCAGAACGCAGATTTGATAGTGGTTCTGCACAATGGAAAGATAAAGGAACA AGGAACTCATCAAGAGCTCCTGAGAAATCGAGACATATATTTTAAGTTAGTGAATGCA CAGTCAGCGAGCAAAGGTCGGACTACAATCGTGGTAGCACACCGACTTTCTACTATTC GAAGTGCAGATTTGATTGTGACCCTAAAGGATGGAATGCTGGCGGAGAAAGGAGCACA TGCTGAACTAATGGCAAAACGAGGTCTATATTATTCACTTGTGATGTCACAGGTAATG CTTATGGGGACTCTTTCAGACTGTGGTAATAGTCTTCCTGAAGTCTCTCTATTAAAAA TTTTAAAGTTAAACAAGCCTGAATGGCCTTTTGTGGTTCTGGGGACATTGGCTTCTGT TCTAAATGGAACTGTTCATCCAGTATTTTCCATCATCTTTGCAAAAATTATAACCGTA ATGTTTGGAAATAATGATCTTTTGTTTTTCCTCAAAATTTTTTTATATTCATTCCTTT
Further analysis of the NOVl 07a protein yielded the following properties shown in Table 107B.
Table 107B. Protein Sequence Properties NOVl 07a
PSort 0.6000 probability located in plasma membrane; 0.4000 probability located in analysis: Golgi body; 0.3000 probability located in endoplasmic reticulum (membrane); 0.3000 probability located in microbody (peroxisome)
A search of the NOVl 07a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 107C.
In a BLAST search of public sequence databases, the NOVl 07a protein was found to have homology to the proteins shown in the BLASTP data in Table 107D.
PFam analysis predicts that the NOVl 07a protein contains the domains shown in the Table 107E.
Example 108.
The NOVl 08 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 108 A.
Further analysis of the NOVl 08a protein yielded the following properties shown in Table 108B.
Table 108B. Protein Sequence Properties NOV108a
PSort 0.6400 probability located in microbody (peroxisome); 0.6000 probability analysis: located in plasma membrane; 0.4500 probability located in cytoplasm; 0.1000 probability located in mitochondrial matrix space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVl 08a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 108C.
In a BLAST search of public sequence databases, the NOVl 08a protein was found to have homology to the proteins shown in the BLASTP data in Table 108D.
PFam analysis predicts that the NOVl 08a protein contains the domains shown in the Table 108E.
Example 109.
The NOVl 09 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 109A.
Further analysis of the NOVl 09a protein yielded the following properties shown in Table 109B.
Table 109B. Protein Sequence Properties NOVl 09a
PSort 0.6500 probability located in cytoplasm; 0.1000 probability located in analysis: mitochondrial matrix space; 0.1000 probability located in lysosome (lumen); 0.0000 probability located in endoplasmic reticulum (membrane)
SignalP Likely cleavage site between residues 19 and 20 analysis:
A search of the NOVl 09a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 109C.
In a BLAST search of public sequence databases, the NOVl 09a protein was found to have homology to the proteins shown in the BLASTP data in Table 109D.
PFam analysis predicts that the NOVl 09a protein contains the domains shown in the Table 109E.
Example 110.
The NOVl 10 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 110A.
Table 110A. NOVl 10 Sequence Analysis
Further analysis of the NOVl 10a protein yielded the following properties shown in Table HOB.
Table HOB. Protein Sequence Properties NOVllOa
PSort 0.4500 probability located in cytoplasm; 0.1547 probability located in microbody analysis: (peroxisome); 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVl 10a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 1 IOC.
In a BLAST search of public sequence databases, the NOVl 10a protein was found to have homology to the proteins shown in the BLASTP data in Table HOD.
PFam analysis predicts that the NOVl 10a protein contains the domains shown in the Table HOE.
Example 111.
The NOVl 11 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 111 A.
Further analysis of the NOVl 1 la protein yielded the following properties shown in Table 11 IB.
Table 11 IB. Protein Sequence Properties NOVl 11a
PSort 0.8500 probability located in endoplasmic reticulum (membrane); 0.4400 analysis: probability located in plasma membrane; 0.1000 probability located in mitochondrial inner membrane; 0.1000 probability located in Golgi body
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVl 1 la protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 11 IC.
In a BLAST search of public sequence databases, the NOVl 1 la protein was found to have homology to the proteins shown in the BLASTP data in Table 11 ID.
PFam analysis predicts that the NOVl 1 la protein contains the domains shown in the Table H IE.
Example 112.
The NOVl 12 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 112 A.
Table 112A. NOVl 12 Sequence Analysis
Further analysis of the NOVl 12a protein yielded the following properties shown in Table 112B.
Table 112B. Protein Sequence Properties NOVl 12a
Psort 0.6400 probability located in plasma membrane; 0.4600 probability located in analysis: Golgi body; 0.3700 probability located in endoplasmic reticulum (membrane); 0.1000 probability located in endoplasmic reticulum (lumen)
SignalP Likely cleavage site between residues 22 and 23 analysis:
A search of the NOVl 12a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 112C.
In a BLAST search of public sequence databases, the NOVl 12a protein was found to have homology to the proteins shown in the BLASTP data in Table 112D.
PFam analysis predicts that the NOVl 12a protein contains the domains shown in the Table 112E.
Example 113.
The NOVl 13 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 113 A.
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 113B.
Table 113B. Comparison of NOVl 13a against NOVl 13b.
Further analysis of the NOVl 13a protein yielded the following properties shown in Table 113C.
Table 113C. Protein Sequence>Properties NOVl 13a
PSort 0.6400 probability located in plasma membrane; 0.4600 probability located in analysis: Golgi body; 0.3700 probability located in endoplasmic reticulum (membrane); 0.1000 probability located in endoplasmic reticulum (lumen)
SignalP Likely cleavage site between residues 59 and 60 analysis:
A search of the NOVl 13a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 113D.
In a BLAST search of public sequence databases, the NOVl 13a protein was found to have homology to the proteins shown in the BLASTP data in Table 113E.
PFam analysis predicts that the NOVl 13a protein contains the domains shown in the Table 113F.
Example 114.
The NOVl 14 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 114 A.
Table 114A. NOVl 14 Sequence Analysis
SEQ ED NO: 319 876 bp
NOVl 14a, AACTTGCTTTTGGGAGCCAGCGGTATGGCGTCGGGCTGCAAGATTGGCCCGTCCATCC
TCAACAGCGACCTGGCCAATTTAGGGGCCGAGTGCTCCCGGATGCTAGACTCTGGGGC
CG59861-01 DNA Sequence CGATTATCTGCACCTGGACGTAATGGACGGGCATTTTGTTCCCAACATCACCTTTGGT CACCCTGTGGTGGAAAGCCTTCGAAAGCAGCTAGGCCAGGACCCTTTCTTTGACATGC ACATGATGGTGTCCAAGCCAGAACAGTGGGTAAAGCCAATGGCTGTAGCAGGAGCCAA TCAGTACACCTTTCATCTCGAGGCTACTGAGAACCCAGGGGCTTTGATTAAAGACATT CGGGAGAATGGGATGAAGGTTGGCCTTGCCATCAAACCAGGAACCTCAGTTGAGTATT TGGCACCATGGGCTAATCAGATAGATATGGCCTTGGTTATGACAGTGGAACCGGGGTT TGGAGGGCAGAAATTCATGGAAGATATGATGCCAAAGGTTCACTGGTTGAGGACCCAG TTCCCATCTTTGGATATAGAGGTCGATGGTGGAGTAGGTCCTGACACTGTCCATAAAT GTGCAGAGGCAGGAGCTAACATGATTGTGTCTGGCAGTGCTATTATGAGGAGTGAAGA CCCCAGATCTGTGATCAATCTATTAAGAAATGTTTGCTCAGAAGCTGCTCAGAAACGT TCTCTTGATCGGTGAAACCATAAGGAGCCCAGTGTTCCTGTTCATGAAATCTCCCTTT
TACTGGAAAACAGGAATATTGACTACCAAATCACAATGCAATTGAAGCCGTACTGCTT
TTTTGAGCAGTTATTCATTCCAGTGATTAAAACTGATTGTGCAGAATAAAAAAAAAAA
AAAAAA
ORF Start: ATG at 25 ORF Stop: TGA at 709
SEQ ED NO: 320 228 aa MW at 24901.4kD
NOVl 14a, MASGCKIGPSILNSDLANLGAECSRMLDSGADYLHLDVMDGHFVPNITFGHPWESLR KQLGQDPFFDMHMMVSKPEQ VKP AVAGANQYTFHLEATENPGALIKDIRENG KVG
CG59861-01 Protein Sequence LAIKPGTSVEYLAP ANQIDMALVMTVEPGFGGQKFMEDMMPKVH LRTQFPSLDIEV DGGVGPDTVHKCAEAGANMIVSGSAIMRSEDPRSVINLLRNVCSEAAQKRSLDR
SEQ ED NO: 321 730 bp
NOVl 14b, AACTTGCTTTTGGGAGCCAGCGGTATGGCGTCGGGCTGCAAGATTGGCCCGTCCATCC
TCAACAGCGACCTGGCCAATTTAGGGGCCGAGTGCCTCCGGATGCTAGACTCTGGGGC
CG59861-02 DNA Sequence CGATTATCTGCACCTGGACGTAATGGACGGGCATTTTGTTCCCAACATCACCTTTGGT CACCCTGTGGTAGAAAGCCTTCGAAAGCAGCTAGGCCAGGACCCTTTCTTTGACATGC ACATGATGGTGTCCAAGCCAGAACAGTGGGTAAAGCCAATGGCTGTAGCAGGAGCCAA TCAGTACACCTTTCATCTCGAGGCTACTGAGAACCCAGGGGCTTTGATTAAAGACATT CGGGAGAATGGGATGAAGGTTGGCCTTGCCATCAAACCAGGAACCTCAGTTGAGTATT TGGCACCATGGGCTAATCAGATAGATATGGCCTTGGTTATGACAGTGGAACCGGGGTT TGGAGGGCAGAAATTCATGGAAGATATGATGCCAAAGGTTCACTGGTTGAGGACCCAG TTCCCATCTTTGGATATAGAGGTCGATGGTGGAGTAGGTCCTGACACTGTCCATAAAT GTGCAGAGGCAGGAGCTAACATGATTGTGTCTGGCAGTGCTATTATGAGGAGTGAAGA CCCCAGATCTGTGATCAATCTATTAAGAAATGTTTGCTCAGAAGCTGCTCAGAAACGT TCTCTTGATCGGTGAAACCATAAGGAGCCCAGTG
ORF Start: ATG at 25 ORF Stop: TGA at 709
SEQ ED NO: 322 228 aa MW at 24927.5kD
NOVl 14b, ASGCKIGPSILNSDLANLGAECLRMLDSGADYLHLDVMDGHFVPNITFGHPWESLR KQLGQDPFFDMHMMVSKPEQ VKPMAVAGANQYTFHLEATENPGALIKDIRENGMKVG
CG59861-02 Protein Sequence LAIKPGTSVEYLAP ANQIDMALVMTVEPGFGGQKFMEDMMPKVHWLRTQFPSLDIEV
DGGVGPDTVHKCAEAGANMIVSGSAIMRSEDPRSVINLLRNVCSEAAQKRSLDR
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 114B.
Further analysis of the NOVl 14a protein yielded the following properties shown in Table 114C.
Table 114C. Protein Sequence Properties NOVl 14a
PSort 0.6500 probability located in cytoplasm; 0.1753 probability located in lysosome analysis: (lumen); 0.1000 probability located in mitochondrial matrix space; 0.0000 probability located in endoplasmic reticulum (membrane)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVl 14a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 114D.
In a BLAST search of public sequence databases, the NOVl 14a protein was found to have homology to the proteins shown in the BLASTP data in Table 114E.
PFam analysis predicts that the NOVl 14a protein contains the domains shown in the Table 114F.
Table 114F. Domain Analysis of NOVl 14a
Identities/
NOVl 14a Match Expect
Pfam Domain Similarities Region Value
Example 115.
The NOVl 15 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 115 A.
Table 115A. NOVl 15 Sequence Analysis
SEQ ID NO: 323 1761 bp
NOVl 15a, AGTGTGGTACCTATCTGTCCCCCCTCTGGAGGGGTTGACAAGGGAAAGGGCACCGGGG
GGCACAGAGATGCAGGACAGATTGCACATCCTGGAGGACCTGAATATGCTCTACATTC
CG59857-01 DNA Sequence GGCAGATGGCACTCAGCCTGGAGGACACGGAGTTGCAGAGGAAGCTAGACCATGAGAT CCGGATGAGGGAAGGGGCCTGTAAGCTGCTGGCAGCCTGCTCCCAGCGAGAGCAGGCT CTGGAGGCCACCAAGAGCCTGCTAGTGTGCAACAGCCGCATCCTCAGCTACATGGGCG AGCTGCAGCGGCGCAAGGAGGCGCAGGTGCTGGGGAAGACAAGCCGGCGGCCTTCTGA CAGTGGCCCGCCCGCTGAGCGCTCCCCCTGCCGCGGCCGGGTCTGCATCTCTGACCTC CGGATTCCACTCATGTGGAAGGACACAGAATATTTCAAGAACAAAGACTTGCACCGCT GGGCTGTGTTCCTGCTGCTGCAGCTGGGGGAACACATCCAGGACACAGAGATGATCCT AGTGGACAGGACCCTCACAGACATCTCCTTTCAGAGCAATGTGCTCTTCGCTGAGGCG GGGCCAGACTTTGAACTGCGGTTAGAGCTGTATGGGGCCTGTGTGGAAGAAGAGGGGG CCCTGACTGGCGGCCCCAAGAGGCTTGCCACCAAACTCAGCAGCTCCCTGGGCCGCTC CTCAGGGAGGCGTGTCCGGGCATCGCTGGACAGTGCTGGGGGTTCAGGGAGCAGTCCC ATCTTGCTCCCCACCCCAGTTGTTGGTGGTCCTCGTTACCACCTCTTGGCTCACACCA CACTCACCCTGGCAGCAGTGCAAGATGGATTCCGCACACATGACCTCACCCTTGCCAG TCATGAGGAGAACCCTGCCTGGCTGCCCCTTTATGGTAGCGTGTGTTGCCGTCTGGCA GCTCAGCCTCTCTGCATGACTCAGCCCACTGCAAGTGGTACCCTCAGGGTGCAGCAAG CTGGGGAGATGCAGAACTGGGCACAAGTGCATGGAGTTCTGAAAGGCACAAACCTCTT CTGTTACCGGCAACCTGAGGATGCAGACACTGGGGAAGAGCCGCTGCTTACTATTGCT GTCAACAAGGAGACTCGAGTCCGGGCAGGGGAGCTGGACCAGGCTCTAGGACGGCCCT TCACCCTAAGCATCAGTAACCAGTATGGGGATGATGAGGTGACACACACCCTTCAGAC AGAAAGTCGGGAAGCACTGCAGAGCTGGATGGAGGCTCTGTGGCAGCTTTTCTTTGAC ATGAGCCAATGGAAGCAGTGCTGTGATGAAATCATGAAAATTGAAACTCCTGCTCCCC GGAAACCACCCCAAGCACTGGCAAAGCAGGGGTCCTTGTACCATGAGATGGCTATTGA GCCGCTGGATGACATCGCAGCGGTGACAGACATCCTGACCCAGCGGGAGGGCGCAAGG CTGGAGACACCCCCACCCTGGCTGGCAATGTTTACAGACCAGCCTGCCCTGCCTAACC CCTGCTCGCCTGCCTCAGTGGCCCCAGCCCCAGACTGGACCCACCCCCTGCCCTGGGG GAGACCCCGAACCTTTTCCCTGGATGCTGTCCCCCCAGACCACTCCCCTAGGGCTCGC TCGGTTGCCCCCCTCCCACCTCAGCGATCCCCACGGACCAGAGGCCTCTGCAGCAAAG GCCAACCTCGCACTTGGCTCCAGTCACCAGTGTGAGAGAGAAAGGTGCTGGCATAGGA TCTGCCCAGAAGAGAAAATGA
ORF Start: ATG at 68 ORF Stop: TGA at 1715
SEQ ED NO: 324 549 aa MW at 61171.0kD
NOVl 15a, MQDRLHILEDLNMLYIRQMALSLEDTELQRKLDHEIR REGACKLLAACSQREQALEA TKSLLVCNSRILSYMGELQRRKEAQVLGKTSRRPSDSGPPAERSPCRGRVCISDLRIP
CG59857-01 Protein Sequence LM KDTEYFKNKDLHRWAVFLLLQLGEHIQDTEMILVDRTLTDISFQSNVLFAEAGPD FELRLELYGACVEEEGALTGGPKRLATKLSSSLGRSSGRRVRASLDSAGGSGSSPILL PTPWGGPRYHLLAHTTLTLAAVQDGFRTHDLTLASHEENPA LPLYGSVCCRLAAQP LCMTQPTASGTLRVQQAGEMQN AQVHGVLKGTNLFCYRQPEDADTGEEPLLTIAVNK ETRVRAGELDQALGRPFTLSISNQYGDDEVTHTLQTESREALQS MEALWQLFFDMSQ WKQCCDEIMKIETPAPRKPPQALAKQGSLYHEMAIEPLDDIAAVTDILTQREGARLET PPPWLAMFTDQPALPNPCSPASVAPAPDWTHPLP GRPRTFSLDAVPPDHSPRARSVA PLPPQRSPRTRGLCSKGQPRTWLQSPV
Further analysis of the NOVl 15a protein yielded the following properties shown in Table 115B.
Table 115B. Protein Sequence Properties NOV115a
PSort 0.4500 probability located in cytoplasm; 0.3000 probability located in microbody analysis: (peroxisome); 0.1707 probability located in lysosome (lumen); 0.1000 probability located in mitochondrial matrix space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVl 15a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 115C.
In a BLAST search of public sequence databases, the NOVl 15a protein was found to have homology to the proteins shown in the BLASTP data in Table 115D.
Table 115D. Public BLASTP Results for NOVllSa
PFam analysis predicts that the NOVl 15a protein contains the domains shown in the Table 115E.
Example 116.
The NOVl 16 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 116 A.
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 116B.
Further analysis of the NOVl 16a protein yielded the following properties shown in Table 116C.
Table 116C. Protein Sequence Properties NOVl 16a
PSort 0.9190 probability located in plasma membrane; 0.3000 probability located in analysis: lysosome (membrane); 0.1888 probability located in microbody (peroxisome); 0.1000 probability located in endoplasmic reticulum (membrane)
SignalP Likely cleavage site between residues 28 and 29 analysis:
A search of the NOVl 16a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 116D.
In a BLAST search of public sequence databases, the NOVl 16a protein was found to have homology to the proteins shown in the BLASTP data in Table 116E.
PFam analysis predicts that the NOVl 16a protein contains the domains shown in the Table 116F.
Example 117.
The NOVl 17 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 117 A.
Further analysis of the NOVl 17a protein yielded the following properties shown in Table 117B.
Table 117B. Protein Sequence Properties NOVl 17a
Psort 0.4500 probability located in cytoplasm; 0.3000 probability located in microbody analysis: (peroxisome); 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
SignalP Likely cleavage site between residues 19 and 20 analysis:
A search of the NOVl 17a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 117C.
En a BLAST search of public sequence databases, the NOVl 17a protein was found to have homology to the proteins shown in the BLASTP data in Table 117D.
PFam analysis predicts that the NOVl 17a protein contains the domains shown in the Table 117E.
Example 118.
The NOVl 18 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 118 A.
Further analysis of the NOVl 18a protein yielded the following properties shown in Table 118B.
Table 118B. Protein Sequence Properties NOVl 18a
PSort 0.4500 probability located in cytoplasm; 0.3796 probability located in microbody analysis: (peroxisome); 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVl 18a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 118C.
In a BLAST search of public sequence databases, the NOVl 18a protein was found to have homology to the proteins shown in the BLASTP data in Table 118D.
PFam analysis predicts that the NOVl 18a protein contains the domains shown in the Table 118E.
The NOVl 19 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 119A.
Table 119A. NOVl 19 Sequence Analysis
SEQ ID NO: 333 1546 bp
NOVl 19a, GCTCAGTAGGCGTCGGGCTGTGATGCCCCAACTGCTCCAGCGTCTGCAGGCGCGCGCG
GGCGCGGTAGGCGTACTCGCTGGCCGGATAGCGCGTGATGATGAACTGGTAGGTCTGC
CG59928-01 DNA Sequence GCCGCATCGACGAACAGGCTCTCGCGCTCCAGGCATTGACCGCGCAGCAGGGAAATCT
CCGGCTGCAGGTAATTGCGTGAGCGGCTCTTGCGCTCGGCCTGCGACAGCTCCAGCGC
GACACGGGCGCAATCGCCTTCGTTGTAGGCGCGATAGGCGTTGTTCAGATGATGGTCG
AGCGAGACACGGGTGCAACCCGCAGCAACCAGGGCCACGGCCAGAATGATCAGGTTAC
GCATGGGCAATTCCTCCAATGAGCAGTGTATCGACAGCCCAGGCAAAAACTGAACAGC
GGCAAGCCGACGACGGTTTTTCTGGCGGCGCCTTGGCATGACGCCACTGCCTCTCATT
TTATCAACGCCAGCGCCACGACCGCTCGTCCTCTCGAACCAGCGCTAAATCCCCTTCT
GCGCTGACCCATATCAATGCCGTTCAGCGCAACAGGGTGTGTAATGTAGGTACAGACT
CCAGGCGAGGACGCTGCCATGAAACTGCAACGACTGTTGGTCGTCATCGACGCCGAAC
ACCAGCAACAACCCGCCCTGCAACGCGCAGCCGATGTGGCACGCAAGACCGGCGCCGA ATTGCACCTGTTGCAGATCGAATACCACCCAAGCCTGGAAAGCGGCCTGCTGGACAGC CATCTGCTCAACCGCGCCCGTGAAACCATCCTGCGACAGAGCCACGAGGCCCTGCGTG CCAGCGTCGCTCACCTGAGCGATGAAGGATTCAAGATCGCAGTGGACGTGCGCTGGGG CAAACGTCGTCATGAAGAAATCCTCGCCCGCGTCGCGGTGTTGCAACCGGACATCCTG TTCAAGTCGACTCATCCCAGCAGTGCGCTGCGCCGCCTGTTGTTCAGTGATACCAGTT GGCAGCTGATTCGCCGCAGCCCGGTGCCGCTGTGGCTGGTACACGACGCCGAGCCCCA TGGTCAGAGCCTGTGCGCTGCGCTCGACCCGCTGCACAGCGCGGACAAACCTGCCGCC CTCGATCATCAGTTGATTGATGCCAGCCAGACCCTGCAGGCCGAGCTCGGCTTACAGG CCCAATACCTGCATGCACAGGCGCCTCTGCCGCGGTCGCTGCTGTTCGACGCCGAGGT AGCGCAGGAATATGAAGACTACGTGACCCAGTGCAGCCGCGAGCACCGCGAAGCCTTC GACAAGCTGATCGCCCAGCACGCCATCGATAGAGCACAGGCCCACCTGTTGGACGGTT TTGCCGAGGAAGTCATCCCGCGTTTCGTGCGTGAGCACAATATAGGCCTGCTGGTGAT GGGCGCCATCGCCCGCGGCCATCTGGACAGCCTGCTGATCGGCCACACCGCAGAACGG GTGCTGGAACGTGTCGAGTGCGATCTGCTGGTGATCAAATCGCACGGCAAAGGGTAGT GCACAGGAACAATGACTACAGCCCGACGCCTACTGAGC
ORF Start: ATG at 599 ORF Stop: TAG at 1505
SEQ ED NO: 334 302 aa MW at 33922.3kD
NOVl 19a, MKLQRLLWIDAEHQQQPALQRAADVARKTGAELHLLQIEYHPSLESGLLDSHLLNRA RETILRQSHEALRASVAHLSDEGFKIAVDVR GKRRHEEILARVAVLQPDILFKSTHP
CG59928-01 Protein Sequence SSALRRLLFSDTSWQLIRRSPVPL LVHDAEPHGQSLCAALDPLHSADKPAALDHQLI DASQTLQAELGLQAQYLHAQAPLPRSLLFDAEVAQEYEDYVTQCSREHREAFDKLIAQ HAIDRAQAHLLDGFAEEVIPRFVREHNIGLLVMGAIARGHLDSLLIGHTAERVLERVE CDLLVIKSHGKG
Further analysis of the NOVl 19a protein yielded the following properties shown in Table 119B.
Table 119B. Protein Sequence Properties NOV119a
PSort 0.3000 probability located in microbody (peroxisome); 0.3000 probability analysis: located in nucleus; 0.2014 probability located in lysosome (lumen); 0.1000 probability located in mitochondrial matrix space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVl 19a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 119C.
Table 119C. Geneseq Results for NOVl 19a
NOVl 19a Identities/
Geneseq Protein/Organism/Length Residues/ Similarities for Expect Identifier [Patent #, Date] Match the Matched Value
Residues Region
No Significant Matches Found
En a BLAST search of public sequence databases, the NOVl 19a protein was found to have homology to the proteins shown in the BLASTP data in Table 119D.
PFam analysis predicts that the NOVl 19a protein contains the domains shown in the Table 119E.
Example 120.
The NOVl 20 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 120 A.
Table 120A. NOV120 Sequence Analysis
SEQ ID NO: 335 2202 bp
NOVl 20a, CACCCTCCCGCCCCGCCCCCCGTCCAATGCTGAGCTCAGTCTGCGTCTCGTCCTTCCG
CGGGCGCCAGGGGGCCAGCAAGCAGCAGCCGGCGCCACCGCCGCAGCCGCCCGAGGTC
CG59947-01 DNA Sequence CCCGGTGGCGACAGCGGCAAGATCGTGATCAACGTGGGCGGCGTGCGCCATGAGACGT ACCGCTCGACGCTGCGCACCCTGCCGGGGACGCGGCTGGCCGGCCTGACGGAGCCCGA GGCGGCGGCACGCTTCGACTACGACCCGGGCGCCGACGAGTTCTTCTTTGACCGGCAC CCGGGAGTCTTCGCGTACGTGCTCAACTACTACCGCACCGGCAAGCTGCACTGCCCAG CCGACGTGTGCGGGCCCCTGTTTGAGGAGGAGCTCGGCTTCTGGGGCATCGACGAGAC CGACGTGGAGGCCTGCTGCTGGATGACCTACCGGCAGCATCGCGACGCTGAGGAGGCG CTCGACTCCTTCGAGGCGCCCGACCCCGCGGGCGCCGCCAACGCCGCCAACGCCGCAG GCGCCCACGACGGAGGCCTGGACGACGAGGCGGGCGCGGGCGGCGGCGGCCTGGACGG AGCGGGCGGCGAGCTCAAGCGCCTCTGCTTCCAGGACGCGGGCGGCGGCGCCGGGGGG CCGCCAGGGGGCGCGGGCGGCGCGGGCGGCACATGGTGGCGCCGCTGGCAGCCCCGCG TGTGGGCGCTCTTCGAGGACCCCTACTCGTCGCGGGCTGCCAGGTATGTGGCCTTCGC CTCCCTCTTCTTCATCCTCATCTCCATCACCACCTTCTGCCTGGAAACCCATGAGGGC TTCATCCATATTAGCAACAAGACGGTGACCCAGGCCTCCCCGATCCCCGGGGCACCTC CGGAGAACATCACCAACGTGGAGGTGGAGACGGAGCCCTTCCTGACCTACGTGGAGGG GGTGTGCGTGGTCTGGTTCACCTTCGAGTTCCTCATGCGCATCACCTTCTGCCCAGAC AAGGTGGAGTTTCTTAAAAGCAGCCTCAACATCATCGACTGTGTGGCCATCCTGCCCT TCTATCTCGAGGTGGGCCTCTCGGGCCTCAGCTCCAAGGCCGCCAAAGACGTGCTGGG CTTCCTGCGGGTGGTCCGCTTCGTCCGCATCCTGCGCATCTTCAAGCTGACCCGGCAC TTCGTGGGGCTGCGCGTGCTGGGACACACGCTCCGCGCCAGCACCAACGAGTTCCTGC TGCTCATCATCTTCCTGGCCCTGGGGGTGCTCATCTTCGCCACCATGATTTACTACGC TGAGCGCATTGGCGCCGACCCCGATGACATCCTGGGCTCCAACCACACCTACTTCAAG AACATCCCCATTGGCTTCTGGTGGGCTGTGGTCACCATGACGACCCTGGGCTATGGAG ACATGTACCCCAAGACGTGGTCGGGGATGCTGGTCGGGGCGCTGTGTGCCCTGGCGGG GGTGCTGACCATCGCCATGCCTGTGCCCGTCATTGTCAACAACTTTGGCATGTACTAT TCGCTGGCCATGGCCAAGCAGAAGCTGCCCAAGAAGAAGAACAAACACATCCCCCGGC CCCCGCAACCGGGCTCGCCCAACTACTGCAAGCCTGACCCACCCCCGCCACCCCCGCC CCACCCGCACCACGGCAGCGGGGGCATCAGCCCGCCGCCACCCATCACCCCACCCTCC ATGGGGGTGACTGTGGCCGGGGCCTACCCAGCGGGGCCCCACACGCACCCCGGGCTGC TCAGGGGGGGAGCGGGTGGGCTGGGGATCATGGGGCTGCCTCCTCTGCCAGCCCCCGG CGAGCCTTGCCCGTTGGCTCAGGAGGAGGTGATTGAGATCAACCGGGCAGATCCTCGC CCCAATGGGGATCCGGCAGCAGCTGCGCTTGCCCACGAGGACTGCCCAGCCATTGACC AGCCTGCCATGTCCCCGGAAGACAAGAGCCCCATCACGCCTGGAAGCCGTGGCCGCTA TAGCCGGGACCGAGCCTGCTTCCTCCTCACCGACTATGCCCCTTCCCCTGATGGCTCC ATCCGAAAAGCCACTGGTGCTCCCCCACTGCCCCCCCAAGACTGGCGTAAGCCAGGCC CCCCAAGCTTCTTGCCCGACCTCAACGCCAACGCCGCGGCCTGGATATCCCCCTAGTG GACGAACCCCCTCCCCCCGGGCTCTTGTCACCGCCTGAGACCTCGCGAGACTTTCG
ORF Start: ATG at 27 ORF Stop: TAG at 2142
SEQ ED NO: 336 705 aa MW at 75590.5kD
NOV120a, MLSSVCVSSFRGRQGASKQQPAPPPQPPEVPGGDSGKIVINVGGVRHETYRSTLRTLP GTRLAGLTEPEAAARFDYDPGADEFFFDRHPGVFAYVL YYRTGKLHCPADVCGPLFE
CG59947-01 Protein Sequence EELGF GIDETDVEACCWMTYRQHRDAEEALDSFEAPDPAGAANAANAAGAHDGGLDD EAGAGGGGLDGAGGELKRLCFQDAGGGAGGPPGGAGGAGGTWWRR QPRVWALFEDPY SSRAARYVAFASLFFILISITTFCLETHEGFIHISNKTVTQASPIPGAPPENIT VEV ETEPFLTYVEGVCW FTFEFLMRITFCPDKVEFLKSSL IIDCVAILPFYLEVGLSG LSSKAAKDVLGFLRWRFVRILRIFKLTRHFVGLRVLGHTLRASTNEFLLLIIFLALG VLIFATMIYYAERIGADPDDILGSNHTYFKNIPIGFW AWTMTTLGYGDMYPKT SG MLVGALCALAGVLTIAMPVPVIVN FGMYYSLAMAKQKLPKKKNKHIPRPPQPGSPNY CKPDPPPPPPPHPHHGSGGISPPPPITPPSMGVTVAGAYPAGPHTHPGLLRGGAGGLG IMGLPPLPAPGEPCPLAQEEVIEINRADPRPNGDPAAAALAHEDCPAIDQPAMSPEDK SPITPGSRGRYSRDRACFLLTDYAPSPDGSIRKATGAPPLPPQD RKPGPPSFLPDLN ANAAA ISP
Further analysis of the NOVl 20a protein yielded the following properties shown in Table 120B.
Table 120B. Protein Sequence Properties NOV120a
PSort 0.6000 probability located in plasma membrane; 0.5071 probability located in analysis: mitochondrial inner membrane; 0.4000 probability located in Golgi body; 0.3000 probability located in endoplasmic reticulum (membrane)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVl 20a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 120C.
AAW42996 Putative mature potassium channel 2 17-510 171/503 (33%) 2e-66 protein - Homo sapiens, 494 aa. 4„425 240/503 (46%) [US5710019-A, 20-JAN-1998]
In a BLAST search of public sequence databases, the NOVl 20a protein was found to have homology to the proteins shown in the BLASTP data in Table 120D.
PFam analysis predicts that the NOVl 20a protein contains the domains shown in the Table 120E.
Example 121.
The NOVl 21 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 121 A.
Table 121 A. NOV121 Sequence Analysis
SEQ TD NO: 337 1943 bp
NOV121a, AGATCCACGTGATCTCCAAAGACCCCTGTTGTGTTGTGTTGGGAGGTGGATCCTGAAT
CCACCCAGAGAAGCCTGATACCAATAAAATCCCTGCTTGCTTTCCAGGAGACCCTTGG
CG59938-01 DNA Sequence TCTTCATGTCTTTGGTGTGTGCACTCTTGAACACATGCCAGGCACACAGGGTGCATGA
CGACAAGCCTAATATTGTCCTAATCATGGTTGATGACCTGGGTATTGGAGATCTGGGC TGCTACGGCAATGACACCATGAGGACGCCTCACATCGACCGCCTTGCCAGGGAAGGCG TGCGACTGACTCAGCACATCTCTGCCGCCTCCCTCTGCAGCCCAAGCCGGTCCGCGTT CTTGACGGGAAGATACCCCATCCGATCAGGTATGGTTTCTAGTGGTAATAGACGTGTC ATCCAAAATCTTGCAGTCCCCGCAGGCCTCCCTCTTAATGAGACAACACTTGCAGCCT TGCTAAAGAAGCAAGGATACAGCACGGGGCTTATAGGTAAGTTAGGCAAATGGCACCT GGGTTTGAGCTGCGCCTCTCGGAATGATCACTGTTACCACCCGCTCAACCATGGTTTT CACTACTTTTACGGGGTGCCTTTTGGACTTTTAAGCGACTGCCAGGCATCCAAGACAC CAGAACTGCACCGCTGGCTCAGGATCAAACTGTGGATCTCCACGGTAGCCCTTGCCCT GGTTCCTTTTCTGCTTCTCATTCCCAAGTTCGCCCGCTGGTTCTCAGTGCCATGGAAG GTCATCTTTGTCTTTGCTCTCCTCGCCTTTCTGTTTTTCACTTCCTGGTACTCTAGTT ATGGATTTACTCGACGTTGGAATTGCATCCTTATGAGGAACCATGAAATTATCCAGCA GCCAATGAAAGAGGAGAAAGTAGCTTCCCTCATGCTGAAGGAGGCACTTGCTTTCATT GAAAGGTACAAAAGGGAACCTTTTCTCCTCTTTTTTTCCTTCCTGCACGTACATACTC CACTCATCTCCAAAAAGAAGTTTGTTGGGCGCAGTAAATATGGCAGGTATGGGGACAA TGTAGAAGAAATGGATTGGATGGTGGGTGGTAAAATCCTGGATGCCCTGGACCAGGAG CGCCTGGCCAACCACACCTTGGTGTACTTCACCTCTGACAACGGGGGCCACCTGGAGC CCCTGGACGGGGCTGTTCAGCTGGGTGGCTGGAACGGGATCTACAAAGGTGGCAAAGG AATGGGAGGATGGGAAGGAGGTATCCGTGTGCCAGGGATATTCCGGTGGCCGTCAGTC TTGGAGGCTGGGAGAGTGATCAATGAGCCCACCAGCTTAATGGACATCTATCCGACGC TGTCTTATATAGGCGGAGGGATCTTGTCCCAGGACAGAGTGATTGACGGCCAGAACCT AATGCCCCTGCTGGAAGGAAGGGCGTCCCACTCCGACCACGAGTTCCTCTTCCACTAC TGTGGGGTCTATCTGCACACGGTCAGGTGGCATCAGAAGGACACTGTGTGGAAAGCTC ATTATGTGACTCCTAAATTCTACCCTGAAGGAACAGGTGCCTGCTATGGGAGTGGAAT ATGTTCATGTTCGGGGGATGTAACCTACCACGACCCACCACTCCTCTTTGACATCTCA AGAGACCCTTCAGAAGCCCTTCCACTGAACCCTGACAATGAGCCATTATTTGACTCCG TGATCAAAAAGATGGAGGCAGCCATAAGAGAGCATCGTAGGACACTAACACCTGTCCC ACAGCAGTTCTCTGTGTTCAACACAATTTGGAAACCATGGCTGCAGCCTTGCTGTGGG ACCTTCCCCTTCTGTGGGTGTGACAAGGAAGATGACATCCTTCCCATGGCTCCCTGAG ACCATGCGGACCACGTGTTACCCACCACAAACTTACTGTTACAATGGTCATAGGAGCA
GAGCTCACCTGACTGATTCATTCCATTTG
ORF Start: ATG at 122 ORF Stop: TGA at 1853
SEQ ID NO: 338 577 aa MW at 65099.5kD
NOV121a, MSLVCALLNTCQAHRVHDDKPNIVLIMVDDLGIGDLGCYGNDTMRTPHIDRLAREGVR LTQHISAASLCSPSRSAFLTGRYPIRSGMVSSGNRRVIQNLAVPAGLPLNETTLAALL
CG59938-01 Protein Sequence KKQGYSTGLIGKLGK HLGLSCASR DHCYHPLNHGFHYFYGVPFGLLSDCQASKTPE LHR LRIKL ISTVALALVPFLLLIPKFARWFSVP KVIFVFALLAFLFFTS YSSYG FTRR CILMRNHEIIQQPMKEEKVASLMLKEALAFIERYKREPFLLFFSFLHVHTPL ISKKKFVGRSKYGRYGD VEEMDWMVGGKILDALDQERLANHTLVYFTSDNGGHLEPL DGAVQLGGWNGIYKGGKGMGG EGGIRVPGIFR PSVLEAGRVINEPTSLMDIYPTLS YIGGGILSQDRVIDGQNLMPLLEGRASHSDHEFLFHYCGVYLHTVR HQKDTVWKAHY VTPKFYPEGTGACYGSGICSCSGDVTYHDPPLLFDISRDPSEALPLNPDNEPLFDSVI KKMEAAIREHRRTLTPVPQQFSVFNTI KPWLQPCCGTFPFCGCDKEDDILPMAP
Further analysis of the NOV121a protein yielded the following properties shown in Table 121B.
Table 121B. Protein Sequence Properties NOV121a
PSort 0.6400 probability located in plasma membrane; 0.4600 probability located in analysis: Golgi body; 0.3700 probability located in endoplasmic reticulum (membrane); 0.1000 probability located in endoplasmic reticulum (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVl 2 la protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 121 C.
In a BLAST search of public sequence databases, the NOV121 a protein was found to have homology to the proteins shown in the BLASTP data in Table 12 ID.
PFam analysis predicts that the NOV121a protein contains the domains shown in the Table 121E.
Example 122.
The NOVl 22 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 122 A.
Table 122A. NOV122 Sequence Analysis
SEQ TD NO: 339 3005 bp
NOV122a, ATTTCTTTGGTGTTGTCTTCACAGCTGAACTTGCAAAACAGATTGGAACTTCAAGATT
ATCAATAATCGGAGATACGTATATTTTATTTGTAAAGAAAACATGGCTGCCCTATTCC
CG59746-01 DNA Sequence TACGTGGTTTTGTCCAAATAGGGAACTGCAAGACTGGGATATCTAAGTCAAAAGAAGC ATTCATTGAAGCAGTGGAAAGAAAGAAGAAAGATAGACTGGTGCTGTATTTCAAAAGT GGAAAATATAGCACTTTTCGGCTAAGTGATAATATTCAAAATGTAGTCCTTAAATCCT ATAGAGGAAACCAAAATCACCTGCATTTAACTTTACAAAATAATAATGGCTTGTTTAT TGAAGGATTATCCTCCACAGATGCTGAACAATTGAAGATATTCTTGGACAGAGTTCAT CAAAACGAGGTTCAGCCACCTGTGAGACCTGGTAAGGGTGGGAGTGTCTTTTCTAGCA CAACACAGAAGGAAATCAACAAAACTTCATTCCACAAAGTTGATGAGAAATCAAGTAG CAAATCTTTTGAGATAGCAAAAGGAAGTGGGACAGGTGTCCTTCAGAGGATGCCTTTG CTTACATCAAAATTGACACTTACTTGCGGAGAGTTATCAGAAAATCAGCACAAGAAGA GGAAAAGAATGCTCTCATCTAGCTCAGAGATGAATGAGGAATTCTTGAAAGAAAATAA TTCTGTAGAATACAAGAAATCCAAGGCAGATTGTTCGAGGTGTGTAAGCTATAATCGA GAGAAACAATTGAAGTTAAAAGAGTTAGAAGAGAATAAGAAATTGGAATGTGAATCTT CATGCATCATGAACGCCACTGGAAATCCTTACCTAGATGACATTGGTCTTCTCCAAGC TCTCACTGAGAAAATGGTTTTGGTATTTCTGTTACAACAAGGGTATAGTGACGGTTAC ACAAAGTGGGATAAATTAAAACTATTTTTTGAATTATTTCCAGAGAAAATATGCCACG GCCTCCCCAATTTGGGAAACACCTGTTATATGAATGCAGTGTTACAGTCTCTACTTTC AATCCCATCGTTTGCTGATGATTTACTTAATCAGAGTTTCCCATGGGGTAAAATTCCC CTTAATGCTCTTACCATGTGCTTGGCACGGCTACTTTTTTTTAAAGATACCTATAATA TAGAAATCAAGGAGATGTTACTCTTGAATCTTAAAAAGGCCATTTCAGCAGCTGCAGA GATATTCCATGGCAATGCACAGAACGATGCTCATGAGTTTTTAGCTCACTGTTTAGAT CAACTGAAAGATAACATGGAAAAACTCAACACAATTTGGAAGCCTAAAAGTGAATTTG GGGAAGATAATTTTCCTAAACAGGTTTTTGCTGATGATCCTGACACCAGTGGGTTTTC TTGCCCTGTCATTACTAATTTTGAGTTAGAGTTGTTGCACTCCATTGCTTGTAAAGCT TGTGGTCAGGTTATTCTCAAGACAGAACTGAATAATTACCTCTCCATCAACCTTCCCC AAAGAATAAAAGCACATCCTTCATCTATTCAGTCTACTTTTGATCTTTTTTTTGGAGC AGAAGAGCTTGAGTATAAATGTGCAAAATGTGAGCACAAGACTTCCGTTGGAGTGCAC TCATTCAGTAGGCTACCTAGAATCCTTATTGTTCACCTCAAACGCTATAGCTTGAATG AGTTTTGTGCATTAAAGAAGAATGACCAGGAAGTCATCATTTCCAAATATTTAAAGGT GTCTTCTCATTGCAATGAAGGCACCAGACCACCTCTTCCCTTGAGTGAGGATGGAGAA ATTACAGATTTCCAATTATTAAAAGTTATTCGAAAGATGACTTCTGGAAACATCAGTG TATCATGGCCTGCAACAAAGGAATCCAAAGATATCCTGGCTCCACACATTGGATCAGA TAAGGAGTCTGAACAAAAAAAAGGCCAGACAGTCTTTAAAGGGGCAAGCAGAAGACAG CAGCAAAAGTACCTTGGAAAAAATTCTAAACCAAATGAGCTAGAATCTGTATACTCAG GAGATCGAGCATTCATTGAAAAAGAACCGTTAGCTCACTTAATGACGTATCTGGAAGA TACCTCACTTTGTCAGTTCCACAAAGCTGGAGGTAAACCTGCCAGCAGCCCAGGCACA CCTCTCTCAAAAGTTGACTTTCAAACAGTGCCCGAAAATCCAAAACGAAAGAAATATG TGAAAACCAGTAAGTTTGTAGCTTTTGATAGGATTATCAATCCTACTAAAGATTTGTA TGAAGATAAAAATATCAGAATTCCAGAAAGATTCCAAAAAGTGTCTGAACAGACTCAG CAGTGTGACGGTATGAGAATCTGTGAACAAGCCCCTCAGCAGGCACTGCCTCAAAGCT TTCCAAAGCCAGGCACCCAGGGGCACACAAAGAACCTCCTAAGACCTACAAAATTAAA TCTACAGAAGTCTAACAGGAATTCCCTACTTGCACTGGGTTCCAATAAGAATCCAAGA AACAAAGACATTTTAGATAAGATAAAATCTAAAGCCAAGGAAACAAAAAGAAATGATG ATAAGGGAGATCATACCTACCGGCTCATTAGTGTTGTCAGCCATCTTGGGAAGACTCT AAAGTCAGGCCATTATATCTGTGATGCCTATGACTTTGAGAAACAGATCTGGTTCACT TACGATGATATGCGGGTGTTAGGTATCCAGGAGGCCCAGATGCAGGAGGATAGGCGTT GCACTGGGTACATCTTCTTTTACATGCATAATGAGATCTTTGAAGAGATGTTGAAAAG AGAAGAGAATGCCCAGCTTAATAGCAAGGAGGTAGAGGAGACCCTTCAGAAGGAATAA GAGGAACGTACTCCTCCTTGTACAGATCTGCCTGACTGTCTCACTCGATACCACTTCC
TCCATGGAAGGAAAACTGTGAACTTTATCCAGAGATGAAAATGCAATTAGTCTAGGAC
CAAAGGTCAAACAGAAACACTTAATGGGGAGATCTGCATTCTAATCC
ORF Start: ATG at 101 ORF Stop: TAA at 2840
SEQ ED NO: 340 913 aa MW at l04046.0kD
NOVl 22a, MAALFLRGFVQIGNCKTGISKSKEAFIEAVERKKKDRLVLYFKSGKYSTFRLSDNIQN WLKSYRGNQNHLHLTLQNNNGLFIEGLSSTDAEQLKIFLDRVHQNEVQPPVRPGKGG
CG59746-01 Protein Sequence SVFSSTTQKEINKTSFHKVDEKSSSKSFEIAKGSGTGVLQRMPLLTSKLTLTCGELSE NQHKKRKRMLSSSSEMNEEFLKENNSVEYKKSKADCSRCVSYNREKQLKLKELEENKK LECESSCIMNATGNPYLDDIGLLQALTEKMVLVFLLQQGYSDGYTK DKLKLFFELFP EKICHGLPNLGNTCYMNAVLQSLLSIPSFADDLLNQSFP GKIPLNALTMCLARLLFF KDTYNIEIKEMLLLNLKKAISAAAEIFHGNAQ DAHEFLAHCLDQLKDNMEKL TIWK PKSEFGEDNFPKQVFADDPDTSGFSCPVITNFELELLHSIACKACGQVILKTELNNYL SINLPQRIKAHPSSIQSTFDLFFGAEELEYKCAKCEHKTSVGVHSFSRLPRILIVHLK RYSLNEFCALKKNDQEVIISKYLKVSSHCNEGTRPPLPLSEDGEITDFQLLKVIRKMT SGNISVSWPATKESKDILAPHIGSDKESEQKKGQTVFKGASRRQQQKYLGKNSKPNEL ESVYSGDRAFIEKEPLAHLMTYLEDTSLCQFHKAGGKPASSPGTPLSKVDFQTVPENP
Further analysis of the NOVl 22a protein yielded the following properties shown in Table 122B.
Table 122B. Protein Sequence Properties NOV122a
PSort 0.7000 probability located in nucleus; 0.4270 probability located in analysis: mitochondrial matrix space; 0.3000 probability located in microbody (peroxisome); 0.1047 probability located in mitochondrial inner membrane
SignalP Likely cleavage site between residues 16 and 17 analysis:
A search of the NOVl 22a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 122C.
In a BLAST search of public sequence databases, the NOVl 22a protein was found to have homology to the proteins shown in the BLASTP data in Table 122D.
PFam analysis predicts that the NOVl 22a protein contains the domains shown in the Table 122E.
Example 123.
The NOVl 23 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 123 A.
Table 123 A. NOV123 Sequence Analysis
SEQ ID NO: 341 2146 bp
NOV123a, GAAGGAGCGGGCATGAGGCGCTGCCCGTGCCGTGGGAGCCTGAACGAGGCGGAGGCCG
GGGCGCTGCCCGCGGCGGCCCGCATGGGACTGGAGGCGCCGCGAGGAGGGCGGCGGCG
CG88613-01 DNA Sequence GCAGCCGGGACAGCAGCGACCTGGGCCCGGCGCAGGGGCCCCGGCGGGGCGGCCGGAG GGGGGCGGGCCCTGGGCCCGGACAGAGGGGTCCAGCCTCCACAGCGAGCCTGAGAGGG CCGGCCTCGGGCCTGCGCCGGGGACAGAGAGTCCGCAGGCAGAATTCTGGACAGACGG ACAGACTGAGCCCGCGGCAGCTGGCCTTGGAGTAGAGACCGAGAGGCCCAAGCAAAAG ACGGAGCCAGACAGGTCCAGCCTCCGGACGCATCTAGAATGGAGCTGGTCAGAGCTGG AGACGACTTGTCTTTGGACGGAGACCGGGACAGATGGCCTTTGGACTGATCCGCACAG GTCCGACCTCCAGTTTCAGCCCGAGGAGGCCAGCCCCTGGACACAGCCAGGGGTTCAT GGGCCCTGGACAGAGCTGGAAACGCATGGGTCACAGACTCAGCCAGAGAGGGTCAAGT CCTGGGCTGATAACCTCTGGACCCACCAGAACAGTTCCAGCCTCCAGACTCACCCAGA AGGAGCCTGTCCCTCAAAAGAGCCAAGTGCTGATGGCTCCTGGAAAGAATTGTATACT GATGGCTCCAGGACACAACAGGATATTGAAGGTCCCTGGACAGAGCCATATACTGATG GCTCCCAGAAAAAACAGGATACTGAAGCAGCCAGGAAACAGCCTGGCACTGGTGGTTT CCAAATACAACAGGATACTGATGGCTCCTGGACACAACCTAGCACTGACGGTTCCCAG ACAGCACCTGGGACAGACTGCCTCTTGGGAGAGCCTGAGGATGGCCCATTAGAGGAAC CAGAGCCTGGAGAATTGCTGACTCACCTGTACTCTCACCTGAAGTGTAGCCCCCTGTG CCCTGTGCCCCGCCTCATCATTACCCCTGAGACCCCTGAGCCTGAGGCCCAGCCAGTG GGACCCCCCTCCCGGGTTGAGGGGGGCAGCGGCGGCTTCTCCTCTGCCTCTTCTTTCG ACGAGTCTGAGGATGACGTGGTGGCCGGGGGCGGAGGTGCCAGCGATCCCGAGGACAG GTCTGGGAGCAAACCCTGGAAGAAGCTGAAGACAGTTCTGAAGTATTCACCCTTTGTG GTCTCCTTCCGAAAACACTACCCTTGGGTCCAGCTTTCTGGACATGCTGGGAACTTCC AGGCAGGAGAGGATGGTCGGATTCTGAAACGTTTCTGTCAGTGTGAGCAGCGCAGCCT GGAGCAGCTGATGAAAGACCCGCTGCGACCTTTCGTGCCTGCCTACTATGGCATGGTG CTGCAGGATGGCCAGACCTTCAACCAGATGGAAGACCTCCTGGCTGACTTTGAGGGCC CCTCCATTATGGACTGCAAGATGGGCAGCAGGACCTATCTGGAAGAGGAGCTAGTGAA GGCACGGGAACGTCCCCGTCCCCGGAAGGACATGTATGAGAAGATGGTGGCTGTGGAC CCTGGGGCCCCTACCCCTGAGGAGCATGCCCAGGGTGCAGTCACCAAGCCCCGCTACA TGCAGTGGAGGGAAACCATGAGCTCCACCTCTACCCTGGGCTTCCGGATCGAGGGCAT CAAGAAGGCAGATGGGACCTGTAACACCAACTTCAAGAAGACGCAGGCACTGGAGCAG GTGACAAAAGTGCTGGAGGACTTCGTGGATGGAGACCACGTCATCCTGCAAAAGTACG TGGCATGCCTAGAAGAACTTCGTGAAGCTCTGGAGATCTCCCCCTTCTTCAAGACCCA CGAGGTGGTAGGCAGCTCCCTCCTCTTCGTGCACGACCACACCGGCCTGGCCAAGGTC TGGATGATAGACTTCGGCAAGACGGTGGCCTTGCCCGACCACCAGACGCTCAGCCACA GGCTGCCCTGGGCTGAGGGCAACCGTGAGGACGGCTACCTCTGGGGCCTGGACAACAT GATCTGCCTCCTGCAGGGGCTGGCACAGAGCTGAGCTGCTCAGCCACCATCAGGTTAA
TTGGATGGCGCCAGTCTGGCTGGAGGAGCCCTGAGATGCCATGGGAGGCCTGAGGTTG
ORF Start: ATG at 13 ORF Stop: TGA at 2062
SEQ ED NO: 342 683 aa MW at 75206.8kD
NOV123a, MRRCPCRGSLNEAEAGALPAAARMGLEAPRGGRRRQPGQQRPGPGAGAPAGRPEGGGP WARTEGSSLHSEPERAGLGPAPGTESPQAEFWTDGQTEPAAAGLGVETERPKQKTEPD
CG88613-01 Protein Sequence RSSLRTHLEWSWSELETTCLWTETGTDGLWTDPHRSDLQFQPEEASP TQPGVHGPWT ELETHGSQTQPERVKSWADNL THQNSSSLQTHPEGACPSKEPSADGS KELYTDGSR TQQDIEGP TEPYTDGSQKKQDTEAARKQPGTGGFQIQQDTDGS TQPSTDGSQTAPG TDCLLGEPEDGPLEEPEPGELLTHLYSHLKCSPLCPVPRLIITPETPEPEAQPVGPPS RVEGGSGGFSSASSFDESEDDWAGGGGASDPEDRSGSKP KKLKTVLKYSPFWSFR KHYP VQLSGHAGNFQAGEDGRILKRFCQCEQRSLEQLMKDPLRPFVPAYYGMVLQDG QTFNQMEDLLADFEGPSIMDCKMGSRTYLEEELVKARERPRPRKDMYEKMVAVDPGAP TPEEHAQGAVTKPRYMQWRETMSSTSTLGFRIEGIKKADGTCNTNFKKTQALEQVTKV LEDFVDGDHVILQKYVACLEELREALEISPFFKTHEWGSSLLFVHDHTGLAKVWMID FGKTVALPDHQTLSHRLP AEGNREDGYL GLDNMICLLQGLAQS
Further analysis of the NOVl 23 a protein yielded the following properties shown in Table 123B.
Table 123B. Protein Sequence Properties NOV123a
PSort 0.5663 probability located in microbody (peroxisome); 0.3000 probability analysis: located in nucleus; 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVl 23 a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 123C.
In a BLAST search of public sequence databases, the NOV123a protein was found to have homology to the proteins shown in the BLASTP data in Table 123D.
PFam analysis predicts that the NOV123a protein contains the domains shown in the Table 123E.
Table 123E. Domain Analysis of NOV123a
Identities/
Pfam Domain NOV123a Match Region Similarities Expect Value for the Matched Region
No Significant Matches Found
Example 124.
The NOVl 24 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 124A.
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 124B.
Table 124B. Comparison of NOV124a against NOVl 24b.
NOVl 24a Residues/ Identities/
Protein Sequence Match Residues Similarities for the Matched Region
NOV124b 1..419 335/419 (79%) 1..419 335/419 (79%)
Further analysis of the NOVl 24a protein yielded the following properties shown in Table 124C.
Table 124C. Protein Sequence Properties NOVl 24a
PSort 0.8202 probability located in mitochondrial inner membrane; 0.6000 probability analysis: located in endoplasmic reticulum (membrane); 0.3500 probability located in nucleus; 0.3034 probability located in mitochondrial intermembrane space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVl 24a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 124D.
In a BLAST search of public sequence databases, the NOVl 24a protein was found to have homology to the proteins shown in the BLASTP data in Table 124E.
PFam analysis predicts that the NOVl 24a protein contains the domains shown in the Table 124F.
Example 125.
The NOVl 25 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 125 A.
NOV125a, GGACCACTTCTGATGCATCTCTGGGTCCCAACACTATCCACTGCAAGGCCTCGAAACA
GGGGGGCCAGATGGGACCCCCATTTAGCACAAGAGAGACGTCCACACTCTGTGAGCCC
CG59991-01 DNA Sequence AAAGGGAGAAGGCTCAGGCCACGGCAGAGACGGAACCAGGAAAACGTCACGAAAAACA GCCTCAAGTTGCCAGGTCCCTTGCAGGAACAGACAGGCCTGGGGCCGCCCCACCTGGG CTCAGAGCTTGGGCTGCATGGAGGTGACACATGGGACTACAAGAGTCACGTGATGACC AAATTCGCTGAGGAGGAGGATGTACGTCGTAGTTTTGAAAACACTGCTGCTGACTGGC CGGAAATGCAAACGTTGGCTGGTGCTTTTGATTCAGACCGGTGGGGCTTCCGGCCTCG CACGGTGGTTCTGCACGGAAAGTCAGGAATTGGGAAATCGGCTCTAGCCAGAAGGATC GTGCTGTGCTGGGCGCAAGGTGGACTCTACCAGGGAATGTTCTCCTACGTCTTCTTCC TCCCCGTTAGAGAGATGCAGCGGAAGAAGGAGAGCAGTGTCACAGAGTTCATCTCCAG GGAGTGGCCAGACTCCCAGGCTCCGGTGACGGAGATCATGTCCCGACCAGAAAGGCTG TTGTTCATCATTGACGGTTTCGATGACCTGGGCTCTGTCCTCAACAATGACACAAAGC TCTGCAAAGACTGGGCTGAGAAGCAGCCTCCGTTCACCCTCATACGCAGTCTGCTGAG GAAGGTCCTGCTCCCTGAGTCCTTCCTGATCGTCACCGTCAGAGACGTGGGCACAGAG AAGCTCAAGTCAGAGGTCGTGTCTCCCCGTTACCTGTTAGTTAGAGGAATCTCCGGGG AACAAAGAATCCACTTGCTCCTTGAGCGCGGGATTGGTGAGCATCAGAAGACACAAGG GTTGCGTGCGATCATGAACAACCGTGAGCTGCTCGACCAGTGCCAGGTGCCCGCCGTG GGCTCTCTCATCTGCGTGGCCCTGCAGCTGCAGGACGTGGTGGGGGAGAGCGTCGCCC CCTTCAACCAAACGCTCACAGGCCTGCACGCCGCTTTTGTGTTTCATCAGCTCACCCC TCGAGGCGTGGTCCGGCGCTGTCTCAATCTGGAGGAAAGAGTTGTCCTGAAGCGCTTC TGCCGTATGGCTGTGGAGGGAGTGTGGAATAGGAAGTCAGTGTTTGACGGTGACGACC TCATGGTTCAAGGACTCGGGGAGTCTGAGCTCCGTGCTCTGTTTCACATGAACATCCT TCTCCCAGACAGCCACTGTGAGGAGTACTACACCTTCTTCCACCTCAGTCTCCAGGAC TTCTGTGCCGCCTTGTACTACGTGTTAGAGGGCCTGGAAATCGAGCCAGCTCTCTGCC CTCTGTACGTTGAGAAGACAAAGAGGTCCATGGAGCTTAAACAGGCAGGCTTCCATAT CCACTCGCTTTGGATGAAGCGTTTCTTGTTTGGCCTCGTGAGCGAAGACGTAAGGAGG CCACTGGAGGTCCTGCTGGGCTGTCCCGTTCCCCTGGGGGTGAAGCAGAAGCTTCTGC ACTGGGTCTCTCTGTTGGGTCAGCAGCCTAATGCCACCACCCCAGGAGACACCCTGGA CGCCTTCCACTGTCTTTTCGAGACTCAAGACAAAGAGTTTGTTCGCTTGGCATTAAAC AGCTTCCAAGAAGTGTGGCTTCCGATTAACCAGAACCTGGACTTGATAGCATCTTCCT TCTGCCTCCAGCACTGTCCGTATTTGCGGAAAATTCGGGTGGATGTCAAAGGGATCTT CCCAAGAGATGAGTCCGCTGAGGCATGTCCTGTGGTCCCTCTATGGATGCGGGATAAG ACCCTCATTGAGGAGCAGTGGGAAGATTTCTGCTCCATGCTTGGCACCCACCCACACC TGCGGCAGCTGGACCTGGGCAGCAGCATCCTGACAGAGCGGGCCATGAAGACCCTGTG TGCCAAGCTGAGGCATCCCACCTGCAAGATACAGACCCTGATGTTTAGAAATGCACAG ATTACCCCTGGTGTGCAGCACCTCTGGAGAATCGTCATGGCCAACCGTAACCTAAGAT CCCTCAACTTGGGAGGCACCCACCTGAAGGAAGAGGATGTAAGGATGGCGTGTGAAGC CTTAAAACACCCAAAATGTTTGTTGGAGTCTTTGAGGCTGGATTGCTGTGGATTGACC CATGCCTGTTACCTGAAGATCTCCCAAATCCTTACGACCTCCCCCAGCCTGAAATCTC TGAGCCTGGCAGGAAACAAGGTGACAGACCAGGGAGTAATGCCTCTCAGTGATGCCTT GAGAGTCTCCCAGTGCGCCCTGCAGAAGCTGATACTGGAGGACTGTGGCATCACAGCC ACGGGTTGCCAGAGTCTGGCCTCAGCCCTCGTCAGCAACCGGAGCTTGACACACCTGT GCCTATCCAACAACAGCCTGGGGAACGAAGGTGTAAATCTACTGTGTCGATCCATGAG GCTTCCCCACTGTAGTCTGCAGAGGCTGATGCTGAATCAGTGCCACCTGGACACGGCT GGCTGTGGTTTTCTTGCACTTGCGCTTATGGGTAACTCATGGCTGACGCACCTGAGCC TTAGCATGAACCCTGTGGAAGACAATGGCGTGAAGCTTCTGTGCGAGGTCATGAGAGA ACCATCTTGTCATCTCCAGGACCTGGAGTTGGTAAAGTGTCATCTCACCGCCGCGTGC TGTGAGAGTCTGTCCTGTGTGATCTCGAGGAGCAGACACCTGAAGAGCCTGGATCTCA CGGACAATGCCCTGGGTGACGGTGGGGTTGCTGCACTGTGCGAGGGACTGAAGCAAAA GAACAGTGTTCTGACGAGACTCGGGTTGAAGGCATGTGGACTGACTTCTGATTGCTGT GAGGCACTCTCCTTGGCCCTTTCCTGCAACCGGCATCTGACCAGTCTAAACCTGGTGC AGAATAACTTCAGTCCCAAAGGAATGATGAAGCTGTGTTCGGCCTTTGCCTGTCCCAC GTCTAACTTACAGATAATTGGGCTGTGGAAATGGCAGTACCCTGTGCAAATAAGGAAG CTGCTGGAGGAAGTGCAGCTACTCAAGCCCCGAGTCGTAATTGACGGTAGTTGGCATT CTTTTGATGAAGATGACCGGTACTGGTGGAAAAACTGAAGATACGGAAACCTGCCCCA CTCACACCCATCTGATGGAGGAACTTTAAACGCTGT
ORF Start: ATG at 69 ORF Stop: TGA at 3168
SEQ ID NO: 348 1033 aa MW at l l6310.7kD
NOV125a, MGPPFSTRETSTLCEPKGRRLRPRQRRNQENVTKNSLKLPGPLQEQTGLGPPHLGSEL GLHGGDTWDYKSHVMTKFAEEEDVRRSFENTAADWPEMQTLAGAFDSDRWGFRPRTW
CG59991-01 Protein Sequence LHGKSGIGKSALARRIVLCWAQGGLYQGMFSYVFFLPVREMQRKKESSVTEFISRE P DSQAPVTEIMSRPERLLFIIDGFDDLGSVLNNDTKLCKDWAEKQPPFTLIRSLLRKVL LPESFLIVTVRDVGTEKLKSEWSPRYLLVRGISGEQRIHLLLERGIGEHQKTQGLRA IMNNRELLDQCQVPAVGSLICVALQLQDWGESVAPFNQTLTGLHAAFVFHQLTPRGV VRRCL LEERWLKRFCRMAVEGVW RKSVFDGDDLMVQGLGESELRALFHMNILLPD SHCEEYYTFFHLSLQDFCAALYYVLEGLEIEPALCPLYVEKTKRSMELKQAGFHIHSL MKRFLFGLVSEDVRRPLEVLLGCPVPLGVKQKLLHWVSLLGQQPNATTPGDTLDAFH CLFETQDKEFVRLALNSFQEVWLPINQNLDLIASSFCLQHCPYLRKIRVDVKGIFPRD ESAEACPWPLWMRDKTLIEEQWEDFCSMLGTHPHLRQLDLGSSILTERAMKTLCAKL RHPTCKIQTLMFRNAQITPGVQHL RIVMANRNLRSLNLGGTHLKEEDVRMACEALKH PKCLLESLRLDCCGLTHACYLKISQILTTSPSLKSLSLAGNKVTDQGVMPLSDALRVS QCALQKLILEDCGITATGCQSLASALVSNRSLTHLCLSNNSLGNEGVNLLCRSMRLPH CSLQRLMLNQCHLDTAGCGFLALALMGNSWLTHLSLSMNPVEDNGVKLLCEVMREPSC HLQDLELVKCHLTAACCESLSCVISRSRHLKSLDLTDNALGDGGVAALCEGLKQKNSV
Further analysis of the NOV125a protein yielded the following properties shown in Table 125B.
Table 125B. Protein Sequence Properties NOVl 25a
PSort 0.7600 probability located in nucleus; 0.3000 probability located in microbody analysis: (peroxisome); 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOV125a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 125C.
In a BLAST search of public sequence databases, the NOV125a protein was found to have homology to the proteins shown in the BLASTP data in Table 125D.
PFam analysis predicts that the NOV125a protein contains the domains shown in the Table 125E.
Example 126.
The NOVl 26 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 126 A.
Table 126A. NOVl 26 Sequence Analysis
SEQ ED NO: 349 2310 bp
NOVl 26a, CCGCGCCTCAGTCCGCCGTCCGCCCTCCGCGCCCGCGCCGCTAGCATGACCGACGCGC
TGTTGCCCGCGGCCCCCCAGCCGCTGGAGAAGGAGAACGACGGCTACTTTCGGAAGGG
CG59987-01 DNA Sequence CTGTAATCCCCTTGCACAAACCGGCCGGAGTAAATTGCAGAATCAAAGAGCTGCTTTG AATCAGCAGATCCTGAAAGCCGTGCGGATGAGGACCGGAGCGGAAAACCTTCTGAAAG TGGCCACAAACTCAAAGGTGCGGGAGCAAGTGCGGCTGGAGCTGAGCTTCGTCAACTC AGACCTGCAGATGCTCAAGGAAGAGCTGGAGGGGCTGAACATCTCGGTGGGCGTCTAT CAGAACACAGAGGAGGCATTTACGATTCCCCTGATTCCTCTTGGCCTGAAGGAAACGA AAGACGTCGACTTTGCAGTCGTCCTCAAGGATTTTATCCTGGAACATTACAGTGAAGA TGGCTATTTATATGAAGATGAAATTGCAGATCTTATGGATCTGAGACAAGCTTGTCGG ACGCCTAGCCGGGATGAGGCCGGGGTGGAACTGCTGATGACATACTTCATCCAGCTGG GCTTTGTCGAGAGTCGATTCTTCCCGCCCACACGGCAGATGGGACTCCTGTTCACCTG GTATGACTCTCTCACCGGGGTTCCGGTCAGCCAGCAGAACCTGCTGCTGGAGAAGGCC AGTGTCCTGTTCAACACTGGGGCCCTCTACACCCAGATTGGGACCCGGTGTGATCGGC AGACGCAGGCTGGGCTGGAGAGTGCCATAGATGCCTTTCAGAGAGCCGCAGGGGTTTT AAATTACCTGAAAGACACATTTACCCATACTCCAAGTTACGACATGAGCCCTGCCATG CTCAGCGTGCTCGTCAAAATGATGCTTGCACAAGCCCAAGAAAGCGTGTTTGAGAAAA TCAGCCTTCCTGGGATCCGGAATGAATTCTTCATGCTGGTGAAGGTGGCTCAGGAGGC TGCTAAGGTGGGAGAGGTCTACCAACAGCTACACGCAGCCATGAGCCAGGCGCCGGTG AAAGAGAACATCCCCTACTCCTGGGCCAGCTTAGCCTGCGTGAAGGCCCACCACTACG CGGCCCTGGCCCACTACTTCACTGCCATCCTCCTCATCGACCACCAGGTGAAGCCAGG CACGGATCTGGACCACCAGGAGAAGTGCCTGTCCCAGCTCTACGACCACATGCCAGAG GGGCTGACACCCTTGGCCACACTGAAGAATGATCAGCAGCGCCGACAGCTGGGGAAGT CCCACTTGCGCAGAGCCATGGCTCATCACGAGGAGTCGGTGCGGGAGGCCAGCCTCTG CAAGAAGCTGCGGAGCATTGAGGTGCTACAGAAGGTGCTGTGTGCCGCACAGGAACGC TCCCGGCTCACGTACGCCCAGCACCAGGAGGAGGATGACCTGCTGAACCTGATCGACG CCCCCAGAGTGTTGTTGCTAAAACTGAGCAAGAGGTTGACATTATATTGCCCCAGTTC TCCAGCTGACAGTCACGGACTTCTTCCAGAAGCTGGGCCCTTATCTGTGCTGTCGGCT AACAAGCGGTGGACGCCTCCTCGAAGCATCCGCTTCACTGCAGAAGAAGGGGACTTGG GGTTCACCTTGAGAGGGAACGCCCCCGTTCAGGTTCACTTCCTGGATCCTTACTGCTC TGCCTCGGTGGCAGGAGCCCGGGAAGGAGATTATATTGTCTCCATTCAGCTTGTGGAT TGTAAGTGGCTGACGCTGAGTGAGGTTATGAAGCTGCTGAAGAGCTTTGGCGAGGACG AGATCGAGATGAAAGTCGTGAGCCTCCTGGACTCCACATCATCCATGCATAATAAGAG TGCCACATACTCCGTGGGAATGCAGAAAACGTACTCCATGATCTGCTTAGCCATTGAT GATGACGACAAAACTGATAAAACCAAGAAAATCTCCAAGAAGCTTTCCTTCCTGAGTT GGGGCACCAACAAGAACAGACAGAAGTCAGCCAGCACCTTGTGCCTCCCATCGGTCGG GGCTGCACGGCCTCAGGTCAAGAAGAAGCTGCCCTCCCCTTTCAGCCTTCTCAACTCA GACAGTTCTTGGTACTAATGTGAGGAAACAAACATGTTCAGGCCCCGAACATTTCCGG
TGCTGACTCGGCCTTAAACGTTTGTGCCATAATGGAAAATATCTATCTATCTGTTCTC
AAATCCTGTTTTTCTCATAGTGTAAACTCACATTTGATGTGTTTTTATGAAGGAAAGT
AACCAAGAAACCTCTAGGAATTAGTGAAAAAAGAACTTTTTTGAGGTG
ORF Start: ATG at 46 ORF Stop: TAA at 2104
SEQ ED NO: 350 |686 aa~ MW at 76812.3kD
NOV126a, MTDALLPAAPQPLEKENDGYFRKGCNPLAQTGRSKLQNQRAALNQQILKAVRMRTGAE NLLKVATNSKVREQVRLELSFV SDLQMLKEELEGLNISVGVYQNTEEAFTIPLIPLG
CG59987-01 Protein Sequence LKETKDVDFAWLKDFILEHYSEDGYLYEDEIADLMDLRQACRTPSRDEAGVELLMTY FIQLGFVESRFFPPTRQMGLLFTWYDSLTGVPVSQQNLLLEKASVLFNTGALYTQIGT RCDRQTQAGLESAIDAFQRAAGVLNYLKDTFTHTPSYDMSPAMLSVLVKMMLAQAQES VFEKISLPGIRNEFFMLVKVAQEAAKVGEVYQQLHAAMSQAPVKENIPYSWASLACVK AHHYAALAHYFTAILLIDHQVKPGTDLDHQEKCLSQLYDHMPEGLTPLATLK DQQRR QLGKSHLRRAMAHHEESVREASLCKKLRSIEVLQKVLCAAQERSRLTYAQHQEEDDLL NLIDAPRVLLLKLSKRLTLYCPSSPADSHGLLPEAGPLSVLSANKR TPPRSIRFTAE
EGDLGFTLRGNAPVQVHFLDPYCSASVAGAREGDYIVSIQLVDCK LTLSEVMKLLKS FGEDEIEMKWSLLDSTSSMHNKSATYSVGMQKTYSMICLAIDDDDKTDKTKKISKKL SFLS GTNKNRQKSASTLCLPSVGAARPQVKKKLPSPFSLLNSDSS Y
SEQ ED NO: 351 2109 bp
NOVl 26b, CGCCGCTAGCATGACCGACGCGCTGTTGCCCGCGGCCCCCCAGCCGCTGGAGAAGGAG
AACGACGGCTACTTTCGGAAGGGCTGTAATCCCCTTGCACAAACCGGCCGGAGTAAAT
CG59987-02 DNA Sequence TGCAGAATCAAAGAGCTGCTTTGAATCAGCAGATCCTGAAAGCCGTGCGGATGAGGAC CGGAGCGGAAAACCTTCTGAAAGTGGCCACAAACTCAAAGGTGCGGGAGCAAGTGCGG CTGGAGCTGAGCTTCGTCAACTCAGACCTGCAGATGCTCAAGGAAGAGCTGGAGGGGC TGAACATCTCGGTGGGCGTCTATCAGAACACAGAGGAGGCATTTACGATTCCCCTGAT TCCTCTTGGCCTGAAGGAAACGAAAGACGTCGACTTTGCAGTCGTCCTCAAGGATTTT ATCCTGGAACATTACAGTGAAGATGGCTATTTATATGAAGATGAAATTGCAGATCTTA TGGATCTGAGACAAGCTTGTCGGACGCCTAGCCGGGATGAGGCCGGGGTGGAACTGCT GATGACATACTTCATCCAGCTGGGCTTTGTCGAGAGTCGATTCTTCCCGCCCACACGG CAGATGGGACTCCTGTTCACCTGGTATGACTCTCTCACCGGGGTTCCGGTCAGCCAGC AGAACCTGCTGCTGGAGAAGGCCAGTGTCCTGTTCAACACTGGGGCCCTCTACACCCA GATTGGGACCCGGTGCGATCGGCAGACGCAGGCTGGGCTGGAGAGTGCCATAGATGCC TTTCAGAGAGCCGCAGGGGTTTTAAATTACCTGAAAGACACATTTACCCATACTCCAA GTTACGACATGAGCCCTGCCATGCTCAGCGTGCTCGTCAAAATGATGCTTGCACAAGC CCAAGAAAGCGTGTTTGAGAAAATCAGCCTTCCTGGGATCCGGAATGAATTCTTCATG CTGGTGAAGGTGGCTCAGGAGGCTGCTAAGGTGGGAGAGGTCTACCAACAGCTACACG CAGCCATGAGCCAGGCGCCGGTGAAAGAGAACATCCCCTACTCCTGGGCCAGCTTAGC CTGCGTGAAGGCCCACCACTACGCGGCCCTGGCCCACTACTTCACTGCCATCCTCCTC ATCGACCACCAGGTGAAGCCAGGCACGGATCTGGACCACCAGGAGAAGTGCCTGTCCC AGCTCTACGACCACATGCCAGAGGGGCTGACACCCTTGGCCACACTGAAGAATGATCA GCAGCGCCGACAGCTGGGGAAGTCCCACTTGCGCAGAGCCATGGCTCATCACGAGGAG TCGGTGCGGGAGGCAAGCCTCTGCAAGAAGCTGCGGAGCATTGAGGTGCTACAGAAGG TGCTGTGTGCCGCACAGGAACGCTCCCGGCTCACGTACGCCCAGCACCAGGAGGAGGA TGACCTGCTGAACCTGATCGACGCCCCCAGTGTTGTTGCTAAAACTGAGCAAGAGGTT GACATTATATTGCCCCAGTTCTCCAAGCTGACAGTCACGGACTTCTTCCAGAAGCTGG GCCCCTTATCTGTGTTTTCGGCTAACAAGCGGTGGACGCCTCCTCGAAGCATCCGCTT CACTGCAGAAGAAGGGGACTTGGGGTTCACCTTGAGAGGGAACGCCCCCGTTCAGGTT CACTTCCTGGATCCTTACTGCTCTGCCTCGGTGGCAGGAGCCCGGGAAGGAGATTATA TTGTCTCCATTCAGCTTGTGGATTGTAAGTGGCTGACGCTGAGTGAGGTTATGAAGCT GCTGAAGAGCTTTGGCGAGGACGAGATCGAGATGAAAGTCGTGAGCCTCCTGGACTCC ACATCATCCATGCATAATAAGAGTGCCACATACTCCGTGGGAATGTAGAAAACGTACT CCATGATCTGCTTAGCCATTGATGATGACGACAAAACTGATAAAACCAAGAAAATCTC
CAAGAAGCTTTCCTTCCTGAGTTGGGGCACCAACAAGAACAGACAGAAGTCAGCCAGC
ACCTTGTGCCTCCCATCGGTCGGGGCTGCACGGCCTCAGGTCAAGAAGAAGCTGCCCT
CCCCTTTCAGCCTTCTCAACTCAGACAGTTCTTGGTACTAATGTGAGGAAACAAACAT
GTTCAGGCCCCGAACATTTCC
ORF Start: ATG at 11 ORF Stop: TAG at 1844
SEQ ED NO: 352 611 aa MW at 68613.9kD
NOV126b, MTDALLPAAPQPLEKENDGYFRKGCNPLAQTGRSKLQNQRAALNQQILKAVRMRTGAE NLLKVATNSKVREQVRLELSFVNSDLQMLKEELEGLNISVGVYQNTEEAFTIPLIPLG
CG59987-02 Protein Sequence LKETKDVDFAWLKDFILEHYSEDGYLYEDEIADLMDLRQACRTPSRDEAGVELLMTY FIQLGFVESRFFPPTRQMGLLFTWYDSLTGVPVSQQNLLLEKASVLFNTGALYTQIGT RCDRQTQAGLESAIDAFQRAAGVL YLKDTFTHTPSYDMSPAMLSVLVKMMLAQAQES VFEKISLPGIRNEFFMLVKVAQEAAKVGEVYQQLHAAMSQAPVKENIPYS ASLACVK AHHYAALAHYFTAILLIDHQVKPGTDLDHQEKCLSQLYDHMPEGLTPLATLKNDQQRR QLGKSHLRRAMAHHEESVREASLCKKLRSIEVLQKVLCAAQERSRLTYAQHQEEDDLL NLIDAPSWAKTEQEVDIILPQFSKLTVTDFFQKLGPLSVFSANKR TPPRSIRFTAE EGDLGFTLRGNAPVQVHFLDPYCSASVAGAREGDYIVSIQLVDCKWLTLSEVMKLLKS FGEDEIEMKWSLLDSTSSMHNKSATYSVGM
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 126B.
Further analysis of the NOVl 26a protein yielded the following properties shown in Table 126C.
Table 126C. Protein Sequence Properties NOV126a
PSort 0.4500 probability located in cytoplasm; 0.3000 probability located in microbody analysis: (peroxisome); 0.1000 probability located in mitochondrial matrix space; 0.1000 probability located in lysosome (lumen)
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVl 26a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 126D.
In a BLAST search of public sequence databases, the NOVl 26a protein was found to have homology to the proteins shown in the BLASTP data in Table 126E.
PFam analysis predicts that the NOVl 26a protein contains the domains shown in the Table 126F.
Example 127.
The NOVl 27 clone was analyzed, and the nucleotide and predicted polypeptide sequences are shown in Table 127 A.
Table 127A. NOVl 27 Sequence Analysis
SEQ ED NO: 353 3351 bp
NOVl 27a, CGTCCCGTGGCCATGACGACCGCTCAGAGGGACTCCCTGTTGTGGAAGCTCGCGGGGT
TGCTGCGGGAGTCCGGTGATGTGGTCCTGTCTGGCTGTAGCACCCTGAGCCTGCTGAC
CG59971-01 DNA Sequence TCCCACACTGCAACAGCTGAACCACGTATTTGAGCTGCACCTGGGGCCATGGGGCCCT GGCCAGACAGGCTTTGTGGCTCTGCCCTCCCATCCTGCCGACTCCCCTGTTATTCTTC AGCTTCAGTTTCTCTTCGATGTGCTGCAGAAAACACTTTCACTCAAGCTGGTCCATGT TGCTGGTCCTGGCCCCACAGGGCCCATCAAGATTTTCCCCTTCAAATCCCTTCGGCAC CTGGAGCTCCGAGGTGTTCCCCTCCACTGTCTGCATGGCCTCCGAGGCATCTACTCCC AGCTGGAGACCCTGATTTGCAGCAGGAGCCTCCAGGCATTAGAGGAGCTCCTCTCAGC CTGCGGCGGCGACTTCTGCTCTGCCCTCCCTTGGCTGGCTCTGCTTTCTGCCAACTTC AGCTACAATGCACTGACCGCCTTAGACAGCTCCCTGCGCCTCTTGTCAGCTCTGCGTT TCTTGAACCTAAGCCACAATCAAGTCCAGGACTGTCAGGGATTCCTGATGGATTTGTG TGAGCTCCACCATCTGGACATCTCCTATAATCGCCTGCATTTGGTGCCAAGAATGGGA CCCTCAGGGGCTGCTCTGGGGGTCCTGATACTGCGAGGCAATGAGCTTCGGAGCCTGC CAGGCCTAGAGCAGCTGAGGAATCTGCGGCACCTGGATTTGGCATACAACCTGCTGGA AGGACACCGGGAGCTGTCACCACTGTGGCTGCTGGCTGAGCTCCGCAAGCTCTACCTG GAGGGGAACCCTCTTTGGTTCCACCCTGAGCACCGAGCAGCCACTGCCCAGTACTTGT CACCCCGGGCCAGGGATGCTGCTACTGGCTTCCTTCTCGATGGCAAGGTCTTGTCACT GACAGATTTTCAGCAGACTCACACATCCTTGGGGCTCAGCCCCATGGGCCCACCTTTG CCCTGGCCAGTGGGGAGTACTCCTGAAACCTCAGGTGGCCCTGACCTGAGTGACAGCC TCTCCTCAGGGGGTGTTGTGACCCAGCCCCTGCTTCATAAGGTTAAGAGCCGAGTCCG TGTGAGGCGGGCAAGCATCTCTGAACCCAGTGATACGGACCCGGAGCCCCGAACTCTG AACCCCTCTCCGGCTGGTTGGTTCGTGCAGCAGCACCCGGAGCTGGAGCTCATGAGCA GCTTCCGGGAACGGTTCGGCCGCAACTGGCTGCAGTACAGGAGTCACCTGGAGCCCTC CGGAAACCCTCTGCCGGCCACCCCCACTACTTCTGCACCCAGTGCACCTCCAGCCAGC TCCCAGGGCCCCGACACTGCACCCAGACCTTCACCCCCGCAGGAGGAAGCCAGAGGCC CCCAGGAGTCACCACAGAAAATGTCAGAGGAGGTCAGGGCGGAGCCACAGGAGGAGGA AGAGGAGAAGGAGGGGAAGGAGGAGAAGGAGGAGGGGGAGATGGTGGAACAGGGAGAA GAGGAGGCAGGAGAGGAGGAAGAAGAGGAGCAGGACCAGAAGGAAGTGGAAGCGGAAC TCTGTCGCCCCTTGTTGGTGTGTCCCCTGGAGGGGCCTGAGGGCGTACGGGGCAGGGA ATGCTTTCTCAGGGTCACTTCTGCCCACCTGTTTGAGGTGGAACTCCAAGCAGCTCGC ACCTTGGAGCGACTGGAGCTCCAGAGTCTGGAGGCAGCTGAGATAGAGCCGGAGGCCC AGGCCCAGGGTCCCCCTCTTGCTGCGCAGGGCTCAGATCTGCTCCCTGGAGCCCCCAT CCTCAGTCTGCGCTTCTCCTACATCTGCCCTGACCGGCAGTTGCGTCGCTATTTGGTG CTGGAGCCTGATGCCCACGCAGCTGTCCAGGAGCTGCTTGCCGTGTTGACCCCAGTCA CCAATGTGGCTCGGGAACAGCTTGGGGAGGCCAGGGACCTCCTGCTGGGTAGATTCCA GTGTCTACGCTGTGGCCATGAGTTCAAGCCAGAGGAGCCCAGGATGGGATTAGACAGT GAGGAAGGCTGGAGGCCTCTGTTCCAAAAGACAGAATCTCCTGCTGTGTGTCCTAACT GTGGTAGTGACCACGTGGTTCTCCTCGCTGTGTCTCGGGGAACCCCCAACAGGGAGCG GAAACAGGGAGAGCAGTCTCTGGCTCCTTCTCCGTCTGCCAGCCCTGTCTGCCACCCT CCTGGCCATGGTGACCACCTTGACAGGGCCAAGAACAGCCCACCTCAGGCACCGAGCA CCCGTGACCATGGTAGTTGGAGCCTCAGTCCCGCCCCTGAGCGCTGTGGCCTCCGCTC TGTGGACCACCGACTCCGGCTCTTCCTGGATGTTGAGGTGTTCAGCGATGCCCAGGAG GAGTTCCAGTGCTGCCTCAAGGTCCCAGTGGCATTGGCAGGCCACACTGGGGAGTTCA TGTGCCTTGTGGTTGTGTCTGACCGCAGGCTGTACCTGTTGAAGGTGACTGGGGAGAT GAGTGAGCCTCCAGCTAGCTGGCTGCAGCTGACCCTGGCTGTTCCCCTGCAGGATCTG AGTGGCATAGAGCTGGGCCTGGCAGGCCAGAGCCTGCGGCTAGAGTGGGCAGCTGGGG CGGGCCGCTGTGTGCTGCTGCCCCGAGATGCCAGGCATTGCCGGGCCTTCCTAGAGGA GCTCCTTGGTGTCTTGCAGTCTCTGCCCCCTGCCTGGAGGAACTGTGTCAGTGCCACA GAGGAGGAGGTCACCCCCCAGCACCGGCTCTGGCCATTGCTGGAAAAAGACTCATCCT TGGAGGCTCGCCAGTTCTTCTACCTTCGGGCGTTCCTGGTTGAAGGTGAAGCCTCTGT GCAGCTGATGCTTCCCTCCACCTGCCTCGTATCCCTGTTGCTGACTCCGTCCACCCTG TTCCTGTTAGATGAGGATGCTGCAGGGTCCCCGGCAGAGCCCTCTCCTCCAGCAGCAT CTGGCGAAGCCTCTGAGAAGGTGCCTCCCTCGGGGCCGGGCCCTGCTGTGCGTGTCAG GGAGCAGCAGCCACTCAGCAGCCTGAGCTCCGTGCTGCTCTACCGCTCAGCCCCTGAG GACTTGCGGCTGCTCTTCTACGATGAGGTGTCCCGGCTGGAGAGCTTTTGGGCACTCC GTGTGGTGTGTCAGGAGCAGCTGACAGCCCTGCTTGCCTGGATCCGGGAACCATGGGA GGAGCTGTTTTCCATCGGACTCCGGACAGTGATCCAAGAGGCGCTGGCCCTTGACCGA TGAGGGTCCCACGCTGACCTTGGCCCTGACCTCAGGAGCCACGCT
ORF Start: ATG at 13JORF Stop: TGA at 3307
SEQ ED NO: 354 1098 aa MW at l21004.1kD
NOVl 27a, MTTAQRDSLL KLAGLLRESGDWLSGCSTLSLLTPTLQQLNHVFELHLGP GPGQTG
CG59971-01 Protein Sequence FVALPSHPADSPVILQLQFLFDVLQKTLSLKLVHVAGPGPTGPIKIFPFKSLRHLELR GVPLHCLHGLRGIYSQLETLICSRSLQALEELLSACGGDFCSALP LALLSANFSYNA LTALDSSLRLLSALRFLNLSHNQVQDCQGFLMDLCELHHLDISY RLHLVPRMGPSGA ALGVLILRGNELRSLPGLEQLR LRHLDLAYNLLEGHRELSPL LLAELRKLYLEGNP LWFHPEHRAATAQYLSPRARDAATGFLLDGKVLSLTDFQQTHTSLGLSPMGPPLP PV GSTPETSGGPDLSDSLSSGGWTQPLLHKVKSRVRVRRASISEPSDTDPEPRTLNPSP AG FVQQHPELELMSSFRERFGRN LQYRSHLEPSGNPLPATPTTSAPSAPPASSQGP DTAPRPSPPQEEARGPQESPQKMSEEVRAEPQEEEEEKEGKEEKEEGEMVEQGEEEAG EEEEEEQDQKEVEAELCRPLLVCPLEGPEGVRGRECFLRVTSAHLFEVELQAARTLER LELQSLEAAEIEPEAQAQGPPLAAQGSDLLPGAPILSLRFSYICPDRQLRRYLVLEPD AHAAVQELLAVLTPVTNVAREQLGEARDLLLGRFQCLRCGHEFKPEEPRMGLDSEEGW RPLFQKTESPAVCPNCGSDHWLLAVSRGTPNRERKQGEQSLAPSPSASPVCHPPGHG DHLDRAKNSPPQAPSTRDHGS SLSPAPERCGLRSVDHRLRLFLDVEVFSDAQEEFQC CLKVPVALAGHTGEFMCLVWSDRRLYLLKVTGEMSEPPASWLQLTLAVPLQDLSGIE LGLAGQSLRLEWAAGAGRCVLLPRDARHCRAFLEELLGVLQSLPPAWRNCVSATEEEV TPQHRL PLLEKDSSLEARQFFYLRAFLVEGEASVQLMLPSTCLVSLLLTPSTLFLLD EDAAGSPAEPSPPAASGEASEKVPPSGPGPAVRVREQQPLSSLSSVLLYRSAPEDLRL LFYDEVSRLESF ALRWCQEQLTALLAWIREP EELFSIGLRTVIQEALALDR
SEQ ED NO: 355 3348 bp
NOVl 27b, CGTCCCGTGGCCATGACGACCGCTCAGAGGGACTCCCTGTTGTGGAAGCTCGCGGGGT
TGCTGCGGGAGTCCGGTGATGTGGTCCTGTCTGGCTGTAGCACCCTGAGCCTGCTGAC
CG59971-02 DNA Sequence TCCCACACTGCAACAGCTGAACCACGTATTTGAGCTGCACCTGGGGCCATGGGGCCCT GGCCAGACAGGCTTTGTGGCTCTGCCCTCCCATCCTGCCGACTCCCCTGTTATTCTTC AGCTTCAGTTTCTCTTCGATGTGCTGCAGAAAACACTTTCACTCAAGCTGGTCCATGT TGCTGGTCCTGGCCCCACAGGGCCCATCAAGATTTTCCCCTTCAAATCCCTTCGGCAC CTGGAGCTCCGAGGTGTTCCCCTCCACTGTCTGCATGGCCTCCGAGGCATCTACTCCC AGCTGGAGACCCTGATTTGCAGCAGGAGCCTCCAGGCATTAGAGGAGCTCCTCTCAGC CTGCGGCGGCGACTTCTGCTCTGCCCTCCCTTGGCTGGCTCTGCTTTCTGCCAACTTC AGCTACAATGCACTGACCGCCTTAGACAGCTCCCTGCGCCTCTTGTCAGCTCTGCGTT TCTTGAACCTAAGCCACAATCAAGTCCAGGACTGTCAGGGATTCCTGATGGATTTGTG TGAGCTCCACCATCTGGACATCTCCTATAATCGCCTGCATTTGGTGCCAAGAATGGGA CCCTCAGGGGCTGCTCTGGGGGTCCTGATACTGCGAGGCAATGAGCTTCGGAGCCTGC CAGGCCTAGAGCAGCTGAGGAATCTGCGGCACCTGGATTTGGCATACAACCTGCTGGA AGGACACCGGGAGCTGTCACCACTGTGGCTGCTGGCTGAGCTCCGCAAGCTCTACCTG GAGGGGAACCCTCTTTGGTTCCACCCTGAGCACCGAGCAGCCACTGCCCAGTACTTGT CACCCCGGGCCAGGGATGCTGCTACTGGCTTCCTTCTCGATGGCAAGGTCTTGTCACT GACAGATTTTCAGCAGACTCACACATCCTTGGGGCTCAGCCCCATGGGCCCACCTTTG CCCTGGCCAGTGGGGAGTACTCCTGAAACCTCAGGTGGCCCTGACCTGAGTGACAGCC TCTCCTCAGGGGGTGTTGTGACCCAGCCCCTGCTTCATAAGGTTAAGAGCCGAGTCCG TGTGAGGCGGGCAAGCATCTCTGAACCCAGTGATACGGACCCGGAGCCCCGAACTCTG AACCCCTCTCCGGCTGGTTGGTTCGTGCAGCAGCACCCGGAGCTGGAGCTCATGAGCA GCTTCCGGGAACGGTTCGGCCGCAACTGGCTGCAGTACAGGAGTCACCTGGAGCCCTC CGGAAACCCTCTGCCGGCCACCCCCACTACTTCTGCACCCAGTGCACCTCCAGCCAGC TCCCAGGGCCCCGACACTGCACCCAGACCTTCACCCCCGCAGGAGGAAGCCAGAGGCC CCCAGGAGTCACCACAGAAAATGTCAGAGGAGGTCAGGGCGGAGCCACAGGAGGAGGA AGAGGAGAAGGAGGGGAAGGAGGAGAAGGAGGAGGGGGAGATGGTGGAACAGGGAGAA GAGGAGGCAGGAGAGGAGGAAGAAGAGGAGCAGGACCAGAAGGAAGTGGAAGCGGAAC TCTGTCGCCCCTTGTTGGTGTGTCCCCTGGAGGGGCCTGAGGGCGTACGGGGCAGGGA ATGCTTTCTCAGGGTCACTTCTGCCCACCTGTTTGAGGTGGAACTCCAAGCAGCTCGC ACCTTGGAGCGACTGGAGCTCCAGAGTCTGGAGGCAGCTGAGATAGAGCCGGAGGCCC AGGCCCAGAGGTCGCCCAGGCCCACGGGCTCAGATCTGCTCCCTGGAGCCCCCATCCT CAGTCTGCGCTTCTCCTACATCTGCCCTGACCGGCAGTTGCGTCGCTATTTGGTGCTG GAGCCTGATGCCCACGCAGCTGTCCAGGAGCTGCTTGCCGTGTTGACCCCAGTCACCA ATGTGGCTCGGGAACAGCTTGGGGAGGCCAGGGACCTCCTGCTGGGTAGATTCCAGTG TCTACGCTGTGGCCATGAGTTCAAGCCAGAGGAGCCCAGGATGGGATTAGACAGTGAG GAAGGCTGGAGGCCTCTGTTCCAAAAGACAGAATCTCCTGCTGTGTGTCCTAACTGTG GTAGTGACCACGTGGTTCTCCTCGCTGTGTCTCGGGGAACCCCCAACAGGGAGCGGAA ACAGGGAGAGCAGTCTCTGGCTCCTTCTCCGTCTGCCAGCCCTGTCTGCCACCCTCCT GGCCATGGTGACCACCTTGACAGGGCCAAGAACAGCCCACCTCAGGCACCGAGCACCC GTGACCATGGTAGTTGGAGCCTCAGTCCCGCCCCTGAGCGCTGTGGCCTCCGCTCTGT GGACCACCGACTCCGGCTCTTCCTGGATGTTGAGGTGTTCAGCGATGCCCAGGAGGAG TTCCAGTGCTGCCTCAAGGTCCCAGTGGCATTGGCAGGCCACACTGGGGAGTTCATGT GCCTTGTGGTTGTGTCTGACCGCAGGCTGTACCTGTTGAAGGTGACTGGGGAGATGAG TGAGCCTCCAGCTAGCTGGCTGCAGCTGACCCTGGCTGTTCCCCTGCAGGATCTGAGT GGCATAGAGCTGGGCCTGGCAGGCCAGAGCCTGCGGCTAGAGTGGGCAGCTGGGGCGG GCCGCTGTGTGCTGCTGCCCCGAGATGCCAGGCATTGCCGGGCCTTCCTAGAGGAGCT CCTTGGTGTCTTGCAGTCTCTGCCCCCTGCCTGGAGGAACTGTGTCAGTGCCACAGAG GAGGAGGTCACCCCCCAGCACCGGCTCTGGCCATTGCTGGAAAAAGACTCATCCTTGG AGGCTCGCCAGTTCTTCTACCTTCGGGCGTTCCTGGTTGAAGGTGAAGCCTCTGTGCA GCTGATGCTTCCCTCCACCTGCCTCGTATCCCTGTTGCTGACTCCGTCCACCCTGTTC CTGTTAGATGAGGATGCTGCAGGGTCCCCGGCAGAGCCCTCTCCTCCAGCAGCATCTG GCGAAGCCTCTGAGAAGGTGCCTCCCTCGGGGCCGGGCCCTGCTGTGCGTGTCAGGGA GCAGCAGCCACTCAGCAGCCTGAGCTCCGTGCTGCTCTACCGCTCAGCCCCTGAGGAC TTGCGGCTGCTCTTCTACGATGAGGTGTCCCGGCTGGAGAGCTTTTGGGCACTCCGTG TGGTGTGTCAGGAGCAGCTGACAGCCCTGCTTGCCTGGATCCGGGAACCATGGGAGGA
Sequence comparison of the above protein sequences yields the following sequence relationships shown in Table 127B.
Further analysis of the NOVl 27a protein yielded the following properties shown in Table 127C.
Table 127C. Protein Sequence Properties NOVl 27a
PSort 0.5163 probability located in mitochondrial matrix space; 0.3000 probability analysis: located in microbody (peroxisome); 0.2442 probability located in mitochondrial inner membrane; 0.2442 probability located in mitochondrial intermembrane space
SignalP No Known Signal Sequence Predicted analysis:
A search of the NOVl 27a protein against the Geneseq database, a proprietary database that contains sequences published in patents and patent publication, yielded several homologous proteins shown in Table 127D.
Table 127D. Geneseq Results for NOV127a
In a BLAST search of public sequence databases, the NOVl 27a protein was found to have homology to the proteins shown in the BLASTP data in Table 127E.
musculus (Mouse), 1072 aa. 1..1072 895/1098 (81%) j
Q9VMK9 CG9044 PROTEIN - Drosophila 12..433 139/459 (30%) j 6e-38 melanogaster (Fruit fly), 1289 aa. 8-463 220/459 (47%)
PFam analysis predicts that the NOVl 27a protein contains the domains shown in the Table 127F.
Example B: Sequencing Methodology and Identification of NOVX Clones
1. GeneCalling™ Technology: This is a proprietary method of performing differential gene expression profiling between two or more samples developed at CuraGen and described by Shimkets, et al., "Gene expression analysis by transcript profiling coupled to a gene database query" Nature Biotechnology 17:198-803 (1999). cDNA was derived from various
human samples representing multiple tissue types, normal and diseased states, physiological states, and developmental states from different donors. Samples were obtained as whole tissue, primary cells or tissue cultured primary cells or cell lines. Cells and cell lines may have been treated with biological or chemical agents that regulate gene expression, for example, growth factors, chemokines or steroids. The cDNA thus derived was then digested with up to as many as 120 pairs of restriction enzymes and pairs of linker-adaptors specific for each pair of restriction enzymes were ligated to the appropriate end. The restriction digestion generates a mixture of unique cDNA gene fragments. Limited PCR amplification is performed with primers homologous to the linker adapter sequence where one primer is biotinylated and the other is fluorescently labeled. The doubly labeled material is isolated and the fluorescently labeled single strand is resolved by capillary gel electrophoresis. A computer algorithm compares the electropherograms from an experimental and control group for each of the restriction digestions. This and additional sequence-derived information is used to predict the identity of each differentially expressed gene fragment using a variety of genetic databases. The identity of the gene fragment is confirmed by additional, gene-specific competitive PCR or by isolation and sequencing of the gene fragment.
2. SeqCalling™ Technology: cDNA was derived from various human samples representing multiple tissue types, normal and diseased states, physiological states, and developmental states from different donors. Samples were obtained as whole tissue, primary cells or tissue cultured primary cells or cell lines. Cells and cell lines may have been treated with biological or chemical agents that regulate gene expression, for example, growth factors, chemokines or steroids. The cDNA thus derived was then sequenced using CuraGen's proprietary SeqCalling technology. Sequence traces were evaluated manually and edited for corrections if appropriate. cDNA sequences from all samples were assembled together, sometimes including public human sequences, using bioinformatic programs to produce a consensus sequence for each assembly. Each assembly is included in CuraGen Coφoration's database. Sequences were included as components for assembly when the extent of identity with another component was at least 95%> over 50 bp. Each assembly represents a gene or portion thereof and includes information on variants, such as splice forms single nucleotide polymorphisms (SNPs), insertions, deletions and other sequence variations.
3. PathCalling™ Technology:
The NOVX nucleic acid sequences are derived by laboratory screening of cDNA library by the two-hybrid approach. cDNA fragments covering either the full length of the DNA sequence, or part of the sequence, or both, are sequenced. In silico prediction was based on sequences available in CuraGen Corporation's proprietary sequence databases or in the public human sequence databases, and provided either the full length DNA sequence, or some portion thereof.
The laboratory screening was performed using the methods summarized below:
cDNA libraries were derived from various human samples representing multiple tissue types, normal and diseased states, physiological states, and developmental states from different donors. Samples were obtained as whole tissue, primary cells or tissue cultured primary cells or cell lines. Cells and cell lines may have been treated with biological or chemical agents that regulate gene expression, for example, growth factors, chemokines or steroids. The cDNA thus derived was then directionally cloned into the appropriate two-hybrid vector (Gal4-activation domain (Gal4-AD) fusion). Such cDNA libraries as well as commercially available cDNA libraries from Clontech (Palo Alto, CA) were then transferred from E.coli into a CuraGen Coφoration proprietary yeast strain (disclosed in U. S. Patents 6,057,101 and 6,083,693, incoφorated herein by reference in their entireties).
Gal4-binding domain (Gal4-BD) fusions of a CuraGen Coφortion proprietary library of human sequences was used to screen multiple Gal4-AD fusion cDNA libraries resulting in the selection of yeast hybrid diploids in each of which the Gal4-AD fusion contains an individual cDNA. Each sample was amplified using the polymerase chain reaction (PCR) using non-specific primers at the cDNA insert boundaries. Such PCR product was sequenced; sequence traces were evaluated manually and edited for corrections if appropriate. cDNA sequences from all samples were assembled together, sometimes including public human sequences, using bioinformatic programs to produce a consensus sequence for each assembly. Each assembly is included in CuraGen Coφoration's database. Sequences were included as components for assembly when the extent of identity with another component was at least 95%o over 50 bp. Each assembly represents a gene or portion thereof and includes information on variants, such as splice forms single nucleotide polymoφhisms (SNPs), insertions, deletions and other sequence variations.
Physical clone: the cDNA fragment derived by the screening procedure, covering the entire open reading frame is, as a recombinant DNA, cloned into pACT2 plasmid (Clontech) used to make the cDNA library. The recombinant plasmid is inserted into the host and selected by the yeast hybrid diploid generated during the screening procedure by the mating of both CuraGen Coφoration proprietary yeast strains N106' and YULH (U. S. Patents 6,057,101 and 6,083,693).
4. RACE: Techniques based on the polymerase chain reaction such as rapid amplification of cDNA ends (RACE), were used to isolate or complete the predicted sequence of the cDNA of the invention. Usually multiple clones were sequenced from one or more human samples to derive the sequences for fragments. Various human tissue samples from different donors were used for the RACE reaction. The sequences derived from these procedures were included in the SeqCalling Assembly process described in preceding paragraphs.
5. Exon Linking: The NOVX target sequences identified in the present invention were subjected to the exon linking process to confirm the sequence. PCR primers were designed by starting at the most upstream sequence available, for the forward primer, and at the most downstream sequence available for the reverse primer. Table BI shows the sequences of the PCR primers used for obtaining different clones. In each case, the sequence was examined, walking inward from the respective termini toward the coding sequence, until a suitable sequence that is either unique or highly selective was encountered, or, in the case of the reverse primer, until the stop codon was reached. Such primers were designed based on in silico predictions for the full length cDNA, part (one or more exons) of the DNA or protein sequence of the target sequence, or by translated homology of the predicted exons to closely related human sequences from other species. These primers were then employed in PCR amplification based on the following pool of human cDNAs: adrenal gland, bone marrow, brain - amygdala, brain - cerebellum, brain - hippocampus, brain - substantia nigra, brain - fhalamus, brain -whole, fetal brain, fetal kidney, fetal liver, fetal lung, heart, kidney, lymphoma - Raji, mammary gland, pancreas, pituitary gland, placenta, prostate, salivary gland, skeletal muscle, small intestine, spinal cord, spleen, stomach, testis, thyroid, trachea, uterus. Usually the resulting amplicons were gel purified, cloned and sequenced to high redundancy. The PCR product derived from exon linking was cloned into the pCR2.1 vector from Invitrogen. The resulting bacterial clone has an insert covering the entire open reading
frame cloned into the pCR2.1 vector. The resulting sequences from all clones were assembled with themselves, with other fragments in CuraGen Coφoration's database and with public ESTs. Fragments and ESTs were included as components for an assembly when the extent of their identity with another component of the assembly was at least 95% over 50 bp. In addition, sequence traces were evaluated manually and edited for corrections if appropriate. These procedures provide the sequence reported herein.
6. Physical Clone:
Exons were predicted by homology and the intron exon boundaries were determined using standard genetic rules. Exons were further selected and refined by means of similarity determination using multiple BLAST (for example, tBlastN, BlastX, and BlastN) searches, and, in some instances, GeneScan and Grail. Expressed sequences from both public and proprietary databases were also added when available to further define and complete the gene sequence. The DNA sequence was then manually corrected for apparent inconsistencies thereby obtaining the sequences encoding the full-length protein.
The PCR product derived by exon linking, covering the entire open reading frame, was cloned into the pCR2.1 vector from Invitrogen to provide clones used for expression and screening puφoses.
Example C: Quantitative expression analysis of clones in various cells and tissues
The quantitative expression of various clones was assessed using microtiter plates containing RNA samples from a variety of normal and pathology-derived cells, cell lines and tissues using real time quantitative PCR (RTQ PCR). RTQ PCR was performed on an Applied
Biosystems ABI PRISM® 7700 or an ABI PRISM® 7900 HT Sequence Detection System.
Various collections of samples are assembled on the plates, and referred to as Panel 1
(containing normal tissues and cancer cell lines), Panel 2 (containing samples derived from tissues from normal and cancer sources), Panel 3 (containing cancer cell lines), Panel 4
(containing cells and cell lines from normal tissues and cells related to inflammatory conditions), Panel 5D/5I (containing human tissues and cell lines with an emphasis on metabolic diseases), AI_comprehensive_panel (containing normal tissue and samples from autoimmune diseases), Panel CNSD.01 (containing central nervous system samples from normal and diseased brains) and CNS_neurodegeneration_panel (containing samples from normal and Alzheimer's diseased brains).
RNA integrity from all samples is controlled for quality by visual assessment of agarose gel electropherograms using 28S and 18S ribosomal RNA staining intensity ratio as a guide (2:1 to 2.5:1 28s: 18s) and the absence of low molecular weight RNAs that would be indicative of degradation products. Samples are controlled against genomic DNA contamination by RTQ PCR reactions run in the absence of reverse transcriptase using probe and primer sets designed to amplify across the span of a single exon.
First, the RNA samples were normalized to reference nucleic acids such as constitutively expressed genes (for example, β-actin and GAPDH). Normalized RNA (5 ul) was converted to cDNA and analyzed by RTQ-PCR using One Step RT-PCR Master Mix Reagents (Applied Biosystems; Catalog No. 4309169) and gene-specific primers according to the manufacturer's instructions.
In other cases, non-normalized RNA samples were converted to single strand cDNA (sscDNA) using Superscript II (Invitrogen Coφoration; Catalog No. 18064-147) and random hexamers according to the manufacturer's instructions. Reactions containing up to 10 μg of total RNA were performed in a volume of 20 μl and incubated for 60 minutes at 42°C. This reaction can be scaled up to 50 μg of total RNA in a final volume of 100 μl. sscDNA samples are then normalized to reference nucleic acids as described previously, using IX TaqMan® Universal Master mix (Applied Biosystems; catalog No. 4324020), following the manufacturer's instructions.
Probes and primers were designed for each assay according to Applied Biosystems Primer Express Software package (version I for Apple Computer's Macintosh Power PC) or a similar algorithm using the target sequence as input. Default settings were used for reaction conditions and the following parameters were set before selecting primers: primer concentration = 250 nM, primer melting temperature (Tm) range = 58°-60°C, primer optimal Tm = 59°C, maximum primer difference = 2°C, probe does not have 5'G, probe Tm must be 10°C greater than primer Tm, amplicon size 75bp to lOObp. The probes and primers selected (see below) were synthesized by Synthegen (Houston, TX, USA). Probes were double purified by HPLC to remove uncoupled dye and evaluated by mass spectroscopy to verify coupling of reporter and quencher dyes to the 5' and 3' ends of the probe, respectively. Their final concentrations were: forward and reverse primers, 900nM each, and probe, 200nM.
PCR conditions: When working with RNA samples, normalized RNA from each tissue and each cell line was spotted in each well of either a 96 well or a 384-well PCR plate (Applied Biosystems). PCR cocktails included either a single gene specific probe and primers set, or two multiplexed probe and primers sets (a set specific for the target clone and another gene-specific set multiplexed with the target probe). PCR reactions were set up using TaqMan® One-Step RT-PCR Master Mix (Applied Biosystems, Catalog No. 4313803) following manufacturer's instructions. Reverse transcription was performed at 48°C for 30 minutes followed by amplification/PCR cycles as follows: 95°C 10 min, then 40 cycles of 95°C for 15 seconds, 60°C for 1 minute. Results were recorded as CT values (cycle at which a given sample crosses a threshold level of fluorescence) using a log scale, with the difference in RNA concentration between a given sample and the sample with the lowest CT value being represented as 2 to the power of delta CT. The percent relative expression is then obtained by taking the reciprocal of this RNA difference and multiplying by 100.
When working with sscDNA samples, normalized sscDNA was used as described previously for RNA samples. PCR reactions containing one or two sets of probe and primers were set up as described previously, using IX TaqMan® Universal Master mix (Applied Biosystems; catalog No. 4324020), following the manufacturer's instructions. PCR amplification was performed as follows: 95°C 10 min, then 40 cycles of 95°C for 15 seconds, 60°C for 1 minute. Results were analyzed and processed as described previously.
Panels 1, 1.1, 1.2, and 1.3D
The plates for Panels 1, 1.1, 1.2 and 1.3D include 2 control wells (genomic DNA control and chemistry control) and 94 wells containing cDNA from various samples. The samples in these panels are broken into 2 classes: samples derived from cultured cell lines and samples derived from primary normal tissues. The cell lines are derived from cancers of the following types: lung cancer, breast cancer, melanoma, colon cancer, prostate cancer, CNS cancer, squamous cell carcinoma, ovarian cancer, liver cancer, renal cancer, gastric cancer and pancreatic cancer. Cell lines used in these panels are widely available through the American Type Culture Collection (ATCC), a repository for cultured cell lines, and were cultured using the conditions recommended by the ATCC. The normal tissues found on these panels are comprised of samples derived from all major organ systems from single adult individuals or
fetuses. These samples are derived from the following organs: adult skeletal muscle, fetal skeletal muscle, adult heart, fetal heart, adult kidney, fetal kidney, adult liver, fetal liver, adult lung, fetal lung, various regions of the brain, the spleen, bone marrow, lymph node, pancreas, salivary gland, pituitary gland, adrenal gland, spinal cord, thymus, stomach, small intestine, colon, bladder, trachea, breast, ovary, uterus, placenta, prostate, testis and adipose.
In the results for Panels 1, 1.1, 1.2 and 1.3D, the following abbreviations are used:
ca. = carcinoma,
* = established from metastasis, met = metastasis, s cell var = small cell variant, non-s = non-sm = non-small, squam = squamous, pi. eff = pi effusion = pleural effusion, glio = glioma, astro = astrocytoma, and neuro = neuroblastoma.
General_screening_panel_vl .4
The plates for Panel 1.4 include 2 control wells (genomic DNA control and chemistry control) and 94 wells containing cDNA from various samples. The samples in Panel 1.4 are broken into 2 classes: samples derived from cultured cell lines and samples derived from primary normal tissues. The cell lines are derived from cancers of the following types: lung cancer, breast cancer, melanoma, colon cancer, prostate cancer, CNS cancer, squamous cell carcinoma, ovarian cancer, liver cancer, renal cancer, gastric cancer and pancreatic cancer. Cell lines used in Panel 1.4 are widely available through the American Type Culture Collection (ATCC), a repository for cultured cell lines, and were cultured using the conditions recommended by the ATCC. The normal tissues found on Panel 1.4 are comprised of pools of samples derived from all major organ systems from 2 to 5 different adult individuals or fetuses. These samples are derived from the following organs: adult skeletal muscle, fetal skeletal muscle, adult heart, fetal heart, adult kidney, fetal kidney, adult liver, fetal liver, adult lung, fetal lung, various regions of the brain, the spleen, bone marrow, lymph node, pancreas, salivary gland, pituitary gland, adrenal gland, spinal cord, thymus, stomach, small intestine, colon, bladder, trachea, breast, ovary, uterus, placenta, prostate, testis and adipose. Abbreviations are as described for Panels 1, 1.1, 1.2, and 1.3D.
Panels 2D and 2.2
The plates for Panels 2D and 2.2 generally include 2 control wells and 94 test samples composed of RNA or cDNA isolated from human tissue procured by surgeons working in close cooperation with the National Cancer Institute's Cooperative Human Tissue Network (CHTN) or the National Disease Research Initiative (NDRI). The tissues are derived from human malignancies and in cases where indicated many malignant tissues have "matched margins" obtained from noncancerous tissue just adjacent to the tumor. These are termed normal adjacent tissues and are denoted "NAT" in the results below. The tumor tissue and the "matched margins" are evaluated by two independent pathologists (the surgical pathologists and again by a pathologist at NDRI or CHTN). This analysis provides a gross histopathological assessment of tumor differentiation grade. Moreover, most samples include the original surgical pathology report that provides information regarding the clinical stage of the patient. These matched margins are taken from the tissue surrounding (i.e. immediately proximal) to the zone of surgery (designated "NAT", for normal adjacent tissue, in Table RR). In addition, RNA and cDNA samples were obtained from various human tissues derived from autopsies performed on elderly people or sudden death victims (accidents, etc.). These tissues were ascertained to be free of disease and were purchased from various commercial sources such as Clontech (Palo Alto, CA), Research Genetics, and Invitrogen.
Panel 3D
The plates of Panel 3D are comprised of 94 cDNA samples and two control samples. Specifically, 92 of these samples are derived from cultured human cancer cell lines, 2 samples of human primary cerebellar tissue and 2 controls. The human cell lines are generally obtained from ATCC (American Type Culture Collection), NCI or the German tumor cell bank and fall into the following tissue groups: Squamous cell carcinoma of the tongue, breast cancer, prostate cancer, melanoma, epidermoid carcinoma, sarcomas, bladder carcinomas, pancreatic cancers, kidney cancers, leukemias/lymphomas, ovarian/uterine/cervical, gastric, colon, lung and CNS cancer cell lines. In addition, there are two independent samples of cerebellum. These cells are all cultured under standard recommended conditions and RNA extracted using the standard procedures. The cell lines in panel 3D and 1.3D are of the most common cell lines used in the scientific literature.
Panels 4D, 4R, and 4.1D
Panel 4 includes samples on a 96 well plate (2 control wells, 94 test samples) composed of RNA (Panel 4R) or cDNA (Panels 4D/4.1D) isolated from various human cell lines or tissues related to inflammatory conditions. Total RNA from control normal tissues such as colon and lung (Stratagene, La Jolla, CA) and thymus and kidney (Clontech) was employed. Total RNA from liver tissue from cirrhosis patients and kidney from lupus patients was obtained from BioChain (Biochain Institute, Inc., Hayward, CA). Intestinal tissue for RNA preparation from patients diagnosed as having Crohn's disease and ulcerative colitis was obtained from the National Disease Research Interchange (NDRI) (Philadelphia, PA).
Astrocytes, lung fibroblasts, dermal fibroblasts, coronary artery smooth muscle cells, small airway epithelium, bronchial epithelium, microvascular dermal endothelial cells, microvascular lung endothelial cells, human pulmonary aortic endothelial cells, human umbilical vein endothelial cells were all purchased from Clonetics (Walkersville, MD) and grown in the media supplied for these cell types by Clonetics. These primary cell types were activated with various cytokines or combinations of cytokines for 6 and/or 12-14 hours, as indicated. The following cytokines were used; IL-1 beta at approximately l-5ng/ml, TNF alpha at approximately 5-lOng/ml, EFN gamma at approximately 20-50ng/ml, IL-4 at approximately 5-10ng/ml, IL-9 at approximately 5-10ng/ml, IL-13 at approximately 5- lOng/ml. Endothelial cells were sometimes starved for various times by culture in the basal media from Clonetics with 0.1 %> serum.
Mononuclear cells were prepared from blood of employees at CuraGen Coφoration, using Ficoll. LAK cells were prepared from these cells by culture in DMEM 5% FCS (Hyclone), lOOμM non essential amino acids (Gibco/Life Technologies, Rockville, MD), ImM sodium pyruvate (Gibco), mercaptoefhanol 5.5xl0"5M (Gibco), and lOmM Hepes (Gibco) and Interleukin 2 for 4-6 days. Cells were then either activated with 10-20ng/ml PMA and l-2μg/ml ionomycin, IL-12 at 5-10ng/ml, EFN gamma at 20-50ng/ml and IL-18 at 5- lOng/ml for 6 hours. In some cases, mononuclear cells were cultured for 4-5 days in DMEM 5% FCS (Hyclone), lOOμM non essential amino acids (Gibco), ImM sodium pyruvate (Gibco), mercaptoefhanol 5.5xl0"5M (Gibco), and lOmM Hepes (Gibco) with PHA (phytohemagglutinin) or PWM (pokeweed mitogen) at approximately 5 μg/ml. Samples were taken at 24, 48 and 72 hours for RNA preparation. MLR (mixed lymphocyte reaction) samples were obtained by taking blood from two donors, isolating the mononuclear cells using Ficoll and mixing the isolated mononuclear cells 1 : 1 at a final concentration of approximately
2xl06cells/ml in DMEM 5% FCS (Hyclone), lOOμM non essential amino acids (Gibco), ImM
sodium pyruvate (Gibco), mercaptoethanol (5.5xl0"5M) (Gibco), and lOmM Hepes (Gibco). The MLR was cultured and samples taken at various time points ranging from 1- 7 days for RNA preparation.
Monocytes were isolated from mononuclear cells using CD 14 Miltenyi Beads, +ve VS selection columns and a Vario Magnet according to the manufacturer's instructions. Monocytes were differentiated into dendritic cells by culture in DMEM 5%> fetal calf serum (FCS) (Hyclone, Logan, UT), lOOμM non essential amino acids (Gibco), ImM sodium pyruvate (Gibco), mercaptoethanol 5.5xl0"5M (Gibco), and lOmM Hepes (Gibco), 50ng/ml GMCSF and 5ng/ml IL-4 for 5-7 days. Macrophages were prepared by culture of monocytes for 5-7 days in DMEM 5%> FCS (Hyclone), lOOμM non essential amino acids (Gibco), ImM sodium pyruvate (Gibco), mercaptoethanol 5.5xl0"5M (Gibco), lOmM Hepes (Gibco) and 10% AB Human Serum or MCSF at approximately 50ng/ml. Monocytes, macrophages and dendritic cells were stimulated for 6 and 12-14 hours with lipopolysaccharide (LPS) at lOOng/ml. Dendritic cells were also stimulated with anti-CD40 monoclonal antibody (Pharmingen) at lOμg/ml for 6 and 12-14 hours.
CD4 lymphocytes, CD8 lymphocytes and NK cells were also isolated from mononuclear cells using CD4, CD8 and CD56 Miltenyi beads, positive VS selection columns and a Vario Magnet according to the manufacturer's instructions. CD45RA and CD45RO CD4 lymphocytes were isolated by depleting mononuclear cells of CD8, CD56, CD14 and CD19 cells using CD8, CD56, CD 14 and CD 19 Miltenyi beads and positive selection. CD45RO beads were then used to isolate the CD45RO CD4 lymphocytes with the remaining cells being
CD45RA CD4 lymphocytes. CD45RA CD4, CD45RO CD4 and CD8 lymphocytes were placed in DMEM 5%> FCS (Hyclone), lOOμM non essential amino acids (Gibco), ImM sodium pyruvate (Gibco), mercaptoethanol 5.5xlO"5M (Gibco), and lOmM Hepes (Gibco) and plated at 106cells/ml onto Falcon 6 well tissue culture plates that had been coated overnight with 0.5μg/ml anti-CD28 (Pharmingen) and 3ug/ml anti-CD3 (OKT3, ATCC) in PBS. After 6 and 24 hours, the cells were harvested for RNA preparation. To prepare chronically activated
CD8 lymphocytes, we activated the isolated CD8 lymphocytes for 4 days on anti-CD28 and anti-CD3 coated plates and then harvested the cells and expanded them in DMEM 5%> FCS
(Hyclone), lOOμM non essential amino acids (Gibco), ImM sodium pyruvate (Gibco), mercaptoethanol 5.5xl0"5M (Gibco), and lOmM Hepes (Gibco) and IL-2. The expanded CD8 cells were then activated again with plate bound anti-CD3 and anti-CD28 for 4 days and expanded as before. RNA was isolated 6 and 24 hours after the second activation and after 4
days of the second expansion culture. The isolated NK cells were cultured in DMEM 5% FCS (Hyclone), lOOμM non essential amino acids (Gibco), ImM sodium pyruvate (Gibco), mercaptoethanol 5.5xlO"5M (Gibco), and lOmM Hepes (Gibco) and IL-2 for 4-6 days before RNA was prepared.
To obtain B cells, tonsils were procured from NDRI. The tonsil was cut up with sterile dissecting scissors and then passed through a sieve. Tonsil cells were then spun down and resupended at 106cells/ml in DMEM 5%> FCS (Hyclone), lOOμM non essential amino acids (Gibco), ImM sodium pyruvate (Gibco), mercaptoethanol 5.5xl0"5M (Gibco), and lOmM Hepes (Gibco). To activate the cells, we used PWM at 5μg/ml or anti-CD40 (Pharmingen) at approximately lOμg/ml and IL-4 at 5-lOng/ml. Cells were harvested for RNA preparation at 24,48 and 72 hours.
To prepare the primary and secondary Thl/Th2 and Tri cells, six-well Falcon plates were coated overnight with lOμg/ml anti-CD28 (Pharmingen) and 2μg/ml OKT3 (ATCC), and then washed twice with PBS. Umbilical cord blood CD4 lymphocytes (Poietic Systems, German Town, MD) were cultured at 105-106cells/ml in DMEM 5% FCS (Hyclone), lOOμM non essential amino acids (Gibco), ImM sodium pyruvate (Gibco), mercaptoethanol 5.5x10" 5M (Gibco), lOmM Hepes (Gibco) and IL-2 (4ng/ml). EL- 12 (5ng/ml) and anti-IL4 (1 μg/ml) were used to direct to Thl, while IL-4 (5ng/ml) and anti-IFN gamma (1 μg/ml) were used to direct to Th2 and IL-10 at 5ng/ml was used to direct to Tri. After 4-5 days, the activated Thl, Th2 and Tri lymphocytes were washed once in DMEM and expanded for 4-7 days in DMEM 5% FCS (Hyclone), lOOμM non essential amino acids (Gibco), ImM sodium pyruvate (Gibco), mercaptoethanol 5.5xlO"5M (Gibco), lOmM Hepes (Gibco) and IL-2 (lng/ml). Following this, the activated Thl, Th2 and Tri lymphocytes were re-stimulated for 5 days with anti-CD28/OKT3 and cytokines as described above, but with the addition of anti-CD95L (1 μg/ml) to prevent apoptosis. After 4-5 days, the Thl, Th2 and Tri lymphocytes were washed and then expanded again with IL-2 for 4-7 days. Activated Thl and Th2 lymphocytes were maintained in this way for a maximum of three cycles. RNA was prepared from primary and secondary Thl, Th2 and Tri after 6 and 24 hours following the second and third activations with plate bound anti-CD3 and anti-CD28 mAbs and 4 days into the second and third expansion cultures in Interleukin 2.
The following leukocyte cells lines were obtained from the ATCC: Ramos, EOL-1, KU-812. EOL cells were further differentiated by culture in O.lmM dbcAMP at 5xl05cells/ml
for 8 days, changing the media every 3 days and adjusting the cell concentration to 5xl05 cells/ml. For the culture of these cells, we used DMEM or RPMI (as recommended by the ATCC), with the addition of 5%> FCS (Hyclone), lOOμM non essential amino acids (Gibco), ImM sodium pyruvate (Gibco), mercaptoethanol 5.5xl0"5M (Gibco), lOmM Hepes (Gibco). RNA was either prepared from resting cells or cells activated with PMA at lOng/ml and ionomycin at 1 μg/ml for 6 and 14 hours. Keratinocyte line CCD106 and an airway epithelial tumor line NCI-H292 were also obtained from the ATCC. Both were cultured in DMEM 5%> FCS (Hyclone), lOOμM non essential amino acids (Gibco), ImM sodium pyruvate (Gibco), mercaptoethanol 5.5xl0"5M (Gibco), and lOmM Hepes (Gibco). CCD1106 cells were activated for 6 and 14 hours with approximately 5 ng/ml TNF alpha and lng/ml IL-1 beta, while NCI-H292 cells were activated for 6 and 14 hours with the following cytokines: 5ng/ml IL-4, 5ng/ml IL-9, 5ng/ml IL-13 and 25ng/ml EFN gamma.
For these cell lines and blood cells, RNA was prepared by lysing approximately 107cells/ml using Trizol (Gibco BRL). Briefly, 1/10 volume of bromochloropropane (Molecular Research Coφoration) was added to the RNA sample, vortexed and after 10 minutes at room temperature, the tubes were spun at 14,000 φm in a Sorvall SS34 rotor. The aqueous phase was removed and placed in a 15ml Falcon Tube. An equal volume of isopropanol was added and left at -20°C overnight. The precipitated RNA was spun down at 9,000 φm for 15 min in a Sorvall SS34 rotor and washed in 70% ethanol. The pellet was redissolved in 300μl of RNAse-free water and 35μl buffer (Promega) 5μl DTT, 7μl RNAsin and 8μl DNAse were added. The tube was incubated at 37°C for 30 minutes to remove contaminating genomic DNA, extracted once with phenol chloroform and re-precipitated with 1/10 volume of 3M sodium acetate and 2 volumes of 100%> ethanol. The RNA was spun down and placed in RNAse free water. RNA was stored at -80°C.
AI_comprehensive panel_vl.0
The plates for AI_comprehensive panel_vl.O include two control wells and 89 test samples comprised of cDNA isolated from surgical and postmortem human tissues obtained from the Backus Hospital and Clinomics (Frederick, MD). Total RNA was extracted from tissue samples from the Backus Hospital in the Facility at CuraGen. Total RNA from other tissues was obtained from Clinomics.
Joint tissues including synovial fluid, synovium, bone and cartilage were obtained from patients undergoing total knee or hip replacement surgery at the Backus Hospital. Tissue samples were immediately snap frozen in liquid nitrogen to ensure that isolated RNA was of optimal quality and not degraded. Additional samples of osteoarthritis and rheumatoid arthritis joint tissues were obtained from Clinomics. Normal control tissues were supplied by Clinomics and were obtained during autopsy of trauma victims.
Surgical specimens of psoriatic tissues and adjacent matched tissues were provided as total RNA by Clinomics. Two male and two female patients were selected between the ages of 25 and 47. None of the patients were taking prescription drugs at the time samples were isolated.
Surgical specimens of diseased colon from patients with ulcerative colitis and Crohns disease and adjacent matched tissues were obtained from Clinomics. Bowel tissue from three female and three male Crohn's patients between the ages of 41-69 were used. Two patients were not on prescription medication while the others were taking dexamethasone, phenobarbital, or tylenol. Ulcerative colitis tissue was from three male and four female patients. Four of the patients were taking lebvid and two were on phenobarbital.
Total RNA from post mortem lung tissue from trauma victims with no disease or with emphysema, asthma or COPD was purchased from Clinomics. Emphysema patients ranged in age from 40-70 and all were smokers, this age range was chosen to focus on patients with cigarette-linked emphysema and to avoid those patients with alpha- lanti-tryp sin deficiencies. Asthma patients ranged in age from 36-75, and excluded smokers to prevent those patients that could also have COPD. COPD patients ranged in age from 35-80 and included both smokers and non-smokers. Most patients were taking corticosteroids, and bronchodilators.
In the labels employed to identify tissues in the AI_comprehensive panel_vl.O panel, the following abbreviations are used:
AI = Autoimmunity
Syn = Synovial
Normal = No apparent disease
Rep22 /Rep20 = individual patients
RA = Rheumatoid arthritis
Backus = From Backus Hospital
OA = Osteoarthritis
(SS) (BA) (MF) = Individual patients
Adj = Adjacent tissue
Match control = adjacent tissues
-M = Male
-F = Female
COPD = Chronic obstructive pulmonary disease
Panels 5D and 51
The plates for Panel 5D and 51 include two control wells and a variety of cDNAs isolated from human tissues and cell lines with an emphasis on metabolic diseases. Metabolic tissues were obtained from patients enrolled in the Gestational Diabetes study. Cells were obtained during different stages in the differentiation of adipocytes from human mesenchymal stem cells. Human pancreatic islets were also obtained.
In the Gestational Diabetes study subjects are young (18 - 40 years), otherwise healthy women with and without gestational diabetes undergoing routine (elective) Caesarean section. After delivery of the infant, when the surgical incisions were being repaired/closed, the obstetrician removed a small sample (<1 cc) of the exposed metabolic tissues during the closure of each surgical level. The biopsy material was rinsed in sterile saline, blotted and fast frozen within 5 minutes from the time of removal. The tissue was then flash frozen in liquid nitrogen and stored, individually, in sterile screw-top tubes and kept on dry ice for shipment to or to be picked up by CuraGen. The metabolic tissues of interest include uterine wall (smooth muscle), visceral adipose, skeletal muscle (rectus) and subcutaneous adipose. Patient descriptions are as follows:
Patient 2 Diabetic Hispanic, overweight, not on insulin
Patient 7-9 Nondiabetic Caucasian and obese (BMI>30)
Patient 10 Diabetic Hispanic, overweight, on insulin
Patient 11 Nondiabetic African American and overweight
Patient 12 Diabetic Hispanic on insulin
Adipocyte differentiation was induced in donor progenitor cells obtained from Osirus (a division of Clonetics/BioWhittaker) in triplicate, except for Donor 3U which had only two replicates. Scientists at Clonetics isolated, grew and differentiated human mesenchymal stem cells (HuMSCs) for CuraGen based on the published protocol found in Mark F. Pittenger, et al., Multilineage Potential of Adult Human Mesenchymal Stem Cells Science Apr 2 1999: 143-147. Clonetics provided Trizol lysates or frozen pellets suitable for mRNA isolation and ds cDNA production. A general description of each donor is as follows:
Donor 2 and 3 U: Mesenchymal Stem cells, Undifferentiated Adipose Donor 2 and 3 AM: Adipose, AdiposeMidway Differentiated Donor 2 and 3 AD: Adipose, Adipose Differentiated
Human cell lines were generally obtained from ATCC (American Type Culture Collection), NCI or the German tumor cell bank and fall into the following tissue groups: kidney proximal convoluted tubule, uterine smooth muscle cells, small intestine, liver HepG2 cancer cells, heart primary stromal cells, and adrenal cortical adenoma cells. These cells are all cultured under standard recommended conditions and RNA extracted using the standard procedures. All samples were processed at CuraGen to produce single stranded cDNA.
Panel 51 contains all samples previously described with the addition of pancreatic islets from a 58 year old female patient obtained from the Diabetes Research Institute at the University of Miami School of Medicine. Islet tissue was processed to total RNA at an outside source and delivered to CuraGen for addition to panel 51.
In the labels employed to identify tissues in the 5D and 51 panels, the following abbreviations are used:
GO Adipose = Greater Omentum Adipose
SK = Skeletal Muscle
UT = Uterus
PL = Placenta
AD = Adipose Differentiated
AM = Adipose Midway Differentiated
U = Undifferentiated Stem Cells
Panel CNSD.01
The plates for Panel CNSD.01 include two control wells and 94 test samples comprised of cDNA isolated from postmortem human brain tissue obtained from the Harvard Brain Tissue Resource Center. Brains are removed from calvaria of donors between 4 and 24 hours after death, sectioned by neuroanatomists, and frozen at -80°C in liquid nitrogen vapor. All brains are sectioned and examined by neuropathologists to confirm diagnoses with clear associated neuropathology.
Disease diagnoses are taken from patient records. The panel contains two brains from each of the following diagnoses: Alzheimer's disease, Parkinson's disease, Huntington's disease, Progressive Supernuclear Palsy, Depression, and "Normal controls". Within each of these brains, the following regions are represented: cingulate gyrus, temporal pole, globus
palladus, substantia nigra, Brodman Area 4 (primary motor strip), Brodman Area 7 (parietal cortex), Brodman Area 9 (prefrontal cortex), and Brodman area 17 (occipital cortex). Not all brain regions are represented in all cases; e.g., Huntington's disease is characterized in part by neurodegeneration in the globus palladus, thus this region is impossible to obtain from confirmed Huntington's cases. Likewise Parkinson's disease is characterized by degeneration of the substantia nigra making this region more difficult to obtain. Normal control brains were examined for neuropathology and found to be free of any pathology consistent with neurodegeneration.
In the labels employed to identify tissues in the CNS panel, the following abbreviations are used:
PSP = Progressive supranuclear palsy Sub Nigra = Substantia nigra Glob Palladus= Globus palladus Temp Pole = Temporal pole Cing Gyr = Cingulate gyrus BA 4 = Brodman Area 4
Panel CNS Neurodegeneration Vl.O
The plates for Panel CNS_Neurodegeneration_Vl .0 include two control wells and 47 test samples comprised of cDNA isolated from postmortem human brain tissue obtained from the Harvard Brain Tissue Resource Center (McLean Hospital) and the Human Brain and Spinal Fluid Resource Center (NA Greater Los Angeles Healthcare System). Brains are removed from calvaria of donors between 4 and 24 hours after death, sectioned by neuroanatomists, and frozen at -80°C in liquid nitrogen vapor. All brains are sectioned and examined by neuropathologists to confirm diagnoses with clear associated neuropathology.
Disease diagnoses are taken from patient records. The panel contains six brains from
Alzheimer's disease (AD) patients, and eight brains from "Normal controls" who showed no evidence of dementia prior to death. The eight normal control brains are divided into two categories: Controls with no dementia and no Alzheimer's like pathology (Controls) and controls with no dementia but evidence of severe Alzheimer's like pathology, (specifically senile plaque load rated as level 3 on a scale of 0-3; 0 = no evidence of plaques, 3 = severe AD senile plaque load). Within each of these brains, the following regions are represented: hippocampus, temporal cortex (Brodman Area 21), parietal cortex (Brodman area 7), and occipital cortex (Brodman area 17). These regions were chosen to encompass all levels of
neurodegeneration in AD. The hippocampus is a region of early and severe neuronal loss in AD; the temporal cortex is known to show neurodegeneration in AD after the hippocampus; the parietal cortex shows moderate neuronal death in the late stages of the disease; the occipital cortex is spared in AD and therefore acts as a "control" region within AD patients. Not all brain regions are represented in all cases.
In the labels employed to identify tissues in the CNS_Neurodegeneration_V1.0 panel, the following abbreviations are used:
AD = Alzheimer's disease brain; patient was demented and showed AD-like pathology upon autopsy
Control = Control brains; patient not demented, showing no neuropathology
Control (Path) = Control brains; pateint not demented but showing sever AD-like pathology
SupTemporal Ctx = Superior Temporal Cortex
Inf Temporal Ctx = Inferior Temporal Cortex
A. CG58522-01: HUMAN PLATELET-ACTIVATING FACTOR ACETYLHYDROLASE IB BETA
Expression of gene CG58522-01 was assessed using the primer-probe set Ag3365, described in Table AA. Results of the RTQ-PCR runs are shown in Table AB.
Table AA. Probe Name Ag3365
Table AB. General_screening_panel_vl.4
CNS neurodegeneration vl.O Summary: Ag3365 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3365 - Significant expression of this gene is seen only in the lung cancer cell line NCI-H23 (CT=33.1). Therefore, expression of this gene may be used to distinguish this sample from the other samples on this panel.
Panel 4D Summary: Ag3365 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
B. CG58520-01: GAMMA-AMINOBUTYRIC-ACID RECEPTOR GAMMA-1
Expression of gene CG58520-01 was assessed using the primer-probe set Ag3364, described in Table BA.
Table BA. Probe Name Ag3364
Start SEQ ID
Primers Sequences Length Position NO:
Forward ■ -ttcttctgcggagtcaaagtag-3 ' 22 43 360
TET-5 ' -ttggtcttcttgttactgaccctgca-3 '
Probe TAMRA 26 75 361
Reverse 5 ' -tcatctgccttatcaacgtttc-3 ' 22 106 362
CNS neurodegeneration vl.O Summary: Ag3364 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3364 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
Panel 4D Summary: Ag3364 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
Panel CNS_1 Summary: Ag3364 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
C. CG58520-03: GAMMA-AMINOBUTYRIC-ACID RECEPTOR GAMMA-1 SUBUNIT PRECURSOR (GABA(A) RECEPTOR)
Expression of gene CG58520-03 was assessed using the primer-probe set Ag5092, described in Table CA.
Table CA. Probe Name Ag5092
CNS neurodegeneration vl.O Summary: Ag5092 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.5 Summary: Ag5092 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
Panel 4.1D Summary: Ag5092 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
D. CG58518-01: GAMMA-AMINOBUTYRIC-ACID RECEPTOR RHO-3 -
Expression of gene CG58518-01 was assessed using the primer-probe sets Ag3363, Agl 130, Agl 198, Agl253 and Agl603, described in Tables DA, DB, DC, DD and DE. Results of the RTQ-PCR runs are shown in Tables DF, DG and DH.
Table DA. Probe Name Ag3363
Table DB. Probe Name Agl 130
Table DC. Probe Name Agl 198
Table DP. Probe Name Agl 253
Table DE. Probe Name Agl 603
Table DF. General_screening_panel_vl.4
Table DG. Panel 1.2
Table PH. Panel 4R
Tissue Name Rel. Exp.(%) | Tissue Name Rel. Exp.(%)
CNS neurodegeneration vl.O Summary: Ag3363 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3363 - Significant expression is seen in lung cancer cell line NCI-H146 (CT=34.5) and lung cancer cell line SHP-77 (CT=34.2). Therefore, expression of this can be used to distinguish these samples from the rest of the samples on this panel.
Panel 1.2 Summary: Agl 130/ Agl 198 - Three different runs using the same primer sequences yield similar results. Significant expression of this gene is seen in testis and a colon cancer sample. Therefore, expression of this gene can be used to differentiate these samples
from other samples on these panels. Results from a third experiment using the probe and primer set Agl253 show low/undetectable levels of expression in all the samples on this panel.
Panel 1.3D Summary: Agl253 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
Panel 2D Summary: Agl603 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
Panel 4D Summary: Agl 130/ Agl 198/Agl253/Ag3363 - Two experiments showed possible experimental difficulties, while the other three runs showed expression of this gene as low/undetectable (CTs > 35) across all of the samples on the panel.
Panel 4R Summary: Agl 198 - Significant expression of this gene is seen only in the EBD colitis 1 sample (CT=34.2). Therefore, expression of this gene can be used to differentiate this sample from others on the panel.
Panel CNS_1 Summary: Agl253/Agl603 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
E. CG58516-01: G-protein beta WD-40 repeats
Expression of gene CG58516-01 was assessed using the primer-probe set Ag3362, described in Table EA. Results of the RTQ-PCR runs are shown in Tables EB and EC.
Table EA. Probe Name Ag3362
Table EB. CNS_neurodegeneration_vl.O
Table EC. General_screening_panel_vl.4
CNS_neurodegeneration_vl.0 Summary: Ag3362 Highest expression of the CG58516-01 gene is seen in the occipital cortex of a control patient and the temporal cortex of an Alzheimer's patient. While the CG58516-01 gene does not appear to be preferentially expressed in Alzheimer's disease, this panel confirms expression of the CG58516-01 gene at moderate/high levels in the brain in an additional set of individuals. Please see Panel 1.4 for discussion of potential utility of this gene in the central nervous system.
General_screening_panel_vl.4 Summary: Ag3362 The CG58516-01 gene is widely expressed in this panel, with highest expression in the breast cancer cell line T47D (CT=29). Significant expression is also seen in cell lines derived from prostate, breast and ovarian cancers. In general, expression of the CG58516-01 gene appears to be greater in the cancer cell lines than in normal tissue. Thus, the expression of this gene could be used to distinguish these cell line types from others in the panel.
Among tissues involved in central nervous system function, this gene is expressed at low but significant levels in all brain regions examined. This gene encodes a protein with a
putative zinc-finger motif. Since these proteins are known to interact with nucleic acids, this suggests that this gene product may play a potential role in transcription. Thus, therapeutic modulation of the CG58516-01 gene product may be used to regulate the transcription of disease-related proteins such as ataxin, huntingtin, or various apoptosis cascade proteins.
Among tissues with metabolic function, this gene is expressed at low levels in pituitary, adipose, adrenal gland, pancreas, thyroid, skeletal muscle, heart, and fetal liver. This widespread expression among these tissues suggests that this gene product may play a role in normal neuroendocrine and metabolic and that disregulated expression of this gene may contribute to neuroendocrine disorders or metabolic diseases, such as obesity and diabetes.
References:
1. Zhu W, Chan EK, Li J, Hemmerich P, Tan EM. (2001) Transcription activating property of autoantigen SG2NA and modulating effect of WD-40 repeats. Exp Cell Res. 269(2):312-21
Panel 4D Summary: Ag3362 Results from one experiment with the CG58516-01 gene are not included because the amp plot corresponding to the run indicates that there were problems with the experiment.
F. CG58473-01: PROTEIN KINASE
Expression of gene CG58473-01 was assessed using the primer-probe set Ag3357, described in Table FA. Results of the RTQ-PCR runs are shown in Tables FB and FC.
Table FA. Probe Name Ag3357
Table FB. General_screening_panel__vl.4
Rel. Exp.(%) Ag3357, Rel. Exp.(%) Ag3357,
Tissue Name Tissue Name Run 216523477 Run 216523477
Table FC. Panel 4D
CNS_neurodegeneration_vl.0 Summary: Ag3357 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3357 This gene is primarily expressed in cancer cell lines, with highest expression in a breast cancer cell line BT 549(CT=32.8). This gene is expressed in the following cell lines but not the corresponding healthy tissue: gastric, brain, colon, lung, breast, ovarian cancer and melanomas. Thus, expression of this gene could be used as a diagnostic marker for the presence of these cancers. Furthermore, therapeutic inhibition using antibodies or small molecule drugs might be of use in the treatment of these cancers.
Panel 4D Summary: Ag3357 Highest expression of the CG58473-01 gene is seen in pokeweed mitogen-activated purified peripheral blood B lymphocytes (CT=33.2). In addition, no expression of the transcript is seen in PBMC that contain normal B cells, but the transcript is induced when PBMC are tieated with the B cell selective pokeweed mitogen. The tianscript is not seen in the B cell lymphoma cell line Ramos regardless of stimulation. Thus, the putative protein encoded by this gene could potentially be used diagnostically to identify activated B cells. Therefore, therapeutics that antagonize the function of this gene product may be useful as therapeutic drugs to reduce or eliminate the symptoms in patients with autoimmune and inflammatory diseases in which B cells play a part in the intiation or progression of the disease process, such as lupus eryfhematosus, Crohn's disease, ulcerative
colitis, multiple sclerosis, chronic obstructive pulmonary disease, asthma, emphysema, rheumatoid arthritis, or psoriasis.
G. CG58470-01: UDP-N-ACETYLHEXOSAMINE PYROPHOSPHORYLASE
Expression of gene CG58470-01 was assessed using the primer-probe set Ag5940, described in Table GA.
Table GA. Probe Name Ag5940
General_screening_panel_vl.5 Summary: Ag5940 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
Panel 5 Islet Summary: Ag5940 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
H. CG58593-01: UBIQUITIN-52
Expression of gene CG58593-01 was assessed using the primer-probe set Ag3421, described in Table HA.
Table HA. Probe Name Ag3421
Start SEQ ID
Primers Sequences Length Position NO:
Forward 5 ' -atctgctgcaagtgctatgc-3 ' 1 20 1 291 390
Probe TET-5 ' -cggtgctatcaactgccacaagaaga-3 ' -TAMRAI 26 j 323 391
Reverse 5 ' -tgaccttcttcctggggtac-3 ' 1 20 | 371 392
CNS neurodegeneration vl.O Summary: Ag3421 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3421 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
Panel 4D Summary: Ag3421 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
I. CG57871-01: TOUSLED-LIKE KINASE
Expression of gene CG57871-01 was assessed using the primer-probe set Ag3351, described in Table IA. Results of the RTQ-PCR runs are shown in Tables EB and IC.
Table IA. Probe Name Ag3351
Table EB. CNS_neurodegeneration_vl.O
Table IC. Panel 4D
CNS_neurodegeneration_vl.0 Summary: Ag3351 - This panel confirms the expression of this gene at low levels in the brain in an independent group of individuals. While no differential expression of this gene is detected between Alzheimer's diseased postmortem brains and those of non-demented controls, the widespread expression of this gene in the brain suggests that therapeutic modulation of the expression or function of this gene may be effective in the tieatment of neurologic disorders such as Parkinson's disease, epilepsy, stroke and multiple sclerosis.
General_screening_panel_vl.4 Summary: Ag3351 - Results from one experiment are not included. The amp plot indicates that there were experimental difficulties with this run.
Panel 4D Summary: Ag3351 The CG57871-01 gene is expressed at high to moderate levels in a wide range of cell types of significance in the immune response in health and disease. These cells include members of the T-cell, B-cell, endothelial cell, macrophage/monocyte, and peripheral blood mononuclear cell family, as well as epithelial and fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product may be involved in homeostatic processes for these and other cell types and tissues. This pattern also suggests a role for the gene product in cell survival and proliferation. Therefore, modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus eryfhematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
J. CG58590-01 and CG58590-02: PALS Guanylate kinase
Expression of gene CG58590-01 and CG58590-02 was assessed using the primer- probe set Ag3380, described in Table JA. Results of the RTQ-PCR runs are shown in Tables
JB, JC and JD. Please note that CG58590-02 represents a full-length physical clone of the CG58590-01 gene, validating the prediction of the gene sequence.
Table JA. Probe Name Ag3380
Table JB. CNS_neurodegeneration_vl.O
Table JD. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3380 This panel does not show differential expression of the CG58590-01 gene in Alzheimer's disease. However, this expression profile confirms the presence of this gene in the brain. Please see Panel 1.3D for discussion of utility of this gene in the central nervous system.
General_screening_panel_vl.4 Summary: Ag3380 - This gene is expressed at low to moderate levels in all samples on this pattern. The highest level of expression is seen in breast cancer cell line T47D (CT=27.8). Based on expression in this panel, this gene may be involved in brain, colon, renal, lung, ovarian and prostate cancer as well as melanomas. Thus, expression of this gene could be used as a diagnostic marker for the presence of these cancers.
Furthermore, therapeutic inhibition using antibodies or small molecule drugs might be of use in the tieatment of these cancers.
This gene product is also expressed in adipose, pancreas, adrenal, thyroid, pituitary, skeletal muscle, heart, and fetal liver. This widespread expression in tissues with metabolic function suggests that this gene product may be important for the pathogenesis, diagnosis, and/or treatment of metabolic and endocrine diseases, including obesity and Types 1 and 2 diabetes. Furthermore, this gene is more highly expressed in fetal (CT=30.9) liver when compared to expression in the adult (CT>35) and may be useful for the differentiation of the fetal and adult sources of this tissue.
In addition, this gene is expressed at moderate levels in the all regions of the CNS examined. Therefore, this gene may play a role in central nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Panel 4D Summary: Ag3380 - This gene is expressed from moderate to low levels across all of the samples on this panel. The highest expression is seen in small airway epithelium tieated with TNFalpha and IL-lbeta (CT=28.7). Interestingly, expression is much lower in untreated small airway epithelium (CT=31.5). There is also a significant difference between mononuclear cells treated with PWM (CT=29.5) and untreated cells (CT=32.7). Therefore, expression of this gene can be used to differentiate tieated and untreated samples.
Expression of this gene is detected at a moderate level (CT=30.2) in normal colon (similar levels for colon are seen on panel 1.4 (CT=30.9), but is significantly lower in the EBD Colitis 2 (CT=34.4) and EBD Crohn's (CT-33.5) samples. Therefore, therapies designed with the protein encoded for by this gene may potentially modulate colon function and play a role in the identification and treatment of inflammatory or autoimmune diseases which effect the colon including Crohn's disease and ulcerative colitis.
K. CG58572-01 and CG58572-02: GLUCOSAMINE-PHOSPHATE N- ACETYLTRANSFERASE
Expression of gene CG58572-01 and full length clone CG58572-02 was assessed using the primer-probe set Ag3375, described in Table KA. Results of the RTQ-PCR runs are shown in Tables KB, KC and KD.
Table KA. Probe Name Ag3375
Table KB. CNS_neurodegeneration_vl.0
Table KC. Panel 1.3D
Table KD. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3375 This panel does not show differential expression of the CG58572-01 gene in Alzheimer's disease. However, this expression profile confirms the presence of this gene in the brain. Please see Panel 1.3D for discussion of utility of this gene in the central nervous system.
Panel 1.3D Summary: Ag3375 - This gene is expressed at moderate to low levels in all samples on this panel, with the highest expression in gastric cancer cell line NCI-N87 (CT=28.8). Based on expression in this panel, this gene may be involved in gastric, pancreatic,
brain, colon, renal, lung, breast, ovarian and prostate cancer as well as melanomas. Thus, expression of this gene could be used as a diagnostic marker for the presence of these cancers. Furthermore, therapeutic modulation of the expression or function of this gene might be of use in the treatment of these cancers.
This gene product is also expressed in adipose, pancreas, adrenal, thyroid, pituitary, skeletal muscle, heart, and liver. This widespread expression in tissues with metabolic function suggests that this gene product may be important for the pathogenesis, diagnosis, and/or tieatment of metabolic and endocrine diseases, including obesity and Types 1 and 2 diabetes.
In addition, this gene is expressed at moderate levels in the CNS. Therefore, this gene may play a role in central nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Panel 4D Summary: Ag3375 The CG58572-01 gene is ubiquitously expressed on this panel, with highest expression in the B cell line Ramos treated with ionomycin (CT=26.2). Significant levels of expression are also seen in pokeweed mitogen-activated B lymphocytes. Therefore, therapies that antagonize the function of this gene product may be useful as therapeutic drugs to reduce or eliminate the symptoms in patients with autoimmune and inflammatory diseases in which B cells play a part in the initiation or progression of the disease process, such as lupus ery hematosus, Crohn's disease, ulcerative colitis, multiple sclerosis, chronic obstructive pulmonary disease, asthma, emphysema, rheumatoid arthritis, or psoriasis.
Interestingly, there is a difference between the levels of expression in resting and activated secondary T cells. The level in activated secondary T cells (CT=28.7-29.2) appears to be higher than in resting T cells (CT=31.3-33.1). Therefore, therapeutics designed with the protein encoded by this transcript could be important in the regulation of T cell function.
L. CG58564-01 and CG58564-02: PROTEIN TYROSINE PHOSPHATASE -
Expression of gene CG58564-01 and full length clone CG58564-02 was assessed using the primer-probe sets Ag3023 and Ag3373, described in Tables LA and LB. Results of the RTQ-PCR runs are shown in Tables LC, LD, LE and LF.
Table LA. Probe Name Ag3023
Table LB. Probe Name Ag3373
Table LC. CNS_neurodegeneration_vl.O
Table LD. General_screening_panel_vl.4
Table LF. Panel 4D
CNS_neurodegeneration_vl.0 Summary: Ag3023/Ag3373 This panel does not show differential expression of the CG58564-01 gene in Alzheimer's disease. However, this
expression profile confirms the presence of this gene in the brain. Please see Panel 1.3D for discussion of utility of this gene in the central nervous system.
General_screening_panel_vl.4 Summary: Ag3373 Highest expression of the CG58564-01 gene is seen in a prostate cancer cell line (CT=27). Overall, this gene is expressed at moderate levels in the cancer cell lines in this panel. A higher level of expression is observed in clusters of cell lines derived from prostate, brain, melanoma, colon, lung, breast and ovarian cancer when compared to expression in normal prostate, brain, colon, lung, breast and ovary. Thus, this gene could potentially be used as a diagnostic marker of cancer in these tissues. Furthermore, inhibition of the activity of this gene product using small molecule drugs may be effective in the tieatment of cancer in these tissues.
Among tissues with metabolic function, this gene product has moderate levels of expression in adipose, heart, skeletal muscle, adrenal, pituitary, thyroid and pancreas. Thus, this gene product may be a small molecule target for the treatment of endocrine and metabolic diseases, including obesity and Types 1 and 2 diabetes.
In addition, this gene appears to be differentially expressed in fetal (CT value = 29) vs adult liver (CT value =33) and may be useful for differentiation between the two sources of this tissue.
This gene is also expressed at moderate levels in all central nervous system samples present on this panel. Please see Panel 1.3D for discussion of utility of this gene in the central nervous system.
Panel 1.3D Summary: Ag3023 The CG58564-01 gene is ubiquitously expressed among the samples on this panel, with highest expression in an ovarian cancer cell line (CT=28.8). Overall, the expression of this gene shows good agreement with panel 1.4. A higher level of expression is observed in prostate, brain, melanoma, colon, lung, pancreatic, breast and ovarian cancer cell lines than the normal prostate, brain, colon, lung, pancreas, breast and ovary. Thus, expression of this gene could be used as a diagnostic marker of cancer in these tissues. Furthermore, inhibition of the activity of this gene product using small molecule drugs may be effective in the treatment of cancer in these tissues.
Among tissues with metabolic function, expression of this gene is widespread, as in the previous panel. Please see Panel 1.4 for discussion of utility of this gene in metabolic disease.
This gene represents a phosphatase that is also expressed at low to moderate levels across the CNS. Some phosphatases comprise a family of MAP kinase regulating enzymes, members of which are upregulated in brains subjected to insults such as ischemia and seizure activity. MAP kinases are kown to regulate neuro trophic and neurotoxic pathways. Consequently, agents that modulate the activity of this gene may have utility in attenuating the apoptotic and neurodegenerative processes following brain insults.
References:
1. Wiessner C. The dual specificity phosphatase PAC-1 is transcriptionally induced in the rat brain following transient forebrain ischemia. Brain Res Mol Brain Res 1995 Feb;28(2):353-6
2. Boschert U, Muda M, Camps M, Dickinson R, Arkinstall S. Induction of the dual specificity phosphatase PAC1 in rat brain following seizure activity. Neuroreport 1997 Sep 29;8(14):3077-80
Panel 4D Summary: Ag3023/Ag3373 The CG585864-01 gene is expressed at high to moderate levels in a wide range of cell types and tissues of significance in the immune response in health and disease. Highest expression of this gene is seen in ionomycin treated Ramos B cells (CT=26.83). Therefore, targeting of this gene product with a small molecule drug or antibody therapeutic may modulate the functions of cells of the immune system as well as resident tissue cells and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, and arthritis, including osteoarthritis and rheumatoid arthritis.
M. CG58564-03: Dual specificity phosphatase
Expression of gene CG58564-03 was assessed using the primer-probe sets Ag3023, Ag3373 and Ag5847, described in Tables MA, MB and MC. Results of the RTQ-PCR runs are shown in Tables MD, ME, MF, MG and MH.
Table MA. Probe Name Ag3023
Table MB. Probe Name Ag3373
Table MC. Probe Name Ag5847
Table MD. CNS_neurodegeneration_vl.O
Table ME. General_screening_panel_vl.4
Table MF. General_screening_panel_vl.5
Table MG. Panel 1.3D
Table MH. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3023/Ag3373 This panel does not show differential expression of the CG56804-03 gene, a splice variant of CG56804-01, in Alzheimer's disease. However, this expression profile confirms the presence of this gene in the brain. Please see Panel 1.3D for discussion of utility of this gene in the central nervous system. Ag5847 - This primer pair recognizes only the splice variant CG58564-03. Expression of this variant is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3373 Highest expression of the CG56804-03 gene is seen in a prostate cancer cell line (CT=27). Overall, this gene is expressed at moderate levels in the cancer cell lines in this panel. A higher level of expression is observed in clusters of cell lines derived from prostate, brain, melanoma, colon, lung, breast and ovarian cancer when compared to expression in normal prostate, brain, colon, lung, breast and ovary. Thus, this gene could potentially be used as a diagnostic marker of cancer in these tissues. Furthermore, inhibition of the activity of this gene product using small molecule drugs may be effective in the treatment of cancer in these tissues.
Among tissues with metabolic function, this gene product has moderate levels of expression in adipose, heart, skeletal muscle, adrenal, pituitary, thyroid and pancreas. Thus, this gene product may be a small molecule target for the treatment of endocrine and metabolic diseases, including obesity and Types 1 and 2 diabetes.
In addition, this gene appears to be differentially expressed in fetal (CT value = 29) vs adult liver (CT value =33) and may be useful for differentiation between the two sources of this tissue.
This gene is also expressed at moderate levels in all central nervous system samples present on this panel. Please see Panel 1.3D for discussion of utility of this gene in the central nervous system.
General_screening_panel_vl.5 Summary: Ag5847 - This primer pair, specific to this splice variant, CG58564-03. Expression of this variant is highest in salivary gland (CT=28.6). Therefore, expression of this gene can be used to differentiate this sample from others on the panel.
Panel 1.3D Summary: Ag3023 The CG56804-03 gene is ubiquitously expressed among the samples on this panel, with highest expression in an ovarian cancer cell line (CT=28.8). Overall, the expression of this gene shows good agreement with panel 1.4. A higher level of expression is observed in prostate, brain, melanoma, colon, lung, pancreatic, breast and ovarian cancer cell lines than the normal prostate, brain, colon, lung, pancreas, breast and ovary. Thus, expression of this gene could be used as a diagnostic marker of cancer in these tissues. Furthermore, inhibition of the activity of this gene product using small molecule drugs may be effective in the treatment of cancer in these tissues.
Among tissues with metabolic function, expression of this gene is widespread, as in the previous panel. Please see Panel 1.4 for discussion of utility of this gene in metabolic disease.
This gene represents a dual specificity phosphatase that is also expressed at low to moderate levels across the CNS. Dual-specificity phosphatases comprise a family of MAP kinase regulating enzymes, members of which are upregulated in brains subjected to insults such as ischemia and seizure activity. MAP kinases are kown to regulate neurotrophic and neurotoxic pathways. Consequently, agents that modulate the activity of this gene may have utility in attenuating the apoptotic and neurodegenerative processes following brain insults.
Panel 4.1D Summary: Ag5847 - This primer pair recognizes a splice variant of CG58564- 03. Expression of this variant is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
Panel 4D Summary: Ag3023/Ag3373 The CG56804-03 gene is expressed at high to moderate levels in a wide range of cell types and tissues of significance in the immune response in health and disease. Highest expression of this gene is seen in ionomycin tieated Ramos B cells (CT=26.83). Therefore, targeting of this gene product with a small molecule drug or antibody therapeutic may modulate the functions of cells of the immune system as well as resident tissue cells and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus eryfhematosus, and arthritis, including osteoarthritis and rheumatoid arthritis.
N. CG58564-04: Dual specificity phosphatase
Expression of gene CG58564-04, a splice variant of CG58564-01, was assessed using the primer-probe sets Ag3023, Ag3373 and Ag5844, described in Tables NA, NB and NC. Results of the RTQ-PCR runs are shown in Tables ND, NE, NF and NG.
Table NA. Probe Name Ag3023
Table NB. Probe Name Ag3373
Table NC. Probe Name Ag5844
Table ND. CNS_neurodegeneration_vl.0
Table NE. General_screening_panel_vl.4
Table NF. Panel 1.3D
Table NG. Panel 4D
Macrophages LPS 9.9 7.1 jThymus 14.4 12.9
HUVEC none 20.6 17.9 jKidney 27.5 19.6
HUVEC starved 43.5 38.4 1 I
CNS_neurodegeneration_vl.O Summary: Ag3023/Ag3373 This panel does not show differential expression of the CG56804-04 gene in Alzheimer's disease. However, this expression profile confirms the presence of this gene in the brain. Please see Panel 1.3D for discussion of utility of this gene in the central nervous system. Ag5847 - This primer pair recognizes a splice variant of CG58564-01 designated CG58564-04. Expression of this variant is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3373 Highest expression of the CG56804-04 gene is seen in a prostate cancer cell line (CT=27). Overall, this gene is expressed at moderate levels in the cancer cell lines in this panel. A higher level of expression is observed in clusters of cell lines derived from prostate, brain, melanoma, colon, lung, breast and ovarian cancer when compared to expression in normal prostate, brain, colon, lung, breast and ovary. Thus, this gene could potentially be used as a diagnostic marker of cancer in these tissues. Furthermore, inhibition of the activity of this gene product using small molecule drugs may be effective in the treatment of cancer in these tissues.
Among tissues with metabolic function, this gene product has moderate levels of expression in adipose, heart, skeletal muscle, adrenal, pituitary, thyroid and pancreas. Thus, this gene product may be a small molecule target for the treatment of endocrine and metabolic diseases, including obesity and Types 1 and 2 diabetes.
In addition, this gene appears to be differentially expressed in fetal (CT value = 29) vs adult liver (CT value =33) and may be useful for differentiation between the two sources of this tissue.
This gene is also expressed at moderate levels in all central nervous system samples present on this panel. Please see Panel 1.3D for discussion of utility of this gene in the central nervous system.
General_screening_panel_vl.5 Summary: Ag5844 - This primer pair recognizes a splice variant of CG58564-01. Expression of this variant is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
Panel 1.3D Summary: Ag3023 The CG56804-04 gene is ubiquitously expressed among the samples on this panel, with highest expression in an ovarian cancer cell line (CT=28.8). Overall, the expression of this gene shows good agreement with panel 1.4. A higher level of expression is observed in prostate, brain, melanoma, colon, lung, pancreatic, breast and ovarian cancer cell lines than the normal prostate, brain, colon, lung, pancreas, breast and ovary. Thus, expression of this gene could be used as a diagnostic marker of cancer in these tissues. Furthermore, inhibition of the activity of this gene product using small molecule drugs may be effective in the treatment of cancer in these tissues.
Among tissues with metabolic function, expression of this gene is widespread, as in the previous panel. Please see Panel 1.4 for discussion of utility of this gene in metabolic disease.
This gene represents a dual specificity phosphatase that is also expressed at low to moderate levels across the CNS. Dual-specificity phosphatases comprise a family of MAP kinase regulating enzymes, members of which are upregulated in brains subjected to insults such as ischemia and seizure activity. MAP kinases are known to regulate neurotiophic and neurotoxic pathways. Consequently, agents that modulate the activity of this gene may have utility in attenuating the apoptotic and neurodegenerative processes following brain insults.
Panel 4.1D Summary: Ag5844 - This primer pair recognizes a splice variant of CG58564- 01. Expression of this variant is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
Panel 4D Summary: Ag3023/Ag3373 The CG56804-04 gene is expressed at high to moderate levels in a wide range of cell types and tissues of significance in the immune response in health and disease. Highest expression of this gene is seen in ionomycin treated Ramos B cells (CT=26.83). Therefore, targeting of ghis gene product with a small molecule drug or antibody therapeutic may modulate the functions of cells of the immune system as well as resident tissue cells and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus eryfhematosus, and arthritis, including osteoarthritis and rheumatoid arthritis.
O. CG57819-01: RPGR-INTERACTING PROTEIN-1
Expression of gene CG57819-01 was assessed using the primer-probe set Ag3338, described in Table OA. Results of the RTQ-PCR runs are shown in Tables OB and OC.
Table OA. Probe Name Ag3338
Table OB. General_screening_panel_vl.4
CNS_neurodegeneration_vl.O Summary: Ag3338 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3338 - Expression of this gene is highest in testis (CT=29.4). Therefore, expression of this gene could be used to distinguish this sample from others on the panel.
There is also low expression in pancreatic cancer cell line CAPAN2, lung cancer cell line HOP-62, breast cancer cell line T47D, and ovarian cancer cell line OVCAR-5. Thus, expression of this gene could be used to differentiate these samples from other samples on this panel.
Panel 4D Summary: Ag3338 - Significant expression of this gene is seen only in resting monocytes (CT=32.3) Therefore, expression of this gene can be used to differentiate between this sample and others on this panel.
P. CG57789-01 and CG57789-02: RAS-LIKE PROTEIN RRP22-like
Expression of gene CG57789-01 and variant CG57789-02 was assessed using the primer-probe set Ag3333, described in Table PA. Results of the RTQ-PCR runs are shown in Tables PB, PC and PD.
Table PA. Probe Name Ag3333
Table PB. CNS_neurodegeneration_vl.0
Table PC. General_screening_panel_vl.4
Kidney Pool j 3.8 Adrenal Gland | 4.7
Fetal Kidney j 7.4 Pituitary gland Pool j 3.7
Renal ca. 786-0 j 0.2 Salivary Gland j 48.0
Renal ca. A498 j 20.9 Thyroid (female) j 1.1
Renal ca. ACHN j 8.5 Pancreatic ca. CAPAN2| 0.0
Renal ca. UO-31 j 3.0 Pancreas Pool | 4.0
Table PP. Panel 4D
CNS_neurodegeneration_vl.0 Summary: This panel confirms the expression of this gene in the brain in an independent group of individuals. However, no differential expression of this
gene was detected between Alzheimer's diseased postmortem brains and those of non- demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in treatment of central nervous system disorders.
General_screening_panel_vl.4 Summary: Ag3333 This gene is expressed at moderate to low levels in many of the samples on this panel, with the highest expression in colon cancer cell line SW480 (CT=27.8). Expression is significantly lower in SW680, a cell line derived from a metastasis of the primary tumor represented by SW480. Thus, expression of this gene could be used to differentiate between these two cell lines and potentially between primary colon cancer and its metastases.
Based on expression in this panel, this gene may be involved in gastric, brain, colon, renal, lung, breast, ovarian and prostate cancer as well as melanomas. Thus, expression of this gene could be used as a diagnostic marker for the presence of these cancers. Furthermore, therapeutic inhibition using antibodies or small molecule drugs might be of use in the treatment of these cancers.
This gene product is also expressed in adipose, pancreas, adrenal, thyroid, pituitary, skeletal muscle, heart, and liver. This widespread expression in tissues with metabolic function suggests that this gene product may be important for the pathogenesis, diagnosis, and/or treatment of metabolic and endocrine diseases, including obesity and Types 1 and 2 diabetes
This gene is expressed at low levels throughout the CNS, including in amygdala, substantia nigra, fhalamus, cerebellum, cerebral cortex, and spinal cord. Therefore, this gene may play a role in central nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Panel 4D Summary: Ag3333 The CG57789-01 gene is expressed at moderate to low levels in several samples on this panel, with the highest expression in resting astrocytes (CT=28.4). Moderate expression of this gene is seen in treated and untreated dermal and lung fibroblasts and the airway epithelial tumor line NCI-H292 cells. Thus, the transcript or the protein it encodes may be involved in pathological and inflammatory skin and lung conditions, including psoriasis, asthma, allergy, emphysema, and COPD.
Q. CG57758-01 and CG57758-02: SODIUM/LITHIUM-DEPENDENT DICARBOXYLATE TRANSPORTER
Expression of gene CG57758-01, a splice variant of CG57758-02, and CG57758-02 was assessed using the primer-probe sets Ag3326 and Ag3692, described in Tables QA and QB. Results of the RTQ-PCR runs are shown in Tables QC, QD, QE and QF.
Table QA. Probe Name Ag3326
Table QB. Probe Name Ag3692
Table QC. CNS_neurodegeneration_vl.O
Table QD. General_screening_panel_vl.4
Table OF. Panel 5 Islet
CNS_neurodegeneration_vl.O Summary: Ag3326/Ag3692 - Three experiments done with two primer pairs (same sequence) are in excellent agreement. This panel confirms the expression of this gene at low levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in tieatment of central nervous system disorders.
General_screening_panel_vl.4 Summary: Ag3326/Ag3692 Two experiments with the smae probe and primer set produce results that are in excellent agreement. This gene is highly expressed in fetal liver (CT=26.5-27.0) and moderately expressed in adult liver (CT=28.5- 28.8) and liver cancer cell line HepG2 (CT=28.4-28.8). This result agrees with the results seen in Panel 5 (expression in HepG2 (CT=29.2). These results are in agreement with published data that show a novel sodium dicarboxylate transporter in brain, choroid plexus kidney, intestine and liver. Thus, expression of this gene could be used to differentiate between these samples and other samples on this panel and as a marker for liver derived tissue.
This gene is expressed at low levels throughout the CNS, including in amygdala, substantia nigra, fhalamus, cerebellum, and cerebral cortex. Therefore, this gene may play a role in central nervous system disorders such as Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Low but significant levels of expression are also seen in the adrenal gland. Thus, this gene product may also be involved in metabolic disorders of this gland, including adrenoleukodystrophy and congenital adrenal hyperplasia.
References:
1. Pajor AM, Gangula R, Yao X. Cloning and functional characterization of a high- affinity Na(+)/dicarboxylate cotransporter from mouse brain. Am J Physiol Cell Physiol 2001 May;280(5):C1215-23.
2. Chen XZ, Shayakul C, Berger UV, Tian W, Hediger MA. Characterization of a rat Na+-dicarboxylate cotransporter. J Biol Chem 1998 Aug 14;273(33):20972-81.
Panel 4.1D Summary: Ag3692 Significant expression of this gene is seen only in kidney and a liver cirrhosis sample (CTs=34.0). These results confirm that this gene is expressed in liver derived samples. The presence in the kidney is also in agreement with published results. Please see Panel 1.4. This gene product may be involved in maintaining or restoring normal function to the kidney during inflammation.
Panel 4D Summary: Ag3326 Results from one experiment are not included. The amp plot indicates that there were experimental difficulties with this run.
Panel 5 Islet Summary: Ag3326 - The highest expression of this gene is in liver cancer cell line HepG2 (CT=29.2). There is also moderate expression in the small intestine (CT=30.5). These results compare well with previously published reports of sodium dicarboxylate transporter expression in mouse and rat (see discussion Panel 1.4).
R. CG57758-04 and CG57758-05: Sodium:sulfate symporter
Expression of gene CG57758-04 and CG57758-05, both splice variants of CG577584- 01, was assessed using the primer-probe sets Ag3326, Ag3692 and Ag5818, described in Tables RA, RB and RC. Results of the RTQ-PCR runs are shown in Tables RD, RE, RF, RG and RH.
Table RA. Probe Name Ag3326
Table RB. Probe Name Ag3692
Table RC. Probe Name Ag5818
Table RD. CNS_neurodegeneration_vl.O
Table RE. General_screening_panel_vl.4
Tissue Name Rel. Exp.(%) Rel. Exp.(%) Tissue Name | Rel. Exp.(%) Rel. Exρ.(%)
Table RF. General_screening_panel_vl.5
Table RG. Panel 4. ID
Table RH. Panel 5 Islet
CNS_neurodegeneration_vl.O Summary: Ag3326/Ag3692 - Three experiments done with two primer pairs (same sequence) are in excellent agreement. This panel confirms the expression of this gene at low levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in tieatment of central nervous system disorders. Ag5818 Results from one experiment are not included. The amp plot indicates that there were experimental difficulties with this run.
General_screening_panel_vl.4 Summary: Ag3326/Ag3692 Two experiments with the same probe and primer set produce results that are in excellent agreement. This gene is highly expressed in fetal liver (CT=26.5-27.0) and moderately expressed in adult liver (CT=28.5- 28.8) and liver cancer cell line HepG2 (CT=28.4-28.8). This result agrees with the results seen in Panel 5 (expression in HepG2 (CT=29.2). These results are in agreement with published data that show a novel sodium dicarboxylate transporter in brain, choroid plexus kidney, intestine and liver. Thus, expression of this gene could be used to differentiate between these samples and other samples on this panel and as a marker for liver derived tissue.
This gene is expressed at low levels throughout the CNS, including in amygdala, substantia nigra, thalamus, cerebellum, and cerebral cortex. Therefore, this gene may play a
role in central nervous system disorders such as Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Low but significant levels of expression are also seen in the adrenal gland. Thus, this gene product may also be involved in metabolic disorders of this gland, including adrenoleukodystrophy and congenital adrenal hyperplasia.
References:
1. Pajor AM, Gangula R, Yao X. Cloning and functional characterization of a high- affinity Na(+)/dicarboxylate cotiansporter from mouse brain. Am J Physiol Cell Physiol 2001 May;280(5):C1215-23.
2. Chen XZ, Shayakul C, Berger TJV, Tian W, Hediger MA. Characterization of a rat Na+-dicarboxylate cotiansporter. J Biol Chem 1998 Aug 14;273(33):20972-81.
General_screening_panel_vl.5 Summary: Ag5818 Results using this primer pair are in excellent agreement with the results seen in panel 1.4. See Panel 1.4 for discussion.
Panel 4.1D Summary: Ag3692 Significant expression of this gene is seen only in kidney and a liver cirrhosis sample (CTs=34.0). These results confirm that this gene is expressed in liver derived samples. The presence in the kidney is also in agreement with published results. Please see Panel 1.4. This gene product may be involved in maintaining or restoring normal function to the kidney during inflammation.
Panel 4D Summary: Ag3326 Results from one experiment are not included. The amp plot indicates that there were experimental difficulties with this run.
Panel 5 Islet Summary: Ag3326 The highest expression of this gene is in liver cancer cell line HepG2 (CT=29.2). There is also moderate expression in the small intestine (CT=30.5). These results compare well with previously published reports of sodium dicarboxylate transporter expression in mouse and rat (see discussion Panel 1.4).
S. CG57732-01 and CG57732-02 and CG57732-03: CA2+/CALMODULIN- DEPENDENT PROTEIN KINASE IV KINASE
Expression of gene CG57732-01 and full length clones CG57732-02 and CG57732-03, was assessed using the primer-probe set Ag3317, described in Table SA. Results of the RTQ-
PCR runs are shown in Tables SB, SC and SD. Please note CG57732-03 represents a splice variant of CG57732-01.
Table SA. Probe Name Ag3317
Table SB. CNS_neurodegeneration_vl.O
Table SC. General_screening_panel_vl.4
Table SD. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3317 - This panel does not show differential expression of this gene in Alzheimer's disease. However, this expression profile confirms the presence of this gene in the brain. Please see Panel 1.4 for discussion of utility of this gene in the central nervous system.
General_screening_panel_vl.4 Summary: Ag3317 - There is low to moderate expression this gene across all samples on this panel. This gene is expressed at moderate levels throughout the CNS, including in amygdala, substantia nigra, thalamus, cerebellum, and
cerebral cortex. Highest expression is observed in the cerebral cortex (CT=29.0). This gene encodes a calmodulin-dependent protein kinase IV homolog, which is known to play a role in Ca2+ signaling in the CNS that controls neuronal growth, differentiation, and plasticity. Mice deficient in calmodulin-dependent protein kinase IV were found to have cerebellar defects. Therefore, this gene may play a role in central nervous system disorders such as Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
This gene product is also expressed in adipose, pancreas, adrenal, thyroid, pituitary, skeletal muscle, heart, and liver. This widespread expression in tissues with metabolic function suggests that this gene product may be important for the pathogenesis, diagnosis, and/or treatment of metabolic and endocrine diseases, including obesity and Types 1 and 2 diabetes.
Based on expression in this panel, this gene may be also be involved in gastric, pancreatic, brain, colon, renal, lung, breast, ovarian and prostate cancer as well as melanomas. Thus, expression of this gene could be used as a diagnostic marker for the presence of these cancers. Furthermore, therapeutic inhibition using antibodies or small molecule drugs might be of use in the tieatment of these cancers.
References:
1. Okuno S, Kitani T, Fujisawa H. Evidence for the existence of Ca2+/calmodulin- dependent protein kinase IV kinase isoforms in rat brain. J Biochem (Tokyo) 1996 Jun;119(6):l 176-81.
2. Ribar TJ, Rodriguiz RM, Khiroug L, Wetsel WC, Augustine GJ, Means AR. Cerebellar defects in Ca2+/calmodulin kinase EV-deficient mice. J Neurosci 2000 Nov 15;20(22):RC107.
Panel 4D Summary: Ag3317 - This gene was found to have low expression across almost all the samples on this panel, with the highest level of expression seen in kidney and resting dermal fibroblasts (CTs=32). Expression of Ca2+/calmodulin-dependent kinase type IV in fhymocytes has been found in mice, where it plays a role in Ca2+-dependent gene transcription.
Reference
1. Raman V, Blaeser F, Ho N, Engle DL, Williams CB, Chatila TA. Requirement for Ca2+/calmodulin-dependent kinase type EV/Gr in setting the thymocyte selection threshold. J Immunol 2001 Dec l;167(l l):6270-8.
T. CG57709-01: Novel mitochondrial protein
Expression of gene CG57709-01 was assessed using the primer-probe set Ag3323, described in Table TA. Results of the RTQ-PCR runs are shown in Tables TB, TC and TD.
Table TA. Probe Name Ag3323
Table TB. CNS_neurodegeneration_vl.O
Table TC. Panel 1.3D
Table TD. Panel 4D
HUVEC starved | 24.8 j
CNS_neurodegeneration_vl.O Summary: Ag3323 This panel does not show differential expression of the CG57709-01 gene in Alzheimer's disease. However, this expression profile confirms the presence of this gene in the brain. Please see Panel 1.3D for discussion of utility of this gene in the central nervous system.
Panel 1.3D Summary: Ag3323 - This gene is expressed at moderate levels in all samples on this panel, with highest expression in a brain cancer cell line. Expression is also seen in all the cancer cell lines on this panel. Thus, expression of this gene could be used to differentiate between this brain cancer cell line sample and other samples on this panel and as a marker for brain cancer.
Among tissues with metabolic function, this gene is expressed at moderate to low levels in pituitary, adipose, adrenal gland, pancreas, thyroid, and adult and fetal skeletal muscle, heart, and liver. This widespread expression among these tissues suggests that this gene product may play a role in normal neuroendocrine and metabolic and that disregulated expression of this gene may contribute to neuroendocrine disorders or metabolic diseases, such as obesity and diabetes.
This molecule is also expressed at moderate to low levels in the CNS and may be a small molecule target for the treatment of neurologic diseases such as Alzheimer's disease, Parkinson's disease, epilepsy, schizophrenia, stroke and multiple sclerosis.
Panel 4D Summary: Ag3323 - This gene is expressed at high to moderate levels in all samples on this panel, with highest expression in B lymphocytes stimulated with polkweed mitogen (CT=24.5). In addition, this gene is expressed at higher levels in ionomycin-activated Ramos B lymphocytes. The highl levels of expression in activated B lymphocytes suggests that therapies that antagonize the function of this gene product may reduce or eliminate the symptoms in patients with autoimmune and inflammatory diseases in which B cells play a part in the initiation or progression of the disease process, such as lupus eryfhematosus, Crohn's disease, ulcerative colitis, multiple sclerosis, chronic obstructive pulmonary disease, asthma, emphysema, rheumatoid arthritis, or psoriasis.
U. CG57700-01: HYDROXYACYLGLUTATHIONE HYDROLASE (GLYOXALASE II)
Expression of gene CG57700-01 was assessed using the primer-probe set Ag3311, described in Table UA. Results of the RTQ-PCR runs are shown in Table UB.
Table UA. Probe Name Ag3311
Table UB. Panel 4D
AI_comprehensive panel_vl.0 Summary: Ag3311 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
CNS neurodegeneration vl.O Summary: Ag3311 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3311 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
Panel 4D Summary: Ag3311 - Significant expression of this gene is seen only in colon (CT=33.9). Therefore, expression of this gene can be used to distinguish between this sample and the others on the panel and between healthy and inflammed bowel. Since expression is not detectable in samples derived from Crohn's and colitis patients, therapeutic modulation of the expression or function of this gene may be useful in the treatment of inflammatory bowel disease.
V. CG58553-01 : vasolpressin receptor
Expression of gene CG58553-01 was assessed using the primer-probe set Ag3372, described in Table VA. Results of the RTQ-PCR runs are shown in Tables VB and VC.
Table VA. Probe Name Ag3372
Table VB. Panel 1.3D
Table VC. Panel 4D
Panel 1.3D Summary: Ag3372 Highest expression of the CG58553-01 gene is seen in the small intestine sample (CT=26.8). This gene encodes a novel vasopressin gene that plays a role in regulating electrolyte transport in the colon. Therefore, regulation of the transcript or the protein it encodes could be important in maintaining normal cellular homeostasis and in the treatment of Crohn's disease and ulcerative colitis.
Among tissues with metabolic function, this gene is expressed in liver and adipose. Thus, this gene product may be involved in disorders that affect these tissues, such as obesity and type II diabetes.
Low, but significant expression is also seen in the hippocampus. The hippocampus is critical for learning and memory. Thus, this gene product may have utility treating CNS disorders involving memory deficits, including Alzheimer's disease and aging.
References:
1. Sato Y, Hanai H, Nogaki A, Hirasawa K, Kaneko E, Hayashi H, Suzuki Y. Role of the vasopressin V(l) receptor in regulating the epithelial functions of the guinea pig distal colon. Am J Physiol 1999 Oct;277(4 Pt l):G819-28.
Panel 4D Summary: Ag3372 In agreement with the results seen in panel 1.4, the highest level of expression of this gene is in the colon sample (CT=27.5). Interestingly, the expression is significantly lower in the IBD colitis 2 (CT>35) and EBD Crohn's (CT=30.9)samples. Therefore, alterations in the expression of this gene may be used in the treatment of Crohn's disease and ulcerative colitis.
In addition, the expression of the CG58553-01 gene in severahpreparations of T lymphocytes suggests that small molecule antagonists, therapeutic antibodies specific for this molecule, or the extiacellular domain of this protein, may be useful to reduce or eliminate the symptoms of Crohn's disease, ulcerative colitis, multiple sclerosis, chronic obstructive pulmonary disease, asthma, emphysema, rheumatoid arthritis, lupus eryfhematosus, or psoriasis.
W. CG58626-01: Phospholipase
Expression of gene CG58626-01 was assessed using the primer-probe set Ag3386, described in Table WA. Results of the RTQ-PCR runs are shown in Tables WB, WC and WD.
Table WA. Probe Name Ag3386
Table WB. CNS_neurodegeneration_vl.0
Table WC. General_screening_panel_vl.4
Table WD. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3386 This panel confirms the expression of this gene at moderate to low levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in treatment of central nervous system disorders.
General_screening_panel_vl.4 Summary: Ag3386 This gene is moderately expressed in most of the samples on this panel. Based on expression in this panel, this gene may be involved in gastric, pancreatic, brain, colon, renal, lung, breast, ovarian and prostate cancer as well as melanomas. Thus, expression of this gene could be used as a diagnostic marker for the presence of these cancers. Furthermore, therapeutic inhibition using antibodies or small molecule drugs might be of use in the tieatment of these cancers.
This gene product is also expressed in adipose, pancreas, adrenal, thyroid, pituitary, skeletal muscle, heart, and liver. This widespread expression in tissues with metabolic function suggests that this gene product may be important for the pathogenesis, diagnosis, and/or tieatment of metabolic and endocrine diseases, including obesity and Types 1 and 2 diabetes.
In addition, this gene is expressed at moderate levels in the CNS. Therefore, this gene may play a role in central nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Panel 4D Summary: Ag3386 The CG58626-01 tianscript is expressed ubiquitously in this panel. Highest expression of this transcript is seen in activated Ramos cells and activated B cells (CTs=27). The expression of this tianscript in activated lymphoid cells when compared to non activated cells suggests that the CG58626-01 gene may be important for the diagnosis or pathogenesis of immune mediated diseases. Therefore, modulation of the expression and/or activity of this gene product might important for the tieatment of autoimmune diseases, allergy, and delayed type hypersensitivity.
X. CG57597-01: Hypothetical protein
Expression of gene CG57597-01 was assessed using the primer-probe set Ag3293, described in Table XA.
Table XA. Probe Name Ag3293
CNS neurodegeneration vl.O Summary: Ag3293 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3293 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
Panel 4D Summary: Ag3293 - Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
Y. CG57804-01: talin
Expression of gene CG57804-01 was assessed using the primer-probe set Ag3337, described in Table YA. Results of the RTQ-PCR runs are shown in Tables YB, YC and YD.
Table YA. Probe Name Ag3337
Table YB. CNS_neurodegeneration_vl.0
Tissue Name j Rel. Exp.(%) Ag3337, j Tissue Name | Rel. Exp.(%) Ag3337,
Table YC. General_screening_panel_vl.4
Table YD. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3337 - This panel confirms the expression of this gene at low to moderate levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in treatment of central nervous system disorders
General_screening_panel_vl.4 Summary: Ag3337 - This gene is expressed in almost all samples on this panel. This gene is expressed at moderate levels in the CNS. Therefore, this gene may play a role in central nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
In addition, this gene is also expressed in adipose, pancreas, adrenal, thyroid, pituitary, skeletal muscle, heart, and liver. This widespread expression in tissues with metabolic function suggests that this gene product may be important for the pathogenesis, diagnosis, and/or treatment of metabolic and endocrine diseases, including obesity and Types 1 and 2 diabetes.
Panel 4D Summary: Ag3337 This gene is most highly expressed in resting astrocytes (CT=28.9). In addition, this gene is highly expressed in a cluster of treated and untreated samples derived from lung and dermal fibroblasts. Thus, therapeutic modulation of the
expression or function of this gene maybe effective in the treatment of pathological and inflammatory lung and skin diseases, such as psoriasis, asthma, emphysema, and allergies.
Z. CG57551-01: NAC-1 Like Gene
Expression of gene CG57551-01 was assessed using the primer-probe set Ag3282, described in Table ZA. Results of the RTQ-PCR runs are shown in Tables ZB, ZC and ZD.
Table ZA. Probe Name Ag3282
Start SEQ ID
Primers Sequences Length Position NO:
Forward|5 ' -cagatcctcagcttctgctaca-3 ' 22 j 269 469 p , JTET-5 ' -accagttcctgctcatgtacacggct-3 ' - ,-,,- 3 roDe JTAMRA 0 J 318 470
Reverse |5 ' -atctcctggatctgcaggaa-3 ' 20 1 347 471
Table ZB. CNS_neurodegeneration_vl.O
Table ZC. General_screening_panel_vl.4
Renal ca. A498 1 16.7 Thyroid (female) ] 3.6
Renal ca. ACHN f 13.9 Pancreatic ca. CAPAN2| 15.5
Renal ca. UO-31 1 17.4 Pancreas Pool | 6.6
Table ZD. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3282 - This panel confirms the expression of this gene at low levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in tieatment of central nervous system disorders.
General_screening_panel_vl.4 Summary: Ag3282 Highest expression of this gene is seen in a brain cancer cell line (CT=24.3). This gene appears to be expressed more highly in the cancer cell lines than in the normal tissue samples on this panel and may be involved in cellular growth and proliferation. Based on this expression profile, this gene may be involved in gastric, pancreatic, brain, colon, renal, lung, breast, ovarian and prostate cancer as well as melanomas. Thus, expression of this gene could be used as a diagnostic marker for the presence of these cancers. Furthermore, therapeutic inhibition using antibodies or small molecule drugs might be of use in the treatment of these cancers.
This gene is also expressed at high levels in all regions of the CNS examined. Therefore, this gene may play a role in central nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
In addition, this gene product is expressed in adipose, pancreas, adrenal, thyroid, pituitary, fetal skeletal muscle, heart, and liver. This widespread expression in tissues with metabolic function suggests that this gene product may be important for the pathogenesis, diagnosis, and/or treatinent of metabolic and endocrine diseases, including obesity and Types 1 and 2 diabetes.
Furthermore, this gene is more highly expressed in fetal skeletal muscle (CT=30.4) and liver (CT=27) when compared to expression in the adult skeletal muscle (CT>35) and liver (CT=30) may be useful for the differentiation of the fetal and adult sources of this tissue.
Panel 4D Summary: Ag3282 This gene is expressed at high to moderate levels in a wide range of cell types of significance in the immune response in health and disease. Highest expression is seen in polkweed mitogen stimulated B lymphocytes (CT=25.7). In addition, expression is seen in members of the T-cell, B-cell, endothelial cell, macrophage/monocyte, and peripheral blood mononuclear cell family, as well as epithelial and fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product may be involved in homeostatic processes for these and other cell types and tissues. This pattern is in agreement with the expression profile in Panel 1.4 and also suggests a role for the gene product in cell survival and proliferation. Therefore, modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory
diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
AA. CG57411-01: KELCH-LIKE PROTEIN KLHL3C
Expression of gene CG57411-01 was assessed using the primer-probe set Ag3229, described in Table AAA. Results of the RTQ-PCR runs are shown in Tables AAB, AAC, AAD and AAE.
Table AAA. Probe Name Ag3229
Table AAB. CNS_neurodegeneration_vl.0
Table AAC. General_screening_panel_vl.4
Table AAD. Panel 2.2
CNS_neurodegeneration_vl.0 Summary: Ag3229 - This panel confirms the expression of this gene at low levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in treatment of cential nervous system disorders.
General_screening_panel_vl.4 Summary: Ag3229 - Highest levels of expression of this gene are seen in breast cancer cell line T47D (CT=28.5). Based on expression in this panel, this gene may be involved in gastric, brain, colon, renal, lung, breast, ovarian and prostate cancer as well as melanomas. Thus, expression of this gene could be used as a diagnostic marker for the presence of these cancers. Furthermore, therapeutic inhibition using antibodies or small molecule drugs might be of use in the tieatment of these cancers.
This gene product is also expressed in adipose, pancreas, adrenal, thyroid, pituitary, skeletal muscle, and heart. This widespread expression in tissues with metabolic function suggests that this gene product may be important for the pathogenesis, diagnosis, and/or treatment of metabolic and endocrine diseases, including obesity and Types 1 and 2 diabetes.
In addition, this gene is expressed at low to moderate levels in all regions of the CNS examined. Therefore, this gene may play a role in central nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Panel 2.2 Summary: Ag3229 Highest expression of the CG57411-01 gene is seen in the kidney (CT=32.2). In addition, significant levels of expression are seen in samples derived from normal lung and breast. Expression in these normal tissues is also higher than in the corresponding malignant tissue. Thus, expression of this gene could be used to differentiate between these samples and other samples on this panel and as a marker to detect the presence of lung, breast and kidney cancer. Furthermore, therapeutic modulation of the expression or function of this gene may be effective in the tieatment of lung, breast and kidney cancer.
Panel 4D Summary: Ag3229 Highest expression of the CG57411-01 gene is seen in IL-4 tieated lung fibroblasts (CT=31.3). Significant levels of expression are seen in activated-NCI-
H292 mucoepidermoid cells as well as untieated NCI-H292 cells. Moderate expression is also detected in IL-9, IL-13 and EFN gamma activated lung fibroblasts, human pulmonary aortic endothelial cells (tieated and untieated), small airway epithelium (treated and untreated), treated bronchial epithelium and lung microvascular endothelial cells (treated and untreated). The expression of this gene in cells derived from or within the lung suggests that this gene may be involved in normal conditions as well as pathological and inflammatory lung disorders that include chronic obstructive pulmonary disease, asthma, allergy and emphysema. Moderate/low expression of this gene is also detected in tieated and untreated HUVECs (endothelial cells) and coronary artery smooth muscle cells (treated and untieated) and normal tissues that include lung, colon, thymus and kidney. Expression in the various immune cell types and tissue samples suggests that therapeutic modulation of this gene product may ameliorate symptoms associated with infectious conditions as well as inflammatory and autoimmune disorders that include psoriasis, allergy, asthma, inflammatory bowel disease, rheumatoid arthritis and osteoarthritis.
AB. CG57399-01 and CG57399-03: PHOSPHOLIPASE ADRAB-B PRECURSOR
Expression of gene CG57399-01 and variant CG57399-03 was assessed using the primer-probe sets Ag3952 and Ag3226, described in Tables ABA and ABB. Results of the RTQ-PCR runs are shown in Tables ABC and ABD.
Table ABA. Probe Name Ag3952
Table ABB. Probe Name Ag3226
Table ABC. General_screening_panel_vl.4
Table ABD. Panel 1.3D
General_screening_panel_vl.4 Summary: Ag3952 Highest expression of this gene is seen in the adrenal gland (CT=29). Thus, this gene product may be a tieatment for Addison's disease and other adrenalopathies. This gene also has low levels of expression in adipose, heart, skeletal muscle, pituitary, thyroid, and pancreas. Therapeutic modulation of this gene product may be important for the diagnosis or treatment of endocrine or metabolic disease, including Types 1 and 2 diabetes, obesity and pancreatitis.
Expression of this gene is also seen in sample derived from colon, gastric, lung and breast cancers. Thus, expression of this gene could be used to differentiate between these samples and other samples on this panel and as a marker to detect the presence of these cancers. Furthermore, therapeutic modulation of the expression or function of this gene may be effective in the treatment of colon, gastric, lung and breast cancers.
Low but significant levels of expression are also seen for all regions of the CNS examined. Thus, this gene product may be useful for treatment of CNS disorders such as Alzheimer's disease, Parkinson's disease, stroke, epilepsy, schizophrenia and multiple sclerosis.
Panel 1.3D Summary: Ag3952 Highest expression of the CG57399-01 gene is seen in a lung cancer cell line (CT=32.5). Low but significant expression is also seen in cell lines derived from breast and colon cancers. Overall, expression is consistent with expression seen in Panel
1.4. Thus, expression of this gene could be used to differentiate between these samples and other samples on this panel and as a marker to detect the presence of these cancers. Furthermore, therapeutic modulation of the expression or function of this gene may be effective in the tieatment of colon, gastric, lung and breast cancers.
Among metabolic tissues, significant levels of expression are seen in adipose and the adrenal gland. Thus, this gene product may be useful for treatment of obesity, Addison's disease and other adrenalopathies.
In addition, this gene is expressed in the hippocampus, and cerebral cortex. Both these regions of the brain undergo degeneration in Alzheimer's disease. Thus, therapeutic modulation of the expression or function of this gene may be effective in the treatment of this disease or any other neurodegenerative disorders.
AC. CG57399-02: PHOSPHOLIPASE ADRAB-B PRECURSOR
Expression of gene CG57399-02 was assessed using the primer-probe set Ag3952, described in Table ACA. Results of the RTQ-PCR runs are shown in Table ACB. Please note that this gene represents a variant of CG57399-01. This sequence however, only corresponds to probe and primer set Ag3952.
Table ACA. Probe Name Ag3952
Table ACB. General_screening_panel_vl.4
General_screening_panel_vl.4 Summary: Ag3952 Highest expression of this gene is seen in the adrenal gland (CT=29). Thus, this gene product may be a tieatment for Addison's disease and other adrenalopathies. This gene also has low levels of expression in adipose, heart, skeletal muscle, pituitary, thyroid, and pancreas. Therapeutic modulation of this gene product may be important for the diagnosis or treatment of endocrine or metabolic disease, including Types 1 and 2 diabetes, obesity and pancreatitis.
Expression of this gene is also seen in cell line samples derived from colon, gastric, lung and breast cancers. Thus, expression of this gene could be used to differentiate between these samples and other samples on this panel and as a marker to detect the presence of these cancers. Furthermore, therapeutic modulation of the expression or function of this gene may be effective in the tieatment of colon, gastric, lung and breast cancers.
Low but significant levels of expression are also seen for all regions of the CNS examined. Thus, this gene product may be useful for tieatment of CNS disorders such as Alzheimer's disease, Parkinson's disease, stroke, epilepsy, schizophrenia and multiple sclerosis.
AD. CG59311-01: ACYL-COENZYME A THIOESTER HYDROLASE bp.
Expression of gene CG59311-01, splice variant CG59311-02, and full length clone CG59311-03, was assessed using the primer-probe set Ag3541, described in Table ADA. Results of the RTQ-PCR runs are shown in Tables ADB and ADC.
Table ADA. Probe Name Ag3541
Table ADB. General_screening_panel_vl.4
Table ADC. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3541 - Expression of this gene is low/undetectable (CTs > 34.5) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3541 Significant expression of this gene is seen only in cerebellum, fetal brain, the breast cancer cell line T47D, and ovarian cancer cell line OVCAR-5 (CTs=32-35). Therefore, expression of this gene can be used to differentiate between these samples and others on this panel.
Panel 4D Summary: Ag3541 - There is significant expression of this gene only in thymus (CT=33.8). Therefore, expression of this gene may be used to identify thymic tissue.
Furthermore, drugs that inhibit the function of this protein may regulate T cell development in the thymus and reduce or eliminate the symptoms of T cell mediated autoimmune or inflammatory diseases, including asthma, allergies, inflammatory bowel disease, lupus erythematosus, or rheumatoid arthritis. Additionally, therapeutics designed against this putative protein may disrupt T cell development in the thymus and function as an immunosuppresant for tissue transplant.
AE. CG59309-01: ACYL-COENZYME A THIOESTER HYDROLASE
Expression of gene CG59309-01 was assessed using the primer-probe set Ag3540, described in Table AEA. Results of the RTQ-PCR runs are shown in Tables AEB, AEC, AED and AEE.
Table AEA. Probe Name Ag3540
Table AEB. CNS_neurodegeneration_vl.0
Table AEC. General_screening_panel_vl.4
Table AED. Panel 4D
HUVEC starved 1.4
Table AEE. Panel 5 Islet
CNS_neurodegeneration_vl.O Summary: Ag3540 - This panel confirms the expression of this gene at low levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment.
General_screening_panel_vl.4 Summary: Ag3540 This gene is most highly expressed in a breast cancer cell line (CT=27.1). Thus, expression of this gene could be used to differentiate between this sample and other samples on this panel and as a marker to detect the presence of breast cancer. Furthermore, therapeutic modulation of the expression or function of this gene may be effective in the treatment of breast cancer.
Among metabolic tissues, this gene, an acyl coA thioesterase homolog, has a low level of expression in adipose, adult and fetal liver, adrenal, thyroid and pancreas. Acyl CoA thioesterases have multiple roles in lipid homeostasis. Therefore, therapeutic modulation of this gene product may be a tieatment for endocrine and metabolic disease, including Types 1 and 2 diabetes and obesity.
In addition, this gene is expressed in all CNS regions examined. Thus, therapeutic modulation of the expression or function of this gene may be effective in the treatment of neurologic disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, stroke, schizophrenia and multiple sclerosis.
References:
1. Hunt MC, Alexson SE. The role Acyl-CoA thioesterases play in mediating intracellular lipid metabolism. Prog Lipid Res. 2002 Mar;41(2):99-130.
2. Hunt MC, Nousiainen SE, Huttunen MK, Orii KE, Svensson LT, Alexson SE. Peroxisome proliferator-induced long chain acyl-CoA thioesterases comprise a highly conserved novel multi-gene family involved in lipid metabolism. J Biol Chem. 1999 Nov 26;274(48):34317-26.
Panel 4D Summary: Ag3540 Highest expression of the CG59309-01 gene is seen in the thymus and colon (CTs=31.5). Significant levels of expression are also seen in a cluter of treated and untreated samples derived from the NCI-H292 mucoepidermoid cell line. Thus, expression of this gene could be used as a marker for thymus and colon. Furthermore, therapeutic modulation of the expression or function of this gene may regulate T cell development in the thymus and reduce or eliminate the symptoms of T cell mediated autoimmune or inflammatory diseases, including asthma, allergies, inflammatory bowel disease, lupus erythematosus, or rheumatoid arthritis. Additionally, small molecule or antibody therapeutics designed against this putative protein may disrupt T cell development in the thymus and function as an immunosuppresant for tissue transplant.
Panel 5 Islet Summary: Ag3540 This gene has moderate expression in skeletal muscle, (highest expression CT=30.5). Acyl CoA thioesterases function in peroxisomal fatty acid oxidation. Therefore, therapeutic modulation of this homolog may increase fatty acid oxidation in muscle and be a tieatment for Type 2 diabetes and obesity.
References:
1. Hunt MC, Solaas K, Kase BF, Alexson SE. Characterization of an acyl-coA thioesterase that functions as a major regulator of peroxisomal lipid metabolism. J Biol Chem. 2002 Jan 11;277(2): 1128-38.
AF. CG57364-01: CG6896
Expression of gene CG57364-01 was assessed using the primer-probe sets Ag3218 and Ag3378, described in Tables AFA and AFB. Results of the RTQ-PCR runs are shown in Tables AFC, AFD, AFE and AFF.
Table AFA. Probe Name Ag3218
Table AFB. Probe Name Ag3378
Table AFC. CNS_neurodegeneration_vl.0
Table AFD. Panel 1.3D
Table AFE. Panel 2.2
Table AFF. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3218/Ag3378 - Two different experiments using probe/primer sets with the same sequence are in very good agreement. This panel confirms the expression of this gene at low levels to moderate levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls
in this experiment. Please see Panel 1.3D for a discussion of the potential utility of this gene in treatment of central nervous system disorders.
Panel 1.3D Summary: Ag3218/Ag3378 - Two different experiments using probe/primer sets with the same sequence are in good agreement. Highest expression is seen in testis and a lung cancer cell line (CTs=30-31). This panel confirms the expression of this gene at low levels in the brain. Therefore, this gene may play a role in central nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
This gene product is also expressed in adipose, pancreas, thyroid, pituitary, heart, and liver. This widespread expression in tissues with metabolic function suggests that this gene product may be important for the pathogenesis, diagnosis, and/or treatment of metabolic and endocrine diseases, including obesity and Types 1 and 2 diabetes.
Based on expression in this panel, this gene may be involved in gastric, pancreatic, brain, colon, renal, lung, breast, ovarian and prostate cancer as well as melanomas. Thus, expression of this gene could be used as a diagnostic marker for the presence of these cancers. Furthermore, therapeutic inhibition using antibodies or small molecule drugs might be of use in the tieatment of these cancers.
Panel 2.2 Summary: Ag3218 - This gene is expressed at low to moderate levels in many samples on this panel, with the highest levels of expression in breast cancer sample OD04590- 01 (CT=30.3). This gene is expressed in a cluster of breast cancer samples with no expression in normal breast (CT>35). Similarly, this gene is expressed in ovarian cancer samples at higher levels than the matched margin samples.
Interestingly, this gene is expressed at higher levels in kidney cancer margin samples than in the matched cancer samples.
This gene is homologous to a mouse myosin phosphatase targeting subunit (MYPT) which have been found to play a role in cell division. MYPT undergoes mitosis-specific phosphorylation which is reversed during cytokinesis.
References:
1. Totsukawa G, Yamakita Y, Yamashiro S, Hosoya H, Hartshorne DJ, Matsumura F. Activation of myosin phosphatase targeting subunit by mitosis-specific phosphorylation. J Cell Biol 1999 Feb 22;144(4):735-44.
Panel 4D Summary: Ag3218/Ag3378 - Two different experiments using probe/primer sets with the same sequence are in very good agreement. Highest expression is seen in the colon and a mucoepidermoid cell line (CTs=30-32). This gene is expressed at low to moderate levels in a wide range of cell types of significance in the immune response in health and disease. These cells include members of the T-cell, B-cell, endothelial cell, macrophage/monocyte, and peripheral blood mononuclear cell family, as well as epithelial and fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product may be involved in homeostatic processes for these and other cell types and tissues. This pattern suggests a role for the gene product in cell survival and proliferation. Therefore, modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
AG. CG59241-01 : Amiloride-sensitive sodium channel
Expression of gene CG59241-01 was assessed using the primer-probe set Ag3407, described in Table AGA. Results of the RTQ-PCR runs are shown in Tables AGB, AGC and AGD.
Table AGA. Probe Name Ag3407
Table AGB. CNS_neurodegeneration_vl.0
Table AGD. Panel 4D
CNS_neurodegeneration_vl.0 Summary: Ag3407 This panel confirms the expression of this gene at low levels in the brains of an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in treatment of central nervous system disorders.
General_screening_panel_vl.4 Summary: Ag3407 Highest expression of the CG59241-01 gene is seen in fetal brain (CT=31.3). Furthermore, low to moderate levels of expression is also observed in CNS cancer cell lines (CTs=32-34). The CG59241-01 gene codes for a putative amiloride-sensitive sodium channel. A similar amiloride-sensitive sodium channel was shown to be highly expressed in malignant glioblastoma multiforme tumors and to be a charachteristic feature of malignant brain tumor cells (Ref.l). Therefore, therapeutic modulation of the activity of the protein encoded by this gene may be beneficial in the treatment of CNS cancer. Significant expression is also seen in a cluster of cell lines derived from brain, colon, breast, and ovarian cancers. Therefore, therapeutic modulation of the activity of this gene or its protein product, through the use of small molecule drugs, protein therapeutics or antibodies, might be beneficial in the tieatment of these cancers.
In addition, this gene is expressed at low levels in all regions of the central nervous system examined, including amygdala, hippocampus, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. Therefore, this gene may play a role in central nervous system
disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
References:
1. Bubien JK, Keeton DA, Fuller CM, Gillespie GY, Reddy AT, Mapstone TB, Benos DJ. (1999) Malignant human gliomas express an amiloride-sensitive Na+ conductance. Am J Physiol 276(6 Pt 1):C1405-10
Panel 4D Summary: Ag3407 Highest expression Of the CG59241-01 gene is detected in PWM treated B lymphocytes (CT=32). Similar expression is also detected in primary activated Thl, Th2 and Tri cells, as well as TNF alpha tieated dermal fibroblast CCD 1070 cells (CTs=32). Therefore, expression of this gene can be used to distinguish these samples from other samples in the panel. Furthermore, this gene is expressed in activated lymphocytes. Likewise, no expression of this gene is seen in PBMC that contain normal B cells (CT=40), but it is induced when PBMC are treated with the pokeweed mitogen or PHA-L (CTs=34). En addition, the transcript is not seen in the B cell lymphoma Ramos regardless of stimulation. Therefore, the gene product could potentially be used therapeutically in the treatment of Crohn's disease, ulcerative colitis, multiple sclerosis, chronic obstructive pulmonary disease, asthma, emphysema, rheumatoid arthritis, lupus erythematosus, psoriasis and in other diseases in which T cells and B cells are activated.
In addition, low expression of this gene is also observed in normal colon, lung, thymus and kidney tissues. The CG59241-01 gene encodes an amiloride-sensitive sodium channel. A similar channel, the amiloride-sensitive epithelial sodium channel (ENaC) constitutes the limiting step for sodium reabsorption in epithelial cells that line the distal nephron, distal colon, ducts of several exocrine glands and lung airways and plays an important role in pathophysiological and clinical conditions such as hypertension or lung edema. ENaC has been implicated in two genetic diseases, Liddle's syndrome and pseudohypoaldosteronism (PHA-1) (Ref.l). Therefore, antibody or small molecule therapies designed with the protein encoded for by CG59241-01 gene could modulate kidney/colon/lung function and be important in the treatment of inflammatory or autoimmune diseases of these tissues in addition to hypertension, lung edema, Liddle's syndrom and PHA-1.
Reference.
1. Hummler E. (1998) Reversal of convention: from man to experimental animal in elucidating the function of the renal amiloride-sensitive sodium channel. Exp Nephrol 1998 Jul-Aug;6(4):265-71
AH. CG58602-01 : FAD binding domain containing protein
Expression of gene CG58602-01 was assessed using the primer-probe set Ag3385, described in Table AHA. Results of the RTQ-PCR runs are shown in Tables AHB, AHC and AHD.
Table AHA. Probe Name Ag3385
Table AHB. CNS_neurodegeneration_ vl.0
Table AHC. General_screening_panel_vl.4
CNS_neurodegeneration_vl.O Summary: Ag3385 This panel confirms the expression of
CG58602-01 gene at low levels in the brains of an independent group of individuals.
However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in treatment of central nervous system disorders.
General_screening_panel_vl.4 Summary: Ag3385 Highest expression of the CG58602-01 gene is seen in a breast cancer cell line (CT=26.3). Significant expression is also seen in an ovarian cancer cell line. Thus, expression of this gene could be used to differentiate between these samples and other samples on this panel and as a marker to detect the presence of breast and ovarian cancers. Furthermore, therapeutic modulation of the expression or function of this gene may be effective in the treatment of breast and ovarian cancers.
Among tissues with metabolic function, this gene is expressed at moderate to low levels in pituitary, adipose, adrenal gland, pancreas, thyroid, and adult and fetal skeletal muscle, heart, and liver. This widespread expression among these tissues suggests that this gene product may play a role in normal neuroendocrine and metabolic and that disregulated expression of this gene may contribute to neuroendocrine disorders or metabolic diseases, such as obesity and diabetes.
Expression of this gene is higher in fetal skeletal muscle (CT=28.3) when compared to expression in adult skeletal muscle (CT=31.5). Thus, expression of this gene could be used to distinguish fetal from adult skeletal muscle.
In addition, this gene is expressed at high levels (CTs=29-30.4) in all regions of the central nervous system examined, including amygdala, hippocampus, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. Therefore, this gene may play a role in cential nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Panel 4D Summary: Ag3385 Highest expression of the CG58602-01 gene is seen in the thymus (CT=28). Thus, the putative protein encoded for by this gene could therefore play an important role in T cell development. Therefore, small molecule therapeutics designed against the proetin encoded by this gene could be utilized to modulate immune function (T cell development) and be important for organ transplant, AEDS treatment or post chemotherapy immune reconstitiution.
AI. CG58468-01: Serum Amyloid P Component
Expression of gene CG58468-01 was assessed using the primer-probe set Ag3356, described in Table ALA. Results of the RTQ-PCR runs are shown in Table AEB.
Table ALA. Probe Name Ag3356
Table AEB. General_screening_panel_vl.4
CNS_neurodegeneration_vl.O Summary: Ag3356 Expression of the CG58468-01 gene is low/undetectable in all the samples on this panel. (CTs>35). (Data not shown.)
General_screening_panel_vl.4 Summary: Ag3356 Expression of the CG58468-01 gene is restricted to the colon (CT=34). Thus, expression of this gene could be used to differentiate between this sample and other samples on this panel.
Panel 4D Summary: Ag3356 Results from one experiment with the CG56003-01 gene are not included. The amp plot indicates that there were experimental difficulties with this run.
AJ. CG58183-01: N-METHYL-D-ASPARTATE RECEPTOR
Expression of gene CG58183-01 was assessed using the primer-probe set Ag3355, described in Table AJA. Results of the RTQ-PCR runs are shown in Tables AJB, AJC and AJD.
Table AJA. Probe Name Ag3355
Table AJB. CNS_neurodegeneration_vl.O
Table AJC. General_screening_panel_vl.4
Table AJD. Panel 4D
CNS_neurodegeneration_vl.0 Summary: Ag3355 This panel confirms the expression of CG58183-01 gene at low levels in the brains of an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in treatment of central nervous system disorders.
General_screening_panel_vl.4 Summary: Ag3355 Highest expression of CG58183-01 gene is detected in fetal brain (Ct=29.2). In addition, this gene is expressed at high levels in all regions of the cential nervous system examined, including amygdala, hippocampus, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord (CTs= 29-32). Therefore, this gene may play a role in central nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
This gene codes for N-mefhyl-D-aspartate (NMD A) receptor 3A protein. In cats and rhodent models competitive NMD A receptor antagonists, such as D-(E)-4-(3-phosphonoprop- 2-enyl)piperazine-2-carboxylic acid, which act at the neurotiansmitter recognition site were shown to be effective in reducing ischaemic brain damage when administered prior to the onset of an ischaemic episode (Ref. 1). Therefore, therapeutic modulation of the activity of the protein encoded by this gene may be beneficial in the treatment of ischaemic brain.
Among tissues with metabolic or endocrine function, this gene is expressed at low levels in pancreas, heart, and the gastrointestinal tract. Therefore, therapeutic modulation of the activity of this gene may prove useful in the tieatment of endocrine/metabolically related diseases, such as obesity and diabetes.
Furthermore, low to moderate expression of this gene is detected in lung cancer, and CNS cancer cell lines. Therefore, therapeutic modulation of the activity of this gene or its protein product, through the use of small molecule drugs, protein therapeutics or antibodies, might be beneficial in the treatment of lung cancer or CNS cancer.
References:
1. McCulloch J. (1991) Ischaemic brain damage— prevention with competitive and non-competitive antagonists of N-methyl-D-aspartate receptors. Arzneimittelforschung 41(3A):319-24.
Panel 4D Summary: Ag3355 Expression of the CG58183-01 gene is limited to a few samples, with highest expression in the thymus (CT=33.5). Thus, expression of this gene may be useful as a marker of thymic tissue. Low, but significant levels of expression are also seen in the kidney, in TNF- alpha and IL-1 beta tieated astrocytes and in the PMA/ionomycin tieated basophil cell line KU-812. Thus, this gene product may be involved in the normal homeostasis of this tissue. Therefore, agonistic antibodies or protein therapeutics may be important in the treatment of inflammatory or autoimmune diseases that affect the kidney, including lupus and glomerulonephritis. In addition, the expression of this transcript in astrocytes treated with TNF-a and IL-1 indicates that therapeutics designed against the protein encoded by this gene may be useful for the tieatment of inflammatory CNS diseases such as multiple sclerosis.
AK. CG59315-01: connexin
Expression of gene CG59315-01 was assessed using the primer-probe set Ag3542, described in Table AKA. Results of the RTQ-PCR runs are shown in Tables AKB and AKC.
Table AKA. Probe Name Ag3542
Table AKC. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3542 Expression of the CG59315-01 gene is low/undetectable in all the samples on this panel. (CTs>35). (Data not shown.)
General_screening_panel_vl.4 Summary: Ag3542 Expression of the CG59315-01 gene is highest in a breast cancer cell line (CT=31.3). Furthermore, there is significant expression in a cluster of cell lines derived from brain cancer, colon cancer and ovarian cancer. Therefore, expression of this gene could be used to differentiate between these samples and other samples on this panel and as a marker to detect the presence of colon, brain, ovarian, and breast cancer. Furthermore, therapeutic modulation of the expression or function of this gene may be effective in the treatment of colon, brain, ovarian, and breast cancers.
Low but significant levels of expression are also seen in the cerebellum. Therefore, therapeutic modulation of the expression or function of this gene may be useful in the tieatment of neurologic disorders, such as Alzheimer's disease, Parkinson's disease, schizophrenia, multiple sclerosis, stroke and epilepsy.
Among metabolic tissues, this gene is expressed at low levels in adipose. Therefore, this gene product may be useful in the treatment of obesity.
Panel 4D Summary: Ag3542 Expression of the CG59315-01 gene is highest in the normal colon (CT=30). Furthermore, expression is undetectable in colon samples from Crohn's and colitis patients. Thus, expression of this gene could be used to differentiate between normal and inflammed colon. This gene encodes a connexin homolog, a gap junction protein involved in intercellular communication.
The expression of this connexin- like protein in several of the resting and activated T lymphocyte preparations and in resting monocytes suggests that small molecule antagonists or therapeutic antibodies that block its function may also be useful in the tieatment of a number of inflammatory and autoimmune diseases in which T cells and monocytes play a pivotal role.
These include, but are not limited to, Crohn's disease, ulcerative colitis, multiple sclerosis, chronic obstructive pulmonary disease, asthma, emphysema, rheumatoid arthritis, lupus erythematosus, or psoriasis.
References:
1. Kwak BR, Mulhaupt F, Veillard N, Gros DB, Mach F. Altered pattern of vascular connexin expression in atherosclerotic plaques. Arterioscler Thromb Vase Biol 2002 Feb l;22(2):225-30
AL. CG59203-01: Lysozyme C-like protein
Expression of gene CG59203-01 was assessed using the primer-probe set Ag3392, described in Table ALA. Results of the RTQ-PCR runs are shown in Tables ALB and ALC.
Table ALA. Probe Name Ag3392
Table ALB. General_screening_panel_vl.4
CΝS_neurodegeneration_vl.0 Summary: Ag3392 Expression of the CG59203-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
General_screening_panel_vl.4 Summary: Ag3392 Highest expression of the CG59203-01 gene is seen in the testis. Thus, expression of this gene could be used as a marker of testicular tissue. Furthermore, therapeutic modulation of the expression or function of this gene may be effective in treating infertility or hypogonadism.
Panel 4D Summary: Ag3392 Significant expression of this gene is detected in a liver cirrhosis sample (CT = 33.8). Furthermore, expression of this gene is not detected in normal liver in Panel 1.3D, suggesting that its expression is unique to liver cirrhosis. Therefore, therapeutic modulation of the expression or function of this gene may reduce or inhibit fibrosis that occurs in liver cirrhosis. In addition, expression of this gene could also be used for the diagnosis of liver cirrhosis.
AM. CG58662-01: cytoplasmic protein
Expression of gene CG58662-01 was assessed using the primer-probe set Ag3387, described in Table AMA. Results of the RTQ-PCR runs are shown in Tables AMB, AMC and AMD.
Table AMA. Probe Name Ag3387
Table AMB. CNS_neurodegeneration_vl.O
Rel. Exp.(%) Ag3387, Rel. Exp.(%) Ag3387,
Tissue Name Tissue Name Run 210155038 Run 210155038
Contiol (Path) 3
AD 1 Hippo 15.7 7.3 Temporal Ctx
Table AMD. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3387 This panel does not show differential expression of the CG58662-01 gene in Alzheimer's disease. However, this expression profile confirms the presence of this gene in the brain. Please see Panel 1.3D for discussion of utility of this gene in the cential nervous system.
General_screening_panel_vl.4 Summary: Ag3387 Expression of the CG58662-01 gene is ubiquitous in this panel, with highest expression in a lung cancer cell line (CT=29.5). In addition, this gene is expressed at higher levels in kidney cancer cell lines when compared to normal kidney expression. Thus, expression of this gene could be used to differentiate these samples from other samples and as a marker for these cancers. Furthermore, therapeutic modulation of the expression of function of this gene may be effective in the treatment of lung and kidney cancer.
Among metabolic tissues this gene is expressed at moderate to low levels in adipose, adrenal gland, pancreas, pituitary, and adult and fetal skeletal muscle, heart and liver. This widespread expression among these tissues suggests that this gene plays a role in normal metabolic and neuroendocrine function and that disregulated expression of this gene may contribute to neuroendocrine diseases or metabolic disorders, such as obesity and diabetes.
In addition, this gene is expressed at moderate to low levels in all CNS regions examinded and may be a small molecule target for the tieatment of neurologic diseases, such as Alzheimer's disease, Parkinson's disease, schizophrenia, multiple sclerosis, stroke and epilepsy.
Panel 4D Summary: Ag3387 The CG58662-01 gene is expressed at high to moderate levels in a wide range of cell types of significance in the immune response in health and disease, with highest expression in the thymus (CT=31). In addition, expression is seen in members of the T-cell, B-cell, endothelial cell, macrophage/monocyte, and peripheral blood mononuclear cell family, as well as epithelial and fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product may be involved in homeostatic processes for these and other cell types and tissues. This pattern is in agreement with the expression profile in General_screening_panel_vl .5 and also suggests a role for the gene product in cell survival and proliferation. Therefore, modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
AN. CG59371-01: Novel cytoplasmic protein
Expression of gene CG59371-01 was assessed using the primer-probe set Ag3558, described in Table ANA. Results of the RTQ-PCR runs are shown in Tables ANB, ANC, AND and ANE.
Table ANA. Probe Name Ag3558
Table ANB. General_screening_panel_vl.4
Kidney Pool j 0.1 Adrenal Gland J 0.1
Fetal Kidney } 4.6 Pituitary gland Pool j 0.0
Renal ca. 786-0 j 44.1 Salivary Gland j 0.0
Renal ca. A498 j 4.2 Thyroid (female) j 0.1
Renal ca. ACHN j 15.2 Pancreatic ca. CAPAN2] 48.3
Renal ca. UO-31 j 20.4 Pancreas Pool j 0.5
Table AND. Panel 2.2
Table ANE. Panel 4D
CNS_neurodegeneration_vl.0 Summary: Ag3558 Expression of the CG59371-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
General_screening_panel_vl.4 Summary: Ag3558 Highest expression of the CG59371-01 gene is seen in a breast cancer cell line (CT=23.4). Overall, expression of this gene is significantly higher in cancer cell lines and fetal derived tissues than in samples derived from normal adult tissues. There are significant levels of expression in clusters of cell lines derived from pancreatic, brain, colon, gastric, renal, lung, ovarian, breast and melanoma cancers. Thus, expression of this gene in could be used to differentiate between the cancer derived samples and fetal tissues from other samples on this panel and as a marker to detect the presence of cancer. Furthermore, the much higher levels of expression in proliferative tissue suggest that this gene may be involved in cell proliferation. Therefore, therapeutic modulation of the expression or function of this gene may be effective in the tieatment of these cancers.
Among tissues with metabolic function, this gene is expressed at moderate to low levels in pituitary, adipose, adrenal gland, pancreas, thyroid, and adult and fetal skeletal muscle, heart, and liver. This widespread expression among these tissues suggests that this gene product may play a role in normal neuroendocrine and metabolic and that disregulated
expression of this gene may contribute to neuroendocrine disorders or metabolic diseases, such as obesity and diabetes.
This molecule is a novel protein phosphatase expressed at moderate to low levels in all regions of the CNS examined. Therefore, therapeutic modulation of the expression or function of this gene may be useful in the treatment of neurologic disorders, such as Alzheimer's disease, Parkinson's disease, schizophrenia, multiple sclerosis, stroke and epilepsy.
General_screening_panel_vl.5 Summary: Ag3558 Results from this experiment are in excellent agreement with results from Panel 1.4. Please see that panel for discussion of utility of this gene in cancer, metabolic disorders and the central nervous system.
Panel 2.2 Summary: Ag3558 Two experiments with the same probe and primer produce results that are in excellent agreement, with highest expression of the CG59371-01 gene in colon cancer (CTs=30). Furthermore, expression is higher in kidney, lung, ovary and colon cancers when compared to normal adjacent tissue. In addition, significant expression is also seen in gastric, breast, and bladder cancer. Thus, , expression of this gene in could be used to differentiate between the cancer derived samples and other samples on this panel and as a marker to detect the presence of cancer. Furthermore, therapeutic modulation of the expression or function of this gene may be effective in the treatment of these cancers.
Panel 4D Summary: Ag3558 The CG59371-01 gene is widely expressed among the samples on this panel, with highest expression in dermal fibroblasts treated with TNF-alpha. Significant levels of expression are also seen in tieated and untieated samples from skin, lung, T-cells and B-cells. Therefore, modulation of the expression or activity of the protein encoded by this tianscript through the application of antibodies or peptides therapeutics may be beneficial for the treatment of lung inflammatory diseases such as asthma, and chronic obstructive pulmonary diseases, inflammatory skin diseases such as psoriasis, atopic dermatitis, ulcerative dermatitis, and ulcerative colitis, autoimmune diseases such as Crohn's disease, lupus erythematosus, rheumatoid arthritis and osteoarthritis and in other diseases in which T cells and B cells are activated.
AO. CG59346-01: Cortactin-binding protein 1
Expression of gene CG59346-01 was assessed using the primer-probe set Ag3550, described in Table AOA. Results of the RTQ-PCR runs are shown in Tables AOB, AOC and AOD.
Table AOA. Probe Name Ag3550
Table AOB. CNS_neurodegeneration_vl.0
Table AOC. General_screening_panel_vl.4
Table AQD. Panel 4D
CNS_neurodegeneration_vl.0 Summary: Ag3550 This panel does not show differential expression of the CG59346-01 gene in Alzheimer's disease. However, this expression profile confirms the presence of this gene in the brain. Please see Panel 1.4 for discussion of utility of this gene in the central nervous system.
General_screening_panel_vl.4 Summary: Ag3550 Highest expression of the CG59346-01 gene is seen in the brain. Expression of this gene is seen at high levels in the cerebellum, cerebral cortex, and thalamus and at moderate levels in the amygdala, hippocampus, and
thalamus. This CG59346-01 gene encodes a homologue of Pro line-rich synapse-associated protein- 1/cortactin binding protein 1 (ProSAPl/CortBPl). ProSAPl is PDZ-domain protein highly enriched in the postsynaptic density (PSD) and involved in in the assembly of the PSD during neuronal differentiation that may function with contactin, in the recruitment and activation of neural intracellular signaling pathways. Therefore, this gene may play a role in central nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
In addition, moderate levels of expression are seen in colon, gastric, renal, pancreatic, lung, ovarian, breast and prostate cancer cell lines. Thus, expression of this gene could be used to detect the presence of cancer. Furthermore, therapeutic modulation of the expression or function of this gene may be effective in the treatment of these cancers.
Among tissues with metabolic function, this gene is expressed at moderate to low levels in pituitary, adipose, adrenal gland, pancreas, thyroid, and adult and fetal and liver. This widespread expression among these tissues suggests that this gene product may play a role in normal neuroendocrine and metabolic and that disregulated expression of this gene may contribute to neuroendocrine disorders or metabolic diseases, such as obesity and diabetes.
In addition, this gene is expressed at higher levels in fetal lung and kidney (CTs=29) when compared to expression in adult lung and kidney (CTs=35-40). Thus, expression of this gene could be used to differentiate between the two sources of lung and kidney tissue.
References:
1. Peles E, Nativ M, Lustig M, Grumet M, Schilling J, Martinez R, Plowman GD, Schlessinger J. Identification of a novel contactin-associated transmembrane receptor with multiple domains implicated in protein-protein interactions. EMBO J 1997 Mar 3;16(5):978- 88.
2. Boeckers TM, Kreutz MR, Winter C, Zuschratter W, Smalla KH, Sanmarti-Vila L, Wex H, Langnaese K, Bockmann J, Gamer CC, Gundelfinger ED. (1999) Proline-rich synapse-associated protein- 1/cortactin binding protein 1 (ProSAPl/CortBPl) is a PDZ-domain protein highly enriched in the postsynaptic density. J Neurosci 1999 Aug 1;19(15):6506-18.
Panel 4D Summary: Ag3550 Highest expression of the CG59346-01 gene is seen in thymus (CT=27). In addition, significant levels of expression are seen in IL-4, IL-9, EL- 13 and EFN gamma activated-NCI-H292 mucoepidermoid cells as well as untreated NCI-H292 cells. Moderate/low expression is also detected in IL-4, IL-9, IL-13 and IFN gamma activated lung fibroblasts, small airway epithelium (treated and untieated), and treated bronchial epithelium. The expression of this gene in cells derived from or within the lung suggests that this gene may be involved in normal conditions as well as pathological and inflammatory lung disorders that include chronic obstructive pulmonary disease, asthma, allergy and emphysema.
In addition, significant levels of expression are seen in treated and untreated dermal fibroblasts and keratinocytes, suggesting that modulation of the expression or function of this gene may also reduce symtptoms in inflammatory skin diseases such as psoriasis, atopic dermatitis, and ulcerative dermatitis.
AP. CG57814-01 and CG57814-02: Basic 1 19 protein
Expression of gene CG57814-01 and varian CG57814-02 was assessed using the primer-probe set Ag791, described in Table APA.
Table APA. Probe Name Ag791
Panel 1.2 Summary: Ag791 Expression of the CG57814-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
AQ. CG59327-01: MONOCARBOXYLATE TRANSPORTER 1 like protein
Expression of gene CG59327-01 was assessed using the primer-probe set Ag3548, described in Table AQA. Results of the RTQ-PCR runs are shown in Tables AQB and AQC.
Table AOA. Probe Name Ag3548
Table AOB. General_screening_panel_vl.4
Table AOC. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3548 Expression of the CG59327-01 gene is low/undetectable in all the samples on the panel (CTs>35). (Data not shown.)
General_screening_panel_vl.4 Summary: Ag3548 Significant expression of the CG59327- 01 gene is restricted to a sample derived from a kidney cancer cell line (CT=33.34). Thus, expression of this gene could be used to differentiate between this sample and other samples on this panel and as a marker to detect the presence of kidney cancer. Furthermore, therapeutic modulation of the expression or function of this gene may be effective in the treatment of kidney cancer.
Panel 4D Summary: Ag3548 Significant expression of the CG59327-01 gene is restricted to a samples derived from untreated microvascular dermal endothelial cells (CT=30.3). Thus, expression of this gene could be used as a marker of these cells.
AR. CG59494-01: Myelin P2
Expression of gene CG59494-01, which represents a full length physical clone, was assessed using the primer-probe set Ag3206, described in Table ARA. Results of the RTQ- PCR runs are shown in Tables ARB and ARC.
Table ARA. Probe Name Ag3206
Table ARB. Panel 1.3D
Table ARC. Panel 4D
Panel 1.3D Summary: Ag3206 Expression of the CG59494-01 gene is restricted to a sample derived from a prostate cancer cell line (CT=34.9). Thus, expression of this gene could be used to differentiate between this sample and other samples on this panel and as a marker to
detect the presence of prostate cancer. Furthermore, therapeutic modulation of the expression or function of this gene may be effective in the tieatment of prostate cancer.
Panel 4D Summary: Ag3206 Expression of the CG59494-01 gene is primarily restricted to a cluster of samples derived from microvasculature of the lung and the dermis suggesting a role for this gene in the maintenance of the integrity of the microvasculature. Therefore, therapeutics designed for this putative protein could be beneficial for the treatment of diseases associated with damaged microvasculature including heart diseases or inflammatory diseases, such as psoriasis, asthma, and chronic obstructive pulmonary diseases.
AS. CG59432-01 and CG59432-02: Chloride Channel
Expression of gene CG59432-01 and CG59432-02 was assessed using the primer- probe set Ag5938, described in Table ASA. Results of the RTQ-PCR runs are shown in Tables ASB and ASC. Please note that CG59432-02 represents a full-length physical clone of CG59432-01 gene, validating the prediction of the gene sequence.
Table ASA. Probe Name Ag5938
Table ASB. General_screening_panel_vl.5
Table ASC. Panel 5 Islet
General_screening_panel_vl.5 Summary: Ag5938 Highest expression of the CG59432-01 gene is seen in a gastric cancer cell line (CT=32.5). Thus, expression of this gene could be used to differentiate between this sample and other samples on this panel. In addition, low expression of this gene is seen in colon cancer CaCo-2, lung cancer NCI-H526, ovarian cancer OVCAR-5, and squamous cell carcinoma SCC-4 cell lines. Therefore, therapeutic modulation of the activity of this gene or its protein product, through the use of small molecule drugs, protein therapeutics or antibodies, might be beneficial in the treatment of these cancers.
Significant expression is also detected in fetal skeletal muscle and adult skeletal muscle (CT=32.5). At least 50 disease-causing mutations in the skeletal muscle voltage-gated chloride channel gene (CLCN1), almost all of which originate from Caucasian families, have been identified. Therefore, therapeutic modulation of this gene product, a chloride channel homolog, may be a tieatment for myotonia congenita and other muscle channelopafhies.
In addition, this gene is expressed at low levels in most regions of the cential nervous system examined, including amygdala, substantia nigra, thalamus, and cerebral cortex.
Therefore, this gene may play a role in central nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
References:
1. Sasaki R, Ito N, Shimamura M, Murakami T, Kuzuhara S, Uchino M, Uyama E. A novel CLCN1 mutation: P480T in a Japanese family with Thomsen's myotonia congenita. Muscle Nerve. 2001 Mar;24(3):357-63.
Panel 5 Islet Summary: Ag5938 Expression of the CG59432-01 is restricted to a sample from small intestine (CT=31.6). Thus, expression of this gene could be used to differentiate between this sample and other samples on this panel and as a marker for this tissue.
AT. CG59383-01: D6MM5E
Expression of gene CG59383-01 was assessed using the primer-probe set Ag3427, described in Table ATA. Results of the RTQ-PCR runs are shown in Tables ATB, ATC and ATD.
Table ATA. Probe Name Ag3427
Table ATB. CNS_neurodegeneration_vl.0
Table ATC. General_screening_panel_vl.4
Table ATP. Panel 4D
CNS_neurodegeneration_vl.0 Summary: Ag3427 This panel confirms the expression of CG59383-01 gene at low levels in the brains of an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in treatment of central nervous system disorders.
General_screening_panel_vl.4 Summary: Ag3427 Highest expression of the CG59383-01 gene is seen in a colon cancer cell line (CT=27.2). Significant expression is also seen in a cluster of samples derived from ovarian cancer cell lines. Thus, expression of this gene could be used to differentiate between these samples and other samples on this panel and as a marker for the presence of these cancers. Furthermore, therapeutic modulation of the expression or function of this gene may be effective in the treatment of ovarian or colon cancers.
This molecule is also expressed at low levels in all regions of the CNS examined. Therefore, therapeutic modulation of the expression or function of this gene may be useful in the treatment of neurologic disorders, such as Alzheimer's disease, Parkinson's disease, schizophrenia, multiple sclerosis, stroke and epilepsy.
Among tissues with metabolic function, this gene is expressed at low levels in adipose and pancreas. This expression suggests that this gene product may play a role in normal neuroendocrine and metabolic and that disregulated expression of this gene may contribute to neuroendocrine disorders or metabolic diseases, such as obesity and diabetes
Panel 4D Summary: Ag3427 Highest expression of the CG59383-01 gene is seen in keratinocytes tieated with the inflammatory cytokines TNF-alpha and IL-1 beta (CT=30.3). Therefore, modulation of the expression or activity of the protein encoded by this tianscript through the application of small molecule therapeutics may be useful in the treatment of asthma, COPD, emphysema, psoriasis and wound healing.
AU. CG58526-01: Scramblase
Expression of gene CG58526-01 was assessed using the primer-probe set Ag3366, described in Table AUA. Results of the RTQ-PCR runs are shown in Table AUB.
Table AUA. Probe Name Ag3366
Table AUB. General_screenmg_panel_vl.4
CNS_neurodegeneration_vl.O Summary: Ag3366 Expression of the CG58526-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
General_screening_panel_vl.4 Summary: Ag3366 Expression of the CG58526-01 gene is restricted to a sample derived from a colon cancer cell line (CT=34.5) and the testis. Thus, expression of this gene could be used to differentiate between these samples and other samples on this panel and as a marker to detect the presence of colon cancer. Furthermore, therapeutic modulation of the expression or function of this gene may be effective in the treatment of colon cancer.
Panel 4D Summary: Ag3366 Results from one experiment with the CG58526-01 gene are not included. The amp plot indicates that there were experimental difficulties with this run.
AV. CG57851-01: sulfotransferase
Expression of gene CG57851-01 was assessed using the primer-probe set Ag3349, described in Table AVA. Results of the RTQ-PCR runs are shown in Tables AVB, AVC and AVD.
Table AVA. Probe Name Ag3349
Table AVB. CNS_neurodegeneration_vl.0
Table AVC. General_screening_panel_vl.4
Table AVD. Panel 4D
Monocytes rest 0.0 EBD Crohn's 0.0
Monocytes LPS 1.3 Colon 0.6
Macrophages rest 0.9 Lung 0.7
Macrophages LPS 0.2 Thymus 100.0
HUVEC none 0.0 Kidney 1.7
HUVEC starved 0.0
CNS_neurodegeneration_vl.O Summary: Ag3349 This panel confirms the expression of CG57851-01 gene at low levels in the brains of an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. The expression of this gene in the brain suggests that therapeutic modulation of the expression or function of this gene may be useful in the treatment of neurologic disorders, such as Alzheimer's disease, Parkinson's disease, schizophrenia, multiple sclerosis, stroke and epilepsy.
General_screening_panel_vl.4 Summary: Ag3349 Highest expression of the CG57851-01 gene is seen in a lung cancer cell line (CT=30). Thus, expression of this gene may be used to differentiate between this sample and other samples on this panel and as a marker for lung cancer. This gene encodes a sulfotransferase homolog. Sulfotiansferases are involved in the metabolism of drugs and endogenous compounds in the body and also synthesize the complex glycoproteins found on the cell surface of cancer cells. Therefore, therapeutic modulation of the expression or function of this gene may be effective in the treatment of lung cancer.
Among tissues with metabolic function, this gene is expressed at moderate to low levels in adipose and pancreas. This expression among these tissues suggests that this gene product may play a role in normal metabolic function and that disregulated expression of this gene may contribute to metabolic diseases, such as obesity and diabetes.
Panel 4D Summary: Ag3349 Highest expression of the CG57851-01 gene is seen in the thymus (CT=29.7). The putative protein encoded by this gene could therefore play an important role in T cell development. Small molecule therapeutics designed against the protein encoded by this gene could be utilized to modulate immune function (T cell development) and be important for organ transplant, AEDS treatment or post chemotherapy immune reconstitiution.
Panel 5 Islet Summary: Ag3349 Expression of the CG57851-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
AW. CG59258-01: KIAA1 08 protein
Expression of gene CG59258-01 was assessed using the primer-probe set Ag3520, described in Table AWA.
Table AWA. Probe Name Ag3520
CNS_neurodegeneration_vl.0 Summary: Ag3520 Expression of the CG59258-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
General_screening_panel_vl.4 Summary: Ag3520 Expression of the CG59258-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
Panel 4D Summary: Ag3520 Expression of the CG59258-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
AX. CG59564-01: Sorting nexin 6
Expression of gene CG59564-01 was assessed using the primer-probe set Ag3471, described in Table AXA. Results of the RTQ-PCR runs are shown in Tables AXB, AXC and AXD.
Table AXA. Probe Name Ag3471
Table AXB. CNS_neurodegeneration_vl.0
Table AXC. General_screening_panel_vl.4
Tissue Name Rel. Exp.(%) Ag3471, Tissue Name Rel. Exp.(%) Ag3471,
CNS_neurodegeneration_vl.0 Summary: Ag3471 This panel does not show differential expression of the CG59564-01 gene in Alzheimer's disease. However, this expression profile confirms the presence of this gene in the brain. Please see Panel 1.4 for discussion of utility of this gene in the central nervous system.
General_screening_panel_vl.4 Summary: Ag3471 The CG59564-01 gene, a sorting nexin homolog, shows highly brain preferential expression. Moderate levels of expression are seen in all brain regions examined, with highest expression in the fetal brain (CT=28.5). Thus, this gene would be useful for distinguishing brain tissue from non-neural tissue, and may be beneficial as a drug target in neurologic disease, such as Alzheimer's disease, Parkinson's disease, schizophrenia, multiple sclerosis, stroke and epilepsy.
Among tissues with metabolic function, this gene is expressed at low levels in pituitary, adipose, adrenal gland, pancreas, and adult and fetal skeletal muscle, heart, and liver. This widespread expression among these tissues suggests that this gene product may play a role in normal neuroendocrine and metabolic and that disregulated expression of this gene may contribute to neuroendocrine disorders or metabolic diseases, such as obesity and diabetes.
In addition, this gene is expressed at significant levels in a breast cancer cell line (CT=28.6). Thus, expression of this gene could be used to differentiate this sample from other samples on this panel and as a marker for breast cancer.
Panel 4D Summary: Ag3471 The CG59564-01 gene, a sorting nexin homolog, is most highly expressed in normal colon (CT=30). In addition, this gene is expressed at high to moderate levels in a wide range of cell types of significance in the immune response in health and disease. These cells include members of the T-cell, B-cell, endothelial cell, macrophage/monocyte, and peripheral blood mononuclear cell family, as well as epithelial and fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product may be involved in homeostatic processes for these and other cell types and tissues. This pattern is in agreement with the expression profile in General_screening_panel_vl.4 and also suggests a role for the gene product in cell survival and proliferation. Therefore, modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
AY. CG59553-01: Secretory protein SEC8
Expression of gene CG59553-01 was assessed using the primer-probe set Ag3465, described in Table AYA. Results of the RTQ-PCR runs are shown in Tables AYB, AYC and AYD.
Table AYA. Probe Name Ag3465
Table AYB. CNS_neurodegeneration_ vl.O
Table AYC. General_screening_panel_vl.4
Table AYD. Panel 4D
CNS_neurodegeneration_vl.0 Summary: Ag3465 This panel does not show differential expression of the CG59553-01 gene in Alzheimer's disease. However, this expression profile confirms the presence of this gene in the brain. Please see Panel 1.4 for discussion of utility of this gene in the cential nervous system.
General_screening_panel_vl.4 Summary: Ag3465 Highest expression of the CG59553-01 gene is seen in a brain cancer cell line (CTs=24). Expression of this gene is ubiquitous throughout this panel, with significant levels of expression in clusters of cell lines derived from brain, renal, colon, lung, breast, ovarian, and melanoma cancers. These high levels of expression in all the samples on this panel suggest a role for this gene in cell growth and proliferation.
This molecule is also expressed at high levels in all regions of the CNS examined. Therefore, therapeutic modulation of the expression or function of this gene may be useful in the tieatment of neurologic disorders, such as Alzheimer's disease, Parkinson's disease, schizophrenia, multiple sclerosis, stroke and epilepsy.
Among tissues with metabolic function, this gene is expressed at high to moderate levels in pituitary, adipose, adrenal gland, pancreas, thyroid, and adult and fetal skeletal muscle, heart, and liver. This widespread expression among these tissues suggests that this gene product may play a role in normal neuroendocrine and metabolic and that disregulated expression of this gene may contribute to neuroendocrine disorders or metabolic diseases, such as obesity and diabetes.
Panel 4D Summary: Ag3465 The CG59553-01 gene is expressed at high to moderate levels in a wide range of cell types of significance in the immune response in health and disease.
These cells include members of the T-cell, B-cell, endothelial cell, macrophage/monocyte, and peripheral blood mononuclear cell family, as well as epithelial and fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product may be involved in homeostatic processes for these and other cell types and tissues. This pattern is in agreement with the expression profile in General_screening_panel_vl.5 and also suggests a role for the gene product in cell survival and proliferation. Therefore, modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
AZ. CG59435-01 and CG59435-02: Human Neddl
Expression of gene CG59435-01 and CG59435-02 was assessed using the primer- probe set Ag3437, described in Table AZA. Results of the RTQ-PCR runs are shown in Tables AZB, AZC and AZD. Please note that CG59435-02 represents a full-length physical clone of the CG59435-01 gene, validating the prediction of the gene sequence.
Table AZA. Probe Name Ag3437
Table AZB. CNS neurodegeneration vl.O
Table AZC. General_screening_panel_yl.4
Table AZD. Panel 4. ID
CNS_neurodegeneration_vl.O Summary: Ag3437 This panel confirms the expression of CG59435-01 gene at low levels in the brains of an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in treatment of central nervous system disorders.
General_screening_panel_vl.4 Summary: Ag3437 The CG59435-01 is gene is ubiquitously exrpressed in this panel, with highest expression in a gastric cancer cell line (CT=26.5). In addition, significant levels of expression are evident in cell lines from brain cancer, colon cancer, ovarian cancer, breast cancer, prostate cancer and lung cancer. Thus, expression of this gene could be used as a marker to detect the presence of these cancers. Furthermore, therapeutic modulation of the expression or function of this gene may be effective in the tieatment of these cancers.
In addition, this gene is expressed at moderate to low levels in pituitary, adipose, adrenal gland, pancreas, thyroid, and adult and fetal skeletal muscle, heart, and liver. This widespread expression among metabolic tissues suggests that this gene product may play a role in normal neuroendocrine and metabolic and that disregulated expression of this gene may contribute to neuroendocrine disorders or metabolic diseases, such as obesity and diabetes.
In addition, the CG59435-01 gene encodes a homologue of mouse NEDD1 protein.
Nedd is an acronym of "neural precursor cell expressed developmentally and down-regulated"
(Ref 1) The developmentally regulated mouse gene Neddl encodes a protein with similarities to the beta subunit of heterotrimeric GTP-binding proteins that has growth suppressing activity when overexpressed in various cultured cell types. Neddl mRNA is shown to be strongly expressed in early embryonic brain and may play a role in the differentiation-coupled growth arrest in neuronal cells (Ref. 2). The moderate to low levels (CT=30-33) in all regions of the central nervous system examined suggest that this gene product may also play a role in the
differentiation-coupled growth arrest in neuronal cells.Furthermore, this gene may play a role in central nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
References:
1. Kumar S, Tomooka Y, Noda M. (1992) Identification of a set of genes with developmentally down-regulated expression in the mouse brain. Biochem Biophys Res Commun 185(3): 1155-61
2. Kumar S, Matsuzaki T, Yoshida Y, Noda M. (1994) Molecular cloning and biological activity of a novel developmentally regulated gene encoding a protein with beta- transducin-like structure. J Biol Chem 269(15):11318-26.
Panel 4.1D Summary: Ag3437 The CG59435-01 is gene is ubiquitously exrpressed in this panel, with highest expression in the basophil cell line KU-812 treated with PMA/ionomycin (CT=27.9). This gene is expressed at high to moderate levels in a wide range of cell types of significance in the immune response in health and disease. These cells include members of the T-cell, B-cell, endothelial cell, macrophage/monocyte, and peripheral blood mononuclear cell family, as well as epithelial and fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product may be involved in homeostatic processes for these and other cell types and tissues. This pattern is in agreement with the expression profile in General_screening__panel_vl.4 and also suggests a role for the gene product in cell survival and proliferation. Therefore, modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
BA. CG59439-01 and CG59439-02: Xenobiotic/medium-chain fatty acid:CoA ligase form XL-HI
Expression of gene CG59439-01 was assessed using the primer-probe set Ag3438, described in Table BAA. Results of the RTQ-PCR runs are shown in Table BAB. Please note that CG59439-02 represents a full-length physical clone of the CG59439-01 gene, validating the prediction of the gene sequence.
Table BAA. Probe Name Ag3438
Table BAB. Panel 4. ID
CNS neurodegeneration vl.O Summary: Ag3438 Expression of the CG59439-02 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
General_screening_panel_vl.4 Summary: Ag3438 Results from one experiment with the CG59439-02 gene are not included. The amp plot indicates that there were experimental difficulties with this run.
Panel 4.1D Summary: Ag3438 Expression of the CG59439-02 gene is restricted to a sample derived from chronically activated Th2 cells (CT=33).
Panel 4D Summary: Ag3438 Results from one experiment with the CG59439-02 gene are not included. The amp plot indicates that there were experimental difficulties with this run.
BB. CG59354-01 and CG59354-02 and CG59354-03: phosducin-like protein
Expression of gene CG59354-01 and variant CG59354-02 was assessed using the primer-probe set Ag3553, described in Table BBA. Results of the RTQ-PCR runs are shown in Tables BBB, BBC and BBD. Please note that CG59354-03 represents a full-length physical clone of the CG59354-01 gene, validating the prediction of the gene sequence.
Table BBB. CNS_neurodegeneration_vl.O
Table BBC. General_screening_panel_vl.4
CNS_neurodegeneration_vl.O Summary: Ag3553 This panel confirms the expression of CG59354-03 gene at low levels in the brains of an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in tieatment of central nervous system disorders.
General_screening_panel_vl.4 Summary: Ag3553 The CG59354-03 gene is ubiquitously expressed in this panel, with highest expression in a brain cancer cell line (CT=25.9). In addition, significant levels of expression are seen in cell lines derived from colon, breast, ovarian, renal, lung, prostate, and melanoma cancers. Furthermore, higher levels of expression are seen in fetal liver and lung (CTs=27-28) when compared to expression in the adult tissues (CTs=30-33). The high levels of expression in fetal tissue and cancer cell lines, both of which are highly proliferative, suggests that this gene product may be involved in cell growth and differentiation. Therefore, therapeutic modulation of the expression or function of this gene may be effective in the treatment of cancer.
Among tissues with metabolic or endocrine function, this gene is expressed at high to moderate levels in pancreas, adipose, adrenal gland, thyroid, pituitary gland, skeletal muscle, heart, liver and the gastrointestinal tract. Therefore, therapeutic modulation of the activity of this gene may prove useful in the tieatment of endocrine/metabolically related diseases, such as obesity and diabetes.
In addition, this gene is expressed at high levels in all regions of the central nervous system examined, including amygdala, hippocampus, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. The CG59354-03 gene encodes a splice variant of
phosphoducin-like protein (PHLP). PDCL is a putative modulator of heterotrimeric G proteins. It was initially isolated as the product of an ethanol-responsive gene in neural cell cultures (Ref. 1). PDCL shares extensive amino acid sequence homology with phosducin (PDC), a phosphoprotein expressed in retina and pineal gland that inhibits several G protein- coupled signaling pathways by binding to the beta-gamma subunits of G proteins. Therefore, this gene may play a role in cential nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
References:
1. Miles MF, Barhite S, Sganga M, Elliott M. (1993) Phosducin-like protein: an ethanol-responsive potential modulator of guanine nucleotide-binding protein function. Proc Natl Acad Sci U S A 90(22): 10831-5
Panel 4D Summary: Ag3553 The CG59354-03 gene is ubiquitously expressed in this panel, with highest expression in B cells treated with polk-weed mitogen (CT=27.2). In addition, this gene is expressd at is expressed at high to moderate levels in a wide range of cell types of significance in the immune response in health and disease. These cells include members of the T-cell, B-cell, endothelial cell, macrophage/monocyte, and peripheral blood mononuclear cell family, as well as epithelial and fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product may be involved in homeostatic processes for these and other cell types and tissues. This pattern is in agreement with the expression profile in General_screening_panel_vl.4 and also suggests a role for the gene product in cell survival and proliferation. Therefore, modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
BC. CG59319-01 and CG59319-02: phosducin-like protein
Expression of gene CG59319-01 was assessed using the primer-probe set Ag3544, described in Table BCA. Results of the RTQ-PCR runs are shown in Tables BCB and BCC. Please note that CG59319-02 represents a full-length physical clone of the CG59319-01 gene, validating the prediction of the gene sequence.
Table BCA. Probe Name Ag3544
Table BCB. General_screening_panel_vl.4
CNS_neurodegeneration_vl.0 Summary: Ag3544 Expression of the CG59319-02 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
General_screening_panel_vl.4 Summary: Ag3544 Expression of the CG59319-02 gene is restricted to a sample derived from the testis (CT=29.8). Thus, expression of this gene could be used to differentiate between this sample and other samples on this panel and as a marker of testicular tissue. Furthermore, therapeutic modulation of the expression or function of this gene may be useful in the tieatment of male infertility or hypogonadism.
Panel 4.1D Summary: Ag3544 Expression of the CG59319-02 gene is restricted to samples derived from the basophil cell line KU-812 (CTs=32). Thus, expression of this gene could be used as a marker of this cell type. Furthermore, the specific pattern of expression of this gene suggests that therapeutic modulation of the expression or function of the protein encoded by
this gene may block or inhibit inflammation or tissue damage due to basophil activation in response to asthma, allergies, hypersensitivity reactions, psoriasis, and viral infections.
BD. CG59576-01: Olfactory Receptor
Expression of gene CG59576-01 was assessed using the primer-probe set Ag3478, described in Table BDA. Results of the RTQ-PCR runs are shown in Table BDB.
Table BDA. Probe Name Ag3478
Table BDB. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3478 Expression of the CG59576-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
General_screening_panel_vl.4 Summary: Ag3478 Expression of the CG59576-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
General_screening_panel_vl.5 Summary: Ag3478 Expression of the CG59576-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
Panel 4D Summary: Ag3478 Expression of the CG59576-01 gene is restricted to a sample derived from liver cirrhosis (CT=32.3). Furthermore, expression of this gene is not detected in normal liver in Panel 1.4, suggesting that its expression is unique to liver cirrhosis. This gene encodes a putative GPCR; therefore, antibodies or small molecule therapeutics could reduce or inhibit fibrosis that occurs in liver cirrhosis. In addition, antibodies to this putative GPCR could also be used for the diagnosis of liver cirrhosis.
Panel 5 Islet Summary: Ag3478 Expression of the CG59576-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
BE. CG59557-01: Olfactory Receptor
Expression of gene CG59557-01 was assessed using the primer-probe set Ag3470, described in Table BEA. Results of the RTQ-PCR runs are shown in Table BEB.
Table BEA. Probe Name Ag3470
Table BEB. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3470 Expression of the CG59557-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
General_screening_panel_vl.4 Summary: Ag3470 Expression of the CG59557-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
Panel 4D Summary: Ag3470 Expression of the CG59557-01 gene is detected in a liver cirrhosis sample (CT = 32.2). Furthermore, expression of this gene is not detected in normal liver in Panel 1.4, suggesting that its expression is unique to liver cirrhosis. This gene encodes
a putative GPCR; therefore, antibodies or small molecule therapeutics could reduce or inhibit fibrosis that occurs in liver cirrhosis. In addition, antibodies to this putative GPCR could also be used for the diagnosis of liver cirrhosis.
BF. CG59555-01: Olfactory Receptor
Expression of gene CG59555-01 was assessed using the primer-probe set Ag3467, described in Table BFA. Results of the RTQ-PCR runs are shown in Tables BFB, BFC and BFD.
Table BFA. Probe Name Ag3467
Table BFB. CNS_neurodegeneration_vl.O
Table BFC. General_screening_panel_vl.4
Kidney Pool j 34.4 Adrenal Gland j 6.3
Fetal Kidney } 76.3 Pituitary gland Pool | 4.5
Renal ca. 786-0 j 28.1 Salivary Gland | 1.8
Renal ca. A498 j 12.1 Thyroid (female) j 13.4
Renal ca. ACHN J 23.0 Pancreatic ca. CAPAN2] 1.0
Renal ca. UO-31 j 25.0 Pancreas Pool j 27.2
Table BFD. Panel 4D
CNS neurodegeneration vl.O Summary: Ag3467 The CG59555-01 gene encodes a putative GPCR. It is expressed at low to moderate levels in most of the samples used in this
panel. This panel confirms the expression of CG59555-01 gene at low levels in the brains of an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in tieatment of central nervous system disorders.
General_screening_panel_vl.4 Summary: Ag3467 The CG59555-01 gene encodes a putative GPCR. It is expressed at low to moderate levels in large number of the samples used in this panel. Highest expression of this gene is detected in fetal lung (CT=28). Interestingly, this gene is expressed at much higher levels in fetal (CT = 28) when compared to adult lung (CT = 31). Therefore, expression of this gene can be used to distinguish fetal lung from adult lung and from other samples used in this panel. In addition, this gene is also expressed at much higher levels in fetal fetal liver (CT=32) as compared to adult liver (CT=38). Thus, expression of this gene can be used to distinguish fetal liver from adult liver.
Among tissues with metabolic or endocrine function, this gene is expressed at low to moderate levels in pancreas, adipose, adrenal gland, thyroid, pituitary gland, skeletal muscle, heart, liver and the gastrointestinal tract. Therefore, therapeutic modulation of the activity of this gene may prove useful in the tieatment of endocrine/metabolically related diseases, such as obesity and diabetes.
This gene is also expressed at low levels in all regions of the central nervous system examined, including amygdala, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. Several neurotransmitter receptors are GPCRs, including the dopamine receptor family, the serotonin receptor family, the GABAB receptor, muscarinic acetylcholine receptors, and others; thus this GPCR may represent a novel neurotransmitter receptor. Targeting various neurotransmitter receptors (dopamine, serotonin) has proven to be an effective therapy in psychiatric illnesses such as schizophrenia, bipolar disorder, and depression. Furthermore, the cerebral cortex and hippocampus are regions of the brain that are known to be involved in Alzheimer's disease, seizure disorders, and in the normal process of memory formation. Therefore, this gene may play a role in central nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Panel 4D Summary: Ag3467 The CG59555-01 gene encodes a putative GPCR. Highest expression of this gene is detected in resting primary Thl cells (CT=27). This gene is expressed at moderate levels in a wide range of cell types of significance in the immune response in health and disease. These cells include members of the T-cell, B-cell, endothelial cell, macrophage/monocyte, and peripheral blood mononuclear cell family, as well as epithelial and fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product may be involved in homeostatic processes for these and other cell types and tissues. This pattern is in agreement with the expression profile in General_screening_panel_vl.4 and also suggests a role for the gene product in cell survival and proliferation. Therefore, modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
BG. CG59551-01: Olfactory Receptor
Expression of gene CG59551-01 was assessed using the primer-probe set Ag3463, described in Table BGA. Results of the RTQ-PCR runs are shown in Tables BGB and BGC.
Table BGA. Probe Name Ag3463
Table BGB. General_screening_panel_vl.4
Table BGC. Panel 4. ID
CNS_neurodegeneration_vl.O Summary: Ag3463 Expression of the CG59551-01 gene is low/undetectable in all the samples on this panel. (Data not shown.)
General_screening_panel_vl.4 Summary: Ag3463 The CG59551-01 gene encodes a putative GPCR. Highest expression of this gene is detected in an ovarian cancer cell line SK- OV-3 (CT=34). In addition, low expression of this gene is also observed in fetal skeletal muscle (CT= 34.4), one of the lung cancer cell line (CT= 34.9), and testis (CT= 34.3). Thus, expression of this gene can be used to distinguish these sample from other samples used in this panel. In addition, therapeutic modulation of the activity of the GPCR encoded by this gene may be useful in the treatment of ovarian and lung cancer, fertility, hypogonadism, and muscle related diseases.
Panel 4.1D Summary: Ag3463 The CG59551-01 gene encodes a putative GPCR. Highest expression of this gene is seen in KU-812 cells treated with PMA/ionomycin (CT=30.86). Thus, expression of this gene can be used to distinguish this sample from other samples used in this panel. In addition, expression of this gene is high in KU-812 (basophils) cells treated with PMA/ionomycin (CT=30.86) as compared to resting KU-812 cells (CT=34.66). Therefore, expression of this gene can be used to distinguish resting from PMA/ionomycin treated- basophils. It is known that GPCR-type receptors are important in multiple physiological responses mediated by basophils (ref. 1). Therefore, antibody or small molecule therapies designed with the protein encoded for by this gene could block or inhibit inflammation or tissue damage due to basophil activation in response to asthma, allergies, hypersensitivity reactions, psoriasis, and viral infections.
References:
1. Heinemann A., Hartnell A., Stubbs V.E., Murakami K., Soler D., LaRosa G., Askenase P.W., Williams T.J., Sabroe I. (2000) Basophil responses to chemokines are regulated by both sequential and cooperative receptor signaling. J. Immunol. 165: 7224-7233.
BH. CG59540-01: OLFACTORY RECEPTOR
Expression of gene CG59540-01 was assessed using the primer-probe sets Ag3460 and Agl 519, described in Tables BHA and BHB. Results of the RTQ-PCR runs are shown in Tables BHC, BHD and BHE.
Table BHA. Probe Name Ag3460
Table BHB. Probe Name Agl 519
Table BHC. Panel 1.2
Table BHD. Panel 1.3D
CNS_neurodegeneration_vl.0 Summary: Ag3460 Expression of the CG59540-01 gene is low/undetectable (CT values > 35) across the samples in this panel.
General_screening_panel_vl.4 Summary: Ag3460 Expression of the CG59540-01 gene is low/undetectable (CT values > 35) across the samples in this panel.
Panel 1.2 Summary: Agl519 The expression of the CG59540-01 gene appears to be highest in a sample derived from a colon cancer cell line (HCC-2998) (CT=28.2). In addition, there is substantial expression associated with normal kidney and bladder. Thus, the expression of this gene could be used to distinguish these tissues from other tissues in the panel. In addition there was noted expression clustered in ovarian, renal and colon cancer cell lines. Therefore, therapeutic modulation of this gene, through the use of small molecule drugs, antibodies or protein therapeutics might be of use in the tieatment of colon cancer, renal cancer or ovarian cancer.
Among tissues with metabolic function, there is moderate expression in fetal and adult heart, adrenal, and pancreas. This expression suggests that therapeutic modulation of the expression or function of the protein encoded by this gene may be useful in the treatment of diseases that involve these tissues, including obesity and diabetes.
In addition, there appears to be higher levels of expression in adult heart (CT=31) when compared to expression in fetal heart (CT=34.4). Thus, expression of this gene could be used to differentiate between adult and fetal heart tissue. Conversely, expression of this gene is
higher in fetal lung (CT=34.5) than in adult lung (CT=40). Thus, expression of this gene could also be used to differentiate between adult and fetal lung.
Panel 1.3D Summary: Agl519 Significant expression the CG59540-01 gene is limited to a sample derived from colorectal tissue (CT=34.3). Thus, expression of this gene could be used to differentiate between this sample and other samples on this panel, and between colorectal tissue and other normal or malignant tissues.
Panel 2D Summary: Agl519 The expression of the CG59540-01 gene in panel 2 appears to be highest in a samples derived from normal kidney tissue (CT=32). In addition there appears to be substantial difference in expression between normal kidney adjacent to cancer tissue and the cancer tissue itself. Thus, the expression of this gene could be used to distinguish normal kidney tissue from other samples in the panel. In addition, the expression of this gene could be used to distinguish normal kidney from malignant tissue. Moreover, therapeutic modulation of this gene, through the use of small molecule drugs, antibodies or protein therapeutics might be of use in the treatment of kidney cancer.
Panel 4D Summary: Ag3460 Expression of the CG59540-01 gene is low/undetectable (CT values > 35) across the samples in this panel.
BI. CG59280-01 and CG59280-02: OLFACTORY RECEPTOR
Expression of gene CG59280-01 and CG59280-02 was assessed using the primer- probe set Ag3527, described in Table BEA. Results of the RTQ-PCR runs are shown in Table BEB. Please note that CG59280-02 represents a full-length physical clone of the CG59280-01 gene, validating the prediction of the gene sequence.
Table BIA. Probe Name Ag3527
Table BEB. Panel 4D
Tissue Name Rel. Exp.(%) Tissue Name Rel. Exp.(%)
CNS_neurodegeneration_vl.0 Summary: Ag3527 Expression of the CG59280-01 gene is low/undetectable (CT values > 35) across the samples in this panel. (Data not shown.)
General_screening_panel_vl.4 Summary: Ag3527 Expression of the CG59280-01 gene is low/undetectable (CT values > 35) across the samples in this panel. (Data not shown.) This gene encodes a G protein-coupled receptor (GPCR), a type of cell surface receptor involved in signal transduction. It is most similar to members of the odorant receptor subfamily of GPCRs. Based on analogy to other odorant receptor genes, we predict that expression of this gene may be highest in nasal epithelium, a sample not represented on this panel.
Panel 4D Summary: Ag3527 Highest expression of the CG59280-01 gene is seen in the liver cirrhosis sample(CT=31.81). Thus, expression of this gene could be used to differentiate between this sample from the other samples on this panel and as a marker to detect the presence of liver cirrhosis. Furthermore, expression of this gene is not detected in normal liver in Panel 1.4, suggesting that its expression is unique to liver cirrhosis. This gene encodes a putative GPCR; therefore, antibodies or small molecule therapeutics could reduce or inhibit fibrosis that occurs in liver cirrhosis. In addition, antibodies to this putative GPCR could also be used for the diagnosis of liver cirrhosis.
BJ. CG59568-01: GPCR
Expression of gene CG59568-01 was assessed using the primer-probe set Ag3474, described in Table BJA. Results of the RTQ-PCR runs are shown in Table BJB.
Table BJA. Probe Name Ag3474
Table BJB. Panel 4D
CNS_neurodegeneration_vl.0 Summary: Ag3474 Expression of the CG59568-01 gene is low/undetectable (CT values > 35) across the samples in this panel. (Data not shown.)
General_screening_panel_vl.4 Summary: Ag3474 Expression of the CG59568-01 gene is low/undetectable (CT values > 35) across the samples in this panel. (Data not shown.) This gene encodes a G protein-coupled receptor (GPCR), a type of cell surface receptor involved in signal transduction. It is most similar to members of the odorant receptor subfamily of GPCRs. Based on analogy to other odorant receptor genes, we predict that expression of this gene may be highest in nasal epithelium, a sample not represented on this panel.
Panel 4D Summary: Ag3474 Highest expression of the CG59280-01 gene is seen in the liver cirrhosis sample(CT=31.37). Thus, expression of this gene could be used to differentiate between this sample from the other samples on this panel and as a marker to detect the presence of liver cirrhosis. Furthermore, expression of this gene is not detected in normal liver in Panel 1.4, suggesting that its expression is unique to liver cirrhosis. This gene encodes a putative GPCR; therefore, antibodies or small molecule therapeutics could reduce or inhibit fibrosis that occurs in liver cirrhosis. In addition, antibodies to this putative GPCR could also be used for the diagnosis of liver cirrhosis.
References:
1. Mark M.D., Wittemarm S., Herlitze S. (2000) G protein modulation of recombinant P/Q-type calcium channels by regulators of G protein signalling proteins. J. Physiol. 528 Pt 1: 65-77.
BK. CG59224-01 and CG59216-01: GPCR
Expression of gene CG59224-01 and variant CG59216-01 was assessed using the primer-probe sets Ag3400 and Ag3405, described in Tables BKA and BKB. Results of the RTQ-PCR runs are shown in Table BKC.
Table BKA. Probe Name Ag3400
Table BKB. Probe Name Ag3405
SEQ ID
Primers Sequences Length Start Position NO:
Forward 5 ' -cacatctgtgctgtgcttatct-3 22 746 587
Probe TET-5 ' -agtgctgccatgctccaccagttt-3 ' -TAMRA 24 785 588
Reverse 5 ' -acgtggatcataggagacacat-3 ' 22 816 589
Table BKC. General_screening_panel_vl.4
CNS_neurodegeneration_vl.O Summary: Ag3400/Ag3405 Expression of the CG59224-01 gene is low/undetectable (CT values > 35) across the samples in this panel. (Data not shown.)
General_screening_panel_vl.4 Summary: Ag3400/Ag3405 Two experiments with two different probe and primer sets produce results that are in excellent agreement, with significant expression of the CG59224-01 gene exclusively in a lung cancer cell line sample (CTs = 30-
33). Therefore, expression of this gene may be used to distinguish this sample from other samples on this panel and as a marker for lung cancer. Furthermore, therapeutic modulation of the activity of the GPCR encoded by this gene may be beneficial in the tieatment of lung cancer.
Panel 4D Summary: Ag3400/Ag3405 Expression of the CG59224-01 gene is low/undetectable (CT values > 35) across the samples in this panel. (Data not shown.) This gene encodes a G protein-coupled receptor (GPCR), a type of cell surface receptor involved in signal tiansduction. It is most similar to members of the odorant receptor subfamily of GPCRs. Based on analogy to other odorant receptor genes, we predict that expression of this gene may be highest in nasal epithelium, a sample not represented on this panel.
BL. CG59214-01 and CG59214-01: GPCR
Expression of gene CG59214-01 and CG59214-01 was assessed using the primer- probe sets Ag3398 and Ag3404, described in Tables BLA and BLB. Results of the RTQ-PCR runs are shown in Tables BLC and BLD.
Table BLA. Probe Name Ag3398
Table BLB. Probe Name Ag3404
Table BLC. General_screening_panel_vl.4
Table BLD. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3398/Ag3404 Expression of the CG59222-01 gene is low/undetectable (CT values > 35) across the samples in this panel. (Data not shown.)
General_screening_panel_vl.4 Summary: Ag3398/Ag3404 Two experiments with two different probe and primer sets produce results that are in excellent agreement, with significant
expression of the CG59222-01 gene exclusively in a lung cancer cell line sample (CT = 33.8). Therefore, expression of this gene may be used to this sample from other samples on this panel and as a marker for lung cancer. Furthermore, therapeutic modulation of the activity of the GPCR encoded by this gene may be beneficial in the treatment of lung cancer.
Panel 4D Summary: Ag3404 Highest expression of the CG59222-01 gene is seen in the liver cirrhosis sample (CT=32.65). Thus, expression of this gene could be used to differentiate between this sample from the other samples on this panel and as a marker to detect the presence of liver cirrhosis. Furthermore, expression of this gene is not detected in normal liver in Panel 1.4, suggesting that its expression is unique to liver cirrhosis. This gene encodes a putative GPCR; therefore, antibodies or small molecule therapeutics could reduce or inhibit fibrosis that occurs in liver cirrhdsis. In addition, antibodies to this putative GPCR could also be used for the diagnosis of liver cirrhosis. Ag3398 Expression of CG59222-01 gene is low/undetectable (CTs > 35) across all of the samples on this panel (Data not shown).
BM. CG59220-01: GPCR
Expression of gene CG59220-01 was assessed using the primer-probe set Ag3402, described in Table BMA. Results of the RTQ-PCR runs are shown in Tables BMB, BMC and BMD.
Table BMA. Probe Name Ag3402
Table BMB. CNSneurodegeneration vl.O
Table BMC. General_screening_panel_vl.4
CNS_neurodegeneration_vl.O Summary: Ag3402 The CG59220-01 gene is expressed at low levels throughout the CNS, including in amygdala, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. The GPCR family of receptors contains a large number of neurotransmitter receptors, including the dopamine, serotonin, a and b-adrenergic, acetylcholine muscarinic, histamine, peptide, and metabotropic glutamate receptors. GPCRs are excellent drug targets in various neurologic and psychiatric diseases. All antipsychotics have been shown to act at the dopamine D2 receptor; similarly novel antipsychotics also act at the serotonergic receptor, and often the muscarinic and adrenergic receptors as well. While the majority of antidepressants can be classified as selective serotonin reuptake inhibitors, blockade of the 5-HT1A and a2 adrenergic receptors increases the effects of these drugs. The GPCRs are also of use as drug targets in the treatment of stroke. Blockade of the glutamate receptors may decrease the neuronal death resulting from excitotoxicity; further more the purinergic receptors have also been implicated as drug targets in the tieatment of cerebral ischemia. The b-adrenergic receptors have been implicated in the tieatment of ADHD with Ritalin, while the a-adrenergic receptors have been implicated in memory. Therefore this gene may be of use as a small molecule target for the treatment of any of the described diseases.
General_screening_panel_vl.4 Summary: Ag3402 The CG59220-01 gene represents a novel G-protein coupled receptor (GPCR) with highest expression in spinal cord sample (CT=31.12) and moderate expression in other samples from brain. Please see Panel CNS_neurodegeneration_vl.O for discussion of utility of this gene in the cential nervous system.
Low levels of expression of the CG59220-01 gene are also observed in areas outside of the cential nervous system such as the, adipose tissue, fetal and adult heart, skeletal muscle, adrenal gland, pituitary gland, and thyroid suggesting the possibility of a wider role in
intercellular signaling. Therapeutic modulation of the expression or function of this gene may therefore be useful in the treatment of metabolic disorders, including obesity and diabetes.
Panel 4D Summary: Ag3402 The CG59220-01 gene represents a novel G-protein coupled receptor (GPCR) with highest expression in colon (CT=33.12). Thus expression of this gene can be used to distinguish these samples from other samples used in this panel. In addition, expression of this gene is low/undetectable (CT values > 35) in samples derived from EBD colitis and EBS Crohn's. Therefore, expression of this gene can be used to distinguish normal colon sample from the IBD colitis and IBD Crohn's sample used in this panel.
BN. CG59218-01: GPCR
Expression of gene CG59218-01 was assessed using the primer-probe set Ag3401, described in Table BNA. Results of the RTQ-PCR runs are shown in Tables BNB.
Table BNA. Probe Name Ag3401
Table BNB. Panel 4D
CNS neurodegeneration vl.O Summary: Ag3401 Expression of the CG59218-01 gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3401 Expression of the CG59218-01 gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown). This gene product is most similar to members of the odorant receptor subfamily of GPCRs. Based on analogy to other odorant receptor genes, we predict that expression of this gene may be highest in nasal epithelium, a sample not represented on this panel.
Panel 4D Summary: Ag3401 Highest expression of the CG59218-01 gene is seen in the liver cirrhosis sample(CT=33.03). Thus, expression of this gene could be used to differentiate between this sample from the other samples on this panel and as a marker to detect the presence of liver cirrhosis. Furthermore, expression of this gene is not detected in normal liver in Panel 1.4, suggesting that its expression is unique to liver cirrhosis. This gene encodes a putative GPCR; therefore, antibodies or small molecule therapeutics could reduce or inhibit fibrosis that occurs in liver cirrhosis. In addition, antibodies to this putative GPCR could also be used for the diagnosis of liver cirrhosis.
BO. CG59211-01: GPCR
Expression of gene CG59211-01 was assessed using the primer-probe set Ag3397, described in Table BOA. Results of the RTQ-PCR runs are shown in Table BOB.
Table BOA. Probe Name Ag3397
Table BOB. General_screening_panel_vl.4
CNS_neurodegeneration_vl.O Summary: Ag3397 Expression of the CG59211-01 gene is low/undetectable (CT values > 35) across the samples in this panel. (Data not shown.) This gene encodes a G protein-coupled receptor (GPCR), a type of cell surface receptor involved in signal tiansduction. It is most similar to members of the odorant receptor subfamily of GPCRs. Based on analogy to other odorant receptor genes, we predict that expression of this gene may be highest in nasal epithelium, a sample not represented on this panel.
General_screening_panel_vl.4 Summary: Ag3397 Significant expression of the CG59211- 01 gene is seen exclusively in one of the lung cancer sample (CT = 32.29). Therefore, expression of this gene may be used to distinguish this sample from other samples on this panel and as a marker for lung cancer. There is an increasing awareness that some GPCRs can regulate proliferative signaling pathways and that chronic stimulation or mutational activation of receptors can lead to oncogenic transformation. Activating mutations in GPCRs are associated with several types of human tumors and some receptors exhibit potent oncogenic activity due to agonist overexpression (Whitehead et al., 2001). Therefore, therapeutic modulation of the activity of the GPCR encoded by this gene may be beneficial in the tieatment of lung cancer.
References:
1. Whitehead EP, Zohn IE, Der CJ. (2001) Rho GTPase-dependent transformation by G protein-coupled receptors. Oncogene 2001 Mar 26;20(13): 1547-55
Panel 4D Summary: Ag3397 Expression of the CG59211-01 gene is low/undetectable (CT values > 35) across the samples in this panel. (Data not shown.)
BP. CG59276-01: Dihydroorotate dehydrogenase
Expression of gene CG59276-01 was assessed using the primer-probe set Ag3524, described in Table BPA. Results of the RTQ-PCR runs are shown in Tables BPB, BPC, BPD, BPE and BPF.
Table BPA. Probe Name Ag3524
Table BPB. CNS_neurodegeneration_vl.O
Table BPC. General_screening_panel_vl.4
Table BPD. Panel 2D
Table BPF. Panel 5 Islet
CNS_neurodegeneration_vl.0 Summary: Ag3524 No differential expression of the CG59276-01 gene is detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. However, as observed in panel 1.4 this gene is expressed at low levels throughout the CNS, including in amygdala, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. Therefore, this gene may play a role in central nervous system disorders such as Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
General_screening_panel_vl.4 Summary: Ag3524 Expression of the CG59276-01 gene is highest in a sample derived from a brain and lung cancer cell lines (CTs = 29). Thus, the expression of this gene could be used to distinguish these samples from the other samples in the panel. The CG59276-01 gene encodes a dihydroorotate dehydrogenase (DHODH)
homolog. DHODH is an enzyme involved in the pathway for pyrimidine production. Drugs known to inhibit DHODH activity, such as brequinar sodium (Dup-785), have been shown to have anti-tumor activities (ref. 1). Therefore, therapeutic modulation of the activity of this gene encoded by this gene may be beneficial in the tieatment of CNS and lung cancer. In addition, low to moderate expression of this gene is seen in all of the samples on this panel. Therefore, this gene may be playing an important role in cellular function.
This gene is expressed at low to moderate levels in a number of tissues with metabolic or endocrine function, including adipose, adrenal gland, gastrointestinal tract, pancreas, and skeletal muscle. Therefore, therapeutic modulation of the activity of this gene may prove useful in the treatment of endocrine/metabolically related diseases, such as obesity and diabetes.
Recently, it has been demonstrated that down regulation of DHODH mRNA using RNA interference (RNAi) may inhibit growth of Plasmodium falciparum (ref 2).
References:
1. Braakhuis BJ, van Dongen GA, Peters GJ, van Walsum M, Snow GB (1990) Antitumor activity of brequinar sodium (Dup-785) against human head and neck squamous cell carcinoma xenografts. Cancer Lett 49(2): 133-7.
2. McRobert L, McConkey GA.(2002) RNA interference (RNAi) inhibits growth of Plasmodium falciparum. Mol Biochem Parasitol 119(2):273-8
Panel 2D Summary: Ag3524 The expression of this gene appears to be highest in a sample derived from a normal liver tissue (CT=30.3). In addition, there appears to be substantial expression in other samples derived from liver cancers and breast cancers. Thus, the expression of this gene could be used to distinguish normal liver tissue from other samples in the panel. Moreover, therapeutic modulation of this gene, through the use of small molecule drugs, protein therapeutics or antibodies could be of benefit in the tieatment of liver or breast cancer.
Panel 4D Summary: Ag3524 Highest expression of the CG59276-01 gene is detected in resting primary Thl cells (CT=30.03). In addition, the expression of this gene is significantly reduced in activated primary Thl cells, suggesting a regulatory role for this gene in T-cell
activation. The CG59276-01 encodes a dihydroorotate dehydrogenase, an enzyme involved in the pathway for pyrimidine production. Recently, an inhibitor of this enzyme, leflunomide has been shown to be an effective tieatment for rheumatoid arthritis (ref 1). Therefore, therapeutics designed with the protein encoded for by this transcript could be important in regulating T cell function and treating T cell mediated diseases such as asthma, rheumatoid arthritis, psoriasis, IBD, and systemic lupus erythematosus.
Overall, this gene is expressed at low to moderate levels in a wide range of cell types of significance in the immune response in health and disease. These cells include members of the T-cell, B-cell, endothelial cell, macrophage/monocyte, and peripheral blood mononuclear cell family, as well as epithelial and fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattem of expression suggests that this gene product may be involved in homeostatic processes for these and other cell types and tissues. This pattern is in agreement with the expression profile in General_screening_panel_vl.4 and also suggests a role for the gene product in cell survival and proliferation.
Therefore, modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
Reference:
1. Schattenkirchner M. (2000) The use of leflunomide in the tieatment of rheumatoid arthritis: an experimental and clinical review. Immunopharmacology 47(2**-3):291-8
Panel 5 Islet Summary: Ag3524 This gene has a low level of expression in adipose tissue (CTs=33-35). Thus, this gene product maybe a small molecule drug for the treatment of obesity and obesity-related diseases, including Type 2 diabetes.
BQ. CG59268-01: KIAA2372
Expression of gene CG59268-01 was assessed using the primer-probe set Ag3523, described in Table BQA. Results of the RTQ-PCR runs are shown in Tables BQB and BQC.
Table BOA. Probe Name Ag3523
Table BOB. General_screening_panel_vl.4
Table BOC. Panel 4D
CNS neurodegeneration vl.O Summary: Ag3523 Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3523 Expression of the CG59268-01 gene is highest in sample derived from liver cancer cell line (CT=32.55). Therefore, expression of this gene may be used to distinguish liver cancers from the other samples on this panel. In addition, low levels of expression of this gene are also observed in one of the ovarian cancer, 2 of the breast cancer, 2 of the renal cancer, bladder, gastric cancer, 3 of the colon cancer, and 4 of the CNS cancer samples. Therefore, therapeutic modulation of the activity of this gene product may be beneficial in the treatment of these cancers.
Among the tissues with metabolic or endocrine function, this gene is expressed at low levels in adipose tissue sample. Adipose tissue has several crucial roles including (i) mobilization from stores of fatty acids as an energy source, (ii) catabolism of lipoproteins such as very-low-density lipoprotein and (iii) synthesis and release of hormonal signals such as leptin and interleukin-6 (Coppack et al., 2001). Therefore, therapeutic modulation of the activity of this gene may prove useful in the treatment of endocrine/metabolically related diseases, such as obesity, hyperlipidemia, and insulin resistance.
References:
1. Coppack SW, Patel JN, Lawrence VJ. (2001) Nutritional regulation of lipid metabolism in human adipose tissue. Exp Clin Endocrinol Diabetes ;109(Suppl 2):S202-S214
Panel 4D Summary: Ag3523 Expression of the CG59268-01 gene is highest in sample derived from colon (CT=31.56). Therefore, expression of this gene may be used to distinguish colon sample from the other samples on this panel. In addition, significant expression of this gene is also observed in IBD Crohn's sample (CT=32.16). Thus, expression of this gene in colon and Crohn's sample can be used to distinguish these two samples from EBD Colitis 2 sample. In addition, therapeutic modulation of the activity of this gene product may be beneficial in the treatment of EBD Crohn's disease.
BR. CG59549-01: H326 like
Expression of gene CG59549-01 was assessed using the primer-probe set Ag3464, described in Table BRA. Results of the RTQ-PCR runs are shown in Tables BRB and BRC.
Table BRA. Probe Name Ag3464
Table BRB. General_screening_panel_vl.4
Table BRC. Panel 4D
CNS neurodegeneration vl.O Summary: Ag3464 Expression of the CG59549-01 gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3464 Expression of the CG59549-01 gene is highest in a CNS cancer (glio) SF-295 sample (CT = 31.15). Thus, the expression of this gene could be used to distinguish this sample from the other samples in the panel. In addition, low to moderate expression of this gene is detected in a melanoma and a CNS cancer sample. Therefore, therapeutic modulation of this gene or its protein product may be beneficial in the treatment of melanoma and CNS cancer.
Panel 4D Summary: Ag3464 Low but significant expression of the CG59549-01 gene is detected exclusively in liver cirrhosis sample (CT=33.4). Therefore, expression of this gene may be used to distinguish liver cirrhosis from the other samples on this panel. Furthermore, expression of this gene is not detected in normal liver in Panel 1.3D, suggesting that its expression is unique to liver cirrhosis. Therefore, antibodies or small molecule therapeutics could reduce or inhibit fibrosis that occurs in liver cirrhosis. In addition, antibodies to this gene product could also be used for the diagnosis of liver cirrhosis.
BS. CG59641-01: ACETYL-COA CARBOXYLASE 2
Expression of gene CG59641-01 was assessed using the primer-probe set Ag3502, described in Table BSA. Results of the RTQ-PCR runs are shown in Table BSB.
Table BSA. Probe Name Ag3502
Table BSB. General_screening_panel_vl.4
General_screening_panel_vl.4 Summary: Ag3502 The CG59641-01 encodes an acetyl- CoA carboxylase 2 (ACC2) protein. Expression of this gene is highest in adipose tissue (CT=25.5). High levels of expression of this gene are also detected in other tissues with metabolic or endocrine function such as pancreas, adrenal gland, gastiointestinal tiact, heart, skeletal muscle, and thyroid. Acetyl-coenzyme A (acetyl-CoA) carboxylase (ACC) catalyzes the synthesis of malonyl-CoA, a metabolite that plays a pivotal role in the synthesis and
oxidation of fatty. Hence, ACC links fatty acid and carbohydrate metabolism through the shared intermediate acetyl-CoA, the product of pyruvate dehydrogenase. It has been shown recently that mutations in ACC2 gene lead to loss of body fat in a normal caloric intake in mouse (Abu-Elheiga et al., 2001). Therefore, therapeutic modulation of the activity of this gene may prove useful in the treatment of endocrine/metabolically related diseases, such as obesity and diabetes.
Low to moderate expression of this gene is also detected in most of the samples used in this panel suggesting the possibility of a wider role in intercellular signaling for this molecule.
Among tissues that originate in the central nervous system, this gene is expressed in all regions represented on this panel. Therefore, therapeutic modulation of the expression or function of this gene may be useful in the tieatment of neurologic disorders, such as Alzheimer's disease, Parkinson's disease, schizophrenia, multiple sclerosis, stroke and epilepsy.
In addition, significantly higher levels of expression are seen in a breast cancer cell line. Thus, expression of this gene could be used to differentiate between this sample and other samples on this panel and as a marker to detect the presence of breast cancer. Furthermore, therapeutic modulation of the expression or function of this gene may be effective in the treatment of breast cancer.
Reference:
1. Abu-Elheiga L, Matzuk MM, Abo-Hashema KA, Wakil SJ. (2001) Continuous fatty acid oxidation and reduced fat storage in mice lacking acetyl-CoA carboxylase 2. Science 2001 Mar 30;291(5513):2613-6
BT. CG59630-01: Midnolin
Expression of gene CG59630-01 was assessed using the primer-probe set Ag3425, described in Table BTA. Results of the RTQ-PCR runs are shown in Tables BTB, BTC and BTD.
Table BTA. Probe Name Ag3425
Primers) Sequences [LengthjStart Position) SEQ ID
Table BTB. CNSneurodegeneration vl.O
Table BTC. General_screening_panel_vl.4
Table BTD. Panel 4. ID
CNS_neurodegeneration_vl.O Summary: Ag3425 This panel confirms the expression of this gene at low levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in treatment of central nervous system disorders.
General_screening_panel_vl.4 Summary: Ag3425 The CG59630-01 gene is a homologue of mouse midnoline (midbrain nucleolar protein). Its expression is moderate to high across all of the samples on this panel, with highest expression in a breast cancer cell line (CT=25.3). The widespread expression suggests that this gene may play an important role in cellular function. In mouse, the expression of this gene is developmentally regulated: it is strongly expressed at the mesencephalon (midbrain) of the embryo and is involved in regulation of genes related to neurogenesis in the nucleolus (Tsukahara et al., 2000). Based on the gene's expression in all CNS regions examined, this gene may therefore play a role in central nervous
system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Among tissues with metabolic function, this gene is expressed at moderate to low levels in pituitary, adipose, adrenal gland, pancreas, thyroid, and adult and fetal skeletal muscle, heart, and liver. This widespread expression among these tissues suggests that this gene product may play a role in normal neuroendocrine and metabolic and that disregulated expression of this gene may contribute to neuroendocrine disorders or metabolic diseases, such as obesity and diabetes.
Reference:
1. Tsukahara M, Suemori H, Noguchi S, Ji ZS, Tsunoo H. (2000) Novel nucleolar protein, midnolin, is expressed in the mesencephalon during mouse development. Gene 2000 Aug 22;254(l-2):45-55
Panel 4.1D Summary: Ag3425 The CG59630-01 gene is a homologue of mouse midnoline (midbrain nucleolar protein). Its expression is moderate to high across all of the samples on this panel, with highest expression in resting neutrophils (CT=29.1). In addition, this gene is expressed at high to moderate levels in a wide range of cell types of significance in the immune response in health and disease. These cells include members of the T-cell, B-cell, endothelial cell, macrophage/monocyte, and peripheral blood mononuclear cell family, as well as epithelial and fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product may be involved in homeostatic processes for these and other cell types and tissues. This pattern is in agreement with the expression profile in General_screening_panel_vl.4 and also suggests a role for the gene product in cell survival and proliferation. Therefore, modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
BU. CG59561-01: CYTOSOLIC ACYL COENZYME A THIOESTER HYDROLASE
Expression of gene CG59561-01 was assessed using the primer-probe set Ag3424, described in Table BUA. Results of the RTQ-PCR runs are shown in Tables BUB, BUC and BUD.
Table BUA. Probe Name Ag3424
Table BUB. CNS_neurodegeneration_vl.0
Table BUC. Panel 4D
CNS_neurodegeneration_vl.0 Summary: Ag3424 This panel confirms the expression of the CG59561-01 gene at low levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. This expression profile suggests that this gene may play a role in central nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
General_screening_panel_vl.4 Summary: Ag3424 Results from one experiment with the CG59561-01 gene are not included. The amp plot indicates that there were experimental difficulties with this run. (Data not shown.)
Panel 4D Summary: Ag3424 The CG59561-01 gene encodes a protein homologous to cytosolic acyl coenzyme A thioester hydrolase (Brain acyl-CoA hydrolase, BACH). Among the tissue samples used in this panel, highest expression of this gene is detected in thymus (CT=29.6). In addition, expression of this gene is stimulated in activated primary and secondary - Thl, Th2 and Tri cells. Therefore, this gene product may play an important role in T cell development. Thus, therapeutics designed with the protein encoded for by this tianscript could be important in regulating T cell function and tieating T cell mediated diseases such as emphysema, asthma, arthritis, psoriasis, EBD, and systemic lupus erythematosus.
Interestingly, expression of this gene is also seen in activated PBMCs (CTs=30) as compared to resting PBMCs (CT=36) suggesting a role for this gene product in B-cell and T- cell proliferation. Therefore, small molecules that antagonize the function of this gene product may be useful as therapeutic drugs to reduce or eliminate the symptoms in patients with
autoimmune and inflammatory diseases in which B cells play a part in the initiation or progression of the disease process, such as systemic lupus erythematosus, Crohn's disease, ulcerative colitis, multiple sclerosis, chronic obstructive pulmonary disease, asthma, emphysema, rheumatoid arthritis, or psoriasis.
Panel 5 Islet Summary: Ag3424 The CG59561-01 gene is expressed at low levels in adipose and placenta, with highest expression in the kidney (CT=30.8). As an enzyme involved in lipid homeostasis, therapeutic modulation of this gene product may be a treatment for obesity and obesity-related diseases, including Type 2 diabetes.
BV. CG59452-01: CELL PROLIFERATION RELATED PROTEIN CAP -
Expression of gene CG59452-01 was assessed using the primer-probe set Ag3443, described in Table BVA. Results of the RTQ-PCR runs are shown in Tables BVB and BVC.
Table BVA. Probe Name Ag3443
Table BVB. CNS_neurodegeneration_vl.0
Table BVC. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3443 This panel confirms the expression of the CG59452-01 gene at significant levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Expression of this gene in the brain suggests that it may play a role in central nervous system disorders other than Alzheimer's disease, such as Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
General_screening_panel_vl.4 Summary: Ag3443 The amp plot indicates that there were experimental difficulties with this run. (Data not shown).
Panel 4D Summary: Ag3443 Highest expression of the CG59452-01 gene is detected in TNFalpha + IL-lbeta treated keratinocytes and PMA/ionomycin treated KU-812 basophil cells (CTs=24.5). Thus, antibody or small molecule therapies designed with the protein encoded for by this gene could block or inhibit inflammation or tissue damage due to basophil activation in response to asthma, allergies, hypersensitivity reactions, psoriasis, and viral infections.
BW. CG59572-01 and CG59572-02: Pseudouridine Synthase 3
Expression of gene CG59572-01 and CG59572-02 was assessed using the primer- probe set Ag3476, described in Table BWA. Results of the RTQ-PCR runs are shown in
Tables BWB, BWC and BWD. Please note that CG59572-02 represents a full-length physical clone of the CG59572-01 gene, validating the prediction of the gene sequence.
Table BWA. Probe Name Ag3476
Table BWB. CNS_neurodegeneration_vl.O
Table BWC. General_screening_panel_vl.4
Table BWD. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3476 This panel confirms the expression of the CG59572-01 gene at low levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in treatment of central nervous system disorders.
General_screening_panel_vl.4 Summary: Ag3476 Highest expression of the CG59572-01 gene is detected in a breast cancer cell line sample (CT=27.4). Furthermore, moderate to high expression of this gene is detected in CNS cancer, colon cancer, gastric cancer, pancreatic
cancer, lung cancer, ovarian cancer, and prostate cancer. Therefore, therapeutic modulation of the activity of this gene or its protein product, through the use of small molecule drugs, protein therapeutics or antibodies, might be beneficial in the tieatment of these cancers.
This gene is expressed at low to moderate levels throughout the CNS, including in amygdala, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. Therefore, this gene may play a role in central nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Among tissues with metabolic function, this gene is expressed at moderate to low levels in pituitary, adipose, adrenal gland, pancreas, thyroid, and adult and fetal skeletal muscle, heart, and liver. This widespread expression among these tissues suggests that this gene product may play a role in normal neuroendocrine and metabolic and that disregulated expression of this gene may contribute to neuroendocrine disorders or metabolic diseases, such as obesity and diabetes.
In addition, this gene is expressed at much higher levels in fetal lung and liver tissue (CTs=30) when compared to expression in the adult counterpart (CTs=33). Thus, expression of this gene may be used to differentiate between the fetal and adult source of these tissues.
Panel 4D Summary: Ag3476 Highest expression of the CG59572-01 gene is detected in TNFalpha + IL-lbeta treated keratinocytes (CT=27.2). Expression of this gene appears to be stimulated in activated secondary Thl, Th2 and Tri cells, PWM tieated PBMCs, PWM tieated B-lymphocytes, IL-2/IL-2+EL-12/IL-2+EFN gamma/IL-2+IL-18 treated LAK cells, and TNFalpha + IL-lbeta treated small airway epithelium (CTs=28-30). Thus, this gene may be important in the activation of T and B cells or the function of activated T and B cells. Therefore, small molecules that antagonize the function of this gene product may reduce or eliminate the symptoms in patients with autoimmune and inflammatory diseases in which B and T cells play a part in the initiation or progression of the disease process, such as systemic lupus erythematosus, Crohn's disease, ulcerative colitis, multiple sclerosis, chronic obstructive pulmonary disease, asthma, emphysema, rheumatoid arthritis, or psoriasis.
BX. CG59522-01: Myosin I
Expression of gene CG59522-01 was assessed using the primer-probe set Ag3456, described in Table BXA. Results of the RTQ-PCR runs are shown in Table BXB.
Table BXA. Probe Name Ag3456
Table BXB. Panel 4D
CNS neurodegeneration vl.O Summary: Ag3456 Expression of CG59522-01 gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3456 Expression of CG59522-01 gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
Panel 4D Summary: Ag3456 Highest expression of the CG59522-01 gene is detected in sample derived from resting primary Thl cells (CT=29.8). Thus, expression of this gene
be used to distinguish this sample from other samples in this panel. This gene is also expressed at low but significant levels in T cells prepared under a number of conditions, LAK cells, macrophages and dendritic cells also express the transcript. The only non-hematopoietic cell type that expresses the transcript detected by this primer and probe at significant levels is dermal fibroblasts. Colon and kidney also express low levels of the transcript. Thus, this transcript or the protein it encodes could be used to detect hernatopoietically-derived cells. Furthermore, therapeutics designed with the protein encoded by this tianscript could be important in the regulation the function of antigen presenting cells (macrophages and dendritic cells)or T cells and be important in the treatment of asthma, emphysema, psoriasis, arthritis, and EBD. Therefore, therapeutics designed with the protein encoded for by this transcript could be important in regulating T cell function and treating T and B cell mediated diseases such as asthma, arthritis, psoriasis, BD, and systemic lupus erythematosus.
BY. CG59520-01: FARNESYL PYROPHOSPHATE SYNTHETASE
Expression of gene CG59520-01 was assessed using the primer-probe set Ag5923, described in Table BYA. Results of the RTQ-PCR runs are shown in Tables BYB and BYC.
Table BYA. Probe Name Ag5923
Table BYB. General screening_panel vl.5
Tissue Name Rel. Exp.(%) Ag5923, Rel. Exp.(%) Ag5923,
Tissue Name Run 247608956 Run 247608956
Table BYC. Panel 4. ID
CNS_neurodegeneration_vl.0 Summary: Ag5923 Expression of the CG59520-01 gene is low/undetectable (CTs > 34.5) across all of the samples on this panel (data not shown).
General_screening_panel_vl.5 Summary: Ag5923 Highest expression of the CG59520-01 gene is detected in sample derived from a pancreatic cancer cell line (CT=31.5). Thus, expression of this gene can be used in distinguishing this sample from other samples from the panel and as a marker for pancreatic cancer. In addition low levels of expression of this gene are associated with samples derived from CNS, colon, gastric, renal, lung, breast, ovarian and melanoma cnacer cell lines. This gene encodes a farnesyl pyrophosphate synthetase, which is involved in cholesterol biosynthesis. It has been suggested that in several types of cancer, activation of p21 would be aided by continuous farnesylation due to stimulation of the cholesterol biosynthetic pathway in tumors (Rao, 1995). Therefore, therapeutic modulation of the activity of protein encoded by this gene may be beneficial in the treatment of these cancers.
In addition, low but significant levels of expression in the pancreas suggest that this gene product may be useful in the tieatment of type II diabetes.
References:
1. Rao KN. (1995) The significance of the cholesterol biosynthetic pathway in cell growth and carcinogenesis (review). Anticancer Res 1995 Mar-Apr;15(2):309-14
Panel 4.1D Summary: Ag5923 High expression of the CG59520-01 gene is detected in sample derived from untreated and IL4 treated NCI-H292 cells (CTs=33). Thus, expression of this gene could be used to distinguish these samples from other samples from the panel. Also, therapeutic modulation of the activity of this gene product may be beneficial in the treatment asthma and emphysema.
Panel 5 Islet Summary: Ag5923 Expression of the CG59520-01 gene is low/undetectable (CTs > 34.5) across all of the samples on this panel (data not shown).
BZ. CG59704-01: Serine/Threonine Kinase
Expression of gene CG59704-01 was assessed using the primer-probe set Ag3509, described in Table BZA. Results of the RTQ-PCR runs are shown in Tables BZB, BZC and BZD.
Table BZA. Probe Name Ag3509
Table BZB. CNS_neurodegeneration_vl.O
Table BZD. Panel 4D
JMacrophages LPS 7.2 JThymus 1 27-5
JHUVEC none 24.1 JKidney j 51.4
JHUVEC starved 33.9 | |
CNS_neurodegeneration_vl.O Summary: Ag3509 This panel confirms the expression of this gene at low levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment.
General_screening_panel_vl.4 Summary: Ag3509 Highest expression of the CG59704-01 gene is detected in a sample derived from a lung cancer cell line (CT=31.69). Thus, expression of this gene can be used in distinguishing this sample from other samples in this panel. Furthermore, moderate expression of this gene is associated with cell lines derived from pancreatic, brain, colon, gastric, renal, lung, breast and ovarian cancers. Therefore, therapeutic modulation of the activity of this gene or its protein product, through the use of small molecule drugs, or antibodies, might be beneficial in the treatment of these cancers.
Panel 4D Summary: Ag3509 Expression of the CG59704-01 gene is stimulated in T cells, LAK cells and B cells, with highest expression in primary activated Tri cells (CT=32). Therefore, therapeutics designed with the protein encoded for by this transcript could be important in regulating T and B cell function and treating T cell/B cell mediated diseases such as asthma, arthritis, psoriasis, EBD, allergies, hypersensitivity reactions, microbial and viral infections systemic lupus erythematosus, multiple sclerosis, chronic obstructive pulmonary disease and systemic lupus erythematosus.
Furthermore, expression of this gene is decreased in colon samples from patients with EBD colitis and Crohn's disease relative to normal colon. Therefore, therapeutic modulation of the activity of this gene product may be useful in the tieatment of inflammatory bowel disease.
Panel 5 Islet Summary: Ag3509 Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
CA. CG59628-01 : short-chain dehydrogenase like homo sapiens
Expression of gene CG59628-01 was assessed using the primer-probe set Ag3500, described in Table CAA. Results of the RTQ-PCR runs are shown in Tables CAB and CAC.
Table CAA. Probe Name Ag3500
Table CAB. General_screening_panel_vl.4
CNS_neurodegeneration_vl.O Summary: Ag3500 Results from one experiment with the CG59628-01 gene are not included. The amp plot indicates that there were experimental difficulties with this run.
General_screening_panel_vl.4 Summary: Ag3500 Highest expression of the CG59628-01 gene is detected in a sample derived from a CNS cancer cell line (CT=31.1). Therefore, expression of this gene may be used to distinguish this sample from the other samples on this panel. In addition, significant expression of this gene is associated with samples derived from colon, ovarian, breast, renal, lung, melanoma, and brain cancer cell lines. Therefore, therapeutic modulation of the activity of this gene or its protein product might be beneficial in the tieatment of these cancers.
Among tissues with metabolic function, this gene is expressed at low but significant levels in pituitary, adipose, adrenal gland, pancreas, thyroid, and adult and fetal skeletal muscle, heart, and liver. This widespread expression among these tissues suggests that this gene product may play a role in normal neuroendocrine and metabolic and that disregulated expression of this gene may contribute to neuroendocrine disorders or metabolic diseases, such as obesity and diabetes.
This molecule is also expressed at low levels in the CNS, including the hippocampus, thalamus, substantia nigra and cerebral cortex. Therefore, therapeutic modulation of the expression or function of this gene may be useful in the treatment of neurologic disorders, such as Alzheimer's disease, Parkinson's disease, schizophrenia, multiple sclerosis, stroke and epilepsy.
Panel 4D Summary: Ag3500 Highest expression of the CG59628-01 gene is detected in colon (CT=30.3). Therefore, expression of this gene may be used to distinguish colon from the other tissues on this panel. Furthermore, expression of this gene is decreased in colon samples from patients with EBD colitis and Crohn's disease relative to normal colon. Therefore, therapeutic modulation of the activity of the GPCR encoded by this gene may be useful in the treatment of inflammatory bowel disease.
CB. CG59671-02: acetyl-coenzyme A synthetase
Expression of gene CG59671-02 was assessed using the primer-probe sets Ag3506 and Ag3581, described in Tables CBA and CBB. Results of the RTQ-PCR runs are shown in Tables CBC, CBD, CBE and CBF.
Table CBA. Probe Name Ag3506
Table CBB. Probe Name Ag3581
Table CBC. CNS_neurodegeneration_vl.0
Table CBD. General_screening_panel_vl.4
Table CBE. Panel 4. ID
CNS_neurodegeneration_vl.O Summary: Ag3506/Ag3581 This panel confirms the expression of the CG59671-02 gene at significant levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment.
This gene is expressed at moderate levels throughout the CNS, including in amygdala, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord as observed in panel 1.4. Therefore, this gene may play a role in other central nervous system disorders such as, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression
General_screening_panel_vl.4 Summary: Ag3506/Ag3581 Two experiments produce results that are in very good agreement. Highest expression of the CG59671-02 gene is observed in samples derived from melanoma cell lines (CTs=23-35). Thus, expression of this gene can be used in distinguishing these samples from other samples in the panel. In addition,
significant levels of expression of this gene are also associated with colon cancer, ovarian cancer, breast cancer, and lung cancer cell lines. Therefore, therapeutic modulation of the activity of this gene or its protein product might be beneficial in the treatment of these cancers.
This gene is also expressed at low to moderate levels in a number of tissues with metabolic or endocrine function, including adipose, adrenal gland, gastrointestinal tract, pancreas, skeletal muscle and thyroid. Therefore, therapeutic modulation of the activity of this gene may prove useful in the treatment of endocrine/metabolically related diseases, such as obesity and diabetes.
This gene is also expressed at high to moderate levels in all regions of the CNS examined. Please see Panel CNS_neurodegeneration_vl.0 for discussion of utility of this gene in the central nervous system.
Panel 4.1D Summary: Ag3581 Highest expression of the CG59671-02 gene is observed in the resting KU-812 sample (CT=29.18). In addition, high expression of this gene is detected in colon, lung, thymus and kidney. Therefore, therapies designed with the protein encoded for by this gene could be important in the treatment of inflammatory or autoimmune diseases that affect the kidney, lung and kidney including, asthma, allergies, lupus and glomerulonephritis. Expression of this gene is decreased in colon samples from patients with EBD colitis and Crohn's disease relative to normal colon. Therefore, therapeutic modulation of the activity of the protein encoded by this gene may also be useful in the tieatment of inflammatory bowel disease.
Expression of this gene appears to be down-regulated in activated primary and secondary Thl, Th2, and Tri cells. Also, TNF alpha stimulates the expression of this gene in resting dermal fibroblasts. Therefore, therapeutics designed with the protein encoded by this transcript could be important in regulating T cell function and treating diseases such as asthma, arthritis, psoriasis, IBD, and systemic lupus erythematosus.
Panel 4D Summary: Ag3506 Highest expression of CG59671-02 is observed colon sample (CT=27.3). Overall, the expression pattern using this probe is in excellent agreement with results in panel 4. ID for Ag3581. Please see that panel for discussion of utility of this gene in inflammation.
CC. CG56870-01: NDR3
Expression of gene CG56870-01 was assessed using the primer-probe set Ag2075, described in Table CCA. Results of the RTQ-PCR runs are shown in Tables CCB, CCC, CCD and CCE. Please note that CG56870-02 represents a full-length physical clone of the CG56870-01 gene, validating the prediction of the gene sequence.
Table CCA. Probe Name Ag2075
Table CCB. Panel 1.3D
Table CCE. Panel 4D
Monocytes LPS 6.5 JColon | 26.8
Macrophages rest 36.1 |Lung 1 21-3
Macrophages LPS 13.3 JThymus 1 41-5
HUVEC none 37.6 (Kidney 1 24.3
HUVEC starved 58.6 j 1
Panel 1.3D Summary: Ag2075 Highest expression of the CG56870-01 gene is detected in the cerebral cortex (CT=24.2). Thus expression of this gene can be used in distinguishing this sample from other samples in the panel. Furthermore, significant expression of this gene is observed throughout the CNS, including in amygdala, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. The CG56870-01 gene encodes an Ndr3 homolog which is a putative member of Ndr family. This family consists of proteins from different gene families: Ndrl/RTP/Drgl/NDRGl, Ndr2, and Ndr3 (PFAM: EPR004142). NDRG1 is a cytoplasmic protein involved in stress responses, hormone responses, cell growth, and differentiation. Mutation of this gene was reported to be causative for hereditary motor and sensory neuropathy-Lorn. Recently, NDRG4, another memember of Ndr family, was shown to be expressed in neurons of the brain and spinal cord. Its expression was markedly decreased in the brain of Alzheimer's disease patient (Zhou et al., 2001). Therefore, this gene may play a role in central nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
This gene also has moderate levels of expression in adipose, adrenal, thyroid, liver, heart, thyroid and skeletal muscle. Thus, this gene product may be important in the pathogenesis, diagnosis and/or treatment of metabolic and endocrine disease, including Types 1 and 2 diabetes and obesity.
In addition, there appears to be substantial expression in other samples derived from breast cancer cell lines, lung cancer cell lines, renal cancer cell lines and colon cancer cell lines. Thus, therapeutic modulation of this gene could be of benefit in the treatment of breast, lung, renal or colon cancer.
References:
1. Zhou RH, Kokame K, Tsukamoto Y, Yutani C, Kato H, Miyata T. (2001) Characterization of the human NDRG gene family: a newly identified member, NDRG4, is specifically expressed in brain and heart. Genomics 73(l):86-97
Ag2075 The expression of this gene appears to be highest in a sample derived from a normal brain tissue. In addition, there appears to be substantial expression in other samples derived from breast cancer cell lines, lung cancer cell lines, renal cancer cell lines and colon cancer cell lines. Thus, the expression of this gene could be used to distinguish normal brain tissue from other samples in the panel. Moreover, therapeutic modulation of this gene could be of benefit in the treatment of breast, lung, renal or colon cancer.
Panel 2.2 Summary: Ag2075 Highest expression of CG56870-01 is detected in breast cancer sample (CT=29.89). Thus expression of this gene can be used in distinguishing this sample from other samples in the panel. In addition, there appears to be substantial expression in other samples derived from breast cancers, kidney cancers and colon cancers. Therefore, therapeutic modulation of this could be of benefit in the treatment of breast, kidney or colon cancer.
Panel 3D Summary: Ag2075 The expression of this gene appears to be highest in a sample derived from a lung cancer cell line (DMS-79)(CT=26.4). In addition, there appears to be substantial expression in other samples derived from pancreatic cancer cell lines, lung cancer cell lines, brain cancer cell lines and cervical cancer cell lines. Thus, the expression of this gene could be used to distinguish DMS-79 cells from other samples in the panel. Moreover, therapeutic modulation of this gene could be of benefit in the treatment of pancreatic, lung, brain or cervical cancer.
Panel 4D Summary: Ag2075 Expression of the CG56870-01 gene is ubiquitous througout this panel, with highest in samples derived from ionomycin treated Ramos (B cell) cells (CT=26.1). Furthermore, expression of this gene is also detected in PWM treated PBMC cells and PWM treated B lymphocytes. Therefore, therapeutic modulation of the expression or function of this gene may reduce or eliminate the symptoms in patients with autoimmune and inflammatory diseases in which B cells play a part in the initiation or progression of the disease process, such as systemic lupus erythematosus, Crohn's disease, ulcerative colitis, multiple sclerosis, chronic obstructive pulmonary disease, asthma, emphysema, rheumatoid arthritis, or psoriasis.
CD. CG56870-04: N-myc downstream-regulated gene 3
Expression of gene CG56870-04 was assessed using the primer-probe sets Ag5279 and Ag2075, described in Tables CDA and CDB. Results of the RTQ-PCR runs are shown in Tables CDC, CDD, CDE, CDF, CDG, CDH and CDI.
Table CD A. Probe Name Ag5279
Table CDB. Probe Name Ag2075
Table CDC. CNS_neurodegeneration_vl.O
Table CDD. General_screening_panel_vl.5
Renal ca. A498 I 11.0 Thyroid (female) j 2.0
Renal ca. ACHN 1 6.0 Pancreatic ca. CAPAN2| 5.6
Renal ca. UO-31 1 8.0 Pancreas Pool | 6.2
Table CDE. Panel 1.3D
Table CDG. Panel 3D
Table CDH. Panel 4. ID
CNS_neurodegeneration_vl.O Summary: Ag5279 This panel confirms the expression of the CG56870-04 gene at low levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.5 for a discussion of the potential utility of this gene in tieatment of central nervous system disorders.
General_screening_panel_vl.5 Summary: Ag5279 Highest expression of the CG56870-01 is detected in cerebral cortex (CT=25.02). Thus, expression of this gene can be used in distinguishing this sample from other samples in the panel. Furthermore, significant expression of this gene is observed throughout the CNS, including in amygdala, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. The CG56870-01 gene encodes a Ndr3 protein homolog. The Ndr family is comprised of members from different gene families:
Ndrl/RTP/Drgl/NDRGl, Ndr2, and Ndr3 (PFAM: EPR004142). NDRG1 is a cytoplasmic protein involved in stress responses, hormone responses, cell growth, and differentiation. Mutation of this gene was reported to be causative for hereditary motor and sensory neuropathy-Lorn. Recently, NDRG4, another memember of Ndr family, was shown to be expressed in neurons of the brain and spinal cord. Its expression was markedly decreased in the brain of Alzheimer's disease patient (Zhou et al., 2001). Therefore, this gene may play a role in cential nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Among metabolic tissues, this gene is moderately expressed in adipose, adrenal, heart, thyroid, liver, pancreas, pituitary, and skeletal muscle. Thus, this gene product may be important for the pathogenesis, diagnosis, and/or treatment of metabolic and endocrine disease, including Types 1 and 2 diabetes and obesity.
In addition, there appears to be substantial expression in other samples derived from brain cancer cell lines, colon cancer cell lines, breast cancer cell lines and ovarian cancer cell lines. Moreover, therapeutic modulation of this gene could be of benefit in the treatment of brain, colon, breast or ovarian cancer.
References:
1. Zhou RH, Kokame K, Tsukamoto Y, Yutani C, Kato H, Miyata T. (2001) Characterization of the human NDRG gene family: a newly identified member, NDRG4, is specifically expressed in brain and heart. Genomics 73(l):86-97
Panel 1.3D Summary: Ag2075 Highest expression of the CG56870-01 gene is detected in the cerebral cortex (CT=24.2). This expression is consistent with expression in Panel 1.5. Please see that panel for discussion of utility of this gene in the central nervous system.
This gene also has moderate levels of expression in adipose, adrenal, thyroid, liver, heart, thyroid and skeletal muscle. Thus, this gene product may be important in the pathogenesis, diagnosis and/or tieatment of metabolic and endocrine disease, including Types 1 and 2 diabetes and obesity.
In addition, there appears to be substantial expression in other samples derived from breast cancer cell lines, lung cancer cell lines, renal cancer cell lines and colon cancer cell
lines. Thus, therapeutic modulation of this gene could be of benefit in the treatment of breast, lung, renal or colon cancer.
Panel 2.2 Summary: Ag2075 Highest expression of CG56870-01 is detected in breast cancer sample (CT=29.89). Thus expression of this gene can be used in distinguishing this sample from other samples in the panel. In addition, there appears to be substantial expression in other samples derived from breast cancers, kidney cancers and colon cancers. Therefore, therapeutic modulation of this gene could be of benefit in the treatment of breast, kidney or colon cancer.
Panel 3D Summary: Ag2075 The expression of this gene appears to be highest in a sample derived from a lung cancer cell line (DMS-79)(CT=26.4). In addition, there appears to be substantial expression in other samples derived from pancreatic cancer cell lines, lung cancer cell lines, brain cancer cell lines and cervical cancer cell lines. Thus, the expression of this gene could be used to distinguish DMS-79 cells from other samples in the panel. Moreover, therapeutic modulation of this gene could be of benefit in the tieatment of pancreatic, lung, brain or cervical cancer.
Panel 4.1D Summary: Ag5279 Expression of the CG56870-01 gene is highest in samples derived from TNF alpha treated dermal fibroblast CCD1070 cells (CT=30.6). Expression of this gene is also prominent in activated secondary and primarey Thl, Th2 and Tri cells when compared expression in the corresponding resting cell lines. Thus, this gene may be involved in T lymphocyte function. Therefore, therapeutic modulation fo the expression or function of this gene may be as anti-inflammatory therapeutics for T cell-mediated autoimmune and inflammatory diseases, such as asthma, athritis, psoriasis, EBD, and lupus.
Panel 4D Summary: Ag2075 Expression of the CG56870-01 gene is ubiquitous througout this panel, with highest in samples derived from ionomycin tieated Ramos (B cell) cells (CT=26.1). Furthermore, expression of this gene is also detected in PWM treated PBMC cells and PWM tieated B lymphocytes. Therefore, therapeutic modulation of the expression or function of this gene may reduce or eliminate the symptoms in patients with autoimmune and inflammatory diseases in which B cells play a part in the initiation or progression of the disease process, such as systemic lupus erythematosus, Crohn's disease, ulcerative colitis, multiple sclerosis, chronic obstructive pulmonary disease, asthma, emphysema, rheumatoid arthritis, or psoriasis.
CE. CG56870-05: N-myc downstream-regulated gene 3
Expression of gene CG56870-05 was assessed using the primer-probe set Ag5265, described in Table CEA. Results of the RTQ-PCR runs are shown in Tables CEB and CEC.
Table CEA. Probe Name Ag5265
Table CEB. CNS_neurodegeneration_vl.0
Table CEC. General_screening_panel_vl.5
CNS_neurodegeneration_vl.0 Summary: Ag5265 This panel confirms the expression of the CG56870-04 gene at low levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.5 for a discussion of the potential utility of this gene in treatment of central nervous system disorders.
General_screening_panel_vl.5 Summary: Ag5265 Highest expression of the CG56870-05 gene is detected in cerebral cortex (CT=28.86). Thus, expression of this gene can be used in distinguishing this sample from other samples in the panel. Furthermore, significant expression of this gene is observed throughout the CNS, including in amygdala, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. The CG56870-05 gene encodes a putative Ndr3 protein. This family consists of proteins from different gene families: Ndrl/RTP/Drgl/NDRGl, Ndr2, and Ndr3 (PFAM: IPR004142). NDRG1 is a cytoplasmic protein involved in stress responses, hormone responses, cell growth, and differentiation. Mutation of this gene was reported to be causative for hereditary motor and sensory neuropathy-Lorn. Recently, NDRG4, another memember of Ndr family, was shown to be expressed in neurons of the brain and spinal cord. Its expression was markedly decreased in the brain of Alzheimer's disease patient (Zhou et al., 2001). Therefore, this gene may play a role in central nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Among metabolic tissues, this gene has low levels of expression in heart, skeletal muscle, adrenal, thyroid, pancreas and pituitary. Therefore, this gene product may be important for the pathogenesis, diagnosis, and/or treatment of metabolic and endocrine disease, including Types 1 and 2 diabetes and obesity.
Overall, this gene is expressed in all the samples on this panel, with slightly higher levels of expression in the cancer cell lines compared to expression in the normal tissues samples.
Panel 4.1D Summary: Ag5265 Expression of this gene is low/undetectable (CTs > 34.5) across all of the samples on this panel (data not shown).
CF. CG59764-01: FERRITIN HEAVY CHAIN like protein
Expression of gene CG59764-01 was assessed using the primer-probe set Ag3578, described in Table CFA. Results of the RTQ-PCR runs are shown in Tables CFB and CFC.
Table CFA. Probe Name Ag3578
Table CFB. CNS_neurodegeneration_vl.O
Table CFC. General_screening_panel_vl.4
CNS_neurodegeneration_vl.0 Summary: Ag3578 This panel confirms the expression of the CG59764-01 gene at low levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in treatment of central nervous system disorders.
General_screening_panel_vl.4 Summary: Ag3578 Highest expression of the CG59764-01 gene is detected in sample derived from skeletal muscle (CT=31.2). Thus expression of this gene can be used to distinguish skeletal muscle sample from other samples used in this panel. This gene is also expressed at low but significant levels in heart and adipose. Thus, this gene product may be useful in the treatment of metabolic disorders that involve these tissues, including obesity.
Significant expression of this gene is also associated with samples derived from breast cancer, pancreatic cancer, colon cancer and lung cancer cell lines. Therefore, therapeutic modulation of the activity of this gene or its protein product might be beneficial in the treatment of these cancers.
In addition, this gene is expressed at low levels throughout the CNS, including in amygdala, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. The CG59764-01 gene encodes a homologue of ferritin heavy chain protein (H-feritin). It has been hypothesized that the up-regulation of the H- ferritin mRNA is part of a mechanism protecting the hippocampus, a seizure-prone area, against a possible overactivation during absence seizures (Lakaye et al., 2000). Therefore, therapeutic modulation of the expression or function of this gene may be useful in the treatment of seizure disorders, such as epilepsy. Furthermore, this gene may play a role in cential nervous system disorders such as Alzheimer's disease, Parkinson's disease, multiple sclerosis, schizophrenia and depression.
References:
1. Lakaye B, de Borman B, Minet A, Arckens L, Vergnes M, Marescaux C, Grisar T. (2000) Increased expression of mRNA encoding ferritin heavy chain in brain structures of a rat model of absence epilepsy. Exp Neurol 162(1): 112-20.
Panel 4.1D Summary: Ag3578 Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
CG. CG59710-01: P14
Expression of gene CG59710-01 was assessed using the primer-probe set Ag3512, described in Table CGA. Results of the RTQ-PCR runs are shown in Tables CGB and CGC.
Table CGA. Probe Name Ag3512
Table CGB. CNS_neurodegeneration_vl.0
CNS_neurodegeneration_vl.0 Summary: Ag3512 This panel confirms the expression of the CG59710-01 gene at low levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. However, as seen in panel 1.4, this gene is expressed at low levels throughout the CNS, including in amygdala, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. Therefore, this gene may play a role in other central nervous system disorders such as Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
General_screening_panel_vl.4 Summary: Ag3512 Highest expression of the CG59710-01 gene is detected in a sample derived from a breast cancer cell line (CT=25.3). Therefore, expression of this gene could be used in distinguishing this sample from other samples in the panel. Overall, expression of this gene appears to be associated with the cancer cell lines suggesting a role for this gene product in cellular growth and proliferation. Specifically, significant expression of this gene is associated with CNS cancer, colon cancer, gastric cancer, renal cancer, lung cancer, breast cancer, ovarian cancer, and melanoma cancer cell lines. Therefore, therapeutic modulation of the activity of this gene or its protein product might be beneficial in the treatment of these cancers.
Panel 4.1D Summary: Ag3512 Results from one experiment with the CG59710-01 gene are not included. The amp plot indicates that there were experimental difficulties with this run.
CH. CG59754-02 and CG59754-01: DOWN SYNDROME CELL ADHESION MOLECULE
Expression of gene CG59754-02 and variant CG59754-01 was assessed using the primer-probe set Agl 305, described in Table CHA.
Table CHA. Probe Name Agl 305
Panel 4D Summary: Agl 305 Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
CI. CG59800-01: HEPARAN SULFATE D-GLUCOSAMINYL 3-O- SULFOTRANSFERASE-3B
Expression of gene CG59800-01 was assessed using the primer-probe set Ag3589, described in Table CIA.
Table CIA. Probe Name Ag3589
Results from Panels CNS_neurodegeneration_vl .0, 1.4, 2.2, and 4. ID are not included. The amp plots corresponding to these runs suggest that there were experimental difficulties with these runs.
CJ. CG59761-01: AXIN 1 (AXIS INHIBITION PROTEIN 1) (HAXIN) - isoforml, submitted to study DDSMT on 03/21/01 by cmiller; clone status=FIS; novelty=Novel; ORF start=97, ORF stop=2833, frame=l; 2949 bp.
Expression of gene CG59761-01 was assessed using the primer-probe set Ag3577, described in Table CJA. Results of the RTQ-PCR runs are shown in Tables CJB, CJC and CJD.
Table CJA. Probe Name Ag3577
Table CJB. CNS_neurodegeneration_vl.O
Table CJC. General_screening_panel_vl.4
Table CJD. Panel 4. ID
CNS_neurodegeneration_vl.O Summary: Ag3577 This panel confirms the expression of the CG59671-01 gene at low levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. As seen in panel 1.4, this gene is expressed at low levels throughout the CNS, including in amygdala, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. Therefore, this gene may play a role in other cential nervous system disorders such as Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
General_screening_panel_vl.4 Summary: Ag3577 Highest expression of the CG59671-01 gene is detected in a gastric cancer cell line sample (CTs=27.3). In addition, significant expression of this gene is associated with clusters of cell lines derived from ovarian cancer, breast cancer, and gastric cancer. Therefore, expression of this gene might be used to differentiate between these samples and other samples on this panel and as a marker for these cancers. The CG59671-01 gene encodes an Axin 1 protein, which is known play an important role in Wnt signalling transduction pathway. The Wnt/Wingless signaling transduction pathway plays an important role in both embryonic development and tumorigenesis. Beta- Catenin, a key component of the Wnt signaling pathway, interacts with the TCF/LEF family of transcription factors and activates transcription of Wnt target genes. A number of proteins such as the tumor suppressor APC and Axin are also involved in the regulation of the Wnt signaling pathway. Furthermore, mutations in APC or beta-catenin have been found to be
responsible for the genesis of human cancers (Akiyama T, 2000). Recently, Dahmen et al. (2001) have shown presence of a single somatic point mutation in exon 1 (Pro255Ser) and deletion of seven large of AXENl (12%) in 86 medulloblastoma (MB) samples and 11 MB cell lines. Therefore, AXINl may play a role as tumor suppressor gene in MBs. Furthermore, therapeutic modulation of the activity of this gene or its protein product might be beneficial in the tieatment of these cancers.
Among tissues with metabolic function, this gene is expressed at moderate to low levels in pituitary, adipose, adrenal gland, pancreas, thyroid, and adult and fetal skeletal muscle, heart, and liver. This widespread expression among these tissues suggests that this gene product may play a role in normal neuroendocrine and metabolic and that disregulated expression of this gene may contribute to neuroendocrine disorders or metabolic diseases, such as obesity and diabetes.
This gene is also expressed in all regions of the CNS examined. Please see Panel CNS_neurodegeneration_vl.0 for discussion of utility of this gene in the cential nervous system.
References:
1. Akiyama T. (2000) Wnt/beta-catenin signaling. Cytokine Growth Factor Rev l l(4):273-82.
2. Dahmen RP, Koch A, Denkhaus D, Tonn JC, Sorensen N, Berthold F, Behrens J, Birchmeier W, Wiestler OD, Pietsch T. (2001) Deletions of AXINl, a component of the WNT/wingless pathway, in sporadic medulloblastomas. Cancer Res 2001 Oct 1;61(19):7039- 43
Panel 4.1D Summary: Ag3577 Highest expression of the CG59671-01 gene is detected in resting NK Cells IL-2 cells (CTs=28.3). In addition, this gene is expressed at high to moderate levels in a wide range of cell types of significance in the immune response in health and disease. These cells include members of the T-cell, B-cell, endothelial cell, macrophage/monocyte, and peripheral blood mononuclear cell family, as well as epithelial and fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product may be involved in homeostatic processes for these and other cell types and tissues. Therefore,
modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
CK. CG59708-01 and CG59708-02 and CG59708-03: Ubiquitin carboxyl-terminal hydrolase 21
Expression of gene CG59708-01, full length clone CG59708-03 and variant CG59708- 02 was assessed using the primer-probe set Ag3511, described in Table CKA. Results of the RTQ-PCR runs are shown in Tables CKB, CKC and CKD. Please note that CG59708-03 represents a full-length physical clone of the CG59708-01 gene, validating the prediction of the gene sequence.
Table CKA. Probe Name Ag3511
Table CKB. CNS_neurodegeneration_vl.O
Table CKC. General_screening_panel_vl.4
Table CKD. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3511 This panel confirms the expression of CG59708-01 gene at low levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. However, as seen in panel 1.4, this gene is expressed at low levels throughout the CNS, including in amygdala, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. Therefore, this gene may play a role in other cential nervous system disorders such as Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
General_screening_panel_vl.4 Summary: Ag3511 Highest expression of the CG59708-01 is detected in a gastric cancer cell line sample (CT=27.1). Thus, expression of this gene can be used to distinguish this sample from other samples in this panel. In addition, high levels of expression of this gene are associated with breast cancer, ovarian cancer, and gastric cancer cell lines. Therefore, therapeutic modulation of the activity of this gene or its protein product might be beneficial in the tieatment of these cancers.
Among tissues with metabolic function, this gene is expressed at moderate to low levels in pituitary, adipose, adrenal gland, pancreas, thyroid, and adult and fetal skeletal muscle, heart, and liver. This widespread expression among these tissues suggests that this gene product may play a role in normal neuroendocrine and metabolic and that disregulated expression of this gene may contribute to neuroendocrine disorders or metabolic diseases, such as obesity and diabetes.
This gene is also expressed at moderate to low levels in all regions of the CNS examined. Please see Panel CNS neurodegeneration vl.O for discussion of utility of this gene in the cential nervous system.
Panel 4D Summary: Ag3511 Highest expression of the CG59708-01 gene is detected in a IL-4 treated NCI-H292 sample (CT=26.4). In addition, this gene is expressed at high to moderate levels in a wide range of cell types of significance in the immune response in health and disease. These cells include members of the T-cell, B-cell, endothelial cell, macrophage/monocyte, and peripheral blood mononuclear cell family, as well as epithelial and
fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product may be involved in homeostatic processes for these and other cell types and tissues. This pattern is in agreement with the expression profile in General_screening_panel_vl.5 and also suggests a role for the gene product in cell survival and proliferation. Therefore, modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
CL. CG59559-01: CPSase-related
Expression of gene CG59559-01 was assessed using the primer-probe set Ag3469, described in Table CLA. Results of the RTQ-PCR runs are shown in Tables CLB, CLC and CLD.
Table CLA. Probe Name Ag3469
Table CLB. CNS_neurodegeneration_vl.0
Table CLC. General_screening_panel_vl.4
Table CLP. Panel 4. ID
CNS_neurodegeneration_vl.0 Summary: Ag3469 This panel confirms the expression of the CG59559-01 gene at low levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. However, as seen in panel 1.4, this gene is expressed at low levels throughout the CNS, including in amygdala, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. Therefore, this gene may play a role in other central nervous system disorders such as Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
General_screening_panel_vl.4 Summary: Ag3469 Highest expression of the CG59559-01 gene is detected in sample derived from a lung cancer cell line (CT=25.6). Thus, expression of this gene can be used to distinguish this sample from other samples used in this panel. Furthermore, significant expression of this gene is associated with pancreatic cancer, CNS cancer and breast cancer cell lines. Therefore, therapeutic modulation of the activity of this gene or its protein product might be beneficial in the treatment of these cancers.
Among tissues with metabolic function, this gene is expressed in pituitary, adipose, adrenal gland, pancreas, thyroid, and adult and fetal skeletal muscle, heart, and liver. This widespread expression among these tissues suggests that this gene product may play a role in normal neuroendocrine and metabolic and that disregulated expression of this gene may contribute to neuroendocrine disorders or metabolic diseases, such as obesity and diabetes.
This gene is also expressed at low but significant levels in all regions of the CNS examined. Please see Panel CNS_neurodegeneration_vl.0 for discussion of utility of this gene in the cential nervous system.
Panel 4.1D Summary: Ag3469 Highest expression of the CG59559-01 gene is detected in sample derived CD40L and IL-4 treated B lymphocytes (CT=27.2). Fur5hermore, this gene is expressed at significant levels in a wide range of cell types of significance in the immune response in health and disease. These cells include members of the T-cell, B-cell, endothelial cell, and peripheral blood mononuclear cell family, as well as epithelial and fibroblast cell
types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product may be involved in homeostatic processes for these and other cell types and tissues. Therefore, modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
CM. CG59669-01: CARBONYL REDUCTASE
Expression of gene CG59669-01 was assessed using the primer-probe set Ag3505, described in Table CMA.
Table CMA. Probe Name Ag3505
CNS neurodegeneration vl.O Summary: Ag3505 Expression of the CG59669-01 gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3505 Results from one experiment with the CG59669-01 gene are not included. The amp plot indicates that there were experimental difficulties with this run.
Panel 4.1D Summary: Ag3505 Expression of the CG59669-01 gene is low/undetectable (CTs > 35) across all of the samples on this panel due to a probable probe or chemistry failure (data not shown).
Panel 5 Islet Summary: Ag3505 Expression of the CG59669-01 gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
CN. CG59679-01: CARBONYL REDUCTASE
Expression of gene CG59679-01 was assessed using the primer-probe set Ag3507, described in Table CNA.
Table CNA. Probe Name Ag3507
CNS neurodegeneration vl.O Summary: Ag3507 Expression of the CG59679-01 gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3507 Expression of the CG59679-01 gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
Panel 4.1D Summary: Ag3507 Expression of the CG59679-01 gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown). The data suggest that there may have been experimental difficulties with this run.
Panel 5 Islet Summary: Ag3507 Expression of CG59679-01 gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
CO. CG59644-01 : Putative protein phosphatase
Expression of gene CG59644-01 was assessed using the primer-probe set Ag3503, described in Table COA. Results of the RTQ-PCR runs are shown in Tables COB, COC and COD.
Table COA. Probe Name Ag3503
Table COB. CNS_neurodegeneration_vl.0
Table COC. General_screening_panel_vl.4
Tissue Name | Rel. Exp.(%) Ag3503, j Tissue Name j Rel. Exp.(%) Ag3503,
Table COD. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3503 This panel confirms the expression of this gene at low levels in the brains of an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in treatment of cential nervous system disorders.
General_screening_panel_vl.4 Summary: Ag3503 Expression of the CG59644-01 gene is highest in adult skeletal muscle (CT = 25.5). Interestingly, expression of this gene is much lower in fetal skeletal muscle (CT = 29.9), suggesting that expression of this gene may be used to distinguish adult and fetal skeletal muscle.
The CG59644-01 gene encodes a protein with homology to protein phosphatases. This gene is expressed at high to moderate levels in the majority of samples on this panel. However, expression of this gene appears to be higher in cancer cell lines when compared to normal adult tissues. This observation is consistent with the potential role for this gene product in cell survival and proliferation.
In addition, this gene is expressed at high levels in all regions of the cential nervous system examined, including amygdala, hippocampus, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. Therefore, this gene may play a role in cential nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Among tissues with metabolic or endocrine function, this gene is expressed at high to moderate levels in pancreas, adipose, adrenal gland, thyroid, pituitary gland, skeletal muscle, heart, liver and the gastiointestinal tiact. Therefore, therapeutic modulation of the activity of this gene may prove useful in the treatment of endocrine/metabolically related diseases, such as obesity and diabetes.
Panel 4D Summary: Ag3503 This gene is expressed at high to moderate levels in a wide range of cell types of significance in the immune response in health and disease. These cells include T cells, B cells, endothelial cells, macrophages, monocytes, dendritic cells, basophils, eosinophils and peripheral blood mononuclear cells, as well as epithelial and fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product may be involved in homeostatic processes for these and other cell types and tissues. This pattern is in agreement with the expression profile in General_screening_panel_vl.5 and also suggests a role for the gene product in cell survival and proliferation. Therefore, therapeutic modulation of the activity of this gene or its protein product may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
CP. CG59662-01: Cyclophilin
Expression of gene CG59662-01 was assessed using the primer-probe set Ag3504, described in Table CPA. Results of the RTQ-PCR runs are shown in Tables CPB and CPC.
Table CPA. Probe Name Ag3504
Table CPB. General_screening_panel_vl.4
Rel. Exp.(%) Ag3504, Rel. Exp.(%) Ag3504,
Tissue Name Tissue Name Run 217236170 Run 217236170
Table CPC. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3504 Expression of the CG59662-01 gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3504 The CG59662-01 gene is expressed at low levels in the majority of samples on this panel, with highest expression in a melanoma cell line (CT = 30). The CG59662-01 gene encodes a protein with homology to cyclophilin, a specific high-affinity binding protein for the immunosuppressant agent cyclosporin A.
Among tissues with metabolic or endocrine function, this gene is expressed at low levels in pancreas, thyroid, pituitary gland, skeletal muscle, heart, liver and the gastrointestinal tract. Therefore, therapeutic modulation of the activity of this gene may prove useful in the treatment of endocrine/metabolically related diseases, such as obesity and diabetes. Interestingly, this gene is expressed at higher levels in fetal liver (CT = 32.5) than in adult liver (CT = 36.4), suggesting that expression of this gene can be used to distinguish fetal and adult liver.
In addition, this gene is expressed at low levels in some regions of the central nervous system, including amygdala, hippocampus, substantia nigra, thalamus, and spinal cord. Therefore, this gene may play a role in central nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Panel 4D Summary: Ag3504 Significant expression of this gene is detected in a liver cirrhosis sample (CT = 34.4). Furthermore, expression of this gene is not detected at
significant levels in normal adult liver in Panel 1.4, suggesting that its expression is unique to liver cirrhosis. This gene encodes a putative cyclophilin; therefore, small molecule therapeutics could reduce or inhibit fibrosis that occurs in liver cirrhosis. In addition, expression of this putative cyclophilin could also be used for the diagnosis of liver cirrhosis.
CQ. CG59773-01: splice variant of myomegalin
Expression of gene CG59773-01 was assessed using the primer-probe set Ag3580, described in Table CQA. Results of the RTQ-PCR runs are shown in Tables CQB, CQC and CQD.
Table COA. Probe Name Ag3580
Table COB. CNS neurodegeneration vl.O
Table CQC. General_screening_panel_vl.4
Table COD. Panel 4. ID
CNS_neurodegeneration_vl.O Summary: Ag3580 Results from two experiments using the same probe/primer set are in excellent agreement. This panel confirms the expression of this gene at high to moderate levels in the brains of an independent group of individuals. This gene is found to be upregulated in the temporal cortex of Alzheimer's disease patients. Therefore, therapeutic modulation of this gene or its protein product may be used to decrease neuronal death and treat Alzheimer's disease.
General_screening_panel_vl.4 Summary: Ag3580 The CG59773-01 gene encodes a splice variant of the myomegalin protein, which is a component of the golgi/centrosome and interacts with a cyclic nucleotide phosphodiesterase (ref. 1). Expression of the CG59773-01 gene is
highest in the cerebellum (CT = 23.8). In addition, this gene is expressed at high levels in all other regions of the central nervous system examined, including amygdala, hippocampus, substantia nigra, thalamus, cerebral cortex, and spinal cord. Therefore, this gene may play a role in central nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Among tissues with metabolic or endocrine function, this gene is expressed at high to moderate levels in pancreas, adipose, adrenal gland, thyroid, pituitary gland, skeletal muscle, heart, liver and the gastrointestinal tract. Therefore, therapeutic modulation of the activity of this gene may prove useful in the tieatment of endocrine/metabolically related diseases, such as obesity and diabetes.
This gene is also expressed at very high levels in a number of melanoma cell lines. Therefore, therapeutic modulation of the activity of this gene or its protein product may be of benefit in the treatinent of melanoma.
References:
1. Verde I, Pahlke G, Salanova M, Zhang G, Wang S, Coletti D, Onuffer J, Jin SL, Conti M. Myomegalin is a novel protein of the golgi/centiosome that interacts with a cyclic nucleotide phosphodiesterase. J Biol Chem 2001 Apr 6;276(14):11189-98
Panel 4.1D Summary: Ag3580 This gene is expressed at high to moderate levels in a wide range of cell types of significance in the immune response in health and disease. These cells include T cells, B cells, endothelial cells, macrophages, monocytes, dendritic cells, basophils, eosinophils and peripheral blood mononuclear cells, as well as epithelial and fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product may be involved in homeostatic processes for these and other cell types and tissues. This pattern is in agreement with the expression profile in General_screening_j»anel_vl.5 and also suggests a role for the gene product in cell survival and proliferation. Therefore, therapeutic modulation of the activity of this gene or its protein product may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
CR. CG57460-01: N-ACETYLTRANSFERASE CAMELLO 2
Expression of gene CG57460-01 was assessed using the primer-probe set Ag3273, described in Table CRA. Results of the RTQ-PCR runs are shown in Tables CRB, CRC and CRD.
Table CRA. Probe Name Ag3273
Table CRB. CNS_neurodegeneration_vl.0
Table CRC. General_screening_panel_vl.4
Table CRP. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3273 Two experiments with the same probe and primer set produce results that are in excellent agreement. This panel confirms the expression of this gene at low to moderate levels in the brains of an independent group of individuals. Expression of this gene is found to be down-regulated in the temporal cortex of Alzheimer's disease patients. Therefore, up-regulation of this gene or its protein product, or treatment with specific agonists for this protein, may be of use in reversing the dementia/memory loss associated with Alzheimer's disease and neuronal death.
General_screening_panel_vl.4 Summary: Ag3273 Highest expression of the CG57460-01 gene is seen in fetal heart (CT=28.6). In addition, this gene is expressed at much higher levels in fetal heart when compared to expression in the adult heart (CT=38). Thus, expression of this gene may be used to differentiate between the fetal and adult source of this tissue. In addition, the higher expression in fetal heart suggests that this protein product may be involved in the development of this organ. Therefore, therapeutic modulation of the expression or function of this gene may be useful in the treatment of heart disease.
This gene also shows highly specific brain expression. Please see Panel CNS_neurodegeneration for discussion of utility of this gene in the central nervous system.
In addition, expression of this gene appears to be upregulated in a number of cancer cell lines when compared to the normal tissues. Specifically, expression of this gene appears to be higher in ovarian, breast, lung and renal cancer cell lines when compared to their respective normal tissues. Therefore, therapeutic modulation of the activity of this gene or its protein may be of benefit in the treatment of ovarian, breast, lung and renal cancer. The CG57460-01 gene encodes a transmembrane protein with homology to N-acetyltransferase Camello 2, a protein involved in cellular adhesion (ref. 1).
References:
1. Popsueva AE, Luchinskaya NN, Ludwig AV, Zinovjeva OY, Poteryaev DA, Feigelman MM, Ponomarev MB, Berekelya L, Belyavsky AV. Overexpression of camello, a member of a novel protein family, reduces blastomere adhesion and inhibits gastrulation in Xenopus laevis. Dev Biol 2001 Jun 15;234(2):483-96
Panel 4D Summary: Ag3273 Highest expression of the CG57460-01 is seen in eosinophils. In addition, differential expression is observed in the eosinophil cell line EOL-1 under resting conditions over that in EOL-1 cells stimulated by phorbol ester and ionomycin. Thus, this gene may be involved in eosinophil function. Therefore, therapeutic modulation of the expression or function of this gene may reduce eosinophil activation and be useful in the treatment of asthma and allergies.
In addition, significant expression in normal colon and thymus suggest a role for this gene in the normal homeostasis of these tissues. Therefore, therapeutic modulation of the expression or function of this gene may modulate immune function (T cell development) and be important for organ transplant, AIDS treatment or post chemotherapy immune reconstitution. Furthermore, since expression of this gene is decreased in colon samples from patients with EBD colitis and Crohn's disease relative to normal colon, therapeutic modulation of the activity of the protein encoded by this gene may be useful in the tieatment of inflammatory bowel disease.
CS. CG57464-01
Expression of gene CG57464-01 was assessed using the primer-probe set Ag3248, described in Table CSA. Results of the RTQ-PCR runs are shown in Tables CSB, CSC, CSD and CSE.
Table CSA. Probe Name Ag3248
Table CSB. CNSneurodegeneration vl.O
Table CSC. General_screening_panel_vl.4
Table CSD. Panel 2.2
CNS_neurodegeneration_vl.O Summary: Ag3248 Results from two experiments using the same probe/primer set gave results that are in excellent agreement. This panel confirms the expression of this gene at low levels in the brains of an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in treatment of central nervous system disorders.
General_screeningjpanel_vl.4 Summary: Ag3248 Expression of the CG57464-01 gene is highest in a breast cancer cell line (CT = 27). This also gene appears to be overexpressed in ovarian and CNS cancer cell lines when compared to the normal tissue controls. Thus, therapeutic modulation of the activity of this gene or its protein may be of benefit in the treatment of breast, ovarian and CNS cancer.
In addition, this gene is expressed at low levels in all regions of the central nervous system examined, including amygdala, hippocampus, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. Therefore, this gene may play a role in cential nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Among tissues with metabolic or endocrine function, this gene is expressed at low levels in pancreas, adipose, adrenal gland, thyroid, pituitary gland, skeletal muscle, heart, liver and the gastrointestinal tract. Therefore, therapeutic modulation of the activity of this gene may prove useful in the tieatment of endocrine/metabolically related diseases, such as obesity and diabetes.
Panel 2.2 Summary: Ag3248 This gene is expressed at low to moderate levels in the majority of samples on this panel, with highest expression detected in a sample derived from normal kidney (CT = 28.6). Expression of the CG57464-01 gene appears to be upregulated in a number of breast cancer samples when compared to normal breast. Thus, therapeutic modulation of the activity of this gene or its protein product may be of benefit in the treatment of breast cancer.
Panel 4D Summary: Ag3248 Expression of the CG57464-01 gene is highest in Ramos B cells tieated with ionomycin (CT = 29). Therefore, expression of this gene may be used as a marker of activated B cells. In addition, this gene is expressed at relatively high levels in lung fibroblasts as well as in the mucoepidermoid cell line NCI-H292 independent of treatment (CTs = 30), suggesting that therapeutic modulation of the activity of this gene or its protein product may be of benefit in the treatment of asthma and emphysema.
This gene is also expressed at low to moderate levels in a wide range of other cell types of significance in the immune response in health and disease. These cells include members of the T-cell, B-cell, endothelial cell, macrophage/monocyte, and peripheral blood mononuclear cell family, as well as epithelial and fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product may be involved in homeostatic processes for these and other cell types and tissues.
This pattern is in agreement with the expression profile in General_screening_panel_vl.5 and also suggests a role for the gene product in cell survival and proliferation.
Therefore, modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma,
allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
CT. CG57466-01 : Acetylglucosaminyltransferase
Expression of gene CG57466-01 was assessed using the primer-probe set Ag3249, described in Table CTA. Results of the RTQ-PCR runs are shown in Tables CTB, CTC and CTD.
Table CTA. Probe Name Ag3249
Table CTB. CNS_neurodegeneration_vl.0
Table CTC. General_screening_panel_vl.4
Fetal Kidney | 4.5 Pituitary gland Pool j 10.5
Renal ca. 786-0 | 2.5 Salivary Gland J 5.7
Renal ca. A498 j 1.4 Thyroid (female) j 5.8
Renal ca. ACHN j 22.2 Pancreatic ca. CAPAN2] 43.8
Renal ca. UO-31 | 32.8 Pancreas Pool j 11.3
Table CTD. Panel 4D
CNS_neurodegeneration_vl .0 Summary: Ag3249 This panel confirms the expression of tthhiiss ggeennee aatt llooww lleevveellss iinn tthhee b br ..a„,i„n„s off an„ i i„ndAea~p<e>„ndAae„n+t g ^r.^o.u.^p of individuals. However, no differential expression of this gene was ; detected between Alzheimer's diseased postmortem
brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in treatment of cential nervous system disorders.
General_screening_panel_vl.4 Summary: Ag3249 The CG57466-01 gene encodes a protein with homology to beta-l,3-galactosyltransferases, which catalyze the formation of type I oligosaccharides (ref. 1). Expression of this gene is highest in a breast cancer cell line (CT = 28.1). In addition, expression of this gene appears to be upregulated in pancreatic and gastric cancer cell lines when compared to their respective normal tissues. Thus, therapeutic modulation of the activity of this gene or its protein product may be of benefit in the treatment of breast, pancreatic and gastric cancer.
This gene also shows significant levels of expression in trachea, bladder and fetal lung. Interestingly, CG57466-01 gene expression is much higher in fetal lung (CT = 28.3) than in adult lung (CT = 32.2), suggesting that expression of this gene can be used to distinguish adult from fetal lung.
In addition, this gene is expressed at low levels in all regions of the central nervous system examined, including amygdala, hippocampus, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. Therefore, this gene may play a role in cential nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Among tissues with metabolic or endocrine function, this gene is expressed at low to moderate levels in pancreas, adipose, adrenal gland, thyroid, pituitary gland, heart, fetal liver and the gastrointestinal tiact. Therefore, therapeutic modulation of the activity of this gene may prove useful in the treatment of endocrine/metabolically related diseases, such as obesity and diabetes.
References:
1. Shiraishi N, Natsume A, Togayachi A, Endo T, Akashima T, Yamada Y, Lmai N, Nakagawa S, Koizumi S, Sekine S, Narimatsu H, Sasaki K. Identification and characterization of three novel beta 1,3-N-acetylglucosaminyltransferases structurally related to the beta 1,3- galactosyltransferase family. J Biol Chem 2001 Feb 2;276(5):3498-507
Panel 4D Summary: Ag3249 This transcript is most highly expressed in a cluster of treated and untreated samples derived from the NCI-H292 cell line, a human airway epithelial cell line that produces mucins (CTs = 30-32). Mucus overproduction is an important feature of bronchial asthma and chronic obstructive pulmonary disease samples. The transcript is also expressed at lower but still significant levels in small airway epithelium treated with IL-1 beta and TNF-alpha. The expression of the transcript in this mucoepidermoid cell line that is often used as a model for airway epithelium (NCI-H292 cells) suggests that this transcript may be important in the proliferation or activation of airway epithelium. Therefore, therapeutics designed with the protein encoded by the tianscript may reduce or eliminate symptoms caused by inflammation in lung epithelia in chronic obstructive pulmonary disease, asthma, allergy, and emphysema.
CU. CG57468-01: multidrug resistance protein 1
Expression of gene CG57468-01 was assessed using the primer-probe set Ag3250, described in Table CUA. Results of the RTQ-PCR runs are shown in Tables CUB.
Table CUA. Probe Name Ag3250
Table CUB. General_screening_panel_vl.4
CNS_neurodegeneration_vl.O Summary: Ag3250 Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3250 Expression of the CG57468-01 gene is highest in normal breast (CT = 23.8). In addition, this gene is highly expressed in fetal/adult kidney and fetal/adult liver (CTs = 26-27). Thus, expression of this gene may be used to distinguish these tissues from the other samples on this panel. Strikingly, expression of this gene is much lower in breast, kidney, and liver cancer cell lines. Therapeutic modulation of the activity of this gene or its protein product may be of benefit in the tieatment of these types of cancers.
Panel 4D Summary: Ag3250 Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
CV. CG59609-01: PEPTIDYL-PROLYL CIS-TRANS ISOMERASE A
Expression of gene CG59609-01 was assessed using the primer-probe set Ag3494, described in Table CVA. Results of the RTQ-PCR runs are shown in Tables CVB and CVC.
Table CVA. Probe Name Ag3494
Start SEQ ID
Primers Sequences Length Position NO:
Forward 5 ' -ccgttctatcagccatggt-3 ' 19 704
TET-5 ' -ccccaccaggttcttagacatcatcg-3 '
Probe TAMRA 26 25 705
Table CVC. Panel 4D
CNS neurodegeneration vl.O Summary: Ag3494 Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3494 Expression of the CG59609-01 gene is highest in testis (CT = 34.3). In addition, low but significant expression of this gene is detected in a breast cancer cell line and an ovarian cancer cell line. Thus, expression of this gene may be used to distinguish these samples from the other samples on this panel. Furthermore, therapeutic modulation of the activity of this gene may be of benefit in the treatment of fertility, breast cancer, and ovarian cancer.
Panel 4D Summary: Ag3494 Expression of the CG59609-01 gene is highest in a liver cirrhosis sample (CT = 34.3). In addition, low but significant expression of this gene is detected in samples from thymus as well as from normal and EBD colon. Thus, expression of this gene may be used to distinguish these samples from the other samples on this panel. Furthermore, therapies designed with the protein encoded for by this gene may potentially modulate liver function and play a role in the identification and tieatment of inflammatory or autoimmune diseases which effect the liver including liver cirrhosis and fibrosis.
CW. CG59613-01: PROLIFERATING CELL NUCLEAR ANTIGEN
Expression of gene CG59613-01 was assessed using the primer-probe set Ag3496, described in Table CWA. Results of the RTQ-PCR runs are shown in Tables CWB and CWC.
Table CWA. Probe Name Ag3496
Table CWB. General_screening_panel_vl.4
Table CWC. Panel 4D
CNS neurodegeneration vl.O Summary: Ag3496 Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3496 Expression of the CG59613-01 gene is highest in fetal and adult kidney (CTs = 31). This gene is also expressed at higher levels in fetal lung (CT = 31.4) than in adult lung (CT = 34.8), suggesting that expression of this gene can be used to distinguish adult and fetal lung and that this gene may play a role in lung
development and regeneration. Differentially higher expression in fetal tissues also occurs in brain and skeletal muscle.
In general, expression of this gene is associated with normal tissues rather than cancer cell lines. Specifically, CG59613-01 gene expression is downregulated in pancreatic, colon, gastric, renal, lung, breast and prostate cancer cell lines when compared to their respective normal tissues. Therefore, therapeutic modulation of the activity of this gene may be of benefit in the tieatment of these cancers.
Among tissues with metabolic or endocrine function, this gene is expressed at low levels in pancreas, adipose, adrenal gland, fetal skeletal muscle, and the gastrointestinal tract. Therefore, therapeutic modulation of the activity of this gene may prove useful in the treatment of endocrine/metabolically related diseases, such as obesity and diabetes.
Panel 4D Summary: Ag3496 Expression of the CG59613-01 gene is highest in small airway epithelium tieated with TNF alpha and IL-1 beta (CT = 29.4). In addition, this gene is substantially upregulated in keratinocytes tieated with TNF alpha and IL-1 beta. Low expression of this gene is also seen in lung and dermal fibroblasts independent of treatment. Therefore, therapeutics designed with the protein encoded by the transcript may reduce or eliminate symptoms caused by inflammation of the lung and skin in chronic obstructive pulmonary disease, asthma, allergy, emphysema, and psoriasis.
CX. CG59619-01: ACTIN, CYTOPLASMIC 2
Expression of gene CG59619-01 was assessed using the primer-probe set Ag3498, described in Table CXA. Results of the RTQ-PCR runs are shown in Tables CXB and CXC.
Table CXA. Probe Name Ag3498
Table CXB. General_screening_panel_vl.4
Tissue Name Rel. Exp.(%) Ag3498, Tissue Name Rel. Exp.(%) Ag3498,
Table CXC. Panel 4. ID
CNS neurodegeneration vl.O Summary: Ag3498 Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
General_screening_panel_vl.4 Summary: Ag3498 The CG59619-01 gene is only expressed at detectable levels in the adult kidney (CT = 34.2). Thus, expression of this gene can be used to distinguish kidney from the other samples on this panel. In addition, expression of this gene is much lower in fetal kidney (CT = 38.7), suggesting that this gene can be used to distinguish between the fetal and adult source of this tissue. Furthermore, this gene is not expressed at detectable levels in renal cancer cell lines. Therefore, therapeutic modulation of this gene may be of use in the treatment of renal cell carcinoma.
Panel 4.1D Summary: Ag3498 Expression of the CG59619-01 gene is highest in activated eosinophils (CT = 25.7), displaying 10-fold upregulation when compared to the control eosinophils. Therefore, therapies designed with the protein encoded for by this gene could block or inhibit inflammation or tissue damage due to eosinophil activation in response to asthma, ulcerative colitis and parasitic diseases.
The CG59619-01 gene is expressed at moderate levels in the majority of samples on this panel, including T cells, B cells, endothelial cells, macrophages, monocytes, dendritic cells, basophils and peripheral blood mononuclear cells, as well as epithelial and fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product may be involved in
homeostatic processes for these and other cell types and tissues. Therefore, modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
CY. CG59621-01: SELENIDE-WATER DIKINASE 1
Expression of gene CG59621-01 was assessed using the primer-probe set Ag3764, described in Table CYA.
Table CYA. Probe Name Ag3764
General_screening_panel_vl.4 Summary: Ag3764 Results from one experiment with the CG59621-01 gene are not included. The amp plot indicates that there were experimental difficulties with this run.
Panel 4.1D Summary: Ag3764 Expression of this gene is low/undetectable (CTs > 35) across all of the samples on this panel (data not shown).
CZ. CG59625-01: GLUCOSE TRANSPORTER TYPE 3
Expression of gene CG59625-01 was assessed using the primer-probe set Ag3499, described in Table CZA. Results of the RTQ-PCR runs are shown in Tables CZB and CZC.
Table CZA. Probe Name Ag3499
Table CZB. CNSneurodegeneration vl.O
Table CZC. Panel 4D
CNS_neurodegeneration_vl.O Summary: Ag3499 This panel confirms the expression of this gene at moderate levels in the brains of an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in treatment of central nervous system disorders.
Panel 4D Summary: Ag3499 Expression of the CG59625-01 gene is highest in PMA/ionomycin-treated lymphokine activated killer (LAK) cells (CT = 24.3). Since these
cells are involved in tumor immunology and tumor cell clearance, as well as virally and bacterial infected cells, therapeutic modulation of this gene product may alter the functions of these cells and lead to improvement in cancer cell killing as well as host immunity to microbial and viral infections.
This gene is also expressed at high levels in stimulated keratinocytes, dendritic cells, monocytes and macrophages, suggesting that small molecule therapeutics designed against the CG59625-01 protein could reduce or inhibit inflammation in asthma, emphysema, allergy, psoriasis, arthritis, or any other condition in which localization/activation of these cell types is important.
This gene is also expressed at moderate levels in a number of other cell types of significance in the immune response in health and disease.
DA. CG59887-01 and CG59887-02: Amino Acid/Metabolite Permease
Expression of gene CG59887-01 and full length clone CG59887-02 was assessed using the primer-probe set Ag4715, described in Table DAA. Please note that CG59887-02 represents a full-length physical clone of the CG59887-02 gene, validating the prediction of the gene sequence.
Table DAA. Probe Name Ag4715
General_screening_panel_vl.4 Summary: Ag4715 Expression of the CG59887-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.) The amp plot indicates that there is a high probability of a probe failure.
DB. CG59857-01: RHOTEKIN
Expression of gene CG59857-01 was assessed using the primer-probe set Ag3622, described in Table DBA. Results of the RTQ-PCR runs are shown in Tables DBB, DBC and DBD.
Table DBA. Probe Name Ag3622
Start 3 SEQ ID
Primers Sequences Length Position j NO:
[Forward 5 ' -acatcctggaggacctgaatat - 3 ' j 22 J 84 j 722
TET- 5 ' - ctctacattcggcagatggcactcag- 3 ' -
Probe TAMRA T 26 j 107 723
(Reverse 5 ' -ggatctcatggtctagcttcct - 3 ' 1 22 1 155 j 724
Table DBB. CNS_neurodegeneration_vl.0
Table DBC. General_screening_panel_vl.4
CNS_neurodegeneration_vl.O Summary: Ag3622 This panel confirms the expression of the CG59857-01 gene at significant levels in the brains of an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in treatment of central nervous system disorders.
General_screening_panel_vl.4 Summary: Ag3622 Two experiments with the same probe and primer set show highest expression of the CG59857-01 gene in spinal cord samples (CTs=26-28). In addition, high levels of expression of this gene are seen in brain derived tissue, including samples from amygdala, hippocampus, substantia nigra, thalamus, cerebellum, cerebral cortex, and CNS cancer cell lines. Therefore, expression of this gene could be used to distinguish between brain derived samples and other samples used in this panel. Furthermore, this gene may play a role in central nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Significant expression is also detected in fetal skeletal muscle (CTs=27-31). Interestingly, this gene is expressed at much higher levels in fetal when compared to adult skeletal muscle (CTs=32-34). This observation suggests that expression of this gene can be used to distinguish fetal from adult skeletal muscle. In addition, the relative overexpression of this gene in fetal skeletal muscle suggests that the protein product may enhance muscular growth or development in the fetus and thus may also act in a regenerative capacity in the adult. Therefore, therapeutic modulation of the protein encoded by this gene could be useful in
treatment of muscle related diseases. More specifically, treatment of weak or dystrophic muscle with the protein encoded by this gene could restore muscle mass or function.
Among tissues with metabolic function, this gene is expressed at moderate to low levels in pituitary, adipose, adrenal gland, pancreas, thyroid, and adult and fetal skeletal muscle, heart, and liver. This widespread expression among these tissues suggests that this gene product may play a role in normal neuroendocrine and metabolic and that disregulated expression of this gene may contribute to neuroendocrine disorders or metabolic diseases, such as obesity and diabetes.
Panel 4.1D Summary: Ag3622 Highest expression of the CG59857-01 gene is seen in IL- 9/EL-13 treated lung fibroblasts (CT=31). In addition, significant expression is seen in clusters of treated and untreated lung and dermal fibroblasts, epithelium and endothelium. Therefore, modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, and psoriasis.
DC. CG59855-01 and CG59855-02: ATP SYNTHASE SUBUNIT C
Expression of gene CG59855-01 and full length clone CG59855-02 was assessed using the primer-probe set Ag3621, described in Table DCA. Results of the RTQ-PCR runs are shown in Tables DCB and DCC. Please note that CG59855-02 represents a full-length physical clone of the CG59855-02 gene, validating the prediction of the gene sequence.
Table DCA. Probe Name Ag3621
Table DCB. General_screening_panel_vl.4
Rel. Exp.(%) Ag3621, Rel. Exp.(%) Ag3621,
Tissue Name Tissue Name Run 217702346 Run 217702346
Adipose 0.0 Renal ca. TK-10 0.0
Table DCC. Panel 4. ID
CNS_neurodegeneration_vl.0 Summary: Ag3621 Expression of the CG59855-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
General_screening_panel_vl.4 Summary: Ag3621 Expression of the CG59855-01 gene is restricted to samples from fetal lung and adult pancrease(CTs=34.5-35). Thus, expression of this gene can be used to distinguish this sample from other samples in the panel.
The CG59855-01 gene encodes a homologue of ATP synthase subunit c, mitochondrial precursor. Subunit c is an intrinsic membrane component of ATP synthase, and in mammals it is encoded by two expressed nuclear genes, PI and P2. Both genes encode the same mature c subunit, but the mitochondrial import pre-sequences in the precursors of subunit c are different (ref. 1). Each ATP synthase complex has multiple copies of subunit C. The mitochondrial ATP synthase uses energy derived from a proton gradient to synthesize ATP. The structure of this complex has been referred to as a 'lollipop,' as the soluble FI catalytic unit is attached to the mitochondrial inner membrane via the F0 unit containing subunit c. F0 subunit C transports protons across the mitochondrial inner membrane to the Fl-ATPase (ref. 2).
Subunit C of the Fo region of the ATP synthase complex of the inner mitochondrial membrane is found in high concentrations in lysosomes in late infantile neuronal ceroid lipofuscinosis (Batten's disease). Kominami et al. (1995, Ref 3) found marked delay of degradation of subunit C in patient fibroblasts with no significant differences between contiol and patient cells with regard to degradation of cytochrome oxidase subunit IV. Furthermore, accumulation of labeled subunit C in the mitochondrial fraction was detected before lysosomal appearance of the radiolabeled subunit, suggesting to the authors a specific failure in the degradation of subunit C after its normal inclusion in mitochondria and its consequent
accumulation in lysosomes. Jolly (1995, ref 4) reported that subunit C represents more than 50% of the accumulated metabolites in the ovine form of the disease and also accumulates significantly in late infantile and juvenile forms of the human disease and several other animal forms. The author suggested that the extreme hydrophobicity and lipophilicity of subunit C may be in part responsible.
References:
1. Dyer MR, Walker JE. (1993) Sequences of members of the human gene family for the c subunit of mitochondrial ATP synthase. Biochem J 293 ( Pt l):51-64
2. OMEM 603192
3. Kominami E, Ezaki J, Wolfe LS. (1995) New insight into lysosomal protein storage disease: delayed catabolism of ATP synthase subunit c in Batten disease. Neurochem Res 20(11)1305-9
4. Jolly RD. (1995) Batten disease (ceroid-lipofuscinosis): the enigma of subunit c of mitochondrial ATP synthase accumulation. Neurochem Res 20(11): 1301-4
Panel 4.1D Summary: Ag3621 Expression of the CG59855-01 gene is exclusively seen in resting monocytes (CT=32). Thus, expression of this gene can be used to distinguish this sample from other samples in the panel. In addition, expression of this gene in monocytes suggests a role for the gene product in their function as antigen-presenting cells. This suggests that antibodies or small molecule therapeutics that block the function of this protein may be useful as anti-inflammatory therapeutics for the tieatment of autoimmune and inflammatory diseases and for the tieatment of immunosupressed individuals.
DD. CG59807-01: Nuclear Hormone Receptor/Zinc Finger
Expression of gene CG59807-01 was assessed using the primer-probe set Ag3591, described in Table DDA. Results of the RTQ-PCR runs are shown in Tables DDB and DDC.
Table PDA. Probe Name Ag3591
Table DDB. General_screening_panel_vl.4
Table DDC. Panel 4. ID
General_screening_panel_vl.4 Summary: Ag3591 Highest expression of the CG59807-01 gene is detected in the gastric cancer cell line(CT=28). In addition, high expression of this gene is seen in samples derived from CNS cancer, colon cancer, breast cancer, ovarian cancer, prostate cancer cell lines (CTs=28-31). Therefore, therapeutic modulation of the activity of this gene or its protein product, through the use of small molecule drugs, protein therapeutics or antibodies, might be beneficial in the treatment of these cancers.
In addition, expression of this gene is higher in fetal liver (CT=31) as compared to the corresponding adult tissues (CTs=34). Thus, expression of this gene can be used to distinguish between the fetal and adults source of this tissue.
Among tissues with metabolic or endocrine function, this gene is expressed at high to moderate levels in pancreas, adipose, adrenal gland, thyroid, pituitary gland, skeletal muscle, heart, liver and the gastrointestinal tract. Therefore, therapeutic modulation of the activity of this gene may prove useful in the treatment of endocrine/metabolically related diseases, such as obesity and diabetes.
This gene is also expressed at high levels in all regions of the central nervous system examined, including amygdala, hippocampus, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. Therefore, this gene may play a role in cential nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Panel 4.1D Summary: Ag3591 Highest expression of the CG59807-01 gene is detected in treated mucoepidermoid NCI-H292 cells. In addition, this gene is expressed at high to moderate levels in a wide range of cell types of significance in the immune response in health and disease. These cells include members of the T-cell, B-cell, endothelial cell, macrophage/monocyte, and peripheral blood mononuclear cell family, as well as epithelial and fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product may be involved in homeostatic processes for these and other cell types and tissues. This pattern is in agreement with the expression profile in General_screening_panel_vl.5 and also suggests a role for the gene product in cell survival and proliferation. Therefore, modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
DE. CG59805-01: Nuclear Hormone Receptor/Zinc Finger
Expression of gene CG59805-01 was assessed using the primer-probe set Ag3590, described in Table DEA. Results of the RTQ-PCR runs are shown in Tables DEB, DEC and DED.
Table PEA. Probe Name Ag3590
Table DEB. CNS_neurodegeneration_vl.0
Table DEC. General_screening_panel_vl.4
Table DED. Panel 4. ID
CNS_neurodegeneration_vl.0 Summary: Ag3590 This panel confirms the expression of the CG59805-01 gene at low levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in treatment of central nervous system disorders.
General_screening_panel_vl.4 Summary: Ag3590 Highest expression of the CG59805-01 gene is detected in one of the breast cancer cell line BT 549 (CT=26). In addition, expression of this gene is high in CNS cancer, gastric cancer, and prostate cancer cell lines. Therefore, expression of this gene can be used to distinguish these samples from other samples in this panel and it can be used as marker for detection of these cancers. Furthermore, therapeutic modulation of the activity of the protein encoded by this gene may be beneficial in the treatment of these cancers.
Among tissues with metabolic or endocrine function, this gene is expressed at high to moderate levels in pancreas, adipose, adrenal gland, thyroid, pituitary gland, skeletal muscle, heart, liver and the gastrointestinal tract. Therefore, therapeutic modulation of the activity of this gene may prove useful in the treatment of endocrine/metabolically related diseases, such as obesity and diabetes.
In addtion, this gene is expressed at high levels in all regions of the central nervous system examined, including amygdala, hippocampus, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. Therefore, this gene may play a role in cential nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression.
Panel 4.1D Summary: Ag3590 Highest expression of the CG59805-01 gene is detected in PMA/ionomycin treated Ku-812 (basophil) cells (CT=29). In addition, this gene is expressed at high to moderate levels in a wide range of cell types of significance in the immune response in health and disease. These cells include members of the T-cell, B-cell, endothelial cell, macrophage/monocyte, and peripheral blood mononuclear cell family, as well as epithelial and fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product maybe involved in homeostatic processes for these and other cell types and tissues. This pattern is in agreement with the expression profile in General_screening_panel_vl.4 and also suggests a role for the gene product in cell survival and proliferation. Therefore, modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
DF. CG59928-01: Novel Universal Stress (USP) Domain Containg Protein
Expression of gene CG59928-01 was assessed using the primer-probe set Ag3636, described in Table DFA. Please note that this sequence is represented by a full length clone.
Table DFA. Probe Name Ag3636
CNS neurodegeneration vl.O Summary: Ag3636 Expression of the CG59928-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.) The amp plot indicates that there is a high probability of a probe failure.
General_screening_panel_vl.4 Summary: Ag3636 Expression of the CG59928-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.) The amp plot indicates that there is a high probability of a probe failure.
Panel 4.1D Summary: Ag3636 Expression of the CG59928-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.) The amp plot indicates that there is a high probability of a probe failure.
DG. CG59947-01: VOLTAGE-GATED POTASSIUM CHANNEL PROTEIN KV3.3
Expression of gene CG59947-01 was assessed using the primer-probe set Ag3635, described in Table DGA. Results of the RTQ-PCR runs are shown in Tables DGB, DGC, DGD and DGE.
Table DGA. Probe Name Ag3635
Table DGB. CNS_neurodegeneration_vl.0
Table DGC. Panel 2.2
Table DGE. Panel CNS 1
CNS_neurodegeneration_vl.O Summary: Ag3635 This panel confirms the expression of CG59947-01 gene at low levels in the brain in an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. This gene encodes a potassium channel protein homolog. The significant levels of expression in the brain may indicate a role for this protein in signal processing in the central nervous system.
References:
1. Rudy B, Chow A, Lau D, Amarillo Y, Ozaita A, Saganich M, Moreno H, Nadal MS, Hernandez-Pineda R, Hernandez-Cruz A, Erisir A, Leonard C, Vega-Saenz de Miera E.
2. Contributions of Kv3 channels to neuronal excitability. Ann N Y Acad Sci 1999 Apr 30;868:304-43
General_screening_panel_vl.4 Summary: Ag3635 Results from one experiment with the CG59947-01 gene are not included. The amp plot indicates that there were experimental difficulties with this run.
Panel 2.2 Summary: Ag3635 Highest expression of the CG59447-01 gene is seen in normal kidney tissue adjacent to a tumor (CT=28). In addition, expression appears to be higher in normal kidney tissue than in the adjacent tumor in six out of nine matched pairs. Conversely expression appears to be higher in breast cancer than in matched normal breast tissue. Thus, expression of this gene could be used to differentiate between these samples and other samples on this panel and as a marker for kidney and breast cancers. Furthermore, therapeutic modulation of the expression or function of this protein may be effective in the tieatment of breast and kidney cancer.
Panel 4.1D Summary: Ag3635 This gene is expressed at high to moderate levels in a wide range of cell types of significance in the immune response in health and disease, with highest expression in anti CD40 dendritic cells (CT=28.1). Other cells that express this protein include members of the T-cell, B-cell, endothelial cell, macrophage/monocyte, and peripheral blood mononuclear cell family, as well as epithelial and fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product may be involved in homeostatic processes for these and other cell types and tissues. Therefore, modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
Panel CNS_1 Summary: Ag3635 Expression in this panel confirms expression of the CG59947-01 gene in the brain. Please see Panel CNS_neurodegeneration_vl.0 for discussion of utility of this gene in the central nervous system.
DH. CG59938-01: arylsulfatase
Expression of gene CG59938-01 was assessed using the primer-probe set Ag3634, described in Table DHA.
Table DHA. Probe Name Ag3634
Start SEQ ID
Primers Sequences Length Position NO:
Forward 5 ' -agccaatgaaagaggagaaagt-3 22 870 740
TET-5 1 -cttccctcatgctgaaggaggcactt-3
Probe TAMRA 26 894 741
Reverse j5 ' -cccttttgtacctttcaatgaa-3 ' 22 j 923 742
CNS_neurodegeneration_vl.O Summary: Ag3634 Expression of the CG55938-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
General_screening_panel_vl.4 Summary: Ag3634 Expression of the CG55938-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
Panel 2.2 Summary: Ag3634 Expression of the CG55938-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
Panel 4.1D Summary: Ag3634 Expression of the CG55938-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
DI. CG59746-01: Ubiquitin Carboxyl-terminal Hydrolase
Expression of gene CG59746-01 was assessed using the primer-probe set Ag3574, described in Table DIA.
Table DIA. Probe Name Ag3574
CNS_neurodegeneration_vl.O Summary: Ag3574 Expression of the CG59746-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
General_screening_panel_vl.4 Summary: Ag3574 Expression of the CG59746-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
Panel 2.2 Summary: Ag3574 Expression of the CG59746-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
Panel 4.1D Summary: Ag3574 Expression of the CG59746-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
Panel CNS_1 Summary: Ag3574 Expression of the CG59746-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown.)
DJ. CG88613-01: INOSITOL 1,4,5-TRISPHOSPHATE 3-KINASE ISOENZYME
Expression of gene CG88613-01 was assessed using the primer-probe set Ag3647, described in Table DJA. Results of the RTQ-PCR runs are shown in Tables DJB, DJC and DJD.
Table DJA. Probe Name Ag3647
Table DJB. CNS neurodegeneration vl.O
Table DJC. General_screening_panel_vl.4
Table DJD. Panel 4. ID
CNS_neurodegeneration_vl.O Summary: Ag3647 This panel confirms the expression of this gene at moderate levels in the brains of an independent group of individuals. However, no differential expression of this gene was detected between Alzheimer's diseased postmortem brains and those of non-demented controls in this experiment. Please see Panel 1.4 for a discussion of the potential utility of this gene in treatment of central nervous system disorders.
General_screening_panel_vl.4 Summary: Ag3647 Expression of the CG88613-01 gene is highest in a gastric cancer cell line (CT = 28). Expression of this gene appears to be upregulated in a number of cancer cell lines when compared to normal tissues. Specifically, CG88613-01 gene expression is somewhat higher in breast and ovarian cancers when compared to their respective normal tissues. Thus, therapeutic modulation of the activity of this gene or its protein product, using small molecule drugs, antibodies or protein therapeutics, may be of benefit in the tieatment of gastric, breast and ovarian cancer.
In addition, this gene is expressed at moderate levels in all regions of the central nervous system examined, including amygdala, hippocampus, substantia nigra, thalamus, cerebellum, cerebral cortex, and spinal cord. The CG88613-01 gene encodes a protein that is identical to a protein now known in the public domain as inositol 1,4, 5 -triphosphate 3-kinase
C (ref. 1). Inositol 1,4,5-trisphosphate 3-kinase (ITPK) catalyzes the phosphorylation of
Ins(l,4,5)P3 to Ins(l,4,5)P4, both of which are modulators of calcium homeostasis. Calcium is one of the most important intracellular messengers in the brain, being essential for neuronal development, synaptic transmission and plasticity, and the regulation of various metabolic pathways (ref. 2). Therefore, this gene may play a role in cential nervous system disorders such as Alzheimer's disease, Parkinson's disease, epilepsy, multiple sclerosis, schizophrenia and depression. Furthermore, this gene is also expressed in tissues with metabolic or endocrine
function, including pancreas, adipose, adrenal gland, thyroid, pituitary gland, skeletal muscle, heart, liver and the gastrointestinal tract. Therefore, therapeutic modulation of the activity of this gene may prove useful in the treatment of endocrine/metabolically related diseases, such as obesity and diabetes.
References:
1. Dewaste V, Pouillon V, Moreau C, Shears S, Takazawa K, Emeux C. Cloning and expression of a cDNA encoding human inositol 1,4,5-trisphosphate 3-kinase C. Biochem J 2000 Dec l;352 Pt 2:343-51
2. Mattson MP, Chan SL. Dysregulation of cellular calcium homeostasis in Alzheimer's disease: bad genes and bad habits. J Mol Neurosci 2001 Oct;17(2):205-24
Panel 4.1D Summary: Ag3647 Results from two experiments using the same probe/primer set are in excellent agreement. Expression of the CG88613-01 gene is highest in keratinocytes tieated with the inflammatory cytokines TNF-a and IL-lb(CT = 29.5). Therefore, modulation of the expression or activity of this protein through the application of small molecule therapeutics may be useful in the treatment of psoriasis and wound healing.
This gene is also expressed at moderate levels in small airway epithelial cells, bronchial epithelium, and lung microvascular endothelial cells. Endothelial cells are known to play important roles in inflammatory responses by altering the expression of surface proteins that are involved in activation and recruitment of effector inflammatory cells (ref. 1). Expression in small airway epithelial cells, bronchial epithelium, lung microvascular endothelial cells suggests that the protein encoded by this tianscript may be involved in lung disorders including asthma, allergies, chronic obstructive pulmonary disease, and emphysema. This gene is homologoust o PI-3-kinase which is involved in cell survival and receptor signaling of a number of cells of importance in the immune response in health and disease, including lung pathologies. Therefore, Small molecule antagonists of this gene product may lead to amelioration of symptoms associated with asthma, allergies, chronic obstructive pulmonary disease, and emphysema.
This gene is expressed at low levels in the remainder of the samples on this panel, suggesting that the gene product may play an important role in homeostasis of a number of cell types.
References:
1. Siddiqui RA, English D. Phosphatidylinositol 3'-kinase-mediated calcium mobilization regulates chemotaxis in phosphatidic acid-stimulated human neutrophils. Biochim Biophys Acta 2000 Jan 3;1483(l):161-73
2. Condliffe AM, Cadwallader KA, Walker TR, Rintoul RC, Cowburn AS, Chilvers ER. Phosphoinositide 3-kinase: a critical signalling event in pulmonary cells. Respir Res 2000;l(l):24-9
DK. CG59993-01 and CG59993-02: synaptotagmin II
Expression of gene CG59993-01 and variant CG59993-02 was assessed using the primer-probe set Ag3645, described in Table DKA. Results of the RTQ-PCR runs are shown in Tables DKB, DKC and DKD.
Table DKA. Probe Name Ag3645
Table DKB. CNS_neurodegeneration_vl.O
Table DKC. General_screening_panel_vl.4
Table DKD. Panel 4. ID
Macrophages LPS j 0.0 JThymus 43.2
HUVEC none j 0.0 JKidney 25.9
HUVEC starved j 0.0 1
CNS_neurodegeneration_vl.O Summary: Ag3645 While no association between the expression of the CG59993-01 gene and the presence of Alzheimer's disease is detected in this panel, these results confirm the expression of this gene in areas that degenerate in Alzheimer's disease, including the cortex and hippocampus. Synaptotagmin expression is altered in the brain of Alzheimer's patients, possibly explaining impaired synaptogenesis and/or synaptosomal loss secondary to neuronal loss observed in the neurodegenerative disorder. It may also represent, reflect or account for the impaired neuronal transmission in Alzheimer's disease (AD), caused by deterioration of the exocytic machinery. Since the this gene is a homolog of synaptotagmin, agents that potentiate the expression or function of the protein encoded by the this gene may be useful in the tieatment of Alzheimer's disease.
General_screening_panel_vl.4 Summary: Ag3645 The CG59993-01 gene is a homolog of synaptotagmin, and shows high to moderate expression across all brain regions with highest expression in the cerebellum (CT = 26.4) Synaptotagmin is a presynaptic protein involved in synaptic vesicle release, making this an ideal drug target for diseases such as epilepsy, in which reduction of neurotransmission is beneficial. Selective inhibition of this gene or its protein product may therefore be useful in the treatment of seizure disorders. Furthermore, selective inhibition of neural transmission through antagonism of the protein encoded by this gene may show therapeutic benefit in psychiatric diseases where it is believed that inappropriate neural connections have been established, such as schizophrenia and bipolar disorder. In addition, antibodies against synaptotagmin may cause Lambert-Eaton myasthenic syndrome. Therefore, peptide fragments of the protein encoded by this gene may serve to block the action of these antibodies and treat Lambert-Eaton myasthenic syndrome.
Panel 4.1D Summary: Ag3645 Expression of the CG59993-01 gene is restricted to a sample derived from astrocytes treated with TNF-alpha and IL-1 beta (CT=33.9). This expression in samples related to the central nervous system is consistent with results of the previous panels and suggests that modulation of this protein could be beneficial in the treatment of CNS disease-associated inflammation or neurodegeneration, including mutliple sclerosis.
DL. CG59991-01: OOPLASM SPECIFIC PROTEIN
Expression of gene CG59991-01 was assessed using the primer-probe set Ag3644, described in Table DLA. Results of the RTQ-PCR runs are shown in Tables DLB and DLC.
Table DLA. Probe Name Ag3644
Table DLB. General_screening_panel_vl.4
Table DLC. Panel 4. ID
CNS_neurodegeneration_vl.O Summary: Ag3644 Expression of the CG59991-01 gene is low/undetectable in all samples on this panel (CTs>35). (Data not shown).
General_screening_panel_vl.4 Summary: Ag3644 Expression of the CG59991-01 gene is restricted to a sample derived from a lung cancer cell line (CT=27.2). Thus, expression of this gene could be used to differentiate between this sample and other samples on this panel and as a marker to detect the presence of lung cancer. Furthermore, therapeutic modulation of the expression or function of this gene may be effective in the tieatment of lung cancer.
Panel 4.1D Summary: Ag3644 Expression of the CG59991-01 gene is restricted to samples derived from the basophil cell line KU-812 (CTs=32). Thus, expression of this gene could be
used as a marker of this cell type. Basophils release histamines and other biological modifiers in repose to allergens and play an important role in the pathology of asthma and hypersensitivity reactions. Therefore, the specific pattern of expression of this gene suggests that therapeutic modulation of the expression or function of the protein encoded by this gene may block or inhibit inflammation or tissue damage due to basophil activation in response to asthma, allergies, hypersensitivity reactions, psoriasis, and viral infections.
DM. CG59987-01 and CG59987-02: RHOPHILIN
Expression of gene CG59987-01 and full length clone CG59987-02 was assessed using the primer-probe set Ag3643, described in Table DMA. Results of the RTQ-PCR runs are shown in Tables DMB and DMC. Please note that CG59987-02 represents a full-length physical clone of the CG59987-01 gene, validating the prediction of the gene sequence.
Table DMA. Probe Name Ag3643
Table DMB. General_screening_panel_vl.4
Table PMC. Panel 4. ID
CNS_neurodegeneration_vl.O Summary: Ag3643 Results from one experiment with the CG59987-01 gene are not included. The amp plot indicates that there were experimental difficulties with this run.
General_screening_panel_vl.4 Summary: Ag3643 Expression of the CG59987-01 gene is highest in a breast cancer cell line (CT=25.3). In addition, significant levels of expression are seen in clusters of cell lines derived from brain, gastric, colon, lung, and ovarian cancers. In addition, expression overall appears to be higher in samples derived from cancer cell lines than in normal tissues. Thus, expression of this gene could be used as a marker to detect the presence of cancer. This gene encodes a homolog of rhophilin, a rho GTPase that is involved in a signaling pathway that regulates cell adhesion, among other functions. Therefore, therapeutic modulation of the expression or function of this gene may be effective in the treatment of these cancers.
Among tissues with metabolic function, this gene is expressed at moderate to low levels in pituitary, adipose, adrenal gland, pancreas, thyroid, skeletal muscle, and adult and fetal heart, and liver. This widespread expression among these tissues suggests that this gene product may play a role in normal neuroendocrine and metabolic and that disregulated expression of this gene may contribute to neuroendocrine disorders or metabolic diseases, such as obesity and diabetes.
This gene is also expressed at moderate to low levels in the CNS and may be a small molecule target for the treatment of neurologic diseases.
Panel 4.1D Summary: Ag3643 Expression of the CG59987-01 gene is highest in NCI-H292 cells stimulated by IL-9(CT=29.2), The gene is also expressed in a cluster of treated and untieated NCI-H292 mucoepidermoid cell line samples. The tianscript is also expressed at lower but still significant levels in both small airway and bronchial epithelium tieated with IL- 1 beta and TNF-alpha. In comparison, expression in the normal lung is relatively low. The expression of the tianscript in activated normal epithelium as well as a cell line that is often used as a model for airway epithelium (NCI-H292 cells) suggests that this transcript may be important in the proliferation or activation of airway epithelium. Therefore, therapuetics designed with the protein encoded by this tianscript could be important in the tieatment of diseases which include lung airway inflammation such as asthma and COPD.
DN. CG59971-01 and CG59971-02: Leucine Rich Repeat protein
Expression of gene CG59971-01 and variant CG59971-02 was assessed using the primer-probe set Ag3639, described in Table DNA. Results of the RTQ-PCR runs are shown in Tables DNB and DNC.
Table DNA. Probe Name Ag3639
Start | SEQ ID
Primers Sequences Length Position j NO:
(Forward 5 ' -ttctgccaacttcagctacaat -3 ' 1 22 I 510 J 755
TET-5 ' -cttagacagctccctgcgcctcttgt - 3 ' -
Probe TAMRA 543 1 756 26
JReverse 5 ' -acttgattgtggcttaggttca-3 ' 1 22 1 584 j 757
Table DNB. General_screening_panel_vl.4
Table DNC. Panel 4. ID
CNS_neurodegeneration_vl.O Summary: Ag3639 Results from one experiment with the CG59971-01 gene are not included. The amp plot indicates that there were experimental difficulties with this run.
General_screening_panel_vl.4 Summary: Ag3639 Expression of the CG59971-02 gene is ubiquitous in this panel, with highest expression in a breast cancer cell line (CT=26.6). Overall, expression of this gene appears to be higher in samples derived from cancer cell lines than in normal tissues. This widespread expression suggests that this gene product is involved in cell growth and prolideration. Thus, expression of this gene could be used as a marker to
detect the presence of cancer. Furthermore, therapeutic modulation of the expression or function of this gene may be useful in the tieatment of cancer.
In addition, this gene is expressed at much higher levels in fetal lung and liver (CTs=29-30) when compared to expression in the adult counterpart (CTs=33). Thus, expression of this gene may be used to differentiate between the fetal and adult sources of these tissue.
Among tissues with metabolic function, this gene is expressed at moderate to low levels in pituitary, adipose, adrenal gland, pancreas, thyroid, and adult and fetal skeletal muscle, heart, and liver. This widespread expression among these tissues suggests that this gene product may play a role in normal neuroendocrine and metabolic and that disregulated expression of this gene may contribute to neuroendocrine disorders or metabolic diseases, such as obesity and diabetes.
This gene is also highly expressed in the brain, with highest expression in the cerebellum (CT = 28.5), with moderate expression in other CNS regions as well including, amygdala, hippocampus, cerebral cortex, substantia nigra and thalamus. This gene encodes a leucine-rich repeat protein. Leucine rich repeats (LRR) mediate reversible protein-protein interactions and have diverse cellular functions, including cellular adhesion and signaling. Several of these proteins, such as connectin, slit, chaoptin, and Toll have pivotal roles in neuronal development in Drosophila and may play significant but distinct roles in neural development and in the adult nervous system of humans (Ref. 1). In Drosophilia, the LRR region of axon guidance proteins has been shown to be critical for their function (especially in axon this gene shows high expression in the brain, it is an excellent candidate neuronal guidance protein for axons, dendrites and/or growth cones in general. Therefore, therapeutic modulation of the levels of this protein, or possible signaling via this protein, may be of utility in enhancing/directing compensatory synaptogenesis and fiber growth in the CNS in response to neuronal death (stroke, head trauma), axon lesion (spinal cord injury), or neurodegeneration (Alzheimer's, Parkinson's, Huntington's, vascular dementia or any neurodegenerative disease).
References:
1. Battye R., Stevens A., Perry R.L., Jacobs J.R. (2001) Repellent signaling by Slit requires the leucine-rich repeats. J. Neurosci. 21 : 4290-4298.
Panel 4.1D Summary: Ag3639 The CG59971-01 gene is expressed at high to moderate levels in a wide range of cell types of significance in the immune response in health and disease. Highest expression of the gene is seen in resting monocytes (CT=28.6). Significant levels of expression are also seen in members of the T-cell, B-cell, endothelial cell, macrophage/monocyte, and peripheral blood mononuclear cell family, as well as epithelial and fibroblast cell types from lung and skin, and normal tissues represented by colon, lung, thymus and kidney. This ubiquitous pattern of expression suggests that this gene product may be involved in homeostatic processes for these and other cell types and tissues. This pattern is in agreement with the expression profile in General_screening_panel_vl.4 and also suggests a role for the gene product in cell survival and proliferation. Therefore, modulation of the gene product with a functional therapeutic may lead to the alteration of functions associated with these cell types and lead to improvement of the symptoms of patients suffering from autoimmune and inflammatory diseases such as asthma, allergies, inflammatory bowel disease, lupus erythematosus, psoriasis, rheumatoid arthritis, and osteoarthritis.
Example D. Identification of Single Nucleotide Polymorphisms in NOVX nucleic acid sequences
Variant sequences are also included in this application. A variant sequence can include a single nucleotide polymorphism (SNP). A SNP can, in some instances, be referred to as a "cSNP" to denote that the nucleotide sequence containing the SNP originates as a cDNA. A SNP can arise in several ways. For example, a SNP may be due to a substitution of one nucleotide for another at the polymorphic site. Such a substitution can be either a transition or a transversion. A SNP can also arise from a deletion of a nucleotide or an insertion of a nucleotide, relative to a reference allele. In this case, the polymoφhic site is a site at which one allele bears a gap with respect to a particular nucleotide in another allele. SNPs occurring within genes may result in an alteration of the amino acid encoded by the gene at the position of the SNP. Intiagenic SNPs may also be silent, when a codon including a SNP encodes the same amino acid as a result of the redundancy of the genetic code. SNPs occurring outside the region of a gene, or in an intron within a gene, do not result in changes in any amino acid sequence of a protein but may result in altered regulation of the expression pattern. Examples include alteration in temporal expression, physiological response regulation, cell type expression regulation, intensity of expression, and stability of transcribed message.
SeqCalling assemblies produced by the exon linking process were selected and extended using the following criteria. Genomic clones having regions with 98% identity to all
or part of the initial or extended sequence were identified by BLASTN searches using the relevant sequence to query human genomic databases. The genomic clones that resulted were selected for further analysis because this identity indicates that these clones contain the genomic locus for these SeqCalling assemblies. These sequences were analyzed for putative coding regions as well as for similarity to the known DNA and protein sequences. Programs used for these analyses include Grail, Genscan, BLAST, HMMER, FASTA, Hybrid and other relevant programs.
Some additional genomic regions may have also been identified because selected SeqCalling assemblies map to those regions. Such SeqCalling sequences may have overlapped with regions defined by homology or exon prediction. They may also be included because the location of the fragment was in the vicinity of genomic regions identified by similarity or exon prediction that had been included in the original predicted sequence. The sequence so identified was manually assembled and then may have been extended using one or more additional sequences taken from CuraGen Corporation's human SeqCalling database. SeqCalling fragments suitable for inclusion were identified by the CuraTools™ program SeqExtend or by identifying SeqCalling fragments mapping to the appropriate regions of the genomic clones analyzed.
The regions defined by the procedures described above were then manually integrated and corrected for apparent inconsistencies that may have arisen, for example, from miscalled bases in the original fragments or from discrepancies between predicted exon junctions, EST locations and regions of sequence similarity, to derive the final sequence disclosed herein. When necessary, the process to identify and analyze SeqCalling assemblies and genomic clones was reiterated to derive the full length sequence (Alderbom et al., Determination of Single Nucleotide Polymorphisms by Real-time Pyrophosphate DNA Sequencing. Genome Research. 10 (8) 1249-1265, 2000).
Variants are reported individually but any combination of all or a select subset of variants are also included as contemplated NOVX embodiments of the invention.
NOV5a has one SNP variant, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ TD NOs: 13 and 14 , respectively. The nucleotide sequence of the NOV5a variant differs as shown in Table SNPl .
Table SNPl. NOV5a variants.
NOV9a has eight SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ TD NOs:21 and 22 , respectively. The nucleotide sequence of the NOV9a variant differs as shown in Table SNP2.
Table SNP2. NOV9a variants.
NOVl 4a has five SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ ID NOs:43 and 44 , respectively. The nucleotide sequence of the NOVl 4a variant differs as shown in Table SNP3.
Table SNP3. NOV14a variants.
NOVl 5a has twelve SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ TD NOs:53 and 54 , respectively. The nucleotide sequence of the NOVl 5a variant differs as shown in Table SNP4.
Table SNP4. NOVl 5a variants.
NOVl 7a has four SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ ID NOs:61 and 62 , respectively. The nucleotide sequence of the NOVl 7a variant differs as shown in Table SNP5.
Table SNP5. NOVl 7a variants.
NOVl 9a has one SNP variant, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ TD NOs:71 and 72 , respectively. The nucleotide sequence of the NOV19a variant differs as shown in Table SNP6.
Table SNP6. NOVl 9a variants.
NOV21a has one SNP variant, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ ED NOs:75 and 76 , respectively. The nucleotide sequence of the NOV2 la variant differs as shown in Table SNP7.
Table SNP7. NOV21a variants.
NOV38a has two SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ ID NOs: 123 and 124 , respectively. The nucleotide sequence of the NOV38a variant differs as shown in Table SNP8.
Table SNP8. NOV38a variants.
NOV39a has three SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ ED NOs: 125 and 126 , respectively. The nucleotide sequence of the NOV39a variant differs as shown in Table SNP9.
Table SNP9. NOV39a variants.
NOV46a has four SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ ED NOs: 143 and 144 , respectively. The nucleotide sequence of the NOV46a variant differs as shown in Table SNP 10.
Table SNP10. NOV46a variants.
NOV49a has two SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ TD NOs: 151 and 152 , respectively. The nucleotide sequence of the NOV49a variant differs as shown in Table SNPl 1.
Table SNP11. NOV49a variants.
NOV50a has one SNP variant, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ TD NOs: 153 and 154 , respectively. The nucleotide sequence of the NOV50a variant differs as shown in Table SNP 12.
Table SNP12. NOV50a variants.
NOV51a has five SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ ED NOs: 155 and 156 , respectively. The nucleotide sequence of the NOV51a variant differs as shown in Table SNP13.
Table SNP13. NOVSla variants.
NOV52a has one SNP variant, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ TD NOs: 157 and 158 , respectively. The nucleotide sequence of the NOV52a variant differs as shown in Table SNP 14.
Table SNP14. NOV52a variants.
NOV55a has one SNP variant, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ ED NOs: 163 and 164 , respectively. The nucleotide sequence of the NOV55a variant differs as shown in Table SNP15.
Table SNP15. NOV55a variants.
NOV60a has one SNP variant, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ TD NOs: 183 and 184 , respectively. The nucleotide sequence of the NOV55a variant differs as shown in Table SNP 16.
Table SNP16. NOV60a variants.
NOV65a has one SNP variant, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ TD NOs: 195 and 196, respectively. The nucleotide sequence of the NOV65a variant differs as shown in Table SNP17.
Table SNPl 7. NOV65a variants.
NOV68a has one SNP variant, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ ED NOs:201 and 202, respectively. The nucleotide sequence of the NOV68a variant differs as shown in Table SNPl 8.
Table SNPl 8. NOV68a variants.
NOV72a has one SNP variant, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ ED NOs:209 and 210, respectively. The nucleotide sequence of the NOV72a variant differs as shown in Table SNP 19.
Table SNP19. NOV72a variants.
NOV80a has four SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ ED NOs:225 and 226, respectively. The nucleotide sequence of the NOV80a variant differs as shown in Table SNP20.
Table SNP20. NOV80a variants.
NOV81a has four SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ ID NOs:229 and 230, respectively. The nucleotide sequence of the NOV8 la variant differs as shown in Table SNP21.
Table SNP21. NOV81a variants.
NOV89a has two SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ TD NOs:249 and 250, respectively. The nucleotide sequence of the NOV89a variant differs as shown in Table SNP22.
Table SNP22. NOV89a variants.
NOV94a has one SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ ID NOs:269 and 270, respectively. The nucleotide sequence of the NOV94a variant differs as shown in Table SNP23.
Table SNP23. NOV94a variants.
NOV96a has two SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ ED NOs:273 and 274, respectively. The nucleotide sequence of the NOV96a variant differs as shown in Table SNP24.
Table SNP24. NOV96a variants.
NOV99a has one SNP variant, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ TD NOs:283 and 284, respectively. The nucleotide sequence of the NOV99a variant differs as shown in Table SNP25.
Table SNP25. NOV99a variants.
NOV105a has three SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ ID NOs:299 and 300, respectively. The nucleotide sequence of the NOV105a variant differs as shown in Table SNP26.
Table SNP26. NOV105a variants.
NOVl 13a has three SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ ED NOs:315 and 316, respectively. The nucleotide sequence of the NOVl 13a variant differs as shown in Table SNP27.
Table SNP27. NOVl 13a variants.
NOVl 14a has two SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ ED NOs:319 and 320, respectively. The nucleotide sequence of the NOVl 14a variant differs as shown in Table SNP28.
Table SNP28. NOVl 14a variants.
NOVl 16a has two SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ ID NOs:325 and 326, respectively. The nucleotide sequence of the NOVl 16a variant differs as shown in Table SNP29.
Table SNP29. NOVllόa variants.
NOVl 17a has three SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ TD NOs:329 and 330, respectively. The nucleotide sequence of the NOVl 17a variant differs as shown in Table SNP30.
Table SNP30. NOV117a variants.
NOVl 24a has six SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ TD NOs:343 and 344, respectively. The nucleotide sequence of the NOVl 24a variant differs as shown in Table SNP31.
Table SNP31. NOV124a variants.
NOVl 26a has two SNP variants, whose variant positions for its nucleotide and amino acid sequences is numbered according to SEQ ID NOs:349 and 350, respectively. The nucleotide sequence of the NOVl 26a variant differs as shown in Table SNP32.
Table SNP32. NOV126a variants.
OTHER EMBODIMENTS
Although particular embodiments have been disclosed herein in detail, this has been done by way of example for purposes of illustration only, and is not intended to be limiting with respect to the scope of the appended claims, which follow. In particular, it is contemplated by the inventors that various substitutions, alterations, and modifications may be made to the invention without departing from the spirit and scope of the invention as defined by the claims. The choice of nucleic acid starting material, clone of interest, or library type is believed to be a matter of routine for a person of ordinary skill in the art with knowledge of the embodiments described herein. Other aspects, advantages, and modifications considered to be within the scope of the following claims. The claims presented are representative of the inventions disclosed herein. Other, unclaimed inventions are also contemplated. Applicants reserve the right to pursue such inventions in later claims.