WO2025019385A1 - Tools for interrogating dynamic organizational principles of protein complexes in vivo - Google Patents

Tools for interrogating dynamic organizational principles of protein complexes in vivo Download PDF

Info

Publication number
WO2025019385A1
WO2025019385A1 PCT/US2024/037959 US2024037959W WO2025019385A1 WO 2025019385 A1 WO2025019385 A1 WO 2025019385A1 US 2024037959 W US2024037959 W US 2024037959W WO 2025019385 A1 WO2025019385 A1 WO 2025019385A1
Authority
WO
WIPO (PCT)
Prior art keywords
dna
caliper
antibody
stranded oligonucleotide
protein
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/US2024/037959
Other languages
French (fr)
Inventor
Alon Goren
Tianyao XU
Christopher Benner
Sven Heinz
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
University of California Berkeley
University of California San Diego UCSD
Original Assignee
University of California Berkeley
University of California San Diego UCSD
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by University of California Berkeley, University of California San Diego UCSD filed Critical University of California Berkeley
Publication of WO2025019385A1 publication Critical patent/WO2025019385A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G01MEASURING; TESTING
    • G01NINVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
    • G01N33/00Investigating or analysing materials by specific methods not covered by groups G01N1/00 - G01N31/00
    • G01N33/48Biological material, e.g. blood, urine; Haemocytometers
    • G01N33/50Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing
    • G01N33/68Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing involving proteins, peptides or amino acids
    • G01N33/6803General methods of protein analysis not limited to specific proteins or families of proteins
    • G01N33/6845Methods of identifying protein-protein interactions in protein mixtures
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • C12Q1/6804Nucleic acid analysis using immunogens
    • GPHYSICS
    • G01MEASURING; TESTING
    • G01NINVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
    • G01N33/00Investigating or analysing materials by specific methods not covered by groups G01N1/00 - G01N31/00
    • G01N33/48Biological material, e.g. blood, urine; Haemocytometers
    • G01N33/50Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing
    • G01N33/68Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing involving proteins, peptides or amino acids
    • G01N33/6875Nucleoproteins
    • GPHYSICS
    • G01MEASURING; TESTING
    • G01NINVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
    • G01N2458/00Labels used in chemical analysis of biological material
    • G01N2458/10Oligonucleotides as tagging agents for labelling antibodies

Definitions

  • BACKGROUND Dynamic protein complexes and transient protein-protein interactions are integral for the vast majority of normal and cancer associated processes such as cellular metabolism, signal transduction networks and regulation of chromatin structure. Current methods cannot fully capture this complexity due to an inability to simultaneously detect multiple dynamic interactions. While current tools to study protein-protein interactions (PPI; e.g., co- immunoprecipitation, proximity ligation assay, two-hybrid and affinity purification–mass spectrometry (AP-MS) may offer means for studying networks of interactors in an unbiased manner, they have several key limitations.
  • a DNA-caliper comprising two single-stranded oligonucleotide arms, wherein the arms are joined at their 5’ ends resulting in each arm having a free 3’ end, wherein each 3’ end comprises a single stranded oligonucleotide anchor’ region, wherein each arm comprises a primer region, wherein at least one arm of the caliper comprises a cleavage site at its 5’ end.
  • at least one arm of the DNA caliper further comprises a capture molecule.
  • the capture molecule comprises biotin or digoxigenin.
  • the single stranded oligonucleotide anchor’ region is complementary to a single stranded oligonucleotide anchor region on another oligonucleotide.
  • each anchor’ region has the same sequence. In other embodiment, each anchor’ region has a different sequence.
  • each arm is about 15 to about 1000 nucleotides long.
  • each single stranded oligonucleotide anchor’ is about 10 to about 50 nucleotides long.
  • each primer region is about 5 to about 45 nucleotides long.
  • the cleavage site is a restriction enzyme site, a deoxy uridine (dU) or a UV light cleavage site.
  • compositions comprising the DNA-caliper described herein.
  • the composition comprises a carrier.
  • One embodiment provides an antibody covalently linked to a 5’ end of a single-stranded oligonucleotide tail, wherein the oligonucleotide tail comprises at its 3’ end a single stranded oligonucleotide anchor region, a unique barcode, and a single stranded oligonucleotide unique molecular identifier (UMI) and at its 5’ end a restriction enzyme site.
  • the single stranded oligonucleotide anchor region is complementary to a single stranded oligonucleotide anchor’ region on another molecule.
  • the barcode is specific to the antibody.
  • the UMI is a random sequence. In one embodiment, the UMI is about 5 to about 45 nucleotides long. In one embodiment, the single-stranded oligonucleotide tail is about 15 to about 1000 nucleotides long. In one embodiment, the anchor is about 10 to about 50 nucleotides long.
  • One embodiment provides a composition comprising the antibody-oligo described herein. In one embodiment, the composition comprises a carrier. One embodiment provides a composition comprising the DNA-caliper described herein and the antibody-oligo described herein. In one embodiment, the composition comprises a carrier.
  • One embodiment provides a method for mapping the genomic co-localization of one or more proteins on a chromatin fragment comprising: (a) incubating chromatin fragments with a plurality of antibodies, each of the plurality of antibodies binding to each of the one or more proteins, wherein each antibody comprises a single stranded oligonucleotide tail comprising a hook region, a single stranded oligonucleotide anchor region, a unique barcode, and a single stranded oligonucleotide unique molecular identifier (UMI); and wherein each chromatin fragment comprises a chromatin DNA molecule ligated to adaptors comprising a chromatin DNA molecule anchor’ region and a chromatin DNA molecule UMI; (b) contacting the chromatin fragments with a DNA-caliper molecule comprising two single-stranded oligonucleotide arms, wherein the arms are joined at their 5’ ends resulting in each arm having a free 3’ end
  • One embodiment further comprises denaturing the extended DNA- caliper from the chromatin fragment.
  • the capture molecule comprises biotin or digoxigenin.
  • the extended DNA-calipers are captured/isolated by the capture molecule.
  • One embodiment further comprises adding a polyC sequence with terminal transferase to the 3’ end of the extended DNA-caliper corresponding to the chromatin fragment and then hybridizing and ligating a double stranded DNA molecule with two overhanging regions to the hook region and the polyC overhang on the extended DNA-caliper, bringing the two ends of the extended DNA-caliper together, wherein the 3’ ends of DNA-caliper are extended with polymerase using the other arm as a template to generate a double stranded construct.
  • the cleavage site is a restriction enzyme site, dU or UV light cleavage site.
  • One embodiment further comprises linearizing the double stranded construct at the cleavage site to generate a linearized fragment.
  • the linearized fragment is PCR amplified using the primer sequences on the DNA-caliper.
  • the PCR amplified products are sequenced.
  • the single stranded oligonucleotide has a nucleic acid sequence of SEQ ID NO: 1.
  • the DNA-caliper has a nucleic acid sequence of SEQ ID NOs: 2 and 3.
  • the single stranded oligonucleotide has a nucleic acid sequence of SEQ ID NO: 4.
  • the permeabilized cells are lysed cells.
  • the cleavage site is a restriction endonuclease site, dU or UV light cleavage site.
  • One embodiment further comprises denaturing the extended DNA-calipers from the single stranded oligonucleotide of the first antibody and the single stranded oligonucleotide of the second antibody or prior to or after adding DNA polymerase remove the antibodies by restriction enzyme cutting and ligating the ends; after DNA polymerase extension of the DNA-caliper ends and the oligonucleotide tails the extended molecule can be linearized, PCR amplified and sequenced.
  • the capture molecule comprises biotin or digoxigenin.
  • the extended DNA-calipers are captured/isolated by the capture molecule.
  • One embodiment further comprises hybridizing and ligating a double stranded DNA molecule with two overhanging regions to the hook regions on the extended DNA-caliper bringing the two ends of the DNA-caliper together, wherein the 3’ ends of DNA-caliper are extended with DNA polymerase using the other arm as a template to generate a double stranded construct.
  • One embodiment further comprises linearizing the double stranded construct at the cleavage site to generate a linearized fragment.
  • the linearized fragment is PCR amplified using the primer sequences on the DNA-caliper.
  • the single stranded oligonucleotide tail comprises one or more unique molecular identifier (UMI) sequences.
  • the single stranded oligonucleotide tail comprises a barcode sequence or molecular tag.
  • the abundance of the target protein comprises detecting the presence of the UMI sequence.
  • One embodiment further comprises separating the proteins in the sample by gel electrophoresis and transferring the separated proteins to a membrane prior to contacting the sample with the one or more antibodies.
  • One embodiment further comprises contacting the one or more antibodies comprising one or more single stranded oligonucleotide tails with a single stranded binding protein prior to incubating the membrane with the one or more antibodies.
  • One embodiment further comprises a protein ladder of known size.
  • One embodiment further comprises contacting the membrane with one or more antibodies and cutting the membrane into segments.
  • the protein ladder of known size is used such that each segment contains about a 10kDa to about a 20kDa range of sizes of protein.
  • the abundance of one or more proteins in the sample is performed by comparing the relative abundance of the one or more UMI sequences from the sequencing results, wherein the one or more UMI sequences are specific to the antibody that targets a protein.
  • the sample is sonicated prior to gel electrophoresis.
  • One embodiment further comprises one or more single stranded oligonucleotide DNA-caliper molecules comprising two single-stranded oligonucleotide arms, wherein the arms are joined at their 5’ ends resulting in each arm having a free 3’ end, a capture molecule, a cleavage site, a first primer and a second primer sequence, wherein one primer sequence is on each side of the cleavage site.
  • One embodiment further comprises a tool for cutting a protein membrane, such as a pair of scissors, blade, knife or cookie cutter type device.
  • One embodiment further comprises a first set of PCR primers and a second set of PCR primers, wherein the first set of PCR primers comprises a first sequence and are formulated for use in a first PCR reaction that converts the single stranded oligonucleotide to a double stranded oligonucleotide, and wherein the second set of PCR primers comprises a second sequence and are formulated for use in a second PCR reaction that generates double stranded oligonucleotides for sequencing.
  • + denotes the proportion of true interactions in the entire dataset
  • ⁇ ⁇ is the readout scaling factor for PQ-seq detection protocol and sequencing
  • ⁇ 6%:4; ⁇ and ⁇ 45%67 denote the readout scaling factors for specific and non-specific PPI detection in Prod- seq, respectively
  • ⁇ % and ⁇ % denote the levels of specific and non-specific antibody-oligo binding for each protein target ⁇ , respectively, wherein + , , ⁇ ⁇ , 6%:4; ⁇ , ⁇ 45%67 , ⁇ A,...,4 , ⁇ A,...,4 , ⁇ ! , ⁇ >5?
  • interacting proteins illustrate PPI between proteins(c,k) under condition(h) (e.g., EZH2 and SUZ12 in hiPSCs); a DNA fragment with multiple proteins bound to it depicts protein-DNA interactions between proteins(b,d,g,h,j) and genomic locus(h) under condition(a). Note, while not represented, the abundance of each protein studied is also measured.
  • FIGS.2A-2B Steps of Prod-seq and WhIP-seq. Simplified depiction of the tools A.
  • Prod- seq (1) Fixed and lysed cells are incubated with multiple Ab-oligos, each antibody is linked to an oligonucleotide with a barcode (BC n, m, etc.), a UMI and an anchor site (Anchor1). At each 3’ end of the DNA-caliper there is an Anchor1’. (2) Hybridization of the DNA-caliper to the Anchor1 (Anc1) on each Ab-oligo captures the barcodes and UMIs by DNA extension (dashed) at both ends. (3) The extended DNA-calipers are converted into a sequencing library. B. WhIP-seq: (1) Sheared chromatin is incubated with multiple Ab-oligos, these are the same ones used for Prod- seq.
  • the chromatin fragments are ligated to WhIP-adaptors that include a UMI and a different anchor (Anchor 2).
  • the DNA-caliper used here contains Anchor2’ on one arm.
  • Target sequences are captured by hybridization of the two DNA-caliper anchors (Anc1’ and Anc2’) to the Ab-oligo and the adaptor, followed by DNA extension (dashed).
  • Extended DNA-calipers are converted into a sequencing library.
  • FIG.2C One embodiment of generation sequence ready constructs continued from Figure 2A is depicted (another embodiment is depicted in Figure 6). The two ends of the extended DNA- caliper are brought together for intramolecular hybridization of the Hook and Hook’.
  • each of the DNA-caliper arms are extended using the other arm as a template.
  • the double stranded fragments are used as a template for generation of a sequencing ready construct by PCR from the primer sequences (P1 and P2) embedded in the original DNA-caliper.
  • Different DNA barcodes are also depicted.
  • FIG. 2D Conversion of the Prod-seq hybridization-extension products into a sequence library.
  • the extended DNA-caliper is used as a template for DNA polymerase to generate a new normal strand (dashed) with P1 (embedded in the oligonucleotides linked to the antibodies) as a primer (bi-directionally extended DNA-caliper).
  • FIGS. 6A-6H Ligation of adaptors to cross-linked chromatin.
  • A Cross-linked chromatin was immobilized with a H3K27ac antibody and ligated to Illumina adaptors, and the product was then amplified by PCR. From left to right: ladder; ligated chromatin; original sheared chromatin.
  • B A representative view from the IGV browser (37) showing the similarity between data generated by chromatin ligation (black) and by ChIP-seq (red) from the same mouse ES cells.
  • the barcodes and UMIs are captured by hybridization of the Anchor 1’ (Anc1’) on the DNA-caliper arms to the Anchor 1 (Anc1) on each Ab-oligo, followed by DNA polymerase extension (dashed lines) at both ends. c-d. Next, the extended DNA-calipers are denatured and pulled-down using the biotin. A Splint molecule with two overhanging Hook’ regions (c, top) is hybridized and ligated to the extended DNA-caliper. e. Following intramolecular hybridization, the 3’ ends of each of the DNA-caliper arms are extended using the other arm as a template. f.
  • the molecule is linearized by USER enzyme that cleaves the dU and the fragments are used as a template for generation of sequencing-ready constructs by PCR from the primer sequences (P1 and P2; The dU and the primers are embedded in the original DNA- caliper, panel a.).
  • e Another embodiment of Prod-seq.
  • the diverse DNA barcodes that are captured by the process are depicted by different color combinations.h. Prod-seq, capture PPI by sequencing.
  • Sheared chromatin is incubated with multiple Ab-oligos, these are the same ones used for Prod-seq (FIG.
  • Target sequences are captured by hybridization of the two DNA-caliper anchors (Anc1’ and Anc2’) to the Ab-oligos and the adaptor 3’ overhang, followed by DNA extension (dashed lines) at both ends.
  • c. The extended DNA-calipers are denatured and pulled-down using the Biotin moiety.
  • d-f Conversion of the extended DNA-calipers into sequencing-ready constructs via intramolecular hybridization-extension.
  • a Splint molecule (left), with 5’ phosphorylated termini is added. Note, this Splint has one arm with Hook’ to match the Ab-oligo, while the other has a Poly(G) overhang at the 3’ end.
  • FIG. 9D depicts protein quantification followed by sequencing (PQ-seq).
  • PQ-seq is a protein quantification method using DNA oligonucleotides conjugated to antibodies and ending with sequencing to get the amount and proportion of various proteins in a sample.
  • Western-seq allows multiplex protein detection with size separation (FIG. 9E).
  • FIG. 9F provides an exemplary protocol for Western-seq.
  • FIG. 9G demonstrates more specific binding at expected section.
  • FIGS.10A-10B for illustration purposes only, provide two example designs for Ab-oligos and DNA-calipers with complementary anchor sequences, hook sequences, UMI, barcode, and splint sequences. (SEQ ID NOS: 10-13)
  • FIGS. 11A-11C FIGS. 11A-11C.
  • FIGS. 12A-12C Improvement of the process for converting the extended DNA-caliper into a sequencing library by replacing splint/whip-adaptor with designed removal of the antibody- oligo components.
  • Prod-seq was performed using either the approach that employs a splint (left 4 samples) or the one detailed here with the restriction enzyme site. The ratio of unique UMI pairs increased by up to ⁇ 20 fold.
  • c Identification of optimal reaction conditions using “mimic” reactions. Oligonucleotides that imitate the output of the reaction were used to evaluate multiple conditions and designs, such as different restriction enzymes (SbfI, NotI, BbsI – sticky; XmnI, blunt); different types of Klenow (exo+ or exo-) and single or nested PCR. Under these mimic conditions, single PCR, exo- and the blunt enzymes show the best ability to capture the correct structure. FIGS.13A-13C.
  • FIG. 14 illustrates a diagrammatic representation of a machine 1400 in the form of a computer system within which a set of instructions may be executed for causing the machine 1400 to perform any one or more of the methodologies discussed herein.
  • DESCRIPTION Reference will now be made in detail to certain embodiments of the disclosed subject matter. While the disclosed subject matter will be described in conjunction with the enumerated claims, it will be understood that the exemplified subject matter is not intended to limit the claims to the disclosed subject matter. The methods described herein provide a means to simultaneously identify interactions between cellular entities.
  • references in the specification to "one embodiment,” “an embodiment,” etc., indicate that the embodiment described may include a particular aspect, feature, structure, moiety, or characteristic, but not every embodiment necessarily includes that aspect, feature, structure, moiety, or characteristic. Moreover, such phrases may, but do not necessarily, refer to the same embodiment referred to in other portions of the specification. Further, when a particular aspect, feature, structure, moiety, or characteristic is described in connection with an embodiment, it is within the knowledge of one skilled in the art to affect or connect such aspect, feature, structure, moiety, or characteristic with other embodiments, whether or not explicitly described.
  • the singular forms “a,” “an,” and “the” include plural reference unless the context clearly dictates otherwise.
  • a reference to "a compound” includes a plurality of such compounds, so that a compound X includes a plurality of compounds X.
  • the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for the use of exclusive terminology, such as “solely,” “only,” and the like, in connection with any element described herein, and/or the recitation of claim elements or use of “negative” limitations.
  • the term “and/or” means any one of the items, any combination of the items, or all of the items with which this term is associated.
  • the phrase “one or more” is readily understood by one of skill in the art, particularly when read in context of its usage.
  • one or more substituents on a phenyl ring refers to one to five, or one to four, for example if the phenyl ring is di-substituted.
  • “or” should be understood to have the same meaning as “and/or” as defined above. For example, when separating a listing of items, “and/or” or “or” shall be interpreted as being inclusive, e.g., the inclusion of at least one, but also including more than one of a number of items, and, optionally, additional unlisted items.
  • “about 50” percent can in some embodiments carry a variation from 45 to 55 percent.
  • the term “about” can include one or two integers greater than and/or less than a recited integer at each end of the range. Unless indicated otherwise herein, the term “about” is intended to include values, e.g., weight percentages, proximate to the recited range that are equivalent in terms of the functionality of the individual ingredient, the composition, or the embodiment. The term about can also modify the endpoints of a recited range as discuss above in this paragraph.
  • a recited range includes each specific value, integer, decimal, or identity within the range. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, or tenths. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, etc. As will also be understood by one skilled in the art, all language such as “up to,” “at least,” “greater than,” “less than,” “more than,” “or more,” and the like, include the number recited and such terms refer to ranges that can be subsequently broken down into sub-ranges as discussed above.
  • provisos may apply to any of the disclosed categories or embodiments whereby any one or more of the recited elements, species, or embodiments, may be excluded from such categories or embodiments, for example, for use in an explicit negative limitation.
  • the term "contacting" refers to the act of touching, making contact, or of bringing to immediate or close proximity, including at the cellular or molecular level, for example, to bring about a physiological reaction, a chemical reaction, or a physical change, e.g., in a solution, in a reaction mixture, in vitro, or in vivo.
  • nucleic acid encompasses RNA as well as single and double stranded DNA and cDNA.
  • nucleic acid encompasses RNA as well as single and double stranded DNA and cDNA.
  • nucleic acid encompasses RNA as well as single and double stranded DNA and cDNA.
  • nucleic acid encompasses RNA as well as single and double stranded DNA and cDNA.
  • nucleic acid DNA
  • RNA and similar terms also include nucleic acid analogs, i.e., analogs having other than a phosphodiester backbone.
  • nucleic acid is meant any nucleic acid, whether composed of deoxyribonucleosides or ribonucleosides, and whether composed of phosphodiester linkages or modified linkages such as phosphotriester, phosphoramidate, siloxane, carbonate, carboxymethylester, acetamidate, carbamate, thioether, bridged phosphoramidate, bridged methylene phosphonate, bridged phosphoramidate, bridged phosphoramidate, bridged methylene phosphonate, phosphorothioate, methylphosphonate, phosphorodithioate, bridged phosphorothioate or sulfone linkages, and combinations of such linkages.
  • nucleic acid also specifically includes nucleic acids composed of bases other than the five biologically occurring bases (adenine, guanine, thymine, cytosine, and uracil).
  • bases other than the five biologically occurring bases
  • Conventional notation is used herein to describe polynucleotide sequences: the left-hand end of a single-stranded polynucleotide sequence is the 5’-end; the left-hand direction of a double-stranded polynucleotide sequence is referred to as the 5’-direction.
  • the direction of 5’ to 3’ addition of nucleotides to nascent RNA transcripts is referred to as the transcription direction.
  • oligonucleotide typically refers to short polynucleotides, generally, no greater than about 100 nucleotides.
  • carrier includes a liquid, such as water, saline solution or other pH balanced liquids, as well as encapsulating agents (e.g., liposomes), and other materials.
  • the tools (Ab-oligos, DNA-calipers etc.) provided herein can be formulated in dry form (e.g., in freeze-dried form), and resuspended with in a convenient liquid.
  • standard refers to something used for comparison.
  • Standard can be a known standard agent or compound which is administered and used for comparing results when administering a test compound, or it can be a standard parameter or function which is measured to obtain a control value when measuring an effect of an agent or compound on a parameter or function.
  • Standard can also refer to an “internal standard”, such as an agent or compound which is added at known amounts to a sample and is useful in determining such things as purification or recovery rates when a sample is processed or subjected to purification or extraction procedures before a marker of interest is measured.
  • Internal standards are often a purified marker of interest which has been labeled, such as with a radioactive isotope, allowing it to be distinguished from an endogenous marker.
  • click reaction is recognized in the art, which describe a collection of reliable and self-directed organic reactions, such as the most recognized copper catalyzed azide-alkyne [3+2] cycloaddition.
  • Non-limiting examples of click chemistry reactions can be found, for example, in H. C. Kolb, M. G. Finn, K. B. Sharpless, Angew. Chem. Int. Ed. 2001, 40, 2004 and E. M. Sletten, C. R. Bertozzi, Angew. Chem. Int. Ed. 2009, 48, 6974, the disclosures of which are herein incorporated by reference in their entireties for all purposes.
  • cross-linking agent is used to describe a compound that is capable of forming a chemical bond between molecular groups on similar or dissimilar molecules so as to covalently bond together the molecules.
  • Examples of common cross-linking agents are known in the art. See, for example, Bioconjugate Techniques (Academic Press, New York, 1996 or later versions) the content of which is herein incorporated by reference in its entirety for all purposes. Indirect attachment of the biomolecule to polymer dots can occur through the use of “linker” molecule, for example, avidin, streptavidin, neutravidin, biotin or a like molecule.
  • a "cell” refers to any type of cell isolated from a prokaryotic, eukaryotic, or archaeon organism, including bacteria, archaea, fungi, protists, plants, and animals, including cells from tissues, organs, and biopsies, as well as recombinant cells, cells from cell lines cultured in vitro, and cellular fragments, cell components, or organelles comprising nucleic acids.
  • the term also encompasses artificial cells, such as nanoparticles, liposomes, polymersomes, or microcapsules encapsulating nucleic acids.
  • the methods described herein can be performed, for example, on a sample comprising a single cell or a population of cells.
  • the term also includes genetically modified cells.
  • biological samples can comprise any commercially available cells, such as HeLa, HEK293, or stem cells, and/or lysates of such cells.
  • hybridize and “hybridization” refer to the formation of complexes between nucleotide sequences which are sufficiently complementary to form complexes via Watson Crick base pairing.
  • label or a detectable label intends a directly or indirectly detectable compound or composition that is conjugated directly or indirectly to the composition to be detected, e.g., N-terminal histidine tags (N-His), magnetically active isotopes, e.g., 115 Sn, 117 Sn and 119 Sn, a non-radioactive isotopes such as 13 C and 15 N, polynucleotide or protein such as an antibody so as to generate a “labeled” composition.
  • N-terminal histidine tags N-His
  • magnetically active isotopes e.g., 115 Sn, 117 Sn and 119 Sn
  • a non-radioactive isotopes such as 13 C and 15 N
  • polynucleotide or protein such as an antibody so as to generate a “labeled” composition.
  • the term also includes sequences conjugated to the polynucleotide that will provide a signal upon expression of the inserted sequence
  • the label may be detectable by itself (e.g., radioisotope labels or fluorescent labels) or, in the case of an enzymatic label, may catalyze chemical alteration of a substrate compound or composition which is detectable.
  • the labels can be suitable for small scale detection or more suitable for high-throughput screening.
  • suitable labels include, but are not limited to magnetically active isotopes, non-radioactive isotopes, radioisotopes, fluorochromes, chemiluminescent compounds, dyes, and proteins, including enzymes.
  • the label may be simply detected, or it may be quantified.
  • a response that is simply detected generally comprises a response whose existence merely is confirmed
  • a response that is quantified generally comprises a response having a quantifiable (e.g., numerically reportable) value such as an intensity, polarization, and/or other property.
  • the detectable response may be generated directly using a luminophore or fluorophore associated with an assay component actually involved in binding, or indirectly using a luminophore or fluorophore associated with another (e.g., reporter or indicator) component.
  • luminescent labels that produce signals include but are not limited to bioluminescence and chemiluminescence.
  • Detectable luminescence response generally comprises a change in, or an occurrence of a luminescence signal.
  • Suitable methods and luminophores for luminescently labeling assay components are known in the art and described for example in Haugland, Richard P. (1996) Handbook of Fluorescent Probes and Research Chemicals (6th ed).
  • Examples of luminescent probes include, but are not limited to, aequorin and luciferases.
  • fluorescent labels include, but are not limited to, fluorescein, rhodamine, tetramethylrhodamine, eosin, erythrosin, coumarin, methyl-coumarins, pyrene, Malacite green, stilbene, Lucifer Yellow, Cascade BlueTM, and Texas Red.
  • suitable optical dyes are described in the Haugland, Richard P. (1996) Handbook of Fluorescent Probes and Research Chemicals (6th ed.).
  • a purification label or maker refers to a label that may be used in purifying the molecule or component that the label is conjugated to, such as an epitope tag (including but not limited to a Myc tag, a human influenza hemagglutinin (HA) tag, a FLAG tag), an affinity tag (including but not limited to a glutathione-S transferase (GST), a poly-Histidine (His) tag, Calmodulin Binding Protein (CBP), or Maltose-binding protein (MBP)), or a fluorescent tag.
  • an epitope tag including but not limited to a Myc tag, a human influenza hemagglutinin (HA) tag, a FLAG tag
  • an affinity tag including but not limited to a glutathione-S transferase (GST), a poly-Histidine (His) tag, Calmodulin Binding Protein (CBP), or Maltose-binding protein (MBP)
  • GST glutathione-
  • Prod-seq, WhIP-seq and Western-seq Although the activity of many dynamic protein complexes is highly dependent on interactions between molecules (e.g., protein ⁇ protein and protein ⁇ DNA), due to the limitations of current methods, most of the existing studies are focused on a single molecule at a time.
  • One goal of the methods described herein is to develop an innovative cross ⁇ disciplinary framework for studying these proteins as a class and use mouse embryonic stem cells, as an example, to study how multiple factors work together to regulate gene activity.
  • the Prod-seq process proceeds as follows (FIG.2A): (i) Fixed and lysed cells are incubated with multiple antibodies, where each specific antibody is covalently linked to a single-stranded oligonucleotide tail (Ab-oligos).
  • the oligonucleotide tail contains an anchor sequence (Anc) complementary to both arms of the DNA-caliper (Anc’; the DNA-caliper is depicted in FIG.2A); a unique molecular identifier (UMI; a short random sequence) and an antibody-specific unique barcode (Barcode n, m, etc.).
  • Each arm of the DNA-caliper contains at each 3’ end a sequence that is complementary to the antibody anchors at their 3’ ends (Anc’).
  • Target sequences are captured by hybridization of the two DNA-caliper anchors to the complementary anchor sequences on the oligonucleotide linked to the antibodies, followed by DNA polymerase extension at both ends (FIG.2A). This extension covalently links the barcodes and UMI sequences of the two antibodies.
  • the extended DNA-calipers are then denatured and captured by, for example, chemical interactions with modifications on the DNA, including but not limited to, biotin or digoxigenin, pull-down, and then converted into sequencing-ready constructs (FIG. 2C).
  • FIG.6 offers a variation on Prod-seq in which a. Fixed and lysed cells are incubated with multiple Ab-oligos.
  • the oligonucleotide tail contains: an anchor sequence (Anchor 1) complementary to both arms of the DNA-caliper; a UMI; an antibody-specific unique barcode (BC n, m, etc.), a restriction enzyme site, and a Hook region.
  • An anchor 1 complementary to both arms of the DNA-caliper
  • UMI an antibody-specific unique barcode
  • BC n, m, etc. an antibody-specific unique barcode
  • the ends of each oligo are non-extendable (labeled by a T block).
  • the DNA-caliper is depicted at the bottom and includes biotinylated arms (circled B) and two identical Anchor 1’. Target proteins are removed from next steps for simplicity.
  • the barcodes and UMIs are captured by hybridization of the Anchor 1’ (Anc1’) on the DNA- caliper arms to the Anchor 1 (Anc1) on each Ab-oligo, followed by DNA polymerase extension (dashed lines) at both ends.
  • c-d Next, the extended DNA-calipers are denatured and pulled-down using the biotin. A Splint molecule with two overhanging Hook’ regions (c, top) is hybridized and ligated to the extended DNA-caliper.
  • each of the DNA-caliper arms is extended using the other arm as a template.
  • the molecule is linearized by USER enzyme that cleaves the dU and the fragments are used as a template for generation of sequencing-ready constructs by PCR from the primer sequences (P1 and P2; The dU and the primers are embedded in the original DNA-caliper, panel a.).
  • the diverse DNA barcodes that are captured by the process are depicted.
  • Prod-seq and WhIP-seq use PCR to convert and amplify the product into, for example, a next generation sequencing (NGS) library.
  • NGS next generation sequencing
  • the novel design disclosed herein incorporates a cleavage site (such as restriction enzyme site, dU or a photo-cleavable site) that separates the two 5’ ends following, for example, the Prod-seq reaction, which enables the release of the two strands while maintaining the proximity detection information intact.
  • a cleavage site such as restriction enzyme site, dU or a photo-cleavable site
  • This approach is highly effective in increasing PCR efficiency and yield of the reaction to generate NGS libraries (FIGS.11A-11B).
  • the restriction enzyme e.g., of barcodes or UMIs
  • an alternative design of the DNA- caliper was made that employs a UV-light sensitive modification.
  • Exposure to UV-light replaces the replaces the need for a restriction enzyme site at the cleavage site (FIG.11C).
  • the process included two rounds of PCR – the first introducing the Illumina adapter and the second adding Illumina indices.
  • the design was modified to reduce the number of PCR reactions, and the primer site of the DNA-caliper is now directly compatible with the Illumina sequencing primers. This further simplifies the library preparation process (only one round of PCR needed), reduces efforts and improves the efficiency as the added PCR was prone to lead to loss of material and add noise.
  • the antibodies are removed through use of a restriction enzyme and the restriction enzyme site on the oligonucleotide tail.
  • a “splint” molecule (or a “whip-adaptor” as referenced in US Pat. No. 10,655,162, incorporated herein by reference), a partially double-stranded DNA, was used to facilitate the cyclization of the extended DNA caliper, which enables joining the two molecule barcodes within the same DNA sequence by a ligation reaction.
  • the annealing of a molecule such as the splint or whip-adaptor is extremely inefficient, potentially the reaction may be outcompeted by the reannealing of the DNA-caliper back to the Ab-oligos.
  • a novel design was made that introduces a restriction enzyme recognition site (such as for XmnI; see example sequences below) at the sequence of the oligonucleotide that is conjugated to the antibodies.
  • the extended DNA-caliper is simply digested with, for example, XmnI, the ends are then directly be ligated together (intra-molecular) without the need of splint or whip-adaptor annealing.
  • This approach was shown to significantly improve the efficiency of the workflow and the quality of the Prod-seq products (FIGS. 12-A-12C).
  • the WhIP-seq method is described schematically in FIGS.2B and 7. Sheared chromatin is incubated with multiple Ab-oligos.
  • the Ab-oligos can be the same ones used for Prod-seq (FIG. 2A), as the oligonucleotide tail has an identical structure (Anchor 1, a UMI, a barcode (BC m, k...n) and a Hook’).
  • the chromatin fragments (Genomic Region) are ligated to adaptors (WhIP- adaptors) that include a UMI and a different anchor (Anchor 2).
  • Each arm of the DNA-caliper contains a specific primer sequence, a capture molecule (such as biotin (B in a circle)), a cleavage site and, at each 3’ end, sequences complementary to the Anchors it targets.
  • a capture molecule such as biotin (B in a circle)
  • Target sequences are captured by hybridization of the two DNA-caliper anchors (Anc1’ and Anc2’) to the Ab-oligos and the adaptor 3’ overhang, followed by DNA extension (dashed lines) at both ends.
  • c. The extended DNA-calipers are denatured and pulled-down using the Biotin moiety.
  • d-f Conversion of the extended DNA-calipers into sequencing-ready constructs via intramolecular hybridization- extension.
  • a Splint molecule (left), with 5’ phosphorylated termini is added. Note, this Splint has one arm with Hook’ to match the Ab-oligo, while the other has a Poly(G) overhang at the 3’ end. Both termini end with cleavable extension-blockers.
  • the other end of the extended DNA- caliper is 3’ modified by terminal transferase to have a Poly(C) stretch.
  • the cleavable blockers are removed, and the Splint is ligated to the extended DNA-caliper.
  • the free 3’ ends of the DNA-caliper arms are extended using the other arm as a template followed by linearization of the molecule by USER enzyme to cleave the dU and PCR using the P1 and P3.
  • primers f. Diverse DNA barcodes are depicted; the genomic regions (Genomic) are represented. Western-seq is described schematically in FIG.10.
  • the methods described comprise a proximity detection (“Prod-seq” hereafter) method comprising contacting or incubating biological samples containing protein and nucleic acids (e.g., cell lysate) with antibody-oligonucleotide conjugates (Ab-oligos hereafter) to detect PPIs and provide detailed measurements of protein abundance.
  • the antibodies in the Ab-oligos are modified to comprise an anchor sequence (Anc) that binds to a DNA-caliper oligonucleotide that serves as a proximity detector by capturing interacting proteins bound to the Ab-oligos on a single molecule.
  • the oligonucleotide tail sequence can comprise about 10-1000 nucleotides.
  • a oligonucleotide tail sequence may have a length of 10-900, 10-800, 10-700, 10-600, 10-500, 10- 400, 10-300, 10-200, 10-100, 10-50 or 10-25 nucleotides.
  • an oligonucleotide tail sequence has a length of 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 225, 250, 275, 300, 350, 400, 450 or 500 nucleotides.
  • an oligonucleotide tail sequence is longer than 1000 nucleotides.
  • the length of an oligonucleotide tail sequence is 10 3 , 10 4 , 10 5 or 10 6 nucleotides.
  • Ab-oligos can be synthesized by conjugating single strand oligonucleotides to antibodies, followed by removal of the non-linked components. The rapid kinetics of the iEDDA click chemistry reaction between tetrazine and trans-cyclooctene (TCO) was used as reported previously in van Buggenum, J., Gerlach, J., Eising, S. et al. A covalent and cleavable antibody-DNA conjugation strategy for sensitive protein detection via immuno-PCR. Sci Rep 6, 22675 (2016), which is incorporated herein by reference.
  • the oligo of the Ab-oligo comprises a barcode (BC n, m, etc., which can be specific to the antibody; further discussed below), a unique molecular identifier (UMI, which can be a randomly generated sequence; further discussed below), one or more restriction enzyme sites and an anchor region (Anchor1).
  • the length of an anchor region may vary.
  • the anchor can be located near the 3’ end of the tail. In some embodiments, an anchor region has a length of 5-50 nucleotides.
  • an anchor region may have a length of 5-45, 5-40, 5-35, 5-30, 5-25, 5-20, 5-15, 5-10, 10-50, 10-45, 10-40, 10-35, 10-30, 10-25, 10-20, 10-15, 15-50, 15-45, 15-40, 15-35, 15-30, 15-25, 15-20, 20-50, 20-45, 20-40, 20-35, 20-30, 20-25, 25-40, 25-35, 25-30, 30-40, 30-35 or 35-40 nucleotides.
  • an anchor region has a length of 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50 nucleotides.
  • an anchor region has a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleotides.
  • An anchor region in some embodiments, is longer than 50 nucleotides, or shorter than 5 nucleotides.
  • Any restriction enzyme site can be used, including, but not limited to, those which 5’ or 3’ overhangs (sticky ends) or blunt ends. These enzymes include Type I, Type II, Type III, Type 4 and Type V enzymes.
  • restriction enzymes include, but are not limited to, EcroRI, EcoRII, BamHI, HindIII, TaqI, NotI, HinFI, Sau3Ai, PvuII, SmaI, HaeII, HgaI, AluI, EcorRV, EcoP15I, KpnI, PstI, SacI, SalI, ScaI, SpeI, SphI, StuI and XbaI.
  • the oligonucleotide tail comprises a hook.
  • the length of a hook region may vary. The hook can be located near the 3’ or 5’ end of the tail.
  • a hook region has a length of 5-50 nucleotides.
  • a hook region may have a length of 5-45, 5-40, 5-35, 5-30, 5-25, 5-20, 5-15, 5-10, 10-50, 10-45, 10-40, 10-35, 10-30, 10-25, 10-20, 10-15, 15-50, 15-45, 15-40, 15-35, 15-30, 15-25, 15-20, 20-50, 20-45, 20-40, 20-35, 20-30, 20-25, 25-40, 25-35, 25-30, 30-40, 30-35 or 35-40 nucleotides.
  • a hook region has a length of 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50 nucleotides.
  • an anchor region has a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleotides.
  • a hook region in some embodiments, is longer than 50 nucleotides, or shorter than 5 nucleotides.
  • the oligonucleotide tail comprises one or more primer domains.
  • a primer domain is a domain to which a primer binds.
  • a primer is a strand of short nucleotide sequence that serves as a starting point for nucleic acid (e.g., DNA) synthesis.
  • an oligonucleotide tail can comprise an internal primer domain (e.g., near the 5′ or 3’ end), which may be used for amplification.
  • the length of a primer domain may vary.
  • a primer domain has a length of 5-50 nucleotides.
  • a primer domain may have a length of 5-45, 5-40, 5-35, 5-30, 5-25, 5-20, 5-15, 5-10, 10-50, 10-45, 10-40, 10-35, 10-30, 10-25, 10-20, 10-15, 15-50, 15-45, 15-40, 15-35, 15-30, 15-25, 15-20, 20-50, 20-45, 20-40, 20-35, 20-30, 20-25, 25-40, 25-35, 25-30, 30-40, 30-35 or 35-40 nucleotides.
  • a primer domain has a length of 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50 nucleotides.
  • a primer domain has a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleotides.
  • a primer domain in some embodiments, is longer than 50 nucleotides, or shorter than 5 nucleotides.
  • DNA-Caliper A DNA-caliper comprises single-stranded oligonucleotide with two free 3' ends that allows bidirectional priming and thus serves as a proximity detector.
  • the DNA-caliper can be used to capture proximity between protein-protein, protein-DNA and also other entities such as protein- RNA, DNA-RNA, protein-metabolite, or cell-cell interactions, and thus for instance improve the characterization of in vitro differentiated cells.
  • a DNA-caliper a single-stranded nucleic acid comprising two 3′ ends.
  • a nucleic acid (a polymer of nucleotides) is “single-stranded” if nucleotides that form the nucleic acid are unpaired.
  • nucleotides of a single-stranded nucleic acid are not base paired (via Watson-Crick base pairs, e.g., guanine-cytosine and adenine-thymine/uracil) to nucleotides of another nucleic acid.
  • a single-stranded nucleic acid may be contrasted with a double-stranded (paired) nucleic acid, a typical example of which is a DNA double helix.
  • Single- stranded nucleic acids may include a contiguous (uninterrupted) sequence of nucleotides or, in some embodiments, a single-stranded nucleic acid may be a conjugate that includes two nucleic acid strands joined together, for example, through a chemical (covalent) linkage.
  • a single strand of a nucleic acid e.g., DNA or RNA
  • the 5′ end typically contains a phosphate group attached to the 5′ carbon of the ribose ring of a nucleotide and a 3′ end, which is unmodified from the ribose - OH substituent.
  • Nucleic acids are synthesized in vivo in the 5′ to 3′ direction. Polymerase relies on the energy produced by breaking nucleoside triphosphate bonds to attach new nucleoside monophosphates to the 3′-hydroxyl (—OH) group, via a phosphodiester bond.
  • An engineered single-stranded nucleic acid of the present disclosure has two 3′ ends (a whip molecule). Each terminus of the single-stranded nucleic acid includes a 3′-hydroxyl (—OH) group.
  • a single-stranded DNA-caliper is formed by joining (linking) the 5′ end of one single-stranded nucleic acid to the 5′ end of another single-stranded nucleic acid.
  • the linkage between two 5′ ends is a covalent linkage. In other embodiments, the linkage is noncovalent.
  • the 5′ ends of single-stranded nucleic acids may be linked to each other using any means in the art for linking nucleic acids to each other.
  • nucleic acids are linked together using ‘click chemistry’ (see, e.g., V. V. Rostovtsev et al. Angew. Chem. Int. Ed., 2002, 41, 2596-2599; F. Himo et al. J. Am. Chem. Soc., 2005, 127, 210-216; and B. C. Boren et al. J. Am. Chem.
  • An example of a click chemistry reaction is the Huisgen 1,3- dipolar cycloaddition of alkynes to azides to form 1,4-disubsituted-1,2,3-triazoles.
  • the copper(I)- catalyzed reaction is mild and very efficient, requiring no protecting groups, and requiring no purification, in many cases.
  • the azide and alkyne functional groups are largely inert towards biological molecules and aqueous environments, which allows the use of the Huisgen 1,3-dipolar cycloaddition in target-guided synthesis and activity-based protein profiling.
  • a whip molecule is formed by linking the 5′ end of one nucleic acid strand that includes an azide group to the 5′ end of another nucleic acid strand that includes an alkyne group.
  • Other linkage reactions are encompassed by the present disclosure and are known in the art.
  • Each arm of the caliper can be about 10 to about 1000 nucleotides.
  • a DNA- caliper arm may have a length of 10-900, 10-800, 10-700, 10-600, 10-500, 10-400, 10-300, 10- 200, 10-100, 10-50 or 10-25 nucleotides.
  • a DNA-caliper arm has a length of 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 225, 250, 275, 300, 350, 400, 450 or 500 nucleotides.
  • a DNA-caliper arm is longer than 1000 nucleotides.
  • the length of a DNA-caliper arm is 10 3 , 10 4 , 10 5 or 10 6 nucleotides.
  • Single-stranded nucleic acid DNA-calipers typically include an “anchor domain” at each 3′ end.
  • a “domain” refers to a discrete, contiguous sequence of nucleotides or nucleotide base pairs, depending on whether the domain is unpaired (single-stranded nucleotides) or paired (double-stranded nucleotide base pairs), respectively.
  • Complementary anchor domains bind to each other.
  • a domain is “complementary to” another domain if one domain contains nucleotides that base pair (hybridize/bind through Watson-Crick nucleotide base pairing) with nucleotides of the other domain such that the two domains form a paired (double-stranded) or partially paired molecular species/structure.
  • Complementary domains need not be perfectly (100%) complementary to form a paired structure, although perfect complementarity is provided, in some embodiments.
  • an anchor domain of a DNA-caliper that is complementary to an anchor domain of a barcoded nucleic acid (described below) binds to that domain, for example, for a time sufficient to initiate polymerization in the presence of polymerase and under conditions appropriate for polymerization.
  • the length of an anchor domain may vary. In some embodiments, an anchor domain has a length of 5-50 nucleotides.
  • an anchor domain may have a length of 5-45, 5-40, 5-35, 5-30, 5-25, 5-20, 5-15, 5-10, 10-50, 10-45, 10-40, 10-35, 10-30, 10-25, 10-20, 10-15, 15-50, 15- 45, 15-40, 15-35, 15-30, 15-25, 15-20, 20-50, 20-45, 20-40, 20-35, 20-30, 20-25, 25-40, 25-35, 25-30, 30-40, 30-35 or 35-40 nucleotides.
  • an anchor domain has a length of 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50 nucleotides.
  • an anchor domain has a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleotides.
  • An anchor domain in some embodiments, is longer than 50 nucleotides, or shorter than 5 nucleotides.
  • the DNA-caliper comprises one or more hooks (e.g., 2, 4, 6 etc). The length of a hook region may vary. The hook can be located near the 3’ or 5’ end of the caliper arm. In some embodiments, a hook region has a length of 5-50 nucleotides.
  • a hook region may have a length of 5-45, 5-40, 5-35, 5-30, 5-25, 5-20, 5-15, 5-10, 10-50, 10-45, 10-40, 10-35, 10-30, 10-25, 10-20, 10-15, 15-50, 15-45, 15-40, 15-35, 15-30, 15-25, 15-20, 20-50, 20-45, 20-40, 20-35, 20-30, 20-25, 25-40, 25-35, 25-30, 30-40, 30-35 or 35-40 nucleotides.
  • a hook region has a length of 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50 nucleotides.
  • an anchor region has a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleotides.
  • a hook region in some embodiments, is longer than 50 nucleotides, or shorter than 5 nucleotides.
  • a single-stranded nucleic acid DNA-caliper comprises primer domains.
  • a primer domain is a domain to which a primer binds.
  • a primer is a strand of short nucleotide sequence that serves as a starting point for nucleic acid (e.g., DNA) synthesis.
  • DNA-calipers comprise a pair of internal primer domains (e.g., near the linked 5′ ends), which may be used for amplification of sequence-ready barcoded constructs produced using the methods of the present disclosure.
  • the length of a primer domain may vary.
  • a primer domain has a length of 5-50 nucleotides.
  • a primer domain may have a length of 5-45, 5-40, 5-35, 5-30, 5-25, 5-20, 5-15, 5-10, 10-50, 10-45, 10-40, 10-35, 10-30, 10-25, 10-20, 10-15, 15-50, 15-45, 15-40, 15-35, 15-30, 15-25, 15-20, 20-50, 20-45, 20-40, 20-35, 20-30, 20-25, 25-40, 25-35, 25-30, 30-40, 30-35 or 35-40 nucleotides.
  • a primer domain has a length of 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50 nucleotides.
  • the cleavage site functions to cleave the molecule in two; separates the two 5’ ends of the DNA-caliper following, for example, the Prod-seq reaction, which enables the release of the two strands while maintaining the proximity detection information intact.
  • the cleavage site can be any restriction enzyme site, deoxy uridine (dU; with, for example, a uracil-DNA glycosylase, which are evolutionarily conserved DNA repair enzymes that initiate the base excision repair pathway and remove uracil from DNA) or a photo-cleavable site (such as a site that separates the molecule upon exposure to UV light, for example, with a photo-cleavable modification that contains a photolabile functional group that is cleavable by UV light of specific wavelength (e.g., 300-350 nm); a photo-cleavable spacer can be purchased from Integrated DNA Technologies (IDT) called IDT-PC and can be placed between DNA bases or between an oligo and a terminal modification).
  • IDT Integrated DNA Technologies
  • An example of a photo-cleavable spacer Any restriction not limited to, those which 5’ or 3’ overhangs (sticky ends) or blunt ends.
  • These enzymes include Type I, Type II, Type III, Type 4 and Type V enzymes.
  • Some examples of restriction enzymes include, but are not limited to, EcroRI, EcoRII, BamHI, HindIII, TaqI, NotI, HinFI, Sau3Ai, PvuII, SmaI, HaeII, HgaI, AluI, EcorRV, EcoP15I, KpnI, PstI, SacI, SalI, ScaI, SpeI, SphI, StuI and XbaI.
  • the PPI bound Ab-oligo sequences are captured by hybridization of the two DNA-caliper anchors to the complementary anchor sequences on the Ab-oligo (Anc), followed by DNA polymerase extension at both ends. This extension covalently links the barcodes and UMI sequences of the two antibodies.
  • the extended DNA-calipers are then denatured and captured by biotin pull-down and then converted into sequencing-ready constructs.
  • the DNA-calipers can also comprise one or more capture molecules (to aid in the isolation of the DNA-calipers, e.g., extended DNA calipers). Further, an extended DNA ⁇ caliper can be used as a template for DNA synthesis regardless of the directionality inversion in mid ⁇ template.
  • UMI and barcodes The use of unique molecular identifier (e.g., a short random DNA sequence) labeled DNA ⁇ barcoded antibodies allows the simultaneous interrogation of multiple proteins from a single sample. For example, two proteins that interact with each other can be captured by using synthetic DNA (oligonucleotides) to barcode each and every target protein (using modified antibodies). The DNA-Caliper can capture the barcodes (that represent proteins) that are near each other in the cell and allows hybridization and extension at both ends. Such extension results in the barcodes of the two interacting proteins in one DNA molecule. In embodiments, interactions of multiple proteins can be captured in a single reaction. DNA sequencing and the cellular data can be analyzed to report all of the interactions and the expression levels of the protein in each condition.
  • DNA sequencing and the cellular data can be analyzed to report all of the interactions and the expression levels of the protein in each condition.
  • the DNA-calipers and/or oligonucleotide tail also comprise a barcode.
  • a “barcoded nucleic acid” is a nucleic acid, typically single-stranded, that includes a barcode domain.
  • a “barcode domain” is a domain that includes a nucleotide sequence that can be used to identify the barcoded nucleic acid or to identify a biomolecule(s) to which the barcoded nucleic acid is linked (directly or indirectly linked).
  • a barcoded nucleic acid may include a barcode domain that is unique to that single nucleic acid (among a population of barcoded nucleic acids, the barcode is specific to that one nucleic acid) or a barcode domain that is unique to a subpopulation of nucleic acids (among multiple populations of barcoded nucleic acids, the barcode is specific to a single subpopulation of barcoded nucleic acids).
  • the length of a barcode domain may vary.
  • a barcode domain may have a length of 5-45, 5-40, 5-35, 5-30, 5-25, 5- 20, 5-15, 5-10, 10-50, 10-45, 10-40, 10-35, 10-30, 10-25, 10-20, 10-15, 15-50, 15-45, 15-40, 15- 35, 15-30, 15-25, 15-20, 20-50, 20-45, 20-40, 20-35, 20-30, 20-25, 25-40, 25-35, 25-30, 30-40, 30-35 or 35-40 nucleotides.
  • a barcode domain has a length of 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50 nucleotides.
  • a barcode domain has a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleotides.
  • a barcode domain in some embodiments, is longer than 50 nucleotides, or shorter than 5 nucleotides.
  • the length of a barcoded nucleic acid itself may vary. In some embodiments, the length of a barcoded nucleic acid is 20-1000 nucleotides.
  • a barcoded nucleic acid may have a length of 20-900, 20-800, 20-700, 20-600, 20-500, 20-400, 20-300, 20-200, 20-100, 20-50 or 20- 25 nucleotides.
  • a barcoded nucleic acid has a length of 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 225, 250, 275, 300, 350, 400, 450 or 500 nucleotides.
  • a barcoded nucleic acid is longer than 1000 nucleotides.
  • a barcoded nucleic acid in some embodiments, further includes a primer domain and an anchor domain that is complementary to one of the anchor domains of, for example, a DNA caliper.
  • barcoded nucleic acid can comprise a primer domain at its 5′ end, an internal (central) barcode domain and an anchor domain at its 3′ end. It should be understood that the barcoded domain need not be centrally located.
  • the 5′ end of the barcoded nucleic acid can be linked to an antibody that recognizes a protein of interest (e.g., a protein associated with genomic DNA).
  • a barcoded nucleic acid may be linked to any biomolecule.
  • the 3′ end of the barcoded nucleic acid can include an anchor domain that that is complementary to one 3′ end of a DNA- caliper such that the two anchor domains bind to each other to form a paired domain.
  • Anchor domains in some embodiments, are used for localizing a target biomolecule(s) of interest. When co-localizing two biomolecules, one of the biomolecules contains an anchor domain complementary to one of the anchor domains of a DNA caliper, and the other biomolecule contains an anchor domain complementary to the other of the anchor domains of a DNA caliper.
  • the barcoded nucleic acid contains an anchor domain complementary to one anchor domain of the DNA caliper, and another biomolecule contains an anchor domain complementary to the other anchor domain of the DNA caliper and has a unique molecular identifier (UMI).
  • the DNA caliper or single stranded oligonucleotide tail comprises a unique molecular identifier (UMI) or other nucleotide sequence or moiety specific to that adaptor molecule, which makes distinct each adaptor molecule in a population (see, e.g., Kivioja, T. et al. Nat Methods 2012, 9, 72-4).
  • the UMI can be about 5-45, 5-40, 5-35, 5-30, 5-25, 5-20, 5-15, 5- 10, 10-50, 10-45, 10-40, 10-35, 10-30, 10-25, 10-20, 10-15, 15-50, 15-45, 15-40, 15-35, 15-30, 15-25, 15-20, 20-50, 20-45, 20-40, 20-35, 20-30, 20-25, 25-40, 25-35, 25-30, 30-40, 30-35 or 35- 40 nucleotides.
  • a UMI has a length of 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50 nucleotides.
  • a UMI has a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleotides.
  • a UMI in some embodiments, is longer than 50 nucleotides, or shorter than 5 nucleotides.
  • Adaptor molecules are used, in some embodiments, to add an anchor domain or other nucleotide domain to another nucleic acid molecule. For example, adaptor molecules can be added at each end of a fragmented piece of chromatin. Each adaptor molecule can include an unpaired 3′ overhang, which includes an anchor domain, Anchor2.
  • an adaptor is single-stranded, double-stranded, or partially double-stranded (partially single-stranded), depends on the target biomolecule to which the adaptor is being added.
  • a double-stranded, or partially, double- stranded adaptor is added to the terminus (or termini) of a double-stranded target nucleic acid, and a single-stranded adaptor is added to the terminus of a single-stranded target nucleic acid.
  • an adaptor molecule can be used, for example, to add a double-stranded domain and a single-stranded homopolymer overhang domain (or other single-stranded nucleotide overhang domain) to a 3′ end of an extended DNA-caliper (e.g., following hybridization to a barcoded nucleic acid, polymerization through the barcoded domain, and dissociation of the resulting partially double-stranded molecule).
  • This adaptor facilitates joining of the two 3′ ends of an extended whip molecule to form a circular double-stranded molecule, which may then be isolated, linearized and sequenced, as discussed below.
  • Design 1 Ab-oligo and DNA-caliper (FIG.10A)
  • the oligo can comprise the following nucleic acid sequence (SEQ ID NO: 1): TACAACTCTTGTATCTACACGCCACGTCGTGGCGANNNNNNNNNNNNNttacaaccag actgatactacag/3AmMO/
  • SEQ ID NO: 1 A DNA-caliper with complementary anchor sequences to SEQ ID NO: 1 is shown in the nucleic acid sequences shown below: DNA-caliper-armA (SEQ ID NO: 2): /5AzideN/ATTAGGTTGCTAGCTCCTGCAGTCcagtctggttgtaa DNA-caliper-armB (SEQ ID NO: 3): /5Hexynyl/ATA
  • a biological sample is isolated from a patient or a test subject, such a mammalian patient/subject, including human, livestock (e.g., cow, horse, goat, pig), laboratory (e.g., mice, rate, rabbit) or companion (e.g., dog, cat) patients/animals.
  • livestock e.g., cow, horse, goat, pig
  • laboratory e.g., mice, rate, rabbit
  • companion e.g., dog, cat
  • the definition also includes samples that have been manipulated in any way after their procurement, such as by treatment with reagents, washed, or enrichment for certain cell populations.
  • the definition also includes samples that have been enriched for particular types of molecules, e.g., DNA.
  • sample encompasses biological samples such as a clinical sample such as saliva, sputum, mucus, nasopharyngeal samples, blood, plasma, serum, aspirate, cerebral spinal fluid (CSF), Stem Cells
  • CSF cerebral spinal fluid
  • hiPSCs The suite of tools was established using hiPSCs and hiPSC-derived neuronal cells to identify optimal conditions for these cells, with the goal of using the novel approach to study cells originating from individuals from diverse genetic backgrounds (e.g., iPSCORE22). Further, the ability to model disease conditions by using hiPSC-derived organoids (32) allows for systematic studies. These can include the biological changes happening during the differentiation (multiple time points), a survey of perturbations (e.g., CRISPR libraries) or therapeutics (FIG.1). hiPSCs can be differentiated to NPCs and neurons, which will allow one to monitor changes in PPIs and protein-DNA interactions during in vitro neurodifferentiation.
  • hiPSCs human induced pluripotent stem cells
  • iPSCORE a collection of hiPSC lines.
  • the iPSCORE includes 222 hiPSCs from individuals from diverse genetic background, age and gender (22). For example, one can focus on two hiPSC lines, from a male and a female of African American ethnicity.
  • hiPSCs can be cultured and fixed in the lab, and used to evaluate conditions, for example, for protein abundances and studies can be extended to survey high number of hiPSC and hiPSC- derived cells or organoids.
  • the following Examples illustrate some of the materials, methods, and experiments that were used or performed in the development of the invention.
  • EXAMPLES Example I - Prod-seq and WhIP-seq Introduction
  • Dynamic protein complexes and transient protein-protein interactions (PPI) are integral in numerous normal and abnormal, such as cancer, associated processes, such as cellular metabolism, signal transduction networks and regulation of chromatin structure. In cancer, this complex organization is involved in processes such as tumor initiation, progression and metastasis.
  • stem cells properties such as pluripotency, response to differentiation signals and chromatin regulation (4), are dependent on protein-protein interactions (PPIs) and the genomic localization of DNA-associated proteins. Yet, there is limited data about the complexity of such interactions due to the inability of available assays to simultaneously detect multiple interactions. Further, as most of the current studies focus on a single or a handful of factors, they do not address the multi-factorial manner in which complexes affect the physiology of health or disease. This gap presents a roadblock in advancing knowledge, for example, of stem cell biology and improving, for example, regenerative medicine, including the study of the impact of genetic mutations on stem cell properties and how aberrations in proteins lead to disease and drug development (5, 6).
  • PPIs proteins
  • RNA nucleic acid sequences
  • the methods and tools allow high-throughput queries of the dynamic organization of molecular entities and shifts the focus from single interactions towards an unprecedented, multi-dimensional interrogation of the cellular complexity.
  • the molecular interactions captured by the assays enables both evaluation of current hypotheses and generation of novel ones with reagent costs and experimental efforts comparable to existing genomic methods, such as ChIP-seq.
  • the Polycomb Group (PcG) complex in human induced pluripotent stem cells (hiPSCs) was studied in the pluripotent cells, as well as during their differentiation.
  • PcG Polycomb Repressive Complex 1 and 2
  • PRC1 and PRC2 Polycomb Repressive Complex 1 and 2
  • PRC1 deposits H2AK119ub1
  • PRC2 adds methyl groups to histone H3 to form H3K27me1-3. Both histone modifications are associated with developmental repression of gene activity.
  • PRC1 and PRC2 were shown to form an array of complexes that vary in the specific combination of the auxiliary subunits associated with the core complex (18).
  • the methods employ antibody-oligonucleotide conjugates (AB-oligos hereafter) to target dozens to hundreds of proteins and a molecular detector that can identify DNA-tagged spatially proximal entities.
  • AB-oligos antibody-oligonucleotide conjugates
  • Ideal detection of proximal nucleic acids requires capturing both of them on a single molecule. This presents a unique challenge, as both must be initiated in a 5' to 3' direction, and thus cannot both be captured using traditional methods.
  • a specialized single-stranded oligo with two free 3' ends the DNA-caliper that allows bidirectional priming and thus serves as a detector of molecular proximity by converting biological information into easily read DNA sequences (FIG.2).
  • Prod-seq and WhIP-seq Provided herein is the use of Prod-seq and WhIP-seq to catalog protein interactions, genomic binding and abundance of proteins, using the polycomb group complex members in hiPSCs as an example.
  • the methods and tools e.g., Ab-oligos, DNA-calipers
  • PcG Polycomb Group
  • a computational analysis pipeline for robustly identifying and quantifying interactions and protein abundances was produced.
  • the novel methods/tools were next used to map PPIs, binding patterns, and abundances of PcG members in the hiPSCs lines.
  • the experiments and data provide: (i) biological understanding by detection of genes and/or genomic loci that serve as key nodes in pluripotency and during neuro-differentiation; (ii) benchmarked Prod-seq and WhIP-seq protocols and a matching computational analysis pipeline; (iii) a set of optimized reagents (e.g., Ab-oligos) and conditions; and (iv) datasets from hiPSCs, NPCs, and neurons that can be used as a reference.
  • optimized reagents e.g., Ab-oligos
  • Prod-seq Proximity detection
  • WhIP-seq Without immunoprecipitation
  • Prod-seq uses oligonucleotide-conjugated antibodies (Ab-oligos) to target dozens to hundreds of proteins of interest and the DNA-caliper, a molecular detector that can identify proximal entities labeled with nucleic acids.
  • Ab-oligos oligonucleotide-conjugated antibodies
  • proximal nucleic acids oligonucleotide-conjugated antibodies
  • Ideal detection of proximal nucleic acids requires the ability to capture both of them on a single molecule. This presents a unique challenge, as both must be initiated in a 5' to 3' direction, and thus cannot both be captured using traditional methods.
  • the broadly applicable invention overcomes this hurdle with the DNA-caliper, a specialized single stranded oligonucleotide with two free 3' ends that allows bidirectional priming and thus serves as a proximity detector.
  • This unique molecule is generated by covalently linking the 5' ends of two oligonucleotides by click chemistry (23).
  • the two 3' ends of the DNA-caliper enable detection of adjacent DNA molecules via hybridization and strand extension.
  • variations in the DNA-caliper features e.g., linker properties or length
  • the Prod-seq process proceeds as follows (FIG.2A): (i) Fixed and lysed cells are incubated with multiple antibodies, where each specific antibody is covalently linked to a single-stranded oligonucleotide tail (Ab-oligos).
  • the oligonucleotide tail contains an anchor sequence (Anchor 1) complementary to both arms of the DNA-caliper (Anchor 1’; the DNA-caliper is depicted in FIG. 2A); a unique molecular identifier (UMI (24); a short random sequence) and an antibody specific unique barcode (Barcode n, m, etc.).
  • Each arm of the DNA-caliper contains at each 3’ end a sequence that is complementary to the antibody anchors at their 3’ ends (Anchor 1’).
  • Target sequences are captured by hybridization of the two DNA-caliper anchors to the complementary anchor sequences on the oligonucleotide linked to the antibodies, followed by DNA polymerase extension at both ends (FIG.2A).
  • the extension covalently links the barcodes and UMI sequences of the two antibodies.
  • the extended DNA-calipers are then denatured and captured by biotin pull-down and then converted into sequencing ready constructs. Sequencing the libraries allows identification of PPI instances as well as counting (using the UMIs (24)). Further, the same Ab- oligos also serve for quantitative detection of protein abundance (similar to other tools (25, 26)). WhIP-seq The second assay, WhIP-seq (FIG. 2B) is built in a similar way and uses the same Ab- oligos.
  • WhIP-seq uses a DNA-caliper with different anchors on each arm: one targets the Ab-oligo and the other targets the proximal genomic DNA (the DNA is sheared, and anchors are added at the 3’ end).
  • Advantages The methods and tools provided herein convey several advantages that will transform the fields of, for example, stem cells and regenerative medicine: (i) The unique design of the DNA-caliper allows its widespread use as a proximity detector (FIG.2).
  • the DNA-caliper can capture proximity between other entities such as protein- RNA, DNA-RNA, protein-metabolite, or cell-cell interactions, and thus for instance improve the characterization of in vitro differentiated cells.
  • UMI labeled oligonucleotide-barcoded antibodies (Ab-oligos) allows the simultaneous interrogation of multiple proteins from a single sample.
  • the assays have a reduced requirement for an a priori selection of a specific protein or genomic locus, such that only a general set of target proteins need be identified.
  • Prod-seq and WhIP-seq can identify interactions with non-protein targets using Ab- oligos to metabolites or other molecules in the cell.
  • the workflow of Prod-seq and WhIP-seq is modular and adding a second reaction that uses a subset of the sample enables quantitative abundance measurement of the studied proteins.
  • the methods provided herein allow, for the first time, detection of relative changes between the abundance of proteins and the strengths of their PPIs.
  • the tendencies for certain proteins or antibodies to participate in non-specific interactions can be more readily detected, leading to increased accuracy in the assessment of PPI rates.
  • WhIP-seq can eliminate the need for the spike-in controls currently required for quantitative ChIP-seq (31). Most differential analysis methods for ChIP-seq normalize samples by relying on the total number of mapped reads. When the level of the targeted protein differs substantially between samples, spike-ins are required to properly normalize signals between conditions (31). Inclusion of Ab-oligos to positive controls (e.g., H3) provides an internal normalization standard, thus eliminating the need for spike-in chromatin. (viii) The framework is designed to allow usage of reagents by multiple assays.
  • the same Ab-oligos can be used for both Prod-seq and WhIP-seq by employing a specific DNA-caliper for each tool.
  • Prod-seq and WhIP-seq can also be used to: (1) Discern near vs. far interactions in Prod-seq and WhIP-seq by varying the properties of the DNA-caliper (e.g., linker arm length); (2) The two tools can be merged as the DNA-calipers can be barcoded as well as tagged differently (e.g., replacing the biotin on the capture arm with, for example, another capture molecule, such as digoxigenin).
  • Prod-seq and WhIP-seq jointly would allow, for instance, to study the dynamics of PPIs between DNA-associated proteins together with their genomic localization as well as the abundance of each protein surveyed.
  • Such a merged approach enables discerning between instances where the PPIs and the genomic organization change in a dependent (e.g., for PPIs of proteins that are bound to the genome) or independent manner (e.g., for PPIs of proteins that are free in the cell/nucleus); (3) Integration of Prod-seq and WhIP-seq with single- cell approaches by generating emulsions of cell-barcoded DNA-calipers.
  • the set also comprises controls, such as histone H3 and H4 that are members of the nucleosome and thus serve as a highly abundant positive control and the FLAG-Tag and H3K27M antibodies that can serve either as negative controls (if not expressed) or as positive controls (if expressed).
  • controls such as histone H3 and H4 that are members of the nucleosome and thus serve as a highly abundant positive control and the FLAG-Tag and H3K27M antibodies that can serve either as negative controls (if not expressed) or as positive controls (if expressed).
  • Initial set of antibodies to establish Prod-seq and WhIP-seq and study the PcG The below table highlights antibodies that target PRC1, PRC2 their target histone modifications and positive and negative controls and a signaling pathway (HER-PIK-mTOR).
  • the canonical PRC1 is composed of RING1A or RING1B and one of the PCGF1-6 proteins as well as a chromobox (CBX2, CBX4, CBX6, CBX7 or CBX8) and one of the PHC1-PHC3 proteins.
  • Variant PRC1 involve RYBP and a PCGF protein that is plays a role in incorporation of the other subunits (e.g., HDAC1 and HDAC2), thus leading to several of PRC1 variants (18).
  • the core of PRC2 is composed of EZH1 or EZH2, EED, RBBP4 or RBBP7 and SUZ (12). Variants of this complex, similar to PRC1, involve interactions with distinct subunits. Thus, the PRC2.2.
  • Recombinant protein complexes were initially used– this allowed one to identify the noise ratios for Ab-oligos whose target is not present while capturing signal from the targets.
  • a cell line with a knockout of EZH2 was then used, the H3K27me3 writer component of the PcG complex (EZH2-KO cells (1)), and thus have reduced levels of the H3K27me3 modification.
  • a multiplex set of Ab-oligos (3 or 6) were used and the abundance of H3K27me3 in KO and wildtype (WT) cells across multiple optimization conditions were compared.
  • Prod-seq (i) Fixed and lysed cells are incubated with multiple Ab-oligos. For each Ab- oligo, the oligonucleotide tail contains: Anchor 1 that is complementary to the DNA-caliper arms; a UMI; a unique barcode; and a Hook that enables creation of a sequencing library. ii) The Anchor 1’ regions on the two arms of the DNA-caliper are hybridized to the complementary Anchor 1 regions on the Ab-oligos, followed by extension of the DNA-caliper in both directions.
  • the extended DNA-calipers are then denatured and captured by biotin pull-down followed by annealing a Splint with 3’-overhanging Hook’ regions on both sides.
  • the Splint is ligated, and the free 3’ ends of each of the DNA-caliper arms are extended using the other arm as a template.
  • the extended DNA-calipers are linearized by USER enzyme excision of the dU (deoxyuridine; part of the DNA-caliper) and serve as a template for PCR using Illumina AmpliSeq kit and primers that match regions P1 and P2.
  • Sequencing the libraries allows detection and counting (using the UMIs (24)) of PPI instances.
  • the same Ab-oligos also serve for quantitative detection of protein abundance (similar to ID-seq (25) or NEAT-seq (26)). This is done by using only the biotinylated arm of the DNA-caliper in a separate hybridization-extension reaction followed by PCR to generate a sequencing library.
  • WhIP-seq The DNA-caliper used for WhIP-seq differs in one arm from the one used for Prod-seq. This design of the DNA-caliper allows capturing the Ab-oligo on one arm and proximal genomic DNA on the other.
  • the process includes the following steps: (i) First, as in ChIP, cellular chromatin is cross-linked to covalently bind proteins to the DNA.
  • each adaptor contains a UMI and a single stranded overhang region (Anchor 2) complementary to the lower arm of the DNA-caliper (Anchor 2’) (FIG. 7A).
  • Anchor 2 complementary to the lower arm of the DNA-caliper (Anchor 2’)
  • FIG. 7A the same Ab-oligos used for Prod-seq are added to the cross-linked chromatin constructs.
  • the Anchor sequences of the DNA-caliper are hybridized to the complementary regions on the Ab-oligos and on the adaptors followed by extension of the DNA-caliper in both directions.
  • the extended DNA-calipers are denatured and captured by biotin pull-down followed by annealing and ligation of a Splint molecule.
  • this Splint has one arm with Hook’ to match the Ab-oligo, while the other has a Poly(G) overhang.
  • the other end of the extended DNA-caliper is 3’-modified by terminal transferase to have a Poly(C). Following intramolecular hybridization, the free 3’ ends of each of the DNA-caliper arms are then extended using the other arm as a template.
  • PQ-seq samples (#A-P): Resuspend beads in 8 ul of rCut+TWEEN+PI Add 2 ul of 1:100 diluted 100uM of calpber oligos (2pmol) from step 14 and incubate RT for 30min. 17.
  • Klenow (exo-) Extension use the same master mix and program as the Prod-seq samples below in step 19; skip the remove supernatant step but do the heat inactivation Add 35 ul of rCut+TWEEN+PI and keep at 4oC after 18.
  • Prod-seq samples (#1-16): resuspend beads in 5 ul rCut+TWEEN+PI and add Calipers; use caliper with ETSSB from step 14. Add 5 ul caliper+ETSSB into each sample and incubate RT for 30 min. 19. Klenow Extension; make Klenow exo- mix according to table below and add to bead suspension.
  • HMGB1 incubation Use 75 ul sample from the previous step + 18 ul master mix; for ligation in rCutSmart buffer. After adding 18 u s. 27. Do ligation with T4 ligase, using 93 ul total from prior step and combing with the following: After adding 5ul pe r sample, incubate at 25 C for 20 minutes, and 65oC for 10 minutes. Move to 4oC until step 28. 28. Cut the 98 ul of sample with EcoRV according to table below:
  • PCR amplification of Caliper/Oligo duplexes Amplify 14.6 ul of the PROD-seq samples with PrimeStar according to table below: Add 33.4ul to stock) for a final volume of 50ul. Run the following PCR conditions 1. 98oC for 30 seconds 2. 98oC for 10 seconds 3. 62oC for 10 seconds 4. 72oC for 10 seconds Repeat 2-419 times 5. 72oC for 2 minutes Hold at 12oC. 30. SpeedBead purification of PCR products Make master mixes as calculated below: Take PCR samples orresponding 7.5% PEG bead mastermix.
  • Prod-seq and WhIP-seq using the PcG complex in iPSCs The PcG complex that is composed of two main subcomplexes, Polycomb Repressive Complex 1 and 2 (PRC1 and PRC2) that are regulators in pluripotency and differentiation (18) were focused on to demonstrate aspects of the invention.
  • iPSCORE_1_14 and iPSCORE_1_13 iPSC lines from African American female and male, respectively.
  • Prod-seq and WhIP-seq use several synthetic components and biochemical reactions, such as the Ab-oligos, the DNA-caliper and enzymes (e.g., DNA Polymerase).
  • the following parameters in designing the different synthetic components were considered: (i) Optimal length and sequence of the DNA-caliper (FIG.3); (ii) GC content of all of the components of the DNA-caliper (anchors, primers and flanking regions); (iii) Tm for high specificity hybridization between the anchors on the DNA-caliper and the Ab-oligos; (iv) Avoidance of unwanted secondary structures; (v) Design of specific barcodes for multiple antibodies, accounting for decoding each one to identify the antibody; (vi) To avoid low diversity in nucleotides during sequencing which may impact base call accuracy, the primers used (P1-P3; FIGS.6-7) contain a “spacer” region at the 5’ end.
  • the oligonucleotides are functionalized with TCO-PEG4-NHS, and carrier free monoclonal antibodies (concentration >1mg/ml) with Methyltetrazine-PEG4-NHS, and incubated together overnight at 4°C. The reaction is quenched with glycine and free antibodies and oligonucleotides are removed. The efficiency of the conjugation was evaluated by absorbance and silver staining. To verify that the antibody functionality is not impacted, the ability of the Ab-oligo to bind to permeabilized human cells was measured. Ab-oligos produced by these methods were used in Prod-seq experiments (FIG. 4).
  • the Ab-oligos will also allow the determination of the ability to capture signal (positive controls) vs. noise (our negative controls).
  • Prod-seq cells are crosslinked using formaldehyde and disuccinimidyl glutarate, lysed in lysis buffer (containing EDTA, Tris-HCl, and SDS), blocked with blocking buffer (containing PBS, Triton X-100, BSA and protease inhibitors), and incubated with Ab-oligos overnight at 4°C. Following washes (in PBS), the DNA-caliper is added, and the hybridization is followed by strand- extension using Klenow fragment.
  • Splint (FIG.6; a double-stranded DNA molecule with two 3’ overhangs and 5’-phosphorylated termini) is added to enable intramolecular annealing and subsequent ligation by T4 DNA ligase and another round of strand-extension by Klenow.
  • the extended DNA- calipers are linearized by USER enzyme, and PCR with P1 and P2 primers is used to generate the PPI fragment pool.
  • the biotinylated arm of the DNA-caliper is added to a subset of the sample that was incubated overnight with the Ab-oligo panel and washed with PBS.
  • the chromatin is incubated with the Ab-oligos overnight, and after washes the DNA-caliper is hybridized (55°C) followed by strand-extension (Klenow). Following denaturation (95°C) and biotin pulldown of the extended DNA-calipers, a Splint is added.
  • This Splint has cleavable extension-blockers on both 3’ termini, 5’ phosphorylated termini, one arm that matches the Ab-oligo, and the other with a Poly(G) overhang.
  • the other end of the extended DNA-caliper is 3’-modified by terminal transferase to have a Poly(C) stretch. The cleavable blockers are removed, and the Splint is ligated to the extended DNA-caliper.
  • interaction counts are likely to scale with the relative abundance of proteins irrespective of interaction strength, (ii) only a fraction of each targeted protein is likely to participate in interactions measured by the Ab-oligo panel used in the experiment, (iii) certain Ab-oligos may have higher levels of non-specific interactions, (iv) specific barcodes and/or UMIs may lead to biased incorporation into the final read pairs. Due to the complexity of these potential factors and their influence on the interpretation of the data, the experiment can include Ab-oligos targeting both positive controls (e.g., histone H3 and H4) and negative controls (e.g., FLAG-Tag, H3K27M) to verify the assay was successful and estimate the prevalence of non-specific interactions.
  • positive controls e.g., histone H3 and H4
  • negative controls e.g., FLAG-Tag, H3K27M
  • PPIs are quantified between each pair of antibodies, yielding (n(n+1))/2 potential pairwise interactions across the experiment, which includes self-interactions (e.g., dimers).
  • the PPI counts can be modeled as a contingency table and the significance of interaction between any two proteins assessed using the Chi-squared test and controlled for multiple hypothesis testing using the Benjamini–Hochberg procedure.
  • the ratio of observed to expected interactions between each protein and negative control Ab-oligos (e.g., FLAG-Tag) in the contingency table establish a lower bound for PPI detection across the dataset.
  • Detected PPIs will be compared to the known PPIs from public databases (e.g., BioGRID (53)) to estimate the relative accuracy. Putative multi-protein complexes can be inferred from the pairwise PPIs by applying the MCODE algorithm (54) as has been implemented (55).
  • WhIP-seq Once the protein and genomic sequence information is extracted from the reads, genomic binding maps for each Ab-oligo in the panel can be determined. Since WhIP-seq interrogates sonicated DNA fragments associated with protein complexes, the data can be analyzed in a similar manner to conventional ChIP-seq pipelines.
  • Ab-oligos encoding negative controls (e.g., FLAG-Tag) will be used in each experiment to estimate non-specific interactions.
  • the input DNA will be sequenced as done in ChIP-seq.
  • WhIP-seq data will first be segregated by Ab-oligo barcode, resulting in collections of DNA fragments specific to each targeted protein. These DNA fragments will be aligned to the genome using BWA (56) and analyzed for significantly enriched peaks by HOMER (50), using both the input and negative control Ab-oligo data as controls.
  • Enriched regions will be subjected to procedures such as annotation to nearby genes and regulatory features (e.g., promoters, ChromHMM (57), etc.) and the enrichment of DNA motifs (50). Detecting changes over time As an example, changes in PPIs, protein-DNA interactions and abundance of PcG members during hiPSCs neural linage commitment can be determined. Multiple lines of evidence show that the PcG complexes present dynamic configurations and interaction patterns (58). These variations in PcG composition are suggested to enable the pleotropic roles these complexes play during development, potentially via fine-tuning of its activity. Thus, for instance, the various configurations of the PcG each bind to a distinct set of genomic loci and act in a different manner (58).
  • Prod-seq and WhIP-seq can be used to compare the changes in PPI and protein-DNA interactions of PcG members between the pluripotent hiPSCs and the committed hiPSC-derived NPCs and cortical neurons.
  • Prod-seq and WhIP-seq on hiPSCs-derived cortical NPCs and neurons will be generated as was previously described (59). Briefly, hiPSC plated in Matrigel-coated dishes are maintained in mTeSR Plus media to confluency. Next, media is switched to NMM (Neurobasal media supplemented with TGFb- inhibitors) for 7 days, and then to NMM media with bFGF for several days.
  • NPCs can differentiate into cortical neurons by withdrawing bFGF from the media for 4-6 weeks (59).
  • Classical markers for cortical NPCs and neurons will be used to monitor effectiveness of the protocol (60, 61).
  • NPCs and neurons will be fixed and collected.
  • an array of Ab-oligos e.g., Table provided above
  • Prod-seq and WhIP-seq will be performed in triplicates (3 rounds of NPC/neuron differentiation), from female and male hiPSCs. In parallel, co-IP and ChIP-seq will be performed.
  • the approach also uses a probability mixture model to assign PPI confidence scores via maximum likelihood estimation on the generated protein quantification and PPI detection profiles. These confidence scores can then be used to filter PPI detection noise and to evaluate the reliability of the detected PPIs. These data provide information about the probability distribution to use in the mixture model to appropriately score true-positive and true-negative PPIs.
  • To assign confidence scores to Prod-seq PPI detection profiles the similarity between Prod-seq and AP-MS for quantitative PPI detection was built on. As a starting point, we adapted the probability mixture model scoring approach of Significance Analysis of INTeractome (SAINT; Choi et al., Curr Protoc Bioinformatics, 2013), one of the widely used AP-MS PPI scoring methods.
  • ⁇ 23 ⁇ 45%67 ⁇ 2 + ⁇ 2 ⁇ # ⁇ 3 + ⁇ 3 ) + ⁇ 6%:4; ⁇ ⁇ 2 ⁇ 3 and the additional + , denotes the proportion of true interactions in the entire dataset;
  • ⁇ ⁇ is the readout scaling factor for PQ-seq detection protocol and sequencing;
  • ⁇ 6%:4; ⁇ and ⁇ 45%67 denote the readout scaling factors for specific and non-specific PPI detection in Prod-seq, respectively;
  • ⁇ % and ⁇ % denote the levels of specific and non-specific antibody-oligo binding for each protein target ⁇ , respectively.
  • a Figure 14 illustrates a diagrammatic representation of a machine 1400 in the form of a computer system within which a set of instructions may be executed for causing the machine 1400 to perform any one or more of the methodologies discussed herein, according to an example implementation.
  • Figure 14 shows a diagrammatic representation of the machine 1400 in the example form of a computer system, within which instructions 1402 (e.g., software, a program, an application, an applet, an app, or other executable code) for causing the machine 1400 to perform any one or more of the methodologies discussed herein may be executed.
  • the instructions 1402 may cause the machine 1400 to implement the computational processes, methods, techniques, and algorithms described herein.
  • the instructions 1402 transform the general, non-programmed machine 1400 into a particular machine 1400 programmed to carry out the described and illustrated functions in the manner described.
  • the machine 1400 operates as a standalone device or may be coupled (e.g., networked) to other machines.
  • the machine 1400 may operate in the capacity of a server machine or a client machine in a server- client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment.
  • the machine 1400 may comprise, but not be limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular telephone, a smart phone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of executing the instructions 1402, sequentially or otherwise, that specify actions to be taken by the machine 1400.
  • PC personal computer
  • PDA personal digital assistant
  • machine 1400 can include logic, one or more components, circuits (e.g., modules), or mechanisms. Circuits are tangible entities configured to perform certain operations. In an example, circuits can be arranged (e.g., internally or with respect to external entities such as other circuits) in a specified manner.
  • one or more computer systems e.g., a standalone, client or server computer system
  • one or more hardware processors can be configured by software (e.g., instructions, an application portion, or an application) as a circuit that operates to perform certain operations as described herein.
  • the software can reside (1) on a non-transitory machine readable medium or (2) in a transmission signal.
  • the software when executed by the underlying hardware of the circuit, causes the circuit to perform the certain operations.
  • a circuit can be implemented mechanically or electronically.
  • a circuit can comprise dedicated circuitry or logic that is specifically configured to perform one or more techniques such as discussed above, such as including a special-purpose processor, a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC).
  • a circuit can comprise programmable logic (e.g., circuitry, as encompassed within a general-purpose processor or other programmable processor) that can be temporarily configured (e.g., by software) to perform the certain operations. It will be appreciated that the decision to implement a circuit mechanically (e.g., in dedicated and permanently configured circuitry), or in temporarily configured circuitry (e.g., configured by software) can be driven by cost and time considerations.
  • circuit is understood to encompass a tangible entity, be that an entity that is physically constructed, permanently configured (e.g., hardwired), or temporarily (e.g., transitorily) configured (e.g., programmed) to operate in a specified manner or to perform specified operations.
  • each of the circuits need not be configured or instantiated at any one instance in time.
  • the circuits comprise a general-purpose processor configured via software
  • the general-purpose processor can be configured as respective different circuits at different times.
  • Software can accordingly configure a processor, for example, to constitute a particular circuit at one instance of time and to constitute a different circuit at a different instance of time.
  • circuits can provide information to, and receive information from, other circuits.
  • the circuits can be regarded as being communicatively coupled to one or more other circuits.
  • communications can be achieved through signal transmission (e.g., over appropriate circuits and buses) that connect the circuits.
  • communications between such circuits can be achieved, for example, through the storage and retrieval of information in memory structures to which the multiple circuits have access.
  • one circuit can perform an operation and store the output of that operation in a memory device to which it is communicatively coupled.
  • a further circuit can then, at a later time, access the memory device to retrieve and process the stored output.
  • circuits can be configured to initiate or receive communications with input or output devices and can operate on a resource (e.g., a collection of information).
  • a resource e.g., a collection of information.
  • the various operations of method examples described herein can be performed, at least partially, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors can constitute processor-implemented circuits that operate to perform one or more operations or functions.
  • the circuits referred to herein can comprise processor-implemented circuits.
  • the methods described herein can be at least partially processor implemented. For example, at least some of the operations of a method can be performed by one or processors or processor-implemented circuits.
  • the performance of certain of the operations can be distributed among the one or more processors, not only residing within a single machine, but deployed across a number of machines.
  • the processor or processors can be located in a single location (e.g., within a home environment, an office environment or as a server farm), while in other examples the processors can be distributed across a number of locations.
  • the one or more processors can also operate to support performance of the relevant operations in a "cloud computing" environment or as a "software as a service” (SaaS).
  • Example implementations can be implemented in digital electronic circuitry, in computer hardware, in firmware, in software, or in any combination thereof.
  • Example implementations can be implemented using a computer program product (e.g., a computer program, tangibly embodied in an information carrier or in a machine readable medium, for execution by, or to control the operation of, data processing apparatus such as a programmable processor, a computer, or multiple computers).
  • a computer program can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a software module, subroutine, or other unit suitable for use in a computing environment.
  • a computer program can be deployed to be executed on one computer or on multiple computers at one site or distributed across multiple sites and interconnected by a communication network.
  • operations can be performed by one or more programmable processors executing a computer program to perform functions by operating on input data and generating output. Examples of method operations can also be performed by, and example apparatus can be implemented as, special purpose logic circuitry (e.g., a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC)).
  • FPGA field programmable gate array
  • ASIC application-specific integrated circuit
  • the computing system can include clients and servers.
  • a client and server are generally remote from each other and generally interact through a communication network.
  • the relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
  • both hardware and software architectures require consideration.
  • the choice of whether to implement certain functionality in permanently configured hardware e.g., an ASIC
  • temporarily configured hardware e.g., a combination of software and a programmable processor
  • a combination of permanently and temporarily configured hardware can be a design choice.
  • hardware e.g., machine 1400
  • software architectures that can be deployed in example implementations.
  • the machine 1400 can operate as a standalone device or the machine 1400 can be connected (e.g., networked) to other machines. In a networked deployment, the machine 1400 can operate in the capacity of either a server or a client machine in server-client network environments. In an example, machine 1400 can act as a peer machine in peer-to-peer (or other distributed) network environments.
  • the machine 1400 can be a personal computer (PC), a tablet PC, a set-top box (STB), a Personal Digital Assistant (PDA), a mobile telephone, a web appliance, a network router, switch or bridge, or any machine capable of executing instructions (sequential or otherwise) specifying actions to be taken (e.g., performed) by the machine 1400.
  • PC personal computer
  • PDA Personal Digital Assistant
  • Example machine 1400 can include a processor 1404 (e.g., a central processing unit CPU), a graphics processing unit (GPU) or both), a main memory 1406 and a static memory 1408, some or all of which can communicate with each other via a bus 1410.
  • the machine 1400 can further include a display unit 1412, an alphanumeric input device 1414 (e.g., a keyboard), and a user interface (UI) navigation device 1416 (e.g., a mouse).
  • a processor 1404 e.g., a central processing unit CPU
  • GPU graphics processing unit
  • main memory 1406 main memory
  • static memory 1408 static memory
  • the machine 1400 can further include a display unit 1412, an alphanumeric input device 1414 (e.g., a keyboard), and a user interface (UI) navigation device 1416 (e.g., a mouse).
  • UI user interface
  • the display unit 1412, input device 1414 and UI navigation device 1416 can be a touch screen display.
  • the machine 1400 can additionally include a storage device (e.g., drive unit) 1418, a signal generation device 1420 (e.g., a speaker), a network interface device 1422, and one or more sensors 1424, such as a global positioning system (GPS) sensor, compass, accelerometer, or another sensor.
  • the storage device 1418 can include a machine readable medium 1426 on which is stored one or more sets of data structures or instructions 1402 (e.g., software) embodying or utilized by any one or more of the methodologies or functions described herein.
  • the instructions 1402 can also reside, completely or at least partially, within the main memory 1406, within static memory 1408, or within the processor 1404 during execution thereof by the machine 1400.
  • one or any combination of the processor 1404, the main memory 1406, the static memory 1408, or the storage device 1418 can constitute machine readable media.
  • the machine readable medium 1426 is illustrated as a single medium, the term "machine readable medium" can include a single medium or multiple media (e.g., a centralized or distributed database, and/or associated caches and servers) that configured to store the one or more instructions 1402.
  • machine readable medium can also be taken to include any tangible medium that is capable of storing, encoding, or carrying instructions for execution by the machine and that cause the machine to perform any one or more of the methodologies of the present disclosure or that is capable of storing, encoding or carrying data structures utilized by or associated with such instructions.
  • machine readable medium can accordingly be taken to include, but not be limited to, solid-state memories, and optical and magnetic media.
  • machine-readable media can include non-volatile memory, including, by way of example, semiconductor memory devices (e.g., Electrically Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM)) and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
  • the instructions 1402 can further be transmitted or received over a communications network 1428 using a transmission medium via the network interface device 1422 utilizing any one of a number of transfer protocols (e.g., frame relay, IP, TCP, UDP, HTTP, etc.).
  • transfer protocols e.g., frame relay, IP, TCP, UDP, HTTP, etc.
  • Example communication networks can include a local area network (LAN), a wide area network (WAN), a packet data network (e.g., the Internet), mobile telephone networks (e.g., cellular networks), Plain Old Telephone (POTS) networks, and wireless data networks (e.g., IEEE 802.11 standards family known as Wi-Fi®, IEEE 802.16 standards family known as WiMax®), peer-to-peer (P2P) networks, among others.
  • the term “transmission medium” shall be taken to include any intangible medium that is capable of storing, encoding or carrying instructions for execution by the machine, and includes digital or analog communications signals or other intangible medium to facilitate communication of such software.
  • Example 3- Western-seq Western blotting limitations Western blotting is a widely used technique developed to detect expression of a protein of interest in a sample.
  • Protein lysate is created from a sample, and along with a protein ladder run through gel electrophoresis to separate all of the proteins present in the sample by size.
  • the proteins are transferred from inside the gel onto a membrane using a voltage, making them accessible to antibody binding.
  • the membrane can be stained.
  • the membrane then undergoes a process called blocking, to reduce any nonspecific binding and background noise for the final product.
  • the membrane is treated with an antibody that binds to the protein of interest, called the primary antibody.
  • Primary antibodies can be pAb, mAb, or even rAb and are generated in small animals, such as mice, rabbits, or rats. Using smaller animals makes the antibody easier to generate and by using an antibody from a different species than the sample, it allows for detection of only the protein of interest.
  • the sample is treated with another antibody, called a secondary antibody.
  • the secondary antibody is specific to the Fc domain of the primary antibody.
  • This secondary antibody can be conjugated to also include a fluorescence tag or horseradish peroxidase (HRP), so as to allow for visualization in imaging. For example, if a researcher were interested in visualizing beta actin protein in a human cell sample, the cell sample would first be lysed and the lysate would be size separated using gel electrophoresis.
  • the sample After the transfer and blocking steps, the sample would be stained with a primary antibody specific for beta actin. The sample would then be incubated with a secondary antibody specific for the Fc domain of the primary antibody. The secondary antibody contains a fluorescence label that allows visualization.
  • the membrane In the final step the membrane is imaged, and the result is a picture of the membrane with a band at the size of the protein of interest, confirmed by the protein ladder standard (8).
  • researchers use this technique to quantify protein levels, such as to validate knockdown or knockout systems. If a specific protein is knocked down or knocked out in a sample and a Western blot is run, the resulting band will be much weaker or entirely absent when compared to a control or wildtype sample.
  • Western blots may be widely used, but they are also known to be challenging and difficult to reproduce.
  • the process typically takes two days as the blocking step and/or primary antibody incubation occurs overnight.
  • the vast number of steps, each requiring various buffers and reagents, give much room for complications to arise with little room for pinpointing exactly which is the problematic step. If a membrane is imaged after blotting and there are no protein bands, it could be an issue with the wash buffer, with the concentration of the primary antibody, with what the primary antibody was diluted in, with the primary antibody itself or with the secondary antibody (8).
  • ELSA Enzyme Linked Immunosorbent Assay
  • the oligos have overlapping complementary sequences, and when both antibodies are bound to the same protein, the oligos are brought into proximity with each other and create a double stranded sequence. If there is non-specific binding of a single antibody, the oligo will remain single stranded and cannot be utilized in downstream processes.
  • qPCR can be used to detect the amount of double stranded DNA for a target protein, and therefore determine the amount of protein in the sample.
  • Adaptors and barcodes are added to the DNA for next generation sequencing (NGS) to increase the throughput of analysis (18).
  • NGS next generation sequencing
  • SomaScan utilizes modified aptamers, which are oligonucleotides that have specificity for target proteins and are used in place of an antibody.
  • the aptamers are referred to as SOMAmers (slow off-rate modified aptamers) and are marketed as having more specificity than polyclonal antibodies.
  • SOMAmers slow off-rate modified aptamers
  • the SOMAmers are eluted and detected by DNA quantification (23).
  • DNA quantification Unlike the single stranded DNA tags of PEA or the single stranded SOMAmers, van Buggenum et al. have developed immunodetection sequencing, or ID-seq, which uses double stranded DNA tags bound to antibodies to quantify protein levels in a sample. After the antibody- dsDNA conjugate binds to the epitope, the dsDNA tag gets cleaved.
  • the dsDNA on the antibody contains a barcode and unique molecular identifier (UMI), a randomized nucleotide sequence used in NGS which allows researchers to discern read duplicates from unique, significant events (26).
  • UMI unique molecular identifier
  • the purpose of the barcode is to associate the DNA with a specific antibody (25). The presence of a UMI increases confidence in this association.
  • the DNA tags each have a barcode for the specific antibody as well as UMI.
  • a range of antibodies are added to a sample to bind to their respective target proteins.
  • the DNA tags are hybridized, extended, and amplified using PCR. In this amplification they are also prepared for sequencing and once sequenced, the barcodes and UMIs are used to detect protein count. (FIGS.9A-9D)
  • Western-seq is a novel method that provides higher confidence by multiplexed detection of protein sizes and quantification Combining Western blotting and PQ-seq into Western-seq yields a novel protein quantification technique that merges the qualitative and quantitative advantages of each method (FIGS.9A-9F).
  • the main advantage of Western blotting is that there is a visual of the protein of interest by the binding of the antibody, which is confirmed by the size in kDa.
  • the biggest advantage of PQ-seq is that one can associate relative abundance of protein to DNA, which is easier to quantify than protein, through DNA sequencing on a high scale.
  • the beginning of the Western-seq protocol is very similar to a standard Western blot. The sample is run on a gel, transferred to a membrane, and blocked.
  • the primary antibodies used are the same as in PQ-seq and are tagged with a DNA oligonucleotide that features a barcode and UMI.
  • the antibodies can be incubated with single stranded binding protein (SSB) to bind the oligonucleotides, improve signal, and decrease background noise.
  • SSB single stranded binding protein
  • a mixture of hundreds of conjugated primary antibodies can be used.
  • the membrane is treated with the traditional secondary antibody for each species of primary antibody used and the membrane is imaged. Due to being blotted with many different antibodies, there will not be a singular band at a singular size for the protein of interest, but many bands at the various sizes confirmed by the protein ladder. At this point, the protocol diverges from a typical Western blot.
  • segments of 10-20 kDa increments are cut out of the membrane, placed in a PCR tube strip or 96-well plate and assigned a barcode label.
  • a caliper arm will be added and annealed to the oligonucleotide and undergo a Klenow extension to extend the sequence.
  • the oligonucleotide sequence is amplified, and in a second PCR, sequencing primers are added to the sequence.
  • the samples are pooled and sequenced, and the analyzed results will inform on the abundance of protein in the sample by relating the abundance of each unique read with the barcodes to which antibodies they came from.
  • FIGS.9A-9F provide an overview of Western-seq. Multiplex protein detection can occur via conjugated antibodies, hybridization and strand extension (FIGS.10A-10B), while barcoding allows detection of different proteins in one reaction (FIGS. 10C-10D). Western-seq allows multiplex protein detection with size separation (FIG. 10E).
  • FIG. 10A-10B conjugated antibodies, hybridization and strand extension
  • FIGS. 10C-10D barcoding
  • 10F provides an exemplary protocol for Western-seq.
  • An Exemplary Method Using Cultured Cells Cell Culture HEK293T cells were grown in a DMEM solution (2mM L-glutamine, 10% HI-FBS, 1% Antimycotic-Antibiotic) on standard culture dishes. Plates were stored in 37 °C incubators with 5% CO 2 . Tazemetostat and DMSO treated cells were treated with l0 uM of their respective medium for 48 hours. The EZH2 knockout cell line was obtained from researchers Wang et al. (27).
  • HeLa-S3 (ATCC CCL-2.2 TM ) cells were grown in DMEM with GlutaMAX TM (DMEM, 4.5 g/L D-Glucose, 110mg/L sodium pyruvate, 2 mM L-glutamine, 10% HI-FBS, 1 % Antibiotic Antimycotic) on standard culture dishes. Plates were stored in 37 °C incubators with 5% CO 2 . Induced pluripotent stem cells (iPSCs) were grown on Matrigel (Corning 354230) coated plates with mTeSR (Stem Cell Technologies 85850). Plates were stored in 37 °C incubators with 5% CO 2 .
  • DMEM 4.5 g/L D-Glucose, 110mg/L sodium pyruvate, 2 mM L-glutamine, 10% HI-FBS, 1 % Antibiotic Antimycotic
  • the harvested cells were crosslinked with 1 % formaldehyde/ PBS solution per 1 million cells and incubated at room temperature for 10 minutes while rotating at 8 rpm. The reaction was quenched by adding 1/20th volume of 2.625 M glycine and 1/20th volume 10%; BSA The cell pellet was resuspended, stored in 0.5% BSA/PBS, and frozen at -80°C.
  • Western blotting Lysing conditions and preparation LDS (NuPAGE TM LDS Sample Buffer NP0007) method: Thaw cell pellet and resuspend in lx PBS. Add an equal amount of 2X LDS, and vortex for two minutes. Aliquot and store at -80 °C.
  • This method was used for preparing the protein lysate for all HEK293T cells and HeLaS3 cells, as well as to prepare some of the HEK293T samples.
  • RIPA 150 mM NaCl, 1.0%) N-P-40, 0.5% sodium deoxycholate, 0.1 % SDS, lmM EDTA, 50 mM Tris, pH 8.0
  • Method Thaw cell pellet and resuspend in lx PBS. Add an equal amount of RIPA buffer with 1 X Protease Inhibitor and incubate for 30 minutes. Spin down at 12,000g for 5-l 0 minutes, aliquot the supernatant and store at -80 °C.
  • NP-40 50mM Tris-HCl pH 8.5, 150mM NaCl, 1 % NP-40, cOmplete EDTA-free Protease Inhibitor 1 X (Roche l 1873580001), PhosSTOP lX (Roche 4906845001), Benzonase 1 :500) method: Thaw cell pellet and resuspend in lx PBS. Add an equal amount of NP-40 buffer with IX Protease Inhibitor and incubate for 30 minutes.
  • Lysis buffer 2 (l0mNI EDTA, 50 mM HEPES, 0.5% SDS, pH 7.3 - 7.5) method: Thaw cell pellet and resuspend in lx PBS. For each 1 million cells, add 100 uL lysis buffer 2 and incubate at 37 °C for 60 minutes.
  • Protein concentration was measured using the ThermoFisher Scientific Qubit protein assay kit and ThermoFisher Scientific Qubit fluorometer the same day a Western blot was done once the cell pellet had thawed.
  • Western blotting Sample preparation, running, and transfer Protein amounts were normalized to 15 ug of protein per lane. 3.3uL of sample buffer which included 9% beta-mercaptoethanol and 6X Laemmli Buffer (375 mM Tris-HCl, 10% SDS, 7.5 ⁇ o glycerol, 0.03%i bromophenol blue) along with 15 ug of protein was added to PCR tubes, and water was added to the samples to bring the total volume to 20 uL.
  • the transfer system consisted of lX transfer buffer (25mM Tris, 125 mM glycine, 20% methanol pH 8.3) and was placed in a bucket filled with ice in the cold room with ice packs in the system. The transfer was an hour at 100V. The membranes were blocked overnight on a shaker at 4 °C in 5% BSA/ lX TBST (Fisher Scientific BP1600-100; 20mM Tris, 150mM NaCl, 0.5% Tween (EMD Millipore 655204)). When salmon sperm DNA (Invitrogen 15-632-011) was added to the blocking buffer, l00ug/ mL of blocking buffer was used.
  • the antibodies were conjugated by the company AlphaThera, using their oYo-Link technology. This technology connects the oligonucleotides to the Fc chain of the antibody with a very specific photo crosslinker labeling called Light Activated Site-specific Conjugation or LASIC. Once conjugated, the antibodies are purified by fast protein liquid chromatography, FPLC, to separate unbound oligos from the antibodies as well as the antibodies with different degrees of labeling.
  • Western-seq Protocol 1. Cut the membrane for each desired band and sample and submerge in 38 ul of T4 ligase buffer (NEB B0202S) + 0.05% Tween. Boil at 95 °C for 5 minutes. 2.
  • a DNA barcoded protein ladder for every 10 kDa increment in size can be used.
  • a standard commercial ladder can be used.
  • SETD5 Regulates Chromatin Methylation State and Preserves Global Transcriptional Fidelity during Brain Development and Neuronal Wiring. Neuron, 104(2), 271- 289.e 13. https:/ /doi.org/10.1016/j.neuron.2019.07.013 15.
  • SETD5-Coordinated Chromatin Reprogramming Regulates Adaptive Resistance to Targeted Pancreatic Cancer Therapy. Cancer Cell, 37(6), 834-849.eB. https:/ /doi.org/10.1016 ⁇ j.ccell.2020.04.014 16. Alhajj, M., Zubair, M., & Farhana, A. (2023). Enzyme Linked Immunosorbent Assay. In StatPearls. StatPearls Publishing. http://w'vvw.ncbi.n1m.nih.gov/books/NBK555922/ 17. Aydin, S. (2015). A short history, principles, and types of ELISA, and our laboratory experience with peptide/protein analyses using ELISA.

Landscapes

  • Life Sciences & Earth Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Chemical & Material Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Molecular Biology (AREA)
  • Immunology (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Physics & Mathematics (AREA)
  • Analytical Chemistry (AREA)
  • Urology & Nephrology (AREA)
  • Hematology (AREA)
  • Biomedical Technology (AREA)
  • General Health & Medical Sciences (AREA)
  • Pathology (AREA)
  • Biotechnology (AREA)
  • Microbiology (AREA)
  • Organic Chemistry (AREA)
  • Biochemistry (AREA)
  • Biophysics (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Cell Biology (AREA)
  • Zoology (AREA)
  • Wood Science & Technology (AREA)
  • Food Science & Technology (AREA)
  • Medicinal Chemistry (AREA)
  • General Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • Genetics & Genomics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Abstract

Provided herein are methods and tools that enable(s) simultaneous identification and characterization of multiple interactions between proteins and between proteins and DNA or RNA within functional complexes. This transformative approach/method allows for high-throughput queries of the dynamic interactions of proteins and expands the focus from single interactions towards an unprecedented, multidimensional interrogation of the cellular complexity.

Description

TOOLS FOR INTERROGATING DYNAMIC ORGANIZATIONAL PRINCIPLES OF PROTEIN COMPLEXES IN VIVO PRIORITY This patent application claims the benefit of priority to U.S. Provisional Application Serial No. 63/513,630, filed July 14, 2023 and to U.S. Provisional Application Serial No. 63/561,661, filed March 5, 2024, which is incorporated by reference herein in its entirety. INCORPORATION BY REFERENCE OF SEQUENCE LISTING This application contains a Sequence Listing which has been submitted electronically in ST26 format and hereby incorporated by reference in its entirety. Said ST26 file, created on July 13, 2024, is name 1133087WO1.xml and is 20,534 bytes in size. BACKGROUND Dynamic protein complexes and transient protein-protein interactions (PPI) are integral for the vast majority of normal and cancer associated processes such as cellular metabolism, signal transduction networks and regulation of chromatin structure. Current methods cannot fully capture this complexity due to an inability to simultaneously detect multiple dynamic interactions. While current tools to study protein-protein interactions (PPI; e.g., co- immunoprecipitation, proximity ligation assay, two-hybrid and affinity purification–mass spectrometry (AP-MS) may offer means for studying networks of interactors in an unbiased manner, they have several key limitations. These limitations include the need for specialized instrumentation (e.g., MS), genetically engineered cell models (e.g., for protein tagging) and large input requirements (i.e., many cells and/or high amounts of proteins). Thus, most of the current work is focused on a single or a handful of factors, and do not address the multi-factorial manner in which they affect biology. SUMMARY One embodiment provides a DNA-caliper comprising two single-stranded oligonucleotide arms, wherein the arms are joined at their 5’ ends resulting in each arm having a free 3’ end, wherein each 3’ end comprises a single stranded oligonucleotide anchor’ region, wherein each arm comprises a primer region, wherein at least one arm of the caliper comprises a cleavage site at its 5’ end. In one embodiment, at least one arm of the DNA caliper further comprises a capture molecule. In another embodiment, the capture molecule comprises biotin or digoxigenin. In one embodiment, the single stranded oligonucleotide anchor’ region is complementary to a single stranded oligonucleotide anchor region on another oligonucleotide. In one embodiment, each anchor’ region has the same sequence. In other embodiment, each anchor’ region has a different sequence. In one embodiment, each arm is about 15 to about 1000 nucleotides long. In one embodiment, each single stranded oligonucleotide anchor’ is about 10 to about 50 nucleotides long. In one embodiment, each primer region is about 5 to about 45 nucleotides long. In one embodiment, the cleavage site is a restriction enzyme site, a deoxy uridine (dU) or a UV light cleavage site. One embodiment provides a composition comprising the DNA-caliper described herein. In one embodiment, the composition comprises a carrier. One embodiment provides an antibody covalently linked to a 5’ end of a single-stranded oligonucleotide tail, wherein the oligonucleotide tail comprises at its 3’ end a single stranded oligonucleotide anchor region, a unique barcode, and a single stranded oligonucleotide unique molecular identifier (UMI) and at its 5’ end a restriction enzyme site. In one embodiment, the single stranded oligonucleotide anchor region is complementary to a single stranded oligonucleotide anchor’ region on another molecule. In one embodiment, the barcode is specific to the antibody. In one embodiment, the UMI is a random sequence. In one embodiment, the UMI is about 5 to about 45 nucleotides long. In one embodiment, the single-stranded oligonucleotide tail is about 15 to about 1000 nucleotides long. In one embodiment, the anchor is about 10 to about 50 nucleotides long. One embodiment provides a composition comprising the antibody-oligo described herein. In one embodiment, the composition comprises a carrier. One embodiment provides a composition comprising the DNA-caliper described herein and the antibody-oligo described herein. In one embodiment, the composition comprises a carrier. One embodiment provides a method for mapping the genomic co-localization of one or more proteins on a chromatin fragment comprising: (a) incubating chromatin fragments with a plurality of antibodies, each of the plurality of antibodies binding to each of the one or more proteins, wherein each antibody comprises a single stranded oligonucleotide tail comprising a hook region, a single stranded oligonucleotide anchor region, a unique barcode, and a single stranded oligonucleotide unique molecular identifier (UMI); and wherein each chromatin fragment comprises a chromatin DNA molecule ligated to adaptors comprising a chromatin DNA molecule anchor’ region and a chromatin DNA molecule UMI; (b) contacting the chromatin fragments with a DNA-caliper molecule comprising two single-stranded oligonucleotide arms, wherein the arms are joined at their 5’ ends resulting in each arm having a free 3’ end, wherein the first 3’ end comprises a first sequence complementary to the single stranded oligonucleotide anchor region and the second 3’ end comprises a second sequence complementary to the chromatin DNA molecule anchor’ region, wherein at least one arm comprises a cleavage site and/or a capture molecule, and wherein each arm comprises a primer sequence; (c) hybridizing the DNA-caliper first sequence with the single stranded oligonucleotide anchor region and the DNA-caliper second sequence with the chromatin DNA molecule anchor’ region to obtain a hybridized product; (d) contacting the hybridized product with a DNA polymerase to obtain an extended DNA-caliper that covalently captures the sequences of the antibody unique barcode, the antibody UMI, and the chromatin DNA molecule. One embodiment further comprises denaturing the extended DNA- caliper from the chromatin fragment. In one embodiment, the capture molecule comprises biotin or digoxigenin. In one embodiment, the extended DNA-calipers are captured/isolated by the capture molecule. One embodiment further comprises adding a polyC sequence with terminal transferase to the 3’ end of the extended DNA-caliper corresponding to the chromatin fragment and then hybridizing and ligating a double stranded DNA molecule with two overhanging regions to the hook region and the polyC overhang on the extended DNA-caliper, bringing the two ends of the extended DNA-caliper together, wherein the 3’ ends of DNA-caliper are extended with polymerase using the other arm as a template to generate a double stranded construct. In one embodiment, the cleavage site is a restriction enzyme site, dU or UV light cleavage site. One embodiment further comprises linearizing the double stranded construct at the cleavage site to generate a linearized fragment. In one embodiment, the linearized fragment is PCR amplified using the primer sequences on the DNA-caliper. In one embodiment, the PCR amplified products are sequenced. In one embodiment, the single stranded oligonucleotide has a nucleic acid sequence of SEQ ID NO: 1. In one embodiment, the DNA-caliper has a nucleic acid sequence of SEQ ID NOs: 2 and 3. In one embodiment, the single stranded oligonucleotide has a nucleic acid sequence of SEQ ID NO: 4. In one embodiment, the DNA-caliper has a nucleic acid sequence of SEQ ID NOs: 5 and 6. One embodiment provides a method for detecting and/or quantifying at least one of multiple protein-protein interactions (PPIs) and/or protein abundance comprising: (a) incubating fixed and permeabilized cells with a plurality of antibody molecules, wherein each antibody molecule is covalently linked to a single stranded oligonucleotide tail comprising an antibody- specific unique barcode, a single stranded oligonucleotide unique molecular identifier (UMI), a restriction enzyme site, a single stranded oligonucleotide anchor’ region complementary to both arms of a DNA-caliper anchor region and optionally a hook region; and (b) contacting the plurality of antibody molecules with one or more single stranded oligonucleotide DNA-caliper molecules comprising two single-stranded oligonucleotide arms, wherein the arms are joined at their 5’ ends resulting in each arm having a free 3’ end, a capture molecule, a cleavage site, a first primer and a second primer sequence, wherein one primer sequence is on each side of the cleavage site, wherein the first 3’ end of the caliper comprises an anchor region that is complementary to the single stranded oligonucleotide anchor’ region of a first antibody and the second 3’ end of the caliper comprises an anchor region that is complementary to the single stranded oligonucleotide anchor’ region of a second antibody; (c) hybridizing the DNA-caliper with the single stranded oligonucleotide anchor’ region of the first antibody and the single stranded oligonucleotide anchor’ region of the second antibody to obtain a hybridized product; (d) contacting the hybridized product with a DNA polymerase to obtain a bidirectionally extended DNA-caliper that covalently links the sequences of the first antibody unique barcode and first antibody UMI and the second antibody unique barcode and second antibody UMI. In one embodiment, the permeabilized cells are lysed cells. In one embodiment, the cleavage site is a restriction endonuclease site, dU or UV light cleavage site. One embodiment further comprises denaturing the extended DNA-calipers from the single stranded oligonucleotide of the first antibody and the single stranded oligonucleotide of the second antibody or prior to or after adding DNA polymerase remove the antibodies by restriction enzyme cutting and ligating the ends; after DNA polymerase extension of the DNA-caliper ends and the oligonucleotide tails the extended molecule can be linearized, PCR amplified and sequenced. In one embodiment, the capture molecule comprises biotin or digoxigenin. In one embodiment, the extended DNA-calipers are captured/isolated by the capture molecule. One embodiment further comprises hybridizing and ligating a double stranded DNA molecule with two overhanging regions to the hook regions on the extended DNA-caliper bringing the two ends of the DNA-caliper together, wherein the 3’ ends of DNA-caliper are extended with DNA polymerase using the other arm as a template to generate a double stranded construct. One embodiment further comprises linearizing the double stranded construct at the cleavage site to generate a linearized fragment. In one embodiment, the linearized fragment is PCR amplified using the primer sequences on the DNA-caliper. In one embodiment, the PCR amplified products are sequenced so as to identify the PPIs and their abundance through UMI counting. One embodiment provides a method of determining abundance of a target protein in a sample (PQ-seq) comprising: a) contacting a sample suspected of comprising one or more target proteins with one or more antibodies that binds to at least one of the target proteins in the sample, wherein each of the one or more antibodies comprises a single stranded oligonucleotide tail; b) detecting, isolating and PCR amplifying the single stranded oligonucleotide of the antibody bound to the protein, wherein the single stranded oligonucleotide tail is a template oligonucleotide to generate a first PCR product comprising a double stranded version of the single stranded oligonucleotide tail; performing a second PCR amplification using the first PCR product as a template, wherein the second PCR generates a second PCR product; sequencing the second PCR product; and determining the abundance of the single stranded oligonucleotide sequence indicating protein abundance of the target protein in the sample. In one embodiment, the single stranded oligonucleotide tail comprises one or more unique molecular identifier (UMI) sequences. In one embodiment, the single stranded oligonucleotide tail comprises a barcode sequence or molecular tag. In one embodiment, the abundance of the target protein comprises detecting the presence of the UMI sequence. One embodiment further comprises separating the proteins in the sample by gel electrophoresis and transferring the separated proteins to a membrane prior to contacting the sample with the one or more antibodies. One embodiment further comprises contacting the one or more antibodies comprising one or more single stranded oligonucleotide tails with a single stranded binding protein prior to incubating the membrane with the one or more antibodies. One embodiment further comprises a protein ladder of known size. One embodiment further comprises contacting the membrane with one or more antibodies and cutting the membrane into segments. In one embodiment, the protein ladder of known size is used such that each segment contains about a 10kDa to about a 20kDa range of sizes of protein. In one embodiment, the abundance of one or more proteins in the sample is performed by comparing the relative abundance of the one or more UMI sequences from the sequencing results, wherein the one or more UMI sequences are specific to the antibody that targets a protein. In one embodiment, the sample is sonicated prior to gel electrophoresis. One embodiment provides a kit comprising: one or more primary antibodies against one or more target proteins, wherein the one or more primary antibodies comprise a single stranded oligonucleotide tail and wherein the single stranded oligonucleotide tail comprises a primer region, an antibody-specific unique barcode, a single stranded oligonucleotide unique molecular identifier (UMI), a 3’ single stranded oligonucleotide anchor’ region, a restriction enzyme site and optionally a single stranded binding protein; and instructions for use thereof for methods of determining protein abundance in a sample. One embodiment further comprises one or more single stranded oligonucleotide DNA-caliper molecules comprising two single-stranded oligonucleotide arms, wherein the arms are joined at their 5’ ends resulting in each arm having a free 3’ end, a capture molecule, a cleavage site, a first primer and a second primer sequence, wherein one primer sequence is on each side of the cleavage site. One embodiment further comprises a tool for cutting a protein membrane, such as a pair of scissors, blade, knife or cookie cutter type device. One embodiment further comprises a first set of PCR primers and a second set of PCR primers, wherein the first set of PCR primers comprises a first sequence and are formulated for use in a first PCR reaction that converts the single stranded oligonucleotide to a double stranded oligonucleotide, and wherein the second set of PCR primers comprises a second sequence and are formulated for use in a second PCR reaction that generates double stranded oligonucleotides for sequencing. One embodiment provides a method comprising: a) perform PQ-seq on the sample and collect protein quantification data, b) perform Prod-seq on the sample and collect protein-protein interaction data, c) determining, by a computing system having processing resources and memory, a true interaction using the following equation: ^^ ∼ ^^^^^^^^ ^^^^^^^^ ^^^ℎ ^^^^ ^^^ ^^^ + ^^ ^ ^^^ ^^^^^^^^^^ ^^^^^^^^^ ^ !
Figure imgf000007_0001
Figure imgf000007_0002
λ23 = ^45%67 ^α2 + ^2 ^3 + ^3) + ^6%:4;<α2α3 wherein ^ >5? as the
Figure imgf000007_0003
wherein +, denotes the proportion of true interactions in the entire dataset; wherein ^^^ is the readout scaling factor for PQ-seq detection protocol and sequencing; ^6%:4;< and ^45%67 denote the readout scaling factors for specific and non-specific PPI detection in Prod- seq, respectively; ^% and ^% denote the levels of specific and non-specific antibody-oligo binding
Figure imgf000007_0004
for each protein target ^, respectively, wherein +,, ^^^, ^6%:4;< , ^45%67 , ^A,...,4 , ^A,...,4 , ^ !, ^ >5? are estimated from the maximum likelihood estimation of "^$, ^| ⋅^. In one embodiment, after fitting the maximum likelihood model, the model calculates the confidence score of each PPI as Score23 = FG ^ X Iλ P#true'X ) = H 23 23 J 23 FG^HX23Iλ23JK^ALFG^^HX23Iθ23J . DESCRIPTIION OF THE DRAWINGS The drawings illustrate generally, by way of example, but not by way of limitation, various embodiments discussed herein. FIG.1. Overview of the framework outputs. Illustration of the interactions captured across different conditions, by the two tools provided herein, with Prod-seq (capture of PPIs) depicted as on the left cube, WhIP-seq (capture of protein-DNA interactions) depicted in the center cube, and other tools, e.g., for Protein-RNA interactions, depicted in the third cube. Each protein, genomic locus and condition (e.g., a differentiation time point) is represented as a dimension within a 3D matrix, where the tools are represented by horizontal shaded cubes. To convey the simultaneous identification of interactions, the axes in the cubes represent the array of studied proteins (1–n), or genomic loci (1–n) and different conditions (1–n). Two examples are shown: interacting proteins illustrate PPI between proteins(c,k) under condition(h) (e.g., EZH2 and SUZ12 in hiPSCs); a DNA fragment with multiple proteins bound to it depicts protein-DNA interactions between proteins(b,d,g,h,j) and genomic locus(h) under condition(a). Note, while not represented, the abundance of each protein studied is also measured. FIGS.2A-2B. Steps of Prod-seq and WhIP-seq. Simplified depiction of the tools A. Prod- seq: (1) Fixed and lysed cells are incubated with multiple Ab-oligos, each antibody is linked to an oligonucleotide with a barcode (BC n, m, etc.), a UMI and an anchor site (Anchor1). At each 3’ end of the DNA-caliper there is an Anchor1’. (2) Hybridization of the DNA-caliper to the Anchor1 (Anc1) on each Ab-oligo captures the barcodes and UMIs by DNA extension (dashed) at both ends. (3) The extended DNA-calipers are converted into a sequencing library. B. WhIP-seq: (1) Sheared chromatin is incubated with multiple Ab-oligos, these are the same ones used for Prod- seq. The chromatin fragments (Genomic Region) are ligated to WhIP-adaptors that include a UMI and a different anchor (Anchor 2). The DNA-caliper used here contains Anchor2’ on one arm. (2). Target sequences are captured by hybridization of the two DNA-caliper anchors (Anc1’ and Anc2’) to the Ab-oligo and the adaptor, followed by DNA extension (dashed). (3) Extended DNA-calipers are converted into a sequencing library. FIG.2C. One embodiment of generation sequence ready constructs continued from Figure 2A is depicted (another embodiment is depicted in Figure 6). The two ends of the extended DNA- caliper are brought together for intramolecular hybridization of the Hook and Hook’. The free 3’ ends of each of the DNA-caliper arms are extended using the other arm as a template. Finally, the double stranded fragments are used as a template for generation of a sequencing ready construct by PCR from the primer sequences (P1 and P2) embedded in the original DNA-caliper. Different DNA barcodes are also depicted. FIG. 2D. Conversion of the Prod-seq hybridization-extension products into a sequence library. Continued from FIG. 2A, the extended DNA-caliper is used as a template for DNA polymerase to generate a new normal strand (dashed) with P1 (embedded in the oligonucleotides linked to the antibodies) as a primer (bi-directionally extended DNA-caliper). Next, the new strand, isolated by denaturation and removal of the extended DNA-caliper by biotin pull-down is used as a template for generation of a sequencing ready library by PCR using the primer sequences (P1 and P2). Different DNA barcodes are depicted. FIG. 2E. Conversion of the WhIP-seq hybridization extension products into a sequence library (another embodiment is depicted in Figures 7 and 8). Continued from Figure 2B: The extended DNA-caliper is used as a template for DNA-polymerase to generate a new, normal strand (dash line) with P1 (embedded in the oligonucleotides linked to the antibodies as a primer). Next, the new strand is isolated by denaturation and removal of the extended DNA-caliper by biotin pull- down, and 3’ modified by terminal transferase to have a stretch of oligo-C. A primer (P2) that ends with a stretch of oligo-G locked nucleic acids is used, together with P1, to generate a sequencing library by PCR. Various DNA barcodes are also depicted. FIGS.3A-3B: Ab-oligo based protein abundance measurements. A. Western blot showing the reduction in H3K27me3 in EZH2 KO HEK293 cells compared to WT (~5-fold). Adapted from Wang et al.1 B. Optimization of Ab-oligo signal-to-noise ratio by preincubating with different single-stranded DNA binding proteins (SSB/ETSSB) at different temperatures to minimize non- specific Ab-oligo interactions. Black line denotes mean of UMI count ratio (KO=0.67; WT=7.5; ~11-fold mean reduction). FIG.4. Prod-seq captures PPIs. A multiplexed pool of 6 Ab-oligos (H3K4me3, H3k27me3, H3k27as, EZH2, EGFR, and P-Tyr). The percent of UMI pairs associated with each pair of epitopes after removing all self-self-epitope interactions detected. FIGS.5A-5C. Ligation of adaptors to cross-linked chromatin. (A) Cross-linked chromatin was immobilized with a H3K27ac antibody and ligated to Illumina adaptors, and the product was then amplified by PCR. From left to right: ladder; ligated chromatin; original sheared chromatin. (B) A representative view from the IGV browser (37) showing the similarity between data generated by chromatin ligation (black) and by ChIP-seq (red) from the same mouse ES cells. (C) Genome-wide analysis showing high correlation (R2=0.75) between read coverage in peak regions as generated by chromatin ligation and ChIP-seq. FIGS. 6A-6H. Interrogation of the protein interactome and abundance by Prod-seq. Illustration of steps (not to scale). a. Fixed and lysed cells are incubated with multiple Ab-oligos. The oligonucleotide tail contains 4 elements: an anchor sequence (Anchor 1) complementary to both arms of the DNA-caliper; a UMI; an antibody-specific unique barcode (BC n, m, etc.), and Hook. The ends of each oligo are non-extendable (labeled by a T block). The DNA-caliper is depicted at the bottom and includes biotinylated arms (circled B) and two identical Anchor 1’. Target proteins are removed from next steps for simplicity. b. The barcodes and UMIs are captured by hybridization of the Anchor 1’ (Anc1’) on the DNA-caliper arms to the Anchor 1 (Anc1) on each Ab-oligo, followed by DNA polymerase extension (dashed lines) at both ends. c-d. Next, the extended DNA-calipers are denatured and pulled-down using the biotin. A Splint molecule with two overhanging Hook’ regions (c, top) is hybridized and ligated to the extended DNA-caliper. e. Following intramolecular hybridization, the 3’ ends of each of the DNA-caliper arms are extended using the other arm as a template. f. The molecule is linearized by USER enzyme that cleaves the dU and the fragments are used as a template for generation of sequencing-ready constructs by PCR from the primer sequences (P1 and P2; The dU and the primers are embedded in the original DNA- caliper, panel a.). e. Another embodiment of Prod-seq. The diverse DNA barcodes that are captured by the process are depicted by different color combinations.h. Prod-seq, capture PPI by sequencing. FIGS.7A-7F. WhIP-seq: Mapping the genomic co-localization and abundance of multiple proteins. a. Sheared chromatin is incubated with multiple Ab-oligos, these are the same ones used for Prod-seq (FIG. 7), as the oligonucleotide tail has an identical structure (Anchor 1, a UMI, a barcode (BC m, k…n) and a Hook’). The chromatin fragments (Genomic Region) are ligated to adaptors (WhIP-adaptors) that include a UMI and a different anchor (Anchor 2). Each arm of the DNA-caliper contains a specific primer sequence, a biotin (B in a circle) and, at each 3’ end, sequences complementary to the Anchors it targets. b. Target sequences are captured by hybridization of the two DNA-caliper anchors (Anc1’ and Anc2’) to the Ab-oligos and the adaptor 3’ overhang, followed by DNA extension (dashed lines) at both ends. c. The extended DNA- calipers are denatured and pulled-down using the Biotin moiety. d-f. Conversion of the extended DNA-calipers into sequencing-ready constructs via intramolecular hybridization-extension. d. A Splint molecule (left), with 5’ phosphorylated termini is added. Note, this Splint has one arm with Hook’ to match the Ab-oligo, while the other has a Poly(G) overhang at the 3’ end. Both termini end with cleavable extension-blockers. The other end of the extended DNA-caliper is 3’ modified by terminal transferase to have a Poly(C) stretch. The cleavable blockers are removed, and the Splint is ligated to the extended DNA-caliper. e. Following intramolecular hybridization, the free 3’ ends of the DNA-caliper arms are extended using the other arm as a template followed by linearization of the molecule by USER enzyme to cleave the dU and PCR using the P1 and P3. primers f. Different color combinations depict the diverse DNA barcodes; the genomic regions (Genomic) are represented in purple. FIG. 8. Capturing the structural details of the product of Prod-seq by paired end sequencing. The length of the fragment is ~240bp. read1 (dashed black) will capture primer 1 (P1), Anchor1 (Anc1), the UMI (that is followed by a constant region) the barcode (BC) of the first antibody and the Hook or part of it. read2 captures primer 2 (P2), the, Anchor 1 (Anc1), the UMI that is followed by a constant region, the barcode (BC) of the second antibody and the Hook’. FIGS.9A-9G. provide an overview of Western-seq. Multiplex protein detection can occur via conjugated antibodies, hybridization and strand extension (FIGS. 9A-9B), while barcoding allows detection of different proteins in one reaction (FIGS. 9C-9D). FIG. 9D depicts protein quantification followed by sequencing (PQ-seq). PQ-seq is a protein quantification method using DNA oligonucleotides conjugated to antibodies and ending with sequencing to get the amount and proportion of various proteins in a sample. Western-seq allows multiplex protein detection with size separation (FIG. 9E). FIG. 9F provides an exemplary protocol for Western-seq. FIG. 9G demonstrates more specific binding at expected section. FIGS.10A-10B, for illustration purposes only, provide two example designs for Ab-oligos and DNA-calipers with complementary anchor sequences, hook sequences, UMI, barcode, and splint sequences. (SEQ ID NOS: 10-13) FIGS. 11A-11C. PCR and library preparation efficiency improvement following linearization of the product by restriction digestion. a. Left (green arrows), PCR (2 replicates of 2 different products) linearized by EcoRV digestion; right (red arrows) same PCR reaction (same cycle number) without linearization. b. Library quality presented as precent of unique (orange) vs duplicated (blue) reads measured by UMIs. c. Demonstration of cleaving the DNA-caliper by UV light – left: untreated DNA-caliper, right: UV light treated DNA-caliper. FIGS. 12A-12C. Improvement of the process for converting the extended DNA-caliper into a sequencing library by replacing splint/whip-adaptor with designed removal of the antibody- oligo components. a. Illustration of the modified library preparation process. The original oligonucleotides conjugated to the antibody include extendable ends, allowing the Klenow reaction to extend both the DNA-caliper and the Ab-oligo arms. Here, instead of the original approach that uses the splint/whip-adaptor, a recognition site was added for the restriction enzyme (XmnI) on the Ab-oligo arm. The enzyme cannot cut single stranded DNA, so only fragments that were extended by the Klenow will be cut, leaving behind a blunt end. Next, the digested fragments are ligated (intramolecularly). b. The cut and ligate process improves the library preparation process. Prod-seq was performed using either the approach that employs a splint (left 4 samples) or the one detailed here with the restriction enzyme site. The ratio of unique UMI pairs increased by up to ~20 fold. c. Identification of optimal reaction conditions using “mimic” reactions. Oligonucleotides that imitate the output of the reaction were used to evaluate multiple conditions and designs, such as different restriction enzymes (SbfI, NotI, BbsI – sticky; XmnI, blunt); different types of Klenow (exo+ or exo-) and single or nested PCR. Under these mimic conditions, single PCR, exo- and the blunt enzymes show the best ability to capture the correct structure. FIGS.13A-13C. Evaluation of the PPI scoring models on example Prod-seq and PQ-seq dataset. Test results for the Poisson model and negative binomial model for PPI scoring on the example Prod-seq and PQ-seq dataset. a. The example dataset. Prod-seq and PQ-seq workflows are performed on recombinant H3K4me3/H3K27ac mono-nucleosomes. Numbers in heatmaps are UMI combination counts (Prod-seq) or UMI counts (PQ-seq). b. PPI confidence score output from the Poisson PPI scoring model. c. PPI confidence score output from the negative binomial PPI scoring model. FIG. 14 illustrates a diagrammatic representation of a machine 1400 in the form of a computer system within which a set of instructions may be executed for causing the machine 1400 to perform any one or more of the methodologies discussed herein. DESCRIPTION Reference will now be made in detail to certain embodiments of the disclosed subject matter. While the disclosed subject matter will be described in conjunction with the enumerated claims, it will be understood that the exemplified subject matter is not intended to limit the claims to the disclosed subject matter. The methods described herein provide a means to simultaneously identify interactions between cellular entities. These advances, in turn, will push the current boundaries of several fields including cell biology and metabolism, epigenomics and proteomics, as well as studies and diagnostics of complex diseases such as multiple types of cancer, psychiatric/neurodevelopmental disorders and metabolic diseases, where chromatin structure and PPI are playing roles. Definitions The following definitions are included to provide a clear and consistent understanding of the specification and claims. As used herein, the recited terms have the following meanings. All other terms and phrases used in this specification have their ordinary meanings as one of skill in the art would understand. Such ordinary meanings may be obtained by reference to technical dictionaries, such as Hawley's Condensed Chemical Dictionary 14th Edition, by R.J. Lewis, John Wiley & Sons, New York, N.Y., 2001. References in the specification to "one embodiment," "an embodiment," etc., indicate that the embodiment described may include a particular aspect, feature, structure, moiety, or characteristic, but not every embodiment necessarily includes that aspect, feature, structure, moiety, or characteristic. Moreover, such phrases may, but do not necessarily, refer to the same embodiment referred to in other portions of the specification. Further, when a particular aspect, feature, structure, moiety, or characteristic is described in connection with an embodiment, it is within the knowledge of one skilled in the art to affect or connect such aspect, feature, structure, moiety, or characteristic with other embodiments, whether or not explicitly described. The singular forms "a," "an," and "the" include plural reference unless the context clearly dictates otherwise. Thus, for example, a reference to "a compound" includes a plurality of such compounds, so that a compound X includes a plurality of compounds X. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for the use of exclusive terminology, such as "solely," "only," and the like, in connection with any element described herein, and/or the recitation of claim elements or use of "negative" limitations. The term "and/or" means any one of the items, any combination of the items, or all of the items with which this term is associated. The phrase "one or more" is readily understood by one of skill in the art, particularly when read in context of its usage. For example, one or more substituents on a phenyl ring refers to one to five, or one to four, for example if the phenyl ring is di-substituted. As used herein, “or” should be understood to have the same meaning as “and/or” as defined above. For example, when separating a listing of items, “and/or” or “or” shall be interpreted as being inclusive, e.g., the inclusion of at least one, but also including more than one of a number of items, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of” or “exactly one of,” or, when used in the claims, “consisting of,” will refer to the inclusion of exactly one element of a number or list of elements. In general, the term “or” as used herein shall only be interpreted as indicating exclusive alternatives (i.e., “one or the other but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of,” or “exactly one of.” As used herein, the terms “including,” “includes,” “having,” “has,” “with,” or variants thereof, are intended to be inclusive similar to the term “comprising.” The term "about" can refer to a variation of ± 5%, ± 10%, ± 20%, or ± 25% of the value specified. For example, "about 50" percent can in some embodiments carry a variation from 45 to 55 percent. For integer ranges, the term "about" can include one or two integers greater than and/or less than a recited integer at each end of the range. Unless indicated otherwise herein, the term "about" is intended to include values, e.g., weight percentages, proximate to the recited range that are equivalent in terms of the functionality of the individual ingredient, the composition, or the embodiment. The term about can also modify the endpoints of a recited range as discuss above in this paragraph. As will be understood by the skilled artisan, all numbers, including those expressing quantities of ingredients, properties such as molecular weight, reaction conditions, and so forth, are approximations and are understood as being optionally modified in all instances by the term "about." These values can vary depending upon the desired properties sought to be obtained by those skilled in the art utilizing the teachings of the descriptions herein. It is also understood that such values inherently contain variability necessarily resulting from the standard deviations found in their respective testing measurements. As will be understood by one skilled in the art, for any and all purposes, particularly in terms of providing a written description, all ranges recited herein also encompass any and all possible sub-ranges and combinations of sub-ranges thereof, as well as the individual values making up the range, particularly integer values. A recited range (e.g., weight percentages or carbon groups) includes each specific value, integer, decimal, or identity within the range. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, or tenths. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, etc. As will also be understood by one skilled in the art, all language such as "up to," "at least," "greater than," "less than," "more than," "or more," and the like, include the number recited and such terms refer to ranges that can be subsequently broken down into sub-ranges as discussed above. In the same manner, all ratios recited herein also include all sub-ratios falling within the broader ratio. Accordingly, specific values recited for radicals, substituents, and ranges, are for illustration only; they do not exclude other defined values or other values within defined ranges for radicals and substituents. One skilled in the art will also readily recognize that where members are grouped together in a common manner, such as in a Markush group, the invention encompasses not only the entire group listed as a whole, but each member of the group individually and all possible subgroups of the main group. Additionally, for all purposes, the invention encompasses not only the main group, but also the main group absent one or more of the group members. The invention therefore envisages the explicit exclusion of any one or more of members of a recited group. Accordingly, provisos may apply to any of the disclosed categories or embodiments whereby any one or more of the recited elements, species, or embodiments, may be excluded from such categories or embodiments, for example, for use in an explicit negative limitation. The term "contacting" refers to the act of touching, making contact, or of bringing to immediate or close proximity, including at the cellular or molecular level, for example, to bring about a physiological reaction, a chemical reaction, or a physical change, e.g., in a solution, in a reaction mixture, in vitro, or in vivo. The use of the word “detect” and its grammatical variants refers to measurement of the species without quantification, whereas use of the word “determine” or “measure” with their grammatical variants are meant to refer to measurement of the species with quantification. The terms “detect” and “identify” are used interchangeably herein. As used herein, the term “nucleic acid” encompasses RNA as well as single and double stranded DNA and cDNA. Furthermore, the terms, “nucleic acid,” “DNA,” “RNA” and similar terms also include nucleic acid analogs, i.e., analogs having other than a phosphodiester backbone. For example, the so called “peptide nucleic acids,” which are known in the art and have peptide bonds instead of phosphodiester bonds in the backbone, are considered within the scope of the present invention. By “nucleic acid” is meant any nucleic acid, whether composed of deoxyribonucleosides or ribonucleosides, and whether composed of phosphodiester linkages or modified linkages such as phosphotriester, phosphoramidate, siloxane, carbonate, carboxymethylester, acetamidate, carbamate, thioether, bridged phosphoramidate, bridged methylene phosphonate, bridged phosphoramidate, bridged phosphoramidate, bridged methylene phosphonate, phosphorothioate, methylphosphonate, phosphorodithioate, bridged phosphorothioate or sulfone linkages, and combinations of such linkages. The term nucleic acid also specifically includes nucleic acids composed of bases other than the five biologically occurring bases (adenine, guanine, thymine, cytosine, and uracil). Conventional notation is used herein to describe polynucleotide sequences: the left-hand end of a single-stranded polynucleotide sequence is the 5’-end; the left-hand direction of a double-stranded polynucleotide sequence is referred to as the 5’-direction. The direction of 5’ to 3’ addition of nucleotides to nascent RNA transcripts is referred to as the transcription direction. The DNA strand having the same sequence as an mRNA is referred to as the “coding strand”; sequences on the DNA strand which are located 5’ to a reference point on the DNA are referred to as “upstream sequences”; sequences on the DNA strand which are 3’ to a reference point on the DNA are referred to as “downstream sequences.” The term “oligonucleotide” typically refers to short polynucleotides, generally, no greater than about 100 nucleotides. It will be understood that when a nucleotide sequence is represented by a DNA sequence (i.e., A, T, G, C), this also includes an RNA sequence (i.e., A, U, G, C) in which “U” replaces “T.” Define “carrier” includes a liquid, such as water, saline solution or other pH balanced liquids, as well as encapsulating agents (e.g., liposomes), and other materials. The tools (Ab-oligos, DNA-calipers etc.) provided herein can be formulated in dry form (e.g., in freeze-dried form), and resuspended with in a convenient liquid. The term “standard,” as used herein, refers to something used for comparison. For example, it can be a known standard agent or compound which is administered and used for comparing results when administering a test compound, or it can be a standard parameter or function which is measured to obtain a control value when measuring an effect of an agent or compound on a parameter or function. Standard can also refer to an “internal standard”, such as an agent or compound which is added at known amounts to a sample and is useful in determining such things as purification or recovery rates when a sample is processed or subjected to purification or extraction procedures before a marker of interest is measured. Internal standards are often a purified marker of interest which has been labeled, such as with a radioactive isotope, allowing it to be distinguished from an endogenous marker. As used herein, the term “click reaction” is recognized in the art, which describe a collection of reliable and self-directed organic reactions, such as the most recognized copper catalyzed azide-alkyne [3+2] cycloaddition. Non-limiting examples of click chemistry reactions can be found, for example, in H. C. Kolb, M. G. Finn, K. B. Sharpless, Angew. Chem. Int. Ed. 2001, 40, 2004 and E. M. Sletten, C. R. Bertozzi, Angew. Chem. Int. Ed. 2009, 48, 6974, the disclosures of which are herein incorporated by reference in their entireties for all purposes. As used herein, the term “cross-linking agent” is used to describe a compound that is capable of forming a chemical bond between molecular groups on similar or dissimilar molecules so as to covalently bond together the molecules. Examples of common cross-linking agents are known in the art. See, for example, Bioconjugate Techniques (Academic Press, New York, 1996 or later versions) the content of which is herein incorporated by reference in its entirety for all purposes. Indirect attachment of the biomolecule to polymer dots can occur through the use of “linker” molecule, for example, avidin, streptavidin, neutravidin, biotin or a like molecule. As used herein, a "cell" refers to any type of cell isolated from a prokaryotic, eukaryotic, or archaeon organism, including bacteria, archaea, fungi, protists, plants, and animals, including cells from tissues, organs, and biopsies, as well as recombinant cells, cells from cell lines cultured in vitro, and cellular fragments, cell components, or organelles comprising nucleic acids. The term also encompasses artificial cells, such as nanoparticles, liposomes, polymersomes, or microcapsules encapsulating nucleic acids. The methods described herein can be performed, for example, on a sample comprising a single cell or a population of cells. The term also includes genetically modified cells. In embodiments, biological samples can comprise any commercially available cells, such as HeLa, HEK293, or stem cells, and/or lysates of such cells. The terms "hybridize" and "hybridization" refer to the formation of complexes between nucleotide sequences which are sufficiently complementary to form complexes via Watson Crick base pairing. As used herein, the term “label” or a detectable label intends a directly or indirectly detectable compound or composition that is conjugated directly or indirectly to the composition to be detected, e.g., N-terminal histidine tags (N-His), magnetically active isotopes, e.g., 115Sn, 117Sn and 119Sn, a non-radioactive isotopes such as 13C and 15N, polynucleotide or protein such as an antibody so as to generate a “labeled” composition. The term also includes sequences conjugated to the polynucleotide that will provide a signal upon expression of the inserted sequences, such as green fluorescent protein (GFP) and the like. The label may be detectable by itself (e.g., radioisotope labels or fluorescent labels) or, in the case of an enzymatic label, may catalyze chemical alteration of a substrate compound or composition which is detectable. The labels can be suitable for small scale detection or more suitable for high-throughput screening. As such, suitable labels include, but are not limited to magnetically active isotopes, non-radioactive isotopes, radioisotopes, fluorochromes, chemiluminescent compounds, dyes, and proteins, including enzymes. The label may be simply detected, or it may be quantified. A response that is simply detected generally comprises a response whose existence merely is confirmed, whereas a response that is quantified generally comprises a response having a quantifiable (e.g., numerically reportable) value such as an intensity, polarization, and/or other property. In luminescence or fluorescence assays, the detectable response may be generated directly using a luminophore or fluorophore associated with an assay component actually involved in binding, or indirectly using a luminophore or fluorophore associated with another (e.g., reporter or indicator) component. Examples of luminescent labels that produce signals include but are not limited to bioluminescence and chemiluminescence. Detectable luminescence response generally comprises a change in, or an occurrence of a luminescence signal. Suitable methods and luminophores for luminescently labeling assay components are known in the art and described for example in Haugland, Richard P. (1996) Handbook of Fluorescent Probes and Research Chemicals (6th ed). Examples of luminescent probes include, but are not limited to, aequorin and luciferases. Examples of suitable fluorescent labels include, but are not limited to, fluorescein, rhodamine, tetramethylrhodamine, eosin, erythrosin, coumarin, methyl-coumarins, pyrene, Malacite green, stilbene, Lucifer Yellow, Cascade Blue™, and Texas Red. Other suitable optical dyes are described in the Haugland, Richard P. (1996) Handbook of Fluorescent Probes and Research Chemicals (6th ed.). As used herein, a purification label or maker refers to a label that may be used in purifying the molecule or component that the label is conjugated to, such as an epitope tag (including but not limited to a Myc tag, a human influenza hemagglutinin (HA) tag, a FLAG tag), an affinity tag (including but not limited to a glutathione-S transferase (GST), a poly-Histidine (His) tag, Calmodulin Binding Protein (CBP), or Maltose-binding protein (MBP)), or a fluorescent tag. Methods involving conventional molecular biology techniques are described herein. Such techniques are generally known in the art and are described in detail in methodology treatises, such as Molecular Cloning: A Laboratory Manual, 2nd ed., vol.1-3, ed. Sambrook et al., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., 1989; and Current Protocols in Molecular Biology, ed. Ausubel et al., Greene Publishing and Wiley-Interscience, New York, 1992 (with periodic updates). Methods for chemical synthesis of nucleic acids are discussed, for example, in Beaucage and Carruthers, Tetra. Letts. 22: 1859-1862, 1981, and Matteucci et al., J. Am. Chem. Soc.103:3185, 1981. Prod-seq, WhIP-seq and Western-seq Although the activity of many dynamic protein complexes is highly dependent on interactions between molecules (e.g., protein‐protein and protein‐DNA), due to the limitations of current methods, most of the existing studies are focused on a single molecule at a time. One goal of the methods described herein is to develop an innovative cross‐disciplinary framework for studying these proteins as a class and use mouse embryonic stem cells, as an example, to study how multiple factors work together to regulate gene activity. Provide herein are methods and compositions for carrying out the processes of Prod-seq, WhIP-seq and Western-seq. The Prod-seq process proceeds as follows (FIG.2A): (i) Fixed and lysed cells are incubated with multiple antibodies, where each specific antibody is covalently linked to a single-stranded oligonucleotide tail (Ab-oligos). The oligonucleotide tail contains an anchor sequence (Anc) complementary to both arms of the DNA-caliper (Anc’; the DNA-caliper is depicted in FIG.2A); a unique molecular identifier (UMI; a short random sequence) and an antibody-specific unique barcode (Barcode n, m, etc.). Each arm of the DNA-caliper contains at each 3’ end a sequence that is complementary to the antibody anchors at their 3’ ends (Anc’). (ii) Target sequences are captured by hybridization of the two DNA-caliper anchors to the complementary anchor sequences on the oligonucleotide linked to the antibodies, followed by DNA polymerase extension at both ends (FIG.2A). This extension covalently links the barcodes and UMI sequences of the two antibodies. (iii) The extended DNA-calipers are then denatured and captured by, for example, chemical interactions with modifications on the DNA, including but not limited to, biotin or digoxigenin, pull-down, and then converted into sequencing-ready constructs (FIG. 2C). Sequencing the libraries allows identification of PPI instances as well as counting (using the UMIs). Further, the same Ab-oligos can also serve for quantitative detection of protein abundance (similar to other tools). FIG.6 offers a variation on Prod-seq in which a. Fixed and lysed cells are incubated with multiple Ab-oligos. The oligonucleotide tail contains: an anchor sequence (Anchor 1) complementary to both arms of the DNA-caliper; a UMI; an antibody-specific unique barcode (BC n, m, etc.), a restriction enzyme site, and a Hook region. The ends of each oligo are non-extendable (labeled by a T block). The DNA-caliper is depicted at the bottom and includes biotinylated arms (circled B) and two identical Anchor 1’. Target proteins are removed from next steps for simplicity. b. The barcodes and UMIs are captured by hybridization of the Anchor 1’ (Anc1’) on the DNA- caliper arms to the Anchor 1 (Anc1) on each Ab-oligo, followed by DNA polymerase extension (dashed lines) at both ends. c-d. Next, the extended DNA-calipers are denatured and pulled-down using the biotin. A Splint molecule with two overhanging Hook’ regions (c, top) is hybridized and ligated to the extended DNA-caliper. e. Following intramolecular hybridization, the 3’ ends of each of the DNA-caliper arms are extended using the other arm as a template. f. The molecule is linearized by USER enzyme that cleaves the dU and the fragments are used as a template for generation of sequencing-ready constructs by PCR from the primer sequences (P1 and P2; The dU and the primers are embedded in the original DNA-caliper, panel a.). The diverse DNA barcodes that are captured by the process are depicted. In a further embodiment, which applies to Prod-seq and WhIP-seq, Prod-seq and WhIP- seq use PCR to convert and amplify the product into, for example, a next generation sequencing (NGS) library. Thus, high efficiency of PCR is needed for the effectiveness of the workflow. The design that was disclosed in US Pat No. 10,655,162B1 was limited, as the double-stranded molecule that is the product of the Prod-seq or WhIP-seq process tends to reanneal to itself (the two strands were complimentary and connected at the 5’). This limited the ability of the PCR primers to anneal to the product, and highly reduced the efficiency of library preparation. To overcome this challenge, the novel design disclosed herein, incorporates a cleavage site (such as restriction enzyme site, dU or a photo-cleavable site) that separates the two 5’ ends following, for example, the Prod-seq reaction, which enables the release of the two strands while maintaining the proximity detection information intact. This approach is highly effective in increasing PCR efficiency and yield of the reaction to generate NGS libraries (FIGS.11A-11B). Further, to simplify the process (reduce efforts, time and costs) and avoid off-target cleavage by the restriction enzyme (e.g., of barcodes or UMIs), an alternative design of the DNA- caliper was made that employs a UV-light sensitive modification. Exposure to UV-light replaces the replaces the need for a restriction enzyme site at the cleavage site (FIG.11C). Previously the process included two rounds of PCR – the first introducing the Illumina adapter and the second adding Illumina indices. The design was modified to reduce the number of PCR reactions, and the primer site of the DNA-caliper is now directly compatible with the Illumina sequencing primers. This further simplifies the library preparation process (only one round of PCR needed), reduces efforts and improves the efficiency as the added PCR was prone to lead to loss of material and add noise. In a further method, the antibodies are removed through use of a restriction enzyme and the restriction enzyme site on the oligonucleotide tail. For example, for Prod-seq, to associate the barcodes and UMIs of the two proteins/molecules in proximity, a “splint” molecule (or a “whip- adaptor” as referenced in US Pat. No. 10,655,162, incorporated herein by reference), a partially double-stranded DNA, was used to facilitate the cyclization of the extended DNA caliper, which enables joining the two molecule barcodes within the same DNA sequence by a ligation reaction. However, the annealing of a molecule such as the splint or whip-adaptor is extremely inefficient, potentially the reaction may be outcompeted by the reannealing of the DNA-caliper back to the Ab-oligos. To overcome this constraint, a novel design was made that introduces a restriction enzyme recognition site (such as for XmnI; see example sequences below) at the sequence of the oligonucleotide that is conjugated to the antibodies. The extended DNA-caliper is simply digested with, for example, XmnI, the ends are then directly be ligated together (intra-molecular) without the need of splint or whip-adaptor annealing. This approach was shown to significantly improve the efficiency of the workflow and the quality of the Prod-seq products (FIGS. 12-A-12C). The WhIP-seq method is described schematically in FIGS.2B and 7. Sheared chromatin is incubated with multiple Ab-oligos. The Ab-oligos can be the same ones used for Prod-seq (FIG. 2A), as the oligonucleotide tail has an identical structure (Anchor 1, a UMI, a barcode (BC m, k…n) and a Hook’). The chromatin fragments (Genomic Region) are ligated to adaptors (WhIP- adaptors) that include a UMI and a different anchor (Anchor 2). Each arm of the DNA-caliper contains a specific primer sequence, a capture molecule (such as biotin (B in a circle)), a cleavage site and, at each 3’ end, sequences complementary to the Anchors it targets. b. Target sequences are captured by hybridization of the two DNA-caliper anchors (Anc1’ and Anc2’) to the Ab-oligos and the adaptor 3’ overhang, followed by DNA extension (dashed lines) at both ends. c. The extended DNA-calipers are denatured and pulled-down using the Biotin moiety. d-f. Conversion of the extended DNA-calipers into sequencing-ready constructs via intramolecular hybridization- extension. d. A Splint molecule (left), with 5’ phosphorylated termini is added. Note, this Splint has one arm with Hook’ to match the Ab-oligo, while the other has a Poly(G) overhang at the 3’ end. Both termini end with cleavable extension-blockers. The other end of the extended DNA- caliper is 3’ modified by terminal transferase to have a Poly(C) stretch. The cleavable blockers are removed, and the Splint is ligated to the extended DNA-caliper. e. Following intramolecular hybridization, the free 3’ ends of the DNA-caliper arms are extended using the other arm as a template followed by linearization of the molecule by USER enzyme to cleave the dU and PCR using the P1 and P3. primers f. Diverse DNA barcodes are depicted; the genomic regions (Genomic) are represented. Western-seq is described schematically in FIG.10. It is the combination of Prod-seq and WhIP-seq technology with Western blots, enabling quantitative measurement of the number of copies of proteins of varying sizes. Western-seq allows multiplex protein detection with size separation. The results can elucidate that each protein of a given size exists in a certain number of copies in the analyte sample. This can be performed at a scale of thousands of proteins at a time (using a single test). Current Western blot techniques can only perform this for a handful of proteins (2-3) at a time, making for a powerful research tool. Tools Ab-Oligos The methods described comprise a proximity detection (“Prod-seq” hereafter) method comprising contacting or incubating biological samples containing protein and nucleic acids (e.g., cell lysate) with antibody-oligonucleotide conjugates (Ab-oligos hereafter) to detect PPIs and provide detailed measurements of protein abundance. The antibodies in the Ab-oligos are modified to comprise an anchor sequence (Anc) that binds to a DNA-caliper oligonucleotide that serves as a proximity detector by capturing interacting proteins bound to the Ab-oligos on a single molecule. The oligonucleotide tail sequence can comprise about 10-1000 nucleotides. For example, a oligonucleotide tail sequence may have a length of 10-900, 10-800, 10-700, 10-600, 10-500, 10- 400, 10-300, 10-200, 10-100, 10-50 or 10-25 nucleotides. In some embodiments, an oligonucleotide tail sequence has a length of 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 225, 250, 275, 300, 350, 400, 450 or 500 nucleotides. In some embodiments, an oligonucleotide tail sequence is longer than 1000 nucleotides. In some embodiments, the length of an oligonucleotide tail sequence is 103, 104, 105 or 106 nucleotides. Ab-oligos can be synthesized by conjugating single strand oligonucleotides to antibodies, followed by removal of the non-linked components. The rapid kinetics of the iEDDA click chemistry reaction between tetrazine and trans-cyclooctene (TCO) was used as reported previously in van Buggenum, J., Gerlach, J., Eising, S. et al. A covalent and cleavable antibody-DNA conjugation strategy for sensitive protein detection via immuno-PCR. Sci Rep 6, 22675 (2016), which is incorporated herein by reference. The oligo of the Ab-oligo comprises a barcode (BC n, m, etc., which can be specific to the antibody; further discussed below), a unique molecular identifier (UMI, which can be a randomly generated sequence; further discussed below), one or more restriction enzyme sites and an anchor region (Anchor1). The length of an anchor region may vary. The anchor can be located near the 3’ end of the tail. In some embodiments, an anchor region has a length of 5-50 nucleotides. For example, an anchor region may have a length of 5-45, 5-40, 5-35, 5-30, 5-25, 5-20, 5-15, 5-10, 10-50, 10-45, 10-40, 10-35, 10-30, 10-25, 10-20, 10-15, 15-50, 15-45, 15-40, 15-35, 15-30, 15-25, 15-20, 20-50, 20-45, 20-40, 20-35, 20-30, 20-25, 25-40, 25-35, 25-30, 30-40, 30-35 or 35-40 nucleotides. In some embodiments, an anchor region has a length of 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50 nucleotides. In some embodiments, an anchor region has a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleotides. An anchor region, in some embodiments, is longer than 50 nucleotides, or shorter than 5 nucleotides. Any restriction enzyme site can be used, including, but not limited to, those which 5’ or 3’ overhangs (sticky ends) or blunt ends. These enzymes include Type I, Type II, Type III, Type 4 and Type V enzymes. Some examples of restriction enzymes include, but are not limited to, EcroRI, EcoRII, BamHI, HindIII, TaqI, NotI, HinFI, Sau3Ai, PvuII, SmaI, HaeII, HgaI, AluI, EcorRV, EcoP15I, KpnI, PstI, SacI, SalI, ScaI, SpeI, SphI, StuI and XbaI. In some embodiment, the oligonucleotide tail comprises a hook. The length of a hook region may vary. The hook can be located near the 3’ or 5’ end of the tail. In some embodiments, a hook region has a length of 5-50 nucleotides. For example, a hook region may have a length of 5-45, 5-40, 5-35, 5-30, 5-25, 5-20, 5-15, 5-10, 10-50, 10-45, 10-40, 10-35, 10-30, 10-25, 10-20, 10-15, 15-50, 15-45, 15-40, 15-35, 15-30, 15-25, 15-20, 20-50, 20-45, 20-40, 20-35, 20-30, 20-25, 25-40, 25-35, 25-30, 30-40, 30-35 or 35-40 nucleotides. In some embodiments, a hook region has a length of 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50 nucleotides. In some embodiments, an anchor region has a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleotides. A hook region, in some embodiments, is longer than 50 nucleotides, or shorter than 5 nucleotides. In some embodiments, the oligonucleotide tail comprises one or more primer domains. A primer domain is a domain to which a primer binds. A primer is a strand of short nucleotide sequence that serves as a starting point for nucleic acid (e.g., DNA) synthesis. In some embodiments, an oligonucleotide tail can comprise an internal primer domain (e.g., near the 5′ or 3’ end), which may be used for amplification. The length of a primer domain may vary. In some embodiments, a primer domain has a length of 5-50 nucleotides. For example, a primer domain may have a length of 5-45, 5-40, 5-35, 5-30, 5-25, 5-20, 5-15, 5-10, 10-50, 10-45, 10-40, 10-35, 10-30, 10-25, 10-20, 10-15, 15-50, 15-45, 15-40, 15-35, 15-30, 15-25, 15-20, 20-50, 20-45, 20-40, 20-35, 20-30, 20-25, 25-40, 25-35, 25-30, 30-40, 30-35 or 35-40 nucleotides. In some embodiments, a primer domain has a length of 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50 nucleotides. In some embodiments, a primer domain has a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleotides. A primer domain, in some embodiments, is longer than 50 nucleotides, or shorter than 5 nucleotides. DNA-Caliper A DNA-caliper comprises single-stranded oligonucleotide with two free 3' ends that allows bidirectional priming and thus serves as a proximity detector. The DNA-caliper can be used to capture proximity between protein-protein, protein-DNA and also other entities such as protein- RNA, DNA-RNA, protein-metabolite, or cell-cell interactions, and thus for instance improve the characterization of in vitro differentiated cells. Provided herein is a single-stranded nucleic acid comprising two 3′ ends (referred to as a DNA-caliper). A nucleic acid (a polymer of nucleotides) is “single-stranded” if nucleotides that form the nucleic acid are unpaired. That is, nucleotides of a single-stranded nucleic acid are not base paired (via Watson-Crick base pairs, e.g., guanine-cytosine and adenine-thymine/uracil) to nucleotides of another nucleic acid. A single-stranded nucleic acid may be contrasted with a double-stranded (paired) nucleic acid, a typical example of which is a DNA double helix. Single- stranded nucleic acids may include a contiguous (uninterrupted) sequence of nucleotides or, in some embodiments, a single-stranded nucleic acid may be a conjugate that includes two nucleic acid strands joined together, for example, through a chemical (covalent) linkage. In nature, a single strand of a nucleic acid (e.g., DNA or RNA) has a 5′ end (five prime end) and a 3′ end (three prime end). The 5′ end typically contains a phosphate group attached to the 5′ carbon of the ribose ring of a nucleotide and a 3′ end, which is unmodified from the ribose - OH substituent. Nucleic acids are synthesized in vivo in the 5′ to 3′ direction. Polymerase relies on the energy produced by breaking nucleoside triphosphate bonds to attach new nucleoside monophosphates to the 3′-hydroxyl (—OH) group, via a phosphodiester bond. An engineered single-stranded nucleic acid of the present disclosure has two 3′ ends (a whip molecule). Each terminus of the single-stranded nucleic acid includes a 3′-hydroxyl (—OH) group. In some embodiments, a single-stranded DNA-caliper is formed by joining (linking) the 5′ end of one single-stranded nucleic acid to the 5′ end of another single-stranded nucleic acid. In some embodiments, the linkage between two 5′ ends is a covalent linkage. In other embodiments, the linkage is noncovalent. The 5′ ends of single-stranded nucleic acids may be linked to each other using any means in the art for linking nucleic acids to each other. In some embodiments, nucleic acids are linked together using ‘click chemistry’ (see, e.g., V. V. Rostovtsev et al. Angew. Chem. Int. Ed., 2002, 41, 2596-2599; F. Himo et al. J. Am. Chem. Soc., 2005, 127, 210-216; and B. C. Boren et al. J. Am. Chem. Soc., 2008, 130, 8923-8930). An example of a click chemistry reaction is the Huisgen 1,3- dipolar cycloaddition of alkynes to azides to form 1,4-disubsituted-1,2,3-triazoles. The copper(I)- catalyzed reaction is mild and very efficient, requiring no protecting groups, and requiring no purification, in many cases. The azide and alkyne functional groups are largely inert towards biological molecules and aqueous environments, which allows the use of the Huisgen 1,3-dipolar cycloaddition in target-guided synthesis and activity-based protein profiling. Thus, in some embodiments, a whip molecule is formed by linking the 5′ end of one nucleic acid strand that includes an azide group to the 5′ end of another nucleic acid strand that includes an alkyne group. Other linkage reactions are encompassed by the present disclosure and are known in the art. Each arm of the caliper can be about 10 to about 1000 nucleotides. For example, a DNA- caliper arm may have a length of 10-900, 10-800, 10-700, 10-600, 10-500, 10-400, 10-300, 10- 200, 10-100, 10-50 or 10-25 nucleotides. In some embodiments, a DNA-caliper arm has a length of 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 225, 250, 275, 300, 350, 400, 450 or 500 nucleotides. In some embodiments, a DNA-caliper arm is longer than 1000 nucleotides. In some embodiments, the length of a DNA-caliper arm is 103, 104, 105 or 106 nucleotides. Single-stranded nucleic acid DNA-calipers typically include an “anchor domain” at each 3′ end. A “domain” refers to a discrete, contiguous sequence of nucleotides or nucleotide base pairs, depending on whether the domain is unpaired (single-stranded nucleotides) or paired (double-stranded nucleotide base pairs), respectively. Complementary anchor domains bind to each other. A domain is “complementary to” another domain if one domain contains nucleotides that base pair (hybridize/bind through Watson-Crick nucleotide base pairing) with nucleotides of the other domain such that the two domains form a paired (double-stranded) or partially paired molecular species/structure. Complementary domains need not be perfectly (100%) complementary to form a paired structure, although perfect complementarity is provided, in some embodiments. Thus, an anchor domain of a DNA-caliper that is complementary to an anchor domain of a barcoded nucleic acid (described below) binds to that domain, for example, for a time sufficient to initiate polymerization in the presence of polymerase and under conditions appropriate for polymerization. The length of an anchor domain may vary. In some embodiments, an anchor domain has a length of 5-50 nucleotides. For example, an anchor domain may have a length of 5-45, 5-40, 5-35, 5-30, 5-25, 5-20, 5-15, 5-10, 10-50, 10-45, 10-40, 10-35, 10-30, 10-25, 10-20, 10-15, 15-50, 15- 45, 15-40, 15-35, 15-30, 15-25, 15-20, 20-50, 20-45, 20-40, 20-35, 20-30, 20-25, 25-40, 25-35, 25-30, 30-40, 30-35 or 35-40 nucleotides. In some embodiments, an anchor domain has a length of 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50 nucleotides. In some embodiments, an anchor domain has a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleotides. An anchor domain, in some embodiments, is longer than 50 nucleotides, or shorter than 5 nucleotides. In some embodiment, the DNA-caliper comprises one or more hooks (e.g., 2, 4, 6 etc). The length of a hook region may vary. The hook can be located near the 3’ or 5’ end of the caliper arm. In some embodiments, a hook region has a length of 5-50 nucleotides. For example, a hook region may have a length of 5-45, 5-40, 5-35, 5-30, 5-25, 5-20, 5-15, 5-10, 10-50, 10-45, 10-40, 10-35, 10-30, 10-25, 10-20, 10-15, 15-50, 15-45, 15-40, 15-35, 15-30, 15-25, 15-20, 20-50, 20-45, 20-40, 20-35, 20-30, 20-25, 25-40, 25-35, 25-30, 30-40, 30-35 or 35-40 nucleotides. In some embodiments, a hook region has a length of 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50 nucleotides. In some embodiments, an anchor region has a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleotides. A hook region, in some embodiments, is longer than 50 nucleotides, or shorter than 5 nucleotides. In some embodiments, a single-stranded nucleic acid DNA-caliper comprises primer domains. A primer domain is a domain to which a primer binds. A primer is a strand of short nucleotide sequence that serves as a starting point for nucleic acid (e.g., DNA) synthesis. In some embodiments, DNA-calipers comprise a pair of internal primer domains (e.g., near the linked 5′ ends), which may be used for amplification of sequence-ready barcoded constructs produced using the methods of the present disclosure. The length of a primer domain may vary. In some embodiments, a primer domain has a length of 5-50 nucleotides. For example, a primer domain may have a length of 5-45, 5-40, 5-35, 5-30, 5-25, 5-20, 5-15, 5-10, 10-50, 10-45, 10-40, 10-35, 10-30, 10-25, 10-20, 10-15, 15-50, 15-45, 15-40, 15-35, 15-30, 15-25, 15-20, 20-50, 20-45, 20-40, 20-35, 20-30, 20-25, 25-40, 25-35, 25-30, 30-40, 30-35 or 35-40 nucleotides. In some embodiments, a primer domain has a length of 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50 nucleotides. In some embodiments, a primer domain has a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleotides. A primer domain, in some embodiments, is longer than 50 nucleotides, or shorter than 5 nucleotides. In some embodiments, a single-stranded nucleic acid DNA-caliper comprises a cleavage site near the 5’ end of one or both arms of the DNA-caliper. The cleavage site functions to cleave the molecule in two; separates the two 5’ ends of the DNA-caliper following, for example, the Prod-seq reaction, which enables the release of the two strands while maintaining the proximity detection information intact. For example, the cleavage site can be any restriction enzyme site, deoxy uridine (dU; with, for example, a uracil-DNA glycosylase, which are evolutionarily conserved DNA repair enzymes that initiate the base excision repair pathway and remove uracil from DNA) or a photo-cleavable site (such as a site that separates the molecule upon exposure to UV light, for example, with a photo-cleavable modification that contains a photolabile functional group that is cleavable by UV light of specific wavelength (e.g., 300-350 nm); a photo-cleavable spacer can be purchased from Integrated DNA Technologies (IDT) called IDT-PC and can be placed between DNA bases or between an oligo and a terminal modification). An example of a photo-cleavable spacer: Any restriction
Figure imgf000028_0001
not limited to, those which 5’ or 3’ overhangs (sticky ends) or blunt ends. These enzymes include Type I, Type II, Type III, Type 4 and Type V enzymes. Some examples of restriction enzymes include, but are not limited to, EcroRI, EcoRII, BamHI, HindIII, TaqI, NotI, HinFI, Sau3Ai, PvuII, SmaI, HaeII, HgaI, AluI, EcorRV, EcoP15I, KpnI, PstI, SacI, SalI, ScaI, SpeI, SphI, StuI and XbaI. The PPI bound Ab-oligo sequences are captured by hybridization of the two DNA-caliper anchors to the complementary anchor sequences on the Ab-oligo (Anc), followed by DNA polymerase extension at both ends. This extension covalently links the barcodes and UMI sequences of the two antibodies. (iii) The extended DNA-calipers are then denatured and captured by biotin pull-down and then converted into sequencing-ready constructs. The DNA-calipers can also comprise one or more capture molecules (to aid in the isolation of the DNA-calipers, e.g., extended DNA calipers). Further, an extended DNA‐caliper can be used as a template for DNA synthesis regardless of the directionality inversion in mid‐template. Further still, in the methods described herein, there is no requirement for an a priori targeting of a specific protein or genomic loci. Thus, while major approaches such as affinity purification‐mass spectrometry (AP‐MS) enable the study of PPI, a specific protein first has to be targeted and tagged by genetically engineered cell models. Further, while extensions of proximity ligation approaches (e.g., in situ proximity ligation assay (PLA)) and mass spectrometric-based methodologies can be used to study the details of DNA associated complexes, these methods also require the targeting of specific genomic loci. UMI and barcodes The use of unique molecular identifier (e.g., a short random DNA sequence) labeled DNA‐ barcoded antibodies allows the simultaneous interrogation of multiple proteins from a single sample. For example, two proteins that interact with each other can be captured by using synthetic DNA (oligonucleotides) to barcode each and every target protein (using modified antibodies). The DNA-Caliper can capture the barcodes (that represent proteins) that are near each other in the cell and allows hybridization and extension at both ends. Such extension results in the barcodes of the two interacting proteins in one DNA molecule. In embodiments, interactions of multiple proteins can be captured in a single reaction. DNA sequencing and the cellular data can be analyzed to report all of the interactions and the expression levels of the protein in each condition. In some embodiments, the DNA-calipers and/or oligonucleotide tail also comprise a barcode. A “barcoded nucleic acid” is a nucleic acid, typically single-stranded, that includes a barcode domain. A “barcode domain” is a domain that includes a nucleotide sequence that can be used to identify the barcoded nucleic acid or to identify a biomolecule(s) to which the barcoded nucleic acid is linked (directly or indirectly linked). A barcoded nucleic acid may include a barcode domain that is unique to that single nucleic acid (among a population of barcoded nucleic acids, the barcode is specific to that one nucleic acid) or a barcode domain that is unique to a subpopulation of nucleic acids (among multiple populations of barcoded nucleic acids, the barcode is specific to a single subpopulation of barcoded nucleic acids). The length of a barcode domain may vary. For example, a barcode domain may have a length of 5-45, 5-40, 5-35, 5-30, 5-25, 5- 20, 5-15, 5-10, 10-50, 10-45, 10-40, 10-35, 10-30, 10-25, 10-20, 10-15, 15-50, 15-45, 15-40, 15- 35, 15-30, 15-25, 15-20, 20-50, 20-45, 20-40, 20-35, 20-30, 20-25, 25-40, 25-35, 25-30, 30-40, 30-35 or 35-40 nucleotides. In some embodiments, a barcode domain has a length of 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50 nucleotides. In some embodiments, a barcode domain has a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleotides. A barcode domain, in some embodiments, is longer than 50 nucleotides, or shorter than 5 nucleotides. The length of a barcoded nucleic acid itself may vary. In some embodiments, the length of a barcoded nucleic acid is 20-1000 nucleotides. For example, a barcoded nucleic acid may have a length of 20-900, 20-800, 20-700, 20-600, 20-500, 20-400, 20-300, 20-200, 20-100, 20-50 or 20- 25 nucleotides. In some embodiments, a barcoded nucleic acid has a length of 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 225, 250, 275, 300, 350, 400, 450 or 500 nucleotides. In some embodiments, a barcoded nucleic acid is longer than 1000 nucleotides. A barcoded nucleic acid, in some embodiments, further includes a primer domain and an anchor domain that is complementary to one of the anchor domains of, for example, a DNA caliper. As an example, barcoded nucleic acid can comprise a primer domain at its 5′ end, an internal (central) barcode domain and an anchor domain at its 3′ end. It should be understood that the barcoded domain need not be centrally located. The 5′ end of the barcoded nucleic acid can be linked to an antibody that recognizes a protein of interest (e.g., a protein associated with genomic DNA). A barcoded nucleic acid may be linked to any biomolecule. The 3′ end of the barcoded nucleic acid can include an anchor domain that that is complementary to one 3′ end of a DNA- caliper such that the two anchor domains bind to each other to form a paired domain. Anchor domains, in some embodiments, are used for localizing a target biomolecule(s) of interest. When co-localizing two biomolecules, one of the biomolecules contains an anchor domain complementary to one of the anchor domains of a DNA caliper, and the other biomolecule contains an anchor domain complementary to the other of the anchor domains of a DNA caliper. When a DNA caliper is used in combination with a barcoded nucleic acid linked to a target biomolecule, typically the barcoded nucleic acid contains an anchor domain complementary to one anchor domain of the DNA caliper, and another biomolecule contains an anchor domain complementary to the other anchor domain of the DNA caliper and has a unique molecular identifier (UMI). In some embodiments, the DNA caliper or single stranded oligonucleotide tail comprises a unique molecular identifier (UMI) or other nucleotide sequence or moiety specific to that adaptor molecule, which makes distinct each adaptor molecule in a population (see, e.g., Kivioja, T. et al. Nat Methods 2012, 9, 72-4). The UMI can be about 5-45, 5-40, 5-35, 5-30, 5-25, 5-20, 5-15, 5- 10, 10-50, 10-45, 10-40, 10-35, 10-30, 10-25, 10-20, 10-15, 15-50, 15-45, 15-40, 15-35, 15-30, 15-25, 15-20, 20-50, 20-45, 20-40, 20-35, 20-30, 20-25, 25-40, 25-35, 25-30, 30-40, 30-35 or 35- 40 nucleotides. In some embodiments, a UMI has a length of 5, 10, 15, 20, 25, 30, 35, 40, 45 or 50 nucleotides. In some embodiments, a UMI has a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleotides. A UMI, in some embodiments, is longer than 50 nucleotides, or shorter than 5 nucleotides. Adaptor molecules are used, in some embodiments, to add an anchor domain or other nucleotide domain to another nucleic acid molecule. For example, adaptor molecules can be added at each end of a fragmented piece of chromatin. Each adaptor molecule can include an unpaired 3′ overhang, which includes an anchor domain, Anchor2. Whether an adaptor is single-stranded, double-stranded, or partially double-stranded (partially single-stranded), depends on the target biomolecule to which the adaptor is being added. Typically, a double-stranded, or partially, double- stranded adaptor is added to the terminus (or termini) of a double-stranded target nucleic acid, and a single-stranded adaptor is added to the terminus of a single-stranded target nucleic acid. As another example, an adaptor molecule can used, for example, to add a double-stranded domain and a single-stranded homopolymer overhang domain (or other single-stranded nucleotide overhang domain) to a 3′ end of an extended DNA-caliper (e.g., following hybridization to a barcoded nucleic acid, polymerization through the barcoded domain, and dissociation of the resulting partially double-stranded molecule). This adaptor facilitates joining of the two 3′ ends of an extended whip molecule to form a circular double-stranded molecule, which may then be isolated, linearized and sequenced, as discussed below. A homopolymer domain is simply a contiguous stretch of the same nucleotides, such as for example, GGGG. A homopolymer may comprise adenines, guanines, cytosines or thymines (or variants thereof). It should be understood that the homopolymer domain is used to join the 3′ ends of the extended whip molecule to each other to permit polymerization to form a circular, double- stranded molecule. Other means of joining the two 3′ ends are encompassed by the present disclosure. Thus, the homopolymer domains may be substituted with other complementary nucleotide domains, for example. Ab-Oligo and DNA-Caliper Examples Below are example designs for Ab-oligos and DNA-calipers with complementary anchor sequences. Within each design, underlined sequences are hooks, bold sequences are barcodes, “N” sequences are UMIs, lowercase sequences are anchor regions, and italicized sequences are primer sequences. Design 1: Ab-oligo and DNA-caliper (FIG.10A) In an example of an Ab-oligo and DNA-caliper with complementary anchor sequences, the oligo can comprise the following nucleic acid sequence (SEQ ID NO: 1): TACAACTCTTGTATCTACACGCCACGTCGTGGCGANNNNNNNNNNNNNNNttacaaccag actgatactacag/3AmMO/ A DNA-caliper with complementary anchor sequences to SEQ ID NO: 1 is shown in the nucleic acid sequences shown below: DNA-caliper-armA (SEQ ID NO: 2): /5AzideN/ATTAGGTTGCTAGCTCCTGCAGTCcagtctggttgtaa DNA-caliper-armB (SEQ ID NO: 3): /5Hexynyl/ATAGGGCGTCAGCTCTACGTATGGCAcagtctggttgtaa Design 2: Ab-oligo and DNA-caliper FIG.10B) In an example of an Ab-oligo and DNA-caliper with complementary anchor sequences, the oligo can comprise the following nucleic acid sequence (SEQ ID NO: 4): CCTTATTCTATGTTAAACAGTGCCATCAATNNNNNNNNNNNNNNNttacaaccagactg/3Am MO/ A DNA-caliper with complementary anchor sequences to SEQ ID NO: 4 is shown in the nucleic acid sequences shown below: DNA-caliper-armA (SEQ ID NO: 5): /5AzideN/ACACTCTTTCCCTACACGACGCTCTTCCGATCTcagtctggttgtaa DNA-caliper-armB (SEQ ID NO: 6): /5Hexynyl/GCTAGCGATATCTGGAGTTCAGACGTGTGCTCTTCCGATCTcagtctggttgtaa In Design 2, a systematic computational approach to design the DNA oligo sequences was used to minimize unspecific DNA strand annealing within the sequence. This approach involves the use of multiple computational algorithms and tools in the following order: (1) break the sequences into k-mers (of desired length; e.g. k=5) and find sets of candidate sequences where k- mers are not repeated; (2) perform local alignment on the candidate sequences of different components (DNA caliper, Ab-oligos, etc.) to filter out undesired self-dimers; (3) use sliding windows and local alignments to rank and filter remaining sequences by sequence similarities and unspecific complementary regions; (4) take experimental data from golden gate assembly for additional filtering of non-specific bindings; (5) cross-validate the design with sequence evaluation webtools (Integrated DNA Technologies: OligoAnalyzer, New England Biolabs: NEBridge Ligase Fidelity tools). In addition, the barcode sequences are also carefully selected to take into consideration of DNA synthesis and sequencing errors. Additional examples of sequence designs: 1. Caliper Arm Azide Example: /5AzideN/ACACTCTTTCCCTACACGACGCTCTTCCGATCTNNNNNNNNCagtctggttgtaagc (SEQ ID NO: 7) Bold: For generating caliper via click chemistry Underline: Illumina primer site Ns: For improved Illumina clustering sequencing quality Lowercase: For annealing with the Antibody-oligos; added GC in the end for improved annealing strength and specificity. 2. Caliper Arm Alkyne Example: /5Hexynyl/GCTAGCGATATCTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNNNN cagtctggttgtaagc (SEQ ID NO: 8) Bold: For generating caliper via click chemistry Underline: spacer for the linearization cutting site for improved cutting efficiency; also a back-up cutting site Italics: The linearization cutting site (EcoRV) Bold underline: Illumina primer site Ns: For improved Illumina clustering sequencing quality Lowercase: For annealing with the Antibody-oligos; added GC in the end for improved annealing strength and specificity 3. Antibody-oligo Example: CCTTGAACCACTTCTCTAAATCGACTCANNNNNNNNNNNNNNNgcttacaaccagactg (SEQ ID NO: 9) Bold: Fixed sequences for spacer Underline: Cut&ligate cutting site (XmnI) Italics: Example protein barcode (variable) Ns: UMI for distinguishing different molecules of the same type Lowercase: For annealing with the caliper arms Biological Samples In some embodiments, a biological sample is isolated from a patient or a test subject, such a mammalian patient/subject, including human, livestock (e.g., cow, horse, goat, pig), laboratory (e.g., mice, rate, rabbit) or companion (e.g., dog, cat) patients/animals. Non-limiting examples of suitable biological samples include saliva, sputum, mucus, nasopharyngeal samples, blood, serum, plasma, urine, aspirate, cerebral spinal fluid (CSF) and other liquid samples of biological origin, solid tissue samples such as a biopsy specimen or tissue obtained by surgical resection or tissue cultures or cells derived therefrom and the progeny thereof, cell supernatants, cell lysates, tissue samples, organs, bone marrow, and the like. A sample containing DNAs can be obtained from such cells (e.g., a cell lysate or other cell extract comprising DNAs). The term "sample" with respect to a patient/subject can include DNA and proteins. The definition also includes samples that have been manipulated in any way after their procurement, such as by treatment with reagents, washed, or enrichment for certain cell populations. The definition also includes samples that have been enriched for particular types of molecules, e.g., DNA. The term "sample" encompasses biological samples such as a clinical sample such as saliva, sputum, mucus, nasopharyngeal samples, blood, plasma, serum, aspirate, cerebral spinal fluid (CSF), Stem Cells The methods and tools are highly applicable for studying and advancing stem cell research and regenerative medicine, both for normal and disease states. The suite of tools was established using hiPSCs and hiPSC-derived neuronal cells to identify optimal conditions for these cells, with the goal of using the novel approach to study cells originating from individuals from diverse genetic backgrounds (e.g., iPSCORE22). Further, the ability to model disease conditions by using hiPSC-derived organoids (32) allows for systematic studies. These can include the biological changes happening during the differentiation (multiple time points), a survey of perturbations (e.g., CRISPR libraries) or therapeutics (FIG.1). hiPSCs can be differentiated to NPCs and neurons, which will allow one to monitor changes in PPIs and protein-DNA interactions during in vitro neurodifferentiation. Monitoring complexes that regulate chromatin during neuronal differentiation, such as PcG, can have implications for human disease as suggested by the high number of chromatin regulators associated with neurocognitive genetic variants (33, 34). A challenge in the field of stem cell biology is the limited ability to precisely characterize and compare the features of cells derived by differentiation protocols. There are high expectations in the field of regenerative medicine to substitute neuronal loss caused by disease or environmental factors by engrafting hESC/hiPSC-derived NPCs. Improving the survival rate, the differentiation potential, and the ability of those cells to integrate into host tissue is imperative as the successful rate in current clinical trials is low. The ability of these methods and tools to easily detect numerous molecular phenotypes such as PPIs and protein-DNA interactions can advance the ability to understand these cells and their regenerative potential. Diversity The under-representation of non-European individuals in genomic studies is a limiting factor in the applicability of some treatments and diagnostics procedures to diverse populations. Thus, there is limited knowledge about the variations of biological processes in different genetic backgrounds. This inequality has already caused multiple challenges in diagnosis, drug development and therapy, including an FDA-approved biomarker that was shown to have a population bias (19-21). The limited information about the ways biological processes vary between different populations may lead to a lack of appreciation of genetic background of underrepresented individuals during drug development, diagnosis or treatment. Thus, the overrepresentation of European individuals in genomic studies is a limiting factor in the applicability of some treatments and diagnostics procedures to other, less represented populations. This may be due to differences in protein abundances or the interactions between proteins and the genomic localization of proteins in a manner that represents the diverse populations. Provided herein are new methods and tools that are designed to allow high-throughput multi-dimension studies of many samples at a low cost (FIG. 1). This approach will advance the systematic understanding of the physiology and pathophysiology using samples obtained from individuals from various genetic backgrounds. In particular, efforts can be focused on, for example, human induced pluripotent stem cells (hiPSCs) as these can be derived from individuals from all populations and can be differentiated to model normal and disease states – including neuronal, lung or other tissues. Prod-seq and WhIP-seq can be established using cells from iPSCORE, a collection of hiPSC lines. The iPSCORE includes 222 hiPSCs from individuals from diverse genetic background, age and gender (22). For example, one can focus on two hiPSC lines, from a male and a female of African American ethnicity. These hiPSCs can be cultured and fixed in the lab, and used to evaluate conditions, for example, for protein abundances and studies can be extended to survey high number of hiPSC and hiPSC- derived cells or organoids. The following Examples illustrate some of the materials, methods, and experiments that were used or performed in the development of the invention. EXAMPLES Example I - Prod-seq and WhIP-seq Introduction Dynamic protein complexes and transient protein-protein interactions (PPI) are integral in numerous normal and abnormal, such as cancer, associated processes, such as cellular metabolism, signal transduction networks and regulation of chromatin structure. In cancer, this complex organization is involved in processes such as tumor initiation, progression and metastasis. As another example, stem cells properties, such as pluripotency, response to differentiation signals and chromatin regulation (4), are dependent on protein-protein interactions (PPIs) and the genomic localization of DNA-associated proteins. Yet, there is limited data about the complexity of such interactions due to the inability of available assays to simultaneously detect multiple interactions. Further, as most of the current studies focus on a single or a handful of factors, they do not address the multi-factorial manner in which complexes affect the physiology of health or disease. This gap presents a roadblock in advancing knowledge, for example, of stem cell biology and improving, for example, regenerative medicine, including the study of the impact of genetic mutations on stem cell properties and how aberrations in proteins lead to disease and drug development (5, 6). Most current studies of PPIs use methods such as proximity ligation assay, two-hybrid and affinity purification-mass spectrometry (AP-MS (7, 8)). Protein-DNA interactions are detected by ChIP-seq (9), CUT&RUN (10) or CUT&Tag (11). Recent tools have further extended the abilities to detect multiple interactions, both for PPIs (PROPER-seq, Prox-seq) (12, 13) and for DNA-associated proteins (Co-ChIP, multi-CUT&Tag) (14, 15). These tools have advanced the understanding of stem cell biology and have shown the importance of measuring cellular interactions at scale in an unbiased manner. Yet, these assays still have limitations, e.g., PROPER-seq (12) uses in vitro transcription, Prox-seq (13) can capture only a subset of the PPIs targeted, Co-ChIP (14) can co-survey only a limited number of epitopes and multi-CUT&Tag (15) requires unfixed samples. These assays cannot easily be connected to measure both PPIs and protein-DNA interactions, and usually do not measure protein abundance. Thus, it is infeasible to study the physiological complexity of interaction networks, limiting the understanding of, for example, stem cells and differentiation biology. To fill this gap, provided herein are methods and tools (Fig. 1) that enable simultaneous identification and deep characterization of multiple interactions both between proteins (PPIs) and between proteins and the genome (or other nucleic acid sequences such as RNA). The methods and tools allow high-throughput queries of the dynamic organization of molecular entities and shifts the focus from single interactions towards an unprecedented, multi-dimensional interrogation of the cellular complexity. Further, the molecular interactions captured by the assays enables both evaluation of current hypotheses and generation of novel ones with reagent costs and experimental efforts comparable to existing genomic methods, such as ChIP-seq. As example, the Polycomb Group (PcG) complex in human induced pluripotent stem cells (hiPSCs) was studied in the pluripotent cells, as well as during their differentiation. This provided proof-of-concept for the methods and tools and expanded the current knowledge regarding PcG interactions and functions in pluripotency and differentiation and show its applicability to regenerative medicine research. The PcG complex provides an ideal target for studying both PPIs and protein-DNA interactions. PcG is composed of Polycomb Repressive Complex 1 and 2 (PRC1 and PRC2). PRC1 deposits H2AK119ub1, while PRC2 adds methyl groups to histone H3 to form H3K27me1-3. Both histone modifications are associated with developmental repression of gene activity. PRC1 and PRC2 were shown to form an array of complexes that vary in the specific combination of the auxiliary subunits associated with the core complex (18). The methods employ antibody-oligonucleotide conjugates (AB-oligos hereafter) to target dozens to hundreds of proteins and a molecular detector that can identify DNA-tagged spatially proximal entities. Ideal detection of proximal nucleic acids requires capturing both of them on a single molecule. This presents a unique challenge, as both must be initiated in a 5' to 3' direction, and thus cannot both be captured using traditional methods. To overcome this hurdle, provided herein is a specialized single-stranded oligo with two free 3' ends, the DNA-caliper that allows bidirectional priming and thus serves as a detector of molecular proximity by converting biological information into easily read DNA sequences (FIG.2). Leveraging this proximity detector, a suite of tools were built that enable (i) comprehensive and quantitative reporting of multiple PPIs; (ii) simultaneous mapping of the genomic co-localization of multiple DNA associated proteins; and (iii) detailed measurements of protein abundance. This approach offers several advantages that will transform the fields of, for example, stem cell biology and regenerative medicine: (a) Since specialized equipment is not required, the assays can be carried out in any molecular biology laboratory. Further, the simplicity of the process make the methods easily accessible; (b) The easy to purchase synthetic components make the approach highly cost-effective and scalable; (c) The small number of experimental steps increases reproducibility and improves sensitivity by limiting loss of material across steps; (d) The conversion of molecular proximity into barcoded DNA fragments that are easily sequenced increases accuracy and reduces costs; (e) The ability to simultaneously query dozens to hundreds of epitopes reduces the need for an a priori selection of specific targets, and enables more robust normalization and quantification of protein-protein and protein-DNA interactions; (f) The tools enable comparative studies of normal vs. disease states or across conditions such as screening of small molecules or time points; and (g) The framework extends to profiling interactions in single cells. Prod-seq and WhIP-seq Provided herein is the use of Prod-seq and WhIP-seq to catalog protein interactions, genomic binding and abundance of proteins, using the polycomb group complex members in hiPSCs as an example. The methods and tools, (e.g., Ab-oligos, DNA-calipers) that focused on targeting Polycomb Group (PcG) members were used to identify optimal conditions capturing known interactions in samples hiPSCs of male and female origin. A computational analysis pipeline for robustly identifying and quantifying interactions and protein abundances was produced. The novel methods/tools were next used to map PPIs, binding patterns, and abundances of PcG members in the hiPSCs lines. Further, changes in PPIs, protein-DNA interactions and abundance of PcG members during hiPSCs neural lineage commitment were studied. Prod-seq and WhIP-seq were used to compare the changes in PPI and protein-DNA interactions of PcG members between the pluripotent hiPSCs and the committed hiPSC-derived neural progenitor cells (NPCs) and cortical neurons. The data provides proof of concept for the powerful abilities of the methods/tools disclosed herein while advancing the understanding of how variations in composition of the PcG machinery provides exact regulation for pluripotency and differentiation. The experiments and data provide: (i) biological understanding by detection of genes and/or genomic loci that serve as key nodes in pluripotency and during neuro-differentiation; (ii) benchmarked Prod-seq and WhIP-seq protocols and a matching computational analysis pipeline; (iii) a set of optimized reagents (e.g., Ab-oligos) and conditions; and (iv) datasets from hiPSCs, NPCs, and neurons that can be used as a reference. Two novel complementary methods were developed: (i) Prod-seq (Proximity detection), that allows the comprehensive and quantitative reporting of multiple PPIs, as well as detailed measurements of protein abundance; and (ii) WhIP-seq (Without immunoprecipitation), for simultaneous mapping of the genomic co-localization of multiple DNA associated proteins at an unprecedented resolution together with protein abundance measurements (FIGS. 1 and 2A). As will be made clear below, many of the steps and molecular components/tools of the two methods are identical or similar. Prod-seq is first described and then the variations needed to capture protein- DNA interactions in WhIP-seq are highlighted. Prod-seq Prod-seq uses oligonucleotide-conjugated antibodies (Ab-oligos) to target dozens to hundreds of proteins of interest and the DNA-caliper, a molecular detector that can identify proximal entities labeled with nucleic acids. Ideal detection of proximal nucleic acids requires the ability to capture both of them on a single molecule. This presents a unique challenge, as both must be initiated in a 5' to 3' direction, and thus cannot both be captured using traditional methods. The broadly applicable invention overcomes this hurdle with the DNA-caliper, a specialized single stranded oligonucleotide with two free 3' ends that allows bidirectional priming and thus serves as a proximity detector. This unique molecule is generated by covalently linking the 5' ends of two oligonucleotides by click chemistry (23). The two 3' ends of the DNA-caliper enable detection of adjacent DNA molecules via hybridization and strand extension. Notably, variations in the DNA-caliper features (e.g., linker properties or length) can provide means for discerning near vs. far interactions. The Prod-seq process proceeds as follows (FIG.2A): (i) Fixed and lysed cells are incubated with multiple antibodies, where each specific antibody is covalently linked to a single-stranded oligonucleotide tail (Ab-oligos). The oligonucleotide tail contains an anchor sequence (Anchor 1) complementary to both arms of the DNA-caliper (Anchor 1’; the DNA-caliper is depicted in FIG. 2A); a unique molecular identifier (UMI (24); a short random sequence) and an antibody specific unique barcode (Barcode n, m, etc.). Each arm of the DNA-caliper contains at each 3’ end a sequence that is complementary to the antibody anchors at their 3’ ends (Anchor 1’). (ii) Target sequences are captured by hybridization of the two DNA-caliper anchors to the complementary anchor sequences on the oligonucleotide linked to the antibodies, followed by DNA polymerase extension at both ends (FIG.2A). The extension covalently links the barcodes and UMI sequences of the two antibodies. (iii) The extended DNA-calipers are then denatured and captured by biotin pull-down and then converted into sequencing ready constructs. Sequencing the libraries allows identification of PPI instances as well as counting (using the UMIs (24)). Further, the same Ab- oligos also serve for quantitative detection of protein abundance (similar to other tools (25, 26)). WhIP-seq The second assay, WhIP-seq (FIG. 2B) is built in a similar way and uses the same Ab- oligos. To detect protein-DNA interactions, WhIP-seq uses a DNA-caliper with different anchors on each arm: one targets the Ab-oligo and the other targets the proximal genomic DNA (the DNA is sheared, and anchors are added at the 3’ end). Advantages The methods and tools provided herein convey several advantages that will transform the fields of, for example, stem cells and regenerative medicine: (i) The unique design of the DNA-caliper allows its widespread use as a proximity detector (FIG.2). Discussed herein is the use of this tool for capturing PPIs and mapping DNA-associated proteins; however, the DNA-caliper can capture proximity between other entities such as protein- RNA, DNA-RNA, protein-metabolite, or cell-cell interactions, and thus for instance improve the characterization of in vitro differentiated cells. (ii) The use of UMI labeled oligonucleotide-barcoded antibodies (Ab-oligos) allows the simultaneous interrogation of multiple proteins from a single sample. (iii) The assays have a reduced requirement for an a priori selection of a specific protein or genomic locus, such that only a general set of target proteins need be identified. Thus, while many approaches such as AP-MS (7, 8) enable the study of PPIs, a specific protein first must be tagged by genetically engineered cell models. Further, while an in vitro method (12) or extensions of proximity-ligation and -labeling assays (27-30) can be used to study PPIs and DNA associated complexes, these methods are either technically challenging or require targeting (e.g., a specific genomic locus);, such methods are complementary to Prod-seq and WhIP-seq, as they (1) allow orthogonal validation of observations obtained using the tools provided herein; (2) provide a list of potential targets that can then be further studied by the novel novel methods and tools provided herein. (iv) Prod-seq and WhIP-seq can identify interactions with non-protein targets using Ab- oligos to metabolites or other molecules in the cell. (v) The workflow of Prod-seq and WhIP-seq is modular and adding a second reaction that uses a subset of the sample enables quantitative abundance measurement of the studied proteins. Thus, the methods provided herein allow, for the first time, detection of relative changes between the abundance of proteins and the strengths of their PPIs. (vi) By detecting all pairwise interactions between targeted proteins in Prod-seq, the tendencies for certain proteins or antibodies to participate in non-specific interactions can be more readily detected, leading to increased accuracy in the assessment of PPI rates. In the case of WhIP- seq, comparison between protein abundance and the number of protein-DNA interactions captured can provide a unique insight into the chromatin bound fraction of proteins in the assay. (vii) WhIP-seq can eliminate the need for the spike-in controls currently required for quantitative ChIP-seq (31). Most differential analysis methods for ChIP-seq normalize samples by relying on the total number of mapped reads. When the level of the targeted protein differs substantially between samples, spike-ins are required to properly normalize signals between conditions (31). Inclusion of Ab-oligos to positive controls (e.g., H3) provides an internal normalization standard, thus eliminating the need for spike-in chromatin. (viii) The framework is designed to allow usage of reagents by multiple assays. For example, the same Ab-oligos can be used for both Prod-seq and WhIP-seq by employing a specific DNA-caliper for each tool. (ix) Prod-seq and WhIP-seq can also be used to: (1) Discern near vs. far interactions in Prod-seq and WhIP-seq by varying the properties of the DNA-caliper (e.g., linker arm length); (2) The two tools can be merged as the DNA-calipers can be barcoded as well as tagged differently (e.g., replacing the biotin on the capture arm with, for example, another capture molecule, such as digoxigenin). Thus, application of Prod-seq and WhIP-seq jointly would allow, for instance, to study the dynamics of PPIs between DNA-associated proteins together with their genomic localization as well as the abundance of each protein surveyed. Such a merged approach enables discerning between instances where the PPIs and the genomic organization change in a dependent (e.g., for PPIs of proteins that are bound to the genome) or independent manner (e.g., for PPIs of proteins that are free in the cell/nucleus); (3) Integration of Prod-seq and WhIP-seq with single- cell approaches by generating emulsions of cell-barcoded DNA-calipers. Conjugating single strand oligonucleotides to antibodies In collaboration with CST (www.cellsignal.com), a pipeline was optimized and implemented for a controlled process of generating and evaluating the Ab-oligos by covalently linking single-stranded oligonucleotides to monoclonal antibodies. The rapid kinetics of the iEDDA click chemistry reaction was used as reported (35) and quality control was performed by Illumina sequencing of the oligonucleotides. As an example, 30 monoclonal antibodies for use as Ab-oligos are provided in Table 1. This set targets components of the Polycomb-group (PcG) and members of other well-studied complexes known to have strong interactions. The set also comprises controls, such as histone H3 and H4 that are members of the nucleosome and thus serve as a highly abundant positive control and the FLAG-Tag and H3K27M antibodies that can serve either as negative controls (if not expressed) or as positive controls (if expressed). Initial set of antibodies to establish Prod-seq and WhIP-seq and study the PcG The below table highlights antibodies that target PRC1, PRC2 their target histone modifications and positive and negative controls and a signaling pathway (HER-PIK-mTOR). The canonical PRC1 is composed of RING1A or RING1B and one of the PCGF1-6 proteins as well as a chromobox (CBX2, CBX4, CBX6, CBX7 or CBX8) and one of the PHC1-PHC3 proteins. Variant PRC1 involve RYBP and a PCGF protein that is plays a role in incorporation of the other subunits (e.g., HDAC1 and HDAC2), thus leading to several of PRC1 variants (18). The core of PRC2 is composed of EZH1 or EZH2, EED, RBBP4 or RBBP7 and SUZ (12). Variants of this complex, similar to PRC1, involve interactions with distinct subunits. Thus, the PRC2.2. variant for example, interacts with AEBP2 and JARID2 (18).
Figure imgf000043_0002
Figure imgf000043_0001
Ab-oligo based protein abundance measurements Both the Prod-seq and WhIP-seq methods build on the activity of the Ab-oligos and use the measurements of protein abundances to properly interpret variations in PPIs, in addition to providing a tool to simultaneously measure protein levels of multiple epitopes by counting UMIs. The conditions to enable the use of Ab-oligos in detecting protein abundance variations were first systematically optimized. This included improvement of protein capture, washes, and blocking of the single strand oligonucleotide to avoid non-specific interactions. Recombinant protein complexes were initially used– this allowed one to identify the noise ratios for Ab-oligos whose target is not present while capturing signal from the targets. A cell line with a knockout of EZH2 was then used, the H3K27me3 writer component of the PcG complex (EZH2-KO cells (1)), and thus have reduced levels of the H3K27me3 modification. A multiplex set of Ab-oligos (3 or 6) were used and the abundance of H3K27me3 in KO and wildtype (WT) cells across multiple optimization conditions were compared. Reduced levels (between 7.5 to 14-fold) of H3K27me3 in the EZH2-KO cells were reproducibly detected using the main steps of the Prod-seq and WhIP- seq processes, including cross-linking of the cells, incubation with a multiplex set of Ab-oligos, washes, hybridization, strand extension, library preparation and Illumina sequencing (FIG.3). Assessment of DNA-caliper constructs and reaction conditions The performance of nine distinct designs was evaluated and twelve different DNA-caliper molecules were tested. Several parameters for generating the DNA-caliper were examined: (1) the length of each of the arms was varied from ~30 to 100 bases; (2) flexible chemical “hinges” were tested: amino modifier C12 was added to the 5’ of the oligonucleotides and then PEGylated linkers using inverse electron-demand Diels-Alder (iEDDA) click chemistry (35); (3) 5’ Azide and 5’ Hexynyl groups were added to the oligonucleotides and the Copper-Catalyzed Azide-Alkyne Cycloaddition (CuAAC) click chemistry (36) was performed; (4) a deoxy Uridine (dU) was added to the 5’ end of one of the caliper arms (see Fig. 7) to enable Uracil-Specific Excision Reagent (USER) enzyme digestion to separate the two extended DNA-caliper arms to increase PCR amplification efficiency. The purification of the DNA-caliper was optimized using Urea gel extraction. Overall, it was found that the DNA-caliper that is generated by CuAAC click chemistry and is cleavable by USER enzyme provides the highest efficiency in capturing barcodes from Ab- oligos and conversion to a sequencing library. Development of Prod-seq To identify optimal conditions for the Prod-seq process, a recombinant nucleosome was used that has two histone modifications targeted by the Ab-oligos, H3K27ac and H3K4me3. This recombinant molecule was used to perform a proof-of-concept experiment of the entire Prod-seq process using a set of six Ab-oligos and the ability of the method to detect the two proximal epitopes was observed (FIG. 4). The experiment successfully identified the interaction between H3K27ac and H3K4me3, demonstrating that the overall workflow can successfully detect molecular interactions. Evaluation of adaptor design and ligation to cross-linked chromatin An adaptor was designed (“whip-adaptor”) and the conditions for its ligation to cross- linked chromatin was evaluated by comparing to commercially available adaptors (NEXTflex DNA Barcoded adaptors). Following the optimization of ligation and washing conditions we achieved comparable ligation efficiency to chromatin as to naked DNA. Additionally, the ability to generate Illumina libraries from cross-linked chromatin was evaluated. Focusing on an antibody that targets H3K27ac, highly similar results to ChIP-seq data was obtained (FIG.5). The DNA-caliper can detect genomic binding of a target proteins To evaluate the ability of the DNA-caliper to perform the association between Ab-oligos and cross-linked chromatin, the Ab-oligo targeting H3K4me3 was pre-annealed to the DNA- caliper and incubated with sheared chromatin from HEK293T cells. Following Klenow strand- extension, the enrichment of specific genomic regions captured by DNA-caliper was measured using qPCR with one primer targeting the promoter of positive or negative control loci and the other the DNA-caliper sequence. Between ~4 to ~6-fold enrichment of the positive controls (UBC and B2M) over background (CRYAA) was observed, evidencing the ability to capture genomic loci in the vicinity of the Ab-oligo. Prod-seq and WhIP-seq to catalog protein interactions, genomic binding and abundance of the polycomb group complex members in hiPSCs as an example As an example of this technology, provided herein is the use of Prod-seq and WhIP-seq to study in detail the PcG complex in hiPSCs. An overview of several embodiments of the process is provided in FIGS.6-7. Prod-seq: (i) Fixed and lysed cells are incubated with multiple Ab-oligos. For each Ab- oligo, the oligonucleotide tail contains: Anchor 1 that is complementary to the DNA-caliper arms; a UMI; a unique barcode; and a Hook that enables creation of a sequencing library. ii) The Anchor 1’ regions on the two arms of the DNA-caliper are hybridized to the complementary Anchor 1 regions on the Ab-oligos, followed by extension of the DNA-caliper in both directions. (iii) The extended DNA-calipers are then denatured and captured by biotin pull-down followed by annealing a Splint with 3’-overhanging Hook’ regions on both sides. (iv) The Splint is ligated, and the free 3’ ends of each of the DNA-caliper arms are extended using the other arm as a template. (v) The extended DNA-calipers are linearized by USER enzyme excision of the dU (deoxyuridine; part of the DNA-caliper) and serve as a template for PCR using Illumina AmpliSeq kit and primers that match regions P1 and P2. (vi) Sequencing the libraries allows detection and counting (using the UMIs (24)) of PPI instances. (vii) The same Ab-oligos also serve for quantitative detection of protein abundance (similar to ID-seq (25) or NEAT-seq (26)). This is done by using only the biotinylated arm of the DNA-caliper in a separate hybridization-extension reaction followed by PCR to generate a sequencing library. WhIP-seq: The DNA-caliper used for WhIP-seq differs in one arm from the one used for Prod-seq. This design of the DNA-caliper allows capturing the Ab-oligo on one arm and proximal genomic DNA on the other. The process includes the following steps: (i) First, as in ChIP, cellular chromatin is cross-linked to covalently bind proteins to the DNA. The cross-linked genomic DNA is sheared and ligated to adaptors; each adaptor contains a UMI and a single stranded overhang region (Anchor 2) complementary to the lower arm of the DNA-caliper (Anchor 2’) (FIG. 7A). (ii) Next, the same Ab-oligos used for Prod-seq are added to the cross-linked chromatin constructs. (iii) The Anchor sequences of the DNA-caliper are hybridized to the complementary regions on the Ab-oligos and on the adaptors followed by extension of the DNA-caliper in both directions. (iv) Then, in a manner similar to Prod-seq, the extended DNA-calipers are denatured and captured by biotin pull-down followed by annealing and ligation of a Splint molecule. Note, this Splint has one arm with Hook’ to match the Ab-oligo, while the other has a Poly(G) overhang. (v) The other end of the extended DNA-caliper is 3’-modified by terminal transferase to have a Poly(C). Following intramolecular hybridization, the free 3’ ends of each of the DNA-caliper arms are then extended using the other arm as a template. (vi) Similar to Prod-seq, the molecule is linearized using USER enzyme and serves as a template for PCR using Illumina TruSeq kit and primers that match regions P1 and P3. (vi) Sequencing the libraries generated allows the identification of genomic regions associated with the specific proteins targeted by the Ab-oligos. As in Prod-seq, the same Ab-oligos also serve for quantitative detection of protein abundance. The genomic localization of DNA-associated proteins together with quantifying protein abundances can provide means to study the ratio between the cellular levels of a target protein and its genomic binding. Materials and Methods Prod-seq: Protocol 1. Sonication: Add Protease Inhibitor (PI) into lysis buffer 2; (960 ul lysis buffer 2 (Tris- free) with 40 ul 25X PI) and resuspend DSG+FA crosslinked cells in Lysis buffer (Tris-free) + PI; each 16M cells put 400 ul lysis buffer 2 + PI; put on ice. 2. Transfer to a 96-well cell plate (each well 100 ul), sonicate for 72 min (circulate till pixel reaches 15C). Load a sample-containing 96-Well PIXUL Plate. Set the desired sonication parameters. Once the run has completed, collect the cell lysates and transfer to tubes, spin down at 18000g at 4C for 10 min. 3. Transfer supernatant into new Eppie tubes, spin down again at 18000g at 4C for 10 min. 4. Prepare NHS beads: Take 64 ul NHS beads for 16 sample pairs, remove storage buffer, wash once with 125 ul ice-cold 1mM hydrochloric acid. After the second spin, mix the sonication supernatant with the beads; each [125 ul PBS+PI+Tritin + 25 ul lysate] to beads and incubate @ RT, 120 min. 5. Make Ab-oligo resuspension solution and resuspend Ab-oligos.
6. Add Ab-oligos together in a pool with ET SSB buffer, note some Ab-oligos (such as H4) are added at lower concentrations to maximize signal across all antibodies. of M % ml
Figure imgf000048_0001
11. Resuspend each 12. Add Ab-oligo pool to each tube according to the table in step 6 and incubate overnight at 4C. DAY 2 13. Prepare rCut Buffer (with Tween and PI) 14. Incubate
Figure imgf000049_0001
10 min and at 37ºC for 120 min.
15. After overnight ead mix:
Figure imgf000050_0001
- 3 times with 100uL of ChIP WBI + PI (6000uL of ChIP WBI + 240uL of 25x PI) - 3 times with 100uL of ChIP WB III + PI (6000uL of ChIP WBIII + 240uL of 25x PI) - 2 times with 100uL of ChIP TET + PI (5000uL of ChIP TET + 200uL of 25x PI) At the second TET+PI wash, split half of the volume (50 ul) as samples for PIQ-seq (#A- P) for step 16 and then remove supernatant for Prod-seq samples for step 18. 16. PQ-seq samples (#A-P): Resuspend beads in 8 ul of rCut+TWEEN+PI Add 2 ul of 1:100 diluted 100uM of calpber oligos (2pmol) from step 14 and incubate RT for 30min. 17. Klenow (exo-) Extension - use the same master mix and program as the Prod-seq samples below in step 19; skip the remove supernatant step but do the heat inactivation Add 35 ul of rCut+TWEEN+PI and keep at 4ºC after 18. Prod-seq samples (#1-16): resuspend beads in 5 ul rCut+TWEEN+PI and add Calipers; use caliper with ETSSB from step 14. Add 5 ul caliper+ETSSB into each sample and incubate RT for 30 min. 19. Klenow Extension; make Klenow exo- mix according to table below and add to bead suspension.
Figure imgf000051_0001
20. Incubate at 25ºC for 20 minutes and immediately proceed to step 21. 21. Remove supernatant, and then each add 15 ul rCut+TWEEN+PI from step 13. 22. Heat inactivation: 75ºC for 20 minutes 23. Add XmnI enzyme according to table below:
Figure imgf000051_0002
24. Incubate samples for 1 hour at 37ºC and then 65ºC for 20 minutes. Move to 4ºC after for storage until step 25. 25. Transfer supernatant to a new tube and use 25 ul rCutSmart buffer+Tween+PI to wash beads and save supernatant. Do two total washes and combine all supernatants for a total of 75 ul (discard beads). 26. HMGB1 incubation: Use 75 ul sample from the previous step + 18 ul master mix; for ligation in rCutSmart buffer. After adding 18 u s.
Figure imgf000052_0001
27. Do ligation with T4 ligase, using 93 ul total from prior step and combing with the following: After adding 5ul pe
Figure imgf000052_0002
r sample, incubate at 25 C for 20 minutes, and 65ºC for 10 minutes. Move to 4ºC until step 28. 28. Cut the 98 ul of sample with EcoRV according to table below:
After adding 5ul per nutes, and 65ºC for 20 minutes. Move to 4ºC until step 29. 29. PCR amplification of Caliper/Oligo duplexes. Amplify 14.6 ul of the PROD-seq samples with PrimeStar according to table below: Add 33.4ul to
Figure imgf000053_0001
stock) for a final volume of 50ul. Run the following PCR conditions 1. 98ºC for 30 seconds 2. 98ºC for 10 seconds 3. 62ºC for 10 seconds 4. 72ºC for 10 seconds Repeat 2-419 times 5. 72ºC for 2 minutes Hold at 12ºC. 30. SpeedBead purification of PCR products Make master mixes as calculated below: Take PCR samples
Figure imgf000054_0001
orresponding 7.5% PEG bead mastermix. Incubate at RT (20ºC–25ºC) for 10 min. 31. Place tubes on a magnet and discard the supernatant. Wash the beads 2x by adding 100µL of 80% ethanol at 20ºC–25ºC. Move tubes on the magnet a few times and collect beads on the magnet and discard the supernatant once cleared. 32. Elute using 20 ul of H2O (pre-warmed to 55ºC). Vortex, then incubate at 20ºC–25ºC for 5–10 min. Pulse-centrifuge then place tubes to the magnet. Transfer the supernatant to a new tube and discard the SpeedBeads. 33. Run small amount of the PCR on an agarose gel to confirm and proceed to sequencing. Prod-seq and WhIP-seq using the PcG complex in iPSCs The PcG complex that is composed of two main subcomplexes, Polycomb Repressive Complex 1 and 2 (PRC1 and PRC2) that are regulators in pluripotency and differentiation (18) were focused on to demonstrate aspects of the invention. iPSCORE_1_14 and iPSCORE_1_13 (iPSC lines from African American female and male, respectively) were used. Prod-seq and WhIP-seq use several synthetic components and biochemical reactions, such as the Ab-oligos, the DNA-caliper and enzymes (e.g., DNA Polymerase). For barcode design and decoding/demultiplexing, a computational approach employing a Levenshtein distance-based algorithm that optimizes for low decoding error rate by taking into account base substitution, insertion, and deletion events, which may occur during DNA synthesis and sequencing, was used (51), building on similar antibody barcoding approaches (25, 52) and previous work (44). The following parameters in designing the different synthetic components were considered: (i) Optimal length and sequence of the DNA-caliper (FIG.3); (ii) GC content of all of the components of the DNA-caliper (anchors, primers and flanking regions); (iii) Tm for high specificity hybridization between the anchors on the DNA-caliper and the Ab-oligos; (iv) Avoidance of unwanted secondary structures; (v) Design of specific barcodes for multiple antibodies, accounting for decoding each one to identify the antibody; (vi) To avoid low diversity in nucleotides during sequencing which may impact base call accuracy, the primers used (P1-P3; FIGS.6-7) contain a “spacer” region at the 5’ end. A combination of approaches was employed to identify optimal experimental conditions for iPSCs, e.g., buffers composition, reagents for cell fixation, hybridization temperatures and salt concentration. as discussed below. For creating Ab-oligos, in an effort to identify monoclonal antibodies for ChIP-seq (39), the CST antibodies used provided ideal signal to noise ratios, and two different antibody lots showed highly matching results (39). Yet, the procedure provided herein is not limited to CST antibodies, as other Ab-oligos were and can be used (3, 25). Each Ab-oligo can be evaluated by Western blot. 5’ amine-modified HPLC-grade oligonucleotides (avoiding truncated molecules) were ordered from Integrated DNA Technologies (IDT). Using a published approach (35), the oligonucleotides are functionalized with TCO-PEG4-NHS, and carrier free monoclonal antibodies (concentration >1mg/ml) with Methyltetrazine-PEG4-NHS, and incubated together overnight at 4°C. The reaction is quenched with glycine and free antibodies and oligonucleotides are removed. The efficiency of the conjugation was evaluated by absorbance and silver staining. To verify that the antibody functionality is not impacted, the ability of the Ab-oligo to bind to permeabilized human cells was measured. Ab-oligos produced by these methods were used in Prod-seq experiments (FIG. 4). The Ab-oligos will also allow the determination of the ability to capture signal (positive controls) vs. noise (our negative controls). For Prod-seq, cells are crosslinked using formaldehyde and disuccinimidyl glutarate, lysed in lysis buffer (containing EDTA, Tris-HCl, and SDS), blocked with blocking buffer (containing PBS, Triton X-100, BSA and protease inhibitors), and incubated with Ab-oligos overnight at 4°C. Following washes (in PBS), the DNA-caliper is added, and the hybridization is followed by strand- extension using Klenow fragment. Next, following denaturation (95°C) and biotin pulldown of the extended DNA-calipers, Splint (FIG.6; a double-stranded DNA molecule with two 3’ overhangs and 5’-phosphorylated termini) is added to enable intramolecular annealing and subsequent ligation by T4 DNA ligase and another round of strand-extension by Klenow. The extended DNA- calipers are linearized by USER enzyme, and PCR with P1 and P2 primers is used to generate the PPI fragment pool. In a parallel reaction, to measure protein abundance, the biotinylated arm of the DNA-caliper is added to a subset of the sample that was incubated overnight with the Ab-oligo panel and washed with PBS. The hybridization is followed by strand-extension (Klenow), denaturation (95°C) and biotin pulldown. PCR is used to generate the protein abundance fragment pool. Illumina adaptors with a designated index are added by PCR (AmpliSeq) to each of the pools (PPI and protein abundance) and sequenced. WhIP-seq is performed in a similar manner. Briefly, cells are crosslinked and lysed as above; chromatin is sheared by sonication (PIXUL) and adaptors are ligated to the termini of sheared genomic DNA fragments using NEBNext Library Prep Kit. The chromatin is incubated with the Ab-oligos overnight, and after washes the DNA-caliper is hybridized (55°C) followed by strand-extension (Klenow). Following denaturation (95°C) and biotin pulldown of the extended DNA-calipers, a Splint is added. This Splint has cleavable extension-blockers on both 3’ termini, 5’ phosphorylated termini, one arm that matches the Ab-oligo, and the other with a Poly(G) overhang. The other end of the extended DNA-caliper is 3’-modified by terminal transferase to have a Poly(C) stretch. The cleavable blockers are removed, and the Splint is ligated to the extended DNA-caliper. Next, in a similar manner to Prod-seq a pool of fragments containing Ab- oligos barcodes and genomic sequences is created by a PCR with P1 and P3. Protein abundance is measured as above. Computational approach for the analysis of Prod-seq and WhIP-seq data A strategy to use the sequencing readout of the tools to: (i) robustly link: (a) two barcoded oligonucleotides, each represents a specific antibody (Prod-seq); (b) a genomic region with a barcode of specific Ab-oligo (WhIP-seq); (ii) use the UMIs to count the number of interactions (PPIs or binding enrichments) and discriminate significant interactions from non-specific background; and (iii) measure the abundance of the proteins detected by the set of Ab-oligos by counting the UMIs. After paired-end sequencing of Prod-seq and WhIP-seq libraries (2x150bp; the example in FIG.8 is ~240bp), reads will be decoded to determine the barcodes and UMIs associated with each Ab-oligo captured by the DNA-caliper. Reads 1 and 2 will separately encode the Ab-identifying barcodes and UMIs for the two Ab-oligos captured. WhIP-seq libraries will be processed in a similar manner. The difference is that the left arm of the DNA-caliper will encode a different anchor sequence (Anc2) adjacent to a variable stretch of genomic DNA sequence and its corresponding UMI. For each read pair, he Ab-oligo barcode, the genomic DNA fragment, and their corresponding UMIs will be extracted. For decoding the barcodes and UMIs, we will take advantage of their designed length and location relative to the fixed sequences of the DNA-caliper to extract their sequences (FIGS.6-7). The methods and tools provided herein allow one to assess metrics such as the percentage of reads containing valid barcodes and whether barcodes occur in the expected location and frequency. Read pairs encoding identical barcodes and UMIs will be assumed to be PCR duplicates and collapsed for subsequent analyses. Ab-oligo interaction counts are then used to identify biologically meaningful interactions in each sample. However, interpretation of the observed interaction counts has several challenges: (i) interaction counts are likely to scale with the relative abundance of proteins irrespective of interaction strength, (ii) only a fraction of each targeted protein is likely to participate in interactions measured by the Ab-oligo panel used in the experiment, (iii) certain Ab-oligos may have higher levels of non-specific interactions, (iv) specific barcodes and/or UMIs may lead to biased incorporation into the final read pairs. Due to the complexity of these potential factors and their influence on the interpretation of the data, the experiment can include Ab-oligos targeting both positive controls (e.g., histone H3 and H4) and negative controls (e.g., FLAG-Tag, H3K27M) to verify the assay was successful and estimate the prevalence of non-specific interactions. Data analysis Prod-seq: PPIs are quantified between each pair of antibodies, yielding (n(n+1))/2 potential pairwise interactions across the experiment, which includes self-interactions (e.g., dimers). Under the null assumption that the measured PPI counts represent non-specific interactions, the PPI counts can be modeled as a contingency table and the significance of interaction between any two proteins assessed using the Chi-squared test and controlled for multiple hypothesis testing using the Benjamini–Hochberg procedure. Further, the ratio of observed to expected interactions between each protein and negative control Ab-oligos (e.g., FLAG-Tag) in the contingency table establish a lower bound for PPI detection across the dataset. Detected PPIs will be compared to the known PPIs from public databases (e.g., BioGRID (53)) to estimate the relative accuracy. Putative multi-protein complexes can be inferred from the pairwise PPIs by applying the MCODE algorithm (54) as has been implemented (55). WhIP-seq: Once the protein and genomic sequence information is extracted from the reads, genomic binding maps for each Ab-oligo in the panel can be determined. Since WhIP-seq interrogates sonicated DNA fragments associated with protein complexes, the data can be analyzed in a similar manner to conventional ChIP-seq pipelines. Similar to Prod-seq, Ab-oligos encoding negative controls (e.g., FLAG-Tag) will be used in each experiment to estimate non-specific interactions. In addition, the input DNA will be sequenced as done in ChIP-seq. To identify enriched peaks, WhIP-seq data will first be segregated by Ab-oligo barcode, resulting in collections of DNA fragments specific to each targeted protein. These DNA fragments will be aligned to the genome using BWA (56) and analyzed for significantly enriched peaks by HOMER (50), using both the input and negative control Ab-oligo data as controls. Enriched regions will be subjected to procedures such as annotation to nearby genes and regulatory features (e.g., promoters, ChromHMM (57), etc.) and the enrichment of DNA motifs (50). Detecting changes over time As an example, changes in PPIs, protein-DNA interactions and abundance of PcG members during hiPSCs neural linage commitment can be determined. Multiple lines of evidence show that the PcG complexes present dynamic configurations and interaction patterns (58). These variations in PcG composition are suggested to enable the pleotropic roles these complexes play during development, potentially via fine-tuning of its activity. Thus, for instance, the various configurations of the PcG each bind to a distinct set of genomic loci and act in a different manner (58). Here, Prod-seq and WhIP-seq can be used to compare the changes in PPI and protein-DNA interactions of PcG members between the pluripotent hiPSCs and the committed hiPSC-derived NPCs and cortical neurons. Prod-seq and WhIP-seq on hiPSCs-derived cortical NPCs and neurons As an example, cortical NPCs and neurons will be generated as was previously described (59). Briefly, hiPSC plated in Matrigel-coated dishes are maintained in mTeSR Plus media to confluency. Next, media is switched to NMM (Neurobasal media supplemented with TGFb- inhibitors) for 7 days, and then to NMM media with bFGF for several days. At this point, neural rosettes should be apparent. NPCs can differentiate into cortical neurons by withdrawing bFGF from the media for 4-6 weeks (59). Classical markers for cortical NPCs and neurons will be used to monitor effectiveness of the protocol (60, 61). NPCs and neurons will be fixed and collected. Using an array of Ab-oligos (e.g., Table provided above), the methods and tools provided herein will be used to measure PPIs, genomic localization, and abundance of members of PRC1 and PRC2. Prod-seq and WhIP-seq will be performed in triplicates (3 rounds of NPC/neuron differentiation), from female and male hiPSCs. In parallel, co-IP and ChIP-seq will be performed. PcG configuration and interaction patterns during hiPSCs neural linage commitment The results will be compared to public datasets of PPIs and genomic localization maps from relevant models (e.g., 62). For instance, one expects to observe a reduction in the abundance of the core PRC2 members (EZH2, SUZ12, EED) following neurodifferentiation (58). This would also include variation in the genomic localization patterns, particularly to developmental genes. Further, multiple reports have shown that developmental genes are labeled by the repressive H3K27me3 deposited by EZH2 (PRC2 member) and H3K4me3, a histone modification associated with actively transcribed genes. This combination was termed as “bivalent domains” and was shown to change during differentiation (63, 64). In this instant, the data will be used to follow up on previous observations as well as provide higher resolution understanding of the association between the changes in the composition of bivalent domains and the PcG complexes that regulate their function. Bibliography 1. Wang, X. et al. Genes Dev 2019. 2. Chen, A.F. et al. Nat Methods 2022. 3. Stoeckius, M. et al. Nat Methods 2017. 4. Yousefi, M. et al. Stem Cell Rev Rep 2012. 5. Scott, D.E. et al. Nat Rev Drug Discov 2016. 6. Arrowsmith, C.H. et al. Nat Rev Drug Discov 2012. 7. Richards, A.L. et al. Mol Syst Biol 2021. 8. Carneiro, D.G. et al. Methods 2016. 9. Furey, T.S. Nat Rev Genet 2012. 10. Skene, P.J. et al. Elife 2017. 11. Kaya-Okur, H.S. et al. Nat Commun 2019. 12. Johnson, K.L. et al. Mol Cell 2021. 44. Rotem, A. et al. Nat Biotechnol 2015. 45. Shalek, A.K. et al. Nature 2013. 46. Duttke, S.H. et al. Genome Res 2019. 47. Heinz, S. et al. Cell 2018. 48. Lam, M.T. et al. Nature 2013. 49. Lin, Y.C. et al. Nat Immunol 2012. 50. Heinz, S. et al. Mol Cell 2010. 51. Costea, P.I. et al. PLoS One 2013. 52. Stoeckius, M. et al. Genome Biol 2018. 53. Oughtred, R. et al. Nucleic Acids Res 2019. 54. Bader, G.D. et al. BMC Bioinformatics 2003. 55. Zhou, Y. et al. Nat Commun 2019. 56. Li, H. et al. Bioinformatics 2009. 57. Ernst, J. et al. Nat Methods 2012. 58. Loubiere, V. et al. Bioessays 2019. 59. Almenar-Queralt, A. et al. Nat Genet 2019. 60. Shi, Y. et al. Nat Protoc 2012. 61. Shi, Y. et al. Nat Neurosci 2012. 62. Kloet, S.L. et al. Nat Struct Mol Biol 2016. 63. Bernstein, B.E. et al. Cell 2006. 64. Voigt, P. et al. Genes Dev 2013. Example 2 - Prod-seq and WhIP-seq to catalog protein interactions, genomic binding and abundance of, for example, the polycomb group complex members in hiPSCs A computational framework accompanying Prod-seq and Whip-seq was developed. Initially Prod-seq was focused on. The framework generates multiplexed profiles of protein quantification and PPI detection from sequencing results. The approach also uses a probability mixture model to assign PPI confidence scores via maximum likelihood estimation on the generated protein quantification and PPI detection profiles. These confidence scores can then be used to filter PPI detection noise and to evaluate the reliability of the detected PPIs. These data provide information about the probability distribution to use in the mixture model to appropriately score true-positive and true-negative PPIs. To assign confidence scores to Prod-seq PPI detection profiles, the similarity between Prod-seq and AP-MS for quantitative PPI detection was built on. As a starting point, we adapted the probability mixture model scoring approach of Significance Analysis of INTeractome (SAINT; Choi et al., Curr Protoc Bioinformatics, 2013), one of the widely used AP-MS PPI scoring methods. We similarly used a probability mixture model setup and the initial assumption that readout follows Poisson distribution but introduced some initial changes: (i) in addition to Prod- seq PPI readout, we added PQ-seq readout (i.e., overall protein abundance) into the maximum likelihood fitting; and (ii) we changed the parameters specific to AP-MS experimental settings into parameters that model Prod-seq and PQ-seq experimental workflows. However, the Poisson model reports numerous false positive and false negative PPI pairs in the example dataset (FIG.13B). It was also observed that almost all probabilities from all the Poisson distributions in the model are extremely close to 0 (data not shown). This is likely due to the wider range and much higher values of Prod-seq UMI combination counts than AP-MS spectral counts, which when modeled with a Poisson distribution, can underestimate the true variance in the data. This observation suggests improving the model by replacing the Poisson distribution with a probability distribution that allows additional parameters to model the variance observed in the count data. The Poisson distributions in the model were thus replaced with negative binomial distributions whose dispersion parameters enable additional variances to be included in the model. In the negative binomial model, it was further assumed that the dispersion parameters account for additional variances introduced in the experimental workflows and thus one dispersion parameter is shared among all PPIs for Prod-seq and another dispersion parameter is shared among all protein targets for PQ-seq (total of two additional parameters for the entire model). Implementing the negative binomial model on the example dataset shows that the negative binomial model assigns PPI confidence scores that align better with ground truth (FIG.13C). In particular, one make assumptions similar to SAINT that the mean values of the readout distributions are proportional to the quantification level of the targets (details below), and one additionally introduces dispersion parameters ϕ that describe the relationship between the mean N and variance OP of the distributions as σP = μ + ϕμP; two dispersion parameters ^ ! and ^ >5? are added where ^ ! is shared for all PQ-seq UMI count distributions and ^ >5? is shared for all Prod-seq UMI count distributions. To formally state the negative binomial model, for a given biological sample with PQ-seq and Prod-seq readouts, one maximizes the combined likelihood of all readout values P^X, Q|^ = 2V3 P#X23' ⋅) U P^QU |^ , where $ denotes the Prod-seq PPI UMI combination count matrix combination count between proteins ^ and W,
Figure imgf000063_0001
and ^ denotes the PQ-seq count vector denotes the UMI count of protein X. For the distributions of QU and X23: ∼ ^^^^^^^^ ^^^^^^^^ ^^^^ ^^^ ^^^^^^^^^^ ^^^^^^^^^ ^ !
Figure imgf000063_0002
Figure imgf000063_0003
λ23 = ^45%672 + ^2^#α3 + ^3) + ^6%:4;<α2α3 and the additional
Figure imgf000063_0004
+, denotes the proportion of true interactions in the entire dataset; ^^^ is the readout scaling factor for PQ-seq detection protocol and sequencing; ^6%:4;< and ^45%67 denote the readout scaling factors for specific and non-specific PPI detection in Prod-seq, respectively; ^% and ^% denote the levels of specific and non-specific antibody-oligo binding for each protein target ^, respectively. Parameters +,, ^^^, ^6%:4;< , ^45%67 , ^A,...,4 , ^A,...,4 , ^ !, ^ >5? are estimated from the maximum likelihood
Figure imgf000063_0005
After fitting the maximum
Figure imgf000063_0006
model, the model calculates the confidence score of F Score G^HX232 J each PPI as 23 = P#true'X = 3 . A
Figure imgf000063_0007
Figure 14 illustrates a diagrammatic representation of a machine 1400 in the form of a computer system within which a set of instructions may be executed for causing the machine 1400 to perform any one or more of the methodologies discussed herein, according to an example implementation. Specifically, Figure 14 shows a diagrammatic representation of the machine 1400 in the example form of a computer system, within which instructions 1402 (e.g., software, a program, an application, an applet, an app, or other executable code) for causing the machine 1400 to perform any one or more of the methodologies discussed herein may be executed. For example, the instructions 1402 may cause the machine 1400 to implement the computational processes, methods, techniques, and algorithms described herein. The instructions 1402 transform the general, non-programmed machine 1400 into a particular machine 1400 programmed to carry out the described and illustrated functions in the manner described. In alternative implementations, the machine 1400 operates as a standalone device or may be coupled (e.g., networked) to other machines. In a networked deployment, the machine 1400 may operate in the capacity of a server machine or a client machine in a server- client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 1400 may comprise, but not be limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular telephone, a smart phone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of executing the instructions 1402, sequentially or otherwise, that specify actions to be taken by the machine 1400. Further, while only a single machine 1400 is illustrated, the term “machine” shall also be taken to include a collection of machines 1400 that individually or jointly execute the instructions 1402 to perform any one or more of the methodologies discussed herein. Examples of machine 1400 can include logic, one or more components, circuits (e.g., modules), or mechanisms. Circuits are tangible entities configured to perform certain operations. In an example, circuits can be arranged (e.g., internally or with respect to external entities such as other circuits) in a specified manner. In an example, one or more computer systems (e.g., a standalone, client or server computer system) or one or more hardware processors (processors) can be configured by software (e.g., instructions, an application portion, or an application) as a circuit that operates to perform certain operations as described herein. In an example, the software can reside (1) on a non-transitory machine readable medium or (2) in a transmission signal. In an example, the software, when executed by the underlying hardware of the circuit, causes the circuit to perform the certain operations. In an example, a circuit can be implemented mechanically or electronically. For example, a circuit can comprise dedicated circuitry or logic that is specifically configured to perform one or more techniques such as discussed above, such as including a special-purpose processor, a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). In an example, a circuit can comprise programmable logic (e.g., circuitry, as encompassed within a general-purpose processor or other programmable processor) that can be temporarily configured (e.g., by software) to perform the certain operations. It will be appreciated that the decision to implement a circuit mechanically (e.g., in dedicated and permanently configured circuitry), or in temporarily configured circuitry (e.g., configured by software) can be driven by cost and time considerations. Accordingly, the term “circuit” is understood to encompass a tangible entity, be that an entity that is physically constructed, permanently configured (e.g., hardwired), or temporarily (e.g., transitorily) configured (e.g., programmed) to operate in a specified manner or to perform specified operations. In an example, given a plurality of temporarily configured circuits, each of the circuits need not be configured or instantiated at any one instance in time. For example, where the circuits comprise a general-purpose processor configured via software, the general-purpose processor can be configured as respective different circuits at different times. Software can accordingly configure a processor, for example, to constitute a particular circuit at one instance of time and to constitute a different circuit at a different instance of time. In an example, circuits can provide information to, and receive information from, other circuits. In this example, the circuits can be regarded as being communicatively coupled to one or more other circuits. Where multiple of such circuits exist contemporaneously, communications can be achieved through signal transmission (e.g., over appropriate circuits and buses) that connect the circuits. In implementations in which multiple circuits are configured or instantiated at different times, communications between such circuits can be achieved, for example, through the storage and retrieval of information in memory structures to which the multiple circuits have access. For example, one circuit can perform an operation and store the output of that operation in a memory device to which it is communicatively coupled. A further circuit can then, at a later time, access the memory device to retrieve and process the stored output. In an example, circuits can be configured to initiate or receive communications with input or output devices and can operate on a resource (e.g., a collection of information). The various operations of method examples described herein can be performed, at least partially, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors can constitute processor-implemented circuits that operate to perform one or more operations or functions. In an example, the circuits referred to herein can comprise processor-implemented circuits. Similarly, the methods described herein can be at least partially processor implemented. For example, at least some of the operations of a method can be performed by one or processors or processor-implemented circuits. The performance of certain of the operations can be distributed among the one or more processors, not only residing within a single machine, but deployed across a number of machines. In an example, the processor or processors can be located in a single location (e.g., within a home environment, an office environment or as a server farm), while in other examples the processors can be distributed across a number of locations. The one or more processors can also operate to support performance of the relevant operations in a "cloud computing" environment or as a "software as a service” (SaaS). For example, at least some of the operations can be performed by a group of computers (as examples of machines including processors), with these operations being accessible via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., Application Program Interfaces (APIs).) Example implementations (e.g., apparatus, systems, or methods) can be implemented in digital electronic circuitry, in computer hardware, in firmware, in software, or in any combination thereof. Example implementations can be implemented using a computer program product (e.g., a computer program, tangibly embodied in an information carrier or in a machine readable medium, for execution by, or to control the operation of, data processing apparatus such as a programmable processor, a computer, or multiple computers). A computer program can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a software module, subroutine, or other unit suitable for use in a computing environment. A computer program can be deployed to be executed on one computer or on multiple computers at one site or distributed across multiple sites and interconnected by a communication network. In an example, operations can be performed by one or more programmable processors executing a computer program to perform functions by operating on input data and generating output. Examples of method operations can also be performed by, and example apparatus can be implemented as, special purpose logic circuitry (e.g., a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC)). The computing system can include clients and servers. A client and server are generally remote from each other and generally interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In implementations deploying a programmable computing system, it will be appreciated that both hardware and software architectures require consideration. Specifically, it will be appreciated that the choice of whether to implement certain functionality in permanently configured hardware (e.g., an ASIC), in temporarily configured hardware (e.g., a combination of software and a programmable processor), or a combination of permanently and temporarily configured hardware can be a design choice. Below are set out hardware (e.g., machine 1400) and software architectures that can be deployed in example implementations. In an example, the machine 1400 can operate as a standalone device or the machine 1400 can be connected (e.g., networked) to other machines. In a networked deployment, the machine 1400 can operate in the capacity of either a server or a client machine in server-client network environments. In an example, machine 1400 can act as a peer machine in peer-to-peer (or other distributed) network environments. The machine 1400 can be a personal computer (PC), a tablet PC, a set-top box (STB), a Personal Digital Assistant (PDA), a mobile telephone, a web appliance, a network router, switch or bridge, or any machine capable of executing instructions (sequential or otherwise) specifying actions to be taken (e.g., performed) by the machine 1400. Further, while only a single machine 1400 is illustrated, the term “computing device” shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein. Example machine 1400 can include a processor 1404 (e.g., a central processing unit CPU), a graphics processing unit (GPU) or both), a main memory 1406 and a static memory 1408, some or all of which can communicate with each other via a bus 1410. The machine 1400 can further include a display unit 1412, an alphanumeric input device 1414 (e.g., a keyboard), and a user interface (UI) navigation device 1416 (e.g., a mouse). In an example, the display unit 1412, input device 1414 and UI navigation device 1416 can be a touch screen display. The machine 1400 can additionally include a storage device (e.g., drive unit) 1418, a signal generation device 1420 (e.g., a speaker), a network interface device 1422, and one or more sensors 1424, such as a global positioning system (GPS) sensor, compass, accelerometer, or another sensor. The storage device 1418 can include a machine readable medium 1426 on which is stored one or more sets of data structures or instructions 1402 (e.g., software) embodying or utilized by any one or more of the methodologies or functions described herein. The instructions 1402 can also reside, completely or at least partially, within the main memory 1406, within static memory 1408, or within the processor 1404 during execution thereof by the machine 1400. In an example, one or any combination of the processor 1404, the main memory 1406, the static memory 1408, or the storage device 1418 can constitute machine readable media. While the machine readable medium 1426 is illustrated as a single medium, the term "machine readable medium" can include a single medium or multiple media (e.g., a centralized or distributed database, and/or associated caches and servers) that configured to store the one or more instructions 1402. The term “machine readable medium” can also be taken to include any tangible medium that is capable of storing, encoding, or carrying instructions for execution by the machine and that cause the machine to perform any one or more of the methodologies of the present disclosure or that is capable of storing, encoding or carrying data structures utilized by or associated with such instructions. The term “machine readable medium” can accordingly be taken to include, but not be limited to, solid-state memories, and optical and magnetic media. Specific examples of machine-readable media can include non-volatile memory, including, by way of example, semiconductor memory devices (e.g., Electrically Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM)) and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The instructions 1402 can further be transmitted or received over a communications network 1428 using a transmission medium via the network interface device 1422 utilizing any one of a number of transfer protocols (e.g., frame relay, IP, TCP, UDP, HTTP, etc.). Example communication networks can include a local area network (LAN), a wide area network (WAN), a packet data network (e.g., the Internet), mobile telephone networks (e.g., cellular networks), Plain Old Telephone (POTS) networks, and wireless data networks (e.g., IEEE 802.11 standards family known as Wi-Fi®, IEEE 802.16 standards family known as WiMax®), peer-to-peer (P2P) networks, among others. The term “transmission medium” shall be taken to include any intangible medium that is capable of storing, encoding or carrying instructions for execution by the machine, and includes digital or analog communications signals or other intangible medium to facilitate communication of such software. Example 3- Western-seq Western blotting limitations Western blotting is a widely used technique developed to detect expression of a protein of interest in a sample. Protein lysate is created from a sample, and along with a protein ladder run through gel electrophoresis to separate all of the proteins present in the sample by size. Next, the proteins are transferred from inside the gel onto a membrane using a voltage, making them accessible to antibody binding. To check if the separation and transfer process worked, the membrane can be stained. The membrane then undergoes a process called blocking, to reduce any nonspecific binding and background noise for the final product. Then, the membrane is treated with an antibody that binds to the protein of interest, called the primary antibody. Primary antibodies can be pAb, mAb, or even rAb and are generated in small animals, such as mice, rabbits, or rats. Using smaller animals makes the antibody easier to generate and by using an antibody from a different species than the sample, it allows for detection of only the protein of interest. At this point, the sample is treated with another antibody, called a secondary antibody. The secondary antibody is specific to the Fc domain of the primary antibody. This secondary antibody can be conjugated to also include a fluorescence tag or horseradish peroxidase (HRP), so as to allow for visualization in imaging. For example, if a researcher were interested in visualizing beta actin protein in a human cell sample, the cell sample would first be lysed and the lysate would be size separated using gel electrophoresis. After the transfer and blocking steps, the sample would be stained with a primary antibody specific for beta actin. The sample would then be incubated with a secondary antibody specific for the Fc domain of the primary antibody. The secondary antibody contains a fluorescence label that allows visualization. In the final step the membrane is imaged, and the result is a picture of the membrane with a band at the size of the protein of interest, confirmed by the protein ladder standard (8). Researchers use this technique to quantify protein levels, such as to validate knockdown or knockout systems. If a specific protein is knocked down or knocked out in a sample and a Western blot is run, the resulting band will be much weaker or entirely absent when compared to a control or wildtype sample. Western blots may be widely used, but they are also known to be challenging and difficult to reproduce. The process typically takes two days as the blocking step and/or primary antibody incubation occurs overnight. The vast number of steps, each requiring various buffers and reagents, give much room for complications to arise with little room for pinpointing exactly which is the problematic step. If a membrane is imaged after blotting and there are no protein bands, it could be an issue with the wash buffer, with the concentration of the primary antibody, with what the primary antibody was diluted in, with the primary antibody itself or with the secondary antibody (8). Not only is troubleshooting in Western blotting cumbersome, but the knowledge gained from a traditional Western blot is limited to qualitative information - visually observing a change in band intensity in a sample makes it difficult and unreliable for a researcher to quantify the difference in protein levels between two samples. While Western blots can give great qualitative information, there is minimal quantitative infomation gained. When trying to determine the knockdown of a protein in a sample compared to a control, the only data points are the band size and visual difference in band intensity. Quantitative Western blotting exists but is an extensive protocol which requires careful plotting and testing of sample loaded and band intensity for it to be fully reliable. Ultimately, it does not provide an exact amount for the antibody that has bound or for the abundance of the protein of interest. Another limitation of Western blots is the number of proteins able to be detected from a single sample, which comes from a limitation of the number of antibodies able to be used for a single sample. To overcome this, researchers either use primary antibodies generated in a multitude of species to allow for the use of several different secondary antibodies, but this is still limited to the number of commonly used small animals for primary antibody generation. There is also a technique called multiplex fluorescence. In multiplex fluorescence Western blotting, multiple primary antibodies, ideally from different species, are used on the same membrane (10). The secondary antibody contains fluorescence with various colors from different species which allows for visualization of the multiple proteins of interest by the different fluorescent colors. In these techniques, visualization is still limited by the number of primary antibodies from different species available and by the number of fluorescent antibody colors. Despite these limitations, Western blots are widely used, and have allowed researchers to gain much information on protein expression in cells. Analysis of alternate protein quantification methods Many researchers have been working towards creating high-throughput and large-scale assays targeting proteins to gain a better understanding of the roles they play in molecular biology, as well as to develop diagnostic tools. One of the long-standing existing diagnostic assays is known as ELSA (Enzyme Linked Immunosorbent Assay). This assay can detect multiple proteins in a sample by using many antibodies at the same time. An antigen is adhered into a single well, with one well prepared per antibody used. Each well is treated with a different enzyme-antibody conjugate and then washed to dilute out any unbound antibody-enzyme. When the substrate for the enzyme is added, a color change occurs, indicating the presence of the protein specific for a certain antibody in the sample. This assay generally tests for a multitude of blood proteins and is done on blood samples (16, 17). ELSA was established in the 1970s and is still one of the primary protein detection assays used today. Olink Proteornics, a Swedish company, has developed proximity extension assay (PEA) technology, as a way to detect proteins in a sample, which can be as little as 50 uL, like a drop of blood. Two antibody-oligonucleotide conjugates targeting the same protein of interest are used. The oligos have overlapping complementary sequences, and when both antibodies are bound to the same protein, the oligos are brought into proximity with each other and create a double stranded sequence. If there is non-specific binding of a single antibody, the oligo will remain single stranded and cannot be utilized in downstream processes. For proteins with both antibodies correctly bound, qPCR can be used to detect the amount of double stranded DNA for a target protein, and therefore determine the amount of protein in the sample. Adaptors and barcodes are added to the DNA for next generation sequencing (NGS) to increase the throughput of analysis (18). Several groups partnering with the UK BioBank have shown this assay to be highly scalable (19-22). Other companies have created comparable technologies to PEA, some of which were evaluated in tandem with PEA by Haslam, et al in 2022 (22). SomaScan from Somalogic was tested alongside Olink's PEA, and also uses 50uL as input. SomaScan utilizes modified aptamers, which are oligonucleotides that have specificity for target proteins and are used in place of an antibody. The aptamers are referred to as SOMAmers (slow off-rate modified aptamers) and are marketed as having more specificity than polyclonal antibodies. In this technique, a SOMAmer added to a sample will bind and create a protein complex to its target protein, with wash steps diluting out any unbound SONAmers. Subsequent steps cause the SOMAmers to unbind and rebind, which increases the specificity of the aptamer. After the final re-complexing step, the SOMAmers are eluted and detected by DNA quantification (23). Unlike the single stranded DNA tags of PEA or the single stranded SOMAmers, van Buggenum et al. have developed immunodetection sequencing, or ID-seq, which uses double stranded DNA tags bound to antibodies to quantify protein levels in a sample. After the antibody- dsDNA conjugate binds to the epitope, the dsDNA tag gets cleaved. The dsDNA on the antibody contains a barcode and unique molecular identifier (UMI), a randomized nucleotide sequence used in NGS which allows researchers to discern read duplicates from unique, significant events (26). The purpose of the barcode is to associate the DNA with a specific antibody (25). The presence of a UMI increases confidence in this association. After the dsDNA is cleaved from the antibody, it is barcoded per sample. Then there is another set of barcodes used for library prep and sequencing. Another method is called protein quantification followed by sequencing (PQ-seq). Using DNA oligonucleotides conjugated to antibodies, the proportion of protein in a sample can be determined. The DNA tags each have a barcode for the specific antibody as well as UMI. A range of antibodies are added to a sample to bind to their respective target proteins. After a few wash steps, the DNA tags are hybridized, extended, and amplified using PCR. In this amplification they are also prepared for sequencing and once sequenced, the barcodes and UMIs are used to detect protein count. (FIGS.9A-9D) These various protein quantification technologies are useful for testing samples for many different proteins in a high throughput manner, but there are limitations. Western-seq is a novel method that provides higher confidence by multiplexed detection of protein sizes and quantification Combining Western blotting and PQ-seq into Western-seq yields a novel protein quantification technique that merges the qualitative and quantitative advantages of each method (FIGS.9A-9F). The main advantage of Western blotting is that there is a visual of the protein of interest by the binding of the antibody, which is confirmed by the size in kDa. The biggest advantage of PQ-seq is that one can associate relative abundance of protein to DNA, which is easier to quantify than protein, through DNA sequencing on a high scale. The beginning of the Western-seq protocol is very similar to a standard Western blot. The sample is run on a gel, transferred to a membrane, and blocked. The primary antibodies used are the same as in PQ-seq and are tagged with a DNA oligonucleotide that features a barcode and UMI. Before the membrane is treated with the primary antibodies, the antibodies can be incubated with single stranded binding protein (SSB) to bind the oligonucleotides, improve signal, and decrease background noise. A mixture of hundreds of conjugated primary antibodies can be used. The membrane is treated with the traditional secondary antibody for each species of primary antibody used and the membrane is imaged. Due to being blotted with many different antibodies, there will not be a singular band at a singular size for the protein of interest, but many bands at the various sizes confirmed by the protein ladder. At this point, the protocol diverges from a typical Western blot. For each sample run on the original gel, segments of 10-20 kDa increments are cut out of the membrane, placed in a PCR tube strip or 96-well plate and assigned a barcode label. With the piece of membrane still in the tube or well, a caliper arm will be added and annealed to the oligonucleotide and undergo a Klenow extension to extend the sequence. During the first PCR, the oligonucleotide sequence is amplified, and in a second PCR, sequencing primers are added to the sequence. The samples are pooled and sequenced, and the analyzed results will inform on the abundance of protein in the sample by relating the abundance of each unique read with the barcodes to which antibodies they came from. Not only does Western-seq offer qualitative and quantitative information to increase confidence in results, but it can suggest a ratio of non-specific binding to specific binding by comparing noise to signal. It can also act as a quality control step if coupled with other antibody- reliant techniques. Overview of Western-seq Traditional Western Blots are used to detect proteins/determine protein abundance. However, generally there is unspecific binding, such as in multiplex, high throughput protein detection. FIGS.9A-9F provide an overview of Western-seq. Multiplex protein detection can occur via conjugated antibodies, hybridization and strand extension (FIGS.10A-10B), while barcoding allows detection of different proteins in one reaction (FIGS. 10C-10D). Western-seq allows multiplex protein detection with size separation (FIG. 10E). FIG. 10F provides an exemplary protocol for Western-seq. An Exemplary Method Using Cultured Cells Cell Culture HEK293T cells were grown in a DMEM solution (2mM L-glutamine, 10% HI-FBS, 1% Antimycotic-Antibiotic) on standard culture dishes. Plates were stored in 37 °C incubators with 5% CO2. Tazemetostat and DMSO treated cells were treated with l0 uM of their respective medium for 48 hours. The EZH2 knockout cell line was obtained from researchers Wang et al. (27). HeLa-S3 (ATCC CCL-2.2TM) cells were grown in DMEM with GlutaMAXTM (DMEM, 4.5 g/L D-Glucose, 110mg/L sodium pyruvate, 2 mM L-glutamine, 10% HI-FBS, 1 % Antibiotic Antimycotic) on standard culture dishes. Plates were stored in 37 °C incubators with 5% CO2. Induced pluripotent stem cells (iPSCs) were grown on Matrigel (Corning 354230) coated plates with mTeSR (Stem Cell Technologies 85850). Plates were stored in 37 °C incubators with 5% CO2. Western blotting: Harvesting and storing cells HeLa-S3 and HEK293T cells were separately harvested with 0.25% trypsin, washed with PBS, centrifuged at 320g for 5 minutes at 4 °C, resuspended in PBS and frozen at -80°C. iPSCs were rinsed with cold PBS, scraped and transferred to a pre-chilled tube, centrifuged at 200g for 5 minutes at 4 °C, had PBS aspirated, and transferred to dry ice and stored at -80 °C. The HeLa-S3 cells and iPSCs were single crosslinked using formaldehyde. Before freezing, the harvested cells were crosslinked with 1 % formaldehyde/ PBS solution per 1 million cells and incubated at room temperature for 10 minutes while rotating at 8 rpm. The reaction was quenched by adding 1/20th volume of 2.625 M glycine and 1/20th volume 10%; BSA The cell pellet was resuspended, stored in 0.5% BSA/PBS, and frozen at -80°C. Western blotting: Lysing conditions and preparation LDS (NuPAGETM LDS Sample Buffer NP0007) method: Thaw cell pellet and resuspend in lx PBS. Add an equal amount of 2X LDS, and vortex for two minutes. Aliquot and store at -80 °C. This method was used for preparing the protein lysate for all HEK293T cells and HeLaS3 cells, as well as to prepare some of the HEK293T samples. RIPA (150 mM NaCl, 1.0%) N-P-40, 0.5% sodium deoxycholate, 0.1 % SDS, lmM EDTA, 50 mM Tris, pH 8.0) method: Thaw cell pellet and resuspend in lx PBS. Add an equal amount of RIPA buffer with 1 X Protease Inhibitor and incubate for 30 minutes. Spin down at 12,000g for 5-l 0 minutes, aliquot the supernatant and store at -80 °C. This method was used for preparing some of the protein lysate for HEK293T cells and HeLa-S3 cells. NP-40 (50mM Tris-HCl pH 8.5, 150mM NaCl, 1 % NP-40, cOmplete EDTA-free Protease Inhibitor 1 X (Roche l 1873580001), PhosSTOP lX (Roche 4906845001), Benzonase 1 :500) method: Thaw cell pellet and resuspend in lx PBS. Add an equal amount of NP-40 buffer with IX Protease Inhibitor and incubate for 30 minutes. Spin down at 12,000g for 5-10 minutes, aliquot the supernatant and store at -80 °C. This method was used for preparing the protein lysate for some HEK293T cells and HeLa-S3 cells. Lysis buffer 2 (l0mNI EDTA, 50 mM HEPES, 0.5% SDS, pH 7.3 - 7.5) method: Thaw cell pellet and resuspend in lx PBS. For each 1 million cells, add 100 uL lysis buffer 2 and incubate at 37 °C for 60 minutes. Transfer to a 96-well cell plate, each well containing 100 uL and around 1 million cells, seal with a PCR plate seal, and sonicate for 24 minutes using the PIXUL sonicator by ActiveMotif. Add 25X protease inhibitor and spin down at 12,000g for 5-10 minutes. Aliquot the supernatant and store at -80 °C. This method was used for preparing some of the protein lysate for all HEK293T cells. When samples using LDS, RIP A, or N-P-40 were sonicated, samples were transferred to a 96-well plate, each well l00 ul and about l million cells, and sonicated for 24 minutes using the PIXUL sonicator. Protein concentration was measured using the ThermoFisher Scientific Qubit protein assay kit and ThermoFisher Scientific Qubit fluorometer the same day a Western blot was done once the cell pellet had thawed. Western blotting: Sample preparation, running, and transfer Protein amounts were normalized to 15 ug of protein per lane. 3.3uL of sample buffer which included 9% beta-mercaptoethanol and 6X Laemmli Buffer (375 mM Tris-HCl, 10% SDS, 7.5~o glycerol, 0.03%i bromophenol blue) along with 15 ug of protein was added to PCR tubes, and water was added to the samples to bring the total volume to 20 uL. Samples were boiled at 95°C for 5 minutes before loading. 7.5% and 12% precast polyacrylamide gels (BioRad #4568025 and #4561045) were loaded into a Bio-rad Mini Protean Tetra cell system and filled with running buffier (25 mM Tris, 190 mM glycine, 0.1%) SDS).12% gels were used when the protein of interest was less than 30 kDa. 10 uL of sample was added into each lane, and Precision Plus Kaleidoscope Prestained Protein Standards (Bio-Rad #1610375) was used as a protein ladder. Proteins were transferred onto PVDF membrane (0.45 urn, Immobilon 1PVH00005). The transfer system consisted of lX transfer buffer (25mM Tris, 125 mM glycine, 20% methanol pH 8.3) and was placed in a bucket filled with ice in the cold room with ice packs in the system. The transfer was an hour at 100V. The membranes were blocked overnight on a shaker at 4 °C in 5% BSA/ lX TBST (Fisher Scientific BP1600-100; 20mM Tris, 150mM NaCl, 0.5% Tween (EMD Millipore 655204)). When salmon sperm DNA (Invitrogen 15-632-011) was added to the blocking buffer, l00ug/ mL of blocking buffer was used. Western blotting: Antibody incubation and imaging All antibodies used were diluted in 5% BSA/ l X TBST. The membranes were incubated with primary rabbit antibodies for 1 hour at room temperature on a shaker. All primary antibodies (Table 1) were diluted at 1: 1000. When conjugated antibodies were used and SSB (Promega M3011) was being tested, 290.9 μL of 5% BSA/ lx TBST, 6.38 uL of SSB, and 2.72 uL of the conjugated antibody were combined in a 1.5 mL tube and incubated at 37 °C for 30 minutes. The contents of the tube were then added to the correct amount of 5% BSA/ IX TBST to create the l:1000 dilution. All membranes were incubated with anti-rabbit secondary antibodies diluted at l :3000 for 45 rninutes at room temperature on a shaker. Proteins were imaged using the ECL Thermo ScientificTM SuperSignalTM West Pico PLUS Chemiluminescent Substrate (PI34577) in a Bio-Rad ChemiDoc XRS+ system in a chemiluminescent image while the ladder was visualized using a colorimetric image. Images were merged and labeled.
Figure imgf000077_0002
Figure imgf000077_0001
y g The antibodies were conjugated by the company AlphaThera, using their oYo-Link technology. This technology connects the oligonucleotides to the Fc chain of the antibody with a very specific photo crosslinker labeling called Light Activated Site-specific Conjugation or LASIC. Once conjugated, the antibodies are purified by fast protein liquid chromatography, FPLC, to separate unbound oligos from the antibodies as well as the antibodies with different degrees of labeling. Western-seq: Protocol 1. Cut the membrane for each desired band and sample and submerge in 38 ul of T4 ligase buffer (NEB B0202S) + 0.05% Tween. Boil at 95 °C for 5 minutes. 2. Add 1 ul of 1: 100 diluted l 00uM of the caliper (GG 940, 2 pmol), to make the final volume 39uL. Incubate at 37 °C for 30 minutes. 3. Make Klenow Mastermix according to the table below & add 10 ul of the Klenow master mix into each tube. Klenow Mastermix - List of reagents and amounts to use per sample in uL. Mastermix shows amounts for 6 samples.
Figure imgf000078_0001
25 °C, 15 minutes, 95 ºC, 5 minutes 5. Prepare PCT 1 Mastermix according to the table below. As a sample of the NTC. The product size should be 171 bp. The total amount per reaction should be 20 uL :1 uL sample (from Klenow extension). 19 uL mastermix and include membrane piece. Mastermix - List of reagents and amounts to use per sample in uL. Mastermix shows amounts for 7 samples: 6 true samples and 1 NTC. 6
Figure imgf000078_0002
. Run the 1st PCR with the following program: 1. 98 °C, 3:00 2. 98 °C, 0:10 3. 58 °C, 0:10 4. 72 °C, 0:10 5. GO TO 2, 17X 6. 72 °C, 2:00 7. 12 °C, infinite 7. Prepare PCR 2 Mastermix according to table below. Add another sample for the NTC. The product size should be 245 bp. The total amount per reaction should be 25 uL. 3uL sample (from 1st PCR), 19.5 uL mastermix, 1.25 uL forward primer and 1.25 uL reverse primer. PCR 2 Mastermix - List of reagents and amounts to use per sample in uL. Mastermix shows amounts for 8 samples: 6 true samples, l NTC from PCR 1, and 1 new NTC. 8
Figure imgf000079_0001
to the table below. PCR 2 Primer Pair Table. List of primer pairs and amounts to use per sample in uL with the identifying sequence. 9
Figure imgf000079_0002
. u e w e o ow g p og a : 1. 98 °C, 3 :00 2. 98 °C, 0: 10 3. 58 °C, 0: 10 4. 72 °C, 0:10 5. GO TO 2, 4X 6. 72 °C, 2:00 7. 12 °C, infinite 10. Prepare a 2% agarose gel with lX GelGreen Nucleic Acid Gel Stain (Biotium 41004) and take 12.5 uL from each sample and 2.5 uL TriTrack (ThermoFisher Scientific R1161) to prepare samples for running. Add 2uL of GeneRuler 50bp DNA Ladder (Fisher Scientific SJVW373) on either side of the samples and run at 130V for around 30 minutes. Image using Bio- Rad Chemi-Doc XRS+ system. 11. Extract each of your samples from the gel and purify using Zymoclean Gel DNA Recovery Kit (Zymo D4007). 12. Sequence your samples. Primer Sequences and Uses- List of primers used with the identifying sequence, lab code, and uses. Sequences shaded show the difference between sequences (SEQ ID NOS: 14-21)
Figure imgf000080_0001
Western-seq: Sequencing Samples were sequenced using Illumina's iSeq 100 benchtop sequencer. Libraries were prepped for sequencing as mentioned above, with paired end indices. The library was pooled to a concentration of 55 pM with 40% PhiX and 235,000 paired end reads were allocated per sample. Sequencing cartridge was thawed and prepared according to the manufacturer's instructions. Western-seq read analysis The Fastq files generated from sequencing were nm through a Python script in Jupyter Notebook. The length of the reads are determined, the barcodes are compared, the UMIs are recorded, and the UMIs are counted which leads to the proportion of UMIs and results in the proportion of each protein. Ladder To confirm size association in the sequence analysis portion of Western-seq., a DNA barcoded protein ladder for every 10 kDa increment in size can be used. For example, a standard commercial ladder can be used. One can also use a ladder system in which the system can include plasmids which express 10, 15, 20, 30, 40, 50, 60, 80 and 100 kD proteins in E. coli. Each protein migrates appropriately on SDS-PAGE gels, is expressed at very high levels (10–50 mg per liter of culture), is easy to purify via histidine tags and can be detected directly on Western blots via engineered immunoglobulin binding domains (see, for example, Santilli et al. The Penn State Protein Ladder System for inexpensive protein molecular weight markers. Scientific Reports volume 11, Article number: 16703 (2021). These can be run separately (in a separate lane on the gel) or with the protein sample itself. Bibliography 1. Janeway, CA J., Travers, P., Walport, M., & Shlomchik, M. J. (2001). The structure of a typical antibody molecule. In Immunobiology: The Immune System in Health and Disease. 5th edition. Garland Science. https://www.ncbi.nlrn.nih.gov/books/NBK27l44/ 2. Kubota, T., Niwa, R, Satoh, M., Akinaga, S, Shitara, K., & 1-fanai, N. (2009). Engineered therapeutic antibodies with improved effector functions. Cancer Science, 100(9), 1566-1572. https:/ /doi.org/10.111 l ~i .1349-7006.2009.0l 222.x 3. Weinstein, A (2021). Antibodies 101: Introduction to Antibodies. Retrieved December 21, 2023, from https://blog.addgene.org/antibodies-l O 1-introduction-to-antibodies. 4. Polyclonal vs Monoclonal Antibodies / Proteintech. (n.d.). Retrieved December 5, 2023, from https://www.ptglab.com/news/blog/polyclonal-vs-monoclonal-antibodies/ 5. Recombinant Antibody - 7 facts about recombinant antibodies. (2023, January 16). https://www.evitria.com/joumal/recombinant-antibodies/recornbinant-antihody/ 6. Immunoprecipitation - an overview / ScienceDirect Topics. (n.d.). Retrieved December 15, 2023. 7. Hussaini, H. M., Seo, B., & Rich, A. M. (2023). Immunohistochemistry and Immunofluorescence. Methods in Molecular Biology (Clifton, N.J.), 2588, 439---450. https://doi.org/10.1007 /978-1-0716-2780-8 ___ 26 8. Mahmood, T., & Yang, P -C (2012). Western Blot: Technique, Theory, and Trouble Shooting. North American Journal of Medical Sciences, 4(9), 429-434. https://doi.org/l 0.4103/194 7- 2714.100998 9. Introduction to Secondary Antibodies - US. (n.d.). Retrieved December 14, 2023, from https://www.thermofisher.com/us/en/home/life-science/antibodies/antihodies-learningcenter/ antibodies-resource-library/antibody-methods/introduction-secondary-antibodies.html 10. Gingrich, J. C, Davis, D. R, & Nguyen, Q. (2000). Multiplex Detection and Quantitation of Proteins on Westem Blots Using Fluorescent Probes. BioTechniques, 29(3), 636-642. https://doi.org/10.2144/00293pf02 l l. Grozeva, D., Carss, K., Spasic-Boskovic, 0., Parker, M. J., Archer, H., Firth, H. V., Park, S.M, Canham, N., Holder, S. E., Wilson, M., Hackett, A , Field, M ., Floyd, l .A. B , UK 1 OK Consortium, Hurles, M., & Raymond, F. L. (2014). De novo loss-of-function mutations in SETDS, encoding a methyltransferase in a 3p25 microdeletion syndrome critical region, cause intellectual disability. American Journal of Human Genetics, 94(4), 618-624. https:/ /doi.org/10.1016/j .ajhg.2014.03.006 12. Li, M., Hou, Y., Zhang, Z., Zhang, B., Huang, T., Sun, A., Shao, G., & Lin, Q. (2023). Structure, activity and function of the lysine methyltransferase SETD5. Frontiers in Endocrinology, 14, 1089527. https://doi.org/10.3389/fendo.2023.1089527 13. Q5XJV7 • SETD5 ___ MOUSE. (n.d.). UniProt. Retrieved January 17, 2024, from https://www.uniprot.org/uniprotkh/Q5XJV7/entry 14. Sessa, A., Fagnocchi, L, Mastrototaro, G., Massimino, L, Zaghi, M., Indrigo, NI., Cattaneo, S, Martini, D, Gabellini, C., Pucci, C, Fasciani, A., Belli, R., Taverna, S., Andreazzoli, M, Zippo, A., & Broccoli, V. (2019). SETD5 Regulates Chromatin Methylation State and Preserves Global Transcriptional Fidelity during Brain Development and Neuronal Wiring. Neuron, 104(2), 271- 289.e 13. https:/ /doi.org/10.1016/j.neuron.2019.07.013 15. Wang, Z., Hausmann, S., Lyu, R., Li, T.-M., Lofgren, S. M., Flores, N. M., Fuentes, M. E., Caporicci, M., Yang, Z., Meiners, JVL l, Cheek, M. A, Howard, S. A, Zhang, L., Elias, J. E., Kim, NI. P., Niaitra, A., Wang, H., Bassik, NI. C., Keogh, M.-C., Sage, J., Gozani, 0., Nfazur, P. K. (2020). SETD5-Coordinated Chromatin Reprogramming Regulates Adaptive Resistance to Targeted Pancreatic Cancer Therapy. Cancer Cell, 37(6), 834-849.eB. https:/ /doi.org/10.1016~j.ccell.2020.04.014 16. Alhajj, M., Zubair, M., & Farhana, A. (2023). Enzyme Linked Immunosorbent Assay. In StatPearls. StatPearls Publishing. http://w'vvw.ncbi.n1m.nih.gov/books/NBK555922/ 17. Aydin, S. (2015). A short history, principles, and types of ELISA, and our laboratory experience with peptide/protein analyses using ELISA. Peptides, 72, 4---15 https://doi.org/10.1016/j .peptides.2015.04.012 18. Wik, L, Nordberg, N., Broberg, l, Bj1)rkesten, J., Assarsson, E., Henriksson, S, Grundberg, I., Pettersson, E., Westerberg, C., Liljeroth, E., Falck, A., & Lundberg, M. (2021). Proximity Extension Assay in Combination with Next-Generation Sequencing for High-throughput Proteorne-wide Analysis. Molecular & Cellular Proteornics • MCP, 20, 100168. https://doi.org/10.1016/j .mcpro.2021.100168 19. Dhindsa, R S., Burren, 0. S., Sun, B. B., Prins, B. P., :Matelska, D., ·wheeler, E., :Mitchell, l, Oerton, E., Hristova, V. A., Smith, K. R., Carss, K., Wasilewski, S., Harper, A. R., Paul, D. S., Fabre, M. A, Runz, H., Viollet, C., Challis, B., Platt, A., AstraZeneca Genomics Initiative, Vitsios, D., Ashley, E. A., \Vhelan, C. D., Pangalos, M. N., \Vang, Q., Petrovski, S. (2023). Rare variant associations with plasma protein levels in the UK Biobank. Nature, 622(7982), 339---347. https:/ /doi .org/10.1038/s4 l 586-023-06547-x 20. Eldjarn, G. H., et al. (2023). Large-scale plasma proteomics comparisons through genetics and disease associations. Nature, 622(7982), 348-358. https://doi.org/10.1038/s41586-023-06563-x 2l. Sun, B. et al., (2023). Plasma proteomic associations with genetics and health in the UK Biobank. Nature, 622(7982), 329---338. https://doi.org/10.1038/s41586-023-06592-6 22. Haslam, D. E., et al. (2022). Stability and reproducibility of proteomic profiles in epidemiological studies: comparing the Olink and SOMAscan platforms. PROTEOMICS, 22(13- --14), 2100170. https://doi .org/10.1002/pmic.202100170 23. Gold, L., et al. (2010). Aptamer-Based Multiplexed Proteomic Technology for Biomarker Discovery. PLoS ONE, 5(12) e15004, https://doi.org/10.1371/journal.pone.0015004 24. van Buggenum, l A.G., Gerlach, J.P., Tanis, S. E. J., Hogeweg, M., Jansen, P. \V. T. C., Middelwijk, J., van der Steen, R., Vermeulen, M., Stunnenberg, H. G., Albers, C. A., & Mulder, K. W. (2018). Immuno-detection by sequencing enables large-scale high-dimensional phenotyping in cells. Nature Communications, 9(1), 2384. https://doi.org/10. l 038/s41467-018-04761-0 25. Kress, W. J., & Erickson, D. L. (2008). DNA barcodes: Genes, genomics, and bioinformatics. Proceedings of the National Academy of Sciences of the United States of America, 105(8), 2761- 2762. https://doi.org/10.1073/pnas.0800476105 26. Unique Molecular Identifiers (UMIs) I For sequencing accuracy. (n.d.). Retrieved January 9, 2024, from https://www.illumina.com/techniques/ sequencing/ngs-libraryprep/multiplexing/uni que-molecular-identifiers.html 27. Wang, X., Long, Y., Paucek, RD., Gooding, A. R., Lee, T., Burdorf, RM., & Cech, T. R. (2019). Regulation of hi sone methylation by automethylation of PRC2. Genes & Development, 33(19- 20), 1416-1427. https:/ /doi.org/10.110 l/gad.328849.119 28. AlQazzaz, M., & Edwards, A. (2024, February 5). We found a major flaw in a scientific reagent used in thousands of neuroscience experiments - and we're trying to fix it. The Transmitter: Neuroscience News and Perspectives. https://www.thetransmitter.org/openneuroscience-and-data- sharing/we-found-a-major-flaw-in-a-scientific-reagent-used-inthousands-of-neuroscience- experiments-and-were-trying-to-fix-it/ All publications, patent applications, patents, Gene Accession numbers and other references mentioned herein are expressly incorporated by reference in their entirety, to the same extent as if each were incorporated by reference individually. In case of conflict, the present specification, including definitions, will control.

Claims

WHAT IS CLAIMED IS: 1. A DNA-caliper comprising two single-stranded oligonucleotide arms, wherein the arms are joined at their 5’ ends resulting in each arm having a free 3’ end, wherein each 3’ end comprises a single stranded oligonucleotide anchor’ region, wherein each arm comprises a primer region, wherein at least one arm of the caliper comprises a cleavage site at its 5’ end.
2. The caliper of claim 1, wherein at least one arm of the DNA caliper further comprises a capture molecule.
3. The caliper of claim 2, wherein the capture molecule comprises biotin or digoxigenin.
4. The caliper of any one of claims 1 to 3, wherein the single stranded oligonucleotide anchor’ region is complementary to a single stranded oligonucleotide anchor region on another oligonucleotide.
5. The caliper of any one of claims 1 to 4, wherein each anchor’ region has the same sequence.
6. The caliper of any one of claims 1 to 4, wherein each anchor’ region has a different sequence.
7. The caliper of any one of claims 1 to 6, wherein each arm is about 15 to about 1000 nucleotides long.
8. The caliper of any one of claims 1 to 7, wherein each single stranded oligonucleotide anchor’ is about 10 to about 50 nucleotides long.
9. The caliper of any one of claims 1 to 8, wherein each primer region is about 5 to about 45 nucleotides long.
10. The caliper of any one of claims 1 to 9, wherein the cleavage site is a restriction enzyme site, a deoxy uridine (dU) or a UV light cleavage site.
11. A composition comprising the DNA-caliper of any one of claims 1 to 10 and a carrier.
12. An antibody covalently linked to a 5’ end of a single-stranded oligonucleotide tail, wherein the oligonucleotide tail comprises at its 3’ end a single stranded oligonucleotide anchor region, a unique barcode, and a single stranded oligonucleotide unique molecular identifier (UMI) and at its 5’ end a restriction enzyme site.
13. The antibody of claim 12, wherein the single stranded oligonucleotide anchor region is complementary to a single stranded oligonucleotide anchor’ region on another molecule.
14. The antibody of claim 12 or 13, wherein the barcode is specific to the antibody.
15. The antibody of any one of claims 12 to 14, wherein the UMI is a random sequence.
16. The antibody of claim 15, wherein the UMI is about 5 to about 45 nucleotides long.
17. The antibody of any one of claims 12 to 16, wherein the single-stranded oligonucleotide tail is about 15 to about 1000 nucleotides long.
18. The antibody of any one of claims 12 to 17, wherein the anchor is about 10 to about 50 nucleotides long.
19. A composition comprising the antibody of any one of claims 12 to 18 and a carrier.
20. A composition comprising the DNA-caliper of any one of claims 1 to 10, the antibody of any one of claims 12-18 and a carrier.
21. A method for mapping the genomic co-localization of one or more proteins on a chromatin fragment comprising: (a) incubating chromatin fragments with a plurality of antibodies, each of the plurality of antibodies binding to each of the one or more proteins, wherein each antibody comprises a single stranded oligonucleotide tail comprising a hook region, a single stranded oligonucleotide anchor region, a unique barcode, and a single stranded oligonucleotide unique molecular identifier (UMI); and wherein each chromatin fragment comprises a chromatin DNA molecule ligated to adaptors comprising a chromatin DNA molecule anchor’ region and a chromatin DNA molecule UMI; (b) contacting the chromatin fragments with a DNA-caliper molecule comprising two single- stranded oligonucleotide arms, wherein the arms are joined at their 5’ ends resulting in each arm having a free 3’ end, wherein the first 3’ end comprises a first sequence complementary to the single stranded oligonucleotide anchor region and the second 3’ end comprises a second sequence complementary to the chromatin DNA molecule anchor’ region, wherein at least one arm comprises a cleavage site and/or a capture molecule, and wherein each arm comprises a primer sequence; (c) hybridizing the DNA-caliper first sequence with the single stranded oligonucleotide anchor region and the DNA-caliper second sequence with the chromatin DNA molecule anchor’ region to obtain a hybridized product; (d) contacting the hybridized product with a DNA polymerase to obtain an extended DNA- caliper that covalently captures the sequences of the antibody unique barcode, the antibody UMI, and the chromatin DNA molecule.
22. The method of claim 21, further comprising denaturing the extended DNA-caliper from the chromatin fragment.
23. The method of claim 21 or 22, wherein the capture molecule comprises biotin or digoxigenin.
24. The method of claim 23, wherein the extended DNA-calipers are captured/isolated by the capture molecule.
25. The method of claim 24, further comprising adding a polyC sequence with terminal transferase to the 3’ end of the extended DNA-caliper corresponding to the chromatin fragment and then hybridizing and ligating a double stranded DNA molecule with two overhanging regions to the hook region and the polyC overhang on the extended DNA-caliper, bringing the two ends of the extended DNA-caliper together, wherein the 3’ ends of DNA-caliper are extended with polymerase using the other arm as a template to generate a double stranded construct.
26. The method of any one of claims 21 to 25, wherein the cleavage site is a restriction enzyme site, dU or UV light cleavage site.
27. The method of claim any one of claims 21 to 26, further comprising linearizing the double stranded construct at the cleavage site to generate a linearized fragment.
28. The method of claim 27, wherein the linearized fragment is PCR amplified using the primer sequences on the DNA-caliper.
29. The method of claim 28, wherein the PCR amplified products are sequenced.
30. The method of any one of claims 21 to 29, wherein the single stranded oligonucleotide has a nucleic acid sequence of SEQ ID NO: 1.
31. The method of any one of claims 21 to 30, wherein the DNA-caliper has a nucleic acid sequence of SEQ ID NOs: 2 and 3.
32. The method of any one of claims 21 to 29 and 31, wherein the single stranded oligonucleotide has a nucleic acid sequence of SEQ ID NO: 4.
33. The method of any one of claims 21 to 29 and 32, wherein the DNA-caliper has a nucleic acid sequence of SEQ ID NOs: 5 and 6.
34. A method for detecting and/or quantifying at least one of multiple protein-protein interactions (PPIs) and/or protein abundance comprising: (a) incubating fixed and permeabilized cells with a plurality of antibody molecules, wherein each antibody molecule is covalently linked to a single stranded oligonucleotide tail comprising an antibody-specific unique barcode, a single stranded oligonucleotide unique molecular identifier (UMI), a restriction enzyme site, a single stranded oligonucleotide anchor’ region complementary to both arms of a DNA-caliper anchor region and optionally a hook region; and (b) contacting the plurality of antibody molecules with one or more single stranded oligonucleotide DNA-caliper molecules comprising two single-stranded oligonucleotide arms, wherein the arms are joined at their 5’ ends resulting in each arm having a free 3’ end, a capture molecule, a cleavage site, a first primer and a second primer sequence, wherein one primer sequence is on each side of the cleavage site, wherein the first 3’ end of the caliper comprises an anchor region that is complementary to the single stranded oligonucleotide anchor’ region of a first antibody and the second 3’ end of the caliper comprises an anchor region that is complementary to the single stranded oligonucleotide anchor’ region of a second antibody; (c) hybridizing the DNA-caliper with the single stranded oligonucleotide anchor’ region of the first antibody and the single stranded oligonucleotide anchor’ region of the second antibody to obtain a hybridized product; (d) contacting the hybridized product with a DNA polymerase to obtain a bidirectionally extended DNA-caliper that covalently links the sequences of the first antibody unique barcode and first antibody UMI and the second antibody unique barcode and second antibody UMI.
35. Them method of claim 34, wherein the permeabilized cells are lysed cells.
36. The method of claim 34 or 35, wherein the cleavage site is a restriction endonuclease site, dU or UV light cleavage site.
37. The method of any one of claims 34 to 36, further comprising denaturing the extended DNA-calipers from the single stranded oligonucleotide of the first antibody and the single stranded oligonucleotide of the second antibody or prior to or after adding DNA polymerase remove the antibodies by restriction enzyme cutting and ligating the ends; after DNA polymerase extension of the DNA-caliper ends and the oligonucleotide tails the extended molecule can be linearized, PCR amplified and sequenced.
38. The method of any one of claims 34 to 37, wherein the capture molecule comprises biotin or digoxigenin.
39. The method of any one of claims 34 to 38, wherein the extended DNA-calipers are captured/isolated by the capture molecule.
40. The method of claim 39, further comprising hybridizing and ligating a double stranded DNA molecule with two overhanging regions to the hook regions on the extended DNA-caliper bringing the two ends of the DNA-caliper together, wherein the 3’ ends of DNA-caliper are extended with DNA polymerase using the other arm as a template to generate a double stranded construct.
41. The method of claim 40, further comprising linearizing the double stranded construct at the cleavage site to generate a linearized fragment.
42. The method of claim 41, wherein the linearized fragment is PCR amplified using the primer sequences on the DNA-caliper.
43. The method of claim 42, wherein the PCR amplified products are sequenced so as to identify the PPIs and their abundance through UMI counting.
44. A method of determining abundance of a target protein in a sample comprising: a) contacting a sample suspected of comprising one or more target proteins with one or more antibodies that binds to at least one of the target proteins in the sample, wherein each of the one or more antibodies comprises a single stranded oligonucleotide tail; b) detecting, isolating and PCR amplifying the single stranded oligonucleotide of the antibody bound to the protein, wherein the single stranded oligonucleotide tail is a template oligonucleotide to generate a first PCR product comprising a double stranded version of the single stranded oligonucleotide tail; performing a second PCR amplification using the first PCR product as a template, wherein the second PCR generates a second PCR product; sequencing the second PCR product; and determining the abundance of the single stranded oligonucleotide sequence indicating protein abundance of the target protein in the sample.
45. The method of claim 44, wherein the single stranded oligonucleotide tail comprises one or more unique molecular identifier (UMI) sequences.
46. The method of claim 44 or 45, wherein the single stranded oligonucleotide tail comprises a barcode sequence or molecular tag.
47. The method of claim 45 or 46, wherein determining the abundance of the target protein comprises detecting the presence of the UMI sequence.
48. The method of any one of claims 44 to 47, further comprising separating the proteins in the sample by gel electrophoresis and transferring the separated proteins to a membrane prior to contacting the sample with the one or more antibodies.
49. The method of claim 48, further comprising contacting the one or more antibodies comprising one or more single stranded oligonucleotide tails with a single stranded binding protein prior to incubating the membrane with the one or more antibodies.
50. The method of claim 48, further comprising a protein ladder of known size.
51. The method of claim 48 or 49, further comprising contacting the membrane with one or more antibodies and cutting the membrane into segments.
52. The method of claims 51, comprising using the protein ladder of known size such that each segment contains about a 10kDa to about a 20kDa range of sizes of protein.
53. The method of claim any one of claims 45 to 52, wherein the abundance of one or more proteins in the sample is performed by comparing the relative abundance of the one or more UMI sequences from the sequencing results, wherein the one or more UMI sequences are specific to the antibody that targets a protein.
54. The method of any one of claims 48 to 53, wherein the sample is sonicated prior to gel electrophoresis.
55. A kit comprising: one or more primary antibodies against one or more target proteins, wherein the one or more primary antibodies comprise a single stranded oligonucleotide tail and wherein the single stranded oligonucleotide tail comprises a primer region, an antibody-specific unique barcode, a single stranded oligonucleotide unique molecular identifier (UMI), a 3’ single stranded oligonucleotide anchor’ region, a restriction enzyme site and optionally a single stranded binding protein; and instructions for use thereof for methods of determining protein abundance in a sample.
56. The kit of claim 55 further comprising one or more single stranded oligonucleotide DNA- caliper molecules comprising two single-stranded oligonucleotide arms, wherein the arms are joined at their 5’ ends resulting in each arm having a free 3’ end, a capture molecule, a cleavage site, a first primer and a second primer sequence, wherein one primer sequence is on each side of the cleavage site.
57. The kit of claim 55 or 56, further comprising a tool for cutting a protein membrane.
58. The kit of claim 57, wherein the tool is a pair of scissors, blade, knife or cookie cutter type device.
59. The kit of any one of claims 55 to 58 further comprising a first set of PCR primers and a second set of PCR primers, wherein the first set of PCR primers comprises a first sequence and are formulated for use in a first PCR reaction that converts the single stranded oligonucleotide to a double stranded oligonucleotide, and wherein the second set of PCR primers comprises a second sequence and are formulated for use in a second PCR reaction that generates double stranded oligonucleotides for sequencing.
60. A method comprising: a) perform PQ-seq on the sample and collect protein quantification data, b) perform Prod-seq on the sample and collect protein-protein interaction data, c) determining, by a computing system having processing resources and memory, a true interaction using the following equation: ^^ ∼ ^^^^^^^^ ^^^^^^^^ ^^^ℎ ^^^^ ^^^ ^^^ + ^^ ^ ^^^ ^^^^^^^^^^ ^^^^^^^^^ ^ !
Figure imgf000093_0001
binomial
Figure imgf000093_0002
λ23 = ^45%67 ^α2 + ^2 ^3 + ^3) + ^6%:4;<α2α3 wherein ^ >5? as
Figure imgf000093_0003
wherein +, denotes the proportion of true interactions in the entire dataset; wherein ^^^ is the readout scaling factor for PQ-seq detection protocol and sequencing; ^6%:4;< and ^45%67 denote the readout scaling factors for specific and non-specific PPI detection in Prod-seq, respectively; ^% and ^% denote the levels of specific and non-specific antibody-oligo binding for each protein target ^, respectively, wherein +,, ^^^, ^6%:4;< , ^45%67, ^A,...,4 , ^A,...,4 , ^ !, ^ >5? are estimated from the maximum likelihood estimation of "^$, ^| ⋅^.
61. The method of claim 60, wherein after fitting the maximum likelihood model, the model X λ calculates the confidence score of each PPI as Score F G ^H 23 I 23 J 23 = P#true'X23 ) = FG^HX23Iλ23JK^ALFG^^HX23Iθ23J .
PCT/US2024/037959 2023-07-14 2024-07-14 Tools for interrogating dynamic organizational principles of protein complexes in vivo Ceased WO2025019385A1 (en)

Applications Claiming Priority (4)

Application Number Priority Date Filing Date Title
US202363513630P 2023-07-14 2023-07-14
US63/513,630 2023-07-14
US202463561661P 2024-03-05 2024-03-05
US63/561,661 2024-03-05

Publications (1)

Publication Number Publication Date
WO2025019385A1 true WO2025019385A1 (en) 2025-01-23

Family

ID=94282652

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2024/037959 Ceased WO2025019385A1 (en) 2023-07-14 2024-07-14 Tools for interrogating dynamic organizational principles of protein complexes in vivo

Country Status (1)

Country Link
WO (1) WO2025019385A1 (en)

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2015118029A1 (en) * 2014-02-04 2015-08-13 Olink Ab Proximity assay with detection based on hybridisation chain reaction (hcr)
US10655162B1 (en) * 2016-03-04 2020-05-19 The Broad Institute, Inc. Identification of biomolecular interactions

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2015118029A1 (en) * 2014-02-04 2015-08-13 Olink Ab Proximity assay with detection based on hybridisation chain reaction (hcr)
US10655162B1 (en) * 2016-03-04 2020-05-19 The Broad Institute, Inc. Identification of biomolecular interactions

Similar Documents

Publication Publication Date Title
US12320813B2 (en) Macromolecule analysis employing nucleic acid encoding
Lei et al. Mitochondrial base editor induces substantial nuclear off-target mutations
Gorbovytska et al. Enhancer RNAs stimulate Pol II pause release by harnessing multivalent interactions to NELF
AU2020346959B2 (en) Methods and compositions for protein and peptide sequencing
US20230236198A1 (en) Kits for analysis using nucleic acid encoding and/or label
Sun et al. Genetically encoded chemical crosslinking of RNA in vivo
US20150011398A1 (en) Methods for quantitative determination of protein-nucleic acid interactions in complex mixtures
US11926820B2 (en) Methods and compositions for protein and peptide sequencing
US11834756B2 (en) Methods and compositions for protein and peptide sequencing
CN107109698B (en) RNA STITCH sequencing: an assay for direct mapping of RNA:RNA interactions in cells
US20240209378A1 (en) Methods and compositions for protein and peptide sequencing
US10655162B1 (en) Identification of biomolecular interactions
Hong et al. ProtSeq: toward high-throughput, single-molecule protein sequencing via amino acid conversion into DNA barcodes
US12560608B2 (en) Detection of molecular associations
US20210102248A1 (en) Methods and compositions for protein and peptide sequencing
US20200407467A1 (en) Protease activity profiling via programmable phage display of comprehensive proteome-scale peptide libraries
Zhong et al. Modular DNA barcoding of nanobodies enables multiplexed in situ protein imaging and high-throughput biomolecule detection
US20240352452A1 (en) Crispr-based protein barcoding and surface assembly
Zaripov et al. Simultaneous Detection of Protein and RNA Using Proximity Ligation of Aptamers and Quantitative Polymerase Chain Reaction
Zaripov et al. Simultaneous Detection of SARS-CoV-2 Nucleocapsid Protein and RNA by Aptamer-Based Proximity Ligation and Quantitative PCR
Tasca Investigation of Functional Protein-RNA Interactions by Mass Spectrometry
Melendez Developments in Proteomics, Trans-Splicing Technology And Endogenous Transcript Manipulation
Bastiaanssen Exploring a new dimension: Single-molecule interaction studies in sequence space
WO2024197298A1 (en) Methods for tagging molecules
Nguyen Development of high-throughput technologies to map RNA structures and interactions

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24843791

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE