WO2023213993A1 - Variants of the gain domain of adhesion g protein-coupled receptors - Google Patents
Variants of the gain domain of adhesion g protein-coupled receptors Download PDFInfo
- Publication number
- WO2023213993A1 WO2023213993A1 PCT/EP2023/061899 EP2023061899W WO2023213993A1 WO 2023213993 A1 WO2023213993 A1 WO 2023213993A1 EP 2023061899 W EP2023061899 W EP 2023061899W WO 2023213993 A1 WO2023213993 A1 WO 2023213993A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- amino acid
- acid position
- seq
- replaced
- cysteines
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C07—ORGANIC CHEMISTRY
- C07K—PEPTIDES
- C07K14/00—Peptides having more than 20 amino acids; Gastrins; Somatostatins; Melanotropins; Derivatives thereof
- C07K14/435—Peptides having more than 20 amino acids; Gastrins; Somatostatins; Melanotropins; Derivatives thereof from animals; from humans
- C07K14/705—Receptors; Cell surface antigens; Cell surface determinants
- C07K14/72—Receptors; Cell surface antigens; Cell surface determinants for hormones
- C07K14/723—G protein coupled receptor, e.g. TSHR-thyrotropin-receptor, LH/hCG receptor, FSH receptor
Definitions
- the present invention relates to variants of the GAIN domain (G-protein-coupled receptor autoproteolysis-inducing domain) of an adhesion G protein-coupled receptor (ADGRG), such as a GAIN domain variant comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG1 (Adhesion G Protein-Coupled Receptor G1) of SEQ ID NO: 1 , wherein (i) one of amino acid positions 3 to 5, preferably amino acid position 4 and one of amino acid positions 47 to 49, preferably amino acid position 48 of SEQ ID NO: 1 are replaced by cysteines, (ii) one of amino acid positions 35 to 37, preferably amino acid position 36 and one of amino acid positions 180 to 182, preferably amino acid position 181 of SEQ ID NO: 1 are replaced by cysteines, (iii) one of amino acid positions 19 to 21 , preferably amino acid position 20 and one of amino acid positions 124 to 126, preferably amino acid position 125 of SEQ ID NO: 1 are replaced by cysteines,
- Cells can sense the mechanical properties of their local environment.
- Cellular mechanosensing is currently thought to primarily involve integrins, which couple the extracellular matrix (ECM) to the contractile cytoskeleton and form force-bearing connections that require tensile forces to mature 1-3 , as well as by mechano-activated transmembrane ion channels 45 .
- ECM extracellular matrix
- Both sensing modalities are currently exploited as drug targets and for bioengineering applications 67 .
- GPCRs G-protein-coupled receptors
- G-protein-coupled receptors have large extracellular domains that can bind to ECM components 8-10 .
- aGPCRs or AGPCRs adhesion GPCRs
- a general feature of these aGPCR molecules is the highly conserved GAIN domain located on the extracellular side of the receptor, between the ligand binding domain (LBD) and the 7 transmembrane (7TM) domain.
- LBD ligand binding domain
- 7TM 7 transmembrane
- Adhesion Receptors Ars, such as integrins and cadherins
- EC extra-cellular
- Kinases associate indirectly with integrins through peripheral adaptor proteins for signal regulation. Due to the complexity of the systems, reprogramming integrin or cadherin functions has remained very challenging.
- Protein topology and the orientation of chemical bonds with respect to the direction of the mechanical force propagating inside the structure are key determinants of protein mechanical resistance, but these principles have yet to be recapitulated in protein design. Forcebearing non-covalent interactions will deform and break eventually under force, and computational methods to engineer specific networks of chemical interactions are now well established and should be directly applicable to the design of mechano-resistant protein structures and protein-ligand binding interactions.
- GPCRs play a very important role as drug target.
- adhesion GPCRs aGPCRs
- GAIN domain of aGPCRs harbors the GPCR proteolysis site at which the receptor is autoproteolytically cleaved into an N- and a C-terminal fragment which trigger signalling cascades. This knowledge on aGPCRS leads to the hypothesis that they “sense” in particular “mechanosense” via the GAIN domain their environment and contribute to a context-dependent signal transduction.
- the present invention aims at the provision of a method to reprogram the response of the human GAIN domains of various aGPCRS to mechanical force as well as the provision of such reprogrammed GAIN domains as well as aGPCRs comprising the same.
- Such methods and GAINS domains are helpful to study context-dependent signal transduction and consequently helpful for applications such as cell signaling studies and drug development (see also Bassilana et al. (2019), Nature Reviews Drug Discovery, 18:869-884).
- the present invention relates to variants of the GAIN domain (G-protein-coupled receptor autoproteolysis-inducing domain) of an adhesion G protein-coupled receptor (ADGRG):
- one of amino acid positions 56 to 58, preferably amino acid position 57 and one of amino acid positions 65 to 67, preferably amino acid position 66 of SEQ ID NO: 1 are replaced by cysteines
- one of amino acid positions 16 to 18, preferably amino acid position 17 and one of amino acid positions 60 to 62, preferably amino acid position 61 of SEQ ID NO: 1 are replaced by cysteines
- amino acid position at position 40 is replaced by glutamic acid
- amino acid position at position 43 is replaced by leucine
- amino acid position at position 181 is replaced by arginine
- amino acid position at position 40 is replaced by glutamic acid
- amino acid position at position 43 is replaced by leucine
- amino acid position at position 135 is replaced by arginine
- amino acid position at position 181 is replaced by arginine
- amino acid position at position 40 is replaced by glutamic acid
- amino acid position at position 43 is replaced by leucine
- amino acid position at position 135 is replaced by arginine
- amino acid position at position 146 is replaced by tyrosine
- amino acid position at position 181 is replaced by arginine
- amino acid position at position 33 is replaced by tryptophane
- amino acid position at position 40 is replaced by glutamic acid
- amino acid position at position 43 is replaced by leucine
- amino acid position at position 135 is replaced by arginine
- amino acid position at position 146 is replaced by tyrosine
- amino acid position at position 178 is replaced by tyrosine
- amino acid position at position 181 is replaced by arginine
- amino acid position at position 184 is replaced by glycine
- amino acid position at position 186 is replaced by glutamine
- amino acid position at position 187 is replaced by glycine
- (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii), (iv) and/or (v), or the amino acid replacement(s) as defined in (A), item (vi), (vii), (viii), (ix), (x) and/or (xi), or the cysteines and/or the amino acid replacement(s) as defined in (A), item (xii) are retained;
- amino acid positions 134 to 136 preferably amino acid position 135 and one of amino acid positions 144 to 146, preferably amino acid position 145 of SEQ ID NO: 2 are replaced by cysteines,
- amino acid positions 128 to 130 preferably amino acid position 129 and one of amino acid positions 138 to 140, preferably amino acid position 139 of SEQ ID NO: 3 are replaced by cysteines,
- V (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRL1 (Adhesion G Protein-Coupled Receptor L1) of SEQ ID NO: 5, wherein
- one of amino acid positions 118 to 120, preferably amino acid position 119 and one of amino acid positions 275 to 277, preferably amino acid position 276 of SEQ ID NO: 5 are replaced by cysteines
- one of amino acid positions 102 to 104, preferably amino acid position 103 and one of amino acid positions 202 to 204, preferably amino acid position 203 of SEQ ID NO: 5 are replaced by cysteines
- VI (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG7 (Adhesion G Protein-Coupled Receptor G7) of SEQ ID NO: 6, wherein
- amino acid positions 90 to 92 preferably amino acid position 91 and one of amino acid positions 246 to 248, preferably amino acid position 247 of SEQ ID NO: 6 are replaced by cysteines,
- one of amino acid positions 74 to 76, preferably amino acid position 75 and one of amino acid positions 186 to 188, preferably amino acid position 187 of SEQ ID NO: 6 are replaced by cysteines
- one of amino acid positions 64 to 66, preferably amino acid position 65 and one of amino acid positions 111 to 113, preferably amino acid position 112 of SEQ ID NO: 6 are replaced by cysteines, or
- VIS (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG6 (Adhesion G Protein-Coupled Receptor G6) of SEQ ID NO: 7, wherein
- amino acid positions 34 to 36 preferably amino acid position 35 and one of amino acid positions 103 to 105, preferably amino acid position 104 of SEQ ID NO: 7 are replaced by cysteines,
- VIII (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG5 (Adhesion G Protein-Coupled Receptor G5) of SEQ ID NO: 8, wherein
- amino acid positions 27 to 29, preferably amino acid position 28 and one of amino acid positions 70 to 72, preferably amino acid position 71 of SEQ ID NO: 8 are replaced by cysteines, or
- amino acid positions 274 to 276, preferably amino acid position 275 and one of amino acid positions 385 to 387, preferably amino acid position 386 of SEQ ID NO: 9 are replaced by cysteines
- amino acid positions 311 to 313, preferably amino acid position 312 and one of amino acid positions 320 to 322, preferably amino acid position 321 of SEQ ID NO: 9 are replaced by cysteines
- (X) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG3 (Adhesion G Protein-Coupled Receptor G3) of SEQ ID NO: 10, wherein
- one of amino acid positions 29 to 31 , preferably amino acid position 30 and one of amino acid positions 59 to 61 , preferably amino acid position 60 of SEQ ID NO: 10 are replaced by cysteines, or
- amino acid positions 72 to 74 preferably amino acid position 73 and one of amino acid positions 81 to 83, preferably amino acid position 82 of SEQ ID NO: 11 are replaced by cysteines,
- amino acid positions 32 to 34 preferably amino acid position 33 and one of amino acid positions 76 to 78, preferably amino acid position 77 of SEQ ID NO: 11 are replaced by cysteines, or
- amino acid positions 126 to 128, preferably amino acid position 127 and one of amino acid positions 135 to 137, preferably amino acid position 136 of SEQ ID NO: 14 are replaced by cysteines
- XV (XV) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF2 (Adhesion G Protein-Coupled Receptor F2) of SEQ ID NO: 15, wherein
- XVI (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF1 (Adhesion G Protein-Coupled Receptor F1) of SEQ ID NO: 16, wherein
- amino acid positions 128 to 130 preferably amino acid position 129 and one of amino acid positions 136 to 138, preferably amino acid position 137 of SEQ ID NO: 16 are replaced by cysteines,
- (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
- (XVII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE5 (Adhesion G Protein-Coupled Receptor E5) of SEQ ID NO: 17, wherein
- amino acid positions 96 to 98 preferably amino acid position 97 and one of amino acid positions 109 to 111 , preferably amino acid position 110 of SEQ ID NO: 17 are replaced by cysteines, or
- (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), (i), (ii) or (iii) are retained;
- (XVIII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE3 (Adhesion G Protein-Coupled Receptor E3) of SEQ ID NO: 18, wherein one of amino acid positions 72 to 74, preferably amino acid position 73 and one of amino acid positions 83 to 85, preferably amino acid position 84 of SEQ ID NO: 18 are replaced by cysteines; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A) are retained;
- (XIX) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE3 (Adhesion G Protein-Coupled Receptor E3) of SEQ ID NO: 19, wherein one of amino acid positions 73 to 75, preferably amino acid position 74 and one of amino acid positions 186 to 188, preferably amino acid position 187 of SEQ ID NO: 19 are replaced by cysteines; or
- XX (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE2 (Adhesion G Protein-Coupled Receptor E2) of SEQ ID NO: 20, wherein one of amino acid positions 72 to 74, preferably amino acid position 73 and one of amino acid positions 228 to 230, preferably amino acid position 229 of SEQ ID NO: 20 are replaced by cysteines,
- amino acid positions 96 to 98 preferably amino acid position 97 and one of amino acid positions 109 to 111 , preferably amino acid position 110 of SEQ ID NO: 20 are replaced by cysteines, or
- (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), (i), (ii) or (iii) are retained;
- XXI (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE1 (Adhesion G Protein-Coupled Receptor E1) of SEQ ID NO: 21 , wherein
- (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
- (XXII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRD2 (Adhesion G Protein-Coupled Receptor D2) of SEQ ID NO: 22, wherein
- amino acid positions 80 to 82 preferably amino acid position 81 and one of amino acid positions 127 to 129, preferably amino acid position 128 of SEQ ID NO: 22 are replaced by cysteines or
- (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (v) are retained;
- (XXIII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRD1 (Adhesion G Protein-Coupled Receptor D1) of SEQ ID NO: 23, wherein (i) one of amino acid positions 36 to 38, preferably amino acid position 37 and one of amino acid positions 92 to 94, preferably amino acid position 93 of SEQ ID NO: 23 are replaced by cysteines,
- amino acid positions 104 to 106 preferably amino acid position 105 and one of amino acid positions 116 to 118, preferably amino acid position 117 of SEQ ID NO: 23 are replaced by cysteines,
- (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
- (XXIV) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRB3 (Adhesion G Protein-Coupled Receptor B3) of SEQ ID NO: 24, wherein
- amino acid positions 133 to 135, preferably amino acid position 134 and one of amino acid positions 142 to 144, preferably amino acid position 143 of SEQ ID NO: 24 are replaced by cysteines
- (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
- (XXV) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRB2 (Adhesion G Protein-Coupled Receptor B2) of SEQ ID NO: 25, wherein
- one of amino acid positions 132 to 134, preferably amino acid position 133 and one of amino acid positions 141 to 143, preferably amino acid position 142 of SEQ ID NO: 25 are replaced by cysteines
- one of amino acid positions 92 to 94, preferably amino acid position 93 and one of amino acid positions 137 to 139, preferably amino acid position 138 of SEQ ID NO: 25 are replaced by cysteines or
- (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
- (XXVI) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRB1 (Adhesion G Protein-Coupled Receptor B1) of SEQ ID NO: 26, wherein
- amino acid positions 128 to 130 preferably amino acid position 129 and one of amino acid positions 136 to 138, preferably amino acid position 137 of SEQ ID NO: 26 are replaced by cysteines,
- (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
- (XXVII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRA3 (Adhesion G Protein-Coupled Receptor A3) of SEQ ID NO: 27, wherein
- (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
- (XXVIII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRA2 (Adhesion G Protein-Coupled Receptor A2) of SEQ ID NO: 28, wherein (i) one of amino acid positions 57 to 59, preferably amino acid position 58 and one of amino acid positions 104 to 106, preferably amino acid position 105 of SEQ ID NO: 28 are replaced by cysteines,
- amino acid positions 92 to 94 preferably amino acid position 93 and one of amino acid positions 294 to 296, preferably amino acid position 295 of SEQ ID NO: 28 are replaced by cysteines,
- XXIX (A) comprising or consisting of the amino acid sequence of GAIN domain of ADGRV1 (Adhesion G Protein-Coupled Receptor V1) of SEQ ID NO: 29, wherein
- one of amino acid positions 199 to 201 preferably amino acid position 200 and one of amino acid positions 239 to 241 , preferably amino acid position 240 of SEQ ID NO: 29 are replaced by cysteines, or
- the present invention relates to a variant of the GAIN domain (G-protein-coupled receptor autoproteolysis-inducing domain) of an adhesion G protein- coupled receptor (ADGRG)
- GAIN domain G-protein-coupled receptor autoproteolysis-inducing domain
- ADGRG adhesion G protein- coupled receptor
- (I) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG1 (Adhesion G Protein-Coupled Receptor G1) of SEQ ID NO: 1 , wherein (i) one of amino acid positions 3 to 5, preferably amino acid position 4 and one of amino acid positions 47 to 49, preferably amino acid position 48 of SEQ ID NO: 1 are replaced by cysteines, (ii) one of amino acid positions 35 to 37, preferably amino acid position 36 and one of amino acid positions 180 to 182, preferably amino acid position 181 of SEQ ID NO: 1 are replaced by cysteines, (iii) one of amino acid positions 19 to 21 , preferably amino acid position 20 and one of amino acid positions 124 to 126, preferably amino acid position 125 of SEQ ID NO: 1 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided
- (II) (A)comprising or consisting of the amino acid sequence of the GAIN domain of ADGRL4 (Adhesion G Protein-Coupled Receptor L4) of SEQ ID NO: 2, wherein(i) one of amino acid positions 87 to 89, preferably amino acid position 88 and one of amino acid positions 123 to 125, preferably amino acid position 124 of SEQ ID NO: 2 are replaced by cysteines, (ii) one of amino acid positions 111 to 113, preferably amino acid position 112 and one of amino acid positions 263 to 265, preferably amino acid position 264 of SEQ ID NO: 2 are replaced by cysteines, (iii) one of amino acid positions 95 to 97, preferably amino acid position 96 and one of amino acid positions 192 to 194, preferably amino acid position 193 of SEQ ID NO: 2 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least
- (III)(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRL3 (Adhesion G Protein-Coupled Receptor L3) of SEQ ID NO: 3, wherein (i) one of amino acid positions 81 to 83, preferably amino acid position 82 and one of amino acid positions 117 to 119, preferably amino acid position 118 of SEQ ID NO: 3 are replaced by cysteines, (ii) one of amino acid positions 105 to 107, preferably amino acid position 106 and one of amino acid positions 261 to 263, preferably amino acid position 262 of SEQ ID NO: 3 are replaced by cysteines, (iii) one of amino acid positions 89 to 91 , preferably amino acid position 90 and one of amino acid positions 188 to 190, preferably amino acid position 189 of SEQ ID NO: 3 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least
- (IV) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRL2 (Adhesion G Protein-Coupled Receptor L2) of SEQ ID NO: 4, wherein(i) one of amino acid positions 94 to 96, preferably amino acid position 95 and one of amino acid positions 130 to 132, preferably amino acid position 131 of SEQ ID NO: 4 are replaced by cysteines, (ii) one of amino acid positions 118 to 120, preferably amino acid position 119 and one of amino acid positions 272 to 274, preferably amino acid position 273 of SEQ ID NO: 4 are replaced by cysteines, (iii) one of amino acid positions 102 to 104, preferably amino acid position 103 and one of amino acid positions 202 to 204, preferably amino acid position 203 of SEQ ID NO: 4 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity
- (V) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRL1 (Adhesion G Protein-Coupled Receptor L1) of SEQ ID NO: 5, wherein (i) one of amino acid positions 94 to 96, preferably amino acid position 95 and one of amino acid positions 130 to 132, preferably amino acid position 131 of SEQ ID NO: 5 are replaced by cysteines, (ii) one of amino acid positions 118 to 120, preferably amino acid position 119 and one of amino acid positions 275 to 277, preferably amino acid position 276 of SEQ ID NO: 5 are replaced by cysteines, (iii) one of amino acid positions 102 to 104, preferably amino acid position 103 and one of amino acid positions 202 to 204, preferably amino acid position 203 of SEQ ID NO: 5 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence
- (VI) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG7 (Adhesion G Protein-Coupled Receptor G7) of SEQ ID NO: 6, wherein (i) one of amino acid positions 67 to 69, preferably amino acid position 68 and one of amino acid positions 102 to 104, preferably amino acid position 103 of SEQ ID NO: 6 are replaced by cysteines, (ii) one of amino acid positions 90 to 92, preferably amino acid position 91 and one of amino acid positions 246 to 248, preferably amino acid position 247 of SEQ ID NO: 6 are replaced by cysteines, (iii) one of amino acid positions 74 to 76, preferably amino acid position 75 and one of amino acid positions 186 to 188, preferably amino acid position 187 of SEQ ID NO: 6 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence
- (VII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG6 (Adhesion G Protein-Coupled Receptor G6) of SEQ ID NO: 7, wherein(i) one of amino acid positions 35 to 37, preferably amino acid position 36 and one of amino acid positions 84 to 86, preferably amino acid position 85 of SEQ ID NO: 7 are replaced by cysteines, (ii) one of amino acid positions 73 to 75, preferably amino acid position 74 and one of amino acid positions 229 to 231 , preferably amino acid position 230 of SEQ ID NO: 7 are replaced by cysteines, (iii) one of amino acid positions 56 to 58, preferably amino acid position 57 and one of amino acid positions 172 to 174, preferably amino acid position 173 of SEQ ID NO: 7 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity
- (VIII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG5 (Adhesion G Protein-Coupled Receptor G5) of SEQ ID NO: 8, wherein(i) one of amino acid positions 45 to 47, preferably amino acid position 46 and one of amino acid positions 193 to 195, preferably amino acid position 194 of SEQ ID NO: 8 are replaced by cysteines, (ii) one of amino acid positions 30 to 32, preferably amino acid position 31 and one of amino acid positions 137 to 139, preferably amino acid position 138 of SEQ ID NO: 8 are replaced by cysteines, or(iii) (i) and (ii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii) or (iii) are retained;
- (IX) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG4 (Adhesion G Protein-Coupled Receptor G4) of SEQ ID NO: 9, wherein (i) one of amino acid positions 251 to 253, preferably amino acid position 252 and one of amino acid positions 302 to 304, preferably amino acid position 303 of SEQ ID NO: 9 are replaced by cysteines, (ii) one of amino acid positions 290 to 292, preferably amino acid position 291 and one of amino acid positions 442 to 444, preferably amino acid position 443 of SEQ ID NO: 9 are replaced by cysteines, (iii) one of amino acid positions 274 to 276, preferably amino acid position 275 and one of amino acid positions 385 to 387, preferably amino acid position 386 of SEQ ID NO: 9 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80%
- (X) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG3 (Adhesion G Protein-Coupled Receptor G3) of SEQ ID NO: 10, wherein(i) one of amino acid positions 16 to 19, preferably amino acid position 17 and one of amino acid positions 45 to 47, preferably amino acid position 46 of SEQ ID NO: 10 are replaced by cysteines, (ii) one of amino acid positions 32 to 34, preferably amino acid position 33 and one of amino acid positions 116 to 118, preferably amino acid position 117 of SEQ ID NO: 10 are replaced by cysteines, or(iii) (i) and (ii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii) or (iii) are retained;
- (XI) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG2 (Adhesion G Protein-Coupled Receptor G2) of SEQ ID NO: 11 , wherein(i) one of amino acid positions 12 to 14, preferably amino acid position 13 and one of amino acid positions 63 to 65, preferably amino acid position 64 of SEQ ID NO: 11 are replaced by cysteines, (ii) one of amino acid positions 51 to 53, preferably amino acid position 52 and one of amino acid positions 205 to 207, preferably amino acid position 206 of SEQ ID NO: 11 are replaced by cysteines, (iii) one of amino acid positions 35 to 37, preferably amino acid position 36 and one of amino acid positions 147 to 149, preferably amino acid position 148 of SEQ ID NO: 11 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (
- (XII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF5 (Adhesion G Protein-Coupled Receptor F5) of SEQ ID NO: 12, wherein(i) one of amino acid positions 82 to 84, preferably amino acid position 83 and one of amino acid positions 117 to 119, preferably amino acid position 118 of SEQ ID NO: 12 are replaced by cysteines, (ii) one of amino acid positions 105 to 107, preferably amino acid position 106 and one of amino acid positions 240 to 242, preferably amino acid position 241 of SEQ ID NO: 12 are replaced by cysteines, (iii) one of amino acid positions 90 to 92, preferably amino acid position 91 and one of amino acid positions 186 to 188, preferably amino acid position 187 of SEQ ID NO: 12 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least
- (XIII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF4 (Adhesion G Protein-Coupled Receptor F4) of SEQ ID NO: 13, wherein(i) one of amino acid positions 53 to 55, preferably amino acid position 54 and one of amino acid positions 86 to 88, preferably amino acid position 87 of SEQ ID NO: 13 are replaced by cysteines, (ii) one of amino acid positions 74 to 76, preferably amino acid position 75 and one of amino acid positions 216 to 218, preferably amino acid position 217 of SEQ ID NO: 13 are replaced by cysteines, (iii) one of amino acid positions 61 to 63, preferably amino acid position 62 and one of amino acid positions 159 to 161 , preferably amino acid position 160 of SEQ ID NO: 13 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity
- (XIV) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF3 (Adhesion G Protein-Coupled Receptor F3) of SEQ ID NO: 14, wherein(i) one of amino acid positions 81 to 83, preferably amino acid position 82 and one of amino acid positions 116 to 118, preferably amino acid position 117 of SEQ ID NO: 14 are replaced by cysteines, (ii) one of amino acid positions 104 to 106, preferably amino acid position 105 and one of amino acid positions 243 to 245, preferably amino acid position 244 of SEQ ID NO: 14 are replaced by cysteines, (iii) one of amino acid positions 89 to 91 , preferably amino acid position 90 and one of amino acid positions 184 to 186, preferably amino acid position 185 of SEQ ID NO: 14 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least
- (XV) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF2 (Adhesion G Protein-Coupled Receptor F2) of SEQ ID NO: 15, wherein(i) one of amino acid positions 73 to 75, preferably amino acid position 74 and one of amino acid positions 106 to 108, preferably amino acid position 107 of SEQ ID NO: 15 are replaced by cysteines, (ii) one of amino acid positions 94 to 96, preferably amino acid position 95 and one of amino acid positions 230 to 232, preferably amino acid position 231 of SEQ ID NO: 15 are replaced by cysteines, (iii) one of amino acid positions 81 to 83, preferably amino acid position 82 and one of amino acid positions 177 to 179, preferably amino acid position 178 of SEQ ID NO: 15 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least
- (XVI) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF1 (Adhesion G Protein-Coupled Receptor F1) of SEQ ID NO: 16, wherein(i) one of amino acid positions 83 to 85, preferably amino acid position 84 and one of amino acid positions 118 to 120, preferably amino acid position 119 of SEQ ID NO: 16 are replaced by cysteines, (ii) one of amino acid positions 106 to 108, preferably amino acid position 107 and one of amino acid positions 243 to 245, preferably amino acid position 244 of SEQ ID NO: 16 are replaced by cysteines, (iii) one of amino acid positions 91 to 93, preferably amino acid position 92 and one of amino acid positions 190 to 192, preferably amino acid position 191 of SEQ ID NO: 16 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80%
- (XVII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE5 (Adhesion G Protein-Coupled Receptor E5) of SEQ ID NO: 17, wherein one of amino acid positions 72 to 74, preferably amino acid position 73 and one of amino acid positions 244 to 246, preferably amino acid position 245 of SEQ ID NO: 17 are replaced by cysteines; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A) are retained;
- (XVIII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE3 (Adhesion G Protein-Coupled Receptor E3) of SEQ ID NO: 19, wherein one of amino acid positions 73 to 75, preferably amino acid position 74 and one of amino acid positions 186 to 188, preferably amino acid position 187 of SEQ ID NO: 19 are replaced by cysteines; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A) are retained;
- (XIX) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE2 (Adhesion G Protein-Coupled Receptor E2) of SEQ ID NO: 20, wherein one of amino acid positions 72 to 74, preferably amino acid position 73 and one of amino acid positions 228 to 230, preferably amino acid position 229 of SEQ ID NO: 20 are replaced by cysteines; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A) are retained;
- (XX) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE1 (Adhesion G Protein-Coupled Receptor E1) of SEQ ID NO: 21 , wherein (i) one of amino acid positions
- XXI (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRD2 (Adhesion G Protein-Coupled Receptor D2) of SEQ ID NO: 22, wherein(i) one of amino acid positions
- cysteines preferably amino acid positions 98 to 100, preferably amino acid position 99 and one of amino acid positions 235 to 237, preferably amino acid position 236 of SEQ ID NO: 22 are replaced by cysteines, or(iii) (i) and (ii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii) or (iii) are retained;
- (XXII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRD1 (Adhesion G Protein-Coupled Receptor D1) of SEQ ID NO: 23, wherein(i) one of amino acid positions 36 to 38, preferably amino acid position 37 and one of amino acid positions 92 to 94, preferably amino acid position 93 of SEQ ID NO: 23 are replaced by cysteines, (ii) one of amino acid positions 80 to 82, preferably amino acid position 81 and one of amino acid positions 247 to 249, preferably amino acid position 248 of SEQ ID NO: 23 are replaced by cysteines, (iii) one of amino acid positions 66 to 68, preferably amino acid position 67 and one of amino acid positions 162 to 164, preferably amino acid position 163 of SEQ ID NO: 23 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least
- (XXIII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRB3 (Adhesion G Protein-Coupled Receptor B3) of SEQ ID NO: 24, wherein(i) one of amino acid positions 87 to 89, preferably amino acid position 88 and one of amino acid positions 121 to 123, preferably amino acid position 122 of SEQ ID NO: 24 are replaced by cysteines, (ii) one of amino acid positions 109 to 111 , preferably amino acid position 110 and one of amino acid positions 261 to 263, preferably amino acid position 262 of SEQ ID NO: 24 are replaced by cysteines, (iii) one of amino acid positions 95 to 97, preferably amino acid position 96 and one of amino acid positions 206 to 208, preferably amino acid position 207 of SEQ ID NO: 24 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least
- (XXIV) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRB2 (Adhesion G Protein-Coupled Receptor B2) of SEQ ID NO: 25, wherein(i) one of amino acid positions 86 to 88, preferably amino acid position 87 and one of amino acid positions 120 to 122, preferably amino acid position 121 of SEQ ID NO: 25 are replaced by cysteines, (ii) one of amino acid positions 108 to 110, preferably amino acid position 109 and one of amino acid positions 295 to 297, preferably amino acid position 296 of SEQ ID NO: 25 are replaced by cysteines, (iii) one of amino acid positions 94 to 96, preferably amino acid position 95 and one of amino acid positions 239 to 241 , preferably amino acid position 240 of SEQ ID NO: 25 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least
- (XXV) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRB1 (Adhesion G Protein-Coupled Receptor B1) of SEQ ID NO: 26, wherein(i) one of amino acid positions 81 to 83, preferably amino acid position 82 and one of amino acid positions 115 to 117, preferably amino acid position 116 of SEQ ID NO: 26 are replaced by cysteines, (ii) one of amino acid positions 103 to 105, preferably amino acid position 104 and one of amino acid positions 253 to 255, preferably amino acid position 254 of SEQ ID NO: 26 are replaced by cysteines, (iii) one of amino acid positions 89 to 91 , preferably amino acid position 90 and one of amino acid positions 198 to 200, preferably amino acid position 199 of SEQ ID NO: 26 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at
- (XXVI) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRA3 (Adhesion G Protein-Coupled Receptor A3) of SEQ ID NO: 27, wherein(i) one of amino acid positions 57 to 59, preferably amino acid position 58 and one of amino acid positions 106 to 108, preferably amino acid position 107 of SEQ ID NO: 27 are replaced by cysteines, (ii) one of amino acid positions 94 to 96, preferably amino acid position 95 and one of amino acid positions 288 to 290, preferably amino acid position 289 of SEQ ID NO: 27 are replaced by cysteines, (iii) one of amino acid positions 81 to 83, preferably amino acid position 82 and one of amino acid positions 225 to 227, preferably amino acid position 226 of SEQ ID NO: 27 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at
- (XXVII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRA2 (Adhesion G Protein-Coupled Receptor A2) of SEQ ID NO: 28, wherein (i) one of amino acid positions 57 to 59, preferably amino acid position 58 and one of amino acid positions 104 to 106, preferably amino acid position 105 of SEQ ID NO: 28 are replaced by cysteines, (ii) one of amino acid positions 92 to 94, preferably amino acid position 93 and one of amino acid positions 294 to 296, preferably amino acid position 295 of SEQ ID NO: 28 are replaced by cysteines, (iii) one of amino acid positions 79 to 81 , preferably amino acid position 80 and one of amino acid positions 229 to 231 , preferably amino acid position 230 of SEQ ID NO: 28 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence
- (XXVIII) (A) comprising or consisting of the amino acid sequence of GAIN domain of ADGRV1 (Adhesion G Protein-Coupled Receptor V1) of SEQ ID NO: 29, wherein(i) one of amino acid positions 171 to 173, preferably amino acid position 172 and one of amino acid positions 229 to 231 , preferably amino acid position 230 of SEQ ID NO: 29 are replaced by cysteines, (ii) one of amino acid positions 127 to 219, preferably amino acid position 218 and one of amino acid positions 436 to 438, preferably amino acid position 437 of SEQ ID NO: 29 are replaced by cysteines, (iii) one of amino acid positions 201 to 203, preferably amino acid position 202 and one of amino acid positions 377 to 379, preferably amino acid position 378 of SEQ ID NO: 29 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing
- the GAIN domain (G-protein-coupled receptor autoproteolysis-inducing domain) is a protein domain found in adhesion GPCRs. The domain is involved in the self-cleavage of these transmembrane receptors, and has been shown to be crucial for their function (Liebscher et al. (2015), Oncotarget, Vol. 6, No. 27).
- Adhesion G protein-coupled receptors are a class of currently 33 human protein receptors with a broad distribution in embryonic and larval cells, cells of the reproductive tract, neurons, leukocytes, and a variety of tumours.
- DGRA1 GPR123
- ADGRA2 GPR124
- ADGRA3 GPR125
- ADGRB1 BAI1
- ADGRB2 BAI2
- ADGRB3 BAI3
- ADGRC1 CELSR1
- ADGRC2 CELSR2
- ADGRC3 CELSR3
- ADGRD1 GPR133
- ADGRD2 GPR144
- ADGRE1 EMR1 , F4/80
- ADGRE2 EMR2
- ADGRE3 EMR3
- ADGRE4 EMR4
- ADGRE5 CD97
- ADGRF1 GPR110
- ADGRF2 GPR111
- ADGRF3 GPR113
- ADGRF4 GPR115
- ADGRF5 GPR116, Ig-Hepta
- ADGRG1 GPR56
- ADGRG2 GPR64, HE6
- ADGRG3 GPR97
- ADGRG4 GPR112
- Each subfamily is assigned to a letter, such as L for the latrophilins, E for the EGF-TM7 receptors, C for the CELSRs, B for the BAIs, and V for VLGR, while the subfamilies with GPR# names have been given a letter in alphabetic order (A, D, F, G).
- Adhesion GPCRs are distinguished from other GPCR families in that they mostly possess large N and C termini and a GAIN domain.
- the GAIN domain is evolutionarily well conserved among AGPCRs and seems to be present from tetrahymena to mammals.
- Adhesion GPCR molecules can be divided into an N-terminal fragment (NTF) and a C-terminal fragment (CTF).
- the cleavage site also marks the beginning of the tethered agonist sequence, coined the Stachel sequence.
- the Stachel sequence is a tethered peptide agonist (Liebscher et al. (2014), cell reports, 9(6):2018-2026) being located within the ectodomain of AGPCRs.
- the N-terminal fragment (NTF) contains adhesion domains and most of the GAIN domain and the C-terminal fragment (CTF) contains the Stachel sequence, that emanates extracellularly from the seven-transmembrane-spanning (7TM) domain.
- the NTF and CTF remain non-covalently associated on the plasma membrane via the association of the GAIN domain and the Stachel sequence.
- SMFS single-molecule force spectroscopy
- the variants of the GAIN domain according to the invention are required to comprise at least two cysteines at specific sites and/or are required to comprise specific amino acid replacements at specific sites.
- the at least two cysteines can form a disulfide bond that cannot be found in the corresponding wild-type (WT) GAIN domains of SEQ ID NOs 1 to 29.
- the specific amino acid replacements introduce amino acid substitutions at specific sites that result in amino acids at these specific sites that cannot be found in the corresponding wild-type (WT) GAIN domains of SEQ ID NOs 1 to 29.
- the specific sites, wherein the two cysteines are to be added and/or the amino acids are to be substituted were determined for ADGRG1 GAIN domain.
- the sites were selected such that the mechanical force that is required to dissociate ADGRG1 GAIN domain from is Stachel sequence is either higher or lower (i.e. the variants have an increased or a decreased resistance to mechanical force).
- such variants of the GAIN domain are advantageous for applications such as cell signaling studies and drug development.
- the specific sites of the ADGRG1 GAIN domain are amino acid position 4 and position 48 of SEQ ID NO: 1.
- This variant GAIN domain is also designated h1 h2 herein and displays a lower force threshold of GAIN-Stachel dissociation (about 67- 135 pN). It is expected that the essentially same results can be achieved within positions 3 to 5 and positions 47 to 49 because the two-dimensional lock by the Cys-Cys bridge is essentially the same as for positions 4 and 38.
- the specific sites of the ADGRG1 GAIN domain are amino acid position 36 and position 181 of SEQ ID NO: 1.
- This variant GAIN domain is also designated h2l10 herein and displays a higher force threshold of GAIN-Stachel dissociation (about 182-236 pN). It is expected that the essentially same results can be achieved within positions 35 to 37 and positions 180 to 182 because the two-dimensional lock by the Cys-Cys bridge is essentially the same as for positions 36 and 181 .
- the specific sites of the ADGRG1 GAIN domain are amino acid position 20 and position 125 of SEQ ID NO: 1.
- This variant GAIN domain is also designated hll7 herein and displays a lower force threshold of GAIN-Stachel dissociation (about 66-117 pN).
- the H1 L7 force propagation is very similar to that of the H1 H2 variant.
- the H1 L7 variant behaves similarly to the H1 H2 variant as was measured by atomic force microscopy (AFM). It is expected that the essentially same results can be achieved within positions 19 to 21 and positions 124 to 126 because the two-dimensional lock by the Cys-Cys bridge is essentially the same as for positions 20 and 125.
- the specific sites of the ADGRG1 GAIN domain are amino acid position 57 and position 66 of SEQ ID NO: 1.
- This variant GAIN domain is also designated blb2 herein and displays a lower force threshold of GAIN-Stachel dissociation.
- This variant GAIN domain is also designated b1 b2 herein and displays a lower force threshold of GAIN- Stachel dissociation (about 39-58 pN). It is expected that the essentially same results can be achieved within positions 56 to 58 and positions 65 to 67 because the two-dimensional lock by the Cys-Cys bridge is essentially the same as for positions 57 and 66.
- the specific sites of the ADGRG1 GAIN domain are amino acid position 17 and position 61 of SEQ ID NO: 1.
- This variant GAIN domain is also designated h1 b1 herein and displays a lower force threshold of GAIN-Stachel dissociation (about 78- 117 pN). It is expected that the essentially same results can be achieved within positions 16 to 18 and positions 60 to 62 because the two-dimensional lock by the Cys-Cys bridge is essentially the same as for positions 17 and 61 .
- the specific amino acid substitutions in the ADGRG1 GAIN domain are the three Q40E/E43L/E181 R.
- This ADGRG1 GAIN domain variant is also designated 3Mut herein.
- the specific amino acid substitutions in the ADGRG1 GAIN domain are the four Q40E/E43L/S135R/E181 R.
- This ADGRG1 GAIN domain variant is also designated 4Mut herein.
- the specific amino acid substitutions in the ADGRG1 GAIN domain are the six Q40E/E43L/S135R/V146/F178Y/E181 R.
- This ADGRG1 GAIN domain variant is also designated 6Mut herein.
- the specific amino acid substitutions in the ADGRG1 GAIN domain are the ten A33W/Q40E/E43L/S135R/V146Y/F178Y/E181 R/T184G/S186Q/S187G.
- This ADGRG1 GAIN domain variant is also designated 10Mut herein.
- the specific amino acid substitution in the ADGRG1 GAIN domain is L174I.
- This ADGRG1 GAIN domain variant is also designated L344I herein (noting that amino acid position 344 of SEQ ID NO: 30 (full ADGRG1) corresponds to amino acid position 174 of SEQ ID NO: 1 (ADGRG1 GAIN domain)).
- the specific amino acid substitution in the ADGRG1 GAIN domain is L218A.
- This ADGRG1 GAIN domain variant is also designated L388A herein (noting that amino acid position 388 of SEQ ID NO: 30 (full ADGRG1) corresponds to amino acid position 218 of SEQ ID NO: 1 (ADGRG1 GAIN domain)).
- the GAIN domain of items (l)(A)(vi) to (xi) all display a lower force threshold of GAIN-Stachel dissociation; see Figure 18.
- ADGRG1 GAIN domain variant with eight of the ten A203W/Q210E/E213L/S305R/V316Y/F348Y/E351 R/T354G/S356Q/S357H substitutions was prepared.
- This 8Mut variant does not display a lower or higher force threshold of GAIN-Stachel dissociation but essentially behaves as the wild-type ADGRG1 GAIN domain. Accordingly, this 8Mut variant is not part of the first aspect.
- the mutational effects in the non-covalent variants are more subtle, which allows for a finer grained modulation of the desired mechanical force.
- the effects on the signaling responses to ligands can be different than pure mechanical resistance because the non-covalent mutants show lower signaling responses to TG2 (transglutaminase 2) binding despite being generally less mechanically resistant.
- the non-covalent design advantageously provides functional effects that can complement those obtained with the S-S variants.
- At least two of items (l)(A)(i) to (xi) apply.
- the at least two are with increasing preference at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, and at least ten.
- ADGRG1 GAIN domain variants that comprise the mutations of alt least one SS variant and at least one non- covalent variant.
- the different pairs of replaced cysteines of (l)(A)(i) to (v) can be combined and/or the specific amino acid substitutions of (l)(A)(vi) to (xi).
- This is illustrated in the examples for the variant h1 h2-h2l10.
- the h1 h2-h2l10 force propagation is very similar to that of the WT. Consistent with these predictions, hi h2-h2l10 displays a force threshold being similar to the WT ADGRG1 GAIN domain.
- the variant GAIN domain comprises or consists of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii), (iv), (v) and/or (vi), or the amino acid replacement(s) as defined in (A), item (vii), (viii), (ix), (x), (xi) an/or (xii), or the cysteines and/or the amino acid replacement(s) as defined in (A), item (xiii) are retained.
- amino acid sequence sharing at least 80% sequence identity displays with increasing preference at least 85% sequence identity, at least 90% sequence identity, at least 95% sequence identity, at least 96% sequence identity, at least 97% sequence identity, at least 98% sequence identity and at least 99% sequence identity.
- the term “percent (%) sequence identity” describes the number of matches (“hits”) of identical nucleotides/amino acids of two or more aligned nucleic acid or amino acid sequences as compared to the number of nucleotides or amino acid residues making up the overall length of the template nucleic acid or amino acid sequences.
- hits the number of matches of identical nucleotides/amino acids of two or more aligned nucleic acid or amino acid sequences as compared to the number of nucleotides or amino acid residues making up the overall length of the template nucleic acid or amino acid sequences.
- using an alignment for two or more sequences or subsequences the percentage of amino acid residues or nucleotides that are the same (e.g.
- Nucleotide and amino acid sequence analysis and alignment in connection with the present invention are preferably carried out using the NCBI BLAST algorithm (Stephen F. Altschul, Thomas L. Madden, Alejandro A. Schaffer, Jinghui Zhang, Zheng Zhang, Webb Miller, and David J. Lipman (1997), Nucleic Acids Res. 25:3389-3402).
- BLAST can be used for nucleotide sequences (nucleotide BLAST) and amino acid sequences (protein BLAST).
- the skilled person is aware of additional suitable programs to align nucleic acid sequences.
- the NCBI BLAST algorithm is available for protein (Protein BLAST) and nucleotides (Nucleotide BLAST).
- the algorithm parameters are preferably: max target sequences: 100, with automatically adjust parameters for short input sequences, expect threshold 0.05, word size 6, Max matches in a query range 0, matrix BLOSUM62, cap cost existence: 10 extension: 1 , and compositional adjustment.
- max target sequences 100, with “automatically adjust parameters for short input sequences”, Expect threshold 0.05, word size 28, Max matches in a query range 0, match/mismatch scores 1 ,-2, cap costs linear, low complexity regions filter, and .mask for look up table only.
- amino acid sequence sharing at least 80% sequence identity display with increasing preference at least 85% sequence identity, at least 90% sequence identity, at least 95% sequence identity, at least 96% sequence identity, at least 97% sequence identity, at least 98% sequence identity and at least 99% sequence identity.
- the present invention relates in a second aspect to an adhesion G protein-coupled receptor (AGPCR) comprising the variant of the GAIN domain the first aspect, wherein the AGPCR preferably comprises of consists of any of SEQ ID NOs 30 to 58 or a sequence being at least 70% identical, preferably at least 80% and most preferably at least 90% identical thereto.
- AGPCR adhesion G protein-coupled receptor
- first aspect as far as being amenable for combination with the second aspect apply mutatis mutandis to the second aspect of the invention.
- the definitions and preferred embodiments of the first and second aspect as far as being amenable for combination with the third aspect apply mutatis mutandis to the third aspect of the invention, etc.
- amino acid sequence sharing at least 70% sequence identity displays with increasing preference at least 75% sequence identity, at least 80% sequence identity, at least 85% sequence identity, at least 90% sequence identity, at least 95% sequence identity, at least 96% sequence identity, at least 97% sequence identity, at least 98% sequence identity and at least 99% sequence identity.
- SEQ ID NOs 30 to 58 are the full-length amino acid sequences of the wild-type AGPCRS as named in items (I) to (XXIX) of the first aspect, in that particular order. It follows that the GAIN domain of SEQ ID NO: 1 is comprised in SEQ ID NO: 30, the GAIN domain of SEQ ID NO: 2 is comprised in SEQ ID NO: 31 , etc. (see table A herein above).
- these wild-type GAIN domains are replaced by a GAIN domain of the first aspect while preferably overall a sequence being of at least 70% identical, preferably of at least 80% and most preferably of at least 90% identical to any of SEQ ID NOs 30 to 58 is retained.
- the wild-type GAIN domains are replaced by a corresponding variant of the GAIN domains.
- the wild-type GAIN domain of SEQ ID NO. 1 is preferably replaced by a variant of the GAIN domain of item (I) of the first aspect
- in SEQ ID NO: 31 the wild-type GAIN domain of SEQ ID NO: 2 is preferably replaced by a variant of the GAIN domain of item (II) of the first aspect, etc.
- a sequence being at least 70% identical, preferably at least 80% and most preferably at least 90% identical to any of SEQ ID NOs 30 to 58 is retained.
- the variant of the GAIN domain of the first aspect and/or the AGPCR of the second aspect is a fusion protein.
- protein as used herein is used interchangeably with the term (poly)peptide and refers to polymers constructed from one or more chains of amino acid residues linked by peptide bonds.
- Poly)peptides describe a group of molecules which comprise the group of peptides consisting of up to 30 amino acids, as well as the group of polypeptides consisting of more than 30 amino acids.
- protein also refers to chemically or post-translationally modified proteins. The terms apply to naturally occurring amino acid polymers, as well as to amino acid polymers in which one or more amino acid residue is an artificial chemical analogue of a corresponding naturally occurring amino acid.
- a “fusion protein” according to the present invention contains at least one additional heterologous amino acid sequence other than the variant of the GAIN domain and/or the AGPCR. Often, but not necessarily, these additional sequences will be located at the N- or C-terminal end of the (poly)peptide. It may e.g. be convenient to initially express the (poly)peptide as a fusion protein from which the additional amino acid residues can be removed, e.g. by a proteinase capable of specifically trimming the fusion protein and releasing the (poly)peptide of the present invention.
- the additional heterologous amino acid sequence can either be directly or indirectly fused to the variant of the GAIN domain and/or the AGPCR of the invention. In case of an indirect fusion generally a peptide linker may be used for the fusion, such as a GS-linker.
- fusion proteins with antibodies are envisioned herein.
- the term “antibody” comprises antibody fragments and derivatives.
- the antibody may be, for example, specific for cell surface markers or may be an antigen-recognizing fragment of said antibodies.
- the poly(peptide) of the invention can be fused to the N-terminus or C-terminus of the light and/or heavy chain(s) of an antibody.
- the poly(peptide) of the invention is preferably fused to the N-terminus of the light and/or heavy chain(s) of an antibody, so that the Fc part of the antibody is free to bind to Fc-receptors.
- antibody as used in accordance with the present invention comprises, for example, polyclonal or monoclonal antibodies. Furthermore, also derivatives or fragments thereof, which still retain the binding specificity to the desired target, e.g. a tumor antigen, are comprised in the term "antibody”.
- Antibody fragments or derivatives comprise, inter alia, Fab or Fab’ fragments, Fd, F(ab')2, Fv or scFv fragments, single domain VH or V-like domains, such as VhH or V-NAR-domains, as well as multimeric formats such as minibodies, diabodies, tribodies or triplebodies, tetrabodies or chemically conjugated Fab’-multimers (see, for example, Harlow and Lane “Antibodies, A Laboratory Manual”, Cold Spring Harbor Laboratory Press, 198; Harlow and Lane “Using Antibodies: A Laboratory Manual” Cold Spring Harbor Laboratory Press, 1999; Altshuler EP, Serebryanaya DV, Katrukha AG. 2010, Biochemistry (Mose)., vol.
- the multimeric formats in particular comprise bispecific antibodies that can simultaneously bind to two different types of antigen.
- the first antigen can be found on the poly(peptide) of the invention.
- the second antigen may, for example, be a tumor marker that is specifically expressed on cancer cells or a certain type of cancer cells.
- bispecific antibodies formats are Biclonics (bispecific, full length human IgG antibodies), DART (Dual- affinity Re-targeting Antibody) and BiTE (consisting of two single-chain variable fragments (scFvs) of different antibodies) molecules (Kontermann and Brinkmann (2015), Drug Discovery Today, 20(7):838-847).
- antibody also includes embodiments such as chimeric (human constant domain, nonhuman variable domain), single chain and humanised (human antibody with the exception of nonhuman CDRs) antibodies.
- the fusion protein may also comprise protein domains known to function in signal transduction and/or known to be involved in protein-protein interaction.
- Examples for such domains are Ankyrin repeats; arm, Bcl-homology, Bromo, CARD, CH, Chr, C1 , C2, DD, DED, DH, EFh, ENTH, F-box, FHA, FYVE, GEL, GYF, hect, LIM, MH2, PDZ, PB1 , PH, PTB, PX, RGS, RING, SAM, SC, SH2, SH3, SOCS, START, TIR, TPR, TRAF, tsnare, Tubby, UBA, VHS, W, WW, and 14-3-3 domains.
- the at least one additional heterologous amino acid sequence of the fusion protein according to the present invention may also comprise or consist of (a) a cytokine, (b) a chemokine, (c) a pro-coagulant factor, (d) a proteinaceous toxic compound, and/or (e) an enzyme for pro-drug activation.
- the cytokine is preferably selected from the group consisting of IL-2, IL-12, TNF-alpha, IFN alpha, IFN beta, IFN gamma, IL-10, IL-15, IL-24, GM-CSF, IL-3, IL-4, IL-5, IL-6, IL-7, IL-9, IL-11 , IL-13, LIF, CD80, B70, TNF beta, LT-beta, CD-40 ligand, Fas-ligand, TGF-beta, IL-1 alpha and IL-1 beta.
- cytokines may favour a pro-inflammatory or an anti-inflammatory response of the immune system.
- fusion proteins with a pro- inflammatory or an anti-inflammatory cytokine may be favored.
- anti-inflammatory cytokines are preferred
- pro-inflammatory cytokines are preferred.
- the chemokine is preferably selected from the group consisting of IL-8, GRO alpha, GRO beta, GRO gamma, ENA-78, LDGF-PBP, GCP-2, PF4, Mig, IP-10, SDF-1alpha/beta, BUNZO/STRC33, l-TAC, BLC/BCA-1 , MIP-1 alpha, MIP-1 beta, MDC, TECK, TARC, RANTES, HCC-1 , HCC-4, DC-CK1 , MIP-3 alpha, MIP-3 beta, MCP-1-5, eotaxin, Eotaxin-2, I-309, MPIF-1 , 6Ckine, CTACK, MEC, lymphotactin and fractalkine.
- chemokines act as a chemoattractant to guide the migration of cells.
- Cells that are attracted by chemokines follow a signal of increasing chemokine concentration towards the source of the chemokine. It follows that within the fusion protein the chemokine can be used to guide the migration of the protein or peptide of the invention, e.g. to a specific cell type or body site.
- the pro-coagulant factor is preferably a tissue factor.
- a pro-coagulant factor promotes the process by which blood changes from a liquid to a gel, forming a blood clot.
- Pro-coagulant factors may, for example, aid in wound healing.
- the proteinaceous toxic compound is preferably Ricin-A chain, modeccin, truncated Pseudomonas exotoxin A, diphtheria toxin or recombinant gelonin.
- Toxic compounds can have a toxic effect on a whole organism as well as on a substructure of the organism, such as a particular cell type. Toxic compounds are frequently used in the treatment of tumors. Tumor cells generally grow faster than normal body cells, so that they preferentially accumulate toxic compounds and in higher amounts.
- the enzyme for pro-drug activation is preferably an enzyme selected from the group consisting of carboxy-peptidases, glucuronidases and glucosidases.
- an enzyme selected from the group consisting of carboxy-peptidases, glucuronidases and glucosidases are especially appealing as they directly complement ongoing clinical chemotherapeutic regimes. These enzymes can activate prodrugs that have low inherent toxicity using both bacterial and yeast enzymes or enhance prodrug activation by mammalian enzymes.
- the fusion protein may also comprise a tag, such as purification tag.
- a tag such as purification tag.
- purification tags are available and an overview of affinity tags for protein purification is available in Kimple et al. (2013), Curr Protoc Protein Sci. 2013; 73: Unit-9.9. In the examples a HA tag (3xHA) is illustrated.
- the variant of the GAIN domain and/or the AGPCR is fused to a heterologous non-proteinaceous compound.
- a heterologous compound is a compound that cannot be found in nature fused to the variant of the GAIN domain and/or the AGPCR.
- the heterologous non-proteinaceous compound can either be directly or indirectly fused to the variant of the GAIN domain and/or the AGPCR.
- a chemical linker may be used.
- Chemical linkers may contain diverse functional groups, such as primary amines, sulfhydryls, acids, alcohols and bromides. Many crosslinkers are functionalized with maleimide (sulfhydryl reactive) and succinimidyl ester (NHS) or isothiocyanate (ITC) groups that react with amines.
- the heterologous non-proteinaceous compound is preferably a pharmaceutically active compound or diagnostically active compound.
- the pharmaceutically active compound or diagnostically active compound is preferably selected from the group consisting of (a) a fluorescent dye, (b) a photosensitizer, (c) a radionuclide, (d) a contrast agent for medical imaging, (e) a toxic compound, or (f) an ACE inhibitor, a Renin inhibitor, an ADH inhibitor, an Aldosteron inhibitor, an Angiotensin receptor blocker, a TSH-receptor, a LH-/HCG-receptor, an oestrogen receptor, a progesterone receptor, an androgen receptor, a GnRH-receptor, a GH (growth hormone) receptor, or a receptor for IGF-I or IGF-II.
- the fluorescent dye is preferably a component selected from Alexa Fluor or Cy dyes.
- the photosensitizer is preferably phototoxic red fluorescent protein KillerRed or haematoporphyrin.
- the radionuclide is preferably either selected from the group of gamma-emitting isotopes, more preferably 99m Tc, 123 l, 111 ln, and/or from the group of positron emitters, more preferably 18 F, 64 Cu, 68 Ga, 86 Y, 124 l, and/or from the group of beta-emitters, more preferably 131 l, 90 Y, 177 Lu, 67 Cu, 90 Sr, and/or from the group of alpha-emitters, preferably 213 Bi, 211 At.
- a contrast agent as used herein is a substance used to enhance the contrast of structures or fluids within the body in medical imaging. Common contrast agents work based on X-ray attenuation and magnetic resonance signal enhancement.
- the toxic compound is preferably a small organic compound, more preferably a toxic compound selected from the group consisting of calicheamicin, maytansinoid, neocarzinostatin, esperamicin, dynemicin, kedarcidin, maduropeptin, doxorubicin, daunorubicin, and auristatin.
- a toxic compound selected from the group consisting of calicheamicin, maytansinoid, neocarzinostatin, esperamicin, dynemicin, kedarcidin, maduropeptin, doxorubicin, daunorubicin, and auristatin.
- these toxic compounds are non-proteinaceous.
- the present invention relates in a third aspect to a nucleic acid molecule, preferably a vector encoding the variant of the GAIN domain of the first aspect or the AGPCR of the second aspect.
- nucleic acid molecule in accordance with the present invention includes DNA, such as cDNA or double or single stranded genomic DNA and RNA.
- DNA deoxyribonucleic acid
- DNA means any chain or sequence of the chemical building blocks adenine (A), guanine (G), cytosine (C) and thymine (T), called nucleotide bases, that are linked together on a deoxyribose sugar backbone.
- DNA can have one strand of nucleotide bases, or two complimentary strands which may form a double helix structure.
- RNA ribonucleic acid
- A adenine
- G guanine
- C cytosine
- U uracil
- RNA typically has one strand of nucleotide bases, such as mRNA. Included are also single- and double-stranded hybrids molecules, i.e., DNA-DNA, DNA- RNA and RNA-RNA.
- the nucleic acid molecule may also be modified by many means known in the art.
- Non-limiting examples of such modifications include methylation, "caps”, substitution of one or more of the naturally occurring nucleotides with an analog, and internucleotide modifications such as, for example, those with uncharged linkages (e.g., methyl phosphonates, phosphotriesters, phosphoroamidates, carbamates, etc.) and with charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.).
- uncharged linkages e.g., methyl phosphonates, phosphotriesters, phosphoroamidates, carbamates, etc.
- charged linkages e.g., phosphorothioates, phosphorodithioates, etc.
- Nucleic acid molecules in the following also referred as polynucleotides, may contain one or more additional covalently linked moieties, such as, for example, proteins (e.g., nucleases, toxins, antibodies, signal peptides, poly-L-lysine, etc.), intercalators (e.g., acridine, psoralen, etc.), chelators (e.g., metals, radioactive metals, iron, oxidative metals, etc.), and alkylators.
- proteins e.g., nucleases, toxins, antibodies, signal peptides, poly-L-lysine, etc.
- intercalators e.g., acridine, psoralen, etc.
- chelators e.g., metals, radioactive metals, iron, oxidative metals, etc.
- alkylators e.g., metals, radioactive metals, iron, oxidative metals, etc.
- nucleic acid mimicking molecules known in the art such as synthetic or semi-synthetic derivatives of DNA or RNA and mixed polymers.
- nucleic acid mimicking molecules or nucleic acid derivatives according to the invention include phosphorothioate nucleic acid, phosphoramidate nucleic acid, 2’-0-methoxyethyl ribonucleic acid, morpholino nucleic acid, hexitol nucleic acid (HNA), peptide nucleic acid (PNA) and locked nucleic acid (LNA) (see Braasch and Corey, Chem Biol 2001 , 8: 1).
- LNA is an RNA derivative in which the ribose ring is constrained by a methylene linkage between the 2’-oxygen and the 4’-carbon.
- nucleic acids containing modified bases for example thio-uracil, thio-guanine and fluoro-uracil.
- a nucleic acid molecule typically carries genetic information, including the information used by cellular machinery to make proteins and/or polypeptides.
- the nucleic acid molecule of the invention may additionally comprise promoters, enhancers, response elements, signal sequences, polyadenylation sequences, introns, 5'- and 3'- non-coding regions, and the like.
- vector in accordance with the invention means preferably a plasmid, cosmid, virus, bacteriophage or another vector used e.g. conventionally in genetic engineering which carries the nucleic acid molecule of the invention.
- the nucleic acid molecule of the invention may, for example, be inserted into several commercially available vectors.
- Non-limiting examples include prokaryotic plasmid vectors, such as of the pUC-series, pBluescript (Stratagene), the pET-series of expression vectors (Novagen) or pCRTOPO (Invitrogen) and vectors compatible with an expression in mammalian cells like pREP (Invitrogen), pcDNA3 (Invitrogen), pCEP4 (Invitrogen), pMCI neo (Stratagene), pXT1 (Stratagene), pSG5 (Stratagene), EBO-pSV2neo, pBPV-1 , pdBPVMMTneo, pRSVgpt, pRSVneo, pSV2-dhfr, plZD35, pLXlN, pSIR (Clontech), pIRES-EGFP (Clontech), pEAK-10 (Edge Biosystems) pTriEx-Hygro
- the nucleic acid molecules inserted into the vector can e.g. be synthesized by standard methods, or isolated from natural sources. Ligation of the coding sequences to transcriptional regulatory elements and/or to other amino acid encoding sequences can also be carried out using established methods.
- Transcriptional regulatory elements parts of an expression cassette
- These elements comprise regulatory sequences ensuring the initiation of transcription (e. g., translation initiation codon, promoters, such as naturally-associated or heterologous promoters and/or insulators; see above), internal ribosomal entry sites (IRES) (Owens, Proc. Natl. Acad. Sci.
- polypeptide/protein or fusion protein of the invention is operatively linked to such expression control sequences allowing expression in prokaryotes or eukaryotic cells.
- the vector may further comprise nucleic acid sequences encoding secretion signals as further regulatory elements. Such sequences are well known to the person skilled in the art.
- leader sequences capable of directing the expressed polypeptide to a cellular compartment may be added to the coding sequence of the polynucleotide of the invention. Such leader sequences are well known in the art.
- the vector comprises a selectable marker.
- selectable markers include genes encoding resistance to neomycin, ampicillin, hygromycine, and kanamycin.
- Specifically-designed vectors allow the shuttling of DNA between different hosts, such as bacteria- fungal cells or bacteria-animal cells (e. g. the Gateway system available at Invitrogen).
- An expression vector according to this invention is capable of directing the replication, and the expression, of the polynucleotide and encoded peptide or fusion protein of this invention.
- vectors such as phage vectors or viral vectors (e.g.
- nucleic acid molecules as described herein above may be designed for direct introduction or for introduction via liposomes into a cell.
- baculoviral systems or systems based on vaccinia virus or Semliki Forest virus can be used as eukaryotic expression systems for the nucleic acid molecules of the invention.
- the vector is preferably a retroviral vector.
- a retroviral vector consists of proviral sequences that can accommodate the gene of interest, to allow incorporation of both into the target cells.
- the present invention relates in a fourth aspect to a cell, preferably an epithelial cell or a lymphocyte and most preferably a T-cell comprising the nucleic acid molecule, preferably the vector of the third aspect.
- the cell is preferably an in vitro cell and not an in vivo cell.
- the cell may also be referred to as a “host cell”.
- host cell means any cell of any organism that is selected, modified, transformed, grown, or used or manipulated in any way, for the production of the GAIN domain and/or the AGPCR of the invention by the cell.
- the host cell of the invention is typically produced by introducing the nucleic acid molecule or vector(s) of the invention into the host cell which upon its/their presence mediates the expression of the nucleic acid molecule of the invention encoding the GAIN domain and/or the AGPCR and/or the fusion proteins of the invention.
- the host from which the host cell is derived or isolated may be any prokaryote or eukaryotic cell or organism, preferably with the exception of human embryonic stem cells that have been derived directly by destruction of a human embryo or are derived from such stem cells.
- Suitable prokaryotes (bacteria) useful as hosts for the invention are, for example, those generally used for cloning and/or expression like E.
- E coli e.g., E coli strains BL21 , HB101 , DH5a, XL1 Blue, Y1090 and JM101
- Salmonella typhimurium Salmonella typhimurium, Serratia marcescens, Burkholderia glumae, Pseudomonas putida, Pseudomonas fluorescens, Pseudomonas stutzeri, Streptomyces lividans, Lactococcus lactis, Mycobacterium smegmatis, Streptomyces coelicolor or Bacillus subtilis.
- Appropriate culture mediums and conditions for the above-described host cells are well known in the art.
- a suitable eukaryotic host cell may be a vertebrate cell, an insect cell, a fungal/yeast cell, a nematode cell or a plant cell.
- the fungal/yeast cell may a Saccharomyces cerevisiae cell, Pichia pastoris cell or an Aspergillus cell.
- Preferred examples for host cell to be genetically engineered with the nucleic acid molecule or the vector(s) of the invention is a cell of yeast, E. coli and/or a species of the genus Bacillus (e.g., B. subtilis).
- the host cell is a yeast cell (e.g. S. cerevisiae).
- the host cell is a mammalian host cell, such as a Chinese Hamster Ovary (CHO) cell, mouse myeloma lymphoblastoid, human embryonic kidney cell (HEK-293), human embryonic retinal cell (Crucell's Per.C6), or human amniocyte cell (Glycotope and CEVEC).
- CHO Chinese Hamster Ovary
- HEK-293 human embryonic kidney cell
- Crucell's Per.C6 human embryonic retinal cell
- human amniocyte cell Glycotope and CEVEC
- the cells are frequently used in the art to produce recombinant proteins.
- CHO cells are the most commonly used mammalian host cells for industrial production of recombinant protein therapeutics for humans.
- the following cell types are also suitable as host cells:
- Epithelial cells make up primary tissues throughout the body. There are many arrangements of epithelial cells such as squamous, cuboidal, and columnar that organize as simple, stratified, pseudostratified, and transitional. Epithelial cells form from ectoderm, mesoderm, and endoderm, which explains why epithelia line body cavities and cover most of the body and organ surfaces. Since epithelial cells are prevalent throughout the body, their function changes based on their location. For example, epithelial cells in the skin provide protection, whereas, in the gut, they have secretory and absorptive properties. Epithelial cells have specialized cytoskeletons comprised of microtubules, actin filaments, and intermediate filaments. Intermediate filaments provide structural resilience to the cytoskeleton and show tissue-specific expression, unlike microtubules and actin filaments. These intermediate filaments are a useful target for histochemical staining techniques.
- a lymphocyte is a type of white blood cell in the immune system of jawed vertebrates. Lymphocytes include, innate lymphoid cells (ILCs, i.e. innate counterparts of T cells that contribute to immune responses by secreting effector cytokines and regulating the functions of other innate and adaptive immune cells), natural killer cells (which function in cell-mediated, cytotoxic innate immunity), T cells (for cell-mediated, cytotoxic adaptive immunity), and B cells (for humoral, antibody-driven adaptive immunity).
- the lymphocyte is preferably an anti-tumor lymphocyte.
- An anti-tumor lymphocyte is a lymphocyte capable of eliciting a cytolytic response that can cause tumor cell death.
- the lymphocytes are specializing in and equipped for tumor cell elimination.
- the first category encompasses clonally expanded T lymphocytes expressing a unique T cell receptor (TCR) and recognizing tumor epitopes in the context of the major histocompatibility complex (MHC) molecules.
- TCR T cell receptor
- MHC major histocompatibility complex
- T cells optionally together with B cells producing tumor-specific antibodies and dendritic cells (DC) processing and presenting tumor epitopes, can mediate an adaptive immunity against tumors.
- the second category of effector cells includes natural killer (NK) cells, NK-T cells, and macrophages (M). These cells are not restricted by the MHC molecules in their interactions with tumor targets, and they mediate innate immunity.
- Each type of effector cells contains subsets of cells at different stages of differentiation and activation. This means that each type of effector potentially able to target tumor cells contains a heterogeneous mix of cells with distinct functional capabilities, depending on their stage of differentiation, maturation, and/or activation (Holland, Frei; Cancer Medicine; 6th edition, chapter “Antitumor Effector Cells in Humans”). All these types of anti-tumor lymphocytes are envisioned in accordance with the present invention.
- the lymphocytes are preferably T-cells or NK cells, whereby T-cells are further preferred.
- a T-cell or T-lymphocyte can be distinguished from other lymphocytes by the presence of a T-cell receptor (TCR) on its cell surface.
- TCR T-cell receptor
- One of the functions of T-cells is mediating immune-mediated cell death, and it is carried out by two major subtypes: CD8+ "killer” and CD4+ "helper” T-cells.
- CD8+ T cells are cytotoxic which means that they are able to directly kill selected cell. These selected cells are in accordance with the invention tumor cells, virus-infected cells, as well as cancer cells.
- CD4+ cells function as "helper cells”.
- these CD4+ helper T-cells function by further activating memory B cells and cytotoxic T-cells, which leads to a larger immune response which is in accordance with the invention directed against tumor cells.
- the specific adaptive immune response regulated by the T-helper cell depends on its subtype, which is distinguished by the types of cytokines they secrete.
- the T-cells are preferably a CD8+ killer T-cells or mixture of CD8+ killer and CD4+ helper T-cells.
- NK cell is a type of cytotoxic lymphocyte being critical to the innate immune system that belong to the rapidly expanding family of innate lymphoid cells (ILC) and represent 5-20% of all circulating lymphocytes in humans.
- the role of NK cells in innate immune system is analogous to that of cytotoxic T-cells in the vertebrate adaptive immune response.
- NK cells provide rapid responses to virus-infected cells and other intracellular pathogens acting at around 3 days after infection and respond to tumor formation.
- immune cells detect the major histocompatibility complex (MHC) presented on infected cell surfaces, triggering cytokine release, causing the death of the infected cell by lysis or apoptosis.
- MHC major histocompatibility complex
- NK cells are unique, however, as they have the ability to recognize and kill stressed cells in the absence of antibodies and MHC, allowing for a much faster immune reaction. They were named “natural killers” because they do not require activation to kill cells that are missing "self markers of MHC class 1 . This role is especially important because harmful cells that are missing MHC I markers cannot be detected and destroyed by other immune cells, such as T-cells.
- the lymphocytes may also be chimeric antigen receptor T-cells (CAR T-cells), T-cell-receptor- engineered T-cells (TCR T-cells), chimeric antigen receptor NK-cells (CAR NK-cells), NK cell receptor- engineered NK cells (NCR NK-cells), TCR/CAR hybrid T-cells, NCR/CAR hybrid NK-cells or tumorinfiltrating lymphocytes (TILs).
- CAR T-cells chimeric antigen receptor T-cells
- TCR T-cells T-cell-receptor- engineered T-cells
- CAR NK-cells chimeric antigen receptor NK-cells
- NCR NK-cells NK cell receptor- engineered NK cells
- TILs tumorinfiltrating lymphocytes
- Chimeric antigen receptor T-cells are T-cells that have been genetically engineered to produce a chimeric T cell receptor (CAR) for use in immunotherapy.
- the receptors are chimeric because they combine both antigen-binding and T-cell activating functions into a single receptor.
- CAR-T cell therapy uses T-cells engineered with CARs for cancer therapy.
- the premise of CAR-T immunotherapy is to modify T-cells to recognize tumor cells in order to more effectively target and destroy them in order to generate CAR T-cells.
- T-cells are harvested from subjects, genetically altered, and then infused into patients to attack a tumor in a subject.
- CAR T-cells can be both CD4+ and/or CD8+ cells. A 1 -to-1 ratio of both cell types is preferred since it provides synergistic antitumor effects.
- CAR T-cells are engineered to transfer arbitrary specificity onto a T cell, which specifically eliminates antigen-bearing tumor cells.
- the CAR may comprise a scFv being derived from an antibody, a CD3£ and a transmembrane domain (so-called first-generation CARs).
- first-generation CARs a transmembrane domain
- the engineered CAR is able to recognize specific tumor associated-antigens. Therefore, the CAR has the ability to bind unprocessed tumor surface antigens without MHC processing while TCRs engage with both tumor intracellular and surface antigenic peptides embedded in MHC.
- TCRs are a/p heterodimers that bind to the MHC-bound antigens.
- CARs recognize tumor antigen which led to T-cell activation with different functions compared with TCRs.
- CAR-T cell therapy has certain disadvantages like off-tumor toxicities when targeting tumorspecific antigen.
- TCRs Compared with CARs, TCRs have several structural advantages in T cell-based therapy, such as more subunits in their receptor structure (ten subunits vs one subunit), greater immunoreceptor tyrosine-based activation motif (ITAMs) (ten vs three), less dependence on antigens (one vs 100), and more co-stimulate receptors (CD3, CD4, CD28, etc.) (Zhao et al. (2021) Front. Immunol.,
- ITAMs immunoreceptor tyrosine-based activation motif
- CAR NK-cells are distinguished from CAR T-cells in that the chimeric antigen receptor is introduced into NK cells instead of T-cells.
- CAR-NK cells can be engineered to target diverse antigens, enhance proliferation and persistence in vivo, increase infiltration into solid tumours, overcome resistant tumour microenvironment, and ultimately achieve an effective anti-tumour response.
- Natural cytotoxicity receptor NK cells are NK-cell that have been genetically engineered to express a NCR. The NCRs have been proposed to bind to many cellular ligands which are implicated in NK cell surveillance of tumor cells.
- NCRs may regulate other anti-tumor pathways.
- NCRs and their ligands can be successfully targeted for cancer immunotherapy.
- NCRs have been classically defined as activating receptors delivering potent signals to NK cells in order to lyse harmful cells and to produce inflammatory cytokines.
- TCR/CAR hybrid T-cells are T-cells that have been genetically engineered to express a TCR and CAR.
- NCR/CAR hybrid NK-cells are T-cells that have been genetically engineered to express a NCR and CAR.
- TILs Tumor-infiltrating lymphocytes
- TILs are white blood cells that have left the bloodstream and migrated towards a tumor. TILs are implicated in killing tumor cells. The presence of lymphocytes in tumors is often associated with better clinical outcomes.
- the tumor-infiltrating lymphocytes are tumor-infiltrating T-cells or tumor-infiltrating NK cells.
- TILs are expanded ex vivo from surgically resected tumors that have been cut into small fragments or from single cell suspensions isolated from the tumor fragments. Multiple individual cultures are established, grown separately and assayed for specific tumor recognition. TILs are typically expanded over the course of a few weeks with a high dose of IL-2 in 24- well plates. Selected TIL lines that presented best tumor reactivity are then further expanded in a "rapid expansion protocol" (REP), which uses anti-CD3 activation for a typical period of two weeks. The final post-REP TIL is infused back into a patient in order to treat a tumor of the patient. This applies mutatis mutandis to adoptive NK-cell transfer with TILs.
- REP rapid expansion protocol
- the anti-tumor lymphocytes are autologous anti-tumor lymphocytes.
- lymphocytes are taken from a subject having a tumor and are genetically engineered (e.g. to produce CAR T-cells) and/or selected and/or expanded (e.g. to produce TILs) ex vivo and then transferred back into the same subject.
- CAR T-cells CAR T-cells
- TILs selected and/or expanded
- the present invention relates in a fifth aspect to a molecular complex comprising the variant of the GAIN domain of the first aspect or the AGPCR of the second aspect in association with a Stachel sequence or fragment thereof being capable of binding to the GAIN domain of claim 1 , wherein the Stachel sequence preferably comprises or consists of any one SEQ ID NOs 59 to 81 or a sequence being at least 70% identical, preferably at least 80% and most preferably at least 90% identical thereto.
- the variant of the GAIN domain or the AGPCR are capable of binding to each other via non-covalent interactions, thereby forming a molecular complex.
- Such molecular complexes are the subject of the sixth aspect of the present invention.
- the molecular complexes may be formed, for example, within a cell and also within in a suitable medium.
- amino acid sequence sharing at least 70% sequence identity displays with increasing preference at least 75% sequence identity, at least 80% sequence identity, at least 85% sequence identity, at least 90% sequence identity, at least 95% sequence identity, at least 96% sequence identity, at least 97% sequence identity, at least 98% sequence identity and at least 99% sequence identity.
- the variant of the GAIN domain or the AGPCR on the one hand and the Stachel sequence on the other hand are preferably derived from the same AGPCRS.
- the Stachel sequence is preferably a Stachel sequence of SEQ ID NO: 59 or a sequence being at least 70% identical, preferably at least 80% and most preferably at least 90% identical thereto
- the Stachel sequence is preferably a Stachel sequence of SEQ ID NO: 60 or a sequence being at least 70% identical, preferably at least 80% and most preferably at least 90% identical thereto., etc.
- the present invention relates in a sixth aspect to a composition, preferably a diagnostic or pharmaceutical composition, or a kit comprising the variant of thee GAIN domain, the AGPCR, the nucleic acid molecule, the vector, the host cell of the invention or a combination thereof.
- the term “pharmaceutical composition” relates to a composition for administration to a patient, preferably a human patient.
- the pharmaceutical composition of the invention comprises the compounds recited above. It may, optionally, comprise further molecules capable of altering the characteristics of the compounds of the invention thereby, for example, stabilizing, modulating and/or activating their function.
- the composition may be in solid, liquid or gaseous form and may be, inter alia, in the form of (a) powder(s), (a) tablet(s), (a) solution(s) or (an) aerosol(s).
- the pharmaceutical composition of the present invention may, optionally and additionally, comprise a pharmaceutically acceptable carrier.
- Suitable pharmaceutical carriers include phosphate buffered saline solutions, water, emulsions, such as oil/water emulsions, various types of wetting agents, sterile solutions, organic solvents including DMSO etc.
- Compositions comprising such carriers can be formulated by well-known conventional methods. These pharmaceutical compositions can be administered to the subject at a suitable dose. The dosage regimen will be determined by the attending physician and clinical factors. As is well known in the medical arts, dosages for any one patient depends upon many factors, including the patient's size, body surface area, age, the particular compound to be administered, sex, time and route of administration, general health, and other drugs being administered concurrently.
- the therapeutically effective amount for a given situation will readily be determined by routine experimentation and is within the skills and judgement of the ordinary clinician or physician.
- the regimen as a regular administration of the pharmaceutical composition should be in the range of 1 pg to 5 g units per day.
- a more preferred dosage might be in the range of 0.01 mg to 100 mg, even more preferably 0.01 mg to 50 mg and most preferably 0.01 mg to 10 mg per day.
- the length of treatment needed to observe changes and the interval following treatment for responses to occur vary depending on the desired effect. The particular amounts may be determined by conventional tests which are well known to the person skilled in the art.
- diagnostic composition relates to a composition, optionally for administration to a patient, preferably a human patient and may comprise the essentially same additional compounds as discussed in connection with the pharmaceutical composition. While pharmaceutical compositions are to cure or prevent a disease, a diagnostic composition is to identify the presence and optionally also the site of a disease in a subject.
- the diagnostic composition preferably comprises a diagnostically active compound as discussed herein above, such as a radiolabel or a fluorophore.
- the various components of the kit may be packaged into one or more containers such as one or more vials.
- the vials may, in addition to the components, comprise preservatives or buffers for storage.
- the kit may comprise instructions how to use the kit, which preferably inform how to use the components of the kit for diagnosing a tumor and/or for grading a tumor and/or for tumor prognosis.
- the present invention relates in a seventh aspect to a computer-implemented method for increasing or decreasing the association of the GAIN domain of an AGPCR and the Stachel sequence under force comprising (a) performing in silico steered molecular dynamics (SMD) simulations on the GAIN domain in association with the Stachel sequence whereby the GAIN domain undergoes topological rearrangements that enable force to propagate through the GAIN domain to the associated Stachel sequence, (b) identifying force propagation pathways from the SMD simulations of (a) by analysing the correlations between alpha carbon motions within the GAIN domain, (c) identifying the arrival angles of the optimal force propagation pathways as obtained in (b) on the Stachel sequence, (d) mutating in silico based on the information as obtained in (a) to (c) one or more pairs of amino acids of the GAIN domain outside of the interface of the GAIN domain that associates with the Stachel sequence to cysteine residues and/or substituting in silico based on the information as obtained in
- a preferred embodiment of the seventh aspect of the invention is directed to a computer-implemented method for increasing or decreasing the association of the GAIN domain of an AGPCR and the Stachel sequence under force comprising (a) performing in silico steered molecular dynamics (SMD) simulations on the GAIN domain in association with the Stachel sequence whereby the GAIN domain undergoes topological rearrangements that enable force to propagate through the GAIN domain to the associated Stachel sequence, (b) identifying force propagation pathways from the SMD simulations of (a) by analysing the correlations between alpha carbon motions within the GAIN domain, (c) identifying the arrival angles of the optimal force propagation pathways as obtained in (b) on the Stachel sequence, (d) mutating in silico based on the information as obtained in (a) to (c) one or more pairs of amino acids of the GAIN domain outside of the interface of the GAIN domain that associates with the Stachel sequence to cysteine residues to obtain a variant of the GAIN domain with a disul
- amino acid substitutions are preferably non-covalent amino acid substitutions and/or the substitution of amino acid(s) other than cysteine(s) by amino acid(s) other than cysteine(s). These amino acid(s) is/are more preferably the naturally-occurring proteinaceous amino acid(s) other than cysteine(s)
- a computer-implemented method as used herein designates a method which involves the use of a computer, computer network or other programmable apparatus, where one or more features are realised wholly or partly by means of a computer program. Since the method is a computer- implemented method, also a computer-implemented program is described herein that when being executed on a computer causes the computer to carry out the method of the first aspect of the invention.
- step (a) of the method of the invention in silico steered molecular dynamics (SMD) simulations are performed on the GAIN domain in association with the Stachel sequence whereby the GAIN domain undergoes topological rearrangements that enable force to propagate through the GAIN domain to the associated Stachel sequence.
- SMD silico steered molecular dynamics
- the Stachel is a structural part of the GAIN domain.
- the GAIN domain is folded, so that it is association with the GAIN domain.
- the Stachel dissociates, the GAIN partially unfolds. The SMD simulations are performed starting from the Stachel bound state.
- SMD Steered molecular dynamics
- step (a) the GAIN domain undergoes topological rearrangements that enable force to propagate through pathways favoring its mechanical resistance.
- the SMD simulation of step (a) is preferably essentially performed as described in the below Example 3, Section “SMD simulations”.
- SMD simulations are preferably performed with GROMACS 2019.4 (https://manual.gromacs.Org/documentation/2019.4/download.html) using the CHARMM36 forcefield (https://www.charmm.org/archive/charmm/) along with the TIP3 (transferable intermolecular potential with 3 points, which is implemented into CHARMM36) water model in an NPT ensemble (also called isothermal-isobaric ensemble; constant temperature and constant pressure ensemble) at about 31 OK (about 36,85°C) and about 1 bar using a velocity rescaling thermostat (with a relaxation time of 0.1 ps) and Berendsen barostat (with a relaxation time of I ps). Equations of motion are integrated with a timestep of about 2 fs using a leap-frog algorithm (e.g. https://www.compphy.com/leap-frog/
- Each system is preferably energy minimized using a steepest descent algorithm (https://people.seas.harvard.edu/ ⁇ yaron/AM221-S16/lecture_notes/AM221_lecture10.pdf), more preferably for about 5000 steps, and then equilibrated with the atoms of the protein restrained using a harmonic restraining force in two steps, first by adding the thermostat for about 50000 steps and then the barostat for about 500000 steps.
- a steepest descent algorithm https://people.seas.harvard.edu/ ⁇ yaron/AM221-S16/lecture_notes/AM221_lecture10.pdf
- SMD simulations are preferably performed using the umbrella sampling method (e.g. http://jeti.uni-freiburg.de/studenten_seminar/term_paper_WS_16_17/Kaestner11 .pdf) with a pulling speed ranging from about 1 A/ns to about 10 A/ns.
- umbrella sampling method e.g. http://jeti.uni-freiburg.de/studenten_seminar/term_paper_WS_16_17/Kaestner11 .pdf
- the 2 alpha carbons of the N-term and C-term residues are moved harmonically in the desired direction with constant velocity.
- a distance cutoff of about 14 A is then used for short-range non-bonded interactions, whereas electrostatic interactions are computed using the Particle-mesh Ewald method with a cubic interpolation of power 4.
- step (b) of the method of the invention force propagation pathways from the SMD simulations of (a) are identified by analysing the correlations between alpha carbon motions within the GAIN domain.
- step (b) When force is propagated through a molecule, soft degrees of freedom will be stretched out along the path of force propagation. These paths of force propagation are identified according to step (b).
- the coordinates of a-carbons are extracted from the time window corresponding to the GAIN “loaded” state (i.e. the GAIN domain in association with the Stachel sequence). Their sequential motions between two consecutive frames of the simulation are then computed and correlation analyses is performed to reconstruct the force propagation pathways from these motions. Pathways are identified by following the route from the N- to C-terminal that maximizes the observed correlations. Finally and preferably, identified pathways between consecutive frames are clustered and a deep analysis of cluster populations enables to discern optimal (primary cluster) from suboptimal pathways (secondary clusters).
- step (c) of the method of the invention the arrival angles of the optimal force propagation pathways as obtained in (b) on the Stachel sequence are identified.
- the arrival angles are determined.
- the arrival angles indicate the orientation of feree on the Stachel mechanical clamp between the GAIN domain and the Stachel sequence.
- the arrival angles are preferably defined as the angles between the force pathway vector and a Stachel network plane.
- step (d) of the method of the invention based on the information as obtained in (a) to (c) one or more pairs of amino acids of the GAIN domain outside of the interface of the GAIN domain that associates with the Stachel sequence are mutated in silico to cysteine residues to obtain a variant of the GAIN domain with a disulfide bond and/or are substituted in silico to other amino acid(s) at to obtain a variant of the GAIN domain with specific amin acid substitution(s) selected site(s) that are expected to decrease or increase the arrival angles as compared to the arrival angles as identified in (c).
- Coupled Moves is a flexible backbone design method meant to be used for designing small-molecule binding sites, protein-protein interfaces, and protein-peptide binding sites (https://www.rosettacommons.org/docs/latest/coupled-moves).
- the Backrub application is useful for creating ensembles of protein backbones, modeling protein flexibility, modeling mutations, and detailed refinement of backbone/side chain conformations.
- the backrub algorithm rotates local segments of the protein backbone as a rigid body about an axis defined by the starting and ending atoms of the segment. Atoms branching of the main chain at the pivot points (side chains, hydrogens, carbonyl oxygens), are updated to minimize the bond angle strain incurred. These moves are accepted or rejected using the Metropolis criterion (https://www.rosettacommons.org/docs/latest/application_documentation/structure_prediction/backru).
- a total of about 100 independent structures is generated for every pair of cysteines and/or every amino acid substitution or set of amino acid substitutions, and the about top 10% lowest energy structures are conserved. All structures are then minimized and relaxed to access lowest energy potential through the relax protocols included in the Rosetta package. Distances between the groups of added cysteines are computed and averaged across the lowest energy structures. To determine which pairs of mutations form a disulfide, a distance cutoff of about 2.15 A between the 2 cysteines is preferably used to conserve or exclude mutants. Similarly, the lowest energy structures can be computed for every amino acid substitution or set of amino acid substitutions.
- steps (a) to (c) are repeated on the variant of the GAIN domain as obtained in (d) thereby identifying the arrival angles of the optimal force propagation pathways on the Stachel sequence for the variant of the GAIN domain in (d), wherein increased arrival angles are indicative for an increased strength of the association of the GAIN domain and the Stachel sequence under force and decreased arrival angles are indicative for a decreased strength of the association of the GAIN domain and the Stachel sequence under force.
- the method of the seventh aspect is advantageously the first computational method integrating dynamics under mechanical stress and systematic calculation of deformation upon sequence variation that can be applied for rationally designing mechanosensing properties.
- AGPCRs play a very important role as drug target. The majority of approved pharmaceutics targets GPCRs and AGPCRs are second largest group of GPCRs. This underlines the important contribution of the method of the seventh aspect to the prior art.
- step (a) two or more different pulling speeds are applied to the GAIN domain, wherein the pulling speeds are preferably selected from about 0.025 nm.ns -1 , 0.1 nm.ns -1 and about 1.0 nm.ns -1 .
- the GAIN domain can be stretched through constant velocity pulling.
- the SMD atom is attached to a dummy atom via a virtual spring. This dummy atom is moved at a constant pulling speed and then the force between both is measured.
- This preferred embodiment results in an increase of the forces that can be accumulated within the SMD system that is commonly defined as the loaded state of the protein (i.e. GAIN domain in association with the Stachel sequence).
- the appliance of two or more different pulling speeds provides the technical advantage of obtaining more information during the SMS simulations that can be used in the subsequent steps of the method of the invention. Thereby a better understanding of how force propagates through the GAIN domain to the Stachel peptide.
- the two pulling speeds as used in the appended examples are 0.1 nm.ns -1 and 1.0 nm.ns -1 .
- step (a) a harmonic force potential is applied to the GAIN domain by pulling simultaneously from the N- and C- terminus of the GAIN domain.
- the stretching between two atoms may be defined by the linear harmonic potential, which describes the energy associated with deviations from the equilibrium distance. Pulling simultaneously from the N- and C-terminus of the GAIN domain determines the direction of the force being applied and ensures that the force can propagate throughout the GAIN domain to the Stachel sequence.
- step (b) for every frame the three-dimensional coordinates for every alpha carbon of the GAIN domain are extracted, for every pair of two consecutive frames of the simulation the distance displacements of the three-dimensional coordinates are obtained, and the distance displacements are correlated to reconstruct the optimal and suboptimal force propagation pathways from the N- to the C-terminus and computing force propagation pathways.
- the above preferred embodiment describes the preferred determination of force propagation pathways.
- SMD simulations simulate the evolution and motions of a biomolecular system as a function of time. Motions are assessed through changes in atomic positions recorded at discrete time steps called “frames”. Consecutive frames correspond to consecutive time steps of a few picoseconds.
- force propagation pathways are essentially determined as described in Example 3, section “SMD simulation analysis and force propagation pathway identification”.
- correlation between alpha carbon motions are more preferably analyzed.
- the [x,y,z] cartesian coordinates of every alpha carbon are extracted. Displacement of these coordinates between 2 consecutive frames are computed and represented as a vector [xo-xi, yo-yi, zo-zi].
- Dijkstra shortest path algorithm e.g. https://www.geeksforgeeks.org/dijkstras- shortest-path-algorithm-greedy-algo-7/
- the shortest path maximizing the total correlation is searched between the N-term of the GAIN domain and the C-term of the Stachel mechanical clamp for every correlation matrix, resulting in / shortest force propagation pathways.
- k-means clustering was performed.
- the optimal number of clusters to use for independent simulations is determined using the inertia attribute to identify the sum of squared distances of samples to the nearest cluster center. Centroids of clusters, corresponding to the optimal force propagation pathway within this cluster were used for later extraction of features.
- step (c) the arrival angles are calculated with respect to a plane that is formed between the N-terminal and C-terminal alpha carbon of the Stachel sequence and a selected alpha carbon in the middle of the Stachel sequence.
- the angles of the force propagation pathways arrival on the Stachel sequence are useful for predicting the effect of the different Cys-Cys mutants and/or non-covalent amino acid substitutions in-silico.
- This fixed plane is in accordance with the above preferred embodiment formed between the N-terminal and C-terminal alpha carbon of the Stachel sequence and a selected alpha carbon in the middle of the Stachel sequence.
- the Stachel sequences are highly conserved among the 33 AGPCRs.
- the Stachel sequences of AGPCRs have a length of 7, 8 or 9 amino acids and the “middle of the Stachel sequence” is in the case of 7 amino acids position 3-5 and preferably position 4, in the case of 8 amino acids position 4 or 5 in the case of 9 amino acids position 4-6 and, preferably position 5.
- the arrival angles are even more preferably essentially determined as described in Example 3, “force propagation pathway projection on Stachel rupture interface”.
- coordinates of all alpha carbons are extracted. These coordinates represent a [f, n r esidues, nresidues] matrix, t being the number of frames of the simulation.
- a plane containing both the Stachel N-term and C-term alpha carbons, with the alpha carbon of the middle of the Stachel sequence is constructed. This results in an ensemble of dynamic planes moving according to local structural rearrangements and corresponding to the rupture interface of the Stachel mechanical clamp.
- the method comprises in addition the identification of the backbone H-bonds that ensure or determine the mechanical resistance of the Stachel mechanical clamp to force.
- Identified H- bonds across the simulation may be 2D-projected in a plane perpendicular to the Stachel axis defined by its N-term and C-term residues.
- the 2D projections can be compared between the native unloaded structure and the loaded structure corresponding to the highest mechanical resistance point observed in the SMD simulations. Projections may be averaged across replicates and plotted.
- the method further comprises producing the variant of the GAIN domain as obtained after steps (a) to (d).
- the final “product” is a modelled variant of the GAIN domain as obtained after steps (a) to (d) that either displays increased arrival angles being indicative for an increased strength of the association of the GAIN domain and the Stachel sequence under force or displays decreased arrival angles being indicative for a decreased strength of the association of the GAIN domain and the Stachel sequence under force.
- variant of the GAIN domain is actually produced by peptide and/or protein synthesis or by site-directed mutagenesis.
- Naturally occurring proteins and peptides site-directed mutagenesis
- the naturally occurring proteins and peptides are engineered to proteins and peptides being based on the naturally occurring proteins and peptides that comprise the engineered Cys-Cys bonds and/or non-covalent amino acid substitutions as discussed herein above.
- the variant of the GAIN domain may be implemented in an AGPCR.
- the variant of the GAIN domain may also be produced as part of a full-length AGPCR or a fragment thereof comprising the variant of the GAIN domain.
- the produced products may be C-terminally amidated in order to avoid unwanted charge effects of the carboxy terminus in further experiments or uses.
- the method further comprises applying single-molecule force spectrometry (SMFS) to the association of the produced variant of the GAIN domain and the Stachel sequence in order to identify the interface of variant of the GAIN domain that associates with the Stachel sequence and/or the force that disrupts the association, wherein the SMFS is preferably atomic force microscopy (AFM) based SMFS.
- SMFS single-molecule force spectrometry
- the SMSF is not part of the computational method. It is used independently to determine experimentally the force required to dissociate the Stachel and validate the computational predictions.
- the interface of the variants of the GAIN domain that associates with the Stachel sequence is determined from the X-ray structures of the variants of GAIN domain.
- the variant of the GAIN-domain is fused to a Fgp domain and a Spytag whereby the Fgp domain connects the variant of the GAIN domain via an isopeptide bond to a surface, preferably a coverglass surface and the Fgp domain connects the variant of the GAIN domain to a cantilever.
- the SMFS is essentially as is shown in Fig. 5 and Example 3.
- Such and similar method are art-established (Yang et al. Front Mol Biosci. 2020; 7: 85.)
- a fingerprint domain preferably a FLN4 domain is located.
- a fingerprint domain in particular a FLN4 domain is advantageous for filtering the obtained data as is described in detail in Example 3, section AFM-SMFS measurement and data analysis.
- the SMFS data traces were filtered by searching for contour length increments that matched the lengths of the fingerprint domains, FLN (about 36 nm).
- ALc 36.9 nm - Lf, where Lf is typically ⁇ 5nm.
- the contour length increment is between first FLN unfolding and final stachel dissociation (Fig. 7).
- two different, preferably three different and most preferably four different pulling speeds are applied in the SMFS, wherein the different pulling speeds are preferably selected from about 400, about 800, about 1600 and about 3200 nm.s -1 .
- the appliance of two different, preferably three different and most preferably four different pulling speeds improved the accuracy of the SMFS since forces over different times are thereby applied.
- pulling speeds of 400, about 800, about 1600 and about 3200 nm.s -1 are used.
- the method further comprises measuring the force that disrupts the association between the variant of the GAIN domain and the Stachel sequence, preferably by an serum response element (SRE) assay wherein shear stress, preferably through vibration or spinning is applied to cells that express an aGPCR comprising the variant of the GAIN domain, wherein the aGPCRs on the cells are coupled via aGPCR ligands, preferably collagen III or transglutaminase 2 onto a plate.
- SRE serum response element
- SMFS is not the only method measuring the force that disrupts the association between the variant of the GAIN domain and the Stachel sequence and many other methods are available as well.
- SRE serum response element
- SRE serum response element
- each embodiment mentioned in a dependent claim is combined with each embodiment of each claim (independent or dependent) said dependent claim depends from.
- a dependent claim 2 reciting 3 alternatives D, E and F and a claim 3 depending from claims 1 and 2 and reciting 3 alternatives G, H and I
- the specification unambiguously discloses embodiments corresponding to combinations A, D, G; A, D, H; A, D, I; A, E, G; A, E, H; A, E, I; A, F, G; A, F, H; A, F, I; B, D, G; B, D, H; B, D, I; B, E, G; B, E, H; B, E, I; B, F, G; B, F, H; B, F, I; C, D, G; C, D, H; C, D, I; C,
- the adhesion GPCR GAIN domain is a bona fide mechanosensing protein domain which dissociates from the Stachel peptide in a single concerted shearing event (A)
- A General framework for the study of the ADGRG1 GAIN domain under force.
- AlphaFold structure prediction UniProt Q9Y653
- FIG. 3A Structural view of the Stachel backbone (cyan) H-bond network (dotted magenta lines) formed with adjacent p-strands b6 and b8 in the human ADGRG1 GAIN domain (see also Fig. 3A). The right-most panel shows a view from C- terminal end of the Stachel.
- C Representative force-extension trace of the WT ADGRG1 GAIN domain showing clear ddFLN4 unfolding events and a single GAIN-Stachel dissociation event at ⁇ 200 nm extension.
- D Dynamic force spectrum of the GAIN domain at cantilever retraction velocities of 0.4 pm.
- Fig. 2. SMD simulations provide metrics to characterize and redesign the structural and dynamic determinants of GAIN mechanical resistance.
- A SMD simulation and data analysis pipeline used to obtain a mechanistic understanding of GAIN domain behavior under force (see methods in Example 3).
- B-E Violin plots of the distributions of the average angles between force pathways and the Stachel hydrogen bond network in the loaded state from SMD simulations at pulling speeds of 0.1 and 1.0 nm.ns -1 , forWT, SS_h2l10, SS_h1 h2 and SS_h1 h2-h2l10 GAIN domains.
- FIG. 3 Modulating the GAIN domain force dissociation threshold through engineered disulfides.
- A Topology chart of the GAIN domain with the positions of disulfide-forming cysteine mutations shown as red dots and dotted lines, a-helices (h) and p-strands (b) are numbered and define the naming nomenclature of the mutants. Loop 10 (110) is shown in purple and connects p- strands 8 and 9. The a-helical portion of the GAIN is colored green and p-strands are colored blue or orange according to the p-sheet they belong to. The C-terminal Stachel is shown in cyan and the autoproteolysis site is marked by a green asterisk.
- Backbone-backbone H-bonds between the Stachel and adjacent p-strands 6 and 8 are represented as dashed magenta lines (see also Fig. 1 B).
- B Crystal structure of the SS-h1 h2-h2l10 GAIN domain variant at 1.9 A resolution, following the coloring scheme in Panel A. The right panel shows the transparent GAIN surface (grey) and Stachel surface (cyan).
- C GAIN variant rupture-force probability density histograms at the indicated cantilever retraction velocities overlaid on the corresponding WT data in transparent red. A Gaussian fit is shown for each distribution.
- (E) 2D scatter plot showing the differences in mechanical and thermodynamic stabilities between variants relative to WT. Data is presented as the difference to WT for each measurement. A single average and standard deviation (n 3) value is plotted for the thermodynamic stability obtained from circular dichroism measurements. The median of the rupture-force probability densities of each variant minus WT is plotted for each cantilever retraction velocity shown in Panel C, giving four values of mechanical stability for each variant.
- Fig. 4 Contrasted effects of the designed disulfides on the activation of full-length ADGRG1 expressed on cells and stimulated by vibration and receptor-specific ligands.
- Fig. 5 Schematic of the cantilever and coverglass surface modification and protein immobilization for AFM-SMFS experiments (see Example 3).
- FIG. AFM-SMFS of GAIN variants for contour length analysis.
- A Schematic illustration of the experimental setup for comparative AFM-SMFS of GAIN-WT and GAIN(4Cys). GAIN-WT and GAIN(4Cys) were immobilized on different spots on the surface. SdrG was immobilized at the cantilever tip. The same cantilever was used to probe each sample at 3200 nm s-1.
- B Contour length of first FLN unfolding and final stachel dissociation using GAIN-WT and GAIN(4Cys) with the same cantilever tip.
- Fig 7. Snapshots from SMD simulations at 1 nm.ns' 1 showing force propagation pathways (red) represented on the native and loaded state structures. Residues involved in each pathway are listed and colored following the color code of the 2D topology chart in Figure 2. The average angle of the force pathway vector on the Stachel mechanical clamp is indicated for each GAIN variant for reference. Fig 8. Thermodynamic stability of purified GAIN domain variants measured by circular dichroism spectroscopy. CD melting curves from 20°C to 96°C at 216 nm and 232 nm are shown for WT and the three GAIN domain disulfide variants.
- FIG. 1 GAIN domain purification for AFM and X-ray crystallography.
- A Gel-filtration (Superdex 200 Increase 10/300 GL) elution profiles for representative purifications of the WT GAIN domain construct for AFM experiments and of the SS_h1 h2-h2l10 GAIN domain construct for X-ray crystallography (suffix “X”). Protein constructs for X-ray crystallography lack the N- and C-terminal AFM handles (see full protein sequences).
- B SDS-PAGE analysis of the purified WT and variant GAIN domains and of the SS_h1 h2-h2l10_X GAIN domain construct. The lower molecular weight bands correspond to the cleaved C-terminal containing the Stachel and tags and/or handles.
- FIG. 10 AFM-SMFS of GAIN-WT and NoGAIN in native or reducing conditions for contour length analysis.
- A Dynamic force spectra of GAIN-WT for stachel dissociation force according to the loading rate in reducing (100 mM DTT) and non-reducing condition (0 mM DTT). Black circles represent the median rupture force/loading rate in non-reducing condition and purple circles represent in reducing condition at each pulling speed of 400, 800, 1600, and 3200 nm s -1 . Error bars are ⁇ 1 s.d. Solid lines are least square fits to the Bell-Evans model.
- H1L7 force propagation angle The H1 L7 force propagation angle is very similar to that of the H1 H2 variant. Consistent with these predictions, the H1 L7 variant is mechanically weaker than WT and behaves similarly to the H1 H2 variant as measured by AFM.
- FIG. 13 AFM - non-covalent designs. Distribution & Median Value of Dissociation Forces. Ruptureforce probability density histograms for velocities of 0.4 pm.s -1 , 0.8 pm.s -1 , 1.6 pm.s -1 and 3.2 pm.s -1 projected onto individual axes. Fig 14. AFM - Disulfide designs. Dynamic force spectrum of the GAIN domain at cantilever retraction velocities of 0.4 pm.s -1 , 0.8 pm.s -1 , 1.6 pm.s -1 and 3.2 pm.s -1 .
- FIG. 15 AFM - Disulfide designs. Distribution & Median Value of Dissociation Forces. Rupture-force probability density histograms for velocities of 0.4 pm.s -1 , 0.8 pm.s -1 , 1.6 pm.s -1 and 3.2 pm.s -1 projected onto individual axes.
- Fig 16. Constitutive activity SRE assays. Constitutive activities measured using the SRE assay for GAIN variants expressed in the context of the full length receptor (dark grey) or the receptor truncated from the adhesion PLL domain (light grey), (top) “raw” assay data, (bottom) normalized to WT.
- TG2-induced activity SRE assays a. TG2 Ligand induced activities measured using the SRE assay for GAIN variants expressed in the context of the full length receptor (dark green) compared to activities without ligand (grey), b. TG2 Ligand induced activities normalized to those without ligand.
- SMFS single-molecule force spectroscopy
- SMD steered molecular dynamics
- computational protein design and cell-based signaling are used assays to characterize and reprogram the response of the human ADGRG1 GAIN domain to mechanical force (Fig. 1A).
- silico SMD simulations and computational protein design are fundamental to arrive reprogrammed human ADGRG1 GAIN domain which are either more or less sensitive to mechanical force and subsequently SMFS as well as cell-based signaling assay can be used to show and measure the changed sensitivity more or less sensitive to mechanical force experimentally.
- SMFS setup that best mimics the predicted orientation and direction of force application on the GAIN domain in the context of the full-length ADGRG1 receptor expressed at the cell-surface (Fig. 1A).
- Force-extension curves were recorded at constant speed using an atomic force microscope and ddFLN4 unfolding fingerprints were used to select single-molecule unfolding events for statistical analysis.
- WT GAIN-Stachel dissociation events were measured in the range of 100-160 pN with a clear dependence on the force loading rate (Fig. 1C-D). Based on the shape of the forceextension curve (Fig.
- GAIN domain unfolding from the Stachel peptide appears to be a single concerted dissociation event with no predicted unfolding intermediates.
- the GAIN domain thus functions as a two-state mechanosensitive system and encodes a high force resistance for a cell-surface receptor involved in adhesion and signaling (Fig. 1 E).
- adhesion GPCRs Unlike other adhesion receptors that respond to mechanical force such as integrins or Notch, adhesion GPCRs have a relatively simple architecture. Furthermore, our results show that a major component of their signaling, i.e. GAIN shedding and Stachel-dependent receptor activation, involves a single conformational transition at a force of 100 pN or above. Even though the exact position of this activation barrier may differ in the context of the full-length receptor at the cell surface, these results suggest that there is a large range available for the reprogramming of the mechanical properties of the GAIN domain, both towards lower and higher mechanical strength. Coupled to the modularity of their adhesion domains, adhesion GPCRs and their GAIN domains represent ideal scaffolds for the design of mechanosensors with specific thresholds of force activation.
- a traditional approach to modify the mechanical stability of this domain would consist in generating mutations predicted to alter the local environment of the mechanical clamp based on analysis of the native (static) structure alone, as reported previously (22, 36).
- the Stachel environment and broader GPS motif encode multiple other functions relevant to the full-length receptor context.
- the GPS motif ensures proper conformation of the fold to enable autoproteolysis and the Stachel sequence is highly conserved among adhesion GPCRs and enables receptor activation by engagement with the orthosteric pocket upon GAIN shedding.
- the consequences on folding of mutations in the domain’s core would be hard to predict, and even mutations that enable proper folding might not affect mechanostability since the mechanical clamp is mediated solely by backbone hydrogen bonds.
- Figure 2B shows the results of the WT GAIN SMD simulations at the two pulling speeds.
- the optimal and suboptimal pathways corresponding to the three largest clusters are represented as violin plots that depict the distribution of the average force angles upon contact with the Stachel hydrogen bond network in the loaded state.
- the relative size of each pathway cluster is given in percentages above each violin plot distribution and indicates how much each pathway is “used” by the GAIN domain to transmit force.
- the residues involved in the suboptimal pathways change between the two pulling speeds, the WT optimal force propagation pathway uses the same residues at both pulling speeds, suggesting a higher robustness.
- the average angle it forms with the Stachel varies with pulling speed (Fig. 2B), indicating speed-specific structural rearrangements of the topology.
- the optimal pathway shows a quick jump from the helix 1 N-terminal pulling point to the p-domain and, unlike the suboptimal pathways, transits to the Stachel without utilizing the a-domain, suggesting a preferred force propagation through the mechanically resistant p-sandwich.
- the lower force angle of the optimal pathway on the Stachel at the higher pulling speed 23° vs 40°
- the higher pulling speed enhanced the relative importance of the optimal path (from 40% to 52%) over the two suboptimal paths.
- the SS_h2l10 variant should have increased mechanical stability and a higher GAIN-Stachel dissociation threshold compared to WT.
- the SS_h1 h2 variant which locks the N-terminal of helix 1 with the C-terminal of helix 2 (Fig. 3A)
- the optimal pathway in this variant displayed much lower average angles of the force vector on the rupture interface at both pulling speeds compared to WT (10° and 2° compared to 23° and 40° for the WT optimal path at 0.1 and 1.0 nm.ns -1 , respectively; Fig. 2D). Due to the increased collinearity of the force vector and rupture interface, we predict a lower force threshold of GAIN-Stachel dissociation in this variant. The combination of the two disulfides in the SS_h1 h2-h2l10 variant also generated WT-like optimal and suboptimal pathways at both pulling speeds. Interestingly, the h2l10 disulfide was not utilized to propagate force in this variant.
- the SS_h1 h2-h2l10 variant should have a GAIN-Stachel force dissociation threshold close to that of the WT GAIN domain.
- DMEM Dulbecco's Modified Eagle's medium
- FBS Fetal Bovine Serum
- Penicillin-Streptomycin cocktail (PenStrep)
- Bovine Serum Albumin (BSA, Sigma)
- AFM handles - SpyCatcher and SdrG proteins for SMFS were prepared as previously reported (7, 2).
- Constructed recombinant plasmids pET28a-ybbR-His-ELP(MV7E2)3-FLN-SpyCatcher (Addgene #157674) and pET28a-SdrG-FLN-ELP-His-ybbr (Addgene #168047) were transformed into E. coli BL21 (DE3) strain. Cells were cultured in 5 ml of Luria-Bertani (LB) medium with 50 pg ml -1 kanamycin at 37 °C overnight.
- LB Luria-Bertani
- the culture was transferred to 100 mL of Terrific Broth (TB) medium with 50 pg ml’ 1 kanamycin and cultivated at 37 °C and 200 rpm until an optical density at 600 nm (GD600) of ⁇ 0.8- 1.0 was reached.
- Recombinant protein expression was induced upon addition of 0.5 mM isopropyl-p- D-thio-galactopyranoside (IPTG) and the culture was further incubated at 20 °C and 200 rpm for 18 h.
- the cells were harvested by centrifugation at 4,000 g for 20 min at 4 °C.
- the cell pellets were stored at -80 °C until further purification.
- the harvested cell pellet was resuspended in a lysis buffer (50 mM Tris, 50 mM NaCI, 0.1% Triton X-100, 5 mM MgCI2; pH 8.0), and disrupted with a sonic dismembrator.
- the lysate was centrifuged at 14,000 g for 20 min at 4 °C.
- the supernatant was collected and incubated with Ni-NTA resin (Thermo Fisher Scientific), loaded onto a column, washed with wash buffer (1x PBS buffer with 25 mM imidazole; pH 7.2), and eluted in elution buffer (1x PBS buffer with 500 mM imidazole; pH 7.2).
- the eluted protein solution was buffer-exchanged and further purified by Superose SEC column, and finally stored in 33% glycerol at -20 °C.
- NoGAIN contruct The same contruct with GAIN-WT but without GAIN domain
- newly constructed recombinant plasmid pET28a-Fgb-NoGAIN-HIS-SpyTag was used. Recombinant protein expression was induced upon addition of 1 mM IPTG and the culture was further incubated at 37 °C and 200 rpm for 9 h. Purification was carried out in the same manner as previously described.
- Ni-NTA resin Due to the small size of NoGAIN, elutant from Ni-NTA resin was directly incubated with purified ybbR-His-ELP(MV7E2)3-FLN- SpyCatcher. The conjugated form was then further purified by Superose SEC column, and finally stored in 33% glycerol at -20 °C.
- GAIN domain - Insect codon-optimized human ADGRG1 ECR sequences (Uniprot: Q9Y653) were synthesized (Twist Biosciences) as part of larger constructs for either AFM or X-ray crystallography experiments (see full protein sequences) in a pFastBac vector backbone.
- a baculovirus expression system was used for expression of proteins in High-Five insect cells.
- pFastBac constructs were transformed into DHWMultiBac cells (Geneva Biotech), white colonies indicating successful bacmid recombination were selected, and bacmids were purified by the alkaline lysis method.
- Sf9 cells were transfected using the X-tremeGENE 9 transfection reagent (Sigma) with the desired bacmid. eYFP-positive cells were observed after 1 week and subjected to one round of viral amplification. Amplified P2 virus was used to infect High-Five cells at a density between 1-2 x 10 6 cells/ml.
- the three or four purest and most concentrated fractions were pooled and concentrated on a 30 kDa Amicon Ultra-4 centrifugal filter (UFC803024, Merck Millipore) down to 500 pl.
- the concentrated proteins were loaded on a Superdex 200 Increase 10/300 GL (28-9909-44, Sigma-Aldrich) connected to an AKTA Pure chromatography system (GE) and eluted in HBS.
- GE AKTA Pure chromatography system
- the fractions containing the peak corresponding to the purified GAIN construct were pooled and quantified using A280 absorbance.
- Purified GAIN proteins were either directly concentrated to 10 mg. ml -1 for crystallization or flash-frozen in liquid nitrogen and stored at -80°C for subsequent AFM experiments.
- cantilevers and coverglasses were cleaned by UV-ozone treatment for 40 min and coverglasses were soaked in piranha etching solution and rinsed with distilled water (DW). Then, cantilevers and coverglasses were treated with 3-Aminopropyl (diethoxy) methylsilane (APDMES, ABCR GmbH, Düsseldorf, Germany) to silanize the surfaces with amine groups.
- APDMES 3-Aminopropyl (diethoxy) methylsilane
- the amine groups subsequently reacted to a NHS group from sulfosuccinimidyl 4-(N-maleimidomethyl)cyclohexane-1- carboxylate (sulfo-SMCC; Thermo Fischer Scientific) in 50 mM HEPES buffer pH 7.5 for 30 min.
- the thiol group from Coenzyme A (CoA, 200 pM) reacted to a maleimide group from sulfo-SMCC in coupling buffer (50 mM sodium phosphate, 50 mM NaCI, 10 mM EDTA, pH 7.2) for 2 hrs.
- the ybbR-tagged proteins SdrG-FLN-ELP-His-ybbr and pre-conjugated ybbR-His-ELP(MV7E2)3-FLN- SpyCatcher-SpyTag-GAIN variants were site-specifically anchored to the surface using SFP-mediated ligation to CoA in Mg2+ supplemented 1x PBS buffer. This resulted in covalent immobilization of SdrG and GAIN variants to cantilever and cover glasses, respectively. Protein-immobilized cantilevers and coverglasses were extensively washed and kept in 1x PBS buffer prior to immediate use.
- SpyTag- SpyCatcher conjugation of ybbR-His-ELP-FLN-SpyCatcher and Fgp-GAIN-His-SpyTag variants were done by mixing two proteins with same molar ratio in PBS buffer and pre-incubation for 1-12 hrs prior to ybbR tag ligation.
- Force spectroscopy measurements with SdrG and Fgp-fused GAIN variants were conducted in the same manner as previously illustrated using automated AFM-based SMFS (Force Robot 300, JPK Instruments). SMFS data were recorded in 1x PBS at room temperature with constant pulling speeds of 400, 800, 1600 and 3200 nm.s 1 . A total of 258,044 force-extension curves were acquired. Forceextension curves were filtered and analyzed by a combination of software available on the AFM instrument and custom python scripts. The majority of data traces contained no interactions, nonspecific interaction, or complex multiplicity of interactions.
- AL C 36.9 nm - Lf, where Lf is typically ⁇ 5nm.
- the total number of force-extension curves matching this criterion was 509 out of 120,780, 1 ,058 out of 68,376, 263 out of 42,295, and 158 out of 18,447 for GAIN-WT, GAIN(2Cys), GAIN(2Cys)-h1 h2, and GAIN(4Cys), respectively.
- the rupture or unfolding forces vs. loading rate was plotted and median forces and loading rates for each pulling speed were fitted to Bell-Evans model (3, 4) to estimate the effective distance to the transition state (Ax) and the intrinsic dissociation rate or unfolding rate (kotr) in the absence of force.
- both GAIN-WT and GAIN(4Cys) were immobilized on different areas of the same surface and a single cantilever was used to probe each spot using constant speed pulling at 3200 nm s-1 . After 500 approach-retraction cycles at one area, the surface was moved to the other area. This cycle was done once for each GAIN-WT and GAIN(4Cys). The total number of force-extension curves matching the criterion was 69 out of 9,484, and 67 out of 4,986 for GAIN-WT and GAIN(4Cys), respectively.
- each GAIN-WT and noGAIN was probed with the same cantilever both in non-reducing condition (0 mM DTT in 1x PBS) and reducing condition (100 mM DTT in 1x PBS).
- SMFS data were recorded in 1x PBS first at room temperature with constant pulling speed of 800 nm s-1 , and then buffer was changed to 100 mM DTT in 1x PBS and SMFS data were recorded with the same pulling speed.
- the total number of force-extension curves matching the criterion was 101 out of 1 ,566, and 116 out of 26,097 for GAIN-WT and GAIN(4Cys), respectively.
- HEK293T Human embryonic kidney cells
- ATCC CRL-3216 Human embryonic kidney cells
- DMEM 10% heat-inactivated FBS, 1% PenStrep
- Cells were passaged every 2-3 days (up to passage 30) by lifting the cells with 0.05% trypsin in DPBS.
- Transfection - Cells were transiently (forward or reverse) transfected with plasmid constructs using the GeneJuice transfection reagent following the supplier’s recommended protocol (see specific conditions for each assay below).
- forward transfection 25x10 3 cells were seeded per well of a 96- well plate and the transfection was carried out 24h later.
- reverse transfection 50x10 3 cells were seeded per well of a 96-well plate and directly transfected.
- ELISA - 25x10 3 HEK293T cells were seeded per well of a 96-well plate pre-coated with 0.1 mg. ml -1 PDL. 24 hours post-seeding, cells were transfected as described previously with 100 ng of the indicated pcDNA-GPR56 construct and 25 ng of dualLUC-SRE plasmid. 24 hours post- transfection, cells were washed with DPBS, fixed with cold 4% PFA for 10 minutes at 4°C and washed again DPBS after PFA removal. This was followed by three consecutive 60-minute incubations with 2% BSA in DPBS, X pg.
- Dual-Gio Luciferase SRE assays 25x10 3 HEK293T cells were seeded per well of a 96-well plate precoated with 0.1 mg. ml -1 PDL. 24 hours post-seeding, cells were transfected as described previously with 100 ng of the indicated pcDNA-GPR56 construct and 25 ng of dualLUC-SRE plasmid. 8-10 hours post-transfection, the culture medium was replaced with FBS-free DMEM for over-night serum starvation. The next morning, cells were induced with DMSO or 50 pM Stachel peptide in two doses at to and t+2h of the 4-hour induction phase. At the end of the induction period, 60 pl of the 100 pl total well volume was removed and the Dual-Glo Luciferase SRE assay was carried out following the manufacturer’s instructions apart from a working volume of 40 pl instead of 100 pl.
- Vibration assays with coated ligands - 96-well plates were coated overnight at 4°C with 40 pl PDL, collagen or transglutaminase at the indicated densities.
- 50x10 3 cells were reverse transfected onto the pre-coated wells as described previously with 100 ng of the indicated pcDNA- GPR56 construct and 25 ng of dualLUC-SRE plasmid. 6-8 hours post-transfection, the culture medium was replaced with FBS-free DMEM for over-night serum starvation.
- cells were placed on a Titramax 100 vibration shaker (Heidolph Instruments) and vibrated at the indicated frequencies and time periods. 4 hours after the beginning of the induction phase, the Dual-Glo Luciferase SRE assay was carried out as described previously.
- the Coupled Moves algorithm from the Rosetta software was used with the backrub movers protocol (7). A total of 100 independent structures was generated for every pair and top 10% lowest energy structures were conserved. All structures were then minimized and relaxed to access lowest energy potential through the relax protocols included in the Rosetta package. Distances between the sulfide groups of added cysteines were computed and averaged across the lowest energy structures. To determine which pairs of mutations could form a disulfide, a distance cutoff of 2.15 A between the 2 cysteines was used to conserve or exclude mutants.
- the structure of the GAIN_SS-h2l10-h1 h2 variant was resolved at 1.92 A resolution and back-mutated using the above protocol to recover the individual disulfide variants as well as the WT structure.
- the system was solvated with a box size of 75 A along X and Y dimensions and varying Z depending on the simulation with 0.1 M of Na + and Cl ⁇ ions.
- the total system size was approximatively 130 k atoms.
- the Wernet-Nilsson algorithm was used (8). Identified H-bonds across the simulation were 2D- projected in a plane perpendicular to the Stachel axis defined by its N-term and C-term residues. To highlight the H-bonds rearrangements occurring during the pulling, the 2D projections were compared between the native unloaded structure and the loaded structure corresponding to the highest mechanical resistance point observed in the SMD simulations. Projections were averaged across replicates and plotted.
- the human GAIN domain sequence used corresponds to native residues N171 to S391.
- C177 found in the native sequence and forming a disulfide bridge with C121 of the PLL domain, was mutated to a serine in all constructs (bold).
- Engineered disulfides are indicated as bold and underlined letters.
Landscapes
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Chemical & Material Sciences (AREA)
- Organic Chemistry (AREA)
- Immunology (AREA)
- General Health & Medical Sciences (AREA)
- Zoology (AREA)
- Gastroenterology & Hepatology (AREA)
- Cell Biology (AREA)
- Biochemistry (AREA)
- Biophysics (AREA)
- Toxicology (AREA)
- Genetics & Genomics (AREA)
- Medicinal Chemistry (AREA)
- Molecular Biology (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Endocrinology (AREA)
- Peptides Or Proteins (AREA)
Abstract
The present invention relates to variants of the GAIN domain (G-protein-coupled receptor autoproteolysis-inducing domain) of an adhesion G protein-coupled receptor (ADGRG), such as a GAIN domain variant comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG1 (Adhesion G Protein-Coupled Receptor G1) of SEQ ID NO: 1, wherein (i) one of amino acid positions (3) to (5), preferably amino acid position (4) and one of amino acid positions (47) to (49), preferably amino acid position (48) of SEQ ID NO: 1 are replaced by cysteines, (ii) one of amino acid positions (35) to (37), preferably amino acid position (36) and one of amino acid positions 180 to (182), preferably amino acid position (181) of SEQ ID NO: 1 are replaced by cysteines, (iii) one of amino acid positions (19) to (21), preferably amino acid position 20 and one of amino acid positions (124) to (126), preferably amino acid position (125) of SEQ ID NO: 1 are replaced by cysteines, (iv) one of amino acid positions (56) to (58), preferably amino acid position (57) and one of amino acid positions (65) to (67), preferably amino acid position (66) of SEQ ID NO: 1 are replaced by cysteines, (v) one of amino acid positions (16) to (18), preferably amino acid position (17) and one of amino acid positions (60) to (62), preferably amino acid position (61) of SEQ ID NO: 1 are replaced by cysteines, (vi) the amino acid position at position (40) is replaced by glutamic acid, the amino acid position at position (43) is replaced by leucine and the amino acid position at position (181) is replaced by arginine, (vii) the amino acid position at position (40) is replaced by glutamic acid, the amino acid position at position (43) is replaced by leucine, the amino acid position at position (135) is replaced by arginine and the amino acid position at position (181) is replaced by arginine, (viii) the amino acid position at position (40) is replaced by glutamic acid, the amino acid position at position (43) is replaced by leucine, the amino acid position at position (135) is replaced by arginine, the amino acid position at position (146) is replaced by tyrosine, and the amino acid position at position (181) is replaced by arginine, (ix) the amino acid position at position (33) is replaced by tryptophane, the amino acid position at position (40) is replaced by glutamic acid, the amino acid position at position (43) is replaced by leucine, the amino acid position at position (135) is replaced by arginine, the amino acid position at position (146) is replaced by tyrosine, the amino acid position at position (178) is replaced by tyrosine, the amino acid position at position (181) is replaced by arginine, the amino acid position at position (184) is replaced by glycine, the amino acid position at position (186) is replaced by glutamine, and the amino acid position at position (187) is replaced by glycine, (x) the amino acid position at position (174) is replaced by isoleucine, (xi) the amino acid position at position (218) is replaced by alanine, or (xii) at least two of (i) to (xi) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii), (iv) and/or (v), or the amino acid replacement(s) as defined in (A), item (vi), (vii), (viii), (ix), (x) and/or (xi), or the cysteines and/or the amino acid replacement(s) as defined in (A), item (xii) are retained.
Description
Variants of the GAIN domain of adhesion G protein-coupled receptors
The present invention relates to variants of the GAIN domain (G-protein-coupled receptor autoproteolysis-inducing domain) of an adhesion G protein-coupled receptor (ADGRG), such as a GAIN domain variant comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG1 (Adhesion G Protein-Coupled Receptor G1) of SEQ ID NO: 1 , wherein (i) one of amino acid positions 3 to 5, preferably amino acid position 4 and one of amino acid positions 47 to 49, preferably amino acid position 48 of SEQ ID NO: 1 are replaced by cysteines, (ii) one of amino acid positions 35 to 37, preferably amino acid position 36 and one of amino acid positions 180 to 182, preferably amino acid position 181 of SEQ ID NO: 1 are replaced by cysteines, (iii) one of amino acid positions 19 to 21 , preferably amino acid position 20 and one of amino acid positions 124 to 126, preferably amino acid position 125 of SEQ ID NO: 1 are replaced by cysteines, (iv) one of amino acid positions 56 to 58, preferably amino acid position 57 and one of amino acid positions 65 to 67, preferably amino acid position 66 of SEQ ID NO: 1 are replaced by cysteines, (v) one of amino acid positions 16 to 18, preferably amino acid position 17 and one of amino acid positions 60 to 62, preferably amino acid position 61 of SEQ ID NO: 1 are replaced by cysteines, (vi) the amino acid position at position 40 is replaced by glutamic acid, the amino acid position at position 43 is replaced by leucine and the amino acid position at position 181 is replaced by arginine, (vii) the amino acid position at position 40 is replaced by glutamic acid, the amino acid position at position 43 is replaced by leucine, the amino acid position at position 135 is replaced by arginine and the amino acid position at position 181 is replaced by arginine, (viii) the amino acid position at position 40 is replaced by glutamic acid, the amino acid position at position 43 is replaced by leucine, the amino acid position at position 135 is replaced by arginine, the amino acid position at position 146is replaced by tyrosine, and the amino acid position at position 181 is replaced by arginine, (ix) the amino acid position at position 33 is replaced by tryptophane, the amino acid position at position 40 is replaced by glutamic acid, the amino acid position at position 43 is replaced by leucine, the amino acid position at position 135 is replaced by arginine, the amino acid position at position 146 is replaced by tyrosine, the amino acid position at position 178 is replaced by tyrosine, the amino acid position at position 181 is replaced by arginine, the amino acid position at position 184 is replaced by glycine, the amino acid position at position 186 is replaced by glutamine, and the amino acid position at position 187 is replaced by glycine, (x) the amino acid position at position 174 is replaced by isoleucine, (xi) the amino acid position at position 218 is replaced by alanine, or (xii) at least two of (i) to (xi) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii), (iv) and/or (v), or the amino acid replacement(s) as defined in (A), item (vi), (vii), (viii), (ix), (x) and/or (xi), or the cysteines and/or the amino acid replacement(s) as defined in (A), item (xii) are retained.
In this specification, a number of documents including patent applications and manufacturer’s manuals are cited. The disclosure of these documents, while not considered relevant for the patentability of this invention, is herewith incorporated by reference in its entirety. More specifically, all referenced documents are incorporated by reference to the same extent as if each individual document was specifically and individually indicated to be incorporated by reference.
Cells can sense the mechanical properties of their local environment. Cellular mechanosensing is currently thought to primarily involve integrins, which couple the extracellular matrix (ECM) to the contractile cytoskeleton and form force-bearing connections that require tensile forces to mature 1-3, as well as by mechano-activated transmembrane ion channels45. Both sensing modalities are currently exploited as drug targets and for bioengineering applications67. Surprisingly though, certain classes of GPCRs (G-protein-coupled receptors) have large extracellular domains that can bind to ECM components8-10. Even though they have been described in the literature, little is known if and how these adhesion GPCRs (aGPCRs or AGPCRs) are involved in mechanosensing and whether this could be employed for a variety of engineering applications. A general feature of these aGPCR molecules is the highly conserved GAIN domain located on the extracellular side of the receptor, between the ligand binding domain (LBD) and the 7 transmembrane (7TM) domain. During folding, a peptide agonist called Stachel is buried and auto-proteolytically cleaved inside the folded GAIN structure. Upon exposure to mechanical stress from the LBD, GAIN is mechanically dissociated from Stachel, which then binds the 7TM domain and triggers intracellular signaling.
The modular topology of aGPCRs is an appealing feature for a variety of engineering applications. By contrast, Adhesion Receptors (Ars, such as integrins and cadherins) do not possess catalytic/signaling domains and behave as scaffold for adhesome assembly connecting extra-cellular (EC) mechanical signals with cytoskeleton. Kinases associate indirectly with integrins through peripheral adaptor proteins for signal regulation. Due to the complexity of the systems, reprogramming integrin or cadherin functions has remained very challenging.
Recent lines of biophysical evidence indicate that protein structures undergo specific deformation upon mechanical load which determine their resistance and response to the applied stress. Yet physical properties underlying the mechanical functions of proteins remain poorly understood. Engineering mechanosensing properties implies manipulating such structures at non-equilibrium which represents a real challenge. So far, steered molecular dynamics (SMD) have shown reasonable agreement with experimental single force spectroscopy and are able to provide atomic level explanation of mutational effects on mechanical responses in a few systems. However, no computational method integrating dynamics under mechanical stress and systematic calculation of deformation upon sequence variation has been developed and applied for rationally designing mechanosensing properties.
How protein structures react to and resist to mechanical stress is a fundamental property of proteins that remains poorly exploited. Protein topology and the orientation of chemical bonds with respect to the direction of the mechanical force propagating inside the structure are key determinants of protein mechanical resistance, but these principles have yet to be recapitulated in protein design. Forcebearing non-covalent interactions will deform and break eventually under force, and computational methods to engineer specific networks of chemical interactions are now well established and should be directly applicable to the design of mechano-resistant protein structures and protein-ligand binding interactions.
In addition, GPCRs play a very important role as drug target. The majority of approved pharmaceutics target this receptor family. About one third of the GPCR family is still orphan, which means that either function or activation are unknown. Deciphering the role of orphan GPCRs is of high interest to basic research and of clinical importance as it has the potential to yield even more therapeutic approaches. With 33 members, adhesion GPCRs (aGPCRs) form the second largest group of the GPCR family and the majority of them are orphan. As discussed, the GAIN domain of aGPCRs harbors the GPCR proteolysis site at which the receptor is autoproteolytically cleaved into an N- and a C-terminal fragment which trigger signalling cascades. This knowledge on aGPCRS leads to the hypothesis that they “sense” in particular “mechanosense” via the GAIN domain their environment and contribute to a context-dependent signal transduction.
In view of the above the present invention aims at the provision of a method to reprogram the response of the human GAIN domains of various aGPCRS to mechanical force as well as the provision of such reprogrammed GAIN domains as well as aGPCRs comprising the same. Such methods and GAINS domains are helpful to study context-dependent signal transduction and consequently helpful for applications such as cell signaling studies and drug development (see also Bassilana et al. (2019), Nature Reviews Drug Discovery, 18:869-884).
Hence, the present invention relates to variants of the GAIN domain (G-protein-coupled receptor autoproteolysis-inducing domain) of an adhesion G protein-coupled receptor (ADGRG):
(I) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG1 (Adhesion G Protein-Coupled Receptor G1) of SEQ ID NO: 1 , wherein
(i) one of amino acid positions 3 to 5, preferably amino acid position 4 and one of amino acid positions 47 to 49, preferably amino acid position 48 of SEQ ID NO: 1 are replaced by cysteines,
(ii) one of amino acid positions 35 to 37, preferably amino acid position 36 and one of amino acid positions 180 to 182, preferably amino acid position 181 of SEQ ID NO: 1 are replaced by cysteines,
(iii) one of amino acid positions 19 to 21 , preferably amino acid position 20 and one of amino acid positions 124 to 126, preferably amino acid position 125 of SEQ ID NO: 1 are replaced by cysteines,
(iv) one of amino acid positions 56 to 58, preferably amino acid position 57 and one of amino acid positions 65 to 67, preferably amino acid position 66 of SEQ ID NO: 1 are replaced by cysteines,
(v) one of amino acid positions 16 to 18, preferably amino acid position 17 and one of amino acid positions 60 to 62, preferably amino acid position 61 of SEQ ID NO: 1 are replaced by cysteines,
(vi) the amino acid position at position 40 is replaced by glutamic acid, the amino acid position at position 43 is replaced by leucine and the amino acid position at position 181 is replaced by arginine,
(vii) the amino acid position at position 40 is replaced by glutamic acid, the amino acid position at position 43 is replaced by leucine, the amino acid position at position 135 is replaced by arginine and the amino acid position at position 181 is replaced by arginine,
(viii) the amino acid position at position 40 is replaced by glutamic acid, the amino acid position at position 43 is replaced by leucine, the amino acid position at position 135 is replaced by arginine, the amino acid position at position 146is replaced by tyrosine, and the amino acid position at position 181 is replaced by arginine,
(ix) the amino acid position at position 33 is replaced by tryptophane, the amino acid position at position 40 is replaced by glutamic acid, the amino acid position at position 43 is replaced by leucine, the amino acid position at position 135 is replaced by arginine, the amino acid position at position 146 is replaced by tyrosine, the amino acid position at position 178 is replaced by tyrosine, the amino acid position at position 181 is replaced by arginine, the amino acid position at position 184 is replaced by glycine, the amino acid position at position 186 is replaced by glutamine, and the amino acid position at position 187 is replaced by glycine,
(x) the amino acid position at position 174 is replaced by isoleucine,
(xi) the amino acid position at position 218 is replaced by alanine, or
(xii) at least two of (i) to (xi) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii), (iv) and/or (v), or the amino acid replacement(s) as defined in (A), item (vi), (vii), (viii), (ix), (x) and/or (xi), or the cysteines and/or the amino acid replacement(s) as defined in (A), item (xii) are retained;
(II) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRL4 (Adhesion G Protein-Coupled Receptor L4) of SEQ ID NO: 2, wherein
(i) one of amino acid positions 87 to 89, preferably amino acid position 88 and one of amino acid positions 123 to 125, preferably amino acid position 124 of SEQ ID NO: 2 are replaced by cysteines,
(ii) one of amino acid positions 111 to 113, preferably amino acid position 112 and one of amino acid positions 263 to 265, preferably amino acid position 264 of SEQ ID NO: 2 are replaced by cysteines,
(iii) one of amino acid positions 95 to 97, preferably amino acid position 96 and one of amino acid positions 192 to 194, preferably amino acid position 193 of SEQ ID NO: 2 are replaced by cysteines,
(iv) one of amino acid positions 134 to 136, preferably amino acid position 135 and one of amino acid positions 144 to 146, preferably amino acid position 145 of SEQ ID NO: 2 are replaced by cysteines,
(v) one of amino acid positions 93 to 95, preferably amino acid position 94 and one of amino acid positions 140 to 142, preferably amino acid position 141 of SEQ ID NO: 2 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with
the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(III) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRL3 (Adhesion G Protein-Coupled Receptor L3) of SEQ ID NO: 3, wherein
(i) one of amino acid positions 81 to 83, preferably amino acid position 82 and one of amino acid positions 117 to 119, preferably amino acid position 118 of SEQ ID NO: 3 are replaced by cysteines,
(ii) one of amino acid positions 105 to 107, preferably amino acid position 106 and one of amino acid positions 261 to 263, preferably amino acid position 262 of SEQ ID NO: 3 are replaced by cysteines,
(iii) one of amino acid positions 89 to 91 , preferably amino acid position 90 and one of amino acid positions 188 to 190, preferably amino acid position 189 of SEQ ID NO: 3 are replaced by cysteines,
(iv) one of amino acid positions 128 to 130, preferably amino acid position 129 and one of amino acid positions 138 to 140, preferably amino acid position 139 of SEQ ID NO: 3 are replaced by cysteines,
(v) one of amino acid positions 87 to 89, preferably amino acid position 88 and one of amino acid positions 134 to 136, preferably amino acid position 135 of SEQ ID NO: 3 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(IV) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRL2 (Adhesion G Protein-Coupled Receptor L2) of SEQ ID NO: 4, wherein
(i) one of amino acid positions 94 to 96, preferably amino acid position 95 and one of amino acid positions 130 to 132, preferably amino acid position 131 of SEQ ID NO: 4 are replaced by cysteines,
(ii) one of amino acid positions 118 to 120, preferably amino acid position 119 and one of amino acid positions 272 to 274, preferably amino acid position 273 of SEQ ID NO: 4 are replaced by cysteines,
(iii) one of amino acid positions 102 to 104, preferably amino acid position 103 and one of amino acid positions 202 to 204, preferably amino acid position 203 of SEQ ID NO: 4 are replaced by cysteines,
(iv) one of amino acid positions 141 to 143, preferably amino acid position 144 and one of amino acid positions 151 to 153, preferably amino acid position 152 of SEQ ID NO: 4 are replaced by cysteines,
(v) one of amino acid positions 99 to 101 , preferably amino acid position 100 and one of amino acid positions 147 to 149, preferably amino acid position 148 of SEQ ID NO: 4 are replaced by cysteines, or
(vi) at least two or all three of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(V) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRL1 (Adhesion G Protein-Coupled Receptor L1) of SEQ ID NO: 5, wherein
(i) one of amino acid positions 94 to 96, preferably amino acid position 95 and one of amino acid positions 130 to 132, preferably amino acid position 131 of SEQ ID NO: 5 are replaced by cysteines,
(ii) one of amino acid positions 118 to 120, preferably amino acid position 119 and one of amino acid positions 275 to 277, preferably amino acid position 276 of SEQ ID NO: 5 are replaced by cysteines,
(iii) one of amino acid positions 102 to 104, preferably amino acid position 103 and one of amino acid positions 202 to 204, preferably amino acid position 203 of SEQ ID NO: 5 are replaced by cysteines,
(iv) one of amino acid positions 141 to 143, preferably amino acid position 144 and one of amino acid positions 151 to 153, preferably amino acid position 152 of SEQ ID NO: 5 are replaced by cysteines,
(v) one of amino acid positions 100 to 102, preferably amino acid position 101 and one of amino acid positions 147 to 149, preferably amino acid position 148 of SEQ ID NO: 5 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(VI) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG7 (Adhesion G Protein-Coupled Receptor G7) of SEQ ID NO: 6, wherein
(i) one of amino acid positions 67 to 69, preferably amino acid position 68 and one of amino acid positions 102 to 104, preferably amino acid position 103 of SEQ ID NO: 6 are replaced by cysteines,
(ii) one of amino acid positions 90 to 92, preferably amino acid position 91 and one of amino acid positions 246 to 248, preferably amino acid position 247 of SEQ ID NO: 6 are replaced by cysteines,
(iii) one of amino acid positions 74 to 76, preferably amino acid position 75 and one of amino acid positions 186 to 188, preferably amino acid position 187 of SEQ ID NO: 6 are replaced by cysteines, (vi) one of amino acid positions 64 to 66, preferably amino acid position 65 and one of amino acid positions 111 to 113, preferably amino acid position 112 of SEQ ID NO: 6 are replaced by cysteines, or
(v) at least two of (i) to (iv apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (v) are retained;
(VII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG6 (Adhesion G Protein-Coupled Receptor G6) of SEQ ID NO: 7, wherein
(i) one of amino acid positions 35 to 37, preferably amino acid position 36 and one of amino acid positions 84 to 86, preferably amino acid position 85 of SEQ ID NO: 7 are replaced by cysteines,
(ii) one of amino acid positions 73 to 75, preferably amino acid position 74 and one of amino acid positions 229 to 231 , preferably amino acid position 230 of SEQ ID NO: 7 are replaced by cysteines,
(iii) one of amino acid positions 56 to 58, preferably amino acid position 57 and one of amino acid positions 172 to 174, preferably amino acid position 173 of SEQ ID NO: 7 are replaced by cysteines,
(iv) one of amino acid positions 34 to 36, preferably amino acid position 35 and one of amino acid positions 103 to 105, preferably amino acid position 104 of SEQ ID NO: 7 are replaced by cysteines,
(v) one of amino acid positions 52 to 54, preferably amino acid position 53 and one of amino acid positions 98 to 100, preferably amino acid position 99 of SEQ ID NO: 7 are replaced by cysteines, or
(vi) at least two or all three of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(VIII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG5
(Adhesion G Protein-Coupled Receptor G5) of SEQ ID NO: 8, wherein
(i) one of amino acid positions 45 to 47, preferably amino acid position 46 and one of amino acid positions 193 to 195, preferably amino acid position 194 of SEQ ID NO: 8 are replaced by cysteines,
(ii) one of amino acid positions 30 to 32, preferably amino acid position 31 and one of amino acid positions 137 to 139, preferably amino acid position 138 of SEQ ID NO: 8 are replaced by cysteines,
(iii) one of amino acid positions 66 to 68, preferably amino acid position 67 and one of amino acid positions 75 to 77, preferably amino acid position 76 of SEQ ID NO: 8 are replaced by cysteines,
(iv) one of amino acid positions 27 to 29, preferably amino acid position 28 and one of amino acid positions 70 to 72, preferably amino acid position 71 of SEQ ID NO: 8 are replaced by cysteines, or
(v) at least two of (i) to (iv) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (v) are retained;
(IX) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG4 (Adhesion G Protein-Coupled Receptor G4) of SEQ ID NO: 9, wherein
(i) one of amino acid positions 251 to 253, preferably amino acid position 252 and one of amino acid positions 302 to 304, preferably amino acid position 303 of SEQ ID NO: 9 are replaced by cysteines,
(ii) one of amino acid positions 290 to 292, preferably amino acid position 291 and one of amino acid positions 442 to 444, preferably amino acid position 443 of SEQ ID NO: 9 are replaced by cysteines,
(iii) one of amino acid positions 274 to 276, preferably amino acid position 275 and one of amino acid positions 385 to 387, preferably amino acid position 386 of SEQ ID NO: 9 are replaced by cysteines,
(iv) one of amino acid positions 311 to 313, preferably amino acid position 312 and one of amino acid positions 320 to 322, preferably amino acid position 321 of SEQ ID NO: 9 are replaced by cysteines,
(v) one of amino acid positions 268 to 270, preferably amino acid position 269 and one of amino acid positions 315 to 317, preferably amino acid position 316 of SEQ ID NO: 9 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(X) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG3 (Adhesion G Protein-Coupled Receptor G3) of SEQ ID NO: 10, wherein
(i) one of amino acid positions 16 to 19, preferably amino acid position 17 and one of amino acid positions 45 to 47, preferably amino acid position 46 of SEQ ID NO: 10 are replaced by cysteines,
(ii) one of amino acid positions 32 to 34, preferably amino acid position 33 and one of amino acid positions 116 to 118, preferably amino acid position 117 of SEQ ID NO: 10 are replaced by cysteines,
(iii) one of amino acid positions 55 to 57, preferably amino acid position 56 and one of amino acid positions 63 to 65, preferably amino acid position 64 of SEQ ID NO: 10 are replaced by cysteines,
(iv) one of amino acid positions 29 to 31 , preferably amino acid position 30 and one of amino acid positions 59 to 61 , preferably amino acid position 60 of SEQ ID NO: 10 are replaced by cysteines, or
(v) at least two of (i) to (iv) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with
the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (v) are retained;
(XI) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG2 (Adhesion G Protein-Coupled Receptor G2) of SEQ ID NO: 11 , wherein
(i) one of amino acid positions 12 to 14, preferably amino acid position 13 and one of amino acid positions 63 to 65, preferably amino acid position 64 of SEQ ID NO: 11 are replaced by cysteines,
(ii) one of amino acid positions 51 to 53, preferably amino acid position 52 and one of amino acid positions 205 to 207, preferably amino acid position 206 of SEQ ID NO: 11 are replaced by cysteines,
(iii) one of amino acid positions 35 to 37, preferably amino acid position 36 and one of amino acid positions 147 to 149, preferably amino acid position 148 of SEQ ID NO: 11 are replaced by cysteines,
(iv) one of amino acid positions 72 to 74, preferably amino acid position 73 and one of amino acid positions 81 to 83, preferably amino acid position 82 of SEQ ID NO: 11 are replaced by cysteines,
(v) one of amino acid positions 32 to 34, preferably amino acid position 33 and one of amino acid positions 76 to 78, preferably amino acid position 77 of SEQ ID NO: 11 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(XII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF5 (Adhesion G Protein-Coupled Receptor F5) of SEQ ID NO: 12, wherein
(i) one of amino acid positions 82 to 84, preferably amino acid position 83 and one of amino acid positions 117 to 119, preferably amino acid position 118 of SEQ ID NO: 12 are replaced by cysteines,
(ii) one of amino acid positions 105 to 107, preferably amino acid position 106 and one of amino acid positions 240 to 242, preferably amino acid position 241 of SEQ ID NO: 12 are replaced by cysteines,
(iii) one of amino acid positions 90 to 92, preferably amino acid position 91 and one of amino acid positions 186 to 188, preferably amino acid position 187 of SEQ ID NO: 12 are replaced by cysteines,
(iv) one of amino acid positions 127 to 129, preferably amino acid position 128 and one of amino acid positions 136 to 138, preferably amino acid position 137 of SEQ ID NO: 12 are replaced by cysteines,
(v) one of amino acid positions 88 to 90, preferably amino acid position 89 and one of amino acid positions 131 to 133, preferably amino acid position 132 of SEQ ID NO: 12 are replaced by cysteines, or (vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (v) are retained;
(XIII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF4 (Adhesion G Protein-Coupled Receptor F4) of SEQ ID NO: 13, wherein
(i) one of amino acid positions 53 to 55, preferably amino acid position 54 and one of amino acid positions 86 to 88, preferably amino acid position 87 of SEQ ID NO: 13 are replaced by cysteines,
(ii) one of amino acid positions 74 to 76, preferably amino acid position 75 and one of amino acid positions 216 to 218, preferably amino acid position 217 of SEQ ID NO: 13 are replaced by cysteines,
(iii) one of amino acid positions 61 to 63, preferably amino acid position 62 and one of amino acid positions 159 to 161 , preferably amino acid position 160 of SEQ ID NO: 13 are replaced by cysteines,
(iv) one of amino acid positions 96 to 98, preferably amino acid position 97 and one of amino acid
positions 105 to 107, preferably amino acid position 106 of SEQ ID NO: 13 are replaced by cysteines,
(v) one of amino acid positions 59 to 61 , preferably amino acid position 60 and one of amino acid positions 100 to 102, preferably amino acid position 102 of SEQ ID NO: 13 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(XIV) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF3 (Adhesion G Protein-Coupled Receptor F3) of SEQ ID NO: 14, wherein
(i) one of amino acid positions 81 to 83, preferably amino acid position 82 and one of amino acid positions 116 to 118, preferably amino acid position 117 of SEQ ID NO: 14 are replaced by cysteines,
(ii) one of amino acid positions 104 to 106, preferably amino acid position 105 and one of amino acid positions 243 to 245, preferably amino acid position 244 of SEQ ID NO: 14 are replaced by cysteines,
(iii) one of amino acid positions 89 to 91 , preferably amino acid position 90 and one of amino acid positions 184 to 186, preferably amino acid position 185 of SEQ ID NO: 14 are replaced by cysteines,
(iv) one of amino acid positions 126 to 128, preferably amino acid position 127 and one of amino acid positions 135 to 137, preferably amino acid position 136 of SEQ ID NO: 14 are replaced by cysteines,
(v) one of amino acid positions 87 to 89, preferably amino acid position 88 and one of amino acid positions 129 to 131 , preferably amino acid position 130 of SEQ ID NO: 14 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(XV) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF2 (Adhesion G Protein-Coupled Receptor F2) of SEQ ID NO: 15, wherein
(i) one of amino acid positions 73 to 75, preferably amino acid position 74 and one of amino acid positions 106 to 108, preferably amino acid position 107 of SEQ ID NO: 15 are replaced by cysteines,
(ii) one of amino acid positions 94 to 96, preferably amino acid position 95 and one of amino acid positions 230 to 232, preferably amino acid position 231 of SEQ ID NO: 15 are replaced by cysteines,
(iii) one of amino acid positions 81 to 83, preferably amino acid position 82 and one of amino acid positions 177 to 179, preferably amino acid position 178 of SEQ ID NO: 15 are replaced by cysteines,
(iv) one of amino acid positions 116 to 118, preferably amino acid position 117 and one of amino acid positions 124 to 126, preferably amino acid position 125 of SEQ ID NO: 15 are replaced by cysteines,
(v) one of amino acid positions 79 to 81 , preferably amino acid position 80 and one of amino acid positions 119 to 121 , preferably amino acid position 120 of SEQ ID NO: 15 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(XVI) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF1
(Adhesion G Protein-Coupled Receptor F1) of SEQ ID NO: 16, wherein
(i) one of amino acid positions 83 to 85, preferably amino acid position 84 and one of amino acid positions 118 to 120, preferably amino acid position 119 of SEQ ID NO: 16 are replaced by cysteines,
(ii) one of amino acid positions 106 to 108, preferably amino acid position 107 and one of amino acid positions 243 to 245, preferably amino acid position 244 of SEQ ID NO: 16 are replaced by cysteines,
(iii) one of amino acid positions 91 to 93, preferably amino acid position 92 and one of amino acid positions 190 to 192, preferably amino acid position 191 of SEQ ID NO: 16 are replaced by cysteines,
(iv) one of amino acid positions 128 to 130, preferably amino acid position 129 and one of amino acid positions 136 to 138, preferably amino acid position 137 of SEQ ID NO: 16 are replaced by cysteines,
(v) one of amino acid positions 89 to 91 , preferably amino acid position 90 and one of amino acid positions 131 to 133, preferably amino acid position 132 of SEQ ID NO: 16 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained; (XVII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE5 (Adhesion G Protein-Coupled Receptor E5) of SEQ ID NO: 17, wherein
(i) one of amino acid positions 72 to 74, preferably amino acid position 73 and one of amino acid positions 244 to 246, preferably amino acid position 245 of SEQ ID NO: 17 are replaced by cysteines,
(ii) one of amino acid positions 96 to 98, preferably amino acid position 97 and one of amino acid positions 109 to 111 , preferably amino acid position 110 of SEQ ID NO: 17 are replaced by cysteines, or
(iii) (i) and (ii) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), (i), (ii) or (iii) are retained;
(XVIII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE3 (Adhesion G Protein-Coupled Receptor E3) of SEQ ID NO: 18, wherein one of amino acid positions 72 to 74, preferably amino acid position 73 and one of amino acid positions 83 to 85, preferably amino acid position 84 of SEQ ID NO: 18 are replaced by cysteines; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A) are retained;
(XIX) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE3 (Adhesion G Protein-Coupled Receptor E3) of SEQ ID NO: 19, wherein one of amino acid positions 73 to 75, preferably amino acid position 74 and one of amino acid positions 186 to 188, preferably amino acid position 187 of SEQ ID NO: 19 are replaced by cysteines; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A) are retained;
(XX) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE2 (Adhesion G Protein-Coupled Receptor E2) of SEQ ID NO: 20, wherein
one of amino acid positions 72 to 74, preferably amino acid position 73 and one of amino acid positions 228 to 230, preferably amino acid position 229 of SEQ ID NO: 20 are replaced by cysteines,
(ii) one of amino acid positions 96 to 98, preferably amino acid position 97 and one of amino acid positions 109 to 111 , preferably amino acid position 110 of SEQ ID NO: 20 are replaced by cysteines, or
(iii) (i) and (ii) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), (i), (ii) or (iii) are retained;
(XXI) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE1 (Adhesion G Protein-Coupled Receptor E1) of SEQ ID NO: 21 , wherein
(i) one of amino acid positions 66 to 68, preferably amino acid position 67 and one of amino acid positions 103 to 105, preferably amino acid position 104 of SEQ ID NO: 21 are replaced by cysteines,
(ii) one of amino acid positions 91 to 93, preferably amino acid position 92 and one of amino acid positions 243 to 245, preferably amino acid position 244 of SEQ ID NO: 21 are replaced by cysteines,
(iii) one of amino acid positions 75 to 77, preferably amino acid position 76 and one of amino acid positions 173 to 175, preferably amino acid position 174 of SEQ ID NO: 21 are replaced by cysteines,
(iv) one of amino acid positions 111 to 113, preferably amino acid position 112 and one of amino acid positions 124 to 126, preferably amino acid position 125 of SEQ ID NO: 21 are replaced by cysteines,
(v) one of amino acid positions 73 to 75, preferably amino acid position 74 and one of amino acid positions 119 to 121 , preferably amino acid position 120 of SEQ ID NO: 21 are replaced by cysteines or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained; (XXII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRD2 (Adhesion G Protein-Coupled Receptor D2) of SEQ ID NO: 22, wherein
(i) one of amino acid positions 67 to 69, preferably amino acid position 68 and one of amino acid positions 110 to 112, preferably amino acid position 111 of SEQ ID NO: 22 are replaced by cysteines,
(ii) one of amino acid positions 98 to 100, preferably amino acid position 99 and one of amino acid positions 235 to 237, preferably amino acid position 236 of SEQ ID NO: 22 are replaced by cysteines,
(iii) one of amino acid positions 121 to 123, preferably amino acid position 121 and one of amino acid positions 131 to 133, preferably amino acid position 132 of SEQ ID NO: 22 are replaced by cysteines,
(iv) one of amino acid positions 80 to 82, preferably amino acid position 81 and one of amino acid positions 127 to 129, preferably amino acid position 128 of SEQ ID NO: 22 are replaced by cysteines or
(v) at least two of (i) to (iv) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (v) are retained; (XXIII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRD1 (Adhesion G Protein-Coupled Receptor D1) of SEQ ID NO: 23, wherein
(i) one of amino acid positions 36 to 38, preferably amino acid position 37 and one of amino acid positions 92 to 94, preferably amino acid position 93 of SEQ ID NO: 23 are replaced by cysteines,
(ii) one of amino acid positions 80 to 82, preferably amino acid position 81 and one of amino acid positions 247 to 249, preferably amino acid position 248 of SEQ ID NO: 23 are replaced by cysteines,
(iii) one of amino acid positions 66 to 68, preferably amino acid position 67 and one of amino acid positions 162 to 164, preferably amino acid position 163 of SEQ ID NO: 23 are replaced by cysteines,
(iv) one of amino acid positions 104 to 106, preferably amino acid position 105 and one of amino acid positions 116 to 118, preferably amino acid position 117 of SEQ ID NO: 23 are replaced by cysteines,
(v) one of amino acid positions 64 to 66, preferably amino acid position 65 and one of amino acid positions 111 to 113, preferably amino acid position 112 of SEQ ID NO: 23 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained; (XXIV) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRB3 (Adhesion G Protein-Coupled Receptor B3) of SEQ ID NO: 24, wherein
(i) one of amino acid positions 87 to 89, preferably amino acid position 88 and one of amino acid positions 121 to 123, preferably amino acid position 122 of SEQ ID NO: 24 are replaced by cysteines,
(ii) one of amino acid positions 109 to 111 , preferably amino acid position 110 and one of amino acid positions 261 to 263, preferably amino acid position 262 of SEQ ID NO: 24 are replaced by cysteines,
(iii) one of amino acid positions 95 to 97, preferably amino acid position 96 and one of amino acid positions 206 to 208, preferably amino acid position 207 of SEQ ID NO: 24 are replaced by cysteines,
(iv) one of amino acid positions 133 to 135, preferably amino acid position 134 and one of amino acid positions 142 to 144, preferably amino acid position 143 of SEQ ID NO: 24 are replaced by cysteines,
(v) one of amino acid positions 93 to 95, preferably amino acid position 94 and one of amino acid positions 138 to 140, preferably amino acid position 139 of SEQ ID NO: 24 are replaced by cysteines or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained; (XXV) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRB2 (Adhesion G Protein-Coupled Receptor B2) of SEQ ID NO: 25, wherein
(i) one of amino acid positions 86 to 88, preferably amino acid position 87 and one of amino acid positions 120 to 122, preferably amino acid position 121 of SEQ ID NO: 25 are replaced by cysteines,
(ii) one of amino acid positions 108 to 110, preferably amino acid position 109 and one of amino acid positions 295 to 297, preferably amino acid position 296 of SEQ ID NO: 25 are replaced by cysteines,
(iii) one of amino acid positions 94 to 96, preferably amino acid position 95 and one of amino acid positions 239 to 241 , preferably amino acid position 240 of SEQ ID NO: 25 are replaced by cysteines,
(iv) one of amino acid positions 132 to 134, preferably amino acid position 133 and one of amino acid positions 141 to 143, preferably amino acid position 142 of SEQ ID NO: 25 are replaced by cysteines,
(v) one of amino acid positions 92 to 94, preferably amino acid position 93 and one of amino acid positions 137 to 139, preferably amino acid position 138 of SEQ ID NO: 25 are replaced by cysteines or
(vi) at least two or all three of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained; (XXVI) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRB1 (Adhesion G Protein-Coupled Receptor B1) of SEQ ID NO: 26, wherein
(i) one of amino acid positions 81 to 83, preferably amino acid position 82 and one of amino acid positions 115 to 117, preferably amino acid position 116 of SEQ ID NO: 26 are replaced by cysteines,
(ii) one of amino acid positions 103 to 105, preferably amino acid position 104 and one of amino acid positions 253 to 255, preferably amino acid position 254 of SEQ ID NO: 26 are replaced by cysteines,
(iii) one of amino acid positions 89 to 91 , preferably amino acid position 90 and one of amino acid positions 198 to 200, preferably amino acid position 199 of SEQ ID NO: 26 are replaced by cysteines,
(iv) one of amino acid positions 128 to 130, preferably amino acid position 129 and one of amino acid positions 136 to 138, preferably amino acid position 137 of SEQ ID NO: 26 are replaced by cysteines,
(v) one of amino acid positions 87 to 89, preferably amino acid position 88 and one of amino acid positions 132 to 134, preferably amino acid position 133 of SEQ ID NO: 26 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained; (XXVII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRA3 (Adhesion G Protein-Coupled Receptor A3) of SEQ ID NO: 27, wherein
(i) one of amino acid positions 57 to 59, preferably amino acid position 58 and one of amino acid positions 106 to 108, preferably amino acid position 107 of SEQ ID NO: 27 are replaced by cysteines,
(ii) one of amino acid positions 94 to 96, preferably amino acid position 95 and one of amino acid positions 288 to 290, preferably amino acid position 289 of SEQ ID NO: 27 are replaced by cysteines,
(iii) one of amino acid positions 81 to 83, preferably amino acid position 82 and one of amino acid positions 225 to 227, preferably amino acid position 226 of SEQ ID NO: 27 are replaced by cysteines,
(iv) one of amino acid positions 115 to 117, preferably amino acid position 116 and one of amino acid positions 125 to 127, preferably amino acid position 126 of SEQ ID NO: 27 are replaced by cysteines,
(v) one of amino acid positions 73 to 75, preferably amino acid position 74 and one of amino acid positions 120 to 122, preferably amino acid position 121 of SEQ ID NO: 27 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained; (XXVIII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRA2 (Adhesion G Protein-Coupled Receptor A2) of SEQ ID NO: 28, wherein
(i) one of amino acid positions 57 to 59, preferably amino acid position 58 and one of amino acid positions 104 to 106, preferably amino acid position 105 of SEQ ID NO: 28 are replaced by cysteines,
(ii) one of amino acid positions 92 to 94, preferably amino acid position 93 and one of amino acid positions 294 to 296, preferably amino acid position 295 of SEQ ID NO: 28 are replaced by cysteines,
(iii) one of amino acid positions 79 to 81 , preferably amino acid position 80 and one of amino acid positions 229 to 231 , preferably amino acid position 230 of SEQ ID NO: 28 are replaced by cysteines,
(iv) one of amino acid positions 113 to 115, preferably amino acid position 114 and one of amino acid positions 123 to 125, preferably amino acid position 124 of SEQ ID NO: 28 are replaced by cysteines,
(v) one of amino acid positions 71 to 73, preferably amino acid position 72 and one of amino acid positions 118 to 120, preferably amino acid position 129 of SEQ ID NO: 28 are replaced by cysteines, or
(vi) at least two or all three of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), items (i) to (vi) are retained; or
(XXIX) (A) comprising or consisting of the amino acid sequence of GAIN domain of ADGRV1 (Adhesion G Protein-Coupled Receptor V1) of SEQ ID NO: 29, wherein
(i) one of amino acid positions 171 to 173, preferably amino acid position 172 and one of amino acid positions 229 to 231 , preferably amino acid position 230 of SEQ ID NO: 29 are replaced by cysteines,
(ii) one of amino acid positions 127 to 219, preferably amino acid position 218 and one of amino acid positions 436 to 438, preferably amino acid position 437 of SEQ ID NO: 29 are replaced by cysteines,
(iii) one of amino acid positions 201 to 203, preferably amino acid position 202 and one of amino acid positions 377 to 379, preferably amino acid position 378 of SEQ ID NO: 29 are replaced by cysteines,
(iv) one of amino acid positions 199 to 201 , preferably amino acid position 200 and one of amino acid positions 239 to 241 , preferably amino acid position 240 of SEQ ID NO: 29 are replaced by cysteines, or
(v) at least two of (i) to (iv) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), items (i) to (v) are retained.
According to a preferred embodiment of the first aspect the present invention relates to a variant of the GAIN domain (G-protein-coupled receptor autoproteolysis-inducing domain) of an adhesion G protein- coupled receptor (ADGRG)
(I) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG1 (Adhesion G Protein-Coupled Receptor G1) of SEQ ID NO: 1 , wherein (i) one of amino acid positions 3 to 5, preferably amino acid position 4 and one of amino acid positions 47 to 49, preferably amino acid position 48 of SEQ ID NO: 1 are replaced by cysteines, (ii) one of amino acid positions 35 to 37, preferably amino acid position 36 and one of amino acid positions 180 to 182, preferably amino acid position 181 of SEQ ID NO: 1 are replaced by cysteines, (iii) one of amino acid positions 19 to 21 , preferably amino acid position 20 and one of amino acid positions 124 to 126, preferably amino acid position 125 of SEQ ID NO: 1 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained;
(II) (A)comprising or consisting of the amino acid sequence of the GAIN domain of ADGRL4 (Adhesion G Protein-Coupled Receptor L4) of SEQ ID NO: 2, wherein(i) one of amino acid positions 87 to 89, preferably amino acid position 88 and one of amino acid positions 123 to 125, preferably amino acid position 124 of SEQ ID NO: 2 are replaced by cysteines, (ii) one of amino acid positions 111 to 113, preferably amino acid position 112 and one of amino acid positions 263 to 265, preferably amino acid position 264 of SEQ ID NO: 2 are replaced by cysteines, (iii) one of amino acid positions 95 to 97, preferably amino acid position 96 and one of amino acid positions 192 to 194, preferably amino acid position 193 of SEQ ID NO: 2 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained;
(III)(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRL3 (Adhesion G Protein-Coupled Receptor L3) of SEQ ID NO: 3, wherein (i) one of amino acid positions 81 to 83, preferably amino acid position 82 and one of amino acid positions 117 to 119, preferably amino acid position 118 of SEQ ID NO: 3 are replaced by cysteines, (ii) one of amino acid positions 105 to 107, preferably amino acid position 106 and one of amino acid positions 261 to 263, preferably amino acid position 262 of SEQ ID NO: 3 are replaced by cysteines, (iii) one of amino acid positions 89 to 91 , preferably amino acid position 90 and one of amino acid positions 188 to 190, preferably amino acid position 189 of SEQ ID NO: 3 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained;
(IV) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRL2 (Adhesion G Protein-Coupled Receptor L2) of SEQ ID NO: 4, wherein(i) one of amino acid positions 94 to 96, preferably amino acid position 95 and one of amino acid positions 130 to 132, preferably amino acid position 131 of SEQ ID NO: 4 are replaced by cysteines, (ii) one of amino acid positions 118 to 120, preferably amino acid position 119 and one of amino acid positions 272 to 274, preferably amino acid position 273 of SEQ ID NO: 4 are replaced by cysteines, (iii) one of amino acid positions 102 to 104, preferably amino acid position 103 and one of amino acid positions 202 to 204, preferably amino acid position 203 of SEQ ID NO: 4 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained;
(V) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRL1 (Adhesion G Protein-Coupled Receptor L1) of SEQ ID NO: 5, wherein (i) one of amino acid positions 94 to 96, preferably amino acid position 95 and one of amino acid positions 130 to 132, preferably amino acid position 131 of SEQ ID NO: 5 are replaced by cysteines, (ii) one of amino acid positions 118 to 120, preferably amino acid position 119 and one of amino acid positions 275 to 277, preferably amino acid position 276 of SEQ ID NO: 5 are replaced by cysteines, (iii) one of amino acid positions
102 to 104, preferably amino acid position 103 and one of amino acid positions 202 to 204, preferably amino acid position 203 of SEQ ID NO: 5 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained;
(VI) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG7 (Adhesion G Protein-Coupled Receptor G7) of SEQ ID NO: 6, wherein (i) one of amino acid positions 67 to 69, preferably amino acid position 68 and one of amino acid positions 102 to 104, preferably amino acid position 103 of SEQ ID NO: 6 are replaced by cysteines, (ii) one of amino acid positions 90 to 92, preferably amino acid position 91 and one of amino acid positions 246 to 248, preferably amino acid position 247 of SEQ ID NO: 6 are replaced by cysteines, (iii) one of amino acid positions 74 to 76, preferably amino acid position 75 and one of amino acid positions 186 to 188, preferably amino acid position 187 of SEQ ID NO: 6 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained;
(VII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG6 (Adhesion G Protein-Coupled Receptor G6) of SEQ ID NO: 7, wherein(i) one of amino acid positions 35 to 37, preferably amino acid position 36 and one of amino acid positions 84 to 86, preferably amino acid position 85 of SEQ ID NO: 7 are replaced by cysteines, (ii) one of amino acid positions 73 to 75, preferably amino acid position 74 and one of amino acid positions 229 to 231 , preferably amino acid position 230 of SEQ ID NO: 7 are replaced by cysteines, (iii) one of amino acid positions 56 to 58, preferably amino acid position 57 and one of amino acid positions 172 to 174, preferably amino acid position 173 of SEQ ID NO: 7 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained;
(VIII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG5 (Adhesion G Protein-Coupled Receptor G5) of SEQ ID NO: 8, wherein(i) one of amino acid positions 45 to 47, preferably amino acid position 46 and one of amino acid positions 193 to 195, preferably amino acid position 194 of SEQ ID NO: 8 are replaced by cysteines, (ii) one of amino acid positions 30 to 32, preferably amino acid position 31 and one of amino acid positions 137 to 139, preferably amino acid position 138 of SEQ ID NO: 8 are replaced by cysteines, or(iii) (i) and (ii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii) or (iii) are retained;
(IX) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG4 (Adhesion G Protein-Coupled Receptor G4) of SEQ ID NO: 9, wherein (i) one of amino acid positions 251 to 253, preferably amino acid position 252 and one of amino acid positions 302 to 304, preferably amino acid position 303 of SEQ ID NO: 9 are replaced by cysteines, (ii) one of amino acid positions 290 to 292, preferably amino acid position 291 and one of amino acid positions 442 to 444, preferably amino acid position 443 of SEQ ID NO: 9 are replaced by cysteines, (iii) one of amino acid positions 274 to 276, preferably amino acid position 275 and one of amino acid positions 385 to 387, preferably
amino acid position 386 of SEQ ID NO: 9 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained;
(X) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG3 (Adhesion G Protein-Coupled Receptor G3) of SEQ ID NO: 10, wherein(i) one of amino acid positions 16 to 19, preferably amino acid position 17 and one of amino acid positions 45 to 47, preferably amino acid position 46 of SEQ ID NO: 10 are replaced by cysteines, (ii) one of amino acid positions 32 to 34, preferably amino acid position 33 and one of amino acid positions 116 to 118, preferably amino acid position 117 of SEQ ID NO: 10 are replaced by cysteines, or(iii) (i) and (ii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii) or (iii) are retained;
(XI) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG2 (Adhesion G Protein-Coupled Receptor G2) of SEQ ID NO: 11 , wherein(i) one of amino acid positions 12 to 14, preferably amino acid position 13 and one of amino acid positions 63 to 65, preferably amino acid position 64 of SEQ ID NO: 11 are replaced by cysteines, (ii) one of amino acid positions 51 to 53, preferably amino acid position 52 and one of amino acid positions 205 to 207, preferably amino acid position 206 of SEQ ID NO: 11 are replaced by cysteines, (iii) one of amino acid positions 35 to 37, preferably amino acid position 36 and one of amino acid positions 147 to 149, preferably amino acid position 148 of SEQ ID NO: 11 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained;
(XII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF5 (Adhesion G Protein-Coupled Receptor F5) of SEQ ID NO: 12, wherein(i) one of amino acid positions 82 to 84, preferably amino acid position 83 and one of amino acid positions 117 to 119, preferably amino acid position 118 of SEQ ID NO: 12 are replaced by cysteines, (ii) one of amino acid positions 105 to 107, preferably amino acid position 106 and one of amino acid positions 240 to 242, preferably amino acid position 241 of SEQ ID NO: 12 are replaced by cysteines, (iii) one of amino acid positions 90 to 92, preferably amino acid position 91 and one of amino acid positions 186 to 188, preferably amino acid position 187 of SEQ ID NO: 12 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained;
(XIII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF4 (Adhesion G Protein-Coupled Receptor F4) of SEQ ID NO: 13, wherein(i) one of amino acid positions 53 to 55, preferably amino acid position 54 and one of amino acid positions 86 to 88, preferably amino acid position 87 of SEQ ID NO: 13 are replaced by cysteines, (ii) one of amino acid positions 74 to 76, preferably amino acid position 75 and one of amino acid positions 216 to 218, preferably amino acid position 217 of SEQ ID NO: 13 are replaced by cysteines, (iii) one of amino acid positions 61 to 63, preferably amino acid position 62 and one of amino acid positions 159 to 161 , preferably amino acid
position 160 of SEQ ID NO: 13 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained;
(XIV) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF3 (Adhesion G Protein-Coupled Receptor F3) of SEQ ID NO: 14, wherein(i) one of amino acid positions 81 to 83, preferably amino acid position 82 and one of amino acid positions 116 to 118, preferably amino acid position 117 of SEQ ID NO: 14 are replaced by cysteines, (ii) one of amino acid positions 104 to 106, preferably amino acid position 105 and one of amino acid positions 243 to 245, preferably amino acid position 244 of SEQ ID NO: 14 are replaced by cysteines, (iii) one of amino acid positions 89 to 91 , preferably amino acid position 90 and one of amino acid positions 184 to 186, preferably amino acid position 185 of SEQ ID NO: 14 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained;
(XV) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF2 (Adhesion G Protein-Coupled Receptor F2) of SEQ ID NO: 15, wherein(i) one of amino acid positions 73 to 75, preferably amino acid position 74 and one of amino acid positions 106 to 108, preferably amino acid position 107 of SEQ ID NO: 15 are replaced by cysteines, (ii) one of amino acid positions 94 to 96, preferably amino acid position 95 and one of amino acid positions 230 to 232, preferably amino acid position 231 of SEQ ID NO: 15 are replaced by cysteines, (iii) one of amino acid positions 81 to 83, preferably amino acid position 82 and one of amino acid positions 177 to 179, preferably amino acid position 178 of SEQ ID NO: 15 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained;
(XVI) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF1 (Adhesion G Protein-Coupled Receptor F1) of SEQ ID NO: 16, wherein(i) one of amino acid positions 83 to 85, preferably amino acid position 84 and one of amino acid positions 118 to 120, preferably amino acid position 119 of SEQ ID NO: 16 are replaced by cysteines, (ii) one of amino acid positions 106 to 108, preferably amino acid position 107 and one of amino acid positions 243 to 245, preferably amino acid position 244 of SEQ ID NO: 16 are replaced by cysteines, (iii) one of amino acid positions 91 to 93, preferably amino acid position 92 and one of amino acid positions 190 to 192, preferably amino acid position 191 of SEQ ID NO: 16 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained;
(XVII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE5 (Adhesion G Protein-Coupled Receptor E5) of SEQ ID NO: 17, wherein one of amino acid positions 72 to 74, preferably amino acid position 73 and one of amino acid positions 244 to 246, preferably amino acid position 245 of SEQ ID NO: 17 are replaced by cysteines; or (B) comprising or consisting of an
amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A) are retained;
(XVIII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE3 (Adhesion G Protein-Coupled Receptor E3) of SEQ ID NO: 19, wherein one of amino acid positions 73 to 75, preferably amino acid position 74 and one of amino acid positions 186 to 188, preferably amino acid position 187 of SEQ ID NO: 19 are replaced by cysteines; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A) are retained;
(XIX) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE2 (Adhesion G Protein-Coupled Receptor E2) of SEQ ID NO: 20, wherein one of amino acid positions 72 to 74, preferably amino acid position 73 and one of amino acid positions 228 to 230, preferably amino acid position 229 of SEQ ID NO: 20 are replaced by cysteines; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A) are retained;
(XX) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE1 (Adhesion G Protein-Coupled Receptor E1) of SEQ ID NO: 21 , wherein (i) one of amino acid positions
66 to 68, preferably amino acid position 67 and one of amino acid positions 103 to 105, preferably amino acid position 104 of SEQ ID NO: 21 are replaced by cysteines, (ii) one of amino acid positions 91 to 93, preferably amino acid position 92 and one of amino acid positions 243 to 245, preferably amino acid position 244 of SEQ ID NO: 21 are replaced by cysteines, (iii) one of amino acid positions 75 to 77, preferably amino acid position 76 and one of amino acid positions 173 to 175, preferably amino acid position 174 of SEQ ID NO: 21 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained;
(XXI) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRD2 (Adhesion G Protein-Coupled Receptor D2) of SEQ ID NO: 22, wherein(i) one of amino acid positions
67 to 69, preferably amino acid position 68 and one of amino acid positions 110 to 112, preferably amino acid position 111 of SEQ ID NO: 22 are replaced by cysteines, (ii) one of amino acid positions 98 to 100, preferably amino acid position 99 and one of amino acid positions 235 to 237, preferably amino acid position 236 of SEQ ID NO: 22 are replaced by cysteines, or(iii) (i) and (ii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii) or (iii) are retained;
(XXII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRD1 (Adhesion G Protein-Coupled Receptor D1) of SEQ ID NO: 23, wherein(i) one of amino acid positions 36 to 38, preferably amino acid position 37 and one of amino acid positions 92 to 94, preferably amino acid position 93 of SEQ ID NO: 23 are replaced by cysteines, (ii) one of amino acid positions 80 to 82, preferably amino acid position 81 and one of amino acid positions 247 to 249, preferably amino acid position 248 of SEQ ID NO: 23 are replaced by cysteines, (iii) one of amino acid positions 66 to 68, preferably amino acid position 67 and one of amino acid positions 162 to 164, preferably amino acid
position 163 of SEQ ID NO: 23 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained;
(XXIII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRB3 (Adhesion G Protein-Coupled Receptor B3) of SEQ ID NO: 24, wherein(i) one of amino acid positions 87 to 89, preferably amino acid position 88 and one of amino acid positions 121 to 123, preferably amino acid position 122 of SEQ ID NO: 24 are replaced by cysteines, (ii) one of amino acid positions 109 to 111 , preferably amino acid position 110 and one of amino acid positions 261 to 263, preferably amino acid position 262 of SEQ ID NO: 24 are replaced by cysteines, (iii) one of amino acid positions 95 to 97, preferably amino acid position 96 and one of amino acid positions 206 to 208, preferably amino acid position 207 of SEQ ID NO: 24 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained;
(XXIV) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRB2 (Adhesion G Protein-Coupled Receptor B2) of SEQ ID NO: 25, wherein(i) one of amino acid positions 86 to 88, preferably amino acid position 87 and one of amino acid positions 120 to 122, preferably amino acid position 121 of SEQ ID NO: 25 are replaced by cysteines, (ii) one of amino acid positions 108 to 110, preferably amino acid position 109 and one of amino acid positions 295 to 297, preferably amino acid position 296 of SEQ ID NO: 25 are replaced by cysteines, (iii) one of amino acid positions 94 to 96, preferably amino acid position 95 and one of amino acid positions 239 to 241 , preferably amino acid position 240 of SEQ ID NO: 25 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained;
(XXV) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRB1 (Adhesion G Protein-Coupled Receptor B1) of SEQ ID NO: 26, wherein(i) one of amino acid positions 81 to 83, preferably amino acid position 82 and one of amino acid positions 115 to 117, preferably amino acid position 116 of SEQ ID NO: 26 are replaced by cysteines, (ii) one of amino acid positions 103 to 105, preferably amino acid position 104 and one of amino acid positions 253 to 255, preferably amino acid position 254 of SEQ ID NO: 26 are replaced by cysteines, (iii) one of amino acid positions 89 to 91 , preferably amino acid position 90 and one of amino acid positions 198 to 200, preferably amino acid position 199 of SEQ ID NO: 26 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained;
(XXVI) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRA3 (Adhesion G Protein-Coupled Receptor A3) of SEQ ID NO: 27, wherein(i) one of amino acid positions 57 to 59, preferably amino acid position 58 and one of amino acid positions 106 to 108, preferably amino acid position 107 of SEQ ID NO: 27 are replaced by cysteines, (ii) one of amino acid positions
94 to 96, preferably amino acid position 95 and one of amino acid positions 288 to 290, preferably amino acid position 289 of SEQ ID NO: 27 are replaced by cysteines, (iii) one of amino acid positions 81 to 83, preferably amino acid position 82 and one of amino acid positions 225 to 227, preferably amino acid position 226 of SEQ ID NO: 27 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained;
(XXVII) (A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRA2 (Adhesion G Protein-Coupled Receptor A2) of SEQ ID NO: 28, wherein (i) one of amino acid positions 57 to 59, preferably amino acid position 58 and one of amino acid positions 104 to 106, preferably amino acid position 105 of SEQ ID NO: 28 are replaced by cysteines, (ii) one of amino acid positions 92 to 94, preferably amino acid position 93 and one of amino acid positions 294 to 296, preferably amino acid position 295 of SEQ ID NO: 28 are replaced by cysteines, (iii) one of amino acid positions 79 to 81 , preferably amino acid position 80 and one of amino acid positions 229 to 231 , preferably amino acid position 230 of SEQ ID NO: 28 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained; or
(XXVIII) (A) comprising or consisting of the amino acid sequence of GAIN domain of ADGRV1 (Adhesion G Protein-Coupled Receptor V1) of SEQ ID NO: 29, wherein(i) one of amino acid positions 171 to 173, preferably amino acid position 172 and one of amino acid positions 229 to 231 , preferably amino acid position 230 of SEQ ID NO: 29 are replaced by cysteines, (ii) one of amino acid positions 127 to 219, preferably amino acid position 218 and one of amino acid positions 436 to 438, preferably amino acid position 437 of SEQ ID NO: 29 are replaced by cysteines, (iii) one of amino acid positions 201 to 203, preferably amino acid position 202 and one of amino acid positions 377 to 379, preferably amino acid position 378 of SEQ ID NO: 29 are replaced by cysteines, or (iv) two or all three of (i) to (iii) apply; or (B) comprising or consisting of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii) or (iv) are retained.
The GAIN domain (G-protein-coupled receptor autoproteolysis-inducing domain) is a protein domain found in adhesion GPCRs. The domain is involved in the self-cleavage of these transmembrane receptors, and has been shown to be crucial for their function (Liebscher et al. (2015), Oncotarget, Vol. 6, No. 27).
Adhesion G protein-coupled receptors (aGPCRs of AGPCRs) are a class of currently 33 human protein receptors with a broad distribution in embryonic and larval cells, cells of the reproductive tract, neurons, leukocytes, and a variety of tumours. There 33 human adhesion GPCRs are DGRA1 (GPR123), ADGRA2 (GPR124), ADGRA3 (GPR125), ADGRB1 (BAI1), ADGRB2 (BAI2), ADGRB3 (BAI3), ADGRC1 (CELSR1), ADGRC2 (CELSR2), ADGRC3 (CELSR3), ADGRD1 (GPR133),
ADGRD2 (GPR144), ADGRE1 (EMR1 , F4/80), ADGRE2 (EMR2), ADGRE3 (EMR3), ADGRE4 (EMR4), ADGRE5 (CD97), ADGRF1 (GPR110), ADGRF2 (GPR111), ADGRF3 (GPR113), ADGRF4 (GPR115), ADGRF5 (GPR116, Ig-Hepta), ADGRG1 (GPR56), ADGRG2 (GPR64, HE6), ADGRG3 (GPR97), ADGRG4 (GPR112), ADGRG5 (GPR114), ADGRG6 (GPR126), ADGRG7 (GPR128), ADGRL1 (latrophilin-1 , CIRL-1 , CL1), ADGRL2 (latrophilin-2, CIRL-2, CL2), ADGRL3 (latrophilin-3, CIRL-3, CL3), ADGRL4 (ELTD1 , ETL), and ADGRV1 (VLGR1 , GPR98) (Hamann et al. (2015). Pharmacological Reviews, 67(2):338-367). Each subfamily is assigned to a letter, such as L for the latrophilins, E for the EGF-TM7 receptors, C for the CELSRs, B for the BAIs, and V for VLGR, while the subfamilies with GPR# names have been given a letter in alphabetic order (A, D, F, G).
Adhesion GPCRs are distinguished from other GPCR families in that they mostly possess large N and C termini and a GAIN domain. The GAIN domain is evolutionarily well conserved among AGPCRs and seems to be present from tetrahymena to mammals.
AGPCRs undergo autoproteolysis at the GPCR proteolysis site (GPS) within their GAIN domain, resulting in a two-partite structure containing the N-terminal fragment (NTF) and the C-terminal fragment (CTF). Hence, owing to the juxtamembranous position of the GAIN domain and its autoproteolytic function, Adhesion GPCR molecules can be divided into an N-terminal fragment (NTF) and a C-terminal fragment (CTF).
The cleavage site also marks the beginning of the tethered agonist sequence, coined the Stachel sequence. The Stachel sequence is a tethered peptide agonist (Liebscher et al. (2014), cell reports, 9(6):2018-2026) being located within the ectodomain of AGPCRs. The N-terminal fragment (NTF) contains adhesion domains and most of the GAIN domain and the C-terminal fragment (CTF) contains the Stachel sequence, that emanates extracellularly from the seven-transmembrane-spanning (7TM) domain. The NTF and CTF remain non-covalently associated on the plasma membrane via the association of the GAIN domain and the Stachel sequence. Upon a certain mechanical force the GAIN domain and the Stachel sequence dissociate from each other. This is known as Stachel mediated activation of the AGPCR or also as NTF shedding (Beata Jastrzebska, Paul S.H. Park (2019), GPCRs: Structure, Function, and Drug Discovery, Part I, Chapter 2, ISBN 0128167289).
It can be taken from the examples herein below that the mechanical force that is required to dissociate the ADGRG1 GAIN domain from its Stachel sequence has been determined by single-molecule force spectroscopy (SMFS) to be in the range of about 100 pN to about 160 pN (Fig. 1C-D).
The variants of the GAIN domain according to the invention are required to comprise at least two cysteines at specific sites and/or are required to comprise specific amino acid replacements at specific sites. The at least two cysteines can form a disulfide bond that cannot be found in the corresponding wild-type (WT) GAIN domains of SEQ ID NOs 1 to 29. The specific amino acid replacements introduce amino acid substitutions at specific sites that result in amino acids at these specific sites that cannot
be found in the corresponding wild-type (WT) GAIN domains of SEQ ID NOs 1 to 29. Based on the computational approach according to the seventh aspect of the invention (that will be described herein below) the specific sites, wherein the two cysteines are to be added and/or the amino acids are to be substituted were determined for ADGRG1 GAIN domain. The sites were selected such that the mechanical force that is required to dissociate ADGRG1 GAIN domain from is Stachel sequence is either higher or lower (i.e. the variants have an increased or a decreased resistance to mechanical force). This has been achieved because the Cys-Cys bridge adds a two-dimensional lock into the variants of the GAIN domain that it not found in the corresponding WT GAIN domain and/or because the specific amino acid substitutions change the 3-dimensional confirmation of the corresponding WT GAIN domain (i.e. without adding a two-dimensional lock or any other additional linkage). As discussed herein above, such variants of the GAIN domain are advantageous for applications such as cell signaling studies and drug development.
In the case of item (l)(A)(i) of the first aspect of the invention the specific sites of the ADGRG1 GAIN domain are amino acid position 4 and position 48 of SEQ ID NO: 1. This variant GAIN domain is also designated h1 h2 herein and displays a lower force threshold of GAIN-Stachel dissociation (about 67- 135 pN). It is expected that the essentially same results can be achieved within positions 3 to 5 and positions 47 to 49 because the two-dimensional lock by the Cys-Cys bridge is essentially the same as for positions 4 and 38.
In the case of item (l)(A)(ii) of the first aspect of the invention the specific sites of the ADGRG1 GAIN domain are amino acid position 36 and position 181 of SEQ ID NO: 1. This variant GAIN domain is also designated h2l10 herein and displays a higher force threshold of GAIN-Stachel dissociation (about 182-236 pN). It is expected that the essentially same results can be achieved within positions 35 to 37 and positions 180 to 182 because the two-dimensional lock by the Cys-Cys bridge is essentially the same as for positions 36 and 181 .
In the case of item (l)(A)(iii) of the first aspect of the invention the specific sites of the ADGRG1 GAIN domain are amino acid position 20 and position 125 of SEQ ID NO: 1. This variant GAIN domain is also designated hll7 herein and displays a lower force threshold of GAIN-Stachel dissociation (about 66-117 pN). The H1 L7 force propagation is very similar to that of the H1 H2 variant. The H1 L7 variant behaves similarly to the H1 H2 variant as was measured by atomic force microscopy (AFM). It is expected that the essentially same results can be achieved within positions 19 to 21 and positions 124 to 126 because the two-dimensional lock by the Cys-Cys bridge is essentially the same as for positions 20 and 125.
In the case of item (l)(A)(iv) of the first aspect of the invention the specific sites of the ADGRG1 GAIN domain are amino acid position 57 and position 66 of SEQ ID NO: 1. This variant GAIN domain is also designated blb2 herein and displays a lower force threshold of GAIN-Stachel dissociation. This
variant GAIN domain is also designated b1 b2 herein and displays a lower force threshold of GAIN- Stachel dissociation (about 39-58 pN). It is expected that the essentially same results can be achieved within positions 56 to 58 and positions 65 to 67 because the two-dimensional lock by the Cys-Cys bridge is essentially the same as for positions 57 and 66.
In the case of item (l)(A)(v) of the first aspect of the invention the specific sites of the ADGRG1 GAIN domain are amino acid position 17 and position 61 of SEQ ID NO: 1. This variant GAIN domain is also designated h1 b1 herein and displays a lower force threshold of GAIN-Stachel dissociation (about 78- 117 pN). It is expected that the essentially same results can be achieved within positions 16 to 18 and positions 60 to 62 because the two-dimensional lock by the Cys-Cys bridge is essentially the same as for positions 17 and 61 .
In the case of items (l)(A)(vi) to (xi) the amino acids at between one and ten specific sites of the ADGRG1 GAIN domain are substituted by other amino acids. The resulting GAIN domain variants are also designated non-covalent variants herein in contrast to the before discussed SS variants.
In the case of item (l)(A)(vi) the specific amino acid substitutions in the ADGRG1 GAIN domain are the three Q40E/E43L/E181 R. This ADGRG1 GAIN domain variant is also designated 3Mut herein.
In the case of item (l)(A)(vii) the specific amino acid substitutions in the ADGRG1 GAIN domain are the four Q40E/E43L/S135R/E181 R. This ADGRG1 GAIN domain variant is also designated 4Mut herein.
In the case of item (l)(A)(viii) the specific amino acid substitutions in the ADGRG1 GAIN domain are the six Q40E/E43L/S135R/V146/F178Y/E181 R. This ADGRG1 GAIN domain variant is also designated 6Mut herein.
In the case of item (l)(A)(ix) the specific amino acid substitutions in the ADGRG1 GAIN domain are the ten A33W/Q40E/E43L/S135R/V146Y/F178Y/E181 R/T184G/S186Q/S187G. This ADGRG1 GAIN domain variant is also designated 10Mut herein.
In the case of item (l)(A)(x) the specific amino acid substitution in the ADGRG1 GAIN domain is L174I. This ADGRG1 GAIN domain variant is also designated L344I herein (noting that amino acid position 344 of SEQ ID NO: 30 (full ADGRG1) corresponds to amino acid position 174 of SEQ ID NO: 1 (ADGRG1 GAIN domain)).
In the case of item (l)(A)(xi) the specific amino acid substitution in the ADGRG1 GAIN domain is L218A. This ADGRG1 GAIN domain variant is also designated L388A herein (noting that amino acid position 388 of SEQ ID NO: 30 (full ADGRG1) corresponds to amino acid position 218 of SEQ ID NO: 1 (ADGRG1 GAIN domain)).
The GAIN domain of items (l)(A)(vi) to (xi) all display a lower force threshold of GAIN-Stachel dissociation; see Figure 18. In this respect it is of note that also a ADGRG1 GAIN domain variant with eight of the ten A203W/Q210E/E213L/S305R/V316Y/F348Y/E351 R/T354G/S356Q/S357H substitutions was prepared. This 8Mut variant does not display a lower or higher force threshold of GAIN-Stachel dissociation but essentially behaves as the wild-type ADGRG1 GAIN domain. Accordingly, this 8Mut variant is not part of the first aspect.
As compared to the mutational effects of the SS variants the mutational effects in the non-covalent variants are more subtle, which allows for a finer grained modulation of the desired mechanical force. In addition, it is expected for the non-covalent mutants that the effects on the signaling responses to ligands can be different than pure mechanical resistance because the non-covalent mutants show lower signaling responses to TG2 (transglutaminase 2) binding despite being generally less mechanically resistant. Hence, the non-covalent design advantageously provides functional effects that can complement those obtained with the S-S variants.
In the case of item (l)(A)(xiii) at least two of items (l)(A)(i) to (xi) apply. The at least two are with increasing preference at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, and at least ten. In this respect particular preference is provided for ADGRG1 GAIN domain variants that comprise the mutations of alt least one SS variant and at least one non- covalent variant.
According to item (l)(A)(xiii) of the first aspect the different pairs of replaced cysteines of (l)(A)(i) to (v) can be combined and/or the specific amino acid substitutions of (l)(A)(vi) to (xi). This is illustrated in the examples for the variant h1 h2-h2l10. Unlike the individual cysteine pairs, the h1 h2-h2l10 force propagation is very similar to that of the WT. Consistent with these predictions, hi h2-h2l10 displays a force threshold being similar to the WT ADGRG1 GAIN domain.
According to item (l)(B) of the first aspect the variant GAIN domain comprises or consists of an amino acid sequence sharing at least 80% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii), (iii), (iv), (v) and/or (vi), or the amino acid replacement(s) as defined in (A), item (vii), (viii), (ix), (x), (xi) an/or (xii), or the cysteines and/or the amino acid replacement(s) as defined in (A), item (xiii) are retained..
The amino acid sequence sharing at least 80% sequence identity displays with increasing preference at least 85% sequence identity, at least 90% sequence identity, at least 95% sequence identity, at least 96% sequence identity, at least 97% sequence identity, at least 98% sequence identity and at least 99% sequence identity.
In accordance with the present invention, the term “percent (%) sequence identity” describes the
number of matches (“hits”) of identical nucleotides/amino acids of two or more aligned nucleic acid or amino acid sequences as compared to the number of nucleotides or amino acid residues making up the overall length of the template nucleic acid or amino acid sequences. In other terms, using an alignment, for two or more sequences or subsequences the percentage of amino acid residues or nucleotides that are the same (e.g. 70%, 75%, 80%, 85%, 90% or 95% identity) may be determined, when the (sub)sequences are compared and aligned for maximum correspondence over a window of comparison, or over a designated region as measured using a sequence comparison algorithm as known in the art, or when manually aligned and visually inspected. This definition also applies to the complement of any sequence to be aligned.
Nucleotide and amino acid sequence analysis and alignment in connection with the present invention are preferably carried out using the NCBI BLAST algorithm (Stephen F. Altschul, Thomas L. Madden, Alejandro A. Schaffer, Jinghui Zhang, Zheng Zhang, Webb Miller, and David J. Lipman (1997), Nucleic Acids Res. 25:3389-3402). BLAST can be used for nucleotide sequences (nucleotide BLAST) and amino acid sequences (protein BLAST). The skilled person is aware of additional suitable programs to align nucleic acid sequences. The NCBI BLAST algorithm is available for protein (Protein BLAST) and nucleotides (Nucleotide BLAST). For Protein BLAST the algorithm parameters are preferably: max target sequences: 100, with automatically adjust parameters for short input sequences, expect threshold 0.05, word size 6, Max matches in a query range 0, matrix BLOSUM62, cap cost existence: 10 extension: 1 , and compositional adjustment. For Nucleotide BLAST the algorithm parameters are preferably: max target sequences: 100, with “automatically adjust parameters for short input sequences”, Expect threshold 0.05, word size 28, Max matches in a query range 0, match/mismatch scores 1 ,-2, cap costs linear, low complexity regions filter, and .mask for look up table only. These are the standard algorithm parameters for protein BLAST and Nucleotide BLAST and they can be adjusted, if needed.
It is furthermore preferred in connection with item (l)(B) that no further amino acids with the exception of the ones as defined in (l)(A), item (i), (ii), (iii), (iv), and/or (v) are mutated to cysteine as compared to SEQ ID NO: 1 , so that it is ensured that no additional cysteine bridges can be formed.
While item (I) of the first aspect of the invention recites specific sites of the ADGRG1 GAIN domain, it has already been discussed above that the amino acid sequence of the GAIN domain is highly conserved throughout the class of AGPCRs. For this reason the above-discussed specific sites of the ADGRG1 GAIN domain have homologues sites in nearly all other GAIN domains of AGPCRs.
Table A - Cys-Cys sites in GAIN domains of AGPCRs corresponding to the GAIN domain variants h1 h2, h2l10 and h1 l7. The amino acid positions are counted based on the amino acid sequences of the GAIN domains (SEQ ID NOs 1 to 29).
Table B - Cys-Cys sites in GAIN domains of AGPCRs corresponding to the GAIN domain variants b1 b2 and h1 b1. The amino acid positions are counted based on the amino acid sequences of the GAIN domains (SEQ ID NOs 1 to 29).
Table C - Non-covalent mutations sites in GAIN domains of the AGPCRG1 domain variants 10Mut, L344I and L388A (noting the variants 3Mut, 4Mut and 6Mut comprise a subset of the mutations of 10 Mut). The amino acid positions are counted based on the amino acid sequences of the GAIN domains (SEQ ID NOs 1 to 29).
The homologous sites as shown in Tables A and B are covered by items (II) to (XXIX) of the first aspect and the above definitions and preferred embodiments of item (I) apply mutatis mutandis to items (II) to (XXIX) of the first aspect of the invention.
For instance, also according to items (II) to (XXIX) (B) of the first aspect the amino acid sequence sharing at least 80% sequence identity display with increasing preference at least 85% sequence identity, at least 90% sequence identity, at least 95% sequence identity, at least 96% sequence identity, at least 97% sequence identity, at least 98% sequence identity and at least 99% sequence identity.
In addition, also according to items (II) to (XXIX) of the first aspect the different pairs of replaced cysteines can be combined.
The present invention relates in a second aspect to an adhesion G protein-coupled receptor (AGPCR)
comprising the variant of the GAIN domain the first aspect, wherein the AGPCR preferably comprises of consists of any of SEQ ID NOs 30 to 58 or a sequence being at least 70% identical, preferably at least 80% and most preferably at least 90% identical thereto.
The definitions and preferred embodiments of the first aspect as far as being amenable for combination with the second aspect apply mutatis mutandis to the second aspect of the invention. This applies mutatis mutandis to the further aspects of the invention as described herein below. For instance, the definitions and preferred embodiments of the first and second aspect as far as being amenable for combination with the third aspect apply mutatis mutandis to the third aspect of the invention, etc.
The amino acid sequence sharing at least 70% sequence identity displays with increasing preference at least 75% sequence identity, at least 80% sequence identity, at least 85% sequence identity, at least 90% sequence identity, at least 95% sequence identity, at least 96% sequence identity, at least 97% sequence identity, at least 98% sequence identity and at least 99% sequence identity.
SEQ ID NOs 30 to 58 are the full-length amino acid sequences of the wild-type AGPCRS as named in items (I) to (XXIX) of the first aspect, in that particular order. It follows that the GAIN domain of SEQ ID NO: 1 is comprised in SEQ ID NO: 30, the GAIN domain of SEQ ID NO: 2 is comprised in SEQ ID NO: 31 , etc. (see table A herein above).
In the AGPCRs of the second aspect these wild-type GAIN domains are replaced by a GAIN domain of the first aspect while preferably overall a sequence being of at least 70% identical, preferably of at least 80% and most preferably of at least 90% identical to any of SEQ ID NOs 30 to 58 is retained.
It is particularly preferred that in the wild-type AGPCRs of SEQ ID NOs 30 to 58 the wild-type GAIN domains are replaced by a corresponding variant of the GAIN domains. For instance, in SEQ ID NO: 30 the wild-type GAIN domain of SEQ ID NO. 1 is preferably replaced by a variant of the GAIN domain of item (I) of the first aspect, in SEQ ID NO: 31 the wild-type GAIN domain of SEQ ID NO: 2 is preferably replaced by a variant of the GAIN domain of item (II) of the first aspect, etc. Again, preferably overall a sequence being at least 70% identical, preferably at least 80% and most preferably at least 90% identical to any of SEQ ID NOs 30 to 58 is retained.
According to a preferred embodiment, the variant of the GAIN domain of the first aspect and/or the AGPCR of the second aspect is a fusion protein.
The term “protein” as used herein is used interchangeably with the term (poly)peptide and refers to polymers constructed from one or more chains of amino acid residues linked by peptide bonds. (Poly)peptides describe a group of molecules which comprise the group of peptides consisting of up to 30 amino acids, as well as the group of polypeptides consisting of more than 30 amino acids. The
term “protein” also refers to chemically or post-translationally modified proteins. The terms apply to naturally occurring amino acid polymers, as well as to amino acid polymers in which one or more amino acid residue is an artificial chemical analogue of a corresponding naturally occurring amino acid.
A “fusion protein” according to the present invention contains at least one additional heterologous amino acid sequence other than the variant of the GAIN domain and/or the AGPCR. Often, but not necessarily, these additional sequences will be located at the N- or C-terminal end of the (poly)peptide. It may e.g. be convenient to initially express the (poly)peptide as a fusion protein from which the additional amino acid residues can be removed, e.g. by a proteinase capable of specifically trimming the fusion protein and releasing the (poly)peptide of the present invention. The additional heterologous amino acid sequence can either be directly or indirectly fused to the variant of the GAIN domain and/or the AGPCR of the invention. In case of an indirect fusion generally a peptide linker may be used for the fusion, such as a GS-linker.
Those at least one additional heterologous amino acid sequence of said fusion proteins includes amino acid sequences which confer desired properties such as modified/enhanced stability, modified/enhanced solubility and/or the ability of targeting one or more specific cell types. For example, fusion proteins with antibodies are envisioned herein. The term “antibody” comprises antibody fragments and derivatives. The antibody may be, for example, specific for cell surface markers or may be an antigen-recognizing fragment of said antibodies. The poly(peptide) of the invention can be fused to the N-terminus or C-terminus of the light and/or heavy chain(s) of an antibody. The poly(peptide) of the invention is preferably fused to the N-terminus of the light and/or heavy chain(s) of an antibody, so that the Fc part of the antibody is free to bind to Fc-receptors.
The term “antibody” as used in accordance with the present invention comprises, for example, polyclonal or monoclonal antibodies. Furthermore, also derivatives or fragments thereof, which still retain the binding specificity to the desired target, e.g. a tumor antigen, are comprised in the term "antibody". Antibody fragments or derivatives comprise, inter alia, Fab or Fab’ fragments, Fd, F(ab')2, Fv or scFv fragments, single domain VH or V-like domains, such as VhH or V-NAR-domains, as well as multimeric formats such as minibodies, diabodies, tribodies or triplebodies, tetrabodies or chemically conjugated Fab’-multimers (see, for example, Harlow and Lane "Antibodies, A Laboratory Manual", Cold Spring Harbor Laboratory Press, 198; Harlow and Lane “Using Antibodies: A Laboratory Manual” Cold Spring Harbor Laboratory Press, 1999; Altshuler EP, Serebryanaya DV, Katrukha AG. 2010, Biochemistry (Mose)., vol. 75(13), 1584; Holliger P, Hudson PJ. 2005, Nat Biotechnol., vol. 23(9), 1126). The multimeric formats in particular comprise bispecific antibodies that can simultaneously bind to two different types of antigen. The first antigen can be found on the poly(peptide) of the invention. The second antigen may, for example, be a tumor marker that is specifically expressed on cancer cells or a certain type of cancer cells. Non-limiting examples of bispecific antibodies formats are Biclonics (bispecific, full length human IgG antibodies), DART (Dual-
affinity Re-targeting Antibody) and BiTE (consisting of two single-chain variable fragments (scFvs) of different antibodies) molecules (Kontermann and Brinkmann (2015), Drug Discovery Today, 20(7):838-847).
The term "antibody" also includes embodiments such as chimeric (human constant domain, nonhuman variable domain), single chain and humanised (human antibody with the exception of nonhuman CDRs) antibodies.
The fusion protein may also comprise protein domains known to function in signal transduction and/or known to be involved in protein-protein interaction. Examples for such domains are Ankyrin repeats; arm, Bcl-homology, Bromo, CARD, CH, Chr, C1 , C2, DD, DED, DH, EFh, ENTH, F-box, FHA, FYVE, GEL, GYF, hect, LIM, MH2, PDZ, PB1 , PH, PTB, PX, RGS, RING, SAM, SC, SH2, SH3, SOCS, START, TIR, TPR, TRAF, tsnare, Tubby, UBA, VHS, W, WW, and 14-3-3 domains. Further information about these and other protein domains is available from the databases InterPro (http://www.ebi.ac.uk/interpro/, Mulder et al., 2003, Nucl. Acids. Res. 31 : 315-318), Pfam (http://www.sanger.ac.uk/Software/Pfam/, Bateman et al., 2002, Nucleic Acids Research 30(1): 276- 280) and SMART (http://smart.embl-heidelberg.de/, Letunic et al., 2002, Nucleic Acids Res. 30(1), 242-244).
The at least one additional heterologous amino acid sequence of the fusion protein according to the present invention may also comprise or consist of (a) a cytokine, (b) a chemokine, (c) a pro-coagulant factor, (d) a proteinaceous toxic compound, and/or (e) an enzyme for pro-drug activation.
The cytokine is preferably selected from the group consisting of IL-2, IL-12, TNF-alpha, IFN alpha, IFN beta, IFN gamma, IL-10, IL-15, IL-24, GM-CSF, IL-3, IL-4, IL-5, IL-6, IL-7, IL-9, IL-11 , IL-13, LIF, CD80, B70, TNF beta, LT-beta, CD-40 ligand, Fas-ligand, TGF-beta, IL-1 alpha and IL-1 beta. As it is well-known in the art, cytokines may favour a pro-inflammatory or an anti-inflammatory response of the immune system. Thus, depending on the disease to be treated either fusion proteins with a pro- inflammatory or an anti-inflammatory cytokine may be favored. For example, for the treatment of inflammatory diseases in general fusion constructs comprising anti-inflammatory cytokines are preferred, whereas for the treatment of cancer in general fusion constructs comprising pro- inflammatory cytokines are preferred.
The chemokine is preferably selected from the group consisting of IL-8, GRO alpha, GRO beta, GRO gamma, ENA-78, LDGF-PBP, GCP-2, PF4, Mig, IP-10, SDF-1alpha/beta, BUNZO/STRC33, l-TAC, BLC/BCA-1 , MIP-1 alpha, MIP-1 beta, MDC, TECK, TARC, RANTES, HCC-1 , HCC-4, DC-CK1 , MIP-3 alpha, MIP-3 beta, MCP-1-5, eotaxin, Eotaxin-2, I-309, MPIF-1 , 6Ckine, CTACK, MEC, lymphotactin and fractalkine. The major role of chemokines is to act as a chemoattractant to guide the migration of cells. Cells that are attracted by chemokines follow a signal of increasing chemokine concentration towards the source of the chemokine. It follows that within the fusion protein the chemokine can be
used to guide the migration of the protein or peptide of the invention, e.g. to a specific cell type or body site.
The pro-coagulant factor is preferably a tissue factor. A pro-coagulant factor promotes the process by which blood changes from a liquid to a gel, forming a blood clot. Pro-coagulant factors may, for example, aid in wound healing.
The proteinaceous toxic compound is preferably Ricin-A chain, modeccin, truncated Pseudomonas exotoxin A, diphtheria toxin or recombinant gelonin. Toxic compounds can have a toxic effect on a whole organism as well as on a substructure of the organism, such as a particular cell type. Toxic compounds are frequently used in the treatment of tumors. Tumor cells generally grow faster than normal body cells, so that they preferentially accumulate toxic compounds and in higher amounts.
The enzyme for pro-drug activation is preferably an enzyme selected from the group consisting of carboxy-peptidases, glucuronidases and glucosidases. Among the broad array of genes that have been evaluated for tumor therapy, those encoding pro-drug activation enzymes are especially appealing as they directly complement ongoing clinical chemotherapeutic regimes. These enzymes can activate prodrugs that have low inherent toxicity using both bacterial and yeast enzymes or enhance prodrug activation by mammalian enzymes.
The fusion protein may also comprise a tag, such as purification tag. Several purification tags are available and an overview of affinity tags for protein purification is available in Kimple et al. (2013), Curr Protoc Protein Sci. 2013; 73: Unit-9.9. In the examples a HA tag (3xHA) is illustrated.
In accordance with a preferred embodiment, the variant of the GAIN domain and/or the AGPCR is fused to a heterologous non-proteinaceous compound.
As used herein, a heterologous compound is a compound that cannot be found in nature fused to the variant of the GAIN domain and/or the AGPCR.
The heterologous non-proteinaceous compound can either be directly or indirectly fused to the variant of the GAIN domain and/or the AGPCR. For example, a chemical linker may be used. Chemical linkers may contain diverse functional groups, such as primary amines, sulfhydryls, acids, alcohols and bromides. Many crosslinkers are functionalized with maleimide (sulfhydryl reactive) and succinimidyl ester (NHS) or isothiocyanate (ITC) groups that react with amines.
The heterologous non-proteinaceous compound is preferably a pharmaceutically active compound or diagnostically active compound. The pharmaceutically active compound or diagnostically active compound is preferably selected from the group consisting of (a) a fluorescent dye, (b) a photosensitizer, (c) a radionuclide, (d) a contrast agent for medical imaging, (e) a toxic compound, or
(f) an ACE inhibitor, a Renin inhibitor, an ADH inhibitor, an Aldosteron inhibitor, an Angiotensin receptor blocker, a TSH-receptor, a LH-/HCG-receptor, an oestrogen receptor, a progesterone receptor, an androgen receptor, a GnRH-receptor, a GH (growth hormone) receptor, or a receptor for IGF-I or IGF-II.
The fluorescent dye is preferably a component selected from Alexa Fluor or Cy dyes.
The photosensitizer is preferably phototoxic red fluorescent protein KillerRed or haematoporphyrin.
The radionuclide is preferably either selected from the group of gamma-emitting isotopes, more preferably 99mTc, 123l, 111ln, and/or from the group of positron emitters, more preferably 18F, 64Cu, 68Ga, 86Y, 124l, and/or from the group of beta-emitters, more preferably 131l, 90Y, 177Lu, 67Cu, 90Sr, and/or from the group of alpha-emitters, preferably 213Bi, 211At.
A contrast agent as used herein is a substance used to enhance the contrast of structures or fluids within the body in medical imaging. Common contrast agents work based on X-ray attenuation and magnetic resonance signal enhancement.
The toxic compound is preferably a small organic compound, more preferably a toxic compound selected from the group consisting of calicheamicin, maytansinoid, neocarzinostatin, esperamicin, dynemicin, kedarcidin, maduropeptin, doxorubicin, daunorubicin, and auristatin. In contrast to the herein above described proteinaceous toxic compound these toxic compounds are non-proteinaceous.
The present invention relates in a third aspect to a nucleic acid molecule, preferably a vector encoding the variant of the GAIN domain of the first aspect or the AGPCR of the second aspect.
The term “nucleic acid molecule” in accordance with the present invention includes DNA, such as cDNA or double or single stranded genomic DNA and RNA. In this regard, "DNA" (deoxyribonucleic acid) means any chain or sequence of the chemical building blocks adenine (A), guanine (G), cytosine (C) and thymine (T), called nucleotide bases, that are linked together on a deoxyribose sugar backbone. DNA can have one strand of nucleotide bases, or two complimentary strands which may form a double helix structure. "RNA" (ribonucleic acid) means any chain or sequence of the chemical building blocks adenine (A), guanine (G), cytosine (C) and uracil (U), called nucleotide bases, that are linked together on a ribose sugar backbone. RNA typically has one strand of nucleotide bases, such as mRNA. Included are also single- and double-stranded hybrids molecules, i.e., DNA-DNA, DNA- RNA and RNA-RNA. The nucleic acid molecule may also be modified by many means known in the art. Non-limiting examples of such modifications include methylation, "caps", substitution of one or more of the naturally occurring nucleotides with an analog, and internucleotide modifications such as, for example, those with uncharged linkages (e.g., methyl phosphonates, phosphotriesters, phosphoroamidates, carbamates, etc.) and with charged linkages (e.g., phosphorothioates,
phosphorodithioates, etc.). Nucleic acid molecules, in the following also referred as polynucleotides, may contain one or more additional covalently linked moieties, such as, for example, proteins (e.g., nucleases, toxins, antibodies, signal peptides, poly-L-lysine, etc.), intercalators (e.g., acridine, psoralen, etc.), chelators (e.g., metals, radioactive metals, iron, oxidative metals, etc.), and alkylators. The polynucleotides may be derivatized by formation of a methyl or ethyl phosphotriester or an alkyl phosphoramidate linkage. Further included are nucleic acid mimicking molecules known in the art such as synthetic or semi-synthetic derivatives of DNA or RNA and mixed polymers. Such nucleic acid mimicking molecules or nucleic acid derivatives according to the invention include phosphorothioate nucleic acid, phosphoramidate nucleic acid, 2’-0-methoxyethyl ribonucleic acid, morpholino nucleic acid, hexitol nucleic acid (HNA), peptide nucleic acid (PNA) and locked nucleic acid (LNA) (see Braasch and Corey, Chem Biol 2001 , 8: 1). LNA is an RNA derivative in which the ribose ring is constrained by a methylene linkage between the 2’-oxygen and the 4’-carbon. Also included are nucleic acids containing modified bases, for example thio-uracil, thio-guanine and fluoro-uracil. A nucleic acid molecule typically carries genetic information, including the information used by cellular machinery to make proteins and/or polypeptides. The nucleic acid molecule of the invention may additionally comprise promoters, enhancers, response elements, signal sequences, polyadenylation sequences, introns, 5'- and 3'- non-coding regions, and the like.
The term “vector” in accordance with the invention means preferably a plasmid, cosmid, virus, bacteriophage or another vector used e.g. conventionally in genetic engineering which carries the nucleic acid molecule of the invention. The nucleic acid molecule of the invention may, for example, be inserted into several commercially available vectors. Non-limiting examples include prokaryotic plasmid vectors, such as of the pUC-series, pBluescript (Stratagene), the pET-series of expression vectors (Novagen) or pCRTOPO (Invitrogen) and vectors compatible with an expression in mammalian cells like pREP (Invitrogen), pcDNA3 (Invitrogen), pCEP4 (Invitrogen), pMCI neo (Stratagene), pXT1 (Stratagene), pSG5 (Stratagene), EBO-pSV2neo, pBPV-1 , pdBPVMMTneo, pRSVgpt, pRSVneo, pSV2-dhfr, plZD35, pLXlN, pSIR (Clontech), pIRES-EGFP (Clontech), pEAK-10 (Edge Biosystems) pTriEx-Hygro (Novagen) and pCINeo (Promega). Examples for plasmid vectors suitable for Pichia pastoris comprise e.g. the plasmids pAO815, pPIC9K and pPIC3.5K (all Invitrogen).
The nucleic acid molecules inserted into the vector can e.g. be synthesized by standard methods, or isolated from natural sources. Ligation of the coding sequences to transcriptional regulatory elements and/or to other amino acid encoding sequences can also be carried out using established methods. Transcriptional regulatory elements (parts of an expression cassette) ensuring expression in prokaryotes or eukaryotic cells are well known to those skilled in the art. These elements comprise regulatory sequences ensuring the initiation of transcription (e. g., translation initiation codon, promoters, such as naturally-associated or heterologous promoters and/or insulators; see above), internal ribosomal entry sites (IRES) (Owens, Proc. Natl. Acad. Sci. USA 98 (2001), 1471-1476) and optionally poly-A signals ensuring termination of transcription and stabilization of the transcript. Additional regulatory elements may include transcriptional as well as translational enhancers.
Preferably, the polynucleotide encoding the polypeptide/protein or fusion protein of the invention is operatively linked to such expression control sequences allowing expression in prokaryotes or eukaryotic cells. The vector may further comprise nucleic acid sequences encoding secretion signals as further regulatory elements. Such sequences are well known to the person skilled in the art. Furthermore, depending on the expression system used, leader sequences capable of directing the expressed polypeptide to a cellular compartment may be added to the coding sequence of the polynucleotide of the invention. Such leader sequences are well known in the art.
Furthermore, it is preferred that the vector comprises a selectable marker. Examples of selectable markers include genes encoding resistance to neomycin, ampicillin, hygromycine, and kanamycin. Specifically-designed vectors allow the shuttling of DNA between different hosts, such as bacteria- fungal cells or bacteria-animal cells (e. g. the Gateway system available at Invitrogen). An expression vector according to this invention is capable of directing the replication, and the expression, of the polynucleotide and encoded peptide or fusion protein of this invention. Apart from introduction via vectors such as phage vectors or viral vectors (e.g. adenoviral, retroviral), the nucleic acid molecules as described herein above may be designed for direct introduction or for introduction via liposomes into a cell. Additionally, baculoviral systems or systems based on vaccinia virus or Semliki Forest virus can be used as eukaryotic expression systems for the nucleic acid molecules of the invention.
The vector is preferably a retroviral vector. A retroviral vector consists of proviral sequences that can accommodate the gene of interest, to allow incorporation of both into the target cells.
The present invention relates in a fourth aspect to a cell, preferably an epithelial cell or a lymphocyte and most preferably a T-cell comprising the nucleic acid molecule, preferably the vector of the third aspect.
The cell is preferably an in vitro cell and not an in vivo cell. The cell may also be referred to as a “host cell”.
The term "host cell" means any cell of any organism that is selected, modified, transformed, grown, or used or manipulated in any way, for the production of the GAIN domain and/or the AGPCR of the invention by the cell.
The host cell of the invention is typically produced by introducing the nucleic acid molecule or vector(s) of the invention into the host cell which upon its/their presence mediates the expression of the nucleic acid molecule of the invention encoding the GAIN domain and/or the AGPCR and/or the fusion proteins of the invention. The host from which the host cell is derived or isolated may be any prokaryote or eukaryotic cell or organism, preferably with the exception of human embryonic stem cells that have been derived directly by destruction of a human embryo or are derived from such stem cells.
Suitable prokaryotes (bacteria) useful as hosts for the invention are, for example, those generally used for cloning and/or expression like E. coli (e.g., E coli strains BL21 , HB101 , DH5a, XL1 Blue, Y1090 and JM101), Salmonella typhimurium, Serratia marcescens, Burkholderia glumae, Pseudomonas putida, Pseudomonas fluorescens, Pseudomonas stutzeri, Streptomyces lividans, Lactococcus lactis, Mycobacterium smegmatis, Streptomyces coelicolor or Bacillus subtilis. Appropriate culture mediums and conditions for the above-described host cells are well known in the art.
A suitable eukaryotic host cell may be a vertebrate cell, an insect cell, a fungal/yeast cell, a nematode cell or a plant cell. The fungal/yeast cell may a Saccharomyces cerevisiae cell, Pichia pastoris cell or an Aspergillus cell. Preferred examples for host cell to be genetically engineered with the nucleic acid molecule or the vector(s) of the invention is a cell of yeast, E. coli and/or a species of the genus Bacillus (e.g., B. subtilis). In one preferred embodiment the host cell is a yeast cell (e.g. S. cerevisiae).
In a different preferred embodiment the host cell is a mammalian host cell, such as a Chinese Hamster Ovary (CHO) cell, mouse myeloma lymphoblastoid, human embryonic kidney cell (HEK-293), human embryonic retinal cell (Crucell's Per.C6), or human amniocyte cell (Glycotope and CEVEC). The cells are frequently used in the art to produce recombinant proteins. CHO cells are the most commonly used mammalian host cells for industrial production of recombinant protein therapeutics for humans.
The following cell types are also suitable as host cells:
Epithelial cells make up primary tissues throughout the body. There are many arrangements of epithelial cells such as squamous, cuboidal, and columnar that organize as simple, stratified, pseudostratified, and transitional. Epithelial cells form from ectoderm, mesoderm, and endoderm, which explains why epithelia line body cavities and cover most of the body and organ surfaces. Since epithelial cells are prevalent throughout the body, their function changes based on their location. For example, epithelial cells in the skin provide protection, whereas, in the gut, they have secretory and absorptive properties. Epithelial cells have specialized cytoskeletons comprised of microtubules, actin filaments, and intermediate filaments. Intermediate filaments provide structural resilience to the cytoskeleton and show tissue-specific expression, unlike microtubules and actin filaments. These intermediate filaments are a useful target for histochemical staining techniques.
A lymphocyte is a type of white blood cell in the immune system of jawed vertebrates. Lymphocytes include, innate lymphoid cells (ILCs, i.e. innate counterparts of T cells that contribute to immune responses by secreting effector cytokines and regulating the functions of other innate and adaptive immune cells), natural killer cells (which function in cell-mediated, cytotoxic innate immunity), T cells (for cell-mediated, cytotoxic adaptive immunity), and B cells (for humoral, antibody-driven adaptive immunity).
The lymphocyte is preferably an anti-tumor lymphocyte. An anti-tumor lymphocyte (or an anti-tumor effector lymphocyte) is a lymphocyte capable of eliciting a cytolytic response that can cause tumor cell death. These lymphocytes are specializing in and equipped for tumor cell elimination. The first category encompasses clonally expanded T lymphocytes expressing a unique T cell receptor (TCR) and recognizing tumor epitopes in the context of the major histocompatibility complex (MHC) molecules. These T cells, optionally together with B cells producing tumor-specific antibodies and dendritic cells (DC) processing and presenting tumor epitopes, can mediate an adaptive immunity against tumors. The second category of effector cells includes natural killer (NK) cells, NK-T cells, and macrophages (M). These cells are not restricted by the MHC molecules in their interactions with tumor targets, and they mediate innate immunity. Each type of effector cells, whether specific or nonspecific, contains subsets of cells at different stages of differentiation and activation. This means that each type of effector potentially able to target tumor cells contains a heterogeneous mix of cells with distinct functional capabilities, depending on their stage of differentiation, maturation, and/or activation (Holland, Frei; Cancer Medicine; 6th edition, chapter “Antitumor Effector Cells in Humans”). All these types of anti-tumor lymphocytes are envisioned in accordance with the present invention.
The lymphocytes are preferably T-cells or NK cells, whereby T-cells are further preferred.
A T-cell or T-lymphocyte can be distinguished from other lymphocytes by the presence of a T-cell receptor (TCR) on its cell surface. One of the functions of T-cells is mediating immune-mediated cell death, and it is carried out by two major subtypes: CD8+ "killer" and CD4+ "helper" T-cells. CD8+ T cells are cytotoxic which means that they are able to directly kill selected cell. These selected cells are in accordance with the invention tumor cells, virus-infected cells, as well as cancer cells. CD4+ cells function as "helper cells". Unlike CD8+ killer T-cells, these CD4+ helper T-cells function by further activating memory B cells and cytotoxic T-cells, which leads to a larger immune response which is in accordance with the invention directed against tumor cells. The specific adaptive immune response regulated by the T-helper cell depends on its subtype, which is distinguished by the types of cytokines they secrete. The T-cells are preferably a CD8+ killer T-cells or mixture of CD8+ killer and CD4+ helper T-cells.
A natural killer (NK) cell is a type of cytotoxic lymphocyte being critical to the innate immune system that belong to the rapidly expanding family of innate lymphoid cells (ILC) and represent 5-20% of all circulating lymphocytes in humans. The role of NK cells in innate immune system is analogous to that of cytotoxic T-cells in the vertebrate adaptive immune response. NK cells provide rapid responses to virus-infected cells and other intracellular pathogens acting at around 3 days after infection and respond to tumor formation. Typically, immune cells detect the major histocompatibility complex (MHC) presented on infected cell surfaces, triggering cytokine release, causing the death of the infected cell by lysis or apoptosis. NK cells are unique, however, as they have the ability to recognize and kill stressed cells in the absence of antibodies and MHC, allowing for a much faster immune
reaction. They were named "natural killers" because they do not require activation to kill cells that are missing "self markers of MHC class 1 . This role is especially important because harmful cells that are missing MHC I markers cannot be detected and destroyed by other immune cells, such as T-cells.
The lymphocytes may also be chimeric antigen receptor T-cells (CAR T-cells), T-cell-receptor- engineered T-cells (TCR T-cells), chimeric antigen receptor NK-cells (CAR NK-cells), NK cell receptor- engineered NK cells (NCR NK-cells), TCR/CAR hybrid T-cells, NCR/CAR hybrid NK-cells or tumorinfiltrating lymphocytes (TILs).
Chimeric antigen receptor T-cells (CAR) T-cells are T-cells that have been genetically engineered to produce a chimeric T cell receptor (CAR) for use in immunotherapy. The receptors are chimeric because they combine both antigen-binding and T-cell activating functions into a single receptor. CAR-T cell therapy uses T-cells engineered with CARs for cancer therapy. The premise of CAR-T immunotherapy is to modify T-cells to recognize tumor cells in order to more effectively target and destroy them in order to generate CAR T-cells. T-cells are harvested from subjects, genetically altered, and then infused into patients to attack a tumor in a subject. CAR T-cells can be both CD4+ and/or CD8+ cells. A 1 -to-1 ratio of both cell types is preferred since it provides synergistic antitumor effects.
CAR T-cells are engineered to transfer arbitrary specificity onto a T cell, which specifically eliminates antigen-bearing tumor cells. The CAR may comprise a scFv being derived from an antibody, a CD3£ and a transmembrane domain (so-called first-generation CARs). In this way, the engineered CAR is able to recognize specific tumor associated-antigens. Therefore, the CAR has the ability to bind unprocessed tumor surface antigens without MHC processing while TCRs engage with both tumor intracellular and surface antigenic peptides embedded in MHC.
In contrast, TCRs are a/p heterodimers that bind to the MHC-bound antigens. As discussed above, CARs recognize tumor antigen which led to T-cell activation with different functions compared with TCRs. CAR-T cell therapy has certain disadvantages like off-tumor toxicities when targeting tumorspecific antigen. Compared with CARs, TCRs have several structural advantages in T cell-based therapy, such as more subunits in their receptor structure (ten subunits vs one subunit), greater immunoreceptor tyrosine-based activation motif (ITAMs) (ten vs three), less dependence on antigens (one vs 100), and more co-stimulate receptors (CD3, CD4, CD28, etc.) (Zhao et al. (2021) Front. Immunol., | https://doi.org/10.3389/fimmu.2021.658753).
CAR NK-cells are distinguished from CAR T-cells in that the chimeric antigen receptor is introduced into NK cells instead of T-cells. Just as CAR T-cells, CAR-NK cells can be engineered to target diverse antigens, enhance proliferation and persistence in vivo, increase infiltration into solid tumours, overcome resistant tumour microenvironment, and ultimately achieve an effective anti-tumour response.
Natural cytotoxicity receptor NK cells (NCR NK-cells) are NK-cell that have been genetically engineered to express a NCR. The NCRs have been proposed to bind to many cellular ligands which are implicated in NK cell surveillance of tumor cells. Many of these interactions have been shown to evoke the cytotoxic and cytokine-secreting functions of NK cells. However, it is also possible that the NCRs may regulate other anti-tumor pathways. NCRs and their ligands can be successfully targeted for cancer immunotherapy. NCRs have been classically defined as activating receptors delivering potent signals to NK cells in order to lyse harmful cells and to produce inflammatory cytokines.
TCR/CAR hybrid T-cells are T-cells that have been genetically engineered to express a TCR and CAR. Similarly, NCR/CAR hybrid NK-cells are T-cells that have been genetically engineered to express a NCR and CAR.
Tumor-infiltrating lymphocytes (TILs) are white blood cells that have left the bloodstream and migrated towards a tumor. TILs are implicated in killing tumor cells. The presence of lymphocytes in tumors is often associated with better clinical outcomes.
In accordance with a more preferred embodiment of the invention, the tumor-infiltrating lymphocytes are tumor-infiltrating T-cells or tumor-infiltrating NK cells.
In adoptive T-cell transfer therapy, TILs are expanded ex vivo from surgically resected tumors that have been cut into small fragments or from single cell suspensions isolated from the tumor fragments. Multiple individual cultures are established, grown separately and assayed for specific tumor recognition. TILs are typically expanded over the course of a few weeks with a high dose of IL-2 in 24- well plates. Selected TIL lines that presented best tumor reactivity are then further expanded in a "rapid expansion protocol" (REP), which uses anti-CD3 activation for a typical period of two weeks. The final post-REP TIL is infused back into a patient in order to treat a tumor of the patient. This applies mutatis mutandis to adoptive NK-cell transfer with TILs.
In accordance with a further preferred embodiment of the invention the anti-tumor lymphocytes are autologous anti-tumor lymphocytes.
In an anti-tumor therapy with autologous lymphocytes the lymphocytes are taken from a subject having a tumor and are genetically engineered (e.g. to produce CAR T-cells) and/or selected and/or expanded (e.g. to produce TILs) ex vivo and then transferred back into the same subject. These autologous therapies are subject-specific because the therapeutic cells are created from a subject's own cells.
The present invention relates in a fifth aspect to a molecular complex comprising the variant of the GAIN domain of the first aspect or the AGPCR of the second aspect in association with a Stachel
sequence or fragment thereof being capable of binding to the GAIN domain of claim 1 , wherein the Stachel sequence preferably comprises or consists of any one SEQ ID NOs 59 to 81 or a sequence being at least 70% identical, preferably at least 80% and most preferably at least 90% identical thereto.
As discussed herein above, the variant of the GAIN domain or the AGPCR are capable of binding to each other via non-covalent interactions, thereby forming a molecular complex. Such molecular complexes are the subject of the sixth aspect of the present invention. The molecular complexes may be formed, for example, within a cell and also within in a suitable medium.
The amino acid sequence sharing at least 70% sequence identity displays with increasing preference at least 75% sequence identity, at least 80% sequence identity, at least 85% sequence identity, at least 90% sequence identity, at least 95% sequence identity, at least 96% sequence identity, at least 97% sequence identity, at least 98% sequence identity and at least 99% sequence identity.
In the molecular complex the variant of the GAIN domain or the AGPCR on the one hand and the Stachel sequence on the other hand are preferably derived from the same AGPCRS. For instance, in the case of the variant of the GAIN domain item (I) of the first aspect the Stachel sequence is preferably a Stachel sequence of SEQ ID NO: 59 or a sequence being at least 70% identical, preferably at least 80% and most preferably at least 90% identical thereto, in in the case of the variant of the GAIN domain item (I) of the first aspect the Stachel sequence is preferably a Stachel sequence of SEQ ID NO: 60 or a sequence being at least 70% identical, preferably at least 80% and most preferably at least 90% identical thereto., etc.
The present invention relates in a sixth aspect to a composition, preferably a diagnostic or pharmaceutical composition, or a kit comprising the variant of thee GAIN domain, the AGPCR, the nucleic acid molecule, the vector, the host cell of the invention or a combination thereof.
In accordance with the present invention, the term “pharmaceutical composition” relates to a composition for administration to a patient, preferably a human patient. The pharmaceutical composition of the invention comprises the compounds recited above. It may, optionally, comprise further molecules capable of altering the characteristics of the compounds of the invention thereby, for example, stabilizing, modulating and/or activating their function. The composition may be in solid, liquid or gaseous form and may be, inter alia, in the form of (a) powder(s), (a) tablet(s), (a) solution(s) or (an) aerosol(s). The pharmaceutical composition of the present invention may, optionally and additionally, comprise a pharmaceutically acceptable carrier. Examples of suitable pharmaceutical carriers are well known in the art and include phosphate buffered saline solutions, water, emulsions, such as oil/water emulsions, various types of wetting agents, sterile solutions, organic solvents including DMSO etc. Compositions comprising such carriers can be formulated by well-known conventional methods. These pharmaceutical compositions can be administered to the subject at a
suitable dose. The dosage regimen will be determined by the attending physician and clinical factors. As is well known in the medical arts, dosages for any one patient depends upon many factors, including the patient's size, body surface area, age, the particular compound to be administered, sex, time and route of administration, general health, and other drugs being administered concurrently. The therapeutically effective amount for a given situation will readily be determined by routine experimentation and is within the skills and judgement of the ordinary clinician or physician. Generally, the regimen as a regular administration of the pharmaceutical composition should be in the range of 1 pg to 5 g units per day. However, a more preferred dosage might be in the range of 0.01 mg to 100 mg, even more preferably 0.01 mg to 50 mg and most preferably 0.01 mg to 10 mg per day. The length of treatment needed to observe changes and the interval following treatment for responses to occur vary depending on the desired effect. The particular amounts may be determined by conventional tests which are well known to the person skilled in the art.
Also the term “diagnostic composition” relates to a composition, optionally for administration to a patient, preferably a human patient and may comprise the essentially same additional compounds as discussed in connection with the pharmaceutical composition. While pharmaceutical compositions are to cure or prevent a disease, a diagnostic composition is to identify the presence and optionally also the site of a disease in a subject. The diagnostic composition preferably comprises a diagnostically active compound as discussed herein above, such as a radiolabel or a fluorophore.
The various components of the kit may be packaged into one or more containers such as one or more vials. The vials may, in addition to the components, comprise preservatives or buffers for storage. The kit may comprise instructions how to use the kit, which preferably inform how to use the components of the kit for diagnosing a tumor and/or for grading a tumor and/or for tumor prognosis.
The present invention relates in a seventh aspect to a computer-implemented method for increasing or decreasing the association of the GAIN domain of an AGPCR and the Stachel sequence under force comprising (a) performing in silico steered molecular dynamics (SMD) simulations on the GAIN domain in association with the Stachel sequence whereby the GAIN domain undergoes topological rearrangements that enable force to propagate through the GAIN domain to the associated Stachel sequence, (b) identifying force propagation pathways from the SMD simulations of (a) by analysing the correlations between alpha carbon motions within the GAIN domain, (c) identifying the arrival angles of the optimal force propagation pathways as obtained in (b) on the Stachel sequence, (d) mutating in silico based on the information as obtained in (a) to (c) one or more pairs of amino acids of the GAIN domain outside of the interface of the GAIN domain that associates with the Stachel sequence to cysteine residues and/or substituting in silico based on the information as obtained in (a) to (c) one or more amino acids, preferably between one and ten amino acids by other amino acids outside of the interface of the GAIN domain that associates with the Stachel sequence to obtain a variant of the GAIN domain with a disulfide bond that is expected to decrease or increase the arrival angles as compared to the arrival angles as identified in (c), (e) repeating steps (a) to (c) on the variant of the
GAIN domain as obtained in (d) thereby identifying the arrival angles of the optimal force propagation pathways on the Stachel sequence for the variant of the GAIN domain in (d), wherein increased arrival angles are indicative for an increased strength of the association of the GAIN domain and the Stachel sequence under force and decreased arrival angles are indicative for a decreased strength of the association of the GAIN domain and the Stachel sequence under force.
A preferred embodiment of the seventh aspect of the invention is directed to a computer-implemented method for increasing or decreasing the association of the GAIN domain of an AGPCR and the Stachel sequence under force comprising (a) performing in silico steered molecular dynamics (SMD) simulations on the GAIN domain in association with the Stachel sequence whereby the GAIN domain undergoes topological rearrangements that enable force to propagate through the GAIN domain to the associated Stachel sequence, (b) identifying force propagation pathways from the SMD simulations of (a) by analysing the correlations between alpha carbon motions within the GAIN domain, (c) identifying the arrival angles of the optimal force propagation pathways as obtained in (b) on the Stachel sequence, (d) mutating in silico based on the information as obtained in (a) to (c) one or more pairs of amino acids of the GAIN domain outside of the interface of the GAIN domain that associates with the Stachel sequence to cysteine residues to obtain a variant of the GAIN domain with a disulfide bond that are expected to decrease or increase the arrival angles as compared to the arrival angles as identified in (c), (e) repeating steps (a) to (c) on the variant of the GAIN domain as obtained in (d) thereby identifying the arrival angles of the optimal force propagation pathways on the Stachel sequence for the variant of the GAIN domain in (d), wherein increased arrival angles are indicative for an increased strength of the association of the GAIN domain and the Stachel sequence under force and decreased arrival angles are indicative for a decreased strength of the association of the GAIN domain and the Stachel sequence under force.
The amino acid substitutions are preferably non-covalent amino acid substitutions and/or the substitution of amino acid(s) other than cysteine(s) by amino acid(s) other than cysteine(s). These amino acid(s) is/are more preferably the naturally-occurring proteinaceous amino acid(s) other than cysteine(s)
A computer-implemented method as used herein designates a method which involves the use of a computer, computer network or other programmable apparatus, where one or more features are realised wholly or partly by means of a computer program. Since the method is a computer- implemented method, also a computer-implemented program is described herein that when being executed on a computer causes the computer to carry out the method of the first aspect of the invention.
The term “in silico" as used herein refers to steps of the method of the invention that are performed on or with the aid of a computer or via computer simulation. The phrase is pseudo-Latin for 'in silicon' (in Latin it would be in silicic), referring to silicon in computer chips.
According to step (a) of the method of the invention in silico steered molecular dynamics (SMD) simulations are performed on the GAIN domain in association with the Stachel sequence whereby the GAIN domain undergoes topological rearrangements that enable force to propagate through the GAIN domain to the associated Stachel sequence.
Regarding the feature “GAIN domain in association with the Stachel sequence” it is noted that the Stachel is a structural part of the GAIN domain. When the GAIN domain is in association with the Stachel sequence the GAIN domain is folded, so that it is association with the GAIN domain. When the Stachel dissociates, the GAIN partially unfolds. The SMD simulations are performed starting from the Stachel bound state.
Steered molecular dynamics (SMD) is art established and is issued to induce unbinding of ligands and conformational changes in biomolecules on time scales accessible to molecular dynamics simulations. Time-dependent external forces are applied to a system, and the responses of the system are analysed (Izrailev et al. (1999), Computational Molecular Dynamics: Challenges, Methods, Ideas pp 39-65, ISBN 978-3-540-63242-9). Also in silico SMD simulations are art established and were, for example, used to study the conformational lock of Cu, Zn-superoxide dismutase (Xiao (2019), Scientific Reports, 9: Article number 4353). The basic idea behind any SMD simulation is to apply an external force to one or more atoms (e.g. alpha carbons), which can be referred as SMD atoms and study the behaviour of the other atoms in the system that are linked to the SMD atoms under various conditions (https://www.ks.uiuc.edu/Training/Tutorials/namd/namd-tutorial-html/node19.html).
During the SMD simulation of step (a) the GAIN domain undergoes topological rearrangements that enable force to propagate through pathways favoring its mechanical resistance.
The SMD simulation of step (a) is preferably essentially performed as described in the below Example 3, Section “SMD simulations”. Hence, SMD simulations are preferably performed with GROMACS 2019.4 (https://manual.gromacs.Org/documentation/2019.4/download.html) using the CHARMM36 forcefield (https://www.charmm.org/archive/charmm/) along with the TIP3 (transferable intermolecular potential with 3 points, which is implemented into CHARMM36) water model in an NPT ensemble (also called isothermal-isobaric ensemble; constant temperature and constant pressure ensemble) at about 31 OK (about 36,85°C) and about 1 bar using a velocity rescaling thermostat (with a relaxation time of 0.1 ps) and Berendsen barostat (with a relaxation time of I ps). Equations of motion are integrated with a timestep of about 2 fs using a leap-frog algorithm (e.g. https://www.compphy.com/leap-frog/).
Each system is preferably energy minimized using a steepest descent algorithm (https://people.seas.harvard.edu/~yaron/AM221-S16/lecture_notes/AM221_lecture10.pdf), more preferably for about 5000 steps, and then equilibrated with the atoms of the protein restrained using a harmonic restraining force in two steps, first by adding the thermostat for about 50000 steps and then
the barostat for about 500000 steps.
Also SMD simulations are preferably performed using the umbrella sampling method (e.g. http://jeti.uni-freiburg.de/studenten_seminar/term_paper_WS_16_17/Kaestner11 .pdf) with a pulling speed ranging from about 1 A/ns to about 10 A/ns.
In SMD simulations, preferably the 2 alpha carbons of the N-term and C-term residues are moved harmonically in the desired direction with constant velocity. A distance cutoff of about 14 A is then used for short-range non-bonded interactions, whereas electrostatic interactions are computed using the Particle-mesh Ewald method with a cubic interpolation of power 4.
According to step (b) of the method of the invention force propagation pathways from the SMD simulations of (a) are identified by analysing the correlations between alpha carbon motions within the GAIN domain.
When force is propagated through a molecule, soft degrees of freedom will be stretched out along the path of force propagation. These paths of force propagation are identified according to step (b).
To identify these force propagation pathways, preferably the coordinates of a-carbons are extracted from the time window corresponding to the GAIN “loaded” state (i.e. the GAIN domain in association with the Stachel sequence). Their sequential motions between two consecutive frames of the simulation are then computed and correlation analyses is performed to reconstruct the force propagation pathways from these motions. Pathways are identified by following the route from the N- to C-terminal that maximizes the observed correlations. Finally and preferably, identified pathways between consecutive frames are clustered and a deep analysis of cluster populations enables to discern optimal (primary cluster) from suboptimal pathways (secondary clusters).
According to step (c) of the method of the invention the arrival angles of the optimal force propagation pathways as obtained in (b) on the Stachel sequence are identified.
For the force propagation pathways parameters such as the amplitudes, path-lengths or arrival angles can be determined. According to step (c) the arrival angles are determined.
The arrival angles indicate the orientation of feree on the Stachel mechanical clamp between the GAIN domain and the Stachel sequence. The arrival angles are preferably defined as the angles between the force pathway vector and a Stachel network plane.
According to step (d) of the method of the invention based on the information as obtained in (a) to (c) one or more pairs of amino acids of the GAIN domain outside of the interface of the GAIN domain that associates with the Stachel sequence are mutated in silico to cysteine residues to obtain a variant of
the GAIN domain with a disulfide bond and/or are substituted in silico to other amino acid(s) at to obtain a variant of the GAIN domain with specific amin acid substitution(s) selected site(s) that are expected to decrease or increase the arrival angles as compared to the arrival angles as identified in (c).
The one or more pairs of amino acids of the GAIN domain enabling the formation of disulfides are preferably identified by using the Coupled Moves algorithm from the Rosetta software with the backrub movers protocol (Ollikainen et al. (2015). PLoS computational biology 11 , e1004335 (2015)). Coupled Moves is a flexible backbone design method meant to be used for designing small-molecule binding sites, protein-protein interfaces, and protein-peptide binding sites (https://www.rosettacommons.org/docs/latest/coupled-moves). The Backrub application is useful for creating ensembles of protein backbones, modeling protein flexibility, modeling mutations, and detailed refinement of backbone/side chain conformations. The backrub algorithm rotates local segments of the protein backbone as a rigid body about an axis defined by the starting and ending atoms of the segment. Atoms branching of the main chain at the pivot points (side chains, hydrogens, carbonyl oxygens), are updated to minimize the bond angle strain incurred. These moves are accepted or rejected using the Metropolis criterion (https://www.rosettacommons.org/docs/latest/application_documentation/structure_prediction/backru).
It is furthermore preferred that a total of about 100 independent structures is generated for every pair of cysteines and/or every amino acid substitution or set of amino acid substitutions, and the about top 10% lowest energy structures are conserved. All structures are then minimized and relaxed to access lowest energy potential through the relax protocols included in the Rosetta package. Distances between the groups of added cysteines are computed and averaged across the lowest energy structures. To determine which pairs of mutations form a disulfide, a distance cutoff of about 2.15 A between the 2 cysteines is preferably used to conserve or exclude mutants. Similarly, the lowest energy structures can be computed for every amino acid substitution or set of amino acid substitutions.
The term “about” as used herein is for each occurrence independently with increasing preference ± 30%, ± 20%, ± 10% and ± 5%.
According to final step (e) of the method of the invention, steps (a) to (c) are repeated on the variant of the GAIN domain as obtained in (d) thereby identifying the arrival angles of the optimal force propagation pathways on the Stachel sequence for the variant of the GAIN domain in (d), wherein increased arrival angles are indicative for an increased strength of the association of the GAIN domain and the Stachel sequence under force and decreased arrival angles are indicative for a decreased strength of the association of the GAIN domain and the Stachel sequence under force.
According to be best knowledge of the inventors the method of the seventh aspect is advantageously the first computational method integrating dynamics under mechanical stress and systematic calculation of deformation upon sequence variation that can be applied for rationally designing mechanosensing properties. In addition, AGPCRs play a very important role as drug target. The majority of approved pharmaceutics targets GPCRs and AGPCRs are second largest group of GPCRs. This underlines the important contribution of the method of the seventh aspect to the prior art.
In accordance with a preferred embodiment of the seventh aspect of the invention in step (a) two or more different pulling speeds are applied to the GAIN domain, wherein the pulling speeds are preferably selected from about 0.025 nm.ns-1, 0.1 nm.ns-1 and about 1.0 nm.ns-1.
In SMD simulations, the GAIN domain can be stretched through constant velocity pulling. In this type of simulation, the SMD atom is attached to a dummy atom via a virtual spring. This dummy atom is moved at a constant pulling speed and then the force between both is measured.
This preferred embodiment results in an increase of the forces that can be accumulated within the SMD system that is commonly defined as the loaded state of the protein (i.e. GAIN domain in association with the Stachel sequence). Hence, the appliance of two or more different pulling speeds provides the technical advantage of obtaining more information during the SMS simulations that can be used in the subsequent steps of the method of the invention. Thereby a better understanding of how force propagates through the GAIN domain to the Stachel peptide. The two pulling speeds as used in the appended examples are 0.1 nm.ns-1 and 1.0 nm.ns-1.
In accordance with further a preferred embodiment of the seventh aspect of the invention in step (a) a harmonic force potential is applied to the GAIN domain by pulling simultaneously from the N- and C- terminus of the GAIN domain.
The stretching between two atoms may be defined by the linear harmonic potential, which describes the energy associated with deviations from the equilibrium distance. Pulling simultaneously from the N- and C-terminus of the GAIN domain determines the direction of the force being applied and ensures that the force can propagate throughout the GAIN domain to the Stachel sequence.
In accordance with a preferred embodiment of the seventh aspect in step (b) for every frame the three-dimensional coordinates for every alpha carbon of the GAIN domain are extracted, for every pair of two consecutive frames of the simulation the distance displacements of the three-dimensional coordinates are obtained, and the distance displacements are correlated to reconstruct the optimal and suboptimal force propagation pathways from the N- to the C-terminus and computing force propagation pathways.
The above preferred embodiment describes the preferred determination of force propagation
pathways.
SMD simulations simulate the evolution and motions of a biomolecular system as a function of time. Motions are assessed through changes in atomic positions recorded at discrete time steps called “frames”. Consecutive frames correspond to consecutive time steps of a few picoseconds.
Even more preferably force propagation pathways are essentially determined as described in Example 3, section “SMD simulation analysis and force propagation pathway identification”. Hence, to extract force propagation pathways from the SMD simulations, correlation between alpha carbon motions are more preferably analyzed. For every frame of the simulation, the [x,y,z] cartesian coordinates of every alpha carbon are extracted. Displacement of these coordinates between 2 consecutive frames are computed and represented as a vector [xo-xi, yo-yi, zo-zi].
To establish a correlation between these displacement vectors, distance correlation most preferably (Equation 1) used:
For a simulation of t frames, this resulted in t correlation matrices of [nresidues, nresidues]. To identify force propagation pathways, Dijkstra shortest path algorithm (e.g. https://www.geeksforgeeks.org/dijkstras- shortest-path-algorithm-greedy-algo-7/) is most preferably used. The shortest path maximizing the total correlation is searched between the N-term of the GAIN domain and the C-term of the Stachel mechanical clamp for every correlation matrix, resulting in / shortest force propagation pathways. To extract relevant information from these pathways, k-means clustering was performed. The optimal number of clusters to use for independent simulations is determined using the inertia attribute to identify the sum of squared distances of samples to the nearest cluster center. Centroids of clusters, corresponding to the optimal force propagation pathway within this cluster were used for later extraction of features.
In accordance with a preferred embodiment of the seventh aspect in step (c) the arrival angles are calculated with respect to a plane that is formed between the N-terminal and C-terminal alpha carbon of the Stachel sequence and a selected alpha carbon in the middle of the Stachel sequence.
As explained above, the angles of the force propagation pathways arrival on the Stachel sequence are useful for predicting the effect of the different Cys-Cys mutants and/or non-covalent amino acid substitutions in-silico. Here, it is helpful to define a fixed plane in Stachel sequence for the calculation of the angles. This fixed plane is in accordance with the above preferred embodiment formed between the N-terminal and C-terminal alpha carbon of the Stachel sequence and a selected alpha carbon in the middle of the Stachel sequence.
Regarding the feature “middle of the Stachel sequence” it is noted that the Stachel sequences are highly conserved among the 33 AGPCRs. The Stachel sequences of AGPCRs have a length of 7, 8 or 9 amino acids and the “middle of the Stachel sequence” is in the case of 7 amino acids position 3-5 and preferably position 4, in the case of 8 amino acids position 4 or 5 in the case of 9 amino acids position 4-6 and, preferably position 5.
The arrival angles are even more preferably essentially determined as described in Example 3, “force propagation pathway projection on Stachel rupture interface". Hence, more preferably in order to extract the angle of pathway projection on the rupture interface, coordinates of all alpha carbons are extracted. These coordinates represent a [f, nresidues, nresidues] matrix, t being the number of frames of the simulation. For each of this coordinate matrix, a plane containing both the Stachel N-term and C-term alpha carbons, with the alpha carbon of the middle of the Stachel sequence is constructed. This results in an ensemble of dynamic planes moving according to local structural rearrangements and corresponding to the rupture interface of the Stachel mechanical clamp. To estimate the angle of force arrival on this rupture interface defined by these different planes, vectors representing the jump of pathway from the GAIN domain onto the Stachel mechanical clamp are extracted. These propagation vectors are computed from the alpha carbon coordinates of the pair of residues involved in the jump of force onto the Stachel mechanical clamp. Angles between the plane normal and the force arrival were computed for every frame corresponding to the loaded state of the system. Finally, angles are preferably concatenated through simulation replicates, and preferably also violin plots may be used to show the distribution of angles across mutants.
In accordance with a preferred embodiment of the seventh aspect the method comprises in addition the identification of the backbone H-bonds that ensure or determine the mechanical resistance of the Stachel mechanical clamp to force.
To identify backbone H-bonds that ensure the mechanical resistance of the Stachel mechanical clamp, preferably the Wernet-Nilsson algorithm (see reference 8 of Example 3) is used. Identified H- bonds across the simulation may be 2D-projected in a plane perpendicular to the Stachel axis defined by its N-term and C-term residues. To highlight the H-bonds rearrangements occurring during the pulling, the 2D projections can be compared between the native unloaded structure and the loaded structure corresponding to the highest mechanical resistance point observed in the SMD simulations. Projections may be averaged across replicates and plotted.
In accordance with a further preferred embodiment of the seventh aspect the method further comprises producing the variant of the GAIN domain as obtained after steps (a) to (d).
As also discussed herein above, so far all steps of the claimed method were carried out in silico and the final “product” is a modelled variant of the GAIN domain as obtained after steps (a) to (d) that either displays increased arrival angles being indicative for an increased strength of the association of
the GAIN domain and the Stachel sequence under force or displays decreased arrival angles being indicative for a decreased strength of the association of the GAIN domain and the Stachel sequence under force.
In accordance with the above preferred embodiment variant of the GAIN domain is actually produced by peptide and/or protein synthesis or by site-directed mutagenesis.
Means and methods for protein and peptide synthesis are known in the art; see, for example, Albericio and Govender (2012), Special Issue "Chemical Protein and Peptide Synthesis", ISSN 1420-3049.
It is also possible to introduce selected amino acid substitutions into naturally occurring proteins and peptides (site-directed mutagenesis), so that the naturally occurring proteins and peptides are engineered to proteins and peptides being based on the naturally occurring proteins and peptides that comprise the engineered Cys-Cys bonds and/or non-covalent amino acid substitutions as discussed herein above.
As discussed above, the variant of the GAIN domain may be implemented in an AGPCR. In accordance the above preferred embodiment the variant of the GAIN domain may also be produced as part of a full-length AGPCR or a fragment thereof comprising the variant of the GAIN domain.
The produced products may be C-terminally amidated in order to avoid unwanted charge effects of the carboxy terminus in further experiments or uses.
In accordance with a more preferred embodiment of the seventh aspect the method further comprises applying single-molecule force spectrometry (SMFS) to the association of the produced variant of the GAIN domain and the Stachel sequence in order to identify the interface of variant of the GAIN domain that associates with the Stachel sequence and/or the force that disrupts the association, wherein the SMFS is preferably atomic force microscopy (AFM) based SMFS.
The SMSF is not part of the computational method. It is used independently to determine experimentally the force required to dissociate the Stachel and validate the computational predictions. The interface of the variants of the GAIN domain that associates with the Stachel sequence is determined from the X-ray structures of the variants of GAIN domain.
Single-molecule force spectroscopy has emerged as a powerful tool to investigate the forces and motions associated with biological molecules (Neuman, and Nagy, Nature Methods volume 5, pages491-505 (2008). SMFS is meanwhile a well-established method that directly probes structural changes of macromolecules under the influence of mechanical force. Since mechanical forces are ubiquitous in biology, insights gleaned from SMFS experiments shed light onto fundamentally important molecular mechanisms by which biological systems are able to sense, transduce and
generate mechanical forces in vivo. The most common force spectroscopy techniques are optical tweezers, magnetic tweezers and atomic force microscopy. Herein, atomic force microscopy is preferred
As is illustrated in the examples herein below the results of the SMSF confirmed that the computational predictions are correct. As predicted, a variant of GAIN domain for which increased arrival angles were computed show an increased strength of the association of the GAIN domain and the Stachel sequence under force and a variant of GAIN domain for which decreased arrival angles were computed show a decreased strength of the association of the GAIN domain and the Stachel sequence under force.
In accordance with an even more preferred embodiment in the SMFS the variant of the GAIN-domain is fused to a Fgp domain and a Spytag whereby the Fgp domain connects the variant of the GAIN domain via an isopeptide bond to a surface, preferably a coverglass surface and the Fgp domain connects the variant of the GAIN domain to a cantilever.
According to this preferred embodiment the SMFS is essentially as is shown in Fig. 5 and Example 3. Such and similar method are art-established (Yang et al. Front Mol Biosci. 2020; 7: 85.)
In accordance with an additional more preferred embodiment between the isopeptide bond and the surface a fingerprint domain, preferably a FLN4 domain is located.
A fingerprint domain, in particular a FLN4 domain is advantageous for filtering the obtained data as is described in detail in Example 3, section AFM-SMFS measurement and data analysis.
Here, the SMFS data traces were filtered by searching for contour length increments that matched the lengths of the fingerprint domains, FLN (about 36 nm). Theoretical contour length increment was calculated based on the equation ALc = (0.365 nm/AA) x (# AAs in POI) - Lf, where ALc is the expected contour length increment and Lf is the end-to-end length of folded protein domain. For FLN, ALc = 36.9 nm - Lf, where Lf is typically < 5nm. The contour length increment is between first FLN unfolding and final stachel dissociation (Fig. 7).
In accordance with an even more preferred embodiment two different, preferably three different and most preferably four different pulling speeds are applied in the SMFS, wherein the different pulling speeds are preferably selected from about 400, about 800, about 1600 and about 3200 nm.s-1.
The appliance of two different, preferably three different and most preferably four different pulling speeds improved the accuracy of the SMFS since forces over different times are thereby applied. In the appended examples pulling speeds of 400, about 800, about 1600 and about 3200 nm.s-1 are used.
In accordance with another more preferred embodiment the method further comprises measuring the force that disrupts the association between the variant of the GAIN domain and the Stachel sequence, preferably by an serum response element (SRE) assay wherein shear stress, preferably through vibration or spinning is applied to cells that express an aGPCR comprising the variant of the GAIN domain, wherein the aGPCRs on the cells are coupled via aGPCR ligands, preferably collagen III or transglutaminase 2 onto a plate.
SMFS is not the only method measuring the force that disrupts the association between the variant of the GAIN domain and the Stachel sequence and many other methods are available as well.
One preferred example is a serum response element (SRE) assay wherein shear stress, preferably through vibration or spinning is applied to cells that express an aGPCR comprising the variant of the GAIN domain, wherein the aGPCRs on the cells are coupled via aGPCR ligands, preferably collagen III or transglutaminase 2 onto a plate.
An example of serum response element (SRE) assay is described as “vibration assay” in the Fig. 4 and the examples. Also results of the vibration assay were in agreement with the computational predictions.
As regards the embodiments characterized in this specification, in particular in the claims, it is intended that each embodiment mentioned in a dependent claim is combined with each embodiment of each claim (independent or dependent) said dependent claim depends from. For example, in case of an independent claim 1 reciting 3 alternatives A, B and C, a dependent claim 2 reciting 3 alternatives D, E and F and a claim 3 depending from claims 1 and 2 and reciting 3 alternatives G, H and I, it is to be understood that the specification unambiguously discloses embodiments corresponding to combinations A, D, G; A, D, H; A, D, I; A, E, G; A, E, H; A, E, I; A, F, G; A, F, H; A, F, I; B, D, G; B, D, H; B, D, I; B, E, G; B, E, H; B, E, I; B, F, G; B, F, H; B, F, I; C, D, G; C, D, H; C, D, I; C, E, G; C, E, H; C, E, I; C, F, G; C, F, H; C, F, I, unless specifically mentioned otherwise.
Similarly, and also in those cases where independent and/or dependent claims do not recite alternatives, it is understood that if dependent claims refer back to a plurality of preceding claims, any combination of subject-matter covered thereby is considered to be explicitly disclosed. For example, in case of an independent claim 1 , a dependent claim 2 referring back to claim 1 , and a dependent claim 3 referring back to both claims 2 and 1 , it follows that the combination of the subject-matter of claims 3 and 1 is clearly and unambiguously disclosed as is the combination of the subject-matter of claims 3, 2 and 1. In case a further dependent claim 4 is present which refers to any one of claims 1 to 3, it follows that the combination of the subject-matter of claims 4 and 1 , of claims 4, 2 and 1 , of claims 4, 3 and 1 , as well as of claims 4, 3, 2 and 1 is clearly and unambiguously disclosed.
The Figures show.
Fig. 1. The adhesion GPCR GAIN domain is a bona fide mechanosensing protein domain which dissociates from the Stachel peptide in a single concerted shearing event (A) General framework for the study of the ADGRG1 GAIN domain under force. Representation of membrane- embedded full-length human ADGRG1 adapted from an AlphaFold structure prediction (UniProt Q9Y653), composed of the 7TM (blue), GAIN (grey, with Stachel in cyan) and PLL (magenta) domains (left panel). Combination of experimental and computational techniques used to probe the behavior of WT and mutant GAIN domains under force (right insets). (B) Structural view of the Stachel backbone (cyan) H-bond network (dotted magenta lines) formed with adjacent p-strands b6 and b8 in the human ADGRG1 GAIN domain (see also Fig. 3A). The right-most panel shows a view from C- terminal end of the Stachel. (C) Representative force-extension trace of the WT ADGRG1 GAIN domain showing clear ddFLN4 unfolding events and a single GAIN-Stachel dissociation event at ~200 nm extension. (D) Dynamic force spectrum of the GAIN domain at cantilever retraction velocities of 0.4 pm. s-1, 0.8 pm.s'1, 1.6 pm.s-1 and 3.2 pm.s-1, with corresponding rupture-force probability density histograms for each velocity projected onto individual axes on the right. A Gaussian fit (dashed line) with the fit parameters are shown for each distribution. (E) Human ADGRG1 GAIN domain position within a spectrum of mechanical proteins with varied mechanical stabilities. Force values are only indicative and given as broad reference points or ranges corresponding to the measured force at which a given protein or protein-protein interaction undergoes a major conformational transition important for its function. These values can vary depending on the experimental technique and setup used, the loading rate applied, the ligand used (if any) and the exact protein type and subdomain considered (especially for the tandem repeat proteins such as titin).
Fig. 2. SMD simulations provide metrics to characterize and redesign the structural and dynamic determinants of GAIN mechanical resistance. (A) SMD simulation and data analysis pipeline used to obtain a mechanistic understanding of GAIN domain behavior under force (see methods in Example 3). (B-E) Violin plots of the distributions of the average angles between force pathways and the Stachel hydrogen bond network in the loaded state from SMD simulations at pulling speeds of 0.1 and 1.0 nm.ns-1, forWT, SS_h2l10, SS_h1 h2 and SS_h1 h2-h2l10 GAIN domains.
Fig. 3. Modulating the GAIN domain force dissociation threshold through engineered disulfides. (A) Topology chart of the GAIN domain with the positions of disulfide-forming cysteine mutations shown as red dots and dotted lines, a-helices (h) and p-strands (b) are numbered and define the naming nomenclature of the mutants. Loop 10 (110) is shown in purple and connects p- strands 8 and 9. The a-helical portion of the GAIN is colored green and p-strands are colored blue or orange according to the p-sheet they belong to. The C-terminal Stachel is shown in cyan and the autoproteolysis site is marked by a green asterisk. Backbone-backbone H-bonds between the Stachel and adjacent p-strands 6 and 8 are represented as dashed magenta lines (see also Fig. 1 B). (B) Crystal structure of the SS-h1 h2-h2l10 GAIN domain variant at 1.9 A resolution, following the coloring
scheme in Panel A. The right panel shows the transparent GAIN surface (grey) and Stachel surface (cyan). (C) GAIN variant rupture-force probability density histograms at the indicated cantilever retraction velocities overlaid on the corresponding WT data in transparent red. A Gaussian fit is shown for each distribution. (D) Cell-surface expression of APLL ADGRG1 constructs determined by ELISA using an anti-AVI antibody (top panel, data normalized to empty vector). Constitutive activity of APLL ADGRG1 constructs measured using an SRE assay (bottom panel, data normalized to empty vector). (E) 2D scatter plot showing the differences in mechanical and thermodynamic stabilities between variants relative to WT. Data is presented as the difference to WT for each measurement. A single average and standard deviation (n = 3) value is plotted for the thermodynamic stability obtained from circular dichroism measurements. The median of the rupture-force probability densities of each variant minus WT is plotted for each cantilever retraction velocity shown in Panel C, giving four values of mechanical stability for each variant.
Fig. 4. Contrasted effects of the designed disulfides on the activation of full-length ADGRG1 expressed on cells and stimulated by vibration and receptor-specific ligands. A. static conditions. B. vibration conditions. C. signaling differences between vibration and static conditions. D. Schematic description of the cellular assay measuring the mechano-induced activation of cells expressing ADGRG1 .
Fig. 5. Schematic of the cantilever and coverglass surface modification and protein immobilization for AFM-SMFS experiments (see Example 3).
Fig 6. AFM-SMFS of GAIN variants for contour length analysis. (A) Schematic illustration of the experimental setup for comparative AFM-SMFS of GAIN-WT and GAIN(4Cys). GAIN-WT and GAIN(4Cys) were immobilized on different spots on the surface. SdrG was immobilized at the cantilever tip. The same cantilever was used to probe each sample at 3200 nm s-1. (B) Contour length of first FLN unfolding and final stachel dissociation using GAIN-WT and GAIN(4Cys) with the same cantilever tip. (C) Contour length increment between first FLN unfolding and final stachel dissociation using GAIN-WT and GAIN(4Cys) with the same cantilever tip. (D) Contour length of first FLN unfolding and final stachel dissociation using GAIN variants. (E) Contour length increment between first FLN unfolding and final Fgp-SdrG dissociation using GAIN variants. (F) Contour length histogram of GAIN-WT (n=201), GAIN(2Cys) (n=476), GAIN(2Cys)-h1 h2 (n=96), and GAIN(4Cys) (n=71).
Fig 7. Snapshots from SMD simulations at 1 nm.ns'1 showing force propagation pathways (red) represented on the native and loaded state structures. Residues involved in each pathway are listed and colored following the color code of the 2D topology chart in Figure 2. The average angle of the force pathway vector on the Stachel mechanical clamp is indicated for each GAIN variant for reference.
Fig 8. Thermodynamic stability of purified GAIN domain variants measured by circular dichroism spectroscopy. CD melting curves from 20°C to 96°C at 216 nm and 232 nm are shown for WT and the three GAIN domain disulfide variants.
Fig 9. GAIN domain purification for AFM and X-ray crystallography. (A) Gel-filtration (Superdex 200 Increase 10/300 GL) elution profiles for representative purifications of the WT GAIN domain construct for AFM experiments and of the SS_h1 h2-h2l10 GAIN domain construct for X-ray crystallography (suffix “X”). Protein constructs for X-ray crystallography lack the N- and C-terminal AFM handles (see full protein sequences). (B) SDS-PAGE analysis of the purified WT and variant GAIN domains and of the SS_h1 h2-h2l10_X GAIN domain construct. The lower molecular weight bands correspond to the cleaved C-terminal containing the Stachel and tags and/or handles.
Fig 10. AFM-SMFS of GAIN-WT and NoGAIN in native or reducing conditions for contour length analysis. (A) Dynamic force spectra of GAIN-WT for stachel dissociation force according to the loading rate in reducing (100 mM DTT) and non-reducing condition (0 mM DTT). Black circles represent the median rupture force/loading rate in non-reducing condition and purple circles represent in reducing condition at each pulling speed of 400, 800, 1600, and 3200 nm s-1. Error bars are ±1 s.d. Solid lines are least square fits to the Bell-Evans model. (B) Contour length of first FLN unfolding and final stachel dissociation using GAIN-WT in reducing (100 mM DTT) and non-reducing condition (0 mM DTT). (C) Contour length increment between first FLN unfolding and final stachel dissociation using GAIN-WT in reducing (100 mM DTT) and non-reducing condition (0 mM DTT). (D) Schematic illustration of experimental configuration for NoGain. (E) Contour length of first FLN unfolding and final Fgp-SdrG dissociation using NoGain in reducing (100 mM DTT) and non-reducing condition (0 mM DTT). (F) Contour length increment between first FLN unfolding and final Fgp-SdrG dissociation using NoGain in reducing (100 mM DTT) and non-reducing condition (0 mM DTT). Black circles represent the median rupture or unfolding force/contour length or contour length increment in non-reducing condition and purple circles represent in reducing condition. Error bars are ±1 s.d.
Fig 11. H1L7 force propagation angle. The H1 L7 force propagation angle is very similar to that of the H1 H2 variant. Consistent with these predictions, the H1 L7 variant is mechanically weaker than WT and behaves similarly to the H1 H2 variant as measured by AFM.
Fig 12. AFM - non-covalent designs. Dynamic force spectrum of the GAIN domain at cantilever retraction velocities of 0.4 pm.s'1, 0.8 pm.s'1, 1.6 pm.s-1 and 3.2 pm.s-1
Fig 13. AFM - non-covalent designs. Distribution & Median Value of Dissociation Forces. Ruptureforce probability density histograms for velocities of 0.4 pm.s-1, 0.8 pm.s-1, 1.6 pm.s-1 and 3.2 pm.s-1 projected onto individual axes.
Fig 14. AFM - Disulfide designs. Dynamic force spectrum of the GAIN domain at cantilever retraction velocities of 0.4 pm.s-1, 0.8 pm.s-1, 1.6 pm.s-1 and 3.2 pm.s-1.
Fig 15. AFM - Disulfide designs. Distribution & Median Value of Dissociation Forces. Rupture-force probability density histograms for velocities of 0.4 pm.s-1, 0.8 pm.s-1, 1.6 pm.s-1 and 3.2 pm.s-1 projected onto individual axes.
Fig 16. Constitutive activity SRE assays. Constitutive activities measured using the SRE assay for GAIN variants expressed in the context of the full length receptor (dark grey) or the receptor truncated from the adhesion PLL domain (light grey), (top) “raw” assay data, (bottom) normalized to WT.
Fig 17. TG2-induced activity SRE assays a. TG2 Ligand induced activities measured using the SRE assay for GAIN variants expressed in the context of the full length receptor (dark green) compared to activities without ligand (grey), b. TG2 Ligand induced activities normalized to those without ligand.
Fig 18. Table summarizing all results. Surf express: surface expression by ELISA, Const act: constitutive activity, TG2-ind ac: TG2-induced activity.
The examples illustrate the invention.
Example 1 - Introduction
In the present examples single-molecule force spectroscopy (SMFS), steered molecular dynamics (SMD) simulations, computational protein design and cell-based signaling are used assays to characterize and reprogram the response of the human ADGRG1 GAIN domain to mechanical force (Fig. 1A). In more detail, in silico SMD simulations and computational protein design are fundamental to arrive reprogrammed human ADGRG1 GAIN domain which are either more or less sensitive to mechanical force and subsequently SMFS as well as cell-based signaling assay can be used to show and measure the changed sensitivity more or less sensitive to mechanical force experimentally.
A protein’s mechanical stability is highly dependent on the point and direction of force application. Thus, we designed an SMFS setup that best mimics the predicted orientation and direction of force application on the GAIN domain in the context of the full-length ADGRG1 receptor expressed at the cell-surface (Fig. 1A). Force-extension curves were recorded at constant speed using an atomic force microscope and ddFLN4 unfolding fingerprints were used to select single-molecule unfolding events for statistical analysis. WT GAIN-Stachel dissociation events were measured in the range of 100-160 pN with a clear dependence on the force loading rate (Fig. 1C-D). Based on the shape of the forceextension curve (Fig. 1C, red portion) and on contour length analysis of unfolding traces, GAIN domain unfolding from the Stachel peptide appears to be a single concerted dissociation event with no predicted unfolding intermediates. The GAIN domain thus functions as a two-state mechanosensitive
system and encodes a high force resistance for a cell-surface receptor involved in adhesion and signaling (Fig. 1 E).
Example 2 - Results and Discussion
Unlike other adhesion receptors that respond to mechanical force such as integrins or Notch, adhesion GPCRs have a relatively simple architecture. Furthermore, our results show that a major component of their signaling, i.e. GAIN shedding and Stachel-dependent receptor activation, involves a single conformational transition at a force of 100 pN or above. Even though the exact position of this activation barrier may differ in the context of the full-length receptor at the cell surface, these results suggest that there is a large range available for the reprogramming of the mechanical properties of the GAIN domain, both towards lower and higher mechanical strength. Coupled to the modularity of their adhesion domains, adhesion GPCRs and their GAIN domains represent ideal scaffolds for the design of mechanosensors with specific thresholds of force activation.
To this end, we first wanted to elucidate which structural features might confer this protein domain with high mechanical resistance against Stachel dissociation. Structural observation of the local Stachel environment shows an extensive network of backbone hydrogen bonds between the Stachel and its two adjacent anti-parallel p-strands b6 and b8 (Fig. 1 B). These features are reminiscent of the “mechanical clamp” found in well-characterized Ig-like, p-sandwich type mechanical proteins such as the I27 module of human cardiac titin (18-21, 35) (Fig. 1 E). In these systems, the patch of backbone hydrogen bonds forming the clamp contributes prominently to the protein domain’s overall force resistance. The single GAIN-Stachel dissociation step observed by AFM suggests that the Stachel Flbond network breaks in a single shearing step once the mechanical force threshold is reached, thus classifying the GAIN domain as a shear mechanical clamp system.
A traditional approach to modify the mechanical stability of this domain would consist in generating mutations predicted to alter the local environment of the mechanical clamp based on analysis of the native (static) structure alone, as reported previously (22, 36). However, the Stachel environment and broader GPS motif encode multiple other functions relevant to the full-length receptor context. The GPS motif ensures proper conformation of the fold to enable autoproteolysis and the Stachel sequence is highly conserved among adhesion GPCRs and enables receptor activation by engagement with the orthosteric pocket upon GAIN shedding. Furthermore, the consequences on folding of mutations in the domain’s core would be hard to predict, and even mutations that enable proper folding might not affect mechanostability since the mechanical clamp is mediated solely by backbone hydrogen bonds. Finally, as reported in other mechanical clamp systems, many other factors can influence a protein’s mechanical stability including core (in particular hydrophobic) packing (37-39) and solvent accessibility of force-bearing hydrogen bonds (40).
Rather than design based on a static picture of the GAIN domain, we opted for design based on the dynamical properties of the domain under force. Several studies have shown that force propagation pathways through a system and the orientation of the force vector on the rupture interface, in our case the Stachel mechanical clamp, are major determinants of the system’s mechanical resistance (41-44). High mechanical resistance requires the propagation of force through pathways facing strong orthogonal components at the rupture interface. To understand how force propagates through the GAIN domain and is applied to the Stachel peptide, SMD simulations were performed on the WT GAIN at pulling speeds of 0.1 nm.ns-1 and 1.0 nm.ns-1. A harmonic force potential was applied to the structure by pulling simultaneously from the N- and C-termini of the GAIN, resulting in an increase of the force accumulated within the system that is commonly defined as the loaded state of the protein. During this state, the protein undergoes topological rearrangements that enable force to propagate through pathways favoring its mechanical resistance. To identify these force propagation pathways, the coordinates of a-carbons were extracted from the time window corresponding to the GAIN loaded state. Their sequential motions between two consecutive frames of the simulation were then computed and correlation analyses were performed to reconstruct the force propagation pathways from these motions. Pathways were identified by following the route from the N- to C-terminal that maximizes the observed correlations. Finally, identified pathways between consecutive frames were clustered and a deep analysis of cluster populations enabled us to discern optimal (primary cluster) from suboptimal pathways (secondary clusters). To extract the orientation of force on the Stachel mechanical clamp, the angles between the force pathway vector and the Stachel H-bond network plane were computed. This data analysis pipeline is represented schematically in Fig 2A and detailed in Example 3.
Figure 2B shows the results of the WT GAIN SMD simulations at the two pulling speeds. The optimal and suboptimal pathways corresponding to the three largest clusters are represented as violin plots that depict the distribution of the average force angles upon contact with the Stachel hydrogen bond network in the loaded state. The relative size of each pathway cluster is given in percentages above each violin plot distribution and indicates how much each pathway is “used” by the GAIN domain to transmit force. We find that there is large variability in the computed average angles among the three observed pathways, suggesting that these pathways are not redundant and that they each carry distinct mechanical information. Interestingly, whereas the residues involved in the suboptimal pathways change between the two pulling speeds, the WT optimal force propagation pathway uses the same residues at both pulling speeds, suggesting a higher robustness. However, the average angle it forms with the Stachel varies with pulling speed (Fig. 2B), indicating speed-specific structural rearrangements of the topology. Furthermore, the optimal pathway shows a quick jump from the helix 1 N-terminal pulling point to the p-domain and, unlike the suboptimal pathways, transits to the Stachel without utilizing the a-domain, suggesting a preferred force propagation through the mechanically resistant p-sandwich. The lower force angle of the optimal pathway on the Stachel at the higher pulling speed (23° vs 40°) would suggest a lower mechanical resistance but may be counteracted by the large increase in the relative importance of this pathway propagating through the mechanically resistant p-domain (42% vs 29%). Overall, the diversity of pathways observed and their change of
relative importance at different pulling speeds shows the modularity of information transmission through the GAIN domain. This is further proof that there is room to redesign the mechanical properties of the GAIN domain based on the dynamical features described in this section. Using the described metrics, we thus set out to design mutations and predict their effects on the domain’s mechanical stability based on the rewiring of its force propagation pathways.
In order to leave the functionally diverse Stachel environment and GAIN domain core untouched, we chose to introduce mutations at sites distal from the rupture interface. We opted to maximize our chances of modifying the system’s mechanical resistance by creating topology locks using disulfides (45). A GAIN domain-wide screen of Ca distances was performed to extract pairs of residues (distal in sequence) that could potentially satisfy the requirements for disulfide formation. We used the SMD force pathway analysis of the WT GAIN and rational structural analysis to narrow down this list to two key residue pairs that, if able to form disulfides, were predicted to significantly affect force propagation and overall GAIN behavior under force: [helix 1 Val174 and helix 2 Ser218] (giving mutant V174C- S218C, referred to as SS_h1 h2), and [helix 2 Ser206 and loop 10 Glu351] (giving mutant S206C- E351C, referred to as SS_h2l10). These two mutants were also combined to give the SS_h1 h2-h2l10 variant. Figure 3A shows a one-dimensional topology chart of the GAIN domain with the positions of the cysteine mutations.
Our SMD simulation and data analysis pipeline was used to characterize the three GAIN domain variants (Fig. 2C-E). The SS_h2l10 variant, which locks loop 10 from the p-domain onto the N-terminal of helix 2 (Fig. 3A), showed an optimal pathway that remained the same at both speeds, propagating along helix 1 and interestingly utilizing the engineered disulfide to transit to the p-domain and reach the Stachel. This rewired pathway generated a significant increase of the angle between the force vector and the rupture interface at both pulling speed (79° and 63° compared to 23° and 40° for the WT optimal path at 0.1 and 1.0 nm.ns-1, respectively; Fig. 2C). Much like the WT GAIN domain, the higher pulling speed enhanced the relative importance of the optimal path (from 40% to 52%) over the two suboptimal paths. We predict from these results that the SS_h2l10 variant should have increased mechanical stability and a higher GAIN-Stachel dissociation threshold compared to WT. Conversely, the SS_h1 h2 variant, which locks the N-terminal of helix 1 with the C-terminal of helix 2 (Fig. 3A), showed optimal and suboptimal pathways that were very similar to WT pathways, with no major rewiring. However, the optimal pathway in this variant displayed much lower average angles of the force vector on the rupture interface at both pulling speeds compared to WT (10° and 2° compared to 23° and 40° for the WT optimal path at 0.1 and 1.0 nm.ns-1, respectively; Fig. 2D). Due to the increased collinearity of the force vector and rupture interface, we predict a lower force threshold of GAIN-Stachel dissociation in this variant. The combination of the two disulfides in the SS_h1 h2-h2l10 variant also generated WT-like optimal and suboptimal pathways at both pulling speeds. Interestingly, the h2l10 disulfide was not utilized to propagate force in this variant. However, the addition of the h1 h2 disulfide to the h2l10 disulfide significantly decreased the average angles of the optimal pathway on the Stachel at both speeds, “rescuing” the high SS_h2l10 angles to WT-like values (from 79° and 63°
in SS_h2l10 down to 36° and 38° in SS_h1 h2-h2l10; Fig. 2E). We predict from this interesting behavior that the SS_h1 h2-h2l10 variant should have a GAIN-Stachel force dissociation threshold close to that of the WT GAIN domain.
In order to test our SMD predictions, the three GAIN variants were characterized experimentally. We first solved the crystal structure of the SS_h1 h2-h2l10 variant of the human ADGRG1 GAIN domain at 1.9 A resolution to confirm the correct formation of the h1 h2 and h2l10 disulfides (Fig. 3B). Circular dichroism spectra were recorded from 20°C to 96°C to measure the thermodynamic stability of WT and variant GAINs and further confirm the correct formation of each disulfide individually: all three disulfide variants showed a significantly enhanced melting temperature from WT. As predicted from the shifts in the mutant force pathway angles upon contact with the Stachel seen in SMD, AFM measurements showed that SS_h1 h2 dissociates at lower forces compared to WT, while SS_h2l10 reaches a force threshold of dissociation almost twice as high as WT (Fig. 3C). Interestingly, we also confirm that the combination of these two disulfides in the SS_h1 h2-h2l10 variant rescues a WT behavior in terms of dissociation forces. Cell-surface expression of these variants was assessed by ELISA after transfection of HEK293T cells with APLL ADGRG1 constructs containing an N-terminal AVI-tag. All three variants showed equivalent expression to WT (Fig. 3D, top). The same transfection was performed to measure the constitutive activity of the variants using an SRE assay. SS_h1 h2 showed a 50% decrease in constitutive activity compared to WT, while SS_h2l10 and SS_h1 h2-h2l10 showed a 75% decrease (Fig. 3D, bottom). These results coincide with the thermodynamic stability values obtained by CD (Fig 3E), suggesting/confirming that spontaneous thermal unfolding is the major determinant of constitutive activity of the full-length receptor expressed at the cell membrane. Figure 3E shows how our GAIN variants populate the 2D space of mechanical versus thermodynamic stability with respect to WT. As reported for other mechanosensing protein systems (46), GAIN mechanical stability is uncoupled from GAIN thermodynamic stability.
We next sought to assess whether the reprogrammed mechanical resistance of the engineered GAINs would translate in distinct signaling responses to mechanical stimulus of full-length receptors incorporating the designed domains. We reasoned that receptors engineered with mechanoresistant disulfide bonds should be less sensitive to shear stress than those incorporating GAINs rewired with mechanosensitive bonds. We applied shear stress through vibration to cells expressing full-length ADGRG1 cultivated on plates coated with specific receptor ligands such as collagen III (collll) and transglutaminase 2 (TG2). In this experiment, mechanical stress is transmitted to the receptor by the adhesion domain through the coated ligand and receptor activation is measured by the SRE reporter (Fig. 4D). By comparing SRE responses for a specific ligand under static (Fig. 4A) or vibration conditions (Fig. 4B), we can selectively measure the impact of the designed bonds on the mechanosignaling responses of the receptors (Fig. 4C). While ligands can trigger different activation responses under the same vibration conditions, we observed a similar trend between the designed ADGRGIs. Consistent with the mechanical properties of the designed GAIN domains, SS_h1 h2 was more
mechanosensitive than the other variants and displayed signaling responses similar to or higher than wr.
Overall, while we did not observe a linear relationship between single molecule behavior of the GAIN and the cellular response of the full length receptors, our results indicate that adhesion GPCRs can be rationally designed to modulate cellular responses to mechanical stress. Since mechanical cues represent a universal class of cellular stimuli, our approach provides a foundation for engineering custom-built mechanosensors that should broadly impact synthetic cell biology and therapeutic cell engineering approaches.
Example 3 - Material and Methods
Reagents
Dulbecco's Modified Eagle's medium (DMEM)
Fetal Bovine Serum (FBS)
Penicillin-Streptomycin cocktail (PenStrep)
DPBS
Bovine Serum Albumin (BSA, Sigma)
GeneJuice transfection reagent (Novagen)
Paraformaldehyde (PFA)
Primary anti-GPR56 antibody
Secondary anti-X antibody
SuperSignal™ West Pico PLUS Chemiluminescent Substrate (ThermoScientific #34080)
Collagen
Transglutaminase
Gene Constructs pcDNA constructs pBB2 constructs
Protein Expression and Purification
AFM handles - SpyCatcher and SdrG proteins for SMFS were prepared as previously reported (7, 2). Constructed recombinant plasmids pET28a-ybbR-His-ELP(MV7E2)3-FLN-SpyCatcher (Addgene #157674) and pET28a-SdrG-FLN-ELP-His-ybbr (Addgene #168047) were transformed into E. coli BL21 (DE3) strain. Cells were cultured in 5 ml of Luria-Bertani (LB) medium with 50 pg ml-1 kanamycin at 37 °C overnight. The culture was transferred to 100 mL of Terrific Broth (TB) medium with 50 pg ml’ 1 kanamycin and cultivated at 37 °C and 200 rpm until an optical density at 600 nm (GD600) of ~0.8-
1.0 was reached. Recombinant protein expression was induced upon addition of 0.5 mM isopropyl-p- D-thio-galactopyranoside (IPTG) and the culture was further incubated at 20 °C and 200 rpm for 18 h. The cells were harvested by centrifugation at 4,000 g for 20 min at 4 °C. The cell pellets were stored at -80 °C until further purification. For purification, the harvested cell pellet was resuspended in a lysis buffer (50 mM Tris, 50 mM NaCI, 0.1% Triton X-100, 5 mM MgCI2; pH 8.0), and disrupted with a sonic dismembrator. The lysate was centrifuged at 14,000 g for 20 min at 4 °C. The supernatant was collected and incubated with Ni-NTA resin (Thermo Fisher Scientific), loaded onto a column, washed with wash buffer (1x PBS buffer with 25 mM imidazole; pH 7.2), and eluted in elution buffer (1x PBS buffer with 500 mM imidazole; pH 7.2). The eluted protein solution was buffer-exchanged and further purified by Superose SEC column, and finally stored in 33% glycerol at -20 °C. For the NoGAIN contruct (The same contruct with GAIN-WT but without GAIN domain), newly constructed recombinant plasmid pET28a-Fgb-NoGAIN-HIS-SpyTag was used. Recombinant protein expression was induced upon addition of 1 mM IPTG and the culture was further incubated at 37 °C and 200 rpm for 9 h. Purification was carried out in the same manner as previously described. Due to the small size of NoGAIN, elutant from Ni-NTA resin was directly incubated with purified ybbR-His-ELP(MV7E2)3-FLN- SpyCatcher. The conjugated form was then further purified by Superose SEC column, and finally stored in 33% glycerol at -20 °C.
GAIN domain - Insect codon-optimized human ADGRG1 ECR sequences (Uniprot: Q9Y653) were synthesized (Twist Biosciences) as part of larger constructs for either AFM or X-ray crystallography experiments (see full protein sequences) in a pFastBac vector backbone. A baculovirus expression system was used for expression of proteins in High-Five insect cells. pFastBac constructs were transformed into DHWMultiBac cells (Geneva Biotech), white colonies indicating successful bacmid recombination were selected, and bacmids were purified by the alkaline lysis method. Sf9 cells were transfected using the X-tremeGENE 9 transfection reagent (Sigma) with the desired bacmid. eYFP-positive cells were observed after 1 week and subjected to one round of viral amplification. Amplified P2 virus was used to infect High-Five cells at a density between 1-2 x 106 cells/ml. Cells were incubated for 72 h at 28°C, pelleted (1000 x g, 10 minutes) and the recovered supernatant was clarified (5000 x g, 30 minutes) and filtered using a Grade GD 1 pm pre-filter (1841-047, Whatman, GE) on top of a Durapore 0.45 pm PVDF Membrane filter (HVLP04700, Merck Milipore). 5 ml of 50% Fastback Ni Advance Resin slurry (Protein Ark) pre-equilibrated in HBS buffer containing 10 mM HEPES (pH 7.2) and 150 mM NaCI was added to the 500 ml clarified supernatant containing the secreted proteins and incubated over-night at 4°C on a rotating wheel. The resin was recovered on a gravity-flow column, equilibrated with 30 ml HBS and washed with 30 ml HBS containing 30 mM imidazole. Proteins were eluted with 15 ml of HBS containing 300 mM imidazole and 1 ml fractions were collected and analyzed by SDS-PAGE. The three or four purest and most concentrated fractions were pooled and concentrated on a 30 kDa Amicon Ultra-4 centrifugal filter (UFC803024, Merck Millipore) down to 500 pl. The concentrated proteins were loaded on a Superdex 200 Increase 10/300 GL (28-9909-44, Sigma-Aldrich) connected to an AKTA Pure chromatography system (GE) and eluted in HBS. The fractions containing the peak corresponding to the purified GAIN construct were pooled
and quantified using A280 absorbance. Purified GAIN proteins were either directly concentrated to 10 mg. ml-1 for crystallization or flash-frozen in liquid nitrogen and stored at -80°C for subsequent AFM experiments.
AFM-SMFS surface preparation and protein immobilization
The surface modification of cantilever and coverglasses and the protein immobilization were done as described previously (fig. 5) (2). Cantilevers were cleaned by UV-ozone treatment for 40 min and coverglasses were soaked in piranha etching solution and rinsed with distilled water (DW). Then, cantilevers and coverglasses were treated with 3-Aminopropyl (diethoxy) methylsilane (APDMES, ABCR GmbH, Karlsruhe, Germany) to silanize the surfaces with amine groups. The amine groups subsequently reacted to a NHS group from sulfosuccinimidyl 4-(N-maleimidomethyl)cyclohexane-1- carboxylate (sulfo-SMCC; Thermo Fischer Scientific) in 50 mM HEPES buffer pH 7.5 for 30 min. The thiol group from Coenzyme A (CoA, 200 pM) reacted to a maleimide group from sulfo-SMCC in coupling buffer (50 mM sodium phosphate, 50 mM NaCI, 10 mM EDTA, pH 7.2) for 2 hrs. Finally, the ybbR-tagged proteins SdrG-FLN-ELP-His-ybbr and pre-conjugated ybbR-His-ELP(MV7E2)3-FLN- SpyCatcher-SpyTag-GAIN variants were site-specifically anchored to the surface using SFP-mediated ligation to CoA in Mg2+ supplemented 1x PBS buffer. This resulted in covalent immobilization of SdrG and GAIN variants to cantilever and cover glasses, respectively. Protein-immobilized cantilevers and coverglasses were extensively washed and kept in 1x PBS buffer prior to immediate use. SpyTag- SpyCatcher conjugation of ybbR-His-ELP-FLN-SpyCatcher and Fgp-GAIN-His-SpyTag variants were done by mixing two proteins with same molar ratio in PBS buffer and pre-incubation for 1-12 hrs prior to ybbR tag ligation.
AFM-SMFS measurements and data analysis
Force spectroscopy measurements with SdrG and Fgp-fused GAIN variants were conducted in the same manner as previously illustrated using automated AFM-based SMFS (Force Robot 300, JPK Instruments). SMFS data were recorded in 1x PBS at room temperature with constant pulling speeds of 400, 800, 1600 and 3200 nm.s 1. A total of 258,044 force-extension curves were acquired. Forceextension curves were filtered and analyzed by a combination of software available on the AFM instrument and custom python scripts. The majority of data traces contained no interactions, nonspecific interaction, or complex multiplicity of interactions. Therefore, the data traces were filtered by searching for contour length increments that matched the lengths of the fingerprint domains, FLN (=36 nm). Theoretical contour length increment was calculated based on the equation ALC = (0.365 nm/AA) x (# AAs in POI) - Lf, where ALC is expected contour length increment and Lf is end-to-end length of folded protein domain. For FLN, ALC = 36.9 nm - Lf, where Lf is typically < 5nm. The total number of force-extension curves matching this criterion was 509 out of 120,780, 1 ,058 out of 68,376, 263 out of 42,295, and 158 out of 18,447 for GAIN-WT, GAIN(2Cys), GAIN(2Cys)-h1 h2, and GAIN(4Cys), respectively.
For dynamic force spectra, the rupture or unfolding forces vs. loading rate was plotted and median forces and loading rates for each pulling speed were fitted to Bell-Evans model (3, 4) to estimate the effective distance to the transition state (Ax) and the intrinsic dissociation rate or unfolding rate (kotr) in the absence of force. Data were fitted using the Dudko-Hummer-Szabo model (5, 6) with a cusp-like barrier (v = 0.5) to estimate Ax, kotf, and energy barrier height (AG).
Table 1
Energy landscape parameters of stachel dissociation from GAIN domain from AFM-SMFS from Bell- Evans (BE) model (3, 4) and Dudko-Hummer-Szabo (DHS) model (5, 6) (v = 0.5 and bin size of 20 PN).
For direct comparison with the same cantilever, both GAIN-WT and GAIN(4Cys) were immobilized on different areas of the same surface and a single cantilever was used to probe each spot using constant speed pulling at 3200 nm s-1 . After 500 approach-retraction cycles at one area, the surface was moved to the other area. This cycle was done once for each GAIN-WT and GAIN(4Cys). The total number of force-extension curves matching the criterion was 69 out of 9,484, and 67 out of 4,986 for GAIN-WT and GAIN(4Cys), respectively.
For contour length analysis in reducing condition, each GAIN-WT and noGAIN was probed with the same cantilever both in non-reducing condition (0 mM DTT in 1x PBS) and reducing condition (100 mM DTT in 1x PBS). SMFS data were recorded in 1x PBS first at room temperature with constant pulling speed of 800 nm s-1 , and then buffer was changed to 100 mM DTT in 1x PBS and SMFS data were recorded with the same pulling speed. The total number of force-extension curves matching the criterion was 101 out of 1 ,566, and 116 out of 26,097 for GAIN-WT and GAIN(4Cys), respectively.
Circular Dichroism measurements
Purified GAIN variant proteins were adjusted to 10-15 pM final concentration in 10 mM HEPES, 15 mM NaCI, pH 7.5. Thermal denaturation CD spectra were recorded on a Chirascan CD spectrometer (AppliedPhotophysics) from 20°C to 96°C using a 2°C step, between 215 nm and 250 nm. Each melting curve was normalized between 0 and 1 across the entire spectral range and the 216 nm and 232 nm normalized spectra from 2 or 3 replicate experiments were extracted and plotted.
Mammalian Cell Assays
Maintenance - Human embryonic kidney cells (HEK293T, ATCC CRL-3216) were cultured in 10 cm diameter tissue culture-treated dishes in a humidified incubator (95% air and 5% CO2) at 37°C in growth media (DMEM, 10% heat-inactivated FBS, 1% PenStrep). Cells were passaged every 2-3 days (up to passage 30) by lifting the cells with 0.05% trypsin in DPBS.
Transfection - Cells were transiently (forward or reverse) transfected with plasmid constructs using the GeneJuice transfection reagent following the supplier’s recommended protocol (see specific conditions for each assay below). For forward transfection, 25x103 cells were seeded per well of a 96- well plate and the transfection was carried out 24h later. For reverse transfection, 50x103 cells were seeded per well of a 96-well plate and directly transfected.
ELISA - 25x103 HEK293T cells were seeded per well of a 96-well plate pre-coated with 0.1 mg. ml-1 PDL. 24 hours post-seeding, cells were transfected as described previously with 100 ng of the indicated pcDNA-GPR56 construct and 25 ng of dualLUC-SRE plasmid. 24 hours post- transfection, cells were washed with DPBS, fixed with cold 4% PFA for 10 minutes at 4°C and washed again DPBS after PFA removal. This was followed by three consecutive 60-minute incubations with 2% BSA in DPBS, X pg. ml-1 primatry anti-GPR56 antibody in 2% BSA-DPBS, and X pg. ml-1 secondary anti-X antibody in 2% BSA-DPBS. After each antibody incubation, cells were thoroughly rinsed with DPBS. The cells were then incubated in chemiluminescent substrate at room temperature for 10 minutes. The Flexstation 3 (Molecular Devices) plate reader and SoftMax Pro 7 software (Molecular Devices) were used to measure luminescence. Assays were performed in triplicate or quadruplicate and the luminescence signals were normalized by the background signal of empty vector transfected cells.
Dual-Gio Luciferase SRE assays - 25x103 HEK293T cells were seeded per well of a 96-well plate precoated with 0.1 mg. ml-1 PDL. 24 hours post-seeding, cells were transfected as described previously with 100 ng of the indicated pcDNA-GPR56 construct and 25 ng of dualLUC-SRE plasmid. 8-10 hours post-transfection, the culture medium was replaced with FBS-free DMEM for over-night serum starvation. The next morning, cells were induced with DMSO or 50 pM Stachel peptide in two doses at to and t+2h of the 4-hour induction phase. At the end of the induction period, 60 pl of the 100 pl total
well volume was removed and the Dual-Glo Luciferase SRE assay was carried out following the manufacturer’s instructions apart from a working volume of 40 pl instead of 100 pl.
Vibration assays with coated ligands - 96-well plates were coated overnight at 4°C with 40 pl PDL, collagen or transglutaminase at the indicated densities. The next morning, 50x103 cells were reverse transfected onto the pre-coated wells as described previously with 100 ng of the indicated pcDNA- GPR56 construct and 25 ng of dualLUC-SRE plasmid. 6-8 hours post-transfection, the culture medium was replaced with FBS-free DMEM for over-night serum starvation. The next morning, cells were placed on a Titramax 100 vibration shaker (Heidolph Instruments) and vibrated at the indicated frequencies and time periods. 4 hours after the beginning of the induction phase, the Dual-Glo Luciferase SRE assay was carried out as described previously.
Computational variant design
To screen for possible pairs of residues enabling the formation of disulfides, the Coupled Moves algorithm from the Rosetta software was used with the backrub movers protocol (7). A total of 100 independent structures was generated for every pair and top 10% lowest energy structures were conserved. All structures were then minimized and relaxed to access lowest energy potential through the relax protocols included in the Rosetta package. Distances between the sulfide groups of added cysteines were computed and averaged across the lowest energy structures. To determine which pairs of mutations could form a disulfide, a distance cutoff of 2.15 A between the 2 cysteines was used to conserve or exclude mutants.
SMD simulations
The structure of the GAIN_SS-h2l10-h1 h2 variant was resolved at 1.92 A resolution and back-mutated using the above protocol to recover the individual disulfide variants as well as the WT structure. The system was solvated with a box size of 75 A along X and Y dimensions and varying Z depending on the simulation with 0.1 M of Na+ and Cl~ ions. The total system size was approximatively 130 k atoms. Simulations were performed with GROMACS 2019.4 (8) using the CHARMM36 forcefield (9) along with the TIP3 water model (10) in an NPT ensemble at 31 OK and 1 bar using a velocity rescaling thermostat (with a relaxation time of 0.1 ps) and Berendsen barostat (with a relaxation time of I ps). Equations of motion were integrated with a timestep of 2 fs using a leap-frog algorithm. Each system was energy minimized using a steepest descent algorithm for 5000 steps, and then equilibrated with the atoms of the protein restrained using a harmonic restraining force in two steps, first by adding the thermostat for 50000 steps and then the barostat for 500000 steps. SMD simulations were performed using the umbrella sampling method with a pulling speed ranging from 1 A/ns to 10 A/ns. In all simulations, the 2 alpha carbons of the N-term and C-term residues were moved harmonically in the desired direction with constant velocity. A distance cutoff of 14 A was used for short-range non-
bonded interactions, whereas electrostatic interactions were computed using the Particle-mesh Ewald method (11) with a cubic interpolation of power 4.
SMD simulation analysis and force propagation pathway identification
To extract force propagation pathways from the SMD simulations, correlation between alpha carbon motions were analyzed. For every frame of the simulation, the [x,y,z] cartesian coordinates of every alpha carbon were extracted. Displacement of these coordinates between 2 consecutive frames were computed and represented as a vector [xo-xi, yo-yi, zo-zi]. To establish a correlation between these displacement vectors, distance correlation was used (Equation 1).
For a simulation of t frames, this resulted in t correlation matrices of [nresidues, nresidues]. To identify force propagation pathways, Dijkstra shortest path algorithm was used (72). The shortest path maximizing the total correlation were searched between the N-term of the GAIN domain and the C-term of the Stachel mechanical clamp for every correlation matrix, resulting in t shortest force propagation pathways. To extract relevant information from these pathways, k-means clustering was performed. The optimal number of clusters to use for independent simulations was determined using the inertia attribute to identify the sum of squared distances of samples to the nearest cluster center. Centroids of clusters, corresponding to the optimal force propagation pathway within this cluster were used for later extraction of features.
Force propagation pathway projection on Stachel rupture interface
To predict the effect of the different mutants in-silico, one criterion used was the angle of force propagation pathway arrival on the Stachel. Force propagation pathways were identified as previously described. To extract the angle of pathway projection on the rupture interface, coordinates of all alpha carbons were extracted. These coordinates represented a [t, nresidues, nresidues] matrix, t being the number of frames of the simulation. For each of this coordinate matrix, we constructed a plane containing both the Stachel N-term and C-term alpha carbons, with the alpha carbon of Val142. This resulted in an ensemble of dynamic planes moving accordingly to local structural rearrangements and corresponding to the rupture interface of the Stachel mechanical clamp. To estimate the angle of force arrival on this rupture interface defined by these different planes, vectors representing the jump of pathway from the GAIN domain onto the Stachel mechanical clamp were extracted. These propagation vectors were computed from the alpha carbon coordinates of the pair of residues involved in the jump of force onto the Stachel mechanical clamp. Angles between the plane normal and the force arrival were computed for every frame corresponding to the loaded state of the system. Finally, angles were concatenated through simulation replicates, and violin plots were used to show the distribution of angles across mutants.
H-bonds 2D projection
To identify backbone H-bonds that ensure the mechanical resistance of the Stachel mechanical clamp, the Wernet-Nilsson algorithm was used (8). Identified H-bonds across the simulation were 2D- projected in a plane perpendicular to the Stachel axis defined by its N-term and C-term residues. To highlight the H-bonds rearrangements occurring during the pulling, the 2D projections were compared between the native unloaded structure and the loaded structure corresponding to the highest mechanical resistance point observed in the SMD simulations. Projections were averaged across replicates and plotted.
Full recombinant protein construct sequences
The human GAIN domain sequence used corresponds to native residues N171 to S391. C177, found in the native sequence and forming a disulfide bridge with C121 of the PLL domain, was mutated to a serine in all constructs (bold). Engineered disulfides are indicated as bold and underlined letters.
FqB-//nker-GAIN WT-//nker-70H/s-Spytaq (AFM)
ATNEEGFFFSARGHRPLDGSGSGSGSAGTGSGNASVDMSELKRDLQLLSQFLKHPQKASRRPSAAP ASQQLQSLESKLTSVRFMGDMVSFEEDRINATVWKLQPTAGLQDLHIHSRQEEEQSEIMEYSVLLPRT LFQRTKGRSGEAEKRLLLVDFSSQALFQDKNSSQVLGEKVLGIWQNTKVANLTEPVVLTFQHQLQPK NVTLQCVFWVEDPTLSSPGHWSSAGCETVRRETQTSCFCNHLTYFAVLMVSGSGSGSGSAG7GSG /7/7/7/7/7/7/7/7/7/7GSAHIVMVDAYKPTK (SEQ ID NO:82)
FqB-//nker-GAIN SS h2l10-//nker-70H/s-Spytaq (AFM)
ATNEEGFFFSARGHRPLDGSGSGSGSAGTGSGNASVDMSELKRDLQLLSQFLKHPQKASRRPSAAP ACQQLQSLESKLTSVRFMGDMVSFEEDRINATVWKLQPTAGLQDLHIHSRQEEEQSEIMEYSVLLPRT LFQRTKGRSGEAEKRLLLVDFSSQALFQDKNSSQVLGEKVLGIWQNTKVANLTEPVVLTFQHQLQPK NVTLQCVFWVCDPTLSSPGHWSSAGCETVRRETQTSCFCNHLTYFAVLMVSGSGSGSGSAG7GSG /7/7/7/7/7/7/7/7/7/7GSAHIVMVDAYKPTK (SEQ ID NO: 83)
FqB-//nker-GAIN SS hi h2-//nker-70H/s-Spytaq (AFM)
ATNEEGFFFSARGHRPLDGSGSGSGSAGTGSGNASCDMSELKRDLQLLSQFLKHPQKASRRPSAAP ASQQLQSLESKLTCVRFMGDMVSFEEDRINATVWKLQPTAGLQDLHIHSRQEEEQSEIMEYSVLLPRT LFQRTKGRSGEAEKRLLLVDFSSQALFQDKNSSQVLGEKVLGIWQNTKVANLTEPVVLTFQHQLQPK NVTLQCVFWVEDPTLSSPGHWSSAGCETVRRETQTSCFCNHLTYFAVLMVSGSGSGSGSAG7GSG /7/7/7/7/7/7/7/7/7/7GSAHIVMVDAYKPTK (SEQ ID NO: 84)
FqB-//nker-GAIN SS h2l1O-h1h2-//nker-lOH/s-Spytaq (AFM)
ATNEEGFFFSARGHRPLDGSGSGSGSAG7GSGNASCDMSELKRDLQLLSQFLKHPQKASRRPSAAP
ACQQLQSLESKLTCVRFMGDMVSFEEDRINATVWKLQPTAGLQDLHIHSRQEEEQSEIMEYSVLLPR
TLFQRTKGRSGEAEKRLLLVDFSSQALFQDKNSSQVLGEKVLGIVVQNTKVANLTEPVVLTFQHQLQP
KNVTLQCVFWVCDPTLSSPGHWSSAGCETVRRETQTSCFCNHLTYFAVLMVSGSGSGSGSAGTGS
GHHHHHHHHHHGSAHIVMVDAYKPTK (SEQ ID NO: 85)
GAIN SS h2!1O-h1h2-//nker-,fOH/s (X-ray crystallography)
ATNASCDMSELKRDLQLLSQFLKHPQKASRRPSAAPACQQLQSLESKLTCVRFMGDMVSFEEDRINA
TVWKLQPTAGLQDLHIHSRQEEEQSEIMEYSVLLPRTLFQRTKGRSGEAEKRLLLVDFSSQALFQDKN
SSQVLGEKVLGIWQNTKVANLTEPVVLTFQHQLQPKNVTLQCVFWVCDPTLSSPGHWSSAGCETVR
RETQTSCFCNHLTYFAVLMVSGSGS/7/7/7/7/7/7/7/7/7/7 (SEQ ID NO: 86) ybbR-6H/s-ELP-FLN-SpyCatcher (AFM surface handle)
MGTDSLEF1ASKLA/7/7/7/7/7/-A/VGSGHGVGVPGMGVPGVGVPGVGVPGVGVPGVGVPGVGVPGVGV
PGVGVPGEGVPGEGVPGVGVPGMGVPGVGVPGVGVPGVGVPGVGVPGVGVPGVGVPGVGVPGE
GVPGEGVPGVGVPGMGVPGVGVPGVGVPGVGVPGVGVPGVGVPGVGVPGVGVPGEGVPGEGVP
GWPSGSADPEKSYAEGPGLDGGECFQPSKFKIHAVDPDGVHRTDGGDGFVVTIEGPAPVDPVMVDN
GDGTYDVEFEPKEAGDYVINLTLDGDNVNGFPKTVTVKPAPGSGSGSGSVDTLSGLSSEQGQSGDM
TIEEDSATHIKFSKRDEDGKELAGATMELRDSSGKTISTWISDGQVKDFYLYPGKYTFVETAAPDGYEV
ATAITFTVNEQGQVTVNGKATKGDAHI (SEQ ID NO: 87)
SdrG-FLN-ELP-6H/s-ybbr (AFM cantilever handle)
MGTEQGSNVNHLIKVTDQSITEGYDDSDGIIKAHDAENLIYDVTFEVDDKVKSGDTMTVNIDKNTVPSD
LTDSFAIPKIKDNSGEIIATGTYDNTNKQITYTFTDYVDKYENIKAHLKLTSYIDKSKVPNNNTKLDVEYKT
ALSSVNKTITVEYQKPNENRTANLQSMFTNIDTKNHTVEQTIYINPLRYSAKETNVNISGNGDEGSTIID
DSTIIKVYKVGDNQNLPDSNRIYDYSEYEDVTNDDYAQLGNNNDVNINFGNIDSPYIIKVISKYDPNKDD
YTTIQQTVTMQTTINEYTGEFRTASYDNTIAFSTSSGQGQGDLPPEGSGSGSGSADPEKSYAEGPGL
DGGECFQPSKFKIHAVDPDGVHRTDGGDGFVVTIEGPAPVDPVMVDNGDGTYDVEFEPKEAGDYVI
NLTLDGDNVNGFPKTVTVKPAPGSGSGSHGVGVPGMGVPGVGVPGVGVPGVGVPGVGVPGVGVP
GVGVPGVGVPGEGVPGEGVPGVGVPGMGVPGVGVPGVGVPGVGVPGVGVPGVGVPGVGVPGVG
VPGEGVPGEGVPGVGVPGMGVPGVGVPGVGVPGVGVPGVGVPGVGVPGVGVPGVGVPGEGVPG
EGVPGWRG/7/7/7/7/7/7GSDSLEFIASKLA (SEQ ID NO: 88)
FqB-linker-lOH/s-Spytaq (noGAIN construct)
MNEEGFFFSARGHRPLDGSGSGSGSAGTGSGGSGSGSGSAGTGSG/7/7/7/7/7/7/7/7/7/7GSAHIVMVD AYKPTK (SEQ ID NO: 89)
Full-length amino acid sequences of indicated human adhesion G protein-coupled receptors
Preferable amino acid positions for replacement by Cys-Cys sites in GAIN domains (underlined) of indicated AGPCRs corresponding to the variants h1 h2 (bold), h2l10 (double-underlined) and h1 l7 (italic) are indicated. Respective Stachel sequences are highlighted in dark grey.
>sp|Q9Y653|AGRG1_HUMAN Adhesion G-protein coupled receptor G1 OS=Homo sapiens OX=9606 GN=ADGRG1 PE=1 SV=2 MTPQSLLQTTLFLLSLLFLVQGAHGRGHREDFRFCSQRNQTHRSSLHYKPTPDLRISIENSEEALTVH APFPAAHPASRSFPDPRGLYHFCLYWNRHAGRLHLLYGKRDFLLSDKASSLLCFQHQEESLAQGPPL LATSVTSWWSPQNISLPSAASFTFSFHSPPHTAAHNASVDMCELKRDLQLLSQF/-KHPQKASRRPSA APASQQLQSLESKLTSVRFMGDMVSFEEDRINATVWKLQPTAGLQDLHIHSRQEEEQSEIMEYSVLLP RTLFQRTKGRSGEAEKRLLLVDFSSQALFQDKNSSQVLGEKVLGIWQNTKVANLTEPVVLTFQHQLQ PKNVTLQCVFWVEDPTLSSPGHWSSAGCETVRRETQTSCFCNHLTYFAVLMVSSVEVDAVHKHYLS LLSYVGCVVSALACLVTIAAYLCSRVPLPCRRKPRDYTIKVHMNLLLAVFLLDTSFLLSEPVALTGSEAG CRASAIFLHFSLLTCLSWMGLEGYNLYRLVVEVFGTYVPGYLLKLSAMGWGFPIFLVTLVALVDVDNY GPIILAVHRTPEGVIYPSMCWIRDSLVSYITNLGLFSLVFLFNMAMLATMVVQILRLRPHTQKWSHVLTL LGLSLVLGLPWALIFFSFASGTFQLVVLYLFSIITSFQGFLIFIWYWSMRLQARGGPSPLKSNSDSARLPI SSGSTSSSRI
>sp|Q9HBW9|AGRL4_HUMAN Adhesion G protein-coupled receptor L4 OS=Homo sapiens OX=9606 GN=ADGRL4 PE=1 SV=3 MKRLPLLVVFSTLLNCSYTQNCTKTPCLPNAKCEIRNGIEACYCNMGFSGNGVTICEDDNECGNLTQS CGENANCTNTEGSYYCMCVPGFRSSSNQDRFITNDGTVCIENVNANCHLDNVCIAANINKTLTKIRSIK EPVALLQEVYRNSVTDLSPTDIITYIEILAESSSLLGYKNNTISAKDTLSNSTLTEFVKTVNNFVQRDTFV VWDKLSVNHRRTHLTKLMHTVEQATLRISQSFQKTTEFDTNSTDIALKVFFFDSYNMKHIHPHMNMDG DYINIFPKRKAAYDSNGNVAVAFVYYKSIGPLLSSSDNFLLKPQNYDNSEEEERVISSVISVSMSSNPPT LYELEKITFTLSHRKVTDRYRSLCAFWNYSPDTMNGSWSSEGCELTYSNETHTSCRCNHLTHFAILMS SGPSIGIKDYNILTRITQLGIIISLICLAICIFTFWFFSEIQSTRTTIHKNLCCSLFLAELVFLVGINTNTNKLF CSIIAGLLHYFFLAAFAWMCIEGIHLYLIVVGVIYNKGFLHKNFYIFGYLSPAVWGFSAALGYRYYGTTK VCWLSTENNFIWSFIGPACLIILVNLLAFGVIIYKVFRHTAGLKPEVSCFENIRSCARGALALLFLLGTTWI FGVLHWHASVVTAYLFTVSNAFQGMFIFLFLCVLSRKIQEEYYRLFKNVPCCFGCLR
>sp|Q9HAR2|AGRL3_HUMAN Adhesion G protein-coupled receptor L3 OS=Homo sapiens OX=9606
GN=ADGRL3 PE=1 SV=2
MWPSQLLIFMMLLAPIIHAFSRAPIPMAVVRRELSCESYPIELRCPGTDVIMIESANYGRTDDKICDSDP
AQMENIRCYLPDAYKIMSQRCNNRTQCAVVAGPDVFPDPCPGTYKYLEVQYECVPYKVEQKVFLCP GLLKGVYQSEHLFESDHQSGAWCKDPLQASDKIYYMPWTPYRTDTLTEYSSKDDFIAGRPTTTYKLP HRVDGTGFWYDGALFFNKERTRNIVKFDLRTRIKSGEAIIANANYHDTSPYRWGGKSDIDLAVDENGL
WVIYATEQNNGKIVISQLNPYTLRIEGTWDTAYDKRSASNAFMICGILYVVKSVYEDDDNEATGNKIDYI
YNTDQSKDSLVDVPFPNSYQYIAAVDYNPRDNLLYVWNNYHVVKYSLDFGPLDSRSGQAHHGQVSYI
SPPIHLDSELERPSVKDISTTGPLGMGSTTTSTTLRTTTLSPGRSTTPSVSGRRNRSTSTPSPAVEVLD DMTTHLPSASSQIPALEESCEAVEAREIMWFKTRQGQIAKQPCPAGTIGVSTYLCLAPDGIWDPQGPD LSNCSSPWVNHITQKLKSGETAANIARELAEQTRNHLNAGDITYSVRAMDQLVGLLDVQLRNLTPGGK
DSAARSLNKAMVETVNNLLQPQALNAWRDLTTSDQLRAATMLLHTVEESAFVLADNLLKTDIVRENTD NIKLEVARLSTEGNLEDLKFPENMGHGSTIQLSANTLKQNGRNGEIRVAFVLYN/VLGPYLSTENASMK LGTEALSTNHSVIVNSPVITAAINKEFSNKVYLADPWFTVKHIKQSEENFNPNCSFWSYSKRTMTGYW
STQGCRLLTTNKTHTTCSCNHLTNFAVLMAHVEVKHSDAVHDLLLDVITWVGILLSLVCLLICIFTFCFF RGLQSDRNTIHKNLCISLFVAELLFLIGINRTDQPIACAVFAALLHFFFLAAFTWMFLEGVQLYIMLVEVF ESEHSRRKYFYLVGYGMPALIVAVSAAVDYRSYGTDKVCWLRLDTYFIWSFIGPATLIIMLNVIFLGIALY
KMFHHTAILKPESGCLDNIKSWVIGAIALLCLLGLTWAFGLMYINESTVIMAYLFTIFNSLQGMFIFIFHCV
LQKKVRKEYGKCLRTHCCSGKSTESSIGSGKTSGSRTPGRYSTGSQSRIRRMWNDTVRKQSESSFIT
GDINSSASLNREGLLNNARDTSVMDTLPLNGNHGNSYSIASGEYLSNCVQIIDRGYNHNETALEKKILK
ELTSNYIPSYLNNHERSSEQNRNLMNKLVNNLGSGREDDAIVLDDATSFNHEESLGLELIHEESDAPLL PPRVYSTENHQPHHYTRRRIPQDHSESFFPLLTNEHTEDLQSPHRDSLYTSMPTLAGVAATESVTTST QTEPPPAKCGDAEDVYYKSMPNLGSRNHVHQLHTYYQLGRGSSDGFIVPPNKDGTPPEGSSKGPAH
LVTSL
>sp|O95490|AGRL2_HUMAN Adhesion G protein-coupled receptor L2 OS=Homo sapiens OX=9606 GN=ADGRL2 PE=1 SV=2
MVSSGCRMRSLWFIIVISFLPNTEGFSRAALPFGLVRRELSCEGYSIDLRCPGSDVIMIESANYGRTDD
KICDADPFQMENTDCYLPDAFKIMTQRCNNRTQCIVVTGSDVFPDPCPGTYKYLEVQYECVPYIFVCP
GTLKAIVDSPCIYEAEQKAGAWCKDPLQAADKIYFMPWTPYRTDTLIEYASLEDFQNSRQTTTYKLPN
RVDGTGFVVYDGAVFFNKERTRNIVKFDLRTRIKSGEAIINYANYHDTSPYRWGGKTDIDLAVDENGL WVIYATEQNNGMIVISQLNPYTLRFEATWETVYDKRAASNAFMICGVLYVVRSVYQDNESETGKNSID YIYNTRLNRGEYVDVPFPNQYQYIAAVDYNPRDNQLYVWNNNFILRYSLEFGPPDPAQVPTTAVTITS
SAELFKTIISTTSTTSQKGPMSTTVAGSQEGSKGTKPPPAVSTTKIPPITNIFPLPERFCEALDSKGIKWP
QTQRGMMVERPCPKGTRGTASYLCMISTGTWNPKGPDLSNCTSHWVNQLAQKIRSGENAASLANEL
AKHTKGPVFAGDVSSSVRLMEQLVDILDAQLQELKPSEKDSAGRSYNKLQKREKTCRAYLKAIVDTVD NLLRPEALESWKHMNSSEQAHTATMLLDTLEEGAFVLADNLLEPTRVSMPTENIVLEVAVLSTEGQIQ DFKFPLGIKGAGSSIQLSANTVKQNSRNGLAKLVFIIYRSLGQFLSTENATIKLGADFIGRNSTIAVNSHV
ISVSINKESSRVYLTDPVLFTLPHIDPDNYFNANCSFWNYSERTMMGYWSTQGCKLVDTNKTRTTCAC
SHLINFA1LMAHREIAYKDGVHELLLTVITWVGIVISLVCLAICIFTFCFFRGLQSDRNTIHKNLCINLFIAE FIFLIGIDKTKYAIACPIFAGLLHFFFLAAFAWMCLEGVQLYLMLVEVFESEYSRKKYYYVAGYLFPATVV GVSAAIDYKSYGTEKACWLHVDNYFIWSFIGPVTFIILLNIIFLVITLCKMVKHSNTLKPDSSRLENIKSWV
LGAFALLCLLGLTWSFGLLFINEETIVMAYLFTIFNAFQGVFIFIFHCALQKKVRKEYGKCFRHSYCCGG
LPTESPHSSVKASTTRTSARYSSGTQSRIRRMWNDTVRKQSESSFISGDINSTSTLNQGMTGNYLLTN
PLLRPHGTNNPYNTLLAETVVCNAPSAPVFNSPGHSLNNARDTSAMDTLPLNGNFNNSYSLHKGDYN
DSVQVVDCGLSLNDTAFEKMIISELVHNNLRGSSKTHNLELTLPVKPVIGGSSSEDDAIVADASSLMHS
DNPGLELHHKELEAPLIPQRTHSLLYQPQKKVKSEGTDSYVSQLTAEAEDHLQSPNRDSLYTSMPNLR
DSPYPESSPDMEEDLSPSRRSENEDIYYKSMPNLGAGHQLQMCYQISRGNSDGYIIPINKEGCIPEGD VREGQMQLVTSL
>sp|O94910|AGRL1_HUMAN Adhesion G protein-coupled receptor L1 OS=Homo sapiens OX=9606
GN=ADGRL1 PE=1 SV=1
MARLAAVLWNLCVTAVLVTSATQGLSRAGLPFGLMRRELACEGYPIELRCPGSDVIMVENANYGRTD
DKICDADPFQMENVQCYLPDAFKIMSQRCNNRTQCWVAGSDAFPDPCPGTYKYLEVQYDCVPYKV
EQKVFVCPGTLQKVLEPTSTHESEHQSGAWCKDPLQAGDRIYVMPWIPYRTDTLTEYASWEDYVAA
RHTTTYRLPNRVDGTGFVVYDGAVFYNKERTRNIVKYDLRTRIKSGETVINTANYHDTSPYRWGGKTD
IDLAVDENGLWVIYATEGNNGRLVVSQLNPYTLRFEGTWETGYDKRSASNAFMVCGVLYVLRSVYVD
DDSEAAGNRVDYAFNTNANREEPVSLTFPNPYQFISSVDYNPRDNQLYVWNNYFVVRYSLEFGPPDP
SAGPATSPPLSTTTTARPTPLTSTASPAATTPLRRAPLTTHPVGAINQLGPDLPPATAPVPSTRRPPAP
NLHVSPELFCEPREVRRVQWPATQQGMLVERPCPKGTRGIASFQCLPALGLWNPRGPDLSNCISP
WVNQVAQKIKSGENAANIASELARHTRGSIYAGDVSSSVKLMEQLLDILDAQLQALRPIERESAGKNY
NKMHKRERTCKDYIKAWETVDNLLRPEALESWKDMNATEQVHTATMLLDVLEEGAFLLADNVREPA
RFLAAKENVVLEVTVLNTEGQVQELVFPQEEYPRKNSIQLSAKTIKQNSRNGVVKVVFILYN/VLGLFLS
TENATVKLAGEAGPGGPGGASLVVNSQVIAASINKESSRVFLMDPVIFTVAHLEDKNHFNANCSFWNY
SERSMLGYWSTQGCRLVESNKTHTTCACSHLTNFAVLMAHREIYQGRINELLLSVITWVGIVISLVCLAI
CISTFCFLRGLQTDRNTIHKNLCINLFLAELLFLVGIDKTQYEIACPIFAGLLHYFFLAAFSWLCLEGVHLY
LLLVEVFESEYSRTKYYYLGGYCFPALVVGIAAAIDYRSYGTEKACWLRVDNYFIWSFIGPVSFVIWNL
VFLMVTLHKMIRSSSVLKPDSSRLDNIKSWALGAIALLFLLGLTWAFGLLFINKESVVMAYLFTTFNAFQ
GVFIFVFHCALQKKVHKEYSKCLRHSYCCIRSPPGGTHGSLKTSAMRSNTRYYTGTQSRIRRMWNDT
VRKQTESSFMAGDINSTPTLNRGTMGNHLLTNPVLQPRGGTSPYNTLIAESVGFNPSSPPVFNSPGS
YREPKHPLGGREACGMDTLPLNGNFNNSYSLRSGDFPPGDGGPEPPRGRNLADAAAFEKMIISELVH
NNLRGSSSAAKGPPPPEPPVPPVPGGGGEEEAGGPGGADRAEIELLYKALEEPLLLPRAQSVLYQSD
LDESESCTAEDGATSRPLSSPPGRDSLYASGANLRDSPSYPDSSPEGPSEALPPPPPAPPGPPEIYYT SRPPALVARNPLQGYYQVRRPSHEGYLAAPGLEGPGPDGDGQMQLVTSL
>sp|Q96K78|AGRG7_HUMAN Adhesion G-protein coupled receptor G7 OS=Homo sapiens OX=9606
GN=ADGRG7 PE=1 SV=2
MASCRAWNLRVLVAVVCGLLTGIILGLGIWRIVIRIQRGKSTSSSSTPTEFCRNGGTWENGRCICTEEW
KGLRCTIANFCENSTYMGFTFARIPVGRYGPSLQTCGKDTPNAGNPMAVRLCSLSLYGEIELQKVTIG
NCNENLETLEKQVKDVTAPLNNISSEVQILTSDANKLTAENITSATRWGQIFNTSRNASPEAKKVAIVT
VSQLLDASEDAFQRVAATANDDALTTLIEQMETYSLSLGNQSWEPNIAIQSANFSSENAVGPSNVRF SVQKGASSSLVSSSTFIHTNVDGLNPDAQTELQVLLNMTKNYTKTCGFVVYQ/VDKLFQSKTFTAKSDF
SQKIISSKTDENEQDQSASVDMVFSPKYNQKEFQLYSYACVYWNLSAKDWDTYGCQKDKGTDGFLR
CRCNHTTNFAVLMTFKKDYQYPKSLDILSNVGCALSVTGLALTVIFQIVTRKVRKTSVTWVLVNLCISML
IFNLLFVFGIENSNKNLQTSDGDINNIDFDNNDIPRTDTINIPNPMCTAIAALLHYFLLVTFTWNALSAAQL
YYLLIRTMKPLPRHFILFISLIGWGVPAIVVAITVGVIYSQNGNNPQWELDYRQEKICWLAIPEPNGVIKS PLLWSFIVPVTIILISNWMFITISIKVLWKNNQNLTSTKKVSSMKKIVSTLSVAVVFGITWILAYLMLVNDD SIRIVFSYIFCLFNTTQGLQIFILYTVRTKVFQSEASKVLMLLSSIGRRKSLPSVTRPRLRVKMYNFLRSL
PTLHERFRLLETSPSTEEITLSESDNAKESI
>sp|Q86SQ4|AGRG6_HUMAN Adhesion G-protein coupled receptor G6 OS=Homo sapiens OX=9606
GN=ADGRG6 PE=1 SV=3
MMFRSDRMWSCHWKWKPSPLLFLFALYIMCVPHSVWGCANCRWLSNPSGTFTSPCYPNDYPNSQ
ACMWTLRAPTGYIIQITFNDFDIEEAPNCIYDSLSLDNGESQTKFCGATAKGLSFNSSANEMHVSFSSD
FSIQKKGFNASYIRVAVSLRNQKVILPQTSDAYQVSVAKSISIPELSAFTLCFEATKVGHEDSDWTAFSY
SNASFTQLLSFGKAKSGYFLSISDSKCLLNNALPVKEKEDIFAESFEQLCLVWNNSLGSIGVNFKRNYE
TVPCDSTISKVIPGNGKLLLGSNQNEIVSLKGDIYNFRLWNFTMNAKILSNLSCNVKGNVVDWQNDFW
NIPNLALKAESNLSCGSYLIPLPAAELASCADLGTLCQATVNSPSTTPPTVTTNMPVTNRIDKQRNDGII
YRISVVIQNILRHPEVKVQSKVAEWLNSTFQNWNYTVYVVNISFHLSAGEDKIKVKRSLEDEPRLVLWA
LLVYNATNNTNLEGKIIQQKLLKNNESLDEGLRLHTVNVRQLGHCLAMEEPKGYYWPSIQPSEYVLPC
PDKPGFSASRICFYNATNPLVTYWGPVDISNCLKEANEVANQILNLTADGQNLTSANITNIVEQVKRIVN
KEENIDITLGSTLMNIFSNILSSSDSDLLESSSEALKTIDELAFKIDLNSTSHVNITTRNLALSVSSLLPGTN
AISNFSIGLPSNNESYFQMDFESGQVDPLASVILPPNLLENLSPEDSVLVRRAQFTFFNKTGLFQDVGP
QRKTLVSYVMACSIGNITIQNLKDPVQIKIKHTRTQEVHHPICAFWDLNKNKSFGGWNTSGCVAHRDS
DASETVCLCNHFTHFGVLMDLPRSASQLDARNTKVLTFISYIGCGISAIFSAATLLTYVAFEKLRRDYPS
KILMNLSTALLFLNLLFLLDGWITSFNVDGLCIAVAVLLHFFLLATFTWMGLEAIHMYIALVKVFNTYIRRY
ILKFCIIGWGLPALVVSVVLASRNNNEVYGKESYGKEKGDEFCWIQDPVIFYVTCAGYFGVMFFLNIAM
FIVVMVQICGRNGKRSNRTLREEVLRNLRSVVSLTFLLGMTWGFAFFAWGPLNIPFMYLFSIFNSLQGL
FIFIFHCAMKENVQKQWRQHLCCGRFRLADNSDWSKTATNIIKKSSDNLGKSLSSSSIGSNSTYLTSKS KSSSTTYFKRNSHTDNVSYEHSFNKSGSLRQCFHGQVLVKTGPC
>sp|Q8IZF4|AGRG5_HUMAN Adhesion G-protein coupled receptor G5 OS=Homo sapiens OX=9606 GN=ADGRG5 PE=2 SV=3
MDHCGALFLCLCLLTLQNATTETWEELLSYMENMQVSRGRSSVFSSRQLHQLEQMLLNTSFPGYNLT
LQTPTIQSLAFKLSCDFSGLSLTSATLKRVPQAGGQHARGQHAMQFPAELTRDACKTRPRELRLICIYF
SA/THFFKDENNSSLLNNYVLGAQLSHGHVNNLRDPVNISFWHNQSLEGYTLTCVFWKEGARKQPWG
GWSPEGCRTEQPSHSQVLCRCNHLTYFAVLMQLSPALVPAELLAPLTYISLVGCSISIVASLITVLLHFH
FRKQSDSLTRIHMNLHASVLLLNIAFLLSPAFAMSPVPGSACTALAAALHYALLSCLTWMAIEGFNLYLL
LGRVYNIYIRRYVFKLGVLGWGAPALLVLLSLSVKSSVYGPCTIPVFDSWENGTGFQNMSICWVRSPV
VHSVLVMGYGGLTSLFNLVVLAWALWTLRRLRERADAPSVRACHDTVTVLGLTVLLGTTWALAFFSF GVFLLPQLFLFTILNSLYGFFLFLWFCSQRCRSEAEAKAQIEAFSSSQTTQ
>sp|Q8IZF6|AGRG4_HUMAN Adhesion G-protein coupled receptor G4 OS=Homo sapiens OX=9606
GN=ADGRG4 PE=2 SV=2
MKEHIIYQKLYGLILMSSFIFLSDTLSLKGKKLDFFGRGDTYVSLIDTIPELSRFTACIDLVFMDDNSRYW
MAFSYITNNALLGREDIDLGLAGDHQQLILYRLGKTFSIRHHLASFQWHTICLIWDGVKGKLELFLNKER
ILEVTDQPHNLTPHGTLFLGHFLKNESSEVKSMMRSFPGSLYYFQLWDHILENEEFMKCLDGNIVSWE
EDVWLVNKIIPTVDRTLRCFVPENMTIQEKSTTVSQQIDMTTPSQITGVKPQNTAHSSTLLSQSIPIFAT
DYTTISYSNTTSPPLETMTAQKILKTLVDETATFAVDVLSTSSAISLPTQSISIDNTTNSMKKTKSPSSES
TKTTKMVEAMATEIFQPPTPSNFLSTSRFTKNSVVSTTSAIKSQSAVTKTTSLFSTIESTSMSTTPCLKQ
KSTNTGALPISTAGQEFIESTAAGTVPWFTVEKTSPASTHVGTASSFPPEPVLISTAAPVDSVFPRNQT
AFPLATTDMKIAFTVHSLTLPTRLIETTPAPRTAETELTSTNFQDVSLPRVEDAMSTSMSKETSSKTFSF
LTSFSFTGTESVQTVIDAEATRTALTPEITLASTVAETMLSSTITGRVYTQNTPTADGHLLTLMSTRSAS
TSKAPESGPTSTTDEAAHLFSSNETIWTSRPDQALLASMNTTTILTFVPNENFTSAFHENTTYTEYLSA
TTNITPLKASPEGKGTTANDATTARYTTAVSKLTSPWFANFSIVSGTTSITNMPEFKLTTLLLKTIPMSTK
PANELPLTPRETVVPSVDIISTLACIQPNFSTEESASETTQTEINGAIVFGGTTTPVPKSATTQRLNATVT
RKEATSHYLMRKSTIAAVAEVSPFSTMLEVTDESAQRVTASVTVSSFPDIEKLSTPLDNKTATTEVRES
WLLTKLVKTTPRSSYNEMTEMFNFNHTYVAHWTSETSEGISAGSPTSGSTHIFGEPLGASTTRISETS
FSTTPTDRTATSLSDGILPPQPTAAHSSATPVPVTHMFSLPVNGSSVVAEETEVTMSEPSTLARAFST
SVLSDVSNLSSTTMTTALVPPLDQTASTTIVIVPTHGDLIRTTSEATVISVRKTSMAVPSLTETPFHSLRL
STPVTAKAETTLFSTSVDTVTPSTHTLVCSKPPPDNIPPASSTHVISTTSTPEATQPISQVEETSTYALS
FPYTFSGGGWASLATGTTETSWDETTPSHISANKLTTSVNSHISSSATYRVHTPVSIQLVTSTSVLSS
DKDQMTISLGKTPRTMEVTEMSPSKNSFISYSRGTPSLEMTDTGFPETTKISSHQTHSPSEIPLGTPSD
GNLASSPTSGSTQITPTLTSSNTVGVHIPEMSTSLGKTALPSQALTITTFLCPEKESTSALPAYTPRTVE
MIVNSTYVTHSVSYGQDTSFVDTTTSSSTRISNPMDINTTFSHLHSLRTQPEVTSVASFISESTQTFPES
LSLSTAGLYNDGFTVLSDRITTAFSVPNVPTMLPRESSMATSTPIYQMSSLPVNVTAFTSKKVSDTPPI
VITKSSKTMHPGCLKSPCTATSGPMSEMSSIPVNNSAFTPATVSSDTSTRVGLFSTLLSSVTPRTTMT
MQTSTLDVTPVIYAGATSKNKMVSSAFTTEMIEAPSRITPTTFLSPTEPTLPFVKTVPTTIMAGIVTPFVG
TTAFSPLSSKSTGAISSIPKTTFSPFLSATQQSSQADEATTLGILSGITNRSLSTVNSGTGVALTDTYSRI
TVPENMLSPTHADSLHTSFNIQVSPSLTSFKSASGPTKNVKTTTNCFSSNTRKMTSLLEKTSLTNYATS
LNTPVSYPPWTPSSATLPSLTSFVYSPHSTEAEISTPKTSPPPTSQMVEFPVLGTRMTSSNTQPLLMT
SWNIPTAEGSQFPISTTINVPTSNEMETETLHLVPGPLSTFTASQTGLVSKDVMAMSSIPMSGILPNHG
LSENPSLSTSLRAITSTLADVKHTFEKMTTSVTPGTTLPSILSGATSGSVISKSPILTWLLSSLPSGSPPA
TVSNAPHVMTSSTVEVSKSTFLTSDMISAHPFTNLTTLPSATMSTILTRTIPTPTLGGITTGFPTSLPMSI
NVTDDIVYISTHPEASSRTTITANPRTVSHPSSFSRKTMSPSTTDHTLSVGAMPLPSSTITSSWNRIPTA
SSPSTLIIPKPTLDSLLNIMTTTSTVPGASFPLISTGVTYPFTATVSSPISSFFETTWLDSTPSFLSTEAST
SPTATKSTVSFYNVEMSFSVFVEEPRIPITSVINEFTENSLNSIFQNSEFSLATLETQIKSRDISEEEMVM
DRAILEQREGQEMATISYVPYSCVCQVIIKASSSLASSELMRKIKSKIHGNFTHGNFTQDQLTLLVNCEH
VAVKKLEPGNCKADETASKYKGTYKWLLTNPTETAQTRCIKNEDGNATRFCSISINTGKSQWEKPKFK
QCKLLQELPDKIVDLANITISDENAEDVAEHILNLINESPALGKEETKIIVSKISDISQCDEISMNLTHVMLQ
IINV'/LEKQNNSASDLHEISNEILRIIERTGHKMEFSGQIANLTVAGLALAVLRGDHTFDGMAFSIHSYEE
GTDPEIFLGNVPVGGILASIYLPKSLTERIPLSNLQTILFNFFGQTSLFKTKNVTKALTTYVVSASISDDMF
IQNLADPWITLQHIGGNQNYGQVHCAFWDFENNNGLGGWNSSGCKVKETNVNYTICQCDHLTHFG VLMDLSRSTVDSVNEQILALITYTGCGISSIFLGVAWTYIAFHKLRKDYPAKILINLCTALLMLNLVFLINS WLSSFQKVGVCITAAVALHYFLLVSFTWMGLEAVHMYLALVKVFNIYIPNYILKFCLVGWGIPAIMVAIT
VSVKKDLYGTLSPTTPFCWIKDDSIFYISVVAYFCLIFLMNLSMFCTVLVQLNSVKSQIQKTRRKMILHDL KGTMSLTFLLGLTWGFAFFAWGPMRNFFLYLFAIFNTLQGFFIFVFHCVMKESVREQWQIHLCCGWL RLDNSSDGSSRCQIKVGYKQEGLKKIFEHKLLTPSLKSTATSSTFKSLGSAQGTPSEISFPNDDFDKDP
YCSSP
>sp|Q86Y34|AGRG3_HUMAN Adhesion G protein-coupled receptor G3 OS=Homo sapiens OX=9606 GN=ADGRG3 PE=1 SV=1
MATPRGLGALLLLLLLPTSGQEKPTEGPRNTCLGSNNMYDIFNLNDKALCFTKCRQSGSDSCNVENL
QRYWLNYEAHLMKEGLTQKVNTPFLKALVQNLSTNTAEDFYFSLEPSQVPRQVMKDEDKPPDRVRL PKSLFRSLPGNRSVVRLAVTILDIGPGTLFKGPRLGLGDGSGVLNNRLVGLSVGQMHVTKLAEPLEIVF SHQRPPPNMTLTCVFWDVTKGTTGDWSSEGCSTEVRPEGTVCCCDHLTFFALLLRPTLDQSTVHILT
RISQAGCGVSMIFLAFTIILYAFLRLSRERFKSEDAPKIHVALGGSLFLLNLAFLVNVGSGSKGSDAACW
ARGAVFHYFLLCAFTWMGLEAFHLYLLAVRVFNTYFGHYFLKLSLVGWGLPALMVIGTGSANSYGLYT
IRDRENRTSLELCWFREGTTMYALYITVHGYFLITFLFGMVVLALWWKIFTLSRATAVKERGKNRKKV
LTLLGLSSLVGVTWGLAIFTPLGLSTVYIFALFNSLQGVFICCWFTILYLPSQSTTVSSSTARLDQAHSA SQE
>sp|Q8IZP9|AGRG2_HUMAN Adhesion G-protein coupled receptor G2 OS=Homo sapiens OX=9606 GN=ADGRG2 PE=1 SV=2
MVFSVRQCGHVGRTEEVLLTFKIFLVIICLHVVLVTSLEEDTDNSSLSPPPAKLSVVSFAPSSNGTPEVE
TTSLNDVTLSLLPSNETEKTKITIVKTFNASGVKPQRNICNLSSICNDSAFFRGEIMFQYDKESTVPQNQ
HITNGTLTGVLSLSELKRSELNKTLQTLSETYFIMCATAEAQSTLNCTFTIKLNNTMNACAVIAALERVKI
RPMEHCCCSVRIPCPSSPEELEKLQCDLQDPIVCLADHPRGPPFSSSQSIPVVPRATVLSQVPKATSF
AEPPDYSPVTHNVPSPIGEIQPLSPQPSAPIASSPAIDMPPQSETISSPMPQTHVSGTPPPVKASFSSP
TVSAPANVNTTSAPPVQTDIVNTSSISDLENQVLQMEKALSLGSLEPNLAGEMINQV'SRLLHSPPDML
APLAQRLLKVVDDIGLQLNFSNTTISLTSPSLALAVIRVNASSFNTTTFVAQDPANLQVSLETQAPENSI GTITLPSSLMNNLPAHDMELASRVQFNFFE7PALFQDPSLENLSLISYVISSSVANLTVRNLTRNVTVTL KHINPSQDELTVRCVFWDLGRNGGRGGWSDNGCSVKDRRLNETICTCSHLTSFGVLLDLSRTSVLPA
QMMALTFITYIGCGLSSIFLSVTLVTYIAFEKIRRDYPSKILIQLCAALLLLNLVFLLDSWIALYKMQGLCIS
VAVFLHYFLLVSFTWMGLEAFHMYLALVKVFNTYIRKYILKFCIVGWGVPAVVVTIILTISPDNYGLGSY
GKFPNGSPDDFCWINNNAVFYITVVGYFCVIFLLNVSMFIWLVQLCRIKKKKQLGAQRKTSIQDLRSIA GLTFLLGITWGFAFFAWGPVNVTFMYLFAIFNTLQGFFIFIFYCVAKENVRKQWRRYLCCGKLRLAENS DWSKTATNGLKKQTVNQGVSSSSNSLQSSSNSTNSTTLLVNNDCSVHASGNGNASTERNGVSFSVQ NGDVCLHDFTGKQHMFNEKEDSCNGKGRMALRRTSKRGSLHFIEQM
>sp|Q8IZF2|AGRF5_HUMAN Adhesion G protein-coupled receptor F5 OS=Homo sapiens OX=9606 GN=ADGRF5 PE=1 SV=3
MKSPRRTTLCLMFIVIYSSKAALNWNYESTIHPLSLHEHEPAGEEALRQKRAVATKSPTAEEYTVNIEIS
FENASFLDPIKAYLNSLSFPIHGNNTDQITDILSINVTTVCRPAGNEIWCSCETGYGWPRERCLHNLICQ
ERDVFLPGHHCSCLKELPPNGPFCLLQEDVTLNMRVRLNVGFQEDLMNTSSALYRSYKTDLETAFRK
GYGILPGFKGVTVTGFKSGSVVVTYEVKTTPPSLELIHKANEQWQSLNQTYKMDYNSFQAVTINESN
FFVTPEIIFEGDTVSLVCEKEVLSSNVSWRYEEQQLEIQNSSRFSIYTALFNNMTSVSKLTIHNITPGDA
GEYVCKLILDIFEYECKKKIDVMPIQILANEEMKVMCDNNPVSLNCCSQGNVNWSKVEWKQEGKINIP
GTPETDIDSSCSRYTLKADGTQCPSGSSGTTVIYTCEFISAYGARGSANIKVTFISVANLTITPDPISVSE
GQNFSIKCISDVSNYDEVYWNTSAGIKIYQRFYTTRRYLDGAESVLTVKTSTREWNGTYHCIFRYKNS
YSIATKDVIVHPLPLKLNIMVDPLEATVSCSGSHHIKCCIEEDGDYKVTFHTGSSSLPAAKEVNKKQVCY
KHNFNASSVSWCSKTVDVCCHFTNAANNSVWSPSMKLNLVPGENITCQDPVIGVGEPGKVIQKLCRF
SNVPSSPESPIGGTITYKCVGSQWEEKRNDCISAPINSLLQMAKALIKSPSQDEMLPTYLKDLSISIDKA
EHEISSSPGSLGAIINILDLLSTVPTQVNSEMMTHVLSTVNVILGKPVLNTWKVLQQQWTNQSSQLLHS
VERFSQALQSGDSPPLSFSQTNVQMSSMVIKSSHPETYQQRFVFPYFDLWGNWIDKSYLENLQSDS
SIVTMAFP7LQAILAQDIQENNFAESLVMTTTVSHNTTMPFRISMTFKNNSPSGGETKCVFWNFRLAN
NTGGWDSSGCYVEEGDGDNVTCICDHLTSFSILMSPDSPDPSSLLGILLDIISYVGVGFSILSLAACLVV
EAVVWKSVTKNRTSYMRHTCIVNIAASLLVANTWFIVVAAIQDNRYILCKTACVAATFFIHFFYLSVFFW
MLTLGLMLFYRLVFILHETSRSTQKAIAFCLGYGCPLAISVITLGATQPREVYTRKNVCWLNWEDTKAL
LAFAIPALIIVVVNITITIWITKILRPSIGDKPCKQEKSSLFQISKSIGVLTPLLGLTWGFGLTTVFPGTNLVF
HIIFAILNVFQGLFILLFGCLWDLKVQEALLNKFSLSRWSSQHSKSTSLGSSTPVFSMSSPISRRFNNLF
GKTGTYNVSTPEATSSSLENSSSASSLLN
>sp|Q8IZF3|AGRF4_HUMAN Adhesion G protein-coupled receptor F4 OS=Homo sapiens OX=9606
GN=ADGRF4 PE=1 SV=3
MKMKSQATMICCLVFFLSTECSHYRSKIHLKAGDKLQSPEGKPKTGRIQEKCEGPCISSSNCSQPCAK
DFHGEIGFTCNQKKWQKSAETCTSLSVEKLFKDSTGASRLSVAAPSIPLHILDFRAPETIESVAQGIRK
NCPFDYACITDMVKSSETTSGNIAFIVELLKNISTDLSDNVTREKMKSYSEVANHILDTAA/SNWAFIPNK
NASSDLLQSVNLFARQLHIHNNSENIVNELFIQTKGFHINHNTSEKSLNFSMSMNNTTEDILGMVQIPR
QELRKLWPNASQAISIAFP7LGAILREAHLQNVSLPRQVNGLVLSVVLPERLQEIILTFEKINKTRNARAQ
CVGWHSKKRRWDEKACQMMLDIRNEVKCRCNYTSVVMSFSILMSSKSMTDKVLDYITCIGLSVSILSL
VLCLIIEATVWSRVVVTEISYMRHVCIVNIAVSLLTANVWFIIGSHFNIKAQDYNMCVAVTFFSHFFYLSL
FFWMLFKALLIIYGILVIFRRMMKSRMMVIGFAIGYGCPLIIAVTTVAITEPEKGYMRPEACWLNWDNTK
ALLAFAIPAFVIVAVNLIVVLVVAVNTQRPSIGSSKSQDVVIIMRISKNVAILTPLLGLTWGFGIATLIEGTS
LTFHIIFALLNAFQGFFILLFGTIMDHKIRDALRMRMSSLKGKSRAAENASLGPTNGSKLMNRQG
>sp|Q8IZF5|AGRF3_HUMAN Adhesion G-protein coupled receptor F3 OS=Homo sapiens OX=9606
GN=ADGRF3 PE=2 SV=1
MVCSAAPLLLLATTLPLLGSPVAQASQPVSETGVRPREGLQRRQWGPLIGRDKAWNERIDRPFPACPI
PLSSSFGRWPKGQTMWAQTSTLTLTEEELGQSQAGGESGSGQLLDQENGAGESALVSVYVHLDFP
DKTWPPELSRTLTLPAASASSSPRPLLTGLRLTTECNVNHKGNFYCACLSGYQWNTSICLHYPPCQSL
HNHQPCGCLVFSHPEPGYCQLLPPGSPVTCLPAVPGILNLNSQLQMPGDTLSLTLHLSQEATNLSWF
LRHPGSPSPILLQPGTQVSVTSSHGQAALSVSNMSHHWAGEYMSCFEAQGFKWNLYEVVRVPLKAT DVARLPYQLSISCATSPGFQLSCCIPSTNLAYTAAWSPGEGSKASSFNESGSQCFVLAVQRCPMADT TYACDLQSLGLAPLRVPISITIIQDGDITCPEDASVLTWNVTKAGHVAQAPCPESKRGIVRRLCGADGV
WGPVHSSCTDARLLALFTRTKLLQAGQGSPAEEVPQILAQLPGQAAEASSPSDLLTLLSTMKYVAKVV AEARIQLDRRALKNLLIATDKVLDMDTRSLWTLAQARKPWAGSTLLLAVETLACSLCPQDHPFAFSLP NVLLQSQLFGPTFPADYSISFPTRPPLQAQIPRHSLAPLVRNGTEISITSLVLRKLDHLLPSNYGQGLGD
SLYATPGLVLVISIMAGDRAFSQGEVIMDFGNTDGSPHCVFWDHSLFQGRGGWSKEGCQAQVASAS PTAQCLCQHLTAFSVLMSPHTVPEEPALALLTQVGLGASILALLVCLGVYWLVWRVVVRNKISYFRHA ALLNMVFCLLAADTCFLGAPFLSPGPRSPLCLAAAFLCHFLYLATFFWMLAQALVLAHQLLFVFHQLAK
HRVLPLMVLLGYLCPLGLAGVTLGLYLPQGQYLREGECWLDGKGGALYTFVGPVLAIIGVNGLVLAMA MLKLLRPSLSEGPPAEKRQALLGVIKALLILTPIFGLTWGLGLATLLEEVSTVPHYIFTILNTLQGVFILLF GCLMDRKIQEALRKRFCRAQAPSSTISLVSCCLQILSCASKSMSEGIPWPSSEDMGTARS
>sp|Q8IZF7|AGRF2_HUMAN Adhesion G-protein coupled receptor F2 OS=Homo sapiens OX=9606 GN=ADGRF2 PE=2 SV=1
MGLTAYGNRRVQPGELPFGANLTLIHTRAQPVICSKLLLTKRVSPISFFLSKFQNSWGEDGWVQLDQL
PSPNAVSSDQVHCSAGCTHRKCGWAASKSKEKVPARPHGVCDGVCTDYSQCTQPCPPDTQGNMG FSCRQKTWHKITDTCQTLNALNIFEEDSRLVQPFEDNIKISVYTGKSETITDMLLQKCPTDLSCVIRNIQ QSPWIPGNIAVIVQLLHNISTAIWTGVDEAKMQSYSTIANHILNSKS/SNWTFIPDRNSSYILLHSVNSFA
RRLFIDKHPVDISDVFIHTMGTTISGDNIGKNFTFSMRINDTSNEVTGRVLISRDELRKVPSPSQVISIAF
P71GAILEASLLENVTVNGLVLSAILPKELKRISLIFEKISKSEERRTQCVGWHSVENRWDQQACKMIQE NSQQAVCKCRPSKLFTSFSILMSPHILESLILTYITYVGLGISICSLILCLSIEVLVWSQVTKTEITYLRHVCI VNIAATLLMADVWFIVASFLSGPITHHKGCVAATFFVHFFYLSVFFWMLAKALLILYGIMIVFHTLPKSVL
VASLFSVGYGCPLAIAAITVAATEPGKGYLRPEICWLNWDMTKALLAFVIPALAIVVVNLITVTLVIVKTQ RAAIGNSMFQEVRAIVRISKNIAILTPLLGLTWGFGVATVIDDRSLAFHIIFSLLNAFQVSPDASDQVQSE RIHEDVL
>sp|Q5T601 |AGRF1 _HUMAN Adhesion G-protein coupled receptor F1 OS=Homo sapiens OX=9606 GN=ADGRF1 PE=1 SV=2
MKVGVLWLISFFTFTDGHGGFLGKNDGIKTKKELIVNKKKHLGPVEEYQLLLQVTYRDSKEKRDLRNFL
KLLKPPLLWSHGLIRIIRAKATTDCNSLNGVLQCTCEDSYTWFPPSCLDPQNCYLHTAGALPSCECHL
NNLSQSVNFCERTKIWGTFKINERFTNDLLNSSSAIYSKYANGIEIQLKKAYERIQGFESVQVTQFRNG SIVAGYEWGSSSASELLSAIEHVAEKAKTALHKLFPLEDGSFRVFGKAQCNDIVFGFGSKDDEYTLPC SSGYRGNITAKCESSGWQVIRETCVLSLLEELNKNFSMIVGNATEAAVSSFVQNLSVIIRQNPSTTVGN
LASVVSILSNISSLSLASHFRVSNSTMEDVISIADNILNSASVTNWTVLLREEKYASSRLLETLENISTLVP
PTALPLNFSRKFIDWKGIPVNKSQLKRGYSYQIKMCPQNTSIPIRGRVLIGSDQFQRSLPETIISMASL7L GNILPVSKNGNAQVNGPVISTVIQNYSINEVFLFFSKIESNLSQPHCVFWDFSHLQWNDAGCHLVNET QDIVTCQCTHLTSFSILMSPFVPSTIFPVVKWITYVGLGISIGSLILCLIIEALFWKQIKKSQTSHTRRICMV
NIALSLLIADVWFIVGATVDTTVNPSGVCTAAVFFTHFFYLSLFFWMLMLGILLAYRIILVFHHMAQHLM MAVGFCLGYGCPLIISVITIAVTQPSNTYKRKDVCWLNWSNGSKPLLAFVVPALAIVAVNFVWLLVLTK
LWRPTVGERLSRDDKATIIRVGKSLLILTPLLGLTWGFGIGTIVDSQNLAWHVIFALLNAFQGFFILCFGI LLDSKLRQLLFNKLSALSSWKQTEKQNSSDLSAKPKFSKPFNPLQNKGHYAFSHTGDSSDNIMLTQFV SNE
>sp|P48960|AGRE5_HUMAN Adhesion G protein-coupled receptor E5 OS=Homo sapiens OX=9606 GN=ADGRE5 PE=1 SV=4
MGGRVFLAFCVWLTLPGAETQDSRGCARWCPQNSSCVNATACRCNPGFSSFSEIITTPTETCDDINE CATPSKVSCGKFSDCWNTEGSYDCVCSPGYEPVSGAKTFKNESENTCQDVDECQQNPRLCKSYGT CVNTLGSYTCQCLPGFKFIPEDPKVCTDVNECTSGQNPCHSSTHCLNNVGSYQCRCRPGWQPIPGS PNGPNNTVCEDVDECSSGQHQCDSSTVCFNTVGSYSCRCRPGWKPRHGIPNNQKDTVCEDMIFST WTPPPGVHSQTLSRFFDKVQDLGRDSKTSSAEVTIQNVIKLVDELMEAPGDVEALAPPVRHLIATQLL SNLEDIMRILAKSLPKGPFTYISPSNTELTLMIQERGDKNVTMGQSSARMKLNWAVAAGAEDPGPAVA GILSIQ/VMTTLLANASLNLHSKKQAELEEIYESSIRGVQLRRLSAVNSIFLSHNNTKELNSPILFAFSHLE
SSDGEAGRDPPAKDVMPGPRQELLCAFWKSDSDRGGHWATEGCQVLGSKNGSTTCQCSHLSSFAI LMAHYDVEDWKLTLITRVGLALSLFCLLLCILTFLLVRPIQGSRTTIHLHLCICLFVGSTIFLAGIENEGGQ
VGLRCRLVAGLLHYCFLAAFCWMSLEGLELYFLVVRVFQGQGLSTRWLCLIGYGVPLLIVGVSAAIYSK GYGRPRYCWLDFEQGFLWSFLGPVTFIILCNAVIFVTTVWKLTQKFSEINPDMKKLKKARALTITAIAQL FLLGCTWVFGLFIFDDRSLVLTYVFTILNCLQGAFLYLLHCLLNKKVREEYRKWACLVAGGSKYSEFTS TTSGTGHNQTRALRASESGI
>sp|Q86SQ3|AGRE4_HUMAN Putative adhesion G protein-coupled receptor E4P OS=Homo sapiens OX=9606 GN=ADGRE4P PE=5 SV=1
MGSRFLLVLLSGASCPPCPKYASCHNSTHCTCEDGFRARSGRTYFHDSSEKCEDINECETGLAKCKY KAYCRNKVGGYICSCLVKYTLFNFLAGIIDYDHPDCYENNSQGTTQSNVDIWVSGVKPGFGKQLPGD KRTKHICVYWEGSEGGWSTEGCSHVHSNGSYTKCKCFHLSSFAVLVALAPKEDPVLTVITQVGLTISL LCLFLAILTFLLCRPIQNTSTSLHLELSLCLFLAHLLFLTGINRTEPEVLCSIIAGLLHFLYLACFTWMLLEG LHLFLTVRNLKVANYTSTGRFKKRFMYPVGYGIPAVIIAVSAIVGPQNYGTFTCWLKLDKGFIWSFMGP
VAVIILINLVFYFQVLWILRSKLSSLNKEVSTIQDTRVMTFKAISQLFILGCSWGLGFFMVEEVGKTIGSIIA YSFTIINTLQGVLLFVVHCLLNRQVRLIILSVISLVPKSN
>sp|Q9BY15|AGRE3_HUMAN Adhesion G protein-coupled receptor E3 OS=Homo sapiens OX=9606 GN=ADGRE3 PE=2 SV=2
MQGPLLLPGLCFLLSLFGAVTQKTKTSCAKCPPNASCVNNTHCTCNHGYTSGSGQKLFTFPLETCNDI NECTPPYSVYCGFNAVCYNVEGSFYCQCVPGYRLHSGNEQFSNSNENTCQDTTSSKTTEGRKELQK IVDKFESLLTNQTLWRTEGRQEISSTATTILRDVESKVLETALKDPEQKVLKIQNDSVAIETQAITDNCSE ERKTFNLNVQMNSMDIRCSDIIQGDTQGPSAIAFISYSSLGNIINATFFEEMDKKDQVYLNSQVVSAAIG PKRNVSLSKSVTLTFQHVKMTPSTKKVFCVYWKSTGQGSQWSRDGCFLIHVNKSHTMCNCSHLSSF AVLMALTSQEEDPVLTVITYVGLSVSLLCLLLAALTFLLCKAIRNTSTSLHLQLSLCLFLAHLLFLVGIDRT EPKVLCSIIAGALHYLYLAAFTWMLLEGVHLFLTARNLTVVNYSSINRLMKWIMFPVGYGVPAVTVAISA
ASWPHLYGTADRCWLHLDQGFMWSFLGPVCAIFSANLVLFILVFWILKRKLSSLNSEVSTIQNTRMLA
FKATAQLFILGCTWCLGLLQVGPAAQVMAYLFTIINSLQGFFIFLVYCLLSQQVQKQYQKWFREIVKSK
SESETYTLSSKMGPDSKPSEGDVFPGQVKRKY
>sp|Q9UHX3|AGRE2_HUMAN Adhesion G protein-coupled receptor E2 OS=Homo sapiens OX=9606
GN=ADGRE2 PE=1 SV=2
MGGRVFLVFLAFCVWLTLPGAETQDSRGCARWCPQDSSCVNATACRCNPGFSSFSEIITTPMETCD
DINECATLSKVSCGKFSDCWNTEGSYDCVCSPGYEPVSGAKTFKNESENTCQDVDECQQNPRLCKS
YGTCVNTLGSYTCQCLPGFKLKPEDPKLCTDVNECTSGQNPCHSSTHCLNNVGSYQCRCRPGWQPI
PGSPNGPNNTVCEDVDECSSGQHQCDSSTVCFNTVGSYSCRCRPGWKPRHGIPNNQKDTVCEDMT
FSTWTPPPGVHSQTLSRFFDKVQDLGRDYKPGLANNTIQSILQALDELLEAPGDLETLPRLQQHCVAS
HLLDGLEDVLRGLSKNLSNGLLNFSYPAGTELSLEVQKQVDRSVTLRQNQAVMQLDWNQAQKSGDP
GPSVVGLVSIPGMGKLLAEAPLVLEPEKQMLLHETHQGLLQDGSPILLSDVISAFLSNNDTQNLSSPVT
FTFSHRSVIPRQKVLCVFWEHGQNGCGHWATTGCSTIGTRDTSTICRCTHLSSFAVLMAHYDVQEED
PVLTVITYMGLSVSLLCLLLAALTFLLCKAIQNTSTSLHLQLSLCLFLAHLLFLVAIDQTGHKVLCSIIAGTL
HYLYLATLTWMLLEALYLFLTARNLTVVNYSSINRFMKKLMFPVGYGVPAVTVAISAASRPHLYGTPSR
CWLQPEKGFIWGFLGPVCAIFSVNLVLFLVTLWILKNRLSSLNSEVSTLRNTRMLAFKATAQLFILGCT
WCLGILQVGPAARVMAYLFTIINSLQGVFIFLVYCLLSQQVREQYGKWSKGIRKLKTESEMHTLSSSAK
ADTSKPSTVN
>sp|Q14246|AGRE1_HUMAN Adhesion G protein-coupled receptor E1 OS=Homo sapiens OX=9606
GN=ADGRE1 PE=2 SV=3
MRGFNLLLFWGCCVMHSWEGHIRPTRKPNTKGNNCRDSTLCPAYATCTNTVDSYYCACKQGFLSSN
GQNHFKDPGVRCKDIDECSQSPQPCGPNSSCKNLSGRYKCSCLDGFSSPTGNDWVPGKPGNFSCT
DINECLTSSVCPEHSDCVNSMGSYSCSCQVGFISRNSTCEDVDECADPRACPEHATCNNTVGNYSC
FCNPGFESSSGHLSFQGLKASCEDIDECTEMCPINSTCTNTPGSYFCTCHPGFAPSNGQLNFTDQGV
ECRDIDECRQDPSTCGPNSICTNALGSYSCGCIAGFHPNPEGSQKDGNFSCQRVLFKCKEDVIPDNK
QIQQCQEGTAVKPAYVSFCAQINNIFSVLDKVCENKTTVVSLKNTTESFVPVLKQ/STWTKFTKEETSS
LATVFLESVESMTLASFWKPSANITPAVRTEYLDIESKVINKECSEENVTLDLVAKGDKMKIGCSTIEES
ESTETTGVAFVSFVGMESVLNERFFKDHQAPLTTSEIKLKMNSRVVGGIMTGEKKDGFSDPIIYTLENI
QPKQKFERPICVSWSTDVKGGRWTSFGCVILEASETYTICSCNQMANLAVIMASGELTMDFSLYI I SH V
GIIISLVCLVLAIATFLLCRSIRNHNTYLHLHLCVCLLLAKTLFLAGIHKTDNKMGCAIIAGFLHYLFLACFF
WMLVEAVILFLMVRNLKVVNYFSSRNIKMLHICAFGYGLPMLVVVISASVQPQGYGMHNRCWLNTET
GFIWSFLGPVCTVIVINSLLLTWTLWILRQRLSSVNAEVSTLKDTRLLTFKAFAQLFILGCSWVLGIFQIG
PVAGVMAYLFTIINSLQGAFIFLIHCLLNGQVREEYKRWITGKTKPSSQSQTSRILLSSMPSASKTG
>sp|Q7Z7M1 |AGRD2_HUMAN Adhesion G-protein coupled receptor D2 OS=Homo sapiens OX=9606
GN=ADGRD2 PE=2 SV=1
MDAPWGAGERWLHGAAVDRSGVSLGPPPTPQVNQGTLGPQVAPVAAGEVVKTAGGVCKFSGQRL
SWWQAQESCEQQFGHLALQPPDGVLASRLRDPVWVGQREAPLRRPPQRRARTTAVLVFDERTADR
AARLRSPLPELAALTACTHVQWDCASPDPAALFSVAAPALPNALQLRAFAEPGGVVRAALWRGQHA
PFLAAFRADGRWHHVCATWEQRGGRWALFSDGRRRAGARGLGAGHPVPSGGILVLGQDQDSLGG
GFSVRHALSGNLTDFHLWARALSPAQLHRARACAPPSEGLLFRWDPGALDVTPSLLPTVWVRLLCPV
PSEECPTWNPGPRSEGSELCLEPQPFLCCYRTEPYRRLQDAQSWPGQDVISRVNALANDIVLLPDPL
SEVHGALSPAEASSFLGLLEHVLAMEMAPLGPAALLAVVRFLKRVVALGAGDPELLL7GPWEQLSQG
VVSVASLVLEEQVADTWLSLREVIGGPMALVASVQRLAPLLSTSMTSERPRMRIQHRHAGLSGVTVIH
SWFTSRVFQHTLEGPDLEPQAPASSEEANRVQRFLSTQVGSAIISSEVWDVTGEVNVAMTFHLQHRA
QSPLFPPHPPSPYTGGAWATTGCSVAALYLDSTACFCNHSTSFAILLQIYEVQRGPEEESLLRTLSFV
GCGVSFCALTTTFLLFLVAGVPKSERTTVHKNLTFSLASAEGFLMTSEWAKANEVACVAVTVAMHFLF
LVAFSWMLVEGLLLWRKWAVSMHPGPGMRLYHATGWGVPVGIVAVTLAMLPHDYVAPGHCWLNV
HTNAIWAFVGPVLFVLTANTCILARVVMITVSSARRRARMLSPQPCLQQQIWTQIWATVKPVLVLLPVL
GLTWLAGILVHLSPAWAYAAVGLNSIQGLYIFLVYAACNEEVRSALQRMAEKKVAEVLRALGVWGGAA
KEHSLPFSVLPLFLPPKPSTPRHPLKAPA
>sp|Q6QNK2|AGRD1_HUMAN Adhesion G-protein coupled receptor D1 OS=Homo sapiens OX=9606
GN=ADGRD1 PE=1 SV=1
MEKLLRLCCWYSWLLLFYYNFQVRGVYSRSQDHPGFQVLASASHYWPLENVDGIHELQDTTGDIVE
GKVNKGIYLKEEKGVTLLYYGRYNSSCISKPEQCGPEGVTFSFFWKTQGEQSRPIPSAYGGQVISNGF
KVCSSGGRGSVELYTRDNSMTWEASFSPPGPYWTHVLFTWKSKEGLKVYVNGTLSTSDPSGKVSR
DYGESNVNLVIGSEQDQAKCYENGAFDEFIIWERALTPDEIAMYFTAAIGKHALLSSTLPSLFMTSTAS
PVMPTDAYHPIITNLTEERKTFQSPGVILSYLQNVSLSLPSKSLSEQTALNLTKTFLKAVGEILL/-PGWIA
LSEDSAVVLSLIDTIDTVMGHVSSNLHGSTPQVTVEGSSAMAEFSVAKILPKTVNSSHYRFPAHGQSFI
QIPHEAFHRHAWSTVVGLLYHSMHYYLNNIWPAHTKIAEAMHHQDCLLFATSHLISLEVSPPPTLSQNL
SGSPLITVHLKHRLTRKQHSEATNSSNRVFVYCAFLDFSSGEGVWSNHGCALTRGNLTYSVCRCTHL
INFA1LMQVVPLELARGHQVALSSISYVGCSLSVLCLVATLVTFAVLSSVSTIRNQRYHIHANLSFAVLV
AQVLLLISFRLEPGTTPCQVMAVLLHYFFLSAFAWMLVEGLHLYSMVIKVFGSEDSKHRYYYGMGWG
FPLLICIISLSFAMDSYGTSNNCWLSLASGAIWAFVAPALFVIWNIGILIAVTRVISQISADNYKIHGDPSA
FKLTAKAVAVLLPILGTSWVFGVLAVNGCAVVFQYMFATLNSLQGLFIFLFHCLLNSEVRAAFKHKTKV
WSLTSSSARTSNAKPFHSDLMNGTRPGMASTKLSPWDKSSHSAHRVDLSAV
>sp|O60242|AGRB3_HUMAN Adhesion G protein-coupled receptor B3 OS=Homo sapiens OX=9606
GN=ADGRB3 PE=1 SV=2
MKAVRNLLIYIFSTYLLVMFGFNAAQDFWCSTLVKGVIYGSYSVSEMFPKNFTNCTWTLENPDPTKYSI
YLKFSKKDLSCSNFSLLAYQFDHFSHEKIKDLLRKNHSIMQLCNSKNAFVFLQYDKNFIQIRRVFPTNFP
GLQKKGEEDQKSFFEFLVLNKVSPSQFGCHVLCTWLESCLKSENGRTESCGIMYTKCTCPQHLGEW
GIDDQSLILLNNVVLPLNEQTEGCLTQELQTTQVCNLTREAKRPPKEEFGMMGDHTIKSQRPRSVHEK
RVPQEQADAAKFMAQTGESGVEEWSQWSTCSVTCGQGSQVRTRTCVSPYGTHCSGPLRESRVCN
NTALCPVHGVWEEWSPWSLCSFTCGRGQRTRTRSCTPPQYGGRPCEGPETHHKPCNIALCPVDGQ
WQEWSSWSQCSVTCSNGTQQRSRQCTAAAHGGSECRGPWAESRECYNPECTANGQWNQWGH
WSGCSKSCDGGWERRIRTCQGAVITGQQCEGTGEEVRRCNEQRCPAPYEICPEDYLMSMVWKRTP
AGDLAFNQCPLNATGTTSRRCSLSLHGVAFWEQPSFARCISNEYRHLQHSIKEHLAKGQRMLAGDG
MSQVTKTLLDLTQRKNFYAGDLLMSVEILRNVTDTFKRASYIPASDGVQNFFQIVSNLLDEENKEKWE DAQQIYPGSIELMQVIEDFIHIVGMGMMDFQNSYLMTGNVVASIQKLPAASVLTDINFPMKGRKGMVD WARNSEDRVVIPKSIFTPVSSKELDESSVFVLGAVLYK/VLDLILPTLRNYTVINSKIIVVTIRPEPKTTDSF LEIELAHLANGTLNPYCVLWDDSKTNESLGTWSTQGCKTVLTDASHTKCLCDRLSTFAILAQQPREI I M ESSGTPSVTLIVGSGLSCLALITLAVVYAALWRYIRSERSIILINFCLSIISSNILILVGQTQTHNKSICTTTT AFLHFFFLASFCWVLTEAWQSYMAVTGKIRTRLIRKRFLCLGWGLPALWATSVGFTRTKGYGTDHYC WLSLEGGLLYAFVGPAAAWLVNMVIGILVFNKLVSRDGILDKKLKHRAGQMSEPHSGLTLKCAKCGV VSTTALSATTASNAMASLWSSCWLPLLALTWMSAVLAMTDKRSILFQILFAVFDSLQGFVIVMVHCILR REVQDAFRCRLRNCQDPINADSSSSFPNGHAQIMTDFEKDVDIACRSVLHKDIGPCRAATITGTLSRIS LNDDEEEKGTNPEGLSYSTLPGNVISKVIIQQPTGLHMPMSMNELSNPCLKKENSELRRTVYLCTDDN LRGADMDIVHPQERMMESDYIVMPRSSVNNQPSMKEESKMNIGMETLPHERLLHYKVNPEFNMNPP VMDQFNMNLEQHLAPQEHMQNLPFEPRTAVKNFMASELDDNAGLSRSETGSTISMSSLERRKSRYS DLDFEKVMHTRKRHMELFQELNQKFQTLDRFRDIPNTSSMENPAPNKNPWDTFKNPSEYPHYTTINV LDTEAKDALELRPAEWEKCLNLPLDVQEGDFQTEV
>sp|O60241 |AGRB2_HUMAN Adhesion G protein-coupled receptor B2 OS=Homo sapiens OX=9606 GN=ADGRB2 PE=1 SV=2
MENTGWMGKGHRMTPACPLLLSVILSLRLATAFDPAPSACSALASGVLYGAFSLQDLFPTIASGCSWT LENPDPTKYSLYLRFNRQEQVCAHFAPRLLPLDHYLVNFTCLRPSPEEAVAQAESEVGRPEEEEAEA AAGLELCSGSGPFTFLHFDKNFVQLCLSAEPSEAPRLLAPAALAFRFVEVLLINNNNSSQFTCGVLCR WSEECGRAAGRACGFAQPGCSCPGEAGAGSTTTTSPGPPAAHTLSNALVPGGPAPPAEADLHSGS SNDLFTTEMRYGEEPEEEPKVKTQWPRSADEPGLYMAQTGDPAAEEWSPWSVCSLTCGQGLQVRT RSCVSSPYGTLCSGPLRETRPCNNSATCPVHGVWEEWGSWSLCSRSCGRGSRSRMRTCVPPQHG GKACEGPELQTKLCSMAACPVEGQWLEWGPWGPCSTSCANGTQQRSRKCSVAGPAWATCTGALT DTRECSNLECPATDSKWGPWNAWSLCSKTCDTGWQRRFRMCQATGTQGYPCEGTGEEVKPCSEK RCPAFHEMCRDEYVMLMTWKKAAAGEIIYNKCPPNASGSASRRCLLSAQGVAYWGLPSFARC1SHEY RYLYLSLREHLAKGQRMLAGEGMSQVVRSLQELLARRTYYSGDLLFSVDILRNVTDTFKRATYVPSAD DVQRFFQVVSFMVDAENKEKWDDAQQVSPGSyHLLRVVEDFIHLVGDALKAFQSSLIVTDNLVISIQR EPVSAVSSDITFPMRGRRGMKDWVRHSEDRLFLPKEVLSLSSPGKPATSGAAGSPGRGRGPGTVPP GPGHSHQRLLPADPDESSYFVIGAVLYR7LGLILPPPRPPLAVTSRVMTVTVRPPTQPPAEPLITVELS YIINGTTDPHCASWDYSRADASSGDWDTENCQTLETQAAHTRCQCQHLSTFAVLAQPPKDLTLELAG SPSVPLVIGCAVSCMALLTLLAIYAAFWRFIKSERSIILLNFCLSILASNILILVGQSRVLSKGVCTMTAAFL
HFFFLSSFCWVLTEAWQSYLAVIGRMRTRLVRKRFLCLGWGLPALVVAVSVGFTRTKGYGTSSYCWL SLEGGLLYAFVGPAAVIVLVNMLIGIIVFNKLMARDGISDKSKKQRAGSERCPWASLLLPCSACGAVPS PLLSSASARNAMASLWSSCVVLPLLALTWMSAVLAMTDRRSVLFQALFAVFNSAQGFVITAVHCFLRR EVQDVVKCQMGVCRADESEDSPDSCKNGQLQILSDFEKDVDLACQTVLFKEVNTCNPSTITGTLSRL SLDEDEEPKSCLVGPEGSLSFSPLPGNILVPMAASPGLGEPPPPQEANPVYMCGEGGLRQLDLTWLR PTEPGSEGDYMVLPRRTLSLQPGGGGGGGEDAPRARPEGTPRRAAKTVAHTEGYPSFLSVDHSGL GLGPAYGSLQNPYGMTFQPPPPTPSARQVPEPGERSRTMPRTVPGSTMKMGSLERKKLRYSDLDFE
KVMHTRKRHSELYHELNQKFHTFDRYRSQSTAKREKRWSVSSGGAAERSVCTDKPSPGERPSLSQ
HRRHQSWSTFKSMTLGSLPPKPRERLTLHRAAAWEPTEPPDGDFQTEV
>sp|O14514|AGRB1_HUMAN Adhesion G protein-coupled receptor B1 OS=Homo sapiens OX=9606
GN=ADGRB1 PE=1 SV=2
MRGQAAAPGPVWILAPLLLLLLLLGRRARAAAGADAGPGPEPCATLVQGKFFGYFSAAAVFPANASR
CSWTLRNPDPRRYTLYMKVAKAPVPCSGPGRVRTYQFDSFLESTRTYLGVESFDEVLRLCDPSAPLA
FLQASKQFLQMRRQQPPQHDGLRPRAGPPGPTDDFSVEYLVVGNRNPSRAACQMLCRWLDACLAG
SRSSHPCGIMQTPCACLGGEAGGPAAGPLAPRGDVCLRDAVAGGPENCLTSLTQDRGGHGATGGW
KLWSLWGECTRDCGGGLQTRTRTCLPAPGVEGGGCEGVLEEGRQCNREACGPAGRTSSRSQSLR
STDARRREELGDELQQFGFPAPQTGDPAAEEWSPWSVCSSTCGEGWQTRTRFCVSSSYSTQCSGP
LREQRLCNNSAVCPVHGAWDEWSPWSLCSSTCGRGFRDRTRTCRPPQFGGNPCEGPEKQTKFCNI
ALCPGRAVDGNWNEWSSWSACSASCSQGRQQRTRECNGPSYGGAECQGHWVETRDCFLQQCPV
DGKWQAWASWGSCSVTCGAGSQRRERVCSGPFFGGAACQGPQDEYRQCGTQRCPEPHEICDED
NFGAVIWKETPAGEVAAVRCPRNATGLILRRCELDEEGIAYWEPPTYIRCVSIDYRNIQMMTREHLAKA
QRGLPGEGVSEVIQTLVEISQDGTSYSGDLLSTIDVLRNMTEIFRRAYYSPTPGDVQNFVQILSNLLAE
ENRDKWEEAQLAGPNAKELFRLVEDFVDVIGFRMKDLRDAYQVTDNLVLSIHKLPASGATDISFPMKG
WRATGDWAKVPEDRVTVSKSVFSTGLTEADEASVFVVGTVLYR/VLGSFLALQRNTTVLNSKVISVTVK
PPPRSLRTPLEIEFAHMYNGTTNQTCILWDETDVPSSSAPPQLGPWSWRGCRTVPLDALRTRCLCDR
LSTFA1LAQLSADANMEKATLPSVTLIVGCGVSSLTLLMLVIIYVSVWRYIRSERSVILINFCLSIISSNALIL
IGQTQTRNKVVCTLVAAFLHFFFLSSFCWVLTEAWQSYMAVTGHLRNRLIRKRFLCLGWGLPALWAI
SVGFTKAKGYSTMNYCWLSLEGGLLYAFVGPAAAVVLVNMVIGILVFNKLVSKDGITDKKLKERAGAS
LWSSCVVLPLLALTWMSAVLAVTDRRSALFQILFAVFDSLEGFVIVMVHCILRREVQDAVKCRVVDRQ
EEGNGDSGGSFQNGHAQLMTDFEKDVDLACRSVLNKDIAACRTATITGTLKRPSLPEEEKLKLAHAK
GPPTNFNSLPANVSKLHLHGSPRYPGGPLPDFPNHSLTLKRDKAPKSSFVGDGDIFKKLDSELSRAQE
KALDTSYVILPTATATLRPKPKEEPKYSIHIDQMPQTRLIHLSTAPEASLPARSPPSRQPPSGGPPEAPP
AQPPPPPPPPPPPPQQPLPPPPNLEPAPPSLGDPGEPAAHPGPSTGPSTKNENVATLSVSSLERRKS
RYAELDFEKIMHTRKRHQDMFQDLNRKLQHAAEKDKEVLGPDSKPEKQQTPNKRPWESLRKAHGTP TWVKKELEPLQPSPLELRSVEWERSGATIPLVGQDIIDLQTEV
>sp|Q8IWK6|AGRA3_HUMAN Adhesion G protein-coupled receptor A3 OS=Homo sapiens OX=9606
GN=ADGRA3 PE=1 SV=2
MEPPGRRRGRAQPPLLLPLSLLALLALLGGGGGGGAAALPAGCKHDGRPRGAGRAAGAAEGKVVCS
SLELAQVLPPDTLPNRTVTLILSNNKISELKNGSFSGLSLLERLDLRNNLISSIDPGAFWGLSSLKRLDLT
NNRIGCLNADIFRGLTNLVRLNLSGNLFSSLSQGTFDYLASLRSLEFQTEYLLCDCNILWMHRWVKEK
NITVRDTRCVYPKSLQAQPVTGVKQELLTCDPPLELPSFYMTPSHRQVVFEGDSLPFQCMASYIDQD
MQVLWYQDGRIVETDESQGIFVEKNMIHNCSLIASALTISNIQAGSTGNWGCHVQTKRGNNTRTVDIV
VLESSAQYCPPERWNNKGDFRWPRTLAGITAYLQCTRNTHGSGIYPGNPQDERKAWRRCDRGGF
WADDDYSRCQYANDVTRVLYMFNQMPLNLTNAVATARQLLAYTVEAANFSDKMDVIFVAEMIEKFGR
FTKEEKSKELGDVMVDIASNIMLADERVLWLAQREAKACSRIVQCLQRIATYRLAGGAHVYSTYSPNIA
LEAYVIKSTGFTGMTCTVFQKVAASDRTGLSDYGRRDPEGNLDKQLSFKCNVSNTFSSLALKNTIVEA
SIQLPPSLFSPKQKRELRPTDDSLYKLQLIAFR/VGKLFPATGNSTNLADDGKRRTWTPVILTKIDGVNV
DTHHIPVNVTLRRIAHGADAVAARWDFDLLNGQGGWKSDGCHILYSDENITTIQCYSLSNYAVLMDLT
GSELYTQAASLLHPVVYTTAIILLLCLLAVIVSYIYHHSLIRISLKSWHMLVNLCFHIFLTCVVFVGGITQTR
NASICQAVGIILHYSTLATVLWVGVTARNIYKQVTKKAKRCQDPDEPPPPPRPMLRFYLIGGGIPIIVCGI
TAAANIKNYGSRPNAPYCWMAWEPSLGAFYGPASFITFVNCMYFLSIFIQLKRHPERKYELKEPTEEQ
QRLAANENGEINHQDSMSLSLISTSALENEHTFHSQLLGASLTLLLYVALWMFGALAVSLYYPLDLVFS
FVFGATSLSFSAFFVVHHCVNREDVRLAWIMTCCPGRSSYSVQVNVQPPNSNGTNGEAPKCPNSSA
ESSCTNKSASSFKNSSQGCKLTNLQAAAAQCHANSLPLNSTPQLDNSLTEHSMDNDIKMHVAPLEVQ
FRTNVHSSRHHKNRSKGHRASRLTVLREYAYDVPTSVEGSVQNGLPKSRLGNNEGHSRSRRAYLAY
RERQYNPPQQDSSDACSTLPKSSRNFEKPVSTTSKKDALRKPAVVELENQQKSYGLNLAIQNGPIKS NGQEGPLLGTDSTGNVRTGLWKHETTV
>sp|Q96PE1 |AGRA2_HUMAN Adhesion G protein-coupled receptor A2 OS=Homo sapiens OX=9606
GN=ADGRA2 PE=1 SV=2
MGAGGRRMRGAPARLLLPLLPWLLLLLAPEARGAPGCPLSIRSCKCSGERPKGLSGGVPGPARRRV
VCSGGDLPEPPEPGLLPNGTVTLLLSNNKITGLRNGSFLGLSLLEKLDLRNNIISTVQPGAFLGLGELKR
LDLSNNRIGCLTSETFQGLPRLLRLNISGNIFSSLQPGVFDELPALKWDLGTEFLTCDCHLRWLLPWA
QNRSLQLSEHTLCAYPSALHAQALGSLQEAQLCCEGALELHTHHLIPSLRQVVFQGDRLPFQCSASYL
GNDTRIRWYHNRAPVEGDEQAGILLAESLIHDCTFITSELTLSHIGVWASGEWECTVSMAQGNASKKV
EIVVLETSASYCPAERVANNRGDFRWPRTLAGITAYQSCLQYPFTSVPLGGGAPGTRASRRCDRAGR
WEPGDYSHCLYTNDITRVLYTFVLMPINASNALTLAHQLRVYTAEAASFSDMMDVVYVAQMIQKFLGY
VDQIKELVEVMVDMASNLMLVDE/7LLWLAQREDKACSRIVGALERIGGAALSPHAQHISVNARNVALE
AYLIKPHSYVGLTCTAFQRREGGVPGTRPGSPGQNPPPEPEPPADQQLRFRCTTGRPNVSLSSFHIK NSVALASIQLPPSLFSSLPAALAPPVPPDCTLQLLVFR/VGRLFHSHSNTSRPGAAGPGKRRGVATPVIF
AGTSGCGVGNLTEPVAVSLRHWAEGAEPVAAWWSQEGPGEAGGWTSEGCQLRSSQPNVSALHCQ
HLGN\/A\/LMELSAFPREVGGAGAGLHPVVYPCTALLLLCLFATIITYILNHSSIRVSRKGWHMLLNLCF
HIAMTSAVFAGGITLTNYQMVCQAVGITLHYSSLSTLLWMGVKARVLHKELTWRAPPPQEGDPALPTP
SPMLRFYLIAGGIPLIICGITAAVNIHNYRDHSPYCWLVWRPSLGAFYIPVALILLITWIYFLCAGLRLRGP
LAQNPKAGNSRASLEAGEELRGSTRLRGSGPLLSDSGSLLATGSARVGTPGPPEDGDSLYSPGVQL
GALVTTHFLYLAMWACGALAVSQRWLPRVVCSCLYGVAASALGLFVFTHHCARRRDVRASWRACCP
PASPAAPHAPPRALPAAAEDGSPVFGEGPPSLKSSPSGSSGHPLALGPCKLTNLQLAQSQVCEAGAA
AGGEGEPEPAGTRGNLAHRHPNNVHHGRRAHKSRAKGHRAGEACGKNRLKALRGGAAGALELLSS
ESGSLHNSPTDSYLGSSRNSPGAGLQLEGEPMLTPSEGSDTSAAPLSEAGRAGQRRSASRDSLKGG GALEKESHRRSYPLNAASLNGAPKGGKYDDVTLMGAEVASGGCMKTGLWKSETTV
>sp|Q8WXG9|AGRV1_HUMAN Adhesion G-protein coupled receptor V1 OS=Homo sapiens OX=9606 GN=ADGRV1 PE=1 SV=2
MSVFLGPGMPSASLLVNLLSALLILFVFGETEIRFTGQTEFVVNETSTTVIRLIIERIGEPANVTAIVSLYG
EDAGDFFDTYAAAFIPAGETNRTVYIAVCDDDLPEPDETFIFHLTLQKPSANVKLGWPRTVTVTILSND
NAFGIISFNMLPSIAVSEPKGRNESMPLTLIREKGTYGMVMVTFEVEGGPNPPDEDLSPVKGNITFPPG RATVIYNLTVLDDEVPENDEIFLIQLKSVEGGAEINTSRNSIEIIIKKNDSPVRFLQSIYLVPEEDHILIIPVV
RGKDNNGNLIGSDEYEVSISYAVTTGNSTAHAQQNLDFIDLQPNTTVVFPPFIHESHLKFQIVDDTIPEI AESFHIMLLKDTLQGDAVLISPSVVQVTIKPNDKPYGVLSFNSVLFERTVIIDEDRISRYEEITWRNGGT HGNVSANWVLTRNSTDPSPVTADIRPSSGVLHFAQGQMLATIPLTVVDDDLPEEAEAYLLQILPHTIRG GAEVSEPAELLFYIQDSDDVYGLITFFPMENQKIESSPGERYLSLSFTRLGGTKGDVRLLYSVLYIPAGA VDPLQAKEGILNISRRNDLIFPEQKTQVTTKLPIRNDAFLQNGAHFLVQLETVELLNIIPLIPPISPRFGEIC NISLLVTPAIANGEIGFLSNLPIILHEPEDFAAEVVYIPLHRDGTDGQATVYWSLKPSGFNSKAVTPDDIG PFNGSVLFLSGQSDTTINITIKGDDIPEMNETVTLSLDRVNVENQVLKSGYTSRDLIILENDDPGGVFEF
SPASRGPYVIKEGESVELHIIRSRGSLVKQFLHYRVEPRDSNEFYGNTGVLEFKPGEREIVITLLARLDG IPELDEHYWVVLSSHGERESKLGSATIVNITILKNDDPHGIIEFVSDGLIVMINESKGDAIYSAVYDVVRN
RGNFGDVSVSWVVSPDFTQDVFPVQGTVVFGDQEFSKNITIYSLPDEIPEEMEEFTVILLNGTGGAKV GNRTTATLRIRRNDDPIYFAEPRVVRVQEGETANFTVLRNGSVDVTCMVQYATKDGKATARERDFIPV EKGETLIFEVGSRQQSISIFVNEDGIPETDEPFYIILLNSTGDTVVYQYGVATVIIEANDDPNGIFSLEPID KAVEEGKTNAFWILRHRGYFGSVSVSWQLFQNDSALQPGQEFYETSGTVNFMDGEEAKPIILHAFPD
KIPEFNEFYFLKLVNISGGSPGPGGQLAETNLQVTVMVPFNDDPFGVFILDPECLEREVAEDVLSEDD MSYITNFTILRQQGVFGDVQLGWEILSSEFPAGLPPMIDFLLVGIFPTTVHLQQHMRRHHSGTDALYFT GLEGAFGTVNPKYHPSRNNTIANFTFSAWVMPNANTNGFIIAKDDGNGSIYYGVKIQTNESHVTLSLHY KTLGSNATYIAKTTVMKYLEESVWLHLLIILEDGIIEFYLDGNAMPRGIKSLKGEAITDGPGILRIGAGING NDRFTGLMQDVRSYERKLTLEEIYELHAMPAKSDLHPISGYLEFRQGETNKSFIISARDDNDEEGEELF ILKLVSVYGGARISEENTTARLTIQKSDNANGLFGFTGACIPEIAEEGSTISCVVERTRGALDYVHVFYTI SQIETDGINYLVDDFANASGTITFLPWQRSEVLNIYVLDDDIPELNEYFRVTLVSAIPGDGKLGSTPTSG
ASIDPEKETTDITIKASDHPYGLLQFSTGLPPQPKDAMTLPASSVPHITVEEEDGEIRLLVIRAQGLLGRV TAEFRTVSLTAFSPEDYQNVAGTLEFQPGERYKYIFINITDNSIPELEKSFKVELLNLEGGVAELFRVDG SGSGDGDMEFFLPTIHKRASLGVASQILVTIAASDHAHGVFEFSPESLFVSGTEPEDGYSTVTLNVIRH HGTLSPVTLHWNIDSDPDGDLAFTSGNITFEIGQTSANITVEILPDEDPELDKAFSVSVLSVSSGSLGAH INATLTVLASDDPYGIFIFSEKNRPVKVEEATQNITLSIIRLKGLMGKVLVSYATLDDMEKPPYFPPNLAR ATQGRDYIPASGFALFGANQSEATIAISILDDDEPERSESVFIELLNSTLVAKVQSRSIPNSPRLGPKVET IAQLIIIANDDAFGTLQLSAPIVRVAENHVGPIINVTRTGGAFADVSVKFKAVPITAIAGEDYSIASSDVVLL
EGETSKAVPIYVINDIYPELEESFLVQLMNETTGGARLGALTEAVIIIEASDDPYGLFGFQITKLIVEEPEF NSVKVNLPIIRNSGTLGNVTVQWVATINGQLATGDLRVVSGNVTFAPGETIQTLLLEVLADDVPEIEEVI QVQLTDASGGGTIGLDRIANIIIPANDDPYGTVAFAQMVYRVQEPLERSSCANITVRRSGGHFGRLLLF
YSTSDIDVVALAMEEGQDLLSYYESPIQGVPDPLWRTWMNVSAVGEPLYTCATLCLKEQACSAFSFF SASEGPQCFWMTSWISPAVNNSDFWTYRKNMTRVASLFSGQAVAGSDYEPVTRQWAIMQEGDEFA NLTVSILPDDFPEMDESFLISLLEVHLMNISASLKNQPTIGQPNISTVVIALNGDAFGVFVIYNISPNTSED GLFVEVQEQPQTLVELMIHRTGGSLGQVAVEWRVVGGTATEGLDFIGAGEILTFAEGETKKTVILTILD DSEPEDDESIIVSLVYTEGGSRILPSSDTVRVNILANDNVAGIVSFQTASRSVIGHEGEILQFHVIRTFPG RGNVTVNWKIIGQNLELNFANFSGQLFFPEGSLNTTLFVHLLDDNIPEEKEVYQVILYDVRTQGVPPAG IALLDAQGYAAVLTVEASDEPHGVLNFALSSRFVLLQEANITIQLFINREFGSLGAINVTYTTVPGMLSLK
NQTVGNLAEPEVDFVPIIGFLILEEGETAAAINITILEDDVPELEEYFLVNLTYVGLTMAASTSFPPRLDSE
GLTAQVIIDANDGARGVIEWQQSRFEVNETHGSLTLVAQRSREPLGHVSLFVYAQNLEAQVGLDYIFT
PMILHFADGERYKNVNIMILDDDIPEGDEKFQLILTNPSPGLELGKNTIALIIVLANDDGPGVLSFNNSEH
FFLREPTALYVQESVAVLYIVREPAQGLFGTVTVQFIVTEVNSSNESKDLTPSKGYIVLEEGVRFKALQI
SAILDTEPEMDEYFVCTLFNPTGGARLGVHVQTLITVLQNQAPLGLFSISAVENRATSIDIEEANRTVYL
NVSRTNGIDLAVSVQWETVSETAFGMRGMDVVFSVFQSFLDESASGWCFFTLENLIYGIMLRKSSVT
VYRWQGIFIPVEDLNIENPKTCEAFNIGFSPYFVITHEERNEEKPSLNSVFTFTSGFKLFLVQTIIILESSQ
VRYFTSDSQDYLIIASQRDDSELTQVFRWNGGSFVLHQKLPVRGVLTVALFNKGGSVFLAISQANARL
NSLLFRWSGSGFINFQEVPVSGTTEVEALSSANDIYLIFAENVFLGDQNSIDIFIWEMGQSSFRYFQSV
DFAAVNRIHSFTPASGIAHILLIGQDMSALYCWNSERNQFSFVLEVPSAYDVASVTVKSLNSSKNLIALV
GAHSHIYELAYISSHSDFIPSSGELIFEPGEREATIAVNILDDTVPEKEESFKVQLKNPKGGAEIGINDSV
TITILSNDDAYGIVAFAQNSLYKQVEEMEQDSLVTLNVERLKGTYGRITIAWEADGSISDIFPTSGVILFT
EGQVLSTITLTILADNIPELSEVVIVTLTRITTEGVEDSYKGATIDQDRSKSVITTLPNDSPFGLVGWRAA
SVFIRVAEPKENTTTLQLQIARDKGLLGDIAIHLRAQPNFLLHVDNQATENEDYVLQETIIIMKENIKEAH
AEVSILPDDLPELEEGFIVTITEVNLVNSDFSTGQPSVRRPGMEIAEIMIEENDDPRGIFMFHVTRGAGE
VITAYEVPPPLNVLQVPVVRLAGSFGAVNVYWKASPDSAGLEDFKPSHGILEFADKQVTAMIEITIIDDA
EFELTETFNISLISVAGGGRLGDDWVTWIPQNDSPFGVFGFEEKTVMIDESLSSDDPDSYVTLTVVR
SPGGKGTVRLEWTIDEKAKHNLSPLNGTLHFDETESQKTIVLHTLQDTVLEEDRRFTIQLISIDEVEISPV
KGSASIIIRGDKRASGEVGIAPSSRHILIGEPSAKYNGTAIISLVRGPGILGEVTVFWRIFPPSVGEFAETS
GKLTMRDEQSAVIWIQALNDDIPEEKSFYEFQLTAVSEGGVLSESSSTANITWASDSPYGRFAFSHE
QLRVSEAQRVNITIIRSSGDFGHVRLWYKTMSGTAEAGLDFVPAAGELLFEAGEMRKSLHVEILDDDY
PEGPEEFSLTITKVELQGRGYDFTIQENGLQIDQPPEIGNISIVRIIIMKNDNAEGIIEFDPKYTAFEVEED
VGLIMIPVVRLHGTYGYVTADFISQSSSASPGGVDYILHGSTVTFQHGQNLSFINISIIDDNESEFEEPIEI
LLTGATGGAVLGRHLVSRIIIAKSDSPFGVIRFLNQSKISIANPNSTMILSLVLERTGGLLGEIQVNWETV
GPNSQEALLPQNRDIADPVSGLFYFGEGEGGVRTIILTIYPHEEIEVEETFIIKLHLVKGEAKLDSRAKDV
TLTIQEFGDPNGVVQFAPETLSKKTYSEPLALEGPLLITFFVRRVKGTFGEIMVYWELSSEFDITEDFLS
TSGFFTIADGESEASFDVHLLPDEVPEIEEDYVIQLVSVEGGAELDLEKSITWFSVYANDDPHGVFALY
SDRQSILIGQNLIRSIQINITRLAGTFGDVAVGLRISSDHKEQPIVTENAERQLWKDGATYKVDVVPIKN
QVFLSLGSNFTLQLVTVMLVGGRFYGMPTILQEAKSAVLPVSEKAANSQVGFESTAFQLMNITAGTSH
VMISRRGTYGALSVAWTTGYAPGLEIPEFIWGNMTPTLGSLSFSHGEQRKGVFLWTFPSPGWPEAF
VLHLSGVQSSAPGGAQLRSGFIVAEIEPMGVFQFSTSSRNIIVSEDTQMIRLHVQRLFGFHSDLIKVSY
QTTAGSAKPLEDFEPVQNGELFFQKFQTEVDFEITIINDQLSEIEEFFYINLTSVEIRGLQKFDVNWSPR
LNLDFSVAVITILDNDDLAGMDISFPETTVAVAVDTTLIPVETESTTYLSTSKTTTILQPTNWAIVTEATG
VSAIPEKLVTLHGTPAVSEKPDVATVTANVSIHGTFSLGPSIVYIEEEMKNGTFNTAEVLIRRTGGFTGN
VSITVKTFGERCAQMEPNALPFRGIYGISNLTWAVEEEDFEEQTLTLIFLDGERERKVSVQILDDDEPE
GQEFFYVFLTNPQGGAQIVEEKDDTGFAAFAMVIITGSDLHNGIIGFSEESQSGLELREGAVMRRLHLI
VTRQPNRAFEDVKVFWRVTLNKTWVLQKDGVNLVEELQSVSGTTTCTMGQTKCFISIELKPEKVPQV
EVYFFVELYEATAGAAINNSARFAQIKILESDESQSLVYFSVGSRLAVAHKKATLISLQVARDSGTGLM
MSVNFSTQELRSAETIGRTIISPAISGKDFVITEGTLVFEPGQRSTVLDVILTPETGSLNSFPKRFQIVLF
DPKGGARIDKVYGTANITLVSDADSQAIWGLADQLHQPVNDDILNRVLHTISMKVATENTDEQLSAMM
HLIEKITTEGKIQAFSVASRTLFYEILCSLINPKRKDTRGFSHFAEVTENFAFSLLTNVTCGSPGEKSKTIL
DSCPYLSILALHWYPQQINGHKFEGKEGDYIRIPERLLDVQDAEIMAGKSTCKLVQFTEYSSQQWFISG
NNLPTLKNKVLSLSVKGQSSQLLTNDNEVLYRIYAAEPRIIPQTSLCLLWNQAAASWLSDSQFCKWEE
TADYVECACSHMSVYAVYARTDNLSSYNEAFFTSGFICISGLCLAVLSHIFCARYSMFAAKLLTHMMAA
SLGTQILFLASAYASPQLAEESCSAMAAVTHYLYLCQFSWMLIQSVNFWYVLVMNDEHTERRYLLFFL
LSWGLPAFWILLIVILKGIYHQSMSQIYGLIHGDLCFIPNVYAALFTAALVPLTCLVVVFVVFIHAYQVKP
QWKAYDDVFRGRTNAAEIPLILYLFALISVTWLWGGLHMAYRHFWMLVLFVIFNSLQGLYVFMVYFILH
NQMCCPMKASYTVEMNGHPGPSTAFFTPGSGMPPAGGEISKSTQNLIGAMEEVPPDWERASFQQG
SQASPDLKPSPQNGATFPSSGGYGQGSLIADEESQEFDDLIFALKTGAGLSVSDNESGQGSQEGGTL
TDSQIVELRRIPIADTHL
Preferable amino acid positions for replacement by Cys-Cys sites in GAIN domains (bold) of indicated
AGPCRs corresponding to the variants b1 b2 (underlined) and h1 b1 (italic) are indicated.
>sp|Q9Y653|AGRG1_HUMAN Adhesion G-protein coupled receptor G1 OS=Homo sapiens OX=9606
GN=ADGRG1 PE=1 SV=2
MTPQSLLQTTLFLLSLLFLVQGAHGRGHREDFRFCSQRNQTHRSSLHYKPTPDLRISIENSEEALTVH
APFPAAHPASRSFPDPRGLYHFCLYWNRHAGRLHLLYGKRDFLLSDKASSLLCFQHQEESLAQGPPL
LATSVTSWWSPQNISLPSAASFTFSFHSPPHTAAHNASVDMCELKRDLQLLSQFLKHPQKASRRPS
AAPASQQLQSLESKLTSVRFMGDMVSFEEDRINATVWKLQPTAGLQDLHIHSRQEEEQSEIMEYSV
LLPRTLFQRTKGRSGEAEKRLLLVDFSSQALFQDKNSSQVLGEKVLGIVVQNTKVANLTEPWLTFQ
HQLQPKNVTLQCVFWVEDPTLSSPGHWSSAGCETVRRETQTSCFCNHLTYFAVLMVSSVEVDAVH
KHYLSLLSYVGCWSALACLVTIAAYLCSRVPLPCRRKPRDYTIKVHMNLLLAVFLLDTSFLLSEPVALT
GSEAGCRASAIFLHFSLLTCLSWMGLEGYNLYRLWEVFGTYVPGYLLKLSAMGWGFPIFLVTLVALV
DVDNYGPIILAVHRTPEGVIYPSMCWIRDSLVSYITNLGLFSLVFLFNMAMLATMVVQILRLRPHTQKW
SHVLTLLGLSLVLGLPWALIFFSFASGTFQLWLYLFSIITSFQGFLIFIWYWSMRLQARGGPSPLKSNS
DSARLPISSGSTSSSRI
>sp|Q9HBW9|AGRL4_HUMAN Adhesion G protein-coupled receptor L4 OS=Homo sapiens OX=9606
GN=ADGRL4 PE=1 SV=3
MKRLPLLVVFSTLLNCSYTQNCTKTPCLPNAKCEIRNGIEACYCNMGFSGNGVTICEDDNECGNLTQS
CGENANCTNTEGSYYCMCVPGFRSSSNQDRFITNDGTVCIENVNANCHLDNVCIAANINKTLTKIRSI
KEPVALLQEVYRNSVTDLSPTDIITYIEILAESSSLLGYKNNTISAKDTLSNSTLTEFVKTVNNFVQRDT
FVVWDKLSVNHRRTHLTKLMHTVEQATLRISQSFQKTTEFDTNSTDIALKVFFFDSYNMKHIHPHMN
MDGDYINIFPKRKAAYDSNGNVAVAFVYYKSIGPLLSSSDNFLLKPQNYDNSEEEERVISSVISVSMS
SNPPTLYELEKITFTLSHRKVTDRYRSLCAFWNYSPDTMNGSWSSEGCELTYSNETHTSCRCNHLT
HFAILMSSGPSIGIKDYNILTRITQLGIIISLICLAICIFTFWFFSEIQSTRTTIHKNLCCSLFLAELVFLVGINT
NTNKLFCSIIAGLLHYFFLAAFAWMCIEGIHLYLIVVGVIYNKGFLHKNFYIFGYLSPAVVVGFSAALGYR
YYGTTKVCWLSTENNFIWSFIGPACLIILVNLLAFGVIIYKVFRHTAGLKPEVSCFENIRSCARGALALLF
LLGTTWIFGVLHVVHASVVTAYLFTVSNAFQGMFIFLFLCVLSRKIQEEYYRLFKNVPCCFGCLR
>sp|Q9HAR2|AGRL3_HUMAN Adhesion G protein-coupled receptor L3 OS=Homo sapiens OX=9606
GN=ADGRL3 PE=1 SV=2
MWPSQLLIFMMLLAPIIHAFSRAPIPMAVVRRELSCESYPIELRCPGTDVIMIESANYGRTDDKICDSDP
AQMENIRCYLPDAYKIMSQRCNNRTQCAVVAGPDVFPDPCPGTYKYLEVQYECVPYKVEQKVFLCP
GLLKGVYQSEHLFESDHQSGAWCKDPLQASDKIYYMPWTPYRTDTLTEYSSKDDFIAGRPTTTYKLP
HRVDGTGFWYDGALFFNKERTRNIVKFDLRTRIKSGEAIIANANYHDTSPYRWGGKSDIDLAVDENGL
WVIYATEQNNGKIVISQLNPYTLRIEGTWDTAYDKRSASNAFMICGILYVVKSVYEDDDNEATGNKIDYI
YNTDQSKDSLVDVPFPNSYQYIAAVDYNPRDNLLYVWNNYHVVKYSLDFGPLDSRSGQAHHGQVSYI
SPPIHLDSELERPSVKDISTTGPLGMGSTTTSTTLRTTTLSPGRSTTPSVSGRRNRSTSTPSPAVEVLD
DMTTHLPSASSQIPALEESCEAVEAREIMWFKTRQGQIAKQPCPAGTIGVSTYLCLAPDGIWDPQGPD
LSNCSSPWVNHITQKLKSGETAANIARELAEQTRNHLNAGDITYSVRAMDQLVGLLDVQLRNLTPG
GKDSAARSLNKAMVETVNNLLQPQALNAWRDLTTSDQLRAATMLLHTVEESAFVLADNLLKTDIVR
ENTDWIKLEVARLSTEGNLEDLKFPENMGHGSTIQLSANTLKQNGRNGEIRVAFVLYNNLGPYLSTE
NASMKLGTEALSTNHSVIVNSPVITAAINKEFSNKVYLADPVVFTVKHIKQSEENFNPNCSFWSYSKR
TMTGYWSTQGCRLLTTNKTHTTCSCNHLTNFAVLMAHVEVKHSDAVHDLLLDVITWVGILLSLVCLLI
CIFTFCFFRGLQSDRNTIHKNLCISLFVAELLFLIGINRTDQPIACAVFAALLHFFFLAAFTWMFLEGVQL
YIMLVEVFESEHSRRKYFYLVGYGMPALIVAVSAAVDYRSYGTDKVCWLRLDTYFIWSFIGPATLIIMLN
VIFLGIALYKMFHHTAILKPESGCLDNIKSWVIGAIALLCLLGLTWAFGLMYINESTVIMAYLFTIFNSLQG
MFIFIFHCVLQKKVRKEYGKCLRTHCCSGKSTESSIGSGKTSGSRTPGRYSTGSQSRIRRMWNDTVR
KQSESSFITGDINSSASLNREGLLNNARDTSVMDTLPLNGNHGNSYSIASGEYLSNCVQIIDRGYNHNE
TALEKKILKELTSNYIPSYLNNHERSSEQNRNLMNKLVNNLGSGREDDAIVLDDATSFNHEESLGLELIH
EESDAPLLPPRVYSTENHQPHHYTRRRIPQDHSESFFPLLTNEHTEDLQSPHRDSLYTSMPTLAGVAA
TESVTTSTQTEPPPAKCGDAEDVYYKSMPNLGSRNHVHQLHTYYQLGRGSSDGFIVPPNKDGTPPE
GSSKGPAHLVTSL
>sp|O95490|AGRL2_HUMAN Adhesion G protein-coupled receptor L2 OS=Homo sapiens OX=9606
GN=ADGRL2 PE=1 SV=2
MVSSGCRMRSLWFIIVISFLPNTEGFSRAALPFGLVRRELSCEGYSIDLRCPGSDVIMIESANYGRTDD
KICDADPFQMENTDCYLPDAFKIMTQRCNNRTQCIVVTGSDVFPDPCPGTYKYLEVQYECVPYIFVCP
GTLKAIVDSPCIYEAEQKAGAWCKDPLQAADKIYFMPWTPYRTDTLIEYASLEDFQNSRQTTTYKLPN
RVDGTGFVVYDGAVFFNKERTRNIVKFDLRTRIKSGEAIINYANYHDTSPYRWGGKTDIDLAVDENGL
WVIYATEQNNGMIVISQLNPYTLRFEATWETVYDKRAASNAFMICGVLYVVRSVYQDNESETGKNSID
YIYNTRLNRGEYVDVPFPNQYQYIAAVDYNPRDNQLYVWNNNFILRYSLEFGPPDPAQVPTTAVTITS
SAELFKTIISTTSTTSQKGPMSTTVAGSQEGSKGTKPPPAVSTTKIPPITNIFPLPERFCEALDSKGIKWP
QTQRGMMVERPCPKGTRGTASYLCMISTGTWNPKGPDLSNCTSHWVNQLAQKIRSGENAASLANE
LAKHTKGPVFAGDVSSSVRLMEQLVDILDAQLQELKPSEKDSAGRSYNKLQKREKTCRAYLKAIVD
TVDNLLRPEALESWKHMNSSEQAHTATMLLDTLEEGAFVLADNLLEPTRVSMPTEWIVLEVAVLSTE
GQIQDFKFPLGIKGAGSSIQLSANTVKQNSRNGLAKLVFIIYRSLGQFLSTENATIKLGADFIGRNSTIA
VNSHVISVSINKESSRVYLTDPVLFTLPHIDPDNYFNANCSFWNYSERTMMGYWSTQGCKLVDTNKT
RTTCACSHLTNFAILMAHREIAYKDGVHELLLTVITWVGIVISLVCLAICIFTFCFFRGLQSDRNTIHKNL
CINLFIAEFIFLIGIDKTKYAIACPIFAGLLHFFFLAAFAWMCLEGVQLYLMLVEVFESEYSRKKYYYVAGY LFPATVVGVSAAIDYKSYGTEKACWLHVDNYFIWSFIGPVTFIILLNIIFLVITLCKMVKHSNTLKPDSSRL
ENIKSWVLGAFALLCLLGLTWSFGLLFINEETIVMAYLFTIFNAFQGVFIFIFHCALQKKVRKEYGKCFRH SYCCGGLPTESPHSSVKASTTRTSARYSSGTQSRIRRMWNDTVRKQSESSFISGDINSTSTLNQGMT
GNYLLTNPLLRPHGTNNPYNTLLAETWCNAPSAPVFNSPGHSLNNARDTSAMDTLPLNGNFNNSYS LHKGDYNDSVQVVDCGLSLNDTAFEKMIISELVHNNLRGSSKTHNLELTLPVKPVIGGSSSEDDAIVAD
ASSLMHSDNPGLELHHKELEAPLIPQRTHSLLYQPQKKVKSEGTDSYVSQLTAEAEDHLQSPNRDSLY TSMPNLRDSPYPESSPDMEEDLSPSRRSENEDIYYKSMPNLGAGHQLQMCYQISRGNSDGYIIPINKE GCIPEGDVREGQMQLVTSL
>sp|O94910|AGRL1_HUMAN Adhesion G protein-coupled receptor L1 OS=Homo sapiens OX=9606 GN=ADGRL1 PE=1 SV=1
MARLAAVLWNLCVTAVLVTSATQGLSRAGLPFGLMRRELACEGYPIELRCPGSDVIMVENANYGRTD
DKICDADPFQMENVQCYLPDAFKIMSQRCNNRTQCWVAGSDAFPDPCPGTYKYLEVQYDCVPYKV
EQKVFVCPGTLQKVLEPTSTHESEHQSGAWCKDPLQAGDRIYVMPWIPYRTDTLTEYASWEDYVAA RHTTTYRLPNRVDGTGFVVYDGAVFYNKERTRNIVKYDLRTRIKSGETVINTANYHDTSPYRWGGKTD
IDLAVDENGLWVIYATEGNNGRLVVSQLNPYTLRFEGTWETGYDKRSASNAFMVCGVLYVLRSVYVD DDSEAAGNRVDYAFNTNANREEPVSLTFPNPYQFISSVDYNPRDNQLYVWNNYFVVRYSLEFGPPDP
SAGPATSPPLSTTTTARPTPLTSTASPAATTPLRRAPLTTHPVGAINQLGPDLPPATAPVPSTRRPPAP NLHVSPELFCEPREVRRVQWPATQQGMLVERPCPKGTRGIASFQCLPALGLWNPRGPDLSNCTSP
WVNQVAQKIKSGENAANIASELARHTRGSIYAGDVSSSVKLMEQLLDILDAQLQALRPIERESAGKN YNKMHKRERTCKDYIKAVVETVDNLLRPEALESWKDMNATEQVHTATMLLDVLEEGAFLLADNVRE PARFLAAKEA/VVLEVTVLNTEGQVQELVFPQEEYPRKNSIQLSAKTIKQNSRNGVVKVVFILYNNLGL FLSTENATVKLAGEAGPGGPGGASLVVNSQVIAASINKESSRVFLMDPVIFTVAHLEDKNHFNANCS
FWNYSERSMLGYWSTQGCRLVESNKTHTTCACSHLTNFAVLMAHREIYQGRINELLLSVITWVGIVIS
LVCLAICISTFCFLRGLQTDRNTIHKNLCINLFLAELLFLVGIDKTQYEIACPIFAGLLHYFFLAAFSWLCLE GVHLYLLLVEVFESEYSRTKYYYLGGYCFPALVVGIAAAIDYRSYGTEKACWLRVDNYFIWSFIGPVSF
VIVVNLVFLMVTLHKMIRSSSVLKPDSSRLDNIKSWALGAIALLFLLGLTWAFGLLFINKESVVMAYLFTT
FNAFQGVFIFVFHCALQKKVHKEYSKCLRHSYCCIRSPPGGTHGSLKTSAMRSNTRYYTGTQSRIRR MWNDTVRKQTESSFMAGDINSTPTLNRGTMGNHLLTNPVLQPRGGTSPYNTLIAESVGFNPSSPPVF
NSPGSYREPKHPLGGREACGMDTLPLNGNFNNSYSLRSGDFPPGDGGPEPPRGRNLADAAAFEKMI ISELVHNNLRGSSSAAKGPPPPEPPVPPVPGGGGEEEAGGPGGADRAEIELLYKALEEPLLLPRAQSV LYQSDLDESESCTAEDGATSRPLSSPPGRDSLYASGANLRDSPSYPDSSPEGPSEALPPPPPAPPGP PEIYYTSRPPALVARNPLQGYYQVRRPSHEGYLAAPGLEGPGPDGDGQMQLVTSL
>sp|Q96K78|AGRG7_HUMAN Adhesion G-protein coupled receptor G7 OS=Homo sapiens OX=9606 GN=ADGRG7 PE=1 SV=2
MASCRAWNLRVLVAVVCGLLTGIILGLGIWRIVIRIQRGKSTSSSSTPTEFCRNGGTWENGRCICTEEW KGLRCTIANFCENSTYMGFTFARIPVGRYGPSLQTCGKDTPNAGNPMAVRLCSLSLYGEIELQKVTIG
NCNENLETLEKQVKDVTAPLNNISSEVQILTSDANKLTAENITSATRWGQIFNTSRNASPEAKKVAIV
TVSQLLDASEDAFQRVAATANDDALTTLIEQMETYSLSLGNQSVVEPNIAIQSANFSSENAVGPSNV
RFSVQKGASSSLVSSSTFIHTNVDGLNPDAQTELQVLLNMTKNYTKTCGFVVYQNDKLFQSKTFTA
KSDFSQKIISSKTDENEQDQSASVDMVFSPKYNQKEFQLYSYACVYWNLSAKDWDTYGCQKDKGT
DGFLRCRCNHTTNFAVLMTFKKDYQYPKSLDILSNVGCALSVTGLALTVIFQIVTRKVRKTSVTWVLVN
LCISMLIFNLLFVFGIENSNKNLQTSDGDINNIDFDNNDIPRTDTINIPNPMCTAIAALLHYFLLVTFTWNA
LSAAQLYYLLIRTMKPLPRHFILFISLIGWGVPAIVVAITVGVIYSQNGNNPQWELDYRQEKICWLAIPEP
NGVIKSPLLWSFIVPVTIILISNVVMFITISIKVLWKNNQNLTSTKKVSSMKKIVSTLSVAVVFGITWILAYL
MLVNDDSIRIVFSYIFCLFNTTQGLQIFILYTVRTKVFQSEASKVLMLLSSIGRRKSLPSVTRPRLRVKMY
NFLRSLPTLHERFRLLETSPSTEEITLSESDNAKESI
>sp|Q86SQ4|AGRG6_HUMAN Adhesion G-protein coupled receptor G6 OS=Homo sapiens OX=9606
GN=ADGRG6 PE=1 SV=3
MMFRSDRMWSCHWKWKPSPLLFLFALYIMCVPHSVWGCANCRWLSNPSGTFTSPCYPNDYPNSQ
ACMWTLRAPTGYIIQITFNDFDIEEAPNCIYDSLSLDNGESQTKFCGATAKGLSFNSSANEMHVSFSSD
FSIQKKGFNASYIRVAVSLRNQKVILPQTSDAYQVSVAKSISIPELSAFTLCFEATKVGHEDSDWTAFSY
SNASFTQLLSFGKAKSGYFLSISDSKCLLNNALPVKEKEDIFAESFEQLCLVWNNSLGSIGVNFKRNYE
TVPCDSTISKVIPGNGKLLLGSNQNEIVSLKGDIYNFRLWNFTMNAKILSNLSCNVKGNVVDWQNDFW
NIPNLALKAESNLSCGSYLIPLPAAELASCADLGTLCQATVNSPSTTPPTVTTNMPVTNRIDKQRNDGII
YRISVVIQNILRHPEVKVQSKVAEWLNSTFQNWNYTVYVVNISFHLSAGEDKIKVKRSLEDEPRLVLWA
LLVYNATNNTNLEGKIIQQKLLKNNESLDEGLRLHTVNVRQLGHCLAMEEPKGYYWPSIQPSEYVLPC
PDKPGFSASRICFYNATNPLVTYWGPVDISNCLKEANEVANQILNLTADGQNLTSANITNIVEQVKRIV
NKEENIDITLGSTL/WNIFSNILSSSDSDLLESSSEALKTIDELAFKIDLNSTSHVNITTRNLALSVSSLLP
GTNAISNFSIGLPSNNESYFQMDFESGQVDPLASVILPPNLLENLSPEDSVLVRRAQFTFFNKTGLFQ
DVGPQRKTLVSYVMACSIGNITIQNLKDPVQIKIKHTRTQEVHHPICAFWDLNKNKSFGGWNTSGCV
AHRDSDASETVCLCNHFTHFGVLMDLPRSASQLDARNTKVLTFISYIGCGISAIFSAATLLTYVAFEKLR
RDYPSKILMNLSTALLFLNLLFLLDGWITSFNVDGLCIAVAVLLHFFLLATFTWMGLEAIHMYIALVKVFN
TYIRRYILKFCIIGWGLPALVVSVVLASRNNNEVYGKESYGKEKGDEFCWIQDPVIFYVTCAGYFGVMF
FLNIAMFIWMVQICGRNGKRSNRTLREEVLRNLRSWSLTFLLGMTWGFAFFAWGPLNIPFMYLFSIF
NSLQGLFIFIFHCAMKENVQKQWRQHLCCGRFRLADNSDWSKTATNIIKKSSDNLGKSLSSSSIGSNS
TYLTSKSKSSSTTYFKRNSHTDNVSYEHSFNKSGSLRQCFHGQVLVKTGPC
>sp|Q8IZF4|AGRG5_HUMAN Adhesion G-protein coupled receptor G5 OS=Homo sapiens OX=9606
GN=ADGRG5 PE=2 SV=3
MDHCGALFLCLCLLTLQNATTETWEELLSYMENMQVSRGRSSVFSSRQLHQLEQMLLNTSFPGYN
LTLQTPTIQSLAFKLSCDFSGLSLTSATLKRVPQAGGQHARGQHAMQFPAELTRDACKTRPRELRLI
CIYFSNTHFFKDENNSSLLNNYVLGAQLSHGHVNNLRDPVNISFWHNQSLEGYTLTCVFWKEGARK
QPWGGWSPEGCRTEQPSHSQVLCRCNHLTYFAVLMQLSPALVPAELLAPLTYISLVGCSISIVASLIT
VLLHFHFRKQSDSLTRIHMNLHASVLLLNIAFLLSPAFAMSPVPGSACTALAAALHYALLSCLTWMAIEG
FNLYLLLGRVYNIYIRRYVFKLGVLGWGAPALLVLLSLSVKSSVYGPCTIPVFDSWENGTGFQNMSICW
VRSPVVHSVLVMGYGGLTSLFNLVVLAWALWTLRRLRERADAPSVRACHDTVTVLGLTVLLGTTWAL AFFSFGVFLLPQLFLFTILNSLYGFFLFLWFCSQRCRSEAEAKAQIEAFSSSQTTQ
>sp|Q8IZF6|AGRG4_HUMAN Adhesion G-protein coupled receptor G4 OS=Homo sapiens OX=9606 GN=ADGRG4 PE=2 SV=2
MKEHIIYQKLYGLILMSSFIFLSDTLSLKGKKLDFFGRGDTYVSLIDTIPELSRFTACIDLVFMDDNSRYW
MAFSYITNNALLGREDIDLGLAGDHQQLILYRLGKTFSIRHHLASFQWHTICLIWDGVKGKLELFLNKER
ILEVTDQPHNLTPHGTLFLGHFLKNESSEVKSMMRSFPGSLYYFQLWDHILENEEFMKCLDGNIVSWE
EDVWLVNKIIPTVDRTLRCFVPENMTIQEKSTTVSQQIDMTTPSQITGVKPQNTAHSSTLLSQSIPIFAT
DYTTISYSNTTSPPLETMTAQKILKTLVDETATFAVDVLSTSSAISLPTQSISIDNTTNSMKKTKSPSSES
TKTTKMVEAMATEIFQPPTPSNFLSTSRFTKNSVVSTTSAIKSQSAVTKTTSLFSTIESTSMSTTPCLKQ
KSTNTGALPISTAGQEFIESTAAGTVPWFTVEKTSPASTHVGTASSFPPEPVLISTAAPVDSVFPRNQT
AFPLATTDMKIAFTVHSLTLPTRLIETTPAPRTAETELTSTNFQDVSLPRVEDAMSTSMSKETSSKTFSF
LTSFSFTGTESVQTVIDAEATRTALTPEITLASTVAETMLSSTITGRVYTQNTPTADGHLLTLMSTRSAS
TSKAPESGPTSTTDEAAHLFSSNETIWTSRPDQALLASMNTTTILTFVPNENFTSAFHENTTYTEYLSA
TTNITPLKASPEGKGTTANDATTARYTTAVSKLTSPWFANFSIVSGTTSITNMPEFKLTTLLLKTIPMSTK
PANELPLTPRETVVPSVDIISTLACIQPNFSTEESASETTQTEINGAIVFGGTTTPVPKSATTQRLNATVT
RKEATSHYLMRKSTIAAVAEVSPFSTMLEVTDESAQRVTASVTVSSFPDIEKLSTPLDNKTATTEVRES
WLLTKLVKTTPRSSYNEMTEMFNFNHTYVAHWTSETSEGISAGSPTSGSTHIFGEPLGASTTRISETS
FSTTPTDRTATSLSDGILPPQPTAAHSSATPVPVTHMFSLPVNGSSVVAEETEVTMSEPSTLARAFST
SVLSDVSNLSSTTMTTALVPPLDQTASTTIVIVPTHGDLIRTTSEATVISVRKTSMAVPSLTETPFHSLRL
STPVTAKAETTLFSTSVDTVTPSTHTLVCSKPPPDNIPPASSTHVISTTSTPEATQPISQVEETSTYALS
FPYTFSGGGWASLATGTTETSWDETTPSHISANKLTTSVNSHISSSATYRVHTPVSIQLVTSTSVLSS
DKDQMTISLGKTPRTMEVTEMSPSKNSFISYSRGTPSLEMTDTGFPETTKISSHQTHSPSEIPLGTPSD
GNLASSPTSGSTQITPTLTSSNTVGVHIPEMSTSLGKTALPSQALTITTFLCPEKESTSALPAYTPRTVE
MIVNSTYVTHSVSYGQDTSFVDTTTSSSTRISNPMDINTTFSHLHSLRTQPEVTSVASFISESTQTFPES
LSLSTAGLYNDGFTVLSDRITTAFSVPNVPTMLPRESSMATSTPIYQMSSLPVNVTAFTSKKVSDTPPI
VITKSSKTMHPGCLKSPCTATSGPMSEMSSIPVNNSAFTPATVSSDTSTRVGLFSTLLSSVTPRTTMT
MQTSTLDVTPVIYAGATSKNKMVSSAFTTEMIEAPSRITPTTFLSPTEPTLPFVKTVPTTIMAGIVTPFVG
TTAFSPLSSKSTGAISSIPKTTFSPFLSATQQSSQADEATTLGILSGITNRSLSTVNSGTGVALTDTYSRI
TVPENMLSPTHADSLHTSFNIQVSPSLTSFKSASGPTKNVKTTTNCFSSNTRKMTSLLEKTSLTNYATS
LNTPVSYPPWTPSSATLPSLTSFVYSPHSTEAEISTPKTSPPPTSQMVEFPVLGTRMTSSNTQPLLMT
SWNIPTAEGSQFPISTTINVPTSNEMETETLHLVPGPLSTFTASQTGLVSKDVMAMSSIPMSGILPNHG
LSENPSLSTSLRAITSTLADVKHTFEKMTTSVTPGTTLPSILSGATSGSVISKSPILTWLLSSLPSGSPPA
TVSNAPHVMTSSTVEVSKSTFLTSDMISAHPFTNLTTLPSATMSTILTRTIPTPTLGGITTGFPTSLPMSI
NVTDDIVYISTHPEASSRTTITANPRTVSHPSSFSRKTMSPSTTDHTLSVGAMPLPSSTITSSWNRIPTA
SSPSTLIIPKPTLDSLLNIMTTTSTVPGASFPLISTGVTYPFTATVSSPISSFFETTWLDSTPSFLSTEAST
SPTATKSTVSFYNVEMSFSVFVEEPRIPITSVINEFTENSLNSIFQNSEFSLATLETQIKSRDISEEEMV
MDRAILEQREGQEMATISYVPYSCVCQVIIKASSSLASSELMRKIKSKIHGNFTHGNFTQDQLTLLVN CEHVAVKKLEPGNCKADETASKYKGTYKWLLTNPTETAQTRCIKNEDGNATRFCSISINTGKSQWE
KPKFKQCKLLQELPDKIVDLANITISDENAEDVAEHILNLINESPALGKEETKIIVSKISDISQCDEISMN LTHVMLQIINVVLEKQNNSASDLHEISNEILRIIERTGHKMEFSGQIANLTVAGLALAVLRGDHTFDGM AFSIHSYEEGTDPEIFLGNVPVGGILASIYLPKSLTERIPLSNLQTILFNFFGQTSLFKTKNVTKALTTYV
VSASISDDMFIQNLADPVVITLQHIGGNQNYGQVHCAFWDFENNNGLGGWNSSGCKVKETNVNYTI
CQCDHLTHFGVLMDLSRSTVDSVNEQILALITYTGCGISSIFLGVAVVTYIAFHKLRKDYPAKILINLCTA LLMLNLVFLINSWLSSFQKVGVCITAAVALHYFLLVSFTWMGLEAVHMYLALVKVFNIYIPNYILKFCLVG WGIPAIMVAITVSVKKDLYGTLSPTTPFCWIKDDSIFYISVVAYFCLIFLMNLSMFCTVLVQLNSVKSQIQ
KTRRKMILHDLKGTMSLTFLLGLTWGFAFFAWGPMRNFFLYLFAIFNTLQGFFIFVFHCVMKESVREQ
WQIHLCCGWLRLDNSSDGSSRCQIKVGYKQEGLKKIFEHKLLTPSLKSTATSSTFKSLGSAQGTPSEI SFPNDDFDKDPYCSSP
>sp|Q86Y34|AGRG3_HUMAN Adhesion G protein-coupled receptor G3 OS=Homo sapiens OX=9606 GN=ADGRG3 PE=1 SV=1 MATPRGLGALLLLLLLPTSGQEKPTEGPRNTCLGSNNMYDIFNLNDKALCFTKCRQSGSDSCNVENL
QRYWLNYEAHLMKEGLTQKVNTPFLKALVQNLSTA/TAEDFYFSLEPSQVPRQVMKDEDKPPDRVR LPKSLFRSLPGNRSVVRLAVTILDIGPGTLFKGPRLGLGDGSGVLNNRLVGLSVGQMHVTKLAEPLE IVFSHQRPPPNMTLTCVFWDVTKGTTGDWSSEGCSTEVRPEGTVCCCDHLTFFALLLRPTLDQSTV
HILTRISQAGCGVSMIFLAFTIILYAFLRLSRERFKSEDAPKIHVALGGSLFLLNLAFLVNVGSGSKGSDA ACWARGAVFHYFLLCAFTWMGLEAFHLYLLAVRVFNTYFGHYFLKLSLVGWGLPALMVIGTGSANSY GLYTIRDRENRTSLELCWFREGTTMYALYITVHGYFLITFLFGMVVLALVVWKIFTLSRATAVKERGKN
RKKVLTLLGLSSLVGVTWGLAIFTPLGLSTVYIFALFNSLQGVFICCWFTILYLPSQSTTVSSSTARLDQ AHSASQE
>sp|Q8IZP9|AGRG2_HUMAN Adhesion G-protein coupled receptor G2 OS=Homo sapiens OX=9606 GN=ADGRG2 PE=1 SV=2
MVFSVRQCGHVGRTEEVLLTFKIFLVIICLHVVLVTSLEEDTDNSSLSPPPAKLSVVSFAPSSNGTPEVE
TTSLNDVTLSLLPSNETEKTKITIVKTFNASGVKPQRNICNLSSICNDSAFFRGEIMFQYDKESTVPQNQ
HITNGTLTGVLSLSELKRSELNKTLQTLSETYFIMCATAEAQSTLNCTFTIKLNNTMNACAVIAALERVKI
RPMEHCCCSVRIPCPSSPEELEKLQCDLQDPIVCLADHPRGPPFSSSQSIPVVPRATVLSQVPKATSF
AEPPDYSPVTHNVPSPIGEIQPLSPQPSAPIASSPAIDMPPQSETISSPMPQTHVSGTPPPVKASFSSP TVSAPANVNTTSAPPVQTDIVNTSSISDLENQVLQMEKALSLGSLEPNLAGEM/NQVSRLLHSPPDML
APLAQRLLKWDDIGLQLNFSNTTISLTSPSLALAVIRVNASSFNTTTFVAQDPANLQVSLETQAPEN SIGTITLPSSLMNNLPAHDMELASRVQFNFFETPALFQDPSLENLSLISYVISSSVANLTVRNLTRNVT VTLKHINPSQDELTVRCVFWDLGRNGGRGGWSDNGCSVKDRRLNETICTCSHLTSFGVLLDLSRTS
VLPAQMMALTFITYIGCGLSSIFLSVTLVTYIAFEKIRRDYPSKILIQLCAALLLLNLVFLLDSWIALYKMQG LCISVAVFLHYFLLVSFTWMGLEAFHMYLALVKVFNTYIRKYILKFCIVGWGVPAVWTIILTISPDNYGL GSYGKFPNGSPDDFCWINNNAVFYITVVGYFCVIFLLNVSMFIVVLVQLCRIKKKKQLGAQRKTSIQDL
RSIAGLTFLLGITWGFAFFAWGPVNVTFMYLFAIFNTLQGFFIFIFYCVAKENVRKQWRRYLCCGKLRL AENSDWSKTATNGLKKQTVNQGVSSSSNSLQSSSNSTNSTTLLVNNDCSVHASGNGNASTERNGVS FSVQNGDVCLHDFTGKQHMFNEKEDSCNGKGRMALRRTSKRGSLHFIEQM
>sp|Q8IZF2|AGRF5_HUMAN Adhesion G protein-coupled receptor F5 OS=Homo sapiens OX=9606 GN=ADGRF5 PE=1 SV=3
MKSPRRTTLCLMFIVIYSSKAALNWNYESTIHPLSLHEHEPAGEEALRQKRAVATKSPTAEEYTVNIEIS
FENASFLDPIKAYLNSLSFPIHGNNTDQITDILSINVTTVCRPAGNEIWCSCETGYGWPRERCLHNLICQ
ERDVFLPGHHCSCLKELPPNGPFCLLQEDVTLNMRVRLNVGFQEDLMNTSSALYRSYKTDLETAFRK
GYGILPGFKGVTVTGFKSGSVVVTYEVKTTPPSLELIHKANEQWQSLNQTYKMDYNSFQAVTINESN
FFVTPEIIFEGDTVSLVCEKEVLSSNVSWRYEEQQLEIQNSSRFSIYTALFNNMTSVSKLTIHNITPGDA
GEYVCKLILDIFEYECKKKIDVMPIQILANEEMKVMCDNNPVSLNCCSQGNVNWSKVEWKQEGKINIP GTPETDIDSSCSRYTLKADGTQCPSGSSGTTVIYTCEFISAYGARGSANIKVTFISVANLTITPDPISVSE GQNFSIKCISDVSNYDEVYWNTSAGIKIYQRFYTTRRYLDGAESVLTVKTSTREWNGTYHCIFRYKNS
YSIATKDVIVHPLPLKLNIMVDPLEATVSCSGSHHIKCCIEEDGDYKVTFHTGSSSLPAAKEVNKKQVCY
KHNFNASSVSWCSKTVDVCCHFTNAANNSVWSPSMKLNLVPGENITCQDPVIGVGEPGKVIQKLCRF
SNVPSSPESPIGGTITYKCVGSQWEEKRNDCISAPINSLLQMAKALIKSPSQDEMLPTYLKDLSISIDK
AEHEISSSPGSLGAIINILDLLSTVPTQVNSEMMTHVLSTVNVILGKPVLNTWKVLQQQWTNQSSQLL
HSVERFSQALQSGDSPPLSFSQT1WQMSSMVIKSSHPETYQQRFVFPYFDLWGNWIDKSYLENLQ SDSSIVTMAFPTLQAILAQDIQENNFAESLVMTTTVSHNTTMPFRISMTFKNNSPSGGETKCVFWNFR LANNTGGWDSSGCYVEEGDGDNVTCICDHLTSFSILMSPDSPDPSSLLGILLDIISYVGVGFSILSLAA
CLVVEAVVWKSVTKNRTSYMRHTCIVNIAASLLVANTWFIVVAAIQDNRYILCKTACVAATFFIHFFYLS VFFWMLTLGLMLFYRLVFILHETSRSTQKAIAFCLGYGCPLAISVITLGATQPREVYTRKNVCWLNWED TKALLAFAIPALIIWVNITITIWITKILRPSIGDKPCKQEKSSLFQISKSIGVLTPLLGLTWGFGLTTVFPGT
NLVFHIIFAILNVFQGLFILLFGCLWDLKVQEALLNKFSLSRWSSQHSKSTSLGSSTPVFSMSSPISRRF NNLFGKTGTYNVSTPEATSSSLENSSSASSLLN
>sp|Q8IZF3|AGRF4_HUMAN Adhesion G protein-coupled receptor F4 OS=Homo sapiens OX=9606 GN=ADGRF4 PE=1 SV=3
MKMKSQATMICCLVFFLSTECSHYRSKIHLKAGDKLQSPEGKPKTGRIQEKCEGPCISSSNCSQPCAK
DFHGEIGFTCNQKKWQKSAETCTSLSVEKLFKDSTGASRLSVAAPSIPLHILDFRAPETIESVAQGIRK
NCPFDYACITDMVKSSETTSGNIAFIVELLKNISTDLSDNVTREKMKSYSEVANHILDTAAISNWAFIP
NKNASSDLLQSVNLFARQLHIHNNSENIVNELFIQTKGFHINHNTSEKSLNFSMSMNNTTEDILGMVQI
PRQELRKLWPNASQAISIAFPTLGAILREAHLQNVSLPRQVNGLVLSVVLPERLQEIILTFEKINKTRN
ARAQCVGWHSKKRRWDEKACQMMLDIRNEVKCRCNYTSVVMSFSILMSSKSMTDKVLDYITCIGLS
VSILSLVLCLIIEATVWSRVVVTEISYMRHVCIVNIAVSLLTANVWFIIGSHFNIKAQDYNMCVAVTFFSHF
FYLSLFFWMLFKALLIIYGILVIFRRMMKSRMMVIGFAIGYGCPLIIAVTTVAITEPEKGYMRPEACWLNW
DNTKALLAFAIPAFVIVAVNLIVVLVVAVNTQRPSIGSSKSQDVVIIMRISKNVAILTPLLGLTWGFGIATLI
EGTSLTFHIIFALLNAFQGFFILLFGTIMDHKIRDALRMRMSSLKGKSRAAENASLGPTNGSKLMNRQG
>sp|Q8IZF5|AGRF3_HUMAN Adhesion G-protein coupled receptor F3 OS=Homo sapiens OX=9606 GN=ADGRF3 PE=2 SV=1
MVCSAAPLLLLATTLPLLGSPVAQASQPVSETGVRPREGLQRRQWGPLIGRDKAWNERIDRPFPACPI
PLSSSFGRWPKGQTMWAQTSTLTLTEEELGQSQAGGESGSGQLLDQENGAGESALVSVYVHLDFP
DKTWPPELSRTLTLPAASASSSPRPLLTGLRLTTECNVNHKGNFYCACLSGYQWNTSICLHYPPCQSL
HNHQPCGCLVFSHPEPGYCQLLPPGSPVTCLPAVPGILNLNSQLQMPGDTLSLTLHLSQEATNLSWF
LRHPGSPSPILLQPGTQVSVTSSHGQAALSVSNMSHHWAGEYMSCFEAQGFKWNLYEVVRVPLKAT
DVARLPYQLSISCATSPGFQLSCCIPSTNLAYTAAWSPGEGSKASSFNESGSQCFVLAVQRCPMADT
TYACDLQSLGLAPLRVPISITIIQDGDITCPEDASVLTWNVTKAGHVAQAPCPESKRGIVRRLCGADGV
WGPVHSSCTDARLLALFTRTKLLQAGQGSPAEEVPQILAQLPGQAAEASSPSDLLTLLSTMKYVAK
VVAEARIQLDRRALKNLLIATDKVLDMDTRSLWTLAQARKPWAGSTLLLAVETLACSLCPQDHPFA
FSLPNVLLQSQLFGPTFPADYSISFPTRPPLQAQIPRHSLAPLVRNGTEISITSLVLRKLDHLLPSNYG
QGLGDSLYATPGLVLVISIMAGDRAFSQGEVIMDFGNTDGSPHCVFWDHSLFQGRGGWSKEGCQA
QVASASPTAQCLCQHLTAFSVLMSPHTVPEEPALALLTQVGLGASILALLVCLGVYWLVWRVWRNKI
SYFRHAALLNMVFCLLAADTCFLGAPFLSPGPRSPLCLAAAFLCHFLYLATFFWMLAQALVLAHQLLFV
FHQLAKHRVLPLMVLLGYLCPLGLAGVTLGLYLPQGQYLREGECWLDGKGGALYTFVGPVLAIIGVNG
LVLAMAMLKLLRPSLSEGPPAEKRQALLGVIKALLILTPIFGLTWGLGLATLLEEVSTVPHYIFTILNTLQG
VFILLFGCLMDRKIQEALRKRFCRAQAPSSTISLVSCCLQILSCASKSMSEGIPWPSSEDMGTARS
>sp|Q8IZF7|AGRF2_HUMAN Adhesion G-protein coupled receptor F2 OS=Homo sapiens OX=9606 GN=ADGRF2 PE=2 SV=1
MGLTAYGNRRVQPGELPFGANLTLIHTRAQPVICSKLLLTKRVSPISFFLSKFQNSWGEDGWVQLDQL
PSPNAVSSDQVHCSAGCTHRKCGWAASKSKEKVPARPHGVCDGVCTDYSQCTQPCPPDTQGNMG
FSCRQKTWHKITDTCQTLNALNIFEEDSRLVQPFEDNIKISVYTGKSETITDMLLQKCPTDLSCVIRNIQ
QSPWIPGNIAVIVQLLHNISTAIWTGVDEAKMQSYSTIANHILNSKSISNWTFIPDRNSSYILLHSVNSF
ARRLFIDKHPVDISDVFIHTMGTTISGDNIGKNFTFSMRINDTSNEVTGRVLISRDELRKVPSPSQVISIA
FPTIGAILEASLLENVTVNGLVLSAILPKELKRISLIFEKISKSEERRTQCVGWHSVENRWDQQACKMI
QENSQQAVCKCRPSKLFTSFSILMSPHILESLILTYITYVGLGISICSLILCLSIEVLVWSQVTKTEITYLR
HVCIVNIAATLLMADVWFIVASFLSGPITHHKGCVAATFFVHFFYLSVFFWMLAKALLILYGIMIVFHTLP
KSVLVASLFSVGYGCPLAIAAITVAATEPGKGYLRPEICWLNWDMTKALLAFVIPALAIWVNLITVTLVIV
KTQRAAIGNSMFQEVRAIVRISKNIAILTPLLGLTWGFGVATVIDDRSLAFHIIFSLLNAFQVSPDASDQV QSERIHEDVL
>sp|Q5T601 |AGRF1 _HUMAN Adhesion G-protein coupled receptor F1 OS=Homo sapiens OX=9606
GN=ADGRF1 PE=1 SV=2
MKVGVLWLISFFTFTDGHGGFLGKNDGIKTKKELIVNKKKHLGPVEEYQLLLQVTYRDSKEKRDLRNFL
KLLKPPLLWSHGLIRIIRAKATTDCNSLNGVLQCTCEDSYTWFPPSCLDPQNCYLHTAGALPSCECHL
NNLSQSVNFCERTKIWGTFKINERFTNDLLNSSSAIYSKYANGIEIQLKKAYERIQGFESVQVTQFRNG
SIVAGYEWGSSSASELLSAIEHVAEKAKTALHKLFPLEDGSFRVFGKAQCNDIVFGFGSKDDEYTLPC
SSGYRGNITAKCESSGWQVIRETCVLSLLEELNKNFSMIVGNATEAAVSSFVQNLSVIIRQNPSTTVG
NLASVVSILSNISSLSLASHFRVSNSTMEDVISIADNILNSASVTNWTVLLREEKYASSRLLETLENIST
LVPPTALPLNFSRKFIDWKGIPVNKSQLKRGYSYQIKMCPQNTSIPIRGRVLIGSDQFQRSLPETIISM
ASLTLGNILPVSKNGNAQVNGPVISTVIQNYSINEVFLFFSKIESNLSQPHCVFWDFSHLQWNDAGCH
LVNETQDIVTCQCTHLTSFSILMSPFVPSTIFPVVKWITYVGLGISIGSLILCLIIEALFWKQIKKSQTSHTR
RICMVNIALSLLIADVWFIVGATVDTTVNPSGVCTAAVFFTHFFYLSLFFWMLMLGILLAYRIILVFHHMA QHLMMAVGFCLGYGCPLIISVITIAVTQPSNTYKRKDVCWLNWSNGSKPLLAFVVPALAIVAVNFWVL LVLTKLWRPTVGERLSRDDKATIIRVGKSLLILTPLLGLTWGFGIGTIVDSQNLAWHVIFALLNAFQGFFI
LCFGILLDSKLRQLLFNKLSALSSWKQTEKQNSSDLSAKPKFSKPFNPLQNKGHYAFSHTGDSSDNIM LTQFVSNE
>sp|P48960|AGRE5_HUMAN Adhesion G protein-coupled receptor E5 OS=Homo sapiens OX=9606 GN=ADGRE5 PE=1 SV=4
MGGRVFLAFCVWLTLPGAETQDSRGCARWCPQNSSCVNATACRCNPGFSSFSEIITTPTETCDDINE
CATPSKVSCGKFSDCWNTEGSYDCVCSPGYEPVSGAKTFKNESENTCQDVDECQQNPRLCKSYGT
CVNTLGSYTCQCLPGFKFIPEDPKVCTDVNECTSGQNPCHSSTHCLNNVGSYQCRCRPGWQPIPGS PNGPNNTVCEDVDECSSGQHQCDSSTVCFNTVGSYSCRCRPGWKPRHGIPNNQKDTVCEDMTFST
WTPPPGVHSQTLSRFFDKVQDLGRDSKTSSAEVTIQNVIKLVDELMEAPGDVEALAPPVRHLIATQL LSNLEDIMRILAKSLPKGPFTYISPSNTELTLMIQERGDKNVTMGQSSARMKLNWAVAAGAEDPGPA VAGILSIQNMTTLLANASLNLHSKKQAELEEIYESSIRGVQLRRLSAVNSIFLSHNNTKELNSPILFAFS
HLESSDGEAGRDPPAKDVMPGPRQELLCAFWKSDSDRGGHWATEGCQVLGSKNGSTTCQCSHL
SSFAILMAHYDVEDWKLTLITRVGLALSLFCLLLCILTFLLVRPIQGSRTTIHLHLCICLFVGSTIFLAGIEN EGGQVGLRCRLVAGLLHYCFLAAFCWMSLEGLELYFLVVRVFQGQGLSTRWLCLIGYGVPLLIVGVS AAIYSKGYGRPRYCWLDFEQGFLWSFLGPVTFIILCNAVIFVTTVWKLTQKFSEINPDMKKLKKARALTI
TAIAQLFLLGCTWVFGLFIFDDRSLVLTYVFTILNCLQGAFLYLLHCLLNKKVREEYRKWACLVAGGSK YSEFTSTTSGTG H N QTRALRASESG I
>sp|Q86SQ3|AGRE4_HUMAN Putative adhesion G protein-coupled receptor E4P OS=Homo sapiens OX=9606 GN=ADGRE4P PE=5 SV=1
MGSRFLLVLLSGASCPPCPKYASCHNSTHCTCEDGFRARSGRTYFHDSSEKCEDINECETGLAKCK
YKAYCRNKVGGYICSCLVKYTLFNFLAGIIDYDHPDCYEWNSQGTTQSNVDIWVSGVKPGFGKQLPG
DKRTKHICVYWEGSEGGWSTEGCSHVHSNGSYTKCKCFHLSSFAVLVALAPKEDPVLTVITQVGLTI
SLLCLFLAILTFLLCRPIQNTSTSLHLELSLCLFLAHLLFLTGINRTEPEVLCSIIAGLLHFLYLACFTWMLL EGLHLFLTVRNLKVANYTSTGRFKKRFMYPVGYGIPAVIIAVSAIVGPQNYGTFTCWLKLDKGFIWSFM GPVAVIILINLVFYFQVLWILRSKLSSLNKEVSTIQDTRVMTFKAISQLFILGCSWGLGFFMVEEVGKTIG
SIIAYSFTIINTLQGVLLFVVHCLLNRQVRLIILSVISLVPKSN
>sp|Q9BY15|AGRE3_HUMAN Adhesion G protein-coupled receptor E3 OS=Homo sapiens OX=9606 GN=ADGRE3 PE=2 SV=2
MQGPLLLPGLCFLLSLFGAVTQKTKTSCAKCPPNASCVNNTHCTCNHGYTSGSGQKLFTFPLETCNDI
NECTPPYSVYCGFNAVCYNVEGSFYCQCVPGYRLHSGNEQFSNSNENTCQDTTSSKTTEGRKELQ
KIVDKFESLLTNQTLWRTEGRQEISSTATTILRDVESKVLETALKDPEQKVLKIQNDSVAIETQAITDN CSEERKTFNLNVQMNSMDIRCSDIIQGDTQGPSAIAFISYSSLGNIINATFFEEMDKKDQVYLNSQVVS AAIGPKRNVSLSKSVTLTFQHVKMTPSTKKVFCVYWKSTGQGSQWSRDGCFLIHVNKSHTMCNCS
HLSSFAVLMALTSQEEDPVLTVITYVGLSVSLLCLLLAALTFLLCKAIRNTSTSLHLQLSLCLFLAHLLFL
VGIDRTEPKVLCSIIAGALHYLYLAAFTWMLLEGVHLFLTARNLTVVNYSSINRLMKWIMFPVGYGVPA
VTVAISAASWPHLYGTADRCWLHLDQGFMWSFLGPVCAIFSANLVLFILVFWILKRKLSSLNSEVSTIQ NTRMLAFKATAQLFILGCTWCLGLLQVGPAAQVMAYLFTIINSLQGFFIFLVYCLLSQQVQKQYQKWFR
EIVKSKSESETYTLSSKMGPDSKPSEGDVFPGQVKRKY
>sp|Q9UHX3|AGRE2_HUMAN Adhesion G protein-coupled receptor E2 OS=Homo sapiens OX=9606
GN=ADGRE2 PE=1 SV=2
MGGRVFLVFLAFCVWLTLPGAETQDSRGCARWCPQDSSCVNATACRCNPGFSSFSEIITTPMETCD
DINECATLSKVSCGKFSDCWNTEGSYDCVCSPGYEPVSGAKTFKNESENTCQDVDECQQNPRLCKS
YGTCVNTLGSYTCQCLPGFKLKPEDPKLCTDVNECTSGQNPCHSSTHCLNNVGSYQCRCRPGWQPI
PGSPNGPNNTVCEDVDECSSGQHQCDSSTVCFNTVGSYSCRCRPGWKPRHGIPNNQKDTVCEDMT
FSTWTPPPGVHSQTLSRFFDKVQDLGRDYKPGLANNTIQSILQALDELLEAPGDLETLPRLQQHCVA
SHLLDGLEDVLRGLSKNLSNGLLNFSYPAGTELSLEVQKQVDRSVTLRQNQAVMQLDWNQAQKSG
DPGPSWGLVSIPGMGKLLAEAPLVLEPEKQMLLHETHQGLLQDGSPILLSDVISAFLSNNDTQNLS
SPVTFTFSHRSVIPRQKVLCVFWEHGQNGCGHWATTGCSTIGTRDTSTICRCTHLSSFAVLMAHYDV
QEEDPVLTVITYMGLSVSLLCLLLAALTFLLCKAIQNTSTSLHLQLSLCLFLAHLLFLVAIDQTGHKVLCSI
IAGTLHYLYLATLTWMLLEALYLFLTARNLTVVNYSSINRFMKKLMFPVGYGVPAVTVAISAASRPHLYG
TPSRCWLQPEKGFIWGFLGPVCAIFSVNLVLFLVTLWILKNRLSSLNSEVSTLRNTRMLAFKATAQLFIL
GCTWCLGILQVGPAARVMAYLFTIINSLQGVFIFLVYCLLSQQVREQYGKWSKGIRKLKTESEMHTLSS SAKADTSKPSTVN
>sp|Q14246|AGRE1_HUMAN Adhesion G protein-coupled receptor E1 OS=Homo sapiens OX=9606 GN=ADGRE1 PE=2 SV=3
MRGFNLLLFWGCCVMHSWEGHIRPTRKPNTKGNNCRDSTLCPAYATCTNTVDSYYCACKQGFLSSN
GQNHFKDPGVRCKDIDECSQSPQPCGPNSSCKNLSGRYKCSCLDGFSSPTGNDWVPGKPGNFSCT
DINECLTSSVCPEHSDCVNSMGSYSCSCQVGFISRNSTCEDVDECADPRACPEHATCNNTVGNYSC
FCNPGFESSSGHLSFQGLKASCEDIDECTEMCPINSTCTNTPGSYFCTCHPGFAPSNGQLNFTDQGV
ECRDIDECRQDPSTCGPNSICTNALGSYSCGCIAGFHPNPEGSQKDGNFSCQRVLFKCKEDVIPDNK
QIQQCQEGTAVKPAYVSFCAQINNIFSVLDKVCENKTTWSLKNTTESFVPVLKQISTWTKFTKEETS
SLATVFLESVESMTLASFWKPSANITPAVRTEYLDIESKVINKECSEENVTLDLVAKGDKMKIGCSTIE
ESESTETTGVAFVSFVGMESVLNERFFKDHQAPLTTSEIKLKMNSRVVGGIMTGEKKDGFSDPIIYTL
ENIQPKQKFERPICVSWSTDVKGGRWTSFGCVILEASETYTICSCNQMANLAVIMASGELTMDFSLYI
ISHVGIIISLVCLVLAIATFLLCRSIRNHNTYLHLHLCVCLLLAKTLFLAGIHKTDNKMGCAIIAGFLHYLFLA
CFFWMLVEAVILFLMVRNLKVVNYFSSRNIKMLHICAFGYGLPMLVVVISASVQPQGYGMHNRCWLNT
ETGFIWSFLGPVCTVIVINSLLLTWTLWILRQRLSSVNAEVSTLKDTRLLTFKAFAQLFILGCSWVLGIFQ IGPVAGVMAYLFTIINSLQGAFIFLIHCLLNGQVREEYKRWITGKTKPSSQSQTSRILLSSMPSASKTG
>sp|Q7Z7M1 |AGRD2_HUMAN Adhesion G-protein coupled receptor D2 OS=Homo sapiens OX=9606
GN=ADGRD2 PE=2 SV=1
MDAPWGAGERWLHGAAVDRSGVSLGPPPTPQVNQGTLGPQVAPVAAGEVVKTAGGVCKFSGQRL
SWWQAQESCEQQFGHLALQPPDGVLASRLRDPVWVGQREAPLRRPPQRRARTTAVLVFDERTADR
AARLRSPLPELAALTACTHVQWDCASPDPAALFSVAAPALPNALQLRAFAEPGGVVRAALWRGQHA
PFLAAFRADGRWHHVCATWEQRGGRWALFSDGRRRAGARGLGAGHPVPSGGILVLGQDQDSLGG
GFSVRHALSGNLTDFHLWARALSPAQLHRARACAPPSEGLLFRWDPGALDVTPSLLPTVWVRLLCPV
PSEECPTWNPGPRSEGSELCLEPQPFLCCYRTEPYRRLQDAQSWPGQDVISRVNALANDIVLLPDP
LSEVHGALSPAEASSFLGLLEHVLAMEMAPLGPAALLAVVRFLKRVVALGAGDPELLLTGPWEQLS
QGWSVASLVLEEQVADTWLSLREVIGGPMALVASVQRLAPLLSTSMTSERPRMRIQHRHAGLSGV
TVIHSWFTSRVFQHTLEGPDLEPQAPASSEEANRVQRFLSTQVGSAIISSEVWDVTGEVNVAMTFHL
QHRAQSPLFPPHPPSPYTGGAWATTGCSVAALYLDSTACFCNHSTSFAILLQIYEVQRGPEEESLLR
TLSFVGCGVSFCALTTTFLLFLVAGVPKSERTTVHKNLTFSLASAEGFLMTSEWAKANEVACVAVTVA
MHFLFLVAFSWMLVEGLLLWRKVVAVSMHPGPGMRLYHATGWGVPVGIVAVTLAMLPHDYVAPGHC
WLNVHTNAIWAFVGPVLFVLTANTCILARVVMITVSSARRRARMLSPQPCLQQQIWTQIWATVKPVLV
LLPVLGLTWLAGILVHLSPAWAYAAVGLNSIQGLYIFLVYAACNEEVRSALQRMAEKKVAEVLRALGV WGGAAKEHSLPFSVLPLFLPPKPSTPRHPLKAPA
>sp|Q6QNK2|AGRD1_HUMAN Adhesion G-protein coupled receptor D1 OS=Homo sapiens OX=9606
GN=ADGRD1 PE=1 SV=1
MEKLLRLCCWYSWLLLFYYNFQVRGVYSRSQDHPGFQVLASASHYWPLENVDGIHELQDTTGDIVE
GKVNKGIYLKEEKGVTLLYYGRYNSSCISKPEQCGPEGVTFSFFWKTQGEQSRPIPSAYGGQVISNGF
KVCSSGGRGSVELYTRDNSMTWEASFSPPGPYWTHVLFTWKSKEGLKVYVNGTLSTSDPSGKVSR
DYGESNVNLVIGSEQDQAKCYENGAFDEFIIWERALTPDEIAMYFTAAIGKHALLSSTLPSLFMTSTAS
PVMPTDAYHPIITNLTEERKTFQSPGVILSYLQNVSLSLPSKSLSEQTALNLTKTFLKAVGEILLLPGW
IALSEDSAVVLSLIDTIDTVMGHVSSNLHGSTPQVTVEGSSAMAEFSVAKILPKTVNSSHYRFPAHGQ
SFIQIPHEAFHRHAWSTWGLLYHSMHYYLNNIWPAHTKIAEAMHHQDCLLFATSHLISLEVSPPPTL
SQNLSGSPLITVHLKHRLTRKQHSEATNSSNRVFVYCAFLDFSSGEGVWSNHGCALTRGNLTYSVC
RCTHLTNFAILMQVVPLELARGHQVALSSISYVGCSLSVLCLVATLVTFAVLSSVSTIRNQRYHIHANLS
FAVLVAQVLLLISFRLEPGTTPCQVMAVLLHYFFLSAFAWMLVEGLHLYSMVIKVFGSEDSKHRYYYG
MGWGFPLLICIISLSFAMDSYGTSNNCWLSLASGAIWAFVAPALFVIVVNIGILIAVTRVISQISADNYKIH
GDPSAFKLTAKAVAVLLPILGTSWVFGVLAVNGCAWFQYMFATLNSLQGLFIFLFHCLLNSEVRAAFK HKTKVWSLTSSSARTSNAKPFHSDLMNGTRPGMASTKLSPWDKSSHSAHRVDLSAV
>sp|O60242|AGRB3_HUMAN Adhesion G protein-coupled receptor B3 OS=Homo sapiens OX=9606 GN=ADGRB3 PE=1 SV=2
MKAVRNLLIYIFSTYLLVMFGFNAAQDFWCSTLVKGVIYGSYSVSEMFPKNFTNCTWTLENPDPTKYSI
YLKFSKKDLSCSNFSLLAYQFDHFSHEKIKDLLRKNHSIMQLCNSKNAFVFLQYDKNFIQIRRVFPTNFP
GLQKKGEEDQKSFFEFLVLNKVSPSQFGCHVLCTWLESCLKSENGRTESCGIMYTKCTCPQHLGEW
GIDDQSLILLNNVVLPLNEQTEGCLTQELQTTQVCNLTREAKRPPKEEFGMMGDHTIKSQRPRSVHEK
RVPQEQADAAKFMAQTGESGVEEWSQWSTCSVTCGQGSQVRTRTCVSPYGTHCSGPLRESRVCN NTALCPVHGVWEEWSPWSLCSFTCGRGQRTRTRSCTPPQYGGRPCEGPETHHKPCNIALCPVDGQ
WQEWSSWSQCSVTCSNGTQQRSRQCTAAAHGGSECRGPWAESRECYNPECTANGQWNQWGH
WSGCSKSCDGGWERRIRTCQGAVITGQQCEGTGEEVRRCNEQRCPAPYEICPEDYLMSMVWKRTP
AGDLAFNQCPLNATGTTSRRCSLSLHGVAFWEQPSFARCISNEYRHLQHSIKEHLAKGQRMLAGDG
MSQVTKTLLDLTQRKNFYAGDLLMSVEILRNVTDTFKRASYIPASDGVQNFFQIVSNLLDEENKEKW
EDAQQIYPGSIELMQVIEDFIHIVGMGMMDFQNSYLMTGA/VVASIQKLPAASVLTDINFPMKGRKGM
VDWARNSEDRVVIPKSIFTPVSSKELDESSVFVLGAVLYKNLDLILPTLRNYTVINSKIIVVTIRPEPKTT
DSFLEIELAHLANGTLNPYCVLWDDSKTNESLGTWSTQGCKTVLTDASHTKCLCDRLSTFAILAQQP
REIIMESSGTPSVTLIVGSGLSCLALITLAVVYAALWRYIRSERSIILINFCLSIISSNILILVGQTQTHNKSIC
TTTTAFLHFFFLASFCWVLTEAWQSYMAVTGKIRTRLIRKRFLCLGWGLPALVVATSVGFTRTKGYGT
DHYCWLSLEGGLLYAFVGPAAAWLVNMVIGILVFNKLVSRDGILDKKLKHRAGQMSEPHSGLTLKCA
KCGVVSTTALSATTASNAMASLWSSCWLPLLALTWMSAVLAMTDKRSILFQILFAVFDSLQGFVIVMV
HCILRREVQDAFRCRLRNCQDPINADSSSSFPNGHAQIMTDFEKDVDIACRSVLHKDIGPCRAATITGT
LSRISLNDDEEEKGTNPEGLSYSTLPGNVISKVIIQQPTGLHMPMSMNELSNPCLKKENSELRRTVYLC
TDDNLRGADMDIVHPQERMMESDYIVMPRSSVNNQPSMKEESKMNIGMETLPHERLLHYKVNPEFN MNPPVMDQFNMNLEQHLAPQEHMQNLPFEPRTAVKNFMASELDDNAGLSRSETGSTISMSSLERRK SRYSDLDFEKVMHTRKRHMELFQELNQKFQTLDRFRDIPNTSSMENPAPNKNPWDTFKNPSEYPHY
TTINVLDTEAKDALELRPAEWEKCLNLPLDVQEGDFQTEV
>sp|O60241 |AGRB2_HUMAN Adhesion G protein-coupled receptor B2 OS=Homo sapiens OX=9606 GN=ADGRB2 PE=1 SV=2 MENTGWMGKGHRMTPACPLLLSVILSLRLATAFDPAPSACSALASGVLYGAFSLQDLFPTIASGCSWT
LENPDPTKYSLYLRFNRQEQVCAHFAPRLLPLDHYLVNFTCLRPSPEEAVAQAESEVGRPEEEEAEA
AAGLELCSGSGPFTFLHFDKNFVQLCLSAEPSEAPRLLAPAALAFRFVEVLLINNNNSSQFTCGVLCR
WSEECGRAAGRACGFAQPGCSCPGEAGAGSTTTTSPGPPAAHTLSNALVPGGPAPPAEADLHSGS SNDLFTTEMRYGEEPEEEPKVKTQWPRSADEPGLYMAQTGDPAAEEWSPWSVCSLTCGQGLQVRT RSCVSSPYGTLCSGPLRETRPCNNSATCPVHGVWEEWGSWSLCSRSCGRGSRSRMRTCVPPQHG
GKACEGPELQTKLCSMAACPVEGQWLEWGPWGPCSTSCANGTQQRSRKCSVAGPAWATCTGALT
DTRECSNLECPATDSKWGPWNAWSLCSKTCDTGWQRRFRMCQATGTQGYPCEGTGEEVKPCSEK
RCPAFHEMCRDEYVMLMTWKKAAAGEIIYNKCPPNASGSASRRCLLSAQGVAYWGLPSFARCISHE
YRYLYLSLREHLAKGQRMLAGEGMSQWRSLQELLARRTYYSGDLLFSVDILRNVTDTFKRATYVP
SADDVQRFFQVVSFMVDAENKEKWDDAQQVSPGSVHLLRWEDFIHLVGDALKAFQSSLIVTDWLVI
SIQREPVSAVSSDITFPMRGRRGMKDWVRHSEDRLFLPKEVLSLSSPGKPATSGAAGSPGRGRGP
GTVPPGPGHSHQRLLPADPDESSYFVIGAVLYRTLGLILPPPRPPLAVTSRVMTVTVRPPTQPPAEP
LITVELSYIINGTTDPHCASWDYSRADASSGDWDTENCQTLETQAAHTRCQCQHLSTFAVLAQPPK
DLTLELAGSPSVPLVIGCAVSCMALLTLLAIYAAFWRFIKSERSIILLNFCLSILASNILILVGQSRVLSKGV
CTMTAAFLHFFFLSSFCWVLTEAWQSYLAVIGRMRTRLVRKRFLCLGWGLPALWAVSVGFTRTKGY
GTSSYCWLSLEGGLLYAFVGPAAVIVLVNMLIGIIVFNKLMARDGISDKSKKQRAGSERCPWASLLLPC
SACGAVPSPLLSSASARNAMASLWSSCVVLPLLALTWMSAVLAMTDRRSVLFQALFAVFNSAQGFVI TAVHCFLRREVQDVVKCQMGVCRADESEDSPDSCKNGQLQILSDFEKDVDLACQTVLFKEVNTCNP STITGTLSRLSLDEDEEPKSCLVGPEGSLSFSPLPGNILVPMAASPGLGEPPPPQEANPVYMCGEGGL
RQLDLTWLRPTEPGSEGDYMVLPRRTLSLQPGGGGGGGEDAPRARPEGTPRRAAKTVAHTEGYPS FLSVDHSGLGLGPAYGSLQNPYGMTFQPPPPTPSARQVPEPGERSRTMPRTVPGSTMKMGSLERK KLRYSDLDFEKVMHTRKRHSELYHELNQKFHTFDRYRSQSTAKREKRWSVSSGGAAERSVCTDKPS PGERPSLSQHRRHQSWSTFKSMTLGSLPPKPRERLTLHRAAAWEPTEPPDGDFQTEV
>sp|O14514|AGRB1_HUMAN Adhesion G protein-coupled receptor B1 OS=Homo sapiens OX=9606 GN=ADGRB1 PE=1 SV=2
MRGQAAAPGPVWILAPLLLLLLLLGRRARAAAGADAGPGPEPCATLVQGKFFGYFSAAAVFPANASR CSWTLRNPDPRRYTLYMKVAKAPVPCSGPGRVRTYQFDSFLESTRTYLGVESFDEVLRLCDPSAPLA FLQASKQFLQMRRQQPPQHDGLRPRAGPPGPTDDFSVEYLVVGNRNPSRAACQMLCRWLDACLAG SRSSHPCGIMQTPCACLGGEAGGPAAGPLAPRGDVCLRDAVAGGPENCLTSLTQDRGGHGATGGW KLWSLWGECTRDCGGGLQTRTRTCLPAPGVEGGGCEGVLEEGRQCNREACGPAGRTSSRSQSLR
STDARRREELGDELQQFGFPAPQTGDPAAEEWSPWSVCSSTCGEGWQTRTRFCVSSSYSTQCSGP LREQRLCNNSAVCPVHGAWDEWSPWSLCSSTCGRGFRDRTRTCRPPQFGGNPCEGPEKQTKFCNI ALCPGRAVDGNWNEWSSWSACSASCSQGRQQRTRECNGPSYGGAECQGHWVETRDCFLQQCPV DGKWQAWASWGSCSVTCGAGSQRRERVCSGPFFGGAACQGPQDEYRQCGTQRCPEPHEICDED NFGAVIWKETPAGEVAAVRCPRNATGLILRRCELDEEGIAYWEPPTYIRCVSIDYRNIQMMTREHLAK
AQRGLPGEGVSEVIQTLVEISQDGTSYSGDLLSTIDVLRNMTEIFRRAYYSPTPGDVQNFVQILSNLL AEENRDKWEEAQLAGPNAKELFRLVEDFVDVIGFRMKDLRDAYQVTDWLVLSIHKLPASGATDISFP MKGWRATGDWAKVPEDRVTVSKSVFSTGLTEADEASVFVVGTVLYRNLGSFLALQRNTTVLNSKVI SVTVKPPPRSLRTPLEIEFAHMYNGTTNQTCILWDETDVPSSSAPPQLGPWSWRGCRTVPLDALRT RCLCDRLSTFAILAQLSADANMEKATLPSVTLIVGCGVSSLTLLMLVIIYVSVWRYIRSERSVILINFCLSI
ISSNALILIGQTQTRNKVVCTLVAAFLHFFFLSSFCWVLTEAWQSYMAVTGHLRNRLIRKRFLCLGWGL PALVVAISVGFTKAKGYSTMNYCWLSLEGGLLYAFVGPAAAVVLVNMVIGILVFNKLVSKDGITDKKLK ERAGASLWSSCWLPLLALTWMSAVLAVTDRRSALFQILFAVFDSLEGFVIVMVHCILRREVQDAVKC RVVDRQEEGNGDSGGSFQNGHAQLMTDFEKDVDLACRSVLNKDIAACRTATITGTLKRPSLPEEEKL KLAHAKGPPTNFNSLPANVSKLHLHGSPRYPGGPLPDFPNHSLTLKRDKAPKSSFVGDGDIFKKLDSE
LSRAQEKALDTSYVILPTATATLRPKPKEEPKYSIHIDQMPQTRLIHLSTAPEASLPARSPPSRQPPSGG PPEAPPAQPPPPPPPPPPPPQQPLPPPPNLEPAPPSLGDPGEPAAHPGPSTGPSTKNENVATLSVSS LERRKSRYAELDFEKIMHTRKRHQDMFQDLNRKLQHAAEKDKEVLGPDSKPEKQQTPNKRPWESLR KAHGTPTWVKKELEPLQPSPLELRSVEWERSGATIPLVGQDIIDLQTEV
>sp|Q8IWK6|AGRA3_HUMAN Adhesion G protein-coupled receptor A3 OS=Homo sapiens OX=9606 GN=ADGRA3 PE=1 SV=2
MEPPGRRRGRAQPPLLLPLSLLALLALLGGGGGGGAAALPAGCKHDGRPRGAGRAAGAAEGKVVCS SLELAQVLPPDTLPNRTVTLILSNNKISELKNGSFSGLSLLERLDLRNNLISSIDPGAFWGLSSLKRLDLT NNRIGCLNADIFRGLTNLVRLNLSGNLFSSLSQGTFDYLASLRSLEFQTEYLLCDCNILWMHRWVKEK NITVRDTRCVYPKSLQAQPVTGVKQELLTCDPPLELPSFYMTPSHRQVVFEGDSLPFQCMASYIDQD MQVLWYQDGRIVETDESQGIFVEKNMIHNCSLIASALTISNIQAGSTGNWGCHVQTKRGNNTRTVDIV
VLESSAQYCPPERWNNKGDFRWPRTLAGITAYLQCTRNTHGSGIYPGNPQDERKAWRRCDRGGF
WADDDYSRCQYANDVTRVLYMFNQMPLNLTNAVATARQLLAYTVEAANFSDKMDVIFVAEMIEKFG
RFTKEEKSKELGDVMVDIASNIMLADERVLWLAQREAKACSRIVQCLQRIATYRLAGGAHVYSTYSP
NIALEAYVIKSTGFTGMTCTVFQKVAASDRTGLSDYGRRDPEGNLDKQLSFKCNVSNTFSSLALKNT
IVEASIQLPPSLFSPKQKRELRPTDDSLYKLQLIAFRNGKLFPATGNSTNLADDGKRRTVVTPVILTKI
DGVNVDTHHIPVNVTLRRIAHGADAVAARWDFDLLNGQGGWKSDGCHILYSDENITTIQCYSLSNYA
VLMDLTGSELYTQAASLLHPVVYTTAIILLLCLLAVIVSYIYHHSLIRISLKSWHMLVNLCFHIFLTCVVFV
GGITQTRNASICQAVGIILHYSTLATVLWVGVTARNIYKQVTKKAKRCQDPDEPPPPPRPMLRFYLIGG
GIPIIVCGITAAANIKNYGSRPNAPYCWMAWEPSLGAFYGPASFITFVNCMYFLSIFIQLKRHPERKYEL KEPTEEQQRLAANENGEINHQDSMSLSLISTSALENEHTFHSQLLGASLTLLLYVALWMFGALAVSLYY PLDLVFSFVFGATSLSFSAFFVVHHCVNREDVRLAWIMTCCPGRSSYSVQVNVQPPNSNGTNGEAPK
CPNSSAESSCTNKSASSFKNSSQGCKLTNLQAAAAQCHANSLPLNSTPQLDNSLTEHSMDNDIKMHV
APLEVQFRTNVHSSRHHKNRSKGHRASRLTVLREYAYDVPTSVEGSVQNGLPKSRLGNNEGHSRSR
RAYLAYRERQYNPPQQDSSDACSTLPKSSRNFEKPVSTTSKKDALRKPAVVELENQQKSYGLNLAIQ NGPIKSNGQEGPLLGTDSTGNVRTGLWKHETTV
>sp|Q96PE1 |AGRA2_HUMAN Adhesion G protein-coupled receptor A2 OS=Homo sapiens OX=9606
GN=ADGRA2 PE=1 SV=2
MGAGGRRMRGAPARLLLPLLPWLLLLLAPEARGAPGCPLSIRSCKCSGERPKGLSGGVPGPARRRV
VCSGGDLPEPPEPGLLPNGTVTLLLSNNKITGLRNGSFLGLSLLEKLDLRNNIISTVQPGAFLGLGELKR
LDLSNNRIGCLTSETFQGLPRLLRLNISGNIFSSLQPGVFDELPALKWDLGTEFLTCDCHLRWLLPWA
QNRSLQLSEHTLCAYPSALHAQALGSLQEAQLCCEGALELHTHHLIPSLRQVVFQGDRLPFQCSASYL
GNDTRIRWYHNRAPVEGDEQAGILLAESLIHDCTFITSELTLSHIGVWASGEWECTVSMAQGNASKKV
EIVVLETSASYCPAERVANNRGDFRWPRTLAGITAYQSCLQYPFTSVPLGGGAPGTRASRRCDRAGR
WEPGDYSHCLYTNDITRVLYTFVLMPINASNALTLAHQLRVYTAEAASFSDMMDVVYVAQMIQKFLG
YVDQIKELVEVMVDMASNLMLVDEHLLWLAQREDKACSRIVGALERIGGAALSPHAQHISVNARNV
ALEAYLIKPHSYVGLTCTAFQRREGGVPGTRPGSPGQNPPPEPEPPADQQLRFRCTTGRPNVSLSS
FHIKNSVALASIQLPPSLFSSLPAALAPPVPPDCTLQLLVFRNGRLFHSHSNTSRPGAAGPGKRRGV
ATPVIFAGTSGCGVGNLTEPVAVSLRHWAEGAEPVAAWWSQEGPGEAGGWTSEGCQLRSSQPNV
SALHCQHLGNVAVLMELSAFPREVGGAGAGLHPVVYPCTALLLLCLFATIITYILNHSSIRVSRKGWHM
LLNLCFHIAMTSAVFAGGITLTNYQMVCQAVGITLHYSSLSTLLWMGVKARVLHKELTWRAPPPQEGD
PALPTPSPMLRFYLIAGGIPLIICGITAAVNIHNYRDHSPYCWLVWRPSLGAFYIPVALILLITWIYFLCAG
LRLRGPLAQNPKAGNSRASLEAGEELRGSTRLRGSGPLLSDSGSLLATGSARVGTPGPPEDGDSLYS
PGVQLGALVTTHFLYLAMWACGALAVSQRWLPRVVCSCLYGVAASALGLFVFTHHCARRRDVRASW
RACCPPASPAAPHAPPRALPAAAEDGSPVFGEGPPSLKSSPSGSSGHPLALGPCKLTNLQLAQSQVC
EAGAAAGGEGEPEPAGTRGNLAHRHPNNVHHGRRAHKSRAKGHRAGEACGKNRLKALRGGAAGAL
ELLSSESGSLHNSPTDSYLGSSRNSPGAGLQLEGEPMLTPSEGSDTSAAPLSEAGRAGQRRSASRD
SLKGGGALEKESHRRSYPLNAASLNGAPKGGKYDDVTLMGAEVASGGCMKTGLWKSETTV
>sp|Q8WXG9|AGRV1_HUMAN Adhesion G-protein coupled receptor V1 OS=Homo sapiens OX=9606 GN=ADGRV1 PE=1 SV=2
MSVFLGPGMPSASLLVNLLSALLILFVFGETEIRFTGQTEFVVNETSTTVIRLIIERIGEPANVTAIVSLYG EDAGDFFDTYAAAFIPAGETNRTVYIAVCDDDLPEPDETFIFHLTLQKPSANVKLGWPRTVTVTILSND
NAFGIISFNMLPSIAVSEPKGRNESMPLTLIREKGTYGMVMVTFEVEGGPNPPDEDLSPVKGNITFPPG RATVIYNLTVLDDEVPENDEIFLIQLKSVEGGAEINTSRNSIEIIIKKNDSPVRFLQSIYLVPEEDHILIIPVV RGKDNNGNLIGSDEYEVSISYAVTTGNSTAHAQQNLDFIDLQPNTTVVFPPFIHESHLKFQIVDDTIPEI AESFHIMLLKDTLQGDAVLISPSVVQVTIKPNDKPYGVLSFNSVLFERTVIIDEDRISRYEEITWRNGGT
HGNVSANWVLTRNSTDPSPVTADIRPSSGVLHFAQGQMLATIPLTVVDDDLPEEAEAYLLQILPHTIRG GAEVSEPAELLFYIQDSDDVYGLITFFPMENQKIESSPGERYLSLSFTRLGGTKGDVRLLYSVLYIPAGA VDPLQAKEGILNISRRNDLIFPEQKTQVTTKLPIRNDAFLQNGAHFLVQLETVELLNIIPLIPPISPRFGEIC NISLLVTPAIANGEIGFLSNLPIILHEPEDFAAEVVYIPLHRDGTDGQATVYWSLKPSGFNSKAVTPDDIG PFNGSVLFLSGQSDTTINITIKGDDIPEMNETVTLSLDRVNVENQVLKSGYTSRDLIILENDDPGGVFEF
SPASRGPYVIKEGESVELHIIRSRGSLVKQFLHYRVEPRDSNEFYGNTGVLEFKPGEREIVITLLARLDG IPELDEHYWVVLSSHGERESKLGSATIVNITILKNDDPHGIIEFVSDGLIVMINESKGDAIYSAVYDVVRN
RGNFGDVSVSWVVSPDFTQDVFPVQGTVVFGDQEFSKNITIYSLPDEIPEEMEEFTVILLNGTGGAKV GNRTTATLRIRRNDDPIYFAEPRVVRVQEGETANFTVLRNGSVDVTCMVQYATKDGKATARERDFIPV EKGETLIFEVGSRQQSISIFVNEDGIPETDEPFYIILLNSTGDTVVYQYGVATVIIEANDDPNGIFSLEPID
KAVEEGKTNAFWILRHRGYFGSVSVSWQLFQNDSALQPGQEFYETSGTVNFMDGEEAKPIILHAFPD KIPEFNEFYFLKLVNISGGSPGPGGQLAETNLQVTVMVPFNDDPFGVFILDPECLEREVAEDVLSEDD
MSYITNFTILRQQGVFGDVQLGWEILSSEFPAGLPPMIDFLLVGIFPTTVHLQQHMRRHHSGTDALYFT GLEGAFGTVNPKYHPSRNNTIANFTFSAWVMPNANTNGFIIAKDDGNGSIYYGVKIQTNESHVTLSLHY KTLGSNATYIAKTTVMKYLEESVWLHLLIILEDGIIEFYLDGNAMPRGIKSLKGEAITDGPGILRIGAGING
NDRFTGLMQDVRSYERKLTLEEIYELHAMPAKSDLHPISGYLEFRQGETNKSFIISARDDNDEEGEELF ILKLVSVYGGARISEENTTARLTIQKSDNANGLFGFTGACIPEIAEEGSTISCVVERTRGALDYVHVFYTI
SQIETDGINYLVDDFANASGTITFLPWQRSEVLNIYVLDDDIPELNEYFRVTLVSAIPGDGKLGSTPTSG ASIDPEKETTDITIKASDHPYGLLQFSTGLPPQPKDAMTLPASSVPHITVEEEDGEIRLLVIRAQGLLGRV TAEFRTVSLTAFSPEDYQNVAGTLEFQPGERYKYIFINITDNSIPELEKSFKVELLNLEGGVAELFRVDG SGSGDGDMEFFLPTIHKRASLGVASQILVTIAASDHAHGVFEFSPESLFVSGTEPEDGYSTVTLNVIRH HGTLSPVTLHWNIDSDPDGDLAFTSGNITFEIGQTSANITVEILPDEDPELDKAFSVSVLSVSSGSLGAH
INATLTVLASDDPYGIFIFSEKNRPVKVEEATQNITLSIIRLKGLMGKVLVSYATLDDMEKPPYFPPNLAR ATQGRDYIPASGFALFGANQSEATIAISILDDDEPERSESVFIELLNSTLVAKVQSRSIPNSPRLGPKVET IAQLIIIANDDAFGTLQLSAPIVRVAENHVGPIINVTRTGGAFADVSVKFKAVPITAIAGEDYSIASSDVVLL EGETSKAVPIYVINDIYPELEESFLVQLMNETTGGARLGALTEAVIIIEASDDPYGLFGFQITKLIVEEPEF
NSVKVNLPIIRNSGTLGNVTVQWVATINGQLATGDLRVVSGNVTFAPGETIQTLLLEVLADDVPEIEEVI QVQLTDASGGGTIGLDRIANIIIPANDDPYGTVAFAQMVYRVQEPLERSSCANITVRRSGGHFGRLLLF
YSTSDIDVVALAMEEGQDLLSYYESPIQGVPDPLWRTWMNVSAVGEPLYTCATLCLKEQACSAFSFF SASEGPQCFWMTSWISPAVNNSDFWTYRKNMTRVASLFSGQAVAGSDYEPVTRQWAIMQEGDEFA
NLTVSILPDDFPEMDESFLISLLEVHLMNISASLKNQPTIGQPNISTVVIALNGDAFGVFVIYNISPNTSED GLFVEVQEQPQTLVELMIHRTGGSLGQVAVEWRVVGGTATEGLDFIGAGEILTFAEGETKKTVILTILD
DSEPEDDESIIVSLVYTEGGSRILPSSDTVRVNILANDNVAGIVSFQTASRSVIGHEGEILQFHVIRTFPG RGNVTVNWKIIGQNLELNFANFSGQLFFPEGSLNTTLFVHLLDDNIPEEKEVYQVILYDVRTQGVPPAG
IALLDAQGYAAVLTVEASDEPHGVLNFALSSRFVLLQEANITIQLFINREFGSLGAINVTYTTVPGMLSLK
NQTVGNLAEPEVDFVPIIGFLILEEGETAAAINITILEDDVPELEEYFLVNLTYVGLTMAASTSFPPRLDSE
GLTAQVIIDANDGARGVIEWQQSRFEVNETHGSLTLVAQRSREPLGHVSLFVYAQNLEAQVGLDYIFT
PMILHFADGERYKNVNIMILDDDIPEGDEKFQLILTNPSPGLELGKNTIALIIVLANDDGPGVLSFNNSEH
FFLREPTALYVQESVAVLYIVREPAQGLFGTVTVQFIVTEVNSSNESKDLTPSKGYIVLEEGVRFKALQI
SAILDTEPEMDEYFVCTLFNPTGGARLGVHVQTLITVLQNQAPLGLFSISAVENRATSIDIEEANRTVYL
NVSRTNGIDLAVSVQWETVSETAFGMRGMDVVFSVFQSFLDESASGWCFFTLENLIYGIMLRKSSVT
VYRWQGIFIPVEDLNIENPKTCEAFNIGFSPYFVITHEERNEEKPSLNSVFTFTSGFKLFLVQTIIILESSQ
VRYFTSDSQDYLIIASQRDDSELTQVFRWNGGSFVLHQKLPVRGVLTVALFNKGGSVFLAISQANARL
NSLLFRWSGSGFINFQEVPVSGTTEVEALSSANDIYLIFAENVFLGDQNSIDIFIWEMGQSSFRYFQSV
DFAAVNRIHSFTPASGIAHILLIGQDMSALYCWNSERNQFSFVLEVPSAYDVASVTVKSLNSSKNLIALV
GAHSHIYELAYISSHSDFIPSSGELIFEPGEREATIAVNILDDTVPEKEESFKVQLKNPKGGAEIGINDSV
TITILSNDDAYGIVAFAQNSLYKQVEEMEQDSLVTLNVERLKGTYGRITIAWEADGSISDIFPTSGVILFT
EGQVLSTITLTILADNIPELSEVVIVTLTRITTEGVEDSYKGATIDQDRSKSVITTLPNDSPFGLVGWRAA
SVFIRVAEPKENTTTLQLQIARDKGLLGDIAIHLRAQPNFLLHVDNQATENEDYVLQETIIIMKENIKEAH
AEVSILPDDLPELEEGFIVTITEVNLVNSDFSTGQPSVRRPGMEIAEIMIEENDDPRGIFMFHVTRGAGE
VITAYEVPPPLNVLQVPVVRLAGSFGAVNVYWKASPDSAGLEDFKPSHGILEFADKQVTAMIEITIIDDA
EFELTETFNISLISVAGGGRLGDDWVTWIPQNDSPFGVFGFEEKTVMIDESLSSDDPDSYVTLTVVR
SPGGKGTVRLEWTIDEKAKHNLSPLNGTLHFDETESQKTIVLHTLQDTVLEEDRRFTIQLISIDEVEISPV
KGSASIIIRGDKRASGEVGIAPSSRHILIGEPSAKYNGTAIISLVRGPGILGEVTVFWRIFPPSVGEFAETS
GKLTMRDEQSAVIWIQALNDDIPEEKSFYEFQLTAVSEGGVLSESSSTANITWASDSPYGRFAFSHE
QLRVSEAQRVNITIIRSSGDFGHVRLWYKTMSGTAEAGLDFVPAAGELLFEAGEMRKSLHVEILDDDY
PEGPEEFSLTITKVELQGRGYDFTIQENGLQIDQPPEIGNISIVRIIIMKNDNAEGIIEFDPKYTAFEVEED
VGLIMIPVVRLHGTYGYVTADFISQSSSASPGGVDYILHGSTVTFQHGQNLSFINISIIDDNESEFEEPIEI
LLTGATGGAVLGRHLVSRIIIAKSDSPFGVIRFLNQSKISIANPNSTMILSLVLERTGGLLGEIQVNWETV
GPNSQEALLPQNRDIADPVSGLFYFGEGEGGVRTIILTIYPHEEIEVEETFIIKLHLVKGEAKLDSRAKDV
TLTIQEFGDPNGVVQFAPETLSKKTYSEPLALEGPLLITFFVRRVKGTFGEIMVYWELSSEFDITEDFLS
TSGFFTIADGESEASFDVHLLPDEVPEIEEDYVIQLVSVEGGAELDLEKSITWFSVYANDDPHGVFALY
SDRQSILIGQNLIRSIQINITRLAGTFGDVAVGLRISSDHKEQPIVTENAERQLWKDGATYKVDVVPIKN
QVFLSLGSNFTLQLVTVMLVGGRFYGMPTILQEAKSAVLPVSEKAANSQVGFESTAFQLMNITAGTSH
VMISRRGTYGALSVAWTTGYAPGLEIPEFIWGNMTPTLGSLSFSHGEQRKGVFLWTFPSPGWPEAF
VLHLSGVQSSAPGGAQLRSGFIVAEIEPMGVFQFSTSSRNIIVSEDTQMIRLHVQRLFGFHSDLIKVSY
QTTAGSAKPLEDFEPVQNGELFFQKFQTEVDFEITIINDQLSEIEEFFYINLTSVEIRGLQKFDVNWSPR
LNLDFSVAVITILDNDDLAGMDISFPETTVAVAVDTTLIPVETESTTYLSTSKTTTILQPTNWAIVTEATG
VSAIPEKLVTLHGTPAVSEKPDVATVTANVSIHGTFSLGPSIVYIEEEMKNGTFNTAEVLIRRTGGFTGN
VSITVKTFGERCAQMEPNALPFRGIYGISNLTWAVEEEDFEEQTLTLIFLDGERERKVSVQILDDDEPE
GQEFFYVFLTNPQGGAQIVEEKDDTGFAAFAMVIITGSDLHNGIIGFSEESQSGLELREGAVMRRLHLI
VTRQPNRAFEDVKVFWRVTLNKTWVLQKDGVNLVEELQSVSGTTTCTMGQTKCFISIELKPEKVPQ
VEVYFFVELYEATAGAAINNSARFAQIKILESDESQSLVYFSVGSRLAVAHKKATLISLQVARDSGTG
LMMSVNFSTQELRSAETIGRTIISPAISGKDFVITEGTLVFEPGQRSTVLDVILTPETGSLNSFPKRFQI
VLFDPKGGARIDKVYGTANITLVSDADSQAIWGLADQLHQPVNDDILNRVLHTISMKVATENTDEQL SAMMHLIEKITTEGKIQAFSVASRTLFYEILCSLINPKRKDTRGFSHFAEVTENFAFSLLTNVTCGSPG
EKSKTILDSCPYLSILALHWYPQQINGHKFEGKEGDYIRIPERLLDVQDAEIMAGKSTCKLVQFTEYS
SQQWFISGNNLPTLKNKVLSLSVKGQSSQLLTNDNEVLYRIYAAEPRIIPQTSLCLLWNQAAASWLS
DSQFCKWEETADYVECACSHMSVYAVYARTDNLSSYNEAFFTSGFICISGLCLAVLSHIFCARYSMF
AAKLLTHMMAASLGTQILFLASAYASPQLAEESCSAMAAVTHYLYLCQFSWMLIQSVNFWYVLVMND
EHTERRYLLFFLLSWGLPAFVVILLIVILKGIYHQSMSQIYGLIHGDLCFIPNVYAALFTAALVPLTCLVVV
FVVFIHAYQVKPQWKAYDDVFRGRTNAAEIPLILYLFALISVTWLWGGLHMAYRHFWMLVLFVIFNSLQ GLYVFMVYFILHNQMCCPMKASYTVEMNGHPGPSTAFFTPGSGMPPAGGEISKSTQNLIGAMEEVPP DWERASFQQGSQASPDLKPSPQNGATFPSSGGYGQGSLIADEESQEFDDLIFALKTGAGLSVSDNES
GQGSQEGGTLTDSQIVELRRIPIADTHL
The 10 amino acid positions (underlined) for replacements by other amino acids in the 10Mut Adhesion G-protein coupled receptor G1 GAIN domain (bold).
>sp|Q9Y653|AGRG1_HUMAN Adhesion G-protein coupled receptor G1 OS=Homo sapiens OX=9606
GN=ADGRG1 PE=1 SV=2
MTPQSLLQTTLFLLSLLFLVQGAHGRGHREDFRFCSQRNQTHRSSLHYKPTPDLRISIENSEEALTVH
APFPAAHPASRSFPDPRGLYHFCLYWNRHAGRLHLLYGKRDFLLSDKASSLLCFQHQEESLAQGPPL
LATSVTSWWSPQNISLPSAASFTFSFHSPPHTAAHNASVDMCELKRDLQLLSQFLKHPQKASRRPS
AAPASQQLQSLESKLTSVRFMGDMVSFEEDRINATVWKLQPTAGLQDLHIHSRQEEEQSEIMEYSV
LLPRTLFQRTKGRSGEAEKRLLLVDFSSQALFQDKNSSQVLGEKVLGIVVQNTKVANLTEPWLTFQ
HQLQPKNVTLQCVFWVEDPTLSSPGHWSSAGCETVRRETQTSCFCNHLTYFAVLMVSSVEVDAVH
KHYLSLLSYVGCWSALACLVTIAAYLCSRVPLPCRRKPRDYTIKVHMNLLLAVFLLDTSFLLSEPVALT
GSEAGCRASAIFLHFSLLTCLSWMGLEGYNLYRLWEVFGTYVPGYLLKLSAMGWGFPIFLVTLVALV
DVDNYGPIILAVHRTPEGVIYPSMCWIRDSLVSYITNLGLFSLVFLFNMAMLATMVVQILRLRPHTQKW
SHVLTLLGLSLVLGLPWALIFFSFASGTFQLWLYLFSIITSFQGFLIFIWYWSMRLQARGGPSPLKSNS
DSARLPISSGSTSSSRI
References
1 . H. H. Lin et al., Autocatalytic cleavage of the EMR2 receptor occurs at a conserved G protein- coupled receptor proteolytic site motif. J Biol Chem 279, 31823-31832 (2004).
2. J. X. Gray et al., CD97 is a processed, seven-transmembrane, heterodimeric receptor associated with inflammation. Journal of Immunology 157 , 5438-5447 (1996).
3. H.-H. Lin, M. Stacey, S. Yona, G.-W. Chang, in Adhesion-GPCRs. (Springer, Boston, MA, 2010), pp. 49-58.
4. H. M. Stoveken, A. G. Hajduczok, L. Xu, G. G. Tall, Adhesion G protein-coupled receptors are activated by exposure of a cryptic tethered agonist. Proc Natl Acad Sei U S A 112, 6194-6199 (2015).
5. I. Liebscher et al., A tethered agonist within the ectodomain activates the adhesion G protein- coupled receptors GPR126 and GPR133. Cell Rep 9, 2018-2026 (2014).
6. S. E. Boyden et al., Vibratory Urticaria Associated with a Missense Variant in ADGRE2. N Engl J Med 374, 656-663 (2016).
7. O. N. Karpus et al., Shear stress-dependent downregulation of the adhesion-G protein-coupled receptor CD97 on circulating leukocytes upon contact with its ligand CD55. J Immunol 190, 3740- 3748 (2013).
8. R. Luo et al., Mechanism for adhesion G protein-coupled receptor GPR56-mediated RhoA activation induced by collagen III stimulation. PLoS One 9, e100043 (2014).
9. S. C. Petersen et al., The adhesion GPCR GPR126 has distinct, domain-dependent functions in Schwann cell development mediated by interaction with laminin-211. Neuron 85, 755-769 (2015).
10. N. Scholz et al., The adhesion GPCR latrophilin/CIRL shapes mechanosensation. Cell Rep 11 , 866-874 (2015).
11. N. Scholz et al., Mechano-dependent signaling by Latrophilin/CIRL quenches cAMP in proprioceptive neurons. Elite 6, e28360 (2017).
12. J. P. White et al., G protein-coupled receptor 56 regulates mechanical overload-induced muscle hypertrophy. Proc Natl Acad Sc/ U S A 111 , 15756-15761 (2014).
13. C. Wilde et al., The constitutive activity of the adhesion GPCR GPR114/ADGRG5 is mediated by its tethered agonist. FASEB J 30, 666-673 (2016).
14. A. Valbuena et al., On the remarkable mechanostability of scaffoldins and the mechanical clamp motif. Proceedings of the National Academy of Sciences 106, 13791-13796 (2009).
15. A. S. Oude Vrielink, T. D. Vance, A. M. de Jong, P. L. Davies, I. K. Voets, Unusually high mechanical stability of bacterial adhesin extender domains having calcium clamps. PloS one 12, e0174682 (2017).
16. C.-L. Chyan et al., Reversible mechanical unfolding of single ubiquitin molecules. Biophysical journal 87, 3995-4006 (2004).
17. M. Carrion-Vazquez et al., The mechanical stability of ubiquitin is linkage dependent. Nature Structural & Molecular Biology A0, 738-743 (2003).
18. P. E. Marszalek et al., Mechanical unfolding intermediates in titin modules. Nature 402, 100-103 (1999).
19. M. Carrion-Vazquez et al., Mechanical and chemical unfolding of a single protein: a comparison. Proceedings of the National Academy of Sciences 96, 3694-3699 (1999).
20. H. Li et al., Reverse engineering of the giant muscle protein titin. Nature 418, 998-1002 (2002).
21. M. Rief, M. Gautel, F. Oesterhelt, J. M. Fernandez, H. E. Gaub, Reversible unfolding of individual titin immunoglobulin domains by AFM. science 276, 1109-1112 (1997).
22. D. Craig, M. Gao, K. Schulten, V. Vogel, Tuning the mechanical stability of fibronectin type III modules through sequence variations. Structure 12, 21-30 (2004).
23. A. F. Oberhauser, C. Badilla-Fernandez, M. Carrion-Vazquez, J. M. Fernandez, The mechanical hierarchies of fibronectin observed with single-molecule AFM. Journal of molecular biology 319, 433-447 (2002).
24. A. F. Oberhauser, P. E. Marszalek, H. P. Erickson, J. M. Fernandez, The molecular elasticity of the extracellular matrix protein tenascin. Nature 393, 181-185 (1998).
25. M. Rief, J. Pascual, M. Saraste, H. E. Gaub, Single molecule force spectroscopy of spectrin repeats: low unfolding forces in helix bundles. Journal of molecular biology 286, 553-561 (1999).
26. J. P. Junker, F. Ziegler, M. Rief, Ligand-dependent equilibrium fluctuations of single calmodulin molecules. Science 323, 633-637 (2009).
27. J. Fritz, A. G. Katopodis, F. Kolbinger, D. Anselmetti, Force-mediated kinetics of single P- selectin/ligand complexes observed by atomic force microscopy. Proceedings of the National Academy of Sciences 95, 12283-12288 (1998).
28. W. Hanley et al., Single molecule characterization of P-selectin/ligand binding. Journal of biological chemistry 278, 10556-10561 (2003).
29. F. Li, S. D. Redick, H. P. Erickson, V. T. Moy, Force measurements of the a5p1 integrin— fibronectin interaction. Biophysical journal 84, 1252-1262 (2003).
30. W. Baumgartner et al., Cadherin interaction probed by atomic force microscopy. Proceedings of the National Academy of Sciences 97, 4005-4010 (2000).
31 . T. Yago et al., Platelet glycoprotein Iba forms catch bonds with human WT vWF but not with type 2B von Willebrand disease vWF. The Journal of clinical investigation 118, 3195-3207 (2008).
32. L. Ju, J.-f. Dong, M. A. Cruz, C. Zhu, The N-terminal flanking region of the A1 domain regulates the force-dependent binding of von Willebrand factor to platelet glycoprotein Iba. Journal of Biological Chemistry 288, 32289-32301 (2013).
33. V. C. Luca et al., Notch-Jagged complex structure implicates a catch bond in tuning ligand sensitivity. Science 355, 1320-1324 (2017).
34. W. R. Gordon et al., Mechanical allostery: evidence for a force requirement in the proteolytic activation of Notch. Developmental cell 33, 729-736 (2015).
35. H. Lu, B. Isralewitz, A. Krammer, V. Vogel, K. Schulten, Unfolding of titin immunoglobulin domains by steered molecular dynamics simulation. Biophysical journal 75, 662-671 (1998).
36. H. Li, M. Carrion-Vazquez, A. F. Oberhauser, P. E. Marszalek, J. M. Fernandez, Point mutations alter the mechanical stability of immunoglobulin modules. Nature structural biology 7, 1117-1120 (2000).
37. S. P. Ng et al., Designing an extracellular matrix protein with enhanced mechanical stability. Proceedings of the National Academy of Sciences 104, 9633-9637 (2007).
38. D. P. Sadler et al. , Identification of a mechanical rheostat in the hydrophobic core of protein L. Journal of molecular biology 393, 237-248 (2009).
39. S. P. Ng et al., Mechanical unfolding of TNfn3: the unfolding pathway of a fnlll domain probed by protein engineering, AFM and MD simulation. Journal of molecular biology 350, 776-789 (2005).
40. D. L. Guzman, A. Randall, P. Baldi, Z. Guan, Computational and single-molecule force studies of a macro domain protein reveal a key molecular determinant for mechanical stability. Proceedings of the National Academy of Sciences 107, 1989-1994 (2010).
41 . J. Seppala et al., Skeletal dysplasia mutations effect on human filamins’ structure and mechanosensing. Scientific reports 7, 1-14 (2017).
42. W. Stacklies, F. Graeter, Force Propagation in Proteins From Molecular Dynamics Simulations. Biophysical Journal 96, 589a (2009).
43. C. Schoeler et al., Mapping mechanical force propagation through biomolecular complexes. Nano letters 15, 7370-7376 (2015).
44. D. K. West, D. J. Brockwell, P. D. Olmsted, S. E. Radford, E. Paci, Mechanical resistance of proteins explained using simple molecular models. Biophysical journal 90, 287-297 (2006).
45. S. R. K. Ainavarapu et al., Contour length and refolding rate of a small protein controlled by engineered disulfide bonds. Biophysical journal 92, 225-233 (2007).
46. M. Carrion-Vazquez et al., Mechanical design of proteins studied by single-molecule force spectroscopy and protein engineering. Progress in biophysics and molecular biology 74, 63-91 (2000).
Additional references in Example 3
1 . Z. Liu et al. , Mapping mechanostable pulling geometries of a therapeutic anticalin/CTLA-4 protein complex. Nano letters, (2021).
2. B. Yang, H. Liu, Z. Liu, R. Doenen, M. A. Nash, Influence of fluorination on single-molecule unfolding and rupture pathways of a mechanostable protein adhesion complex. Nano letters 20, 8940-8950 (2020).
3. R. Merkel, P. Nassoy, A. Leung, K. Ritchie, E. Evans, Energy landscapes of receptor-ligand bonds explored with dynamic force spectroscopy. Nature 397, 50-53 (1999).
4. E. Evans, K. Ritchie, Dynamic strength of molecular adhesion bonds. Biophysical journal 72, 1541-1555 (1997).
5. O. K. Dudko, Decoding the mechanical fingerprints of biomolecules. Quarterly reviews of biophysics 49 , (2016).
6. O. K. Dudko, G. Hummer, A. Szabo, Theory, analysis, and interpretation of single-molecule force spectroscopy experiments. Proceedings of the National Academy of Sciences 105, 15755-15760 (2008).
7. N. Ollikainen, R. M. de Jong, T. Kortemme, Coupling protein side-chain and backbone flexibility improves the re-design of protein-ligand specificity. PLoS computational biology 11 , e1004335 (2015).
8. P. Wernet et al., The structure of the first coordination shell in liquid water. Science 304, 995-999 (2004).
9. J. B. Klauda et al., Update of the CHARMM all-atom additive force field for lipids: validation on six lipid types. The journal of physical chemistry B 114, 7830-7843 (2010).
10. P. Mark, L. Nilsson, Structure and dynamics of the TIP3P, SPC, and SPC/E water models at 298 K. The Journal of Physical Chemistry A 105, 9954-9960 (2001).
11. U. Essmann et al., A smooth particle mesh Ewald method. The Journal of chemical physics 103, 8577-8593 (1995). 12. E. W. Dijkstra, A short introduction to the art of programming. (Technische Hogeschool
Eindhoven Eindhoven, 1971), vol. 4.
Claims
CLAIMS A variant of the GAIN domain (G-protein-coupled receptor autoproteolysis-inducing domain) of an adhesion G protein-coupled receptor (ADGRG):
(I)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG1 (Adhesion G Protein-Coupled Receptor G1) of SEQ ID NO: 1 , wherein
(i) one of amino acid positions 3 to 5, preferably amino acid position 4 and one of amino acid positions 47 to 49, preferably amino acid position 48 of SEQ ID NO: 1 are replaced by cysteines,
(ii) one of amino acid positions 35 to 37, preferably amino acid position 36 and one of amino acid positions 180 to 182, preferably amino acid position 181 of SEQ ID NO: 1 are replaced by cysteines,
(iii) one of amino acid positions 19 to 21 , preferably amino acid position 20 and one of amino acid positions 124 to 126, preferably amino acid position 125 of SEQ ID NO: 1 are replaced by cysteines,
(iv) one of amino acid positions 56 to 58, preferably amino acid position 57 and one of amino acid positions 65 to 67, preferably amino acid position 66 of SEQ ID NO: 1 are replaced by cysteines,
(v) one of amino acid positions 16 to 18, preferably amino acid position 17 and one of amino acid positions 60 to 62, preferably amino acid position 61 of SEQ ID NO: 1 are replaced by cysteines,
(vi) the amino acid position at position 40 is replaced by glutamic acid, the amino acid position at position 43 is replaced by leucine and the amino acid position at position 181 is replaced by arginine,
(vii) the amino acid position at position 40 is replaced by glutamic acid, the amino acid position at position 43 is replaced by leucine, the amino acid position at position 135 is replaced by arginine and the amino acid position at position 181 is replaced by arginine, (viii) the amino acid position at position 40 is replaced by glutamic acid, the amino acid position at position 43 is replaced by leucine, the amino acid position at position 135 is replaced by arginine, the amino acid position at position 146is replaced by tyrosine, and the amino acid position at position 181 is replaced by arginine,
(ix) the amino acid position at position 33 is replaced by tryptophane, the amino acid position at position 40 is replaced by glutamic acid, the amino acid position at position 43 is replaced by leucine, the amino acid position at position 135 is replaced by arginine, the amino acid position at position 146 is replaced by tyrosine, the amino acid position at position 178 is replaced by tyrosine, the amino acid position at position 181 is replaced by arginine, the amino acid position at position 184 is replaced by glycine, the amino acid position at position 186 is replaced by glutamine, and the amino acid position at position
187 is replaced by glycine,
(x) the amino acid position at position 174 is replaced by isoleucine,
(xi) the amino acid position at position 218 is replaced by alanine, or
(xii) at least two of (i) to (xi) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in (A), item (i), (ii),
(iii), (iv) and/or (v), or the amino acid replacement(s) as defined in (A), item (vi), (vii), (viii), (ix), (x) and/or (xi), or the cysteines and/or the amino acid replacement(s) as defined in (A), item (xii) are retained;
(II)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRL4 (Adhesion G Protein-Coupled Receptor L4) of SEQ ID NO: 2, wherein
(i) one of amino acid positions 87 to 89, preferably amino acid position 88 and one of amino acid positions 123 to 125, preferably amino acid position 124 of SEQ ID NO: 2 are replaced by cysteines,
(ii) one of amino acid positions 111 to 113, preferably amino acid position 112 and one of amino acid positions 263 to 265, preferably amino acid position 264 of SEQ ID NO: 2 are replaced by cysteines,
(iii) one of amino acid positions 95 to 97, preferably amino acid position 96 and one of amino acid positions 192 to 194, preferably amino acid position 193 of SEQ ID NO: 2 are replaced by cysteines,
(iv) one of amino acid positions 134 to 136, preferably amino acid position 135 and one of amino acid positions 144 to 146, preferably amino acid position 145 of SEQ ID NO: 2 are replaced by cysteines,
(v) one of amino acid positions 93 to 95, preferably amino acid position 94 and one of amino acid positions 140 to 142, preferably amino acid position 141 of SEQ ID NO: 2 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(HI)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRL3 (Adhesion G Protein-Coupled Receptor L3) of SEQ ID NO: 3, wherein
(i) one of amino acid positions 81 to 83, preferably amino acid position 82 and one of amino acid positions 117 to 119, preferably amino acid position 118 of SEQ ID NO: 3 are replaced by cysteines,
(ii) one of amino acid positions 105 to 107, preferably amino acid position 106 and one of amino acid positions 261 to 263, preferably amino acid position 262 of SEQ ID NO: 3 are replaced by cysteines,
(iii) one of amino acid positions 89 to 91 , preferably amino acid position 90 and one of amino acid positions 188 to 190, preferably amino acid position 189 of SEQ ID NO: 3 are replaced by cysteines,
(iv) one of amino acid positions 128 to 130, preferably amino acid position 129 and one of amino acid positions 138 to 140, preferably amino acid position 139 of SEQ ID NO: 3 are replaced by cysteines,
(v) one of amino acid positions 87 to 89, preferably amino acid position 88 and one of amino acid positions 134 to 136, preferably amino acid position 135 of SEQ ID NO: 3 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(IV)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRL2 (Adhesion G Protein-Coupled Receptor L2) of SEQ ID NO: 4, wherein
(i) one of amino acid positions 94 to 96, preferably amino acid position 95 and one of amino acid positions 130 to 132, preferably amino acid position 131 of SEQ ID NO: 4 are replaced by cysteines,
(ii) one of amino acid positions 118 to 120, preferably amino acid position 119 and one of amino acid positions 272 to 274, preferably amino acid position 273 of SEQ ID NO: 4 are replaced by cysteines,
(iii) one of amino acid positions 102 to 104, preferably amino acid position 103 and one of amino acid positions 202 to 204, preferably amino acid position 203 of SEQ ID NO: 4 are replaced by cysteines,
(iv) one of amino acid positions 141 to 143, preferably amino acid position 144 and one of amino acid positions 151 to 153, preferably amino acid position 152 of SEQ ID NO: 4 are replaced by cysteines,
(v) one of amino acid positions 99 to 101 , preferably amino acid position 100 and one of amino acid positions 147 to 149, preferably amino acid position 148 of SEQ ID NO: 4 are replaced by cysteines, or
(vi) at least two or all three of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(V)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRL1 (Adhesion G Protein-Coupled Receptor L1) of SEQ ID NO: 5, wherein
(i) one of amino acid positions 94 to 96, preferably amino acid position 95 and one of amino acid positions 130 to 132, preferably amino acid position 131 of SEQ ID NO: 5 are
replaced by cysteines,
(ii) one of amino acid positions 118 to 120, preferably amino acid position 119 and one of amino acid positions 275 to 277, preferably amino acid position 276 of SEQ ID NO: 5 are replaced by cysteines,
(iii) one of amino acid positions 102 to 104, preferably amino acid position 103 and one of amino acid positions 202 to 204, preferably amino acid position 203 of SEQ ID NO: 5 are replaced by cysteines,
(iv) one of amino acid positions 141 to 143, preferably amino acid position 144 and one of amino acid positions 151 to 153, preferably amino acid position 152 of SEQ ID NO: 5 are replaced by cysteines,
(v) one of amino acid positions 100 to 102, preferably amino acid position 101 and one of amino acid positions 147 to 149, preferably amino acid position 148 of SEQ ID NO: 5 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(VI)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG7 (Adhesion G Protein-Coupled Receptor G7) of SEQ ID NO: 6, wherein
(i) one of amino acid positions 67 to 69, preferably amino acid position 68 and one of amino acid positions 102 to 104, preferably amino acid position 103 of SEQ ID NO: 6 are replaced by cysteines,
(ii) one of amino acid positions 90 to 92, preferably amino acid position 91 and one of amino acid positions 246 to 248, preferably amino acid position 247 of SEQ ID NO: 6 are replaced by cysteines,
(iii) one of amino acid positions 74 to 76, preferably amino acid position 75 and one of amino acid positions 186 to 188, preferably amino acid position 187 of SEQ ID NO: 6 are replaced by cysteines,
(vi) one of amino acid positions 64 to 66, preferably amino acid position 65 and one of amino acid positions 111 to 113, preferably amino acid position 112 of SEQ ID NO: 6 are replaced by cysteines, or
(v) at least two of (i) to (iv apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (v) are retained;
(VII)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG6 (Adhesion G Protein-Coupled Receptor G6) of SEQ ID NO: 7, wherein
(i) one of amino acid positions 35 to 37, preferably amino acid position 36 and one of
amino acid positions 84 to 86, preferably amino acid position 85 of SEQ ID NO: 7 are replaced by cysteines,
(ii) one of amino acid positions 73 to 75, preferably amino acid position 74 and one of amino acid positions 229 to 231 , preferably amino acid position 230 of SEQ ID NO: 7 are replaced by cysteines,
(iii) one of amino acid positions 56 to 58, preferably amino acid position 57 and one of amino acid positions 172 to 174, preferably amino acid position 173 of SEQ ID NO: 7 are replaced by cysteines,
(iv) one of amino acid positions 34 to 36, preferably amino acid position 35 and one of amino acid positions 103 to 105, preferably amino acid position 104 of SEQ ID NO: 7 are replaced by cysteines,
(v) one of amino acid positions 52 to 54, preferably amino acid position 53 and one of amino acid positions 98 to 100, preferably amino acid position 99 of SEQ ID NO: 7 are replaced by cysteines, or
(vi) at least two or all three of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(VIII)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG5 (Adhesion G Protein-Coupled Receptor G5) of SEQ ID NO: 8, wherein
(i) one of amino acid positions 45 to 47, preferably amino acid position 46 and one of amino acid positions 193 to 195, preferably amino acid position 194 of SEQ ID NO: 8 are replaced by cysteines,
(ii) one of amino acid positions 30 to 32, preferably amino acid position 31 and one of amino acid positions 137 to 139, preferably amino acid position 138 of SEQ ID NO: 8 are replaced by cysteines,
(iii) one of amino acid positions 66 to 68, preferably amino acid position 67 and one of amino acid positions 75 to 77, preferably amino acid position 76 of SEQ ID NO: 8 are replaced by cysteines,
(iv) one of amino acid positions 27 to 29, preferably amino acid position 28 and one of amino acid positions 70 to 72, preferably amino acid position 71 of SEQ ID NO: 8 are replaced by cysteines, or
(v) at least two of (i) to (iv) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (v) are retained;
(IX)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG4 (Adhesion G Protein-Coupled Receptor G4) of SEQ ID NO: 9, wherein
(i) one of amino acid positions 251 to 253, preferably amino acid position 252 and one of amino acid positions 302 to 304, preferably amino acid position 303 of SEQ ID NO: 9 are replaced by cysteines,
(ii) one of amino acid positions 290 to 292, preferably amino acid position 291 and one of amino acid positions 442 to 444, preferably amino acid position 443 of SEQ ID NO: 9 are replaced by cysteines,
(iii) one of amino acid positions 274 to 276, preferably amino acid position 275 and one of amino acid positions 385 to 387, preferably amino acid position 386 of SEQ ID NO: 9 are replaced by cysteines,
(iv) one of amino acid positions 311 to 313, preferably amino acid position 312 and one of amino acid positions 320 to 322, preferably amino acid position 321 of SEQ ID NO: 9 are replaced by cysteines,
(v) one of amino acid positions 268 to 270, preferably amino acid position 269 and one of amino acid positions 315 to 317, preferably amino acid position 316 of SEQ ID NO: 9 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(X)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG3 (Adhesion G Protein-Coupled Receptor G3) of SEQ ID NO: 10, wherein
(i) one of amino acid positions 16 to 19, preferably amino acid position 17 and one of amino acid positions 45 to 47, preferably amino acid position 46 of SEQ ID NO: 10 are replaced by cysteines,
(ii) one of amino acid positions 32 to 34, preferably amino acid position 33 and one of amino acid positions 116 to 118, preferably amino acid position 117 of SEQ ID NO: 10 are replaced by cysteines,
(iii) one of amino acid positions 55 to 57, preferably amino acid position 56 and one of amino acid positions 63 to 65, preferably amino acid position 64 of SEQ ID NO: 10 are replaced by cysteines,
(iv) one of amino acid positions 29 to 31 , preferably amino acid position 30 and one of amino acid positions 59 to 61 , preferably amino acid position 60 of SEQ ID NO: 10 are replaced by cysteines, or
(v) at least two of (i) to (iv) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (v) are retained;
(XI)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRG2
(Adhesion G Protein-Coupled Receptor G2) of SEQ ID NO: 11 , wherein
(i) one of amino acid positions 12 to 14, preferably amino acid position 13 and one of amino acid positions 63 to 65, preferably amino acid position 64 of SEQ ID NO: 11 are replaced by cysteines,
(ii) one of amino acid positions 51 to 53, preferably amino acid position 52 and one of amino acid positions 205 to 207, preferably amino acid position 206 of SEQ ID NO: 11 are replaced by cysteines,
(iii) one of amino acid positions 35 to 37, preferably amino acid position 36 and one of amino acid positions 147 to 149, preferably amino acid position 148 of SEQ ID NO: 11 are replaced by cysteines,
(iv) one of amino acid positions 72 to 74, preferably amino acid position 73 and one of amino acid positions 81 to 83, preferably amino acid position 82 of SEQ ID NO: 11 are replaced by cysteines,
(v) one of amino acid positions 32 to 34, preferably amino acid position 33 and one of amino acid positions 76 to 78, preferably amino acid position 77 of SEQ ID NO: 11 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(XII)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF5 (Adhesion G Protein-Coupled Receptor F5) of SEQ ID NO: 12, wherein
(i) one of amino acid positions 82 to 84, preferably amino acid position 83 and one of amino acid positions 117 to 119, preferably amino acid position 118 of SEQ ID NO: 12 are replaced by cysteines,
(ii) one of amino acid positions 105 to 107, preferably amino acid position 106 and one of amino acid positions 240 to 242, preferably amino acid position 241 of SEQ ID NO: 12 are replaced by cysteines,
(iii) one of amino acid positions 90 to 92, preferably amino acid position 91 and one of amino acid positions 186 to 188, preferably amino acid position 187 of SEQ ID NO: 12 are replaced by cysteines,
(iv) one of amino acid positions 127 to 129, preferably amino acid position 128 and one of amino acid positions 136 to 138, preferably amino acid position 137 of SEQ ID NO: 12 are replaced by cysteines,
(v) one of amino acid positions 88 to 90, preferably amino acid position 89 and one of amino acid positions 131 to 133, preferably amino acid position 132 of SEQ ID NO: 12 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence
identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (v) are retained;
(XIII)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF4 (Adhesion G Protein-Coupled Receptor F4) of SEQ ID NO: 13, wherein
(i) one of amino acid positions 53 to 55, preferably amino acid position 54 and one of amino acid positions 86 to 88, preferably amino acid position 87 of SEQ ID NO: 13 are replaced by cysteines,
(ii) one of amino acid positions 74 to 76, preferably amino acid position 75 and one of amino acid positions 216 to 218, preferably amino acid position 217 of SEQ ID NO: 13 are replaced by cysteines,
(iii) one of amino acid positions 61 to 63, preferably amino acid position 62 and one of amino acid positions 159 to 161 , preferably amino acid position 160 of SEQ ID NO: 13 are replaced by cysteines,
(iv) one of amino acid positions 96 to 98, preferably amino acid position 97 and one of amino acid positions 105 to 107, preferably amino acid position 106 of SEQ ID NO: 13 are replaced by cysteines,
(v) one of amino acid positions 59 to 61 , preferably amino acid position 60 and one of amino acid positions 100 to 102, preferably amino acid position 102 of SEQ ID NO: 13 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(XIV)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF3 (Adhesion G Protein-Coupled Receptor F3) of SEQ ID NO: 14, wherein
(i) one of amino acid positions 81 to 83, preferably amino acid position 82 and one of amino acid positions 116 to 118, preferably amino acid position 117 of SEQ ID NO: 14 are replaced by cysteines,
(ii) one of amino acid positions 104 to 106, preferably amino acid position 105 and one of amino acid positions 243 to 245, preferably amino acid position 244 of SEQ ID NO: 14 are replaced by cysteines,
(iii) one of amino acid positions 89 to 91 , preferably amino acid position 90 and one of amino acid positions 184 to 186, preferably amino acid position 185 of SEQ ID NO: 14 are replaced by cysteines,
(iv) one of amino acid positions 126 to 128, preferably amino acid position 127 and one of amino acid positions 135 to 137, preferably amino acid position 136 of SEQ ID NO: 14 are replaced by cysteines,
(v) one of amino acid positions 87 to 89, preferably amino acid position 88 and one of
amino acid positions 129 to 131 , preferably amino acid position 130 of SEQ ID NO: 14 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(XV)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF2 (Adhesion G Protein-Coupled Receptor F2) of SEQ ID NO: 15, wherein
(i) one of amino acid positions 73 to 75, preferably amino acid position 74 and one of amino acid positions 106 to 108, preferably amino acid position 107 of SEQ ID NO: 15 are replaced by cysteines,
(ii) one of amino acid positions 94 to 96, preferably amino acid position 95 and one of amino acid positions 230 to 232, preferably amino acid position 231 of SEQ ID NO: 15 are replaced by cysteines,
(iii) one of amino acid positions 81 to 83, preferably amino acid position 82 and one of amino acid positions 177 to 179, preferably amino acid position 178 of SEQ ID NO: 15 are replaced by cysteines,
(iv) one of amino acid positions 116 to 118, preferably amino acid position 117 and one of amino acid positions 124 to 126, preferably amino acid position 125 of SEQ ID NO: 15 are replaced by cysteines,
(v) one of amino acid positions 79 to 81 , preferably amino acid position 80 and one of amino acid positions 119 to 121 , preferably amino acid position 120 of SEQ ID NO: 15 are replaced by cysteines, or
(vi)at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(XVI)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRF1 (Adhesion G Protein-Coupled Receptor F1) of SEQ ID NO: 16, wherein
(i) one of amino acid positions 83 to 85, preferably amino acid position 84 and one of amino acid positions 118 to 120, preferably amino acid position 119 of SEQ ID NO: 16 are replaced by cysteines,
(ii) one of amino acid positions 106 to 108, preferably amino acid position 107 and one of amino acid positions 243 to 245, preferably amino acid position 244 of SEQ ID NO: 16 are replaced by cysteines,
(iii) one of amino acid positions 91 to 93, preferably amino acid position 92 and one of amino acid positions 190 to 192, preferably amino acid position 191 of SEQ ID NO: 16
are replaced by cysteines,
(iv) one of amino acid positions 128 to 130, preferably amino acid position 129 and one of amino acid positions 136 to 138, preferably amino acid position 137 of SEQ ID NO: 16 are replaced by cysteines,
(v) one of amino acid positions 89 to 91 , preferably amino acid position 90 and one of amino acid positions 131 to 133, preferably amino acid position 132 of SEQ ID NO: 16 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(XVII)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE5 (Adhesion G Protein-Coupled Receptor E5) of SEQ ID NO: 17, wherein
(i) one of amino acid positions 72 to 74, preferably amino acid position 73 and one of amino acid positions 244 to 246, preferably amino acid position 245 of SEQ ID NO: 17 are replaced by cysteines,
(ii) one of amino acid positions 96 to 98, preferably amino acid position 97 and one of amino acid positions 109 to 111 , preferably amino acid position 110 of SEQ ID NO: 17 are replaced by cysteines, or
(iii) (i) and (ii) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in (A), (i), (ii) or (iii) are retained;
(XVIII)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE3 (Adhesion G Protein-Coupled Receptor E3) of SEQ ID NO: 18, wherein one of amino acid positions 72 to 74, preferably amino acid position 73 and one of amino acid positions 83 to 85, preferably amino acid position 84 of SEQ ID NO: 18 are replaced by cysteines; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in (A) are retained;
(XIX)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE3 (Adhesion G Protein-Coupled Receptor E3) of SEQ ID NO: 19, wherein one of amino acid positions 73 to 75, preferably amino acid position 74 and one of amino acid positions 186 to 188, preferably amino acid position 187 of SEQ ID NO: 19 are replaced by cysteines; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in (A) are retained;
(XX)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE2 (Adhesion G Protein-Coupled Receptor E2) of SEQ ID NO: 20, wherein one of amino acid positions 72 to 74, preferably amino acid position 73 and one of amino acid positions 228 to 230, preferably amino acid position 229 of SEQ ID NO: 20 are replaced by cysteines,
(ii) one of amino acid positions 96 to 98, preferably amino acid position 97 and one of amino acid positions 109 to 111 , preferably amino acid position 110 of SEQ ID NO: 20 are replaced by cysteines, or
(iii) (i) and (ii) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in (A), (i), (ii) or (iii) are retained;
(XXI)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRE1 (Adhesion G Protein-Coupled Receptor E1) of SEQ ID NO: 21 , wherein
(i) one of amino acid positions 66 to 68, preferably amino acid position 67 and one of amino acid positions 103 to 105, preferably amino acid position 104 of SEQ ID NO: 21 are replaced by cysteines,
(ii) one of amino acid positions 91 to 93, preferably amino acid position 92 and one of amino acid positions 243 to 245, preferably amino acid position 244 of SEQ ID NO: 21 are replaced by cysteines,
(iii) one of amino acid positions 75 to 77, preferably amino acid position 76 and one of amino acid positions 173 to 175, preferably amino acid position 174 of SEQ ID NO: 21 are replaced by cysteines,
(iv) one of amino acid positions 111 to 113, preferably amino acid position 112 and one of amino acid positions 124 to 126, preferably amino acid position 125 of SEQ ID NO: 21 are replaced by cysteines,
(v) one of amino acid positions 73 to 75, preferably amino acid position 74 and one of amino acid positions 119 to 121 , preferably amino acid position 120 of SEQ ID NO: 21 are replaced by cysteines or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(XXII)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRD2 (Adhesion G Protein-Coupled Receptor D2) of SEQ ID NO: 22, wherein
(i) one of amino acid positions 67 to 69, preferably amino acid position 68 and one of amino acid positions 110 to 112, preferably amino acid position 111 of SEQ ID NO: 22
are replaced by cysteines,
(ii) one of amino acid positions 98 to 100, preferably amino acid position 99 and one of amino acid positions 235 to 237, preferably amino acid position 236 of SEQ ID NO: 22 are replaced by cysteines,
(iii) one of amino acid positions 121 to 123, preferably amino acid position 121 and one of amino acid positions 131 to 133, preferably amino acid position 132 of SEQ ID NO: 22 are replaced by cysteines,
(iv) one of amino acid positions 80 to 82, preferably amino acid position 81 and one of amino acid positions 127 to 129, preferably amino acid position 128 of SEQ ID NO: 22 are replaced by cysteines or
(v) at least two of (i) to (iv) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (v) are retained;
(XXIII)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRD1 (Adhesion G Protein-Coupled Receptor D1) of SEQ ID NO: 23, wherein
(i) one of amino acid positions 36 to 38, preferably amino acid position 37 and one of amino acid positions 92 to 94, preferably amino acid position 93 of SEQ ID NO: 23 are replaced by cysteines,
(ii) one of amino acid positions 80 to 82, preferably amino acid position 81 and one of amino acid positions 247 to 249, preferably amino acid position 248 of SEQ ID NO: 23 are replaced by cysteines,
(iii) one of amino acid positions 66 to 68, preferably amino acid position 67 and one of amino acid positions 162 to 164, preferably amino acid position 163 of SEQ ID NO: 23 are replaced by cysteines,
(iv) one of amino acid positions 104 to 106, preferably amino acid position 105 and one of amino acid positions 116 to 118, preferably amino acid position 117 of SEQ ID NO: 23 are replaced by cysteines,
(v) one of amino acid positions 64 to 66, preferably amino acid position 65 and one of amino acid positions 111 to 113, preferably amino acid position 112 of SEQ ID NO: 23 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(XXIV)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRB3 (Adhesion G Protein-Coupled Receptor B3) of SEQ ID NO: 24, wherein
(i) one of amino acid positions 87 to 89, preferably amino acid position 88 and one of
amino acid positions 121 to 123, preferably amino acid position 122 of SEQ ID NO: 24 are replaced by cysteines,
(ii) one of amino acid positions 109 to 111 , preferably amino acid position 110 and one of amino acid positions 261 to 263, preferably amino acid position 262 of SEQ ID NO: 24 are replaced by cysteines,
(iii) one of amino acid positions 95 to 97, preferably amino acid position 96 and one of amino acid positions 206 to 208, preferably amino acid position 207 of SEQ ID NO: 24 are replaced by cysteines,
(iv) one of amino acid positions 133 to 135, preferably amino acid position 134 and one of amino acid positions 142 to 144, preferably amino acid position 143 of SEQ ID NO: 24 are replaced by cysteines,
(v) one of amino acid positions 93 to 95, preferably amino acid position 94 and one of amino acid positions 138 to 140, preferably amino acid position 139 of SEQ ID NO: 24 are replaced by cysteines or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(XXV)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRB2 (Adhesion G Protein-Coupled Receptor B2) of SEQ ID NO: 25, wherein
(i) one of amino acid positions 86 to 88, preferably amino acid position 87 and one of amino acid positions 120 to 122, preferably amino acid position 121 of SEQ ID NO: 25 are replaced by cysteines,
(ii) one of amino acid positions 108 to 110, preferably amino acid position 109 and one of amino acid positions 295 to 297, preferably amino acid position 296 of SEQ ID NO: 25 are replaced by cysteines,
(iii) one of amino acid positions 94 to 96, preferably amino acid position 95 and one of amino acid positions 239 to 241 , preferably amino acid position 240 of SEQ ID NO: 25 are replaced by cysteines,
(iv) one of amino acid positions 132 to 134, preferably amino acid position 133 and one of amino acid positions 141 to 143, preferably amino acid position 142 of SEQ ID NO: 25 are replaced by cysteines,
(v) one of amino acid positions 92 to 94, preferably amino acid position 93 and one of amino acid positions 137 to 139, preferably amino acid position 138 of SEQ ID NO: 25 are replaced by cysteines or
(vi) at least two or all three of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(XXVI)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRB1 (Adhesion G Protein-Coupled Receptor B1) of SEQ ID NO: 26, wherein
(i) one of amino acid positions 81 to 83, preferably amino acid position 82 and one of amino acid positions 115 to 117, preferably amino acid position 116 of SEQ ID NO: 26 are replaced by cysteines,
(ii) one of amino acid positions 103 to 105, preferably amino acid position 104 and one of amino acid positions 253 to 255, preferably amino acid position 254 of SEQ ID NO: 26 are replaced by cysteines,
(iii) one of amino acid positions 89 to 91 , preferably amino acid position 90 and one of amino acid positions 198 to 200, preferably amino acid position 199 of SEQ ID NO: 26 are replaced by cysteines,
(iv) one of amino acid positions 128 to 130, preferably amino acid position 129 and one of amino acid positions 136 to 138, preferably amino acid position 137 of SEQ ID NO: 26 are replaced by cysteines,
(v) one of amino acid positions 87 to 89, preferably amino acid position 88 and one of amino acid positions 132 to 134, preferably amino acid position 133 of SEQ ID NO: 26 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(XXVII)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRA3 (Adhesion G Protein-Coupled Receptor A3) of SEQ ID NO: 27, wherein
(i) one of amino acid positions 57 to 59, preferably amino acid position 58 and one of amino acid positions 106 to 108, preferably amino acid position 107 of SEQ ID NO: 27 are replaced by cysteines,
(ii) one of amino acid positions 94 to 96, preferably amino acid position 95 and one of amino acid positions 288 to 290, preferably amino acid position 289 of SEQ ID NO: 27 are replaced by cysteines,
(iii) one of amino acid positions 81 to 83, preferably amino acid position 82 and one of amino acid positions 225 to 227, preferably amino acid position 226 of SEQ ID NO: 27 are replaced by cysteines,
(iv) one of amino acid positions 115 to 117, preferably amino acid position 116 and one of amino acid positions 125 to 127, preferably amino acid position 126 of SEQ ID NO: 27 are replaced by cysteines,
(v) one of amino acid positions 73 to 75, preferably amino acid position 74 and one of amino acid positions 120 to 122, preferably amino acid position 121 of SEQ ID NO: 27 are replaced by cysteines, or
(vi) at least two of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in any one of (A), items (i) to (vi) are retained;
(XXVIII)
(A) comprising or consisting of the amino acid sequence of the GAIN domain of ADGRA2 (Adhesion G Protein-Coupled Receptor A2) of SEQ ID NO: 28, wherein
(i) one of amino acid positions 57 to 59, preferably amino acid position 58 and one of amino acid positions 104 to 106, preferably amino acid position 105 of SEQ ID NO: 28 are replaced by cysteines,
(ii) one of amino acid positions 92 to 94, preferably amino acid position 93 and one of amino acid positions 294 to 296, preferably amino acid position 295 of SEQ ID NO: 28 are replaced by cysteines,
(iii) one of amino acid positions 79 to 81 , preferably amino acid position 80 and one of amino acid positions 229 to 231 , preferably amino acid position 230 of SEQ ID NO: 28 are replaced by cysteines,
(iv) one of amino acid positions 113 to 115, preferably amino acid position 114 and one of amino acid positions 123 to 125, preferably amino acid position 124 of SEQ ID NO: 28 are replaced by cysteines,
(v) one of amino acid positions 71 to 73, preferably amino acid position 72 and one of amino acid positions 118 to 120, preferably amino acid position 129 of SEQ ID NO: 28 are replaced by cysteines, or
(vi) at least two or all three of (i) to (v) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in (A), items (i) to (vi) are retained; or
(XXIX)
(A) comprising or consisting of the amino acid sequence of GAIN domain of ADGRV1 (Adhesion G Protein-Coupled Receptor V1) of SEQ ID NO: 29, wherein
(i) one of amino acid positions 171 to 173, preferably amino acid position 172 and one of amino acid positions 229 to 231 , preferably amino acid position 230 of SEQ ID NO: 29 are replaced by cysteines,
(ii) one of amino acid positions 127 to 219, preferably amino acid position 218 and one of amino acid positions 436 to 438, preferably amino acid position 437 of SEQ ID NO: 29 are replaced by cysteines,
(iii) one of amino acid positions 201 to 203, preferably amino acid position 202 and one of amino acid positions 377 to 379, preferably amino acid position 378 of SEQ ID NO: 29 are replaced by cysteines,
(iv) one of amino acid positions 199 to 201 , preferably amino acid position 200 and one of amino acid positions 239 to 241 , preferably amino acid position 240 of SEQ ID NO: 29
are replaced by cysteines, or
(v) at least two of (i) to (iv) apply; or
(B) comprising or consisting of an amino acid sequence sharing at least 90% sequence identity with the variant of (A), provided that the cysteines as defined in (A), items (i) to (v) are retained.
2. An adhesion G protein-coupled receptor (AGPCR) comprising the variant of the GAIN domain of claim 1 , wherein the AGPCR preferably comprises of consists of any of SEQ ID NOs 30 to 58 or a sequence being at least 90% identical thereto.
3. A nucleic acid molecule, preferably a vector encoding the variant of the GAIN domain of claim 1 or the AGPCR of claim 2.
4. A cell, preferably an epithelial cell or a lymphocyte and most preferably a T-cell comprising the nucleic acid molecule, preferably the vector of claim 3.
5. A molecular complex comprising the variant of the GAIN domain of claim 1 or the AGPCR of claim 2 in association with a Stachel sequence or fragment thereof being capable of binding to the GAIN domain of claim 1 , wherein the Stachel sequence preferably comprises or consists of any one of SEQ ID NOs 59 to 81 or a sequence being at least 70% identical, preferably at least 80% and most preferably at least 90% identical thereto.
6. A composition, preferably a diagnostic or pharmaceutical composition, or a kit comprising the variant of claim 1 or the AGPCR of claim 2.
7. A computer-implemented method for increasing or decreasing the association of the GAIN domain of an AGPCR and the Stachel sequence under force comprising
(a) performing in silico steered molecular dynamics (SMD) simulations on the GAIN domain in association with the Stachel sequence whereby the GAIN domain undergoes topological rearrangements that enable force to propagate through the GAIN domain to the associated Stachel sequence,
(b) identifying force propagation pathways from the SMD simulations of (a) by analysing the correlations between alpha carbon motions within the GAIN domain,
(c) identifying the arrival angles of the optimal force propagation pathways as obtained in (b) on the Stachel sequence,
(d) mutating in silico based on the information as obtained in (a) to (c) one or more pairs of amino acids of the GAIN domain outside of the interface of the GAIN domain that associates with the Stachel sequence to cysteine residues and/or substituting in silico based on the information as obtained in (a) to (c) one or more amino acids, preferably
between one and ten amino acids by other amino acids outside of the interface of the GAIN domain that associates with the Stachel sequence to obtain a variant of the GAIN domain with a disulfide bond that is expected to decrease or increase the arrival angles as compared to the arrival angles as identified in (c),
(e) repeating steps (a) to (c) on the variant of the GAIN domain as obtained in (d) thereby identifying the arrival angles of the optimal force propagation pathways on the Stachel sequence for the variant of the GAIN domain in (d), wherein increased arrival angles are indicative for an increased strength of the association of the GAIN domain and the Stachel sequence under force and decreased arrival angles are indicative for a decreased strength of the association of the GAIN domain and the Stachel sequence under force.
8. The method of claim 7, wherein in step (a) two or more different pulling speeds are applied to the GAIN domain, wherein the pulling speeds are preferably selected from about 0.025 nm.ns-1, 0.1 nm.ns-1 and about 1.0 nm.ns-1.
9. The method of claim 7 or 8, wherein in step (a) a harmonic force potential is applied to the GAIN domain by pulling simultaneously from the N- and C-terminus of the GAIN domain.
10. The method of any one of claims 7 to 9, wherein in step (b) for every frame the three-dimensional coordinates for every alpha carbon of the GAIN domain are extracted, for every pair of two consecutive frames of the simulation the distance displacements of the three-dimensional coordinates are obtained, and the distance displacements are correlated to reconstruct the optimal and suboptimal force propagation pathways from the N- to the C-terminus and computing force propagation pathways.
11. The method of any one of claims 7 to 10, wherein in step (c) the arrival angles are calculated with respect to a plane that is formed between the N-terminal and C-terminal alpha carbon of the Stachel sequence and a selected alpha carbon in the middle of the Stachel sequence.
12. The method of any one of claims 7 to 11 , further comprising producing the variant of the GAIN domain as obtained by any one of the methods 7 to 10.
13. The method of claim 12, further comprising applying single-molecule force spectrometry (SMFS) to the association of the produced variant of the GAIN domain and the Stachel sequence in order to identify the interface of variant of the GAIN domain that associates with the Stachel sequence and/or the force that disrupts the association, wherein the SMFS is preferably atomic force microscopy (AFM) based SMFS.
The method of claim 13, wherein for the SMFS the variant of the GAIN-domain is fused to a Fgp domain and a Spytag whereby the Fgp domain connects the variant of the GAIN domain via an isopeptide bond to a surface, preferably a coverglass surface and the Fgp domain connects the variant of the GAIN domain to a cantilever. The method of any one of claims 12 to 14, further comprising measuring the force that disrupts the association between the variant of the GAIN domain and the Stachel sequence, preferably by an serum response element (SRE) assay wherein shear stress, preferably through vibration or spinning is applied to cells that express an aGPCR comprising the variant of the GAIN domain of step (d), wherein the aGPCRs on the cells are coupled via aGPCR ligands, preferably collagen III or transglutaminase 2 onto a plate.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP22172182 | 2022-05-06 | ||
| EP22172182.2 | 2022-05-06 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2023213993A1 true WO2023213993A1 (en) | 2023-11-09 |
Family
ID=81585330
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/EP2023/061899 Ceased WO2023213993A1 (en) | 2022-05-06 | 2023-05-05 | Variants of the gain domain of adhesion g protein-coupled receptors |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2023213993A1 (en) |
-
2023
- 2023-05-05 WO PCT/EP2023/061899 patent/WO2023213993A1/en not_active Ceased
Non-Patent Citations (92)
| Title |
|---|
| A. F. OBERHAUSERC. BADILLA-FERNANDEZM. CARRION-VAZQUEZJ. M. FERNANDEZ: "The mechanical hierarchies of fibronectin observed with single-molecule AFM", JOURNAL OF MOLECULAR BIOLOGY, vol. 319, 2002, pages 433 - 447 |
| A. F. OBERHAUSERP. E. MARSZALEKH. P. ERICKSONJ. M. FERNANDEZ: "The molecular elasticity of the extracellular matrix protein tenascin", NATURE, vol. 393, 1998, pages 181 - 185 |
| A. S. OUDE VRIELINKT. D. VANCEA. M. DE JONGP. L. DAVIESI. K. VOETS: "Unusually high mechanical stability of bacterial adhesin extender domains having calcium clamps", PLOS ONE, vol. 12, 2017, pages e0174682 |
| A. VALBUENA ET AL.: "On the remarkable mechanostability of scaffoldins and the mechanical clamp motif", PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES, vol. 106, 2009, pages 13791 - 13796 |
| ALBERICIOGOVENDER, CHEMICAL PROTEIN AND PEPTIDE SYNTHESIS, 2012 |
| ALTSHULER EPSEREBRYANAYA DVKATRUKHA AG, BIOCHEMISTRY (MOSC)., vol. 75, no. 13, 2010, pages 1584 |
| B. YANGH. LIUZ. LIUR. DOENENM. A. NASH: "Influence of fluorination on single-molecule unfolding and rupture pathways of a mechanostable protein adhesion complex", NANO LETTERS, vol. 20, 2020, pages 8940 - 8950 |
| BARROS-ÁLVAREZ XIMENA ET AL: "The tethered peptide activation mechanism of adhesion GPCRs", NATURE, NATURE PUBLISHING GROUP UK, LONDON, vol. 604, no. 7907, 13 April 2022 (2022-04-13), pages 757 - 762, XP037805438, ISSN: 0028-0836, [retrieved on 20220413], DOI: 10.1038/S41586-022-04575-7 * |
| BASSILANA ET AL., NATURE REVIEWS DRUG DISCOVERY, vol. 18, 2019, pages 869 - 884 |
| BASSILANA FREDERIC ET AL: "Adhesion G protein-coupled receptors: opportunities for drug discovery", NATURE REVIEWS DRUG DISCOVERY, NATURE PUBLISHING GROUP, GB, vol. 18, no. 11, 28 August 2019 (2019-08-28), pages 869 - 884, XP036917468, ISSN: 1474-1776, [retrieved on 20190828], DOI: 10.1038/S41573-019-0039-Y * |
| BATEMAN ET AL., NUCLEIC ACIDS RESEARCH, vol. 30, no. 1, 2002, pages 276 - 280, Retrieved from the Internet <URL:http://smart.embl-heidelberg.de> |
| BEATA JASTRZEBSKAPAUL S.H. PARK, GPCRS: STRUCTURE, FUNCTION, AND DRUG DISCOVERY, 2019 |
| BELIU GERTI ET AL: "Tethered agonist exposure in intact adhesion/class B2 GPCRs through intrinsic structural flexibility of the GAIN domain", MOLECULAR CELL, vol. 81, no. 5, 1 March 2021 (2021-03-01), AMSTERDAM, NL, pages 905 - 921.e5, XP055975098, ISSN: 1097-2765, DOI: 10.1016/j.molcel.2020.12.042 * |
| C. SCHOELER ET AL.: "Mapping mechanical force propagation through biomolecular complexes", NANO LETTERS, vol. 15, 2015, pages 7370 - 7376 |
| C. WILDE ET AL.: "The constitutive activity of the adhesion GPCR GPR114/ADGRG5 is mediated by its tethered agonist", FASEB J, vol. 30, 2016, pages 666 - 673 |
| C.-L. CHYAN ET AL.: "Reversible mechanical unfolding of single ubiquitin molecules", BIOPHYSICAL JOURNAL, vol. 87, 2004, pages 3995 - 4006 |
| D. CRAIGM. GAOK. SCHULTENV. VOGEL: "Tuning the mechanical stability of fibronectin type III modules through sequence variations", STRUCTURE, vol. 12, 2004, pages 21 - 30 |
| D. K. WESTD. J. BROCKWELLP. D. OLMSTEDS. E. RADFORDE. PACI: "Mechanical resistance of proteins explained using simple molecular models", BIOPHYSICAL JOURNAL, vol. 90, 2006, pages 287 - 297 |
| D. L. GUZMANA. RANDALLP. BALDIZ. GUAN: "Computational and single-molecule force studies of a macro domain protein reveal a key molecular determinant for mechanical stability", PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES, vol. 107, 2010, pages 1989 - 1994 |
| D. P. SADLER ET AL.: "Identification of a mechanical rheostat in the hydrophobic core of protein L", JOURNAL OF MOLECULAR BIOLOGY, vol. 393, 2009, pages 237 - 248, XP026796885, DOI: 10.1016/j.jmb.2009.08.015 |
| DAOLAI ZHANG ET AL: "Function and therapeutic potential of G protein-coupled receptors in epididymis", BRITISH JOURNAL OF PHARMACOLOGY, WILEY-BLACKWELL, UK, vol. 177, no. 24, 29 October 2020 (2020-10-29), pages 5489 - 5508, XP071026690, ISSN: 0007-1188, DOI: 10.1111/BPH.15252 * |
| DUMAS L. ET AL: "Uncovering and engineering the mechanical properties of the adhesion GPCR ADGRG1 GAIN domain", BIORXIV, 6 April 2023 (2023-04-06) - 6 March 2023 (2023-03-06), pages 1 - 28, XP093061403, Retrieved from the Internet <URL:https://www.biorxiv.org/content/10.1101/2023.04.05.535724v1.full.pdf> [retrieved on 20230705], DOI: 10.1101/2023.04.05.535724 * |
| E. EVANSK. RITCHIE: "Dynamic strength of molecular adhesion bonds", BIOPHYSICAL JOURNAL, vol. 72, 1997, pages 1541 - 1555, XP000989495 |
| E. W. DIJKSTRA: "A short introduction to the art of programming", vol. 4, 1971, TECHNISCHE HOGESCHOOL EINDHOVEN |
| F. LIS. D. REDICKH. P. ERICKSONV. T. MOY: "Force measurements of the a5p1 integrin-fibronectin interaction", BIOPHYSICAL JOURNAL, vol. 84, 2003, pages 1252 - 1262 |
| H. H. LIN ET AL.: "Autocatalytic cleavage of the EMR2 receptor occurs at a conserved G protein-coupled receptor proteolytic site motif.", J BIOL CHEM, vol. 279, 2004, pages 31823 - 31832 |
| H. LI ET AL.: "Reverse engineering of the giant muscle protein titin", NATURE, vol. 418, 2002, pages 998 - 1002 |
| H. LIM. CARRION-VAZQUEZA. F. OBERHAUSERP. E. MARSZALEKJ. M. FERNANDEZ: "Point mutations alter the mechanical stability of immunoglobulin modules", NATURE STRUCTURAL BIOLOGY, vol. 7, 2000, pages 1117 - 1120 |
| H. LUB. ISRALEWITZA. KRAMMERV. VOGELK. SCHULTEN: "Unfolding of titin immunoglobulin domains by steered molecular dynamics simulation", BIOPHYSICAL JOURNAL, vol. 75, 1998, pages 662 - 671 |
| H. M. STOVEKENA. G. HAJDUCZOKL. XUG. G. TALL: "Adhesion G protein-coupled receptors are activated by exposure of a cryptic tethered agonist", PROC NATL ACAD SCI U S A, vol. 112, 2015, pages 6194 - 6199 |
| HAMANN ET AL., PHARMACOLOGICAL REVIEWS, vol. 67, no. 2, 2015, pages 338 - 367 |
| HOLLIGER PHUDSON PJ, NAT BIOTECHNOL., vol. 23, no. 9, 2005, pages 1126 |
| I. LIEBSCHER ET AL.: "A tethered agonist within the ectodomain activates the adhesion G protein-coupled receptors GPR126 and GPR133", CELL REP, vol. 9, 2014, pages 2018 - 2026, XP055974541, DOI: 10.1016/j.celrep.2014.11.036 |
| IZRAILEV ET AL., COMPUTATIONAL MOLECULAR DYNAMICS: CHALLENGES, METHODS, IDEAS, 1999, pages 198 - 65 |
| J. B. KLAUDA ET AL.: "Update of the CHARMM all-atom additive force field for lipids: validation on six lipid types", THE JOURNAL OF PHYSICAL CHEMISTRY, vol. 8, no. 114, 2010, pages 7830 - 7843 |
| J. FRITZA. G. KATOPODISF. KOLBINGERD. ANSELMETTI: "Force-mediated kinetics of single P-selectin/ligand complexes observed by atomic force microscopy", PROCEEDINGS OF THE NATIONAL, vol. 95, 1998, pages 12283 - 12288 |
| J. P. JUNKERF. ZIEGLERM. RIEF: "Ligand-dependent equilibrium fluctuations of single calmodulin molecules", SCIENCE, vol. 323, 2009, pages 633 - 637 |
| J. P. WHITE ET AL.: "G protein-coupled receptor 56 regulates mechanical overload-induced muscle hypertrophy", PROC NATL ACAD SCI U S A, vol. 111, 2014, pages 15756 - 15761 |
| J. SEPPALA ET AL.: "Skeletal dysplasia mutations effect on human filamins' structure and mechanosensing", SCIENTIFIC REPORTS, vol. 7, 2017, pages 1 - 14 |
| J. X. GRAY ET AL.: "CD97 is a processed, seven-transmembrane, heterodimeric receptor associated with inflammation", JOURNAL OF IMMUNOLOGY, vol. 157, 1996, pages 5438 - 5447, XP002061690 |
| KIMPLE ET AL., CURR PROTOC PROTEIN SCI., vol. 73, 2013 |
| KONTERMANNBRINKMANN, DRUG DISCOVERY TODAY, vol. 20, no. 7, 2015, pages 838 - 847 |
| L. JUJ.-F. DONGM. A. CRUZC. ZHU: "The N-terminal flanking region of the A1 domain regulates the force-dependent binding of von Willebrand factor to platelet glycoprotein Iba", JOURNAL OF, vol. 288, 2013, pages 32289 - 32301 |
| LETUNIC ET AL., NUCLEIC ACIDS RES., vol. 30, no. 1, 2002, pages 242 - 244 |
| LIANE FISCHER ET AL: "Functional relevance of naturally occurring mutations in adhesion G protein-coupled receptor ADGRD1 (GPR133)", BMC GENOMICS, BIOMED CENTRAL LTD, LONDON, UK, vol. 17, no. 1, 11 August 2016 (2016-08-11), pages 1 - 9, XP021266249, DOI: 10.1186/S12864-016-2937-2 * |
| LIEBSCHER ET AL., CELL REPORTS, vol. 9, no. 6, 2014, pages 2018 - 2026 |
| LIEBSCHER ET AL., ONCOTARGET, vol. 6, no. 27, 2015 |
| LIEBSCHER INES ET AL: "A Tethered Agonist within the Ectodomain Activates the Adhesion G Protein-Coupled Receptors GPR126 and GPR133", vol. 9, no. 6, 1 December 2014 (2014-12-01), US, pages 2018 - 2026, XP055974541, ISSN: 2211-1247, Retrieved from the Internet <URL:https://www.cell.com/cell-reports/pdf/S2211-1247(14)01002-X.pdf> DOI: 10.1016/j.celrep.2014.11.036 * |
| M. CARRION-VAZQUEZ ET AL.: "Mechanical and chemical unfolding of a single protein: a comparison", PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES, vol. 96, 1999, pages 3694 - 3699, XP055943968, DOI: 10.1073/pnas.96.7.3694 |
| M. CARRION-VAZQUEZ ET AL.: "Mechanical design of proteins studied by single-molecule force spectroscopy and protein engineering", PROGRESS IN BIOPHYSICS AND MOLECULAR BIOLOGY, vol. 74, 2000, pages 63 - 91 |
| M. CARRION-VAZQUEZ ET AL.: "The mechanical stability of ubiquitin is linkage dependent", NATURE STRUCTURAL & MOLECULAR BIOLOGY, vol. 10, 2003, pages 738 - 743 |
| M. RIEFJ. PASCUALM. SARASTEH. E. GAUB: "Single molecule force spectroscopy of spectrin repeats: low unfolding forces in helix bundles", JOURNAL OF MOLECULAR BIOLOGY, vol. 286, 1999, pages 553 - 561, XP004459771, DOI: 10.1006/jmbi.1998.2466 |
| M. RIEFM. GAUTELF. OESTERHELTJ. M. FERNANDEZH. E. GAUB: "Reversible unfolding of individual titin immunoglobulin domains by AFM", SCIENCE, vol. 276, 1997, pages 1109 - 1112 |
| MASER ROBIN L ET AL: "Adhesion GPCRs as a paradigm for understanding polycystin-1 G protein regulation", CELLULAR SIGNALLING, ELSEVIER SCIENCE LTD, GB, vol. 72, 16 April 2020 (2020-04-16), XP086163807, ISSN: 0898-6568, [retrieved on 20200416], DOI: 10.1016/J.CELLSIG.2020.109637 * |
| MULDER ET AL., NUCL. ACIDS. RES., vol. 31, 2003, pages 315 - 318, Retrieved from the Internet <URL:http://www.sanger.ac.uk/Software/Pfam> |
| N. OLLIKAINENR. M. DE JONGT. KORTEMME: "Coupling protein side-chain and backbone flexibility improves the re-design of protein-ligand specificity", PLOS COMPUTATIONAL BIOLOGY, vol. 11, 2015, pages e1004335 |
| N. SCHOLZ ET AL.: "Mechano-dependent signaling by Latrophilin/CIRL quenches cAMP in proprioceptive neurons", ELIFE, vol. 6, 2017, pages e28360 |
| N. SCHOLZ ET AL.: "The adhesion GPCR latrophilin/CIRL shapes mechanosensation", CELL REP, vol. 11, 2015, pages 866 - 874 |
| NEUMANNAGY, NATURE METHODS, vol. 5, 2008, pages 491 - 505 |
| O. K. DUDKO: "Decoding the mechanical fingerprints of biomolecules", QUARTERLY REVIEWS OF BIOPHYSICS, 2016, pages 49 |
| O. K. DUDKOG. HUMMERA. SZABO: "Theory, analysis, and interpretation of single-molecule force spectroscopy experiments", PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES, vol. 105, 2008, pages 15755 - 15760 |
| O. N. KARPUS ET AL.: "Shear stress-dependent downregulation of the adhesion-G protein-coupled receptor CD97 on circulating leukocytes upon contact with its ligand CD55", J IMMUNOL, vol. 190, 2013, pages 3740 - 3748 |
| OWENS, PROC. NATL. ACAD. SCI. USA, vol. 98, 2001, pages 1471 - 1476 |
| P. E. MARSZALEK ET AL.: "Mechanical unfolding intermediates in titin modules", NATURE, vol. 402, 1999, pages 100 - 103 |
| P. MARKL. NILSSON: "Structure and dynamics of the TIP3P, SPC, and SPC/E water models at 298 K", THE JOURNAL OF PHYSICAL CHEMISTRY A, vol. 105, 2001, pages 9954 - 9960 |
| P. WERNET ET AL.: "The structure of the first coordination shell in liquid water", SCIENCE, vol. 304, 2004, pages 995 - 999, XP002431041, DOI: 10.1126/science.1096205 |
| PING YU-QI ET AL: "Structural basis for the tethered peptide activation of adhesion GPCRs", NATURE, NATURE PUBLISHING GROUP UK, LONDON, vol. 604, no. 7907, 13 April 2022 (2022-04-13), pages 763 - 770, XP037805453, ISSN: 0028-0836, [retrieved on 20220413], DOI: 10.1038/S41586-022-04619-Y * |
| QU XIANGLI ET AL: "Structural basis of tethered agonism of the adhesion GPCRs ADGRD1 and ADGRF1", NATURE, NATURE PUBLISHING GROUP UK, LONDON, vol. 604, no. 7907, 13 April 2022 (2022-04-13), pages 779 - 785, XP037805436, ISSN: 0028-0836, [retrieved on 20220413], DOI: 10.1038/S41586-022-04580-W * |
| R. LUO ET AL.: "Mechanism for adhesion G protein-coupled receptor GPR56-mediated RhoA activation induced by collagen III stimulation", PLOS ONE, vol. 9, 2014, pages e100043 |
| R. MERKELP. NASSOYA. LEUNGK. RITCHIEE. EVANS: "Energy landscapes of receptor-ligand bonds explored with dynamic force spectroscopy", NATURE, vol. 397, 1999, pages 50 - 53 |
| S. C. PETERSEN ET AL.: "The adhesion GPCR GPR126 has distinct, domain-dependent functions in Schwann cell development mediated by interaction with laminin-211", NEURON, vol. 85, 2015, pages 755 - 769, XP029198958, DOI: 10.1016/j.neuron.2014.12.057 |
| S. E. BOYDEN ET AL.: "Vibratory Urticaria Associated with a Missense Variant in ADGRE2", N ENGL J MED, vol. 374, 2016, pages 656 - 663 |
| S. P. NG ET AL.: ", Designing an extracellular matrix protein with enhanced mechanical stability", PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES, vol. 104, 2007, pages 9633 - 9637 |
| S. P. NG ET AL.: "Mechanical unfolding of TNfn3: the unfolding pathway of a fnlll domain probed by protein engineering, AFM and MD simulation", JOURNAL OF MOLECULAR BIOLOGY, vol. 350, 2005, pages 776 - 789, XP004951615, DOI: 10.1016/j.jmb.2005.04.070 |
| S. R. K. AINAVARAPU ET AL.: ", Contour length and refolding rate of a small protein controlled by engineered disulfide bonds", BIOPHYSICAL JOURNAL, vol. 92, 2007, pages 225 - 233 |
| SPECA DAVID J ET AL: "Whole exome sequencing reveals a functional mutation in the GAIN domain of the Bai2 receptor underlying a forward mutagenesis hyperactivity QTL", MAMMALIAN GENOME, SPRINGER NEW YORK LLC, US, vol. 28, no. 11, 12 September 2017 (2017-09-12), pages 465 - 475, XP036357537, ISSN: 0938-8990, [retrieved on 20170912], DOI: 10.1007/S00335-017-9716-5 * |
| STEPHEN F. ALTSCHULTHOMAS L. MADDENALEJANDRO A. SCHAFFERJINGHUI ZHANGZHENG ZHANGWEBB MILLERDAVID J. LIPMAN, NUCLEIC ACIDS RES., vol. 25, 1997, pages 3389 - 3402 |
| T. YAGO ET AL.: "Platelet glycoprotein Iba forms catch bonds with human WT vWF but not with type 2B von Willebrand disease vWF", THE JOURNAL OF CLINICAL INVESTIGATION, vol. 118, 2008, pages 3195 - 3207 |
| U. ESSMANN ET AL.: "A smooth particle mesh Ewald method", THE JOURNAL OF CHEMICAL PHYSICS, vol. 103, 1995, pages 8577 - 8593, XP055441882, DOI: 10.1063/1.470117 |
| V. C. LUCA ET AL.: "Notch-Jagged complex structure implicates a catch bond in tuning ligand sensitivity", SCIENCE, vol. 355, 2017, pages 1320 - 1324 |
| VETRINI FRANCESCO ET AL: "Bi-allelic Mutations inPKD1L1Are Associated with Laterality Defects in Humans", THE AMERICAN JOURNAL OF HUMAN GENETICS, AMERICAN SOCIETY OF HUMAN GENETICS , CHICAGO , IL, US, vol. 99, no. 4, 8 September 2016 (2016-09-08), pages 886 - 893, XP029761017, ISSN: 0002-9297, DOI: 10.1016/J.AJHG.2016.07.011 * |
| VIZURRAGA ALEXANDER ET AL: "Mechanisms of adhesion G protein?coupled receptor activation", JOURNAL OF BIOLOGICAL CHEMISTRY, vol. 295, no. 41, 6 August 2020 (2020-08-06), US, pages 14065 - 14083, XP055777502, ISSN: 0021-9258, Retrieved from the Internet <URL:http://dx.doi.org/10.1074/jbc.REV120.007423> DOI: 10.1074/jbc.REV120.007423 * |
| W. BAUMGARTNER ET AL.: "Cadherin interaction probed by atomic force microscopy", PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES, vol. 97, 2000, pages 4005 - 4010, XP055175231, DOI: 10.1073/pnas.070052697 |
| W. HANLEY ET AL.: "Single molecule characterization of P-selectin/ligand binding", JOURNAL OF BIOLOGICAL CHEMISTRY, vol. 278, 2003, pages 10556 - 10561 |
| W. R. GORDON ET AL.: "Mechanical allostery: evidence for a force requirement in the proteolytic activation of Notch", DEVELOPMENTAL CELL, vol. 33, 2015, pages 729 - 736, XP029179656, DOI: 10.1016/j.devcel.2015.05.004 |
| W. STACKLIESF. GRAETER: "Force Propagation in Proteins From Molecular Dynamics Simulations", BIOPHYSICAL JOURNAL, vol. 96, 2009 |
| XIAO PENG ET AL: "Tethered peptide activation mechanism of the adhesion GPCRs ADGRG2 and ADGRG4", NATURE, NATURE PUBLISHING GROUP UK, LONDON, vol. 604, no. 7907, 13 April 2022 (2022-04-13), pages 771 - 778, XP037805439, ISSN: 0028-0836, [retrieved on 20220413], DOI: 10.1038/S41586-022-04590-8 * |
| XIAO, SCIENTIFIC REPORTS, vol. 9, 2019 |
| YANG ET AL., FRONT MOL BIOSCI, vol. 7, 2020, pages 85 |
| YING LIU ET AL: "G protein-coupled receptors as promising cancer targets", CANCER LETTERS, vol. 376, no. 2, 1 July 2016 (2016-07-01), US, pages 226 - 239, XP055448597, ISSN: 0304-3835, DOI: 10.1016/j.canlet.2016.03.031 * |
| Z. LIU ET AL.: "Mapping mechanostable pulling geometries of a therapeutic anticalin/CTLA-4 protein complex", NANO LETTERS, 2021 |
| ZHAO ET AL., FRONT. IMMUNOL., 2021, Retrieved from the Internet <URL:https://doi.org/10.3389/fimmu.2021.658753> |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| KR102499753B1 (en) | Engineering t-cell receptors | |
| Saotome et al. | Structural analysis of cancer-relevant TCR-CD3 and peptide-MHC complexes by cryoEM | |
| Sumida et al. | Structural basis of leader peptide recognition in lasso peptide biosynthesis pathway | |
| US8741814B2 (en) | T cell receptor display | |
| Hickman et al. | Structural basis of hAT transposon end recognition by Hermes, an octameric DNA transposase from Musca domestica | |
| Moi et al. | Discovery of archaeal fusexins homologous to eukaryotic HAP2/GCS1 gamete fusion proteins | |
| Feng et al. | Fusion surface structure, function, and dynamics of gamete fusogen HAP2 | |
| Madura et al. | TCR‐induced alteration of primary MHC peptide anchor residue | |
| Onesti et al. | Structural studies of lysyl-tRNA synthetase: conformational changes induced by substrate binding | |
| Jones et al. | Engineering and characterization of a stabilized α1/α2 module of the class I major histocompatibility complex product Ld | |
| Lluis et al. | Protein engineering methods applied to membrane protein targets | |
| CN115461068A (en) | Identification of biomimetic viral peptides and uses thereof | |
| Barber et al. | Structure‐guided stabilization of pathogen‐derived peptide‐HLA‐E complexes using non‐natural amino acids conserves native TCR recognition | |
| US12134647B2 (en) | Binding molecules | |
| AU2017212788B2 (en) | G proteins | |
| WO2006023248A2 (en) | Processes for producing and crystallizing g-protein coupled receptors | |
| Ma et al. | Crystal structure and functional analysis of Drosophila Wind, a protein-disulfide isomerase-related protein | |
| Ebersbach et al. | Antigen generation and display in therapeutic antibody drug discovery–a neglected but critical player | |
| US20250210128A1 (en) | Computer-implemented design of peptide:receptor signaling complexes for enhanced chemotaxis | |
| JP2023116826A (en) | antigen binding molecule | |
| US20220411472A1 (en) | Self-assembling circular tandem repeat proteins with increased stability | |
| Sjöström et al. | Tuning the binding interface between Machupo virus glycoprotein and human transferrin receptor | |
| Tanaka et al. | Structural insights into biased signaling at chemokine receptor CCR7 | |
| Du et al. | A general platform for targeting MHC-II antigens via a single loop | |
| Ganguly et al. | Molecular basis of promiscuous chemokine-engagement by the Duffy antigen receptor |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23725188 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 23725188 Country of ref document: EP Kind code of ref document: A1 |




