WO2019196300A1 - 电子装置、鼻咽癌筛查分析方法和计算机可读存储介质 - Google Patents

电子装置、鼻咽癌筛查分析方法和计算机可读存储介质 Download PDF

Info

Publication number
WO2019196300A1
WO2019196300A1 PCT/CN2018/102108 CN2018102108W WO2019196300A1 WO 2019196300 A1 WO2019196300 A1 WO 2019196300A1 CN 2018102108 W CN2018102108 W CN 2018102108W WO 2019196300 A1 WO2019196300 A1 WO 2019196300A1
Authority
WO
WIPO (PCT)
Prior art keywords
feature
customer
tag
various
nasopharyngeal cancer
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2018/102108
Other languages
English (en)
French (fr)
Inventor
卢少烽
洪博然
徐亮
阮晓雯
肖京
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Ping An Technology Shenzhen Co Ltd
Original Assignee
Ping An Technology Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Ping An Technology Shenzhen Co Ltd filed Critical Ping An Technology Shenzhen Co Ltd
Priority to SG11202008389RA priority Critical patent/SG11202008389RA/en
Publication of WO2019196300A1 publication Critical patent/WO2019196300A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H50/00ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
    • G16H50/20ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for computer-aided diagnosis, e.g. based on medical expert systems
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H50/00ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
    • G16H50/30ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for calculating health indices; for individual health risk assessment
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H50/00ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
    • G16H50/70ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for mining of medical data, e.g. analysing previous cases of other patients

Definitions

  • the present application relates to the field of big data analysis application technologies, and in particular, to an electronic device, a nasopharyngeal cancer screening analysis method, and a computer readable storage medium.
  • Nasopharyngeal carcinoma is a high-grade malignant tumor with the highest incidence of otolaryngology and malignant tumors.
  • the etiology is complicated.
  • the cause of nasopharyngeal carcinoma in traditional medicine is still unclear.
  • Most of the existing researches come from clinical observations (related to heredity, environment, and viruses). However, clinical observations have limited access to personal information.
  • existing research usually relies on the professional medical knowledge and personal experience of the researchers. Therefore, the accuracy of the research cannot meet the requirements, and it is impossible to form an objective and accurate screening system for the cause of the disease. .
  • the present invention provides an electronic device and a nasopharyngeal cancer screening analysis method, which aims to achieve objective, accurate and efficient screening of high-risk nasopharyngeal cancer medical judgment objects, thereby enabling customers to prevent or treat them early.
  • a first aspect of the present application provides an electronic device including a memory and a processor, wherein the memory stores a nasopharyngeal cancer screening analysis system operable on the processor, the nasopharyngeal cancer screening analysis system is The processor implements the following steps when executed:
  • nasopharyngeal cancer screening request After receiving a nasopharyngeal cancer screening request from a client to be screened, obtaining a characteristic label of the customer to be screened from the nasopharyngeal cancer screening request;
  • the feature tag fails to be obtained from the nasopharyngeal cancer screening request, obtaining the customer attribute data of the to-screen customer from the nasopharyngeal cancer screening request;
  • the characteristic tag of the customer to be screened is substituted into the pre-trained analysis model to analyze whether the customer to be screened is a medical object of high risk nasopharyngeal cancer.
  • the feature tag includes basic feature information, preference habit information, behavior information, and/or social relationship information.
  • a second aspect of the present application provides a method for screening and analyzing nasopharyngeal cancer, the method comprising the steps of:
  • nasopharyngeal cancer screening request After receiving a nasopharyngeal cancer screening request from a client to be screened, obtaining a characteristic label of the customer to be screened from the nasopharyngeal cancer screening request;
  • the feature tag fails to be obtained from the nasopharyngeal cancer screening request, obtaining the customer attribute data of the to-screen customer from the nasopharyngeal cancer screening request;
  • the characteristic tag of the customer to be screened is substituted into the pre-trained analysis model to analyze whether the customer to be screened is a medical object of high risk nasopharyngeal cancer.
  • a third aspect of the present application provides a computer readable storage medium storing a nasopharyngeal cancer screening analysis system, the nasopharyngeal cancer screening analysis system being executable by at least one processor such that The at least one processor performs the following steps:
  • nasopharyngeal cancer screening request After receiving a nasopharyngeal cancer screening request from a client to be screened, obtaining a characteristic label of the customer to be screened from the nasopharyngeal cancer screening request;
  • the feature tag fails to be obtained from the nasopharyngeal cancer screening request, obtaining the customer attribute data of the to-screen customer from the nasopharyngeal cancer screening request;
  • the characteristic tag of the customer to be screened is substituted into the pre-trained analysis model to analyze whether the customer to be screened is a medical object of high risk nasopharyngeal cancer.
  • the technical solution of the present application after receiving the nasopharyngeal cancer screening request of the client to be screened, obtains the characteristic label of the customer to be screened according to the nasopharyngeal cancer screening request, specifically screening directly from the nasopharyngeal cancer Obtaining a feature tag of the to-be-screened client in the request, and when the acquisition fails, acquiring various types of the customer attribute information from each predetermined service server according to the customer attribute information in the nasopharyngeal cancer screening request Feature data, and extracting feature tags of the customer to be screened from the feature data according to a preset extraction rule; after obtaining the feature tag, inputting the obtained feature tag into the pre-trained analysis model, thereby analyzing Whether the client to be screened is a medical judgment target for high-risk nasopharyngeal cancer.
  • the present application analyzes the characteristic tags of the various characteristic data to be screened by the customer into the pre-trained analysis model to analyze whether the customer to be screened is a high-risk nasopharyngeal medical judgment object, and the customer is accurate.
  • the high-efficiency screening of high-risk nasopharyngeal cancer medical judgment objects allows customers to prevent or treat them early.
  • FIG. 1 is a schematic flow chart of an embodiment of a method for screening and analyzing nasopharyngeal cancer according to the present application
  • FIG. 2 is a schematic diagram of a training process of an analysis model in a nasopharyngeal cancer screening analysis method according to the present application;
  • FIG. 3 is a schematic diagram of an operating environment of an embodiment of a nasopharyngeal cancer screening and analysis system of the present application;
  • FIG. 4 is a block diagram showing the program of an embodiment of the nasopharyngeal carcinoma screening analysis system of the present application.
  • the present application proposes a method for screening and analyzing nasopharyngeal cancer.
  • FIG. 1 is a schematic flow chart of an embodiment of a method for screening and analyzing nasopharyngeal cancer according to the present application.
  • the nasopharyngeal cancer screening analysis method comprises:
  • Step S10 after receiving a nasopharyngeal cancer screening request of the client to be screened, obtaining a feature label of the customer to be screened from the nasopharyngeal cancer screening request;
  • the nasopharyngeal cancer screening request includes customer attribute data (eg, a certificate number, or a name and a document number) of the customer to be screened and/or a feature tag of the customer to be screened; wherein the feature tag includes basic feature information (eg, The city belongs to the northwestern part of China, the eating habits of the cities in which it belongs are spicy, the PM2.5 (fine particles) of the city is often exceeded, the bad weather in the city is too high, and the harmful gas positions are engaged, etc., and the preference information is preferred (for example, prefer to sleep late) , preference for tobacco and alcohol, preference for grilled food, preference for outdoor sports, etc.), behavioral information (for example, more games, more overtime, more take-outs, more medical visits, etc.) and/or social relationship information (for example, unmarried, Living alone, not always in contact with friends, etc.).
  • the screening server After receiving the nasopharyngeal cancer screening request from the client to be screened, the screening server first checks whether the
  • Step S20 if the feature tag fails to be obtained from the nasopharyngeal cancer screening request, obtaining the customer attribute data of the to-screen customer from the nasopharyngeal cancer screening request;
  • the nasopharyngeal cancer screening request does not include the customer's feature tag
  • the feature tag is not obtained from the nasopharyngeal cancer screening request (ie, the feature tag is failed to be acquired)
  • the nasopharyngeal cancer screening is performed.
  • the customer attribute data of the to-screen customer is obtained from the request, to obtain the feature label of the customer to be screened according to the customer attribute data of the screening customer.
  • Step S30 extracting, from a plurality of predetermined service servers, various feature data corresponding to the customer attribute data of the to-screen customer;
  • the screening server communicates with a plurality of predetermined business servers (eg, a bank server, a medical server, an insurance server, an instant messaging server, a game server, a weather server, a takeaway server, and/or a resume server, etc.); After the customer attribute data of the customer is to be screened, the screening server extracts various characteristic data corresponding to the customer attribute data of the to-screen customer from a plurality of predetermined business servers (for example, bank loan amount and repayment information) , outpatient medical record information "for example, the number of visits in a preset time, the type of disease, the duration of each illness, etc.”, insurance information "for example, the industry, gender, age, marital status, occupation, etc.”
  • Information about the use of the instant messaging tool account for example, information such as the communication tool daily login time information, daily online duration, etc.”, game information "for example, daily game login time information, daily game online duration, etc.”, weather information "for example , in the last three years, the number of days in which PM2.5 (fine
  • Step S40 Perform feature tag analysis on the extracted feature data according to a preset feature tag extraction rule to analyze the feature tag of the to-screen customer;
  • the extraction rule of the feature tag is preset in the screening server, and after extracting various feature data corresponding to the to-screen customer, the extraction rule is used to perform feature tag analysis on the extracted feature data, thereby analyzing The feature tag of the customer to be screened.
  • step S50 the feature tag of the customer to be screened is substituted into the pre-trained analysis model to analyze whether the customer to be screened is a high-risk nasopharyngeal cancer medical judgment object.
  • the screening server has a pre-trained analysis model, that is, the analysis model is trained through a large number of known customer data (including whether the customer has nasopharyngeal cancer and the corresponding feature tag of the customer); the screening server is currently waiting After screening the customer's feature label, the obtained feature label is input into the pre-trained analysis model, and the analysis model analyzes whether the customer to be screened is a high-risk nasopharyngeal medical judgment object according to the input characteristic label.
  • the technical solution of the embodiment after receiving the nasopharyngeal cancer screening request of the client to be screened, according to the nasopharyngeal cancer screening request, obtaining the characteristic label of the customer to be screened, specifically directly screening from the nasopharyngeal cancer Obtaining a feature tag of the to-screen customer in the request, and, when the acquisition fails, acquiring, according to the customer attribute information in the nasopharyngeal cancer screening request, each of the customer attribute information corresponding to each of the predetermined service servers
  • the feature data is extracted from the feature data according to the preset extraction rule, and the feature tag of the to-screen customer is extracted after the feature tag is obtained, and then the obtained feature tag is input into the pre-trained analysis model, thereby analyzing Whether the client to be screened is a medical judgment target for high-risk nasopharyngeal cancer.
  • the feature tags representing the various feature data of the customer to be screened are input into the pre-trained analysis model for analysis, to analyze whether the customer to be screened is a high-risk nasopharyngeal medical judgment object, and the client is realized. Accurate and efficient screening of high-risk nasopharyngeal cancer medical judgment objects, allowing customers to prevent or treat early.
  • FIG. 2 is a schematic diagram of a training process of an analysis model in the nasopharyngeal cancer screening analysis method of the present application.
  • the analysis model is a multiple linear regression model
  • the training process of the analysis model includes:
  • Step E1 extract, according to the predetermined customer attribute data, various feature data of the first preset number of customers from the plurality of predetermined service servers;
  • Customer attribute data such as a document number, or name and ID number; a screening server and a plurality of predetermined business servers (eg, a bank server, a medical server, an insurance server, an instant messaging server, a game server, a weather server, a takeaway server, and / / resume server, etc.) communication, each predetermined business server has a mapping relationship between customer attribute data and feature data; the screening server first selects a first preset number (for example, 500,000) of customers, and then, for each A selected customer separately extracts feature data corresponding to the customer attribute data from each of the predetermined service servers according to the customer attribute data of the customer, thereby obtaining various characteristic data of the selected first predetermined number of customers. .
  • a first preset number for example, 500,000
  • Step E2 analyzing the various feature data of each extracted customer according to a preset feature tag extraction rule to determine a feature tag of each client;
  • the extraction rule of the feature tag is preset in the screening server, and after extracting the first preset number of customer characteristic data, performing various feature data of each customer according to the preset extraction rule Feature tag analysis, the final analysis to get the feature tag of each customer.
  • Step E3 determining, according to a predetermined mapping relationship between the nasopharyngeal cancer and the customer attribute data, an abnormal customer having a nasopharyngeal cancer among the first predetermined number of customers, determining a feature tag corresponding to each abnormal customer, and each normal The feature tag corresponding to the customer and the feature tag corresponding to each abnormal client are used as training samples of the preset model;
  • the customer's disease record is recorded in the customer database of the screening server, and the screening server can determine which of the first predetermined number of customers is suffering from nasopharyngeal cancer by searching the customer database according to the customer's customer attribute data.
  • Abnormal customers which are normal customers who do not have nasopharyngeal cancer; after determining the abnormal customer and normal customer among the first preset number of customers, the characteristic label corresponding to each normal customer is taken as a training sample, and each abnormal customer The corresponding feature tag is also used as a training sample.
  • Step E4 the training sample is divided into a first percentage training set and a second percentage verification set, and the sum of the first percentage and the second percentage is less than or equal to 100%;
  • the first percentage of all training samples is used as the training set
  • the second percentage of all training samples is used as the validation set, for example, the first percentage is 65% and the second percentage is 35%.
  • both the training set and the verification set include training samples corresponding to some normal customers and training samples corresponding to some abnormal customers.
  • step E5 the analysis model is trained by using feature tags of each normal client in the training set and feature tags of each abnormal client, and after the training is completed, the feature tags of each normal client in the verification set and the feature tags of each abnormal client are utilized. Verifying the accuracy of the trained analytical model;
  • the training model is trained by the training set.
  • the verification set is used to verify the accuracy of the analysis result of the analysis model.
  • step E6 if the accuracy rate is greater than the preset threshold, the model training ends;
  • the accuracy threshold is preset in the screening server (ie, the preset threshold, for example, 98%). If the accuracy of the analysis result of the analysis model is determined to be greater than the accuracy threshold according to the verification set, the training of the analysis model is achieved.
  • step E7 if the accuracy rate is less than or equal to the preset threshold, the above steps E1, E2, and E3 are performed to increase the number of training samples, and the above steps E4 and E5 are re-executed based on the added training samples.
  • the screening server continues to perform steps E1, E2, and E3 at this time. And adding a first preset number of values (for example, increasing 50,000 each time), and re-executing the above steps E4 and E5 based on the added training samples until the requirement of step E6 is reached.
  • a first preset number of values for example, increasing 50,000 each time
  • the labeling threshold of the number of days in which the city PM2.5 (fine particles) exceeds the standard in the last three years may be 60 days.
  • the tag threshold corresponding to the number of days above the blue warning level of the city in the last three years may be 55 days.
  • the threshold is "55 days", it means that the bad weather in the city is too much; the number of days in the last year after sleeping at 23:00 in the last year may be 100 days.
  • the number of days in the last year is more than 23:00, the number of days is greater than the corresponding one.
  • the label threshold is "100 days", it means that it is preferred to sleep late; the labeling threshold of the number of take-outs of the barbecue in the most recent year may be 80 times.
  • the label threshold corresponding to the number of medical treatments in the last year may be 30 times, when the number of medical treatments in the most recent year is greater than the corresponding label threshold "30 times” , representing the number of medical treatments.
  • the label range corresponding to the city includes: a collection of cities in the northwest region, a collection of cities in the North China region, a collection of cities in the Central China region, a collection of cities in the South China region, etc., when the city to which the customer belongs belongs to the city in the northwest region city collection,
  • the label information corresponding to the city to which the customer belongs is “belonging to Northwest China”;
  • the label range corresponding to the eating habits of the city includes: spicy city collection, partial greasy city collection, partial light city collection, partial sweet/salty city combination, etc.
  • the label information corresponding to the city to which the customer belongs is “the eating habit is spicy”; the label range corresponding to the position includes: the collection of harmful gas positions, non-volatile and harmful The collection of gas positions, the collection of positions that are prone to generate volatile harmful gases, and the collection of harmless gas positions, etc.
  • the label information corresponding to the industry in which the customer is engaged is “working with harmful gases”. post”.
  • the tag information corresponding to the feature data of various continuous values of each client is determined, and the mapping relationship between various feature data types and tag ranges according to non-continuous values is determined.
  • the tag information corresponding to the feature data of various non-continuous values of each client is determined.
  • the present application also proposes a nasopharyngeal cancer screening analysis system.
  • FIG. 3 is a schematic diagram of an operating environment of a preferred embodiment of the nasopharyngeal cancer screening analysis system 10 of the present application.
  • the nasopharyngeal cancer screening analysis system 10 is installed and operated in the electronic device 1.
  • the electronic device 1 may be a computing device such as a desktop computer, a notebook, a palmtop computer, and a server.
  • the electronic device 1 may include, but is not limited to, a memory 11, a processor 12, and a display 13.
  • Figure 3 shows only the electronic device 1 with components 11-13, but it should be understood that not all illustrated components may be implemented, and more or fewer components may be implemented instead.
  • the memory 11 may be an internal storage unit of the electronic device 1 in some embodiments, such as a hard disk or memory of the electronic device 1.
  • the memory 11 may also be an external storage device of the electronic device 1 in other embodiments, such as a plug-in hard disk equipped on the electronic device 1, a smart memory card (SMC), and a secure digital (SD). Card, flash card, etc.
  • the memory 11 may also include both an internal storage unit of the electronic device 1 and an external storage device.
  • the memory 11 is used to store application software and various types of data installed in the electronic device 1, such as program codes of the nasopharyngeal cancer screening analysis system 10.
  • the memory 11 can also be used to temporarily store data that has been output or is about to be output.
  • the processor 12 may be a Central Processing Unit (CPU), microprocessor or other data processing chip for running program code or processing data stored in the memory 11, such as performing nasopharyngeal carcinoma Screening analysis system 10 and the like.
  • CPU Central Processing Unit
  • microprocessor or other data processing chip for running program code or processing data stored in the memory 11, such as performing nasopharyngeal carcinoma Screening analysis system 10 and the like.
  • the display 13 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch sensor, or the like in some embodiments.
  • the display 13 is for displaying information processed in the electronic device 1 and a user interface for displaying visualization.
  • the components 11-13 of the electronic device 1 communicate with one another via a system bus.
  • FIG. 4 is a program block diagram of a preferred embodiment of the nasopharyngeal cancer screening analysis system 10 of the present application.
  • the nasopharyngeal cancer screening analysis system 10 can be segmented into one or more modules, one or more modules being stored in the memory 11 and processed by one or more processors (this embodiment is processed The device 12) is executed to complete the application.
  • the nasopharyngeal cancer screening analysis system 10 can be divided into a first acquisition module 101, a second acquisition module 102, an extraction module 103, a first analysis module 104, and a second analysis module 105.
  • module refers to a series of computer program instructions that are capable of performing a particular function, and are more suitable than the program for describing the execution of the nasopharyngeal cancer screening analysis system 10 in the electronic device 1, wherein:
  • the first obtaining module 101 is configured to obtain, after receiving a nasopharyngeal cancer screening request of the client to be screened, the feature label of the to-screen customer from the nasopharyngeal cancer screening request;
  • the nasopharyngeal cancer screening request includes customer attribute data (eg, a certificate number, or a name and a document number) of the customer to be screened and/or a feature tag of the customer to be screened; wherein the feature tag includes basic feature information (eg, The city belongs to the northwestern part of China, the eating habits of the cities in which it belongs are spicy, the PM2.5 (fine particles) of the city is often exceeded, the bad weather in the city is too high, and the harmful gas positions are engaged, etc., and the preference information is preferred (for example, prefer to sleep late) , preference for tobacco and alcohol, preference for grilled food, preference for outdoor sports, etc.), behavioral information (for example, more games, more overtime, more take-outs, more medical visits, etc.) and/or social relationship information (for example, unmarried, Living alone, not always in contact with friends, etc.).
  • the screening server After receiving the nasopharyngeal cancer screening request from the client to be screened, the screening server first checks whether the
  • the second obtaining module 102 is configured to obtain, after the failure of acquiring the feature tag from the nasopharyngeal cancer screening request, the customer attribute data of the to-screen customer to be obtained from the nasopharyngeal cancer screening request;
  • the nasopharyngeal cancer screening request does not include the customer's feature tag
  • the feature tag is not obtained from the nasopharyngeal cancer screening request (ie, the feature tag is failed to be acquired)
  • the nasopharyngeal cancer screening is performed.
  • the customer attribute data of the to-screen customer is obtained from the request, to obtain the feature label of the customer to be screened according to the customer attribute data of the screening customer.
  • the extracting module 103 is configured to extract, from a plurality of predetermined service servers, various feature data corresponding to the customer attribute data of the to-be-screened client;
  • the screening server communicates with a plurality of predetermined business servers (eg, a bank server, a medical server, an insurance server, an instant messaging server, a game server, a weather server, a takeaway server, and/or a resume server, etc.); After the customer attribute data of the customer is to be screened, the screening server extracts various characteristic data corresponding to the customer attribute data of the to-screen customer from a plurality of predetermined business servers (for example, bank loan amount and repayment information) , outpatient medical record information "for example, the number of visits in a preset time, the type of disease, the duration of each illness, etc.”, insurance information "for example, the industry, gender, age, marital status, occupation, etc.”
  • Information about the use of the instant messaging tool account for example, information such as the communication tool daily login time information, daily online duration, etc.”, game information "for example, daily game login time information, daily game online duration, etc.”, weather information "for example , in the last three years, the number of days in which PM2.5 (fine
  • the first analysis module 104 is configured to perform feature tag analysis on the extracted feature data according to a preset feature tag extraction rule to analyze the feature tag of the to-screen customer;
  • the extraction rule of the feature tag is preset in the screening server, and after extracting various feature data corresponding to the to-screen customer, the extraction rule is used to perform feature tag analysis on the extracted feature data, thereby analyzing The feature tag of the customer to be screened.
  • the second analysis module 105 is configured to substitute the feature tag of the customer to be screened into the pre-trained analysis model to analyze whether the customer to be screened is a high-risk nasopharyngeal cancer medical judgment object.
  • the screening server has a pre-trained analysis model, that is, the analysis model is trained through a large number of known customer data (including whether the customer has nasopharyngeal cancer and the corresponding feature tag of the customer); the screening server is currently waiting After screening the customer's feature label, the obtained feature label is input into the pre-trained analysis model, and the analysis model analyzes whether the customer to be screened is a high-risk nasopharyngeal medical judgment object according to the input characteristic label.
  • the technical solution of the embodiment after receiving the nasopharyngeal cancer screening request of the client to be screened, according to the nasopharyngeal cancer screening request, obtaining the characteristic label of the customer to be screened, specifically directly screening from the nasopharyngeal cancer Obtaining a feature tag of the to-screen customer in the request, and, when the acquisition fails, acquiring, according to the customer attribute information in the nasopharyngeal cancer screening request, each of the customer attribute information corresponding to each of the predetermined service servers
  • the feature data is extracted from the feature data according to the preset extraction rule, and the feature tag of the to-screen customer is extracted after the feature tag is obtained, and then the obtained feature tag is input into the pre-trained analysis model, thereby analyzing Whether the client to be screened is a medical judgment target for high-risk nasopharyngeal cancer.
  • the feature tags representing the various feature data of the customer to be screened are input into the pre-trained analysis model for analysis, to analyze whether the customer to be screened is a high-risk nasopharyngeal medical judgment object, and the client is realized. Accurate and efficient screening of high-risk nasopharyngeal cancer medical judgment objects, allowing customers to prevent or treat early.
  • the analysis model is a multiple linear regression model, and the training process of the analysis model is (refer to FIG. 2):
  • Step E1 extract, according to the predetermined customer attribute data, various feature data of the first preset number of customers from the plurality of predetermined service servers;
  • Customer attribute data such as a document number, or name and ID number; a screening server and a plurality of predetermined business servers (eg, a bank server, a medical server, an insurance server, an instant messaging server, a game server, a weather server, a takeaway server, and / / resume server, etc.) communication, each predetermined business server has a mapping relationship between customer attribute data and feature data; the screening server first selects a first preset number (for example, 500,000) of customers, and then, for each A selected customer separately extracts feature data corresponding to the customer attribute data from each of the predetermined service servers according to the customer attribute data of the customer, thereby obtaining various characteristic data of the selected first predetermined number of customers. .
  • a first preset number for example, 500,000
  • Step E2 analyzing the various feature data of each extracted customer according to a preset feature tag extraction rule to determine a feature tag of each client;
  • the extraction rule of the feature tag is preset in the screening server, and after extracting the first preset number of customer characteristic data, performing various feature data of each customer according to the preset extraction rule Feature tag analysis, the final analysis to get the feature tag of each customer.
  • Step E3 determining, according to a predetermined mapping relationship between the nasopharyngeal cancer and the customer attribute data, an abnormal customer having a nasopharyngeal cancer among the first predetermined number of customers, determining a feature tag corresponding to each abnormal customer, and each normal The feature tag corresponding to the customer and the feature tag corresponding to each abnormal client are used as training samples of the preset model;
  • the customer's disease record is recorded in the customer database of the screening server, and the screening server can determine which of the first predetermined number of customers is suffering from nasopharyngeal cancer by searching the customer database according to the customer's customer attribute data.
  • Abnormal customers which are normal customers who do not have nasopharyngeal cancer; after determining the abnormal customer and normal customer among the first preset number of customers, the characteristic label corresponding to each normal customer is taken as a training sample, and each abnormal customer The corresponding feature tag is also used as a training sample.
  • Step E4 the training sample is divided into a first percentage training set and a second percentage verification set, and the sum of the first percentage and the second percentage is less than or equal to 100%;
  • the first percentage of all training samples is used as the training set
  • the second percentage of all training samples is used as the validation set, for example, the first percentage is 65% and the second percentage is 35%.
  • both the training set and the verification set include training samples corresponding to some normal customers and training samples corresponding to some abnormal customers.
  • step E5 the analysis model is trained by using feature tags of each normal client in the training set and feature tags of each abnormal client, and after the training is completed, the feature tags of each normal client in the verification set and the feature tags of each abnormal client are utilized. Verifying the accuracy of the trained analytical model;
  • the training model is trained by the training set.
  • the verification set is used to verify the accuracy of the analysis result of the analysis model.
  • step E6 if the accuracy rate is greater than the preset threshold, the model training ends;
  • the accuracy threshold is preset in the screening server (ie, the preset threshold, for example, 98%). If the accuracy of the analysis result of the analysis model is determined to be greater than the accuracy threshold according to the verification set, the training of the analysis model is achieved.
  • step E7 if the accuracy rate is less than or equal to the preset threshold, the above steps E1, E2, and E3 are performed to increase the number of training samples, and the above steps E4 and E5 are re-executed based on the added training samples.
  • the screening server continues to perform steps E1, E2, and E3 at this time. And adding a first preset number of values (for example, increasing 50,000 each time), and re-executing the above steps E4 and E5 based on the added training samples until the requirement of step E6 is reached.
  • a first preset number of values for example, increasing 50,000 each time
  • the preset feature tag extraction rules in the nasopharyngeal cancer screening analysis system of the present application are:
  • the labeling threshold of the number of days in which the city PM2.5 (fine particles) exceeds the standard in the last three years may be 60 days.
  • the tag threshold corresponding to the number of days above the blue warning level of the city in the last three years may be 55 days.
  • the threshold is "55 days", it means that the bad weather in the city is too much; the number of days in the last year after sleeping at 23:00 in the last year may be 100 days.
  • the number of days in the last year is more than 23:00, the number of days is greater than the corresponding one.
  • the label threshold is "100 days", it means that it is preferred to sleep late; the labeling threshold of the number of take-outs of the barbecue in the most recent year may be 80 times.
  • the label threshold corresponding to the number of medical treatments in the last year may be 30 times, when the number of medical treatments in the most recent year is greater than the corresponding label threshold "30 times” , representing the number of medical treatments.
  • the label range corresponding to the city includes: a collection of cities in the northwest region, a collection of cities in the North China region, a collection of cities in the Central China region, a collection of cities in the South China region, etc., when the city to which the customer belongs belongs to the city in the northwest region city collection,
  • the label information corresponding to the city to which the customer belongs is “belonging to Northwest China”;
  • the label range corresponding to the eating habits of the city includes: spicy city collection, partial greasy city collection, partial light city collection, partial sweet/salty city combination, etc.
  • the label information corresponding to the city to which the customer belongs is “the eating habit is spicy”; the label range corresponding to the position includes: the collection of harmful gas positions, non-volatile and harmful The collection of gas positions, the collection of positions that are prone to generate volatile harmful gases, and the collection of harmless gas positions, etc.
  • the label information corresponding to the industry in which the customer is engaged is “working with harmful gases”. post”.
  • the tag information corresponding to the feature data of various continuous values of each client is determined, and the mapping relationship between various feature data types and tag ranges according to non-continuous values is determined.
  • the tag information corresponding to the feature data of various non-continuous values of each client is determined.
  • the present application further provides a computer readable storage medium storing a nasopharyngeal cancer screening analysis system, the nasopharyngeal cancer screening analysis system being executable by at least one processor, Having the at least one processor perform the nasopharyngeal cancer screening assay method of any of the above embodiments.

Landscapes

  • Engineering & Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Public Health (AREA)
  • Medical Informatics (AREA)
  • Data Mining & Analysis (AREA)
  • Biomedical Technology (AREA)
  • Pathology (AREA)
  • Databases & Information Systems (AREA)
  • Epidemiology (AREA)
  • General Health & Medical Sciences (AREA)
  • Primary Health Care (AREA)
  • Medical Treatment And Welfare Office Work (AREA)
  • Investigating Or Analysing Biological Materials (AREA)

Abstract

一种电子装置、鼻咽癌筛查分析方法和计算机可读存储介质,该方法包括:在收到一个待筛查客户的鼻咽癌筛查请求后,从鼻咽癌筛查请求中获取该待筛查客户的特征标签(S10);若获取特征标签失败,则从鼻咽癌筛查请求中获取该待筛查客户的客户属性数据(S20);从多个预先确定的业务服务器提取出与该待筛查客户的客户属性数据对应的各种特征数据(S30);按照预设的特征标签提取规则,从提取出的各种特征数据中分析出该待筛查客户的特征标签(S40);将该待筛查客户的特征标签代入预先训练好的分析模型中,以分析出该待筛查客户是否是高危鼻咽癌医学判断对象(S50),其实现了客观准确的、高效的筛查出高危鼻咽癌医学判断对象,从而使客户及早进行预防或治疗。

Description

电子装置、鼻咽癌筛查分析方法和计算机可读存储介质
本申请基于巴黎公约申明享有2018年4月9日递交的申请号为CN2018103116296、名称为“电子装置、鼻咽癌筛查分析方法和计算机可读存储介质”中国专利申请的优先权,该中国专利申请的整体内容以参考的方式结合在本申请中。
技术领域
本申请涉及大数据分析应用技术领域,特别涉及一种电子装置、鼻咽癌筛查分析方法和计算机可读存储介质。
背景技术
鼻咽癌是一种高发恶性肿瘤,发病率为耳鼻咽喉恶性肿瘤之首,病因复杂。传统医学上对鼻咽癌的发病原因尚不明确,现有的研究大多来自于临床观察(与遗传、环境、病毒有关)。而临床观察获取个人信息有限,同时,现有的研究通常依赖于研究人员专业医学知识和个人经验,因此,研究的准确性无法满足要求,无法形成客观准确的、指标明确的发病原因筛查体系。
因此,如何实现对高危鼻咽癌医学判断对象客观准确的、高效的筛查,进而实现及早预防或治疗,已经成为一个亟待解决的技术问题。
发明内容
本申请提供一种电子装置、鼻咽癌筛查分析方法,旨在实现客观准确的、高效的筛查出高危鼻咽癌医学判断对象,从而使客户及早进行预防或治疗。
本申请第一方面提供一种电子装置,包括存储器和处理器,所述存储器上存储有可在所述处理器上运行的鼻咽癌筛查分析系统,所述鼻咽癌筛查分析系统被所述处理器执行时实现如下步骤:
在收到一个待筛查客户的鼻咽癌筛查请求后,从所述鼻咽癌筛查请求中获取该待筛查客户的特征标签;
若从所述鼻咽癌筛查请求中获取特征标签失败,则从所述鼻咽癌筛查请求中获取该待筛查客户的客户属性数据;
从多个预先确定的业务服务器提取出与该待筛查客户的客户属性数据对应的各种特征数据;
按照预设的特征标签提取规则,对提取出的各种特征数据进行特征标签分析,以分析出该待筛查客户的特征标签;
将该待筛查客户的特征标签代入预先训练好的分析模型中,以分析出该待筛查客户是否是高危鼻咽癌医学判断对象。
优选地,所述特征标签包括基本特征信息、偏好习惯信息、行为信息及/或社会关系信息。
本申请第二方面提供一种鼻咽癌筛查分析方法,该方法包括步骤:
在收到一个待筛查客户的鼻咽癌筛查请求后,从所述鼻咽癌筛查请求中获取该待筛查客户的特征标签;
若从所述鼻咽癌筛查请求中获取特征标签失败,则从所述鼻咽癌筛查请求中获取该待筛查客户的客户属性数据;
从多个预先确定的业务服务器提取出与该待筛查客户的客户属性数据对应的各种特征数据;
按照预设的特征标签提取规则,对提取出的各种特征数据进行特征标签分析,以分析出该待筛查客户的特征标签;
将该待筛查客户的特征标签代入预先训练好的分析模型中,以分析出该待筛查客户是否是高危鼻咽癌医学判断对象。
本申请第三方面提供一种计算机可读存储介质,所述计算机可读存储介质存储有鼻咽癌筛查分析系统,所述鼻咽癌筛查分析系统可被至少一个处理器执行,以使所述至少一个处理器执行如下步骤:
在收到一个待筛查客户的鼻咽癌筛查请求后,从所述鼻咽癌筛查请求中获取该待筛查客户的特征标签;
若从所述鼻咽癌筛查请求中获取特征标签失败,则从所述鼻咽癌筛查请求中获取该待筛查客户的客户属性数据;
从多个预先确定的业务服务器提取出与该待筛查客户的客户属性数据对应的各种特征数据;
按照预设的特征标签提取规则,对提取出的各种特征数据进行特征标签分析,以分析出该待筛查客户的特征标签;
将该待筛查客户的特征标签代入预先训练好的分析模型中,以分析出该待筛查客户是否是高危鼻咽癌医学判断对象。
本申请技术方案,在接收到待筛查客户的鼻咽癌筛查请求后,根据该鼻咽癌筛查请求去获取该待筛查客户的特征标签,具体为先直接从鼻咽癌筛查请求中获取该待筛查客户的特征标签,在获取失败时,则根据该鼻咽癌筛查请求中的客户属性信息,去从各个预先确定的业务服务器中获取该客户属性 信息对应的各种特征数据,再按照预设的提取规则从这些特征数据中提取出该待筛查客户的特征标签;在获得特征标签后,将获得的特征标签输入预先训练好的分析模型中,从而分析得出该待筛查客户是否为高危鼻咽癌医学判断对象。本申请通过将分别代表待筛查客户的各种特征数据的特征标签,输入预先训练好的分析模型中进行分析,以分析出待筛查客户是否为高危鼻咽癌医学判断对象,实现客户准确的、高效的筛查出高危鼻咽癌医学判断对象,让客户可以及早的进行预防或治疗。
附图说明
图1为本申请鼻咽癌筛查分析方法一实施例的流程示意图;
图2为本申请鼻咽癌筛查分析方法中分析模型的训练流程示意图;
图3为本申请鼻咽癌筛查分析系统一实施例的运行环境示意图;
图4为本申请鼻咽癌筛查分析系统一实施例的程序模块图。
具体实施方式
以下结合附图对本申请的原理和特征进行描述,所举实例只用于解释本申请,并非用于限定本申请的范围。
本申请提出一种鼻咽癌筛查分析方法。
如图1所示,图1为本申请鼻咽癌筛查分析方法一实施例的流程示意图。
本实施例中,该鼻咽癌筛查分析方法包括:
步骤S10,在收到一个待筛查客户的鼻咽癌筛查请求后,从所述鼻咽癌筛查请求中获取该待筛查客户的特征标签;
鼻咽癌筛查请求中包含待筛查客户的客户属性数据(例如,证件号码,或者,姓名和证件号码)及/或待筛查客户的特征标签;其中,特征标签包括基本特征信息(例如,所属城市属于中国西北地区、所属城市饮食习惯偏辛辣、所属城市PM2.5(细颗粒物)经常超标、所属城市恶劣天气偏多、从事有害气体岗位等)、偏好习惯信息(例如,偏好晚睡、偏好烟酒、偏好烧烤食物、偏好户外运动等)、行为信息(例如,游戏次数多、加班次数多、点外卖的次数多、就医次数多等)及/或社会关系信息(例如,未婚、独居、不经常与朋友联系等)。筛查服务器在收到一个待筛查客户的鼻咽癌筛查请求后,先看该请求中是否包含了该客户的特征标签,即从中去获取该客户的特征标签。
步骤S20,若从所述鼻咽癌筛查请求中获取特征标签失败,则从所述鼻咽癌筛查请求中获取该待筛查客户的客户属性数据;
如果该鼻咽癌筛查请求中没有包含该客户的特征标签,则从该鼻咽癌筛查请求中获取不到特征标签(即获取特征标签失败),那么,则从该鼻咽癌筛查请求中获取该待筛查客户的客户属性数据,以根据该带筛查客户的客户属性数据去获得该待筛查客户的特征标签。
步骤S30,从多个预先确定的业务服务器提取出与该待筛查客户的客户属性数据对应的各种特征数据;
筛查服务器与多个预先确定的业务服务器(例如,银行服务器、医疗服务器、保险服务器、即时通讯服务器、游戏服务器、天气服务器、外卖服务器及/或简历服务器等)之间通讯;在获取到该待筛查客户的客户属性数据后,筛查服务器从多个预先确定的业务服务器提取出与该待筛查客户的客户属性数据对应的各种特征数据(例如,银行贷款额度及还款情况信息、门诊病历信息“例如,预设时间内的看病次数、所患的疾病种类、每次患病的持续时间等”、保险信息“例如,所处行业,性别、年龄、婚姻状况、职业等”、即时通讯工具账号的使用信息“例如,通讯工具每天登陆时间信息、每天在线时长等信息”等等、游戏信息“例如,每天游戏登陆时间信息、每天游戏在线时长等信息”、天气信息“例如,最近三年内,PM2.5(细颗粒物)严重超标的天数”、外卖点餐信息“例如,每天点外卖的时间信息、每天所点外卖的外卖类型等”、求职简历上填写的信息“例如,兴趣爱好、性格、工作经历等信息”)。
步骤S40,按照预设的特征标签提取规则,对提取出的各种特征数据进行特征标签分析,以分析出该待筛查客户的特征标签;
筛查服务器中预先设置了特征标签的提取规则,在提取出该待筛查客户对应的各种特征数据后,利用该提取规则,对提取出的各种特征数据进行特征标签分析,从而分析得到该待筛查客户的特征标签。
步骤S50,将该待筛查客户的特征标签代入预先训练好的分析模型中,以分析出该待筛查客户是否是高危鼻咽癌医学判断对象。
筛查服务器中具有预先训练好的分析模型,即该分析模型是经过了大量已知客户数据(包括客户是否患鼻咽癌与客户对应的特征标签)的训练;筛查服务器在得到了当前待筛查客户的特征标签后,将得到的特征标签输入到该预先训练好的分析模型中,分析模型根据输入的各个特征标签,分析出该待筛查客户是否为高危鼻咽癌医学判断对象。
本实施例技术方案,在接收到待筛查客户的鼻咽癌筛查请求后,根据该鼻咽癌筛查请求去获取该待筛查客户的特征标签,具体为先直接从鼻咽癌筛 查请求中获取该待筛查客户的特征标签,在获取失败时,则根据该鼻咽癌筛查请求中的客户属性信息,去从各个预先确定的业务服务器中获取该客户属性信息对应的各种特征数据,再按照预设的提取规则从这些特征数据中提取出该待筛查客户的特征标签;在获得特征标签后,将获得的特征标签输入预先训练好的分析模型中,从而分析得出该待筛查客户是否为高危鼻咽癌医学判断对象。本实施例通过将分别代表待筛查客户的各种特征数据的特征标签,输入预先训练好的分析模型中进行分析,以分析出待筛查客户是否为高危鼻咽癌医学判断对象,实现客户准确的、高效的筛查出高危鼻咽癌医学判断对象,让客户可以及早的进行预防或治疗。
如图2所示,图2为本申请鼻咽癌筛查分析方法中分析模型的训练流程示意图。
本实施例中,所述分析模型为多元线性回归模型,所述分析模型的训练过程包括:
步骤E1,根据预先确定的客户属性数据,从多个预先确定的业务服务器提取出第一预设数量的客户的各种特征数据;
客户属性数据例如为证件号码,或姓名和证件号码;筛查服务器与多个预先确定的业务服务器(例如,银行服务器、医疗服务器、保险服务器、即时通讯服务器、游戏服务器、天气服务器、外卖服务器及/或简历服务器等)之间通讯,各个预先确定的业务服务器中具有客户属性数据与特征数据的映射关系;筛查服务器先选取第一预设数量(例如50万)的客户,然后,对于每一个选取的客户,分别根据该客户的客户属性数据去预先确定的各个业务服务器中,分别提取出其客户属性数据对应的特征数据,从而得到选取的第一预设数量的客户的各种特征数据。
步骤E2,按照预设的特征标签提取规则,对提取的各个客户的各种特征数据进行分析,以确定各个客户的特征标签;
筛查服务器中预先设置了特征标签的提取规则,在提取出第一预设数量的客户的各种特征数据后,根据该预先设置的提取规则,分别对每个客户的的各种特征数据进行特征标签分析,最终分析得到每一个客户的特征标签。
步骤E3,根据预先确定的鼻咽癌与客户属性数据的映射关系,确定所述第一预设数量的客户中患有鼻咽癌的异常客户,确定各个异常客户对应的特征标签,将各个正常客户对应的特征标签和各个异常客户对应的特征标签作为预设模型的训练样本;
筛查服务器的客户数据库中记录了每个客户的患病记录,筛查服务器根据客户的客户属性数据通过查找客户数据库就可确定该第一预设数量的客户中,哪些为患有鼻咽癌的异常客户,哪些为未患鼻咽癌的正常客户;在确定第一预设数量的客户中的异常客户和正常客户后,将每一个正常客户对应的特征标签作为一个训练样本,每一个异常客户对应的特征标签也作为一个训练样本。
步骤E4,将所述训练样本分为第一百分比的训练集和第二百分比的验证集,所述第一百分比和第二百分比之和小于或者等于100%;
将所有的训练样本的第一百分比作为训练集,以及所有的训练样本的第二百分比作为验证集,例如,第一百分比为65%,第二百分比为35%。当然,训练集和验证集中均包括部分正常客户对应的训练样本和部分异常客户对应的训练样本。
步骤E5,利用训练集中的各个正常客户的特征标签和各个异常客户的特征标签对所述分析模型进行训练,并在训练完成后利用验证集中的各个正常客户的特征标签和各个异常客户的特征标签对训练的所述分析模型的准确率进行验证;
先利用训练集训练分析模型,在训练完成后,再用验证集验证分析模型的分析结果的准确率。
步骤E6,若准确率大于预设阈值,则模型训练结束;
筛查服务器中预先设置了准确率阈值(即所述预设阈值,例如98%),如果根据验证集确定分析模型的分析结果的准确率大于该准确率阈值,则说明分析模型的训练达到了预期的标准,筛查服务器则确认分析模型的训练结束。
步骤E7,若准确率小于或者等于预设阈值,则执行上述步骤E1、E2、E3以增加训练样本的数量,并基于增加后的训练样本重新执行上述步骤E4和E5。
如果根据验证集确定分析模型的分析结果的准确率小于或等于该准确率阈值,则说明分析模型的训练还没有达到预期的标准,筛查服务器此时,则继续执行步骤E1、E2、E3,并增加第一预设数量的值(例如每次增大5万),并基于增加后的训练样本重新执行上述步骤E4和E5,直至达到步骤E6的要求。
本申请鼻咽癌筛查分析方法中所述预设的特征标签提取规则为:
对于连续数值的各种特征数据种类设置对应的标签阈值;
例如,所属城市PM2.5(细颗粒物)最近三年每年超标的天数对应的标签阈值可以为60天,当所属城市PM2.5最近三年每年超标的天数大于对应的标签阈值“60天”时,代表所属城市PM2.5经常超标;所属城市最近三年每年蓝色以上预警等级的天数对应的标签阈值可以为55天,当所属城市最近三年每年蓝色以上预警等级的天数大于对应的标签阈值“55天”时,代表所属城市恶劣天气偏多;最近一年晚于23:00睡觉的天数对应的标签阈值可以是100天,当最近一年晚于23:00睡觉的天数大于对应的标签阈值“100天”时,代表偏好晚睡;最近一年点烧烤类的外卖次数对应的标签阈值可以是80次,当最近一年点烧烤类的外卖次数大于对应的标签阈值“80次”时,代表偏好烧烤食物;最近一年就医次数对应的标签阈值可以是30次,当最近一年就医次数大于对应的标签阈值“30次”时,代表就医次数多。
对于非为连续数值的各种特征数据种类设置对应的标签范围;
例如,所属城市对应的标签范围包括:西北地区城市集合、华北地区城市集合、华中地区城市集合、华南地区城市集合等等,当客户所属城市属于所述西北地区城市集合中的城市时,代表该客户所属城市对应的标签信息是“属于中国西北地区”;所属城市饮食习惯对应的标签范围包括:偏辛辣城市集合、偏油腻城市集合、偏清淡城市集合、偏甜/咸城市结合等,当客户所属城市属于所述偏辛辣城市集合中的城市时,代表该客户所属城市对应的标签信息是“饮食习惯偏辛辣”;从事岗位对应的标签范围包括:有害气体岗位集合、非易产生挥发性有害气体的岗位集合、易产生挥发性有害气体的岗位集合、无害气体岗位集合等,当客户从事岗位属于有害气体岗位集合中的岗位,代表该客户所从事行业对应的标签信息是“从事有害气体岗位”。
根据连续数值的各种特征数据种类与标签阈值的映射关系,确定出各个客户的各种连续数值的特征数据对应的标签信息,及根据非连续数值的各种特征数据种类与标签范围的映射关系,确定出各个客户的各种非连续数值的特征数据对应的标签信息。
此外,本申请还提出一种鼻咽癌筛查分析系统。
请参阅图3,是本申请鼻咽癌筛查分析系统10较佳实施例的运行环境示意图。
在本实施例中,鼻咽癌筛查分析系统10安装并运行于电子装置1中。电子装置1可以是桌上型计算机、笔记本、掌上电脑及服务器等计算设备。该电子装置1可包括,但不仅限于,存储器11、处理器12及显示器13。图3 仅示出了具有组件11-13的电子装置1,但是应理解的是,并不要求实施所有示出的组件,可以替代的实施更多或者更少的组件。
存储器11在一些实施例中可以是电子装置1的内部存储单元,例如该电子装置1的硬盘或内存。存储器11在另一些实施例中也可以是电子装置1的外部存储设备,例如电子装置1上配备的插接式硬盘,智能存储卡(Smart Media Card,SMC),安全数字(Secure Digital,SD)卡,闪存卡(Flash Card)等。进一步地,存储器11还可以既包括电子装置1的内部存储单元也包括外部存储设备。存储器11用于存储安装于电子装置1的应用软件及各类数据,例如鼻咽癌筛查分析系统10的程序代码等。存储器11还可以用于暂时地存储已经输出或者将要输出的数据。
处理器12在一些实施例中可以是一中央处理器(Central Processing Unit,CPU),微处理器或其他数据处理芯片,用于运行存储器11中存储的程序代码或处理数据,例如执行鼻咽癌筛查分析系统10等。
显示器13在一些实施例中可以是LED显示器、液晶显示器、触控式液晶显示器以及OLED(Organic Light-Emitting Diode,有机发光二极管)触摸器等。显示器13用于显示在电子装置1中处理的信息以及用于显示可视化的用户界面。电子装置1的部件11-13通过系统总线相互通信。
请参阅图4,是本申请鼻咽癌筛查分析系统10较佳实施例的程序模块图。在本实施例中,鼻咽癌筛查分析系统10可以被分割成一个或多个模块,一个或者多个模块被存储于存储器11中,并由一个或多个处理器(本实施例为处理器12)所执行,以完成本申请。例如,在图4中,鼻咽癌筛查分析系统10可以被分割成第一获取模块101、第二获取模块102、提取模块103、第一分析模块104及第二分析模块105。本申请所称的模块是指能够完成特定功能的一系列计算机程序指令段,比程序更适合于描述鼻咽癌筛查分析系统10在电子装置1中的执行过程,其中:
第一获取模块101,用于在收到一个待筛查客户的鼻咽癌筛查请求后,从所述鼻咽癌筛查请求中获取该待筛查客户的特征标签;
鼻咽癌筛查请求中包含待筛查客户的客户属性数据(例如,证件号码,或者,姓名和证件号码)及/或待筛查客户的特征标签;其中,特征标签包括基本特征信息(例如,所属城市属于中国西北地区、所属城市饮食习惯偏辛辣、所属城市PM2.5(细颗粒物)经常超标、所属城市恶劣天气偏多、从事有害气体岗位等)、偏好习惯信息(例如,偏好晚睡、偏好烟酒、偏好烧烤食物、偏好户外运动等)、行为信息(例如,游戏次数多、加班次数多、点外卖 的次数多、就医次数多等)及/或社会关系信息(例如,未婚、独居、不经常与朋友联系等)。筛查服务器在收到一个待筛查客户的鼻咽癌筛查请求后,先看该请求中是否包含了该客户的特征标签,即从中去获取该客户的特征标签。
第二获取模块102,用于在从所述鼻咽癌筛查请求中获取特征标签失败后,从所述鼻咽癌筛查请求中获取该待筛查客户的客户属性数据;
如果该鼻咽癌筛查请求中没有包含该客户的特征标签,则从该鼻咽癌筛查请求中获取不到特征标签(即获取特征标签失败),那么,则从该鼻咽癌筛查请求中获取该待筛查客户的客户属性数据,以根据该带筛查客户的客户属性数据去获得该待筛查客户的特征标签。
提取模块103,用于从多个预先确定的业务服务器提取出与该待筛查客户的客户属性数据对应的各种特征数据;
筛查服务器与多个预先确定的业务服务器(例如,银行服务器、医疗服务器、保险服务器、即时通讯服务器、游戏服务器、天气服务器、外卖服务器及/或简历服务器等)之间通讯;在获取到该待筛查客户的客户属性数据后,筛查服务器从多个预先确定的业务服务器提取出与该待筛查客户的客户属性数据对应的各种特征数据(例如,银行贷款额度及还款情况信息、门诊病历信息“例如,预设时间内的看病次数、所患的疾病种类、每次患病的持续时间等”、保险信息“例如,所处行业,性别、年龄、婚姻状况、职业等”、即时通讯工具账号的使用信息“例如,通讯工具每天登陆时间信息、每天在线时长等信息”等等、游戏信息“例如,每天游戏登陆时间信息、每天游戏在线时长等信息”、天气信息“例如,最近三年内,PM2.5(细颗粒物)严重超标的天数”、外卖点餐信息“例如,每天点外卖的时间信息、每天所点外卖的外卖类型等”、求职简历上填写的信息“例如,兴趣爱好、性格、工作经历等信息”)。
第一分析模块104,用于按照预设的特征标签提取规则,对提取出的各种特征数据进行特征标签分析,以分析出该待筛查客户的特征标签;
筛查服务器中预先设置了特征标签的提取规则,在提取出该待筛查客户对应的各种特征数据后,利用该提取规则,对提取出的各种特征数据进行特征标签分析,从而分析得到该待筛查客户的特征标签。
第二分析模块105,用于将该待筛查客户的特征标签代入预先训练好的分析模型中,以分析出该待筛查客户是否是高危鼻咽癌医学判断对象。
筛查服务器中具有预先训练好的分析模型,即该分析模型是经过了大量已知客户数据(包括客户是否患鼻咽癌与客户对应的特征标签)的训练;筛 查服务器在得到了当前待筛查客户的特征标签后,将得到的特征标签输入到该预先训练好的分析模型中,分析模型根据输入的各个特征标签,分析出该待筛查客户是否为高危鼻咽癌医学判断对象。
本实施例技术方案,在接收到待筛查客户的鼻咽癌筛查请求后,根据该鼻咽癌筛查请求去获取该待筛查客户的特征标签,具体为先直接从鼻咽癌筛查请求中获取该待筛查客户的特征标签,在获取失败时,则根据该鼻咽癌筛查请求中的客户属性信息,去从各个预先确定的业务服务器中获取该客户属性信息对应的各种特征数据,再按照预设的提取规则从这些特征数据中提取出该待筛查客户的特征标签;在获得特征标签后,将获得的特征标签输入预先训练好的分析模型中,从而分析得出该待筛查客户是否为高危鼻咽癌医学判断对象。本实施例通过将分别代表待筛查客户的各种特征数据的特征标签,输入预先训练好的分析模型中进行分析,以分析出待筛查客户是否为高危鼻咽癌医学判断对象,实现客户准确的、高效的筛查出高危鼻咽癌医学判断对象,让客户可以及早的进行预防或治疗。
本实施例的鼻咽癌筛查分析系统中,所述分析模型为多元线性回归模型,该分析模型的训练过程为(可参照图2):
步骤E1,根据预先确定的客户属性数据,从多个预先确定的业务服务器提取出第一预设数量的客户的各种特征数据;
客户属性数据例如为证件号码,或姓名和证件号码;筛查服务器与多个预先确定的业务服务器(例如,银行服务器、医疗服务器、保险服务器、即时通讯服务器、游戏服务器、天气服务器、外卖服务器及/或简历服务器等)之间通讯,各个预先确定的业务服务器中具有客户属性数据与特征数据的映射关系;筛查服务器先选取第一预设数量(例如50万)的客户,然后,对于每一个选取的客户,分别根据该客户的客户属性数据去预先确定的各个业务服务器中,分别提取出其客户属性数据对应的特征数据,从而得到选取的第一预设数量的客户的各种特征数据。
步骤E2,按照预设的特征标签提取规则,对提取的各个客户的各种特征数据进行分析,以确定各个客户的特征标签;
筛查服务器中预先设置了特征标签的提取规则,在提取出第一预设数量的客户的各种特征数据后,根据该预先设置的提取规则,分别对每个客户的的各种特征数据进行特征标签分析,最终分析得到每一个客户的特征标签。
步骤E3,根据预先确定的鼻咽癌与客户属性数据的映射关系,确定所述 第一预设数量的客户中患有鼻咽癌的异常客户,确定各个异常客户对应的特征标签,将各个正常客户对应的特征标签和各个异常客户对应的特征标签作为预设模型的训练样本;
筛查服务器的客户数据库中记录了每个客户的患病记录,筛查服务器根据客户的客户属性数据通过查找客户数据库就可确定该第一预设数量的客户中,哪些为患有鼻咽癌的异常客户,哪些为未患鼻咽癌的正常客户;在确定第一预设数量的客户中的异常客户和正常客户后,将每一个正常客户对应的特征标签作为一个训练样本,每一个异常客户对应的特征标签也作为一个训练样本。
步骤E4,将所述训练样本分为第一百分比的训练集和第二百分比的验证集,所述第一百分比和第二百分比之和小于或者等于100%;
将所有的训练样本的第一百分比作为训练集,以及所有的训练样本的第二百分比作为验证集,例如,第一百分比为65%,第二百分比为35%。当然,训练集和验证集中均包括部分正常客户对应的训练样本和部分异常客户对应的训练样本。
步骤E5,利用训练集中的各个正常客户的特征标签和各个异常客户的特征标签对所述分析模型进行训练,并在训练完成后利用验证集中的各个正常客户的特征标签和各个异常客户的特征标签对训练的所述分析模型的准确率进行验证;
先利用训练集训练分析模型,在训练完成后,再用验证集验证分析模型的分析结果的准确率。
步骤E6,若准确率大于预设阈值,则模型训练结束;
筛查服务器中预先设置了准确率阈值(即所述预设阈值,例如98%),如果根据验证集确定分析模型的分析结果的准确率大于该准确率阈值,则说明分析模型的训练达到了预期的标准,筛查服务器则确认分析模型的训练结束。
步骤E7,若准确率小于或者等于预设阈值,则执行上述步骤E1、E2、E3以增加训练样本的数量,并基于增加后的训练样本重新执行上述步骤E4和E5。
如果根据验证集确定分析模型的分析结果的准确率小于或等于该准确率阈值,则说明分析模型的训练还没有达到预期的标准,筛查服务器此时,则继续执行步骤E1、E2、E3,并增加第一预设数量的值(例如每次增大5万),并基于增加后的训练样本重新执行上述步骤E4和E5,直至达到步骤E6的要求。
本申请鼻咽癌筛查分析系统中所述预设的特征标签提取规则为:
对于连续数值的各种特征数据种类设置对应的标签阈值;
例如,所属城市PM2.5(细颗粒物)最近三年每年超标的天数对应的标签阈值可以为60天,当所属城市PM2.5最近三年每年超标的天数大于对应的标签阈值“60天”时,代表所属城市PM2.5经常超标;所属城市最近三年每年蓝色以上预警等级的天数对应的标签阈值可以为55天,当所属城市最近三年每年蓝色以上预警等级的天数大于对应的标签阈值“55天”时,代表所属城市恶劣天气偏多;最近一年晚于23:00睡觉的天数对应的标签阈值可以是100天,当最近一年晚于23:00睡觉的天数大于对应的标签阈值“100天”时,代表偏好晚睡;最近一年点烧烤类的外卖次数对应的标签阈值可以是80次,当最近一年点烧烤类的外卖次数大于对应的标签阈值“80次”时,代表偏好烧烤食物;最近一年就医次数对应的标签阈值可以是30次,当最近一年就医次数大于对应的标签阈值“30次”时,代表就医次数多。
对于非为连续数值的各种特征数据种类设置对应的标签范围;
例如,所属城市对应的标签范围包括:西北地区城市集合、华北地区城市集合、华中地区城市集合、华南地区城市集合等等,当客户所属城市属于所述西北地区城市集合中的城市时,代表该客户所属城市对应的标签信息是“属于中国西北地区”;所属城市饮食习惯对应的标签范围包括:偏辛辣城市集合、偏油腻城市集合、偏清淡城市集合、偏甜/咸城市结合等,当客户所属城市属于所述偏辛辣城市集合中的城市时,代表该客户所属城市对应的标签信息是“饮食习惯偏辛辣”;从事岗位对应的标签范围包括:有害气体岗位集合、非易产生挥发性有害气体的岗位集合、易产生挥发性有害气体的岗位集合、无害气体岗位集合等,当客户从事岗位属于有害气体岗位集合中的岗位,代表该客户所从事行业对应的标签信息是“从事有害气体岗位”。
根据连续数值的各种特征数据种类与标签阈值的映射关系,确定出各个客户的各种连续数值的特征数据对应的标签信息,及根据非连续数值的各种特征数据种类与标签范围的映射关系,确定出各个客户的各种非连续数值的特征数据对应的标签信息。
进一步地,本申请还提出一种计算机可读存储介质,所述计算机可读存储介质存储有鼻咽癌筛查分析系统,所述鼻咽癌筛查分析系统可被至少一个处理器执行,以使所述至少一个处理器执行上述任一实施例中的鼻咽癌筛查 分析方法。
以上所述仅为本申请的优选实施例,并非因此限制本申请的专利范围,凡是在本申请的发明构思下,利用本申请说明书及附图内容所作的等效结构变换,或直接/间接运用在其他相关的技术领域均包括在本申请的专利保护范围内。

Claims (20)

  1. 一种电子装置,其特征在于,所述电子装置包括存储器和处理器,所述存储器上存储有可在所述处理器上运行的鼻咽癌筛查分析系统,所述鼻咽癌筛查分析系统被所述处理器执行时实现如下步骤:
    在收到一个待筛查客户的鼻咽癌筛查请求后,从所述鼻咽癌筛查请求中获取该待筛查客户的特征标签;
    若从所述鼻咽癌筛查请求中获取特征标签失败,则从所述鼻咽癌筛查请求中获取该待筛查客户的客户属性数据;
    从多个预先确定的业务服务器提取出与该待筛查客户的客户属性数据对应的各种特征数据;
    按照预设的特征标签提取规则,对提取出的各种特征数据进行特征标签分析,以分析出该待筛查客户的特征标签;
    将该待筛查客户的特征标签代入预先训练好的分析模型中,以分析出该待筛查客户是否是高危鼻咽癌医学判断对象。
  2. 如权利要求1所述的电子装置,其特征在于,所述预设的特征标签提取规则为:
    对于连续数值的各种特征数据种类设置对应的标签阈值;
    对于非为连续数值的各种特征数据种类设置对应的标签范围;
    根据连续数值的各种特征数据种类与标签阈值的映射关系,确定出各个客户的各种连续数值的特征数据对应的标签信息,及根据非连续数值的各种特征数据种类与标签范围的映射关系,确定出各个客户的各种非连续数值的特征数据对应的标签信息。
  3. 如权利要求1所述的电子装置,其特征在于,所述分析模型为多元线性回归模型,所述分析模型的训练步骤包括:
    E1、根据预先确定的客户属性数据,从多个预先确定的业务服务器提取出第一预设数量的客户的各种特征数据;
    E2、按照预设的特征标签提取规则,对提取的各个客户的各种特征数据进行分析,以确定各个客户的特征标签;
    E3、根据预先确定的鼻咽癌与客户属性数据的映射关系,确定所述第一预设数量的客户中患有鼻咽癌的异常客户,确定各个异常客户对应的特征标签,将各个正常客户对应的特征标签和各个异常客户对应的特征标签作为预设模型的训练样本;
    E4、将所述训练样本分为第一百分比的训练集和第二百分比的验证集, 所述第一百分比和第二百分比之和小于或者等于100%;
    E5、利用训练集中的各个正常客户的特征标签和各个异常客户的特征标签对所述分析模型进行训练,并在训练完成后利用验证集中的各个正常客户的特征标签和各个异常客户的特征标签对训练的所述分析模型的准确率进行验证;
    E6、若准确率大于预设阈值,则模型训练结束;
    E7、若准确率小于或者等于预设阈值,则执行上述步骤E1、E2、E3以增加训练样本的数量,并基于增加后的训练样本重新执行上述步骤E4和E5。
  4. 如权利要求3所述的电子装置,其特征在于,所述预设的特征标签提取规则为:
    对于连续数值的各种特征数据种类设置对应的标签阈值;
    对于非为连续数值的各种特征数据种类设置对应的标签范围;
    根据连续数值的各种特征数据种类与标签阈值的映射关系,确定出各个客户的各种连续数值的特征数据对应的标签信息,及根据非连续数值的各种特征数据种类与标签范围的映射关系,确定出各个客户的各种非连续数值的特征数据对应的标签信息。
  5. 如权利要求1所述的电子装置,其特征在于,所述特征标签包括基本特征信息、偏好习惯信息、行为信息及/或社会关系信息。
  6. 如权利要求5所述的电子装置,其特征在于,所述预设的特征标签提取规则为:
    对于连续数值的各种特征数据种类设置对应的标签阈值;
    对于非为连续数值的各种特征数据种类设置对应的标签范围;
    根据连续数值的各种特征数据种类与标签阈值的映射关系,确定出各个客户的各种连续数值的特征数据对应的标签信息,及根据非连续数值的各种特征数据种类与标签范围的映射关系,确定出各个客户的各种非连续数值的特征数据对应的标签信息。
  7. 如权利要求5所述的电子装置,其特征在于,所述分析模型为多元线性回归模型,所述分析模型的训练步骤包括:
    E1、根据预先确定的客户属性数据,从多个预先确定的业务服务器提取出第一预设数量的客户的各种特征数据;
    E2、按照预设的特征标签提取规则,对提取的各个客户的各种特征数据进行分析,以确定各个客户的特征标签;
    E3、根据预先确定的鼻咽癌与客户属性数据的映射关系,确定所述第一 预设数量的客户中患有鼻咽癌的异常客户,确定各个异常客户对应的特征标签,将各个正常客户对应的特征标签和各个异常客户对应的特征标签作为预设模型的训练样本;
    E4、将所述训练样本分为第一百分比的训练集和第二百分比的验证集,所述第一百分比和第二百分比之和小于或者等于100%;
    E5、利用训练集中的各个正常客户的特征标签和各个异常客户的特征标签对所述分析模型进行训练,并在训练完成后利用验证集中的各个正常客户的特征标签和各个异常客户的特征标签对训练的所述分析模型的准确率进行验证;
    E6、若准确率大于预设阈值,则模型训练结束;
    E7、若准确率小于或者等于预设阈值,则执行上述步骤E1、E2、E3以增加训练样本的数量,并基于增加后的训练样本重新执行上述步骤E4和E5。
  8. 如权利要求5所述的电子装置,其特征在于,所述预设的特征标签提取规则为:
    对于连续数值的各种特征数据种类设置对应的标签阈值;
    对于非为连续数值的各种特征数据种类设置对应的标签范围;
    根据连续数值的各种特征数据种类与标签阈值的映射关系,确定出各个客户的各种连续数值的特征数据对应的标签信息,及根据非连续数值的各种特征数据种类与标签范围的映射关系,确定出各个客户的各种非连续数值的特征数据对应的标签信息。
  9. 一种鼻咽癌筛查分析方法,其特征在于,该方法包括步骤:
    在收到一个待筛查客户的鼻咽癌筛查请求后,从所述鼻咽癌筛查请求中获取该待筛查客户的特征标签;
    若从所述鼻咽癌筛查请求中获取特征标签失败,则从所述鼻咽癌筛查请求中获取该待筛查客户的客户属性数据;
    从多个预先确定的业务服务器提取出与该待筛查客户的客户属性数据对应的各种特征数据;
    按照预设的特征标签提取规则,对提取出的各种特征数据进行特征标签分析,以分析出该待筛查客户的特征标签;
    将该待筛查客户的特征标签代入预先训练好的分析模型中,以分析出该待筛查客户是否是高危鼻咽癌医学判断对象。
  10. 如权利要求9所述的鼻咽癌筛查分析方法,其特征在于,所述预设的特征标签提取规则为:
    对于连续数值的各种特征数据种类设置对应的标签阈值;
    对于非为连续数值的各种特征数据种类设置对应的标签范围;
    根据连续数值的各种特征数据种类与标签阈值的映射关系,确定出各个客户的各种连续数值的特征数据对应的标签信息,及根据非连续数值的各种特征数据种类与标签范围的映射关系,确定出各个客户的各种非连续数值的特征数据对应的标签信息。
  11. 如权利要求9所述的鼻咽癌筛查分析方法,其特征在于,所述分析模型为多元线性回归模型,所述分析模型的训练步骤包括:
    E1、根据预先确定的客户属性数据,从多个预先确定的业务服务器提取出第一预设数量的客户的各种特征数据;
    E2、按照预设的特征标签提取规则,对提取的各个客户的各种特征数据进行分析,以确定各个客户的特征标签;
    E3、根据预先确定的鼻咽癌与客户属性数据的映射关系,确定所述第一预设数量的客户中患有鼻咽癌的异常客户,确定各个异常客户对应的特征标签,将各个正常客户对应的特征标签和各个异常客户对应的特征标签作为预设模型的训练样本;
    E4、将所述训练样本分为第一百分比的训练集和第二百分比的验证集,所述第一百分比和第二百分比之和小于或者等于100%;
    E5、利用训练集中的各个正常客户的特征标签和各个异常客户的特征标签对所述分析模型进行训练,并在训练完成后利用验证集中的各个正常客户的特征标签和各个异常客户的特征标签对训练的所述分析模型的准确率进行验证;
    E6、若准确率大于预设阈值,则模型训练结束;
    E7、若准确率小于或者等于预设阈值,则执行上述步骤E1、E2、E3以增加训练样本的数量,并基于增加后的训练样本重新执行上述步骤E4和E5。
  12. 如权利要求11所述的鼻咽癌筛查分析方法,其特征在于,所述预设的特征标签提取规则为:
    对于连续数值的各种特征数据种类设置对应的标签阈值;
    对于非为连续数值的各种特征数据种类设置对应的标签范围;
    根据连续数值的各种特征数据种类与标签阈值的映射关系,确定出各个客户的各种连续数值的特征数据对应的标签信息,及根据非连续数值的各种特征数据种类与标签范围的映射关系,确定出各个客户的各种非连续数值的特征数据对应的标签信息。
  13. 如权利要求9所述的鼻咽癌筛查分析方法,其特征在于,所述特征标签包括基本特征信息、偏好习惯信息、行为信息及/或社会关系信息。
  14. 如权利要求13所述的鼻咽癌筛查分析方法,其特征在于,所述预设的特征标签提取规则为:
    对于连续数值的各种特征数据种类设置对应的标签阈值;
    对于非为连续数值的各种特征数据种类设置对应的标签范围;
    根据连续数值的各种特征数据种类与标签阈值的映射关系,确定出各个客户的各种连续数值的特征数据对应的标签信息,及根据非连续数值的各种特征数据种类与标签范围的映射关系,确定出各个客户的各种非连续数值的特征数据对应的标签信息。
  15. 如权利要求13所述的鼻咽癌筛查分析方法,其特征在于,所述分析模型为多元线性回归模型,所述分析模型的训练步骤包括:
    E1、根据预先确定的客户属性数据,从多个预先确定的业务服务器提取出第一预设数量的客户的各种特征数据;
    E2、按照预设的特征标签提取规则,对提取的各个客户的各种特征数据进行分析,以确定各个客户的特征标签;
    E3、根据预先确定的鼻咽癌与客户属性数据的映射关系,确定所述第一预设数量的客户中患有鼻咽癌的异常客户,确定各个异常客户对应的特征标签,将各个正常客户对应的特征标签和各个异常客户对应的特征标签作为预设模型的训练样本;
    E4、将所述训练样本分为第一百分比的训练集和第二百分比的验证集,所述第一百分比和第二百分比之和小于或者等于100%;
    E5、利用训练集中的各个正常客户的特征标签和各个异常客户的特征标签对所述分析模型进行训练,并在训练完成后利用验证集中的各个正常客户的特征标签和各个异常客户的特征标签对训练的所述分析模型的准确率进行验证;
    E6、若准确率大于预设阈值,则模型训练结束;
    E7、若准确率小于或者等于预设阈值,则执行上述步骤E1、E2、E3以增加训练样本的数量,并基于增加后的训练样本重新执行上述步骤E4和E5。
  16. 如权利要求15所述的鼻咽癌筛查分析方法,其特征在于,所述预设的特征标签提取规则为:
    对于连续数值的各种特征数据种类设置对应的标签阈值;
    对于非为连续数值的各种特征数据种类设置对应的标签范围;
    根据连续数值的各种特征数据种类与标签阈值的映射关系,确定出各个客户的各种连续数值的特征数据对应的标签信息,及根据非连续数值的各种特征数据种类与标签范围的映射关系,确定出各个客户的各种非连续数值的特征数据对应的标签信息。
  17. 一种计算机可读存储介质,其特征在于,所述计算机可读存储介质存储有鼻咽癌筛查分析系统,所述鼻咽癌筛查分析系统可被至少一个处理器执行,以使所述至少一个处理器执行如下步骤:
    在收到一个待筛查客户的鼻咽癌筛查请求后,从所述鼻咽癌筛查请求中获取该待筛查客户的特征标签;
    若从所述鼻咽癌筛查请求中获取特征标签失败,则从所述鼻咽癌筛查请求中获取该待筛查客户的客户属性数据;
    从多个预先确定的业务服务器提取出与该待筛查客户的客户属性数据对应的各种特征数据;
    按照预设的特征标签提取规则,对提取出的各种特征数据进行特征标签分析,以分析出该待筛查客户的特征标签;
    将该待筛查客户的特征标签代入预先训练好的分析模型中,以分析出该待筛查客户是否是高危鼻咽癌医学判断对象。
  18. 如权利要求17所述的计算机可读存储介质,其特征在于,所述特征标签包括基本特征信息、偏好习惯信息、行为信息及/或社会关系信息。
  19. 如权利要求17所述的计算机可读存储介质,其特征在于,所述分析模型为多元线性回归模型,所述分析模型的训练步骤包括:
    E1、根据预先确定的客户属性数据,从多个预先确定的业务服务器提取出第一预设数量的客户的各种特征数据;
    E2、按照预设的特征标签提取规则,对提取的各个客户的各种特征数据进行分析,以确定各个客户的特征标签;
    E3、根据预先确定的鼻咽癌与客户属性数据的映射关系,确定所述第一预设数量的客户中患有鼻咽癌的异常客户,确定各个异常客户对应的特征标签,将各个正常客户对应的特征标签和各个异常客户对应的特征标签作为预设模型的训练样本;
    E4、将所述训练样本分为第一百分比的训练集和第二百分比的验证集,所述第一百分比和第二百分比之和小于或者等于100%;
    E5、利用训练集中的各个正常客户的特征标签和各个异常客户的特征标签对所述分析模型进行训练,并在训练完成后利用验证集中的各个正常客户 的特征标签和各个异常客户的特征标签对训练的所述分析模型的准确率进行验证;
    E6、若准确率大于预设阈值,则模型训练结束;
    E7、若准确率小于或者等于预设阈值,则执行上述步骤E1、E2、E3以增加训练样本的数量,并基于增加后的训练样本重新执行上述步骤E4和E5。
  20. 如权利要求17所述的计算机可读存储介质,其特征在于,所述预设的特征标签提取规则为:
    对于连续数值的各种特征数据种类设置对应的标签阈值;
    对于非为连续数值的各种特征数据种类设置对应的标签范围;
    根据连续数值的各种特征数据种类与标签阈值的映射关系,确定出各个客户的各种连续数值的特征数据对应的标签信息,及根据非连续数值的各种特征数据种类与标签范围的映射关系,确定出各个客户的各种非连续数值的特征数据对应的标签信息。
PCT/CN2018/102108 2018-04-09 2018-08-24 电子装置、鼻咽癌筛查分析方法和计算机可读存储介质 Ceased WO2019196300A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
SG11202008389RA SG11202008389RA (en) 2018-04-09 2018-08-24 Electronic apparatus, nasopharyngeal cancer screening and analyzing method, and computer readable medium

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201810311629.6 2018-04-09
CN201810311629.6A CN108831552A (zh) 2018-04-09 2018-04-09 电子装置、鼻咽癌筛查分析方法和计算机可读存储介质

Publications (1)

Publication Number Publication Date
WO2019196300A1 true WO2019196300A1 (zh) 2019-10-17

Family

ID=64154424

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2018/102108 Ceased WO2019196300A1 (zh) 2018-04-09 2018-08-24 电子装置、鼻咽癌筛查分析方法和计算机可读存储介质

Country Status (3)

Country Link
CN (1) CN108831552A (zh)
SG (1) SG11202008389RA (zh)
WO (1) WO2019196300A1 (zh)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN110956628B (zh) * 2019-12-13 2023-05-09 广州达安临床检验中心有限公司 图片等级分类方法、装置、计算机设备和存储介质

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20150213302A1 (en) * 2014-01-30 2015-07-30 Case Western Reserve University Automatic Detection Of Mitosis Using Handcrafted And Convolutional Neural Network Features
CN106529110A (zh) * 2015-09-09 2017-03-22 阿里巴巴集团控股有限公司 一种用户数据分类的方法和设备
CN107220506A (zh) * 2017-06-05 2017-09-29 东华大学 基于深度卷积神经网络的乳腺癌风险评估分析系统
CN107247887A (zh) * 2017-07-27 2017-10-13 点内(上海)生物科技有限公司 基于人工智能帮助肺癌筛查的方法及系统

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN105744005A (zh) * 2016-04-30 2016-07-06 平安证券有限责任公司 客户定位分析方法及服务器

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20150213302A1 (en) * 2014-01-30 2015-07-30 Case Western Reserve University Automatic Detection Of Mitosis Using Handcrafted And Convolutional Neural Network Features
CN106529110A (zh) * 2015-09-09 2017-03-22 阿里巴巴集团控股有限公司 一种用户数据分类的方法和设备
CN107220506A (zh) * 2017-06-05 2017-09-29 东华大学 基于深度卷积神经网络的乳腺癌风险评估分析系统
CN107247887A (zh) * 2017-07-27 2017-10-13 点内(上海)生物科技有限公司 基于人工智能帮助肺癌筛查的方法及系统

Also Published As

Publication number Publication date
SG11202008389RA (en) 2020-09-29
CN108831552A (zh) 2018-11-16

Similar Documents

Publication Publication Date Title
Adegoke et al. Cervical cancer trends in the United States: a 35-year population-based analysis
Carter et al. Correcting for bias in psychology: A comparison of meta-analytic methods
Johnston et al. The prevalence of chronic fatigue syndrome/myalgic encephalomyelitis: a meta-analysis
Diaz-Ordaz et al. Are missing data adequately handled in cluster randomised trials? A systematic review and guidelines
CA2991230C (en) Genetic and genealogical analysis for identification of birth location and surname information
WO2019061994A1 (zh) 电子装置、保险产品推荐方法、系统及计算机可读存储介质
WO2019071965A1 (zh) 数据处理的方法、数据处理装置及计算机可读存储介质
JP6038727B2 (ja) 分析システム及び分析方法
Wunsch et al. Validation of intensive care and mechanical ventilation codes in Medicare data
US20140006044A1 (en) System and method for preparing healthcare service bundles
WO2019071906A1 (zh) 金融产品推荐装置、方法及计算机可读存储介质
CN109712712B (zh) 一种健康评估方法、健康评估装置及计算机可读存储介质
US11366927B1 (en) Computing system for de-identifying patient data
Soljak et al. Does higher quality primary health care reduce stroke admissions? A national cross-sectional study
Lim et al. Smoking cessation and mortality among middle-aged and elderly Chinese in Singapore: the Singapore Chinese Health Study
Mittmann et al. Utilization and costs of home care for patients with colorectal cancer: a population-based study
CN114756669A (zh) 问题意图的智能分析方法、装置、电子设备及存储介质
CN114840660A (zh) 业务推荐模型训练方法、装置、设备及存储介质
WO2019062186A1 (zh) 糖尿病分析方法、应用服务器和计算机可读存储介质
Checa et al. Social risk and mortality: a cohort study in patients with advanced heart failure
JP7121276B2 (ja) データ管理レベル判定プログラム、およびデータ管理レベル判定方法
de Courville et al. Secondary healthcare resource utilization and related costs associated with influenza-related hospital admissions in adult patients, England 2016–2020
WO2019196300A1 (zh) 电子装置、鼻咽癌筛查分析方法和计算机可读存储介质
Liu et al. Generalized linear mixed models for multi-reader multi-case studies of diagnostic tests
Hulíková Tesárková et al. The age structure of cases as the key of COVID-19 severity: Longitudinal population-based analysis of European countries during 150 days

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 18914786

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 26/01/2021)

122 Ep: pct application non-entry in european phase

Ref document number: 18914786

Country of ref document: EP

Kind code of ref document: A1