WO2020015216A1 - 参数阈值确定方法、装置及计算机存储介质 - Google Patents
参数阈值确定方法、装置及计算机存储介质 Download PDFInfo
- Publication number
- WO2020015216A1 WO2020015216A1 PCT/CN2018/111123 CN2018111123W WO2020015216A1 WO 2020015216 A1 WO2020015216 A1 WO 2020015216A1 CN 2018111123 W CN2018111123 W CN 2018111123W WO 2020015216 A1 WO2020015216 A1 WO 2020015216A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- parameter
- threshold
- candidate
- tuning
- parameter threshold
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q30/00—Commerce
- G06Q30/02—Marketing; Price estimation or determination; Fundraising
- G06Q30/0201—Market modelling; Market analysis; Collecting market data
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F17/00—Digital computing or data processing equipment or methods, specially adapted for specific functions
- G06F17/10—Complex mathematical operations
- G06F17/18—Complex mathematical operations for evaluating statistical data, e.g. average values, frequency distributions, probability functions, regression analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q40/00—Finance; Insurance; Tax strategies; Processing of corporate or income taxes
- G06Q40/08—Insurance
Definitions
- the present application relates to the field of information technology, and in particular, to a method, a device, and a computer storage medium for determining a parameter threshold.
- a rule engine is built before a rating model is built.
- the rule engine uses a series of parameter reduction conditions in the rating evaluation model, and then the rating evaluation model evaluates each step to be performed according to the parameter reduction conditions until a conclusion is reached.
- Parameter reduction conditions are usually constructed from parameters and parameter thresholds. For example, for the amount consumption scenario, when assessing the risk level of the amount consumption, the parameters involved in the amount consumption can have parameter a: the amount of consumption; parameter b: the number of consumption items.
- the corresponding parameter specification conditions can be Is a> x 1 and b> y 1 ; when the level of money consumption is low, the corresponding parameter reduction conditions a ⁇ x 2 and b ⁇ y 2 ; when the risk level of money consumption is medium, the corresponding parameter reduction conditions can be : A> x 1 and y 1 >b> y 2 , or x 1 >a> x 2 and b> y 1 or x 1 > a and y 1 >b> y 2 .
- x 1 and x 2 are parameter thresholds of parameter a
- y 1 and y 2 are parameter thresholds of parameter b.
- the parameter threshold for rating is usually determined based on the business experience of the business party or the expert, that is, the parameter threshold given by the business party or the expert based on the business experience is directly determined as the parameter threshold for the rating.
- the level assessed according to the above parameter threshold may not meet the actual situation of the data, resulting in lower accuracy of the parameter threshold of the level assessment and data level The accuracy of the assessment is low.
- the present application provides a method, a device, and a computer storage medium for determining a parameter threshold, mainly to enable the parameter threshold given based on business experience to be optimized based on the actual situation of data distribution, and to improve the accuracy of the parameter threshold for rating. This can improve the accuracy of data grade assessment.
- a method for determining a parameter threshold including:
- the parameter threshold of the rating is determined according to the adjusted parameter threshold.
- a device for determining a parameter threshold including:
- a determining unit configured to determine a parameter candidate threshold of each parameter according to the data distribution of the sample data on each parameter
- a tuning unit configured to use a preset tuning hypothesis and the candidate parameter threshold to tune a parameter threshold given according to business experience
- the determining unit is configured to determine a parameter threshold value of the rating according to the adjusted parameter threshold value.
- a computer non-volatile readable storage medium on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the following steps are implemented:
- the parameter candidate thresholds for each parameter determine the parameter candidate thresholds for each parameter; use a preset tuning assumption and the candidate parameter thresholds to tune the parameter thresholds given based on business experience;
- the preference parameter threshold determines the parameter threshold of the rating.
- a computer device including a memory, a processor, and computer-readable instructions stored on the memory and executable on the processor.
- the processor executes the computer-readable instructions, Implement the following steps:
- the parameter candidate thresholds for each parameter determine the parameter candidate thresholds for each parameter; use a preset tuning assumption and the candidate parameter thresholds to tune the parameter thresholds given based on business experience;
- the preference parameter threshold determines the parameter threshold of the rating.
- Determine the parameter candidate thresholds of each parameter can use a preset tuning hypothesis and the candidate parameter thresholds to tune the parameter thresholds given based on business experience, and determine the graded rating based on the tuned parameter thresholds
- Parameter thresholds can be used to optimize the parameter thresholds based on business experience based on the actual situation of data distribution, and to improve the accuracy of parameter thresholds for grade evaluation, thereby improving the accuracy of data grade evaluation.
- FIG. 1 shows a flowchart of a method for determining a parameter threshold according to an embodiment of the present application
- FIG. 2 shows a flowchart of another method for determining a parameter threshold according to an embodiment of the present application
- FIG. 3 is a schematic structural diagram of a parameter threshold determining apparatus according to an embodiment of the present application.
- FIG. 4 is a schematic structural diagram of another parameter threshold determining apparatus according to an embodiment of the present application.
- FIG. 5 is a schematic structural diagram of a computer device according to an embodiment of the present application.
- the parameter threshold value of the rating is usually determined according to the business experience of the business party or the expert, that is, the parameter threshold value given by the business party or the expert based on the business experience is directly determined as the parameter threshold value of the rating.
- the level assessed according to the above parameter threshold may not conform to the actual situation of the data, resulting in lower accuracy of the parameter threshold of the level assessment and data level. The accuracy of the assessment is low.
- an embodiment of the present application provides a method for determining a parameter threshold. As shown in FIG. 1, the method includes:
- the sample data is the amount of consumption data in the scenario of bank card theft
- the parameters corresponding to the amount of consumption data may include: parameter a: consumption amount; parameter b: consumption amount, then the amount consumption data is divided between the consumption amount and the consumption amount.
- parameter a consumption amount
- parameter b consumption amount
- the amount consumption data is divided between the consumption amount and the consumption amount.
- the sample data is human attribute data
- the parameters corresponding to the human attribute data may include: parameter a: weight; parameter b: height
- the human attribute data will show a certain distribution on the two parameters of weight and height, such as human attributes
- a statistical line chart of each parameter may be drawn according to the data distribution on each parameter, and then each inflection point in the statistical line chart is found as the candidate parameter threshold.
- the parameter candidate threshold determined for each parameter may have a set.
- the parameter candidate threshold determined for parameter a may include ;
- the parameter candidate threshold determined for parameter b may include .
- the preset tuning assumptions can be set according to specific application scenarios or actual requirements, which are not limited in the embodiment of the present application. Specifically, the tuning assumption may be set based on the number of levels and the number of parameters of the data. For example, the preset tuning assumptions may include:
- the R hp represents a percentile of a parameter threshold of the p-th level of the h-th parameter
- the consumption value threshold corresponding to the adjusted high-risk level is 48525 yuan and the corresponding consumption number is 10, so the consumption amount threshold value is 48525 yuan and the consumption number is 10 as The parameter threshold for assessing the high-risk level, that is, when the consumption amount is greater than 48,525 yuan, and the number of consumptions is greater than 10, the level determined as the amount of consumption is the high-risk level.
- the consumption amount threshold and the number of consumption items correspond to the consumption amount threshold value of 50,000 yuan corresponding to the current high-risk level given based on business experience, and the corresponding consumption amount is 10, which is more in line with the consumption data.
- the actual distribution of consumption amount and consumption number is 48525 yuan and the corresponding consumption number is 10.
- a method for determining a parameter threshold provided in the embodiment of the present application compared with the parameter threshold currently determined for rating based on the business experience of a business party or an expert, the embodiment of the present application can determine all parameters according to the data distribution of sample data on each parameter.
- the parameter candidate thresholds of each parameter are described; the parameter thresholds given based on business experience can be tuned using preset tuning assumptions and the candidate parameter thresholds, and the parameter thresholds for rating are determined based on the tuned parameter thresholds , So that the parameter threshold given based on business experience can be optimized based on the actual situation of data distribution, and the accuracy of the parameter threshold for grade evaluation can be improved, thereby improving the accuracy of data grade evaluation.
- an embodiment of the present application provides another method for determining a parameter threshold, as shown in FIG. 2, the method includes :
- the statistic may be one or more of quantile, intra-group dispersion, or inter-group distance.
- each parameter can be directly calculated in the 0-100 quantile; if the statistic is an intra-group deviation or an inter-group distance, you can follow a certain step size Threshold divides the sample data into two groups of data, and then calculates the intra-group dispersion according to the intra-group dispersion calculation formula or the inter-group distance calculation formula to calculate the inter-group distance between the two groups; finally, according to a certain step
- the long movement threshold is used to calculate the intra-group dispersion or inter-group distance for each movement to obtain the intra-group dispersion or inter-group distance.
- the step 202 may specifically be: drawing a statistic line chart corresponding to each parameter according to the statistics, and the abscissa of the statistic line chart may be a specific parameter value, and the ordinate may be Statistics.
- the statistic line chart is a quantile line chart
- the quantile line chart can be drawn according to the 0-100 quantile calculated in step 201.
- the abscissa of the quantile line chart is the specific value of the parameter a and the ordinate. Is the quantile of parameter a.
- the statistic line chart is the intra-group dispersion line chart
- the inflection point in the statistical line chart may be a point with a large difference in front-to-back slope.
- the point in the statistical line chart with a positive slope and a negative backward slope may be determined as the required inflection point.
- the inflection point can be found using the following formula:
- x i can represent the abscissa value of the quantile line graph
- y i can represent the ordinate value of the quantile line graph.
- the parameter candidate threshold of the parameter a determined through the step 203 may include:
- the parameter b parameter candidate threshold may include
- the preset tuning assumptions include:
- the R hp represents a percentile of a parameter threshold of the p-th level of the h-th parameter
- the risk level of the risk level is 3 levels, which can be high risk level, medium risk level, and low risk level.
- the sample data has 2 parameters and the number of consumption items.
- the tuning assumptions can be:
- the step 204 may specifically include:
- the sample data is 100 pieces of consumption data
- the given high-risk level consumption amount threshold x 1 is 50,000 yuan
- the number of consumption y 1 is 10
- the consumption amount threshold is 50,000 yuan
- the consumption is 10
- the hypothesis 2 According to the combined threshold division table corresponding to the hypothesis 3, the hypothesis 2, and the candidate parameter threshold, select a percentage of the sample that is approximately equal to the proportion of the sample before the tuning, and the percentage of the corresponding candidate parameter threshold.
- Candidate parameter threshold combinations with approximately equal digits are used as the first set of parameter thresholds after tuning.
- the step 2 specifically includes: determining a combination threshold division table corresponding to the candidate parameter threshold, and the combination threshold division table stores a correspondence between the candidate parameter threshold and a percentile thereof, A sample ratio corresponding to the candidate parameter threshold combination of each parameter, and a mapping relationship between the corresponding relationship and the sample ratio; according to the hypothesis 3, look up from the combination threshold partition table from the combination threshold partition table before the tuning And the first correspondence corresponding to the proportions of the multiple samples found; according to the hypothesis 2, find each corresponding corresponding percentile approximately from the first correspondence.
- the candidate parameter thresholds of the parameters and according to the candidate parameter thresholds found determine the first set of parameter thresholds after tuning.
- the step of determining a combined threshold partition table corresponding to the candidate parameter threshold may specifically include: determining the candidate parameter threshold combination of each parameter according to the candidate parameter threshold; and calculating the candidate parameter threshold combination of each parameter
- the ratio of the corresponding amount of sample data to the total amount of sample data is used as the proportion of samples corresponding to the candidate parameter threshold combination of each of the parameters; establishing a correspondence between the candidate parameter threshold and its percentile.
- a mapping relationship between the comparison relationship and the sample proportion and based on the correspondence relationship and the mapping relationship, a combined threshold division table corresponding to the candidate parameter threshold is constructed.
- parameter candidate thresholds for parameter a include Parameter b parameter candidate threshold includes Can be combined into an n ⁇ m candidate parameter threshold combination matrix, as follows:
- the corresponding sample ratio can be calculated, indicating the ratio of the amount of sample data to the total sample data amount for which both parameters are greater than the corresponding candidate parameter threshold:
- the correspondence between the candidate parameter threshold of parameter a and its percentile can be:
- the correspondence between the candidate parameter threshold of parameter b and its percentile can be:
- step 2 can be based on Hypothesis 3 to find from Table 1 the proportion of samples that is approximately equal to the proportion of samples before tuning H.
- H i-1j , H ij , H ij-1 , and corresponding correspondences can be found according to H i-1j :
- the values of r i and r j can be found to be relatively close from the above corresponding relationship, then Is the candidate parameter threshold of the parameter a to be found, Is the candidate parameter threshold of the parameter b to be found; finally That is, the first set of candidate thresholds after tuning, which is one of the results of tuning a given parameter threshold (x 1 , y 1 ).
- the combined threshold division table, and the hypothesis 1 determine other tuned parameter thresholds.
- the step 3 specifically includes: calculating the sample proportions corresponding to the thresholds of other groups of parameters according to the tuned sample proportion corresponding to the first set of parameter thresholds and the hypothesis 1;
- the combination threshold value division table searches for a second correspondence relationship corresponding to the sample proportion corresponding to the other group parameter threshold value, and determines the other group parameter threshold value according to the second correspondence relationship.
- a first set of candidate sample accounted threshold corresponding to P i-1 (P i- 1 with the same steps here above in the meaning of H ij 2, so to avoid confusion with P i-1 denotes the first time after tuning
- the proportion of samples corresponding to a set of candidate thresholds H ij denotes the first time after tuning
- a tuning inequality can be set. It is assumed that the proportion of samples corresponding to the level of the next set of parameter thresholds is P i . :
- P i can be calculated, and P i can be returned to the combined threshold division table to find the corresponding sample proportion.
- the candidate parameter threshold where the two parameters are located can be selected.
- the selected candidate parameter thresholds can be used to determine the next group of parameter thresholds and other group parameter thresholds.
- the corresponding risk level can be determined Determine the consumption amount threshold and consumption amount threshold corresponding to the high risk level, such as
- the embodiment of the present application is described by using only two parameters as an example. If there are multiple parameters, the above steps are repeated and iterated until each set of thresholds is obtained, and details are not described herein.
- the sample data amount, percentiles, and sample proportion are usually not integers, and may be numerical values with multiple decimals, and sample data amounts of different levels.
- the embodiment of the present application can determine the data according to the data distribution of each parameter on the sample data.
- Parameter candidate thresholds for each parameter can use the preset tuning assumptions and the candidate parameter thresholds to tune the parameter thresholds given based on business experience, and determine the parameters for rating based on the tuned parameter thresholds Thresholds, so that parameter thresholds based on business experience can be adjusted based on the actual situation of data distribution, and the accuracy of parameter thresholds for level evaluation can be improved, thereby improving the accuracy of data level evaluation.
- an embodiment of the present application provides a parameter threshold determining device. As shown in FIG. 3, the device includes a determining unit 31 and an tuning unit 32.
- the determining unit 31 may be configured to determine a parameter candidate threshold of each parameter according to the data distribution of the sample data on each parameter.
- the determining unit 31 is a data distribution of the parameter on each parameter according to the sample data in the device.
- the tuning unit 32 may be configured to use a preset tuning hypothesis and the candidate parameter threshold to tune a parameter threshold given according to business experience.
- the tuning unit 32 is a main function module of the device for tuning a parameter threshold given according to business experience by using a preset tuning assumption and the candidate parameter threshold, and is also a core module.
- the determining unit 31 may be configured to determine a parameter threshold value of the rating according to the adjusted parameter threshold value.
- the determining unit 31 is also a main function module for determining a parameter threshold value of the rating according to the parameter threshold value after optimization in the present device.
- the determining unit 31 may be specifically configured to calculate the statistics of the sample data on each parameter according to the data distribution of the sample data on each parameter; and determine the various parameters according to the statistics
- the corresponding statistical line chart, and the candidate parameter thresholds of the respective parameters are determined according to threshold values corresponding to respective inflection points in the statistical line chart.
- the statistic may be one or more of quantile, intra-group dispersion, or inter-group distance.
- the preset tuning assumptions include:
- the R hp represents a percentile of a parameter threshold of the p-th level of the h-th parameter
- the tuning unit 32 may include a calculation module 321, a selection module 322, and a determination module 322, as shown in FIG. 4.
- the calculation module 321 may be configured to calculate a ratio of a sample data amount to a total sample data amount of the given parameter threshold corresponding to a parameter threshold combination, as a sample proportion before tuning.
- the selection module 322 may be configured to select, based on the combined threshold division table corresponding to the hypothesis 3, the hypothesis 2, and the candidate parameter threshold, a sample ratio that is approximately equal to the sample ratio before the tuning and The candidate parameter threshold combinations whose percentiles corresponding to the candidate parameter thresholds are approximately equal are used as the first set of parameter thresholds after tuning.
- the determining module 323 may be configured to determine the thresholds of other groups of parameters after tuning according to the first group of parameter thresholds after adjustment, the combination threshold division table, and the hypothesis 1.
- the selection module 322 may include: a determination sub-module 3221 and a search sub-module 3222.
- the determination submodule 3221 may be used to determine a combination threshold division table corresponding to the candidate parameter threshold.
- the combination threshold division table stores a correspondence relationship between the candidate parameter threshold and a percentile of the candidate parameter threshold.
- a sample ratio corresponding to the candidate parameter threshold combination and a mapping relationship between the correspondence relationship and the sample ratio.
- the search submodule 3222 may be configured to find, from the combination threshold partition table according to the hypothesis 3, a sample ratio that is approximately equal to the sample ratio before the tuning, and a ratio of multiple samples to the search. Corresponding first correspondence.
- the determining sub-module 3221 may be further configured to find, from the first correspondence relationship, candidate parameter thresholds corresponding to parameters whose percentiles are approximately equal to each other according to the hypothesis 2, and determine tuning according to the candidate parameter thresholds found. After the first set of parameter thresholds.
- the determining unit 31 may be specifically configured to calculate the sample proportions corresponding to the thresholds of other groups of parameters according to the sample proportion corresponding to the first set of parameter thresholds after the tuning and the hypothesis 1. Searching for a second correspondence relationship corresponding to the sample proportion corresponding to the other group parameter threshold value from the combined threshold division table, and determining the other group parameter threshold value according to the second correspondence relationship.
- the determination submodule 3221 may be specifically configured to determine the candidate parameter threshold combination of each parameter according to the candidate parameter threshold; calculate the various parameters The ratio of the sample data amount corresponding to the candidate parameter threshold value combination to the total sample data amount is used as the sample proportion corresponding to the candidate parameter threshold value combination for each of the parameters; establishing the candidate parameter threshold value and its percentile A corresponding relationship between them, and a mapping relationship between the comparison relationship and the sample proportion; and a combined threshold division table corresponding to the candidate parameter threshold is constructed according to the correspondence relationship and the mapping relationship.
- an embodiment of the present application further provides a computer non-volatile readable storage medium that stores computer-readable instructions.
- the computer-readable instructions are executed by a processor, The following steps are implemented: determining the parameter candidate thresholds of each parameter according to the data distribution of the sample data on each parameter; using a preset tuning assumption and the candidate parameter threshold to adjust the parameter threshold given according to business experience Excellent; the parameter threshold of the rating is determined according to the adjusted parameter threshold.
- an embodiment of the present application further provides a physical structure diagram of a computer device.
- the device includes a processor. 41.
- the following steps are implemented: determining the parameter candidate thresholds of each parameter according to the data distribution of the sample data on each parameter; using a preset tuning assumption and the candidate parameter threshold to adjust the parameter threshold given according to business experience Excellent; the parameter threshold of the rating is determined according to the adjusted parameter threshold.
- the parameter candidate thresholds of each parameter can be determined according to the data distribution of the sample data on each parameter; the preset tuning hypothesis and the candidate parameter threshold can be used to give a set value based on business experience. Tune the parameter thresholds and determine the parameter thresholds for the rating based on the tuned parameter thresholds, so that the parameter thresholds based on business experience can be adjusted based on the actual situation of the data distribution, and the parameter thresholds for the rating are improved. The accuracy of the data level can be improved.
- modules or steps of the present application may be implemented by a general-purpose computing device, and they may be concentrated on a single computing device or distributed in a network composed of multiple computing devices.
- they can be implemented with computer-readable instruction code executable by the computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, can be different from this
- the steps shown or described are performed in sequence, either by making them into individual integrated circuit modules, or by making multiple modules or steps into a single integrated circuit module. As such, this application is not limited to any particular combination of hardware and software.
Landscapes
- Engineering & Computer Science (AREA)
- Business, Economics & Management (AREA)
- Physics & Mathematics (AREA)
- Accounting & Taxation (AREA)
- Finance (AREA)
- General Physics & Mathematics (AREA)
- Data Mining & Analysis (AREA)
- Development Economics (AREA)
- Strategic Management (AREA)
- Theoretical Computer Science (AREA)
- Mathematical Physics (AREA)
- Marketing (AREA)
- General Business, Economics & Management (AREA)
- Economics (AREA)
- Computational Mathematics (AREA)
- Pure & Applied Mathematics (AREA)
- Entrepreneurship & Innovation (AREA)
- Mathematical Optimization (AREA)
- Mathematical Analysis (AREA)
- Life Sciences & Earth Sciences (AREA)
- Evolutionary Biology (AREA)
- Bioinformatics & Computational Biology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Operations Research (AREA)
- Probability & Statistics with Applications (AREA)
- Technology Law (AREA)
- Algebra (AREA)
- Game Theory and Decision Science (AREA)
- Databases & Information Systems (AREA)
- Software Systems (AREA)
- General Engineering & Computer Science (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
- Complex Calculations (AREA)
Abstract
本申请公开了一种参数阈值确定方法、装置及计算机存储介质,涉及信息技术领域,主要目的在于能够实现基于数据分布的实际情况对依据业务经验给定的参数阈值进行调优,提升等级评定的参数阈值的精确度,从而能够提升数据等级评定的精确度。所述方法包括:根据样本数据在各个参数上的数据分布,确定所述各个参数的参数候选阈值;利用预先设定的调优假设和所述候选参数阈值对依据业务经验给定的参数阈值进行调优;根据调优后的参数阈值确定等级评定的参数阈值。本申请适用于参数阈值的确定。
Description
本申请要求与2018年7月19日提交中国专利局、申请号为2018107970543、申请名称为“参数阈值确定方法、装置及计算机存储介质”的中国专利申请的优先权,其全部内容通过引用结合在申请中。
本申请涉及信息技术领域,尤其是涉及一种参数阈值确定方法、装置及计算机存储介质。
在某些场景中,为了评定数据的等级,通常需要建立等级评定模型。一般来说,建立等级评定模型之前会构建规则引擎。规则引擎会即在等级评定模型中使用一系列的参数规约条件,然后等级评定模型根据参数规约条件评定要执行的每一步步骤,直到得出等级的结论。通常参数规约条件是由参数和参数阈值构建的。例如,针对金额消费场景,评定金额消费的风险等级时,金额消费涉及的参数可以有参数a:消费金额;参数b:消费笔数,金额消费的风险等级为高时,对应的参数规约条件可以为a>x
1 and b>y
1;金额消费的等级为低时,对应的参数规约条件a<x
2 and b<y
2;金额消费的风险等级为中时,对应的参数规约条件可以为:a>x
1 and y
1>b>y
2、或者x
1>a>x
2 and b>y
1或者x
1>a and y
1>b>y
2。其中,x
1、x
2为参数a的参数阈值;y
1、y
2为参数b的参数阈值。
目前,通常根据业务方或者专家的业务经验确定等级评定的参数阈值,即直接将业务方或者专家的依据业务经验给出的参数阈值确定为等级评定的参数阈值。然而,在业务方或者专家的业务经验不足或者针对新场景的情况下,可能会造成根据上述参数阈值评定的等级不符合数据的实际情况,导致等级评定的参数阈值的精确度较低,数据等级评定的精确度较低。
发明内容
本申请提供了一种参数阈值确定方法、装置及计算机存储介质,主要在于能够实现基于数据分布的实际情况对依据业务经验给定的参数阈值进行调优,提升等级评定的参数阈值的精确度,从而能够提升数据等级评定的精确度。
根据本申请的第一个方面,提供一种参数阈值确定方法,包括:
根据样本数据在各个参数上的数据分布,确定所述各个参数的参数候选阈值;
利用预先设定的调优假设和所述候选参数阈值对依据业务经验给定的参数阈值进行调优;
根据调优后的参数阈值确定等级评定的参数阈值。
根据本申请的第二个方面,提供一种参数阈值确定装置,包括:
确定单元,用于根据样本数据在各个参数上的数据分布,确定所述各个参数的参数候选阈值;
调优单元,用于利用预先设定的调优假设和所述候选参数阈值对依据业务经验给定的参数阈值进行调优;
所述确定单元,用于根据调优后的参数阈值确定等级评定的参数阈值。
根据本申请的第三个方面,提供一种计算机非易失性可读存储介质,其上存储有计算机可读指令,该计算机可读指令被处理器执行时实现以下步骤:
根据样本数据在各个参数上的数据分布,确定所述各个参数的参数候选阈值;利用预先设定的调优假设和所述候选参数阈值对依据业务经验给定的参数阈值进行调优;根据调优后的参数阈值确定等级评定的参数阈值。
根据本申请的第四个方面,提供一种计算机设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机可读指令,所述处理器执行所述计算机可读指令时实现以下步骤:
根据样本数据在各个参数上的数据分布,确定所述各个参数的参数候选阈值;利用预先设定的调优假设和所述候选参数阈值对依据业务经验给定的参数阈值进行调优;根据调优后的参数阈值确定等级评定的参数阈值。
本申请提供的一种参数阈值确定方法、装置及计算机存储介质,与目前根据业务方或者专家的业务经验确定等级评定的参数阈值相比,本申请能够根据样本数据在各个参数上的数据分布,确定所述各个参数的参数候选阈值;能够利用预先设定的调优假设和所述候选参数阈值对依据业务经验给定的参数阈值进行调优,并根据调优后的参数阈值确定等级评定的参数阈值,从而能够实现基于数据分布的实际情况对依据业务经验给定的参数阈值进行调优,提升等级评定的参数阈值的精确度,从而能够提升数据等级评定的精确度。
此处所说明的附图用来提供对本申请的进一步理解,构成本申请的一部分,本申请的 示意性实施例及其说明用于解释本申请,并不构成对本申请的不当限定。在附图中:
图1示出了本申请实施例提供的一种参数阈值确定方法流程图;
图2示出了本申请实施例提供的另一种参数阈值确定方法流程图;
图3示出了本申请实施例提供的一种参数阈值确定装置的结构示意图;
图4示出了本申请实施例提供的另一种参数阈值确定装置的结构示意图;
图5示出了本申请实施例提供的一种计算机设备的实体结构示意图。
下文中将参考附图并结合实施例来详细说明本申请。需要说明的是,在不冲突的情况下,本申请中的实施例及实施例中的特征可以相互组合。
如背景技术,目前,通常根据业务方或者专家的业务经验确定等级评定的参数阈值,即直接将业务方或者专家的依据业务经验给出的参数阈值确定为等级评定的参数阈值。然而,在业务方或者专家的业务经验不足或者针对新场景的情况下,可能会造成根据上述参数阈值评定的等级不符合数据的实际情况,导致等级评定的参数阈值的精确度较低,数据等级评定的精确度较低。
为了解决上述问题,本申请实施例提供了一种参数阈值确定方法,如图1所示,所述方法包括:
101、根据样本数据在各个参数上的数据分布,确定所述各个参数的参数候选阈值。
例如,样本数据为银行卡盗刷场景中的金额消费数据,金额消费数据对应的参数可以包括:参数a:消费金额;参数b:消费笔数,那么金额消费数据在消费金额和消费笔数两个参数上都会呈现一定的分布,如金额消费数据为100条,消费笔数为10笔以上的金额消费数据可能有10条,消费笔数为10笔以下的金额消费数据可能有90条;消费金额在5万元以上的金额消费数据可能有20条,消费金额在在5万以下的金额消费数据可能有80条等。又例如,样本数据为人体属性数据,人体属性数据对应的参数可以包括:参数a:体重;参数b:身高,那么人体属性数据在体重和身高两个参数上都会呈现一定的分布,如人体属性数据有100条,体重在100斤以上的人体属性数据可能有80条,体重在100斤以下的人体属性数据可能有20条,身高在1.6m以上的人体属性数据可能有70条,身高在1.6m以下的人体属性数据可能有20条等。
在本申请实施例中,可以根据在各个参数上的述数据分布,绘制各个参数的统计量折线图,然后通过查找所述统计量折线图中的各个拐点作为候选参数阈值。针对每个参数确 定的参数候选阈值可以有一组,例如,针对参数a确定的参数候选阈值可以包括
;针对参数b确定的参数候选阈值可以包括
。
102、利用预先设定的调优假设和所述候选参数阈值对依据业务经验给定的参数阈值进行调优。
其中,所述预先设定的调优假设可以根据具体应用场景或者实际需求进行设置,本申请实施例在此不做限定。具体地,所述调优假设可以为根据等级评定的级数,数据的参数个数设定的。例如,所述预先设定的调优假设可以包括:
假设1:若待评定的等级有p级,则N
p:N
p-1≈N
p-1:N
p-2≈...≈N
2:N
1,所述N
p表示第p个等级所包含的样本数据量;
假设2:若参数个数为h,则调优后的每个参数的各级参数阈值符合
所述R
hp表示第h个参数的第p个等级的参数阈值所在百分位数;
假设3:调优后的参数阈值组合对应的样本占比约等于调优前的样本占比。
103、根据调优后的参数阈值确定等级评定的参数阈值。
例如,在银行卡盗刷场景中,调优后的高风险等级对应的消费金额阈值为48525元,对应的消费笔数为10笔,则将消费金额阈值48525元,消费笔数10笔确定为评定高风险等级的参数阈值,即在消费金额大于48525元,且消费笔数大于10笔时确定为金额消费的等级为高风险等级。本申请实施例中的消费金额阈值和消费笔数与目前依据业务经验给定的高风险等级对应的消费金额阈值为5万元,对应的消费笔数为10笔相比,更符合金额消费数据在消费金额和消费笔数上分布的实际情况。
本申请实施例提供的一种参数阈值确定方法,与目前根据业务方或者专家的业务经验确定等级评定的参数阈值相比,本申请实施例能够根据样本数据在各个参数上的数据分布,确定所述各个参数的参数候选阈值;能够利用预先设定的调优假设和所述候选参数阈值对依据业务经验给定的参数阈值进行调优,并根据调优后的参数阈值确定等级评定的参数阈值,从而能够实现基于数据分布的实际情况对依据业务经验给定的参数阈值进行调 优,提升等级评定的参数阈值的精确度,从而能够提升数据等级评定的精确度。
进一步的,为了更好的说明上述参数阈值确定的过程,作为对上述实施例的细化和扩展,本申请实施例提供了另一种参数阈值确定方法,如图2所示,所述方法包括:
201、根据样本数据在各个参数上的数据分布,计算所述样本数据在各个参数上的统计量。
其中,所述统计量可以为分位数、组内离差或者组间距离中的一种或者多种。例如,若所述统计量为分位数,则可以直接计算所述各个参数在0-100分位数;若所述统计量为组内离差或者组间距离,则可以按照一定的步长阈值将所述样本数据划分为两组数据,然后根据组内离差计算公式,计算两组的组内离差或者根据组间距离计算公式,计算两组的组间距离;最后按照一定的步长移动阈值,计算每次移动的组内离差或者组间距离,得到一组组内离差或者组间距离。
202、根据所述统计量确定所述各个参数对应的统计量折线图。
对于本申请实施例,所述步骤202具体可以为:根据所述统计量绘制所述各个参数对应的统计量折线图,所述统计量折线图的横坐标可以为参数具体数值,纵坐标可以为统计量。例如,若统计量折线图为分位数折线图,可以根据步骤201计算的0-100分位数绘制分位数折线图,分位数折线图的横坐标为参数a的具体数值,纵坐标为参数a的分位数。若统计量折线图为组内离差折线图,可以根据步骤201计算的一组组内离差折线图绘制组内离差折线图;若统计量折线图为组间距离折线图,可以根据步骤201计算的一组组间距离折线图绘制组间距离折线图。
203、根据所述统计量折线图中的各个拐点对应的阈值,确定所述各个参数的的候选参数阈值。
对于本申请实施例,所述统计量折线图中的拐点可以为前后斜率差异较大的点,可以将统计量折线图中前斜率为正,后斜率为负的点确定为为所需要的拐点。具体地,当所述统计量折线图为分位数折线图时,可以利用如下公式查找拐点:
204、利用预先设定的调优假设和所述候选参数阈值对依据业务经验给定的参数阈值进行调优。
其中,所述预先设定的调优假设包括:
假设1:若待评定的等级有p级,则N
p:N
p-1≈N
p-1:N
p-2≈...≈N
2:N
1,所述N
p表示第p个等级所包含的样本数据量;
假设2:若参数个数为h,则调优后的每个参数的各级参数阈值符合
所述R
hp表示第h个参数的第p个等级的参数阈值所在百分位数;
假设3:调优后的参数阈值组合对应的样本占比约等于调优前的样本占比。
例如,针对银行卡盗刷的场景,风险等级评定的风险等级有3级,可以分别为高风险等级,中风险等级,低风险等级,样本数据的参数有2个参数,消费笔数;设定的调优假设可以为:
假设1:低风险等级的样本数据量:中风险等级的样本数据量≈中风险等级的样本数据量:高风险等级的样本数据量;或者低风险等级的样本占比:中风险等级的样本占比≈中风险等级的样本占比:高风险等级的样本占比;
假设2:消费金额阈值所在百分位数≈消费笔数阈值所在百分位数;
假设3:调优后的消费金额阈值和消费笔数阈值组合对应的样本占比约等于调优前的样本占比。
对于本申请实施例,所述步骤204具体可以包括:
1、计算所述给定的参数阈值对应参数阈值组合的样本数据量与总样本数据量的比值,作为调优前的样本占比。
需要说明的是,所述样本占比H通过如下方式表示:
例如,样本数据为100条金额消费数据,给定的高风险等级的消费金额阈值x
1为5万 元,消费笔数y
1为10笔,消费金额阈值为5万元且消费笔数为10笔的金额消费数据为20条,则调优前的样本占比为:20/100=0.2。
2、根据所述假设3、所述假设2和所述候选参数阈值对应的组合阈值划分表,选择样本占比与所述调优前的样本占比约等于的且对应候选参数阈值所在百分位数约等于的候选参数阈值组合,作为调优后的第一组参数阈值。
对于本申请实施例,所述步骤2具体包括:确定所述候选参数阈值对应的组合阈值划分表,所述组合阈值划分表中保存有所述候选参数阈值与其所在百分位数的对应关系,所述各个参数的候选参数阈值组合对应的样本占比,以及所述对应关系与所述样本占比的映射关系;根据所述假设3从所述组合阈值划分表中查找与所述调优前的样本占比约等于的样本占比,以及与查找的多个样本占比对应的第一对应关系;根据所述假设2从所述第一对应关系中查找对应百分位数约等于的各个参数的候选参数阈值并根据查找的候选参数阈值,确定调优后的第一组参数阈值。
此外,所述确定所述候选参数阈值对应的组合阈值划分表的步骤具体可以包括:根据所述候选参数阈值,确定所述各个参数的候选参数阈值组合;计算所述各个参数的候选参数阈值组合对应的样本数据量与总样本数据量的比值,作为所述所述各个参数的候选参数阈值组合对应的样本占比;建立所述候选参数阈值及其所在百分位数的之间的对应关系,以及所述对比关系与所述样本占比之间的映射关系;根据所述对应关系和所述映射关系,构建所述候选参数阈值对应的组合阈值划分表。
针对组合矩阵中的每个元素(候选参数阈值组合),可以计算对应的样本占比,表示两个参数均大于对应候选参数阈值的样本数据量与总样本数据量的比值:
参数a的候选参数阈值及其所在百分位数的之间的对应关系可以为:
参数b的候选参数阈值及其所在百分位数的之间的对应关系可以为:
所述对比关系与所述样本占比之间的映射关系可以为:H
11分别与
和
映射;…;H
ij分别与
和
…H
nm分别与
和
映射;确定的所述组合阈值划分表可以如表1所示为:
表1
因此,关于所述步骤2可以为根据假设3从表1中寻找与调优前的样本占比H约等于的样本占比,如与调优前的样本占比H约等于的样本占比有H
i-1j、H
ij、H
ij-1、根据H
i-1j可以查找对应的对应关系:
根据H
ij可以查找对应的对应关系:
根据H
ij-1可以查找对应的对应关系:
根据所述假设2可以从上述对应关系找到r
i和r
j的值比较接近,则
即为所要查找的参数a的候选参数阈值,
即为所要查找的参数b的候选参数阈值;最后
即为调优后的第一组候选阈值,即为对给定的参数阈值(x
1,y
1)进行调优的其中一个结果。在银行盗刷应用场景中,
可以为中风险等级对应的消费金额阈值和消费笔数阈值。
3、根据所述调优后的第一组参数阈值、所述组合阈值划分表和所述假设1,确定调优后的其他组参数阈值。
对于本申请实施例,所述步骤3具体包括:根据所述调优后的第一组参数阈值对应的的样本占比和所述假设1,计算其他组参数阈值对应的样本占比;从所述组合阈值划分表中查找与所述其他组参数阈值对应的样本占比对应的第二对应关系,并根据所述第二对应关系,确定所述其他组参数阈值。
例如,第一组候选阈值对应的样本占比为P
i-1(此处的P
i-1与上述步骤2中H
ij含义相同,为了避免混淆因此用P
i-1表示调优后的第一组候选阈值对应的样本占比H
ij),在本申请实施例中根据假设1,可以设定调优不等式,假设下组参数阈值对应等级的样本占比为P
i,调优不等式可以满足:
根据上述不等式即可以计算出P
i,根据P
i重新回到所述组合阈值划分表可以找对应的样本占比;根据找到的样本占比,可以选取出两个参数所在的候选参数阈值;根据选取的候选参数阈值,即可以确定下组参数阈值以及其他组参数阈值,在银行盗刷应用场景中,可以根据中风险等级对应的
确定出高风险等级对应的消费金额阈值和消费笔数阈值,如为
需要说明的是,本申请实施例仅是以2个参数的情况进行举例说明,若有多个参数,重复迭代上述步骤,直到求出每一组阈值,在此不进行赘述。此外,由于在实际各种应用场景中的参数阈值调优计算中,样本数据量、百分位数、样本占比通常并非为整数,可能为有多位小数的数值,不同等级的样本数据量、参数阈值的百分位数、调优前后的样本占比,很难做到相等,因此本申请实施例中的用约等于设定各种假设,本申请实施例涉及的约等于或者接近,并非含义不清楚的情况。
205、根据调优后的参数阈值确定等级评定的参数阈值。
本申请实施例提供的另一种参数阈值确定方法,与目前根据业务方或者专家的业务经验确定等级评定的参数阈值相比,本申请实施例能够根据样本数据在各个参数上的数据分布,确定所述各个参数的参数候选阈值;能够利用预先设定的调优假设和所述候选参数阈值对依据业务经验给定的参数阈值进行调优,并根据调优后的参数阈值确定等级评定的参数阈值,从而能够实现基于数据分布的实际情况对依据业务经验给定的参数阈值进行调优,提升等级评定的参数阈值的精确度,从而能够提升数据等级评定的精确度。
进一步地,作为图1的具体实现,本申请实施例提供了一种参数阈值确定装置,如图3所示,所述装置包括:确定单元31和调优单元32。
所述确定单元31,可以用于根据样本数据在各个参数上的数据分布,确定所述各个参数的参数候选阈值,所述确定单元31是本装置中根据样本数据在各个参数上的数据分布,确定所述各个参数的参数候选阈值的主要功能模块。
所述调优单元32,可以用于利用预先设定的调优假设和所述候选参数阈值对依据业务经验给定的参数阈值进行调优。所述调优单元32是本装置利用预先设定的调优假设和所述候选参数阈值对依据业务经验给定的参数阈值进行调优的主要功能模块,也是核心模块。
所述确定单元31,可以用于根据调优后的参数阈值确定等级评定的参数阈值。所述确定单元31还是本装置中根据调优后的参数阈值确定等级评定的参数阈值的主要功能模块。
对于本申请实施例,所述确定单元31,具体可以用于根据样本数据在各个参数上的数据分布,计算所述样本数据在各个参数上的统计量;根据所述统计量确定所述各个参数对应的统计量折线图,并根据所述统计量折线图中的各个拐点对应的阈值,确定所述各个参数的的候选参数阈值。所述统计量可以为分位数、组内离差或者组间距离中的一种或者多种。
对于本申请实施例,所述预先设定的调优假设包括:
假设1:若待评定的等级有p级,则N
p:N
p-1≈N
p-1:N
p-2≈...≈N
2:N
1,所述N
p表示第p个等级所包含的样本数据量;
假设2:若参数个数为h,则调优后的每个参数的各级参数阈值符合
所述R
hp表示第h个参数的第p个等级的参数阈值所在百分位数;
假设3:调优后的参数阈值组合对应的样本占比约等于调优前的样本占比。
对于本申请实施例,所述调优单元32可以包括:计算模块321、选择模块322和确定模块322,如图4所示。
所述计算模块321,可以用于计算所述给定的参数阈值对应参数阈值组合的样本数据量与总样本数据量的比值,作为调优前的样本占比。所述选择模块322,可以用于根据所述假设3、所述假设2和所述候选参数阈值对应的组合阈值划分表,选择样本占比与所述 调优前的样本占比约等于的且对应候选参数阈值所在百分位数约等于的候选参数阈值组合,作为调优后的第一组参数阈值。所述确定模块323,可以用于根据所述调优后的第一组参数阈值、所述组合阈值划分表和所述假设1,确定调优后的其他组参数阈值。
在具体应用场景中,所述选择模块322可以包括:确定子模块3221和查找子模块3222。
确定子模块3221,可以用于确定所述候选参数阈值对应的组合阈值划分表,所述组合阈值划分表中保存有所述候选参数阈值与其所在百分位数的对应关系,所述各个参数的候选参数阈值组合对应的样本占比,以及所述对应关系与所述样本占比的映射关系。所述查找子模块3222,可以用于根据所述假设3从所述组合阈值划分表中查找与所述调优前的样本占比约等于的样本占比,以及与查找的多个样本占比对应的第一对应关系。所述确定子模块3221,还可以用于根据所述假设2从所述第一对应关系中查找对应百分位数约等于的各个参数的候选参数阈值并根据查找的候选参数阈值,确定调优后的第一组参数阈值。
对于本申请实施例,所述确定单元31,具体可以用于根据所述调优后的第一组参数阈值对应的的样本占比和所述假设1,计算其他组参数阈值对应的样本占比;从所述组合阈值划分表中查找与所述其他组参数阈值对应的样本占比对应的第二对应关系,并根据所述第二对应关系,确定所述其他组参数阈值。
此外,为了确定所述候选参数阈值对应的组合阈值划分表,所述确定子模块3221,具体可以用于根据所述候选参数阈值,确定所述各个参数的候选参数阈值组合;计算所述各个参数的候选参数阈值组合对应的样本数据量与总样本数据量的比值,作为所述所述各个参数的候选参数阈值组合对应的样本占比;建立所述候选参数阈值及其所在百分位数的之间的对应关系,以及所述对比关系与所述样本占比之间的映射关系;根据所述对应关系和所述映射关系,构建所述候选参数阈值对应的组合阈值划分表。
需要说明的是,本申请实施例提供的一种参数阈值确定装置所涉及各功能模块的其他相应描述,可以参考图1所示方法的对应描述,在此不再赘述。
基于上述如图1所示方法,相应的,本申请实施例还提供了一种计算机非易失性可读存储介质,其上存储有计算机可读指令,该计算机可读指令被处理器执行时实现以下步骤:根据样本数据在各个参数上的数据分布,确定所述各个参数的参数候选阈值;利用预先设定的调优假设和所述候选参数阈值对依据业务经验给定的参数阈值进行调优;根据调优后的参数阈值确定等级评定的参数阈值。
基于上述如图1所示方法和如图3所示参数阈值确定装置的实施例,本申请实施例还提供了一种计算机设备的实体结构图,如图5所示,该装置包括:处理器41、存储器42、及存储在存储器42上并可在处理器上运行的计算机可读指令,其中存储器42和处理器41 均设置在总线43上所述处理器41执行所述计算机可读指令时实现以下步骤:根据样本数据在各个参数上的数据分布,确定所述各个参数的参数候选阈值;利用预先设定的调优假设和所述候选参数阈值对依据业务经验给定的参数阈值进行调优;根据调优后的参数阈值确定等级评定的参数阈值。
通过本申请的技术方案,能够根据样本数据在各个参数上的数据分布,确定所述各个参数的参数候选阈值;能够利用预先设定的调优假设和所述候选参数阈值对依据业务经验给定的参数阈值进行调优,并根据调优后的参数阈值确定等级评定的参数阈值,从而能够实现基于数据分布的实际情况对依据业务经验给定的参数阈值进行调优,提升等级评定的参数阈值的精确度,从而能够提升数据等级评定的精确度。
显然,本领域的技术人员应该明白,上述的本申请的各模块或各步骤可以用通用的计算装置来实现,它们可以集中在单个的计算装置上,或者分布在多个计算装置所组成的网络上,可选地,它们可以用计算装置可执行的计算机可读指令代码来实现,从而,可以将它们存储在存储装置中由计算装置来执行,并且在某些情况下,可以以不同于此处的顺序执行所示出或描述的步骤,或者将它们分别制作成各个集成电路模块,或者将它们中的多个模块或步骤制作成单个集成电路模块来实现。这样,本申请不限制于任何特定的硬件和软件结合。
以上所述仅为本申请的优选实施例而已,并不用于限制本申请,对于本领域的技术人员来说,本申请可以有各种更改和变化。凡在本申请的精神和原则之内,所作的任何修改、等同替换、改进等,均应包括在本申请的保护范围之内。
Claims (20)
- 一种参数阈值确定方法,其特征在于,包括:根据样本数据在各个参数上的数据分布,确定所述各个参数的参数候选阈值;利用预先设定的调优假设和所述候选参数阈值对依据业务经验给定的参数阈值进行调优;根据调优后的参数阈值确定等级评定的参数阈值。
- 根据权利要求1所述的方法,其特征在于,所述根据样本数据在各个参数上的数据分布,确定所述各个参数的候选参数阈值,包括:根据样本数据在各个参数上的数据分布,计算所述样本数据在各个参数上的统计量;根据所述统计量确定所述各个参数对应的统计量折线图;根据所述统计量折线图中的各个拐点对应的阈值,确定所述各个参数的候选参数阈值。
- 根据权利要求3所述的方法,其特征在于,所述利用预先设定的调优假设和所述候选参数阈值对依据业务经验给定的参数阈值进行调优,包括:计算所述给定的参数阈值对应参数阈值组合的样本数据量与总样本数据量的比值,作为调优前的样本占比;根据所述假设3、所述假设2和所述候选参数阈值对应的组合阈值划分表,选择样本占比与所述调优前的样本占比约等于的且对应候选参数阈值所在百分位数约等于的候选参数阈值组合,作为调优后的第一组参数阈值;根据所述调优后的第一组参数阈值、所述组合阈值划分表和所述假设1,确定调优后的其他组参数阈值。
- 根据权利要求4所述的方法,其特征在于,所述根据所述假设3、所述假设2和所述候选参数阈值对应的组合阈值划分表,选择样本占比与所述调优前的样本占比约等于的且对应候选参数阈值所在百分位数约等于的候选参数阈值组合,作为调优后的第一组参数阈值,包括:确定所述候选参数阈值对应的组合阈值划分表,所述组合阈值划分表中保存有所述候选参数阈值与其所在百分位数的对应关系,所述各个参数的候选参数阈值组合对应的样本占比,以及所述对应关系与所述样本占比的映射关系;根据所述假设3从所述组合阈值划分表中查找与所述调优前的样本占比约等于的样本占比,以及与查找的多个样本占比对应的第一对应关系;根据所述假设2从所述第一对应关系中查找对应百分位数约等于的各个参数的候选参数阈值并根据查找的候选参数阈值,确定调优后的第一组参数阈值。
- 根据权利要求5所述的方法,其特征在于,所述根据所述调优后的第一组参数阈值、所述组合阈值划分表和所述假设1,确定调优后的其他组参数阈值,包括:根据所述调优后的第一组参数阈值对应的的样本占比和所述假设1,计算其他组参数阈值对应的样本占比;从所述组合阈值划分表中查找与所述其他组参数阈值对应的样本占比对应的第二对应关系,并根据所述第二对应关系,确定所述其他组参数阈值。
- 根据权利要求5所述的方法,其特征在于,所述确定所述候选参数阈值对应的组合阈值划分表,包括:根据所述候选参数阈值,确定所述各个参数的候选参数阈值组合;计算所述各个参数的候选参数阈值组合对应的样本数据量与总样本数据量的比值,作为所述各个参数的候选参数阈值组合对应的样本占比;建立所述候选参数阈值及其所在百分位数的之间的对应关系,以及所述对比关系与所述样本占比之间的映射关系;根据所述对应关系和所述映射关系,构建所述候选参数阈值对应的组合阈值划分表。
- 一种参数阈值确定装置,其特征在于,包括:确定单元,用于根据样本数据在各个参数上的数据分布,确定所述各个参数的参数候选阈值;调优单元,用于利用预先设定的调优假设和所述候选参数阈值对依据业务经验给定的参数阈值进行调优;所述确定单元,用于根据调优后的参数阈值确定等级评定的参数阈值。
- 根据权利要求8所述的装置,其特征在于,所述确定单元,具体用于根据样本数据在各个参数上的数据分布,计算所述样本数据在各个参数上的统计量;根据所述统计量确定所述各个参数对应的统计量折线图,并根据所述统计量折线图中的各个拐点对应的阈值,确定所述各个参数的的候选参数阈值。
- 根据权利要求10所述的装置,其特征在于,所述调优单元包括:计算模块,用于计算所述给定的参数阈值对应参数阈值组合的样本数据量与总样本数据量的比值,作为调优前的样本占比;选择模块,用于根据所述假设3、所述假设2和所述候选参数阈值对应的组合阈值划分表,选择样本占比与所述调优前的样本占比约等于的且对应候选参数阈值所在百分位数约等于的候选参数阈值组合,作为调优后的第一组参数阈值。确定模块,用于根据所述调优后的第一组参数阈值、所述组合阈值划分表和所述假设1,确定调优后的其他组参数阈值。
- 根据权利要求11所述的装置,其特征在于,所述选择模块包括:确定子模块,用于确定所述候选参数阈值对应的组合阈值划分表,所述组合阈值划分表中保存有所述候选参数阈值与其所在百分位数的对应关系,所述各个参数的候选参数阈值组合对应的样本占比,以及所述对应关系与所述样本占比的映射关系;查找子模块,用于根据所述假设3从所述组合阈值划分表中查找与所述调优前的样本占比约等于的样本占比,以及与查找的多个样本占比对应的第一对应关系。确定子模块,还用于根据所述假设2从所述第一对应关系中查找对应百分位数约等于的各个参数的候选参数阈值并根据查找的候选参数阈值,确定调优后的第一组参数阈值。
- 根据权利要求12所述的装置,其特征在于,所述确定单元,具体用于根据所述调优后的第一组参数阈值对应的的样本占比和所述假设1,计算其他组参数阈值对应的样本占比;从所述组合阈值划分表中查找与所述其他组参数阈值对应的样本占比对应的第二对应关系,并根据所述第二对应关系,确定所述其他组参数阈值。
- 根据权利要求12所述的装置,其特征在于,所述确定子模块,具体用于根据所述候选参数阈值,确定所述各个参数的候选参数阈值组合;计算所述各个参数的候选参数阈值组合对应的样本数据量与总样本数据量的比值,作为所述各个参数的候选参数阈值组合对应的样本占比;建立所述候选参数阈值及其所在百分位数的之间的对应关系,以及所述对比关系与所述样本占比之间的映射关系;根据所述对应关系和所述映射关系,构建所述候选参数阈值对应的组合阈值划分表。
- 一种计算机非易失性可读存储介质,其上存储有计算机可读指令,其特征在于,所述计算机可读指令被处理器执行时实现参数阈值确定方法,包括:根据样本数据在各个参数上的数据分布,确定所述各个参数的参数候选阈值;利用预先设定的调优假设和所述候选参数阈值对依据业务经验给定的参数阈值进行调优;根据调优后的参数阈值确定等级评定的参数阈值。
- 根据权利要求15所述的计算机非易失性可读存储介质,其特征在于,所述计算机可读指令被处理器执行时实现所述根据样本数据在各个参数上的数据分布,确定所述各个参数的候选参数阈值,包括:根据样本数据在各个参数上的数据分布,计算所述样本数据在各个参数上的统计量;根据所述统计量确定所述各个参数对应的统计量折线图;根据所述统计量折线图中的各个拐点对应的阈值,确定所述各个参数的候选参数阈值。
- 一种计算机设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机可读指令,其特征在于,所述计算机可读指令被处理器执行时实现参数阈值确定方法,包括:根据样本数据在各个参数上的数据分布,确定所述各个参数的参数候选阈值;利用预先设定的调优假设和所述候选参数阈值对依据业务经验给定的参数阈值进行调优;根据调优后的参数阈值确定等级评定的参数阈值。
- 根据权利要求18所述的计算机设备,其特征在于,所述处理器执行所述计算机可读指令时实现所述根据样本数据在各个参数上的数据分布,确定所述各个参数的候选参数阈值,包括:根据样本数据在各个参数上的数据分布,计算所述样本数据在各个参数上的统计量;根据所述统计量确定所述各个参数对应的统计量折线图;根据所述统计量折线图中的各个拐点对应的阈值,确定所述各个参数的候选参数阈值。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201810797054.3A CN109272340B (zh) | 2018-07-19 | 2018-07-19 | 参数阈值确定方法、装置及计算机存储介质 |
| CN201810797054.3 | 2018-07-19 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020015216A1 true WO2020015216A1 (zh) | 2020-01-23 |
Family
ID=65152893
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2018/111123 Ceased WO2020015216A1 (zh) | 2018-07-19 | 2018-10-21 | 参数阈值确定方法、装置及计算机存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN109272340B (zh) |
| WO (1) | WO2020015216A1 (zh) |
Families Citing this family (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111178595B (zh) * | 2019-12-11 | 2023-03-24 | 深圳平安医疗健康科技服务有限公司 | 项目控制参数生成方法、装置、计算机设备和存储介质 |
| CN112258093B (zh) * | 2020-11-25 | 2024-06-21 | 京东城市(北京)数字科技有限公司 | 风险等级的数据处理方法及装置、存储介质、电子设备 |
| CN114971106B (zh) * | 2021-02-25 | 2025-07-08 | 腾讯科技(深圳)有限公司 | 目标指标值确定方法、装置、计算机设备和存储介质 |
| CN114565231B (zh) * | 2022-02-07 | 2024-07-12 | 三一汽车制造有限公司 | 作业方量确定方法、装置、设备、存储介质及作业机械 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106096226A (zh) * | 2016-05-27 | 2016-11-09 | 腾讯科技(深圳)有限公司 | 一种数据评估方法、装置及服务器 |
| CN107316205A (zh) * | 2017-05-27 | 2017-11-03 | 银联智惠信息服务(上海)有限公司 | 识别持卡人属性的方法、装置、计算机可读介质及系统 |
| CN107316204A (zh) * | 2017-05-27 | 2017-11-03 | 银联智惠信息服务(上海)有限公司 | 识别持卡人属性的方法、装置、计算机可读介质及系统 |
| US20170372229A1 (en) * | 2016-06-22 | 2017-12-28 | Fujitsu Limited | Method and apparatus for managing machine learning process |
Family Cites Families (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN107944156B (zh) * | 2017-11-29 | 2018-11-06 | 中国海洋大学 | 波高阈值的选取方法 |
-
2018
- 2018-07-19 CN CN201810797054.3A patent/CN109272340B/zh active Active
- 2018-10-21 WO PCT/CN2018/111123 patent/WO2020015216A1/zh not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106096226A (zh) * | 2016-05-27 | 2016-11-09 | 腾讯科技(深圳)有限公司 | 一种数据评估方法、装置及服务器 |
| US20170372229A1 (en) * | 2016-06-22 | 2017-12-28 | Fujitsu Limited | Method and apparatus for managing machine learning process |
| CN107316205A (zh) * | 2017-05-27 | 2017-11-03 | 银联智惠信息服务(上海)有限公司 | 识别持卡人属性的方法、装置、计算机可读介质及系统 |
| CN107316204A (zh) * | 2017-05-27 | 2017-11-03 | 银联智惠信息服务(上海)有限公司 | 识别持卡人属性的方法、装置、计算机可读介质及系统 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN109272340B (zh) | 2023-04-18 |
| CN109272340A (zh) | 2019-01-25 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2020015216A1 (zh) | 参数阈值确定方法、装置及计算机存储介质 | |
| CN106844781A (zh) | 数据处理的方法及装置 | |
| JP6243045B2 (ja) | グラフデータクエリ方法および装置 | |
| US20240176657A1 (en) | Task processing method and apparatus, electronic device, storage medium and program product | |
| US20160328445A1 (en) | Data Query Method and Apparatus | |
| CN112639841B (zh) | 用于在多方策略互动中进行策略搜索的采样方案 | |
| CN107688605A (zh) | 跨平台数据匹配方法、装置、计算机设备和存储介质 | |
| WO2020034593A1 (zh) | 人群绩效特征预测中的缺失特征处理方法及装置 | |
| CN108989122A (zh) | 虚拟网络请求映射方法、装置及实现装置 | |
| CN117973545A (zh) | 一种基于大语言模型的推荐方法、装置、设备及存储介质 | |
| CN112001786A (zh) | 基于知识图谱的客户信用卡额度配置方法及装置 | |
| CN114429195A (zh) | 混合专家模型训练的性能优化方法和装置 | |
| CN108304404B (zh) | 一种基于改进的Sketch结构的数据频率估计方法 | |
| TWI768857B (zh) | 模型分類的過程中防止模型竊取的方法及裝置 | |
| CN107656989A (zh) | 云存储系统中基于数据分布感知的近邻查询方法 | |
| CN108228896B (zh) | 一种基于密度的缺失数据填补方法及装置 | |
| US9152688B2 (en) | Summarizing a stream of multidimensional, axis-aligned rectangles | |
| CN113012336A (zh) | 银行业务的排队预约方法及其装置、存储介质和设备 | |
| US20160292300A1 (en) | System and method for fast network queries | |
| CN111260384A (zh) | 服务订单处理方法、装置、电子设备及存储介质 | |
| CN115392206B (zh) | 基于wps/excel快速查询数据方法、装置、设备及存储介质 | |
| WO2024045725A1 (zh) | 用于目标保单的处理方法、电子设备和可读存储介质 | |
| CN107818347A (zh) | Gga数据质量的评定预测方法 | |
| CN116680327A (zh) | 基于产品属性的数据结构化方法、装置、终端及存储介质 | |
| CN103700097A (zh) | 一种背景分割方法及系统 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 18926731 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 18926731 Country of ref document: EP Kind code of ref document: A1 |








