WO2020258918A1 - 一种面向非正态分布水质观测数据的幂变换分析方法 - Google Patents

一种面向非正态分布水质观测数据的幂变换分析方法 Download PDF

Info

Publication number
WO2020258918A1
WO2020258918A1 PCT/CN2020/078258 CN2020078258W WO2020258918A1 WO 2020258918 A1 WO2020258918 A1 WO 2020258918A1 CN 2020078258 W CN2020078258 W CN 2020078258W WO 2020258918 A1 WO2020258918 A1 WO 2020258918A1
Authority
WO
WIPO (PCT)
Prior art keywords
transformation
value
water quality
quality observation
likelihood function
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2020/078258
Other languages
English (en)
French (fr)
Inventor
赵铜铁钢
陈浩玲
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Sun Yat Sen University
Original Assignee
Sun Yat Sen University
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Sun Yat Sen University filed Critical Sun Yat Sen University
Publication of WO2020258918A1 publication Critical patent/WO2020258918A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/20Information retrieval; Database structures therefor; File system structures therefor of structured data, e.g. relational data
    • G06F16/24Querying
    • G06F16/245Query processing
    • G06F16/2458Special types of queries, e.g. statistical queries, fuzzy queries or distributed queries
    • G06F16/2462Approximate or statistical queries
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/20Information retrieval; Database structures therefor; File system structures therefor of structured data, e.g. relational data
    • G06F16/24Querying
    • G06F16/245Query processing
    • G06F16/2458Special types of queries, e.g. statistical queries, fuzzy queries or distributed queries
    • G06F16/2474Sequence data queries, e.g. querying versioned data
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F17/00Digital computing or data processing equipment or methods, specially adapted for specific functions
    • G06F17/10Complex mathematical operations
    • G06F17/15Correlation function computation including computation of convolution operations
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F17/00Digital computing or data processing equipment or methods, specially adapted for specific functions
    • G06F17/10Complex mathematical operations
    • G06F17/18Complex mathematical operations for evaluating statistical data, e.g. average values, frequency distributions, probability functions, regression analysis
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q50/00Information and communication technology [ICT] specially adapted for implementation of business processes of specific business sectors, e.g. utilities or tourism
    • G06Q50/06Energy or water supply
    • YGENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02ATECHNOLOGIES FOR ADAPTATION TO CLIMATE CHANGE
    • Y02A20/00Water conservation; Efficient water supply; Efficient water use
    • Y02A20/152Water filtration

Definitions

  • the invention relates to the technical field of environmental engineering, in particular to a power transformation analysis method for non-normally distributed water quality observation data.
  • the transformation method commonly used in water quality series is logarithmic transformation.
  • some variables are still skewed after logarithmic transformation, especially negative skew data, which will increase their skewness after logarithmic transformation. degree.
  • a single type of transformation is not suitable for all observation variable sequences; and when different transformation methods are selected through the subjective judgment of analysts, The selection criteria are different. It is difficult to choose the most suitable transformation method based on the characteristics of the observed variables of the water plant. As a result, the transformed data cannot meet the requirements of linearity, uniformity of variance and normality required by common data mining and statistical analysis. Some important properties of the data that are lost when these transformed data are used in actual analysis and application affect the analysis effect.
  • the present invention solves the problems of poor data transformation effect due to the selected transformation method being not adapted to the characteristics of the water plant observation variable during data transformation of the existing non-normally distributed water quality observation data, and provides a non-normal distribution oriented Power transformation analysis method of water quality observation data.
  • a power transformation analysis method for non-normally distributed water quality observation data including the following steps:
  • the estimated values of the corresponding parameters include: After normal transformation processing , The mean, standard deviation and transformation parameters of the distribution of water quality observation data;
  • step S1 For the different normal transformation methods described in step S1, respectively calculate the minimum negative log likelihood function value, AIC value and BIC value corresponding to each normal transformation method;
  • the different normal transformation methods in step S1 include identity transformation, logarithmic transformation, Box-Cox transformation and Yeo-Johnson transformation.
  • the estimated parameters described in step S1 adopt the method of maximum likelihood function, and adopt the downhill simplex method to solve.
  • the specific steps of calculating the estimated value of the corresponding parameter after the normal transformation processing of the water quality observation data through the Box-Cox transformation is as follows:
  • the transformation parameter ⁇ is estimated by the maximum likelihood method
  • the density of x is:
  • J( ⁇ ; x) is the transformed Jacobian matrix:
  • the specific steps of calculating the estimated value of the corresponding parameter after the normal transformation of the water quality observation data through the Yeo-Johnson transformation in the step S1 are:
  • the transformation parameter ⁇ is estimated by the maximum likelihood method
  • the density of x is:
  • J( ⁇ ; x) is the transformed Jacobian matrix:
  • sgn( ⁇ ) is a sign function, when the variable x i is positive, the value is 1, and when the variable x i is negative, it is -1, otherwise the value is 0;
  • the calculated minimum negative log likelihood function value, AIC value, and BIC value first select the minimum negative log likelihood function value, AIC value and BIC value that are lower than the original water quality observation data.
  • the normal transformation method corresponding to the parameter value otherwise it is considered that the original water quality observation data meets the normality assumption, and it is not transformed, and this step is ended;
  • the transformation method is the optimal transformation method
  • the minimum negative log likelihood function value is expressed as -L;
  • k is the number of estimated parameters
  • L is the maximum likelihood function value
  • n is the number of water quality observation data.
  • the method of the present invention determines the transformation parameters through the information carried by the water quality observation data itself, sets specific measurement indicators to calculate and compares in a variety of normal transformation methods, thereby selecting the optimal normal transformation method according to the data characteristics of the water quality observation Finally, the optimal transformation method transfers the sequence into a space that obeys or approximately obeys the normal distribution function, and obtains a new sequence corresponding to the original sequence to eliminate possible nonlinearity, heteroscedasticity and non-normality in the data sequence
  • the data is directly transformed by the power transformation method. After the transformation, the variable sequence will not change relative to the original value sequence, so the probability density of a specific value in the variable is not changed.
  • the transformation process converges or diverges from the original sequence Achieve changes in the overall distribution of variables.
  • the method of the present invention can make the transformed data have better normality, is convenient for further data analysis, and solves the problems of poor data transformation effect due to the selected transformation method not adapting to the characteristics of the water plant observation variable.
  • Figure 1 is a general flow chart of the method of the present invention.
  • Figure 2 is a transformation effect diagram of the Box-Cox transformation method used in the present invention under different parameters.
  • Figure 3 is a transformation effect diagram of the Yeo-Johnson transformation method used in the present invention under different parameters.
  • Figure 4 is a Q-Q diagram of the original water quality observation sequence in Example 2.
  • Figure 5 is the Q-Q diagram of the water quality observation sequence after Box-Cox transformation in Example 2.
  • Figure 6 is a Q-Q diagram of the water quality observation sequence after Yeo-Johnson transformation in Example 2.
  • Figure 7 is the Q-Q diagram of the water quality observation sequence after logarithmic transformation in Example 2.
  • Example 8 is a schematic diagram of the original water quality observation sequence in Example 2.
  • Figure 9 is a schematic diagram of the water quality observation sequence after logarithmic transformation in Example 2.
  • Figure 10 is a distribution diagram of the original water quality observation data in Example 2.
  • Fig. 11 is a distribution diagram of water quality observation data after logarithmic transformation in Example 2.
  • Embodiment 12 is a diagram showing the relationship between the water quality observation data after inverse transformation and the original water quality observation data in Embodiment 2.
  • Example 13 is a schematic diagram of comparison between the sequence obtained by inverse transformation of the autoregressive statistical analysis result in Example 2 and the original water quality observation sequence.
  • a power transformation analysis method for non-normally distributed water quality observation data including the following steps:
  • the estimated values of the corresponding parameters include: After normal transformation processing , The mean value, standard deviation, and transformation parameters of the distribution of water quality observation data; among them, in this embodiment 1, different normal transformation methods include identity transformation, logarithmic transformation, Box-Cox transformation and Yeo-Johnson transformation; among them, the estimated parameters are The method of maximum likelihood function, and the downhill simplex method is used to solve;
  • step S1 For the Box-Cox transformation, the specific steps in step S1 for calculating the estimated values of the corresponding parameters after the normal transformation of the water quality observation data through the Box-Cox transformation are:
  • the transformation parameter ⁇ is estimated by the maximum likelihood method
  • the density of x is:
  • J( ⁇ ; x) is the transformed Jacobian matrix:
  • step S1 For the Yeo-Johnson transformation, the specific steps in step S1 for calculating the estimated values of the corresponding parameters after the normal transformation of the water quality observation data through the Yeo-Johnson transformation are:
  • the transformation parameter ⁇ is estimated by the maximum likelihood method
  • the density of x is:
  • J( ⁇ ; x) is the transformed Jacobian matrix:
  • sgn( ⁇ ) is a sign function, when the variable x i is positive, the value is 1, and when the variable x i is negative, it is -1, otherwise the value is 0;
  • step S1 For the different normal transformation methods described in step S1, respectively calculate the minimum negative log likelihood function value, AIC value and BIC value corresponding to each normal transformation method;
  • the calculated minimum negative log-likelihood function value, AIC value and BIC value first select the minimum negative log-likelihood function value, AIC value and BIC value which is lower than the corresponding parameter value of the original water quality observation data.
  • Corresponding normal transformation method otherwise it is considered that the original water quality observation data satisfies the normality assumption, and it is not transformed, and this step is ended;
  • the transformation method is the optimal transformation method
  • the maximum likelihood function value obtained by the method of maximum likelihood function when defining estimated parameters is L
  • k is the number of estimated parameters
  • L is the maximum likelihood function value
  • n is the number of water quality observation data.
  • Example 2 is based on the method of Example 1.
  • the daily observation sequence of chemical oxygen demand (COD) in a sewage treatment plant is used as the experimental data.
  • the length of the water quality observation sequence is 655 days.
  • the relevant statistical parameters are shown in Table 1. ;
  • this embodiment 2 further draws a quantile-quantile diagram (QQ diagram) according to the transformation results.
  • QQ diagram uses a graphical method to identify whether the sample data is close to a normal distribution.
  • the QQ diagram can be used to obtain data distribution information intuitively. Used to assist in determining the transformation effect.
  • the point (x, y) on the QQ chart reflects the quantile of the empirical distribution of one of the sample data and the same quantile of the normal distribution. If the point on the QQ chart is approximately a diagonal straight line, it can be regarded as a data point It is normally distributed.
  • the QQ diagram of the water quality observation sequence after transformation in this embodiment 2 is shown in Figure 4-7, and the transformation parameter estimation results are shown in Table 2.
  • the Log item in the table refers to logarithmic transformation, and nllf refers to the negative log likelihood function value. .
  • the resulting transformation sequence is shown in Figure 9.
  • Figure 10 and Figure 11 the normality of the distribution of the water quality observation sequence after the transformation has been significantly improved.
  • the original water quality observation sequence and the inverse transformation sequence are drawn, as shown in Figure 12.
  • the transformed data is completely consistent with the original data, indicating that the transformation process will not lose the original data information.
  • the time series autoregressive analysis of the transformed sequence shows that the COD sequence of the water plant has significant autocorrelation.
  • the fitting result is inversely transformed.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Data Mining & Analysis (AREA)
  • Mathematical Physics (AREA)
  • Mathematical Analysis (AREA)
  • General Engineering & Computer Science (AREA)
  • Computational Mathematics (AREA)
  • Mathematical Optimization (AREA)
  • Probability & Statistics with Applications (AREA)
  • Databases & Information Systems (AREA)
  • Software Systems (AREA)
  • Pure & Applied Mathematics (AREA)
  • Business, Economics & Management (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Economics (AREA)
  • Algebra (AREA)
  • Fuzzy Systems (AREA)
  • Public Health (AREA)
  • Human Resources & Organizations (AREA)
  • Marketing (AREA)
  • Primary Health Care (AREA)
  • Strategic Management (AREA)
  • Tourism & Hospitality (AREA)
  • General Business, Economics & Management (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Evolutionary Biology (AREA)
  • Operations Research (AREA)
  • General Health & Medical Sciences (AREA)
  • Water Supply & Treatment (AREA)
  • Computing Systems (AREA)
  • Investigating Or Analysing Materials By Optical Means (AREA)
  • Testing And Monitoring For Control Systems (AREA)

Abstract

本发明公开了一种面向非正态分布水质观测数据的幂变换分析方法,首先分别计算通过不同的正态变换方法对水质观测数据进行正态变换处理后对应参数的估计值,设置具体的衡量指标并进行计算和比对,从而根据水质观测的数据特征选择最优的正态变换方法,使得变换后的数据具有更好的正态性,最后经最优变换方法进行正态变换处理后的水质观测数据作为输入数据进行水质观测的统计分析,提升了分析的效果。本发明方法能使得变换后的数据具有更好的正态性,便于进一步的数据分析,解决了由于选择的变换方法不适应水厂观测变量自身特征导致数据的变换效果差等问题。

Description

一种面向非正态分布水质观测数据的幂变换分析方法 技术领域
本发明涉及环境工程技术领域,尤其涉及一种面向非正态分布水质观测数据的幂变换分析方法。
背景技术
水质观测序列的挖掘和统计分析,往往要求数据呈正态分布,而实际操作中,很多原始水质序列不呈正态分布,需要在不丢失信息的前提下进行数据的正态变换。
目前常用于水质序列的变换方法为对数变换,而水厂实际运行中,一些变量经对数变换后仍是偏态分布,尤其是负偏态数据,经对数变换后反而会增加其偏度。同时,由于污水处理厂进出水观测变量多、序列长且分布不一,单一类型的变换并不适用于所有的观测变量序列;而通过分析人员的主观判断对不同的变换方法进行选择时,由于选择标准不一,难以根据水厂观测变量自身特征选择最合适的变换方法,导致变换后的数据也无法满足常用数据挖掘和统计分析所要求的线性、方差齐性和正态性的要求,利用这些变换后的数据进行实际分析应用时丢失数据的一些重要性质,影响了分析效果。
发明内容
本发明为解决现有的非正态分布水质观测数据进行数据变换时,由于选择的变换方法不适应水厂观测变量自身特征导致数据的变换效果差等问题,提供了一种面向非正态分布水质观测数据的幂变换分析方法。
为实现以上发明目的,而采用的技术手段是:
一种面向非正态分布水质观测数据的幂变换分析方法,包括以下步骤:
S1.获取非正态分布的水质观测数据,分别计算通过不同的正态变换方法对水质观测数据进行正态变换处理后对应参数的估计值,对应参数的估计值包括:进行正态变换处理后,水质观测数据分布的均值、标准差以及变换参数;
S2.对于步骤S1中所述的不同的正态变换方法,分别计算每种正态变换方法相对应的最小负对数似然函数值、AIC值及BIC值;
S3.根据计算得到的最小负对数似然函数值、AIC值及BIC值,与预设的选 择标准进行比对,根据比对结果从所述不同的正态变换方法中选择得到最优变换方法;
S4.将经所述最优变换方法进行正态变换处理后的水质观测数据作为输入数据进行水质观测的统计分析,将统计分析得到的结果进行逆变换,从而得到最终的分析结果,降低了分析过程的复杂性,提高了分析的准确性。
上述方案中,首先分别计算通过不同的正态变换方法对水质观测数据进行正态变换处理后对应参数的估计值,设置具体的衡量指标并进行计算和比对,从而根据水质观测的数据特征选择最优的正态变换方法,使得变换后的数据具有更好的正态性,最后经最优变换方法进行正态变换处理后的水质观测数据作为输入数据进行水质观测的统计分析,提升了分析的效果。
优选的,步骤S1中所述不同的正态变换方法包括有恒等变换、对数变换、Box-Cox变换及Yeo-Johnson变换。
优选的,步骤S1中所述的估计参数采用最大似然函数的方法,并采用下山单纯形法进行求解。
优选的,所述步骤S1中计算通过Box-Cox变换对水质观测数据进行正态变换处理后对应参数的估计值的具体步骤为:
定义获取得到的非正态分布的水质观测数据序列为x={x 1,x 2,...,x n},λ为变换参数,y={y 1,y 2,...,y n}为输出序列;
若x中各项均为正数,则Box-Cox变换的函数形式为:
Figure PCTCN2020078258-appb-000001
若x中存在x i≤0,则对整个水质观测数据序列进行平移ε,使x i+ε>0,对应的Box-Cox变换的函数形式如下:
Figure PCTCN2020078258-appb-000002
其中变换参数λ通过最大似然法进行估计;
定义经过变换后,水质观测数据服从均值为μ,方差为σ 2的正态分布,则变换后输出的第i个水质观测数据y i的密度为:
Figure PCTCN2020078258-appb-000003
x的密度为:
Figure PCTCN2020078258-appb-000004
其中,J(λ;x)为变换的雅可比矩阵:
Figure PCTCN2020078258-appb-000005
Figure PCTCN2020078258-appb-000006
若x中各项均为正数,获取对数似然函数为:
Figure PCTCN2020078258-appb-000007
令其中的logσ=s,μ/σ=v,同时去掉常数项(-nlog(2π)/2),得到:
Figure PCTCN2020078258-appb-000008
对上式的对数似然函数取负值,然后采用数值法求解使对数似然函数的函数值最小的参数组合,得到最小负对数似然函数值-L,则最大似然函数值为L;
若x中存在x i≤0,获取对数似然函数为:
Figure PCTCN2020078258-appb-000009
令其中的logσ=s,μ/σ=v,同时去掉常数项(-nlog(2π)/2),得到:
Figure PCTCN2020078258-appb-000010
对上式的对数似然函数取负值,然后采用数值法求解使对数似然函数的函数值最小的参数组合,得到最小负对数似然函数值-L,则最大似然函数值为L。
优选的,所述步骤S1中计算通过Yeo-Johnson变换对水质观测数据进行正 态变换处理后对应参数的估计值的具体步骤为:
定义获取得到的非正态分布的水质观测数据序列为x={x 1,x 2,...,x n},λ为变换参数,y={y 1,y 2,...,y n}为输出序列;
则Yeo-Johnson变换的函数形式为:
Figure PCTCN2020078258-appb-000011
其中变换参数λ通过最大似然法进行估计;
定义经过变换后,水质观测数据服从均值为μ,方差为σ 2的正态分布,则变换后输出的第i个水质观测数据y i的密度为:
Figure PCTCN2020078258-appb-000012
x的密度为:
Figure PCTCN2020078258-appb-000013
其中,J(λ;x)为变换的雅可比矩阵:
Figure PCTCN2020078258-appb-000014
Figure PCTCN2020078258-appb-000015
Figure PCTCN2020078258-appb-000016
获取对数似然函数为:
Figure PCTCN2020078258-appb-000017
其中,sgn(·)为符号函数,当其中的变量x i为正时取值为1,当其中的变量取值x i为负时为-1,否则取值为0;
令其中的logσ=s,μ/σ=v,同时去掉常数项(-nlog(2π)/2)后,得到:
Figure PCTCN2020078258-appb-000018
对上式的对数似然函数取负值,然后采用数值法求解使对数似然函数的函数值最小的参数组合,得到最小负对数似然函数值-L,则最大似然函数值为L。
优选的,根据计算得到的最小负对数似然函数值、AIC值及BIC值,首先选出最小负对数似然函数值、AIC值及BIC值三者同时低于原始水质观测数据的对应参数值所对应的正态变换方法,否则认为原始的水质观测数据满足正态性假设,不对其进行变换,结束本步骤;
若最小负对数似然函数值、AIC值及BIC值三者同时低于原始水质观测数据的对应参数值所对应的正态变换方法有多个,则其中最低的BIC值所对应的正态变换方法为最优变换方法;
其中最小负对数似然函数值表示为-L;
AIC值表示为:AIC=2k-2ln(L);
其中k是估计的参数数量,L是最大似然函数值;
BIC值表示为:BIC=ln(n)k-2ln(L);
其中k是估计的参数数量,L是最大似然函数值,n为水质观测数据的个数。
与现有技术相比,本发明技术方案的有益效果是:
本发明方法通过对水质观测数据自身的携带信息确定变换参数,设置具体的衡量指标在多种正态变换方法中进行计算和比对,从而根据水质观测的数据特征选择最优的正态变换方法,最后经最优变换方法将序列转入一个服从或近似服从正态分布函数的空间内,得到与原序列相应的新序列,以排除数据序列中可能的非线性、异方差性和非正态性;通过幂变换方法直接对数据进行变换,变换后变量序列相对于原始值的序列不会改变,也就没有改变变量中某个特定值的概率密度,变换过程通过将原序列进行收敛或发散实现变量整体分布的改变。本发明方法能使得变换后的数据具有更好的正态性,便于进一步的数据分析,解决了由于 选择的变换方法不适应水厂观测变量自身特征导致数据的变换效果差等问题。
附图说明
图1为本发明方法的总流程图。
图2为本发明中使用的Box-Cox变换方法在不同参数下的变换效果图。
图3为本发明中使用的Yeo-Johnson变换方法在不同参数下的变换效果图。
图4为实施例2中原始水质观测序列的Q-Q图。
图5为实施例2中经Box-Cox变换后水质观测序列的Q-Q图。
图6为实施例2中经Yeo-Johnson变换后水质观测序列的Q-Q图。
图7为实施例2中经对数变换后水质观测序列的Q-Q图。
图8为实施例2中原始水质观测序列的示意图。
图9为实施例2中经对数变换后水质观测序列的示意图。
图10为实施例2中原始水质观测数据的分布图。
图11为实施例2中经对数变换后水质观测数据分布图。
图12为实施例2中逆变换后的水质观测数据与原始水质观测数据的关系图。
图13为实施例2中自回归统计分析结果经逆变换所得序列与原始水质观测序列对比示意图。
具体实施方式
附图仅用于示例性说明,不能理解为对本专利的限制;
为了更好说明本实施例,附图某些部件会有省略、放大或缩小,并不代表实际产品的尺寸;
对于本领域技术人员来说,附图中某些公知结构及其说明可能省略是可以理解的。
下面结合附图和实施例对本发明的技术方案做进一步的说明。
实施例1
一种面向非正态分布水质观测数据的幂变换分析方法,包括以下步骤:
S1.获取非正态分布的水质观测数据,分别计算通过不同的正态变换方法对水质观测数据进行正态变换处理后对应参数的估计值,对应参数的估计值包括:进行正态变换处理后,水质观测数据分布的均值、标准差以及变换参数;其中本实施例1中,不同的正态变换方法包括有恒等变换、对数变换、Box-Cox变换及Yeo-Johnson变换;其中估计参数采用最大似然函数的方法,并采用下山单纯形 法进行求解;
对于Box-Cox变换,步骤S1中计算通过Box-Cox变换对水质观测数据进行正态变换处理后对应参数的估计值的具体步骤为:
定义获取得到的非正态分布的水质观测数据序列为x={x 1,x 2,...,x n},λ为变换参数,y={y 1,y 2,...,y n}为输出序列;
若x中各项均为正数,则Box-Cox变换的函数形式为:
Figure PCTCN2020078258-appb-000019
若x中存在x i≤0,则对整个水质观测数据序列进行平移ε,使x i+ε>0,对应的Box-Cox变换的函数形式如下:
Figure PCTCN2020078258-appb-000020
其中变换参数λ通过最大似然法进行估计;
定义经过变换后,水质观测数据服从均值为μ,方差为σ 2的正态分布,则变换后输出的第i个水质观测数据y i的密度为:
Figure PCTCN2020078258-appb-000021
x的密度为:
Figure PCTCN2020078258-appb-000022
其中,J(λ;x)为变换的雅可比矩阵:
Figure PCTCN2020078258-appb-000023
Figure PCTCN2020078258-appb-000024
若x中各项均为正数,获取对数似然函数为:
Figure PCTCN2020078258-appb-000025
令其中的logσ=s,μ/σ=v,同时去掉常数项(-nlog(2π)/2),得到:
Figure PCTCN2020078258-appb-000026
对上式的对数似然函数取负值,然后采用数值法求解使对数似然函数的函数值最小的参数组合,得到最小负对数似然函数值-L,则最大似然函数值为L;
若x中存在x i≤0,获取对数似然函数为:
Figure PCTCN2020078258-appb-000027
令其中的logσ=s,μ/σ=v,同时去掉常数项(-nlog(2π)/2),得到:
Figure PCTCN2020078258-appb-000028
对上式的对数似然函数取负值,然后采用数值法求解使对数似然函数的函数值最小的参数组合,得到最小负对数似然函数值-L,则最大似然函数值为L。
对于Yeo-Johnson变换,步骤S1中计算通过Yeo-Johnson变换对水质观测数据进行正态变换处理后对应参数的估计值的具体步骤为:
定义获取得到的非正态分布的水质观测数据序列为x={x 1,x 2,...,x n},λ为变换参数,y={y 1,y 2,...,y n}为输出序列;
则Yeo-Johnson变换的函数形式为:
Figure PCTCN2020078258-appb-000029
其中变换参数λ通过最大似然法进行估计;
定义经过变换后,水质观测数据服从均值为μ,方差为σ 2的正态分布,则变换后输出的第i个水质观测数据y i的密度为:
Figure PCTCN2020078258-appb-000030
x的密度为:
Figure PCTCN2020078258-appb-000031
其中,J(λ;x)为变换的雅可比矩阵:
Figure PCTCN2020078258-appb-000032
Figure PCTCN2020078258-appb-000033
Figure PCTCN2020078258-appb-000034
获取对数似然函数为:
Figure PCTCN2020078258-appb-000035
其中,sgn(·)为符号函数,当其中的变量x i为正时取值为1,当其中的变量取值x i为负时为-1,否则取值为0;
令其中的logσ=s,μ/σ=v,同时去掉常数项(-nlog(2π)/2)后,得到:
Figure PCTCN2020078258-appb-000036
对上式的对数似然函数取负值,然后采用数值法求解使对数似然函数的函数值最小的参数组合,得到最小负对数似然函数值-L,则最大似然函数值为L。
S2.对于步骤S1中所述的不同的正态变换方法,分别计算每种正态变换方法相对应的最小负对数似然函数值、AIC值及BIC值;
S3.根据计算得到的最小负对数似然函数值、AIC值及BIC值,与预设的选择标准进行比对,根据比对结果从所述不同的正态变换方法中选择得到最优变换方法;具体如下:
根据计算得到的最小负对数似然函数值、AIC值及BIC值,首先选出最小 负对数似然函数值、AIC值及BIC值三者同时低于原始水质观测数据的对应参数值所对应的正态变换方法,否则认为原始的水质观测数据满足正态性假设,不对其进行变换,结束本步骤;
若最小负对数似然函数值、AIC值及BIC值三者同时低于原始水质观测数据的对应参数值所对应的正态变换方法有多个,则其中最低的BIC值所对应的正态变换方法为最优变换方法;
定义估计参数时采用最大似然函数的方法获得的最大似然函数值为L,
则负对数似然函数值表示为-L
AIC值表示为:AIC=2k-2ln(L)
其中k是估计的参数数量,L是最大似然函数值;
BIC值表示为:BIC=ln(n)k-2ln(L)
其中k是估计的参数数量,L是最大似然函数值,n为水质观测数据的个数。
S4.将经所述最优变换方法进行正态变换处理后的水质观测数据作为输入数据进行水质观测的统计分析,将统计分析得到的结果进行逆变换,从而得到最终的分析结果;记统计分析后获得的序列为z={z 1,...,z m},其逆变换序列为
Figure PCTCN2020078258-appb-000037
其中对于经Box-Cox变换后的水质观测数据的逆变换形式如下:
对于Box-Cox单参数变换,即变换时没有对水质观测数据进行平移:
Figure PCTCN2020078258-appb-000038
其中
Figure PCTCN2020078258-appb-000039
为逆变换后的水质数据,z i为经Box-Cox变换后的水质数据,λ为变换参数;
对于Box-Cox双参数变换,即变换时对整个水质观测数据序列平移了ε:
Figure PCTCN2020078258-appb-000040
其中
Figure PCTCN2020078258-appb-000041
为逆变换后的水质数据,z i为经Box-Cox变换后的水质数据,λ为变换参数;
其中对于经Yeo-Johnson变换后的水质观测数据的逆变换形式如下:
Figure PCTCN2020078258-appb-000042
其中
Figure PCTCN2020078258-appb-000043
为逆变换后的水质数据,z i为经Yeo-Johnson变换后的水质数据,λ为变换参数。
不同参数的Box-Cox变换及Yeo-Johnson变换效果分别如图2、3所示,两者对变量偏度的改变明显,在一定程度上甚至会改变偏移的方向。本发明采取对观测变量同时进行不同变换,选取合适的变换方法,获得便于实际统计分析计算的变换结果。
实施例2
本实施例2基于实施例1的方法进行实验,以某污水处理厂入水化学需氧量(COD)日观测序列为实验数据,该水质观测序列长度为655日,相关统计参数如表1所示;
Figure PCTCN2020078258-appb-000044
表1
根据表1对数正态分布及正态分布K-S检验pvalue,可以认为该序列服从对数正态分布,作为输入的非正态分布的水质观测数据水观测序列,分别对这些数据进行恒等、Box-Cox、Yeo-Johnson和对数变换参数估计,获得不同的变换结果。另外本实施例2还进一步根据变换结果绘制quantile-quantile图(Q-Q图),Q-Q图采用图形的方法鉴别样本数据是否近似于正态分布,通过Q-Q图可以比较直观的获取数据分布信息,其主要用于辅助判断变换效果。Q-Q图上的点(x,y)反映出其中一个样本数据的经验分布的分位数和正态分布的相同分位数,若Q-Q图上的点近似一条对角直线,则可认为数据点呈正态分布。本实施例2变换后水质观测序列的Q-Q图如图4-7所示,变换参数估计结果如表2所示;其中表格中的Log项指对数变换,nllf指负对数似然函数值。
Parameter Identity Log Box-Cox Yeo-Johnson
λ / / 0.57 0.04
μ 398.41 5.89 50.91 6.72
σ 182.82 0.45 14.06 0.58
nllf 3,739.07 3,663.61 3,687.02 3,663.46
AIC 7,478.14 7,327.23 7,376.05 7,328.92
BIC 7,478.14 7,327.23 7,380.53 7,333.41
表2
由图4-7和表2可以得出,相比于原始的水质观测数据,即表格中的Identity项,经Box-Cox、Yeo-Johnson和对数变换后,数据的负对数似然函数值、AIC和BIC值均有所降低,其中对数变换的BIC值是所有变换方法中最低的,为7327.23,故采用对数变换作为最优变换方法对水质观测数据进行变换。
采用对数变换作为最优变换方法,对原始水质观测序列进行变换,得到的变换序列如图9所示,与图8的原始水质观测序列相比,整体波动更为平稳。通过图10和图11得到,变换后水质观测序列分布正态性得到明显提高,直接对变换序列进行逆变换后,绘制原始水质观测序列与逆变换序列散点图,如图12所示,逆变换后的数据与原始数据完全一致,说明变换过程并不会丢失原始数据信息。对变换后序列进行时间序列自回归分析,发现该水厂COD序列存在显著的自相关性,对该拟合结果进行逆变换,所得逆变换序列与原始序列对比结果如图13所示;从图13可以看到,自回归拟合序列结果基本消除了原始序列中存在的噪声,同时能够较好地概括原始序列的存在的变化趋势,利用该逆变换序列,可以进一步预测该水厂未来入水COD的变化趋势。
附图中描述位置关系的用语仅用于示例性说明,不能理解为对本专利的限制;
显然,本发明的上述实施例仅仅是为清楚地说明本发明所作的举例,而并非是对本发明的实施方式的限定。对于所属领域的普通技术人员来说,在上述说明的基础上还可以做出其它不同形式的变化或变动。这里无需也无法对所有的实施方式予以穷举。凡在本发明的精神和原则之内所作的任何修改、等同替换和改进等,均应包含在本发明权利要求的保护范围之内。

Claims (6)

  1. 一种面向非正态分布水质观测数据的幂变换分析方法,其特征在于:包括以下步骤:
    S1.获取非正态分布的水质观测数据,分别计算通过不同的正态变换方法对水质观测数据进行正态变换处理后对应参数的估计值,对应参数的估计值包括:进行正态变换处理后,水质观测数据分布的均值、标准差以及变换参数;
    S2.对于步骤S1中所述的不同的正态变换方法,分别计算每种正态变换方法相对应的最小负对数似然函数值、AIC值及BIC值;
    S3.根据计算得到的最小负对数似然函数值、AIC值及BIC值,与预设的选择标准进行比对,根据比对结果从所述不同的正态变换方法中选择得到最优变换方法;
    S4.将经所述最优变换方法进行正态变换处理后的水质观测数据作为输入数据进行水质观测的统计分析,将统计分析得到的结果进行逆变换,从而得到最终的分析结果。
  2. 根据权利要求1所述的面向非正态分布水质观测数据的幂变换分析方法,其特征在于,步骤S1中所述不同的正态变换方法包括有恒等变换、对数变换、Box-Cox变换及Yeo-Johnson变换。
  3. 根据权利要求2所述的面向非正态分布水质观测数据的幂变换分析方法,其特征在于,步骤S1中所述的估计参数采用最大似然函数的方法,并采用下山单纯形法进行求解。
  4. 根据权利要求3所述的面向非正态分布水质观测数据的幂变换分析方法,其特征在于,所述步骤S1中计算通过Box-Cox变换对水质观测数据进行正态变换处理后对应参数的估计值的具体步骤为:
    定义获取得到的非正态分布的水质观测数据序列为x={x 1,x 2,...,x n},λ为变换参数,y={y 1,y 2,...,y n}为输出序列;
    若x中各项均为正数,则Box-Cox变换的函数形式为:
    Figure PCTCN2020078258-appb-100001
    若x中存在x i≤0,则对整个水质观测数据序列进行平移ε,使x i+ε>0,对应的Box-Cox变换的函数形式如下:
    Figure PCTCN2020078258-appb-100002
    其中变换参数λ通过最大似然法进行估计;
    定义经过变换后,水质观测数据服从均值为μ,方差为σ 2的正态分布,则变换后输出的第i个水质观测数据y i的密度为:
    Figure PCTCN2020078258-appb-100003
    x的密度为:
    Figure PCTCN2020078258-appb-100004
    其中,J(λ;x)为变换的雅可比矩阵:
    Figure PCTCN2020078258-appb-100005
    Figure PCTCN2020078258-appb-100006
    若x中各项均为正数,获取对数似然函数为:
    Figure PCTCN2020078258-appb-100007
    令其中的logσ=s,μ/σ=v,同时去掉常数项(-n log(2π)/2),得到:
    Figure PCTCN2020078258-appb-100008
    对上式的对数似然函数取负值,然后采用数值法求解使对数似然函数的函数值最小的参数组合,得到最小负对数似然函数值-L,则最大似然函数值为L;
    若x中存在x i≤0,获取对数似然函数为:
    Figure PCTCN2020078258-appb-100009
    令其中的logσ=s,μ/σ=v,同时去掉常数项(-n log(2π)/2),得到:
    Figure PCTCN2020078258-appb-100010
    对上式的对数似然函数取负值,然后采用数值法求解使对数似然函数的函数值最小的参数组合,得到最小负对数似然函数值-L,则最大似然函数值为L。
  5. 根据权利要求3所述的面向非正态分布水质观测数据的幂变换分析方法,其特征在于,所述步骤S1中计算通过Yeo-Johnson变换对水质观测数据进行正态变换处理后对应参数的估计值的具体步骤为:
    定义获取得到的非正态分布的水质观测数据序列为x={x 1,x 2,...,x n},λ为变换参数,y={y 1,y 2,...,y n}为输出序列;
    则Yeo-Johnson变换的函数形式为:
    Figure PCTCN2020078258-appb-100011
    其中变换参数λ通过最大似然法进行估计;
    定义经过变换后,水质观测数据服从均值为μ,方差为σ 2的正态分布,则变换后输出的第i个水质观测数据y i的密度为:
    Figure PCTCN2020078258-appb-100012
    x的密度为:
    Figure PCTCN2020078258-appb-100013
    其中,J(λ;x)为变换的雅可比矩阵:
    Figure PCTCN2020078258-appb-100014
    Figure PCTCN2020078258-appb-100015
    Figure PCTCN2020078258-appb-100016
    获取对数似然函数为:
    Figure PCTCN2020078258-appb-100017
    其中,sgn(·)为符号函数,当其中的变量x i为正时取值为1,当其中的变量取值x i为负时为-1,否则取值为0;
    令其中的logσ=s,μ/σ=v,同时去掉常数项(-n log(2π)/2)后,得到:
    Figure PCTCN2020078258-appb-100018
    对上式的对数似然函数取负值,然后采用数值法求解使对数似然函数的函数值最小的参数组合,得到最小负对数似然函数值-L,则最大似然函数值为L。
  6. 根据权利要求1~5任一项所述的面向非正态分布水质观测数据的幂变换分析方法,其特征在于,根据计算得到的最小负对数似然函数值、AIC值及BIC值,首先选出最小负对数似然函数值、AIC值及BIC值三者同时低于原始水质观测数据的对应参数值所对应的正态变换方法,否则认为原始的水质观测数据满足正态性假设,不对其进行变换,结束本步骤;
    若最小负对数似然函数值、AIC值及BIC值三者同时低于原始水质观测数据的对应参数值所对应的正态变换方法有多个,则其中最低的BIC值所对应的正态变换方法为最优变换方法;
    其中最小负对数似然函数值表示为-L;
    AIC值表示为:AIC=2k-2ln(L);
    其中k是估计的参数数量,L是最大似然函数值;
    BIC值表示为:BIC=ln(n)k-2ln(L);
    其中k是估计的参数数量,L是最大似然函数值,n为水质观测数据的个数。
PCT/CN2020/078258 2019-06-24 2020-03-06 一种面向非正态分布水质观测数据的幂变换分析方法 Ceased WO2020258918A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201910550362.0A CN110309199B (zh) 2019-06-24 2019-06-24 一种面向非正态分布水质观测数据的幂变换分析方法
CN201910550362.0 2019-06-24

Publications (1)

Publication Number Publication Date
WO2020258918A1 true WO2020258918A1 (zh) 2020-12-30

Family

ID=68076514

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2020/078258 Ceased WO2020258918A1 (zh) 2019-06-24 2020-03-06 一种面向非正态分布水质观测数据的幂变换分析方法

Country Status (2)

Country Link
CN (1) CN110309199B (zh)
WO (1) WO2020258918A1 (zh)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114239294A (zh) * 2021-12-21 2022-03-25 中国人民解放军国防科技大学 基于原点矩偏导的k分布杂波参数估计方法和装置
CN114626008A (zh) * 2022-03-15 2022-06-14 中铁二院工程集团有限责任公司 一种基于幂相关随机过程的铁路路基沉降预测方法和装置

Families Citing this family (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN110309199B (zh) * 2019-06-24 2021-09-28 中山大学 一种面向非正态分布水质观测数据的幂变换分析方法
CN111259554B (zh) * 2020-01-20 2022-03-15 山东大学 推土机变矩变速装置螺栓装配大数据检测方法及系统
CN116602435A (zh) * 2023-07-03 2023-08-18 上海神齐信息科技有限公司 一种基于机器学习的制丝单机内烟丝水分变化分析方法
CN116955993B (zh) * 2023-08-24 2024-03-12 中国长江电力股份有限公司 一种混凝土性态多元时序监测数据补全方法
CN118230850B (zh) * 2024-05-13 2025-07-22 农业农村部环境保护科研监测所 一种机器学习驱动的金属改性生物炭除磷性能预测方法

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20090157351A1 (en) * 2007-12-12 2009-06-18 Xerox Corporation System and method for improving print shop operability
CN102855757A (zh) * 2012-03-05 2013-01-02 浙江大学 基于排队检测器信息瓶颈状态识别方法
CN104899419A (zh) * 2015-04-28 2015-09-09 清华大学 一种淡水水体中氮和/或磷含量检测的方法
CN110309199A (zh) * 2019-06-24 2019-10-08 中山大学 一种面向非正态分布水质观测数据的幂变换分析方法

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20090157351A1 (en) * 2007-12-12 2009-06-18 Xerox Corporation System and method for improving print shop operability
CN102855757A (zh) * 2012-03-05 2013-01-02 浙江大学 基于排队检测器信息瓶颈状态识别方法
CN104899419A (zh) * 2015-04-28 2015-09-09 清华大学 一种淡水水体中氮和/或磷含量检测的方法
CN110309199A (zh) * 2019-06-24 2019-10-08 中山大学 一种面向非正态分布水质观测数据的幂变换分析方法

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114239294A (zh) * 2021-12-21 2022-03-25 中国人民解放军国防科技大学 基于原点矩偏导的k分布杂波参数估计方法和装置
CN114626008A (zh) * 2022-03-15 2022-06-14 中铁二院工程集团有限责任公司 一种基于幂相关随机过程的铁路路基沉降预测方法和装置
CN114626008B (zh) * 2022-03-15 2023-03-21 中铁二院工程集团有限责任公司 一种基于幂相关随机过程的铁路路基沉降预测方法和装置

Also Published As

Publication number Publication date
CN110309199A (zh) 2019-10-08
CN110309199B (zh) 2021-09-28

Similar Documents

Publication Publication Date Title
WO2020258918A1 (zh) 一种面向非正态分布水质观测数据的幂变换分析方法
Albers et al. Estimation in Shewhart control charts: effects and corrections
Houpt et al. The statistical properties of the survivor interaction contrast
Almanidis et al. The skewness issue in stochastic frontiers models: Fact or fiction?
CN105676817A (zh) 不同大小样本均值-标准偏差控制图的统计过程控制方法
CN110442911B (zh) 一种基于统计机器学习的高维复杂系统不确定性分析方法
CN104792350B (zh) 一种大坝监测自动化比测方法
CN103971022B (zh) 基于t2控制图的飞机零部件质量稳定性控制算法
CN104267610A (zh) 高精度的高炉冶炼过程异常数据检测及修补方法
Ademola et al. Exponentiated gompertz exponential (egoe) distribution: Derivation, properties and applications
CN108519760A (zh) 一种基于变点检测理论的制丝过程稳态识别方法
CN116738222B (zh) 多晶硅生产设备数据预测方法、装置、服务器及存储介质
Cangul et al. Testing treatment effects in unconfounded studies under model misspecification: Logistic regression, discretization, and their combination
CN103995985B (zh) 基于Daubechies小波变换和弹性网的故障检测方法
CN111651876B (zh) 工业循环冷却水腐蚀状况在线分析方法及检测系统
Afify et al. A new skewed discrete model: properties, inference, and applications
CN109887253A (zh) 石油化工装置报警的关联分析方法
CN110826836A (zh) 确定评价分值的方法和装置
CN112257958A (zh) 一种电力饱和负荷预测方法及装置
CN117951695A (zh) 一种工业未知威胁检测方法及系统
Tailor Exponential weighted moving average (EWMA) chart under the assumption of moderateness and its 3Δ control limits
CN111583990B (zh) 一种结合稀疏回归和淘汰规则的基因调控网络推断方法
CN115310260B (zh) 疲劳寿命分布模型建模方法、系统、装置、计算机可读介质
Hanna Some information measures for testing stochastic models
Habiger A method for modifying multiple testing procedures

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20831983

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 20831983

Country of ref document: EP

Kind code of ref document: A1

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205 DATED 15/09/2022)

122 Ep: pct application non-entry in european phase

Ref document number: 20831983

Country of ref document: EP

Kind code of ref document: A1