WO2017185576A1 - 一种多流流式数据的处理方法、系统、存储介质及设备 - Google Patents

一种多流流式数据的处理方法、系统、存储介质及设备 Download PDF

Info

Publication number
WO2017185576A1
WO2017185576A1 PCT/CN2016/097416 CN2016097416W WO2017185576A1 WO 2017185576 A1 WO2017185576 A1 WO 2017185576A1 CN 2016097416 W CN2016097416 W CN 2016097416W WO 2017185576 A1 WO2017185576 A1 WO 2017185576A1
Authority
WO
WIPO (PCT)
Prior art keywords
data
calculation
window
stream
semantic
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2016/097416
Other languages
English (en)
French (fr)
Inventor
项连志
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Baidu Online Network Technology Beijing Co Ltd
Original Assignee
Baidu Online Network Technology Beijing Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Baidu Online Network Technology Beijing Co Ltd filed Critical Baidu Online Network Technology Beijing Co Ltd
Publication of WO2017185576A1 publication Critical patent/WO2017185576A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00Arrangements for program control, e.g. control units
    • G06F9/06Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/46Multiprogramming arrangements
    • G06F9/466Transaction processing
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/90Details of database functions independent of the retrieved data types
    • G06F16/95Retrieval from the web
    • G06F16/951Indexing; Web crawling techniques

Definitions

  • the embodiments of the present invention relate to a big data processing technology, and in particular, to a multi-streaming data processing method, system, storage medium, and device.
  • the embodiment of the invention provides a multi-streaming data processing method, system, storage medium and device, so as to realize real-time processing of massive multi-stream data.
  • an embodiment of the present invention provides a method for processing multi-streaming data, where the method includes:
  • the semantic calculation of the obtained window calculation result is performed to obtain a semantic structure recombination data stream
  • the semantic flow recombination data stream is subjected to a recombination flow calculation to obtain a data processing result.
  • an embodiment of the present invention provides a multi-streaming data processing system, where the system includes:
  • a windowing processing module configured to perform window processing on the acquired multi-stream streaming data
  • a window calculation module configured to perform window calculation on the windowed multi-stream stream data based on a window calculation formula
  • the data stream recombining module is configured to perform semantic structure recombination on the obtained window calculation result according to the set data processing purpose, to obtain a semantic structure recombination data stream;
  • a recombination flow calculation module configured to recompose the data stream based on the recombination flow calculation formula The recombination flow calculation is performed to obtain the data processing result.
  • an embodiment of the present invention further provides a storage medium including computer executable instructions for executing a multi-stream streaming data processing method when executed by a computer processor, Methods include:
  • the semantic calculation of the obtained window calculation result is performed to obtain a semantic structure recombination data stream
  • the semantic flow recombination data stream is subjected to a recombination flow calculation to obtain a data processing result.
  • the embodiment of the present invention further provides a device, where the device includes:
  • One or more processors are One or more processors;
  • One or more programs the one or more programs being stored in the memory, and when executed by the one or more processors, performing the following operations:
  • the semantic calculation of the obtained window calculation result is performed to obtain a semantic structure recombination data stream
  • the semantic flow recombination data stream is subjected to a recombination flow calculation to obtain a data processing result.
  • the present invention provides a multi-streaming data processing method, system, storage medium and device.
  • the processing method firstly performs window processing on the data in the data stream after obtaining the data stream.
  • the window calculation is performed on the window calculation formula; then the obtained window calculation result is reconstructed based on the data processing purpose to form the semantic structure reorganization flow; then the semantic structure reorganization flow is calculated based on the recombination flow calculation formula to obtain the data processing result.
  • the processing method effectively realizes multi-stream flow calculation involving massive entity data in real time, thereby not only solving the problem of long time in big data processing and high cost of abnormal processing, but also solving the problem that multi-data stream cannot be processed.
  • the calculation problem further improves the implementation process of processing and calculating massive data streams, thus satisfying users' demands for real-time exploration of deep information or reasons based on data processing results.
  • FIG. 1 is a flowchart of a method for processing multi-streaming data according to Embodiment 1 of the present invention
  • FIG. 2 is a flowchart of a method for processing multi-streaming data according to Embodiment 2 of the present invention
  • FIG. 3 is a flowchart of a method for processing multi-streaming data according to Embodiment 3 of the present invention.
  • FIG. 4 is a flowchart of a method for processing multi-streaming data according to Embodiment 4 of the present invention.
  • FIG. 5a is a schematic diagram of a preferred embodiment of a method for processing multi-streaming data according to Embodiment 5 of the present invention.
  • FIG. 5b is a schematic diagram of an application effect of a multi-streaming data processing method in keyword consumption diagnosis according to Embodiment 5 of the present invention.
  • FIG. 6a is a structural block diagram of a multi-streaming data processing system according to Embodiment 6 of the present invention.
  • FIG. 6b is a structural diagram of a multi-streaming data processing according to Embodiment 6 of the present invention.
  • FIG. 7 is a schematic structural diagram of a device according to Embodiment 7 of the present invention.
  • FIG. 1 is a flowchart of a method for processing multi-streaming data according to Embodiment 1 of the present invention.
  • the method is applicable to data processing of massive multi-stream data, and may be performed by multi-stream data.
  • Processing system execution wherein the processing system can be implemented by software and/or hardware, and can generally be integrated on a framework platform for streaming data analysis, which can be a server.
  • a method for processing multi-streaming data according to Embodiment 1 of the present invention specifically includes:
  • the multi-streaming data can be specifically understood as streaming data from different data streams that acts on the same data processing calculation.
  • input of a data stream is required.
  • input processes of different data streams are independent of each other.
  • the data stream in this embodiment can generally be regarded as an unordered data stream, that is, the input timing of the data stream is different from its service timing.
  • the purpose of the windowing process is to divide the data in the acquired data stream into a data calculation window.
  • the data calculation window may refer to a data set having the same data tag (the data tag is primarily determined by one or more attribute values of the data).
  • the value can determine its own data mark.
  • the data calculation window to which the data belongs can be determined. After determining the data calculation window to which the data belongs, the data in the data stream can be divided into corresponding data calculations. window.
  • the window calculation may be performed on the data in the data calculation window in real time based on the window calculation formula, and the window calculation result is output.
  • the window calculation formula can be specifically understood as a formula according to the calculation and calculation of the divided data in the data calculation window.
  • common window calculation formulas include: cumulative summation of specified attributes of data, maximum calculation of specified attributes of data, minimum calculation of specified attributes of data, or calculation of specified attributes of data. Refers to calculations, etc.
  • the output window calculation result can be stored in the cache for use in subsequent calculations.
  • the process may be regarded as separately processing a plurality of data streams separately, so as to reorganize the results of the separate processing calculations, and finally based on the settings.
  • Multi-flow calculation formula realizes data processing of multi-stream data.
  • the process of performing a single stream processing operation on the acquired one data stream is: the acquired data stream is divided into different data calculation windows based on the time attribute value (at this time, the data is
  • the calculation window can be simply referred to as time window); and the time window can be divided based on the time point to realize the windowing of the data (common time point division can be 0 to 1 point, 1 to 2 points, ..., 23 points) ⁇ 0 points division); in any time window (such as 1 to 2 points), based on the set window calculation formula (such as cumulative summation), a certain attribute (such as consumption attribute) of the data is calculated in real time.
  • the corresponding value in the time window (such as the cumulative value of consumption).
  • S120 Perform semantic structure recombination on the obtained window calculation result according to the set data processing purpose to obtain a semantic structure recombination data stream.
  • S110 only performs the calculation of separately processing each data stream in the multi-streaming data processing.
  • the window calculation result of the multi-data stream needs to be based on the specific flow form.
  • the conditions are merged and recombined to get a data stream based on a specific semantic structure.
  • the data processing purpose can be specifically understood as the goal of the semantic structure reorganization phase of multi-stream data processing.
  • the data processing purpose generally includes processing parameters required for the semantic structure reorganization operation and processing information obtained based on the user input instruction, wherein the processing parameter can be specifically understood as a practical application scenario required for synthesizing the semantic structure reorganization flow.
  • the relevant context dynamically obtained in the external environment; the processing information may be determined in a pre-processing stage of the multi-streaming data processing, and the pre-processing stage may be specifically understood as analyzing and disassembling the instructions input by the user in the actual application scenario. process.
  • the window calculation result stored in the cache After obtaining the window calculation result stored in the cache, first determining the processing information in the data processing purpose, and then determining the data required for the semantic structure reorganization based on the processing information. After the flow, and after determining the required multiple data streams, the window calculation results corresponding to the data streams are reconstructed in a stream form based on the processing parameters in the data processing purpose, and finally the semantic structure recombination flow is obtained.
  • the semantic structure reorganization flow can be calculated, and when the calculation is performed, the formula based on the recombination flow calculation is needed.
  • the recombination flow calculation formula is obtained in advance in the pre-processing stage of multi-stream flow data processing, and can be directly called when performing the semantic structure reorganization flow calculation operation.
  • the technical solution of the embodiment firstly performs window processing on the data in the data stream after acquiring the data stream, and performs window calculation based on the window calculation formula; and then calculates a result base for the obtained window.
  • the semantic structure reorganization is performed in the data processing purpose to form the semantic structure reorganization flow; then the semantic structure reorganization flow is calculated based on the recombination flow calculation formula to obtain the data processing result.
  • the processing method effectively realizes multi-stream flow calculation involving massive entity data in real time, thereby not only solving the problem of long time in big data processing and high cost of abnormal processing, but also solving the problem that multiple data streams cannot be performed.
  • the problem of processing is processed, and the realization process of processing and calculating the massive data stream is further improved, thereby satisfying the user's request for real-time exploration of deep information or reason based on the data processing result.
  • FIG. 2 is a flowchart of a method for processing multi-streaming data according to a second embodiment of the present invention.
  • the embodiment of the present invention is optimized based on the foregoing embodiment 1.
  • the optimization further includes: determining the window calculation formula, the recombination flow calculation formula, and the semantic structure based on an instruction input by the user.
  • a method for processing multi-streaming data according to Embodiment 2 of the present invention specifically includes:
  • S210 Determine, according to an instruction input by a user, the window calculation formula, the recombination flow calculation formula, and a semantic structure.
  • an instruction input by a user or a technician may be acquired, and multi-streaming data processing is preprocessed based on the acquired instruction.
  • the instruction may be a code instruction
  • the preprocessing process may be summarized as: 1) extracting a multi-flow calculation formula and a semantic structure required for multi-stream streaming processing by a semantic analyzer based on the acquired code instruction; 2) Formula disassembly strategy for the multi-flow calculation formula The formula disassembly is performed to obtain the window calculation formula for single stream calculation and the recombination flow calculation formula for multi-flow calculation.
  • the multi-stream flow calculation formula is specifically applicable to the calculation of multi-stream flow data; the semantic structure may be used as processing information in the data processing purpose for obtaining a semantic structure reorganization flow.
  • the multi-flow calculation formula After determining the multi-flow calculation formula based on the above step 1), since the data information required for the multi-flow calculation formula is included in the plurality of data streams, the multi-flow calculation formula cannot be directly processed and calculated, and the multi-flow is required.
  • the calculation formula is disassembled, split into multiple calculation formulas, and then multiple calculation formulas are processed and calculated.
  • the above step 2) implements the disassembly of the multi-flow calculation formula, wherein the formula disassembly strategy can be specifically understood as a formula disassembly algorithm with a set priority calculation rule.
  • the priority calculation rule can be understood as: first consider the calculation of a single data stream, and then consider the calculation of multiple data streams. Based on the split strategy of the formula, the determined multi-flow calculation formula can be split into a window calculation formula required for single data stream calculation, and a recombination flow calculation formula required for calculation of multiple data streams.
  • the window calculation formula and the recombination flow calculation formula obtained above are all stored in the task configuration file for use in subsequent steps.
  • the method may further include: storing the semantic structure, the window calculation formula, and the reorganization flow calculation formula into a task configuration file. in.
  • the task configuration file may be specifically understood as a configuration file for storing relevant operation data information required in the multi-streaming data processing process.
  • Task scheduling can be performed based on the operational data information in the task configuration file throughout the multi-streaming data processing process.
  • S230 Perform semantic structural reorganization on the obtained window calculation result according to the set data processing purpose to obtain a semantic structure recombination data stream.
  • the data processing purpose may specifically include a semantic structure and a processing parameter, where the semantic structure is obtained based on the pre-processing of S210, and is pre-stored in the task configuration file, and may be from the task when the semantic structure is reorganized. Called directly in the configuration file; the processing parameters can be dynamically obtained from the external environment when the semantic structure is reorganized. Specifically, first, the data flow required for the semantic structure reorganization is determined based on the semantic structure, and after determining the required multiple data streams, the window calculation result corresponding to each data stream is processed in a stream form based on the data processing. The processing parameters in the purpose are structurally reorganized, and finally the semantic structure reorganization flow is obtained.
  • the technical solution of this embodiment further increases the pre-processing operation for multi-streaming data processing.
  • the calculation formula and semantic structure required for multi-stream streaming data processing can be determined before the processing of the multi-data stream, so as to realize the pre-processing of the multi-flow calculation formula, thereby ensuring multi-stream stream data processing. Work properly.
  • FIG. 3 is a flowchart of a method for processing multi-streaming data according to Embodiment 3 of the present invention.
  • the embodiment of the present invention is optimized based on the foregoing embodiment 1.
  • “the acquired Multi-streaming data is windowed” optimized to: divide the acquired multi-stream streaming data into respective data calculation windows based on the acquired data attributes of the multi-stream streaming data; calculate windows for the respective data A window mark is assigned, wherein the data calculation window assigned to the window mark is set to reject subsequent data inflows belonging to the data calculation window in the data stream.
  • the optimization further includes: collecting data of subsequent requests flowing into the respective data calculation windows as expiration data of the respective data calculation windows; performing window calculation on the expired data, and performing semantic structure reorganization on the window calculation result to obtain a semantic structure Recombining the data compensation stream; performing a recombination flow calculation on the semantic structure recombination data compensation stream to obtain a compensation calculation result; and correcting the data processing result according to the compensation calculation result.
  • a method for processing multi-streaming data according to Embodiment 3 of the present invention specifically includes:
  • the multi-streaming data can be understood as data of multiple data streams, and the acquired data stream is mainly an unordered data stream input in real time.
  • the data attribute may specifically refer to information capable of representing data characteristics in the data stream. Illustrative, such as the date, time, or ID number of the data.
  • the data attribute is generally used to divide data in the data stream, and the acquired data stream can be divided into a plurality of data sets having different data marks, thereby forming a plurality of data calculation windows including different data.
  • the respective data calculation windows may be given window markers based on the marking policy.
  • the marking strategy can be specifically understood as an ending condition of the current data calculation window calculation operation based on the dynamic setting of the inflow of the data stream.
  • the marking policy may be based on a time setting or a data flow value setting based on the set time.
  • the marking policy can be set to +10 minutes at the T time point, and when the time point of the T time point +10 minutes is reached, the T time period window can reject the receiving genus. The data in this window flows into this window.
  • the acquired data stream is an unordered data stream input in real time
  • dividing the data in the data stream into the corresponding data calculation window is also performed in real time, but any data calculation window is not unconditionally received for the inflow data.
  • the timing of reception is conditionally limited.
  • the conditional limitation is the marking policy provided by the embodiment, that is, after marking any data computing window based on the marking policy, the data computing window no longer receives subsequent data belonging to the data computing window in the data stream.
  • the operation of S320 can ensure the overall timeliness of multi-stream data processing. Because the real-time processing calculation of the data in the data calculation window is only a part of the operation of the entire multi-stream streaming data processing, the subsequent semantic structure reorganization operation needs to be based on the window calculation result before the execution can be started, if the data of the data stream is always divided into For the operation of the corresponding data calculation window, it is necessary to perform real-time window calculation on the data in the data calculation window, thereby delaying the entire data processing time.
  • the real-time window calculation is performed on the data in the data calculation window based on the window calculation formula, and the window calculation result is output to the cache in real time.
  • the window calculation formula is determined based on the formula split strategy proposed in the second embodiment, and is pre-stored in the task configuration file.
  • the window calculation result of real-time window calculation on the data in the data calculation window is stored in real time to the corresponding position of the cache, and the window calculation result can be used as the data flow of the semantic structure reorganization operation part, and the window is The result of the calculation is stored in the cache for quick recall during subsequent use.
  • step S340 and S350 are completely the same as steps S120 and S130 in the first embodiment, and will not be described in detail herein.
  • S360-S390 in the third embodiment of the present invention embodies another optimization of the embodiment of the present invention, that is, the correction operation of the obtained data processing result is added.
  • S360 Collect data of subsequent requests flowing into the respective data calculation windows as expired data of the respective data calculation windows.
  • the input data stream is an unordered data stream when the window calculation operation is performed on each data calculation window, data delay input occurs, thereby assigning a window mark to the respective data calculation windows.
  • data flows into the corresponding data calculation window. Since the data calculation window given to the window mark is set to reject the subsequent data inflow belonging to the data calculation window in the data stream, when the above situation occurs, the new inflow data cannot be continued in the original data calculation. Incremental processing is performed in the window. However, in order to ensure the correctness of the final data processing result, special processing of the newly flowing delayed data is required.
  • the delayed data requested to flow into the data calculation window after the window is marked is referred to as expired data, and the following steps S370 to S390 embodies the special processing performed on the expired data.
  • the window calculation is performed again on the expired data, and the window calculation result is output within the set time to ensure the subsequent steps are performed. It can be understood that after the window is marked, the data calculation window needs to be monitored in real time, and the expired data rejected by the data calculation window is collected and processed in real time to ensure real-time correction of the final data processing result.
  • the processing of the expired data may be equivalent to re-executing the three operational steps described in Embodiment 1 of the present invention. Therefore, after performing the window calculation operation data window calculation result on the expired data, the semantics of the window calculation result are still performed based on the set data processing purpose. Structural reorganization, recombination data compensation flow with synthetic semantic structure.
  • the recombination flow calculation can be further implemented, and the compensation calculation result is achieved.
  • the compensation calculation result calculated by the processing may be used to correct the result of the previously obtained data processing result.
  • the implementation of the correction may set a corresponding correction formula based on actual conditions.
  • the timing of correcting the data processing result based on the compensation calculation result can be understood as being performed after the data processing result is obtained.
  • the correction operation is started.
  • the data processing result of the data can be corrected in real time. . This not only ensures the accuracy of multi-stream data processing results, but also ensures the timeliness of multi-stream data processing.
  • the technical solution of the embodiment proposes an optimization scheme for correcting the data processing result, thereby realizing the correction of the data processing result while ensuring efficient processing and calculation of the massive multi-data stream, thereby ensuring the multi-stream flow type.
  • the efficiency and accuracy of data processing methods proposes an optimization scheme for correcting the data processing result, thereby realizing the correction of the data processing result while ensuring efficient processing and calculation of the massive multi-data stream, thereby ensuring the multi-stream flow type.
  • Embodiment 4 is a flowchart of a method for processing multi-streaming data according to Embodiment 4 of the present invention.
  • the embodiment of the present invention is optimized based on the foregoing Embodiment 1.
  • the optimization further includes: modifying the value of the first semantic collaboration metadata, wherein the value of the first semantic collaboration metadata is used to identify whether to trigger the semantic structure reorganization operation.
  • the semantic structure reorganization of the obtained window calculation result based on the set data processing purpose to obtain the semantic structure recombination data stream is optimized to: the value of the first semantic collaboration metadata
  • the window calculation result is obtained from the acquisition cache in a stream form; based on the set data processing purpose, the obtained window calculation result stream is semantically reorganized to obtain the recombined data. flow.
  • the optimizing includes: modifying a value of the second semantic collaboration metadata, where the value of the second semantic collaboration metadata is used to identify whether to trigger a recombination flow calculation operation. Further, when the value of the second semantic collaboration metadata meets the trigger condition of the recombination flow calculation operation, the recombination flow calculation is performed on the semantic reorganization data stream based on the recombination flow calculation formula, and the data processing result is obtained.
  • a method for processing multi-streaming data according to Embodiment 4 of the present invention specifically includes:
  • S420 Modify a value of the first semantic collaboration metadata, where the value of the first semantic collaboration metadata is used to identify whether to trigger a semantic structure reorganization operation.
  • the calculation result of the real-time processing calculation window is directly stored in the corresponding position of the data calculation window in the cache, and can be used as the subsequent semantics.
  • the trigger condition can be considered as the value of the first semantic collaboration metadata set after the data calculation window is marked.
  • the data calculation window when the data calculation window is marked, no data flows into the window, and the corresponding window calculation result value is no longer changed in real time, and the multi-stream stream data needs to be further processed, that is, the subsequent semantic structure is performed.
  • Reorganization prior to the semantic reorganization of multi-streaming data, specific triggering operations are required to initiate semantic reorganization of multi-streaming data. This embodiment may determine whether to trigger the initiation of semantic structural reorganization based on modifying the value of the first semantic collaboration metadata.
  • the value of the first semantic collaboration metadata may be specifically regarded as a flag bit that triggers the start of the next processing step.
  • the value of the first semantic collaboration metadata may be allowed to be 0 or 1. When the value is 0, the start of the next step may not be triggered. When the value is 1, the value may be considered as The start of the next step can be triggered.
  • the window calculation result is obtained in a stream form.
  • the obtaining the window calculation result in a stream form may be summarized as: determining at least two input streams required for performing multi-stream stream data processing, where the at least two input streams may be specifically regarded as passing through a data calculation window.
  • the calculated window calculation result stream of at least two data streams may be summarized as: determining at least two input streams required for performing multi-stream stream data processing, where the at least two input streams may be specifically regarded as passing through a data calculation window.
  • window calculation results of different data streams are respectively stored in specified locations of the cache, and different data streams may refer to data sets generated by different source assemblies at the same time or at different times.
  • a data set generated by a source assembly on a different date Illustratively, a data set generated yesterday from the same source assembly and a data set generated today may belong to different data streams.
  • the set data processing purpose can be specifically understood as the purpose that the currently processed multi-stream streaming data needs to achieve.
  • the semantic structural reorganization of the obtained window calculation result stream is specifically understood as: performing semantic structural reorganization on the window calculation result of the at least two data streams.
  • the semantic structure is reorganized based on the data processing purpose, thereby synthesizing at least two data streams into a recombined data stream based on the window calculation result,
  • the reorganized data stream can be used to calculate the final data results for multi-stream streaming data.
  • S450 Modify a value of the second semantic collaboration metadata, where the value of the second semantic collaboration metadata is used to identify whether to trigger a recombination flow calculation operation.
  • the second semantic collaboration metadata has the same function as the first semantic collaboration metadata, and can be understood as triggering the initiation of the next processing step.
  • the second semantic metadata is specifically used to trigger a recombination flow calculation operation. It should be noted that the operation of the window calculation, the operation of the semantic structure reorganization, and the reorganization flow calculation operation are related to each other and cooperate with each other, and the execution of the latter operation can be triggered only after the previous operation is completed.
  • the value of the second semantic collaboration metadata may also be regarded as a trigger flag, and the value may also be preferably 0 or 1, and when the value is 0, it indicates that the next operation is not triggered, and the value is 1 indicates that the next step can be triggered.
  • the recombination flow calculation operation on the semantic structure reorganization data stream is started.
  • the weight The group flow calculation operation is mainly based on a recombination flow calculation formula, and the recombination flow calculation formula is specifically obtained based on a formula disassembly strategy. After performing the recombination stream calculation, the data processing result expected by the multi-stream stream data processing can be obtained.
  • the technical solution in this embodiment embodies the implementation steps in the first embodiment. Based on the processing method provided in this embodiment, the real-time processing of massive data can be realized, and the processing and calculation of multiple data streams can be realized, and the processing is further improved.
  • the implementation process of processing and calculating massive data streams satisfies the user's desire to explore deep-level information or causes in real time based on data processing results.
  • the three large operation steps S410, S420-S440, and S450-S460 in this embodiment are mutually triggered to cooperate with each other.
  • the major operation steps may be between the respective operations.
  • the associated semantic collaboration metadata is triggered, and the semantic structures acquired in the preprocessing stage are used to cooperate with each other to jointly perform multi-stream data processing.
  • the data flow of the initial window computing operation is an unconditional running flow.
  • the semantic structure reorganization flow is based on the first The value of a semantic collaboration metadata is triggered and started, and then the reorganization flow calculation operation is triggered and started based on the value of the second semantic collaboration metadata, and finally the required data processing result is obtained.
  • FIG. 5 is a preferred embodiment of a multi-streaming data processing method according to Embodiment 5 of the present invention.
  • the application scenario of the fifth embodiment is a keyword consumption diagnosis function in a Baidu compass product, which is required by the application scenario.
  • the data is the click consumption data generated by the netizens based on the keywords displayed by Baidu promotion, and the data is called Fengchao click consumption data.
  • Fengchao click consumption data the information content of each data can be summarized as: keyword, click date, click time and consumption amount generated by click.
  • the method can be determined that the purpose of data processing is to obtain the sudden drop value of the accumulated consumption of all keywords in the hour of today's hour T, so that the user can query the account in real time based on the obtained data processing result.
  • the initial input stream is a Fengchao click consumption data stream
  • the data stream is characterized by: 1) large amount of data and 2) strong real-time performance.
  • the data stream is characterized by: 1) large amount of data and 2) strong real-time performance.
  • the preferred embodiment provided by the present embodiment also embodies the feature of multi-stream processing, that is, for the ring ratio calculation, data information based on today and yesterday is required, and thus conforms to the scope of multi-stream processing.
  • a preferred method for processing multi-streaming data according to Embodiment 5 of the present invention specifically includes:
  • the disassembly process of the multi-flow calculation formula can be expressed as:
  • the present embodiment since the present embodiment performs multi-stream streaming data processing, it is required to determine a multi-flow calculation formula for multi-stream streaming data processing based on the acquired input instruction before processing the multi-stream streaming data.
  • the semantic structure in this embodiment, it can be determined that the semantic structure is the today's hourly T cumulative consumption value of the calculated keyword and the yesterday's hourly T cumulative consumption value. Then, based on the input command, the multi-flow calculation formula can be determined as follows: calculating the integral point T cumulative consumption value of the keyword, and then calculating the current day of the keyword. The difference between the cumulative consumption value of point T and the cumulative consumption value of yesterday's hourly T.
  • the above calculation formula needs to be calculated step by step, so the above multi-flow calculation formula is split into a window calculation formula for single data flow calculation and used for multi-data based on the formula split strategy.
  • the formula for calculating the split of the above formula is: the cumulative sum of the consumption values of the keywords flowing in the time period of the whole point T; the formula for calculating the split flow is: calculating the keyword hour today T The cumulative consumption value of the time period and the cumulative consumption value of the keyword T time period of the hour yesterday. It should be noted that the calculation of the keyword consumption value mentioned in this embodiment specifically refers to the calculation of the same keyword consumption value.
  • the semantic structure determined based on the input instruction and the calculation formula separated based on the formula splitting strategy are correspondingly stored in the task configuration file, and when the data processing is performed based on the processing method, the operation can be quickly called when the corresponding operation step is performed.
  • the calculation formula required for the step is correspondingly stored in the task configuration file, and when the data processing is performed based on the processing method, the operation can be quickly called when the corresponding operation step is performed.
  • the calculation formula required for the step is performed.
  • the phoenix nest click consumption data stream may be acquired as an input stream, and after the phoenix nest click consumption data stream is acquired, the hour is obtained.
  • the time attribute of the level performs data calculation window division on the data of the data stream.
  • the hourly time attribute may specifically refer to a time period divided into 24 time segments, each time period is an hour-level window, and then determining which hour-level window to divide the data based on the click time of the phoenix nest click data. For example, if the currently acquired keyword click time is 0:24, the keyword and its data information can be divided into 0 hour window.
  • S530 Perform real-time keyword total point consumption value calculation on the data in the hourly window, and output the calculation result to the cache.
  • the data in the data calculation window needs to be calculated in real time based on the window calculation formula, and the window calculation result is output to the cache in real time.
  • the cumulative calculation of the consumption value is performed on the keywords classified into this window.
  • a click is generated at 0:1, 0:4, and 0:8.
  • the consumption amount is 1 yuan, 1 yuan, and 1 yuan respectively, and as of 0:9, the cumulative value of the consumption value of the keyword A is 3 yuan, and the calculation result stored in the cache is 3 yuan.
  • the calculation process of other hourly windows is the same as described above and will not be described in detail here.
  • S540 Determine whether to mark the hourly window, if yes, execute S550 and S551 respectively; if not, return to execute S530.
  • a freeze timing of freezing the hourly window may be set (the freezing may be understood as rejecting an inflow request of subsequent data belonging to the window in the data stream), and the The freeze timing setting can be based on the window mark.
  • the time when the input data stream is divided into the hour-level window can be roughly determined to be between 0 and 1 point, after 1 point Most of the data streams obtained will be divided into 1 hour-level windows based on time attributes.
  • the 0 o'clock level window can be window marked by setting a marking policy to identify that the hourly level window can be frozen.
  • the setting of the marking policy is generally based on the application scenario setting in which the processing method of the embodiment is located.
  • the marking policy can be set to timing.
  • the tag for example, sets the tagging policy to: window mark the 0 o'clock level window at 1 o'clock + 5 minutes, i.e., 1 o'clock, thereby freezing the hourly window.
  • the window marking it is determined whether the hour-level window is marked based on the marking policy. If the window marking is not performed, it is required to return to S530 to continue the real-time incremental calculation. If the marking is performed in accordance with the marking policy, the window marking may be performed separately. Two branching steps, which can be summarized as a compensation processing calculation of the expired data based on S551, and a multi-stream processing calculation of the real-time data based on S550 to S580.
  • the data that is requested to flow into the hourly window after the window is marked is determined as expired data, and the real-time fractal compensation calculation is performed on the expired data, and the compensation calculation result is obtained, and then S590 is executed.
  • the process of real-time fractal compensation calculation for expired data can be summarized as: obtaining expired data, performing real-time incremental calculation on the expired data based on the window calculation formula, and obtaining a window calculation result of real-time incremental calculation;
  • the today's keyword integral point consumption value of the expired data stream and the yesterday's keyword integral point consumption value corresponding to the above-mentioned expired data stream are subjected to semantic structure reorganization based on data processing purpose, and the semantic structure reorganization compensation flow is synthesized; finally, based on the recombination calculation formula
  • the semantic structure recombination compensation stream performs a recombination flow calculation, and finally obtains a compensation calculation result.
  • the acquisition process of the compensation calculation result is performed in real time, and the above operation is performed as long as the expired data appears.
  • S550 Obtain a data stream of today's keyword hourly consumption cumulative value in the cache and yesterday's keyword hourly consumption cumulative value data stream.
  • S560 based on the purpose of data processing, reconstructs the semantics of the two data streams obtained, and synthesizes the semantic structure recombination stream.
  • S550 to S570 are another branch step that needs to be executed after the hour-level window is marked.
  • the value of the first semantic collaboration metadata associated with the semantic structural reorganization needs to be modified to determine whether to trigger the semantic based on the value of the first semantic collaboration metadata.
  • the operation of the structural reorganization, and after triggering the execution of the semantic structure reorganization, the data stream of the today's keyword hourly consumption cumulative value stored in the cache and the data stream of the yesterday's keyword whole point consumption cumulative value are obtained, based on the data in the application scenario.
  • the two data streams obtained are reconstructed by semantic structure to obtain a synthesized semantic structure reorganization flow; then, after synthesizing the semantic structure reorganization flow, the second semantic cooperation element associated with the reorganization flow calculation operation needs to be modified.
  • the value of the data is determined based on the value of the second semantic collaboration metadata to determine whether to perform the operation of performing the recombination flow calculation, and after the triggering of the execution of the reorganization flow calculation, the recombination flow calculation is performed on the synthesized semantic reorganization flow.
  • the recombination flow calculation is to calculate the difference between the cumulative value of today's keyword hourly consumption and the cumulative value of yesterday's keyword hourly consumption, and the calculated difference can be regarded as multi-stream data processing in the application scenario. Data processing results.
  • first semantic collaboration metadata and the second semantic collaboration metadata proposed in this embodiment are not specifically embodied in the flowchart of the fifth embodiment, but are based on semantic collaboration metadata and corresponding operation steps.
  • the exchange of data and information between the two is reflected, but the role played and the meaning to be expressed are not changed by the change of the textual expression.
  • the data initial processing result can be obtained.
  • the data processing result can be regarded as the processing result of the multi-stream streaming data in the application scenario.
  • the data processing result correction needs to be performed based on the compensation calculation result obtained by the real-time calculation, and Finally, the corrected processing result is output.
  • the data processing knot For example, the triggering process may be performed based on the data information exchange between the semantic collaboration metadata and the reorganization computing operation in FIG. 5a.
  • the final output result is the cumulative consumption value of each keyword at the whole point T, and the cumulative consumption of each keyword at the whole point T.
  • the value of the ring is the sudden drop value.
  • FIG. 5b is an application effect diagram of a multi-stream data processing method in keyword consumption diagnosis according to Embodiment 5 of the present invention.
  • the display result is a curve distribution of the accumulated consumption value of the data of any specified keyword in any user ID in the time period of the whole point T, and the accumulated consumption value of the keyword in the time period of the whole point T.
  • the curve distribution of the ring drop values where T ⁇ [0 points, 23 points].
  • the information content shown in FIG. 5b can only be regarded as an intercepted segment of the data processing result of a certain keyword at a certain time point. In actual application, if the keyword has a compensation calculation result, then At different times, the data processing results obtained by the keyword will change based on the correction of the compensation calculation result.
  • the technical solution of the present embodiment provides a preferred embodiment of a multi-streaming data processing method. Based on the preferred embodiment, the operation steps of the multi-streaming data processing method are more clear. Based on the processing method, Efficiently realize real-time processing and calculation of massive data, save data collection time, speed up data processing progress, and further ensure the accuracy of multi-flow calculation results by setting compensation processing calculation.
  • FIG. 6 is a structural block diagram of a multi-streaming data processing system according to Embodiment 6 of the present invention.
  • the processing system of this embodiment may be implemented by software and/or hardware, and may be applicable to massive multi-streaming data.
  • the processing system includes a windowing processing module 60, a window computing module 61, a data stream recombining module 62, and a recombination stream computing module 63.
  • the windowing processing module 60 is configured to perform windowing processing on the acquired multi-stream streaming data.
  • the window calculation module 61 is configured to perform window calculation on the windowed multi-stream stream data based on the window calculation formula.
  • the data stream recombination module 62 is configured to perform semantic structure recombination on the obtained window calculation result based on the set data processing purpose to obtain a semantic structure recombination data stream.
  • the recombination flow calculation module 63 is configured to perform a recombination flow calculation on the semantic structure recombination data stream based on the recombination flow calculation formula to obtain a data processing result.
  • the processing system first performs windowing processing on the acquired multi-stream streaming data through the windowing processing module 60, and the window computing module 61 performs window calculation based on the window calculation formula; and then, through the data stream recombining module. 62, based on the set data processing purpose, perform semantic structural reorganization on the obtained window calculation result to obtain a semantic structure recombination data stream; finally, the reorganization flow calculation module 63 reorganizes the semantic structure based on the recombination flow calculation formula The stream performs a recombination stream calculation to obtain a data processing result.
  • the technical solution of the embodiment effectively implements multi-stream flow calculation involving massive entities in real-time processing, thereby not only solving the problem of long time consumption in big data processing and high cost of abnormal processing, but also solving the problem that multi-data stream cannot be performed.
  • the problem of processing is processed, and the realization process of processing and calculating the massive data stream is further improved, thereby satisfying the user's request for real-time exploration of deep information or reason based on the data processing result.
  • processing system may further include:
  • a first pre-processing module configured to determine the window calculation publicly based on an instruction input by a user before windowing the acquired multi-stream streaming data and performing window calculation based on the window calculation formula Equation, the recombination flow calculation formula and semantic structure.
  • processing system may further include:
  • a second pre-processing module configured to store the semantic structure, the window calculation formula, and the reorganization flow calculation formula into a task configuration file.
  • the windowing processing module 60 may be specifically configured to:
  • processing system may further include:
  • a first metadata value modification module configured to modify a value of the first semantic collaboration metadata after the window calculation is performed based on the window calculation formula, where the value of the first semantic collaboration metadata is used to identify whether to trigger semantic structural reorganization operating.
  • the data stream recombining module 62 may be specifically configured to:
  • processing system may further include:
  • a second metadata value modification module configured to modify a value of the second semantic collaboration metadata, where the value of the second semantic collaboration metadata is used to identify whether to trigger a recombination flow calculation operation.
  • the reassembly flow calculation module 63 is specifically configured to:
  • the recombination flow calculation formula is used to perform a recombination flow calculation on the semantic structure recombination data stream to obtain a data processing result.
  • the processing system may further include an expiration data compensation processing module, configured to: after the window identifiers are given to the respective data calculation windows, collect data of subsequent requests flowing into the respective data calculation windows, as the respective data Calculating the outdated data of the window; performing window calculation on the expired data, and performing semantic structure reorganization on the window calculation result to obtain a semantic structure recombination data compensation stream; performing recombination flow calculation on the semantic structure recombination data compensation stream to obtain a compensation calculation result And correcting the data processing result according to the compensation calculation result.
  • an expiration data compensation processing module configured to: after the window identifiers are given to the respective data calculation windows, collect data of subsequent requests flowing into the respective data calculation windows, as the respective data Calculating the outdated data of the window; performing window calculation on the expired data, and performing semantic structure reorganization on the window calculation result to obtain a semantic structure recombination data compensation stream; performing recombination flow calculation on the semantic structure recombination data compensation stream to obtain a compensation calculation result And
  • FIG. 6b is a structural diagram of a multi-streaming data processing system according to Embodiment 7 of the present invention, and a framework diagram shown in FIG. 6b, which illustrates a workflow of a multi-stream streaming data processing system according to an embodiment of the present invention.
  • the framework diagram may be specifically divided into four interrelated semantic streams, namely: a window statistical analysis stream 64, a semantic structure reorganization stream 65, a semantic structure calculation stream 66, and a fractal compensation stream 67.
  • the window statistical analysis stream 64 corresponds to the window processing module 60 and the window computing module 61 of the multi-stream stream data processing system, and the remaining three semantic streams respectively correspond to the data stream recombining module of the multi-stream stream data processing system. 62.
  • each semantic stream independently completes the functional steps of the respective corresponding modules on the basis of mutual cooperation, and finally realizes processing of the multi-stream streaming data.
  • Embodiment 7 of the present invention provides an apparatus, including a multi-streaming data processing system provided by any embodiment of the present invention, and the device can serve as a server.
  • an embodiment of the present invention provides a device, where the device includes: a processor 71, a memory 72, an input device 73, and an output device 74.
  • the number of processors 71 in the device may be one or more.
  • the processor 71, the memory 72, the input device 73, and the output device 74 in the device may be connected by a bus or other means, and the bus connection is taken as an example in FIG.
  • the memory 72 is used as a computer readable storage medium, and can be used to store a software program, a computer executable program, and a module, such as a program instruction/module corresponding to the processing method of the multi-streaming data in the embodiment of the present invention (for example, the drawing).
  • the windowing processing module 60, the window calculation module 61, the data stream recombining module 62, and the recombination stream computing module 63 in the processing system of the multi-streaming data shown in Fig. 6a.
  • the processor 71 executes various functional applications and data processing of the device by executing software programs, instructions, and modules stored in the memory 72, that is, the processing method of the multi-stream streaming data in the foregoing method embodiments.
  • the memory 72 may include a storage program area and a storage data area, wherein the storage program area may store an operating system, an application required for at least one function; the storage data area may store data created according to usage of the device, and the like.
  • memory 72 can include high speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other non-volatile solid state storage device.
  • memory 72 may further include memory remotely located relative to processor 71, which may be connected to the device over a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
  • Input device 73 can be used to receive input numeric or character information and to generate key signal inputs related to user settings and function control of the device.
  • Output device 74 can include an output port or the like.
  • the above apparatus may include one or more programs, the one or more programs being stored in the memory 72, and when executed by the one or more processors 71, the program performs the following operations:
  • Semantic structure reorganization of the obtained window calculation result based on the set data processing purpose To obtain a semantic structure reorganization data stream;
  • the semantic flow recombination data stream is subjected to a recombination flow calculation to obtain a data processing result.
  • the embodiment of the present invention further provides a storage medium including computer executable instructions for executing a multi-stream streaming data processing method when executed by a computer processor, the method comprising:
  • the semantic calculation of the obtained window calculation result is performed to obtain a semantic structure recombination data stream
  • the semantic flow recombination data stream is subjected to a recombination flow calculation to obtain a data processing result.
  • the present invention can be implemented by software and necessary general hardware, and can also be implemented by hardware, but in many cases, the former is a better implementation. .
  • the technical solution of the present invention which is essential or contributes to the prior art, may be embodied in the form of a software product, which may be stored in a computer readable storage medium, such as a floppy disk of a computer. , Read-Only Memory (ROM), Random Access Memory (RAM), Flash (FLASH), hard disk or optical disk, etc., including a number of instructions to make a computer device (can be a personal computer)
  • the server, or network device, etc. performs the methods described in various embodiments of the present invention.
  • each of the included The units and modules are only divided according to the functional logic, but are not limited to the above-mentioned divisions, as long as the corresponding functions can be implemented; in addition, the specific names of the functional units are only for the purpose of facilitating mutual differentiation, and are not intended to limit the present invention. The scope of protection.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Databases & Information Systems (AREA)
  • Software Systems (AREA)
  • Data Mining & Analysis (AREA)
  • Information Transfer Between Computers (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

一种多流流式数据的处理方法、系统、存储介质及设备。该处理方法包括:对所获取的多流流式数据进行窗口化处理并基于窗口计算公式进行窗口计算(S110);基于设定的数据处理目的,对所得到的窗口计算结果进行语义结构重组,以得到语义结构重组数据流(S120);基于重组流计算公式,对所述语义结构重组数据流进行重组流计算,得到数据处理结果(S130)。利用该处理方法,不仅解决了大数据处理中耗时长,异常处理成本高的问题,还解决了不能对多数据流进行处理计算的问题,进而更加完善了对海量数据流进行处理计算的实现过程,从而满足了用户基于数据处理结果实时探索深层次信息或原因的诉求。

Description

一种多流流式数据的处理方法、系统、存储介质及设备
本专利申请要求于2016年4月25日提交的、申请号为201610262701.1、申请人为百度在线网络技术(北京)有限公司、发明名称为“一种多流流式数据的处理方法和系统”的中国专利申请的优先权,该申请的全文以引用的方式并入本申请中。
技术领域
本发明实施例涉及大数据处理技术,尤其涉及一种多流流式数据的处理方法、系统、存储介质及设备。
背景技术
随着时代的进步和经济的发展,人们日常生活中对信息的需求量越来越大,尤其是随着互联网的日益普及,每天都有海量的信息在互联网上发布和传播,对于数据计算和分析的技术人员来说,传统的数据计算分析系统已不能承受对海量数据的计算分析,由此出现了对大规模数据的处理系统。
目前,常见的大数据处理技术有两种:一种是将已获取的整体数据集作为输入,然后通过批量分析计算模型进行处理计算,最终输出所需的结果集;另一种是以数据流的形式输入,然后对流形式的数据进行实时的分析计算,并实时获取计算结果。上述两种数据处理技术,尽管可以处理海量数据,但存在一定的不足:针对第一种数据处理技术,存在耗时较长,异常处理成本高等问题,且也不能满足实时性分析计算大数据的要求。针对后一种数据处理方法,尽管能够满足实时性分析,但仅擅长处理单条数据流的计算处理(如对数据流的统计求和计算),却不能很好地支持多条数据流的分析计算(如,多条数据流之 间进行计算,或单条数据流进行环比计算等)。
综上所述,现有技术中实时处理多条海量数据流的问题依旧不能解决,无法满足实际需求。
发明内容
本发明实施例提供了一种多流流式数据的处理方法、系统、存储介质及设备,以实现对海量的多流流式数据的实时处理。
本发明实施例采用以下技术方案:
第一方面,本发明实施例提供了一种多流流式数据的处理方法,该方法包括:
对所获取的多流流式数据进行窗口化处理并基于窗口计算公式进行窗口计算;
基于设定的数据处理目的,对所得到的窗口计算结果进行语义结构重组,以得到语义结构重组数据流;
基于重组流计算公式,对所述语义结构重组数据流进行重组流计算,得到数据处理结果。
第二方面,本发明实施例提供了一种多流流式数据的处理系统,该系统包括:
窗口化处理模块,用于对所获取的多流流式数据进行窗口化处理;
窗口计算模块,用于基于窗口计算公式对经过窗口化处理后的多流流式数据进行窗口计算;
数据流重组模块,用于基于设定的数据处理目的,对所得到的窗口计算结果进行语义结构重组,以得到语义结构重组数据流;
重组流计算模块,用于基于重组流计算公式,对所述语义结构重组数据流 进行重组流计算,得到数据处理结果。
第三方面,本发明实施例还提供了一种包含计算机可执行指令的存储介质,所述计算机可执行指令在由计算机处理器执行时用于执行一种多流流式数据的处理方法,该方法包括:
对所获取的多流流式数据进行窗口化处理并基于窗口计算公式进行窗口计算;
基于设定的数据处理目的,对所得到的窗口计算结果进行语义结构重组,以得到语义结构重组数据流;
基于重组流计算公式,对所述语义结构重组数据流进行重组流计算,得到数据处理结果。
第四方面,本发明实施例又提供了一种设备,该设备包括:
一个或者多个处理器;
存储器;
一个或者多个程序,所述一个或者多个程序存储在所述存储器中,当被所述一个或者多个处理器执行时,进行如下操作:
对所获取的多流流式数据进行窗口化处理并基于窗口计算公式进行窗口计算;
基于设定的数据处理目的,对所得到的窗口计算结果进行语义结构重组,以得到语义结构重组数据流;
基于重组流计算公式,对所述语义结构重组数据流进行重组流计算,得到数据处理结果。
本发明提供了一种多流流式数据的处理方法、系统、存储介质及设备,该处理方法,首先在获取到数据流后,对数据流中的数据进行窗口化处理,并基 于窗口计算公式进行窗口计算;然后对所得的窗口计算结果基于数据处理目的进行语义结构重组形成语义结构重组流;之后对语义结构重组流基于重组流计算公式进行计算得到数据处理结果。利用该处理方法,有效的实现了实时处理涉及海量实体数据的多流流式计算,由此不仅解决大数据处理中耗时长,异常处理成本高的问题,还解决了不能对多数据流进行处理计算的问题,进而更加完善了对海量数据流进行处理计算的实现过程,从而满足了用户基于数据处理结果实时探索深层次信息或原因的诉求。
附图说明
图1为本发明实施例一提供的一种多流流式数据的处理方法的流程图;
图2为本发明实施例二提供的一种多流流式数据的处理方法的流程图;
图3为本发明实施例三提供的一种多流流式数据的处理方法的流程图;
图4为本发明实施例四提供的一种多流流式数据的处理方法的流程图;
图5a为本发明实施例五提供的一种多流流式数据的处理方法的优选实施例;
图5b为本发明实施例五提供的基于多流流式数据处理方法在关键词消费诊断中体现出的应用效果图;
图6a为本发明实施例六提供的一种多流流式数据的处理系统的结构框图;
图6b为本发明实施例六提供的一种多流流式数据处理的构架图;
图7为本发明实施例七提供的一种设备的结构示意图。
具体实施方式
下面结合附图并通过具体实施方式来进一步说明本发明的技术方案。可以理解的是,此处所描述的具体实施例仅仅用于解释本发明,而非对本发明的限定。
另外还需要说明的是,为了便于描述,附图中仅示出了与本发明相关的部 分而非全部内容。在更加详细地讨论示例性实施例之前应当提到的是,一些示例性实施例被描述成作为流程图描绘的处理或方法。虽然流程图将各项操作(或步骤)描述成顺序的处理,但是其中的许多操作可以被并行地、并发地或者同时实施。此外,各项操作的顺序可以被重新安排。当其操作完成时所述处理可以被终止,但是还可以具有未包括在附图中的附加步骤。所述处理可以对应于方法、函数、规程、子例程、子程序等等。
实施例一
图1为本发明实施例一提供的一种多流流式数据的处理方法的流程图,该方法可适用于对海量的多流流式数据进行数据处理的情况,可以由多流流式数据的处理系统执行,其中该处理系统可由软件和/或硬件实现,并一般可集成在流式数据分析的框架平台上,所述框架平台可以是一种服务器。
如图1所示,本发明实施例一提供的一种多流流式数据的处理方法,具体包括:
S110、对所获取的多流流式数据进行窗口化处理并基于窗口计算公式进行窗口计算。
在本实施例中,所述多流流式数据具体可理解为来自不同数据流的、作用于同一个数据处理计算的流式数据。在本实施例中,在对多流流式数据进行处理时,需要有数据流的输入,一般地,不同数据流的输入过程相互独立。此外,本实施例中的数据流一般可看作无序数据流,即所述数据流的输入时序与其业务时序存在不同的情况。
所述窗口化处理的目的为将获取的数据流中的数据划分到数据计算窗口中。具体的,所述数据计算窗口可指具有相同数据标记(数据标记主要由数据的一个或多个属性值决定)的数据集合。对于流式数据而言,基于数据自身的属性 值就可以确定自身的数据标记,基于确定的数据标记,就可以确定数据自身所属的数据计算窗口,在确定出数据所属的数据计算窗口后,就可以将数据流中数据划分到相应的数据计算窗口。
在本实施例中,将数据流中的数据划分到相应的数据计算窗口后,可以基于窗口计算公式对所述数据计算窗口中的数据实时进行窗口计算,并输出窗口计算结果。其中,所述窗口计算公式具体可理解为在数据计算窗口中对所划分的数据进行处理计算时所依据的公式。一般的,常见的窗口计算公式有:对数据的指定属性进行累计求和、对数据的指定属性进行最大值计算、对数据的指定属性进行最小值计算、又或者对数据的指定属性进行的均指计算等。在本实施例中,所输出的窗口计算结果可以存储在高速缓存中,以便于在后续计算中使用。
在本实施例中,对于多流流式数据的数据处理,其过程可看作先对多个数据流分别进行单独处理计算,以便于后续将单独处理计算的结果进行重组,最终基于设定的多流计算公式实现多流流式数据的数据处理。
以将数据的时间属性值作为数据标记为例,对所获取的一条数据流进行单流处理操作的过程为:所获取的数据流基于时间属性值划分到不同的数据计算窗口(此时该数据计算窗口可简称为时间窗口);且可基于时间点进行时间窗口划分,实现数据的窗口化(常见的时间点划分可以是0点~1点、1点至2点,...,23点~0点的划分);在任一个时间窗口(如1点~2点)中,又可基于设定的窗口计算公式(如累计求和)实时计算数据的某个属性(如消费属性)在该时间窗口内的对应值(如消费累计值)。
S120、基于设定的数据处理目的,对所得到的窗口计算结果进行语义结构重组,以得到语义结构重组数据流。
在本实施例中,S110仅完成了多流流式数据处理中对各数据流进行单独处理的计算,为了实现多数据流的相互计算,需要将多数据流的窗口计算结果以流形式基于特定的条件进行汇合重组,以得到一条基于特定语义结构的数据流。
在本实施例中,所述数据处理目的具体可理解为多流流式数据处理的语义结构重组阶段所需要达到的目的。该数据处理目的一般包括语义结构重组操作时所需的处理参数以及基于用户输入指令获得的处理信息,其中,所述处理参数具体可理解为合成语义结构重组流时所需的从实际应用场景的外部环境中动态获取的相关上下文;所述处理信息可以在多流流式数据处理的预处理阶段确定,所述预处理阶段具体可理解为对实际应用场景中用户输入的指令进行分析拆解的过程。
在实际的多流流式数据处理过程中,获取到高速缓存中存放的窗口计算结果后,首先确定数据处理目的中的处理信息,然后基于所述处理信息来确定进行语义结构重组所需要的数据流,并在确定出所需要的多条数据流之后,将上述各数据流对应的窗口计算结果以流形式基于所述数据处理目的中的处理参数进行结构重组,最终得到语义结构重组流。
S130、基于重组流计算公式,对所述语义结构重组数据流进行重组流计算,得到数据处理结果。
在本实施例中,基于S120得到语义结构重组流之后,可以对语义结构重组流进行计算,在进行计算时,需要基于重组流计算公式。一般地所述重组流计算公式在进行多流流式数据处理的预处理阶段提前获得,只需在进行语义结构重组流计算操作时直接调用即可。
本实施例的技术方案,首先在获取到数据流后,对数据流中的数据进行窗口化处理,并基于窗口计算公式进行窗口计算;然后对所得的窗口计算结果基 于数据处理目的进行语义结构重组形成语义结构重组流;之后对语义结构重组流基于重组流计算公式进行计算得到数据处理结果。利用该处理方法,有效的实现了实时处理涉及海量实体数据的多流流式计算,由此不仅解决大数据处理中耗时长,异常处理成本高的问题,还解决了不能对多条数据流进行处理计算的问题,进而更加完善了对海量数据流进行处理计算的实现过程,从而满足了用户基于数据处理结果实时探索深层次信息或原因的诉求。
实施例二
图2为本发明实施例二提供的一种多流流式数据的处理方法的流程图,本发明实施例以上述实施例一为基础进行优化,在本实施例中,在“对所获取的多流流式数据进行窗口化处理并基于窗口计算公式进行窗口计算”之前,还优化包括了:基于用户输入的指令,确定所述窗口计算公式、所述重组流计算公式和语义结构。
如图2所示,本发明实施例二提供的一种多流流式数据的处理方法,具体包括:
S210、基于用户输入的指令,确定所述窗口计算公式、所述重组流计算公式和语义结构。
一般的,在进行流式数据处理(无论是单流流式数据处理还是多流流式数据处理)之前,首先需要对上述流式数据处理进行预处理,以获得流式数据处理所需的计算公式和语义结构。
在本实施例中,可获取用户或技术人员输入的指令,并基于所获取的指令对多流流式数据处理进行预处理。具体的,所述指令可以是代码指令,该预处理过程可以概括为:1)基于获取的代码指令通过语义分析器提取多流流式处理所需的多流计算公式以及语义结构;2)基于公式拆解策略对所述多流计算公式 进行公式拆解,得到用于单流计算的所述窗口计算公式和用于多流计算的所述重组流计算公式。
在本实施例中,所述多流流式计算公式具体可用于多流流式数据的计算;所述语义结构可以作为所述数据处理目的中的处理信息,以用于得到语义结构重组流。
在基于上述步骤1)确定出多流计算公式后,由于多流计算公式所需的数据信息包含在多个数据流中,不能对所述多流计算公式直接处理计算,需要对所述多流计算公式进行拆解,将其拆分成多个计算公式,然后分别对多个计算公式进行处理计算。
上述步骤2)实现了对多流计算公式的拆解,其中,所述公式拆解策略具体可理解为具有设定优先计算规则的公式拆解算法。其优先计算规则可理解为:先考虑单条数据流的计算,再考虑多条数据流的计算。基于所述公式拆分策略,可将确定的多流计算公式拆分成单条数据流计算所需的窗口计算公式,以及多条数据流计算所需的重组流计算公式。上述所得到的窗口计算公式和重组流计算公式均存放至任务配置文件中,以便于后续步骤中使用。
进一步的,在确定所述窗口计算公式、所述重组流计算公式和语义结构之后,还可以优选包括:将所述语义结构、所述窗口计算公式和所述重组流计算公式存放至任务配置文件中。
在本实施例中,所述任务配置文件具体可理解为存储多流流式数据处理过程中所需要的相关操作数据信息的配置文件。在整个多流流式数据处理过程中可基于任务配置文件中的操作数据信息进行任务调度。
S220、对所获取的多流流式数据进行窗口化处理并基于窗口计算公式进行窗口计算。
S230、基于设定的数据处理目的,对所得到的窗口计算结果进行语义结构重组,以得到语义结构重组数据流。
在本实施例中,所述数据处理目的具体可以包括语义结构以及处理参数,其中,所述语义结构基于S210的预处理获得,预先存储在任务配置文件中,在进行语义结构重组时可以从任务配置文件中直接调用;所述处理参数可以在在进行语义结构重组时从外部环境中动态获取。具体的,首先基于所述语义结构确定进行语义结构重组所需要的数据流,并在确定所需要的多条数据流之后,将上述各数据流对应的窗口计算结果以流形式基于所述数据处理目的中的处理参数进行结构重组,最终得到语义结构重组流。
S240、基于重组流计算公式,对所述语义结构重组数据流进行重组流计算,得到数据处理结果。
本实施例的技术方案,进一步增加了对多流流式数据处理的预处理操作。利用该处理方法,可以在对多数据流进行处理之前,确定多流流式数据处理所需的计算公式和语义结构,以实现多流计算公式的预处理,进而保证多流流式数据处理的正常进行。
实施例三
图3为本发明实施例三提供的一种多流流式数据的处理方法的流程图,本发明实施例以上述实施例一为基础进行优化,在本实施例中,将“对所获取的多流流式数据进行窗口化处理”优化为:基于所获取的多流流式数据的数据属性,将所获取的多流流式数据划分到各个数据计算窗口中;为所述各个数据计算窗口赋予窗口标记,其中被赋予窗口标记的数据计算窗口被设定为拒绝所述数据流中属于所述数据计算窗口的后续数据流入。
在上述优化的基础上,在“为所述各个数据计算窗口赋予窗口标记”之后, 还优化包括了:收集后续请求流入所述各个数据计算窗口的数据,作为所述各个数据计算窗口的过期数据;对所述过期数据进行窗口计算,并对窗口计算结果进行语义结构重组得到语义结构重组数据补偿流;对所述语义结构重组数据补偿流进行重组流计算,得到补偿计算结果;根据所述补偿计算结果对所述数据处理结果进行修正。
如图3所示,本发明实施例三提供的一种多流流式数据的处理方法,具体包括:
S310、基于所获取的多流流式数据的数据属性,将所获取的多流流式数据划分到各个数据计算窗口中。
在本实施例中,可以将所述多流流式数据理解为多条数据流的数据,且所获取的数据流主要为实时输入的无序数据流。所述数据属性具体可指能够表示数据流中的数据特性的信息。示例性的,如数据所具有的日期、时间或者ID号等。
所述数据属性一般用来划分数据流中的数据,可将所获取的数据流划分成多个具有不同数据标记的数据集,由此就可以形成多个包含不同数据的数据计算窗口。
S320、为所述各个数据计算窗口赋予窗口标记,其中被赋予窗口标记的数据计算窗口被设定为拒绝所述数据流中属于所述数据计算窗口的后续数据流入。
在本实施例中,可以基于标记策略对所述各个数据计算窗口赋予窗口标记。所述标记策略具体可理解为基于数据流的流入动态设定的当前数据计算窗口计算操作的结束条件。一般地,所述标记策略可基于时间设置,也可基于设定时间内的数据流流量值设置。示例性的,所述标记策略可设置为在T时间点+10分钟,当到达T时间点+10分钟这个时间点后,T时间段窗口就可拒绝接收属 于该窗口的数据流入到该窗口中。
具体的,由于所获取的数据流为实时输入的无序数据流,所以将数据流中数据划分到相应数据计算窗口也是实时进行的,但是任一数据计算窗口对所流入的数据并不是无条件接收的,其接收时机是有条件限定的。该条件限定就是本实施例所提供的标记策略,即在基于标记策略对任一数据计算窗口进行标记后,该数据计算窗口不再接收数据流中属于该数据计算窗口的后续数据。
在本实施例中,进行S320的操作可以保证多流流式数据处理的整体时效性。因为对数据计算窗口中数据的实时处理计算仅是整个多流流式数据处理的一部分操作,后续的语义结构重组操作需要基于窗口计算结果才可开始执行,若一直执行将数据流的数据划分到相应的数据计算窗口的操作,则就需要一直对数据计算窗口中的数据进行实时窗口计算,进而延误整个数据处理的时间。
S330、基于窗口计算公式进行窗口计算。
在本实施例中,可以基于所述窗口计算公式,对所述数据计算窗口中的数据进行实时窗口计算,并实时输出窗口计算结果到高速缓存中。
其中,所述窗口计算公式基于上述实施例二中提出的公式拆分策略确定,并预先存储在任务配置文件中。此外,对所述数据计算窗口中数据进行实时窗口计算的窗口计算结果是实时存放到高速缓存的相应位置上的,所述窗口计算结果可作为语义结构重组操作部分的数据流,将所述窗口计算结果存放至高速缓存中可以在后续使用时实现快速调用。
S340、基于设定的数据处理目的,对所得到的窗口计算结果进行语义结构重组,以得到语义结构重组数据流;
S350、基于重组流计算公式,对所述语义结构重组数据流进行重组流计算,得到数据处理结果。
在本实施例中,上述S340和S350与实施例一中的步骤S120和S130完全相同,这里不再详述。本发明实施例三中的S360~S390体现了本发明实施例的又一个优化,即增加了对已得出的数据处理结果的修正操作。
S360、收集后续请求流入所述各个数据计算窗口的数据,作为所述各个数据计算窗口的过期数据。
在本实施例中,由于对各个数据计算窗口进行窗口计算操作时输入的数据流为无序数据流,所以就会出现数据延迟输入的情况,由此在为所述各个数据计算窗口赋予窗口标记后仍存在数据流入对应数据计算窗口的情况。由于被赋予窗口标记的数据计算窗口被设定为拒绝所述数据流中属于所述数据计算窗口的后续数据流入,所以在出现上述情况时,无法对上述新流入的数据继续在原来的数据计算窗口中进行增量处理,然而,为了保证最终数据处理结果的正确性,需要对新流入的延迟数据进行特殊处理。
本实施例将窗口标记后请求流入数据计算窗口的延迟数据称为过期数据,下述步骤S370~S390体现出了对所述过期数据进行的特殊处理。
S370、对所述过期数据进行窗口计算,并对窗口计算结果进行语义结构重组得到语义结构重组数据补偿流。
在本实施例中,在确定过期数据后,对过期数据再次重新进行窗口计算,并在设定时间内输出窗口计算结果,以保证后续步骤的进行。可以理解的是,窗口标记后,需要实时监控所述数据计算窗口,并对所述数据计算窗口拒绝接收的过期数据进行实时收集和处理,以确保对最终数据处理结果的实时修正。
在本实施例中,对所述过期数据的处理过程可相当于对过期数据重新进行本发明实施例一中描述的三个操作步骤。由此,在对过期数据进行窗口计算操作数据窗口计算结果后,仍基于设定的数据处理目的对窗口计算结果进行语义 结构重组,以合成语义结构重组数据补偿流。
S380、对所述语义结构重组数据补偿流进行重组流计算,得到补偿计算结果。
在本实施例中,基于所合成的语义结构重组数据补偿流以及相应的重组计算公式,可以进一步实现重组流计算,并达到补偿计算结果。
S390、根据所述补偿计算结果对所述数据处理结果进行修正。
在本实施例中,处理计算得出的补偿计算结果可用于对之前所得出的数据处理结果的结果修正,具体的,所述修正的实现可基于实际情况设定相应的修正公式。
此外,在本实施例中,基于补偿计算结果对数据处理结果的修正时机可理解为在得到数据处理结果之后进行。具体的,对于任意两窗口计算结果流合成的语义结构重组流而言,在对该语义结构重组流中的数据处理结果进行修正时,无须等到该语义结构重组流中的所有数据都计算得到数据处理结果后才开始进行修正操作,而是只要该语义结构重组流中存在已得到数据处理结果的数据,且该数据又存在相应的补偿计算结果,就可以将该数据的数据处理结果进行实时修正。由此既保证了多流流式数据处理结果的准确性,还保证了多流流式数据处理的时效性。
本实施例的技术方案,提出了对数据处理结果进行修正的优化方案,由此在使得对海量多数据流进行高效处理计算的同时,实现对数据处理结果的修正,进而保证了多流流式数据处理方法的高效性和准确性。
实施例四
图4为本发明实施例四提供的一种多流流式数据的处理方法的流程图,本发明实施例以上述实施例一为基础进行优化,在本实施例中,在“基于窗口计 算公式进行窗口计算”之后,还优化包括了:修改第一语义协作元数据的取值,所述第一语义协作元数据的取值用于标识是否触发语义结构重组操作。
进一步的,本实施例将“基于设定的数据处理目的对所得到的窗口计算结果进行语义结构重组,以得到语义结构重组数据流”优化为:在所述第一语义协作元数据的取值符合语义结构重组操作的触发条件时,以流形式从获取高速缓存中获取所述窗口计算结果;基于设定的数据处理目的,对所获取的窗口计算结果流进行语义结构重组,以得到重组数据流。
此外,还优化包括了:修改第二语义协作元数据的取值,所述第二语义协作元数据的取值用于标识是否触发重组流计算操作。进一步的,在所述第二语义协作元数据的取值符合重组流计算操作的触发条件时,基于重组流计算公式,对所述语义结构重组数据流进行重组流计算,得到数据处理结果。
如图4所示,本发明实施例四提供的一种多流流式数据的处理方法,具体包括:
S410、对所获取的多流流式数据进行窗口化处理并基于窗口计算公式进行窗口计算。
在本实施例中,上述S410与实施例一中的步骤S110完全相同,这里不再详述。
S420、修改第一语义协作元数据的取值,所述第一语义协作元数据的取值用于标识是否触发语义结构重组操作。
在本实施例中,由于对数据计算窗口中的数据进行的是实时处理计算,且实时处理计算所得窗口计算结果直接存放于所述数据计算窗口在高速缓存中对应的位置,并可作为后续语义结构重组时的输入流。然而,对于语义结构重组的操作,该操作需要相应的触发条件才能进行。
在本实施例中,其触发条件可认为是在数据计算窗口被标记后,所设定的第一语义协作元数据的取值。在本实施例中,当数据计算窗口被标记后,不再有数据流入该窗口,相应的窗口计算结果值也不再实时改变,需要对多流流式数据进一步处理,即进行后续的语义结构重组,但是,在进行多流流式数据的语义结构重组之前,需要特定的触发操作来启动多流流式数据的语义结构重组。本实施例可以基于修改第一语义协作元数据的取值来确定是否触发启动语义结构重组。
在本实施例中,所述第一语义协作元数据的取值具体可看作一个触发下一处理步骤启动的标记位。示例性的,可允许所述第一语义协作元数据的取值为0或1,当取值为0时,则可认为无法触发下一步骤的启动,当取值为1时,则可认为可以触发下一步骤的启动。
S430、在所述第一语义协作元数据的取值符合语义结构重组操作的触发条件时,以流形式获取所述窗口计算结果。
在本实施例中,当S420中的第一语义协作元数据的取值被修改后,满足语义结构重组的触发条件时,以流形式获取所述窗口计算结果。
具体的,所述以流形式获取所述窗口计算结果可概述为:确定进行多流流式数据处理所需的至少两条输入流,所述至少两条输入流具体可看作通过数据计算窗口计算得到的至少两条数据流的窗口计算结果流。
需要说明的是,不同数据流的窗口计算结果分别存放在高速缓存的指定位置中,且不同的数据流除了指由不同源程序集在同一时间或不同时间产生的数据集合外,还可以指由一个源程序集在不同日期产生的数据集,示例性的,如同一源程序集在昨天产生的数据集合与今天产生的数据集合就可以属于不同数据流。
S440、基于设定的数据处理目的,对所获取的窗口计算结果流进行语义结构重组,以得到重组数据流。
在本实施例中,所述设定的数据处理目的具体可理解为当前所处理的多流流式数据需要达到的目的。所述对所获取的窗口计算结果流进行语义结构重组具体可理解为:对所述至少两条数据流的窗口计算结果进行语义结构重组。
在本实施例中,获取到至少两条数据流的窗口计算结果后,需要基于数据处理目的进行语义结构的重组,由此将至少两条数据流基于窗口计算结果合成一条重组数据流,所述重组数据流可用于计算多流流式数据的最终数据结果。
S450、修改第二语义协作元数据的取值,所述第二语义协作元数据的取值用于标识是否触发重组流计算操作。
在本实施例中,所述第二语义协作元数据与所述第一语义协作元数据的作用相同,都可理解为用于触发下一处理步骤的启动。不同的是,所述第二语义元数据具体可用于触发重组流计算操作。需要说明的是,所述窗口计算的操作、语义结构重组的操作以及重组流计算操作之间是相互关联,相互协作的,且只有在完成前一个操作后,才能触发后一个操作的执行。
示例性的,所述第二语义协作元数据的取值同样可看作一个触发标记位,其取值也可优选为0或1,且值为0时,表明不触发下一步操作,值为1时表明可以触发下一步操作。
S460、在所述第二语义协作元数据的取值符合重组流计算操作的触发条件时,基于重组流计算公式,对所述语义结构重组数据流进行重组流计算,得到数据处理结果。
在本实施例中,当第二语义协作元数据的取值符合重组流计算操作的触发条件后,就会启动对语义结构重组数据流的重组流计算操作。具体的,所述重 组流计算操作主要基于重组流计算公式,所述重组流计算公式具体为基于公式拆解策略所得。在进行重组流计算后,就可得到多流流式数据处理期望得到的数据处理结果。
本实施例的技术方案,具体化了实施例一中的实现步骤,基于本实施例提供的处理方法,既能够实现海量数据的实时处理,还能实现多数据流的处理计算,进而更加完善了对海量数据流进行处理计算的实现过程,从而满足了用户基于数据处理结果实时探索深层次信息或原因的诉求。
此外,需要说明的是,本实施例中的S410、S420~S440、以及S450~S460这3个大操作步骤之间是相互触发相互协作的,具体的,各大操作步骤之间可以按照各自间关联的语义协作元数据进行触发,并按照预处理阶段获取的语义结构进行相互间的协作,共同协作完成多流流式数据处理。其中,最初的窗口计算操作的数据流为无条件运行流,只要有数据流输入,就会持续运行,实时计算窗口计算结果,且窗口计算结果会实时存放至高速缓存中;语义结构重组流基于第一语义协作元数据的取值触发并启动,之后重组流计算操作基于第二语义协作元数据的取值触发并启动,最终得到所需的数据处理结果。
实施例五
图5a为本发明实施例五提供的一种多流流式数据的处理方法的优选实施例,本实施例五的应用场景是对百度罗盘产品中的关键词消费诊断功能,该应用场景所需要的数据是网民基于百度推广显示的关键词进行点击所产生的点击消费数据,该数据被称为凤巢点击消费数据。对于凤巢点击消费数据,每条数据所包含的信息内容可概括为:关键词、点击日期、点击时间以及点击产生的消费金额等。
对于应用场景中所提的关键词消费诊断,基于本发明实施例五所提供的处 理方法,可以确定其进行数据处理的目的是:获得所有关键词在今日整点T累计消费环比昨日整点T累计消费的突降值,以使用户基于得到的数据处理结果可以实时查询自身账户下所有关键词在整点T累计消费情况以及环比情况。
需要说明的是,本实施例五提供的优选实施例,最初的输入流为凤巢点击消费数据流,该数据流的特点是:1)数据量大和2)实时性强。作为全球最大的中文搜索引擎,每时每刻都有网民在百度搜索引擎中基于关键词搜索,因此,时刻就有用户基于搜索的关键词点击百度推广中所显示的广告,用户在基于关键词点击时,后台就产生了点击消费数据。由此可以看出作为数据流的凤巢点击数据的海量的实时的数据流。此外,本实施例所提供的优选实施例还体现了对多数据流处理的特点,即,对于环比计算,需要基于今天和昨天的数据信息,因此符合多数据流处理的范畴。
如图5a所示,本发明实施例五提供的一种优选的多流流式数据的处理方法,具体包括:
S510、基于公式拆解策略拆解多流计算的计算公式。
在本实施例中,所述多流计算公式的拆解过程可表述为:
1)基于获取的关键词整点T累计消费环比突降值这个输入指令,确定用于多流流式数据处理的多流计算公式和语义结构;2)对所述环比突降值的计算公式进行公式拆解,以得到窗口计算公式和重组流计算公式。
示例性的,由于本实施例进行的是多流流式数据处理,需要在对多流流式数据进行处理之前,先基于获取的输入指令确定用于多流流式数据处理的多流计算公式和语义结构,在本实施例中,可以确定其语义结构是计算关键词的今日整点T累计消费值和昨日整点T累计消费值。之后还可基于输入指令确定其多流计算公式是:计算关键词的整点T累计消费值,然后计算关键词的今日整 点T累计消费值与昨日整点T累计消费值的差。
分析上述确定的多流计算公式,可知上述计算公式需要分步计算才能实现,因此基于公式拆分策略将上述多流计算公式拆分为用于单数据流计算的窗口计算公式和用于多数据流计算的重组流计算公式。示例性的,上述公式拆分的窗口计算公式为:对整点T时间段流入的关键词进行消费值的累计求和;其拆分出的重组流计算公式为:计算关键词今日整点T时间段的累计消费值与关键词昨日整点T时间段的累计消费值。需要说明的是,本实施例中所提的关键词消费值的计算具体指同一关键词消费值的计算。此外,基于输入指令确定的语义结构和基于公式拆分策略拆分出的计算公式均对应存储到任务配置文件中,当基于处理方法进行数据处理时可以在进行相应的操作步骤时快速调用该操作步骤所需的计算公式。
S520、按小时级时间属性划分凤巢点击消费数据流中的数据到小时级窗口。
在本实施例中,在对基于输入指令确定多流计算公式进行公式拆分操作后,可以将凤巢点击消费数据流作为输入流进行获取,并在获取到凤巢点击消费数据流之后按小时级的时间属性对数据流的数据进行数据计算窗口划分。
示例性的,所述小时级时间属性具体可指一天划分为24个时间段,每个时间段为一个小时级窗口,然后基于凤巢点击数据的点击时间确定将数据划分到哪个小时级窗口。例如,当前获取的关键词点击时间为0点24份,则可将该关键词及其数据信息划分到0点这个小时级窗口中。
S530、对小时级窗口中的数据实时进行关键词整点累计消费值计算,并输出计算结果到高速缓存中。
示例性的,在进行数据计算窗口划分后,需要基于窗口计算公式对数据计算窗口中的数据进行实时计算,并将窗口计算结果实时输出到高速缓存中。例 如,在0点这个小时级窗口中,对划分至这个窗口的关键词进行消费值的累计计算,对于关键词A,在0点1分,0点4分,以及0点8分产生了点击消费金额分别为1元,1元,和1元,则截止到0点9分,关键词A的消费值累计计算值为3元,此时高速缓存中存放的计算结果就为3元。对于其他关键词,其他小时级窗口的计算过程与上述描述相同,这里不再详述。
S540、判断是否对小时级窗口进行标记,若是,则分别执行S550和S551;若否,则返回执行S530。
在本实施例中,为了节省多流流式数据处理时间,可以设定冻结小时级窗口的冻结时机(所述冻结可理解为拒绝数据流中属于该窗口的后续数据的流入请求),而该冻结时机的设定可以基于窗口标记进行。示例性的,由于是实时的数据计算,对0点小时级窗口来说,输入的数据流划分到该小时级窗口的时间可以大致确定为0点到1点之间,在过了1点之后,所获取的绝大部分数据流将会基于时间属性划分到1点小时级窗口中。由此,可以通过设定标记策略对0点小时级窗口进行窗口标记,以标识可以冻结该小时级窗口。
一般地,在设定标记策略时通常基于本实施例的处理方法所处的应用场景设定,例如,针对本实施例中提到的关键词消费诊断这个应用场景,可以将标记策略设置为定时标记,例如,设置标记策略为:在1点+5分钟时,即1点05分时对0点小时级窗口进行窗口标记,由此冻结该小时级窗口。
在本实施例中,判断小时级窗口是否基于标记策略的进行了窗口标记,如果没有进行窗口标记,则需要返回S530继续进行实时增量计算,如果符合标记策略进行了窗口标记,则可分别执行两个分支步骤,这两个分支步骤可概括为基于S551进行过期数据的补偿处理计算,以及基于S550~S580进行的实时数据的多流处理计算。
S551、将窗口标记后请求流入小时级窗口的数据确定为过期数据,对过期数据进行实时分形补偿计算,得到实施补偿计算结果,之后执行S590。
在本实施例中,在基于标记策略对小时级窗口进行标记后,由于存在数据延迟的情况,所以还会有属于所述小时级窗口的数据没有被及时划分到其中,进而在窗口标记后又无法流入到该小时级窗口中,可称这些数据为过期数据。然而,由于这些过期数据不能在原来的小时级窗口中继续进行窗口计算,但为了保证多流流式数据处理结果的准确性,也不能直接舍弃这些过期数据,因此需要设置特定的步骤对这些过期数据进行分形补偿处理。
示例性的,对过期数据进行实时分形补偿计算的过程可概括为:获取过期数据,对过期数据重新基于窗口计算公式进行实时增量计算,并得到实时增量计算的窗口计算结果;然后,对过期数据流的今日关键词整点累计消费值以及上述过期数据流对应的昨日关键词整点累计消费值进行基于数据处理目的进行语义结构重组,合成语义结构重组补偿流;最后基于重组计算公式对所述语义结构重组补偿流进行重组流计算,并最终得到补偿计算结果。在本实施例中,补偿计算结果的获得过程是实时进行的,只要有过期数据出现,就进行上述操作。
S550、获取高速缓存中的今日关键词整点消费累计值数据流和昨日关键词整点消费累计值数据流。
S560、基于数据处理目的,将获取的两条数据流进行语义结构重组,合成语义结构重组流。
S570、基于拆分出的环比计算公式,对语义结构重组流进行重组流计算,获得数据处理结果。
在本实施例中,S550~S570是小时级窗口标记后需要执行的另一个分支步 骤,首先,在基于S530确定小时级窗口进行标记后,需要修改与语义结构重组相关联的第一语义协作元数据的取值,以基于第一语义协作元数据的取值确定是否触发进行语义结构重组的操作,并在触发启动执行语义结构重组后,再获取高速缓存中存放的今日关键词整点消费累计值数据流和昨日关键词整点消费累计值数据流,基于应用场景中的数据处理目的,将获取的两条数据流进行语义结构重组,获得合成后的语义结构重组流;然后,在合成语义结构重组流之后,还需要修改与重组流计算操作相关联的第二语义协作元数据的取值,以基于第二语义协作元数据的取值确定是否触发进行重组流计算的操作,并在触发启动执行重组流计算之后,对合成的语义结构重组流进行重组流计算。示例性的,重组流计算就是进行今日关键词整点消费累计值与昨日关键词整点消费累计值的差计算,计算出的差值就可看作该应用场景下进行多流流式数据处理的数据处理结果。
需要注意的是,本实施例中提出的第一语义协作元数据和第二语义协作元数据并没有具体体现在本实施例五的流程图中,而是以语义协作元数据与相应操作步骤之间的数据信息的交流来体现,但所起的作用和要表达的意思并不因文字表述的改变而改变。
S580、在数据处理结果输出后,基于实时补偿计算结果修正数据处理结果,并实时输出修正后处理结果。
在本实施例中,基于窗口计算操作、语义结构重组流的合成操作以及重组流计算操作进行多流流式数据处理之后,就可以数据初始的数据处理结果。所述数据处理结果可看作该应用场景下的多流流式数据的处理结果,然而,为了保证数据处理结果的准确性,需要基于实时计算得出的补偿计算结果进行数据处理结果修正,并最终输出修正后的处理结果。需要说明的是,对数据处理结 果的修正操作,也需要基于与结果修正操作关联的语义协作元数据的触发,示例性的,其触发过程可参考图5a中语义协作元数据与重组计算操作之间的数据信息交流来完成。
在关键词消费诊断这个应用场景中,基于上述提供的多流流式数据处理方法,最终输出的结果是各关键词在整点T的累计消费值,以及各关键词在整点T的累计消费值的环比突降值。
图5b为本发明实施例五基于多流流式数据处理方法在关键词消费诊断中体现出的应用效果图。如图5b所示,其显示结果为任一用户ID中任一指定关键词在整点T时间段内的数据累计消费值的曲线分布,以及该关键词在整点T时间段累计消费值的环比突降值的曲线分布,其中,T∈[0点,23点]。需要说明的是,图5b中所显示的信息内容仅可看作在某个时间点对某一关键词的数据处理结果的截取片段,在实际应用中,若该关键词存在补偿计算结果,则在不同时刻,该关键词所得到的数据处理结果会基于补偿计算结果的修正而产生变化。
本实施例的技术方案,提供了一种多流流式数据的处理方法的优选实施例,基于该优选实施例,更清楚了多流流式数据的处理方法的操作步骤,基于该处理方法,高效的实现了海量数据的实时处理计算,节省了数据的收集时间,加快了数据处理进度,且通过设定补偿处理计算,进一步保证多流计算结果的准确性。
实施例六
图6a为本发明实施例六提供的一种多流流式数据的处理系统的结构框图,本实施例的处理系统可由软件和/或硬件实现,可适用于对海量的多流流式数据进行数据处理的情况,并一般可集成在流式数据分析的框架平台上,所述框架 平台可以是一种服务器。如图6a所示,该处理系统包括:窗口化处理模块60、窗口计算模块61、数据流重组模块62以及重组流计算模块63。
其中,窗口化处理模块60用于对所获取的多流流式数据进行窗口化处理。
窗口计算模块61,用于基于窗口计算公式对经过窗口化处理后的多流流式数据进行窗口计算。
数据流重组模块62用于基于设定的数据处理目的,对所得到的窗口计算结果进行语义结构重组,以得到语义结构重组数据流。
重组流计算模块63用于基于重组流计算公式,对所述语义结构重组数据流进行重组流计算,得到数据处理结果。
在本实施例中,该处理系统首先通过窗口化处理模块60将对所获取的多流流式数据进行窗口化处理,窗口计算模块61基于窗口计算公式进行窗口计算;然后,通过数据流重组模块62基于设定的数据处理目的,对所得到的窗口计算结果进行语义结构重组,以得到语义结构重组数据流;最终,通过重组流计算模块63基于重组流计算公式,对所述语义结构重组数据流进行重组流计算,得到数据处理结果。
本实施例的技术方案,有效的实现了实时处理涉及海量实体的多流流式计算,由此不仅解决大数据处理中耗时长,异常处理成本高的问题,还解决了不能对多数据流进行处理计算的问题,进而更加完善了对海量数据流进行处理计算的实现过程,从而满足了用户基于数据处理结果实时探索深层次信息或原因的诉求。
进一步的,所述处理系统还可以包括:
第一预处理模块,用于在对所获取的多流流式数据进行窗口化处理并基于窗口计算公式进行窗口计算之前,基于用户输入的指令,确定所述窗口计算公 式、所述重组流计算公式和语义结构。
进一步的,所述处理系统还可以包括:
第二预处理模块,用于将所述语义结构、所述窗口计算公式和所述重组流计算公式存放至任务配置文件中。
在上述实施例的技术上,所述窗口化处理模块60具体可以用于:
基于所获取的多流流式数据的数据属性,将所获取的多流流式数据划分到各个数据计算窗口中;为所述各个数据计算窗口赋予窗口标记,其中被赋予窗口标记的数据计算窗口被设定为拒绝所述数据流中属于所述数据计算窗口的后续数据流入。
进一步的,所述处理系统还可以包括:
第一元数据值修改模块,用于在基于窗口计算公式进行窗口计算之后,修改第一语义协作元数据的取值,所述第一语义协作元数据的取值用于标识是否触发语义结构重组操作。
在上述实施例的基础上,所述数据流重组模块62具体可以用于:
在所述第一语义协作元数据的取值符合语义结构重组操作的触发条件时,以流形式获取所述窗口计算结果;基于设定的数据处理目的,对所获取的窗口计算结果流进行语义结构重组,以得到重组数据流。
进一步的,所述处理系统还可以包括:
第二元数据值修改模块,用于修改第二语义协作元数据的取值,所述第二语义协作元数据的取值用于标识是否触发重组流计算操作。
相应的,所述重组流计算模块63具体用于:
在所述第二语义协作元数据的取值符合重组流计算操作的触发条件时,基 于重组流计算公式,对所述语义结构重组数据流进行重组流计算,得到数据处理结果。
进一步的,所述处理系统还可以包括过期数据补偿处理模块,用于:在为所述各个数据计算窗口赋予窗口标记之后,收集后续请求流入所述各个数据计算窗口的数据,作为所述各个数据计算窗口的过期数据;对所述过期数据进行窗口计算,并对窗口计算结果进行语义结构重组得到语义结构重组数据补偿流;对所述语义结构重组数据补偿流进行重组流计算,得到补偿计算结果;根据所述补偿计算结果对所述数据处理结果进行修正。
此外,图6b为本发明实施例七提供的一种多流流式数据处理的构架图,图6b所示的构架图,将本发明实施例提供的多流流式数据处理系统的工作流程及其之间的协作关系更加清晰化。如图6b所示,该构架图具体可分为4条相互关联的语义流,分别是:窗口统计分析流64、语义结构重组流65、语义结构计算流66、以及分形补偿流67。上述4条语义流,窗口统计分析流64对应多流流式数据处理系统的窗口化处理模块60和窗口计算模块61,其余3条语义流分别对应多流流式数据处理系统的数据流重组模块62、重组流计算模块63以及过期数据补偿处理模块。图6b可以看出,各语义流在相互协作的基础上又独立完成各自对应模块的功能步骤,最终实现对多流流式数据的处理。
实施例七
本发明实施例七提供了一种设备,包括本发明任意实施例所提供的多流流式数据的处理系统,且该设备可以作为服务器。具体的,如图7所示,本发明实施例提供一种设备,该设备包括:处理器71、存储器72、输入装置73和输出装置74;设备中处理器71的数量可以是一个或多个,图7中以一个处理器 71为例;所述设备中的处理器71、存储器72、输入装置73和输出装置74可以通过总线或其他方式连接,图7中以通过总线连接为例。
存储器72作为一种计算机可读存储介质,可用于存储软件程序、计算机可执行程序以及模块,如本发明实施例中的多流流式数据的处理方法对应的程序指令/模块(例如,附图6a所示的多流流式数据的处理系统中的窗口化处理模块60、窗口计算模块61、数据流重组模块62以及重组流计算模块63)。处理器71通过运行存储在存储器72中的软件程序、指令以及模块,从而执行设备的各种功能应用以及数据处理,即实现上述方法实施例中的多流流式数据的处理方法。
存储器72可包括存储程序区和存储数据区,其中,存储程序区可存储操作系统、至少一个功能所需的应用程序;存储数据区可存储根据设备的使用所创建的数据等。此外,存储器72可以包括高速随机存取存储器,还可以包括非易失性存储器,例如至少一个磁盘存储器件、闪存器件、或其他非易失性固态存储器件。在一些实例中,存储器72可进一步包括相对于处理器71远程设置的存储器,这些远程存储器可以通过网络连接至设备。上述网络的实例包括但不限于互联网、企业内部网、局域网、移动通信网及其组合。
输入装置73可用于接收输入的数字或字符信息,以及产生与设备的用户设置以及功能控制有关的键信号输入。输出装置74可包括输出端口等。
并且,上述设备可以包括一个或者多个程序,所述一个或者多个程序存储在存储器72中,当被所述一个或者多个处理器71执行时,程序进行如下操作:
对所获取的多流流式数据进行窗口化处理并基于窗口计算公式进行窗口计算;
基于设定的数据处理目的,对所得到的窗口计算结果进行语义结构重组, 以得到语义结构重组数据流;
基于重组流计算公式,对所述语义结构重组数据流进行重组流计算,得到数据处理结果。
本发明实施例还提供了一种包含计算机可执行指令的存储介质,所述计算机可执行指令在由计算机处理器执行时用于执行一种多流流式数据的处理方法,该方法包括:
对所获取的多流流式数据进行窗口化处理并基于窗口计算公式进行窗口计算;
基于设定的数据处理目的,对所得到的窗口计算结果进行语义结构重组,以得到语义结构重组数据流;
基于重组流计算公式,对所述语义结构重组数据流进行重组流计算,得到数据处理结果。
通过以上关于实施方式的描述,所属领域的技术人员可以清楚地了解到,本发明可借助软件及必需的通用硬件来实现,当然也可以通过硬件实现,但很多情况下前者是更佳的实施方式。基于这样的理解,本发明的技术方案本质上或者说对现有技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品可以存储在计算机可读存储介质中,如计算机的软盘、只读存储器(Read-Only Memory,ROM)、随机存取存储器(Random Access Memory,RAM)、闪存(FLASH)、硬盘或光盘等,包括若干指令用以使得一台计算机设备(可以是个人计算机,服务器,或者网络设备等)执行本发明各个实施例所述的方法。
值得注意的是,上述多流流式数据的处理系统的实施例中,所包括的各个 单元和模块只是按照功能逻辑进行划分的,但并不局限于上述的划分,只要能够实现相应的功能即可;另外,各功能单元的具体名称也只是为了便于相互区分,并不用于限制本发明的保护范围。
注意,上述仅为本发明的较佳实施例及所运用技术原理。本领域技术人员会理解,本发明不限于这里所述的特定实施例,对本领域技术人员来说能够进行各种明显的变化、重新调整和替代而不会脱离本发明的保护范围。因此,虽然通过以上实施例对本发明进行了较为详细的说明,但是本发明不仅仅限于以上实施例,在不脱离本发明构思的情况下,还可以包括更多其他等效实施例,而本发明的范围由所附的权利要求范围决定。

Claims (18)

  1. 一种多流流式数据的处理方法,其特征在于,包括:
    对所获取的多流流式数据进行窗口化处理并基于窗口计算公式进行窗口计算;
    基于设定的数据处理目的,对所得到的窗口计算结果进行语义结构重组,以得到语义结构重组数据流;
    基于重组流计算公式,对所述语义结构重组数据流进行重组流计算,得到数据处理结果。
  2. 根据权利要求1所述的方法,其特征在于,对所获取的多流流式数据进行窗口化处理并基于窗口计算公式进行窗口计算之前,所述方法还包括:
    基于用户输入的指令,确定所述窗口计算公式、所述重组流计算公式和语义结构。
  3. 根据权利要求2所述的方法,其特征在于,所述方法还包括:
    将所述语义结构、所述窗口计算公式和所述重组流计算公式存放至任务配置文件中。
  4. 根据权利要求1所述的方法,其特征在于,对所获取的多流流式数据进行窗口化处理包括:
    基于所获取的多流流式数据的数据属性,将所获取的多流流式数据划分到各个数据计算窗口中;
    为所述各个数据计算窗口赋予窗口标记,其中被赋予窗口标记的数据计算窗口被设定为拒绝所述数据流中属于所述数据计算窗口的后续数据流入。
  5. 根据权利要求1-4任一项所述的方法,其特征在于,基于窗口计算公式进行窗口计算之后,所述方法还包括:
    修改第一语义协作元数据的取值,所述第一语义协作元数据的取值用于标 识是否触发语义结构重组操作。
  6. 根据权利要求5所述的方法,其特征在于,基于设定的数据处理目的对所得到的窗口计算结果进行语义结构重组,以得到语义结构重组数据流,包括:
    在所述第一语义协作元数据的取值符合语义结构重组操作的触发条件时,以流形式获取所述窗口计算结果;
    基于设定的数据处理目的,对所获取的窗口计算结果流进行语义结构重组,以得到重组数据流。
  7. 根据权利要求6所述的方法,其特征在于,所述方法还包括:
    修改第二语义协作元数据的取值,所述第二语义协作元数据的取值用于标识是否触发重组流计算操作,以及
    其中,在所述第二语义协作元数据的取值符合重组流计算操作的触发条件时,基于重组流计算公式,对所述语义结构重组数据流进行重组流计算,得到数据处理结果。
  8. 根据权利要求4所述的方法,其特征在于,在为所述各个数据计算窗口赋予窗口标记之后,所述方法还包括:
    收集后续请求流入所述各个数据计算窗口的数据,作为所述各个数据计算窗口的过期数据;
    对所述过期数据进行窗口计算,并对窗口计算结果进行语义结构重组得到语义结构重组数据补偿流;
    对所述语义结构重组数据补偿流进行重组流计算,得到补偿计算结果;
    根据所述补偿计算结果对所述数据处理结果进行修正。
  9. 一种多流流式数据的处理系统,其特征在于,包括:
    窗口化处理模块,用于对所获取的多流流式数据进行窗口化处理;
    窗口计算模块,用于基于窗口计算公式对窗口化处理后的多流流式数据进行窗口计算;
    数据流重组模块,用于基于设定的数据处理目的,对所得到的窗口计算结果进行语义结构重组,以得到语义结构重组数据流;
    重组流计算模块,用于基于重组流计算公式,对所述语义结构重组数据流进行重组流计算,得到数据处理结果。
  10. 根据权利要求9所述的系统,其特征在于,所述系统还包括:
    第一预处理模块,用于在对所获取的多流流式数据进行窗口化处理并基于窗口计算公式进行窗口计算之前,基于用户输入的指令,确定所述窗口计算公式、所述重组流计算公式和语义结构。
  11. 根据权利要求10所述的系统,其特征在于,所述系统还包括:
    第二预处理模块,用于将所述语义结构、所述窗口计算公式和所述重组流计算公式存放至任务配置文件中。
  12. 根据权利要求9所述的系统,其特征在于,所述窗口化处理模块,具体用于:
    基于所获取的多流流式数据的数据属性,将所获取的多流流式数据划分到各个数据计算窗口中;
    为所述各个数据计算窗口赋予窗口标记,其中被赋予窗口标记的数据计算窗口被设定为拒绝所述数据流中属于所述数据计算窗口的后续数据流入。
  13. 根据权利要求9-12任一所述的系统,其特征在于,所述系统还包括:
    第一元数据值修改模块,用于在基于窗口计算公式进行窗口计算之后,修改第一语义协作元数据的取值,所述第一语义协作元数据的取值用于标识是否触发语义结构重组操作。
  14. 根据权利要求13所述的系统,其特征在于,所述数据流重组模块,具体用于:
    在所述第一语义协作元数据的取值符合语义结构重组操作的触发条件时,以流形式获取所述窗口计算结果;
    基于设定的数据处理目的,对所获取的窗口计算结果流进行语义结构重组,以得到重组数据流。
  15. 根据权利要求14所述的系统,其特征在于,所述系统还包括:
    第二元数据值修改模块,用于修改第二语义协作元数据的取值,所述第二语义协作元数据的取值用于标识是否触发重组流计算操作;
    相应的,所述重组流计算模块具体用于:
    在所述第二语义协作元数据的取值符合重组流计算操作的触发条件时,基于重组流计算公式,对所述语义结构重组数据流进行重组流计算,得到数据处理结果。
  16. 根据权利要求12所述的系统,其特征在于,所述系统还包括过期数据补偿处理模块,用于:
    在为所述各个数据计算窗口赋予窗口标记之后,收集后续请求流入所述各个数据计算窗口的数据,作为所述各个数据计算窗口的过期数据;
    对所述过期数据进行窗口计算,并对窗口计算结果进行语义结构重组得到语义结构重组数据补偿流;
    对所述语义结构重组数据补偿流进行重组流计算,得到补偿计算结果;根据所述补偿计算结果对所述数据处理结果进行修正。
  17. 一种包含计算机可执行指令的存储介质,所述计算机可执行指令在由 计算机处理器执行时用于执行一种多流流式数据的处理方法,其特征在于,该方法包括:
    对所获取的多流流式数据进行窗口化处理并基于窗口计算公式进行窗口计算;
    基于设定的数据处理目的,对所得到的窗口计算结果进行语义结构重组,以得到语义结构重组数据流;
    基于重组流计算公式,对所述语义结构重组数据流进行重组流计算,得到数据处理结果。
  18. 一种设备,其特征在于,包括:
    一个或者多个处理器;
    存储器;
    一个或者多个程序,所述一个或者多个程序存储在所述存储器中,当被所述一个或者多个处理器执行时,进行如下操作:
    对所获取的多流流式数据进行窗口化处理并基于窗口计算公式进行窗口计算;
    基于设定的数据处理目的,对所得到的窗口计算结果进行语义结构重组,以得到语义结构重组数据流;
    基于重组流计算公式,对所述语义结构重组数据流进行重组流计算,得到数据处理结果。
PCT/CN2016/097416 2016-04-25 2016-08-30 一种多流流式数据的处理方法、系统、存储介质及设备 Ceased WO2017185576A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201610262701.1 2016-04-25
CN201610262701.1A CN107305501B (zh) 2016-04-25 2016-04-25 一种多流流式数据的处理方法和系统

Publications (1)

Publication Number Publication Date
WO2017185576A1 true WO2017185576A1 (zh) 2017-11-02

Family

ID=60150458

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2016/097416 Ceased WO2017185576A1 (zh) 2016-04-25 2016-08-30 一种多流流式数据的处理方法、系统、存储介质及设备

Country Status (2)

Country Link
CN (1) CN107305501B (zh)
WO (1) WO2017185576A1 (zh)

Cited By (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111143415A (zh) * 2019-12-26 2020-05-12 政采云有限公司 一种数据处理方法、装置和计算机可读存储介质
CN111367951A (zh) * 2020-02-29 2020-07-03 深圳前海微众银行股份有限公司 一种流数据处理的方法及装置
CN112202607A (zh) * 2020-09-28 2021-01-08 中移(杭州)信息技术有限公司 日志消息的统计计算方法、服务器及存储介质
CN113850929A (zh) * 2021-09-18 2021-12-28 广州文远知行科技有限公司 一种标注数据流处理的展示方法、装置、设备和介质
CN114185885A (zh) * 2021-11-05 2022-03-15 中国科学院计算技术研究所 一种基于列存数据库的流式数据处理方法及系统
WO2023077451A1 (zh) * 2021-11-05 2023-05-11 中国科学院计算技术研究所 一种基于列存数据库的流式数据处理方法及系统

Families Citing this family (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN108038085A (zh) * 2017-12-27 2018-05-15 世纪龙信息网络有限责任公司 实时任务的数据校准方法及装置
CN109345377B (zh) * 2018-09-28 2020-03-27 北京九章云极科技有限公司 一种数据实时处理系统及数据实时处理方法
CN109375923B (zh) * 2018-10-26 2022-05-03 网易(杭州)网络有限公司 变更数据处理方法、装置、存储介质、处理器及服务器
CN113127512B (zh) * 2020-01-15 2023-09-29 百度在线网络技术(北京)有限公司 多数据流的数据拼接触发方法、装置、电子设备和介质
CN116881610B (zh) * 2023-09-08 2024-01-09 国网信息通信产业集团有限公司 能源设备量测项数据流式计算方法、装置、设备及介质

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2009199507A (ja) * 2008-02-25 2009-09-03 Nippon Telegr & Teleph Corp <Ntt> 類似部分シーケンス検出方法、類似部分シーケンス検出プログラム、および、類似部分シーケンス検出装置
CN103136217A (zh) * 2011-11-24 2013-06-05 阿里巴巴集团控股有限公司 一种分布式数据流处理方法及其系统
CN103914531A (zh) * 2014-03-31 2014-07-09 百度在线网络技术(北京)有限公司 数据的处理方法及装置
CN104915247A (zh) * 2015-04-29 2015-09-16 上海瀚银信息技术有限公司 一种实时数据计算方法及系统

Family Cites Families (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN103997464B (zh) * 2005-04-29 2018-07-24 艾利森电话股份有限公司 用于控制分组交换数据流中的实时连续数据的系统和方法
US8056065B2 (en) * 2007-09-26 2011-11-08 International Business Machines Corporation Stable transitions in the presence of conditionals for an advanced dual-representation polyhedral loop transformation framework
US8478775B2 (en) * 2008-10-05 2013-07-02 Microsoft Corporation Efficient large-scale filtering and/or sorting for querying of column based data encoded structures
US8180801B2 (en) * 2009-07-16 2012-05-15 Sap Ag Unified window support for event stream data management
CN103473636B (zh) * 2013-09-03 2017-08-08 沈效国 一种收集、分析和分发网络商业信息的系统数据组件
CN105335218A (zh) * 2014-07-03 2016-02-17 北京金山安全软件有限公司 一种基于本地的流式计算方法及流式计算系统

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2009199507A (ja) * 2008-02-25 2009-09-03 Nippon Telegr & Teleph Corp <Ntt> 類似部分シーケンス検出方法、類似部分シーケンス検出プログラム、および、類似部分シーケンス検出装置
CN103136217A (zh) * 2011-11-24 2013-06-05 阿里巴巴集团控股有限公司 一种分布式数据流处理方法及其系统
CN103914531A (zh) * 2014-03-31 2014-07-09 百度在线网络技术(北京)有限公司 数据的处理方法及装置
CN104915247A (zh) * 2015-04-29 2015-09-16 上海瀚银信息技术有限公司 一种实时数据计算方法及系统

Cited By (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111143415A (zh) * 2019-12-26 2020-05-12 政采云有限公司 一种数据处理方法、装置和计算机可读存储介质
CN111143415B (zh) * 2019-12-26 2023-12-29 政采云有限公司 一种数据处理方法、装置和计算机可读存储介质
CN111367951A (zh) * 2020-02-29 2020-07-03 深圳前海微众银行股份有限公司 一种流数据处理的方法及装置
CN112202607A (zh) * 2020-09-28 2021-01-08 中移(杭州)信息技术有限公司 日志消息的统计计算方法、服务器及存储介质
CN112202607B (zh) * 2020-09-28 2022-06-14 中移(杭州)信息技术有限公司 日志消息的统计计算方法、服务器及存储介质
CN113850929A (zh) * 2021-09-18 2021-12-28 广州文远知行科技有限公司 一种标注数据流处理的展示方法、装置、设备和介质
CN113850929B (zh) * 2021-09-18 2023-05-26 广州文远知行科技有限公司 一种标注数据流处理的展示方法、装置、设备和介质
CN114185885A (zh) * 2021-11-05 2022-03-15 中国科学院计算技术研究所 一种基于列存数据库的流式数据处理方法及系统
WO2023077451A1 (zh) * 2021-11-05 2023-05-11 中国科学院计算技术研究所 一种基于列存数据库的流式数据处理方法及系统

Also Published As

Publication number Publication date
CN107305501B (zh) 2020-11-17
CN107305501A (zh) 2017-10-31

Similar Documents

Publication Publication Date Title
CN107305501B (zh) 一种多流流式数据的处理方法和系统
US8442863B2 (en) Real-time-ready behavioral targeting in a large-scale advertisement system
CN104298771B (zh) 一种海量web日志数据查询与分析方法
US20190230000A1 (en) Intelligent analytic cloud provisioning
JP6190255B2 (ja) グラフデータの再帰クエリを用いたストリームデータ処理方法
Kolchinsky et al. Lazy evaluation methods for detecting complex events
US10019675B2 (en) Actuals cache for revenue management system analytics engine
CN113568938A (zh) 数据流处理方法、装置、电子设备及存储介质
CN107623639A (zh) 基于emd距离的数据流分布式相似性连接方法
CN106446134B (zh) 基于谓词规约和代价估算的局部多查询优化方法
CN106296286A (zh) 广告点击率的预估方法和预估装置
Backman et al. C-MR: continuously executing MapReduce workflows on multi-core processors
CN113407587B (zh) 用于联机分析处理引擎的数据处理方法、装置、设备
CN104951509A (zh) 一种大数据在线交互式查询方法及系统
US20200065412A1 (en) Predicting queries using neural networks
CN119046124B (zh) 分布式系统的代价评估方法、装置、设备、介质及产品
CN104598474B (zh) 云环境下基于数据语义的信息推荐方法
CN118503512A (zh) 一种面向大规模网络舆情的Elasticsearch检索优化系统
CN116450673B (zh) 数据处理方法、电子设备及计算机存储介质
CN106599122A (zh) 一种基于垂直分解的并行频繁闭序列挖掘方法
Huang et al. Burst topic discovery and trend tracing based on Storm
CN108932334A (zh) 一种基于时间序列存储模型扩展以及匹配优化方法
CN108804224A (zh) 一种基于Spark框架的中间数据权重设置方法
CN116126238A (zh) 数据存储方法、系统、装置及非易失性存储介质
CN103500219B (zh) 一种标签自适应精准匹配的控制方法

Legal Events

Date Code Title Description
NENP Non-entry into the national phase

Ref country code: DE

121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 16900085

Country of ref document: EP

Kind code of ref document: A1

122 Ep: pct application non-entry in european phase

Ref document number: 16900085

Country of ref document: EP

Kind code of ref document: A1