WO2020119064A1 - 互联网信息链式存储方法、装置、计算机设备及存储介质 - Google Patents
互联网信息链式存储方法、装置、计算机设备及存储介质 Download PDFInfo
- Publication number
- WO2020119064A1 WO2020119064A1 PCT/CN2019/092551 CN2019092551W WO2020119064A1 WO 2020119064 A1 WO2020119064 A1 WO 2020119064A1 CN 2019092551 W CN2019092551 W CN 2019092551W WO 2020119064 A1 WO2020119064 A1 WO 2020119064A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- information
- file
- text file
- recognition model
- data information
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/90—Details of database functions independent of the retrieved data types
- G06F16/95—Retrieval from the web
- G06F16/958—Organisation or management of web site content, e.g. publishing, maintaining pages or automatic linking
Definitions
- the present application relates to the field of computer technology, and in particular to an Internet information chain storage method, device, computer equipment, and storage medium.
- Massive data information is stored on each web page on the Internet, and the newly added data information will gradually replace the data information saved on the web page, causing the data information on the web page to change and change. Therefore, the existing data information on the Internet
- the storage method cannot obtain data information that has been deleted or modified on the Internet. In judicial practice, it is extremely difficult to obtain evidence from relevant data information published on the Internet. Therefore, existing data information storage methods cannot obtain deleted data information.
- Embodiments of the present application provide an Internet information chain storage method, device, computer device, and storage medium, which are intended to solve the problem that the data information storage method in the prior art cannot obtain deleted data information.
- an embodiment of the present application provides an Internet information chain storage method, which includes:
- an internet information chain storage device which includes:
- the webpage monitoring unit is used to obtain the webpage information of the webpage to be monitored, and real-time monitoring the data information published in the webpage to be monitored according to the webpage information of the webpage to be monitored to obtain new data information; the judgment unit is used to add new data information Whether the file in is a text file is judged; the information conversion unit is used to convert the non-text file into a text file through a preset information recognition model if the file in the newly added data information is a non-text file; the information storage unit is used to Save the text file in the newly added data information and/or the converted text file to the preset data list.
- an embodiment of the present application further provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor executing the computer
- the program implements the Internet information chain storage method described in the first aspect above.
- an embodiment of the present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor causes the processor to execute the first On the one hand, the Internet information chain storage method.
- FIG. 1 is a schematic flowchart of an Internet information chain storage method provided by an embodiment of this application.
- FIG. 2 is a schematic diagram of a sub-process of an Internet information chain storage method provided by an embodiment of the present application
- FIG. 3 is a schematic diagram of another sub-process of the Internet information chain storage method provided by an embodiment of the present application.
- FIG. 4 is a schematic diagram of another sub-process of the Internet information chain storage method provided by an embodiment of the present application.
- FIG. 5 is a schematic diagram of another sub-process of the Internet information chain storage method provided by an embodiment of the present application.
- FIG. 6 is a schematic block diagram of an internet information chain storage device provided by an embodiment of the present application.
- FIG. 7 is a schematic block diagram of a subunit of an Internet information chain storage device provided by an embodiment of the present application.
- FIG. 8 is a schematic block diagram of another subunit of an Internet information chain storage device provided by an embodiment of the present application.
- FIG. 9 is a schematic block diagram of another subunit of an Internet information chain storage device provided by an embodiment of this application.
- FIG. 10 is a schematic block diagram of another subunit of an Internet information chain storage device provided by an embodiment of this application.
- FIG. 11 is a schematic block diagram of a computer device provided by an embodiment of the present application.
- FIG. 1 is a schematic flowchart of an Internet information chain storage method provided by an embodiment of the present application.
- the Internet information chain storage method is applied to terminal devices with information storage functions, such as desktop computers, notebook computers, tablet computers, or mobile phones.
- the method includes steps S110-S140.
- S110 Obtain the website information of the webpage to be monitored, and perform real-time monitoring on the data information published in the webpage to be monitored according to the webpage information of the webpage to be monitored to obtain newly added data information.
- the webpage information to be monitored is the URL information of the webpage to be monitored input by the user.
- the webpage to be monitored may be all data information published on the Internet such as Weibo, WeChat, corporate website, government website, etc.
- the publisher may be an individual, An enterprise, organization or government department, for example, monitors the information published by a celebrity in its Weibo, then the web page information to be monitored is the URL information of the celebrity's Weibo page.
- the data information published on the web page to be monitored may contain files in various formats, such as text format information, video format information, audio format, and picture format information. Since the data information published in the webpage to be monitored is published in real time, the webpage to be monitored needs to be monitored to obtain the latest data information published in the webpage in real time.
- step S110 includes sub-steps S111, S112, and S113.
- the source information is generated according to the website information of the webpage to be monitored and the publisher of the data information.
- the source information includes the website information of the webpage to be monitored and the publisher of the data information.
- the website information of the webpage to be monitored is also the webpage information to be monitored entered by the user; the publisher is also the publication of the newly added data information.
- the main body, the publisher can be an individual, enterprise, organization or government department.
- the corresponding release time stamp needs to be generated according to the release time of the data information. It is changed, that is to ensure that the release time of the newly added data information is recorded in time and cannot be changed.
- the webpage to be monitored is a celebrity's microblog page
- each microblog information release includes a release time
- the release time of acquiring the microblog information is the release time stamp of the corresponding newly added data information.
- the new data information may include one or more files.
- the webpage to be monitored is a celebrity’s Weibo page
- a text message and a video message are posted on the celebrity’s Weibo page
- the published text message and video information, source information, and timestamp are included.
- New data information for a text file and a video file is included.
- the format information of each file in the newly added data information is obtained to determine whether the file is a text file.
- Each file has its own format information.
- the format information of different files matches the corresponding type, and the specific type of the file can be judged by the format information.
- Whether each file is a text file is judged according to the format information of each file, and the specific type of the file can be judged by the format information of each file.
- the file is a text file
- the format information of a file is wav, mp3, wma
- the file is an audio file
- the format information of a file is avi, flv, rmvb
- the file is a video file.
- non-text files are also a type of file, and non-text files include audio files, video files, and pictures.
- the information recognition model is a model for identifying and converting non-text files. Among them, the information recognition model includes an audio recognition model and a picture recognition model.
- step S130 includes sub-steps S131, S132, and S133.
- the format information of the file and determine whether the file is an audio file. If the file is an audio file, identify the file by the audio recognition model in the information recognition model to obtain the corresponding text file. Through the audio recognition model, the voice information in the audio file can be recognized and converted to obtain a corresponding text file containing text information. After conversion, each audio file corresponds to a text file.
- the audio recognition model includes an acoustic model, a speech feature dictionary and a semantic analysis model.
- step S131 includes sub-steps S1311 and S1312.
- S1311 Divide the voice information in the audio file according to the acoustic model in the audio recognition model to obtain multiple phonemes contained in the voice information.
- the speech information in the audio file is segmented according to the acoustic model in the audio recognition model to obtain multiple phonemes contained in the speech information.
- the voice information is composed of phonemes for pronunciation of multiple characters, and the phonemes of a character include the frequency and timbre of the pronunciation of the character.
- the acoustic model contains phonemes for pronunciation of all characters.
- Matching the obtained phonemes according to the speech feature dictionary in the audio recognition model can convert all phonemes into pinyin information.
- the phonetic feature dictionary contains the phoneme information corresponding to the pinyin of all characters.
- the phoneme of a single character can be converted into the pinyin of the character in the phonetic feature dictionary that matches the phoneme , To convert all phonemes contained in the voice information into pinyin information.
- the obtained Pinyin information is semantically analyzed to convert the Pinyin information into a corresponding text file.
- the semantic analysis model contains the mapping relationship between the pinyin information and the text information. Through the mapping relationship contained in the semantic analysis model, the obtained pinyin information can be semantically analyzed to convert the pinyin information into text containing text information file.
- the text template is the template information used to identify the text in the picture.
- a text template corresponds to a text in the picture.
- a text template contains multiple fonts corresponding to the corresponding text. The text in the picture can be It matches with a certain font in the corresponding text template.
- S133 Obtain the format information of the non-text file and determine whether the file is a video file. If the file is a video file, the file is identified by the audio recognition model and the picture recognition model in the information recognition model to obtain the corresponding text. file.
- the file is a video file. If the file is a video file, the file is identified by the audio recognition model and the picture recognition model in the information recognition model to obtain the corresponding text file. If the file is a video file, first obtain the voice information in the video file, and through the audio recognition model, the voice information in the video file can be recognized and converted to obtain the text information corresponding to the voice information, specific recognition And the conversion method is the same as the step S131; obtain each frame picture included in the video file, and identify each frame picture included in the video file through the picture recognition model to obtain each frame picture The specific identification method of the included text information is the same as that in step S132. Obtain the text information corresponding to the voice information of the video file and the text information contained in each frame of the video file to finally obtain the text file corresponding to the video file, that is, each video file after conversion Each corresponds to a text file.
- the information recognition model cannot process the file, and an alarm prompt message is generated to remind the user that the file cannot be processed.
- the data link table is a database for storing information preset in the terminal device.
- the data link table is a database that stores text files included in the newly added data information according to the time axis, and the data information stored in the data link table
- the logical order of is implemented by the pointer linking order in the data linked list.
- the timestamp of the newly added data information is used as the logical order of the data linked list, that is, the new data is added by using the time information as the pointer linking order.
- the text file in the message is stored in the data link list.
- step S140 includes sub-steps S141, S142, and S143.
- S141 Obtain the release source information and release time stamp of the newly added data information in the web page to be monitored.
- the release source information includes the URL information of the web page to be monitored And the publisher of the data.
- S142 Classify the text files included in the newly added data information and/or the converted text files according to the source information.
- the text files included in the newly added data information and/or the converted text files are classified according to the source information. Specifically, in order to classify and store the newly added data information, it is necessary to classify the newly added data information according to the publisher in the release source information, where each publisher corresponds to a category, and each category is associated with a child in the data list Corresponding to the linked list, the new data information released by the same publisher is divided into the sub-linked list corresponding to the release source information to be saved, and the text contained in the new data information can be saved by the publisher in the release source information Files and/or converted text files are classified to classify and save the newly added data information.
- S143 Save the text file included in the newly added data information and/or the converted text file according to the release timestamp to the sub-link list corresponding to the release source information in the preset data link list.
- a publisher corresponds to a category, each category corresponds to a sub-linked list in the data linked list, and because the data information in the data linked list is stored according to the time axis, it is necessary to add the new data information according to the release timestamp
- the text file corresponding to the newly added data information is stored in the sub-linked list corresponding to the category of the publisher in the data linked list, and the newly added data information can be saved.
- the historical data information published on the web page to be monitored can be saved to facilitate later forensics on the historical data information published by the corresponding publisher.
- the non-text file file is converted to a text file, and all text files are stored in the data link list to realize the chaining of Internet information Storage can ensure that the stored text files cannot be deleted and modified, and can facilitate users to obtain deleted data information on the Internet to assist users in forensics of related data information, which has great practical value.
- An embodiment of the present application further provides an Internet information chain storage device, which is used to execute any embodiment of the foregoing Internet information chain storage method.
- FIG. 6, is a schematic block diagram of an Internet information chain storage device provided by an embodiment of the present application.
- the Internet information chain storage device can be configured in a terminal device such as a desktop computer, a notebook computer, a tablet computer or a mobile phone.
- the Internet information chain storage device 100 includes a webpage monitoring unit 110, a judgment unit 120, an information conversion unit 130, and an information storage unit 140.
- the webpage monitoring unit 110 is configured to obtain the webpage information of the webpage to be monitored, and perform real-time monitoring on the data information published in the webpage to be monitored according to the webpage information of the webpage to be monitored to obtain newly added data information.
- the webpage monitoring unit 110 includes subunits: a publishing source information generating unit 111, a publishing timestamp generating unit 112, and a newly added data information acquiring unit 113.
- the publishing source information generating unit 111 is configured to generate publishing source information according to the website information of the web page to be monitored and the publisher of the data information if the web information to be monitored is published.
- the publishing time stamp generating unit 112 is configured to generate a publishing time stamp according to the publishing time of the data information.
- the newly added data information obtaining unit 113 is configured to obtain all the files in the published data information, the published source information, and the published timestamp to obtain the newly added data information.
- the judging unit 120 is used to judge whether the file in the newly added data information is a text file.
- the information conversion unit 130 is configured to convert the non-text file into a text file through a preset information recognition model if the file in the newly added data information is a non-text file.
- the information conversion unit 130 includes subunits: a first text file acquisition unit 131, a second text file acquisition unit 132, and a third text file acquisition unit 133.
- the first text file obtaining unit 131 is used to obtain the format information of the non-text file and determine whether the file is an audio file. If the file is an audio file, the file is recognized by the audio recognition model in the information recognition model to obtain The corresponding text file.
- the first text file acquisition unit 131 includes subunits: a phoneme segmentation unit 1311, a phoneme conversion unit 1312, and a speech analysis unit 1313.
- the phoneme segmentation unit 1311 is configured to segment the voice information in the audio file according to the acoustic model in the audio recognition model to obtain multiple phonemes contained in the voice information.
- the phoneme conversion unit 1312 is configured to match the obtained phonemes according to the speech feature dictionary in the audio recognition model to convert all phonemes into pinyin information.
- the speech analysis unit 1313 is configured to perform semantic analysis on the obtained pinyin information according to the semantic analysis model in the audio recognition model to obtain a text file containing text information.
- the second text file obtaining unit 132 is used to obtain the format information of the non-text file and determine whether the file is a picture. If the file is a picture, the file is recognized by the picture recognition model in the information recognition model to obtain the corresponding Text file.
- the third text file obtaining unit 133 is used to obtain the format information of the non-text file and determine whether the file is a video file. If the file is a video file, the audio recognition model and the picture recognition model in the information recognition model The file is identified to obtain the corresponding text file.
- the information storage unit 140 is configured to save the text file in the newly added data information and/or the converted text file to a preset data list.
- the information storage unit 140 includes subunits: an information acquisition unit 141, a file classification unit 142, and a file storage unit 143.
- the information obtaining unit 141 is used to obtain the release source information and the release time stamp of the newly added data information in the web page to be monitored.
- the file classification unit 142 is used to classify the text files included in the newly added data information and/or the converted text files according to the distribution source information.
- the file storage unit 143 is configured to save the text file included in the newly added data information and/or the converted text file to the sub-link list corresponding to the release source information in the preset data link list according to the release time stamp.
- the above-mentioned Internet information chain storage device may be implemented in the form of a computer program, and the computer program may run on a computer device as shown in FIG. 11.
- FIG. 11 is a schematic block diagram of a computer device provided by an embodiment of the present application.
- the computer device 500 includes a processor 502, a memory, and a network interface 505 connected through a system bus 501, where the memory may include a non-volatile storage medium 503 and an internal memory 504.
- the non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032.
- the processor 502 can execute the Internet information chain storage method.
- the processor 502 is used to provide computing and control capabilities and support the operation of the entire computer device 500.
- the internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503.
- the processor 502 can cause the processor 502 to execute the Internet information chain storage method.
- the network interface 505 is used for network communication, such as the transmission of data information.
- the network interface 505 is used for network communication, such as the transmission of data information.
- FIG. 11 is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied.
- the specific computer device 500 may include more or less components than shown in the figure, or combine certain components, or have a different arrangement of components.
- the processor 502 is used to run the computer program 5032 stored in the memory to implement the Internet information chain storage method of this embodiment.
- the embodiment of the computer device shown in FIG. 11 does not constitute a limitation on the specific configuration of the computer device.
- the computer device may include more or fewer components than shown in the figure. Or combine certain components, or arrange different components.
- the computer device may only include a memory and a processor. In such an embodiment, the structures and functions of the memory and the processor are consistent with the embodiment shown in FIG. 11, and details are not described herein again.
- the processor 502 may be a central processing unit (Central Processing Unit, CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), Application specific integrated circuit (Application Specific Integrated Circuit, ASIC), ready-made programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
- the general-purpose processor may be a microprocessor or the processor may be any conventional processor.
- a computer-readable storage medium may be a non-volatile computer-readable storage medium.
- the computer-readable storage medium stores a computer program, where the computer program is executed by the processor to implement the Internet information chain storage method of the embodiments of the present application.
- the storage medium may be an internal storage unit of the foregoing device, such as a hard disk or a memory of the device.
- the storage medium may also be an external storage device of the device, such as a plug-in hard disk equipped on the device, a smart memory card (Smart) Card (SMC), a secure digital (SD) card, or a flash memory card (Flash Card) etc.
- the storage medium may also include both an internal storage unit of the device and an external storage device.
Landscapes
- Engineering & Computer Science (AREA)
- Databases & Information Systems (AREA)
- Theoretical Computer Science (AREA)
- Data Mining & Analysis (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Information Transfer Between Computers (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
本申请公开了互联网信息链式存储方法、装置、计算机设备及存储介质。方法包括:获取待监控网页的网址信息,根据待监控网页的网址信息对待监控网页中所发布的数据信息进行实时监控以获取新增数据信息;对新增数据信息中的文件是否为文字文件进行判断;若新增数据信息中的文件为非文字文件,通过预设信息识别模型将非文字文件转换为文字文件;将新增数据信息中的文字文件和/或转换得到的文字文件保存至预设数据链表中。
Description
本申请要求于2018年12月13日提交中国专利局、申请号为201811526834.0、申请名称为“互联网信息链式存储方法、装置、计算机设备及存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本申请涉及计算机技术领域,尤其涉及一种互联网信息链式存储方法、装置、计算机设备及存储介质。
互联网中各网页上保存有海量的数据信息,且新增数据信息会逐渐更替网页中已保存的数据信息,造成网页中的数据信息发生更迭变化的情况,因而现有对互联网中的数据信息进行存储方法无法对互联网上已删除或已修改的数据信息进行获取,在司法实践中对互联网上所发布的相关数据信息进行取证存在极大的困难。因此,现有的数据信息存储方法无法获取已删除数据信息。
发明内容
本申请实施例提供了一种互联网信息链式存储方法、装置、计算机设备及存储介质,旨在解决现有技术中数据信息存储方法无法获取已删除数据信息的问题。
第一方面,本申请实施例提供了一种互联网信息链式存储方法,其包括:
获取待监控网页的网址信息,根据待监控网页的网址信息对待监控网页中所发布的数据信息进行实时监控以获取新增数据信息;对新增数据信息中的文件是否为文字文件进行判断;若新增数据信息中的文件为非文字文件,通过预设信息识别模型将非文字文件转换为文字文件;将新增数据信息中的文字文件和/或转换得到的文字文件保存至预设数据链表中。
第二方面,本申请实施例提供了一种互联网信息链式存储装置,其包括:
网页监控单元,用于获取待监控网页的网址信息,根据待监控网页的网址 信息对待监控网页中所发布的数据信息进行实时监控以获取新增数据信息;判断单元,用于对新增数据信息中的文件是否为文字文件进行判断;信息转换单元,用于若新增数据信息中的文件为非文字文件,通过预设信息识别模型将非文字文件转换为文字文件;信息存储单元,用于将新增数据信息中的文字文件和/或转换得到的文字文件保存至预设数据链表中。
第三方面,本申请实施例又提供了一种计算机设备,其包括存储器、处理器及存储在所述存储器上并可在所述处理器上运行的计算机程序,所述处理器执行所述计算机程序时实现上述第一方面所述的互联网信息链式存储方法。
第四方面,本申请实施例还提供了一种计算机可读存储介质,其中所述计算机可读存储介质存储有计算机程序,所述计算机程序当被处理器执行时使所述处理器执行上述第一方面所述的互联网信息链式存储方法。
为了更清楚地说明本申请实施例技术方案,下面将对实施例描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图是本申请的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1为本申请实施例提供的互联网信息链式存储方法的流程示意图;
图2为本申请实施例提供的互联网信息链式存储方法的子流程示意图;
图3为本申请实施例提供的互联网信息链式存储方法的另一子流程示意图;
图4为本申请实施例提供的互联网信息链式存储方法的另一子流程示意图;
图5为本申请实施例提供的互联网信息链式存储方法的另一子流程示意图;
图6为本申请实施例提供的互联网信息链式存储装置的示意性框图;
图7为本申请实施例提供的互联网信息链式存储装置的子单元示意性框图;
图8为本申请实施例提供的互联网信息链式存储装置的另一子单元示意性框图;
图9为本申请实施例提供的互联网信息链式存储装置的另一子单元示意性框图;
图10为本申请实施例提供的互联网信息链式存储装置的另一子单元示意性框图;
图11为本申请实施例提供的计算机设备的示意性框图。
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例是本申请一部分实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本申请保护的范围。
应当理解,当在本说明书和所附权利要求书中使用时,术语“包括”和“包含”指示所描述特征、整体、步骤、操作、元素和/或组件的存在,但并不排除一个或多个其它特征、整体、步骤、操作、元素、组件和/或其集合的存在或添加。
还应当理解,在此本申请说明书中所使用的术语仅仅是出于描述特定实施例的目的而并不意在限制本申请。如在本申请说明书和所附权利要求书中所使用的那样,除非上下文清楚地指明其它情况,否则单数形式的“一”、“一个”及“该”意在包括复数形式。
还应当进一步理解,在本申请说明书和所附权利要求书中使用的术语“和/或”是指相关联列出的项中的一个或多个的任何组合以及所有可能组合,并且包括这些组合。
请参阅图1,图1是本申请实施例提供的互联网信息链式存储方法的流程示意图。该互联网信息链式存储方法应用于具有信息存储功能的终端设备中,例如台式电脑、笔记本电脑、平板电脑或手机等。
如图1所示,该方法包括步骤S110~S140。
S110、获取待监控网页的网址信息,根据待监控网页的网址信息对待监控网页中所发布的数据信息进行实时监控以获取新增数据信息。
获取待监控网页的网址信息,根据待监控网页的网址信息对待监控网页中所发布的数据信息进行实时监控以获取新增数据信息。其中,待监控网页信息为用户所输入的待监控网页的网址信息,待监控网页可以是微博、微信、企业网址、政府网站等所有在互联网上所发布的数据信息,发布人可以是个人、企业、组织或政府部门,例如对某一名人在其微博中所发布的信息进行监控,则待监控网页信息即是该名人微博网页的网址信息。
待监控网页中所发布的数据信息中可包含多种格式的文件,例如文字格式的信息、视频格式的信息、音频格式、图片格式的信息等。由于待监控网页中所发布的数据信息为实时发布,因此需对待监控网页进行监控以实时获取该网页中所发布的最新的数据信息。
在一实施例中,如图2所示,步骤S110包括子步骤S111、S112和S113。
S111、若监控到待监控网页中发布数据信息,根据待监控网页的网址信息及所述数据信息的发布人生成发布源信息。
若监控到待监控网页中发布数据信息,根据待监控网页的网址信息及所述数据信息的发布人生成发布源信息。为获取新增数据信息的发布人,需根据待监控网页的网址信息及所述数据信息的发布人生成相应的发布源信息。发布源信息中包括待监控网页的网址信息以及该数据信息的发布人,待监控网页的网址信息也即是用户所输入的待监控网页信息;发布人也即是发布该新增数据信息的发布主体,发布人可以是个人、企业、组织或政府部门。
S112、根据所述数据信息的发布时间生成发布时间戳。
根据所述数据信息的发布时间生成新增数据信息的发布时间戳,为对新增数据信息的发布时间进行记录,需根据数据信息的发布时间生成相应的发布时间戳,发布时间戳生成后无法被更改,也即是确保新增数据信息的发布时间被及时记录且无法更改。
例如,待监控网页为某一名人的微博网页,每一条微博信息的发布均包含一个发布时间,获取该微博信息的发布时间即为相应新增数据信息的发布时间戳。
S113、获取所述发布数据信息中的所有文件及发布源信息、发布时间戳以得到新增数据信息。
获取所述发布数据信息中的所有文件及发布源信息、发布时间戳以得到新增数据信息。获取所发布数据信息中的所有文件作为新增数据信息,并获取所得到的发布源信息及发布时间戳即可得到新增数据信息,新增数据信息中可包含一个或多个文件。
例如,待监控网页为某一名人的微博网页,该名人微博网页中发布了一段文字信息及一个视频信息,则获取所发布的文字信息及视频信息、发布源信息、发布时间戳得到包含一个文字文件及一个视频文件的新增数据信息。
S120、对新增数据信息中的文件是否为文字文件进行判断。
对新增数据信息中的文件是否为文字文件进行判断,为对新增数据信息中各种格式的文件进行保存,需先对新增数据信息中的文件是否为文字文件进行判断。具体的,通过获取新增数据信息中各文件的格式信息以判断该文件是否为文字文件。
获取新增数据信息中各文件的格式信息。每一个文件都拥有各自的格式信息,不同文件的格式信息与相应的类型相匹配,通过格式信息即可对文件的具体类型进行判断。根据各文件的格式信息对各文件是否为文字文件进行判断,通过各文件的格式信息即可对文件的具体类型进行判断。
例如,若某一文件的格式信息为txt、string,则该文件为文字文件;若某一文件的格式信息为wav、mp3、wma,则该文件为音频文件;若某一文件的格式信息为avi、flv、rmvb,则该文件为视频文件。
S130、若新增数据信息中的文件为非文字文件,通过信息识别模型将非文字文件转换为文字文件。
若新增数据信息中的文件为非文字文件,通过预设信息识别模型将非文字文件转换为文字文件。具体的,非文字文件也是文件中的一种,非文字文件包括音频文件、视频文件、图片。信息识别模型即是用于对非文字文件进行识别及转换的模型,其中,信息识别模型中包括音频识别模型及图片识别模型。
在一实施例中,如图3所示,步骤S130包括子步骤S131、S132和S133。
S131、获取所述非文字文件的格式信息并判断该文件是否为音频文件,若该文件为音频文件则通过信息识别模型中的音频识别模型对该文件进行识别以得到相应的文字文件。
获取文件的格式信息并判断该文件是否为音频文件,若该文件为音频文件则通过信息识别模型中的音频识别模型对该文件进行识别以得到相应的文字文件。通过音频识别模型即可对音频文件中的语音信息进行识别并转换,以得到相应的包含文字信息的文字文件,进行转换后每一个音频文件均对应得到一个文字文件。其中,音频识别模型包括声学模型、语音特征词典及语义解析模型。
在一实施例中,如图4所示,步骤S131包括子步骤S1311和S1312。
S1311、根据音频识别模型中的声学模型对音频文件中的语音信息进行切分以得到语音信息中所包含的多个音素。
根据音频识别模型中的声学模型对音频文件中的语音信息进行切分以得到语音信息中所包含的多个音素。具体的,语音信息由多个字符发音的音素而组成,一个字符的音素包括该字符发音的频率和音色。声学模型中包含所有字符发音的音素,通过将语音信息与声学模型中所有的音素进行匹配,即可对语音信息中单个字符的音素进行切分,通过切分最终得到该语音信息中所包含的多个音素。
S1312、根据音频识别模型中的语音特征词典对所得到的音素进行匹配以将所有音素转换为拼音信息。
根据音频识别模型中的语音特征词典对所得到的音素进行匹配,即可将所有音素转换为拼音信息。语音特征词典中包含所有字符拼音对应的音素信息,通过将所得到的音素与字符拼音对应的音素信息进行匹配,即可将单个字符的音素转换为语音特征词典中与该音素相匹配的字符拼音,以实现将语音信息中所包含的所有音素转换为拼音信息。
S1313、根据音频识别模型中的语义解析模型对所得到的拼音信息进行语义解析以得到包含文字信息的文字文件。
根据音频识别模型中的语义解析模型对所得到的拼音信息进行语义解析,以实现将拼音信息转换为对应的文字文件。语义解析模型中包含拼音信息与文字信息之间所对应的映射关系,通过语义解析模型中所包含的映射关系即可对所得到的拼音信息进行语义解析以将拼音信息转换为包含文字信息的文字文件。
S132、获取所述非文字文件的格式信息并判断该文件是否为图片,若该文件为图片则通过信息识别模型中图片识别模型的对该文件进行识别以得到相应的文字文件。
获取所述非文字文件的格式信息并判断该文件是否为图片,若该文件为图片则通过信息识别模型中图片识别模型的对该文件中所包含的文字进行识别以得到相应的文字文件。具体的,文字模板即是用于对图片中文字进行识别的模板信息,一个文字模板与图片中的一个文字相对应,一个文字模板包含相应文字所对应的多种字体,图片中的文字均能与对应文字模板中的某一种字体相匹配,通过图片识别模型中的文字模板与图片进行匹配,即可对该图片中所包含的文字进行识别以得到相应的文字文件。
S133、获取所述非文字文件的格式信息并判断该文件是否为视频文件,若该文件为视频文件则通过信息识别模型中的音频识别模型及图片识别模型对该文件进行识别以得到相应的文字文件。
获取所述非文字文件的格式信息并判断该文件是否为视频文件,若该文件为视频文件则通过信息识别模型中的音频识别模型及图片识别模型对该文件进行识别以得到相应的文字文件。若该文件为视频文件,则先获取该视频文件中的语音信息,并通过音频识别模型即可对该视频文件中的语音信息进行识别并转换以得到该语音信息对应的文字信息,具体的识别及转换方法与所述步骤S131相同;获取该视频文件中所包含的每一帧图片,并通过图片识别模型对该视频文件中所包含的每一帧图片进行识别,以得到每一帧图片中所包含的文字信息,具体的识别方法与所述步骤S132相同。获取视频文件的语音信息所对应的文字信息,及该视频文件中每一帧图片所包含的文字信息,即可最终得到该视频文件所对应的文字文件,也即是进行转换后每一个视频文件均对应得到一个文字文件。
此外,若非文字文件不是视频文件、音频文件及图片中的任意一种,信息识别模型无法对该文件进行处理,则生成报警提示信息以提示用户无法对该文件进行处理。
S140、将新增数据信息中的文字文件和/或转换得到的文字文件保存至预设数据链表中。
将新增数据信息中的文字文件和/或转换得到的文字文件保存至预设数据链表中以对新增数据信息进行保存。数据链表为终端设备中所预设的用于存储信息的数据库,具体的,数据链表为根据时间轴对新增数据信息中所包含的文字文件进行存储的数据库,数据链表中所存储的数据信息的逻辑顺序是通过数据链表中的指针链接次序实现的,在本实施例中以新增数据信息的发布时间戳作为数据链表的逻辑顺序,也即是通过时间信息为指针链接次序将新增数据信息中的文字文件存储至数据链表中。通过以时间顺序作为链表的逻辑顺序对新增数据信息进行存储,用户可通过数据链表获取到以时间信息为顺序的文字文件列表,数据链表所存储的信息具有无法删除的特性。
此外,由于将其他非文字文件转换为包含文字信息的文字文件进行存储,因此可极大地压缩对相应数据信息进行存储所需的存储空间,便于用户进行使用。
在一实施例中,如图5所示,步骤S140包括子步骤S141、S142和S143。
S141、获取待监控网页中新增数据信息的发布源信息及发布时间戳。
获取待监控网页中新增数据信息的发布源信息及发布时间戳。为方便对新增数据信息进行存储以后期对所存储的信息数据信息进行检索,需获取新增数据信息的发布源信息及发布时间戳,具体的,发布源信息中包括待监控网页的网址信息以及该数据信息的发布人。
S142、根据发布源信息将新增数据信息中所包含的文字文件和/或转换得到的文字文件进行分类。
根据发布源信息将新增数据信息中所包含的文字文件和/或转换得到的文字文件进行分类。具体的,为实现对新增数据信息进行分类存储,需根据发布源信息中的发布人对新增数据信息进行分类,其中,一个发布人对应一个类别,每一个类别与数据链表中的一个子链表相对应,相同发布人所发布的新增数据信息则分至该发布源信息对应的子链表中进行保存,通过发布源信息中的发布人,即可将新增数据信息中所包含的文字文件和/或转换得到的文字文件进行分类以对新增数据信息进行分类保存。
S143、根据发布时间戳将新增数据信息中所包含的文字文件和/或转换得到的文字文件保存至预设数据链表中与发布源信息对应的子链表。
根据发布时间戳将新增数据信息中所包含的文字文件和/或转换得到的文字文件保存至预设数据链表中相应的子链表中进行保存。一个发布人对应一个类别,每一个类别与数据链表中的一个子链表相对应,且由于数据链表中的数据信息均即是根据时间轴进行存储,因此需根据新增数据信息的发布时间戳将新增数据信息中所对应的文字文件存储至数据链表中与发布人类别相对应的子链表中,即可实现对新增数据信息进行保存。
由于数据链表所存储的文字文件无法删除和修改,因此可实现对待监控网页中所发布的历史数据信息进行保存,以方便后期进行对相应发布人所发布的历史数据信息进行取证。
通过对网页中所发布的数据信息进行监控并判断其中的文件是否为文字文 件,将非文字文件的文件转换为文字文件,并对所有文字文件存储至数据链表中以实现对互联网信息进行链式存储,能够确保所存储的文字文件无法删除和修改,能够方便用户获取互联网上已删除的数据信息以协助用户对相关数据信息进行取证,具有极大的实用价值。
本申请实施例还提供一种互联网信息链式存储装置,该互联网信息链式存储装置用于执行前述互联网信息链式存储方法的任一实施例。具体地,请参阅图6,图6是本申请实施例提供的互联网信息链式存储装置的示意性框图。该互联网信息链式存储装置可以配置于台式电脑、笔记本电脑、平板电脑或手机等终端设备中。
如图6所示,互联网信息链式存储装置100包括网页监控单元110、判断单元120、信息转换单元130、信息存储单元140。
网页监控单元110,用于获取待监控网页的网址信息,根据待监控网页的网址信息对待监控网页中所发布的数据信息进行实时监控以获取新增数据信息。
其他申请实施例中,如图7所示,所述网页监控单元110包括子单元:发布源信息生成单元111、发布时间戳生成单元112和新增数据信息获取单元113。
发布源信息生成单元111,用于若监控到待监控网页中发布数据信息,根据待监控网页的网址信息及所述数据信息的发布人生成发布源信息。
发布时间戳生成单元112,用于根据所述数据信息的发布时间生成发布时间戳。
新增数据信息获取单元113,用于获取所述发布数据信息中的所有文件及发布源信息、发布时间戳以得到新增数据信息。
判断单元120,用于对新增数据信息中的文件是否为文字文件进行判断。
信息转换单元130,用于若新增数据信息中的文件为非文字文件,通过预设信息识别模型将非文字文件转换为文字文件。
其他申请实施例中,如图8所示,所述信息转换单元130包括子单元:第一文字文件获取单元131、第二文字文件获取单元132和第三文字文件获取单元133。
第一文字文件获取单元131,用于获取所述非文字文件的格式信息并判断该文件是否为音频文件,若该文件为音频文件则通过信息识别模型中的音频识别模型对该文件进行识别以得到相应的文字文件。
其他申请实施例中,如图9所示,所述第一文字文件获取单元131包括子单元:音素切分单元1311、音素转换单元1312和语音解析单元1313。
音素切分单元1311,用于根据音频识别模型中的声学模型对音频文件中的语音信息进行切分以得到语音信息中所包含的多个音素。
音素转换单元1312,用于根据音频识别模型中的语音特征词典对所得到的音素进行匹配以将所有音素转换为拼音信息。
语音解析单元1313,用于根据音频识别模型中的语义解析模型对所得到的拼音信息进行语义解析以得到包含文字信息的文字文件。
第二文字文件获取单元132,用于获取所述非文字文件的格式信息并判断该文件是否为图片,若该文件为图片则通过信息识别模型中图片识别模型的对该文件进行识别以得到相应的文字文件。
第三文字文件获取单元133,用于获取所述非文字文件的格式信息并判断该文件是否为视频文件,若该文件为视频文件则通过信息识别模型中的音频识别模型及图片识别模型对该文件进行识别以得到相应的文字文件。
信息存储单元140,用于将新增数据信息中的文字文件和/或转换得到的文字文件保存至预设数据链表中。
其他申请实施例中,如图10所示,所述信息存储单元140包括子单元:信息获取单元141、文件分类单元142和文件存储单元143。
信息获取单元141,用于获取待监控网页中新增数据信息的发布源信息及发布时间戳。
文件分类单元142,用于根据发布源信息将新增数据信息中所包含的文字文件和/或转换得到的文字文件进行分类。
文件存储单元143,用于根据发布时间戳将新增数据信息中所包含的文字文件和/或转换得到的文字文件保存至预设数据链表中与发布源信息对应的子链表。
上述互联网信息链式存储装置可以实现为计算机程序的形式,该计算机程序可以在如图11所示的计算机设备上运行。
请参阅图11,图11是本申请实施例提供的计算机设备的示意性框图。
参阅图11,该计算机设备500包括通过系统总线501连接的处理器502、存储器和网络接口505,其中,存储器可以包括非易失性存储介质503和内存储器504。该非易失性存储介质503可存储操作系统5031和计算机程序5032。该 计算机程序5032被执行时,可使得处理器502执行互联网信息链式存储方法。该处理器502用于提供计算和控制能力,支撑整个计算机设备500的运行。该内存储器504为非易失性存储介质503中的计算机程序5032的运行提供环境,该计算机程序5032被处理器502执行时,可使得处理器502执行互联网信息链式存储方法。该网络接口505用于进行网络通信,如提供数据信息的传输等。本领域技术人员可以理解,图11中示出的结构,仅仅是与本申请方案相关的部分结构的框图,并不构成对本申请方案所应用于其上的计算机设备500的限定,具体的计算机设备500可以包括比图中所示更多或更少的部件,或者组合某些部件,或者具有不同的部件布置。
其中,所述处理器502用于运行存储在存储器中的计算机程序5032,以实现本实施例的互联网信息链式存储方法。
本领域技术人员可以理解,图11中示出的计算机设备的实施例并不构成对计算机设备具体构成的限定,在其他实施例中,计算机设备可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件布置。例如,在一些实施例中,计算机设备可以仅包括存储器及处理器,在这样的实施例中,存储器及处理器的结构及功能与图11所示实施例一致,在此不再赘述。
应当理解,在本申请实施例中,处理器502可以是中央处理单元(Central Processing Unit,CPU),该处理器502还可以是其他通用处理器、数字信号处理器(Digital Signal Processor,DSP)、专用集成电路(Application Specific Integrated Circuit,ASIC)、现成可编程门阵列(Field-Programmable Gate Array,FPGA)或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件等。其中,通用处理器可以是微处理器或者该处理器也可以是任何常规的处理器等。
在本申请的另一实施例中提供计算机可读存储介质。该计算机可读存储介质可以为非易失性的计算机可读存储介质。该计算机可读存储介质存储有计算机程序,其中计算机程序被处理器执行时实现本申请实施例的互联网信息链式存储方法。
所述存储介质可以是前述设备的内部存储单元,例如设备的硬盘或内存。所述存储介质也可以是所述设备的外部存储设备,例如所述设备上配备的插接式硬盘,智能存储卡(Smart Media Card,SMC),安全数字(Secure Digital,SD)卡,闪存卡(Flash Card)等。进一步地,所述存储介质还可以既包括所述设备 的内部存储单元也包括外部存储设备。
所属领域的技术人员可以清楚地了解到,为了描述的方便和简洁,上述描述的设备、装置和单元的具体工作过程,可以参考前述方法实施例中的对应过程,在此不再赘述。
以上所述,仅为本申请的具体实施方式,但本申请的保护范围并不局限于此,任何熟悉本技术领域的技术人员在本申请揭露的技术范围内,可轻易想到各种等效的修改或替换,这些修改或替换都应涵盖在本申请的保护范围之内。因此,本申请的保护范围应以权利要求的保护范围为准。
Claims (20)
- 一种互联网信息链式存储方法,包括:获取待监控网页的网址信息,根据待监控网页的网址信息对待监控网页中所发布的数据信息进行实时监控以获取新增数据信息;对新增数据信息中的文件是否为文字文件进行判断;若新增数据信息中的文件为非文字文件,通过预设信息识别模型将非文字文件转换为文字文件;将新增数据信息中的文字文件和/或转换得到的文字文件保存至预设数据链表中。
- 根据权利要求1所述的互联网信息链式存储方法,其中,所述根据待监控网页的网址信息对待监控网页中所发布的数据信息进行实时监控以获取新增数据信息,包括:若监控到待监控网页中发布数据信息,根据待监控网页的网址信息及所述数据信息的发布人生成发布源信息;根据所述数据信息的发布时间生成发布时间戳;获取所述发布数据信息中的所有文件及发布源信息、发布时间戳以得到新增数据信息。
- 根据权利要求1所述的互联网信息链式存储方法,其中,所述通过预设信息识别模型将非文字文件转换为文字文件,包括:获取所述非文字文件的格式信息并判断该文件是否为音频文件,若该文件为音频文件则通过信息识别模型中的音频识别模型对该文件进行识别以得到相应的文字文件;获取所述非文字文件的格式信息并判断该文件是否为图片,若该文件为图片则通过信息识别模型中图片识别模型的对该文件进行识别以得到相应的文字文件;获取所述非文字文件的格式信息并判断该文件是否为视频文件,若该文件为视频文件则通过信息识别模型中的音频识别模型及图片识别模型对该文件进行识别以得到相应的文字文件。
- 根据权利要求3所述的互联网信息链式存储方法,其中,所述若该文件 为音频文件则通过信息识别模型中的音频识别模型对该文件进行识别以得到相应的文字文件,包括:根据音频识别模型中的声学模型对音频文件中的语音信息进行切分以得到语音信息中所包含的多个音素;根据音频识别模型中的语音特征词典对所得到的音素进行匹配以将所有音素转换为拼音信息;根据音频识别模型中的语义解析模型对所得到的拼音信息进行语义解析以得到包含文字信息的文字文件。
- 根据权利要求2所述的互联网信息链式存储方法,其中,所述将新增数据信息中的文字文件和/或转换得到的文字文件保存至预设数据链表中,包括:获取待监控网页中新增数据信息的发布源信息及发布时间戳;根据发布源信息将新增数据信息中所包含的文字文件和/或转换得到的文字文件进行分类;根据发布时间戳将新增数据信息中所包含的文字文件和/或转换得到的文字文件保存至预设数据链表中与发布源信息对应的子链表。
- 根据权利要求1所述的互联网信息链式存储方法,其中,所述对新增数据信息中的文件是否为文字文件进行判断,包括:根据新增数据信息中各文件的格式信息对各文件是否为文字文件进行判断。
- 根据权利要求3所述的互联网信息链式存储方法,其中,所述通过预设信息识别模型将非文字文件转换为文字文件,还包括:若所述非文字文件不是视频文件、音频文件及图片中的任意一种,生成报警提示信息以提示用户无法对所述文件进行处理。
- 一种互联网信息链式存储装置,包括:网页监控单元,用于获取待监控网页的网址信息,根据待监控网页的网址信息对待监控网页中所发布的数据信息进行实时监控以获取新增数据信息;判断单元,用于对新增数据信息中的文件是否为文字文件进行判断;信息转换单元,用于若新增数据信息中的文件为非文字文件,通过预设信息识别模型将非文字文件转换为文字文件;信息存储单元,用于将新增数据信息中的文字文件和/或转换得到的文字文件保存至预设数据链表中。
- 根据权利要求8所述的互联网信息链式存储装置,其中,所述网页监控单元,包括:发布源信息生成单元,用于若监控到待监控网页中发布数据信息,根据待监控网页的网址信息及所述数据信息的发布人生成发布源信息;发布时间戳生成单元,用于根据所述数据信息的发布时间生成发布时间戳;新增数据信息获取单元,用于获取所述发布数据信息中的所有文件及发布源信息、发布时间戳以得到新增数据信息。
- 根据权利要求8所述的互联网信息链式存储装置,其中,所述信息转换单元,包括:第一文字文件获取单元,用于获取所述非文字文件的格式信息并判断该文件是否为音频文件,若该文件为音频文件则通过信息识别模型中的音频识别模型对该文件进行识别以得到相应的文字文件;第二文字文件获取单元,用于获取所述非文字文件的格式信息并判断该文件是否为图片,若该文件为图片则通过信息识别模型中图片识别模型的对该文件进行识别以得到相应的文字文件;第三文字文件获取单元,用于获取所述非文字文件的格式信息并判断该文件是否为视频文件,若该文件为视频文件则通过信息识别模型中的音频识别模型及图片识别模型对该文件进行识别以得到相应的文字文件。
- 一种计算机设备,包括存储器、处理器及存储在所述存储器上并可在所述处理器上运行的计算机程序,其中,所述处理器执行所述计算机程序时实现以下步骤:获取待监控网页的网址信息,根据待监控网页的网址信息对待监控网页中所发布的数据信息进行实时监控以获取新增数据信息;对新增数据信息中的文件是否为文字文件进行判断;若新增数据信息中的文件为非文字文件,通过预设信息识别模型将非文字文件转换为文字文件;将新增数据信息中的文字文件和/或转换得到的文字文件保存至预设数据链表中。
- 根据权利要求11所述的计算机设备,其中,所述根据待监控网页的网址信息对待监控网页中所发布的数据信息进行实时监控以获取新增数据信息, 包括:若监控到待监控网页中发布数据信息,根据待监控网页的网址信息及所述数据信息的发布人生成发布源信息;根据所述数据信息的发布时间生成发布时间戳;获取所述发布数据信息中的所有文件及发布源信息、发布时间戳以得到新增数据信息。
- 根据权利要求11所述的计算机设备,其中,所述通过预设信息识别模型将非文字文件转换为文字文件,包括:获取所述非文字文件的格式信息并判断该文件是否为音频文件,若该文件为音频文件则通过信息识别模型中的音频识别模型对该文件进行识别以得到相应的文字文件;获取所述非文字文件的格式信息并判断该文件是否为图片,若该文件为图片则通过信息识别模型中图片识别模型的对该文件进行识别以得到相应的文字文件;获取所述非文字文件的格式信息并判断该文件是否为视频文件,若该文件为视频文件则通过信息识别模型中的音频识别模型及图片识别模型对该文件进行识别以得到相应的文字文件。
- 根据权利要求13所述的计算机设备,其中,所述若该文件为音频文件则通过信息识别模型中的音频识别模型对该文件进行识别以得到相应的文字文件,包括:根据音频识别模型中的声学模型对音频文件中的语音信息进行切分以得到语音信息中所包含的多个音素;根据音频识别模型中的语音特征词典对所得到的音素进行匹配以将所有音素转换为拼音信息;根据音频识别模型中的语义解析模型对所得到的拼音信息进行语义解析以得到包含文字信息的文字文件。
- 根据权利要求12所述的计算机设备,其中,所述将新增数据信息中的文字文件和/或转换得到的文字文件保存至预设数据链表中,包括:获取待监控网页中新增数据信息的发布源信息及发布时间戳;根据发布源信息将新增数据信息中所包含的文字文件和/或转换得到的文字 文件进行分类;根据发布时间戳将新增数据信息中所包含的文字文件和/或转换得到的文字文件保存至预设数据链表中与发布源信息对应的子链表。
- 根据权利要求11所述的计算机设备,其中,所述对新增数据信息中的文件是否为文字文件进行判断,包括:根据各文件的格式信息对各文件是否为文字文件进行判断。
- 根据权利要求13所述的计算机设备,其中,所述通过预设信息识别模型将非文字文件转换为文字文件,还包括:若所述非文字文件不是视频文件、音频文件及图片中的任意一种,生成报警提示信息以提示用户无法对所述文件进行处理。
- 一种计算机可读存储介质,其中,所述计算机可读存储介质存储有计算机程序,所述计算机程序当被处理器执行时使所述处理器执行以下操作:获取待监控网页的网址信息,根据待监控网页的网址信息对待监控网页中所发布的数据信息进行实时监控以获取新增数据信息;对新增数据信息中的文件是否为文字文件进行判断;若新增数据信息中的文件为非文字文件,通过预设信息识别模型将非文字文件转换为文字文件;将新增数据信息中的文字文件和/或转换得到的文字文件保存至预设数据链表中。
- 根据权利要求18所述的计算机可读存储介质,其中,所述根据待监控网页的网址信息对待监控网页中所发布的数据信息进行实时监控以获取新增数据信息,包括:若监控到待监控网页中发布数据信息,根据待监控网页的网址信息及所述数据信息的发布人生成发布源信息;根据所述数据信息的发布时间生成发布时间戳;获取所述发布数据信息中的所有文件及发布源信息、发布时间戳以得到新增数据信息。
- 根据权利要求18所述的计算机可读存储介质,其中,所述通过预设信息识别模型将非文字文件转换为文字文件,包括:获取所述非文字文件的格式信息并判断该文件是否为音频文件,若该文件 为音频文件则通过信息识别模型中的音频识别模型对该文件进行识别以得到相应的文字文件;获取所述非文字文件的格式信息并判断该文件是否为图片,若该文件为图片则通过信息识别模型中图片识别模型的对该文件进行识别以得到相应的文字文件;获取所述非文字文件的格式信息并判断该文件是否为视频文件,若该文件为视频文件则通过信息识别模型中的音频识别模型及图片识别模型对该文件进行识别以得到相应的文字文件。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201811526834.0A CN109657181B (zh) | 2018-12-13 | 2018-12-13 | 互联网信息链式存储方法、装置、计算机设备及存储介质 |
| CN201811526834.0 | 2018-12-13 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020119064A1 true WO2020119064A1 (zh) | 2020-06-18 |
Family
ID=66113068
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2019/092551 Ceased WO2020119064A1 (zh) | 2018-12-13 | 2019-06-24 | 互联网信息链式存储方法、装置、计算机设备及存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN109657181B (zh) |
| WO (1) | WO2020119064A1 (zh) |
Families Citing this family (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109657181B (zh) * | 2018-12-13 | 2024-05-14 | 平安科技(深圳)有限公司 | 互联网信息链式存储方法、装置、计算机设备及存储介质 |
| CN110490538B (zh) * | 2019-07-04 | 2023-08-22 | 平安科技(深圳)有限公司 | 信息链生成方法、装置、计算机设备和存储介质 |
| CN111125345B (zh) * | 2019-12-24 | 2024-04-16 | 南京三百云信息科技有限公司 | 数据应用方法和装置 |
| CN112104747B (zh) * | 2020-10-30 | 2021-02-26 | 广州市玄武无线科技股份有限公司 | 一种基于链式处理的请求响应系统 |
| CN112784077A (zh) * | 2021-03-17 | 2021-05-11 | 陕西省大数据集团有限公司 | 一种分类提取数据资产价值方法及装置 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101364955A (zh) * | 2008-09-28 | 2009-02-11 | 杭州电子科技大学 | 一种分析和提取电子邮件客户端证据的方法 |
| CN101882162A (zh) * | 2010-06-29 | 2010-11-10 | 北京搜狗科技发展有限公司 | 一种网络信息推送方法及系统 |
| WO2017206739A1 (zh) * | 2016-06-01 | 2017-12-07 | 广州市动景计算机科技有限公司 | 一种截图方法及装置 |
| CN109657181A (zh) * | 2018-12-13 | 2019-04-19 | 平安科技(深圳)有限公司 | 互联网信息链式存储方法、装置、计算机设备及存储介质 |
Family Cites Families (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20090157407A1 (en) * | 2007-12-12 | 2009-06-18 | Nokia Corporation | Methods, Apparatuses, and Computer Program Products for Semantic Media Conversion From Source Files to Audio/Video Files |
| CN103942639B (zh) * | 2014-03-21 | 2017-07-25 | 宁波中小在线信息服务有限公司 | 用于政策咨询服务系统的政策管理系统及其方法 |
| US10332506B2 (en) * | 2015-09-02 | 2019-06-25 | Oath Inc. | Computerized system and method for formatted transcription of multimedia content |
| CN106412678A (zh) * | 2016-09-14 | 2017-02-15 | 安徽声讯信息技术有限公司 | 一种视频新闻实时转写存储方法及系统 |
| CN107680602A (zh) * | 2017-08-24 | 2018-02-09 | 平安科技(深圳)有限公司 | 语音欺诈识别方法、装置、终端设备及存储介质 |
| CN108829765A (zh) * | 2018-05-29 | 2018-11-16 | 平安科技(深圳)有限公司 | 一种信息查询方法、装置、计算机设备及存储介质 |
-
2018
- 2018-12-13 CN CN201811526834.0A patent/CN109657181B/zh active Active
-
2019
- 2019-06-24 WO PCT/CN2019/092551 patent/WO2020119064A1/zh not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101364955A (zh) * | 2008-09-28 | 2009-02-11 | 杭州电子科技大学 | 一种分析和提取电子邮件客户端证据的方法 |
| CN101882162A (zh) * | 2010-06-29 | 2010-11-10 | 北京搜狗科技发展有限公司 | 一种网络信息推送方法及系统 |
| WO2017206739A1 (zh) * | 2016-06-01 | 2017-12-07 | 广州市动景计算机科技有限公司 | 一种截图方法及装置 |
| CN109657181A (zh) * | 2018-12-13 | 2019-04-19 | 平安科技(深圳)有限公司 | 互联网信息链式存储方法、装置、计算机设备及存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN109657181B (zh) | 2024-05-14 |
| CN109657181A (zh) | 2019-04-19 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2020119064A1 (zh) | 互联网信息链式存储方法、装置、计算机设备及存储介质 | |
| US11176141B2 (en) | Preserving emotion of user input | |
| US10108698B2 (en) | Common data repository for improving transactional efficiencies of user interactions with a computing device | |
| WO2019095586A1 (zh) | 会议纪要生成方法、应用服务器及计算机可读存储介质 | |
| WO2020253399A1 (zh) | 日志分类规则的生成方法、装置、设备及可读存储介质 | |
| US11468107B2 (en) | Enhance a mail application to format a long email conversation for easy consumption | |
| WO2021218069A1 (zh) | 基于场景动态配置的交互处理方法、装置、计算机设备 | |
| CN114461790B (zh) | 新闻事件主题自动生成方法、装置、电子设备及存储介质 | |
| EP3175375A1 (en) | Image based search to identify objects in documents | |
| US12411876B2 (en) | Answer information generation method | |
| WO2020103447A1 (zh) | 视频信息链式存储方法、装置、计算机设备及存储介质 | |
| CN111400361A (zh) | 数据实时存储方法、装置、计算机设备和存储介质 | |
| CN111639157A (zh) | 音频标记方法、装置、设备及可读存储介质 | |
| CN116796758A (zh) | 对话交互方法、对话交互装置、设备及存储介质 | |
| WO2020037921A1 (zh) | 表情图片提示方法、装置、计算机设备及存储介质 | |
| CN112101003A (zh) | 语句文本的切分方法、装置、设备和计算机可读存储介质 | |
| CN108846098B (zh) | 一种信息流摘要生成及展示方法 | |
| CN117131093A (zh) | 基于人工智能的业务数据处理方法、装置、设备及介质 | |
| CN111354349A (zh) | 一种语音识别方法及装置、电子设备 | |
| CN105786929A (zh) | 一种信息监测方法及装置 | |
| WO2019000697A1 (zh) | 信息检索方法、系统、服务器及可读存储介质 | |
| US20250124086A1 (en) | Systems and Methods For Structured Bayesian Classification For Content Management | |
| CN109977423B (zh) | 一种生词处理方法、装置、电子设备和可读存储介质 | |
| CN118916453A (zh) | 基于自研发gpt模型的智能运维方法及其相关设备 | |
| CN116629236A (zh) | 一种待办事项提取方法、装置、设备及存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19896008 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19896008 Country of ref document: EP Kind code of ref document: A1 |