WO2007147359A1 - Système et procédé permettant de rectifier les informations d'un fichier multimédia - Google Patents
Système et procédé permettant de rectifier les informations d'un fichier multimédia Download PDFInfo
- Publication number
- WO2007147359A1 WO2007147359A1 PCT/CN2007/070114 CN2007070114W WO2007147359A1 WO 2007147359 A1 WO2007147359 A1 WO 2007147359A1 CN 2007070114 W CN2007070114 W CN 2007070114W WO 2007147359 A1 WO2007147359 A1 WO 2007147359A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- information
- multimedia file
- name
- corrected
- module
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/50—Information retrieval; Database structures therefor; File system structures therefor of still image data
- G06F16/58—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/60—Information retrieval; Database structures therefor; File system structures therefor of audio data
- G06F16/68—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
- G06F16/686—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using information manually generated, e.g. tags, keywords, comments, title or artist information, time, location or usage information, user ratings
Definitions
- the present invention relates to the technical field of correcting multimedia file information, in particular, a multimedia file information correction system and a multimedia file information correction method. Background of the invention
- FIG. 1 is a schematic diagram showing the results obtained by searching using the above prior art. As shown in FIG. 1, the following problems exist:
- the "dancing mother” listed in the song name should be the album name, not part of the song name;
- "Love Entertainment" in the song name is the name of the website that provides the song, not the name of the song, and belongs to the unwanted advertisement text;
- the song name is not completely correct, and this entry lacks the artist name and album name;
- Another method is to use the music website to crawl the music file information, and then use the manually maintained template to extract the music file information on the page. Although the method can obtain more accurate information, the method uses manual maintenance.
- the template to extract the music file information on the page, and different templates usually need to maintain different templates for different websites, and the workload is very large, which leads to low efficiency and can only collect music file information on fewer websites. Summary of the invention
- the embodiment of the invention provides a multimedia file information correction system, which aims to effectively provide more accurate multimedia file information.
- the embodiment of the invention also proposes a multimedia file information correction method for effectively providing relatively accurate multimedia file information.
- an embodiment of the present invention provides a multimedia file information correction system, where the system includes an information storage module and an information correction module, where:
- the information saving module is configured to save the determined information of the multimedia file
- the information correction module is configured to search for the determined information related to the multimedia file to be corrected in the information saving module, and replace the to-be-corrected multimedia file information with the found determined information.
- the embodiment of the invention further provides a multimedia file information correction method, the method comprising: A. checking the determined information of the multimedia file according to the information of the multimedia file to be corrected Finding the determined information related to the information of the multimedia file to be corrected;
- the present invention utilizes the determined information of the multimedia file, the related determined information is found according to the multimedia file to be corrected, and then the corrected modified multimedia file is replaced with the found determined information.
- the information of the embodiment of the present invention is more efficient, so that the implementation of the above technical solution does not depend on the manually maintained template.
- the technical solution of the embodiment of the present invention can be applied to various aspects such as correcting music file information of a search engine, and correcting multimedia file information in a device such as a local computer.
- FIG. 1 is a schematic diagram of music file information obtained by searching using the prior art
- FIG. 2 is a schematic structural diagram of a multimedia file information correction system according to an embodiment of the present invention
- FIG. 3 is a schematic flowchart of a method for modifying multimedia file information according to an embodiment of the present invention
- FIG. 2 is a schematic structural diagram of a multimedia file information correction system according to an embodiment of the present invention. As shown in FIG. 2, the method includes at least an information storage module and an information modification module, and may further include a word segmentation module, a filtering module, an output module, and the like.
- the multimedia file information correction system of the embodiment of the present invention can be used in various aspects, such as correcting information of a multimedia file such as a music file searched by an existing search engine, correcting information of a multimedia file in a device such as a local computer, and More for other occasions
- the information of the media file is corrected.
- the information of the music file searched by the search engine is modified as an example. Those skilled in the art can replace the music file with other multimedia files and replace the music file information with the information corresponding to other multimedia files. .
- the multimedia file information correction system is further connected with a multimedia file search system that utilizes existing search technology from the Internet device and / or the local device obtains multimedia file information, for example, crawling the Internet with a crawler, extracting a link related to the music file, saving corresponding text information such as an anchor text, a title, a tag, and the like as the music to be corrected.
- File information provided to the multimedia file information correction system.
- the music file information includes a song name, an artist name, an album name, and the like.
- the multimedia file search system is prior art, it is not mentioned here.
- the information saving module stores the determined information of the music file, that is, the music file information such as the determined song name, the song name corresponding to the song, and the album name.
- Table 1 shows an implementation of the music file information in the information saving module: Table 1 No. Song name Artist name Album name
- the music information may be saved in another manner, for example: storing information of all the albums, including the song name and corresponding singer name in the album; or storing information of all the singers, including the singer name, singing The song name and the album name corresponding to the song. It is also possible not to save the correspondence between the above song name, artist name, and album name, but to save a number of song names, artist names, and album names.
- the information saving module can be a single module or a plurality of saving modules. That is, the song information can be saved in a save module or saved in multiple modules.
- the word segmentation module in FIG. 2 is used for word segmentation processing of music file information searched by the multimedia file search system, that is, word segmentation processing of anchor text, web page title and tag content saved in text form.
- the word segmentation of the above word segmentation module is divided into two cases:
- the word segmentation module uses the space and / or punctuation as a separator to segment the text, such as "Jay Chou - Nocturne” will be divided into “Jay Chou” and "Nocturne”.
- the second case is that the anchor text, the page title, and the tag content are separated by no separators, such as "Jay Chou nocturne.” Because there is no separator, so use the separator above.
- the word segmentation method cannot implement the word segmentation.
- the embodiment of the present invention uses the song name, the artist name, and the album name in the information storage module as a dictionary, and uses, for example, a reverse maximum matching method, a forward maximum matching method, a statistical-based word segmentation method, and the like.
- the word segmentation method is used for word segmentation.
- multiple word segmentation methods can be combined to implement word segmentation.
- the forward maximum matching method and the inverse maximum matching method are combined to form a two-way matching method to implement word segmentation.
- singer names, song names, album names, etc. are all proper nouns, which are less ambiguous.
- word segmentation can achieve better results. For example, “Jay Chou nocturne” can be correctly divided into “ Jay Chou “and “nocturne” two words.
- the page title and the tag content saved in the text form are processed by word segmentation, the anchor text, the page title and the tag content in the music file information to be corrected are formally converted into one or a group of words, as in the above example. It was converted into the words "Jay Chou” and "Nocturne".
- the filtering module shown in Fig. 2 is used for filtering processing the music files to be corrected, and removing unnecessary information such as advertisements, fraudulent information and the like which have no meaning to the user.
- the filtering module can directly filter the information of the multimedia file to be corrected without word segmentation, or filter the information of the multimedia file to be corrected after the word segmentation processing, and then filter the information of the multimedia file to be corrected after the word segmentation processing into The example is explained.
- Filtering conditions are stored in the filtering module.
- the filtering module filters unnecessary information such as advertisements or spoofing information according to the filtering conditions. If the filtering conditions are found, the related items in the multimedia file information to be corrected are removed.
- the filtering condition mainly includes two parts: an action and a comparison value, wherein the action may be include, the same, the length is greater than, the length is less than, and the like, and the contrast value may be various information.
- the filtering conditions in this embodiment are, for example, "including WWW”, "including XX” (XX is some pornographic vocabulary). If one or more information in the multimedia file information to be corrected satisfies the filtering condition, the filtering module removes related information in the multimedia file information to be corrected. For example, to be repaired If the multimedia file information includes "WWW" or "XX", etc., the filtering module removes the corresponding information.
- the filtering module in the present invention is also used to filter such information according to other filtering conditions, and the processing flow is as follows:
- the ratio between the two is judged. If it exceeds a certain threshold (for example, 50%), it can be determined as fraudulent information, and the corresponding entry in the multimedia file information to be corrected is removed.
- a certain threshold for example, 50%
- the fraud information, the advertisement information, and other information that is not required by the user are basically filtered, and a set of filtered information is obtained.
- the filtering condition is saved in the filtering module.
- the filtering condition may also be saved in the information saving module, and the filtering module is connected to the information saving module for filtering.
- the filter module can call the filter condition to filter.
- the filtering module filters out the information of the partial search items for the anchor text, the page title and the tag content of the word segmentation, a set of filtered information is obtained.
- the information correction module shown in FIG. 2 is configured to search for the determined information related to the multimedia file to be corrected in the information saving module, and replace the to-be-corrected one with the found determined information.
- the multimedia file information to be corrected here may be the original multimedia file information without word segmentation and filtering, and may be multimedia information that is only subjected to word segmentation or only filtering, and may be multimedia file information after word segmentation and filtering. The following is an example of the multimedia file information after segmentation and filtering.
- the song name determining step sorting the word segmentation and the filtered information in the order of the anchor text, the title, and the tag, and then matching the song names in the information saving module to find whether there is an exact matching song name, if any,
- the first matching song name is used as the song name of the music file, otherwise the song name with the highest similarity above the similarity criterion is used as the song name of the music file. That is to say, the found song name, artist name, and album name in the embodiment of the present invention may be exactly matched, or may be the highest similarity.
- the similarity is defined as: the ratio of the number of characters of the two information S1 and S2 to the average length of S1 and S2, where S1 and S2 are the two pieces of information.
- S1 and S2 are the two pieces of information.
- the similarity between "ABC” and “BCA” is 100%, while the similarity between "ABC” and “BCD” is 67%, and the similarity between "ABC” and “BA” is 80%.
- the criterion for similarity should be set to an appropriate value, such as 70%. If the similarity is less than 70%, it cannot be used as the song name.
- the singer name determining step after determining the song name, if the singer name corresponding to the song name in the music information saving module is unique, the singer name can be determined at the same time. If it corresponds to multiple singer names, it means that this is a song of the same name that has been sung by many people. Therefore, the word segmentation and filtered information are sequentially performed in the order of anchor text, title, and tag. Match the search to see if there is an exact match of the artist name. If there is, the first matching artist name will be the singer name of the music file, otherwise the highest similarity song will be judged above the standard.
- the name of the hand is the name of the artist of the music file. If it is not found, leave the name of the artist name blank.
- the album name determining step after determining the song name and the artist name, if the album name corresponding to the song name in the music information saving module is unique, the album name can be determined. If there are multiple album names, the word segmentation and the filtered words are matched and matched with the album name corresponding to the song name in the order of the anchor text, the title, and the tag, to see if there is an exact matching album name, and if so, A matching album name is used as the album name of the music file, otherwise the album name with the highest similarity above the judging criterion is used as the album name of the music file, and if not found, the album name is left blank.
- the output module shown in Figure 2 is used to output the multimedia file information after the information correction module is modified.
- the information stored in the information saving module may also be a correspondence between a hash code of the multimedia file and the determined information of the multimedia file. Then, the information correction module can search for the corresponding determined information according to the hash code of the multimedia file to be corrected as the determined information related to the multimedia file to be corrected, and replace the information to be corrected with the determined information, thereby realizing the modification of the multimedia file information.
- the data saving format in the information saving module can be as shown in Table 2: Table 2
- the file hash code can be a 32-bit or 64-bit integer. Commonly used hash algorithms include Cyclic Redundancy Check 32 (CRC32), Message Digest Algorithm 5 (MD5), and so on. If the two files F1 and F2 are equal in hash code calculated by the same algorithm, the contents of F1 and F2 can be considered to be completely equal.
- CRC32 Cyclic Redundancy Check 32
- MD5 Message Digest Algorithm 5
- the music file information correction method mainly includes the following steps: Step S1, using the existing search technology to acquire a music file to be modified from the Internet, step S2, and performing word segmentation processing on the modified multimedia file information, and the specific word segmentation processing is as follows. As mentioned before, it will not be traced here.
- step S3 the multimedia file information to be corrected after the word segmentation is filtered to remove unnecessary information.
- Step S4 the ⁇ participle word, the filtered multimedia file information to be corrected, find the determined information related to the to-be-corrected multimedia file, and replace the to-be-corrected multimedia file information with the found determined information.
- the step may further include: if the determined song name is found, replacing the song name of the music file to be corrected with the found song name; if the found song name has a unique corresponding artist name , the name of the singer to be corrected is replaced by the singer name, otherwise the singer name of the music file to be corrected is replaced with the name of the singer found; if the name of the song and/or the singer name found has a unique corresponding album name, then The album name replaces the music file album name to be corrected, otherwise the album name of the music file to be corrected is replaced with the found album name.
- the searched song name, artist name, and album name described here are exactly matched.
- Step S5 Output the corrected multimedia file information, and further display the information to the user.
- various inaccurate multimedia file information can be effectively corrected by using the determined multimedia file information to obtain accurate information of the multimedia file.
- the multimedia file modification system and method of the embodiments of the present invention are also highly efficient because no tools such as manual templates are required.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Library & Information Science (AREA)
- Data Mining & Analysis (AREA)
- Databases & Information Systems (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Multimedia (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
- Information Transfer Between Computers (AREA)
Description
多媒体文件信息修正系统及方法
技术领域
本发明涉及对多媒体文件信息进行修正的技术领域, 尤其是多媒体 文件信息修正系统和多媒体文件信息修正方法。 发明背景
目前的音乐文件信息采集主要有以下两种方式: 使用爬虫对整个互 联网进行爬行和对专门的音乐网站进行爬行。
通过使用爬虫对整个互联网进行爬行 , 并从中提取出与音乐文件有 关的链接, 保存对应的锚文本、 标题等文本信息, 并对这部分文本建立 索引, 用户通过输入关键字在其中进行检索, 然后通过搜索结果输出系 统展现给用户。 目前, 网络上大多数的搜索引擎都是采用这种方式进行 处理。
通过使用爬虫对整个互联网进行爬行来采集音乐文件信息 , 虽然得 到的可搜索的链接数量较多, 并且需要较少的人工干预, 但是这种方式 得到信息的准确率较低, 常出现以下的错误: 信息不完整, 如缺少歌曲 名、 歌手名、 专辑名等中的一个或多个; 信息的内容不准确, 比如出现 错别字、 错误写法等; 歌曲描述信息为无意义的广告或乱码; 歌曲描述 信息与真实内容不符; 故意堆砌大量广告关键字或热门歌曲名进行欺 骗。
图 1所示为利用上述现有技术进行搜索后得到的结果示意图, 如图 1所示, 存在如下问题:
第 2条搜索信息中, 被列在歌曲名中的 "舞娘" 应该是专辑名, 而 不是歌曲名的一部分;
第 10条搜索信息中,歌曲名中的 "爱娱乐 "是提供该歌曲的网站名, 不是歌曲名, 属于不需要的广告文字;
第 11条搜索信息中,歌曲名写法不完全正确, 并且此条目缺少歌手 名与专辑名;
第 16条, 搜索信息中, 歌曲名中的最后一个字是乱码, 且缺少歌手 名与专辑名。
另一种方法是利用对专门的音乐网站进行爬行获取音乐文件信息 , 然后利用人工维护的模板来提取页面上的音乐文件信息, 虽然该方法可 以获得比较准确的信息, 但是由于该方法使用人工维护的模板来提取页 面上的音乐文件信息, 而针对不同的网站通常需要维护不同的模板, 工 作量非常大, 这就导致了效率较低, 只能采集较少网站上的音乐文件信 息。 发明内容
本发明实施例提出了一种多媒体文件信息修正系统, 其目的在于有 效地提供较为准确的多媒体文件信息。 本发明实施例还提出了一种多媒 体文件信息修正方法, 用来有效地提供较为准确的多媒体文件信息。
为实现上述目的 , 本发明实施例提供了一种多媒体文件信息修正系 统, 该系统包括信息保存模块和信息修正模块, 其中:
所述信息保存模块用于保存多媒体文件的已确定信息;
所述信息修正模块用于在所述信息保存模块中查找与待修正多媒体 文件相关的已确定信息, 并用查找到的已确定信息替换待修正多媒体文 件信息。
本发明实施例还提供了一种多媒体文件信息修正方法,该方法包括: A. 根据待修正多媒体文件的信息在多媒体文件的已确定信息中查
找与待修正多媒体文件信息相关的已确定信息;
B. 用查找到的已确定信息替换待修正多媒体文件信息。
从上述技术方案可以看出, 由于本发明利用多媒体文件的已确定信 息, 根据待修正多媒体文件在其中查找到相关的已确定信息, 然后用所 查找到的已确定信息替换待修正多媒文件原来的信息, 从而提供了较为 准确的多媒体文件信息, 并且上述技术方案的实施不依赖于人工维护的 模板, 因此本发明实施例的技术方案还具有较高的效率。
另外, 本发明实施例的技术方案可以应用于对搜索引擎的音乐文件 信息的修正、 对本地计算机等设备中的多媒体文件信息进行修正等多个 方面。 附图简要说明
图 1为利用现有技术进行搜索得到的音乐文件信息的示意图; 图 2为本发明实施例的多媒体文件信息修正系统的结构示意图; 图 3为本发明实施例的多媒体文件信息修正方法的流程示意图。 实施本发明的方式
为使本发明的目的、 技术方案和优点更加清楚, 以下举实施例对本 发明进一步详细说明。
图 2为本发明实施例的多媒体文件信息修正系统的结构示意图, 如 图 2所示, 其至少包括信息保存模块和信息修正模块, 还可以进一步包 括分词模块、 过滤模块、 输出模块等。
本发明实施例的多媒体文件信息修正系统可以用在多个方面, 例如 对现有搜索引擎搜索到的音乐文件等多媒体文件的信息进行修正、 对本 地计算机等设备中的多媒体文件的信息进行修正以及对其它场合的多
媒体文件的信息进行修正。 在下面的描述中, 以修正搜索引擎搜索得到 的音乐文件的信息为例进行说明, 本领域技术人员只要将音乐文件替换 为其它多媒体文件、 将音乐文件信息替换为其它多媒体文件对应的信息 即可。
由于要对搜索引擎搜索得到的音乐文件的信息进行修正, 如图 2所 示, 多媒体文件信息修正系统进一步与多媒体文件搜索系统相连接, 该 多媒体文件搜索系统利用现有的搜索技术从互联网设备和 /或本地设备 获取多媒体文件信息, 例如利用爬虫对互联网爬行, 从中提取与音乐文 件相关的链接, 保存对应的锚文本、 标题、 标签(Tag )等文本信息, 并将这些信息作为待修正的音乐文件信息 , 提供给多媒体文件信息修正 系统。 在这里, 所述的音乐文件信息包括歌曲名、 歌手名、 专辑名等。 另外, 由于多媒体文件搜索系统为现有技术, 在此不再赞述。
下面结合附图对本发明的多媒体文件信息修正系统进行详细描述。 如图 2所示, 信息保存模块中保存有音乐文件的已确定信息, 即: 已经得到确定的歌曲名、 歌曲对应的歌手名及专辑名等音乐文件信息。 表 1给出了信息保存模块中音乐文件信息的一种实现方式: 表 1 序号 歌曲名 歌手名 专辑名
1 A1 A2 A3
B21 B31
2 B1
B22 B32
C31
3 C1 C2
C32
在表 1中描述了多种情况:
1、 歌曲 Al, 唯一由歌手 A2演唱, 其对应的专辑为 A3;
2、 歌曲 Bl, 歌手 B21和 B22都演唱过, 其对应的专辑分别为 B31 和 B32;
3、 歌曲 Cl, 唯一由歌手 C2演唱, 但该歌在两张专辑 C31和 C32 中出现。
当然上述表 1只是一种示范,一首歌也可以由两位以上的歌手演唱, 也可以出现在 2张以上的专辑中。
在本发明实施例中 , 也可以利用另外的方式实现音乐信息的保存 , 例如: 存储所有专辑的信息, 包括专辑中的歌曲名和对应的歌手名; 或 者存储所有歌手的信息, 包括歌手名、 演唱过的歌曲名和歌曲对应的专 辑名。 也可以不保存上述歌曲名、 歌手名以及专辑名之间的对应关系, 只是保存若干个歌曲名、 歌手名、 专辑名。
该信息保存模块可以是一个单独的模块, 也可以由多个保存模块组 成。 即, 歌曲信息可以保存在一个保存模块中, 也可以分多个模块进行 保存。
图 2中的分词模块用于对多媒体文件搜索系统搜索得到的音乐文件 信息进行分词处理, 即对以文本形式保存的锚文本、 网页标题和 Tag内 容进行分词处理。
上述分词模块的分词处理分为两种情况:
第一种情况是锚文本、 网页标题和 Tag内容中存在空格和 /或标点符 号。 此时, 分词模块以空格和 /或标点符号作为分隔符对文本进行分词, 如 "周杰伦-夜曲" 会被切分为 "周杰伦" 和 "夜曲" 两个词。
第二种情况是锚文本、 网页标题和 Tag内容中没有分隔符隔开, 如 "周杰伦夜曲"。 因为其中不带有分隔符, 因此使用上述的利用分隔符
分词的方法无法实现分词, 在此, 本发明实施例使用信息保存模块中的 歌曲名、 歌手名及专辑名作为词典, 利用诸如逆向最大匹配法、 正向最 大匹配法、 基于统计的分词方法等分词方法进行分词, 当然也可以多种 分词方法结合起来实现分词, 如将正向最大匹配法和逆向最大匹配法结 合起来构成双向匹配法实现分词。
一般来说, 歌手名、 歌曲名、 专辑名等均为专有名词, 较少产生歧 义, 使用上述方法进行分词即可达到较好的效果, 例如完全可以将 "周 杰伦夜曲" 正确地分为 "周杰伦" 和 "夜曲" 两个词。
对文本形式保存的锚文本、 网页标题和 Tag内容进行分词处理后, 待修正的音乐文件信息中的锚文本、 网页标题和 Tag内容在形式上都转 化为一个或一组词语, 如上面的例子中就被转化为 "周杰伦"和 "夜曲" 两个词语。
如图 2所示的过滤模块用于对待修正的音乐文件中进行过滤处理, 去除不需要的信息, 例如广告、 欺骗信息等对用户来说没有任何意义的 信息。 过滤模块可以直接对未经分词处理的待修正多媒体文件信息进行 过滤 , 也可以是对经过分词处理后的待修正多媒体文件信息进行过滤 , 下面以对分词处理后的待修正多媒体文件信息进行过滤为例进行说明。
过滤模块中保存有过滤条件, 过滤模块根据过滤条件对广告或欺骗 信息等不需要的信息进行过滤, 如果发现符合过滤条件, 则将待修正多 媒体文件信息中的相关条目去除。
具体来说, 过滤条件主要包括两部分: 动作和对比值, 其中动作可 以是包括、 相同、 长度大于、 长度小于等, 对比值可以是各种信息。 本 实施例中的过滤条件例如为 "包括 WWW"、 "包括 XX" ( XX为某些色 情词汇)。 如果待修正多媒体文件信息中的一个或多个信息满足过滤条 件, 则过滤模块将待修正多媒体文件信息中的相关信息去除。 如, 待修
正多媒体文件信息中包括 "WWW"或 "XX" 等, 则过滤模块将相应的 信息去除。
另外, 上述的过滤条件可随时进行修改、 删除、 增加等操作。
同时, 在待修正多媒体文件信息中, 还有可能出现以下的情况: 由 于音乐网站会在标题中罗列大量的热门歌曲名来提高自己的排名, 但并 不提供真实的内容。 对于这种情况, 本发明中的过滤模块还用于根据另 外的过滤条件过滤这种信息, 处理流程如下:
首先, 统计待修正多媒体文件信息中来自同一网站的搜索条目的总 数;
其次, 统计分词后各信息在该网站的信息中出现的次数;
最后, 判断二者之间的比例, 如果超过一定的阔值(如 50% ) 时, 即可判定为欺骗信息 , 将待修正多媒体文件信息中的相应条目去除。
如从某网站下共采集了 1000条搜索条目, 其中 "七里香"在搜索条 目的标题中出现了超过 500次, 则可判断这个词语被该网站用来当作欺 骗信息, 因此删除相关条目的信息。
通过上述的处理后, 基本过滤了欺骗信息、 广告信息以及其他一些 不是用户所需要的信息, 得到了一组过滤后的信息。
结合图 2所示, 上述的对过滤模块的描述中, 过滤条件保存在过滤 模块中, 当然, 也可以将过滤条件保存在信息保存模块中, 并将过滤模 块与信息保存模块连接, 在进行过滤处理的时候, 由过滤模块调用过滤 条件进行过滤即可实现。
在过滤模块对分词后的锚文本、 网页标题和 Tag内容的词组过滤掉 部分搜索条目的信息后, 得到了一组过滤后的信息。
如图 2所示的信息修正模块, 用于在信息保存模块中查找与待修正 多媒体文件相关的已确定信息 , 并用查找到的已确定信息替换待修正多
媒体文件的信息。 这里的待修正多媒体文件信息可以是未经分词、 过滤 的原始多媒体文件信息, 可以是只经过分词或只经过过滤的多媒体信 息, 可以是经过分词和过滤的多媒体文件信息。 下面以经过分词和过滤 的多媒体文件信息为例进行说明。
现在以修正音乐文件的歌曲名、 歌手名、 专辑名为例对信息修正模 块的处理进行伴细描述, 其包括歌曲名确定步骤、 歌手名确定步骤和专 辑名确定步骤, 其中:
歌曲名确定步骤, 将分词、 过滤后的信息按锚文本、 标题、 Tag 的 顺序排序, 然后依次与信息保存模块中的歌曲名进行匹配查找, 看是否 有完全匹配的歌曲名, 如果有, 将第一个匹配的歌曲名作为该音乐文件 的歌曲名 , 否则将相似度评判标准之上的相似度最高的歌曲名作为该音 乐文件的歌曲名。 也就是说, 本发明实施例中的查找到的歌曲名、 歌手 名、 专辑名可以是精确匹配的, 也可以是相似度最高的。
在此, 相似度的定义为: 两个信息 S1与 S2相同的字符数与 S1和 S2的平均长度的比值,其中, S1与 S2与比对的两个信息。例如, "ABC" 与 "BCA"的相似度为 100%, 而 "ABC"与 "BCD" 的相似度为 67%, 而 "ABC" 与 " BA" 的相似度为 80%。
在此, 相似度的评判标准应该设置一个合适的值, 如 70%, 如果相 似度低于 70%则不能将其作为歌曲名。
歌手名确定步骤, 在确定歌曲名之后, 如果音乐信息保存模块中该 歌曲名对应的歌手名是唯一的, 则可同时确定歌手名。 如果对应多个歌 手名, 则说明这是一首曾被多人翻唱过的同名歌曲, 因此, 将分词、 过 滤后的信息按锚文本、 标题、 Tag 的顺序依次与歌曲名对应的歌手名进 行匹配查找, 看是否有完全匹配的歌手名, 如果有, 将第一个匹配的歌 手名作为该音乐文件的歌手名 , 否则将评判标准之上的相似度最高的歌
手名作为该音乐文件的歌手名, 如果找不到, 则将歌手名这一项留空。 专辑名确定步骤, 在确定了歌曲名和歌手名之后, 如果音乐信息保 存模块中该歌曲名对应的专辑名是唯一的, 则可确定专辑名。 如果对应 多个专辑名, 则将分词、 过滤后的词语按锚文本、 标题、 Tag 的顺序依 次与歌曲名对应的专辑名进行匹配查找, 看是否有完全匹配的专辑名 , 如果有, 将第一个匹配的专辑名作为该音乐文件的专辑名, 否则将评判 标准之上的相似度最高的专辑名作为该音乐文件的专辑名, 如果找不 到, 则将专辑名这一项留空。
在每个音乐文件的歌曲名、 歌手名和专辑名确定后, 将待修正多媒 体文件信息中的对应信息进行替换。
如图 2所示的输出模块, 用于输出信息修正模块修正之后的多媒体 文件信息。
另一方面, 信息保存模块中保存的也可以是多媒体文件的哈希码与 该多媒体文件的已确定信息之间的对应关系。 那么信息修正模块可以根 据待修正多媒体文件的哈希码查找对应的已确定信息作为与待修正多 媒体文件相关的已确定信息, 并利用已确定信息替换待修正信息, 从而 实现多媒体文件信息的修正。
信息保存模块中的数据保存格式可以如表 2所示: 表 2
文件哈希码 歌曲名 歌手名 专辑名
( 234ABCD A1 B1 C1
0x5678CDEF A2 B2 C2
其中, 文件哈希码可以是一个 32位或 64位的整数, 常用的哈希算 法有循环冗余校验 32 ( CRC32 )、 报文摘要算法 5 ( MD5 )等。 如果两 个文件 F1和 F2用同一算法所计算出来的哈希码相等, 则可认为 F1和 F2的内容完全相等。
一般来说, 互联网上的音乐文件都存在一个文件多份拷贝的情况。 假设可以确定某音乐文件 F的准确信息, 则可以将这一信息保存在信息 保存模块中。 这样, 在以后遇到与音乐文件 F的哈希码相同的文件时, 可以直接从信息保存模块中查找到该音乐文件的准确信息。
根据本发明实施例的音乐文件信息修正方法主要包括如下步骤: 步骤 Sl, 利用现有的搜索技术从互联网获取音乐文件的待修正信 步骤 S2, 对待修正多媒体文件信息进行分词处理, 具体分词处理如 前所述, 这里不再追述。
步骤 S3, 对分词处理后的待修正多媒体文件信息进行过滤处理, 去 除不需要的信息。
步骤 S4, ^居分词、 过滤后的待修正多媒体文件信息, 查找与待修 正多媒体文件相关的已确定信息, 并用查找到的已确定信息替换待修正 多媒体文件信息。
以音乐文件信息为例, 该步骤可以进一步包括: 如果查找到已确定 歌曲名, 则用所查找到的歌曲名替换待修正音乐文件歌曲名; 如果所查 找到的歌曲名存在唯一对应的歌手名, 则用该歌手名替换待修正音乐文 件歌手名, 否则用所查找到的歌手名替换待修正音乐文件歌手名; 如果 所查找到的歌曲名和 /或歌手名存在唯一对应的专辑名,则用该专辑名替 换待修正音乐文件专辑名, 否则用所查找到的专辑名替换待修正音乐文 件专辑名。 这里所述的所查找到的歌曲名、 歌手名和专辑名为精确匹配
的歌曲名、 歌手名和专辑名或相似度最高的歌曲名、 歌手名和专辑名。 如果预先保存有多媒体文件的哈希码与该多媒体文件的已确定信息 之间的对应关系 , 则本步骤中根据待修正多媒体文件的哈希码查找对应 的已确定信息作为与待修正多媒体文件相关的已确定信息, 并用查找到 的已确定信息替换待修正多媒体文件信息。
步骤 S5 , 输出修正后的多媒体文件信息, 并进一步展现给用户。 这样,通过本发明实施例的多媒体文件信息修正系统和方法的使用 , 可以利用已经确定的多媒体文件信息对各种不准确的多媒体文件信息 进行有效的修正, 得到多媒体文件的准确信息。 由于不需要人工模板等 工具, 本发明实施例的多媒体文件修正系统和方法还具有很高的效率。
以上所述仅是本发明的优选实施方式, 应当指出, 对于本技术领域 的普通技术人员来说, 在不脱离本发明原理的前提下, 还可以作出若干 改进和润饰, 这些改进和润饰也应视为本发明的保护范围。
Claims
1、一种多媒体文件信息修正系统, 其特征在于, 该系统包括信息保 存模块和信息修正模块, 其中:
所述信息保存模块用于保存多媒体文件的已确定信息;
所述信息修正模块用于在所述信息保存模块中查找与待修正多媒体 文件相关的已确定信息, 并用查找到的已确定信息替换待修正多媒体文 件信息。
2、根据权利要求 1所述的系统, 其特征在于, 该系统进一步包括分 词模块和 /或过滤模块, 其中:
所述分词模块用于对待修正多媒体文件信息进行分词处理; 所述过滤模块用于对待修正多媒体文件信息进行过滤, 删除不需要 的信息。
3、 根据权利要求 1所述的系统, 其特征在于,
所述信息保存模块中保存有多媒体文件的哈希码与该多媒体文件的 已确定信息之间的对应关系;
所述信息修正模块根据待修正多媒体文件的哈希码查找对应的已确 定信息作为与待修正多媒体文件相关的已确定信息。
4、根据权利要求 1所述的系统, 其特征在于, 该系统进一步包括输 出模块, 所述输出模块用于输出多媒体文件修正后的信息。
5、根据权利要求 1所述的系统, 其特征在于, 所述多媒体文件信息 修正系统与多媒体文件搜索系统相连, 所述多媒体文件搜索系统用于从 互联网设备和 /或本地设备获取多媒体文件信息 ,并提供给多媒体文件信 息修正系统作为待修正多媒体文件的信息。
6、 一种多媒体文件信息修正方法, 其特征在于, 该方法包括:
A. 根据待修正多媒体文件的信息在多媒体文件的已确定信息中查 找与待修正多媒体文件信息相关的已确定信息;
B. 用查找到的已确定信息替换待修正多媒体文件信息。
7、 根据权利要求 6所述的方法, 其特征在于, 步骤 A之前进一步 包括: 对待修正多媒体文件信息进行分词处理; 和 /或,
对待修正多媒体文件信息进行过滤, 删除不需要的信息。
8、根据权利要求 6所述的方法, 其特征在于, 该方法预先保存有多 媒体文件的哈希码与该多媒体文件的已确定信息之间的对应关系;
在步骤 B中, 根据待修正多媒体文件的哈希码查找对应的已确定信 息作为与待修正多媒体文件相关的已确定信息。
9、 根据权利要求 6所述的方法, 其特征在于, 该方法进一步包括: 输出多媒体文件修正后的信息。
10、 根据权利要求 6所述的方法, 其特征在于, 所述待修正多媒体 文件信息为从互联网设备和 /或本地设备获取的多媒体文件信息。
11、 根据权利要求 6所述的方法, 其特征在于, 所述多媒体文件为 音乐文件;
所述多媒体文件的信息包括歌曲名、 和 /或歌手名、 和 /或专辑名。
12、 根据权利要求 6所述的方法, 其特征在于, 所述多媒体文件为 音乐文件, 所述多媒体文件的信息包括歌曲名、 歌手名和专辑名;
步骤 B包括:
如果查找到已确定歌曲名 , 则用所查找到的歌曲名替换待修正音乐 文件歌曲名;
如果所查找到的歌曲名存在唯一对应的歌手名 , 则用该歌手名替换 待修正音乐文件歌手名; 否则用所查找到的歌手名替换待修正音乐文件 歌手名;
如果所查找到的歌曲名和 /或歌手名存在唯一对应的专辑名 ,则用该 专辑名替换待修正音乐文件专辑名; 否则用所查找到的专辑名替换待修 正音乐文件专辑名。
13、根据权利要求 12所述的方法, 其特征在于, 所查找到的歌曲名 为精确匹配的歌曲名或相似度最高的歌曲名; 和 /或
所查找到的歌手名为精确匹配的歌手名或相似度最高的歌手名;和 / 或
所查找到的专辑名为精确匹配的专辑名或相似度最高的专辑名。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN200610061188.6 | 2006-06-15 | ||
| CN2006100611886A CN101071422B (zh) | 2006-06-15 | 2006-06-15 | 一种音乐文件搜索处理系统及方法 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2007147359A1 true WO2007147359A1 (fr) | 2007-12-27 |
Family
ID=38833082
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2007/070114 Ceased WO2007147359A1 (fr) | 2006-06-15 | 2007-06-14 | Système et procédé permettant de rectifier les informations d'un fichier multimédia |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN101071422B (zh) |
| WO (1) | WO2007147359A1 (zh) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110827074A (zh) * | 2019-10-31 | 2020-02-21 | 夏振宇 | 采用视频语音分析进行广告投放评估的方法 |
| CN115080523A (zh) * | 2021-03-16 | 2022-09-20 | 上海擎感智能科技有限公司 | 一种文件信息的显示方法及装置 |
Families Citing this family (16)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102207941A (zh) * | 2010-03-29 | 2011-10-05 | 上海博泰悦臻电子设备制造有限公司 | 车载音乐的提供、获取方法和装置以及车载音乐传输系统 |
| CN102289440A (zh) * | 2010-06-18 | 2011-12-21 | 上海博泰悦臻电子设备制造有限公司 | 音乐文件提供方法及其提供系统 |
| CN102289439B (zh) * | 2010-06-18 | 2014-10-22 | 上海博泰悦臻电子设备制造有限公司 | 音乐文件提供方法及其提供系统 |
| CN102929874A (zh) * | 2011-08-08 | 2013-02-13 | 深圳市快播科技有限公司 | 检索数据的排序方法及装置 |
| CN103150595B (zh) * | 2011-12-06 | 2016-03-09 | 腾讯科技(深圳)有限公司 | 数据处理系统中的自动配对选择方法和装置 |
| CN102662957B (zh) * | 2012-03-02 | 2015-02-18 | 百度在线网络技术(北京)有限公司 | 用于优化浏览器的搜索结果页面的装置及方法 |
| CN103049578A (zh) * | 2013-01-15 | 2013-04-17 | 深圳市宜搜科技发展有限公司 | 一种获取歌曲信息的方法及系统 |
| CN104484379B (zh) * | 2014-12-09 | 2018-06-12 | 百度在线网络技术(北京)有限公司 | 确定音乐实体关系的方法和装置及查询处理方法和装置 |
| CN105808627A (zh) * | 2014-12-31 | 2016-07-27 | 高德软件有限公司 | Poi信息更新、检索、poi数据包生成方法及装置 |
| CN105608129B (zh) * | 2015-12-16 | 2019-04-26 | 北京奇虎科技有限公司 | 文件清理方法、装置及系统、移动终端 |
| WO2018094689A1 (zh) * | 2016-11-25 | 2018-05-31 | 深圳前海达闼云端智能科技有限公司 | 一种改进浏览体验的方法、装置和设备 |
| EP3350726B1 (en) | 2016-12-09 | 2019-02-20 | Google LLC | Preventing the distribution of forbidden network content using automatic variant detection |
| CN108304367B (zh) * | 2017-04-07 | 2021-11-26 | 腾讯科技(深圳)有限公司 | 分词方法及装置 |
| CN108319635A (zh) * | 2017-12-15 | 2018-07-24 | 海南智媒云图科技股份有限公司 | 一种多平台音乐资源整合播放的方法、电子设备及存储介质 |
| CN110717062B (zh) * | 2018-07-11 | 2024-03-22 | 斑马智行网络(香港)有限公司 | 音乐搜索及车载音乐播放方法、装置、设备以及存储介质 |
| CN109543064B (zh) * | 2018-11-30 | 2020-12-18 | 北京微播视界科技有限公司 | 歌词显示处理方法、装置、电子设备及计算机存储介质 |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN1501395A (zh) * | 1995-07-26 | 2004-06-02 | 更新小型光盘变换器及记录介质播放器中的存储器的方法 | |
| US20050065912A1 (en) * | 2003-09-02 | 2005-03-24 | Digital Networks North America, Inc. | Digital media system with request-based merging of metadata from multiple databases |
| JP2005309712A (ja) * | 2004-04-21 | 2005-11-04 | Sharp Corp | 楽曲検索システムおよび楽曲検索方法 |
Family Cites Families (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP4189758B2 (ja) * | 2004-06-30 | 2008-12-03 | ソニー株式会社 | コンテンツ記憶装置、コンテンツ記憶方法、コンテンツ記憶プログラム、コンテンツ転送装置、コンテンツ転送プログラム及びコンテンツ転送記憶システム |
-
2006
- 2006-06-15 CN CN2006100611886A patent/CN101071422B/zh active Active
-
2007
- 2007-06-14 WO PCT/CN2007/070114 patent/WO2007147359A1/zh not_active Ceased
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN1501395A (zh) * | 1995-07-26 | 2004-06-02 | 更新小型光盘变换器及记录介质播放器中的存储器的方法 | |
| US20050065912A1 (en) * | 2003-09-02 | 2005-03-24 | Digital Networks North America, Inc. | Digital media system with request-based merging of metadata from multiple databases |
| JP2005309712A (ja) * | 2004-04-21 | 2005-11-04 | Sharp Corp | 楽曲検索システムおよび楽曲検索方法 |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110827074A (zh) * | 2019-10-31 | 2020-02-21 | 夏振宇 | 采用视频语音分析进行广告投放评估的方法 |
| CN115080523A (zh) * | 2021-03-16 | 2022-09-20 | 上海擎感智能科技有限公司 | 一种文件信息的显示方法及装置 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN101071422A (zh) | 2007-11-14 |
| CN101071422B (zh) | 2010-10-13 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2007147359A1 (fr) | Système et procédé permettant de rectifier les informations d'un fichier multimédia | |
| Haveliwala et al. | Scalable Techniques for Clustering the Web. | |
| CN101876981B (zh) | 一种构建知识库的方法及装置 | |
| CN1664818B (zh) | 用于单词拆分的新词收集方法和系统 | |
| CN102171683B (zh) | 从查询日志中挖掘新词用于输入方法编辑器 | |
| US20220147526A1 (en) | Keyword and business tag extraction | |
| US11487816B2 (en) | Facilitating video search | |
| CN103198079B (zh) | 相关搜索的实现方法和装置 | |
| US20050210017A1 (en) | Error model formation | |
| US20070250501A1 (en) | Search result delivery engine | |
| US20120166414A1 (en) | Systems and methods for relevance scoring | |
| CN104850574A (zh) | 一种面向文本信息的敏感词过滤方法 | |
| WO2002041161A1 (en) | Method and apparatus for efficient identification of duplicate and near-duplicate documents and text spans using high-discriminability text fragments | |
| US7730316B1 (en) | Method for document fingerprinting | |
| CN101154228A (zh) | 一种分段模式匹配方法及其装置 | |
| CN101477527A (zh) | 一种检索多媒体资源的方法及装置 | |
| CN103838798A (zh) | 页面分类系统及页面分类方法 | |
| CN108846016A (zh) | 一种面向中文分词的搜索算法 | |
| CN101714147B (zh) | 相同或相似文件的过滤方法 | |
| US20100161615A1 (en) | Index anaysis apparatus and method and index search apparatus and method | |
| CN112100318B (zh) | 一种多维度信息合并方法、装置、设备及存储介质 | |
| CN102799586B (zh) | 一种用于搜索结果排序的转义度确定方法和装置 | |
| CN103646017B (zh) | 用于命名的缩略词生成系统及其工作方法 | |
| Liu | Information retrieval and Web search | |
| CN103136212A (zh) | 一种类别新词的挖掘方法及装置 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 07721735 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS EPO FORM 1205A DATED 04.05.2009. |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 07721735 Country of ref document: EP Kind code of ref document: A1 |