WO2018094952A1 - 一种内容推荐方法与装置 - Google Patents

一种内容推荐方法与装置 Download PDF

Info

Publication number
WO2018094952A1
WO2018094952A1 PCT/CN2017/079624 CN2017079624W WO2018094952A1 WO 2018094952 A1 WO2018094952 A1 WO 2018094952A1 CN 2017079624 W CN2017079624 W CN 2017079624W WO 2018094952 A1 WO2018094952 A1 WO 2018094952A1
Authority
WO
WIPO (PCT)
Prior art keywords
voice
user
information
feature information
recommendation
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2017/079624
Other languages
English (en)
French (fr)
Inventor
崔宝宏
王方舟
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Baidu Online Network Technology Beijing Co Ltd
Original Assignee
Baidu Online Network Technology Beijing Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Baidu Online Network Technology Beijing Co Ltd filed Critical Baidu Online Network Technology Beijing Co Ltd
Publication of WO2018094952A1 publication Critical patent/WO2018094952A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/70Information retrieval; Database structures therefor; File system structures therefor of video data
    • G06F16/73Querying
    • G06F16/735Filtering based on additional data, e.g. user or group profiles
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/60Information retrieval; Database structures therefor; File system structures therefor of audio data
    • G06F16/63Querying
    • G06F16/635Filtering based on additional data, e.g. user or group profiles
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/48Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
    • G10L25/51Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
    • G10L25/54Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination for retrieval

Definitions

  • the present invention relates to the field of voice information search technology, and in particular, to a content recommendation technology.
  • voice search schemes are based on the recognition of the user's voice content. After identifying the user's voice content, the search engine provides the user with various content corresponding to the voice content. Therefore, these voice search schemes are still traditional content search in their essence, but the user's input mode has changed, instead of text input, voice input can be used.
  • a content recommendation method comprising the following steps:
  • a content recommendation apparatus comprising:
  • the present invention can recommend matching voice materials, such as songs, video clips, etc., according to the user's voice feature information.
  • the present invention uses a fun operation method to increase the user's willingness to use the voice automatically in the gradual outbreak of the voice interaction application scenario, which is beneficial to cultivating the habit of the user using the voice to search. For example, based on the social and entertainment needs of the current voice audience, when the user searches for music or uses voice to automatically analyze various voice feature information of the user, such as sound color, audio, etc., the present invention appropriately recommends to the user to sing through voice big data analysis. Songs or singers suitable for imitation.
  • FIG. 1 shows a flow chart of a method for performing content recommendation in accordance with one embodiment of the present invention
  • FIG. 2 shows a schematic diagram of an apparatus for performing content recommendation in accordance with one embodiment of the present invention.
  • Computer equipment also known as “computer” means that it can be transported
  • An intelligent electronic device that schedules a program or instruction to perform a predetermined process such as numerical calculations and/or logic calculations, which may include a processor and a memory, the processor executing program instructions pre-stored in the memory to perform a predetermined process, or The predetermined processing is performed by hardware such as an ASIC, an FPGA, or a DSP, or a combination of the two.
  • Computer devices include, but are not limited to, servers, personal computers (PCs), notebook computers, tablets, smart phones, and the like.
  • the computer device includes, for example, a user device and a network device.
  • the user equipment includes, but is not limited to, a personal computer (PC), a notebook computer, a mobile terminal, etc., and the mobile terminal includes, but is not limited to, a smart phone, a PDA, etc.;
  • the network device includes but is not limited to a single network server, and more A network of server servers or a cloud-based cloud consisting of a large number of computers or network servers, where cloud computing is a type of distributed computing, a super-virtual set of a loosely coupled set of computers. computer.
  • the computer device can be operated separately to implement the present invention, and can also access the network and implement the present invention by interacting with other computer devices in the network.
  • the network in which the computer device is located includes, but is not limited to, the Internet, a wide area network, a metropolitan area network, a local area network, a VPN network, and the like.
  • the user equipment, the network equipment, the network, and the like are merely examples, and other existing or future possible computer equipment or networks, such as those applicable to the present invention, are also included in the scope of the present invention. It is included here by reference.
  • the invention can be implemented by a computer device.
  • the present invention can be implemented by a network device, but those skilled in the art will appreciate that the solution of the present invention can be implemented by a user device as long as it has the computing/processing capabilities required by the present invention.
  • the following description of the implementation of the network device is exemplified in the present specification, but those skilled in the art should understand that the examples are only for explaining the purpose of the present invention, and should not be construed as limiting the present invention.
  • FIG. 1 shows a method flow diagram in accordance with one embodiment of the present invention, in which a content recommendation process is specifically illustrated.
  • step S1 the network device receives the voice information submitted by the user; in step S2, the network device extracts the voice feature information of the user according to the voice information; in step S3, the network device is configured according to The sound feature information queries a preset first voice database to obtain a matched first voice data; in step S4, the network device recommends the first voice data to the user.
  • the present invention is typically applicable to search scenes or music interaction class scenarios.
  • the first voice material recommended to the user may be any voice material that is matched by the voice feature information, including but not limited to songs and video clips.
  • voice data is not intended to limit its expression, such as various audio materials, but may include any material having audio content, so that it can include not only songs but also songs.
  • a video clip can be included.
  • the network device recommends a first voice material according to the voice feature information of the user, for example, a video segment suitable for the user to imitate.
  • the network device also needs to determine whether the user has entertainment requirements, such as identifying the voice content submitted by the user, to analyze whether the content delivered by the voice has video, music, and other related entertainment needs, for example, by using a preset keyword table.
  • Entertainment needs, such as whether the voice content is a lyric, a line of words, and the like.
  • the network device can recommend the first voice material according to the voice feature information of the user, for example, a song suitable for the user to sing.
  • the content recommendation scheme of the present invention may be embodied as a “matching” function, which may provide the user with the first voice material that matches the voice feature information.
  • the client may present a "match" function button to the user, such as placed on the right side of the voice search box or at a specific location in the interface of the music class APP.
  • a "match" function button to the user, such as placed on the right side of the voice search box or at a specific location in the interface of the music class APP.
  • the "match" function can be directly prompted to the user to remind the user that voice information can be recommended for the previously submitted voice information.
  • match function button is only used to illustrate the purpose of the present invention, and its expression in the actual application may be different, for example, it may be presented as a "main song” function button, "large The “star” function button or the like, as long as the function buttons are intended to provide the user with a content recommendation scheme as in the present invention, it should fall within the scope of the patent protection of the present invention.
  • network device in the present invention is not limited to one computer device, and may be a plurality of computer devices, typically such as multiple servers, which cooperate with each other to implement the present invention.
  • Content recommendation program is not limited to one computer device, and may be a plurality of computer devices, typically such as multiple servers, which cooperate with each other to implement the present invention.
  • the interface server is responsible for interacting with the client and other servers, wherein the interaction with the client is such as receiving voice information sent by the client and returning the recommended first voice data to the client, and interacting with the matching server, such as submitting voice information. And matching the server and obtaining the matched first voice data from the matching server; the matching server is responsible for performing matching query on the voice information submitted by the interface server in the first voice database to obtain the matched first voice data.
  • These servers are considered as a whole relative to other external devices and are generally considered as "network devices" in their entirety.
  • step S1 the network device receives the voice information submitted by the user.
  • the present invention requires the network device to interact with the client to implement.
  • the user end is, for example, various functional entities running in the user equipment, such as a webpage in a PC, an APP in a mobile phone, and the like.
  • the user submits a piece of voice information through the user equipment, and the user equipment sends the voice information to the network device, so that the network device receives the voice information submitted by the user.
  • step S2 the network device extracts the voice feature information of the user according to the voice information submitted by the user.
  • the user's voice feature information can be defined from at least the following two dimensions:
  • the perceived dimension is intended to be defined from the perceived characteristics of the sound.
  • the sound feature information is specifically in this dimension such as pitch feature information, pitch feature information, melody feature information, rhythm feature information, audio feature information, speech rate feature information, and the like.
  • the acoustic dimension is intended to be expressed from various features defined by the acoustic angle of the sound.
  • the sound feature information is specifically such as energy feature information, zero-crossing rate, LPC (Linear Prediction Coefficient) parameter information, and the like in this dimension.
  • the zero-crossing rate means the ratio of the sign change of the sound signal, such as the sound signal changing from a positive number to a negative number or a reverse direction.
  • the network device may extract the voice feature information by using any known or future feasible technology, which is not limited by the present invention.
  • step S3 the network device queries the preset first voice data database according to the user's voice feature information to obtain the matched first voice data.
  • the network device may pre-establish a first voice database in which various sound feature information of each voice material is stored.
  • the first voice database may also be established and maintained by other devices, and the network device may have access rights thereto.
  • a song or a participating video segment of each star is stored in the first voice database, and the voice materials are labeled with various sound feature information.
  • the network device performs a matching query in the first voice database according to the sound feature information recognized by the user to obtain a matched song or video segment.
  • step S4 the network device recommends the matched first voice material to the user.
  • the network device recommends the matched songs or video clips and the like to the user, so that the present invention can provide the user with a song or video dubbing suitable for singing or imitating, so as to customize the "title song” or "famous song” for the user. And help users find their singer that is suitable for imitation. This is practical in commercial applications and significantly increases user stickiness and length of use.
  • the invention automatically matches singers, poets, celebrities, etc. by guiding the user through various voice inputs such as speaking, shouting, singing two sentences, reading two poems and the like. Further, the present invention can also Continue to receive the user's singing of the recommended songs, and then score the user's own singing relative to the original singing, and by sharing the friends, inviting friends to PK, etc. to enhance the fun to help users find the songs that are most suitable for singing, create users I imitate the famous songs, and gradually guide the user to form a voice search habit and a second interaction with the search engine to provide user stickiness and duration of use. Accordingly, the present invention can also be derived from a search engine to a social field by providing a sharing function.
  • the content recommendation process shown in FIG. 1 further includes a step S5.
  • step S5 the network device queries the preset second voice data database according to the content of the user voice information to obtain the matched second voice data; and then the network device recommends the matched second voice data to the user.
  • step of recommending the second voice material by the network device may be merged with the recommendation of step S4, so that step S5 occurs before step S4, and the network device may recommend the first voice data and the second to the user in step S4. Voice data.
  • the recommendation of the network device for the second voice material may also be independent of the recommendation of the first voice material in step S4.
  • the acquisition of the second voice material in step S5 and the acquisition of the second voice material in step S6 may be performed before, after or simultaneously with the two steps. recommend.
  • the network device After identifying the content in the voice information of the user, the network device performs a matching query in the second voice database to obtain second voice data with the same content, such as voice data with the same content but different sounds. If you use the same content without dialect.
  • the voice information submitted by the user is "Hello” in Mandarin
  • the network device recognizes that the content information is “Hello” and performs a matching query in the second voice database to obtain “Hello” in various dialects. For example, “Hello” in Shanghai dialect and Cantonese.
  • the first voice database can be integrated with the second voice library. That is, in the present invention, only one voice database is used to perform matching query of the first voice data and the second voice data, so that at least the voice feature information and the content information of each voice material are stored in the voice database.
  • the content recommendation device 20 is installed in the network device 200, and specifically includes a receiving device 21, an extracting device 22, a matching device 23, and a recommending device 24.
  • the receiving device 21 receives the voice information submitted by the user; the extracting device 22 extracts the sound feature information of the user according to the voice information; the matching device 23 queries the preset first voice database according to the sound feature information. A matching first voice material is obtained; the recommendation device 24 recommends the first voice material to the user.
  • the present invention is typically applicable to search scenes or music interaction class scenarios.
  • the first voice material recommended to the user may be any voice material that is matched by the voice feature information, including but not limited to songs and video clips.
  • voice data is not intended to limit its expression, such as various audio materials, but may include any material having audio content, so that it can include not only songs but also songs.
  • a video clip can be included.
  • the content recommendation device 20 recommends a first voice material for the user according to the voice feature information of the user, for example, a video segment suitable for the user to imitate.
  • the content recommendation device 20 or other devices in the network device also need to determine whether the user has entertainment requirements, such as identifying the voice content submitted by the user, to analyze whether the content delivered by the voice has video, music, and other related entertainment needs, for example, Identify the entertainment needs through a preset keyword list, such as whether the voice content is a lyric, a line, etc.
  • the user currently uses the music class APP.
  • the content recommendation device 20 may recommend the first voice material according to the voice feature information of the user, for example, a song suitable for the user to sing. .
  • the content recommendation scheme of the present invention may be embodied as a “matching” function, which may provide the user with the first voice material that matches the voice feature information.
  • the client may present a "match" function button to the user, such as placed on the right side of the voice search box or at a specific location in the interface of the music class APP.
  • the user sends the voice information submitted by the user to the network device, so that the content recommendation device 20 recommends the corresponding first according to the user's voice feature information.
  • Voice data such as placed on the right side of the voice search box or at a specific location in the interface of the music class APP.
  • the "match" function can be directly prompted to the user to remind the user that voice information can be recommended for the previously submitted voice information.
  • match function button is only used to illustrate the purpose of the present invention, and its expression in the actual application may be different, for example, it may be presented as a "main song” function button, "large The “star” function button or the like, as long as the function buttons are intended to provide the user with a content recommendation scheme as in the present invention, it should fall within the scope of the patent protection of the present invention.
  • network device in the present invention is not limited to one computer device, and may be a plurality of computer devices, typically such as multiple servers, which cooperate with each other to implement the present invention.
  • Content recommendation program is not limited to one computer device, and may be a plurality of computer devices, typically such as multiple servers, which cooperate with each other to implement the present invention.
  • the interface server is responsible for interacting with the client and other servers, wherein the interaction with the client is such as receiving voice information sent by the client and returning the recommended first voice data to the client, and interacting with the matching server, such as submitting voice information. And matching the server and obtaining the matched first voice data from the matching server; the matching server is responsible for performing matching query on the voice information submitted by the interface server in the first voice database to obtain the matched first voice data.
  • These servers are considered as a whole relative to other external devices and are generally considered as "network devices" in their entirety.
  • the receiving device 21 receives the voice information submitted by the user.
  • the present invention requires the network device to interact with the client to implement.
  • the user end is, for example, various functional entities running in the user equipment, such as a webpage in a PC, an APP in a mobile phone, and the like.
  • the user submits a piece of voice information through the user equipment, and the user equipment sends the voice information to the network device, so that the receiving device 21 receives the voice information submitted by the user.
  • the extracting means 22 extracts the sound feature information of the user based on the voice information submitted by the user.
  • the user's voice feature information can be defined from at least the following two dimensions:
  • the perceived dimension is intended to be defined from the perceived characteristics of the sound.
  • the sound feature information is specifically in this dimension such as pitch feature information, pitch feature information, melody feature information, rhythm feature information, audio feature information, speech rate feature information, and the like.
  • the acoustic dimension is intended to be expressed from various features defined by the acoustic angle of the sound.
  • the sound feature information is specifically such as energy feature information, zero-crossing rate, LPC (Linear Prediction Coefficient) parameter information, and the like in this dimension.
  • the zero-crossing rate means the ratio of the sign change of the sound signal, such as the sound signal changing from a positive number to a negative number or a reverse direction.
  • the extracting device 22 may extract the sound feature information by using any known or future feasible technology, which is not limited by the present invention.
  • the matching device 23 queries the preset first voice data database according to the user's voice feature information to obtain a matched first voice data.
  • the network device may pre-establish a first voice database in which various sound feature information of each voice material is stored.
  • the first voice database may also be established and maintained by other devices, and the matching device 23 may have access rights thereto.
  • a song or a participating video segment of each star is stored in the first voice database, and the voice materials are labeled with various sound feature information.
  • the matching device 23 performs a matching query in the first voice database according to the sound feature information recognized by the user to obtain a matched song or video segment.
  • the recommendation device 24 recommends the matched first voice material to the user.
  • the recommendation device 24 recommends the matched songs or video clips and the like to the user, so that the present invention can provide the user with a song or video dubbing suitable for singing or imitating, in order to customize the user's "title song” or “famous". "Song" and help users find their singer suitable for imitation. . This is practical in commercial applications and significantly increases user stickiness and length of use.
  • the invention automatically matches singers, poets, celebrities, etc. by guiding the user through various voice inputs such as speaking, shouting, singing two sentences, reading two poems and the like. Further, the present invention can continue to receive the user's singing of the recommended song, and then the user's own singing is relative to the original singing. Score and share the fun with friends, invite friends to PK, etc. to help users find the songs that they are most suitable to sing, create their own imitations, and gradually guide users to form voice search habits and search engines. A second interaction that provides user stickiness and length of use. Accordingly, the present invention can also be derived from a search engine to a social field by providing a sharing function.
  • the content recommendation device 20 shown in FIG. 2 further includes another matching device (hereinafter referred to as a second matching device).
  • the second matching device queries the preset second voice database according to the content of the user voice information to obtain the matched second voice data; subsequently, the content recommendation device 20 recommends the matched second voice data to the user.
  • the recommendation of the second voice material by the content recommendation device 20 may be combined with the above recommendation for the first voice material, so that the operation performed by the second matching device occurs before the operation performed by the recommendation device 24, and the recommendation device 24 may The first voice material and the second voice data are recommended to the user.
  • the recommendation of the second voice material by the content recommendation device 20 may also be independent of the recommendation of the first voice material by the recommendation device 24.
  • the second matching device may perform the second voice data before, after, or while performing operations of the two devices.
  • the content recommendation device 20 also needs to include another recommendation device (hereinafter referred to as a second recommendation device). To perform a recommendation for the second voice material.
  • the second matching device can still be integrated with the matching device 23, and the second recommending device is integrated with the recommendation device 24. At this time, each integrated device performs two matching operations or two recommended operations, respectively.
  • the second matching device performs a matching query in the second voice database to obtain the same content.
  • Two voice data such as voice data with the same content but different sounds, such as using the same content without dialect.
  • the voice message submitted by the user is “hello” in Mandarin
  • the second matching device is in the first
  • match queries are made to obtain “hello” in various dialects, such as “Hello” in Shanghai dialect and Cantonese.
  • the first voice database can be integrated with the second voice library. That is, in the present invention, only one voice database is used to perform matching query of the first voice data and the second voice data, so that at least the voice feature information and the content information of each voice material are stored in the voice database.
  • the present invention can be implemented in software and/or a combination of software and hardware.
  • the various devices of the present invention can be implemented using an application specific integrated circuit (ASIC) or any other similar hardware device.
  • the software program of the present invention may be executed by a processor to implement the steps or functions described above.
  • the software program (including related data structures) of the present invention can be stored in a computer readable recording medium such as a RAM memory, a magnetic or optical drive or a floppy disk and the like.
  • some of the steps or functions of the present invention may be implemented in hardware, for example, as a circuit that cooperates with a processor to perform various steps or functions.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Data Mining & Analysis (AREA)
  • Databases & Information Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Computational Linguistics (AREA)
  • Human Computer Interaction (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Acoustics & Sound (AREA)
  • Health & Medical Sciences (AREA)
  • Signal Processing (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

一种内容推荐方法与装置。其中,网络设备接收用户提交的语音信息(S1);根据所述语音信息,提取所述用户的声音特征信息(S2);根据所述声音特征信息,查询预置的第一语音资料库,以获得相匹配的第一语音资料(S3);将所述第一语音资料推荐给所述用户(S4)。本技术方案可以根据用户的声音特征信息为其推荐相匹配的语音资料,诸如歌曲、视频片段等。在如今逐步爆发的语音交互应用场景中,能够使用趣味性运营方法来增加用户自动使用语音的意愿,有利于培养用户使用语音进行搜索的习惯。例如,基于当前语音受众的社交及娱乐需求,当用户搜索音乐或者使用语音时自动分析用户的各项声音特征信息,诸如声色、音频等,通过语音大数据分析来适度推荐给用户适合演唱的歌曲或者适合模仿的歌手。

Description

一种内容推荐方法与装置
本申请以一中国专利申请作为优先权申请,该中国专利申请的申请日为2016年11月22日,申请号为201611036906.4,发明名称为“一种内容推荐方法与装置”。
技术领域
本发明涉及语音信息搜索技术领域,尤其涉及一种内容推荐技术。
背景技术
现有的各种语音搜索方案,均是基于对用户的语音内容的识别来进行。在识别出用户的语音内容后,搜索引擎为用户提供各种与该语音内容相应的内容。因此,这些语音搜索方案从其本质而言,仍是传统的内容搜索,只是用户的输入方式发生了变化,不再是文字输入,而可以采用语音输入。
发明内容
本发明的目的是提供一种内容推荐方法与装置。
根据本发明的一个方面,提供了一种内容推荐方法,其中,该方法包括以下步骤:
-接收用户提交的语音信息;
-根据所述语音信息,提取所述用户的声音特征信息;
-根据所述声音特征信息,查询预置的第一语音资料库,以获得相匹配的第一语音资料;
-将所述第一语音资料推荐给所述用户。
根据本发明的另一个方面,还提供了一种内容推荐装置,其中,该装置包括:
-用于接收用户提交的语音信息的装置;
-用于根据所述语音信息,提取所述用户的声音特征信息的装置;
-用于根据所述声音特征信息,查询预置的第一语音资料库,以获得相匹配的第一语音资料的装置;
-用于将所述第一语音资料推荐给所述用户的装置。
与现有技术相比,本发明可以根据用户的声音特征信息为其推荐相匹配的语音资料,诸如歌曲、视频片段等。本发明在如今逐步爆发的语音交互应用场景中,使用趣味性运营方法来增加用户自动使用语音的意愿,有利于培养用户使用语音进行搜索的习惯。例如,基于当前语音受众的社交及娱乐需求,当用户搜索音乐或者使用语音时自动分析用户的各项声音特征信息,诸如声色、音频等,本发明通过语音大数据分析来适度推荐给用户适合演唱的歌曲或者适合模仿的歌手。
附图说明
通过阅读参照以下附图所作的对非限制性实施例所作的详细描述,本发明的其它特征、目的和优点将会变得更明显:
图1示出根据本发明一个实施例的进行内容推荐的方法流程图;
图2示出根据本发明一个实施例的进行内容推荐的装置示意图。
附图中相同或相似的附图标记代表相同或相似的部件。
具体实施方式
在更加详细地讨论示例性实施例之前应当提到的是,一些示例性实施例被描述成作为流程图描绘的处理或方法。虽然流程图将各项操作描述成顺序的处理,但是其中的许多操作可以被并行地、并发地或者同时实施。此外,各项操作的顺序可以被重新安排。当其操作完成时所述处理可以被终止,但是还可以具有未包括在附图中的附加步骤。所述处理可以对应于方法、函数、规程、子例程、子程序等等。
在上下文中所称“计算机设备”,也称为“电脑”,是指可以通过运 行预定程序或指令来执行数值计算和/或逻辑计算等预定处理过程的智能电子设备,其可以包括处理器与存储器,由处理器执行在存储器中预存的程序指令来执行预定处理过程,或是由ASIC、FPGA、DSP等硬件执行预定处理过程,或是由上述二者组合来实现。计算机设备包括但不限于服务器、个人电脑(PC)、笔记本电脑、平板电脑、智能手机等。
所述计算机设备例如包括用户设备与网络设备。其中,所述用户设备包括但不限于个人电脑(PC)、笔记本电脑、移动终端等,所述移动终端包括但不限于智能手机、PDA等;所述网络设备包括但不限于单个网络服务器、多个网络服务器组成的服务器组或基于云计算(Cloud Computing)的由大量计算机或网络服务器构成的云,其中,云计算是分布式计算的一种,由一群松散耦合的计算机集组成的一个超级虚拟计算机。其中,所述计算机设备可单独运行来实现本发明,也可接入网络并通过与网络中的其他计算机设备的交互操作来实现本发明。其中,所述计算机设备所处的网络包括但不限于互联网、广域网、城域网、局域网、VPN网络等。
需要说明的是,所述用户设备、网络设备和网络等仅为举例,其他现有的或今后可能出现的计算机设备或网络如可适用于本发明,也应包含在本发明保护范围以内,并以引用方式包含于此。
本文后面所讨论的方法(其中一些通过流程图示出)可以通过硬件、软件、固件、中间件、微代码、硬件描述语言或者其任意组合来实施。当用软件、固件、中间件或微代码来实施时,用以实施必要任务的程序代码或代码段可以被存储在机器或计算机可读介质(比如存储介质)中。(一个或多个)处理器可以实施必要的任务。
这里所公开的具体结构和功能细节仅仅是代表性的,并且是用于描述本发明的示例性实施例的目的。但是本发明可以通过许多替换形式来具体实现,并且不应当被解释成仅仅受限于这里所阐述的实施例。
应当理解的是,虽然在这里可能使用了术语“第一”、“第二”等等 来描述各个单元,但是这些单元不应当受这些术语限制。使用这些术语仅仅是为了将一个单元与另一个单元进行区分。举例来说,在不背离示例性实施例的范围的情况下,第一单元可以被称为第二单元,并且类似地第二单元可以被称为第一单元。这里所使用的术语“和/或”包括其中一个或更多所列出的相关联项目的任意和所有组合。
应当理解的是,当一个单元被称为“连接”或“耦合”到另一单元时,其可以直接连接或耦合到所述另一单元,或者可以存在中间单元。与此相对,当一个单元被称为“直接连接”或“直接耦合”到另一单元时,则不存在中间单元。应当按照类似的方式来解释被用于描述单元之间的关系的其他词语(例如“处于...之间”相比于“直接处于...之间”,“与...邻近”相比于“与...直接邻近”等等)。
这里所使用的术语仅仅是为了描述具体实施例而不意图限制示例性实施例。除非上下文明确地另有所指,否则这里所使用的单数形式“一个”、“一项”还意图包括复数。还应当理解的是,这里所使用的术语“包括”和/或“包含”规定所陈述的特征、整数、步骤、操作、单元和/或组件的存在,而不排除存在或添加一个或更多其他特征、整数、步骤、操作、单元、组件和/或其组合。
还应当提到的是,在一些替换实现方式中,所提到的功能/动作可以按照不同于附图中标示的顺序发生。举例来说,取决于所涉及的功能/动作,相继示出的两幅图实际上可以基本上同时执行或者有时可以按照相反的顺序来执行。
本发明可由计算机设备实现。典型地,本发明可由网络设备实现,但本领域技术人员应能理解,本发明的方案同样可由用户设备实现,只要其具备本发明所要求的计算/处理能力。为便于说明,本说明书中以下多用网络设备的实现进行举例,但本领域技术人员应能理解,此等举例仅用于解释说明本发明之目的,而不应被理解为对本发明的任何限制。
下面结合附图对本发明作进一步详细描述。
图1示出根据本发明一个实施例的方法流程图,其中具体示出一种内容推荐过程。
如图1所示,在步骤S1中,网络设备接收用户提交的语音信息;在步骤S2中,网络设备根据所述语音信息,提取所述用户的声音特征信息;在步骤S3中,网络设备根据所述声音特征信息,查询预置的第一语音资料库,以获得相匹配的第一语音资料;在步骤S4中,网络设备将所述第一语音资料推荐给所述用户。
具体地,本发明典型地适用于搜索场景或音乐交互类场景。
其中,推荐给用户的第一语音资料可以是任何通过声音特征信息来进行匹配的语音资料,包括但不限于歌曲、视频片段。
在此,本领域技术人员应能理解,“语音资料”的表述并不意在限定其表现形式,如各种音频资料,而是可以包括任何具有音频内容的资料,因此其不仅可以包括歌曲,也可以包括视频片段。
例如,当用户进行语音搜索时,网络设备根据用户的声音特征信息为其推荐第一语音资料,例如,适合该用户模仿的视频片段。此时,网络设备还需判断用户是否具有娱乐需求,如对用户提交的语音内容进行识别,以分析语音传递的内容是否具备视频、音乐等相关娱乐需求,例如通过预置的关键词表来识别娱乐需求,具体如该语音内容是否为一句歌词、一段台词等。
又如,用户当前使用音乐类APP,当用户在该APP内提交一段语音信息时,网络设备可以根据该用户的声音特征信息为其推荐第一语音资料,例如,适合该用户唱的歌曲。
优选地,本发明的内容推荐方案可以具体表现为一“匹配”功能,该“匹配”功能可以为用户提供与其声音特征信息相匹配的第一语音资料。
例如,用户端可以向用户呈现一“匹配”功能按钮,如放置于语音搜索框的右侧或设置于音乐类APP的界面中的特定位置。当用户点击该 “匹配”功能按钮后,用户端即将用户提交的语音信息发送至网络设备,以由网络设备根据该用户的声音特征信息为其推荐相应的第一语音资料。
可替代地,在用户提交语音信息之后,“匹配”功能可以被直接提示给用户,以提醒用户可以对之前提交的语音信息做语音资料推荐。
在此,本领域技术人员应能理解,“匹配”功能按钮仅用于举例说明本发明之目的,其在实际应用中的表现形式可以不同,例如可以呈现为“主打歌”功能按钮、“大明星”功能按钮等,只要此等功能按钮意在为用户提供如本发明一般的内容推荐方案,其即应落入本发明的专利保护范围之内。
此外,本领域技术人员还应能理解,本发明中的“网络设备”并不限于一台计算机设备,其可以是多台计算机设备,典型地如多台服务器,这些服务器相互配合来实现本发明的内容推荐方案。
例如,接口服务器负责与用户端以及其他服务器的交互,其中,与用户端的交互如接收用户端发送的语音信息以及向用户端返回推荐的第一语音资料,与匹配服务器的交互如将语音信息提交至匹配服务器并从匹配服务器获得匹配的第一语音资料;匹配服务器负责对接口服务器提交的语音信息在第一语音资料库中进行匹配查询,以获得相匹配的第一语音资料。这些服务器相对于其他外部设备而言,被视为一个整体,并被通常整体视为“网络设备”。
仍返回图1,在步骤S1中,网络设备接收用户提交的语音信息。
在此,本发明需要网络设备与用户端进行交互来实现。其中,用户端诸如各种在用户设备中运行的功能实体,具体如PC中的网页、手机中的APP等。
用户通过其用户设备提交一段语音信息,用户设备将该语音信息发送至网络设备,从而网络设备接收该用户提交的语音信息。
在步骤S2中,网络设备根据用户提交的语音信息,提取该用户的声音特征信息。
在此,用户的声音特征信息可以至少从以下两个维度来定义:
1)感知维度
感知维度意在从对声音的感知特征来进行定义。声音特征信息在该维度下具体诸如音调特征信息、音高特征信息、旋律特征信息、节奏特征信息、音频特征信息、语速特征信息等。
2)声学维度
声学维度意在从声音在声学角度定义的各种特征来进行表达。声音特征信息在该维度下具体诸如能量特征信息、过零率、LPC(线性预测系数)参数信息等。其中,过零率意指声音信号的符号变化的比率,如声音信号从正数变成负数或反向。
其中,网络设备可以采用任何已知或将来可行的技术对上述声音特征信息进行提取,本发明对此不作限定。
在步骤S3中,网络设备根据用户的声音特征信息,查询预置的第一语音资料库,以获得相匹配的第一语音资料。
在此,网络设备可以预先建立第一语音资料库,其中存储有各语音资料的各项声音特征信息。可替代地,第一语音资料库也可以由其他设备来建立及维护,网络设备对其具有访问权限即可。
在一个示例性的应用场景中,第一语音资料库中例如存储有各明星(包括歌手和演员)的歌曲或参演视频片段,并且这些语音资料均已标注有各项声音特征信息。网络设备根据用户所识别出的声音特征信息在该第一语音资料库中进行匹配查询,以获得相匹配的歌曲或视频片段。
在步骤S4中,网络设备将所匹配的第一语音资料推荐给用户。
在此,网络设备将所匹配的如歌曲或视频片段等推荐给用户,从而本发明可以为用户提供其适合演唱或模仿的歌曲或视频配音,以为用户定制自己的“主打歌”或“成名歌”并帮助用户找到其适合模仿的歌手。这在商业应用中是有实际意义的,并显著提升了用户粘性及使用时长。
本发明通过引导用户通过说话、大声喊、唱两句、读两句诗等各类语音输入来自动匹配出歌手、诗人、名人等。进一步地,本发明还可以 继续接收用户对推荐歌曲的演唱,进而对用户自己的演唱相对与原唱来进行打分,并通过分享到好友、邀请好友一起来PK等提升趣味性帮助用户找到自己最适合唱的歌曲,打造用户自己的模仿成名曲,从而逐步引导用户形成使用语音搜索习惯以及与搜索引擎的二次交互,提供用户粘性及使用时长。据此,本发明还可以通过提供分享功能来从搜索引擎衍生至社交领域。
根据本发明的一个优选实施例,图1所示的内容推荐过程还包括一步骤S5。在步骤S5中,网络设备根据用户语音信息的内容,查询预置的第二语音资料库,以获得相匹配的第二语音资料;随后,网络设备将所匹配的第二语音资料推荐给用户。
在此,网络设备对第二语音资料的推荐步骤可以与上述步骤S4的推荐合并,从而步骤S5发生在步骤S4之前,网络设备可以在步骤S4中向用户一并推荐第一语音资料和第二语音资料。
可替代地,网络设备对第二语音资料的推荐也可以独立于步骤S4中对第一语音资料的推荐。例如,网络设备在执行步骤S3和S4的操作时,可以在这两个步骤之前、之后或与这两个步骤同时来执行步骤S5对第二语音资料的获取以及步骤S6对第二语音资料的推荐。
在此,第二语音资料库中至少存储有语音资料的内容信息或内容摘要信息。据此,网络设备在识别出用户的语音信息中的内容后,在该第二语音资料库中进行匹配查询,以获得内容相同的第二语音资料,如内容相同但声音不同的语音资料,具体如使用不用方言来讲的同一内容。
例如,用户提交的语音信息为普通话的“你好”,网络设备识别出其中的内容信息为“你好”并在第二语音资料库中进行匹配查询,以获得各种方言的“你好”,诸如上海话、广东话的“你好”等。
优选地,第一语音资料库可以与第二语音资料库集成在一起。也即,本发明中仅使用一个语音资料库来进行第一语音资料和第二语音资料的匹配查询,从而该语音资料库中需至少存储有各语音资料的各项声音特征信息和内容信息。
图2示出根据本发明一个实施例的装置示意图,其中具体示出一种内容推荐装置。如图2所示,内容推荐装置20装置于网络设备200中,并具体包括接收装置21、提取装置22、匹配装置23和推荐装置24。
其中,接收装置21接收用户提交的语音信息;提取装置22根据所述语音信息,提取所述用户的声音特征信息;匹配装置23根据所述声音特征信息,查询预置的第一语音资料库,以获得相匹配的第一语音资料;推荐装置24将所述第一语音资料推荐给所述用户。
具体地,本发明典型地适用于搜索场景或音乐交互类场景。
其中,推荐给用户的第一语音资料可以是任何通过声音特征信息来进行匹配的语音资料,包括但不限于歌曲、视频片段。
在此,本领域技术人员应能理解,“语音资料”的表述并不意在限定其表现形式,如各种音频资料,而是可以包括任何具有音频内容的资料,因此其不仅可以包括歌曲,也可以包括视频片段。
例如,当用户进行语音搜索时,内容推荐装置20根据用户的声音特征信息为其推荐第一语音资料,例如,适合该用户模仿的视频片段。此时,内容推荐装置20或网络设备中的其他装置还需判断用户是否具有娱乐需求,如对用户提交的语音内容进行识别,以分析语音传递的内容是否具备视频、音乐等相关娱乐需求,例如通过预置的关键词表来识别娱乐需求,具体如该语音内容是否为一句歌词、一段台词等
又如,用户当前使用音乐类APP,当用户在该APP内提交一段语音信息时,内容推荐装置20可以根据该用户的声音特征信息为其推荐第一语音资料,例如,适合该用户唱的歌曲。
优选地,本发明的内容推荐方案可以具体表现为一“匹配”功能,该“匹配”功能可以为用户提供与其声音特征信息相匹配的第一语音资料。
例如,用户端可以向用户呈现一“匹配”功能按钮,如放置于语音搜索框的右侧或设置于音乐类APP的界面中的特定位置。当用户点击该“匹配”功能按钮后,用户端即将用户提交的语音信息发送至网络设备,以由内容推荐装置20根据该用户的声音特征信息为其推荐相应的第一 语音资料。
可替代地,在用户提交语音信息之后,“匹配”功能可以被直接提示给用户,以提醒用户可以对之前提交的语音信息做语音资料推荐。
在此,本领域技术人员应能理解,“匹配”功能按钮仅用于举例说明本发明之目的,其在实际应用中的表现形式可以不同,例如可以呈现为“主打歌”功能按钮、“大明星”功能按钮等,只要此等功能按钮意在为用户提供如本发明一般的内容推荐方案,其即应落入本发明的专利保护范围之内。
此外,本领域技术人员还应能理解,本发明中的“网络设备”并不限于一台计算机设备,其可以是多台计算机设备,典型地如多台服务器,这些服务器相互配合来实现本发明的内容推荐方案。
例如,接口服务器负责与用户端以及其他服务器的交互,其中,与用户端的交互如接收用户端发送的语音信息以及向用户端返回推荐的第一语音资料,与匹配服务器的交互如将语音信息提交至匹配服务器并从匹配服务器获得匹配的第一语音资料;匹配服务器负责对接口服务器提交的语音信息在第一语音资料库中进行匹配查询,以获得相匹配的第一语音资料。这些服务器相对于其他外部设备而言,被视为一个整体,并被通常整体视为“网络设备”。
仍返回图2,接收装置21接收用户提交的语音信息。
在此,本发明需要网络设备与用户端进行交互来实现。其中,用户端诸如各种在用户设备中运行的功能实体,具体如PC中的网页、手机中的APP等。
用户通过其用户设备提交一段语音信息,用户设备将该语音信息发送至网络设备,从而接收装置21接收该用户提交的语音信息。
随后,提取装置22根据用户提交的语音信息,提取该用户的声音特征信息。
在此,用户的声音特征信息可以至少从以下两个维度来定义:
1)感知维度
感知维度意在从对声音的感知特征来进行定义。声音特征信息在该维度下具体诸如音调特征信息、音高特征信息、旋律特征信息、节奏特征信息、音频特征信息、语速特征信息等。
2)声学维度
声学维度意在从声音在声学角度定义的各种特征来进行表达。声音特征信息在该维度下具体诸如能量特征信息、过零率、LPC(线性预测系数)参数信息等。其中,过零率意指声音信号的符号变化的比率,如声音信号从正数变成负数或反向。
其中,提取装置22可以采用任何已知或将来可行的技术对上述声音特征信息进行提取,本发明对此不作限定。
接着,匹配装置23根据用户的声音特征信息,查询预置的第一语音资料库,以获得相匹配的第一语音资料。
在此,网络设备可以预先建立第一语音资料库,其中存储有各语音资料的各项声音特征信息。可替代地,第一语音资料库也可以由其他设备来建立及维护,匹配装置23对其具有访问权限即可。
在一个示例性的应用场景中,第一语音资料库中例如存储有各明星(包括歌手和演员)的歌曲或参演视频片段,并且这些语音资料均已标注有各项声音特征信息。匹配装置23根据用户所识别出的声音特征信息在该第一语音资料库中进行匹配查询,以获得相匹配的歌曲或视频片段。
随后,推荐装置24将所匹配的第一语音资料推荐给用户。
在此,推荐装置24将所匹配的如歌曲或视频片段等推荐给用户,从而本发明可以为用户提供其适合演唱或模仿的歌曲或视频配音,以为用户定制自己的“主打歌”或“成名歌”并帮助用户找到其适合模仿的歌手。。这在商业应用中是有实际意义的,并显著提升了用户粘性及使用时长。
本发明通过引导用户通过说话、大声喊、唱两句、读两句诗等各类语音输入来自动匹配出歌手、诗人、名人等。进一步地,本发明还可以继续接收用户对推荐歌曲的演唱,进而对用户自己的演唱相对与原唱来 进行打分,并通过分享到好友、邀请好友一起来PK等提升趣味性帮助用户找到自己最适合唱的歌曲,打造用户自己的模仿成名曲,从而逐步引导用户形成使用语音搜索习惯以及与搜索引擎的二次交互,提供用户粘性及使用时长。据此,本发明还可以通过提供分享功能来从搜索引擎衍生至社交领域。
根据本发明的一个优选实施例,图2所示的内容推荐装置20还包括另一匹配装置(以下称为第二匹配装置)。第二匹配装置根据用户语音信息的内容,查询预置的第二语音资料库,以获得相匹配的第二语音资料;随后,内容推荐装置20将所匹配的第二语音资料推荐给用户。
在此,内容推荐装置20对第二语音资料的推荐可以与上述对第一语音资料的推荐合并,从而第二匹配装置所执行的操作发生在推荐装置24所执行的操作之前,推荐装置24可以向用户一并推荐第一语音资料和第二语音资料。
可替代地,内容推荐装置20对第二语音资料的推荐也可以独立于推荐装置24对第一语音资料的推荐。例如,在匹配装置23和推荐装置24执行其各自的操作时,第二匹配装置可以在这两个装置的操作之前、之后或与这两个装置执行操作的同时,来执行对第二语音资料的获取,从而内容推荐装置20还需包括另一推荐装置(以下称为第二推荐装置)。来执行对第二语音资料的推荐。
优选地,第二匹配装置仍可以与匹配装置23集成在一起,以及第二推荐装置与推荐装置24集成在一起。此时,集成后的各装置分别执行两次匹配操作或两次推荐操作。
在此,第二语音资料库中至少存储有语音资料的内容信息或内容摘要信息。据此,在第二匹配装置或内容推荐装置20中的其他装置识别出用户的语音信息中的内容后,第二匹配装置在该第二语音资料库中进行匹配查询,以获得内容相同的第二语音资料,如内容相同但声音不同的语音资料,具体如使用不用方言来讲的同一内容。
例如,用户提交的语音信息为普通话的“你好”,第二匹配装置在第 二语音资料库中进行匹配查询,以获得各种方言的“你好”,诸如上海话、广东话的“你好”等。
优选地,第一语音资料库可以与第二语音资料库集成在一起。也即,本发明中仅使用一个语音资料库来进行第一语音资料和第二语音资料的匹配查询,从而该语音资料库中需至少存储有各语音资料的各项声音特征信息和内容信息。
需要注意的是,本发明可在软件和/或软件与硬件的组合体中被实施,例如,本发明的各个装置可采用专用集成电路(ASIC)或任何其他类似硬件设备来实现。在一个实施例中,本发明的软件程序可以通过处理器执行以实现上文所述步骤或功能。同样地,本发明的软件程序(包括相关的数据结构)可以被存储到计算机可读记录介质中,例如,RAM存储器,磁或光驱动器或软磁盘及类似设备。另外,本发明的一些步骤或功能可采用硬件来实现,例如,作为与处理器配合从而执行各个步骤或功能的电路。
对于本领域技术人员而言,显然本发明不限于上述示范性实施例的细节,而且在不背离本发明的精神或基本特征的情况下,能够以其他的具体形式实现本发明。因此,无论从哪一点来看,均应将实施例看作是示范性的,而且是非限制性的,本发明的范围由所附权利要求而不是上述说明限定,因此旨在将落在权利要求的等同要件的含义和范围内的所有变化涵括在本发明内。不应将权利要求中的任何附图标记视为限制所涉及的权利要求。此外,显然“包括”一词不排除其他单元或步骤,单数不排除复数。系统权利要求中陈述的多个单元或装置也可以由一个单元或装置通过软件或者硬件来实现。第一,第二等词语用来表示名称,而并不表示任何特定的顺序。

Claims (15)

  1. 一种内容推荐方法,其中,该方法包括以下步骤:
    -接收用户提交的语音信息;
    -根据所述语音信息,提取所述用户的声音特征信息;
    -根据所述声音特征信息,查询预置的第一语音资料库,以获得相匹配的第一语音资料;
    -将所述第一语音资料推荐给所述用户。
  2. 根据权利要求1所述的方法,其中,该方法还包括:
    -根据所述语音信息的内容,查询预置的第二语音资料库,以获得相匹配的第二语音资料;
    其中,所述推荐包括对所述第二语音资料的推荐。
  3. 根据权利要求2所述的方法,其中,所述第一语音资料库与所述第二语音资料库集成在一起。
  4. 根据权利要求1至3中任一项所述的方法,其中,所述声音特征信息基于以下至少任一维度来提取:
    -感知维度;
    -声学维度。
  5. 根据权利要求1至4中任一项所述的方法,其中,所述第一语音资料包括以下至少任一项:
    -歌曲;
    -视频片段。
  6. 根据权利要求1至5中任一项所述的方法,其中,所述接收用户提交的步骤用于搜索场景或音乐交互类场景。
  7. 一种内容推荐装置,其中,该装置包括:
    -用于接收用户提交的语音信息的装置;
    -用于根据所述语音信息,提取所述用户的声音特征信息的装置;
    -用于根据所述声音特征信息,查询预置的第一语音资料库,以获 得相匹配的第一语音资料的装置;
    -用于将所述第一语音资料推荐给所述用户的装置。
  8. 根据权利要求7所述的装置,其中,该装置还包括:
    -用于根据所述语音信息的内容,查询预置的第二语音资料库,以获得相匹配的第二语音资料的装置;
    其中,所述推荐包括对所述第二语音资料的推荐。
  9. 根据权利要求8所述的装置,其中,所述第一语音资料库与所述第二语音资料库集成在一起。
  10. 根据权利要求7至9中任一项所述的装置,其中,所述声音特征信息基于以下至少任一维度来提取:
    -感知维度;
    -声学维度。
  11. 根据权利要求7至10中任一项所述的装置,其中,所述第一语音资料包括以下至少任一项:
    -歌曲;
    -视频片段。
  12. 根据权利要求7至11中任一项所述的装置,其中,所述接收用户提交的操作用于搜索场景或音乐交互类场景。
  13. 一种计算机可读存储介质,所述计算机可读存储介质包括计算机指令,当所述计算机指令被执行时,如权利要求1至6中任一项所述的方法被执行。
  14. 一种计算机程序产品,当所述计算机程序产品被执行时,如权利要求1至6中任一项所述的方法被执行。
  15. 一种计算机设备,所述计算机设备包括存储器和处理器,所述存储器中存储有计算机指令,所述处理器被配置来通过执行所述计算机指令以执行如权利要求1至6中任一项所述的方法。
PCT/CN2017/079624 2016-11-22 2017-04-06 一种内容推荐方法与装置 Ceased WO2018094952A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201611036906.4 2016-11-22
CN201611036906.4A CN108090081A (zh) 2016-11-22 2016-11-22 一种内容推荐方法与装置

Publications (1)

Publication Number Publication Date
WO2018094952A1 true WO2018094952A1 (zh) 2018-05-31

Family

ID=62169766

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2017/079624 Ceased WO2018094952A1 (zh) 2016-11-22 2017-04-06 一种内容推荐方法与装置

Country Status (2)

Country Link
CN (1) CN108090081A (zh)
WO (1) WO2018094952A1 (zh)

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113486208A (zh) * 2021-06-09 2021-10-08 安徽沐峰数据科技有限公司 一种基于人工智能的语音搜索设备及其搜索方法
CN114697759A (zh) * 2022-04-25 2022-07-01 中国平安人寿保险股份有限公司 虚拟形象视频生成方法及其系统、电子设备、存储介质
CN116662494A (zh) * 2023-04-26 2023-08-29 华南师范大学 一种基于人工智能的辅助教学方法及装置

Families Citing this family (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109657162A (zh) * 2018-12-17 2019-04-19 掌阅科技股份有限公司 阅读地图的生成方法、电子设备及计算机存储介质
CN110083772A (zh) * 2019-04-29 2019-08-02 北京小唱科技有限公司 基于演唱技巧的歌手推荐方法及装置
CN110688586B (zh) * 2019-09-30 2023-05-09 上海掌门科技有限公司 一种为用户推荐社交活动或好友的方法与设备

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN102347060A (zh) * 2010-08-04 2012-02-08 鸿富锦精密工业(深圳)有限公司 电子记录装置及方法
CN102404278A (zh) * 2010-09-08 2012-04-04 盛乐信息技术(上海)有限公司 一种基于声纹识别的点歌系统及其应用方法
CN103685520A (zh) * 2013-12-13 2014-03-26 深圳Tcl新技术有限公司 基于语音识别的歌曲推送的方法和装置

Family Cites Families (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2014026603A (ja) * 2012-07-30 2014-02-06 Hitachi Ltd 音楽選択支援システム、音楽選択支援方法、および音楽選択支援プログラム
CN102833582B (zh) * 2012-08-02 2015-06-17 四川长虹电器股份有限公司 采用语音搜索音视频资源的方法
CN102880693A (zh) * 2012-09-20 2013-01-16 浙江大学 一种基于个体发声能力的音乐推荐方法
CN104657438A (zh) * 2015-02-02 2015-05-27 联想(北京)有限公司 信息处理方法及电子设备
CN105095406A (zh) * 2015-07-09 2015-11-25 百度在线网络技术(北京)有限公司 一种基于用户特征的语音搜索方法及装置
CN105070283B (zh) * 2015-08-27 2019-07-09 百度在线网络技术(北京)有限公司 为歌声语音配乐的方法和装置
CN105575393A (zh) * 2015-12-02 2016-05-11 中国传媒大学 一种基于人声音色的个性化点唱歌曲推荐方法
CN106095925B (zh) * 2016-06-12 2018-07-03 北京邮电大学 一种基于声乐特征的个性化歌曲推荐方法

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN102347060A (zh) * 2010-08-04 2012-02-08 鸿富锦精密工业(深圳)有限公司 电子记录装置及方法
CN102404278A (zh) * 2010-09-08 2012-04-04 盛乐信息技术(上海)有限公司 一种基于声纹识别的点歌系统及其应用方法
CN103685520A (zh) * 2013-12-13 2014-03-26 深圳Tcl新技术有限公司 基于语音识别的歌曲推送的方法和装置

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113486208A (zh) * 2021-06-09 2021-10-08 安徽沐峰数据科技有限公司 一种基于人工智能的语音搜索设备及其搜索方法
CN114697759A (zh) * 2022-04-25 2022-07-01 中国平安人寿保险股份有限公司 虚拟形象视频生成方法及其系统、电子设备、存储介质
CN114697759B (zh) * 2022-04-25 2024-04-09 中国平安人寿保险股份有限公司 虚拟形象视频生成方法及其系统、电子设备、存储介质
CN116662494A (zh) * 2023-04-26 2023-08-29 华南师范大学 一种基于人工智能的辅助教学方法及装置

Also Published As

Publication number Publication date
CN108090081A (zh) 2018-05-29

Similar Documents

Publication Publication Date Title
CN115082602B (zh) 生成数字人的方法、模型的训练方法、装置、设备和介质
CN111667811B (zh) 语音合成方法、装置、设备和介质
JP6896690B2 (ja) マルチメディアコンテンツにおける文脈探索
JP6505903B2 (ja) 会話型相互作用システムの検索入力におけるユーザ意図を推定する方法およびそのためのシステム
US11017010B2 (en) Intelligent playing method and apparatus based on preference feedback
CN110995569B (zh) 一种智能互动方法、装置、计算机设备和存储介质
US10572602B2 (en) Building conversational understanding systems using a toolset
US10586541B2 (en) Communicating metadata that identifies a current speaker
US8972265B1 (en) Multiple voices in audio content
CN110825835B (zh) 从先前会话检索情境
CN109165302B (zh) 多媒体文件推荐方法及装置
US20150179170A1 (en) Discriminative Policy Training for Dialog Systems
US9190052B2 (en) Systems and methods for providing information discovery and retrieval
US20140236570A1 (en) Exploiting the semantic web for unsupervised spoken language understanding
US20140379323A1 (en) Active learning using different knowledge sources
CN107526809B (zh) 基于人工智能推送音乐的方法和装置
CN107221323A (zh) 语音点歌方法、终端及存储介质
CN110765270B (zh) 用于口语交互的文本分类模型的训练方法及系统
US12511338B2 (en) Conference information query method and apparatus, storage medium, terminal device, and server
CN112799630A (zh) 使用网络可寻址设备创建电影化的讲故事体验
CN111324626B (zh) 基于语音识别的搜索方法、装置、计算机设备及存储介质
WO2022224584A1 (ja) 情報処理装置、情報処理方法、端末装置及び表示方法
CN107608799B (zh) 一种用于执行交互指令的方法、设备及存储介质
CN108255798A (zh) 一种拉泰赫格式公式的输入方法及其装置
JP2022163217A (ja) 映像コンテンツに対する合成音のリアルタイム生成を基盤としたコンテンツ編集支援方法およびシステム

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 17873702

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 17873702

Country of ref document: EP

Kind code of ref document: A1