WO2025251807A1 - 生成音乐的方法、装置、设备和存储介质 - Google Patents
生成音乐的方法、装置、设备和存储介质Info
- Publication number
- WO2025251807A1 WO2025251807A1 PCT/CN2025/091595 CN2025091595W WO2025251807A1 WO 2025251807 A1 WO2025251807 A1 WO 2025251807A1 CN 2025091595 W CN2025091595 W CN 2025091595W WO 2025251807 A1 WO2025251807 A1 WO 2025251807A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- lyrics
- content
- lyric
- editing
- editing control
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H1/00—Details of electrophonic musical instruments
- G10H1/0008—Associated control or indicating means
- G10H1/0025—Automatic or semi-automatic music composition, e.g. producing random music, applying rules from music theory or modifying a musical piece
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H2220/00—Input/output interfacing specifically adapted for electrophonic musical tools or instruments
- G10H2220/005—Non-interactive screen display of musical or status data
- G10H2220/011—Lyrics displays, e.g. for karaoke applications
Definitions
- the exemplary embodiments disclosed herein generally relate to the Internet field, and particularly to methods, apparatus, devices, storage media, and products for generating music.
- a method for generating music includes: presenting an editing interface including a lyrics editing control; in response to selection of first lyrics content in the lyrics editing control, presenting a set of candidate lyrics content, the set of candidate lyrics content being generated based on context information associated with the lyrics editing control; determining target lyrics based on selection of second lyrics content from the set of candidate lyrics content; and in response to receiving a music generation request, providing music content generated based on the target lyrics.
- an apparatus for generating music includes: an interface presentation module configured to present an editing interface including a lyrics editing control; a lyrics presentation module configured to present a set of candidate lyrics content in response to selection of first lyrics content in the lyrics editing control, the set of candidate lyrics content being generated based on context information associated with the lyrics editing control; a lyrics determination module configured to determine target lyrics based on selection of a second lyrics content from the set of candidate lyrics content; and a music providing module configured to provide music content generated based on the target lyrics in response to receiving a music generation request.
- an electronic device in a third aspect of this disclosure, includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.
- a computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.
- a computer program product includes a computer program executable by a processing unit, the computer program including instructions for performing the method of the first aspect.
- Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure may be implemented
- FIGS. 2A to 2E illustrate example interfaces according to some embodiments of the present disclosure
- Figure 3 illustrates a flowchart of an example process for generating music according to some embodiments of the present disclosure
- Figure 4 shows a schematic structural block diagram of an example apparatus for generating music according to some embodiments of the present disclosure.
- Figure 5 shows a block diagram of an electronic device capable of implementing several embodiments of the present disclosure.
- the term “comprising” and similar terms should be understood as open-ended inclusion, i.e., “including but not limited to”.
- the term “based on” should be understood as “at least partially based on”.
- the term “one embodiment” or “the embodiment” should be understood as “at least one embodiment”.
- the term “some embodiments” should be understood as “at least some embodiments”.
- Other explicit and implicit definitions may also be included below.
- the terms “first”, “second”, etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
- the embodiments of this disclosure may involve user data, data acquisition, and/or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and/or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
- any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon.
- a user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
- Embodiments of this disclosure propose a scheme for generating music.
- an editing interface is presented, including a lyrics editing control; in response to the selection of a first lyric content in the lyrics editing control, a set of candidate lyric content is presented, the set of candidate lyric content being generated based on context information associated with the lyrics editing control; based on the selection of a second lyric content from the set of candidate lyric content, a target lyric is determined; and in response to receiving a music generation request, music content generated based on the target lyric is provided.
- embodiments of the present disclosure can help users edit lyrics based on contextual information, thereby improving editing efficiency and the quality of the input lyrics, and thus improving the quality of the generated music content.
- Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented.
- the example environment 100 may include an electronic device 110.
- electronic device 110 may run an application 120 that supports user interface interaction.
- Application 120 may be any suitable type of application for user interface interaction, examples of which may include, but are not limited to, music applications or other suitable applications.
- User 140 may interact with application 120 via electronic device 110 and/or its attached devices.
- electronic device 110 can use application 120 to present interface 150 for supporting interface interaction.
- electronic device 110 communicates with server 130 to provide services to application 120.
- Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR/AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio/video players, digital cameras/camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof.
- electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).
- Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms.
- Server 130 may include, for example, computing systems/servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc.
- Server 130 can provide backend services for applications 120 that support virtual scenarios in electronic devices 110.
- a communication connection can be established between server 130 and electronic device 110.
- This communication connection can be established via wired or wireless means.
- the communication connection may include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections; the embodiments of this disclosure are not limited in this respect.
- server 130 and electronic device 110 can achieve signaling interaction through the communication connection between them.
- FIGS 2A to 2E illustrate example interfaces 200A to 200E according to some embodiments of the present disclosure. Interfaces 200A to 200E may be provided, for example, by the electronic device 110 shown in Figure 1.
- Figure 2A illustrates an editing interface 200A according to some embodiments of the present disclosure.
- the editing interface 200A may include one or more controls for editing lyric creation parameters.
- the editing interface 200A may include a lyrics editing control 202.
- This lyrics editing control 202 can support users directly inputting lyrics, or it can generate lyrics by inputting prompts.
- the electronic device 110 can obtain the prompt item 204, such as a text prompt item, via the lyrics editing control 202.
- the prompt item 204 may be a text prompt item corresponding to a preset template.
- the electronic device 110 can display the text prompt item corresponding to the preset template in the lyrics editing control 202.
- the electronic device 110 can also support users to quickly edit one or more items in the prompt item 204. Specifically, as shown in FIG2A, the electronic device 110 can, for example, provide an editing entry corresponding to the content "History” and "Ancient Style”. After receiving an editing entry selection for content 206 (e.g., "History"), the electronic device 110 can provide a set of candidate content 208 in association, such as "Family” and "Friendship”. Further, the electronic device 110 can, for example, replace the content 206 in the prompt item 204 based on the user's selection of the candidate content 208.
- embodiments of this disclosure can provide users with templated prompts, thereby improving the quality of prompts entered by the user and increasing editing efficiency.
- the user can edit the prompt item 204 by inserting, modifying, or deleting it.
- the electronic device 110 can also acquire text prompt items directly entered by the user in the lyrics editing control 202.
- electronic device 110 may, for example, trigger the generation of lyrics corresponding to prompt 204.
- electronic device 110 may, for instance, provide prompt 204 to a language model to obtain the corresponding lyrics. It should be understood that such a language model may be implemented based on appropriate machine learning model techniques, and this disclosure is not intended to limit it.
- the electronic device 110 may display lyrics 220 in the lyrics editing control 202, for example.
- the lyrics 220 may be generated based on the prompt 204 mentioned above.
- the lyrics 220 may also be lyrics directly entered by the user using the lyrics editing control 202.
- the electronic device 110 can also receive preset operations from the user on the completed musical content (also known as the target musical content) and obtain the user's secondary creation request. Accordingly, the electronic device 110 can present the editing interface 200B as shown in FIG2B based on the preset operations.
- the lyrics 220 displayed by the lyrics editing control 202 can be, for example, the lyrics of the selected target musical content, to support the user to further edit the lyrics.
- one or more of the attribute editing controls 212, 214, and 216 in the editing interface 200B can display the attributes corresponding to the target music content by default, such as style, timbre, and name.
- the electronic device 110 can receive the user's selection of the first lyric content 222 in the lyrics 220 (e.g., "the obsession in the eyes"), and can accordingly present a set of candidate lyric content 224.
- the first lyric content 222 in the lyrics 220 e.g., "the obsession in the eyes”
- the set of candidate lyrics 224 may be generated based on context information associated with the lyrics editing control 202.
- context information may include, for example, existing lyrics in the lyrics editing control 202, so that the provided candidate lyrics 224 can be adapted to existing lyrics.
- such context information may include at least one lyric parameter.
- such lyric parameters may include rhyme parameters, theme parameters, etc.
- such lyric parameters may be automatically determined based on existing lyric content.
- the electronic device 110 may also provide controls for editing lyric parameters in the lyric editing control 202, such as controls 228 and 230.
- the set of candidate lyrics 224 provided by the electronic device 110 can match the lyrics parameters associated with the lyrics editing control 202, such as having matching rhymes and matching themes.
- the electronic device 110 may receive, for example, the user's selection of a second lyric content (e.g., "longing in the eyes") for the group of candidate lyric content 224, and may use the second lyric content to replace the first lyric content 222 in the lyrics 220, thereby completing the rapid editing of the lyrics.
- a second lyric content e.g., "longing in the eyes”
- embodiments of this disclosure enable users to efficiently optimize lyrics, thereby improving the efficiency of lyrics editing. Furthermore, the provided candidate lyrics can match contextual information, allowing users to easily create high-quality lyrics.
- the lyrics editing control 202 can also receive user selections of control 228 or control 230 and update the lyrics 220 accordingly. For example, when the user changes a new rhyme through control 228, the electronic device 110 can provide updated lyrics matching the new rhyme. Alternatively, when the user specifies a new theme through control 230, the electronic device 110 can provide updated lyrics matching the new theme.
- the lyrics editing control 202 may further include a continuation control 226. Further, upon receiving a selection of the continuation control 226, the electronic device 110 may provide additional lyrics content created based on existing lyrics 220 within the lyrics editing control 202.
- such additional lyrics content can be generated based on context information associated with the lyrics editing control.
- context information may be the same as or different from the context information used to generate candidate lyrics content 224.
- a user can input several lines of lyrics through the lyrics editing control 202, and the electronic device 110 can use a language model to automatically continue writing the subsequent lyrics based on the existing lyrics.
- the continued lyrics can, for example, match the rhymes and/or themes of the existing lyrics.
- the electronic device 110 may also provide a segmentation control 232 in the lyrics editing control 202. Upon receiving a trigger for the segmentation control 232, the electronic device 110 may accordingly present a set of indicator elements associated with the lyrics 220 in the lyrics editing control based on the received segmentation request, such as indicator element 234-1 and indicator element 234-2 (referred to individually or collectively as indicator element 234).
- such an indicator element 234 can be used to indicate the structural type of a corresponding lyric segment in the lyrics 220, such as verse, chorus, etc. In some embodiments, such an indicator element 234 can characterize the structural information of the lyrics 220. As an example, such structural information can be provided for the generation of corresponding musical content.
- the user can edit the indicator element 234. For example, the user can insert a new indicator element at a specific position in the lyrics to indicate the structural type of the corresponding lyric segment. Alternatively, the user can delete one or more existing indicator elements 234. Or, the user can modify the type of the indicator element 234 to, for example, change the corresponding lyric segment from a first structural type to a second structural type.
- the electronic device 110 may also support user-editable indicator elements. Specifically, when the electronic device 110 receives a preset identifier 236 (e.g., "/") input by the user in the lyrics editing control 202, the electronic device 110 may provide a set of candidate indicator elements 238 accordingly.
- a preset identifier 236 e.g., "/”
- the electronic device 110 may provide a set of candidate indicator elements 238 accordingly.
- the electronic device 110 can receive the user's selection of the target indicator element in the group of candidate indicator elements 238, and can insert the target indicator element into the target position corresponding to the preset identifier 236 to indicate the structural type of the corresponding lyrics segment.
- embodiments of the present disclosure can more effectively edit the structural information of lyrics, thereby further improving the quality of the generated musical content.
- the electronic device 110 can obtain the final lyrics in the lyrics editing control 202 and the parameter information input in the other attribute editing controls 212 to 216, and can trigger the generation process of the corresponding music content based on the user's selection of the input control 218.
- electronic device 110 can provide such creation parameters to a music generation model to generate corresponding music content.
- creation parameters may include, but are not limited to, the lyrics, musical style, timbre, song title, and lyrical structure information mentioned above.
- the music generation model can utilize appropriate existing or future available machine learning models, and this disclosure is not intended to limit it.
- the electronic device 110 can correspondingly present the interface 200E as shown in FIG2E.
- the interface 200E may include a first area 240 for displaying creation parameters, a second area 250 for displaying creation history, and a third area 260 for displaying descriptive information of the currently playing music content.
- the electronic device 110 can present the content item 255 corresponding to the generated music content in the second area 250, and can trigger the display of descriptive information (e.g., cover art, title, lyrics, etc.) of the music content in the third area 260 based on the selection of the content item 255.
- the electronic device 110 can also trigger the playback of the music content accordingly, and can provide a playback control 270 to control the playback process of the music content.
- the embodiments of this disclosure can improve the efficiency of music creation and enhance the quality of the generated music content.
- FIG 3 shows a flowchart of an example process 300 for generating music according to some embodiments of the present disclosure.
- Process 300 can be implemented at electronic device 110.
- Process 300 will now be described with reference to Figure 1.
- the electronic device 110 presents an editing interface, which includes lyrics editing controls.
- the electronic device 110 in response to the selection of a first lyric content in the lyrics editing control, the electronic device 110 presents a set of candidate lyric content, which is generated based on contextual information associated with the lyrics editing control.
- electronic device 110 determines the target lyrics based on the selection of a second lyric from a set of candidate lyric contents.
- electronic device 110 in response to receiving a music generation request, provides music content generated based on the target lyrics.
- the context information includes: at least one lyric parameter, which includes a rhyme parameter and/or a theme parameter; and/or existing lyric content in the lyric editing control.
- the lyrics editing control includes: a first configuration control for configuring rhyme parameters; and/or a second configuration control for configuring theme parameters.
- the method further includes: in response to obtaining an updated rhyme parameter via a first configuration control, providing updated lyrics content corresponding to the updated rhyme parameter in a lyrics editing control; or in response to obtaining an updated theme parameter via a second configuration control, providing updated lyrics content corresponding to the updated theme parameter in a lyrics editing control.
- the context information is first context information
- the method further includes: in response to a received continuation request, presenting additional lyrics content in the lyrics editing control, the additional lyrics content being generated based on second context information associated with the lyrics editing control.
- the method further includes: in response to a received segmentation request, presenting a set of indicator elements associated with existing lyrics content in a lyrics editing control, the set of indicator elements being used to characterize the structural type of the corresponding lyrics segment.
- the music content is also generated based on the structural information of the target lyrics, which characterizes the structural type of a set of lyric segments of the target lyrics.
- the method further includes: modifying the structural information of existing lyrics content based on editing operations on a set of indicator elements.
- the editing operation includes at least one of the following: modifying the type of an existing indicator element to change the corresponding lyric segment from a first structure type to a second structure type; adding a new indicator element to indicate the structure type of the corresponding lyric segment; and deleting at least one indicator element.
- adding a new indicator element includes: presenting a set of candidate indicator elements in response to inputting a preset identifier in the lyrics editing control; and inserting a target indicator element at a target position in the lyrics editing control based on the selection of a target indicator element from the set of candidate indicator elements.
- the first lyrics content is a portion of reference lyrics
- the method further includes: presenting a text prompt in a lyrics editing control; and providing reference lyrics generated based on the text prompt in the lyrics editing control based on a received lyrics generation request.
- presenting a text prompt in the lyrics editing control includes: presenting a text prompt corresponding to a preset template in the lyrics editing control; and presenting a set of candidate contents for replacing the target portion based on the selection of target content in the text prompt.
- presenting the editing interface includes: presenting the editing interface based on a preset operation on the target music content; and presenting the lyrics of the target music content in the lyrics editing control of the editing interface.
- the editing interface also includes a property editing control for configuring at least one property of the music content to be generated.
- FIG. 4 shows a schematic structural block diagram of an example apparatus 400 for generating music according to certain embodiments of this disclosure.
- Apparatus 400 may be implemented as or included in electronic device 110.
- the various modules/components in apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.
- the device 400 includes an interface presentation module 410 configured to present an editing interface, which includes a lyrics editing control; a lyrics presentation module 420 configured to present a set of candidate lyrics content in response to the selection of a first lyric content in the lyrics editing control, the set of candidate lyrics content being generated based on context information associated with the lyrics editing control; a lyrics determination module 430 configured to determine a target lyric based on the selection of a second lyric content from the set of candidate lyric content; and a music providing module 440 configured to provide music content generated based on the target lyric in response to receiving a music generation request.
- an interface presentation module 410 configured to present an editing interface, which includes a lyrics editing control
- a lyrics presentation module 420 configured to present a set of candidate lyrics content in response to the selection of a first lyric content in the lyrics editing control, the set of candidate lyrics content being generated based on context information associated with the lyrics editing control
- a lyrics determination module 430 configured to determine a target lyric based on the selection
- the context information includes: at least one lyric parameter, which includes a rhyme parameter and/or a theme parameter; and/or existing lyric content in the lyric editing control.
- the lyrics editing control includes: a first configuration control for configuring rhyme parameters; and/or a second configuration control for configuring theme parameters.
- the music providing module 440 is further configured to: in response to obtaining the updated rhyme parameter via the first configuration control, provide updated lyrics content corresponding to the updated rhyme parameter in the lyrics editing control; or in response to obtaining the updated theme parameter via the second configuration control, provide updated lyrics content corresponding to the updated theme parameter in the lyrics editing control.
- the context information is first context information
- the method further includes: in response to a received continuation request, presenting additional lyrics content in the lyrics editing control, the additional lyrics content being generated based on second context information associated with the lyrics editing control.
- the interface presentation module 410 is further configured to: in response to a received segmentation request, present a set of indicator elements associated with existing lyrics content in the lyrics editing control, wherein the set of indicator elements is used to characterize the structural type of the corresponding lyrics segment.
- the music content is also generated based on the structural information of the target lyrics, which characterizes the structural type of a set of lyric segments of the target lyrics.
- the lyrics determination module 430 is further configured to modify the structural information of existing lyrics content based on editing operations on a set of indicator elements.
- the editing operation includes at least one of the following: modifying the type of an existing indicator element to change the corresponding lyric segment from a first structure type to a second structure type; adding a new indicator element to indicate the structure type of the corresponding lyric segment; and deleting at least one indicator element.
- adding a new indicator element includes: presenting a set of candidate indicator elements in response to inputting a preset identifier in the lyrics editing control; and inserting a target indicator element at a target position in the lyrics editing control based on the selection of a target indicator element from the set of candidate indicator elements.
- the first lyrics content is a portion of reference lyrics
- the method further includes: presenting a text prompt in a lyrics editing control; and providing reference lyrics generated based on the text prompt in the lyrics editing control based on a received lyrics generation request.
- presenting a text prompt in the lyrics editing control includes: presenting a text prompt corresponding to a preset template in the lyrics editing control; and presenting a set of candidate contents for replacing the target portion based on the selection of target content in the text prompt.
- presenting the editing interface includes: presenting the editing interface based on a preset operation on the target music content; and presenting the lyrics of the target music content in the lyrics editing control of the editing interface.
- the editing interface also includes a property editing control for configuring at least one property of the music content to be generated.
- Figure 5 shows a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 500 shown in Figure 5 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The electronic device 500 shown in Figure 5 can be used to implement the electronic device 110 of Figure 1.
- the electronic device 500 is in the form of a general-purpose electronic device.
- Components of the electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 730, one or more communication units 540, one or more input devices 550, and one or more output devices 560.
- the processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to the program stored in memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 500.
- Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media.
- Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof.
- Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and/or data and can be accessed within electronic device 500.
- Electronic device 500 may further include additional removable/non-removable, volatile/non-volatile storage media.
- disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided.
- each drive may be connected to a bus (not shown) via one or more data media interfaces.
- Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
- Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
- PCs network personal computers
- Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc.
- Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc.
- Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input/output (I/O) interface (not shown).
- I/O input/output
- a computer-readable storage medium that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above.
- a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
- These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions/actions specified in one or more blocks of the flowchart and/or block diagram.
- These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and/or other device to operate in a particular manner.
- the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions/actions specified in one or more blocks of the flowchart and/or block diagram.
- Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions/actions specified in one or more boxes of a flowchart and/or block diagram.
- each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function.
- the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved.
- each block in the block diagrams and/or flowcharts, and combinations of blocks in the block diagrams and/or flowcharts may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- User Interface Of Digital Computer (AREA)
Abstract
一种生成音乐的方法、装置、设备和存储介质。该方法包括:呈现编辑界面,编辑界面包括歌词编辑控件(310);响应于对歌词编辑控件中的第一歌词内容的选择,呈现一组候选歌词内容,一组候选歌词内容基于与歌词编辑控件相关联的上下文信息所生成(320);基于对一组候选歌词内容中的第二歌词内容的选择,确定目标歌词(330);以及响应于接收到音乐生成请求,提供基于目标歌词生成的音乐内容(340)。该方法能够提高在音乐生成过程中的歌词编辑效率,从而提高音乐生成的效率。
Description
本申请要求2024年06月04日递交的、标题为“生成音乐的方法、装置、设备和存储介质”、申请号为202410718366.6的中国发明专利申请的优先权,该申请的全部内容通过引用结合在本申请中。
本公开的示例实施例总体涉及互联网领域,特别地涉及生成音乐的方法、装置、设备、存储介质和产品。
近年来,随着互联网的快速发展,生成式人工智能技术已经被逐渐应用于各种类型的媒体内容的创作。与视频内容和图片内容等媒体内容相比,音乐内容的生成质量对歌词的依赖性较强,这需要人们具有较强的乐理知识。
在本公开的第一方面,提供了一种生成音乐的方法。该方法包括:呈现编辑界面,编辑界面包括歌词编辑控件;响应于对歌词编辑控件中的第一歌词内容的选择,呈现一组候选歌词内容,一组候选歌词内容基于与歌词编辑控件相关联的上下文信息所生成;基于对一组候选歌词内容中的第二歌词内容的选择,确定目标歌词;以及响应于接收到音乐生成请求,提供基于目标歌词生成的音乐内容。
在本公开的第二方面,提供了一种用于生成音乐的装置。该装置包括:界面呈现模块,被配置为呈现编辑界面,编辑界面包括歌词编辑控件;歌词呈现模块,被配置为响应于对歌词编辑控件中的第一歌词内容的选择,呈现一组候选歌词内容,一组候选歌词内容基于与歌词编辑控件相关联的上下文信息所生成;歌词确定模块,被配置为基于对一组候选歌词内容中的第二歌词内容的选择,确定目标歌词;以及音乐提供模块,被配置为响应于接收到音乐生成请求,提供基于目标歌词生成的音乐内容。
在本公开的第三方面,提供了一种电子设备。该设备包括至少一个处理单元;以及至少一个存储器,至少一个存储器被耦合到至少一个处理单元并且存储用于由至少一个处理单元执行的指令。指令在由至少一个处理单元执行时使设备执行第一方面的方法。
在本公开的第四方面,提供了一种计算机可读存储介质。该计算机可读存储介质上存储有计算机程序,计算机程序可由处理器执行以实现第一方面的方法。
在本公开的第五方面,提供了一种计算机程序产品。该一种计算机程序产品包括可由处理单元执行的计算机程序,计算机程序包括用于执行第一方面的方法的指令。
应当理解,本内容部分中所描述的内容并非旨在限定本公开的实施例的关键特征或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的描述而变得容易理解。
结合附图并参考以下详细说明,本公开各实施例的上述和其他特征、优点及方面将变得更加明显。在附图中,相同或相似的附图标记表示相同或相似的元素,其中:
图1示出了其中可以实施根据本公开的实施例的示例环境的示意图;
图2A至图2E示出了根据本公开的一些实施例的示例界面;
图3示出了根据本公开的一些实施例的生成音乐的示例过程的流程图;
图4示出了根据本公开的一些实施例的用于生成音乐的示例装置的示意性结构框图;以及
图5示出了能够实施本公开的多个实施例的电子设备的框图。
下面将参照附图更详细地描述本公开的实施例。虽然附图中示出了本公开的某些实施例,然而应当理解的是,本公开可以通过各种形式来实现,而且不应该被解释为限于这里阐述的实施例,相反,提供这些实施例是为了更加透彻和完整地理解本公开。应当理解的是,本公开的附图及实施例仅用于示例性作用,并非用于限制本公开的保护范围。
需要注意的是,本文中所提供的任何节/子节的标题并不是限制性的。本文通篇描述了各种实施例,并且任何类型的实施例都可以包括在任何节/子节下。此外,在任一节/子节中描述的实施例可以以任何方式与同一节/子节和/或不同节/子节中描述的任何其他实施例相结合。
在本公开的实施例的描述中,术语“包括”及其类似用语应当理解为开放性包含,即“包括但不限于”。术语“基于”应当理解为“至少部分地基于”。术语“一个实施例”或“该实施例”应当理解为“至少一个实施例”。术语“一些实施例”应当理解为“至少一些实施例”。下文还可能包括其他明确的和隐含的定义。术语“第一”、“第二”等可以指代不同的或相同的对象。下文还可能包括其他明确的和隐含的定义。
本公开的实施例中可能涉及用户的数据、数据的获取和/或使用等。这些方面均遵循相应的法律法规及相关规定。在本公开的实施例中,所有数据的采集、获取、处理、加工、转发、使用等,都是在用户知晓并且确认的前提下进行的。相应地,在实现本公开的各实施例时,均应根据相关法律法规通过适当的方式,将可能所涉及的数据或信息的类型、使用范围、使用场景等告知用户并获得用户的授权。具体的告知和/或授权方式可以根据实际情况和应用场景而变化,本公开的范围在此方面不受限制。
本说明书及实施例中方案,如涉及个人信息处理,则均会在具备合法性基础(例如征得个人信息主体同意,或者为履行合同所必需等)的前提下进行处理,且仅会在规定或者约定的范围内进行处理。用户拒绝处理基本功能所需必要信息以外的个人信息,不会影响用户使用基本功能。
如上文所提及的,与视频内容和图片内容等媒体内容相比,音乐内容的生成质量对歌词的依赖性较强,这需要人们具有较强的乐理知识。由于歌词创作的专业性,普通用户难以提供高质量的歌词输入,这极大地影响了所生成的音乐内容的质量。
本公开的实施例提出了一种生成音乐的方案。根据该方案,呈现编辑界面,编辑界面包括歌词编辑控件;响应于对歌词编辑控件中的第一歌词内容的选择,呈现一组候选歌词内容,一组候选歌词内容基于与歌词编辑控件相关联的上下文信息所生成;基于对一组候选歌词内容中的第二歌词内容的选择,确定目标歌词;以及响应于接收到音乐生成请求,提供基于目标歌词生成的音乐内容。
以此方式,本公开的实施例能够基于上下文信息来帮助用户编辑歌词内容,从而提高编辑效率和所输入歌词的质量,进而提高所生成的音乐内容的质量。
以下进一步结合附图来详细描述该方案的各种示例实现。
示例环境
图1示出了本公开的实施例能够在其中实现的示例环境100的示意图。如图1所示,示例环境100可以包括电子设备110。
在该示例环境100中,电子设备110可以运行有支持界面交互的应用120。应用120可以是用于界面交互的任何适当类型应用,其示例可以包括但不限于:音乐应用或其它适当的应用。用户140可以经由电子设备110和/或其附接设备来与应用120进行交互。
在图1的环境100中,如果应用120处于活动状态,电子设备110可以通过应用120呈现用于支持界面交互的界面150。
在一些实施例中,电子设备110与服务器130通信,以实现对应用120的服务的供应。电子设备110可以是任意类型的移动终端、固定终端或便携式终端,包括移动手机、台式计算机、膝上型计算机、笔记本计算机、上网本计算机、平板计算机、媒体计算机、多媒体平板、掌上电脑、便携式游戏终端、VR/AR设备、个人通信系统(Personal Communication System,PCS)设备、个人导航设备、个人数字助理(Personal Digital Assistant,PDA)、音频/视频播放器、数码相机/摄像机、定位设备、电视接收器、无线电广播接收器、电子书设备、游戏设备或者前述各项的任意组合,包括这些设备的配件和外设或者其任意组合。在一些实施例中,电子设备110也能够支持任意类型的针对用户的接口(诸如“可佩戴”电路等)。
服务器130可以是独立的物理服务器,也可以是多个物理服务器构成的服务器集群或者分布式系统,还可以是提供云服务、云数据库、云计算、云函数、云存储、网络服务、云通信、中间件服务、域名服务、安全服务、内容分发网络、以及大数据和人工智能平台等基础云计算服务的云服务器。服务器130例如可以包括计算系统/服务器,诸如大型机、边缘计算节点、云环境中的计算设备,等等。服务器130可以为电子设备110中支持虚拟场景的应用120提供后台服务。
服务器130与电子设备110之间可以建立有通信连接。通信连接可以通过有线方式或无线方式建立。通信连接可以包括但不限于蓝牙连接、移动网络连接、通用串行总线(Universal Serial Bus,USB)连接、无线保真(Wireless Fidelity,WiFi)连接等,本公开的实施例在此方面不受限制。在本公开的实施例中,服务器130与电子设备110可以通过二者之间的通信连接实现信令交互。
应当理解,仅出于示例性的目的描述环境100中各个元素的结构和功能,而不暗示对于本公开的范围的任何限制。
以下将继续参考附图描述本公开的一些示例实施例。
示例过程
图2A至图2E示出了根据本公开的一些实施例的示例界面200A至界面200E。界面200A至界面200E例如可以由图1所示的电子设备110所提供。
图2A示出了根据本公开的一些实施例的编辑界面200A。如图2A所示,该编辑界面200A可以包括用于编辑歌词创作参数的一个或多个控件。
具体地,编辑界面200A可以包括歌词编辑控件202。该歌词编辑控件202可以支持用户直接输入歌词内容,或者还可以通过输入提示项来生成歌词内容。
以图2A作为示例,电子设备110可以经由歌词编辑控件202来获取提示项204,例如,文本提示项。在一些实施例中,该提示项204可以是与预设模板所对应的文本提示项。例如,在接收到用户的预设请求时,电子设备110可以在歌词编辑控件202中显示与预设模板所对应的文本提示项。
在一些实施例中,电子设备110还可以支持用户快速地编辑与提示项204中的一项或多项内容。具体地,如图2A所示,电子设备110例如可以提供与内容“历史”和“古风”对应的编辑入口。在接收到针对内容206(例如,“历史”)的编辑入口选择后,电子设备110可以关联地提供一组候选内容208,例如,“亲情”和“友情”。进一步地,电子设备110例如可以基于用户对候选内容208的选择来替换提示项204中的内容206。
基于这样的方式,本公开的实施例能够为用户提供模板化的提示项,从而提高用户所输入的提示项的质量,并提高编辑效率。
在一些实施例中,用户例如还可以通过插入、修改、删除等方式来对提示项204进行编辑。在又一些实施例中,电子设备110例如还可以获取用户直接在歌词编辑控件202中输入的文本提示项。
进一步地,在接收到对于按钮210的选择后,电子设备110例如可以触发生成与提示项204所对应的歌词。作为示例,电子设备110例如可以向语言模型提供提示项204,以获取对应的歌词。应当理解的是,这样的语言模型例如可以基于适当的机器学习模型技术来实现,本公开不旨在对此进行限定。
进一步地,如图2B所示,电子设备110例如可以在歌词编辑控件202中展示歌词220。作为示例,歌词220例如看可以是基于上文所提及的提示项204所生成。作为另一示例,歌词220例如还可以是用户利用歌词编辑控件202所直接输入的歌词。
作为又一示例,电子设备110例如还可以接收用户对于已完成创作的音乐内容(也称为目标音乐内容)的预设操作,并获取用户的二次创作请求。相应地,电子设备110可以基于该预设操作来呈现如图2B所示的编辑界面200B。在这种情况下,歌词编辑控件202所显示的歌词220例如可以是所选择的目标音乐内容的歌词,以支持用户进一步对歌词进行编辑。
在基于目标音乐内容的二次创作场景下,编辑界面200B中的属性编辑控件212、214、216中的一项或多项属性编辑控件可以默认地显示目标音乐内容所对应的属性,例如,曲风、音色、名称等。
继续参考图2B,电子设备110可以接收用户对歌词220中的第一歌词内容222(例如,“眸中的执念”)的选择,并可以相应地呈现一组候选歌词内容224。
在一些实施例中,该组候选歌词内容224可以是基于与歌词编辑控件202相关联的上下文信息所生成。在一些实施例中,这样的上下文信息例如可以包括歌词编辑控件202中已有的歌词内容,以使得所提供候选歌词内容224能够与已有的歌词内容适配。
在一些实施例中,这样的上下文信息可以包括至少一项歌词参数。在一些实施例中,这样的歌词参数可以包括韵脚参数、主题参数等。在一些实施例中,这样的歌词参数可以是基于已有的歌词内容所自动确定。备选地,电子设备110还可以在歌词编辑控件202中提供用于编辑歌词参数的控件,例如,控件228和控件230。
以此方式,电子设备110所提供的该组候选歌词内容224能够匹配与歌词编辑控件202相关联的歌词参数,例如,具有匹配的韵脚和匹配的主题等。
在一些实施例中,电子设备110例如可以接收用户对于该组候选歌词内容224的第二歌词内容(例如,“眼中的思念”)的选择,并可以利用第二歌词内容替换歌词220中的第一歌词内容222,从而完成歌词的快速编辑。
以此方式,本公开的实施例能够支持用户高效地优化歌词的内容,从而提高歌词编辑的效率。此外,所提供的候选歌词内容能够匹配上下文信息,这使得用户可以简单地创作出高质量的歌词。
在一些实施例中,歌词编辑控件202还可以接收用户对控件228或控件230的选择,并可以相应地更新歌词220。例如,在用户经由控件228更换了新的韵脚时,电子设备110可以提供与新的韵脚匹配的更新歌词。或者,在用户经由控件230指定了新的主题时,电子设备110可以提供与新的主题匹配的更新歌词。
在一些实施例中,歌词编辑控件202还可以包括续写控件226。进一步地,在接收到对续写控件226的选择后,电子设备110可以在歌词编辑控件202中提供基于已有的歌词220所创作的附加歌词内容。
在一些实施例中,这样的附加歌词内容可以基于与歌词编辑控件相关联的上下文信息所生成。这样的上下文信息例如可以与用于生成候选歌词内容224的上下文信息相同或不同。
例如,用户可以通过歌词编辑控件202输入几句歌词,电子设备110可以利用语言模型来基于已有的歌词内容来自动地续写出后续的歌词。为了保证歌词的匹配性,续写的歌词内容例如可以与已有歌词的韵脚和/或主题匹配。
在一些实施例中,如图2C所示,电子设备110还可以在歌词编辑控件202中提供分段控件232。在接收到对于分段控件232的触发后,电子设备110可以基于接收到的分段请求来相应地在歌词编辑控件中呈现与歌词220相关联的一组指示元素,例如,指示元素234-1和指示元素234-2(单独或统一称为指示元素234)。
在一些实施例中,这样的指示元素234可以用于指示歌词220中的对应歌词分段的结构类型,例如,主歌部分、副歌部分等等。在一些实施例中,这样的指示元素234可以表征歌词220的结构信息。作为示例,这样的结构信息可以被提供以用于对应音乐内容的生成。
在一些实施例中,用户例可以编辑指示元素234。例如,用户可以在歌词的特定位置中插入新的指示元素,以指示对应歌词分段的结构类型。或者,用户还可以删除已有的一个或多个指示元素234。或者,用户也可以修改指示元素234的类型,以例如将对应的歌词分段从第一结构类型修改为第二结构类型。
一些实施例中,如图2D所示,电子设备110还可以支持用户自由编辑指示元素。具体地,电子设备110在接收到用户在歌词编辑控件202中输入的预设标识符236(例如,“/”)时,电子设备110可以相应地提供一组候选指示元素238。
进一步地,电子设备110可以接收用户对该组候选指示元素238中的目标指示元素的选择,并可以将该目标指示元素插入到预设标识符236对应的目标位置处,以指示对应的歌词分段的结构类型。
以此方式,本公开的实施例能够更为有效地编辑歌词的结构信息,从而进一步提高所生成的音乐内容的质量。
进一步地,如图2B至图2D所示,电子设备110可以获取歌词编辑控件202中最终的歌词和其他属性编辑控件212至216中所输入的参数信息,并可以基于用户对输入控件218的选择,来触发对应音乐内容的生成过程。
作为示例,电子设备110可以向音乐生成模型提供这样的创作参数,以生成对应的音乐内容。这样的创作参数可以包括但不限于上文所提及的歌词、曲风、音色、歌曲名称、歌词的结构信息等。此外,音乐生成模型可以利用已有或未来可用的适当机器学习模型来实现,本公开不旨在对此进行限定。
示例性地,在音乐内容生成后,电子设备110可以相应地呈现如图2E所示的界面200E。如图所示,界面200E可以包括用于显示创作参数的第一区域240、用于显示创作历史的第二区域250和用于显示当前播放的音乐内容的描述信息的第三区域260。
作为示例,电子设备110可以在第二区域250中呈现所生成的音乐内容所对应的内容项255,并可以基于对内容项255的选择来触发在第三区域260中显示该音乐内容的描述信息(例如,封面、标题、歌词等)。此外,电子设备110还可以相应地触发音乐内容的播放,并可以提供播放控件270,以控制音乐内容的播放过程。
以此方式,本公开的实施例能够提高音乐创作的效率,并改善所生成的音乐内容的质量。
示例过程
图3示出了根据本公开的一些实施例的生成音乐的示例过程300的流程图。过程300可以被实现在电子设备110处。下面参考图1来描述过程300。
如图3所示,在框310,电子设备110呈现编辑界面,编辑界面包括歌词编辑控件。
在框320,电子设备110响应于对歌词编辑控件中的第一歌词内容的选择,呈现一组候选歌词内容,一组候选歌词内容基于与歌词编辑控件相关联的上下文信息所生成。
在框330,电子设备110基于对一组候选歌词内容中的第二歌词内容的选择,确定目标歌词。
在框340,电子设备110响应于接收到音乐生成请求,提供基于所述目标歌词生成的音乐内容。
在一些实施例中,上下文信息包括:至少一项歌词参数,至少一项歌词参数包括韵脚参数和/或主题参数;和/或歌词编辑控件中的已有歌词内容。
在一些实施例中,歌词编辑控件包括:用于配置韵脚参数的第一配置控件;和/或用于配置主题参数的第二配置控件。
在一些实施例中,该方法还包括:响应于经由第一配置控件获取更新韵脚参数,在歌词编辑控件中提供与更新韵脚参数对应的更新歌词内容;或响应于经由第二配置控件获取更新主题参数,在歌词编辑控件中提供与更新主题参数对应的更新歌词内容。
在一些实施例中,上下文信息为第一上下文信息,方法还包括:响应于接收到的续写请求,在歌词编辑控件中呈现附加歌词内容,附加歌词内容是基于与歌词编辑控件相关联的第二上下文信息所生成。
在一些实施例中,该方法还包括:响应于接收到的分段请求,在歌词编辑控件中呈现与已有歌词内容相关联的一组指示元素,一组指示元素用于表征对应歌词分段的结构类型。
在一些实施例中,音乐内容还基于目标歌词的结构信息所生成,结构信息表征目标歌词的一组歌词分段的结构类型。
在一些实施例中,该方法还包括:基于对一组指示元素的编辑操作,修改已有歌词内容的结构信息。
在一些实施例中,编辑操作包括以下至少一项:修改已有的指示元素的类型,以将对应的歌词分段从第一结构类型修改为第二结构类型;增加新的指示元素,以指示对应的歌词分段的结构类型;删除至少一个指示元素。
在一些实施例中,增加新的指示元素包括:响应于在歌词编辑控件中输入预设标识符,呈现一组候选指示元素;以及基于对一组候选指示元素中的目标指示元素的选择,在歌词编辑控件中的目标位置处插入目标指示元素。
在一些实施例中,第一歌词内容为参考歌词的一部分,方法还包括:在歌词编辑控件中呈现文本提示项;以及基于接收到的歌词生成请求,在歌词编辑控件中提供基于文本提示项所生成的参考歌词。
在一些实施例中,在歌词编辑控件中呈现文本提示项包括:在歌词编辑控件中呈现与预设模板对应的文本提示项;以及基于对文本提示项中的目标内容的选择,呈现用于替换目标部分的一组候选内容。
在一些实施例中,呈现编辑界面包括:基于对目标音乐内容的预设操作,呈现编辑界面;以及在编辑界面的歌词编辑控件中呈现目标音乐内容的歌词。
在一些实施例中,编辑界面还包括用于配置待生成的音乐内容的至少一项属性的属性编辑控件。
示例装置和设备
本公开的实施例还提供了用于实现上述方法或过程的相应装置。图4示出了根据本公开的某些实施例的用于生成音乐的示例装置400的示意性结构框图。装置400可以被实现为或者被包括在电子设备110中。装置400中的各个模块/组件可以由硬件、软件、固件或者它们的任意组合来实现。
如图4所示,装置400包括界面呈现模块410,被配置为呈现编辑界面,编辑界面包括歌词编辑控件;歌词呈现模块420,被配置为响应于对歌词编辑控件中的第一歌词内容的选择,呈现一组候选歌词内容,一组候选歌词内容基于与歌词编辑控件相关联的上下文信息所生成;歌词确定模块430,被配置为基于对一组候选歌词内容中的第二歌词内容的选择,确定目标歌词;以及音乐提供模块440,被配置为响应于接收到音乐生成请求,提供基于目标歌词生成的音乐内容。
在一些实施例中,上下文信息包括:至少一项歌词参数,至少一项歌词参数包括韵脚参数和/或主题参数;和/或歌词编辑控件中的已有歌词内容。
在一些实施例中,歌词编辑控件包括:用于配置韵脚参数的第一配置控件;和/或用于配置主题参数的第二配置控件。
在一些实施例中,音乐提供模块440还被配置为:响应于经由第一配置控件获取更新韵脚参数,在歌词编辑控件中提供与更新韵脚参数对应的更新歌词内容;或响应于经由第二配置控件获取更新主题参数,在歌词编辑控件中提供与更新主题参数对应的更新歌词内容。
在一些实施例中,上下文信息为第一上下文信息,方法还包括:响应于接收到的续写请求,在歌词编辑控件中呈现附加歌词内容,附加歌词内容是基于与歌词编辑控件相关联的第二上下文信息所生成。
在一些实施例中,界面呈现模块410还被配置为:响应于接收到的分段请求,在歌词编辑控件中呈现与已有歌词内容相关联的一组指示元素,一组指示元素用于表征对应歌词分段的结构类型。
在一些实施例中,音乐内容还基于目标歌词的结构信息所生成,结构信息表征目标歌词的一组歌词分段的结构类型。
在一些实施例中,歌词确定模块430还被配置为:基于对一组指示元素的编辑操作,修改已有歌词内容的结构信息。
在一些实施例中,编辑操作包括以下至少一项:修改已有的指示元素的类型,以将对应的歌词分段从第一结构类型修改为第二结构类型;增加新的指示元素,以指示对应的歌词分段的结构类型;删除至少一个指示元素。
在一些实施例中,增加新的指示元素包括:响应于在歌词编辑控件中输入预设标识符,呈现一组候选指示元素;以及基于对一组候选指示元素中的目标指示元素的选择,在歌词编辑控件中的目标位置处插入目标指示元素。
在一些实施例中,第一歌词内容为参考歌词的一部分,方法还包括:在歌词编辑控件中呈现文本提示项;以及基于接收到的歌词生成请求,在歌词编辑控件中提供基于文本提示项所生成的参考歌词。
在一些实施例中,在歌词编辑控件中呈现文本提示项包括:在歌词编辑控件中呈现与预设模板对应的文本提示项;以及基于对文本提示项中的目标内容的选择,呈现用于替换目标部分的一组候选内容。
在一些实施例中,呈现编辑界面包括:基于对目标音乐内容的预设操作,呈现编辑界面;以及在编辑界面的歌词编辑控件中呈现目标音乐内容的歌词。
在一些实施例中,编辑界面还包括用于配置待生成的音乐内容的至少一项属性的属性编辑控件。
图5示出了其中可以实施本公开的一个或多个实施例的电子设备500的框图。应当理解,图5所示出的电子设备500仅仅是示例性的,而不应当构成对本文所描述的实施例的功能和范围的任何限制。图5所示出的电子设备500可以用于实现图1的电子设备110。
如图5所示,电子设备500是通用电子设备的形式。电子设备500的组件可以包括但不限于一个或多个处理器或处理单元510、存储器520、存储设备730、一个或多个通信单元540、一个或多个输入设备550以及一个或多个输出设备560。处理单元510可以是实际或虚拟处理器并且能够根据存储器520中存储的程序来执行各种处理。在多处理器系统中,多个处理单元并行执行计算机可执行指令,以提高电子设备500的并行处理能力。
电子设备500通常包括多个计算机存储介质。这样的介质可以是电子设备500可访问的任何可以获取的介质,包括但不限于易失性和非易失性介质、可拆卸和不可拆卸介质。存储器520可以是易失性存储器(例如寄存器、高速缓存、随机访问存储器(RAM))、非易失性存储器(例如,只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、闪存)或它们的某种组合。存储设备530可以是可拆卸或不可拆卸的介质,并且可以包括机器可读介质,诸如闪存驱动、磁盘或者任何其他介质,其可以能够用于存储信息和/或数据并且可以在电子设备500内被访问。
电子设备500可以进一步包括另外的可拆卸/不可拆卸、易失性/非易失性存储介质。尽管未在图5中示出,可以提供用于从可拆卸、非易失性磁盘(例如“软盘”)进行读取或写入的磁盘驱动和用于从可拆卸、非易失性光盘进行读取或写入的光盘驱动。在这些情况中,每个驱动可以由一个或多个数据介质接口被连接至总线(未示出)。存储器520可以包括计算机程序产品525,其具有一个或多个程序模块,这些程序模块被配置为执行本公开的各种实施例的各种方法或动作。
通信单元540实现通过通信介质与其他电子设备进行通信。附加地,电子设备500的组件的功能可以以单个计算集群或多个计算机器来实现,这些计算机器能够通过通信连接进行通信。因此,电子设备500可以使用与一个或多个其他服务器、网络个人计算机(PC)或者另一个网络节点的逻辑连接来在联网环境中进行操作。
输入设备550可以是一个或多个输入设备,例如鼠标、键盘、追踪球等。输出设备560可以是一个或多个输出设备,例如显示器、扬声器、打印机等。电子设备500还可以根据需要通过通信单元540与一个或多个外部设备(未示出)进行通信,外部设备诸如存储设备、显示设备等,与一个或多个使得用户与电子设备500交互的设备进行通信,或者与使得电子设备500与一个或多个其他电子设备通信的任何设备(例如,网卡、调制解调器等)进行通信。这样的通信可以经由输入/输出(I/O)接口(未示出)来执行。
根据本公开的示例性实现方式,提供了一种计算机可读存储介质,其上存储有计算机可执行指令,其中计算机可执行指令被处理器执行以实现上文描述的方法。根据本公开的示例性实现方式,还提供了一种计算机程序产品,计算机程序产品被有形地存储在非瞬态计算机可读介质上并且包括计算机可执行指令,而计算机可执行指令被处理器执行以实现上文描述的方法。
这里参照根据本公开实现的方法、装置、设备和计算机程序产品的流程图和/或框图描述了本公开的各个方面。应当理解,流程图和/或框图的每个方框以及流程图和/或框图中各方框的组合,都可以由计算机可读程序指令实现。
这些计算机可读程序指令可以提供给通用计算机、专用计算机或其他可编程数据处理装置的处理单元,从而生产出一种机器,使得这些指令在通过计算机或其他可编程数据处理装置的处理单元执行时,产生了实现流程图和/或框图中的一个或多个方框中规定的功能/动作的装置。也可以把这些计算机可读程序指令存储在计算机可读存储介质中,这些指令使得计算机、可编程数据处理装置和/或其他设备以特定方式工作,从而,存储有指令的计算机可读介质则包括一个制造品,其包括实现流程图和/或框图中的一个或多个方框中规定的功能/动作的各个方面的指令。
可以把计算机可读程序指令加载到计算机、其他可编程数据处理装置、或其他设备上,使得在计算机、其他可编程数据处理装置或其他设备上执行一系列操作步骤,以产生计算机实现的过程,从而使得在计算机、其他可编程数据处理装置、或其他设备上执行的指令实现流程图和/或框图中的一个或多个方框中规定的功能/动作。
附图中的流程图和框图显示了根据本公开的多个实现的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段或指令的一部分,模块、程序段或指令的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个连续的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或动作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
以上已经描述了本公开的各实现,上述说明是示例性的,并非穷尽性的,并且也不限于所公开的各实现。在不偏离所说明的各实现的范围和精神的情况下,对于本技术领域的普通技术人员来说许多修改和变更都是显而易见的。本文中所用术语的选择,旨在最好地解释各实现的原理、实际应用或对市场中的技术的改进,或者使本技术领域的其他普通技术人员能理解本文公开的各个实现方式。
Claims (18)
- 一种生成音乐的方法,包括:呈现编辑界面,所述编辑界面包括歌词编辑控件;响应于对所述歌词编辑控件中的第一歌词内容的选择,呈现一组候选歌词内容,所述一组候选歌词内容基于与所述歌词编辑控件相关联的上下文信息所生成;基于对所述一组候选歌词内容中的第二歌词内容的选择,确定目标歌词;以及响应于接收到音乐生成请求,提供基于所述目标歌词生成的音乐内容。
- 根据权利要求1所述的方法,其中所述上下文信息包括:至少一项歌词参数,所述至少一项歌词参数包括韵脚参数和/或主题参数;和/或所述歌词编辑控件中的已有歌词内容。
- 根据权利要求2所述的方法,其中所述歌词编辑控件包括:用于配置所述韵脚参数的第一配置控件;和/或用于配置所述主题参数的第二配置控件。
- 根据权利要求3所述的方法,还包括:响应于经由所述第一配置控件获取更新韵脚参数,在所述歌词编辑控件中提供与所述更新韵脚参数对应的更新歌词内容;或响应于经由所述第二配置控件获取更新主题参数,在所述歌词编辑控件中提供与所述更新主题参数对应的更新歌词内容。
- 根据权利要求1-4中任一项所述的方法,其中所述上下文信息为第一上下文信息,所述方法还包括:响应于接收到的续写请求,在所述歌词编辑控件中呈现附加歌词内容,所述附加歌词内容是基于与所述歌词编辑控件相关联的第二上下文信息所生成。
- 根据权利要求1-4中任一项所述的方法,还包括:响应于接收到的分段请求,在所述歌词编辑控件中呈现与已有歌词内容相关联的一组指示元素,所述一组指示元素用于表征对应歌词分段的结构类型。
- 根据权利要求6所述的方法,其中所述音乐内容还基于所述目标歌词的结构信息所生成,所述结构信息表征所述目标歌词的一组歌词分段的所述结构类型。
- 根据权利要求6所述的方法,还包括:基于对所述一组指示元素的编辑操作,修改所述已有歌词内容的结构信息。
- 根据权利要求8所述的方法,其中所述编辑操作包括以下至少一项:修改已有的指示元素的类型,以将对应的歌词分段从第一结构类型修改为第二结构类型;增加新的指示元素,以指示对应的歌词分段的结构类型;删除至少一个指示元素。
- 根据权利要求9所述的方法,其中增加新的指示元素包括:响应于在所述歌词编辑控件中输入预设标识符,呈现一组候选指示元素;以及基于对所述一组候选指示元素中的目标指示元素的选择,在所述歌词编辑控件中的目标位置处插入所述目标指示元素。
- 根据权利要求1-4和7-10中任一项所述的方法,其中所述第一歌词内容为参考歌词的一部分,所述方法还包括:在所述歌词编辑控件中呈现文本提示项;以及基于接收到的歌词生成请求,在所述歌词编辑控件中提供基于所述文本提示项所生成的所述参考歌词。
- 根据权利要求11所述的方法,其中在所述歌词编辑控件中呈现文本提示项包括:在所述歌词编辑控件中呈现与预设模板对应的所述文本提示项;以及基于对所述文本提示项中的目标内容的选择,呈现用于替换所述目标部分的一组候选内容。
- 根据权利要求1-4、7-10和12中任一项所述的方法,其中呈现编辑界面包括:基于对目标音乐内容的预设操作,呈现所述编辑界面;以及在所述编辑界面的所述歌词编辑控件中呈现所述目标音乐内容的歌词。
- 根据权利要求1-4、7-10和12中任一项所述的方法,其中所述编辑界面还包括用于配置待生成的所述音乐内容的至少一项属性的属性编辑控件。
- 一种用于生成音乐的装置,包括:界面呈现模块,被配置为呈现编辑界面,所述编辑界面包括歌词编辑控件;歌词呈现模块,被配置为响应于对所述歌词编辑控件中的第一歌词内容的选择,呈现一组候选歌词内容,所述一组候选歌词内容基于与所述歌词编辑控件相关联的上下文信息所生成;歌词确定模块,被配置为基于对所述一组候选歌词内容中的第二歌词内容的选择,确定目标歌词;以及音乐提供模块,被配置为响应于接收到音乐生成请求,提供基于所述目标歌词生成的音乐内容。
- 一种电子设备,包括:至少一个处理单元;以及至少一个存储器,所述至少一个存储器被耦合到所述至少一个处理单元并且存储用于由所述至少一个处理单元执行的指令,所述指令在由所述至少一个处理单元执行时使所述电子设备执行根据权利要求1至14中任一项所述的方法。
- 一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序可由处理器执行以实现根据权利要求1至14中任一项所述的方法。
- 一种计算机程序产品,所述计算机程序产品被有形地存储在计算机存储介质中并且包括计算机可执行指令,所述计算机可执行指令在由设备执行时使所述设备执行根据权利要求1至14中任一项所述的方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202410718366.6 | 2024-06-04 | ||
| CN202410718366.6A CN118486282A (zh) | 2024-06-04 | 2024-06-04 | 生成音乐的方法、装置、设备和存储介质 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025251807A1 true WO2025251807A1 (zh) | 2025-12-11 |
Family
ID=92197174
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2025/091595 Pending WO2025251807A1 (zh) | 2024-06-04 | 2025-04-27 | 生成音乐的方法、装置、设备和存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN118486282A (zh) |
| WO (1) | WO2025251807A1 (zh) |
Families Citing this family (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN118486282A (zh) * | 2024-06-04 | 2024-08-13 | 北京字跳网络技术有限公司 | 生成音乐的方法、装置、设备和存储介质 |
| CN121640966A (zh) * | 2024-09-06 | 2026-03-10 | 北京字跳网络技术有限公司 | 用于音乐生成的方法、装置、设备和存储介质 |
| CN119207350A (zh) * | 2024-09-18 | 2024-12-27 | 北京字跳网络技术有限公司 | 音乐生成方法、装置、计算机设备及存储介质 |
| CN118897663B (zh) * | 2024-10-08 | 2025-02-28 | 北京字跳网络技术有限公司 | 音乐生成方法、电子设备、存储设备和产品 |
Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112632906A (zh) * | 2020-12-30 | 2021-04-09 | 北京达佳互联信息技术有限公司 | 歌词生成方法、装置、电子设备、计算机可读存储介质 |
| US20210312897A1 (en) * | 2018-10-11 | 2021-10-07 | WaveAI Inc. | Method and system for interactive song generation |
| US20210335334A1 (en) * | 2019-10-11 | 2021-10-28 | WaveAI Inc. | Methods and systems for interactive lyric generation |
| US20220208156A1 (en) * | 2020-12-30 | 2022-06-30 | Beijing Dajia Internet Information Technology Co., Ltd. | Method for generating song melody and electronic device |
| US20220223125A1 (en) * | 2019-06-14 | 2022-07-14 | Microsoft Technology Licensing, Llc | Song generation based on a text input |
| CN117012170A (zh) * | 2022-04-29 | 2023-11-07 | 脸萌有限公司 | 一种音乐生成方法、装置、系统及存储介质 |
| CN118486282A (zh) * | 2024-06-04 | 2024-08-13 | 北京字跳网络技术有限公司 | 生成音乐的方法、装置、设备和存储介质 |
-
2024
- 2024-06-04 CN CN202410718366.6A patent/CN118486282A/zh active Pending
-
2025
- 2025-04-27 WO PCT/CN2025/091595 patent/WO2025251807A1/zh active Pending
Patent Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20210312897A1 (en) * | 2018-10-11 | 2021-10-07 | WaveAI Inc. | Method and system for interactive song generation |
| US20220223125A1 (en) * | 2019-06-14 | 2022-07-14 | Microsoft Technology Licensing, Llc | Song generation based on a text input |
| US20210335334A1 (en) * | 2019-10-11 | 2021-10-28 | WaveAI Inc. | Methods and systems for interactive lyric generation |
| CN112632906A (zh) * | 2020-12-30 | 2021-04-09 | 北京达佳互联信息技术有限公司 | 歌词生成方法、装置、电子设备、计算机可读存储介质 |
| US20220208156A1 (en) * | 2020-12-30 | 2022-06-30 | Beijing Dajia Internet Information Technology Co., Ltd. | Method for generating song melody and electronic device |
| CN117012170A (zh) * | 2022-04-29 | 2023-11-07 | 脸萌有限公司 | 一种音乐生成方法、装置、系统及存储介质 |
| CN118486282A (zh) * | 2024-06-04 | 2024-08-13 | 北京字跳网络技术有限公司 | 生成音乐的方法、装置、设备和存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN118486282A (zh) | 2024-08-13 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2025251807A1 (zh) | 生成音乐的方法、装置、设备和存储介质 | |
| WO2025168004A1 (zh) | 创建虚拟对象的方法、装置、设备和存储介质 | |
| WO2025168002A1 (zh) | 发布虚拟对象的方法、装置、设备和存储介质 | |
| WO2025157280A1 (zh) | 交互方法、装置、设备和存储介质 | |
| WO2025256646A1 (zh) | 信息展示方法、装置、设备和存储介质 | |
| WO2026040862A1 (zh) | 创建模板的方法、装置、设备和存储介质 | |
| WO2026007522A1 (zh) | 生成媒体内容的方法、装置、设备和存储介质 | |
| WO2026012482A1 (zh) | 消息交互的方法、装置、设备和存储介质 | |
| WO2026051838A1 (zh) | 用于媒体编辑方法、装置、设备和存储介质 | |
| WO2025256532A1 (zh) | 发布内容的方法、装置、设备和存储介质 | |
| WO2026040869A1 (zh) | 用于媒体编辑的方法、装置、设备和存储介质 | |
| WO2026037307A1 (zh) | 创建虚拟对象的方法、装置、设备和存储介质 | |
| WO2026021505A1 (zh) | 处理媒体内容的方法、装置、设备和存储介质 | |
| WO2025261520A1 (zh) | 交互方法、装置、设备和存储介质 | |
| WO2025261521A1 (zh) | 生成媒体内容的方法、装置、设备和存储介质 | |
| WO2025195387A1 (zh) | 发布作品和查看作品的方法、装置、设备和存储介质 | |
| WO2025251806A1 (zh) | 生成媒体的方法、装置、设备和存储介质 | |
| WO2025176202A1 (zh) | 媒体资源生成方法、装置、设备和存储介质 | |
| WO2025194818A1 (zh) | 生成查询指令的方法、装置、设备和存储介质 | |
| CN118550625A (zh) | 生成媒体内容方法、装置、设备和存储介质 | |
| WO2025228337A1 (zh) | 生成媒体内容的方法、装置、设备和存储介质 | |
| WO2026026950A1 (zh) | 媒体编辑方法、装置、设备和存储介质 | |
| WO2026012290A1 (zh) | 创建应用的方法、装置、设备和存储介质 | |
| WO2026046036A1 (zh) | 特效编辑的方法、装置、设备和存储介质 | |
| WO2026001452A1 (zh) | 生成媒体内容的方法、装置、设备和存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25818867 Country of ref document: EP Kind code of ref document: A1 |