WO2023051651A1 - 音乐生成方法、装置、设备、存储介质及程序 - Google Patents

音乐生成方法、装置、设备、存储介质及程序 Download PDF

Info

Publication number
WO2023051651A1
WO2023051651A1 PCT/CN2022/122334 CN2022122334W WO2023051651A1 WO 2023051651 A1 WO2023051651 A1 WO 2023051651A1 CN 2022122334 W CN2022122334 W CN 2022122334W WO 2023051651 A1 WO2023051651 A1 WO 2023051651A1
Authority
WO
WIPO (PCT)
Prior art keywords
scale
target
music
action
information
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2022/122334
Other languages
English (en)
French (fr)
Inventor
于晨磊
康铭全
陈海东
白蓉
侯志文
谭宵
陈卓文
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Zitiao Network Technology Co Ltd
Original Assignee
Beijing Zitiao Network Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Zitiao Network Technology Co Ltd filed Critical Beijing Zitiao Network Technology Co Ltd
Priority to US18/571,086 priority Critical patent/US20240290305A1/en
Publication of WO2023051651A1 publication Critical patent/WO2023051651A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H1/00Details of electrophonic musical instruments
    • G10H1/0008Associated control or indicating means
    • G10H1/0025Automatic or semi-automatic music composition, e.g. producing random music, applying rules from music theory or modifying a musical piece
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H1/00Details of electrophonic musical instruments
    • G10H1/0008Associated control or indicating means
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2210/00Aspects or methods of musical processing having intrinsic musical character, i.e. involving musical theory or musical parameters or relying on musical knowledge, as applied in electrophonic musical tools or instruments
    • G10H2210/031Musical analysis, i.e. isolation, extraction or identification of musical elements or musical parameters from a raw acoustic signal or from an encoded audio signal
    • G10H2210/081Musical analysis, i.e. isolation, extraction or identification of musical elements or musical parameters from a raw acoustic signal or from an encoded audio signal for automatic key or tonality recognition, e.g. using musical rules or a knowledge base
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2210/00Aspects or methods of musical processing having intrinsic musical character, i.e. involving musical theory or musical parameters or relying on musical knowledge, as applied in electrophonic musical tools or instruments
    • G10H2210/101Music Composition or musical creation; Tools or processes therefor
    • G10H2210/111Automatic composing, i.e. using predefined musical rules
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2210/00Aspects or methods of musical processing having intrinsic musical character, i.e. involving musical theory or musical parameters or relying on musical knowledge, as applied in electrophonic musical tools or instruments
    • G10H2210/101Music Composition or musical creation; Tools or processes therefor
    • G10H2210/131Morphing, i.e. transformation of a musical piece into a new different one, e.g. remix
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2210/00Aspects or methods of musical processing having intrinsic musical character, i.e. involving musical theory or musical parameters or relying on musical knowledge, as applied in electrophonic musical tools or instruments
    • G10H2210/395Special musical scales, i.e. other than the 12-interval equally tempered scale; Special input devices therefor
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2210/00Aspects or methods of musical processing having intrinsic musical character, i.e. involving musical theory or musical parameters or relying on musical knowledge, as applied in electrophonic musical tools or instruments
    • G10H2210/555Tonality processing, involving the key in which a musical piece or melody is played
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2220/00Input/output interfacing specifically adapted for electrophonic musical tools or instruments
    • G10H2220/155User input interfaces for electrophonic musical instruments
    • G10H2220/201User input interfaces for electrophonic musical instruments for movement interpretation, i.e. capturing and recognizing a gesture or a specific kind of movement, e.g. to control a musical instrument
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2220/00Input/output interfacing specifically adapted for electrophonic musical tools or instruments
    • G10H2220/155User input interfaces for electrophonic musical instruments
    • G10H2220/441Image sensing, i.e. capturing images or optical patterns for musical purposes or musical control purposes
    • G10H2220/455Camera input, e.g. analyzing pictures from a video camera and using the analysis results as control data

Definitions

  • Embodiments of the present disclosure relate to the technical field of artificial intelligence, and in particular to a music generation method, device, device, electronic device, computer-readable storage medium, computer program product, and computer program.
  • piano keys may be displayed on a screen of a terminal device, and a user may simulate playing a piano by touching the keys on the screen, thereby realizing music creation.
  • Embodiments of the present disclosure provide a music generation method, device, device, electronic device, computer-readable storage medium, computer program product, and computer program, so as to solve the problem of poor interactivity in music creation methods.
  • an embodiment of the present disclosure provides a method for generating music, including:
  • action information of a first action performed by the first user within the preset duration according to the first video where the action information includes an action type and an action duration
  • target music corresponding to the preset duration is generated.
  • an embodiment of the present disclosure provides a music generating device, including:
  • An acquisition module configured to acquire the first video obtained by collecting the first user within a preset duration
  • a determining module configured to determine action information of a first action performed by the first user within the preset duration according to the first video, the action information including action type and action duration;
  • a generating module configured to generate target music corresponding to the preset duration according to the action information of the first action.
  • an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;
  • the memory stores computer-executable instructions
  • the processor executes the computer-executed instructions to implement the music generation method in the first aspect and various possible implementation manners of the first aspect.
  • an embodiment of the present disclosure provides a computer-readable storage medium, where computer-executable instructions are stored in the computer-readable storage medium, and when the processor executes the computer-executable instructions, the above first aspect and the first A music generation method in various possible implementations of the aspect.
  • an embodiment of the present disclosure provides a computer program product, including a computer program.
  • the computer program is executed by a processor, the music generation method in the above first aspect and various possible implementation manners of the first aspect is implemented.
  • an embodiment of the present disclosure provides a computer program, when the computer program is executed by a processor, implements the music generation method in the above first aspect and various possible implementation manners of the first aspect.
  • the music generation method, device, device, electronic device, computer-readable storage medium, computer program product and computer program provided by the embodiments of the present disclosure includes: obtaining the first user's collected data within a preset time period A video, according to the first video, determine the action information of the first action performed by the first user within the preset duration, the action information includes the action type and action duration, and then according to the first action action information to generate the target music corresponding to the preset duration.
  • FIG. 1 is a schematic diagram of an application scenario provided by an embodiment of the present disclosure
  • FIG. 2 is a schematic flow diagram of a music generation method provided by an embodiment of the present disclosure
  • FIG. 3 is a schematic flowchart of another music generation method provided by an embodiment of the present disclosure.
  • FIG. 4 is a schematic diagram of a display interface provided by an embodiment of the present disclosure.
  • FIG. 5 is a schematic diagram of another display interface provided by an embodiment of the present disclosure.
  • FIG. 6 is a schematic diagram of another display interface provided by an embodiment of the present disclosure.
  • FIG. 7 is a schematic structural diagram of a music generating device provided by an embodiment of the present disclosure.
  • FIG. 8 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure.
  • Scale Arranging the tones in the mode in a ladder form from low to high (called ascending) or from high to low (called descending) from the beginning of the tonic to the end of the tonic is called a scale.
  • the natural heptatonic scale is the most widely used heptatonic scale. Its interval organization is that there are 5 whole tones in each octave, which are divided into two strings and three strings, and the two strings are separated by semitones. .
  • the 7 scales are: 1(Do), 2(Re), 3(Mi), 4(Fa), 5(So), 6(La), 7(Si).
  • the embodiment of the present disclosure does not limit the scale system, and may be applied to any scale system.
  • the natural seven-tone scale is used as an example when referring to illustrations.
  • Timbre Different sounds always have distinctive characteristics in terms of waveforms, and different objects vibrate with different characteristics. Different sounding bodies have different timbres due to their different materials and structures. For example, the sounds produced by pianos and violins are different from those produced by people, and the sounds produced by each person are also different. Therefore, timbre can be understood as the characteristic of sound.
  • FIG. 1 is a schematic diagram of an application scenario provided by an embodiment of the present disclosure.
  • this application scenario involves terminal equipment.
  • the terminal equipment is provided with a video acquisition device such as a camera.
  • the terminal equipment is also provided with an audio playback device such as a loudspeaker.
  • the terminal device is also provided with a music generating device.
  • the music generation device may be in the form of software and/or hardware.
  • the music generation device may be a processor, a chip, a chip module, a module, a unit, an application program, etc. in a terminal device.
  • the image capture device can capture user actions (such as body movements, facial expressions, gestures, etc.) to obtain videos.
  • the image acquisition device sends the video to the music generation device.
  • the music generating means can generate music based on the video.
  • the music generation device may recognize a user's action sequence from a video, and map the user's action into a musical scale to obtain a musical scale sequence. In turn, music can be generated from sequences of scales.
  • the music generating device sends the generated music to the audio playing device, and the audio playing device plays the music.
  • the user can make the terminal device generate different music to realize music creation.
  • the user can perform a series of actions to interact with the terminal device to realize music creation, which makes the interaction and interest higher, and can improve the user experience.
  • the terminal device in the embodiments of the present disclosure may be any electronic device with a video acquisition device and an audio playback device, including but not limited to: smart phones, tablet computers, notebook computers, smart TVs, smart wearable devices, smart home devices, etc.
  • Fig. 2 is a schematic flowchart of a method for generating music provided by an embodiment of the present disclosure.
  • the method in this embodiment may be executed by a terminal device, or may be executed by a music generating apparatus in the terminal device.
  • the method of this embodiment includes:
  • S201 Acquire a first video obtained by collecting a first user within a preset time period.
  • This embodiment is applied to a scenario where the first user creates target music.
  • the preset duration represents the execution granularity of this embodiment in the time dimension.
  • the first video captured within a preset duration is acquired, and the target music corresponding to the first preset duration is generated based on the first video.
  • the preset duration may also be referred to as a time window.
  • the preset duration can be set to a small value, for example, 1 second, 100 milliseconds, and so on.
  • the preset duration may be the duration for the user to perform an action.
  • the user usually needs to repeat this embodiment for many times.
  • the corresponding preset durations may be the same or different. In this way, the target music corresponding to multiple preset durations is combined into the final complete music.
  • Method 1 The music generation process is carried out synchronously with the video capture process. That is to say, the method of this embodiment is executed synchronously during the process of capturing the video for the first user, so as to generate music based on the video while capturing the video. Specifically, with reference to Fig. 1, whenever the video capture device captures a first video with a preset duration, it sends the first video to the music generation device, and the music generation device generates the target music corresponding to the preset duration based on the first video. .
  • Mode 2 the music generation process and the video capture process are executed asynchronously.
  • the video capture device captures the user to obtain a second video, for example, the duration of the second video may be 3 minutes or 5 minutes.
  • the second video may be a complete video of the user dancing a dance.
  • the second video collected above is stored in a preset storage space.
  • the music generation device obtains the second video from the preset storage space, and takes the preset duration as the time granularity, and obtains the first video corresponding to the preset duration from the second video, according to the first The video generates target music corresponding to the preset duration.
  • S202 According to the first video, determine action information of a first action performed by the first user within the preset duration, where the action information includes an action type and an action duration.
  • the first action is an action performed by a preset body part of the first user.
  • the first action may be a body action, such as raising a leg, raising an arm, twisting, bending, and the like.
  • the first action may be a gesture action, for example: an applause gesture, an OK gesture, a scissors gesture, and the like.
  • the first action may also be an emoticon, such as: smiling emoticon, laughing emoticon, pouting emoticon, surprised emoticon, and the like.
  • the action information of the first action may be determined in the following manner: determine a target body part, where the target body part includes at least one of the following parts of the first user: limbs, hands, face ; respectively detecting feature information of the target body part in multiple image frames of the first video; determining the first action according to the feature information of the target body part detected in the multiple image frames action information.
  • the action information of the first action includes: the action type of the first action and the action duration of the first action.
  • the action duration of the first action may refer to the hold time of the first action, or refer to the sum of the time taken to complete the first action and the hold time of the first action.
  • the target body part is a limb as an example.
  • the first video includes 12 image frames
  • the target recognition algorithm uses the target recognition algorithm to recognize the limbs in the image frame to obtain the feature information of the limbs.
  • the left arm hangs naturally in image frame 1
  • the angle between the left arm and the torso in image frame 2 is 30 degrees
  • the angle between the left arm and the torso in image frame 3 is 60 degrees
  • the left arm in image frame 4 is 60 degrees.
  • the angle between the arm and the torso is 90 degrees
  • the angle between the left arm and the torso in image frame 5 is 120 degrees.
  • the included angle between the left arm and the torso in image frame 6 to image frame 12 is 180 degrees.
  • the action type of the first action is determined as "raise the left arm” according to the feature information of the limbs recognized from the above image frames.
  • the time interval between image frame 1 and image frame 12 may be used as the action duration of the first action, or the time interval between image frame 6 and image frame 12 may be taken as the action duration of the first action.
  • the first video can be sampled according to a preset sampling rule, and the above-mentioned identification process of the target body part can be performed on each image frame after sampling, which can improve the performance of the second action.
  • the real-time performance of the detection result of an action can be improved.
  • S203 Generate target music corresponding to the preset duration according to the action information of the first action.
  • a piece of music is usually formed by a variety of musical elements, such as: scales, timbres, tunes, etc.
  • the correspondence between different actions and music elements may be defined in advance. In this way, using the above correspondence, the first action can be mapped to the relevant information of the music element, so as to generate the target music.
  • a correspondence between body movements and musical scales may be defined, that is, different types of body movements correspond to different types of musical scales.
  • a correspondence between facial expressions and timbres may be defined, that is, different facial expressions correspond to different timbres.
  • a correspondence between gesture actions and tunes may be defined, that is, different gesture actions correspond to different tunes. It should be noted that the above correspondences are only some possible examples. The different examples above can be used in combination with each other.
  • the following will take the mapping of different actions to different scales as an example for illustration, and the scale information of the target scale corresponding to the first action may be determined according to the action information of the first action.
  • Information includes scale type and duration.
  • the scale type of the target scale may be determined according to the action type of the first action and a preset corresponding relationship, and the preset corresponding relationship is used to indicate the corresponding relationship between different action types and different scale types .
  • the preset corresponding relationship may be as shown in Table 1.
  • the sound length of the target scale may be determined according to the action duration of the first action.
  • the target music may be generated according to the scale information of the target scale.
  • the method may further include: playing the target music.
  • the first user can hear the effect of the target music in time.
  • the first user can adjust the action in real time to correct the target music when he thinks that the effect of the target music is not good by listening to the target music Or adjust to improve the efficiency of users creating music and improve user experience.
  • the music generation method provided in this embodiment includes: acquiring a first video obtained by collecting the first user within a preset time period, and determining the first user’s time in the preset time period according to the first video.
  • the action information of the first action executed within the first action the action information includes the action type and the action duration, and according to the action information of the first action, the target music corresponding to the preset duration is generated.
  • the user can perform a series of actions to interact with the terminal device to realize music creation, which makes the interactivity and interest high, and can improve the user experience.
  • Fig. 3 is a schematic flowchart of another method for generating music provided by an embodiment of the present disclosure. As shown in Figure 3, the method of this embodiment includes:
  • the target music is the music to be composed/generated.
  • the timbre type of the target music can be any one of the following: piano timbre, accordion timbre, violin timbre, harmonica timbre, and so on.
  • the generation mode is the first generation mode based on free creation, or the second generation mode based on reference music.
  • the first user can freely organize the sequence of actions and the action duration of each action. That is to say, the first user is not restricted during the music creation process, and can create music completely according to his own preference.
  • the first user needs to organize the sequence of actions based on the scale sequence in the reference music, and adjust the duration of each scale in the reference music through the action duration of each action, so as to generate the target music.
  • the generated target music is equivalent to adapting the rhythm of the reference music.
  • the terminal device may receive the timbre type of the target music input by the first user, and receive the generation mode of the target music input by the first user.
  • FIG. 4 is a schematic diagram of a display interface provided by an embodiment of the present disclosure.
  • the terminal device can display the first interface as shown in Figure 4(a) to the user, in the first interface provide options for multiple timbre types, and the user can select the target in the first interface according to the creative needs
  • the tone type of the music For example, suppose the user selects "piano tone" in the first interface. After the user clicks the next step, the terminal device may display the second interface as shown in FIG. 4( b ).
  • the terminal device provides the option of generating mode in the second interface, and the user can select the generating mode of the target music in the second interface according to the creation requirement.
  • S303 Obtain a first video obtained by collecting the first user within a preset time period.
  • S304 According to the first video, determine action information of a first action performed by the first user within the preset duration, where the action information includes an action type and an action duration.
  • S305 Determine, according to the action information of the first action, scale information of a target scale corresponding to the first action, where the scale information includes a scale type and a sound length.
  • S303 to S307 may be repeatedly executed for multiple rounds.
  • the user performs the first action, and the terminal device determines the scale information of the target scale according to the action information of the first action.
  • the musical scale information of the target musical scale determined during multiple rounds of execution forms the target music created by the user.
  • FIG. 5 is a schematic diagram of another display interface provided by an embodiment of the present disclosure.
  • the terminal device displays the third interface as shown in FIG. 5( b ).
  • the current creation progress is displayed on the third interface.
  • the target scale determined by the terminal device is "1"; assuming that the second action performed by the user is “raise the right leg”, the terminal device The determined target scale is “2"; assuming that the third action performed by the user is “raise the right arm (over the shoulder)", the target scale determined by the terminal device is "4"; assuming that the user performs the fourth action The action is “raise the left arm (over the shoulder)", and the target scale determined by the terminal device is "5"; as shown in Figure 5(b), the current creation progress is "1 2 4 5".
  • the creation progress can be displayed in various ways, for example, in the form of musical scale sequence, numbered musical notation, etc., which is not limited in this embodiment.
  • the terminal device receives the completion instruction and determines that the creation of the target music is completed.
  • the reference music is the music that needs to be adapted to generate the target music.
  • the reference music may be specified by the user, or may be randomly determined by the terminal device.
  • S310 Perform scale analysis processing on the reference music to obtain a scale sequence corresponding to the reference music.
  • the scale sequence includes multiple reference scales, and the multiple reference scales are arranged in sequence according to the order in which they appear in the reference music.
  • the terminal device may receive reference music input by the first user.
  • FIG. 6 is a schematic diagram of another display interface provided by an embodiment of the present disclosure.
  • the terminal device displays the fourth interface as shown in FIG. 6( b ).
  • the terminal device displays selection controls for the user to select reference music on the fourth interface.
  • the user can input reference music to the terminal device by selecting the control.
  • the terminal device can also display an input control on the fourth interface, so that the user can also input the name of the reference music to the terminal device through the input control.
  • the user can also input reference music to the terminal device by voice. This embodiment does not limit it.
  • the terminal device performs scale analysis processing on the reference music to obtain
  • the reference scales form the following scale sequence in the order in which the reference scales appear.
  • the terminal device may display the fifth interface shown in FIG. 6(c), in which the above-mentioned scale sequence is displayed. In this way, the user can perform actions corresponding to the reference scales according to the order of the reference scales in the scale sequence displayed on the fifth interface.
  • S311 Determine a target reference scale according to the order of the reference scales in the scale sequence.
  • S311 to S318 may be repeatedly executed for multiple rounds.
  • the first reference scale in the scale sequence is determined as the target reference scale.
  • the second reference scale in the scale sequence is determined as the target reference scale. and so on.
  • the user needs to perform an action corresponding to the target reference scale.
  • the target reference scale can also be highlighted (for example, the target reference scale is located in the rectangular box in Figure 6(c)), so that it is more intuitive Remind the user of the current creation progress and the actions that need to be performed.
  • S312 Acquire a first video obtained by collecting the first user within a preset time period.
  • S313 According to the first video, determine action information of a first action performed by the first user within a preset duration, where the action information includes an action type and an action duration.
  • S314 According to the action information of the first action, determine the scale information of the target scale corresponding to the first action, where the scale information includes a scale type and a sound length.
  • S315 Determine whether the scale type of the target scale is the same as the scale type of the target reference scale.
  • the user when the user selects the second generating mode, the user needs to perform corresponding actions according to the order of the reference scales in the reference music. Therefore, in each round of execution, it is necessary to determine whether the scale type of the target scale is the same as the scale type of the target reference scale. If they are the same, S316 can be executed. If not, return to S312 to re-detect the actions performed by the user.
  • the terminal device may also display a prompt message on the fifth interface shown in FIG. 6(c), such as "the current action is incorrect, Please perform the action again" to prompt the user to make timely adjustments.
  • S317 Play the target music corresponding to the preset duration.
  • S318 Determine whether the target reference scale is the last scale in the scale sequence.
  • the scale sequence in the generated target music is the same as the scale sequence in the reference music.
  • the difference between the two is that the length of each scale is different.
  • the rhythm of the music is different. Therefore, the target music can be regarded as obtained by adapting the rhythm of the reference music.
  • the user can perform a series of actions to interact with the terminal device to realize music creation, so that interactivity and interest are high, and user experience can be improved. Furthermore, the user can also use the first generation mode based on free creation, or the second generation mode based on reference music to create music, which further improves the fun of creating music.
  • Fig. 7 is a schematic structural diagram of a music generating device provided by an embodiment of the present disclosure.
  • the means may be in the form of software and/or hardware.
  • the apparatus may be a terminal device, or a processor, chip, chip module, module, unit, application program, etc. integrated into the terminal device.
  • the music generation device 700 provided in this embodiment includes: an acquisition module 701 , a determination module 702 and a generation module 703 .
  • the obtaining module 701 is configured to obtain the first video obtained by collecting the first user within a preset duration
  • a determining module 702 configured to determine action information of a first action performed by the first user within the preset duration according to the first video, where the action information includes an action type and an action duration;
  • a generating module 703, configured to generate target music corresponding to the preset duration according to the action information of the first action.
  • the generating module 703 is specifically used to:
  • the scale information of the target scale corresponding to the first action determines the scale information of the target scale corresponding to the first action, and the scale information includes a scale type and a sound length;
  • the target music is generated according to the scale information of the target scale.
  • the generating module 703 is specifically used to:
  • the pitch length of the target scale is determined according to the action duration of the first action.
  • the obtaining module 701 is also used to obtain the generation mode of the target music, the generation mode is the first generation mode based on free creation, or the second generation mode based on reference music;
  • the generation module 703 is specifically configured to: generate the target music according to the generation mode of the target music and the scale information of the target scale.
  • the generating module 703 is specifically used to:
  • the generation mode of the target music is the first generation mode, then generate the target music according to the scale information of the target scale; or,
  • the generation mode of the target music is the second generation mode, then determine the target reference scale from the reference music, and when the scale type of the target scale is the same as the scale type of the target reference scale, according to the The scale information of the target scale is used to generate the target music.
  • the acquiring module 701 is further configured to: acquire the reference music, perform scale analysis processing on the reference music, and obtain a scale sequence corresponding to the reference music, and the scale sequence includes multiple reference music Scales, the plurality of reference scales are arranged sequentially according to the order in which they appear in the reference music;
  • the generating module 703 is specifically configured to: determine the target reference scale according to the order of the reference scales in the scale sequence.
  • the generating module 703 is specifically used to:
  • the target music is generated according to the tone color type and scale information of the target scale.
  • the device further includes:
  • the playing module is used to play the target music.
  • the determining module 702 is specifically configured to:
  • the target body part comprising at least one of the following parts of the first user: limbs, hands, face;
  • Action information of the first action is determined according to the feature information of the target body part detected in the plurality of image frames.
  • the music generating device provided in this embodiment can be used to execute the music generating method provided in any of the above method embodiments, and its implementation principle and technical effect are similar, and will not be repeated here.
  • the embodiments of the present disclosure further provide an electronic device.
  • the electronic device 800 may be a terminal device or a server.
  • the terminal equipment may include but not limited to mobile phones, notebook computers, digital broadcast receivers, personal digital assistants (Personal Digital Assistant, PDA for short), tablet computers (Portable Android Device, PAD for short), portable multimedia players (Portable Mobile terminals such as Media Player (PMP for short), vehicle-mounted terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital television (Digital Television, digital TV for short), desktop computers, etc.
  • PDA Personal Digital Assistant
  • PDA Personal Digital Assistant
  • PAD Personal Android Device
  • portable multimedia players Portable Mobile terminals such as Media Player (PMP for short
  • vehicle-mounted terminals such as vehicle navigation terminals
  • fixed terminals such as digital television (Digital Television, digital TV for short), desktop computers, etc.
  • the electronic device shown in FIG. 8 is only an example, and should not limit the functions and scope of use of the embodiments of the present disclosure.
  • an electronic device 800 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 808 loads programs in random access memory (Random Access Memory, RAM for short) 803 to execute various appropriate actions and processes. In the RAM 803, various programs and data necessary for the operation of the electronic device 800 are also stored.
  • the processing device 801, ROM 802, and RAM 803 are connected to each other through a bus 804.
  • An input/output (Input/Output, I/O for short) interface 805 is also connected to the bus 804 .
  • an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; ), a speaker, a vibrator, etc.
  • a storage device 808 including, for example, a magnetic tape, a hard disk, etc.
  • the communication means 809 may allow the electronic device 800 to communicate with other devices wirelessly or by wire to exchange data. While FIG. 8 shows electronic device 800 having various means, it is to be understood that implementing or having all of the means shown is not a requirement. More or fewer means may alternatively be implemented or provided.
  • embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, where the computer program includes program codes for executing the methods shown in the flowcharts.
  • the computer program may be downloaded and installed from a network via communication means 809, or from storage means 808, or from ROM 802.
  • the processing device 801 the above-mentioned functions defined in the methods of the embodiments of the present disclosure are executed.
  • the above-mentioned computer-readable medium in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two.
  • a computer readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof.
  • Computer-readable storage media may include, but are not limited to, electrical connections with one or more wires, portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable Programming read-only memory (Erasable Programmable Read Only Memory, referred to as EPROM or flash memory), optical fiber, portable compact disk read-only memory (Compact Disk Read Only Memory, referred to as CD-ROM), optical storage device, magnetic storage device, or any of the above the right combination.
  • a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
  • a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave carrying computer-readable program code therein. Such propagated data signals may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing.
  • a computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can transmit, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device .
  • the program code contained on the computer readable medium can be transmitted by any appropriate medium, including but not limited to: electric wire, optical cable, radio frequency (Radio Frequency, RF for short), etc., or any suitable combination of the above.
  • the above-mentioned computer-readable medium may be included in the above-mentioned electronic device, or may exist independently without being incorporated into the electronic device.
  • the above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed by the electronic device, the electronic device is made to execute the methods shown in the above-mentioned embodiments.
  • Computer program code for carrying out the operations of the present disclosure can be written in one or more programming languages, or combinations thereof, including object-oriented programming languages—such as Java, Smalltalk, C++, and conventional Procedural Programming Language - such as "C" or a similar programming language.
  • the program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server.
  • the remote computer can be connected to the user's computer through any kind of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or it can be connected to an external A computer (connected via the Internet, eg, using an Internet service provider).
  • LAN Local Area Network
  • WAN Wide Area Network
  • each block in a flowchart or block diagram may represent a module, program segment, or portion of code that contains one or more logical functions for implementing specified executable instructions.
  • the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or they may sometimes be executed in the reverse order, depending upon the functionality involved.
  • each block of the block diagrams and/or flowchart illustrations, and combinations of blocks in the block diagrams and/or flowchart illustrations can be implemented by a dedicated hardware-based system that performs the specified functions or operations , or may be implemented by a combination of dedicated hardware and computer instructions.
  • the units involved in the embodiments described in the present disclosure may be implemented by software or by hardware. Wherein, the name of the unit does not constitute a limitation of the unit itself under certain circumstances, for example, the first obtaining unit may also be described as "a unit for obtaining at least two Internet Protocol addresses".
  • exemplary types of hardware logic components include: Field Programmable Gate Array (Field Programmable Gate Array, FPGA for short), Application Specific Integrated Circuit (ASIC for short), Application Specific Standard Products ( Application Specific Standard Parts (ASSP for short), System on Chip (SOC for short), Complex Programmable Logic Device (CPLD for short), etc.
  • a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device.
  • a machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium.
  • a machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing.
  • machine-readable storage media would include one or more wire-based electrical connections, portable computer discs, hard drives, random access memory (RAM), read only memory (ROM), erasable programmable read only memory (EPROM or flash memory), optical fiber, compact disk read only memory (CD-ROM), optical storage, magnetic storage, or any suitable combination of the foregoing.
  • RAM random access memory
  • ROM read only memory
  • EPROM or flash memory erasable programmable read only memory
  • CD-ROM compact disk read only memory
  • magnetic storage or any suitable combination of the foregoing.
  • a music generation method including:
  • action information of a first action performed by the first user within the preset duration according to the first video where the action information includes an action type and an action duration
  • target music corresponding to the preset duration is generated.
  • generating the target music corresponding to the preset duration includes:
  • the scale information of the target scale corresponding to the first action determines the scale information of the target scale corresponding to the first action, and the scale information includes a scale type and a sound length;
  • the target music is generated according to the scale information of the target scale.
  • determining the scale information of the target scale corresponding to the first action includes:
  • the pitch length of the target scale is determined according to the action duration of the first action.
  • before obtaining the first video obtained by capturing the first user within a preset time period further includes:
  • the generation mode is a first generation mode based on free creation, or a second generation mode based on reference music;
  • generating the target music according to the scale information of the target scale includes:
  • the target music is generated according to the generation mode of the target music and the scale information of the target scale.
  • generating the target music according to the generation mode of the target music and the scale information of the target scale includes:
  • the generation mode of the target music is the first generation mode, then generate the target music according to the scale information of the target scale; or,
  • the generation mode of the target music is the second generation mode, then determine the target reference scale from the reference music, and when the scale type of the target scale is the same as the scale type of the target reference scale, according to the The scale information of the target scale is used to generate the target music.
  • the target reference scale from the reference music before determining the target reference scale from the reference music, it also includes:
  • determining the target reference scale from the reference music includes:
  • the target reference scale is determined according to the order of the reference scales in the scale sequence.
  • generating the target music according to the scale information of the target scale includes:
  • the target music is generated according to the tone color type and scale information of the target scale.
  • after generating the target music according to the scale information of the target scale further includes:
  • determining the action information of the first action performed by the first user within the preset duration includes:
  • the target body part comprising at least one of the following parts of the first user: limbs, hands, face;
  • Action information of the first action is determined according to the feature information of the target body part detected in the plurality of image frames.
  • a music generation device including:
  • An acquisition module configured to acquire the first video obtained by collecting the first user within a preset duration
  • a determining module configured to determine action information of a first action performed by the first user within the preset duration according to the first video, the action information including action type and action duration;
  • a generating module configured to generate target music corresponding to the preset duration according to the action information of the first action.
  • the generating module is specifically configured to:
  • the scale information of the target scale corresponding to the first action determines the scale information of the target scale corresponding to the first action, and the scale information includes a scale type and a sound length;
  • the target music is generated according to the scale information of the target scale.
  • the generating module is specifically configured to:
  • the pitch length of the target scale is determined according to the action duration of the first action.
  • the acquisition module is further configured to acquire the generation mode of the target music, the generation mode is the first generation mode based on free creation, or the first generation mode based on reference music Two generation mode;
  • the generation module is specifically configured to: generate the target music according to the generation mode of the target music and the scale information of the target scale.
  • the generating module is specifically configured to:
  • the generation mode of the target music is the first generation mode, then generate the target music according to the scale information of the target scale; or,
  • the generation mode of the target music is the second generation mode, then determine the target reference scale from the reference music, and when the scale type of the target scale is the same as the scale type of the target reference scale, according to the The scale information of the target scale is used to generate the target music.
  • the acquiring module is further configured to acquire the reference music, perform scale analysis processing on the reference music, and obtain a scale sequence corresponding to the reference music, in the scale sequence including a plurality of reference scales, the plurality of reference scales are arranged sequentially according to the order of their appearance in the reference music;
  • the generating module is specifically configured to determine the target reference scale according to the order of the reference scales in the scale sequence.
  • the generating module is specifically configured to:
  • the target music is generated according to the tone color type and scale information of the target scale.
  • the device further includes:
  • the playing module is used to play the target music.
  • the determining module is specifically configured to:
  • the target body part comprising at least one of the following parts of the first user: limbs, hands, face;
  • Action information of the first action is determined according to the feature information of the target body part detected in the plurality of image frames.
  • an electronic device including: at least one processor and a memory;
  • the memory stores computer-executable instructions
  • the at least one processor executes the computer-executed instructions stored in the memory, so that the at least one processor executes the music generation method described in the above first aspect and various possible implementation manners of the first aspect.
  • a computer-readable storage medium stores computer-executable instructions, and when a processor executes the computer-executable instructions, Realize the music generation method described in the above first aspect and various possible implementation manners of the first aspect.
  • a computer program product including a computer program, when the computer program is executed by a processor, various possible implementations of the first aspect and the first aspect can be realized The music generation method described in the manner.
  • a computer program is provided, and when the computer program is executed by a processor, the music described in the first aspect and various possible implementation modes of the first aspect is realized. generate method.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Theoretical Computer Science (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

一种音乐生成方法、装置、设备、电子设备、计算机可读存储介质、计算机程序产品及计算机程序,该方法包括:获取在预设时长内对第一用户进行采集得到的第一视频(S201),根据第一视频,确定第一用户在预设时长内执行的第一动作的动作信息,该动作信息包括动作类型和动作时长(S202),进而根据该第一动作的动作信息,生成预设时长对应的目标音乐(S203)。通过上述过程,用户可以通过执行一系列的动作与终端设备进行交互以实现音乐创作,提高了音乐创作过程的互动性和趣味性,进而提升用户体验。

Description

音乐生成方法、装置、设备、存储介质及程序
相关申请的交叉引用
本公开要求于2021年09月28日提交中国专利局、申请号为202111140485.0、申请名称为“音乐生成方法、装置、设备、存储介质及程序”的中国专利申请的优先权,其全部内容通过引用结合在本公开中。
技术领域
本公开实施例涉及人工智能技术领域,尤其涉及一种音乐生成方法、装置、设备、电子设备、计算机可读存储介质、计算机程序产品及计算机程序。
背景技术
随着终端技术的发展,用户希望能够通过终端设备来创作音乐,从而增加音乐创作的趣味性。
相关技术中,可以在终端设备的屏幕中显示钢琴按键,用户通过触摸屏幕中的按键来模拟弹钢琴,从而实现音乐创作。
然而,上述创作音乐的方式较为单调,互动性较差。
发明内容
本公开实施例提供一种音乐生成方法、装置、设备、电子设备、计算机可读存储介质、计算机程序产品及计算机程序,以解决音乐创作方式互动性较差问题。
第一方面,本公开实施例提供一种音乐生成方法,包括:
获取在预设时长内对第一用户进行采集得到的第一视频;
根据所述第一视频,确定所述第一用户在所述预设时长内执行的第一动作的动作信息,所述动作信息包括动作类型和动作时长;
根据所述第一动作的动作信息,生成所述预设时长对应的目标音乐。
第二方面,本公开实施例提供一种音乐生成装置,包括:
获取模块,用于获取在预设时长内对第一用户进行采集得到的第一视频;
确定模块,用于根据所述第一视频,确定所述第一用户在所述预设时长内执行的第一动作的动作信息,所述动作信息包括动作类型和动作时长;
生成模块,用于根据所述第一动作的动作信息,生成所述预设时长对应的目标音乐。
第三方面,本公开实施例提供一种电子设备,包括:处理器和存储器;
所述存储器存储计算机执行指令;
所述处理器执行所述计算机执行指令,实现如第一方面以及第一方面各种可能的实现方式中的音乐生成方法。
第四方面,本公开实施例提供一种计算机可读存储介质,所述计算机可读存储介质中存储有计算机执行指令,当处理器执行所述计算机执行指令时,实现如上第一方面以及第一方面各种可能的实现方式中的音乐生成方法。
第五方面,本公开实施例提供一种计算机程序产品,包括计算机程序,所述计算机程序被处理器执行时实现如上第一方面以及第一方面各种可能的实现方式中的音乐生成方法。
第六方面,本公开实施例提供一种计算机程序,所述计算机程序被处理器执行时实现如上第一方面以及第一方面各种可能的实现方式中的音乐生成方法。
本公开实施例提供的音乐生成方法、装置、设备、电子设备、计算机可读存储介质、计算机程序产品及计算机程序,所述方法包括:获取在预设时长内对第一用户进行采集得到的第一视频,根据所述第一视频,确定所述第一用户在所述预设时长内执行的第一动作的动作信息,所述动作信息包括动作类型和动作时长,进而根据所述第一动作的动作信息,生成所述预设时长对应的目标音乐。通过上述过程,用户可以通过执行一系列的动作与终端设备进行交互以实现音乐创作,提高了音乐创作过程的互动性和趣味性,进而提升用户体验。
附图说明
为了更清楚地说明本公开实施例或相关技术中的技术方案,下面将对实施例或相关技术描述中所需要使用的附图作一简单地介绍,显而易见地,下面描述中的附图是本公开的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1为本公开实施例提供的一种应用场景的示意图;
图2为本公开实施例提供的一种音乐生成方法的流程示意图;
图3为本公开实施例提供的另一种音乐生成方法的流程示意图;
图4为本公开实施例提供的一种显示界面的示意图;
图5为本公开实施例提供的另一种显示界面的示意图;
图6为本公开实施例提供的又一种显示界面的示意图;
图7为本公开实施例提供的一种音乐生成装置的结构示意图;
图8为本公开实施例提供的一种电子设备的结构示意图。
具体实施方式
为使本公开实施例的目的、技术方案和优点更加清楚,下面将结合本公开实施例中的附图,对本公开实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例是本公开一部分实施例,而不是全部的实施例。基于本公开中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本公开保护的范围。
首先对本公开实施例中涉及的概念和术语进行解释说明。
音阶:将调式中的音,从以主音开始到以主音结束,由低到高(叫做上行),或者由高到低(叫做下行)以阶梯状排列起来,就叫做音阶。
不同地域通常采用不同的音阶系统。目前,自然七声音阶是应用最广的七声音阶, 其音程组织是,每个八度之内有5处全音,分成两个一串和3个一串,两串之间以半音隔开。7个音阶分别为:1(Do)、2(Re)、3(Mi)、4(Fa)、5(So)、6(La)、7(Si)。
需要说明的是,本公开实施例对于音阶系统不做限定,可以应用于任意音阶系统。实施例中为了便于理解,在涉及举例说明时,以自然七声音阶为例进行示例。
音色:不同声音表现在波形方面总是有与众不同的特性,不同的物体振动都有不同的特点。不同的发声体由于其材料、结构不同,则发出声音的音色也不同。例如钢琴、小提琴和人发出的声音不一样,每一个人发出的声音也不一样。因此,可以把音色理解为声音的特征。
为了便于对本公开技术方案的理解,下面结合图1对本公开实施例的应用场景进行介绍。
图1为本公开实施例提供的一种应用场景的示意图。如图1所示,该应用场景涉及终端设备。终端设备中设置有摄像头等视频采集装置。终端设备还设置有扬声器等音频播放装置。终端设备中还设置有音乐生成装置。音乐生成装置可以为软件和/或硬件的形式,示例性的,音乐生成装置可以为终端设备中的处理器、芯片、芯片模组、模块、单元、应用程序等。
参见图1,图像采集装置可以对用户的动作(例如肢体动作、表情动作、手势动作等)进行采集得到视频。图像采集装置向音乐生成装置发送视频。音乐生成装置可以基于视频生成音乐。示例性的,参见图1,音乐生成装置可以从视频中识别得到用户的动作序列,并将用户的动作映射为音阶,得到音阶序列。进而可以根据音阶序列生成音乐。音乐生成装置将生成的音乐发送至音频播放装置,音频播放装置对音乐进行播放。
上述应用场景中,用户通过执行不同的动作序列,即可使终端设备生成不同的音乐,实现音乐创作。在用户创作音乐的过程中,用户可以通过执行一系列的动作与终端设备进行互动以实现音乐创作,使得互动性和趣味性较高,能够提升用户体验。
本公开实施例中的终端设备可以是具有视频采集装置和音频播放装置的任意电子设备,包括但不限于:智能手机、平板电脑、笔记本电脑、智能电视、智能穿戴设备、智能家居设备等。
下面以具体地实施例对本公开的技术方案进行详细说明。下面这几个具体的实施例可以相互结合,对于相同或相似的概念或过程可能在某些实施例中不再赘述。
图2为本公开实施例提供的一种音乐生成方法的流程示意图。本实施例的方法可以由终端设备执行,或者,由终端设备中的音乐生成装置执行。如图2所示,本实施例的方法包括:
S201:获取在预设时长内对第一用户进行采集得到的第一视频。
本实施例应用于第一用户创作目标音乐的场景。
其中,预设时长表示本实施例在时间维度的执行粒度。本实施例执行时,按照预设时长,获取一个预设时长内采集到的第一视频,并基于第一视频生成该第一预设时长对应的目标音乐。预设时长也可以称为时间窗口。通常,预设时长可以设置为一个较小的值,例如,1秒、100毫秒等。示例性的,预设时长可以是用户执行一个动作的时长。
能够理解的是,用户创作一首音乐通常需要重复执行本实施例多次。本实施例在不 同次执行时,对应的预设时长可以相同也可以不同。这样,多个预设时长对应的目标音乐组合成最终的完整音乐。
本实施例可以采用如下两种执行方式:
方式一:音乐生成过程与视频采集过程同步进行。也就是说,在对第一用户采集视频的过程中同步执行本实施例的方法,从而实现一边采集视频一边基于视频生成音乐。具体而言,结合图1,视频采集装置每采集到预设时长的第一视频,则把第一视频发送至音乐生成装置,由音乐生成装置基于第一视频生成该预设时长对应的目标音乐。
方式二:音乐生成过程与视频采集过程异步执行。举例而言,视频采集装置对用户进行采集得到第二视频,例如第二视频的持续时长可以为3分钟或者5分钟等。第二视频可以为用户跳一支舞的完整视频。上述采集到的第二视频存储至预设存储空间中。在视频采集完成之后的某个时间,音乐生成装置从预设存储空间中获取第二视频,以预设时长为时间粒度,从第二视频中获取预设时长对应的第一视频,根据第一视频生成预设时长对应的目标音乐。
S202:根据所述第一视频,确定所述第一用户在所述预设时长内执行的第一动作的动作信息,所述动作信息包括动作类型和动作时长。
本实施例中,第一动作是由第一用户的预设身体部位所执行的动作。举例而言,第一动作可以为肢体动作,例如:抬腿、抬胳膊、扭腰、弯腰等。第一动作可以为手势动作,例如:鼓掌手势、OK手势、剪刀手势等。第一动作还可以为表情动作,例如:微笑表情、大笑表情、嘟嘴表情、惊讶表情等。
一种可能的实现方式中,可以采用如下方式确定第一动作的动作信息:确定目标身体部位,目标身体部位包括所述第一用户的下述部位中的至少一种:肢体、手部、面部;在所述第一视频的多个图像帧中分别检测所述目标身体部位的特征信息;根据所述多个图像帧中检测得到的所述目标身体部位的特征信息,确定所述第一动作的动作信息。
其中,第一动作的动作信息包括:第一动作的动作类型和第一动作的动作时长。其中,第一动作的动作时长可以是指第一动作的保持时长,或者是指完成第一动作所花费时长与第一动作的保持时长之和。
示例性的,以目标身体部位为肢体为例。假设第一视频包括12个图像帧,针对每个图像帧,分别利用目标识别算法对该图像帧中的肢体进行识别,得到肢体的特征信息。例如,图像帧1中左胳膊自然下垂,图像帧2中左胳膊与躯干之间的夹角为30度,图像帧3中左胳膊与躯干之间的夹角为60度,图像帧4中左胳膊与躯干之间的夹角为90度,图像帧5中左胳膊与躯干之间的夹角为120度。图像帧6至图像帧12中左胳膊与躯干之间的夹角为180度。这样,根据从上述各图像帧中识别得到的肢体的特征信息,确定出第一动作的动作类型为“抬起左胳膊”。进而,可以将图像帧1与图像帧12之间的时间间隔作为第一动作的动作时长,或者,将图像帧6与图像帧12之间的时间间隔作为第一动作的动作时长。
可选的,为了提高第一动作的检测效率,可以按照预设的采样规则,对第一视频进行采样处理,对采样后的各图像帧执行上述的目标身体部分的识别过程,这样可以提高第一动作的检测结果的实时性。
S203:根据所述第一动作的动作信息,生成所述预设时长对应的目标音乐。
一首音乐通常由多种音乐元素形成,例如:音阶、音色、曲调等。本实施例中,可以事先定义不同动作与音乐元素之间的对应关系。这样,利用上述对应关系,可以将第一动作映射为音乐元素的相关信息,从而生成目标音乐。
一些示例中,可以定义肢体动作与音阶之间的对应关系,即,不同的肢体动作类型对应不同的音阶类型。另一些示例中,可以定义表情动作与音色之间的对应关系,即不同的表情动作对应不同的音色。又一些示例中,可以定义手势动作与曲调之间的对应关系,即,不同的手势动作对应不同的曲调。需要说明的是,上述对应关系仅为一些可能的示例。上述不同的示例可以相互结合使用。
一种可能的实现方式中,下面将以不同动作映射为不同音阶为例进行说明,可以根据所述第一动作的动作信息,确定所述第一动作对应的目标音阶的音阶信息,所述音阶信息包括音阶类型和音长。
示例性的,可以根据所述第一动作的动作类型和预设对应关系,确定所述目标音阶的音阶类型,所述预设对应关系用于指示不同动作类型与不同音阶类型之间的对应关系。其中,预设对应关系可以如表1所示。
表1
动作类型 音阶类型
抬起左腿 1(Do)
抬起右腿 2(Re)
抬起左胳膊 3(Mi)
抬起右胳膊 4(Fa)
抬起左胳膊过肩 5(So)
抬起右胳膊过肩 6(La)
T字型(即左右胳膊与肩平齐,左右腿竖直站立) 7(Si)
示例性的,可以根据第一动作的动作时长,确定所述目标音阶的音长。
进一步的,确定出第一动作对应的目标音阶的音阶信息之后,可以根据所述目标音阶的音阶信息,生成所述目标音乐。
一种可能的实现方式中,在S203生成目标音乐之后,还可以包括:播放所述目标音乐。这样,使得第一用户可以及时听到目标音乐的效果。尤其是在上述方式一(即音乐生成过程与视频采集过程同步进行)的场景中,第一用户通过收听目标音乐,在认为目标音乐的效果不佳时,可以实时调整动作以对目标音乐进行修正或者调整,提升用户创作音乐的效率,提升用户体验。
本实施例提供的音乐生成方法,包括:获取在预设时长内对所述第一用户进行采集得到的第一视频,根据所述第一视频,确定所述第一用户在所述预设时长内执行的第一动作的动作信息,所述动作信息包括动作类型和动作时长,根据所述第一动作的动作信息,生成所述预设时长对应的目标音乐。上述过程中,用户可以通过执行一系列的动作与终端设备进行交互以实现音乐创作,使得互动性和趣味性较高,能够提升用户体验。
在上述实施例的基础上,下面结合一个更具体的实施例对本公开技术方案进行更详 细的描述。
图3为本公开实施例提供的另一种音乐生成方法的流程示意图。如图3所示,本实施例的方法包括:
S301:获取目标音乐的音色类型。
目标音乐为待创作/待生成的音乐。目标音乐的音色类型可以为下述中的任意一种:钢琴音色、手风琴音色、小提琴音色、口琴音色等。
S302:获取目标音乐的生成模式。
其中,所述生成模式为基于自由创作的第一生成模式,或者,为基于参考音乐的第二生成模式。
在第一生成模式中,第一用户可以自由组织动作的先后顺序、以及各动作的动作时长。也就是说,第一用户在音乐创作过程中不受到约束,可以完全按照自己的喜好来创作音乐。
在第二生成模式中,第一用户需要基于参考音乐中的音阶序列来组织动作的先后顺序,并通过各动作的动作时长来调整参考音乐中各音阶的音长,从而生成目标音乐。这样,生成的目标音乐相当于是对参考音乐的节奏进行改编得到的。
一种可能的实现方式中,终端设备可以接收第一用户输入的目标音乐的音色类型,以及,接收第一用户输入的目标音乐的生成模式。
举例而言,图4为本公开实施例提供的一种显示界面的示意图。如图4所示,终端设备可以向用户展示如图4(a)所示的第一界面,在第一界面中提供多个音色类型的选项,用户可以根据创作需求在第一界面中选择目标音乐的音色类型。例如,假设用户在第一界面中选择“钢琴音色”。当用户点击下一步后,终端设备可以显示如图4(b)所示的第二界面。终端设备在第二界面中提供生成模式的选项,用户可以根据创作需求在第二界面中选择目标音乐的生成模式。
下面针对两种生成模式下的音乐生成过程分别进行介绍。
(一)若目标音乐的生成模式为第一生成模式,则执行下述的S303至S308。
S303:获取在预设时长内对第一用户进行采集得到的第一视频。
S304:根据所述第一视频,确定所述第一用户在所述预设时长内执行的第一动作的动作信息,所述动作信息包括动作类型和动作时长。
S305:根据所述第一动作的动作信息,确定所述第一动作对应的目标音阶的音阶信息,所述音阶信息包括音阶类型和音长。
应理解,S303至S305的实现方式与图2所示实施例类似,此处不做赘述。
S306:根据目标音乐的音色类型以及目标音阶的音阶信息,生成预设时长对应的目标音乐。
S307:播放预设时长对应的目标音乐。
能够理解的是,上述S303至S307可以重复执行多轮。在每轮执行过程中,用户执行第一动作,终端设备根据第一动作的动作信息,确定出目标音阶的音阶信息。这样,多轮执行过程中确定出的目标音阶的音阶信息,形成用户创作的目标音乐。
S308:判断是否接收到完成指令。
若是,则结束。若否,则返回执行S303。
举例而言,图5为本公开实施例提供的另一种显示界面的示意图。如图5所示,假设用户在如图5(a)所示的第二界面中选择“第一生成模式”。当用户点击下一步后,终端设备显示如图5(b)所示的第三界面。在第三界面中显示当前的创作进度。例如,假设用户执行的第一个动作为“抬起左腿”,则终端设备确定出的目标音阶为“1”;假设用户执行的第二个动作为“抬起右腿”,则终端设备确定出的目标音阶为“2”;假设用户执行的第三个动作为“抬起右胳膊(不过肩)”,则终端设备确定出的目标音阶为“4”;假设用户执行的第四个动作为“抬起左胳膊(过肩)”,终端设备确定出的目标音阶为“5”;则参见图5(b)所示,当前的创作进度为“1 2 4 5”。
需要说明的是,图5(b)所示的第三界面中,可以采用多种方式显示创作进度,例如,音阶序列形式、简谱形式等,本实施例对此不作限定。
当用户创作完成时,用户可以在图5(b)所示的第三界面中点击“完成”按钮。这样,终端设备接收到完成指令,确定目标音乐创作完成。
(二)若目标音乐的生成模式为上述第二生成模式,则执行下述的S309至S318。
S309:获取参考音乐。
其中,参考音乐为需要对其进行改编以生成目标音乐的音乐。参考音乐可以为用户指定的,还可以是由终端设备随机确定的。
S310:对参考音乐进行音阶解析处理,得到参考音乐对应的音阶序列,音阶序列中包括多个参考音阶,多个参考音阶按照各自在参考音乐中出现顺序依次排列。
一种可能的实现方式中,终端设备可以接收第一用户输入的参考音乐。示例性的,图6为本公开实施例提供的又一种显示界面的示意图。如图6所示,假设用户在图6(a)所示的第二界面中,选择“第二生成模式”。当用户点击下一步后,终端设备显示如图6(b)所示的第四界面。终端设备在第四界面中显示可供用户选择参考音乐的选择控件。用户可以通过选择控件向终端设备输入参考音乐。可选的,终端设备还可以在第四界面中显示输入控件,这样,用户还可以通过输入控件向终端设备输入参考音乐的名称。可选的,用户还可以通过语音方式向终端设备输入参考音乐。本实施例对此不作限定。
继续参见图6,假设用户在图6(b)所示的第四界面中,选择参考音乐“两只老虎”,则终端设备对该参考音乐进行音阶解析处理,得到该参考音乐中依次出现的参考音阶,按照参考音阶出现的次序形成如下音阶序列。
1、2、3、1、1、2、3、1、3、4、5、3、4、5、5、6、5、4、3、1、5、6、5、4、3、1、2、5、1、2、5、1
继续参见图6,用户选择参考音乐“两只老虎”,并点击下一步后,终端设备可以显示图6(c)所示的第五界面,在第五界面中显示上述音阶序列。这样,用户可以根据第五界面中显示的音阶序列中各参考音阶的顺序,执行参考音阶对应的动作。
S311:按照所述音阶序列中各参考音阶的顺序,确定目标参考音阶。
本实施例中,S311至S318可以重复执行多轮。第一次执行时,将音阶序列中的第一个参考音阶,确定为目标参考音阶。第二次执行时,将音阶序列中的第二个参考音阶,确定为目标参考音阶。以此类推。相应的,在每轮执行过程中,用户需要执行目标参考音阶对应的动作。
可选的,在图6(c)所示的第五界面中,还可以对目标参考音阶进行突出显示(例 如,图6(c)中位于矩形框内的为目标参考音阶),以便更加直观的提醒用户当前的创作进度,以及当前需要执行的动作。
S312:获取在预设时长内对第一用户进行采集得到的第一视频。
S313:根据第一视频,确定第一用户在预设时长内执行的第一动作的动作信息,动作信息包括动作类型和动作时长。
S314:根据第一动作的动作信息,确定第一动作对应的目标音阶的音阶信息,音阶信息包括音阶类型和音长。
应理解,S312至S314的实现方式与图2所示实施例类似,此处不做赘述。
S315:判断目标音阶的音阶类型与目标参考音阶的音阶类型是否相同。
若相同,则执行S316。
若不同,则返回执行S312。
本实施例中,当用户选择第二生成模式时,用户需要按照参考音乐中各参考音阶的顺序,来执行相应的动作。因此,在每一轮执行过程中,需要判断目标音阶的音阶类型与目标参考音阶的音阶类型是否相同。如果相同,则可以执行S316。如果不同,则返回执行S312,重新检测用户所执行的动作。
可选的,在目标音阶的音阶类型与目标参考音阶的音阶类型不同的情况下,终端设备还可以在图6(c)所示的第五界面中显示提示信息,例如“当前动作不正确,请重新执行动作”,以便提示用户及时进行调整。
S316:根据目标音乐的音色类型以及目标音阶的音阶信息,生成预设时长对应的目标音乐。
S317:播放预设时长对应的目标音乐。
S318:判断目标参考音阶是否为音阶序列中的最后一个音阶。
若是,则说明本次基于参考音乐的创作完成,结束。
若否,则返回执行S311。
需要说明的是,在第二生成模式下,生成的目标音乐中的音阶序列与参考音乐中的音阶序列是相同的,二者不同之处在于各音阶的音长不同,即,目标音乐与参考音乐的节奏不同。因此,目标音乐可以视为对参考音乐的节奏进行改编得到的。
本实施例中,用户可以通过执行一系列的动作与终端设备进行交互以实现音乐创作,使得互动性和趣味性较高,能够提升用户体验。进一步的,用户还可以采用基于自由创作的第一生成模式,或者采用基于参考音乐的第二生成模式,来进行音乐创作,进一步提高了创作音乐的趣味性。
图7为本公开实施例提供的一种音乐生成装置的结构示意图。该装置可以为软件和/或硬件的形式。该装置可以为终端设备,或者为集成到终端设备中的处理器、芯片、芯片模组、模块、单元、应用程序等。
如图7所示,本实施例提供的音乐生成装置700,包括:获取模块701、确定模块702和生成模块703。
其中,获取模块701,用于获取在预设时长内对第一用户进行采集得到的第一视频;
确定模块702,用于根据所述第一视频,确定所述第一用户在所述预设时长内执行的第一动作的动作信息,所述动作信息包括动作类型和动作时长;
生成模块703,用于根据所述第一动作的动作信息,生成所述预设时长对应的目标音乐。
一种可能的实现方式中,生成模块703具体用于:
根据所述第一动作的动作信息,确定所述第一动作对应的目标音阶的音阶信息,所述音阶信息包括音阶类型和音长;
根据所述目标音阶的音阶信息,生成所述目标音乐。
一种可能的实现方式中,生成模块703具体用于:
根据所述第一动作的动作类型和预设对应关系,确定所述目标音阶的音阶类型,所述预设对应关系用于指示不同动作类型与不同音阶类型之间的对应关系;
根据所述第一动作的动作时长,确定所述目标音阶的音长。
一种可能的实现方式中,获取模块701还用于获取所述目标音乐的生成模式,所述生成模式为基于自由创作的第一生成模式,或者,为基于参考音乐的第二生成模式;
生成模块703具体用于:根据所述目标音乐的生成模式,以及所述目标音阶的音阶信息,生成所述目标音乐。
一种可能的实现方式中,生成模块703具体用于:
若所述目标音乐的生成模式为所述第一生成模式,则根据所述目标音阶的音阶信息,生成所述目标音乐;或者,
若所述目标音乐的生成模式为所述第二生成模式,则从所述参考音乐中确定目标参考音阶,在所述目标音阶的音阶类型与所述目标参考音阶的音阶类型相同时,根据所述目标音阶的音阶信息,生成所述目标音乐。
一种可能的实现方式中,获取模块701还用于:获取所述参考音乐,对所述参考音乐进行音阶解析处理,得到所述参考音乐对应的音阶序列,所述音阶序列中包括多个参考音阶,所述多个参考音阶按照各自在所述参考音乐中出现顺序依次排列;
生成模块703具体用于:按照所述音阶序列中各参考音阶的顺序,确定所述目标参考音阶。
一种可能的实现方式中,生成模块703具体用于:
获取所述目标音乐的音色类型;
根据所述音色类型和所述目标音阶的音阶信息,生成所述目标音乐。
一种可能的实现方式中,所述装置还包括:
播放模块,用于播放所述目标音乐。
一种可能的实现方式中,确定模块702具体用于:
确定目标身体部位,所述目标身体部位包括所述第一用户的下述部位中的至少一种:肢体、手部、面部;
在所述第一视频的多个图像帧中分别检测所述目标身体部位的特征信息;
根据所述多个图像帧中检测得到的所述目标身体部位的特征信息,确定所述第一动作的动作信息。
本实施例提供的音乐生成装置,可用于执行上述任一方法实施例提供的音乐生成方法,其实现原理和技术效果类似,此处不作赘述。
为了实现上述实施例,本公开实施例还提供了一种电子设备。
参考图8,其示出了适于用来实现本公开实施例的电子设备800的结构示意图,该电子设备800可以为终端设备或服务器。其中,终端设备可以包括但不限于诸如移动电话、笔记本电脑、数字广播接收器、个人数字助理(Personal Digital Assistant,简称PDA)、平板电脑(Portable Android Device,简称PAD)、便携式多媒体播放器(Portable Media Player,简称PMP)、车载终端(例如车载导航终端)等等的移动终端以及诸如数字电视(Digital Televison,简称数字TV)、台式计算机等等的固定终端。图8示出的电子设备仅仅是一个示例,不应对本公开实施例的功能和使用范围带来任何限制。
如图8所示,电子设备800可以包括处理装置(例如中央处理器、图形处理器等)801,其可以根据存储在只读存储器(Read Only Memory,简称ROM)802中的程序或者从存储装置808加载到随机访问存储器(Random Access Memory,简称RAM)803中的程序来执行各种适当的动作和处理。在RAM 803中,还存储有电子设备800操作所需的各种程序和数据。处理装置801、ROM 802以及RAM 803通过总线804彼此相连。输入/输出(Input/Output,简称I/O)接口805也连接至总线804。
通常,以下装置可以连接至I/O接口805:包括例如触摸屏、触摸板、键盘、鼠标、摄像头、麦克风、加速度计、陀螺仪等的输入装置806;包括例如液晶显示器(Liquid Crystal Display,简称LCD)、扬声器、振动器等的输出装置807;包括例如磁带、硬盘等的存储装置808;以及通信装置809。通信装置809可以允许电子设备800与其他设备进行无线或有线通信以交换数据。虽然图8示出了具有各种装置的电子设备800,但是应理解的是,并不要求实施或具备所有示出的装置。可以替代地实施或具备更多或更少的装置。
特别地,根据本公开的实施例,上文参考流程图描述的过程可以被实现为计算机软件程序。例如,本公开的实施例包括一种计算机程序产品,其包括承载在计算机可读介质上的计算机程序,该计算机程序包含用于执行流程图所示的方法的程序代码。在这样的实施例中,该计算机程序可以通过通信装置809从网络上被下载和安装,或者从存储装置808被安装,或者从ROM 802被安装。在该计算机程序被处理装置801执行时,执行本公开实施例的方法中限定的上述功能。
需要说明的是,本公开上述的计算机可读介质可以是计算机可读信号介质或者计算机可读存储介质或者是上述两者的任意组合。计算机可读存储介质例如可以是——但不限于——电、磁、光、电磁、红外线、或半导体的系统、装置或器件,或者以上的任意组合。计算机可读存储介质的更具体的例子可以包括但不限于:具有一个或多个导线的电连接、便携式计算机磁盘、硬盘、随机访问存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(Erasable Programmable Read Only Memory,简称EPROM或闪存)、光纤、便携式紧凑磁盘只读存储器(Compact Disk Read Only Memory,简称CD-ROM)、光存储器件、磁存储器件、或者上述的任意合适的组合。在本公开中,计算机可读存储介质可以是任何包含或存储程序的有形介质,该程序可以被指令执行系统、装置或者器件使用或者与其结合使用。而在本公开中,计算机可读信号介质可以包括在基带中或者作为载波一部分传播的数据信号,其中承载了计算机可读的程序代码。这种传播的数据信号可以采用多种形式,包括但不限于电磁信号、光信号或上述的任意合适的组合。计算机可读信号介质还可以是计算机可读存储介质以外的任何计算机可读介质,该计算机可读信号介质可以发送、传播或者传输用于由指令执行系统、装置或者器件使 用或者与其结合使用的程序。计算机可读介质上包含的程序代码可以用任何适当的介质传输,包括但不限于:电线、光缆、射频(Radio Frequency,简称RF)等等,或者上述的任意合适的组合。
上述计算机可读介质可以是上述电子设备中所包含的;也可以是单独存在,而未装配入该电子设备中。
上述计算机可读介质承载有一个或者多个程序,当上述一个或者多个程序被该电子设备执行时,使得该电子设备执行上述实施例所示的方法。
可以以一种或多种程序设计语言或其组合来编写用于执行本公开的操作的计算机程序代码,上述程序设计语言包括面向对象的程序设计语言—诸如Java、Smalltalk、C++,还包括常规的过程式程序设计语言—诸如“C”语言或类似的程序设计语言。程序代码可以完全地在用户计算机上执行、部分地在用户计算机上执行、作为一个独立的软件包执行、部分在用户计算机上部分在远程计算机上执行、或者完全在远程计算机或服务器上执行。在涉及远程计算机的情形中,远程计算机可以通过任意种类的网络——包括局域网(Local Area Network,简称LAN)或广域网(Wide Area Network,简称WAN)—连接到用户计算机,或者,可以连接到外部计算机(例如利用因特网服务提供商来通过因特网连接)。
附图中的流程图和框图,图示了按照本公开各种实施例的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段、或代码的一部分,该模块、程序段、或代码的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。也应当注意,在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个接连地表示的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或操作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
描述于本公开实施例中所涉及到的单元可以通过软件的方式实现,也可以通过硬件的方式来实现。其中,单元的名称在某种情况下并不构成对该单元本身的限定,例如,第一获取单元还可以被描述为“获取至少两个网际协议地址的单元”。
本文中以上描述的功能可以至少部分地由一个或多个硬件逻辑部件来执行。例如,非限制性地,可以使用的示范类型的硬件逻辑部件包括:现场可编程门阵列(Field Programmable Gate Array,简称FPGA)、专用集成电路(Application Specific Integrated Circuit,简称ASIC)、专用标准产品(Application Specific Standard Parts,简称ASSP)、片上系统(System on Chip,简称SOC)、复杂可编程逻辑设备(Complex Programmable Logic Device,简称CPLD)等等。
在本公开的上下文中,机器可读介质可以是有形的介质,其可以包含或存储以供指令执行系统、装置或设备使用或与指令执行系统、装置或设备结合地使用的程序。机器可读介质可以是机器可读信号介质或机器可读储存介质。机器可读介质可以包括但不限于电子的、磁性的、光学的、电磁的、红外的、或半导体系统、装置或设备,或者上述内容的任何合适组合。机器可读存储介质的更具体示例会包括基于一个或多个线的电气 连接、便携式计算机盘、硬盘、随机存取存储器(RAM)、只读存储器(ROM)、可擦除可编程只读存储器(EPROM或快闪存储器)、光纤、便捷式紧凑盘只读存储器(CD-ROM)、光学储存设备、磁储存设备、或上述内容的任何合适组合。
第一方面,根据本公开的一个或多个实施例,提供了一种音乐生成方法,包括:
获取在预设时长内对第一用户进行采集得到的第一视频;
根据所述第一视频,确定所述第一用户在所述预设时长内执行的第一动作的动作信息,所述动作信息包括动作类型和动作时长;
根据所述第一动作的动作信息,生成所述预设时长对应的目标音乐。
根据本公开的一个或多个实施例,根据所述第一动作的动作信息,生成所述预设时长对应的目标音乐,包括:
根据所述第一动作的动作信息,确定所述第一动作对应的目标音阶的音阶信息,所述音阶信息包括音阶类型和音长;
根据所述目标音阶的音阶信息,生成所述目标音乐。
根据本公开的一个或多个实施例,根据所述第一动作的动作信息,确定所述第一动作对应的目标音阶的音阶信息,包括:
根据所述第一动作的动作类型和预设对应关系,确定所述目标音阶的音阶类型,所述预设对应关系用于指示不同动作类型与不同音阶类型之间的对应关系;
根据所述第一动作的动作时长,确定所述目标音阶的音长。
根据本公开的一个或多个实施例,获取在预设时长内对第一用户进行采集得到的第一视频之前,还包括:
获取所述目标音乐的生成模式,所述生成模式为基于自由创作的第一生成模式,或者,为基于参考音乐的第二生成模式;
相应的,根据所述目标音阶的音阶信息,生成所述目标音乐,包括:
根据所述目标音乐的生成模式,以及所述目标音阶的音阶信息,生成所述目标音乐。
根据本公开的一个或多个实施例,根据所述目标音乐的生成模式,以及所述目标音阶的音阶信息,生成所述目标音乐,包括:
若所述目标音乐的生成模式为所述第一生成模式,则根据所述目标音阶的音阶信息,生成所述目标音乐;或者,
若所述目标音乐的生成模式为所述第二生成模式,则从所述参考音乐中确定目标参考音阶,在所述目标音阶的音阶类型与所述目标参考音阶的音阶类型相同时,根据所述目标音阶的音阶信息,生成所述目标音乐。
根据本公开的一个或多个实施例,从所述参考音乐中确定目标参考音阶之前,还包括:
获取所述参考音乐;
对所述参考音乐进行音阶解析处理,得到所述参考音乐对应的音阶序列,所述音阶序列中包括多个参考音阶,所述多个参考音阶按照各自在所述参考音乐中出现顺序依次排列;
相应的,从所述参考音乐中确定目标参考音阶,包括:
按照所述音阶序列中各参考音阶的顺序,确定所述目标参考音阶。
根据本公开的一个或多个实施例,根据所述目标音阶的音阶信息,生成所述目标音乐,包括:
获取所述目标音乐的音色类型;
根据所述音色类型和所述目标音阶的音阶信息,生成所述目标音乐。
根据本公开的一个或多个实施例,根据所述目标音阶的音阶信息,生成所述目标音乐之后,还包括:
播放所述目标音乐。
根据本公开的一个或多个实施例,根据所述第一视频,确定所述第一用户在所述预设时长内执行的第一动作的动作信息,包括:
确定目标身体部位,所述目标身体部位包括所述第一用户的下述部位中的至少一种:肢体、手部、面部;
在所述第一视频的多个图像帧中分别检测所述目标身体部位的特征信息;
根据所述多个图像帧中检测得到的所述目标身体部位的特征信息,确定所述第一动作的动作信息。
第二方面,根据本公开的一个或多个实施例,提供了一种音乐生成装置,包括:
获取模块,用于获取在预设时长内对第一用户进行采集得到的第一视频;
确定模块,用于根据所述第一视频,确定所述第一用户在所述预设时长内执行的第一动作的动作信息,所述动作信息包括动作类型和动作时长;
生成模块,用于根据所述第一动作的动作信息,生成所述预设时长对应的目标音乐。
根据本公开的一个或多个实施例,所述生成模块具体用于:
根据所述第一动作的动作信息,确定所述第一动作对应的目标音阶的音阶信息,所述音阶信息包括音阶类型和音长;
根据所述目标音阶的音阶信息,生成所述目标音乐。
根据本公开的一个或多个实施例,所述生成模块具体用于:
根据所述第一动作的动作类型和预设对应关系,确定所述目标音阶的音阶类型,所述预设对应关系用于指示不同动作类型与不同音阶类型之间的对应关系;
根据所述第一动作的动作时长,确定所述目标音阶的音长。
根据本公开的一个或多个实施例,所述获取模块还用于,获取所述目标音乐的生成模式,所述生成模式为基于自由创作的第一生成模式,或者,为基于参考音乐的第二生成模式;
所述生成模块具体用于:根据所述目标音乐的生成模式,以及所述目标音阶的音阶信息,生成所述目标音乐。
根据本公开的一个或多个实施例,所述生成模块具体用于:
若所述目标音乐的生成模式为所述第一生成模式,则根据所述目标音阶的音阶信息,生成所述目标音乐;或者,
若所述目标音乐的生成模式为所述第二生成模式,则从所述参考音乐中确定目标参考音阶,在所述目标音阶的音阶类型与所述目标参考音阶的音阶类型相同时,根据所述目标音阶的音阶信息,生成所述目标音乐。
根据本公开的一个或多个实施例,所述获取模块还用于,获取所述参考音乐,对所 述参考音乐进行音阶解析处理,得到所述参考音乐对应的音阶序列,所述音阶序列中包括多个参考音阶,所述多个参考音阶按照各自在所述参考音乐中的出现顺序依次排列;
所述生成模块具体用于,按照所述音阶序列中各参考音阶的顺序,确定所述目标参考音阶。
根据本公开的一个或多个实施例,所述生成模块具体用于:
获取所述目标音乐的音色类型;
根据所述音色类型和所述目标音阶的音阶信息,生成所述目标音乐。
根据本公开的一个或多个实施例,所述装置还包括:
播放模块,用于播放所述目标音乐。
根据本公开的一个或多个实施例,所述确定模块具体用于:
确定目标身体部位,所述目标身体部位包括所述第一用户的下述部位中的至少一种:肢体、手部、面部;
在所述第一视频的多个图像帧中分别检测所述目标身体部位的特征信息;
根据所述多个图像帧中检测得到的所述目标身体部位的特征信息,确定所述第一动作的动作信息。
第三方面,根据本公开的一个或多个实施例,提供了一种电子设备,包括:至少一个处理器和存储器;
所述存储器存储计算机执行指令;
所述至少一个处理器执行所述存储器存储的计算机执行指令,使得所述至少一个处理器执行如上第一方面以及第一方面各种可能的实现方式所述的音乐生成方法。
第四方面,根据本公开的一个或多个实施例,提供了一种计算机可读存储介质,所述计算机可读存储介质中存储有计算机执行指令,当处理器执行所述计算机执行指令时,实现如上第一方面以及第一方面各种可能的实现方式所述的音乐生成方法。
第五方面,根据本公开的一个或多个实施例,提供了一种计算机程序产品,包括计算机程序,所述计算机程序被处理器执行时实现如第一方面以及第一方面各种可能的实现方式所述的音乐生成方法。
第六方面,根据本公开的一个或多个实施例,提供了一种计算机程序,所述计算机程序被处理器执行时实现如第一方面以及第一方面各种可能的实现方式所述的音乐生成方法。
以上描述仅为本公开的较佳实施例以及对所运用技术原理的说明。本领域技术人员应当理解,本公开中所涉及的公开范围,并不限于上述技术特征的特定组合而成的技术方案,同时也应涵盖在不脱离上述公开构思的情况下,由上述技术特征或其等同特征进行任意组合而形成的其它技术方案。例如上述特征与本公开中公开的(但不限于)具有类似功能的技术特征进行互相替换而形成的技术方案。
此外,虽然采用特定次序描绘了各操作,但是这不应当理解为要求这些操作以所示出的特定次序或以顺序次序来执行。在一定环境下,多任务和并行处理可能是有利的。同样地,虽然在上面论述中包含了若干具体实现细节,但是这些不应当被解释为对本公开的范围的限制。在单独的实施例的上下文中描述的某些特征还可以组合地实现在单个实施例中。相反地,在单个实施例的上下文中描述的各种特征也可以单独地或以任何合 适的子组合的方式实现在多个实施例中。
尽管已经采用特定于结构特征和/或方法逻辑动作的语言描述了本主题,但是应当理解所附权利要求书中所限定的主题未必局限于上面描述的特定特征或动作。相反,上面所描述的特定特征和动作仅仅是实现权利要求书的示例形式。

Claims (14)

  1. 一种音乐生成方法,包括:
    获取在预设时长内对第一用户进行采集得到的第一视频;
    根据所述第一视频,确定所述第一用户在所述预设时长内执行的第一动作的动作信息,所述动作信息包括动作类型和动作时长;
    根据所述第一动作的动作信息,生成所述预设时长对应的目标音乐。
  2. 根据权利要求1所述的方法,其中,根据所述第一动作的动作信息,生成所述预设时长对应的目标音乐,包括:
    根据所述第一动作的动作信息,确定所述第一动作对应的目标音阶的音阶信息,所述音阶信息包括音阶类型和音长;
    根据所述目标音阶的音阶信息,生成所述目标音乐。
  3. 根据权利要求2所述的方法,其中,根据所述第一动作的动作信息,确定所述第一动作对应的目标音阶的音阶信息,包括:
    根据所述第一动作的动作类型和预设对应关系,确定所述目标音阶的音阶类型,所述预设对应关系用于指示不同动作类型与不同音阶类型之间的对应关系;
    根据所述第一动作的动作时长,确定所述目标音阶的音长。
  4. 根据权利要求2或3所述的方法,其中,获取在预设时长内对第一用户进行采集得到的第一视频之前,还包括:
    获取所述目标音乐的生成模式,所述生成模式为基于自由创作的第一生成模式,或者,为基于参考音乐的第二生成模式;
    相应的,根据所述目标音阶的音阶信息,生成所述目标音乐,包括:
    根据所述目标音乐的生成模式,以及所述目标音阶的音阶信息,生成所述目标音乐。
  5. 根据权利要求4所述的方法,其中,根据所述目标音乐的生成模式,以及所述目标音阶的音阶信息,生成所述目标音乐,包括:
    若所述目标音乐的生成模式为所述第一生成模式,则根据所述目标音阶的音阶信息,生成所述目标音乐;或者,
    若所述目标音乐的生成模式为所述第二生成模式,则从所述参考音乐中确定目标参考音阶,在所述目标音阶的音阶类型与所述目标参考音阶的音阶类型相同时,根据所述目标音阶的音阶信息,生成所述目标音乐。
  6. 根据权利要求5所述的方法,其中,从所述参考音乐中确定目标参考音阶之前,还包括:
    获取所述参考音乐;
    对所述参考音乐进行音阶解析处理,得到所述参考音乐对应的音阶序列,所述音阶序列中包括多个参考音阶,所述多个参考音阶按照各自在所述参考音乐中出现顺序依次排列;
    相应的,从所述参考音乐中确定目标参考音阶,包括:
    按照所述音阶序列中各参考音阶的顺序,确定所述目标参考音阶。
  7. 根据权利要求2至6中任一项所述的方法,其中,根据所述目标音阶的音阶信息,生成所述目标音乐,包括:
    获取所述目标音乐的音色类型;
    根据所述音色类型和所述目标音阶的音阶信息,生成所述目标音乐。
  8. 根据权利要求2至7中任一项所述的方法,其中,根据所述目标音阶的音阶信息,生成所述目标音乐之后,还包括:
    播放所述目标音乐。
  9. 根据权利要求1至8中任一项所述的方法,其中,根据所述第一视频,确定所述第一用户在所述预设时长内执行的第一动作的动作信息,包括:
    确定目标身体部位,所述目标身体部位包括所述第一用户的下述部位中的至少一种:肢体、手部、面部;
    在所述第一视频的多个图像帧中分别检测所述目标身体部位的特征信息;
    根据所述多个图像帧中检测得到的所述目标身体部位的特征信息,确定所述第一动作的动作信息。
  10. 一种音乐生成装置,包括:
    获取模块,用于获取在预设时长内对第一用户进行采集得到的第一视频;
    确定模块,用于根据所述第一视频,确定所述第一用户在所述预设时长内执行的第一动作的动作信息,所述动作信息包括动作类型和动作时长;
    生成模块,用于根据所述第一动作的动作信息,生成所述预设时长对应的目标音乐。
  11. 一种电子设备,包括:处理器和存储器;
    所述存储器存储计算机执行指令;
    所述处理器执行所述计算机执行指令,实现如权利要求1至9中任一项所述的音乐生成方法。
  12. 一种计算机可读存储介质,所述计算机可读存储介质中存储有计算机执行指令,当处理器执行所述计算机执行指令时,实现如权利要求1至9中任一项所述的音乐生成方法。
  13. 一种计算机程序产品,包括计算机程序,所述计算机程序被处理器执行时实现如权利要求1至9中任一项所述的音乐生成方法。
  14. 一种计算机程序,所述计算机程序被处理器执行时实现如权利要求1至9中任一项所述的音乐生成方法。
PCT/CN2022/122334 2021-09-28 2022-09-28 音乐生成方法、装置、设备、存储介质及程序 Ceased WO2023051651A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US18/571,086 US20240290305A1 (en) 2021-09-28 2022-09-28 Music generation method and apparatus, device, storage medium, and program

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202111140485.0A CN115881064A (zh) 2021-09-28 2021-09-28 音乐生成方法、装置、设备、存储介质及程序
CN202111140485.0 2021-09-28

Publications (1)

Publication Number Publication Date
WO2023051651A1 true WO2023051651A1 (zh) 2023-04-06

Family

ID=85763262

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2022/122334 Ceased WO2023051651A1 (zh) 2021-09-28 2022-09-28 音乐生成方法、装置、设备、存储介质及程序

Country Status (3)

Country Link
US (1) US20240290305A1 (zh)
CN (1) CN115881064A (zh)
WO (1) WO2023051651A1 (zh)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN118283343A (zh) * 2024-03-27 2024-07-02 北京度友信息技术有限公司 基于配乐的视频生成方法、装置以及设备

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP4660994A4 (en) * 2024-04-24 2025-12-10 Beijing Zitiao Network Technology Co Ltd Music generation method, music generation apparatus, and computer readable storage medium
CN119668456A (zh) * 2024-12-13 2025-03-21 北京字跳网络技术有限公司 一种媒体内容生成方法、装置、设备、介质及程序产品

Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN103885663A (zh) * 2014-03-14 2014-06-25 深圳市东方拓宇科技有限公司 一种生成和播放音乐的方法及其对应终端
CN105700808A (zh) * 2016-02-18 2016-06-22 广东欧珀移动通信有限公司 音乐播放方法、装置及终端设备
CN109119057A (zh) * 2018-08-30 2019-01-01 Oppo广东移动通信有限公司 音乐创作方法、装置及存储介质和穿戴式设备
CN109413351A (zh) * 2018-10-26 2019-03-01 平安科技(深圳)有限公司 一种音乐生成方法及装置
CN110827789A (zh) * 2019-10-12 2020-02-21 平安科技(深圳)有限公司 音乐生成方法、电子装置及计算机可读存储介质
CN110874171A (zh) * 2018-08-31 2020-03-10 阿里巴巴集团控股有限公司 音频信息处理方法及装置
CN110944085A (zh) * 2019-11-12 2020-03-31 南京邮电大学 一种晃动智能手机产生音乐的方法
CN112698757A (zh) * 2020-12-25 2021-04-23 北京小米移动软件有限公司 界面交互方法、装置、终端设备及存储介质

Family Cites Families (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN108053831A (zh) * 2017-12-05 2018-05-18 广州酷狗计算机科技有限公司 音乐生成、播放、识别方法、装置及存储介质
WO2020154422A2 (en) * 2019-01-22 2020-07-30 Amper Music, Inc. Methods of and systems for automated music composition and generation
CN111276122B (zh) * 2020-01-14 2023-10-27 广州酷狗计算机科技有限公司 音频生成方法及装置、存储介质
WO2021159203A1 (en) * 2020-02-10 2021-08-19 1227997 B.C. Ltd. Artificial intelligence system & methodology to automatically perform and generate music & lyrics
CN112927665B (zh) * 2021-01-22 2022-08-30 咪咕音乐有限公司 创作方法、电子设备和计算机可读存储介质

Patent Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN103885663A (zh) * 2014-03-14 2014-06-25 深圳市东方拓宇科技有限公司 一种生成和播放音乐的方法及其对应终端
CN105700808A (zh) * 2016-02-18 2016-06-22 广东欧珀移动通信有限公司 音乐播放方法、装置及终端设备
CN109119057A (zh) * 2018-08-30 2019-01-01 Oppo广东移动通信有限公司 音乐创作方法、装置及存储介质和穿戴式设备
CN110874171A (zh) * 2018-08-31 2020-03-10 阿里巴巴集团控股有限公司 音频信息处理方法及装置
CN109413351A (zh) * 2018-10-26 2019-03-01 平安科技(深圳)有限公司 一种音乐生成方法及装置
CN110827789A (zh) * 2019-10-12 2020-02-21 平安科技(深圳)有限公司 音乐生成方法、电子装置及计算机可读存储介质
CN110944085A (zh) * 2019-11-12 2020-03-31 南京邮电大学 一种晃动智能手机产生音乐的方法
CN112698757A (zh) * 2020-12-25 2021-04-23 北京小米移动软件有限公司 界面交互方法、装置、终端设备及存储介质

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN118283343A (zh) * 2024-03-27 2024-07-02 北京度友信息技术有限公司 基于配乐的视频生成方法、装置以及设备
CN118283343B (zh) * 2024-03-27 2025-01-21 北京度友信息技术有限公司 基于配乐的视频生成方法、装置以及设备

Also Published As

Publication number Publication date
US20240290305A1 (en) 2024-08-29
CN115881064A (zh) 2023-03-31

Similar Documents

Publication Publication Date Title
US11158102B2 (en) Method and apparatus for processing information
US20210029305A1 (en) Method and apparatus for adding a video special effect, terminal device and storage medium
US11749246B2 (en) Systems and methods for music simulation via motion sensing
US11514923B2 (en) Method and device for processing music file, terminal and storage medium
WO2023051651A1 (zh) 音乐生成方法、装置、设备、存储介质及程序
CN111798821B (zh) 声音转换方法、装置、可读存储介质及电子设备
WO2020119150A1 (zh) 节奏点识别方法、装置、电子设备及存储介质
CN112380362B (zh) 基于用户交互的音乐播放方法、装置、设备及存储介质
JP2022505118A (ja) 画像処理方法、装置、ハードウェア装置
WO2021129628A1 (zh) 视频特效处理方法及装置
EP4604064A1 (en) Image processing method and apparatus, electronic device, and storage medium
CN115691544A (zh) 虚拟形象口型驱动模型的训练及其驱动方法、装置和设备
CN111833460A (zh) 增强现实的图像处理方法、装置、电子设备及存储介质
WO2020151491A1 (zh) 图像形变的控制方法、装置和硬件装置
WO2023061229A1 (zh) 视频生成方法及设备
CN115774539B (zh) 和声处理方法、装置、设备及介质
WO2023160713A1 (zh) 音乐生成方法、装置、设备、存储介质及程序
KR102637788B1 (ko) 이동 단말기용 악기 장치 및 그의 제어 방법 및 프로그램
EP4597488A1 (en) Audio processing method and apparatus, and electronic device
CN116737994A (zh) 视频、唱谱音频和曲谱同步播放方法、装置、设备和介质
WO2025194881A1 (zh) 音频播放方法、装置及终端设备
CN118567468A (zh) 交互方法、装置、设备及存储介质
CN107404581B (zh) 移动终端的乐器模拟方法、装置及存储介质和移动终端
CN116932812A (zh) 乐谱更新方法、装置、电子设备和计算机可读介质
CN115138062A (zh) 设备游戏互动方法、电子设备及可读存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 22875034

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 18571086

Country of ref document: US

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 05.07.2024)

122 Ep: pct application non-entry in european phase

Ref document number: 22875034

Country of ref document: EP

Kind code of ref document: A1