WO2022100103A1 - 歌词视频展示方法、装置、电子设备及计算机可读介质 - Google Patents

歌词视频展示方法、装置、电子设备及计算机可读介质 Download PDF

Info

Publication number
WO2022100103A1
WO2022100103A1 PCT/CN2021/102438 CN2021102438W WO2022100103A1 WO 2022100103 A1 WO2022100103 A1 WO 2022100103A1 CN 2021102438 W CN2021102438 W CN 2021102438W WO 2022100103 A1 WO2022100103 A1 WO 2022100103A1
Authority
WO
WIPO (PCT)
Prior art keywords
lyrics
target
target object
data
display
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2021/102438
Other languages
English (en)
French (fr)
Inventor
郑霓雯
孙磊
瞿佳
吴昊
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Zitiao Network Technology Co Ltd
Original Assignee
Beijing Zitiao Network Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Zitiao Network Technology Co Ltd filed Critical Beijing Zitiao Network Technology Co Ltd
Priority to US17/438,732 priority Critical patent/US12549801B2/en
Publication of WO2022100103A1 publication Critical patent/WO2022100103A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/431Generation of visual interfaces for content selection or interaction; Content or additional data rendering
    • H04N21/4312Generation of visual interfaces for content selection or interaction; Content or additional data rendering involving specific graphical features, e.g. screen layout, special fonts or colors, blinking icons, highlights or animations
    • H04N21/4316Generation of visual interfaces for content selection or interaction; Content or additional data rendering involving specific graphical features, e.g. screen layout, special fonts or colors, blinking icons, highlights or animations for displaying supplemental content in a region of the screen, e.g. an advertisement in a separate window
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H1/00Details of electrophonic musical instruments
    • G10H1/36Accompaniment arrangements
    • G10H1/361Recording/reproducing of accompaniment for use with an external source, e.g. karaoke systems
    • G10H1/368Recording/reproducing of accompaniment for use with an external source, e.g. karaoke systems displaying animated or moving pictures synchronized with the music or audio part
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/435Processing of additional data, e.g. decrypting of additional data, reconstructing software from modules extracted from the transport stream
    • H04N21/4355Processing of additional data, e.g. decrypting of additional data, reconstructing software from modules extracted from the transport stream involving reformatting operations of additional data, e.g. HTML pages on a television screen
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/439Processing of audio elementary streams
    • H04N21/4398Processing of audio elementary streams involving reformatting operations of audio signals
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/44Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
    • H04N21/4402Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving reformatting operations of video signals for household redistribution, storage or real-time display
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/442Monitoring of processes or resources, e.g. detecting the failure of a recording device, monitoring the downstream bandwidth, the number of times a movie has been viewed, the storage space available from the internal hard disk
    • H04N21/44213Monitoring of end-user related data
    • H04N21/44222Analytics of user selections, e.g. selection of programmes or purchase activity
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2220/00Input/output interfacing specifically adapted for electrophonic musical tools or instruments
    • G10H2220/005Non-interactive screen display of musical or status data
    • G10H2220/011Lyrics displays, e.g. for karaoke applications

Definitions

  • the present application relates to the technical field of video processing, and in particular, to a method, apparatus, electronic device, and computer-readable medium for displaying lyrics video.
  • the lyrics will appear under the video by scrolling or panning.
  • Some technologies also have the function of coloring the lyrics, but these are just simple overlays between the lyrics and the video.
  • the entry and exit are some basic special effects, and for the lyrics video with background, the lyrics and the background are completely separated from the background, there is no correlation, resulting in a poor experience for users to watch the lyrics video, and because the lyrics are just some simple mechanical Basic special effects and lyrics display form is relatively simple, resulting in poor user experience.
  • the purpose of this application is to solve at least one of the above-mentioned technical defects, especially in the prior art, there is a technical problem that the display form of lyrics is single, and the lyrics are completely separated from the video background, resulting in poor user experience.
  • a method for displaying lyrics video includes:
  • the multimedia data includes image data
  • the music data includes audio data and lyrics
  • the target lyrics are displayed within the preset range of the position of the target object in the target image, and based on the depth information of the target object, the display special effects of the target lyrics are adjusted, and the corresponding lyrics of the target lyrics are played at the same time. audio data.
  • a lyrics video display device comprising:
  • a data acquisition module for playing multimedia data and music data to be displayed based on the user's lyrics video display operation, wherein the multimedia data includes image data, and the music data includes audio data and lyrics;
  • a lyrics determination module used for determining a target time point, determining a target object corresponding to the target time point in the image data, and determining a target lyrics corresponding to the target time point in the lyrics;
  • the lyrics display module is used to display the target lyrics within the preset range of the position of the target object in the target image, and based on the depth information of the target object, adjust the display special effects of the target lyrics, and at the same time Play the audio data corresponding to the target lyrics.
  • an electronic device comprising:
  • processors one or more processors
  • one or more application programs wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs are configured to: execute The above lyrics video display method.
  • a computer-readable medium is provided, the readable storage is at least one instruction, at least one piece of program, code set or instruction set, the at least one instruction, the at least one piece of program, the code set or The instruction set is loaded and executed by the processor to implement the above-mentioned method for displaying lyrics video.
  • This embodiment of the present application plays the multimedia data and music data to be displayed based on the user's video display operation of lyrics, wherein the multimedia data includes image data, and the music data includes audio data and lyrics.
  • the target lyrics are displayed in the vicinity of the target object, and the display special effects of the lyrics are adjusted based on the depth information of the target object, and the lyrics are embedded in the real space of the image data, giving users a sensory experience of virtual display, user experience better.
  • FIG. 1 is a schematic flowchart of a method for displaying lyrics video provided by an embodiment of the present application
  • FIG. 2 is a schematic diagram of a display interface provided by an embodiment of the present application.
  • FIG. 3 is a schematic diagram of a target movement process provided by an embodiment of the present application.
  • FIG. 4 is a schematic flowchart of a method for acquiring multimedia data according to an embodiment of the present application
  • FIG. 5 is a schematic flowchart of a method for capturing multimedia data by a user according to an embodiment of the present application
  • FIG. 6 is a schematic diagram of a multimedia data selection interface provided by an embodiment of the present application.
  • FIG. 7 is a schematic flowchart of a method for generating a lyrics patch provided by an embodiment of the present application.
  • FIG. 8 is a schematic flowchart of a method for adjusting lyrics position provided by an embodiment of the present application.
  • FIG. 9 is a schematic diagram of a special effect adjustment provided by an embodiment of the present application.
  • FIG. 10 is a schematic diagram of adjusting the size of lyrics provided by an embodiment of the application.
  • FIG. 11 is a schematic structural diagram of a lyrics video display device provided by an embodiment of the application.
  • FIG. 12 is a schematic structural diagram of an electronic device according to an embodiment of the present application.
  • the term “including” and variations thereof are open-ended inclusions, ie, "including but not limited to”.
  • the term “based on” is “based at least in part on.”
  • the term “one embodiment” means “at least one embodiment”; the term “another embodiment” means “at least one additional embodiment”; the term “some embodiments” means “at least some embodiments”. Relevant definitions of other terms will be given in the description below.
  • the lyrics video display method, device, electronic device and computer-readable medium provided by the present application are intended to solve the above technical problems in the prior art.
  • An embodiment of the present application provides a method for displaying lyrics video, which is applied to a user terminal.
  • the user terminal may be a mobile terminal such as a mobile phone or a tablet computer.
  • An APP Application, application program
  • a function in the application can implement the lyrics video display method provided by the embodiment of the present application. As shown in FIG. 1 , the method includes:
  • Step S101 based on the user's lyric video display operation, play multimedia data and music data to be displayed, the multimedia data includes image data, and the music data includes audio data and lyrics;
  • Step S102 determine the target time point, determine the target object corresponding to the target time point in the image data, and determine the target lyrics corresponding to the target time point in the lyrics;
  • step S103 the target lyrics are displayed within a preset range of the position of the target object in the target image, and based on the depth information of the target object, the display special effects of the target lyrics are adjusted, and the audio data corresponding to the target lyrics is played at the same time.
  • the multimedia data includes image data and music data
  • the image data can be pictures, animations, and videos with depth information
  • the music data includes audio data and lyrics
  • the depth information is used to represent image data
  • the information of the distance between the target object and the image data acquisition device is different from the imaging ratio displayed in the image data of the target object with different distances from the image data acquisition device.
  • multimedia data and music to be synthesized when acquiring multimedia data and music to be synthesized, it may be multimedia data and music stored locally, multimedia data and music on the network, or multimedia data and music recorded by the user himself.
  • the multimedia data includes image data
  • the image data may be picture data or video data.
  • the method before adjusting the display special effects of the target lyrics based on the depth information of the target object, the method further includes:
  • the depth information of the target object in the image data is acquired, wherein the depth information of the target object is used for information representing the distance of the target object from the video capture device.
  • the depth information of the target object in the image data is obtained, wherein the target object refers to the object in the image data that is combined with the lyrics, which may be an object in the image data or a module area.
  • Depth information refers to the information used to indicate the distance between the target object and the video capture device. The deeper the depth, the farther the target object is from the video capture device, and the shallower the depth is, the closer the target object is to the video capture device. .
  • the display position of the target object is adjusted based on the display position of the target object when the image data is playing. showing the location.
  • the lyrics are displayed in the position corresponding to the target object, that is, the upper position in the image data, or the target object. The location near the object.
  • the special effects of the corresponding lyrics patches can be adjusted according to different depth information of different target objects.
  • the image data can be displayed through the first interface 201.
  • the data is video data
  • the video is played through the first interface 201
  • the music is played at the same time.
  • the position of the corresponding lyrics A1 is adjusted.
  • the target object is at the upper left in the video, then The position where the corresponding lyrics A1 appear is the upper left of the video.
  • the target object moves from the upper left of the video to the lower right of the video, the corresponding lyrics It will also move to the bottom right of the video with the target object.
  • its depth information becomes deeper and deeper, that is, the distance between the target object A and the video capture device is further and further, and the special effects of the lyrics A1 can be adjusted accordingly.
  • This embodiment of the present application plays the multimedia data and music data to be displayed based on the user's video display operation of lyrics, wherein the multimedia data includes image data, and the music data includes audio data and lyrics.
  • the target lyrics are displayed in the vicinity of the target object, and the display special effects of the lyrics are adjusted based on the depth information of the target object, and the lyrics are embedded in the real space of the image data, giving users a sensory experience of virtual display, user experience better.
  • the method further includes:
  • Step S401 receiving a user's multimedia data selection operation
  • Step S402 Determine multimedia data based on the multimedia data selection operation.
  • the user can select the desired multimedia data from the locally stored multimedia data.
  • the selection can be made through touch selection, or selection through voice, motion, or the like.
  • a specific embodiment is taken as an example, as shown in FIG. 5 , a user's multimedia data selection operation is received, and a multimedia data selection interface 501 is displayed, and the multimedia data selection interface 501 has available
  • the selected multimedia data 502 based on the user's selection operation, determines the multimedia data that the user wants to synthesize the lyric video.
  • the user can select the multimedia data that he wants to synthesize the lyric video, and the user can select different multimedia data according to his own needs, so as to improve the user experience.
  • the method further includes:
  • Step S601 based on the user's multimedia data capture operation, start the multimedia data capture device
  • Step S602 acquiring multimedia data captured by the multimedia data capturing device.
  • a user can capture multimedia data through a user terminal, and the user terminal should be provided with a multimedia data capture device, or an external multimedia data capture device, optionally an image capture device.
  • the user terminal is a mobile phone with a camera
  • the user's multimedia data capture operation is received through the mobile phone.
  • the multimedia data capture operation may be performed by the user.
  • Touch operation Based on the touch operation, the built-in camera of the collection is turned on, and image data is captured by the camera.
  • the user can shoot the video by himself, and synthesize the video, the video selection range is wider, and the user experience is good.
  • the target lyrics are displayed within a preset range of the position of the target object in the target image, including:
  • Step S701 generating a corresponding lyrics patch based on the target lyrics
  • Step S702 displaying the lyrics patch within a preset range of the position of the target object in the target image.
  • the lyric patch is a display form corresponding to the lyrics.
  • it can be a picture showing the lyrics, or a moving picture, which can display the content of the lyrics.
  • it can be a sentence of lyrics.
  • a lyric patch it can also be a long lyric corresponding to multiple lyric patches, or multiple short lyrics corresponding to a lyric patch. Different corresponding methods will present different display effects.
  • one lyric corresponds to one lyric patch.
  • generating a corresponding lyrics patch based on the target lyrics including:
  • Sentence or word segmentation processing is performed on the target lyrics to generate corresponding lyrics patches.
  • lyrics content is displayed in the lyrics patch, and when the corresponding lyrics patch is generated, it also includes:
  • the lyric content displayed in the lyric tile is typeset, including:
  • the lyrics content displayed in the lyrics patch is typeset according to at least one of the number of words, font, color, alignment, and display position of the lyrics in the lyrics patch.
  • a lyric patch when a lyric patch is generated according to the lyrics of the music, the lyrics in the music are obtained, and the lyrics are divided into sentences and words according to preset rules, such as forming a lyric into a lyric patch, or Certain words in a lyric form a lyric patch, and multiple lyrics can be combined to form a lyric patch.
  • the lyrics when forming a lyric patch, the lyrics can be typeset.
  • the lyrics of music are obtained, the lyrics are processed into sentences, and a lyrics patch is correspondingly generated for each lyrics.
  • the lyrics patch The content of the lyrics is displayed in the video; optional, the lyrics can be typeset according to the preset rules, such as specifying the number of characters displayed in each line, the font of the lyrics, the alignment, etc.
  • each lyric patch can display There are different numbers of lyrics, the display position of the lyrics in each lyrics patch can be different, and the alignment of the lyrics in each lyrics patch can be different.
  • the target lyrics are displayed within a preset range of the position of the target object in the target image, including:
  • Step S801 determining the target image data in the multimedia data based on the playing time period of the target lyrics in the music
  • Step S802 Adjust the display position of the target lyrics in the image data based on the position of the target object in the target image data, wherein the display position is within a preset range of the position of the target object in the target image data.
  • the time period in which the lyrics corresponding to the lyrics appear in the music must correspond to the time period in which the target object appears in the image data to ensure that the lyrics appear.
  • the target object can also appear in the image data.
  • the time period in which the corresponding lyrics appear in the music is 35S ⁇ 38S, then determine the target in the 35S ⁇ 38S in the image data. image, determine the target object B1 that has always existed in the target image, determine the target object B1 as the target object corresponding to the lyrics B, and determine the position of the lyrics patch based on the position information of the target object B1 in the image data.
  • the location where the lyrics appear may also be a region near the display location of the target object corresponding to the lyrics.
  • the corresponding relationship between the lyrics patch and the target object is determined by the time when the lyrics corresponding to the lyrics patch appear in the music and the time when the target object appears in the image data, so as to ensure that the lyrics can appear in the data image during playback, Users can see the lyrics.
  • adjusting the display special effects of the target lyrics based on the depth information of the target object in the image data including:
  • the display size of the target lyrics is adjusted.
  • adjusting the display special effect of the lyrics patch may be adjusting the display size of the lyrics.
  • the target object C1 is in the process of playing the image data.
  • the distance from the image capturing device is getting farther and the depth is getting deeper and deeper, the display size of the lyrics C corresponding to the target object C1 can be adjusted to be larger and larger.
  • adjusting the display special effects of the target lyrics based on the depth information of the target object including:
  • the adjustment method of the special effect may also be to adjust the display color of the corresponding lyrics patch according to the depth information of the target object, and the like.
  • the display size of the lyrics corresponding to the target object is adjusted according to the depth information of the target object, so as to provide the user with a visual experience that is far small and near large.
  • the image data includes video data and/or panoramic image data with depth-of-field information, wherein, when the image data is panoramic image data with depth-of-field information, a pre-determined location of the target object in the target image Display target lyrics within a set range, including:
  • the display positions of the target lyrics corresponding to the different target objects in the panoramic image data are adjusted respectively.
  • the panoramic image 1001 there are a target object D1 and a target object E1, wherein the target object D1 is at the lower left of the panoramic image. If the target object E1 is in the middle of the panoramic image, the target object D1 corresponds to the lyrics D, and the target object E1 corresponds to the lyrics E.
  • the positions of the lyrics D and E are adjusted according to the positions of the target objects D1 and E1 respectively, wherein the target object D1
  • the depth information of the target object E1 is relatively shallow, and the depth information of the target object E1 is relatively deep, you can adjust the lyric D to display larger, and the lyric E to display smaller.
  • the size of the corresponding lyrics is adjusted by the position and depth information of different target objects in the panoramic image, so as to give the user a real spatial experience.
  • the embodiment of the present application generates a corresponding lyric patch based on the lyrics of the music, obtains the depth information of the target object in the image data based on the image data in the multimedia data, and adjusts the target object based on the position information of the target object in the image data.
  • the display position of the corresponding lyrics patch in the image data, and based on the depth information of the target object in the image data, the display special effects of the corresponding lyrics patch are adjusted, and the lyrics are embedded in the real space of the image data.
  • a sensory experience of virtual display, the user experience is better.
  • the lyrics video display device 110 may include: a data acquisition module 1110, a lyrics determination module 1120, and a lyrics display module 1130, wherein,
  • the data acquisition module 1110 is used to play multimedia data and music data to be displayed based on the user's lyrics video display operation, the multimedia data includes image data, and the music data includes audio data and lyrics;
  • the lyrics determination module 1120 is used to determine the target time point, determine the target object corresponding to the target time point in the image data, and determine the target lyrics corresponding to the target time point in the lyrics;
  • the lyrics display module 1130 is used to display the target lyrics within the preset range of the position of the target object in the target image, adjust the display special effects of the target lyrics based on the depth information of the target object, and play the audio data corresponding to the target lyrics at the same time.
  • the data acquisition module 1110 when the data acquisition module 1110 acquires multimedia data to be synthesized, it can be used to:
  • the data acquisition module 1110 when the data acquisition module 1110 acquires multimedia data to be synthesized, it can be used to:
  • the multimedia data captured by the multimedia data capture device is acquired.
  • the lyrics display module 1130 when the lyrics display module 1130 displays the target lyrics within a preset range of the position of the target object in the target image, it can be used to: generate a corresponding lyrics patch based on the target lyrics; Lyrics tiles are displayed within a preset range of positions.
  • the lyrics display module 1130 may be used to: perform sentence or word segmentation processing on the target lyrics to generate the corresponding lyrics patch.
  • lyric content is displayed in the lyric patch, and when the lyrics display module 1130 generates the corresponding lyric patch, it can also be used to: typeset the lyric content displayed in the lyric patch.
  • the lyrics display module 1130 when the lyrics display module 1130 typesets the lyrics content displayed in the lyrics patch, it can be used to: according to at least one of the number of words, font, color, alignment, and display position of the lyrics in the lyrics patch. One, typesetting the lyrics content displayed in the lyrics patch.
  • the lyrics display module 1130 adjusts the display special effects of the target lyrics based on the depth information of the target object in the image data, it can be used to:
  • the lyrics display module 1130 when the lyrics display module 1130 displays the target lyrics within a preset range of the position of the target object in the target image, it can be used to:
  • the display position of the target lyrics in the image data is adjusted based on the position of the target object in the target image data, wherein the display position is within a preset range of the position of the target object in the target image data.
  • the lyrics display module 1130 when the lyrics display module 1130 adjusts the display position of the target lyrics in the image data based on the position of the target object in the image data, it can be used to:
  • the display size of the target lyrics is adjusted.
  • the image data includes video data and/or panoramic image data with depth-of-field information, wherein, when the image data is panoramic image data with depth-of-field information, the lyrics display module 1130 displays the location of the target object in the target image.
  • the lyrics display module 1130 displays the location of the target object in the target image.
  • the display positions of the lyrics patches corresponding to the different target objects in the panoramic image data are adjusted respectively.
  • the lyrics display module 1130 may further be used to: before adjusting the display special effects of the target lyrics based on the depth information of the target object:
  • Acquire depth information of the target object in the image data wherein the depth information of the target object is used for information representing the distance of the target object from the video capture device.
  • the embodiment of the present application generates a corresponding lyric patch based on the lyrics of the music, obtains the depth information of the target object in the image data based on the image data in the multimedia data, and adjusts the target object based on the position information of the target object in the image data.
  • the display position of the corresponding lyrics patch in the image data, and based on the depth information of the target object in the image data, the display special effects of the corresponding lyrics patch are adjusted, and the lyrics are embedded in the real space of the image data.
  • a sensory experience of virtual display, the user experience is better.
  • Terminal devices in the embodiments of the present application may include, but are not limited to, such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablets), PMPs (portable multimedia players), vehicle-mounted terminals (such as mobile terminals such as in-vehicle navigation terminals), etc., and stationary terminals such as digital TVs, desktop computers, and the like.
  • the electronic device shown in FIG. 12 is only an example, and should not impose any limitations on the functions and scope of use of the embodiments of the present application.
  • the electronic device includes: a memory and a processor, wherein the processor here may be referred to as the processing device 1201 hereinafter, and the memory may include a read-only memory (ROM) 1202, a random access memory (RAM) 1203, and a storage device 1208 hereinafter. at least one of the following:
  • an electronic device 1200 may include a processing device (eg, a central processing unit, a graphics processor, etc.) 1201 that may be loaded into random access according to a program stored in a read only memory (ROM) 1202 or from a storage device 1208 Various appropriate actions and processes are executed by the programs in the memory (RAM) 1203 . In the RAM 1203, various programs and data required for the operation of the electronic device 1200 are also stored.
  • the processing device 1201, the ROM 1202, and the RAM 1203 are connected to each other through a bus 1204.
  • An input/output (I/O) interface 1205 is also connected to bus 1204 .
  • the following devices may be connected to the I/O interface 1205: input devices 1206 including, for example, a touch screen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; including, for example, a liquid crystal display (LCD), speakers, vibration An output device 1207 of a computer, etc.; a storage device 1208 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1209. Communication means 1209 may allow electronic device 1200 to communicate wirelessly or by wire with other devices to exchange data.
  • FIG. 12 shows an electronic device 1200 having various means, it should be understood that not all of the illustrated means are required to be implemented or provided. More or fewer devices may alternatively be implemented or provided.
  • embodiments of the present application include a computer program product comprising a computer program carried on a non-transitory computer readable medium, the computer program containing program code for performing the method illustrated in the flowchart.
  • the computer program may be downloaded and installed from the network via the communication device 1209, or from the storage device 1208, or from the ROM 1202.
  • the processing device 1201 the above-mentioned functions defined in the methods of the embodiments of the present application are executed.
  • the computer-readable medium mentioned above in the present application may be a computer-readable signal medium or a computer-readable medium, or any combination of the above two.
  • the computer readable medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or a combination of any of the above. More specific examples of computer readable media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read only memory (ROM), erasable programmable Read only memory (EPROM or flash memory), fiber optics, portable compact disk read only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
  • a computer-readable medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
  • a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code therein. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing.
  • a computer-readable signal medium can also be any computer-readable medium other than a computer-readable medium that can transmit, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device.
  • Program code embodied on a computer readable medium may be transmitted using any suitable medium including, but not limited to, electrical wire, optical fiber cable, RF (radio frequency), etc., or any suitable combination of the foregoing.
  • the client and server can use any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol) to communicate, and can communicate with digital data in any form or medium Communication (eg, a communication network) interconnects.
  • HTTP HyperText Transfer Protocol
  • Examples of communication networks include local area networks (“LAN”), wide area networks (“WAN”), the Internet (eg, the Internet), and peer-to-peer networks (eg, ad hoc peer-to-peer networks), as well as any currently known or future development network of.
  • the above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or may exist alone without being assembled into the electronic device.
  • the above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed by the electronic device, the electronic device: acquires multimedia data and music to be synthesized, and the multimedia data includes image data; acquires an image The depth information of the target object in the data; the lyrics patch is generated according to the lyrics of the music; the display position of the lyrics patch in the image data is adjusted based on the position of the target object in the image data; based on the depth information of the target object in the image data, adjustment Display special effects of lyrics patch, and generate lyrics video.
  • Computer program code for performing the operations of the present application may be written in one or more programming languages, including but not limited to object-oriented programming languages—such as Java, Smalltalk, C++, and This includes conventional procedural programming languages - such as the "C" language or similar programming languages.
  • the program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server.
  • the remote computer may be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (eg, using an Internet service provider through Internet connection).
  • LAN local area network
  • WAN wide area network
  • each block in the flowchart or block diagrams may represent a module, segment, or portion of code that contains one or more logical functions for implementing the specified functions executable instructions.
  • the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved.
  • each block of the block diagrams and/or flowchart illustrations, and combinations of blocks in the block diagrams and/or flowchart illustrations can be implemented in dedicated hardware-based systems that perform the specified functions or operations , or can be implemented in a combination of dedicated hardware and computer instructions.
  • modules or units involved in the embodiments of the present application may be implemented in a software manner, and may also be implemented in a hardware manner.
  • exemplary types of hardware logic components include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chips (SOCs), Complex Programmable Logical Devices (CPLDs) and more.
  • FPGAs Field Programmable Gate Arrays
  • ASICs Application Specific Integrated Circuits
  • ASSPs Application Specific Standard Products
  • SOCs Systems on Chips
  • CPLDs Complex Programmable Logical Devices
  • a machine-readable medium may be a tangible medium that may contain or store the program for use by or in connection with the instruction execution system, apparatus or device.
  • the machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium.
  • Machine-readable media may include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or devices, or any suitable combination of the foregoing.
  • machine-readable storage media would include one or more wire-based electrical connections, portable computer disks, hard disks, random access memory (RAM), read only memory (ROM), erasable programmable read only memory (EPROM or flash memory), fiber optics, compact disk read only memory (CD-ROM), optical storage, magnetic storage, or any suitable combination of the foregoing.
  • RAM random access memory
  • ROM read only memory
  • EPROM or flash memory erasable programmable read only memory
  • CD-ROM compact disk read only memory
  • magnetic storage or any suitable combination of the foregoing.
  • a method for displaying lyrics video comprising:
  • the multimedia data and music data to be displayed the multimedia data includes image data, and the music data includes audio data and lyrics; determine the target time point, and determine the target object corresponding to the target time point in the image data , determine the target lyrics corresponding to the target time point in the lyrics; display the target lyrics within the preset range of the position of the target object in the target image, and adjust the display special effects of the target lyrics based on the depth information of the target object, and play the target lyrics at the same time corresponding audio data.
  • the method further includes: receiving a user's multimedia data selection operation; and determining multimedia data based on the multimedia data selection operation.
  • the method further includes: based on the user's multimedia data capture operation, enabling a multimedia data capture device; and acquiring multimedia data captured by the multimedia data capture device. data.
  • displaying the target lyrics within a preset range of the position of the target object in the target image includes: generating a corresponding lyric tile based on the target lyrics; displaying within the preset range of the position of the target object in the target image Lyrics patch.
  • generating a corresponding lyrics patch based on the target lyrics including:
  • Sentence or word segmentation processing is performed on the target lyrics to generate corresponding lyrics patches.
  • lyrics content is displayed in the lyrics patch, and when the corresponding lyrics patch is generated, it also includes:
  • the lyric content displayed in the lyric tile is typeset, including:
  • the lyrics content displayed in the lyrics patch is typeset according to at least one of the number of words, font, color, alignment, and display position of the lyrics in the lyrics patch.
  • adjusting the display special effects of the target lyrics based on the depth information of the target object including:
  • displaying the target lyrics within a preset range of the position of the target object in the target image includes: determining the target image data in the multimedia data based on the playing time period of the target lyrics in the music; The position of the target object in the target image data adjusts the display position of the target lyrics in the image data, wherein the display position is within a preset range of the position of the target object in the target image data.
  • adjusting the display special effects of the target lyrics based on the depth information of the target object including:
  • the display size of the target lyrics is adjusted.
  • the image data includes video data and/or panoramic image data with depth-of-field information, wherein, when the image data is panoramic image data with depth-of-field information, a pre-determined location of the target object in the target image Displaying the target lyrics within the set range includes: based on the positions of different target objects in the panoramic image data, respectively adjusting the display positions of the target lyrics corresponding to the different target objects in the panoramic image data.
  • the method before adjusting the display effect of the target lyrics based on the depth information of the target object, the method further includes: acquiring depth information of the target object in the image data, wherein the depth information of the target object is used to indicate the distance from the target object to the video capture. Information about the distance of the device.
  • a video display device for lyrics comprising:
  • the data acquisition module is used to play multimedia data and music data to be displayed based on the user's lyrics video display operation, the multimedia data includes image data, and the music data includes audio data and lyrics;
  • the lyrics determination module is used to determine the target time point, determine the target object corresponding to the target time point in the image data, determine the target lyrics corresponding to the target time point in the lyrics;
  • the lyrics display module is used to display the target lyrics within the preset range of the position of the target object in the target image, and based on the depth information of the target object, adjust the display special effects of the target lyrics, and play the audio data corresponding to the target lyrics at the same time.
  • the data acquisition module when acquiring the multimedia data to be synthesized, may be configured to: receive a user's multimedia data selection operation; and determine the multimedia data based on the multimedia data selection operation.
  • the data acquisition module when the data acquisition module acquires the multimedia data to be synthesized, it can be used to: start the multimedia data capture device based on the user's multimedia data capture operation; and acquire the multimedia data captured by the multimedia data capture device.
  • the lyrics display module when the lyrics display module displays the target lyrics within a preset range of the position of the target object in the target image, it can be used to: generate a corresponding lyrics patch based on the target lyrics; display the position of the target object in the target image The lyric patch is displayed within the preset range of .
  • the lyrics display module when the lyrics display module generates a corresponding lyrics patch based on the target lyrics, it can be used to:
  • Sentence or word segmentation processing is performed on the target lyrics to generate corresponding lyrics patches.
  • the lyric content is displayed in the lyric patch, and when the lyrics display module generates the corresponding lyric patch, the lyric display module can also be used to: typeset the lyric content displayed in the lyric patch.
  • the lyrics display module when the lyrics display module typesets the lyrics content displayed in the lyrics patch, it can be used to: according to at least one of the number of words, font, color, alignment, and display position of the lyrics in the lyrics patch Typesetting the lyrics content displayed in the lyrics patch.
  • the lyrics display module when adjusting the display special effects of the target lyrics based on the depth information of the target object in the image data, can be used to:
  • the lyrics display module when the lyrics display module displays the target lyrics within a preset range of the position of the target object in the target image, it can be used to:
  • the display position of the target lyrics in the image data is adjusted based on the position of the target object in the target image data, wherein the display position is within a preset range of the position of the target object in the target image data.
  • the lyrics display module when the lyrics display module adjusts the display position of the lyrics tile in the image data based on the position of the target object in the image data, it can be used to:
  • the display size of the target lyrics is adjusted.
  • the lyrics display module when adjusting the display special effects of the lyrics patch based on the depth information of the target object in the image data, can be used to:
  • the image data includes video data and/or panoramic image data with depth-of-field information
  • the lyrics display module displays the location of the target object in the target image.
  • the display positions of the lyrics patches corresponding to the different target objects in the panoramic image data are adjusted respectively.
  • the lyrics display module before adjusting the display special effects of the target lyrics based on the depth information of the target object, may also be used to:
  • Acquire depth information of the target object in the image data wherein the depth information of the target object is used for information representing the distance of the target object from the video capture device.
  • an electronic device comprising: one or more processors; a memory; one or more application programs, wherein the one or more application programs are stored in The one or more programs are stored in the memory and configured to be executed by one or more processors, and the one or more programs are configured to: execute the lyrics video display method of the above embodiment.
  • a computer-readable medium stores at least one instruction, at least one piece of program, code set or instruction set, the at least one instruction, the at least one A piece of program, the code set or the instruction set is loaded and executed by the processor to implement the lyrics video display method described in the above embodiment.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Social Psychology (AREA)
  • Business, Economics & Management (AREA)
  • Marketing (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Databases & Information Systems (AREA)
  • Acoustics & Sound (AREA)
  • Physics & Mathematics (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

本申请提供一种歌词视频展示方法、装置、电子设备及计算机可读介质,该方法包括:基于用户的歌词视频展示操作,播放待展示的多媒体数据和音乐数据,确定图像数据中与目标时间点对应的目标对象,确定歌词中与目标时间点对应的目标歌词;在目标图像中目标对象所在的位置的预设范围内展示目标歌词,基于目标对象的深度信息,调节目标歌词的显示特效,播放待展示的多媒体数据和音乐数据。

Description

歌词视频展示方法、装置、电子设备及计算机可读介质
本申请要求北京字跳网络技术有限公司于2020年11月10日提交的,申请名称为“歌词视频展示方法、装置、电子设备及计算机可读介质”的、中国专利申请号为“202011247956.3”的优先权,该中国专利申请的全部内容通过引用结合在本申请中。
技术领域
本申请涉及视频处理技术领域,具体而言,本申请涉及一种歌词视频展示方法、装置、电子设备及计算机可读介质。
背景技术
随着视频技术的发展,人们对音乐视频的要求也越来越高,音乐视频中出现音乐歌词已经是一个十分常见的功能。
现有的音乐视频中,在播放音乐时会在视频下方滚动或平移的出现歌词,有些技术还会有给歌词进行染色的功能,但是这些都只是简单的将歌词与视频进行简单的叠加,歌词的进场、出场都是一些基础特效,并且对于有背景的歌词视频,歌词与背景之间完全脱离,没有相关性,导致用户观看歌词视频的体验不佳,并且由于歌词仅仅只是一些简单机械的基础特效,歌词展现形式比较单一,导致用户体验不佳。
由此可见,现有技术中存在歌词展示形式单一,且歌词与视频背景完全脱离,导致用户体验不佳的技术问题,需要解决。
技术问题
本申请的目的旨在至少能解决上述的技术缺陷之一,特别是现有技术中存在歌词展示形式单一,且歌词与视频背景完全脱离,导致用户体验不佳的技术问题。
技术解决方案
第一方面,提供了一种歌词视频展示方法,该方法包括:
基于用户的歌词视频展示操作,播放待展示的多媒体数据和音乐数据,所述多媒体数据中包含图像数据,所述音乐数据包括音频数据和歌词;
确定目标时间点,确定所述图像数据中与所述目标时间点对应的目标对象,确定所述歌词中所述目标时间点对应的目标歌词;
在所述目标图像中所述目标对象所在的位置的预设范围内展示所述目标歌词,并基于所述目标对象的深度信息,调节所述目标歌词的显示特效,同时播放所述目标歌词对应的音频数据。
第二方面,提供了一种歌词视频展示装置,该装置包括:
数据获取模块,用于基于用户的歌词视频展示操作,播放待展示的多媒体数据和音乐数据,所述多媒体数据中包含图像数据,所述音乐数据包括音频数据和歌词;
歌词确定模块,用于确定目标时间点,确定所述图像数据中与所述目标时间点对应的目标对象,确定所述歌词中所述目标时间点对应的目标歌词;
歌词展示模块,用于在所述目标图像中所述目标对象所在的位置的预设范围内展示所述目标歌词,并基于所述目标对象的深度信息,调节所述目标歌词的显示特效,同时播放所述目标歌词对应的音频数据。第三方面,提供了一种电子设备,该电子设备包括:
一个或多个处理器;
存储器;
一个或多个应用程序,其中所述一个或多个应用程序被存储在所述存储器中并被配置为由所述一个或多个处理器执行,所述一个或多个程序配置用于:执行上述的歌词视频展示方法。
第四方面,提供了一种计算机可读介质,所述可读存储有至少一条指令、至少一段程序、代码集或指令集,所述至少一条指令、所述至少一段程序、所述代码集或指令集由所述处理器加载并执行以实现上述的歌词视频展示方法。
有益效果
本申请实施例基于用户的歌词视频展示操作,播放待展示的多媒体数据和音乐数据,其中,多媒体数据包括图像数据,音乐数据包括音频数据和歌词,同时,确定图像数据中的目标对象和歌词中的目标歌词,并在目标对象的附近展示目标歌词,并基于目标对象的深度信息调节歌词的显示特效,将歌词嵌入到图像数据的现实空间中,给用户一种虚拟显示的感官体验,用户体验更好。
附图说明
为了更清楚地说明本申请实施例中的技术方案,下面将对本申请实施例描述中所需要使用的附图作简单地介绍。
图1为本申请实施例提供的一种歌词视频展示方法的流程示意图;
图2为本申请实施例提供的一种显示界面示意图;
图3为本申请实施例提供的一种目标移动过程示意图;
图4为本申请实施例提供的一种获取多媒体数据方法的流程示意图;
图5为本申请实施例提供的一种用户捕捉多媒体数据的方法的流程示意图;
图6为本申请实施例提供的一种多媒体数据选择界面示意图;
图7为本申请实施例提供的一种生成歌词贴片方法的流程示意图;
图8为本申请实施例提供的一种调节歌词位置的方法的流程示意图;
图9为本申请实施例提供的一种特效调节示意图;
图10为本申请实施例提供的一种调节歌词大小的示意图;
图11为本申请实施例提供的一种歌词视频展示装置的结构示意图;
图12为本申请实施例提供的一种电子设备的结构示意图。
结合附图并参考以下具体实施方式,本申请各实施例的上述和其他特征、优点及方面将变得更加明显。贯穿附图中,相同或相似的附图标记表示相同或相似的元素。应当理解附图是示意性的,原件和元素不一定按照比例绘制。
本发明的实施方式
下面将参照附图更详细地描述本申请的实施例。虽然附图中显示了本申请的某些实施例,然而应当理解的是,本申请可以通过各种形式来实现,而且不应该被解释为限于这里阐述的实施例,相反提供这些实施例是为了更加透彻和完整地理解本申请。应当理解的是,本申请的附图及实施例仅用于示例性作用,并非用于限制本申请的保护范围。
应当理解,本申请的方法实施方式中记载的各个步骤可以按照不同的顺序执行,和/或并行执行。此外,方法实施方式可以包括附加的步骤和/或省略执行示出的步骤。本申请的范围在此方面不受限制。
本文使用的术语“包括”及其变形是开放性包括,即“包括但不限于”。术语“基于”是“至少部分地基于”。术语“一个实施例”表示“至少一个实施例”;术语“另一实施例”表示“至少一个另外的实施例”;术语“一些实施例”表示“至少一些实施例”。其他术语的相关定义将在下文描述中给出。
需要注意,本申请中提及的“第一”、“第二”等概念仅用于对装置、模块或单元进行区分,并非用于限定这些装置、模块或单元一定为不同的装置、模块或单元,也并非用于限定这些装置、模块或单元所执行的功能的顺序或者相互依存关系。
需要注意,本申请中提及的“一个”、“多个”的修饰是示意性而非限制性的,本领域技术人员应当理解,除非在上下文另有明确指出,否则应该理解为“一个或多个”。
本申请实施方式中的多个装置之间所交互的消息或者信息的名称仅用于说明性的目的,而并不是用于对这些消息或信息的范围进行限制。
本申请提供的歌词视频展示方法、装置、电子设备和计算机可读介质,旨在解决现有技术的如上技术问题。
下面以具体地实施例对本申请的技术方案以及本申请的技术方案如何解决上述技术问题进行详细说明。下面这几个具体的实施例可以相互结合,对于相同或相似的概念或过程可能在某些实施例中不再赘述。下面将结合附图,对本申请的实施例进行描述。
本申请实施例中提供了一种歌词视频展示方法,应用于用户终端,该用户终端可以是手机、平板电脑等移动终端,该用户终端中安装有APP(Application,应用程序),通过该应用程序或者该应用程序中的某个功能可以实现本申请实施例提供的歌词视频展 示方法,如图1所示,该方法包括:
步骤S101,基于用户的歌词视频展示操作,播放待展示的多媒体数据和音乐数据,多媒体数据中包含图像数据,音乐数据包括音频数据和歌词;
步骤S102,确定目标时间点,确定图像数据中与目标时间点对应的目标对象,确定歌词中与目标时间点对应的目标歌词;
步骤S103,在目标图像中目标对象所在的位置的预设范围内展示目标歌词,并基于目标对象的深度信息,调节目标歌词的显示特效,同时播放目标歌词对应的音频数据。
在本申请实施例中,多媒体数据包括图像数据和音乐数据,其中图像数据可以是带有景深信息的图片、动图以及视频等,音乐数据包括音频数据和歌词,深度信息是用于表示图像数据中目标对象距离图像数据采集装置的距离的信息,与图像数据采集装置距离不同的目标对象在图像数据中显示的成像比例不同。
对于本申请实施例,在获取待合成的多媒体数据和音乐时,可以是获取本地存储的多媒体数据和音乐,也可以是网络上的多媒体数据和音乐,还可以是用户自己录制的多媒体数据和音乐,其中,该多媒体数据中包含有图像数据,可选的,该图像数据可以是图片数据,也可以是视频数据。
在一些实施例中,在基于目标对象的深度信息,调节目标歌词的显示特效之前,还包括:
获取图像数据中目标对象的深度信息,其中,目标对象的深度信息用于表示目标对象距离视频捕捉装置的距离的信息。
在本申请实施例中,获取图像数据中目标对象的深度信息,其中,目标对象是指图像数据中与歌词结合的对象,可以是图像数据中某个物体,也可以是模块区域,目标对象的深度信息是指用于表示该目标对象距离视频捕捉装置的距离的信息,深度越深,表示该目标对象距离视频捕捉装置越远,深度越浅,表示该目标对象距离视频捕捉装置的距离越近。
在本申请实施例中,在基于目标对象在图像数据中的位置调节歌词在图像数据中的显示位置时,是基于图像数据在播放时,目标对象的显示位置调节与该目标对象对应的歌词的显示位置的。可选的,如通过第一界面显示图像数据时,目标对象显示在图像数据中上部的位置,则将歌词显示在该目标对象对应的位置,即图像数据中上部的位置,也可以是该目标对象附近的位置。
本申请实施例中,在基于目标对象在图像数据中的深度信息,调节歌词的显示特效时,调节特效时,可以根据不同目标对象的不同的深度信息,调节对应的歌词贴片的特效。
对于本申请实施例,为方便说明,以一个具体实施例为例说明本申请实施例提供的展示效果,如图2所示,可以通过第一界面201对图像数据进行展示,可选的,图像数据为视频数据,通过第一界面201播放视频,同时播放音乐,基于视频中目标对象A的位置,调节对应的歌词A1的位置,如图2所示,目标对象在视频中的左上方,则对应的歌词A1出现的位置为视频的左上方,可选的,如图3所示,在视频和音乐的播放过程中,目标对象从视频的左上方移动到视频的右下方,则对应的歌词也会随着目标对象移动到视频的右下方。可选的,目标对象A在移动过程中,其深度信息越来越深,即目标对象A距离视频捕捉装置的距离越来越远,可以对应的调节歌词A1的特效。
本申请实施例基于用户的歌词视频展示操作,播放待展示的多媒体数据和音乐数据,其中,多媒体数据包括图像数据,音乐数据包括音频数据和歌词,同时,确定图像数据中的目标对象和歌词中的目标歌词,并在目标对象的附近展示目标歌词,并基于目标对象的深度信息调节歌词的显示特效,将歌词嵌入到图像数据的现实空间中,给用户一种虚拟显示的感官体验,用户体验更好。
在一些实施例中,如图4所示,基于用户的歌词视频展示操作,播放待展示的多媒体数据和音乐数据之前,还包括:
步骤S401,接收用户的多媒体数据选择操作;
步骤S402,基于多媒体数据选择操作确定多媒体数据。
在本申请实施例中,用户可以从本地存储的多媒体数据中选择想要的多媒体数据,可选的,可以是通过触控选择的方式选择,也可以是通过语音、动作等方式进行选择。
对于本申请实施例,为方便说明,以一个具体实施例为例,如图5所示,接收用户的多媒体数据选择操作,并显示多媒体数据选择界面501,该多媒体数据选择界面501中有可供选择的多媒体数据502,基于用户的选择操作,确定用户想要合成歌词视频的多媒体数据。
本申请实施例中用户可以自己选择想要合成歌词视频的多媒体数据,用户可以根据自己需要,选择不同的多媒体数据,提升用户体验。
在一些实施例中,如图6所示,基于用户的歌词视频展示操作,播放待展示的多媒体数据和音乐数据之前,还包括:
步骤S601,基于用户的多媒体数据捕捉操作,开启多媒体数据捕捉装置;
步骤S602,获取多媒体数据捕捉装置捕捉的多媒体数据。
在本申请实施例中,用户可以通过用户终端捕捉多媒体数据,该用户终端应该设置有多媒体数据捕捉装置、或外界有多媒体数据捕捉装置,可选的,为图像捕捉装置。
对于本申请实施例,为方便说明,以一个具体实施例为例,用户终端为带有摄像头 的手机,通过手机接收用户的多媒体数据捕捉操作,可选的,该多媒体数据捕捉操作可以是用户的触控操作,基于该触控操作,开启收集自带的摄像头,并通过摄像头捕捉图像数据。
本申请实施例中用户可以自己拍摄视频,并对视频进行合成,视频选择范围更广,用户体验佳。
在一些实施例中,如图7所示,在目标图像中目标对象所在的位置的预设范围内展示目标歌词,包括:
步骤S701,基于目标歌词生成对应的歌词贴片;
步骤S702,在目标图像中目标对象的位置的预设范围内展示歌词贴片。
在本申请实施例中,歌词贴片是歌词对应的一种展现形式,如可以是显示有歌词的一张图片,或者一张动图,能给显示歌词内容,可选的,可以是一句歌词对应一个歌词贴片,也可以是一句长歌词对应多个歌词贴片,还可以是多句短歌词对应一个歌词贴片,不同的对应方式,会呈现不同的展示效果。可选的,一句歌词对应一个歌词贴片。
在一些实施例中,基于目标歌词生成对应的歌词贴片,包括:
对目标歌词进行分句或者分词处理,以生成对应的歌词贴片。
在一些实施例中,歌词贴片中显示有歌词内容,在生成对应的歌词贴片时,还包括:
对歌词贴片中显示的歌词内容进行排版。
在一些实施例中,对歌词贴片中显示的歌词内容进行排版,包括:
根据歌词贴片中的歌词的字数、字体、颜色、对齐方式、显示位置中的至少一种,对歌词贴片中显示的歌词内容进行排版。
在本申请实施例中,在根据音乐的歌词生成歌词贴片时,获取该音乐中的歌词,按照预设的规则对歌词进行分句、分词,如将一句歌词形成一个歌词贴片,或者将一句歌词中的某几个词语形成一个歌词贴片,还可以将多句歌词联合起来排版形成一个歌词贴片,可选的,在形成歌词贴片时,对歌词进行排版,可选的,可以调节歌词的字体、颜色等,最终形成一个能够显示歌词内容的图片或者动图。
对于本申请实施例,为方便说明,以一个具体实施例为例,获取音乐的歌词,并对歌词进行分句处理,对每一句歌词,对应生成一个歌词贴片,可选的,该歌词贴片中显示有歌词内容;可选的,可以对歌词按照预设的规则进行排版,如规定每行显示歌词的字数、歌词的字体、对齐方式等,可选的,每个歌词贴片可以显示有不同数量的歌词,每个歌词贴片中歌词的显示位置可以不同,每个歌词贴片中歌词的对齐方式可以不同。
本申请实施例通过获取音乐的歌词,并对每一句歌词都生成对应的歌词贴片,保证歌词贴片的多样性,展示效果更丰富。
在一些实施例中,如图8所示,在目标图像中目标对象所在的位置的预设范围内展示目标歌词,包括:
步骤S801,基于目标歌词在音乐中的播放时间段确定多媒体数据中的目标图像数据;
步骤S802,基于目标图像数据中的目标对象在目标图像数据中的位置调节目标歌词在图像数据中的显示位置,其中显示位置在目标对象在目标图像数据中的位置的预设范围内。
在本申请实施例中,在将歌词和图像数据中的目标对象进行结合时,歌词对应的歌词在音乐中出现的时间段和目标对象在图像数据中出现的时间段要相对应,保证歌词出现时,目标对象也能在图像数据中出现。
对于本申请实施例,为方便说明,以一个具体实施例为例,对于歌词B,其对应的歌词在音乐中出现的时间段为35S~38S,则确定图像数据中在第35S~38S的目标图像,确定该目标图像中一直存在的目标对象B1,将目标对象B1确定为与该歌词B对应的目标对象,并基于该目标对象B1在图像数据中的位置信息,确定歌词贴片的位置。作为本申请一个实施例,歌词出现的位置,也可以是与该歌词对应的目标对象显示位置的附近区域。
本申请实施例通过歌词贴片对应的歌词在音乐中出现的时间和目标对象在图像数据中出现的时间确定歌词贴片与目标对象的对应关系,保证歌词在播放时能够在数据图像中出现,用户可以看到歌词。
在一些实施例中,基于目标对象在图像数据中的深度信息,调节目标歌词的显示特效,包括:
基于目标对象在目标图像中的深度信息,调节目标歌词的显示大小。
在本申请实施例中,调节歌词贴片的显示特效可以是调节歌词的显示大小,为方便说明,以一个具体实施例为例,如图9所示,目标对象C1在图像数据播放的过程中,距离图像捕捉装置越来越远,其深度越来越深,则可以调节该目标对象C1对应的歌词C的显示大小越来越大。
在一些实施例中,基于目标对象的深度信息,调节目标歌词的显示特效,包括:
基于目标对象在目标图像中的深度信息,调节歌词贴片的显示颜色。
可选的,特效的调节方式还可以是根据目标对象的深度信息调节对应的歌词贴片的显示颜色等。
本申请实施例通过根据目标对象的深度信息调节与之对应的歌词的显示大小,给用户一种远小近大的视觉体验。
在一些实施例中,图像数据包括视频数据和/或带有景深信息的全景图像数据,其中,当图像数据为带有景深信息的全景图像数据时,在目标图像中目标对象所在的位置的预设范围内展示目标歌词,包括:
基于全景图像数据中不同目标对象的位置,分别调节不同目标对象对应的目标歌词在全景图像数据中的显示位置。
在本申请实施例中,为方便说明,以一个具体实施例为例,如图10所示,在全景图像1001中,有目标对象D1和目标对象E1,其中,目标对象D1在全景图像的左下方,目标对象E1在全景图像的中间,其中,目标对象D1对应歌词D,目标对象E1对应歌词E,则分别根据目标对象D1和E1的位置调节歌词D和E的位置,其中,目标对象D1的深度信息较浅,目标对象E1的深度信息较深,则可以调节歌词D显示较大,歌词E显示较小。
本申请实施例通过全景图像中不同目标对象的位置和深度信息调节对应歌词的大小,给用户真实的空间感受。
本申请实施例基于音乐的歌词生成对应的歌词贴片,并基于多媒体数据中的图像数据获取图像数据中目标对象的深度信息,基于该目标对象在该图像数据中的位置信息调节与该目标对象对应的歌词贴片在该图像数据中的显示位置,并基于该目标对象在该图像数据中的深度信息调节对应的歌词贴片的显示特效,将歌词嵌入到图像数据的现实空间中,给用户一种虚拟显示的感官体验,用户体验更好。
本申请实施例提供了一种歌词视频展示装置,如图11所示,该歌词视频展示装置110可以包括:数据获取模块1110、歌词确定模块1120以及歌词展示模块1130,其中,
数据获取模块1110,用于基于用户的歌词视频展示操作,播放待展示的多媒体数据和音乐数据,多媒体数据中包含图像数据,音乐数据包括音频数据和歌词;
歌词确定模块1120,用于确定目标时间点,确定图像数据中与目标时间点对应的目标对象,确定歌词中目标时间点对应的目标歌词;
歌词展示模块1130,用于在目标图像中目标对象所在的位置的预设范围内展示目标歌词,并基于目标对象的深度信息,调节目标歌词的显示特效,同时播放目标歌词对应的音频数据。
在一些实施例中,数据获取模块1110在获取待合成的多媒体数据时,可以用于:
接收用户的多媒体数据选择操作;基于多媒体数据选择操作确定多媒体数据。
在一些实施例中,数据获取模块1110在获取待合成的多媒体数据时,可以用于:
基于用户的多媒体数据捕捉操作,开启多媒体数据捕捉装置;
获取多媒体数据捕捉装置捕捉的多媒体数据。
在一些实施例中,歌词展示模块1130在目标图像中目标对象所在的位置的预设范围内展示目标歌词时,可以用于:基于目标歌词生成对应的歌词贴片;在目标图像中目标对象的位置的预设范围内展示歌词贴片。
在一些实施例中,歌词展示模块1130在基于目标歌词生成对应的歌词贴片时,可以用于:对目标歌词进行分句或者分词处理,以生成对应的歌词贴片。
在一些实施例中,歌词贴片中显示有歌词内容,歌词展示模块1130在生成对应的歌词贴片时,还可以用于:对歌词贴片中显示的歌词内容进行排版。
在一些实施例中,歌词展示模块1130在对歌词贴片中显示的歌词内容进行排版时,可以用于:根据歌词贴片中的歌词的字数、字体、颜色、对齐方式、显示位置中的至少一种,对歌词贴片中显示的歌词内容进行排版。
在一些实施例中,歌词展示模块1130在基于目标对象在图像数据中的深度信息,调节目标歌词的显示特效时,可以用于:
基于目标对象在目标图像中的深度信息,调节歌词贴片的显示颜色。
在一些实施例中,歌词展示模块1130在目标图像中目标对象所在的位置的预设范围内展示目标歌词时,可以用于:
基于目标歌词在音乐中的播放时间段确定多媒体数据中的目标图像数据;
基于目标图像数据中的目标对象在目标图像数据中的位置调节目标歌词在图像数据中的显示位置,其中显示位置在目标对象在目标图像数据中的位置的预设范围内。
在一些实施例中,歌词展示模块1130在基于目标对象在图像数据中的位置调节目标歌词在图像数据中的显示位置时,可以用于:
基于目标对象在目标图像中的深度信息,调节目标歌词的显示大小。
在一些实施例中,图像数据包括视频数据和/或带有景深信息的全景图像数据,其中,当图像数据为带有景深信息的全景图像数据时,歌词展示模块1130在目标图像中目标对象所在的位置的预设范围内展示目标歌词时,可以用于:
基于全景图像数据中不同目标对象的位置,分别调节不同目标对象对应的歌词贴片在全景图像数据中的显示位置。
在一些实施例中,歌词展示模块1130在所述基于所述目标对象的深度信息,调节所述目标歌词的显示特效之前,还可以用于:
获取所述图像数据中所述目标对象的深度信息,其中,所述目标对象的深度信息用于表示所述目标对象距离视频捕捉装置的距离的信息。
本申请实施例基于音乐的歌词生成对应的歌词贴片,并基于多媒体数据中的图像数据获取图像数据中目标对象的深度信息,基于该目标对象在该图像数据中的位置信息调 节与该目标对象对应的歌词贴片在该图像数据中的显示位置,并基于该目标对象在该图像数据中的深度信息调节对应的歌词贴片的显示特效,将歌词嵌入到图像数据的现实空间中,给用户一种虚拟显示的感官体验,用户体验更好。
下面参考图12,其示出了适于用来实现本申请实施例的电子设备1200的结构示意图。本申请实施例中的终端设备可以包括但不限于诸如移动电话、笔记本电脑、数字广播接收器、PDA(个人数字助理)、PAD(平板电脑)、PMP(便携式多媒体播放器)、车载终端(例如车载导航终端)等等的移动终端以及诸如数字TV、台式计算机等等的固定终端。图12示出的电子设备仅仅是一个示例,不应对本申请实施例的功能和使用范围带来任何限制。
电子设备包括:存储器以及处理器,其中,这里的处理器可以称为下文的处理装置1201,存储器可以包括下文中的只读存储器(ROM)1202、随机访问存储器(RAM)1203以及存储装置1208中的至少一项,具体如下所示:
如图12所示,电子设备1200可以包括处理装置(例如中央处理器、图形处理器等)1201,其可以根据存储在只读存储器(ROM)1202中的程序或者从存储装置1208加载到随机访问存储器(RAM)1203中的程序而执行各种适当的动作和处理。在RAM 1203中,还存储有电子设备1200操作所需的各种程序和数据。处理装置1201、ROM 1202以及RAM 1203通过总线1204彼此相连。输入/输出(I/O)接口1205也连接至总线1204。
通常,以下装置可以连接至I/O接口1205:包括例如触摸屏、触摸板、键盘、鼠标、摄像头、麦克风、加速度计、陀螺仪等的输入装置1206;包括例如液晶显示器(LCD)、扬声器、振动器等的输出装置1207;包括例如磁带、硬盘等的存储装置1208;以及通信装置1209。通信装置1209可以允许电子设备1200与其他设备进行无线或有线通信以交换数据。虽然图12示出了具有各种装置的电子设备1200,但是应理解的是,并不要求实施或具备所有示出的装置。可以替代地实施或具备更多或更少的装置。
特别地,根据本申请的实施例,上文参考流程图描述的过程可以被实现为计算机软件程序。例如,本申请的实施例包括一种计算机程序产品,其包括承载在非暂态计算机可读介质上的计算机程序,该计算机程序包含用于执行流程图所示的方法的程序代码。在这样的实施例中,该计算机程序可以通过通信装置1209从网络上被下载和安装,或者从存储装置1208被安装,或者从ROM 1202被安装。在该计算机程序被处理装置1201执行时,执行本申请实施例的方法中限定的上述功能。
需要说明的是,本申请上述的计算机可读介质可以是计算机可读信号介质或者计算机可读介质或者是上述两者的任意组合。计算机可读介质例如可以是——但不限于——电、磁、光、电磁、红外线、或半导体的系统、装置或器件,或者任意以上的组合。 计算机可读介质的更具体的例子可以包括但不限于:具有一个或多个导线的电连接、便携式计算机磁盘、硬盘、随机访问存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(EPROM或闪存)、光纤、便携式紧凑磁盘只读存储器(CD-ROM)、光存储器件、磁存储器件、或者上述的任意合适的组合。在本申请中,计算机可读介质可以是任何包含或存储程序的有形介质,该程序可以被指令执行系统、装置或者器件使用或者与其结合使用。而在本申请中,计算机可读信号介质可以包括在基带中或者作为载波一部分传播的数据信号,其中承载了计算机可读的程序代码。这种传播的数据信号可以采用多种形式,包括但不限于电磁信号、光信号或上述的任意合适的组合。计算机可读信号介质还可以是计算机可读介质以外的任何计算机可读介质,该计算机可读信号介质可以发送、传播或者传输用于由指令执行系统、装置或者器件使用或者与其结合使用的程序。计算机可读介质上包含的程序代码可以用任何适当的介质传输,包括但不限于:电线、光缆、RF(射频)等等,或者上述的任意合适的组合。
在一些实施方式中,客户端、服务器可以利用诸如HTTP(HyperText Transfer Protocol,超文本传输协议)之类的任何当前已知或未来研发的网络协议进行通信,并且可以与任意形式或介质的数字数据通信(例如,通信网络)互连。通信网络的示例包括局域网(“LAN”),广域网(“WAN”),网际网(例如,互联网)以及端对端网络(例如,ad hoc端对端网络),以及任何当前已知或未来研发的网络。
上述计算机可读介质可以是上述电子设备中所包含的;也可以是单独存在,而未装配入该电子设备中。
上述计算机可读介质承载有一个或者多个程序,当上述一个或者多个程序被该电子设备执行时,使得该电子设备:获取待合成的多媒体数据和音乐,多媒体数据中包含图像数据;获取图像数据中目标对象的深度信息;根据音乐的歌词生成歌词贴片;基于目标对象在图像数据中的位置调节歌词贴片在图像数据中的显示位置;基于目标对象在图像数据中的深度信息,调节歌词贴片的显示特效,并生成歌词视频。
可以以一种或多种程序设计语言或其组合来编写用于执行本申请的操作的计算机程序代码,上述程序设计语言包括但不限于面向对象的程序设计语言—诸如Java、Smalltalk、C++,还包括常规的过程式程序设计语言—诸如“C”语言或类似的程序设计语言。程序代码可以完全地在用户计算机上执行、部分地在用户计算机上执行、作为一个独立的软件包执行、部分在用户计算机上部分在远程计算机上执行、或者完全在远程计算机或服务器上执行。在涉及远程计算机的情形中,远程计算机可以通过任意种类的网络——包括局域网(LAN)或广域网(WAN)—连接到用户计算机,或者,可以连接到外部计算机(例如利用因特网服务提供商来通过因特网连接)。
附图中的流程图和框图,图示了按照本申请各种实施例的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段、或代码的一部分,该模块、程序段、或代码的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。也应当注意,在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个接连地表示的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或操作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
描述于本申请实施例中所涉及到的模块或单元可以通过软件的方式实现,也可以通过硬件的方式来实现。
本文中以上描述的功能可以至少部分地由一个或多个硬件逻辑部件来执行。例如,非限制性地,可以使用的示范类型的硬件逻辑部件包括:现场可编程门阵列(FPGA)、专用集成电路(ASIC)、专用标准产品(ASSP)、片上系统(SOC)、复杂可编程逻辑设备(CPLD)等等。
在本申请的上下文中,机器可读介质可以是有形的介质,其可以包含或存储以供指令执行系统、装置或设备使用或与指令执行系统、装置或设备结合地使用的程序。机器可读介质可以是机器可读信号介质或机器可读储存介质。机器可读介质可以包括但不限于电子的、磁性的、光学的、电磁的、红外的、或半导体系统、装置或设备,或者上述内容的任何合适组合。机器可读存储介质的更具体示例会包括基于一个或多个线的电气连接、便携式计算机盘、硬盘、随机存取存储器(RAM)、只读存储器(ROM)、可擦除可编程只读存储器(EPROM或快闪存储器)、光纤、便捷式紧凑盘只读存储器(CD-ROM)、光学储存设备、磁储存设备、或上述内容的任何合适组合。
根据本申请提供的一个或多个实施例,提供了一种歌词视频展示方法,该方法包括:
基于用户的歌词视频展示操作,播放待展示的多媒体数据和音乐数据,多媒体数据中包含图像数据,音乐数据包括音频数据和歌词;确定目标时间点,确定图像数据中与目标时间点对应的目标对象,确定歌词中与目标时间点对应的目标歌词;在目标图像中目标对象所在的位置的预设范围内展示目标歌词,并基于目标对象的深度信息,调节目标歌词的显示特效,同时播放目标歌词对应的音频数据。
在一些实施例中,基于用户的歌词视频展示操作,播放待展示的多媒体数据和音乐数据之前,还包括:接收用户的多媒体数据选择操作;基于多媒体数据选择操作确定多媒体数据。
在一些实施例中,基于用户的歌词视频展示操作,播放待展示的多媒体数据和音乐数据之前,还包括:基于用户的多媒体数据捕捉操作,开启多媒体数据捕捉装置;获取多媒体数据捕捉装置捕捉的多媒体数据。
在一些实施例中,在目标图像中目标对象所在的位置的预设范围内展示目标歌词,包括:基于目标歌词生成对应的歌词贴片;在目标图像中目标对象的位置的预设范围内展示歌词贴片。
在一些实施例中,基于目标歌词生成对应的歌词贴片,包括:
对目标歌词进行分句或者分词处理,以生成对应的歌词贴片。
在一些实施例中,歌词贴片中显示有歌词内容,在生成对应的歌词贴片时,还包括:
对歌词贴片中显示的歌词内容进行排版。
在一些实施例中,对歌词贴片中显示的歌词内容进行排版,包括:
根据歌词贴片中的歌词的字数、字体、颜色、对齐方式、显示位置中的至少一种,对歌词贴片中显示的歌词内容进行排版。
在一些实施例中,基于目标对象的深度信息,调节目标歌词的显示特效,包括:
基于目标对象在目标图像中的深度信息,调节歌词贴片的显示颜色。
在一些实施例中,在目标图像中目标对象所在的位置的预设范围内展示目标歌词,包括:基于目标歌词在音乐中的播放时间段确定多媒体数据中的目标图像数据;基于目标图像数据中的目标对象在目标图像数据中的位置调节目标歌词在图像数据中的显示位置,其中显示位置在目标对象在目标图像数据中的位置的预设范围内。
在一些实施例中,基于目标对象的深度信息,调节目标歌词的显示特效,包括:
基于目标对象在目标图像中的深度信息,调节目标歌词的显示大小。
在一些实施例中,图像数据包括视频数据和/或带有景深信息的全景图像数据,其中,当图像数据为带有景深信息的全景图像数据时,在目标图像中目标对象所在的位置的预设范围内展示目标歌词,包括:基于全景图像数据中不同目标对象的位置,分别调节不同目标对象对应的目标歌词在全景图像数据中的显示位置。
在一些实施例中,在基于目标对象的深度信息,调节目标歌词的显示特效之前,还包括:获取图像数据中目标对象的深度信息,其中,目标对象的深度信息用于表示目标对象距离视频捕捉装置的距离的信息。
根据本申请提供的一个或多个实施例,提供了一种歌词视频展示装置,该装置包括:
数据获取模块,用于基于用户的歌词视频展示操作,播放待展示的多媒体数据和音乐数据,多媒体数据中包含图像数据,音乐数据包括音频数据和歌词;
歌词确定模块,用于确定目标时间点,确定图像数据中与目标时间点对应的目标对 象,确定歌词中目标时间点对应的目标歌词;
歌词展示模块,用于在目标图像中目标对象所在的位置的预设范围内展示目标歌词,并基于目标对象的深度信息,调节目标歌词的显示特效,同时播放目标歌词对应的音频数据。
在一些实施例中,数据获取模块在获取待合成的多媒体数据时,可以用于:接收用户的多媒体数据选择操作;基于多媒体数据选择操作确定多媒体数据。
在一些实施例中,数据获取模块在获取待合成的多媒体数据时,可以用于:基于用户的多媒体数据捕捉操作,开启多媒体数据捕捉装置;获取多媒体数据捕捉装置捕捉的多媒体数据。
在一些实施例中,歌词展示模块在目标图像中目标对象所在的位置的预设范围内展示目标歌词时,可以用于:基于目标歌词生成对应的歌词贴片;在目标图像中目标对象的位置的预设范围内展示歌词贴片。
在一些实施例中,歌词展示模块在基于目标歌词生成对应的歌词贴片时,可以用于:
对目标歌词进行分句或者分词处理,以生成对应的歌词贴片。
在一些实施例中,歌词贴片中显示有歌词内容,歌词展示模块在生成对应的歌词贴片时,还可以用于:对歌词贴片中显示的歌词内容进行排版。
在一些实施例中,歌词展示模块在对歌词贴片中显示的歌词内容进行排版时,可以用于:根据歌词贴片中的歌词的字数、字体、颜色、对齐方式、显示位置中的至少一种,对歌词贴片中显示的歌词内容进行排版。
在一些实施例中,歌词展示模块在基于目标对象在图像数据中的深度信息,调节目标歌词的显示特效时,可以用于:
基于目标对象在目标图像中的深度信息,调节歌词贴片的显示颜色。
在一些实施例中,歌词展示模块在目标图像中目标对象所在的位置的预设范围内展示目标歌词时,可以用于:
基于目标歌词在音乐中的播放时间段确定多媒体数据中的目标图像数据;
基于目标图像数据中的目标对象在目标图像数据中的位置调节目标歌词在图像数据中的显示位置,其中显示位置在目标对象在目标图像数据中的位置的预设范围内。
在一些实施例中,歌词展示模块在基于目标对象在图像数据中的位置调节歌词贴片在图像数据中的显示位置时,可以用于:
基于目标对象在目标图像中的深度信息,调节目标歌词的显示大小。
在一些实施例中,歌词展示模块在基于目标对象在图像数据中的深度信息,调节歌词贴片的显示特效时,可以用于:
基于目标对象在图像数据中的深度信息,调节歌词贴片的显示大小。
在一些实施例中,图像数据包括视频数据和/或带有景深信息的全景图像数据,其中,当图像数据为带有景深信息的全景图像数据时,歌词展示模块在在目标图像中目标对象所在的位置的预设范围内展示目标歌词时,可以用于:
基于全景图像数据中不同目标对象的位置,分别调节不同目标对象对应的歌词贴片在全景图像数据中的显示位置。
在一些实施例中,歌词展示模块在所述基于所述目标对象的深度信息,调节所述目标歌词的显示特效之前,还可以用于:
获取所述图像数据中所述目标对象的深度信息,其中,所述目标对象的深度信息用于表示所述目标对象距离视频捕捉装置的距离的信息。
根据本申请提供的一个或多个实施例,提供了一种电子设备,该电子设备包括:一个或多个处理器;存储器;一个或多个应用程序,其中一个或多个应用程序被存储在存储器中并被配置为由一个或多个处理器执行,一个或多个程序配置用于:执行上述实施例的歌词视频展示方法。
根据本申请实施例提供的一个或多个实施例,提供了一种计算机可读介质,该介质存储有至少一条指令、至少一段程序、代码集或指令集,所述至少一条指令、所述至少一段程序、所述代码集或指令集由所述处理器加载并执行以实现上述实施例所述的歌词视频展示方法。
以上描述仅为本申请的较佳实施例以及对所运用技术原理的说明。本领域技术人员应当理解,本申请中所涉及的公开范围,并不限于上述技术特征的特定组合而成的技术方案,同时也应涵盖在不脱离上述公开构思的情况下,由上述技术特征或其等同特征进行任意组合而形成的其它技术方案。例如上述特征与本申请中公开的(但不限于)具有类似功能的技术特征进行互相替换而形成的技术方案。
此外,虽然采用特定次序描绘了各操作,但是这不应当理解为要求这些操作以所示出的特定次序或以顺序次序执行来执行。在一定环境下,多任务和并行处理可能是有利的。同样地,虽然在上面论述中包含了若干具体实现细节,但是这些不应当被解释为对本申请的范围的限制。在单独的实施例的上下文中描述的某些特征还可以组合地实现在单个实施例中。相反地,在单个实施例的上下文中描述的各种特征也可以单独地或以任何合适的子组合的方式实现在多个实施例中。
尽管已经采用特定于结构特征和/或方法逻辑动作的语言描述了本主题,但是应当理解所附权利要求书中所限定的主题未必局限于上面描述的特定特征或动作。相反,上面所描述的特定特征和动作仅仅是实现权利要求书的示例形式。

Claims (15)

  1. 一种歌词视频展示方法,其包括:
    基于用户的歌词视频展示操作,播放待展示的多媒体数据和音乐数据,所述多媒体数据包含图像数据,所述音乐数据包括音频数据和歌词;
    确定目标时间点,确定所述图像数据中与所述目标时间点对应的目标对象,确定所述歌词中与所述目标时间点对应的目标歌词;
    在所述目标图像中所述目标对象所在的位置的预设范围内展示所述目标歌词,并基于所述目标对象的深度信息,调节所述目标歌词的显示特效,同时播放所述目标歌词对应的音频数据。
  2. 根据权利要求1所述的方法,其中,所述基于用户的歌词视频展示操作,播放待展示的多媒体数据和音乐数据之前,还包括:
    接收用户的多媒体数据选择操作;
    基于所述多媒体数据选择操作确定多媒体数据。
  3. 根据权利要求1所述的方法,其中,所述基于用户的歌词视频展示操作,播放待展示的多媒体数据和音乐数据之前,还包括:
    基于用户的多媒体数据捕捉操作,开启多媒体数据捕捉装置;
    获取所述多媒体数据捕捉装置捕捉的多媒体数据。
  4. 根据权利要求1所述的方法,其中,所述在所述目标图像中所述目标对象所在的位置的预设范围内展示所述目标歌词,包括:
    基于所述目标歌词生成对应的歌词贴片;
    在所述目标图像中所述目标对象的位置的预设范围内展示所述歌词贴片。
  5. 根据权利要求4所述的方法,其中,所述基于所述目标歌词生成对应的歌词贴片,包括:
    对所述目标歌词进行分句或者分词处理,以生成对应的歌词贴片。
  6. 根据权利要求5所述的方法,其中,所述歌词贴片中显示有歌词内容,所述在生成对应的歌词贴片时,还包括:
    对所述歌词贴片中显示的所述歌词内容进行排版。
  7. 根据权利要求6所述的方法,其中,对所述歌词贴片中显示的所述歌词内容进行排版,包括:
    根据所述歌词贴片中的歌词的字数、字体、颜色、对齐方式、显示位置中的至少一种,对所述歌词贴片中显示的所述歌词内容进行排版。
  8. 根据权利要求4所述的方法,其中,所述基于所述目标对象的深度信息,调节 所述目标歌词的显示特效,包括:
    基于所述目标对象在所述目标图像中的深度信息,调节所述歌词贴片的显示颜色。
  9. 根据权利要求1所述的方法,其中,所述在所述目标图像中所述目标对象所在的位置的预设范围内展示所述目标歌词,包括:
    基于所述目标歌词在所述音乐中的播放时间段确定所述多媒体数据中的目标图像;
    基于所述目标图像中的目标对象在所述目标图像中的位置调节所述目标歌词在所述目标图像中的显示位置,其中所述显示位置在所述目标对象在所述目标图像中的位置的预设范围内。
  10. 根据权利要求1所述的方法,其中,所述基于所述目标对象的深度信息,调节所述目标歌词的显示特效,包括:
    基于所述目标对象在所述目标图像中的深度信息,调节所述目标歌词的显示大小。
  11. 根据权利要求1所述的方法,其中,所述图像数据包括视频数据和/或带有景深信息的全景图像数据,其中,当所述图像数据为带有景深信息的全景图像数据时,所述在所述目标图像中所述目标对象所在的位置的预设范围内展示所述目标歌词,包括:
    基于所述全景图像数据中不同目标对象的位置,分别调节所述不同目标对象对应的目标歌词在所述全景图像数据中的显示位置。
  12. 根据权利要求1所述的方法,其中,在所述基于所述目标对象的深度信息,调节所述目标歌词的显示特效之前,还包括:
    获取所述图像数据中所述目标对象的深度信息,其中,所述目标对象的深度信息用于表示所述目标对象距离视频捕捉装置的距离的信息。
  13. 一种歌词视频展示装置,其包括:
    数据获取模块,用于基于用户的歌词视频展示操作,播放待展示的多媒体数据和音乐数据,所述多媒体数据中包含图像数据,所述音乐数据包括音频数据和歌词;
    歌词确定模块,用于确定目标时间点,确定所述图像数据中与所述目标时间点对应的目标对象,确定所述歌词中所述目标时间点对应的目标歌词;
    歌词展示模块,用于在所述目标图像中所述目标对象所在的位置的预设范围内展示所述目标歌词,并基于所述目标对象的深度信息,调节所述目标歌词的显示特效,同时播放所述目标歌词对应的音频数据。
  14. 一种电子设备,其包括:
    一个或多个处理器;
    存储器;
    一个或多个应用程序,其中所述一个或多个应用程序被存储在所述存储器中并被配 置为由所述一个或多个处理器执行,所述一个或多个程序配置用于:执行根据权利要求1~12任一项所述的歌词视频展示方法。
  15. 一种计算机可读介质,其中,所述介质存储有至少一条指令、至少一段程序、代码集或指令集,所述至少一条指令、所述至少一段程序、所述代码集或指令集由所述处理器加载并执行以实现如权利要求1~12任一所述的歌词视频展示方法。
PCT/CN2021/102438 2020-11-10 2021-06-25 歌词视频展示方法、装置、电子设备及计算机可读介质 Ceased WO2022100103A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US17/438,732 US12549801B2 (en) 2020-11-10 2021-06-25 Lyric video display method and device, electronic apparatus and computer-readable medium

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202011247956.3 2020-11-10
CN202011247956.3A CN112383810A (zh) 2020-11-10 2020-11-10 歌词视频展示方法、装置、电子设备及计算机可读介质

Publications (1)

Publication Number Publication Date
WO2022100103A1 true WO2022100103A1 (zh) 2022-05-19

Family

ID=74579297

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2021/102438 Ceased WO2022100103A1 (zh) 2020-11-10 2021-06-25 歌词视频展示方法、装置、电子设备及计算机可读介质

Country Status (3)

Country Link
US (1) US12549801B2 (zh)
CN (1) CN112383810A (zh)
WO (1) WO2022100103A1 (zh)

Families Citing this family (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112383810A (zh) 2020-11-10 2021-02-19 北京字跳网络技术有限公司 歌词视频展示方法、装置、电子设备及计算机可读介质
CN112380378B (zh) * 2020-11-17 2022-09-02 北京字跳网络技术有限公司 歌词特效展示方法、装置、电子设备及计算机可读介质
CN116801050A (zh) * 2023-07-07 2023-09-22 北京字跳网络技术有限公司 视频处理方法、装置、设备、存储介质和程序产品

Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2011007868A (ja) * 2009-06-23 2011-01-13 Daiichikosho Co Ltd 背景映像中の顔画像を避けるように歌詞字幕を表示するカラオケ装置
CN104219559A (zh) * 2013-05-31 2014-12-17 奥多比公司 在视频内容中投放不明显叠加
CN107944397A (zh) * 2017-11-27 2018-04-20 腾讯音乐娱乐科技(深圳)有限公司 视频录制方法、装置及计算机可读存储介质
CN107943964A (zh) * 2017-11-27 2018-04-20 腾讯音乐娱乐科技(深圳)有限公司 歌词显示方法、装置及计算机可读存储介质
CN108089830A (zh) * 2017-12-11 2018-05-29 维沃移动通信有限公司 歌曲信息显示方法、装置及移动终端
WO2020105847A1 (ko) * 2018-11-23 2020-05-28 삼성전자주식회사 전자 장치 및 그 제어 방법
CN112383810A (zh) * 2020-11-10 2021-02-19 北京字跳网络技术有限公司 歌词视频展示方法、装置、电子设备及计算机可读介质

Family Cites Families (13)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US7106381B2 (en) * 2003-03-24 2006-09-12 Sony Corporation Position and time sensitive closed captioning
US8269821B2 (en) * 2009-01-27 2012-09-18 EchoStar Technologies, L.L.C. Systems and methods for providing closed captioning in three-dimensional imagery
JP2011029849A (ja) * 2009-07-23 2011-02-10 Sony Corp 受信装置、通信システム、立体画像への字幕合成方法、プログラム、及びデータ構造
US20110107216A1 (en) * 2009-11-03 2011-05-05 Qualcomm Incorporated Gesture-based user interface
US9591374B2 (en) * 2010-06-30 2017-03-07 Warner Bros. Entertainment Inc. Method and apparatus for generating encoded content using dynamically optimized conversion for 3D movies
KR20120004203A (ko) * 2010-07-06 2012-01-12 삼성전자주식회사 디스플레이 방법 및 장치
KR101830656B1 (ko) * 2011-12-02 2018-02-21 엘지전자 주식회사 이동 단말기 및 이의 제어방법
US9456170B1 (en) * 2013-10-08 2016-09-27 3Play Media, Inc. Automated caption positioning systems and methods
US20170371498A1 (en) * 2016-06-28 2017-12-28 International Business Machines Corporation Displaying ui components
US10599916B2 (en) * 2017-11-13 2020-03-24 Facebook, Inc. Methods and systems for playing musical elements based on a tracked face or facial feature
CN108008930B (zh) * 2017-11-30 2020-06-30 广州酷狗计算机科技有限公司 确定k歌分值的方法和装置
CN110620946B (zh) * 2018-06-20 2022-03-18 阿里巴巴(中国)有限公司 字幕显示方法及装置
US11263817B1 (en) * 2019-12-19 2022-03-01 Snap Inc. 3D captions with face tracking

Patent Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2011007868A (ja) * 2009-06-23 2011-01-13 Daiichikosho Co Ltd 背景映像中の顔画像を避けるように歌詞字幕を表示するカラオケ装置
CN104219559A (zh) * 2013-05-31 2014-12-17 奥多比公司 在视频内容中投放不明显叠加
CN107944397A (zh) * 2017-11-27 2018-04-20 腾讯音乐娱乐科技(深圳)有限公司 视频录制方法、装置及计算机可读存储介质
CN107943964A (zh) * 2017-11-27 2018-04-20 腾讯音乐娱乐科技(深圳)有限公司 歌词显示方法、装置及计算机可读存储介质
CN108089830A (zh) * 2017-12-11 2018-05-29 维沃移动通信有限公司 歌曲信息显示方法、装置及移动终端
WO2020105847A1 (ko) * 2018-11-23 2020-05-28 삼성전자주식회사 전자 장치 및 그 제어 방법
CN112383810A (zh) * 2020-11-10 2021-02-19 北京字跳网络技术有限公司 歌词视频展示方法、装置、电子设备及计算机可读介质

Also Published As

Publication number Publication date
US20220394325A1 (en) 2022-12-08
US12549801B2 (en) 2026-02-10
CN112383810A (zh) 2021-02-19

Similar Documents

Publication Publication Date Title
CN112259062B (zh) 特效展示方法、装置、电子设备及计算机可读介质
CN112380379B (zh) 歌词特效展示方法、装置、电子设备及计算机可读介质
CN111970571B (zh) 视频制作方法、装置、设备及存储介质
CN114168250B (zh) 页面显示方法、装置、电子设备和存储介质
CN111833460B (zh) 增强现实的图像处理方法、装置、电子设备及存储介质
WO2022100103A1 (zh) 歌词视频展示方法、装置、电子设备及计算机可读介质
CN112954441B (zh) 视频编辑及播放方法、装置、设备、介质
CN112423107B (zh) 歌词视频展示方法、装置、电子设备及计算机可读介质
WO2021135626A1 (zh) 菜单项选择方法、装置、可读介质及电子设备
WO2022048504A1 (zh) 视频处理方法、终端设备及存储介质
WO2022057348A1 (zh) 音乐海报生成方法、装置、电子设备及介质
US12019669B2 (en) Method, apparatus, device, readable storage medium and product for media content processing
WO2022105245A1 (zh) 歌词特效展示方法、装置、电子设备及计算机可读介质
CN110798327B (zh) 消息处理方法、设备及存储介质
WO2020259130A1 (zh) 精选片段处理方法、装置、电子设备及可读介质
WO2021197024A1 (zh) 视频特效配置文件生成方法、视频渲染方法及装置
US20250380023A1 (en) Method and apparatus for interaction in live-streaming room, and device and medium
WO2022206335A1 (zh) 图像显示方法、装置、设备及介质
WO2022042290A1 (zh) 一种虚拟模型处理方法、装置、电子设备和存储介质
WO2023088006A1 (zh) 云游戏交互方法、装置、可读介质和电子设备
CN114125358A (zh) 云会议字幕显示方法、系统、装置、电子设备和存储介质
CN116527993A (zh) 视频的处理方法、装置、电子设备、存储介质和程序产品
CN114245218A (zh) 音视频播放方法、装置、计算机设备及存储介质
CN114554292A (zh) 视角的切换方法、装置、电子设备、存储介质和程序产品
WO2022022363A1 (zh) 视频生成及播放方法、装置、电子设备及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 21890634

Country of ref document: EP

Kind code of ref document: A1

REG Reference to national code

Ref country code: BR

Ref legal event code: B01A

Ref document number: 112021018273

Country of ref document: BR

ENP Entry into the national phase

Ref document number: 112021018273

Country of ref document: BR

Kind code of ref document: A2

Effective date: 20210914

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 21890634

Country of ref document: EP

Kind code of ref document: A1

WWG Wipo information: grant in national office

Ref document number: 17438732

Country of ref document: US