WO2022247293A1 - 视频直播方法和视频直播装置 - Google Patents

视频直播方法和视频直播装置 Download PDF

Info

Publication number
WO2022247293A1
WO2022247293A1 PCT/CN2022/070254 CN2022070254W WO2022247293A1 WO 2022247293 A1 WO2022247293 A1 WO 2022247293A1 CN 2022070254 W CN2022070254 W CN 2022070254W WO 2022247293 A1 WO2022247293 A1 WO 2022247293A1
Authority
WO
WIPO (PCT)
Prior art keywords
live
background
video
live video
region
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2022/070254
Other languages
English (en)
French (fr)
Inventor
田园
李鑫
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Dajia Internet Information Technology Co Ltd
Original Assignee
Beijing Dajia Internet Information Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Dajia Internet Information Technology Co Ltd filed Critical Beijing Dajia Internet Information Technology Co Ltd
Publication of WO2022247293A1 publication Critical patent/WO2022247293A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/20Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
    • H04N21/21Server components or server architectures
    • H04N21/218Source of audio or video content, e.g. local disk arrays
    • H04N21/2187Live feed
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/20Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
    • H04N21/23Processing of content or additional data; Elementary server operations; Server middleware
    • H04N21/234Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
    • H04N21/23424Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving splicing one content stream with another content stream, e.g. for inserting or substituting an advertisement

Definitions

  • the present disclosure relates to the technical field of video processing, and in particular, to a video live broadcast method and a video live broadcast device.
  • the live video service has become the current trend.
  • the anchor can conduct live video broadcasting through the live broadcast software, and online interaction can also be realized between the anchors.
  • the current live broadcast interaction method is relatively simple.
  • the present disclosure provides a video live broadcast method and a video live broadcast device.
  • the video live broadcast method may include: acquiring the first live data collected by the first device, wherein the first live data includes the first live data collected by the first device At least one of the first region of interest in the first live video and the first background; obtaining second live data collected by the second device, wherein the second live data includes the second live video collected by the second device at least one of the second region of interest and the second background; generate a target live video based on the first live data and the second live data; and send the target live video.
  • the region of interest may be obtained by performing target region extraction on each frame of the first live video and/or the second live video.
  • the region of interest may be a portrait region.
  • the step of generating the target live video based on the first live data and the second live data may include: fusing the first background and the second background to generate a fusion background as the background in the target live video; A first region of interest and a second region of interest are displayed in the fused background.
  • the step of generating the target live video based on the first live data and the second live data may include: selecting the first background or the second background as the background in the target live video, and in the selected background Displays the first and second ROIs.
  • the first ROI and/or the second ROI may change position or size in the target live video based on user input.
  • a live video device which may include: an acquisition module configured to: acquire the first live data collected by the first device and the data collected by the second device The second live data, wherein the first live data includes at least one of the first region of interest and the first background in the first live video collected by the first device, and the second live data includes at least one of the first background collected by the second device At least one of the second region of interest and the second background in the second live video; a processing module configured to generate a target live video based on the first live data and the second live data; and a sending module configured to send The target live video.
  • the region of interest may be obtained by performing target region extraction on each frame of the first live video and/or the second live video.
  • the region of interest may be a portrait region.
  • the processing module may be configured to: fuse the first background and the second background to generate a fusion background as the background in the target live video; display the first region of interest and Second region of interest.
  • the processing module may be configured to: select the first background and/or the second background as the background in the target live video, and display the first region of interest and the second region of interest in the selected background area.
  • the first ROI and/or the second ROI may change position or size in the target live video based on user input.
  • an electronic device may include: at least one processor; at least one memory storing computer-executable instructions, wherein the computer-executable instructions are executed by the When the at least one processor is running, the at least one processor is prompted to execute the video live broadcast method as described above.
  • a non-volatile computer-readable storage medium on which instructions are stored, and when the instructions are executed by at least one processor, the at least one processor is prompted to perform the above The method for live video broadcasting.
  • a computer program product is provided, where instructions in the computer program product are executed by at least one processor in an electronic device to execute the video live broadcast method as described above.
  • the present disclosure can fuse regions of interest and/or backgrounds in different live videos, which enhances live broadcast interactivity and enriches live broadcast scenes, giving users more choices.
  • FIG. 1 is a diagram of an application environment for live interaction according to some embodiments of the present disclosure
  • FIG. 2 is a flow chart of a video live broadcast method according to some embodiments of the present disclosure
  • 3 to 5 are schematic diagrams of live images according to some embodiments of the present disclosure.
  • Fig. 6 is a schematic structural diagram of a live video device according to some embodiments of the present disclosure.
  • Fig. 7 is a block diagram of a live video device according to some embodiments of the present disclosure.
  • Figure 8 is a block diagram of an electronic device according to some embodiments of the present disclosure.
  • Fig. 9 is a flow chart of a video live broadcast method according to some embodiments of the present disclosure.
  • Fig. 10 is a schematic flowchart of live video interaction according to some embodiments of the present disclosure.
  • Fig. 1 is a diagram of an application environment for live interaction according to an embodiment of the present disclosure.
  • the application environment 100 includes a terminal 110 , a terminal 120 and a server 130 .
  • the terminal 110 may be the terminal where the user is located, for example, the terminal used by the anchor for live broadcast.
  • the terminal 110 may be at least one of a smart phone, a tablet computer, a portable computer, a desktop computer, and the like.
  • the terminal 110 may be installed with a target application, which is used for, for example, receiving data from an external device and sending data to an external device, performing matting processing on a collected video, and the like.
  • the terminal 120 may be a terminal where the user is located, for example, a terminal used by an anchor for live broadcasting.
  • the terminal 120 may be at least one of a smart phone, a tablet computer, a portable computer, a desktop computer, and the like.
  • the terminal 120 may be installed with a target application, which is used for, for example, receiving data from an external device and sending data to an external device, and performing matting processing on a collected video, and the like.
  • a target application which is used for, for example, receiving data from an external device and sending data to an external device, and performing matting processing on a collected video, and the like.
  • the terminal 110 may be connected to the terminal 120 through a wireless network, so that data interaction between the terminal 110 and the terminal 120 may be performed.
  • a network may include Bluetooth, a local area network (LAN), a wide area network (WAN), a wireless link, an intranet, the Internet, combinations thereof, and the like.
  • terminal 110 is taken as an example.
  • Terminal 110 may receive live data from terminal 120 and fuse the live data of terminal 110 with the live data of terminal 120 .
  • the live data received by terminal 110 may be, for example, an area of interest (such as a portrait area) in a video collected by terminal 120 .
  • An embodiment in this regard will be described in detail below with reference to FIG. 2 .
  • terminal 120 may receive live data from terminal 110 and fuse the live data of terminal 120 with the live data of terminal 110 .
  • the terminal 110 and the terminal 120 may send the live data captured by them to the server 130, and the server 130 may fuse the received live data. An embodiment in this regard will be described in detail below with reference to FIG. 9 .
  • the terminal 110 may be connected to the server 130 through a wireless network, so that data interaction between the terminal 110 and the server 130 may be performed.
  • a network may comprise a local area network (LAN), a wide area network (WAN), wireless links, an intranet, the Internet, combinations thereof, and the like.
  • the terminal 110 may also be connected to the server 130 through a wired network for data interaction.
  • the server 130 may be a server for parsing and processing the received data.
  • the server 130 can receive the live data from the terminal 110 , and push the received live data to the audience of the terminal 110 .
  • the anchor uses the terminal 110 to collect live data, and receives the region of interest data in the live data collected by the terminal 120 in real time via the network. sent to the server 130 via the network.
  • the server 130 may forward the received fusion data to the viewer of the terminal 110 . In this way, the viewer can watch the live content from the terminal 110 and the terminal 120 at the same time.
  • terminal 110 may receive live data transmitted by terminal 120 . Before terminal 120 transmits the live broadcast data to terminal 110, terminal 120 may firstly perform matting processing on the collected video picture, so that the interested area in the picture (such as the portrait of user 2) is separated from the background, and then the separated interested area The zone data is transmitted to the terminal 110 .
  • the terminal 110 can merge the video collected by itself with the received data, so as to display the region of interest from the terminal 120 in the video collected by the terminal 110, such as The portrait of user 1 and the portrait of user 2 may be simultaneously displayed on terminal 110 , or other regions of interest in a picture taken by user 2 using terminal 120 .
  • the terminal 110 can return to the upper layer through the abstract interface layer, and the upper layer can push the relevant data.
  • the above-mentioned examples are only exemplary, and the present disclosure is not limited thereto.
  • Fig. 2 is a flow chart of a video live broadcast method according to an embodiment of the present disclosure.
  • the video live broadcast method in FIG. 2 may be executed by a first device (such as terminal 110). Before executing the video live broadcast method in FIG. 2 , the first device may first be communicatively connected to an external device (such as the second device or the terminal 120 ) to realize data interaction.
  • the video live broadcast method in Fig. 2 includes S201-S203.
  • At S201 acquire a first live video collected by a first device.
  • user 1 may use a first device to perform a live broadcast to obtain a first live video.
  • the second live data includes a second region of interest in a second live video collected by the second device.
  • the second region of interest may be obtained by performing target region extraction on each frame of the second live video.
  • the second region of interest may be a portrait region or a certain part of the portrait.
  • the second device Before the second device sends data to the first device, it can first extract the target area of the second live video collected by itself, such as image matting processing, so as to separate the second region of interest and the second background in the second live video . Thereafter, the second device may send the second region of interest data to the first device. For example, user 2 uses a second device to perform a live broadcast to obtain a second live video, and the second device performs cutout processing on the second live video to separate the portrait area from the background area, and then sends the separated portrait area to the second live video. a device.
  • the second device may use a method based on a ternary map (Trimap) to perform matting processing on the second live video, or may use a neural network to separate the second region of interest from the second background in the second live video.
  • the second device may encode the extracted second region-of-interest data, and then send it to the first device.
  • Trimap ternary map
  • the second device may perform image segmentation processing on each video frame in the second live video collected by itself, so as to extract the second region of interest and the second background region. For example, deep learning techniques can be used to implement image segmentation processing. The second device then sends the second region of interest to the first device.
  • a target live video is generated based on the first live video and the second live data.
  • the electronic device can analyze the data.
  • the first device may display the image formed by the second live data at a predetermined position in the first live video, so that the first device can At least part of the video data from the second device is displayed while displaying the first live video.
  • the first device may receive a portrait (such as the portrait of the host) in the second live video captured by the second device, and then superimpose the received portrait on the first live video to generate the target live video.
  • the portrait of the first live video and the portrait of the second live video can be displayed in the video at the same time, as if two anchors are communicating and interacting in the same background.
  • Fig. 3 shows the live image collected by user 1 using the first device
  • Fig. 4 shows the live image collected by user 2 using the second device
  • the second device can carry out the live image collected by itself
  • the image matting process for example, separates the puppy in the live broadcast picture as the second region of interest from the second background, and then sends the live broadcast data related to the puppy to the first device.
  • the first device After the first device receives the relevant data, it displays the image formed by the received data on the live screen captured by the first device, as shown in FIG. 5 .
  • the background in the target live video displayed in the first device can be changed arbitrarily according to user selection. For example, backgrounds in different live videos may be fused to obtain a fused background, or backgrounds in other videos may be replaced with backgrounds in the target live video, or desired backgrounds selected by users may be used as backgrounds in the target live video.
  • the first device when the first device receives the second region of interest and the second background in the second live video, the first device can obtain the first The first region of interest and the first background in the live video, and then the first background and the second background are fused to generate a fused background as the background in the target live video.
  • the first device may use the neural network model to perform image segmentation processing on the live broadcast picture in Figure 3 to obtain the corresponding first region of interest and the first background, and then fuse the background in Figure 3 and the background in Figure 4 to obtain the fusion background, and then display the pigeon in Figure 3 and the puppy in Figure 4 against the blended background.
  • the first device may extract the target region by performing target region extraction on each frame in the first live video To obtain the first ROI in the first live video, use the second background as the background in the target live video, and display the first ROI and the second ROI in the background.
  • the first device may replace the background shown in FIG. 4 with the live broadcast picture in FIG. 3 , that is, the first device displays the pigeon in FIG. 3 , the background and the puppy in FIG. 4 .
  • the user can select other backgrounds as the background in the target live video, and then integrate the selected background with the pigeon in FIG. 3 and the puppy in FIG. 4 to generate the target live video.
  • the first device may change the size or position of the ROI in the target live video according to user input.
  • the user can move the area of interest corresponding to the dog, such as moving up or moving to the left, to change the position of the area.
  • the user can move the region of interest corresponding to the person to change the location of the region.
  • users can change the size of the region of interest by stretching it.
  • the user can zoom in or zoom out the region of interest corresponding to the dog, or zoom in or zoom out the region of interest corresponding to the person via the touch screen of the first device.
  • the target live video is sent.
  • the first device can send the generated target live video to the server, so that the target live video can be forwarded to the audience of the first device via the server. In this way, the viewer can watch the data collected by multiple electronic devices at the same time.
  • the live content of multiple devices can be displayed, which greatly enriches the scene of the live broadcast and enhances the interactivity of the live broadcast.
  • Fig. 9 is a flow chart of a video live broadcast method according to an embodiment of the present disclosure.
  • the video live broadcast method in FIG. 9 can be executed by a terminal or a server, and includes S901-S904.
  • the first live data collected by the first device is obtained, wherein the first live data may include the first region of interest in the first live video collected by the first device and the first background at least one.
  • the first ROI may be a portrait area.
  • the first device may obtain the first region of interest and the first background by performing matting processing on each frame of the first live video.
  • the first device may send the first live video to the server, and the server obtains the first region of interest and the first background by performing matting processing on the first live video.
  • the second device acquire second live data collected by the second device, where the second live data includes at least one of a second region of interest and a second background in the second live video collected by the second device.
  • the second ROI may be a portrait area.
  • the second device may obtain the second region of interest and the second background by performing matting processing on each frame of the second live video.
  • the second device may send the second live video to the server, and the server obtains the second region of interest and the second background by performing matting processing on the second live video.
  • a target live video is generated based on the first live data and the second live data.
  • the first background and the second background may be fused to generate a fused background as the background in the target live video, and the first region of interest and the second region of interest are displayed in the fused background.
  • the first background or the second background can be selected as the background in the target live video, and the first region of interest and the second region of interest are displayed in the selected background.
  • the server may perform cutout processing on the first live video and the second live video, and perform fusion processing.
  • the first device and the second device perform matting processing on the live video collected by themselves, and then the server may receive the first region of interest and/or the first background and the second region of interest from the first device and the second device. region of interest and/or the second background, and perform fusion processing on the received video data.
  • the target live video is sent.
  • the first device and/or the second device may change the size or position of the ROI in the target live video according to user input.
  • regions of interest and/or backgrounds in different live videos can be fused, enhancing the interactivity of the live broadcast and enriching the live broadcast scene, thereby giving users more choices.
  • Fig. 10 is a schematic flowchart of live video interaction according to an embodiment of the present disclosure.
  • anchor A can shoot a live video via the first device, and the first device can perform video processing on the captured video, for example, perform beautification processing, filter processing or matting processing on the video (such as combining the background and the portrait area separation), etc., and then the first device can compress and encode the processed video data, and then send the encoded data to a server, such as an MCU streaming media server.
  • a server such as an MCU streaming media server.
  • the host B can shoot and shoot live video via the second device, and the second device can perform video processing on the captured video, for example, perform beauty treatment, filter processing or matting processing on the video (such as separating the background from the portrait area), etc. , and then the second device can compress and encode the processed video data, and then send the encoded data to the streaming media server.
  • video processing on the captured video for example, perform beauty treatment, filter processing or matting processing on the video (such as separating the background from the portrait area), etc.
  • the second device can compress and encode the processed video data, and then send the encoded data to the streaming media server.
  • the streaming media server can perform fusion processing on the received video data.
  • the streaming server can fuse the background shot by anchor A with the background shot by anchor B, and then display the portrait of anchor A and the portrait of anchor B on the fused background to generate the target video, and then send the target video to The first device where anchor A is located and the second device where anchor B is located.
  • the streaming media server can generate different target videos for different devices. For example, for the first device where the anchor A is located, the streaming media server may display the portrait of the anchor A and the portrait of the anchor B in the background shot by the anchor A, so as to generate target video data for the anchor A. For the second device where the anchor B is located, the streaming media server can display the portrait of the anchor A and the portrait of the anchor B in the background shot by the anchor B, so as to generate target video data for the anchor B.
  • the above examples are only exemplary, and different target videos may be generated according to user selections.
  • the first device where anchor A is located and the second device where anchor B is located will respectively perform video compression encoding on the received target video, and encapsulate and push the stream to the CDN distribution network.
  • the viewers at anchor A and the audience at anchor B can respectively obtain the corresponding target videos via the CND distribution network.
  • FIG. 6 is a schematic structural diagram of a live video device in a hardware operating environment according to an embodiment of the disclosure.
  • the live video device 600 may include: a processing component 601 , a communication bus 602 , a network interface 603 , an input and output interface 604 , a memory 605 and a power supply component 606 .
  • the communication bus 602 is used to realize connection and communication between these components.
  • the input and output interface 604 may include a video display (such as a liquid crystal display), a microphone and a speaker, and a user interaction interface (such as a keyboard, mouse, touch input device, etc.), and in some embodiments, the input and output interface 604 may also include a standard Wired interface, wireless interface.
  • the network interface 603 may include a standard wired interface and a wireless interface (such as a wireless fidelity interface).
  • the memory 605 can be a high-speed random access memory, or a stable non-volatile memory.
  • the memory 605 may also be a storage device independent of the aforementioned processing component 601 .
  • FIG. 6 does not constitute a limitation to the live video device 600, and may include more or less components than those shown in the figure, or combine some components, or arrange different components.
  • memory 605 as a storage medium may include an operating system (such as a MAC operating system), a data storage module, a network communication module, a user interface module, a live video program, and a database.
  • an operating system such as a MAC operating system
  • the network interface 603 is mainly used for data communication with external electronic devices/terminals; the input and output interface 604 is mainly used for data interaction with users; the processing component 601 in the live video device 600 .
  • the memory 605 can be set in the live video device 600.
  • the live video device 600 calls the live video program stored in the memory 605 and various APIs provided by the operating system through the processing component 601 to execute the live video method provided by the embodiment of the present disclosure. .
  • the processing component 601 may include at least one processor, and the memory 605 stores a set of computer-executable instructions. When the set of computer-executable instructions is executed by the at least one processor, the video live broadcast method according to the embodiment of the present disclosure is executed. Additionally, the processing component 601 can perform encoding operations, decoding operations, and the like. However, the above-mentioned examples are only exemplary, and the present disclosure is not limited thereto.
  • the input-output interface 604 may receive an input or instruction for searching for a target device such as a second device. Based on the first input or instruction, the network interface 603 may search for a target device, and communicate and connect the live video device 600 to the searched target device.
  • the input-output interface 604 may display a user interface including identifications of the searched external electronic devices, and receive an input or instruction for selecting a target device from the searched external electronic devices via the user interface.
  • the network interface 603 can communicate the live video device 600 with the target device.
  • the processing component 601 can obtain the first live video collected by the live video device, and receive the second live data sent by the target device, wherein the second live data can include the data collected by the target device.
  • the second region of interest in the second live video of the target live video is generated based on the first live video and the second live data, and the target live video is sent to the server, so that the target live video can be forwarded to the live video via the server The viewer end of the device 600 .
  • the target device may obtain the second region of interest by extracting the target region from each frame of the second live video collected by itself.
  • the target device may send the second live video collected by itself to the live video device 600, and then the live video device 600 performs cutout processing on the received second live video to extract the content of the second live video.
  • Second region of interest For example, the second region of interest may be a portrait region.
  • the processing component 601 can obtain the first live video by extracting the target area from each frame in the first live video The first region of interest and the first background in , and then fuse the first background and the second background to generate a fused background as the background in the target live video.
  • the processing component 601 can extract the target region by performing target region extraction on each frame in the first live video To obtain the first ROI in the first live video, use the second background as the background in the target live video, and display the first ROI and the second ROI in the background.
  • the target live video including the user portrait of the video live broadcast device 600 , the user portrait of the target device, and the fusion background can be displayed via the input and output interface 604 .
  • the target live video including the user portrait of the live video device 600 , the user portrait of the target device and the video background collected by the target device can be displayed via the input and output interface 604 .
  • the background in the target live video can also be changed arbitrarily according to user selection/input.
  • the live video device 600 may receive user input, and change the size or position of the region of interest (such as the user portrait of the video live device 600 and the user portrait of the target device) displayed in the target live video according to the user input.
  • the region of interest such as the user portrait of the video live device 600 and the user portrait of the target device
  • the processing component 601 can realize the control of the components included in the live video device 600 by executing a program.
  • the live video device 600 can receive or output video and/or audio via the input and output interface 604 .
  • the user can output the merged live content via the input/output interface 604 to share with viewers.
  • the live video device 600 may be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above-mentioned set of instructions.
  • the video live broadcast device 600 is not necessarily a single electronic device, but may also be any assembly of devices or circuits capable of individually or jointly executing the above-mentioned instructions (or instruction sets).
  • the live video device 600 may also be part of an integrated control system or system manager, or may be configured as a portable electronic device that interfaces locally or remotely (eg, via wireless transmission).
  • the processing component 601 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller or a microprocessor.
  • the processing component 601 may also include analog processors, digital processors, microprocessors, multi-core processors, processor arrays, network processors, and the like.
  • the processing component 601 can execute instructions or codes stored in the memory, wherein the memory 605 can also store data. Instructions and data can also be sent and received through the network via the network interface 603, wherein the network interface 603 can adopt any known transmission protocol.
  • the memory 605 may be integrated with the processing component 601, for example, by placing RAM or flash memory within an integrated circuit microprocessor or the like. Additionally, storage 605 may comprise a separate device, such as an external disk drive, storage array, or any other storage device usable by the database system.
  • the memory and processing component 601 may be operatively coupled, or may communicate with each other, eg, through I/O ports, network connections, etc., such that the processing component 601 can read data stored in the memory 605 .
  • FIG. 7 is a block diagram of a live video device according to an embodiment of the present disclosure.
  • a video live broadcast device 700 may include an input module 701 , a communication module 702 , a receiving module 703 , a processing module 704 , a display module 705 and a collection module 706 .
  • Each module in the video live broadcast device 700 may be implemented by one or more modules, and the names of the corresponding modules may vary according to the type of the module. In various embodiments, some modules in the live video device 700 may be omitted, or additional modules may also be included. Also, modules/elements according to various embodiments of the present disclosure may be combined to form a single entity, and thus may equivalently perform the functions of the corresponding modules/elements before combination.
  • the collection module 706 is configured to collect live data (such as the first live video).
  • the input module 701 may be configured to receive user input or instructions.
  • the user input may be, for example, a touch input, a button input, a hover input, a gesture input, and the like.
  • the communication module 702 may be configured to search for and communicate with external electronic devices.
  • the receiving module 703 may be configured to receive data from an external electronic device.
  • the processing module 704 may be configured to process data received from external electronic devices and data collected/obtained by itself.
  • the display module 705 may be configured to display data received from external electronic devices and data collected/obtained by itself, for example, the final synthesized target live video.
  • the input module 701 may receive an input or instruction for searching for a target device. Based on an input or instruction, the communication module 702 can search for a target device, and communicate and connect the live video device 700 with the searched target device.
  • the receiving module 703 may receive the second live data sent by the second device (ie, the target device), where the second live data includes the second region of interest in the second live video collected by the second device.
  • the second region of interest may be obtained by performing target region extraction on each frame of the second live video.
  • the second region of interest may be a portrait region.
  • the processing module 704 can generate a target live video based on the first live video and the second live data.
  • the sending module (not shown) can send the target live video. In some embodiments, sending the target live video may be implemented by the processing module 704 .
  • the processing module 704 can extract the target area from each frame of the first live video to obtain the first live video In the first region of interest and the first background, the first background and the second background are fused to generate a fused background as the background in the target live video.
  • the processing module 704 can extract the target area from each frame of the first live video to obtain the first live video the first region of interest in ; use the second background as the background in the target live video, and display the first region of interest and the second region of interest in the background.
  • the input module 701 may receive user input, and then the processing module 704 changes the position or size of the ROI in the target live video according to the user input. For example, the user may drag a second region of interest (such as a portrait in the second live video) in the target live video to change the position or size of the second region of interest.
  • a second region of interest such as a portrait in the second live video
  • the live video device 700 may include an acquisition module, a processing module and a sending module.
  • the acquisition module can receive live data from different external devices.
  • the processing module fuses the received live data to generate a target live video, and then the sending module sends the target live video to a server or a corresponding external device.
  • multiple devices can be used for data collection, and the data collected by multiple devices can be merged and streamed, thereby realizing a distributed live broadcast service.
  • an electronic device may include at least one memory 802 and at least one processor 801.
  • the at least one memory 802 stores a set of computer-executable instructions.
  • the computer-executable instructions When the set is executed by at least one processor 801, the video live broadcast method according to the embodiment of the present disclosure is executed.
  • Processor 801 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a special purpose processor system, a microcontroller, or a microprocessor.
  • the processor 801 may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, and the like.
  • the memory 802 as a storage medium may include an operating system (eg, MAC operating system), a data storage module, a network communication module, a user interface module, a live video program, and a database.
  • an operating system eg, MAC operating system
  • the memory 802 can be integrated with the processor 801, for example, RAM or flash memory can be arranged in an integrated circuit microprocessor or the like. Additionally, storage 802 may comprise a separate device, such as an external disk drive, storage array, or any other storage device usable by the database system.
  • the memory 802 and the processor 801 may be operatively coupled, or may communicate with each other, eg, through an I/O port, network connection, etc., such that the processor 801 can read files stored in the memory 802 .
  • the electronic device 800 may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, mouse, touch input device, etc.). All components of the electronic device 800 may be connected to each other via a bus and/or a network.
  • a video display such as a liquid crystal display
  • a user interaction interface such as a keyboard, mouse, touch input device, etc.
  • FIG. 8 does not constitute a limitation, and may include more or less components than shown in the figure, or combine some components, or arrange different components.
  • a non-volatile computer-readable storage medium on which instructions are stored, wherein, when the instructions are executed by at least one processor, at least one processor is caused to execute the method according to the present disclosure. Live video method.
  • Non-volatile computer-readable storage medium examples include: Read Only Memory (ROM), Random Access Programmable Read Only Memory (PROM), Electrically Erasable Programmable Read Only Memory (EEPROM), Random Access Memory (RAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Flash, Nonvolatile Memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW , DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or Disc memory, hard disk drive ( HDD), solid state drive (SSD), card memory (such as MultiMediaCard, Secure Digital (SD) or Extreme Digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, Solid state disks and any other device configured to store a computer program and any associated data, data files and data structures in a non-trans
  • the computer program in the above-mentioned non-transitory computer-readable storage medium can run in an environment deployed in computer equipment such as a client, a host, an agent device, a server, etc.
  • the computer program and any associated The data, data files and data structures are distributed over networked computer systems so that the computer programs and any associated data, data files and data structures are stored, accessed and executed in a distributed fashion by one or more processors or computers.
  • a computer program product may also be provided, and instructions in the computer program product may be executed by a processor of a computer device to complete the above video live broadcast method.
  • any references to memory, storage, database or other media used in the various embodiments provided in the present application may include non-volatile and/or volatile memory.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Databases & Information Systems (AREA)
  • Business, Economics & Management (AREA)
  • Marketing (AREA)
  • Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)

Abstract

本公开提供一种视频直播方法和视频直播装置,涉及视频处理技术领域。所述视频直播方法可以由第一设备执行,所述视频直播方法包括:获取由第一设备采集的第一直播数据,其中第一直播数据包括由第一设备采集的第一直播视频中的第一感兴趣区域和第一背景中的至少一个;获取由第二设备采集的第二直播数据,其中第二直播数据包括由第二设备采集的第二直播视频中的第二感兴趣区域和第二背景中的至少一个;基于第一直播数据和第二直播数据来生成目标直播视频;发送所述目标直播视频。

Description

视频直播方法和视频直播装置
相关申请的交叉引用
本公开基于申请日为2021年5月27日、申请号为202110584596.4的中国专利申请,并要求该中国专利申请的优先权,在此全文引用上述中国专利申请公开的内容以作为本公开的一部分。
技术领域
本公开涉及视频处理技术领域,尤其涉及一种视频直播方法和视频直播装置。
背景技术
近来,随着互联网技术的迅猛发展,视频直播业务已成为当今潮流。在直播过程中,主播可以通过直播软件进行视频直播,并且主播之间也可以实现在线互动。然而,目前的直播互动方式较为单一。
发明内容
本公开提供一种视频直播方法和视频直播装置。
根据本公开实施例的第一方面,提供一种视频直播方法,所述视频直播方法可以包括:获取由第一设备采集的第一直播数据,其中,第一直播数据包括由第一设备采集的第一直播视频中的第一感兴趣区域和第一背景中的至少一个;获取由第二设备采集的第二直播数据,其中,第二直播数据包括由第二设备采集的第二直播视频中的第二感兴趣区域和第二背景中的至少一个;基于第一直播数据和第二直播数据来生成目标直播视频;发送所述目标直播视频。
在一些实施例中,感兴趣区域可以是通过对第一直播视频和/或第二直播视频中的每一帧进行目标区域提取而获得的。
在一些实施例中,感兴趣区域可以是人像区域。
在一些实施例中,基于第一直播数据和第二直播数据来生成目标直播视频的步骤可以包括:将第一背景与第二背景进行融合来生成融合背景作为所述目标直播视频中的背景;在所述融合背景中显示第一感兴趣区域和第二感兴趣区域。
在一些实施例中,基于第一直播数据和第二直播数据来生成目标直播视频的步骤可以包括:选择第一背景或第二背景作为所述目标直播视频中的背景,并且在选择的背景中显示第一感兴趣区域和第二感兴趣区域。
在一些实施例中,第一感兴趣区域和/或第二感兴趣区域可以基于用户输入在所述目标直播视频中改变位置或尺寸。
根据本公开实施例的第二方面,提供一种视频直播装置,所述视频直播装置可以包括: 获取模块,被配置为:获取由第一设备采集的第一直播数据和由第二设备采集的第二直播数据,其中,第一直播数据包括由第一设备采集的第一直播视频中的第一感兴趣区域和第一背景中的至少一个,并且第二直播数据包括由第二设备采集的第二直播视频中的第二感兴趣区域和第二背景中的至少一个;处理模块,被配置为基于第一直播数据和第二直播数据来生成目标直播视频;以及发送模块,被配置为发送所述目标直播视频。
在一些实施例中,感兴趣区域可以是通过对第一直播视频和/或第二直播视频中的每一帧进行目标区域提取而获得的。
在一些实施例中,感兴趣区域可以是人像区域。
在一些实施例中,处理模块可以被配置为:将第一背景与第二背景进行融合来生成融合背景作为所述目标直播视频中的背景;在所述融合背景中显示第一感兴趣区域和第二感兴趣区域。
在一些实施例中,处理模块可以被配置为:选择第一背景和/或第二背景作为所述目标直播视频中的背景,并且在选择的背景中显示第一感兴趣区域和第二感兴趣区域。
在一些实施例中,第一感兴趣区域和/或第二感兴趣区域可以基于用户输入在所述目标直播视频中改变位置或尺寸。
根据本公开实施例的第三方面,提供一种电子设备,所述电子设备可以包括:至少一个处理器;至少一个存储计算机可执行指令的存储器,其中,所述计算机可执行指令在被所述至少一个处理器运行时,促使所述至少一个处理器执行如上所述的视频直播方法。
根据本公开实施例的第四方面,提供一种非易失性计算机可读存储介质,其上存储有指令,当所述指令被至少一个处理器运行时,促使所述至少一个处理器执行如上所述的视频直播方法。
根据本公开实施例的第五方面,提供一种计算机程序产品,所述计算机程序产品中的指令被电子装置中的至少一个处理器运行以执行如上所述的视频直播方法。
本公开的实施例提供的技术方案至少带来以下有益效果:
本公开可以将不同直播视频中的感兴趣区域和/或背景进行融合,增强了直播的交互性并且丰富了直播场景,给用户更多选择。
应当理解的是,以上的一般描述和后文的细节描述仅是示例性和解释性的,并不能限制本公开。
附图说明
此处的附图被并入说明书中并构成本说明书的一部分,示出了符合本公开的实施例,并与说明书一起用于解释本公开的原理,并不构成对本公开的不当限定。
图1是根据本公开的一些实施例的用于直播交互的应用环境的示图;
图2是根据本公开的一些实施例的视频直播方法的流程图;
图3至图5是根据本公开的一些实施例的直播画面的示意图;
图6是根据本公开的一些实施例的视频直播设备的结构示意图;
图7是根据本公开的一些实施例的视频直播装置的框图;
图8是根据本公开的一些实施例的电子设备的框图;
图9是根据本公开的一些实施例的视频直播方法的流程图;
图10是根据本公开的一些实施例的视频直播交互的流程示意图。
在整个附图中,应注意,相同的参考标号用于表示相同或相似的元件、特征和结构。
具体实施方式
为了使本领域普通人员更好地理解本公开的技术方案,下面将结合附图,对本公开实施例中的技术方案进行清楚、完整地描述。
提供参照附图的以下描述以帮助对由权利要求及其等同物限定的本公开的实施例的全面理解。包括各种特定细节以帮助理解,但这些细节仅被视为是示例性的。因此,本领域的普通技术人员将认识到在不脱离本公开的范围和精神的情况下,可对描述于此的实施例进行各种改变和修改。此外,为了清楚和简洁,省略对公知的功能和结构的描述。
以下描述和权利要求中使用的术语和词语不限于书面含义,而仅由发明人用来实现本公开的清楚且一致的理解。因此,本领域的技术人员应清楚,本公开的各种实施例的以下描述仅被提供用于说明目的而不用于限制由权利要求及其等同物限定的本公开的目的。
需要说明的是,本公开的说明书和权利要求书及上述附图中的术语“第一”、“第二”等是用于区别类似的对象,而不必用于描述特定的顺序或先后次序。应该理解这样使用的数据在适当情况下可以互换,以便这里描述的本公开的实施例能够以除了在这里图示或描述的那些以外的顺序实施。以下示例性实施例中所描述的实施方式并不代表与本公开相一致的所有实施方式。相反,它们仅是与如所附权利要求书中所详述的、本公开的一些方面相一致的装置和方法的例子。
在下文中,根据本公开的各种实施例,将参照附图对本公开的方法和装置进行详细描述。
图1是根据本公开的实施例的用于直播交互的应用环境的示图。
参照图1,该应用环境100包括终端110、终端120和服务器130。
终端110可以是用户所在终端,例如,主播进行直播时所使用的终端。终端110可以是智能手机、平板电脑、便携式计算机和台式计算机等中的至少一种。终端110可以安装有目标应用,用于诸如从外部设备接收数据和向外部设备发送数据、对采集的视频进行抠图处理等。
终端120可以是用户所在终端,例如,主播进行直播时所使用的终端。终端120可以是智能手机、平板电脑、便携式计算机和台式计算机等中的至少一种。终端120可以安装有目标应用,用于诸如从外部设备接收数据和向外部设备发送数据、对采集的视频进行抠 图处理等。虽然本实施例仅示出一个终端120进行说明,但是本领域技术人员可知晓,终端120的数量可以为一个或两个以上。本公开实施例不对终端120的数量和设备类型进行任何限定。
终端110可以通过无线网络与终端120连接,使得终端110与终端120之间可以进行数据交互。例如,网络可以包含蓝牙、局域网(LAN)、广域网(WAN)、无线链路、内联网、互联网或其组合等。
在图1中,以终端110作为示例,终端110可以从终端120接收直播数据,并且将终端110的直播数据与终端120的直播数据进行融合。在本公开中,终端110接收的直播数据可以是例如由终端120采集的视频中的感兴趣区域(诸如人像区域)。下面将参照图2详细描述关于这方面的实施例。在一些实施例中,终端120可以从终端110接收直播数据,并且将终端120的直播数据与终端110的直播数据进行融合。作为又一示例,终端110和终端120可以将各自拍摄的直播数据发送到服务器130,服务器130可以将对接收到的直播数据进行融合。下面将参照图9详细描述关于这方面的实施例。
终端110可以通过无线网络与服务器130连接,使得终端110与服务器130之间可以进行数据交互。例如,网络可以包含局域网(LAN)、广域网(WAN)、无线链路、内联网、互联网或其组合等。此外,终端110也可以通过有线网络与服务器130连接,以进行数据交互。
服务器130可以是用于对接收到的数据进行解析处理的服务器。服务器130可以从终端110接收直播数据,将接收的直播数据推流给终端110的观众端。例如,主播利用终端110采集直播数据,并且经由网络接收终端120实时采集的直播数据中的感兴趣区域数据,终端110可以将自身采集的数据和从终端120接收的数据进行融合,最终将融合数据经由网络发送至服务器130。服务器130可以将接收的融合数据转发给终端110的观众端。这样,观众端可以同时观看来自终端110和终端120的直播内容。
假设用户1使用终端110进行直播并且用户2使用终端120进行直播,终端110可以接收由终端120传输的直播数据。在终端120向终端110传输直播数据之前,终端120可以首先对采集的视频画面进行抠图处理,使得画面中的感兴趣区域(诸如用户2的人像)与背景分离,然后将分离出的感兴趣区域数据传输给终端110。在终端110接收到来自终端120的感兴趣区域数据后,终端110可以将自身采集的视频与接收的数据进行合流,以将来自终端120的感兴趣区域显示在由终端110采集的视频中,诸如在终端110中可以同时显示用户1的人像和用户2的人像、或者由用户2使用终端120拍摄的画面中的其他感兴趣区域。最后终端110可以通过抽象的接口层返回给上层,上层把相关的数据进行推流即可。然而,上述示例仅是示例性的,本公开不限于此。
图2是根据本公开的实施例的视频直播方法的流程图。图2的视频直播方法可以由第一设备(诸如终端110)执行。在执行图2的视频直播方法前,可以首先将第一设备与外部设备(诸如第二设备或者终端120)进行通信连接以实现数据交互。图2的视频直播方 法包括S201~S203。
在S201,获取由第一设备采集的第一直播视频。例如,用户1可以利用第一设备进行直播以获得第一直播视频。
在S202,接收由第二设备发送的第二直播数据,其中,第二直播数据包括由第二设备采集的第二直播视频中的第二感兴趣区域。该第二感兴趣区域可以通过对第二直播视频中的每一帧进行目标区域提取而获得的。例如,该第二感兴趣区域可以是人像区域或者人像的某个部位区域。
第二设备在向第一设备发送数据前,可以首先对自身采集的第二直播视频进行目标区域提取,如抠图处理,以将第二直播视频中的第二感兴趣区域和第二背景分离。之后,第二设备可以向第一设备发送第二感兴趣区域数据。例如,用户2使用第二设备进行直播以获得第二直播视频,第二设备对第二直播视频进行抠图处理,以将人像区域和背景区域分离开,然后将分离出的人像区域发送给第一设备。第二设备可以利用基于三元图(Trimap)的方法对第二直播视频进行抠图处理,或者可以利用神经网络来实现对第二直播视频中的第二感兴趣区域和第二背景的分离。第二设备可以将提取的第二感兴趣区域数据进行编码,然后发送给第一设备。
在一些实施例中,第二设备可以对自身采集的第二直播视频中的每个视频帧进行图像分割处理,以提取出第二感兴趣区域和第二背景区域。例如,可以利用深度学习技术来实现图像分割处理。第二设备然后将第二感兴趣区域发送给第一设备。
在S203,基于第一直播视频和第二直播数据来生成目标直播视频。电子设备在接收到第二设备的数据后,可以对该数据进行解析。在第二直播数据仅包括第二直播视频中的第二感兴趣区域数据时,第一设备可以将由第二直播数据形成的图像显示在第一直播视频的预定位置处,使得第一设备可以在显示第一直播视频的同时显示来自第二设备的至少部分视频数据。例如,第一设备可以接收由第二设备采集的第二直播视频中的人像(诸如主播人像),然后将接收到的人像叠加在第一直播视频上以生成目标直播视频,这样,在目标直播视频中可以同时显示第一直播视频中的人像和第二直播视频中的人像,犹如两个主播在同一背景下进行交流互动。
参照图3至图5,图3示出了用户1使用第一设备采集的直播画面,图4示出了用户2使用第二设备采集的直播画面,第二设备可以对自身采集的直播画面进行抠图处理,例如,将该直播画面中的小狗作为第二感兴趣区域与第二背景进行分离,然后将与小狗相关的直播数据发送给第一设备。第一设备在接收到相关数据后,将由所接收的数据形成的图像显示在第一设备采集的直播画面中,如图5所示。
根据本公开的实施例,在第一设备中显示的目标直播视频中的背景可以根据用户选择任意改变。例如,可以将不同直播视频中的背景进行融合来获得融合背景,或者可以将其他视频中的背景替换为目标直播视频中的背景,或者可以将用户选择的期望背景作为目标直播视频中的背景。
作为示例,在第一设备接收第二直播视频中的第二感兴趣区域和第二背景的情况下,第一设备可以通过对第一直播视频中的每一帧进行目标区域提取来获得第一直播视频中的第一感兴趣区域和第一背景,然后将第一背景与第二背景进行融合来生成融合背景作为目标直播视频中的背景。例如,第一设备可以利用神经网络模型对图3的直播画面进行图像分割处理以获得相应的第一感兴趣区域和第一背景,然后将图3的背景与图4的背景进行融合以获得融合背景,然后在融合背景下显示图3中的鸽子和图4中的小狗。
在一些实施例中,在第二直播数据包括第二直播视频中的第二感兴趣区域和第二背景的情况下,第一设备可以通过对第一直播视频中的每一帧进行目标区域提取来获得第一直播视频中的第一感兴趣区域,使用第二背景作为目标直播视频中的背景,并且在该背景中显示第一感兴趣区域和第二感兴趣区域。例如,第一设备可以将图4所示的背景替换到图3的直播画面中,即第一设备显示图3的鸽子、图4的背景和小狗。在一些实施例中,用户可以选择其他背景作为目标直播视频中的背景,然后将选择的背景与图3中的鸽子和图4中的小狗进行整合以生成目标直播视频。
此外,在显示目标直播视频的过程中,第一设备可以根据用户输入来改变目标直播视频中的感兴趣区域的尺寸或位置。例如,参照图5,用户可以移动与小狗相应的感兴区域,如向上移动、向左移动,以改变该区域的位置。或者,用户可以移动与人相应的感兴趣区域,以改变该区域的位置。此外,用户可以通过拉伸感兴趣区域来改变该区域的尺寸。例如,用户可以经由第一设备的触摸屏来放大或缩小与小狗相应的感兴趣区域,或者放大或缩小与人物相应的感兴趣区域。
在S204,发送目标直播视频。第一设备可以将生成的目标直播视频发送至服务器,使得目标直播视频可以经由服务器转发给第一设备的观众端。这样,观众端可以同时观看多个电子设备采集的数据。
根据本公开的实施例,可以显示多个设备直播的内容,极大地丰富了直播的场景,增强了直播的互动性。
图9是根据本公开的实施例的视频直播方法的流程图。图9的视频直播方法可以由终端或者服务器执行,并且包括S901~S904。
参照图9,在S901,获取由第一设备采集的第一直播数据,其中,第一直播数据可以包括由第一设备采集的第一直播视频中的第一感兴趣区域和第一背景中的至少一个。这里,第一感兴趣区域可以是人像区域。例如,第一设备可以通过对第一直播视频的每一帧进行抠图处理来获得第一感兴趣区域和第一背景。在一些实施例中,第一设备可以将第一直播视频发送到服务器,由服务器通过对第一直播视频进行抠图处理来获得第一感兴趣区域和第一背景。
在S902,获取由第二设备采集的第二直播数据,其中,第二直播数据包括由第二设备采集的第二直播视频中的第二感兴趣区域和第二背景中的至少一个。这里,第二感兴趣区域可以是人像区域。例如,第二设备可以通过对第二直播视频的每一帧进行抠图处理来 获得第二感兴趣区域和第二背景。在一些实施例中,第二设备可以将第二直播视频发送到服务器,由服务器通过对第二直播视频进行抠图处理来获得第二感兴趣区域和第二背景。
在S903,基于第一直播数据和第二直播数据来生成目标直播视频。作为示例,可以将第一背景与第二背景进行融合来生成融合背景作为目标直播视频中的背景,并且在融合背景中显示第一感兴趣区域和第二感兴趣区域。在一些实施例中,可以选择第一背景或第二背景作为目标直播视频中的背景,并且在选择的背景中显示第一感兴趣区域和第二感兴趣区域。
当由服务器执行上述S901至S903时,服务器可以对第一直播视频和第二直播视频进行抠图处理,并且执行融合处理。在一些实施例中,第一设备和第二设备对自身采集的直播视频进行抠图处理,然后服务器可以从第一设备和第二设备接收第一感兴趣区域和/或第一背景以及第二感兴趣区域和/或第二背景,并且对接收到的视频数据进行融合处理。
在S904,发送目标直播视频。
此外,在显示目标直播视频的过程中,第一设备和/或第二设备可以根据用户输入来改变目标直播视频中的感兴趣区域的尺寸或位置。
通过上述处理,可以将不同直播视频中的感兴趣区域和/或背景进行融合,增强了直播的交互性并且丰富了直播场景,从而给用户更多选择。
图10是根据本公开的实施例的视频直播交互的流程示意图。
参照图10,主播A可以经由第一设备拍摄直播视频,第一设备可以对拍摄的视频进行视频处理,例如,对视频进行美颜处理、滤镜处理或抠图处理(诸如将背景与人像区域分离)等,然后第一设备可以将视频处理后的数据进行压缩编码,然后将编码数据发送到服务器,诸如MCU流媒体服务器。
主播B可以经由第二设备拍摄拍摄直播视频,第二设备可以对拍摄的视频进行视频处理,例如,对视频进行美颜处理、滤镜处理或抠图处理(诸如将背景与人像区域分离)等,然后第二设备可以将视频处理后的数据进行压缩编码,然后将编码数据发送到流媒体服务器。
流媒体服务器可以对接收到的视频数据进行融合处理。例如。流媒体服务器可以将主播A拍摄的背景与主播B拍摄的背景进行融合,然后将主播A的人像和主播B的人像显示在融合后的背景上,以生成目标视频,然后将目标视频分别发送给主播A所在的第一设备和主播B所在的第二设备。
在一些实施例中,流媒体服务器可以针对不同的设备生成不同的目标视频。例如,针对主播A所在的第一设备,流媒体服务器可以将主播A的人像和主播B的人像显示在主播A拍摄的背景中,以生成针对主播A的目标视频数据。针对主播B所在第二设备,流媒体服务器可以将主播A的人像和主播B的人像显示在主播B拍摄的背景中,以生成针对主播B的目标视频数据。然而,上述示例仅是示例性的,可以根据用户选择来生成不同的目标视频。
接下来,主播A所在的第一设备和主播B所在的第二设备分别将对接收的目标视频进行视频压缩编码,并且进行封装推流至CDN分发网络。
主播A端的观众和主播B端的观众可以分别经由CND分发网络获得相应的目标视频。
图6是本公开实施例的硬件运行环境的视频直播设备的结构示意图。
如图6所示,视频直播设备600可包括:处理组件601、通信总线602、网络接口603、输入输出接口604、存储器605以及电源组件606。其中,通信总线602用于实现这些组件之间的连接通信。输入输出接口604可以包括视频显示器(诸如,液晶显示器)、麦克风和扬声器以及用户交互接口(诸如,键盘、鼠标、触摸输入装置等),在一些实施例中,输入输出接口604还可以包括标准的有线接口、无线接口。网络接口603可选的可包括标准的有线接口、无线接口(如无线保真接口)。存储器605可以是高速的随机存取存储器,也可以是稳定的非易失性存储器。存储器605可选的还可以是独立于前述处理组件601的存储装置。
本领域技术人员可以理解,图6中示出的结构并不构成对视频直播设备600的限定,可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件布置。
如图6所示,作为一种存储介质的存储器605中可以包括操作系统(诸如MAC操作系统)、数据存储模块、网络通信模块、用户接口模块、视频直播程序以及数据库。
在图6所示的视频直播设备600中,网络接口603主要用于与外部电子设备/终端进行数据通信;输入输出接口604主要用于与用户进行数据交互;视频直播设备600中的处理组件601、存储器605可以被设置在视频直播设备600中,视频直播设备600通过处理组件601调用存储器605中存储的视频直播程序以及由操作系统提供的各种API,执行本公开实施例提供的视频直播方法。
处理组件601可以包括至少一个处理器,存储器605中存储有计算机可以执行指令集合,当计算机可以执行指令集合被至少一个处理器执行时,执行根据本公开实施例的视频直播方法。此外,处理组件601可执行编码操作和解码操作等。然而,上述示例仅是示例性的,本公开不限于此。
输入输出接口604可以接收用于搜索目标设备(诸如第二设备)的输入或指令。基于第一输入或指令,网络接口603可以搜索目标设备,并且将视频直播设备600与搜索到的目标设备进行通信连接。
作为一些实施方式,基于搜索结果,输入输出接口604可以显示包括搜索到的外部电子设备的标识的用户界面,经由该用户界面接收用于从搜索到的外部电子设备中选择目标设备的输入或指令,基于该输入或指令,网络接口603可以将视频直播设备600与目标设备进行通信连接。
在视频直播设备600与目标设备连接后,处理组件601可以获取由视频直播装置采集的第一直播视频,接收由目标设备发送的第二直播数据,其中,第二直播数据可以包括由目标设备采集的第二直播视频中的第二感兴趣区域,基于第一直播视频和第二直播数据来 生成目标直播视频,并将目标直播视频发送至服务器,从而可以将目标直播视频经由服务器转发给视频直播设备600的观众端。
根据本公开的一些实施例,在目标设备发送第二直播数据之前,目标设备可以通过对由自身采集的第二直播视频中的每一帧进行目标区域提取来获得第二感兴趣区域。在一些实施例中,目标设备可以将自身采集的第二直播视频发送给视频直播设备600,然后视频直播设备600对接收到的第二直播视频进行抠图处理,以提取第二直播视频中的第二感兴趣区域。例如,第二感兴趣区域可以是人像区域。
在第二直播数据包括第二直播视频中的第二感兴趣区域和第二背景的情况下,处理组件601可以通过对第一直播视频中的每一帧进行目标区域提取来获得第一直播视频中的第一感兴趣区域和第一背景,然后将第一背景与第二背景进行融合来生成融合背景作为目标直播视频中的背景。
在一些实施例中,在第二直播数据包括第二直播视频中的第二感兴趣区域和第二背景的情况下,处理组件601可以通过对第一直播视频中的每一帧进行目标区域提取来获得第一直播视频中的第一感兴趣区域,使用第二背景作为目标直播视频中的背景,并且在该背景中显示第一感兴趣区域和第二感兴趣区域。例如,在视频直播设备600中,可以经由输入输出接口604显示包括视频直播设备600的用户人像、目标设备的用户人像以及融合背景的目标直播视频。在一些实施例中,在视频直播设备600中,可以经由输入输出接口604显示包括视频直播设备600的用户人像、目标设备的用户人像以及由目标设备采集的视频背景的目标直播视频。此外,也可以根据用户选择/输入来任意改变目标直播视频中的背景。
此外,视频直播设备600可以接收用户输入,并根据该用户输入来改变显示在目标直播视频中的感兴趣区域(诸如视频直播设备600的用户人像和目标设备的用户人像)的尺寸或位置。
处理组件601可以通过执行程序来实现对视频直播设备600所包括的组件的控制。
视频直播设备600可以经由输入输出接口604接收或输出视频和/或音频。例如,用户可经由输入输出接口604输出合流后的直播内容以分享给观看者。
作为示例,视频直播设备600可以是PC计算机、平板装置、个人数字助理、智能手机、或其他能够执行上述指令集合的装置。这里,视频直播设备600并非必须是单个的电子设备,还可以是任何能够单独或联合执行上述指令(或指令集)的装置或电路的集合体。视频直播设备600还可以是集成控制系统或系统管理器的一部分,或者可以被配置为与本地或远程(例如,经由无线传输)以接口互联的便携式电子设备。
在视频直播设备600中,处理组件601可以包括中央处理器(CPU)、图形处理器(GPU)、可编程逻辑装置、专用处理器系统、微控制器或微处理器。作为示例而非限制,处理组件601还可以包括模拟处理器、数字处理器、微处理器、多核处理器、处理器阵列、网络处理器等。
处理组件601可以运行存储在存储器中的指令或代码,其中,存储器605还可以存储数据。指令和数据还可以经由网络接口603而通过网络被发送和接收,其中,网络接口603可以采用任何已知的传输协议。
存储器605可以与处理组件601集成为一体,例如,将RAM或闪存布置在集成电路微处理器等之内。此外,存储器605可以包括独立的装置,诸如,外部盘驱动、存储阵列或任何数据库系统可以使用的其他存储装置。存储器和处理组件601可以在操作上进行耦合,或者可以例如通过I/O端口、网络连接等互相通信,使得处理组件601能够读取存储在存储器605中的数据。
图7是根据本公开的实施例的视频直播装置的框图。
参照图7,视频直播装置700可以包括输入模块701、通信模块702、接收模块703、处理模块704、显示模块705以及采集模块706。视频直播装置700中的每个模块可以由一个或多个模块来实现,并且对应模块的名称可以根据模块的类型而变化。在各种实施例中,可省略视频直播装置700中的一些模块,或者还可以包括另外的模块。此外,根据本公开的各种实施例的模块/元件可以被组合以形成单个实体,并且因此可等以效地执行相应模块/元件在组合之前的功能。
采集模块706被配置为采集直播数据(诸如第一直播视频)。
输入模块701可以被配置为接收用户输入或指令。这里,用户输入可以是例如触摸输入、按钮输入、悬停输入和手势输入等。通信模块702可以被配置为搜索外部电子设备以及与外部电子设备进行通信连接。接收模块703可以被配置为接收外部电子设备的数据。处理模块704可以被配置为处理从外部电子设备接收的数据以及自身采集/获得的数据。显示模块705可以被配置为显示从外部电子设备接收的数据以及自身采集/获得的数据,例如,最终合成的目标直播视频。
输入模块701可以接收用于搜索目标设备的输入或指令。基于输入或指令,通信模块702可以搜索目标设备,并且将视频直播装置700与搜索到的目标设备进行通信连接。
接收模块703可以接收由第二设备(即目标设备)发送的第二直播数据,其中,第二直播数据包括由第二设备采集的第二直播视频中的第二感兴趣区域。该第二感兴趣区域可以通过对第二直播视频中的每一帧进行目标区域提取而获得的。例如,该第二感兴趣区域可以是人像区域。
处理模块704可以基于第一直播视频和第二直播数据来生成目标直播视频。发送模块(未示出)可以发送目标直播视频。在一些实施例中,发送目标直播视频可以由处理模块704实现。
在第二直播数据包括第二直播视频中的第二感兴趣区域和第二背景的情况下,处理模块704可以通过对第一直播视频中的每一帧进行目标区域提取来获得第一直播视频中的第一感兴趣区域和第一背景,将第一背景与第二背景进行融合来生成融合背景作为目标直播视频中的背景。
在第二直播数据包括第二直播视频中的第二感兴趣区域和第二背景的情况下,处理模块704可以通过对第一直播视频中的每一帧进行目标区域提取来获得第一直播视频中的第一感兴趣区域;使用第二背景作为目标直播视频中的背景,并且在该背景中显示第一感兴趣区域和第二感兴趣区域。
此外,输入模块701可以接收用户输入,然后处理模块704根据该用户输入在目标直播视频中改变视频中的感兴趣区域的位置或尺寸。例如,用户可以在目标直播视频中拖拽第二直播视频中的第二感兴趣区域(诸如第二直播视频中的人像)来改变该第二感兴趣区域的位置或者尺寸。
根据本公开的一些实施例,视频直播装置700可以包括获取模块、处理模块和发送模块。获取模块可以从不同的外部设备接收直播数据。处理模块对接收的直播数据进行融合处理以生成目标直播视频,然后发送模块将目标直播视频发送到服务器或者相应的外部设备。
根据本公开的实施例,可以使用多个设备进行数据采集,并且将多个设备采集的数据进行合流并推流,从而实现了分布式直播业务。
根据本公开的实施例,可以提供一种电子设备。图8是根据本公开实施例的电子设备的框图,该电子设备800可以包括至少一个存储器802和至少一个处理器801,所述至少一个存储器802存储有计算机可执行指令集合,当计算机可执行指令集合被至少一个处理器801执行时,执行根据本公开实施例的视频直播方法。
处理器801可以包括中央处理器(CPU)、图形处理器(GPU)、可编程逻辑装置、专用处理器系统、微控制器或微处理器。作为示例而非限制,处理器801还可包括模拟处理器、数字处理器、微处理器、多核处理器、处理器阵列、网络处理器等。
作为一种存储介质的存储器802可以包括操作系统(例如,MAC操作系统)、数据存储模块、网络通信模块、用户接口模块、视频直播程序以及数据库。
存储器802可以与处理器801集成为一体,例如,可以将RAM或闪存布置在集成电路微处理器等之内。此外,存储器802可包括独立的装置,诸如,外部盘驱动、存储阵列或任何数据库系统可以使用的其他存储装置。存储器802和处理器801可以在操作上进行耦合,或者可以例如通过I/O端口、网络连接等互相通信,使得处理器801能够读取存储在存储器802中的文件。
此外,电子设备800还可以包括视频显示器(诸如,液晶显示器)和用户交互接口(诸如,键盘、鼠标、触摸输入装置等)。电子设备800的所有组件可以经由总线和/或网络而彼此连接。
本领域技术人员可以理解,图8中示出的结构并不构成对的限定,可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件布置。
根据本公开的实施例,还可以提供一种非易失性计算机可读存储介质,其上存储有指令,其中,当指令被至少一个处理器运行时,促使至少一个处理器执行根据本公开的视频 直播方法。这里的非易失性计算机可读存储介质的示例包括:只读存储器(ROM)、随机存取可编程只读存储器(PROM)、电可擦除可编程只读存储器(EEPROM)、随机存取存储器(RAM)、动态随机存取存储器(DRAM)、静态随机存取存储器(SRAM)、闪存、非易失性存储器、CD-ROM、CD-R、CD+R、CD-RW、CD+RW、DVD-ROM、DVD-R、DVD+R、DVD-RW、DVD+RW、DVD-RAM、BD-ROM、BD-R、BD-R LTH、BD-RE、蓝光或光盘存储器、硬盘驱动器(HDD)、固态硬盘(SSD)、卡式存储器(诸如,多媒体卡、安全数字(SD)卡或极速数字(XD)卡)、磁带、软盘、磁光数据存储装置、光学数据存储装置、硬盘、固态盘以及任何其他装置,所述任何其他装置被配置为以非暂时性方式存储计算机程序以及任何相关联的数据、数据文件和数据结构并将所述计算机程序以及任何相关联的数据、数据文件和数据结构提供给处理器或计算机使得处理器或计算机能执行所述计算机程序。上述非易失性计算机可读存储介质中的计算机程序可以在诸如客户端、主机、代理装置、服务器等计算机设备中部署的环境中运行,此外,在一个示例中,计算机程序以及任何相关联的数据、数据文件和数据结构分布在联网的计算机系统上,使得计算机程序以及任何相关联的数据、数据文件和数据结构通过一个或多个处理器或计算机以分布式方式存储、访问和执行。
根据本公开的实施例中,还可以提供一种计算机程序产品,该计算机程序产品中的指令可以由计算机设备的处理器执行以完成上述视频直播方法。
本公开所有实施例均可以单独被执行,也可以于其他实施例相结合被执行,均视为本公开要求的保护范围。
本领域普通技术人员可以理解实现上述实施例方法中的全部或部分流程,是可以通过计算机程序来指令相关的硬件来完成,该计算机程序可存储于非易失性计算机可读取存储介质中,该计算机程序在执行时,可以包括如上述各方法的实施例的流程。其中,本申请所提供的各实施例中所使用的对存储器、存储、数据库或其它介质的任何引用,均可以包括非易失性和/或易失性存储器。

Claims (20)

  1. 一种视频直播方法,包括:
    获取由第一设备采集的第一直播数据,其中,第一直播数据包括由第一设备采集的第一直播视频中的第一感兴趣区域和第一背景中的至少一个;
    获取由第二设备采集的第二直播数据,其中,第二直播数据包括由第二设备采集的第二直播视频中的第二感兴趣区域和第二背景中的至少一个;
    基于第一直播数据和第二直播数据来生成目标直播视频;
    发送所述目标直播视频。
  2. 如权利要求1所述的视频直播方法,其中,所述第一感兴趣区域是通过对所述第一直播视频中的每一帧进行目标区域提取而获得的,且所述第二感兴趣区域是通过对所述第二直播视频中的每一帧进行目标区域提取而获得的。
  3. 如权利要求1所述的视频直播方法,其中,所述第一感兴趣区域和/或所述第二感兴趣区域为人像区域。
  4. 如权利要求1所述的视频直播方法,其中,所述基于第一直播数据和第二直播数据来生成目标直播视频的步骤包括:
    将所述第一背景与所述第二背景进行融合来生成融合背景作为所述目标直播视频中的背景;
    在所述融合背景中显示所述第一感兴趣区域和所述第二感兴趣区域。
  5. 如权利要求1所述的视频直播方法,其中,所述基于第一直播数据和第二直播数据来生成目标直播视频的步骤包括:
    选择所述第一背景或所述第二背景作为所述目标直播视频中的背景,并且在选择的背景中显示所述第一感兴趣区域和所述第二感兴趣区域。
  6. 如权利要求1-5中任一项所述的视频直播方法,其中,所述第一感兴趣区域和/或所述第二感兴趣区域基于用户输入在所述目标直播视频中改变位置或尺寸。
  7. 一种视频直播装置,包括:
    获取模块,被配置为:获取由第一设备采集的第一直播数据和由第二设备采集的第二直播数据,其中,第一直播数据包括由第一设备采集的第一直播视频中的第一感兴趣区域和第一背景中的至少一个,并且第二直播数据包括由第二设备采集的第二直播视频中的第二感兴趣区域和第二背景中的至少一个;
    处理模块,被配置为基于第一直播数据和第二直播数据来生成目标直播视频;以及
    发送模块,被配置为发送所述目标直播视频。
  8. 如权利要求7所述的视频直播装置,其中,所述第一感兴趣区域是通过对所述第一直播视频中的每一帧进行目标区域提取而获得的,且所述第二感兴趣区域是通过对所述第二直播视频中的每一帧进行目标区域提取而获得的。
  9. 如权利要求7所述的视频直播装置,其中,所述第一感兴趣区域和/或所述第二感 兴趣区域为人像区域。
  10. 如权利要求7所述的视频直播装置,其中,处理模块被配置为:
    将所述第一背景与所述第二背景进行融合来生成融合背景作为所述目标直播视频中的背景;
    在所述融合背景中显示所述第一感兴趣区域和所述第二感兴趣区域。
  11. 如权利要求7所述的视频直播装置,其中,处理模块被配置为:
    选择所述第一背景和/或所述第二背景作为所述目标直播视频中的背景,并且在选择的背景中显示所述第一感兴趣区域和所述第二感兴趣区域。
  12. 如权利要求7-11中任一项所述的视频直播装置,其中,所述第一感兴趣区域和/或所述第二感兴趣区域基于用户输入在所述目标直播视频中改变位置或尺寸。
  13. 一种电子设备,包括:
    至少一个处理器;
    至少一个存储计算机可执行指令的存储器,
    其中,所述至少一个处理器被配置为:
    获取由第一设备采集的第一直播数据,其中,第一直播数据包括由第一设备采集的第一直播视频中的第一感兴趣区域和第一背景中的至少一个;
    获取由第二设备采集的第二直播数据,其中,第二直播数据包括由第二设备采集的第二直播视频中的第二感兴趣区域和第二背景中的至少一个;
    基于第一直播数据和第二直播数据来生成目标直播视频;
    发送所述目标直播视频。
  14. 根据权利要求13所述的电子设备,其中,所述第一感兴趣区域是通过对所述第一直播视频中的每一帧进行目标区域提取而获得的,且所述第二感兴趣区域是通过对所述第二直播视频中的每一帧进行目标区域提取而获得的。
  15. 根据权利要求13所述的电子设备,其中,所述第一感兴趣区域和/或所述第二感兴趣区域为人像区域。
  16. 根据权利要求13所述的电子设备,其中,所述至少一个处理器被配置为:
    将所述第一背景与所述第二背景进行融合来生成融合背景作为所述目标直播视频中的背景;
    在所述融合背景中显示所述第一感兴趣区域和所述第二感兴趣区域。
  17. 根据权利要求13所述的电子设备,其中,所述至少一个处理器被配置为:
    选择所述第一背景或所述第二背景作为所述目标直播视频中的背景,并且在选择的背景中显示所述第一感兴趣区域和所述第二感兴趣区域。
  18. 根据权利要求13-17中任一项所述的电子设备,其中,所述第一感兴趣区域和/或所述第二感兴趣区域基于用户输入在所述目标直播视频中改变位置或尺寸。
  19. 一种非易失性计算机可读存储介质,当所述指令被至少一个处理器运行时,促使 所述至少一个处理器执行视频直播方法,
    其中,所述视频直播方法包括:
    获取由第一设备采集的第一直播数据,其中,第一直播数据包括由第一设备采集的第一直播视频中的第一感兴趣区域和第一背景中的至少一个;
    获取由第二设备采集的第二直播数据,其中,第二直播数据包括由第二设备采集的第二直播视频中的第二感兴趣区域和第二背景中的至少一个;
    基于第一直播数据和第二直播数据来生成目标直播视频;
    发送所述目标直播视频。
  20. 一种计算机程序产品,所述计算机程序产品中的指令被电子装置中的至少一个处理器运行以执行视频直播方法,
    其中,所述视频直播方法包括:
    获取由第一设备采集的第一直播数据,其中,第一直播数据包括由第一设备采集的第一直播视频中的第一感兴趣区域和第一背景中的至少一个;
    获取由第二设备采集的第二直播数据,其中,第二直播数据包括由第二设备采集的第二直播视频中的第二感兴趣区域和第二背景中的至少一个;
    基于第一直播数据和第二直播数据来生成目标直播视频;
    发送所述目标直播视频。
PCT/CN2022/070254 2021-05-27 2022-01-05 视频直播方法和视频直播装置 Ceased WO2022247293A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202110584596.4A CN113315987A (zh) 2021-05-27 2021-05-27 视频直播方法和视频直播装置
CN202110584596.4 2021-05-27

Publications (1)

Publication Number Publication Date
WO2022247293A1 true WO2022247293A1 (zh) 2022-12-01

Family

ID=77375551

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2022/070254 Ceased WO2022247293A1 (zh) 2021-05-27 2022-01-05 视频直播方法和视频直播装置

Country Status (2)

Country Link
CN (1) CN113315987A (zh)
WO (1) WO2022247293A1 (zh)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN116112728A (zh) * 2023-01-09 2023-05-12 北京达佳互联信息技术有限公司 信息展示方法、装置、设备及存储介质
CN116389827A (zh) * 2023-04-14 2023-07-04 广州播丫科技有限公司 一种将直播画面推送到异地融合的方法、系统、终端以及存储介质

Families Citing this family (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113315987A (zh) * 2021-05-27 2021-08-27 北京达佳互联信息技术有限公司 视频直播方法和视频直播装置
CN113891044B (zh) * 2021-09-29 2023-03-24 天翼物联科技有限公司 视频直播方法、装置、计算机设备及计算机可读存储介质
CN113902989B (zh) * 2021-09-30 2025-05-23 腾讯音乐娱乐科技(深圳)有限公司 直播场景检测方法、存储介质及电子设备
CN115695711B (zh) * 2022-11-01 2025-10-28 联想(北京)有限公司 一种处理方法及采集设备

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN107105315A (zh) * 2017-05-11 2017-08-29 广州华多网络科技有限公司 直播方法、主播客户端的直播方法、主播客户端及设备
CN108965746A (zh) * 2018-07-26 2018-12-07 北京竞业达数码科技股份有限公司 视频合成方法及系统
CN110719416A (zh) * 2019-09-30 2020-01-21 咪咕视讯科技有限公司 一种直播方法、通信设备及计算机可读存储介质
CN112752116A (zh) * 2020-12-30 2021-05-04 广州繁星互娱信息科技有限公司 直播视频画面的显示方法、装置、终端及存储介质
CN113315987A (zh) * 2021-05-27 2021-08-27 北京达佳互联信息技术有限公司 视频直播方法和视频直播装置

Family Cites Families (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN107124662B (zh) * 2017-05-10 2022-03-18 腾讯科技(上海)有限公司 视频直播方法、装置、电子设备及计算机可读存储介质
CN107682729A (zh) * 2017-09-08 2018-02-09 广州华多网络科技有限公司 一种基于直播的互动方法及直播系统、电子设备
CN111083507B (zh) * 2019-12-09 2021-11-23 广州酷狗计算机科技有限公司 连麦方法及系统、第一主播端、观众端及计算机存储介质
CN111432235A (zh) * 2020-04-01 2020-07-17 网易(杭州)网络有限公司 直播视频生成方法、装置、计算机可读介质及电子设备
CN112291579A (zh) * 2020-10-26 2021-01-29 北京字节跳动网络技术有限公司 数据处理方法、装置、设备和存储介质

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN107105315A (zh) * 2017-05-11 2017-08-29 广州华多网络科技有限公司 直播方法、主播客户端的直播方法、主播客户端及设备
CN108965746A (zh) * 2018-07-26 2018-12-07 北京竞业达数码科技股份有限公司 视频合成方法及系统
CN110719416A (zh) * 2019-09-30 2020-01-21 咪咕视讯科技有限公司 一种直播方法、通信设备及计算机可读存储介质
CN112752116A (zh) * 2020-12-30 2021-05-04 广州繁星互娱信息科技有限公司 直播视频画面的显示方法、装置、终端及存储介质
CN113315987A (zh) * 2021-05-27 2021-08-27 北京达佳互联信息技术有限公司 视频直播方法和视频直播装置

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN116112728A (zh) * 2023-01-09 2023-05-12 北京达佳互联信息技术有限公司 信息展示方法、装置、设备及存储介质
CN116389827A (zh) * 2023-04-14 2023-07-04 广州播丫科技有限公司 一种将直播画面推送到异地融合的方法、系统、终端以及存储介质

Also Published As

Publication number Publication date
CN113315987A (zh) 2021-08-27

Similar Documents

Publication Publication Date Title
WO2022247293A1 (zh) 视频直播方法和视频直播装置
CN109547819B (zh) 直播列表展示方法、装置以及电子设备
CN106658200B (zh) 直播视频分享和获取的方法、装置及其终端设备
CN108401192B (zh) 视频流处理方法、装置、计算机设备及存储介质
CN101778257B (zh) 用于数字视频点播中的视频摘要片断的生成方法
US20190253474A1 (en) Media production system with location-based feature
CN109618224B (zh) 视频数据处理方法、装置、计算机可读存储介质和设备
WO2021114708A1 (zh) 多人视频直播业务实现方法、装置、计算机设备
WO2018045927A1 (zh) 一种基于三维虚拟技术的网络实时互动直播方法及装置
US11581018B2 (en) Systems and methods for mixing different videos
CN107436921B (zh) 视频数据处理方法、装置、设备及存储介质
CN114845149B (zh) 视频片段的剪辑方法、视频推荐方法、装置、设备及介质
CN108449631B (zh) 用于媒体处理的方法、装置及可读介质
US10897658B1 (en) Techniques for annotating media content
KR20230026321A (ko) 복합 비디오 촬상, 전자 장치 및 컴퓨터 판독가능 매체를 위한 방법 및 장치
US9872056B1 (en) Methods, systems, and media for detecting abusive stereoscopic videos by generating fingerprints for multiple portions of a video frame
CN114139491A (zh) 一种数据处理方法、装置及存储介质
CN104618741A (zh) 一种基于视频内容的信息推送系统及方法
WO2017157135A1 (zh) 媒体信息处理方法及媒体信息处理装置、存储介质
CN105814905B (zh) 用于使使用信息在装置与服务器之间同步的方法和系统
CN106060573A (zh) 基于终端屏幕内容的直播方法及装置
CN108268139A (zh) 虚拟场景交互方法及装置、计算机装置及可读存储介质
RU2764375C1 (ru) Способ формирования изображений с дополненной и виртуальной реальностью с возможностью взаимодействия внутри виртуального мира, содержащего данные виртуального мира
KR102400733B1 (ko) 이미지에 내재된 코드를 이용한 컨텐츠 확장 장치
CN116980637A (zh) 直播数据的处理系统、电子设备、存储介质及程序产品

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 22810032

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 28.03.2024)

122 Ep: pct application non-entry in european phase

Ref document number: 22810032

Country of ref document: EP

Kind code of ref document: A1