WO2025002428A1 - 直播处理方法、设备及存储介质 - Google Patents

直播处理方法、设备及存储介质 Download PDF

Info

Publication number
WO2025002428A1
WO2025002428A1 PCT/CN2024/102664 CN2024102664W WO2025002428A1 WO 2025002428 A1 WO2025002428 A1 WO 2025002428A1 CN 2024102664 W CN2024102664 W CN 2024102664W WO 2025002428 A1 WO2025002428 A1 WO 2025002428A1
Authority
WO
WIPO (PCT)
Prior art keywords
video frame
latest
target
live broadcast
current video
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2024/102664
Other languages
English (en)
French (fr)
Other versions
WO2025002428A9 (zh
Inventor
邓卓尧
曲喆麒
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Zitiao Network Technology Co Ltd
Original Assignee
Beijing Zitiao Network Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Zitiao Network Technology Co Ltd filed Critical Beijing Zitiao Network Technology Co Ltd
Publication of WO2025002428A1 publication Critical patent/WO2025002428A1/zh
Anticipated expiration legal-status Critical
Publication of WO2025002428A9 publication Critical patent/WO2025002428A9/zh
Ceased legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/44Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
    • H04N21/44016Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving splicing one content stream with another content stream, e.g. for substituting a video clip
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/431Generation of visual interfaces for content selection or interaction; Content or additional data rendering
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/431Generation of visual interfaces for content selection or interaction; Content or additional data rendering
    • H04N21/4312Generation of visual interfaces for content selection or interaction; Content or additional data rendering involving specific graphical features, e.g. screen layout, special fonts or colors, blinking icons, highlights or animations
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/433Content storage operation, e.g. storage operation in response to a pause request, caching operations
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/433Content storage operation, e.g. storage operation in response to a pause request, caching operations
    • H04N21/4331Caching operations, e.g. of an advertisement for later insertion during playback
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/44Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/44Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
    • H04N21/44004Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving video buffer management, e.g. video decoder buffer or video display buffer
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/47End-user applications
    • H04N21/478Supplemental services, e.g. displaying phone caller identification, shopping application
    • H04N21/4788Supplemental services, e.g. displaying phone caller identification, shopping application communicating with other users, e.g. chatting

Definitions

  • the embodiments of the present disclosure relate to the field of computer and network communication technology, and in particular, to a live broadcast processing method, device, and storage medium.
  • multi-person connection is a method often used by anchors to interact with other anchors. Based on this method, the anchor in the current live broadcast room can invite other anchors to broadcast live together.
  • the embodiments of the present disclosure provide a live broadcast processing method, device and storage medium to reduce the performance pressure on the live broadcast client in a multi-person live broadcast scenario and improve the live broadcast fluency.
  • an embodiment of the present disclosure provides a live broadcast processing method, which is applied to a first live broadcast client, and the method includes: receiving a live broadcast stream of a second live broadcast client, and caching the latest video frame of the live broadcast stream of the second live broadcast client; when obtaining the current video frame of the first live broadcast client, reading the cached latest video frame; based on a single rendering view, synthesizing the current video frame and the latest video frame into a target video frame according to preset layout information; and displaying the target video frame.
  • an embodiment of the present disclosure provides a live broadcast processing device, which is applied to a first live broadcast and microphone connection client, and the device includes: a receiving unit, which is used to receive a live broadcast stream of a second live broadcast and microphone connection client; A cache unit is used to cache the latest video frame of the live stream of the second live broadcast and microphone connection client; when obtaining the current video frame of the first live broadcast and microphone connection client, the cached latest video frame is read; a rendering unit is used to synthesize the current video frame and the latest video frame into a target video frame based on a single rendering view and according to preset layout information; a display unit is used to display the target video frame.
  • an embodiment of the present disclosure provides an electronic device, comprising: at least one processor and a memory; the memory stores computer execution instructions; the at least one processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the live broadcast processing method described in the first aspect and various possible designs of the first aspect.
  • an embodiment of the present disclosure provides a computer-readable storage medium, in which computer execution instructions are stored.
  • a processor executes the computer execution instructions, the live broadcast processing method described in the first aspect and various possible designs of the first aspect is implemented.
  • an embodiment of the present disclosure provides a computer program product, including computer execution instructions.
  • a processor executes the computer execution instructions, the live broadcast processing method described in the first aspect and various possible designs of the first aspect is implemented.
  • FIG1 is a scene example diagram of a live broadcast processing method provided by an embodiment of the present disclosure.
  • FIG2 is a schematic diagram of a live broadcast processing method according to an embodiment of the present disclosure.
  • FIG3 is a schematic flow chart of a live broadcast processing method provided by another embodiment of the present disclosure.
  • FIG4a is a schematic diagram of an interface of a layout method provided by an embodiment of the present disclosure.
  • FIG4b is a schematic diagram of an interface of a layout method provided by another embodiment of the present disclosure.
  • FIG5 is a schematic diagram of aligning a layer corresponding to the latest video frame of each live stream with a preset list view control provided by an embodiment of the present disclosure
  • FIG6 is a schematic flow chart of a live broadcast processing method provided by another embodiment of the present disclosure.
  • FIG7 is a schematic diagram of a target template provided by an embodiment of the present disclosure.
  • FIG8 is a schematic flow chart of a live broadcast processing method provided by another embodiment of the present disclosure.
  • FIG9 is a schematic diagram of determining a target template according to an embodiment of the present disclosure.
  • FIG10 is a structural block diagram of a live broadcast processing device provided by an embodiment of the present disclosure.
  • FIG. 11 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure.
  • multi-person connection is a method often used by anchors to interact with other anchors. Based on this method, the anchor in the current live broadcast room can invite other anchors to broadcast live together.
  • a multi-person live broadcast scenario for any first live broadcast client, different rendering views, different GL rendering threads, different frame buffer objects, and different surface views will be used for the current video frame of the first live broadcast client and the video frame of the live stream of each other second live broadcast client.
  • the rendering view, GL rendering thread frame buffer object, and surface view cannot be reused, thus occupying a large amount of resources of the live broadcast client.
  • the live broadcast fluency is seriously affected, the performance of the live broadcast client is seriously affected, and it may be seriously heated.
  • the present disclosure provides a live broadcast processing method, for any first live broadcast client in a multi-person live broadcast room, the live broadcast client receives the live broadcast stream of the second live broadcast client, and caches the latest video frame of each live broadcast stream; when obtaining the current video frame of the first live broadcast client, read the cached latest video frame; based on a single rendering view, synthesize the current video frame and the latest video frame into a target video frame according to preset layout information; and display the target video frame.
  • rendering can be achieved using only a single rendering view, which reduces resource usage, reduces performance pressure on the live broadcast client, and improves live broadcast fluency.
  • the execution subject is the first live broadcast client.
  • the first live broadcast client can receive the live streams of one or more second live broadcast clients, and cache the latest video frame of each live stream to the corresponding layer Layer.
  • the first live broadcast client can collect the current video frame as the initial layer OriginLayer, and then when the first live broadcast client collects the current video frame, read the latest video frame of each cached live stream, and synthesize the current video frame and the latest video frame of each live stream into a target video frame for display.
  • the user information and data involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
  • the method of this embodiment can be applied to any live broadcast and microphone client (referred to as the first live broadcast and microphone client) in a multi-person live broadcast room.
  • the live broadcast processing method includes:
  • S201 Receive a live stream from a second live broadcast and microphone connection client, and cache the latest video frame of the live stream from the second live broadcast and microphone connection client.
  • the first live broadcast client can invite one or more second live broadcast clients to connect to the microphone, so that the screen of the first live broadcast client and the screen of one or more second live broadcast clients can be viewed simultaneously in the interfaces of the first live broadcast client and the second live broadcast client.
  • the first live broadcast client can receive the live stream of one or more second live broadcast clients.
  • one or more second live broadcast clients can upload their own live streams to the server, and the server can send the live streams of one or more second live broadcast clients to the first live broadcast client.
  • the live stream of each second live broadcast client is parsed to obtain the video frames in the live stream, and the latest video frame is cached. Specifically, the latest video frame of each live stream can be drawn and cached in a corresponding layer.
  • the first live broadcast and microphone client can collect the first live broadcast and microphone client in real time.
  • the current video frame is obtained from the screen on the side, and the synthesis of the current video frame and the latest video frame of the live stream of each second live broadcast and microphone client is driven based on the current video frame. Since the current video frame collected by the first live broadcast and microphone client does not need to be transmitted through the network, the output is relatively stable. Therefore, the current video frame collected by the first live broadcast and microphone client can be driven to ensure the stability of the live picture.
  • the latest video frame of each cached live stream can be read. Specifically, the latest video frame of each live stream can be read from each layer Layer for synthesis of the target video frame.
  • S203 Based on a single rendering view, synthesize the current video frame and the latest video frame of each live stream into a target video frame according to preset layout information.
  • the current video frame collected by the first live broadcast and the latest video frame of the live stream of each second live broadcast and the microphone client can be synthesized and drawn on the same texture to obtain a target video frame, and a rendering view (RenderView) is shared to achieve rendering of the current video frame and the latest video frame of each live stream in one view, without using different rendering views for the current video frame of the first live broadcast and the video frame of the live stream of each second live broadcast and the microphone client, thereby reducing the performance pressure on the first live broadcast and the microphone client.
  • how the current video frame and the latest video frame of each live stream are laid out can be determined according to the preset layout information.
  • the target video frame can be rendered into a single frame buffer object (Frame Buffer Object, FBO), avoiding the use of different FBOs to cache the current video frame of the first live broadcast client and the video frame of the live stream of each second live broadcast client, further reducing the performance pressure on the first live broadcast client.
  • FBO Frame Buffer Object
  • the target video frame is synthesized and rendered, it is displayed in the display interface of the first live broadcast and microphone connection client.
  • the rendered target video frames cached in the FBO are displayed using a single surface view.
  • the live broadcast processing method receives the live broadcast stream of the second live broadcast client through the first live broadcast client, and caches the latest video frame of the live broadcast stream of the second live broadcast client; when obtaining the current video frame of the first live broadcast client, read the cached latest video frame; based on a single rendering view, synthesize the current video frame and the latest video frame into a target video frame according to preset layout information; and display the target video frame.
  • the step of synthesizing the current video frame and the latest video frame into a target video frame according to the preset layout information in S203 includes:
  • S301 Determine the size and position of a first layer corresponding to the current video frame and a second layer corresponding to the latest video frame in the target video frame according to preset layout information.
  • S302 Draw the current video frame into the first layer, draw the latest video frames into the second layers respectively, and synthesize the first layer and the second layer into the target video frame.
  • the current video frame of the first live broadcast and microphone client and the latest video frame of the live stream of each second live broadcast and microphone client can be determined according to the preset layout information.
  • the first layer (Layer) of the current video frame of the first live broadcast and microphone client is located on the left side of the target video frame, showing the picture of the current video frame of the first live broadcast and microphone client
  • the second layer (Layer) of the latest video frame of the live stream of each second live broadcast and microphone client is arranged in a row vertically and located on the right side of the target video frame, showing the picture of each second live broadcast and microphone client.
  • the first layer (Layer) of the current video frame of the first live broadcast and microphone client is displayed in full screen in the target video frame, showing the picture of the current video frame of the first live broadcast and microphone client, and the second layer (Layer) of the latest video frame of the live stream of each second live broadcast and microphone client is located on the current video frame of the first live broadcast and microphone client in the form of a floating window, and the picture of each second live broadcast and microphone client is suspended above the picture of the current video frame; of course, other layout methods can also be used.
  • the size and position of the first layer corresponding to the current video frame and the second layer corresponding to the latest video frame of each live stream can be determined according to the preset layout information in the target video frame, and then the current video frame is drawn to the first layer corresponding to the current video frame, and the latest video frame of each live stream is drawn to the second layer corresponding to the latest video frame of each live stream, so that each layer is synthesized into the target video frame.
  • the synthesizing the current video frame and the latest video frame into a target video frame based on a single rendering view according to preset layout information in S203 it also includes: obtaining the size and position of each view control in the preset list view control, and using the size and position of any view control as the size and position of the second layer corresponding to the latest video frame of the live stream of any second live broadcast client in the target video frame, and obtaining the preset layout information so that the second layer corresponding to the latest video frame of the live stream is aligned with the view control in the preset list view control.
  • a preset list view control can be set in the display interface of the multi-person live broadcast scene.
  • the preset list view control includes multiple view controls (ViewHolder).
  • ViewHolder the preset list view control RecyclerView shown in Figure 5 is a nine-grid list view control, which includes nine view controls ViewHolders.
  • the layers corresponding to the latest video frames of each second live broadcast client live stream are aligned with each view control ViewHolder, respectively, so that the corresponding live stream can be interacted with through each view control ViewHolder.
  • the size and position of each view control in the preset list view control can be obtained, and the size and position of each view control are respectively used as the size and position of the layer corresponding to the latest video frame of each live stream in the target video frame, thereby serving as the preset layout information.
  • the current video frame and the latest video frame of each live stream of the second live broadcast client are synthesized into a target video frame according to the preset layout information, so that the second layer corresponding to the latest video frame of each live stream in the target video frame is aligned with each view control in the preset list view control respectively.
  • the current video frame of the first live broadcast and microphone client and the latest video frame of the live broadcast stream of the second live broadcast and microphone client may have a certain overlapping area, so it is necessary to set the display level value, with the highest level placed at the top and the lowest level placed at the bottom. If there is an overlapping area, the video frame with a higher level covers the video frame with a lower level.
  • synthesizing the current video frame and the latest video frame into a target video frame according to the preset layout information in S203 includes:
  • a depth value at an area corresponding to the current video frame in the target template is a level value corresponding to the current video frame
  • a depth value at an area corresponding to the latest video frame is a level value corresponding to the latest video frame
  • the rendering pipeline of the OpenGL fragment shader includes a stencil test link.
  • the stencil test determines whether a fragment is retained or discarded based on the template. If a fragment fails the stencil test, the fragment is discarded, that is, the fragment is not drawn. If a fragment passes the stencil test, the fragment is retained.
  • a target template for stencil testing in the fragment shader rendering pipeline can be pre-configured.
  • the depth value of the corresponding area of the current video frame and the latest video frame of each second live broadcast client live stream in the target template is the level value zOrder corresponding to each area. As shown in Figure 7, the higher the level value zOrder, the larger the number, that is, the higher the depth value.
  • the latest video frame of the live stream of the second live broadcast client is suspended above the current video frame of the first live broadcast client, that is, the level value zOrder of the latest video frame of the live stream of the second live broadcast client is higher than the level value zOrder of the current video frame of the first live broadcast client.
  • the level value zOrder of the latest video frame of the live stream of the second live broadcast client can be set to 3, and the level value zOrder of the current video frame of the first live broadcast client is set to 2.
  • the depth value of each position in the corresponding area (inside the box) of the latest video frame of the live stream is 3, and the depth value of each position in the area not covered by the current video frame of the first live broadcast client is 2.
  • a template test is performed according to the level value of the current video frame, the level value of the latest video frame and the target template, and the level value of the current video frame and the position that passes the template test in the latest video frame are rendered and drawn, and the position that fails the template test is not rendered and drawn, and finally the to the synthesized target video frame.
  • the specific process of the template test may include:
  • the hierarchical relationship during synthesis is achieved. After the positions that pass the template test in the current video frame and the latest video frame of each second live broadcast client live stream are rendered, the synthesized target video frame is obtained.
  • the step of determining the target template for template testing in the fragment shader rendering pipeline in S401 may include:
  • S502 According to the level value of the latest video frame and the preset layout information, update the depth value of the area corresponding to the latest video frame in the target template to the level value of the latest video frame.
  • the preset level value zOrder of the current video frame is obtained.
  • the current The hierarchy value zOrder of the video frame is 2, and the depth values at all positions of the target template are set to the hierarchy value of the current video frame, that is, the depth values at all positions of the target template are 2, as shown in the upper side of Figure 9; obtain the hierarchy value zOrder of the latest video frame of any second live broadcast client live stream that is preset, assuming that the hierarchy value zOrder of the latest video frame of any second live broadcast client live stream is 3, and the position and size of the latest video frame of the live stream in the target video frame can be obtained from the preset layout information.
  • the corresponding area of the latest video frame of any second live broadcast client live stream in the target template can be determined, that is, the box area in the figure. Further, the depth value at the corresponding area of the latest video frame of any second live broadcast client live stream in the target template is updated to the hierarchy value of the latest video frame of the second live broadcast client live stream, that is, the depth value of each position in the box area of the target template is updated to 3, as shown in the lower side of Figure 9, and finally the target template is obtained.
  • the target template is configured with a buffer (cache), and S501-S502 are implemented in the buffer of the target template, that is, the depth values at all positions of the target template in the buffer of the target template are set to the level values of the current video frame; in the buffer of the target template, the depth value at the corresponding area of the latest video frame of the live stream of any second live broadcast client in the target template is updated to the level value of the latest video frame of the live stream of any second live broadcast client in the buffer of the target template.
  • a buffer cache
  • S501-S502 are implemented in the buffer of the target template, that is, the depth values at all positions of the target template in the buffer of the target template are set to the level values of the current video frame; in the buffer of the target template, the depth value at the corresponding area of the latest video frame of the live stream of any second live broadcast client in the target template is updated to the level value of the latest video frame of the live stream of any second live broadcast client in the buffer of the target template.
  • the live broadcast processing method, device and storage medium provided by the embodiments of the present disclosure receive the live broadcast stream of the second live broadcast client through the first live broadcast client, and cache the latest video frame of the live broadcast stream of the second live broadcast client; when obtaining the current video frame of the first live broadcast client, read the cached latest video frame; based on a single rendering view, synthesize the current video frame and the latest video frame into a target video frame according to preset layout information; and display the target video frame.
  • rendering can be achieved using only a single rendering view, which reduces resource usage, reduces the performance pressure on the live broadcast client, and improves the fluency of the live broadcast.
  • FIG10 is a structural block diagram of the live broadcast processing device provided by the embodiment of the present disclosure.
  • the live broadcast processing device 600 includes: a receiving unit 601, a cache unit 602, a rendering unit 603, and a display unit 604.
  • the receiving unit 601 is used to receive the live stream of the second live broadcast and microphone client; the buffering unit 602 is used to buffer the latest video frame of the live stream of the second live broadcast and microphone client. Storage; when obtaining the current video frame of the first live broadcast and microphone connection client, read the cached latest video frame; a rendering unit 603 is used to synthesize the current video frame and the latest video frame into a target video frame based on a single rendering view and according to preset layout information; a display unit 604 is used to display the target video frame.
  • the rendering unit 603 when the rendering unit 603 synthesizes the current video frame and the latest video frame into a target video frame according to preset layout information, it is used to: determine the size and position of the first layer corresponding to the current video frame and the second layer corresponding to the latest video frame in the target video frame according to the preset layout information; draw the current video frame into the first layer, draw the latest video frame into the second layer respectively, and synthesize the first layer and the second layer into the target video frame.
  • the rendering unit 603 when the rendering unit 603 synthesizes the current video frame and the latest video frame into a target video frame according to preset layout information, it is used to: determine a target template for template testing in a fragment shader rendering pipeline, wherein a depth value at an area corresponding to the current video frame in the target template is a level value corresponding to the current video frame, and a depth value at an area corresponding to the latest video frame is a level value corresponding to the latest video frame; perform a template test according to the level value of the current video frame, the level value of the latest video frame and the target template, and synthesize the current video frame and the latest video frame based on the template test result.
  • the rendering unit 603 when performing a template test based on the level value of the current video frame, the level value of the latest video frame of each live stream, and the target template, the rendering unit 603 is used to: determine the level value at any first position of the current video frame as the actual depth value at the first position, compare the actual depth value at the first position of the current video frame with the depth value at the corresponding position in the target template, if the actual depth value at the first position is higher than or equal to the depth value at the corresponding position in the target template, the template test of the first position passes; determine the level value at any second position of the latest video frame as the actual depth value at the second position, compare the actual depth value at the second position of the latest video frame with the depth value at the corresponding position in the target template, if the actual depth value at the second position is higher than or equal to the depth value at the corresponding position in the target template, the template test of the second position passes.
  • the rendering unit 603 when the rendering unit 603 synthesizes the current video frame and the latest video frame based on the template test result, it is used to: If the template test of any first position of the frame passes, the first position of the current video frame is rendered; if the template test of any second position of the latest video frame passes, the second position of the latest video frame is rendered; after the positions that pass the template test in the current video frame and the latest video frame are rendered, the synthesized target video frame is obtained.
  • the level value of the latest video frame is higher than the level value of the current video frame.
  • the rendering unit 603 when determining the target template for template testing in the fragment shader rendering pipeline, is used to: set the depth values at all positions of the target template to the level value of the current video frame; and update the depth value at the area corresponding to the latest video frame in the target template to the level value of the latest video frame according to the level value of the latest video frame and the preset layout information.
  • the rendering unit 603 before the rendering unit 603 synthesizes the current video frame and the latest video frame of each live stream into a target video frame based on a single rendering view and according to preset layout information, it is also used to: obtain the size and position of each view control in the preset list view control, use the size and position of any view control as the size and position of the second layer corresponding to the latest video frame in the target video frame, and obtain the preset layout information so that the second layer corresponding to the latest video frame is aligned with any view control in the preset list view control.
  • the rendering unit 603 when the rendering unit 603 synthesizes the current video frame and the latest video frame into a target video frame based on a single rendering view and according to preset layout information, it is used to: synthesize the current video frame and the latest video frame into a target video frame based on a single rendering view and a single rendering thread according to preset layout information.
  • the display unit 604 when displaying the target video frame, is configured to: display the target video frame using a single surface view.
  • the rendering unit 603 when the rendering unit 603 synthesizes the current video frame and the latest video frame into a target video frame based on a single rendering view and according to preset layout information, it is also used to: render the target video frame into a single frame buffer object; when the display unit 604 displays the target video frame using a single surface view, it is used to: display the rendered target video frame cached in the frame buffer object using a single surface view.
  • the live broadcast processing device provided in this embodiment can be used to execute the technical solution of the above-mentioned method embodiment. Its implementation principle and technical effects are similar, and this embodiment will not be repeated here.
  • FIG. 11 it shows a schematic diagram of the structure of an electronic device 700 suitable for implementing the embodiment of the present disclosure
  • the electronic device 700 may be a terminal device or a server.
  • the terminal device may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.
  • PDAs personal digital assistants
  • PADs Portable Android Devices
  • PMPs portable multimedia players
  • vehicle terminals such as vehicle navigation terminals
  • fixed terminals such as digital TVs, desktop computers, etc.
  • the electronic device shown in FIG. 11 is only an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.
  • the electronic device 700 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 to a random access memory (RAM) 703.
  • a processing device 701 e.g., a central processing unit, a graphics processing unit, etc.
  • RAM random access memory
  • Various programs and data required for the operation of the electronic device 700 are also stored in the RAM 703.
  • the processing device 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704.
  • An input/output (I/O) interface 705 is also connected to the bus 704.
  • the following devices may be connected to the I/O interface 705: input devices 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 708 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 709.
  • the communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or by wire to exchange data.
  • FIG. 11 shows an electronic device 700 having various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
  • an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart.
  • the computer program can be downloaded and installed from the network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702.
  • the processing device 701 the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
  • the computer-readable medium of the present disclosure may be a computer-readable signal medium.
  • the computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
  • a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device.
  • a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above.
  • the computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device.
  • the program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
  • the computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
  • the computer-readable medium carries one or more programs.
  • the electronic device executes the method shown in the above embodiment.
  • Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages.
  • the program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server.
  • the remote computer may be connected to the user's computer through any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., using the Internet). Internet connection through an Internet service provider).
  • LAN Local Area Network
  • WAN Wide Area Network
  • Internet connection through an Internet service provider).
  • each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function.
  • the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved.
  • each square box in the block diagram and/or flow chart, and the combination of the square boxes in the block diagram and/or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
  • the units involved in the embodiments described in the present disclosure may be implemented by software or hardware.
  • the name of a unit does not limit the unit itself in some cases.
  • the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses".
  • exemplary types of hardware logic components include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
  • FPGAs field programmable gate arrays
  • ASICs application specific integrated circuits
  • ASSPs application specific standard products
  • SOCs systems on chips
  • CPLDs complex programmable logic devices
  • a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment.
  • a machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium.
  • a machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing.
  • a more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
  • RAM random access memory
  • ROM read-only memory
  • EPROM or flash memory erasable programmable read-only memory
  • CD-ROM portable compact disk read-only memory
  • CD-ROM compact disk read-only memory
  • magnetic storage device or any suitable combination of the foregoing.
  • a live broadcast processing method is provided, which is applied to a first live broadcast and microphone connection client, the method comprising: receiving a live broadcast from a second live broadcast and microphone connection client; The method comprises the steps of: performing a stream broadcast and caching the latest video frame of the live stream of the second live broadcast and microphone connection client; when obtaining the current video frame of the first live broadcast and microphone connection client, reading the cached latest video frame; based on a single rendering view, synthesizing the current video frame and the latest video frame into a target video frame according to preset layout information; and displaying the target video frame.
  • synthesizing the current video frame and the latest video frame into a target video frame according to preset layout information includes: determining the size and position of a first layer corresponding to the current video frame and a second layer corresponding to the latest video frame in the target video frame according to the preset layout information; drawing the current video frame into the first layer, drawing the latest video frame into the second layer, respectively, and synthesizing the first layer and the second layer into the target video frame.
  • synthesizing the current video frame and the latest video frame into a target video frame according to preset layout information includes: determining a target template for template testing in a fragment shader rendering pipeline, wherein a depth value at an area corresponding to the current video frame in the target template is a level value corresponding to the current video frame, and a depth value at an area corresponding to the latest video frame is a level value corresponding to the latest video frame; performing a template test according to the level value of the current video frame, the level value of the latest video frame, and the target template, and synthesizing the current video frame and the latest video frame based on the template test result.
  • the template test is performed based on the level value of the current video frame, the level value of the latest video frame and the target template, including: determining the level value at any first position of the current video frame as the actual depth value at the first position, comparing the actual depth value at the first position of the current video frame with the depth value at the corresponding position in the target template, if the actual depth value at the first position is higher than or equal to the depth value at the corresponding position in the target template, the template test of the first position passes; determining the level value at any second position of the latest video frame as the actual depth value at the second position, comparing the actual depth value at the second position of the latest video frame with the depth value at the corresponding position in the target template, if the actual depth value at the second position is higher than or equal to the depth value at the corresponding position in the target template, the template test of the second position passes.
  • the synthesizing the current video frame and the latest video frame based on the template test result includes: if the template test of any first position of the current video frame passes, rendering the first position of the current video frame; if the latest If the template test of any second position of the video frame passes, the second position of the latest video frame is rendered; after the positions that pass the template test in the current video frame and the latest video frame are rendered, the synthesized target video frame is obtained.
  • the level value of the latest video frame is higher than the level value of the current video frame.
  • the method of determining a target template for template testing in a fragment shader rendering pipeline includes: setting depth values at all positions of the target template to the level value of the current video frame; and updating the depth value at the area corresponding to the latest video frame in the target template to the level value of the latest video frame according to the level value of the latest video frame and the preset layout information.
  • the present disclosure before synthesizing the current video frame and the latest video frame into a target video frame based on a single rendering view according to preset layout information, it also includes: obtaining the size and position of each view control in a preset list view control, using the size and position of any view control as the size and position of the second layer corresponding to the latest video frame in the target video frame, and obtaining the preset layout information to align the second layer corresponding to the latest video frame with any view control in the preset list view control.
  • the synthesizing the current video frame and the latest video frame into a target video frame based on a single rendering view and according to preset layout information includes: based on a single rendering view and a single rendering thread, synthesizing the current video frame and the latest video frame into a target video frame according to preset layout information.
  • displaying the target video frame includes: displaying the target video frame using a single surface view.
  • the synthesizing the current video frame and the latest video frame into a target video frame based on a single rendering view according to preset layout information also includes: rendering the target video frame into a single frame buffer object; the displaying the target video frame using a single surface view includes: displaying the rendered target video frame cached in the frame buffer object using a single surface view.
  • a live broadcast processing device which is applied to a first live broadcast and microphone client, and the device includes: a receiving unit, which is used to receive a live broadcast stream of a second live broadcast and microphone client; a cache unit, which is used to cache the live broadcast stream of the second live broadcast and microphone client; The new video frame is cached; when the current video frame of the first live broadcast and microphone connection client is obtained, the cached latest video frame is read; a rendering unit is used to synthesize the current video frame and the latest video frame into a target video frame based on a single rendering view and according to preset layout information; a display unit is used to display the target video frame.
  • the rendering unit when the rendering unit synthesizes the current video frame and the latest video frame into a target video frame according to preset layout information, it is used to: determine the size and position of the first layer corresponding to the current video frame and the second layer corresponding to the latest video frame in the target video frame according to the preset layout information; draw the current video frame into the first layer, draw the latest video frame into the second layer, respectively, and synthesize the first layer and the second layer into the target video frame.
  • the rendering unit when the rendering unit synthesizes the current video frame and the latest video frame into a target video frame according to preset layout information, it is used to: determine a target template for template testing in a fragment shader rendering pipeline, wherein a depth value at an area corresponding to the current video frame in the target template is a level value corresponding to the current video frame, and a depth value at an area corresponding to the latest video frame is a level value corresponding to the latest video frame; perform a template test according to the level value of the current video frame, the level value of the latest video frame and the target template, and synthesize the current video frame and the latest video frame based on the template test result.
  • the rendering unit when the rendering unit performs a template test based on the level value of the current video frame, the level value of the latest video frame of each live stream, and the target template, the rendering unit is used to: determine the level value at any first position of the current video frame as the actual depth value at the first position, compare the actual depth value at the first position of the current video frame with the depth value at the corresponding position in the target template, if the actual depth value at the first position is higher than or equal to the depth value at the corresponding position in the target template, the template test of the first position passes; determine the level value at any second position of the latest video frame as the actual depth value at the second position, compare the actual depth value at the second position of the latest video frame with the depth value at the corresponding position in the target template, if the actual depth value at the second position is higher than or equal to the depth value at the corresponding position in the target template, the template test of the second position passes.
  • the rendering unit when the rendering unit synthesizes the current video frame and the latest video frame based on the template test result, If the template test of any first position passes, the first position of the current video frame is rendered; if the template test of any second position of the latest video frame passes, the second position of the latest video frame is rendered; after the positions that pass the template test in the current video frame and the latest video frame are rendered, the synthesized target video frame is obtained.
  • the level value of the latest video frame is higher than the level value of the current video frame.
  • the rendering unit when determining a target template for template testing in a fragment shader rendering pipeline, is used to: set the depth values at all positions of the target template to the level value of the current video frame; and update the depth value at the area corresponding to the latest video frame in the target template to the level value of the latest video frame according to the level value of the latest video frame and the preset layout information.
  • the rendering unit synthesizes the current video frame and the latest video frame of each live stream into a target video frame based on a single rendering view and according to preset layout information
  • the rendering unit is also used to: obtain the size and position of each view control in a preset list view control, use the size and position of any view control as the size and position of the second layer corresponding to the latest video frame in the target video frame, and obtain the preset layout information so that the second layer corresponding to the latest video frame is aligned with any view control in the preset list view control.
  • the rendering unit when the rendering unit synthesizes the current video frame and the latest video frame into a target video frame based on a single rendering view and according to preset layout information, it is used to: synthesize the current video frame and the latest video frame into a target video frame based on a single rendering view and a single rendering thread according to preset layout information.
  • the display unit when displaying the target video frame, is configured to: display the target video frame using a single surface view.
  • the rendering unit when the rendering unit synthesizes the current video frame and the latest video frame into a target video frame based on a single rendering view and according to preset layout information, the rendering unit is also used to: render the target video frame into a single frame buffer object; when the display unit displays the target video frame using a single surface view, the display unit is used to: display the rendered target video frame cached in the frame buffer object using a single surface view.
  • an electronic device comprising: at least one processor and a memory; the memory stores computer-executable instructions; the at least A processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the live broadcast processing method as described in the first aspect and various possible designs of the first aspect.
  • a computer-readable storage medium stores computer execution instructions.
  • the live broadcast processing method described in the first aspect and various possible designs of the first aspect is implemented.
  • a computer program product comprising computer execution instructions.
  • a processor executes the computer execution instructions, the live broadcast processing method as described in the first aspect and various possible designs of the first aspect is implemented.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • General Engineering & Computer Science (AREA)
  • Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)

Abstract

本公开实施例提供一种直播处理方法、设备及存储介质,通过第一直播连麦客户端接收第二直播连麦客户端的直播流,并对第二直播连麦客户端的直播流的最新视频帧进行缓存;在获取所述第一直播连麦客户端的当前视频帧时,读取所缓存的所述最新视频帧;基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧;对所述目标视频帧进行显示。在多人连线直播场景中通过将第一直播连麦客户端的当前视频帧以及其他的第二直播连麦客户端直播流的最新视频帧合成为一帧目标视频帧,只需要采用单个渲染视图即可实现渲染,降低了资源占用,降低对直播连麦客户端的性能压力,提升直播流畅度。

Description

直播处理方法、设备及存储介质
本申请要求2023年6月28日递交的、标题为“直播处理方法、设备及存储介质”、申请号为:202310782876.5的中国发明专利申请的优先权,该申请的全部内容通过引用结合在本申请中。
技术领域
本公开实施例涉及计算机与网络通信技术领域,尤其涉及一种直播处理方法、设备及存储介质。
背景技术
在网络直播的过程中,多人连线是一种主播经常采用的与其它主播进行互动直播的方式,基于这种方式,当前直播间的主播可以邀请其它主播一起进行直播。
然而随着多人连线直播场景中主播人数的增加,直播流畅度及直播连麦客户端性能受到严重影响、且可能严重发热。
发明内容
本公开实施例提供一种直播处理方法、设备及存储介质,以在多人连麦直播场景中降低对直播连麦客户端的性能压力,提升直播流畅度。
第一方面,本公开实施例提供一种直播处理方法,应用于第一直播连麦客户端,所述方法包括:接收第二直播连麦客户端的直播流,并对第二直播连麦客户端的直播流的最新视频帧进行缓存;在获取所述第一直播连麦客户端的当前视频帧时,读取所缓存的所述最新视频帧;基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧;对所述目标视频帧进行显示。
第二方面,本公开实施例提供一种直播处理装置,应用于第一直播连麦客户端,所述装置包括:接收单元,用于接收第二直播连麦客户端的直播流; 缓存单元,用于对第二直播连麦客户端的直播流的最新视频帧进行缓存;在获取所述第一直播连麦客户端的当前视频帧时,读取所缓存的所述最新视频帧;渲染单元,用于基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧;显示单元,用于对所述目标视频帧进行显示。
第三方面,本公开实施例提供一种电子设备,包括:至少一个处理器和存储器;所述存储器存储计算机执行指令;所述至少一个处理器执行所述存储器存储的计算机执行指令,使得所述至少一个处理器执行如上第一方面以及第一方面各种可能的设计所述的直播处理方法。
第四方面,本公开实施例提供一种计算机可读存储介质,所述计算机可读存储介质中存储有计算机执行指令,当处理器执行所述计算机执行指令时,实现如上第一方面以及第一方面各种可能的设计所述的直播处理方法。
第五方面,本公开实施例提供一种计算机程序产品,包括计算机执行指令,当处理器执行所述计算机执行指令时,实现如上第一方面以及第一方面各种可能的设计所述的直播处理方法。
附图说明
为了更清楚地说明本公开实施例或现有技术中的技术方案,下面将对实施例或现有技术描述中所需要使用的附图作一简单地介绍,显而易见地,下面描述中的附图是本公开的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动性的前提下,还可以根据这些附图获得其他的附图。
图1为本公开一实施例提供的直播处理方法的场景示例图;
图2为本公开一实施例提供的直播处理方法流程示意图;
图3为本公开另一实施例提供的直播处理方法流程示意图;
图4a为本公开一实施例提供的一种布局方式的界面示意图;
图4b为本公开另一实施例提供的一种布局方式的界面示意图;
图5为本公开一实施例提供的每一直播流的最新视频帧对应的图层与预设列表视图控件对齐的示意图;
图6为本公开另一实施例提供的直播处理方法流程示意图;
图7为本公开一实施例提供的目标模板的示意图;
图8为本公开另一实施例提供的直播处理方法流程示意图;
图9为本公开一实施例提供的确定目标模板的示意图;
图10为本公开一实施例提供的直播处理设备的结构框图;
图11为本公开一实施例提供的电子设备的硬件结构示意图。
具体实施方式
为使本公开实施例的目的、技术方案和优点更加清楚,下面将结合本公开实施例中的附图,对本公开实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例是本公开一部分实施例,而不是全部的实施例。基于本公开中的实施例,本领域普通技术人员在没有作出创造性劳动前提下所获得的所有其他实施例,都属于本公开保护的范围。
在网络直播的过程中,多人连线是一种主播经常采用的与其它主播进行互动直播的方式,基于这种方式,当前直播间的主播可以邀请其它主播一起进行直播。
在多人连线直播场景中,对于任一第一直播连麦客户端而言,会对第一直播连麦客户端的当前视频帧以及每一其他第二直播连麦客户端直播流的视频帧分别采用不同的渲染视图、不同的GL渲染线程、不同的帧缓存对象、以及不同的表面视图,渲染视图、GL渲染线程帧缓存对象、以及表面视图无法复用,因此占用了直播连麦客户端大量资源。然而随着多人连线直播场景中主播人数的增加,直播流畅度受到严重影响,直播连麦客户端性能受到严重影响、且可能严重发热。
为解决上述技术问题,本公开提供一种直播处理方法,对于多人连麦直播间中任一第一直播连麦客户端,通过直播连麦客户端接收第二直播连麦客户端的直播流,并对每一直播流的最新视频帧进行缓存;在获取第一直播连麦客户端的当前视频帧时,读取所缓存的最新视频帧;基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧;对所述目标视频帧进行显示。在多人连线直播场景中通过将第一直播连麦客户端的当前视频帧以及其他第二直播连麦客户端直播流的最新视频帧合成为一帧目标视频帧,只需要采用单个渲染视图即可实现渲染,降低了资源占用,降低对直播连麦客户端的性能压力,提升直播流畅度。
本公开提供的直播处理方法应用场景如图1所示,执行主体为第一直播连麦客户端,第一直播连麦客户端可以接收一个或多个第二直播连麦客户端的直播流,并将每一直播流的最新视频帧缓存到对应的图层Layer,而第一直播连麦客户端可以采集当前视频帧,作为初始图层OriginLayer,进而在获取第一直播连麦客户端采集到当前视频帧时,读取所缓存的每一直播流的最新视频帧,将所述当前视频帧和所述每一直播流的最新视频帧合成为一个目标视频帧进行显示。
需要说明的是,本申请所涉及的用户信息和数据,均为经用户授权或者经过各方充分授权的信息和数据,并且相关数据的收集、使用和处理需要遵守相关国家和地区的相关法律法规和标准,并提供有相应的操作入口,供用户选择授权或者拒绝。
下面将结合具体实施例对本公开的直播处理方法进行详细介绍。
参考图2,图2为本公开一实施例提供的直播处理方法流程示意图。本实施例的方法可以应用在多人连麦直播间中的任一直播连麦客户端(记为第一直播连麦客户端)中,该直播处理方法包括:
S201、接收第二直播连麦客户端的直播流,并对第二直播连麦客户端的直播流的最新视频帧进行缓存。
在本实施例中,在多人连线直播场景中第一直播连麦客户端可邀请一个或多个第二直播连麦客户端连麦,从而在第一直播连麦客户端和第二直播连麦客户端的界面中同时观看到第一直播连麦客户端的画面以及一个或多个第二直播连麦客户端的画面。在多人连麦直播场景中第一直播连麦客户端可以接收到一个或多个第二直播连麦客户端的直播流,具体的,一个或多个第二直播连麦客户端可以将各自的直播流上传服务端,服务端可以将一个或多个第二直播连麦客户端的直播流发送给第一直播连麦客户端。
进一步的,对于每一第二直播连麦客户端的直播流进行解析处理,得到直播流中的视频帧,并对其中最新的视频帧进行缓存,具体的,可以将每一直播流的最新视频帧分别绘制并缓存到对应的一个图层Layer中。
S202、在获取所述第一直播连麦客户端的当前视频帧时,读取所缓存的所述最新视频帧。
在本实施例中,第一直播连麦客户端可以实时采集第一直播连麦客户端 侧的画面,得到当前视频帧,基于当前视频帧来驱动当前视频帧和每一第二直播连麦客户端的直播流的最新视频帧的合成,由于第一直播连麦客户端采集的当前视频帧无需经过网络传输,因此输出较为稳定第一直播连麦客户端采集的当前视频帧,因此由第一直播连麦客户端采集的当前视频帧来驱动可以保证直播画面的稳定。在获取到第一直播连麦客户端采集到的当前视频帧时,可以读取所缓存的每一直播流的最新视频帧,具体的,可从各图层Layer中读取各直播流的最新视频帧,以用于目标视频帧的合成。
S203、基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述每一直播流的最新视频帧合成为目标视频帧。
在本实施例中,在获取到第一直播连麦客户端采集的当前视频帧以及每一第二直播连麦客户端的直播流的最新视频帧后,可将第一直播连麦客户端采集的当前视频帧以及每一直播流的最新视频帧合成绘制到同一张纹理上,得到一个目标视频帧,共用一个渲染视图(RenderView),来实现在一个视图中渲染当前视频帧和每一直播流的最新视频帧,而不需要对第一直播连麦客户端的当前视频帧以及每一第二直播连麦客户端直播流的视频帧分别采用不同的渲染视图,降低了对第一直播连麦客户端的性能压力。在合成过程中根据预设布局信息可以确定当前视频帧和每一直播流的最新视频帧如何布局。
此外,在基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述每一直播流的最新视频帧合成为一个目标视频帧时,只需要采用一个GL渲染线程,而不需要对第一直播连麦客户端的当前视频帧以及每一第二直播连麦客户端直播流的视频帧分别采用一个GL渲染线程,进一步降低了对第一直播连麦客户端的性能压力。
可选的,本实施例中可将目标视频帧渲染到单个帧缓存对象(Frame Buffer Object,FBO)中,避免了对第一直播连麦客户端的当前视频帧以及每一第二直播连麦客户端直播流的视频帧分别采用不同FBO进行缓存,进一步降低了对第一直播连麦客户端的性能压力。
S204、对所述目标视频帧进行显示。
在本实施例中,在对目标视频帧合成并渲染之后在第一直播连麦客户端的显示界面中进行显示。
由于第一直播连麦客户端的当前视频帧以及每一第二直播连麦客户端直 播流的最新视频帧合成为一帧目标视频帧,因此本实施例中可以只需要采用一个表面视图SurfaceView即可对目标视频帧进行显示,而不需要对第一直播连麦客户端的当前视频帧以及每一第二直播连麦客户端直播流的视频帧分别采用不同SurfaceView进行显示,进一步降低了对第一直播连麦客户端的性能压力。
可选的,若目标视频帧渲染到单个FBO中,则将FBO中缓存的经过渲染的目标视频帧采用单个表面视图进行显示。
本实施例提供的直播处理方法,通过第一直播连麦客户端接收第二直播连麦客户端的直播流,并对第二直播连麦客户端的直播流的最新视频帧进行缓存;在获取所述第一直播连麦客户端的当前视频帧时,读取所缓存的所述最新视频帧;基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧;对所述目标视频帧进行显示。在多人连线直播场景中通过将第一直播连麦客户端的当前视频帧以及其他的第二直播连麦客户端直播流的最新视频帧合成为一帧目标视频帧,只需要采用单个渲染视图即可实现渲染,降低了资源占用,降低对直播连麦客户端的性能压力,提升直播流畅度。
在上述实施例的基础上,如图3所示,S203所述的根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧,包括:
S301、根据预设布局信息确定所述当前视频帧对应的第一图层和所述最新视频帧对应的第二图层在所述目标视频帧中的尺寸和位置。
S302、将所述当前视频帧绘制到所述第一图层中,将所述最新视频帧分别绘制到第二图层中,将所述第一图层和所述第二图层合成为所述目标视频帧。
在本实施例中,在合成过程中根据预设布局信息可以确定第一直播连麦客户端的当前视频帧以及每一第二直播连麦客户端直播流的最新视频帧在目标视频帧中如何布局。例如,如图4a所示,第一直播连麦客户端的当前视频帧的第一图层(Layer)位于目标视频帧中的左侧,显示第一直播连麦客户端的当前视频帧的画面,每一第二直播连麦客户端直播流的最新视频帧的第二图层(Layer)纵向排成一列位于目标视频帧中的右侧,显示每一第二直播连 麦客户端的画面;再如,如图4b所示,第一直播连麦客户端的当前视频帧的第一图层(Layer)在目标视频帧中全屏显示,显示第一直播连麦客户端的当前视频帧的画面,每一第二直播连麦客户端直播流的最新视频帧的第二图层(Layer)以浮窗形式位于第一直播连麦客户端的当前视频帧之上,每一第二直播连麦客户端的画面悬浮于当前视频帧的画面之上;当然也可以采用其他的布局方式,在每种布局方式中,均可根据预设布局信息确定当前视频帧对应的第一图层和每一直播流的最新视频帧对应的第二图层在所述目标视频帧中的尺寸和位置,进而将当前视频帧绘制到当前视频帧对应的第一图层中,将每一直播流的最新视频帧分别绘制到每一直播流的最新视频帧对应的第二图层中,从而将各图层合成为所述目标视频帧。
可选的,在S203所述基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧前,还包括:获取预设列表视图控件中每一视图控件的尺寸和位置,将任一视图控件的尺寸和位置分别作为任一第二直播连麦客户端直播流的最新视频帧对应的第二图层在目标视频帧中的尺寸和位置,得到预设布局信息,以使该直播流的最新视频帧对应的第二图层分别与预设列表视图控件中该视图控件对齐。
在本实施例中,由于目标视频帧最终采用单个表面视图,而目标视频帧中每一直播流还需要进行UI交互,在多人连线直播场景的显示界面中可设置预设列表视图控件(RecyclerView),预设列表视图控件中包括多个视图控件(ViewHolder),例如图5所示预设列表视图控件RecyclerView为九宫格的列表视图控件,则包括九个视图控件ViewHolder,将每一第二直播连麦客户端直播流的最新视频帧对应的图层分别与每一视图控件ViewHolder对齐,可实现通过每一视图控件ViewHolder对对应的直播流进行交互。为了实现每一第二直播连麦客户端直播流的最新视频帧对应的第二图层分别与每一视图控件对齐,可获取预设列表视图控件中每一视图控件的尺寸和位置,将每一视图控件的尺寸和位置分别作为每一直播流的最新视频帧对应的图层在所述目标视频帧中的尺寸和位置,从而作为所述预设布局信息,进而在合成时,根据预设布局信息将当前视频帧和每一第二直播连麦客户端直播流的最新视频帧合成为一个目标视频帧,这样目标视频帧中的每一直播流的最新视频帧对应的第二图层分别与预设列表视图控件中每一视图控件对齐。
在上述任一实施例的基础上,在合成时,第一直播连麦客户端的当前视频帧以及第二直播连麦客户端直播流的最新视频帧可能存在一定的重叠区域,因此需要设置显示的层级值,层级最高的置于顶层,层级最低的置于底层,若存在重叠区域,则层级高的视频帧覆盖层级低的视频帧。可选的,如图6所示,在S203所述根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧,包括:
S401、确定片元着色器渲染管线中用于模板测试的目标模板,其中所述目标模板中所述当前视频帧对应区域处的深度值为所述当前视频帧对应的层级值,所述最新视频帧对应区域处的深度值为所述最新视频帧对应的层级值;
S402、根据所述当前视频帧的层级值、所述最新视频帧的层级值以及所述目标模板进行模板测试,基于模板测试结果对所述当前视频帧和所述最新视频帧进行合成。
在本实施例中,在OpenGL片元着色器的渲染管线中包括模板测试(Stencil Test)环节,模板测试是基于模板判断片段是否保留或丢弃,若某一片段未通过模板测试则丢弃该片段,也即不绘制该片段,若某一片段通过模板测试则保留该片段。本实施例中可预先配置片元着色器渲染管线中用于模板测试的目标模板,目标模板中当前视频帧以及每一第二直播连麦客户端直播流的最新视频帧对应区域处的深度值为各区域对应的层级值zOrder。如图7所示,层级值zOrder越高,数字越大,也即深度值越高。例如第二直播连麦客户端直播流的最新视频帧悬浮于第一直播连麦客户端的当前视频帧之上,也即第二直播连麦客户端直播流的最新视频帧的层级值zOrder高于第一直播连麦客户端的当前视频帧的层级值zOrder,可第二直播连麦客户端直播流的最新视频帧的层级值zOrder设置为3,第一直播连麦客户端的当前视频帧的层级值zOrder设置为2,在目标模板中直播流的最新视频帧对应区域(方框内)处每一个位置的深度值为3,第一直播连麦客户端的当前视频帧未被覆盖的区域处每一个位置的深度值为2。
在对第一直播连麦客户端的当前视频帧、第二直播连麦客户端直播流的最新视频帧进行合成时,根据当前视频帧的层级值、最新视频帧的层级值以及目标模板进行模板测试,对当前视频帧的层级值、最新视频帧中通过模板测试的位置进行渲染绘制,未通过模板测试的位置不进行渲染绘制,最终得 到合成后的目标视频帧。
在一些实施例中,模板测试的具体过程可包括:
将所述当前视频帧的任意第一位置处的层级值确定为所述第一位置处的实际深度值,将所述当前视频帧的所述第一位置处的实际深度值与所述目标模板中对应位置处的深度值进行比对,若所述第一位置处的实际深度值高于或等于所述目标模板中对应位置处的深度值,则所述第一位置的模板测试通过,可对所述当前视频帧的所述第一位置进行渲染;而若所述第一位置处的实际深度值低于所述目标模板中对应位置处的深度值,则所述第一位置的模板测试未通过,不对所述当前视频帧的所述第一位置进行渲染;
将任一第二直播连麦客户端直播流的最新视频帧的任意第二位置处的层级值确定为所述第二位置处的实际深度值,将所述最新视频帧的所述第二位置处的实际深度值与所述目标模板中对应位置处的深度值进行比对,若所述第二位置处的实际深度值高于或等于所述目标模板中对应位置处的深度值,则所述第二位置的模板测试通过,可对所述最新视频帧的所述第二位置进行渲染;若所述第二位置处的实际深度值低于所述目标模板中对应位置处的深度值,则所述第二位置的模板测试未通过,不对所述最新视频帧的所述第二位置进行渲染。
通过对当前视频帧中通过模板测试的所有第一位置进行渲染、以及对每一直播流的最新视频帧中通过模板测试的所有第二位置进行渲染,从而实现合成时的层级关系,在当前视频帧和每一第二直播连麦客户端直播流的最新视频帧中通过模板测试的位置完成渲染后,得到合成后的所述目标视频帧。
在上述实施例的基础上,若所述任一直播流的最新视频帧的层级值高于所述当前视频帧的层级值,则S401所述确定片元着色器渲染管线中用于模板测试的目标模板,如图8所示,可包括:
S501、将所述目标模板所有位置处的深度值设置为所述当前视频帧的层级值;
S502、根据所述最新视频帧的层级值以及所述预设布局信息,将所述目标模板中所述最新视频帧对应区域处的深度值更新为所述最新视频帧的层级值。
在本实施例中,获取预先设置的当前视频帧的层级值zOrder,假设当前 视频帧的层级值zOrder为2,将目标模板所有位置处的深度值设置为当前视频帧的层级值,也即目标模板所有位置处的深度值均为2,如图9上侧所示;获取预先设置的任一第二直播连麦客户端直播流的最新视频帧的层级值zOrder,假设任一第二直播连麦客户端直播流的最新视频帧的层级值zOrder为3,且该直播流的最新视频帧在目标视频帧中的位置和尺寸可从预设布局信息中获得,因此,可确定任一第二直播连麦客户端直播流的最新视频帧在目标模板中的对应区域,也即图中方框区域,进一步的,将目标模板中任一第二直播连麦客户端直播流的最新视频帧对应区域处的深度值更新为该第二直播连麦客户端直播流的最新视频帧的层级值,也即将目标模板方框区域内各位置的深度值更新为3,如图9下侧所示,最终得到目标模板。
在一些实施例中,目标模板配置有缓冲区(缓存),S501-S502是在目标模板的缓冲区内实现,也即在目标模板的缓冲区中将所述目标模板所有位置处的深度值设置为所述当前视频帧的层级值;在目标模板的缓冲区中将所述目标模板中所述任一第二直播连麦客户端直播流的最新视频帧对应区域处的深度值更新为所述任一第二直播连麦客户端直播流的最新视频帧的层级值。
本公开实施例提供的直播处理方法、设备及存储介质,通过第一直播连麦客户端接收第二直播连麦客户端的直播流,并对第二直播连麦客户端的直播流的最新视频帧进行缓存;在获取所述第一直播连麦客户端的当前视频帧时,读取所缓存的所述最新视频帧;基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧;对所述目标视频帧进行显示。在多人连线直播场景中通过将第一直播连麦客户端的当前视频帧以及其他的第二直播连麦客户端直播流的最新视频帧合成为一帧目标视频帧,只需要采用单个渲染视图即可实现渲染,降低了资源占用,降低对直播连麦客户端的性能压力,提升直播流畅度。
对应于上文实施例的直播处理方法,图10为本公开实施例提供的直播处理装置的结构框图。为了便于说明,仅示出了与本公开实施例相关的部分。参照图10,所述直播处理装置600包括:接收单元601、缓存单元602、渲染单元603、显示单元604。
在一些实施例中,接收单元601,用于接收第二直播连麦客户端的直播流;缓存单元602,用于对第二直播连麦客户端的直播流的最新视频帧进行缓 存;在获取所述第一直播连麦客户端的当前视频帧时,读取所缓存的所述最新视频帧;渲染单元603,用于基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧;显示单元604,用于对所述目标视频帧进行显示。
在本公开的一个或多个实施例中,所述渲染单元603在根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧时,用于:根据预设布局信息确定所述当前视频帧对应的第一图层和所述最新视频帧对应的第二图层在所述目标视频帧中的尺寸和位置;将所述当前视频帧绘制到所述第一图层中,将所述最新视频帧分别绘制到第二图层中,将所述第一图层和所述第二图层合成为所述目标视频帧。
在本公开的一个或多个实施例中,所述渲染单元603在根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧时,用于:确定片元着色器渲染管线中用于模板测试的目标模板,其中所述目标模板中所述当前视频帧对应区域处的深度值为所述当前视频帧对应的层级值,所述最新视频帧对应区域处的深度值为所述最新视频帧对应的层级值;根据所述当前视频帧的层级值、所述最新视频帧的层级值以及所述目标模板进行模板测试,基于模板测试结果对所述当前视频帧和所述最新视频帧进行合成。
在本公开的一个或多个实施例中,所述渲染单元603在根据所述当前视频帧的层级值、所述每一直播流的最新视频帧的层级值以及所述目标模板进行模板测试时,用于:将所述当前视频帧的任意第一位置处的层级值确定为所述第一位置处的实际深度值,将所述当前视频帧的所述第一位置处的实际深度值与所述目标模板中对应位置处的深度值进行比对,若所述第一位置处的实际深度值高于或等于所述目标模板中对应位置处的深度值,则所述第一位置的模板测试通过;将所述最新视频帧的任意第二位置处的层级值确定为所述第二位置处的实际深度值,将所述最新视频帧的所述第二位置处的实际深度值与所述目标模板中对应位置处的深度值进行比对,若所述第二位置处的实际深度值高于或等于所述目标模板中对应位置处的深度值,则所述第二位置的模板测试通过。
在本公开的一个或多个实施例中,所述渲染单元603在基于模板测试结果对所述当前视频帧和所述最新视频帧进行合成时,用于:若所述当前视频 帧任意第一位置的模板测试通过,则对所述当前视频帧的所述第一位置进行渲染;若所述最新视频帧任意第二位置的模板测试通过,则对所述最新视频帧的所述第二位置进行渲染;在所述当前视频帧和所述最新视频帧中通过模板测试的位置完成渲染后,得到合成后的所述目标视频帧。
在本公开的一个或多个实施例中,所述最新视频帧的层级值高于所述当前视频帧的层级值。
在本公开的一个或多个实施例中,所述渲染单元603在确定片元着色器渲染管线中用于模板测试的目标模板时,用于:将所述目标模板所有位置处的深度值设置为所述当前视频帧的层级值;根据所述最新视频帧的层级值以及所述预设布局信息,将所述目标模板中所述最新视频帧对应区域处的深度值更新为所述最新视频帧的层级值。
在本公开的一个或多个实施例中,所述渲染单元603在基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述每一直播流的最新视频帧合成为目标视频帧前,还用于:获取预设列表视图控件中每一视图控件的尺寸和位置,将任一视图控件的尺寸和位置作为所述最新视频帧对应的第二图层在所述目标视频帧中的尺寸和位置,得到所述预设布局信息,以使所述最新视频帧对应的第二图层与所述预设列表视图控件中所述任一视图控件对齐。
在本公开的一个或多个实施例中,所述渲染单元603在基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧时,用于:基于单个渲染视图和单个渲染线程,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧。
在本公开的一个或多个实施例中,所述显示单元604在对所述目标视频帧进行显示时,用于:采用单个表面视图对所述目标视频帧进行显示。
在本公开的一个或多个实施例中,所述渲染单元603在基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧时,还用于:将所述目标视频帧渲染到单个帧缓存对象中;所述显示单元604在采用单个表面视图对所述目标视频帧进行显示时,用于:将所述帧缓存对象中缓存的经过渲染的目标视频帧采用单个表面视图进行显示。
本实施例提供的直播处理装置,可用于执行上述方法实施例的技术方案,其实现原理和技术效果类似,本实施例此处不再赘述。
参考图11,其示出了适于用来实现本公开实施例的电子设备700的结构示意图,该电子设备700可以为终端设备或服务器。其中,终端设备可以包括但不限于诸如移动电话、笔记本电脑、数字广播接收器、个人数字助理(Personal Digital Assistant,简称PDA)、平板电脑(Portable Android Device,简称PAD)、便携式多媒体播放器(Portable Media Player,简称PMP)、车载终端(例如车载导航终端)等等的移动终端以及诸如数字TV、台式计算机等等的固定终端。图11示出的电子设备仅仅是一个示例,不应对本公开实施例的功能和使用范围带来任何限制。
如图11所示,电子设备700可以包括处理装置(例如中央处理器、图形处理器等)701,其可以根据存储在只读存储器(Read Only Memory,简称ROM)702中的程序或者从存储装置708加载到随机访问存储器(Random Access Memory,简称RAM)703中的程序而执行各种适当的动作和处理。在RAM 703中,还存储有电子设备700操作所需的各种程序和数据。处理装置701、ROM 702以及RAM 703通过总线704彼此相连。输入/输出(I/O)接口705也连接至总线704。
通常,以下装置可以连接至I/O接口705:包括例如触摸屏、触摸板、键盘、鼠标、摄像头、麦克风、加速度计、陀螺仪等的输入装置706;包括例如液晶显示器(Liquid Crystal Display,简称LCD)、扬声器、振动器等的输出装置707;包括例如磁带、硬盘等的存储装置708;以及通信装置709。通信装置709可以允许电子设备700与其他设备进行无线或有线通信以交换数据。虽然图11示出了具有各种装置的电子设备700,但是应理解的是,并不要求实施或具备所有示出的装置。可以替代地实施或具备更多或更少的装置。
特别地,根据本公开的实施例,上文参考流程图描述的过程可以被实现为计算机软件程序。例如,本公开的实施例包括一种计算机程序产品,其包括承载在计算机可读介质上的计算机程序,该计算机程序包含用于执行流程图所示的方法的程序代码。在这样的实施例中,该计算机程序可以通过通信装置709从网络上被下载和安装,或者从存储装置708被安装,或者从ROM 702被安装。在该计算机程序被处理装置701执行时,执行本公开实施例的方法中限定的上述功能。
需要说明的是,本公开上述的计算机可读介质可以是计算机可读信号介 质或者计算机可读存储介质或者是上述两者的任意组合。计算机可读存储介质例如可以是——但不限于——电、磁、光、电磁、红外线、或半导体的系统、装置或器件,或者任意以上的组合。计算机可读存储介质的更具体的例子可以包括但不限于:具有一个或多个导线的电连接、便携式计算机磁盘、硬盘、随机访问存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(EPROM或闪存)、光纤、便携式紧凑磁盘只读存储器(CD-ROM)、光存储器件、磁存储器件、或者上述的任意合适的组合。在本公开中,计算机可读存储介质可以是任何包含或存储程序的有形介质,该程序可以被指令执行系统、装置或者器件使用或者与其结合使用。而在本公开中,计算机可读信号介质可以包括在基带中或者作为载波一部分传播的数据信号,其中承载了计算机可读的程序代码。这种传播的数据信号可以采用多种形式,包括但不限于电磁信号、光信号或上述的任意合适的组合。计算机可读信号介质还可以是计算机可读存储介质以外的任何计算机可读介质,该计算机可读信号介质可以发送、传播或者传输用于由指令执行系统、装置或者器件使用或者与其结合使用的程序。计算机可读介质上包含的程序代码可以用任何适当的介质传输,包括但不限于:电线、光缆、RF(射频)等等,或者上述的任意合适的组合。
上述计算机可读介质可以是上述电子设备中所包含的;也可以是单独存在,而未装配入该电子设备中。
上述计算机可读介质承载有一个或者多个程序,当上述一个或者多个程序被该电子设备执行时,使得该电子设备执行上述实施例所示的方法。
可以以一种或多种程序设计语言或其组合来编写用于执行本公开的操作的计算机程序代码,上述程序设计语言包括面向对象的程序设计语言—诸如Java、Smalltalk、C++,还包括常规的过程式程序设计语言—诸如“C”语言或类似的程序设计语言。程序代码可以完全地在用户计算机上执行、部分地在用户计算机上执行、作为一个独立的软件包执行、部分在用户计算机上部分在远程计算机上执行、或者完全在远程计算机或服务器上执行。在涉及远程计算机的情形中,远程计算机可以通过任意种类的网络——包括局域网(Local Area Network,简称LAN)或广域网(Wide Area Network,简称WAN)—连接到用户计算机,或者,可以连接到外部计算机(例如利用因特 网服务提供商来通过因特网连接)。
附图中的流程图和框图,图示了按照本公开各种实施例的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段、或代码的一部分,该模块、程序段、或代码的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。也应当注意,在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个接连地表示的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或操作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
描述于本公开实施例中所涉及到的单元可以通过软件的方式实现,也可以通过硬件的方式来实现。其中,单元的名称在某种情况下并不构成对该单元本身的限定,例如,第一获取单元还可以被描述为“获取至少两个网际协议地址的单元”。
本文中以上描述的功能可以至少部分地由一个或多个硬件逻辑部件来执行。例如,非限制性地,可以使用的示范类型的硬件逻辑部件包括:现场可编程门阵列(FPGA)、专用集成电路(ASIC)、专用标准产品(ASSP)、片上系统(SOC)、复杂可编程逻辑设备(CPLD)等等。
在本公开的上下文中,机器可读介质可以是有形的介质,其可以包含或存储以供指令执行系统、装置或设备使用或与指令执行系统、装置或设备结合地使用的程序。机器可读介质可以是机器可读信号介质或机器可读储存介质。机器可读介质可以包括但不限于电子的、磁性的、光学的、电磁的、红外的、或半导体系统、装置或设备,或者上述内容的任何合适组合。机器可读存储介质的更具体示例会包括基于一个或多个线的电气连接、便携式计算机盘、硬盘、随机存取存储器(RAM)、只读存储器(ROM)、可擦除可编程只读存储器(EPROM或快闪存储器)、光纤、便捷式紧凑盘只读存储器(CD-ROM)、光学储存设备、磁储存设备、或上述内容的任何合适组合。
第一方面,根据本公开的一个或多个实施例,提供了一种直播处理方法,应用于第一直播连麦客户端,所述方法包括:接收第二直播连麦客户端的直 播流,并对第二直播连麦客户端的直播流的最新视频帧进行缓存;在获取所述第一直播连麦客户端的当前视频帧时,读取所缓存的所述最新视频帧;基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧;对所述目标视频帧进行显示。
根据本公开的一个或多个实施例,所述根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧,包括:根据预设布局信息确定所述当前视频帧对应的第一图层和所述最新视频帧对应的第二图层在所述目标视频帧中的尺寸和位置;将所述当前视频帧绘制到所述第一图层中,将所述最新视频帧分别绘制到第二图层中,将所述第一图层和所述第二图层合成为所述目标视频帧。
根据本公开的一个或多个实施例,所述根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧,包括:确定片元着色器渲染管线中用于模板测试的目标模板,其中所述目标模板中所述当前视频帧对应区域处的深度值为所述当前视频帧对应的层级值,所述最新视频帧对应区域处的深度值为所述最新视频帧对应的层级值;根据所述当前视频帧的层级值、所述最新视频帧的层级值以及所述目标模板进行模板测试,基于模板测试结果对所述当前视频帧和所述最新视频帧进行合成。
根据本公开的一个或多个实施例,所述根据所述当前视频帧的层级值、所述最新视频帧的层级值以及所述目标模板进行模板测试,包括:将所述当前视频帧的任意第一位置处的层级值确定为所述第一位置处的实际深度值,将所述当前视频帧的所述第一位置处的实际深度值与所述目标模板中对应位置处的深度值进行比对,若所述第一位置处的实际深度值高于或等于所述目标模板中对应位置处的深度值,则所述第一位置的模板测试通过;将所述最新视频帧的任意第二位置处的层级值确定为所述第二位置处的实际深度值,将所述最新视频帧的所述第二位置处的实际深度值与所述目标模板中对应位置处的深度值进行比对,若所述第二位置处的实际深度值高于或等于所述目标模板中对应位置处的深度值,则所述第二位置的模板测试通过。
根据本公开的一个或多个实施例,所述基于模板测试结果对所述当前视频帧和所述最新视频帧进行合成,包括:若所述当前视频帧任意第一位置的模板测试通过,则对所述当前视频帧的所述第一位置进行渲染;若所述最新 视频帧任意第二位置的模板测试通过,则对所述最新视频帧的所述第二位置进行渲染;在所述当前视频帧和所述最新视频帧中通过模板测试的位置完成渲染后,得到合成后的所述目标视频帧。
根据本公开的一个或多个实施例,所述最新视频帧的层级值高于所述当前视频帧的层级值。
根据本公开的一个或多个实施例,所述确定片元着色器渲染管线中用于模板测试的目标模板,包括:将所述目标模板所有位置处的深度值设置为所述当前视频帧的层级值;根据所述最新视频帧的层级值以及所述预设布局信息,将所述目标模板中所述最新视频帧对应区域处的深度值更新为所述最新视频帧的层级值。
根据本公开的一个或多个实施例,所述基于单个渲染视图根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧前,还包括:获取预设列表视图控件中每一视图控件的尺寸和位置,将任一视图控件的尺寸和位置作为所述最新视频帧对应的第二图层在所述目标视频帧中的尺寸和位置,得到所述预设布局信息,以使所述最新视频帧对应的第二图层与所述预设列表视图控件中所述任一视图控件对齐。
根据本公开的一个或多个实施例,所述基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧,包括:基于单个渲染视图和单个渲染线程,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧。
根据本公开的一个或多个实施例,所述对所述目标视频帧进行显示,包括:采用单个表面视图对所述目标视频帧进行显示。
根据本公开的一个或多个实施例,所述基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧,还包括:将所述目标视频帧渲染到单个帧缓存对象中;所述采用单个表面视图对所述目标视频帧进行显示,包括:将所述帧缓存对象中缓存的经过渲染的目标视频帧采用单个表面视图进行显示。
第二方面,根据本公开的一个或多个实施例,提供了一种直播处理设备,应用于第一直播连麦客户端,所述装置包括:接收单元,用于接收第二直播连麦客户端的直播流;缓存单元,用于对第二直播连麦客户端的直播流的最 新视频帧进行缓存;在获取所述第一直播连麦客户端的当前视频帧时,读取所缓存的所述最新视频帧;渲染单元,用于基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧;显示单元,用于对所述目标视频帧进行显示。
根据本公开的一个或多个实施例,所述渲染单元在根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧时,用于:根据预设布局信息确定所述当前视频帧对应的第一图层和所述最新视频帧对应的第二图层在所述目标视频帧中的尺寸和位置;将所述当前视频帧绘制到所述第一图层中,将所述最新视频帧分别绘制到第二图层中,将所述第一图层和所述第二图层合成为所述目标视频帧。
根据本公开的一个或多个实施例,所述渲染单元在根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧时,用于:确定片元着色器渲染管线中用于模板测试的目标模板,其中所述目标模板中所述当前视频帧对应区域处的深度值为所述当前视频帧对应的层级值,所述最新视频帧对应区域处的深度值为所述最新视频帧对应的层级值;根据所述当前视频帧的层级值、所述最新视频帧的层级值以及所述目标模板进行模板测试,基于模板测试结果对所述当前视频帧和所述最新视频帧进行合成。
根据本公开的一个或多个实施例,所述渲染单元在根据所述当前视频帧的层级值、所述每一直播流的最新视频帧的层级值以及所述目标模板进行模板测试时,用于:将所述当前视频帧的任意第一位置处的层级值确定为所述第一位置处的实际深度值,将所述当前视频帧的所述第一位置处的实际深度值与所述目标模板中对应位置处的深度值进行比对,若所述第一位置处的实际深度值高于或等于所述目标模板中对应位置处的深度值,则所述第一位置的模板测试通过;将所述最新视频帧的任意第二位置处的层级值确定为所述第二位置处的实际深度值,将所述最新视频帧的所述第二位置处的实际深度值与所述目标模板中对应位置处的深度值进行比对,若所述第二位置处的实际深度值高于或等于所述目标模板中对应位置处的深度值,则所述第二位置的模板测试通过。
根据本公开的一个或多个实施例,所述渲染单元在基于模板测试结果对所述当前视频帧和所述最新视频帧进行合成时,用于:若所述当前视频帧任 意第一位置的模板测试通过,则对所述当前视频帧的所述第一位置进行渲染;若所述最新视频帧任意第二位置的模板测试通过,则对所述最新视频帧的所述第二位置进行渲染;在所述当前视频帧和所述最新视频帧中通过模板测试的位置完成渲染后,得到合成后的所述目标视频帧。
根据本公开的一个或多个实施例,所述最新视频帧的层级值高于所述当前视频帧的层级值。
根据本公开的一个或多个实施例,所述渲染单元在确定片元着色器渲染管线中用于模板测试的目标模板时,用于:将所述目标模板所有位置处的深度值设置为所述当前视频帧的层级值;根据所述最新视频帧的层级值以及所述预设布局信息,将所述目标模板中所述最新视频帧对应区域处的深度值更新为所述最新视频帧的层级值。
根据本公开的一个或多个实施例,所述渲染单元在基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述每一直播流的最新视频帧合成为目标视频帧前,还用于:获取预设列表视图控件中每一视图控件的尺寸和位置,将任一视图控件的尺寸和位置作为所述最新视频帧对应的第二图层在所述目标视频帧中的尺寸和位置,得到所述预设布局信息,以使所述最新视频帧对应的第二图层与所述预设列表视图控件中所述任一视图控件对齐。
根据本公开的一个或多个实施例,所述渲染单元在基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧时,用于:基于单个渲染视图和单个渲染线程,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧。
根据本公开的一个或多个实施例,所述显示单元在对所述目标视频帧进行显示时,用于:采用单个表面视图对所述目标视频帧进行显示。
根据本公开的一个或多个实施例,所述渲染单元在基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧时,还用于:将所述目标视频帧渲染到单个帧缓存对象中;所述显示单元在采用单个表面视图对所述目标视频帧进行显示时,用于:将所述帧缓存对象中缓存的经过渲染的目标视频帧采用单个表面视图进行显示。
第三方面,根据本公开的一个或多个实施例,提供了一种电子设备,包括:至少一个处理器和存储器;所述存储器存储计算机执行指令;所述至少 一个处理器执行所述存储器存储的计算机执行指令,使得所述至少一个处理器执行如上第一方面以及第一方面各种可能的设计所述的直播处理方法。
第四方面,根据本公开的一个或多个实施例,提供了一种计算机可读存储介质,所述计算机可读存储介质中存储有计算机执行指令,当处理器执行所述计算机执行指令时,实现如上第一方面以及第一方面各种可能的设计所述的直播处理方法。
第五方面,根据本公开的一个或多个实施例,提供了一种计算机程序产品,包括计算机执行指令,当处理器执行所述计算机执行指令时,实现如上第一方面以及第一方面各种可能的设计所述的直播处理方法。
以上描述仅为本公开的较佳实施例以及对所运用技术原理的说明。本领域技术人员应当理解,本公开中所涉及的公开范围,并不限于上述技术特征的特定组合而成的技术方案,同时也应涵盖在不脱离上述公开构思的情况下,由上述技术特征或其等同特征进行任意组合而形成的其它技术方案。例如上述特征与本公开中公开的(但不限于)具有类似功能的技术特征进行互相替换而形成的技术方案。
此外,虽然采用特定次序描绘了各操作,但是这不应当理解为要求这些操作以所示出的特定次序或以顺序次序执行来执行。在一定环境下,多任务和并行处理可能是有利的。同样地,虽然在上面论述中包含了若干具体实现细节,但是这些不应当被解释为对本公开的范围的限制。在单独的实施例的上下文中描述的某些特征还可以组合地实现在单个实施例中。相反地,在单个实施例的上下文中描述的各种特征也可以单独地或以任何合适的子组合的方式实现在多个实施例中。
尽管已经采用特定于结构特征和/或方法逻辑动作的语言描述了本主题,但是应当理解所附权利要求书中所限定的主题未必局限于上面描述的特定特征或动作。相反,上面所描述的特定特征和动作仅仅是实现权利要求书的示例形式。

Claims (15)

  1. 一种直播处理方法,应用于第一直播连麦客户端,所述方法包括:
    接收第二直播连麦客户端的直播流,并对第二直播连麦客户端的直播流的最新视频帧进行缓存;
    在获取所述第一直播连麦客户端的当前视频帧时,读取所缓存的所述最新视频帧;
    基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧;
    对所述目标视频帧进行显示。
  2. 根据权利要求1所述的方法,其中,所述根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧,包括:
    根据预设布局信息确定所述当前视频帧对应的第一图层和所述最新视频帧对应的第二图层在所述目标视频帧中的尺寸和位置;
    将所述当前视频帧绘制到所述第一图层中,将所述最新视频帧分别绘制到第二图层中,将所述第一图层和所述第二图层合成为所述目标视频帧。
  3. 根据权利要求1所述的方法,其中,所述根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧,包括:
    确定片元着色器渲染管线中用于模板测试的目标模板,其中所述目标模板中所述当前视频帧对应区域处的深度值为所述当前视频帧的层级值,所述最新视频帧对应区域处的深度值为所述最新视频帧的层级值;
    根据所述当前视频帧的层级值、所述最新视频帧的层级值以及所述目标模板进行模板测试,基于模板测试结果对所述当前视频帧和所述最新视频帧进行合成。
  4. 根据权利要求3所述的方法,其中,所述根据所述当前视频帧的层级值、所述最新视频帧的层级值以及所述目标模板进行模板测试,包括:
    将所述当前视频帧的任意第一位置处的层级值确定为所述当前视频帧的所述第一位置处的实际深度值,将所述第一位置处的实际深度值与所述目标模板中对应位置处的深度值进行比对,若所述第一位置处的实际深度值高于或等于所述目标模板中对应位置处的深度值,则所述第一位置的模板测试通过;
    将所述最新视频帧的任意第二位置处的层级值确定为所述最新视频帧的所述第二位置处的实际深度值,将所述第二位置处的实际深度值与所述目标模板中对应位置处的深度值进行比对,若所述第二位置处的实际深度值高于或等于所述目标模板中对应位置处的深度值,则所述第二位置的模板测试通过。
  5. 根据权利要求4所述的方法,其中,所述基于模板测试结果对所述当前视频帧和所述最新视频帧进行合成,包括:
    若所述当前视频帧任意第一位置的模板测试通过,则对所述当前视频帧的所述第一位置进行渲染;
    若所述最新视频帧任意第二位置的模板测试通过,则对所述最新视频帧的所述第二位置进行渲染;
    在所述当前视频帧和所述最新视频帧中通过模板测试的位置完成渲染后,得到合成后的所述目标视频帧。
  6. 根据权利要求3-5任一项所述的方法,其中,所述最新视频帧的层级值高于所述当前视频帧的层级值。
  7. 根据权利要求6所述的方法,其中,所述确定片元着色器渲染管线中用于模板测试的目标模板,包括:
    将所述目标模板所有位置处的深度值设置为所述当前视频帧的层级值;
    根据所述最新视频帧的层级值以及所述预设布局信息,将所述目标模板中所述最新视频帧对应区域处的深度值更新为所述最新视频帧的层级值。
  8. 根据权利要求1所述的方法,其中,所述基于单个渲染视图根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧前,还包括:
    获取预设列表视图控件中每一视图控件的尺寸和位置,将任一视图控件的尺寸和位置作为所述最新视频帧对应的第二图层在所述目标视频帧中的尺寸和位置,得到所述预设布局信息,以使所述最新视频帧对应的第二图层与所述预设列表视图控件中所述任一视图控件对齐。
  9. 根据权利要求1-5任一项所述的方法,其中,所述基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧,包括:
    基于单个渲染视图和单个渲染线程,根据预设布局信息将所述当前视频 帧和所述最新视频帧合成为目标视频帧。
  10. 根据权利要求9所述的方法,其中,所述对所述目标视频帧进行显示,包括:
    采用单个表面视图对所述目标视频帧进行显示。
  11. 根据权利要求10所述的方法,其中,所述基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧,还包括:
    将所述目标视频帧渲染到单个帧缓存对象中;
    所述采用单个表面视图对所述目标视频帧进行显示,包括:
    将所述帧缓存对象中缓存的经过渲染的目标视频帧采用单个表面视图进行显示。
  12. 一种直播处理装置,应用于第一直播连麦客户端,所述装置包括:
    接收单元,用于接收第二直播连麦客户端的直播流;
    缓存单元,用于对第二直播连麦客户端的直播流的最新视频帧进行缓存;在获取所述第一直播连麦客户端的当前视频帧时,读取所缓存的所述最新视频帧;
    渲染单元,用于基于单个渲染视图,根据预设布局信息将所述当前视频帧和所述最新视频帧合成为目标视频帧;
    显示单元,用于对所述目标视频帧进行显示。
  13. 一种电子设备,包括:至少一个处理器和存储器;
    所述存储器存储计算机执行指令;
    所述至少一个处理器执行所述存储器存储的计算机执行指令,使得所述至少一个处理器执行如权利要求1-11任一项所述的方法。
  14. 一种计算机可读存储介质,所述计算机可读存储介质中存储有计算机执行指令,当处理器执行所述计算机执行指令时,实现如权利要求1-11任一项所述的方法。
  15. 一种计算机程序产品,包括计算机执行指令,当处理器执行所述计算机执行指令时,实现如权利要求1-11任一项所述的方法。
PCT/CN2024/102664 2023-06-28 2024-06-28 直播处理方法、设备及存储介质 Ceased WO2025002428A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202310782876.5 2023-06-28
CN202310782876.5A CN119233011A (zh) 2023-06-28 2023-06-28 直播处理方法、设备及存储介质

Publications (2)

Publication Number Publication Date
WO2025002428A1 true WO2025002428A1 (zh) 2025-01-02
WO2025002428A9 WO2025002428A9 (zh) 2026-01-22

Family

ID=93937709

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2024/102664 Ceased WO2025002428A1 (zh) 2023-06-28 2024-06-28 直播处理方法、设备及存储介质

Country Status (2)

Country Link
CN (1) CN119233011A (zh)
WO (1) WO2025002428A1 (zh)

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113507641A (zh) * 2021-09-09 2021-10-15 山东亚华电子股份有限公司 一种基于客户端的多路视频混屏方法、系统及设备
CN113840170A (zh) * 2020-06-23 2021-12-24 武汉斗鱼网络科技有限公司 连麦直播的方法及装置
WO2023045651A1 (zh) * 2021-09-26 2023-03-30 北京字跳网络技术有限公司 连麦直播方法、装置、电子设备、介质及程序产品
CN116016977A (zh) * 2022-12-29 2023-04-25 广州方硅信息技术有限公司 基于直播的虚拟同台连麦互动方法、计算机设备及介质

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113840170A (zh) * 2020-06-23 2021-12-24 武汉斗鱼网络科技有限公司 连麦直播的方法及装置
CN113507641A (zh) * 2021-09-09 2021-10-15 山东亚华电子股份有限公司 一种基于客户端的多路视频混屏方法、系统及设备
WO2023045651A1 (zh) * 2021-09-26 2023-03-30 北京字跳网络技术有限公司 连麦直播方法、装置、电子设备、介质及程序产品
CN116016977A (zh) * 2022-12-29 2023-04-25 广州方硅信息技术有限公司 基于直播的虚拟同台连麦互动方法、计算机设备及介质

Also Published As

Publication number Publication date
CN119233011A (zh) 2024-12-31
WO2025002428A9 (zh) 2026-01-22

Similar Documents

Publication Publication Date Title
US12572337B2 (en) Method and apparatus of control editing, device, and readable storage medium
US12294758B2 (en) Control setting method and apparatus, electronic device and interaction system
CN112258622B (zh) 图像处理方法、装置、可读介质及电子设备
CN111427528A (zh) 显示方法、装置和电子设备
WO2024131621A1 (zh) 特效生成方法、装置、电子设备及存储介质
US12368902B2 (en) Functional component loading method and data processing method for video live-streaming, and device
CN113535105B (zh) 媒体文件处理方法、装置、设备、可读存储介质及产品
WO2023000805A1 (zh) 视频蒙层显示方法、装置、设备及介质
WO2025002428A1 (zh) 直播处理方法、设备及存储介质
CN113382293A (zh) 内容显示的方法、装置、设备及计算机可读存储介质
WO2025055988A1 (zh) 视频显示方法、装置、介质及电子设备
WO2024198952A1 (zh) 图像超分辨率方法、设备、存储介质及程序产品
CN111199569A (zh) 数据处理的方法、装置、电子设备及计算机可读介质
CN117788669A (zh) 图像处理方法、装置、终端和存储介质
CN118741242A (zh) 视频编辑方法、装置、设备及介质
CN116028740A (zh) 页面显示方法、装置、存储介质和电子设备
WO2024188322A1 (zh) 直播界面的处理方法、装置、设备及存储介质
WO2021004171A1 (zh) 水波纹图像实现方法及装置
WO2025044969A1 (zh) 页面切换方法、装置、设备及存储介质
WO2025194828A1 (zh) 多媒体素材显示方法、装置、设备及存储介质
CN115705134A (zh) 图像处理方法、装置及设备
WO2025201506A1 (zh) 图片加载方法、设备及存储介质
WO2025139936A1 (zh) 直播间界面处理方法、设备及存储介质
WO2024255894A1 (zh) 界面切换方法、装置、电子设备及存储介质
CN118286682A (zh) 图像处理方法、装置、终端和存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24831052

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE