WO2020228482A1 - 视频处理方法、装置及系统 - Google Patents
视频处理方法、装置及系统 Download PDFInfo
- Publication number
- WO2020228482A1 WO2020228482A1 PCT/CN2020/085263 CN2020085263W WO2020228482A1 WO 2020228482 A1 WO2020228482 A1 WO 2020228482A1 CN 2020085263 W CN2020085263 W CN 2020085263W WO 2020228482 A1 WO2020228482 A1 WO 2020228482A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- image
- sub
- code stream
- encoded data
- terminal
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/83—Generation or processing of protective or descriptive data associated with content; Content structuring
- H04N21/845—Structuring of content, e.g. decomposing content into time segments
- H04N21/8456—Structuring of content, e.g. decomposing content into time segments by decomposing the content in the time domain, e.g. in time segments
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/21—Server components or server architectures
- H04N21/218—Source of audio or video content, e.g. local disk arrays
- H04N21/2187—Live feed
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/234—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
- H04N21/2343—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving reformatting operations of video signals for distribution or compliance with end-user requests or end-user device requirements
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/25—Management operations performed by the server for facilitating the content distribution or administrating data related to end-users or client devices, e.g. end-user or client device authentication, learning user preferences for recommending movies
- H04N21/266—Channel or content management, e.g. generation and management of keys and entitlement messages in a conditional access system, merging a VOD unicast channel into a multicast channel
- H04N21/2662—Controlling the complexity of the video stream, e.g. by scaling the resolution or bitrate of the video stream based on the client capabilities
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/44—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/44—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
- H04N21/44016—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving splicing one content stream with another content stream, e.g. for substituting a video clip
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/44—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
- H04N21/4402—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving reformatting operations of video signals for household redistribution, storage or real-time display
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/47—End-user applications
- H04N21/472—End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/47—End-user applications
- H04N21/472—End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content
- H04N21/47202—End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content for requesting content on demand, e.g. video on demand
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/60—Network structure or processes for video distribution between server and client or between remote clients; Control signalling between clients, server and network components; Transmission of management data between server and client, e.g. sending from server to client commands for recording incoming content stream; Communication details between server and client
- H04N21/65—Transmission of management data between client and server
- H04N21/658—Transmission by the client directed to the server
- H04N21/6587—Control parameters, e.g. trick play commands, viewpoint selection
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/83—Generation or processing of protective or descriptive data associated with content; Content structuring
- H04N21/845—Structuring of content, e.g. decomposing content into time segments
Definitions
- the embodiments of the present application relate to the field of video processing, and in particular to a video processing method, device, and system.
- the terminal when the terminal is playing video, if the viewing angle is switched, it may happen that the switching time does not reach the random entry point corresponding to the tile in the new viewing angle. Therefore, the terminal cannot be When Tile decodes, it can only temporarily play the lower-quality data layer or a black screen appears, causing a problem of poor user experience.
- the present application provides a video processing method, device, and system, which can avoid the low data layer that appears during the viewing angle switching process to a certain extent, which causes the problem of poor user experience.
- an embodiment of the present application provides a video processing method, the method includes: a terminal (or a client, or a decoder) decodes the encoded data of a first sub-image covered by a first view, thereby obtaining the first sub-image . Subsequently, when the terminal switches from the first view to the second view, it acquires the encoded data of the second sub-image covered by the second view and the encoded data of the first reference image of the second sub-image, where the encoding of the first reference image The data is independent of the code stream where the encoded data of the second sub-image is located. Next, the terminal decodes the encoded data of the first reference image to obtain the first reference image. And, the terminal may decode the encoded data of the second sub-image according to the first reference image, thereby obtaining the second sub-image.
- the server side prepares two independent code streams in advance, including the code stream containing the coded data of the first reference image and the code stream containing the coded data of the second sub-image (and/or the first sub-image).
- Code stream so that after the viewing angle is switched, the terminal can decode the second sub-image covered by the switched viewing angle based on the first reference image to obtain the second sub-image, so that the terminal can realize seamless switching during the viewing angle switching process ,
- the display screen of the terminal is always maintained at the high-quality layer to which the second sub-image belongs, without the low-quality layer or black screen appearing in the viewing angle coverage, effectively improving the user experience.
- the code stream where the coded data of the second sub-image is located is the segment code stream where the coded data of the second sub-image is located.
- the code stream may be a segment code stream, that is, at least one segment code after the server encapsulates the code
- the stream is sent to the terminal for the terminal to decode and display.
- the second sub-image is an image other than the first image in the decoding order in the image sequence described in the code stream where the encoded data of the second sub-image is located, or the second sub-image The first picture and the pictures outside the second picture in the decoding order in the picture sequence described by the code stream where the coded data of the second sub picture is located.
- the switching time is the playback time corresponding to the second image in the decoding order in the image sequence described by the code stream where the encoded data of the second sub-image is located, correspondingly, the first reference image and The image described by the reference image (ie, the image frame at the first position of the segment stream) on which the second sub-image depends is consistent.
- the switching time is the playback time corresponding to the image except the first image and the second image, for example: the third image or the fourth image, correspondingly, the first reference image can be the same as the third image.
- the image described by the reference image on which the image or the fourth image depends is the same.
- the image content of the first reference image is the same as the image content of the second reference image of the second sub-image, and the second reference image is described by the code stream in which the encoded data of the second sub-image is located.
- the images in the image sequence are the same as the image content of the second reference image of the second sub-image, and the second reference image is described by the code stream in which the encoded data of the second sub-image is located.
- the server side encodes the first reference image and the second sub-image based on the same image content when encoding.
- the encoding rate during encoding may be the same or different.
- the second reference image is an image other than the first image in the decoding order in the image sequence described by the code stream in which the encoded data of the second sub-image is located.
- the code stream obtained by the terminal is switched from the code stream where the coded data of the first sub-image is located, for example, the code stream where the coded data of the first sub-image is located, to the second sub-image.
- the code stream where the coded data of the second sub-image is located for example, when the coded data of the second sub-image is located, the image content of the first reference image and the image content of the reference image (ie, the second reference image) on which the second sub-image depends It is consistent, where the second reference image may be an image other than the first image in the decoding order in the image sequence corresponding to the segment code stream where the encoded data of the second sub-image is located.
- the second reference image is the previous image of the second sub-image in the image sequence.
- the image content of the first reference image can be consistent with the image content of the previous image (ie, the second reference image) of the second sub-image in the image sequence.
- the code stream in which the encoded data of the second sub-image is located is a code stream obtained by performing inter-frame prediction on all image frames described in the code stream.
- the prediction mode of all image frames in the code stream where the encoded data of the second sub-image is located can be inter-frame prediction.
- all the image frames in the segment stream where it is located can be P frames.
- the code stream in which the encoded data of the first reference image is located is a code stream obtained by performing intra-frame prediction on all image blocks of all image frames described in the code stream, and the encoding of the first reference image
- the code stream where the data is located is independent of the code stream where the encoded data of the second sub-image is located.
- the prediction mode of all image frames in the code stream where the encoded data of the first reference image is located can be intra prediction.
- all image frames in each segment stream are I frames.
- the first reference image is a pure random access CRA image.
- the prediction mode of all image frames in the code stream where the encoded data of the first reference image is located can be intra prediction, and all the image frames can be CRA type I frames.
- the second sub-image is not covered by the first viewing angle.
- the second sub-image acquired by the terminal may not be covered by the first viewing angle, that is, only displayed in the second viewing angle, that is, this application can be applied to the new sub-image (that is, only covering It is the decoding process of sub-images that are not covered by the view before switching).
- the second perspective covers the area of the first sub-image
- the method further includes: the terminal decodes the next frame of sub-image in the code stream where the encoded data of the first sub-image is located, so as to obtain the next Frame sub-images; and, the terminal splices the next frame image with the first reference image to obtain a spliced image; and plays the spliced image.
- the switching time is the previous image in the image sequence where the second sub-image is located
- the image content of the first reference image acquired by the terminal is consistent with the image content of the previous image
- the second reference image is used to decode the second reference image.
- the sub-image is decoded, and the first reference image and other subsequent sub-images including the second sub-image and the code stream in which it is located are displayed in the coverage of the second view.
- the second perspective covers the area of the first sub-image
- the method further includes: decoding the next frame of sub-image in the code stream where the encoded data of the first sub-image is located, so as to obtain the next frame Sub-image; splicing the next frame image with the second sub-image to obtain the spliced image; playing the spliced image.
- the terminal obtains The image content of the received first reference image is consistent with the image content of the reference image on which the second sub-image depends (for example, the second reference image in this application), and the terminal decodes the first reference image based on the A reference image decodes the second sub-image, and the coverage of the second view is displayed in the second sub-image and other subsequent sub-images including the second sub-image and the code stream in which it is located.
- the playback time corresponding to the first sub-image is t 1
- the playback time corresponding to the second sub-image is t 2
- time t 1 and time t 2 are adjacent playback times
- the sub-image and the second sub-image differ by N frames in the playback order, and N is 1, 2, 3, or 4.
- the switching moment described in this application can be: Sub-image (It should be noted that the code stream where the first sub-image is located is different from the code stream where the second sub-image is located, or it can be understood that the segment code stream displayed by the terminal is switched from one segment code stream when the viewing angle is switched to another segment stream) switching the playback can be started at time t 1, and the playback time of the second sub-image displayed after switching the Perspective of t 2, wherein, at time t 1 and time t 2 adjacent to playback time.
- the first sub-image and the second sub-image may differ by N frames in the playback order, and N is 1, 2, 3, or 4.
- the playback time corresponding to the first sub-image is t 1
- the playback time corresponding to the first reference image is t 2
- time t 1 and time t 2 are adjacent playback times
- the sub-image and the first reference image differ by N frames in the playback order, and N is 1, 2, 3, or 4.
- the switching moment described in this application can be that the terminal starts from the first sub-image being played (It should be noted that the code stream where the first sub-image is located is not the same as the code stream where the second sub-image is located, or it can be understood that the segment code stream displayed by the terminal switches from one segment code stream to another when the viewing angle is switched.
- One segment stream can be switched at time t 1 during playback, and the playback time of the first reference image displayed in the switched viewing angle is t 2 , where time t 1 and time t 2 are adjacent playback times.
- the first sub-image and the first reference image differ by N frames in the playback order, and N is 1, 2, 3, or 4.
- the second perspective covers the first sub-image area
- the first perspective also covers the third sub-image area
- the encoded data of the first sub-image covered by the first perspective is decoded to obtain the first sub-image Including: combining the first code stream and the third code stream to obtain the first combined code stream, the first code stream is the code stream where the coded data of the first sub-image is located, and the third code stream is the code of the third sub-image
- the code stream where the data is located; decoding the first combined code stream to obtain an image including the first sub-image and the third sub-image; decoding the encoded data of the second sub-image according to the first reference image to obtain the second sub-image includes: Combine the first code stream and the second code stream to obtain a second combined code stream.
- the second code stream is the code stream where the coded data of the second sub-image is located.
- the sub-image described in the first code stream corresponds to the first code stream.
- the position in the image described by the combined code stream is consistent with the position in the image described by the second combined code stream corresponding to the sub-image described in the first code stream; the first combined code stream is decoded according to the first reference image to obtain the The second sub-image and the image of the sub-image described in the first code stream.
- the decoding process for the first sub-image existing in the coverage of the first viewing angle and the second viewing angle and the second sub-image only covering the second viewing angle is realized due to the first
- the reference image from the image displayed by the sub-image stream at the current playback time is buffered in the terminal. Therefore, when the terminal decodes the code stream (ie, the first code stream), the code stream is combined with the code stream ( That is, the position in the image described by the second combined code stream can be the same as the position of the image described by the first combined code stream, thereby avoiding this type of code (for example, the segment in the overlapping tile involved in the embodiment of this application).
- the decoding failed during the decoding process of the stream.
- an embodiment of the present application provides a video processing method.
- the method includes: a server (or an encoding end) sends to the terminal encoding data of a first sub-image covered by a first view of the terminal; When the angle of view is switched to the second angle of view, the coded data of the second sub-image covered by the second angle of view and the coded data of the first reference image of the second sub-image are sent to the terminal.
- the coded data of the first reference image is independent of the second sub-image.
- the code stream where the encoded data of the second sub-image is located is the segment code stream where the encoded data of the second sub-image is located.
- the second sub-image is an image other than the first image in the decoding order in the image sequence described in the code stream where the encoded data of the second sub-image is located, or the second sub-image The first picture and the pictures outside the second picture in the decoding order in the picture sequence described by the code stream where the coded data of the second sub picture is located.
- the image content of the first reference image is the same as the image content of the second reference image of the second sub-image, and the second reference image is described by the code stream in which the encoded data of the second sub-image is located.
- the images in the image sequence are the same as the image content of the second reference image of the second sub-image, and the second reference image is described by the code stream in which the encoded data of the second sub-image is located.
- the second reference image is an image other than the first image in the decoding order in the image sequence described by the code stream in which the encoded data of the second sub-image is located.
- the second reference image is the previous image of the second sub-image in the image sequence.
- the code stream in which the encoded data of the second sub-image is located is a code stream obtained by performing inter-frame prediction on all image frames described in the code stream.
- the code stream in which the encoded data of the first reference image is located is a code stream obtained by performing intra-frame prediction on all image blocks of all image frames described in the code stream, and the encoding of the first reference image
- the code stream where the data is located is independent of the code stream where the encoded data of the second sub-image is located.
- the first reference image is a pure random access CRA image.
- the second sub-image is not covered by the first viewing angle.
- an embodiment of the present application provides a video processing device.
- the device may include a decoding module and an acquisition module, where the decoding module may be used to decode the encoded data of the first sub-image covered by the first view to obtain the first sub-image.
- a sub-image; the acquisition module can be used to acquire the encoded data of the second sub-image covered by the second perspective and the encoded data of the first reference image of the second sub-image when the first perspective is switched to the second perspective.
- the coded data of the image is independent of the code stream where the coded data of the second sub-image is located; and the decoding module may be further used to decode the coded data of the first reference image to obtain the first reference image; the decoding module may also be used to obtain the first reference image according to the first reference image.
- the image decodes the encoded data of the second sub-image, thereby obtaining the second sub-image.
- the code stream where the coded data of the second sub-image is located is the segment code stream where the coded data of the second sub-image is located.
- the second sub-image is an image other than the first image in the decoding order in the image sequence described in the code stream where the encoded data of the second sub-image is located, or the second sub-image The first picture and the pictures outside the second picture in the decoding order in the picture sequence described by the code stream where the coded data of the second sub picture is located.
- the image content of the first reference image is the same as the image content of the second reference image of the second sub-image, and the second reference image is described by the code stream in which the encoded data of the second sub-image is located.
- the images in the image sequence are the same as the image content of the second reference image of the second sub-image, and the second reference image is described by the code stream in which the encoded data of the second sub-image is located.
- the second reference image is an image other than the first image in the decoding order in the image sequence described by the code stream in which the encoded data of the second sub-image is located.
- the second reference image is the previous image of the second sub-image in the image sequence.
- the code stream in which the encoded data of the second sub-image is located is a code stream obtained by performing inter-frame prediction on all image frames described in the code stream.
- the code stream in which the encoded data of the first reference image is located is a code stream obtained by performing intra-frame prediction on all image blocks of all image frames described in the code stream, and the encoding of the first reference image
- the code stream where the data is located is independent of the code stream where the encoded data of the second sub-image is located.
- the first reference image is a pure random access CRA image.
- the second sub-image is not covered by the first viewing angle.
- the second perspective covers the first sub-image area
- the first perspective also covers the third sub-image area
- the decoding module can also be used to combine the first code stream and the third code stream. Merge to obtain the first combined code stream, the first code stream is the code stream where the encoded data of the first sub-image is located, and the third code stream is the code stream where the encoded data of the third sub-image is located; decode the first combined code stream , So as to obtain an image including the first sub-image and the third sub-image; and, the decoding module may also be used to combine the first code stream with the second code stream to obtain a second combined code stream, the second code stream being the first The code stream where the encoded data of the two sub-images are located, the sub-image described in the first code stream corresponds to the position in the image described in the first combined code stream, and the sub-image described in the first code stream corresponds to the second combined code The positions in the images described by the stream are consistent; the first combined code stream is
- an embodiment of the present application provides a video processing device.
- the device includes: a sending module, which can be used to send to the terminal the encoded data of the first sub-image covered by the first view of the terminal; and, it can also be used for When the first view of the terminal is switched to the second view, the terminal sends the encoded data of the second sub-image covered by the second view, and the encoded data of the first reference image of the second sub-image, and the encoded data of the first reference image It is independent of the code stream of the second sub-image.
- the code stream where the encoded data of the second sub-image is located is the segment code stream where the encoded data of the second sub-image is located.
- the second sub-image is an image other than the first image in the decoding order in the image sequence described in the code stream where the encoded data of the second sub-image is located, or the second sub-image The first picture and the pictures outside the second picture in the decoding order in the picture sequence described by the code stream where the coded data of the second sub picture is located.
- the image content of the first reference image is the same as the image content of the second reference image of the second sub-image, and the second reference image is described by the code stream in which the encoded data of the second sub-image is located.
- the images in the image sequence are the same as the image content of the second reference image of the second sub-image, and the second reference image is described by the code stream in which the encoded data of the second sub-image is located.
- the second reference image is an image other than the first image in the decoding order in the image sequence described by the code stream in which the encoded data of the second sub-image is located.
- the second reference image is the previous image of the second sub-image in the image sequence.
- the code stream in which the encoded data of the second sub-image is located is a code stream obtained by performing inter-frame prediction on all image frames described in the code stream.
- the code stream in which the encoded data of the first reference image is located is a code stream obtained by performing intra-frame prediction on all image blocks of all image frames described in the code stream, and the encoding of the first reference image
- the code stream where the data is located is independent of the code stream where the encoded data of the second sub-image is located.
- the first reference image is a pure random access CRA image.
- the second sub-image is not covered by the first viewing angle.
- an embodiment of the present application provides a video processing system, including a server and a terminal, wherein the server sends to the terminal the encoded data of the first sub-image covered by the first viewing angle of the terminal; the terminal receives the first sub-image And decode the encoded data of the first sub-image to obtain the first sub-image; when the first view of the terminal is switched to the second view, the server sends the encoded data of the second sub-image covered by the second view to the terminal , And the encoded data of the first reference image of the second sub-image, the encoded data of the first reference image is independent of the code stream where the encoded data of the second sub-image is located; the terminal receives the encoded data of the second sub-image and the first reference image And decode the encoded data of the first reference image to obtain the first reference image; and the terminal decodes the encoded data of the second sub-image according to the first reference image to obtain the second sub-image.
- the embodiments of the present application provide a computer-readable medium for storing a computer program, the computer program including instructions for executing the first aspect or any possible implementation of the first aspect.
- an embodiment of the present application provides a computer-readable medium for storing a computer program, and the computer program includes instructions for executing the second aspect or any possible implementation of the second aspect.
- embodiments of the present application provide a computer program, which includes instructions for executing the first aspect or any possible implementation of the first aspect.
- an embodiment of the present application provides a computer program, which includes instructions for executing the second aspect or any possible implementation of the second aspect.
- an embodiment of the present application provides a chip, which includes a processing circuit and transceiver pins.
- the transceiver pin and the processing circuit communicate with each other through an internal connection path, and the processor executes the method in the first aspect or any one of the possible implementations of the first aspect to control the receiving pin to receive signals, and Control the sending pin to send signals.
- an embodiment of the present application provides a chip, which includes a processing circuit and transceiver pins.
- the transceiver pin and the processing circuit communicate with each other through an internal connection path, and the processor executes the method in the second aspect or any one of the possible implementations of the second aspect to control the receiving pin to receive signals, and Control the sending pin to send signals.
- Fig. 1 is a schematic diagram showing a transmission system of a video processing method exemplarily
- Fig. 2 is a schematic diagram of a transmission process exemplarily shown
- Fig. 3 is a schematic flowchart of a video processing method shown by way of example
- FIG. 4 is a schematic structural diagram of a video processing system provided by an embodiment of the present application.
- FIG. 5 is a schematic flowchart of a video processing method provided by an embodiment of the present application.
- Fig. 6 is a schematic diagram of a handover process provided by an embodiment of the present application.
- FIG. 7 is a schematic diagram of the function mode of the reference frame provided by an embodiment of the present application.
- FIG. 8 is a schematic flowchart of a decoding method provided by an embodiment of the present application.
- FIG. 9 is a schematic flowchart of a video processing method provided by an embodiment of the present application.
- FIG. 10 is a schematic structural diagram of a video processing device provided by an embodiment of the present application.
- FIG. 11 is a schematic structural diagram of a video processing device provided by an embodiment of the present application.
- FIG. 12 is a schematic structural diagram of a video processing device provided by an embodiment of the present application.
- first and second in the description and claims of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order of objects.
- first target object and the second target object are used to distinguish different target objects, rather than to describe the specific order of the target objects.
- words such as “exemplary” or “for example” are used as examples, illustrations, or illustrations. Any embodiment or design solution described as “exemplary” or “for example” in the embodiments of the present application should not be construed as being more preferable or advantageous than other embodiments or design solutions. To be precise, words such as “exemplary” or “for example” are used to present related concepts in a specific manner.
- multiple means two or more.
- multiple processing units refer to two or more processing units; multiple systems refer to two or more systems.
- panoramic video services are gradually launched on various platforms and become an important part of network video traffic.
- the panoramic video service has the characteristics of ultra-high resolution and ultra-high bit rate.
- a dynamic adaptive streaming over HTTP Dynamic Adaptive Streaming over HTTP
- HTTP HyperText Transfer Protocol
- the server can be used to provide a code stream generated based on panoramic video, where the code stream includes but is not limited to: an enhancement layer, or can be called a high-quality layer or a high-quality code stream, and a base layer, or It is called low-quality layer or low-quality code stream.
- the low-quality layer refers to the encoding of panoramic video at a lower bit rate or smaller resolution
- the high-quality layer refers to the encoding of the panoramic video at a higher bit rate or larger resolution.
- the high-quality layer includes multiple tiles divided by the server according to different areas of the panoramic video content.
- the server can transmit the code stream (including the low-quality layer and the high-quality layer) to the terminal through technologies such as DASH or HTTP-based live streaming ((HTTP Live Streaming, HLS)).
- the terminal After the terminal obtains the low-quality layer and the high-quality layer, it merges multiple tiles in the high-quality layer, decodes the merged Tile, and the terminal decodes the low-quality layer. Then, the terminal displays the high-quality layer and the low-quality layer, as shown in Figure 1, that is, the low-quality layer is the data layer covering all the panoramic video, and the high-quality layer can be the terminal’s perspective (including those within and beyond the perspective). Specified range part) video data part. Thus, a clearer, high-quality video image can be presented in the user's perspective.
- the terminal will continue to request the low-quality layer during use to ensure that when the terminal's perspective is rotated, the data frame corresponding to the video data of the high-quality layer will not be decoded successfully. There are areas with no screen or black screen.
- the server divides the panoramic video into multiple tiles.
- the server is further used to divide each Tile into multiple segments according to time, and the starting position of each segment is a random access point (RAP), such as an I frame.
- RAP random access point
- the terminal only needs to request the high-quality tiles and low-quality layers in the user's perspective (including the part within the perspective that is outside the specified range) of the high-quality tiles and low-quality layers, which can reduce bandwidth requirements and consumption while ensuring high-quality images in the user's perspective .
- the basic transmission unit of the tile is the segment, and under different application requirements, the segments have different time lengths, such as 1 second or 3 seconds.
- the starting position of each segment is usually coded as an I frame, which is a random access point for segment switching, that is to say, the switching of tiles can only be done in RAP (the reason is that the frames in each segment need to be based on the previous One or more frames can be used as reference frames for decoding, and the terminal needs to use the I frame corresponding to the random access point of each segment as a reference to continue decoding the subsequent frames).
- I frame which is a random access point for segment switching
- the terminal can only arrive at the next random access point. Only then can Tile n (ie Seg t+1 as shown in Figure 2) be decoded, and a new high-quality tile (ie Tile n) can be displayed. That is, as shown in Figure 2, when the perspective switch occurs in Seg t, the new high-quality tile can only be accessed until Seg t+1. During this period, the user will watch the low-quality data layer or can be called low-quality data layer. Quality video picture.
- the terminal since the terminal can only switch to the content of the high-quality layer when the random access point arrives, even if the terminal switches the perspective at the time corresponding to the first P frame after the I frame, Therefore, the terminal needs to wait for a time period corresponding to one segment to switch to high-quality content, which seriously affects the user experience.
- the prior art proposes a feasible implementation scheme. Specifically, the prior art divides each tile frame by frame shift according to the size of the group of pictures (GOP) on the server side, and Encode according to the divided Tile, as shown in Figure 3. Therefore, when the user's perspective changes, the terminal can immediately request that the next frame is an I-frame segment at the moment of perspective switching and decode it, thereby reducing the length of the segment to reduce the waiting time for switching and realizing the purpose of switching high-quality tiles in real time .
- GOP group of pictures
- the prior art has added multiple backup segments, that is, a large amount of backup data is generated on the server side, thereby increasing the overhead of transcoder and storage on the server side, and reducing the video encoding efficiency of the server, and increasing The amount of data transmitted over the network.
- the server will send a whole piece of new data, and when the bandwidth is not ideal, it still cannot make the video displayed in the terminal meet the expectations.
- the switching moment of the user's perspective usually occurs at a certain position in the segment.
- the original tile in the perspective that is, the tile that appears in the old perspective and the new perspective, which can also be called overlapped tile
- the first frame where new tiles participate in fusion is the I frame.
- performing tile fusion according to the position of the tile in the new view will cause the problem that the P frame of the existing tile loses its reference frame and cannot be decoded due to the position change.
- this application proposes a video transmission method, which aims to realize real-time switching of high-quality video data.
- FIG. 4 is a schematic diagram of a video transmission system provided by an embodiment of this application.
- the video transmission system includes a server 100 and a terminal 200.
- the terminal 200 may be a computer, a smart phone, a VR or other equipment.
- the number of servers and terminals may both be one or more.
- the number of servers and terminals in the video transmission system shown in FIG. 4 is only an example of adaptability, which is not limited in this application.
- FIG. 5 is a schematic flowchart of a video processing method in an embodiment of the application, and in FIG. 5:
- Step 101 The terminal obtains the encoded data of the sub-image covered by the first view from the server.
- the terminal sends a request message generated based on a user instruction to the server.
- the request message can be used to request the server for the video data to be displayed by the terminal, and can also be called the encoded data corresponding to the video screen (or called the video image, or image content, etc.).
- the server encodes and encapsulates the image content into bitstreams with different coding rates, including but not limited to: high-quality bitstream (or called high-quality layer) and low-quality bitstream (or called Is a low-quality layer).
- high-quality layer the server divides the image content into multiple sub-images in the spatial dimension.
- the sub-images may be tiles.
- the server encodes the image content corresponding to each Tile to obtain the encoded data corresponding to each Tile.
- the terminal may request the server for the encoding data of at least one tile covered by the current view of the terminal (for example, the first view in the embodiment of this application).
- the tiles covered by the viewing angle include all tiles within the viewing angle range or partially within the viewing angle range.
- the basic transmission unit of Tile in the high-quality layer (specifically refers to the transmission of the encoded data of the Tile) is the segment, that is, the basic transmission unit when the terminal interacts with the server is the segment, or it can be understood as this
- the code streams involved in the application (including the code stream generated by the server, the code stream transmitted by the server to the terminal, and the code stream received by the terminal) all refer to the segment code stream.
- the terminal may send to the server the identification information carrying the required segments in the tile, so as to request the server for the encoding data of each image in the target segment code stream in the corresponding tile.
- the identification information can be an identification method such as a frame number or a time stamp.
- the identification information may also be the storage location of the segment, that is, the storage location of the segment may be used to uniquely identify the segment.
- the requested information can be the "absolute path of the stream segment in the server" (for example, http://192.168.0.2/dash/Tile0/4.m4s (where 4.m4s represents a segment)).
- the identification information can also be the starting position and size of the segment, that is, the starting position and size can be used to uniquely identify the segment.
- the requested segment is in the 10000th byte of the file 1.mp4, and the size is 2000 bytes.
- the terminal may continue to send the request to the server to request the encoded data of other segments in the tile.
- the video data (also referred to as video images or video images) required by the above terminal can be pre-downloaded video data in the on-demand scene, or it can also be currently displayed by the terminal in the live scene (including the viewing angle). Inside and outside the viewing angle) video data.
- the terminal can request video data from the server, and the requested video data can be the predicted video content to be displayed by the terminal, that is, the pre-loading function is realized, and the decoding efficiency of the terminal is improved through pre-downloading.
- the server prepares in advance a high-quality code stream (also called a high-quality layer) and a low-quality code stream (also called a low-quality layer) corresponding to the video data.
- the terminal can request the server for the high-quality layer within the predicted view range (including the part within and outside the specified range), and the low-quality layer corresponding to the video data to be displayed by the terminal requesting prediction (such as As mentioned above, the low-quality layer is the encoded data corresponding to the panoramic video).
- the terminal can request video data from the server in real time, and the requested video data is the video data (or video screen) that the terminal currently wants to display.
- the server encodes the corresponding video data based on the video data requested by the terminal and sends it to the terminal.
- the terminal may request the server for the high-quality layer within the viewing angle (including the part within and outside the specified range), and the low-quality layer corresponding to the video data to be displayed by the terminal (such as As mentioned above, the low-quality layer is the encoded data corresponding to the panoramic video).
- the terminal requests the server for required video data (for example: encoded data corresponding to the Tile covered by the first view) by detecting the size of the view angle.
- required video data for example: encoded data corresponding to the Tile covered by the first view
- the shape of the viewing angle can be circular or rectangular, and the size can also be different. Therefore, the high-quality layers required for viewing angles of different sizes are not the same. For example: a larger viewing angle requires 10 tiles in a high-quality layer, while a relatively small viewing angle may only require 6 corresponding tiles.
- the server after receiving the request sent by the terminal, the server responds to the request and sends the encoded data of the tile required by the terminal to the terminal.
- the server encodes and encapsulates the image content into bit streams with different coding rates.
- the high-quality bit stream generated by the server is: the panoramic video is divided into multiple rectangular tiles in the spatial dimension, Then, each Tile is divided into multiple segments in the time dimension, and the duration of the segment can be 1 second (which can be divided according to actual needs).
- the spatial size of each tile can be the same or different, and the duration of each segment in the tile can also be the same or different.
- Encode each segment of each Tile, where the first image of each segment in the encoded data of each Tile obtained after encoding (it should be noted that the first image can also be located at the head of each segment)
- the prediction mode of the frame) may be intra-frame prediction.
- the first frame of each segment may be an I frame
- the other frames may be P frames (that is, frames decoded by using inter-frame prediction).
- the server may not refer to the time-domain motion vector when encoding the P frames in the segment, that is, in the high-quality code stream, except for the frame at the first position of each segment (for example, I frame), the subsequent The frame only refers to the inserted frame, that is, it depends on the frame except the reference position and does not depend on the preamble frame.
- the server will pre-encode the Tile to obtain the encoded data of the Tile, and encapsulate the encoded data into a code stream (or data stream).
- the server determines the coded data corresponding to the tile required by the terminal, and sends the coded data of the tile to the terminal.
- the server determines the tile required by the terminal based on the request sent by the terminal, and encodes the tile required by the terminal to obtain the corresponding encoded data, and The encoded data of the Tile is sent to the terminal.
- Step 102 The terminal decodes the acquired encoded data of the sub-image.
- the terminal decodes the received high-quality code stream (ie, the encoded data of the Tile) and the low-quality code stream separately (the decoding and display process of the low-quality layer can refer to the prior art embodiment. I will not repeat it here), and display the video picture composed of the base layer (ie, the low-quality layer) and the high-quality layer after decoding.
- the terminal decodes and merges the encoded data to obtain the Tile (or it may be called Tile). The image sequence described).
- Step 103 When the first perspective is switched to the second perspective, the terminal obtains the encoded data of the sub-image covered by the second perspective and the encoded data of the corresponding reference image from the server.
- the video data (ie Tile) required by the terminal will be updated, and the terminal sends a request message to the server to request the encoded data of the Tile covered by the new viewing angle.
- a first terminal displays the viewing angle is switched to the second perspective, then, the encoded data Tile covering a second request to the server terminal perspective, wherein the second viewing angle covering Tile (Example embodiment of the present application means the second sub-image) playback time for the time t 2, and, t 1 and t 2 of the adjacent playing time.
- the tiles covered by the second perspective that is, the new perspective after switching
- the tiles covered by the new perspective include all the tiles within the scope of the new perspective or part of the tiles within the scope of the new perspective
- Tiles can be divided into two types: one is a new tile (that is, it does not appear in the old view, but only in the coverage of the new view), and the other is an overlapping tile (hereinafter referred to as overlapping tile, that is, includes Tile within the coverage of the old view and the coverage of the new view).
- the terminal stops sending the request message corresponding to the old tile (excluding the overlapping tile) in the old perspective to the server, but the terminal still sends the request message corresponding to the overlapping tile to the server , And cache the overlapping tiles locally.
- the terminal also sends a request message carrying the identification information of the new Tile to the server.
- the terminal may also send a request message carrying identification information of the reference frame corresponding to the new Tile to the server. That is, in this application, after the viewing angle is switched, the terminal no longer obtains the coded data of the old tile, but only obtains the coded data of the new tile and the coded data of the overlapped tile.
- the high-quality layers in view 1 include Tile1 and Tile2.
- the terminal requests the server for the coded data of Tile1 and the coded data of Tile2.
- the viewing angle displayed by the terminal is switched to viewing angle 2, and the high-quality layers in viewing angle 2 include Tile2 and Tile3.
- Tile2 is the overlapping Tile mentioned above.
- the terminal stops the request for Tile1 (referring to the request for the coded data of Tile1), and stops the decoding operation for the coded data of Tile1.
- the terminal continues to request the encoded data of Tile2 and caches the encoded data previously requested.
- a terminal obtains from the server to Tile3 (Tile3 terminal is acquired from the subsequent switching time t 1 corresponding to the time point P frame and the subsequent Segmenet later (i.e., t 2 and time t 2 the image corresponding to the time Corresponding image), as shown by the arrow in FIG. 6) and the encoded data of the reference image corresponding to the encoded data of Tile3 (the generation and transmission of the encoded data of the reference image will be described in detail in the following embodiments).
- the image corresponding to time t 2 in the new Tile may be the image sequence where the Tile is located.
- the image other than the first image in the GOP to which it belongs, or can be understood as the image corresponding to time t 2 in the new Tile (for example, the second sub-image in the embodiment of this application) can be in its code stream (or It can be understood as an image other than the first image in the decoding order in the image sequence described by the segment).
- the image corresponding to time t 2 in the new Tile can also be The first picture and the pictures outside the second picture in the decoding order in the picture sequence described by the code stream.
- the image corresponding to time t 2 may be an image corresponding to frames other than the first frame in the segment to which it belongs.
- the image corresponding to time t 2 in the new Tile may also be an image other than the first image and the second image in the GOP described in the image sequence where the Tile is located, that is, referring to FIG. 6, the image in the new perspective
- the starting frame (ie, the solid black frame) played by Tile3 in the new perspective can be any frame other than the original first frame and the second frame of the segment where it is located.
- the starting frame at the moment of viewing angle playback may be data frame 2 or data frame 3 or data frame 4, which is not limited in this application.
- the server receives the request from the terminal, and sends the corresponding tile coded data and reference image coded data to the terminal.
- the terminal may not be aware of whether the viewing angle is switched, that is, the request message sent by the terminal to the server may carry identification information of the required tile.
- the server side determines whether the terminal is currently switched from perspective based on the identification information carried in the request message. If a perspective switch occurs, the server sends to the terminal the tiles covered by the new perspective (including new tiles and overlapping tiles, The definitions of new tiles and overlapping tiles are as described above, and will not be repeated here).
- the terminal can detect whether the perspective is switched. If the perspective is switched, the request message sent by the terminal to the server can carry the identification information of the tile covered by the new perspective and the part or part covered by the new perspective. Identification information of reference images corresponding to all tiles.
- the reference image may be referred to as a reference frame.
- the identification information of the reference frame ie, the reference image
- the reference image may be an identification method such as the frame number of the reference frame or a time stamp.
- the server side is based on panoramic video, in addition to generating the high-quality bitstream and low-quality bitstream as described above, it can also generate each image (or each frame) included.
- a code stream for independent decoding hereinafter referred to as a reference code stream.
- independent decoding means that all image blocks in the frame are decoded by intra-frame prediction, that is, the reference code stream is the code stream obtained by performing intra-frame prediction on all image blocks of all image frames described in it, and
- the reference code stream is independent of the high-quality code stream.
- the reference code stream may be a RARF stream.
- each frame in the reference code stream may be a pure random access (Clean Random Access, CRA) type I frame.
- each frame in the reference code stream may also be a P frame, where all P frames are I blocks (that is, the P frames that include the I blocks are still decoded by intra-frame reference prediction).
- Frame, the corresponding frame type can be called P slice (slice)).
- the image content corresponding to the reference code stream corresponds to the corresponding tile (that is, the tile decoded based on the reference frame in the reference code stream, such as the second sub-image in the embodiment of this application)
- the reference code stream hereinafter referred to as the RARF stream
- Figure 7 after the view angle is switched, the Tile n in the view angle 2 is the new Tile.
- the terminal performs viewing angle switching at the time corresponding to the P 2 frame (that is, the above-mentioned time t).
- the P 3 frame of the Tile n (that is, the image corresponding to the above-mentioned time t 2 ) performs inter-frame prediction
- the reference frame that the decoding relies on is the P 2 frame, but because the terminal switches at the P 2 frame time, there is no P 2 frame in the terminal (the reason is that the terminal obtains the corresponding image at t 2 and after from the server) then, in the present application, RARF stream frame (i.e. frame shown in FIG. RARF2) P can be used as a reference frame 3, so that the terminal can decode three of P frames based RARF, may continue The P frames in the segment are sequentially decoded.
- the server may send one or more RARF frames to the terminal, so that the terminal decodes multiple frames in the segment based on the RARF frame as the reference frame of the segment.
- the image at time t 2 can be the first image in the GOP, or the first image and the second image (or called the first frame in the segment).
- the image content of the RARF reference frame can be the same as the image content of the first frame (wherein, the decoding type of the first frame can be intra-frame decoding or inter-frame decoding). Consistent, where the decoding type of the RARF reference frame is intra-frame decoding.
- the RARF reference frame can be used as the reference frame of the second frame and decoded.
- the image content of the reference frame may be consistent with the image content of a frame whose decoding type is inter-frame decoding, such as the second frame or the third frame, that is, as described in the above example.
- the image content of the RARF reference frame may be the same as the image content of the reference frame on which the decoding of the frame at time t 2 (that is, the second sub-image in this embodiment of the application) depends on.
- the second sub-reference frame picture decoding may depend at time t 1, i.e. the previous frame corresponding to the time a handover, i.e., RARF frame image corresponding to the new content may Tile The image content of the frame corresponding to the moment before the switching moment is the same.
- the terminal switches the perspective at the time corresponding to the P 2 frame of Tile n, then the image content corresponding to the RARF frame can be the same as the image content corresponding to the P 2 frame, then the P 3 frame can be based on the RARF frame
- the second view of the terminal at time t 2 plays the image described in frame P 3 .
- the image content corresponding to the RARF frame may also be the same as the image content of the nth frame before the switching time.
- P 2 corresponding to a frame image contents may also be the same as the contents of an image of P.
- the second view of the terminal played at time t 2 may also be an image described in the RARF frame.
- the terminal switches the perspective at the time corresponding to the P 2 frame of Tile n, then the image content corresponding to the RARF frame can be the same as the image content corresponding to the P 3 frame, then the terminal responds to the RARF frame after decoding, at time t 2 is an image playback RARF frames (frames or P 3) as described, and then, P 4 frame may continue to decode the frame RARF a reference frame.
- the server can generate three code streams in advance: a low-quality code stream, a high-quality code stream, and a RARF stream corresponding to the high-quality code stream.
- the server can send the corresponding low-quality code stream, the tiles in the high-quality code stream and the corresponding reference frame to the terminal based on the received request message. For example: In a live broadcast scenario, each time the server receives a request message from the terminal, it generates a low-quality code stream, a high-quality code stream, and a RARF stream corresponding to the required video data based on the request message, and combines the low-quality code stream, The high-quality code stream is sent to the terminal.
- the image corresponding to the playback time t 2 of the switched perspective is the image corresponding to frame P 3 in the image sequence of the segment.
- Terminal sends a request message to the server carrying Tile n have the frame number 3, the P frames.
- the server receives the request message and determines that the terminal has a perspective switch, the server generates Tile n based on the request message (the generated Tile n starts from the P 3 frame, refer to Figure 7) and generates the RARF frame of the P 3 frame. encoded data for frames P 3 RARF reference frames to encode the decoded data.
- the server transmits Tile n frames starting from P 3 and RARF Segment1 frame corresponding to the terminal.
- the terminal can request the same RARF frame as the image content of the P 3 frame and the P 4 frame and subsequent P frames like the server.
- the server generates the coded data of the RARF frame and the code of the P 4 frame based on the request message. Data, etc., and encapsulate the encoded data of the above frame into segments, and send them to the terminal.
- the server can generate corresponding high-quality code streams, low-quality code streams, and RARF streams based on the video data, and store them locally on the server. Subsequently, the server may send the corresponding low-quality code stream and the tile in the high-quality code stream to the terminal based on the received request message of the terminal. Similarly, if the server detects that the viewing angle of the terminal has been switched, the server can send the Tile in the high-quality code stream corresponding to the new viewing angle and the corresponding RARF frame to the terminal.
- the server may also generate the corresponding RARF frame after detecting the terminal viewing angle switch, or after receiving the request message carrying the RARF stream. For example: before the viewing angle is switched, in the live broadcast scene, the server generates the corresponding high-quality tile encoded data based on the terminal's request message, and sends it to the terminal. Perspective terminal after the handover, the terminal sends a request message to the server carrying Tile n have the frame number 3, the P frames.
- the server After the server receives the request message and determines that the terminal has a perspective switch, the server generates Tile n based on the request message (the generated Tile n starts from the P 3 frame, refer to Figure 7) and generates the RARF frame of the P 3 frame. P 3 for the reference frame, for RARF decoded frame. The server then transmits Tile n frames starting from P 3 and RARF Segment1 frame corresponding to the terminal.
- the terminal while requesting the encoded data of the new Tile, the terminal will continue to request the old Tile (ie, the overlapping Tile described above) that appears in the new perspective. Encode data. Therefore, while the server sends the encoded data of the new tile corresponding to the new perspective to the terminal, it will continue to send the encoded data of the overlapping tile to the terminal.
- the old Tile ie, the overlapping Tile described above
- the terminal sends the frame number of the I 0 frame carrying Tile m to the server, that is, the terminal requests the server for the encoding of the first segment of Tile m (assuming that each segment has a duration of 3 seconds) Data, then, at 1 minute and 32 seconds (that is, at time t 1 ), the terminal's perspective is switched, then the terminal requests the server for the P 15 frame of Tile n (that is, the frame corresponding to the playback at time t 2 ) (belonging to Tile n The first Segment) and the RARF frame corresponding to the P 15 frame.
- the server After the server receives the request, it generates Tile n and the corresponding RARF frame (it should be noted that in this embodiment, the generated RARF frame is a reference frame corresponding to the P 15 frame, that is, the client can be based on a RARF frame pair P 15 performs decoding), and starts from the P 15 frame and transmits Tile n and the RARF frame corresponding to the P 15 frame to the terminal.
- the server will continue to send the remaining frames in the first segment of Time m to the terminal, and at 1 minute and 33 seconds, the terminal will continue to request the server for the encoded data of the second segment of Tile m, and the server will respond to the terminal. Send the encoded data of other segments of Tile m to the terminal.
- the server returns a response message to the terminal.
- the response message also carries information indicating the actual location of each tile in the image data or video image.
- Step 104 The terminal decodes the sub-image in the second view based on the encoded data of the reference image.
- the terminal obtains the encoded data of the Tile in the high-quality code stream sent by the server (specifically, the encoded data of each image frame in the Segment in the Tile), and the encoded data of the reference image (for example, The coded data of the RARF reference frame), and the location information of each tile (the location information obtained by the terminal from the server is the actual location information of the tile in the screen). Then, the terminal can decode the encoded data of the RARF reference frame to obtain the RARF reference frame. The terminal can decode the new tile in the new view based on the decoded RARF reference frame. In addition, the terminal may continue to decode the overlapped tiles based on the received reference frames (I frame and/or P frame).
- the Tile including Tile1, Tile2, Tile3, Tile4, Tile8, Tile9, Tile10, Tile11
- View 1 Old View
- View 2 Tiles in the new perspective
- Tile4 and Tile11 are overlapping tiles.
- the number of tiles shown in FIG. 8 is only an example of suitability. For example, there may be one or more overlapping tiles (for example, 5), and the specific number is determined according to the viewing angle, the rotation angle, and the speed.
- the terminal decodes the code stream (including the encoded data of the new Tile and the overlapping Tile) as follows:
- the terminal merges the coded data of the new tile with the coded data of the overlapping tile.
- Tile Merge Tile Merge, or Tile Rewrite (bitstream rewrite)
- Tile Rewrite bitstream rewrite
- the reference frame on which the encoded data of Tile4 and the encoded data of Tile11 are decoded is still located on the rightmost side of the TileMerge module. Therefore, the prior art Tile4 and Tile11 cannot be decoded because they cannot find the reference frame. .
- Tile4 and Tile11 can still be placed on the far right side of the Tile Merge module, that is, keep their original positions unchanged. Then, the coded data of Tile4 and the coded data of Tile11 can be decoded based on the received reference frame (I frame or P frame).
- the terminal places new tiles (including Tile5, Tile6, Tile7, Tile12, Tile13, and Tile14) in empty positions other than overlapping tiles in Tile Merge.
- the placement position can be as shown in Figure 8. It should be noted that the placement of the new Tile in Figure 8 is only an example of suitability, and the new Tile can be placed in any empty position except for overlapping tiles.
- the terminal merges multiple tiles in the Tile Merge. For a specific fusion manner, refer to the technical solutions in the prior art embodiments, which will not be repeated in this application.
- the terminal stores the actual location information and the fusion location information of the Tile in the local cache.
- the fusion position of the Tile in the Tile Merge (that is, the placement position of the Tile in the Tile Merge), and the actual position of the Tile in the video image (that is, the position information sent by the server in step 105) can be used Represented in an array.
- the information in the two arrays in view 1 is [1, 2, 3, 4, 8, 9, 10, 11].
- the information in array 1 is updated to [4,5,6,7,11,12,13,14], and the information in array 2 is [5,6,7,4] ,12,13,14,11].
- the above-mentioned array may also be maintained by the server, that is, the server generates a corresponding array in real time according to the fusion position of the tile in the terminal and the actual position of the tile, and sends it to the client when the client switches the perspective.
- the terminal decodes the encoded data of the merged Tile through a single decoder to obtain the Tile.
- the terminal decodes the encoded data of the merged Tile through a single decoder to obtain the Tile.
- the terminal may project each Tile based on the fused location information and the actual location information to obtain the video image to be displayed.
- the array 2 corresponding to the viewing angle 2 stored in the terminal is: [5,6,7,4,12,13,14,11], based on the array 2, and referring to the array 1: [4,5,6,7 ,11,12,13,14], when projecting, just place Tile4 and Tile11 in the position represented by array 1, and then the terminal will display the projected Tile on the screen of the terminal.
- the technical solution in the embodiment of the present application sets the reference code stream in which each frame is an intra-frame decoding frame as the reference frame of the high-quality code stream, so that the terminal can switch the view based on the reference frame.
- Decoding the new tiles in the perspective quickly switching the code stream without backing up a large amount of data, significantly reducing server-side storage overhead.
- the overlapping tiles can continue to be decoded at the original position, thereby reducing repeated downloads of the downloaded stream to the terminal, thereby reducing bandwidth overhead and improving resource utilization.
- the terminal may request the server for the new Tile, the RARF frame corresponding to the new Tile (that is, the RARF frame corresponding to the frame with the new Tile in the first position), and overlapping tiles and overlapping tiles.
- the RARF frame corresponding to the Tile (the RARF frame corresponding to the frame with the overlapping Tile in the first position).
- the server may also send the new tile, the RARF frame corresponding to the new tile, and the overlapping tile and the RARF frame corresponding to the overlapping tile to the terminal. Then, in this embodiment, the positions of the overlapping Tile and the new Tile in the Tile Merge can be arbitrarily placed.
- the terminal can decode the frame with the overlapping Tile in the first position based on the RARF frame corresponding to the overlapping Tile, and , Based on the RARF frame corresponding to the new Tile, decode the new Tile. And based on the location information, each Tile is reorganized to obtain the source video image.
- FIG. 9 is a schematic flowchart of a video processing method in an embodiment of the application, and in FIG. 9:
- Step 201 The terminal obtains the coded data of the sub-image covered by the first view and the coded data of the corresponding reference image from the server.
- the terminal sends a request message generated based on a user instruction to the server.
- the request message can be used to request the low-quality layer displayed by the terminal and the high-quality layer within the perspective of the terminal from the server.
- the server can encode and encapsulate the image into a high-quality code stream, a low-quality code stream, and a RARF stream.
- the high-quality code stream generated by the server is: the panoramic video is divided into multiple rectangular tiles in the spatial dimension, and each tile is divided into multiple segments in the time dimension.
- the duration can be 1 second (it can be divided according to actual needs).
- the spatial size of each tile can be the same or different, and the duration of each segment in the tile can also be the same or different.
- the prediction mode of the image frame in each segment of each Tile is inter-frame prediction, that is, the code stream obtained by performing inter-frame prediction on all the image frames described by the segment, for example: all image frames in the segment can be It is a P frame.
- the server encodes the high-quality code stream frame by frame into the RARF code stream according to the video picture corresponding to the Tile, if the first RARF frame of each segment is used as the reference frame of each segment of the high-quality code stream.
- the code stream corresponding to the video screen data sent by the server to the terminal includes: a low-quality code stream, a high-quality code stream, and the first RARF frame of each segment in the RARF code stream.
- the server pre-encodes the video data to generate code streams of different qualities, including but not limited to: high-quality code streams, low-quality code streams, and RARF streams.
- the server receives the request message sent by the terminal, determines the video data required by the terminal (the required data can be the current data to be displayed and/or predicted video data), and sends the low-quality code stream corresponding to the video data to the terminal.
- the server after the server receives the request message from the terminal, it determines the video data required by the terminal, and encodes the video data required by the terminal to obtain the encoded low-quality stream and high-quality stream And RARF stream. Then, the server sends the low-quality code stream, the Tile in the high-quality code stream corresponding to the video data, and the first RARF frame of each segment in the RARF stream corresponding to the high-quality code stream to the terminal.
- the tiles in the sent high-quality code stream are determined based on the identification information carried in the request message sent by the terminal.
- the server may only generate the RARF frame used as the reference frame of the high-quality code stream, that is, only generate the first RARF frame of each segment, thereby reducing overhead.
- step 101 please refer to step 101, which will not be repeated here.
- Step 202 The terminal decodes the acquired encoded data of the sub-image based on the encoded data of the reference image.
- the terminal separately decodes the low-quality code stream and the high-quality code stream to obtain the base layer (or can be called the low-quality layer) corresponding to the low-quality code stream and the high-quality layer corresponding to the high-quality code stream .
- the terminal receives tiles corresponding to multiple high-quality code streams, and the high-quality code streams combine and merge the multiple tiles to realize the use of a single decoder to decode the merged multiple tiles. Then, the terminal can decode the low-quality code stream based on the I frame of the low-quality code stream. And, the terminal decodes the high-quality code stream based on the received RARF frame.
- the terminal receives the encoded data of the first RARF frame (RARF1 frame for short) of the first segment of the RARF n stream, where the RARF n stream and the tile n correspond to the same image content, that is, the RARF n is according to the tile n If it is generated by encoding frame by frame, the terminal uses the RARF1 frame as the reference frame of the first segment of the tile, and decodes the first image in the image sequence of the first segment.
- RARF1 frame the reference frame of the first segment of the tile
- Step 203 When the first view is switched to the second view, the terminal obtains the encoded data of the sub-image covered by the second view and the encoded data of the corresponding reference image from the server.
- the video data required by the terminal needs to be updated, and the terminal sends a request message to the server to request low-quality layers and new perspectives (including new perspectives and new perspectives).
- the part of the specified range outside the viewing angle) video data is not limited to the viewing angle.
- step 103 For specific details, refer to step 103, which will not be repeated here.
- Step 204 The terminal decodes the sub-image in the second view based on the encoded data of the reference image.
- the server sends the low-quality code stream, high-quality code stream, and coded data of the RARF frame covered by the new view to the terminal (it should be noted that, in this embodiment, the RARF frame refers to The first frame in the tile in the new view (the first frame refers to the RARF frame of the reference frame of the segment of the tile in the new view and the frame corresponding to the view switching moment).
- the server continues to send the segment in the overlapping tile to the terminal based on the request of the terminal, and, based on the request of the terminal, sends the encoded data of the new tile and the RARF frame corresponding to the new tile to the terminal (It should be noted that the number of RARF frames sent can be one or more than one. Among them, if one RARF frame is sent, the RARF frame is as described above, which is used to perform data frames in the first place in the Tile. Decoded reference frame).
- the server returns a response message to the terminal.
- the response message also carries information indicating the actual location of each tile in the video image.
- the actual position information can be represented by an array, and the specific representation method will be illustrated in the following embodiments.
- the terminal obtains the tile coded data, the corresponding reference frame, and the location information of each tile in the high-quality code stream sent by the server. Then, the terminal can decode the encoded data of the image frame in the first position in the segment of the new tile in the new view based on the RARF frame. In addition, the terminal can also decode the first image frame in the overlapping Tile based on the received encoded data of the RARF frame, and continue to decode other frames.
- step 104 For specific details, refer to step 104, which will not be repeated here.
- the technical solution in the embodiment of the present application has no I-frame in the high-quality code stream, thereby further reducing the storage overhead on the server side.
- the terminal may request the server for a new tile, a RARF frame corresponding to the new tile, and an overlapping tile and an RARF frame corresponding to the overlapping tile.
- the server may also send the new tile, the RARF frame corresponding to the new tile, and the overlapping tile and the RARF frame corresponding to the overlapping tile to the terminal.
- the positions of the overlapped Tile and the new Tile in the Tile Merge can be arbitrarily placed.
- the terminal can decode the overlapped Tile based on the RARF frame corresponding to the overlapped Tile, and based on the new Tile.
- the corresponding RARF frame decodes the new Tile. And based on the location information, map each tile to obtain the video image to be displayed.
- the high-quality code stream may further include multiple code streams of different quality levels (ie, different coding rates).
- high-quality code streams can include three high-quality code streams with quality level A, quality level B, and quality level C. Among them, quality grade A is higher than quality grade B and higher than quality grade C.
- the server may determine the quality level corresponding to the Tile sent to the terminal based on external conditions.
- External conditions include but are not limited to: the processing capacity of the terminal, the transmission bandwidth between the server and the terminal, etc. For example: if the load of the current terminal is large or the bandwidth is small, the server can transmit to the terminal the coded data of the tile corresponding to the code stream of the quality level C. Conversely, if the external conditions are good, the server can transmit to the terminal the encoded data of the Tile corresponding to the code stream of quality level A. Generally, the server can transmit one or more streams of quality level A, quality level B, and quality level C to the terminal.
- the server can send the human face image to the terminal with a code stream of A quality level, and other than the face image within the viewing angle range
- the bit stream of the image generation quality level B and/or C is sent to the terminal, so that the terminal can display the human face image more clearly.
- the video processing device includes a hardware structure and/or software module corresponding to each function.
- the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed by hardware or computer software-driven hardware depends on the specific application and design constraint conditions of the technical solution. Professionals and technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered beyond the scope of this application.
- the embodiment of the present application may divide the video processing device into functional modules according to the foregoing method examples.
- each functional module may be divided corresponding to each function, or two or more functions may be integrated into one processing module.
- the above-mentioned integrated modules can be implemented in the form of hardware or software functional modules. It should be noted that the division of modules in the embodiments of the present application is illustrative, and is only a logical function division, and there may be other division methods in actual implementation.
- FIG. 10 shows a type of the video processing apparatus 300 (such as the terminal 200) involved in the above-mentioned embodiment.
- the video processing apparatus may include: a decoding module 301 and an acquisition module 302.
- the decoding module 301 can be used for the step of "decoding the encoded data of the first sub-image covered by the first view to obtain the first sub-image".
- this module can support the video processing device 300 to execute step 102 and step 104 in the foregoing method embodiment.
- the acquiring module 302 may be used for the step of “acquiring the encoded data of the second sub-image covered by the second perspective when the first perspective is switched to the second perspective”.
- this module can support the video processing device 300 to execute step 103 in the foregoing method embodiment.
- the decoding module 301 can also be used to “decode the encoded data of the first reference image to obtain a first reference image” and “decode the encoded data of the second sub-image according to the first reference image to obtain the The second sub-image” step.
- this module can support the video processing device 300 to execute step 104 in the above method embodiment.
- FIG. 11 shows a schematic diagram of a possible structure of the video processing device 400 (for example, the server 100) involved in the above embodiment.
- the video processing device 400 may include: a sending module 401 .
- the sending module 401 can be used to "send to the terminal the encoded data of the first sub-image covered by the first view of the terminal" and "send to the terminal when the first view of the terminal is switched to the second view The step of encoding data of the second sub-image covered by the second view.
- the device includes a processing module 501 and a communication module 502.
- the device further includes a storage module 503.
- the processing module 501, the communication module 502, and the storage module 503 are connected by a communication bus.
- the communication module 502 may be a device with a transceiver function for communicating with other network devices or communication networks.
- the storage module 503 may include one or more memories, and the memories may be devices for storing programs or data in one or more devices or circuits.
- the storage module 503 may exist independently and is connected to the processing module 501 through a communication bus.
- the storage module may also be integrated with the processing module 501.
- the apparatus 500 may be used in network equipment, circuits, hardware components, or chips.
- the apparatus 500 may be a terminal in an embodiment of the present application, such as the terminal 200.
- the communication module 502 of the apparatus 500 may include an antenna and a transceiver of the terminal.
- the device 500 may be a chip in the terminal in the embodiment of the present application.
- the communication module 502 may be an input or output interface, pin or circuit, or the like.
- the storage module may store a computer execution instruction of the method on the terminal side, so that the processing module 501 executes the method on the terminal side in the foregoing embodiment.
- the storage module 503 can be a register, a cache or RAM, etc.
- the storage module 503 can be integrated with the processing module 501; the storage module 503 can be a ROM or other types of static storage devices that can store static information and instructions.
- the storage module 503 can be integrated with The processing module 501 is independent.
- the device 500 may implement the method executed by the terminal in the above-mentioned embodiment.
- the apparatus 500 may be a server in an embodiment of the present application, such as the server 100.
- the device 500 may be a chip in the server in the embodiment of the present application.
- the communication module 502 may be an input or output interface, pin or circuit, or the like.
- the storage module may store computer execution instructions of the server-side method, so that the processing module 501 executes the server-side method in the foregoing embodiment.
- the storage module 503 can be a register, a cache or RAM, etc.
- the storage module 503 can be integrated with the processing module 501; the storage module 503 can be a ROM or other types of static storage devices that can store static information and instructions.
- the storage module 503 can be integrated with The processing module 501 is independent.
- the device 500 is the server or the chip in the server in the embodiment of the present application, the method executed by the server in the above embodiment can be implemented.
- the embodiment of the present application also provides a computer-readable storage medium.
- the methods described in the foregoing embodiments may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. If implemented in software, the functions can be stored as one or more instructions or codes on a computer-readable medium or transmitted on the computer-readable medium.
- Computer-readable media may include computer storage media and communication media, and may also include any media that can transfer a computer program from one place to another.
- the storage medium may be any available medium that can be accessed by a computer.
- the computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used for carrying or with instructions or data structures
- the required program code is stored in the form of and can be accessed by the computer.
- any connection is properly termed a computer-readable medium.
- coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wireless technology such as infrared, radio and microwave
- coaxial cable, fiber optic cable , Twisted pair, DSL or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium.
- Magnetic disks and optical disks as used herein include compact disks (CDs), laser disks, optical disks, digital versatile disks (DVDs), floppy disks and blu-ray disks, where disks usually reproduce data magnetically, while optical disks use lasers to optically reproduce data. Combinations of the above should also be included in the scope of computer-readable media.
- the embodiment of the present application also provides a computer program product.
- the methods described in the foregoing embodiments may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. If it is implemented in software, it can be fully or partially implemented in the form of a computer program product.
- the computer program product includes one or more computer instructions. When the above computer program instructions are loaded and executed on the computer, the processes or functions described in the above method embodiments are generated in whole or in part.
- the above-mentioned computer may be a general-purpose computer, a special-purpose computer, a computer network, network equipment, user equipment, or other programmable devices.
- the steps of the method or algorithm described in combination with the disclosure of the embodiments of the present application may be implemented in a hardware manner, or may be implemented in a manner in which a processor executes software instructions.
- Software instructions can be composed of corresponding software modules, which can be stored in random access memory (Random Access Memory, RAM), flash memory, read-only memory (Read Only Memory, ROM), and erasable programmable read-only memory ( Erasable Programmable ROM (EPROM), Electrically Erasable Programmable Read-Only Memory (Electrically EPROM, EEPROM), register, hard disk, mobile hard disk, CD-ROM or any other form of storage medium known in the art.
- RAM Random Access Memory
- ROM read-only memory
- EPROM Erasable Programmable ROM
- EPROM Electrically Erasable Programmable Read-Only Memory
- register hard disk, mobile hard disk, CD-ROM or any other form of storage medium known in the art.
- An exemplary storage medium is coupled to the processor, so that the processor can read information from the storage medium and can write information to the storage medium.
- the storage medium may also be an integral part of the processor.
- the processor and the storage medium may be located in the ASIC.
- the ASIC may be located in a network device.
- the processor and the storage medium may also exist as discrete components in the network device.
- the functions described in the embodiments of the present application may be implemented by hardware, software, firmware, or any combination thereof. When implemented by software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on the computer-readable medium.
- the computer-readable medium includes a computer storage medium and a communication medium, where the communication medium includes any medium that facilitates the transfer of a computer program from one place to another.
- the storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Databases & Information Systems (AREA)
- Human Computer Interaction (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
本申请实施例提供了一种视频处理方法、装置及系统,该方法包括:解码第一视角覆盖的第一子图像的编码数据,从而得到第一子图像;在第一视角切换到第二视角时,获取第二视角覆盖的第二子图像的编码数据,以及第二子图像的第一参考图像的编码数据,第一参考图像的编码数据独立于第二子图像的编码数据所在的码流;解码第一参考图像的编码数据获得第一参考图像;根据第一参考图像解码第二子图像的编码数据,从而获得第二子图像。从而使终端在视角切换过程中,实现无缝切换,始终显示第二子图像所属的高质量层,有效提升了用户使用体验。
Description
本申请要求于2019年05月13日提交中国专利局、申请号为201910415208.2、申请名称为“视频处理方法、装置及系统”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本申请实施例涉及视频处理领域,尤其涉及一种视频处理方法、装置及系统。
目前,在已有技术中,终端进行视频播放时,如果发生视角切换,则,可能出现由于切换时刻未到达新视角中的切片(Tile)对应的随机切入点,因此,导致终端无法对新的Tile进行解码,仅能暂时播放质量较低的数据层或出现黑屏,造成用户使用体验差的问题。
发明内容
本申请提供一种视频处理方法、装置及系统,能够在一定程度上避免视角切换过程中出现的低数据层,造成用户使用体验差的问题。
为达到上述目的,本申请采用如下技术方案:
第一方面,本申请实施例提供一种视频处理方法,所述方法包括:终端(或客户端,或解码端)解码第一视角覆盖的第一子图像的编码数据,从而得到第一子图像。随后,终端在第一视角切换到第二视角时,获取第二视角覆盖的第二子图像的编码数据,以及第二子图像的第一参考图像的编码数据,其中,第一参考图像的编码数据独立于第二子图像的编码数据所在的码流。接着,终端解码第一参考图像的编码数据获得第一参考图像。以及,终端可根据第一参考图像解码第二子图像的编码数据,从而获得第二子图像。
通过上述方式,实现了服务器端预先准备两个独立的码流,其中,包括包含第一参考图像的编码数据的码流以及包含第二子图像(和/或第一子图像)的编码数据的码流,以使终端在视角切换后,可基于第一参考图像对切换后的视角覆盖的第二子图像进行解码,从而获得第二子图像,使终端在视角切换过程中,实现无缝切换,终端的显示画面始终保持在第二子图像所属的高质量层,而不会在视角覆盖范围内出现低质量层或黑屏的现象,有效提升了用户使用体验。
在一种可能的实现方式中,第二子图像的编码数据所在的码流为第二子图像的编码数据所在的段(Segment)码流。
具体的,在本申请中,服务器与终端的交互过程中,是以码流的形式进行数据传输,可选地,码流可以为段码流,即,服务器将编码封装后的至少一个段码流发送给终端,以供终端进行解码与显示。
在一种可能的实现方式中,第二子图像为第二子图像的编码数据所在的码流所描述的图像序列中在解码顺序上的第一幅图像之外的图像,或者第二子图像为第二子图像的编码数据所在的码流所描述的图像序列中在解码顺序上的第一幅图像以及第二幅图像之外的图像。
通过上述方式,实现了当切换时刻为第二子图像的编码数据所在的码流所描述的图像序列中在解码顺序上的第二幅图像对应的播放时刻时,相应的,第一参考图像与第二子图像所依赖的参考图像(即,位于段码流首位的图像帧)所描述的图像一致。可选地,当切换时刻为除第一幅图像以及第二幅图像,例如:第三幅图像或第四幅图像等图像对应的播放时刻时,相应的,第一参考图像可以与第三副图像或第四幅图像所依赖的参考图像所描述的图像相同。
在一种可能的实现方式中,第一参考图像的图像内容与第二子图像的第二参考图像的图像内容相同,第二参考图像为第二子图像的编码数据所在的码流所描述的图像序列中的图像。
通过上述方式,实现了服务器端在编码时,基于相同的图像内容对第一参考图像与第二子图像进行编码。可选地,编码时的编码率可相同,可不同。
在一种可能的实现方式中,第二参考图像为第二子图像的编码数据所在的码流所描述的图像序列中在解码顺序上的第一幅图像之外的图像。
通过上述方式,实现了当视角切换时,终端所获取的码流从第一子图像的编码数据所在的码流,例如第一子图像的编码数据所在的段码流,切换到第二子图像的编码数据所在的码流,例如第二子图像的编码数据所在的段码流时,第一参考图像的图像内容与第二子图像所依赖的参考图像(即第二参考图像)的图像内容是一致的,其中,第二参考图像可以为第二子图像的编码数据所在的段码流对应的图像序列中在解码顺序上的第一幅图像之外的图像。
在一种可能的实现方式中,第二参考图像为第二子图像在图像序列中的前一幅图像。
通过上述方式,实现了第一参考图像的图像内容可以与第二子图像在图像序列中的前一幅图像(即第二参考图像)的图像内容一致。
在一种可能的实现方式中,第二子图像的编码数据所在的码流为通过对码流中描述的所有图像帧执行帧间预测得到的码流。
通过上述方式,实现了第二子图像的编码数据所在的码流中的所有图像帧的预测方式可以为帧间预测。例如:其所在的段码流中的所有图像帧均可以为P帧。
在一种可能的实现方式中,第一参考图像的编码数据所在的码流为通过对码流中描述的所有图像帧的所有图像块执行帧内预测得到的码流,第一参考图像的编码数据所在 的码流独立于第二子图像的编码数据所在的码流。
通过上述方式,实现了第一参考图像的编码数据所在的码流中的所有图像帧的预测方式可以为帧内预测。例如:每个段码流中的所有图像帧均为I帧。
在一种可能的实现方式中,第一参考图像为纯随机接入CRA图像。
通过上述方式,实现了第一参考图像的编码数据所在的码流中的所有图像帧的预测方式可以为帧内预测,并且,所有图像帧可以为CRA型I帧。
在一种可能的实现方式中,第二子图像的不被第一视角覆盖。
通过上述方式,实现了终端获取到的第二子图像可不被第一视角覆盖,即,仅显示于第二视角中,也就是说,本申请可应用于对新的子图像(即,仅覆盖于切换后的视角而不被切换前的视角覆盖的子图像)的解码过程。
在一种可能的实现方式中,第二视角覆盖第一子图像的区域,方法还包括:终端解码第一子图像的编码数据所在的码流中的下一帧子图像,,从而得到下一帧子图像;以及,终端将下一帧图像与第一参考图像进行拼接,从而得到拼接后的图像;播放拼接后的图像。
通过上述方式,实现了第一参考图像被解码后,可显示于第二视角覆盖范围内,也就是说,在本申请中,若切换时刻为第二子图像所在图像序列中的前一幅图像对应的播放时刻,则,终端获取到的第一参考图像的图像内容与所述前一幅图像的图像内容一致,并且,终端对第一参考图像进行解码后,基于第一参考图像对第二子图像进行解码,以及,第二视角的覆盖范围内显示的为第一参考图像以及包括第二子图像及其所在的码流的其它后续的子图像。
在一种可能的实现方式中,第二视角覆盖第一子图像的区域,方法还包括:解码第一子图像的编码数据所在的码流中的下一帧子图像,,从而得到下一帧子图像;将下一帧图像与第二子图像进行拼接,从而得到拼接后的图像;播放拼接后的图像。
通过上述方式,实现了第一参考图像被解码后,不显示于第二视角覆盖范围内,也就是说,在本申请中,若切换时刻为第二子图像对应的播放时刻,则,终端获取到的第一参考图像的图像内容与第二子图像所依赖的参考图像(例如:本申请中的第二参考图像)的图像内容一致,并且,终端对第一参考图像进行解码后,基于第一参考图像对第二子图像进行解码,以及,第二视角的覆盖范围内显示的为包括第二子图像及其所在的码流的其它后续的子图像。
在一种可能的实现方式中,第一子图像对应的播放时刻为t
1,第二子图像对应的播放时刻为t
2,t
1时刻和t
2时刻为相邻播放时刻,或者,第一子图像和第二子图像在播放顺序上相差N帧,N为1,2,3或4。
通过上述方式,实现了当第一参考图像不作为显示时(不作为显示的情况如上文所 示,此处不赘述),本申请中所述的切换时刻可以为,终端从正在播放的第一子图像(需要说明的是,第一子图像所在的码流与第二子图像所在的码流不相同,或者可以理解为,视角切换时,终端所显示的段码流从一个段码流切换到另一个段码流)的播放时可t
1时刻开始切换,并且,切换后的视角所显示的第二子图像的播放时刻为t
2,其中,t
1时刻和t
2时刻为相邻播放时刻。或者,也可以是第一子图像和第二子图像在播放顺序上相差N帧,N为1,2,3或4。
在一种可能的实现方式中,第一子图像对应的播放时刻为t
1,第一参考图像对应的播放时刻为t
2,t
1时刻和t
2时刻为相邻播放时刻,或者,第一子图像和第一参考图像在播放顺序上相差N帧,N为1,2,3或4。
通过上述方式,实现了当第一参考图像作为显示时(作为显示的情况如上文所示,此处不赘述),本申请中所述的切换时刻可以为,终端从正在播放的第一子图像(需要说明的是,第一子图像所在的码流与第二子图像所在的码流不相同,或者可以理解为,视角切换时,终端所显示的段码流从一个段码流切换到另一个段码流)的播放时可t
1时刻开始切换,并且,切换后的视角所显示的第一参考图像的播放时刻为t
2,其中,t
1时刻和t
2时刻为相邻播放时刻。或者,也可以是第一子图像和第一参考图像在播放顺序上相差N帧,N为1,2,3或4。
在一种可能的实现方式中,第二视角覆盖第一子图像区域,第一视角还覆盖第三子图像区域,解码第一视角覆盖的第一子图像的编码数据,从而得到第一子图像包括:将第一码流和第三码流合并,从而得到第一合并码流,第一码流为第一子图像的编码数据所在的码流,第三码流为第三子图像的编码数据所在的码流;解码第一合并码流,从而得到包括第一子图像和第三子图像的图像;根据第一参考图像解码第二子图像的编码数据,从而获得第二子图像包括:将第一码流与第二码流合并,从而得到第二合并码流,第二码流为第二子图像的编码数据所在的码流,第一码流所描述的子图像对应在第一合并码流描述的图像中的位置,与第一码流所描述的子图像对应在第二合并码流描述的图像中的位置一致;根据第一参考图像解码第一合并码流,从而得到包括第二子图像和第一码流所描述的子图像的图像。
通过上述方式,实现了终端的视角切换后,对于存在于第一视角覆盖与第二视角覆盖范围内的第一子图像以及仅覆盖于第二视角的第二子图像的解码过程,由于第一子图像所在码流在当前播放时刻所显示的图像所以来的参考图像以缓存在终端中,因此,终端对该码流(即第一码流)进行解码时,该码流在合并码流(即第二合并码流)所描述的图像中的位置可以与第一合并码流所描述的图像位置相同,从而避免该类(例如本申请实施例中所涉及到的重叠Tile中的Segment)码流在解码过程中出现解码失败的问题。
第二方面,本申请实施例提供了一种视频处理方法,该方法包括:服务器(或编码端)向终端发送终端的第一视角覆盖的第一子图像的编码数据;服务器在终端的第一视角切换到第二视角时,向终端发送第二视角覆盖的第二子图像的编码数据,以及第二子 图像的第一参考图像的编码数据,第一参考图像的编码数据独立于第二子图像的编码数据所在的码流。
在一种可能的实现方式中,第二子图像的编码数据所在的码流为第二子图像的编码数据所在的段码流。
在一种可能的实现方式中,第二子图像为第二子图像的编码数据所在的码流所描述的图像序列中在解码顺序上的第一幅图像之外的图像,或者第二子图像为第二子图像的编码数据所在的码流所描述的图像序列中在解码顺序上的第一幅图像以及第二幅图像之外的图像。
在一种可能的实现方式中,第一参考图像的图像内容与第二子图像的第二参考图像的图像内容相同,第二参考图像为第二子图像的编码数据所在的码流所描述的图像序列中的图像。
在一种可能的实现方式中,第二参考图像为第二子图像的编码数据所在的码流所描述的图像序列中在解码顺序上的第一幅图像之外的图像。
在一种可能的实现方式中,第二参考图像为第二子图像在图像序列中的前一幅图像。
在一种可能的实现方式中,第二子图像的编码数据所在的码流为通过对码流中描述的所有图像帧的执行帧间预测得到的码流。
在一种可能的实现方式中,第一参考图像的编码数据所在的码流为通过对码流中描述的所有图像帧的所有图像块执行帧内预测得到的码流,第一参考图像的编码数据所在的码流独立于第二子图像的编码数据所在的码流。
在一种可能的实现方式中,第一参考图像为纯随机接入CRA图像。
在一种可能的实现方式中,第二子图像不被第一视角覆盖。
第三方面,本申请实施例提供了一种视频处理装置,装置可以包括:解码模块、以及获取模块,其中,解码模块可用于解码第一视角覆盖的第一子图像的编码数据,从而得到第一子图像;获取模块可用于在第一视角切换到第二视角时,获取第二视角覆盖的第二子图像的编码数据,以及第二子图像的第一参考图像的编码数据,第一参考图像的编码数据独立于第二子图像的编码数据所在的码流;以及,解码模块可以进一步用于解码第一参考图像的编码数据获得第一参考图像;解码模块还可以用于根据第一参考图像 解码第二子图像的编码数据,从而获得第二子图像。
在一种可能的实现方式中,第二子图像的编码数据所在的码流为第二子图像的编码数据所在的段(Segment)码流。
在一种可能的实现方式中,第二子图像为第二子图像的编码数据所在的码流所描述的图像序列中在解码顺序上的第一幅图像之外的图像,或者第二子图像为第二子图像的编码数据所在的码流所描述的图像序列中在解码顺序上的第一幅图像以及第二幅图像之外的图像。
在一种可能的实现方式中,第一参考图像的图像内容与第二子图像的第二参考图像的图像内容相同,第二参考图像为第二子图像的编码数据所在的码流所描述的图像序列中的图像。
在一种可能的实现方式中,第二参考图像为第二子图像的编码数据所在的码流所描述的图像序列中在解码顺序上的第一幅图像之外的图像。
在一种可能的实现方式中,第二参考图像为第二子图像在图像序列中的前一幅图像。
在一种可能的实现方式中,第二子图像的编码数据所在的码流为通过对码流中描述的所有图像帧执行帧间预测得到的码流。
在一种可能的实现方式中,第一参考图像的编码数据所在的码流为通过对码流中描述的所有图像帧的所有图像块执行帧内预测得到的码流,第一参考图像的编码数据所在的码流独立于第二子图像的编码数据所在的码流。
在一种可能的实现方式中,第一参考图像为纯随机接入CRA图像。
在一种可能的实现方式中,第二子图像的不被第一视角覆盖。
在一种可能的实现方式中,第二视角覆盖第一子图像区域,第一视角还覆盖第三子图像区域,相应的,解码模块还可以用于:将第一码流和第三码流合并,从而得到第一合并码流,第一码流为第一子图像的编码数据所在的码流,第三码流为第三子图像的编码数据所在的码流;解码第一合并码流,从而得到包括第一子图像和第三子图像的图像;以及,解码模块还可以用于将第一码流与第二码流合并,从而得到第二合并码流,第二码流为第二子图像的编码数据所在的码流,第一码流所描述的子图像对应在第一合并码流描述的图像中的位置,与第一码流所描述的子图像对应在第二合并码流描述的图像中的位置一致;根据第一参考图像解码第一合并码流,从而得到包括第二子图像和第一码 流所描述的子图像的图像。
第四方面,本申请实施例提供一种视频处理装置,装置包括:发送模块,该模块可以用于向终端发送终端的第一视角覆盖的第一子图像的编码数据;以及,还可以用于在终端的第一视角切换到第二视角时,向终端发送第二视角覆盖的第二子图像的编码数据,以及第二子图像的第一参考图像的编码数据,第一参考图像的编码数据独立于第二子图像的编码数据所在的码流。
在一种可能的实现方式中,第二子图像的编码数据所在的码流为第二子图像的编码数据所在的段码流。
在一种可能的实现方式中,第二子图像为第二子图像的编码数据所在的码流所描述的图像序列中在解码顺序上的第一幅图像之外的图像,或者第二子图像为第二子图像的编码数据所在的码流所描述的图像序列中在解码顺序上的第一幅图像以及第二幅图像之外的图像。
在一种可能的实现方式中,第一参考图像的图像内容与第二子图像的第二参考图像的图像内容相同,第二参考图像为第二子图像的编码数据所在的码流所描述的图像序列中的图像。
在一种可能的实现方式中,第二参考图像为第二子图像的编码数据所在的码流所描述的图像序列中在解码顺序上的第一幅图像之外的图像。
在一种可能的实现方式中,第二参考图像为第二子图像在图像序列中的前一幅图像。
在一种可能的实现方式中,第二子图像的编码数据所在的码流为通过对码流中描述的所有图像帧的执行帧间预测得到的码流。
在一种可能的实现方式中,第一参考图像的编码数据所在的码流为通过对码流中描述的所有图像帧的所有图像块执行帧内预测得到的码流,第一参考图像的编码数据所在的码流独立于第二子图像的编码数据所在的码流。
在一种可能的实现方式中,第一参考图像为纯随机接入CRA图像。
在一种可能的实现方式中,第二子图像不被第一视角覆盖。
第五方面,本申请实施例提供了一种视频处理系统,包括服务器端与终端,其中,服务器端向终端发送终端的第一视角覆盖的第一子图像的编码数据;终端接收第一子图 像的编码数据,并解码第一子图像的编码数据,得到第一子图像;在终端的第一视角切换到第二视角时,服务器端向终端发送第二视角覆盖的第二子图像的编码数据,以及第二子图像的第一参考图像的编码数据,第一参考图像的编码数据独立于第二子图像的编码数据所在的码流;终端接收第二子图像的编码数据以及第一参考图像的编码数据,并解码第一参考图像的编码数据获得第一参考图像;以及,终端根据第一参考图像解码第二子图像的编码数据,从而获得第二子图像。
第六方面,本申请实施例提供了一种计算机可读介质,用于存储计算机程序,该计算机程序包括用于执行第一方面或第一方面的任意可能的实现方式中的方法的指令。
第七方面,本申请实施例提供了一种计算机可读介质,用于存储计算机程序,该计算机程序包括用于执行第二方面或第二方面的任意可能的实现方式中的方法的指令。
第八方面,本申请实施例提供了一种计算机程序,该计算机程序包括用于执行第一方面或第一方面的任意可能的实现方式中的方法的指令。
第九方面,本申请实施例提供了一种计算机程序,该计算机程序包括用于执行第二方面或第二方面的任意可能的实现方式中的方法的指令。
第十方面,本申请实施例提供了一种芯片,该芯片包括处理电路、收发管脚。其中,该收发管脚、和该处理电路通过内部连接通路互相通信,该处理器执行第一方面或第一方面的任一种可能的实现方式中的方法,以控制接收管脚接收信号,以控制发送管脚发送信号。
第十一方面,本申请实施例提供了一种芯片,该芯片包括处理电路、收发管脚。其中,该收发管脚、和该处理电路通过内部连接通路互相通信,该处理器执行第二方面或第二方面的任一种可能的实现方式中的方法,以控制接收管脚接收信号,以控制发送管脚发送信号。
为了更清楚地说明本申请实施例的技术方案,下面将对本申请实施例的描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本申请的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动性的前提下,还可以根据这些附图获得其他的附图。
图1是示例性示出的一种视频处理方法的传输系统示意图;
图2是示例性示出的一种传输流程示意图;
图3是示例性示出的一种视频处理方法的流程示意图;
图4是本申请实施例提供的一种视频处理系统的结构示意图;
图5是本申请实施例提供的一种视频处理方法的流程示意图;
图6是本申请实施例提供的切换过程示意图;
图7是本申请实施例提供的参考帧的作用方式示意图;
图8是本申请实施例提供的解码方式的流程示意图;
图9是本申请实施例提供的一种视频处理方法的流程示意图;
图10是本申请实施例提供的一种视频处理装置的结构示意图;
图11是本申请实施例提供的一种视频处理装置的结构示意图;
图12是本申请实施例提供的一种视频处理装置的结构示意图。
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例是本申请一部分实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员在没有作出创造性劳动前提下所获得的所有其他实施例,都属于本申请保护的范围。
本文中术语“和/或”,仅仅是一种描述关联对象的关联关系,表示可以存在三种关系,例如,A和/或B,可以表示:单独存在A,同时存在A和B,单独存在B这三种情况。
本申请实施例的说明书和权利要求书中的术语“第一”和“第二”等是用于区别不同的对象,而不是用于描述对象的特定顺序。例如,第一目标对象和第二目标对象等是用于区别不同的目标对象,而不是用于描述目标对象的特定顺序。
在本申请实施例中,“示例性的”或者“例如”等词用于表示作例子、例证或说明。本申请实施例中被描述为“示例性的”或者“例如”的任何实施例或设计方案不应被解释为比其它实施例或设计方案更优选或更具优势。确切而言,使用“示例性的”或者“例如”等词旨在以具体方式呈现相关概念。
在本申请实施例的描述中,除非另有说明,“多个”的含义是指两个或两个以上。例如,多个处理单元是指两个或两个以上的处理单元;多个系统是指两个或两个以上的系统。
为使本领域技术人员更好地理解本申请,下面对本申请中可能涉及到的背景技术进行简单介绍:
伴随着虚拟现实(Virtual Reality,VR)应用的兴起,全景视频业务逐渐在各个平台上线,成为网络视频流量的重要组成部分。其中,全景视频业务具有超高分辨率和超高码率的特点。为在保证用户体验质量的同时尽量降低带宽消耗,目前,在全景视频点播业务中通常采用针对Tile的基于超文本传输协议(HyperText Transfer Protocol,HTTP)的动态自适应流(Dynamic Adaptive Streaming over HTTP,DASH)传输系统对全景视频进行传输。一种典型的传输系统图1所示:
在图1中,服务器端可用于提供基于全景视频生成的码流,其中,码流包括但不限于:增强层,或者可以称为高质量层或高质量码流,以及,基础层,或者可以称为低质量层或低质量码流。其中,低质量层是指以较低的码率或较小的分辨率对全景视频进行 编码,高质量层则是指以较高的码率或较大的分辨率对全景视频进行编码,其中,高质量层包括服务器按照全景视频内容的不同区域划分的多个Tile。随后,服务器可通过DASH,或基于HTTP的直播流((HTTP Live Streaming,HLS)等技术向终端传输码流(包括低质量层与高质量层)。
仍参照图1,终端获取到低质量层与高质量层后,将高质量层中的多个Tile进行融合,并对融合后的Tile进行解码,以及,终端对低质量层进行解码。接着,终端显示高质量层与低质量层,如图1所示,即,低质量层为覆盖全部全景视频的数据层,而高质量层则可以为终端视角上(包括视角内以及超出视角的指定范围部分)的视频数据部分。从而使用户视角内能够呈现出较为清晰、高质量的视频画面。需要说明的是,在已有技术中,终端在使用过程中,将会持续请求低质量层,以确保终端的视角转动时,不会因高质量层的视频数据对应的数据帧未成功解码而出现无画面或黑屏的区域。
在已有技术中,如上文所述,服务器将全景视频划分为多个Tile。其中,服务器进一步用于将每个Tile按照时间划分为多个段(Segment),每个段的起始位置为随机接入点RAP(Random Access Point),例如:I帧。终端只需请求用户视角上(包括视角内即视角外的指定范围内的部分)的高质量Tile和低质量层,即可在保证用户视角内的高质量画面的同时,降低带宽的要求和消耗。
但是,在已有技术中,Tile的基本传输单元为Segment,在不同的应用需求下,Segment具有不同的时间长度,如1秒或3秒等。每个Segment的起始位置通常编码为I帧,是进行段切换的随机接入点,也就是说,对Tile的切换就只能在RAP进行(原因为每个Segment中的帧需要以前面的一个或多个帧作为参考帧才能进行解码,而终端需要对每个Segment的随机接入点对应的I帧作为参考才能对后面的帧继续解码)。如图2所示(其中,图2中的Seg即为Segment的缩写)。当用户视角发生变化,并且,视角内出现新的Tile时(即,Tile m为旧视角内的Tile,Tile n为新视角内出现的Tile),终端只能在下一随机接入点到达时刻,才能对Tile n(即如图2所示的Seg t+1)进行解码,并显示高质量的新Tile(即Tile n)。即如图2所示,在Seg t中发生视角切换时,只能等到Seg t+1才能接入高质量的新Tile,在此期间,用户将观看到低质量的数据层或可称为低质量的视频画面。
因此,在已有技术中,由于终端只能在随机接入点到达时刻才能切换到高质量层的内容,甚至,假如终端在I帧过后的第一个P帧对应的时间点进行视角切换,则,终端需要等待接近一个Segment对应的时长,才能切换到高质量内容,严重影响了用户使用体验。
为解决上述问题,已有技术提出一种可行的实施方案,具体的,已有技术在服务器端按照图像组(Group of picture,GOP)大小对每个Tile进行逐帧移位的划分段,并且按照划分后的Tile进行编码,如图3所示。因此,当用户视角发生变化时,终端可立即请求视角切换时刻下一帧为I帧的段并进行解码,从而可以通过减少段的长度来降低切换等待的时长,实现实时切换高质量Tile的目的。
但是,已有技术增加了多个备份的Segment,即,在服务器端生成大量的备份数据,从而增大了服务器端的转码器和存储的开销,并且降低了服务器的视频编码效率,以及, 增加了网络传输的数据量。并且,在视角切换后,服务器将发送整段新的数据,而当带宽不理想时,则仍然无法使终端中显示的视频画面达到预期。
以及,在已有技术中的解码阶段,即,终端接收到所有Tile后,为实现使用单解码器对所有Tile进行实时解码,需要在终端首先对下载的Tile进行拼接融合来得到一个待解码的流。然而,用户视角的切换时刻通常发生在段中的某一位置,此时视角中原有Tile(即出现在旧视角与新视角中的Tile,也可以称为重叠Tile)参与融合的是P帧,新出现Tile参与融合的第一帧为I帧。此时按照新视角中Tile的位置进行Tile融合会导致已有Tile的P帧因位置改变而丢失其参考帧无法进行解码的问题。
针对已有技术中存在问题,本申请提出一种视频传输方法,旨在实现对高质量视频数据的实时切换。
在对本申请实施例的技术方案说明之前,首先结合附图对本申请实施例的视频传输系统进行说明。参见图4,为本申请实施例提供的一种视频传输系统示意图。该视频传输系统中包括服务器100、终端200。在本申请实施例具体实施的过程中,终端200可以为电脑、智能手机、VR等设备。需要说明的是,在实际应用中,服务器与终端的数量均可以为一个或多个,图4所示视频传输系统的服务器与终端的数量仅为适应性举例,本申请对此不做限定。
结合上述如图4所示的应用场景示意图,下面介绍本申请的具体实施方案:
场景一
结合图4,如图5所示为本申请实施例中的视频处理方法的流程示意图,在图5中:
步骤101,终端从服务器端获取第一视角覆盖的子图像的编码数据。
具体的,在本申请的实施例中,终端向服务器发送基于用户指令生成的请求消息。请求消息可用于向服务器请求终端所要显示的视频数据也可以称为视频画面(或者称为视频图像,或图像内容等)所对应的编码数据。
可选的,在本申请中,服务器将图像内容编码并封装为具有不同编码率的码流,包括但不限于:高质量码流(或称为高质量层)及低质量码流(或称为低质量层)。其中,对于高质量层,服务器将图像内容在空间维度上划分为多个子图像,可选的,在本申请中,子图像可以为Tile。以及,服务器对每个Tile所对应的图像内容进行编码,以获取与每个Tile对应的编码数据。相应的,在本申请中,终端可向服务器请求终端当前视角(例如:本申请实施例中的第一视角)覆盖的至少一个Tile的编码数据。需要说明的是,在本申请中,视角覆盖的Tile包括全部在视角范围内的Tile或者部分在视角范围内的Tile。
如上文所述,高质量层中的Tile(具体指Tile的编码数据的传输)的基本传输单元为Segment,即,终端与服务器端进行交互时的基本传输单元为Segment,或者可以理解为,本申请中所涉及到的码流(包括服务器端生成的码流、服务端向终端传输的码流以及终端接收到的码流)均指Segment码流。举例说明:终端可向服务器发送携带有所需的Tile中的Segment的标识信息,以向服务器请求相应的Tile中的目标Segment码流中的每个图像的编码数据。其中,标识信息可以为帧号或时间戳等标识方式。
可选的,标识信息还可以为Segment的存储位置,即,Segment的存储位置可以用于唯一标识该Segment。举例说明:请求的信息可以是在“服务器中码流segment的绝对路 径”(比如http://192.168.0.2/dash/Tile0/4.m4s(其中4.m4s表示一个segment))。
可选的,标识信息还可以为Segment的起始位置和大小,即,起始位置和大小可用于唯一标示该Segment。举例说明:请求的segment在文件1.mp4的第10000个字节,大小为2000个字节。
可选的,在本申请中,终端向服务器请求第一个Segment之后,可继续向服务器发送请求,以请求Tile n中的其它Segment的编码数据。
需要说明的是,上述终端所需的视频数据(也可以称为视频画面或视频图像)可以为点播场景中预下载的视频数据,或者,还可以为直播场景中的终端当前所显示(包括视角内与视角外)的视频数据。
举例说明:若在点播场景中,终端可向服务器请求视频数据,请求的视频数据可以为预测的终端所要显示的视频内容,即,实现预加载功能,通过预下载,提升终端的解码效率。相应的,在点播场景中,服务器预先准备有与视频数据对应的高质量码流(也可以称为高质量层)与低质量码流(也可以称为低质量层),可选的,在点播场景中,终端可向服务器请求预测的视角范围内(包括视角内与视角外的指定范围内的部分)的高质量层,以及请求预测的终端所要显示的视频数据对应的低质量层(如前所述,低质量层为对应于全景视频的编码数据)。
若在直播场景中,终端可实时向服务器请求视频数据,请求的视频数据则为终端当前所要显示的视频数据(或视频画面)。相应的,在直播场景中,服务器基于终端所请求的视频数据,将对应的视频数据进行编码后,发送给终端。可选的,在直播场景中,终端可向服务器请求视角范围内(包括视角内与视角外的指定范围内的部分)的高质量层,以及终端所要显示的视频数据对应的低质量层(如前所述,低质量层为对应于全景视频的编码数据)。
可选的,在本申请的实施例中,终端通过检测视角的尺寸,向服务器请求所需视频数据(例如:第一视角覆盖的Tile对应的编码数据)。举例说明:视角的形状可以为圆形或矩形,并且大小也可以不相同,因此,不同尺寸的视角所需要的高质量层则不相同。举例说明:较大视角需要10个高质量层中的Tile,而相对较小的视角则可能只需要6个对应的Tile即可。
接着,在本申请中,服务器接收到终端发送的请求后,响应该请求,向终端发送终端所需的Tile的编码数据。
如上文所述,在本申请中,服务器将图像内容编码并封装为具有不同编码率的码流,服务器所生成的高质量码流为:将全景视频在空间维度上划分为多个矩形Tile,再对每个Tile在时间维度上划分为多个段,段的时长可以为1秒(可根据实际需求进行划分)。其中,每个Tile在空间上的大小可以相同或不同,Tile中的每个段的时长也可以相同或不同。对每个Tile的每个段进行编码,其中,编码后获取到的每个Tile的编码数据中的每个段的第一图像(需要说明的是,第一图像也可以为位于每个段首位的帧)的预测方式可以为帧内预测。例如:每个段的第一帧可以为I帧,其它帧可以为P帧(即采用帧间预测进行解码的帧)。可选的,在本申请中,服务器在对段中P帧的编码时,可以不参考时域运动矢量,即,高质量码流中除位于每段首位的帧(例如I帧)外,后续的帧 只参考插入帧,即依赖于除了参考位置的帧外,不依赖于前序的帧。
可选的,在本申请中,若应用场景为点播场景,则,服务器将预先对Tile进行编码,以获取Tile的编码数据,并将编码数据封装为码流(或是数据流)。在点播场景中,服务器接收到终端发送的请求消息后,确定终端所需要的Tile对应的编码数据,并将Tile的编码数据发送给终端。
可选的,在本申请中,若应用场景为直播场景,则服务器基于终端发送的请求,确定终端所需的Tile,并对终端所需的Tile进行编码,以获取对应的编码数据,并将Tile的编码数据发送给终端。
对于低质量层的编码、传输及解码方式,可参照已有技术实施例,本申请不再赘述。
步骤102,终端解码获取到的子图像的编码数据。
具体的,在本申请中,终端对接收到的高质量码流(即Tile的编码数据)与低质量码流分别进行解码(低质量层的解码及显示过程可参照已有技术实施例,此处不赘述),并在解码后显示基础层(即低质量层)与高质量层所构成的视频画面。具体的,终端接收到当前视角(即本申请实施例中的第一视角)覆盖的一个或一个以上Tile的编码数据后,对编码数据进行解码与融合,从而获取到Tile(或者可称为Tile所描述的图像序列)。
步骤103,在第一视角切换到第二视角时,终端从服务器端获取第二视角覆盖的子图像的编码数据及对应的参考图像的编码数据。
具体的,在本申请中,终端的视角切换后,终端所需的视频数据(即Tile)将会更新,终端向服务器发送请求消息,以请求新的视角覆盖的Tile的编码数据。举例说明:在t
1时刻,终端显示的第一视角切换为第二视角,则,终端向服务器请求第二视角覆盖的Tile的编码数据,其中,第二视角覆盖的Tile(指本申请实施例中的第二子图像)的播放时刻为t
2时刻,并且,t
1与t
2为相邻的播放时刻。
可选的,在本申请中,第二视角(即切换后的新视角)覆盖的Tile(需要说明的是,新视角覆盖的Tile包括全部在新视角范围内的Tile或者部分在新视角范围内的Tile)可分为两种:一种为新的Tile(即在旧视角中未出现,仅在新视角的覆盖范围内),另一种重叠的Tile(以下简称为重叠Tile,即,包括在旧视角覆盖范围内与新视角覆盖范围内的Tile)。
具体的,在本申请中,视角切换后,终端停止向服务器发送与旧的视角中的旧Tile(不包括重叠Tile)对应的请求消息,但是,终端依旧向服务器发送与重叠Tile对应的请求消息,并将重叠Tile缓存到本地。并且,终端还向服务器发送携带有新的Tile的标识信息的请求消息。可选的,终端还可以向服务器发送携带有与新的Tile对应的参考帧的标识信息的请求消息。即,在本申请中,视角切换后,终端不再获取旧Tile的编码数据,仅获取新的Tile的编码数据及重叠Tile的编码数据。
举例说明:如图6所示,视角1中的高质量层包括Tile1与Tile2,视角未切换前,终端向服务器请求Tile1的编码数据与Tile2的编码数据。视角切换时刻,终端显示的视角切换到视角2,视角2中的高质量层包括Tile2与Tile3。Tile2即为上文所述重叠Tile。其中,终端停止对Tile1的请求(指对Tile1的编码数据的请求),以及,停止对Tile1的编码数据的解码操作。同时,终端继续请求Tile2的编码数据,并将之前请求的编码数 据进行缓存。以及,终端从服务器端获取到Tile3(终端获取到的为Tile3中从切换时刻t
1对应的时间点往后的P帧及后续的Segmenet(即t
2时刻所对应的图像及t
2时刻以后所对应的图像),如图6中的箭头所示)以及与Tile3的编码数据对应的参考图像的编码数据(参考图像的编码数据的生成及传输方式将在下面的实施例中进行详细说明)。
还需要说明的是,在本申请中,可选的,新的Tile中的t
2时刻对应的图像(即新视角内新的Tile的编码数据中位于首位的图像)可以为Tile所在的图像序列所属的GOP中的第一幅图像以外的图像,或者可以理解为,新的Tile中的t
2时刻对应的图像(例如:本申请实施例中的第二子图像)可以其所在码流(或者可以理解为所在的Segment)所描述的图像序列中在解码顺序上的第一幅图像之外的图像,可选的,在本申请中,新的Tile中的t
2时刻对应的图像还可以为其所在码流所描述的图像序列中在解码顺序上的第一幅图像以及第二幅图像之外的图像。举例说明:t
2时刻对应的图像可以为其所属Segment中除第一帧以外的其它帧对应的图像。或者,新的Tile中的t
2时刻对应的图像还可以为Tile所在的图像序列所述的GOP中的第一幅图像和第二幅图像以外的图像,即,参照图6,新视角内的Tile3在新视角内播放的起始帧(即实心黑色框)可以为所在Segment的原第一帧与第二帧以外的任意一帧,举例说明:Tile3的Segment1在编码后其所描述的图像序列包括数据帧1、数据帧2、数据帧3、数据帧4所对应的图像。在切换时刻,位于视角播放时刻的起始帧可以为数据帧2或数据帧3或数据帧4,本申请不做限定。
具体的,在本申请中,服务器接收终端的请求,并向终端发送相应的Tile的编码数据及参考图像的编码数据。
可选的,在本申请中,终端可不感知视角是否切换,即,终端向服务器发送的请求消息中可携带有所需Tile的标识信息。服务器接收到请求消息后,由服务器端基于请求消息中携带的标识信息,判断终端当前是否发生视角切换,若发生视角切换,则服务器向终端发送新视角覆盖的Tile(包括新Tile及重叠Tile,新Tile及重叠Tile的定义如前所述,此处不赘述)。
可选的,在本申请中,终端可检测到视角是否发生切换,若视角切换,则终端向服务器发送的请求消息中可携带有新视角覆盖的Tile的标识信息以及与新视角覆盖的部分或全部Tile对应的参考图像的标识信息。可选的,在本申请中,参考图像可以称为参考帧。可选的,在本申请中,参考帧(即参考图像)的标识信息可以为参考帧的帧号或时间戳等标识方式。
下面对参考帧的生成方式进行详细阐述。具体的,在本申请中,服务器端基于全景视频,除生成如前所述的高质量码流、低质量码流外,还可生成包含的每一图像(或者称为每一帧)都可以进行独立解码的码流(以下简称为参考码流)。其中,独立解码即为帧内所有图像块均采用帧内预测进行解码,也就是说,参考码流为中其所描述的所有图像帧的所有图像块执行帧内预测所得到的码流,以及,在本申请中,参考码流独立于高质量码流。可选的,在本申请中,参考码流可以为RARF流。可选的,在本申请中,参考码流中的每一帧可以为纯随机接入(Clean Random Access,CRA)型I帧。可选的,在本申请中,参考码流中的每一帧也可以为P帧,其中,P帧中均为I块(即,包含I块的P帧仍为帧内参考预测进行解码的帧,对应的帧类型可以称为P slice(条带))。
可选的,在本申请中,参考码流对应的图像内容与对应的Tile(即可基于参考码流中的参考帧进行解码的Tile,例如本申请实施例中的第二子图像)所对应的图像内容部分相同或全部相同。举例说明:参考码流(以下简称为RARF流)作为高质量层的Tile的参考帧的方式如图7所示,在图7中,视角切换后,视角2中的Tile n为新的Tile,终端在P
2帧对应的时刻(即上文所述t时刻)进行视角切换,在已有技术中,Tile n的P
3帧(即上文所述t
2时刻对应的图像)进行帧间预测解码时所依赖的参考帧为P
2帧,但是,由于终端在P
2帧时刻进行切换,终端中不存在P
2帧(原因为终端从服务器端获取到t
2时刻及之后所对应的图像),则,在本申请中,RARF流中的帧(即为图中所示的RARF2帧)可作为P
3帧的参考帧,从而使终端可基于RARF帧对P
3帧进行解码,并可继续对该Segment中的P帧依次进行解码。也就是说,服务器可以向终端发送一个或一个以上RARF帧,以使终端基于RARF帧作为Segment的参考帧,对Segment中的多个帧进行解码。需要说明的是,在本申请中,RARF中的参考帧(例如RARF2帧)的图像内容与新Tile中t
2时刻对应的帧解码所依赖的参考帧(例如P
2帧)的图像内容相同。还需要说明的是,如上文所述,t
2时刻的图像可以为GOP中的除第一幅图像,或者是除第一幅图像与第二幅图像(或者称为Segment中的第一帧)以外的任一图像,因此,在本申请中,可选的,RARF参考帧的图像内容可以与第一帧(其中,第一帧的解码类型可以为帧内解码或帧间解码)的图像内容一致,其中,RARF参考帧的解码类型为帧内解码,在该情况下,RARF参考帧可以作为第二帧的参考帧,并进行解码。可选的,参考帧的图像内容可以与第二帧或第三帧等解码类型为帧间解码的帧的图像内容一致,即如上述举例所述。
可选的,在本申请中,如上所述,RARF参考帧的图像内容可以与t
2时刻的帧(即本申请实施例中的第二子图像)进行解码所依赖的参考帧的图像内容相同。可选的,在本申请中,第二子图像解码所依赖的参考帧可以为t
1时刻,即切换前一时刻对应的帧,也就是说,RARF帧对应的图像内容可以与新的Tile中与切换时刻的前一时刻对应的帧的图像内容相同。例如:仍参照图7,终端在Tile n的P
2帧对应的时刻进行视角切换,则,RARF帧对应的图像内容可以与P
2帧对应的图像内容相同,则,P
3帧可基于RARF帧进行解码,需要说明的是,在该实施例中,终端的第二视角在t
2时刻播放的为P
3帧所描述的图像。可选的,RARF帧对应的图像内容还可以为切换时刻之前的第n个帧的图像内容相同。例如:在图7中,P
2帧对应的图像内容还可以与P
1帧的图像内容相同。
可选的,在本申请中,终端的第二视角在t
2时刻播放的还可以为RARF帧所描述的图像。举例说明:仍以图7为例,终端在Tile n的P
2帧对应的时刻进行视角切换,则,RARF帧对应的图像内容可以与P
3帧对应的图像内容相同,则,终端对RARF帧进行解码后,在t
2时刻播放的是RARF帧(或P
3帧)所描述的图像,接着,P
4帧可继续以RARF帧为参考帧进行解码。
在本申请中,对于生成参考码流的时机,可选的,服务器可预先生成三种码流:低质量码流、高质量码流以及与高质量码流对应的RARF流。在该实施例中,服务器可基于接收到的请求消息,向终端发送对应的低质量码流、高质量码流中的Tile以及对应的参考帧。举例说明:在直播场景中,服务器每次接收到终端的请求消息,均基于请求消 息生成与所需视频数据对应的低质量码流、高质量码流以及RARF流,并将低质量码流、高质量码流发送给终端。终端的视角切换后,切换后的视角的播放时刻t
2对应的图像为Segment的图像序列中的P
3帧所对应的图像。终端向服务器发送携带有Tile n的P
3帧的帧号的请求消息。服务器接收到请求消息后,确定终端发生视角切换,则,服务器基于请求消息,生成Tile n(生成的Tile n从P
3帧开始,可参照图7)以及生成P
3帧的RARF帧,即可供P
3帧的编码数据参考以进行解码的RARF帧的编码数据。随后,服务器将Tile n中从P
3帧起始的Segment1以及对应的RARF帧发送给终端。或者,终端还可以像服务器请求与P
3帧的图像内容相同的RARF帧以及P
4帧及其后的P帧,相应的,服务器基于请求消息,生成RARF帧的编码数据及P
4帧的编码数据等,并将上述帧的编码数据封装为Segment,发送给终端。
可选的,在点播场景中,服务器可基于视频数据生成对应的高质量码流、低质量码流以及RARF流,并存储在服务器本地。随后,服务器可基于接收到的终端的请求消息,向终端发送对应的低质量码流、高质量码流中的Tile。同样,若服务器检测到终端发生视角切换,则,服务器可将新视角上对应的高质量码流中的Tile以及对应的RARF帧发送给终端。
可选的,在本申请中,为降低服务器存储的开销,服务器还可以在检测到终端视角切换,或者,接收到携带有RARF流的请求消息后,再生成对应的RARF帧。举例说明:在视角为切换前,在直播场景中,服务器基于终端的请求消息,生成对应的高质量Tile的编码数据,并发送给终端。终端的视角切换后,终端向服务器发送携带有Tile n的P
3帧的帧号的请求消息。服务器接收到请求消息后,确定终端发生视角切换,则,服务器基于请求消息,生成Tile n(生成的Tile n从P
3帧开始,可参照图7)以及生成P
3帧的RARF帧,即可供P
3帧参考,以进行解码的RARF帧。随后,服务器将Tile n中从P
3帧起始的Segment1以及对应的RARF帧发送给终端。
可选地,在本申请中,如前所述,终端在请求新的Tile的编码数据的同时,还会继续请求出现在新的视角中的旧Tile(即上文所述的重叠Tile)的编码数据。因此,服务器向终端发送新视角对应的新Tile的编码数据的同时,还会继续向终端发送重叠Tile的编码数据。举例说明:在1分30秒时,终端向服务器发送携带有Tile m的I
0帧的帧号,即终端向服务器请求Tile m的第一个Segment(假设每个Segment时长为3秒)的编码数据,随后,在1分32秒时(即为t
1时刻),终端视角切换,则,终端向服务器请求Tile n的P
15帧(即为t
2时刻对应播放的帧)(属于Tile n的第一个Segment)及P
15帧对应的RARF帧。服务器接收到请求后,生成Tile n及对应的RARF帧(需要说明的是,在该实施例中,生成的RARF帧为与P
15帧对应的参考帧,即,客户端可基于一个RARF帧对P
15进行解码),并从P
15帧开始向终端传输Tile n和对应于P
15帧的RARF帧。与此同时,服务器会继续向终端发送Time m的第一个Segment中的剩余帧,并且在1分33秒时,终端会继续向服务器请求Tile m的第二个Segment的编码数据,服务器响应终端的请求,向终端发送Tile m的其它Segment的编码数据。
可选地,在本申请中,服务器向终端返回响应消息,响应消息中除携带有Tile的标识信息以外,还携带有用于指示各Tile在图像数据或者视频图像中的实际位置信息。
步骤104,终端基于参考图像的编码数据,解码第二视角中的子图像。
具体的,在本申请的实施例中,终端获取服务器发送的高质量码流中的Tile的编码数据(具体可以Tile中Segment中的每个图像帧的编码数据)、参考图像的编码数据(例如RARF参考帧的编码数据),以及各个Tile的位置信息(终端从服务器端获取的位置信息为Tile在画面中的实际位置信息)。接着,终端可解码RARF参考帧的编码数据,以获取到RARF参考帧。终端可基于解码后的RARF参考帧对新视角中的新Tile进行解码。并且,终端还可以基于已接收到的参考帧(I帧和/或P帧)对重叠Tile继续进行解码。
具体的,在本申请中,如图8所示,在图8中,视角1(旧视角)中的Tile(包括Tile1、Tile2、Tile3、Tile4、Tile8、Tile9、Tile10、Tile11),视角2(新视角)中的Tile(包括Tile4、Tile5、Tile6、Tile7、Tile11、Tile12、Tile13、Tile14)。其中,Tile4与Tile11即为重叠Tile。需要说明的是,图8中所示的Tile的数量仅为适宜性举例,例如:重叠Tile可以为一个或一个以上(例如5个),具体数量根据视角大小以及旋转角度及速度等确定。
参照图8,终端对码流(包括新Tile以及重叠Tile的编码数据)进行解码的过程如下:
1)终端将新Tile的编码数据与重叠Tile的编码数据进行融合。
具体的,在本申请中,如图8所示,重叠Tile:Tile4与Tile11在旧视角中进行融合时,其在Tile拼接(Tile Merge,或者是Tile重写(bitstream rewrite))模块(或可称为码流拼接模块或区块拼接模块)中的位置如图8中所示,即,Tile4与Tile11位于Tile Merge的最右侧。需要说明的是,在已有技术中,按照新视角中的视频数据对应的Tile的位置在Tile Merge中放置待融合的Tile,则,Tile4和Tile11可以位于新视角对应的Tile Merge中的最左侧,然而,Tile4的编码数据与Tile11的编码数据在解码时所依赖的参考帧仍然位于Tile Merge模块的最右侧,因此,已有技术中的Tile4与Tile11由于找不到参考帧而无法解码。
继续参照图8,可选的,在本申请的实施例中,Tile4与Tile11仍然可放置于Tile Merge模块的最右侧,即,保持原位置不变。则,Tile4的编码数据与Tile11的编码数据可基于已接收到的参考帧(I帧或P帧)进行解码。
可选的,在本申请中,终端将新Tile(包括Tile5、Tile6、Tile7、Tile12、Tile13、Tile14)放置在Tile Merge中除重叠Tile以外的空位置上。放置的位置可以如图8所示,需要说明的是,图8中新Tile的放置方式仅为适宜性举例,新Tile可以放置于除重叠Tile以外的任意空位置上。随后,终端对Tile Merge中的多个Tile进行融合,具体融合方式可参照已有技术实施例中的技术方案,本申请不再赘述。
可选的,在本申请中,终端在本地缓存中存储Tile的实际位置信息以及融合位置信息。可选的,Tile在Tile Merge中的融合位置(即为Tile在Tile Merge中的放置位置),以及Tile在视频图像中的实际位置(即步骤105中所述的服务器发送的位置信息)可以用数组的方式表示。举例说明:以图8中所示的各个Tile的编码数据在Tile Merge中的融合位置为例,设置两个8位的数组逐行记录TileMerge中各Tile的位置信息,其中, 数组1用于表征Tile在画面中的真实位置,数组2用于表征Tile的融合位置信息。具体的,视角1中两个数组中的信息均为[1,2,3,4,8,9,10,11]。当视角切换为转向视角2时,数组1中的信息更新为[4,5,6,7,11,12,13,14],以及,数组2中的信息为[5,6,7,4,12,13,14,11]。可选的,上述数组也可以由服务器进行维护,即,服务器实时根据Tile在终端中的融合位置,以及Tile的实际位置生成对应的数组,并在客户端进行视角切换时,发送给客户端。
2)对融合后的Tile的编码数据进行解码。
具体的,终端对融合后的Tile的编码数据,通过单解码器进行解码,以获取Tile。具体解码的细节可参照已有技术实施例中的技术方案,本申请不再赘述。
3)按照记录的位置信息(包括实际位置信息以及融合位置信息),对Tile进行投影。
具体的,在本申请中,终端可基于融合位置信息以及实际位置信息,将各个Tile进行投影,以获取将要显示的视频图像。举例说明:终端中存储的视角2对应的数组2为:[5,6,7,4,12,13,14,11],基于数组2,并参考数组1:[4,5,6,7,11,12,13,14],则投影时,将Tile4与Tile11放置如数组1表征的位置即可,随后终端将投影后的Tile在终端的屏幕中显示。
综上所述,本申请实施例中的技术方案,通过设置每一帧都均为帧内解码帧的参考码流作为高质量码流的参考帧,使终端进行视角切换后,能够基于参考帧对视角中的新Tile进行解码,无需备份大量数据即可实现码流的快速切换,显著降低了服务器端的存储开销。以及,在解码过程中,重叠Tile可在原有位置上继续进行解码,从而减少了对终端的对已下载码流的重复下载,进而降低了带宽开销,提高了资源利用率。
可选的,在本申请中,视角切换后,终端可以向服务器请求新Tile、与新Tile对应的RARF帧(即与新Tile位于首位的帧对应的RARF帧),以及,重叠Tile以及与重叠Tile对应的RARF帧(与重叠Tile位于首位的帧对应的RARF帧)。或者,服务器在检测到视角切换后,也可以向终端发送新Tile、与新Tile对应的RARF帧,以及,重叠Tile以及与重叠Tile对应的RARF帧。则,在该实施例中,重叠Tile与新Tile在Tile Merge中的位置均可以任意放置,在解码过程中,终端可基于重叠Tile对应的RARF帧,对重叠Tile位于首位的帧进行解码,以及,基于新Tile对应的RARF帧,对新Tile进行解码。并基于位置信息,将各个Tile进行重组,以获取源视频图像。
场景二
结合图4,如图9所示为本申请实施例中的视频处理方法的流程示意图,在图9中:
步骤201,终端从服务器端获取第一视角覆盖的子图像的编码数据及对应的参考图像的编码数据。
具体的,在本申请的实施例中,终端向服务器发送基于用户指令生成的请求消息。请求消息可用于向服务器请求终端所显示的低质量层以及终端的视角范围内的高质量层。
可选地,在本申请中,服务器可将图像编码并封装为高质量码流、低质量码流及RARF流。可选的,在本申请中,服务器所生成的高质量码流为:将全景视频在空间维度上划分为多个矩形Tile,再对每个Tile在时间维度上划分为多个段,段的时长可以为1秒(可 根据实际需求进行划分)。其中,每个Tile在空间上的大小可以相同或不同,Tile中的每个段的时长也可以相同或不同。每个Tile的每个段中的图像帧的预测方式均为帧间预测,即,Segment为其所描述的所有图像帧执行帧间预测得到的码流,例如:Segment中的所有图像帧均可以为P帧。
可选的,在本申请中,服务器将高质量码流逐帧按照Tile对应的视频画面编码为RARF码流,要是每段的第一个RARF帧作为高质量码流的每段的参考帧。以及,服务器向终端发送的视屏数据对应的码流包括:低质量码流、高质量码流,RARF码流中的每段的第一个RARF帧。
可选的,在点播场景中,服务器预先将视频数据进行编码,生成不同质量的码流,其中,包括但不限于:高质量码流、低质量码流及RARF流。服务器接收终端发送的请求消息,确定终端所需要的视频数据(所需要的数据可以为当前所要显示的数据和/或预测的视频数据),并向终端发送视频数据所对应的低质量码流、与视频数据对应的高质量码流中的Tile、以及与高质量码流对应的RARF流中的每段的第一个RARF帧。其中,发送的高质量码流中的Tile为基于终端发送的请求消息中携带的标识信息确定的。
可选的,在直播场景中,服务器接收到终端的请求消息后,确定终端所需的视频数据,并对终端所需的视频数据进行编码,获取编码后的低质量码流、高质量码流及RARF流。接着,服务器将低质量码流、与视频数据对应的高质量码流中的Tile、以及与高质量码流对应的RARF流中的每段的第一个RARF帧发送给终端。其中,发送的高质量码流中的Tile为基于终端发送的请求消息中携带的标识信息确定的。
可选的,服务器在生成RARF帧的过程中,可以仅生成用于作为高质量码流的参考帧的RARF帧,即仅生成每段的第一个RARF帧,从而降低开销。
其它细节可参照步骤101,此处不赘述。
步骤202,终端基于参考图像的编码数据,解码获取到的子图像的编码数据。
具体的,终端对低质量码流与高质量码流分别进行解码,以获取与低质量码流对应的基础层(或可称为低质量层),以及与高质量码流对应的高质量层。
在本申请中,终端接收到多个高质量码流对应的Tile,高质量码流将多个Tile进行拼接融合,以实现使用单解码器对融合后的多个Tile进行解码。接着,终端可基于低质量码流的I帧对低质量码流进行解码。以及,终端基于接收到的RARF帧对高质量码流进行解码。例如:终端接收到RARF n流的第一个Segment的第一个RARF帧(简称RARF1帧)的编码数据,其中,RARF n流与Tile n对应的图像内容相同,即,RARF n为按照Tile n逐帧编码生成的,则,终端将RARF1帧作为Tile n的第一个Segment的参考帧,对第一个Segment的图像序列中的第一图像进行解码。具体融合以及解码过程可参照已有技术实施例中的技术方案,本申请不再赘述。
步骤203,在第一视角切换到第二视角时,终端从服务器端获取第二视角覆盖的子图像的编码数据及对应的参考图像的编码数据。
具体的,在本申请的实施例中,终端的视角切换后,终端所需的视频数据需要更新,终端向服务器发送请求消息,以请求低质量层以及新视角上的(包括新视角内以及新视角外的指定范围内的部分)视频数据。
具体细节参照步骤103,此处不赘述。
步骤204,终端基于参考图像的编码数据,解码第二视角中的子图像。
具体的,在本申请中,服务器向终端发送新视角覆盖的低质量码流、高质量码流以及RARF帧的编码数据(需要说明的是,在本实施例中,所述RARF帧是指作为新视角中的Tile中位于首位的帧(位于首位的帧是指新视角中的Tile的Segment与视角切换时刻对应的帧)的参考帧的RARF帧)。
可选的,在本申请中,服务器继续依据终端的请求,向终端发送重叠Tile中的Segment,以及,基于终端的请求,向终端发送新的Tile的编码数据以及与新的Tile对应的RARF帧的编码数据(需要说明的是,发送的RARF帧的数量可以为一个或一个以上,其中,若发送一个RARF帧,则该RARF帧如上所述,为用于对Tile中位于首位的数据帧进行解码的参考帧)。
可选地,在本申请中,服务器向终端返回响应消息,响应消息中除携带有Tile的标识信息以外,还携带有用于指示各Tile在视频图像中的实际位置信息。其中,实际位置信息可以用数组表示,具体表示方式将在下面的实施例中举例说明。
接着,在本申请中,终端获取服务器发送的高质量码流中的Tile的编码数据、对应的参考帧,以及各个Tile的位置信息。接着,终端可基于RARF帧,对新视角中的新Tile的Segment中位于首位的图像帧的编码数据进行解码。并且,终端还可以基于已接收到的RARF帧的编码数据,对重叠Tile中位于首位的图像帧进行解码,并继续对其它帧进行解码。
具体细节可参照步骤104,此处不赘述。
综上所述,本申请实施例中的技术方案,由于高质量码流没有I帧,从而进一步降低了服务器端的存储开销。
可选的,在本申请中,视角切换后,终端可以向服务器请求新Tile、与新Tile对应的RARF帧,以及,重叠Tile以及与重叠Tile对应的RARF帧。或者,服务器在检测到视角切换后,也可以向终端发送新Tile、与新Tile对应的RARF帧,以及,重叠Tile以及与重叠Tile对应的RARF帧。则,在该实施例中,重叠Tile与新Tile在Tile Merge中的位置均可以任意放置,在解码过程中,终端可基于重叠Tile对应的RARF帧,对重叠Tile进行解码,以及,基于新Tile对应的RARF帧,对新Tile进行解码。并基于位置信息,将各个Tile进行映射,以获取将要显示的视频图像。
可选的,可应用于场景一与场景二的实施例中,在本申请中,高质量码流还可以进一步包括不同质量等级(即不同编码率)的多个码流。举例说明:高质量码流可以包括质量等级为A、质量等级为B以及质量等级为C的三种高质量码流。其中,质量等级A高于质量等级B高于质量等级C。
可选的,在本申请中,服务器可基于外部条件,确定向终端发送的Tile对应的质量等级。外部条件包括但不限于:终端的处理能力、服务器与终端之间的传输带宽大小等。举例说明:若当前终端的负载较大,或带宽较小,则,服务器可向终端传输质量等级为C的码流对应的Tile的编码数据。反之,若外部条件良好,则,服务器可向终端传输质 量等级为A的码流对应的Tile的编码数据。通常情况下,服务器可向终端传输质量等级A、质量等级B以及质量等级C中的一个或一个以上码流。举例说明:以包含人的脸部图像的视频图像为例,则,服务器可将人的脸部图像生成质量等级为A的码流发送给终端,并将视角范围内除脸部图像以外的其它图像生成质量等级为B和/或C的码流发送给终端,从而使终端能够更加清晰的显示人脸部图像。
上述主要从各个网元之间交互的角度对本申请实施例提供的方案进行了介绍。可以理解的是,视频处理装置(包括终端与服务器端)为了实现上述功能,其包含了执行各个功能相应的硬件结构和/或软件模块。本领域技术人员应该很容易意识到,结合本文中所公开的实施例描述的各示例的单元及算法步骤,本申请实施例能够以硬件或硬件和计算机软件的结合形式来实现。某个功能究竟以硬件还是计算机软件驱动硬件的方式来执行,取决于技术方案的特定应用和设计约束条件。专业技术人员可以对每个特定的应用来使用不同方法来实现所描述的功能,但是这种实现不应认为超出本申请的范围。
本申请实施例可以根据上述方法示例对视频处理装置进行功能模块的划分,例如,可以对应各个功能划分各个功能模块,也可以将两个或两个以上的功能集成在一个处理模块中。上述集成的模块既可以采用硬件的形式实现,也可以采用软件功能模块的形式实现。需要说明的是,本申请实施例中对模块的划分是示意性的,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式。
在采用对应各个功能划分各个功能模块的情况下,在采用对应各个功能划分各个功能模块的情况下,图10示出了上述实施例中所涉及的视频处理装置300(例如终端200)的一种可能的结构示意图,如图10所示,视频处理装置可以包括:解码模块301、获取模块302。其中,解码模块301可用于“解码第一视角覆盖的第一子图像的编码数据,从而得到所述第一子图像”的步骤。例如:该模块可支持视频处理装置300执行上述方法实施例中的步骤102、步骤104。获取模块302可用于“在所述第一视角切换到第二视角时,获取所述第二视角覆盖的第二子图像的编码数据”的步骤。例如:该模块可支持视频处理装置300执行上述方法实施例中的步骤步骤103。以及,解码模块301还可以用于“解码所述第一参考图像的编码数据获得第一参考图像”以及“根据所述第一参考图像解码所述第二子图像的编码数据,从而获得所述第二子图像”的步骤。例如:该模块可支持视频处理装置300执行上述方法实施例中的步骤步骤104。
在本申请中,图11示出了上述实施例中所涉及的视频处理装置400(例如服务器100)的一种可能的结构示意图,如图11所示,视频处理装置400可以包括:发送模块401。其中,发送模块401可用于“向终端发送所述终端的第一视角覆盖的第一子图像的编码数据”以及“在所述终端的第一视角切换到第二视角时,向所述终端发送所述第二视角覆盖的第二子图像的编码数据”的步骤。
下面介绍本申请实施例提供的一种装置。如图12所示:
该装置包括处理模块501和通信模块502。可选的,该装置还包括存储模块503。处理模块501、通信模块502和存储模块503通过通信总线相连。
通信模块502可以是具有收发功能的装置,用于与其他网络设备或者通信网络进行通信。
存储模块503可以包括一个或者多个存储器,存储器可以是一个或者多个设备、电路中用于存储程序或者数据的器件。
存储模块503可以独立存在,通过通信总线与处理模块501相连。存储模块也可以与处理模块501集成在一起。
装置500可以用于网络设备、电路、硬件组件或者芯片中。
装置500可以是本申请实施例中的终端,例如终端200。可选的,装置500的通信模块502可以包括终端的天线和收发机。
装置500可以是本申请实施例中的终端中的芯片。通信模块502可以是输入或者输出接口、管脚或者电路等。可选的,存储模块可以存储终端侧的方法的计算机执行指令,以使处理模块501执行上述实施例中终端侧的方法。存储模块503可以是寄存器、缓存或者RAM等,存储模块503可以和处理模块501集成在一起;存储模块503可以是ROM或者可存储静态信息和指令的其他类型的静态存储设备,存储模块503可以与处理模块501相独立。
当装置500是本申请实施例中的终端或者终端中的芯片时,装置500可以实现上述实施例中终端执行的方法。
装置500可以是本申请实施例中的服务器,例如服务器100。
装置500可以是本申请实施例中的服务器中的芯片。通信模块502可以是输入或者输出接口、管脚或者电路等。可选的,存储模块可以存储服务器侧的方法的计算机执行指令,以使处理模块501执行上述实施例中服务器侧的方法。存储模块503可以是寄存器、缓存或者RAM等,存储模块503可以和处理模块501集成在一起;存储模块503可以是ROM或者可存储静态信息和指令的其他类型的静态存储设备,存储模块503可以与处理模块501相独立。
当装置500是本申请实施例中的服务器或者服务器中的芯片时,可以实现上述实施例中服务器执行的方法。
本申请实施例还提供了一种计算机可读存储介质。上述实施例中描述的方法可以全部或部分地通过软件、硬件、固件或者其任意组合来实现。如果在软件中实现,则功能可以作为一个或多个指令或代码存储在计算机可读介质上或者在计算机可读介质上传输。计算机可读介质可以包括计算机存储介质和通信介质,还可以包括任何可以将计算机程序从一个地方传送到另一个地方的介质。存储介质可以是可由计算机访问的任何可用介质。
作为一种可选的设计,计算机可读介质可以包括RAM,ROM,EEPROM,CD-ROM或其它光盘存储器,磁盘存储器或其它磁存储设备,或可用于承载的任何其它介质或以指令或数据结构的形式存储所需的程序代码,并且可由计算机访问。而且,任何连接被适当地称为计算机可读介质。例如,如果使用同轴电缆,光纤电缆,双绞线,数字用户 线(DSL)或无线技术(如红外,无线电和微波)从网站,服务器或其它远程源传输软件,则同轴电缆,光纤电缆,双绞线,DSL或诸如红外,无线电和微波之类的无线技术包括在介质的定义中。如本文所使用的磁盘和光盘包括光盘(CD),激光盘,光盘,数字通用光盘(DVD),软盘和蓝光盘,其中磁盘通常以磁性方式再现数据,而光盘利用激光光学地再现数据。上述的组合也应包括在计算机可读介质的范围内。
本申请实施例还提供了一种计算机程序产品。上述实施例中描述的方法可以全部或部分地通过软件、硬件、固件或者其任意组合来实现。如果在软件中实现,可以全部或者部分得通过计算机程序产品的形式实现。计算机程序产品包括一个或多个计算机指令。在计算机上加载和执行上述计算机程序指令时,全部或部分地产生按照上述方法实施例中描述的流程或功能。上述计算机可以是通用计算机、专用计算机、计算机网络、网络设备、用户设备或者其它可编程装置。
结合本申请实施例公开内容所描述的方法或者算法的步骤可以硬件的方式来实现,也可以是由处理器执行软件指令的方式来实现。软件指令可以由相应的软件模块组成,软件模块可以被存放于随机存取存储器(Random Access Memory,RAM)、闪存、只读存储器(Read Only Memory,ROM)、可擦除可编程只读存储器(Erasable Programmable ROM,EPROM)、电可擦可编程只读存储器(Electrically EPROM,EEPROM)、寄存器、硬盘、移动硬盘、只读光盘(CD-ROM)或者本领域熟知的任何其它形式的存储介质中。一种示例性的存储介质耦合至处理器,从而使处理器能够从该存储介质读取信息,且可向该存储介质写入信息。当然,存储介质也可以是处理器的组成部分。处理器和存储介质可以位于ASIC中。另外,该ASIC可以位于网络设备中。当然,处理器和存储介质也可以作为分立组件存在于网络设备中。
本领域技术人员应该可以意识到,在上述一个或多个示例中,本申请实施例所描述的功能可以用硬件、软件、固件或它们的任意组合来实现。当使用软件实现时,可以将这些功能存储在计算机可读介质中或者作为计算机可读介质上的一个或多个指令或代码进行传输。计算机可读介质包括计算机存储介质和通信介质,其中通信介质包括便于从一个地方向另一个地方传送计算机程序的任何介质。存储介质可以是通用或专用计算机能够存取的任何可用介质。
上面结合附图对本申请的实施例进行了描述,但是本申请并不局限于上述的具体实施方式,上述的具体实施方式仅仅是示意性的,而不是限制性的,本领域的普通技术人员在本申请的启示下,在不脱离本申请宗旨和权利要求所保护的范围情况下,还可做出很多形式,均属于本申请的保护之内。
Claims (28)
- 一种视频处理方法,其特征在于,所述方法包括:解码第一视角覆盖的第一子图像的编码数据,从而得到所述第一子图像;在所述第一视角切换到第二视角时,获取所述第二视角覆盖的第二子图像的编码数据,以及第二子图像的第一参考图像的编码数据,所述第一参考图像的编码数据独立于所述第二子图像的编码数据所在的码流;解码所述第一参考图像的编码数据获得第一参考图像;根据所述第一参考图像解码所述第二子图像的编码数据,从而获得所述第二子图像。
- 根据权利要求1所述的方法,其特征在于,所述第二子图像的编码数据所在的码流为所述第二子图像的编码数据所在的段(Segment)码流。
- 根据权利要求1或2所述的方法,其特征在于,所述第二子图像为所述第二子图像的编码数据所在的码流所描述的图像序列中在解码顺序上的第一幅图像之外的图像,或者所述第二子图像为所述第二子图像的编码数据所在的码流所描述的图像序列中在解码顺序上的第一幅图像以及第二幅图像之外的图像。
- 根据权利要求1至3任一项所述的方法,其特征在于,所述第一参考图像的图像内容与所述第二子图像的第二参考图像的图像内容相同,所述第二参考图像为所述第二子图像的编码数据所在的码流所描述的图像序列中的图像。
- 根据权利要求4所述的方法,其特征在于,所述第二参考图像为所述第二子图像的编码数据所在的码流所描述的图像序列中在解码顺序上的第一幅图像之外的图像。
- 根据权利要求5所述的方法,其特征在于,所述第二参考图像为所述第二子图像在所述图像序列中的前一幅图像。
- 根据权利要求1至6任一项所述的方法,其特征在于,所述第二子图像的编码数据所在的码流为通过对码流中所述描述的所有图像帧执行帧间预测得到的码流。
- 根据权利要求1至7任一项所述的方法,其特征在于,所述第一参考图像的编码数据所在的码流为通过对码流中所述描述的所有图像帧的所有图像块执行帧内预测得到的码流,所述第一参考图像的编码数据所在的码流独立于所述第二子图像的编码数据所在的码流。
- 根据权利要求8所述的方法,其特征在于,所述第一参考图像为纯随机接入CRA图像。
- 根据权利要求1至9任一项所述的方法,其特征在于,所述第二子图像的不被所述第一视角覆盖。
- 根据权利要求1至10任一项所述的方法,其特征在于,所述第二视角覆盖所述第一子图像区域,所述第一视角还覆盖第三子图像区域,所述解码第一视角覆盖的第一子图像的编码数据,从而得到所述第一子图像包括:将第一码流和第三码流合并,从而得到第一合并码流,所述第一码流为所述第一子图像的编码数据所在的码流,所述第三码流为所述第三子图像的编码数据所在的码流;解码所述第一合并码流,从而得到包括第一子图像和所述第三子图像的图像;所述根据所述第一参考图像解码所述第二子图像的编码数据,从而获得所述第二子图像包括:将所述第一码流与第二码流合并,从而得到第二合并码流,所述第二码流为所述第二子图像的编码数据所在的码流,所述第一码流所描述的子图像对应在所述第一合并码流所述描述的图像中的位置,与所述第一码流所描述的子图像对应在所述第二合并码流所述描述的图像中的位置一致;根据所述第一参考图像解码所述第一合并码流,从而得到包括第二子图像和所述第一码流所描述的子图像的图像。
- 一种视频处理方法,其特征在于,所述方法包括:向终端发送所述终端的第一视角覆盖的第一子图像的编码数据;在所述终端的第一视角切换到第二视角时,向所述终端发送所述第二视角覆盖的第二子图像的编码数据,以及第二子图像的第一参考图像的编码数据,所述第一参考图像的编码数据独立于所述第二子图像的编码数据所在的码流。
- 根据权利要求12所述的方法,其特征在于,所述第二子图像的编码数据所在的码流为所述第二子图像的编码数据所在的段码流。
- 根据权利要求12或13所述的方法,其特征在于,所述第二子图像为所述第二子图像的编码数据所在的码流所描述的图像序列中在解码顺序上的第一幅图像之外的图像,或者所述第二子图像为所述第二子图像的编码数据所在的码流所描述的图像序列中在解码顺序上的第一幅图像以及第二幅图像之外的图像。
- 根据权利要求12至14任一项所述的方法,其特征在于,所述第一参考图像的图像内容与所述第二子图像的第二参考图像的图像内容相同,所述第二参考图像为所述第二子图像的编码数据所在的码流所描述的图像序列中的图像。
- 根据权利要求15所述的方法,其特征在于,所述第二参考图像为所述第二子图 像的编码数据所在的码流所描述的图像序列中在解码顺序上的第一幅图像之外的图像。
- 根据权利要求16所述的方法,其特征在于,所述第二参考图像为所述第二子图像在所述图像序列中的前一幅图像。
- 根据权利要求12至17任一项所述的方法,其特征在于,所述第二子图像的编码数据所在的码流为通过对码流中所述描述的所有图像帧的执行帧间预测得到的码流。
- 根据权利要求12至18任一项所述的方法,其特征在于,所述第一参考图像的编码数据所在的码流为通过对码流中所述描述的所有图像帧的所有图像块执行帧内预测得到的码流,所述第一参考图像的编码数据所在的码流独立于所述第二子图像的编码数据所在的码流。
- 根据权利要求19所述的方法,其特征在于,所述第一参考图像为纯随机接入CRA图像。
- 根据权利要求12至20任一项所述的方法,其特征在于,所述第二子图像不被所述第一视角覆盖。
- 一种视频处理装置,其特征在于,所述装置包括:解码模块,用于解码第一视角覆盖的第一子图像的编码数据,从而得到所述第一子图像;获取模块,用于在所述第一视角切换到第二视角时,获取所述第二视角覆盖的第二子图像的编码数据,以及第二子图像的第一参考图像的编码数据,所述第一参考图像的编码数据独立于所述第二子图像的编码数据所在的码流;所述解码模块进一步用于解码所述第一参考图像的编码数据获得第一参考图像;所述解码模块进一步用于根据所述第一参考图像解码所述第二子图像的编码数据,从而获得所述第二子图像。
- 一种视频处理装置,其特征在于,所述装置包括:发送模块,用于向终端发送所述终端的第一视角覆盖的第一子图像的编码数据;所述发送模块进一步用于在所述终端的第一视角切换到第二视角时,向所述终端发送所述第二视角覆盖的第二子图像的编码数据,以及第二子图像的第一参考图像的编码数据,所述第一参考图像的编码数据独立于所述第二子图像的编码数据所在的码流。
- 一种视频处理系统,其特征在于,包括服务器端与终端,其中,所述服务器端向所述终端发送所述终端的第一视角覆盖的第一子图像的编码数据;所述终端接收所述第一子图像的编码数据,并解码所述第一子图像的编码数据,得 到所述第一子图像;在所述终端的第一视角切换到第二视角时,所述服务器端向所述终端发送所述第二视角覆盖的第二子图像的编码数据,以及第二子图像的第一参考图像的编码数据,所述第一参考图像的编码数据独立于所述第二子图像的编码数据所在的码流;所述终端接收所述第二子图像的编码数据以及所述第一参考图像的编码数据,并解码所述第一参考图像的编码数据获得第一参考图像;以及,所述终端根据所述第一参考图像解码所述第二子图像的编码数据,从而获得所述第二子图像。
- 一种计算机可读存储介质,所述计算机可读存储介质存储有计算机程序,所述计算机程序包含至少一段代码,该至少一段代码可由计算机执行,以控制所述计算机执行权利要求1至11任一项所述的方法。
- 一种计算机可读存储介质,所述计算机可读存储介质存储有计算机程序,所述计算机程序包含至少一段代码,该至少一段代码可由计算机执行,以控制所述计算机执行权利要求12至21任一项所述的方法。
- 一种计算机程序,当所述计算机程序被计算机执行时,用于执行权利要求1至11任一项所述的方法。
- 一种计算机程序,当所述计算机程序被计算机执行时,用于执行权利要求12至21任一项所述的方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201910415208.2 | 2019-05-13 | ||
| CN201910415208.2A CN111935557B (zh) | 2019-05-13 | 2019-05-13 | 视频处理方法、装置及系统 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020228482A1 true WO2020228482A1 (zh) | 2020-11-19 |
Family
ID=73282865
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2020/085263 Ceased WO2020228482A1 (zh) | 2019-05-13 | 2020-04-17 | 视频处理方法、装置及系统 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN111935557B (zh) |
| WO (1) | WO2020228482A1 (zh) |
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113660529A (zh) * | 2021-07-19 | 2021-11-16 | 镕铭微电子(济南)有限公司 | 基于Tile编码的视频拼接、编码、解码方法及装置 |
| CN114268835A (zh) * | 2021-11-23 | 2022-04-01 | 北京航空航天大学 | 一种低传输流量的vr全景视频时空切片方法 |
| CN114598853A (zh) * | 2020-11-20 | 2022-06-07 | 中国移动通信有限公司研究院 | 视频数据的处理方法、装置及网络侧设备 |
| CN115834899A (zh) * | 2022-08-23 | 2023-03-21 | 北京博雅睿视科技有限公司 | 基于视角的vr视频编码方法、装置、存储介质及电子设备 |
Families Citing this family (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114584769A (zh) * | 2020-11-30 | 2022-06-03 | 华为技术有限公司 | 一种视角切换方法及装置 |
| CN112770051B (zh) * | 2021-01-04 | 2022-01-14 | 聚好看科技股份有限公司 | 一种基于视场角的显示方法及显示设备 |
| CN115499634B (zh) * | 2021-06-18 | 2025-08-29 | 华为技术有限公司 | 一种视频处理方法及装置 |
| CN114189696B (zh) * | 2021-11-24 | 2024-03-08 | 阿里巴巴(中国)有限公司 | 一种视频播放方法及设备 |
| CN115174942A (zh) * | 2022-07-08 | 2022-10-11 | 叠境数字科技(上海)有限公司 | 一种自由视角切换方法及交互式自由视角播放系统 |
| CN115834966A (zh) * | 2022-11-07 | 2023-03-21 | 抖音视界有限公司 | 一种视频播放方法、装置、设备和存储介质 |
Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102055967A (zh) * | 2009-10-28 | 2011-05-11 | 中国移动通信集团公司 | 多视点视频的视角切换以及编码方法和装置 |
| CN105075269A (zh) * | 2013-04-04 | 2015-11-18 | 夏普株式会社 | 图像解码装置以及图像编码装置 |
| US20150346832A1 (en) * | 2014-05-29 | 2015-12-03 | Nextvr Inc. | Methods and apparatus for delivering content and/or playing back content |
| CN107439010A (zh) * | 2015-05-27 | 2017-12-05 | 谷歌公司 | 流传输球形视频 |
| CN108616758A (zh) * | 2016-12-15 | 2018-10-02 | 北京三星通信技术研究有限公司 | 多视点视频编码、解码方法及编码器、解码器 |
| CN108810636A (zh) * | 2017-04-28 | 2018-11-13 | 华为技术有限公司 | 视频播放方法、设备及系统 |
| CN109691113A (zh) * | 2016-07-15 | 2019-04-26 | 皇家Kpn公司 | 流式传输虚拟现实视频 |
Family Cites Families (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109698949B (zh) * | 2017-10-20 | 2020-08-21 | 腾讯科技(深圳)有限公司 | 基于虚拟现实场景的视频处理方法、装置和系统 |
-
2019
- 2019-05-13 CN CN201910415208.2A patent/CN111935557B/zh active Active
-
2020
- 2020-04-17 WO PCT/CN2020/085263 patent/WO2020228482A1/zh not_active Ceased
Patent Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102055967A (zh) * | 2009-10-28 | 2011-05-11 | 中国移动通信集团公司 | 多视点视频的视角切换以及编码方法和装置 |
| CN105075269A (zh) * | 2013-04-04 | 2015-11-18 | 夏普株式会社 | 图像解码装置以及图像编码装置 |
| US20150346832A1 (en) * | 2014-05-29 | 2015-12-03 | Nextvr Inc. | Methods and apparatus for delivering content and/or playing back content |
| CN107439010A (zh) * | 2015-05-27 | 2017-12-05 | 谷歌公司 | 流传输球形视频 |
| CN109691113A (zh) * | 2016-07-15 | 2019-04-26 | 皇家Kpn公司 | 流式传输虚拟现实视频 |
| CN108616758A (zh) * | 2016-12-15 | 2018-10-02 | 北京三星通信技术研究有限公司 | 多视点视频编码、解码方法及编码器、解码器 |
| CN108810636A (zh) * | 2017-04-28 | 2018-11-13 | 华为技术有限公司 | 视频播放方法、设备及系统 |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114598853A (zh) * | 2020-11-20 | 2022-06-07 | 中国移动通信有限公司研究院 | 视频数据的处理方法、装置及网络侧设备 |
| CN114598853B (zh) * | 2020-11-20 | 2025-02-25 | 中国移动通信有限公司研究院 | 视频数据的处理方法、装置及网络侧设备 |
| CN113660529A (zh) * | 2021-07-19 | 2021-11-16 | 镕铭微电子(济南)有限公司 | 基于Tile编码的视频拼接、编码、解码方法及装置 |
| CN114268835A (zh) * | 2021-11-23 | 2022-04-01 | 北京航空航天大学 | 一种低传输流量的vr全景视频时空切片方法 |
| CN115834899A (zh) * | 2022-08-23 | 2023-03-21 | 北京博雅睿视科技有限公司 | 基于视角的vr视频编码方法、装置、存储介质及电子设备 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN111935557B (zh) | 2022-06-28 |
| CN111935557A (zh) | 2020-11-13 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2020228482A1 (zh) | 视频处理方法、装置及系统 | |
| CN111052754B (zh) | 将空间元素的帧流传输到客户端装置 | |
| CN112789854B (zh) | 用于在暂停期间在360º沉浸式视频中提供质量控制的系统和方法 | |
| CN112752115B (zh) | 直播数据传输方法、装置、设备及介质 | |
| CN102986218B (zh) | 用于串流视频数据的视频切换 | |
| US8818021B2 (en) | Watermarking of digital video | |
| CA2965484C (en) | Adaptive bitrate streaming latency reduction | |
| KR101737325B1 (ko) | 멀티미디어 시스템에서 멀티미디어 서비스의 경험 품질 감소를 줄이는 방법 및 장치 | |
| CN102598617B (zh) | 用于从移动设备向无线显示器传送内容的系统和方法 | |
| US11109092B2 (en) | Synchronizing processing between streams | |
| US10277927B2 (en) | Movie package file format | |
| KR102076064B1 (ko) | Dash의 강건한 라이브 동작 | |
| CN110519640B (zh) | 视频处理方法、编码器、cdn服务器、解码器、设备及介质 | |
| CN111726657A (zh) | 直播视频的播放处理方法、装置及服务器 | |
| CN110582012A (zh) | 视频切换方法、视频处理方法、装置及存储介质 | |
| JP7553679B2 (ja) | タイルベースの没入型ビデオをエンコードするためのエンコーダおよび方法 | |
| TW202423095A (zh) | 回應於網路中斷的視訊內容的自動產生 | |
| US11997366B2 (en) | Method and apparatus for processing adaptive multi-view streaming | |
| US10002644B1 (en) | Restructuring video streams to support random access playback | |
| CN115834899A (zh) | 基于视角的vr视频编码方法、装置、存储介质及电子设备 | |
| CN114513658B (zh) | 一种视频加载方法、装置、设备及介质 | |
| CN115883855A (zh) | 播放数据处理方法、装置、计算机设备和存储介质 | |
| CN107148779A (zh) | 自适应比特率流送时延减少 | |
| CN117440174A (zh) | 一种视频处理方法、装置、电子设备和存储介质 | |
| HK40084134A (zh) | 播放数据处理方法、装置、计算机设备和存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20805909 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20805909 Country of ref document: EP Kind code of ref document: A1 |