WO2023045649A1 - 视频帧播放方法、装置、设备、存储介质及程序产品 - Google Patents
视频帧播放方法、装置、设备、存储介质及程序产品 Download PDFInfo
- Publication number
- WO2023045649A1 WO2023045649A1 PCT/CN2022/113526 CN2022113526W WO2023045649A1 WO 2023045649 A1 WO2023045649 A1 WO 2023045649A1 CN 2022113526 W CN2022113526 W CN 2022113526W WO 2023045649 A1 WO2023045649 A1 WO 2023045649A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- video frame
- video
- frame
- resolution
- video frames
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T3/00—Geometric image transformations in the plane of the image
- G06T3/40—Scaling of whole images or parts thereof, e.g. expanding or contracting
- G06T3/4053—Scaling of whole images or parts thereof, e.g. expanding or contracting based on super-resolution, i.e. the output image resolution being higher than the sensor resolution
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/40—Scenes; Scene-specific elements in video content
- G06V20/46—Extracting features or characteristics from the video content, e.g. video fingerprints, representative shots or key frames
-
- A—HUMAN NECESSITIES
- A63—SPORTS; GAMES; AMUSEMENTS
- A63F—CARD, BOARD, OR ROULETTE GAMES; INDOOR GAMES USING SMALL MOVING PLAYING BODIES; VIDEO GAMES; GAMES NOT OTHERWISE PROVIDED FOR
- A63F13/00—Video games, i.e. games using an electronically generated display having two or more dimensions
- A63F13/30—Interconnection arrangements between game servers and game devices; Interconnection arrangements between game devices; Interconnection arrangements between game servers
- A63F13/35—Details of game servers
- A63F13/355—Performing operations on behalf of clients with restricted processing capabilities, e.g. servers transform changing game scene into an encoded video stream for transmitting to a mobile phone or a thin client
-
- A—HUMAN NECESSITIES
- A63—SPORTS; GAMES; AMUSEMENTS
- A63F—CARD, BOARD, OR ROULETTE GAMES; INDOOR GAMES USING SMALL MOVING PLAYING BODIES; VIDEO GAMES; GAMES NOT OTHERWISE PROVIDED FOR
- A63F13/00—Video games, i.e. games using an electronically generated display having two or more dimensions
- A63F13/30—Interconnection arrangements between game servers and game devices; Interconnection arrangements between game devices; Interconnection arrangements between game servers
- A63F13/35—Details of game servers
- A63F13/358—Adapting the game course according to the network or server load, e.g. for reducing latency due to different connection speeds between clients
-
- A—HUMAN NECESSITIES
- A63—SPORTS; GAMES; AMUSEMENTS
- A63F—CARD, BOARD, OR ROULETTE GAMES; INDOOR GAMES USING SMALL MOVING PLAYING BODIES; VIDEO GAMES; GAMES NOT OTHERWISE PROVIDED FOR
- A63F13/00—Video games, i.e. games using an electronically generated display having two or more dimensions
- A63F13/50—Controlling the output signals based on the game progress
- A63F13/52—Controlling the output signals based on the game progress involving aspects of the displayed game scene
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0464—Convolutional networks [CNN, ConvNet]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0475—Generative networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T3/00—Geometric image transformations in the plane of the image
- G06T3/40—Scaling of whole images or parts thereof, e.g. expanding or contracting
- G06T3/4007—Scaling of whole images or parts thereof, e.g. expanding or contracting based on interpolation, e.g. bilinear interpolation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T3/00—Geometric image transformations in the plane of the image
- G06T3/40—Scaling of whole images or parts thereof, e.g. expanding or contracting
- G06T3/4092—Image resolution transcoding, e.g. by using client-server architectures
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/40—Extraction of image or video features
- G06V10/44—Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersections; Connectivity analysis, e.g. of connected components
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N13/00—Stereoscopic video systems; Multi-view video systems; Details thereof
- H04N13/30—Image reproducers
- H04N13/349—Multi-view displays for displaying three or more geometrical viewpoints without viewer tracking
- H04N13/351—Multi-view displays for displaying three or more geometrical viewpoints without viewer tracking for displaying simultaneously
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/44—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
- H04N21/44008—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving operations for analysing video streams, e.g. detecting features or characteristics in the video stream
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/44—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
- H04N21/4402—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving reformatting operations of video signals for household redistribution, storage or real-time display
- H04N21/440281—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving reformatting operations of video signals for household redistribution, storage or real-time display by altering the temporal resolution, e.g. by frame skipping
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/442—Monitoring of processes or resources, e.g. detecting the failure of a recording device, monitoring the downstream bandwidth, the number of times a movie has been viewed, the storage space available from the internal hard disk
- H04N21/44209—Monitoring of downstream path of the transmission network originating from a server, e.g. bandwidth variations of a wireless network
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/47—End-user applications
- H04N21/478—Supplemental services, e.g. displaying phone caller identification, shopping application
- H04N21/4781—Games
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/60—Network structure or processes for video distribution between server and client or between remote clients; Control signalling between clients, server and network components; Transmission of management data between server and client, e.g. sending from server to client commands for recording incoming content stream; Communication details between server and client
- H04N21/63—Control signaling related to video distribution between client, server and network components; Network processes for video distribution between server and clients or between remote clients, e.g. transmitting basic layer and enhancement layers over different transmission paths, setting up a peer-to-peer communication via Internet between remote STB's; Communication protocols; Addressing
- H04N21/637—Control signals issued by the client directed to the server or network components
Definitions
- the present application relates to the field of computer technology and cloud games, and in particular to a method, device, equipment, storage medium and program product for playing video frames.
- the terminal does not need to perform rendering operations, and the cloud game server can render the game scene based on the control information sent by the terminal to obtain video frames.
- the cloud game server sends the video frame to the terminal, and the terminal only needs to display the video frame.
- the cloud game server when the network status of the terminal is not good, the cloud game server will reduce the resolution of the video frame to ensure the smoothness of the cloud game, but reducing the resolution of the video frame will cause the display effect of the cloud game to deteriorate.
- Embodiments of the present application provide a video frame playback method, device, device, storage medium, and program product, which can improve the clarity of cloud game screens displayed on terminals while ensuring the smoothness of cloud games.
- An embodiment of the present application provides a video frame playback method, the method comprising:
- the resolutions of the plurality of first video frames meet the resolution adjustment conditions, the resolutions of the first video frames are respectively adjusted to obtain corresponding second video frames, and the resolution of the second video frames is higher than the resolution of the corresponding first video frame;
- An embodiment of the present application provides a video frame playback device, the device comprising:
- the video frame acquisition module is configured to acquire a plurality of first video frames, and the plurality of first video frames are video frames obtained by rendering the target virtual scene by the cloud game server;
- the resolution adjustment module is configured to adjust the resolution of each of the first video frames respectively to obtain corresponding second video frames when the resolutions of the plurality of first video frames meet the resolution adjustment conditions.
- the resolution of the second video frame is higher than the resolution of the corresponding first video frame;
- the playing module is configured to play multiple second video frames obtained through resolution adjustment.
- An embodiment of the present application provides a computer device, the computer device includes one or more processors and one or more memories, at least one computer program is stored in the one or more memories, and the computer program is executed by the The one or more processors are loaded and executed, so as to implement the video frame playing method provided in the embodiment of the present application.
- An embodiment of the present application provides a computer-readable storage medium, where at least one computer program is stored in the computer-readable storage medium, and the computer program is loaded and executed by a processor to implement the described Video frame playback method.
- An embodiment of the present application provides a computer program product or computer program, the computer program product or computer program includes program code, the program code is stored in a computer-readable storage medium, and the processor of the computer device reads from the computer-readable storage medium The program code is fetched, and the processor executes the program code, so that the computer device executes the video frame playing method provided by the embodiment of the present application.
- the resolution of the multiple first video frames is judged, and when the resolution meets the resolution adjustment condition, the Adjusting the resolution of each first video frame to increase the resolution of the first video frame to obtain a corresponding second video frame, and then playing a plurality of second video frames obtained by increasing the resolution of the first video frame, Compared with directly playing multiple first video frames sent by the cloud game server, the clarity of the cloud game screen is improved, thereby improving the display effect of the cloud game on the premise of ensuring the smoothness of the cloud game.
- FIG. 1 is a schematic diagram of an implementation environment of a video frame playback method provided in an embodiment of the present application
- Fig. 2 is a basic flowchart of a cloud game provided by the embodiment of the present application.
- Fig. 3 is a flow chart of a video frame playback method provided by an embodiment of the present application.
- FIG. 4 is a flow chart of a video frame playback method provided by an embodiment of the present application.
- FIG. 5 is a flow chart of a video frame playback method provided by an embodiment of the present application.
- FIG. 6 is a first frame delay change diagram provided by an embodiment of the present application.
- Fig. 7 is a diagram of the variation of the freezing rate provided by the embodiment of the present application.
- FIG. 8 is a schematic structural diagram of a video frame playback device provided by an embodiment of the present application.
- FIG. 9 is a schematic structural diagram of a terminal provided by an embodiment of the present application.
- first and second are used to distinguish the same or similar items with basically the same function and function. It should be understood that “first”, “second” and “nth” There are no logical or timing dependencies, nor are there restrictions on quantity or order of execution.
- the term "at least one" refers to one or more, and the meaning of “multiple” refers to two or more.
- multiple reference face images refer to two or more reference faces. image.
- Cloud Gaming also known as Gaming on Demand
- Cloud gaming technology enables thin clients with relatively limited graphics processing and data computing capabilities to run high-quality games.
- the game is not run on the player's game terminal, but on the cloud game server, and the cloud game server renders the game scene into a video and audio stream, which is transmitted to the player's game terminal through the network.
- the player's game terminal does not need to have powerful graphics computing and data processing capabilities, but only needs to have basic streaming media playback capabilities, the ability to obtain player input instructions and send them to the cloud game server, and basic data processing capabilities.
- Virtual scene it is a virtual scene displayed (or provided) when the application program is running on the terminal.
- the virtual scene can be a simulation environment of the real world, a semi-simulation and semi-fictional virtual environment, or a purely fictional virtual environment.
- the virtual scene may be any one of a two-dimensional virtual scene, a 2.5-dimensional virtual scene, or a three-dimensional virtual scene, and the embodiment of the present application does not limit the dimensions of the virtual scene.
- the virtual scene may include sky, land, ocean, etc.
- the land may include environmental elements such as deserts and cities, and the user may control virtual objects to move in the virtual scene.
- Virtual object refers to the movable object in the virtual scene.
- the movable object may be a virtual character, a virtual animal, an animation character, etc., such as: a character, an animal, a plant, an oil drum, a wall, a stone, etc. displayed in a virtual scene.
- the virtual object may be a virtual avatar representing the user in the virtual scene.
- the virtual scene may include multiple virtual objects, and each virtual object has its own shape and volume in the virtual scene and occupies a part of the space in the virtual scene.
- the virtual object is a user character controlled by an operation on the client, or an artificial intelligence (Artificial Intelligence, AI) set in a virtual scene battle through training, or an artificial intelligence (AI) set in a virtual scene Non-Player Character (NPC).
- AI Artificial Intelligence
- AI artificial Intelligence
- AI artificial intelligence
- NPC Non-Player Character
- the virtual object is a virtual character competing in a virtual scene.
- the number of virtual objects participating in the interaction in the virtual scene is preset, or dynamically determined according to the number of clients participating in the interaction.
- the user can control the virtual object to fall freely in the sky of the virtual scene, glide or open the parachute to fall, etc., run, jump, crawl, bend forward, etc. on the land, and can also control The virtual object swims, floats or dives in the ocean.
- the user can also control the virtual object to move in the virtual scene on a virtual vehicle.
- the virtual vehicle can be a virtual car, a virtual aircraft, a virtual yacht, etc.
- the above-mentioned scenario is used as an example, and this embodiment of the present application does not limit it.
- Users can also control the interaction between virtual objects and other virtual objects through interactive props, such as fighting. It can also be shooting interactive props such as virtual machine guns, virtual pistols, virtual rifles, etc. This application does not limit the types of interactive props.
- Display resolution mainly refers to the number of pixels that the display can display, which can be classified from two directions: display resolution and image resolution.
- Display resolution (screen resolution) is the precision of the screen image, which refers to the number of pixels that the display can display. Since points, lines and planes on the screen are all composed of pixels, the more pixels the display can display, the finer the picture, and the more information can be displayed in the same screen area, so resolution is a very important performance One of the indicators. The entire image can be imagined as a large chessboard, and the resolution is expressed as the number of intersections of all longitude and latitude lines. When the display resolution is constant, the smaller the display screen is, the clearer the image will be. Conversely, when the display screen size is fixed, the higher the display resolution will be, the clearer the image will be. Image resolution is the number of pixels contained in a unit inch, and its definition is closer to the definition of resolution itself.
- Frame insertion add one frame to every two frames displayed on the original screen, shorten the display time between each frame, and double the time, for example, increase the video frame rate from the original 30Hz to 60HZ . Correct the illusion caused by the persistence of human vision and effectively improve the stability of the picture.
- super-resolution the method of improving the image resolution, based on the motion prediction or motion compensation of the temporally adjacent frame-assisted reference frame and related deep learning models, providing upsampling of any low resolution to Nx (such as 2x) times resolution technical solutions, such as super-resolution of 2K to 4K resolution.
- MOS Mel Opinion Score
- FPS Framework Per Second
- Hz freshness rate
- FPS is the definition in the image field, which refers to the number of frames per second transmitted by the screen. Generally speaking, it refers to Number of frames for animation or video. FPS is a measure of the amount of information used to save and display dynamic video. The more frames per second, the smoother the displayed motion will be.
- SPS Sequence Parameter Set
- a set of global parameters of the coded video sequence (Coded Video Sequence) is saved in the SPS.
- the so-called coded video sequence is a sequence composed of the encoded structure of the pixel data of the original video frame by frame.
- FIG. 1 is a schematic diagram of an implementation environment of a video frame playback method provided by an embodiment of the present application.
- the implementation environment may include a terminal 110 and a cloud game server 140 .
- the terminal 110 and the cloud game server 140 are nodes in the blockchain system, and the data transmitted between the terminal 110 and the server 140 is stored on the blockchain.
- the terminal 110 is connected to the cloud game server 140 through a wireless network or a wired network.
- the terminal 110 is a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto.
- the terminal 110 is installed and runs a client supporting virtual scene display.
- Cloud game server 140 is an independent physical cloud game server, or a cloud game server cluster or distributed system composed of multiple physical cloud game servers, or provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network Cloud game servers for basic cloud computing services such as cloud communications, middleware services, domain name services, security services, distribution networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms.
- the cloud gaming server 140 is also referred to as an edge computing node.
- the terminal 110 generally refers to one of multiple terminals, and the embodiment of the present application only uses the terminal 110 as an example for illustration.
- the number of the foregoing terminals may be more or less.
- the embodiment of the present application does not limit the number of terminals and device types.
- the terminal that is, It is the above-mentioned terminal 110
- the cloud game server is also the above-mentioned cloud game server 140.
- the video frame playback method provided by the embodiment of the present application can be applied in various cloud game scenarios, such as first-person shooting (First-person Shooting, FPS) games, or third-person shooting games (Third-Personal Shooting) Shooting, TPS) games, or in multiplayer online battle arena (Multiplayer Online Battle Arena, MOBA), or in war chess games or auto chess games, the embodiment of the present application does not limit this.
- the user starts the cloud game client on the terminal, and logs in the user account in the cloud game client, that is, the user is in the cloud game client. Enter the user account and corresponding password, and click the login control to log in.
- the terminal sends a login request to the cloud game server, and the login request carries a user account and a corresponding password.
- the cloud game server After the cloud game server receives the login request, it obtains the user account and the corresponding password from the login request, and verifies the user account and the corresponding password.
- the cloud game server After the cloud game server passes the verification of the user account and the corresponding password, it sends a login success message to the terminal. After receiving the successful login information, the terminal sends a cloud game acquisition request to the cloud game server, and the cloud game acquisition request carries the user account. After the cloud game server obtains the cloud game acquisition request, it performs a query based on the user account carried in the cloud game acquisition request, obtains multiple cloud games corresponding to the user account, sends the identifiers of the multiple cloud games to the terminal, and the terminal The identifiers of multiple cloud games are displayed in the cloud game client.
- the user selects the logo of the FPS game he wants to play from among the logos of the cloud games displayed on the cloud game client, that is, he selects the FPS game he wants to play.
- the terminal sends a game start command to the cloud game server, and the game start command carries the user account, the logo of the FPS game and the hardware information of the terminal, wherein the hardware information of the terminal includes the terminal.
- the screen resolution of the terminal, the model of the terminal, etc., are not limited in this embodiment of the present application.
- the cloud game server receives the game start instruction, it obtains the user account, the identifier of the FPS game and the hardware information of the terminal from the game start instruction.
- the cloud game server initializes the FPS game based on the hardware information of the terminal, so as to realize the matching between the rendered game screen and the terminal.
- the cloud game server starts the FPS game.
- the user can control the controlled virtual object in the FPS to move through the terminal, that is, the terminal sends control information on the controlled virtual object to the cloud game server, and the cloud game server based on the control
- the information renders the virtual scene of the FPS to obtain the first video frame.
- the terminal transmits interactive operations (control information) to the cloud server (cloud game server) in real time.
- the cloud server performs rendering calculations based on the received interactive operations, and returns compressed audio and video streams to the terminal.
- the terminal decodes and plays the received audio and video streams.
- the cloud game server when the cloud game server renders the virtual scene, in addition to referring to the screen resolution and model of the terminal, it also refers to the network delay information of the terminal. If the network delay information of the terminal indicates that the current network delay of the terminal is relatively high, then the cloud game server can render the virtual scene with lower quality, and the resolution of the obtained first video frame is also lower. Correspondingly, the first video frame The network bandwidth occupied by the frame is also small, ensuring the smoothness of the FPS game running on the terminal to the greatest extent. For example, when the terminal network delay is low, the resolution of the first video frame rendered by the cloud game server is 1080p; When the terminal network delay is high, the resolution of the first video frame rendered by the cloud game server can be reduced to 720p.
- the terminal can adjust the resolution of the received first video frame by using the video frame playback method provided in the embodiment of the present application, so as to improve the resolution of the first video frame , to get the second video frame.
- the second video frame obtained after the resolution adjustment has a higher resolution, that is to say, the second video frame has a better display effect.
- the display effect of the PFS can be improved under the premise of ensuring the fluency of the FPS.
- it can improve the user's gaming experience.
- the cloud game server when the cloud game server sends the first video frame to the terminal, it does not send it frame by frame, but renders a sequence of video frames, and the sequence of video frames includes multiple first video frames.
- the server sends the video frame sequence to the terminal each time for the terminal to display.
- the cloud game server can not only reduce the resolution of the rendered first video frame, but also reduce the number of first video frames in the video frame sequence, reducing the The number of the first video frame in the frame sequence is to reduce the frame rate of the terminal displaying the first video frame.
- the video frame sequence carries 60 first video frames, and the 60 The first video frame is evenly displayed by the terminal within 1 second, and the frame rate at this time is 60; when the network delay of the terminal is low, the number of the first video frame carried by the video frame sequence is reduced to 30, these 3 The first video frame is uniformly displayed by the terminal within 1 second, and the frame rate at this time is 30.
- the number of transmitted video frames can be reduced by reducing the frame rate, which also reduces the bandwidth occupation when transmitting video frames, but reducing the frame rate will lead to "frustration" in the display.
- the terminal can use the video frame playing method provided in the embodiment of the present application to insert frames in the received video frame sequence, so as to increase the number of video frames in the video frame sequence. Eliminates "jerking" and improves playback of video frames.
- the video frame playback method provided by the embodiment of the present application can be applied to other types of cloud games besides the above-mentioned FPS game, MOBA game, war chess game or auto chess game. There is no limit to this.
- Fig. 3 is a flow chart of a video frame playback method provided by the embodiment of the present application. Referring to Fig. 3, the method includes:
- the terminal acquires multiple first video frames, where the multiple first video frames are video frames obtained by rendering a target virtual scene by a cloud game server.
- the terminal running on the terminal, such as a game client or other clients with game functions (such as an instant messaging client).
- game client running on the terminal
- game functions such as an instant messaging client
- the cloud game server can render the target virtual scene from the perspective of the controlled virtual object in the target virtual scene to obtain video frames to be displayed by the terminal.
- the cloud game server renders the target virtual scene to obtain multiple video frames, and the multiple video frames may exist in the form of a video frame sequence, and send the video frame sequence to the terminal.
- the controlled virtual object is also the virtual object controlled by the terminal, that is, the virtual object corresponding to the user account logged in by the client.
- rendering the target virtual scene refers to rendering the picture observed by the controlled virtual object in the target virtual scene to obtain the first video frame, and the controlled virtual object is in the target virtual scene.
- the picture observed in the virtual scene is also the picture seen by the user.
- the terminal adjusts the resolution of each first video frame to obtain a corresponding second video frame, and the resolution of the second video frame is higher than The resolution of the corresponding first video frame.
- the terminal after receiving the multiple video frames (video frame sequence) sent by the cloud game server, acquires the resolutions of the multiple video frames.
- the resolution of the multiple video frames is Similarly, the terminal compares the resolutions of the plurality of video frames with the resolution threshold to obtain a comparison result, and when the comparison result indicates that the resolutions of the plurality of first video frames are less than or equal to the resolution threshold, It is determined that the resolutions of the plurality of first video frames meet the resolution adjustment condition.
- the process of the terminal adjusting the resolution of the multiple first video frames is also the process of performing super-resolution on the multiple first video frames.
- the super-resolution can improve the resolution of the first video frames, so that the terminal compared Directly playing multiple first video frames sent by the cloud game server improves the clarity of the cloud game screen, that is, improves the display effect of the video frames.
- the terminal plays the multiple second video frames obtained through resolution adjustment.
- the second video frame is obtained after the terminal adjusts the resolution of the first video frame, then the second video frame also has the same image content as the corresponding first video frame, compared to displaying the first video frame In other words, since the second video frame has a higher resolution, the terminal displays the second video frame with higher definition, and thus has a better display effect.
- the judgment is made based on the resolutions of the multiple first video frames, and when the resolution meets the resolution adjustment condition, the A plurality of first video frames are adjusted in resolution, and the purpose of resolution adjustment is to improve the resolution of the first video frames to obtain a plurality of second video frames, and the plurality of second video frames also have higher resolution, thereby
- the display effect of cloud games has been improved.
- the method includes:
- the terminal sends network delay information to the cloud game server, so that the cloud game server generates multiple first video frames based on the network delay information.
- the first video frame is a video frame obtained by rendering the target virtual scene from the perspective of the controlled virtual object in the target virtual scene by the cloud game server.
- the network delay information is used to represent the network delay between the terminal and the cloud game server.
- the target virtual scene is the game scene of the cloud game selected by the user.
- the angle of view of the controlled virtual object is also the angle of view of the virtual camera of the controlled virtual object.
- the virtual camera of the controlled virtual object is located in the virtual When the user controls the controlled virtual object to move in the target virtual scene through the terminal, the virtual camera will also move with the movement of the controlled virtual object, and the picture captured by the virtual camera is also the controlled virtual object. The picture observed by the virtual object in the target virtual scene.
- the virtual camera of the controlled virtual object is located above the controlled virtual object.
- the virtual camera When the user controls the controlled virtual object to move in the target virtual scene through the terminal, the virtual camera will also move with the controlled virtual object. While moving, the picture captured by the virtual camera is also the picture observed above the controlled virtual object.
- the images captured by the virtual camera are rendered by the cloud game server. Since there are multiple sequential frames during the game, the multiple frames are called video frames in this application.
- the terminal starts the cloud game client, obtains the network delay information between the terminal and the cloud game server through the cloud game client, and the terminal sends the network delay information to the cloud game server.
- the cloud game server determines the corresponding rendering parameters according to the network delay information, uses the rendering parameters to render the target virtual scene from the perspective of the controlled virtual object in the target virtual scene, and obtains multiple A video frame.
- the terminal can send the network delay information to the cloud game server, and the cloud game server can determine the rendering parameters based on the network delay information.
- the obtained video frames are also adapted to the terminal network conditions, thereby ensuring the smoothness of the terminal running cloud games.
- Example 1 The terminal starts the cloud game client, and sends a detection packet to the cloud game server through the cloud game client.
- the detection packet is used to request the cloud game server to return a confirmation packet.
- the terminal determines the time difference between receiving the confirmation data packet and sending the detection data packet as the network delay information, and the terminal sends the network delay information to the cloud game server.
- the cloud game server determines a rendering parameter corresponding to the network delay information, where the rendering parameter is used to indicate the resolution of the rendered video frame.
- the cloud game server uses the rendering parameters to render the target virtual scene from the perspective of the controlled virtual object in the target virtual scene to obtain a plurality of first video frames.
- the rendering work is completed by the Graphics Processing Unit (GPU) of the cloud game server.
- GPU Graphics Processing Unit
- the graphics processor of the cloud game server renders the target virtual scene, it will store the rendered multiple game screens in the video memory.
- the graphics processor of the cloud game server will directly encode multiple game screens in the video memory to obtain multiple first video frames.
- the graphics processor of the cloud game server can encode multiple game screens in the video memory into VP8 (a video format developed and launched by Google)/VP9 (a video format developed and launched by Google)/H.264/ H.265/Advanced Video Coding (Advanced Video Coding, AVC)/Audio Video Coding Standard (Audio Video Coding Standard, AVS) and other formats of the first video frame, which is not limited in this embodiment of the present application.
- the cloud game server can also encode the audio data into Silk (an audio format developed by Microsoft)/Opus (an open source audio format)/Advanced Audio Coding (Advanced Audio Coding) , AAC) and other formats of audio data streams.
- the terminal starts the cloud game client, and sends a test data download request to the cloud game server through the cloud game client.
- the test data download request is used to request to download test data from the cloud game server.
- the terminal downloads the test data from the cloud game server through the cloud game client, and the download time is set by the technician according to the actual situation, such as 1s or 3s, etc., which is not limited in the embodiment of the present application.
- the terminal divides the downloaded data amount by the download time to obtain the network delay information.
- the terminal sends the network delay information to the cloud game server.
- the cloud game server determines a rendering parameter corresponding to the network delay information, where the rendering parameter is used to indicate the resolution of the rendered video frame.
- the cloud game server uses the rendering parameters to render the target virtual scene from the perspective of the controlled virtual object in the target virtual scene to obtain a plurality of first video frames.
- the terminal when running the cloud game on the cloud game client, sends control information on the controlled virtual object and network delay information to the cloud game server through the cloud game client.
- the cloud game server receives the control information and the network delay information, it determines the viewing angle of the controlled virtual object in the target virtual scene based on the control information, and determines the corresponding rendering parameters based on the network delay information.
- the cloud game server renders the target virtual scene based on the rendering parameters and the viewing angle of the controlled virtual object to obtain the first video frame.
- control information on the controlled virtual object is used to change the position, orientation and action of the controlled virtual object in the target scene.
- the control information can control the controlled virtual object to move forward, backward, left, and right in the target virtual scene, or It can control the controlled virtual object to rotate left or right in the target virtual scene, or can control the controlled virtual object to perform actions such as squatting, crawling and using virtual props in the target virtual scene.
- the cloud game server controls the controlled virtual object to move or perform actions in the target virtual scene based on the control information
- the virtual camera bound to the controlled virtual object will also move along with the controlled virtual
- the movement or execution of actions by the controlled virtual object will cause the angle of view of the controlled virtual object to observe the target virtual scene to change, and the virtual camera bound to the controlled virtual object can record this change.
- the terminal sends the network delay information to the cloud game server, and the cloud game server determines the corresponding rendering parameters based on the network delay information.
- the way to determine the rendering parameters is explained.
- the terminal starts the cloud game client, and obtains the network delay information between the terminal and the cloud game server through the cloud game client.
- the terminal determines video stream information based on the network delay information, where the video stream information includes the resolution, bit rate, and frame rate of the video stream, where the video stream includes multiple video sequences, and each video sequence includes multiple first video frames .
- the first video frames in the same video sequence are rendered by the cloud game server using the same rendering parameters, that is, the first video frames in the same video sequence have the same resolution, and the first video frames in different video sequences have the same resolution.
- the first video frame may be rendered by the cloud game server using different rendering parameters, that is, the resolutions of the first video frames in different video sequences may be different.
- the terminal sends the video stream information to the cloud game server, and the cloud game server receives the video stream information, and determines corresponding rendering parameters based on the video stream information.
- the cloud game server uses the rendering parameters to render the target virtual scene from the perspective of the controlled virtual object in the target virtual scene to obtain a plurality of first video frames.
- the terminal after the terminal obtains the network delay information, it can directly determine the video stream information based on the network delay information, and the cloud game server can quickly determine the corresponding rendering parameters directly based on the video stream information, which is more efficient.
- the terminal starts the cloud game client, and obtains the network delay information between the terminal and the cloud game server through the cloud game client. Based on the network delay information, the terminal displays a video stream information selection page.
- the video stream information selection page displays multiple candidate video stream information, and the multiple video stream information is video stream information matching the network delay information.
- the terminal sends the target video stream information to the cloud game server, and the cloud game server receives the target video stream information, and determines corresponding rendering parameters based on the target video stream information .
- the cloud game server uses the rendering parameters to render the target virtual scene from the perspective of the controlled virtual object in the target virtual scene to obtain a plurality of first video frames.
- the terminal can provide the user with multiple video stream information to choose from based on the network delay information.
- the process for the user to select the video stream information is to select the resolution, frame Rate and bit rate process, which also provides users with higher autonomy.
- users can also adjust the selected video stream information at any time through the cloud game client, and the cloud game server can also adjust the rendering parameters accordingly.
- the terminal acquires multiple first video frames.
- the multiple first video frames belong to the same video frame sequence, and the multiple first video frames are obtained after the cloud game server renders the target virtual scene based on the same rendering parameters, that is, the multiple first video frames have same resolution. Since the multiple first video frames are obtained after the cloud game server encodes the game screen, the video frame sequence is also a coded video sequence (Coded Video Sequence, CVS).
- CVS Coded Video Sequence
- the cloud game server when the cloud game server acquires the encoded video sequence, it also acquires a sequence parameter set (Sequence Parameter Set, SPS) corresponding to the encoded video sequence from the server, and the sequence parameter set is used to instruct the terminal how to encode the encoded video sequence.
- SPS Sequence Parameter Set
- the video sequence is decoded.
- the terminal Based on the sequence parameter set, the terminal decodes the coded video frame sequence to obtain multiple first video frames.
- the terminal adjusts the resolution of each first video frame to obtain a corresponding second video frame, and the resolution of the second video frame is higher than The resolution of the corresponding first video frame.
- the terminal after receiving the first video frame sequence sent by the cloud game server, the terminal obtains the resolution of each first video frame in the first video frame sequence.
- the resolutions of the first video frames are the same, and the terminal compares the resolutions of the plurality of video frames with the resolution threshold to obtain a comparison result, and when the comparison result represents the resolution of the plurality of first video frames When it is less than or equal to the resolution threshold, it is determined that the resolutions of the plurality of first video frames meet the resolution adjustment condition.
- the terminal inserts reference pixels between every two pixels in each first video frame to obtain the second pixels corresponding to each first video frame.
- the reference pixel is generated based on every two pixels.
- every two pixels in the first video frame described here refers to two spatially adjacent pixels in the first video frame.
- the terminal can insert reference pixel points between every two pixel points in the first video frame.
- the obtained second video frame will have a higher resolution, and the terminal will display the second video frame with higher resolution than the corresponding first video frame.
- One video frame for better results.
- the resolution conforming to the resolution adjustment condition means that the resolution is less than or equal to the resolution threshold, and the resolution threshold is set by the technician according to the actual situation, or set by the user according to the computing capability of the terminal. This application The embodiment does not limit this.
- the terminal may adjust the resolution of the first video frame in the following manner to obtain the corresponding second video frame:
- the terminal uses the nearest neighbor interpolation method to insert the reference pixel between every two pixels in each first video frame to obtain a second video frame corresponding to each first video frame.
- the terminal when the resolution of the multiple first video frames is less than or equal to the resolution threshold, taking inserting a reference pixel between two pixels in the first video frame as an example, the terminal A reference pixel point is inserted between , and the pixel value of the reference pixel point is an initial value, such as 0.
- the terminal updates the pixel value of any one of the two pixel points to the pixel value of the reference pixel point.
- This process embodies the idea of "nearest", that is, the pixel value of the reference pixel is determined as the pixel value of the nearest pixel, so as to quickly complete the resolution adjustment of the first video frame to improve the resolution of the first video frame.
- the resolution of the video frame to get the second video frame.
- the terminal can process each first video frame in the above manner to obtain a second video frame corresponding to each first video frame.
- the terminal can also insert multiple reference pixels. Insert two pixel points between every two pixel points in the frame as an example for illustration.
- the terminal A first reference pixel point and a second reference pixel point are inserted between the pixel points, and the pixel values of the first reference pixel point and the second reference pixel point are both initial values, such as 0.
- the terminal updates the pixel value of the first reference pixel by using the pixel value of the previous pixel of the two pixel points, and updates the pixel value of the second reference pixel by using the pixel value of the latter pixel, wherein the two pixel points
- the previous pixel is the pixel that is closer to the first reference pixel, and correspondingly, the latter pixel is the closest to the second reference pixel. of pixels.
- the manner of distinguishing the "previous pixel point" and the "next pixel point” in the above description will be described below. If these two pixels are arranged from left to right on the first video frame, then the "previous pixel” is the leftmost pixel among the two pixels; The rightmost pixel of the two pixels. If these two pixels are arranged from top to bottom on the first video frame, then the "previous pixel” is the upper pixel of the two pixels; The lower pixel of the two pixels.
- the terminal may adjust the resolution of the first video frame in the following manner to obtain the corresponding second video frame:
- the terminal inserts the reference pixel between every two pixels in each first video frame by using a bilinear interpolation method to obtain a second video frame corresponding to each first video frame.
- the terminal when the resolution of the plurality of first video frames is less than or equal to the resolution threshold, the terminal inserts the reference pixel between every two pixels in the first video frame, and the reference pixel
- the pixel value of the point is the initial pixel value, such as 0.
- the terminal updates the pixel value of the reference pixel based on the distance between the reference pixel and every two pixel points and the pixel values of the two pixel points.
- the terminal inserts the first reference pixel point and the second reference pixel point between the two pixel points, and the first reference pixel point and the The pixel values of the second reference pixel point are all initial pixel values, such as 0.
- the terminal determines two first weights between the first reference pixel and the two pixels based on the distance between the first reference pixel and the two pixels, and the first weight is positively correlated with the distance.
- the terminal determines two second weights between the second reference pixel and the two pixels based on the distance between the second reference pixel and the two pixels, and the second weight is positively correlated with the distance.
- the terminal Based on the two first weights, the terminal performs weighted summation of the pixel values of the two pixel points to obtain a first pixel value, and uses the first pixel value to update the pixel value of the first reference pixel point. Based on the two second weights, the terminal performs weighted summation of the pixel values of the two pixel points to obtain a second pixel value, and uses the second pixel value to update the pixel value of the second reference pixel point.
- the terminal can use the above method to Three or more reference pixel points are inserted between the pixel points, and the embodiment of the present application does not limit the number of inserted reference pixel points.
- the terminal may adjust the resolution of the first video frame in the following manner to obtain the corresponding second video frame:
- the terminal inserts the reference pixel point between every two pixel points in each first video frame by using a mean value interpolation method to obtain a second video frame corresponding to each first video frame.
- the terminal when the resolution of the plurality of first video frames is less than or equal to the resolution threshold, the terminal inserts the reference pixel between every two pixels in the first video frame, and the reference pixel
- the pixel value of the point is the initial pixel value, such as 0.
- the terminal updates the pixel value of the reference pixel based on the average value of the pixel values of every two pixels.
- the terminal when the resolutions of the plurality of first video frames meet the resolution adjustment condition, the terminal inputs the plurality of first video frames into the super-resolution model, and the super-resolution model determines the resolution of the plurality of first video frames.
- a video frame is up-sampled to obtain the plurality of second video frames.
- Example 1 For any first video frame among the plurality of first video frames, the terminal performs feature extraction on the first video frame through the super-resolution model to obtain the first video frame features of the first video frame.
- the terminal performs nonlinear mapping on the features of the first video frame through the super-resolution model to obtain the features of the second video frame of the first video frame.
- the terminal reconstructs the features of the second video frame through the super-resolution model to obtain the second video frame corresponding to the first video frame.
- the upsampling method provided in Example 1 is also called post-sampling super-resolution. Using post-sampling super-resolution can enable the super-resolution model to learn the upsampling process adaptively, and can also enable the feature extraction process in low-dimensional space It can greatly reduce the calculation burden, and the training speed and inference speed are faster.
- the terminal inputs the first video frame into the super-resolution model, and performs convolution processing on the first video frame through at least one convolution layer of the super-resolution model to obtain the first video frame features of the first video frame.
- the terminal performs full connection and non-linear activation on the features of the first video frame through the fully connected layer and the activation layer of the super-resolution model to obtain the features of the second video frame of the first video frame.
- the terminal performs reconstruction based on the features of the second video frame through the reconstruction layer of the super-resolution model.
- the reconstruction here is upsampling, and upsampling can increase the size of the first video frame, that is, increase the size of the first video frame.
- the number of pixels in the frame means that the resolution of the obtained second video frame is higher than that of the corresponding first video frame.
- the reconstruction layer of the super-resolution model is a deconvolution layer or a sub-pixel convolution layer, and the terminal can deconvolute the features of the second video frame through the deconvolution layer to obtain the second video frame , it is also possible to perform deconvolution processing on the features of the second video frame through the sub-pixel convolution layer to obtain the second video frame, which is not limited in this embodiment of the present application.
- the purpose of using the super-resolution model is to improve the resolution of video frames, then in the training process, multiple high-resolution images and corresponding multiple low-resolution images are used as training samples to The super-resolution model is trained, wherein the low-resolution image is obtained by down-sampling the corresponding high-resolution image, and in some embodiments, the low-resolution image is also called a damaged image.
- the terminal initializes the model parameters of the super-resolution model, inputs the low-resolution image into the super-resolution model, and convolves the low-resolution image through at least one convolution layer of the super-resolution model to obtain the low-resolution image First sample image features.
- the terminal performs full connection and non-linear activation on the first sample image feature through the fully connected layer and the activation layer of the super-resolution model to obtain the second sample image feature of the low-resolution image.
- the terminal performs deconvolution or sub-pixel convolution on the features of the second sample image through the reconstruction layer of the super-resolution model, and outputs a super-resolution image corresponding to the low-resolution image.
- the terminal adjusts the model parameters of the super-resolution model based on the difference information between the super-resolution image and the high-resolution image corresponding to the low-resolution image.
- the difference information between the super-resolution image and the high-resolution image corresponding to the low-resolution image is the pixel value difference between the super-resolution image and the high-resolution image corresponding to the low-resolution image At least one of , image feature difference, and texture difference.
- the terminal can also use the method of Generative Adversarial (GA, Generative Adversarial) to train the super-resolution model, that is, a discriminator is introduced in the training process, and the discriminator is used for the super-resolution model.
- G Generative Adversarial
- the output image is scored, and the score is used to indicate the degree of fidelity of the generated high-resolution image. The higher the score, the higher the probability that the discriminator believes that the corresponding image is the generated image; the higher the score, it means the discriminant The higher the probability that the image processor considers the image to be a native image.
- the terminal After the terminal inputs the low-resolution image into the super-resolution model, it inputs the super-resolution image output by the super-resolution model into the discriminator, and the discriminator performs scoring based on the super-resolution image, and outputs the score corresponding to the super-resolution image. Based on the score, the terminal adjusts the model parameters of the super-resolution model. In the next iteration, the terminal adjusts the parameters of the discriminator according to the difference information between the super-resolution image output by the super-resolution model and the corresponding high-resolution image. Through the "confrontation" between the super-resolution model and the discriminator, the upsampling effect of the super-resolution model is improved.
- the above is an example of training the super-resolution model by the terminal.
- the super-resolution model can also be obtained by cloud training, and the terminal directly obtains the super-resolution model from the cloud. That is, the embodiment of the present application does not limit this.
- Example 2 For any first video frame among the plurality of first video frames, the terminal performs multiple upsampling on the first video frame through the super-resolution model to obtain a second video frame corresponding to the first video frame,
- the upsampling method provided in Example 2 is also called stepwise upsampling superresolution. Using stepwise upsampling superresolution can decompose difficult tasks into simple tasks.
- the superresolution model under this framework not only greatly reduces the learning difficulty, but also obtains better performance.
- the terminal inputs the first video frame into the super-resolution model, and performs convolution processing on the first video frame through at least one convolution layer of the super-resolution model to obtain the first video frame features of the first video frame.
- the terminal performs upsampling on the features of the first video frame through the upsampling layer of the super-resolution model to obtain the first upsampling features.
- the terminal performs convolution processing on the first upsampling feature through at least one convolutional layer of the super-resolution model to obtain the second video frame feature of the first video frame.
- the terminal performs reconstruction based on the features of the second video frame through the reconstruction layer of the super-resolution model.
- the reconstruction here is upsampling, and upsampling can increase the size of the first video frame, that is, increase the size of the first video frame.
- the number of pixels in the frame means that the resolution of the obtained second video frame is higher than that of the corresponding first video frame.
- the reconstruction layer of the super-resolution model is a deconvolution layer or a sub-pixel convolution layer, and the terminal can deconvolute the features of the second video frame through the deconvolution layer to obtain the second video frame , it is also possible to perform deconvolution processing on the features of the second video frame through the sub-pixel convolution layer to obtain the second video frame, which is not limited in this embodiment of the present application.
- the terminal upsamples the first video frame twice through the super-resolution model.
- the terminal can also use the super-resolution model to up-sample the first video frame. Upsampling is performed three times or more, which is not limited in this embodiment of the present application.
- Example 3 For any first video frame among the plurality of first video frames, the terminal upsamples the first video frame through the super-resolution model to obtain an upsampled video frame corresponding to the first video frame. The number of pixels in the sampled video frame is greater than the number of pixels in the first video frame.
- the terminal performs feature extraction on the upsampled video frame through the super-resolution model to obtain the first video frame feature of the upsampled video frame.
- the terminal performs nonlinear mapping on the features of the first video frame through the super-resolution model to obtain the features of the second video frame of the first video frame.
- the terminal performs deconvolution on the features of the second video frame by using the super-resolution model to obtain the second video frame corresponding to the first video frame.
- the upsampling method provided in Example 3 is also called pre-upsampling super-resolution, which can reduce learning difficulty and obtain images of any scale.
- the terminal inputs the first video frame into the super-resolution model, and upsamples the first video frame through an upsampling layer of the super-resolution model to obtain an upsampled video frame corresponding to the first video frame.
- the upsampling layer can upsample the first video frame by any one of the nearest interpolation method, the bilinear interpolation method and the mean value interpolation method to obtain the upsampled video frame.
- the terminal performs convolution processing on the upsampled video frame through at least one convolutional layer of the super-resolution model, to obtain the first video frame feature of the upsampled video frame.
- the terminal performs full connection and non-linear activation on the features of the first video frame through the fully connected layer and the activation layer of the super-resolution model to obtain the features of the second video frame of the upsampled video frame.
- the terminal performs deconvolution processing on the features of the second video frame through the deconvolution layer of the super-resolution model to obtain the second video frame.
- the terminal initializes the model parameters of the super-resolution model, inputs the low-resolution image into the super-resolution model, and upsamples the low-resolution image through the upsampling layer of the super-resolution model to obtain the corresponding Upsample the image.
- the terminal performs convolution on the low-resolution image through at least one convolutional layer of the super-resolution model to obtain the first sample image features of the low-resolution image.
- the terminal performs full connection and non-linear activation on the first sample image feature through the fully connected layer and the activation layer of the super-resolution model to obtain the second sample image feature of the low-resolution image.
- the terminal performs deconvolution on the features of the second sample image through the deconvolution layer of the super-resolution model, and outputs a super-resolution image corresponding to the low-resolution image.
- the terminal adjusts the model parameters of the super-resolution model based on the difference information between the super-resolution image and the high-resolution image corresponding to the low-resolution image.
- the difference information between the super-resolution image and the high-resolution image corresponding to the low-resolution image is the pixel value difference between the super-resolution image and the high-resolution image corresponding to the low-resolution image At least one of , image feature difference, and texture difference.
- the terminal can also use a generative confrontation method to train the super-resolution model, that is, a discriminator is introduced during the training process, and the discriminator is used to score the images output by the super-resolution model, The score is used to indicate the fidelity of the generated high-resolution image.
- a discriminator is introduced during the training process, and the discriminator is used to score the images output by the super-resolution model, The score is used to indicate the fidelity of the generated high-resolution image.
- the higher the score the higher the probability that the discriminator considers the corresponding image to be the generated image; the higher the score, the higher the discriminator considers the image to be native The higher the probability of the image.
- the terminal After the terminal inputs the low-resolution image into the super-resolution model, it inputs the super-resolution image output by the super-resolution model into the discriminator, and the discriminator performs scoring based on the super-resolution image, and outputs the score corresponding to the super-resolution image. Based on the score, the terminal adjusts the model parameters of the super-resolution model. In the next iteration, the terminal adjusts the parameters of the discriminator according to the difference information between the super-resolution image output by the super-resolution model and the corresponding high-resolution image. Through the "confrontation" between the super-resolution model and the discriminator, the upsampling effect of the super-resolution model is improved.
- the above is an example of training the super-resolution model by the terminal.
- the super-resolution model can also be obtained by cloud training, and the terminal directly obtains the super-resolution model from the cloud. That is, the embodiment of the present application does not limit this.
- the terminal can also use iterative up-down sampling and variants of the above-mentioned sampling methods to obtain the second video frame based on the first video frame.
- EDSR Enhanced Deep Residual Networks for Single Image Super-Resolution
- WDSR Wide Activation for Efficient and Accurate Image Super-Resolution
- the terminal performs frame interpolation among the multiple second video frames to obtain multiple second video frames after frame interpolation.
- the frame rate refers to the number of second video frames played by the terminal per second.
- the cloud game server can reduce the transmission of the first video frame by reducing the resolution of the first video frame.
- step 404 is to A method of interpolating frames at low times to increase the frame rate.
- the terminal when the frame rate of the plurality of second video frames is less than or equal to the frame rate threshold, the terminal inserts a reference video frame between every two second video frames in the plurality of second video frames , to obtain the plurality of second video frames after the frame interpolation.
- every two second video frames described here refers to two second video frames that are adjacent in time sequence.
- the terminal when the frame rate of multiple second video frames is low, the terminal can perform frame interpolation between every two second video frames, so as to increase the frame rate of multiple second video frames.
- the improvement of the rate can eliminate the "frustration" and improve the playback effect of multiple second video frames.
- the terminal determines any second video frame in every two second video frames as the reference video frame.
- the terminal can directly determine any second video frame in every two second video frames as the reference video frame, without requiring the terminal to perform additional calculations, and the efficiency of determining the reference video frame is high.
- the terminal can directly determine the second video frame A or the second video frame B as the reference video frame, and then directly determine the reference video frame Frames are added between the second video frame A and the second video frame B. For example, if the terminal determines the second video frame A as the reference video frame, then after adding the reference video frame between the second video frame A and the second video frame B, ⁇ second video frame A
- the terminal determines the average video frame of every two second video frames as the reference video frame, and the pixel value of the pixel in the average video frame is the corresponding pixel in the every two second video frames The average value of the pixel values.
- the terminal can directly determine the reference pixel point by calculating the average value, the calculation amount is small, and the efficiency of determining the reference video frame is high.
- the terminal For example, for the second video frame A and the second video frame B that are temporally adjacent, the terminal generates an average video frame based on the second video frame A and the second video frame B, and the average video frame is The reference video frame.
- the terminal obtains the pixel value matrix M corresponding to the second video frame A and the pixel value matrix N corresponding to the second video frame B, and obtains the average pixel value matrix O of the pixel value matrix M and the pixel value matrix N. .
- the terminal generates a blank video frame, the number and distribution of pixels in the blank video frame are the same as those of the second video frame A and the second video frame B, and the terminal determines the value in the average pixel value matrix O as the corresponding value in the blank video frame The pixel value of the pixel to get the reference video frame.
- Subsequent terminals only need to add the reference video frame to the second video frame A and the second video frame B, that is, change the original ⁇ second video frame A
- the terminal inputs the second video frame sequence into the frame interpolation model, and the frame interpolation model generates a reference video frame based on two adjacent second video frames in the second video frame sequence, that is, the interpolation frame
- the model performs processing based on the every two second video frames to obtain the reference video frame.
- Example 1 The terminal obtains the backward optical flow and forward optical flow of two adjacent second video frames in the second video frame sequence through the frame interpolation model, for example, obtains the third video frame to the fourth video frame The backward optical flow of the frame, the third video frame is the previous second video frame in the every two second video frames, and the fourth video frame is the next second video frame in the every two second video frames .
- the terminal obtains the forward optical flow from the fourth video frame to the third video frame through the frame interpolation model.
- the terminal generates the reference video frame based on the backward optical flow and the forward optical flow through the frame interpolation model.
- the terminal can generate the reference video frame based on the optical flow method, and the quality of the reference video frame is better, thereby improving the playback effect of the terminal on the video frame.
- the terminal obtains the first backward optical flow of the third video frame to the fourth video frame through the frame interpolation model, and determines the first moment to the middle moment of the third video frame based on the first backward optical flow The second backward optical flow of , wherein the intermediate moment is the moment between the third video frame and the fourth video frame.
- the terminal acquires the feature map and edge image of the third video frame through the frame interpolation model.
- the terminal performs forward mapping on the third video frame and the feature map and edge image of the third video frame based on the second backward optical flow through the frame interpolation model to obtain the first forward mapped video frame and the first forward Mapping reference information, the first forward mapping reference information includes feature maps and edge images of the first forward mapping video frame.
- the terminal obtains the first forward optical flow from the fourth video frame to the third video frame through the frame interpolation model, and determines the second forward optical flow from the second moment to the middle moment of the fourth video frame based on the first forward optical flow. direction optical flow, wherein the intermediate moment is the moment between the fourth video frame and the fourth video frame.
- the terminal obtains the feature map and edge image of the fourth video frame through the frame interpolation model.
- the terminal performs forward mapping on the fourth video frame and the feature map and edge image of the fourth video frame to obtain the second forward mapped video frame and the second forward Mapping reference information
- the second forward mapping reference information includes feature maps and edge images of the second forward mapping video frame.
- the terminal Through the frame interpolation model, the terminal combines the first forward mapping video frame, the first forward mapping reference information, the second forward mapping video frame, and the second forward mapping reference information into a forward mapping result.
- the terminal determines the third backward optical flow from the intermediate moment to the second moment based on the first forward optical flow based on the frame interpolation model, and based on the third backward optical flow, the fourth video frame and the fourth video frame.
- the feature map and the edge image are back-mapped to obtain a first back-mapped video frame and first back-mapping reference information, where the first back-mapping reference information includes the feature map and the edge image of the first back-mapped video frame.
- the terminal determines the third forward optical flow from the intermediate moment to the first moment based on the first backward optical flow based on the frame interpolation model, and based on the third forward optical flow, the third video frame and the third video frame
- the feature map and the edge image are back-mapped to obtain a second back-mapping video frame and second back-mapping reference information.
- the second back-mapping reference information includes the feature map and the edge image of the second back-mapping video frame.
- the terminal combines the first backward mapping video frame, the first backward mapping reference information, the second backward mapping video frame and the second backward mapping reference information into a backward mapping result.
- the terminal fuses the forward mapping result and the backward mapping result through the frame interpolation model to obtain the reference video frame.
- the terminal when the terminal fuses the forward mapping result and the backward mapping result through the frame interpolation model to obtain the reference video frame, the terminal encodes the forward mapping result through the frame interpolation model to obtain the forward intermediate feature To encode the result of the backward mapping to obtain the backward intermediate features.
- the terminal fuses the forward intermediate features and the backward intermediate features through the frame interpolation model to obtain the fused intermediate features.
- the terminal decodes the fused intermediate features through the frame interpolation model to obtain the reference video frame.
- Example 2 The terminal obtains the motion vectors of multiple image blocks in the third video frame in the fourth video frame through the frame interpolation model, and the third video frame is the previous second video frame of the two adjacent second video frames. A video frame, the fourth video frame is the last second video frame in two adjacent second video frames. The terminal generates the reference video frame based on the motion vector.
- the terminal can obtain the motion vector of the image block in the video frame, and generate a reference video frame based on the motion vector.
- the generated reference video frame has high quality, and the terminal can play the video frame better.
- the terminal obtains the encoding information of the third video frame and the fourth video frame through the frame interpolation model, and the encoding information is used to indicate the manner of dividing the third video frame and the fourth video frame and the third Motion vectors of each image block in the video frame from the third video frame to the fourth video frame.
- the terminal divides the third video frame and the fourth video frame into multiple image blocks based on the encoding information of the third video frame through the frame interpolation model, and determines the motion of each image block from the third video frame to the fourth video frame vector.
- the terminal processes the motion vector corresponding to each image block to obtain the target motion vector corresponding to each image block.
- the terminal generates the reference video frame according to each image block and the target motion vector of each image block.
- the terminal processes the motion vector corresponding to each image block through the frame interpolation model, including dividing the motion vector corresponding to each image block by the target value, that is, shortening the motion vector of the image block while ensuring that the motion direction of the image block remains unchanged.
- the movement distance and the target value are set by technicians according to the actual situation, which is not limited in this embodiment of the present application.
- the frame interpolation model can also be other types of frame interpolation models, such as Real-Time Intermediate Flow Estimation for Video Frame Interpolation, RIFE), Video Frame Interpolation via Residue Refinement, RRIN, and Multiple Video Frame Interpolation via Enhanced Deformable Separable Convolution, EDSC), etc., which are not limited in this embodiment of the present application.
- RIFE Real-Time Intermediate Flow Estimation for Video Frame Interpolation
- RRIN Video Frame Interpolation via Residue Refinement
- EDSC Enhanced Deformable Separable Convolution
- the terminal first performs super-resolution on the first video frame (step 403), and then performs frame interpolation on multiple second video frames after super-resolution (step 404).
- the terminal may also first perform frame interpolation on multiple first video frames, and then perform super-resolution on the multiple first video frames after frame interpolation, which is not limited in this embodiment of the present application.
- the terminal plays the multiple second video frames after frame insertion.
- the multiple second video frames played by the terminal not only have a higher resolution, but also have a larger number, so that the effect of playing video frames is better .
- the video frame playback method provided by the embodiment of the present application will be described below in conjunction with FIG. 5 and the above steps 401-405.
- the video frame playback method provided by the embodiment of the present application includes:
- Step 501 The terminal starts the cloud game client.
- Step 502 The terminal determines corresponding video stream information according to the network delay information.
- the video stream information includes resolution, frame rate and bit rate.
- Step 503 Determine whether the network status of the terminal is good, if yes, execute step 505, otherwise, execute step 504.
- the cloud game server determines rendering parameters based on the video stream information, renders the target virtual scene, and obtains the first video frame. Determine whether the network status of the terminal is good according to the network delay information.
- the network delay information indicates that the current network status of the terminal is not good (that is, the network delay is greater than or equal to the network delay threshold)
- step 504 is triggered.
- the network delay information indicates that the current network status of the terminal is good (that is, when the network delay is less than the network delay threshold)
- step 505 is triggered.
- Step 504 The cloud game server notifies the terminal to reduce the resolution and frame rate.
- Step 505 super-resolution and frame interpolation.
- the terminal makes judgments based on the acquired multiple first video frames, but when super-resolution and frame interpolation are required, super-resolution and frame interpolation are performed on the acquired first video frames to obtain multiple second video frames after frame interpolation , the terminal plays the plurality of second video frames, and the plurality of second video frames are cloud game screens after super-substituting and interpolating frames.
- the current network or terminal configuration is not good, and the video stream received from the cloud game server is 480P/15FPS, 720P/30FPS or 1080P/60FPS, etc. Users can choose whether to oversubscribe and insert frames according to their own experience habits and terminal computing power to 720P/30FPS, 1080P/60FPS or 4K/90FPS etc.
- Figure 6 and Figure 7 provide the first frame delay and freeze rate changes after using the video frame playback method provided by the embodiment of the present application. It can be seen from Figure 6 that after using the video frame playback method provided by the embodiment of the present application , the first frame delay has been reduced by more than 100ms, and the freeze rate has been reduced by more than 50%.
- the judgment is made based on the resolutions of the multiple first video frames, and when the resolution meets the resolution adjustment condition, the A plurality of first video frames are adjusted in resolution, and the purpose of resolution adjustment is to improve the resolution of the first video frames to obtain a plurality of second video frames, and the plurality of second video frames also have higher resolution, thereby
- the display effect of cloud games has been improved.
- the cloud game server can switch to low-resolution and low-FPS acquisition and encoding, and the low-resolution and low-FPS can greatly reduce the code rate of cloud game server acquisition and encoding, reducing network
- the network bandwidth pressure of unstable terminals makes the transmission more stable and smooth.
- cloud game server capture and encoding 720P/30FPS can save more than 400% of the network bandwidth compared with 1080P/60FPS bit rate, and then the terminal can save more than 400% of the network bandwidth through frame insertion and Super resolution technology super resolution interpolation frame to 1080P/60FPS, cloud game image quality experience is within 5% of the subjective experience MOS (Mean Opinion Score, subjective score) compared with capture and encoding 1080P/60FPS, but the bit rate is saved by more than 400%.
- MOS Mean Opinion Score, subjective score
- FIG. 8 is a schematic structural diagram of a video frame playback device provided by an embodiment of the present application.
- the device includes: a video frame acquisition module 801 , a resolution adjustment module 802 and a playback module 803 .
- the video frame acquisition module 801 is configured to acquire a plurality of first video frames, and the plurality of first video frames are video frames obtained by rendering the target virtual scene by the cloud game server.
- the resolution adjustment module 802 is configured to adjust the resolution of each first video frame respectively when the resolutions of the plurality of first video frames meet the resolution adjustment conditions, to obtain corresponding second video frames, and the resolution of the second video frames is The resolution is higher than that of the corresponding first video frame.
- the playing module 803 is configured to play the plurality of second video frames obtained through resolution adjustment.
- the resolution adjustment module 802 is further configured to, for each of the first video frames, insert reference pixels between every two pixels in each first video frame to obtain each first video frame Corresponding to the second video frame, the reference pixel is generated based on every two pixels.
- the resolution adjustment module 802 is configured to perform any of the following:
- the reference pixel is inserted between every two pixels in each first video frame by using the nearest neighbor interpolation method.
- the reference pixel is inserted between every two pixels in each first video frame by using a bilinear interpolation method.
- the reference pixel is inserted between every two pixels in each first video frame by using a mean value interpolation method.
- the resolution adjustment module 802 is further configured to input the respective first video frames into the super-resolution model, and the super-resolution model performs up-sampling on each first video frame to obtain the corresponding second video frame.
- the resolution adjustment module 802 is configured to, for any first video frame in the plurality of first video frames, perform feature extraction on the first video frame through the super-resolution model to obtain the first video
- the first video frame feature of the frame is nonlinearly mapped to the first video frame feature to obtain the second video frame feature of the first video frame, and the video frame reconstruction is performed based on the second video frame feature to obtain the first video frame corresponding to the second video frame of .
- the device also includes:
- the frame interpolation module is configured to perform frame interpolation between multiple second video frames to obtain multiple second video frames after frame interpolation.
- the playing module 803 is further configured to play the multiple second video frames after frame insertion.
- the frame interpolation module is configured to insert a reference video frame between every two second video frames when the frame rate of the plurality of second video frames is less than or equal to the frame rate threshold, to obtain the interpolation The number of second video frames after the frame.
- the apparatus further includes a reference frame determination module configured to perform any of the following:
- Any second video frame in every two second video frames is determined as the reference video frame.
- the average video frame of every two second video frames is determined as the reference video frame, and the pixel value of the pixel point in the average video frame is the average value of the pixel value of the corresponding pixel point in every two second video frames.
- a second video frame sequence composed of a plurality of second video frames is input into a frame interpolation model, and the frame interpolation model generates a reference video frame based on two adjacent second video frames in the second video frame sequence.
- the reference frame determination module is configured to use the frame interpolation model, the backward optical flow and the forward optical flow of two adjacent second video frames in the second video frame sequence, based on the backward optical flow Forward optical flow and the forward optical flow, generate the reference video frame; for example, obtain the backward optical flow of the third video frame to the fourth video frame, the third video frame is the front of each two second video frames A second video frame, the fourth video frame is the last second video frame in every two second video frames. Obtain the forward optical flow from the fourth video frame to the third video frame. Based on the backward optical flow and the forward optical flow, the reference video frame is generated.
- the reference frame determination module is configured to obtain motion vectors of multiple image blocks in the third video frame in the fourth video frame through the frame interpolation model, and the third video frame is two adjacent The previous second video frame in the second video frame, the fourth video frame is the next second video frame in the two adjacent second video frames. Based on the motion vector, the reference video frame is generated.
- the device also includes:
- the sending module is configured to send network delay information to the cloud game server, so that the cloud game server generates the plurality of first video frames based on the network delay information.
- the division of the above-mentioned functional modules is used as an example for illustration.
- the above-mentioned function allocation can be completed by different functional modules according to needs , that is, divide the internal structure of the computer device into different functional modules, so as to complete all or part of the functions described above.
- the device for playing video frames provided by the above embodiments and the embodiments of the method for playing video frames belong to the same concept, and the implementation process thereof is detailed in the method embodiments, and will not be repeated here.
- the judgment is made based on the resolutions of the multiple first video frames, and when the resolution meets the resolution adjustment condition, the The resolution of each first video frame is adjusted, and the purpose of the resolution adjustment is to increase the resolution of the first video frame to obtain the corresponding second video frame, and then play multiple first video frames obtained by increasing the resolution of the first video frame.
- the second video frame improves the clarity of the cloud game screen, thereby improving the display effect of the cloud game while ensuring the fluency of the cloud game.
- the embodiment of the present application provides a computer device configured to execute the video frame playback method provided in the embodiment of the present application.
- the computer device can be implemented as a terminal or a server.
- the structure of the terminal is firstly introduced below:
- FIG. 9 is a schematic structural diagram of a terminal provided by an embodiment of the present application.
- the terminal 900 may be: a smart phone, a tablet computer, a notebook computer or a desktop computer.
- the terminal 900 may also be called user equipment, portable terminal, laptop terminal, desktop terminal and other names.
- a terminal 900 includes: one or more processors 901 and one or more memories 902 .
- the processor 901 may include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like.
- the processor 901 can adopt at least one hardware form in Digital Signal Processing (Digital Signal Processing, DSP), Field-Programmable Gate Array (Field-Programmable Gate Array, FPGA), Programmable Logic Array (Programmable Logic Array, PLA) accomplish.
- the processor 901 may also include a main processor and a coprocessor, the main processor is a processor for processing data in the wake-up state, and is also called a central processing unit (Central Processing Unit, CPU); the coprocessor is Low-power processor for processing data in standby state.
- CPU Central Processing Unit
- the processor 901 may be integrated with a resolution adjuster (Graphics Processing Unit, GPU), and the GPU is used for rendering and drawing the content that needs to be displayed on the display screen.
- the processor 901 may also include an artificial intelligence (Artificial Intelligence, AI) processor, where the AI processor is used to process computing operations related to machine learning.
- AI Artificial Intelligence
- Memory 902 may include one or more computer-readable storage media, which may be non-transitory.
- the memory 902 may also include high-speed random access memory and non-volatile memory, such as one or more magnetic disk storage devices and flash memory storage devices.
- the non-transitory computer-readable storage medium in the memory 902 is used to store at least one computer program, and the at least one computer program is used to be executed by the processor 901 to implement the methods provided by the method embodiments in this application. Video frame playback method.
- the terminal 900 further includes: a peripheral device interface 903 and at least one peripheral device.
- the processor 901, the memory 902, and the peripheral device interface 903 may be connected through buses or signal lines.
- Each peripheral device can be connected to the peripheral device interface 903 through a bus, a signal line or a circuit board.
- the peripheral equipment includes: at least one of a radio frequency circuit 904 , a display screen 905 , a camera component 906 , an audio circuit 907 , a positioning component 908 and a power supply 909 .
- the peripheral device interface 903 may be used to connect at least one peripheral device related to input/output (Input/Output, I/O) to the processor 901 and the memory 902.
- the processor 901, memory 902 and peripheral device interface 903 are integrated on the same chip or circuit board; in some other embodiments, any one of the processor 901, memory 902 and peripheral device interface 903 or The two can be implemented on a separate chip or circuit board, which is not limited in this embodiment.
- the radio frequency circuit 904 is used to receive and transmit radio frequency (Radio Frequency, RF) signals, also called electromagnetic signals.
- the radio frequency circuit 904 communicates with the communication network and other communication devices through electromagnetic signals.
- the radio frequency circuit 904 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals.
- the radio frequency circuit 904 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and the like.
- the display screen 905 is used to display a user interface (User Interface, UI).
- the UI can include graphics, text, icons, video, and any combination thereof.
- the display screen 905 also has the ability to collect touch signals on or above the surface of the display screen 905 .
- the touch signal can be input to the processor 901 as a control signal for processing.
- the display screen 905 can also be used to provide virtual buttons and/or virtual keyboards, also called soft buttons and/or soft keyboards.
- the camera assembly 906 is used to capture images or videos.
- the camera assembly 906 includes a front camera and a rear camera.
- the front camera is set on the front panel of the terminal, and the rear camera is set on the back of the terminal.
- Audio circuitry 907 may include a microphone and speakers.
- the microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals and input them to the processor 901 for processing, or input them to the radio frequency circuit 904 to realize voice communication.
- the positioning component 908 is used to locate the current geographic location of the terminal 900, so as to realize navigation or location-based service (Location Based Service, LBS).
- LBS Location Based Service
- the power supply 909 is used to supply power to various components in the terminal 900 .
- the power source 909 can be alternating current, direct current, disposable batteries or rechargeable batteries.
- FIG. 9 does not constitute a limitation on the terminal 900, and may include more or less components than shown in the figure, or combine certain components, or adopt different component arrangements.
- the embodiment of the present application also provides a computer-readable storage medium, for example, a memory including a computer program, and the above computer program can be executed by a processor to complete the video frame playing method in the above embodiment.
- the computer-readable storage medium can be a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a read-only optical disc (Compact Disc Read-Only Memory, CD-ROM), Magnetic tapes, floppy disks, and optical data storage devices, etc.
- the embodiment of the present application also provides a computer program product or computer program, the computer program product or computer program includes program code, the program code is stored in a computer-readable storage medium, and the processor of the computer device reads from the computer-readable storage medium The program code is read, and the processor executes the program code, so that the computer device executes the above video frame playing method.
- the computer programs involved in the embodiments of the present application can be deployed and executed on one computer device, or executed on multiple computer devices at one location, or distributed in multiple locations and communicated Executed on multiple computer devices interconnected by the network, multiple computer devices distributed in multiple locations and interconnected through a communication network can form a blockchain system.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Signal Processing (AREA)
- Evolutionary Computation (AREA)
- Computing Systems (AREA)
- Biomedical Technology (AREA)
- Biophysics (AREA)
- Computational Linguistics (AREA)
- Data Mining & Analysis (AREA)
- Life Sciences & Earth Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Molecular Biology (AREA)
- Artificial Intelligence (AREA)
- General Engineering & Computer Science (AREA)
- Mathematical Physics (AREA)
- Software Systems (AREA)
- Health & Medical Sciences (AREA)
- Computer Networks & Wireless Communication (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Databases & Information Systems (AREA)
- Processing Or Creating Images (AREA)
Abstract
本申请公开了一种视频帧播放方法、装置、设备、存储介质及程序产品,能够应用在云游戏、人工智能等领域。本申请实施例获取多个第一视频帧,所述多个第一视频帧是云游戏服务器对目标虚拟场景进行渲染得到的视频帧;当所述多个第一视频帧的分辨率符合分辨率调整条件时,分别对各所述第一视频帧进行分辨率调整,得到相应的第二视频帧,所述第二视频帧的分辨率高于对应的第一视频帧的分辨率;播放进行分辨率调整所得到的多个所述第二视频帧。
Description
相关申请的交叉引用
本申请基于申请号为202111130391.5、申请日为2021年09月26日的中国专利申请提出,并要求该中国专利申请的优先权,该中国专利申请的全部内容在此引入本申请作为参考。
本申请涉及计算机技术及云游戏领域,尤其涉及一种视频帧播放方法、装置、设备、存储介质及程序产品。
随着云计算技术的成熟,用户可以通过云计算来实现终端难以完成的任务。例如,在云游戏的领域,终端无需执行渲染操作,云游戏服务器能够基于终端发送的控制信息对游戏场景进行渲染,得到视频帧。云游戏服务器将该视频帧发送给终端,终端显示该视频帧即可。
相关技术中,当终端所处的网络状态不佳时,云游戏服务器会降低视频帧的分辨率以保证云游戏的流畅程度,但是降低视频帧的分辨率会导致云游戏的显示效果变差。
发明内容
本申请实施例提供了一种视频帧播放方法、装置、设备、存储介质及程序产品,可以在保证云游戏的流畅程度的前提下,提高终端所显示云游戏画面的清晰度。
本申请实施例提供了一种视频帧播放方法,所述方法包括:
获取多个第一视频帧,所述多个第一视频帧是云游戏服务器对目标虚拟场景进行渲染得到的视频帧;
当所述多个第一视频帧的分辨率符合分辨率调整条件时,分别对各所述第一视频帧进行分辨率调整,得到相应的第二视频帧,所述第二视频帧的分辨率高于对应的第一视频帧的分辨率;
播放进行分辨率调整所得到的多个第二视频帧。
本申请实施例提供了一种视频帧播放装置,所述装置包括:
视频帧获取模块,配置为获取多个第一视频帧,所述多个第一视频帧是云游戏服务器对所述目标虚拟场景进行渲染得到的视频帧;
分辨率调整模块,配置为当所述多个第一视频帧的分辨率符合分辨率调整条件时,分别对各所述第一视频帧进行分辨率调整,得到相应的第二视频帧,所述第二视频帧的分辨率高于对应的第一视频帧的分辨率;
播放模块,配置为播放进行分辨率调整所得到的多个第二视频帧。
本申请实施例提供了一种计算机设备,所述计算机设备包括一个或多个处理器和一个或多个存储器,所述一个或多个存储器中存储有至少一条计算机程序,所述计算机程序由所述一个或多个处理器加载并执行,以实现本申请实施例提供的所述视频帧播放方法。
本申请实施例提供了一种计算机可读存储介质,所述计算机可读存储介质中存储有至少一条计算机程序,所述计算机程序由处理器加载并执行,以实现本申请实施例提供的所述视频帧播放方法。
本申请实施例提供了一种计算机程序产品或计算机程序,该计算机程序产品或计算机程序包括程序代码,该程序代码存储在计算机可读存储介质中,计算机设备的处理器从计算机可读存储介质读取该程序代码,处理器执行该程序代码,使得该计算机设备执行本申请实施例提供的视频帧播放方法。
本申请实施例具有以下有益效果:
通过本申请实施例提供的技术方案,获取到云游戏服务器发送的多个第一视频帧之后,对多个第一视频帧的分辨率进行判断,当该分辨率符合分辨率调整条件时,对各第一视频帧分别进行分辨率调整,以提高第一视频帧的分辨率,得到相应的第二视频帧,进而播放调高第一视频帧的分辨率所得到的多个第二视频帧,相较于直接播放云游戏服务器发送的多个第一视频帧,提高了云游戏画面的清晰度,从而在保证云游戏流畅度的前提下,提高了云游戏的显示效果。
图1是本申请实施例提供的一种视频帧播放方法的实施环境的示意图;
图2是本申请实施例提供的一种云游戏的基本流程图;
图3是本申请实施例提供的一种视频帧播放方法的流程图;
图4是本申请实施例提供的一种视频帧播放方法的流程图;
图5是本申请实施例提供的一种视频帧播放方法的流程图;
图6是本申请实施例提供的一种首帧延时变化图;
图7是本申请实施例提供的一种卡顿率变化图;
图8是本申请实施例提供的一种视频帧播放装置的结构示意图;
图9是本申请实施例提供的一种终端的结构示意图。
为使本申请的目的、技术方案和优点更加清楚,下面将结合附图对本申请实施方式进行详细描述。
本申请中术语“第一”“第二”等字样用于对作用和功能基本相同的相同项或相似项进行区分,应理解,“第一”、“第二”、“第n”之间不具有逻辑或时序上的依赖关系,也不对数量和执行顺序进行限定。
本申请中术语“至少一个”是指一个或多个,“多个”的含义是指两个或两个以上,例如,多个参照人脸图像是指两个或两个以上的参照人脸图像。
在以下的描述中,涉及到“一些实施例”,其描述了所有可能实施例的子集,但是可以理解,“一些实施例”可以是所有可能实施例的相同子集或不同子集,并且可以在不冲突的情况下相互结合。
对本申请实施例进行进一步详细说明之前,对本申请实施例中涉及的名词和术语进行说明,本申请实施例中涉及的名词和术语适用于如下的解释。
1)、云游戏(Cloud Gaming)又可称为游戏点播(Gaming on Demand),是一种以云计算技术为基础的在线游戏技术。云游戏技术使图形处理与数据运算能力相对有限的轻端设备(Thin Client)能运行高品质游戏。在云游戏场景下,游戏并不在玩家 游戏终端,而是在云游戏服务器中运行,并由云游戏服务器将游戏场景渲染为视频音频流,通过网络传输给玩家游戏终端。玩家游戏终端无需拥有强大的图形运算与数据处理能力,仅需拥有基本的流媒体播放能力、获取玩家输入指令并发送给云游戏服务器的能力以及基础的数据处理能力即可。
2)、虚拟场景:是应用程序在终端上运行时显示(或提供)的虚拟场景。该虚拟场景可以是对真实世界的仿真环境,也可以是半仿真半虚构的虚拟环境,还可以是纯虚构的虚拟环境。虚拟场景可以是二维虚拟场景、2.5维虚拟场景或者三维虚拟场景中的任意一种,本申请实施例对虚拟场景的维度不加以限定。例如,虚拟场景可以包括天空、陆地、海洋等,该陆地可以包括沙漠、城市等环境元素,用户可以控制虚拟对象在该虚拟场景中进行移动。
3)、虚拟对象:是指在虚拟场景中的可活动对象。该可活动对象可以是虚拟人物、虚拟动物、动漫人物等,比如:在虚拟场景中显示的人物、动物、植物、油桶、墙壁、石块等。该虚拟对象可以是该虚拟场景中的一个虚拟的用于代表用户的虚拟形象。虚拟场景中可以包括多个虚拟对象,每个虚拟对象在虚拟场景中具有自身的形状和体积,占据虚拟场景中的一部分空间。
在一些实施例中,该虚拟对象是通过客户端上的操作进行控制的用户角色,或者是通过训练设置在虚拟场景对战中的人工智能(Artificial Intelligence,AI),或者是设置在虚拟场景中的非用户角色(Non-Player Character,NPC)。在一些实施例中,该虚拟对象是在虚拟场景中进行竞技的虚拟人物。在一些实施例中,该虚拟场景中参与互动的虚拟对象的数量是预先设置的,或者是根据加入互动的客户端的数量动态确定的。
以射击类游戏为例,用户能够控制虚拟对象在该虚拟场景的天空中自由下落、滑翔或者打开降落伞进行下落等,在陆地上中跑动、跳动、爬行、弯腰前行等,也可以控制虚拟对象在海洋中游泳、漂浮或者下潜等,当然,用户也可以控制虚拟对象乘坐虚拟载具在该虚拟场景中进行移动,例如,该虚拟载具可以是虚拟汽车、虚拟飞行器、虚拟游艇等,在此仅以上述场景进行举例说明,本申请实施例对此不作限定。用户也可以控制虚拟对象通过互动道具与其他虚拟对象进行战斗等方式的互动,例如,该互动道具可以是虚拟手雷、虚拟集束雷、虚拟粘性手雷(简称“粘雷”)等投掷类互动道具,也可以是虚拟机枪、虚拟手枪、虚拟步枪等射击类互动道具,本申请对互动道具的类型不作限定。
4)、显示分辨率:分辨率主要是指显示器所能显示的像素的多少,可以从显示分辨率与图像分辨率两个方向来分类。显示分辨率(屏幕分辨率)是屏幕图像的精密度,是指显示器所能显示的像素的多少。由于屏幕上的点、线和面都是由像素组成的,显示器可显示的像素越多,画面就越精细,同样的屏幕区域内能显示的信息也越多,所以分辨率是个非常重要的性能指标之一。可以把整个图像想象成是一个大型的棋盘,而分辨率的表示方式就是所有经线和纬线交叉点的数目。显示分辨率一定的情况下,显示屏越小图像越清晰,反之,显示屏大小固定时,显示分辨率越高图像越清晰。图像分辨率则是单位英寸中所包含的像素点数,其定义更趋近于分辨率本身的定义。
5)、插帧:即在原有画面显示的每两帧画面中增加一帧,缩短每帧之间的显示时间,使时间得到双倍提高,比如将视频帧率从原有的30Hz提升到60HZ。修正人眼视觉暂留形成的错觉,有效提高画面稳定性。
6)、超分:提高图像分辨率的方法,基于时序近邻帧辅助参考帧运动预估或运运补偿以及相关深度学习模型,将任意低分辨率提供向上采样到Nx(如2x)倍分辨率的技术方案,如将2K超分到4K分辨率。
7)、MOS(Mean Opinion Score):平均意见分,是主观评价实验之后,得到的主观分数,取值0-100,值越大,代表主观感受越好。
8)、FPS(Frames Per Second):也可以理解为我们常说的“刷新率(单位为Hz)”,FPS是图像领域中的定义,是指画面每秒传输帧数,通俗来讲就是指动画或视频的画面数。FPS是测量用于保存、显示动态视频的信息数量。每秒钟帧数越多,所显示的动作就会越流畅。
9)、SPS(Sequence Parameter Set):又称作序列参数集。SPS中保存了一组编码视频序列(Coded Video Sequence)的全局参数。所谓的编码视频序列即原始视频的一帧一帧的像素数据经过编码之后的结构组成的序列。
图1是本申请实施例提供的一种视频帧播放方法的实施环境示意图,参见图1,该实施环境中可以包括终端110和云游戏服务器140。在一些实施例中,终端110和云游戏服务器140为区块链系统中的节点,终端110和服务器140之间相互传输的数据存储在区块链上。
终端110通过无线网络或有线网络与云游戏服务器140相连。在一些实施例中,终端110是智能手机、平板电脑、笔记本电脑、台式计算机、智能音箱、智能手表等,但并不局限于此。终端110安装和运行有支持虚拟场景显示的客户端。
云游戏服务器140是独立的物理云游戏服务器,或者是多个物理云游戏服务器构成的云游戏服务器集群或者分布式系统,或者是提供云服务、云数据库、云计算、云函数、云存储、网络服务、云通信、中间件服务、域名服务、安全服务、分发网络(Content Delivery Network,CDN)、以及大数据和人工智能平台等基础云计算服务的云游戏服务器。在一些实施例中,云游戏服务器140也被称为边缘计算节点。
在一些实施例中,终端110泛指多个终端中的一个,本申请实施例仅以终端110来举例说明。
本领域技术人员可以知晓,上述终端的数量可以更多或更少。比如上述终端仅为一个,或者上述终端为几十个或几百个,或者更多数量,此时上述实施环境中还包括其他终端。本申请实施例对终端的数量和设备类型不加以限定。
在介绍完本申请实施例提供的视频帧播放方法的应用场景之后,下面将结合上述实施环境,对本申请实施例提供的视频帧播放方法的应用场景进行说明,在下述说明过程中,终端也即是上述终端110,云游戏服务器也即是上述云游戏服务器140。
本申请实施例提供的视频帧播放方法能够应用在各类云游戏的场景下,比如应用在第一人称射击类(First-person Shooting,FPS)游戏中,或者应用在第三人称射击游戏(Third-Personal Shooting,TPS)游戏中,或者应用在多人在线战术竞技游戏(Multiplayer Online Battle Arena,MOBA)中,或者应用在战棋游戏或者自走棋游戏中,本申请实施例对此不做限定。
以本申请实施例提供的视频帧播放方法应用在FPS游戏中为例,用户在终端上启动云游戏客户端,在云游戏客户端中登录用户账号,也即是,用户在云游戏客户端中输入用户账号和对应的密码,点击登录控件来进行登录。响应于检测到对登录控件的点击操作,终端向云游戏服务器发送登录请求,登录请求中携带有用户账号和对应的密码。云游戏服务器接收到登录请求之后,从登录请求中获取用户账号和对应的密码,对用户账号和对应的密码进行验证。云游戏服务器对用户账号和对应的密码验证通过之后,向终端发送登录成功信息。终端接收到登录成功信息之后,向云游戏服务器发送云游戏获取请求,云游戏获取请求中携带有用户账号。云游戏服务器获取云游戏获取请求之后,基于云游戏获取请求中携带的用户账号进行查询,获取该用户账号对应的多个云游戏,将该多个云游戏的标识发送给终端,由终端将该多个云游戏的标 识展现在云游戏客户端中。用户通过终端,在云游戏客户端中展示的多个云游戏的标识中选择想玩的FPS游戏的标识,也即是选择想玩的FPS游戏。用户在云游戏客户端中选择FPS游戏之后,终端向云游戏服务器发送游戏启动指令,游戏启动指令中携带有用户账号、该FPS游戏的标识以及终端的硬件信息,其中,终端的硬件信息包括终端的屏幕分辨率、终端的型号等,本申请实施例对此不做限定。云游戏服务器接收到该游戏启动指令之后,从该游戏启动指令中获取用户账号、该FPS游戏的标识以及终端的硬件信息。云游戏服务器基于终端的硬件信息对该FPS游戏进行初始化,以实现渲染得到的游戏画面与终端之间的匹配。云游戏服务器启动该FPS游戏。在运行该FPS游戏的过程中,用户能够通过终端控制该FPS中的被控虚拟对象进行移动,也即是终端将对被控虚拟对象的控制信息发送给云游戏服务器,云游戏服务器基于该控制信息对该FPS的虚拟场景进行渲染,得到第一视频帧。参见图2,终端向云端服务器(云游戏服务器)实时传输交互操作(控制信息)。云端服务器基于接收到的交互操作进行渲染计算,向终端返回压缩的音视频流。终端对接收到的音视频流进行解码和播放。
在一些实施例中,云游戏服务器在渲染虚拟场景时,除了参考终端的屏幕分辨率以及型号之外,还会参考终端的网络延迟信息。若终端的网络延迟信息指示终端当前的网络延迟较高,那么云游戏服务器能够以较低的质量来渲染虚拟场景,得到的第一视频帧的分辨率也就较低,相应的,第一视频帧占用的网络带宽也就较小,最大程度的保证终端上运行的FPS游戏的流畅程度,比如,当终端网络延迟较低时,云游戏服务器渲染出的第一视频帧的分辨率为1080p;当终端网络延迟较高时,云游戏服务器渲染出的第一视频帧的分辨率能够降低到720p。当出现第一视频帧的分辨率降低的情况时,终端能够采用本申请实施例提供的视频帧播放方法,对接收到的第一视频帧进行分辨率调整,以提高第一视频帧的分辨率,得到第二视频帧。相较于第一视频帧来说,分辨率调整后得到的第二视频帧具有更高的分辨率,也就是说第二视频帧具有更好的显示效果。这样就能够在终端网络出现波动时,在保证该FPS流畅度的前提下,提高该PFS的显示效果。尤其是对于网络不稳定但是具有一定算力的终端来说,能够提高用户的游戏体验。
在一些实施例中,云游戏服务器向终端发送第一视频帧时,不是一帧一帧发送的,而是会渲染出一个视频帧序列,该视频帧序列包括多个第一视频帧,云游戏服务器每次会向终端发送该视频帧序列以供终端进行显示。当终端的网络延迟信息指示终端当前的网络延迟较高,云游戏服务器除了能够降低渲染得到的第一视频帧的分辨率之外,还能够减少视频帧序列中第一视频帧的数量,减少视频帧序列中第一视频帧的数量也即是降低终端在显示第一视频帧的帧率,比如,当终端的网络延迟较低时,该视频帧序列携带60个第一视频帧,这60个第一视频帧由终端在1s内均匀显示,此时的帧率也即是60;当终端的网络延迟较低时,该视频帧序列携带的第一视频帧的数量降低至30,这3个第一视频帧由终端在1s内均匀显示,此时的帧率也即是30。可以看出,通过降低帧率的方式能够减少传输视频帧的数量,也就减少了传输视频帧时对带宽的占用,但是降低帧率会导致显示的“顿挫感”。在这种情况下,终端能够采用本申请实施例提供的视频帧播放方法,在接收到的视频帧序列中进行插帧,以增加该视频帧序列中视频帧的数量。消除“顿挫感”,提高视频帧的播放效果。
另外,对于MOBA游戏、TPS游戏、战棋游戏以及自走棋游戏来说,均能够采用上述步骤来进行处理,在此不再赘述。
还有,本申请实施例提供的视频帧播放方法除了能够应用在上述FPS游戏、MOBA游戏、战棋游戏或者自走棋游戏之外,也能够应用在其他类型的云游戏中, 本申请实施例对此不做限定。
在介绍完本申请实施例提供的视频帧播放方法的实施环境和应用场景之后,下面对本申请实施例提供的视频帧播放方法进行说明。
图3是本申请实施例提供的一种视频帧播放方法的流程图,参见图3,方法包括:
301、终端获取多个第一视频帧,该多个第一视频帧是云游戏服务器对目标虚拟场景进行渲染得到的视频帧。
这里,终端上运行有客户端,如游戏客户端或具有游戏功能的其它客户端(如即时通讯客户端),用户在基于客户端玩游戏的过程中,终端所显示的游戏画面通过云游戏服务器对虚拟场景进行渲染所得到。在实际应用中,云游戏服务器可以以目标虚拟场景中被控虚拟对象的视角,对该目标虚拟场景进行渲染,得到终端待展示的视频帧。
云游戏服务器对目标虚拟场景进行渲染得到多个视频帧,该多个视频帧可以以视频帧序列的形式存在,并发送视频帧序列给终端。
其中,被控虚拟对象也即是终端控制的虚拟对象,即客户端所登录的用户账号对应的虚拟对象。以目标虚拟场景中被控虚拟对象的视角,对该目标虚拟场景进行渲染是指:对目标虚拟场景中,被控虚拟对象观察到的画面进行渲染,得到第一视频帧,被控虚拟对象在虚拟场景中观察到的画面,也即是用户看到的画面。
302、终端在该多个第一视频帧的分辨率符合分辨率调整条件时,分别对各个第一视频帧进行分辨率调整,得到相应的第二视频帧,第二视频帧的分辨率高于对应的第一视频帧的分辨率。
在一些实施例中,终端在接收到云游戏服务器发送的多个视频帧(视频帧序列)后,获取该多个视频帧的分辨率,在一些实施例中,该多个视频帧的分辨率相同,终端将该多个视频帧的分辨率与分辨率阈值进行比对,得到比对结果,并当该比对结果表征该多个第一视频帧的分辨率小于或等于分辨率阈值时,确定该多个第一视频帧的分辨率符合分辨率调整条件。
其中,终端对多个第一视频帧进行分辨率调整的过程,也即是对多个第一视频帧进行超分的过程,超分能够提高第一视频帧的分辨率,使得终端相较于直接播放云游戏服务器发送的多个第一视频帧,提高了云游戏画面的清晰度,也即提高了视频帧的显示效果。
303、终端播放进行分辨率调整所得到的多个第二视频帧。
其中,第二视频帧是终端对第一视频帧进行分辨率调整后得到的,那么第二视频帧也就与对应的第一视频帧具有相同的图像内容,相较于显示第一视频帧来说,由于第二视频帧具有更高的分辨率,终端显示第二视频帧时清晰度会更高,也就具有更好的显示效果。
通过本申请实施例提供的技术方案,获取到云游戏服务器发送的多个第一视频帧之后,基于多个第一视频帧的分辨率进行判断,当该分辨率符合分辨率调整条件时,对多个第一视频帧进行分辨率调整,分辨率调整的目的是提高第一视频帧的分辨率,得到多个第二视频帧,多个第二视频帧也就具有更高的分辨率,从而在保证云游戏流畅度的前提下,提高了云游戏的显示效果。
下面将结合一些例子,对本申请实施例提供的技术方案进行更加清楚地描述,参见图4,方法包括:
401、终端向云游戏服务器发送网络延迟信息,以使云游戏服务器基于网络延迟信息生成多个第一视频帧。
这里,该第一视频帧是该云游戏服务器以目标虚拟场景中被控虚拟对象的视角,对该目标虚拟场景进行渲染得到的视频帧。
其中,网络延迟信息用于表示终端与云游戏服务器之间的网络延迟。目标虚拟场景也即是用户选择的云游戏的游戏场景,被控虚拟对象的视角也即是被控虚拟对象的虚拟摄像机的视角,在FPS游戏中,被控虚拟对象的虚拟摄像机位于被控虚拟对象的头部,用户通过终端控制被控虚拟对象在目标虚拟场景中进行移动时,该虚拟摄像机也会随着被控虚拟对象的移动而移动,该虚拟摄像机拍摄到的画面也即是被控虚拟对象在目标虚拟场景中观察到的画面。在TPS游戏中,被控虚拟对象的虚拟摄像机位于被控虚拟对象的上方,用户通过终端控制被控虚拟对象在目标虚拟场景中进行移动时,该虚拟摄像机也会随着被控虚拟对象的移动而移动,该虚拟摄像机拍摄到的画面也即是在被控虚拟对象上方观察到的画面。在云游戏场景下,该虚拟摄像机拍摄到的画面是由云游戏服务器渲染的,由于游戏过程中存在多个时序上连续的多个画面,该多个画面在本申请中被称为视频帧。
在一些实施例中,终端启动云游戏客户端,通过云游戏客户端获取终端与云游戏服务器之间的网络延迟信息,终端将该网络延迟信息发送给云游戏服务器。云游戏服务器接收到该网络延迟信息之后,根据该网络延迟信息确定对应的渲染参数,采用该渲染参数以目标虚拟场景中被控虚拟对象的视角,对该目标虚拟场景进行渲染,得到多个第一视频帧。
在这种实施方式下,终端能够将网络延迟信息发送给云游戏服务器,云游戏服务器能够基于该网络延迟信息来确定渲染参数,该网络延迟信息能够反映终端当前的网络状况,采用该渲染参数渲染得到的视频帧也就与终端网络状况相适应,从而保证终端运行云游戏的流畅程度。
下面通过两个例子对上述实施方式进行说明。
例1、终端启动云游戏客户端,通过云游戏客户端向云游戏服务器发送探测数据包,该探测数据包用于请求云游戏服务器返回确认数据包。终端通过云游戏客户端,将接收到确认数据包和发送探测数据包之间的时间差值,确定为该网络延迟信息,终端向云游戏服务器发送该网络延迟信息。云游戏服务器接收到该网络延迟信息之后,确定该网络延迟信息对应的渲染参数,该渲染参数用于指示渲染得到的视频帧的分辨率。云游戏服务器采用该渲染参数,以目标虚拟场景中被控虚拟对象的视角对该目标虚拟场景进行渲染,得到多个第一视频帧。
其中,在云游戏服务器采用该渲染参数,以目标虚拟场景中被控虚拟对象的视角对该目标虚拟场景进行渲染时,渲染工作由云游戏服务器的图形处理器(Graphics Processing Unit,GPU)完成。云游戏服务器的图形处理器在对该目标虚拟场景渲染时,会将渲染得到的多个游戏画面存储在显存中。为了提高处理效率和降低延迟,云游戏服务器的图形处理器会直接对显存中的多个游戏画面进行编码,得到多个第一视频帧。在一些实施例中,云游戏服务器的图形处理器能够将显存中的多个游戏画面编码为VP8(谷歌研发且推出的视频格式)/VP9(谷歌研发且推出的视频格式)/H.264/H.265/高级视频编码(Advanced Video Coding,AVC)/音视频编码标准(Audio Video Coding Standard,AVS)等格式的第一视频帧,本申请实施例对此不做限定。另外,对于游戏画面对应的音频数据来说,云游戏服务器也能够将音频数据编码为Silk(微软开发的一种音频格式)/Opus(一种开源的音频格式)/高级音频编码(Advanced Audio Coding,AAC)等格式的音频数据流。
例2、终端启动云游戏客户端,通过云游戏客户端向云游戏服务器发送测试数据下载请求,该测试数据下载请求用于请求从云游戏服务器下载测试数据。终端通过云 游戏客户端,从云游戏服务器下载该测试数据,下载的时长由技术人员根据实际情况设置,比如设置为1s或3s等,本申请实施例对此不做限定。终端将下载的数据量与下载时间相除,得到该网络延迟信息。终端向云游戏服务器发送该网络延迟信息。云游戏服务器接收到该网络延迟信息之后,确定该网络延迟信息对应的渲染参数,该渲染参数用于指示渲染得到的视频帧的分辨率。云游戏服务器采用该渲染参数,以目标虚拟场景中被控虚拟对象的视角对该目标虚拟场景进行渲染,得到多个第一视频帧。
在上一种实施方式中,是以终端刚启动云游戏客户端为例进行说明的,下面对终端通过云游戏客户端运行云游戏的过程进行说明。
在一些实施例中,在云游戏客户端上运行云游戏时,终端通过云游戏客户端向云游戏服务器发送对被控虚拟对象的控制信息以及网络延迟信息。云游戏服务器接收到该控制信息和该网络延迟信息之后,基于该控制信息确定目标虚拟场景中被控虚拟对象的视角,基于该网络延迟信息确定对应的渲染参数。云游戏服务器基于该渲染参数以及该被控虚拟对象的视角,对该目标虚拟场景进行渲染,得到第一视频帧。
其中,对被控虚拟对象的控制信息用于改变被控虚拟对象在目标场景中的位置、朝向以及动作,比如,控制信息能够控制被控虚拟对象在目标虚拟场景中向前后左右进行移动,或者能够控制被控虚拟对象在目标虚拟场景中向左或者向右旋转,或者能够控制被控虚拟对象在目标虚拟场景中执行下蹲、匍匐以及使用虚拟道具等动作。当然,在云游戏服务器基于控制信息,控制被控虚拟对象在目标虚拟场景中运动或者执行动作时,与被控虚拟对象绑定的虚拟摄像机也会随着被控虚拟对象的运动而运动,被控虚拟对象进行运动或者执行动作会导致被控虚拟对象观察目标虚拟场景的视角发生变化,与被控虚拟对象绑定的虚拟摄像机能够记录这一变化。
在上述两种实施方式中,是以由终端将网络延迟信息发送给云游戏服务器,云游戏服务器基于该网络延迟信息来确定对应的渲染参数为例进行说明的,下面对云游戏服务器基于其他方式来确定渲染参数进行说明。
在一些实施例中,终端启动云游戏客户端,通过云游戏客户端获取终端与云游戏服务器之间的网络延迟信息。终端基于该网络延迟信息,确定视频流信息,该视频流信息包括视频流的分辨率、码率以及帧率,其中,视频流包括多个视频序列,每个视频序列包括多个第一视频帧。在一些实施例中,同一视频序列中的第一视频帧是云游戏服务器采用相同的渲染参数渲染得到的,也即,同一视频序列中的第一视频帧的分辨率相同,不同视频序列中的第一视频帧可能是云游戏服务器采用不同渲染参数渲染得到的,也即,不同视频序列中的第一视频帧的分辨率可以不同,。终端将该视频流信息发送给云游戏服务器,该云游戏服务器接收该视频流信息,基于该视频流信息确定对应的渲染参数。云游戏服务器采用该渲染参数以目标虚拟场景中被控虚拟对象的视角,对该目标虚拟场景进行渲染,得到多个第一视频帧。
在这种实施方式下,终端在获取到网络延迟信息之后,能够直接基于网络延迟信息确定视频流信息,云游戏服务器直接基于该视频流信息就能够快速确定对应的渲染参数,效率较高。
在一些实施例中,终端启动云游戏客户端,通过云游戏客户端获取终端与云游戏服务器之间的网络延迟信息。终端基于该网络延迟信息,显示视频流信息选择页面,该视频流信息选择页面中显示有候选的多个视频流信息,该多个视频流信息为与该网络延迟信息匹配的视频流信息。响应于多个视频流信息中目标视频流信息被选中,终端将该目标视频流信息发送给云游戏服务器,该云游戏服务器接收该目标视频流信息,基于该目标视频流信息确定对应的渲染参数。云游戏服务器采用该渲染参数以目标虚拟场景中被控虚拟对象的视角,对该目标虚拟场景进行渲染,得到多个第一视频 帧。
在这种实施方式下,终端在获取到网络延迟信息之后,能够基于该网络延迟信息为用户提供多个可供选择的视频流信息,用户选择视频流信息的过程也即是选择分辨率、帧率以及码率的过程,这样也就为用户提供的更高的自主性。
当然,在进行云游戏的过程中,用户也能够通过云游戏客户端随时调整选择的视频流信息,云游戏服务器也就能够相应的调整渲染参数。
402、终端获取多个第一视频帧。
其中,多个第一视频帧属于同一个视频帧序列,多个第一视频帧是云游戏服务器基于相同的渲染参数对目标虚拟场景进行渲染后得到的,也即是多个第一视频帧具有相同的分辨率。由于多个第一视频帧是云游戏服务器对游戏画面进行编码后得到的,该视频帧序列也即是一个编码视频序列(Coded Video Sequence,CVS)。
在一些实施例中,云游戏服务器在获取编码视频序列的同时,还会从服务器获取该编码视频序列对应的序列参数集(Sequence Parameter Set,SPS),该序列参数集用于指示终端如何对编码视频序列进行解码。终端基于该序列参数集,对该编码视频帧序列进行解码,得到多个第一视频帧。
403、终端在该多个第一视频帧的分辨率符合分辨率调整条件时,分别对各第一视频帧进行分辨率调整,得到相应的第二视频帧,第二视频帧的分辨率高于对应的第一视频帧的分辨率。
在实际应用中,终端在接收到云游戏服务器发送的第一视频帧序列后,获取第一视频帧序列中各第一视频帧的分辨率,在一些实施例中,该第一视频帧序列中各个第一视频帧的分辨率相同,终端将该多个视频帧的分辨率与分辨率阈值进行比对,得到比对结果,并当该比对结果表征该多个第一视频帧的分辨率小于或等于分辨率阈值时,确定该多个第一视频帧的分辨率符合分辨率调整条件。
在该多个第一视频帧的分辨率符合分辨率调整条件的情况下,终端在各个第一视频帧中每两个像素点之间插入参考像素点,得到各个第一视频帧对应的第二视频帧,该参考像素点是基于该每两个像素点生成的。其中,这里描述的第一视频帧中每两个像素点是指,第一视频帧中在空间上相邻的两个像素点。
在这种实施方式下,对于每个第一视频帧来说,终端能够在该第一视频帧中每两个像素点之间插入参考像素点。通过这种插入参考像素点的方式来实现对第一视频帧的超分,得到的第二视频帧也就具有更高的分辨率,终端显示该第二视频帧时也就具有比对应的第一视频帧更好的效果。
在一些实施例中,分辨率符合分辨率调整条件是指,分辨率小于或等于分辨率阈值,该分辨率阈值由技术人员根据实际情况进行设置,或者由用户根据终端的运算能力进行设置本申请实施例对此不做限定。
在一些实施例中,终端可采用如下方式对第一视频帧进行分辨率调整,得到相应的第二视频帧:
终端采用最临近插值法在各个第一视频帧中每两个像素点之间插入该参考像素点,得到各个第一视频帧对应的第二视频帧。
比如,在该多个第一视频帧的分辨率小于或等于分辨率阈值的情况下,以在第一视频帧中两个像素点之间插入参考像素点为例,终端在这两个像素点之间插入参考像素点,该参考像素点的像素值为初始值,比如为0。终端将这两个像素点中任一像素点的像素值更新该参考像素点的像素值。这一过程体现了“最临近”的思想,也即是将参考像素点的像素值确定为最临近的像素点的像素值,从而快速完成对第一视频帧的分辨率调整,以提高第一视频帧的分辨率,得到第二视频帧。终端能够采用上述方式 对各个第一视频帧进行处理,得到各个第一视频帧对应的第二视频帧。
需要说明的是,终端除了能够通过上述方式在各个第一视频帧中每两个像素点之间插入一个参考像素点之外,还能够插入多个参考像素点,下面以终端在各个第一视频帧中每两个像素点之间插入两个像素点为例进行说明。
比如,在该多个第一视频帧的分辨率小于或等于分辨率阈值的情况下,以在第一视频帧中两个像素点之间插入两个参考像素点为例,终端在这两个像素点之间插入第一参考像素点和第二参考像素点,该第一参考像素点和该第二参考像素点的像素值均为初始值,比如为0。终端采用这两个像素点中前一个像素点的像素值更新第一参考像素点的像素值,采用后一个像素点的像素值更新第二参考像素点的像素值,其中,这两个像素点中前一个像素点也即是与第一参考像素点之间距离较近的像素点,相应的,这两个像素点中后一个像素点也即是与第二参考像素点之间距离较近的像素点。下面对上述说明中“前一个像素点”和“后一个像素点”的区分方式进行说明。若这两个像素点在第一视频帧上从左至右排列,那么“前一个像素点”也即是这两个像素点中靠左的像素点;“后一个像素点”也即是这两个像素点中靠右的像素点。若这两个像素点在第一视频帧上从上至下排列,那么“前一个像素点”也即是这两个像素点中靠上的像素点;“后一个像素点”也即是这两个像素点中靠下的像素点。
在一些实施例中,终端可采用如下方式对第一视频帧进行分辨率调整,得到相应的第二视频帧:
终端采用双线性插值法在各个第一视频帧中每两个像素点之间插入该参考像素点,得到各个第一视频帧对应的第二视频帧。
在一些实施例中,在该多个第一视频帧的分辨率小于或等于分辨率阈值的情况下,终端在第一视频帧中每两个像素点之间插入该参考像素点,该参考像素点的像素值为初始像素值,比如为0。终端基于该参考像素点与每两个像素点之间的距离以及这两个像素点的像素值,更新该参考像素点的像素值。
以在第一视频帧中两个像素点之间插入两个参考像素点为例,终端在这两个像素点之间插入第一参考像素点和第二参考像素点,第一参考像素点和第二参考像素点的像素值均为初始像素值,比如为0。终端基于第一参考像素点与这两个像素点之间的距离,确定第一参考像素点与这两个像素点之间的两个第一权重,第一权重与距离正相关。终端基于第二考像素点与这两个像素点之间的距离,确定第二参考像素点与这两个像素点之间的两个第二权重,第二权重与距离正相关。终端基于两个第一权重,将这两个像素点的像素值进行加权求和,得到第一像素值,采用第一像素值更新第一参考像素点的像素值。终端基于两个第二权重,将这两个像素点的像素值进行加权求和,得到第二像素值,采用第二像素值更新第二参考像素点的像素值。
需要说明的是,上述是以在第一视频帧的每两个像素点之间插入两个参考像素点为例进行说明的,在其他可能的实施方式中,终端能够采用上述方式在每两个像素点之间插入三个或者更多参考像素点,本申请是实施例对于插入参考像素点的数量不做限定。
在一些实施例中,终端可采用如下方式对第一视频帧进行分辨率调整,得到相应的第二视频帧:
终端采用均值插值法在各个第一视频帧中每两个像素点之间插入该参考像素点,得到各个第一视频帧对应的第二视频帧。
在一些实施例中,在该多个第一视频帧的分辨率小于或等于分辨率阈值的情况下,终端在第一视频帧中每两个像素点之间插入该参考像素点,该参考像素点的像素值为初始像素值,比如为0。终端基于每两个像素点的像素值的平均值,更新该参考 像素点的像素值。
在一些实施例中,在该多个第一视频帧的分辨率符合分辨率调整条件的情况下,终端将该多个第一视频帧输入超分模型,由该超分模型对该多个第一视频帧进行上采样,得到该多个第二视频帧。
下面通过几个例子对上述实施方式中超分模型的上采样过程进行说明。
例1、对于该多个第一视频帧中的任一第一视频帧,终端通过该超分模型对该第一视频帧进行特征提取,得到该第一视频帧的第一视频帧特征。终端通过该超分模型对该第一视频帧特征进行非线性映射,得到该第一视频帧的第二视频帧特征。终端通过该超分模型对该第二视频帧特征进行重构,得到该第一视频帧对应的第二视频帧。在一些实施例中,例1提供的上采样方式也被称为后采样超分,采用后采样超分能够让超分模型自适应的学习上采样过程,还能让特征提取过程在低维空间上进行,极大的降低计算负担,训练速度和推理速度较快。
比如,终端将该第一视频帧输入该超分模型,通过该超分模型的至少一个卷积层对第一视频帧进行卷积处理,得到该第一视频帧的第一视频帧特征。终端通过该超分模型的全连接层和激活层,对该第一视频帧特征进行全连接和非线性激活,得到该第一视频帧的第二视频帧特征。终端通过该超分模型的重构层,基于该第二视频帧特征进行重构,这里的重构也即是上采样,上采样能够增加第一视频帧的尺寸,也即是增加第一视频帧中像素点的数量,得到的第二视频帧的分辨率也就高于对应的第一视频帧。在一些实施例中,超分模型的重构层为一个反卷积层或者亚像素卷积层,终端能够通过反卷积层对第二视频帧特征进行反卷积处理,得到第二视频帧,也能够通过亚像素卷积层对第二视频帧特征进行反卷积处理,得到第二视频帧,本申请实施例对此不做限定。
下面对例1中超分模型的训练过程进行说明。
在一些实施例中,由于使用该超分模型的目的是提高视频帧的分辨率,那么在训练过程中,采用多个高分辨率图像和分别对应的多个低分辨率图像作为训练样本来对超分模型进行训练,其中,低分辨率图像是对对应的高分辨率图像进行下采样得到的,在一些实施例中,低分辨率图像也被称为受损图像。终端对该超分模型的模型参数进行初始化,将低分辨率图像输入该超分模型,通过该超分模型的至少一个卷积层对低分辨率图像进行卷积,得到该低分辨率图像的第一样本图像特征。终端通过该超分模型的全连接层和激活层对该第一样本图像特征进行全连接和非线性激活,得到该低分辨率图像的第二样本图像特征。终端通过该超分模型的重构层,对第二样本图像特征进行反卷积或者亚像素卷积,输出该低分辨率图像对应的超分图像。终端基于该超分图像与该低分辨率图像对应的高分辨率图像之间的差异信息,对该超分模型的模型参数进行调整。在一些实施例中,该超分图像与该低分辨率图像对应的高分辨率图像之间的差异信息为该超分图像与该低分辨率图像对应的高分辨率图像之间的像素值差异、图像特征差异、纹理差异中的至少一项。
在一些实施例中,终端还能够采用生成对抗(GA,Generative Adversarial)的方式来对该超分模型进行训练,也即是在训练过程中引入一个判别器,该判别器用于对该超分模型输出的图像进行打分,分数用于表示生成的高分辨率图像的逼真程度,该分数越高,也就表示判别器认为对应图像为生成图像的概率越高;该分数越高,也就表示判别器认为该图像为原生图像的概率越高。终端将低分辨率图像输入超分模型之后,将超分模型输出的超分图像输入判别器,由判别器基于该超分图像进行打分,输出该超分图像对应的分数。终端基于该分数,对该超分模型的模型参数进行调整。在下一轮迭代中,终端根据超分模型输出的超分图像与对应的高分辨率图像之间的差异 信息,对判别器的参数进行调整。通过该超分模型与判别器之间的“对抗”,提高超分模型的上采样效果。
需要说明的是,上述是以终端对该超分模型进行训练为例进行说明的,在其他可能的实施方式中,该超分模型也可以由云端训练得到,终端直接从云端获取该超分模型即可,本申请实施例对此不做限定。
例2、对于该多个第一视频帧中的任一第一视频帧,终端通过该超分模型对该第一视频帧进行多次上采样,得到第一视频帧对应的第二视频帧,例2提供的上采样方式也被称为逐步上采样超分,采用逐步上采样超分能够将困难的任务分解为简单的任务,该框架下的超分模型不仅大大降低了学习难度,而且获得了更好的性能。
比如,终端将该第一视频帧输入该超分模型,通过该超分模型的至少一个卷积层对第一视频帧进行卷积处理,得到该第一视频帧的第一视频帧特征。终端通过该超分模型的上采样层,对该第一视频帧特征进行上采样,得到第一上采样特征。终端通过该超分模型的至少一个卷积层对第一上采样特征进行卷积处理,得到该第一视频帧的第二视频帧特征。终端通过该超分模型的重构层,基于该第二视频帧特征进行重构,这里的重构也即是上采样,上采样能够增加第一视频帧的尺寸,也即是增加第一视频帧中像素点的数量,得到的第二视频帧的分辨率也就高于对应的第一视频帧。在一些实施例中,超分模型的重构层为一个反卷积层或者亚像素卷积层,终端能够通过反卷积层对第二视频帧特征进行反卷积处理,得到第二视频帧,也能够通过亚像素卷积层对第二视频帧特征进行反卷积处理,得到第二视频帧,本申请实施例对此不做限定。
需要说明的是,上述是以终端通过超分模型对第一视频帧进行两次上采样为例进行说明的,在其他可能的实施方式中,终端也能够通过该超分模型对第一视频帧进行三次或三次以上的上采样,本申请实施例对此不做限定。
例3、对于该多个第一视频帧中的任一第一视频帧,终端通过该超分模型对该第一视频帧进行上采样,得到第一视频帧对应的上采样视频帧,该上采样视频帧中像素点的数量大于第一视频帧中像素点的数量。终端通过该超分模型,对该上采样视频帧进行特征提取,得到该上采样视频帧的第一视频帧特征。终端通过该超分模型对该第一视频帧特征进行非线性映射,得到该第一视频帧的第二视频帧特征。终端通过该超分模型对该第二视频帧特征进行反卷积,得到该第一视频帧对应的第二视频帧。在一些实施例中,例3提供的上采样方式也被称为预上采样超分,采用预上采样超分能够减少学习难度,且可以得到任意比例图像。
比如,终端将该第一视频帧输入该超分模型,通过该超分模型的上采样层对该第一视频帧进行上采样,得到该第一视频帧对应的上采样视频帧。其中,该上采样层能够通过最临近插值法、双线性插值法以及均值插值法中的任一个来对第一视频帧进行上采样,得到该上采样视频帧,本申请实施例对此不做限定。终端通过该超分模型的至少一个卷积层对上采样视频帧进行卷积处理,得到该上采样视频帧的第一视频帧特征。终端通过该超分模型的全连接层和激活层,对该第一视频帧特征进行全连接和非线性激活,得到该上采样视频帧的第二视频帧特征。终端通过该超分模型的反卷积层对第二视频帧特征进行反卷积处理,得到第二视频帧。
下面对例3中超分模型的训练过程进行说明。
由于使用该超分模型的目的是提高视频帧的分辨率,那么在训练过程中,采用多个高分辨率图像和分别对应的多个低分辨率图像作为训练样本来对超分模型进行训练,其中,低分辨率图像是对对应的高分辨率图像进行下采样得到的,在一些实施例中,低分辨率图像也被称为受损图像。终端对该超分模型的模型参数进行初始化,将低分辨率图像输入该超分模型,通过该超分模型的上采样层对该低分辨率图像进行上 采样,得到该低分辨率图像对应的上采样图像。终端通过该超分模型的至少一个卷积层对低分辨率图像进行卷积,得到该低分辨率图像的第一样本图像特征。终端通过该超分模型的全连接层和激活层对该第一样本图像特征进行全连接和非线性激活,得到该低分辨率图像的第二样本图像特征。终端通过该超分模型的反卷积层,对第二样本图像特征进行反卷积,输出该低分辨率图像对应的超分图像。终端基于该超分图像与该低分辨率图像对应的高分辨率图像之间的差异信息,对该超分模型的模型参数进行调整。在一些实施例中,该超分图像与该低分辨率图像对应的高分辨率图像之间的差异信息为该超分图像与该低分辨率图像对应的高分辨率图像之间的像素值差异、图像特征差异、纹理差异中的至少一项。
在一些实施例中,终端还能够采用生成对抗的方式来对该超分模型进行训练,也即是在训练过程中引入一个判别器,该判别器用于对该超分模型输出的图像进行打分,分数用于表示生成的高分辨率图像的逼真程度,该分数越高,也就表示判别器认为对应图像为生成图像的概率越高;该分数越高,也就表示判别器认为该图像为原生图像的概率越高。终端将低分辨率图像输入超分模型之后,将超分模型输出的超分图像输入判别器,由判别器基于该超分图像进行打分,输出该超分图像对应的分数。终端基于该分数,对该超分模型的模型参数进行调整。在下一轮迭代中,终端根据超分模型输出的超分图像与对应的高分辨率图像之间的差异信息,对判别器的参数进行调整。通过该超分模型与判别器之间的“对抗”,提高超分模型的上采样效果。
需要说明的是,上述是以终端对该超分模型进行训练为例进行说明的,在其他可能的实施方式中,该超分模型也可以由云端训练得到,终端直接从云端获取该超分模型即可,本申请实施例对此不做限定。
另外,终端除了能够通过上述三个例子中描述的上采样方法之外,还能够采用迭代上下采样以及上述各个采样方法的变种来基于第一视频帧获取第二视频帧,比如,终端采用用于单一图像超分辨率的增强型深度残差网络(Enhanced Deep Residual Networks for Single Image Super-Resolution,EDSR)、广泛激活实现高效准确的图像超分辨率(Wide Activation for Efficient and Accurate Image Super-Resolution,WDSR)等,本申请实施例对此不做限定。
404、在多个第二视频帧的帧率符合帧率条件的情况下,终端在多个第二视频帧中进行插帧,得到插帧后的多个第二视频帧。
其中,帧率是指终端每秒播放第二视频帧的数量,帧率越高,播放第二视频帧形成的视频也就越流畅;帧率越低,播放第二视频帧形成的视频也就越卡顿。正如上述步骤401所描述的,当终端当前网络延迟较高,也即是终端当前的网络状况不佳时,云游戏服务器除了能够通过降低第一视频帧的分辨率的方式来减少传输第一视频帧所占用的带宽之外,还能够减少传输第一视频帧的数量,也即是将多个第一视频帧的帧率来减少传输时所占用的带宽,步骤404也即是在帧率较低时进行插帧,以提高帧率的方法。
在一些实施例中,在该多个第二视频帧的帧率小于或等于帧率阈值的情况下,终端在该多个第二视频帧中每两个第二视频帧之间插入参考视频帧,得到该插帧后的该多个第二视频帧。其中,这里描述的每两个第二视频帧是指,在时序上相邻的两个第二视频帧。
在这种实施方式下,当多个第二视频帧的帧率较低时,终端能够在每两个第二视频帧之间进行插帧,以提高多个第二视频帧的帧率,帧率的提高能够消除“顿挫感”,提高多个第二视频帧的播放效果。
在上述实施方式的基础上,下面对终端确定参考视频帧的方法进行说明。
在一些实施例中,终端将每两个第二视频帧中任一第二视频帧确定为该参考视频帧。
在这种实施方式下,终端能够直接将每两个第二视频帧中任一第二视频帧确定为该参考视频帧,无需终端进行额外的运算,确定参考视频帧的效率较高。
举例来说,对于时序上相邻的第二视频帧A和第二视频帧B来说,终端能够直接将第二视频帧A或者第二视频帧B确定为参考视频帧,后续直接将参考视频帧添加之间第二视频帧A和第二视频帧B。比如,终端将第二视频帧A确定为该参考视频帧,那么将该参考视频帧添加到第二视频帧A和第二视频帧B之间后,得到{第二视频帧A|第二视频帧A|第二视频帧B},原本的两个视频帧也就扩展为了三个视频帧,终端也就能够在相同时间内显示更多的视频帧,缩短视频帧之间的显示时间,从而提高播放的流畅程度。
在一些实施例中,终端将该每两个第二视频帧的平均视频帧确定为该参考视频帧,该平均视频帧中像素点的像素值为该每两个第二视频帧中对应像素点的像素值的平均值。
在这种实施方式下,终端能够直接通过求取平均值的方式来确定参考像素点,运算量较小,确定参考视频帧的效率较高。
举例来说,对于时序上相邻的第二视频帧A和第二视频帧B来说,终端基于第二视频帧A和第二视频帧B生成一个平均视频帧,该平均视频帧也即是该参考视频帧。终端在生成该平均视频帧时,获取第二视频帧A对应的像素值矩阵M和第二视频帧B对应的像素值矩阵N,获取像素值矩阵M和像素值矩阵N的平均像素值矩阵O。终端生成一个空白视频帧,该空白视频帧中像素点的数量和分布与第二视频帧A和第二视频帧B均相同,终端将平均像素值矩阵O中的数值确定为空白视频帧中对应像素点的像素值,得到参考视频帧。后续终端将该参考视频帧添加至第二视频帧A和第二视频帧B即可,也即是将原本的{第二视频帧A|第二视频帧B}变为{第二视频帧A|参考视频帧|第二视频帧B},原本的两个视频帧也就扩展为了三个视频帧,终端也就能够在相同时间内显示更多的视频帧,缩短视频帧之间的显示时间,从而提高播放的流畅程度。
在一些实施例中,终端将该第二视频帧序列输入插帧模型,由插帧模型基于第二视频帧序列中相邻的两个第二视频帧,生成参考视频帧,即由该插帧模型基于该每两个第二视频帧进行处理,得到该参考视频帧。
为了对上述实施方式进行更加清楚的说明,下面将通过两个例子对终端通过插帧模型来获取参考视频帧的方法进行说明。
例1、终端通过该插帧模型,获取所述第二视频帧序列中相邻的两个第二视频帧的后向光流及前向光流,例如,获取第三视频帧到第四视频帧的后向光流,该第三视频帧为该每两个第二视频帧中前一个第二视频帧,该第四视频帧为该每两个第二视频帧中后一个第二视频帧。终端通过该插帧模型,获取该第四视频帧到该第三视频帧的前向光流。终端通过该插帧模型,基于该后向光流和该前向光流,生成该参考视频帧。
在这种实施方式下,终端能够基于光流法来生成参考视频帧,参考视频帧的质量较好,从而提高了终端对视频帧的播放效果。
举例来说,终端过该插帧模型,获取第三视频帧到第四视频帧的第一后向光流,基于该第一后向光流,确定第三视频帧的第一时刻到中间时刻的第二后向光流,其中,中间时刻是第三视频帧和第四视频帧之间的时刻。终端通过该插帧模型,获取第三视频帧的特征图和边缘图像。终端通过该插帧模型,基于该第二后向光流,对第三视频帧以及第三视频帧的特征图和边缘图像进行前向映射,得到第一前向映射视频帧以及 第一前向映射参考信息,第一前向映射参考信息包括第一前向映射视频帧的特征图和边缘图像。终端过该插帧模型,获取第四视频帧到第三视频帧的第一前向光流,基于该第一前向光流,确定第四视频帧的第二时刻到中间时刻的第二前向光流,其中,中间时刻是第四视频帧和第四视频帧之间的时刻。终端通过该插帧模型,获取第四视频帧的特征图和边缘图像。终端通过该插帧模型,基于该第二前向光流,对第四视频帧以及第四视频帧的特征图和边缘图像进行前向映射,得到第二前向映射视频帧以及第二前向映射参考信息,第二前向映射参考信息包括第二前向映射视频帧的特征图和边缘图像。终端通过该插帧模型,将第一前向映射视频帧、第一前向映射参考信息、第二前向映射视频帧以及第二前向映射参考信息组合为前向映射结果。终端通过该插帧模型,基于第一前向光流,确定从中间时刻到第二时刻的第三后向光流,基于第三后向光流,对第四视频帧以及第四视频帧的特征图和边缘图像进行后向映射,得到第一后向映射视频帧以及第一后向映射参考信息,第一后向映射参考信息包括第一后向映射视频帧的特征图和边缘图像。终端通过该插帧模型,基于第一后向光流,确定从中间时刻到第一时刻的第三前向光流,基于第三前向光流,对第三视频帧以及第三视频帧的特征图和边缘图像进行后向映射,得到第二后向映射视频帧以及第二后向映射参考信息,第二后向映射参考信息包括第二后向映射视频帧的特征图和边缘图像。终端通过该插帧模型,将第一后向映射视频帧、第一后向映射参考信息、第二后向映射视频帧以及第二后向映射参考信息组合为后向映射结果。终端通过该插帧模型,将前向映射结果和后向映射结果进行融合,得到该参考视频帧。
其中,终端在通过该插帧模型将前向映射结果和后向映射结果进行融合,得到该参考视频帧时,终端通过该插帧模型,对该前向映射结果进行编码,得到前向中间特征来对该后向映射结果进行编码,得到后向中间特征。终端通过该插帧模型,前向中间特征和后向中间特征进行融合,得到融合中间特征。终端通过该插帧模型,对融合中间特征进行解码,得到该参考视频帧。
例2、终端通过该插帧模型,获取第三视频帧中多个图像块在第四视频帧中的运动矢量,该第三视频帧为相邻的两个第二视频帧中前一个第二视频帧,该第四视频帧为相邻的两个第二视频帧中后一个第二视频帧。终端基于该运动矢量,生成该参考视频帧。
在这种实施方式下,终端能够获取视频帧中图像块的运动矢量,基于运动矢量来生成参考视频帧,生成的参考视频帧的质量较高,终端的播放视频帧的效果较好。
举例来说,终端通过该插帧模型,获取该第三视频帧和第四视频帧的编码信息,该编码信息用于指示对第三视频帧和第四视频帧进行分块的方式以及第三视频帧中各个图像块由第三视频帧到第四视频帧的运动矢量。终端通过该插帧模型,基于该第三视频帧的编码信息,将第三视频帧和第四视频帧划分为多个图像块,确定各个图像块从第三视频帧到第四视频帧的运动矢量。终端通过该插帧模型,对各个图像块对应的运动矢量进行处理,得到各个图像块对应的目标运动矢量。终端通过各个图像块以及各个图像块的目标运动矢量,生成该参考视频帧。其中,终端通过该插帧模型对各个图像块对应的运动矢量进行处理包括将各个图像块对应的运动矢量除以目标数值,也即是保证图像块运动方向不变的前提下,缩短图像块的运动距离,目标数值由技术人员根据实际情况进行设置,本申请实施例对此不做限定。
需要说明的是,该插帧模型除了为上述例1和例2中描述插帧模型之外,还可以为其他类型的插帧模型,比如实时中间流估计算法(Real-Time Intermediate Flow Estimation for Video Frame Interpolation,RIFE)、基于残差细化的视频帧插值(Video Frame Interpolation via Residue Refinement,RRIN)以及基于增强可变形可分离卷积 的多视频帧插值(Multiple Video Frame Interpolation via Enhanced Deformable Separable Convolution,EDSC)等,本申请实施例对此不做限定。
另外,上述是以终端先对第一视频帧进行超分(步骤403),随后对超分后的多个第二视频帧进行插帧(步骤404)这样的顺序进行说明的,在其他可能的实施方式中,终端也能够先对多个第一视频帧进行插帧,随后在对插帧后的多个第一视频帧进行超分,本申请实施例对此不做限定。
405、终端播放插帧后的多个第二视频帧。
通过上述步骤401-405,相较于获取到的多个第一视频帧来说,终端播放的多个第二视频帧不仅分辨率更高,而且数量更多,从而播放视频帧的效果更好。
下面将结合图5和上述步骤401-405,对本申请实施例提供的视频帧播放方法进行说明,参见图5,本申请实施例提供的视频帧播放方法包括:
步骤501:终端启动云游戏客户端。
通过云游戏客户端获取终端当前的网络延迟信息。
步骤502:终端根据网络延迟信息,确定对应的视频流信息。
这里,视频流信息包括分辨率、帧率和码率。
步骤503:判断终端的网络状态是否良好,如果是,执行步骤505,否则,执行步骤504。
云游戏服务器基于该视频流信息确定渲染参数,对目标虚拟场景进行渲染,得到第一视频帧。依据网络延迟信息判断终端的网络状态是否良好,当网络延迟信息指示终端当前网络状态不佳(即网络延迟大于或等于网络延迟阈值)时,触发步骤504,当网络延迟信息指示终端当前网络状态良好(即网络延迟小于网络延迟阈值)时,触发步骤505。
步骤504:云游戏服务器通知终端降低分辨率和帧率。
步骤505:超分和插帧。
终端基于获取到的多个第一视频帧进行判断,但需要进行超分和插帧时,对获取到的第一视频帧进行超分和插帧,得到插帧后的多个第二视频帧,终端播放该多个第二视频帧,该多个第二视频帧也即是超分插帧后的云游戏画面。比如,当前网络或终端配置不好,从云游戏服务器接收到的视频流是480P/15FPS、720P/30FPS或1080P/60FPS等,用户可以根据自己体验习惯和终端算力情况选择是否超分插帧到720P/30FPS、1080P/60FPS或4K/90FPS等。
上述所有可选技术方案,可以采用任意结合形成本申请的可选实施例,在此不再一一赘述。
图6和图7提供了采用本申请实施例提供的视频帧播放方法后首帧延时变化情况和卡顿率变化情况,通过图6可以看出,采用本申请实施例提供的视频帧播放方法后,首帧延时有着100ms以上的下降,卡顿率降低了50%以上。
通过本申请实施例提供的技术方案,获取到云游戏服务器发送的多个第一视频帧之后,基于多个第一视频帧的分辨率进行判断,当该分辨率符合分辨率调整条件时,对多个第一视频帧进行分辨率调整,分辨率调整的目的是提高第一视频帧的分辨率,得到多个第二视频帧,多个第二视频帧也就具有更高的分辨率,从而在保证云游戏流畅度的前提下,提高了云游戏的显示效果。
通过本申请实施例提供的技术方案,对于网络不稳定的终端,云游戏服务器可以切到低分辨率低FPS采集编码,低分辨率低FPS可以大幅降低云游戏服务器采集编码的码率,减少网络不稳定终端的网络带宽压力,让传输更稳定流畅,比如对于网络不稳定终端,云游戏服务器采集编码720P/30FPS相对1080P/60FPS码率可以节省 大于400%的网络带宽,然后终端通过插帧和超分技术超分插帧到1080P/60FPS,云游戏画质体验相对采集编码1080P/60FPS主观体验MOS(Mean Opinion Score,主观评分)相差在5%以内,但码率节省了400%以上,对于网络不稳定但是终端算力不错的情况具有极大的提升。
图8是本申请实施例提供的一种视频帧播放装置的结构示意图,参见图8,装置包括:视频帧获取模块801、分辨率调整模块802以及播放模块803。
视频帧获取模块801,配置为获取多个第一视频帧,该多个第一视频帧是云游戏服务器,对目标虚拟场景进行渲染得到的视频帧。
分辨率调整模块802,配置为在多个第一视频帧的分辨率符合分辨率调整条件时,分别对各第一视频帧进行分辨率调整,得到相应的第二视频帧,第二视频帧的分辨率高于对应的第一视频帧的分辨率。
播放模块803,配置为播放进行分辨率调整所得到的多个第二视频帧。
在一些实施例中,该分辨率调整模块802,还配置为针对各所述第一视频帧,在各个第一视频帧中每两个像素点之间插入参考像素点,得到各个第一视频帧对应的第二视频帧,该参考像素点是基于该每两个像素点生成的。
在一些实施例中,该分辨率调整模块802,配置为执行下述任一项:
采用最临近插值法在各个第一视频帧中每两个像素点之间插入该参考像素点。
采用双线性插值法在各个第一视频帧中每两个像素点之间插入该参考像素点。
采用均值插值法在各个第一视频帧中每两个像素点之间插入该参考像素点。
在一些实施例中,该分辨率调整模块802,还配置为将该各个第一视频帧分别输入超分模型,由该超分模型对各第一视频帧进行上采样,得到该相应的第二视频帧。
在一些实施例中,该分辨率调整模块802,配置为对于该多个第一视频帧中的任一第一视频帧,通过该超分模型对第一视频帧进行特征提取,得到第一视频帧的第一视频帧特征,对第一视频帧特征进行非线性映射,得到第一视频帧的第二视频帧特征,基于第二视频帧特征进行视频帧重构,得到该第一视频帧对应的第二视频帧。
在一些实施例中,该装置还包括:
插帧模块,配置为在多个第二视频帧间进行插帧,得到插帧后的多个第二视频帧。
该播放模块803,还配置为播放插帧后的多个第二视频帧。
在一些实施例中,该插帧模块配置为在多个第二视频帧的帧率小于或等于帧率阈值的情况下,在每两个第二视频帧之间插入参考视频帧,得到该插帧后的多个第二视频帧。
在一些实施例中,该装置还包括参考帧确定模块,配置为执行下述任一项:
将该每两个第二视频帧中任一第二视频帧确定为该参考视频帧。
将该每两个第二视频帧的平均视频帧确定为该参考视频帧,该平均视频帧中像素点的像素值为该每两个第二视频帧中对应像素点的像素值的平均值。
将由多个所述第二视频帧构成的第二视频帧序列输入插帧模型,由该插帧模型基于第二视频帧序列中相邻的两个第二视频帧,生成参考视频帧。
在一些实施例中,该参考帧确定模块配置为通过该插帧模型,所述第二视频帧序列中相邻的两个第二视频帧的后向光流及前向光流,基于该后向光流和该前向光流,生成该参考视频帧;例如,获取第三视频帧到第四视频帧的后向光流,该第三视频帧为该每两个第二视频帧中前一个第二视频帧,该第四视频帧为该每两个第二视频帧中后一个第二视频帧。获取该第四视频帧到该第三视频帧的前向光流。基于该后向光流和该前向光流,生成该参考视频帧。
在一些实施例中,该参考帧确定模块配置为通过该插帧模型,获取第三视频帧中 多个图像块在第四视频帧中的运动矢量,该第三视频帧为相邻的两个第二视频帧中前一个第二视频帧,该第四视频帧为相邻的两个第二视频帧中后一个第二视频帧。基于该运动矢量,生成该参考视频帧。
在一些实施例中,该装置还包括:
发送模块,配置为向该云游戏服务器发送网络延迟信息,以使该云游戏服务器基于该网络延迟信息生成该多个第一视频帧。
需要说明的是:上述实施例提供的视频帧播放装置在播放视频帧时,仅以上述各功能模块的划分进行举例说明,实际应用中,可以根据需要而将上述功能分配由不同的功能模块完成,即将计算机设备的内部结构划分成不同的功能模块,以完成以上描述的全部或者部分功能。另外,上述实施例提供的视频帧播放装置与视频帧播放方法实施例属于同一构思,其实现过程详见方法实施例,这里不再赘述。
通过本申请实施例提供的技术方案,获取到云游戏服务器发送的多个第一视频帧之后,基于多个第一视频帧的分辨率进行判断,当该分辨率符合分辨率调整条件时,对各第一视频帧进行分辨率调整,分辨率调整的目的是提高第一视频帧的分辨率,得到相应的第二视频帧,进而播放调高第一视频帧的分辨率所得到的多个第二视频帧,相较于直接播放云游戏服务器发送的多个第一视频帧,提高了云游戏画面的清晰度,从而在保证云游戏流畅度的前提下,提高了云游戏的显示效果。
本申请实施例提供了一种计算机设备,配置为执行本申请实施例提供的视频帧播放方法,该计算机设备可以实现为终端或者服务器,下面先对终端的结构进行介绍:
图9是本申请实施例提供的一种终端的结构示意图。该终端900可以是:智能手机、平板电脑、笔记本电脑或台式电脑。终端900还可能被称为用户设备、便携式终端、膝上型终端、台式终端等其他名称。
通常,终端900包括有:一个或多个处理器901和一个或多个存储器902。
处理器901可以包括一个或多个处理核心,比如4核心处理器、8核心处理器等。处理器901可以采用数字信号处理(Digital Signal Processing,DSP)、现场可编程门阵列(Field-Programmable Gate Array,FPGA)、可编程逻辑阵列(Programmable Logic Array,PLA)中的至少一种硬件形式来实现。处理器901也可以包括主处理器和协处理器,主处理器是用于对在唤醒状态下的数据进行处理的处理器,也称中央处理器(Central Processing Unit,CPU);协处理器是用于对在待机状态下的数据进行处理的低功耗处理器。在一些实施例中,处理器901可以在集成有分辨率调整器(Graphics Processing Unit,GPU),GPU用于负责显示屏所需要显示的内容的渲染和绘制。一些实施例中,处理器901还可以包括人工智能(Artificial Intelligence,AI)处理器,该AI处理器用于处理有关机器学习的计算操作。
存储器902可以包括一个或多个计算机可读存储介质,该计算机可读存储介质可以是非暂态的。存储器902还可包括高速随机存取存储器,以及非易失性存储器,比如一个或多个磁盘存储设备、闪存存储设备。在一些实施例中,存储器902中的非暂态的计算机可读存储介质用于存储至少一个计算机程序,该至少一个计算机程序用于被处理器901所执行以实现本申请中方法实施例提供的视频帧播放方法。
在一些实施例中,终端900还包括有:外围设备接口903和至少一个外围设备。处理器901、存储器902和外围设备接口903之间可以通过总线或信号线相连。各个外围设备可以通过总线、信号线或电路板与外围设备接口903相连。在实际应用中,外围设备包括:射频电路904、显示屏905、摄像头组件906、音频电路907、定位组件908和电源909中的至少一种。
外围设备接口903可被用于将输入/输出(Input/Output,I/O)相关的至少一个外 围设备连接到处理器901和存储器902。在一些实施例中,处理器901、存储器902和外围设备接口903被集成在同一芯片或电路板上;在一些其他实施例中,处理器901、存储器902和外围设备接口903中的任意一个或两个可以在单独的芯片或电路板上实现,本实施例对此不加以限定。
射频电路904用于接收和发射射频(Radio Frequency,RF)信号,也称电磁信号。射频电路904通过电磁信号与通信网络以及其他通信设备进行通信。射频电路904将电信号转换为电磁信号进行发送,或者,将接收到的电磁信号转换为电信号。在一些实施例中,射频电路904包括:天线系统、RF收发器、一个或多个放大器、调谐器、振荡器、数字信号处理器、编解码芯片组、用户身份模块卡等等。
显示屏905用于显示用户界面(User Interface,UI)。该UI可以包括图形、文本、图标、视频及其它们的任意组合。当显示屏905是触摸显示屏时,显示屏905还具有采集在显示屏905的表面或表面上方的触摸信号的能力。该触摸信号可以作为控制信号输入至处理器901进行处理。此时,显示屏905还可以用于提供虚拟按钮和/或虚拟键盘,也称软按钮和/或软键盘。
摄像头组件906用于采集图像或视频。在一些实施例中,摄像头组件906包括前置摄像头和后置摄像头。通常,前置摄像头设置在终端的前面板,后置摄像头设置在终端的背面。
音频电路907可以包括麦克风和扬声器。麦克风用于采集用户及环境的声波,并将声波转换为电信号输入至处理器901进行处理,或者输入至射频电路904以实现语音通信。
定位组件908用于定位终端900的当前地理位置,以实现导航或基于位置的服务(Location Based Service,LBS)。
电源909用于为终端900中的各个组件进行供电。电源909可以是交流电、直流电、一次性电池或可充电电池。
本领域技术人员可以理解,图9中示出的结构并不构成对终端900的限定,可以包括比图示更多或更少的组件,或者组合某些组件,或者采用不同的组件布置。
本申请实施例还提供了一种计算机可读存储介质,例如包括计算机程序的存储器,上述计算机程序可由处理器执行以完成上述实施例中的视频帧播放方法。例如,该计算机可读存储介质可以是只读存储器(Read-Only Memory,ROM)、随机存取存储器(Random Access Memory,RAM)、只读光盘(Compact Disc Read-Only Memory,CD-ROM)、磁带、软盘和光数据存储设备等。
本申请实施例还提供了一种计算机程序产品或计算机程序,该计算机程序产品或计算机程序包括程序代码,该程序代码存储在计算机可读存储介质中,计算机设备的处理器从计算机可读存储介质读取该程序代码,处理器执行该程序代码,使得该计算机设备执行上述视频帧播放方法。
在一些实施例中,本申请实施例所涉及的计算机程序可被部署在一个计算机设备上执行,或者在位于一个地点的多个计算机设备上执行,又或者,在分布在多个地点且通过通信网络互连的多个计算机设备上执行,分布在多个地点且通过通信网络互连的多个计算机设备可以组成区块链系统。
本领域普通技术人员可以理解实现上述实施例的全部或部分步骤可以通过硬件来完成,也可以通过程序来指令相关的硬件完成,该程序可以存储于一种计算机可读存储介质中,上述提到的存储介质可以是只读存储器,磁盘或光盘等。
上述仅为本申请的可选实施例,并不用以限制本申请,凡在本申请的精神和原则之内,所做的任何修改、等同替换、改进等,均应包含在本申请的保护范围之内。
Claims (15)
- 一种视频帧播放方法,所述方法由计算机设备执行,包括:获取多个第一视频帧,所述多个第一视频帧是云游戏服务器对目标虚拟场景进行渲染得到的视频帧;当所述多个第一视频帧的分辨率符合分辨率调整条件时,分别对各所述第一视频帧进行分辨率调整,得到相应的第二视频帧,所述第二视频帧的分辨率高于对应的第一视频帧的分辨率;播放进行分辨率调整所得到的多个所述第二视频帧。
- 根据权利要求1所述的方法,其中,所述分别对各所述第一视频帧进行分辨率调整,得到相应的第二视频帧,包括:针对各所述第一视频帧,在所述第一视频帧中每两个像素点之间插入参考像素点,得到所述第一视频帧对应的第二视频帧;其中,所述参考像素点是基于所述每两个像素点生成的。
- 根据权利要求2所述的方法,其中,所述在所述第一视频帧中每两个像素点之间插入参考像素点,包括以下至少之一:采用最临近插值法在所述第一视频帧中每两个像素点之间插入所述参考像素点;采用双线性插值法在所述第一视频帧中每两个像素点之间插入所述参考像素点;采用均值插值法在所述第一视频帧中每两个像素点之间插入所述参考像素点。
- 根据权利要求1所述的方法,其中,所述分别对各所述第一视频帧进行分辨率调整,得到相应的第二视频帧,包括:将各所述第一视频帧分别输入超分模型,由所述超分模型对所述第一视频帧进行上采样,得到相应的第二视频帧。
- 根据权利要求4所述的方法,其中,所述由所述超分模型对所述第一视频帧进行上采样,得到相应的第二视频帧,包括:通过所述超分模型对所述第一视频帧进行特征提取,得到所述第一视频帧的第一视频帧特征;对所述第一视频帧特征进行非线性映射,得到所述第一视频帧的第二视频帧特征;基于所述第二视频帧特征,进行视频帧重构,得到所述第一视频帧对应的第二视频帧。
- 根据权利要求1所述的方法,其中,所述分别对各所述第一视频帧进行分辨率调整,得到相应的第二视频帧之后,所述方法还包括:在多个所述第二视频帧间进行插帧,得到插帧后的多个第二视频帧;所述播放进行分辨率调整所得到的多个所述第二视频帧,包括:播放所述插帧后的所述多个第二视频帧。
- 根据权利要求6所述的方法,其中,所述在多个所述第二视频帧间进行插帧,得到插帧后的多个第二视频帧,包括:在多个所述第二视频帧的帧率小于或等于帧率阈值的情况下,在每两个所述第二视频帧之间插入参考视频帧,得到插帧后的多个所述第二视频帧。
- 根据权利要求7所述的方法,其中,所述方法还包括:将所述每两个第二视频帧中任一第二视频帧确定为所述参考视频帧;或者,将所述每两个第二视频帧的平均视频帧确定为所述参考视频帧,所述平均视频帧中像素点的像素值为,所述每两个第二视频帧中对应像素点的像素值的平均值;或者,将由多个所述第二视频帧构成的所述第二视频帧序列,输入插帧模型,由所述插帧模型基于所述第二视频帧序列中相邻的两个第二视频帧,生成所述参考视频帧。
- 根据权利要求8所述的方法,其中,所述由所述插帧模型基于所述第二视频帧序列中相邻的两个第二视频帧,生成所述参考视频帧,包括:通过所述插帧模型,获取所述第二视频帧序列中相邻的两个第二视频帧的后向光流及前向光流;基于所述后向光流和所述前向光流,生成所述参考视频帧。
- 根据权利要求8所述的方法,其中,所述由所述插帧模型基于所述第二视频帧序列中相邻的两个第二视频帧,生成所述参考视频帧,包括:通过所述插帧模型,获取第三视频帧中多个图像块在第四视频帧中的运动矢量;所述第三视频帧为相邻的两个第二视频帧中前一个第二视频帧,所述第四视频帧为相邻的两个第二视频帧中后一个第二视频帧;基于所述运动矢量,生成所述参考视频帧。
- 根据权利要求1-10任一项所述的方法,其中,所述获取多个第一视频帧之前,所述方法还包括:向所述云游戏服务器发送网络延迟信息;所述网络延迟信息,用于所述云游戏服务器基于所述网络延迟信息生成所述多个第一视频帧。
- 一种视频帧播放装置,所述装置包括:视频帧获取模块,配置为获取多个第一视频帧,所述多个第一视频帧是云游戏服务器对所述目标虚拟场景进行渲染得到的视频帧;分辨率调整模块,配置为当所述多个第一视频帧的分辨率符合分辨率调整条件时,分别对各所述第一视频帧进行分辨率调整,得到相应的第二视频帧,所述第二视频帧的分辨率高于对应的第一视频帧的分辨率;播放模块,配置为播放进行分辨率调整所得到的多个所述第二视频帧。
- 一种计算机设备,所述计算机设备包括一个或多个处理器和一个或多个存储器,所述一个或多个存储器中存储有至少一条计算机程序,所述计算机程序由所述一个或多个处理器加载并执行以实现如权利要求1至权利要求11任一项所述的视频帧播放方法。
- 一种计算机可读存储介质,所述计算机可读存储介质中存储有至少一条计算机程序,所述计算机程序由处理器加载并执行以实现如权利要求1至权利要求11任一项所述的视频帧播放方法。
- 一种计算机程序产品,包括计算机程序,所述计算机程序被处理器执行时实现权利要求1至权利要求11任一项所述的视频帧播放方法。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US18/139,273 US20230260084A1 (en) | 2021-09-26 | 2023-04-25 | Video frame playing method and apparatus, device, storage medium, and program product |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202111130391.5 | 2021-09-26 | ||
| CN202111130391.5A CN115883853B (zh) | 2021-09-26 | 2021-09-26 | 视频帧播放方法、装置、设备以及存储介质 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US18/139,273 Continuation US20230260084A1 (en) | 2021-09-26 | 2023-04-25 | Video frame playing method and apparatus, device, storage medium, and program product |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2023045649A1 true WO2023045649A1 (zh) | 2023-03-30 |
Family
ID=85720024
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2022/113526 Ceased WO2023045649A1 (zh) | 2021-09-26 | 2022-08-19 | 视频帧播放方法、装置、设备、存储介质及程序产品 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20230260084A1 (zh) |
| CN (1) | CN115883853B (zh) |
| WO (1) | WO2023045649A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP4468247A4 (en) * | 2023-04-12 | 2024-11-27 | Jiaqi Guo | Rendering acceleration method and system for three-dimensional animation |
Families Citing this family (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116489457B (zh) * | 2023-05-26 | 2025-11-28 | 西安诺瓦星云科技股份有限公司 | 视频的显示控制方法、装置、设备、系统和存储介质 |
| CN116440501B (zh) * | 2023-06-16 | 2023-08-29 | 瀚博半导体(上海)有限公司 | 自适应云游戏视频画面渲染方法和系统 |
| CN116886744B (zh) * | 2023-09-08 | 2023-12-29 | 深圳云天畅想信息科技有限公司 | 一种串流分辨率的动态调整方法、装置、设备及存储介质 |
| CN116931864B (zh) * | 2023-09-18 | 2024-02-09 | 广东保伦电子股份有限公司 | 屏幕共享方法及智能交互平板 |
| CN117291810B (zh) * | 2023-11-27 | 2024-03-12 | 腾讯科技(深圳)有限公司 | 视频帧的处理方法、装置、设备及存储介质 |
| US20250239005A1 (en) * | 2024-01-20 | 2025-07-24 | Lemon Inc. | Methods and systems for generating a multi-dimensional image using cross-view correspondences |
| CN119011964B (zh) * | 2024-07-26 | 2026-03-17 | 深圳Tcl数字技术有限公司 | 云游戏画质调整方法、装置、云端服务器及终端 |
| CN119516566B (zh) * | 2024-11-18 | 2025-10-21 | 上海交通大学 | 文本生成三维内容的质量评价方法、系统、介质及终端 |
Citations (10)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106331750A (zh) * | 2016-10-08 | 2017-01-11 | 中山大学 | 一种基于感兴趣区域的云游戏平台自适应带宽优化方法 |
| CN108694376A (zh) * | 2017-04-01 | 2018-10-23 | 英特尔公司 | 包括静态场景确定、阻塞检测、帧率变换和调整压缩率的视频运动处理 |
| CN110149371A (zh) * | 2019-04-24 | 2019-08-20 | 深圳市九和树人科技有限责任公司 | 设备连接方法、装置及终端设备 |
| CN110881136A (zh) * | 2019-11-14 | 2020-03-13 | 腾讯科技(深圳)有限公司 | 视频帧率控制方法、装置、计算机设备及存储介质 |
| US20200162789A1 (en) * | 2018-11-19 | 2020-05-21 | Zhan Ma | Method And Apparatus Of Collaborative Video Processing Through Learned Resolution Scaling |
| CN111681167A (zh) * | 2020-06-03 | 2020-09-18 | 腾讯科技(深圳)有限公司 | 画质调整方法和装置、存储介质及电子设备 |
| CN111970513A (zh) * | 2020-08-14 | 2020-11-20 | 成都数字天空科技有限公司 | 一种图像处理方法、装置、电子设备及存储介质 |
| WO2021065629A1 (ja) * | 2019-09-30 | 2021-04-08 | 株式会社ソニー・インタラクティブエンタテインメント | 画像表示システム、動画配信サーバ、画像処理装置、および動画配信方法 |
| CN113842635A (zh) * | 2021-09-07 | 2021-12-28 | 山东师范大学 | 提升云游戏流畅度的方法及系统 |
| CN114240749A (zh) * | 2021-12-13 | 2022-03-25 | 网易(杭州)网络有限公司 | 图像处理方法、装置、计算机设备及存储介质 |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20140321561A1 (en) * | 2013-04-26 | 2014-10-30 | DDD IP Ventures, Ltd. | System and method for depth based adaptive streaming of video information |
| CN104216783B (zh) * | 2014-08-20 | 2017-07-11 | 上海交通大学 | 云游戏中虚拟gpu资源自主管理与控制方法 |
| CN113055742B (zh) * | 2021-03-05 | 2023-06-09 | Oppo广东移动通信有限公司 | 视频显示方法、装置、终端及存储介质 |
| CN113423018B (zh) * | 2021-08-24 | 2021-11-02 | 腾讯科技(深圳)有限公司 | 一种游戏数据处理方法、装置及存储介质 |
-
2021
- 2021-09-26 CN CN202111130391.5A patent/CN115883853B/zh active Active
-
2022
- 2022-08-19 WO PCT/CN2022/113526 patent/WO2023045649A1/zh not_active Ceased
-
2023
- 2023-04-25 US US18/139,273 patent/US20230260084A1/en active Pending
Patent Citations (10)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106331750A (zh) * | 2016-10-08 | 2017-01-11 | 中山大学 | 一种基于感兴趣区域的云游戏平台自适应带宽优化方法 |
| CN108694376A (zh) * | 2017-04-01 | 2018-10-23 | 英特尔公司 | 包括静态场景确定、阻塞检测、帧率变换和调整压缩率的视频运动处理 |
| US20200162789A1 (en) * | 2018-11-19 | 2020-05-21 | Zhan Ma | Method And Apparatus Of Collaborative Video Processing Through Learned Resolution Scaling |
| CN110149371A (zh) * | 2019-04-24 | 2019-08-20 | 深圳市九和树人科技有限责任公司 | 设备连接方法、装置及终端设备 |
| WO2021065629A1 (ja) * | 2019-09-30 | 2021-04-08 | 株式会社ソニー・インタラクティブエンタテインメント | 画像表示システム、動画配信サーバ、画像処理装置、および動画配信方法 |
| CN110881136A (zh) * | 2019-11-14 | 2020-03-13 | 腾讯科技(深圳)有限公司 | 视频帧率控制方法、装置、计算机设备及存储介质 |
| CN111681167A (zh) * | 2020-06-03 | 2020-09-18 | 腾讯科技(深圳)有限公司 | 画质调整方法和装置、存储介质及电子设备 |
| CN111970513A (zh) * | 2020-08-14 | 2020-11-20 | 成都数字天空科技有限公司 | 一种图像处理方法、装置、电子设备及存储介质 |
| CN113842635A (zh) * | 2021-09-07 | 2021-12-28 | 山东师范大学 | 提升云游戏流畅度的方法及系统 |
| CN114240749A (zh) * | 2021-12-13 | 2022-03-25 | 网易(杭州)网络有限公司 | 图像处理方法、装置、计算机设备及存储介质 |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP4468247A4 (en) * | 2023-04-12 | 2024-11-27 | Jiaqi Guo | Rendering acceleration method and system for three-dimensional animation |
Also Published As
| Publication number | Publication date |
|---|---|
| US20230260084A1 (en) | 2023-08-17 |
| CN115883853B (zh) | 2024-04-05 |
| CN115883853A (zh) | 2023-03-31 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2023045649A1 (zh) | 视频帧播放方法、装置、设备、存储介质及程序产品 | |
| US12214277B2 (en) | Method and device for generating video frames | |
| US11458393B2 (en) | Apparatus and method of generating a representation of a virtual environment | |
| US11325037B2 (en) | Apparatus and method of mapping a virtual environment | |
| JP6310073B2 (ja) | 描画システム、制御方法、及び記憶媒体 | |
| US8403757B2 (en) | Method and apparatus for providing gaming services and for handling video content | |
| JP5987060B2 (ja) | ゲームシステム、ゲーム装置、制御方法、プログラム及び記録媒体 | |
| JP6576245B2 (ja) | 情報処理装置、制御方法及びプログラム | |
| US20120270652A1 (en) | System for servicing game streaming according to game client device and method | |
| JP2016528563A (ja) | 画像処理装置、画像処理システム、画像処理方法、及び記憶媒体 | |
| JP6379107B2 (ja) | 情報処理装置並びにその制御方法、及びプログラム | |
| CN113242440A (zh) | 直播方法、客户端、系统、计算机设备以及存储介质 | |
| CN111249723B (zh) | 游戏中的显示控制的方法、装置、电子设备及存储介质 | |
| CN112604279A (zh) | 一种特效显示方法及装置 | |
| CN110860084B (zh) | 一种虚拟画面处理方法及装置 | |
| WO2025071681A1 (en) | Systems and methods for artificial intelligence (ai)-driven 2d-to-3d video stream conversion | |
| CN111672132B (zh) | 游戏的控制方法、控制装置、服务器和存储介质 | |
| US20250218114A1 (en) | Interpolated translatable audio for virtual experience | |
| US12499609B2 (en) | Video generating device and method | |
| EP4395329A1 (en) | Method for allowing streaming of video content between server and electronic device, and server and electronic device for streaming video content | |
| HK40084139A (zh) | 视频帧播放方法、装置、设备以及存储介质 | |
| HK40084139B (zh) | 视频帧播放方法、装置、设备以及存储介质 | |
| CN120431223A (zh) | 渲染方法、装置及端云协同系统 | |
| CN113923398A (zh) | 一种视频会议实现方法及装置 | |
| EP4687344A1 (en) | Content streaming system and method |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 22871705 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 16.08.2024) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 22871705 Country of ref document: EP Kind code of ref document: A1 |