WO2022111140A1 - Video encoding through non-saliency compression for live streaming of high definition videos in low-bandwidth transmission - Google Patents
Video encoding through non-saliency compression for live streaming of high definition videos in low-bandwidth transmission Download PDFInfo
- Publication number
- WO2022111140A1 WO2022111140A1 PCT/CN2021/124733 CN2021124733W WO2022111140A1 WO 2022111140 A1 WO2022111140 A1 WO 2022111140A1 CN 2021124733 W CN2021124733 W CN 2021124733W WO 2022111140 A1 WO2022111140 A1 WO 2022111140A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- data
- salient
- salient data
- computer
- implemented method
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/20—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using video object coding
- H04N19/23—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using video object coding with coding of regions that are present throughout a whole video segment, e.g. sprites, background or mosaic
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0475—Generative networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/09—Supervised learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/094—Adversarial learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/40—Extraction of image or video features
- G06V10/46—Descriptors for shape, contour or point-related descriptors, e.g. scale invariant feature transform [SIFT] or bags of words [BoW]; Salient regional features
- G06V10/462—Salient features, e.g. scale invariant feature transforms [SIFT]
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/30—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using hierarchical techniques, e.g. scalability
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/537—Motion estimation other than block-based
- H04N19/54—Motion estimation other than block-based using feature points or meshes
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/597—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding specially adapted for multi-view video sequence encoding
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/234—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
- H04N21/23418—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving operations for analysing video streams, e.g. detecting features or characteristics
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/234—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
- H04N21/2343—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving reformatting operations of video signals for distribution or compliance with end-user requests or end-user device requirements
- H04N21/234318—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving reformatting operations of video signals for distribution or compliance with end-user requests or end-user device requirements by decomposing into objects, e.g. MPEG-4 objects
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/234—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
- H04N21/2343—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving reformatting operations of video signals for distribution or compliance with end-user requests or end-user device requirements
- H04N21/234327—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving reformatting operations of video signals for distribution or compliance with end-user requests or end-user device requirements by decomposing into layers, e.g. base layer and one or more enhancement layers
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/25—Management operations performed by the server for facilitating the content distribution or administrating data related to end-users or client devices, e.g. end-user or client device authentication, learning user preferences for recommending movies
- H04N21/251—Learning process for intelligent management, e.g. learning user preferences for recommending movies
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/47—End-user applications
- H04N21/472—End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content
- H04N21/4728—End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content for selecting a Region Of Interest [ROI], e.g. for requesting a higher resolution version of a selected region
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/60—Network structure or processes for video distribution between server and client or between remote clients; Control signalling between clients, server and network components; Transmission of management data between server and client, e.g. sending from server to client commands for recording incoming content stream; Communication details between server and client
- H04N21/63—Control signaling related to video distribution between client, server and network components; Network processes for video distribution between server and clients or between remote clients, e.g. transmitting basic layer and enhancement layers over different transmission paths, setting up a peer-to-peer communication via Internet between remote STB's; Communication protocols; Addressing
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/60—Network structure or processes for video distribution between server and client or between remote clients; Control signalling between clients, server and network components; Transmission of management data between server and client, e.g. sending from server to client commands for recording incoming content stream; Communication details between server and client
- H04N21/65—Transmission of management data between client and server
- H04N21/658—Transmission by the client directed to the server
- H04N21/6587—Control parameters, e.g. trick play commands, viewpoint selection
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/81—Monomedia components thereof
- H04N21/816—Monomedia components thereof involving special video data, e.g 3D video
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/088—Non-supervised learning, e.g. competitive learning
Definitions
- the present disclosure generally relates to video compression, and more particularly, to the techniques for video enhancement in low-bandwidth transmission applications.
- video coding/decoding techniques such as H. 264 can effectively compress video size for videos having a significant amount of temporal redundancy.
- video coding/decoding techniques can effectively compress video size for videos having a significant amount of temporal redundancy.
- problems with such types of video coding/decoding techniques with lost information occurring during the compression-decompression processes that reduces the quality of the video.
- Another issue with such types of video coding/decoding techniques is their computational complexity. Powerful hardware is involved to implement such video coding/decoding techniques, which poses a problem regarding implementation on devices such as mobile phones.
- Some attempts at addressing the bandwidth cost problem include adaptive video compression of a graphic user interface using the application metadata.
- a structural portion or semantic portion of a video signal is the object of an adaptive coding unit of an identified image region.
- This protocol still requires user end devices capable of performing complex decompression and smoothing in accordance with the analysis of the application metadata.
- Other attempts include data pruning for video compression using example-based super-resolution. Patches of video are extracted from an input video, grouped in a clustering method, and representative patches are packed into patch frames. The original video is downsized and sent along with, or in addition to, patch frames. At the decoding end, regular video frames are upsized and the low-resolution patches are replaced by patches from a patch library. Replacement is only made if there is an appropriate patch available.
- AI artificial intelligence
- down sampling has been carried out in high-definition video at a video source end to obtain a low-definition video.
- the low-definition video is compressed in an existing video coding mode and transmitted, greatly reducing the video traffic.
- the user receives and reconstructs the low-definition video by applying deep learning to a super-resolution image reconstruction method to restore the low-definition video into a high-resolution video at a 50%reduction in video transmission bandwidth cost.
- compression and reconstruction on an entire video is performed without a knowledge of salient and non-salient information.
- a computer-implemented method of encoding video streams for low-bandwidth transmissions includes identifying a salient data and a non-salient data in a high-resolution video stream.
- the salient data and the non-salient data are segmented, and the non-salient data is compressed to a lower resolution.
- the salient data and the compressed non-salient data are transmitted in a low-bandwidth transmission.
- the computer-implemented method advantageously permits the transmission of high-resolution data in a low-bandwidth transmission with a less complicated process that compresses the non-salient data.
- the computer-implemented method further includes encoding the non-salient data prior to performing the compressing of the non-salient data.
- the encoding puts the data in a format suitable for transmission in the low-bandwidth.
- the computer-implemented method further includes the salient data at a lower compression ratio than the non-salient data prior to transmitting the salient data and the compressed non-salient data.
- the salient data is often the data most closely watched, and if not transmitted in its high-resolution form because of bandwidth issues, a compression that is less than the non-salient data can facilitate reconstruction at the receiving end.
- the computer-implemented method further includes identifying at least one of the non-salient data and the salient in the video stream by a machine learning model.
- the use of the machine learning model brings increased efficiency and identifying of salient data and non-salient data using domain knowledge.
- the machine learning model is a General Adversarial Network (GAN)
- the computer-implemented method further includes training the GAN machine learning model to perform identifying the non-salient data with data of non-salient features from previously recorded video streams.
- GAN General Adversarial Network
- the GAN machine learning model is particularly effective in performing accurate identification of the salient and non-salient data.
- the computer-implemented method further includes providing the GAN machine learning model to a user device prior to transmitting the salient data and the compressed non-salient data of the video stream to the user device.
- the user receives access to the GAN model to have an advantage in reconstructing the lower-resolution non-salient data to high-resolution non-salient data and for combining with the salient data to reconstruct the high-resolution video.
- the identifying of the salient data includes identifying domain-specific characteristics of objects in the video stream.
- the characteristics of certain objects can increase the speed and accuracy of identifying salient data.
- the identifying of the salient data includes applying a domain-specific Artificial Intelligence (AI) model for one or more of facial recognition or object recognition.
- AI Artificial Intelligence
- the AI model for facial recognition increases the efficiency and speed of the identification operation of the salient and non-salient data.
- the applying of the domain-specific AI model includes identifying a remainder of the information of the video stream as the non-salient data.
- a plurality of video streams are received having respectively different views of one or more objects, and the identifying and segmenting of the salient data and non-salient data is performed individually for at least two respectively different views that are transmitted.
- the different camera views bring greater flexibility to user views, and performing an individual identifying and segmenting of the video data increases efficiency and the selection of a particular view.
- a computer-implemented method of decoding video data in multiple resolution formats includes receiving a video stream having salient data and non-salient data.
- the salient data is in a higher resolution format than the non-salient data.
- Reconstructing is performed on the non-salient data to increase the resolution format.
- the salient data and the reconstructed non-salient data are recombined to form a video stream in the higher-resolution format of the salient data.
- the decoding permits the received compressed non-salient data to have its resolution increased to be combined with the salient data in a high-resolution video.
- the computer-implemented method further includes receiving one or more of a link to access or executable code for loading a Generative Adversarial Network (GAN) machine learning model trained to identify non-salient features based on previously recorded video streams.
- GAN Generative Adversarial Network
- the non-salient data is reconstructed at an increased resolution using the GAN machine learning model, and the GAN machine model has increased efficiencies at reconstructing the video into a high-definition resolution.
- GAN Generative Adversarial Network
- the received video stream includes salient data and non-salient data captured from multiple viewpoints
- the GAN machine learning model is trained to identify the salient data based on the multiple viewpoints.
- the non-salient data is reconstructed to the higher resolution of the salient data using the GAN machine learning model trained on the multiple viewpoints.
- the computer-implemented method further includes receiving multiple transmissions of the salient data and the non-salient data for each respective viewpoint, reconstructing a particular viewpoint for display in response to a selection.
- the selectability of different viewpoints makes for an increased usefulness of data viewing.
- the computer-implemented method further includes sharing location information with one or more registered users; and receiving selectable views of the salient data and the non-salient data captured by the one or more registered users.
- the users advantageously can share views amongst themselves from different positions in an arena, theater, etc.
- a computing device for encoding video streams for low-bandwidth transmissions includes a processor; a memory coupled to the processor, the memory storing instructions to cause the processor to perform acts including identifying a salient data and a non-salient data in a video stream, and segmenting the video stream into salient data and the non-salient data.
- the non-salient data is encoding and compressed, and the salient data and the compressed non-salient data is transmitted.
- the computer device advantageously permits the transmission of high-resolution data in a low-bandwidth transmission with a less complicated operation to compress the non-salient data. There can be savings in processing power and required bandwidth for transmission.
- the computing device includes a General Adversarial Network (GAN) machine learning model in communication with the memory, and the instructions cause the processor to perform additional acts including training the GAN machine learning model with training data of non-salient features based on previously recorded video streams to perform the identifying of at least the non-salient data.
- GAN General Adversarial Network
- the GAN machine learning model makes for a more efficient operation with reduced processing and power requirements.
- the computing device causes the processor to perform additional acts including the identifying of the salient data includes applying a domain-specific Artificial Intelligence (AI) model for one or more of facial recognition or object recognition.
- AI Artificial Intelligence
- the use of AI in facial or object recognition provides for increased accuracy and efficiency in identifying the salient and non-salient data.
- the computing device includes additional instructions to cause the processor to perform additional acts including transmitting different camera views of the salient data and the non-salient data to respective recipient devices.
- the different camera views increase the effectiveness of any associated user device by providing the different views of an event being captured.
- FIG. 1 provides an architectural overview of a system for encoding video streams for low-bandwidth transmissions, consistent with an illustrative embodiment.
- FIG. 2 illustrates a data segmentation operation of the video of a first sporting event in which salient data is identified, consistent with an illustrative embodiment.
- FIG. 3 illustrates a data segmentation operation of the video of a second sporting event in which salient data is identified, consistent with an illustrative embodiment.
- FIG. 4A illustrates an operation of detecting salient data from multiple viewpoints, consistent with an illustrative embodiment.
- FIG. 4B illustrates the decoding and reconstruction of the multiple viewpoints of salient data detected in FIG. 4A, consistent with an illustrative embodiment.
- FIG. 5 illustrates user end decoding that includes multi-view saliency enhancement, consistent with an illustrative embodiment.
- FIG. 6 is a flowchart illustrating a computer-implemented method of encoding video streams for low-bandwidth transmissions, consistent with an illustrated embodiment.
- FIG. 7 is a flowchart illustrating the use of machine learning models for a computer-implemented method of encoding video streams for high-definition video in a low-bandwidth transmission, consistent with an illustrated embodiment.
- FIG. 8 is a flowchart illustrating operations for decoding and reconstruction, consistent with an illustrative embodiment.
- FIG. 9 is a functional block diagram illustration of a computer hardware platform that can communicate with agents in performing a collaborative task, consistent with an illustrative embodiment.
- FIG. 10 depicts an illustrative cloud computing environment, consistent with an illustrative embodiment.
- FIG. 11 depicts a set of functional abstraction layers provided by a cloud computing environment, consistent with an illustrative embodiment.
- a “Low bandwidth” corresponds to wireless communication at about 2kbps (e.g., 1G) .
- a “High bandwidth range” corresponds to wired/wireless communications up to 1Gbps or higher (e.g., Ethernet or 5G) .
- References herein to video resolutions correspond to QVGA (240X320-pixels) for low resolution, and 4K (3840 x 2160-pixels) for high resolution.
- the computer-implemented method and device of the present disclosure provide for an improvement in the fields of image processing and video transmission, in particular by transmitting the salient data portion of a high definition video data over a low bandwidth transmission without compressing the salient data and without resulting in a loss in quality at a user end.
- By compressing the non-salient data and leaving the salient data for transmission in its high-definition form the efficiency and quality of the video data are increased.
- the system and method of the present disclosure are less complicated, which results in reduced power usage and reduced processing capability required as compared with compressing entire video streams for transmission.
- the video quality will not suffer from a loss in the way a conventional compression of the entire video stream would suffer.
- the domain-specific information (e.g. salient information) is highly relevant for the end user and is therefore kept in its original resolution.
- the other information from the video is compressed, transmitted and reconstructed at the user end.
- the present disclosure is an improvement over methods which extract patches without domain knowledge and uses these patches to up-sample the video.
- FIG. 1 provides an architectural overview 100 of a system for encoding video streams for low-bandwidth transmissions, consistent with an illustrative embodiment.
- FIG. 1 shows a server end 105 that encodes the high definition data for low bandwidth transmission according to the present disclosure.
- the high definition video capture 110 is typically a camera, but if the event were previously recorded, the captured video could be provided by a storage device or a video player.
- data segmentation takes place to segment the data into salient data 120 and non-salient data 125.
- Salient data can include domain-relevant data, such as objects of interest, or user marked regions of interest. Salient data may also be objects in motion. For example, in a soccer match, the players and the ball would at least be considered salient data, whereas the crowd and the arena would be considered non-salient data.
- Non-Salient data is data with little significance, such as, without limitation, static information in video frames, crowd scenes, backgrounds, etc.
- An encoder 130 is configured to perform encoding and compression on the non-salient data.
- the encoded and compressed non-salient data is now low-resolution non-salient data, particularly due to the compression process.
- the salient data in this illustrative embodiment remains in the form of high-resolution salient data.
- the salient data does not suffer from compression losses that can occur when data is compressed, and the perceived quality of the video will remain high, as viewers typically watch the salient data and often do not focus on the background data.
- the reduction is sufficient to transmit the video using low-bandwidth streaming.
- the non-salient data often tends to occupy a large majority of the viewing area (as shown in FIGs.
- the server encoding and compression of the present disclosure provides an efficient way to transmit high definition video by low-bandwidth streaming.
- the server end encoding and compression as discussed herein above does not require the large computational resources that are required of conventional compression of high definition video.
- the user end 155 receives the low-bandwidth transmission 140 and performs decoding and reconstruction.
- the user end device will decode the video stream into the non-salient data in a lower-resolution format and the salient data in the higher resolution format (presuming the salient data was encoded for transmission but not compressed) .
- the non-salient data is reconstructed to the higher resolution format of the salient data.
- the salient data and the reconstructed non-salient data are combined to form a video stream 185 in the higher-resolution format of the salient data that is output.
- Artificial intelligence has a role at the server-end and/or at the user end.
- a machine learning model is trained to identify salient data and non-salient data (e.g., data segmentation) .
- the machine learning model can be trained with previously recorded videos/images of non-salient information. For example, in the event that a soccer match is being streamed, previously recording of the crowd, the arena, the field, etc., can be used to train the machine learning model as to which captured video data is non-salient data, as well as training the machine learning model to identify the salient data.
- One way to detect salient data is by detecting movement. For example, at a soccer match the players, the soccer ball and the referees are usually in motion.
- the salient data corresponds to domain-specific characteristics (e.g., players in a soccer match) , which can be provided to the system through a user interface (e.g., highlights/annotations on the video) , or it can be detected automatically through domain-specific AI-models for facial/object recognition.
- the remaining information in the video is regarded as non-salient or background.
- a machine learning model of a General Adversarial Network (GAN) is trained to detect non-salient features (e.g., a crowd in the arena) .
- GAN General Adversarial Network
- the system when a user registers with the system, the system sends the trained model (GAN) to the user so that, subsequently, the non-salient features can be reconstructed.
- GAN trained model
- Another way the user may access the GAN is through a link, as the user-end 155 may not have the storage space or processing power to receive and operate the GAN.
- multiple cameras 160 can be used for multi-view saliency enhancement 165 in conjunction with a deep learning 170 process.
- the multi-view saliency enhances 165 occur by combining salient information of the multiple viewpoints and training an AI model (deep learning model 170) to improve the image quality of the salient data.
- the data collected from multiple viewpoints from the cameras 180 can also be combined together to train the deep learning model 170 to improve the reconstruction of the image from low resolution to high resolution.
- FIG. 2 illustrates a data segmentation operation 200 of the video of a first sporting event in which salient data is identified, consistent with an illustrative embodiment.
- FIG. 2 shows an image 205 of a soccer match with the players 215 circled for ease of understanding.
- the players are the salient data
- the background crowd 225 and the arena are non-salient data.
- the salient data 260 is identified for data segmentation.
- the salient data is extracted for transmission in its high definition format, whereas the background data is subject to encoding and compression.
- the salient and non-salient data is transmitted to one or more user devices via low-bandwidth transmission.
- the non-salient data 225 I the vast majority of the image as compared to the players 215 (the salient data) , so the encoding and compression of the non-salient data will result in a significant reduction of the data size of the image.
- FIG. 3 illustrates a data segmentation operation 300 of the video of a second sporting event in which salient data is identified, consistent with an illustrative embodiment.
- FIG. 3 shows in 305 a tennis match with the two players 315 circled.
- the two players are the salient data extracted for transmission with a change in formatting, whereas the remainder of the image is non-salient data 365.
- the non-salient data is encoded and compression for transmission in a low-transmission bandwidth.
- FIG. 4A illustrates an operation of detecting salient data from multiple viewpoints 400A, consistent with an illustrative embodiment. It is shown that there are three viewpoints, the first viewpoint 405, the second viewpoint 410, and the third viewpoint 415.
- the first viewpoint 405 appears at about a 45-degree angle relative to the second viewpoint 410
- the third viewpoint 415 appears at about a 90-degree angle relative to the second viewpoint 410.
- Respective cameras 406, 407, 408 each captured a viewpoint 405, 410, 415.
- the salient data for each viewpoint is circled. Below in 435, 445, and 450 is the non-salient data that is compressed.
- FIG. 4B illustrates the user end decoding and reconstruction 400B of the multiple viewpoints of salient data detected in FIG. 4A, consistent with an illustrative embodiment.
- the salient points 455 are shown in FIG. 4B. It can be seen that the amount of salient data 455 (six objects in the example of FIG. 4B) is the same as shown in FIG. 4A.
- the multi-view saliency enhancement occurs by training an AI model (e.g., deep learning 460) with the salient data, and it is shown that different views 465 of the object are output.
- FIG. 4B also shows how the deep learning 460 is used for multi-view background reconstruction. Views 470, 475, 480 are input to the deep learning 460, and a resultant image 485 based on the reconstruction is shown.
- AI model e.g., deep learning 460
- FIG. 5 illustrates user end decoding that includes multi-view saliency enhancement, consistent with an illustrative embodiment.
- FIG. 5 is an illustrative embodiment that is configured for multiple camera view transmission and sharing.
- a server 505 There is shown a server 505 and three user end devices 510, 515, and 520. It is to be understood that the number of user end devices 510, 515, 520 can be more or less than shown. The user end devices may be at different locations during the same event.
- Each of the user end devices 510, 515, 520 can communicate with the server 505, as well as each other.
- the user end devices 510, 515, 520 can use WiFi or Bluetooth to communicate with each other, and cellular (4G) to communicate with the server 505.
- WiFi or Bluetooth to communicate with each other
- cellular (4G) to communicate with the server 505.
- the server 505 sends one or more views to each user, and the user end devices share multiple views locally to improve the reconstruction of the video.
- Data collected from multiple camera viewpoints can be combined to train an AI model to improve the reconstruction of a non-salient image from low resolution to high resolution.
- a salient image can have its quality improved by combining salient information from different camera viewpoints and training an AI model to improve the image quality of the salient data.
- the user end device 510, 515, and 520 can discover each other and establish a channel available in a high bandwidth network through negotiation (e. g., WiFi, Bluetooth) .
- the user end devices can display any view, e.g., a user in geographic proximity of one location can choose any camera of the other users end devices and enjoy any desired view.
- the server may dynamically create user groups based on user’s mobility and network bandwidth availability.
- FIGs. 6, 7, and 8 depict flowcharts 600, 700, and 800, illustrating various aspects of a computer-implemented method, consistent with an illustrative embodiment.
- Processes 600, 700, and 800 are each illustrated as a collection of blocks, in a logical order, which represents a sequence of operations that can be implemented in hardware, software, or a combination thereof.
- the blocks represent computer-executable instructions that, when executed by one or more processors, perform the recited operations.
- computer-executable instructions may include routines, programs, objects, components, data structures, and the like that perform functions or implement abstract data types.
- routines programs, objects, components, data structures, and the like that perform functions or implement abstract data types.
- any number of the described blocks can be combined in any order and/or performed in parallel to implement the process.
- FIG. 6 is a flowchart 600 illustrating a computer-implemented method of encoding video streams for low-bandwidth transmissions, consistent with an illustrated embodiment.
- salient data is identified in a high-resolution video stream.
- the high-resolution video includes but is not limited to a sporting event, musical event, etc.
- An AI model can be used to identify the objects in the image that constitute salient data.
- the salient data could be people, places, objects, etc., as discussed with regard to FIGs. 2 and 3.
- the non-salient data is encoded and compressed to a lower resolution than in the captured high-resolution video.
- the salient data is not compressed and may be encoded.
- the salient data could also compressed, but at a lower rate of compression than the non-salient data.
- the compression can affect the image quality, which is why in this illustrative embodiment the non-salient data is compressed while the salient data is not compressed.
- FIG. 1 provides an overview of an example server end process as well.
- FIG. 7 is a flowchart illustrating the use of machine learning models for a computer-implemented method of encoding video streams for high-definition video in a low-bandwidth transmission, consistent with an illustrated embodiment.
- a General Adversarial Network (GAN) machine learning model is trained with data of non-salient features previously recorded to assist in the identification of the non-salient information.
- the non-salient information may include background information, and/or static information.
- a domain-specific machine learning model for one or more of facial recognition or object recognition is applied to the video data to identify salient data.
- the facial recognition can be used, for example, to identify tennis players in a tennis match.
- the object recognition can be the tennis rackets and a tennis ball.
- one of the non-salient data and the salient data in the video stream are identified by operation of a respective machine learning model.
- the non-salient data can be encoded and compressed, and the salient data made ready for transmission.
- FIG. 8 is a flowchart illustrating operations for decoding and reconstruction, consistent with an illustrative embodiment.
- a user device receives a video data stream containing salient data in a high-resolution format, and non-salient data in a low-resolution format.
- the video stream is decoded and decompressed, and the video data is segmented into non-salient data and salient data.
- An AI model such as a GAN model, or deep learning may be used to identify and segment the data.
- the non-salient data is reconstructed to the higher resolution format of the salient data.
- the user end device may use deep learning or a GAN to assist in this process. There may or may not be multiple camera views that can be used by the deep learning model to assist in the reconstruction.
- the salient data and the reconstructed non-salient data are recombined to form a video stream in the higher-resolution format of the salient data.
- the high definition salient video data can be received by the user end using a low bandwidth without being compressed because the non-salient information is encoded and compressed.
- FIG. 9 provides a functional block diagram illustration 900 of a computer hardware platform.
- FIG. 9 illustrates a particularly configured network or host computer platform 900, as may be used to implement the methods shown in FIGs. 6, 7, and 8.
- the computer platform 900 may include a central processing unit (CPU) 904, a hard disk drive (HDD) 906, random access memory (RAM) and/or read-only memory (ROM) 908, a keyboard 910, a mouse 912, a display 914, and a communication interface 916, which are connected to a system bus 902.
- the HDD 906 can include data stores.
- the HDD 906 has capabilities that include storing a program that can execute various processes, such as encoding module 920 for low-bandwidth transmission, as discussed in a manner described herein above, and is configured to manage the overall process.
- the data segmentation module 925 is configured to segment identified salient and non-salient data in high-resolution videos.
- the data segmentation module can include a machine learning model, such as a General Adversarial Network (GAN) machine learning model.
- GAN General Adversarial Network
- the compression module 930 compresses the identified non-salient data for transmission with the salient data.
- the salient data may remain in its high resolution form, and both the salient and non-salient data can be transmitted together to one or more users.
- the compression of the non-salient data reduces the resolution of the non-salient data to a lower resolution. As it is often that there is significantly more non-salient data than salient data, compressing only the non-salient data reduces the size of the video data so that a low bandwidth transmission can occur.
- the salient data may also be compressed by the compression module 930 to the same compression ratio or a lower compression ratio than the non-salient data.
- the machine learning model (MLM) module 935 is configured to identify one or more of salient data and non-salient data. While the present disclosure is applicable to machine learning modules of various types, as discussed herein above, a General Adversarial Network (GAN) machine learning model is used, consistent with an illustrative embodiment.
- the training of the MLM module 935 can be performed with training data 945 of previously recorded scenes in which there is non-salient data similar to a video stream. For example, in the streaming of live sporting events, previous images of crowds at soccer matches, basketball games, tennis matches can be used to train the machine learning model. For example, at a tennis match, the salient data would be at least the two players and their rackets, the tennis ball, and possibly the net.
- the remainder can be non-salient data that can be compressed to a lower resolution for transmission in a low bandwidth transmission. It is to be understood that other types of machine learning, such as deep learning, can also be used to reconstruct received streams of images back to high resolution at a user end.
- the decoding 940 is configured to decode the video stream into the non-salient data in a lower-resolution format and the salient data in the higher-resolution format.
- the reconstruction module 945 is configured to reconstruct the non-salient data to the higher resolution format of the salient data, and to combine the salient data and the reconstructed non-salient data to form a video stream in the higher-resolution format of the salient data.
- Machine learning is used in an illustrative embodiment to reconstruct the non-salient data into the higher resolution of the salient data and combine the reconstructed non-salient data with the salient data.
- multiple transmissions of the salient data and the non-salient data are received for each respective viewpoint.
- the reconstruction module 945 reconstructs a particular viewpoint or viewpoints for display. The construction of a particular viewpoint may be performed in response to a selection. The viewpoints may not be displayed upon reconstruction, and may be stored for future selection.
- functions relating to the low bandwidth transmission of high definition video data may include a cloud. It is to be understood that although this disclosure includes a detailed description of cloud computing as discussed herein below, implementation of the teachings recited herein is not limited to a cloud computing environment. Rather, embodiments of the present disclosure are capable of being implemented in conjunction with any other type of computing environment now known or later developed.
- Cloud computing is a model of service delivery for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with a provider of the service.
- This cloud model may include at least five characteristics, at least three service models, and at least four deployment models.
- On-demand self-service a cloud consumer can unilaterally provision computing capabilities, such as server time and network storage, as needed automatically without requiring human interaction with the service’s provider.
- Resource pooling the provider’s computing resources are pooled to serve multiple consumers using a multi-tenant model, with different physical and virtual resources dynamically assigned and reassigned according to demand. There is a sense of location independence in that the consumer generally has no control or knowledge over the exact location of the provided resources but may be able to specify location at a higher level of abstraction (e.g., country, state, or datacenter) .
- Rapid elasticity capabilities can be rapidly and elastically provisioned, in some cases automatically, to quickly scale out and rapidly released to quickly scale in. To the consumer, the capabilities available for provisioning often appear to be unlimited and can be purchased in any quantity at any time.
- Measured service cloud systems automatically control and optimize resource use by leveraging a metering capability at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts) . Resource usage can be monitored, controlled, and reported, providing transparency for both the provider and consumer of the utilized service.
- level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts) .
- SaaS Software as a Service: the capability provided to the consumer is to use the provider’s applications running on a cloud infrastructure.
- the applications are accessible from various client devices through a thin client interface such as a web browser (e.g., web-based e-mail) .
- a web browser e.g., web-based e-mail
- the consumer does not manage or control the underlying cloud infrastructure including network, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.
- PaaS Platform as a Service
- the consumer does not manage or control the underlying cloud infrastructure including networks, servers, operating systems, or storage, but has control over the deployed applications and possibly application hosting environment configurations.
- IaaS Infrastructure as a Service
- the consumer does not manage or control the underlying cloud infrastructure but has control over operating systems, storage, deployed applications, and possibly limited control of select networking components (e.g., host firewalls) .
- Private cloud the cloud infrastructure is operated solely for an organization. It may be managed by the organization or a third party and may exist on-premises or off-premises.
- Public cloud the cloud infrastructure is made available to the general public or a large industry group and is owned by an organization selling cloud services.
- Hybrid cloud the cloud infrastructure is a composition of two or more clouds (private, community, or public) that remain unique entities but are bound together by standardized or proprietary technology that enables data and application portability (e.g., cloud bursting for load-balancing between clouds) .
- a cloud computing environment is service-oriented with a focus on statelessness, low coupling, modularity, and semantic interoperability.
- An infrastructure that includes a network of interconnected nodes.
- cloud computing environment 1000 includes cloud 1050 having one or more cloud computing nodes 1010 with which local computing devices used by cloud consumers, such as, for example, personal digital assistant (PDA) or cellular telephone 1054A, desktop computer 1054B, laptop computer 1054C, and/or automobile computer system 1054N may communicate.
- Nodes 1010 may communicate with one another. They may be grouped (not shown) physically or virtually, in one or more networks, such as Private, Community, Public, or Hybrid clouds as described hereinabove, or a combination thereof.
- cloud computing environment 1000 to offer infrastructure, platforms, and/or software as services for which a cloud consumer does not need to maintain resources on a local computing device. It is understood that the types of computing devices 1054A-N shown in FIG. 10 are intended to be illustrative only and that computing nodes 1010 and cloud computing environment 1050 can communicate with any type of computerized device over any type of network and/or network addressable connection (e.g., using a web browser) .
- FIG. 11 a set of functional abstraction layers 1100 provided by cloud computing environment 1000 (FIG. 10) is shown. It should be understood in advance that the components, layers, and functions shown in FIG. 11 are intended to be illustrative only and embodiments of the disclosure are not limited thereto. As depicted, the following layers and corresponding functions are provided:
- Hardware and software layer 1160 include hardware and software components.
- hardware components include: mainframes 1161; RISC (Reduced Instruction Set Computer) architecture based servers 1162; servers 1163; blade servers 1164; storage devices 1165; and networks and networking components 1166.
- software components include network application server software 1167 and database software 1168.
- Virtualization layer 1170 provides an abstraction layer from which the following examples of virtual entities may be provided: virtual servers 1171; virtual storage 1172; virtual networks 1173, including virtual private networks; virtual applications and operating systems 1174; and virtual clients 1175.
- management layer 1180 may provide the functions described below.
- Resource provisioning 1181 provides dynamic procurement of computing resources and other resources that are utilized to perform tasks within the cloud computing environment.
- Metering and Pricing 1182 provide cost tracking as resources are utilized within the cloud computing environment, and billing or invoicing for consumption of these resources. In one example, these resources may include application software licenses.
- Security provides identity verification for cloud consumers and tasks, as well as protection for data and other resources.
- User portal 1183 provides access to the cloud computing environment for consumers and system administrators.
- Service level management 1184 provides cloud computing resource allocation and management such that required service levels are met.
- Service Level Agreement (SLA) planning and fulfillment 1185 provide pre-arrangement for, and procurement of, cloud computing resources for which a future requirement is anticipated in accordance with an SLA.
- SLA Service Level Agreement
- Workloads layer 1190 provides examples of functionality for which the cloud computing environment may be utilized. Examples of workloads and functions which may be provided from this layer include: mapping and navigation 1191; software development and lifecycle management 1192; virtual classroom education delivery 1193; data analytics processing 1194; transaction processing 1195; and a data identification and encoding module 1196 configured to identify salient and non-salient data, and to encode high-resolution video for low bandwidth transmission as discussed herein.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Computing Systems (AREA)
- Databases & Information Systems (AREA)
- Software Systems (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Mathematical Physics (AREA)
- Artificial Intelligence (AREA)
- Evolutionary Computation (AREA)
- General Engineering & Computer Science (AREA)
- Data Mining & Analysis (AREA)
- Biomedical Technology (AREA)
- Computational Linguistics (AREA)
- Biophysics (AREA)
- General Health & Medical Sciences (AREA)
- Molecular Biology (AREA)
- Life Sciences & Earth Sciences (AREA)
- Health & Medical Sciences (AREA)
- Human Computer Interaction (AREA)
- Medical Informatics (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
- Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)
- Image Analysis (AREA)
Abstract
Description
Claims (21)
- A computer-implemented method of encoding a video stream for high-definition video in a low-bandwidth transmission, the method comprising:identifying a salient data and a non-salient data in a high-resolution video stream;segmenting the salient data and the non-salient data;compressing the non-salient data to a lower resolution; andtransmitting the salient data and the compressed non-salient data.
- The computer-implemented method of claim 1, further comprising encoding the non-salient data prior to performing the compressing of the non-salient data.
- The computer-implemented method according to claim 1, further comprising compressing the salient data at a lower compression ratio than the non-salient data prior to transmitting the salient data and the compressed non-salient data.
- The computer-implemented method of claim 1, further comprising identifying at least one of the non-salient data and the salient data in the video stream by a machine learning model.
- The computer-implemented method of claim 4, wherein the machine learning model comprises a General Adversarial Network (GAN) machine learning model; and further comprising:training the GAN machine learning model with data of one or more non-salient features from previously recorded video streams to identify the non-salient data.
- The computer-implemented method of claim 5, further comprising providing to a user device one or more of a link to access or a code to execute the GAN machine learning model prior to transmitting the salient data and the compressed non-salient data of the video stream to the user device.
- The computer-implemented method of claim 1, wherein the identifying of the salient data includes identifying one or more domain-specific characteristics of objects in the video stream.
- The computer-implemented method of claim 1, wherein the identifying of the salient data includes applying a domain-specific Artificial Intelligence (AI) model for one or more of facial recognition or object recognition.
- The computer-implemented method of claim 8, wherein the applying of the domain-specific AI model includes further comprises identifying a remainder of information of the video stream as the non-salient data.
- The computer-implemented method of claim 1, further comprising receiving a plurality of video streams, each video stream having respectively different views of one or more objects, wherein the identifying and segmenting of the salient data and non-salient data is performed individually for at least two respectively different views that are transmitted.
- A computer-implemented method of decoding video data in multiple resolution formats, the computer-implemented method comprising:receiving an encoded video stream including a salient data and a non-salient data, the salient data having a higher resolution format than the non-salient data;decoding the video stream into the non-salient data in a lower-resolution format and the salient data in the higher-resolution format;reconstructing the non-salient data to a higher resolution format; andcombining the salient data and the reconstructed non-salient data to form a video stream in the higher-resolution format of the salient data.
- The computer-implemented method of claim 11, further comprising:receiving one or more of a link to access operation of, or load executable code for, a Generative Adversarial Network (GAN) machine learning model trained to identify non-salient features based on previously recorded video streams; andreconstructing the non-salient data at an increased resolution using the GAN machine learning model.
- The computer-implemented method of claim 12, wherein the received video stream includes a salient data and a non-salient data captured from multiple viewpoints; the computer-implemented method further comprising:training the GAN machine learning model to identify the salient data based on the multiple viewpoints; andreconstructing the non-salient data to the higher resolution of the salient data using the GAN machine learning model trained on the multiple viewpoints.
- The computer-implemented method of claim 13, further comprising:receiving multiple transmissions of the salient data and the non-salient data for each viewpoint, respectively; andreconstructing a particular viewpoint for display in response to a selection.
- The computer-implemented method of claim 14, further comprising:sharing location information with one or more registered users; andreceiving selectable views of the salient data and the non-salient data captured by the one or more registered users.
- A computing device for encoding video streams for high-definition video in a low-bandwidth transmission, the computing device comprising:a processor;a memory coupled to the processor, the memory storing instructions to cause the processor to perform acts comprising:identifying a salient data and a non-salient data in a video stream;segmenting the salient data and the non-salient data;encoding and compressing the non-salient data; andtransmitting the salient data and the compressed non-salient data.
- The computing device of claim 16, further comprising:a General Adversarial Network (GAN) machine learning model in communication with the memory; andwherein the instructions cause the processor to perform an additional act comprising training the GAN machine learning model with training data of non-salient features based on previously recorded video streams to perform the identifying of at least the non-salient data.
- The computing device of claim 17, wherein the instructions cause the processor to perform additional acts comprising:receiving refined results from the selected agents based on the fused parameters; andgenerating a global training model based on the refined results.
- The computing device of claim 16, wherein the instructions cause the processor to perform an additional act comprising:applying a domain-specific Artificial Intelligence (AI) model comprising one or more of facial recognition or object recognition to identify the salient data.
- The computing device of claim 16, wherein the instructions cause the processor to perform an additional act comprising transmitting different camera views of the salient data and the non-salient data to a plurality of recipient devices.
- A computer program comprising program code adapted to perform the method steps of any of claims 1 to 15 when said program is run on a computer.
Priority Applications (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202180076756.2A CN116457819A (en) | 2020-11-25 | 2021-10-19 | Non-parallel compressed video coding of high definition video real-time streams in low bandwidth transmission |
| GB2309315.6A GB2616998B (en) | 2020-11-25 | 2021-10-19 | Video encoding through non-saliency compression for live streaming of high definition videos in low-bandwidth transmission |
| DE112021006157.7T DE112021006157T5 (en) | 2020-11-25 | 2021-10-19 | VIDEO CODING USING NON-SALIENCY COMPRESSION FOR LIVE STREAMING OF HIGH-RESOLUTION VIDEO IN A LOW-BANDWIDTH TRANSMISSION |
| JP2023530212A JP7706861B2 (en) | 2020-11-25 | 2021-10-19 | Video encoding with non-saliency compression for live streaming of high definition video in low bandwidth transmissions - Patents.com |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US17/105,356 | 2020-11-25 | ||
| US17/105,356 US11758182B2 (en) | 2020-11-25 | 2020-11-25 | Video encoding through non-saliency compression for live streaming of high definition videos in low-bandwidth transmission |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2022111140A1 true WO2022111140A1 (en) | 2022-06-02 |
Family
ID=81657630
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2021/124733 Ceased WO2022111140A1 (en) | 2020-11-25 | 2021-10-19 | Video encoding through non-saliency compression for live streaming of high definition videos in low-bandwidth transmission |
Country Status (6)
| Country | Link |
|---|---|
| US (2) | US11758182B2 (en) |
| JP (1) | JP7706861B2 (en) |
| CN (1) | CN116457819A (en) |
| DE (1) | DE112021006157T5 (en) |
| GB (1) | GB2616998B (en) |
| WO (1) | WO2022111140A1 (en) |
Families Citing this family (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US12058312B2 (en) * | 2021-10-06 | 2024-08-06 | Kwai Inc. | Generative adversarial network for video compression |
| EP4418659A4 (en) | 2021-12-21 | 2025-03-05 | Samsung Electronics Co., Ltd. | Ai-based image providing apparatus and method therefor, and ai-based display apparatus and method therefor |
| US11895344B1 (en) * | 2022-12-09 | 2024-02-06 | International Business Machines Corporation | Distribution of media content enhancement with generative adversarial network migration |
| CN116781912B (en) * | 2023-08-17 | 2023-11-14 | 瀚博半导体(上海)有限公司 | Video transmission method, device, computer equipment and computer readable storage medium |
| WO2026009040A1 (en) * | 2024-06-30 | 2026-01-08 | Four Drobotics Corporation | System and method for real-time artificial intelligence-based video compression and decompression |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20050013479A1 (en) * | 2003-07-16 | 2005-01-20 | Rong Xiao | Robust multi-view face detection methods and apparatuses |
| WO2014008541A1 (en) * | 2012-07-09 | 2014-01-16 | Smart Services Crc Pty Limited | Video processing method and system |
| CN107194927A (en) * | 2017-06-13 | 2017-09-22 | 天津大学 | The measuring method of stereo-picture comfort level chromaticity range based on salient region |
| US20200074589A1 (en) * | 2018-09-05 | 2020-03-05 | Toyota Research Institute, Inc. | Systems and methods for saliency-based sampling layer for neural networks |
Family Cites Families (25)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP4127204B2 (en) | 2003-12-17 | 2008-07-30 | セイコーエプソン株式会社 | Manufacturing method of liquid crystal display device |
| US7548657B2 (en) | 2005-06-25 | 2009-06-16 | General Electric Company | Adaptive video compression of graphical user interfaces using application metadata |
| EP1936566A1 (en) | 2006-12-22 | 2008-06-25 | Thomson Licensing | Method for creating the saliency map of an image and system for creating reduced pictures of video frames |
| US8339456B2 (en) * | 2008-05-15 | 2012-12-25 | Sri International | Apparatus for intelligent and autonomous video content generation and streaming |
| JP5156982B2 (en) * | 2008-06-05 | 2013-03-06 | 富士フイルム株式会社 | Image processing system, image processing method, and program |
| US8605795B2 (en) * | 2008-09-17 | 2013-12-10 | Intel Corporation | Video editing methods and systems |
| US8750645B2 (en) * | 2009-12-10 | 2014-06-10 | Microsoft Corporation | Generating a composite image from video frames |
| KR101791919B1 (en) | 2010-01-22 | 2017-11-02 | 톰슨 라이센싱 | Data pruning for video compression using example-based super-resolution |
| US20120300048A1 (en) * | 2011-05-25 | 2012-11-29 | ISC8 Inc. | Imaging Device and Method for Video Data Transmission Optimization |
| US9275300B2 (en) * | 2012-02-24 | 2016-03-01 | Canon Kabushiki Kaisha | Method and apparatus for generating image description vector, image detection method and apparatus |
| US8977582B2 (en) * | 2012-07-12 | 2015-03-10 | Brain Corporation | Spiking neuron network sensory processing apparatus and methods |
| US20140177706A1 (en) * | 2012-12-21 | 2014-06-26 | Samsung Electronics Co., Ltd | Method and system for providing super-resolution of quantized images and video |
| US9324161B2 (en) * | 2013-03-13 | 2016-04-26 | Disney Enterprises, Inc. | Content-aware image compression method |
| US9440152B2 (en) * | 2013-05-22 | 2016-09-13 | Clip Engine LLC | Fantasy sports integration with video content |
| US9807411B2 (en) * | 2014-03-18 | 2017-10-31 | Panasonic Intellectual Property Management Co., Ltd. | Image coding apparatus, image decoding apparatus, image processing system, image coding method, and image decoding method |
| GB201603144D0 (en) | 2016-02-23 | 2016-04-06 | Magic Pony Technology Ltd | Training end-to-end video processes |
| US9936208B1 (en) * | 2015-06-23 | 2018-04-03 | Amazon Technologies, Inc. | Adaptive power and quality control for video encoders on mobile devices |
| WO2017055609A1 (en) | 2015-09-30 | 2017-04-06 | Piksel, Inc | Improved video stream delivery via adaptive quality enhancement using error correction models |
| CN105959705B (en) | 2016-05-10 | 2018-11-13 | 武汉大学 | A kind of net cast method towards wearable device |
| CN106791927A (en) | 2016-12-23 | 2017-05-31 | 福建帝视信息科技有限公司 | A kind of video source modeling and transmission method based on deep learning |
| US10163227B1 (en) * | 2016-12-28 | 2018-12-25 | Shutterstock, Inc. | Image file compression using dummy data for non-salient portions of images |
| CN107423740A (en) * | 2017-05-12 | 2017-12-01 | 西安万像电子科技有限公司 | The acquisition methods and device of salient region of image |
| US10176405B1 (en) * | 2018-06-18 | 2019-01-08 | Inception Institute Of Artificial Intelligence | Vehicle re-identification techniques using neural networks for image analysis, viewpoint-aware pattern recognition, and generation of multi- view vehicle representations |
| US20210006730A1 (en) * | 2019-07-07 | 2021-01-07 | Tangible Play, Inc. | Computing device |
| EP4136848A4 (en) * | 2020-04-16 | 2024-04-03 | INTEL Corporation | Patch based video coding for machines |
-
2020
- 2020-11-25 US US17/105,356 patent/US11758182B2/en active Active
-
2021
- 2021-10-19 GB GB2309315.6A patent/GB2616998B/en active Active
- 2021-10-19 CN CN202180076756.2A patent/CN116457819A/en active Pending
- 2021-10-19 WO PCT/CN2021/124733 patent/WO2022111140A1/en not_active Ceased
- 2021-10-19 DE DE112021006157.7T patent/DE112021006157T5/en active Pending
- 2021-10-19 JP JP2023530212A patent/JP7706861B2/en active Active
-
2023
- 2023-07-30 US US18/361,887 patent/US12126828B2/en active Active
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20050013479A1 (en) * | 2003-07-16 | 2005-01-20 | Rong Xiao | Robust multi-view face detection methods and apparatuses |
| WO2014008541A1 (en) * | 2012-07-09 | 2014-01-16 | Smart Services Crc Pty Limited | Video processing method and system |
| CN107194927A (en) * | 2017-06-13 | 2017-09-22 | 天津大学 | The measuring method of stereo-picture comfort level chromaticity range based on salient region |
| US20200074589A1 (en) * | 2018-09-05 | 2020-03-05 | Toyota Research Institute, Inc. | Systems and methods for saliency-based sampling layer for neural networks |
Also Published As
| Publication number | Publication date |
|---|---|
| US12126828B2 (en) | 2024-10-22 |
| GB2616998B (en) | 2025-04-02 |
| US11758182B2 (en) | 2023-09-12 |
| CN116457819A (en) | 2023-07-18 |
| US20240022759A1 (en) | 2024-01-18 |
| US20220167005A1 (en) | 2022-05-26 |
| JP2023551158A (en) | 2023-12-07 |
| JP7706861B2 (en) | 2025-07-14 |
| GB2616998A (en) | 2023-09-27 |
| DE112021006157T5 (en) | 2023-10-12 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12126828B2 (en) | Video encoding through non-saliency compression for live streaming of high definition videos in low-bandwidth transmission | |
| US11122332B2 (en) | Selective video watching by analyzing user behavior and video content | |
| US11172225B2 (en) | Aerial videos compression | |
| US20240371084A1 (en) | System and method for dynamic images virtualisation | |
| Noor et al. | Orchestrating image retrieval and storage over a cloud system | |
| EP4289137B1 (en) | Video inter/intra compression using mixture of experts | |
| Satheesh Kumar et al. | A novel video compression model based on GPU virtualization with CUDA platform using bi-directional RNN | |
| US12561829B2 (en) | Method, computer device, and computer program for providing high-quality image of region of interest by using single stream | |
| US10666954B2 (en) | Audio and video multimedia modification and presentation | |
| US10609368B2 (en) | Multiple image storage compression tree | |
| US20240144425A1 (en) | Image compression augmented with a learning-based super resolution model | |
| US20240185388A1 (en) | Method, electronic device, and computer program product for image processing | |
| WO2023226504A1 (en) | Media data processing methods and apparatuses, device, and readable storage medium | |
| CN116527839A (en) | Digital conference interaction method, device, system and related device | |
| CN117242421A (en) | Smart client for streaming of scene-based immersive media | |
| US20220060750A1 (en) | Freeview video coding | |
| US20250004991A1 (en) | Block-level, bit-mapped binary data access for parallel processing | |
| Seligmann | Web-based client for remote rendered virtual reality | |
| Amezcua Aragon | Real-time neural network based video super-resolution as a service: Design and implementation of a real-time video super-resolution service using public cloud services | |
| HK40098413A (en) | Image encoding and decoding method, apparatus, computer, storage medium, and program product | |
| CN116980589A (en) | Image coding and decoding methods, devices, computers, storage media and program products | |
| CN116416483A (en) | Computer-implemented method, apparatus and computer program product |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 21896642 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 202180076756.2 Country of ref document: CN |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2023530212 Country of ref document: JP |
|
| ENP | Entry into the national phase |
Ref document number: 202309315 Country of ref document: GB Kind code of ref document: A Free format text: PCT FILING DATE = 20211019 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2309315.6 Country of ref document: GB |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 112021006157 Country of ref document: DE |
|
| WWP | Wipo information: published in national office |
Ref document number: 2309315.6 Country of ref document: GB |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 21896642 Country of ref document: EP Kind code of ref document: A1 |
|
| WWG | Wipo information: grant in national office |
Ref document number: 2309315.6 Country of ref document: GB |